Evolutionary dual-task feature selection method based on hybrid initialization particle swarm optimization
Through the evolutionary dual-task feature selection method optimized by mixing initialized particle swarm, the feature selection problem of high-dimensional small sample data is solved, and efficient and low-cost feature subset screening and classification performance improvement is achieved, which is suitable for high-dimensional data sets.
Patent Information
- Application Number
- CN202310900112.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-21
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2043-07-21
AI Technical Summary
When handling feature selection of high-dimensional small sample data, the prior art has the problem of "dimensional disaster", which has low classification performance, high calculation cost, and the feature selection process is prone to local optimization, and cannot effectively utilize feature correlation and redundancy.
The evolutionary dual-task feature selection method based on hybrid initialization particle swarm optimization is adopted. Through feature probability initialization, inflection point strategy division tasks, and global optimal location knowledge sharing, subsets of feature with high correlation are selected and redundant features are reduced, and knowledge transfer is used to accelerate convergence and reduce computational costs.
Achieve high classification performance with a smaller subset of features in a short time, reducing calculation costs, and is suitable for any high-dimensional data set. The feature selection process is rigorous, reducing waste of computing resources, and improving classification accuracy.
Smart Images

Figure CN117035000B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning, and in particular relates to an evolutionary dual-task feature selection method based on hybrid-initialized particle swarm optimization. Background Art
[0002] Data classification is a major research topic in data mining. It involves training existing data to construct classification models and extract available knowledge to better describe or predict new test data. The demand for data mining in the medical field is increasing rapidly. Simultaneously, the development of bioinformatics has led to a dramatic increase in the dimensionality of medical data, resulting in high-dimensional data. High-dimensional, small-sample data refers to data that exhibits both high dimensionality and a relatively small number of samples containing labeled information (e.g., categories). This type of data is prevalent in the biomedical field. Classifying high-dimensional, small-sample data is subject to a severe "curse of dimensionality." Noisy features in high-dimensional feature data can negatively impact the final classification results. Consequently, existing classification techniques suffer from reduced classification performance due to the large number of irrelevant or redundant features. Furthermore, the high feature dimensionality reduces classification efficiency. Removing these irrelevant and noisy features can improve classification accuracy.
[0003] Currently, there are three main approaches to feature selection: filtering, wrapper, and embedded. Filtering methods typically select feature subsets by calculating intrinsic information between features (such as distance, correlation, and information gain). Because these methods typically manually set thresholds to select the appropriate features, they often miss some important features. While faster, they often suffer from suboptimal classification accuracy. Wrapper methods primarily evaluate the quality of a selected feature subset using specific classifiers, such as decision trees, support vector machines, K-nearest neighbors, or artificial neural networks. However, these methods involve feature subset evaluation and, due to the greater number of features in high-dimensional datasets, require significant computational overhead. Embedded feature selection methods combine feature selection with the training of a learning model to determine the importance of each feature. While this approach can save some computational overhead, its classification results are often influenced by the classifier.
[0004] Evolutionary algorithms are widely used in feature selection due to their powerful search capabilities. In the paper "Evolutionary Machine Learning with Minions: A Case Study in Feature Selection," an algorithm-centric feature selection method based on evolutionary multi-tasks is proposed. First, a set of small-data agents is created by subsampling a small portion of a large dataset. Then, the secondary tasks are combined with the primary task in a single multi-task optimization framework, facilitating evolutionary search by rapidly optimizing the large dataset using small data. However, the following technical issues exist: 1. Because the number of auxiliary tasks to construct is uncertain, extensive experimentation is required before experimentation to determine the optimal number of agents; 2. The large number of tasks involved incurs significant computational overhead; 3. The generated auxiliary tasks are random and fail to consider features that are highly relevant to the label, resulting in low classification accuracy; and a significant amount of resources are consumed in the search process.
[0005] In summary, continuously optimizing the evolutionary dual-task feature selection method based on hybrid initialization particle swarm optimization has become a key research direction for researchers in this field. Summary of the Invention
[0006] The present invention aims to provide an evolutionary dual-task feature selection method based on hybrid initialization particle swarm optimization to overcome the shortcomings of the existing technology, namely, the complexity of the high-dimensional feature space, the susceptibility of the feature subset search process to fall into local optimality, and the high computational cost of the feature subset evaluation process.
[0007] In order to achieve the above object, the present invention provides a technical solution: an evolutionary dual-task feature selection method based on hybrid initialization particle swarm optimization, comprising the following steps:
[0008] Step 1: Hybrid initialization based on feature probability: By comparing the probability correlation of each feature with the size of the random number, we decide whether the feature should be added to the initial population. The proposed initialization strategy is used to initialize half of the population, and the other half is initialized using a random strategy to ensure the activity of the population.
[0009] Step 2: Task division strategy based on feature importance: The frontier feature set and the remaining feature set are distinguished by inflection points. The search for the frontier feature subset is considered the main task. The redundant features in the remaining feature subset are then removed and merged with the frontier feature set to form a non-redundant feature set. The search for the non-redundant feature subset is considered an auxiliary task.
[0010] Step 3. Dual-task knowledge sharing strategy based on global optimal position: First, based on the task division strategy in step 2, the original dataset is divided into a promising feature subset and a non-redundant feature subset. The search process for these two feature sets is regarded as the main task and auxiliary task, respectively. The main task is used to guide the search to promising areas, and the auxiliary task helps the algorithm escape from the local optimum by reducing the possibility of the algorithm falling into the local optimum; then a fitness evaluation operation is performed: the linear relationship between the classification accuracy and the number of selected features is used to represent the fitness function, and finally the feature subset selected by the main task is used as the final returned feature subset.
[0011] Furthermore, the above step 1 specifically includes:
[0012] Step 101: Calculate the maximum information coefficient (MIC) value of the mutual information of each feature based on the label, recorded as M_relevance;
[0013] Step 102: Calculate M_prob of each feature through M_relevance; where for feature x i , its M_prob is recorded as p i The calculation process is:
[0014] Step 103: Compare M_prob with the random number. If the M_prob of the feature is greater than the random number, the feature is selected; otherwise, the feature is not selected into the initial population, and half of the population is initialized using this method.
[0015] Step 104: Initialize the remaining half of the population using a random initialization strategy;
[0016] MIC is based on the mutual information calculation feature f k and the correlation of category c, by adding feature f k and category c are divided into m and n different intervals to obtain an m*n grid G; under the specified grid G, the empirical joint probability density and the empirical marginal probability density are calculated by the proportion of the number of samples in each grid and the number of samples in the interval in the sample capacity, and then the mutual information is estimated, and then the value of the mutual information is normalized in the interval [0,1]; the above steps are repeated using multiple different grids; then the maximum mutual information values obtained from different grids are compared, and the maximum value is selected as the minimum interference ratio.
[0017] Furthermore, the above step 2 specifically includes:
[0018] 201. According to the inflection point selection strategy, find the inflection point:
[0019] 1: Calculate the relevance of features on tags based on MIC;
[0020] 2: Arrange the features in descending order according to the value of M_relevance to form the feature importance curve S;
[0021] 3: Connect the two ends of the curve into a line, record it as line L, and calculate the straight-line distance d from each point on the curve S to line L; when the distance d is the largest, the corresponding point on the curve S is the inflection point;
[0022] 202 takes the inflection point as the dividing point and selects the features whose M_relevance is greater than this point as the frontier feature subset. The original dataset is divided into the frontier feature subset D1 and the remaining feature subset D2; where D1 is the search space of task 1;
[0023] 203 Use mRMR (max-Relevance and Min-Redundancy) to remove redundant features in the remaining feature set D2;
[0024] 204: Combine the frontier feature subset and the remaining feature set D3 after excluding redundant features into a new feature set D4, which becomes a non-redundant feature set; D4 is the search space of task 2;
[0025] MIC is based on the mutual information calculation feature f k and the correlation of category c, by adding feature f k and category c are divided into m and n different intervals to obtain an m*n grid G; under the specified grid G, the empirical joint probability density and the empirical marginal probability density are calculated by the proportion of the number of samples in each grid and the number of samples in the interval in the sample capacity, and then the mutual information is estimated, and then the value of the mutual information is normalized in the interval [0,1]; the above steps are repeated using multiple different grids; then the maximum mutual information values obtained from different grids are compared, and the maximum value is selected as the minimum interference ratio.
[0026] Furthermore, the above step three specifically includes:
[0027] 301. Use skill factors to assign particles to each task;
[0028] 302. If the value of the skill factor is 1, the particle is used to solve task 1; otherwise, it is used to solve task 2;
[0029] 303. Define rmp to determine when the transfer process can be executed;
[0030] 304. If rand>rmp, activate the knowledge transfer operation and use the position information of the current task to update the position of the particle; otherwise, use the position information from other tasks to update the position of the particle;
[0031] 305. Use the fitness function to evaluate the selected feature subset.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] 1. The present invention searches for the optimal solution by transferring knowledge between two related tasks, thereby realizing knowledge transfer between related tasks, thereby accelerating the population convergence speed, achieving higher classification performance with a smaller feature subset in a relatively short time, and reducing expensive computational costs.
[0034] 2. The inflection point selection strategy proposed in the present invention is different from the existing feature selection algorithms. The existing feature selection methods usually manually set thresholds in the initial stage to remove some irrelevant or less relevant feature subsets. Such methods usually lose some important features and require relevant prior knowledge. The present invention automatically determines the inflection point based on the characteristics of the data set itself, thereby screening out those features with greater correlation, making the feature selection process more rigorous, and this method is applicable to any high-dimensional data set.
[0035] 3. The present invention determines two feature subsets based on feature correlation and redundancy, which are used as search spaces for the primary and auxiliary tasks, respectively. The feature space of the primary task is a frontier feature set generated by selecting an inflection point strategy, consisting of features with high correlation. Since features often have redundancy, this redundancy is taken into account when determining the search space for the auxiliary task. The search space for the auxiliary task is formed by removing features that are highly redundant with the frontier feature subset and then merging them with the frontier.
[0036] 4. In the knowledge sharing strategy of the present invention, since the main focus is on the frontier feature subset area of the main task, the auxiliary task is used to help the main task escape the local optimum, and the feature subsets involved in the two tasks are much smaller than the number of features in the original feature subsets, a large amount of computational cost is reduced during fitness evaluation.
[0037] 5. The present invention adopts an initialization strategy based on feature probability when initializing the population. By comparing the probability of feature correlation and the size of the random number, the features are selected to be used for initialization. Therefore, those features with high correlation with the label are selected into the initial population at the beginning of evolution, which is convenient for evolving high-quality populations in the subsequent evolution process. In order to further enhance the activity of the population, the present invention only uses this strategy to initialize half of the population, and the remaining population is still initialized by the random initialization strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is the overall flow chart of the present invention;
[0039] Figure 2 Illustration of the initialized population based on feature probability;
[0040] Figure 3 It is the generation process of two related tasks. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0042] The present invention uses the SRBCT dataset as an example, specifically:
[0043] The SRBCT dataset is used to train and evaluate the performance of the proposed algorithm. The SRBCT dataset is a public dataset obtained from Kaggle. The dataset has 2308 features and 83 samples, which conforms to the characteristics of high-dimensional small sample data. It is a 4-category dataset.
[0044] Data preprocessing operations include:
[0045] Step 1: Check whether the dataset has missing values. If there are missing values, use replacement or filling to handle the missing values.
[0046] Step 2: Check the dataset for outliers.
[0047] For the above prepared dataset, the present invention provides an evolutionary dual-task feature selection method based on hybrid initialization particle swarm optimization, see Figure 1 The overall process of the present invention comprises the following steps:
[0048] Step 1: Hybrid initialization based on feature probability: By comparing the probability correlation of each feature with the size of the random number, the feature is determined whether to be added to the initial population. The proposed initialization strategy is used to initialize half of the population, and the other half is initialized using a random strategy to ensure the activity of the population. Specifically, it includes:
[0049] Step 101: Calculate the maximum information coefficient (MIC) value of the mutual information of each feature based on the label, recorded as M_relevance;
[0050] Step 102: Calculate M_prob of each feature through M_relevance;
[0051] Step 103: Compare M_prob with the random number. If the M_prob of the feature is greater than the random number, the feature is selected; otherwise, the feature is not selected into the initial population. This method is used to initialize half of the population.
[0052] See the above steps for Figure 2 Probability-based initialization strategy.
[0053] Step 104: Initialize the remaining half of the population using a random initialization strategy.
[0054] Step 2: Task division based on feature importance: A frontier feature set is distinguished from a remaining feature set by using inflection points. Searching for the frontier feature subset is considered the primary task. Redundant features in the remaining feature subset are then removed and merged with the frontier feature set to form a non-redundant feature set. Searching for the non-redundant feature subset is considered an auxiliary task.
[0055] See also Figure 3 , specifically including:
[0056] 201. According to the inflection point selection strategy, find the inflection point:
[0057] 1: Calculate the relevance of features on tags based on MIC;
[0058] 2: Arrange the features in descending order according to the value of M_relevance to form the feature importance curve S;
[0059] 3: Connect the two ends of the curve into a line, called line L, and calculate the straight-line distance d from each point on curve S to line L. When distance d is maximum, the corresponding point on curve S is the inflection point.
[0060] 202 takes the inflection point as the dividing point and selects features with M_relevance greater than this point as the frontier feature subset. The original dataset is divided into the frontier feature subset D1 and the remaining feature subset D2; where D1 is the search space of task 1.
[0061] 203 mRMR (max-Relevance and Min-Redundancy) is used to remove redundant features in the remaining feature set D2.
[0062] 204: Combine the frontier feature subset and the remaining feature set D3 after excluding redundant features into a new feature set D4, which becomes a non-redundant feature set. D4 is the search space of task 2.
[0063] MIC is based on the mutual information calculation feature f k and the correlation of category c, by adding feature f kThe sample and category c are divided into m and n intervals, resulting in an m x n grid G. Within a given grid G, the empirical joint probability density and the empirical marginal probability density are calculated based on the number of samples in each grid and the proportion of the number of samples in the interval to the sample capacity, respectively. Mutual information is then estimated and normalized to the interval [0, 1]. Repeat the above steps using multiple different grids. The maximum mutual information values obtained from different grids are then compared and the maximum value is selected as the minimum interference ratio.
[0064] Step 3. Dual-task knowledge sharing strategy based on global optimal position: First, based on the task division strategy of step 2, the original data set is divided into promising feature subsets and non-redundant feature subsets. The search process for these two feature sets is regarded as the main task and auxiliary task respectively. The main task is used to guide the search to promising areas, and the auxiliary task helps the algorithm escape the local optimum by reducing the possibility of the algorithm falling into the local optimum. Then, the fitness evaluation operation is performed: the linear relationship between the classification accuracy and the number of selected features is used to represent the fitness function. Finally, the feature subset selected by the main task is used as the final feature subset returned. Specifically, the following steps are included:
[0065] Specifically include:
[0066] 301. Use skill factors to assign particles to each task.
[0067] 302. If the value of the skill factor is 1, the particle is used to solve task 1. Otherwise, it is used to solve task 2.
[0068] 303. Define rmp to determine when the transfer process can be executed.
[0069] 304. If rand>rmp, activate the knowledge transfer operation and use the position information of the current task to update the position of the particle; otherwise, use the position information from other tasks to update the position of the particle.
[0070] 305. Use the fitness function to evaluate the selected feature subset.
[0071] See Table 1 for basic information of the 12 high-dimensional small sample data sets used in this invention.
[0072]
[0073]
[0074] See Table 2 and Table 3 for experimental results on classification accuracy of the method of the present invention and six intelligent optimization algorithms;
[0075] Table 2. Experimental results on classification accuracy with three intelligent optimization algorithms
[0076]
[0077] Table 3. Experimental results on classification accuracy with three intelligent optimization algorithms
[0078]
[0079] See Table 4 and Table 5, experimental results of the classification accuracy of the method of the present invention and 5 hybrid particle swarm optimization algorithms; Table 4, experimental results of the classification accuracy of the method of the present invention and 3 hybrid particle swarm optimization algorithms
[0080]
[0081] Table 5. Experimental results on classification accuracy with two hybrid particle swarm optimization algorithms
[0082]
[0083] From Table 2 to Table 5, it can be seen that the present invention shows good classification effect on 12 high-dimensional small sample data sets.
[0084] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. An evolutionary dual-task feature selection method based on hybrid-initialized particle swarm optimization, characterized by: The following steps are involved: Step 1: Hybrid initialization based on feature probability: By comparing the probability correlation of each feature with the size of the random number, we decide whether to add the feature to the initial population. The proposed initialization strategy is used to initialize half of the population, and the other half is initialized using a random strategy to ensure the activity of the population. Step 2: Task division strategy based on feature importance: The frontier feature set and the remaining feature set are distinguished by inflection points. The search for the frontier feature subset is considered the main task. The redundant features in the remaining feature subset are then removed and merged with the frontier feature set to form a non-redundant feature set. The search for the non-redundant feature subset is considered an auxiliary task. Step 3: Dual-task knowledge sharing strategy based on global optimal position: First, based on the task division strategy in Step 2, the original dataset is divided into a promising feature subset and a non-redundant feature subset. The search process for these two feature sets is treated as the main task and auxiliary task, respectively. The main task is used to guide the search to promising areas, and the auxiliary task helps the algorithm escape the local optimum by reducing the possibility of falling into the local optimum. Then, a fitness evaluation operation is performed: the fitness function is expressed as a linear relationship between classification accuracy and the number of selected features. Finally, the feature subset selected by the main task is used as the final feature subset returned. The original dataset is the SRBCT dataset.
2. The evolutionary dual-task feature selection method based on hybrid initialization particle swarm optimization according to claim 1, characterized in that: The step 1 specifically includes: Step 101: Calculate the maximum information coefficient MIC value of the mutual information of each feature based on the label, recorded as M_relevance; Step 102: Calculate M_prob of each feature through M_relevance; , its M_prob is recorded as The calculation process is: Step 103: Compare M_prob with the random number. If the M_prob of the feature is greater than the random number, the feature is selected; otherwise, the feature is not selected into the initial population. This method is used to initialize half of the population. Step 104: Initialize the remaining half of the population using a random initialization strategy.
3. The evolutionary dual-task feature selection method based on hybrid initialization particle swarm optimization according to claim 2, characterized in that: The second step specifically includes:
201. According to the inflection point selection strategy, find the inflection point: 1: Calculate the relevance of features on tags based on MIC; 2: Arrange the features in descending order according to the value of M_relevance to form the feature importance curve S; 3: Connect the two ends of the curve into a line, record it as line L, and calculate the straight-line distance d from each point on the curve S to line L; when the distance d is the largest, the corresponding point on the curve S is the inflection point; 202 Taking the inflection point as the dividing point, the features with M_relevance greater than this point are selected as the frontier feature subset; the original data set is divided into the frontier feature subset D1 and the remaining feature subset D2; where D1 is the search space of task 1; 203 Use mRMR to remove redundant features in the remaining feature set D2; 204: The frontier feature subset and the remaining feature set D3 after excluding redundant features are merged into a new feature set D4, which becomes a non-redundant feature set. D4 is the search space of task 2.
4. The evolutionary dual-task feature selection method based on hybrid initialization particle swarm optimization according to claim 3, characterized in that: The step three specifically includes:
301. Use skill factors to assign particles to each task; 302. If the value of the skill factor is 1, the particle is used to solve task 1; otherwise, it is used to solve task 2; 303. Define rmp to determine when the transfer process can be executed; 304. If rand>rmp, activate the knowledge transfer operation and use the position information of the current task to update the position of the particle; otherwise, use the position information from other tasks to update the position of the particle; 305. Use the fitness function to evaluate the selected feature subset.
Citation Information
Patent Citations
Image classification method and device based on continuous learning
CN114387486A
Method for optimizing support vector machine on basis of particle swarm optimization algorithm
WO2018072351A1