Multi-task multi-target feature selection method for solving unbalanced data

By constructing two tasks that optimize overall classification performance and minority category performance in the multi-task multi-objective feature selection method, and introducing a knowledge sharing mechanism, the high-dimensional and class non-balance problems are solved, and efficient feature selection and classification performance improvement is achieved.

CN120030316APending Publication Date: 2025-05-23SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510176900.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively solve the problems of high-dimensionality and class non-balance, resulting in poor classification performance in a few categories and the multi-objective feature selection method fails to make full use of the knowledge sharing mechanism between tasks.

Method used

A multi-task multi-objective feature selection method is proposed, by constructing two tasks: the first task optimizes the overall classification performance and the number of selected features, the second task prioritizes the classification performance of a few categories, and introduces a knowledge sharing mechanism to promote the migration and sharing of subsets of feature between different tasks.

Benefits of technology

Effectively improve the overall classification performance and the performance of a few categories, significantly reduce the number of selected features, while maintaining or improving classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030316A_ABST
    Figure CN120030316A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task multi-target feature selection method for solving unbalanced data. The method comprises the following steps: (1) acquiring an unbalanced data set; (2) evaluating a classification error rate using all features, and initializing a population; (3) selecting a mating pool and carrying out cross-task migration among feature subsets; (4) evaluating the classification performance of the new population; (5) respectively updating the two populations based on objective functions among different tasks; (6) repeatedly iterating until a stop condition is met, and outputting a non-dominated feature subset; according to the method, the overall classification performance and the minority classification performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data mining, and in particular to a multi-task and multi-objective feature selection method for solving unbalanced data. Background Art

[0002] In imbalanced data, the number of majority class samples far exceeds the number of minority class samples. Due to the accuracy-oriented design of standard classification methods and data difficulty factors such as intra-class imbalance, class overlap, and sparse representative samples, this imbalance causes learning models to usually perform poorly on minority classes. As a result, accurate predictions for minority groups (usually the group of primary interest) are greatly compromised. In the era of big data, the high dimensionality of datasets often exacerbates the challenges of class imbalance. These two problems are often interrelated. Specifically, high dimensionality easily leads to overfitting, especially for underrepresented minority samples. Conversely, class imbalance exacerbates the sparsity of data in high-dimensional space, making it difficult for classifiers to identify enough minority class samples for learning, ultimately affecting generalization and model performance. Class imbalance learning aims to address the challenges posed by datasets with highly skewed class distributions. To address the class imbalance problem, existing methods mainly include data-level and algorithmic approaches to class imbalance learning.

[0003] Among them, data-level resampling methods: undersampling and oversampling techniques balance the category distribution by adjusting the number of samples, but undersampling may lead to the loss of majority class information, while oversampling may cause overfitting problems. In addition, the simple resampling method fails to fully utilize the potential of feature selection to improve classification performance.

[0004] Algorithm-level methods: Although techniques such as cost-sensitive learning and ensemble learning have improved the classification performance of minority classes to some extent, they usually rely on modifying existing algorithms and have difficulty in optimizing multiple objectives simultaneously during feature selection, especially in complex environments with high dimensions and class imbalance.

[0005] Multi-objective feature selection methods: Although existing multi-objective feature selection methods such as the multi-objective fireworks algorithm and the feature selection method based on Jaccard similarity have achieved a certain balance between optimizing classification performance and the number of features, they often lack special designs for improving the performance of minority classes when dealing with class imbalance problems, and fail to fully utilize the knowledge sharing mechanism between tasks to enhance the generalization ability of feature subsets.

[0006] Evolutionary Multi-Task Feature Selection: Although methods combining multi-task learning with evolutionary computation have performed well in improving the performance of multiple optimization tasks, such as multi-task learning methods based on particle swarm optimization (PSO), which enhance knowledge transfer by sharing the best solutions, these methods usually fail to specifically optimize for class imbalance problems during the feature selection process, resulting in limited performance improvement for minority classes. Summary of the invention

[0007] Purpose of the invention: The purpose of the present invention is to provide a multi-task and multi-objective feature selection method for solving imbalanced data, to solve the high-dimensionality and class imbalance problems often faced in real data sets, while taking into account the overall classification performance of all categories and the classification accuracy of a few categories.

[0008] Technical solution: A multi-task and multi-objective feature selection method for solving unbalanced data according to the present invention comprises the following steps:

[0009] (1) Obtain an unbalanced dataset;

[0010] (2) Evaluate the classification error rate using all features and initialize the population;

[0011] (3) Select the mating pool and perform cross-task transfer between feature subsets;

[0012] (4) Evaluate the classification performance of the new population;

[0013] (5) Update the two populations based on the objective functions between different tasks;

[0014] (6) Repeat the iteration until the stopping condition is met and output the non-dominated feature subset.

[0015] Furthermore, step (2) includes the following steps:

[0016] (21) Construct 2 tasks: First Task is the overall classification performance, that is, the classification error rate f err (x) and the selected feature ratio f ratio The minimization of (x) is regarded as a constrained multi-objective optimization problem, and the formula is as follows:

[0017]

[0018] Where D is the total number of features in the dataset, x i is a binary vector representing a solution, i.e., a feature subset, where x i =1 and x i = 0 means selecting and discarding the i-th feature respectively; f err (x full) is the classification accuracy using all features; the g(x) constraint is used to ensure that the classification error rate after feature selection is less than or equal to the error rate when all features are used;

[0019] f err (x) is the balanced classification error rate, which is given by:

[0020]

[0021] Where c is the number of categories in the dataset. i represents the number of samples correctly classified in the i-th category, |S i | represents the sample size of the i-th class.

[0022] f ratio (x) is the ratio of the number of selected features to the total number of features, and the formula is as follows:

[0023]

[0024] Second Task is to prioritize the classification performance of minority categories; the second task τ 2 The objective function formula is as follows:

[0025]

[0026] Among them, w i is the weight of the i-th category;

[0027] (22) Two key perspectives, data and algorithm, were considered comprehensively when determining the weights for calculating each type of error rate;

[0028] (23) CIL-MTFS is used to evaluate the classification performance of all features, and then it is used as the constraint boundary of formula (1); then, a population is randomly generated and evaluated according to the objective function defined in formula (1) and formula (4), denoted as P 1 and P 2 .

[0029] Furthermore, in step (22), from the data perspective, the smaller the number of minority class samples, the greater its weight; from the algorithm perspective, the lower the classification accuracy of the minority class in the algorithm, the higher the assigned weight; wherein, the weight of the i-th class is set as:

[0030]

[0031] Furthermore, step (3) is as follows: First, in each iteration, 1 and P 2 In the example, we use the binary tournament strategy to create a mating pool MatingPool1 and MatingPool 2 Among them, MatingPool 1 Created based on non-dominated sorting, MatingPool 2 Created based on a cost-sensitive error rate; then, if MatingPool 1 If the minority class error rate of an individual in is greater than the average minority class error rate of the population, then MatingPool 2 Replace MatingPool with the corresponding individual 1 In order to transfer the individuals with better performance in the minority class to MatingPool 1 In, promote the task To the task Migration and sharing of useful features. Then, the offspring generation operator is based on MatingPool 1 Generate a progeny population O, where the progeny population O is calculated according to formula (1) According to formula (4), Evaluate; then respectively from P 1 and O, P 2 and O to update P respectively. 1 and P 2 , until the stopping condition is met, that is, the predetermined maximum number of function evaluations, and return P 1 The non-dominated feature subset in .

[0032] Furthermore, the specific process of the binary tournament strategy is as follows: different environment selection strategies are adopted: First, the solutions are compared according to the degree to which they violate the constraints; if both solutions are feasible, the non-dominated sorting and crowding distance commonly used in NSGA-II are applied to select the solution from the parent population P. 1 and the offspring population O, select the N best solutions; otherwise, compare the degree to which each solution violates the constraints; for The environment selection is based solely on the cost-sensitive classification error rate, defined by equation (4).

[0033] Furthermore, step (4) is as follows: for each individual x in the offspring population O, according to formula (1) According to formula (4), Evaluation is performed and the evaluation results are used for environmental selection to ensure that the population evolves towards the optimization goal.

[0034] Furthermore, step (5) is as follows: 1 and O, select the optimal solution based on the constraint dominance criterion;2 Combined with O, the optimal solution is selected based on the cost-sensitive classification error rate.

[0035] A multi-task and multi-objective feature selection system for solving unbalanced data according to the present invention comprises:

[0036] Acquisition module: used to obtain unbalanced data sets;

[0037] Initialization module: used to evaluate the classification error rate using all features and initialize the population;

[0038] Migration module: used to select the mating pool and perform cross-task migration between feature subsets;

[0039] Evaluation module: used to evaluate the classification performance of new populations;

[0040] Update module: used to update the two populations based on the objective functions between different tasks;

[0041] Iteration module: used to repeat iterations until the stopping condition is met and output the non-dominated feature subset.

[0042] An electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, it implements any one of the multi-task and multi-objective feature selection methods for solving unbalanced data.

[0043] A storage medium described in the present invention stores a computer program, and when the computer program is executed by a processor, it implements any one of the multi-task and multi-objective feature selection methods for solving unbalanced data.

[0044] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: The present invention proposes to formulate the feature selection task in class imbalanced learning as a multi-task problem, in which each task has a different focus. The first task is a constrained multi-objective optimization problem, which aims to optimize the overall classification performance of all categories and the number of selected features. The second task is specifically aimed at improving the classification performance of minority classes. In addition, the present invention also introduces a knowledge sharing mechanism to promote the transfer of information and universal features between different tasks. This mechanism can effectively reuse feature subsets evolved through cross-task genetic migration, which helps to improve the overall classification performance and the performance of minority classes. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a schematic diagram of the process of the present invention;

[0046] Figure 2 It is the code diagram of Algorithm 1 of the present invention;

[0047] Figure 3 It is the cross-task genetic migration graph of Algorithm 2 of the present invention;

[0048] Figure 4 It is the characteristic of the unbalanced classification data set of Table 1 of the present invention;

[0049] Figure 5 It is the comparison result of the whole set of Table II of the present invention and CIL-MTFS on the test data;

[0050] Figure 6 It is the comparison result of the average F1 scores obtained by NSGA-II, DEAEA, MFFS, BSOEA, PRDH and CIL-MTFS on the test data of the present invention.

[0051] Figure 7 is the comparison result of the average G average values ​​obtained by NSGA-II, DEAEA, MFFS, BSOEA, PRDH and CIL-MTFS on the test data of the present invention in Table IV;

[0052] Figure 8 Table V is a comparison result of the Mcer average values ​​obtained on the test data by NSGA-II, DEAEA, MFFS, BSOEA, PRDH and CIL-MTFS of the present invention. DETAILED DESCRIPTION

[0053] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0054] like Figure 1-Figure 8 As shown, an embodiment of the present invention provides a multi-task multi-objective feature selection method for solving unbalanced data, comprising the following steps:

[0055] (1) Obtaining imbalanced datasets; Table I shows 14 real-world imbalanced classification datasets. These datasets cover a variety of classification tasks with varying degrees of class imbalance and feature dimensions. The number of features ranges from 22 in the Parkinson dataset to 22,283 in the GLI-85 dataset, while the number of samples ranges from 32 in the lung dataset to 1,593 in the Semeion dataset. There are also significant differences in the imbalance ratio (IR), with Semeion having the highest IR of 9.08 and WBCD having the lowest IR of 1.68. Therefore, they can comprehensively test the performance of different comparison methods.

[0056] (2) Evaluate the classification error rate using all features and initialize the population; including the following steps:

[0057] (21) Construct 2 tasks: First Task is the overall classification performance, that is, the classification error rate f err (x) and the selected feature ratio f ratio The minimization of (x) is regarded as a constrained multi-objective optimization problem, and the formula is as follows:

[0058]

[0059] Where D is the total number of features in the dataset, x i is a binary vector representing a solution, i.e., a feature subset, where x i =1 and x i = 0 means selecting and discarding the i-th feature respectively; the g(x) constraint is used to ensure that the classification error rate after feature selection is less than or equal to the error rate when all features are used;

[0060] f err (x) is the balanced classification error rate, which is given by:

[0061]

[0062] Where c is the number of categories in the dataset. i represents the number of samples correctly classified in the i-th category, |S i | represents the sample size of the i-th class.

[0063] f ratio (x) is the ratio of the number of selected features to the total number of features, and the formula is as follows:

[0064]

[0065] Second Task is to prioritize the classification performance of minority categories; the second task τ 2 The objective function formula is as follows:

[0066]

[0067] Among them, w i is the weight of the i-th category;

[0068] (22) When determining the weight for calculating each type of error rate, two key perspectives, data and algorithm, are considered comprehensively. From the data perspective, the smaller the number of minority class samples, the greater its weight; from the algorithm perspective, the lower the classification accuracy of the minority class in the algorithm, the higher the weight assigned. The weight of the i-th class is set as follows:

[0069]

[0070] (23) CIL-MTFS is used to evaluate the classification performance of all features, and then it is used as the constraint boundary of formula (1); then, a population is randomly generated and evaluated according to the objective function defined in formula (1) and formula (4), denoted as P 1 and P 2 .

[0071] (3) Select the mating pool and perform cross-task migration between feature subsets; the details are as follows: First, in each iteration, 1 and P 2 In the example, we use the binary tournament strategy to create a mating pool MatingPool 1 and MatingPool 2 Among them, MatingPool 1 Created based on non-dominated sorting, MatingPool 2 Based on the cost-sensitive error rate of formula (4), if MatingPool 1 If the minority class error rate of an individual in is greater than the average minority class error rate of the population, then MatingPool 2 Replace MatingPool with the corresponding individual 1 In order to transfer the individuals with better performance in the minority class to MatingPool 1 In, promote the task To the task Migration and sharing of useful features. Then, the offspring generation operator is based on MatingPool 1 Generate a progeny population O, where the progeny population O is calculated according to formula (1) According to formula (4), Evaluate; then respectively from P 1 and O, P 2 and O to update P respectively. 1 and P 2 , until the stopping condition is met, that is, the predetermined maximum number of function evaluations, and return P 1 The non-dominated feature subset in .

[0072] The specific process of the binary tournament strategy is as follows: Different environment selection strategies are used: First, the solutions are compared according to the degree to which they violate the constraints; if both solutions are feasible, the non-dominated sorting and crowding distance commonly used in NSGA-II are applied to select the solution from the parent population P. 1 and the offspring population O, select the N best solutions; otherwise, compare the degree to which each solution violates the constraints; for The environment selection is based solely on the cost-sensitive classification error rate, defined by equation (4).

[0073] (4) Evaluate the classification performance of the new population; specifically: for each individual x in the offspring population O, use formula (1) to According to formula (4), Evaluation is performed and the evaluation results are used for environmental selection to ensure that the population evolves towards the optimization goal.

[0074] (5) Update the two populations based on the objective functions between different tasks; specifically: 1 and O, select the optimal solution based on the constraint dominance criterion; 2 Combined with O, the optimal solution is selected based on the cost-sensitive classification error rate.

[0075] (6) Repeat the iteration until the stopping condition is met and output the non-dominated feature subset.

[0076] High-dimensional data and class imbalance are very common in practical applications, such as medical diagnosis (rare diseases), fraud detection (a small number of abnormal samples), text classification (scarce data in some categories), etc. The CIL-MTFS method proposed in this paper can effectively handle the data challenges brought by class imbalance and high dimensionality.

[0077] The present invention can maintain or even improve the classification performance while significantly reducing the number of selected features. Table II lists the results of the full set and CIL-MTFS in terms of MCER, F1-score, G-mean and FN indicators. Among the 14 datasets, the invented CIL-MTFS method outperforms the full feature set in classification performance (i.e., MCER, F1-score and G-mean) in at least 10 datasets while significantly reducing the number of features used. For example, on the Yeoh-2002v1 dataset, CIL-MTFS only selected about 51 features from 5469 features, but the classification accuracy was improved by 4.76%. Similarly, on the high-dimensional CNS dataset, CIL-MTFS selected about 1053 features from the original 7129 features, and the classification accuracy was improved by 3.85%. These results show that the CIL-MTFS method can effectively identify a small number of valuable features from the original feature set, and it can even improve the classification performance compared to using all features.

[0078] Table III lists the average F1 score comparisons achieved by NSGA-II, DAEA, MFFS, BSOEA, PRDH, and CIL-MTFS on the test data, with the best results for each dataset marked in bold. As shown in Table III, among the six state-of-the-art methods compared, the proposed CIL-MTFS method achieves the best results on 7 of the 14 datasets. The Wilcoxon rank sum test results show that CIL-MTFS significantly outperforms NSGA-II, DAEA, MFFS, BSOEA, and PRDH on 6, 4, 3, 5, and 3 datasets, respectively. In contrast, none of the competing algorithms performs significantly better than CIL-MTFS on more than one dataset. In addition, the Friedman test results at the bottom of Table III confirm that CIL-MTFS ranks first among the six algorithms, further highlighting its superior performance in terms of the F1 score metric.

[0079] Tables IV and V show similar patterns for the G-mean and MCER metrics, respectively. A closer look at these results reveals that CIL-MTFS performs particularly well on high-dimensional datasets such as Yeoh-2002-v1, DLBCL, CNS, and GLI-85, which contain 2526, 5469, 7129, and 22283 features, respectively. The number of instances in these datasets is also much smaller than the number of features, resulting in a highly sparse data space, which poses a huge challenge to various feature selection methods. Despite this, CIL-MTFS achieves the best overall classification performance on these datasets, demonstrating its effectiveness in handling high-dimensional data.

[0080] Experimental process: Since the invented CIL-MTFS method is based on multi-objective optimization, the present invention compares the performance of CILMTFS with five multi-objective feature selection methods: NSGA-II, DAEA, MFFS, BSOEA and PRDH. All these methods take the balanced classification error rate in formula (2) as the primary goal, so that they can handle unbalanced classification tasks. Specifically, MFFS also adopts a dual-task approach similar to CIL-MTFS, but its auxiliary task focuses on enhancing the diversity of feature subsets of the target task. We use NSGA-II as a competitor because the environment selection of task T1 in the present invention is based on NSGA-II. DAEA, BSOEA and PRDH are listed as the most advanced feature selection methods, providing a powerful comparison for evaluating the effectiveness of CIL-MTFS.

[0081] Parameter setting: During feature subset search, each dataset is randomly divided into a training data subset (about 70%) and a test data subset (about 30%). KNN (K=5) with five-fold cross validation is used as the basic classification algorithm, and the classification error rate of the training data subset is calculated. Each feature selection method is run 30 times independently, and the average performance is reported. The maximum number of generations for all methods is set to 100. The population size of each algorithm is equal to the number of features (N=D), but the upper limit is 200 (if D>200, then N=200) to prevent the computational cost of high-dimensional datasets from being too high. Single-point crossover and bit flip mutation operators are used to generate offspring solutions, with crossover and mutation probabilities of 1.0 and 1 / (D), respectively. The non-dominated feature subsets obtained by each multi-objective-based method during training are evaluated on an unseen test set. The custom parameters of each algorithm are set according to the values ​​used in their original papers, which have demonstrated their effectiveness. All experiments were conducted in the Linux environment, and the CPU was Intel Core i7-11700K with a main frequency of 3.6 GHz.

[0082] Performance Indicators: Based on different experiments, the following six performance indicators are used to evaluate the performance of each method, as they can capture different aspects of the multi-objective feature selection results:

[0083] a Minimum balanced classification error rate (MCER): Among multiple non-dominant feature subsets, select the feature subset with the smallest training classification error rate (Formula (2)), and record its test classification error rate as MCER.

[0084] bF1-score: This is a widely used metric for unbalanced data, which calculates the harmonic mean of precision and recall. Among multiple non-dominant feature subsets, select the feature subset with the largest F1 score and record its test F1 score.

[0085]

[0086] in,

[0087] c Geometric mean (G-mean): This is also a commonly used metric for unbalanced data. It calculates the geometric mean of the accuracy of each category. Among multiple non-dominant feature subsets, select the feature subset with the largest training G-mean value and record its test G-mean value.

[0088]

[0089] d Selected feature number (FN): represents the number of features selected by the feature subset with the minimum classification error rate.

[0090] eHypervolume (HV): This is a widely used metric in the field of multi-objective optimization to measure the diversity and convergence of the non-dominant solutions obtained.

[0091] fTraining time: measures the efficiency of each compared algorithm on the training data.

[0092] It should be noted that among these six performance indicators, MCER, F1score, G-mean, and FN evaluate the classification performance of a single feature subset, which is of most interest to decision makers. HV evaluates the performance of multiple non-dominated feature subsets.

[0093] Results and data analysis:

[0094] Comparison with the full set: The invented CIL-MTFS method outperforms the full feature set in terms of classification performance (i.e., MCER, F1-score, and G-mean) in at least 10 of the 14 datasets, while significantly reducing the number of features used. For example, on the Yeoh-2002v1 dataset, CIL-MTFS selected only about 51 features out of 5469, but the classification accuracy was improved by 4.76%. Similarly, on the high-dimensional CNS dataset, CIL-MTFS selected about 1053 features out of the original 7129 features, improving the classification accuracy by 3.85%. These results show that the CIL-MTFS method can effectively identify a small number of valuable features from the original feature set, and it can even improve the classification performance compared to using all features.

[0095] Comparison of Classification Performance: Among the six state-of-the-art methods compared, the proposed CIL-MTFS method achieves the best results on 7 out of 14 datasets. The Wilcoxon rank sum test results show that CIL-MTFS significantly outperforms NSGA-II, DAEA, MFFS, BSOEA, and PRDH on 6, 4, 3, 5, and 3 datasets, respectively. In contrast, none of the competing algorithms performs significantly better than CIL-MTFS on more than one dataset. In addition, the Friedman test results confirm that CIL-MTFS ranks first among the six algorithms, further highlighting its superior performance in terms of the F1 score metric. CIL-MTFS performs particularly well on high-dimensional datasets such as Yeoh-2002-v1, DLBCL, CNS, and GLI-85, which contain 2526, 5469, 7129, and 22283 features, respectively. The number of instances in these datasets is also much smaller than the number of features, resulting in a highly sparse data space, which poses a great challenge to various feature selection methods. Nevertheless, CIL-MTFS achieves the best overall classification performance on these datasets, demonstrating its effectiveness in handling high-dimensional data.

[0096] The embodiment of the present invention further provides a multi-task multi-objective feature selection system for solving unbalanced data, comprising:

[0097] Acquisition module: used to obtain unbalanced data sets;

[0098] Initialization module: used to evaluate the classification error rate using all features and initialize the population;

[0099] Migration module: used to select the mating pool and perform cross-task migration between feature subsets;

[0100] Evaluation module: used to evaluate the classification performance of new populations;

[0101] Update module: used to update the two populations based on the objective functions between different tasks;

[0102] Iteration module: used to repeat iterations until the stopping condition is met and output the non-dominated feature subset.

[0103] An embodiment of the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, the computer program implements any one of the multi-task and multi-objective feature selection methods for solving unbalanced data.

[0104] An embodiment of the present invention further provides a storage medium storing a computer program, wherein the computer program, when executed by a processor, implements any one of the multi-task and multi-objective feature selection methods for solving unbalanced data.

Claims

1. A multi-task and multi-objective feature selection method for solving unbalanced data, characterized in that: The following steps are involved: (1) Obtain an unbalanced dataset; (2) Evaluate the classification error rate using all features and initialize the population; (3) Select the mating pool and perform cross-task transfer between feature subsets; (4) Evaluate the classification performance of the new population; (5) Update the two populations based on the objective functions between different tasks; (6) Repeat the iteration until the stopping condition is met and output the non-dominated feature subset.

2. A multi-task and multi-objective feature selection method for solving unbalanced data according to claim 1, characterized in that: Step (2) comprises the following steps: (21) Create two tasks: Task First Task is the overall classification performance, that is, the classification error rate f eeeeee (x) and the selected feature ratio f eerrrrrrrr The minimization of (x) is regarded as a constrained multi-objective optimization problem, and the formula is as follows: Where D is the total number of features in the dataset, x i is a binary vector representing a solution, i.e., a feature subset, where x i =1 and x i = 0 means selecting and discarding the i-th feature respectively; f err (x full ) is the classification accuracy using all features; the g(x) constraint is used to ensure that the classification error rate after feature selection is less than or equal to the error rate when all features are used; f eeeeee (x) is the balanced classification error rate, which is given by: Where c is the number of categories in the dataset. rr represents the number of samples correctly classified in the i-th category, |S i | represents the sample size of the i-th class. f eerrrrrrrr (x) is the ratio of the number of selected features to the total number of features, and the formula is as follows: Task Second Task It gives priority to the classification performance of minority categories; the objective function formula of the second task τ2 is as follows: Among them, w rr is the weight of the i-th category; (22) Two key perspectives, data and algorithm, were considered comprehensively when determining the weights for calculating each type of error rate; (23) CIL-MTFS is used to evaluate the classification performance of all features, and then used as the constraint boundary of formula (1); subsequently, a population is randomly generated and evaluated according to the objective functions defined in formula (1) and formula (4), denoted as P1 and P2.

3. A multi-task and multi-objective feature selection method for solving unbalanced data according to claim 2, characterized in that: In step (22), from the data perspective, the smaller the number of minority class samples, the greater its weight; from the algorithm perspective, the lower the average classification accuracy of the minority class in the population, the higher the assigned weight; where the weight of the i-th class is set:

4. A multi-task and multi-objective feature selection method for solving unbalanced data according to claim 2, characterized in that: Step (3) is as follows: First, in each iteration, the mating groups MatingPool1 and MatingPool2 are created from P1 and P2 respectively using the binary tournament strategy; Among them, MatingPool1 is created based on constrained non-dominated sorting, and MatingPool2 is created based on the cost-sensitive error rate of formula (4); then, if the minority class error rate of an individual in MatingPool1 is greater than the average minority class error rate of the population, the corresponding individual in MatingPool2 replaces the individual in MatingPool1 to migrate the individual with better minority class performance to MatingPool1, which promotes the task To the task The migration and sharing of useful features. Then, the offspring population O is generated based on MatingPool1 through the offspring generation operator, where the offspring population O is based on formula (1) According to formula (4), Evaluate; then update P1 and P2 respectively by selecting from the combined population of P1 and O, and P2 and O, until the stopping condition, that is, the predetermined maximum number of function evaluations, is met, and the non-dominated feature subset in P1 is returned.

5. A multi-task and multi-objective feature selection method for solving unbalanced data according to claim 4, characterized in that: The specific process of the binary tournament strategy is as follows: Different environment selection strategies are adopted: First, the solutions are compared according to the degree to which they violate the constraints. If both solutions are feasible, the non-dominated sorting and crowding distance commonly used in NSGA-II are applied to select the N best solutions from the combination of the parent population P1 and the child population O. Otherwise, each solution is compared according to the degree to which it violates the constraints. The environment selection is based solely on the cost-sensitive classification error rate, defined by equation (4).

6. A multi-task and multi-objective feature selection method for solving unbalanced data according to claim 1, characterized in that: Step (4) is as follows: for each individual x in the offspring population O, the constrained multi-objective feature selection objective function of formula (1) is To evaluate, according to the cost-sensitive objective function of formula (4), Evaluation is performed and the evaluation results are used for environmental selection to ensure that the population evolves towards different preferences for overall classification performance and minority class performance.

7. A multi-task and multi-objective feature selection method for solving unbalanced data according to claim 1, characterized in that: Step (5) is as follows: merge P1 and O, and select the optimal solution based on the constraint dominance criterion; merge P2 and O, and select the optimal solution based on the cost-sensitive classification error rate.

8. A multi-task and multi-objective feature selection system for solving unbalanced data, characterized in that: include: Acquisition module: used to obtain unbalanced data sets; Initialization module: used to evaluate the classification error rate using all features and initialize the population; Migration module: used to select the mating pool and perform cross-task migration between feature subsets; Evaluation module: used to evaluate the classification performance of new populations; Update module: used to update the two populations based on the objective functions between different tasks; Iteration module: used to repeat iterations until the stopping condition is met and output the non-dominated feature subset.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, a multi-task and multi-objective feature selection method for solving unbalanced data is implemented according to any one of claims 1 to 7.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a multi-task and multi-objective feature selection method for solving unbalanced data is implemented according to any one of claims 1 to 7.