Bioinformatics multi-label feature selection method and system based on weighted correlation
By calculating label weights and fuzzy conditional mutual information, and combining the joint information of candidate features and selected features, the problem of inaccurate feature correlation measurement in bioinformatics multi-label feature selection is solved, achieving more efficient feature selection and improved classification performance.
Patent Information
- Application Number
- CN202511052950.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-18
AI Technical Summary
Existing bioinformatics multi-label feature selection methods fail to accurately consider the combined effects of different candidate features and selected features, as well as the label relationships, leading to inaccurate feature correlation measurements and affecting classification performance.
We employ a bioinformatics multi-label feature selection method based on weighted correlation. By calculating label weights and fuzzy conditional mutual information, and combining the joint information content of candidate features and selected features, we calculate the importance score of features and select a high-quality feature subset.
It improves the accuracy of feature selection and the performance of classification models, reduces redundant information, and enhances the classification accuracy and efficiency of multi-label learning models.
Smart Images

Figure CN120977398A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of machine learning and data mining technology, specifically relating to a bioinformatics multi-label feature selection method and system based on weighted correlation. Background Technology
[0002] In the field of bioinformatics, with the continuous increase in multi-label data, not only is the feature space of the data highly dimensional, but the size of the label set is also constantly expanding. High-dimensional data often contains a large number of irrelevant features, which increases the computational complexity of classification models and leads to a decline in classification performance. As an effective data preprocessing technique, feature selection refers to selecting features that are highly relevant to the label information from the numerous features of high-dimensional data. This removes redundant and irrelevant features, extracts important features relevant to the current classification task, thereby reducing data dimensionality. This alleviates the conflict between the large number of features and labels and limited storage space. By providing high-quality input data for the classification model, it helps improve classification accuracy, accelerates training speed, and prevents overfitting. Therefore, feature selection is a key preprocessing step in machine learning, data mining, and pattern recognition, and feature selection in bioinformatics provides a solid foundation for subsequent bioinformatics analysis.
[0003] Information theory is a commonly used evaluation metric for measuring linear and nonlinear relationships between variables and is widely applied in feature selection. However, information-theoretic metrics are suitable for quantifying the correlation between discrete features and cannot directly handle continuous numerical features. For bioinformatics data, the uniform feature type is often mixed, containing both discrete and continuous features. Feature selection methods based on fuzzy mutual information allow for handling mixed discrete and continuous data, thus better adapting to the diversity and complexity of bioinformatics data. However, multi-label feature selection methods based on fuzzy mutual information utilize the fuzzy mutual information between candidate features and all labels, or consider the fuzzy conditional mutual information under the influence of selected features, to measure feature correlation. In measuring the correlation between different candidate features, existing methods typically treat the influence of selected features only as a condition, making the influence of selected features imprecise in the selection process of different candidate features. In reality, the combined effect of different candidate features and selected features on the same label has different relevance, meaning it is necessary to consider the combined effect of different pairs of candidate features and selected features under different labels. Furthermore, these existing methods do not consider combining the importance of different labels with the correlation between features, ignoring the role of label relationships, resulting in inaccurate measurement of feature correlation. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the technical problem this invention aims to solve is to provide a method and system for selecting multi-label bioinformatics features based on weighted correlation.
[0005] The present invention solves the aforementioned technical problem by adopting the following technical solution:
[0006] A feature selection system includes a memory and a processor; the memory stores a computer program that, when run on the processor, implements a bioinformatics multi-label feature selection method based on weighted correlation; the method includes the following steps:
[0007] Step 1: Obtain the candidate feature set and label set, and initialize the selected feature subset to empty;
[0008] Step 2: Calculate the label weights. Based on the label weights and the fuzzy mutual information between the candidate features and the labels, calculate the amount of classification information provided by each candidate feature for the label set. Add the candidate feature with the largest amount of classification information to the selected feature subset and remove the candidate feature from the candidate feature set.
[0009] The formula for calculating tag weight is:
[0010]
[0011] In the formula, w(l) i ) indicates label l i The weights, I(·) represent the mutual information between the two labels, l i l j and l k L represents the tag set, and M represents the number of tags.
[0012] The formula for calculating the amount of categorized information is:
[0013]
[0014] In the formula, relL(f n ;L) represents candidate feature f n The amount of classification information provided by the label set, FMI(f n ;l i ) represents candidate feature f n With label l i The fuzzy mutual information between them;
[0015] Step 3: When the selected feature subset is not empty, calculate the weighted feature correlation between each remaining candidate feature in the candidate feature set and the label set;
[0016] The joint conditional information content of the candidate features and the selected feature subset is calculated according to the following formula:
[0017] relS(f n ,S,l i )=w(l i )·FCMI(f n;S|l i (8)
[0018] In the formula, relS(f n ,S,l i ) indicates in the label l i Candidate features f under the action n The joint conditional information content with the selected feature subset S, FCMI(f n ;S|l i ) indicates in the label l i Candidate features f under the action n Fuzzy conditional mutual information between the selected feature subset S and the selected feature subset S;
[0019] The weighted correlation between candidate features and the label set is calculated using the following formula:
[0020]
[0021] Step 4: Calculate the amount of redundant information between the remaining candidate features and the selected feature subsets;
[0022] MS(f n ,S)=FMI(f n ;S)(13)
[0023] In the formula, MS(f n S) represents the candidate feature f n The amount of redundant information in the selected feature subset, FMI(f n S) represents candidate feature f n The fuzzy mutual information with the selected feature subset;
[0024] Step 5: Calculate the importance score of each remaining candidate feature in the candidate feature set;
[0025]
[0026] In the formula, J(f) n ) represents candidate feature f n Importance score;
[0027] Add the candidate feature with the highest importance score to the selected feature subset, and remove the candidate feature from the candidate feature set;
[0028] Step 6: If the number of features in the selected feature subset is less than the set dimension, repeat steps 3 to 5 until the number of features in the selected feature subset equals the set dimension; if the number of features in the selected feature subset equals the set dimension, stop; at this point, feature selection is complete.
[0029] Furthermore, the formula for calculating fuzzy mutual information is as follows:
[0030]
[0031] In the formula, ∩ represents the minimum value, and |·| represents the sum of the membership degrees of the sample with all other samples based on the candidate features. and They represent samples x respectively u With candidate features f n Tag l i The fuzzy equivalence relation is specifically represented as follows:
[0032]
[0033] In the formula, Indicates sample x u and x v Based on candidate features f n membership degree Indicates sample x u and x v Based on label i Membership degree is used to measure fuzzy relations in fuzzy mutual information; Representing candidate features f n Regarding sample x v The fuzzy relationship Indicates label l i Regarding sample x v The fuzzy relationship, f nu f nv Indicates sample x u and x v Candidate features, l iu l iv Indicates sample x u and x v The i-th label represents the fuzzy relation connector, and p is the number of samples.
[0034] Furthermore, the formula for calculating fuzzy conditional mutual information is as follows:
[0035]
[0036] In the formula, [x u ] S Indicates sample x u Fuzzy equivalence relation with the selected feature subset, Indicates sample x u and x v Based on the membership degree of the selected feature subset, ||S u -S v || 2 Indicates sample x u With xv The square of the Euclidean distance between sample vectors consisting of all selected feature descriptions in the selected feature subset, where X represents the training set.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] (1) For any tag in the tag set, the present invention calculates the weight of the tag considering the relationship between the tags. The larger the weight, the higher the importance of the tag to the tag set, and the higher the feature importance of the tag.
[0039] (2) Existing feature selection methods based on fuzzy mutual information only focus on the correlation between candidate features and the label set, or use selected features as conditional aids, neglecting the effect of the joint information of candidate and selected features on feature relevance, leading to inaccurate feature importance measurement. To solve this problem, this invention introduces fuzzy conditional mutual information for feature relevance measurement, considering the joint effect of paired candidate and selected features to obtain more diverse and accurate classification information. By combining label weights with fuzzy conditional mutual information, a new weighted feature relevance measurement is proposed.
[0040] (3) By utilizing the fuzzy mutual information between candidate features and selected features to measure feature redundancy, a new importance score calculation method is proposed, which combines the weighted feature correlation between candidate features and the label set, and the amount of redundant information between candidate features and subsets of selected features, as a feature evaluation standard. This method fully considers the different roles of candidate features with labels and selected features, and selects the feature subset that provides high-quality input data for the multi-label learning model. Attached Figure Description
[0041] Figure 1 This is the overall flowchart of the present invention. Detailed Implementation
[0042] Specific embodiments are given below. These specific embodiments are only used to describe the technical solution of the present invention in detail, and are not intended to limit the scope of protection of this application.
[0043] This invention provides a bioinformatics multi-label feature selection method based on weighted correlation, comprising the following steps:
[0044] Step 1: Obtain the candidate feature set F = {f1, f2, ..., f N The corresponding tag set is L = {l1, l2, ..., l}. M}; where f1, f2, ..., f N Let l1, l2, ..., l represent candidate features, and N represent the number of candidate features. MLet M represent the number of labels; initialize the selected feature subset S to be empty, and set the dimension of the selected feature subset to K.
[0045] Step 2: Calculate the weight of each label in the label set based on mutual information. According to the weight of the label and the fuzzy mutual information between the candidate feature and the label, calculate the amount of classification information provided by each candidate feature for the label set. Take the candidate feature with the largest amount of classification information as the first selected feature, add it to the selected feature subset S, and remove the candidate feature from the candidate feature set.
[0046] The formula for calculating the weight of a tag is:
[0047]
[0048] In the formula, w(l) i ) indicates label l i The weights, 0 < w(l) i <1, the weight reflects the importance of the label, the more important the label, the greater the weight; I(·) represents the mutual information between two labels, characterizing the correlation between the two labels; l i l j and l k Indicates a label.
[0049] The formula for calculating the amount of classification information provided by the candidate features for the label set is as follows:
[0050]
[0051] In the formula, relL(f n ;L) represents candidate feature f n The amount of classification information provided for the label set L refers to the candidate feature f n Weighted fuzzy mutual information with all labels; FMI(f n ;l i ) represents candidate feature f n With label l i The fuzzy mutual information between them;
[0052] Let X = {x1, x2, ..., x} p Let} represent the training set, and each sample be described by a set of candidate features. Then, the sample is represented as x. u =[f 1u ,f 2u ,...,f Nu ]; u = 1, 2, ..., p; where p is the sample size, f 1u ,f 2u ,...,f Nu Indicates sample x u Candidate features;
[0053] The formula for calculating the fuzzy mutual information between candidate features and labels is:
[0054]
[0055] In the formula, ∩ represents the minimum value, and |·| represents the sum of the membership degrees of the sample with all other samples based on the candidate features. and They represent samples x respectively u With candidate features f n Tag l i The fuzzy equivalence relation is specifically represented as follows:
[0056]
[0057]
[0058] In the formula, Indicates sample x u and x v Based on candidate features f n membership degree Indicates sample x u and x v Based on label i Membership degree is used to measure fuzzy relations in fuzzy mutual information; Representing candidate features f n Regarding sample x v The fuzzy relationship Indicates label l i Regarding sample x v The fuzzy relationship, f nu f nv Indicates sample x u and x v Candidate features, l iu l iv Indicates sample x u and x v The i-th label represents the fuzzy relation connector.
[0059] Step 3: When the selected feature subset S is not empty, calculate the weighted feature correlation between each remaining candidate feature in the candidate feature set and the label set;
[0060] First, the joint conditional information of the candidate features and the selected feature subset is calculated based on the label weights and fuzzy conditional mutual information.
[0061] relS(f n ,S,l i )=w(l i )·FCMI(fn ;S|l i (8)
[0062] In the formula, relS(f n ,S,l i ) indicates in the label l i Candidate features f under the action n The joint conditional information content with the selected feature subset S, FCMI(f n ;S|l i ) indicates in the label l i Candidate features f under the action n The fuzzy conditional mutual information between the selected feature subset S and the selected feature subset S is calculated using the following formula:
[0063]
[0064] In the formula, [x u ] S Indicates sample x u The fuzzy equivalence relation with the selected feature subset S is defined as follows:
[0065]
[0066] In the formula, Indicates sample x u and x v Based on the membership degree of the selected feature subset S; ||S u -S v || 2 Indicates sample x u With x v The square of the Euclidean distance between sample vectors composed of all selected feature descriptions in the selected feature subset S represents the distance between two samples based on the feature measure in the selected feature subset S. The closer the distance, the more similar the two are, and the greater the corresponding membership degree.
[0067] Based on the amount of classification information provided by the candidate features to the label set and the amount of joint conditional information between the candidate features and the selected feature subset, the weighted feature correlation between the candidate features and the label set is calculated by Equation (12), thereby achieving an accurate measurement of the correlation between the candidate features and the label set.
[0068]
[0069] In the formula, FI(f) n L) represents candidate feature f n The weighted feature correlation with the label set L.
[0070] Step 4: Calculate the amount of redundant information between each remaining candidate feature in the candidate feature set and the selected feature subset;
[0071] MS(f n ,S)=FMI(f n ;S)(13)
[0072] In the formula, MS(f n S) represents the candidate feature f n The amount of redundant information in the selected feature subset, FMI(f n S) represents candidate feature f n The fuzzy mutual information with the selected feature subset;
[0073] Step 5: Based on the weighted feature correlation between the candidate features and the label set, and the amount of redundant information between the candidate features and the selected feature subset, calculate the importance score of each remaining candidate feature in the candidate feature set using Equation (14).
[0074]
[0075] In the formula, J(f) n ) represents candidate feature f n Importance score;
[0076] The candidate feature with the highest importance score is added to the selected feature subset, and then removed from the candidate feature set F.
[0077] Step 6: If the number of features in the selected feature subset is less than the dimension K, repeat steps 3 to 5 until the number of features in the selected feature subset equals the dimension K, thus completing feature selection; if the number of features in the selected feature subset equals the dimension K, stop, thus completing feature selection; use the selected feature subset for downstream classification tasks.
[0078] The present invention also provides a feature selection system, including a memory and a processor; the memory stores a computer program, which runs on the processor, and the processor implements the above-described feature selection method when running the computer program.
[0079] Example
[0080] This embodiment uses the Yeast dataset as an example. The Yeast dataset is a typical bioinformatics multi-label dataset, which includes 2417 samples, 103 features, and 14 labels; the training set consists of 1500 samples, and the remaining 917 samples constitute the test set.
[0081] To verify the effectiveness of the method of this invention, the selected feature subset obtained by the method of this invention and the selected feature subset obtained by the existing feature selection method based on fuzzy mutual information (MUCO) were used to train a multi-label classification model (MLKNN). (AveragePrecision, AP), (CoverageError, CE), (ZeroOneLoss, ZOL), (HammingLoss, HL), and (RankingLoss, RL) were used as evaluation metrics to assess the performance of the multi-label classification model, as shown in Table 1. A higher AP value indicates better model classification performance, while lower values for CE, ZOL, HL, and RL indicate better model classification performance.
[0082] Table 1 Comparison of classification performance between the method of this invention and MUCO
[0083]
[0084] The experimental results in the table show that the selected feature subset obtained by this method outperforms the MUCO method in all metrics. Compared with the MUCO method, this method utilizes fuzzy conditional mutual information to measure the effect of the joint information of label-weighted candidate features and selected features on feature relevance, retaining more diverse and effective classification information in the selected feature subset. Therefore, the selected feature subset obtained by the method of this invention helps to improve the classification performance of the model.
[0085] Any aspects not covered in this invention are applicable to existing technologies.
Claims
1. A feature selection system, comprising a memory and a processor; the memory stores a computer program, which, when executed on the processor, implements a bioinformatics multi-label feature selection method based on weighted correlation; characterized in that, The method includes the following steps: Step 1: Obtain the candidate feature set and label set, and initialize the selected feature subset to empty; Step 2: Calculate the label weights. Based on the label weights and the fuzzy mutual information between the candidate features and the labels, calculate the amount of classification information provided by each candidate feature for the label set. Add the candidate feature with the largest amount of classification information to the selected feature subset and remove the candidate feature from the candidate feature set. The formula for calculating tag weight is: In the formula, w(l) i ) indicates label l i The weights, I(·) represent the mutual information between the two labels, l i l j and l k L represents the tag set, and M represents the number of tags. The formula for calculating the amount of categorical information is: In the formula, relL(f n ;L) represents candidate feature f n The amount of classification information provided by the label set, FMI(f n ;l i ) represents candidate feature f n With label l i The fuzzy mutual information between them; Step 3: When the selected feature subset is not empty, calculate the weighted feature correlation between each remaining candidate feature in the candidate feature set and the label set; The joint conditional information content of the candidate features and the selected feature subset is calculated according to the following formula: relS(f n ,S,l i )=w(l i )·FCMI(f n ;S|l i ) (8) In the formula, relS(f n ,S,l i ) indicates in the label l i Candidate features f under the action n The joint conditional information content with the selected feature subset S, FCMI(f n ;S|l i ) indicates in the label l i Candidate features f under the action n Fuzzy conditional mutual information between the selected feature subset S and the selected feature subset S; The weighted correlation between candidate features and the label set is calculated using the following formula: Step 4: Calculate the amount of redundant information between the remaining candidate features and the selected feature subsets; MS(f n ,S)=FMI(f n ;S)(13) In the formula, MS(f n S) represents the candidate feature f n The amount of redundant information in the selected feature subset, FMI(f n S) represents candidate feature f n The fuzzy mutual information with the selected feature subset; Step 5: Calculate the importance score of each remaining candidate feature in the candidate feature set; In the formula, J(f) n ) represents candidate feature f n Importance score; Add the candidate feature with the highest importance score to the selected feature subset, and remove the candidate feature from the candidate feature set; Step 6: If the number of features in the selected feature subset is less than the set dimension, repeat steps 3 to 5 until the number of features in the selected feature subset equals the set dimension; if the number of features in the selected feature subset equals the set dimension, stop; at this point, feature selection is complete.
2. The feature selection system according to claim 1, characterized in that, The formula for calculating fuzzy mutual information is: In the formula, ∩ represents the minimum value, and |·| represents the sum of the membership degrees of the sample with all other samples based on the candidate features. and They represent samples x respectively u With candidate features f n Tag l i The fuzzy equivalence relation is specifically represented as follows: In the formula, Indicates sample x u and x v Based on candidate features f n membership degree Indicates sample x u and x v Based on label i Membership degree is used to measure fuzzy relations in fuzzy mutual information; Representing candidate features f n Regarding sample x v The fuzzy relationship Indicates label l i Regarding sample x v The fuzzy relationship, f nu f nv Indicates sample x u and x v Candidate features, l iu l iv Indicates sample x u and x v The i-th label represents the fuzzy relation connector, and p is the number of samples.
3. The feature selection system according to claim 1 or 2, characterized in that, The formula for calculating fuzzy conditional mutual information is: In the formula, [x u ] S Indicates sample x u Fuzzy equivalence relation with the selected feature subset, Indicates sample x u and x v Based on the membership degree of the selected feature subset, ||S u -S v || 2 Indicates sample x u With x v The square of the Euclidean distance between sample vectors consisting of all selected feature descriptions in the selected feature subset, where X represents the training set.