Characteristic selection method based on variance analysis and hierarchical clustering
Through feature selection methods based on variance analysis and hierarchical clustering, important features are screened out, and the problems of complex feature dimensions and poor model fitting effects in the existing technology are solved, and feature selection is simplified and model performance is improved.
Patent Information
- Application Number
- CN202510453607.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
In specific application scenarios, existing feature selection algorithms have problems such as complex feature dimensions and poor model fitting effects. In particular, MRMD lacks direct analysis of positive and negative sample differences and ANOVA does not consider the similarity between features.
A feature selection method based on variance analysis and hierarchical clustering is adopted. By calculating the contribution F value of each feature dimension and the hierarchical clustering algorithm, important features are selected to reduce the redundancy of feature space, and combined with support vector machines, random forests or XGboost as prediction models, feature selection and model optimization are performed.
It effectively reduces the dimension of feature descriptors, reduces the operation process of feature selection, improves the fitting speed and accuracy of the model, and is suitable for a variety of feature descriptors, taking into account the similarity between features and the contribution to the distinction between positive and negative samples.
Smart Images

Figure CN120336795A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data clustering analysis, and particularly relates to a feature selection method based on variance analysis and hierarchical clustering. Background Art
[0002] Existing feature selection algorithms rely on distance formulas or statistical methods to screen out features that are beneficial to sample discrimination. Maximum relevance maximum distance (MRMD) measures the contribution degree of features by calculating the distance between features and sample labels; Analysis of variance (ANOVA) measures the contribution degree of features by calculating the ratio of features between groups and within groups through variance analysis. The above methods can effectively reduce the feature dimension in specific application scenarios, but MRMD lacks a direct analysis of the differences between positive and negative samples, ANOVA does not consider the similarity between features, and it is necessary to model or determine the retained feature dimension through thresholds, resulting in problems such as the complication of the feature selection process and poor model fitting effect.
[0003] In order to solve the deficiencies of the existing technology, people have conducted long-term explorations and proposed various solutions. For example, a Chinese patent document discloses an improved K-means clustering algorithm for determining the K value using variance analysis [201610708116.X], which determines the clustering hierarchy division and data summary; selects the clustering center and initializes the K value; then finds the classes with the number of internal members greater than 1, and conducts variance analysis respectively to test whether there is significance between the clustering members of each class; and conducts clustering analysis and variance test; finally determines the number of clusters and the clustering members of each class; if the significance level test of variance analysis is passed between the internal members of all classes, then determines the number of clusters and the clustering members of each class.
[0004] The above solution solves the problem of clustering analysis to a certain extent, but there are still many deficiencies in this solution, such as the problem of a relatively high dimension of feature descriptors. Summary of the Invention
[0005] The purpose of the present invention is to provide a feature selection method based on variance analysis and hierarchical clustering with reasonable design and effective reduction of the dimension of feature descriptors for the above problems.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A feature selection method based on variance analysis and hierarchical clustering includes the following steps: S1: Construct a standard data set; S2: Feature descriptor; S3: Feature selection; S4: Construct a prediction model; S5: Train a classification model; S6: Evaluate the model performance; S7: Practical application.
[0007] In the above feature selection method based on analysis of variance and hierarchical clustering, step S1 includes the following steps: S11: Download DNA sequences with existing annotation information from a public database; S12: Delete sequences containing ambiguous bases in positive and negative samples; S13: Use the CD-hit software to reduce the sequence identity of positive samples to 80% and that of negative samples to 60%, and perform random sampling to balance the number of positive and negative samples. Subsequently, save the processed dataset as a standard dataset in FASTA format; S14: Randomly divide the standard dataset, with 60% of the samples as the training set, 20% as the validation set, and the remaining 20% as the test set, and ensure that the number of positive and negative samples in the training set, validation set, and test set is balanced.
[0008] In the above feature selection method based on analysis of variance and hierarchical clustering, step S2 includes the following steps: S21: Split the obtained sample sequence s into dinucleotide form; S22: Analyze the distribution characteristics of the physicochemical properties of dinucleotides.
[0009] In the above feature selection method based on analysis of variance and hierarchical clustering, step S3 includes the following steps: S31: Use analysis of variance and hierarchical clustering algorithms to delete feature descriptors; S32: Calculate the between-group and within-group differences for each feature dimension, and evaluate the contribution of the feature to sample division through the ratio of between-group difference to within-group difference; S33: Cluster similar features into one cluster through hierarchical clustering; S34: Loop through step S33 and continuously iterate to merge the clusters with the closest distance until all features are clustered into one whole from top to bottom and no changes occur in all clusters and all features including internal child nodes. Then stop the iteration and save the hierarchical clustering result from bottom to top; S35: Traverse the bottom-level child nodes of the clustering result in S34 and delete features with lower F values within the cluster; if in the bottom-level nodes, a single feature forms an independent cluster, directly retain the feature of this dimension without feature screening.
[0010] In the above feature selection method based on analysis of variance and hierarchical clustering, step S4 selects a machine learning classification algorithm according to the distribution characteristics of the filtered feature subset.
[0011] In the above feature selection method based on analysis of variance and hierarchical clustering, step S4 uses a support vector machine, random forest, or XGboost as the prediction model.
[0012] In the above feature selection method based on analysis of variance and hierarchical clustering, step S5 includes the following steps: S51: Use 60% of all samples as the training set, 20% as the validation set, and 20% as the test set; S52: Optimize the model parameters.
[0013] In the above feature selection method based on analysis of variance and hierarchical clustering, step S6 includes the following steps: S61: Evaluate the performance of the classification model; S62: Apply the prediction model obtained in S5 to the training set and test set. Both the training set and test set use the dimensional features corresponding to the training set as the final input, and then use the evaluation parameters in step S61 to evaluate and verify the performance of the model on the validation set and test set.
[0014] In the above feature selection method based on analysis of variance and hierarchical clustering, step S6 uses four evaluation parameters to measure the performance of the model for binary classification problems, namely sensitivity, specificity, accuracy, and Matthews correlation coefficient.
[0015] In the above feature selection method based on analysis of variance and hierarchical clustering, step S7 constructs promoter and non-promoter sequence samples, and performs feature descriptor conversion and feature selection.
[0016] Compared with the existing technology, the advantages of the present invention are as follows: By using a feature selection algorithm based on analysis of variance and hierarchical clustering, it can effectively reduce the dimension of feature descriptors and reduce the redundancy of the feature space; there is no need to artificially set thresholds, and there is no need to build a model during the feature selection process, reducing the operation process of feature selection; it takes into account the similarity between features and the contribution of features to distinguishing positive and negative samples, and can more efficiently delete redundant features in the feature descriptors; it is applicable to various feature description features and can effectively reduce the feature dimension. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flow diagram of the present invention; Figure 2 is a distribution diagram of the physicochemical properties of dinucleotides in the feature descriptors provided by the embodiment of the present invention; Figure 3 is a schematic diagram of the feature selection algorithm provided by the embodiment of the present invention; Figure 4 is the feature selection process of the practical application provided by the embodiment of the present invention. Detailed implementation manners
[0018] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0019] As Figure 1 shown, a feature selection method based on variance analysis and hierarchical clustering combines variance analysis and a bottom-up hierarchical clustering algorithm. By comparing the contribution degrees of features within the first-layer clusters and screening important features, the dimension of the feature space is reduced. First, the F value of the contribution degree of each feature dimension is calculated using ANOVA. Then, the features are clustered by the hierarchical clustering algorithm. Next, the features to be retained are determined according to the F values of the features within each cluster, realizing the reduction of the feature dimension. The contribution degrees of features to positive and negative samples are measured by ANOVA, and the similarity of features is measured by hierarchical clustering, taking into account the advantages of multiple feature selection algorithms, reducing the redundancy of the original feature space, and improving the fitting speed of traditional models. The method includes the following steps: S1: Construct a standard data set, collect DNA sequences with existing annotation information. Taking promoters as an example, the DNA sequences are divided into two categories, namely promoters and non-promoters; the CD-hit software is used to reduce the consistency of the positive and negative sample sequences to less than 80%; the positive and negative samples are divided into a training set, a validation set, and a test set according to the ratio of 6:2:2 and stored in the FASTA format; S2: Feature descriptors, select suitable feature descriptors according to the inherent biological characteristics of the samples. For example, convert the "A", "T", "C", "G" in the DNA sequence into one-hot encoding; for promoter sequences, the physical and chemical properties can be used to represent the bases in the sequence; S3: Feature selection, perform feature selection on the feature descriptors, delete redundant information in the feature space, calculate the F value of each feature dimension, perform hierarchical clustering on all features, and delete the features with lower F values within the clusters; S4: Construct a prediction model, select a suitable machine learning classification algorithm according to the distribution characteristics of the screened feature subset; if the features are continuous and the distribution differences are small, use a support vector machine as the prediction model; if the features are discrete or the feature distributions have large differences, use a random forest or XGboost as the prediction model; S5: Train the classification model, optimize the model using the grid search method on the training set; use 60% of the sample sequences as the training set and use 10-fold cross-validation to determine the best candidate parameters; 20% of the sample sequences are used to evaluate the model performance; finally, the remaining 20% of the samples are used to verify the robustness of the model; the performance evaluation of the test set and the validation set ensures the stability and accuracy of the model; S6: Model performance evaluation. For binary classification problems, four evaluation parameters are used to measure the model performance, namely sensitivity, specificity, accuracy, and Matthews correlation coefficient; S7: Practical application. Apply the feature selection algorithm to specific biological problems, such as the identification of promoters, enhancers, etc. While reducing the feature dimension, improve the prediction performance of the model, so as to achieve the rapid identification of regulatory elements in DNA sequences.
[0020] Specifically, step S1 includes the following steps: S11: Download DNA sequences with existing annotation information from public databases; Taking promoters as an example, the DNA sequence of 80bp upstream and 20bp downstream of the gene transcription start site, a total of 81bp, is used as the positive sample; Non-promoter sequences are regions such as gene coding regions, non-gene coding regions, introns, exons, and alternative splicing, and the sequence fragment length is 81bp; S12: Delete the sequences containing ambiguous bases in the positive and negative samples; Such as "N", "X", "B", etc., so that the bases in the samples only contain four bases: "A", "T", "C", "G"; S13: Use the CD-hit software to reduce the sequence identity of the positive samples to 80% and the sequence identity of the negative samples to 60%, and perform random sampling to balance the number of positive and negative samples. Subsequently, save the processed dataset as a standard dataset in FASTA format; S14: Randomly divide the standard dataset. 60% of the samples are used as the training set, 20% of the samples are used as the validation set, and the remaining 20% are used as the test set, and the number of positive and negative samples in the training set, validation set, and test set is balanced.
[0021] Specifically, step S2 includes the following steps: S21: Split the obtained sample sequence s into dinucleotide forms; The 81bp sequence contains a total of 80 dinucleotide fragments, which are defined as follows: Four bases can be combined in pairs to form 16 permutations and combinations, a i Represents the arrangement of dinucleotides in the sequence, which is one of the 16 dinucleotides.
[0022] S22: Analyze the distribution characteristics of the physicochemical properties of 90 dinucleotides; As Figure 2 shown, the five parameters of the maximum value (max), minimum value (min), average value (mean), variance (var), and sum (sum) of the 16 dinucleotides have obvious differences in the distribution of the 16 dinucleotides. Therefore, the above five statistical parameters are selected to characterize the dinucleotides in the sequence, which are defined as follows: The DNA sequence with a length of 81 bp is converted into a feature vector of (81 - 1) × 5 = 400 dimensions.
[0023] In addition, step S3 includes the following steps: S31: To reduce the similarity of the feature descriptors in step S2, the feature descriptors are pruned using the analysis of variance and hierarchical exemplar algorithm; First, the analysis of variance algorithm (ANOVA) is introduced to measure the contribution of each dimension feature to distinguishing positive and negative samples, and the definition is as follows: S2 B used to evaluate the difference of the current feature between positive and negative samples, S2W while is to evaluate the difference within the overall samples of the current feature. The larger the F value, the smaller the within-group difference and the larger the between-group difference, indicating that the current feature dimension makes a greater contribution to distinguishing positive and negative samples; S32: Calculate the between-group and within-group differences of each feature dimension; First, calculate the mean of all samples. For the between-group difference, the squared difference between the means within positive and negative samples and the total sample mean needs to be calculated and divided by the degrees of freedom; while the within-group difference is calculated by taking the squared difference between this feature dimension and the overall sample mean, and then divided by the degrees of freedom. And the contribution of the feature to sample partitioning is evaluated through the ratio of the between-group difference to the within-group difference; The calculation methods of S2B and S2W are as follows: f(j) represents the feature value of each sample under the current feature dimension, M is the total number of samples, K is the sample category. For a binary classification problem, K the value of is 2.
[0024] S33: Since the F value of each dimension feature obtained in S32 depends on a specific threshold or modeling to delete features with low F values, the choice of the threshold causes uncertainty in the final feature space, and modeling increases the computational cost of feature selection. Therefore, similar features are clustered into a cluster through hierarchical clustering, and features with low F values within the cluster are deleted; The bottom-up hierarchical clustering algorithm is a clustering algorithm that constructs a hierarchical structure by continuously merging sub-nodes; First, each dimension feature is regarded as an independent cluster, and then the Euclidean distance or Manhattan distance between all clusters is calculated; Subsequently, the two closest clusters are merged into one cluster. The features in the merged cluster have high similarity, so these two features represent the new cluster, which also means that each feature in the cluster can represent the cluster; When two clusters are merged, the distances between the newly formed cluster and all other clusters are recalculated.
[0025] S34: Repeat step S33 and iterate continuously to merge the clusters with the closest distance until all features are aggregated into one whole from top to bottom and no changes occur to all clusters and the features shown within the internal child nodes. Then, stop the iteration and save the hierarchical clustering result from bottom to top. S35: Traverse the bottom-level child nodes of the S34 clustering result and delete the features with lower F values within the clusters. For example, if F2 and F3 are clustered into one cluster and the value of F3 is less than that of F2, the feature dimension corresponding to F3 is deleted. If a single feature forms an independent cluster among the bottom-level nodes, the dimension feature is directly retained without feature screening. This feature selection algorithm avoids the selection of thresholds and comprehensively considers the similarity between features and the ability of features to distinguish between positive and negative samples.
[0026] Meanwhile, as Figure 3 shown, step S4 selects a machine learning classification algorithm based on the distribution characteristics of the filtered feature subset. Step S4 uses support vector machine, random forest, or XGboost as the prediction model. Given that the used feature descriptors are composed of statistical parameters such as maximum value, minimum value, and mean, and there are significant differences in the parameter distributions, the XGboost algorithm is good at dealing with classification problems with larger feature information entropy. Therefore, the prediction model is built using XGboost in Python.
[0027] Obviously, step S5 includes the following steps: S51: Use 60% of all samples as the training set, 20% as the validation set, and 20% as the test set. S52: Optimize the model parameters; use the grid search method to optimize the parameters of XGBoost such as "max_depth", "min_child_weight", "gamma", etc. Use the training set as the input and the accuracy of ten-fold cross-validation as the evaluation result for model optimization.
[0028] Preferably, step S6 includes the following steps: S61: Evaluate the performance of the classification model; sensitivity ( Sn ) is used to evaluate the model's ability to identify true positive samples; specificity ( Sp ) is used to evaluate the model's ability to identify true negative samples; accuracy ( Acc ) and Matthews correlation coefficient ( MCC ) are used to measure the model's prediction ability for the overall samples.
[0029] TP and TN represent true negative samples, and FP and FN represent the mispredicted positive and negative samples, respectively.
[0030] S62: Apply the prediction model obtained in S5 to the training set and the test set. Both the training set and the test set use the dimensional features corresponding to the training set as the final input, and then use the evaluation parameters in step S61 to evaluate and verify the performance of the model for the validation set and the test set.
[0031] Obviously, in step S6, four evaluation parameters are used to measure the performance of the model for binary classification problems, namely sensitivity, specificity, accuracy, and Matthews correlation coefficient.
[0032] Meanwhile, in step S7, promoter and non-promoter sequence samples are constructed, and feature descriptor conversion and feature selection are performed. As Figure 4 shown, the 400-dimensional features are reduced to 185 dimensions. The features connected by dotted lines indicate that they are clustered into one cluster, and the features with lower F values within the cluster in blue are deleted. In the validation set and the independent set, after feature selection, the model accuracy has improved as above.
[0033] In summary, the principle of this embodiment is as follows: Combining analysis of variance and bottom-up hierarchical clustering algorithm, by comparing the contribution degrees of features within the first-layer clusters and screening important features, the dimensionality of the feature space is reduced. First, use ANOVA to calculate the contribution degree F value of each feature dimension, then cluster the features through the hierarchical clustering algorithm, and then determine the retained features according to the F values of the features within each cluster, realizing the reduction of the feature dimension.
[0034] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
[0035] Although terms such as analysis of variance and hierarchical clustering are used more frequently in this article, the possibility of using other terms is not excluded. Using these terms is only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitation is contrary to the spirit of the present invention.
Claims
1. A feature selection method based on analysis of variance and hierarchical clustering, characterized in that It includes the following steps: S1: Construct a standard data set; S2: Feature descriptors; S3: Feature selection; S4: Construct a prediction model; S5: Train a classification model; S6: Model performance evaluation; S7: Practical application.
2. The feature selection method based on variance analysis and hierarchical clustering according to claim 1, wherein The step S1 includes the following steps: S11: Download DNA sequences with existing annotation information from a public database; S12: Delete sequences containing ambiguous bases in positive and negative samples; S13: Use the CD-hit software to reduce the sequence identity of positive samples to 80% and that of negative samples to 60%, and perform random sampling to balance the number of positive and negative samples. Subsequently, save the processed data set as a standard data set in FASTA format; S14: Randomly divide the standard data set. 60% of the samples are used as the training set, 20% as the validation set, and the remaining 20% as the test set, and the number of positive and negative samples in the training set, validation set, and test set is balanced.
3. A feature selection method based on analysis of variance and hierarchical clustering according to claim 1, characterized in that The step S2 includes the following steps: S21: Divide the obtained sample sequence s into dinucleotide forms; S22: Analyze the distribution characteristics of the physicochemical properties of dinucleotides.
4. A feature selection method based on analysis of variance and hierarchical clustering according to claim 1, characterized in that The step S3 includes the following steps: S31: Use analysis of variance and hierarchical clustering algorithms to delete feature descriptors; S32: Calculate the between-group and within-group differences of each feature dimension, and evaluate the contribution of features to sample division through the ratio of between-group differences to within-group differences; S33: Cluster similar features into a cluster through hierarchical clustering; S34: Loop through step S33, iterate continuously, merge the clusters with the closest distance, so that all features are clustered into a whole from top to bottom, and when all clusters and all features including internal child nodes no longer change, stop the iteration and save the hierarchical clustering result from bottom to top; S35: Traverse the bottom-level child nodes of the clustering result in S34, and delete features with lower F values within the cluster; if in the bottom-level nodes, a single feature forms an independent cluster, then directly retain the above feature dimension without feature screening.
5. A feature selection method based on variance analysis and hierarchical clustering according to claim 1, characterized in that The step S4 selects a machine learning classification algorithm according to the distribution characteristics of the filtered feature subset.
6. The feature selection method based on variance analysis and hierarchical clustering according to claim 5, wherein, The step S4 uses a support vector machine, random forest, or XGboost as a prediction model.
7. A feature selection method based on variance analysis and hierarchical clustering according to claim 1, characterized in that, The step S5 includes the following steps: S51: Use 60% of all samples as the training set, 20% as the validation set, and 20% as the test set; S52: Optimize model parameters.
8. A feature selection method based on analysis of variance and hierarchical clustering according to claim 1, characterized in that, The step S6 includes the following steps: S61: Evaluate the performance of the classification model; S62: Apply the prediction model obtained in S5 to the training set and test set. Both the training set and test set use the dimension features corresponding to the training set as the final input, and then use the evaluation parameters in step S61 to evaluate and verify the performance of the model on the validation set and test set.
9. A feature selection method based on analysis of variance and hierarchical clustering according to claim 7, characterized in that For binary classification problems, the step S6 uses four evaluation parameters to measure the performance of the model, namely sensitivity, specificity, accuracy, and Matthews correlation coefficient.
10. A feature selection method based on variance analysis and hierarchical clustering according to claim 1, characterized in that The step S7 constructs promoter and non-promoter sequence samples, and performs feature descriptor conversion and feature selection.
Citation Information
Patent Citations
Improved K-means clustering algorithm capable of determining value of K by using variance analysis
CN106384119A
An Android Malware Family Clustering Method Based on SinglePass Algorithm
CN109145605A
Network flow classification method based on SOM and K-means fusion algorithm
CN111211994A
Fault feature parameter selection method based on fuzzy preference relationship and adaptive hierarchical clustering
CN111898705A
High-dimensional data feature selection method based on improved L1 regularization and clustering
CN113177604A