A method for image feature dimensionality reduction selection based on feature class distance and machine learning

Through the image feature dimensionality reduction selection method based on inter-feature distance and machine learning, the problem of insufficient data redundancy and generalization capabilities of feature selection methods in image recognition is solved, and efficient feature subset selection and classification recognition performance improvement is achieved.

CN115359283BActive Publication Date: 2025-06-06CHONGQING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210736668.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-06-06
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

The existing feature selection methods have problems such as data redundancy, overfitting, low training efficiency and low generalization ability in image recognition, and insufficient consideration of the correlation between features.

Method used

The image feature dimensionality reduction selection method based on feature class distance and machine learning is adopted. By calculating the evaluation value of features, combining t-test, correlation analysis and machine learning classifiers, a subset of features with high classification performance is selected.

Benefits of technology

It effectively reduces the dimension of the feature data set, reduces the impact of irrelevant features on accuracy, improves the efficiency and accuracy of classification recognition, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359283B_ABST
    Figure CN115359283B_ABST
Patent Text Reader

Abstract

The present invention discloses an image feature dimensionality reduction selection method based on feature inter-class distance and machine learning: 1) dividing a feature data set into a training set, a validation set and a test set according to a ratio; 2) calculating the inter-class distance, inter-feature correlation and t-test p value of the feature data set, obtaining a feature evaluation value, and arranging each feature in descending order according to the evaluation value; 3) arranging all the first N features to obtain different feature combination sets, selecting a classifier to train and learn each feature combination, and obtaining a feature combination with the best classification performance according to the accuracy sorting; 4) outputting a machine learning classification model with the highest accuracy and weight sum, and inputting the test data set into the classification model to obtain the test accuracy of the best feature combination. The present invention uses inter-class distance, correlation and statistical test, combined with machine learning technology to achieve effective dimensionality reduction of high-dimensional feature data sets, and finally selects a feature combination with high classification performance, thereby improving the accuracy and efficiency of image classification and recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, specifically to the field of image feature engineering, and in particular to an image feature dimensionality reduction selection method based on feature class distance and machine learning. Background Art

[0002] Image feature extraction is an important technology for image classification and recognition. Selecting appropriate features as training data sets can effectively improve the accuracy of recognition and classification. If too few or too single features are selected, other important features in the image will be missed, resulting in problems such as low interpretability and low recognition accuracy. However, selecting too many features will make the dimension of the feature data set too high, thus causing problems such as overfitting, low training efficiency, and data redundancy. Regarding the research on feature selection methods, the results after consulting literature and patent searches are as follows.

[0003] 1) The current feature selection methods can be divided into three categories: filtering, encapsulation and embedding. The filtering method mainly relies on the independent discrimination criteria between categories in the feature data set, and performs feature evaluation without considering the interference of other machine learning algorithms, such as using F test, chi-square test and significance test. This method can use statistical tests to quickly exclude the influence of non-critical features, but it cannot be combined with machine learning algorithms to consider the influence of the selected feature subset on the training accuracy, and has certain limitations.

[0004] 2) Feature selection algorithms based on encapsulation and embedding methods are combined with machine learning algorithms. The encapsulation method selects a feature subset from the feature data set and trains it according to the specified feature evaluation criteria. During the training process, it can select a feature subset with excellent performance. The algorithm performance is directly related to the classifier used. Commonly used filtering feature selection algorithms include genetic algorithms, forward search methods, and GA-Fisher methods that combine neural networks with linear discriminant analysis. The performance of the encapsulation method mainly depends on the selected feature evaluation criteria and classifier model, so it has a high algorithm complexity. At the same time, this method relies too much on the selection of the classification model, resulting in low generalization ability, and it is also unable to give the contribution of each feature to the classification.

[0005] 3) The embedding method is similar to the encapsulation method. It can use machine learning algorithms to combine the relationship between features to obtain the weight coefficients of each feature, and then sort them in descending order according to the obtained weight coefficients, and select the features with higher contribution values, i.e., weight coefficients, to the model classification as subsets. Commonly used embedding feature selection algorithms include decision trees, ridge regression, etc., which also have high algorithm complexity, and their performance is related to the selected machine learning algorithm.

[0006] 4) The Chinese invention patent "Unsupervised feature selection method and system based on multi-label learning" (application number: 201911312573.7) provides an unsupervised feature selection method and system based on multi-label learning, constructs an unsupervised feature selection objective function based on multi-label learning for the feature data set, and learns the binary multi-label matrix and the feature selection matrix. In the method, a discrete optimization method based on the augmented Lagrange multiplier method is used to solve the unsupervised feature selection objective function based on multi-label learning to obtain a feature selection matrix; the feature selection matrix is ​​sorted to determine the target features to be selected. At the same time, multi-labels for semantic guidance and feature selection are learned, and binary constraints are imposed in spectral embedding to obtain multi-labels to guide the final feature selection process. This method is aimed at unsupervised multi-label feature selection methods, and its performance is related to the data set. It is not suitable for supervised classification and recognition problems in image recognition.

[0007] 5) The Chinese invention patent "Method for selecting main symptoms of traditional Chinese medicine based on feature groups" (application number: 201710445511.8) provides a method for selecting main symptoms of traditional Chinese medicine based on feature groups. First, the original feature set is screened, and the features with too low frequency are removed, and the screened feature set is clustered using a feature clustering algorithm to obtain the corresponding feature group. A hidden variable is introduced into each feature group to obtain the corresponding hidden class model, and the correlation between the hidden variable and the label is calculated, and the feature groups are sorted from large to small according to the correlation between the hidden variable and the label. The sorted feature groups are added to the selected feature subset in turn, a Bayesian network containing hidden variables is established, and the classification accuracy of the Bayesian network is calculated, and then a curve of the number of added feature groups and the classification accuracy is obtained, and the corresponding optimal feature subset is obtained by judging the convergence of the curve or the highest accuracy. This patent uses a filtering method to first exclude individual features and then cluster to obtain each feature combination. Sort according to the correlation between feature combinations.

[0008] 6) The Chinese invention patent "Method, device, equipment and storage medium for feature selection based on machine learning" (application number: 201910342060.4) provides a feature selection method based on machine learning: presetting multiple reference feature selection models to extract reference feature information from the transaction data; scoring the reference feature selection model according to the selected reference feature information to obtain a model scoring result; selecting a target feature selection model according to the model scoring result, and using the reference feature information selected by the target feature selection model as the target feature information, thereby combining multiple models to select the optimal feature selection model for feature selection. This patent uses multiple machine learning models such as regression models, correlation analysis models, and cluster analysis models to score features, and selects the model with the highest score as output.

[0009] In summary, this patent adopts a supervised learning method, combined with feature class distance, correlation test, t-test and feature selection method of machine learning technology, to provide a new method for fast and effective dimensionality reduction of high-dimensional feature data sets. Summary of the invention

[0010] In the process of image recognition and feature extraction, more and more feature information is mined from the image, which improves the classification performance while also bringing about the problem of data redundancy. Too high a dimension of a data set will lead to low model training efficiency, excessive running memory, and reduced model generalization capabilities. Existing feature selection algorithms have problems such as dependence on clustering algorithms, being greatly influenced by training data sets, and insufficient consideration of the correlation between features. In order to reduce the dimensionality of high-dimensional feature data sets, reduce the impact of irrelevant features on accuracy, select feature subsets with high classification performance, and improve the efficiency and accuracy of classification recognition, this patent provides an image feature dimensionality reduction selection method based on feature class distance and machine learning.

[0011] The invention patent content mainly includes the following steps:

[0012] 1) Establishment of feature dataset

[0013] Create an original feature data set U(m*n) consisting of m samples, each with n features, and divide the original feature data set U into a training set, a validation set, and a test set.

[0014] 2) Calculation of feature evaluation value

[0015] The evaluation value S of each feature is calculated according to formula 1. The evaluation value S is related to the distance d between feature classes, the low correlation threshold w, the low correlation ratio θ, and the p value in the t-test:

[0016]

[0017] Calculate the inter-class distance d, and the calculation method is shown in Formula 2. In the formula, μ 1 and μ 0 are the average values ​​of each feature in the two classification features, σ 1 and σ 0 is the variance of the two categorical features;

[0018]

[0019] By performing a t-test analysis on each dimension feature, the p value can be obtained. In formula 3, μ x is the overall mean of feature x; , s, n are the sample mean, sample standard deviation and sample capacity of feature x respectively;

[0020]

[0021] Calculate the Pearson Correlation Coefficient r between each feature. In the calculation formula 4, Cov(x,y) is the covariance between feature x and feature y; σ x and σ y Represent the standard deviation of feature x and feature y respectively.

[0022]

[0023] A low correlation threshold w is introduced. When the correlation coefficient r of a feature is not greater than w, the feature is considered to be a low correlation feature, and the proportion of low correlation features θ in the feature is calculated.

[0024] For features with a low correlation ratio θ of not less than 0.5 and a t-test p value of not more than 0.05, they have lower correlation and higher confidence and can be considered as key features affecting classification. The features within this value range are calculated according to Formula 1, and the features outside this range are given an evaluation value of 0. The features are arranged in descending order according to S.

[0025] 3) Calculation of feature combination accuracy based on machine learning

[0026] After the first N features are arranged in descending order, all permutations and combinations are performed to obtain 2 N -1 feature combinations with different numbers. According to the data set, a machine learning classifier (such as the commonly used binary SVM classifier) ​​is selected as the classifier model, and the training data set containing different feature combinations is input into the classifier model for training to obtain a machine learning classification model, and each feature combination is arranged in descending order according to the training accuracy.

[0027] 4) Determination of the optimal feature combination

[0028] Select the feature combination with the highest accuracy. If two feature combinations have the same accuracy, select the feature combination with the higher weight. Assign weight coefficients to the first N features in descending order in step (2), where the weight of the feature ranked i is calculated by formula 5:

[0029]

[0030] The test set data is input into the machine learning classification model trained by the feature combination, and finally the best feature combination with high classification performance is determined. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Specific implementation method flow for the case;

[0032] Figure 2 This is an example of a picture dataset;

[0033] Figure 3 ROC curves of LBP and grayscale features under different feature set conditions;

[0034] Figure 4 is the ROC curve of GLCM features under different parameters;

[0035] Figure 5 It is a scatter plot of feature distribution. DETAILED DESCRIPTION

[0036] In order to verify the validity of the present invention, a specific implementation case is given here to explain the present invention. The image dataset used in the embodiment contains 351 pictures with smoke ( Figure 2 (a)) and 49 smoke-free ( Figure 2 (b)) A dataset of 400 images, filtering and graying the original images ( Figure 2 (c)) Processing. The test image needs to be input into the model to determine whether the image is a smoke image, and the accuracy of the test set is used as the recognition accuracy indicator.

[0037] The specific implementation steps are as follows:

[0038] 1) Establishment of feature dataset

[0039] The color features include the maximum, minimum, mean, mode, variance, standard deviation, median, lower quartile, upper quartile, skewness and slope of the image gray value (Gray), and the texture features include local binary pattern (LBP) and gray level co-occurrence matrix (GLCM). A total of 6 eigenvalues ​​of LBP energy and entropy under different parameter combinations and 14 Haralick features and their means in GLCM under different angles and distances, totaling 182 eigenvalues, were extracted.

[0040] The above 199 feature values ​​are extracted for each image in the image dataset, and a 400*199 feature dataset is obtained. The feature dataset is divided into training set, validation set and test set according to the ratio of 0.8 / 0.1 / 0.1.

[0041] 2) Calculation of each feature evaluation value

[0042] The inter-class distance, correlation coefficient and t-test of each feature in the training set are calculated, and the low correlation threshold is set to 0-0.5, so that the low correlation ratio θ and the p-value in the t-test can be obtained. For features with θ not less than 0.5 and p-value not greater than 0.05, the evaluation value S is calculated. For features that are not within the range, their S is set to 0, and each feature is sorted in descending order according to S.

[0043] 3) Calculation of feature combination accuracy based on machine learning

[0044] Selecting the first 10 key features for full permutation and combination, 1023 feature combinations can be obtained. These feature combinations are input into the SVM classifier for training to obtain a machine learning classification model, and each combination is sorted in descending order according to the training accuracy.

[0045] 4) Determination of the optimal feature combination

[0046] We select the feature combination with the highest accuracy. If two feature combinations have the same accuracy, we select the feature combination with the higher sum of weights. Finally, we can obtain the machine learning training model trained by the feature combination. Using the test set data as input to the machine learning training model, we can obtain the best feature combination and test accuracy.

[0047] The specific feature names and accuracy of the feature combination output by this method in this example are shown in Table 1. In order to verify the effectiveness of the method and determine the best feature combination in this case, the classification performance of various types of features is compared and analyzed. If the ROC curve is closer to the upper left and the curve coverage area (Area Under Curve, AUC) is larger, it indicates that the classification performance of the feature is better. First, the classification performance of LBP features and grayscale features under different feature set conditions is analyzed. The ROC curve comparison results are shown in Figure 1. Figure 3 As shown. Secondly, the classification performance of each parameter combination in GLCM features and grayscale features is analyzed, and the ROC curve comparison results are shown as follows Figure 4 (where parameter a is the angle, d is the distance, and the feature mean is the average value of 14 Haralick features in GLCM at 4 angles and 3 distances).

[0048] Table 1 Feature combination information

[0049]

[0050] from Figure 3 It can be seen that the ROC curve of the LBP feature with parameters R=3, P=24 has the largest AUC, and its curve is closest to the upper left distribution. The AUC of the Gray feature in comparison is close to it; while the AUC of the LBP feature with R=1, P=8 is only higher than the AUC of the LBP feature with R=2, P=16. This result comparison shows that the LBP feature and Gray feature under the parameters R=3, P=24 have higher classification performance and improve the results of classification recognition.

[0051] from Figure 4It can be seen from the figure that at four different angles, the AUC of the GLCM feature at distance d = 3 is significantly greater than that at d = 1 and d = 2, and at angle a = 45° ( Figure 4 (b)) and a=90°( Figure 4 The ROC curve in (c) can surround the ROC curve of GLCM average; while the ROC curves of GLCM features at d = 1 and d = 2 are very close at the four angles, with no significant difference, and all have smaller AUCs; a = 135° ( Figure 4 The four ROC curves in (d)) have the same trend, so it is impossible to judge the features with high classification performance from this angle. These comparison results show that the GLCM feature with d=3 has better classification performance than the other two distances, and this feature has an improving effect on the classification and recognition results of the model.

[0052] In order to more intuitively feel the classification performance of each feature in the feature combination, the two-dimensional distribution scatter plot of some features in the feature combination is as follows: Figure 5 As shown in the figure (where (a) LBP entropy when R=3, P=24 and HOM of GLCM when d=3, a=45, (b) LBP entropy when R=3, P=24 and Contrast of GLCM when d=3, a=90, (c) lower quartile of grayscale and Contrast of GLCM when d=3, a=45, (d) LBP energy when R=3, P=24 and Contrast of GLCM when d=3, a=90), the smoke-free classification sample points have obvious clustering in the scatter plot, so these features can make the optimal hyperplane found by SVM in classification recognition training have better classification performance, thereby obtaining a better classifier model. Therefore, it can be considered that this feature combination is the best feature combination for this case data set.

Claims

1. A method for selecting image feature dimensionality reduction based on feature class distance and machine learning. It is characterized in that The following steps are included: 1) Establishment of feature dataset Create an original feature data set U consisting of m samples, each with n features, and divide the original feature data set U into a training set, a validation set, and a test set; 2) Calculation of feature evaluation value The evaluation value S of each feature is calculated according to formula 1. The evaluation value S is related to the distance d between feature classes, the low correlation threshold w, the low correlation ratio θ, and the p value in the t-test: Calculate the inter-class distance d. The calculation method is shown in Formula 2, where μ 1 and μ 0 are the average values ​​of each feature in the two classification features, σ 1 and σ 0 is the variance of the two categorical features; By performing t-test analysis on each dimension feature, we can get the p value. In formula 3, μ x is the overall mean of feature x; s and n are the sample mean, sample standard deviation and sample size of feature x respectively; Calculate the Pearson Correlation Coefficient r between each feature. In the calculation formula 4, Cov(x,y) is the covariance between feature x and feature y; s x and y Represent the standard deviation of feature x and feature y respectively; A low correlation threshold w is introduced. When the correlation coefficient r of a feature is not greater than w, the feature is considered to be a low correlation feature, and the proportion of low correlation features θ in the feature is calculated; For features with a low correlation ratio θ of not less than 0.5 and a t-test p value of not more than 0.05, they have lower correlation and higher confidence, and can be considered as key features affecting classification. The features within the value range are calculated according to Formula 1, and the features outside this range are set to have an evaluation value of 0. The features are arranged in descending order according to S; 3) Calculation of feature combination accuracy based on machine learning After the first N features are arranged in descending order, all permutations and combinations are performed to obtain 2 N -1 feature combinations with different numbers, select a machine learning classifier according to the data set, such as the commonly used binary SVM classifier, as the classifier model, input the above training data sets containing different feature combinations into the classifier model for training, obtain the machine learning classification model, and arrange each feature combination in descending order according to the training accuracy; 4) Determination of the optimal feature combination The feature combination with the highest accuracy is selected as the best feature combination. If two feature combinations have the same accuracy, the feature combination with a higher weight is selected. The weight coefficients are assigned to the first N features in descending order in step (2), where the weight of the feature ranked i is calculated by formula 5: The test set data is input into the machine learning classification model trained by the feature combination, and finally the best feature combination with high classification performance is determined.

2. According to claim 1, a method for selecting image feature dimensionality reduction based on inter-class distance and machine learning, It is characterized in that The evaluation values ​​calculated for each feature in steps 2 and 3 are combined with the inter-class distance, correlation coefficient, confidence test and machine learning method of the feature, fully considering the influence within the feature and between categories, and have strong robustness.

3. According to claim 1, a method for selecting image feature dimensionality reduction based on inter-class distance and machine learning, It is characterized in that The feature data set in step 1 includes a label column, which is a supervised feature dimensionality reduction method guided by labels. The invented method can effectively select feature combinations with high classification performance, thereby improving the efficiency of classifier model training and the accuracy of classification recognition.

Citation Information

Patent Citations

  • A TCM syndrome selection method based on feature groups

    CN107292097B

  • Feature selection method, device and equipment based on machine learning and storage medium

    CN110276369A

  • An Unsupervised Feature Selection Method and System Based on Multi-Label Learning

    CN111027636B