A Feature Selection Method Based on Multiclass Logistic Regression

Through the feature selection method based on multi-classification logistic regression, the weighted deviation of the feature regression coefficient is maximized and the regularization term is introduced, which solves the problem that the feature selection method in the prior art is not sensitive to noise and has poor interpretation, and achieves stronger feature representation ability and higher feature selection accuracy.

CN114821210BActive Publication Date: 2025-06-03NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210265435.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-06-03
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

The embedded feature selection method based on least squares regression in the prior art is noise sensitive and not very interpretable, and it is difficult to consider the importance of features from a global perspective.

Method used

The feature selection method based on multi-classification logistic regression is adopted. By constructing the data matrix, label vector and regression coefficient matrix, the weighted deviation of the feature regression coefficients of all categories is maximized, and the L2, p-norm regularization term is introduced. The gradient descent method is used to iterate the regression deviation and regression coefficients, and the index of the feature is finally extracted.

Benefits of technology

The global representation ability of selected features is improved, the data discrimination after regression is enhanced, overfitting is avoided, and the accuracy of feature selection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821210B_ABST
    Figure CN114821210B_ABST
Patent Text Reader

Abstract

The present invention discloses a feature selection method based on multi-class logistic regression. First, a data matrix, a label vector, and a regression coefficient matrix are constructed. Next, a feature selection model based on multi-class logistic regression is constructed. The gradient descent method is used to obtain the optimal solution of the feature selection model, and the regression bias and regression coefficients are iteratively updated. Finally, the indices of the extracted features are obtained. The operation process of the present invention is simple and easy to understand, and the objective function is differentiable at any order and can be applied to a variety of numerical calculation methods, which can improve the global representation ability of the selected features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to a feature selection method. Background Art

[0002] Feature selection technology is an important research topic in the field of pattern recognition, aiming to select an optimal feature subset from the original data features. In classification problems, its goal is to obtain better classification performance with as few features as possible. Feature selection can not only effectively reduce data redundancy, avoid accuracy loss and waste of computing resources caused by high dimensionality, but also retain the original physical meaning of the processed data features, so as to achieve the feature selection task at the data acquisition end. At present, feature selection has been widely applied in fields such as computer vision, medicine, and remote sensing image processing.

[0003] Zhu Xingyu, Chen Xiuhong (in "Unsupervised Feature Selection Combining Uncorrelated Regression and Nonnegative Spectral Analysis", CAAI Transactions on Intelligent Systems: 1-11 [December 20, 2021]. http: / / kns.cnki.net / kcms / detail / 23.1538.TP.20211015.0035.002.html.) conducted embedded feature selection based on least squares regression combined with spectral analysis, and adopted an uncorrelated constraint to avoid trivial solutions. However, least squares regression is sensitive to noise and it is not easy to obtain the optimal classification surface, and the interpretability of the regression coefficients is not strong. Logistic regression has better robustness to sample noise, and its regression coefficients for features have strong probability interpretations. In addition, from the perspective of model solving, logistic regression can more easily integrate more training data into the model quickly through online gradient descent, and the model is differentiable to any order, and various numerical solution algorithms can be used. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the present invention provides a feature selection method based on multi-class logistic regression. First, a data matrix, a label vector, and a regression coefficient matrix are constructed; next, a feature selection model based on multi-class logistic regression is constructed; the gradient descent method is used to obtain the optimal solution of the feature selection model, and the regression bias and regression coefficients are iteratively updated; finally, the indices of the extracted features are obtained. The operation process of the present invention is simple and easy to understand, and the objective function is differentiable to any order and can be applied to a variety of numerical calculation methods, which can improve the global representation ability of the selected features.

[0005] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0006] Step 1: Construct a data matrix, a label vector, and a regression coefficient matrix;

[0007] Obtain n samples with a feature dimension of d, and construct a data matrix x i Each element value of is the eigenvalue of the sample, and the class label vector of n samples is where y i = 1, 2,..., c represents the class of the i-th sample, and c is the total number of sample classes;

[0008] Regression coefficient matrix The element w in the j-th row and k-th column of represents the regression coefficient of the j-th feature to the k-th class of samples; Regression deviation vector jk The k-th element b of represents the regression deviation of the k-th class of samples; The k-th element b of k represents the regression deviation of the k-th class of samples;

[0009] Step 2: Construct a feature selection model based on multi-class logistic regression;

[0010] Maximize the weighted deviation of the feature regression coefficients for all classes:

[0011]

[0012] This problem is equivalent to the following minimization problem:

[0013]

[0014] Then impose an L 2,p norm regularization term on the model to improve the coefficient degree of the regression coefficient matrix;

[0015] Finally, the objective function of the feature selection model is constructed as follows:

[0016]

[0017] In the formula, l represents another sample class different from y i ; represents the feature regression coefficient vector of the y i -th class of samples, that is, the y i -th column vector of the regression coefficient matrix W; w l represents the feature regression coefficient vector of the l-th class of samples, that is, the l-th column vector of the regression coefficient matrix W; b l respectively represent the regression deviations of the y i -th and l-th classes of samples, p is a constant, and γ is a fixed value;

[0018] Step 3: Solve the feature selection model;

[0019] Use the gradient descent method to find the optimal solution, and iteratively update the regression deviation and regression coefficient according to the following steps:

[0020]

[0021]

[0022] Among them, b(t + 1) and b(t) represent the regression deviation vectors at the (t + 1)-th step and the t-th step respectively, W(t + 1) and W(t) represent the regression coefficient matrices at the (t + 1)-th step and the t-th step respectively, and η is the update step size;

[0023] Specifically:

[0024] ① Let t = 0, randomly initialize W(0) and b(0), and set the convergence precision ε;

[0025] ② Update the regression deviation vector b element by element:

[0026]

[0027] Among them:

[0028]

[0029]

[0030] Among them, X (k) represents the set of samples of the k-th class, n k represents the total number of samples of the k-th class, X (q) represents the set of samples of the q-th class, n q represents the total number of samples of the q-th class;

[0031] ③ Update the regression coefficient matrix W element by element:

[0032]

[0033] Among them

[0034]

[0035]

[0036] Among them, x ji represents the element at the i-th column and the j-th row of the sample matrix X, w j represents the j-th row vector of the regression coefficient matrix W;

[0037] ④ Calculate the difference between the objective function values of the last two iterations: e = J(W(t), b(t)) - J(W(t + 1), b(t + 1)). If e < ε, then end. Otherwise, let t increment by 1 and return to step ② until e < ε;

[0038] Step 4: Extract the indices of the features;

[0039] Calculate ||wj || 2 , where \(j = 1, 2, \ldots, d\), and then select the indices of the \(m\) maximum values as the selected features.

[0040] Preferably, the value range of \(p\) is \(0 \lt p \lt 2\).

[0041] The beneficial effects of the present invention are as follows:

[0042] (1) The present invention establishes a feature selection method model based on multi-class logistic regression, and improves the global representation ability of the selected features by maximizing the deviation of the regression coefficients of all class samples.

[0043] (2) The present invention introduces the \(L\) 2,p norm as a regularization term, which can not only effectively avoid overfitting but also control the sparsity degree of the regression coefficient matrix to improve the accuracy of feature selection.

[0044] (3) The feature selection method based on multi-class logistic regression of the present invention has a simple operation process and is easy to understand, and the objective function is differentiable at any order and can be applied to a variety of numerical calculation methods. Description of the Drawings

[0045] Figure 1 is the flowchart of the method of the present invention.

[0046] Figure 2 is the grayscale image of the actual hyperspectral image scene of the embodiment of the present invention.

[0047] Figure 3 is the result graph of the ground object classification accuracy of the embodiment of the present invention. Detailed Embodiments

[0048] The present invention will be further described below in conjunction with the drawings and embodiments.

[0049] The technical problem solved by the present invention is: aiming at the problem that the interpretability of the existing embedded feature selection method based on least squares regression is not strong and it is sensitive to noise, the present invention proposes a feature selection method based on multi-class logistic regression. The regression coefficients of multi-class logistic regression can highlight the importance of each feature for the discrimination of different class samples. The goal is to make the deviation of the feature regression coefficients of different classes as large as possible, which can maximize the global representation ability of all features and make the discriminability of the regression data stronger. Therefore, the present invention maximizes the deviation of the feature regression coefficients of all classes, and can better consider the importance of discriminating a single feature from a global perspective. Therefore, the present invention can better realize feature selection, thereby reducing the difficulty of data storage and improving the data processing speed.

[0050] A feature selection method based on multi-class logistic regression includes the following steps:

[0051] Step 1: Construct a data matrix, a label vector, and a regression coefficient matrix;

[0052] Obtain n samples with a feature dimension of d and construct a data matrix x i Each element value of is the eigenvalue of the sample, and the class label vector of the n samples is where y i = 1, 2,..., c represents the class of the i-th sample, and c is the total number of sample classes;

[0053] Regression coefficient matrix The element w in the j-th row and k-th column of represents the regression coefficient of the j-th feature to the k-th class of samples; the regression bias vector jk The k-th element b of represents the regression bias of the k-th class of samples; The k-th element b of k represents the regression bias of the k-th class of samples;

[0054] Step 2: Construct a feature selection model based on multi-class logistic regression;

[0055] Maximize the weighted deviation of the feature regression coefficients for all classes:

[0056]

[0057] This problem is equivalent to the following minimization problem:

[0058]

[0059] Then impose an L 2,p norm regularization term on the model to improve the coefficient degree of the regression coefficient matrix and make the model more conducive to the feature selection task. Finally, the objective function of the feature selection model is constructed as follows:

[0060]

[0061] In the formula, l represents another sample class different from y i , b l respectively represent the regression biases of the y i -th and l-th class of samples, p is a constant, 0 < p < 2, and γ is a fixed value;

[0062] Step 3: Solve the feature selection model;

[0063] Use the gradient descent method to find the optimal solution and iteratively update the regression bias and regression coefficient according to the following steps:

[0064]

[0065]

[0066] Among them, b(t + 1) and b(t) respectively represent the regression deviation vectors at the (t + 1)-th step and the t-th step, W(t + 1) and W(t) respectively represent the regression coefficient matrices at the (t + 1)-th step and the t-th step, and η is the update step size;

[0067] Specifically:

[0068] ① Let t = 0, randomly initialize W(0) and b(0), and set the convergence accuracy ε;

[0069] ② Update the regression deviation vector b element by element:

[0070]

[0071] Among them:

[0072]

[0073]

[0074] ③ Update the regression coefficient matrix W element by element:

[0075]

[0076] Among them

[0077]

[0078]

[0079] ④ Calculate the difference between the objective function values of the last two iterations: e = J(W(t), b(t)) - J(W(t + 1), b(t + 1)). If e < ε, then end. Otherwise, let t increase by 1 and return to step ② until e < ε;

[0080] Step 4: Extract the indices of the features;

[0081] Calculate ||w j || 2 , j = 1, 2, …, d, and then select the indices of the m maximum values as the selected features. Specific embodiment:

[0083] The basic process of the multi-class logistic regression feature selection method based on the present invention is as Figure 1 shown. The following describes the specific implementation manner of the present invention in combination with the example of ground object classification of hyperspectral images in an actual scenario, but the technical content of the present invention is not limited to the described scope.

[0084] The present invention provides a hyperspectral image ground object classification method based on a feature selection method of multi-class logistic regression, comprising the following steps:

[0085] 1. Obtain a group of hyperspectral images with a feature dimension of d (i.e., the total number of hyperspectral bands is d). For example, Figure 2 , in the actual ground object dataset used, the feature dimension d is 103. The feature value is the gray value of each pixel corresponding to each band. The total number of pixels in a single band is n = 10370, and the ground object class labels of all pixels are obtained, with a total of 10 classes. Then, a data matrix, a label vector, a regression coefficient matrix, and a regression bias vector are constructed.

[0086] For a group of hyperspectral images with a feature dimension of d (the feature value is the gray value after graying a single band), the total number of pixels in a single band is n. Normalize the pixels corresponding to all bands according to the feature, that is, the value of each pixel is equal to the value divided by the sum of the squares of all pixel values in a certain band. All features of the i-th pixel are expressed as where i = 1, 2,..., n, and the j-th element of x i represents the value of the j-th feature of the i-th pixel. denotes the label vector of all data, where y i = 1, 2,..., c, and c is the total number of pixel ground object classes. The regression coefficient matrix The element w jk in its j-th row and k-th column represents the regression coefficient of the j-th feature to the k-th class of samples. The regression bias vector The k-th element b k of it represents the regression bias of the model to the k-th class.

[0087] 2. Establish an optimization problem, solve the regression coefficient matrix and the regression bias vector, and obtain the index of the finally selected features. It is mainly divided into the following three processes:

[0088] (1) Construct a feature selection model based on multi-class logistic regression

[0089] Maximize the weighted deviation of the feature regression coefficients for all classes. The objective function is as follows:

[0090]

[0091] The objective function includes multi-class logistic regression and an L 2,p -norm regularization term, and p = 1, γ = 1 can be set.

[0092] (2) Solve the feature selection model

[0093] The optimization problem to be solved has no constraints. When p = 2, the global optimal solution is directly obtained by using the gradient descent method. The regression coefficients are iteratively updated according to the following steps:

[0094]

[0095]

[0096] where η is the update step size set to 0.00001. In addition, the maximum number of iterations of the algorithm is set to 50 and the minimum number of iterations is set to 10. The regression coefficient matrix and the regression bias vector are specifically updated according to the following steps:

[0097] ① Randomly initialize W(0) and b(0), let t = 0, η = 0.00001, and set the convergence accuracy ε = 10 -3 ;

[0098] ② Update the regression bias vector b element by element:

[0099]

[0100] where:

[0101]

[0102]

[0103] ③ Update the regression coefficient matrix W element by element:

[0104]

[0105] where

[0106]

[0107]

[0108] ④ Calculate the difference e = J(W(t), b(t)) - J(W(t + 1), b(t + 1)) of the objective function in the last two iterations. If e < ε or the maximum number of iterations exceeds 50, then jump out of the loop. Otherwise, let t = t + 1 and return to step ② until e < ε and the number of iterations is equal to 20.

[0109] (3) Index of the extracted features

[0110] Calculate ||w j || 2 , (j = 1, 2,..., d), and then select the indices of the m maximum values (i.e., the serial numbers of the hyperspectral bands) as the features (selected hyperspectral bands) of the selected hyperspectral image.

[0111] 3. Classify all the hyperspectral image pixels with unknown labels, that is, all the samples that make up the sample matrix, a total of 10,370 pixels with 103 dimensions, mainly divided into the following two processes:

[0112] (1) Use the feature indices obtained in step 2 to select the gray values of the corresponding bands of all pixels to form a new data matrix. Each column represents a set of values of the selected features of a hyperspectral image pixel with an unknown label. The total number of new features is m.

[0113] (2) Classify each column of Z as all the feature sequences of the pixel samples corresponding to the new ground objects. Classify the samples with known labels in the projected new pixel samples using classification algorithms (such as K-nearest neighbor classifier, support vector machine, etc.).

[0114] Figure 3 It is the average classification accuracy when assuming that 20% of the samples with known labels are used to train the K-nearest neighbor classifier and the number of neighbors is set to 2, and the feature selection and classification processes are executed 10 times. Baseline is the average result of classifying the data with unknown labels 10 times by the K-nearest neighbor classifier trained with the original known label data. Our Method is the average result of 10 times of the recognition accuracy of classifying the data with unknown labels after feature selection by the K-nearest neighbor classifier trained with the known label data after feature selection. It can be seen from the classification results that "Baseline" is the result calculated using the original data, "Our Method" is the result calculated using the features selected for all pixels after the present invention performs feature selection on the original data, "ACC" is the classification accuracy, and "m" is the number of selected features. The results show that when p = 1, γ = 1, and the number of selected features is greater than 14, when the number of neighbors of the K-nearest neighbor classifier is set to 2, the feature selection method and classification method of the present invention can both obtain high precision, especially when more than half of the selected features can exceed the classification accuracy of all features.

Claims

1. A feature selection method based on multi-class logistic regression for object classification of hyperspectral images, characterized in that, it includes the following steps: Step 1: Construct a data matrix, a label vector, and a regression coefficient matrix; Obtain a set of hyperspectral images with each feature dimension being , where the total number of pixels in a single band is . All the features of the -th pixel are represented as . Each element value of is the feature value of the sample. The class label vector of samples is , where represents the class of the -th sample, and is the total number of pixel ground object classes; Regression coefficient matrix The element in the th row and th column of represents the regression coefficient of the th feature with respect to the th class of samples; The th element of the regression deviation vector represents the regression deviation of the th class of samples; Step 2: Construct a feature selection model based on multi-class logistic regression; Maximize the weighted deviation of the feature regression coefficients for all classes: This problem is equivalent to the following minimization problem: Apply the norm regularization term to the model to improve the coefficient degree of the regression coefficient matrix; The objective function of the final feature selection model is constructed as follows: (1) In the formula, represents another sample category different from ; represents the characteristic regression coefficient vector of the -th class of samples, that is, the -th column vector of the regression coefficient matrix represents the characteristic regression coefficient vector of the -th class of samples, that is, the -th column vector of the regression coefficient matrix , respectively represent the regression deviations of the -th class and the -th class of samples, is a constant, is a fixed value; Step 3: Solve the feature selection model; Use the gradient descent method to find the optimal solution, and iteratively update the regression deviation and regression coefficients according to the following steps: (2) Among them and respectively represent the regression deviation vectors of the -th and -th steps, and respectively represent the regression coefficient matrices of the -th and -th steps, is the update step size; Specifically: ①Let , randomly initialize , and set the convergence accuracy ; ②Update the regression deviation vector element by element : Where: (3) (4) Among them, represents the set of samples of the th class, represents the total number of samples of the th class, represents the set of samples of the th class, represents the total number of samples of the th class; ③Update the regression coefficient matrix element by element : Where (5) (6) Among them, represents the element in the th row and th column of the sample matrix represents the th row vector of the regression coefficient matrix ④ Calculate the difference between the objective function values of the last two iterations: , if then end, otherwise let increment by 1, and go back to step ② until ; Step 4: Extract the indices of the features; Calculate , and then select the index of the maximum value as the selected feature.

2. The feature selection method based on multi-class logistic regression for object classification of hyperspectral images according to claim 1, characterized in that, The said has a value range of .