Feature selection method and system for label fuzzy relaxation

The class structure of high-dimensional feature space is learned through the fuzzy unsupervised learning method, and the label relaxation is used to solve the shortcomings of hard labels and unsupervised learning in the existing technology, and better generalization performance and interpretability are achieved.

CN120180064APending Publication Date: 2025-06-20THE HONG KONG POLYTECHNIC UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311743438.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The use of hard labels in the model learning process weakens the correlation between categories and ignores potential semantics; unsupervised learning methods cannot utilize important information carried by sample tags; the binary label matrix or linear model used in the classification process is too tough to reflect the characteristics of real-world data.

Method used

The class structure of high-dimensional feature space is learned based on fuzzy unsupervised learning method, the class structure is expressed using fuzzy membership, and the label relaxation is performed through fuzzy membership, and the supervised embedded feature selection framework is included for optimization and solution.

Benefits of technology

By learning class structure and utilizing fuzzy membership for label relaxation, the overfitting problem in traditional methods is solved, the generalization performance of the model is improved, and interpretability is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180064A_ABST
    Figure CN120180064A_ABST
Patent Text Reader

Abstract

The invention provides a feature selection method based on label fuzzy relaxation, and the method comprises the following steps: a, obtaining sample data, carrying out the feature extraction of the sample data, and obtaining a feature data matrix; b, learning the fuzzy membership degree to obtain a fuzzy membership degree matrix; c, performing soft relaxation on the label matrix by using the fuzzy membership matrix, and constraining the feature selection matrix as row sparsity; and d, based on the steps b and c, obtaining a target function of feature selection based on label fuzzy relaxation, and solving the target function. In addition, the invention also provides a feature selection system based on label fuzzy relaxation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to a feature selection method and system for label fuzzy relaxation. Background Art

[0002] Nowadays, in order to more comprehensively describe the target object, the collected data usually has high-dimensional feature expressions. When the dimension of the data is very high, many machine learning problems become quite difficult. The number of different possible configurations of a set of variables will increase exponentially with the increase in the number of variables. In order to avoid the curse of dimensionality and reduce the data noise problem in high-dimensional samples, researchers have proposed feature selection methods and techniques. Feature selection mainly filters and sorts the high-dimensional features of the samples, selects important features, and improves the accuracy of the classifier.

[0003] In patent CN111652271A, Zhu Jianyong et al. invented a nonlinear feature selection method based on neural network. This method proposes to change the linear error function in sparse regularization to a neural network error function, and to perform group sparse constraints on the weights of the neural network input layer according to the complexity of the neural network weights, so as to improve the prediction accuracy of nonlinear problems of the sparse regularization model. In addition, when solving the neural network, L 2,1 The norm is used to solve the error function and reduce the impact of outliers on the feature selection results.

[0004] In patent CN111931562A, Zhu Lei et al. invented an unsupervised feature selection method and system based on soft label regression. The method is based on fuzzy clustering to learn the soft labels of data samples, learns the feature selection matrix through a sparse regression model, establishes a connection between soft label learning and feature selection matrix learning, solves for a more discriminative feature subset in the sample data, and improves the accuracy of the prediction model.

[0005] In patent CN113869454A, Wang Jingyu et al. invented a method for sparse feature selection of hyperspectral images based on fast embedded spectral analysis. This method introduces the F-norm regularization term to maintain the manifold structure of the data and keep the class information of the subspace as much as possible; and introduces L 2,0 Norm constraints,enforcing the sparse constraints of the subspace helps to obtain the feature subset with the richest class information.

[0006] The above patents are incorporated herein by reference. Although some patents have disclosed the use of regularization learning and soft label learning for feature selection of data, it can be seen from their technical solutions that there are still some deficiencies: 1) In the model learning process, hard labels are usually used, which weakens the association between classes and ignores a lot of potential semantics; 2) Unsupervised learning methods only analyze the relationship between features and cannot utilize the important information carried by sample labels; 3) In the classification process, strict binary label matrices or linear models are used for sample points on the decision boundary, and such processing methods are too strong and cannot well reflect the characteristics of real-world data. Summary of the Invention

[0007] The present invention provides a method for learning the class structure of training samples in a high-dimensional feature space by using a fuzzy unsupervised learning method, expressing the class structure by using fuzzy membership degrees, performing label relaxation by using fuzzy membership degrees, and incorporating it into a supervised embedded feature selection framework for optimization and solution.

[0008] According to an aspect of the present invention, a feature selection method based on label fuzzy relaxation is provided, including the following steps:

[0009] a. Obtain sample data, perform feature extraction on the sample data to obtain a feature data matrix;

[0010] b. Learn the fuzzy membership degree to obtain a fuzzy membership degree matrix;

[0011] c. Use the fuzzy membership degree matrix to perform soft relaxation on the label matrix, and constrain the feature selection matrix to be row-sparse;

[0012] d. Based on steps b and c, obtain an objective function for feature selection based on label fuzzy relaxation, and solve the objective function.

[0013] In some embodiments, in step b, the following method is used to learn the fuzzy membership degree:

[0014]

[0015]

[0016] where s.t. is the constraint condition, the first term is the term for learning the fuzzy membership degree, x i is the i-th sample data, o j is the j-th class center of the sample data, h ij is the membership degree between the i-th sample and the j-th class center, the second term uses the square F norm to perform regularization constraint on the membership degree matrix, H ∈ R n×c is the membership degree matrix of the sample data features, and α is the regularization parameter of the second term of the above formula.

[0017] In some embodiments, in step c, the objective function for soft relaxation of the label matrix is as follows:

[0018]

[0019] where the first term is the loss function, X ∈ R n×d is the feature data matrix, P ∈ R d×c is the feature selection matrix, Y ∈ R n ×c is the label matrix, H ∈ R n×c is the membership degree matrix of the sample data features, the second term imposes an L 2,1 norm penalty on the feature selection matrix, and γ is the regularization parameter of the second term in the above formula.

[0020] In some embodiments, in step d, the objective function for feature selection based on label fuzzy relaxation is as follows:

[0021]

[0022]

[0023] where s.t. is the constraint condition, the first term is for learning the fuzzy membership degree matrix, x i is the i-th sample data, o j is the j-th class center of the sample data, h ij is the membership degree between the i-th sample and the j-th class center, H ∈ R n×c is the membership degree matrix of the sample data features, the second term is the loss function and the generalization term, X ∈ R n×d is the feature data matrix, P ∈ R d×c is the feature selection matrix, Y ∈ R n×c is the label matrix, and α, γ, λ are regularization parameters.

[0024] In some embodiments, the objective function in step d is solved using the alternating optimization method.

[0025] In some embodiments, in the alternating optimization method, the update rule for P is:

[0026]

[0027] Convert this formula to the following equivalent form:

[0028]

[0029] By taking the derivative of the above formula with respect to P and setting the derivative to 0, we get P = (X T X + γΓ) -1 XT F, where F = (Y + H).

[0030] In some embodiments, in the alternating optimization method, o j The update rule for is:

[0031]

[0032] By taking the derivative of the above formula with respect to o j and setting the derivative to 0, we get

[0033] In some embodiments, in the alternating optimization method, the update rule for H is:

[0034]

[0035]

[0036] where s.t. are the constraint conditions. Simplifying the formula and setting R = XP - Y, the formula becomes

[0037]

[0038]

[0039] Furthermore, the formula is changed to the following vector form:

[0040]

[0041]

[0042] After organizing it, the Lagrangian function of this formula is expressed as follows:

[0043]

[0044] where η and θ are both Lagrangian coefficients. According to the Karush-Kuhn-Tucker conditions, the solution of H can be obtained in the following way:

[0045]

[0046] where the function () + indicates (a) + = max(0, a).

[0047] In some embodiments, the objective function for solving step d adopts the following pseudocode:

[0048] Input: X ∈ R n×d , Y ∈ R n×c, regularization parameters α, γ, and λ;

[0049] Output: Feature selection matrix P ∈ R d×c and nSel features;

[0050] Repeat the following steps (1)-(4):

[0051] Step (1): Update the fuzzy class center o j ;

[0052] Step (2): Update the membership matrix H;

[0053] Step (3): Update the feature selection matrix P until (||obj(t) - obj(t - 1)|| ≤ 10 -5 );

[0054] Step (4): Output the feature selection matrix P,

[0055] According to the feature selection matrix P, calculate ||p i ||2, (i = 1, 2,..., d), and then sort them in descending order, and take the top nSel features as the selected features.

[0056] According to another aspect of the present invention, a feature selection system based on label fuzzy relaxation is provided. The system includes:

[0057] A dataset acquisition module, which acquires sample data, extracts features from the sample data, and obtains a feature data matrix;

[0058] A fuzzy membership learning module, which learns the fuzzy membership to obtain a fuzzy membership matrix;

[0059] A label matrix soft relaxation module, which uses the fuzzy membership matrix to perform soft relaxation on the label matrix and constrains the feature selection matrix to be row-sparse;

[0060] An objective function solving module, which, based on the fuzzy membership learning module and the label matrix soft relaxation module, obtains the objective function of feature selection based on label fuzzy relaxation and solves the objective function.

[0061] In some embodiments, the system uses the methods mentioned in the above first aspect and its various embodiments.

[0062] The present invention content is provided to introduce the selection of concepts in a simplified form, and these concepts will be further described in the following detailed description. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Other aspects and advantages of the present invention will be illustrated by the following embodiments. Brief Description of the Drawings

[0063] The accompanying drawings include figures for further illustrating and clarifying the above and other aspects, advantages, and features of the present disclosure. It will be understood that these drawings only depict certain embodiments of the present disclosure and are not intended to limit its scope. The present disclosure will now be described and explained with additional specificity and detail by using the accompanying drawings, in which:

[0064] Figure 1 A flowchart of a supervised feature selection method for label fuzzy relaxation according to an embodiment of the present invention is shown;

[0065] Figure 2 The experimental results of an ablation experiment according to an embodiment of the present invention are shown;

[0066] Figure 3 The generalization performance results according to an embodiment of the present invention are shown. Detailed implementation manners

[0067] The present disclosure proposes a feature selection method and system for label fuzzy relaxation based on fuzzy theory and label relaxation technology.

[0068] Benefits, advantages, solutions to problems, and any elements that may cause any benefit, advantage, or solution to occur or become more obvious should not be construed as key, essential, or fundamental features or elements of any or all claims. The present invention is defined only by the appended claims, which include any modifications made during the pendency of this application and all equivalents of those claims that result therefrom.

[0069] In the following claims and the previous description of the present invention, unless the context requires otherwise due to the express language or necessary implication, the word "comprising" or variants such as "comprises" are used in an inclusive sense, that is, to specify the presence of the stated features, but not to exclude the presence or addition of other features in various embodiments of the present invention.

[0070] Feature selection refers to the process of selecting a subset of relevant features (i.e., attributes, metrics) for model construction. The goal of feature selection is to find the optimal feature subset. Feature selection can eliminate irrelevant or redundant features, thereby reducing the number of features, improving the model accuracy, reducing the running time, and enhancing the generalization ability of the model. By choosing different evaluation metrics, feature selection algorithms can be divided into three categories: wrapper methods, filter methods, and embedded methods. Wrapper methods score feature subsets, and the score of a subset is obtained by calculating the number of errors (the error rate of the model) committed by the model trained with the subset on the held-out set. Filter methods use surrogate metrics rather than the error rate to score feature subsets. Embedded methods perform feature selection during model construction, and the learning algorithm utilizes its own variable selection process, with feature selection and algorithm training carried out simultaneously. For example, when training with machine learning algorithms and models, the weight coefficients of each feature are obtained, and the features with large weight coefficients are selected according to the descending order of the weight coefficients. These weight coefficients often represent a certain contribution or importance of the features to the model.

[0071] A core problem in machine learning is to design algorithms that perform well not only on the training data but also generalize well to new inputs. Many strategies are explicitly designed to reduce the test error (possibly at the cost of increasing the training error). These strategies are collectively referred to as regularization. Regularization is one of the commonly used tools to reduce the risk of overfitting. Regularization can be defined as "a modification of the learning algorithm that aims to reduce the generalization error rather than the training error". For example, adding additional constraints that limit the parameter values to the machine learning model or adding additional terms to the objective function to impose soft constraints on the parameter values.

[0072] Many regularization methods limit the learning ability of the model (such as neural networks, linear regression, or logistic regression) by adding a parameter norm penalty Ω(θ) to the objective function J. We denote the regularized objective function as

[0073]

[0074] where X is an m×n design matrix, y is the associated target, and θ contains all the parameters (weights and biases). The weights can be understood as how each feature affects the prediction. α ∈ [0, ∞) is a hyperparameter that balances the relative contributions of the norm penalty term Ω and the standard objective function J(X; θ). Setting α to 0 means no regularization. The larger α is, the greater the corresponding regularization penalty. When our training algorithm minimizes the regularized objective function it reduces the error of the original objective J with respect to the training data and simultaneously reduces the scale of the parameters θ (or a subset of the parameters) under certain metrics. Choosing different parameter norms Ω will prefer different solutions. A norm can be understood as a function that maps an object to a non-negative real number.

[0075] In the norms of vectors, the L0 norm refers to the number of non-zeros in vector x. For example, if x = [0, 1, 1, 0, 0, 1], then ||x||0 = 3. L P The p-norm is defined as follows, where p ∈ R, p ≥ 1:

[0076]

[0077] Intuitively, the norm of vector x measures the distance from the origin to point x. When p = 1, the L1 norm is the sum of the absolute values of the elements in the vector:

[0078]

[0079] Both the L0 norm and the L1 norm can describe the sparsity of vectors. Sparsity means that some parameters in the optimal value are 0. When p = 2, the L2 norm is called the Euclidean norm:

[0080]

[0081] The L2 norm can be simplified as ||x||, that is, the subscript 2 is omitted. In addition, the squared L2 norm is also often used to measure the magnitude of a vector, calculated through the dot product x T x, denoted as

[0082]

[0083] The L2 parameter norm penalty can shrink the weights of features with a smaller covariance with the output target by adding a regularization term to the objective function.

[0084] When it is necessary to measure the magnitude of matrix A ∈ R m×n , the Frobenius norm of the following formula can be used, where

[0085] a ij is the element in the i-th row and j-th column of matrix A:

[0086]

[0087] In addition, the squared Frobenius norm is defined as:

[0088]

[0089] For matrix A ∈ R m×n , the L 2,1 norm is defined as follows:

[0090]

[0091] Matrix B ∈ R n×nThe inverse (or reciprocal) refers to the sum of the elements on the main diagonal and can be expressed as:

[0092]

[0093] Therefore, for matrix A ∈ R m×n , and ||A|| 2,1 = tr(A T ΓA) holds. Where Γ is a diagonal matrix and the i-th diagonal element is ε is a sufficiently small constant used to avoid the case where ||A i: ||2 is 0. A i: is the row vector of matrix A. Regarding the inverse of a matrix, it has the properties tr(A + C) = tr(A) + tr(C) and tr(AC) = tr(C)tr(A).

[0094] In addition, machine learning algorithms are roughly classified into unsupervised algorithms and supervised algorithms. Unsupervised learning uses unlabeled data for training and learns useful structural properties on this dataset. Unsupervised learning generates the entire probability distribution of the dataset, such as density estimation, synthesis or denoising, or clustering. Supervised learning algorithms are given a training set of inputs and outputs and learn how to associate the inputs and outputs. Among them, it is necessary to manually annotate the samples and add labels (markings) to the samples. These labels can be certain properties or targets of the inputs. A common way to represent a dataset is to design a matrix X ∈ R n×d . Each row of the design matrix X contains a different sample. Each column corresponds to a different feature. The selected features correspond to a feature selection matrix P ∈ R d×c . In supervised learning, a sample contains a label or target and a set of features. Usually, when dealing with a dataset containing an observed feature design matrix X, a label matrix Y ∈ R n×c can be designed.

[0095] The above introduces some concepts helpful for understanding the present invention. Much information in the real world is ambiguous and there is no clear judgment basis between many things. To more accurately describe the uncertain relationship between target objects, it is necessary to find a label relaxation matrix that is more in line with the real scenario and can flexibly represent the belonging probability. Assume that among n samples in the dataset, there are c potential semantic classes, and the label relaxation matrix of the samples is learned by calculating the membership degree of each sample to each potential semantic class. Theoretically speaking, the shorter the distance between a sample and the learned class center, the greater the membership degree should be assigned, that is, the higher the similarity. In the present invention, the Euclidean distance is used to measure the correlation between a sample and the class center. The formula is as follows:

[0096]

[0097]

[0098] Among them, s.t. is the constraint condition. Indicates the minimum value of the subsequent expression obtained by adjusting the independent variable o j , H. The first term is the learning fuzzy membership term, x i Represents the i-th sample data, o j Is the j-th class center of the sample data, h ij Is the membership degree between the i-th sample and the j-th class center. The second term is to use the square F-norm to regularize and constrain the membership matrix, H ∈ R n×c Is the membership matrix of the sample data features, and α is the regularization parameter of this term.

[0099] In the process of pattern recognition, a strict binary label matrix or a linear model is used for the sample points on the decision boundary (the boundary used to divide classes in statistical classification). Such processing methods are too rigid and cannot well reflect the characteristics of real-world data. Based on the fuzzy relaxation technology, the present invention uses the fuzzy membership matrix to perform soft relaxation on the label matrix. Under the potential semantic information, the association relationship between classes can be better utilized to improve the accuracy of the model. The formula for soft relaxation of the label matrix is as follows:

[0100]

[0101] Among them, Indicates the minimum value of the subsequent expression obtained by adjusting the independent variable P. The first term is the loss function. X ∈ R n×d Is the set of feature data. Feature extraction is performed on each obtained sample data x i To obtain the feature data matrix. P ∈ R d×c Is the feature selection matrix. Y ∈ R n×c Is the label matrix. H ∈ R n×c Is the membership matrix of the sample data features. In the second term, an L 2,1 Norm penalty is imposed on the feature selection matrix P to constrain it to be row-sparse. γ is the regularization parameter of this term.

[0102] Based on formulas (1) and (2), the objective function of feature selection based on label fuzzy relaxation provided by the present invention is as follows:

[0103]

[0104]

[0105] Among them, s.t. is the constraint condition. Indicates the minimum value of the subsequent expression obtained by adjusting the independent variable o j,The minimum value of the subsequent expression obtained by H, P. The first term is for the learning of the fuzzy membership matrix. The second term is the loss function and the generalization term. λ is the regularization parameter, and the remaining parameters are the same as those in the description part of equations (1) and (2) above.

[0106] Figure 1 The flowchart of the supervised feature selection method for label fuzzy relaxation according to an embodiment of the present invention is shown. In the process of solving formula (3), an alternating optimization method can be adopted:

[0107] (1) Update rule of P

[0108]

[0109] Formula (4) is non-differentiable, so formula (4) is transformed into the following equivalent form:

[0110]

[0111] Where represents the value of the independent variable P that makes the subsequent expression obtain the minimum value.

[0112] Let F = (Y + H)

[0113]

[0114] By taking the derivative of the above formula and setting the derivative to 0, we can get:

[0115]

[0116] (2) Update rule of o j

[0117]

[0118] By taking the derivative of the above formula and setting the derivative to 0. Where I is the identity matrix, that is, a matrix with elements on the main diagonal all being 1 and other elements all being 0. The following deduction can be obtained:

[0119]

[0120] (3) Update rule of H

[0121]

[0122]

[0123] To simplify the formula, set R = XP - Y, the formula can be changed to:

[0124]

[0125]

[0126] Furthermore, the formula can be changed to the following vector form (corresponding to the row vectors of the matrix):

[0127]

[0128]

[0129] Therefore, after organizing it, the Lagrangian function expression of this formula is as follows:

[0130]

[0131] Among them, both η and θ are Lagrangian coefficients. According to the Karush-Kuhn-Tucker conditions, the solution of H can be obtained in the following way:

[0132]

[0133] Among them, the function () + indicates (a) + = max(0, a).

[0134] Although the present invention is solved by the above-mentioned alternating optimization method. It can be understood that other methods in the art can also be used for solving.

[0135] Next, the pseudo-code of the algorithm is introduced:

[0136] Input: X ∈ R n×d , Y ∈ R n×c , regularization parameters α, γ, and λ;

[0137] Output: Feature selection matrix P ∈ R d×c and nSel features.

[0138] The following are updated successively according to the relevant update rules, and the following steps 1-4 are repeated:

[0139] Step 1: Update the fuzzy class center o j ;

[0140] Step 2: Update the membership matrix H;

[0141] Step 3: Update the feature selection matrix P until (||obj(t) - obj(t - 1)|| ≤ 10 -5 );

[0142] Step 4: Output the feature selection matrix P.

[0143] According to the feature selection matrix P, ||p i||2, (i = 1, 2, …, d), and then sort them in descending order. Select the first nSel features as the selected features.

[0144] The present invention utilizes a fuzzy unsupervised learning method to learn the class structure of training samples in a high-dimensional feature space, uses fuzzy membership degrees to represent the class structure, and uses fuzzy membership degrees for label relaxation, and incorporates it into a supervised embedded feature selection framework for optimization and solution. The present invention solves the following problems in traditional methods: For example, traditional label relaxation methods only consider the criterion of "infinitely expanding the class interval", and this way is prone to overfitting. The present invention solves the overfitting problem by learning the class structure of training samples and using the fuzzy membership degree matrix reflecting the class structure for label relaxation. In addition, traditional label relaxation methods have poor interpretability, while the label relaxation method based on fuzzy membership degrees adopted by the present invention has better interpretability because it combines the class structure information of training samples.

[0145] The present invention has at least the following advantages over the prior art:

[0146] 1) The present invention learns the class structure of training samples based on fuzzy theory and uses fuzzy membership degrees for representation. This representation method is more in line with the complex data scenarios in the real world.

[0147] 2) The present invention has better generalization performance by learning the class structure of training samples and using the fuzzy membership degree matrix reflecting the class structure for label relaxation.

[0148] 3) The label relaxation method based on fuzzy membership degrees adopted by the present invention has better interpretability because it combines the class structure information of training samples.

[0149] The present invention has been effectively verified. From 2012 to 2015, an analysis of whether radiotherapy plans need to be reset was carried out on 310 nasopharyngeal carcinoma patients who received radiotherapy at Queen Elizabeth Hospital in Hong Kong, China. For each patient, the inventor obtained 4000 features. These features were used as the original data for testing the present invention.

[0150] The inventor conducted experiments and obtained significant technical results as Figures 2 - 3 shown. Figure 2 This is the experimental result of the ablation experiment, where n_LR (no label relaxation) is the device without using label relaxation, and LLR (luxury matrix label relaxation) is the traditional device using the luxury matrix for label relaxation. From Figure 2It can be seen that the classification performance (evaluated by AUC) of the features selected by the Fuzzy Label Relaxation (FLR) device (or label fuzzy relaxation) used in the present invention is better than that of n_LR and LLR. Among them, the evaluation index AUC is the Area Under Curve. Figure 3 is the generalization performance result, and the adopted index is the absolute value of the error between the training accuracy and the test accuracy. The smaller it is, the better the generalization ability. From Figure 3 it can be seen that the generalization performance of the device FLR of the present invention is better than that of the device without using label relaxation and the traditional label relaxation device.

[0151] The present invention can also be applied to feature selection before the construction of various prediction or diagnosis models, such as feature selection before the construction of clinical diagnosis / prediction models, radiotherapy prognosis prediction models, etc. In addition to being applicable to the medical field, the present invention can also be applied to feature selection applications in other fields such as industrial processes and image processing.

Claims

1. A feature selection method based on label fuzzy relaxation, the method comprising the following steps: a. Obtain sample data, perform feature extraction on the sample data to obtain a feature data matrix; b. Learn the fuzzy membership degree to obtain a fuzzy membership degree matrix; c. Use the fuzzy membership degree matrix to perform soft relaxation on the label matrix, and constrain the feature selection matrix to be row-sparse; d. Based on steps b and c, obtain the objective function of the feature selection based on label fuzzy relaxation, and solve the objective function.

2. The method according to claim 1, wherein, In step b, the following method is used to learn the fuzzy membership degree: Among them, s.t. are the constraint conditions. The first term is the learning fuzzy membership term, and x i is the i-th sample data, and o j is the j-th class center of the sample data, and h ij is the membership degree between the i-th sample and the j-th class center. The second term uses the square F-norm to regularize and constrain the membership matrix. H ∈ R n×c is the membership matrix of the sample data features, and α is the regularization parameter of the second term in the above formula.

3. The method according to claim 2, wherein, In step c, the objective function for performing soft relaxation on the label matrix is: Among them, the first item is the loss function, \(X\in R\) n×d is the feature data matrix, \(P\in R\) d×c is the feature selection matrix, \(Y\in R\) n×c is the label matrix, \(H\in R\) n×c is the membership degree matrix of the sample data features. The second item applies an \(L\) 2,1 norm penalty to the feature selection matrix, and \(\gamma\) is the regularization parameter of the second item in the above formula.

4. The method according to claim 1, wherein, In step d, the objective function of the feature selection based on label fuzzy relaxation is: where s.t. is the constraint condition, the first term is for the learning of the fuzzy membership matrix, x i is the i-th sample data, o j is the j-th class center of the sample data, h ij is the membership degree between the i-th sample and the j-th class center, H ∈ R n×c is the membership matrix of the sample data features, the second term is the loss function and the generalization term, X ∈ R n×d is the feature data matrix, P ∈ R d ×c is the feature selection matrix, Y ∈ R n×c is the label matrix, and α, γ, λ are regularization parameters.

5. The method according to claim 4, wherein, The alternating optimization method is used to solve the objective function in step d.

6. The method according to claim 5, wherein, In the alternating optimization method, the update rule of P is: Convert this formula into the following equivalent form: By taking the derivative of the above equation with respect to P and setting the derivative to 0, we get P = (X T X + γΓ) -1 X T F, where F = (Y + H).

7. The method according to claim 5, wherein, In the alternating optimization method, o j The update rule is: By taking the derivative of the above equation with respect to o j and setting the derivative to 0, we obtain 8. The method according to claim 5, wherein, In the alternating optimization method, the update rule of H is: where s.t. is the constraint condition. Simplify the formula and set R = XP - Y, and the formula becomes Furthermore, the formula is changed to the following vector form: After sorting it out, the Lagrangian function of this formula is expressed as follows: Among them, both η and θ are Lagrangian coefficients. According to the Karush-Kuhn-Tucker conditions, the solution of H can be obtained in the following way: Among them, the function () + indicates that (a) + = max(0, a).

9. The method according to claim 4, wherein, The following pseudocode is used to solve the objective function in step d: Input: X ∈ R n×d , Y ∈ R n×c , regularization parameters α, γ, and λ; Output: Feature selection matrix P ∈ R d×c and nSel features; Repeat the following steps (1)-(4): Step (1): Update the fuzzy class center o j ; Step (2): Update the membership degree matrix H; Step (3): Update the feature selection matrix P until (||obj(t) - obj(t - 1)|| ≤ 10 -5 ); Step (4): Output the feature selection matrix P, Select matrix P according to the said feature, calculate ||p i ||2, (i = 1, 2, …, d), then sort them in descending order, and take the first nSel features as the selected features.

10. A feature selection system based on label fuzzy relaxation, the system comprising: A dataset acquisition module, which acquires sample data, performs feature extraction on the sample data to obtain a feature data matrix; A fuzzy membership degree learning module, which learns the fuzzy membership degree to obtain a fuzzy membership degree matrix; A label matrix soft relaxation module, which uses the fuzzy membership degree matrix to perform soft relaxation on the label matrix, and constrains the feature selection matrix to be row-sparse; An objective function solving module, which, based on the fuzzy membership degree learning module and the label matrix soft relaxation module, obtains the objective function of the feature selection based on label fuzzy relaxation, and solves the objective function.

11. The system according to claim 10, wherein, The system uses the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Hyperspectral image sparse feature selection method based on fast embedded spectrum analysis

    CN113869454A