A data dimensionality reduction method based on fuzzy local discriminant analysis
By fuzzy clustering of each category in the optimal subspace and introducing orthogonal constraints for regularizing the maximum population divergence, the existing local discriminant analysis methods have solved the problems of high computational complexity and high parameter redundancy, and the efficiency and robustness of data dimensionality reduction are achieved.
Patent Information
- Application Number
- CN202210673065.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-06-14
AI Technical Summary
The existing local discriminant analysis methods have high computational complexity, difficult parameter adjustment, are easily disturbed by noise and redundant features, and are difficult to apply in large quantities in actual scenarios.
A data dimensionality reduction method based on fuzzy local discriminant analysis is proposed. By fuzzy clustering of each category in the optimal subspace, it adapts to multimodal data, and introduces regularized maximum population divergence to impose orthogonal constraints on the projection matrix, reducing the computational complexity and parameter redundancy.
This method reduces the computational complexity and parameter redundancy, retains the clustering structure of each category, overcomes the influence of noise and redundant characteristics, and enhances the global information representation ability of the data.
Smart Images

Figure CN115169436B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning and relates to a data dimensionality reduction method based on fuzzy local discriminant analysis. Background Art
[0002] With the development of computer science, the raw data obtained by people from various fields has characteristics such as high dimensionality, high redundancy, and complex distribution. This not only results in low computational efficiency but also causes the "curse of dimensionality" problem. Data dimensionality reduction methods map the original high-dimensional data to a low-dimensional space and retain the structural characteristics of the original space, thereby reducing the computational burden and improving the generalization performance. Currently, data dimensionality reduction techniques have been widely applied in fields such as computer vision, pattern recognition, and medicine. Recently, supervised data dimensionality reduction methods based on local discriminant analysis have received great attention from researchers because they are robust to noisy data and non-Gaussian distributed data, can explore both the local and global structures of samples simultaneously, and have better practical promotion performance compared to traditional linear discriminant analysis methods (Linear Discriminant Analysis, LDA). They have achieved successful applications in scenarios such as hyperspectral image processing and remote sensing image classification.
[0003] Yao Yu et al. ("Robust Nonnegative Supervised Low-Rank Discriminant Embedding Algorithm", Journal of Intelligent Science and Technology, 2021, 3(03): 342-350.) combined the discriminant information of the divergence matrix with non-negative matrix factorization, retained the local and global features of the data, and enhanced the sparsity and robustness of the noise through L1 norm constraints. In addition, this method also introduced graph embedding theory and low-rank representation to characterize local information, avoiding the influence of artificially selecting neighbor parameters. However, this model itself is a Non-deterministic Polynomial (NP) problem. When solving it, it is approximately reduced to a convex optimization problem, and it is difficult to obtain an accurate optimal solution. In addition, the entire algorithm process requires alternating iterative optimization of seven variables, and the model parameters are complex, with high computational difficulty, making it impossible to be widely applied in actual scenarios.
[0004] Most existing local discriminant analysis methods construct a similarity graph between all sample points in the original space through Gaussian kernel functions or k-nearest neighbor methods. This not only makes parameter adjustment difficult and has high computational complexity but is also easily interfered by noise and redundant features in the original space. Summary of the Invention
[0005] Technical Problems to be Solved
[0006] To avoid the deficiencies of the prior art, the present invention proposes a data dimensionality reduction method based on fuzzy local discriminant analysis, which reduces the computational complexity and parameter redundancy while preserving the clustering structure of each category. The algorithm performs fuzzy clustering on each category in the optimal subspace to adapt to the multi-modal data of the same category and overcome the influence of noise and redundant features. In addition, by introducing the regularized maximum total scatter, an orthogonality constraint is imposed on the projection matrix to enhance the algorithm's ability to represent the global information of the data.
[0007] Technical solution
[0008] A data dimensionality reduction method based on fuzzy local discriminant analysis, characterized by the following steps:
[0009] Step 1: Perform data preprocessing on the data matrix and the label matrix:
[0010] The original data matrix is where n is the number of sample points and d is the dimension of the sample points; the label vector is where the element y i represents the category number, usually an integer between 1 and c, and c is the number of categories of the sample points;
[0011] According to the categories represented by the values of the label vector, the data matrix is rearranged and centralized so that the sum of the rows of the data matrix is 0, that is, X1 n = 0, where is a column vector with all elements being 1; record the processed data matrix as X;
[0012] Step 2: Establish a data dimensionality reduction model based on fuzzy local discriminant analysis:
[0013]
[0014]
[0015] where c k represents the number of clustering centers of the k-th category. represents the clustering center of the low-dimensional projection space, satisfying The matrix is the clustering center matrix of the original space, and each represents the clustering center coordinates of a small category in the original space, is the total number of clustering centers. S t is the total scatter matrix. When the data matrix is centralized, S t = XX T . The positive integer q is the fuzzy clustering parameter, P1 n = 1 n represents the elements in P is 0 or and each sample point corresponds to q fuzzy clustering centers Obviously, for all classes k, q < c k . λ is a balance parameter, and its value is generally large. The purpose is to separate sample points as much as possible so that the model can learn the local features of the samples more accurately;
[0016] The is a block diagonal matrix composed of P (1) , P (2) ,..., P (c) as diagonal elements, that is:
[0017]
[0018] where is the membership matrix of the data points in the k-th class and their respective small-class fuzzy clustering centers;
[0019] Step 3, solve the data dimensionality reduction model:
[0020] ① Fix W and P, and optimize M
[0021]
[0022] ② Fix W and M, and optimize P
[0023]
[0024] where r t * is the index of the k-nearest neighbor points of the point in the subspace.
[0025] ③ Fix M and P, and optimize W
[0026] The optimal solution W is composed of the eigenvectors corresponding to the smallest d1 eigenvalues of the matrix . Among them, is the joint matrix of the data and the clustering centers, L S = D S - S, and:
[0027]
[0028]
[0029] At this time, the three variables M, P, and W are updated. Next, re-perform the next iteration calculation according to Step 3 until the objective function value converges; take the matrix obtained in the last iteration as the final projection matrix, then the dimensionality-reduced data matrix is Take Decentralization is performed to obtain the final projection result Z.
[0030] Advantageous Effects
[0031] A data dimensionality reduction method based on fuzzy local discriminant analysis proposed by the present invention constructs a data matrix, a label matrix and performs data preprocessing, establishes a data dimensionality reduction model based on fuzzy local discriminant analysis, solves the data dimensionality reduction model, and takes the matrix obtained in the last iteration as the final projection matrix. Then the dimensionality-reduced data matrix is The Decentralization is performed to obtain the final projection result Z. The present invention reduces the computational complexity and parameter redundancy, while retaining the clustering structure of each category. The algorithm performs fuzzy clustering on each category in the optimal subspace to adapt to the multi-modal data of the same category, and overcomes the influence of noise and redundant features. In addition, by introducing the regularized maximum total scatter, an orthogonal constraint is imposed on the projection matrix, enhancing the global information representation ability of the algorithm for data.
[0032] The beneficial effects of adopting the method of the present invention mainly include:
[0033] (1) A new method for calculating the within-class scatter matrix is proposed. By introducing the small-class fuzzy clustering center for each class of data, the calculation of the within-class scatter matrix is simplified, reducing the computational complexity of data dimensionality reduction.
[0034] (2) Using the global scatter matrix in the low-dimensional space as a balancing term, while minimizing the distance between samples of each class in the subspace, all sample points can be as scattered as possible, enabling the model to adaptively learn different local features of the sample points, not only avoiding the trivial solution, but also improving the learning performance of data dimensionality reduction.
[0035] (3) Discrete constraints are imposed on the membership matrix for fuzzy clustering, and adaptive iterative updates are performed in the low-dimensional space to automatically assign the best q clustering centers to each sample, reducing the influence of noise in the original space. Description of the Drawings
[0036] Figure 1 is the algorithm flow chart
[0037] Figure 2 is the grayscale image on the Yale_32×32 face dataset Detailed Embodiments
[0038] The present invention will be further described in combination with embodiments and drawings:
[0039] The present invention proposes a data dimensionality reduction method based on fuzzy local discriminant analysis, and the specific steps are as follows:
[0040] Step 1: Construct a data matrix, a label matrix, and perform data preprocessing.
[0041] Assume the original data matrix is where n is the number of sample points and d is the dimension of the sample points. The label vector is where the element y i represents the class serial number, usually an integer between 1 and c, and c is the number of classes of the sample points. For the convenience of subsequent data processing, the data matrix is rearranged according to the class order and centralized so that the sum of the rows of the data matrix is 0, that is, X1 n = 0, where is a column vector with all elements being 1. Record the processed data matrix as X.
[0042] Step 2: Establish a data dimensionality reduction model based on fuzzy local discriminant analysis.
[0043] In Step 1, the data matrix X has been arranged according to the label order, that is, X = [X (1) , X (2) ,..., X (c) , where represents the data matrix composed of the i-th class of samples, and n i represents the number of samples in the i-th class. Then, let the projection matrix be d1 is the dimension of the low-dimensional space. The following model is used to explore the local relationship between data points in each class:
[0044]
[0045] where, is a cell array, arranged according to the class order, and is the membership matrix between all data pairs in each class. P (k) is its k-th element, and there are c in total. represents the (i, j)-th element of P (k) , which reflects the and adjacency relationship between them.
[0046] Compared with the traditional LDA method, Equation (1) studies the distribution between data points in each class and can better learn the local structure of the samples. In addition, by imposing an orthogonality constraint, the projection vectors are linearly independent and the data reconstruction is simpler. However, this model depends on the distances between all data pairs in each class, has a high time complexity, and is prone to model redundancy. Therefore, the present invention uses a method based on subspace fuzzy clustering to improve this model, and at the same time introduces a subspace total scatter matrix as a balancing term to avoid trivial solutions, obtaining the following objective function:
[0047]
[0048] Among them, c k represents the number of cluster centers of the k-th class. represents the cluster center in the low-dimensional projection space, satisfying Matrix is the cluster center matrix of the original space, and each represents the cluster center coordinates of a small class in the original space. is the total number of cluster centers. S t is the total scatter matrix. When the data matrix is centered, S t = XX T . The positive integer q is the fuzzy clustering parameter. P1 n = 1 n indicates that the elements in P are 0 or and each sample point corresponds to q fuzzy cluster centers Obviously, for all classes k, q < c k . λ is the balance parameter, and its value is generally large, aiming to separate the sample points as much as possible so that the model can learn the local features of the samples more accurately.
[0049] Note that the definition of matrix P at this time is different from that in Equation (1). Equation (2)'s is a block diagonal matrix composed of P (1) , P (2) ,..., P (c) as diagonal elements, that is:
[0050]
[0051] Among them, is the membership matrix of the k-th class data points and their respective small-class fuzzy cluster centers. This reduces the dimension of matrix P and greatly reduces the computational burden of subsequent data processing.
[0052] Step 3: Solve the data dimensionality reduction model.
[0053] The objective function (2) has a total of 3 optimization variables and is optimized and solved using the alternating iteration method. First, initialize according to the constraint conditions to obtain an arbitrary unitary orthogonal matrix W0 and membership matrix P0. Then, perform iterative optimization, and the specific steps are as follows:
[0054] ① Fix W and P, and optimize M.
[0055] At this time, there is only one variable M, and the optimization function is:
[0056]
[0057] Since the data of each category is independent, it can be optimized separately to obtain the following formula:
[0058]
[0059] Take the partial derivative of the objective function in Equation (5) with respect to and set the partial derivative to 0 to obtain the equation:
[0060]
[0061] We can get:
[0062]
[0063] The above formula has obtained the clustering center of each subspace in the k-th category Perform the same operation for each category, and we can obtain M (1) , M (2) ,..., M (c) respectively. Finally, merge them to obtain all subspace clustering center matrices M = [M (1) , M (2) ,..., M (c) .
[0064] ② Fix W and M, and optimize P.
[0065] At this time, only P is the variable, and the objective function is:
[0066]
[0067] Similar to Step ①, since each category is independent, first consider only the data of the k-th category. From the definition of P, the constraint P1 n = 1 n can be transformed into That is, each row of P (k) has q elements as and the other elements are all 0. Then, Equation (8) is transformed into optimizing P (k) separately, and its vector form is expressed as:
[0068]
[0069] Among them, represents the i-th row vector of P (k) . Assume that the subscripts of the non-zero elements in the vector are {r1, r2,..., r t ,..., r ck-q}(1 ≤ r t ≤ c k - q), Equation (9) can be simplified to:
[0070]
[0071] Obviously, the optimal solution r of Equation (10) t * is the index of the k-nearest neighbor points of the points in the subspace. Therefore, the optimal solution of Problem (9) is:
[0072]
[0073] It can be obtained through Equation (11) and then P (k) is obtained. The operations in Step ② are performed separately for each category to obtain P (1) , P (2) ,..., P (c) , and finally the membership matrix P of all sample points is obtained.
[0074] ③ Fix M and P, and optimize W.
[0075] The objective function in this case is:
[0076]
[0077] According to the embedding expression of the Laplacian matrix, the first term of Equation (12) is transformed into the form of the matrix trace. After the transformation, the above equation becomes:
[0078]
[0079] Among them, is the joint matrix of the data and the cluster centers, L S is the Laplacian matrix, which is derived from the similarity graph matrix S of the joint matrix . The definition of the matrix S is as follows:
[0080]
[0081] The degree matrix is:
[0082]
[0083] Then the Laplacian matrix L S is calculated by the formula L S = D S - S. Equation (13) can be equivalently transformed into:
[0084]
[0085] Equation (16) has only one orthogonality constraint, and the optimal solution of the constraint variable can be obtained by eigenvalue decomposition. The obtained optimal solution W is composed of the matrix Composed of the eigenvectors corresponding to the smallest d1 eigenvalues.
[0086] So far, the three variables M, P, and W have been updated. Next, perform the next iteration calculation according to step 3 again until the objective function value converges. Take the matrix obtained in the last iteration as the final projection matrix, then the data matrix after dimensionality reduction is Will De-center to obtain the final projection result Z.
[0087] The basic flowchart of the embodiment of the present invention is as Figure 1 Shown. Next, take the Yale_32×32 face dataset applied to the data dimensionality reduction problem as an example to introduce the specific implementation method, including the following steps:
[0088] Step 1: Construct a data matrix, a label matrix, and perform data preprocessing.
[0089] Obtain the Yale_32×32 face image dataset. The number of images n = 165, the image resolution is 32×32. Stretch each image into a vector with d = 1024 dimensions. There are a total of c = 15 categories. Thus, the original data matrix is The label vector is Element y i (i = 1, 2,..., 165) is an integer between 1 and 15, representing the sample category. Arrange the data points in the sample matrix in the category order and perform centering processing. Record the processed data matrix as X.
[0090] Step 2: Establish a data dimensionality reduction model based on fuzzy local discriminant analysis.
[0091] The data matrix X is arranged in the label order as X = [X (1) , X (2) ,..., X (c) , where represents the data matrix composed of the i-th class of samples, and n i represents the number of samples in the i-th class. In this example, n i are all 13. The projection matrix is d1 is the dimension of the low-dimensional space. The objective function of the model is:
[0092]
[0093] Among them, c k represents the number of clustering centers in the k-th class, and generally takes an integer between 2 and 5. S t is the total scatter matrix. When the data matrix is centered, S t = XX T. The positive integer q is generally set to 2 or 3 (q < c k ). λ is a balance parameter and can be set to 2.
[0094] Step 3: Solve the data dimensionality reduction model.
[0095] The objective function (17) has a total of 3 optimization variables, and the alternating iteration method is used for optimization and solution. First, perform initialization to obtain any unitary orthogonal matrix W0 and membership matrix P0, and then perform iterative optimization. The specific steps are as follows: ① Fix W and P, and optimize M.
[0096] At this time, there is only one variable M, and the optimization function is:
[0097]
[0098] Since the data of each category are independent of each other, they can be optimized separately to obtain the following formula:
[0099]
[0100] Take the partial derivative of the objective function in Equation (19) with respect to and set the partial derivative to 0 to obtain the equation:
[0101]
[0102] It can be obtained that:
[0103]
[0104] Perform the same operation for each category to obtain M (1) , M (2) ,..., M (15) , and finally merge to obtain all subspace clustering center matrices M = [M (1) , M (2) ,..., M (15) .
[0105] ② Fix W and M, and optimize P.
[0106] At this time, only P is the variable, and the objective function is:
[0107]
[0108] First, only consider the data of the k-th category, and Equation (22) is transformed into the separate optimization of P (k) , and its vector form is expressed as:
[0109]
[0110] Among them, represents P (k)The i-th row vector. Assume the vector The subscripts of non-zero elements in t ,..., r ck-q}(1 ≤ r t ≤ c k - q), Equation (23) is simplified to:
[0111]
[0112] Obviously, the optimal solution r t * of Equation (24) is the index of the k-nearest neighbor points of the point in the subspace. Then the optimal solution of Problem (23) is:
[0113]
[0114] Through Equation (25), can be obtained, and then P (k) is obtained. Perform the operations in Step ② above for each category respectively to obtain P (1) , P (2) ,..., P (c) , and finally the membership matrix P of all sample points is obtained.
[0115] ③ Fix M and P, and optimize W.
[0116] The objective function in this case is:
[0117]
[0118] Equation (26) can be equivalently transformed into:
[0119]
[0120] where L S is the Laplacian matrix, and the expression is L S = D S - S, where:
[0121]
[0122]
[0123] The optimal solution of the constraint variable in Equation (26) is obtained by eigenvalue decomposition. The finally obtained optimal solution W is composed of the eigenvectors corresponding to the smallest d1 eigenvalues of the matrix .
[0124] So far, the three variables M, P, and W have been updated. Next, perform the next iteration calculation according to Step 3 again until the objective function value converges (take the deviation value ε = 10-4 )。Take the matrix obtained in the last iteration as the final projection matrix W.
[0125] Step 4: Classify and recognize the obtained low-dimensional results.
[0126] Calculate the low-dimensional projection result using the projection matrix W obtained in Step 3 The Decentralize to obtain the final projection result Take each column of Z as the new face image sample data, and classify it using a classification algorithm (such as K-nearest neighbor classifier, support vector machine, etc.). Finally, compare the classification result with the original sample label to obtain the final recognition accuracy.
Claims
1. A data dimensionality reduction method based on fuzzy local discriminant analysis, characterized in that The steps are as follows: Step 1: Perform data preprocessing on the data matrix and the label matrix: The original data matrix is where n is the number of sample points and d is the dimension of the sample points; the label vector is where the element y i represents the class serial number, usually an integer between 1 and c, and c is the number of classes of the sample points; The original data is face image data; according to the categories represented by the values of the label vectors, the data matrix is rearranged and centralized so that the sum of the rows of the data matrix is 0, that is, X1 n = 0, where is a column vector with all elements being 1; record the processed data matrix as X; Step 2: Establish a data dimensionality reduction model based on fuzzy local discriminant analysis: Among them, c k represents the number of cluster centers of the k-th class; represents the cluster center of the low-dimensional projection space, satisfying The matrix is the cluster center matrix of the original space, and each represents the cluster center coordinates of a small class in the original space, is the total number of cluster centers; S t is the total scatter matrix. When the data matrix is centered, S t = XX T ; The positive integer q is the fuzzy clustering parameter, P1 n = 1 n indicates that the element in P is 0 or and each sample point corresponds to q fuzzy cluster centers Obviously, for all classes k, q < c k ; λ is the balance parameter, and its value is generally large, aiming to separate sample points as much as possible so that the model can learn the local features of the samples more accurately; The said is a block diagonal matrix composed of P (1) , P (2) ,..., P (c) as diagonal elements, that is: Among them, is the membership degree matrix between the data points of the k-th category and the fuzzy clustering centers of their respective sub-categories; Step 3: Solve the data dimensionality reduction model: ① Fix W and P, and optimize M ② Fix W and M, and optimize P where r t * is the index of the k nearest neighbor points of the point in the subspace; ③ Fix M and P, and optimize W The optimal solution W consists of the eigenvectors corresponding to the smallest d1 eigenvalues of the matrix ; where is the joint matrix of data and cluster centers, L S = D S - S, and: At this time, the three variables M, P, and W are updated. Next, perform the next iteration calculation according to step 3 again until the objective function value converges; take the matrix obtained in the last iteration as the final projection matrix, then the data matrix after dimensionality reduction is Subtract the mean to obtain the final projection result Z.
Citation Information
Patent Citations
Sparse non-negative matrix under-approximation-based hyperspectral polychrome cultural relic sketch extraction method
CN108428237A
Brain function network multi-core fuzzy clustering method based on stacked encoder
CN111310787A