Nuclear transpose projection envelope LDA data dimension reduction method and system
Through the kernel transpose projection envelope LDA method, envelope samples are generated and dimensionality reduction is performed, which solves the problem of performance limitations of existing LDA algorithms in high-dimensional feature space, and achieves more efficient data dimensionality reduction and classification accuracy improvement.
Patent Information
- Application Number
- CN202510073288.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
Existing LDA algorithms are prone to intra-class divergence matrix singularity problems when dealing with high-dimensional feature spaces, and are sensitive to dimensionality reduction dimension selection and noise, and cannot effectively utilize the correlation information between similar samples, resulting in limited performance.
The nuclear transposed projection envelope LDA method is used to envelop the original sample through the nuclear transposed projection envelope transformation, and the envelope sample is generated, and the dimensionality reduction is used to fuse the correlation information between similar samples.
The performance of the LDA algorithm is improved, the problems of small sample size and noise sensitivity are solved, and more effective data dimensionality reduction and classification accuracy are improved.
Smart Images

Figure CN119988952A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing for machine learning, and in particular to a kernel transposed projection envelope (LDA) data dimension reduction method and system. Background Art
[0002] Data sets used for machine learning usually have multiple features. If these features are used directly for modeling without processing, good results may not be achieved. Studies have shown that problems such as correlation between features, redundant features, and noise often reduce the performance of machine learning models, and too many features will also bring heavy computational overhead. Data dimensionality reduction methods map the original high-dimensional features of the data to a new low-dimensional subspace and retain the discriminant information as much as possible, thereby eliminating the correlation between features, removing redundant features, and reducing computational overhead.
[0003] According to whether the sample category label information is used, data dimensionality reduction methods can be divided into unsupervised dimensionality reduction and supervised dimensionality reduction. Unsupervised dimensionality reduction methods refer to dimensionality reduction methods that do not use category label information. LDA (Linear Discriminant Analysis) is a classic supervised dimensionality reduction method with the advantages of low computational overhead and strong generalization. It has been widely used in pattern classification tasks such as finger vein recognition and face recognition. So far, it has been a research hotspot in academia and industry. Although LDA is widely used, it still has many defects, such as: 1. The optimization target solution of LDA depends on the non-singularity of the intra-class scatter matrix. When the feature dimension is much larger than the sample dimension, the intra-class scatter matrix usually becomes a singular matrix, which will cause LDA to fail to solve normally, that is, the small sample size problem; 2. LDA assumes that each category of input data obeys a multi-dimensional Gaussian distribution. For multimodal data that is more complex than Gaussian distribution data, LDA's performance is unsatisfactory. 3. LDA is sensitive to the choice of dimensionality reduction and noise.
[0004] Many scholars have made improvements to LDA to varying degrees to address the above defects. Although significant improvements have been made, there is still room for improvement. One of the main reasons is that the existing improved LDA algorithm model still has limitations: it focuses on dimensionality reduction projection of the original samples without effectively utilizing the correlation information between similar samples, resulting in limited performance. Summary of the invention
[0005] The present invention provides a kernel transposed projected envelope LDA data dimensionality reduction method and system, and solves the technical problem of how to improve the LDA algorithm by considering the association information between similar samples to achieve high-performance data dimensionality reduction.
[0006] In order to solve the above technical problems, the present invention provides a kernel transposed projected envelope LDA data dimensionality reduction method, which comprises the steps of:
[0007] Perform envelope transformation on each original sample using kernel transposed projection envelope transformation to obtain corresponding envelope samples; the kernel transposed projection envelope transformation is to introduce kernel method into transposed projection for kernelization; the transposed projection obtains similar samples of the original sample by neighbor distance measurement, forms sample envelope with the original sample and its similar samples, extracts correlation information by linear weighting of samples inside the sample envelope and fuses it with the original sample to obtain envelope samples corresponding to each original sample;
[0008] The LDA method based on envelope sample modeling is used to reduce the dimension of all envelope samples to obtain envelope samples after dimension reduction.
[0009] Furthermore, the original sample set is represented as X = [x1, x2, ..., x n ]∈R d×n , where d is the number of features of each original sample, n is the total number of samples, and the i-th original sample x i and its K nearest neighbor samples x i,1 ,...,x i,K Form x i The sample envelope E i =[x i ,x i,1 ,...,x i,K ]∈R d×(K+1) , i = 1:n; the purpose of transpose projection is to find a mapping vector p∈R (K+1)×1 , so that E i After mapping E i Generate sample x after p i The corresponding envelope sample
[0010] Furthermore, the transposed projection searches for the mapping vector p∈R (K+1)×1 The specific process is:
[0011] Based on the sample envelope E i Constructing the Matrix and centralize it Custom 1 (n×d)×1 represents a matrix of all 1s of size (n×d)×1, and the superscript T represents the matrix transpose;
[0012] right Perform eigenvalue decomposition, take the eigenvector corresponding to the maximum eigenvalue and record it as p.
[0013] Furthermore, the kernel method is introduced in the transposed projection for kernelization, specifically:
[0014] E i Each row (E i )j,: After nonlinear mapping, φ(·) becomes φ((E i ) j,: T ) T Then follow the subsequent steps of transposed projection to model.
[0015] Furthermore, the purpose of the kernel transposition projection envelope transform is to find a mapping vector α such that E i After α mapping, the sample x is generated i The corresponding envelope sample Superscript right It means that each row of the matrix is mapped to the feature space through a nonlinear mapping φ(·).
[0016] Furthermore, the process of finding the mapping vector α is as follows:
[0017] Based on the sample envelope E i Constructing the Matrix And calculate the kernel matrix through the Gaussian kernel function
[0018] calculate
[0019] Take the generalized eigenvalue decomposition problem The eigenvector corresponding to the maximum generalized eigenvalue is α, and λ represents the generalized eigenvalue.
[0020] Furthermore, the envelope sample after dimension reduction is input into the classifier to obtain the classification label, thereby obtaining the classification label of the corresponding original sample;
[0021] In the process of using the classifier for prediction, the envelope sample is generated Generate the sample x to be predicted in the same way test The corresponding envelope sample
[0022] Furthermore, the loss function used in the process of inputting the reduced dimension envelope sample into the classifier for training is: Represents the new training set consisting of all envelope samples, W∈R d×d′ Represents the dimension reduction matrix used by the linear discriminant analysis method, d′ is the sample dimension after dimensionality reduction, Error_term() represents the error term in the original LDA optimization objective, and Regular_term() represents the regular term in the original LDA optimization objective.
[0023] The present invention also provides a kernel transposition projection envelope LDA data dimension reduction system, the key of which is that the system applies the kernel transposition projection envelope LDA data dimension reduction method, and is provided with a kernel transposition projection envelope transformation module and an LDA module, the kernel transposition projection envelope transformation module is used to use the kernel transposition projection envelope transformation to perform envelope transformation on each original sample to obtain the corresponding envelope sample, and the LDA module is used to use the LDA method based on envelope sample modeling to reduce the dimension of all envelope samples to obtain the envelope samples after dimension reduction.
[0024] The present invention proposes the idea of mining the correlation information between samples for subsequent projection dimension reduction, thereby designing a kernel transposition projection envelope LDA data dimension reduction method and system (referred to as the kernel transposition projection envelope LDA mode). First, the present invention designs a transposition projection envelope transformation algorithm and kernelizes it to obtain a better kernel transposition projection envelope transformation algorithm, which is used to transform the original sample to generate an envelope sample, and the correlation information between similar samples is mined as much as possible. Secondly, the present invention loads the envelope sample to the LDA input end, thereby realizing the projection dimension reduction on the correlation information between similar samples. The experimental part uses 7 data sets and 7 representative LDA algorithms to verify the effectiveness of the LDA mode of this embodiment. The results show that after the introduction of the kernel transposition projection envelope LDA mode, the classification accuracy of various LDA algorithms has been significantly improved, which shows that the kernel transposition projection envelope LDA mode is better than the original LDA mode, realizes the projection dimension reduction on the correlation information between similar samples, and makes up for the defect that the original LDA mode ignores or destroys the correlation information between similar samples in the modeling process. Moreover, the kernel transposed projected envelope LDA model is not a specific LDA improved algorithm, so it has good universality. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of a kernel transposed projected envelope LDA data dimensionality reduction method and system provided by an embodiment of the present invention;
[0026] Figure 2 is a schematic diagram of a transposed projection envelope transform provided by an embodiment of the present invention;
[0027] Figure 3 is a distribution diagram of envelope samples generated by transposed projection provided by an embodiment of the present invention;
[0028] Figure 4 is a distribution diagram of envelope samples generated by nuclear transposition projection provided by an embodiment of the present invention;
[0029] Figure 5 It is a line graph of classification accuracy-dimension corresponding to the same LDA on the Pima dataset in different modes and different dimensionality reduction dimensions provided by an embodiment of the present invention;
[0030] Figure 6 It is a line graph of classification accuracy-dimension corresponding to the same LDA on the Australian dataset in different modes and different dimensionality reduction dimensions provided by an embodiment of the present invention;
[0031] Figure 7 It is a classification accuracy-dimension line graph corresponding to the same LDA on the Wisconsin dataset in different modes and different dimensionality reduction dimensions provided by an embodiment of the present invention;
[0032] Figure 8 It is a line graph of classification accuracy-dimension corresponding to the same LDA on the German dataset in different modes and different dimensionality reduction dimensions provided by an embodiment of the present invention;
[0033] Fig. 9 : is a diagram showing the influence of the transposed projection parameter K on the classification accuracy provided by an embodiment of the present invention;
[0034] Fig.10 : is a diagram showing the influence of kernel transposition projection parameters K and σ on classification accuracy on a German data set provided by an embodiment of the present invention;
[0035] Fig.11 : is a diagram showing the influence of kernel transposition projection parameters K and σ on classification accuracy on the Sonar dataset provided by an embodiment of the present invention;
[0036] Fig.12 : is a graph showing the influence of kernel transposition projection parameters K and σ on classification accuracy on a Control data set provided by an embodiment of the present invention;
[0037] Fig.13 This is a diagram showing the influence of kernel transposition projection parameters K and σ on classification accuracy on the Wdbc dataset provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The following specifically illustrates the implementation mode of the present invention in conjunction with the accompanying drawings. The embodiments are provided for illustrative purposes only and cannot be understood as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0039] The kernel transposed projected envelope LDA data dimensionality reduction method provided by the embodiment of the present invention is as follows: Figure 1 As shown, the steps include:
[0040] Perform envelope transformation on each original sample using kernel transposition projection envelope transformation to obtain the corresponding envelope sample;
[0041] The LDA (Linear Discriminant Analysis) method based on envelope sample modeling is used to reduce the dimension of all envelope samples to obtain envelope samples after dimension reduction.
[0042] In specific applications, the reduced-dimensional envelope sample is input into the classifier to obtain the classification label, thereby obtaining the classification label of the corresponding original sample. Since each original sample corresponds to an envelope sample, and each envelope sample corresponds to a reduced-dimensional envelope sample, the classification label of the corresponding original sample can be obtained through the classification label of a reduced-dimensional envelope sample obtained by the classifier.
[0043] Transposed projection obtains similar samples of the original sample through the nearest neighbor distance measurement, forms a sample envelope with the original sample and its similar samples, extracts the associated information through the linear weighting of the samples inside the sample envelope and fuses it into the original sample, and obtains the envelope sample corresponding to each original sample. By projecting the samples inside the sample envelope in the sample dimension direction to retain the information of each feature dimension, the projection result obtained in this way fuses the information of the samples inside the sample envelope as much as possible and contains the associated information between them. Compared with the conventional projection in the feature dimension, projection in the sample dimension contains a transposition operation, so it is called transposed projection.
[0044] Assume that the original sample set is represented by X = [x1, x2, ..., x n ]∈R d×n , where d is the number of features of each original sample and n is the total number of samples. The schematic diagram of the transposed projection envelope transform is as follows Figure 2 As shown, first, for each original sample, find the K nearest neighbor samples with the same label as its original sample. For example, the i-th original sample x i The K nearest neighbor samples of x are i The distance from near to far is denoted as x i,1 ,…,x i,K , these samples are considered to be i Similar samples, x i and its K nearest neighbor samples constitute x i The sample envelope E i =[x i ,x i,1 ,…,x i,K ]∈R d×(K+1) .
[0045] The purpose of transpose projection is to find a mapping vector p∈R (K+1)×1 , so that E i After mapping E i Generate sample x after p i The corresponding envelope sample Contains sample envelope E iThe correlation information between internal samples, that is, the sample envelope E i The main information of the envelope is therefore also called the main sample. The label is still the label of the original sample.
[0046] E i It can be regarded as a data set, with one line as a sample. Then the feature of the jth sample (E i The jth row of (E i ) j,: =[(x i ) j ,…,(x i,K ) j ], that is, x i and its K nearest neighbor samples x i,1 ,…,x i,K The jth feature of . It can be seen as data set E i The result of reducing to one dimension is that the jth feature satisfies Since principal component analysis (PCA) has the characteristic of retaining sample information to the greatest extent, this embodiment uses E i Use principal component analysis to reduce dimensionality, which is equivalent to Each dimension of x is preserved as much as possible. i Its nearest neighbor sample x i,1 ,…,x i,K The information in the corresponding dimension, i.e. To maximize the integration of x i Its nearest neighbor sample x i,1 ,…,x i,K The envelope samples generated by the information contain the correlation information between them. Principal component analysis requires the reconstruction error to be minimum, that is, the dimension reduction vector p should make Minimum (||·|| F represents the F-norm of the matrix, the superscript T represents the matrix transpose), and p T p=1. In this embodiment, the corresponding envelope samples are generated in the same way for all samples. Then the optimization target of the transposed projection can be written as follows:
[0047]
[0048] Min means minimize, and st means satisfy.
[0049] remember The above formula is equivalent to the following:
[0050]
[0051] tr() represents the trace of a matrix.
[0052] In fact, principal component analysis also requires that the data set be centered first, that is, E1,…,E n Centralize them separately. For simplicity, this embodiment directly uses After centralization To replace, (1 (n×d)×1 represents a matrix of all 1s of size (n×d)×1), we have The generation method of E i -1 d×1 ·μ)p. The optimization objective is finally written as follows:
[0053]
[0054] Using the KKT condition, the solution p is the matrix The eigenvector corresponding to the largest eigenvalue.
[0055] The transposed projection algorithm specifically includes the following steps:
[0056] Input training set X = [x1, x2, ..., x n ]∈R d×n , the number of neighbor samples K;
[0057] For the i-th sample x i Find its K nearest neighbor samples and construct the sample envelope E i , i=1:n;
[0058] Based on the sample envelope E i Constructing the Matrix and centralize it
[0059] right Perform eigenvalue decomposition, take the eigenvector corresponding to the maximum eigenvalue and record it as p;
[0060] Based on p, the sample x i Transformed into the corresponding envelope sample
[0061] Output a new training set consisting of envelope samples
[0062] In order to use the correlation information between similar samples in the model prediction process to improve the prediction accuracy, it is also necessary to transform the test sample into the corresponding envelope sample. The main operation is: let the sample to be tested be x test , find the same value as x in the training set test The most recent K samples, x test The K nearest neighbor samples of x are testThe distance from near to far is denoted as x test,1 ,…,x test,K , remember E test =[x test ,x test,1 ,…,x test,K ]. Then x test The corresponding envelope sample The way to generate is
[0063] In the transposed projection, the envelope sample It can be seen as data set E i The result of reducing the dimension to one dimension is a linear dimension reduction, which makes it difficult to mine the complex nonlinear correlation information between samples. In the above process, the kernel method can be introduced to achieve nonlinear dimension reduction. Suppose a nonlinear mapping φ(·) is given, which can map the input data to the feature space. In this embodiment, E i Each row (E i ) j,: After mapping φ(·) becomes φ((E i ) j,: T ) T Then it is reduced to one dimension, and the subsequent steps of transposed projection are continued to be modeled on this basis to form a nuclear transposed projection. The nuclear transposed projection uses nonlinear operations in the process of fusing the correlation information between the original sample and its similar samples, which can extract nonlinear correlation information and has better effect than the transposed projection.
[0064] make Superscript right Indicates that each row of the matrix is mapped to the feature space through a nonlinear mapping φ(·). Similarly, Centralized processing becomes Right now
[0065] make Then the optimization objective of the kernel transposition projection is written as follows:
[0066]
[0067] Define the kernel function make:
[0068] is the kernel matrix, then the objective function and constraints in the optimization problem (4) can be simplified as follows:
[0069]
[0070] 1 (n×d)×(n×d) Represents a (n×d)×(n×d) matrix of all ones.
[0071] make The final optimization goal is:
[0072]
[0073] Using the KKT condition, we can see that the optimization of problem (7) eventually becomes solving the generalized eigenvalue decomposition problem α is the eigenvector corresponding to its maximum generalized eigenvalue, and λ represents the generalized eigenvalue. i , whose envelope sample The generation method becomes 1 d×1 represents a matrix of all 1s of size d×1. This embodiment uses a Gaussian kernel function, that is, e represents the natural base, σ represents the Gaussian kernel parameter, (E i ) j,: Indicates E i The jth row of (E p ) q,: Indicates E p For the qth row of , i,p=1,2,…,n, j,q=1,2,…,d.
[0074] Specifically, the nuclear transposition projection algorithm includes the following steps:
[0075] Input training set X = [x1, x2, ..., x n ]∈R d×n , the number of neighbor samples K, the Gaussian kernel parameter σ;
[0076] For the i-th sample x i Find its K nearest neighbor samples and construct the sample envelope E i , i=1:n;
[0077] Based on the sample envelope E i Constructing the Matrix And calculate the kernel matrix through the Gaussian kernel function
[0078] calculate
[0079] Take the generalized eigenvalue decomposition problem The eigenvector corresponding to the maximum generalized eigenvalue is α;
[0080] Based on α, x i Transformed to
[0081] Output a new training set consisting of envelope samples
[0082] Similarly, the sample x to be predictedtest The corresponding envelope sample The way to generate is E test Represents the prediction sample x test The generated sample envelope.
[0083] The LDA loss function using the kernel transposed projected envelope LDA mode is constructed as:
[0084]
[0085] Among them, Error_term(·) represents the error term in the original LDA optimization objective, Regular_term() represents the regular term in the original LDA optimization objective, and W∈R d×d′ represents the dimension reduction matrix of LDA, and d′ is the sample dimension after dimension reduction. The error term here is a composite function based on the original sample, which helps to perceive the association information (deep information) between similar samples, and thus helps to more accurately obtain the category label.
[0086] Corresponding to the above-mentioned kernel transposition projection envelope LDA data dimensionality reduction method, this embodiment also provides a kernel transposition projection envelope LDA data dimensionality reduction system, which is provided with a kernel transposition projection envelope transformation module and an LDA module. The kernel transposition projection envelope transformation module is used to use the kernel transposition projection envelope transformation to perform envelope transformation on each original sample to obtain the corresponding envelope sample, and the LDA module is used to use the LDA method based on envelope sample modeling to reduce the dimension of all envelope samples to obtain the envelope samples after dimension reduction. The kernel transposition projection envelope transformation module and the LDA module refer to electronic entities that exist in physical form and can realize corresponding functions.
[0087] In order to verify the effectiveness of the proposed kernel transposed projection envelope LDA data dimensionality reduction method and system, this embodiment used seven real-world data sets and multiple representative LDA algorithms for experiments. In the first set of experiments, this embodiment performed a visualization analysis of the effectiveness of transposed projection and kernel transposed projection on two toy data sets, thereby indicating that envelope samples are more helpful in improving subsequent LDA modeling performance than original samples. In the second set of experiments, this embodiment applies the LDA model of this embodiment to multiple LDA algorithms, thereby illustrating that the model of this embodiment can effectively improve the LDA modeling performance. In the third set of experiments, this embodiment analyzes the relationship between key parameters and method performance, thereby providing a reference for parameter optimization for interested readers. In the fourth set of experiments, this embodiment analyzes the runtime cost of reconstructing envelope samples using the LDA model of this embodiment, thereby demonstrating the practicality of this model.
[0088] Most of the datasets used in the experiment are common datasets in current LDA-related papers, covering a variety of types, widely used, and authoritative verification significance. The specific information of the dataset is shown in Table 1, which gives the name of the dataset, the number of samples, the number of features, the number of categories, the proportion of the number of training samples to the total number of samples, and the type.
[0089] Table 1 Dataset information
[0090]
[0091] In order to verify the effectiveness of the LDA model of this embodiment, this embodiment selects some representative improved LDA algorithms in recent years, and verifies its effectiveness by observing the performance changes of the LDA model of this embodiment applied to these algorithms. The Maximum Margin Criterion (MMC) changes the trace ratio optimization problem into a trace difference optimization problem to solve the problem of small sample size. In this embodiment, the parameter μ in the algorithm is set to 10 -4 . Trace Ratio LDA (TRLDA) transforms the original LDA optimization problem into another equivalent optimization problem, and designs an iterative solution algorithm to solve it. Like the purpose of MMC, it is also to solve the problem of small sample size. Since its parameter α can be set to an arbitrary constant, this embodiment directly sets it to 1 and the number of iterations is set to 100. Under this setting, convergence can be basically achieved. The non-monotonicity of the optimization objective of Optimal Dimensionality LDA (ODLDA) enables it to automatically obtain the optimal dimensionality reduction dimension. For each dimensionality reduction dimension, the number of iterations is set to 100. Under this setting, convergence can be basically achieved. Robust Sparse Linear Discriminant Analysis (RobustSLDA) uses l for the mapping matrix in the optimization objective. 2,1 The constraint is added to solve the problem that LDA is sensitive to the choice of dimension reduction. At the same time, the constraint of minimum reconstruction error is added, and the reconstruction error is constrained by the matrix l1 norm. This allows the data after dimension reduction to not only maintain the main information, but also reduce the influence of noise. According to the parameter analysis part in the original paper, this embodiment sets the search range of its parameters λ1 and λ2 to [10 -4 ,10 -3 ,10 -2 ,10 -1 ,1] to achieve the best results. In addition, the parameter μ is set to 10 -4, ρ is 1.01, and the number of iterations is 50, which is consistent with the original paper. Under this setting, convergence can be basically achieved. Ratio Sum for Linear Discriminant Analysis (RSLDA) changes the original trace sum optimization problem into a trace ratio sum optimization problem to solve the shortcoming that LDA tends to select features with smaller variance. As in the original paper, this embodiment sets the parameter γ to tr(S W )×10(S W is the intra-class divergence matrix), the external iteration number is 100, and the internal iteration number is 50. Under this setting, convergence can be basically completed. Ratio Sum Minimization based Linear Discriminant Analysis (RSM-LDA) also solves the disadvantage that LDA tends to select features with smaller variance. Unlike RSLDA, RSM-LDA has been improved in the optimization target and a new solution algorithm has been proposed. This embodiment uses the solution method based on the greedy algorithm in the original text to solve it. Adaptive Local LDA (ALLDA) can automatically learn the weights in the adjacency matrix in the process of solving the optimization target, thereby introducing the local structure information of the sample to solve the disadvantage that LDA performs poorly on non-Gaussian complex distribution data. In this embodiment, the parameter r is set to 2, the search range of the number of neighbor samples K is set to [1,2,3,4,5], and the number of iterations is set to 50 times. In this case, convergence can be basically completed.
[0092] The parameters of the transposed projection are only the number of neighbor samples K, and the search range is set to [1, 2, 3, 4, 5] in this embodiment; the parameters of the kernel transposed projection are the Gaussian kernel parameter σ in addition to the number of neighbor samples K, where the search range of K is also set to [1, 2, 3, 4, 5], and the search range of σ is set to [10 -2 ,10 -1 ,1,10 1 ,10 2]. The parameter search method of the algorithm uses grid search, and the best results are recorded uniformly. In this embodiment, the mean and standard deviation of the classification accuracy of 10 random experiments are used as evaluation indicators. The training set of each experiment is composed of 20% of the samples randomly selected from the original data set, and the remaining samples constitute the test set. Of course, for the sake of fairness, the training set and test set used by all algorithms in each experiment of the 10 random experiments are exactly the same. Like the settings in other LDA-related papers, the classifier uses the nearest neighbor classifier, which can well evaluate the discriminative performance of the algorithm. The computer configuration used in the experiment is Intel (R) Core (TM) i7-12700F CPU, 64GB RAM, the operating system is Windows 10 20H2, and the software used is MATLAB R2019a.
[0093] One of the key points of the LDA model in this embodiment is to mine the correlation information between samples to form a new sample - the envelope sample. In order to verify the advantages of the envelope sample over the original sample, a sample visualization analysis is performed here. The experiment is conducted on two two-dimensional toy data sets, Gaussian and Moon, to visualize the distribution of envelope samples generated by transposed projection and kernel transposed projection under different numbers of neighbor samples, as shown in Figure 2. Figure 3 and Figure 4 As shown, the original sample distribution and the corresponding envelope sample distribution when the number of neighbor samples is 1 to 5 are shown. In this experiment, the parameter σ in the kernel transpose projection is fixed to 10.
[0094] From the visualization results of the Gaussian dataset, we can see that as K increases, the envelope samples generated by transposed projection and kernel transposed projection are more concentrated towards the center of the class, the number of heterogeneous overlapping samples decreases, and the classification boundaries become more obvious. From the visualization results of the Moon dataset, we can see that the envelope samples generated by transposed projection and kernel transposed projection can still roughly maintain the distribution shape of the original samples, that is, two crossed "moon" shapes, and as K increases, the "moons" become thinner and thinner, and the boundary between the two moons becomes more and more obvious. From the perspective of human vision, whether it is the Gaussian or Moon dataset, the separability of the envelope samples is significantly better than the original samples, which is more conducive to the subsequent dimensionality reduction projection. In other words, transposed projection and kernel transposed projection can better reduce the negative impact of low-quality samples such as noisy samples and outlier samples, thereby providing higher quality samples for subsequent LDA projection dimensionality reduction.
[0095] For samples far away from the class center or even in the heterogeneous sample area, they may be abnormal samples polluted by more noise, and their neighboring samples of the same category are often closer to the class center. At this time, the transposed projection and the kernel transposed projection can fuse the correlation information between them and their neighboring samples (the abnormal samples and the neighboring samples are both of the same category, and should be closer to the position of the neighboring samples), excavating their deep essential information, making them more "similar" to the samples of this category, that is, closer to their own class center, so there will be the above experimental results. This is also in line with human cognitive habits: when seeing those samples that are at the edge and distributed more dispersedly or even in the heterogeneous overlapping area, this embodiment will always subconsciously look for similar samples that are close to (similar to) them and have better position distribution and move these abnormal samples in their direction, so that the boundaries between the distribution areas of samples of different categories are more obvious.
[0096] Figure 3 , Figure 4 It has been verified that there is correlation information between similar samples. Extracting this correlation information and using it to transform samples will achieve good results to a certain extent.
[0097] In order to verify the effectiveness of the present invention, this embodiment further applies it to multiple representative LDA algorithms, and verifies its effect through a large number of comparative experiments conducted on multiple data sets. For each data set, this embodiment tests and records the classification accuracy of the same LDA in the original LDA mode (Ori), the transposed projection envelope LDA mode (Envelope), and the kernel transposed projection envelope LDA mode (Kenvelope) under different dimensionality reduction dimensions, as well as the accuracy of direct classification based on the original sample, the transposed projection envelope sample, and the kernel transposed projection envelope sample. Table 2 records the best classification accuracy and direct classification accuracy in all dimensionality reduction dimensions on the data sets Pima, Australian, Wisconsin, and German. None is direct classification without dimensionality reduction. The results are displayed in accuracy (%) ± standard deviation (%). The display in brackets is the dimension of the data set after dimensionality reduction when the best result is obtained. The font of the maximum value of the best classification accuracy of the same LDA in different modes on the same data set is bolded. Figure 5 , Figure 6 , Figure 7 , Figure 8 The classification accuracy-dimension line graphs of the same LDA in different modes and different dimensionality reduction dimensions on the Pima, Australian, Wisconsin, and German datasets are plotted respectively. Since None is a direct classification without dimensionality reduction, and ODLDA can automatically obtain the optimal dimensionality reduction dimension, there are no corresponding dimension records for the two. The corresponding contents in brackets in Table 2 are replaced by "-". Figures 5 to 8 In order to facilitate observation, a horizontal line is directly used instead, and the vertical axis of the horizontal line is its classification accuracy.
[0098] Table 2 The best classification accuracy of each LDA in different modes (accuracy % ± standard deviation %)
[0099]
[0100] From Table 2, we can see that the best classification accuracy of the same LDA in the transposed projection envelope LDA mode and the kernel transposed projection envelope LDA mode is always higher than the best classification accuracy in the original LDA mode; the improvement values on most data sets are basically in the range of 1% to 3%, which is quite obvious. Figures 5 to 8 It can be seen that in most cases, the classification accuracy of the same LDA in each dimension reduction dimension in the transposed projection envelope LDA mode and the kernel transposed projection envelope LDA mode is higher than the classification accuracy in the original LDA mode. As the dimension increases, the gap between them may be further reduced. Therefore, transposed projection and kernel transposed projection can indeed extract the correlation information between similar samples, making the generated envelope samples better suitable for LDA modeling, and the final classification accuracy is further improved, that is, the transposed projection envelope LDA mode and the kernel transposed projection envelope LDA mode are better than the original LDA mode.
[0101] From Table 2, we can see that the best classification accuracy of the same LDA in the kernel transposed projection envelope LDA mode is almost always higher than the best classification accuracy in the transposed projection envelope LDA mode. Figures 5 to 8 It can be seen that in most cases, the classification accuracy of the same LDA in all dimensionality reduction dimensions under the kernel transposed projection envelope LDA mode is higher than that under the transposed projection envelope LDA mode. Therefore, the kernel transposed projection is better than the transposed projection, which means that the nonlinear operation in the kernel transposed projection can indeed extract more correlation information (nonlinear correlation information), making the generated envelope samples more suitable for LDA modeling than the envelope samples generated using the transposed projection, that is, the kernel transposed projection envelope LDA mode is better than the transposed projection envelope LDA mode.
[0102] From Table 2, we can see that for the same data set, in the above three LDA modes, there is always a suitable LDA that makes its optimal classification accuracy greater than the corresponding direct classification accuracy, indicating that both the original sample and the envelope sample may have redundant features or feature noise, etc. Using a suitable LDA to eliminate or reduce them can effectively improve the classification accuracy. Figures 5 to 8It can be seen that in most cases, when the dimension is small, the overall trend of all broken lines is basically that the broken lines generally increase with the increase of the dimension. After reaching a maximum value, the overall trend of the broken lines is that the broken lines generally decrease with the increase of the dimension. The main reason is that when the dimension reduction dimension is too small, there is less useful information. As the dimension increases, the useful information begins to increase and the accuracy increases. When it reaches a certain level, too many dimensions will introduce redundant features and noise, and the accuracy begins to decline continuously.
[0103] The only parameter that needs to be adjusted for the transposed projection is the number of neighbor samples K, and the parameters that need to be adjusted for the kernel transposed projection are the number of neighbor samples K and the Gaussian kernel parameter σ. Here, this embodiment studies their impact on classification accuracy. The experiment was conducted on the German, Sonar, Control and Wdbc data sets. The remaining parameters remained fixed during the experiment. LDA selected TRLDA, RobSLDA, RSLDA and ALLDA.
[0104] Fig. 9 This is the effect of the transposed projection parameter K on the classification accuracy. Fig. 9 It can be seen that in the above four data sets, the classification accuracy when K is between 1 and 5 is higher than the classification accuracy when K = 0 (without transposed projection), which is consistent with the previous experimental results, indicating that the envelope samples generated by transposed projection are more suitable for LDA. In practical applications, K needs to be adjusted to achieve the best results.
[0105] Fig.10 , Fig.11 , Fig.12 , Fig.13 The following are the effects of kernel transposition projection parameters K and σ on classification accuracy on German, Sonar, Control, and Wdbc datasets. Figures 10 to 13 It can be seen that when K>0, for the same K value, σ takes [10 -2 ,10 -1 ] is lower than when σ is [1, 10 1 ,10 2 ], so in practical applications, the lower bound of the search range of parameter σ should be no less than 1. Different data sets are also sensitive to parameters K and σ. For example, in the Wdbc data set, when σ is [1, 10 1 ,10 2 ], K>0, the classification accuracy obtained under different K values is similar, and the algorithm has good stability of parameters. For the above four data sets, when K>0 and σ is [1, 10 1 ,10 2], its classification accuracy is basically higher than that when K = 0 (no kernel transposition projection is used), indicating that the envelope samples generated by the kernel transposition projection with appropriate parameter settings are more suitable for LDA. In practical applications, the parameters K and σ need to be adjusted to achieve the best results.
[0106] The transposed projection and kernel transposed projection envelope transformation proposed in this embodiment can be used to construct envelope samples. Here, their time cost is tested to verify their practicality. The time taken by the transposed projection and kernel transposed projection to convert all samples in the Pima, Australian, Wisconsin, and German datasets into envelope samples under different numbers of neighbor samples is counted. Since the difference in the parameter σ of the kernel transposed projection hardly affects the calculation time, it is directly set to a fixed value in all tests. All experiments were repeated ten times, and the mean and standard deviation of the corresponding calculation time of the 10 experiments were recorded. The experimental results are shown in Table 3.
[0107] Table 3 Transposition projection and core transposition projection transformation sample consumption time (unit: second)
[0108]
[0109] It can be seen from the experimental data in Table 3 that for the same data set, the different values of K do not have a very obvious effect on the time consumed by the transposed projection and the nuclear transposed projection to transform the data set into a new data set. For the transposed projection, the time consumed increases with the increase of K, which is in line with expectations. However, for the nuclear transposed projection, the time consumed on some data sets does not necessarily increase with the increase of K. This embodiment speculates that this may be related to the eigenvalue decomposition step in the algorithm. Although K is increasing, the amount of calculation of the remaining steps will increase and the time consumption will increase, but the eigenvalue decomposition algorithm may complete the convergence of the algorithm with fewer iterations when K is large, which reduces the total time consumed. Of course, under the same K value of the same data set, the time consumed by the nuclear transposed projection is much higher than that of the transposed projection. In short, both envelope transformation algorithms can meet the needs of real applications. For scenarios with high accuracy requirements but low time cost requirements, this embodiment recommends using nuclear transposed projection to construct envelope samples, corresponding to the nuclear transposed projection envelope LDA mode. If there are certain requirements for time cost, you can consider using transposed projection to construct envelope samples, corresponding to the transposed projection envelope LDA mode.
[0110] In summary, the existing LDA model is based on the modeling of the original sample individuals themselves, and does not consider the correlation information between similar samples. When the sample separability is low and the noise is high, its performance is often poor. In response to this problem, this embodiment proposes the idea of mining the correlation information between samples for subsequent projection dimensionality reduction, thereby designing a kernel transposition projection envelope LDA data dimensionality reduction method and system (referred to as the kernel transposition projection envelope LDA model). First, this embodiment designs a transposed projection envelope transformation algorithm and kernelizes it to obtain a better kernel transposition projection envelope transformation algorithm, which is used to transform the original sample to generate an envelope sample, and the correlation information between similar samples is mined as much as possible. Secondly, the envelope sample is loaded to the LDA input, thereby realizing the projection dimensionality reduction on the correlation information between similar samples. The experimental part uses 7 data sets and 7 representative LDA algorithms to verify the effectiveness of the LDA model of this embodiment. The results show that after the introduction of the kernel transposed projection envelope LDA model, the classification accuracy of various LDA algorithms has been significantly improved, which shows that the kernel transposed projection envelope LDA model is better than the original LDA model, and realizes the projection dimension reduction on the correlation information between similar samples, making up for the defect of the original LDA model that ignores or destroys the correlation information between similar samples during the modeling process. In addition, the kernel transposed projection envelope LDA model is not a specific LDA improved algorithm, so it has good universality.
[0111] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. Kernel transposition projected envelope LDA data dimensionality reduction method, characterized in that: The method comprises the steps of: Perform envelope transformation on each original sample using kernel transposed projection envelope transformation to obtain corresponding envelope samples; the kernel transposed projection envelope transformation is to introduce kernel method into transposed projection for kernelization; the transposed projection obtains similar samples of the original sample by neighbor distance measurement, forms sample envelope with the original sample and its similar samples, extracts correlation information by linear weighting of samples inside the sample envelope and fuses it with the original sample to obtain envelope samples corresponding to each original sample; The LDA method based on envelope sample modeling is used to reduce the dimension of all envelope samples to obtain envelope samples after dimension reduction.
2. The kernel transposed projected envelope LDA data dimensionality reduction method according to claim 1, characterized in that: The original sample set is represented as X = [x1, x2, ..., x n ]∈R d×n , where d is the number of features of each original sample, n is the total number of samples, and the i-th original sample x i Sample x similar to its K nearest neighbors i,1 ,...,x i,K Form x i The sample envelope E i =[x i ,x i,1 ,...,x i,K ]∈R d ×(K+1) , i = 1:n; the purpose of transpose projection is to find a mapping vector p∈R (K+1)×1 , so that E i After mapping E i Generate sample x after p i The corresponding envelope sample 3. The kernel transposed projected envelope LDA data dimensionality reduction method according to claim 2, characterized in that: The transposed projection finds the mapping vector p∈R (K+1)×1 The specific process is: Based on the sample envelope E i Constructing the Matrix and centralize it Custom 1 (n×d)×1 represents a matrix of all 1s of size (n×d)×1, and the superscript T represents the matrix transpose; right Perform eigenvalue decomposition, take the eigenvector corresponding to the maximum eigenvalue and record it as p.
4. The kernel transposed projected envelope LDA data dimensionality reduction method according to claim 3, characterized in that: The kernel method is introduced in the transposed projection for kernelization, specifically: E i Each row (E i ) j,: After nonlinear mapping, φ(·) becomes φ((E i ) j,: T ) T Then follow the subsequent steps of transposed projection to model.
5. The kernel transposed projected envelope LDA data dimensionality reduction method according to claim 4, characterized in that: The purpose of the kernel transposition projection envelope transform is to find a mapping vector α such that E i After α mapping, the sample x is generated i The corresponding envelope sample Superscript right It means that each row of the matrix is mapped to the feature space through a nonlinear mapping φ(·).
6. The kernel transposed projected envelope LDA data dimensionality reduction method according to claim 5, characterized in that: The process of finding the mapping vector α is as follows: Based on the sample envelope E i Constructing the Matrix And calculate the kernel matrix through the Gaussian kernel function calculate Take the generalized eigenvalue decomposition problem The eigenvector corresponding to the maximum generalized eigenvalue is α, and λ represents the generalized eigenvalue.
7. The kernel transposed projected envelope LDA data dimensionality reduction method according to claim 6, characterized in that: Input the envelope sample after dimension reduction into the classifier to obtain the classification label, thereby obtaining the classification label of the corresponding original sample; In the process of using the classifier for prediction, the envelope sample is generated Generate the sample x to be predicted in the same way test The corresponding envelope sample 8. The kernel transposed projected envelope LDA data dimensionality reduction method according to claim 7, characterized in that: The loss function used in the process of inputting the reduced dimension envelope sample into the classifier for training is: Represents the new training set consisting of all envelope samples, W∈R d×d′ represents the dimension reduction matrix used by the linear discriminant analysis method, d′ is the sample dimension after dimension reduction, Error_term(·) represents the error term in the original LDA optimization objective, and Regular_term(·) represents the regular term in the original LDA optimization objective.
9. Kernel transposition projected envelope LDA data dimensionality reduction system, characterized by: The system applies the kernel transposition projection envelope LDA data dimension reduction method described in any one of claims 1 to 8, and is provided with a kernel transposition projection envelope transformation module and an LDA module. The kernel transposition projection envelope transformation module is used to use the kernel transposition projection envelope transformation to perform envelope transformation on each original sample to obtain the corresponding envelope sample, and the LDA module is used to use the LDA method based on envelope sample modeling to reduce the dimension of all envelope samples to obtain the envelope samples after dimension reduction.