A weighted collaborative representation based on complementary subspace for imbalanced classification method
By introducing the regular terms of complement space and adaptive weights in the CRC method, a weighted collaborative representation classification model is established, which solves the problem of inaccurate classification of a few classes of samples on the uneven data sets of CRC, and achieves higher classification accuracy.
Patent Information
- Application Number
- CN202210356037.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-04-06
AI Technical Summary
The existing collaborative representation classification method (CRC) cannot accurately classify a few class samples on uneven data sets, especially when the category distribution is uneven, the classification accuracy is low.
By introducing regular terms of the complement space and applying weights to each class of samples, a weighted coordinated representation classification model based on the complement space is established, and the optimal representation coefficient reconstruction error is used for classification.
It significantly improves the classification accuracy of a few types of samples and improves the overall classification performance on unbalanced data sets.
Smart Images

Figure CN114742149B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of imbalanced data classification, and particularly to a weighted collaborative representation imbalanced classification method based on complementary subspaces. Background Art
[0002] Classification problems of imbalanced data sets are widely involved in application fields such as disease diagnosis, fault detection, and information security. The recognition rate of existing classification technologies for the minority class is much lower than that of the majority class. Especially for severely imbalanced data, the recognition accuracy of the minority class is even 0. In fact, it is particularly important to achieve accurate classification for the minority class. Taking disease diagnosis as an example, the cost of misdiagnosing critically ill patients as healthy people is much higher than that of misdiagnosing healthy people as critically ill patients. Therefore, it is crucial to design an efficient and accurate imbalanced classification method.
[0003] Existing imbalanced classification methods are generally divided into two major types: data-level-based and algorithm-level-based. The core idea of data-level-based is to achieve the balance of class distribution through sampling techniques. However, this technique will destroy the relationship between the original data, thus limiting its development and application. The present invention is committed to proposing a new algorithm-level-based imbalanced classification method. Among many such methods, the collaborative representation based classification (CRC) method has been widely applied to various classification fields due to its simplicity, high efficiency, easy operation, and low complexity. However, its success largely depends on the class distribution, and the imbalance of the class distribution will seriously affect its classification accuracy. Therefore, the CRC method and its variants still cannot achieve accurate recognition for minority class samples.
[0004] In summary, there is no effective solution to the imbalanced data classification problem. Considering the great advantages of the CRC method in theory and operation, the present invention selects it as the basic model, and on the basis of continuing its existing advantages, focuses on making up for its deficiencies in classifying the minority class. Summary of the Invention
[0005] The purpose of the present invention is to provide a weighted collaborative representation imbalanced classification method based on complementary subspaces to solve the problem that CRC in the prior art cannot accurately classify the minority class on imbalanced data sets.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] A weighted collaborative representation imbalanced classification method based on complementary subspaces, including the following:
[0008] Preprocess an imbalanced data set with a total of c classes to obtain a target training sample set X and a target test sample set;
[0009] Obtain the weights of various training sample subsets in the target training sample set X;
[0010] Determine a weighted collaborative representation classification model based on the complementary subspace using the weights of various training sample subsets;
[0011] Solve the weighted collaborative representation classification model to obtain the optimal representation coefficients, and predict the category of the target test sample set according to the reconstruction error of the optimal representation coefficients.
[0012] Preferably, the step of preprocessing an imbalanced data set with a total number of categories c is:
[0013] Randomly divide the imbalanced data set into n parts using the cross-validation method, where a parts are used as the original test set and b parts are used as the original training set; a + b = n; perform random cross-validation on the original test set and the original training set m times to obtain m groups of training sample sets and m groups of test sample sets;
[0014] Convert the m groups of training sample sets and the m groups of test sample sets into column vector data respectively, and perform normalization processing to obtain the target training sample set X and the target test sample set.
[0015] Preferably, the method for obtaining the weights of various training sample subsets in X is:
[0016] Obtain the mean E(X i ) and variance D(X i ) of the i-th type of training sample subset;
[0017] Determine the class j where the minimum variance is located, and assign a weight to the j-th type of training sample subset, let ω j = 1;
[0018] Determine the remaining training sample subset class weights ω i based on the correlation between the training sample subset classes;
[0019] Among them, the expression of j is: j(i) = D(X i );
[0020] Among them, represents the value of the variable i when the function j(i) takes the minimum value.
[0021] Preferably, the method for determining the remaining training sample subset class weights ω i based on the correlation between the training sample subset classes is:
[0022] When i ≠ j is determined, the mean E(X i) and the mean E(X j ) of the j-th subset of training samples, the correlation coefficient ρ i ;
[0023] According to the correlation coefficient ρ i assign a value to ω i , and the assignment expression is: ω i = 1 + σ|ρ i |;
[0024] where σ represents the scaling parameter for regulating the weight.
[0025] Preferably, the expression of the complementary subspace is:
[0026] S - S i = span{X -i};
[0027] where i ≤ c, representing the i-th class; S represents the entire space generated by the target training sample set X; X i represents the i-th subset of training samples; S i represents the subspace of S generated by X i ; S - S i represents the complementary subspace of S i ; X -i represents the set of remaining training samples after removing the i-th subset of training samples from X.
[0028] Preferably, the expression of the weighted collaborative representation classification model is:
[0029]
[0030] where y represents a test sample vector of size d*1, d represents the dimension of the vector, M represents a training sample matrix of size d*J, J represents the total number of training samples, λ and γ are regularization parameters, a is the representation coefficient of M, a * is the optimal representation coefficient, M -i is the matrix corresponding to X -i , a -i is the representation coefficient of M -i .
[0031] Preferably, the method for predicting the category of the target test sample set according to the error reconstructed from the optimal representation coefficient is:
[0032] Classify the test samples using the minimum reconstruction error criterion:
[0033]
[0034] where label(y) is the category of the test sample vector y, Mi It represents the matrix corresponding to the i-th class of training sample subsets. It represents the coefficient a * The corresponding M i The optimal representation coefficient sub-vector on
[0035] Preferably, it further includes: evaluating the performance of the classification model according to the prediction result.
[0036] From the above content, it can be seen that compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] The present invention introduces a more discriminative complementary subspace regular term into the collaborative representation modeling process, and adaptively obtains the weight of each class according to the original class distribution information of the imbalanced data set, thereby giving a greater weight to the minority class and effectively solving the problem that the CRC method cannot correctly classify the minority class. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0039] Figure 1 It is a classification instance diagram of the existing CRC method on the imbalanced data set.
[0040] Figure 2 It is a flowchart of the present invention.
[0041] Figure 3 It is a classification instance diagram of the present invention on the imbalanced data set.
[0042] Figure 4 It is a classification comparison diagram between the present invention and CRC. Detailed Embodiments
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment 1
[0045] This embodiment discloses a weighted collaborative representation imbalanced classification method based on complementary subspaces, including the following content:
[0046] Preprocess the imbalanced dataset with a total of c categories to obtain the target training sample set X and the target test sample set; specifically: use the cross-validation method to randomly divide the imbalanced dataset into n parts, where a parts are used as the original test set and b parts are used as the original training set; a + b = n; perform random cross-validation on the original test set and the original training set m times to obtain m groups of training sample sets and m groups of test sample sets; convert the m groups of training sample sets and m groups of test sample sets into column vector data respectively, and perform normalization processing to obtain the target training sample set X and the target test sample set.
[0047] Obtain the weights of each type of training sample subset in the target training sample set X; specifically: obtain the mean E(X i ) and variance D(X i ) of the i-th type of training sample subset; determine the class j where the minimum variance is located, and assign a weight to the j-th type of training sample subset, let ω j = 1; determine the weights ω i of the remaining training sample subsets based on the correlation between the training sample subset classes; among them, the expression of j is: j(i) = D(X i ); among them, represents the value of the variable i when the function j(i) takes the minimum value. Among them, the method for determining the weights ω i of the remaining training sample subsets based on the correlation between the training sample subset classes is: when i ≠ j, determine the correlation coefficient ρ i between the mean E(X j ) of the i-th type of training sample subset and the mean E(X i ) of the j-th type of training sample subset; assign a value to ω i according to the correlation coefficient ρ i , and the assignment expression is: ω i = 1 + σ|ρ i |; where σ represents the scaling parameter for adjusting the weights.
[0048] Use the weights of each type of training sample subset to determine the weighted collaborative representation classification model based on the complementary subspace; specifically: the expression of the complementary subspace is: S - S i = span{X -i} where i ≤ c, representing the i-th type; S represents the full space generated by the target training sample set X; X i represents the i-th type of training sample subset; S i represents the subspace of S generated by X i ; S - S i represents the complementary subspace of S i ; X -iDenote the set of remaining training samples after removing the subset of the \(i\)-th class of training samples from \(X\). The expression of the weighted collaborative representation classification model is: where \(y\) represents a test sample vector of size \(d\times1\), \(d\) represents the dimension of the vector, \(M\) represents a training sample matrix of size \(d\times J\), \(J\) represents the total number of training samples, \(\lambda\) and \(\gamma\) are regularization parameters, \(a\) is the representation coefficient of \(M\), and \(a\) * is the optimal representation coefficient, and \(M\) -i is the matrix corresponding to \(X\) -i , and \(a\) -i is the representation coefficient of \(M\) -i .
[0049] Solve the weighted collaborative representation classification model to obtain the optimal representation coefficient, and predict the class of the target test sample set according to the reconstruction error of the optimal representation coefficient; specifically: classify the test samples using the minimum reconstruction error criterion: where \(label(y)\) is the class of the test sample vector \(y\), and \(M\) i represents the matrix corresponding to the subset of the \(i\)-th class of training samples, and the subvector of the optimal representation coefficient on \(M\) * corresponding to the representation coefficient \(a\) i .
[0050] Furthermore, the following content is specifically disclosed in this Embodiment 1: The specific solution method for obtaining the optimal representation coefficient is:
[0051] Set the candidate sets of the parameters \(\lambda\), \(\gamma\), and \(\sigma\) as \(\{10\) -5 , \(10\) -4 , \(10\) -3 , \(10\) -2 , \(10\) -1 , \(1, 10, 10\) 2 , \(10\) 3 \}, and use the grid search algorithm to linearly combine these parameters so that the weighted collaborative representation classification model obtains the best classification effect on each data set.
[0052] Calculate the value of \(a\) corresponding to when the partial derivative of the model with respect to \(a\) is zero. Specifically, let the partial derivative of the model with respect to \(a\) be zero, and we get:
[0053] Then the optimal representation coefficient \(a\) * of this model is a vector of size \(J\times1\), that is: \(a\) * = Py
[0054] where is a projection matrix of size \(J\times J\), denotes a matrix of size d*J obtained by setting the column vector of the i-th class training samples in the training sample matrix M to the zero vector, and I is the identity matrix of size J*J. The projection matrix P is only related to the training sample matrix M and has nothing to do with the test sample y. Therefore, before calculating a * save P first. Once the test sample y is input, then use M -i and the class weight ω i to further obtain the optimal representation coefficient a of J*1 using the fifth step * .
[0055] Embodiment 2
[0056] On the basis of Embodiment 1, it further includes: evaluating the performance of the weighted collaborative representation classification model based on the complementary subspace according to the classification result of Embodiment 1. The specific method is as follows:
[0057] For binary classification problems, the metrics used are F-measure and G-mean. Specifically:
[0058]
[0059] where represents the recall rate in binary classification problems; represents the precision in binary classification problems; TP represents the number of samples in the minority class that are correctly classified; FP represents the number of samples in the majority class that are misclassified; FN represents the number of samples in the minority class that are misclassified; TN represents the number of samples in the majority class that are correctly classified;
[0060] For multi-class classification problems, calculate the macro-averages F-macro and G-macro of the metrics as follows:
[0061]
[0062] where represents the recall rate in multi-class classification problems; represents the precision in multi-class classification problems.
[0063] In this embodiment, the number of random cross-validations is set to 10, and the average value of the metrics of the test sample set is taken; by comparing the metric means of existing imbalanced classification algorithms under the same experimental conditions, the classification performance of the present invention for imbalanced data sets is evaluated. Specifically:
[0064] Nine binary datasets and five multi-class imbalanced datasets from the UCI public database are selected to verify the effectiveness of the method of the present invention. Among them, the binary datasets are Vehicle1, Newthyroid1, Glass6, Yeast3, Ecoli3, Shuttle0, Glass5, Yeast6, Abalone19, and the multi-class datasets are Penbased, Wine, Newthyroid, Dermatology, Ecoli. Table 1 details the characteristics of these datasets. Among them, the imbalanced rate (IR) is an important indicator to measure the class distribution, and it is usually expressed as the ratio of the largest number of classes to the smallest number of classes. According to Table 1, the imbalance rates IR of the selected imbalanced datasets vary widely, ranging from 1.10 to 129.44. According to the common general knowledge in this field, the higher the IR, the greater the difficulty of correct identification. Therefore, the datasets used are very challenging.
[0065]
[0066] Table 1 Details of Nine Binary Datasets and Five Multi-class Imbalanced Datasets
[0067] In addition, in this embodiment, the average classification results of the metric indicators of the mildly imbalanced datasets with IR ≤ 10 and the severely imbalanced datasets with IR > 10 in the selected datasets are also calculated respectively. The average classification comparison results between the method of the present invention and other imbalanced classification methods under the two degrees of imbalance are given, as shown in Table 2 specifically.
[0068]
[0069] Table 2 Comparison between the Method of the Present Invention and Other Imbalanced Classification Methods
[0070] According to the content shown in Table 2, it can be seen that the classification performance of the method of the present invention is much higher than that of other imbalanced classification methods whether on mildly imbalanced or severely imbalanced datasets.
[0071] Such as Figure 1As shown, the two confusion matrices respectively show the classification results of CRC on the imbalanced datasets Yeast3 and Penbased. Among them, Yeast3 is a two-class dataset with a class distribution of 163:1321, and Penbased is a multi-class dataset with a class distribution of 115:114:114:105:114:105:115:106:105:106. Then, the IR values of Yeast3 and Penbased are 8.10 and 1.10 respectively. The diagonal elements of the confusion matrix show the accuracy of each class of test sample sets in the dataset being correctly classified. The closer the diagonal element is to 1, the higher the classification accuracy. The two confusion matrix diagrams clearly show that the classification accuracy of CRC on the minority classes is much lower than that on the majority classes. This fully demonstrates that CRC cannot accurately identify the minority classes when dealing with imbalanced problems. The main reason is that CRC regards the entire training sample set as a whole and does not consider the differences in sample classes.
[0072] To more clearly and intuitively illustrate the classification performance of the method of the present invention, Figure 3 the two confusion matrices respectively show the classification results of the present invention on the imbalanced datasets Yeast3 and Penbased. Compared with Figure 1 , it can be clearly seen that the method of the present invention can greatly improve the classification accuracy of the minority classes. Figure 4 is the histogram of the reconstruction error of the present invention and the CRC method on the first-class test samples of the balanced dataset ORL. It is not difficult to find that both methods have achieved the correct classification of the first-class test samples. However, compared with CRC, the reconstruction error of the method of the present invention on the first class is significantly smaller than that of other classes. That is to say, the method of the present invention is more discriminative.
[0073] In summary, the technical method provided by the present invention preprocesses the training samples and test samples of the imbalanced dataset, calculates the complementary subspace of each class of training sample space, adaptively obtains the weights of each class of training sample sets, incorporates the complementary subspace and the weights of each class into the CRC model to establish a weighted collaborative representation classification model based on the complementary subspace, and solves the problem that the CRC method cannot correctly classify the minority classes by the technical means of solving the classification model to classify the test samples.
[0074] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A weighted collaborative representation based unbalanced classification method using complementary subspaces, characterized in that It includes the following contents: Preprocess the imbalanced dataset with a total of c categories to obtain a target training sample set X and a target test sample set; Obtain the weights of each type of training sample subset in the target training sample set X; Determine a weighted collaborative representation classification model based on the complementary subspace using the weights of each type of training sample subset; Solve the weighted collaborative representation classification model to obtain the optimal representation coefficients, and predict the categories of the target test sample set based on the reconstruction error of the optimal representation coefficients; The method for obtaining the weights of each type of training sample subset in X is as follows: Obtain the mean E(Xi) and variance D(Xi) of the i-th type of training sample subset; Determine the class j where the minimum variance is located, and assign weights to the training sample subset of the j-th class, and let ω j = 1; Determine the class weight ω of the remaining training sample subsets based on the inter-class correlation of the training sample subsets i ; Among them, the expression of j is: j(i) = D(X i ); Among them, represents the value of variable i when the function j(i) takes the minimum value; Method for determining the class weight ω of the remaining training sample subsets based on the inter-class correlation of the training sample subsets i is as follows: Determine the correlation coefficient ρ between the mean E(Xi) of the subset of training samples of the i-th class and the mean E(Xj) of the subset of training samples of the j-th class when i ≠ j i ; According to the correlation coefficient ρ i assign a value to ω i The assignment expression is: ω i = 1 + σ|ρ i |; where σ represents the scaling parameter for regulating the weights; The expression of the complementary subspace is: S-S i = span{X -i}; where i ≤ c, representing the i-th type; S represents the full space generated by the target training sample set X; Xi represents the i-th type of training sample subset; Si represents the subspace of S generated by Xi; S - Si represents the complementary subspace of Si; X-i represents the remaining training sample set after removing the i-th type of training sample subset from X; The specific solution method for obtaining the optimal representation coefficients is: Set the candidate sets of parameters λ, γ, and σ as {10 -5 , 10 -4 , 10 -3 , 10 -2 , 10 -1 , 1, 10, 10 2 , 10 3}, and use the grid search algorithm to linearly combine these parameters so that the weighted collaborative representation classification model achieves the best classification effect on each dataset; The value of a corresponding to when the partial derivative of the calculation model with respect to a is zero. Specifically, let the partial derivative of the model with respect to a be zero, and we get: Then the optimal representation coefficient a of the model * is a vector of J*1, that is: a * = Py; Among them, is a projection matrix of size J*J, represents a matrix of size d*J obtained by setting the column vector of the i-th class training samples in the training sample matrix M to the zero vector, and I is the identity matrix of size J*J; the projection matrix P is only related to the training sample matrix M and has nothing to do with the test sample y; so save P before calculating a * ; once the test sample y is input, then use M -i and the class weight ω i to obtain the optimal representation coefficient a of size J*1 * ; The expression of the weighted collaborative representation classification model is: Among them, y represents a test sample vector of size d*1, d represents the dimension of the vector, M represents a training sample matrix of size d*J, J represents the total number of training samples, λ and γ are regularization parameters, a is the representation coefficient of M, and a * is the optimal representation coefficient, and M -i is the matrix corresponding to X-i, and a-i is the representation coefficient of M -i .
2. The weighted collaborative representation based unbalanced classification method using complementary subspaces according to claim 1, wherein The steps for preprocessing the imbalanced dataset with a total of c categories are: Use the cross-validation method to randomly divide the imbalanced dataset into n parts, where a parts are used as the original test set and b parts are used as the original training set; a + b = n; perform random cross-validation on the original test set and the original training set m times to obtain m groups of training sample sets and m groups of test sample sets; Convert the m groups of training sample sets and m groups of test sample sets into column vector data respectively, and perform normalization processing to obtain the target training sample set X and the target test sample set.
3. A weighted collaborative representation based non-equilibrium classification method using complementary subspaces according to claim 1, characterized in that, The method for predicting the categories of the target test sample set based on the reconstruction error of the optimal representation coefficients is: Classify the test samples using the minimum reconstruction error criterion; Among them, label(y) is the category of the test sample vector y, and M i represents the matrix corresponding to the subset of the i-th class of training samples, represents the coefficient a * corresponding to M i the optimal representation coefficient sub-vector on it.
4. A weighted collaborative representation based non-equilibrium classification method using complementary subspaces according to claim 1, characterized in that It also includes: Evaluate the performance of the classification model according to the prediction results.
Citation Information
Patent Citations
Competition and collaborative representation method and system for face or scene recognition data classification
CN109299684A