Basic learner weight distribution method in ensemble learning

By using shrink robust Tyler-M estimator and principal component analysis in ensemble learning, the performance degradation problem of principal component regression algorithm on high-dimensional and non-Gaussian data is solved and applied to classification tasks, achieving higher prediction accuracy and robustness.

CN120012960APending Publication Date: 2025-05-16SHENZHEN TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510092472.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing principal component regression algorithms have degraded performance in high-dimensional data, finite sample size data, or non-Gaussian distributed data and are not suitable for classification tasks.

Method used

Multiple base learners of integrated learning model are used to predict the sample data set, and a shrinkage robust Tyler-M estimator is combined with the Tyler-M estimator and the linear shrinkage estimator is built to determine the optimal shrinkage parameters and optimal principal components through cross-validation. The optimal weight of the base learner is calculated using principal component analysis method or two-dimensional principal component analysis method.

Benefits of technology

The regression accuracy of the integrated regressor on high-dimensional, finite sample and non-Gaussian data is improved, and extended to classification tasks, improving the classification accuracy of the integrated classifier.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012960A_ABST
    Figure CN120012960A_ABST
Patent Text Reader

Abstract

The invention discloses a base learner weight distribution method in ensemble learning, and the method comprises the steps: constructing a shrinkage robust Tyler-M estimator in combination with a Tyler-M estimator and a linear shrinkage estimator, and introducing a diagonal matrix into the shrinkage robust Tyler-M estimator when the method is applied to a classification task; determining an optimal shrinkage parameter and an optimal main component number by adopting a cross validation method; based on a shrinkage robust Tyler-M estimator or a shrinkage robust Tyler-M estimator introducing a diagonal matrix, the optimal shrinkage parameter and the optimal principal component number, PCA or 2DPCA is adopted to carry out principal component analysis on a sample data set, and then the optimal weight of a base learning device is calculated. According to the method, a novel covariance matrix estimator having robustness for high-dimensional data, limited sample size data and non-Gaussian distribution data is constructed, the regression accuracy of an integrated regression device is improved, the method is suitable for classification tasks, and the classification precision of an integrated classifier is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ensemble learning, and in particular to a method for allocating weights of base learners in ensemble learning. Background Art

[0002] Basic ensemble methods usually use unweighted averaging of the prediction results of multiple learners to reduce estimation errors. However, since learners often have high correlations, the effectiveness of this method in practical applications is often limited. To solve this problem, generalized ensemble methods use the sample covariance matrix of the learner prediction results to optimize the weights, thereby further reducing the estimation error. However, generalized ensemble methods may show instability when dealing with multicollinearity problems, affecting their actual effectiveness.

[0003] To address the problem of multicollinearity, the principal component regression method (PCR) uses principal component analysis (PCA) to reduce the impact of collinearity. In PCR, the covariance matrix of the regressor needs to be estimated by the predicted value of each regressor, and the sample covariance matrix is ​​used here. However, when the data dimension is high, the number of samples is limited, or the data distribution is non-Gaussian, the accuracy of the sample covariance matrix in estimating the true covariance matrix decreases, which affects the performance of PCR and limits its effectiveness in complex application scenarios. In addition, PCR is an ensemble learning algorithm designed for regression tasks and cannot be applied to classification tasks.

[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention

[0005] The main purpose of the present invention is to provide a method for weighting base learners in ensemble learning, aiming to solve the problem that the performance of the existing principal component regression algorithm degrades in the case of high-dimensional data, limited sample size data or non-Gaussian distribution data, and is not applicable to classification tasks.

[0006] To achieve the above-mentioned object of the invention, the present invention provides a method for allocating weights of base learners in ensemble learning, wherein the method for allocating weights of base learners in ensemble learning is applied to regression tasks, and the method for allocating weights of base learners in ensemble learning comprises:

[0007] Predicting the sample data set using multiple base learners of the ensemble learning model to obtain a first prediction matrix;

[0008] Combining the Tyler-M estimator and the linear shrinkage estimator, a shrinkage robust Tyler-M estimator is constructed;

[0009] The sample data set is divided into a training set and a test set by a cross-validation method, the overall mean square error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall mean square error among multiple overall mean square errors are determined as the optimal shrinkage parameter and the optimal number of principal components;

[0010] Based on the first prediction matrix, the shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, the principal component analysis method is used to perform principal component analysis on the sample data set, and according to the principal component analysis results, the least squares method is used to calculate the optimal weights of multiple base learners.

[0011] Optionally, the method of using multiple base learners of the ensemble learning model to predict the sample data set to obtain a first prediction matrix specifically includes:

[0012] Using multiple base learners of the ensemble learning model to train the sample data set (x i ,y i ) to make predictions and obtain the first prediction matrix A∈R n×N , A ij =f j (x i );

[0013] Among them, n is the number of samples, N is the number of base learners, R represents a real number matrix, x i represents the i-th sample, y i represents the i-th sample x i The true value of A ij represents the value of the i-th row and j-th column in the first prediction matrix A, f j (x i ) represents the j-th base learner for the i-th sample x i The predicted value of .

[0014] Optionally, the combining of the Tyler-M estimator and the linear shrinkage estimator to construct a shrinkage robust Tyler-M estimator specifically includes:

[0015] Combining the Tyler-M estimator and the linear shrinkage estimator, a shrinkage robust Tyler-M estimator is constructed, and according to the first prediction matrix, the first covariance matrix is ​​obtained by using the shrinkage robust Tyler-M estimator:

[0016]

[0017] in, represents the first covariance matrix corresponding to the shrinkage parameter ρ, express The reciprocal of f(x i ) represents the sample x iThe predicted value, x i′ represents the i′th sample, f(x i′ ) represents the sample x i′ The predicted value of represents the decentralized f(x i ), T represents matrix transpose, I N Represents the identity matrix.

[0018] Optionally, the cross-validation method is used to divide the sample data set into a training set and a test set, calculate the overall mean square error corresponding to the test set, and determine the shrinkage parameter and the number of principal components corresponding to the minimum overall mean square error among multiple overall mean square errors as the optimal shrinkage parameter and the optimal number of principal components, specifically including:

[0019] Using a cross-validation method, the sample data set is divided into P partitions, each of the partitions is used as a test set in turn, and the remaining partitions are used as training sets, to obtain different combinations of training sets and test sets;

[0020] For each candidate shrinkage parameter ρ and principal component number K, the principal component analysis method is used to perform principal component analysis on the training set to obtain the training set projection matrix G of K principal components. train ∈R (n-s)×N , where s represents the number of test set samples;

[0021] Perform linear regression weight calculation according to the training set projection matrix to obtain a training set principal component weight vector and a training set model weight vector w′ corresponding to the training set;

[0022] According to the training set model weight vector w′, the test set prediction matrix A corresponding to the test set is test ∈R s ×N Perform weight optimization to obtain the test set prediction matrix after weight optimization

[0023]

[0024] According to the test set prediction matrix A test And the test set prediction matrix after the weight optimization Calculate the mean square error (MSE) of the test set under the current partition partition :

[0025]

[0026] y test,i ∈A test ;

[0027]

[0028] Among them, partition represents partition, y test,i is the predicted value of the i-th sample in the test set, is the weight-optimized predicted value of the i-th sample in the test set;

[0029] For each principal component number K, the mean square error of all partitioned test sets is cumulatively summed to obtain the overall mean square error MSE(ρ,K):

[0030]

[0031] Among them, MSE partition,p represents the mean square error of the test set under the p-th partition, P represents the total number of partitions, and p represents the p-th partition;

[0032] For each shrinkage parameter ρ, record the number of principal components K that minimizes the overall mean square error MSE(ρ, K) * , obtain the intermediate parameter combination (ρ, K * );

[0033] From all intermediate parameter combinations (ρ, K * ) is selected so that the overall mean square error MSE(ρ, K * )The minimum corresponding intermediate parameter combination is taken as the optimal parameter combination (ρ * , K * ):

[0034]

[0035] Among them, ρ * represents the optimal shrinkage parameter, is the predicted value of the i-th sample in the test set under the p-th partition, is the weight-optimized predicted value of the i-th sample in the test set under the p-th partition, represents ρ and K that minimize the objective function.

[0036] Optionally, the method of performing principal component analysis on the sample data set based on the first prediction matrix, the shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, and calculating the optimal weights of the plurality of base learners using the least squares method according to the result of the principal component analysis, specifically includes:

[0037] Based on the optimal shrinkage parameter, according to the first prediction matrix, the optimal shrinkage parameter ρ is obtained using the shrinkage robust Tyler-M estimator. * The corresponding first covariance matrix

[0038]

[0039] For the optimal shrinkage parameter ρ * The corresponding first covariance matrix Perform eigenvalue decomposition and sort according to the size of the eigenvalue to obtain the corresponding principal component matrix G = {g1, ..., g N}:

[0040] G i,j =γ j,1 f1(x i )+...+γ j,N f N (x i );

[0041] g k =Y k,1 f1+...+γ k,N f N ;

[0042] γ j =[γ j,1 , ..., γ j,N ] T ;

[0043] Among them, g N represents the principal component corresponding to the Nth base learner, G i,j represents the value of the i-th row and j-th column in the principal component matrix G, γ j Denotes the eigenvalue λ j The corresponding eigenvector, γ j,N Denotes the eigenvector γ j The Nth component in g k represents the principal component corresponding to the kth base learner, γ k,N Denotes the eigenvector γ k The Nth component in N Represents the predicted value of the Nth base learner for the sample data set;

[0044] Based on the optimal number of principal components K * , for the principal component matrix G = {g1, ..., g N} to reduce the dimension and get K * principal components

[0045] Assume that the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the principal component linear combination prediction matrix f RPCR :

[0046]

[0047] in, Indicates K * The optimal principal component weights, Indicates K * principal components;

[0048] According to the least squares method, the square error function L(β) is defined to measure the linear combination prediction matrix fR of the principal components. PCR The difference from the real matrix y:

[0049] L(β)=||y-Gβ|| 2 =(y-Gβ) T (y-Gβ);

[0050] Where y represents the true matrix of the sample data set;

[0051] The square error function L(β) is differentiated with respect to β, and the derivative is set to zero to obtain the optimal principal component weight vector β:

[0052]

[0053] y=[y1,...y n ] T ;

[0054] Among them, y n represents the nth sample x n The true value of

[0055] according to The optimal weights of the multiple base learners of the integrated learning model are calculated:

[0056]

[0057] Among them, w j represents the optimal weight of the j-th base learner, f j represents the prediction matrix of the j-th base learner for the sample data set, β k represents the kth optimal principal component weight, γ k,j The eigenvector γ representing the kth principal component k The value corresponding to the j-th base learner in .

[0058] To achieve the above-mentioned object of the invention, the present invention further provides a method for allocating weights of base learners in ensemble learning, wherein the method for allocating weights of base learners in ensemble learning is applied to classification tasks, and the method for allocating weights of base learners in ensemble learning comprises:

[0059] Using multiple base learners of the ensemble learning model to perform category prediction on the sample data set to obtain a second prediction matrix;

[0060] The Tyler-M estimator and the linear shrinkage estimator are combined to construct a shrinkage robust Tyler-M estimator, and the diagonal matrix is ​​introduced as the weighting matrix to obtain a new shrinkage robust Tyler-M estimator.

[0061] The sample data set is divided into a training set and a test set by a cross-validation method, the overall prediction error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall prediction error among multiple overall prediction errors are determined as the optimal shrinkage parameter and the optimal number of principal components;

[0062] Based on the second prediction matrix, the new shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, a two-dimensional principal component analysis method is used to perform a two-dimensional principal component analysis on the sample data set. According to the results of the two-dimensional principal component analysis, the optimal weights of multiple base learners are calculated by minimizing the cross entropy loss using the stochastic gradient descent method.

[0063] Optionally, the method of using multiple base learners of the ensemble learning model to perform category prediction on the sample data set to obtain a second prediction matrix specifically includes:

[0064] Using multiple base learners of the ensemble learning model to train the sample data set (x i ,y i ) to perform category prediction and obtain the second prediction matrix A i ∈R m×N ;

[0065] Among them, m is the number of categories, N is the number of base learners, R represents a real number matrix, A i Represents multiple base learners for the i-th sample x i The second prediction matrix.

[0066] Optionally, the Tyler-M estimator and the linear shrinkage estimator are combined to construct a shrinkage robust Tyler-M estimator, and a diagonal matrix is ​​introduced as a weighting matrix to obtain a new shrinkage robust Tyler-M estimator, which specifically includes:

[0067] Combining the Tyler-M estimator and the linear shrinkage estimator, a shrinkage robust Tyler-M estimator is constructed;

[0068] In the shrinkage robust Tyler-M estimator, the diagonal matrix D is introduced N As the weighting matrix, we get the new contraction robust Tyler-M estimator, where the diagonal matrix D N Each element on the diagonal of The corresponding diagonal elements of are calculated;

[0069] According to the second prediction matrix, the second covariance matrix is ​​obtained using the new shrinkage robust Tyler-M estimator:

[0070]

[0071] in, represents the second covariance matrix corresponding to the shrinkage parameter ρ, express The reciprocal n represents the number of samples. A after decentralization i , I N represents the identity matrix and T represents the matrix device.

[0072] Optionally, the cross-validation method is used to divide the sample data set into a training set and a test set, calculate the overall prediction error corresponding to the test set, and determine the shrinkage parameter and the number of principal components corresponding to the minimum overall prediction error among multiple overall prediction errors as the optimal shrinkage parameter and the optimal number of principal components, specifically including:

[0073] Using a cross-validation method, the sample data set is divided into P partitions, each of the partitions is used as a test set in turn, and the remaining partitions are used as training sets, to obtain different combinations of training sets and test sets;

[0074] For each candidate shrinkage parameter ρ and principal component number K, a two-dimensional principal component analysis method is used to perform a two-dimensional principal component analysis on the training set to obtain a training set model weight vector;

[0075] According to the training set model weight vector, weight optimization is performed on the test set prediction matrix corresponding to the test set to obtain a weight-optimized test set prediction matrix;

[0076] According to the test set prediction matrix and the weight-optimized test set prediction matrix, calculate the prediction error of the test set under the current partition corresponding to each parameter combination (ρ, K), wherein the prediction error is 1 minus the ACC accuracy rate;

[0077] Cumulatively sum the prediction errors of all partitioned test sets to obtain the overall prediction error;

[0078] By traversing the shrinkage parameters and gradually increasing the number of principal components, the parameter combination (ρ, K) that minimizes the overall prediction error is taken as the optimal parameter combination (ρ * ,K * ), where ρ * represents the optimal shrinkage parameter, K * represents the optimal number of principal components.

[0079] Optionally, the performing a two-dimensional principal component analysis on the sample data set by a two-dimensional principal component analysis method based on the second prediction matrix, the new shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, and calculating the optimal weights of the plurality of base learners by a stochastic gradient descent method according to the result of the two-dimensional principal component analysis by minimizing the cross entropy loss, specifically includes:

[0080] Based on the optimal shrinkage parameter, according to the second prediction matrix, the optimal shrinkage parameter ρ is obtained using the new shrinkage robust Tyler-M estimator. * The corresponding second covariance matrix

[0081]

[0082] For the optimal shrinkage parameter ρ * The corresponding second covariance matrix Perform eigenvalue decomposition to obtain the eigenvalue corresponding to the eigenvalue vector, denoted as p1, p2, ..., p N ∈R N×N , for the i-th sample x i The second prediction matrix A i , get the sample x i The corresponding principal component matrix

[0083]

[0084] Among them, p N represents the eigenvector of the Nth principal component, p k represents the eigenvector of the kth principal component, Represents sample x i The principal component corresponding to the kth base learner;

[0085] Based on the optimal number of principal components K * , for the principal component matrix Perform dimensionality reduction and obtain K * principal components

[0086] Assume that the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the final estimate

[0087] in, Indicates K * The optimal principal component weights, Represents sample xi The corresponding K * principal components;

[0088] By minimizing the cross entropy loss L, the optimal principal component weight vector β is obtained using the stochastic gradient descent method:

[0089]

[0090] in, Represents sample x i The true label on category j; Represents sample x i The predicted probability of belonging to category j, express The logarithm of

[0091] according to The optimal weights of the multiple base learners of the integrated learning model are calculated:

[0092]

[0093] Among them, w j represents the optimal weight of the j-th base learner, β k represents the kth optimal principal component weight, p k,j The eigenvector p representing the kth principal component k The value corresponding to the j-th base learner in .

[0094] In the present invention, a plurality of base learners of an integrated learning model are used to predict a sample data set to obtain a prediction matrix; a shrinkage robust Tyler-M estimator is constructed by combining a Tyler-M estimator and a linear shrinkage estimator, and a diagonal matrix is ​​introduced into the shrinkage robust Tyler-M estimator when applied to a classification task; a cross-validation method is used to determine an optimal shrinkage parameter and an optimal number of principal components; based on the shrinkage robust Tyler-M estimator or the shrinkage robust Tyler-M estimator with a diagonal matrix introduced, the optimal shrinkage parameter and the optimal number of principal components, a principal component analysis method PCA or a two-dimensional principal component analysis method 2DPCA (Two-Dimensional Principal Component Analysis) is used to perform principal component analysis or two-dimensional principal component analysis on the sample data set, and according to the principal component analysis result or the two-dimensional principal component analysis result, a least square method or a stochastic gradient descent method is used to calculate the optimal weights of the plurality of base learners. The present invention constructs a new covariance matrix estimator that is robust to high-dimensional data, limited sample size data, and non-Gaussian distribution data, namely, the shrinkage robust Tyler-M estimator, which improves the regression accuracy of the ensemble regressor and further extends the algorithm to classification tasks, so that the algorithm can also perform weight allocation for ensemble learning of classifiers, thereby improving the classification accuracy of the ensemble classifier. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 It is a flow chart of a preferred embodiment of the method for allocating weights of base learners in ensemble learning of the present invention applied to regression tasks;

[0096] Figure 2 It is a flow chart of a preferred embodiment of the method for allocating weights of base learners in ensemble learning of the present invention applied to classification tasks;

[0097] Figure 3 It is a flow chart of a preferred embodiment of the method for allocating weights of base learners in ensemble learning of the present invention. DETAILED DESCRIPTION

[0098] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0099] Ensemble learning has received extensive attention in the field of machine learning because it can improve the accuracy and robustness of model predictions. By integrating the prediction results of multiple basic learners, ensemble learning can usually significantly surpass the performance of a single model. However, how to effectively manage the correlation between learners and reasonably allocate weights remains one of the main challenges currently faced.

[0100] Basic ensemble methods usually use unweighted averaging of the prediction results of multiple learners to reduce estimation errors. However, due to the high correlation between learners, the effectiveness of this method in practical applications is often limited. To solve this problem, generalized ensemble methods use the sample covariance matrix of the learner prediction results to optimize the weights, thereby further reducing the estimation error. However, generalized ensemble methods may show instability when dealing with multicollinearity problems, affecting their actual effectiveness.

[0101] To address the problem of multicollinearity, the principal component regression method PCR uses principal component analysis PCA to reduce the impact of collinearity. In PCR, the covariance matrix of the regressor needs to be estimated by the predicted value of each regressor. Here, the sample covariance matrix is ​​used. However, when the data dimension is high, the number of samples is limited, or the data distribution is non-Gaussian, the accuracy of the sample covariance matrix in estimating the true covariance matrix decreases, which affects the performance of PCR and limits its effectiveness in complex application scenarios. In addition, PCR is an ensemble learning algorithm designed for regression tasks and cannot be applied to classification tasks.

[0102] In order to solve the above technical problems, the present invention provides a method for allocating weights of base learners in ensemble learning, which uses multiple base learners of an ensemble learning model to predict a sample data set to obtain a prediction matrix; combines a Tyler-M estimator and a linear shrinkage estimator to construct a shrinkage robust Tyler-M estimator, and introduces a diagonal matrix into the shrinkage robust Tyler-M estimator when applied to a classification task; uses a cross-validation method to determine an optimal shrinkage parameter and an optimal number of principal components; based on the shrinkage robust Tyler-M estimator or the shrinkage robust Tyler-M estimator with a diagonal matrix, the optimal shrinkage parameter and the optimal number of principal components, uses a principal component analysis method PCA or a two-dimensional principal component analysis method 2DPCA (Two-Dimensional Principal Component Analysis) to perform principal component analysis or two-dimensional principal component analysis on the sample data set, and according to the principal component analysis result or the two-dimensional principal component analysis result, uses a least squares method or a stochastic gradient descent method to calculate the optimal weights of multiple base learners. The present invention constructs a new covariance matrix estimator that is robust to high-dimensional data, limited sample size data, and non-Gaussian distribution data, namely, the shrinkage robust Tyler-M estimator, which improves the regression accuracy of the ensemble regressor and further extends the algorithm to classification tasks, so that the algorithm can also perform weight allocation for ensemble learning of classifiers, thereby improving the classification accuracy of the ensemble classifier.

[0103] The application content is further explained below through the description of embodiments in conjunction with the accompanying drawings.

[0104] The preferred embodiment of the method for allocating weights of base learners in ensemble learning of the present invention is applied to regression tasks, such as Figure 1 and Figure 3 As shown, specifically including:

[0105] S1. Use multiple base learners of the ensemble learning model to predict the sample data set to obtain a first prediction matrix.

[0106] In an implementation of this embodiment, the method of using multiple base learners of the ensemble learning model to predict the sample data set to obtain a first prediction matrix specifically includes:

[0107] Use multiple base learners (here refers to the base regressor, i.e., the integrated regressor) of the ensemble learning model to train the sample data set (x i ,y i ) to make predictions and obtain the first prediction matrix A∈R n×N , A ij =f j (x i );

[0108] Among them, n is the number of samples, N is the number of base learners, R represents a real number matrix, x i represents the i-th sample, y i represents the i-th sample x i The true value of A ij represents the value of the i-th row and j-th column in the first prediction matrix A, f j (x i ) represents the j-th base learner for the i-th sample x i The predicted value of .

[0109] S2. Combine the Tyler-M estimator and the linear shrinkage estimator to construct a shrinkage robust Tyler-M estimator.

[0110] In an implementation of this embodiment, the Tyler-M estimator and the linear shrinkage estimator are combined to construct a shrinkage robust Tyler-M estimator, which specifically includes:

[0111] Combining the Tyler-M estimator and the linear shrinkage estimator, a shrinkage robust Tyler-M estimator is constructed, and according to the first prediction matrix, the first covariance matrix is ​​obtained by using the shrinkage robust Tyler-M estimator:

[0112]

[0113]

[0114] in, represents the first covariance matrix corresponding to the shrinkage parameter ρ, express The reciprocal of f(x i ) represents the sample x i The predicted value of f(x i ) are all base learners for sample x i The predicted value f j (x i ), x i′ represents the i′th sample, f(x i′ ) represents the sample x i′ The predicted value of represents the decentralized f(x i ), T represents matrix transpose, I N Represents the identity matrix.

[0115] Specifically, the algorithms for regression tasks and classification tasks of the present invention are named Robust Principle Components Regression (RPCR) and Robust Principle Components Classification (RPCC) respectively. The RPCR algorithm combines the Tyler-M estimator and the linear shrinkage estimator. The robust covariance matrix estimator estimated by Tyler-M can estimate the covariance matrix more robustly in the presence of noise or outliers. The Tyler-M estimator is the solution of the following fixed-point equation:

[0116]

[0117] in,

[0118] Linear shrinkage estimator combined with sample covariance matrix And the identity matrix IN, shrink the estimator through linear combination, as follows:

[0119]

[0120] Here, ρ is the shrinkage parameter.

[0121] Introducing the idea of ​​contraction, by The calculation process of is embedded in the contraction coefficient (1-ρ), and ρI is added N The weighted part of the identity matrix of is combined with the two estimators to obtain a covariance calculation method with better robustness and accuracy when processing high-dimensional data, small sample data, and non-Gaussian distribution data, called the shrinkage robust Tyler-M estimator, as shown in the following formula:

[0122]

[0123] S3. Use a cross-validation method to divide the sample data set into a training set and a test set, calculate the overall mean square error corresponding to the test set, and determine the shrinkage parameter and the number of principal components corresponding to the minimum overall mean square error among multiple overall mean square errors as the optimal shrinkage parameter and the optimal number of principal components.

[0124] In an implementation of this embodiment, the cross-validation method is used to divide the sample data set into a training set and a test set, calculate the overall mean square error corresponding to the test set, and determine the shrinkage parameter and the number of principal components corresponding to the minimum overall mean square error among multiple overall mean square errors as the optimal shrinkage parameter and the optimal number of principal components, which specifically includes:

[0125] Using a cross-validation method, the sample data set is divided into P partitions, each of the partitions is used as a test set in turn, and the remaining partitions are used as training sets, to obtain different combinations of training sets and test sets;

[0126] For each candidate shrinkage parameter ρ (range 0 to 1) and number of principal components K (range 1 to N), the principal component analysis method is used to perform principal component analysis on the training set to obtain the training set projection matrix G of K principal components train ∈R (n-s)×N , where s represents the number of samples in the test set; that is, the prediction matrix A for the training set train ∈R (n-s)×N , based on the number of principal components selected by the current ρ value, perform principal component analysis PCA dimensionality reduction to obtain the projection matrix G of K principal components train ∈R (n-s)×N ;

[0127] Perform linear regression weight calculation according to the training set projection matrix to obtain the training set principal component weight vector and the training set model weight vector w′ corresponding to the training set (the specific principle process can be referred to step S4, and the calculation of the principal component weight vector and the model weight vector here is consistent with the principle of step S4);

[0128] According to the training set model weight vector w′, the test set prediction matrix A corresponding to the test set is test ∈R s ×N Perform weight optimization to obtain the test set prediction matrix after weight optimization

[0129]

[0130] According to the test set prediction matrix A test And the test set prediction matrix after the weight optimization Calculate the mean square error (MSE) of the test set under the current partition partition :

[0131]

[0132] y test,i ∈A test ;

[0133]

[0134] Among them, partition represents partition, y test,i is the predicted value of the i-th sample in the test set, is the weight-optimized predicted value of the i-th sample in the test set;

[0135] For each principal component number K, the mean square error of the test set of all partitions is accumulated and summed to obtain the overall mean square error MSE(ρ, K):

[0136]

[0137] Among them, MSE partition,p represents the mean square error of the test set under the p-th partition, P represents the total number of partitions, and p represents the p-th partition;

[0138] For each shrinkage parameter ρ, record the number of principal components K that minimizes the overall mean square error MSE(ρ, K) * , obtain the intermediate parameter combination (ρ, K * );

[0139] From all intermediate parameter combinations (ρ, K * ) is selected so that the overall mean square error MSE(ρ, K * )The minimum corresponding intermediate parameter combination is taken as the optimal parameter combination (ρ * , K * ):

[0140]

[0141] Among them, ρ * represents the optimal shrinkage parameter, is the predicted value of the i-th sample in the test set under the p-th partition, is the weight-optimized predicted value of the i-th sample in the test set under the p-th partition, represents ρ and K that minimize the objective function.

[0142] Specifically, in the regression task, in order to optimize the combination weights of the principal components, the mean square error (MSE) is used as the loss function, and the MSE value of each parameter combination is evaluated one by one to measure the performance of the model under different settings. For each ρ, the K value that minimizes MSE (ρ, K) is recorded, and finally the combination with the smallest mean square error among all ρ is selected as the best parameter combination (ρ * ,K * ), ensuring an optimal balance between complexity and accuracy, avoiding underfitting or overfitting.

[0143] S4. Based on the first prediction matrix, the shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, the principal component analysis method is used to perform principal component analysis on the sample data set, and according to the principal component analysis results, the least squares method is used to calculate the optimal weights of multiple base learners.

[0144] In an implementation of this embodiment, based on the first prediction matrix, the shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, a principal component analysis method is used to perform a principal component analysis on the sample data set, and according to the principal component analysis result, a least squares method is used to calculate the optimal weights of the plurality of base learners, specifically including:

[0145] Based on the optimal shrinkage parameter, according to the first prediction matrix, the optimal shrinkage parameter ρ is obtained using the shrinkage robust Tyler-M estimator. * The corresponding first covariance matrix

[0146]

[0147] For the optimal shrinkage parameter ρ * The corresponding first covariance matrix Perform eigenvalue decomposition and sort according to the size of the eigenvalue to obtain the corresponding principal component matrix G = {g1, ..., g N}, each principal component is a column vector of the principal component matrix G:

[0148] G i,j =γ j,1 f1(x i )+...+γ j,N f N (x i );

[0149] g k =γ k,1 f1+...+γ k,N f N ;

[0150] γ j =[γ j,1 , ..., γ j,N ] T ;

[0151] Among them, g N represents the principal component corresponding to the Nth base learner, G i,j represents the value of the i-th row and j-th column in the principal component matrix G, γ j Denotes the eigenvalue λ j The corresponding eigenvector, γ j,N Denotes the eigenvector γ j The Nth component in g k represents the principal component corresponding to the kth base learner, γ k,N Denotes the eigenvector γ k The Nth component in N Represents the predicted value of the Nth base learner for the sample data set;

[0152] Based on the optimal number of principal components K * , for the principal component matrix G = {g1, ..., g N} to reduce the dimension and get K * principal components

[0153] The ultimate goal of principal component regression is to establish a linear weight combination model based on the original learning model. In order to make the linear combination of principal components closest to the true value, it is necessary to find the optimal principal component weight vector. Assume that the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the principal component linear combination prediction matrix f RPCR :

[0154]

[0155] in, Indicates K * The optimal principal component weights, Indicates K * principal components;

[0156] According to the least squares method, the square error function L(β) is defined to measure the linear combination prediction matrix f RPCR The difference from the real matrix y:

[0157] L(β)=||y-Gβ|| 2 =(y-Gβ) T (y-Gβ);

[0158] Where y represents the true matrix of the sample data set;

[0159] The square error function L(β) is differentiated with respect to β, and the derivative is set to zero to obtain the optimal principal component weight vector β:

[0160]

[0161] y=[y1,...y n ] T ;

[0162] Among them, y n represents the nth sample x n The true value of

[0163] according to and g k =γ k,1 f1+...+γ k,N f N Three formulas are used to calculate the optimal weights of multiple base learners in the integrated learning model:

[0164]

[0165] Among them, w j represents the optimal weight of the j-th base learner, f j represents the prediction matrix of the j-th base learner for the sample data set, β k represents the kth optimal principal component weight, γ k,j The eigenvector γ representing the kth principal component k The value corresponding to the j-th base learner in .

[0166] In another implementation of this embodiment, step S4 is to divide the sample data set into a training set and a test set, process the training set, obtain the optimal weight of the base learner, and then use the test set to verify the obtained optimal weight of the base learner. Specifically, after calculating the weight w, the final integrated learning model can be constructed. The prediction matrix of the test set is multiplied by the weight w, to generate the test set prediction value of the integrated learning model, and the mean square error between these prediction values ​​and the true value is calculated to evaluate the performance and prediction accuracy of the model. This method effectively verifies the generalization ability of the model on the test data and ensures that it maintains good performance on unknown data.

[0167] In addition, based on the above-mentioned method for allocating weights of base learners in ensemble learning, the present invention also provides a method for allocating weights of base learners in ensemble learning applied to classification tasks, and a preferred embodiment of the method for allocating weights of base learners in ensemble learning applied to classification tasks, such as Figure 2 and Figure 3 As shown, specifically including:

[0168] B1. Use multiple base learners of the ensemble learning model to predict the category of the sample data set to obtain a second prediction matrix.

[0169] In an implementation of this embodiment, the method of using multiple base learners of the ensemble learning model to perform category prediction on the sample data set to obtain a second prediction matrix specifically includes:

[0170] Use multiple base learners (here refers to the base classifier, i.e., the integrated classifier) ​​of the integrated learning model to classify the sample data set (x i ,y i ) to perform category prediction and obtain the second prediction matrix A i ∈R m×N ;

[0171] Among them, m is the number of categories, N is the number of base learners, R represents a real number matrix, A i Represents multiple base learners for the i-th sample x i The second prediction matrix.

[0172] B2. Combining the Tyler-M estimator and the linear shrinkage estimator, a shrinkage robust Tyler-M estimator is constructed, and a diagonal matrix is ​​introduced as a weighting matrix to obtain a new shrinkage robust Tyler-M estimator.

[0173] In an implementation of this embodiment, the Tyler-M estimator and the linear shrinkage estimator are combined to construct a shrinkage robust Tyler-M estimator, and a diagonal matrix is ​​introduced as a weighting matrix to obtain a new shrinkage robust Tyler-M estimator, which specifically includes:

[0174] Combining the Tyler-M estimator and the linear shrinkage estimator, a shrinkage robust Tyler-M estimator is constructed;

[0175] In the shrinkage robust Tyler-M estimator, the diagonal matrix D is introduced N As the weighting matrix, we get the new contraction robust Tyler-M estimator, where the diagonal matrix D N Each element on the diagonal of The corresponding diagonal elements of are calculated;

[0176] According to the second prediction matrix, the second covariance matrix is ​​obtained using the new shrinkage robust Tyler-M estimator:

[0177]

[0178] in, represents the second covariance matrix corresponding to the shrinkage parameter ρ, express The reciprocal of , n represents the number of samples, A after decentralization i , I N represents the identity matrix and T represents the matrix device.

[0179] Specifically, unlike the regression task, the category prediction result of multiple base learners, i.e., base classifiers, for each sample is a two-dimensional matrix A i , it is not possible to directly use the shrinkage robust Tyler-M covariance matrix To estimate the covariance matrix of the base learner. Therefore, RPCC chooses to modify the estimation method of the covariance matrix. Specifically, the diagonal matrix D is introduced. N As a weighting matrix, it can better adapt to the covariance structure between samples. N is a dynamically adjusted diagonal matrix, each element on its diagonal is represented by The corresponding diagonal elements of are calculated, where Represents the decentralized matrix. The purpose of this step is to introduce an adaptive weight adjustment mechanism to estimate the covariance of each sample more accurately. This diagonal weighting method is similar to the idea of ​​optimizing the projection direction by calculating the divergence of the sample after projection in the two-dimensional principal component analysis method 2DPCA, that is, when calculating the covariance matrix, the contribution weight of each sample matrix is ​​dynamically adjusted, so that the covariance matrix estimator can better adapt to the structural characteristics of the data, thereby improving the robustness to outliers and noise. At the same time, the covariance matrix estimator The construction of introduces a shrinkage parameter ρ to balance the data divergence and estimation robustness. Finally, the 2DST estimator based on two-dimensional matrix samples is obtained, which is the new shrinkage robust Tyler-M estimator, and its calculation method is as follows:

[0180]

[0181] Among them, I N is the N×N identity matrix.

[0182] B3. Using the cross-validation method, the sample data set is divided into a training set and a test set, the overall prediction error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall prediction error are determined as the optimal shrinkage parameter and the optimal number of principal components.

[0183] In an implementation of this embodiment, the cross-validation method is used to divide the sample data set into a training set and a test set, calculate the overall prediction error corresponding to the test set, and determine the shrinkage parameter and the number of principal components corresponding to the minimum overall prediction error as the optimal shrinkage parameter and the optimal number of principal components, which specifically includes:

[0184] Using a cross-validation method, the sample data set is divided into P partitions, each of the partitions is used as a test set in turn, and the remaining partitions are used as training sets, to obtain different combinations of training sets and test sets;

[0185] For each candidate shrinkage parameter ρ (range 0 to 1) and principal component number K (range 1 to N), a two-dimensional principal component analysis is performed on the training set to obtain a training set model weight vector (the specific principle process can be referred to step B4, and the calculation of the model weight vector here is consistent with the principle of step B4);

[0186] According to the training set model weight vector, weight optimization is performed on the test set prediction matrix corresponding to the test set to obtain a weight-optimized test set prediction matrix;

[0187] According to the test set prediction matrix and the weight-optimized test set prediction matrix, calculate the prediction error of the test set under the current partition corresponding to each parameter combination (ρ, K), wherein the prediction error is 1 minus the ACC accuracy rate;

[0188] Cumulatively sum the prediction errors of all partitioned test sets to obtain the overall prediction error;

[0189] By traversing the shrinkage parameters and gradually increasing the number of principal components, cross-validating the cumulative error, and evaluating the performance of the model on the test set, the parameter combination (ρ, K) that minimizes the overall prediction error is finally selected as the optimal parameter combination (ρ * ,K * ), where ρ * represents the optimal shrinkage parameter, K * It indicates the optimal number of principal components and determines the best parameter settings for the model, ensuring the best balance between complexity and accuracy and avoiding underfitting or overfitting.

[0190] B4. Based on the new shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, a two-dimensional principal component analysis method is used to perform a two-dimensional principal component analysis on the sample data set. According to the results of the two-dimensional principal component analysis, the optimal weights of multiple base learners are calculated by minimizing the cross entropy loss using the stochastic gradient descent method.

[0191] In an implementation of this embodiment, based on the new shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, a two-dimensional principal component analysis method is used to perform a two-dimensional principal component analysis on the sample data set, and according to the two-dimensional principal component analysis result, the optimal weights of the plurality of base learners are calculated by minimizing the cross entropy loss using the stochastic gradient descent method, specifically including:

[0192] Based on the optimal shrinkage parameter, according to the second prediction matrix, the optimal shrinkage parameter ρ is obtained using the new shrinkage robust Tyler-M estimator. * The corresponding second covariance matrix

[0193]

[0194] For the optimal shrinkage parameter ρ * The corresponding second covariance matrix Perform eigenvalue decomposition to obtain the eigenvectors corresponding to the eigenvalues, denoted as p1, p2, ..., p N ∈R N×1 , for the i-th sample x i The second prediction matrix A i , get the sample x i The corresponding principal component matrix

[0195]

[0196] Among them, p N represents the eigenvector of the Nth principal component, p k represents the eigenvector of the kth principal component, Represents sample x i The principal component corresponding to the kth base learner;

[0197] Based on the optimal number of principal components K * , for the principal component matrix Perform dimensionality reduction and obtain K * principal components

[0198] Assume that the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the final estimate That is, the principal component linear combination prediction matrix f RPCC :

[0199]

[0200] in, Indicates K * The optimal principal component weights, Represents sample x i The corresponding K * principal components;

[0201] By minimizing the cross entropy loss L, the optimal principal component weight vector β is obtained using the stochastic gradient descent method:

[0202]

[0203] in, Represents sample x i The true label on category j; Represents sample x i The predicted probability of belonging to category j, express The logarithm of ; It should be noted that j here represents the category index, and the j in the previous text is the base learner index. In the cross entropy function, the meaning of j is switched from the base learner index to the category index, and the two do not conflict.

[0204] according to and Three formulas are used to calculate the optimal weights of multiple base learners in the integrated learning model:

[0205]

[0206] Among them, w j represents the optimal weight of the j-th base learner, β k represents the kth optimal principal component weight, p k,j The eigenvector p representing the kth principal component k The value corresponding to the j-th base learner in .

[0207] It should be noted that the final estimate That is, the principal component linear combination prediction matrix f RPCC .

[0208] In another implementation of this embodiment, step B4 is to divide the sample data set into a training set and a test set, process the training set to obtain the optimal weight of the base learner, and then use the test set to verify the obtained optimal weight of the base learner.

[0209] The present invention aims to design a new integrated learning algorithm to solve the problem of performance degradation of PCR when facing high-dimensional data, limited sample size data or non-Gaussian distribution data, and can be applied to classification and regression tasks at the same time, with the following advantages: 1) RPCR and RPCC can effectively process data in high-dimensional feature space and reduce the impact of multicollinearity on model performance by combining PCA / 2DPCA with an improved shrinkage robust Tyler-M covariance matrix estimator, namely, a shrinkage robust Tyler-M estimator. By optimizing the selection of principal components and the calculation of weights, the present invention has the advantages of high performance in regression and multi-classification. 1) Improve the accuracy of the model in the problem; 2) Enhance the robustness to high-dimensional, small sample and non-Gaussian distribution data: The improved shrinkage robust Tyler-M covariance matrix estimator combines Tyler's M estimate and Ledoit-Wolf's shrinkage estimate, so that the RPCR and RPCC methods can maintain stable performance when the data dimension is high, the sample size is limited and the data distribution deviates from the Gaussian distribution. This robustness is especially important for processing abnormal data and noisy data in practical applications; 3) Optimize the selection of principal components: Select the optimal number of principal components K through v-fold cross validation * According to different tasks, different loss functions are selected to optimize the combination weights of the principal components. The loss function of the classification task is the cross entropy, and the loss function of the regression task is the mean square error, which ensures that the selection of the principal components is more consistent with the goals of different tasks, thereby improving the prediction ability of the model.

[0210] To assist in the explanation, the following specific embodiment is provided, a method for allocating weights of base learners in ensemble learning, comprising the following steps:

[0211] Step 1: Divide the data set into a training set and a test set in a ratio of 4:1. Taking the random forest model as an example, train and select the decision trees as base learners (50), and construct a base learner prediction matrix based on the prediction value of each base learner for the sample data;

[0212] Step 2: Obtain ρ through cross-validation * , K * , ρ traverses from 0 to 1 with a step size of 0.1, and divides the data set into 5 partitions. Each partition is used as a new test set in turn. The remaining new training sets are first subjected to principal component analysis to obtain the principal components, and then the number of principal components is gradually increased. The overall mean square error under each combination of ρ and K is accumulated and calculated, and the ρ and K combination corresponding to the minimum overall mean square error is selected as ρ * , K * ;

[0213] Step 3: Get the best combination ρ * , K *After that, the training set can be used to calculate the corresponding ensemble learning model, optimize the required weights β and w, and construct the optimized ensemble learner;

[0214] Step 4: Calculate the error between the predicted value and the true value of the ensemble learner on the test set and evaluate the model.

[0215] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or terminal including the element.

[0216] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.

[0217] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for allocating weights of base learners in ensemble learning, characterized in that: The weight allocation method of base learners in ensemble learning is applied to regression tasks, and the weight allocation method of base learners in ensemble learning includes: Predicting the sample data set using multiple base learners of the ensemble learning model to obtain a first prediction matrix; Combining the Tyler-M estimator and the linear shrinkage estimator, a shrinkage robust Tyler-M estimator is constructed; The sample data set is divided into a training set and a test set by a cross-validation method, the overall mean square error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall mean square error among multiple overall mean square errors are determined as the optimal shrinkage parameter and the optimal number of principal components; Based on the first prediction matrix, the shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, the principal component analysis method is used to perform principal component analysis on the sample data set, and according to the principal component analysis results, the least squares method is used to calculate the optimal weights of multiple base learners.

2. The method for allocating weights of base learners in ensemble learning according to claim 1, characterized in that: The method of using multiple base learners of the ensemble learning model to predict the sample data set to obtain a first prediction matrix specifically includes: Using multiple base learners of the ensemble learning model to train the sample data set (x i ,y i ) to make predictions and obtain the first prediction matrix A∈R n×N , A ij =f j (x i ); Among them, n is the number of samples, N is the number of base learners, R represents a real number matrix, x i represents the i-th sample, y i represents the i-th sample x i The true value of A ij represents the value of the i-th row and j-th column in the first prediction matrix A, f j (x i ) represents the j-th base learner for the i-th sample x i The predicted value of .

3. The method for allocating weights of base learners in ensemble learning according to claim 2, characterized in that: The method combines the Tyler-M estimator and the linear shrinkage estimator to construct a shrinkage robust Tyler-M estimator, which specifically includes: Combining the Tyler-M estimator and the linear shrinkage estimator, a shrinkage robust Tyler-M estimator is constructed, and according to the first prediction matrix, the first covariance matrix is ​​obtained by using the shrinkage robust Tyler-M estimator: in, represents the first covariance matrix corresponding to the shrinkage parameter ρ, express The reciprocal of f(x i ) represents the sample x i The predicted value, x i′ represents the i′th sample, f(x i′ ) represents the sample x i′ The predicted value of represents the decentralized f(x i ), T represents matrix transpose, I N Represents the identity matrix.

4. The method for allocating weights of base learners in ensemble learning according to claim 3, characterized in that: The cross-validation method is used to divide the sample data set into a training set and a test set, calculate the overall mean square error corresponding to the test set, and determine the shrinkage parameter and the number of principal components corresponding to the minimum overall mean square error among multiple overall mean square errors as the optimal shrinkage parameter and the optimal number of principal components, specifically including: Using a cross-validation method, the sample data set is divided into P partitions, each of the partitions is used as a test set in turn, and the remaining partitions are used as training sets, to obtain different combinations of training sets and test sets; For each candidate shrinkage parameter ρ and principal component number K, the principal component analysis method is used to perform principal component analysis on the training set to obtain the training set projection matrix G of K principal components. train ∈R (n-s)×N , where s represents the number of test set samples; Perform linear regression weight calculation according to the training set projection matrix to obtain a training set principal component weight vector and a training set model weight vector w′ corresponding to the training set; According to the training set model weight vector w′, the test set prediction matrix A corresponding to the test set is test ∈R s×N Perform weight optimization to obtain the test set prediction matrix after weight optimization According to the test set prediction matrix A test And the test set prediction matrix after weight optimization Calculate the mean square error (MSE) of the test set under the current partition partition : and test,i ∈A test ; Among them, partition represents partition, y test,i is the predicted value of the i-th sample in the test set, is the weight-optimized predicted value of the i-th sample in the test set; For each principal component number K, the mean square error of all partitioned test sets is cumulatively summed to obtain the overall mean square error MSE(ρ,K): Among them, MSE partition,p represents the mean square error of the test set under the p-th partition, P represents the total number of partitions, and p represents the p-th partition; For each shrinkage parameter ρ, record the number of principal components K that minimizes the overall mean square error MSE(ρ,K) * , obtain the intermediate parameter combination (ρ,K * ); From all the intermediate parameter combinations (ρ, K * ) is selected so that the overall mean square error MSE(ρ,K * )The minimum corresponding intermediate parameter combination is taken as the optimal parameter combination (ρ * ,K * ): Among them, ρ * represents the optimal shrinkage parameter, is the predicted value of the i-th sample in the test set under the p-th partition, is the weight-optimized predicted value of the i-th sample in the test set under the p-th partition, represents ρ and K that minimize the objective function.

5. The method for allocating weights of base learners in ensemble learning according to claim 4, characterized in that: The method of performing principal component analysis on the sample data set based on the first prediction matrix, the shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, and calculating the optimal weights of the plurality of base learners using the least squares method according to the principal component analysis result, specifically includes: Based on the optimal shrinkage parameter, according to the first prediction matrix, the optimal shrinkage parameter ρ is obtained using the shrinkage robust Tyler-M estimator. * The corresponding first covariance matrix For the optimal shrinkage parameter ρ * The corresponding first covariance matrix Perform eigenvalue decomposition and sort according to the size of the eigenvalue to obtain the corresponding principal component matrix G = {g1,...,g N }: G i,j =γ j,1 f1(x i )+...+γ j,N f N (x i ); g k =c k,1 f1+...+c k,N f N ; c j =[γ j,1 ,…,c j,N ] T ; Among them, g N represents the principal component corresponding to the Nth base learner, G i,j represents the value of the i-th row and j-th column in the principal component matrix G, γ j Denotes the eigenvalue λ j The corresponding eigenvector, γ j,N Denotes the eigenvector γ j The Nth component in g k represents the principal component corresponding to the kth base learner, γ k,N Denotes the eigenvector γ k The Nth component in N Represents the predicted value of the Nth base learner for the sample data set; Based on the optimal number of principal components K * , for the principal component matrix G = {g1,...,g N } to reduce the dimension and get K * principal components Assume that the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the principal component linear combination prediction matrix f RPCR : in, Indicates K * The optimal principal component weights, Indicates K * principal components; According to the least squares method, the square error function L(β) is defined to measure the linear combination prediction matrix f of the principal components. RPCR The difference with the real matrix y: L(β)=||y-Gβ|| 2 =(y-Gβ) T (y-Gβ); Where y represents the true matrix of the sample data set; The square error function L(β) is differentiated with respect to β, and the derivative is set to zero to obtain the optimal principal component weight vector β: y=[y1,...y n ] T ; Among them, y n represents the nth sample x n The true value of according to The optimal weights of the multiple base learners of the integrated learning model are calculated: Among them, w j represents the optimal weight of the j-th base learner, f j represents the prediction matrix of the j-th base learner for the sample data set, β k represents the kth optimal principal component weight, γ k,j The eigenvector γ representing the kth principal component k The value corresponding to the j-th base learner in .

6. A method for allocating weights of base learners in ensemble learning, characterized in that: The method for allocating weights of base learners in ensemble learning is applied to classification tasks, and the method for allocating weights of base learners in ensemble learning includes: Using multiple base learners of the ensemble learning model to perform category prediction on the sample data set to obtain a second prediction matrix; The Tyler-M estimator and the linear shrinkage estimator are combined to construct a shrinkage robust Tyler-M estimator, and the diagonal matrix is ​​introduced as the weighting matrix to obtain a new shrinkage robust Tyler-M estimator. The sample data set is divided into a training set and a test set by a cross-validation method, the overall prediction error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall prediction error among multiple overall prediction errors are determined as the optimal shrinkage parameter and the optimal number of principal components; Based on the second prediction matrix, the new shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, a two-dimensional principal component analysis method is used to perform a two-dimensional principal component analysis on the sample data set. According to the results of the two-dimensional principal component analysis, the optimal weights of multiple base learners are calculated by minimizing the cross entropy loss using the stochastic gradient descent method.

7. The method for allocating weights of base learners in ensemble learning according to claim 6, characterized in that: The method of using multiple base learners of the ensemble learning model to perform category prediction on the sample data set to obtain a second prediction matrix specifically includes: Using multiple base learners of the ensemble learning model to train the sample data set (x i ,y i ) to perform category prediction and obtain the second prediction matrix A i ∈R m×N ; Among them, m is the number of categories, N is the number of base learners, R represents a real number matrix, A i Represents multiple base learners for the i-th sample x i The second prediction matrix.

8. The method for allocating weights of base learners in ensemble learning according to claim 7, characterized in that: The Tyler-M estimator and the linear shrinkage estimator are combined to construct a shrinkage robust Tyler-M estimator, and a diagonal matrix is ​​introduced as a weighting matrix to obtain a new shrinkage robust Tyler-M estimator, which specifically includes: Combining the Tyler-M estimator and the linear shrinkage estimator, a shrinkage robust Tyler-M estimator is constructed; In the shrinkage robust Tyler-M estimator, the diagonal matrix D is introduced N As the weighting matrix, we get the new contraction robust Tyler-M estimator, where the diagonal matrix D N Each element on the diagonal of The corresponding diagonal elements of are calculated; According to the second prediction matrix, the second covariance matrix is ​​obtained using the new shrinkage robust Tyler-M estimator: in, represents the second covariance matrix corresponding to the shrinkage parameter ρ, express The reciprocal of , n represents the number of samples, A after decentralization i , I N represents the identity matrix and T represents the matrix device.

9. The method for allocating weights of base learners in ensemble learning according to claim 8, characterized in that: The cross-validation method is adopted to divide the sample data set into a training set and a test set, calculate the overall prediction error corresponding to the test set, and determine the shrinkage parameter and the number of principal components corresponding to the minimum overall prediction error among multiple overall prediction errors as the optimal shrinkage parameter and the optimal number of principal components, specifically including: Using a cross-validation method, the sample data set is divided into P partitions, each of the partitions is used as a test set in turn, and the remaining partitions are used as training sets, to obtain different combinations of training sets and test sets; For each candidate shrinkage parameter ρ and principal component number K, a two-dimensional principal component analysis method is used to perform a two-dimensional principal component analysis on the training set to obtain a training set model weight vector; According to the training set model weight vector, weight optimization is performed on the test set prediction matrix corresponding to the test set to obtain a weight-optimized test set prediction matrix; According to the test set prediction matrix and the weight-optimized test set prediction matrix, calculate the prediction error of the test set under the current partition corresponding to each parameter combination (ρ, K), wherein the prediction error is 1 minus the ACC accuracy rate; Cumulatively sum the prediction errors of all partitioned test sets to obtain the overall prediction error; By traversing the shrinkage parameters and gradually increasing the number of principal components, the parameter combination (ρ, K) that minimizes the overall prediction error is taken as the optimal parameter combination (ρ * ,K * ), where ρ * represents the optimal shrinkage parameter, K * represents the optimal number of principal components.

10. The method for allocating weights of base learners in ensemble learning according to claim 9, characterized in that: The method of performing a two-dimensional principal component analysis on the sample data set based on the second prediction matrix, the new shrinkage robust Tyler-M estimator, the optimal shrinkage parameter and the optimal number of principal components, and calculating the optimal weights of the plurality of base learners using a stochastic gradient descent method by minimizing the cross entropy loss according to the result of the two-dimensional principal component analysis, specifically includes: Based on the optimal shrinkage parameter, according to the second prediction matrix, the optimal shrinkage parameter ρ is obtained using the new shrinkage robust Tyler-M estimator. * The corresponding second covariance matrix For the optimal shrinkage parameter ρ * The corresponding second covariance matrix Perform eigenvalue decomposition to obtain the eigenvectors corresponding to the eigenvalues, denoted as p1, p2, ..., p N ∈R N×1 , for the i-th sample x i The second prediction matrix A i , get the sample x i The corresponding principal component matrix Among them, p N represents the eigenvector of the Nth principal component, p k represents the eigenvector of the kth principal component, Represents sample x i The principal component corresponding to the kth base learner; Based on the optimal number of principal components K * , for the principal component matrix Perform dimensionality reduction and obtain K * principal components Assume that the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the final estimate in, Indicates K * The optimal principal component weights, Represents sample x i The corresponding K * principal components; By minimizing the cross entropy loss L, the optimal principal component weight vector β is obtained using the stochastic gradient descent method: in, Represents sample x i The true label on category j; Represents sample x i The predicted probability of belonging to category j, express The logarithm of according to The optimal weights of the multiple base learners of the integrated learning model are calculated: Among them, w j represents the optimal weight of the j-th base learner, β k represents the kth optimal principal component weight, p k,j The eigenvector p representing the kth principal component k The value corresponding to the j-th base learner in .