Integrated learning weight distribution method based on learner accuracy and correlation
By adopting an integrated learning weight allocation method based on learner accuracy and correlation in regression and classification tasks, combining decision tree partitioning and principal component regression algorithm, the problem of single weight allocation in the existing technology is solved, and higher prediction accuracy and model generalization ability are achieved.
Patent Information
- Application Number
- CN202510092440.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
In the regression and classification tasks, the weight allocation is single, and the model accuracy and data correlation are not fully combined.
The integrated learning weight allocation method based on learner accuracy and correlation is adopted, and data partitioning is partitioned by building a decision tree, combining the relative root mean square error algorithm and principal component regression algorithm, and the weight set is fused for weighted merge.
The prediction accuracy and generalization ability of the model in complex and heterogeneous high data is improved, accuracy and correlation are balanced, and the robustness and adaptability of the model is enhanced.
Smart Images

Figure CN120012959A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ensemble learning, and in particular to an ensemble learning weight allocation method based on learner accuracy and correlation. Background Art
[0002] In the current field of machine learning and statistics, regression and classification algorithms are widely used in various prediction and decision-making tasks. Traditional regression methods usually rely on a single model or a simple integration method to fit and predict data. However, faced with complex and highly heterogeneous data sets, a single model often fails to capture the full picture of the data. In classification algorithms, models based on accuracy or a single evaluation criterion are prone to fall into local optimality, ignoring the characteristics of global data, resulting in reduced classification accuracy. To this end, a variety of integrated learning methods have emerged in the prior art, such as bagging algorithms, boosting methods, etc., which integrate the prediction results of multiple weak learners to improve model performance. However, these methods often do not consider the local structural characteristics of the data, that is, the partitioning of the data and the differences in data distribution in different regions.
[0003] In addition, common weight distribution methods, such as simple weighted average or fixed ratio fusion, are difficult to take into account both model accuracy and data relevance when dealing with multi-model predictions. Especially in regression tasks, many methods only consider accuracy during the fusion process and ignore the correlation characteristics between data. In classification tasks, principal component analysis (PCA) is often used for dimensionality reduction, but lacks a full combination of classification accuracy and correlation characteristics. Summary of the invention
[0004] The main purpose of the present invention is to provide an integrated learning weight allocation method based on learner accuracy and relevance, aiming to solve the problem that the existing weight allocation methods for regression and classification tasks are single and fail to fully combine model accuracy and data relevance.
[0005] To achieve the above-mentioned object of the invention, the present invention provides an integrated learning weight allocation method based on learner accuracy and correlation, the integrated learning weight allocation method based on learner accuracy and correlation is applied to regression tasks, and the integrated learning weight allocation method based on learner accuracy and correlation includes:
[0006] Constructing a decision tree to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition of the input feature space of the sample data set, and using multiple base learners to predict the sample data set to obtain a first prediction matrix;
[0007] Based on the first prediction matrix, using a relative root mean square error algorithm to obtain a first weight set corresponding to the partition;
[0008] Based on the first prediction matrix, obtaining a second weight set using a principal component regression algorithm;
[0009] The first weight set and the second weight set are fused to obtain a final weight set, and the prediction values in the first prediction matrix are weighted and combined using the final weight set to obtain a final prediction set.
[0010] Optionally, the step of constructing a decision tree to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition of an input feature space of the sample data set, and using multiple base learners to predict the sample data set to obtain a first prediction matrix, specifically includes:
[0011] Construct a tree with a depth of d and a number of leaf nodes of L = 2 d A decision tree is used to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition l of the input feature space of the sample data set;
[0012] Use N base learners to predict the sample data set and obtain the first prediction matrix A∈R n×N , A ij =f j (x i );
[0013] Among them, n is the number of samples, N is the number of base learners, R is a real number matrix, A ij represents the value of the i-th row and j-th column in the first prediction matrix A, f j (x i ) represents the jth base learner for sample x i The predicted value of , i represents the sample index, and j represents the base learner index.
[0014] Optionally, the obtaining, based on the first prediction matrix, a first weight set corresponding to the partition by using a relative root mean square error algorithm specifically includes:
[0015] For each base learner, the relative root mean square error RRMSE is used to evaluate the performance of the base learner on partition l, and a constant term constant is introduced as the denominator:
[0016]
[0017] Among them, RRMSE l,j represents the relative root mean square error of the j-th base learner in partition l, n l is the number of samples in partition l, y iRepresents sample x i The corresponding true value, is the mean of all true values in partition l, where l represents the partition index;
[0018] According to the relative root mean square error RRMSE, a weight is assigned to each base learner in partition l:
[0019]
[0020] Among them, w l,j represents the weight of the j-th base learner in partition l, and ε represents the smoothing factor;
[0021] Normalize the weight of each base learner in partition l to obtain the normalized weight:
[0022]
[0023] in, represents the normalized weight of the j-th base learner in partition l, w l,k represents the weight of the k-th base learner in partition l, and k represents the k-th base learner;
[0024] The normalized weights of all base learners in partition l form the first weight set of partition l
[0025]
[0026] in, Represents the normalized weight of the Nth base learner in partition l.
[0027] Optionally, obtaining the second weight set based on the first prediction matrix by using a principal component regression algorithm specifically includes:
[0028] Construct the sample covariance matrix
[0029]
[0030] Among them, f(x i ) represents all base learners for sample x i The prediction vector of represents the normalized f(x i ), Represents the transposed T represents transpose, μ represents the mean vector, σ represents the standard deviation vector, μ j represents the jth component in the mean vector μ, μ N represents the Nth component in the mean vector μ, σ jrepresents the jth component of the standard deviation vector σ, σ N represents the Nth component in the standard deviation vector σ;
[0031] Using principal component analysis, the sample covariance matrix As input, perform eigenvalue decomposition and sort according to the eigenvalue size λ1>λ2>...>λ N , output principal component matrix G = {g1, ..., g N}:
[0032] G i,j =γ j,1 f1(x i )+...+γ j,N f N (x i );
[0033] g k =γ k,1 f1+...+γ k,N f N ;
[0034] γ j =[γ j,1 , ..., γ j,N ] T ;
[0035] Among them, λ N represents the Nth eigenvalue, g N represents the Nth principal component, G i,j represents the value of the i-th row and j-th column in the principal component matrix G, γ j Denotes the eigenvalue λ j The corresponding eigenvector, γ j,N Denotes the eigenvector γ j The Nth component in N (x i ) represents the Nth base learner for sample x i The predicted value, g k represents the kth principal component, γ k,N Denotes the eigenvector γ k The Nth component in N Represents the prediction vector of the Nth base learner for all samples;
[0036] The sample data set is divided into a training set and a test set by using the cross-validation method, the overall mean square error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall mean square error among multiple overall mean square errors are determined as the optimal shrinkage parameter and the optimal number of principal components K * ;
[0037] Based on the optimal number of principal components K * , for the principal component matrix G = {g1, ..., g N} to reduce the dimension and get K * principal components
[0038] If the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the principal component linear combination prediction matrix f PCR :
[0039]
[0040] in, Indicates K * The optimal principal component weights, β j represents the jth optimal principal component weight, g j represents the jth principal component;
[0041] According to the least squares method, the square error function L(β) is defined to measure the linear combination prediction matrix f PCR The difference with the real matrix y:
[0042] L(β)=||y-Gβ|| 2 =(y-Gβ) T (y-Gβ);
[0043] The square error function L(β) is differentiated with respect to β, and the derivative is set to zero to obtain the optimal principal component weight vector β:
[0044]
[0045] y=[y1,...y n ] T ;
[0046] Among them, y n Represents sample x n The corresponding true value;
[0047] according to Calculate and obtain the second weight set v of the multiple base learners of the original learning model * =[v1, ..., v N ]:
[0048]
[0049] Among them, v j represents the weight of the j-th base learner, f jrepresents the prediction vector of the j-th base learner for all samples, β k represents the kth optimal principal component weight, γk j represents the eigenvector γ k The jth component in .
[0050] Optionally, the fusing the first weight set and the second weight set to obtain a final weight set, and using the final weight set to weightedly combine the prediction values in the first prediction matrix to obtain a final prediction set, specifically includes:
[0051] The first weight set and the second weight set of partition l are combined to obtain the final weight set of partition l
[0052]
[0053] in, and Respectively represent the first weight set of partition l The maximum and minimum weight values in v max and v min Respectively represent the second weight set v * The maximum and minimum weight values in ;
[0054] For each base learner, use the final set of weights for the partition l Perform weighted summation on the predicted values in the first prediction matrix to obtain the final predicted value of partition l
[0055] in, represents the predicted value of the j-th base learner on partition l, represents the final weight of the j-th base learner on partition l;
[0056] Merge the final prediction values of each partition to get the final prediction set
[0057]
[0058] in, Represents the final predicted value of partition L.
[0059] To achieve the above-mentioned object of the invention, the present invention also provides an integrated learning weight allocation method based on learner accuracy and relevance, the integrated learning weight allocation method based on learner accuracy and relevance is applied to classification tasks, and the integrated learning weight allocation method based on learner accuracy and relevance includes:
[0060] Constructing a decision tree to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition of the input feature space of the sample data set, and using multiple base learners to perform category prediction on the sample data set to obtain a second prediction matrix;
[0061] Based on the second prediction matrix, obtaining a first weight set corresponding to the partition using an accuracy algorithm;
[0062] Based on the second prediction matrix, obtaining a second weight set using a principal component classification algorithm;
[0063] The first weight set and the second weight set are integrated to obtain a final weight set, and the prediction probabilities in the second prediction matrix are weighted and summed using the final weight set to obtain a final prediction probability, and the category corresponding to the largest final prediction probability is taken as the final prediction result.
[0064] Optionally, the step of constructing a decision tree to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition of an input feature space of the sample data set, and using multiple base learners to perform category prediction on the sample data set to obtain a second prediction matrix, specifically includes:
[0065] Construct a tree with a depth of d and a number of leaf nodes of L = 2 d A decision tree is used to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition l of the input feature space of the sample data set;
[0066] Use N base learners to perform category prediction on the sample data set to obtain a second prediction matrix A i ∈R m×N ;
[0067] Among them, m is the number of categories, N is the number of base learners, R is a real number matrix, A i Represents N base learners for sample x i The second prediction matrix of , i represents the sample index.
[0068] Optionally, the obtaining, based on the second prediction matrix, a first weight set corresponding to the partition by using an accuracy algorithm specifically includes:
[0069] For each base learner pair sample x i For predictions, define the indicator function for correct predictions:
[0070]
[0071] Among them, TP ij is the jth base learner for sample x i The correct predictive indicator, is the jth base learner for sample x i The predicted category, y i is the sample x i The corresponding true category, j represents the base learner index;
[0072] For partition l, calculate the prediction accuracy of each base learner in partition l:
[0073]
[0074] Among them, Accuracy j (l) represents the prediction accuracy of the j-th base learner in partition l, n l is the number of samples in partition l, X l Represents the sample data subset corresponding to partition l;
[0075] The prediction accuracy of each base learner in partition l is normalized to obtain the normalized weight:
[0076]
[0077] Among them, w l,j represents the normalized weight of the j-th base learner in partition l;
[0078] The normalized weights of all base learners in partition l form the first weight set of partition l
[0079]
[0080] Among them, w l,N Represents the normalized weight of the Nth base learner in partition l.
[0081] Optionally, obtaining a second weight set based on the second prediction matrix by using a principal component regression algorithm specifically includes:
[0082] Construct a first covariance matrix C1 and a second covariance matrix C2, and combine the first covariance matrix C1 and the second covariance matrix C2 to obtain a comprehensive covariance matrix
[0083] Among them, n represents the number of samples, p(x i ) are all base learners for sample x i The predicted probability vector on the true category, p N (x i ) represents the Nth base learner for sample x iThe predicted probability on the real category, μ1 is the mean vector of the predicted probability of the real category; Q(x i ) represents all base learners for sample x i The predicted probability vector on the non-true class, q N (x i ) represents the Nth base learner for sample x i The predicted probability on the non-true class, It means that Q(x i ) Expand the high-dimensional tensor into a two-dimensional matrix Q by row priority flat The i-th row of ; μ2 is the predicted probability mean vector of the non-true category; tr(·) is the trace operation of the matrix;
[0084] Using the two-dimensional principal component analysis method, the comprehensive covariance matrix As input, perform eigenvalue decomposition to obtain the eigenvectors corresponding to the eigenvalues, denoted as u1, u2, ..., u N ∈R N×N , for the i-th sample x i The second prediction matrix A i , get the sample x i The corresponding principal component matrix
[0085]
[0086] Among them, uN represents the Nth eigenvector, u k represents the kth eigenvector, Represents sample x i The corresponding kth principal component;
[0087] The sample data set is divided into a training set and a test set by a cross-validation method, the overall prediction error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall prediction error among multiple overall prediction errors are determined as the optimal shrinkage parameter and the optimal number of principal components K. * , where the prediction error is 1 minus the accuracy;
[0088] Based on the optimal number of principal components K * , for the principal component matrix Perform dimensionality reduction and obtain K * principal components
[0089] If the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the final estimate
[0090]
[0091] in, Indicates K * The optimal principal component weights, Represents sample x i The corresponding K * principal components;
[0092] By minimizing the cross entropy loss L, the optimal principal component weight vector β is obtained using the stochastic gradient descent method:
[0093]
[0094] in, Represents sample x i The true label on category j; Represents sample x i The predicted probability of belonging to category j, express The logarithm of
[0095] according to Calculate and obtain the second weight set v of the multiple base learners of the original learning model * =[v1, ..., v N ]:
[0096]
[0097] Among them, v j represents the weight of the j-th base learner, β k represents the kth optimal principal component weight, u k,j Denotes the feature vector u k The jth component in .
[0098] Optionally, the fusing the first weight set and the second weight set to obtain a final weight set, using the final weight set to perform weighted summation on the prediction probabilities in the second prediction matrix to obtain a final prediction probability, and taking the category corresponding to the largest final prediction probability as the final prediction result, specifically includes:
[0099] The first weight set and the second weight set of partition l are combined to obtain the final weight set of partition l
[0100]
[0101] in, and Respectively represent the first weight set of partition l The maximum and minimum weight values in v max and vmin Respectively represent the second weight set v * The maximum and minimum weight values in ;
[0102] For each base learner, use the final set of weights for the partition l Performing weighted summation on the prediction probabilities in the second prediction matrix to obtain the final prediction probability of partition l, and taking the category corresponding to the largest final prediction probability as the final prediction result;
[0103]
[0104] in, Represents the sample x i The final prediction result, f j (x i )[c] represents the jth base learner for sample x i The predicted probability of category c is, represents the final weight of the j-th base learner on partition l.
[0105] The present invention partitions the data by similarity and places more similar data in the same area, thereby performing more refined modeling; for regression problems, the relevance of the principal component regression algorithm (Principal Components Regression, PCR) and the accuracy of the relative root mean square error algorithm (Relative Root Mean square error, RRMSE) are combined, and the weights of the two are reasonably allocated by a fusion formula, so that the model achieves a balance between accuracy and relevance; for classification problems, the principal component classification algorithm (Principal Components Classification, PCC) and the accuracy algorithm (Accuracy, Acc) are used to improve the classification effect with a similar fusion method. The present invention can more effectively process complex and highly heterogeneous data, ensuring that the model of the local area is more targeted; through the comprehensive weight allocation of accuracy and relevance, the present invention can achieve excellent performance on both local and global data; the present invention is suitable for regression and classification tasks, and provides stronger adaptability and model generalization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0106] Figure 1 It is a flow chart of a preferred embodiment of the present invention of an integrated learning weight allocation method based on learner accuracy and correlation applied to a regression task;
[0107] Figure 2 It is a flow chart of a preferred embodiment of the present invention of the integrated learning weight allocation method based on learner accuracy and correlation applied to classification tasks;
[0108] Figure 3 It is a flow chart of a preferred embodiment of the integrated learning weight allocation method based on learner accuracy and correlation of the present invention. DETAILED DESCRIPTION
[0109] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0110] The application content is further explained below through the description of embodiments in conjunction with the accompanying drawings.
[0111] The preferred embodiment of the present invention is to apply the integrated learning weight allocation method based on learner accuracy and correlation to regression tasks, such as Figure 1 and Figure 3 As shown, specifically including:
[0112] S1. Construct a decision tree to partition a sample data set, wherein each leaf node of the decision tree corresponds to a partition of an input feature space of the sample data set, and use multiple base learners to predict the sample data set to obtain a first prediction matrix.
[0113] In an implementation of this embodiment, the construction of a decision tree performs partition processing on the sample data set, each leaf node of the decision tree corresponds to a partition of the input feature space of the sample data set, and the sample data set is predicted using multiple base learners to obtain a first prediction matrix, which specifically includes:
[0114] Construct a tree with a depth of d and a number of leaf nodes of L = 2 d A decision tree is used to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition l of the input feature space of the sample data set;
[0115] Use N base learners to predict the sample data set and obtain the first prediction matrix A∈R n×N , A ij =f j (x i );
[0116] Among them, n is the number of samples, N is the number of base learners, R is a real number matrix, A ij represents the value of the i-th row and j-th column in the first prediction matrix A, f j (x i ) represents the jth base learner for sample x i The predicted value of , i represents the sample index, and j represents the base learner index.
[0117] Specifically, the decision tree is used as the training data set, i.e., the sample data set (X train, y train ) to perform partitioning. Create a tree with a depth of d and a number of leaf nodes of L = 2 d Each leaf node of a decision tree corresponds to a region in the input feature space. i ,y i ) According to the input feature x i Divided into 2 d On leaf nodes, each leaf node l (leaf nodes and partitions have a one-to-one correspondence, so l can represent both leaf nodes and partitions) corresponds to a data subset X l It should be noted that the sample data set can be understood as (X train ,y train ), which can also be understood as X train , the sample can be understood as (x i ,y i ), which can also be understood as x i , which is understood according to the context and does not conflict. The prior art usually does not perform clear partitioning of data, while the present invention partitions data through decision trees or similar algorithms, and classifies more similar data into the same area for modeling. This strategy helps to better capture local features, especially in complex or highly heterogeneous datasets.
[0118] S2. Based on the first prediction matrix, use a relative root mean square error algorithm to obtain a first weight set corresponding to the partition.
[0119] In an implementation of this embodiment, obtaining the first weight set corresponding to the partition by using a relative root mean square error algorithm based on the first prediction matrix specifically includes:
[0120] For each base learner, the relative root mean square error RRMSE is used to evaluate the performance of the base learner on partition l, and in order to prevent the true value of the denominator from being 0, a constant term constant is introduced into the denominator:
[0121]
[0122] Among them, RRMSE l,j represents the relative root mean square error of the j-th base learner in partition l, n l is the number of samples in partition l, y i Represents sample x i The corresponding true value, is the mean of all true values in partition l, where l represents the partition index;
[0123] This processing method ensures that there will be no calculation errors even when the sample value is small or zero; the core of this formula is to evaluate the performance of the model in a specific partition by calculating the deviation between the true value and the predicted value;
[0124] According to the relative root mean square error RRMSE, a weight is assigned to each base learner in partition l:
[0125]
[0126] Among them, w l,j represents the weight of the j-th base learner in partition l, ε represents the smoothing factor, which is used to prevent division by zero errors;
[0127] The weight of each base learner in partition l is normalized to ensure that the sum of the weights of all base learners is 1, and the normalized weight is obtained:
[0128]
[0129] in, represents the normalized weight of the j-th base learner in partition l, w l,k represents the weight of the k-th base learner in partition l, and k represents the k-th base learner;
[0130] The normalized weights of all base learners in partition l form the first weight set of partition l
[0131]
[0132] in, Represents the normalized weight of the Nth base learner in partition l.
[0133] Specifically, the first weight set It reflects the relative effectiveness of the base learner in a specific partition and provides an important basis for subsequent model integration and prediction. This method not only optimizes the performance of the model, but also enhances its adaptability to different partition characteristics.
[0134] S3. Based on the first prediction matrix, obtain a second weight set using a principal component regression algorithm.
[0135] In an implementation of this embodiment, obtaining the second weight set based on the first prediction matrix by using a principal component regression algorithm specifically includes:
[0136] Construct the sample covariance matrix
[0137]
[0138] Among them, f(x i ) represents all base learners for sample x i The prediction vector of represents the normalized f(x i ), Represents the transposed T represents transpose, μ represents the mean vector, σ represents the standard deviation vector, μ j represents the jth component in the mean vector μ, μ N represents the Nth component in the mean vector μ, σ j represents the jth component of the standard deviation vector σ, σ N represents the Nth component in the standard deviation vector σ;
[0139] Using principal component analysis, the sample covariance matrix As input, perform eigenvalue decomposition and sort according to the eigenvalue size λ1>λ2>...>λ N , output principal component matrix G = {g1, ..., g N}:
[0140] G i,j =γ j,1 f1(x i )+...+γ j,N f N (x i );
[0141] g k =γ k,1 f1+...+γ k,N f N ;
[0142] γ j =[γ j,1 , ..., γ j,N ] T ;
[0143] Among them, λ N represents the Nth eigenvalue, g N represents the Nth principal component, G i,j represents the value of the i-th row and j-th column in the principal component matrix G, γ j Denotes the eigenvalue λ j The corresponding eigenvector, γ j,N Denotes the eigenvector γ j The Nth component in N (x i ) represents the Nth base learner for sample x i The predicted value, g k represents the kth principal component, γ k,NDenotes the eigenvector γ k The Nth component in N Represents the prediction vector of the Nth base learner for all samples;
[0144] The sample data set is divided into a training set and a test set by using the cross-validation method, the overall mean square error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall mean square error among multiple overall mean square errors are determined as the optimal shrinkage parameter and the optimal number of principal components K * ;
[0145] The ultimate goal of principal component regression is to establish a linear weight combination model based on the original learning model, based on the optimal number of principal components K * , for the principal component matrix G = {g1, ..., g N} to reduce the dimension and get K * principal components
[0146] In order to make the linear combination of principal components closest to the true value, it is necessary to find the optimal principal component weight vector. Assume that the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the principal component linear combination prediction matrix f PCR :
[0147]
[0148] in, Indicates K * The optimal principal component weight, βj represents the jth optimal principal component weight, g j represents the jth principal component; this formula is a regression formula;
[0149] According to the least squares method, the square error function L(β) is defined to measure the linear combination prediction matrix f of the principal components. PCR The difference with the real matrix y:
[0150] L(β)=||y-Gβ|| 2 =(y-Gβ) T (y-Gβ);
[0151] The square error function L(β) is differentiated with respect to β, and the derivative is set to zero to obtain the optimal principal component weight vector β:
[0152]
[0153] y=[y1,...y n ] T ;
[0154] Among them, y n Represents sample x n The corresponding true value;
[0155] according to Calculate and obtain the second weight set v of the multiple base learners of the original learning model * =[v1, ..., v N ]:
[0156]
[0157] Among them, v j represents the weight of the j-th base learner, f j represents the prediction vector of the j-th base learner for all samples, β k represents the kth optimal principal component weight, γ k,j Denotes the eigenvector γ k The jth component in .
[0158] Specifically, the PCR algorithm uses the sample covariance matrix To process the predicted value of the base learner, we first construct the sample covariance matrix Then, principal component analysis (PCA) is performed, and then the optimal number of principal components K for building the regression model needs to be determined. * ,The search process will be described in detail below, and finally the regression model is established to obtain the weights.
[0159] The core of parameter search is to select the optimal number of principal components K * The basic idea is to gradually incorporate the principal components into the prediction estimate f according to the regression formula (regression formula above) PCR until all N principal components are included. The order of principal components is based on their ability to explain the variance in the prediction matrix A: the first few principal components capture the most model consistency information, while the subsequent principal components reflect the differences in model predictions. * It can balance the complexity and accuracy of the model. Too little may lead to underfitting, while too much may lead to overfitting. The cross-validation method is used to evaluate the error of each candidate parameter combination (ρ, K). The specific process is as follows:
[0160] Data partitioning and dimensionality reduction: Divide the sample data set into P parts, each part has s samples, and select one part each time as the new test set (i.e. the previous test set), and the rest as the new training set (i.e. the previous training set); the prediction matrix A of the new training set newtrain ∈R (n-s)×N Perform principal component analysis (PCA) to generate the projection matrix G of the first K principal components extracted newtrain ;
[0161] Model training and prediction: Calculate regression weights based on new training sets (The same as step S3 to calculate the second weight set), and the prediction matrix A of the new test set newtest Apply the weights to get the predicted values for the new test set
[0162] Error calculation and accumulation: Based on the predicted value and the true value Calculate the mean square error (MSE); accumulate the errors of all partitions to obtain the overall mean square error MSE (ρ, K) of the parameter combination (ρ, K);
[0163] Optimal parameter selection: traverse all combinations of ρ and K and select the one that minimizes the overall mean square error MSE(ρ, K) (ρ * , K * ):
[0164]
[0165] This process accurately evaluates the impact of different parameter combinations on the model's predictive performance through cross-validation, thereby effectively optimizing the model weights, balancing accuracy and complexity, and avoiding underfitting or overfitting.
[0166] S4. Fusing the first weight set and the second weight set to obtain a final weight set, and using the final weight set to weightedly combine the prediction values in the first prediction matrix to obtain a final prediction set.
[0167] In an implementation of this embodiment, the fusing the first weight set and the second weight set to obtain a final weight set, and using the final weight set to weightedly merge the prediction values in the first prediction matrix to obtain a final prediction set, specifically includes:
[0168] The first weight set and the second weight set of partition l are combined to obtain the final weight set of partition l
[0169]
[0170] in, and Respectively represent the first weight set of partition l The maximum and minimum weight values in v max and v min Respectively represent the second weight set v * The maximum and minimum weight values in ;
[0171] For each base learner, use the final set of weights for the partition l Perform weighted summation on the predicted values in the first prediction matrix to obtain the final predicted value of partition l
[0172] in, represents the predicted value of the j-th base learner on partition l, represents the final weight of the j-th base learner on partition l;
[0173] Merge the final prediction values of each partition to get the final prediction set
[0174]
[0175] in, Represents the final predicted value of partition L.
[0176] Specifically, after applying the RRMSE and PCR algorithms, two weight sets are generated on each partition l: RRMSE weight and PCR weight v * By fusing these two weight sets, the final weight of partition l is calculated and obtained After that, the test set (or sample data set) is predicted to generate the final output. For each partition l, combined with its corresponding weight And the prediction results of each base learner, calculate the final prediction value of the partition After getting the prediction results for each partition Later, merge the partitions to get the final complete prediction set of the test set
[0177] In summary, for regression tasks, the present invention has the following effects: 1) Improving the prediction accuracy of the model: Through the partition regression strategy, complex data is decomposed into smaller subsets, so that the model in each subset can more accurately fit the local data characteristics, avoiding the defect that the global model is difficult to capture details; 2) Balancing accuracy and relevance: Through the weights calculated by RRMSE and PCR, the prediction error and relevance of the model are taken into account, which effectively improves the generalization ability of the model, especially when processing high-dimensional data, it can reduce the risk of overfitting; 3) Ensemble learning enhances stability: The weighted integration of multiple base learners can effectively improve the robustness of the model, reduce the deviation that may be caused by a single model, and enhance the adaptability of the model to different partitioned data.
[0178] In addition, based on the above-mentioned ensemble learning weight assignment method based on learner accuracy and correlation applied to regression tasks, the present invention also provides an ensemble learning weight assignment method based on learner accuracy and correlation applied to classification tasks, and a preferred embodiment of the ensemble learning weight assignment method based on learner accuracy and correlation applied to classification tasks, such as Figure 2 and Figure 3 As shown, specifically including:
[0179] B1. Construct a decision tree to partition the sample data set, where each leaf node of the decision tree corresponds to a partition of the input feature space of the sample data set, and use multiple base learners to perform category prediction on the sample data set to obtain a second prediction matrix.
[0180] In an implementation of this embodiment, the construction of a decision tree performs partition processing on the sample data set, each leaf node of the decision tree corresponds to a partition of the input feature space of the sample data set, and a plurality of base learners are used to perform category prediction on the sample data set to obtain a second prediction matrix, specifically including:
[0181] Construct a tree with a depth of d and a number of leaf nodes of L = 2 d A decision tree is used to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition l of the input feature space of the sample data set;
[0182] Use N base learners to perform category prediction on the sample data set to obtain a second prediction matrix A i ∈R m×N ;
[0183] Among them, m is the number of categories, N is the number of base learners, R is a real number matrix, A i Represents N base learners for sample x i The second prediction matrix of , i represents the sample index.
[0184] Specifically, in the classification task, the partitioning strategy still uses decision trees, whose core principle is based on recursive splitting and information gain. This method uses decision trees to train the input training data set (X train ,y train ) to perform spatial segmentation to establish a mapping relationship between features and category labels. By constructing a decision tree classifier T with a depth of d, the feature space is recursively divided into L = 2 d regions, each region corresponds to a leaf node in the decision tree. Dataset X train The distribution of sample points in the feature space is The decision tree maps samples to a discretized leaf node index set {l1, l2, ..., lk}, where k≤2 d , each leaf node l corresponds to a data subset X l .
[0185] B2. Based on the second prediction matrix, use an accuracy algorithm to obtain a first weight set corresponding to the partition.
[0186] In an implementation of this embodiment, obtaining the first weight set corresponding to the partition by using an accuracy algorithm based on the second prediction matrix specifically includes:
[0187] For each base learner pair sample x i For predictions, define the indicator function for correct predictions:
[0188]
[0189] Among them, TP ij is the jth base learner for sample x i The correct predictive indicator, is the jth base learner for sample x i The predicted category, y i is the sample x i The corresponding true category, j represents the base learner index;
[0190] For partition l, calculate the prediction accuracy of each base learner in partition l:
[0191]
[0192] Among them, Accuracy j (l) represents the prediction accuracy of the j-th base learner in partition l, n l is the number of samples in partition l, X l Represents the sample data subset corresponding to partition l;
[0193] The prediction accuracy of each base learner in partition l is normalized to obtain the normalized weight:
[0194]
[0195] Among them, w l,j represents the normalized weight of the j-th base learner in partition l;
[0196] The normalized weights of all base learners in partition l form the first weight set of partition l
[0197]
[0198] Among them, wl,N Represents the normalized weight of the Nth base learner in partition l.
[0199] B3. Based on the second prediction matrix, obtain a second weight set using a principal component regression algorithm.
[0200] In an implementation of this embodiment, obtaining the second weight set based on the second prediction matrix by using a principal component regression algorithm specifically includes:
[0201] Construct a first covariance matrix C1 and a second covariance matrix C2, and combine the first covariance matrix C1 and the second covariance matrix C2 to obtain a comprehensive covariance matrix
[0202] Where n represents the number of samples, p(x i ) are all base learners for sample x i The predicted probability vector on the true category, and p(x i )∈R 1×N , for the i-th sample x i , the prediction result of the j-th base learner is f j (x i )∈R m×N , then the base learner is i The predicted probability can be defined as p(x i ) = f j (x i )[y i ]; p N (x i ) represents the Nth base learner for sample x i The predicted probability on the real category, μ1 is the mean vector of the predicted probability of the real category; Q(x i ) represents all base learners for sample x i The predicted probability vector on the non-true class, and Q(x i )∈R (m-1)×N ,q N (x i ) represents the Nth base learner for sample x i The predicted probability on the non-true class, It means that Q(x i ) Expand the high-dimensional tensor into a two-dimensional matrix Q by row priority flat ; μ2 is the mean vector of predicted probabilities of non-true categories; tr(·) is the trace operation of the matrix (i.e., the sum of the diagonal elements);
[0203] Using two-dimensional principal component analysis (2DPCA), the comprehensive covariance matrix As input, perform eigenvalue decomposition to obtain the eigenvalue corresponding to the eigenvalue vector, denoted as u1, u2, ..., u N ∈R N×N In other words, 2DPCA is used to reduce the dimension of the sample data set. The projection vector of 2DPCA is The feature vector of i The second prediction matrix A i , get the sample x i The corresponding principal component matrix
[0204]
[0205] Among them, u N represents the Nth eigenvector, u k represents the kth eigenvector, Represents sample x i The corresponding kth principal component;
[0206] The sample data set is divided into a training set and a test set by a cross-validation method, the overall prediction error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall prediction error among multiple overall prediction errors are determined as the optimal shrinkage parameter and the optimal number of principal components K. * , where the prediction error is 1 minus the accuracy;
[0207] Based on the optimal number of principal components K * , for the principal component matrix Perform dimensionality reduction and obtain K * principal components
[0208] Assume that the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the final estimate
[0209] in, Indicates K * The optimal principal component weights, Represents sample x i The corresponding K * principal components;
[0210] Different from the linear least squares regression in the regression problem to derive the optimal principal component weight vector, here we minimize the cross entropy loss L and use the stochastic gradient descent method to obtain the optimal principal component weight vector β:
[0211]
[0212] in, Represents sample x i The true label on category j; Represents sample x i The predicted probability of belonging to category j, express The logarithm of
[0213] according to Calculate and obtain the second weight set v of the multiple base learners of the original learning model * =[v1, ..., v N ]:
[0214]
[0215] Among them, v j represents the weight of the j-th base learner, β k represents the kth optimal principal component weight, u k,j Denotes the feature vector u k The jth component in .
[0216] Specifically, the PCC algorithm selects two covariance matrix calculation methods and combines them. The first covariance matrix C1 is used to describe the covariance relationship between the prediction probabilities of the base learners for the true category, and the second covariance matrix C2 describes the covariance relationship between the prediction probabilities of the base learners for the non-true category. In order to facilitate subsequent operations, Q(x i ) Expand the high-dimensional tensor into a two-dimensional matrix by row priority, i.e. Q flat ∈R n(m-1)×N , Indicates Q flat The i-th row of is a vector of dimension 1×N. In order to integrate the characteristics of covariance matrices C1 and C2 at the same time, the weight parameter of the PCC algorithm is fixed to 0.5 and normalized so that the sum of its main diagonal is equal to the number of base learners N. The final covariance matrix combines the information of the two covariance matrices, allowing the algorithm to better capture the structure of the data, effectively integrate the predictions of the base learners, and improve the accuracy and robustness of the model.
[0217] Specifically, the PCC algorithm first constructs the comprehensive covariance matrix Then two-dimensional principal component analysis 2DPCA is performed, and then the optimal number of principal components K for building the regression model needs to be determined.* ,The search process will be described in detail below, and finally the regression model is established to obtain the weights.
[0218] The core of parameter search is to select the optimal number of principal components K * The specific process is to divide the sample data set into P parts, each part is used as a new test set in turn, and the rest is used as a new training set. newtrain ,y newtrain ) calculates the covariance matrix and extracts the first K principal components. After the training data is projected into the principal component space, an optimization model is constructed and the weight vector β is trained by minimizing the cross entropy loss. The weight coefficient is calculated by combining the optimized weights with the principal component eigenvectors Used to make a weighted combination of predictions on the test data.
[0219] For the new test set data (X newtest ,y newtest ), generate the final prediction result based on the calculated weighted coefficients, and calculate the error (i.e. 1-ACC). Accumulate the errors of all partitions at each K value, and finally select the K that minimizes the cumulative error. * . Evaluate all candidate principal component numbers K∈{1, 2, ..., N}, calculate the cumulative error through cross-validation, and select the K with the smallest error * This process ensures that the selected K * Strike a balance between complexity and model performance to avoid underfitting or overfitting:
[0220]
[0221] B4. Fusing the first weight set and the second weight set to obtain a final weight set, using the final weight set to perform weighted summation on the prediction probabilities in the second prediction matrix to obtain a final prediction probability, and taking the category corresponding to the largest final prediction probability as the final prediction result.
[0222] In an implementation of this embodiment, the fusing of the first weight set and the second weight set to obtain a final weight set, using the final weight set to perform weighted summation on the prediction probabilities in the second prediction matrix to obtain a final prediction probability, and taking the category corresponding to the largest final prediction probability as the final prediction result, specifically includes:
[0223] The first weight set and the second weight set of partition l are combined to obtain the final weight set of partition l
[0224]
[0225] in, and Respectively represent the first weight set of partition l The maximum and minimum weight values in v max and v min Respectively represent the second weight set v * The maximum and minimum weight values in ;
[0226] For each base learner, use the final set of weights for the partition l Performing weighted summation on the prediction probabilities in the second prediction matrix to obtain the final prediction probability of partition l, and taking the category corresponding to the largest final prediction probability as the final prediction result;
[0227]
[0228] in, Represents the sample x i The final prediction result, f j (x i )[c] represents the jth base learner for sample x i The predicted probability of category c is, represents the final weight of the j-th base learner on partition l.
[0229] Specifically, after applying the Accuracy and PCC algorithms, two weight sets can be obtained: Accuracy weight and PCC weight v * The two weight sets are linearly combined, and the weight contribution of the Accuracy algorithm and the PCC algorithm is balanced by normalizing the maximum and minimum values to obtain the final weight of each base learner in each partition. Multiply the predicted probability of each base learner in the partition by the corresponding fusion weight and sum them up, and finally select the category with the largest predicted probability as x i The prediction results
[0230] In summary, for classification tasks, the present invention has the following effects: 1) Improving classification accuracy: Through the partitioning strategy, the features of each subset are more consistent, so that the classifier can make predictions more accurately in each local partition, avoiding the performance degradation of the global model on complex distribution data; 2) Combining Accuracy and PCC to improve the rationality of weight allocation: Combining the accuracy of the classifier in each partition with the PCC algorithm can not only reflect the accuracy of the model, but also optimize the weights through the covariance matrix to ensure better classification results; 3) Combining partitions with weights to enhance the generalization ability of the model: By integrating multiple base learners and combining the characteristics of different partitions with the classifier weights, the model's adaptability to different categories and different data distributions is improved; 4) Combining two covariance matrices: Comprehensively considering the influence of different data characteristics and distributions, the model's robustness to noise and outliers is enhanced. C1 reflects the stability under specific conditions, while C2 provides a grasp of the global trend. This combination effectively improves the adaptability of the classification algorithm in high-dimensional data and reduces the risk of overfitting.
[0231] In another implementation of this embodiment, whether it is a regression task or a classification task, the sample data set is divided into a training data set and a test data set, and the weights are calculated. The training data set is used to obtain the weights Then, use the weight Perform weighted calculation on the prediction results corresponding to the test data set.
[0232] To assist in the explanation, the following specific embodiment is provided, a method for allocating weights of base learners in ensemble learning, comprising the following steps:
[0233] Step 1: Data preprocessing.
[0234] Data preprocessing includes data cleaning, removing outliers, processing missing data and other operations to ensure the integrity and consistency of the data set. Feature engineering is then performed to standardize or normalize each feature value to ensure that features of different dimensions have the same scale to avoid affecting the training effect of the model due to dimensional differences. After processing, the data set is divided into a training set and a test set. The training set is used for model training, and the test set is used to evaluate the generalization performance of the classification model.
[0235] Step 2: Build a partition tree and partition the data.
[0236] For both classification and regression tasks, the training set is partitioned by decision trees. By setting the maximum depth, a partition tree is generated based on feature selection and data partitioning to ensure the consistency of sample categories in classification tasks or the similarity of numerical features in regression tasks. The test set is partitioned using the same partition tree to maintain consistency with the training set partition structure.
[0237] Step 3: Train the base learner and calculate the first weight.
[0238] In each partition, the base learners are trained independently. For regression tasks, the regressors are assigned as follows:
[0239] 5 decision trees: composed of multiple independent decision trees, which improve the model stability and prediction ability by fitting the training data multiple times;
[0240] Linear Regression 1: Evaluates the degree of linear relationship in the data and provides a baseline performance for the overall model;
[0241] 1 support vector machine: It uses radial basis function kernel to effectively handle nonlinear relationships in high-dimensional space and is suitable for modeling complex data structures;
[0242] 3 K-nearest neighbor algorithms: By setting different numbers of neighbors, we can explore the impact of the number of neighbors on prediction accuracy and adjust the local sensitivity of the model;
[0243] 5 multilayer perceptron models: They have good function approximation capabilities by learning nonlinear features layer by layer, and can flexibly control model complexity by adjusting the number of neurons;
[0244] In the classification experiment, the base learner can select a commonly used classification algorithm. The present invention selects 15 base learners, which are distributed as follows:
[0245] 6 decision trees: By splitting the data hierarchically, it can effectively select features and achieve classification;
[0246] 1 support vector machine: Using radial basis function (RBF) kernel, it can effectively handle nonlinear classification problems; the performance of SVM in high-dimensional space is particularly important for complex classification tasks and can improve classification accuracy;
[0247] There are three K-nearest neighbor algorithms: The KNN algorithm explores the effect of the number of neighbors on the classification accuracy by setting different numbers of neighbors; different numbers of neighbors (3, 5, and 7) are selected here to study the sensitivity and performance of the number of neighbors on the classification results;
[0248] 5 multilayer perceptron models: This model has good classification performance by learning the nonlinear characteristics of data layer by layer; in the classification task, the present invention adjusts the number of neurons and the network structure to explore the influence of networks of different complexity on the classification effect;
[0249] For regression tasks, the RRMSE indicator is used to measure the performance of the base learner in the regression task, and this indicator will be used for subsequent weight calculations. For classification tasks, after training is completed, the prediction results of the base learner are evaluated using the data set in the partition, and the classification accuracy of the base learner on the current partition is calculated.
[0250] Step 4: Calculate the second weight.
[0251] When calculating weights, regression and classification tasks use different weight calculation methods:
[0252] For regression tasks, the PCR algorithm is used to calculate weights. In this process, the data is first subjected to principal component analysis (PCA) to extract the main features, and the model weights are optimized in combination with the regression error. The goal of the PCR algorithm is to optimize the accuracy of regression prediction by reducing the covariance matrix error.
[0253] For classification tasks, the PCC algorithm is used to calculate weights. Within each partition, the PCC weights are calculated to measure the performance of the classifier on global data. The PCC algorithm generates weights for each base learner, reflecting its linear correlation with the classification target.
[0254] Step 5: Fusion weights.
[0255] Whether it is classification or regression, the final weight is a fusion of two parts: the local weight based on the performance of the base learner in each partition (such as the ACC weight for classification or the RRMSE weight for regression), and the global weight calculated by the global method (PCC or PCR); the specific fusion method is to normalize the two weights to ensure that the weights from different sources are in the same scale range; then use a weighted method to fuse them to obtain the final weight of each base learner; the final fused weight takes into account both the local performance in each partition and the performance of the base learner on the global data.
[0256] Step 6: Perform partition prediction on the test set.
[0257] After completing model training and weight calculation, the test set is partitioned using the previously constructed partition tree. For each partition, the trained base learner is called to predict the data of the partition. In the classification task, the classification results of each partition are weighted and integrated using the fused classifier weights to generate the final classification prediction result. For the regression task, similarly, the regression results of each partition are weighted and predicted based on the fused regressor weights to obtain the final regression prediction result.
[0258] The key innovation of the present invention is to partition the data by similarity and then perform weight-based fusion processing on the regression and classification tasks in different partitions. This partitioning strategy makes the data features more localized, thereby improving the model's adaptability to heterogeneous data. Compared with the prior art, the main differences and innovations of the present invention are as follows:
[0259] (1) Innovation of partitioning strategy: The present invention partitions data through decision trees or similar algorithms, and classifies more similar data into the same area for modeling; this strategy helps to better capture local features, especially in highly heterogeneous data sets.
[0260] (2) Weight fusion in regression tasks: In regression problems, the present invention innovatively combines the PCR algorithm and the RRMSE algorithm and proposes a unique weight allocation formula; this formula comprehensively considers the relevance weight of PCR and the accuracy weight of RRMSE, and by adjusting their respective weight intervals, ensures the optimal balance between accuracy and relevance of the model; this multi-factor fusion scheme is relatively rare in the prior art and significantly improves the accuracy of the regression model;
[0261] (3) Weight fusion in classification tasks: For classification problems, the present invention combines the PCC principal component classification algorithm and classification accuracy, and adopts a similar weight fusion strategy; by calculating the PCC weight of global data and the accuracy weight of local partitions, and performing reasonable fusion, the present invention effectively improves the accuracy and robustness of the model in the classification model of different partitions;
[0262] (4) Local and global combined weight allocation strategy: Another key innovation of the present invention is that the RRMSE weight in the regression task and the Accuracy weight (i.e., the first weight) in the classification task are calculated locally for each partition, while the PCR weight and PCC weight (i.e., the second weight) are calculated based on global data; this local and global combined weight allocation strategy can make full use of the overall structural information of the data while retaining the sensitivity to local features, significantly improving the generalization ability of the model.
[0263] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or terminal including the element.
[0264] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.
[0265] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A method for allocating weights for ensemble learning based on learner accuracy and relevance, characterized in that: The integrated learning weight allocation method based on learner accuracy and correlation is applied to regression tasks, and the integrated learning weight allocation method based on learner accuracy and correlation includes: Constructing a decision tree to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition of the input feature space of the sample data set, and using multiple base learners to predict the sample data set to obtain a first prediction matrix; Based on the first prediction matrix, using a relative root mean square error algorithm to obtain a first weight set corresponding to the partition; Based on the first prediction matrix, obtaining a second weight set using a principal component regression algorithm; The first weight set and the second weight set are fused to obtain a final weight set, and the prediction values in the first prediction matrix are weighted and combined using the final weight set to obtain a final prediction set.
2. The ensemble learning weight allocation method based on learner accuracy and correlation according to claim 1, characterized in that: The step of constructing a decision tree to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition of the input feature space of the sample data set, and using multiple base learners to predict the sample data set to obtain a first prediction matrix, specifically includes: Construct a tree with a depth of d and a number of leaf nodes of L = 2 d A decision tree is used to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition l of the input feature space of the sample data set; Use N base learners to predict the sample data set and obtain the first prediction matrix A∈R n×N , a ij =f j (x i ); Among them, n is the number of samples, N is the number of base learners, R is a real number matrix, A ij represents the value of the i-th row and j-th column in the first prediction matrix A, f j (x i ) represents the jth base learner for sample x i The predicted value of , i represents the sample index, and j represents the base learner index.
3. The ensemble learning weight allocation method based on learner accuracy and correlation according to claim 2, characterized in that: The obtaining, based on the first prediction matrix, a first weight set corresponding to the partition by using a relative root mean square error algorithm specifically includes: For each base learner, the relative root mean square error RRMSE is used to evaluate the performance of the base learner on partition l, and a constant term constant is introduced as the denominator: Among them, RRMSE l,j represents the relative root mean square error of the j-th base learner in partition l, n l is the number of samples in partition l, y i Represents sample x i The corresponding true value, is the mean of all true values in partition l, where l represents the partition index; According to the relative root mean square error RRMSE, a weight is assigned to each base learner in partition l: Among them, w l,j represents the weight of the j-th base learner in partition l, and ε represents the smoothing factor; Normalize the weight of each base learner in partition l to obtain the normalized weight: in, represents the normalized weight of the j-th base learner in partition l, w l,k represents the weight of the k-th base learner in partition l, and k represents the k-th base learner; The normalized weights of all base learners in partition l form the first weight set of partition l in, Represents the normalized weight of the Nth base learner in partition l.
4. The ensemble learning weight assignment method based on learner accuracy and correlation according to claim 3 is characterized in that: The obtaining a second weight set based on the first prediction matrix by using a principal component regression algorithm specifically includes: Construct the sample covariance matrix μ=[μ1,μ2,...,μ N ] T , σ=[σ1,σ2,...,σ N ] T , Among them, f(x i ) represents all base learners for sample x i The prediction vector of represents the normalized f(x i ), Represents the transposed T represents transpose, μ represents the mean vector, σ represents the standard deviation vector, μ j represents the jth component in the mean vector μ, μ N represents the Nth component in the mean vector μ, σ j represents the jth component of the standard deviation vector σ, σ N represents the Nth component in the standard deviation vector σ; Using principal component analysis, the sample covariance matrix As input, perform eigenvalue decomposition and sort according to the eigenvalue size λ1>λ2>...>λ N , output principal component matrix G = {g1,...,g N }: G i,j =γ j,1 f1(x i )+...+γ j,N f N (x i ); g k =c k,1 f1+...+c k,N f N ; c j =[γ j,1 ,...,c j,N ] T ; Among them, λ N represents the Nth eigenvalue, g N represents the Nth principal component, G i,j represents the value of the i-th row and j-th column in the principal component matrix G, γ j Denotes the eigenvalue λ j The corresponding eigenvector, γ j,N Denotes the eigenvector γ j The Nth component in N (x i ) represents the Nth base learner for sample x i The predicted value, g k represents the kth principal component, γ k,N Denotes the eigenvector γ k The Nth component in N Represents the prediction vector of the Nth base learner for all samples; The sample data set is divided into a training set and a test set by using the cross-validation method, the overall mean square error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall mean square error among multiple overall mean square errors are determined as the optimal shrinkage parameter and the optimal number of principal components K * ; Based on the optimal number of principal components K * , for the principal component matrix G = {g1,...,g N } to reduce the dimension and get K * principal components If the optimal principal component weight vector is According to the linear combination of the optimal principal component weight vector and K* principal components, the principal component linear combination prediction matrix f is obtained. PCR : in, Indicates K * The optimal principal component weights, β j represents the jth optimal principal component weight, g j represents the jth principal component; According to the least squares method, the square error function L(β) is defined to measure the linear combination prediction matrix f of the principal components. PCR The difference with the real matrix y: L(β)=||y-Gβ|| 2 =(y-Gβ) T (y-Gβ); The square error function L(β) is differentiated with respect to β, and the derivative is set to zero to obtain the optimal principal component weight vector β: y=[y1,...y n ] T ; Among them, y n Represents sample x n The corresponding true value; according to Calculate and obtain the second weight set v of the multiple base learners of the original learning model * =[v1,...,v N ]: Among them, v j represents the weight of the j-th base learner, f j represents the prediction vector of the j-th base learner for all samples, β k represents the kth optimal principal component weight, γ k,j Denotes the eigenvector γ k The jth component in .
5. The ensemble learning weight allocation method based on learner accuracy and correlation according to claim 4 is characterized in that: The fusing the first weight set and the second weight set to obtain a final weight set, and using the final weight set to weightedly combine the prediction values in the first prediction matrix to obtain a final prediction set, specifically includes: The first weight set and the second weight set of partition l are combined to obtain the final weight set of partition l in, and Respectively represent the first weight set of partition l The maximum and minimum weight values in v max and v min Respectively represent the second weight set v * The maximum and minimum weight values in ; For each base learner, use the final set of weights for the partition l Perform weighted summation on the predicted values in the first prediction matrix to obtain the final predicted value of partition l in, represents the predicted value of the j-th base learner on partition l, represents the final weight of the j-th base learner on partition l; Merge the final prediction values of each partition to get the final prediction set in, Represents the final predicted value of partition L.
6. A method for allocating weights for ensemble learning based on learner accuracy and correlation, characterized in that: The ensemble learning weight assignment method based on learner accuracy and relevance is applied to classification tasks, and the ensemble learning weight assignment method based on learner accuracy and relevance includes: Constructing a decision tree to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition of the input feature space of the sample data set, and using multiple base learners to perform category prediction on the sample data set to obtain a second prediction matrix; Based on the second prediction matrix, obtaining a first weight set corresponding to the partition using an accuracy algorithm; Based on the second prediction matrix, obtaining a second weight set using a principal component classification algorithm; The first weight set and the second weight set are integrated to obtain a final weight set, and the prediction probabilities in the second prediction matrix are weighted and summed using the final weight set to obtain a final prediction probability, and the category corresponding to the largest final prediction probability is taken as the final prediction result.
7. The method for allocating weights for integrated learning based on learner accuracy and correlation according to claim 6, characterized in that: The step of constructing a decision tree to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition of the input feature space of the sample data set, and using multiple base learners to perform category prediction on the sample data set to obtain a second prediction matrix, specifically includes: Construct a tree with a depth of d and a number of leaf nodes of L = 2 d A decision tree is used to partition the sample data set, wherein each leaf node of the decision tree corresponds to a partition l of the input feature space of the sample data set; Use N base learners to perform category prediction on the sample data set to obtain a second prediction matrix A i ∈R m×N ; Among them, m is the number of categories, N is the number of base learners, R is a real number matrix, A i Represents N base learners for sample x i The second prediction matrix of , i represents the sample index.
8. The method for allocating weights for integrated learning based on learner accuracy and correlation according to claim 7, characterized in that: The obtaining, based on the second prediction matrix, a first weight set corresponding to the partition by using an accuracy algorithm specifically includes: For each base learner pair sample x i For predictions, define the indicator function for correct predictions: Among them, TP ij is the jth base learner for sample x i The correct prediction indicator, is the jth base learner for sample x i The predicted category, y i is the sample x i The corresponding true category, j represents the base learner index; For partition l, calculate the prediction accuracy of each base learner in partition l: Among them, Accuracy j (l) represents the prediction accuracy of the j-th base learner in partition l, n l is the number of samples in partition l, X l Represents the sample data subset corresponding to partition l; The prediction accuracy of each base learner in partition l is normalized to obtain the normalized weight: Among them, w l,j represents the normalized weight of the j-th base learner in partition l; The normalized weights of all base learners in partition l form the first weight set of partition l Among them, w l,N Represents the normalized weight of the Nth base learner in partition l.
9. The method for allocating weights for integrated learning based on learner accuracy and correlation according to claim 8, characterized in that: The obtaining a second weight set based on the second prediction matrix by using a principal component regression algorithm specifically includes: Construct a first covariance matrix C1 and a second covariance matrix C2, and combine the first covariance matrix C1 and the second covariance matrix C2 to obtain a comprehensive covariance matrix p(x i )=[p1(x i ),p2(x i ),...,p N (x i )]; Q(x i )=[q1(x i ),q2(x i ),...,q N (x i )]; Where n represents the number of samples, p(x i ) are all base learners for sample x i The predicted probability vector on the true category, p N (x i ) represents the Nth base learner for sample x i The predicted probability on the real category, μ1 is the mean vector of the predicted probability of the real category; Q(x i ) represents all base learners for sample x i The predicted probability vector on the non-true class, q N (x i ) represents the Nth base learner for sample x i The predicted probability on the non-true class, It means that Q(x i ) Expand the high-dimensional tensor into a two-dimensional matrix Q by row priority flat The i-th row of ; μ2 is the predicted probability mean vector of the non-true category; tr(·) is the trace operation of the matrix; Using the two-dimensional principal component analysis method, the comprehensive covariance matrix As input, perform eigenvalue decomposition to obtain the eigenvalue corresponding to the eigenvalue vector, denoted as u1,u2,...,u N ∈R N×1 , for the i-th sample x i The second prediction matrix A i , get the sample x i The corresponding principal component matrix Among them, u N represents the Nth eigenvector, u k represents the kth eigenvector, Represents sample x i The corresponding kth principal component; The sample data set is divided into a training set and a test set by a cross-validation method, the overall prediction error corresponding to the test set is calculated, and the shrinkage parameter and the number of principal components corresponding to the minimum overall prediction error among multiple overall prediction errors are determined as the optimal shrinkage parameter and the optimal number of principal components K. * , where the prediction error is 1 minus the accuracy; Based on the optimal number of principal components K * , for the principal component matrix Perform dimensionality reduction and obtain K * principal components If the optimal principal component weight vector is According to the optimal principal component weight vector and K * The principal components are linearly combined to obtain the final estimate in, Indicates K * The optimal principal component weights, Represents sample x i The corresponding K * principal components; By minimizing the cross entropy loss L, the optimal principal component weight vector β is obtained using the stochastic gradient descent method: in, Represents sample x i The true label on category j; Represents sample x i The predicted probability of belonging to category j, express The logarithm of according to Calculate and obtain the second weight set v of the multiple base learners of the original learning model * =[v1,...,v N ]: Among them, v j represents the weight of the j-th base learner, β k represents the kth optimal principal component weight, u k,j Denotes the feature vector u k The jth component in .
10. The method for allocating weights for integrated learning based on learner accuracy and correlation according to claim 9, characterized in that: The fusing the first weight set and the second weight set to obtain a final weight set, using the final weight set to perform weighted summation on the prediction probabilities in the second prediction matrix to obtain a final prediction probability, and taking the category corresponding to the largest final prediction probability as the final prediction result, specifically includes: The first weight set and the second weight set of partition l are combined to obtain the final weight set of partition l in, and Respectively represent the first weight set of partition l The maximum and minimum weight values in v max and v min Respectively represent the second weight set v * The maximum and minimum weight values in ; For each base learner, use the final set of weights for the partition l Performing weighted summation on the prediction probabilities in the second prediction matrix to obtain the final prediction probability of partition l, and taking the category corresponding to the largest final prediction probability as the final prediction result; in, Represents the sample x i The final prediction result, f j (x i )[c] represents the jth base learner for sample x i The predicted probability of category c is, represents the final weight of the j-th base learner on partition l.