Relational data regression value prediction method based on composite ensemble learning
Through the composite ensemble learning method, data dimensionality reduction is performed by combining principal component analysis and uniform manifold approximate projection, and the base learner output is optimized using stacked ensemble learning model and gradient descent method, which solves the prediction accuracy and stability of large-scale, high-dimensional relational data, and is suitable for medical health and financial market analysis and other fields.
Patent Information
- Application Number
- CN202510330303.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
When the prior art processes large-scale, high-dimensional relational data, traditional methods are difficult to meet the requirements of accuracy and efficiency, and the diversity and output combination strategies of the basic learners in the integrated learning model are insufficient.
The composite ensemble learning method is adopted, combined with principal component analysis and uniform manifold approximation projection to reduce the data dimensions, and the base learner output is optimized by stacked ensemble learning model and gradient descent method, and prediction accuracy and stability are improved through multi-level ensemble learning strategies.
Significantly improves the accuracy and robustness of relational data regression predictions, and is suitable for a variety of data-intensive fields, especially in medical health and financial market analysis.
Smart Images

Figure CN120257230A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data analysis and machine learning, and relates to a method based on ensemble learning. Specifically, it relates to a method for predicting regression values of relational data based on composite ensemble learning, which is used for predicting regression values in relational data. This method is applicable to a variety of data-intensive fields, including healthcare, financial market analysis, environmental monitoring, bioinformatics, etc. Background Art
[0002] In the past few decades, the rapid development of information technology has made data an indispensable resource in modern society. Especially the large amount of structured data stored in relational databases plays a crucial role in multiple fields such as healthcare, finance, environmental monitoring, and bioinformatics, and has become a key factor driving the progress of these fields. With the explosive growth of data volume, how to extract valuable information from the vast and complex data has become a major challenge faced by the research and industrial communities. Traditional data analysis methods such as statistical analysis and linear regression perform well in dealing with small-scale and low-dimensional data sets, but when faced with modern large-scale and high-dimensional data sets, these methods often fall short. Especially in the context of relational data, the complexity and diversity of data make it difficult for traditional methods to meet the requirements in terms of accuracy and efficiency, and they lack flexibility in dealing with unknown patterns and hidden relationships.
[0003] In recent years, with the rapid development of machine learning technology, especially the emergence of ensemble learning methods, new ideas have been provided for solving the above problems. Ensemble learning is an advanced machine learning method that improves the overall prediction performance by combining multiple learning models. Its core idea is to jointly complete tasks through the cooperation of multiple learners to achieve performance that is difficult for a single learner to achieve. Compared with a single model, ensemble learning has proven to be able to significantly improve the accuracy and robustness of prediction in many practical applications. Common ensemble learning methods include Random Forest, Gradient Boosting, AdaBoost, etc., which perform well in dealing with complex and high-dimensional data sets. Especially in the regression field, they focus on predicting numerical outputs, and improve the prediction accuracy by reducing prediction errors and increasing the consistency of results.
[0004] Although ensemble learning has achieved remarkable achievements in many fields, in the specific implementation process, how to ensure the diversity among the individual base learners that make up the ensemble learning model, and how to effectively combine the outputs of these base learners, are still problems that need to be solved urgently. Summary of the Invention
[0005] In view of the problems existing in the prior art, the present invention proposes a method for predicting regression values of relational data based on composite ensemble learning. The data processed in the present invention is relational data, including but not limited to medical records, financial transaction records, social media interaction data, and environmental monitoring data, etc. These relational data usually have the characteristics of high structurality, presented in the form of tables, where each row represents an independent observation instance, and each column represents a specific attribute or feature. A significant feature of relational data is its multidimensionality, that is, each observation instance may contain dozens to hundreds of features, covering various types such as numerical data, categorical data, and time series data. Facing such high-dimensional data, directly performing machine learning modeling will encounter challenges such as the curse of dimensionality, data sparsity, and complex correlations between features. Based on this, the present invention aims to improve the efficiency of data preprocessing, increase the diversity between base learners, and significantly improve the accuracy of regression prediction through an optimized learner combination strategy by combining advanced feature engineering methods and multiple different types of learners, thereby enhancing the model's prediction ability for unknown data and making it more robust and reliable in real-world applications. In addition, the present invention also adopts efficient model training methods, such as k-fold cross-validation and gradient descent method, to further enhance the generalization ability and robustness of the model. These innovative technical combinations can not only effectively address the challenges of large-scale, high-dimensional relational data, but also show broad application potential in related data-intensive fields.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A method for predicting regression values of relational data based on composite ensemble learning, aiming to improve the accuracy and efficiency of data analysis and machine learning techniques. First, use principal component analysis and uniform manifold approximation and projection to reduce the dimension and select features of the data, which can effectively extract the main features of the data and reduce the dimension of the data, reduce the computational complexity, and at the same time retain the main information in the data. Then, use two stacked ensemble learning models to predict the dimension-reduced data respectively. These two models are trained and predicted using different base learners, making full use of the advantages of each base learner to improve the prediction accuracy. At the same time, the gradient descent method is also used to optimize the output weights of the base learners to improve the prediction accuracy. Finally, use the third stacked ensemble learning model to integrate the output results of the first two models. This model also uses the gradient descent method to optimize the output weights to improve the prediction accuracy. This multi-level ensemble learning strategy can effectively improve the stability of the prediction and the generalization ability of the model. The method of the present invention has generality and scalability and can be applied to a variety of data-intensive fields. The method includes the following steps:
[0008] In the first step, for relational data with regression value labels, principal component analysis (PCA) is used for feature dimensionality reduction to generate a reduced-dimensional dataset. Subsequently, a customized stacking ensemble learning model is employed to process the dataset reduced by PCA for regression value prediction, obtaining the first predicted regression value result in the overall architecture. This stacking ensemble learning model combines two different neural network models and a gradient boosting decision tree (GBDT) as sub-models to provide a comprehensive prediction result. Here, a sub-model refers to an individual machine learning method in stacking ensemble learning, each making independent predictions, and the final output result of the stacking ensemble learning model is a set or optimized version of these predictions; a neural network is a node-layer-based model capable of learning complex patterns in data; a gradient boosting decision tree is an ensemble technique that optimizes prediction performance by gradually building decision trees. The combination of these two methods improves the adaptability and accuracy of the model. While processing the data in the first step, the method in the following second step is simultaneously used to process the data, and the first step and the second step run in parallel. Through the above first step, the first predicted regression value result is generated.
[0009] In the second step, which runs in parallel with the first step, for the same relational data with regression value labels as in the previous step, uniform manifold approximation and projection (UMAP) is used to perform non-linear dimensionality reduction on the same relational data, generating another reduced-dimensional dataset. Subsequently, another customized stacking ensemble learning model is used to process the dataset reduced by UMAP for regression value prediction, obtaining the second predicted regression value result in the entire architecture. This stacking ensemble learning model is specifically optimized for the characteristics of UMAP data and integrates specially selected base learners, including two different gradient boosting decision trees (GBDT) and a neural network (NeuralNet), to output a comprehensive prediction result. Compared with the first model, the significant difference in this model is the increase in the number of gradient boosting decision trees, and the structure of each tree is different, mainly to optimize the specific features of the UMAP-reduced dataset processed by each tree, thereby improving the prediction performance and accuracy of the overall model. Through this step, the second predicted regression value result is generated.
[0010] Step 3: After Steps 1 and 2 are completed, the present invention uses a third customized Stacking integrated learning model to integrate the regression value results output by the two Stacking integrated learning models in Steps 1 and 2. This final Stacking model further analyzes and fuses the prediction results from the previous two customized models to generate the final output of the composite integrated learning framework of the present invention. This design not only enhances the model's adaptability to data diversity but also optimizes prediction accuracy and robustness. In this step, the Stacking integrated learning model used includes three sub-models: Random Forest Regression (RandomForestMSE), Neural Network (NeuralNetFastAI), and Extreme Gradient Boosting (XGBoost). Among them, Random Forest Regression is an ensemble method based on decision trees that improves the accuracy and stability of the model through the average prediction values of multiple decision trees; while Extreme Gradient Boosting (XGBoost) improves the model's accuracy by optimizing the speed and performance of the traditional gradient boosting algorithm and is suitable for processing large-scale datasets. The combination of these sub-models optimizes the regression value prediction results of Steps 1 and 2 and finally provides a more accurate and comprehensive prediction result output.
[0011] The following part will elaborate on the specific structure of the present invention, as Figure 1 shown.
[0012] Step 1: Use Principal Component Analysis (PCA) for dimensionality reduction to generate a dataset with reduced dimensions. Subsequently, use the first customized Stacking integrated learning model to process the dataset with reduced dimensions by PCA for regression value prediction, specifically as follows:
[0013] Step 1.1) Use Principal Component Analysis (PCA) for dimensionality reduction to generate a dataset with reduced dimensions. The Principal Component Analysis (PCA) is a commonly used method for data dimensionality reduction and feature selection. Its purpose is to find a new set of orthogonal coordinate axes such that the projection of the data on these axes can maximize the variance of the data, thereby retaining the main information in the data. The mathematical principle of Principal Component Analysis is to perform eigenvalue decomposition on the covariance matrix of the data to obtain a set of eigenvectors and eigenvalues. These eigenvectors are the new coordinate axes, and the eigenvalues represent the variance of the data on these axes. Usually, select the eigenvectors corresponding to the largest several eigenvalues as the principal components, and then project the data onto these principal components to achieve data dimensionality reduction and feature selection. The steps of Principal Component Analysis are as follows:
[0014] First, obtain the original data through a database or relevant data files. The original data is relational data, including but not limited to medical records, financial transaction records, social media interaction data, and environmental monitoring data, etc., and assume the original data matrix is X ∈ R n×d, where n is the number of samples and d is the feature dimension. Centralize the original data, that is where is the mean vector of the data.
[0015] Subsequently, calculate the covariance matrix of the data:
[0016]
[0017] Finally, perform eigenvalue decomposition on the covariance matrix to obtain C = WΛW T , where W ∈ R d×d is the eigenvector matrix, and Λ ∈ R d×d is a diagonal matrix, and its diagonal elements are eigenvalues. Subsequently, select the eigenvectors corresponding to the largest k eigenvalues to form a matrix W k ∈ R d×k , where k < d. Project the data onto W k to obtain the dimension-reduced data Y = XW k ∈ R n×k .
[0018] Through the above steps, the process of reducing the data from d dimensions to k dimensions is completed, where d represents the total number of features or variables in the original dataset, that is, the dimension of the data. And k represents the number of principal components selected to be retained after principal component analysis. It is an integer less than d and represents the new dimension after dimension reduction. In this way, the dimension reduction process from d dimensions to k dimensions not only reduces the complexity of the data but also retains the most important information for more effective subsequent analysis and learning.
[0019] Step 1.2) Use the first independent stacked ensemble learning model (Stacking) to process the dimension-reduced dataset obtained from the first step through PCA. This customized stacked ensemble learning model is specifically trained to optimize and adapt to the characteristics of the PCA dimension-reduced data. Through this arrangement, the model can maximize the utilization of the characteristics of its input dataset and produce a highly targeted regression value prediction result.
[0020] The above stacked ensemble learning model combines two different machine learning algorithms, namely neural network (NeuralNet) and gradient boosting decision tree (GBDT), and three models as sub-models. At the same time and in parallel, it uses the different characteristics of these algorithms to process the data after PCA dimension reduction, and respectively predicts the regression value prediction result for each piece of data based on the PCA dimension-reduced data, in order to use the subsequent gradient descent method to integrate their respective results to obtain the prediction regression value result of the first stacked ensemble learning model.
[0021] The following content will introduce in detail the implementation details of two different types of sub - models in the first Stacking integrated learning model: Neural Network and Gradient Boosting Decision Trees (GBDT) in the above process.
[0022] Neural Network In this model, the neural network realizes data learning and prediction through several important mathematical steps. In the network, the output z(l) of each layer is transformed by the activation function σ of the output of the previous layer, and the calculation process is as follows:
[0023] z(l + 1)=σ(W(l)z(l)+b(l))(2)
[0024] Among them, W(l) and b(l) represent the weight and bias of the l - th layer respectively. The activation function σ is usually selected such as ReLU or Sigmoid to introduce non - linearity so that the network can learn complex data patterns. The loss function adopts the mean square error (MSE), and the specific form is:
[0025]
[0026] where, y i is the target value, is the network prediction value. The weight update is performed through the backpropagation algorithm, and the weight update formula is:
[0027]
[0028] In the present invention, based on the above - mentioned basic principle of the neural network, two different neural networks are used as sub - models of the first Stacking integrated learning model: NeuralNet - FastAI and NeuralNet - pyTorch. NeuralNet - FastAI is based on the FastAI library, with high - level encapsulation and convenient APIs, suitable for rapid development and experimentation. NeuralNet - pyTorch is based on the PyTorch framework, providing high flexibility and detailed control, suitable for the customization and optimization of complex models. The combination of the two enables the model to give full play to their respective advantages in different scenarios, improving the prediction performance and accuracy.
[0029] Gradient Boosting Decision Trees (GBDT) is a powerful algorithm that reduces the prediction error by gradually constructing decision trees. First, the initial model f0(x) provides a baseline prediction based on the training data. The goal of each subsequent step is to find a tree model h t (x) to fit the residual r i,t :
[0030]
[0031] The learning objective of each tree is to minimize the following loss function, usually the squared error loss, to fit these residuals:
[0032]
[0033] During the iteration process, the model update is expressed as:
[0034] f t(x) = f (t-1) (x)+vh t (x) (7)
[0035] where v is the learning rate, which adjusts the step size contributed by each tree. This cumulative model can effectively reduce the overall training error and improve the prediction accuracy by integrating multiple simple models.
[0036] Step 1.3) After the above process, two sub-models, namely the Neural Network and the Gradient Boosting Decision Tree (GBDT), independently learn and predict the data after PCA dimensionality reduction, and obtain the regression value results for each piece of data prediction. Subsequently, the output weights of the sub-models are allocated by the gradient descent method, which is an optimization algorithm that iteratively adjusts parameters to minimize the error of a function and thus find the minimum value. Specifically, the present invention uses the gradient descent method to minimize this loss function according to the loss function, thereby obtaining the output weights of each model. The above loss function is a function used to measure the difference between the model prediction result and the actual result, and is a mathematical expression for quantifying the error degree of the model prediction. In machine learning, common loss functions include the mean squared error (MSE), the root mean squared error (RMSE), and cross-entropy, etc. The purpose is to reduce this difference through the optimization process to optimize the parameters of the model. In the present invention, the loss functions that can be selected mainly include the mean squared error (MSE) and the root mean squared error (RMSE), and these loss functions are particularly suitable for regression tasks to help optimize and evaluate the performance of the model when dealing with continuous numerical predictions. This process can be expressed by the following formula:
[0037]
[0038] where w is the output weight of the model, η is the learning rate, is the gradient of the loss function J(w) with respect to the weight w.
[0039] At the same time, during the training process of each of the above sub-models, the present invention also adopts the method of k-fold cross-validation. Specifically, the data set D is divided into k subsets, denoted as D1, D2,..., D k, and then perform k times of training and validation. In the i-th (i = 1, 2,..., k) training and validation, use other subsets except D_i as training data, that is, D train = D - D i , and use D i as validation data, that is, D valid = D i . In this way, it can be ensured that the performance of the model will not fluctuate greatly due to different data partitioning methods. This process can be expressed by the following formula:
[0040] D train = D - D i , i = 1, 2,…, k (9)
[0041] D valid = D i , i = 1, 2,…, k (10)
[0042] where D is the entire dataset, D i is the i-th subset, D train and D valid are training data and validation data respectively.
[0043] After the above training is completed, the first stacked ensemble learning model outputs the first comprehensive regression value prediction result, which will be used as the first part of the final ensemble result.
[0044] Second step, which is carried out in parallel with the first step above. First, use the Uniform Manifold Approximation and Projection (UMAP) method to perform non-linear dimensionality reduction on the original input relational data, generating a second piece of dimensionality-reduced data, whose goal is to preserve the local and global structures of the high-dimensional data as much as possible in the low-dimensional space. Then, use another customized stacked ensemble learning model (Stacking) to process the dataset after dimensionality reduction by the Uniform Manifold Approximation and Projection (UMAP) to predict the regression value, obtaining the first predicted regression value result in the overall architecture, specifically as follows:
[0045] Step 2.1) Use the Uniform Manifold Approximation and Projection (UMAP) method to perform non-linear dimensionality reduction on the original input relational data. The mathematical model of the Uniform Manifold Approximation and Projection (UMAP) is based on a complex topological structure, and the following is its implementation process:
[0046] First, construct a local neighborhood graph for each point in the high-dimensional space. The so-called high-dimensional space refers to a space with a very large number of features or attributes, where each feature represents a dimension; the local neighborhood graph is a structure that shows the local relationships between points in the high-dimensional space by connecting each data point and its nearest neighbors.
[0047] Subsequently, a data representation is sought in a low-dimensional space to minimize the topological differences between the high-dimensional space and the low-dimensional space. The low-dimensional space is a space with fewer features than the original high-dimensional space. In this space, the data is simplified to reduce the dimension, aiming to retain the most critical information while reducing the computational burden and improving the efficiency of data processing. The following objective function is usually optimized using the gradient descent method:
[0048]
[0049] Among them, d high and d low Represent the distance in high-dimensional space and low-dimensional space respectively. Here, 'distance' refers to the measure of the difference or similarity between two data points, which is usually calculated by a specific distance formula (such as Euclidean distance or Manhattan distance) to evaluate the proximity or similarity between data points. Through the above process, the original input relational data is reduced in dimension. After processing, the low-dimensional data set generated by UMAP retains the key structural features of the high-dimensional data. This process optimizes the objective function and accurately adjusts the low-dimensional representation to reflect the local and global topological structure of the high-dimensional data, so that similar data points are also kept close in the low-dimensional space, while dissimilar data points are separated by a long distance. Among them, the topological structure refers to the relative position relationship between data points and their connection mode. These relationships remain unchanged as much as possible during the dimensionality reduction process to ensure that the low-dimensional representation can accurately reflect the essence of the high-dimensional data. Another set of reduced-dimensional data after feature engineering is obtained by the above process. Using this data, the present invention adopts the following steps for further processing.
[0050] Step 2.2) After the original input relational data is subjected to nonlinear dimensionality reduction processing using the Uniform Manifold Approximation Projection (UMAP) method to obtain the second dimensionality reduction data, another independent stacking ensemble learning model (Stacking) is used to process the second dimensionality reduction data set. This customized stacking ensemble learning model is specially designed to optimize and adjust to the characteristics of the dimensionality reduction data using the Uniform Manifold Approximation Projection (UMAP) method. Through this arrangement, the model can maximize the use of the characteristics of its input data set to produce a second highly targeted regression value prediction result.
[0051] Among them, the second stacked integrated learning model consists of two different Gradient Boosting Decision Trees (GBDT) and a Neural Net. Compared with the first model, the significant difference of this model lies not only in the increase in the number of gradient boosting decision trees, but more in the differences in the structures of these trees. These structural differences mainly stem from the differences in the training data sets used by each tree, enabling each tree to be optimized specifically for the specific features of its data set. The following are the implementation processes of XGBoost and LightGBM and their specific differences in the implementation:
[0052] The objective function of XGBoost is
[0053]
[0054] where y i is the true value of the i-th data point, is the predicted value, n is the total number of data points, is the loss function, which is used to calculate the difference between the true value and the predicted value. K represents the number of tree models used, f k is the k-th tree model, and Ω(f k ) represents the complexity of this tree model.
[0055] The objective function of LightGBM is expressed as
[0056]
[0057] It also includes the same true value, predicted value, total number of data, loss function as XGBoost, as well as the regularization coefficient λ and the regularization term Regularization. The latter usually involves penalty terms for model parameters, such as L1 or L2 norms, to control the model complexity and avoid overfitting. These two models achieve an efficient learning process through different regularization methods and tree structure optimizations, thus providing customized optimization strategies for different data set characteristics, which also reflects a part of the flexible and targeted integrated learning strategy of the present invention.
[0058] In addition, the principle of the sub-model Neural Net in this step is the same as that in the first step and is implemented based on pyTorch.
[0059] The integration weights of the regression value prediction results of the three sub-models in the second step are also assigned by the gradient descent method.
[0060] Through the optimization process of the second stacked ensemble learning model described above, the model can effectively utilize the dimensionality-reduced dataset obtained by the Uniform Manifold Approximation and Projection (UMAP) method, generating an accurate regression value prediction result specifically tailored to the characteristics of the dimensionality-reduced data. This prediction result not only captures the local and global structural features of the data, but also, due to the model's adoption of different learning algorithms and structural optimizations, can better adapt to the diversity and complexity of the data. The obtained regression value prediction result will be passed as an independent output to the next stacked ensemble learning model (Stacking), and this output, together with the outputs of other models, will be further integrated and optimized to form the final comprehensive prediction result. This integration process aims to further improve the prediction accuracy and the robustness of the model.
[0061] In the third step, the output results of the first and second steps are used as inputs, and these results are integrated through a third stacked ensemble learning model (Stacking). In this step, the stacked ensemble learning model used includes three sub-models: Random Forest Regression (RandomForestMSE), Neural Network (NeuralNetFastAI), and Extreme Gradient Boosting (XGBoost). These three sub-models operate in parallel, independently processing the input data and outputting prediction results.
[0062] Random Forest Regression (RandomForestMSE) is an ensemble method based on decision trees, which improves the prediction accuracy and stability by constructing multiple decision trees and averaging their results. Its objective function is expressed as:
[0063]
[0064] where yi is the true value of the i-th data point, is the predicted value, and n is the total number of data points. Random Forest trains by constructing a large number of decision trees (each tree using different random subsets and features), and the final output is the average of the prediction results of all trees, thus reducing the overfitting problem of a single model.
[0065] Extreme Gradient Boosting (XGBoost) is an improved gradient boosting method that optimizes the tree structure by performing a second-order Taylor expansion on the objective function. Its objective function is expressed as:
[0066]
[0067] where, is the loss function, which is used to calculate the difference between the true value and the predicted value, n is the total number of data points, K represents the number of tree models used, and f k is the k-th tree model, and Ω(fk ) represents the complexity of the tree model, usually including a regularization term to control the complexity of the model and prevent overfitting. By using these three sub-models in parallel, each model independently processes the input data and outputs prediction results. Subsequently, these regression value prediction results are optimized through weighted combination by the gradient descent method to further improve the overall prediction accuracy and robustness. Finally, the model in the third step outputs a comprehensive and optimized regression value prediction result, which integrates the prediction capabilities and characteristics of all models in the first two steps and provides solid data support for the final decision-making.
[0068] The beneficial effects of the present invention are mainly reflected in the following aspects:
[0069] (1) First of all, by comprehensively using a variety of machine learning methods and technologies, the present invention effectively improves the accuracy and efficiency of the prediction task. In particular, the present invention adopts two advanced feature engineering methods, namely principal component analysis (PCA) and uniform manifold approximation and projection (UMAP), to reduce the dimension and select features of the data, thereby improving the effectiveness and processing efficiency of the data. In addition, the present invention also adopts a stacking ensemble learning model (Stacking), and further improves the prediction accuracy and robustness by combining multiple different base learners.
[0070] (2) Secondly, during the process of training each model, the present invention adopts the method of k-fold cross-validation to ensure that the performance of the model will not fluctuate greatly due to different data partitioning methods. This method can effectively prevent the model from overfitting and improve the generalization ability of the model.
[0071] (3) Thirdly, the method of the present invention has good versatility and scalability. It can be applied to a variety of data-intensive fields, such as healthcare, financial market analysis, environmental monitoring, etc., and has significant advantages in dealing with large-scale and high-dimensional data sets. At the same time, since the present invention adopts a modular design, it can be flexibly adjusted and replaced according to specific task requirements and data characteristics to achieve the best prediction effect. The method of the present invention has versatility and scalability and is applied to a variety of data-intensive fields. For example, in the prediction task of the disease severity of a chronic obstructive pulmonary disease patient, the present invention shows superior performance, and there are significant improvements in various evaluation indicators compared with a single stacking ensemble learning model.
[0072] Generally speaking, the present invention provides an effective method for predicting regression values of relational data based on composite ensemble learning, which has important theoretical significance and practical value for promoting the development of data analysis and machine learning technologies and facilitating their wide application in various practical applications. Description of the Drawings
[0073] Figure 1 It is a schematic diagram of the overall structure of the model. Specific implementation manners
[0074] The present invention will be further described below in conjunction with specific embodiments and the accompanying drawings. The present invention relates to a method based on ensemble learning for predicting regression values in relational data. In the present invention, the input data adopts a structured data format extracted from a relational database or a related file, where each row represents a unique instance and each column corresponds to a different feature or attribute. Based on the above data, in the first step, the principal component analysis (PCA) method is first used to reduce the dimension and select features of the data, generating the first reduced-dimension data. Subsequently, the first customized stacking ensemble learning model (Stacking) is used to predict the PCA-reduced data, obtaining the first predicted regression value result. In the second step, the uniform manifold approximation and projection (UMAP) method is used to reduce the dimension and select features of the original data, generating the second reduced-dimension data. Subsequently, the second customized stacking ensemble learning model (Stacking) is used to predict the UMAP-reduced data, obtaining the second predicted regression value result. Finally, in the third step, a further customized stacking ensemble learning model (Stacking) is used to integrate the above two outputs to obtain the final result, and the gradient descent method is also used to allocate the weights of the sub-models in this step.
[0075] The experimental dataset of the present invention is derived from the clinical trial data and baseline data of patients in the respiratory department of a certain hospital, and the above data type is relational data. The research data of the present invention was collected with the informed consent of all participants. The dataset has a total of 329 patients, including 229 patients with chronic obstructive pulmonary disease (COPD) and 100 control patients. Among them, there are 188 males and 141 females. The average age of all patients is 60.22 years old, the median age is 62 years old, the standard deviation is 13.33 years old, the maximum age is 86 years old, and the minimum age is 21 years old. The information of each patient includes the following 19 indicators and a diagnosis: gender, age, smoking history, rhinitis history, white blood cell count, eosinophil count, eosinophil percentage, C-reactive protein (CRP), percentage of predicted forced vital capacity (FVC% predicted < 80), ratio of maximum expiratory flow at 75 / 25 (MEF75 / 25 < 65%), maximum expiratory flow at 50 (MEF50), maximum expiratory flow at 25 (MEF25), maximum ventilation volume (MVV 80%), maximum vital capacity (VC MAX > 40%), tidal volume (VT), blood pH value (PH value), partial pressure of oxygen (PO2), partial pressure of carbon dioxide (PCO2), whether suffering from bronchitis or emphysema, and the diagnostic label value is the ratio of forced expiratory volume in one second to the predicted value (FEV1 / predicted value).
[0076] The diagnostic criteria for chronic obstructive pulmonary disease (COPD) are usually based on the results of pulmonary function tests. When the ratio of forced expiratory volume in one second (FEV1) to forced vital capacity (FVC) is less than 0.7, it indicates the presence of COPD. According to the severity grading of the World Health Organization, for mild COPD, the FEV1 / FVC ratio is less than 0.7 and FEV1 ≥ 80% predicted value; for moderate COPD, FEV1 / FVC < 0.7 and 50% ≤ FEV1 < 80% predicted value; for severe COPD, FEV1 / FVC < 0.7 and 30% ≤ FEV1 < 50% predicted value; for very severe COPD, FEV1 / FVC < 0.7 and FEV1 < 30% predicted value. However, although pulmonary function tests can accurately diagnose COPD, some patients are unable to undergo such examinations for various reasons. For example, for patients with severe heart diseases, performing pulmonary function tests may increase their health risks; in addition, patients with impaired cognitive function and those in the intensive care unit may also have difficulty undergoing pulmonary function tests. Based on the regression value prediction method of the present invention, using the relevant data of relevant patients to complete the prediction of the FEV1 regression value of patients can assist doctors in better judging the severity of the disease, especially for those patients who are unable to perform standard pulmonary function tests.
[0077] The present invention has been verified in terms of the disease severity of the above-mentioned COPD patients, but the application scenario is not limited to this field. This method is applicable to a variety of data-intensive fields, including medical and health, financial market analysis, environmental monitoring, bioinformatics, etc. The following is the specific implementation manner of the method of the present invention in the COPD severity dataset. It includes the following steps:
[0078] (1) The first part of the present invention is to use principal component analysis (PCA) to complete feature dimensionality reduction, and use a customized stacking ensemble learning model (Stacking) to process the dataset after dimensionality reduction by principal component analysis (PCA) to perform preliminary regression value prediction. Through this step, the first predicted regression value result is obtained. Next, this result is used as input and passed to the multi-layer stacking ensemble learning model (Stacking) in the third step for further integration.
[0079] Specifically, for principal component analysis (PCA): First, calculate the covariance matrix of the data, and then calculate the eigenvectors and eigenvalues of this covariance matrix. These eigenvectors are the principal components as described in the present invention. In the present invention, the eigenvectors corresponding to the 15 largest eigenvalues are selected, and then the original data is projected onto these 15 eigenvectors to obtain the dimensionality-reduced data. Let the original data matrix be X, the covariance matrix be C = XTX, and the eigenvector matrix be W, then the dimensionality-reduced data Y can be calculated by the following formula: Y = XW. In this way, the process of reducing the features to 15 principal components is completed. In specific implementation, the fit_transform method in the sklearn.decomposition.PCA library is used to reduce the dimensionality of the feature matrix, where the number of dimensions of the reduced features parameter (n_components) is 15.
[0080] Subsequently, the first stacked ensemble learning model (Stacking) is used to process the data generated by principal component analysis (PCA). The first-layer structure of this model contains three machine learning models: LightGBMXT (Gradient Boosting Decision Tree), NeuralNet-FastAI (Neural Network 1), and NeuralNet-pyTorch (Neural Network 2), which are implemented based on the LightGBM (3.3.5), fastai (2.7.12), and PyTorch (1.13.1) libraries respectively.
[0081] The training of the above models is completed using the k-fold cross-validation method, and the value of k is selected as 5. Finally, the gradient descent method is used to output the results of the above models. This process can be represented by the following formula: where w is the output weight of the model, η is the learning rate, is the gradient of the loss function J(w) with respect to the weight w. The learning rate is set to 0.001, and the final weights assigned to the above three models are: Gradient Boosting Decision Tree: 0.113, Neural Network 1: 0.465, Neural Network 2: 0.422.
[0082] (2) The second part of the present invention is to use Uniform Manifold Approximation and Projection (UMAP) to complete feature dimensionality reduction, and use a customized stacked ensemble learning model (Stacking) to process the dataset dimensionality-reduced by Uniform Manifold Approximation and Projection (UMAP) for preliminary regression value prediction. Through this step, the second predicted regression value result is obtained. Next, this result is also passed as an input to the multi-layer stacked ensemble learning model (Stacking) in the third step for further processing.
[0083] Specifically, Uniform Manifold Approximation and Projection (UMAP): First, create a UMAP instance with adjustable parameters including the target number of dimensions after dimensionality reduction. Then, use the UMAP instance to fit the data and transform it into a low-dimensional representation. After this step, a new array is returned, where each row is a low-dimensional representation of the original data. The goal of UMAP is to find a mapping function f: X → Y, where X represents the high-dimensional space and Y represents the low-dimensional space (in this example, the dimension is 9). Through the above steps, the present invention achieves dimensionality reduction from 20 features of the original dataset to 9 features while preserving the topological structure in the original data as much as possible. The above process is implemented through the UMAP method in the umap.umap_ library, and the parameter of the number of dimensions of the features after dimensionality reduction (n_components) is 9.
[0084] After the above process, a second Stacking ensemble learning model is used to process the data generated by UMAP. The first layer structure of this model includes three machine learning models: LightGBM, XGBoost, and NeuralNet-pyTorch, which are implemented based on LightGBM (3.3.5), XGBoost (2.7.12), and PyTorch (1.13.1) respectively. Subsequently, the k-fold cross-validation method is used to complete the training of the above models, and the value of k is selected as 5. Finally, the gradient descent method is also used to complete the result output of the above models, and the learning rate is set to 0.001; the final weights assigned to the above three models are: LightGBM: 0.290, XGBoost: 0.058, NeuralNet: 0.652.
[0085] (3) The third part of the present invention further uses a Stacking ensemble learning model to integrate the two output results in the second step. The sub-models of the Stacking ensemble learning model used in the third step include RandomForestMSE, NeuralNet-FastAI, and XGBoost, which are implemented based on the scikit-learn (1.2.1), fastai (2.7.12), and XGBoost (2.7.12) libraries respectively. The gradient descent method is also used to integrate the above outputs, and the learning rate is set to 0.001; the final output parameters of each model are: RandomForestMSE: 0.293, XGBoost: 0.173, NeuralNet: 0.534.
[0086] Combined with the solution of the present invention, the experimental results are presented and analyzed as follows:
[0087] (1) Experimental data and evaluation indicators
[0088] The present invention selects the following evaluation indicators to evaluate the performance of the model:
[0089] Mean Absolute Error (MAE): This is a commonly used evaluation indicator for regression problems, which calculates the average absolute difference between the predicted value and the true value. Its formula is:
[0090]
[0091] where, y i is the true value of the i-th sample, is the predicted value of the model for the i-th sample, and n is the number of samples, the same hereinafter.
[0092] Root Mean Squared Error (RMSE): This is another commonly used evaluation indicator for regression problems, which calculates the square root of the average squared difference between the predicted value and the true value. Its formula is:
[0093]
[0094] Mean Squared Error (MSE): This is a commonly used evaluation indicator for regression problems, which calculates the average squared difference between the predicted value and the true value. Its formula is:
[0095]
[0096] Coefficient of determination (R-squared, R 2 ): This is a commonly used evaluation indicator for regression problems, which represents the proportion of the data variability explained by the model. Its formula is:
[0097]
[0098] where, SS res is the sum of squared residuals, and SS tot is the total sum of squares.
[0099] (1) Experimental results and analysis
[0100] Table 1 Results of each sub-model
[0101]
[0102] As shown in Table 1, the experimental results of the present invention are presented. The present invention selects four evaluation metrics to evaluate the performance of the model, namely Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Mean Square Error (MSE), and Coefficient of Determination (R 2 ). These metrics can reflect the prediction accuracy, stability, and interpretability of the model.
[0103] The method of the present invention is compared with a single Stacking ensemble learning model. The single Stacking ensemble learning model makes predictions based on the data generated by Principal Component Analysis (PCA) and Uniform Manifold Approximation and Projection (UMAP) respectively. The results show that the method of the present invention has significant improvements in all evaluation metrics, indicating that the method of the present invention can effectively improve the accuracy of the prediction task.
[0104] The present invention proposes a method for predicting regression values of relational data based on composite ensemble learning. By carefully designing three different Stacking ensemble learning models, the accuracy and robustness of regression prediction for relational data are effectively improved. First, Principal Component Analysis (PCA) and Uniform Manifold Approximation and Projection (UMAP) are used for data dimensionality reduction, providing an optimized input feature set for subsequent processing by the ensemble learning model. Then, three independently designed Stacking ensemble learning models are used to process the datasets reduced by PCA and UMAP respectively.
[0105] The first Stacking ensemble learning model integrates a neural network and Gradient Boosting Decision Tree (GBDT), and uses their different processing capabilities to predict the regression value for the data processed by PCA. This model particularly emphasizes improving the prediction accuracy while maintaining the essential features of the data. The second model is specifically designed for the data processed by UMAP, and also adopts a combination of two Gradient Boosting Decision Trees and a neural network to optimize the processing of these non-linearly dimensionally reduced data. This customized strategy ensures that the model can make the most of the characteristics of the data after UMAP dimensionality reduction, thereby improving the prediction performance.
[0106] Finally, the third model is a more advanced Stacking ensemble learning framework that integrates the outputs of the first two models. By further analyzing and optimizing the combination of these results, it forms the final comprehensive regression value prediction. This level of integration not only enhances the model's adaptability to data diversity but also greatly improves the prediction accuracy and robustness by combining the advantages of multiple models.
[0107] The method of the present invention demonstrates excellent capabilities in dealing with complex and multi-dimensional relational data through the design of these multiple levels of models, showing strong application potential in multiple data-intensive fields such as healthcare and financial market analysis. The successful application of this composite integrated learning method fully demonstrates the important role of advanced data analysis technologies in modern society, especially their application value in precision medicine and personalized decision support systems.
[0108] The above-described embodiments only express the implementation modes of the present invention, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A relational data regression value prediction method based on composite ensemble learning, characterized in that The relational data regression value prediction method The following steps are involved: The first step is to use principal component analysis (PCA) to reduce the dimension of the relational data with regression value labels and generate a reduced-dimensional data set. The stacked ensemble learning model is used to process the reduced-dimensional data set, perform regression value prediction, and obtain the regression value result of the first prediction; In the second step, the same relational data with regression value labels as in the first step is used to perform nonlinear dimensionality reduction using the uniform manifold approximation projection (UMAP) to generate another reduced-dimensional data set. The goal is to maintain the local and global structure of the high-dimensional data as much as possible in the low-dimensional space. Another stacked ensemble learning model is used to process the reduced-dimensional data set and perform regression value prediction to obtain the second predicted regression value result; In the third step, the third stacking ensemble learning model Stacking is used to integrate the regression value results output by the two stacking ensemble learning models in the first and second steps, analyze and fuse the two regression value prediction results, and obtain the final output of the composite ensemble learning framework.
2. The relational data regression value prediction method based on composite integrated learning according to claim 1, characterized in that: The first step is as follows: Step 1.1) Use principal component analysis PCA to reduce the dimension and generate a reduced-dimensional data set. The steps of principal component analysis are as follows: First, obtain the original data from a database or relevant data files. The original data is relational data, and let the original data matrix be X ∈ R n×d , where n is the number of samples and d is the feature dimension; Center the original data, that is where is the mean vector of the data; Then, calculate the covariance matrix of the data: Finally, perform eigenvalue decomposition on the covariance matrix to obtain \(C = W\Lambda W^T\), where \(W\in\mathbb{R}^{d\times d}\) is the matrix of eigenvectors, \(\Lambda\in\mathbb{R}^{d\times d}\) is a diagonal matrix whose diagonal elements are the eigenvalues; then select the eigenvectors corresponding to the largest \(k\) eigenvalues to form a matrix \(W_k\in\mathbb{R}^{d\times k}\), where \(k < d\); project the data onto \(W_k\) to obtain the dimensionality-reduced data \(Y = XW_k\in\mathbb{R}^{n\times k}\); T , where \(W\in\mathbb{R}^{d\times d}\) d×d is the matrix of eigenvectors, \(\Lambda\in\mathbb{R}^{d\times d}\) d×d is a diagonal matrix whose diagonal elements are the eigenvalues; then select the eigenvectors corresponding to the largest \(k\) eigenvalues to form a matrix \(W_k\in\mathbb{R}^{d\times k}\), where \(k < d\); project the data onto \(W_k\) to obtain the dimensionality-reduced data \(Y = XW_k\in\mathbb{R}^{n\times k}\); k \(\in\mathbb{R}^{d\times k}\), where \(k < d\); project the data onto \(W_k\) d×k to obtain the dimensionality-reduced data \(Y = XW_k\in\mathbb{R}^{n\times k}\); k to get the dimensionality-reduced data \(Y = XW_k\in\mathbb{R}^{n\times k}\); k \(\in\mathbb{R}^{n\times k}\); n×k ; It should be noted that in the original text, there is a missing transpose symbol \(W^T\) in the formula \(C = W\Lambda W\). The above translation has been corrected according to the correct mathematical formula. Through the above steps, the data is reduced from d dimensions to k dimensions, where d represents the total number of features or variables in the original data set, that is, the dimension of the data; and k represents the number of principal components selected and retained after principal component analysis; Step 1.2) The first stacking ensemble learning model Stacking is used to process the dimension reduction data set obtained by PCA in the first step; the stacking ensemble learning model combines two different neural network models NeuralNet and a gradient boosting decision tree GBDT, and the three models are used as sub-models to provide a comprehensive prediction result; Step 1.3) The two sub-models, NeuralNet and GBDT, independently learn and predict the data after PCA dimension reduction, and obtain the regression value results predicted for each data; then, the output weight of the sub-model is assigned by the gradient descent method, and the error of the function is minimized by iteratively adjusting the parameters to find the minimum value; at the same time, in the training process of each sub-model, the k-fold cross-validation method is used to enhance the generalization ability of the model; After the above training is completed, the first stacked ensemble learning model outputs the first comprehensive regression value prediction result.
3. A relational data regression value prediction method based on composite integrated learning according to claim 1, characterized in that: The second step is as follows: Step 2.1) Using the uniform manifold approximate projection (UMAP) method, nonlinear dimensionality reduction is performed on the original input relational data; First, construct a local neighborhood graph for each point in high-dimensional space; Afterwards, a data representation is sought in the low-dimensional space to minimize the topological structure difference between the high-dimensional space and the low-dimensional space; the following objective function is optimized using the gradient descent method: where d high and d low represent the distances in the high-dimensional space and the low-dimensional space respectively. Here, 'distance' refers to a measure of the difference or similarity between two data points, usually calculated by a specific distance formula and used to evaluate the proximity or similarity between data points; Step 2.2) After using the Uniform Manifold Approximation and Projection (UMAP) method to perform non-linear dimensionality reduction on the original input relational data to obtain the dimensionality-reduced data, another independent Stacking ensemble learning model is used to process the dimensionality-reduced data set; the Stacking ensemble learning model is optimized for the characteristics of UMAP data and integrates base learners, including two different Gradient Boosting Decision Trees (GBDTs) and a Neural Network (NeuralNet); The integration weights of the regression value prediction results of the three sub-models in the second step are also assigned by the gradient descent method; Through the optimization process of the second Stacking ensemble learning model, the dimensionality-reduced data set obtained by the UMAP method can be effectively utilized to generate an accurate regression value prediction result specifically for the characteristics of this dimensionality-reduced data; the obtained regression value prediction result will be passed as an independent output to the next Stacking ensemble learning model. This output will be combined with the outputs of other models for further synthesis and optimization to form the final comprehensive prediction result.
4. A relational data regression value prediction method based on composite integrated learning according to claim 1, characterized in that: The specific content of the third step is as follows: First, the output results of the first and second steps are used as inputs and integrated through a third Stacking ensemble learning model; the Stacking ensemble learning model includes three sub-models: Random Forest Regression, Neural Network, and Extreme Gradient Boosting; The three sub-models are carried out in parallel, independently processing the input data and outputting prediction results; Subsequently, these regression value prediction results are optimized and weighted combined by the gradient descent method to further improve the overall prediction accuracy and robustness; Finally, the Stacking ensemble learning model in the third step outputs a comprehensive and optimized regression value prediction result, which integrates the prediction capabilities and characteristics of all models in the previous two steps, providing solid data support for the final decision-making.
Citation Information
Cited By
Element detection method and device based on model fusion migration, medium and equipment
CN120609806A
Element detection method and device based on model fusion migration, medium and equipment
CN120609806B
Rapid multi-omics prediction method and system based on improved tensor ridge regression
CN121096447A