Disease early screening method based on Stacking integrated model
Through the Stacking integrated model combined with the prediction results of multiple basic learners, the problem of insufficient diagnostic efficiency and accuracy in early disease screening is solved, and more efficient disease screening and early recognition is achieved.
Patent Information
- Application Number
- CN202311461501.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively combine the prediction results of multiple machine learning models in early disease screening, resulting in insufficient diagnostic efficiency and accuracy.
The disease premature screening method based on Stacking integrated model is adopted, and data preprocessing (including outlier value removal and missing value filling), feature dimensionality reduction, and a variety of basic learners (such as SVM, KNN, AdaBoost, GBDT and MLP) are constructed and parameter adjustment is performed. Finally, logistic regression is used as a meta-learner to fusion of prediction results.
It improves the classification prediction performance of early disease screening and enhances the ability to identify early diseases, especially in diabetes screening.
Smart Images

Figure CN119943423A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence medical disease diagnosis, and in particular to an early disease screening method based on a Stacking integrated model. Background Art
[0002] Diseases, such as diabetes, are harmful to human health, especially type 2 diabetes. It is necessary to find a reliable screening mechanism and establish an effective diagnostic model. The earlier the disease is discovered and the earlier the treatment is received, the less harmful it is to human health. Therefore, when conditions permit, collective physical examinations should be conducted to screen out high-risk groups who may be ill, and early diagnosis and timely treatment of diabetes can greatly reduce the burden of diabetes. The development of artificial intelligence in medical diagnosis is in full swing. The use of machine learning as a means of medical auxiliary diagnosis is already very common. Using machine learning algorithms to build disease diagnosis models can not only make up for the subjective defects of traditional medical diagnosis, but also improve diagnostic efficiency.
[0003] Support Vector Machines (SVM) is a supervised machine learning method based on statistical learning and is widely used in binary classification problems. For small sample, high dimensional and nonlinear classification problems, SVM can largely solve the problems of "dimensionality curse" and overfitting. The core mechanism of SVM is: for the input data samples, find a hyperplane to separate these samples in the best way. The optimal hyperplane should satisfy the maximum interval, that is, the distance between the data point closest to the hyperplane and the hyperplane in the two types of data points is the largest.
[0004] KNN (K-nearest Neighbor) is an efficient, simple, non-parametric classification method. It is widely used due to its convenient implementation. The basic idea of KNN is to determine the k samples closest to the predicted sample by calculating the distance between the predicted sample and other samples, and then predict the label of the sample based on the attributes of these k "neighbors".
[0005] Multilayer Perception (MLP) is a feedforward neural network structure consisting of an input layer, an output layer and multiple hidden layers. It has excellent nonlinear mapping capabilities and is very suitable for multi-category nonlinear classification problems. It has achieved good results in image classification, image processing, pattern recognition and other fields.
[0006] Ensemble learning is to build multiple base learners and combine their prediction results as the final prediction result. Ensemble learning combines the prediction results of multiple and various learners. Compared with a single learner, the prediction results of ensemble learning are more accurate. Generally speaking, ensemble learning can be divided into bagging, boosting and stacking.
[0007] Boosting is an important branch of ensemble learning. It refers to an ensemble learning method that trains a series of weak learners to become strong learners. Its idea comes from the concepts of weak learnability and strong learnability proposed by Valiant and Kearns when studying the probable approximately correct (PAC) learning model. They raised the question of whether weak learnability and strong learnability are equivalent in the PAC learning framework. If the two are equivalent, then as long as a weak learning algorithm that is slightly better than random guessing is found, it can be promoted to a strong learning algorithm, without having to find a strong learning algorithm that is difficult to obtain. At present, the two most popular boosting algorithms are adaptive boosting (Adaptive Boosting, AdaBoost) and gradient boosting (Gradient Boosting, GB).
[0008] Different from Bagging and Boosting, the basic idea of stacking is to integrate multiple heterogeneous learners. Stacking is a hierarchical fusion algorithm proposed by Wolpert in 1992. The goal of stacking generalization is to minimize the generalization error rate of one or more learners. The core idea of stacking is not to generate base learners, but to focus on how to combine the results of multiple base learners, use the predicted values of the base learners as the input of the next layer model, and train the combined algorithm for the final prediction. It has been successfully applied to complex classification problems. Compared with single models and other integrated models, stacking usually produces better results. Summary of the invention
[0009] In view of the problem of disease prediction, the purpose of the present invention is to provide a disease early screening method based on the Stacking integrated model. The Stacking integrated model has better classification prediction performance and can provide auxiliary decision support for early disease screening.
[0010] In order to achieve the above-mentioned purpose and other advantages according to the present invention, a method for early disease screening based on a Stacking integrated model is provided, characterized in that it includes the following steps:
[0011] S1. Eliminate data outliers based on box plot theory and relevant medical knowledge;
[0012] S2. Data missing values are processed according to the stratified mean imputation method. First, the data samples are classified according to their disease conditions, and then the mean of each class is used to fill in the missing values;
[0013] S3. Calculate the Pearson correlation coefficient between features In the formula, cov(X i ,X j ) is the covariance of the two features, are the characteristic standard deviations respectively. The correlation analysis is used to determine whether to perform dimensionality reduction, and finally the cleaned data set is obtained;
[0014] S4. Five models including SVM, KNN, AdaBoost, GBDT and MLP were constructed and applied to the preliminary prediction of diseases after adjusting the parameters;
[0015] S5, after applying S4 to the preliminary prediction of the disease, the best three models are confirmed as the base learners of the Stacking ensemble model based on the AUC index values calculated;
[0016] S6. Divide the data set into a training set and a test set in a ratio of 8:2. Then, for the training set, use K-fold cross validation to train three base learners. The data set is divided into K parts, and one of them is taken as a sub-test set each time, and the remaining K-1 parts are taken as sub-training sets. Iterate K times to obtain the probability prediction results of the base learners to form a new feature training set. For the test set, use the iterative K models for probability prediction, and after averaging, obtain a new feature test set.
[0017] S7. Select logistic regression as the meta-learner, use the probability prediction results of the base learner to form a new feature space to train the meta-learner to obtain the final prediction results, and use accuracy, sensitivity, specificity, F1 score and AUC value as evaluation indicators of the prediction method to evaluate the performance of the model. After that, this method can be used to provide auxiliary decision support for early disease screening.
[0018] Preferably, in step S1, the upper limit Q of the box plot theoretical calculation is max =Q3+k′(Q3-Q1) and lower limit Q min =Q1-k′(Q3-Q1), where Q1 is the 25% quantile and Q3 is the 75% quantile. Taking into account the high variability of medical characteristics, the k′ value in the upper and lower limit formulas is taken as 3. If the variable exceeds the upper limit or is less than the lower limit, it is regarded as an outlier, and extreme outliers exceeding the upper and lower limits are excluded.
[0019] Preferably, in step S3, determining whether to perform dimensionality reduction processing according to correlation analysis includes:
[0020] If the Pearson coefficient of two features is greater than 0.7, the variables are strongly correlated, and the dimensionality reduction process removes one of the features with a larger Pearson coefficient than the other features;
[0021] If the coefficient is less than or equal to 0.7, the variables are not strongly correlated and no dimensionality reduction is performed.
[0022] Preferably, the SVM algorithm model parameter adjustment in step S4 includes: first using normalization to process the data of the SVM algorithm model and then applying a grid search algorithm to adjust the parameters to find the optimal parameters of the penalty coefficient and the kernel function parameters.
[0023] Preferably, the KNN algorithm model parameter adjustment in step S4 includes: applying a grid search algorithm to adjust the KNN algorithm model parameters to find the optimal parameters of the number of neighbor points and classification rules.
[0024] Preferably, the parameter adjustment of the AdaBoost algorithm model in step S4 includes: using random search to narrow the search range of parameters for the AdaBoost algorithm model, first finding a set of approximate optimal parameters, and then using grid search around this set of approximate optimal parameters to further accurately determine the optimal parameters, selecting the CART decision tree as the basic evaluator of the AdaBoost algorithm model, finding the maximum number of basic evaluators, the learning rate, the maximum depth considered when constructing the decision tree, the maximum number of features, and the optimal parameter combination of the minimum number of samples for node division, and finally determining the optimal parameter combination of the AdaBoost algorithm model.
[0025] Preferably, the parameter adjustment of the GBDT algorithm model in step S4 includes: adjusting the parameters of the GBDT algorithm model using random search and grid search, finding the optimal parameter combination of the maximum number of basic evaluators, learning rate, subsampling ratio, maximum depth considered when building a decision tree, maximum number of features, and minimum number of samples for node division, and finally determining the optimal parameter combination of the GBDT algorithm model.
[0026] Preferably, the MLP algorithm model parameter adjustment in step S4 includes: adjusting the MLP algorithm model by normalization, random search and grid search, finding the optimal parameter combination of the number of hidden layer neurons and regularization parameters after normalizing the data, and finally determining the optimal parameter combination of the MLP algorithm model.
[0027] Preferably, the disease comprises diabetes.
[0028] Preferably, a method for early screening of diabetes based on the Stacking integrated model comprises the following steps:
[0029] R1. Eliminate data outliers based on box plot theory and relevant medical knowledge, where the box plot theory calculates the upper limit Q max =Q3+k′(Q3-Q1) and lower limit Q min =Q1-k′(Q3-Q1), where Q1 is the 25% quantile and Q3 is the 75% quantile. Considering the high variability of medical characteristics, the k′ value in the upper and lower limit formulas is taken as 3. If the variable exceeds the upper limit or is less than the lower limit, it is regarded as an outlier and is removed;
[0030] R2, using the stratified mean imputation method to handle missing data values, first classify the data samples according to their disease conditions, and then use the mean of each class to fill in the missing data values;
[0031] R3. Calculate the Pearson correlation coefficient between features In the formula, cov(X i ,X j ) is the covariance of the two features, are the feature standard deviations, respectively. According to the correlation analysis, (1) if the Pearson coefficient of two features is greater than 0.7, the variables are strongly correlated, and dimensionality reduction is performed to remove one of the features with a larger Pearson coefficient with the other features; (2) if the Pearson coefficient of two features is less than or equal to 0.7, the variables are not strongly correlated, and dimensionality reduction is not performed;
[0032] R4. Five models, including SVM, KNN, AdaBoost, GBDT and MLP, were constructed. The SVM algorithm model was first processed by normalization and then the grid search algorithm was applied to adjust the parameters to find the optimal parameters of the penalty coefficient and kernel function parameters. The KNN algorithm model was adjusted by the grid search algorithm to find the optimal parameters for the number of neighbor points and classification rules. The AdaBoost algorithm model was used to use random search to narrow the search range of parameters, first find a set of approximate optimal parameters, and then use grid search around this set of approximate optimal parameters to further accurately determine the optimal parameters. The CART decision tree was selected as the basic evaluator of the AdaBoost algorithm model to find the optimal parameter combination of the maximum number of basic evaluators, learning rate, maximum depth considered when building the decision tree, maximum number of features, and minimum number of samples for node division, and finally determine the optimal parameter combination of the AdaBoost algorithm model. ; Random search is used to narrow the search range of parameters for the GBDT algorithm model, and a set of approximate optimal parameters is first found. Then, grid search is used around this set of approximate optimal parameters to further accurately determine the optimal parameters, and the optimal parameter combination of the maximum number of basic evaluators, learning rate, subsampling ratio, maximum depth considered when building a decision tree, maximum number of features, and minimum number of samples for node division is found, and finally the optimal parameter combination of the GBDT algorithm model is determined; Normalization, random search, and grid search are used to adjust the parameters of the MLP algorithm model, and random search is used to narrow the search range of parameters after normalizing the data, and a set of approximate optimal parameters is first found. Then, grid search is used around this set of approximate optimal parameters to further accurately determine the optimal parameters, and the optimal parameter combination of the number of hidden layer neurons and regularization parameters is found, and finally the optimal parameter combination of the MLP algorithm model is determined. After adjusting the parameters of the five models, they are applied to the prediction of diabetes;
[0033] R5, after applying S4 to the prediction of diabetes, the best three models are confirmed as the base learners of the Stacking ensemble model based on the AUC index values calculated;
[0034] R6. Divide the data set into a training set and a test set in a ratio of 8:2. Then, for the training set, use K-fold cross validation to train three base learners. The data set is divided into K parts, and one of them is taken as a sub-test set each time, and the remaining K-1 parts are taken as sub-training sets. Iterate K times to obtain the predicted probability results of the base learners to form a new feature training set. For the test set, use the iterative K models for probability prediction. After averaging, a new feature test set is obtained.
[0035] R7. Logistic regression is selected as the meta-learner. The predicted probability results of the base learner are used to form a new feature space to train the meta-learner to obtain the final prediction results. Accuracy, sensitivity, specificity, F1 score and AUC value are used as evaluation indicators of the prediction method to evaluate the performance of the model and verify the effectiveness of the method proposed in the present invention. After that, this method can be used to provide auxiliary decision support for early diabetes screening.
[0036] Compared with the prior art, the present invention has the following beneficial effects: using box plot theory to process outliers, using stratified mean imputation method to process missing values, using correlation analysis to reduce dimensionality to obtain a preprocessed data set, using AUC value as an evaluation index to select a base learner, and finally using the Stacking strategy to integrate the selected base learners, and using K-fold cross-validation to obtain new feature training sets and test sets, and selecting logistic regression as a meta-learner to predict diseases. By using Stacking integration, different models can learn different features of the data, and the fused features often have better performance. According to the present invention, this method is applied to the Pima Indian diabetes data set for experimental analysis, and the relevant evaluation indicators of the Stacking integration model and the intermediate machine learning model are compared. The obtained Stacking integration model has better classification prediction performance and can provide auxiliary decision support for early diseases, especially diabetes screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram showing the principle of prediction of the disease early screening method based on the Stacking integrated model according to the present invention;
[0038] Figure 2 A box plot of the skinfold thickness, insulin and body mass index characteristics of the Pima Indian diabetes dataset according to the disease early screening method based on the Stacking integrated model of the present invention;
[0039] Figure 3 It is a characteristic heat map of the correlation analysis of the Pima Indian diabetes data set by the disease early screening method based on the Stacking integrated model according to the present invention;
[0040] Figure 4 This is a graph showing the evaluation index results of the Pima Indian diabetes dataset using the disease early screening method based on the Stacking integrated model according to the present invention. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0042] Reference Figure 1-4 , early disease screening method based on Stacking ensemble model for diabetes prediction.
[0043] This embodiment performs predictive analysis on the Pima Indian diabetes dataset. The dataset used in this embodiment is the Pima Indian Diabetes dataset from the National Center for Diabetes, Digestive and Kidney Diseases Research in the University of California, Irvine (UCI) machine learning public database. This dataset is for type II diabetes, and the sample subjects are all women with Pima Indian ancestry who are at least 21 years old. The dataset has a total of 768 data, of which 268 samples have diabetes. It contains 8 medical features and 1 label. The 8 medical features are pregnancies, 2-hour oral glucose tolerance test (OGTT) plasma glucose concentration (Glucose), diastolic blood pressure (BloodPressure), triceps skinfold thickness (SkinThickness), insulin concentration (Insulin), body mass index (BMI), diabetes pedigree function (DiabetesPedigreeFunction) and age (Age). The label value (Outcome) is 0 or 1, 0 means no diabetes, and 1 means diabetes. Common statistical indicators of these features are shown in Table 1.
[0044] Table 1 Feature analysis table of diabetes data set
[0045]
[0046] According to the BMI standard published by the World Health Organization (WHO), it is generally believed that the normal fluctuation range of BMI is 18.5-24.9 kg / m 2 , when BMI ≥ 25 kg / m 2 When BMI ≥ 30 kg / m 2When the patient was diagnosed with obesity, the maximum value of the body mass index (BMI) of 67.1 in Table 1 is obviously far beyond the normal range and may be an abnormal data. In addition, for the four features of number of pregnancies, glucose concentration, diastolic blood pressure, and age, there is a certain gap between their maximum values and the mean and median. Considering the high variability of medical feature values, although the maximum value and the mean have a large gap, in the real world, this gap is allowed, and it is not possible to judge whether it is abnormal based solely on the information provided in Table 1. Therefore, we only cleaned the abnormal values of skinfold thickness, insulin concentration, and body mass index.
[0047] according to Figure 2 The box plot shows that samples with insulin concentrations of 543, 846, 744, 680, 545, 540, and 510 are abnormal samples. These 7 samples and the sample with a BMI of 67.1 are deleted as abnormal samples. After outlier processing, 8 abnormal samples are deleted from the original data set, and the data set size becomes 760, of which 263 are diseased samples and 497 are non-diseased samples.
[0048] From Table 1, we can see that the minimum value of the number of pregnancies, glucose concentration, diastolic blood pressure, skinfold thickness, insulin concentration and body mass index is 0. However, in actual medical situations, glucose concentration, diastolic blood pressure, skinfold thickness, insulin concentration and body mass index will not be 0. The main cause of patients with type 2 diabetes is insulin resistance. Although patients may also have insufficient insulin secretion, it will not be 0. The number of pregnancies can be 0 in reality. For data sets in the medical field, the missing values of data are processed according to the stratified mean imputation method. First, the data samples are classified according to the disease status, and then the mean is filled in each class. The mean values of diseased samples and non-disease samples under each feature are shown in Table 2.
[0049] Table 2. Mean values of characteristics of diabetic patients and normal people
[0050]
[0051] Calculate the Pearson correlation coefficient between each feature, analyze the correlation between each feature and the label, and obtain the following Figure 3 The heat map shown. Figure 3It can be seen that the Pearson correlation coefficient between age and number of pregnancies is 0.55, the Pearson correlation coefficient between BMI and skinfold thickness is 0.57, and the Pearson correlation coefficient between label and glucose concentration is 0.5. Except for these three sets of correlation coefficients, the correlations between other features are all less than 0.5. For age and number of pregnancies, it is in line with the actual situation that the number of pregnancies increases with age; for BMI and skinfold thickness, the higher the BMI, the higher the possibility that the sample suffers from obesity, and the greater the skinfold thickness; for label and glucose concentration, the higher the glucose concentration in the blood, the more likely it is that the sample lacks insulin or insulin may not play a role in lowering blood sugar, and the more likely it is to suffer from diabetes. However, these three sets of correlations are only greater than 0.5. We generally believe that when the correlation coefficient of two variables is greater than 0.7, the two variables have a strong correlation. Therefore, this embodiment continues to use the original 8 features without dimensionality reduction.
[0052] Five models, including SVM, KNN, AdaBoost, GBDT and MLP, were constructed. The SVM algorithm model was first processed by normalization and then the grid search algorithm was applied to adjust the parameters. The AUC value was used as the evaluation index to obtain the optimal parameter combination of penalty coefficient C=1 and kernel function parameter gamma=1. A KNN model was constructed and the grid search algorithm was applied to adjust the parameters of the KNN algorithm model. The AUC value was used as the evaluation index to obtain the number of neighbor points k=27 and the "distance" classification rule as its optimal parameter combination. An AdaBoost algorithm model was constructed and random search was used to narrow the search range of parameters for the AdaBoost algorithm model. A set of approximate optimal parameters was first found, and then a grid search was used around this set of approximate optimal parameters to finally determine the optimal parameter combination of the AdaBoost algorithm model as shown in Table 3.
[0053] Table 3 Optimal parameter combination of AdaBoost algorithm model
[0054]
[0055] The GBDT algorithm model was constructed, and random search and grid search were used to adjust the parameters of the GBDT algorithm. Random search was used to narrow the search range of parameters, and then grid search was used to further accurately determine the optimal parameters. Finally, the optimal parameter combination of the GBDT algorithm model was determined as shown in Table 4.
[0056] Table 4 Optimal parameter combination of GBDT
[0057]
[0058] An MLP model was constructed, and normalization, random search, and grid search were used to adjust the parameters of the MLP. Normalization was first used for data processing, and then random search was used to narrow the search range of parameters. Grid search was then used to further accurately determine the optimal parameters. Finally, the optimal parameter combination of the MLP model was determined, namely, the number of hidden layer neurons hidden_layer_sizes = 200 and the regularization parameter alpha = 0.001.
[0059] The AUC values of each base learner obtained through experiments are shown in Table 5. The AUC values of each base learner are sorted from large to small as GBDT, AdaBoost, MLP, KNN, and SVM. The AUC values of MLP, SVM, and KNN are not much different, so just one of them can be selected. In this embodiment, AdaBoost, GBDT, and MLP are selected as the three base learners of the Stacking integration framework according to the AUC values. Among them, AdaBoost and GBDT have no requirements for data normalization, and MLP requires data set normalization.
[0060] Table 5 Main parameters of base learners and their AUC values
[0061]
[0062] The Stacking ensemble model in this example takes K=5, and divides the training set and the test set in a ratio of 8:2. On the training set, 5-fold cross validation is used to train the GBDT, AdaBoost and MLP base learners. The training set is further divided into 5 parts, and 1 part is taken as the sub-test set each time, and the remaining 4 parts are taken as sub-training sets. It is iterated 5 times. The probability prediction results of each base learner on the 5 groups of sub-test sets constitute the predicted probability results of the base learner. The new feature training set composed of the predicted probability results of the three base learners will be used as the training data set of the second-layer meta-learner to train the meta-learner logistic regression.
[0063] In order to verify the performance of the Stacking ensemble model, various evaluation indicators of the Stacking ensemble model are calculated and compared with the five intermediate machine learning models, and Table 6 is obtained.
[0064] Table 6 Comparison of classification prediction performance between Stacking ensemble model and intermediate machine learning model
[0065]
[0066] As shown in Table 6, the indicators of the Stacking ensemble model are improved compared with SVM, KNN and MLP. The accuracy of the Stacking ensemble model is 0.9211, and the AUC value is 0.9413. The accuracy and AUC value of the GBDT model are increased by 0.7% and 0.1%, respectively, and the AUC value of the Adaboost model is increased by about 2.1%, respectively, and the accuracy and AUC value of the MLP model are increased by about 6.9% and 4.5%, respectively, indicating that the Stacking ensemble model has better classification prediction performance. For sensitivity and specificity, the sensitivity of the Stacking ensemble model is 0.8043, which is about 5.7% higher than that of the GBDT model and about 2.8% higher than that of the MLP model. The increase in sensitivity indicates that the missed diagnosis rate of the Stacking ensemble model is reduced, and more diabetic patients can be identified in the tested population. The specificity of the Stacking ensemble model is reduced by about 1.0% compared with GBDT. The decrease in specificity indicates that the probability of the Stacking ensemble model misjudging people who do not have diabetes as having diabetes has increased.
[0067] Sensitivity and specificity are a pair of contradictory evaluation indicators. Ideally, we hope that the sensitivity and specificity of the classifier are both high, but in actual situations, we often find a balance between sensitivity and specificity. The main research purpose of this embodiment is to establish an effective diabetes prediction model that can help people predict the risk of diabetes in early physical examination screening. Therefore, the sensitivity of the classifier is required to be high enough to identify as many people as possible who may have diabetes. The earlier the discovery, the earlier the treatment, and the higher the possibility of disease cure. Although the Stacking integrated model proposed in this embodiment has a lower specificity than the GBDT model, in the early diabetes risk screening, we hope that the classifier will not miss the diagnosis, identify all patients who may have diabetes risks as much as possible, and remind them to adjust their lifestyles and go to the hospital for further examination in time. The reduction in specificity indicates that the classifier increases the possibility of misdiagnosing normal people as diabetic patients. However, it should be noted that the patient's true clinical diagnosis cannot be judged solely by the prediction results of the classifier, and a professional doctor must give the final clinical diagnosis based on the patient's medical examination results, which can largely avoid misdiagnosis. Therefore, a proper reduction in specificity is acceptable.
[0068] The Stacking integrated model established by the present invention can be applied to the early risk screening of diabetes, and can efficiently screen out high-risk groups who may suffer from diabetes from the population. This not only improves the medical treatment effect, but also can promptly inform relevant people, reminding them to pay attention to their physical health and suggesting them to go to the hospital for further examination.
[0069] In summary, the Stacking integrated model proposed in the present invention has better classification prediction performance than a single intermediate machine learning model, which shows that the Stacking-based disease prediction method proposed in the present invention is effective and can be applied to early risk screening of diabetes.
[0070] Through example analysis, the present invention is applicable and effective in diabetes prediction analysis.
[0071] The number of devices and processing scales described here are used to simplify the description of the present invention, and the application, modification and variation of the present invention will be obvious to those skilled in the art.
[0072] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the implementation modes, and they can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and the illustrations shown and described herein.
Claims
1. A method for early disease screening based on a Stacking integrated model, characterized in that: The following steps are involved: S1. Eliminate data outliers based on box plot theory and relevant medical knowledge; S2. Data missing values are processed according to the stratified mean imputation method. First, the data samples are classified according to their disease conditions, and then the mean of each class is used to fill in the missing values; S3. Calculate the Pearson correlation coefficient between features Where covX( i X, j is the covariance of the two features, are the characteristic standard deviations respectively. The correlation analysis is used to determine whether to perform dimensionality reduction, and finally the cleaned data set is obtained; S4. Five models including SVM, KNN, AdaBoost, GBDT and MLP were constructed and applied to the preliminary prediction of diseases after adjusting the parameters; S5, after applying S4 to the preliminary prediction of the disease, the best three models are confirmed as the base learners of the Stacking ensemble model based on the AUC index values calculated; S6. Divide the data set into a training set and a test set in a ratio of 8:
2. Then, for the training set, use K-fold cross validation to train three base learners. The data set is divided into K parts, and one of them is taken as a sub-test set each time, and the remaining K-1 parts are taken as sub-training sets. Iterate K times to obtain the probability prediction results of the base learners to form a new feature training set. For the test set, use the iterative K models for probability prediction, and after averaging, obtain a new feature test set. S7. Select logistic regression as the meta-learner, use the probability prediction results of the base learner to form a new feature training set to train the meta-learner to obtain the final prediction results, and use accuracy, sensitivity, specificity, F1 score and AUC value as evaluation indicators of the prediction method to evaluate the performance of the model. After that, this method can be used to provide auxiliary decision support for early disease screening.
2. A disease early screening method based on a Stacking integrated model as claimed in claim 1, characterized in that: The upper limit Q of the box plot theoretical calculation in step S1 max =Q3+k′(Q3-Q1) and lower limit Q min =Q1-k′(Q3-Q1), where Q1 is the 25% quantile and Q3 is the 75% quantile. Taking into account the high variability of medical characteristics, the k′ value in the upper and lower limit formulas is taken as 3. If the variable exceeds the upper limit or is less than the lower limit, it is regarded as an outlier, and extreme outliers exceeding the upper and lower limits are excluded.
3. A disease early screening method based on a Stacking integrated model as claimed in claim 1, characterized in that: In step S3, determining whether to perform dimensionality reduction processing according to correlation analysis includes: If the Pearson coefficient of two features is greater than 0.7, the variables are strongly correlated, and the dimensionality reduction process removes one of the features with a larger Pearson coefficient than the other features; If the coefficient is less than or equal to 0.7, the variables are not strongly correlated and no dimensionality reduction is performed.
4. A disease early screening method based on a Stacking integrated model as claimed in claim 1, characterized in that: The SVM algorithm model parameter adjustment in step S4 includes: first using normalization to process the data of the SVM algorithm model and then applying a grid search algorithm to adjust the parameters to find the optimal parameters of the penalty coefficient and the kernel function parameters.
5. The disease early screening method based on the Stacking integrated model according to claim 1, characterized in that: The KNN algorithm model parameter adjustment in step S4 includes: applying a grid search algorithm to adjust the parameters of the KNN algorithm model to find the optimal parameters of the number of neighbor points and the classification rules.
6. The disease early screening method based on the Stacking integrated model according to claim 1, characterized in that: The parameter adjustment of the AdaBoost algorithm model in step S4 includes: using random search to narrow the search range of parameters for the AdaBoost algorithm model, first finding a set of approximate optimal parameters, and then using grid search around this set of approximate optimal parameters to further accurately determine the optimal parameters, selecting the CART decision tree as the basic evaluator of the AdaBoost algorithm model, finding the maximum number of basic evaluators, the learning rate, the maximum depth considered when building the decision tree, the maximum number of features, and the minimum number of samples for node division. The optimal parameter combination, finally determining the optimal parameter combination of the AdaBoost algorithm model.
7. The disease early screening method based on the Stacking integrated model according to claim 1, characterized in that: The parameter adjustment of the GBDT algorithm model in step S4 includes: adjusting the parameters of the GBDT algorithm model using random search and grid search to find the optimal parameter combination of the maximum number of basic evaluators, learning rate, subsampling ratio, maximum depth considered when building a decision tree, maximum number of features, and minimum number of samples for node division, and finally determine the optimal parameter combination of the GBDT algorithm model.
8. The disease early screening method based on the Stacking integrated model according to claim 1, characterized in that: The MLP algorithm model parameter adjustment in step S4 includes: adjusting the MLP algorithm model by normalization, random search and grid search, finding the optimal parameter combination of the number of hidden layer neurons and regularization parameters after normalizing the data, and finally determining the optimal parameter combination of the MLP algorithm model.
9. The method for early disease screening based on the Stacking integrated model according to claim 1, characterized in that: The disease includes diabetes.
Citation Information
Patent Citations
Establishment method of severe spinal cord injury prognosis prediction model
CN112992346A
Method and device for disease prediction, computer equipment and storage medium
CN113096817A
Railway freight customer loss prediction method based on ensemble learning
CN115147155A
Cerebral infarction operation patient survival risk classification method based on machine learning
CN115206527A
Plateau area child electrocardiogram prediction method
CN116602688A
Cited By
Method and system for optimizing diabetic complication prediction model
CN121768690A