Plateau pulmonary edema risk prediction method based on resampling and ensemble learning

By constructing a risk prediction model for plateau pulmonary edema based on resampling and integrated learning, the problem of relying on subjective judgment and data imbalance in traditional methods is solved, and the prediction accuracy and robustness are improved, and it is suitable for health management in plateau environments.

CN120260911APending Publication Date: 2025-07-04XINJIANG MEDICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510328789.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has problems in the risk prediction of plateau pulmonary edema with reliance on doctors’ subjective judgment, lack of personalization and data imbalance, resulting in delayed or inaccurate diagnosis. Especially when medical resources are limited in high altitude areas, traditional methods are not sufficient to provide accurate results.

Method used

A HAPE incidence risk prediction model is constructed based on resampling and integrated learning. By integrating demographic, vital signs and biochemical indicators, the data is balanced using SMOTE, ADASYN and SMOTE+ENN sampling methods, and the integrated model of multi-layer perceptron and gradient-enhanced decision tree is optimized.

Benefits of technology

It improves the accuracy and robustness of the risk prediction of plateau pulmonary edema, reduces morbidity and health risks, provides strong support for clinical intervention and prevention strategies, and is suitable for health management in plateau environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260911A_ABST
    Figure CN120260911A_ABST
Patent Text Reader

Abstract

The invention provides a plateau pulmonary edema risk prediction method based on resampling and ensemble learning, and relates to the technical field of plateau pulmonary edema disease risk prediction, and the method comprises the following steps: S1, information collection; the HAPE onset risk prediction model oriented to the plateau crowd is constructed by combining ensemble learning and a data resampling method, the global expression ability of the model is improved by integrating different types of features including demographic statistics, vital signs and biochemical indexes, and when the problem of data imbalance is solved, various sampling technologies are adopted, so that the prediction accuracy of the HAPE onset risk prediction model is improved. The prediction performance of the model is effectively improved through a grid search technology in feature optimization, in addition, the model integrating a multi-layer perceptron and a gradient boosting decision tree is further optimized through a soft voting mechanism to further optimize a final prediction result, health management in the plateau environment is improved, the incidence rate of HAPE and related health risks are reduced, and the prediction efficiency of the plateau environment is improved. The powerful support is provided for the formulation of clinical intervention and prevention strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of prediction of the risk of high altitude pulmonary edema, and specifically to a method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning. Background Art

[0002] High altitude pulmonary edema (HAPE) is an acute lung disease caused by high altitude environment, whose main features are pulmonary fluid accumulation and dyspnea. The occurrence of HAPE is related to an individual's ability to adapt to high altitude, and is common among travelers and residents in high altitude areas (above 2500m). If not treated in time, HAPE may lead to serious complications and even death. Since the occurrence of HAPE is related to multiple complex factors, including an individual's physiological condition, genetic factors, and the time and speed of exposure to high altitude, its early prediction and intervention have always been the focus of medical research.

[0003] In clinical practice, traditional HAPE risk prediction mainly relies on clinical symptoms, such as cough, dyspnea, chest tightness, hemoptysis, etc.; personal medical history and previous altitude reaction experiences; the adaptation of patients in high altitude environment (such as tolerance to hypoxia, ability to cope with rapid ascent); physiological indicators, such as changes in physical signs like blood oxygen saturation (SpO2), heart rate, and respiratory rate, etc. Relevant research points out that the quick sequential organ failure assessment (qSOFA) score and the national early warning score (NEWS) scoring model can be used as auxiliary tools, and a modified new model that combines clinical symptoms with qSOFA and NEWS scores can help doctors evaluate the severity of the condition and predict poor prognosis.

[0004] However, these traditional methods may have certain limitations in early identifying HAPE high-risk patients. First of all, they rely heavily on doctors' subjective judgments. Doctors make decisions based on reported symptoms and patients' medical histories, which may overlook high-risk groups. Secondly, the early symptoms of HAPE are usually similar to those of acute mountain sickness (AMS), resulting in delayed or inaccurate diagnosis. In addition, these methods lack the ability of personalized prediction and cannot consider individual differences such as genetic or health conditions. Finally, in high altitude areas, the limitations of medical resources and equipment will affect the effectiveness of diagnostic tools, which makes traditional HAPE scoring tools may not be sufficient to provide accurate results under such conditions. Therefore, finding a more accurate and easy-to-implement early HAPE prediction method has become a hot topic in current research.

[0005] In recent years, the rise of electronic medical records has provided new possibilities for the early prediction and risk assessment of HAPE. As an important part of modern medical systems, it records a large amount of patient information, including medical history, diagnosis, treatment, test results, etc., and can provide rich structured and unstructured medical information, which can be used to capture the core characteristics of diseases. These data not only provide a basis for the comprehensive evaluation of patients, but also provide a rich source of features for the construction of machine learning models. For example, by analyzing patients' age, gender, laboratory test results, etc., features valuable for predicting the risk of high-altitude pulmonary edema can be extracted. Reasonable utilization of these features can improve the accuracy and reliability of the prediction model. In recent years, with the rapid development of big data technology and machine learning algorithms, disease prediction models based on machine learning have been widely used in the medical field. These models can automatically learn potential patterns and associations from massive clinical data, thus providing strong support for the early diagnosis and personalized treatment of diseases. Currently, some researchers have mined a wider range of disease risk factors based on electronic medical record data and carried out disease risk prediction and assessment. However, in the risk prediction of clinical HAPE, due to the imbalance in the number ratio of HAPE positive cases (patients with high-altitude pulmonary edema) to negative cases (patients without high-altitude pulmonary edema), this poses a challenge to traditional machine learning models. When facing imbalanced data, traditional machine learning algorithms tend to predict samples in the majority class, resulting in a low prediction accuracy for high-risk HAPE patients. Therefore, how to effectively handle imbalanced data has become a key issue in HAPE risk prediction research. Summary of the Invention

[0006] The object of the present invention is to provide a high-altitude pulmonary edema risk prediction method based on resampling and ensemble learning. By combining ensemble learning and data resampling methods, a risk prediction model for HAPE onset in high-altitude populations is constructed. By integrating different types of features, including demographics, vital signs, and biochemical indicators, the global expression ability of the model is improved. When dealing with the problem of data imbalance, various sampling techniques such as SMOTE sampling method, ADASYN sampling method, and SMOTE+ENN sampling method are adopted, and the prediction performance of the model is effectively improved through grid search technology in feature optimization. In addition, a model integrating a multi-layer perceptron and a gradient boosting decision tree further optimizes the final prediction result through a soft voting mechanism, which helps to improve health management in high-altitude environments, reduce the incidence of HAPE and related health risks, and provides strong support for the formulation of clinical intervention and prevention strategies, having important clinical application prospects.

[0007] The present invention provides the following technical solution: A high-altitude pulmonary edema risk prediction method based on resampling and ensemble learning, comprising the following steps:

[0008] Step 1. Information collection: Determine the HAPE research objects, and collect the original data of the HAPE research objects, where the HAPE research objects include HAPE inpatients and healthy physical examination populations who come to the hospital during the same period;

[0009] Step 2. Data preprocessing: Clean and normalize the original data;

[0010] The following steps are also included:

[0011] Step 3. Construct a dataset: Use the HAPE inpatients as positive samples and the healthy physical examination populations who come to the hospital during the same period as negative samples, and construct a HAPE dataset based on the positive samples and the negative samples;

[0012] Step 4. Resampling processing: Based on the data imbalance in the HAPE dataset, then select a matching resampling method for processing, and adjust the sample distribution between categories multiple times;

[0013] Step 5. Establish an ensemble model: Based on the ensemble learning algorithm, construct multiple HAPE disease onset risk prediction models according to the data in the HAPE dataset, and predict the HAPE disease onset risk based on multiple HAPE disease onset risk prediction models.

[0014] Further, in Step 1, the original data of the HAPE research objects includes patient basic information data, coagulation and biochemical test data, blood routine and five-category data, and novel coronavirus nucleic acid test data.

[0015] Further, in Step 2, the data preprocessing includes deleting outliers in the original data, filling in missing values in the original data, uniformly encoding variables in the original data, and deleting duplicate samples in the original data.

[0016] Further, in Step 3, the HAPE dataset includes a demographic information dataset, a vital sign dataset, and a biochemical test index dataset.

[0017] Further, in Step 3, the steps for constructing the HAPE dataset are as follows:

[0018] S301. Collect the clinical electronic medical record data of the HAPE research objects;

[0019] S302. Delete outliers and duplicate samples in the electronic medical record data;

[0020] S303. Fill in the missing values in the clinical electronic medical record data, and perform unified feature encoding and data annotation on the filled clinical electronic medical record data, and then obtain the HAPE dataset;

[0021] Further, in step four, the resampling method includes downsampling method, SMOTE sampling method, ADASYN sampling method, and SMOTE+ENN sampling method.

[0022] Further, in step four, the downsampling method balances the dataset by reducing the number of samples in the majority class. The SMOTE sampling method balances the data distribution by generating synthetic minority class samples, thereby maximizing the utilization of data. The ADASYN sampling method focuses on the low-density boundary regions in the minority class samples to improve the generalization ability of the model on different distributions. The SMOTE+ENN sampling method can achieve a balance between oversampling and downsampling.

[0023] Further, in step five, the multiple HAPE incidence risk prediction models include support vector machine model, logistic regression model, random forest model, gradient boosting model, multi-layer perceptron model, extreme gradient boosting model, and multi-layer perceptron-gradient boosting ensemble model.

[0024] Further, in step five, a voting classifier is used to construct an ensemble model, integrating the gradient boosting model and the multi-layer perceptron as base learners. The voting classifier uses the soft voting method to generate the final prediction result by weighted averaging the prediction results of each base learner.

[0025] Further, the grid search method is used in step five to ensure the optimal performance of the model under different configurations.

[0026] The present invention provides a method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning, which has the following beneficial effects: By combining ensemble learning with data resampling methods, a prediction model for the incidence risk of HAPE for high altitude populations is constructed. By integrating different types of features, including demographics, vital signs, and biochemical indicators, the global expression ability of the model is improved. When dealing with the problem of data imbalance, various sampling techniques such as SMOTE sampling method, ADASYN sampling method, and SMOTE+ENN sampling method are adopted, and the prediction performance of the model is effectively improved through grid search technology in feature optimization. In addition, the model integrating the multi-layer perceptron and the gradient boosting decision tree further optimizes the final prediction result through a soft voting mechanism, which helps to improve health management in high altitude environments, reduce the incidence rate of HAPE and related health risks, provides strong support for the formulation of clinical intervention and prevention strategies, and has important clinical application prospects. Description of the Drawings

[0027] Figure 1 It is a flowchart of the method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning of the present invention;

[0028] Figure 2This is the overall framework diagram for constructing the model of the high altitude pulmonary edema risk prediction method based on resampling and ensemble learning in the present invention;

[0029] Figure 3 This is the flow chart for constructing the data set of the high altitude pulmonary edema risk prediction method based on resampling and ensemble learning in the present invention;

[0030] Figure 4 This is the data processing flow chart of the high altitude pulmonary edema risk prediction method based on resampling and ensemble learning in the present invention;

[0031] Figure 5 This is the SHAP value distribution diagram of the high altitude pulmonary edema risk prediction method based on resampling and ensemble learning in the present invention;

[0032] Figure 6 This is the arrangement diagram of the average absolute SHAP values of the high altitude pulmonary edema risk prediction method based on resampling and ensemble learning in the present invention. Detailed implementation manners

[0033] Please refer to Figure 1-6 , the present invention provides a technical solution: a high altitude pulmonary edema risk prediction method based on resampling and ensemble learning, including the following steps:

[0034] Step 1, information collection: Determine the HAPE research objects, and collect the original data of the HAPE research objects. The HAPE research objects include HAPE inpatients and healthy physical examination populations who come to the hospital during the same period;

[0035] Step 2, data preprocessing: Clean and normalize the original data;

[0036] It further includes the following steps:

[0037] Step 3, construct a data set: Use the HAPE inpatients as positive samples and the healthy physical examination populations who come to the hospital during the same period as negative samples, and construct a HAPE data set based on the positive samples and the negative samples;

[0038] Step 4, resampling processing: Based on the data imbalance situation in the HAPE data set, then select a matching resampling method for processing, and adjust the sample distribution between categories multiple times;

[0039] Step 5, establish an ensemble model: Based on the ensemble learning algorithm, construct multiple HAPE onset risk prediction models according to the data in the HAPE data set, and predict the HAPE onset risk based on the multiple HAPE onset risk prediction models.

[0040] Specifically, in Step 1, the original data of the HAPE research subjects includes patient basic information data, coagulation and biochemical test data, blood routine and five-classification data, and novel coronavirus nucleic acid test data.

[0041] Specifically, in Step 2, the data preprocessing includes deleting outliers in the original data, filling in missing values in the original data, unifying variable coding of the original data, and deleting duplicate samples in the original data.

[0042] Specifically, in Step 3, the HAPE dataset includes a demographic information dataset, a vital signs dataset, and a biochemical test index dataset.

[0043] Specifically, in Step 3, the construction steps of the HAPE dataset are as follows:

[0044] S301: Collect the clinical electronic medical record data of the HAPE research subjects;

[0045] S302: Delete outliers and duplicate samples in the electronic medical record data;

[0046] S303: Fill in the missing values in the clinical electronic medical record data, and perform unified feature coding and data annotation on the filled clinical electronic medical record data, and then obtain the HAPE dataset;

[0047] Specifically, in Step 4, the resampling methods include the undersampling method, the SMOTE sampling method, the ADASYN sampling method, and the SMOTE+ENN sampling method.

[0048] Specifically, in Step 4, the undersampling method balances the dataset by reducing the number of samples in the majority class, the SMOTE sampling method balances the data distribution by generating synthetic minority class samples, thereby maximizing the utilization of data, the ADASYN sampling method focuses on the boundary regions with lower density in the minority class samples, improving the generalization ability of the model on different distributions, and the SMOTE+ENN sampling method can achieve a balance between oversampling and undersampling.

[0049] Specifically, in Step 5, multiple HAPE incidence risk prediction models include a support vector machine model, a logistic regression model, a random forest model, a gradient boosting model, a multi-layer perceptron model, an extreme gradient boosting model, and a multi-layer perceptron-gradient boosting integration model.

[0050] Specifically, in Step 5, a voting classifier is used to construct an ensemble model, integrating the gradient boosting model and the multi-layer perceptron as the base learners, and the voting classifier uses the soft voting method to generate the final prediction result by weighted averaging the prediction results of each base learner.

[0051] Specifically, the grid search method is used in Step 5 to ensure the optimal performance of the model under different configurations.

[0052] The present invention provides a method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning:

[0053] Solution

[0054] I. Overall framework

[0055] Construct a machine learning prediction model for evaluating the risk of high altitude pulmonary edema. The overall framework is as Figure 2 shown, mainly including the following five modules: dataset construction, data preprocessing, feature extraction, feature optimization, and classifier construction and evaluation.

[0056] 1. Dataset construction: Construct the HAPE dataset, covering relevant demographic information, vital signs, and biochemical test indicators. Use inpatients with HAPE as positive samples and healthy physical examination populations in the same period as negative samples.

[0057] 2. Data preprocessing: mainly including (1) deletion of outliers to ensure the accuracy of the data; (2) filling of missing values to ensure the integrity of the data; (3) unified variable encoding to facilitate subsequent processing of the data; (4) deletion of duplicate samples to prevent overfitting during model training.

[0058] 3. Feature extraction: Integrate vital signs, demographic features, and biochemical indicators to form the input feature set of the model, and improve the overall expression ability of the model through the fusion of multiple features.

[0059] 4. Feature optimization: To address the problem of data imbalance, various sampling methods are adopted, including undersampling, SMOTE, ADASYN, and SMOTE+ENN. In addition, to further improve the model performance, the grid search technique is combined to optimize the feature selection process.

[0060] 5. Classifier construction and evaluation: In the model construction stage, various machine learning classifiers are tried, including support vector machine (SVM), random forest (RF), logistic regression (Logistic), gradient boosting decision tree (GBDT), multi-layer perceptron (MLP), and an ensemble model based on MLP and GBDT (MLP+GB). Through the comprehensive evaluation of each classifier, the optimal HAPE risk prediction model is determined.

[0061] II. Research objects and data collection

[0062] The data for this study were sourced from hospital subjects. Among them, 118 cases had the observed outcome of high altitude pulmonary edema. The information collected included the patients' basic information, coagulation and biochemical tests, blood routine and five-classification, and novel coronavirus nucleic acid testing during their hospital stay. The data were aligned and anonymized according to the patient ID.

[0063] Inclusion criteria: (1) Age ≥ 18 years; (2) Complete medical records.

[0064] Exclusion criteria: (1) Non-death samples; (2) Abnormal values.

[0065] III. Clinical characteristic variables

[0066] The data initially collected from the electronic medical records consisted of 118 variables, including patient demographics, clinical measurements, and laboratory results. Mainly:

[0067] (1) Patients' basic information: Patient gender (SEX), patient age (AGE);

[0068] (2) Liver function tests: Indirect Bilirubin (IBIL), Total Bilirubin (TBIL), Albumin / Globulin Ratio (A / G), Albumin (Alb), Total Protein (TP), etc.;

[0069] (3) Kidney function tests: Uric Acid (UA), Creatinine (Cr), Urea Nitrogen (UN); (4) Routine tests: Hemoglobin (HGB), Mean Platelet Volume (MPV), Basophils Count (Bas), Monocyte Percentage (Mon), Lymphocyte Count (Lymph), Lymphocyte Percentage (Lymph), Red Blood Cell Count (RBC), Plateletcrit (PCT), Eosinophil Percentage (Eos), Platelet Distribution Width (PDW), Platelet Count (PLT), etc.

[0070] Handling of missing values in variables: If the missing values exceed 30%, the variable is deleted; for discrete data, 0 is used for filling. For continuous data, in the case of normal distribution or approximately normal distribution, mean filling can maintain the overall distribution characteristics of the data and reduce the bias introduced by filling; in the case of skewed data or the presence of obvious outliers, median filling can better represent the central position of the data, reduce the impact on the overall data distribution, and avoid the influence of outliers on the filled values.

[0071] The dataset construction is as Figure 3 shown.

[0072] IV. Statistical method analysis

[0073] To further evaluate the quality of the dataset, Student's t-test is used to compare components of features that conform to the normal distribution, Mann-Whitney U-test is used to analyze features that do not conform to the normal distribution, and chi-square test is used to evaluate the statistical analysis of the differences between groups of categorical variables. All tests are two-sided tests, and the significance level is set at p < 0.05. Statistical analysis of all variables is carried out in the Python 3.10.6 environment.

[0074] V. Sampling method design

[0075] In this study, four representative sampling methods are explored, namely undersampling, SMOTE, ADASYN, and SMOTE+ENN. These methods each have their unique mechanisms and applicable scenarios, and adjust the sample distribution between classes in different ways to improve the performance of the model on imbalanced datasets.

[0076] When dealing with the problem of sample imbalance, undersampling is a common method that balances the dataset by reducing the number of samples in the majority class. Although undersampling can effectively reduce the dataset size, computational complexity, and storage cost, randomly deleting samples from the majority class may lead to information loss, especially for majority-class samples containing important information. In contrast, the SMOTE method balances the data distribution by generating synthetic minority-class samples, thus maximizing the utilization of data. Although SMOTE can improve the model's ability to recognize minority classes, linear interpolation may generate unrealistic samples in the boundary region, increasing the classifier's confusion. Therefore, it is necessary to combine other methods such as data cleaning strategies. ADASYN is an improved method based on SMOTE that pays more attention to the low-density boundary regions in minority-class samples to improve the model's generalization ability across different distributions. However, it should be noted that oversynthesizing samples in sparse regions may introduce noise. The ENN method purifies the dataset by removing outliers in the majority class that cross with minority-class samples. Therefore, by combining the SMOTE and ENN methods, a balance can be achieved between oversampling and undersampling. First, SMOTE is used to generate new minority-class samples, and then the ENN algorithm is used to clean the outliers in the majority class. This combined method can effectively reduce the model bias caused by data imbalance when dealing with complex imbalanced datasets.

[0077] VI. Machine Learning Model Construction

[0078] Machine learning models are widely used in tasks such as classification, regression, clustering, and dimensionality reduction. In this study, seven models for predicting the risk of HAPE onset are constructed based on multiple machine learning algorithms, including Support Vector Machine (SVM), Logistic Regression, Random Forest (RF), Gradient Boosting, Multi-Layer Perceptron (MLP), Extreme Gradient Boosting (XGBoost), and the Multi-Layer Perceptron - Gradient Boosting Ensemble Model (MLP-GradientBoosting). To improve the prediction performance of the models, a Voting Classifier is used to construct an ensemble model, integrating the Gradient Boosting Classifier and the Multi-Layer Perceptron (MLP Classifier) as the base learners. During the hyperparameter optimization process, the grid search method is used to ensure the optimal performance of the model under different configurations. The Voting Classifier adopts the soft voting method, generating the final prediction result by weighted averaging the prediction results of each base learner.

[0079] To compare the performance of different learning algorithms in dealing with data imbalance problems, the data was randomly divided into a training set and a test set at a ratio of 4:1, that is, 80% of the data was used for training, and the remaining 20% of the data was used for independent testing. Subsequently, a five-fold cross-validation method was used to evaluate the model. Finally, all the generated models were evaluated in the independent test set, and the mean values of the corresponding evaluation metrics were calculated to ensure the robustness and generalization ability of the model.

[0080] Experimental Design and Result Analysis

[0081] I. Feature Screening

[0082] To ensure the accuracy and consistency of the research results, the original data was strictly screened and cleaned. First, 3 patients under 18 years old were excluded. Second, considering the importance of data integrity for the training of machine learning models, 55 patients with incomplete data were further excluded. Finally, after strict inclusion and exclusion criteria, a total of 674 eligible subjects were included in the study. These patients were further divided into two groups: 560 cases in the control group (accounting for about 83.0%) and 114 cases in the experimental group (high altitude pulmonary edema, accounting for about 17.0%). Since the number of patients in the high altitude pulmonary edema group is relatively small, the dataset shows obvious imbalance and needs further processing in the follow-up.

[0083] The data processing flow is shown as Figure 4 follows.

[0084] The results of statistical analysis showed that the p-values of variables such as age (AGE), IBIL (indirect bilirubin), TBIL (total bilirubin), AG (albumin / globulin ratio), Alb (albumin), ALP (alkaline phosphatase), AST (aspartate aminotransferase), TP (total protein), and DBIL (direct bilirubin) between the two groups were all less than 0.001, indicating that these variables may play an important role in the risk of HAPE. In addition, some hematological indexes such as PLCC (platelet large cell ratio), HCT (hematocrit), HGB (hemoglobin), RBC (red blood cell count), Mon (monocyte ratio), and Bas (basophil ratio) also showed significant differences (p < 0.001). By screening out a series of important features related to the onset of HAPE in this study, it laid a foundation for the construction of subsequent machine learning models. The statistical analysis of specific variables is shown in Table 1.

[0085] Feature Label = 0 (n = 560) Label = 1 (n = 114) p-value AGE 30.000[25.000,35.000] 25.000[23.000,28.750] <0.001 IBIL 13.240[10.037,15.260] 11.360[7.863,13.240] <0.001 TBIL 20.660[16.155,23.512] 17.725[11.922,20.660] <0.001 AG 1.700[1.610,1.800] 1.635[1.520,1.715] <0.001 Alb 45.930[45.500,47.725] 43.400[40.200,45.930] <0.001 ALP 75.890[66.000,83.000] 72.500[59.000,75.890] <0.001 AST 21.190[17.575,22.500] 19.050[15.225,21.190] <0.001 TP 73.350[72.100,76.225] 70.850[66.000,73.350] <0.001 DBIL 7.420[5.758,8.600] 6.135[4.287,7.420] <0.001 PLCC 60.000[50.000,69.000] 62.000[58.000,76.750] <0.001 HCT 53.150[50.193,57.000] 49.000[44.365,52.970] <0.001 HGB 188.000[178.000,199.000] 170.000[159.000,186.090] <0.001 RBC 5.730[5.380,6.183] 5.285[4.912,5.710] <0.001 Mon 0.380[0.300,0.470] 0.410[0.400,0.558] <0.001 Bas 0.300[0.100,0.400] 0.330[0.300,0.475] <0.001

[0086] Table 1 Statistical Analysis of Variables

[0087] II. Evaluation Metrics

[0088] In the evaluation metrics section, four key metrics used to evaluate the model performance will be introduced in detail: Recall, Accuracy, Area Under the Curve (AUC), Area Under the Precision-Recall Curve (AUPR), and the calculation of the Calibration Curve. The comprehensive application of these metrics helps to comprehensively evaluate the performance of the model in predicting high altitude pulmonary edema under the condition of imbalance between positive and negative sample ratios, thus providing a reliable basis for clinical decision-making. The specific metric formulas are shown in Table 2.

[0089]

[0090] Table 2 Model Evaluation Metrics

[0091] Among them, TP (True Positive): the number of samples where HAPE actually occurred and was correctly predicted as HAPE; FN (False Negative): the number of samples where HAPE actually occurred but was wrongly predicted as not occurring HAPE; TN (True Negative): the number of samples where HAPE did not actually occur and was correctly predicted as not occurring HAPE; FP (False Positive): the number of samples where HAPE did not actually occur but was wrongly predicted as occurring HAPE.

[0092] The calibration curve is used to measure the accuracy of the model's predicted probabilities. For imbalanced data, the traditional ROC curve may be biased towards the performance of the minority class, while the calibration curve can more accurately evaluate the model's performance on various types of samples.

[0093] III. Prediction Performance Analysis

[0094] In this study, 46 eligible features were screened out from the original 118 clinical features, and 7 different machine learning models were constructed using these features. Each model used the same clinical features to ensure a fair performance comparison between different models. To evaluate the performance of these models in predicting high altitude pulmonary edema events, 4 resampling methods were used to evaluate the performance of each model based on the test set data.

[0095] 1. Downsampling

[0096] When the downsampling method is used, the results are shown in Table 3. The random forest (RF), gradient boosting (GB), and MLP+GB ensemble models perform relatively well in terms of AUC, Accuracy, Recall, and AUPR. In particular, the MLP+GB model shows high prediction performance in terms of AUC (0.917±0.048) and Recall (0.858±0.098). This indicates that the downsampling method can improve the prediction ability of the model to a certain extent in the case of data imbalance. However, the performance of the SVM model is significantly poor, with an AUC of only 0.640±0.188, which may be due to its sensitivity to imbalanced data, resulting in a bias towards the majority class.

[0097]

[0098] Table 3 Results of the downsampling method

[0099] 2. SMOTE

[0100] When the SMOTE method is used, the results are shown in Table 4. The prediction performance of each model is significantly improved. In particular, the MLP+GB ensemble model stands out in terms of AUC (0.984±0.012) and AUPR (0.985±0.011), indicating that the SMOTE method significantly enhances the model's ability to identify the minority class, which further proves the effectiveness of the SMOTE method in dealing with imbalanced data. In addition, the AUC of the SVM model increases from 0.640±0.188 in downsampling to 0.818±0.034, showing a large improvement.

[0101]

[0102] Table 4 Results of the SMOTE method

[0103] 3. ADASYN

[0104] When the ADASYN method is used, the results are shown in Table 5. The MLP+GB ensemble model performs excellently, with an AUC of 0.987±0.006 and a Recall of 0.960±0.009. These results show that compared with SMOTE, the ADASYN method can significantly improve the model's ability to identify the minority class and the overall prediction performance, especially in datasets with high complexity and severe imbalance. The excellent performance of the MLP+GB model further proves the advantages of the ensemble learning method, especially in a complex high-dimensional data environment, where it effectively improves the robustness of prediction by integrating the advantages of multiple models.

[0105]

[0106] Table 5 Results of the ADASYN method

[0107] 4. SMOTE+ENN4

[0108] When using the SMOTE+ENN method, the results are shown in Table 6. Since the SMOTE+ENN method not only increases the number of minority class samples but also effectively cleans the noise samples in the data, thus improving the generalization ability of the model. The MLP+GB ensemble model utilizes the SMOTE+ENN method to achieve optimal classification performance, especially on high-complexity and imbalanced datasets, further enhancing the performance of the ensemble model. Therefore, this paper proposes a method that combines the SMOTE+ENN sampling method with the MLP+GB ensemble model for HAPE risk prediction. The experimental results show that the MLP+GB ensemble model performs outstandingly in various indicators, especially reaching the optimal performance in terms of AUC (0.995±0.005), Recall (0.987±0.007), and AUPR (0.996±0.004), enhancing the prediction robustness and accuracy of HAPE risk.

[0109]

[0110] Table 6 Results of the SMOTE+ENN Method

[0111] 5. Ablation Experiment

[0112] In this study, a hybrid sampling method based on SMOTE+ENN is proposed and compared with the no-sampling method to systematically verify the effectiveness of this method in dealing with imbalanced data. The experimental results on various classification models (such as SVM, Logistic regression, RF, XGBoost, GB, MLP, MLP+GB) show that the SMOTE+ENN method can significantly improve the recognition ability of the model on minority class samples, especially outstanding in the Recall indicator.

[0113] The experimental results show that after using SMOTE+ENN, the Recall metrics of all models have been significantly improved. For example, the Recall of SVM has increased from 0.509±0.082 without sampling method to 0.825±0.067, the Recall of MLP has increased from 0.404±0.197 to 0.961±0.038, and the AUPR of MLP+GB has increased from 0.822±0.046 to 0.996±0.004. At the same time, SMOTE+ENN improves the model's ability to identify minority classes while ensuring the relative stability of Accuracy, thus achieving better classification balance. The SMOTE+ENN method can better improve the AUPR of the model in complex models (such as MLP and ensemble models), while for simple models (such as SVM, Logistic), due to the simplicity of the model decision boundary and the interference of noise samples, the AUPR improvement is not obvious or even slightly decreases. Therefore, for dealing with imbalanced datasets, SMOTE+ENN is more suitable for complex models and can significantly enhance the model's identification ability and stability.

[0114]

[0115] Table 7 Results without sampling method

[0116] IV. Calibration Curve Results

[0117] By analyzing the calibration effect of different machine learning models under the SMOTE+ENN sampling method through calibration curves, the results of the 7 proposed models are mainly compared. The MLP+GB method is relatively close to the ideal calibration line (Perfectly Calibrated Line) on multiple segments of the predicted probability, showing good calibration ability. In contrast, traditional models such as logistic regression, random forest, and XGBoost show larger deviations, especially being unstable in the middle and high probability segments. This indicates that the MLP+GB ensemble method has significant advantages in dealing with class imbalance problems.

[0118] V. Model Interpretability Analysis

[0119] To further enhance the interpretability of the HAPE prediction model, this paper uses the SHAP (Shapley Additive Explanations) algorithm to analyze the importance of model input features. SHAP provides an explanation value for each feature, quantifying its impact on the prediction result. The SHAP algorithm is used to conduct interpretability analysis on the MLP+GB ensemble model and generate SHAP summary plots. These plots vividly show the impact of clinical features on the model output results. Figure 5Show the distribution of SHAP values for the top 20 clinical features, where each row represents a feature. The position of the dot represents the SHAP value of the feature, and the higher the value, the greater the contribution of the feature to the model output. Red dots indicate higher feature values, and blue dots indicate lower feature values. The darker the color, the stronger the impact of the feature on the target variable. Figure 6 Then, a bar chart shows the contribution of each feature to the model by arranging them in descending order of the average absolute SHAP value of the feature. The larger the absolute value of SHAP, the greater the impact of the feature on the model output result.

[0120] SHAP values identify the top ten clinical features that have the greatest impact on the model's predictive performance. These features show significant predictive ability in distinguishing patients with the disease from those without, specifically including: Hemoglobin (HGB), Age (AGE), Basophil count (Bas#), Albumin (Alb), Mean Platelet Volume (MPV), Total Protein (TP), Monocyte percentage (MON%), Urea (UN), Red Blood Cell count (RBC), and Lymphocyte percentage (Lymph%).

[0121] High Altitude Pulmonary Edema (HAPE) is a severe high altitude disease, and early intervention has a significant impact on the prognosis of patients. In a high altitude environment, the hemoglobin concentration usually increases to adapt to the hypoxic environment. Currently, certain progress has been made in the research on the pathogenesis and risk prediction of HAPE. In relevant literature, by analyzing the impact of the high altitude environment on the human body, the relationships between physiology, genomics, and environmental factors and HAPE are explored. However, most existing studies rely on traditional statistical analysis methods, mainly focusing on specific risk factor analysis or correlation studies of characteristic variables, lacking effective global prediction models. Additionally, these models usually show certain limitations when dealing with complex high-dimensional data and imbalanced datasets, and are unable to accurately capture the potential mechanisms of HAPE occurrence. Therefore, developing more efficient and accurate machine learning models to predict the onset risk of HAPE is of great significance.

[0122] Consistent with previous studies in terms of the selected indicators, the machine learning model constructed in this study can provide a quantitative assessment of the HAPE onset risk at an early stage, helping medical decision-makers take preventive measures in a timely manner and reducing the harm of HAPE to the health of the high altitude population.

[0123] The SMOTE+ENN method significantly improves the Recall of minority class samples, which is of great significance in clinical disease prediction research. It reflects the proportion of correctly identified actual positive samples by the model. Therefore, in a clinical context, an increase in Recall means that the model can better identify patients with a specific disease; a high Recall means that the model can more comprehensively capture diseased cases and reduce the missed diagnosis rate; when dealing with minority class samples, a model with a high Recall can effectively identify those minority class samples that are easily overlooked; an increase in Recall also means that the model shows stronger sensitivity in identifying high altitude pulmonary edema and can better provide decision-making support for doctors.

[0124] In summary, the SMOTE+ENN method provides more reliable support for such applications by balancing the dataset and improving the performance of minority class samples.

[0125] This study combines multiple machine learning algorithms and data resampling methods to construct a prediction model for the incidence risk of HAPE for the plateau population. By integrating different types of features, including demographics, vital signs, and biochemical indicators, the global expression ability of the model is improved. When dealing with the problem of data imbalance, multiple sampling techniques such as SMOTE, ADASYN, and SMOTE+ENN are adopted, and the prediction performance of the model is effectively improved through grid search technology in feature optimization. In addition, a model integrating a multi-layer perceptron (MLP) and a gradient boosting decision tree (GBDT) further optimizes the final prediction result through a soft voting mechanism. The model in this paper helps to improve health management in the plateau environment, reduce the incidence rate of HAPE and related health risks, provides strong support for the formulation of clinical intervention and prevention strategies, and has important clinical application prospects.

[0126] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning, comprising the following steps: S1. Information collection: Determine the HAPE research objects, and collect the original data of the HAPE research objects. The HAPE research objects include HAPE inpatients and healthy physical examination populations who come to the hospital during the same period. S2. Data preprocessing: Clean and normalize the original data. It is characterized in that It also includes the following steps: S3. Construct a dataset: Use the HAPE inpatients as positive samples and the healthy physical examination populations who come to the hospital during the same period as negative samples, and construct a HAPE dataset based on the positive samples and the negative samples. S4. Resampling processing: Based on the data imbalance in the HAPE dataset, then select a matching resampling method for processing, and adjust the sample distribution between categories multiple times. S5. Establish an ensemble model: Based on the ensemble learning algorithm, construct multiple HAPE onset risk prediction models according to the data in the HAPE dataset, and predict the HAPE onset risk based on multiple HAPE onset risk prediction models.

2. The method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning according to claim 1, wherein In step S1, the original data of the HAPE research objects includes patient basic information data, coagulation and biochemical test data, blood routine and five-category data, and novel coronavirus nucleic acid test data.

3. The method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning according to claim 1, wherein In step S2, the data preprocessing includes deleting outliers in the original data, filling in missing values in the original data, uniformly encoding variables in the original data, and deleting duplicate samples in the original data.

4. The method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning according to claim 1, wherein In step S3, the HAPE dataset includes a demographic information dataset, a vital signs dataset, and a biochemical test index dataset.

5. The method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning according to claim 4, characterized in that, In step S3, the construction steps of the HAPE dataset are as follows: S301. Collect the clinical electronic medical record data of the HAPE research objects. S302. Delete outliers and duplicate samples in the electronic medical record data. S303. Fill in the missing values in the clinical electronic medical record data, and perform unified feature encoding and data annotation on the filled clinical electronic medical record data, and then obtain the HAPE dataset.

6. The method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning according to claim 1, wherein In step S4, the resampling methods include undersampling methods, SMOTE sampling methods, ADASYN sampling methods, and SMOTE+ENN sampling methods.

7. The method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning according to claim 6, wherein In step S4, the undersampling method balances the dataset by reducing the number of samples in the majority class. The SMOTE sampling method balances the data distribution by generating synthetic minority class samples, thereby maximizing the utilization of data. The ADASYN sampling method focuses on the boundary regions with low density in the minority class samples, improving the generalization ability of the model on different distributions. The SMOTE+ENN sampling method can achieve a balance between oversampling and undersampling.

8. The method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning according to claim 1, wherein In step S5, multiple HAPE onset risk prediction models include support vector machine models, logistic regression models, random forest models, gradient boosting models, multi-layer perceptron models, extreme gradient boosting models, and multi-layer perceptron-gradient boosting ensemble models.

9. The method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning according to claim 1, wherein In the step S5, a voting classifier is used to construct an ensemble model, integrating a gradient boosting model and a multi-layer perceptron as base learners. The voting classifier adopts the method of soft voting and generates the final prediction result by weighted averaging the prediction results of each base learner.

10. The method for predicting the risk of high altitude pulmonary edema based on resampling and ensemble learning according to claim 1, wherein, In the step S5, a grid search method is used to ensure the optimal performance of the model under different configurations.

Citation Information

Cited By

  • Plateau pulmonary edema risk prediction method and system based on multi-modal data

    CN121075663A