Prediction model for diagnosis and severity judgment of pulmonary arterial hypertension based on pulmonary arterial angiography, construction method and application
By using pulmonary angiography data and laboratory information, a pulmonary hypertension diagnostic model based on Lasso regression and Logistic regression was constructed, which solved the problem of non-invasive diagnosis, achieved efficient diagnosis and severity assessment of pulmonary hypertension, and is suitable for early screening and diagnosis in remote areas, thus improving the accuracy and coverage of diagnosis.
Patent Information
- Application Number
- CN202411110796.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-08-14
AI Technical Summary
Current diagnostic methods for pulmonary hypertension mainly rely on invasive right heart catheterization, which limits its widespread application, especially in critically ill patients and remote areas where there is a lack of non-invasive and efficient diagnostic methods.
By collecting pulmonary angiography (CTA) data and test results, key variables were screened using Lasso regression and Logistic regression to construct a predictive model for the diagnosis and severity assessment of pulmonary hypertension. These variables included albumin level, ascending aortic diameter, mean platelet volume, pulmonary artery width, and right ventricular output. The model was then validated and visualized using machine learning methods.
It provides a non-invasive diagnostic tool with high sensitivity and specificity, which can identify pulmonary hypertension patients at an early stage, expand the scope of diagnosis, improve survival rate, and is suitable for screening special populations and diagnosis in remote areas. The effectiveness of the model is verified through a variety of machine learning methods.
Smart Images

Figure CN119108095B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information technology, specifically to a predictive model, construction method, and application for the diagnosis and severity assessment of pulmonary hypertension based on pulmonary angiography. Background Technology
[0002] Pulmonary hypertension (PH) is a disease characterized by pulmonary vasoconstriction and pulmonary vascular remodeling caused by various factors. Its global prevalence is approximately 1%, reaching as high as 10% in people over 65 years of age. Furthermore, the 7-year survival rate from the date of diagnosis via right heart catheterization is only 49%. Its high morbidity and mortality rates make early diagnosis of PH a top priority in clinical practice. Currently, the gold standard for PH diagnosis is right heart catheterization, but its invasiveness greatly limits its clinical application. Right heart catheterization is not feasible for some critically ill patients or in some remote areas. Therefore, finding a novel non-invasive diagnostic model for PH is urgently needed. Summary of the Invention
[0003] Pulmonary angiography (CTA) is a simple, easy-to-perform, and non-invasive chest imaging examination. Addressing the shortcomings of existing technologies, this invention discovers that by combining patient CTA and laboratory information, a novel predictive model for the clinical diagnosis and severity assessment of pulmonary hypertension can be constructed. This will provide new clinical clues for the non-invasive diagnosis of pulmonary hypertension. Therefore, this invention provides a predictive model, construction method, and application for the diagnosis and severity assessment of pulmonary hypertension based on pulmonary angiography.
[0004] To achieve the above objectives, the technical solution of the present invention is as follows:
[0005] This invention collects pulmonary angiography (CTA) data and test information from patients and controls with pulmonary hypertension. Lasso regression was used to identify 21 variables that showed differences between the case and control groups. Logistic regression was then used to further identify 5 variables that influence the outcome (whether or not the patient has pulmonary hypertension), and a clinical diagnostic prediction model for pulmonary hypertension was constructed.
[0006] Albumin level: OR 1.258, 95% CI 1.087-1.499, p = 0.005;
[0007] Ascending aortic diameter: OR 0.752, 95% CI 0.641-0.853, p<0.001;
[0008] Mean platelet volume: OR 2.193, 95% CI 1.359-3.856, p = 0.003;
[0009] Pulmonary artery width: OR 1.387, 95% CI 1.204-1.68, p<0.001;
[0010] Right ventricular output: OR value 0.739, 95% CI 0.585-0.903, p = 0.005.
[0011] The area under the ROC curve of this model is AUC 0.92.
[0012] This invention provides new clues for the non-invasive diagnosis of pulmonary hypertension. Through this model, pulmonary hypertension patients can be predicted using only pulmonary artery CTA and blood tests, providing new clues for the non-invasive diagnosis of pulmonary hypertension.
[0013] Based on this, in a first aspect, the present invention provides a method for constructing a pulmonary hypertension diagnostic prediction model based on pulmonary angiography, comprising the following steps:
[0014] S1: Collect the age, gender, pulmonary angiography parameters and test information of the case group and the control group. Use the age, gender, pulmonary angiography parameters and test information of the case group and the control group as variables to construct a sample and obtain the first sample set.
[0015] S2: Remove variables and samples with too many missing values or containing outliers from the first sample set to obtain the second sample set. Specifically, remove variables and samples with more than 30% missing values. Imput the remaining variables using the KMN method. Finally, select useful variables for subsequent analysis.
[0016] S3: Using Lasso regression, multiple potential predictor variables are selected from the second sample set. A third sample set is then constructed by combining all predictor variables from the case and control groups.
[0017] Lasso regression analysis identified 21 potential predictive variables: right ventricular output, pulmonary artery width, ascending aorta width, right ventricular enlargement, right atrial enlargement, absolute neutrophil count, carbon dioxide concentration, lipoprotein level, total bile acid level, direct bilirubin level, albumin level, creatinine level, age, sex, phosphorus level, high-density lipoprotein level, mean platelet volume, atrial septal defect, ventricular septal defect, phospholipid level, and calcium ion level.
[0018] S4: Key variables influencing the outcome are screened from the third sample set using binary logistic regression, and a predictive model is constructed. The outcome refers to whether or not the patient has pulmonary hypertension. The key variables influencing the outcome are the several independent predictive variables finally selected. A fourth sample set is constructed by forming a sample consisting of all key variables from the case group and the control group. The model parameters are obtained by statistically analyzing the key variables in the fourth sample set and visualized using a nomogram. Specifically:
[0019] The key variables that affect the outcome are the five independent predictors that were finally selected: albumin level, ascending aortic width, mean platelet volume, pulmonary artery width, and right ventricular output.
[0020] The formula for the prediction model obtained by the binary logistic regression method is: the probability that a patient has pulmonary hypertension is p.
[0021] logit(p) = -16.217 - 0.303×RCO+ 0.327×PA + 0.23×Alb + 0.785×PltVol- 0.285×AA;
[0022] Where logit(p) is the value of the logistic function; RCO is the right ventricular output; PA is the pulmonary artery width; Alb is the albumin level; PltVol is the mean platelet volume; and AA is the ascending aorta.
[0023] In the formula of the prediction model:
[0024] Albumin level: OR 1.258, 95% CI 1.087-1.499, p = 0.005;
[0025] Ascending aortic width: OR 0.752, 95% CI 0.641-0.853, p<0.001;
[0026] Mean platelet volume: OR 2.193, 95% CI 1.359-3.856, p = 0.003;
[0027] Pulmonary artery width: OR 1.387, 95% CI 1.204-1.68, p<0.001;
[0028] Right ventricular output: OR value 0.739, 95% CI 0.585-0.903, p = 0.005.
[0029] S5: Evaluate the clinical predictive ability of the prediction model through subject characteristic curves, decision curves, and calibration curves;
[0030] S6: Validate the model using random forest and demonstrate its implementation using SHAP;
[0031] S7: The prediction model was validated using 10 machine learning methods: decision tree, gradient boosting tree, LightGBM, logistic classifier, random forest, Naive Bayes, CatBoost, XGBoost, support vector machine and classification perceptron.
[0032] Secondly, the present invention provides a pulmonary hypertension diagnosis prediction model based on pulmonary angiography. The prediction model is constructed using the method described above. The prediction model is visualized using a nomogram, which includes a score scale, a total score scale, and predictive variables. The score scale ranges from 0 to 100 points, the total score scale ranges from 0 to 240 points, and the predictive variables are albumin level, ascending aortic width, mean platelet volume, pulmonary artery width, and right ventricular output.
[0033] Thirdly, this invention provides a predictive model for assessing the severity of pulmonary hypertension based on pulmonary angiography. This predictive model is constructed using multinomial Lasso regression and ordered multi-category logistic regression analysis, with the independent predictive variables being: right ventricular enlargement, left ventricular enlargement, hemoglobin content, atrial septal defect, and fibrinogen content. The predictive model is expressed as follows:
[0034] In the fitting model, different values are assigned to different degrees of pulmonary hypertension: mild pulmonary hypertension is assigned a value of 1, moderate pulmonary hypertension is assigned a value of 2, and severe pulmonary hypertension is assigned a value of 3; therefore, the formula for the fitting model is:
[0035] logit(p 轻 = -6.667 - (0.028 × Hb - 0.012 × FBC - 1.577 × RVE - 3.860 × LVE - 2.392 × DAT),
[0036] Logit(p 轻 +p 中 = -4.796 - (0.028 × Hb - 0.012 × FBC - 1.577 × RVE - 3.860 × LVE - 2.392 × DAT),
[0037] p 重 =1-p 轻 -p 中 ;
[0038] The probability that a sample has mild pulmonary hypertension is p. 轻 The probability that the sample has moderate pulmonary hypertension is p. 中 The probability that the sample has severe pulmonary hypertension is p. 重 Hb is the hemoglobin content, FBC is the fibrinogen content, RVE is the right ventricular enlargement, LVE is the left ventricular enlargement, and DAT is the decreased activity tolerance.
[0039] Fourthly, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for constructing the prediction model as described above.
[0040] The application of the predictive model of this invention is not limited to the conventional clinical diagnosis of pulmonary hypertension, but can also include the following extended applications:
[0041] 1) Early screening for special populations: This model can be used to screen people with risk factors for pulmonary hypertension, such as people who are chronically in a hypoxic environment at high altitudes, and patients with chronic obstructive pulmonary disease, interstitial lung disease, etc. that cause chronic hypoxia.
[0042] 2) A more comprehensive assessment of pulmonary embolism patients: Pulmonary thromboembolism is a disease with insidious symptoms but a dangerous course. Chronic thromboembolism can cause thromboembolic pulmonary hypertension, and the diagnosis of pulmonary thromboembolism often relies on pulmonary angiography. Based on the diagnostic model constructed in this invention, when performing pulmonary angiography on suspected pulmonary thromboembolism patients, in addition to understanding the pulmonary thromboembolism situation, this model can also predict the probability of the patient developing pulmonary hypertension, providing more clues for the patient's subsequent diagnosis and treatment.
[0043] 3) As a medical supplement: For people in remote areas who do not have the conditions to carry out right heart floating catheter examination, pulmonary angiography and laboratory information examination are simple and easy to perform. This model can be used for the diagnosis and screening of pulmonary hypertension, thus expanding the scope of diagnosis and treatment of pulmonary hypertension.
[0044] 4) Scientific research and education: This model can be used as a scientific research tool to find the pathogenesis and outcome factors related to pulmonary hypertension, explore the pathogenesis of pulmonary hypertension, and the interaction between different factors.
[0045] Fifthly, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for constructing the prediction model as described above.
[0046] This invention discloses a diagnostic prediction system for pulmonary hypertension, which constructs the prediction model using any of the methods described above; and implements the application of the prediction model described above in the clinical diagnosis and severity prediction of pulmonary hypertension through a computer program executed by a processor.
[0047] The present invention provides a method for constructing a clinical diagnostic prediction model for pulmonary hypertension patients using pulmonary angiography information, as well as blood routine, liver and kidney function, and electrolyte test information. This model uses various statistical methods, such as Lasso regression analysis and Logistic regression analysis, to screen variables that differ between the pulmonary hypertension group and the control group, thereby identifying independent predictive variables for pulmonary hypertension and constructing a clinical diagnostic model for pulmonary hypertension with high sensitivity and specificity.
[0048] It is particularly important to note that when using the model of this invention, attention must be paid to data privacy protection and ethical review to ensure that all operations comply with relevant laws, regulations, and medical ethics standards. Furthermore, since this model is built based on data from a specific population, differences in factors such as region and ethnicity should be considered when promoting its application, which may require appropriate adjustments and validation of the model.
[0049] The above model comprehensively utilizes pulmonary angiography information, and through statistical methods and machine learning techniques, it screens out variables that have a significant impact on the diagnosis of pulmonary hypertension, and constructs a highly accurate prediction model.
[0050] The technical principle of this invention is as follows:
[0051] In the clinical diagnostic model of pulmonary hypertension:
[0052] First, the samples were divided into case and control groups based on whether they had the disease. Age, gender, pulmonary angiography parameters, and various laboratory indicators were statistically analyzed for both groups. Disease status was used as the dependent variable (a binary variable), and statistical indicators were used as independent variables (including continuous and categorical variables). Lasso regression was used to select the independent variables with the greatest impact on the dependent variable. Lasso regression generates a penalty function to compress the coefficients of variables in the regression model, preventing overfitting and resolving severe multicollinearity. This allows for the selection of the most relevant indicators and variables, establishing an efficient predictive model. In this invention, Lasso regression selected 21 potential predictive variables from 84 statistically analyzed variables.
[0053] Furthermore, using disease status as the dependent variable (a binary variable) and the 21 variables selected by Lasso regression as independent variables, a binary logistic regression analysis was performed. The independent variable with the greatest impact on the dependent variable was selected based on the odds ratio (OR): a larger OR value indicates a stronger association between the independent and dependent variables. Simultaneously, the p-value was used to select independent variables: p < 0.05 was considered to indicate a statistically significant difference in the influence of the independent variable on the dependent variable. Ultimately, logistic regression identified 5 independent variables with high OR values and p < 0.05, constructing an efficient predictive model. The predictive model constructed in this invention is as follows: the probability of a patient having pulmonary hypertension is p, logit(p) = -16.217 - 0.303×RCO + 0.327×PA + 0.23×Alb + 0.785×PltVol - 0.285×AA; where logit(p) is the value of the logistic function; RCO is the right ventricular output; PA is the pulmonary artery width; Alb is the albumin level; PltVol is the mean platelet volume; and AA is the ascending aortic width.
[0054] Finally, the model is visualized using a nomogram. A nomogram assigns a score to each level of each independent variable based on its contribution to the dependent variable (the magnitude of the regression coefficient), then sums the scores of all independent variables to obtain a total score. Finally, the predicted value of the individual's outcome event is calculated by using the functional relationship between the total score and the probability of the outcome event. The nomogram transforms the complex regression equation into a visual graph, making the predictive model results more readable and easier to apply in clinical practice.
[0055] After the model is built, the variables are incorporated into the random forest machine learning model and visualized using SHAP to obtain parameter information such as the ranking of variable importance to the model, thus further understanding the model. Furthermore, the model is validated using 10 machine learning methods, and its performance is assessed using ROC and other methods, providing a comprehensive performance evaluation.
[0056] In the model for assessing the severity of pulmonary hypertension:
[0057] In constructing the logistic regression model, the sample of the pulmonary hypertension case group was analyzed. According to the International Classification of Pulmonary Hypertension Severity, the case group was divided into mild pulmonary hypertension (20 mmHg < mPAP ≤ 40 mmHg), moderate pulmonary hypertension (40 mmHg < mPAP ≤ 55 mmHg), and severe pulmonary hypertension (55 mmHg < mPAP) based on the mean pulmonary artery pressure (mPAP) level. With the severity of pulmonary hypertension as the dependent variable, correlation analysis was used to screen variables that were correlated with the severity of pulmonary hypertension. Then, lasso multinomial regression analysis was used to further screen potential outcome predictors. Finally, ordered ternary logistic regression was used to screen independent predictors and construct a predictive model for the severity of pulmonary hypertension.
[0058] In constructing the linear regression equation, the sample of the pulmonary hypertension case group was analyzed. Correlation analysis was performed between all continuous variables and pulmonary hypertension levels. Based on p < 0.05 and correlation coefficient r > 0.3, six highly correlated variables were selected. These six variables were then subjected to stepwise regression (both) – linear regression. Based on p < 0.05, the regression equation was finally constructed using two variables: pulmonary artery width and red blood cell level. This equation can predict pulmonary artery pressure levels using pulmonary artery width and red blood cell level.
[0059] The advantages and beneficial effects of this invention are as follows:
[0060] 1. Strict inclusion and exclusion criteria: Strict inclusion and exclusion criteria were set in the study design. First, all patients with pulmonary hypertension were diagnosed by the gold standard for pulmonary hypertension diagnosis: mean pulmonary artery pressure measured by right heart catheterization >20 mmHg; all pulmonary angiography results of the included controls were: no obvious abnormalities were found in pulmonary artery CTA.
[0061] 2. Application of Statistical Methods: In the pulmonary hypertension (PH) diagnostic prediction model, Lasso regression eliminated collinear variables and screened out potential predictors of PH, while binary logistic regression identified independent predictors of PH. In the PH severity prediction model, multinomial Lasso regression and ordered ternary logistic regression screened out independent predictors of PH severity. The application of these statistical methods effectively screened key variables related to PH, identified independent predictors of the disease, and constructed high-precision prediction models, which is of great significance for early diagnosis and risk assessment.
[0062] 3. Construction of a clinical diagnostic prediction model for pulmonary hypertension: The final prediction model includes five independent predictive variables: right ventricular output, pulmonary artery width, albumin level, mean platelet volume, and ascending aortic width; and the prediction formula was calculated: the probability of a patient having pulmonary hypertension is p,
[0063] logit(P) = -16.217 - 0.303×RVCO+ 0.327×PA + 0.23×Alb + 0.785×PltVol- 0.285×AA;
[0064] Where logit(p) is the value of the logistic function; RVCO is the right ventricular output; PA is the pulmonary artery width; Alb is the albumin level; PltVol is the mean platelet volume; and AA is the ascending aortic diameter.
[0065] The predictive model is visualized using nomograms, making its clinical application simple and easy. Multi-dimensional evaluations, including ROC (AUC 0.92), calibration curve (mean squared absolute error 0.049), and DCA (good benefits across the prevalence range of 0-1), demonstrate that the model possesses excellent predictive capabilities and high clinical value. It can help clinicians identify pulmonary hypertension patients earlier, enabling timely intervention.
[0066] 4. Validation of machine learning methods: This invention validated the model using 10 different machine learning methods, including decision tree, gradient boosting tree, LightGBM, Logistic classifier, random forest, Naive Bayes, CatBoost, XGBoost, support vector machine, and classification perceptron. The average AUC of the ROC of the 10 machine learning methods for model validation was 0.879, which means that the model showed good predictive performance under different algorithms, enhancing confidence in its practical application.
[0067] 5. Construction of a predictive model for assessing the severity of pulmonary hypertension:
[0068] In the fitting model, different values are assigned to different degrees of pulmonary hypertension: mild pulmonary hypertension is assigned a value of 1, moderate pulmonary hypertension is assigned a value of 2, and severe pulmonary hypertension is assigned a value of 3; therefore, the formula for the fitting model is:
[0069] logit(p 轻 = -6.667 - (0.028 × Hb - 0.012 × FBC - 1.577 × RVE - 3.860 × LVE - 2.392 × DAT),
[0070] Logit(p 轻 +p 中= -4.796 - (0.028 × Hb - 0.012 × FBC - 1.577 × RVE - 3.860 × LVE - 2.392 × DAT),
[0071] p 重 =1-p 轻 -p 中 ;
[0072] The probability that a sample has mild pulmonary hypertension is p. 轻 The probability that the sample has moderate pulmonary hypertension is p. 中 The probability that the sample has severe pulmonary hypertension is p. 重 Hb is the hemoglobin content, FBC is the fibrinogen content, RVE is right ventricular enlargement, LVE is left ventricular enlargement, and DAT is decreased exercise tolerance.
[0073] The model constructs predictive formulas for mild, moderate, and severe pulmonary hypertension, which can help clinicians quickly determine the severity of pulmonary hypertension and provide important basis for subsequent targeted diagnosis and treatment.
[0074] 6. Construction of the linear regression equation for mean pulmonary artery pressure: Through correlation analysis and linear regression analysis, a regression equation with mean pulmonary artery pressure as the dependent variable was constructed: y = -15.949 + 1.123× PA + 8.808× RBC +ε, where y is the mean pulmonary artery pressure level, PA is the pulmonary artery width, and RBC is the red blood cell level.
[0075] In summary, this invention, through strict inclusion and exclusion criteria, precise statistical methods, effective clinical prediction model construction, and multi-faceted model validation, demonstrates the significant advantages and potential beneficial effects of pulmonary angiography and related laboratory information in the diagnosis and prediction of pulmonary hypertension. These findings are expected to be widely applied in clinical practice, providing a new tool for pulmonary hypertension diagnosis for patients with contraindications to right ventricular catheterization and those in remote areas where such examination is not available. This is of great significance for early diagnosis and treatment of pulmonary hypertension and improving survival rates. Attached Figure Description
[0076] Figure 1 This is a flowchart of the model construction process for the present invention.
[0077] Figure 2 This describes the Lasso regression screening process of the present invention. Wherein: A: Lasso coefficient path diagram; B: Lasso regression analysis cross-validation curve.
[0078] Figure 3 This is the forest diagram of the prediction model of this invention.
[0079] Figure 4 This is a nodal plot of the prediction model of this invention.
[0080] Figure 5 For the multi-dimensional evaluation of the prediction model of this invention, A: ROC curve; B: DCA curve; C: calibration curve.
[0081] Figure 6 To visualize the prediction model validated by random forest using SHAP (SHapley Additive exPlanations) in this invention; where A: SHAP bee graph; B: SHAP importance ranking graph; C: SHAP heatmap; D: SHAP waterfall plot; E: SHAP heatmap.
[0082] Figure 7 The following is a validation graph for the 10 machine learning multi-prediction models of this invention, where A: ROC of multiple machine learning models; B: summary of ROC curves and confidence intervals of multiple machine learning models; C: precision and completeness curves of multiple machine learning models; and D: DCA curves of multiple machine learning models.
[0083] Figure 8 This describes the Lasso multinomial regression screening process of the present invention. Wherein: A: Lasso coefficient path diagram; B: Lasso coefficient path diagram; C: Lasso coefficient path diagram; D: Lasso regression analysis cross-validation curve.
[0084] Figure 9 This is a validation graph for the prediction model of the present invention. Wherein, A: heatmap of three-class logistic regression; B: ROC of the prediction model.
[0085] Figure 10 Correlation analysis of pulmonary artery hypertension levels in this invention includes: A: Correlation analysis of pulmonary artery width (PA) and mean pulmonary artery pressure; B: Correlation analysis of hemoglobin level and mean pulmonary artery pressure; C: Correlation analysis of red blood cell level (RBC) and mean pulmonary artery pressure; D: Correlation analysis of hematocrit (HCT) and mean pulmonary artery pressure; E: Correlation analysis of elevated body mass index (BMI) and mean pulmonary artery pressure; and F: Correlation analysis of fibrinogen content and mean pulmonary artery pressure. Detailed Implementation
[0086] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0087] Example 1: Construction of a clinical prediction model for the diagnosis of pulmonary hypertension based on pulmonary angiography (CTA) and laboratory information.
[0088] S1. Collect general information, CTA parameters, and laboratory information of patients with pulmonary hypertension (PH) and control groups who underwent pulmonary angiography (CTA) at our hospital from April 2022 to April 2024. All PH patients were diagnosed with pulmonary hypertension (mean pulmonary artery pressure >20 mmHg) via right heart catheterization. The conclusions of all control CTA examinations were: no obvious abnormalities were found in the pulmonary artery CTA.
[0089] The general information collected includes: gender and age;
[0090] CTA information includes: right atrial enlargement, right ventricular enlargement, left atrial enlargement, left ventricular enlargement, atrial septal defect, ventricular septal defect, right ventricular transverse diameter, right ventricular wall thickness, right atrial superior-inferior diameter, right atrial left-right diameter, left atrial anteroposterior diameter, left atrial superior-inferior diameter, left atrial left-right diameter, left ventricular transverse diameter, left ventricular end-diastolic volume (LEDV), left ventricular end-systolic volume (LESV), left ventricular stroke volume (LSV), left ventricular ejection fraction (LEF), left ventricular mass (Leftmass), left ventricular output (LeftCO), right ventricular end-diastolic volume (REDV), right ventricular end-systolic volume (RESV), right ventricular stroke volume (RSV), right ventricular ejection fraction (REF), and right ventricular output (Right). CO), pulmonary artery width (PA), left pulmonary artery diameter (LPA), right pulmonary artery diameter (RPA), McGoon ratio, diameter of the aortosinic junction, ascending aorta diameter, descending aorta diameter, descending aorta diameter at the diaphragmatic level, and total calcification score.
[0091] Inspection information includes:
[0092] Complete blood count: white blood cells, red blood cells, hemoglobin, platelets, neutrophil percentage, lymphocyte percentage, monocyte percentage, eosinophil percentage, basophil percentage, absolute neutrophil count, absolute lymphocyte count, absolute monocyte count, absolute eosinophil count, absolute basophil count, hematocrit, mean corpuscular volume, mean corpuscular hemoglobin content, mean corpuscular hemoglobin concentration, red blood cell distribution width (cv), red blood cell distribution width (SD), mean platelet volume;
[0093] Coagulation function: prothrombin time, international normalized ratio, prothrombin activity, activated partial thromboplastin time, thrombin time, fibrinogen content;
[0094] D2 dimer;
[0095] NTProBNP;
[0096] Procalcitonin (PCT);
[0097] Liver and kidney function electrolytes: ALT, AST, ALT / AST ratio, direct bilirubin, indirect bilirubin, total bilirubin, total protein, albumin, globulin, albumin / globulin ratio, gamma-glutamyl transferase, alkaline phosphatase, total bile acids, creatinine, blood urea nitrogen, uric acid, CO2, serum cystatin C, potassium, sodium, chloride, calcium, magnesium, phosphorus, total cholesterol, triglycerides, high-density lipoprotein, low-density lipoprotein, small low-density lipoprotein, lipoprotein A, free fatty acids, phospholipids;
[0098] Thyroid function: free T3, free T4, thyroid-stimulating hormone;
[0099] Immune function: Anticardiolipin antibody IgM, anticardiolipin antibody IgG, anticardiolipin antibody IgA, complement C3, complement C4, immunoglobulin IgG, immunoglobulin IgA, immunoglobulin IgM, immunoglobulin IgE, anti-streptolysin (ASO), rheumatoid factor (RF), rheumatoid factor antibody IgM, rheumatoid factor antibody IgG, rheumatoid factor antibody IgA;
[0100] Blood gas analysis: pH value, oxygen partial pressure, oxygen saturation, CO2 partial pressure, body temperature, lactate, standard bicarbonate, actual base excess, standard base excess, anion gap;
[0101] S2. Variables and samples with missing values >30% were removed from the sample. The remaining variables were imputed using the KMN method. Finally, 84 variables, including 59 cases of pulmonary hypertension and 43 cases of control, were included in the subsequent analysis, as shown in Table 1.
[0102] Table 1. Samples and variables included in the model
[0103]
[0104]
[0105]
[0106] S3. Use Lasso regression to screen for variables that show differences between the case group and the control group. Figure 2 China A, Figure 2 (B) Variables exhibiting collinearity were also excluded. A total of 21 potential predictive variables were identified: right ventricular output, pulmonary artery width, ascending aortic diameter, right ventricular enlargement, right atrial enlargement, absolute neutrophil count, carbon dioxide concentration, lipoproteins, total bile acids, direct bilirubin, albumin, pulmonary artery width, age, sex, phosphorus level, high-density lipoprotein, mean platelet volume, presence of atrial septal defect, presence of ventricular septal defect, phospholipids, and calcium ion level.
[0107] S4. Further, logistic regression was used to screen for variables that influence the outcome (whether or not the disease occurs), and a clinical prediction model was constructed. The predictive ability of the clinical model was assessed using the area under the curve. Five independent predictors were ultimately selected (Table 2): albumin level, ascending aortic diameter, mean platelet volume, pulmonary artery width, and right ventricular output. See the forest plot for details. Figure 3 ).
[0108] The probability of the sample having pulmonary hypertension is p.
[0109] logit(p) = -16.217 - 0.303×RCO+ 0.327×PA + 0.23×Alb + 0.785×PltVol- 0.285×AA;
[0110] Where logit(p) is the value of the logistic function; RCO is right ventricular output; PA is pulmonary artery width; Alb is albumin level; PltVol is mean platelet volume; and AA is ascending aortic diameter. The model is visualized using a nomogram. Figure 4 ).
[0111] Table 2. Five independent predictors selected by binary logistic regression.
[0112]
[0113] S5. Evaluate the clinical predictive ability, net benefit, and accuracy of the prediction model using ROC, DCA, and calibration curves; after the prediction model is constructed, evaluate the model through multi-dimensional parameter calculations. The area under the constructed ROC curve (Receiver Operating Characteristic curve) is 0.92 ( Figure 5 The result (A) indicates that the prediction model has good accuracy. Decision curve analysis (DCA) ( Figure 5 (B) indicates that the model has good net benefit when the probability of a patient having pulmonary hypertension is within the range of 0-1. The predictive model was internally validated using a computer-simulated bootstrap sampling method. After 300 repeated samplings, the calibration curve was used. Figure 5 In section C), the model's predictive performance was evaluated. As shown in the figure, the calibration curve follows the y=x curve, and the mean absolute error is 0.049, indicating that the model's predicted probability is highly consistent with the actual observed probability.
[0114] S6. Validate the model using random forest and visualize it using SHAP; In this embodiment, random forest is used to validate the constructed model, and SHAP is used for intuitive description, through SHAP bee graph ( Figure 6 (A) and importance ranking chart ( Figure 6 As shown in Figure B, the variables with the greatest to least impact on pulmonary hypertension outcomes are, in descending order: pulmonary artery width, right atrial size, E / E', mean platelet volume, left ventricular posterior wall thickness, and pulmonary valve orifice velocity. According to the SHAP heatmap (… Figure 6 As shown in Figure C, variables such as pulmonary artery width, right atrial size, and left ventricular posterior wall thickness have a negative impact on the deviation of the SHAP value from the baseline, while variables such as pulmonary valve velocity, E / E', and mean platelet volume have a positive impact; and right atrial size and left ventricular posterior wall thickness have the greatest impact on the SHAP value. Figure 6 (D), while the greater the pulmonary artery width, the lower the SHAP value ( Figure 6 (E).
[0115] S7. The model is validated using 10 machine learning methods: Decision Tree, Gradient Boosting Decision Tree (GBDTTs), Light GBM (LGBMTs), Logistic Classifier, Random Forest, Naive Bayesian Trees (NBTs), CatBoost, XGBoost, Support Vector Machine (SVMTs), and Multilayer Perceptron (MLPs). In this embodiment, the test set ratio is set to 0.3, and the training and test sets are split using a random number of 1. The methods used are Decision Trees (Decision Trees), Gradient Boosting Decision Trees (GBDTTs), Light GBM (LGBMTs), Logistic Classifier (Logistic Trees), Random Forest (RFTs), Naive Bayesian Trees (NBTs), CatBoost, XGBoost, Support Vector Machine (SVMTs), and Multilayer Perceptron (MLPs). The constructed model was validated multiple times using 10 machine learning methods, including Trees (Table 3). The results are shown in the figure. The model exhibits good sensitivity and specificity under various machine learning methods. Figure 7 (A) The average area under the ROC curve for 10 machine learning validation models is 0.879. Figure 7 (B), Precision curve (PR curve) Figure 7(C) indicates that this model has good precision and recall under all 11 machine learning methods (Table 3), DCA curve ( Figure 7 (D) indicates that the application of this model has good net benefits. In summary, the verification results show that the pulmonary hypertension diagnosis and prediction model constructed in this invention has good predictive ability.
[0116] Table 3. Validation of the model using various machine learning methods
[0117]
[0118] Example 2. A predictive model for assessing the severity of pulmonary hypertension was constructed based on pulmonary angiography and laboratory information of a group of pulmonary hypertension cases.
[0119] S1. Extract and analyze the information of the pulmonary hypertension case group separately. Extract the information of the pulmonary hypertension case group (n=59) compiled in step S2 of Example 1. According to the International Classification of Pulmonary Hypertension Severity, the case group is divided into mild pulmonary hypertension (20mmHg<mPAP≤40mmHg), moderate pulmonary hypertension (40mmHg<mPAP≤55mmHg), and severe pulmonary hypertension (55mmHg<mPAP) according to the mean pulmonary artery pressure (mPAP) level.
[0120] S2. Variables collected from pulmonary angiography and laboratory data were used as the first part of the independent variables. The following statistics were used as the second part of the independent variables: basal metabolic rate, height, weight, BMI (body mass index), heart rate, symptom duration (days), shortness of breath after activity (binary variable), chest tightness (binary variable), dyspnea (binary variable), palpitations (binary variable), decreased exercise tolerance (binary variable), dizziness (binary variable), fatigue (binary variable), bilateral lower extremity edema (binary variable), and cough (binary variable). The first and second parts of the independent variables together constitute the independent variables for the entire sample. The outcome variable (mild, moderate, and severe pulmonary hypertension) was used as the dependent variable. Spearman correlation analysis was used to screen variables that were correlated with the outcome variable; p < 0.05 was considered a correlation. A total of 9 variables were obtained (units of the variables are as shown in Table 4).
[0121] Table 4. Variables correlated with the severity of pulmonary hypertension
[0122]
[0123] S3. The obtained 9 variables were included in Lasso multinomial regression analysis to screen out potential predictor variables. Figure 8The variables A, B, C, and D were included in the ordered trinomial logistic regression model. Based on p < 0.05, the final prediction model (Table 5) was constructed from the following 5 independent predictor variables (the units of the variables are the same as in the table): hemoglobin content, fibrinogen content, right ventricular enlargement, left ventricular enlargement, and decreased exercise tolerance.
[0124] Table 5. Independent predictors and model parameters selected by ordered ternary logistic regression
[0125]
[0126] In the fitting model, different values are assigned to different degrees of pulmonary hypertension: mild pulmonary hypertension is assigned a value of 1, moderate pulmonary hypertension is assigned a value of 2, and severe pulmonary hypertension is assigned a value of 3; therefore, the formula for the fitting model is:
[0127] logit(p 轻 = -6.667 - (0.028 × Hb - 0.012 × FBC - 1.577 × RVE - 3.860 × LVE - 2.392 × DAT),
[0128] Logit(p 轻 +p 中 = -4.796 - (0.028 × Hb - 0.012 × FBC - 1.577 × RVE - 3.860 × LVE - 2.392 × DAT),
[0129] p 重 =1-p 轻 -p 中 ;
[0130] The probability that a sample has mild pulmonary hypertension is p. 轻 The probability that the sample has moderate pulmonary hypertension is p. 中 The probability that the sample has severe pulmonary hypertension is p. 重 Hb is the hemoglobin content, FBC is the fibrinogen content, RVE is right ventricular enlargement, LVE is left ventricular enlargement, and DAT is decreased exercise tolerance.
[0131] According to the estimated coefficients, the hemoglobin content has a positive coefficient, indicating that the higher the hemoglobin content, the more severe the pulmonary hypertension in the sample is likely. On the other hand, the logistic regression coefficients for fibrinogen content, right ventricular enlargement, left ventricular enlargement, and decreased exercise tolerance are negative, indicating that the lower the fibrinogen content, the more severe the pulmonary hypertension is likely to be, if right ventricular enlargement, left ventricular enlargement, or decreased exercise tolerance is present.
[0132] S4. Heatmap of multi-class logistic regression matrix ( Figure 9(A) and ROC curve ( Figure 9 As shown in Figure B), the area under the ROC curve for predicting mild pulmonary hypertension is 0.842, for moderate pulmonary hypertension it is 0.55, and for severe pulmonary hypertension it is 0.918. This indicates that the model has a relatively weak predictive ability for moderate pulmonary hypertension, but has good predictive ability for both mild and severe pulmonary hypertension.
[0133] Example 3: Construction of a Linear Regression Model
[0134] S1. In the variables statistically analyzed in Example 2, correlation analysis was performed between all continuous variables and mean pulmonary artery pressure. Based on p < 0.05 and correlation coefficient r > 0.3, six variables with high correlation were selected: pulmonary artery width ( Figure 10 (A) Hemoglobin content ( Figure 10 (B) Red blood cell level ( Figure 10 C), hematocrit ( Figure 10 D), Body Mass Index (BMI) Figure 10 E), fibrinogen content ( Figure 10 (Middle F).
[0135] S2. Perform stepwise regression (both) - linear regression on the above 6 variables. Based on p < 0.05, the final two variables PA and RBC are obtained, and the regression equation is constructed as follows: y = -15.949 + 1.123 × PA + 8.808 × RBC + ε, where y is the mean pulmonary artery pressure level, PA is the pulmonary artery width, and RBC is the red blood cell level.
Claims
1. A method for constructing a diagnostic and predictive model for pulmonary hypertension based on pulmonary angiography, characterized in that: Includes the following steps: S1: Collect the age, gender, pulmonary angiography parameters and test information of the case group and the control group. Use the age, gender, pulmonary angiography parameters and test information of the case group and the control group as variables to construct a sample and obtain the first sample set. S2: Remove variables and samples with too many missing values or containing outliers from the first sample set to obtain the second sample set; S3: Using Lasso regression, multiple potential predictor variables are screened out in the second sample set. The sample is composed of all predictor variables from the case group and the control group to form the third sample set. Among them, 21 potential predictive variables were screened out by Lasso regression analysis, namely: right ventricular output, pulmonary artery width, ascending aorta width, right ventricular enlargement, right atrial enlargement, absolute neutrophil count, carbon dioxide concentration, lipoprotein level, total bile acid level, direct bilirubin level, albumin level, creatinine level, age, sex, phosphorus level, high-density lipoprotein level, mean platelet volume, atrial septal defect, ventricular septal defect, phospholipid level, and calcium ion level. S4: Key variables that affect the outcome are screened in the third sample set by binary logistic regression and a predictive model is constructed. The outcome refers to whether or not the patient has pulmonary hypertension. The key variables that affect the outcome are the several independent predictive variables that are finally screened. A fourth sample set is constructed by forming a sample of all key variables in the case group and the control group. The model parameters are obtained by statistically analyzing the key variables in the fourth sample set and visualized by a nomogram. Among them, the key variables that affect the outcome are the five independent predictors that were finally selected: albumin level, ascending aortic width, mean platelet volume, pulmonary artery width, and right ventricular output. The formula for the prediction model obtained by the binary logistic regression method is: the probability that a patient has pulmonary hypertension is p. logit(p) = -16.217 - 0.303×RCO+ 0.327×PA + 0.23×Alb + 0.785×PltVol-0.285×AA; Where logit(p) is the value of the logistic function; RCO is the right ventricular output; PA is the pulmonary artery width; Alb is the albumin level; PltVol is the mean platelet volume; and AA is the ascending aortic diameter. S5: Evaluate the clinical predictive ability of the prediction model through subject characteristic curves, decision curves, and calibration curves; S6: Validate the model using random forest and demonstrate its implementation using SHAP; S7: The prediction model was validated using 10 machine learning methods: decision tree, gradient boosting tree, lightweight gradient boosting machine (LightGBM), logistic classifier, random forest, Naive Bayes, classification boosting (CatBoost), extreme gradient boosting (XGBoost), support vector machine, and classification perceptron.
2. The construction method according to claim 1, characterized in that: Step S2 specifically involves: removing variables and samples with missing values greater than 30% from the sample, imputing missing values for the remaining variables using the KMN method, and finally selecting useful variables for subsequent analysis.
3. The construction method according to claim 1, characterized in that: In the formula of the prediction model in step S4: Albumin level: OR 1.258, 95% CI 1.087-1.499, p = 0.005; Ascending aortic diameter: OR 0.752, 95% CI 0.641-0.853, p<0.001; Mean platelet volume: OR 2.193, 95% CI 1.359-3.856, p = 0.003; Pulmonary artery width: OR 1.387, 95% CI 1.204-1.68, p<0.001; Right ventricular output: OR value 0.739, 95% CI 0.585-0.903, p = 0.
005.
4. A diagnostic and predictive model for pulmonary hypertension based on pulmonary angiography, characterized in that: The prediction model is constructed using the method described in claim 3. The prediction model is visualized using a nomogram, which includes a score scale, a total score scale, and predictive variables. The score scale is 0-100 points, the total score scale is 0-240 points, and the predictive variables are albumin level, ascending aortic width, mean platelet volume, pulmonary artery width, and right ventricular output.
5. A predictive model for assessing the severity of pulmonary hypertension based on pulmonary angiography, characterized in that: The predictive model was constructed using multinomial Lasso regression and ordered multi-category logistic regression analysis. Its independent predictive variables are: right ventricular enlargement, left ventricular enlargement, hemoglobin content, atrial septal defect, and fibrinogen content. The predictive model is expressed as follows: In the fitted model, different degrees of pulmonary hypertension are assigned different values: mild pulmonary hypertension is assigned a value of 1, moderate pulmonary hypertension is assigned a value of 2, and severe pulmonary hypertension is assigned a value of 3. The formula for the fitted model is: logit(p 轻 )=-6.667-(0.028×Hb-0.012×FBC-1.577×RVE-3.860×LVE-2.392×DAT), Logit(p 轻 +p 中 )=-4.796-(0.028×Hb-0.012×FBC-1.577×RVE-3.860×LVE-2.392×DAT), p 重 =1-p 轻 -p 中 ; The probability that a sample has mild pulmonary hypertension is p. 轻 The probability that the sample has moderate pulmonary hypertension is p. 中 The probability that the sample has severe pulmonary hypertension is p. 重 Hb is the hemoglobin content, FBC is the fibrinogen content, RVE is right ventricular enlargement, LVE is left ventricular enlargement, and DAT is decreased exercise tolerance.
6. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the method for constructing the prediction model as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method for constructing the prediction model as described in any one of claims 1 to 4.
Citation Information
Patent Citations
SAP prediction model construction method and device based on machine learning, and product
CN118296457A
Construction method and application of early-stage acute kidney injury prediction model after cardiac arrest resuscitation
CN118430836A