Model for predicting pulmonary embolism risk of patient with severe community-acquired pneumonia and establishment method
By constructing a nomogram model based on LASSO regression and multivariate logistic regression, combined with data processing and nonlinear fitting, the problem of insufficient accuracy in PE risk prediction in ICU-SCAP patients was solved, efficient individualized risk assessment and early intervention were achieved, and the stability and clinical practicality of the model were improved.
Patent Information
- Application Number
- CN202511000070.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-17
Smart Images

Figure CN120809225A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of clinical medical technology, and particularly relates to a severe community-acquired pneumonia patient pulmonary embolism risk prediction model and a method for establishing the same. BACKGROUND
[0002] Severe community-acquired pneumonia (SCAP) has the characteristics of high morbidity and high mortality, and its mortality can be as high as 30%-50%, which is one of the main burdens of medical and health resources. Pulmonary embolism (PE) is a common and potentially fatal disease with high morbidity, high mortality and high disability rate, which can significantly prolong the hospitalization time of patients. Studies have shown that in the community environment, acute infection is associated with a transient increase in the risk of venous thromboembolic events (including deep vein thrombosis and PE). Domestic and foreign clinical experience also suggests that PE is one of the common complications of SCAP patients. However, due to the insufficient detection rate of CT pulmonary angiography (CTPA), the incidence of SCAP-related PE events may be underestimated. In addition, there is currently a lack of systematic research on the epidemiological characteristics, risk of onset and related risk factors of SCAP-related PE events at home and abroad. Therefore, the present application aims to construct and verify a nomogram model for predicting the risk of PE in SCAP patients in the intensive care unit (ICU), to improve the detection efficiency of CTPA, and to provide an effective tool for early identification and intervention in clinical practice, so as to delay or even prevent the progression of such diseases, and reduce the morbidity and mortality.
[0003] Currently, the models for predicting the risk of PE in SCAP patients can be divided into three categories: clinical scoring models, statistical models, and machine learning models. Clinical scoring models are mostly based on traditional scales such as the Wells score and the modified Geneva score, which are used to assess the probability of PE in the general population. However, these scoring tools have not been optimized for the specific characteristics of ICU-SCAP patients, and therefore their predictive accuracy and applicability in this population are still limited. Statistical models usually use single-factor or multi-factor logistic regression methods, combining patient clinical symptoms, physiological parameters, and laboratory indicators such as D-dimer levels to construct PE risk prediction models. These methods have good interpretability and operability, and are suitable for preliminary risk assessment, but they have limited ability to handle high-dimensional data, variable interactions, and nonlinear relationships. In recent years, with the development of artificial intelligence technology, machine learning models have been increasingly applied in medical prediction. These models can automatically identify complex nonlinear relationships and potential interactions by training on large amounts of clinical and imaging data. For example, studies have used random forest and deep neural network algorithms to construct PE risk models for general inpatients, showing better predictive performance than traditional statistical models. Other studies have combined radiomics features with patient clinical parameters to significantly improve the accuracy of identifying high-risk PE patients.
[0004] Currently, there are no large-scale prediction model studies targeting the risk of PE in ICU-SCAP patients. Existing PE risk prediction models are mostly based on the general population and do not fully consider the clinical characteristics and complexity of disease evolution in ICU-SCAP patients, resulting in significant shortcomings in applicability, specificity, interpretability, and clinical utility in this specific population. Nomograms, as a visual risk assessment tool, have the ability to integrate multiple key clinical variables and quantify individual risk, and have been widely used in the construction of prediction models for various diseases. When constructing a prediction model for the risk of PE in ICU-SCAP patients, a reasonable feature selection strategy, statistical modeling method, or machine learning algorithm should be combined to balance the predictive performance and clinical interpretability of the model. At the same time, the stability and generalizability of the model should be evaluated through model validation. Therefore, how to establish a PE risk prediction model that is suitable for ICU-SCAP patients, has good predictive performance, and is easy to use in clinical practice is a problem that needs to be solved in this technical field.
[0005] To address the above problems, a risk prediction model for ICU-SCAP related PE patients is needed to solve the problems existing in the prior art and provide important evidence-based support for early identification of high-risk groups of SCAP patients with PE, optimization of treatment decisions, and improvement of prognosis. SUMMARY
[0006] The present application aims to provide a risk prediction model for ICU-SCAP related PE patients and a method for establishing the same. The method is based on the Critical Illness Database Market-IV (version 2.2), and the data set included in the analysis is randomly divided into a training set and a validation set in a ratio of 7:3. Independent risk factors are screened by using least absolute shrinkage and selection operator (LASSO) regression and multivariate logistic regression, and a nomogram model is constructed based on R software. The nomogram prediction model shows good discrimination, calibration, and clinical applicability in the area under the receiver operating characteristic curve (AUC), calibration curve, Hosmer-Lemeshow test, and decision curve analysis (DCA). According to the existing literature retrieval, the present application first develops and verifies a nomogram model for predicting the risk of SCAP patients complicated with PE. The tool realizes rapid risk assessment by integrating important clinical parameters, and provides important evidence-based support for early identification of high-risk patients, optimization of anticoagulation decisions, and improvement of prognosis.
[0007] The pulmonary embolism risk prediction model for severe community-acquired pneumonia patients and the method for establishing the same according to the present application specifically include the following steps: Step 1. Input patient data from the Critical Illness Database Market-IV (version 2.2); Step 2. Organize and clean the data, especially handle missing values and outliers; Step 3. Randomly divide the data into a training set and a validation set in a ratio, and analyze the relationship between population and baseline characteristics of the divided data set; Step 4. Determine the independent risk factors of pulmonary embolism (PE) in ICU-SCAP patients by using LASSO regression and multivariate logistic regression, and construct a risk prediction model based on the training set; Step 5. Perform internal validation of the risk prediction model based on the validation set, and evaluate its discrimination, calibration, and clinical benefit.
[0008] Preferably, the patient data inclusion criteria in Step 1 are: patients diagnosed with SCAP in the ICU, i.e. community-acquired pneumonia patients who require tracheal intubation or septic shock after active fluid resuscitation still require treatment with vasoactive drugs, with International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) codes: 4800, 4801, 4802, 4803, 4809, 481, 4822, 48231, 48283, 48284, 4830, 4831, 486; International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) codes: J120, J121, J122, J1281, J129, J13, J14, J156, J157, J160, J189; patients aged ≥18 and ≤89 years old, multiple ICU admissions, and patients with first ICU admission data and first ICU admission time exceeding 24 hours; The patient data exclusion criteria are: non-first ICU admission records, patients aged <18 or >89 years old, patients who died or were discharged within 24 hours of first ICU admission, pregnant and lactating women, or total hospitalization duration shorter than ICU hospitalization duration.
[0009] Preferably, the data sorting and cleaning process in Step 2 includes: Step 2-1, delete data in Step 1 that is incorrect in content, format, logic, or unrelated to the study; Step 2-2, use R software package "ggplot2" to draw a specific missing proportion level bar chart of the variable, then remove variables with a missing rate >60%, and then use R software package "mice" for Bayesian multiple imputation to handle missing values; The specific implementation is as follows: Step 2-2-1, missing rate calculation and summary generation: The following R code is used to calculate the missing proportion in the original data: Open the data view View(scap) (only applicable to interactive environment) to ensure that the data format is data.frame; scap <- as.data.frame(scap) Calculate the missing rate and number of each variable missing_summary <- data.frame( variable = colnames(scap), missing_rate = colMeans(is.na(scap)) * 100, # missing rate (percentage) missing_count = colSums(is.na(scap)) # missing number ) View the missing statistics summary head(missing_summary) View(missing_summary) illustrate: colMeans(is.na(scap)) * 100 is used to calculate the missing rate (percentage) of each variable; colSums(is.na(scap)) is used to calculate the number of missing values; The results are stored in the missing_summary data frame, which contains the following fields: variable: variable name; missing_rate: missing rate (%); missing_count: missing number; This table provides the basis for subsequent visualization and variable screening.
[0010] Step 2-2-2, draw a horizontal bar chart of the variable missing rate: Use the ggplot2 package to draw a horizontal bar chart of missing rates to visually demonstrate the quality of the variables: library(ggplot2) ggplot(missing_summary, aes(x = reorder(variables, missing_count), y= missing_count)) + geom_bar(stat = "identity", fill = "skyblue") + labs(x = "variables", y = "Number") + ggtitle("Horizontal Bar Chart of Missing Rates") + theme_minimal() + coord_flip() + # horizontal bar chart illustrate: Sort by missing rate using the reorder() function; coord_flip() is used to convert the bar chart to a horizontal display; Horizontal bar chart helps to quickly identify high missing variables. Step 2-2-3, remove variables with missing rate higher than 60% To ensure data quality, remove variables with missing rate higher than 60%. R code as follows: library(dplyr) # Filter variables with missing rate ≤ 60% filtered_variables <- missing_summary %>% filter(missing_rate <= 60) %>% pull(variable) Keep the filtered variables filtered_scap <- scap %>% select(all_of(filtered_variables)) Check the processing result dim(filtered_scap) head(filtered_scap) Explanation: Only keep variables with missing rate not higher than 60% to form a new dataset filtered_scap; Use the dplyr package to achieve filtering and reconstruction; Provide data basis for subsequent imputation. Step 2-2-4, Bayesian multiple imputation of missing values Use the mice package to perform Bayesian multiple imputation (Multiple Imputation by Chained Equations) on the remaining missing values in the filtered_scap dataset: library(mice) Check variable type str(filtered_scap) Set the imputation method method <- make.method(filtered_scap) method[categorical_vars] <- "logreg" # Use logistic regression for binary variables method[!names(filtered_scap) %in% categorical_vars] <- "pmm" # continuous variables use PMM Perform multiple imputation set.seed(123) mice_result <- mice(filtered_scap, m = 5, method = method, seed =123) Explanation: categorical_vars is a set of pre-defined binary variable names; Logistic regression (logreg) is used for binary variables, and predictive mean matching (PMM) is used for continuous variables; Set the number of imputations m = 5 to generate 5 imputed datasets; The imputation results are saved in the mice_result object, which can be used for subsequent model development and training.
[0011] Step2-3, use boxplot to identify potential outliers in continuous variables, and perform upper and lower 1% tail processing on abnormal data.
[0012] The corresponding R implementation is as follows: # Tail function definition winsorize <- function(x, lower = 0.01, upper = 0.99) { q <- quantile(x, probs = c(lower, upper), na.rm = TRUE) x[x < q[1]] <- q[1] x[x > q[2]] <- q[2] return(x) } Tail processing on variables other than "GCS" numeric_vars_to_winsor <- setdiff(numeric_vars, "GCS") complete_data[numeric_vars_to_winsor] <- lapply(complete_data[numeric_vars_to_winsor], winsorize) Explanation: 1. Input: each continuous variable x; Treatment: Calculate the 1st percentile value q 1% and the 99th percentile value q 99% ; Replace values less than q 1% with q 1% ; Replace values greater than q 99% with q 99% ; 2. Output: More robust continuous variable data after truncation.
[0013] Preferably, Step 3 randomly splits the data into a training set and a validation set in a ratio of 7:3.
[0014] Preferably, Step 4 includes the following steps: Step 4-1, the patient data comes from the training set; Step 4-2, apply the R package "glmnet" for LASSO regression to preliminarily screen the variables; at the same time, introduce L1 regularization and 10-fold cross-validation to select the best lambda (λ) value, and use lambda. min as the best λ value to screen out the variables with non-zero regression coefficients; The specific implementation is as follows: Step 4-2-1, train the LASSO regression model lasso_model <- glmnet(X, y, family = "binomial", alpha = 1) Explanation: X: Independent variable matrix (constructed by model.matrix(), intercept term removed); y: Response variable, 0 / 1 binary classification result; family = "binomial" indicates using a logistic regression model; alpha = 1 indicates using L1 regularization (LASSO).
[0015] Step 4-2-2, use 10-fold cross-validation to select the best lambda value cv_lasso <- cv.glmnet(X, y, family = "binomial", alpha = 1, nfolds =10) Explanation: Automatically divide the data into 10 subsets, and rotate them as the validation set to evaluate the model performance under different λ; Internal automatic computation of cross-validation error (e.g. mean squared error, MSE, or deviance).
[0016] Step4-2-3, Extract the optimal lambda value (lambda.min and lambda.1se) lambda_min <- cv_lasso$lambda.min lambda_1se <- cv_lasso$lambda.1se Explanation: lambda.min: the lambda with the minimum cross-validation error; lambda.1se: the lambda with the minimum error and a sparser model.
[0017] Step4-2-4, Extract the non-zero coefficients at lambda.min: coefs_min <- as.matrix(coef(cv_lasso, s = "lambda.min")) # Convert to a normal matrix selected_vars_min <- rownames(coefs_min)[coefs_min!= 0] # Non-zero variable names selected_coefs_min <- coefs_min[coefs_min!= 0, 1] # Non-zero variable coefficients Explanation: coef() extracts the coefficients of the model at the specified lambda value; Non-zero coefficients are the variables selected by LASSO into the model; The first line is the intercept, which can be excluded if not involved in variable selection.
[0018] Step4-2-5 Extract the number of non-zero coefficients at lambda.min nonzero_vars_min <- sum(coef(cv_lasso, s = "lambda.min")[-1, ]!= 0) print(nonzero_vars_min) Explanation: -1 excludes the intercept; The output is the actual number of variables in the model.
[0019] Step 4-3: Perform multivariate logistic regression analysis on the variables with non-zero regression coefficients to obtain independent risk factors for SCAP-related PE; Step 4-4, construct a prediction model that integrates the nonlinear effects of continuous variables.
[0020] Preferably, the variables with non-zero regression coefficients in Step 4-2 include: age, gender, smoking history, ICU length of stay, total hospital stay, body temperature, heart rate, creatinine, neutrophil-to-lymphocyte ratio, prothrombin time, activated partial thromboplastin time, lactate, oxygen partial pressure, sepsis, respiratory failure, chronic lung disease, severe liver disease, kidney disease, Glasgow Coma Scale, and use of mechanical ventilation.
[0021] Preferably, the independent risk factors described in Step 4-3 include: prolonged ICU stay, increased heart rate, female sex, no respiratory failure, no severe liver disease and no kidney disease.
[0022] Preferably, the specific process of Step 4-3 includes: Step 4-3-1. Perform multivariate logistic regression analysis on the variables with non-zero regression coefficients based on LASSO, and select all variables with p < 0.05 based on the multivariate regression results; The specific implementation is as follows: Construct a logistic regression model formula, where outcome_var is the dependent variable name and variables is the vector of independent variable names. formula <- as.formula(paste(outcome_var, "~", paste(variables,collapse = " + "))); Use the formula and training data to build the initial full model full_model <- glm(formula, data = training_data, family = binomial); A stepwise backward regression based on AIC was used to eliminate insignificant variables. final_model <- step(full_model, direction = "backward", trace = 0); illustrate: Dynamically concatenate model formulas through paste and as.formula to achieve flexible modeling of any set of variables; training_data contains the dependent variable and all independent variables filtered by LASSO; glm(..., family = binomial) builds a logistic regression model; The step() function uses the AIC indicator to perform stepwise variable screening and retain variables that contribute significantly to the model; trace = 0 turns off the step-by-step process log to keep the results concise; The final model final_model can be used for further analysis and risk assessment.
[0023] Step 4-3-2: Delete variables with p>0.05. The variables initially retained include age, ICU hospitalization time, heart rate, gender, respiratory failure, severe liver disease, and kidney disease; Step 4-3-3, generalized additive model (GAM) was used to test the nonlinear association between three continuous variables, namely age, ICU hospitalization time, and heart rate, and the risk of PE. Age and ICU hospitalization time showed significant nonlinear effects, while heart rate showed a linear association. The specific implementation is as follows: Install and load the required packages for GAM install.packages("mgcv") library(mgcv); Construct a GAM model and specify the s() function for nonlinear smooth fitting model_gam <- gam(outcome ~ s(age) + s(ICU_length_of_stay) + s(heart_rate) + sex, family = binomial, data = training_data); View model results summary(model_gam); Visualize the nonlinear relationship curve of each variable plot(model_gam, pages = 1); illustrate: outcome: whether PE occurred (1 / 0); s(age), s(ICU_length_of_stay), s(heart_rate): construct smooth functions for these three continuous variables; Sex: included as a linear covariate control; family = binomial indicates the use of a logistic regression framework; summary(model_gam) outputs the EDF (effective degrees of freedom) and p-value of each variable to evaluate the significance of nonlinearity; plot(model_gam) visually displays the functional form of the variable and PE risk (whether it is a straight line or a curve).
[0024] Step 4-3-4: Perform multivariate logistic regression again, using restricted cubic splines and RCS to perform nonlinear fitting on age and ICU length of stay.
[0025] The specific implementation is as follows: Load the rms package: library(rms); Construct a logistic regression model and use rcs() to perform nonlinear fitting on ICU_length_of_stay: logistic_model <- lrm(outcome ~ rcs(ICU_length_of_stay, 4) + Heart_rate + Sex + Respiratory_failure + Severe_liver_disease + Kidney_disease, data = training_data); Output model summary to view variable coefficients, p-values, nonlinear trends, etc. summary(logistic_model); illustrate: lrm(): logistic regression modeling function from the rms package, suitable for building stable medical regression models; rcs(variable, 4): performs a restricted cubic spline fit with 4 nodes on the variable. The node positions can be specified automatically or manually. outcome: whether PE occurred (1 / 0, dependent variable); Step 4-3-5: Delete the variable with p>0.05, namely age; finally determine that the independent risk factors for PE in ICU-SCAP patients include prolonged ICU stay, increased heart rate, female sex, no respiratory failure, no severe liver disease and no kidney disease; perform collinearity diagnosis on the above 6 independent risk factors, and the VIF of all variables are <5.0, indicating no collinearity.
[0026] Preferably, the verification method of Step 5 is internal verification, specifically: 1000 times Bootstrap method in intensive care medicine information market-IV (version 2.2) database of 944 patients for internal verification. That is, the area under the curve (AUC), calibration curve and decision curve analysis (DCA) are used to evaluate the discrimination, calibration and clinical benefit of the model.
[0027] Compared with the prior art, the beneficial effects of the present application are: Data quality is reliable: the present application uses Bayesian multiple imputation (mice) to systematically process missing values, identifies potential outliers using box plots, and uses upper and lower 1% Winsorization method to process extreme values, which significantly improves data integrity and robustness, effectively reduces the influence of noise interference on model performance.
[0028] High efficiency of variable screening: the present application introduces LASSO regression, combining L1 regularization and 10-fold cross-validation, which can quickly identify variables with actual predictive value for PE risk in high-dimensional features, avoid multicollinearity, simplify model structure, and enhance model generalization ability.
[0029] Rational processing of nonlinear relationship: for the possible nonlinear relationship between some continuous variables and the outcome variable, the present application introduces generalized additive model (GAM) and restricted cubic spline (RCS) modeling method, which improves the fitting performance of the model for complex clinical variables, and can more accurately reflect the real influence trend of the outcome variable on PE risk.
[0030] Strong explanation and high applicability: the final nomogram tool can realize individualized prediction and early intervention of PE risk in SCAP patients in ICU clinical environment.
[0031] Effectively solve the problems of the prior art: compared with the traditional prediction method based on common variables or simple logic modeling, the present application solves the problems of single variable screening method, weak nonlinear processing ability and rough data preprocessing. The proposed modeling process is clear, reproducible, and has good universality and generalizability.
[0032] In summary, based on the systematic data processing strategy, combined with LASSO, GAM, RCS and other advanced modeling methods, the present application realizes the accurate evaluation of PE risk of ICU-SCAP patients, and has good technical advancement, stability and clinical practical value. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1The present invention provides a flow chart for establishing a risk prediction model for ICU-SCAP-related PE patients. Figure 2 As: a horizontal bar chart of the specific missing ratio of the initial variables of the present invention; Figure 3 : Box plot of potential outliers in the initial continuous variables of the present invention; Figure 4 Schematic diagram of the process of using LASSO regression model to screen variables in the present invention; Figure 5 The present invention provides a nomogram model for predicting the risk of PE in ICU-SCAP patients. Figure 6 : ROC curve diagram of the training set and the validation set of the present invention; Figure 7 : calibration curve diagrams of the training set (a) and the validation set (b) of the present invention; Figure 8 : DCA diagram of the training set and validation set of the present invention. DETAILED DESCRIPTION
[0034] The present invention is described in further detail below in conjunction with specific embodiments.
[0035] Example 1: As shown in the attached Figure 1 A risk prediction model for ICU-SCAP-related PE patients and a method for establishing the same include the following steps: Step 1. Determine the data source and research population. The data sources are as follows: Based on the inclusion and exclusion criteria, we retrospectively analyzed clinical data from 3148 ICU-SCAP patients from the Medical Information Marketplace for Intensive Care Medicine IV (MIMIC-IV 2.2). MIMIC-IV is a large, freely available, public database containing comprehensive, reliable medical records of 73,181 patients admitted to the ICU of Beth Israel Deaconess Medical Center from 2008 to 2019. The researchers participated in a series of online training courses offered by the National Institutes of Health, signed a data use agreement, and obtained access to the MIMIC-IV database (certificate number: 59283397) after passing the course evaluation. All patient identity information in the database was de-identified. Furthermore, the database was approved by the Institutional Review Boards of the Massachusetts Institute of Technology and Beth Israel Deaconess Medical Center at its inception. Therefore, no ethical review was required. All methods were performed in accordance with relevant guidelines and regulations, including the Declaration of Helsinki and the policies of the PhysioNet platform.
[0036] The inclusion and exclusion criteria for the study population are as follows: Inclusion criteria: (1) Patients diagnosed as SCAP in ICU, i.e. patients with community-acquired pneumonia who need tracheal intubation or septic shock after active fluid resuscitation still need treatment of vasoactive drugs, the International Classification of Diseases Ninth Clinical Revision (ICD-9-CM) code is: 4800, 4801, 4802, 4803, 4809, 481, 4822, 48231, 48283, 48284, 4830, 4831, 486; the International Classification of Diseases Tenth Clinical Revision (ICD-10-CM) code is: J120, J121, J122, J1281, J129, J13, J14, J156, J157, J160, J189; (2) Patients with community-acquired pneumonia aged 18-89 years old; (3) Patients with multiple ICU admissions, the first ICU admission data were selected; (4) The first ICU hospitalization time was more than 24 hours.
[0037] Exclusion criteria: (1) Non-first ICU admission records; (2) Patients aged <18 or >89 years old; (3) Death or discharge within 24 hours after the first ICU admission; (4) Pregnant and lactating women. In addition, after multiple tests, we excluded data with total hospitalization time shorter than ICU hospitalization time to improve data reliability.
[0038] Step 1-1, extract the clinical data required by Step 1; Using subject_id, hadm_id, icustay_id of SCAP patients, the following information variables were extracted from the database: age, gender, smoking history, ICU stay time, total hospital stay time, vital signs, first laboratory test results, comorbidities, vasopressor use, mechanical ventilation, pathogens and disease severity scoring system. The vital sign variables are the average values of the data recorded at the first time of entering the ICU, including body temperature, respiratory rate, heart rate, systolic blood pressure, diastolic blood pressure, mean arterial pressure (MAP), arterial oxygen saturation (SpO2). The first laboratory test results include serum albumin, creatinine, urea nitrogen, urea nitrogen to creatinine ratio (UCR); white blood cells, neutrophils, lymphocytes, neutrophil to lymphocyte ratio (NLR), platelets, C-reactive protein (CRP); international normalized ratio (INR), prothrombin time (PT), activated partial thromboplastin time (APTT), fibrinogen (Fib), D-dimer; N-terminal brain natriuretic peptide precursor (NT-ProBNP); lactic acid, bicarbonate, partial pressure of carbon dioxide (PCO2), oxygen partial pressure. Comorbidities include sepsis, respiratory failure, congestive heart failure (CHF), diabetes, hypertension, chronic lung disease, liver disease, kidney disease, and malignant tumor. The use of vasopressors is defined as: using vasopressors for any reason during ICU stay. Pathogens include: Streptococcus pneumoniae, group A streptococcus, viruses (adenovirus, respiratory syncytial virus, parainfluenza virus, coronavirus), Haemophilus influenzae, Mycoplasma pneumoniae, chlamydia, legionella, other gram-negative bacteria. Severity scoring systems include systemic inflammatory response syndrome score (SIRS), simplified acute physiology score II (SAPS II), Glasgow coma score (GCS) and sequential organ failure assessment score (SOFA).
[0039] Step 2, data cleaning and processing of missing values and outliers from the data obtained in Step 1-1; The data cleaning and processing steps are as follows: (1) Clean and merge the data, delete the data in the data set extracted in Step 1-1 that contains content errors, format errors, logical errors, and data unrelated to the study; (2) Use the R software package "ggplot2" to draw a specific missing proportion level bar chart of the variable (see Appendix Figure 2 ), and then remove variables with a missing rate of >60%, and then use the R software package "mice" to perform Bayesian multiple imputation to handle missing values; (3) Use the box plot to identify potential outliers in continuous variables (see Appendix Figure 3 ), and perform 1% tail processing on abnormal data.
[0040] Step 3, randomly split the data set into training set and validation set in the ratio of 7:3; Step 3-1, population and baseline characteristics analysis of the pre-processed and split data set; the training set (2204 cases) is used to construct the nomogram, and the validation set (944 cases) is used to evaluate the model performance. In the standard analysis, bilateral p<0.05 is considered statistically significant. We use the Shapiro-Wilk normality test and Bartlett's variance homogeneity test to test the data. The measurement data that meet the normal distribution and variance homogeneity are expressed as mean ± standard deviation (SD), and t-test is used for inter-group comparison; non-normal distribution or variance data are expressed as median (interquartile range, IQR), and Mann-Whitney U rank sum test is used to analyze the differences between groups. The frequency (n) and percentage (%) of categorical variables are expressed, and the chi-square test or Fisher's exact test is used to compare the differences between groups. The analysis results show that the median age of the overall population is 66.5 years (IQR 55.0, 76.0), and males account for 58.2%. Among them, 152 cases (4.8%) of SCAP patients occurred PE during ICU hospitalization, and in the training cohort and validation cohort, there were 110 cases (5.0%) and 42 cases (4.4%) respectively. There was no significant difference in the main baseline characteristics of the two data sets (P>0.05), indicating that the data was randomly assigned and balanced. The relationship between the patient's population and baseline characteristics is shown in Table 1.
[0041] Table 1. Demographic and clinical characteristics of the training and validation sets
[0042]
[0043] Note: Continuous variables are presented as medians (interquartile range, IQR), and categorical variables are presented as frequencies (percentages (%)). Abbreviations: MAP, mean arterial pressure; SpO2, arterial oxygen saturation; UCR, blood urea nitrogen-to-serum creatinine ratio; NLR, neutrophil-to-lymphocyte ratio; INR, international normalized ratio; PT, prothrombin time; APTT, activated partial thromboplastin time; PCO2, partial pressure of carbon dioxide; CHF, congestive heart failure; SIRS, systemic inflammatory response syndrome score; SAPSII, simplified acute physiology score II; GCS, Glasgow Coma Scale; SOFA, sequential organ failure assessment score.
[0044] Step 4: LASSO regression and multivariate logistic regression were used to determine the independent risk factors for ICU-SCAP-related PE, and a risk prediction model was constructed based on the training set.
[0045] This example uses the R software package "glmnet" to perform LASSO regression and conduct preliminary screening of predictor variables (see Appendix Figure 4 ). The optimal lambda (λ) value was selected by introducing L1 regularization and 10-fold cross validation to improve the predictive performance and interpretability of the model. In this example, lambda.min was selected as the optimal λ value for variable screening. The variables retained by LASSO regression included age, gender, smoking history, ICU length of stay, total hospital stay, body temperature, heart rate, creatinine, NLR, PT, APTT, lactate, oxygen partial pressure, sepsis, respiratory failure, chronic lung disease, severe liver disease, kidney disease, GCS, and use of mechanical ventilation. Subsequently, multivariate logistic regression was used to further screen variables, and the model was optimized in combination with the Akaike Information Criterion (AIC). The results showed that age, ICU length of stay, heart rate, gender, respiratory failure, severe liver disease, and kidney disease were independent influencing factors for the occurrence of PE in SCAP patients in the ICU (p<0.05). The OR values and 95% CI values are shown in Table 2.
[0046] Table 2 Multivariate logistic regression analysis before RCS adjustment
[0047] Note: *P<0.05, *P<0.001. Further, considering the assumption of linear relationship of continuous variables in traditional regression models, GAM was used to test the non-linear association of age, ICU length of stay and heart rate with the risk of PE. The results showed that age (estimated degrees of freedom edf = 2.555, p = 0.019) and ICU length of stay (edf = 5.481, p = 0.006) had significant non-linear effects, while heart rate (edf = 1.000, p = 0.011) was consistent with linear relationship.
[0048] Further, in the final model, RCS was used to non-linearly fit age and ICU length of stay, and the node positions were set to the 5th, 35th, 65th and 95th percentiles based on clinical significance. Model comparison showed that the goodness of fit was significantly improved after introducing RCS (likelihood ratio test: χ²= 16.4, p = 0.003). The independent risk factors for ICU-SCAP patients complicated with PE included prolonged ICU length of stay, increased heart rate, female, no respiratory failure, no severe liver disease and no kidney disease, and the OR values and 95% CI values are shown in Table 3.
[0049] Table 3 Multivariate logistic regression analysis after RCS adjustment
[0050] Note: *P <0.05, **P <0.001. The collinearity diagnosis was performed on the 6 variables of prolonged ICU length of stay, increased heart rate, female, no respiratory failure, no severe liver disease and no kidney disease, and the variance inflation factor (VIF) of all variables was <5.0, indicating no significant multiple collinearity problem. Based on the optimized final model, a Nomogram integrating the non-linear effects of continuous variables was constructed to evaluate the individualized risk of PE in ICU-SCAP patients (see Appendix Figure 5 ).
[0051] Step 5, internal validation of the prediction model based on the validation set to evaluate its discrimination, calibration and clinical benefit.
[0052] Internal validation was performed on 944 patients in the Critical-Care Medicine Information Market-IV (version 2.2) database through 1000 times of Bootstrap method. The ROC curves of the training cohort and the validation cohort are shown in Appendix Figure 6 . The AUC of the training cohort was 0.702 (95% CI = 0.655-0.749), and the AUC of the validation cohort was 0.690 (95% CI = 0.611-0.769), indicating that the model had certain discrimination ability. When the predicted probability was ≥0.060, it was suggested that the patient might have PE, and it was recommended to combine clinical practice for comprehensive evaluation and intervention.Figure 7 In the middle, the calibration chart shows a high consistency, and the Hosmer-Lemeshow goodness-of-fit test results have no significant statistical significance (training cohort: χ²= 7.184, p=0.517, validation cohort: χ²= 6.775, p= 0.561), indicating that the predicted probability is close to the actual probability, and the model fits well. The DCA compares the net benefit of the model under different thresholds, and the results show that the model can obtain higher net benefit within the risk threshold of 0.01-0.20, and has better clinical application value (see attached Figure 8 )。
[0053] This embodiment constructs and verifies the first PE risk prediction model for ICU-SCAP patients, providing a quantitative tool for early identification of high-risk patients with PE related to ICU-SCAP in clinical practice. After more external verification, the model is expected to be widely used in clinical practice.
[0054] The above is only part of the specific embodiments of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A pulmonary embolism risk prediction model and its establishment method for patients with severe community-acquired pneumonia, characterized by The following steps are involved: Step 1. Enter patient data; Step 2. Organize and clean the data, especially deal with missing values and outliers; Step 3. Randomly split the data into training and validation sets in proportion, and analyze the relationship between demographic and baseline characteristics of the split datasets; Step 4. LASSO regression and multivariate logistic regression were used to identify independent risk factors for pulmonary embolism (PE) in ICU-SCAP patients, and a risk prediction model was constructed based on the training set. Step 5. Perform internal validation of the risk prediction model based on the validation set to evaluate its discrimination, calibration, and clinical benefit.
2. The pulmonary embolism risk prediction model and establishment method for patients with severe community-acquired pneumonia according to claim 1, characterized in that: The inclusion criteria for the patient data described in Step 1 were as follows: patients diagnosed with SCAP in the ICU, i.e., community-acquired pneumonia requiring endotracheal intubation or septic shock requiring vasoactive drug treatment after active fluid resuscitation, patients aged ≥18 years and ≤89 years, with multiple ICU admissions, and patients with the first ICU admission and the first ICU stay of more than 24 hours were selected; Exclusion criteria for patient data were: non-first ICU hospitalization record, patients aged <18 or >89 years, death or discharge within 24 h after first ICU admission, pregnant or breastfeeding women, or total hospitalization duration shorter than the ICU stay duration.
3. The pulmonary embolism risk prediction model and establishment method for patients with severe community-acquired pneumonia according to claim 1, characterized in that: The data organization and cleaning process described in Step 2 includes: Step 2-1. Delete data with content errors, format errors, logical errors, and data irrelevant to the research in Step 1; Step 2-2: Use the R package "ggplot2" to draw a horizontal bar chart of the specific missing ratio of the variables, then remove variables with a missing rate greater than 60%, and then use the R package "mice" to perform Bayesian multiple imputation to handle missing values; Step 2-3: Use box plots to identify potential outliers in continuous variables and perform winsoring on the abnormal data by 1% above and below.
4. The pulmonary embolism risk prediction model and establishment method for patients with severe community-acquired pneumonia according to claim 1, characterized in that: As described in Step 3, the data is randomly split into a training set and a validation set in a ratio of 7:
3.
5. The pulmonary embolism risk prediction model and establishment method for patients with severe community-acquired pneumonia according to claim 1, characterized in that: Step 4 includes: Step 4-1, patient data comes from the training set; Step 4-2: Use the R package "glmnet" to perform LASSO regression and conduct preliminary variable screening. At the same time, introduce L1 regularization and 10-fold cross-validation to select the optimal lambda (λ) value. Use lambda.min as the optimal λ value to screen out variables with non-zero regression coefficients. Step 4-3: Perform multivariate logistic regression analysis on variables with non-zero regression coefficients to obtain independent risk factors for the occurrence of severe community-acquired pneumonia-related pulmonary embolism; Step 4-4, construct a prediction model that integrates the nonlinear effects of continuous variables.
6. The pulmonary embolism risk prediction model and establishment method for patients with severe community-acquired pneumonia according to claim 5, characterized in that: The variables with non-zero regression coefficients described in Step 4-2 include: age, gender, smoking history, ICU length of stay, total hospital stay, body temperature, heart rate, creatinine, neutrophil-to-lymphocyte ratio, prothrombin time, activated partial thromboplastin time, lactate, oxygen partial pressure, sepsis, respiratory failure, chronic lung disease, severe liver disease, kidney disease, Glasgow Coma Scale, and use of mechanical ventilation.
7. The pulmonary embolism risk prediction model and establishment method for patients with severe community-acquired pneumonia according to claim 5, characterized in that: The independent risk factors described in Step 4-3 include: prolonged ICU stay, increased heart rate, female sex, absence of respiratory failure, absence of severe liver disease, and absence of kidney disease.
8. The pulmonary embolism risk prediction model and establishment method for patients with severe community-acquired pneumonia according to claim 5, characterized in that: The specific process of Step 4-3 includes: Step 4-3-1. Perform multivariate logistic regression analysis on the variables with non-zero regression coefficients based on LASSO, and select all variables with p < 0.05 based on the multivariate regression results; Step 4-3-2: Delete variables with p>0.
05. The variables initially retained include age, ICU hospitalization time, heart rate, gender, respiratory failure, severe liver disease, and kidney disease; Step 4-3-3, generalized additive model (GAM) was used to test the nonlinear association between three continuous variables, namely age, ICU hospitalization time, and heart rate, and the risk of pulmonary embolism. Age and ICU hospitalization time showed significant nonlinear effects, while heart rate showed a linear association. Step 4-3-4: Perform multivariate logistic regression again, using restricted cubic splines and RCS for nonlinear fitting of age and ICU length of stay; Step 4-3-5. Delete the variable with p>0.05, namely age; finally determine that the independent risk factors for pulmonary embolism in patients with severe community-acquired pneumonia in the ICU include prolonged ICU hospitalization, increased heart rate, female sex, no respiratory failure, no severe liver disease and no kidney disease; perform collinearity diagnosis on the above 6 independent risk factors, and the VIF of all variables are <5.0, indicating no collinearity.
9. The pulmonary embolism risk prediction model and establishment method for patients with severe community-acquired pneumonia according to claim 1, characterized in that: The verification method described in Step 5 is internal verification.