Mycoplasma pneumoniae pneumonia severity prediction model and construction method thereof

By constructing machine learning prediction models for protein content of HMGB1, S100A8 and S100A9, the problem of early identification of Mycoplasma pneumonia in children in the prior art is solved, and efficient and accurate severity prediction and risk stratification are achieved to assist clinical decision-making.

CN120376161APending Publication Date: 2025-07-25WUHAN CHILDRENS HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510489700.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

There is a lack of effective tools in the prior art to identify the severity of Mycoplasma pneumonia in children in early stages. The applicability and accuracy of the existing evaluation system to children is limited, and the specificity and sensitivity of traditional laboratory indicators are insufficient, making it difficult to meet the clinical needs of early precision intervention.

Method used

A machine learning prediction model based on the protein content of HMGB1, S100A8 and S100A9 was constructed. Combined with clinical characteristics and conventional laboratory detection indicators, a machine learning algorithm was used to establish a prediction model suitable for the severity of Mycoplasma pneumonia in children.

Benefits of technology

It provides efficient and accurate decision-making support tools that can identify the minor and severe cases of Mycoplasma pneumonia pneumonia in the early stage, improve the accuracy of risk stratification, assist clinical decision-making, reduce non-essential ICU transfers, guide treatment choices and reduce mortality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376161A_ABST
    Figure CN120376161A_ABST
Patent Text Reader

Abstract

The invention discloses a model for predicting the severity of mycoplasma pneumonia and a construction method thereof, and the method comprises the following steps: acquiring a serum sample of a mycoplasma pneumonia patient, establishing a data set by using the content of high mobility group B (High Mobility Group Box 1, HMGB1), S100A8 and S100A9 proteins in the serum sample and the clinical laboratory detection data of the patient, and predicting the severity of mycoplasma pneumonia by using a light gradient boosting machine (Light GBM), XGBoost, Logistic and random forest RF). According to the method, nine machine learning models, namely, a K-nearest neighbor classification (KNN), a support vector machine (SVM), a Gaussian # imgabs0 # bayes, a decision tree (DT) and a Catboost, are respectively constructed, the efficiency of the models is evaluated, the optimal model is screened, the screened model is explained, and the contribution degree of each index in the models is evaluated. The prediction model and method provided by the invention have good prediction performance and good clinical applicability, and can effectively improve the accuracy of risk stratification of mycoplasma pneumonia patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to the technical field of medical disease prediction, and relates to a method for constructing a pneumonia severity prediction model and a prediction model, and particularly relates to a method for constructing a Mycoplasma pneumoniae pneumonia severity prediction model based on HMGB1, S100A8, and S100A9. Background Art:

[0002] Mycoplasma pneumoniae pneumonia (MPP) is an acute respiratory infectious disease caused by Mycoplasma pneumoniae. According to the global epidemiological data in 2021, the annual number of MPP cases is approximately 25.3 million, and the incidence rate in school-age children (5 - 14 years old) is significantly higher than that in other age groups. The national surveillance data in China from 2009 to 2019 also shows that the detection rate of MP in acute respiratory infection patients among school-age children is 18.6%, while in pneumonia patients, the MP infection rate is as high as 49.9%, which is much higher than that in non-pneumonia patients (19.1%) of the same age group. This data fully demonstrates that MPP has become one of the main bacterial causes of respiratory infections in school-age children. In the past, it was believed that some children infected with MP had a certain degree of self-healing ability and could recover even without treatment, and there were no complications or sequelae. However, due to factors such as the irregular use of antibacterial drugs, Mycoplasma colonization, drug resistance and increased pathogenicity, and excessive immune responses, some children develop severe diseases. Even if macrolide drugs are used early and in sufficient doses for symptomatic treatment, the conditions of some children will still progress. Some severely ill children will be accompanied by respiratory distress syndrome, with intrapulmonary and extrapulmonary complications, and may eventually have sequelae, resulting in a certain proportion of disability and fatality rates. The surveillance data after the COVID-19 pandemic shows that the incidence rate of Severe Mycoplasma pneumoniae pneumonia (SMPP) has shown an abnormal upward trend (exceeding 20% at one time). The statistical data of Hebei Children's Hospital shows that this proportion has even reached 50%, bringing a heavy burden to patients and society. Therefore, early and accurate identification of severe MPP and taking targeted intervention measures are the keys to improving the prognosis of children.

[0003] Current clinical practice faces two major challenges. First, existing evaluation systems such as the Clinical Pulmonary Infection Score (CPIS) are mainly designed based on adult respiratory physiological parameters, and there are significant limitations in their applicability to pediatric patients. Although Chinese Patent CN112183572A constructs a pneumonia severity prediction model with low model complexity and high prediction accuracy based on clinical symptoms and imaging features, due to factors such as early non-obvious symptoms, non-specific laboratory indicators, and the difficulty of performing lung CT examinations on critically ill patients, its applicability and accuracy in the pediatric population are limited. Existing studies have proposed independent risk factors for SMPP (such as persistent high fever for more than 10 days, pleural effusion, extrapulmonary complications, CRP > 40 mg / L, etc.), which mostly focus on the manifestations in the middle and late stages of the disease, with a lagging prediction window period and difficulty in meeting the clinical needs of early precise intervention. Second, traditional laboratory indicators used to evaluate the severity of infectious diseases include blood routine (such as white blood cell count, neutrophil ratio), CRP, procalcitonin, interleukin 6, serum amyloid A (SAA), ferritin, etc., which can indicate the type or severity of infection, but the specificity and sensitivity of single indicators or multi-factor analysis are insufficient, and there is a lack of prospective validation, making it difficult to distinguish mild and severe MPP in a timely and effective manner.

[0004] Recent studies have shown that the high or low pathogen load, virulence, and imbalance of the host immune state may be important reasons for the occurrence and development of SMPP. Therefore, dynamically monitoring the body's immune-inflammatory indicators and finding specific circulating markers are of great significance for evaluating the severity of MPP. Currently, a large number of studies focus on exploring the immune response of the host after infection, and a series of damage-associated molecular patterns (DAMPs) that activate the innate immune response and cytokines related to "cytokine storm" have been discovered. Among them, HMGB1, as a nuclear protein, is released extracellularly during infection and inflammation, participates in regulating multiple inflammatory signaling pathways, and its level is closely related to the severity and prognosis of pneumonia. In addition, S100A8 and S100A9, members of the S100 calcium-regulatory protein family (also known as MRP8 and MRP14), as important DAMPs, are mainly secreted by neutrophils, monocytes, and macrophages, and regulate the inflammatory response and immune cell migration by binding to Toll-like receptors (TLRs) and receptor for advanced glycation end products (RAGE). Studies have shown that S100A8 / A9 is significantly elevated in various inflammatory diseases (such as infection, inflammatory bowel disease, rheumatoid arthritis, etc.) and is positively correlated with the disease severity. However, current studies on HMGB1, S100A8, and S100A9 in pediatric MPP and their expression differences and combined prediction value in mild and severe MPP have not been fully explored.

[0005] Based on the above background, the present invention innovatively constructs a combined prediction model of HMGB1, S100A8, and S100A9, combines clinical features and conventional laboratory test indicators, and uses machine learning algorithms to establish a model suitable for predicting the severity of Mycoplasma pneumoniae pneumonia (MPP) in children, so as to solve the problems of lack of prediction tools and difficulty in early identification in the prior art, and provide an efficient and accurate decision-making support tool for clinical practice. Summary of the Invention:

[0006] (1) Technical problems to be solved

[0007] In view of the above background, the present invention provides a prediction model for the severity of Mycoplasma pneumoniae pneumonia and a method for constructing the same. The prediction model has good prediction performance, good clinical applicability, and can improve the accuracy of risk stratification of patients with Mycoplasma pneumoniae pneumonia, and solves the problems of lack of prediction tools and difficulty in early identification in the prior art.

[0008] (2) Technical solutions

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] The application of the protein contents of HMGB1, S100A8, and S100A9 in constructing a prediction model for the severity of pneumonia, exemplarily, in constructing a prediction model for the severity of Mycoplasma pneumoniae pneumonia, preferably, in constructing a prediction model for the severity of Mycoplasma pneumoniae pneumonia in children.

[0011] The technical solution adopted by the present invention is: collecting blood samples from patients with Mycoplasma pneumoniae pneumonia of different severities, quantitatively detecting the protein contents of HMGB1, S100A8, and S100A9 and laboratory test results as input features, and using a machine learning classifier as a prediction model to predict the severity of patients with Mycoplasma pneumoniae pneumonia.

[0012] The present invention provides a method for constructing a prediction model for the severity of Mycoplasma pneumoniae pneumonia, comprising the following steps:

[0013] Step 1: Data collection, processing, and data establishment

[0014] Serum samples of patients with Mycoplasma pneumoniae pneumonia were collected according to inclusion and exclusion criteria. The contents of HMGB1, S100A8, and S100A9 proteins in the serum samples were quantitatively detected by ELISA, and the clinical baseline data of the patients were obtained. The clinical baseline data included the patients' symptoms, signs, laboratory test results, and imaging examination results, specifically including: gender, age, body temperature at admission, respiratory rate at admission, blood routine at admission, high-sensitivity C-reactive protein, procalcitonin, ferritin, IL-6, blood coagulation, liver and kidney function electrolytes, myocardial enzyme spectrum, blood immune indexes, coagulation profile, blood immune indexes, TBNK lymphocyte subsets.

[0015] Specific laboratory test result indicators included inflammatory indicators [hsCRP, procalcitonin (PCT)], ferritin, IL-6, blood routine indicators [white blood cells (WBC), neutrophil percentage (NEUT%), lymphocyte percentage (LYMPH%), red blood cells (RBC), hemoglobin (Hgb), platelets (PLT)], liver function indicators [aspartate aminotransferase (AST), alanine aminotransferase (ALT), gamma-glutamyl transpeptidase (gamma-GT), albumin (ALB), globulin (GLB), total bilirubin (TBIL), direct bilirubin (DBIL), alkaline phosphatase (ALP)]; myocardial enzyme markers [lactate dehydrogenase (LDH), creatine kinase (CK), creatine phosphokinase isoenzyme (CK-MB)]; coagulation function markers [D-Dimer, fibrinogen, antithrombin, prothrombin time, partial thromboplastin time, thrombin clotting time (TT)] and renal function indicators (blood urea nitrogen, creatinine, cystatin C), serum electrolytes; humoral immune indicators (IgA, IgG, IgM, C3, C4) and cellular immune function indicators (TBNK lymphocyte subsets);

[0016] According to the preset severe classification criteria, the cases were divided into mild and severe datasets, and the data were normalized; they were divided into a training set and a test set in a ratio of 7:3; the training set included a mild training set and a severe training set, and the test set included a mild test set and a severe test set;

[0017] Step 2. Variable screening:

[0018] First, use the Boruta algorithm to screen variables for the training set in Step 1. (1) Generate a copy for each original feature in the feature matrix and splice all the copies to the feature matrix; (2) Randomly shuffle the copies of all features, that is, randomly scramble the feature values to generate a shadow feature. There is no correlation between the shadow feature and the outcome. The shadow feature and the original feature form a new feature matrix; (3) Run a random forest on the new feature matrix and collect the Z-score of each feature; (4) Find the maximum Z-score among the shadow features (Maximum Z-score Among Shadow Attributes, MZSA). Then, mark each original feature with a Z-score higher than MZSA as "hit" or "selected"; (5) Repeat the above steps N times. Then, for each original feature whose importance is not yet determined, based on the binomial distribution B(N, 0.5), conduct a two-sided hypothesis test on the number of times the original feature is "selected" to determine its importance; (6) According to the results of the hypothesis test, consider the features with significantly lower importance than MZSA as "unimportant" and permanently delete them from the feature set; (7) According to the results of the hypothesis test, consider the features with significantly higher importance than MZSA as "important"; (8) Delete all shadow features; (9) Repeat all the above processes until the importance is assigned to all features or the algorithm reaches the upper limit of the preset number of loops. The box plot shown by the Boruta algorithm is a box plot showing the importance of each feature variable during the loop process. In this study, the maximum number of loops is set to 100. Finally, plot the change in the importance scores of each variable during the running of the Boruta algorithm and select the top 12 variables with determined differences.

[0019] Then, use the Recursive Feature Elimination (RFE) method to screen variables for the training set in Step 1. Start searching for a feature subset from all the features in the training dataset and successfully delete features until the required number is retained. In this study, by fitting the random forest algorithm, based on all the initially included variables, repeatedly construct models and select the best features, and then repeat this process on the remaining features in turn until all features are traversed.

[0020] Finally, select the 10 key variables commonly identified by the Boruta and RFE algorithms as the model inputs. The key variables are: HMGB1, S100A8, S100A9, IL-6, ferritin, procalcitonin (PCT), lactic dehydrogenase (LDH), thrombin coagulation time (TT), CD16+CD56+ cells, and D-dimer. Evaluate the independent prediction efficacy of each screened variable by plotting the univariate receiver operating characteristic curve (ROC), and further select and optimize the indicators.

[0021] Step 3: Model construction and efficacy comparison:

[0022] Use the obtained training set as the training input, and input the training samples into 9 machine learning models: Lightgbm, XGBoost, Logistic, RF, KNN, SVM, GNB, DT, and Catboost. The parameter settings of the classification models are as follows:

[0023] Optimal hyperparameters of the LightGBM model: n_estimators = 70, learning_rate = 0.4, max_depth = 13, number_leaves = 15;

[0024] Optimal hyperparameters of the XGBoost model: max_depth = 3, random_state = 42, objective = reg:squarederror, eval_metric = rmse, min_child_weight = 10, gamma = 0.1, reg_alpha = 0.5, reg_lambda = 1.0;

[0025] Optimal hyperparameters of the Logistic model: defplot_decision_boundary(x,y,model):x_min,x_max = X[:,0].min()-1,X[:,0].max()+1,y_min,y_max =

[0026] X[:,1].min()-1,X[:,1].max()+1; xx,yy = np.meshgrid(np.linspace(x_min,

[0027] x_max, 100), np.linspace(y_min, y_max, 100)), Z =

[0028] model.predict(np.c_[xx.ravel(), yy.ravel()]), Z = Z.reshape(xx.shape);

[0029] Optimal hyperparameters of the RF model: ScoreAll = [] for i in range(100, 300, 10): model =

[0030] RandomForestClassifier(max_features = 2, n_estimators = i, random_state = 10) model.fit(X_train, y_train) score = model.score(X_test, y_test)

[0031] ScoreAll.append([i, score]) ScoreAll = np.array(ScoreAll);

[0032] Optimal hyperparameters of the KNN model: for k in range(1, 10): s_knn = S_KNN(k)

[0033] s_knn.fit(X_train, y_train) score = s_knn.score(X_test, y_test) k_list.append(k) score_list.append(score) if score > best_score: best_score = score best_k = k;

[0034] Optimal hyperparameters of the SVM model:

[0035] mp.tune <- tune.svm(Degree ~., data = train, kernal = "linear", cost = c(0.001, 0.01, 0.1, 1, 5, 10, 15)), bestmp <- mp.tune$best.model;

[0036] Optimal hyperparameters of the GNB model: def bayesian_optimization(X_train, y_train, X_test): gpr.fit(X_train, y_train) y_pred, sigma = gpr.predict(X_test, return_std = True) gap = y_pred - f(X_test) improvement = (gap + np.sqrt(sigma**2 + 1e-6) * np.abs(gap).mean())

[0037] *0.5 result = minimize(gpr.kernel_, np.zeros(gpr.kernel_.shape[0]),

[0038] method='L-BFGS-B') hyperparameters = result.x return hyperparameters,

[0039] improvement.max() n_iter = 20, n_samples = 5, results = [],

[0040] for i in range(n_iter): X_train = np.random.uniform(-2 * np.pi, 2 * np.pi,

[0041] (n_samples, 1)) y_train = f(X_train), result = bayesian_optimization(X_train, y_train, X_test = np.random.uniform(-2 * np.pi, 2 * np.pi,

[0042] (100, 1))) results.append(result);

[0043] Optimal hyperparameters of the DT model: learning_rate = 0.6, gamma = 0.9, max_depth = 3, subsample = 0.799, min_child_weight = 1, n_estimators = 2000;

[0044] Optimal hyperparameters of the Catboost model: iterations = 100, learning rate = 0.1, random state = 42, verbose = True, silent fit = True;

[0045] The model is trained using the optimized hyperparameters, with the 10 key variables in Step 2 input, and the classification results of mild or severe cases of Mycoplasma pneumoniae pneumonia are output; and the performance of the model in the training set and the test set is evaluated, and the prediction performance of the model is verified according to the accuracy, F1 score, Kappa value, and the area under the receiver operator characteristic curve (AUC) index, and a prediction model for the severity of Mycoplasma pneumoniae pneumonia with better performance is selected.

[0046] After performance evaluation, the random forest (RF) model showed the best prediction performance in the test set, and its specific indicators were: accuracy of 0.90, Kappa value of 0.80, and area under the curve AUC of 0.90 (95% CI 0.82 - 0.98). And 10-fold cross-validation and Bootstrap resampling were used to optimize the model, and the optimal prediction model was constructed by adjusting the parameters: (train.control <- trainControl(method = "repeatedcv", number = 10, repeats = 500)).

[0047] The calibration curve of the RF model for the training set and the test set was used to evaluate the consistency between the predicted severe risk of MPP and the actual empirical observation results. A confusion matrix for the prediction model of the severity of Mycoplasma pneumoniae pneumonia was constructed to show the corresponding relationship between the model prediction results and the actual situation distribution, and the evaluation of the prediction model of the severity of Mycoplasma pneumoniae pneumonia was completed. The metric index formula for the confusion matrix is:

[0048] Confusion Matrix Predicted Severe Predicted Mild Actual Severe TP (True Positive) FN (False Negative) Actual Mild FP (False Positive) TN (True Negative)

[0049]

[0050] Among them, TP represents predicted severe and actually severe; FP represents predicted severe and actually mild; FN represents predicted mild and actually severe; TN represents predicted mild and actually mild;

[0051] Step Four: Model Interpretation:

[0052] To improve the interpretability of the classifier model, this study used SHAP (Shapley Additive Explanation) to analyze the model features. SHAP values are a method used in machine learning to explain model decisions. It assigns weights by judging the importance of model features, helping to understand the important features of the model. SHAP was used to interpret the selected model and evaluate the contribution of each indicator in the model. Specifically: HMGB1 = 0.1956, S100A8 = 0.0701, S100A9 = 0.0672, CD16+CD56+ cells = 0.0557, PCT = 0.0419, Ferritin = 0.0391, TT = 0.0281, LDH = 0.0216, IL-6 = 0.0150, D-Dimer = 0.0115.

[0053] The SHAP dependence graph was used to quantitatively visualize the relationship between the main risk factors and the outcome, showing the top 4 most important features (HMGB1, S100A8, S100A9, CD16+CD56+ cells). It was also possible to determine the cut-off values of each variable to distinguish high-risk (SHAP value > 0) and low-risk (SHAP value < 0) of severe MPP.

[0054] The present invention also provides the application of the Mycoplasma pneumoniae pneumonia severity prediction model in predicting the severity of Mycoplasma pneumoniae pneumonia, especially in predicting the severity of Mycoplasma pneumoniae pneumonia in children.

[0055] Furthermore, the application includes predicting the severity of Mycoplasma pneumoniae pneumonia, especially in children, by detecting the protein expression levels of HMGB1, S100A8, and S100A9 in serum through the prediction model.

[0056] (III) Beneficial effects

[0057] The beneficial effects produced by the present invention are:

[0058] 1. The present invention has mined an innovative biomarker combination of HMGB1, S100A8, and S100A9 for the prediction and evaluation of the severity of Mycoplasma pneumoniae pneumonia in children, and combined with commonly used clinical laboratory test indicators, providing a method for constructing a pneumonia severity prediction model and a prediction model with high efficiency, high specificity, high sensitivity, and high accuracy.

[0059] 2. The present invention has been trained and verified in the training set and the test set, ensuring that the prediction model established by the present invention can early, simply, and accurately predict mild and severe Mycoplasma pneumoniae pneumonia, and thus assist in the formulation of clinical decisions.

[0060] 3. The RF machine learning prediction model provided by the present invention is easy to promote in clinical practice, optimizes the disease stratification and clinical diagnosis and treatment path of Mycoplasma pneumoniae pneumonia, is conducive to early identification of high-risk groups for severe and critical illness and sequelae, and reduces unnecessary ICU transfers; it can be used to guide the selection of clinical treatment drugs for children with Mycoplasma pneumoniae pneumonia, prevent the occurrence of severe illness, and provide evidence-based basis for reducing the fatality rate. Description of the Drawings:

[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained by extending the provided drawings.

[0062] Figure 1 It is a flowchart of the method for constructing the prediction model of the severity of Mycoplasma pneumoniae pneumonia provided by the present invention;

[0063] Figure 2 It is a schematic diagram of variable screening. Among them, A is the feature for variable screening using the Boruta algorithm, and a box plot showing the change of the importance scores of each variable during the operation of the Boruta algorithm is drawn (green are important variables, red are unimportant variables, blue are shadow variables, and yellow are Tentative variables); B is the variable screening using the recursive feature elimination method (RFE), and a distribution diagram of the accuracy rate under different numbers of variables in the model is drawn. When the total number of variables in the model is 52 (i.e., the blue solid dot in ), the model performance is the best; C is by screening the common variables of Boruta and RFE, and finally 10 indicators are included in the modeling study, namely the variables HMGB1, S100A8, S100A9, IL-6, Ferritin, PCT, LDH, TT, CD16+CD56+ cells, D-Dimer.

[0064] Figure 3 It is a schematic diagram for evaluating the predictive value of single variables. Among them, A is the ROC curve of the variables HMGB1, S100A8, S100A9, IL-6, Ferritin, PCT, LDH, TT, CD16+CD56+ cells, and D-Dimer predicting the severity of MPP alone in the training set; B is the ROC curve of the variables HMGB1, S100A8, S100A9, IL-6, Ferritin, PCT, LDH, TT, CD16+CD56+ cells, and D-Dimer predicting the severity of MPP alone in the test set;

[0065] Figure 4Performance of 9 prediction models established based on machine learning algorithms in the validation set and test set. Among them, A is the performance of 9 machine learning-based predictions of MPP severity in the training set; B is the performance of 9 machine learning-based predictions of MPP severity in the test set;

[0066] Figure 5 Predicting the severity of MPP by the prediction model established based on machine learning algorithms in the validation set and test set. Among them, A - B. ROC curves of 9 machine learning-based predictions of MPP severity in the training set (A) and test set (B); C - D. Calibration curves of the RF model in the training set (C) and test set (D) (i.e., the scatter plot and curve of the model prediction probability and the actual occurrence probability, with the horizontal axis being the prediction probability and the vertical axis being the actual situation); E - F. Confusion matrix of the prediction situation and the actual situation distribution in the training set (E) and test set (F) under the RF model.

[0067] Figure 6 SHAP dot plot for the interpretability analysis of the established prediction model RF model. Among them, A - B. The SHAP summary plot shows the feature importance of each predictor variable in the RF model in the training set (A) and test set (B) in descending order. The upper-level predictor variables are more important for the prediction results of the model. Create a point for each feature attribute value of the RF model for each patient. The farther the point is from the baseline SHAP value of zero, the stronger its impact on the model output. The dots are colored according to the feature values. Red indicates a higher feature value, and blue indicates a lower feature value. C - D. Bar chart showing the average absolute SHAP values of each predictor variable in the RF model in descending order. E and F. Force plots providing personalized feature attributions using two representative examples (C: mild MPP, D: severe MPP).

[0068] Figure 7 SHAP dependence scatter plot of the top 4 main indicators with the highest contribution in the prediction RF model. SHAP values of 4 indicators HMGB1 (A), S100A8 (B), S100A9 (C), and CD16 + CD56 + cells (D) at different values. A positive SHAP value indicates an increased probability of the outcome being severe, and a negative SHAP value indicates the opposite. As the values of HMGB1\S100A8\S100A9 increase, the SHAP values all increase, while the CD16 + CD56 + cells are the opposite. Specific implementation method:

[0069] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0070] An embodiment of the present invention provides a prediction model for the severity of Mycoplasma pneumoniae pneumonia. The main components of this model include: data collection, preprocessing and dataset establishment, variable screening, establishment and performance evaluation of a machine learning model prediction, and prediction model interpretation. The embodiment of the present invention is applicable to improving the accuracy of predicting the severe risk of Mycoplasma pneumoniae pneumonia patients, assisting doctors in clinical decision-making, and improving the prognosis of patients. To this end, the embodiment of the present invention aims to design to collect the blood of Mycoplasma pneumoniae pneumonia patients with different severities, quantitatively measure the contents of HMGB1, S100A8, and S100A9 proteins, and fuse the laboratory test results as input features, and use a machine learning classifier to build a prediction system for the severity of Mycoplasma pneumoniae pneumonia patients. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0071] Example 1: Construction of a prediction model for the severity of Mycoplasma pneumoniae pneumonia

[0072] As Figure 1 shown, the overall framework is divided into data collection and dataset establishment, construction of a MPP severity prediction model, model performance evaluation, and model visualization interpretation. The prediction of the severity of Mycoplasma pneumoniae pneumonia includes the following steps:

[0073] Step 1: Sample collection and dataset establishment:

[0074] The construction of the dataset and sample collection are the prerequisites for carrying out technical research. In this invention, serum samples of 277 patients with Mycoplasma pneumoniae pneumonia who visited Wuhan Children's Hospital from January 2023 to January 2024 were collected. Among them, 149 cases were mild and 128 cases were severe. The research subjects were screened strictly in accordance with the inclusion and exclusion criteria to exclude the influence of interference factors such as other diseases and treatment methods. The inclusion criteria for patients with Mycoplasma pneumoniae pneumonia are: a) Clinical diagnosis of pneumonia: respiratory symptoms and signs such as fever, cough, expectoration, shortness of breath, and pulmonary rales, and chest imaging showing lung involvement; b) Evidence of Mycoplasma pneumoniae infection: positive serum-specific MP-IgM antibody and / or positive sputum MP-RNA, positive MP-DNA, and / or tNGS indicating Mycoplasma pneumoniae infection. The exclusion criteria are: a) Having symptoms of respiratory tract infectious diseases and clinically diagnosed as non-Mycoplasma pneumoniae pneumonia: having clear evidence of being mainly infected with other pathogens, such as viruses, Streptococcus pneumoniae, and Mycobacterium tuberculosis infection; b) Having non-infectious pneumonia, such as bronchial foreign bodies, primary ciliary dyskinesia syndrome, allergic asthma, cystic granulomatous inflammation, etc.; c) Having underlying diseases, such as tumors, blood diseases, endocrine system diseases, congenital immunodeficiency diseases, congenital heart diseases, etc.; d) Congenital dysplasia or deformity, such as tracheal dysplasia, pulmonary artery sling, etc.; e) Patients within 3 months after surgery; f) Having a history of repeated respiratory tract infections within the past 3 months, and the number of respiratory tract infections ≥ 3 times.

[0075] The collection of clinical data of children with Mycoplasma pneumoniae pneumonia includes: patient age, gender, BMI; laboratory indicators include blood routine, high-sensitivity C-reactive protein, procalcitonin, ferritin, IL-6, blood coagulation, liver and kidney function electrolytes, myocardial enzyme spectrum, blood immune indicators, coagulation profile, blood immune indicators, and TBNK lymphocyte subsets at the time of admission.

[0076] Measurement of HMGB1, S100A8, and S100A9 protein levels: Specifically: The venous blood of the collected patients was placed at room temperature of 25°C for 30 minutes, centrifuged at a centrifugal force of 3000g for 5 minutes to separate the serum, and the serum was immediately stored at -80°C until measurement. HMGB1, S100A8, and S100A9 ELISA kits were used to detect the HMGB1, S100A8, and S100A9 protein contents in the serum samples.

[0077] Establishment of the dataset: The clinical data of patients with Mycoplasma pneumoniae pneumonia and the contents of HMGB1, S100A8, and S100A9 proteins in serum samples. All cases of Mycoplasma pneumoniae pneumonia were used to establish a dataset according to their disease severity (severe, mild). Variables were excluded with a missing ratio > 40%, and multiple imputation was used to impute the remaining missing variables. To ensure the scientific nature of the research method and the reliability of the research results, it is necessary to ensure that there are no significant differences in basic indicators such as age, gender, and body mass index (BMI) between the collected mild MPP patients and severe MPP patients, and to exclude the influence of basic indicators on the research results. Therefore, in this invention, analysis of variance and chi-square test were used to verify whether there are significant differences in basic indicators between mild MPP and severe MPP patients. The dataset was normalized and divided into a training set and a test set according to a ratio of 7:3. The training set includes a mild training set and a severe training set; the validation set includes a mild training set and a severe validation set. Among them, the severe dataset of Mycoplasma pneumoniae pneumonia: MPP patients meet the inclusion criteria of severe or critical illness; meet any one of the clinical manifestations of severe illness: a) persistent high fever above 39°C for more than 5 days or fever ≥ 7 days, and the peak body temperature shows no downward trend; b) hypoxemia, at rest, when breathing air, the finger pulse oxygen saturation ≤ 0.92; c) increased respiratory and pulse rates, with or without elevated PaCO2, and clinical evidence of respiratory distress and failure; d) signs of pulmonary infection, such as moderate to large pleural effusion, pulmonary consolidation, plastic bronchitis, pulmonary embolism, necrotizing pneumonia, and acute asthma exacerbation; e) the occurrence of extrapulmonary complications, such as meningoencephalitis, myocarditis, erythema multiforme, autoimmune hemolytic anemia, hemophagocytic syndrome, or disseminated intravascular coagulation, but not reaching the critical illness standard. f) Imaging manifestations are one of the following: f1) ≥ 2 / 3 of a single lung lobe is involved, with uniform high-density consolidation or high-density consolidation in 2 or more lung lobes, which may be accompanied by moderate to large pleural effusion and may also be accompanied by manifestations of localized bronchiolitis; f2) diffuse involvement of a single lung or bilateral ≥ 4 / 5 lung lobes shows bronchiolitis, which may be combined with bronchitis and mucus plug formation leading to atelectasis; g) the clinical symptoms gradually worsen, and the imaging shows that the lesion range progresses by more than 50% within 24 - 48 hours; h) one of CRP, LDH, D-Dimer is significantly elevated; Meeting the critical illness manifestations: referring to the presence of respiratory failure and / or life-threatening severe extrapulmonary complications that require life support such as mechanical ventilation. The remaining cases were classified into the mild dataset.

[0078] Step 2. Variable screening: The obtained training set samples were used for variable screening. Specifically:

[0079] First, use the Boruta algorithm to screen variables for the training set in Step 1. The specific steps are as follows: (1) Generate a copy for each original feature in the feature matrix and splice all copies to the feature matrix; (2) Randomly shuffle all copies of the features, that is, randomly shuffle the feature values to generate a shadow feature. There is no correlation between the shadow feature and the outcome. The shadow feature and the original features form a new feature matrix; (3) Run a random forest on the new feature matrix and collect the Z-score of each feature; (4) Find the maximum Z-score (Maximum Z-score Among Shadow Attributes, MZSA) in the shadow features. Then, mark each original feature with a Z-score higher than MZSA as "hit" or "selected"; (5) Repeat the above steps N times. Then, for each original feature with undetermined importance, based on the binomial distribution B(N, 0.5), perform a two-sided hypothesis test on the number of times the original feature is "selected" to determine its importance; (6) According to the results of the hypothesis test, consider the features with significantly lower importance than MZSA as "unimportant" and permanently delete them from the feature set; (7) According to the results of the hypothesis test, consider the features with significantly higher importance than MZSA as "important"; (8) Delete all shadow features; (9) Repeat all the above processes until importance is assigned to all features or the algorithm reaches the upper limit of the preset number of loops. The Boruta algorithm shows the importance of each feature variable during the loop process in a box plot. In this study, the maximum number of loops is set to 100. Finally, draw a box plot of the changes in the importance scores of each variable during the running of the Boruta algorithm, and select the top 12 determined differential variables (S100A8, S100A9, HMGB1, LDH, CysC, Na, D-Dimer, TT, PCT, Ferritin, IL-6, CD16+CD56+ cells).

[0080] Then, use the Recursive Feature Elimination (RFE) method to screen variables for the training set in Step 1. Start searching for feature subsets from all features in the training dataset and successfully delete features until the required number is retained. In this study, by fitting the random forest algorithm, based on all the initially included variables, repeatedly construct models and select the best features, and then repeat this process on the remaining features in turn until all features are traversed. Screen the top 12 variables (HMGB1, S100A8, S100A9, IL-6, Ferritin, PCT, LDH, hsCRP, TT, CD16+CD56+ cells, D-Dimer, Ca) for inclusion in the modeling study.

[0081] Finally, by screening the common variables of Boruta and RFE, the best 10 variables (HMGB1, S100A8, S100A9, IL-6, Ferritin, PCT, LDH, TT, CD16+CD56+ cells, D-Dimer) were finally selected and included in the modeling study;

[0082] And the univariate predictive value was evaluated, and the univariate predictive value of the 10 variables was evaluated respectively to further select the optimal variables.

[0083] Step 3: Construction of the Mycoplasma pneumoniae pneumonia severity prediction model. The 10 variables screened in Step 2 were used as the input of the model, and 9 machine learning methods including Logistic, LightGbm, KNN, SVM, XGBoost, RF, GNB, DT, and Catboost were used for training input in the training set in Step 1. The classification model parameters are set as follows:

[0084] First, determine the optimal hyperparameters of the LightGBM model: n_estimators = 70, learning_rate = 0.4, max_depth = 13, number_leaves = 15,; the optimal hyperparameters of the XGBoost model: max_depth = 3, random_state = 42, objective = reg:squarederror, eval_metric = rmse, min_child_weight = 10, gamma = 0.1, reg_alpha = 0.5, reg_lambda = 1.0; the optimal hyperparameters of the Logistic model: def plot_decision_boundary(x, y, model): x_min, x_max = X[:,0].min() - 1, X[:,0].max() + 1, y_min, y_max = X[:,1].min() - 1, X[:,1].max() + 1; xx, yy = np.meshgrid(np.linspace(x_min, x_max, 100), np.linspace(y_min, y_max, 100)), Z = model.predict(np.c_[xx.ravel(), yy.ravel()]), Z = Z.reshape(xx.shape); the optimal hyperparameters of the RF model: ScoreAll = [] for i in range(100, 300, 10): model = RandomForestClassifier(max_features = 2, n_estimators = i, random_state = 10) model.fit(X_train, y_train) score = model.score(X_test, y_test) ScoreAll.append([i, score]) ScoreAll = np.array(ScoreAll); the optimal hyperparameters of the KNN model: for k in range(1, 10): s_knn = S_KNN(k) s_knn.fit(X_train, y_train) score = s_knn.score(X_test, y_test) k_list.append(k) score_list.append(score) if score > best_score: best_score = score best_k = k; the optimal hyperparameters of the SVM model: mp.tune <- tune.svm(Degree~., data = train, kernel = "linear", cost = c(0.001, 0.01, 0.1, 1, 5, 10, 15)), bestmp <- mp.tune$best.model; Optimal hyperparameters of the GNB model: def bayesian_optimization(X_train, y_train, X_test): gpr.fit(X_train, y_train) y_pred, sigma = gpr.predict(X_test, return_std = True) gap = y_pred - f(X_test) improvement = (gap + np.sqrt(sigma ** 2 + 1e - 6) * np.abs(gap).mean()) * 0.5 result = minimize(gpr.kernel_, np.zeros(gpr.kernel_.shape[0]), method = 'L - BFGS - B') hyperparameters = result.x return hyperparameters, improvement.max() n_iter = 20, n_samples = 5, results = [], for i in range(n_iter): X_train = np.random.uniform(-2 * np.pi, 2 * np.pi, (n_samples, 1)) y_train = f(X_train), result = bayesian_optimization(X_train, y_train, X_test = np.random.uniform(-2 * np.pi, 2 * np.pi, (100, 1))) results.append(result); Optimal hyperparameters of the DT model: learning_rate = 0.6, gamma = 0.9, max_depth = 3, subsample = 0.799, min_child_weight = 1, n_estimators = 2000; Optimal hyperparameters of the Catboost model: iterations = 100, learningrate = 0. It should be noted that there seem to be some errors or incomplete expressions in the original text, such as some commas and parentheses not being properly paired, and the "learningrate = 0." at the end is incomplete. But the translation is done according to the rules.1. randomstate = 42, verbose = True, silentfit = True; The model is trained using optimized hyperparameters, with the 10 key variables described in Step 2 input, and outputs the classification results of mild or severe cases of Mycoplasma pneumoniae pneumonia; and the performance of the model in the training set and test set is evaluated, and the prediction performance of the model is verified according to the accuracy, F1 score, Kappa value, and the area under the receiver operating characteristic curve (AUC) index, and a prediction model for the severity of Mycoplasma pneumoniae pneumonia with better performance is selected. In the training set, Lightgbm (accuracy: 1.00, F1 score: 1.00, Kappa value: 1.00, and AUC: 1.00), RF (accuracy: 0.99, F1 score: 0.99, Kappa value: 0.99, and AUC: 0.99), and CatBoost model (accuracy: 0.97, F1 score: 0.98, Kappa value: 0.97, and AUC: 0.98) all have good performance; in the test set, the performance of each model has decreased to varying degrees, but the RF model is generally more robust. Combining the accuracy (0.90), Kappa value (0.80), and AUC (0.90), a relatively excellent RF model is selected. 10-fold cross-validation and Boostrap resampling are used to optimize the model, and the optimal prediction model - the RF model is constructed by adjusting the parameters: (train.control <- trainControl(method = "repeatedcv", number = 10, repeats = 500)).

[0085] The calibration curve of the RF model using the training set and test set provides insights for model calibration, which centrally reflects the consistency between the predicted severe risk of MPP and the actual empirical observation results. A confusion matrix of the prediction model for the severity of Mycoplasma pneumoniae pneumonia is constructed to show the corresponding relationship between the model prediction results and the actual situation distribution. The evaluation of the prediction model for the severity of Mycoplasma pneumoniae pneumonia is completed according to the confusion matrix. The metric formulas are shown below.

[0086] Table 1 Confusion matrix

[0087] Confusion Matrix Predicted Severe Predicted Mild Actual Severe TP (True Positive) FN (False Negative) Actual Mild FP (False Positive) TN (True Negative)

[0088]

[0089] Step 4: Interpretability analysis of the pneumonia severity prediction model. The SHAP tool was used to perform interpretability analysis on the constructed model. SHAP is a tool for interpreting "black-box models" by calculating the marginal contribution of features to the model. The core of SHAP is the SHAP value. SHAP attempts to provide an intuitive visualization of how different features affect individual predictions. In all samples, each feature has a SHAP value. It represents the degree to which the feature affects the prediction result in that sample. When the SHAP value is 0, the feature in the sample does not affect the prediction result; when the absolute value of the SHAP value is large, it indicates that the feature in the sample has a large impact on the prediction result. When the SHAP value is positive, it represents a positive gain to the prediction, and vice versa for a negative gain. A feature density scatter plot was drawn. At the same time, SHAP was also used to present examples of how to locally interpret individual predictions, attempting to provide personalized feature attributions using a representative example (such as randomly selecting a mild MMP patient and a severe MPP patient), and illustrating how to use SHAP to interpret individual model predictions. It represents an intuitive way to guide the decisions of clinicians and patients and improve their understanding of how the developed model makes specific predictions. It starts from the base value, which is the average of all predictions. Then, each predictor (along with its corresponding Shapley value) is represented by an arrow, which indicates whether the predicted value of the model increases (red) or decreases (blue) relative to the base value. The importance of the predictor is reflected by the length of its arrow, and the longer the arrow, the more important the predictor. The feature values are listed at the bottom of the figure. Finally, for a specific patient, the predicted output value of the model is obtained, represented by the intersection point of the red and blue arrows. In addition, the SHAP dependence plot was used to quantitatively visualize the relationship between the main risk factors and the outcome, showing the top 4 most important features (HMGB1, S100A8, S100A9, and CD16+CD56+ cells), and it was also possible to determine the cut-off values of each variable to distinguish high-risk (SHAP value > 0) and low-risk (SHAP value < 0) severe MPP. A positive SHAP value indicates an increased probability of the outcome being severe, and vice versa for a negative SHAP value. As the values of HMGB1, S100A8, and S100A9 increase, the SHAP values all increase, while the CD16+CD56+ cells are the opposite.

[0090] Example 2: Determination of the protein contents of HMGB1, S100A8, and S100A9 in the blood samples of 277 children with Mycoplasma pneumoniae pneumonia

[0091] The kits used in the following examples:

[0092] Human S100 Calcium Binding Protein A8 (S100A8) ELISA Kit, Human S100 Calcium Binding Protein A9 (S100A9) ELISA Kit, Human High Mobility Group Protein B1 (HMGB-1) ELISA Kit

[0093] Product Catalog Number: ELK1125, ELK2197, ELK1438

[0094] Operation Steps:

[0095] 1) Reagent and Sample Preparation: Before use, equilibrate the HMGB1, S100A8, and S100A9 ELISA kits at room temperature for 30 min; Prepare the washing solution (dilute the concentrated washing solution with distilled water at a ratio of 1:20).

[0096] 2) Sample Incubation: Set up the standard wells, 0-value well, blank well, and sample wells. Add 100 μL of standards with different concentrations to the standard wells, 100 μL of sample diluent to the 0-value well, add nothing to the blank well, and add 100 μL of the sample to be tested to the sample wells; Incubate at 37 °C for 80 minutes.

[0097] 3) Add Capture Antibody Discard the liquid in the ELISA plate. Add 200 μL of washing buffer to each well and wash 3 times. After patting dry, add 100 μL of biotinylated antibody working solution to each well and incubate at 37 °C for 50 minutes.

[0098] 4) Add Enzyme-Labeled Secondary Antibody Discard the liquid in the ELISA plate. Add 200 μL of washing buffer to each well and wash 3 times. After patting dry, except for the blank well, add 100 μL of HRP enzyme working solution to the standard wells, 0-value well, and sample wells, and incubate at 37 °C for 50 minutes

[0099] 5) Substrate Incubation Discard the liquid in the ELISA plate. Add 200 μL of washing buffer to each well and wash 5 times. After patting dry, add 90 μL of TMB to each well and incubate at 37 °C for 20 min.

[0100] 6) Terminate the Reaction: Add 50 μL of stop solution to each well;

[0101] 7) Colorimetric Measurement: Immediately measure the absorbance at 450 nm after adding the stop solution. Before colorimetric measurement, gently shake the microplate to evenly disperse the liquid, and read the absorbance (OD value) at 450 nm for each well on the microplate reader.

[0102] Results:

[0103] Table 2 HMGB1, S100A8, and S100A9 Protein Contents in Serum Samples of 277 Children with Mycoplasma Pneumonia

[0104]

[0105]

[0106]

[0107]

[0108]

[0109] Example 3: Analysis of the General Clinical Characteristics of 277 Patients with Mycoplasma Pneumoniae Pneumonia

[0110] Analysis method: Collect the clinical relevant indicators of the patients, including demographic characteristics, clinical manifestations, biochemical indicators, etc. This invention is based on a single-center clinical study, and 277 children with Mycoplasma pneumoniae pneumonia were recruited from Wuhan Children's Hospital. In this invention, the ratio of mild to severe cases is 1.16:1 (149 vs 128), and the male to female ratio is 1:1.13 (130 vs 147).

[0111] Results: The main demographic and clinical characteristics of the study population are shown in Table 3.

[0112] Table 3 The Main Demographic and Clinical Characteristics of the Study Population

[0113]

[0114]

[0115]

[0116]

[0117] Example 4: Establishment of the Prediction Dataset, Analysis of the Correlation between Each Index and Prediction Grouping

[0118] All cases were used to establish a dataset according to their disease severity (mild and severe), and the children were randomly divided into a training set and a test set at a ratio of 7:3. There were 105 cases of mild MPP and 90 cases of severe MPP in the training set, and 44 cases of mild MPP and 38 cases of severe MPP in the test set. The basic characteristics of each variable in the training set and the test set are shown in Table 4.

[0119] Table 4 Basic Information of the Cases in the Training Cohort

[0120]

[0121]

[0122]

[0123]

[0124] Basic information of cases in the test cohort in Table 5

[0125]

[0126]

[0127]

[0128]

[0129]

[0130] Example 5: Screening of variables for predicting the severity of Mycoplasma pneumoniae pneumonia

[0131] Analysis method: First, the Boruta algorithm is used for variable screening, and a box plot of the changes in the importance scores of each variable during the operation of the Boruta algorithm is drawn ( Figure 2 A). The green ones are important variables, the red ones are unimportant variables, the blue ones are shadow variables, and the yellow ones are Tentative variables. As a result, 12 determined differential variables (S100A8, S100A9, HMGB1, LDH, CysC, Na, Ddimer, TT, PCT, Ferritin, IL-6, CD16+CD56+ cells) and 8 candidate variables that may be different (DBIL, ALP, ALB, ALT, AST, Cl, Ca, hsCRP) are screened, and the remaining variables are excluded.

[0132] Then, the recursive feature elimination method (RFE) is used to screen variables. Figure 2 B is the accuracy distribution diagram for different numbers of variables in the model. When the total number of variables in the model is 52 (i.e., the blue solid dot in the figure), the model performance is the best. The accuracy of the model reaches 0.6368 after the first variable is incorporated. After the first 10 variables are incorporated, the accuracy of the model rises to 0.8010. Subsequently, as the number of variables increases, the accuracy of the model fluctuates. When the number of incorporated variables is 12, the accuracy of the model is approximately 0.8044. After that, as the number of incorporated variables increases, the accuracy of the model fluctuates again. In order to improve the practical application value of the model, we only screen the first 12 variables for the modeling study. That is, the variables HMGB1, S100A8, S100A9, IL-6, Ferritin, PCT, LDH, hsCRP, TT, CD16+CD56+ cells, D-Dimer, and Ca are incorporated into the modeling.

[0133] Finally, Figure 2Through screening the common variables of Boruta and RFE, finally 10 indicators were included in the modeling study. That is, the variables HMGB1, S100A8, S100A9, IL-6, Ferritin, PCT, LDH, TT, CD16+CD56+ cells, and D-Dimer were included in the subsequent modeling.

[0134] See Figure 3 , further use the ROC curve to compare the predictive values of the 10 selected variables HMGB1, S100A8, S100A9, IL-6, Ferritin, PCT, LDH, TT, CD16+CD56+ cells, and D-Dimer when used alone. When HMGB1, S100A8, and S100A9 are used alone, they have good discriminant values in both the training set ( Figure 3 A) and the test set ( Figure 3 B), especially HMGB1 (AUC ≥ 0.8).

[0135] Example 6: Establishment of a prediction model for the severity of Mycoplasma pneumoniae pneumonia patients;

[0136] Analysis method: Using the above 10 selected variables as independent variables, 9 machine learning methods were used to establish a prediction model in the training set. Taking AUC as the main evaluation index, and indicators such as accuracy, sensitivity, specificity, PPV, NPV, F1 score, and Kappa score as secondary evaluation indexes to evaluate the prediction efficacy of the model ( Figure 4 A). Except for KNN and GNB, the accuracy and AUC of the remaining models in the test set are both 0.80 and above. Especially, the accuracy and AUC of the Lightgbm, RF, and CatBoost models in the training set are both 0.90 and above, showing good efficacy ( Figure 4 A and 5A). In the test set ( Figure 4 B and 5B), the efficacies of each model decreased to varying degrees, but the RF model had the highest AUC (0.90), accuracy (0.90), and Kappa score (0.80), indicating that the model is well-recognized and robust. The training set ( Figure 5 C) and the test set ( Figure 5 D) calibration curves of the RF model provided insights for model calibration, which concentratedly reflected the consistency between the predicted MPP severe risk and the actual empirical observation results. The calibration curves of the model on the training set and the test set showed that the predicted probability and the actual probability had good consistency. Therefore, the RF model was considered the optimal model. Figure 5 E-F shows the confusion matrices of the RF model in the training set ( Figure 5 E) and the test set ( Figure 5 F) to further emphasize the prediction ability of the model.

[0137] Example 7: Interpretability Analysis of the Prediction Model for the Severity of Mycoplasma Pneumoniae Pneumonia in Children Based on HMGB1, S100A8, and S100A9

[0138] SHAP values can provide more insights into how the RF model predicts the severity of MPP. Feature importance is summarized by the SHAP summary plot in Figure 6 . In the figure, the abscissa represents the SHAP value, and the ordinate represents each feature in the classification task. The points in the figure represent each sample, and the color of the points represents the magnitude of the sample value of the feature. Figure 6 The SHAP summary plot of A - B illustrates the entire distribution of the impact of each feature on the model output in the training set ( Figure 6 A) and the test set ( Figure 6 B). The color enables us to understand how the change in the feature value affects the change in the result. Red represents a high feature value, and blue represents a low feature value. The farther a point is from the baseline SHAP value of 0, the stronger its impact on the output. Figure 6 C and D depict the standard bar charts of the average absolute SHAP values of each predictor in descending order, which helps to better understand the relationship between the features and the SHAP values (as well as the prediction output). Among the image features, the absolute values of the SHAP values at the midpoints of HMGB1, S100A8, and S100A9 are relatively large, indicating that they play important roles in the classification task. The CD16+CD56+ cell and PCT features rank fourth and fifth respectively. In addition, the SHAP values of the red sample points in the CD16+CD56+ cell and TT features are less than 0, which will have a negative impact on the prediction. That is, when the CD16+CD56+ cell and TT feature values are more significant, it is less likely to cause the patient to become a severe MPP. This figure allows clinical experts to further understand the prediction process of the model. Figure 6 The force plots in E and F use a representative example (such as a randomly selected mild MPP patient ( Figure 6 E) and a severe MPP patient ( Figure 6 F)) to provide personalized feature attributions, illustrating how to use SHAP to explain individual model predictions. Figure 6 For the child in E, the HMGB1 was 1507.98 pg / mL, IL - 6 was 5.26 pg / mL, PCT was 0.06 ng / mL, and CD16+CD56+ cells were 207 within 24 hours of admission, indicating that the patient was mild, but the D - Dimer was 1.18, which was positively correlated with the prediction of severe cases. Figure 6 For the child in F, the IL - 6 was 14.35 pg / mL, S100A9 was 208.3 pg / mL, and S100A8 was 682.78 pg / mL within 24 hours of admission, which greatly promoted the child to develop severe MPP. Figure 7A-D also quantitatively visualizes the relationship between the main risk factors and the outcomes. The SHAP dependence plot shows the top 4 most important features, with higher HMGB1, S100A8, and S100A9 values and lower CD16+CD56+ cell levels leading to an increased risk of severe MPP. At the same time, the cut-off values of each variable can also be determined to distinguish between high-risk (SHAP value > 0) and low-risk (SHAP value < 0) severe MPP.

[0139] In summary, the present invention provides a prediction model for the severity of Mycoplasma pneumoniae pneumonia and a method for constructing the same. The prediction model has good prediction performance, good clinical applicability, and can improve the accuracy of risk stratification for patients with Mycoplasma pneumoniae pneumonia, solving the problems of lack of prediction tools and difficulty in early identification in the prior art, and can effectively improve the prediction accuracy in Mycoplasma pneumoniae pneumonia, especially in children with Mycoplasma pneumoniae pneumonia.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the present invention and not to limit the protection scope of the present invention. Additionally, after reading the technical content of the present invention, those skilled in the art can make various changes, modifications, or variations to the present invention, and all these equivalent forms also fall within the protection scope defined by this application.

Claims

1. A method for constructing a prediction model for the severity of Mycoplasma pneumoniae pneumonia, characterized in that: The prediction model is a risk prediction model obtained based on the contents of HMGB1, S100A8, and S100A9 proteins and laboratory test indicators in the serum samples of patients with Mycoplasma pneumoniae pneumonia, so as to predict the severity of Mycoplasma pneumoniae pneumonia; The construction method of the prediction model includes the following steps: Step 1: Data collection, processing, and dataset establishment: According to the preset clinical inclusion and exclusion criteria, collect venous blood samples of patients with Mycoplasma pneumoniae pneumonia confirmed by laboratory tests, detect the contents of HMGB1, S100A8, and S100A9 proteins in the serum samples, and synchronously collect the clinical baseline data of the patients when they are admitted to the hospital; Divide the cases into mild and severe datasets according to the severe classification criteria, and normalize the data; use the stratified random sampling method to divide the training set and the test set in a ratio of 7:3; the training set includes a mild training set and a severe training set, and the test set includes a mild test set and a severe test set; Step 2: Variable screening: Use the Boruta algorithm and the recursive feature elimination method to screen variables for the training set in Step 1 respectively, and select 10 key variables commonly identified by the two algorithms as the model inputs. The key variables are: HMGB1, S100A8, S100A9, IL-6, Ferritin, procalcitonin PCT, lactate dehydrogenase LDH, thrombin clotting time TT, CD16+CD56+ cells, D-dimer D-Dimer; and evaluate the independent prediction efficacy of each screened variable by drawing a univariate receiver operating characteristic curve ROC, and further select and optimize the indicators; Step 3: Model construction and efficacy comparison: Use 9 machine learning algorithms, namely LightGBM, XGBoost, Logistic, RF, KNN, SVM, GNB, DT, and Catboost, to construct prediction models respectively. Input the 10 key variables in Step 2, and output the classification results of mild or severe cases of Mycoplasma pneumoniae pneumonia; And evaluate and screen the efficacy of the models in the training set and the test set, and evaluate the consistency between the model prediction results and the actual observation results through calibration curves and confusion matrices; Step 4: Model interpretation: Use SHAP additive interpretation values to analyze the features of the selected model and evaluate the contribution degree of each variable index in the model. Finally, the contribution degrees of each index are obtained as follows: HMGB1 = 0.1956, S100A8 = 0.0701, S100A9 = 0.0672, CD16+CD56+ cells = 0.0557, PCT = 0.0419, Ferritin = 0.0391, TT = 0.0281, LDH = 0.0216, IL-6 = 0.0150, D-Dimer = 0.0115.

2. The construction method of a Mycoplasma pneumoniae pneumonia severity prediction model according to claim 1, wherein The preparation method of the serum sample includes: (a) Collect the patient's venous blood in a vacuum blood collection tube and let it stand at 25°C for 30 minutes; (b) Centrifuge at a centrifugal force of 3000×g for 5 minutes, and take the supernatant as the serum sample.

3. The construction method of a Mycoplasma pneumoniae pneumonia severity prediction model according to claim 1, characterized in that, In the first step, the clinical baseline data includes the patient's symptoms, signs, laboratory test results, and imaging examination results, specifically including: gender, age, body temperature at admission, respiratory rate at admission, blood routine at admission, hypersensitive C-reactive protein (hsCRP), procalcitonin, ferritin, IL-6, blood coagulation, liver and kidney function electrolytes, myocardial enzyme spectrum, blood immune indexes, coagulation profile, blood immune indexes, TBNK lymphocyte subsets.

4. The method for constructing a prediction model for the severity of Mycoplasma pneumoniae pneumonia according to claim 1, wherein The specific method for establishing the data set is as follows: According to the clinical baseline data of patients with Mycoplasma pneumoniae pneumonia and the contents of HMGB1, S100A8, and S100A9 proteins in the serum samples, all cases of Mycoplasma pneumoniae pneumonia are used to establish a data set according to the severity of their conditions. The missing value processing adopts a two-step method: First, the feature variables with a missing rate > 40% are excluded, and the retained variables are imputed using the multiple imputation by chained equations (MICE) method. The number of imputations is set to 5 times. Analyze the general characteristics of each variable in the data set, and use the normalization formula of the maximum and minimum values to normalize the data: Where X is the current value, Xmin is the minimum value of this feature, Xmax is the maximum value of this feature, and the normalized value Xscale is calculated.

5. The construction method of a Mycoplasma pneumoniae pneumonia severity prediction model according to claim 1, characterized in that, The severe disease classification criteria include at least one of the following: a) Persistent high fever above 39°C for more than 5 days or fever ≥ 7 days, and the peak body temperature shows no downward trend; b) Hypoxemia, with a finger pulse oxygen saturation ≤ 0.92 while breathing air at rest; c) Increased respiratory and pulse rates, with or without an increase in PaCO2, and clinical evidence of respiratory distress and failure; d) Signs of pulmonary infection, including moderate to large-area pleural effusion, pulmonary consolidation, plastic bronchitis, pulmonary embolism, necrotizing pneumonia, and acute exacerbation of asthma; e) The occurrence of extrapulmonary complications, such as meningoencephalitis, myocarditis, erythema multiforme, autoimmune hemolytic anemia, hemophagocytic syndrome, or disseminated intravascular coagulation, but not reaching the critical illness standard; f) One of the following imaging manifestations: f1) ≥ 2 / 3 of a single lung lobe is involved, with uniform high-density consolidation or high-density consolidation in 2 or more lung lobes, which may be accompanied by moderate to large pleural effusion and may also be accompanied by manifestations of localized bronchiolitis; f2) Diffuse involvement of a single lung or bilateral ≥ 4 / 5 lung lobes with bronchiolitis manifestations, which may be combined with bronchitis and accompanied by mucus plug formation, leading to atelectasis; g) Progressive aggravation of clinical symptoms, and imaging shows that the lesion range progresses by more than 50% within 24 - 48 hours; h) A significant increase in any one of CRP, LDH, and D-Dimer.

6. The construction method of a Mycoplasma pneumoniae pneumonia severity prediction model according to claim 1, characterized in that, In the third step, nine machine learning models are trained simultaneously using the training set. The models are trained with optimized hyperparameters, with the 10 key variables in the second step as the input, and the classification results of mild or severe Mycoplasma pneumoniae pneumonia are output. The performance of the models in the training set and the test set is evaluated, and the prediction performance of the models is verified according to the accuracy, F1 score, Kappa value, and the area under the receiver operating characteristic curve (AUC) index. The prediction model for the severity of Mycoplasma pneumoniae pneumonia with good performance is selected, and 10-fold cross-validation and Bootstrap resampling are used to optimize the model. The optimal prediction model is constructed by adjusting the parameters.

7. The method for constructing a prediction model for the severity of Mycoplasma pneumoniae pneumonia according to claim 6, wherein, The optimal prediction model is a random forest (RF) model.

8. The construction method of a Mycoplasma pneumoniae pneumonia severity prediction model according to claim 1, characterized in that, In the third step, the calibration curve and confusion matrix of the RF prediction model are constructed, and the consistency between the model prediction results and the actual observation results is evaluated according to the calibration curve and the confusion matrix, completing the evaluation of the prediction model for the severity of Mycoplasma pneumoniae pneumonia. The metric formula for the confusion matrix is as follows: Among them, TP represents predicted severe and actually severe; FP represents predicted severe and actually mild; FN represents predicted mild and actually severe; TN represents predicted mild and actually mild.

9. Application of the prediction model for the severity of Mycoplasma pneumoniae pneumonia according to claim 1 in the prediction of the severity of Mycoplasma pneumoniae pneumonia.

10. The application according to claim 9, characterized in that, The application includes predicting the severity of Mycoplasma pneumoniae pneumonia, especially in children, by detecting the protein expression levels of HMGB1, S100A8, and S100A9 in serum through the prediction model.

Citation Information

Patent Citations

  • Method and device for generating prediction model for predicting pneumonia severity

    CN112183572A