Construction method of pneumonia severity prediction model and prediction model

By constructing a multivariate logistic regression model based on S100A8/A9 protein content, the problem of low sensitivity in the assessment of mycoplasma pneumonia in children in the existing technology was solved, and efficient and accurate identification of severe cases and individualized treatment were achieved, thus optimizing the clinical diagnosis and treatment of mycoplasma pneumonia in children.

CN121237390APending Publication Date: 2025-12-30WUHAN CHILDRENS HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410523120.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-28
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing pneumonia scoring systems have low sensitivity in assessing the severity of mycoplasma pneumonia in children, making it difficult to accurately predict severe cases in the early stages. The lack of unified assessment tools and suitable circulating biomarkers makes it difficult to identify and prevent severe cases.

Method used

A multivariate logistic regression model based on S100A8/A9 protein content was constructed. Combined with commonly used clinical inflammation assessment indicators, variables were screened using lasso regression and random forest models to establish a predictive model for the severity of mycoplasma pneumonia, which can be used to identify the severity of mycoplasma pneumonia in children at an early stage.

Benefits of technology

It provides an efficient, highly specific, and highly sensitive pneumonia severity prediction model, which can promptly identify severe and critical cases, guide individualized treatment, reduce mortality, and optimize disease stratification and clinical treatment pathways.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237390A_ABST
    Figure CN121237390A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of medicine, and relates to a construction method of a pneumonia severity prediction model and the prediction model.The construction method comprises the steps that 1, a serum sample of a patient infected by mycoplasma pneumonia is obtained; 2) detecting the content of S100A8 / A9 protein in the serum sample; 3) establishing a data set according to the patient infected by mycoplasma pneumonia, wherein the data set comprises a training set and a verification set; 4) constructing a regression model for the content of the S100A8 / A9 protein by adopting a lasso regression mode; and 5) training the regression model by using the training set to obtain a pneumonia severity prediction model. The invention provides the construction method of the pneumonia severity prediction model which has good prediction performance and good clinical applicability and can improve the accuracy of risk stratification of mycoplasma pneumonia patients, and the prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of medicine, and relates to a construction method of a pneumonia severity prediction model and the prediction model, in particular to a construction method of a S100A8 / A9-based Mycoplasma pneumoniae pneumonia severity prediction model. BACKGROUND

[0002] Mycoplasma pneumoniae pneumonia (MPP) is an acute respiratory infectious disease caused by Mycoplasma pneumoniae. Due to the immature development of the immune system of children, there is a large difference between cases. Children with mild MPP are mostly manifested as upper respiratory tract infection, and have certain self-limiting nature, and have a good prognosis; while children with severe MPP are mostly caused by multiple factors such as combined bacterial and viral infection, and induced autoimmune reaction, resulting in severe lung infection, and even can cause myocarditis, nephritis, encephalitis and hemolytic anemia and other extrapulmonary complications, causing multiple organ dysfunction, increasing the risk of sequelae, and seriously endangering the life and health of children. Therefore, in clinical practice, how to early and accurately predict and identify severe and critical cases, reasonably diagnose and treat, shorten the course of disease, and avoid death and the occurrence of sequelae is the core and key problem of the diagnosis and treatment of children with MPP.

[0003] Reasonable risk assessment is the first step for early identification and prevention of severe or critical MPP. Among the existing pneumonia scoring systems, only the clinical pulmonary infection score (CPIS) can be applied to the severity assessment of children with pneumonia. The CPIS score integrates clinical, imaging and microbiological criteria to assess the severity of infection, and the main purpose is to reduce unnecessary antibiotic exposure, but since the scoring system includes few indicators and the results do not exclude age difference bias, it is not widely used in clinical practice. At present, there is no uniform scale or tool to evaluate and analyze the condition of children with MPP in clinical practice, and the research on predicting severe Mycoplasma pneumoniae pneumonia (SMPP) is mainly focused on finding independent risk factors for SMPP, such as fever for more than 10 days, pleural effusion, extrapulmonary complications, and C-reactive protein (CRP) > 40 mg / l. Although the specificity of SMPP prediction is high, the sensitivity of early identification is low, and it is difficult to accurately predict the occurrence of SMPP in an early stage. Therefore, it is necessary to establish a scoring model that can timely predict the severity of MPP, so as to prevent the development of SMPP.

[0004] The pathogenesis of SMPP is not clear, and studies have shown that the pathogen load, virulence, and imbalance of the host immune status may be important reasons for the development of the disease. Therefore, dynamic monitoring of immune and inflammatory indicators or searching for specific circulating markers is of great significance for evaluating the severity of MPP. Laboratory indicators commonly used to evaluate the severity of infectious diseases include white blood cell count, neutrophil ratio, acute phase reactant protein CRP, procalcitonin (PCT), interleukin 6 (IL-6), serum amyloid A (SAA), and ferritin (FER). The values of these indicators can indicate the type or severity of infection. However, the above indicators have limited effect on early prediction of the severity of MPP in children. Therefore, finding circulating biomarkers suitable for the severity of pneumonia in children can greatly benefit the precise diagnosis and treatment of MPP in children. Domestic and foreign researches focus on the immune response of the host after infection and have discovered a series of damage-associated molecular patterns (DAMPs) that activate innate immune responses and a series of cytokines related to "inflammatory storm". As DAMPs, the S100 calcium-regulating protein family plays an important role in the development of inflammatory diseases. In recent years, S100A8 / A9, a member of the S100 protein family, has attracted widespread attention as a key protein in regulating inflammatory responses.

[0005] S100A8 and S100A9 proteins (also known as MRP8 and MRP14, respectively) are acidic calcium-binding proteins with a molecular weight of about 15 kDa, which are abundant in innate immune cells and account for almost two-thirds of the soluble cell membrane protein content of all neutrophils. S100A8 / A9 proteins mainly mediate intracellular inflammatory signaling pathways by binding and activating Toll-like receptors and glycosylated end product receptors, and play an important role in the inflammatory response and the body's innate immune process. A series of studies have shown that S100A8 / A9 has important clinical value in diagnosis, treatment, and prognosis in various inflammatory diseases (such as rheumatoid arthritis, myocarditis, systemic lupus erythematosus, etc.). Currently, S100A8 / A9 has also been shown to be elevated in various respiratory diseases (such as asthma, idiopathic pulmonary fibrosis, and pulmonary heart disease). However, the study of S100A8 / A9 in pediatric MPP is still lacking. Given its special significance in the diagnosis and prognosis of infectious diseases in adults, further exploration of the expression differences between S100A8 / A9 in children with mild and severe MPP, combined with routine laboratory indicators for admission to analyze the risk factors associated with the development of pediatric SMPP, and based on the results of the analysis, a nomogram model can be constructed and verified for early identification and intervention of severe MPP. SUMMARY

[0006] In order to solve the above technical problems in the background art, the present application provides a method for constructing a pneumonia severity prediction model with good prediction performance, good clinical applicability and high accuracy in risk stratification of patients with mycoplasma pneumonia, and the pneumonia severity prediction model.

[0007] In order to achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0008] The application of S100A8 / A9 protein content in constructing a pneumonia severity prediction model is exemplarily applied in constructing a mycoplasma pneumonia severity prediction model, and preferably applied in constructing a pediatric mycoplasma pneumonia severity prediction model.

[0009] A method for constructing a mycoplasma pneumonia severity prediction model, characterized in that the method comprises the following steps:

[0010] 1) obtaining a serum sample of a patient infected with mycoplasma pneumonia;

[0011] 2) detecting the content of S100A8 / A9 protein in the serum sample in step 1);

[0012] 3) establishing a data set of patients infected with mycoplasma pneumonia according to step 1), wherein the data set comprises a training set and a validation set; the ratio of the training set to the validation set is 7:3; the training set comprises a mild training set and a severe training set; and the validation set comprises a mild training set and a severe validation set;

[0013] 4) constructing a regression model for the content of S100A8 / A9 protein obtained in step 2) by using lasso regression;

[0014] 5) training the regression model in step 4) by using the training set in step 3) to obtain a multi-factor logistic regression model; the multi-factor logistic regression model is a low-order version of the mycoplasma pneumonia severity prediction model.

[0015] Preferably, the specific implementation of step 4) adopted by the present application is:

[0016] 4.1) using lasso regression to establish a lambda value model for the content of S100A8 / A9 protein obtained in step 2);

[0017] 4.2) screening the variables of the model obtained in step 4.1);

[0018] 4.3) using a random forest model to verify the variables obtained in step 4.2) to obtain a regression model.

[0019] Preferably, the method for constructing the mycoplasmal pneumonia severity prediction model used in the present application further comprises the following steps after step 5):

[0020] 6) using the verification set in step 3) to verify the multi-factor logistic regression model obtained in step 5), and determining the verification result as the mycoplasmal pneumonia severity prediction model.

[0021] The mycoplasmal pneumonia severity prediction model constructed by the method for constructing the mycoplasmal pneumonia severity prediction model as described above.

[0022] Preferably, the mycoplasmal pneumonia severity prediction model used in the present application is a multi-factor logistic regression model, and the expression of the mycoplasmal pneumonia severity prediction model is:

[0023] Ln = -3.355 × a1 + 0.682 × a2 + 1.170 × a3 + 3.177 × a4 + 78.416 × a5

[0024] wherein:

[0025]

[0026]

[0027]

[0028]

[0029]

[0030] wherein:

[0031] AGE is the age of the patient infected with mycoplasmal pneumonia, and the unit is month;

[0032] EOS is the number of eosinophils;

[0033] MONO is the number of monocytes;

[0034] NEU is the number of neutrophils;

[0035] S100A8 / A9 is the content of S100A8 / A9 protein in the serum sample of the patient infected with mycoplasmal pneumonia.

[0036] The application of the mycoplasmal pneumonia severity prediction model as described above in predicting the severity of mycoplasmal pneumonia, especially in predicting the severity of mycoplasmal pneumonia in children.

[0037] Compared with the prior art, the present application has the following advantages:

[0038] The present application mines a new marker S100A8 / A9 for grading diagnosis of mycoplasma pneumonia severity in children, and combines commonly used clinical inflammatory evaluation indicators to provide a construction method of a high-efficiency, high-specificity, high-sensitivity and high-accuracy pneumonia severity prediction model and the prediction model. Meanwhile, the application of S100A8 / A9 to the degree of illness of children with mycoplasma pneumonia has good clinical applicability; the prediction model provided by the present application is easier to promote in clinical practice, optimizes the mycoplasma pneumonia disease stratification and clinical diagnosis and treatment path, is conducive to early identification of severe and critical cases and high-risk groups of sequelae, is conducive to guiding individualized treatment; can be used to guide the selection of clinical treatment drugs for children with mycoplasma pneumonia, prevent the occurrence of severe cases and provide evidence-based basis for reducing mortality. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 is a flowchart of the construction method of the pneumonia severity prediction model provided by the present application;

[0040] Figure 2 is a basic data analysis related graph of 168 cases of mycoplasma pneumonia children;

[0041] Figure 3 is a correlation of each index of 168 cases of mycoplasma pneumonia patients and a training set prediction grouping related graph;

[0042] Figure 4 is a Lasso regression analysis and random forest verification related graph of 168 cases of mycoplasma pneumonia patients;

[0043] Figure 5 is a training set model performance evaluation related graph;

[0044] Figure 6 is a training set model visualization verification related graph;

[0045] Figure 7 is a distribution and prediction related graph of the verification set sample;

[0046] Figure 8 is a linear regression equation of the standard curve in Example 1. DETAILED DESCRIPTION

[0047] In order to make the person in the art better further understand the present application scheme, the characteristics and advantages of the present application are further described by combining with the embodiments and the drawings. Obviously, the described embodiments are only a part of the explanation of the present application, not all the embodiments.

[0048] The application provides an application of S100A8 / A9 protein content in constructing a mycoplasma pneumonia severity prediction model, and in particular, an application of S100A8 / A9 protein content in constructing a mycoplasma pneumonia severity prediction model for children.

[0049] Based on the application as described above, the application provides a method for constructing a mycoplasma pneumonia severity prediction model, which comprises the following steps:

[0050] 1) Obtain serum samples of patients infected with mycoplasma pneumonia; wherein the inclusion criteria for patients infected with mycoplasma pneumonia are: a) clinical diagnosis of pneumonia: respiratory symptoms and signs such as fever, cough, expectoration, tachypnea, lung rales, and chest imaging showing lung involvement; b) evidence of mycoplasma pneumonia infection: positive serum specific MP-IgM antibody and / or positive sputum MP-RNA and / or tNGS suggesting mycoplasma pneumonia infection. The exclusion criteria are: a) symptoms of respiratory infectious diseases, and clinical diagnosis of non-mycoplasma pneumonia: clear evidence of infection with other pathogens as the main cause, such as viruses, Streptococcus pneumoniae, and tuberculosis infection; b) non-infectious pneumonia, such as bronchial foreign body, primary ciliary immotility syndrome, allergic asthma, cystic granulomatous inflammation, etc.; c) underlying diseases, such as tumors, hematological diseases, endocrine system diseases, congenital immunodeficiency diseases, congenital heart diseases, etc.; d) congenital dysplasia or malformation, such as tracheal dysplasia, pulmonary artery sling, etc.; e) patients within 3 months after surgery; f) history of repeated respiratory tract infections within the past 3 months, with the number of respiratory tract infections ≥3 times. In addition, the clinical data collection of mycoplasma pneumonia patients includes patient age and gender; laboratory indicators include C-reactive protein, high-sensitivity C-reactive protein, procalcitonin, white blood cell count, neutrophil count, monocyte count, and eosinophil count.

[0051] 2) Detect the content of S100A8 / A9 protein in the serum sample in step 1), specifically: collect the patient's venous blood and store it in a test tube, place it at room temperature 25℃ for 10 minutes, centrifuge the venous blood in a centrifuge at a centrifugal force of 3000g for 5 minutes to obtain centrifuged blood, and immediately store the serum at -80℃ until determination. The content of S100A8 / A9 protein in the serum sample is detected using an S100A8 / A9 ELISA kit, and the specific detection method is: a) reagent preparation: the S100A8 / A9 ELESA kit is warmed at room temperature for 120 min before use; concentrated washing solution is configured (concentrated washing solution is diluted with distilled water at a ratio of 1:20); substrate solution A and B are mixed at a volume ratio of 1:1 before use, and the mixture is used within 15 min;

[0052] b) Perform a one-in-three determination for each standard or sample according to the standard curve of the recombinant S100A8 / A9 protein dilution. The detection limit is 0.01 ng / mL; the coefficient of variation between determination and determination is 4.5%-10% and 5%-15% respectively for the concentration range of 0.125-4 ng / mL.

[0053] c) Freeze-thaw the sample to be tested at 4°C, prepare the required items for the experiment according to the ELISA kit, prepare the washing buffer according to the proportion, set up double-hole detection for the standard, and set up water blank control. According to the principle of double-antibody sandwich method, detect the sample to be tested according to the kit operation steps. Read the OD value of the standard and the sample to be tested at 450 nm wavelength with an enzyme-labeled instrument, draw a standard curve according to the standard concentration and obtain a linear regression equation, and calculate the concentration of each sample to be tested.

[0054] 3) According to step 1), a data set of patients infected with mycoplasma pneumonia is established, including a training set and a validation set; the ratio of the training set to the validation set is 7:3; the training set includes a mild training set and a severe training set; the validation set includes a mild training set and a severe validation set; wherein the mycoplasma pneumonia severe data set: meets the severe or critical standard; any one of the following clinical manifestations meets the severe standard: a) persistent high fever (above 39°C) for more than 5 days or fever for more than 7 days, with no downward trend in the peak of body temperature; b) one of the following: wheezing, shortness of breath, dyspnea, chest pain, hemoptysis, etc. These manifestations are related to severe lesions, combined with plastic bronchitis, asthma attack, pleural effusion and pulmonary embolism, etc.; c) pulmonary complications occur, but do not meet the critical standard; d) in a resting state, the oxygen saturation of the finger pulse is less than or equal to 0.93 when inhaling air; e) one of the following imaging manifestations: e1) a single lung lobe is more than 2 / 3 involved, with uniform and consistent high-density consolidation or 2 or more lung lobes with high-density consolidation (regardless of the size of the involved area), which can be accompanied by moderate to large pleural effusion, or can be accompanied by localized bronchiolitis; e2) diffuse or bilateral lung lobes have bronchiolitis, which can be combined with bronchiolitis and have mucus plug formation leading to atelectasis. f) Clinical symptoms are progressively worsening, and imaging shows that the lesion range has progressed by more than 50% in 24-48h; g) one of CRP, LDH, D-dimer is significantly elevated. The critical performance is: there is respiratory failure and / or severe life-threatening pulmonary complications, and mechanical ventilation or other life support is required. The rest of the cases are classified into a mild data set.

[0055] 4) Use lasso regression to construct a regression model for the S100A8 / A9 protein content obtained in step 2), specifically:

[0056] 4.1) Use lasso regression to establish a lambda value under model for the S100A8 / A9 protein content obtained in step 2);

[0057] 4.2) Screening the variables of the model obtained in step 4.1);

[0058] 4.3) Verifying the variables obtained in step 4.2) using a random forest model to obtain a regression model.

[0059] Exemplarily, the specific construction of step 4) can be according to the following: introducing the clinical data of the children and the content of S100A8 / A9 protein in the serum sample into R software; variable screening, i.e. using lasso regression to establish models under different lambda values (min and 1se), and comparing the prediction values of the two models in the training set, selecting the more valuable min model, and analyzing the variables in the model at this time. Variable screening verification, i.e. using a random forest model to compare the contribution value (Mean decrease gini coefficient) of each index to the model in lasso regression.

[0060] 5) Training the regression model in step 4) using the training set in step 3) to obtain a multi-factor logistic regression model; the multi-factor logistic regression model is a low-order version of the mycoplasma pneumonia severity prediction model, and exemplarily, the training method is to use the 5 variables (S100A8 / A9, NEU, AGE, MONO, EOS) screened by lasso, use the training set, take the severity of mycoplasma pneumonia in children as the dependent variable Y, and take the 5 screening indexes as the independent variable X to establish a multi-factor logistic regression model:

[0061] and the results are displayed using a nomograph.

[0062] 6) Verifying the multi-factor logistic regression model obtained in step 5) using the verification set in step 3), and determining the verification result as the mycoplasma pneumonia severity prediction model, exemplarily, the verification process is: using the established logistic regression model to display the ROC curve and AUC of the model in the training set, and comparing S100A8 / A9, Others and Complex, drawing a DCA (Decision Curve Analysis) curve and a clinical impact curve (Clinical Impact Curve) in the verification set to show the prediction value. The established multi-factor logistic regression model is used for verification set test, and a heat map is used for effect display.

[0063] Based on the foregoing construction method, the present application also obtains a mycoplasma pneumonia severity prediction model, which is a multi-factor logistic regression model, and the expression of the mycoplasma pneumonia severity prediction model is:

[0064] Ln = -3.355 * a1 + 0.682 * a2 + 1.170 * a3 + 3.177 * a4 + 78.416 * a5

[0065] wherein:

[0066]

[0067]

[0068]

[0069]

[0070]

[0071] wherein:

[0072] AGE is the age of the patient infected with mycoplasma pneumonia, in units of: month;

[0073] EOS is the number of eosinophils;

[0074] MONO is the number of monocytes;

[0075] NEU is the number of neutrophils;

[0076] S100A8 / A9 is the content of S100A8 / A9 protein in the serum sample of the patient infected with mycoplasma pneumonia.

[0077] The application of the mycoplasma pneumonia severity prediction model in predicting the severity of mycoplasma pneumonia, especially in predicting the severity of mycoplasma pneumonia in children.

[0078] The technical solutions provided by the application will be described in detail below in combination with the drawings:

[0079] The kits used in the following examples are as follows:

[0080] Human S100 calcium-binding protein A8; A9 complex (S100A8; A9) quantitative detection kit (ELISA)

[0081] Product item number: HY10208

[0082] Example 1: Determination of the content of S100A8 / A9 protein in the blood samples of 168 children with mycoplasma pneumonia

[0083] Operation steps

[0084] 1) Reagent preparation: The S100A8 / A9 ELESA kit should be brought to room temperature for 120 min before use; Washing buffer preparation (concentrated washing buffer and distilled water, diluted at 1:20);

[0085] 2) Sample incubation: Set up standard wells, 0-value wells, blank wells, and sample wells. Add 50 μL of standard at different concentrations to each standard well, add 50 μL of sample diluent to the 0-value wells, do not add any diluent to the blank wells, and add 50 μL of the sample to be tested to the sample wells. Except for the blank wells, add 100 μL of horseradish peroxidase (HRP)-labeled detection antibody to the standard wells, 0-value wells, and sample wells. Cover with sealing film and incubate at 37°C for 60 min.

[0086] 3) Cleaning: Pour out the liquid in the microplate, flatten it on absorbent paper, fill each well with diluted washing solution, let it stand for 30 seconds, shake off the washing solution, and pat it dry on absorbent paper. Repeat this process 5 times.

[0087] 4) Substrate incubation: Mix substrates A and B thoroughly at a volume ratio of 1:1, add 100 μL of substrate mixture to each microwell, and incubate at 37°C for 15 min.

[0088] 5) Termination of reaction: Add 50 μL of the termination solution to each microwell at the same rate and in the same order as the substrate solution was added;

[0089] 6) Colorimetric analysis: Perform colorimetric analysis within 30 minutes after adding the stop solution. Before colorimetric analysis, gently shake the microplate to ensure uniform diffusion of the liquid. Read the absorbance (OD value) of each well at 450 nm on the microplate reader.

[0090] Result Calculation

[0091] A standard curve was constructed with the concentration of the standard samples as the ordinate (6 standard wells plus 1 zero-value well, for a total of 7 concentration points) and the corresponding absorbance (OD value) as the abscissa, using the linear regression equation Y = 162.94X. 2 +1610.3X+65.747, where R 2 =0.9968, the regression equation is as follows Figure 8 As shown in Tables 1 and 2, the concentration of the sample was calculated using an equation based on its absorbance (OD value). The detection limit of this kit is 0.01 ng / mL; for a concentration range of 0.125 to 4 ng / mL, the coefficients of variation within and between assays were 4.5%–10% and 5%–15%, respectively.

[0092] Table 1. Concentration of Standards and Corresponding OD Values

[0093] Concentration (ng / mL) 0 0.125 0.25 0.5 1 2 4 OD value 0.033833 0.054667 0.089833 0.193500 0.520000 1.121833 2.018833 Corrected OD value — 0.020834 0.056000 0.159667 0.486167 0.088000 1.985000

[0094] Table 2 S100A8 / A9 protein content in serum samples of 168 children with mycoplasma pneumonia

[0095] Sample No. S100A8 / A9 (ng / ml) Sample No. S100A8 / A9 (ng / ml) Sample No. S100A8 / A9 (ng / ml) Sample No. S100A8 / A9 (ng / ml) 1 0.79 43 0.36 85 0.32 127 0.94 2 0.70 44 0.70 86 0.22 128 0.21 3 0.48 45 1.28 87 3.79 129 0.54 4 0.39 46 1.52 88 0.16 130 0.57 5 0.63 47 0.64 89 0.16 131 0.59 6 0.30 48 0.47 90 0.60 132 0.52 7 0.53 49 0.33 91 1.24 133 0.34 8 0.24 50 1.49 92 0.47 134 0.51 9 1.03 51 0.47 93 1.54 135 0.83 10 0.43 52 3.65 94 1.76 136 0.78 11 1.19 53 0.80 95 3.44 137 0.32 12 0.55 54 0.16 96 0.68 138 0.54 13 1.05 55 0.54 97 0.71 139 1.06 14 4.04 56 1.03 98 0.77 140 0.62 15 0.24 57 3.16 99 0.13 141 0.55 16 0.53 58 0.76 100 0.54 142 0.65 17 3.37 59 2.68 101 1.20 143 0.25 18 1.06 60 2.33 102 0.95 144 0.72 19 0.21 61 1.71 103 0.20 145 0.52 20 1.06 62 0.33 104 0.72 146 0.96 21 0.91 63 3.99 105 0.35 147 1.12 22 1.31 64 0.84 106 0.52 148 0.99 23 0.86 65 0.92 107 0.39 149 0.77 24 0.47 66 3.29 108 0.74 150 2.45 25 0.58 67 0.31 109 0.11 151 0.13 26 0.32 68 0.45 110 0.75 152 0.71 27 0.33 69 0.30 111 0.78 153 0.50 28 1.25 70 0.18 112 0.24 154 0.43 29 0.40 71 0.84 113 0.71 155 0.68 30 0.32 72 0.87 114 0.22 156 0.25 31 2.08 73 0.17 115 0.70 157 0.85 32 0.21 74 3.84 116 0.55 158 0.49 33 0.39 75 1.84 117 0.49 159 0.46 34 1.16 76 0.35 118 0.21 160 1.59 35 0.80 77 0.47 119 0.31 161 0.61 36 0.25 78 0.34 120 1.48 162 0.68 37 0.31 79 0.23 121 0.21 163 0.41 38 2.27 80 0.62 122 1.20 164 0.29 39 0.38 81 0.39 123 0.74 165 0.24 40 0.75 82 1.16 124 0.19 166 0.49 41 0.37 83 0.58 125 0.54 167 0.73 42 0.22 84 0.71 126 0.78 168 0.46

[0096] Example 2: Analysis and grouping of general clinical characteristics of 168 patients with mycoplasma pneumonia

[0097] Analysis method: Collect the clinical indicators of patients, including demographic characteristics, clinical manifestations, biochemical indicators, etc. Based on single-center clinical research, 168 children with mycoplasma pneumonia were recruited from Wuhan Children's Hospital in the present application. The ratio of mild to severe in the present application is 1:1.3 (73 vs 95), and the ratio of male to female is 1:0.61 (104 vs 64). The original data table of the age of 168 patients with mycoplasma pneumonia was established, and the segmentation point was established. In the present application, the distribution of patients in 0-12, 13-24, 25-36, 37-48, 49-60, 61-72, 73-84, 85-96, 97-108, 109-120, 121-132, 133-144, 145-156, 157-168 age groups (age based on month) was analyzed. The difference in the distribution of S100A8 / A9 in different age groups and different genders was compared through probability density curve, scatter plot, heat map, and further comparison of the correlation of each index of the patients.

[0098] Results: The main demographic and clinical characteristics of the study population are shown in Table 3. The results of the remaining analysis are shown in Figure 2 Figure 2 a The age distribution histogram of the patients suggests that the age of the patients is skewed distribution, and the median distribution is 62 [38.50, 95.50], Figure 2 b Box plot distribution of each index of patients with different disease severity, suggesting that the CRP, WBC, NEU, S100A8 / A9 index distribution of patients with severe disease is higher than that of patients with mild disease, and the difference is statistically significant. Figure 2 c Scatter plot and correlation analysis of S100A8 / A9 index distribution in patients of different ages, suggesting that there is no correlation between S100A8 / A9 index and age (r = -0.05, p = 0.52), and there is no statistical difference in the distribution of S100A8 / A9 index between male and female patients (U = 0.90, p = 0.37). Figure 2 d Probability density distribution of S100A8 / A9 index in male and female patients, suggesting that S100A8 / A9 index is left-skewed distribution, and the median distribution is 0.6 [0.3, 0.9], the peak value of male and female patients is around 0.5, and the distribution pattern is consistent.

[0099] Table 3 Main demographic and clinical characteristics of the study population

[0100] ​ Characteristic Overall Mild (n=73) Severe (n=95) p Gender Male 104 47 57 0.562 Female 64 26 38 Age (month) 62[38.25,95.75] 53[32,100] 65[42,94] 0.46 CRP 13.88[6,14.77] 10[4,14.77] 14.02[7,16] 0.07 hsCRP 18.63[5.6,18.63] 14.3[2.82,18.63] 18.63[8.39,21.7] <0.01 PCT 0.12[0.07,0.26] 0.09[0.06,0.17] 0.14[0.08,0.3] <0.01 WBC 7.68[5.73,10.7] 6.92[5.57,10.03] 8.16[5.82,11.79] 0.23 NEU 4.53[3.07,6.65] 4.03[2.72,5.89] 5.01[3.36,7.2] <0.05 MONO 0.62[0.4,0.8] 0.6[0.4,0.78] 0.63[0.41,0.81] 0.68 EOS 0.08[0.03,0.18] 0.08[0.03,0.15] 0.08[0.02,0.21] 0.98 S100A8 / A9 0.6[0.35,0.95] 0.34[0.24,0.47] 0.85[0.7,1.31] <0.01

[0101] Example 3: Establishment of diagnostic dataset and correlation and predictive grouping analysis of various indicators

[0102] A dataset was created based on the severity of all cases (mild or severe). The dataset was then divided into a training set and a validation set in a 7:3 ratio. The differences in variables between the training set and the validation set were analyzed using the chi-square test and the t-test, respectively. The basic characteristics of each variable in the training set and the validation set are shown in Table 4.

[0103] Table 4. Basic information of cases in the training and testing cohorts.

[0104]

[0105] Using R language, correlation calculations were performed and correlation heatmaps with significant stars were plotted to analyze the correlation between various indicators. The predicted grouping of each indicator was performed with the actual group as the dependent variable Y and each indicator as the independent variable X, using univariate logistic regression.

[0106] The results are as follows Figure 3 As shown, Figure 3 The correlation distribution plot of various indicators in patients (a) shows that there is a strong positive correlation among NEU, MONO, and WBC, and all three are strongly positively correlated with EOS. NEU is weakly positively correlated with CRP and S100A8 / A9. Figure 3 b is a heatmap showing the actual values ​​of each indicator and the predicted groupings for different disease severity levels in the training set patients. It indicates that only indicator S100A8 / A9 has a relatively good predictive effect.

[0107] Example 4: Screening and validation of variables for predicting the severity of mycoplasma pneumonia patients

[0108] Based on the initial variable screening, lasso regression was used to evaluate the diagnostic value of the eight variables for severe mycoplasma pneumonia. (See Table 5 and...) Figure 4 As can be seen, the number of variables included varies between 3 and 5 under the two lambda value modes (min / 1se). Further comparison of the classification performance on the training set shows that there is no significant difference in the effectiveness of the included variables in disease discrimination regardless of whether the lambda value is min or 1se (p>0.05). Analysis of the box plots predicted under different lambda value modes reveals that the model includes more variables under the min mode, with the final variables AGE, NEU, MONO, EOS, and S100A8 / A9 included in the model.

[0109] We further used the Mean decrease gini coefficient (node ​​purity) in the random forest model to evaluate the contribution of the included variables to the diagnostic model.Figure 4 c, 4d are the importance distribution of each index under the random forest model; the results show that the contribution of each variable to the disease severity model in descending order is: S100A8 / A9, NEU, AGE, MONO, EOS. The Mean decrease gini coefficient in the random forest model indicates that the contribution of S100A8 / A9 is more than 30%, and the contribution of the remaining indicators is less than 10%.

[0110] The results are shown in Figure 4 Figure 4 a is the change of the included variables under different lambda values in lasso regression, which shows the change of lambda value and the number of included variables in lasso regression; Figure 4 b is the box plot of the prediction effect under different lambda value modes; Figure 4 c, 4d are the importance distribution of each index under the random forest model.

[0111] Table 5 Lasso regression variable selection under different modes

[0112] Variable Min 1se Intercept -5.85 -3.68 AGE 0.01 0.00 Gender . . CRP . . PCT . . WBC . . NEU 0.12 0.05 MONO -0.17 . EOS 0.12 . S100A8A9 8.74 5.80

[0113] Example 5: Regression model for predicting the severity of mycoplasma pneumonia patients;

[0114] Analysis method: five variables in Table 5 lasso regression (min) are selected as independent variables, and a multivariate logistic regression model is established in the training set. Table 6 is the logistic regression model.

[0115] Table 6 Logistic regression model

[0116] Variable OR (95% CI) p AGE 1.01(1.00,1.03) 0.16 NEU 1.24(1.01,1.57) 0.04 MONO 0.47(0.07,3.18) 0.44 EOS 1.19(0.03,64.04) 0.93 S100A8A9 155830(4689.80,14968000.00) 0.00

[0117] Figure 6 a is the evaluation of the prediction effect of the training set multivariate logistic regression model using the ROC curve, and the results show that the AUC area under the curve is 94.6%, and the sensitivity and specificity are 92.9% and 95.5% when the best cutoff point is selected Y >= 0; Figure 6 b is the Auc area of each index evaluated separately in the training set using ROC, combined with the random forest results, to further evaluate the performance of each index for disease evaluation. The results show that S100A8 / A9 0.95 (0.91, 0.99), NEU 0.65 (0.55, 0.75), EOS 0.51 (0.40, 0.61), MONO 0.60 (0.50, 0.71), and the Auc areas of the four indicators are tested by delong method. The index S100A8 / A9 is significantly higher than the other three indexes, and the difference is statistically significant.​

[0118] To further facilitate clinical application, ROC curves were established to predict the severity of mycoplasma pneumoniae pneumonia independently for each indicator. Based on the closest-to-left method, the cutoff point for the optimal predictive effect of each indicator was determined, thereby transforming continuous indicators into categorical variables. The established regression model is as follows:

[0119]

[0120] Using the maximum likelihood estimation method, the risk scores for each indicator are calculated as follows:

[0121] Ln (Severe / Mild Cases) = -3.355×a1 + 0.682×a2 + 1.170×a3 + 3.177×a4 + 78.416×a5

[0122] Ln(severe / mild) is the Logit transformation of the probability ratio of a child's condition to that of a child with a mild condition.

[0123]

[0124]

[0125]

[0126]

[0127]

[0128] In the logistic regression model, the p-value of the indicator S100A8 / A9 is less than 0.05. The prediction effects of each indicator after transformation are shown in Table 7. When S100A8 / A9 is applied alone, it has a good prediction effect.

[0129] Table 7. Individual Prediction Effects of Each Indicator After Conversion to Categorical Variables

[0130]

[0131]

[0132] Example 6: Demonstration of the predictive performance of the S100A8 / A9 childhood mycoplasma pneumonia severity prediction model

[0133] The predictive efficacy of the model was analyzed using nomograms, clinical decision curves, and clinical impact curves. Figure 6 a is the nomogram of the results of the multivariate logistic regression; Figure 6b is the decision curve analysis (DCA) graph of different variables into the mode. The results show that when the threshold is >0.1 (i.e. the proportion of severe cases in the test population is >10%), the benefits of the complex model and the single S100A8 / A9 model (Ln(severe / mild) = 78.416 x a5) are consistent, and both are better than the single other indicator (Others) model and the net benefit (true positive rate-false positive rate) is more than 0.1. It shows that the single S100A8 / A9 has good clinical application value for disease prediction. Figure 6 c is the clinical impact curve (CIC) graph. The horizontal coordinate of the graph is the probability threshold, and the vertical coordinate is the number of people. The purple line represents the number of people determined as high risk by the model at different probability thresholds; the red line represents the number of people determined as high risk by the model at different probability thresholds and actually having an outcome event. The clinical impact curve results show that at different risk thresholds, the loss: benefit ratio of the model's confirmed positive (red) and true positive (blue) is as follows: when the incidence threshold reaches 0.6 or more, the model's prediction value is close to the actual severe condition probability. Figure 7 a is a heat map of the actual value and predicted grouping of each indicator in patients with different conditions in the validation set. The results show that S100A8 / A9 and Complex mode have similar prediction effects in patients with different conditions in the validation set. Combined with Table 8, the results show that the accuracy of S100A8 / A9 alone reaches 90%, and the accuracy of the Complex mode combined with 4 indicators (AGE, NEU, MONO, EOS) has limited improvement (92% vs 90%, no statistically significant difference).

[0134] Table 8 Sensitivity and specificity of the constructed model

[0135] Model Sensitivity Specificity Kappa Yoden Accuracy Others 50.00% 71.43% 0.22 21.43% 62% S100A8A9 90.91% 89.29% 0.80 80.19% 90% Complex 86.36% 96.43% 0.84 82.79% 92%

[0136] The three models are as follows:

[0137] Complex Model

[0138] Ln(severe / mild) = -3.355 x a1 + 0.682 x a2 + 1.170 x a3 + 3.177 x a4 + 78.416 x a5

[0139] S100A8 / A9 Model

[0140] Ln(severe / mild) = 78.416 x a5

[0141] Others Model

[0142] Ln (severe / mild) = -3.355 x a1+ 0.682 x a2+ 1.170 x a3+ 3.177 x a4

[0143]

[0144]

[0145]

[0146]

[0147]

[0148] The above descriptions are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. Application of S100A8 / A9 protein content in constructing a prediction model for the severity of pneumonia.

2. Application of S100A8 / A9 protein content in constructing a prediction model for the severity of mycoplasma pneumonia.

3. Application of S100A8 / A9 protein content in constructing a prediction model for the severity of mycoplasma pneumonia in children.

4. A method for constructing a Mycoplasma pneumoniae severity prediction model, the method comprising: The method for constructing the prediction model for the severity of mycoplasma pneumonia comprises the following steps: ​ 1) Obtain serum samples from patients infected with mycoplasma pneumonia; 2) Detect the content of S100A8 / A9 protein in the serum samples in step 1); 3) Establish a data set for patients infected with mycoplasma pneumonia according to step 1), which includes a training set and a validation set; the ratio of the training set to the validation set is 7:3; the training set includes a mild training set and a severe training set; the validation set includes a mild training set and a severe validation set; 4) Construct a regression model for the content of S100A8 / A9 protein obtained in step 2) using lasso regression; 5) Train the regression model in step 4) using the training set in step 3) to obtain a multi-factor logistic regression model; the multi-factor logistic regression model is a low-order version of the prediction model for the severity of mycoplasma pneumonia.

5. The method of claim 4, wherein the Mycoplasma pneumonia severity prediction model is constructed by: The specific implementation of step 4) is: 4.1) Use lasso regression to establish a lambda value model for the content of S100A8 / A9 protein obtained in step 2); 4.2) Screen the variables of the model obtained in step 4.1); 4.3) Use a random forest model to verify the variables obtained in step 4.2) to obtain a regression model.

6. The method of claim 4 or 5, wherein the method further comprises: determining the severity of the mycoplasmal pneumonia of the patient based on the determined value of the at least one biomarker. The method for constructing the prediction model for the severity of mycoplasma pneumonia further comprises the following steps after step 5): 6) Verify the multi-factor logistic regression model obtained in step 5) using the validation set in step 3), and determine the verification result as the prediction model for the severity of mycoplasma pneumonia.

7. The prediction model for the severity of mycoplasma pneumonia constructed by the method according to any one of claims 4-6.

8. The Mycoplasma pneumoniae severity prediction model of claim 7, wherein: The prediction model for the severity of mycoplasma pneumonia is a multi-factor logistic regression model, and the expression of the prediction model for the severity of mycoplasma pneumonia is: Ln = -3.355 × a1 + 0.682 × a2 + 1.170 × a3 + 3.177 × a4 + 78.416 × a5 wherein: AGE is the age of the patient infected with mycoplasma pneumonia, in months; EOS is the number of eosinophils; MONO is the number of monocytes; NEU is the number of neutrophils; S100A8 / A9 is the content of S100A8 / A9 protein in the serum sample of the patient infected with mycoplasma pneumonia.

9. Application of the prediction model for the severity of mycoplasma pneumonia according to claim 7 in predicting the severity of mycoplasma pneumonia.

10. Application of the prediction model for the severity of mycoplasma pneumonia according to claim 7 in predicting the severity of mycoplasma pneumonia in children. ​