A macrolide-resistant Mycoplasma pneumoniae pneumonia prediction model and its construction method

By constructing a macrolide-resistant mycoplasma pneumoniae pneumonia prediction model, and utilizing multi-dimensional clinical data and machine learning algorithms, the problem of early identification of high-risk patients for MUMPP was solved, enabling early adjustment of treatment strategies and reduction of RMPP risk, thus improving prediction accuracy and clinical applicability.

CN120809258BActive Publication Date: 2026-04-03SHANGHAI CHILDRENS MEDICAL CENT AFFILIATED TO SHANGHAI JIAOTONG UNIV SCHOOL OF MEDICINE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Current technologies cannot accurately identify high-risk patients with macrolide-resistant Mycoplasma pneumoniae pneumonia (MUMPP) in the early stages of hospital admission, leading to a delay in treatment strategy adjustments and increasing the risk of refractory Mycoplasma pneumoniae pneumonia (RMPP). Furthermore, existing predictive models fail to adequately consider treatment modalities and drug resistance information, resulting in insufficient generalization ability and predictive accuracy.

Method used

A macrolide-resistant mycoplasma pneumoniae pneumonia (MUMPP) prediction model was constructed. By acquiring multi-dimensional clinical data from patients, systematic preprocessing and feature selection were performed. Combined with ensemble learning and deep learning algorithms, the model parameters were optimized to achieve early identification of high-risk MUMPP patients.

Benefits of technology

Accurately identifying high-risk patients for MUMPP within 24 hours of admission and making a diagnosis 48 hours earlier reduces the risk of RMPP, provides quantitative basis for treatment decisions, and improves predictive accuracy and clinical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809258B_ABST
    Figure CN120809258B_ABST
Patent Text Reader

Abstract

This invention discloses a predictive model and construction method for macrolide-resistant Mycoplasma pneumoniae pneumonia (MUMPP), belonging to the field of artificial intelligence technology. It aims to address the problem of delayed identification of MUMPP in clinical practice. The method includes: constructing a standardized dataset; preprocessing the data, including handling missing values ​​using chain equation multiple imputation and converting continuous variables into categorical variables based on clinical significance or ROC curves; training the model using random forest and neural network models, and selecting the optimal feature subset using recursive feature elimination, and optimizing the model hyperparameters using a grid search algorithm; finally, evaluating the model from multiple dimensions including discriminative power, calibration, and clinical effectiveness. This invention, through a systematic modeling process, constructs a highly accurate predictive model capable of early warning of MUMPP, allowing clinicians to adjust treatment strategies promptly and prevent disease progression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a macrolide-based non-reactive Mycoplasma pneumoniae pneumonia prediction model and its construction method. Background Technology

[0002] Mycoplasma pneumoniae (MP) is a common pathogen of community-acquired pneumonia (CAP) in children, and is also frequently seen in infants aged 1 to 3 years. Mycoplasma pneumoniae pneumonia (MPP), as a common type of CAP in children, has significant public health implications in clinical practice.

[0003] Macrolides are currently the first-line drugs recommended by domestic and international guidelines for the treatment of mycoplasma pneumoniae (MPP). However, with the widespread use of macrolides, an increasing number of cases of macrolide-unresponsive mycoplasma pneumoniae pneumonia (MUMPP) have emerged in clinical practice. MUMPP refers to MPP patients who, after 72 hours of standard macrolide treatment, still experience persistent fever, and whose clinical signs and lung imaging show no improvement or even further deterioration. Currently, the incidence of MUMPP has exceeded 30% in some regions and is showing an increasing trend year by year.

[0004] The emergence of MUMPP poses a serious challenge to clinical treatment. If MUMPP is not treated properly, approximately 20% of patients may develop refractory mycoplasma pneumoniae pneumonia (RMPP). RMPP is defined as MPP patients who, despite regular treatment with macrolide antibiotics for 7 days or more, continue to experience fever, worsening clinical signs and lung imaging findings, and the development of extrapulmonary complications. Compared to non-RMPP patients, RMPP patients have more severe clinical symptoms, a significantly higher risk of serious extrapulmonary complications, require longer hospital stays, have a higher proportion of ICU admissions, and in severe cases, may even die. Recent clinical evidence indicates that antibiotic treatment failure has become one of the important reasons for the increased mortality rate in children with MP infection.

[0005] Currently, the identification of MUMPP in clinical practice mainly relies on post-treatment clinical observation. The vast majority of MUMPP cases are diagnosed only three days after the initiation of macrolide antibiotic treatment. This diagnostic lag severely impacts the timely adjustment of treatment strategies. The later the treatment strategy for MUMPP is adjusted, the higher the risk of the patient developing RMPP and the worse the prognosis. Therefore, accurately identifying high-risk MUMPP patients at the initial stage of treatment has become a crucial clinical problem that urgently needs to be solved.

[0006] Although researchers have attempted to build predictive models for RMPP in recent years, existing technologies still have many shortcomings. First, most existing predictive models do not fully consider the impact of treatment methods on disease outcome variables, limiting their predictive accuracy. Second, due to limitations in sample size and algorithms, the generalization ability and predictive accuracy of existing models need improvement. Third, existing models mostly remain at the level of risk prediction, lacking substantial guidance for clinical treatment decisions. Most importantly, there is currently no research specifically on building models for early prediction of MUMPP, failing to meet the urgent clinical need for early identification and intervention of MUMPP.

[0007] Macrolide resistance is a significant cause of MUMPP. The increasing prevalence of MP-resistant strains further expands the MUMPP patient population and increases the risk of RMPP. However, existing predictive methods fail to adequately integrate resistance-related information and effectively utilize multidimensional clinical data from patient admission for comprehensive analysis. Summary of the Invention

[0008] The purpose of this invention is to provide a predictive model for macrolide-resistant Mycoplasma pneumoniae pneumonia and its construction method, which can accurately identify high risk of MUMPP in the early stage of patient admission (e.g., within 24 hours), providing a scientific basis for clinicians to adjust treatment strategies in a timely manner and prevent disease progression to RMPP, thereby improving the prognosis of children and reducing the medical burden.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] A method for constructing a predictive model for macrolide-resistant Mycoplasma pneumoniae pneumonia, comprising:

[0011] S1: Obtain clinical data of patients with Mycoplasma pneumoniae pneumonia, screen predictive variables, and construct the original dataset;

[0012] S2: Preprocess the original dataset, including imputing missing data, classifying continuous variables, and encoding categorical variables to obtain a standardized dataset;

[0013] S3: Based on the standardized dataset, a prediction model is constructed using machine learning algorithms, and the optimal prediction model is obtained through feature selection and parameter optimization;

[0014] S4: The performance of the optimal prediction model is evaluated to obtain a macrolide-based non-reactive Mycoplasma pneumoniae pneumonia prediction model.

[0015] Further: In step S1, the predictive variables include at least two categories of demographic variables, clinical symptom variables, laboratory test variables, and imaging variables; the acquisition of clinical data includes: extracting patient data that meet the diagnostic criteria for Mycoplasma pneumoniae pneumonia from a community-acquired pneumonia database, and dividing patients into a macrolide non-responsive group and a macrolide responsive group based on their clinical response 72 hours after macrolide treatment.

[0016] Furthermore: the missing data interpolation process in step S2 adopts the chain equation multiple interpolation method;

[0017] The classification transformation of the continuous variables includes:

[0018] For continuous variables with clinically recognized thresholds, they are directly converted into categorical variables based on the clinical thresholds.

[0019] For continuous variables without a clear clinical threshold, the optimal cut-off value is determined through receiver operating characteristic (ROC) curve analysis, and the variable is then converted into a categorical variable based on the optimal cut-off value.

[0020] Further: The determination of the optimal cut-off value through receiver operating characteristic curve analysis includes: taking macrolide-resistant Mycoplasma pneumoniae pneumonia as the outcome variable, calculating the Youden index corresponding to each possible value of the continuous variable, and selecting the value that maximizes the Youden index as the optimal cut-off value; wherein, Youden index = sensitivity + specificity - 1.

[0021] Furthermore: the categorical variable encoding process in step S2 employs one-hot encoding, specifically including:

[0022] When a categorical variable has only one option, it is directly converted into a binary variable;

[0023] When a categorical variable contains n options and n≥2, if the n options are mutually exclusive, they are converted into n independent binary variables; if the n options are not mutually exclusive, each option is split into a separate column and converted into a binary variable.

[0024] Furthermore: the machine learning algorithm in step S3 includes ensemble learning algorithm and / or deep learning algorithm; the feature selection adopts recursive feature elimination method, and the performance of different feature subsets is evaluated by cross-validation to determine the optimal feature subset; the parameter optimization adopts grid search algorithm to systematically optimize the hyperparameters of the model.

[0025] Furthermore: the ensemble learning algorithm is a random forest algorithm, and the deep learning algorithm is a feedforward neural network algorithm; the hyperparameters of the random forest algorithm include the number of trees, the maximum depth, and the minimum number of samples required for node splitting; the feedforward neural network algorithm includes an input layer, at least one hidden layer, and an output layer, wherein the hidden layer uses the ReLU activation function, and the output layer uses the Sigmoid activation function.

[0026] Further: the performance evaluation in step S4 includes at least two of the following: discrimination evaluation, calibration evaluation, and clinical effectiveness evaluation; the discrimination evaluation is achieved by calculating the C-index or the area under the receiver operating characteristic curve; the calibration evaluation includes at least one of the Brier score, goodness-of-fit test, or calibration curve analysis; and the clinical effectiveness evaluation uses decision curve analysis.

[0027] Furthermore: the demographic variables include at least two of age, sex, weight, and body mass index; the clinical symptom variables include at least two of the following: number of days of fever, highest body temperature, cough, and wheezing; the laboratory test variables include at least three of the following: C-reactive protein, procalcitonin, white blood cell count, neutrophil percentage, lactate dehydrogenase, ferritin, albumin, and interleukin-6; and the imaging variables include the presence of large areas of consolidation and / or the presence of pleural effusion.

[0028] The present invention also provides a macrolide-based non-reactive Mycoplasma pneumoniae pneumonia prediction model, which is constructed using the above-described method.

[0029] Compared with the prior art, the present invention has the following advantages:

[0030] I. This invention provides for the first time a method for constructing a predictive model specifically for macrolide-resistant Mycoplasma pneumoniae pneumonia, which can accurately identify high-risk patients for MUMPP in the early stages of hospital admission (e.g., within 24 hours). Compared with the traditional method that requires 72 hours of treatment for diagnosis, this method advances the identification time by at least 48 hours, gaining valuable time for timely adjustment of clinical treatment strategies and effectively reducing the risk of patients progressing to refractory Mycoplasma pneumoniae pneumonia.

[0031] Second, through a systematic data preprocessing workflow, including chain equation multiple imputation to handle missing values, determination of the optimal tangent point for continuous variables based on ROC curve analysis, and processing of categorical variables using one-hot encoding, combined with feature selection through recursive feature elimination and hyperparameter optimization through grid search, the constructed prediction model has improved prediction accuracy and significantly enhanced the accuracy of clinical diagnosis.

[0032] Third, through a multi-dimensional evaluation system of discrimination, calibration and clinical effectiveness, especially the decision curve analysis, it is shown that the net benefit of model-assisted decision-making is better than that of extreme strategies within a wide risk threshold range of 20%-80%, which proves that the model is not only statistically significant, but also has practical clinical application value and can provide doctors with quantitative and reliable treatment decision-making basis.

[0033] Fourth, this invention provides a complete technical process from data acquisition, preprocessing, model building to performance evaluation. Each step has clear technical parameters and implementation details, ensuring the reproducibility of the method and its scalability in different medical institutions, which is conducive to its widespread application in clinical practice. Attached Figure Description

[0034] Figure 1 This is a schematic flowchart illustrating the steps of constructing a macrolide-based non-reactive Mycoplasma pneumoniae pneumonia prediction model in one embodiment of the present invention.

[0035] Figure 2 This is a flowchart illustrating the steps of constructing the original dataset in S1 of one embodiment of the present invention.

[0036] Figure 3 This is a flowchart illustrating the steps in S2 of an embodiment of the present invention to preprocess the original dataset to obtain a standardized dataset.

[0037] Figure 4 This is a flowchart illustrating the steps of constructing the optimal prediction model based on the standardized dataset in S3 of one embodiment of the present invention.

[0038] Figure 5 This is a flowchart illustrating the steps in S4 of an embodiment of the present invention to evaluate the performance of the optimal prediction model and obtain a macrolide-based non-reactive Mycoplasma pneumoniae pneumonia prediction model. Detailed Implementation

[0039] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0041] This invention provides a method for constructing a predictive model for macrolide-resistant Mycoplasma pneumoniae pneumonia. For example... Figure 1 As shown, the method includes the following steps:

[0042] S1: Obtain clinical data of patients with Mycoplasma pneumoniae pneumonia, screen predictive variables, and construct the original dataset.

[0043] like Figure 2 As shown, in some embodiments, patient medical record data can be extracted from a community-acquired pneumonia database of a medical institution. Patients need to meet the diagnostic criteria for both community-acquired pneumonia and Mycoplasma pneumoniae infection. The diagnosis of community-acquired pneumonia requires meeting three conditions: evidence of acute infection, such as fever, hypothermia, or abnormal white blood cell count; acute respiratory symptoms, such as new cough, dyspnea, tachypnea, and chest pain; and chest X-ray showing new infiltration or consolidation. The diagnosis of Mycoplasma pneumoniae infection requires a positive test for Mycoplasma pneumoniae DNA or RNA in the respiratory tract, or a serum Mycoplasma IgM titer ≥1:160.

[0044] In other embodiments, during data screening, it is necessary to exclude Mycoplasma pneumoniae carriers, patients with community-acquired pneumonia not infected with Mycoplasma pneumoniae, and patients with underlying diseases such as chronic lung disease, immunodeficiency, airway malformation, and cancer, to ensure the accuracy and representativeness of the data. The predictive variables screened should cover multiple dimensions, including demographic variables such as age, sex, weight, and body mass index; clinical symptom variables such as number of days of fever, highest body temperature, cough characteristics, and presence of wheezing; physical signs such as inspiratory retractions, lung rales, and wheezing; laboratory test variables such as C-reactive protein, procalcitonin, white blood cell count, neutrophil percentage, lactate dehydrogenase, ferritin, albumin, and interleukin-6; and imaging variables such as the presence of large areas of consolidation and pleural effusion. Furthermore, if conditions permit, data on Mycoplasma pneumoniae DNA drug resistance sites and mixed infections may also be included.

[0045] S2: Preprocess the original dataset, including imputing missing data, classifying continuous variables, and encoding categorical variables to obtain a standardized dataset.

[0046] Specifically, such as Figure 3 As shown, after obtaining the original dataset, it needs to be systematically preprocessed to obtain a standardized dataset. In some embodiments, missing data is handled using a chain equation multiple imputation method. This method can generate multiple complete imputation datasets based on the inherent relationships of the data, effectively ensuring data integrity and unbiasedness. In practice, this process can be implemented using the `mice` package in R, typically generating five imputation datasets, from which the subsequent model training results will be aggregated.

[0047] The categorical transformation of continuous variables is a crucial step in data preprocessing. In some embodiments, continuous variables with clinically recognized thresholds can be directly transformed based on these thresholds. For example, fever is typically defined as a body temperature above 38.5°C, so this continuous variable can be directly transformed into a binary variable indicating whether or not a person has a fever. For continuous variables without clearly defined clinical thresholds, receiver operating characteristic (ROC) curve analysis is needed to determine the optimal cutoff value. Specifically, using macrolide-resistant Mycoplasma pneumoniae pneumonia as the outcome variable, the Youden index is calculated for each possible value of the continuous variable. The Youden index is calculated as sensitivity plus specificity minus 1. The value that maximizes the Youden index is selected as the optimal cutoff value, and then the continuous variable is transformed into a categorical variable based on this cutoff value. For example, for the variable C-reactive protein, if the optimal cutoff value is determined to be 48.5 mg / L, it can be transformed into a binary variable indicating whether the protein level is above or below 48.5 mg / L.

[0048] In other embodiments, one-hot encoding is used for categorical variables. When a categorical variable has only one option, it is directly converted to a binary variable. When a categorical variable contains multiple options, different processing is required based on the relationships between the options. If the options are mutually exclusive, they are converted to the corresponding number of independent binary variables; if the options are not mutually exclusive, each option needs to be split into a separate column and converted to a binary variable. For example, for the variable of lung imaging manifestations, if it contains three mutually exclusive options—segmental consolidation, lobar consolidation, and interstitial changes—it needs to be converted to three independent binary variables, with each sample having only one 1 and the rest 0 in these three variables.

[0049] S3: Based on the standardized dataset, a prediction model is constructed using machine learning algorithms, and the optimal prediction model is obtained through feature selection and parameter optimization.

[0050] Specifically, such as Figure 4 As shown, after data preprocessing, the model building stage begins. In this stage, the invention employs a rigorous multi-step process, combining advanced feature selection methods and two different types of machine learning models: an ensemble learning model and a deep learning model. Through systematic hyperparameter optimization, the optimal prediction model is obtained.

[0051] First, feature selection is performed using recursive feature elimination to filter out the combination of clinical variables most informative for the prediction target. All preprocessed features are used as the initial feature set. A random forest algorithm is chosen as the base evaluator in the recursive feature elimination process because it effectively assesses feature importance and has a good ability to capture non-linear relationships in the data. The random forest model is trained on the training data using the complete initial feature set. After training, the importance score of each feature is obtained by calculating the average reduction in Gini impurity or ranking importance scores, and all features are sorted from highest to lowest importance. Then, one or more features with the lowest importance scores are removed from the current feature set, and the model is retrained using the remaining feature subset, and the least important features are evaluated and removed again. This process is iterated until the number of remaining features in the feature set reaches a preset minimum, such as one feature. In each iteration, a feature subset and its corresponding model performance are obtained. 5-fold or 10-fold cross-validation is used to evaluate the performance of the model built on each feature subset, and the performance evaluation metric is the area under the receiver operating characteristic (ROC) curve. Finally, the feature subset that maximizes the average AUC value of cross-validation is selected as the optimal feature subset.

[0052] A random forest prediction model is constructed based on the selected optimal feature subset. Random forests improve the stability and accuracy of the model by constructing multiple decision trees and integrating their prediction results. Specifically, a bootstrap method is first used to randomly select multiple sample subsets of the same size as the original training set from the training dataset with replacement. For each sample subset, a decision tree is constructed independently. When splitting at each node of the decision tree, instead of selecting the optimal splitting feature from all features, a subset of features is randomly selected, typically the square root of the total number of features, and then the optimal feature is selected from this subset for splitting. This double randomness ensures low correlation between the decision trees. For classification tasks, the final prediction result is determined by voting from all decision trees, and the class with the most votes becomes the model's prediction output.

[0053] To achieve optimal performance for the random forest model, a grid search combined with cross-validation is used to systematically optimize its key hyperparameters. The hyperparameters to be optimized include the number of decision trees (searchable values ​​of 100, 300, 500, 800, etc.); the maximum depth of the decision trees (searchable values ​​of 5, 10, 15, 20, or unlimited depth); and the minimum number of samples required for node splits (searchable values ​​of 2, 5, 10, etc.). The grid search traverses all possible combinations of hyperparameters, and cross-validation is used to evaluate the performance of each combination, using AUC as the evaluation metric. Finally, the optimal parameter combination is selected to construct the final random forest model.

[0054] To explore deeper nonlinear relationships in the data, this invention also constructs a feedforward neural network model, also known as a multilayer perceptron. The network architecture includes an input layer, hidden layers, and an output layer. The number of neurons in the input layer equals the number of features in the optimal feature subset, and each neuron receives the input value of one feature. For example, if the optimal features are 18, then the input layer has 18 neurons. At least one hidden layer is included, typically two fully connected hidden layers, to enhance the model's expressive power. For example, a network structure can be constructed where the first hidden layer contains 32 neurons and the second hidden layer contains 16 neurons. Each neuron in the hidden layer uses a modified linear unit as its activation function, with the functional form: This effectively alleviates the vanishing gradient problem and accelerates model convergence. The output layer contains one neuron, which outputs the probability of predicting a MUMPP, using the sigmoid activation function. It can map any real number output to the interval between 0 and 1, which conforms to the definition of probability.

[0055] The neural network model is trained using binary cross-entropy as the loss function to measure the difference between the model's predicted probabilities and the true labels. An adaptive moment estimation optimizer is used to update the network's weights and biases. This optimizer combines the advantages of momentum and RMSProp algorithms, adaptively adjusting the learning rate of each parameter for efficient and stable training. Similarly, a grid search combined with cross-validation method is used to optimize the key hyperparameters of the neural network model, including the learning rate (searching for values ​​of 0.01, 0.001, and 0.0001) and batch size (searching for values ​​of 16, 32, and 64), which determines the number of samples used for each weight update.

[0056] Through the above steps, a random forest model and a feedforward neural network model, after feature selection and parameter optimization, were constructed respectively. These two models capture patterns in the data from different perspectives: the random forest provides stable predictions by integrating multiple decision trees, while the neural network can learn more complex nonlinear relationships in the data. In the subsequent evaluation phase, a comprehensive performance comparison of the two models will be conducted, and the model with the best overall performance will be selected as the prediction model for macrolide-resistant Mycoplasma pneumoniae pneumonia.

[0057] S4: The performance of the optimal prediction model is evaluated to obtain a macrolide-based non-reactive Mycoplasma pneumoniae pneumonia prediction model.

[0058] Specifically, such as Figure 5 As shown, after the model is built, a comprehensive performance evaluation is required. In some embodiments, discrimination is assessed by calculating the C-index or the area under the receiver operating characteristic curve, which reflects the model's ability to distinguish between high-risk and low-risk patients. Calibration is assessed using the Brier score, goodness-of-fit test, or calibration curve analysis, which evaluate the consistency between the model's predicted probabilities and actual probabilities. Clinical effectiveness is assessed using decision curve analysis, which calculates the net benefit of using the model-assisted decision-making approach compared to a universal treatment or no-treatment strategy at different risk thresholds, thereby evaluating the model's clinical utility.

[0059] In another embodiment, the present invention also provides a macrolide-based non-reactive Mycoplasma pneumoniae pneumonia prediction model, which is constructed using the above-described method.

[0060] The implementation process of this invention will be further illustrated below through a specific implementation case. A retrospective collection of medical records from 1200 children diagnosed with Mycoplasma pneumoniae pneumonia and hospitalized between January 2020 and December 2024 was conducted from a children's hospital's community-acquired pneumonia database. Based on the clinical response 72 hours after standardized intravenous infusion of macrolides such as azithromycin, the children were divided into a MUMPP group and a non-MUMPP group. Those whose body temperature remained above 38.5℃ and whose clinical symptoms and lung imaging showed no improvement or worsened were classified as the MUMPP group, serving as positive samples; conversely, those with normal body temperature were classified as the non-MUMPP group, serving as negative samples.

[0061] Sixty-eight candidate predictive variables were extracted from medical records, covering demographic characteristics, clinical symptoms and signs, laboratory test results, and imaging findings. The dataset was randomly divided into a training set of 840 cases and a test set of 360 cases in a 7:3 ratio. For approximately 15% of the missing data in the training set, a chain equation multiple imputation method was used to generate five complete imputation datasets. For continuous variables without clearly defined clinical thresholds, such as C-reactive protein and lactate dehydrogenase, optimal cutoff values ​​were determined through ROC curve analysis. For example, the optimal cutoff value for C-reactive protein was determined to be 48.5 mg / L, and the optimal cutoff value for lactate dehydrogenase was determined to be 380 U / L. All categorical variables were processed using one-hot encoding.

[0062] Feature selection was performed using recursive feature elimination, with random forest as the base evaluator. The performance of different feature subsets was evaluated using 5-fold cross-validation. Results showed that the model achieved a peak cross-validation AUC of 0.89 when the number of features was 18. These 18 optimal features mainly included the number of days of fever before admission, highest body temperature, C-reactive protein level, lactate dehydrogenase level, ferritin level, albumin level, presence of large areas of consolidation, presence of pleural effusion, and Mycoplasma pneumoniae drug resistance mutations.

[0063] Based on 18 selected features, a random forest model and a feedforward neural network model were constructed. The optimal parameters for the random forest model, determined through grid search, were: 300 trees, a maximum depth of 10, and a minimum number of samples required for node splitting of 5. The neural network model adopted an architecture of: 18 nodes in the input layer, 32 nodes in the first hidden layer, 16 nodes in the second hidden layer, and 1 node in the output layer; the optimal learning rate was 0.001, and the batch size was 32.

[0064] Evaluations on independent test sets showed that the neural network model achieved a C-index of 0.91 with a 95% confidence interval of 0.88–0.94, while the random forest model achieved a C-index of 0.88. At the optimal probability threshold of 0.42, the neural network model demonstrated an accuracy of 89.2%, a sensitivity of 92.5%, and a specificity of 87.9%. A Brier score of 0.11 indicates good model calibration. Decision curve analysis showed that within a risk threshold range of 20% to 80%, the net benefit of using this model for decision-making was higher than both universal treatment and non-treatment strategies, demonstrating the model's high clinical value.

[0065] In contrast, the traditional Logistic regression model built using the same data and features has an AUC of only 0.78, which is significantly lower than the model built by the method of this invention, demonstrating the advantages of this invention in processing complex medical data.

[0066] The method of this invention can not only be used to build predictive models, but the resulting models can also be integrated into hospital information systems to provide clinicians with real-time risk assessments. Furthermore, the method can be stored on a computer-readable storage medium or deployed on specialized computer equipment, facilitating its widespread application in different medical institutions.

[0067] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent transformations or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for constructing a predictive model for macrolide-resistant Mycoplasma pneumoniae pneumonia, characterized in that, include: S1: Obtain clinical data of patients with Mycoplasma pneumoniae pneumonia, screen predictive variables, and construct the original dataset; S2: Preprocess the original dataset, including imputing missing data, classifying continuous variables, and encoding categorical variables to obtain a standardized dataset; Missing data imputation is handled using a chain equation multiple imputation method; The classification transformation of continuous variables includes: for continuous variables with clinically recognized thresholds, direct conversion to categorical variables based on the clinical thresholds; for continuous variables without explicit clinical thresholds, determining the optimal cut-off value through receiver operating characteristic (ROC) curve analysis, and converting to categorical variables based on the optimal cut-off value. The determination of the optimal cut-off value through receiver operating characteristic (ROC) curve analysis includes: using macrolide-resistant mycoplasma pneumonia as the outcome variable, calculating the Youden index corresponding to each possible value of the continuous variable, and selecting the value that maximizes the Youden index as the optimal cut-off value; where, Youden index = sensitivity + specificity - 1; S3: Based on the standardized dataset, a prediction model is constructed using machine learning algorithms, and the optimal prediction model is obtained through feature selection and parameter optimization; The machine learning algorithm includes ensemble learning algorithms and / or deep learning algorithms; the feature selection adopts recursive feature elimination, and the performance of different feature subsets is evaluated through cross-validation to determine the optimal feature subset; the parameter optimization adopts grid search algorithm to systematically optimize the hyperparameters of the model. The ensemble learning algorithm is a random forest algorithm, and the deep learning algorithm is a feedforward neural network algorithm. The hyperparameters of the random forest algorithm include the number of trees, the maximum depth, and the minimum number of samples required for node splitting. The feedforward neural network algorithm includes an input layer, at least one hidden layer, and an output layer, wherein the hidden layer uses the ReLU activation function, and the output layer uses the Sigmoid activation function. S4: The performance of the optimal prediction model is evaluated to obtain a macrolide-based non-reactive Mycoplasma pneumoniae pneumonia prediction model.

2. The construction method according to claim 1, characterized in that, In step S1, the predictive variables include at least two categories of demographic variables, clinical symptom variables, laboratory test variables, and imaging variables; the acquisition of clinical data includes: extracting patient data that meet the diagnostic criteria for Mycoplasma pneumoniae pneumonia from a community-acquired pneumonia database, and dividing patients into a macrolide non-responsive group and a macrolide responsive group based on their clinical response 72 hours after macrolide treatment.

3. The construction method according to claim 2, characterized in that, The demographic variables include The clinical symptom variables include at least two of the following: age, sex, weight, and body mass index; the clinical symptom variables include at least two of the following: number of days of fever, highest body temperature, cough, and wheezing; the laboratory test variables include at least three of the following: C-reactive protein, procalcitonin, white blood cell count, neutrophil percentage, lactate dehydrogenase, ferritin, albumin, and interleukin-6; and the imaging variables include the presence of large areas of consolidation and / or the presence of pleural effusion.

4. The construction method according to claim 1, characterized in that, The categorical variable encoding process in step S2 uses one-hot encoding, specifically including: When a categorical variable has only one option, it is directly converted into a binary variable; When a categorical variable has n options and n≥2, if the n options are mutually exclusive, they are converted into n independent binary variables; if the n options are not mutually exclusive, each option is split into a separate column and converted into a binary variable.

5. The construction method according to claim 1, characterized in that, The performance evaluation in step S4 includes at least two of the following: discrimination evaluation, calibration evaluation, and clinical effectiveness evaluation; the discrimination evaluation is achieved by calculating the C-index or the area under the receiver operating characteristic curve; the calibration evaluation includes at least one of the Brier score, goodness-of-fit test, or calibration curve analysis; and the clinical effectiveness evaluation uses decision curve analysis.

6. A macrolide-based non-reactive Mycoplasma pneumoniae pneumonia prediction model, characterized in that, It is constructed by the construction method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Prediction and analysis method for complicated pulmonary embolism of patient with lower limb deep venous thrombosis

    CN113053534A