Instant interpretable viral pneumonia condition grading discrimination model construction method based on symptoms and signs, discrimination system and application
By constructing a disease grading model based on various machine learning methods and the SHAP algorithm based on symptoms and signs, the problem of inaccurate disease grading for novel coronavirus infection was solved, achieving immediacy and interpretability, making it suitable for primary healthcare environments and guiding personalized treatment.
Patent Information
- Application Number
- CN202510913255.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-21
AI Technical Summary
In existing technologies, the severity of COVID-19 infection lacks precise differentiation, relies on laboratory and imaging data, has insufficient interpretability, and is difficult to popularize and apply in primary healthcare settings. Furthermore, existing models lack immediacy and interpretability.
A symptom- and sign-based model for grading viral pneumonia is constructed. By acquiring clinical information, key symptom and sign feature vectors are selected. The model is then combined with various machine learning methods and the SHAP algorithm to establish a grading system. A visual explanation is provided through a nomograph.
It enables real-time assessment of disease severity in the early stages of infection, reduces reliance on medical resources, improves model interpretability, is applicable to primary healthcare environments, guides personalized treatment plans, and optimizes the allocation of medical resources.
Smart Images

Figure CN120824028A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical technology, and in particular to a method for constructing an instant and interpretable viral pneumonia disease grading discrimination model based on symptoms and signs, a discrimination system, and an application. Background Art
[0002] Viral pneumonia is an acute respiratory infection caused by a variety of viruses (such as influenza, respiratory syncytial virus, novel coronavirus, and adenovirus). It is characterized by sudden onset, rapid spread, and significant harm. Its clinical manifestations (such as fever, cough, and dyspnea) and progression patterns are highly similar across infections caused by different pathogens. Because viral pneumonia often progresses rapidly and rapidly, early identification of high-risk patients and intervention are crucial.
[0003] The novel coronavirus (COVID-19) is caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Numerous studies have shown that the clinical prognosis of patients infected with the novel coronavirus is influenced by multiple factors, including individual differences in age, gender, and underlying medical conditions. Furthermore, clinical manifestations in the early stages of infection can also provide some indication of the severity of the disease.
[0004] Novel coronavirus infection can be divided into four categories according to the severity of the disease: mild, moderate, severe, and critical. Different patients have significant differences in clinical manifestations, and the severity of their disease is directly related to the patient's prognosis. Its clinical manifestations are highly heterogeneous, ranging from asymptomatic infections and mild patients to the development of severe viral pneumonia and even multiple organ failure. The clinical manifestations are diverse and involve multiple systems throughout the body. Therefore, how to accurately judge the severity of viral pneumonia in patients in the clinic, especially to identify patients at risk of progression in the early stages, so as to effectively prevent mild or moderate patients from transforming into severe or critical cases, is an important challenge facing current clinical work.
[0005] In recent years, machine learning (ML), leveraging its advantages in feature extraction and pattern recognition, has been widely applied in disease risk prediction and clinical decision support. For example, existing research has focused on the early diagnosis of novel coronavirus infection, differential diagnosis from other types of pneumonia, and the classification of severe and non-severe patients. However, it is unable to precisely differentiate the specific severity of novel coronavirus infection. Furthermore, most existing models rely on laboratory test data and imaging data. For example, one study applied image segmentation technology to analyze chest CT images of novel coronavirus patients and assessed their severity by calculating lung lesion scores. Other researchers have incorporated laboratory indicators such as high-sensitivity C-reactive protein, lymphocyte count, and D-dimer into models to predict the prognosis of novel coronavirus infection, achieving good results.
[0006] However, as mentioned above, the clinical presentation and intervention needs of patients infected with the novel coronavirus vary significantly depending on the severity of their illness. Current research has insufficiently focused on nuanced differentiation of these specific illnesses, and there is still a lack of discriminant models that can rapidly and accurately stratify the disease in its early stages. Early identification of the severity of the disease is crucial for guiding treatment plans, optimizing medical resource allocation, and preventing mild and moderate cases from progressing to severe disease.
[0007] Furthermore, most existing models rely too heavily on laboratory tests and imaging data, resulting in relatively complex operational procedures. This makes them difficult to popularize and apply, particularly in settings where primary healthcare resources are scarce, limiting their clinical practicality. Furthermore, some machine learning models are "black box" models. While they offer high predictive performance, they lack clear decision-making logic, making it difficult to provide clinicians with interpretable judgments. This hinders their acceptance and effectiveness in actual diagnosis and treatment.
[0008] Therefore, developing a viral pneumonia grading and discrimination model based on clinical symptom and sign data that is immediate, well-interpretable, and easy to use in clinical practice has important research value and application prospects. Summary of the Invention
[0009] In view of the above problems, the present invention provides a method for constructing an instant and interpretable viral pneumonia disease grading discrimination model based on symptoms and signs. Through this method, a viral pneumonia disease grading discrimination model that is both interpretable and immediate and practical can be constructed, which is used to quickly determine the disease grade in the early stages of the disease, so as to overcome the above problems or at least partially solve the above problems.
[0010] In a first aspect, the present invention provides a method for constructing an instant interpretable viral pneumonia disease grading model based on symptoms and signs, the method comprising: Obtain clinical information of the examinee and construct a sample data set based on the clinical information; Based on the clinical information of the examinee, N preliminary clinical manifestation variables are determined; M variables are selected from N preliminary clinical manifestation variables as clinical symptom and sign feature vectors; Build multiple machine learning models separately, input the clinical symptom and sign feature vectors into each machine learning model, train each model on the sample data set, and determine the optimal model; The M clinical symptom and sign feature vectors were incorporated into the optimal model to establish a viral pneumonia disease grading and discrimination model; Optionally, the method of screening out M variables from N preliminary clinical manifestation variables as clinical symptom and sign feature vectors includes: Based on the univariate logistic regression analysis method, N preliminary clinical manifestation variables were screened to determine P statistically significant characteristic variables; Sort the P statistically significant characteristic variables screened out from large to small according to the OR value; Based on the correlation analysis method, the characteristic variables with correlation greater than the correlation coefficient threshold and smaller OR value are screened out from the sorted characteristic variables; The remaining M feature variables are used as clinical symptom and sign feature vectors.
[0011] Optionally, the clinical symptom and sign feature vector includes: Vomiting, hoarseness, irritability, abdominal pain, shortness of breath, sputum, fullness and distension in the ribs, cough, blurred vision, heavy breathing, fullness and stuffiness, and fever.
[0012] Optionally, constructing multiple machine learning models respectively, inputting the clinical symptom and sign feature vectors into each machine learning model respectively, performing model training on the sample data set, and determining the optimal model includes: Divide the sample data set into a training set and a test set, and expand the sample data of mild and severe patients in the training set; Build logistic regression model, support vector machine model, random forest model, neural network model and extreme gradient boosting model respectively; On the training set, the above models were trained based on stratified 5-fold cross-validation, and the hyperparameters of each model were optimized in combination with grid search; Sort the M clinical symptom and sign feature vectors by OR value from large to small, and gradually input different numbers of clinical symptom and sign feature vectors into each of the above trained models; The performance of the trained model is evaluated on the test set, and the model with the best performance is selected as the optimal model.
[0013] Optionally, the optimal model is a random forest model.
[0014] Optionally, the method further includes: SHAP interpretability analysis was performed on the optimal model to determine the contribution of each clinical symptom and sign feature vector to the prediction results; Based on M clinical symptom and sign feature vectors, an ordered logistic regression model was constructed and a nomogram was drawn. The corresponding scores of different clinical symptom and sign feature vectors were displayed according to the nomogram.
[0015] In a second aspect, an embodiment of the present application provides a viral pneumonia disease grading system, the system comprising: A data acquisition module is used to obtain clinical manifestation information of target patients; The nomogram construction module uses ordered logistic regression to construct a nomogram based on the M characteristic variables included in the viral pneumonia disease grading discriminant model. The nomogram is used to obtain the corresponding scores of different clinical symptom and sign feature vectors and the probability of the disease grading. The viral pneumonia disease grading discriminant model is constructed according to the method for constructing a symptom- and sign-based instant and interpretable viral pneumonia disease grading discriminant model. The disease classification module is used to calculate the total score of the patient's clinical symptom and sign feature vectors and determine the patient's disease classification based on the total score. The disease classification includes mild, common and severe.
[0016] In a third aspect, an embodiment of the present application provides a readable storage medium storing a program or instruction. When the program or instruction is executed by a processor, a method for constructing an instantly interpretable viral pneumonia disease grading model based on symptoms and signs is implemented.
[0017] Fourthly, the viral pneumonia disease grading and discrimination model is applied to the disease grading of patients infected with the new coronavirus.
[0018] Fifthly, the viral pneumonia disease grading discrimination model is applied to the disease grading of patients with influenza virus pneumonia.
[0019] The specific beneficial effects are: First, this application has constructed a viral pneumonia disease grading and discrimination model based on 12 clinical symptoms and signs by comprehensively applying multiple machine learning methods. The model is based only on the symptoms and signs data shown by the patient in the early stage of infection, without relying on laboratory or imaging examinations. It significantly reduces the dependence on medical resources, making the model highly applicable in primary medical environments, especially in areas with limited resources. Through real-time monitoring and analysis of clinical symptoms and signs, this model can immediately determine the patient's disease grade in the early stage of infection, providing support for early intervention and treatment decisions of the disease; Second, this application uses the SHAP algorithm to analyze the contribution of each symptom and sign in the optimal machine learning model to the prediction results. The SHAP algorithm can quantify the contribution of each feature to the classification of different disease conditions and demonstrate its direction of action in different types of disease conditions (mild, common, and severe). The interpretability of the model is improved, allowing clinicians to clearly understand the decision-making basis of the model. Unlike traditional "black box" machine learning models, the model of the present invention provides a transparent decision-making path, which helps to improve doctors' trust in and willingness to adopt the judgment results, thereby gaining wider application in actual diagnosis and treatment; Third, the present invention uses an ordered logistic regression method to visualize the results of the prediction model through a nomogram. This nomogram clearly shows the discriminative effect of 12 symptoms and signs on different disease grades, and the relationship between the score of each symptom or sign and the disease grade is intuitive and clear. The application of the nomogram not only makes the discrimination process of the model more intuitive, but also helps doctors make quick judgments in the clinic to determine whether the patient is mild, ordinary or severe. Through this visualization tool, doctors can quickly understand the degree of impact of each symptom and the disease grade, and formulate personalized treatment plans for patients based on this. The nomogram is simple to understand, easy to operate, instant and fast, suitable for primary medical scenarios, and has broad potential for promotion and application; In summary, this application addresses the problems of existing COVID-19 related prediction models, such as lack of accurate differentiation of disease classification, reliance on laboratory and imaging data, and insufficient model interpretability. An instant and interpretable viral pneumonia disease classification and discrimination model based on symptoms and signs is proposed. This model is based on multi-system clinical manifestation data, comprehensively considers the various symptoms manifested by patients in the early stage of infection, and combines multiple machine learning algorithms and ordered logistic regression methods. It can identify the early risks of COVID-19 patients, assist clinical prognosis judgment, and thus intervene in potential severe or critical cases in advance. It is of great significance for optimizing the allocation of medical resources and guiding individualized treatment strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A flowchart of the method for constructing an interpretable discriminant model for viral pneumonia grading; Figure 2 is the confusion matrix of the training set and test set of the optimal model (random forest).
[0022] Figure 3 ROC curves for the training set and test set of the optimal model (random forest).
[0023] Figure 4 This is a nomogram for the grading of pneumonia caused by COVID-19 infection based on ordered logistic regression.
[0024] Figure 5 (a)- Figure 5 (c) is the interpretability analysis result of the random forest model based on SHAP. DETAILED DESCRIPTION
[0025] The exemplary embodiments of the present application will be described in more detail below in conjunction with the accompanying drawings in the embodiments of the present application. Although the accompanying drawings show exemplary embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0026] Example 1 In this embodiment, taking the novel coronavirus infection (COVID-19) as the research object, in response to the problems in existing related prediction model research, such as the lack of accurate distinction in the classification of novel coronavirus infection disease, reliance on laboratory and imaging data, and insufficient model interpretability, a method for constructing a COVID-19 disease classification and discrimination model based on symptoms and signs is provided, so as to achieve accurate and rapid identification of the disease in the early stage and convenient clinical operability.
[0027] According to the background technology, COVID-19 can be divided into four categories according to the severity of the disease: mild, common, severe and critical. The classification criteria revolve around clinical manifestations, imaging manifestations, oxygenation function, organ failure, etc. This "classification" is essentially a graded evolution of severity, and the guidelines still use the term "classification", which is mainly based on the classification management needs of acute infectious diseases. "Classification" focuses on the essential differences in disease categories, while "grading" focuses on the continuous spectrum of severity of the same disease. Although the discrimination of this model is based on the "classification" standard (mild / common / severe / critical), its essence focuses on the grading of disease severity. Therefore, in the present invention, "grading" is used to more accurately reflect the discrimination target of this model - the severity of mild, moderate and severe viral pneumonia.
[0028] Reference Figure 1 , Figure 1 A flowchart of a method for constructing a symptom- and sign-based, instantly interpretable viral pneumonia disease grading model is provided in an embodiment of the present application. The method may include: Step 1: Obtain the clinical information of the examinee and construct a sample data set based on the clinical information.
[0029] First, determine the research subjects (i.e., the subjects); The subjects in this example were 320 patients diagnosed with novel coronavirus infection who were admitted to multiple tertiary-level A hospitals in JS Province and HB Province from 2019 to 2020. All collected cases were approved by the hospital's ethics review, and the patient information was anonymized.
[0030] According to the set inclusion and exclusion criteria, the original case data were screened and quality controlled, and finally 289 patient data were included for model construction and analysis. Specifically, the inclusion criteria: meet the diagnostic criteria for confirmed cases of pneumonia infected with the new coronavirus in the relevant national guidelines; the disease classification is clear and meets any of the types of mild, ordinary, severe or critical; the patient has a complete hospitalization record and complete clinical data to meet the needs of feature variable extraction. Exclusion criteria: missing key data (such as missing or unclear disease classification), a total of 12 cases were excluded; combined with serious underlying diseases that may significantly interfere with the progression of COVID-19 (such as advanced malignant tumors, end-stage renal disease, etc.), a total of 9 cases were excluded; although confirmed, there was no clinical symptom record in the case, and symptom feature analysis could not be performed, a total of 10 cases were excluded.
[0031] Secondly, clinical information of the research subjects was collected; Based on 289 confirmed patients infected with the novel coronavirus, clinical information relevant to the analysis subjects was extracted using the hospital's electronic medical record system. This information included general demographic characteristics, medical history, subjective symptoms, and objective physical signs to ensure data integrity and accuracy. A sample dataset was constructed based on the symptoms, signs, and disease severity of all cases. The dataset was randomly divided into stratified groups at a 7:3 ratio, with 70% used as a training set for model construction and 30% as a test set for model validation.
[0032] Step 2: Based on the clinical information of the subject, determine N preliminary clinical manifestation variables.
[0033] Combining the clinical characteristics of novel coronavirus infection and its multi-system impact, we systematically extract information on symptoms and signs that may be related to the disease from the perspectives of Traditional Chinese Medicine (TCM). This covers the following eight system categories: Systemic manifestations: aversion to cold, aversion to wind, fever, fatigue, sweating; Respiratory system: shortness of breath, gurgling phlegm, rough breathing, cough, sputum, and shortness of breath; Head, face, and five sense organs: hoarseness, nasal congestion, runny nose, thirst, bland mouth without thirst, bitter taste in the mouth, sticky mouth, itchy throat, dry throat, sore throat, blurred vision, red eyes, swollen eyes, tinnitus, and deafness; Musculoskeletal system: limb pain, heaviness, limb edema, weakness in waist and knees, cold hands and feet; Gastrointestinal system: abdominal distension, fullness in the ribs, abdominal distension, abdominal pain, poor appetite, nausea, vomiting, sticky stools, dry stools, loose stools; Circulatory system: palpitations, chest pain, fullness and stuffiness; Urinary system: frequent urination, scanty urination, long and clear urination, dark or yellow urine; Nervous system: headache, heaviness of the head, irritability, drowsiness, delirium, coma.
[0034] A total of 54 clinical manifestation variables were extracted as model feature candidates.
[0035] The 54 clinical manifestation variables were labeled and each was uniquely numbered S1–S54, as shown in Table 1. All clinical manifestation variables were binary-coded, with "0" indicating the absence of the manifestation and "1" indicating the presence of the manifestation. During data cleaning, some symptoms were found to be absent in all samples (study subjects) and lacked analytical value. Therefore, these clinical manifestation variables were removed: S3 (delirium), S4 (coma), S22 (limb edema), and S36 (deafness). Finally, 50 clinical manifestation variables were retained as the preliminary set of clinical manifestation variables, with an N of 50.
[0036]
[0037] Step 3: Based on univariate logistic regression analysis and correlation analysis, M variables were selected from N preliminary clinical manifestation variables as clinical symptom and sign feature vectors; Optionally, step 3 includes the following sub-steps: Step 301: Screening N preliminary clinical manifestation variables based on univariate logistic regression analysis to determine P statistically significant characteristic variables; To identify variables that showed statistically significant differences in disease severity (mild, moderate, and severe), we first used univariate logistic regression analysis to screen each variable. Univariate analysis was performed using the patient's disease grade as the dependent variable and N preliminary clinical manifestation variables as the independent variables. Statistically significant clinical manifestations were identified based on the P value of each clinical manifestation variable, with a P value of < 0.05 as the significance standard. During the specific screening process, S2 (somnolence) and S51 (frequent urination) were not included in the subsequent analysis because their frequency in the sample was low (≤1), making it impossible to calculate an OR. The results showed that there were statistically significant differences in 14 variables, including S47 (vomiting), S8 (hoarseness), S1 (irritability), S43 (abdominal pain), S5 (shortness of breath), S6 (sputum), S41 (fullness in the ribs), S13 (cough), S32 (dizziness), S7 (heavy breathing), S39 (fullness and stuffiness), S10 (fever), S15 (shortness of breath), and S46 (nausea). These 14 variables were regarded as statistically significant characteristic variables, that is, P=14.
[0038] The univariate logistic regression analysis method can screen the characteristic variables with statistical differences among mild, common and severe patients.
[0039] Step 302: Sort the P statistically significant characteristic variables selected from large to small according to their OR values (odds ratios).
[0040] By ranking the multiple selected characteristic variables, their relative contribution to the disease classification can be reflected.
[0041] Step 303: Set a correlation coefficient threshold r, screen out characteristic variables with correlations greater than r and smaller OR values, and obtain M characteristic variables as clinical symptom and sign characteristic vectors.
[0042] To further enhance model robustness and mitigate the impact of multicollinearity between features, correlation analysis was performed on the P statistically significant selected feature variables. Correlation analysis was used to assess inter-feature correlations, with a correlation coefficient threshold of r = 0.7. Among variable pairs with correlations greater than 0.7, features S15 (shortness of breath) and S46 (nausea) with small OR values were removed. Finally, 12 feature variables were selected as the input feature combination for the final model: S47 (vomiting), S8 (hoarseness), S1 (irritability), S43 (abdominal pain), S5 (shortness of breath), S6 (sputum sound), S41 (fullness in the ribs), S13 (cough), S32 (dizziness), S7 (heavy breathing), S39 (fullness and stuffiness), and S10 (fever).
[0043] Step 4: Expand the sample data of mild and severe patients in the sample dataset based on the SMOTE algorithm.
[0044] In the original training set of the embodiment of the present application, there is a significant imbalance in the distribution of patient categories, with the sample numbers of mild, common, and severe patients being 22:24:3:24. Among them, the sample numbers of mild and severe patients are obviously too small. In order to improve the impact of class imbalance on model performance and improve the model's prediction effect on minority classes, the training set data is processed by the Synthetic Minority Over-sampling Technique (SMOTE). SMOTE generates new minority class samples by interpolation in the feature space, expands the minority class data, and thus adjusts the class ratio. The class ratio of mild, common, and severe patients in the training set is adjusted from the original ratio to 1:2:1. On the basis of maintaining a moderate advantage in the number of common patients, the SMOTE method is used to oversample the samples of mild and severe patients, increase their sample numbers, and make the proportions of the three types of patients tend to be balanced, which helps the model learn the characteristic laws of each category and improves the model's prediction performance and classification effect on minority class samples.
[0045] Step 5: Build multiple machine learning models respectively, input the clinical symptom and sign feature vectors into each machine learning model respectively, perform model training on the training set, and use the optimal model as the viral pneumonia disease grading discrimination model.
[0046] To identify the most suitable machine learning algorithm for this invention, we constructed the following five machine learning models: Logistic Regression, Support Vector Machine (SVM), Random Forest, Neural Network, and Extreme Gradient Boosting (XGBoost). We trained each of these five models on the training set to determine the optimal model.
[0047] Optionally, step 5 may include the following sub-steps: Step 501 , constructing a logistic regression model, a support vector machine model, a random forest model, a neural network model and an extreme gradient boosting model respectively.
[0048] Specifically, the random forest model was constructed by constructing a manual hyperparameter grid, including mtry (the number of variables randomly selected for each tree): 1, 2, 3, 4, 5, 6; ntree (the number of decision trees): 300, 400, 500, 600; maxnodes (the maximum number of terminal nodes for each tree): 10, 15; nodesize was fixed to 1; When building the support vector machine model, the radial basis function (RBF) was used as the kernel function. The hyperparameter grid included: sigma (width parameter of the RBF kernel function): 0.01, 0.015, 0.02, 0.025, C (penalty coefficient): 1, 5, 10, 15, 20; When building the XGBoost model, the hyperparameter grid includes: nrounds (number of training rounds): 50, 100; max_depth (maximum depth of the tree): 2, 3, 4; eta (learning rate): 0.01, 0.05, 0.1; gamma (minimum split loss): 0, 1; colsample_bytree (column sampling ratio): 0.7, 0.8; min_child_weight (minimum subsample weight sum): 1; subsample (subsample ratio) 0.8,; lambda (L2 regularization) 1; alpha (L1 regularization) 0.1; In the construction of the neural network model, the neural network type selected is "single hidden layer feedforward neural network", and the hyperparameter combinations include: size: 5, 10, 15; decay: 0.1, 0.01, 0.001; In the construction of the logistic regression model, the elastic net regularized logistic regression model was implemented through the glmnet package, and the hyperparameter combinations included: alpha = seq(0, 1, length = 5), lambda = seq(0.01, 0.2, length = 10).
[0049] Step 502: Train the above models based on stratified 5-fold cross validation, and optimize the hyperparameters of each model in combination with grid search.
[0050] During model training, stratified 5-fold cross-validation (CV) was used to train all parameter combinations, combined with grid search to optimize hyperparameters. All data preprocessing and hyperparameter tuning were performed only on the training set to avoid information leakage. The final model was evaluated for performance on the test set.
[0051] Step 503: Sort the M clinical symptom and sign feature vectors from large to small according to the OR value, and input different numbers of clinical symptom and sign feature vectors into each of the above trained models in a step-by-step manner.
[0052] Specifically, to evaluate the impact of different feature combinations on model performance, the 12 selected clinical symptom and sign feature vectors were first sorted from largest to smallest according to their OR values, and the number of features input into each model was gradually reduced. For example, all 12 feature variables were input into each machine learning model for the first time, and all 11 feature variables were input into each machine learning model for the second time. Similarly, 10, 9, and 8 feature variables were input into each model in step 501 in a gradually decreasing manner, thereby observing the impact of different feature numbers on model performance. In addition, the model performance was compared using data before and after incorporating different feature numbers and data balancing.
[0053] Step 504: Evaluate the performance of the trained models on the test set, and select the model with the best performance as the optimal model.
[0054] In the examples of this application, the model performance evaluation metrics include accuracy, precision, recall, F1 score, receiver operating characteristic (ROC) curve, and area under the curve (AUC). The macro-average AUC of the test set is used as the primary evaluation metric, and the test set confusion matrix is used to represent the difference between the final model's predicted results and the actual results.
[0055] Specifically, for random forest, after comparing the macro-average AUC of the test set before and after balancing and with different numbers of features, the macro-average AUC was the highest when all 12 features were included in the test set after SMOTE oversampling. At this time, the optimal hyperparameter combination was: mtry: 24, ntree: 600, maxnodes: 15, and the test set performance indicators were as follows: accuracy 0.632 (95%CI: 0.529-0.736), precision 0.807 (95%CI: 0.705-0.905), recall rate 0.632 (95%CI: 0.529-0.736), F1 score 0.687 (95%CI: 0.595-0.777); the macro-average AUC was 0.813 (95%CI: 0.663-0.897), and the micro-average AUC was 0.852 (95%CI: 0.761-0.922). The confusion matrix of the random forest model training set and test set is as follows Figure 2 (a)- Figure 2 (b) shows that the ROC curve is as follows Figure 3 (a)- Figure 3 (b)
[0056] For the support vector machine, after comparing its macro-average AUC before and after balancing and under different numbers of features, the test set has the highest AUC when all 12 features are included after SMOTE oversampling. At this time, the optimal hyperparameter combination is: sigma: 0.025; C: 1. The performance indicators of the test set are as follows: accuracy 0.644 (95% CI: 0.540–0.747), precision 0.806 (95% CI: 0.712–0.902), recall 0.644 (95% CI: 0.540–0.747), and F1 score 0.695 (95% CI: 0.595–0.779). The macro-average AUC is 0.774 (95% CI: 0.624–0.898), and the micro-average AUC was 0.774 (95% CI: 0.569–0.898).
[0057] For XGBoost, after comparing the macro-average AUC before and after balancing and under different numbers of features, the test set has the highest AUC when all 11 features are included after SMOTE oversampling. At this time, the optimal hyperparameter combination is: nrounds (number of iterations): 100; max_depth (maximum depth of the tree): 3; eta (learning rate): 0.1; gamma (threshold of minimum loss function descent value): 0; colsample_bytree (column sampling ratio of each tree): 0.7; min_child_weight (minimum sum of child node sample weights): 1; subsample (sample sampling ratio of each tree): 0.8, lambda (L2 regularization coefficient): 1, alpha (L1 regularization coefficient): 0.1; the test set performance indicators are as follows: accuracy 0.609 (95% CI: 0.504–0.705), precision 0.798 (95% CI: 0.702–0.904), recall 0.609 (95% CI: 0.609–0.705), and accuracy 0.798 (95% CI: 0.702–0.904). CI: 0.506–0.712), and the F1 score was 0.665 (95% CI: 0.571–0.758). The macro-average AUC was 0.790 (95% CI: 0.651–0.900), and the micro-average AUC was 0.832 (95% CI: 0.721–0.920).
[0058] For the neural network, after comparing the macro-average AUC before and after balancing and under different numbers of features, the test set has the highest AUC when the first 8 features are included after SMOTE oversampling. At this time, the optimal hyperparameter combination is: size: 10, .decay: 0.1. The performance indicators of the test set are as follows: accuracy 0.552 (95% CI: 0.447–0.652), precision 0.783 (95% CI: 0.674–0.897), recall 0.552 (95% CI: 0.448–0.655), and F1 score 0.614 (95% CI: 0.516–0.711). The macro-average AUC is 0.765 (95% CI: 0.619–0.871), and the micro-average AUC was 0.803 (95% CI: 0.671–0.898).
[0059] For the logistic regression model, after comparing the macro-average AUC before and after balancing and under different numbers of features, the test set had the highest AUC when the first 11 features were included after SMOTE oversampling. The optimal hyperparameter combination at this time was: alpha = 1, lambda = 0.01. The test set performance indicators were: accuracy 0.586 (95% CI: 0.447–0.652), precision 0.783 (95% CI: 0.674–0.897), recall 0.552 (95% CI: 0.448–0.655), and F1 score 0.614 (95% CI: 0.516–0.711). The macro-average AUC is 0.765 (95% CI: 0.619–0.871), and the micro-average AUC was 0.803 (95% CI: 0.671–0.898).
[0060] According to the above results, it can be seen that the random forest has the best performance when all 12 features are included after SMOTE balancing (AUC=0.813, 95%CI: 0.663-0.897), which is significantly higher than the optimal AUC values of other models (XGBoost: 0.790(95%CI: 0.651 - 0.900), SVM: 0.774 (95%CI: 0.624 - 0.898), LR: 0.785 (95%CI: 0.632 - 0.881), ANN: 0.765 (95%CI: 0.619 - 0.871)). Therefore, the random forest is determined to be the optimal model, and all 12 feature variables are included at this time.
[0061] Step 6: Perform SHAP interpretability analysis on the optimal model to determine the contribution of each clinical symptom and sign feature vector to the prediction results.
[0062] Optionally, step 6 may include the following sub-steps: Step 601: pre-processing the training set data; Extract the feature data and target variables from the training set. Specifically, extract 12 feature variables for each subject in the training set as feature data, and obtain the severity of the subject's condition (mild, moderate, severe) corresponding to the extracted feature variables. This is used as the target variable. The target variable is a categorical variable named Grade, which includes three categories: Mild, Moderate, and Severe, corresponding to mild, moderate, and severe, respectively. Then, use stratified sampling to extract a subset of 100 data points from the training set as background data for calculating the SHAP value. The background data is used to calculate the baseline contribution of the feature to the prediction in the SHAP analysis (i.e., the prediction when the feature takes the reference value).
[0063] Step 602: Calculate the SHAP value of each feature under different disease categories based on the target variable and feature data; Calculate the SHAP value for each category (mild, moderate, severe). Construct a pred_fun function for each category to predict the probability of each category. Then, use the kernelshap function to calculate the SHAP value of each feature under each category.
[0064] Step 603: Based on the SHAP values of different disease categories, generate a hierarchical interpretability object under each category and generate a visualization chart; The shapviz package in R was used to combine the SHAP value of each category with its corresponding feature data to generate independent hierarchical interpretability objects, one for each disease category. Then, based on the generated hierarchical interpretability objects, the drawing function in the shapviz package was used to draw a bee swarm diagram to intuitively display the contribution and influence direction of each clinical feature in predicting the category.
[0065] The Beeswarm Plot is used to show the influence of each feature, sorted by feature contribution, which can help understand how different features affect the model output. Figure 5 (a)- Figure 5 (c) Bee swarm diagrams corresponding to mild, moderate, and severe conditions.
[0066] Step 604: Determine the contribution of each clinical symptom and sign feature vector to the prediction result based on the visualization object.
[0067] The optimal model was analyzed using the SHapley Additive exPlanations (SHAP) algorithm, which improved the interpretability of the model and quantified the contribution of each feature variable to the prediction results, thereby exploring its impact in different disease classifications.
[0068] Step 7: Based on the M clinical symptom and sign feature vectors, construct an ordered logistic regression model and draw a nomogram. Statistically calculate the total score and probability of the patient's condition classification corresponding to the different clinical symptom and sign feature vectors displayed in the nomogram, and output the patient's discriminant condition classification category.
[0069] Specifically, in order to enhance the applicability of the model, an ordered logistic regression model was constructed based on the 12 characteristic variables screened out by univariate logistic regression and correlation analysis, and a nomogram was drawn, as shown in the figure below: Figure 4 As shown, it can intuitively demonstrate the discriminative role of different features in classifying patient conditions.
[0070] A nomogram is a graphical tool. The nomogram constructed by the present invention can be used to determine the disease grade of patients infected with the new coronavirus based on specific symptoms and signs. Figure 4 The scores corresponding to different symptoms and signs are shown in the table. The score of each symptom is assigned based on "presence" or "absence", "0" = "absence" and "1" = "presence". The patient's condition is finally graded based on the total score of the 12 symptoms and signs. The specific steps are: First, score each symptom and sign based on the patient's clinical presentation (symptoms and signs). The horizontal scale at the top of the nomogram represents the corresponding score when each of the 12 clinical characteristic variables is present or absent. This scale clearly shows the score for each symptom or sign, such as vomiting (S47) with a score of 100, hoarseness (S8) with a score of 51, irritability (S1) with a score of 10, and abdominal pain (S43) with a score of 79. Based on the patient's symptoms and signs, refer to the scores on the nomogram and record the scores of the corresponding symptoms / signs. The nomogram drawn in this application using ordered logistic regression is based on cumulative probability representation, that is, the probability of the disease reaching moderate (common type) and above (Y ≥ 1) and high (severe) and above (Y ≥ 2). The predicted probability of low (mild) is the probability of 'non-common type', that is, P(Y = 0) = 1-P(Y ≥ 1), which means that mild is the remainder of the sum of the probabilities. If the total score falls on the left outer end of the ProbY ≥ 1 probability interval (i.e., to the left of 0.8, with a total score between 0 and 20 points), it can be predicted that the patient's condition is low (mild).
[0071] For example: If the patient has the symptom of "vomiting (S47)", the score is 100 points; the symptom of "shortness of breath (S5)" is 17.5 points; the symptom of "heavy breathing" is 0 points, and there are no other related manifestations, the score is 0, and the total score of all symptoms is 117.5. According to the total score, refer to the probability distribution below the figure to make a judgment on the disease classification. At this time, the total score is outside the probability of Prob Y ≥ 1 (common type) 0.99, and within the probability of Prob Y ≥ 2 (severe) 0.15, indicating that the probability of the patient being a common type is greater than 99%, and the probability of being a severe type is less than 15%. The patient is more likely to develop into a common type in the future.
[0072] In summary, this application screened five machine learning methods to determine that random forest is the optimal model when the first 12 feature combinations are included. However, since the random forest model is a black box model, the interpretability of the random forest model is provided through SHAP analysis. Moreover, based on the above feature screening and the 12 feature combinations determined by the optimal model random forest, it is applied to the ordered logistic regression and the nomogram is drawn to achieve the visualization and applicability of the disease grading discrimination model.
[0073] Example 2 Based on the above-mentioned viral pneumonia disease grading and discrimination model, the present invention also provides a viral pneumonia disease grading and discrimination system, which includes the following modules: A data acquisition module is used to obtain clinical manifestation information of target patients; The nomogram construction module is based on the clinical characteristic variables included in the viral pneumonia disease grading discriminant model and combines ordered logistic regression to construct a nomogram. The nomogram displays the corresponding scores of different clinical symptom and sign feature vectors and the probability of their disease grading. The disease classification module is used to calculate the total score of the patient's clinical symptom and sign feature vectors and determine the patient's disease classification based on the total score. The disease classification includes low, medium, and high, which correspond to mild, common, and severe respectively.
[0074] Experimental Case Case 1 A 77-year-old female patient was classified as having mild disease. Based on the 12 symptoms and signs required by the nomogram, the patient had no S47 (vomiting), with a score of 0; no S8 (hoarseness), with a score of 0; no S1 (irritability), with a score of 0; no S43 (abdominal pain), with a score of 0; no S5 (shortness of breath), with a score of 0; no S6 (sputum), with a score of 0; no S41 (fullness in the ribs), with a score of 0; no S13 (cough), with a score of 0; no S32 (dizziness), with a score of 10; no S7 (rough breathing), with a score of 2.5; no S39 (fullness), with a score of 0; and no S10 (fever), with a score of 0. The total score of the 12 features corresponding to the nomogram "Total Points" was 12.5. According to the nomogram mapping, the patient's total score of 12.5 did not correspond to "common and above disease (Prob Y ≥ 1)" and "severe and above disease (Prob Y The cumulative probability interval of "≥2)" indicates that the patient has a low probability of developing a common or higher condition. According to the transformation calculation of the ordered logistic regression model, the predicted probability of mild (Y=0) is significantly higher than that of common and severe types. Therefore, the model predicts that the patient's condition is mild, which is consistent with the clinical diagnosis.
[0075] Case 2 A 61-year-old male patient was classified as of the common type. According to the 12 symptoms and signs required by the nomogram, the patient had no S47 (vomiting), score 0; no S8 (hoarseness), score 0; no S1 (irritability), score 0; no S43 (abdominal pain), score 0; S5 (shortness of breath), score 17.5; no S6 (sputum), score 0; no S41 (fullness in the ribs), score 0; S13 (cough), score 50; no S32 (dizziness), score 10; S7 (rough breathing), score 0; no S39 (fullness and stuffiness), score 0; and S10 (fever), score 10. According to the nomogram, the total score of the 12 features corresponds to the nomogram "Total Points" of 87.5. The total score corresponds to the 95%-99% probability interval of "Prob Y≥1 Common Type" and the 5%-10% probability interval of "Prob Y≥2 Severe Type". This indicates that the probability of this case reaching the common type is much higher than the probability of severe type. The model judges that the patient's condition is classified as common type, which is consistent with the clinical diagnosis results.
[0076] Case 3 A 50-year-old female patient was classified as having severe pneumonia. Based on the 12 symptoms and signs required by the nomogram, the patient had no S47 (vomiting), a score of 0; had S8 (hoarseness), a score of 52.5; had no S1 (irritability), a score of 0; had no S43 (abdominal pain), a score of 0; had S5 (shortness of breath), a score of 17.5; had no S6 (sputum), a score of 0; had S41 (fullness in the ribs), a score of 47.5; had S13 (cough), a score of 50; had no S32 (dizziness), a score of 10; had no S7 (heavy breathing), a score of 2.5; had S39 (fullness), a score of 10; and had S10 (fever), a score of 10. Based on the nomogram, the total score for the 12 features, corresponding to the "Total Points" nomogram of 200, corresponds to a probability range of 70%-80% for "Prob Y ≥ 2 severe," indicating that the patient's predicted probability of having severe pneumonia was significantly higher than that of other classifications. Therefore, the model classified the patient's condition as severe, which is consistent with the clinical diagnosis.
[0077] Example 3 An application of a viral pneumonia disease grading and discrimination model, which is applied to the disease grading of patients infected with the new coronavirus.
[0078] Example 4 The invention discloses an application of a viral pneumonia disease grading and discrimination model, which is applied to the disease grading of patients with influenza virus pneumonia.
[0079] The reason why this model can be applied to influenza virus pneumonia is that: Viral pneumonia is an inflammation of the lungs caused by a viral infection of the upper respiratory tract that spreads downward. It can occur year-round, but is more common in winter and spring, and can occur in outbreaks or sporadic epidemics. Both novel coronavirus infection and influenza virus pneumonia fall under the category of viral pneumonia. Although the two differ in etiology, they share significant similarities in clinical manifestations, pathophysiological mechanisms, and disease grading criteria. This supports the rationale of applying a disease grading model based on the clinical symptoms and signs of novel coronavirus infection to the grading of influenza virus infection.
[0080] At the level of susceptible populations, both novel coronavirus infection and influenza pneumonia demonstrate widespread susceptibility (Influenza Diagnosis and Treatment Protocol (2025 Edition)). The model proposed in this paper is based on a cohort of COVID-19 patients, whose age range is 47.53 ± 14.45 years (mean ± standard deviation), encompassing young to elderly individuals, and highly overlapping with the population susceptible to influenza pneumonia. Therefore, the resulting data characterizes a common susceptibility spectrum for both types of viral pneumonia, providing a population-based basis for stratifying influenza pneumonia severity.
[0081] In terms of clinical manifestations, the systemic and respiratory manifestations of novel coronavirus infection and influenza virus pneumonia exhibit similarities. Direct viral infection causes damage to respiratory and alveolar epithelial cells, accompanied by local and systemic inflammatory responses, resulting in common symptoms such as cough, fever, and sore throat. The massive release of inflammatory mediators increases alveolar membrane permeability, leading to alveolar edema and impaired oxygenation, resulting in shortness of breath and rasping breaths. Hypersecretion of mucus leads to sputum retention and airway obstruction, resulting in gurgling phlegm. These are all respiratory manifestations of viral infection. From a disease course perspective, excessive immune activation and subsequent inflammatory responses triggered by viral infection can lead to complications related to other systems, most commonly affecting the gastrointestinal and nervous systems, resulting in vomiting, abdominal pain, and irritability. These pathological processes drive similar clinical manifestations, leading to overlapping clinical manifestations of the two viral pneumonias. The features required for prediction in this model focus on the patient's symptoms and signs, rather than specific biomarkers for a single virus. Essentially, they represent a macroscopic manifestation of downstream pathophysiological changes. This homogeneity at the pathophysiological level determines the universality of the model's feature engineering.
[0082] In addition, the disease grading of both diseases is constructed around indicators such as the patient's clinical manifestations, imaging infiltration range, oxygenation function, and organ failure. The disease grading system is highly consistent with that of pneumonia caused by the new coronavirus infection. Both use imaging pneumonia manifestations as the dividing point between mild and common types. In the mild stage, only upper respiratory tract infection-related symptoms are manifested, and there are no imaging manifestations of pneumonia; in the common stage, both have imaging manifestations of pneumonia, but no hypoxemia (resting blood oxygen saturation >93%); in the severe stage, both are judged based on four key core indicators: respiratory rate (≥30 times / min), blood oxygen saturation (≤93%), oxygenation index (≤300 mmHg), and imaging progression rate (>50%); critical type uses life support needs (mechanical ventilation / shock / multiple organ failure) as the core standard. It can be seen that the disease grading framework, core indicators and thresholds of the two diseases are highly similar, reflecting the common pathological process of viral pneumonia.
[0083] In summary, the clinical manifestations, disease progression patterns, and susceptible populations of novel coronavirus infection and influenza virus pneumonia are consistent across pathogens. Therefore, the method of using this model to identify the disease grading of these two viral pneumonias has the potential to be generalized. This provides a theoretical basis for the migration and application of the disease grading discrimination model constructed based on COVID-19 symptoms and signs to influenza virus pneumonia, so that the viral pneumonia disease grading discrimination model proposed in this invention can be applied not only to the disease grading of patients infected with novel coronavirus infection, but also to the disease grading of patients with influenza virus pneumonia.
[0084] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0085] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements that are inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0086] The above is a detailed introduction to the construction method, discrimination system and application of an instantly interpretable viral pneumonia disease grading discrimination model based on symptoms and signs provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for constructing a viral pneumonia disease grading and discrimination model based on immediate and interpretable symptoms and signs, characterized by: The method comprises: Obtain clinical information of the examinee and construct a sample data set based on the clinical information; Based on the clinical information of the examinee, N preliminary clinical manifestation variables are determined; M variables are selected from N preliminary clinical manifestation variables as clinical symptom and sign feature vectors; Build multiple machine learning models separately, input the clinical symptom and sign feature vectors into each machine learning model, train each model on the sample data set, and determine the optimal model; M clinical symptom and sign feature vectors were incorporated into the optimal model to establish a viral pneumonia disease grading and discrimination model.
2. The method according to claim 1, characterized in that The M variables are screened out from the N preliminary clinical manifestation variables as the clinical symptom and sign feature vector, including: Based on the univariate logistic regression analysis method, N preliminary clinical manifestation variables were screened to determine P statistically significant characteristic variables; Sort the P statistically significant characteristic variables screened out from large to small according to the OR value; Based on the correlation analysis method, the characteristic variables with correlation greater than the correlation coefficient threshold and smaller OR value are screened out from the sorted characteristic variables; The remaining M feature variables are used as clinical symptom and sign feature vectors.
3. The method according to claim 1, characterized in that The clinical symptom and sign feature vector includes: Vomiting, hoarseness, irritability, abdominal pain, shortness of breath, sputum, fullness and distension in the ribs, cough, blurred vision, heavy breathing, fullness and stuffiness, and fever.
4. The method according to claim 1, wherein The method of constructing multiple machine learning models, inputting clinical symptom and sign feature vectors into each machine learning model, performing model training on a sample data set, and determining the optimal model includes: Divide the sample data set into a training set and a test set, and expand the sample data of mild and severe patients in the training set; Build logistic regression model, support vector machine model, random forest model, neural network model and extreme gradient boosting model respectively; On the training set, the above models were trained based on stratified 5-fold cross-validation, and the hyperparameters of each model were optimized in combination with grid search; Sort the M clinical symptom and sign feature vectors by OR value from large to small, and gradually input different numbers of clinical symptom and sign feature vectors into each of the above trained models; The performance of the trained model is evaluated on the test set, and the model with the best performance is selected as the optimal model.
5. The method according to claim 1, wherein The optimal model is the random forest model.
6. The method according to claim 1, characterized in that The method further comprises: Perform SHAP interpretability analysis on the optimal model to determine the contribution of clinical symptom and sign feature vectors to the prediction results; Based on M clinical symptom and sign feature vectors, an ordered logistic regression model was constructed and a nomogram was drawn. The corresponding scores of different clinical symptom and sign feature vectors were displayed according to the nomogram.
7. An immediate and interpretable viral pneumonia grading system based on symptoms and signs, characterized by: The system comprises: A data acquisition module is used to obtain clinical manifestation information of target patients; A nomogram construction module, based on the M characteristic variables included in the viral pneumonia disease grading discriminant model, combines ordered logistic regression to construct a nomogram, and obtains the corresponding scores of different clinical symptom and sign characteristic vectors and the probability of the disease grade according to the nomogram; wherein the viral pneumonia disease grading discriminant model is constructed based on the method described in any one of claims 1-6; The disease classification module is used to calculate the total score of the patient's clinical symptom and sign feature vectors and determine the patient's disease classification based on the total score. The disease classification includes mild, common and severe.
8. A readable storage medium storing a program or instruction, wherein the program or instruction, when executed by a processor, implements the method according to any one of claims 1 to 6.
9. An application of a viral pneumonia disease grading and discrimination model, characterized in that: The viral pneumonia disease grading and discrimination model is applied to the disease grading of patients infected with the new coronavirus.
10. An application of a viral pneumonia disease grading and discrimination model, characterized in that: The viral pneumonia disease grading discriminant model is applied to the disease grading of patients with influenza virus pneumonia.