Diagnostic method for predicting tuberculosis risk by using blood routine indexes
By using conventional blood indicators and machine learning algorithms, key indicators are screened out and logistic regression model prediction is carried out, the problems of long diagnosis and detection time, high cost and high requirements of equipment professionals in the existing technology are solved, and rapid and effective tuberculosis risk prediction is achieved, which promotes early detection and precise diagnosis and treatment.
Patent Information
- Application Number
- CN202510225047.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing tuberculosis diagnosis technology has problems such as long testing time, high cost and high requirements for equipment and professionals.
Using blood conventional indicators combined with machine learning algorithms, key blood conventional indicators are screened out through LASSO regression and analysis of multiple machine learning models, and the logistic regression model is used to predict it quickly and efficiently.
It overcomes the problems of long testing time, high cost and high requirements of equipment professionals in the prior art, and provides a fast and effective method for predicting tuberculosis risk, promotes early detection and precise diagnosis and treatment, and alleviates the pressure of medical resources.
Smart Images

Figure CN120148823A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical diagnosis, and particularly to a diagnostic method for predicting the risk of tuberculosis using blood routine indicators. Background Art
[0002] Tuberculosis (TB) is an infectious disease that seriously threatens global public health caused by Mycobacterium tuberculosis (M. tuberculosis). According to the statistics of the World Health Organization (WHO), more than 2 million people die from tuberculosis globally every year. Despite the availability of various vaccines and treatment methods, due to the emergence of drug-resistant tuberculosis strains and the limitations of diagnostic techniques, the early detection and treatment of tuberculosis still face challenges. Therefore, developing a new auxiliary diagnostic method for tuberculosis has important practical significance for improving the diagnostic efficiency, reducing the transmission risk, and improving the prognosis of patients.
[0003] Currently, the diagnosis of tuberculosis mainly relies on traditional bacteriological detection methods, such as sputum smear examination and culture, and molecular biology methods, such as polymerase chain reaction (PCR). However, these methods have certain limitations, including long detection time, high cost, and the need for professional laboratory equipment and technical personnel. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide a diagnostic method for predicting the risk of tuberculosis using blood routine indicators, and to solve the problems of long detection time, high cost, and high requirements for equipment and professionals existing in the existing tuberculosis diagnostic techniques.
[0005] Technical Solution: A diagnostic method for predicting the risk of tuberculosis using blood routine indicators according to the present invention includes the following steps: (1) Collect open-source blood routine data and perform preprocessing, and divide it into a training set, a test set, and a validation set; (2) Screen 25 indicators of blood routine and its derivatives, and obtain a preliminary screening result through LASSO regression analysis; (3) Input the preliminary screening results into seven machine learning models, including a logistic regression model, a random forest model, a naive Bayes model, a K-nearest neighbor model, a support vector machine model, an XGBoost model, and a GBM model, respectively, for analysis to obtain the final variable combination and the best model; (4) Input the validation set into the best model to obtain a DCA decision curve and a calibration curve; Further, in step (1), the preprocessing is specifically to delete or fill missing values, and process outliers; standardize all features (Z-score standardization) to make the features have the same scale.
[0006] Further, in step (2), the preliminary screening results obtained through LASSO regression analysis include: hemoglobin Hb, platelet count PLT, mean corpuscular hemoglobin MCH, mean platelet volume MPV, mean corpuscular hemoglobin concentration MCHC, platelet distribution width PDW, lymphocyte count LYM, monocyte percentage MONO%, lymphocyte percentage LYM%, platelet-lymphocyte ratio PLR, neutrophil-platelet ratio NPR, and derived neutrophil-lymphocyte ratio dNLR.
[0007] Further, in step (3), the final variable combination is hemoglobin Hb, platelet count PLT, mean corpuscular hemoglobin content MCH, mean corpuscular hemoglobin concentration MCHC, platelet distribution width PDW, lymphocyte count LYM, monocyte percentage MONO%, lymphocyte percentage LYM%, neutrophil-platelet ratio NPR; and the corresponding machine learning model is a logistic regression model.
[0008] Further, step (4) also includes: using SHAP to explain individual sample predictions.
[0009] A diagnostic system for predicting tuberculosis risk using blood routine indicators according to the present invention includes: A data acquisition module: used to collect open-source blood routine data, perform preprocessing, and divide it into a training set, a test set, and a validation set; A LASSO regression module: used to screen 25 indicators of blood routine and its derivatives, and obtain preliminary screening results through LASSO regression analysis; A machine learning module: used to input the preliminary screening results into seven machine learning models respectively, including a logistic regression model, a random forest model, a naive Bayes model, a K-nearest neighbor model, a support vector machine model, an XGBoost model, and a GBM model for analysis, and obtain the final variable combination and the best model; A curve module: used to input the validation set into the best model to obtain a DCA decision curve and a calibration curve; Further, in the data acquisition module, the preprocessing is specifically to perform a cleaning operation on the data.
[0010] Further, in the LASSO regression module, the preliminary screening results obtained through LASSO regression analysis include: hemoglobin Hb, platelet count PLT, mean corpuscular hemoglobin MCH, mean platelet volume MPV, mean corpuscular hemoglobin concentration MCHC, platelet distribution width PDW, lymphocyte count LYM, monocyte percentage MONO%, lymphocyte percentage LYM%, platelet-lymphocyte ratio PLR, neutrophil-platelet ratio NPR, and derived neutrophil-lymphocyte ratio dNLR.
[0011] Further, in the machine learning module, the final variable combination is hemoglobin Hb, platelet count PLT, mean corpuscular hemoglobin content MCH, mean corpuscular hemoglobin concentration MCHC, platelet distribution width PDW, lymphocyte count LYM, monocyte percentage MONO%, lymphocyte percentage LYM%, neutrophil-platelet ratio NPR; the corresponding machine learning model is a logistic regression model.
[0012] Further, the curve module also includes: using SHAP to explain individual sample predictions.
[0013] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: The auxiliary diagnosis tool for predicting the risk of tuberculosis using blood routine indicators proposed in the patent combines machine learning algorithms and can quickly and efficiently predict the onset risk of tuberculosis; the present invention overcomes the problems of long detection time, high cost, and high requirements for equipment and professional personnel existing in the existing tuberculosis diagnosis technology, provides strong support for clinical diagnosis, promotes the early detection and precise diagnosis and treatment of tuberculosis, and effectively relieves the pressure on medical resources. In the future, the present invention is expected to be widely applied in aspects such as early screening, precise diagnosis, and treatment effect monitoring of tuberculosis, providing important support for the whole-cycle management of tuberculosis. Description of the Drawings
[0014] Figure 1 is the flow chart of the present invention; Figure 2 is the LASSO regression path diagram of the present invention; Figure 3 is the ten-fold cross-validation result of the LASSO regression of the present invention; Figure 4 is the ROC curve of the best variable combination obtained after stepwise regression of the seven machine learning models of the present invention; Figure 5 is the ROC curve of the seven machine learning models of the present invention on the external validation set; Figure 6 is the DCA decision curve of the logistic regression model of the present invention; Figure 7 The DCA decision curve of the logistic regression model of the present invention Figure 8 The SHAP graph of the logistic regression model of the present invention; Figure 9 The nomogram constructed based on the logistic regression model of the present invention; Figure 10 The web deployment interface of the logistic regression prediction model implemented based on the shiny package of the present invention. Detailed implementation manners
[0015] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0016] As Figures 1 - 10 shown, an embodiment of the present invention provides a diagnostic method for predicting tuberculosis risk using blood routine indexes, including the following steps: S1, Data collection: A total of 3,598 cases of blood routine test results from multiple hospitals were collected. After data cleaning, 3,446 cases were retained, including 728 tuberculosis patients and 2,718 normal people. They were randomly divided into a training set, a test set, and a validation set according to a ratio of 6:2:2.
[0017] S2, As Figures 2 - 3 shown, Index screening: A total of 25 blood routine and its derived indexes were used for screening. Through LASSO regression analysis, the preliminary screening results were obtained: hemoglobin Hb, platelet count PLT, mean corpuscular hemoglobin MCH, mean platelet volume MPV, mean corpuscular hemoglobin concentration MCHC, platelet distribution width PDW, lymphocyte count LYM, monocyte percentage MONO%, lymphocyte percentage LYM%, platelet-lymphocyte ratio PLR, neutrophil-platelet ratio NPR, and derived neutrophil-lymphocyte ratio dNLR. The specific process is as follows: Select the regularization parameter: Use cross-validation to select the best regularization parameter; Fit the LASSO regression model on the training set. Retain the features of the trained LASSO regression model.
[0018] S3, As Figures 4 - 5As shown in the figure, model training: The preliminary screening results were respectively input into the stepwise regression algorithms of seven machine learning models (including logistic regression model, random forest model, naive Bayes model, K-nearest neighbor model, support vector machine model, XGBoost model, GBM model). Finally, the variable combination adopted was determined as hemoglobin (Hb), platelet count (PLT), mean corpuscular hemoglobin (MCH), mean corpuscular hemoglobin concentration (MCHC), platelet distribution width (PDW), lymphocyte count (LYM), monocyte percentage (MONO%), lymphocyte percentage (LYM%), neutrophil-to-platelet ratio (NPR). The corresponding machine learning model was the logistic regression model, whose AUC on the test set reached 0.875, the F1 index was 0.895, and the Youden index was 0.529. The specific process is as follows: a Handle missing values and outliers; standardize the features to ensure they have similar scales. Divide the dataset into a training set, a test set, and a validation set.
[0019] b Select seven machine learning models for analysis: logistic regression model, random forest model, naive Bayes model, K-nearest neighbor model, support vector machine model, XGBoost model, GBM model c Model training and evaluation: Model training: Train the above seven models on the training set and optimize the model parameters using cross-validation. Model evaluation: Evaluate the model performance on the test set, further screen key variables, and the evaluation metrics include: area under the ROC curve (AUC), sensitivity, specificity, accuracy.
[0020] d Determine the best model: Compare the test set performances of the seven models, and select the model with the highest AUC and strong generalization ability as a candidate. Use grid search (GridSearchCV) or random search to tune the hyperparameters of the seven models; select the best model according to the tuned performance.
[0021] As Figures 6 - 8 shown, model external validation: Use the logistic regression model to make predictions on the validation set. The AUC reached 0.809, which was still the best compared to the other six machine learning models. The F1 index was 0.905, and the Youden index was 0.432.
[0022] As Figure 9As shown in the figure, model evaluation and interpretability analysis: The DCA decision curve and calibration curve of the logistic regression model are made on the test set and the external validation set. The DCA decision curve shows that within a range of multiple threshold probabilities, the net benefit of the logistic regression model is higher than the "do nothing" strategy, indicating that the model has stable clinical application value under the discount threshold; the calibration curve shows that the predicted probability of the logistic regression model is relatively consistent with the actual occurrence probability, indicating that the model has good and stable calibration performance. At the same time, in order to study the interpretability of the model, the SHAP swarm plot is used to explain the global importance of each indicator to understand the general impact of various indicators on all samples, and the SHAP force plot is used to explain the prediction of individual samples.
[0023] S4, as Figure 10 As shown in the figure, clinical practical application of the model: A nomogram is made based on the regression coefficients of each predictive factor in the logistic regression model. By mapping each predictive variable to the corresponding score and accumulating these scores to obtain the total score (Total Points), the total score is mapped to the scale of the linear predictor and converted into the probability of disease (Probability of species). This graphical tool not only simplifies the model application process but also enhances the transparency and usability of model prediction, facilitating clinicians to quickly make risk assessments and decisions based on the specific characteristics of patients; on the other hand, the diagnostic model is deployed to a web application through the shiny package. The diagnostic model can be used by logging in to https: / / nana2379723224.shinyapps.io / TBmodel / . By inputting the actual values of the 9 indicators required by the model, this application can automatically predict the risk probability of a single patient having tuberculosis. The development of this online tool not only greatly simplifies the model application process but also significantly improves the efficiency of clinical diagnosis, enabling doctors to more quickly and conveniently conduct risk assessments and decision support for patients.
Claims
1. A diagnostic method for predicting tuberculosis risk using blood routine indicators, characterized in that: The following steps are involved: (1) Collect and preprocess open-source blood routine data and divide them into training set, test set, and validation set; (2) Screening of 25 blood routine tests and their derivatives, and obtaining preliminary screening results through LASSO regression analysis; (3) The preliminary screening results were input into seven machine learning models including logistic regression model, random forest model, naive Bayes model, K-nearest neighbor model, support vector machine model, XGBoost model, and GBM model for analysis to obtain the final variable combination and the best model; (4) Input the validation set into the best model to obtain the DCA decision curve and calibration curve.
2. A diagnostic method for predicting tuberculosis risk using blood routine indicators according to claim 1, characterized in that: In step (1), preprocessing specifically involves deleting or filling missing values and handling outliers; standardizing all features so that they have the same scale.
3. The diagnostic method for predicting tuberculosis risk using blood routine indicators according to claim 1, characterized in that: In step (2), the preliminary screening results obtained through LASSO regression analysis include: hemoglobin Hb, platelet count PLT, mean corpuscular hemoglobin MCH, mean platelet volume MPV, mean corpuscular hemoglobin concentration MCHC, platelet distribution width PDW, lymphocyte count LYM, monocyte percentage MONO%, lymphocyte percentage LYM%, platelet to lymphocyte ratio PLR, neutrophil to platelet ratio NPR and derived neutrophil to lymphocyte ratio dNLR.
4. The diagnostic method for predicting tuberculosis risk using blood routine indicators according to claim 1, characterized in that: In step (3), the final variable combination is hemoglobin Hb, platelet count PLT, mean corpuscular hemoglobin content MCH, mean corpuscular hemoglobin concentration MCHC, platelet distribution width PDW, lymphocyte count LYM, monocyte percentage MONO%, lymphocyte percentage LYM%, and neutrophil to platelet ratio NPR; the corresponding machine learning model is a logistic regression model.
5. The diagnostic method for predicting tuberculosis risk using blood routine indicators according to claim 1, characterized in that: Step (4) also includes: using SHAP to explain single sample predictions.
6. A diagnostic system for predicting tuberculosis risk using blood routine indicators, characterized in that: include: Data collection module: used to collect open-source blood routine data and pre-process it, and divide it into training set, test set, and validation set; LASSO regression module: used to screen blood routine and its derived 25 indicators, and obtain preliminary screening results through LASSO regression analysis; Machine learning module: used to input the preliminary screening results into seven machine learning models including logistic regression model, random forest model, naive Bayes model, K-nearest neighbor model, support vector machine model, XGBoost model, and GBM model for analysis to obtain the final variable combination and the best model; Curve module: used to input the validation set into the best model to obtain the DCA decision curve and calibration curve.
7. A diagnostic system for predicting tuberculosis risk using blood routine indicators according to claim 6, characterized in that: In the data acquisition module, preprocessing specifically involves cleaning the data.
8. A diagnostic system for predicting tuberculosis risk using blood routine indicators according to claim 6, characterized in that: In the LASSO regression module, the preliminary screening results obtained after LASSO regression analysis included: hemoglobin Hb, platelet count PLT, mean corpuscular hemoglobin MCH, mean platelet volume MPV, mean corpuscular hemoglobin concentration MCHC, platelet distribution width PDW, lymphocyte count LYM, monocyte percentage MONO%, lymphocyte percentage LYM%, platelet to lymphocyte ratio PLR, neutrophil to platelet ratio NPR and derived neutrophil to lymphocyte ratio dNLR.
9. A diagnostic system for predicting tuberculosis risk using blood routine indicators according to claim 6, characterized in that: In the machine learning module, the final variable combination is hemoglobin Hb, platelet count PLT, mean corpuscular hemoglobin content MCH, mean corpuscular hemoglobin concentration MCHC, platelet distribution width PDW, lymphocyte count LYM, monocyte percentage MONO%, lymphocyte percentage LYM%, and neutrophil to platelet ratio NPR; the corresponding machine learning model is the logistic regression model.
10. A diagnostic system for predicting tuberculosis risk using blood routine indicators according to claim 6, characterized in that: Also included in the curve module: Using SHAP to explain single sample predictions.
Citation Information
Cited By
Emotional health data processing system and method and computer equipment
CN120511079A
Data processing device and equipment based on blood routine examination and storage medium
CN120954519A