Ankylosing spondylitis early diagnosis and prediction system based on hematological indexes and demographic characteristics and construction method
By combining hematological indicators and demographic characteristics, using machine learning and SHAP value analysis, an efficient and interpretable AS diagnostic model was developed, which solved the problems of low sensitivity, insufficient specificity and lack of interpretability of existing diagnostic methods, and achieved high accuracy and transparency AS diagnosis.
Patent Information
- Application Number
- CN202510185208.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-30
AI Technical Summary
The existing AS diagnostic methods are low in sensitivity, insufficient specificity and lack of interpretability, making it difficult to achieve accurate diagnosis in the early stages of the disease.
By combining hematological indicators and demographic characteristics, using machine learning technology to screen key features and introducing SHAP value analysis, an efficient and interpretable AS diagnostic model is developed.
It improves the accuracy and transparency of AS diagnosis, provides global and individualized visualization of the importance of characteristics, helps clinicians understand the diagnostic results and formulate personalized treatment plans.
Smart Images

Figure CN120072269A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of medical diagnosis and artificial intelligence, and in particular to a machine learning model and method for the auxiliary diagnosis of ankylosing spondylitis (AS) based on hematological indexes and demographic characteristics. Background Art
[0002] Ankylosing spondylitis (AS) is a chronic inflammatory disease, and early diagnosis is of great significance for the prognosis of patients. However, traditional diagnostic methods such as imaging examinations and laboratory indexes have low sensitivity and specificity, especially in the early stage of the disease. In the prior art, although machine learning algorithms have been widely applied to medical diagnosis, most models lack transparency and interpretability, which limits their clinical application value.
[0003] The present invention combines hematological indexes (such as blood cell count) with demographic characteristics (such as age and gender), and introduces SHAP (Shapley Additive Explanations) value analysis to develop an efficient and interpretable AS diagnosis model, providing diagnostic decision support for clinicians. Summary of the Invention
[0004] The present invention aims to provide an AS diagnosis model based on hematological indexes and demographic characteristics and its implementation method, so as to solve the problems of low sensitivity, insufficient specificity and lack of interpretability of existing diagnostic methods. By using machine learning technology to screen key features and combining SHAP value analysis, the accuracy and transparency of diagnosis are improved.
[0005] Technical Solution
[0006] To solve the above technical problems, the present invention provides an ankylosing spondylitis diagnosis model and method based on hematological indexes and demographic characteristics; the construction of the model includes the collection and processing of data, the model construction and optimization method, the interpretability analysis, and the integration of a diagnosis and decision guidance system;
[0007] In the above solution, for data collection, it is characterized in that: the cases are from the blood samples of 19,693 suspected AS patients in Guangxi Medical University. After excluding 2,186 interfering samples, 17,507 valid data are included. Sampling is carried out in the early morning fasting state to reduce the interference of diet and metabolic factors, and the patient's blood sample is obtained by venous blood collection, and the collection volume is usually 2-5 milliliters. The collected blood sample is placed in a vacuum blood collection tube containing anticoagulant, immediately mixed and sent to the laboratory. The blood sample needs to be refrigerated and stored at 2-8°C after collection, and the maximum storage time does not exceed 24 hours to ensure data accuracy.
[0008] Feature construction includes hematological indicators and demographic characteristics. The hematological indicators include RBC, HGB, HCT, MCV, MCH, MCHC, RDW, WBC, BASO, BASO%, EO, EO%, MONO, MONO%, NEUT, NEUT%, LYMPH, LYMPH%, PCT, MPV, PDW, PLT, etc.; the demographic characteristics include age and gender.
[0009] Data preprocessing, characterized in that: standardize the hematological indicators, handle missing values and outliers to ensure data quality.
[0010] The present invention also provides a model construction and optimization method, characterized in that: Model screening: Train the model and compare the performance of 5 constructed machine learning models; Feature screening: Use the performance-optimal random forest algorithm to rank the feature importance and screen out the 10 features with the optimal diagnostic efficacy; Binary regression analysis: Use the above features to construct an optimized training model and compare the performance of 5 constructed optimized machine learning models, optimize the model parameters to ensure the best model performance; Model performance verification: Through cross-validation and external validation set testing, the model performs excellently in indicators such as AUC, sensitivity and specificity.
[0011] The present invention also provides an interpretability analysis, characterized in that Global feature importance: Use SHAP values to generate a global feature importance map to clarify the contribution ranking of hematological indicators and demographic characteristics in the model; Individualized interpretation: Generate a SHAP interpretation map for each patient and label the positive or negative contribution of the features to the prediction result; Clinical application: The SHAP map helps clinicians understand the model diagnosis logic and provide individualized diagnosis and treatment plans for patients.
[0012] The present invention also provides a system to realize the early auxiliary diagnosis of AS and guide doctors' diagnosis and treatment decisions Data input module: Receive the patient's hematological indicators, age and gender; Diagnostic model module: Call the trained machine learning model and output the AS diagnosis prediction result; Interpretability analysis module: Generate global and individualized SHAP visualization charts; provide diagnostic conclusions and detailed feature contribution analyses.
[0013] Beneficial effects: Multi-dimensional feature fusion: For the first time, hematological indicators are combined with demographic characteristics (age and gender) and applied to AS diagnosis to improve the diagnostic efficacy of the model; Enhanced interpretability: Through SHAP value analysis, provide visualizations of global and individual feature importance to assist clinicians in understanding diagnostic results; Efficient diagnosis: The AUC of the model on the validation set reaches 0.853, with high sensitivity and high accuracy; Early diagnosis: Rely on routine blood tests, without the need for complex imaging or antibody tests; Excellent performance: The model has high sensitivity and specificity, suitable for screening and diagnosis; Personalized support: Improve the transparency of the model through interpretability analysis, facilitating clinical application and patient communication. Brief Description of the Drawings
[0014] Figure 1 : Data processing flow chart: Show the processes of data collection, preprocessing, model training, and testing.
[0015] Figure 2 : ROC curve of the diagnostic model.
[0016] Figure 3 : SHAP global feature importance graph: Show the contribution ranking of key features.
[0017] Figure 4 : SHAP value graph: Show the feature contribution situation.
[0018] Figure 5 : Model performance evaluation graph: Show the AUC curve and the performance of the model on the validation set. Detailed Implementation Manner
[0019] Example 1: Data Collection and Preprocessing
[0020] Collected 19,693 blood samples of suspected AS patients from a certain tertiary hospital. After excluding interfering samples, 17,507 cases of data were included for the study. The blood samples were obtained by venous blood collection, with 2 - 5 milliliters collected for each case and stored in anticoagulant tubes. Detect and record the hematological indexes of all samples, including RBC, HGB, HCT, MCV, MCH, MCHC, RDW, WBC, BASO, BASO%, EO, EO%, MONO, MONO%, NEUT, NEUT%, LYMPH, LYMPH%, PCT, MPV, PDW, PLT, etc., and construct a complete data set in combination with the demographic characteristics (age and gender) of the patients. Standardize the data, exclude outliers, and fill in missing values.
[0021] The differences in age, gender, and various blood cell indices in peripheral blood among groups were compared. The results (Table 1) showed that there were significant differences in the levels of blood cell contents in peripheral blood between AS patients and the non-AS group, both in the training set and the validation set.
[0022]
[0023] Note: Continuous variables are expressed as median (IQR), and categorical variables are expressed as number (%)). Abbreviations: RBC: red blood cell count; HGB: hemoglobin concentration; HCT: hematocrit percentage; MCV: mean corpuscular volume; MCH: mean corpuscular hemoglobin; MCHC: mean corpuscular hemoglobin concentration; RDW: red blood cell distribution width; WBC: white blood cell count; BASO: basophil count; BASO%: basophil percentage; EO: eosinophil count; EO%: eosinophil percentage; MONO: monocyte count; MONO%: monocyte percentage; NEUT: neutrophil count; NEUT%: neutrophil percentage; LYMPH: lymphocyte count, LYMPH%: lymphocyte percentage; PCT: plateletcrit; MPV: mean platelet volume; PDW: platelet distribution width; PLT: platelet count; IQR, interquartile range.
[0024] Example 2: Feature Screening and Model Training
[0025] The preprocessed data was divided into a training set and a validation set. The importance of each feature was calculated using the random forest algorithm, and 10 features with the most diagnostic value were selected in combination with regression analysis. An optimized training model was constructed using these features, and the performances of 5 optimized machine learning models were compared. A binary logistic regression model was constructed and its parameters were optimized. The results (Table 2) showed that the performance of RF was the best, and its AUCs on the training set and the validation set reached 0.884 and 0.853 respectively.
[0026]
[0027] AUC = area under the curve; CART = classification and regression tree; RF = random forest; XGB = extreme gradient boosting; GLM = regularized generalized linear model; NB = naive Bayes.
[0028] Example 3: Model Validation and Performance Evaluation
[0029] The performances of all optimized models were tested using the full-feature validation set. The results (Table 3) showed that the ACC of the RF model with the best performance reached 0.818, the sensitivity was 93.8%, and the specificity was 50.3%.
[0030] Table 3. Performance evaluation metrics of different classification models trained with all variables on the training set and the test set.
[0031]
[0032] NB = Naive Bayes; RF = Random Forest; CART = Classification and Regression Tree; XGB = Extreme Gradient Boosting; GLM = Regularized Generalized Linear Model. ACC = Accuracy; PRE = Precision; SEN = Sensitivity; SPE = Specificity.
[0033] The performance of all optimized models was tested using the validation set of the top ten important features. The results (Table 4) showed that the ACC of the RF model with the best performance reached 0.812, the sensitivity was 91.4%, and the specificity was 54.5%.
[0034]
[0035] Finally, through SHAP value analysis, a global feature importance ranking graph was generated, showing that age, gender, PDW, PCT, PLT, MCH, RBC, MCV, MCHC, and RDW were the top 10 most important diagnostic features. The optimized diagnostic model constructed using the above key features can effectively predict the risk of AS and assist doctors in early screening and diagnosis of AS patients. Applying the model to actual clinical cases, a SHAP explanation graph was generated for each patient to visually display the impact of each feature on the diagnostic result.
[0036] The present invention has been described in detail above. For those skilled in the art, without departing from the gist and scope of the present invention and without unnecessary experiments, the present invention can be implemented within a relatively wide range under equivalent parameters, concentrations, and conditions. Although specific embodiments of the present invention are given, it should be understood that the present invention can be further improved. In short, according to the principle of the present invention, this application intends to cover any modification, use, or improvement of the present invention, including changes made using conventional techniques known in the art that depart from the scope disclosed in this application. Some basic features can be applied according to the scope of the appended claims below.
Claims
1. An early diagnosis and prediction system for ankylosing spondylitis based on hematological indicators and demographic characteristics and a construction method, characterized in that It consists of three modules: (1) Data input module: receiving and storing patients’ hematological indicators and demographic characteristics; (2) Diagnostic model module: The key features are selected through machine learning algorithms, and then a diagnostic model is constructed based on the key features. The model can process the input data and output the predicted diagnosis results of ankylosing spondylitis. (3) Explanatory analysis module: It uses an interpretable machine learning model to provide a SHAP global feature importance ranking diagram and a SHAP dependency diagram, which intuitively displays the impact of each feature on the diagnostic results and provides diagnostic decision support for clinicians.
2. A data input module according to claim 1, characterized in that Combined hematological indicators include but are not limited to RBC, HGB, HCT, MCV, MCH, MCHC, RDW, WBC, BASO, BASO%, EO, EO%, MONO, MONO%, NEUT, NEUT%, LYMPH, LYMPH%, PCT, MPV, PDW, PLT; demographic characteristics include age and gender.
3. A diagnostic model module of an ankylosing spondylitis diagnostic system based on the diagnostic model of claim 1, characterized in that The following steps are involved: (1) Enter the patient's hematological indicators, age, and gender; (2) Use machine learning models to screen important features; (3) Explain the best performing machine learning model and select 10 key features through the SHAP global feature importance ranking graph; (4) Reconstruct five machine learning models using the 10 key features and select the model with the best performance; (5) Binary logistic regression analysis was performed on the blood cell indices of the top 10 features that had a greater impact on model performance to illustrate the association between key features and diseases; (6) The model can process the input data and output the predicted ankylosing spondylitis diagnosis results.
4. An explanatory analysis module of an ankylosing spondylitis diagnosis system based on the diagnosis model of claim 1, characterized in that: (1) Generate SHAP global feature importance graph and SHAP dependency graph; (2) Output diagnostic conclusions and feature contribution analysis.
5. The system according to claim 1, characterized in that The generated SHAP graph can be used to guide clinicians' diagnostic decisions.
Citation Information
Cited By
Method and device for mining related markers of input malaria
CN120833921A