Method for auxiliary diagnosis of multiple types of blood diseases by adopting multi-modal data and machine learning
By combining multimodal data and machine learning methods with routine blood parameters and internal control parameters, an automated diagnostic model is constructed, which solves the problems of time-consuming and labor-intensive initial screening of blood diseases and misdiagnosis and missed diagnosis in existing technologies. It achieves efficient and accurate diagnosis of multiple categories of blood diseases, and has significant advantages, especially in the identification of difficult cases.
Patent Information
- Application Number
- CN202511116594.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-21
AI Technical Summary
Existing methods for initial screening of blood diseases are time-consuming and labor-intensive, and are prone to missed or incorrect detections. They lack efficient and accurate automated classification methods, and are particularly inadequate in identifying difficult blood diseases.
Using multimodal data and machine learning methods, combined with routine blood parameters, internal control parameters and derived indicators, we constructed random forest, XGBoost, LightGBM and neural network models. Through voting ensemble and stacked ensemble strategies, we optimized hyperparameters and established a visual blood analysis interface to assist in the diagnosis of multiple types of blood diseases.
It improves the accuracy and specificity of blood disease diagnosis, significantly reduces the risk of missed diagnosis, provides non-invasive and rapid auxiliary diagnostic support, reduces the interpretation burden on physicians, and improves diagnostic and treatment efficiency.
Smart Images

Figure CN120998462A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of medical diagnosis and artificial intelligence, and particularly relates to a method for auxiliary diagnosis of multi-class blood diseases by using multi-modal data and machine learning. BACKGROUND
[0002] Blood diseases are diseases originating from the hematopoietic system or affecting the hematopoietic system with abnormal changes in blood, characterized by anemia, bleeding, fever, and hepatosplenomegaly and lymphadenopathy. Common blood diseases include anemia, immune thrombocytopenia and other benign blood diseases, and malignant blood diseases including leukemia, lymphoma and multiple myeloma (MM). It is particularly important to improve the understanding of blood diseases in order to detect and treat them early, so as to avoid unnecessary loss of health.
[0003] Currently, the initial screening of blood diseases still takes the "routine peripheral blood + artificial peripheral blood smear microscopy" as the core path: blood routine can quickly indicate abnormality of red blood cells, platelets or white blood cells, etc. The laboratory personnel identify the blood routine results to determine whether to smear microscopy, and further evaluate the cell morphology. From the identification of blood routine data, to the smear microscopy, and then to the report morphology, it requires higher identification ability and morphology experience of the laboratory personnel. Since the artificial microscopy is time-consuming and laborious, and the subtle differences between cells are easy to cause identification errors, and subjective factors will cause differences in identification by different personnel, resulting in missed diagnosis and misdiagnosis of blood diseases. Therefore, there is an urgent need for an efficient, accurate and objective automatic classification method to assist clinicians in improving the diagnostic accuracy and reducing the risk of missed diagnosis. SUMMARY
[0004] The purpose of this section is to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments.
[0005] Modern high-throughput peripheral blood cell analyzers can complete single-tube whole blood testing within 60s, output the following high-dimensional data: blood routine parameters: white blood cell five classification absolute value: WBC, NEUT#, LYMPH#, MONO#, EO#, BASO#, NEUT%, LYMPH%, MONO%, EO%, BASO%; red blood cell series: RBC, HGB, HCT, MCV, MCH, MCHC, RDW-CV, RDW-SD; internal control parameters (parameters calculated based on blood cell analyzer scatter plot or histogram distribution): DlFF channel: Neu subgroup (D_Neu_SSC_P, D_Neu_SSC_W, D_Neu_SSC_CV, D_Neu_SFL_P, D_Neu_SFL_W, D_Neu_SFL_CV, D_Neu_FSC_P, D_Neu_FSC_W, D_Neu_FSC_CV, D_Neu_SSFL_Area, D_Neu_FSFL_Area), Lym subgroup (D_Lym_SSC_P, D_Lym_SSC_W, D_Lym_SSC_CV, D_Lym_SFL_P, D_Lym_SFL_W, D_Lym_SFL_CV, D_Lym_FSC_P, D_Lym_FSC_W, D_Lym_FSC_CV, D_Lym_SSFL_Area, D_Lym_FSFL_Area), Mon subgroup (D_Mon_SSC_P, D_Mon_SFL_P, D_Mon_FSC_P), LymMon subgroup (D_LymMon_SSC_P, D_LymMon_SFL_P, D_LymMon_FSC_P), subgroup separation index (D_Lym_Ratio_SSC, D_Lym_Ratio_SFL, D_Lym_Ratio_FSC, D_Lym_Mon_Dist); WNB channel (N_WBC_SSC_P, N_WBC_SSC_W, N_WBC_SSC_CV, N_WBC_SFL_P, N_WBC_SFL_W, N_WBC_SFL_CV, N_WBC_FSC_P, N_WBC_FSC_W, N_WBC_FSC_CV, N_WBC_FLFS_Area, N_WBC_FLSS_Area, N_WBC_SSFS_Area); RET channel (R_RBC_SSC_P, R_RBC_SSC_W, R_RBC_SSC_CV, R_RBC_SFL_P, R_RBC_SFL_W, R_RBC_SFL_CV, R_RBC_FSC_P, R_RBC_FSC_W, R_RBC_FSC_CV, R_RBC_FLFS_Area, R_RBC_FLSS_Area, R_RBC_SSFS_Area);Impedance channels (I_RBC_MFV, I_RDW_SD, I_RDW_CV, I_PLT_MFV, I_PDW_SD, I_PDW_CV). Derived indicators (parameters calculated based on blood routine indicators) Neu_Lym_Ratio (neutrophil / lymphocyte ratio), WBC_RBC_Ratio (white blood cell / red blood cell ratio), HGB_MCH_Ratio (hemoglobin / mean corpuscular hemoglobin ratio).
[0006] The present application provides a method for diagnosing multiple categories of blood diseases with the aid of multi-modal data and machine learning. The purpose of the present application is to provide an automated classification method for multiple blood tests including leukemia, atypical lymphocytosis, infection, anemia, thrombocytopenia, and granulocytopenia. Without the need for additional blood smears and staining imaging equipment, the method can utilize conventional blood test numerical indicators to enhance the discrimination ability based on blood routine parameters, internal control parameters, and derived indicators. The method combines the super parameter optimization and voting integration of multiple single models, and gives appropriate weights to key or rare categories in the integration process, thereby improving the overall category performance and the accuracy and specificity of multi-category abnormal blood test classification. In particular, the method has significant advantages in identifying difficult blood diseases and improves the diagnostic ability of clinically significant and difficult blood diseases.
[0007] Specifically, the present application provides a method for diagnosing multiple categories of blood diseases with the aid of multi-modal data and machine learning, which comprises the following steps,
[0008] Step 1: Collecting gender, age, blood routine parameters, internal control parameters, and blood routine derived indicator data, preprocessing and standardizing the collected data, and obtaining a feature matrix;
[0009] Step 2: Building random forest model, XGBoost model, LightGBM model, and neural network model, and simultaneously building voting integration model and stacking integration model, optimizing super parameters, and establishing evaluation indicators;
[0010] Step 3: Establishing a visual blood analysis interface based on the trained model in step 2 to assist in predicting and identifying blood diseases;
[0011] The internal control parameters are parameters calculated based on the scatter plot or histogram distribution of the blood cell analyzer, including the barycenter position, width, coefficient of variation, area, and distance of cell subgroups. The blood routine derived indicators include the ratio of neutrophils to lymphocytes, the ratio of white blood cells to red blood cells, and the ratio of hemoglobin to mean corpuscular hemoglobin.
[0012] As a preferred scheme of the method for assisting in diagnosing multiple categories of blood diseases by using multi-modal data and machine learning according to the application: in step 1, the preprocessing includes removing abnormally large values by using the quartile range method; the standardization conversion includes performing one-hot encoding on the gender parameter, converting it into a numerical feature, first standardizing it by using StandardScaler, and then processing non-normal distribution data by using PowerTransformer to eliminate the dimension effect; the data after the preprocessing and the standardization conversion are divided according to the training set: verification set = 8:2, the SMOTETomek oversampling strategy is used on the training set to balance the samples of the minority classes of granulocyte deficiency and atypical lymphocytes, and a feature matrix is obtained.
[0013] As a preferred scheme of the method for assisting in diagnosing multiple categories of blood diseases by using multi-modal data and machine learning according to the application: in step 1, the blood routine parameters include white blood cell count, lymphocyte percentage, monocyte percentage, neutrophil percentage, eosinophil percentage, basophil percentage, lymphocyte absolute value, monocyte absolute value, neutrophil absolute value, eosinophil absolute value, basophil absolute value, red blood cell count, hemoglobin concentration, hematocrit, mean corpuscular volume, mean corpuscular hemoglobin content, mean corpuscular hemoglobin concentration, red blood cell distribution width SD, red blood cell distribution width, platelet count, platelet distribution width, mean platelet volume, platelet hematocrit, and large platelet ratio.
[0014] As a preferred scheme of the method for assisting in diagnosing multiple categories of blood diseases by using multi-modal data and machine learning according to the application: in step 2, the random forest model includes setting the maximum depth of the tree to 15, the minimum number of samples leaves required for splitting the internal node to 5, giving 5 times weight to the leukemia samples, and setting the random number seed to 42.
[0015] As a preferred scheme of the method for assisting in diagnosing multiple categories of blood diseases by using multi-modal data and machine learning according to the application: in step 2, the XGBoost model includes setting the learning rate to 0.05, the maximum depth of the tree to 8, the minimum sample weight of each leaf node to 3, setting the random number seed to 42, and giving 3 times weight to the leukemia samples.
[0016] As a preferred scheme of the method for assisting in diagnosing multiple categories of blood diseases by using multi-modal data and machine learning according to the application: in step 2, the LightGBM model includes setting the maximum depth of the tree to 10, the minimum number of samples of the leaf node to 2, not allowing zero gain splitting, the number of leaf nodes to 15, the learning rate to 0.05, the proportion of features used in each iteration to 80%, and setting the random number seed to 42.
[0017] As a preferred scheme of the method for assisting in diagnosing multi-class blood diseases by using multi-modal data and machine learning provided by the application, in step 2, the neural network model comprises two layers of hidden layers (128, 64), an adaptive optimizer (Adam), and is prevented from overfitting by an early stopping mechanism; wherein the two layers of hidden layers have 128 and 64 neurons respectively, use a ReLU activation function, use an Adam optimizer, adopt an L2 regularization coefficient of 0.001, have a batch size of 64, have an initial learning rate of 0.001, and then the learning rate is set to be adaptively adjusted, the early stopping mechanism is adopted, the maximum number of iterations is 500, and the random number seed is set to be 42.
[0018] As a preferred scheme of the method for assisting in diagnosing multi-class blood diseases by using multi-modal data and machine learning provided by the application, in step 2, the voting ensemble model comprises a random forest model, an XGBoost model, a LightGBM model and a neural network model, and adopts a soft voting strategy with weights being weights=[1, 1, 1, 0.8] in sequence.
[0019] As a preferred scheme of the method for assisting in diagnosing multi-class blood diseases by using multi-modal data and machine learning provided by the application, in step 2, the stacking ensemble model takes the three tree models, i.e., the random forest model, the XGBoost model, the LightGBM model and the neural network model, as base learners, takes the XGBoost model as a meta-learner, and adopts five-fold stratified cross-validation as an internal validation mechanism.
[0020] As a preferred scheme of the method for assisting in diagnosing multi-class blood diseases by using multi-modal data and machine learning provided by the application, in step 3, the establishment of the visual blood analysis interface comprises establishment of a simple PYQT blood analysis interface.
[0021] The application has the following beneficial effects: The application constructs a multi-stage modeling framework covering data enhancement (SMOTETomek mixed sampling), grid search super parameter optimization and ensemble learning (Stacking Ensemble) around blood routine test data, instrument internal control parameters and derived indicators. Experimental results show that the overall classification accuracy is 93.97% (95% CI: 92.1-95.6%); the recall rate of leukemia category is 100%, and the F1-score is 100%.
[0022] Based on the above optimal model, the application further develops a lightweight blood analysis interface based on PyQt, which can real-time predict 6 types of blood abnormalities including leukemia, atypical lymphocytosis, infection, anemia, thrombocytopenia and granulocytopenia by only inputting blood routine parameters. The system can be directly embedded into the clinical workflow to provide non-invasive and rapid auxiliary diagnosis support for doctors, significantly reducing the burden of manual interpretation and improving the efficiency of diagnosis and treatment. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed to be used in the embodiment description will be briefly introduced as follows:
[0024] Figure 1 The overall flowchart of the application.
[0025] Figure 2 The comparative radar chart of the performance difference of the random forest, XGBoost, LightGBM, neural network, voting ensemble and stacking ensemble model established by the application on the leukemia detection index.
[0026] Figure 3 The cross-model feature importance heat map of the application to explore the difference in abnormal blood index of different models.
[0027] Figure 4 The F1-score comparative analysis chart of the application for evaluating the recognition ability of the model on different blood diseases.
[0028] Figure 5 The width, barycenter position and SS signal distribution histogram example of the DIFF channel Neu. DETAILED DESCRIPTION
[0029] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the specific embodiments of the application will be described in detail below.
[0030] Embodiment 1:
[0031] The application provides a method for auxiliary diagnosis of multi-class blood diseases by using multi-modal data and machine learning, comprising the following steps,
[0032] Step 1, data acquisition (90 index data of gender, age, blood routine parameters, internal control parameters and derived indexes), pre-processing and standardization conversion of the collected data to obtain a feature matrix that can be input into the model.
[0033] The data set used in the embodiment contains more than 1000 samples of leukemia, heterotypic lymphocytosis, infection, anemia, thrombocytopenia, granulocytopenia and healthy controls, including 90 index data of gender, age, blood routine parameters, internal control parameters, derived indicators in data characteristics, and different blood counts are respectively divided into value labels.
[0034] Among them, the blood routine detection indexes include: white blood cell count (WBC), lymphocyte percentage (LYMPH%), monocyte percentage (MONO%), neutrophil percentage (NEUT%), eosinophil percentage (EO%), basophil percentage (BASO%), lymphocyte absolute value (LYMPH#), monocyte absolute value (MONO#), neutrophil absolute value (NEUT#), eosinophil absolute value (EO#), basophil absolute value (BASO#), red blood cell count (RBC), hemoglobin concentration (HGB), hematocrit (HCT), mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH), mean corpuscular hemoglobin concentration (MCHC), red blood cell distribution width SD (RDW-SD), red blood cell distribution width CV (RDW-CV), platelet count (PLT), platelet distribution width (PDW), mean platelet volume (MPV), platelet hematocrit (PCT), platelet-large cell ratio (P-LCR) blood routine detection data.
[0035] Internal control parameters: Internal control parameters are parameters calculated based on the scatter plot or histogram distribution of the blood cell analyzer, including the barycenter position (P), width (W), coefficient of variation (CV), area (Area) and distance (Dist) of the cell subpopulation, wherein: the barycenter position is the average value of the signal in a specific direction of the cell subpopulation; the width is the distribution width of the signal in a specific direction of the cell subpopulation; the coefficient of variation is the ratio of the distribution width to the barycenter position; the area is the area of the cell subpopulation in the scatter plot; the distance is the distance between the barycenters of two different cell subpopulations.
[0036] Internal control parameters in DIFF channel:
[0037] 1) D_Neu_SSC_P: Barycenter position of DIFF channel Neu particle group in SS direction
[0038] 2) D_Neu_SSC_W: Distribution width of DIFF channel Neu particle group in SS direction
[0039] 3) D_Neu_SSC_CV: Coefficient of variation of DIFF channel Neu particle group in SS direction
[0040] 4) D_Neu_SFL_P: Barycenter position of DIFF channel Neu particle group in FL direction
[0041] 5) D_Neu_SFL_W: Width of distribution of Neu particle clusters in the FL direction for the DIFF channel
[0042] 6) D_Neu_SFL_CV: Coefficient of variation of Neu particle clusters in the FL direction for the DIFF channel
[0043] 7) D_Neu_FSC_P: Position of center of mass of Neu particle clusters in the FS direction for the DIFF channel
[0044] 8) D_Neu_FSC_W: Width of distribution of Neu particle clusters in the FS direction for the DIFF channel
[0045] 9) D_Neu_FSC_CV: Coefficient of variation of Neu particle clusters in the FS direction for the DIFF channel
[0046] 10) D_Lym_SSC_P: Position of center of mass of Lym particle clusters in the SS direction for the DIFF channel
[0047] 11) D_Lym_SSC_W: Width of distribution of Lym particle clusters in the SS direction for the DIFF channel
[0048] 12) D_Lym_SSC_CV: Coefficient of variation of Lym particle clusters in the SS direction for the DIFF channel
[0049] 13) D_Lym_SFL_P: Position of center of mass of Lym particle clusters in the FL direction for the DIFF channel
[0050] 14) D_Lym_SFL_W: Width of distribution of Lym particle clusters in the FL direction for the DIFF channel
[0051] 15) D_Lym_SFL_CV: Coefficient of variation of Lym particle clusters in the FL direction for the DIFF channel
[0052] 16) D_Lym_FSC_P: Position of center of mass of Lym particle clusters in the FS direction for the DIFF channel
[0053] 17) D_Lym_FSC_W: Width of distribution of Lym particle clusters in the FS direction for the DIFF channel
[0054] 18) D_Lym_FSC_CV: Coefficient of variation of Lym particle clusters in the FS direction for the DIFF channel
[0055] 19) D_Mon_SSC_P: Position of center of mass of Mon particle clusters in the SS direction for the DIFF channel
[0056] 20) D_Mon_SFL_P: DIFF channel Mon cluster's barycenter position in FL direction
[0057] 21) D_Mon_FSC_P: DIFF channel Mon cluster's barycenter position in FS direction
[0058] 22) D_LymMon_SSC_P: DIFF channel LymMon cluster's barycenter position in SS direction
[0059] 23) D_LymMon_SFL_P: DIFF channel LymMon cluster's barycenter position in FL direction
[0060] 24) D_LymMon_FSC_P: DIFF channel LymMon cluster's barycenter position in FS direction
[0061] 25) D_Neu_SSFL_Area: DIFF channel Neu cluster's projected area in SSFL direction
[0062] 26) D_Neu_FSFL_Area: DIFF channel Neu cluster's projected area in FSFL direction
[0063] 27) D_Lym_SSFL_Area: DIFF channel Lym cluster's projected area in SSFL direction
[0064] 28) D_Lym_FSFL_Area: DIFF channel Lym cluster's projected area in FSFL direction
[0065] 29) D_Lym_Ratio_SSC: Ratio of Lym's distribution width in SS direction to the distance between LymMon's barycenter position in SS direction, reflecting the separation degree of two clusters in SS direction
[0066] 30) D_Lym_Ratio_SFL: Ratio of Lym's distribution width in FL direction to the distance between LymMon's barycenter position in FL direction, reflecting the separation degree of two clusters in FL direction
[0067] 31) D_Lym_Ratio_FSC: Ratio of Lym's distribution width in FS direction to the distance between LymMon's barycenter position in FS direction, reflecting the separation degree of two clusters in FS direction
[0068] 32) D_Lym_Mon_Dist: DIFF channel LymMon cluster's barycenter distance
[0069] WNB channel internal control parameters:
[0070] 1) N_WBC_SSC_P: WBC particle cluster in WNB channel in SS direction of the center of gravity position
[0071] 2) N_WBC_SSC_W: WBC particle cluster in WNB channel in SS direction of the distribution width
[0072] 3) N_WBC_SSC_CV: WBC particle cluster in WNB channel in SS direction of the coefficient of variation
[0073] 4) N_WBC_SFL_P: WBC particle cluster in WNB channel in FL direction of the center of gravity position
[0074] 5) N_WBC_SFL_W: WBC particle cluster in WNB channel in FL direction of the distribution width
[0075] 6) N_WBC_SFL_CV: WBC particle cluster in WNB channel in FL direction of the coefficient of variation
[0076] 7) N_WBC_FSC_P: WBC particle cluster in WNB channel in FS direction of the center of gravity position
[0077] 8) N_WBC_FSC_W: WBC particle cluster in WNB channel in FS direction of the distribution width
[0078] 9) N_WBC_FSC_CV: WBC particle cluster in WNB channel in FS direction of the coefficient of variation
[0079] 10) N_WBC_FLFS_Area: WBC particle cluster in WNB channel in FSFL direction of the projection area
[0080] 11) N_WBC_FLSS_Area: WBC particle cluster in WNB channel in SSFL direction of the projection area
[0081] 12) N_WBC_SSFS_Area: WBC particle cluster in WNB channel in SSFS direction of the projection area.
[0082] RET channel internal control parameters:
[0083] 1) R_RBC_SSC_P: RBC particle cluster in RET channel in SS direction of the center of gravity position
[0084] 2) R_RBC_SSC_W: RBC particle cluster in RET channel in SS direction of the distribution width
[0085] 3) R_RBC_SSC_CV: RBC particle cluster in RET channel in SS direction of the coefficient of variation
[0086] 4) R_RBC_SFL_P: RET channel RBC particle group center of gravity position in FL direction
[0087] 5) R_RBC_SFL_W: RET channel RBC particle group distribution width in FL direction
[0088] 6) R_RBC_SFL_CV: RET channel RBC particle group coefficient of variation in FL direction
[0089] 7) R_RBC_FSC_P: RET channel RBC particle group center of gravity position in FS direction
[0090] 8) R_RBC_FSC_W: RET channel RBC particle group distribution width in FS direction
[0091] 9) R_RBC_FSC_CV: RET channel RBC particle group coefficient of variation in FS direction
[0092] 10) R_RBC_FLFS_Area: RET channel RBC particle group projection area in FSFL direction
[0093] 11) R_RBC_FLSS_Area: RET channel RBC particle group projection area in SSFL direction
[0094] 12) R_RBC_SSFS_Area: RET channel RBC particle group projection area in SSFS direction.
[0095] Impedance channel internal control parameters:
[0096] 1) I_RBC_MFV: RBC histogram peak position corresponding volume
[0097] 2) I_RDW_SD: Impedance channel red blood cell distribution width standard deviation
[0098] 3) I_RDW_CV: Impedance channel red blood cell distribution width coefficient of variation
[0099] 4) I_PLT_MFV: PLT histogram peak position corresponding volume
[0100] 5) I_PDW_SD: Impedance channel platelet distribution width standard deviation
[0101] 6) I_PDW_CV: Impedance channel platelet distribution width coefficient of variation.
[0102] Three derived indicators: neutrophil to lymphocyte ratio (Neu_Lym_Ratio), white blood cell to red blood cell ratio (WBC_RBC_Ratio), hemoglobin to mean corpuscular hemoglobin ratio (HGB_MCH_Ratio).
[0103] The single label column is set in this study, and there are a total of 7 types of labels as follows: healthy control (Label=0), leukemia (1), atypical lymphocytes (2), infection (3), anemia (4), thrombocytopenia (5), and granulocytopenia (6).
[0104] Step 1 includes:
[0105] (S1.1) Collect 90 indicators including gender, age, blood routine parameters, internal control parameters, and derived indicators, and divide the different indicators into value labels respectively.
[0106] (S1.2) Create three derived indicators according to the blood routine parameters: neutrophil to lymphocyte ratio (Neu_Lym_Ratio), white blood cell to red blood cell ratio (WBC_RBC_Ratio), and hemoglobin to mean corpuscular hemoglobin ratio (HGB_MCH_Ratio).
[0107] (S1.3) Abnormal value processing: remove abnormal large values greater than Q3+3IQR in the feature using the IQR (interquartile range) method, and retain abnormal small values that may have clinical significance.
[0108] (S1.4) One-hot encoding is performed on the gender parameter to convert it into a numerical feature.
[0109] (S1.5) Standardize first, then use PowerTransformer to process non-normal distribution data to eliminate the influence of dimension.
[0110] (S1.6) Divide the above obtained data set according to the training set: validation set = 8:2, and use SMOTETomek oversampling strategy on the training set to balance the minority class granulocytopenia and atypical lymphocyte samples, and obtain the feature matrix that can be input into the model.
[0111] Step 2, build 4 single models, and build an integrated model based on voting and stacking strategies, optimize the hyperparameters of the integrated model through cross-validation and grid search, establish evaluation indicators, and evaluate the model effect through visualization means.
[0112] Step 2 includes:
[0113] (S2.1) Construct a random forest model to balance class weights by including tuning parameters such as n_estimators, max_depth, and min_samples_split.
[0114] Specifically, set n_estimators = 200: train 200 trees, max_depth = 15: the maximum depth of the tree is 15, min_samples_leaf = 5: the minimum number of samples required to split an internal node is 5, class_weight = {0: 1, 1: 5}: handle class imbalance, assign 5 times weight to leukemia samples, random_state = 42, set the random number seed to 42.
[0115] (S2.2) Construct an XGBoost model to optimize learning rate, tree depth, subsampling rate, and other parameters, and introduce scale_pos_weight to improve the attention of minority classes, especially in the identification of rare blood disease categories such as leukemia and granulocytopenia.
[0116] Specifically, set n_estimators = 200: the number of weak learners is 200, learninq_rate = 0.05: the learning rate is set to 0.05, max_depth = 8: the maximum depth of the tree is set to 8, min_child_weight = 3: the minimum sample weight of each leaf node is 3, gamma = 0.1: the minimum loss reduction required for node splitting, subsample = 0.8: use 80% of the samples to train each tree, colsample_bytree = 0.8: use 80% of the features for each tree, random_state = 42, set the random number seed to 42, scale_pos_weight = 3: assign 3 times weight to leukemia samples to handle class imbalance.
[0117] (S2.3) Construct a LightGBM model to adjust num_leaves, feature_fraction, min_child_samples, and other parameters, and calculate feature importance based on gain.
[0118] Specifically, n_estimators = 200: 200 trees are trained, max_depth = 10: the maximum depth of the tree is 10, min_child_samples = 2: the minimum number of samples for a leaf node is 2, min_gain_to_split = 0: zero gain splitting is not allowed, num_leaves = 15: the number of leaf nodes is 15, learning_rate = 0.05: the learning rate is set to 0.05, feature_fraction = 0.8: the proportion of features used in each iteration is 80%, the random number seed is set to 42, the internal warning output is turned off, and the feature importance evaluation is based on information gain.
[0119] (S2.4) A feedforward neural network model is constructed with two hidden layers (128, 64), an adaptive optimizer (Adam), and an early stopping mechanism to prevent overfitting.
[0120] Specifically, hidden_layer_sizes = (128, 64): two hidden layers with 128 and 64 neurons respectively, activation ='relu': ReLU activation function is used, solver = 'adam': Adam optimizer is used, alpha = 0.001: L2 regularization coefficient 0.001 is used, batch_size = 64: batch size is 64, learning_rate_init = 0.001: initial learning rate is 0.001, learning_rate = 'adaptive': learning rate is then set to adaptive adjustment, early stopping strategy is used, maximum number of iterations is 500, and random number seed is set to 42.
[0121] (S2.5) A voting ensemble model is constructed, combining random forest, XGBoost, LightGBM, and neural network, using a soft voting strategy, giving tree models higher weights. The weights are weights = [1, 1, 1, 0.8] in order.
[0122] (S2.6) A stacked ensemble model is constructed, taking three tree models and a neural network as base learners, and XGBoost as a meta-learner, using five-fold stratified cross-validation as an internal validation mechanism.
[0123] Specifically, GridSearchCV is used in combination with StratifiedKFold(n_splits=5) to search for the optimal parameter combination with f1_macro (macro-averaged F1 score) as the optimization index. Set n_estimators: [150, 200, 250], i.e. the model will try different configurations of 150, 200 and 250 trees, and select the most suitable one. max_depth: [5, 7, 10], i.e. the model will try different configurations of the maximum depth of the tree as 5, 7 and 10. min_samples_split: [2, 5, 10], i.e. try the minimum number of split samples required for each internal node split as 2, 5 and 10. min_samples_leaf: [1, 2, 4], i.e. try the minimum number of leaf samples for each leaf node as 1, 2 and 4, to ensure the optimal solution of the overall and each category performance and get the optimal model as the internal validation mechanism to get the optimal model.
[0124] (S2.7) Construct visual evaluation indicators such as multi-index radar chart and feature importance heat map, establish leukemia recall rate formula and specificity formula, balance precision and recall, and establish the required model with F1-score as the optimization target.
[0125] Specifically,
[0126] Establish the leukemia recall rate formula and measure the detection ability of the model for leukemia samples to avoid missed diagnosis, where TP: true positive, FN: false negative. Establish the specificity formula reflect the correct identification rate of the model for non-leukemia samples to reduce misdiagnosis, where TN: true negative, FP: false positive. Establish the accuracy formula Evaluate the overall classification accuracy of the model, suitable for balanced class data. Establish the F1-score formula
[0127] where
[0128] Recall = Sensitivity, balance precision and recall, especially suitable for evaluation of rare diseases such as leukemia, and the optimal stacked ensemble strategy model is obtained after multi-index evaluation, with an accuracy of 93.97%, 100% leukemia recall rate, and 100% leukemia F1-Score score as the optimal training model.
[0129] Step 3, according to the trained model to establish a visual blood analysis interface, realize the prediction and discrimination of the corresponding blood disease classification of the sample, and assist medical diagnosis.
[0130] Step 3 includes:
[0131] (S3.1) using the trained model to establish a simple PYQT blood analysis interface.
[0132] (S3.2) inputting routine blood parameters, internal control parameters, age and gender data.
[0133] (S3.3) real-time prediction of six types of blood abnormalities, including leukemia, atypical lymphocytosis, infection, anemia, thrombocytopenia and granulocytopenia.
[0134] To sum up, the application firstly receives 90 index data including gender, age, routine blood parameters, internal control parameters and derived indicators, and performs preprocessing operations such as missing value processing, abnormal value cleaning, standardization conversion, category coding and construction of internal derived features, i.e. internal control parameters; secondly, the SMOTETomek oversampling strategy is used to alleviate the class imbalance problem; then single models such as random forest, XGBoost, LightGBM and neural network are constructed, and a fusion model is constructed based on voting integration and stacking integration strategy, and hyperparameter optimization is performed through cross-validation and grid search; the key feature contribution and classification performance are visualized by heat map and radar chart, and the accuracy of the stacking integration model reaches 0.9397, the leukemia recall rate reaches 1.000, and the leukemia index F1-score reaches 1.000, which shows that the model method has the advantages of strong identification ability for minority classes of major difficult blood diseases, and can be widely applied to the clinical intelligent auxiliary diagnosis scene of blood diseases and related abnormal blood.
[0135] Figure 1 The present application is a general flowchart. After data collection and input, the model is established, the reliability is evaluated according to the indicators to obtain the best model, and the function of auxiliary medical diagnosis for new data is realized.
[0136] Figure 2 The present application is a comparison radar chart of the performance difference of the random forest, XGBoost, LightGBM, neural network, voting integration and stacking integration model for leukemia detection indicators. The evaluation indicators include Accuracy, Leukemia Recall, Leukemia F1-Score, and the results show that the stacking integration model leads in the above three indicators, with a recall rate of 0.935, an accuracy of 1.000 and an F1 score of 0.996. The voting integration and random forest model have the same performance of 0.923 in the recall rate and accuracy dimensions, and the neural network has a significantly lower leukemia recall rate, indicating that it has limitations in identifying sensitive diseases. Therefore, the best stacking integration model is finally adopted.
[0137] Figure 3The cross-model feature importance heat map of the difference of abnormal blood index indicators in different models is discussed. In order to identify the core index difference of different models in identifying abnormal blood, the first ten normalized importance of 24 blood index data of random forest model, XGBoost model and LightGBM model are evaluated by the graph, and it is found that: the importance of Neu# neutrophil count in random forest model is 1.00, which is 0.78 in XGB and 0.68 in LGB, which reflects the pathological characteristics of abnormal increase of neutrophils in leukemia patients. The PLT blood platelet count index has the highest importance of 1.00 in the LGB model, and the second highest is 0.89 in the random forest model, which confirms the mechanism of platelet reduction caused by leukemia bone marrow suppression. The specificity of the model also exists: the dependence of XGB on Nrbc% nucleated red blood cell percentage reaches 1.00, while the importance of this feature in random forest and LGB is 0, which indicates the advantage of XGB in identifying special cell types; random forest and LGB mainly rely on cell count indicators (Neu# / PLT / HGB), and XGB pays more attention to artificial construction ratio features (such as Neu_Lym_Ratio). The proportion of derived indicators in the top 5 feature importance is 60%, which shows that the contribution of the derived parameters of neutrophil and lymphocyte ratio (Neu_Lym_Ratio), white blood cell and red blood cell ratio (WBC_RBC_Ratio) and hemoglobin and mean hemoglobin content ratio (HGB_MCH_Ratio) in the model is significantly higher than that of the conventional original indicators. The graph represents the innovation of the internal control parameter theory and the derived features added for research.
[0138] Figure 4The F1-SCOre comparison analysis chart of the model for evaluating the recognition ability of different blood diseases of the application. From left to right, in order are "healthy control", "leukemia", "atypical lymphocytes", "infection", "anemia", "thrombocytopenia", "granulocytopenia", and in the comparison, "Stackinq" is observed, that is, the stacking integrated model can be seen that the finally determined model, that is, the stacking model reaches 99% in the F1-Score of leukemia detection classification, which is 19 percentage points higher than the single model XGBoost of 80%, the voting integrated model is 12.3% higher than the single model in the F1-Score of atypical lymphocytes, infection and other categories, and some single models also have some specific performances: the recognition rate of random forest for healthy control reaches 100.0% (F1-score = 1.0), which is due to the centralized feature distribution of healthy samples; the neural network has a F1-score of 0.87 in the anemia category, which is better than other models (average 0.82), which is due to the nonlinear feature fitting capability, and the granulocytopenia pathological feature is less than 90% in all models, while the stacking integrated model can reach 89.8%, which shows the advantage of the stacking integrated model.
[0139] Figure 5 The width, barycenter position and SS signal distribution histogram example of the DIFF channel Neu.
[0140] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.
Claims
1. A method for diagnosing multiple classes of blood diseases with the assistance of multi-modal data and machine learning, characterized in that: The method comprises the following steps: Step 1: Collecting gender, age, blood routine parameters, internal control parameters and blood routine derived index data, preprocessing and standardizing the collected data, and obtaining a feature matrix; Step 2: Constructing a random forest model, an XGBoost model, a LightGBM model and a neural network model, simultaneously constructing a voting ensemble model and a stacking ensemble model, optimizing hyperparameters and establishing evaluation indexes; Step 3: Establishing a visual blood analysis interface according to the model trained in step 2 to assist in predicting and identifying blood diseases; The internal control parameters are parameters calculated based on the scatter plot or histogram distribution of the blood cell analyzer, including the barycenter position, width, coefficient of variation, area and distance of the cell subpopulation; the blood routine derived indexes include the neutrophil-to-lymphocyte ratio, white blood cell-to-red blood cell ratio and hemoglobin-to-mean hemoglobin content ratio.
2. The method of claim 1, wherein the method is for diagnosing multiple classes of blood diseases with multi-modal data and machine learning assistance. In step 1, the preprocessing includes removing abnormal large values by using the interquartile range method; the standardization conversion includes one-hot encoding the gender parameter to convert it into a numerical feature, first standardizing it by StandardScaler, then processing the non-normal distribution data by PowerTransformer to eliminate the dimension effect; the data after preprocessing and standardization conversion are divided into training set: validation set = 8:2, the training set is subjected to SMOTETomek oversampling strategy to balance the samples of the minority class granulocyte deficiency and atypical lymphocytes, and a feature matrix is obtained.
3. The method for diagnosing multiple categories of blood diseases with the aid of multi-modal data and machine learning according to claim 1 or 2, characterized in that: In step 1, the blood routine parameters include white blood cell count, lymphocyte percentage, monocyte percentage, neutrophil percentage, eosinophil percentage, basophil percentage, lymphocyte absolute value, monocyte absolute value, neutrophil absolute value, eosinophil absolute value, basophil absolute value, red blood cell count, hemoglobin concentration, hematocrit, mean corpuscular volume, mean hemoglobin content, mean hemoglobin concentration, red blood cell distribution width SD, red blood cell distribution width, platelet count, platelet distribution width, mean platelet volume, platelet hematocrit and large platelet ratio.
4. The method for diagnosing multiple categories of blood diseases with the aid of multi-modal data and machine learning according to claim 1 or 2, characterized in that: In step 2, the random forest model includes setting the maximum depth of the tree to 15, the minimum number of samples leaves required for splitting an internal node to 5, assigning 5 times weight to the leukemia samples, and setting the random number seed to 42.
5. The method for diagnosing multi-class blood diseases with the aid of multi-modal data and machine learning according to claim 1 or 2, characterized in that: In step 2, the XGBoost model includes setting the learning rate to 0.05, the maximum depth of the tree to 8, the minimum sample weight of each leaf node to 3, setting the random number seed to 42, and assigning 3 times weight to the leukemia samples.
6. The method for diagnosing multiple categories of blood diseases with the aid of multi-modal data and machine learning according to claim 1 or 2, characterized in that: In step 2, the LightGBM model includes setting the maximum depth of the tree to 10, the minimum number of samples of the leaf node to 2, not allowing zero gain splitting, the number of leaf nodes to 15, the learning rate to 0.05, the proportion of features used in each iteration to 80%, and setting the random number seed to 42.
7. The method for diagnosing multiple categories of blood diseases with the aid of multi-modal data and machine learning according to claim 1 or 2, characterized in that: In step 2, the neural network model comprises two layers of hidden layers (128, 64), an adaptive optimizer (Adam), and an early stopping mechanism to prevent overfitting; wherein the two layers of hidden layers have 128 and 64 neurons respectively, use a ReLU activation function, use an Adam optimizer, use an L2 regularization coefficient of 0.001, a batch size of 64, an initial learning rate of 0.001, and then an adaptive learning rate adjustment, an early stopping mechanism, a maximum number of iterations of 500, and a random number seed of 42.
8. The method for diagnosing multiple categories of blood diseases with the aid of multi-modal data and machine learning according to claim 1 or 2, characterized in that: In step 2, the voting ensemble model comprises a combination of a random forest model, an XGBoost model, a LightGBM model, and a neural network model, and uses a soft voting strategy with weights in the order of weights = [1, 1, 1, 0.8].
9. The method for diagnosing multiple categories of blood diseases with the aid of multi-modal data and machine learning according to claim 1 or 2, characterized in that: In step 2, the stacking ensemble model uses the aforementioned three tree models, i.e., a random forest model, an XGBoost model, and a LightGBM model, as base learners, uses an XGBoost model as a meta-learner, and uses five-fold stratified cross-validation as an internal validation mechanism.
10. The method for diagnosing multiple categories of blood diseases with the aid of multi-modal data and machine learning according to claim 1 or 2, characterized in that: In step 3, the establishment of the visual blood analysis interface comprises the establishment of a simple PYQT blood analysis interface.