Systems and methods for classifying blood cells
A machine learning-based digital staining and classification system using GANs addresses the inefficiencies of traditional chemical staining by improving accuracy and safety in blood cell classification.
Patent Information
- Application Number
- JP2025528633
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-09
- Filing Date
- 2023-11-17
- Publication Date
- 2026-01-06
AI Technical Summary
Traditional chemical staining methods for blood cell classification are time-consuming, costly, and pose health risks, necessitating the development of safer and more efficient alternatives.
A machine learning-based approach using a trained model for digital staining and classification, employing a generative adversarial network (GAN) to extract intermediate features for classification without chemical stains.
This method improves classification accuracy and reduces laboratory costs and risks by utilizing digital staining and classification systems, enhancing feature extraction from unstained images.
Smart Images

Figure 2026500096000001_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to the classification of biological samples, and more particularly to the classification of blood cells. [Background technology]
[0002] Microscopic analysis of blood cells is a fundamental tool in clinical diagnosis and medical research. Chemical staining of blood cells increases the contrast between cells and their respective subcomponents, facilitating the identification and differentiation of various cell types and morphologies. For example, structures necessary for the classification of pathological red blood cells or eosinophil cytoplasmic granules may be invisible in unstained images but clearly identifiable in chemically stained images.
[0003] While effective, traditional chemical staining methods often require multi-step procedures that are time-consuming and prone to error, and employ chemical reagents that can be expensive, pose health risks to laboratory staff, and require proper disposal. Summary of the Invention [Problem to be solved by the invention]
[0004] Therefore, there is a need for improved methods and systems for classifying blood cells without the need for the use of chemical stains. [Means for solving the problem]
[0005] In some embodiments, a method for classifying elements of a blood sample is provided, the method including: digitally staining an image of the blood sample using a trained machine learning model to generate a digitally stained image; extracting one or more intermediate features generated by the trained machine learning model during the digital staining of the image; providing the one or more extracted intermediate features to a trained multi-class classifier; and employing the trained multi-class classifier to classify at least one element in the blood sample based on the one or more extracted intermediate features.
[0006] In some embodiments, a machine learning-based digital staining and classification system is provided, comprising: a processor; a memory coupled to the processor, the memory including a trained machine learning model and a multi-class classifier coupled to the trained machine learning model; and computer program instructions stored in the memory that, when executed by the processor, cause the processor to: digitally stain an image of a blood sample using the trained machine learning model to generate a digitally stained image; extract one or more intermediate features generated by the trained machine learning model during digital staining of the image; provide the one or more extracted intermediate features to a trained multi-class classifier; and employ the trained multi-class classifier to classify at least one entity in the blood sample based on the one or more extracted intermediate features.
[0007] Other features and aspects of the present invention will become more apparent from the following detailed description, the appended claims, and the accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1]Figure 1A is an exemplary flow diagram of a method for classifying elements of a blood sample according to embodiments described herein, and Figure 1B is a diagram illustrating an exemplary computer on which the method of Figure 1A may be implemented, according to one or more embodiments. [Figure 2] 2A and 2B illustrate an exemplary embodiment of a first digital staining and classification system having a GAN generator, a GAN classifier, and a multi-class classifier, according to embodiments described herein. [Figure 3-1] 3A and 3B show categorized digital staining criteria for white blood cells for four models, according to embodiments described herein. [Figure 3-2] Figure 3C illustrates categorical digital staining criteria for white blood cells for four models, and Figure 3D illustrates the accuracy of white blood cell image quality results between a stained model, an unstained model, and three leading models, according to embodiments described herein. [Figure 4] 1 is a flowchart of a method for classifying elements of a blood sample according to one or more embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0009] Regardless of grammatical usage, the term includes individuals of male, female, or other gender identities.
[0010] Embodiments described herein include methods and systems for digitally staining and classifying one or more aspects of a blood sample, such as red blood cell size, shape, hemoglobin distribution, inclusion bodies, etc., and white blood cell size, shape, granules, morphology, type, etc.
[0011] In some embodiments, a trained machine learning (ML) model is employed to digitally stain an image of a blood sample. One or more intermediate features (e.g., one or more feature vectors) generated by the trained ML model (during digital staining of the image) are extracted and provided to a trained multi-class classifier. The trained multi-class classifier can then classify one or more elements in the blood sample based on the one or more extracted intermediate features. As used herein, "intermediate feature" refers to a feature, such as a feature vector, generated in an intermediate layer of an ML model, such as an intermediate layer of an ML model used during the generation of the digital stained image. The intermediate feature is typically not visible in the final digital stained image output from the ML model.
[0012] In some embodiments, the one or more intermediate features extracted from the trained ML model may include a feature vector extracted from the final bottleneck layer of the generator (e.g., a layer of a GAN generator between the encoder and decoder layers) or the decoder layer of the generator. Further, in one or more embodiments, the feature vector may be added to the feature vector of a multi-class classifier. Other intermediate features may also be employed. As noted above, the one or more extracted intermediate features may not be directly visible in the digitally stained image (although the intermediate features may affect the digitally stained image output by the generator).
[0013] In one or more embodiments, the trained ML model employed for digital dyeing can include a generative adversarial network (GAN) generator. In this way, multi-output learning is combined with GAN-based style transfer. Other machine learning models can also be employed.
[0014] Exemplary blood sample images employed include bright-field and / or dark-field microscopy captured under reflected and / or transmitted illumination conditions, computerized microscopy generated from a series of images in a Fourier imaging mode such as Fourier ptychography images, blood sample smears measured by differential interference contrast microscopy, interference reflection microscopy, phase contrast microscopy, Hoffman modulation contrast microscopy, or similar processes. Other image types can also be employed.
[0015] By using a trained ML model (for digital staining) and a multi-class classifier, features that are not readily visible but are present in the original image can be employed to classify elements of blood samples without the use of chemical staining. That is, features that enable correct classification are known to be present in the unstained image. For example, in the two datasets analyzed, all features required for classification were present in the unstained image of a healthy blood sample. Furthermore, in some embodiments, the use of extracted intermediate features (from the GAN generator) by the multi-class classifier improved classification performance, and the class information improved the quality of the generated images in terms of accuracy, F1 score, mean square error (MSE), and structural similarity index measure (SSIM).
[0016] The above and other embodiments described herein will be described below with reference to FIGS. 1A to 4. FIG.
[0017] MES is described, for example, in Wang, Z., Bovik, A.C., "Mean Squared Error: Love it or Leave it? A New Look at Signal Fidelity Measures," IEEE Signal Processing Magazine, 26(1), 98-117 (2009). SSIM is described, for example, in Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P., "Image Quality Assessment: From Error Visibility to Structural Similarity," IEEE Transactions on Image Processing, 13(4), 600-612 (2004). LPIPS is described, for example, in Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O., "The Unreasonable Effectiveness of Deep Features as a Perceptual Metric," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 586-595 (2018).
[0018] FIG. 1A illustrates an exemplary flow diagram 100 of a method for classifying elements of a blood sample, according to embodiments described herein. As described in more detail below with reference to FIG. 1A, the method of flow diagram 100 includes obtaining a training dataset 102 of blood sample images (e.g., unstained peripheral blood smear images and chemically stained peripheral blood smear images 104a-104n (e.g., differential interference contrast images or another image type)). The training dataset 102 is then employed to train a machine learning (ML) model 106 to digitally stain the unstained images of the blood sample. For example, in some embodiments, the ML model 106 may include a generative adversarial network (GAN) including a generator 108 and a classifier 110 trained on both the unstained and chemically stained images, as described in more detail below. Once trained, the generator 108 may be employed to generate a digitally stained image from the original unstained image.
[0019] According to flow diagram 100, multi-class classifier 112 can employ one or more intermediate features (e.g., feature vectors from generator 108) from the digital stain pipeline of digital stain model 106 during classification of elements of a blood sample (e.g., characteristics of red blood cells or white blood cells). As described in more detail below, in some embodiments, the use of these intermediate features can improve classification by multi-class classifier 112. Classification by multi-class classifier 112 (e.g., during training) can also improve the digital staining performed by digital stain model 106.
[0020] FIG. 1B illustrates an exemplary computer 120 upon which the method of FIG. 1A may be implemented, according to one or more embodiments. Referring to FIG. 1B, the computer 120 includes a processor 122 coupled to a memory 124. The memory 124 may include a training dataset 102 (including images 104a-104n), a digital stain ML model 106, and a multi-class classifier 112. The memory 124 may also include one or more programs 126 that, when executed by the processor 122, perform methods described herein, such as training the digital stain ML model 106 and the multi-class classifier 112 based on the training dataset 102, employing the trained digital stain ML model 106 to generate a digital stained image from an unstained image, and employing the multi-class classifier 112 to classify the image using one or more intermediate features from the digital stain ML model 106. The memory 124 may include multiple memory units and / or memory types. In some embodiments, all or a portion of the memory 124 may be external and / or remote to the computer 120. Additionally, in some embodiments, multiple processors may be employed.
[0021] As mentioned above, blood analysis is one of the most common prerequisites for the diagnostic process. A standard hematology test includes a complete blood count, which measures the number of red blood cells (RBCs), white blood cells (WBCs), platelets, hemoglobin concentration, and hematocrit. A standard hematology test outputs parameters such as the average size and hemoglobin concentration of each red blood cell. Additionally, white blood cells are differentiated into subcategories to aid the physician in understanding the sample.
[0022] Several characteristics of a blood sample can be used to determine the class of blood cells (e.g., healthy vs. diseased) including size, shape, hemoglobin distribution, inclusions, granules, morphology, etc. Such characteristics are typically determined by capturing and examining images of a peripheral blood smear under a microscope.
[0023] Not all features necessary for classification of pathological red blood cells are present in unstained images. For this reason, blood samples are chemically stained. For example, dots representing basophilic spots may be invisible in unstained images but clearly visible in chemically stained images. On the other hand, in healthy samples, the nuclei and granules in the cytoplasm of eosinophils may be visible in chemically stained samples but not recognizable in unstained pathological myelocytes. To recognize these features in stained images, hematologists usually need to use time-consuming and costly chemical stains before making a diagnosis.
[0024] According to embodiments described herein, sample staining is treated as an image-to-image translation task where deep learning enables the replacement of chemical stains with digital stains. This approach significantly reduces laboratory costs associated with chemical dyes, time, and waste management. Neural networks designed for artificial staining are combined with classifiers to enable complete blood counts and diseased cell labeling in some embodiments.
[0025] As described in more detail below, the use of auxiliary classifiers allows for the enhancement of classifier performance with additional feature information from the digital staining pipeline. In some embodiments, paired datasets are provided so that accurate representations of blood cells in both domains are available. Because such datasets are rarely available, there has been limited effort in the supervised setting of image-to-image transformation, much less in combination with auxiliary classification tasks. The embodiments described herein demonstrate the relationship between hematological image classification and supervised image-to-image transformation.
[0026] In this disclosure, we focus on two issues: are all the features required for classification present in the unstained image? Does classification affect the digital staining and vice versa?
[0027] As previously mentioned, classifying unstained blood cells is a complex task, as many features are only visible in stained cells. In some embodiments described herein, extracting features from the digital staining pipeline allows the classifier to have more information about each cell, thereby improving model performance. Furthermore, in some embodiments, incorporating the classifier into the digital staining pipeline allows the ML model used for digital staining and classification to achieve high classification standards and generate accurate digital stained images from unstained images.
[0028] By focusing on improving the classification of unstained blood cells, the embodiments described herein introduce a supervised image-to-image translation network, where features from an ML model (e.g., a GAN generator) provide auxiliary information to the classifier so that it can predict the correct class of unstained blood cells.
[0029] In some embodiments, the novel image transformation techniques described herein are based on a generative adversarial network (GAN) that includes two neural networks: a generator responsible for transforming random vectors into a chosen distribution, and a classifier trained to distinguish between generated samples from the same distribution and real samples (e.g., images). In a conditional GAN, the generated output is conditional on a class value. For example, the ML model described in Isola, P., Zhu, JY, Zhou, T., Efros, AA, "Image-to-Image Translation with Conditional Adversarial Networks," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1125-1134 (2017) (hereinafter "Isola et al.") and an improved variant described in Wang, TC, Liu, MY, Zhu, JY, Tao, A., Kautz, J., Catanzaro, B., "High Resolution Image Synthesis and Semantic Manipulation with Conditional Gans," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8798-8807 (2018) (hereinafter "Wang et al.") are based on a conditional GAN architecture and can synthesize high-resolution images from semantic label maps.
[0030] According to embodiments described herein, for digital staining tasks, class labels are fed into the digital staining pipeline so that the class information influences the generated image. In some embodiments, the label information is not known a priori but is predicted after (or in parallel with) image generation. As described below, the class information can be used to improve the artificial staining of an image (e.g., by a generator in a conditional GAN) even if it is not initially known (e.g., during testing).
[0031] In some embodiments, a GAN is provided having a generator and a classifier, where the classifier is incorporated into the generator but not the classifier. For example, Figures 2A and 2B show exemplary embodiments of a first digital staining and classification system 200a (Figure 2A) and a second digital staining and classification system 200b (Figure 2B), respectively, having a GAN generator 204, a GAN classifier 206, and a multi-class classifier 208, according to embodiments described herein. Information generated by generator 204 during digital staining of an image is provided to multi-class classifier 208, as described below.
[0032] 2A and 2B , in some embodiments, the generator 204 may include convolution, instance normalization, and activation function (e.g., ReLU) layers 210a, 210b, 210c, and 210d (e.g., each of layers 210a-210d may include a convolution layer, an instance normalization layer, and an activation function layer), and residual blocks 212a-212n. Similarly, the classifier 206 may include convolution, instance normalization, and activation function layers 210e and 210f, and residual blocks 214a-214n, etc. Additionally, one or more flattening layers (indicated by layers 216a and 216b) and fully connected layers may be employed to predict whether an image input to the classifier 206 is genuine or fake (R / F). Other numbers, types, and / or arrangements of layers may be employed in the generator 204 and / or classifier 206.
[0033] In one or more embodiments, classifier 208 may include an input layer 218a and various hidden layers 218b-218n, and may employ one or more flattening layers (illustrated by layers 220a and 220b) and fully connected layers to provide class indications and / or class probabilities. Other numbers, types, and / or arrangements of layers employed in classifier 208 are also possible.
[0034] Additionally, any suitable GAN, GAN generator, GAN discriminator, and / or classifier architecture may be employed. In one or more embodiments, the generator 204 may include one or more of a convolutional layer, a transposed convolutional layer, a batch normalization layer, an activation function layer (e.g., a ReLU or leaky ReLU activation function), a fully connected layer, and / or similar layers. In some embodiments, the discriminator 206 is formed by a neural network including one or more of a convolutional layer, a batch normalization layer, an activation function layer (e.g., a ReLU or leaky ReLU activation function), a fully connected layer, and / or similar layers.
[0035] In some embodiments, the multi-class classifier is formed by a neural network having an input layer, a hidden layer, and an output layer. Exemplary neural networks include convolutional neural networks, recurrent neural networks, etc. The input layer receives input from the generator (e.g., one or more intermediate features, such as one or more feature vectors as described above) and passes it to the hidden layer. The hidden layer can employ an activation function, such as a sigmoid or ReLU function, to convert the classifier input into classifier output classes. The output layer can then output class information (e.g., in some embodiments, a softmax layer can be used to output a probability for each class).
[0036] In some embodiments, ResNet (see He, K., X. Zhang, S. Ren, and J. Sun (2016), "Deep Residual Learning for Image Recognition," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770-778), EfficientNet (see Tan, M. and Q. Le (2019), "Efficientnet: Rethinking Model Scaling for Convolutional Neural Networks," International Conference on Machine Learning, PMLR, pp. 6105-6114), Visual Geometry Group (VGG) (see Simonyan, K. and A. Zisserman (2014), "Very Deep Convolutional Networks for Large-Scale Image Recognition," arXiv preprint One or more known network architectures can be modified for use in classifier 208, such as ConvNext (see Liu, Z., H. Mao, C.-Y. Wu, C. Feichtenhofe, T. Darrell, and S. Xie (2022), “A ConvNet for the 2020s,” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR)) or other known network architectures.
[0037] In a first embodiment shown in FIG. 2A , the input of the classifier 208 is the first intermediate feature (e.g., a feature vector, such as a one-dimensional or multidimensional vector) in the digital staining pipeline. In some embodiments, the first intermediate feature is from the beginning of the first decoder of the generator 204. For example, in some embodiments, the first intermediate feature is from the last residual block 212 n (e.g., the final layer of the bottleneck formed by residual blocks 212 a-212 n). Other intermediate features can also be used. In this manner, the generator 204 can assist the classifier 208 by extracting relevant class information stored in the bottleneck layer of the generator 204. Furthermore, backpropagation of gradients from the classifier 208 can help improve the image generation performed by the generator 204 (as described in more detail below).
[0038] In a second embodiment shown in FIG. 2B, the classifier 208 can employ both the first intermediate features and the second additional intermediate features (e.g., feature vector). In some embodiments, the second intermediate features are from one or more upsampling layers of the generator 204. For example, in some embodiments, the first layer of the decoder of the generator 204 (e.g., layer 210c in FIG. 1B) can provide the second intermediate feature information to the classifier 208.
[0039] In one or more embodiments, the first and / or second intermediate features can be concatenated with hidden vectors of the classifier 208 that have the same height and width values (e.g., the intermediate feature values can be stacked together in a layer of the classifier, such as layer 218b in FIG. 2B). In this manner, the classifier 208 can employ features from within the digital staining pipeline of the generator 204 that are relevant to the image-to-image transformation and that may contain more information than the unstained image. These intermediate features are typically not directly visible and / or accessible by the user (but do affect the final digital stained image).
[0040] As mentioned above, in the embodiment of FIG. 2B, the first intermediate feature is fed from the last residual block 212n of the bottleneck of the generator 204 to the input layer 218a of the classifier 208, and the second intermediate feature is fed from the first layer (e.g., layer 210c) of the decoder of the generator 204 to the hidden layer 218b of the classifier 208 (e.g., a convolutional layer followed by an activation function, etc.).
[0041] The embodiment of FIG. 2A is referred to herein as "CM_1F" (current model 1 with one intermediate feature), and the embodiment of FIG. 2B is referred to herein as "CM_2F" (current model 2 with two intermediate features). In both embodiments, the ML models employed for digital staining (e.g., generator 204 and classifier 206) are trained by employing any suitable loss term (e.g., that minimizes the loss of the generator and classifier). For example, in some embodiments, the loss term employed in Wang et al. may be employed.
[0042] According to one or more embodiments described herein, improved generator losses L G is written as follows: L G =L G_GAN +L G_GAN_FEAT +L G_VGG +L CE where L G_GAN is the probability that the generated image is genuine (e.g., the fake passivity loss), and L G_GAN_FEAT represents the GAN feature matching loss, and L G_VGG is V GG represents the feature matching loss, L CE represents an additional classification loss term (e.g., a cross-entropy (CE) loss function that depends on the blood element being classified, as described below).
[0043] In one embodiment, white blood cells (WBCs) are labeled with only one label from 7 or 14 classes, so the classification loss L CE is the cross entropy loss (L CE_WBC) for red blood cells (RBCs). For RBCs, the classifier implemented for the RBC dataset is expected to be more complex, as the corresponding labels of the cells belong to four categories (e.g., size, shape, hemoglobin distribution, and inclusions). For the first three categories, separate cross-entropy losses are applied, and finally, a binary cross-entropy loss is applied to the inclusion category. These losses are then jointly combined to result in a single classification loss term (L CE_RBC =L CE_SIZE +L CE_SHAPE +L CE_HEMO +LB CE_INCL ). In some embodiments, no regularizer is implemented because the use of the loss terms described above stabilizes the training of the network. In other embodiments, one or more regularization terms may be employed.
[0044] In some embodiments, a set of neural network classifiers (e.g., multiple multi-class classifiers 208) are trained and their performance in various configurations is analyzed to determine whether features necessary for successful classification are present in unstained images. Additionally, a multi-task model is employed to determine how image generation affects classification, and vice versa. Given a dataset containing both unstained and chemically stained images of blood cells with corresponding classes, the goal of this disclosure is to classify the unstained images. In some embodiments, to supplement classical classifier architectures, a neural network is provided that combines a generator and a classifier that share features and are jointly optimized. As described in more detail below, these embodiments demonstrate that the classifier module can be extended to output multiple characteristics of a given cell.
[0045] As mentioned above, in some embodiments, the GAN generator can include a combination of convolution, instance normalization, and ReLU layers, and residual blocks, which can be used to perform style transfer (e.g., creating a digitally stained image from an unstained image).
[0046] In some embodiments, architectures similar to the generator and / or discriminator described in Wang et al. may be employed. In one or more embodiments, the classifier may be based on EfficientNet and / or ConvNext architectures.
[0047] To evaluate the embodiments described herein, a set of experiments was performed on two different datasets. Deep learning models were trained using Python 3.7.13, PyTorch 1.12.0, and Cudatoolkit 11.3.1 on a Titan X GPU with 12GB RAM, available from Nvidia Corporation of Santa Clara, CA. Other software packages and / or computer components may also be employed.
[0048] A personal white blood cell and red blood cell dataset containing images acquired by two different scanners from 175 and 43 patients, respectively, was prepared and annotated by two hematologists. Tables 1(A)–1(E) show the label distribution in the white blood cell and red blood cell (categorical) datasets. We used train_test_split to train the image-to-image translation network. An additional separate validation set was used during training of the classification network. While red blood cell images had a size of 120 × 120 pixels, white blood cells had a size of 360 × 360 pixels. Prior to input to the network, they were resized to 120 × 120 pixels by cropping to 240 × 240 pixels.
[0049] To classify leukocytes, two datasets were created from the dataset introduced above: first, all 14 classes were included (see both Table 1(A) and Table 1(B)); second, only the five normal leukocyte types (basophils (BA), eosinophils (EO), segmented neutrophil granulocytes (SNE), monocytes (MO), lymphocytes (LY)) as well as artifacts (ART) and smudges (SMU) were used (see Table 1(A)).
[0050] For red blood cell classification, labels corresponding to four categories were available: inclusion body, shape, size, and hemoglobin distribution, as shown in Table 1(C), Table 1(D), Table 1(E), and Table 1(F), respectively. We also observed how categories influence each other, since the labels of categories are dependent and auxiliary information from different categories can improve the performance of the classifier.
[0051] [Table 1]
[0052] [Table 2]
[0053] [Table 3]
[0054] [Table 4]
[0055] [Table 5]
[0056] [Table 6]
[0057] We compared classifiers with different numbers of heads: First, a one-head (multi-output) classifier was trained separately for each category. Second, a four-head classifier was trained that included all categories. We expected the four-head classifier to be superior to the one-head classifier because in this case, the classifier looked at the whole cell and could recognize possible class pairs in various categories. In the inclusion body category, the labels were highly unbalanced (see Table 1(C)). For this reason, embodiments described herein can include the use of two labels for the classification task: normal and pathological, which facilitates classification. We did not use balancing techniques during the experiments. We also did not apply data augmentation, such as rotation or inversion, because these augmented images would not correspond to the actual data in the test set. We also did not apply oversampling of underrepresented classes because oversampling rare cell types in one category could cause a large imbalance in the other category.
[0058] In the first series of experiments, we observed how staining affected classifier performance. Chemical staining is used worldwide as a means to arrive at a diagnosis. Because more information is available in stained images, we expect classifiers trained and evaluated on stained cells to perform better. In the present disclosure, we used three different baselines with pre-trained weights for the classification models: EfficientNet (69 million parameters), VGG19 (136 million parameters), and ConvNeXt (201 million parameters). Therefore, the model size is the same order of magnitude as the digital staining pipeline. To prevent overfitting of the models, we trained the white blood cell classifier for 45 epochs, and the one-head and four-head red blood cell classifiers for 35 and 75 epochs, respectively. All digital staining models were trained for 100 epochs.
[0059] To examine the performance of the ML models described herein, four ML models (CM_1F, CM_2F, PM_1, and PM_2) were compared: (1) CM_1F: ML model in which one intermediate feature is fed to classifier 208 (FIG. 2A) (current model with one feature, i.e., “CM_1F”); (2) CM_2F: an ML model in which the classifier 208 (FIG. 2B) is fed with two intermediate features from the generator 204 (current model with two intermediate features, i.e., “CM_2F”); (3) PM_1: the ML model described in Wang et al. (referred to as Prior Model 1, i.e., “PM_1”); (4) PM_2: Odena, A., Olah, C., Shlens, J., "Conditional Image Synthesis with Auxiliary Classifier Gans," International Conference on Machine Learning, pp. 2642-2651. PMLR (2017) ML model (referred to as "Previous Model 2" or "PM_2")
[0060] In Table 2(A) and Table 2(B), we select the benchmarks for the best models (PM_2, CM_1F, and CM_2F) and compare the accuracy and F1 scores of the multiclass classifiers on stained and unstained images using the "macro" setting (which calculates the F1 score for each label and outputs the unweighted average). The results are consistent with expectations, with a gap between stained and unstained images being observed for all models.
[0061] [Table 7]
[0062] [Table 8]
[0063] 3A-3C show categorical digital staining criteria for white blood cells for four models (PM_1, PM_2, CM_1F, and CM_2F) according to embodiments described herein. Artifacts and smudges are treated separately from blood cells due to their variable appearance. FIG. 3D shows the accuracy of white blood cell image quality results between stained, unstained, and three leading models (PM_2, CM_1F, and CM_2F) according to embodiments described herein. In general, all models were observed to perform well on normal cells.
[0064] In Figure 3D, we observe that the difference in classification accuracy between stained and unstained healthy white blood cells is 1% and 20% for diseased cells. For the RBC model, the separation of healthy and diseased cells does not provide relevant information. This is because there is only one healthy label for each category, and if a cell represents a healthy class in one category, other labels corresponding to other categories may be diseased. Overall, the criteria for the model trained on the white blood cell dataset are more accurate due to the dataset's stability compared to the unbalanced and smaller dataset for red blood cells. In some embodiments, size and shape are expected to be the winning categories because all the necessary features for these two categories are present in the unstained image. Because the shape category has 11 classes, and some classes have similar appearances, such as oval and ellipsoidal red blood cells, the classifier does not correctly classify the underrepresented diseased classes.
[0065] As described below with reference to Tables 3(A)-3(F), an embodiment of the model described herein was compared to two prior models. As noted above, the first prior model is referred to as PM_1, and the second prior model is referred to as PM_2. In the second prior model (PM_2), an auxiliary classifier was incorporated into the classifier of the first prior model (PM_1). Additionally, the first model of the present disclosure is referred to as "CM_1F" (the current model with one intermediate feature), while the second model of the present disclosure is referred to as "CM_2F" (the current model with two intermediate features).
[0066] By comparing the accuracy and F1 score of the classifiers for the three multi-task learning models, the digital staining pipeline is expected to outperform the classifier trained on unstained images. This can be explained by the fact that the digital staining pipeline learns the features required for staining, and these features play an important role in classification.
[0067] Improved classification performance is observed for the white blood cell dataset (see Tables 2(A) and 2(B)). Here, the disclosed model outperforms classifiers trained on unstained images. The more significant increase in F1 score indicates that the features from the generator can capture relevant characteristics of pathological cells. Based on detailed evaluation of white blood cells (Figures 3A-3D), it is observed that the digital staining of pathological cells has poorer image quality than normal cells, indicating that important features of pathological cells are missing. Meanwhile, both of these models are expected to perform worse than classifiers trained on stained images.
[0068] In the red blood cell dataset (see Table 2(A) and Table 2(B)), the proposed model of the present disclosure with multi-class classification in size and hemoglobin distribution categories outperforms classifiers trained on unstained images, but the shape classifier underperforms its similar classes. In a comparison of the baselines of the one-head and four-head models (Table 2(A) vs. Table 2(B)), only small variations are observed, indicating that the classifiers were unable to learn dependencies between categories. This may be due to a lack of data from the pathological class.
[0069] These results demonstrate that, using the embodiments described herein, features in unstained images are sufficient for classifying normal cells (e.g., by using intermediate features extracted from a digital staining pipeline), but lack detail for classifying pathological cells. For the dataset of the present disclosure, the CM_1F feature vector extracted from generator 204 of FIG. 2A (e.g., a single intermediate feature embodiment) performed best. Furthermore, the placement of the classifier in the network is also relevant (e.g., integrating the classifier with the generator rather than the discriminator).
[0070] The introduced models were also evaluated in terms of image quality as a comparison against chemically stained ground truth images. In embodiments of the present disclosure, image quality is measured by MSE (Mean Square Error), SSIM (Structured Similarity Indexing Method), and LPIPS (Learned Perceptual Image Patch Similarity). Improved image quality is expected, as the classifier can assist the generator in exploring similarities within each class. As mentioned above, the image quality metrics of the disclosed multi-task learning models (CM_1F and CM_2F) are compared with the prior models PM_1 and PM_2. These image quality metrics are shown in Tables 3(A) through 3(F).
[0071] For white blood cells, the quality of the generated images is of the same order of magnitude, with only slight variations observed. Only a slight improvement is observed for the first of the two models in the present disclosure. This result is reliable because the ground truth chemically stained images are often blurred and may only show slight differences in color tone. Therefore, it is beneficial that the image quality criteria do not change significantly, despite the fact that the digitally stained images in the present disclosure are clearer.
[0072] Overall, the embodiment of the present disclosure that employs two intermediate features from the generator (CM_2F as shown in FIG. 2B) produces the least accurate generated images because this model has the most complex classifier, robbing the generator of computational power. Therefore, the generator fails to recognize important features from the unstained image and one intermediate feature model (CM_1F as shown in FIG. 2A).
[0073] Based on these results, the multi-task learning model according to the embodiments described herein can improve the quality of classification results and generated images.
[0074] [Table 9]
[0075] [Table 10]
[0076] [Table 11]
[0077] [Table 12]
[0078] [Table 13]
[0079] [Table 14]
[0080] Based on the class-specific evaluation, in some embodiments, the networks described herein perform accurately on normal cells. While the features required for classification are scarce for visual inspection on unstained images, embodiments of the networks described herein, such as those shown in Figures 2A and 2B, are able to recover auxiliary information from stained images, improving classification performance. Based on this, in some embodiments, the methods described herein can be used in clinical contexts, such as performing complete blood counts. Different image modalities that separately capture details about unstained cells may enable the use of the ML networks described herein in more complex clinical contexts, such as determining a patient's condition. For example, other methods may enable the capture and extraction of features from unstained images, with the goal of determining a patient's diagnosis with similar accuracy as the use of chemically stained images by the networks described herein.
[0081] As described above, in one or more embodiments, an ML model that simultaneously classifies and digitally stains red blood cell and white blood cell datasets is shown to achieve higher classification performance on unstained images. The ML model generates high-quality and accurate digitally stained images of normal cells for both datasets. The ML model outperforms other classifiers and can reduce the classification performance gap between stained and unstained images.
[0082] The embodiments described herein demonstrate how multi-task learning affects classification and digital staining performance. Furthermore, the embedded placement of the classifier can play a significant role in performance. In some embodiments, the most efficient framework is the disclosed model, where the classifier input is a feature vector from the bottleneck end of the GAN generator (e.g., FIG. 2A). In some embodiments, it is further observed that an image-to-image transformation pipeline improves classification criteria while improving the quality of the generated images. In one or more embodiments, employing the ML models described herein can eliminate the need for chemical staining in healthy samples.
[0083] 4 is a flowchart of a method 400 for classifying elements of a blood sample, according to one or more embodiments. Referring to FIG. 4, at block 402, the method 400 includes generating a digitally stained image by digitally staining an image of the blood sample using a trained machine learning model. For example, in some embodiments, an unstained image is input to an ML model (e.g., generator 204) that has been trained to digitally stain the image.
[0084] During the digital staining process, intermediate features are generated at various layers of the ML model (e.g., convolutional layers, residual blocks, activation function layers, etc.). At block 404, method 400 includes extracting one or more intermediate features generated by the trained machine learning model during digital staining of the image. In some embodiments, the one or more intermediate features are feature vectors extracted from a bottleneck layer (e.g., residual block) and / or a decoder layer of the GAN generator (e.g., generator 204 of FIG. 2A or FIG. 2B ).
[0085] At block 406, method 400 includes providing the one or more extracted intermediate features to a trained multi-class classifier. As shown in Figures 2A and 2B, in one or more embodiments, the intermediate features of the digital staining pipeline (e.g., classifier 208) are fed to one or more layers (e.g., input layer, deeper layers, and / or hidden layers) of classifier 208. Thereafter, at block 408, method 400 includes employing the trained multi-class classifier to classify at least one element in the blood sample based on the one or more extracted intermediate features.
[0086] All publications and patents cited in this disclosure, including those listed above, are herein incorporated by reference in their entirety for all purposes to the same extent as if each individual publication or patent was specifically and individually indicated to be incorporated by reference.
[0087] The foregoing description discloses merely exemplary embodiments of the present invention. Modifications of the above-disclosed apparatus and methods which fall within the scope of the present invention will be readily apparent to those skilled in the art.
[0088] Although the present invention has been disclosed in terms of exemplary embodiments thereof, it is to be understood that other embodiments may fall within the spirit and scope of the invention, as defined by the following claims.
[0089] Illustrative Embodiments The following provides a non-limiting list of exemplary embodiments of the present disclosure:
[0090] Exemplary Embodiment 1. A method for classifying elements of a blood sample, comprising: generating a digital stained image by digitally staining an image of the blood sample using the trained machine learning model; extracting one or more intermediate features generated by the trained machine learning model during digital staining of the image; providing one or more extracted intermediate features to a trained multi-class classifier; classifying at least one element in the blood sample based on the one or more extracted intermediate features by employing a trained multi-class classifier; A method comprising:
[0091] Exemplary Embodiment 2. The method of exemplary embodiment 1, wherein the image comprises a differential interference contrast image.
[0092] Exemplary Embodiment 3. The method of exemplary embodiment 1 or 2, wherein the trained machine learning model includes a trained generator of a generative adversarial network.
[0093] Exemplary Embodiment 4. The method of any one of Exemplary Embodiments 1-3, wherein the one or more intermediate features generated by the trained generator include a feature vector.
[0094] Exemplary Embodiment 5. The method according to any one of exemplary embodiments 1 to 4, wherein the feature vector is extracted from the final layer of the bottleneck of the generator.
[0095] Exemplary Embodiment 6. The method of any one of exemplary embodiments 1-5, further comprising adding the feature vector to a multi-class classifier.
[0096] Exemplary Embodiment 7. The method of any one of Exemplary Embodiments 1-6, wherein adding the feature vectors includes concatenating feature vectors from the trained generator with vectors of a multi-class classifier that have the same height and width values.
[0097] Exemplary Embodiment 8. The method of any one of exemplary embodiments 1-7, wherein one or more extracted intermediate features are not visible in the digital stained image.
[0098] Exemplary Embodiment 9. The method of any one of Exemplary Embodiments 1-8, wherein classifying at least one element in the blood sample based on one or more extracted intermediate features by employing a trained multi-class classifier includes classifying at least one of red blood cells and white blood cells.
[0099] Exemplary Embodiment 10. The method of any one of Exemplary Embodiments 1-9, wherein classifying at least one of red blood cells and white blood cells includes classifying at least one of size, shape, hemoglobin distribution, inclusion bodies, granules, and morphology.
[0100] Exemplary embodiment 11. A machine learning based digital staining and classification system comprising: processor and; a memory coupled to the processor, the memory including a trained machine learning (ML) model and a multi-class classifier coupled to the trained ML model; When stored in memory and executed by a processor: generating a digital stained image by digitally staining an image of the blood sample using the trained ML model; extracting one or more intermediate features generated by the trained machine learning model during digital staining of the image; providing one or more extracted intermediate features to a trained multi-class classifier; classifying at least one element in the blood sample based on the one or more extracted intermediate features by employing a trained multi-class classifier; and computer program instructions for causing a processor to perform the steps of: Including, the system.
[0101] Exemplary Embodiment 12. The system of any one of Exemplary Embodiments 1-11, wherein the image comprises a differential interference contrast image.
[0102] Exemplary Embodiment 13. The system of any one of Exemplary Embodiments 1-12, wherein the trained machine learning model includes a trained generator of a generative adversarial network.
[0103] Exemplary Embodiment 14. The system of any one of Exemplary Embodiments 1-13, wherein the one or more intermediate features generated by the trained generator include a feature vector.
[0104] Exemplary Embodiment 15. The system of any one of exemplary embodiments 1-14, wherein the feature vector is extracted from the final bottleneck layer of the trained generator.
[0105] Exemplary Embodiment 16. The system of any one of exemplary embodiments 1-15, further comprising computer program instructions stored in a memory and that, when executed by a processor, cause the processor to add the feature vector to a multi-class classifier.
[0106] Exemplary Embodiment 17. The system of any one of exemplary embodiments 1-16, further comprising computer program instructions that, when stored in a memory and executed by a processor, cause the processor to concatenate feature vectors from the trained generator with vectors of a multi-class classifier having the same height and width values.
[0107] Exemplary Embodiment 18. The system of any one of exemplary embodiments 1-17, wherein one or more extracted intermediate features are not visible in the digital stained image.
[0108] Exemplary Embodiment 19. The system of any one of exemplary embodiments 1-18, further comprising computer program instructions that, when stored in the memory and executed by the processor, cause the processor to classify at least one of red blood cells and white blood cells using a multi-class classifier.
[0109] Exemplary Embodiment 20. The system of any one of Exemplary Embodiments 1-19, further comprising computer program instructions that, when stored in a memory and executed by a processor, cause the processor to classify at least one of size, shape, hemoglobin distribution, inclusion bodies, granules, and morphology using a multi-class classifier. [Explanation of symbols]
[0110] 102 training datasets 104a~104n Statue 106 Machine Learning (ML) Model, Digital Dyeing ML Model 108 Generator 110 Classifier 112 Multi-class Classifier 120 Computer 122 processors 124 memory 126 Programs 200a The first digital staining and classification system 200b Second Digital Staining and Classification System 204 GAN generator 206 GAN discriminator 208 Multi-class Classifier 210a, 210b, 210c, 210d Convolution, instance normalization, and activation function layers 212a~212n Residual Blocks 210e, 210f Convolution, instance normalization, and activation function layers 214a~214n Residual Block 216a Planarization layer 216b Fully connected layer 218a Input layer 218b~218n Hidden layer 220a Planarization layer 220b fully connected layer
Claims
1. 1. A method for classifying elements of a blood sample, comprising: generating a digital stained image by digitally staining an image of the blood sample using the trained machine learning model; extracting one or more intermediate features generated by the trained machine learning model during digital staining of the image; providing one or more extracted intermediate features to a trained multi-class classifier; classifying at least one entity in the blood sample based on one or more extracted intermediate features by employing a trained multi-class classifier; The method comprising:
2. The method of claim 1 , wherein the image comprises a differential interference contrast image.
3. The method of claim 1 , wherein the trained machine learning model comprises a trained generator of a generative adversarial network.
4. The method of claim 3 , wherein the one or more intermediate features generated by the trained generator include a feature vector.
5. The method of claim 4 , wherein the feature vector is extracted from the final layer of the generator bottleneck.
6. The method of claim 4 , further comprising adding the feature vector to a multi-class classifier.
7. The method of claim 6 , wherein adding feature vectors comprises concatenating feature vectors from a trained generator with vectors of a multi-class classifier that have the same height and width values.
8. The method of claim 1 , wherein one or more extracted intermediate features are not visible in the digital stained image.
9. 10. The method of claim 1, wherein classifying at least one element in the blood sample based on the one or more extracted intermediate features by employing a trained multi-class classifier comprises classifying at least one of red blood cells and white blood cells.
10. Sorting at least one of red blood cells and white blood cells includes:
10. The method of claim 9, comprising classifying at least one of size, shape, hemoglobin distribution, inclusion bodies, granules, and morphology.
11. 1. A machine learning based digital staining and classification system comprising: a processor; a memory coupled to the processor, the memory including a trained machine learning (ML) model and a multi-class classifier coupled to the trained ML model; When stored in memory and executed by a processor: generating a digital stained image by digitally staining an image of the blood sample using the trained ML model; extracting one or more intermediate features generated by the trained machine learning model during digital staining of the image; providing one or more extracted intermediate features to a trained multi-class classifier; classifying at least one entity in the blood sample based on one or more extracted intermediate features by employing a trained multi-class classifier; computer program instructions to cause a processor to: The system comprising:
12. The system of claim 11 , wherein the image comprises a differential interference contrast image.
13. The system of claim 11 , wherein the trained machine learning model comprises a trained generator of a generative adversarial network.
14. The system of claim 13 , wherein the one or more intermediate features generated by the trained generator include a feature vector.
15. The system of claim 14 , wherein the feature vector is extracted from the final bottleneck layer of the trained generator.
16. 15. The system of claim 14, further comprising computer program instructions stored in the memory that, when executed by the processor, cause the processor to add the feature vector to a multi-class classifier.
17. 17. The system of claim 16, further comprising computer program instructions stored in the memory that, when executed by the processor, cause the processor to concatenate feature vectors from the trained generator with vectors of a multi-class classifier that have the same height and width values.
18. The system of claim 11 , wherein the one or more extracted intermediate features are not visible in the digital stained image.
19. 12. The system of claim 11, further comprising computer program instructions stored in the memory that, when executed by the processor, cause the processor to classify at least one of red blood cells and white blood cells using a multi-class classifier.
20. 20. The system of claim 19, further comprising computer program instructions stored in the memory that, when executed by the processor, cause the processor to classify at least one of size, shape, hemoglobin distribution, inclusion bodies, granules, and morphology using a multi-class classifier.
Citation Information
Cited By
Nut turning tool and fitting nut turning method
US12553553B2
Fitting nut, fitting, fluid pressure device, and fluid control system
US12565953B2