Machine learning enabled histological analysis
Through machine learning cell classification model, using confidence threshold processing annotated by expert, distinguishing macrophage types in cancer tissues, solving the problem of inaccurate identification in the prior art and improving the accuracy of treatment response prediction.
Patent Information
- Application Number
- CN202380088226.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-18
- Filing Date
- 2023-11-17
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art is difficult to accurately distinguish different types of macrophages in cancer tissues, especially non-pigmented stromal macrophages and fibroblasts, resulting in inaccurate prediction of treatment responses.
Using machine-learning cell classification model, we distinguish foamy macrophages, in-alveolar macrophages and pigmented stromal macrophages from stromal cells, including fibroblasts and non-pigmented stromal macrophages to improve recognition accuracy through the confidence threshold processing annotated by expert in the training data.
Improves the accuracy of identification of macrophage types in cancer tissues, enhances the predictive ability of combined immunotherapy responders, and provides more accurate prediction of treatment response.
Smart Images

Figure CN120418795A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 384,364, entitled "Machine Learning Enabled Histological Analysis," filed on November 18, 2022, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] The subject matter described herein generally relates to digital and computational pathology, and more particularly to machine - learning - based methods for performing cell type classification. Background Art
[0004] Cancer cells trigger significant molecular, cellular, and physical changes within their host tissue to support further growth and metastasis. Tumors are typically heterogeneous collections of cells that include infiltrating and resident host cells as well as secreted factors and the extracellular matrix. On a broader scale, the tumor microenvironment can include immune cells, endothelial cells, and fibroblasts present in the vicinity of cancer cells. There is a strong correlation between the composition of the tumor, the tumor microenvironment, and the ability of the tumor to maintain itself at its primary site, evade the immune response, resist drug intervention, and proliferate to various secondary locations. Summary of the Invention
[0005] Systems, methods, and articles of manufacture for enabling machine - learning - based histological analysis are provided, including computer program products. In one aspect, a system for identifying one or more cell types present in an image of a biological sample is provided. The system can include at least one processor and at least one memory. The at least one memory can include program code that, when executed by the at least one processor, provides operations. The operations can include: receiving an image of a biological sample; applying a cell classification model to identify one or more cell types present in the biological sample based at least on the image of the biological sample, the cell classification model being trained to distinguish between multiple cell types, the multiple cell types including a first cell type and a second cell type, the probability that the first cell type is a macrophage meeting a threshold, and the probability that the second cell type is a macrophage not meeting the threshold; and generating a compositional profile for the biological sample based at least on the one or more cell types identified in the biological sample.
[0006] In another aspect, a method for identifying one or more cell types present in an image of a biological sample is provided. The method may include: receiving an image of the biological sample; applying a cell classification model to identify one or more cell types present in the biological sample based at least on the image of the biological sample, the cell classification model being trained to distinguish between multiple cell types, the multiple cell types including a first cell type and a second cell type, the probability that the first cell type is a macrophage meeting a threshold, and the probability that the second cell type is a macrophage not meeting the threshold; and generating a composition profile for the biological sample based at least on the one or more cell types identified in the biological sample.
[0007] In another aspect, a computer program product for identifying one or more cell types present in an image of a biological sample is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations. The operations may include: receiving an image of the biological sample; applying a cell classification model to identify one or more cell types present in the biological sample based at least on the image of the biological sample, the cell classification model being trained to distinguish between multiple cell types, the multiple cell types including a first cell type and a second cell type, the probability that the first cell type is a macrophage meeting a threshold, and the probability that the second cell type is a macrophage not meeting the threshold; and generating a composition profile for the biological sample based at least on the one or more cell types identified in the biological sample.
[0008] Specific implementations of the present subject matter may include, but are not limited to, methods consistent with the description provided herein and articles of manufacture including tangible, machine-readable media operable to cause one or more machines (e.g., computers, etc.) to cause operations implementing one or more of the described features. Similarly, a computer system that may include one or more processors and one or more memories coupled to the one or more processors is also described. The memory, which may include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. A computer-implemented method consistent with one or more implementations of the present subject matter may be implemented by one or more data processors present in a single computing system or multiple computing systems. Such multiple computing systems may be connected and may exchange data and / or commands or other instructions, etc., via one or more connections, including, for example, via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.) or via a direct connection between one or more of the multiple computing systems, etc.
[0009] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will become apparent from the description, the drawings, and the claims. Although certain features of the presently disclosed subject matter are described for illustrative purposes in connection with the identification of a cell type associated with low-confidence expert annotations, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure are intended to define the scope of the protected subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings are incorporated into and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein, and together with the description, help explain some of the principles associated with the disclosed embodiments. In the drawings,
[0011] Figure 1 a system diagram depicting an example of a digital pathology system according to some exemplary embodiments is shown;
[0012] Figure 2A a flowchart depicting an example of a process for identifying one or more cell types present in an image of a biological sample according to some exemplary embodiments is shown;
[0013] Figure 2B a flowchart depicting another example of a process for identifying one or more cell types present in an image of a biological sample according to some exemplary embodiments is shown;
[0014] Figure 3A a flowchart depicting an example of a process for training a cell classification model according to some exemplary embodiments is shown;
[0015] Figure 3B a flowchart depicting an example of a process for generating a training set for training a cell classification model according to some exemplary embodiments is shown;
[0016] Figure 4 a schematic diagram depicting an example of a tumor microenvironment according to some exemplary embodiments is shown;
[0017] Figure 5 an example of a macrophage associated with high-confidence expert annotations according to some exemplary embodiments is shown; and
[0018] Figure 6 a block diagram depicting an example of a computing system according to some exemplary embodiments is shown.
[0019] When actually applied, like reference numerals represent like structures, features, or elements. DETAILED DESCRIPTION
[0020] In highly heterogeneous diseases such as cancer, insights into the cell types that form diseased tissue and the surrounding microenvironment can be essential for accurately diagnosing disease subtypes, prognosing disease progression, and predicting responses to various treatments. For example, patients with non-small cell lung cancer (NSCLC) who are likely to respond to combination immunotherapy, such as the combination of T-cell immunoreceptor with Ig and ITIM domains (TIGIT) and atezolizumab, can be identified based on their cancer cells expressing high levels of the programmed death ligand-1 (PD-L1) gene. Within the population of patients with non-small cell lung cancer whose cancer cells are PD-L1 negative or express low levels of the PD-L1 gene, it is desirable to further distinguish between responders and non-responders to combination immunotherapy.
[0021] Interactions between macrophages and lymphocytes (e.g., CD8-positive T cells), which trigger the release of anti-inflammatory cytokines (e.g., IL-6, etc.), can inhibit a patient's response to combination immunotherapy (such as a combination of a T-cell immunoreceptor (TIGIT) with Ig and ITIM domains and atezolizumab). Therefore, macrophages can be used as an indicator of treatment response. However, macrophage detection alone is not a sufficient indicator. It may be important to understand what types of macrophages are present along with the context of the tumor environment. The present disclosure provides systems and methods for distinguishing responders from non-responders based on the presence of tumor-associated macrophages in the tumor environment. In particular, the present disclosure describes a method for distinguishing low-confidence macrophages from high-confidence macrophages. In addition, the present disclosure presents systems and methods for identifying and distinguishing various macrophages, including stromal macrophages and alveolar macrophages. When macrophages are classified at this level of refinement, the disclosed systems and methods can provide predictions of treatment response or clinical outcomes. In some cases, refined macrophage type classification and detection can be used as a proxy or alternative to gene expression analysis in predicting treatment response.
[0022] In some exemplary embodiments, an image depicting a biological sample may undergo histological analysis to identify the one or more cell types present therein. For example, in some cases, a machine learning-based cell classification model may be applied to a whole slide image of a biological sample (e.g., a whole slide image stained with hematoxylin and eosin (H&E), a whole slide image stained with multiplex immunofluorescence (MxIF), a whole slide image stained with immunohistochemistry (IHC), etc.) to identify one or more tumor cells, lymphocytes, plasma cells, fibroblasts, macrophages, endothelial cells, adipocytes, and neutrophils present in the biological sample. A compositional profile for the biological sample may be generated based at least on the one or more cell types identified within the biological sample. At least one of a disease diagnosis, disease progression, and treatment of a patient associated with the biological sample may be determined based on the compositional profile. For example, in some cases, a patient associated with the biological sample may be identified as a responder (or non-responder) to a particular treatment (e.g., a combination of T cell immunoglobulin and immunoreceptor tyrosine-based inhibitory motif (TIGIT) and atezolizumab) based at least on whether macrophages are identified as being present within the biological sample.
[0023] In some exemplary embodiments, the cell classification model may be trained to distinguish various different cells, which include, for example, tumor cells, lymphocytes, plasma cells, fibroblasts, macrophages, endothelial cells, adipocytes, neutrophils, and the like. In some cases, the cell classification model may be trained based on a training set containing one or more images, each of the one or more images being associated with one or more ground truth labels identifying the cell types present therein. Additionally, in some cases, the ground truth labels associated with each image in the training set may be determined based on expert annotations. Thus, the performance of the cell classification model may be limited by the accuracy of the expert annotations. In cases where the expert annotations fail to accurately identify a particular cell type within the one or more images, the trained cell classification model may perform poorly when encountering that particular cell type. For example, certain types of macrophages (including non-pigmented stromal macrophages) may be difficult to visually decipher in an image. Non-pigmented stromal macrophages are often confused with other stromal cells such as fibroblasts, while foamy macrophages, alveolar macrophages, and pigmented stromal macrophages may be more readily distinguishable. Thus, the ground truth labels identifying non-pigmented stromal macrophages and fibroblasts tend to be less reliable. Therefore, training the cell classification model to identify a single monolithic class of macrophages may degrade the performance of the trained cell classification model in correctly identifying macrophages that may be present in a biological sample.
[0024] In some exemplary embodiments, a cell classification model can be trained to identify: macrophages of a first category, the macrophages of the first category including foamy macrophages, alveolar macrophages (including intra-alveolar macrophages), and pigmented stromal macrophages; and stromal cells of a second category, the stromal cells of the second category including non-pigmented stromal macrophages and fibroblasts. Training the cell classification model to identify separate categories of stromal cells (including non-pigmented stromal macrophages that may be confused with fibroblasts) can increase the performance of the cell classification model in identifying other macrophages, such as foamy macrophages, alveolar macrophages (e.g., intra-alveolar macrophages), pigmented stromal macrophages, and the like. For example, in some cases, the cell classification model can be trained to distinguish between multiple cell types, including a first cell type and a second cell type, where the probability that the first cell type is a macrophage meets a threshold, and the probability that the second cell type is a macrophage does not exceed the threshold. The first cell type (where the probability that the first cell type is a macrophage meets the threshold) can include stromal cells such as fibroblasts and non-pigmented stromal macrophages, while the second cell type (where the probability that the second cell type is a macrophage does not meet the threshold) can include foamy macrophages, intra-alveolar macrophages, and pigmented stromal macrophages. In some cases, in addition to the first cell type and the second cell type, the cell classification model can be further trained to distinguish tumor cells (e.g., non-small cell lung cancer (NSCLC) tumor cells, etc.), lymphocytes, plasma cells, endothelial cells, adipocytes, and neutrophils.
[0025] In some exemplary embodiments, a cell classification model can be trained based on training data that includes one or more images annotated with ground truth labels identifying at least one cell type present in each image. When training a cell classification model to identify stromal cells including separate classes of fibroblasts and non-pigmented stromal macrophages, the uncertainty present in the expert annotations associated with non-pigmented stromal macrophages can be captured. For example, in some cases, training data can be generated by at least assigning a first ground truth label identifying a first cell type (e.g., stromal cells) to the one or more images based on a first expert annotation identifying one or more macrophages being associated with a confidence value that fails to meet one or more thresholds. Additionally, in some cases, training data can be generated by at least assigning a second ground truth label identifying a second cell type (e.g., macrophages) to the one or more images based on a second expert annotation identifying the one or more macrophages being associated with a confidence value that meets the one or more thresholds. Thus, if the corresponding expert annotation is reliable enough (e.g., a confidence value that meets one or more thresholds), the ground truth label of the image can indicate the presence of macrophages. In cases where the expert annotation indicating the presence of macrophages is not reliable enough (e.g., a confidence value that fails to meet one or more thresholds), the ground truth label of the image can indicate the presence of stromal cells rather than the presence of macrophages.
[0026] In some exemplary embodiments, a cell classification model can classify each cell depicted in an image of a biological sample by at least assigning a "hard label" of a specific cell type or a "soft label" of the probability that the cell is positive for each of multiple different cell types (e.g., stromal cells, macrophages, tumor cells, lymphocytes, and plasma cells) to each cell. In some cases, the cell classification model can operate on a segmented image that has undergone segmentation (such as watershed cell segmentation, machine learning-based cell segmentation, etc.) to locate the individual cells present therein. Additionally, in some cases, the cell classification model can include a first machine learning model that is trained to identify one or more visible features within an image of a biological sample, and the one or more visible features can be identified, located, interpreted, inferred, and / or otherwise detected by, for example, visual inspection of the image by a person, machine, algorithm, etc. Cell classification can further include a second machine learning model that is trained to determine one or more cell types present in the image based at least on the one or more visible features extracted from the image of the biological sample. Alternatively or and / or additionally, the cell classification model can be implemented as an end-to-end model that determines the cell types present in the image based on one or more hidden features extracted from the image, and the one or more hidden features do not necessarily correspond to the aforementioned visible features.
[0027] Figure 1 FIG. shows a system diagram depicting an example of a digital pathology system 100 according to some exemplary embodiments. Referring to Figure 1 , the digital pathology system 100 may include a digital pathology platform 110, an imaging system 120, and a client device 130. As Figure 1 shown, the digital pathology platform 110, the imaging system 120, and the client device 130 may be communicatively coupled via a network 140. The network 140 may be a wired network and / or a wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc. The imaging system 120 may include one or more imaging devices (including, for example, a microscope, a digital camera, a whole slide scanner, an automated microscope, etc.). The client device 130 may be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc.
[0028] Referring again to Figure 1 , the digital pathology platform 110 may include a training engine 112, a cell classification engine 114, and a diagnosis and treatment engine 116. As Figure 1 shown, the cell classification engine 114 may include a cell classification model 115 that is trained to identify one or more cell types present in an image 117 depicting a biological sample. In some cases, the image 117 may be a stained whole slide image (WSI), including, for example, a hematoxylin and eosin (H&E) stained whole slide image, a multiplex immunofluorescence (MxIF) stained whole slide image, an immunohistochemistry (IHC) stained whole slide image, etc. Additionally, in some cases, the image 117 may have undergone segmentation (such as watershed cell segmentation, machine learning-based cell segmentation, etc.) to locate the individual cells present therein. The cell classification model 115 may be trained to identify: a macrophage class that includes foamy macrophages, alveolar macrophages, and pigmented stromal macrophages; and a separate stromal cell class that includes non-pigmented stromal macrophages and fibroblasts. Additionally, in some cases, the cell classification model 115 may be trained to identify tumor cells, plasma cells, lymphocytes, endothelial cells, adipocytes, and neutrophils.
[0029] In some exemplary embodiments, the training engine 112 may train the cell classification model 115 to distinguish multiple cell types based at least on the training set 113, the multiple cell types including, for example, macrophages, stromal cells, tumor cells, plasma cells, lymphocytes, endothelial cells, adipocytes, and neutrophils. In some cases, the training engine 112 may generate the training set 113 including one or more images, and the ground truth labels identifying the cell types present therein are determined based on expert annotations. The ground truth labels assigned to the images may be determined based at least on the confidence values of the corresponding expert annotations. For example, in a case where an expert annotation indicating the presence of macrophages in an image is associated with a confidence value that meets one or more thresholds (e.g., the probability that the expert annotation is accurate exceeds the threshold), a ground truth label indicating the presence of macrophages in the image may be assigned to the image. Alternatively, in a case where an expert annotation indicating the presence of macrophages in an image is associated with a confidence value that fails to meet the one or more thresholds (e.g., the probability that the annotation is accurate fails to exceed the threshold), a ground truth label indicating the presence of stromal cells in the image may be assigned to the image.
[0030] In a case where an image has been segmented to locate the individual cells present therein, the ground truth labels may be assigned to one or more corresponding pixels. For example, in a case where an expert annotation indicating the presence of macrophages in an image is associated with a confidence value that meets one or more thresholds (e.g., the probability that the annotation is accurate exceeds the threshold), the ground truth label assigned to the image may identify the one or more pixels corresponding to macrophages. Alternatively, in a case where an expert annotation indicating the presence of macrophages in an image is associated with a confidence value that fails to meet the one or more thresholds (e.g., the probability that the annotation is accurate fails to exceed the threshold), the ground truth label assigned to the image may identify the one or more pixels corresponding to stromal cells.
[0031] In some exemplary embodiments, the cell classification engine 114 may generate a composition profile 119 of the biological sample depicted in the image based at least on the one or more cell types identified within the image 117. For example, in some cases, the composition profile 119 of the biological sample may include one or more cell types among cell types such as macrophages, stromal cells, lymphocytes, tumor cells, plasma cells, endothelial cells, adipocytes, and neutrophils, which are identified as present in the biological sample. Alternatively or and / or additionally, the composition profile 119 of the biological sample may include the quantity, relative proportion, density, and / or spatial distribution of the one or more cell types present in the biological sample.
[0032] In some exemplary embodiments, a compositional profile of the biological sample depicted in image 117 can be generated to include an indication of whether cells identified as one cell type are within a threshold distance of cells identified as another cell type. For example, in some cases, a compositional profile of the biological sample can be generated to include an indication of whether cells identified as a second cell type are within a threshold distance of cells identified as lymphocytes. Alternatively or and / or additionally, a compositional profile of the biological sample depicted in image 117 can be generated to include the density and / or spatial distribution of one or more cell types across tumor and / or non-tumor regions of the biological sample. For example, in some cases, a compositional profile of the biological sample can be generated to include a first indication of whether cells identified as one cell type are present in the tumor region of the biological sample. Additionally, in some cases, a compositional profile of the biological sample can be generated to include a second indication of whether cells identified as another cell type are also present in the tumor region and / or non-tumor region of the biological sample. Thus, in some cases, a compositional profile of the biological sample is generated to include an indication of whether cells identified as a first cell type and / or a second cell type are present in the tumor region of the biological sample. Alternatively or and / or additionally, the compositional profile of the biological sample is generated to include an indication of whether cells identified as a first cell type and / or a second cell type are present in the non-tumor region of the biological sample.
[0033] In some cases, the diagnostic and treatment engine 116 can determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response for a patient associated with a biological sample based at least on the compositional profile 119 of the biological sample depicted in the image 117. For example, in some cases, the patient can be a non-small cell lung cancer (NSCLC) patient who is identified as a responder (or non-responder) to a combination immunotherapy (such as a combination therapy of T cell immunoglobulin and immunoreceptor tyrosine-based inhibitory motif (TIGIT) and atezolizumab) based at least on macrophages being within a threshold distance of lymphocytes (e.g., CD8-positive T cells) in the biological sample. In this context, treatment response can be evaluated based on various clinical outcomes, including, for example, progression-free survival, overall survival, mortality, and the like. To further illustrate, FIG. 3 depicts a schematic diagram showing an example of a tumor microenvironment filled with tumor cells, lymphocytes, stromal cells, macrophages, plasma cells, endothelial cells, adipocytes, and neutrophils. In some cases, the proximity between macrophages and lymphocytes (e.g., CD8-positive T cells) within the tumor microenvironment (e.g., a distance less than 40 μm) can be used as a biomarker to distinguish responders and non-responders to a combination immunotherapy (such as a combination therapy of T cell immunoglobulin and immunoreceptor tyrosine-based inhibitory motif (TIGIT) and atezolizumab), particularly in a population of non-small cell lung cancer (NSCLC) patients whose cancer cells are PD-L1 negative or exhibit a low expression level of the PD-L1 gene.
[0034] Figure 2A A flowchart depicting an example of a process 200 for identifying one or more cell types present within an image of a biological sample according to some exemplary embodiments is shown. Referring Figures 1 to 2A to, the process 200 can be performed by the digital pathology platform 110 (e.g., by the cell sorting engine 114) to identify, for example, the one or more cell types present within the image 117.
[0035] At 202, the digital pathology platform 110 can receive an image of a biological sample. For example, in some cases, the digital pathology platform 110 can receive the image 117 from the imaging system 120. In some cases, the image 117 can be a whole-slide image depicting the biological sample. Additionally, in some cases, the image 117 can be a stained whole-slide image, including, for example, a hematoxylin and eosin (H&E)-stained whole-slide image, a multiplex immunofluorescence (MxIF)-stained whole-slide image, an immunohistochemistry (IHC)-stained whole-slide image, and the like.
[0036] At 204, the digital pathology platform 110 may apply a cell classification model to identify one or more cell types present in a biological sample based at least on an image of the biological sample. In some exemplary embodiments, the digital pathology platform 110 (e.g., the cell segmentation engine 114) may apply a cell segmentation model 115 to determine one or more cell types present in the biological sample depicted in the image 117 based at least on the image 117. The cell segmentation model 115 may be trained to distinguish various different cell types, including for example a first cell type and a second cell type, where the likelihood that the first cell type is a macrophage meets a threshold and the likelihood that the second cell type is a macrophage does not meet the threshold. For example, in some cases, the cell segmentation model 115 may be trained to distinguish: a stromal cell class that includes fibroblasts and non-pigmented stromal macrophages; and a macrophage class that includes foamy macrophages, alveolar macrophages, and pigmented stromal macrophages. Additionally, in some cases, in addition to stromal cells and macrophages, the cell segmentation model 115 may also be trained to distinguish tumor cells, plasma cells, lymphocytes, endothelial cells, adipocytes, and neutrophils.
[0037] In some exemplary embodiments, the cell classification model 115 may classify each individual cell depicted in the image 117 by at least assigning a label to each cell present in the biological sample depicted in the image 117. In some cases, the label assigned to each cell may be a "hard label" indicating a specific cell type. Alternatively, the cell classification model 115 may determine a "soft label" for each cell, where the soft label may be a probability distribution of the cell being positive for each of different cell types (e.g., stromal cells, macrophages, tumor cells, lymphocytes, and plasma cells) p0, p1, …, p n of the probability distribution.
[0038] In some exemplary embodiments, the cell classification model 115 may operate on a segmented version of the image 117. That is, before the cell classification model 115 is applied to the image 117, the image 117 may undergo segmentation (e.g., watershed cell segmentation, machine learning-based cell segmentation, etc.) to locate the individual cells present therein. Thus, in some cases, the labels assigned to the image 117 by the cell classification model 115 may include labels for one or more of the pixels that are identified by the segmentation as part of a cell present in the image 117. For example, in the case where the image 117 is segmented to locate one or more of the cells present therein, the cell classification model 115 may assign a label indicating the corresponding cell type (e.g., macrophage, stromal cell, plasma cell, lymphocyte, or tumor cell) to each pixel associated with the cell.
[0039] In some exemplary embodiments, various machine learning models can be used to implement the cell classification model 115, and the various machine learning models include, for example, gradient boosting tree binary classifiers, random forests, naive Bayes classifiers, neural networks, k-means clustering models, logistic regression models, and the like. In some cases, the cell classification model 115 can include: a first machine learning model that is trained to identify one or more visible features within an image 117 of a biological sample; and a second machine learning model that is trained to identify the one or more cell types based at least on the one or more visible features extracted from the image 117. As used herein, the term "visible feature" can refer to a feature that can be identified, located, interpreted, inferred, and / or otherwise detected by, for example, visual inspection of an image by a person, machine, algorithm, or the like. Alternatively and / or additionally, the cell classification model 115 can be implemented as an end-to-end model that determines the cell types present in the image 117 based on one or more hidden features extracted from the image 117, and the one or more hidden features do not necessarily correspond to the aforementioned visible features.
[0040] At 206, the digital pathology platform 110 can generate a composition profile for the biological sample based at least on the one or more cell types identified in the biological sample. In some exemplary embodiments, the digital pathology platform 110 (e.g., the cell classification engine 114) can generate a composition profile 119 of the biological sample depicted in the image 117 based at least on the one or more cell types identified as being present in the image 117. For example, in some cases, the composition profile 119 can indicate the one or more cell types present in the biological sample, and the one or more cell types include, for example, stromal cells, macrophages, lymphocytes, plasma cells, tumor cells, endothelial cells, adipocytes, and neutrophils. Alternatively and / or additionally, the composition profile 119 can indicate one or more of the number, relative proportion, density, and spatial distribution of the one or more cell types present in the image 117.
[0041] At 208, the digital pathology platform 110 can determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response for a patient associated with a biological sample based at least on a compositional profile of the biological sample. In some exemplary embodiments, the digital pathology platform 110 (e.g., the diagnostic and treatment engine 116) can determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response for a patient associated with a biological sample based at least on the compositional profile 119 of the biological sample depicted in the image 117. In some cases, the compositional profile 119 of the biological sample can be generated by the cell classification engine 114 to indicate one or more of the cell types present in the biological sample. Alternatively and / or additionally, the cell classification engine 114 can generate the compositional profile 119 to indicate one or more of the quantity, relative proportion, density, and spatial distribution of the one or more cell types in the biological sample. As Figures 2A to 2B shown, for a patient with non-small cell lung cancer (NSCLC), the distance between macrophages and lymphocytes (e.g., CD8-positive T cells) in the tumor microenvironment can indicate whether the patient is a responder (or non-responder) to a combination immunotherapy (such as a combination of the T cell immunoreceptor with Ig and ITIM domains (TIGIT) and atezolizumab).
[0042] Figure 2B A flowchart depicting another example of a process 250 for identifying one or more cell types present within an image of a biological sample in accordance with some exemplary embodiments is shown. Referring to Figure 1 and Figures 2A to 2B , the process 250 can be performed by the digital pathology platform 110 (e.g., by the cell classification engine 114) to identify one or more macrophages and stromal cells present in the biological sample depicted in the image 117.
[0043] At 252, the digital pathology platform 110 can receive an image of the biological sample. For example, in some cases, the digital pathology platform 110 can receive the image 117 from the imaging system 120. As described above, in some cases, the image 117 can be a whole slide image depicting the biological sample. Additionally, in some cases, the image 117 can be a stained whole slide image, which can include, for example, a hematoxylin and eosin (H&E) stained whole slide image, a multiplex immunofluorescence (MxIF) stained whole slide image, an immunohistochemistry (IHC) stained whole slide image, and the like.
[0044] At 254, the digital pathology platform 110 can apply a cell classification model to identify a first cell type and / or a second cell type in a biological sample depicted in an image, where the likelihood that the first cell type is a macrophage meets a threshold and the likelihood that the second cell type is a macrophage fails to meet the threshold. In an exemplary embodiment, the digital pathology platform 110 (e.g., the cell segmentation engine 114) can apply a cell segmentation model 115 to determine one or more cell types present in a biological sample depicted in the image 117, at least based on the image 117. In some cases, the cell segmentation model 115 can be trained to distinguish macrophages from stromal cells. That is, the cell segmentation model 115 can be trained to identify a first cell type and a second cell type within the biological sample depicted in the image 117, where the likelihood that the first cell type is a macrophage meets a threshold and the likelihood that the second cell type is a macrophage fails to meet the threshold. In some cases, the first cell type can correspond to a macrophage category including foamy macrophages, alveolar macrophages, and pigmented stromal macrophages, while the second cell type can correspond to a stromal cell category including fibroblasts and non-pigmented stromal macrophages. In some cases, the cell segmentation model 115 can be further trained to identify additional cell types in the biological sample depicted in the image 117, where the additional cell types include, for example, tumor cells, plasma cells, lymphocytes, endothelial cells, adipocytes, neutrophils, and the like.
[0045] At 256, the digital pathology platform 110 can generate a compositional profile for a biological sample based at least on a first cell type and / or a second cell type identified in the biological sample. In some exemplary embodiments, the digital pathology platform 110 (e.g., the cell sorting engine 114) can generate a compositional profile 119 of the biological sample depicted in the image 117 to include an indication of whether the first cell type and / or the second cell type is present in the biological sample. That is, in some cases, the compositional profile 119 can be generated to include an indication of whether macrophages and / or stromal cells are present in the biological sample depicted in the image 117. In some cases, the cell sorting engine 114 can generate the compositional profile 119 to include one or more of the number, relative proportion, density, and spatial distribution of the first cell type (e.g., macrophages) and / or the second cell type (e.g., stromal cells) within the biological sample depicted in the image 117. Alternatively or and / or additionally, the cell sorting engine 114 can generate the compositional profile 119 to include the density and / or spatial distribution of the first cell type (e.g., macrophages) and / or the second cell type (e.g., stromal cells) across the tumor region and / or the non-tumor region of the biological sample. For example, in some cases, the compositional profile 119 of the biological sample can be generated to include a first indication of whether the first cell type (e.g., macrophages) and / or the second cell type (e.g., stromal cells) is present in the tumor region of the biological sample. Additionally, in some cases, the compositional profile 119 of the biological sample can be generated to include a second indication of whether the first cell type (e.g., macrophages) and / or the second cell type (e.g., stromal cells) is present in the non-tumor region of the biological sample.
[0046] At 258, the digital pathology platform 110 can determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response for a patient associated with the biological sample based at least on the compositional profile of the biological sample. In some exemplary embodiments, the digital pathology platform 110 (e.g., the diagnostic and treatment engine 116) can determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response for a patient associated with the biological sample based at least on the compositional profile 119 of the biological sample depicted in the image 117. For example, for a patient with non-small cell lung cancer (NSCLC), the distance between macrophages and lymphocytes (e.g., CD8-positive T cells) in the tumor microenvironment can indicate whether the patient is a responder (or non-responder) to a combination immunotherapy (such as the combination of T cell immunoglobulin and immunoreceptor with ITIM domains (TIGIT) and atezolizumab).
[0047] Figure 3AFIG. 0 depicts a flow diagram illustrating an example of a process 300 for training a cell classification model according to some exemplary embodiments. Referring Figure 1 to FIGS. 1 and 3, the process 300 may be performed by a digital pathology platform 110 (e.g., by a training engine 112) to train a cell classification model 115 to identify one or more cell types present within a biological sample depicted in an image 117.
[0048] At 302, the digital pathology platform 110 may generate an annotated training set. In some exemplary embodiments, the digital pathology platform 110 (e.g., the training engine 112) may train the cell classification model 115 based at least on a training set 113. In some cases, the training set 113 may be an annotated training set, where each image of the biological sample is annotated with one or more ground truth labels of the cell types present in the biological sample. For example, in the case where the images in the training set 113 are segmented to locate the individual cells present therein, each pixel depicting a cell may be associated with a ground truth label indicating the corresponding cell type.
[0049] In some exemplary embodiments, the training engine 112 can generate a training set by assigning at least one or more ground truth labels to each image included in the training set 113 based at least on expert annotations. For example, in a case where an expert annotation indicating the presence of macrophages is associated with a confidence value that meets one or more thresholds (e.g., the probability that the annotation is accurate exceeds the threshold), the training engine 112 can assign a ground truth label that identifies one or more corresponding pixels as depicting macrophages. Alternatively, in a case where an expert annotation indicating the presence of macrophages is associated with a confidence value that fails to meet the one or more thresholds (e.g., the probability that the annotation is accurate fails to exceed the threshold), the training engine 112 can assign a ground truth label that identifies one or more corresponding pixels as depicting stromal cells. In doing so, the cell classification engine 115 can be trained to distinguish between: a macrophage class that includes foamy macrophages, alveolar macrophages, and pigmented stromal macrophages; and a stromal cell class that includes fibroblasts and non-pigmented stromal macrophages. Non-pigmented stromal macrophages can be included in a separate stromal cell class together with fibroblasts in order to account for the uncertainty associated with the expert annotations. That is, non-pigmented stromal macrophages tend to be visually confused with fibroblasts. Thus, expert annotations associated with non-pigmented stromal macrophages are not reliable enough for the operation of training the cell classification model 115 to identify non-pigmented stromal macrophages together with other types of macrophages (such as foamy macrophages, alveolar macrophages, and pigmented stromal macrophages) in a single macrophage class, where the other types of macrophages can be identified with a much higher degree of certainty. Accordingly, the performance of the cell classification model 115, particularly with respect to the identification of macrophages with high-confidence expert annotations, can be improved by training the cell classification model 115 to recognize non-pigmented stromal macrophages as part of a separate stromal cell class that also includes fibroblasts.
[0050] In a conventional paradigm, the one or more images of the biological sample forming the training set 113 would be annotated based on a single overall class of macrophages that includes foamy macrophages, alveolar macrophages, and pigmented stromal macrophages associated with high-confidence expert annotations, as well as non-pigmented stromal macrophages associated with low-confidence expert annotations. Fibroblasts are often confused with non-pigmented stromal macrophages, but the fibroblasts are part of a separate fibroblast class, which means that training the cell classification model 115 to identify cancer cells, lymphocytes, fibroblasts, plasma cells, and macrophages without distinguishing between high-confidence macrophages and low-confidence macrophages may perpetuate the errors present in the less reliable expert annotations associated with fibroblasts and non-pigmented stromal macrophages.
[0051] Instead of five cell types, the cell classification model 115 can be trained to distinguish stromal cells, macrophages, lymphocytes, plasma cells, tumor cells, endothelial cells, adipocytes, and neutrophils. In particular, the cell classification model 115 can be trained to recognize non-pigmented stromal macrophages associated with low-confidence expert annotations as part of a separate stromal cell category (along with fibroblasts), while foamy macrophages, alveolar macrophages, and pigmented stromal macrophages associated with high-confidence expert annotations are part of the macrophage category. Figure 5 Three types of macrophages (e.g., foamy macrophages, alveolar macrophages, and pigmented stromal macrophages) are shown, which can be identified with confidence values (e.g., accurate probability exceeding a threshold) that meet one or more thresholds. Training the cell classification model 115 to recognize the macrophage category (which includes foamy macrophages, alveolar macrophages, and pigmented stromal macrophages but not non-pigmented stromal macrophages) can prevent the uncertainty in the expert annotations associated with non-pigmented stromal macrophages from degrading the performance of the cell classification model 115 in recognizing foamy macrophages, alveolar macrophages, and pigmented stromal macrophages.
[0052] At 304, the digital pathology platform 110 can train the cell classification model to recognize one or more cell types present in an image of a biological sample, at least based on the annotated training set. In some exemplary embodiments, the digital pathology platform 110 (e.g., the training engine 112) can train the cell classification model 115 at least based on the training set 113. As described above, the cell classification model 115 can be trained to distinguish multiple cell types, including stromal cells, macrophages, plasma cells, lymphocytes, tumor cells, endothelial cells, adipocytes, and neutrophils. Additionally, training the cell classification model 115 can include adjusting the learnable parameters of the cell classification model 115 until convergence, where the loss of the output of the cell classification model 115 stabilizes within a certain error range. For example, in some cases, training the cell classification model 115 can include adjusting one or more of the weights and biases of the cell classification model 115 by backpropagation of the loss present in the output of the cell classification model 115. In this context, the loss of the output of the cell classification model 115 can be a quantity corresponding to the difference between the label (e.g., cell type label) assigned to the image of the biological sample by the cell classification model 115 and the ground truth label (e.g., true cell type) associated with the image.
[0053] Figure 3B A flowchart depicting an example of a process 350 for generating a training set for training a cell classification model according to some exemplary embodiments is shown. Refer toFigure 1 and Figures 3A to 3B ,process 350 may be performed by the digital pathology platform 110 (e.g., by the training engine 112) to generate a training set 113 for training the cell classification model 115 to identify one or more cell types present within a biological sample depicted in the image 117. In some cases, process 350 may be performed to implement Figure 3A operation 302 of process 300 shown in
[0054] At 352, the digital pathology platform 110 may receive one or more user inputs corresponding to expert annotations indicating the presence of macrophages in the biological sample depicted in the image. In some exemplary embodiments, the digital pathology platform 110 (e.g., the training engine 112) may receive one or more user inputs from the client device 130 corresponding to expert annotations indicating the presence of macrophages in the biological sample depicted in the image. For example, in some cases, the expert annotations may identify one or more pixels in the image as depicting macrophages.
[0055] At 354, the digital pathology platform 110 may assign a first ground truth label indicating the presence of macrophages in the biological sample depicted in the image based at least on a confidence metric associated with the expert annotations that meet one or more thresholds. In some exemplary embodiments, if the confidence metric associated with the expert annotations meets one or more thresholds, the digital pathology platform 110 (e.g., the training engine 112) may assign the first ground truth label to the image. For example, in some cases, the training engine 112 may assign the first ground truth label to the image where the likelihood that the expert annotations accurately identify macrophages exceeds a threshold. In some cases, the first ground truth label may identify one or more corresponding pixels in the image as depicting macrophages (e.g., foamy macrophages, alveolar macrophages, and pigmented stromal macrophages).
[0056] At 356, the digital pathology platform 110 can assign a second ground truth label indicating the presence of stromal cells in the biological sample depicted in the image based at least on a confidence metric associated with an expert annotation that fails to meet the one or more thresholds. In some exemplary embodiments, if the confidence metric associated with the expert annotation fails to meet the one or more thresholds, the digital pathology platform 110 (e.g., the training engine 112) can assign the second ground truth label to the image. For example, in some cases, where the likelihood that the expert annotation accurately identifies macrophages fails to exceed a threshold, the training engine 112 can assign the second ground truth label to the image. In some cases, the second ground truth label can identify one or more corresponding pixels in the image as depicting stromal cells (e.g., fibroblasts and non-pigmented stromal macrophages).
[0057] At 358, the digital pathology platform 110 can generate a training sample that includes an image of a biological sample and a first ground truth label or a second ground truth label assigned to the image. In some exemplary embodiments, the digital pathology platform 110 can generate (for inclusion in the training set 113) a training sample that includes an image of a biological sample and either a first ground truth label or a second ground truth label assigned to the image.
[0058] Figure 6 A block diagram depicting an example of a computing system 600 according to some exemplary embodiments is shown. Referring Figures 1 to 6 to, the computing system 600 can be used to implement the digital pathology platform 110, the imaging system 120, the client device 130, and / or any components thereof.
[0059] As Figure 6 shown, the computing system 600 can include a processor 610, a memory 620, a storage device 630, and an input / output device 640. The processor 610, the memory 620, the storage device 630, and the input / output device 640 can be interconnected via a system bus 650. The processor 610 is capable of processing instructions for execution within the computing system 600. Such executed instructions can implement one or more components of, for example, the digital pathology platform 110, the imaging system 120, the client device 130, etc. In some exemplary embodiments, the processor 610 can be a single-threaded processor. Alternatively, the processor 610 can be a multi-threaded processor. The processor 610 is capable of processing instructions stored in the memory 620 and / or the storage device 630 to display graphical information for a user interface provided via the input / output device 640.
[0060] The memory 620 is a computer-readable medium for storing information within the computing system 600, such as a volatile or non-volatile computer-readable medium. For example, the memory 620 may store data structures representing a configuration object database. The storage device 630 is capable of providing persistent storage for the computing system 600. The storage device 630 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device or other suitable persistent storage means. The input / output device 640 provides input / output operations for the computing system 600. In some exemplary embodiments, the input / output device 640 includes a keyboard and / or a pointing device. In various specific implementations, the input / output device 640 includes a display unit for displaying a graphical user interface.
[0061] According to some exemplary embodiments, the input / output device 640 may provide input / output operations for a network device. For example, the input / output device 640 may include an Ethernet port or other networking ports to communicate with one or more wired and / or wireless networks (e.g., local area network (LAN), wide area network (WAN), Internet).
[0062] In some exemplary embodiments, the computing system 600 may be used to execute various interactive computer software applications that may be used to organize, analyze, and / or store data in various formats. Alternatively, the computing system 600 may be used to execute any type of software application. These applications may be used to perform various functions, such as scheduling functions (e.g., generating, managing, editing spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functions, communication functions, etc. The applications may include various additional functions or may be stand-alone computing products and / or functions. After being activated within the application, the functions may be used to generate a user interface provided via the input / output device 640. The user interface may be generated by the computing system 600 and presented to the user (e.g., on a computer screen monitor, etc.).
[0063] One or more aspects or features of the subject matter described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor (which can be dedicated or general purpose and coupled to receive data and instructions from, and to send data and instructions to, a storage system, at least one input device, and at least one output device). The programmable system or computing system can include clients and servers. Typically, the clients and servers are remotely located from each other and generally interact via a communication network. The relationship between the client and server is generated by computer programs running on respective computers and the client-server relationship between them.
[0064] These computer programs may also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer product, apparatus, and / or device (such as, for example, a disk, optical disk, memory, and programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor. A machine-readable medium can non-transitorily store such machine instructions (such as, for example, in non-transitory solid state memory or a magnetic hard disk drive or any equivalent storage medium). A machine-readable medium can alternatively or additionally store such machine instructions in a transitory manner (such as, for example, in a processor cache or other random access memory associated with one or more physical processor cores).
[0065] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device such as, for example, a cathode ray tube (CRT), or a liquid crystal display (LCD), or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device such as, for example, a mouse or a trackball by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback such as, for example, visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form including, but not limited to, voice, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single-point or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.
[0066] In the foregoing description and claims, phrases such as "at least one" or "one or more" may appear, followed by a list of elements or features. The term "and / or" may also appear in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, such phrases are intended to mean any of the recited elements or features individually, or any combination of any other recited elements or features. For example, the phrases "at least one of A and B"; "one or more of A and B"; "A and / or B" are each intended to mean "A alone, B alone, or A and B together". Similar interpretations apply to lists including three or more items. For example, the phrases "at least one of A, B, and C"; "one or more of A, B, and C" and "A, B, and / or C" are each intended to mean "A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together". The use of the term "based on" in the foregoing and the claims is intended to mean "at least in part based on", such that an unrecited feature or element is also permissible.
[0067] Exemplary embodiments
[0068] Embodiments disclosed herein may include:
[0069] 1. A computer-implemented method, comprising:
[0070] Receiving an image of a biological sample;
[0071] Apply a cell classification model to identify one or more cell types present in a biological sample based at least on an image of the biological sample, the cell classification model being trained to distinguish between multiple cell types, the multiple cell types including a first cell type and a second cell type, the probability that the first cell type is a macrophage meeting a threshold, and the probability that the second cell type is a macrophage not meeting the threshold; and
[0072] Generate a compositional profile for the biological sample based at least on the one or more cell types identified in the biological sample.
[0073] 2. The method according to embodiment 1, wherein the first cell type is a macrophage.
[0074] 3. The method according to embodiment 1 or embodiment 2, wherein the first cell type includes foamy macrophages, alveolar macrophages, and pigmented stromal macrophages.
[0075] 4. The method according to any one of embodiments 1 to 3, wherein the second cell type is a stromal cell.
[0076] 5. The method according to any one of embodiments 1 to 4, wherein the second cell type includes fibroblasts and non-pigmented stromal macrophages.
[0077] 6. The method according to any one of embodiments 1 to 5, wherein the multiple cell types further include tumor cells, lymphocytes, plasma cells, endothelial cells, adipocytes, and neutrophils.
[0078] 7. The method according to any one of embodiments 1 to 6, wherein the cell classification model includes a first machine learning model that is trained to extract one or more features from an image of the biological sample, and wherein the cell classification model further includes a second machine learning model that is trained to identify the one or more cell types present in the biological sample based at least on the one or more features extracted from the image of the biological sample.
[0079] 8. The method according to any one of embodiments 1 to 7, wherein the cell classification model includes an end-to-end machine learning model that is trained to identify the one or more cell types present in the biological sample.
[0080] 9. The method according to any one of embodiments 1 to 8, wherein the image is a whole-slide image.
[0081] 10. The method according to any one of embodiments 1 to 9, wherein the image is a whole-slide image stained with hematoxylin and eosin (H&E) or an immunohistochemistry (IHC)-stained whole-slide image.
[0082] 11. The method according to any one of embodiments 1 to 10, wherein the biological sample comprises one or more tissue fragments, free cells, and / or body fluids.
[0083] 12. The method according to any one of embodiments 1 to 11, wherein the biological sample comprises tumor tissue.
[0084] 13. The method according to any one of embodiments 1 to 12, wherein the compositional profile of the biological sample is generated to include the density and / or spatial distribution of the one or more cell types present in the biological sample.
[0085] 14. The method according to any one of embodiments 1 to 13, wherein the compositional profile of the biological sample is generated to include an indication of whether cells identified as one cell type are within a threshold distance of cells identified as another cell type.
[0086] 15. The method according to any one of embodiments 1 to 14, wherein the compositional profile of the biological sample is generated to include an indication of whether cells identified as a second cell type are within a threshold distance of cells identified as lymphocytes.
[0087] 16. The method according to any one of embodiments 1 to 15, wherein the compositional profile of the biological sample is generated to include the number and / or relative proportions of the one or more cell types present in the biological sample.
[0088] [[ID=z19]]17. The method according to any one of embodiments 1 to 16, wherein the compositional profile of the biological sample is generated to include a first indication of whether cells identified as one cell type are present in the tumor region of the biological sample.
[0089] [[ID=z22]]18. The method according to embodiment 17, wherein the compositional profile of the biological sample is further generated to include a second indication of whether cells identified as another cell type are also present in the tumor region and / or non-tumor region of the biological sample.
[0090] 19. The method according to any one of embodiments 1 to 18, wherein the compositional profile of the biodistribution is generated to include the spatial distribution of the one or more cell types across the tumor region and / or non-tumor region of the biological sample.
[0091] 20. The method according to any one of embodiments 1 to 19, wherein the compositional profile of the biological sample is generated to include an indication of whether cells identified as a first cell type and / or a second cell type are present in the tumor region of the biological sample.
[0092] Note: There seems to be a typo in the original text where "z19" and "z22" are used instead of "19" and "22" in the tags. The translation has been done as per the provided text with the tags as they are.21. The method according to any one of embodiments 1 to 20, wherein the compositional profile of the biological sample is generated to include an indication of whether cells identified as the first cell type and / or the second cell type are present in the non-tumor region of the biological sample.
[0093] 22. The method according to any one of embodiments 1 to 21, further comprising:
[0094] Training a cell classification model based at least on training data to identify the one or more cell types present in the biological sample.
[0095] 23. The method according to embodiment 22, further comprising:
[0096] Generating training data to include one or more images annotated with ground truth labels identifying at least one cell type present in each image.
[0097] 24. The method according to embodiment 23, wherein generating the training data includes: assigning a first ground truth label identifying a first cell type to the one or more images based at least on a first expert annotation identifying one or more macrophages being associated with a first confidence value meeting one or more thresholds.
[0098] 25. The method according to embodiment 24, wherein generating the training data further includes: assigning a second ground truth label identifying a second cell type to the one or more images based at least on a second expert annotation identifying the one or more macrophages being associated with a second confidence value failing to meet the one or more thresholds.
[0099] 26. The method according to any one of embodiments 1 to 25, further comprising:
[0100] Determining at least one of a disease diagnosis, disease progression, disease burden, and treatment response for a patient associated with the biological sample based at least on the compositional profile of the biological sample.
[0101] 27. The method according to any one of embodiments 1 to 26, further comprising:
[0102] Identifying a patient associated with the biological sample as a responder or non-responder to treatment based at least on the compositional profile of the biological sample.
[0103] 28. The method according to any one of embodiments 1 to 27, further comprising:
[0104] Determine (i) a first likelihood that a patient associated with a biological sample will respond to treatment, (ii) a second likelihood that the patient will relapse after treatment, and / or (iii) the persistence of the patient's response to treatment, based at least on the compositional profile of the biological sample.
[0105] 29. A system comprising:
[0106] at least one data processor; and
[0107] at least one memory storing instructions that, when executed by the at least one data processor, cause operations including the method according to any one of embodiments 1 to 28.
[0108] 30. A non - transitory computer - readable medium storing instructions that, when executed by at least one data processor, cause operations including the method according to any one of embodiments 1 to 28.
[0109] Depending on the desired configuration, the subject matter described herein can be embodied in a system, an apparatus, a method, and / or an article of manufacture. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Instead, they are only some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, other features and / or variations can be provided in addition to those features and / or variations set forth herein. For example, the above - described specific embodiments can be directed to various combinations and sub - combinations of the disclosed features and / or to combinations and sub - combinations of several further features disclosed above. Additionally, the logical flows depicted in the figures and / or described herein do not necessarily require the particular order or sequential order shown to achieve the desired result. Other specific embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method, comprising: Receiving an image of a biological sample; Applying a cell classification model to identify one or more cell types present in the biological sample based at least on the image of the biological sample, the cell classification model being trained to distinguish multiple cell types, the multiple cell types including a first cell type and a second cell type, the probability that the first cell type is a macrophage meeting a threshold, and the probability that the second cell type is the macrophage not meeting the threshold; And Generating a composition profile for the biological sample based at least on the one or more cell types identified in the biological sample.
2. The method according to claim 1, wherein the first cell type is a macrophage.
3. The method according to claim 1 or claim 2, wherein the first cell type includes foamy macrophages, alveolar macrophages, and pigmented stromal macrophages.
4. The method according to any one of claims 1 to 3, wherein the second cell type is a stromal cell.
5. The method according to any one of claims 1 to 4, wherein the second cell type includes fibroblasts and non-pigmented stromal macrophages.
6. The method according to any one of claims 1 to 5, wherein the multiple cell types further include tumor cells, lymphocytes, plasma cells, endothelial cells, adipocytes, and neutrophils.
7. The method according to any one of claims 1 to 6, wherein the cell classification model includes a first machine learning model trained to extract one or more features from the image of the biological sample, and wherein the cell classification model further includes a second machine learning model trained to identify the one or more cell types present in the biological sample based at least on the one or more features extracted from the image of the biological sample.
8. The method according to any one of claims 1 to 7, wherein the cell classification model includes an end-to-end machine learning model trained to identify the one or more cell types present in the biological sample.
9. The method according to any one of claims 1 to 8, wherein the image is a whole slide image.
10. The method according to any one of claims 1 to 9, wherein the image is a whole slide image stained with hematoxylin and eosin (H&E) or an immunohistochemistry (IHC)-stained whole slide image.
11. The method according to any one of claims 1 to 10, wherein the biological sample includes one or more tissue fragments, free cells, and / or body fluids.
12. The method according to any one of claims 1 to 11, wherein the biological sample includes tumor tissue.
13. The method according to any one of claims 1 to 12, wherein the composition profile of the biological sample is generated to include the density and / or spatial distribution of the one or more cell types present in the biological sample.
14. The method according to any one of claims 1 to 13, wherein the compositional profile of the biological sample is generated to include an indication of whether a cell identified as one cell type is within a threshold distance of a cell identified as another cell type.
15. The method according to any one of claims 1 to 14, wherein the compositional profile of the biological sample is generated to include an indication of whether a cell identified as the second cell type is within a threshold distance of a cell identified as a lymphocyte.
16. The method according to any one of claims 1 to 15, wherein the compositional profile of the biological sample is generated to include the number and / or relative proportion of the one or more cell types present in the biological sample.
17. The method according to any one of claims 1 to 16, wherein the compositional profile of the biological sample is generated to include a first indication of whether a cell identified as one cell type is present in a tumor region of the biological sample.
18. The method according to claim 17, wherein the compositional profile of the biological sample is further generated to include a second indication of whether a cell identified as another cell type is also present in the tumor region and / or a non-tumor region of the biological sample.
19. The method according to any one of claims 1 to 18, wherein the compositional profile of the biological profile is generated to include the spatial distribution of the one or more cell types across a tumor region of the biological sample and / or a non-tumor region of the biological sample.
20. The method according to any one of claims 1 to 19, wherein the compositional profile of the biological sample is generated to include an indication of whether a cell identified as the first cell type and / or the second cell type is present in a tumor region of the biological sample.
21. The method according to any one of claims 1 to 20, wherein the compositional profile of the biological sample is generated to include an indication of whether a cell identified as the first cell type and / or the second cell type is present in a non-tumor region of the biological sample.
22. The method according to any one of claims 1 to 21, further comprising: training the cell classification model based at least on training data to identify the one or more cell types present in the biological sample.
23. The method according to claim 22, further comprising: generating the training data to include one or more images annotated with ground truth labels identifying at least one cell type present in each image.
24. The method according to claim 23, wherein the generation of the training data comprises: assigning a first ground truth label identifying the first cell type to the one or more images based at least on a first expert annotation identifying one or more macrophages being associated with a first confidence value meeting one or more thresholds.
25. The method according to claim 24, wherein the generation of the training data further comprises: Assign a second ground truth label identifying the second cell type to the one or more images, based at least on a second expert annotation identifying the one or more macrophages and associated with a second confidence value that fails to meet the one or more thresholds.
26. The method according to any one of claims 1 to 25, further comprising: Determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response for a patient associated with the biological sample, based at least on the compositional profile of the biological sample.
27. The method according to any one of claims 1 to 26, further comprising: Identify a patient associated with the biological sample as a responder or non-responder to a treatment, based at least on the compositional profile of the biological sample.
28. The method according to any one of claims 1 to 27, further comprising: Determine (i) a first likelihood that a patient associated with the biological sample will respond to a treatment, (ii) a second likelihood that the patient will relapse after the treatment, and / or (iii) the persistence of the patient's response to the treatment, based at least on the compositional profile of the biological sample.
29. A system, comprising: At least one data processor; And At least one memory storing instructions that, when executed by the at least one data processor, cause operations including the method according to any one of claims 1 to 28.
30. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations including the method according to any one of claims 1 to 28.