Histological analysis using machine learning

A machine learning-based cell classification model distinguishes between specific macrophage classes using confidence thresholds in expert annotations, improving treatment response prediction in non-small cell lung cancer by addressing the limitations of unreliable expert annotations.

JP2025537582APending Publication Date: 2025-11-18GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025528622
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-18
Filing Date
2023-11-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing methods for identifying cell types in biological samples, particularly macrophages, are limited by unreliable expert annotations, leading to poor performance in distinguishing between different cell types, especially non-pigmented interstitial macrophages and fibroblasts, which affects treatment response prediction in diseases like non-small cell lung cancer.

Method used

A machine learning-based cell classification model is trained to distinguish between distinct classes of macrophages, including foamy, alveolar, and pigmented interstitial macrophages, and a separate class for non-pigmented interstitial macrophages and fibroblasts, using ground truth labels with confidence thresholds to improve accuracy.

Benefits of technology

Enhances the ability to predict treatment response by accurately identifying macrophage types, providing a more reliable indicator for immunotherapy outcomes in non-small cell lung cancer patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025537582000001_ABST
    Figure 2025537582000001_ABST
Patent Text Reader

Abstract

The method may include applying a cell classification model to identify one or more cell types present in the biological sample based at least on an image of the biological sample. The cell classification model may be trained to distinguish between multiple cell types, including a first cell type that meets a threshold likelihood of being a macrophage and a second cell type that does not meet the threshold likelihood of being a macrophage. A compositional profile of the biological sample may be generated based on the one or more cell types identified in the biological sample. At least one of a disease diagnosis, disease progression, disease burden, and treatment response of a patient associated with the biological sample may be determined based on the compositional profile of the biological sample. Related systems and computer program products are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 384,364, entitled "Machine Learning Enabled Histological Analysis," filed November 18, 2022, the disclosure of which is incorporated herein by reference in its entirety.

[0002] Technical Field The subject matter described herein relates generally to digital and computational pathology, and more specifically to machine learning-based approaches to cell type classification. [Background technology]

[0003] introduction Cancer cells induce significant molecular, cellular, and physical changes within host tissues to support further growth and metastasis. Tumors are typically heterogeneous collections of cells, including infiltrating and resident host cells, as well as secreted factors and extracellular matrix. On a broader scale, the tumor microenvironment may include immune cells, endothelial cells, and fibroblasts present in the vicinity of cancer cells. There is a strong correlation between tumor composition, tumor microenvironment, and the ability of tumors to maintain themselves at their primary site, evade immune responses, resist drug intervention, and grow to various secondary locations. Summary of the Invention

[0004] overview Systems, methods, and products, including computer program products, are provided for histological analysis using machine learning. In one embodiment, a system for identifying one or more cell types present in an image of a biological sample is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code that, when executed by the at least one processor, provides operations. The operations may include receiving an image of the biological sample; applying a cell classification model to identify one or more cell types present in the biological sample based at least on the image of the biological sample, the cell classification model being trained to distinguish between a plurality of cell types, including a first cell type that meets a threshold likelihood of being a macrophage and a second cell type that does not meet a threshold likelihood of being a macrophage; and generating a composition profile for the biological sample based at least on the one or more cell types identified in the biological sample.

[0005] In another aspect, a method for identifying one or more cell types present in an image of a biological sample is provided. The method may include receiving an image of the biological sample, applying a cell classification model to identify one or more cell types present in the biological sample based at least on the image of the biological sample, where the cell classification model is trained to distinguish between a plurality of cell types, including a first cell type that meets a threshold likelihood of being a macrophage and a second cell type that does not meet the threshold likelihood of being a macrophage, and generating a composition profile for the biological sample based at least on the one or more cell types identified in the biological sample.

[0006] In another aspect, a computer program product is provided for identifying one or more cell types present in an image of a biological sample. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations to occur. The operations may include receiving an image of the biological sample; applying a cell classification model to identify one or more cell types present in the biological sample based at least on the image of the biological sample, where the cell classification model is trained to distinguish between a plurality of cell types, including a first cell type that meets a threshold likelihood of being a macrophage and a second cell type that does not meet the threshold likelihood of being a macrophage; and generating a compositional profile for the biological sample based at least on the one or more cell types identified in the biological sample.

[0007] Implementations of the present subject matter can include, but are not limited to, methods according to the description provided herein, as well as articles comprising tangibly embodied machine-readable media operable to cause one or more machines (e.g., computers, etc.) to perform operations that implement one or more of the described features. Similarly, computer systems are described that may include one or more processors and one or more memories coupled to the one or more processors. The memory, which may include a non-transitory computer-readable or machine-readable storage medium, may include, encode, or store one or more programs that cause the one or more processors to perform one or more of the operations described herein. Computer-implemented methods consistent with one or more implementations of the present subject matter can be implemented by one or more data processors present in a single computing system or in multiple computing systems. Such multiple computing systems can be connected, for example, via one or more connections, including connections over a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), via a direct connection between one or more of the multiple computing systems, or the like, and can exchange data and / or instructions or other instructions, etc.

[0008] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the subject matter of the present disclosure are described for illustrative purposes in connection with identifying cell types associated with low-confidence expert annotations, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure define the scope of the protected subject matter. [Brief explanation of the drawings]

[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, serve to explain some of the principles associated with the disclosed embodiments.

[0010] [Figure 1] 1 depicts a system diagram showing an example of a digital pathology system, according to some exemplary embodiments.

[0011] [Figure 2A] 1 shows a flowchart illustrating an example of a process for identifying one or more cell types present in an image of a biological sample, according to some exemplary embodiments.

[0012] [Figure 2B] 10 shows a flowchart illustrating another example of a process for identifying one or more cell types present in an image of a biological sample, according to some exemplary embodiments.

[0013] [Figure 3A] 1 depicts a flowchart showing an example of a process for training a cell classification model, according to some exemplary embodiments.

[0014] [Figure 3B] 1 depicts a flowchart showing an example of a process for generating a training set for training a cell classification model, according to some exemplary embodiments.

[0015] [Figure 4] 1 depicts a schematic diagram showing an example of a tumor microenvironment, according to some exemplary embodiments.

[0016] [Figure 5] 1 depicts an example of a macrophage associated with a high-confidence expert annotation, according to some exemplary embodiments.

[0017] [Figure 6] 1 depicts a block diagram illustrating an example of a computing system, according to some illustrative embodiments.

[0018] Wherever practical, like reference numerals refer to like structures, features, or elements. DETAILED DESCRIPTION OF THE INVENTION

[0019] Detailed Description In highly heterogeneous diseases such as cancer, insight into the cell types and surrounding microenvironment that make up disease tissue can be essential for accurate diagnosis of disease subtypes, prognosis of disease progression, and prediction of response to various treatments. For example, non-small cell lung cancer (NSCLC) patients likely to respond to immunotherapy combinations, such as the T cell immunoreceptor with Ig and ITIM domain (TIGIT) and atezolizumab, can be identified based on cancer cells that exhibit high expression levels of the programmed death ligand-1 (PD-L1) gene. Within the population of NSCLC patients whose cancer cells are PD-L1 negative or express low levels of the PD-L1 gene, further differentiation between responders and non-responders to immunotherapy combinations is desirable.

[0020] Interactions between macrophages and lymphocytes (e.g., CD8-positive T cells) that result in the release of anti-inflammatory cytokines (e.g., IL-6 and / or the like) can inhibit patient response to immunotherapy combinations, such as the combination of T cell immunoreceptor with Ig and ITIM domain (TIGIT) and atezolizumab. Thus, macrophages can serve as an indicator of treatment response. However, macrophage detection alone is not a sufficient indicator. Knowing the type of macrophage present, along with the context of the tumor environment, can be important. This disclosure provides systems and methods for distinguishing responders from non-responders based on the presence of tumor-associated macrophages in the tumor environment. In particular, this disclosure describes methods for distinguishing between low-confidence and high-confidence macrophages. Additionally, this disclosure presents systems and methods for identifying and differentiating various macrophages, including interstitial macrophage cells and alveolar macrophages. When classifying macrophages at this level of sophistication, the disclosed systems and methods can provide predictions of treatment response or clinical outcome. In some cases, sophisticated macrophage type classification and detection may serve as a proxy or alternative to gene expression analysis in predicting treatment response.

[0021] In some exemplary embodiments, an image representing a biological sample may undergo histological analysis to identify one or more cell types present therein. For example, in some cases, a machine learning-based cell classification model may be applied to a whole-slide image of the biological sample (e.g., a hematoxylin and eosin (H&E)-stained whole-slide image, a multiplex immunofluorescence (MxIF)-stained whole-slide image, an immunohistochemistry (IHC)-stained whole-slide image, and / or the like) to identify one or more tumor cells, lymphocytes, plasma cells, fibroblasts, macrophages, endothelial cells, adipocytes, and neutrophils present in the biological sample. A compositional profile of the biological sample may be generated based at least on the one or more cell types identified within the biological sample. At least one of a disease diagnosis, disease progression, and treatment of a patient associated with the biological sample may be determined based on the compositional profile. For example, in some cases, a patient associated with a biological sample can be identified as a responder (or non-responder) to a particular treatment (e.g., a combination of T-cell immunoreceptor with Ig and ITIM domain (TIGIT) and atezolizumab) based on whether at least macrophages are identified as present within the biological sample.

[0022] In some exemplary embodiments, a cell classification model may be trained to distinguish between a variety of different cells, including, for example, tumor cells, lymphocytes, plasma cells, fibroblasts, macrophages, endothelial cells, adipocytes, neutrophils, and / or the like. In some cases, the cell classification model may be trained based on a training set including one or more images, each of which is associated with one or more ground truth labels that identify the cell type present therein. Furthermore, in some cases, the ground truth labels associated with each image in the training set may be determined based on expert annotations. Thus, the performance of the cell classification model may be limited by the accuracy of the expert annotations. In instances where the expert annotations cannot accurately identify a particular cell type within one or more images, the trained cell classification model may perform poorly when encountering that particular cell type. For example, certain types of macrophages, including non-pigmented stromal macrophages, may be difficult to visually decipher in images. While non-pigmented interstitial macrophages are often confused with other interstitial cells such as fibroblasts, foamy macrophages, alveolar macrophages, and pigmented interstitial macrophages are more easily distinguishable. Therefore, ground truth labels identifying non-pigmented interstitial macrophages and fibroblasts tend to be unreliable. Therefore, training a cell classification model to recognize a single, monolithic class of macrophages may reduce the performance of the trained cell classification model in correctly identifying macrophages that may be present in a biological sample.

[0023] In some exemplary embodiments, a cell classification model may be trained to recognize a first class of macrophages, including foamy macrophages, alveolar macrophages (including intraalveolar macrophages), and pigmented interstitial macrophages, and a second class of interstitial cells, including non-pigmented interstitial macrophages and fibroblasts. Training a cell classification model to recognize a distinct class of interstitial cells, including non-pigmented interstitial macrophages, which are likely to be confused with fibroblasts, may enhance the performance of the cell classification model in distinguishing between other macrophages, including, for example, foamy macrophages, alveolar macrophages (e.g., intraalveolar macrophages), pigmented interstitial macrophages, and / or the like. For example, in some cases, a cell classification model may be trained to distinguish between multiple cell types, including a first cell type whose likelihood of being a macrophage meets a threshold and a second cell type whose likelihood of being a macrophage does not exceed a threshold. The first cell type that meets a threshold likelihood of being a macrophage may include stromal cells such as fibroblasts and non-pigmented interstitial macrophages, while the second cell type that does not meet a threshold likelihood of being a macrophage may include foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages. In some cases, in addition to the first and second cell types, the cell classification model may be further trained to distinguish between tumor cells (e.g., non-small cell lung cancer (NSCLC) tumor cells and / or the like), lymphocytes, plasma cells, endothelial cells, adipocytes, and neutrophils.

[0024] In some exemplary embodiments, a cell classification model may be trained based on training data including one or more images annotated with ground truth labels that identify at least one cell type present in each image. Uncertainty present in expert annotations related to non-pigmented stromal macrophages may be captured when training a cell classification model to recognize distinct classes of stromal cells, including fibroblasts and non-pigmented stromal macrophages. For example, in some cases, the training data may be generated by assigning one or more images a first ground truth label that identifies a first cell type (e.g., stromal cell) based at least on a first expert annotation that identifies one or more macrophages associated with a confidence value that does not meet one or more thresholds. Furthermore, in some cases, the training data may be generated by assigning one or more images a second ground truth label that identifies a second cell type (e.g., macrophage) based at least on a second expert annotation that identifies one or more macrophages associated with a confidence value that meets one or more thresholds. Thus, the ground truth label of an image may indicate that macrophages are present if the corresponding expert annotations are sufficiently reliable (e.g., confidence values ​​that meet one or more thresholds). If the expert annotations indicating that macrophages are present are not sufficiently reliable (e.g., confidence values ​​that do not meet one or more thresholds), the ground truth label of an image may indicate that stromal cells are present instead of macrophages.

[0025] In some exemplary embodiments, the cell classification model may classify each cell shown in an image of a biological sample by assigning to each cell at least a "hard label" of a particular cell type or a "soft label" of the probability that the cell is positive for each of a plurality of different cell types (e.g., stromal cell, macrophage, tumor cell, lymphocyte, and plasma cell). In some cases, the cell classification model may operate on segmented images that have undergone watershed cell segmentation, machine learning-based cell segmentation, and / or the like, to localize individual cells present therein. Further, in some cases, the cell classification model may include a first machine learning model trained to identify one or more visible features in the image of the biological sample that can be identified, localized, interpreted, inferred, and / or otherwise detected through visual inspection of the image by, for example, a human, a machine, an algorithm, and / or the like. The cell classification may further include a second machine learning model trained to determine one or more cell types present in the image based at least on one or more visible features extracted from the image of the biological sample. Alternatively and / or additionally, the cell classification model may be implemented as an end-to-end model that determines the cell types present in an image based on one or more hidden features extracted from the image, which may not necessarily correspond to the visible features mentioned above.

[0026] FIG. 1 depicts a system diagram illustrating an example of a digital pathology system 100, according to some exemplary embodiments. Referring to FIG. 1, digital pathology system 100 may include a digital pathology platform 110, an imaging system 120, and a client device 130. As shown in FIG. 1, digital pathology platform 110, imaging system 120, and client device 130 may be communicatively coupled via a network 140. Network 140 may be a wired and / or wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and / or the like. Imaging system 120 may include one or more imaging devices, including, for example, a microscope, a digital camera, a whole slide scanner, a robotic microscope, and / or the like. Client device 130 may be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, and / or the like.

[0027] Referring again to FIG. 1, the digital pathology platform 110 may include a training engine 112, a cell classification engine 114, and a diagnosis and treatment engine 116. As shown in FIG. 1, the cell classification engine 114 may include a cell classification model 115 trained to identify one or more cell types present in an image 117 depicting a biological sample. In some cases, the image 117 may be a stained whole slide image (WSI), including, for example, a hematoxylin and eosin (H&E)-stained whole slide image, a multiplex immunofluorescence (MxIF)-stained whole slide image, an immunohistochemistry (IHC)-stained whole slide image, and / or the like. Further, in some cases, the image 117 may undergo watershed cell segmentation, machine learning-based cell segmentation, and / or the like to localize individual cells present therein. The cell classification model 115 can be trained to recognize macrophage classes including foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages, as well as distinct stromal cell classes including non-pigmented interstitial macrophages and fibroblasts. Additionally, in some cases, the cell classification model 115 can be trained to recognize tumor cells, plasma cells, lymphocytes, endothelial cells, adipocytes, and neutrophils.

[0028] In some exemplary embodiments, the training engine 112 may train the cell classification model 115 based at least on the training set 113 to distinguish between multiple cell types, including, for example, macrophages, stromal cells, tumor cells, plasma cells, lymphocytes, endothelial cells, adipocytes, and neutrophils. In some cases, the training engine 112 may generate the training set 113 including one or more images for which ground truth labels identifying the cell types present therein are determined based on expert annotations. The ground truth label assigned to an image may be determined based at least on the confidence value of the corresponding expert annotation. For example, if an expert annotation indicating the presence of macrophages in the image is associated with a confidence value that meets one or more thresholds (e.g., the probability that the expert annotation is accurate exceeds a threshold), the image may be assigned a ground truth label indicating the presence of macrophages in the image. Alternatively, if an expert annotation indicating the presence of macrophages in an image is associated with a confidence value that does not meet one or more thresholds (e.g., the probability that the annotation is accurate does not exceed a threshold), the image may be assigned a ground truth label indicating the presence of stromal cells in the image.

[0029] When an image is segmented to localize individual cells present therein, ground truth labels may be assigned to one or more corresponding pixels. For example, if an expert annotation indicating the presence of macrophages in an image is associated with a confidence value that meets one or more thresholds (e.g., the probability that the annotation is accurate exceeds the thresholds), the ground truth label assigned to the image may identify one or more pixels that correspond to macrophages. Alternatively, if an expert annotation indicating the presence of macrophages in an image is associated with a confidence value that does not meet one or more thresholds (e.g., the probability that the annotation is accurate does not exceed the thresholds), the ground truth label assigned to the image may identify one or more pixels that correspond to stromal cells.

[0030] In some exemplary embodiments, the cell classification engine 114 may generate a composition profile 119 of the biological sample depicted in the image based at least on one or more cell types identified within the image 117. For example, in some cases, the composition profile 119 of the biological sample may include one or more of the following cell types identified as being present within the biological sample: macrophages, stromal cells, lymphocytes, tumor cells, plasma cells, endothelial cells, adipocytes, and neutrophils. Alternatively and / or additionally, the composition profile 119 of the biological sample may include the amount, relative proportion, density, and / or spatial distribution of one or more cell types present in the biological sample.

[0031] In some exemplary embodiments, the compositional profile of the biological sample shown in image 117 may be generated to include an indication of whether cells identified as one cell type are present within a threshold distance of cells identified as another cell type. For example, in some cases, the compositional profile of the biological sample may be generated to include an indication of whether cells identified as a second cell type are present within a threshold distance of cells identified as lymphocytes. Alternatively and / or additionally, the compositional profile of the biological sample shown in image 117 may be generated to include the density and / or spatial distribution of one or more cell types across tumor and / or non-tumor regions of the biological sample. For example, in some cases, the compositional profile of the biological sample may be generated to include a first indication of whether cells identified as one cell type are present in tumor regions of the biological sample. Furthermore, in some cases, the compositional profile of the biological sample may be generated to include a second indication of whether cells identified as another cell type are also present in tumor regions of the biological sample and / or non-tumor regions of the biological sample. Thus, in some cases, a compositional profile of the biological sample is generated that includes an indication of whether cells identified as the first cell type and / or the second cell type are present in tumorous regions of the biological sample. Alternatively and / or additionally, a compositional profile of the biological sample is generated that includes an indication of whether cells identified as the first cell type and / or the second cell type are present in non-tumorous regions of the biological sample.

[0032] In some cases, the diagnostic and therapeutic engine 116 may determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response of a patient associated with the biological sample based at least on the composition profile 119 of the biological sample shown in the image 117. For example, in some cases, the patient may be a non-small cell lung cancer (NSCLC) patient identified as a responder (or non-responder) to an immunotherapy combination, such as a T-cell immunoreceptor with Ig and ITIM domains (TIGIT) and atezolizumab combination therapy, based at least on the presence of macrophages within a threshold distance of lymphocytes (e.g., CD8-positive T cells) in the biological sample. In this context, the treatment response may be assessed based on various clinical outcomes, including, for example, progression-free survival, overall survival, mortality, and / or the like. To further illustrate, Figure 3 shows a schematic diagram illustrating an example of a tumor microenvironment in which tumor cells, lymphocytes, stromal cells, macrophages, plasma cells, endothelial cells, adipocytes, and neutrophils reside. In some cases, proximity (e.g., a distance of less than 40 μm) between macrophages and lymphocytes (e.g., CD8-positive T cells) within the tumor microenvironment can serve as a biomarker to distinguish between responders and non-responders to immunotherapy combinations, such as T-cell immunoreceptor with Ig and ITIM domains (TIGIT) and atezolizumab combination therapy, particularly among populations of non-small cell lung cancer (NSCLC) patients whose cancer cells are PD-L1 negative or exhibit low expression levels of the PD-L1 gene.

[0033] 2A shows a flowchart illustrating an example of a process 200 for identifying one or more cell types present in an image of a biological sample, according to some exemplary embodiments. Referring to FIGS. 1-2A, process 200 may be performed by digital pathology platform 110, for example, by cell classification engine 114, to identify one or more cell types present in image 117, for example.

[0034] At 202, the digital pathology platform 110 may receive an image of a biological specimen. For example, in some cases, the digital pathology platform 110 may receive an image 117 from the imaging system 120. In some cases, the image 117 may be a whole-slide image depicting the biological specimen. Further, in some cases, the image 117 may be a stained whole-slide image, including, for example, a hematoxylin and eosin (H&E)-stained whole-slide image, a multiplex immunofluorescence (MxIF)-stained whole-slide image, an immunohistochemistry (IHC)-stained whole-slide image, and / or the like.

[0035] At 204, the digital pathology platform 110 may apply a cell classification model to identify one or more cell types present in the biological sample based on at least the image of the biological sample. In some exemplary embodiments, the digital pathology platform 110, e.g., the cell segmentation engine 114, may apply the cell segmentation model 115 to determine, based at least on the image 117, one or more cell types present in the biological sample shown in the image 117. The cell segmentation model 115 may be trained to distinguish between various different cell types, including, for example, a first cell type whose likelihood of being a macrophage meets a threshold and a second cell type whose likelihood of being a macrophage does not meet a threshold. For example, in some cases, the cell segmentation model 115 may be trained to distinguish between a stromal cell class, including fibroblasts and non-pigmented interstitial macrophages, and a macrophage class, including foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages. Furthermore, in some cases, the cell segmentation model 115 may be trained to distinguish between tumor cells, plasma cells, lymphocytes, endothelial cells, adipocytes, and neutrophils in addition to stromal cells and macrophages.

[0036] In some exemplary embodiments, the cell classification model 115 may classify each cell shown in the image 117 by assigning at least a label to each individual cell present in the biological sample shown in the image 117. In some cases, the label assigned to each cell may be a "hard label" indicating a particular cell type. Alternatively, the cell classification model 115 may determine a "soft label" for each cell, which is the probability p0, p1, ..., p of the cell being positive for each of n different cell types (e.g., stromal cell, macrophage, tumor cell, lymphocyte, and plasma cell). n It may be a probability distribution of

[0037] In some exemplary embodiments, cell classification model 115 may operate on a segmented version of image 117. That is, before cell classification model 115 is applied to image 117, image 117 may undergo segmentation (e.g., watershed cell segmentation, machine learning-based cell segmentation, and / or the like) to localize individual cells present therein. Thus, in some cases, the label assigned to image 117 by cell classification model 115 may include labels for one or more of the pixels identified by the segmentation as being part of a cell present in image 117. For example, if image 117 is segmented to localize one or more of the cells present therein, cell classification model 115 may assign to each pixel associated with a cell a label indicating the corresponding cell type (e.g., macrophage, stromal cell, plasma cell, lymphocyte, or tumor cell).

[0038] In some exemplary embodiments, the cell classification model 115 may be implemented using various machine learning models, including, for example, a gradient-boosted tree binary classifier, a random forest, a naive Bayes classifier, a neural network, a k-means clustering model, a logistic regression model, and / or the like. In some cases, the cell classification model 115 may include a first machine learning model trained to identify one or more visible features within the image 117 of the biological sample and a second machine learning model trained to identify one or more cell types based at least on one or more visible features extracted from the image 117. As used herein, the term "visible feature" may refer to a feature that can be identified, localized, interpreted, inferred, and / or otherwise detected through visual inspection of the image, for example, by a human, a machine, an algorithm, and / or the like. Alternatively and / or additionally, the cell classification model 115 may be implemented as an end-to-end model that determines the cell type present in the image 117 based on one or more hidden features extracted from the image 117, which may not necessarily correspond to the aforementioned visible features.

[0039] At 206, the digital pathology platform 110 may generate a compositional profile of the biological sample based at least on one or more cell types identified in the biological sample. In some exemplary embodiments, the digital pathology platform 110, e.g., the cell classification engine 114, may generate a compositional profile 119 of the biological sample shown in the image 117 based at least on one or more cell types identified as present in the image 117. For example, in some cases, the compositional profile 119 may indicate one or more cell types present in the biological sample, including, e.g., stromal cells, macrophages, lymphocytes, plasma cells, tumor cells, endothelial cells, adipocytes, and neutrophils. Alternatively and / or additionally, the compositional profile 119 may indicate one or more of the amounts, relative proportions, densities, and spatial distributions of one or more cell types present in the image 117.

[0040] At 208, the digital pathology platform 110 may determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response of a patient associated with the biological sample based at least on the composition profile of the biological sample. In some exemplary embodiments, the digital pathology platform 110, e.g., the diagnosis and treatment engine 116, may determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response of a patient associated with the biological sample based at least on the composition profile 119 of the biological sample shown in the image 117. In some cases, the composition profile 119 of the biological sample may be generated by the cell classification engine 114 to indicate one or more cell types present in the biological sample. Alternatively and / or additionally, the cell classification engine 114 may generate the composition profile 119 to indicate one or more of the amount, relative proportion, density, and spatial distribution of one or more cell types in the biological sample. As shown in Figures 2A-2B, for patients with non-small cell lung cancer (NSCLC), the distance between macrophages and lymphocytes (e.g., CD8-positive T cells) in the tumor microenvironment can indicate whether a patient is a responder (or non-responder) to immunotherapy combinations, such as the combination of T cell immunoreceptor bearing Ig and ITIM domain bearing atezolizumab (TIGIT).

[0041] 2B shows a flowchart illustrating another example of a process 250 for identifying one or more cell types present in an image of a biological sample, according to some exemplary embodiments. With reference to FIG. 1 and FIGS. 2A-2B, process 250 may be performed by digital pathology platform 110, for example, by cell classification engine 114, to identify one or more macrophages and stromal cells present in a biological sample shown in image 117.

[0042] At 252, the digital pathology platform 110 may receive an image of a biological specimen. For example, in some cases, the digital pathology platform 110 may receive an image 117 from the imaging system 120. As previously mentioned, in some cases, the image 117 may be a whole-slide image depicting the biological specimen. Further, in some cases, the image 117 may be a stained whole-slide image, including, for example, a hematoxylin and eosin (H&E)-stained whole-slide image, a multiplex immunofluorescence (MxIF)-stained whole-slide image, an immunohistochemistry (IHC)-stained whole-slide image, and / or the like.

[0043] At 254, the digital pathology platform 110 may apply a cell classification model to identify, in the biological sample shown in the image, a first cell type whose likelihood of being a macrophage meets a threshold and / or a second cell type whose likelihood of being a macrophage does not meet a threshold. In an exemplary embodiment, the digital pathology platform 110, e.g., the cell segmentation engine 114, may apply the cell segmentation model 115 to determine, based at least on the image 117, one or more cell types present in the biological sample shown in the image 117. In some cases, the cell segmentation model 115 may be trained to distinguish between macrophages and stromal cells. That is, the cell segmentation model 115 may be trained to identify, in the biological sample shown in the image 117, a first cell type whose likelihood of being a macrophage meets a threshold and a second cell type whose likelihood of being a macrophage does not meet a threshold. In some cases, the first cell type may correspond to a macrophage class including foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages, while the second cell type may correspond to a stromal cell class including fibroblasts and non-pigmented interstitial macrophages. In some cases, cell segmentation model 115 may be further trained to identify additional cell types within the biological sample shown in image 117, including, for example, tumor cells, plasma cells, lymphocytes, endothelial cells, adipocytes, neutrophils, and / or the like.

[0044] At 256, the digital pathology platform 110 may generate a compositional profile of the biological sample based on at least the first cell type and / or the second cell type identified in the biological sample. In some exemplary embodiments, the digital pathology platform 110, e.g., the cell classification engine 114, may generate the compositional profile 119 of the biological sample shown in the image 117 to include an indication of whether the first cell type and / or the second cell type are present in the biological sample. That is, in some cases, the compositional profile 119 may be generated to include an indication of whether macrophages and / or stromal cells are present in the biological sample shown in the image 117. In some cases, the cell classification engine 114 may generate the compositional profile 119 to include one or more of the amount, relative proportion, density, and spatial distribution of the first cell type (e.g., macrophages) and / or the second cell type (e.g., stromal cells) within the biological sample shown in the image 117. Alternatively and / or additionally, the cell classification engine 114 may generate the composition profile 119 to include the density and / or spatial distribution of a first cell type (e.g., macrophages) and / or a second cell type (e.g., stromal cells) across tumor and / or non-tumor regions of the biological sample. For example, in some cases, the composition profile 119 of the biological sample may be generated to include a first indication of whether the first cell type (e.g., macrophages) and / or the second cell type (e.g., stromal cells) are present in tumor regions of the biological sample. Further, in some cases, the composition profile 119 of the biological sample may be generated to include a second indication of whether the first cell type (e.g., macrophages) and / or the second cell type (e.g., stromal cells) are present in non-tumor regions of the biological sample.

[0045] At 258, the digital pathology platform 110 may determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response of a patient associated with the biological sample based at least on the composition profile of the biological sample. In some exemplary embodiments, the digital pathology platform 110, e.g., the diagnostic and treatment engine 116, may determine at least one of a disease diagnosis, disease progression, disease burden, and treatment response of a patient associated with the biological sample based at least on the composition profile 119 of the biological sample shown in the image 117. In the case of a non-small cell lung cancer (NSCLC) patient, for example, the distance between macrophages and lymphocytes (e.g., CD8-positive T cells) in the tumor microenvironment may indicate whether the patient is a responder (or non-responder) to an immunotherapy combination, such as a combination of T-cell immunoreceptor with Ig and ITIM domain (TIGIT) and atezolizumab.

[0046] 3A depicts a flowchart showing an example of a process 300 for training a cell classification model, according to some exemplary embodiments. With reference to FIGS. 1 and 3, process 300 may be performed by digital pathology platform 110, for example, by training engine 112, to train cell classification model 115 to identify one or more cell types present in a biological sample shown in image 117.

[0047] At 302, the digital pathology platform 110 may generate an annotated training set. In some exemplary embodiments, the digital pathology platform 110, e.g., the training engine 112, may train the cell classification model 115 based on at least the training set 113. In some cases, the training set 113 may be an annotated training set in which each image of a biological sample is annotated with one or more ground truth labels of cell types present in the biological sample. For example, if the images of the training set 113 are segmented to localize individual cells present therein, each pixel representing a cell may be associated with a ground truth label indicating the corresponding cell type.

[0048] In some exemplary embodiments, the training engine 112 may generate the training set by at least assigning one or more ground truth labels to each image included in the training set 113 based at least on the expert annotations. For example, if an expert annotation indicating the presence of a macrophage is associated with a confidence value that meets one or more thresholds (e.g., the probability that the annotation is accurate exceeds the threshold), the training engine 112 may assign ground truth labels that identify one or more corresponding pixels as depicting a macrophage. Alternatively, if an expert annotation indicating the presence of a macrophage is associated with a confidence value that does not meet one or more thresholds (e.g., the probability that the annotation is accurate does not exceed the threshold), the training engine 112 may assign ground truth labels that identify one or more corresponding pixels as depicting an interstitial cell. In doing so, the cell classification engine 115 may be trained to distinguish between a macrophage class, including foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages, and a stromal cell class, including fibroblasts and non-pigmented interstitial macrophages. Non-pigmented interstitial macrophages may be included in a separate stromal cell class along with fibroblasts to account for the uncertainty associated with expert annotations. That is, non-pigmented interstitial macrophages tend to be visually confused with fibroblasts. Therefore, the expert annotations associated with non-pigmented interstitial macrophages are not sufficiently reliable to train the cell classification model 115 to recognize non-pigmented interstitial macrophages in a single macrophage class, along with other types of macrophages, such as foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages, which can be identified with a much higher degree of certainty. Therefore, the performance of the cell classification model 115, particularly with regard to identifying those macrophages with high-confidence expert annotations, may be improved by training the cell classification model 115 to recognize non-pigmented interstitial macrophages as part of a separate stromal cell class that also includes fibroblasts.

[0049] In a traditional paradigm, one or more images of a biological specimen forming a training set 113 are annotated based on a single, monolithic class of macrophages, including foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages, which are associated with high-confidence expert annotations, and non-pigmented interstitial macrophages, which are associated with low-confidence expert annotations. Fibroblasts are often confused with non-pigmented interstitial macrophages, but being part of a distinct fibroblast class means that training a cell classification model 115 to recognize cancer cells, lymphocytes, fibroblasts, plasma cells, and macrophages without distinguishing between high-confidence and low-confidence macrophages may perpetuate errors present in the less-confident expert annotations associated with fibroblasts and non-pigmented interstitial macrophages.

[0050] Instead of five cell types, the cell classification model 115 can be trained to distinguish between stromal cells, macrophages, lymphocytes, plasma cells, tumor cells, endothelial cells, adipocytes, and neutrophils. In particular, the cell classification model 115 can be trained to recognize non-pigmented interstitial macrophages, which are associated with low-confidence expert annotations, as part of a separate stromal cell class along with fibroblasts, while foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages, which are associated with high-confidence expert annotations, are part of the macrophage class. Figure 5 shows three types of macrophages (e.g., foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages) that can be identified with confidence values ​​that meet one or more thresholds (e.g., probability of being correct, above a threshold). Training the cell classification model 115 to recognize macrophage classes that include foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages, but not non-pigmented interstitial macrophages, may prevent uncertainty in expert annotations associated with non-pigmented interstitial macrophages from degrading the performance of the cell classification model 115 in identifying foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages.

[0051] At 304, the digital pathology platform 110 may train a cell classification model based at least on the annotated training set to identify one or more cell types present in the image of the biological sample. In some exemplary embodiments, the digital pathology platform 110, e.g., the training engine 112, may train the cell classification model 115 based at least on the training set 113. As described above, the cell classification model 115 may be trained to distinguish between multiple cell types, including stromal cells, macrophages, plasma cells, lymphocytes, tumor cells, endothelial cells, adipocytes, and neutrophils. Furthermore, training the cell classification model 115 may include adjusting learnable parameters of the cell classification model 115 until convergence is reached, where a loss in the output of the cell classification model 115 settles within a specified error range. For example, in some cases, training the cell classification model 115 may include adjusting one or more of the weights and biases of the cell classification model 115 through backpropagation of a loss present in the output of the cell classification model 115. In this context, loss in the output of the cell classification model 115 may be an amount corresponding to the discrepancy between the label assigned by the cell classification model 115 to the image of the biological sample (e.g., cell type label) and the ground truth label associated with the image (e.g., ground truth cell type).

[0052] 3B depicts a flowchart showing an example of a process 350 for generating a training set for training a cell classification model, according to some exemplary embodiments. With reference to FIG. 1 and FIGS. 3A-3B, process 350 may be performed by digital pathology platform 110, e.g., by training engine 112, to generate a training set 113 for training cell classification model 115 to identify one or more cell types present within a biological sample shown in image 117. In some cases, process 350 may be performed to implement operation 302 of process 300 shown in FIG. 3A.

[0053] At 352, the digital pathology platform 110 may receive one or more user inputs corresponding to expert annotations indicating the presence of macrophages in the biological specimen shown in the image. In some exemplary embodiments, the digital pathology platform 110, e.g., the training engine 112, may receive one or more user inputs from the client device 130 corresponding to expert annotations indicating the presence of macrophages in the biological specimen shown in the image. For example, in some cases, the expert annotations may identify one or more pixels of the image as depicting macrophages.

[0054] At 354, the digital pathology platform 110 may assign a first ground truth label indicating the presence of macrophages in the biological specimen shown in the image based at least on a confidence metric associated with the expert annotations that meets one or more thresholds. In some exemplary embodiments, the digital pathology platform 110, e.g., the training engine 112, may assign the first ground truth label to the image if the confidence metric associated with the expert annotations meets one or more thresholds. For example, in some cases, the training engine 112 may assign the first ground truth label to the image if the likelihood that the expert annotations correctly identify macrophages exceeds a threshold. In some cases, the first ground truth label may identify one or more corresponding pixels of the image as indicative of macrophages (e.g., foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages).

[0055] At 356, the digital pathology platform 110 may assign a second ground truth label indicating the presence of stromal cells in the biological sample shown in the image based at least on the confidence metrics associated with the expert annotations that do not meet one or more thresholds. In some exemplary embodiments, the digital pathology platform 110, e.g., the training engine 112, may assign the second ground truth label to the image if the confidence metrics associated with the expert annotations do not meet one or more thresholds. For example, in some cases, the training engine 112 may assign the second ground truth label to the image if the likelihood that the expert annotations accurately identify macrophages does not exceed a threshold. In some cases, the second ground truth label may identify one or more corresponding pixels of the image as indicative of stromal cells (e.g., fibroblasts and non-pigmented stromal macrophages).

[0056] At 358, the digital pathology platform 110 may generate training samples including images of the biological sample and either the first ground truth label or the second ground truth label assigned to the image. In some exemplary embodiments, the digital pathology platform 110 may generate training samples including images of the biological sample and either the first ground truth label or the second ground truth label assigned to the image for inclusion in the training set 113.

[0057] 6 is a block diagram depicting an example of a computing system 600 according to some exemplary embodiments. Referring to FIGS. 1 and 6, computing system 600 may be used to implement digital pathology platform 110, imaging system 120, client device 130, and / or any components therein.

[0058] As shown in FIG. 6 , computing system 600 may include processor 610, memory 620, storage device 630, and input / output device 640. Processor 610, memory 620, storage device 630, and input / output device 640 may be interconnected via system bus 650. Processor 610 is capable of processing instructions for execution within computing system 600. Such executed instructions may implement, for example, one or more components of digital pathology platform 110, imaging system 120, client device 130, and / or the like. In some exemplary embodiments, processor 610 may be a single-threaded processor. Alternatively, processor 610 may be a multi-threaded processor. Processor 610 is capable of processing instructions stored in memory 620 and / or storage device 630 to display graphical information for a user interface provided via input / output device 640.

[0059] Memory 620 is a computer-readable medium, such as a volatile or non-volatile medium, that stores information within computing system 600. Memory 620 may store, for example, data structures representing a configuration object database. Storage device 630 may provide persistent storage for computing system 600. Storage device 630 may be a floppy disk drive, a hard disk drive, an optical disk drive, or a tape drive, or other suitable persistent storage means. Input / output device 640 provides input / output operations for computing system 600. In some exemplary embodiments, input / output device 640 includes a keyboard and / or a pointing device. In various implementations, input / output device 640 includes a display device for displaying a graphical user interface.

[0060] According to some demonstrative embodiments, the input / output devices 640 may provide input / output operations for network devices. For example, the input / output devices 640 may include an Ethernet port or other networking port for communicating with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0061] In some exemplary embodiments, computing system 600 may be used to execute various interactive computer software applications that may be used for organizing, analyzing, and / or storing various forms of data. Alternatively, computing system 600 may be used to execute any type of software application. These applications may be used to perform various functions, such as planning functions (e.g., creating, managing, editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, etc. Applications may include various add-in functions or may be standalone computing products and / or functions. When active within an application, functionality may be used to generate a user interface that is provided via input / output devices 640. The user interface may be generated by computing system 600 and presented to a user (e.g., on a computer screen monitor, etc.).

[0062] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs, field programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0063] These computer programs, which may also be referred to as programs, software, software applications, applications, components, or code, contain machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium may non-transitory store such machine instructions, such as, for example, a non-transitory solid-state memory, a magnetic hard drive, or any equivalent storage medium. Alternatively or additionally, a machine-readable medium may temporarily store such machine instructions, such as, for example, a processor cache or other random access memory associated with one or more physical processor cores.

[0064] To provide for user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having a display device, such as, for example, a cathode ray tube (CRT) or liquid crystal display (LCD) or light-emitting diode (LED) monitor, for displaying information to a user, and a keyboard and pointing device, such as, for example, a mouse or trackball, by which a user may provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, tactile feedback, etc., and input from the user may be received in any form, including acoustic input, voice input, and tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, etc.

[0065] In the above specification and claims, phrases such as "at least one of" or "one or more of" may appear before a list of consecutive elements or features. The term "and / or" may also be used in listings of two or more elements or features. Unless otherwise implicitly or explicitly stated by the context in which it is used, such phrases are intended to refer to any of the listed elements or features individually, or any of the listed elements or features in combination with any of the other listed elements or features. For example, the phrases "at least one of A and B," "one or more of A and B," and "A and / or B" are intended to mean "A only, B only, or A and B together," respectively. A similar interpretation is intended for lists containing more than two items. For example, the phrases "at least one of A, B, C," "one or more of A, B, C," and "A, B, and / or C" are intended to mean "A only, B only, C only, A and B together, A and C together, B and C together, or A, B and C together," respectively. Use of the term "based on" above and in the claims means "based at least in part on," and implies that unrecited features or elements are also permitted.

[0066] Illustrative Embodiments Embodiments disclosed herein may include the following. 1. A computer-implemented method comprising: receiving an image of the biological sample; applying a cell classification model to identify one or more cell types present in the biological sample based at least on the image of the biological sample, the cell classification model being trained to distinguish between a plurality of cell types including a first cell type that meets a threshold likelihood of being a macrophage and a second cell type that does not meet the threshold likelihood of being a macrophage; generating a composition profile for the biological sample based at least on one or more cell types identified in the biological sample. 2. The method of embodiment 1, wherein the first cell type is a macrophage. 3. The method of embodiment 1 or embodiment 2, wherein the first cell type comprises foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages. 4. The method of any one of embodiments 1-3, wherein the second cell type is a stromal cell. 5. The method of any one of embodiments 1-4, wherein the second cell type comprises fibroblasts and non-pigmented interstitial macrophages. 6. The method of any one of embodiments 1-5, wherein the plurality of cell types further comprises tumor cells, lymphocytes, plasma cells, endothelial cells, adipocytes, and neutrophils. 7. The method of any one of embodiments 1 to 6, wherein the cell classification model comprises a first machine learning model trained to extract one or more features from the image of the biological sample, and the cell classification model further comprises a second machine learning model trained to identify one or more cell types present in the biological sample based at least on the one or more features extracted from the image of the biological sample. 8. The method of any one of embodiments 1-7, wherein the cell classification model comprises an end-to-end machine learning model trained to identify one or more cell types present in the biological sample. 9. The method of any one of embodiments 1 to 8, wherein the image is an image of the whole slide. 10. The method of any one of embodiments 1-9, wherein the image is a hematoxylin and eosin (H&E) stained whole slide image or an immunohistochemistry (IHC) stained whole slide image. 11. The method of any one of embodiments 1 to 10, wherein the biological sample comprises one or more tissue fragments, free cells, and / or body fluids. 12. The method of any one of embodiments 1 to 11, wherein the biological sample comprises tumor tissue. 13. The method of any one of embodiments 1 to 12, wherein a compositional profile of the biological sample is generated that includes the density and / or spatial distribution of one or more cell types present in the biological sample. 14. The method of any one of embodiments 1-13, wherein a compositional profile of the biological sample is generated to include an indication of whether cells identified as one cell type are within a threshold distance of cells identified as another cell type. 15. The method of any one of embodiments 1 to 14, wherein a compositional profile of the biological sample is generated to include an indication of whether cells identified as the second cell type are within a threshold distance of cells identified as lymphocytes. 16. The method of any one of embodiments 1 to 15, wherein a composition profile of the biological sample is generated to include the amount and / or relative proportion of one or more cell types present in the biological sample. 17. The method of any one of embodiments 1-16, wherein a compositional profile of the biological sample is generated to include a first indication of whether cells identified as one cell type are present in a tumor region of the biological sample. 18. The method of embodiment 17, wherein the compositional profile of the biological sample is further generated to include a second indication of whether cells identified as another cell type are also present in tumor regions of the biological sample and / or non-tumor regions of the biological sample. 19. The method of any one of embodiments 1 to 18, wherein the composition profile of the biological profile is generated to include the spatial distribution of one or more cell types across a tumor region of the biological sample and / or a non-tumor region of the biological sample. 20. The method of any one of embodiments 1-19, wherein a compositional profile of the biological sample is generated to include an indication of whether cells identified as the first cell type and / or the second cell type are present in a tumor region of the biological sample. 21. The method of any one of embodiments 1 to 20, wherein a compositional profile of the biological sample is generated to include an indication of whether cells identified as the first cell type and / or the second cell type are present in non-tumorous regions of the biological sample. 22. The method of any one of embodiments 1 to 21, further comprising training a cell classification model to identify one or more cell types present in the biological sample based at least on the training data. 23. The method of embodiment 22, further comprising generating training data to include one or more images annotated with ground truth labels that identify at least one cell type present in each image. 24. The method of embodiment 23, wherein generating training data includes assigning a first ground truth label identifying a first cell type to one or more images based at least on a first expert annotation identifying one or more macrophages associated with a first confidence value that satisfies one or more thresholds. 25. The method of embodiment 24, wherein generating training data further includes assigning a second ground truth label identifying a second cell type to one or more images based at least on a second expert annotation identifying one or more macrophages associated with a second confidence value that does not meet one or more thresholds. 26. The method of any one of embodiments 1 to 25, further comprising determining at least one of a disease diagnosis, disease progression, disease burden, and treatment response of a patient associated with the biological sample based at least on the composition profile of the biological sample. 27. The method of any one of embodiments 1 to 26, further comprising identifying a patient associated with the biological sample as a responder to the treatment or a non-responder to the treatment based at least on the composition profile of the biological sample. 28. The method of any one of embodiments 1 to 27, further comprising determining, based at least on the composition profile of the biological sample, (i) a first likelihood that the patient associated with the biological sample will respond to treatment, (ii) a second likelihood that the patient will relapse after treatment, and / or (iii) the durability of the patient's response to treatment. 29. A system comprising: at least one data processor; At least one memory storing instructions that, when executed by at least one data processor, result in operations including the method according to any one of embodiments 1 to 28; A system comprising: 30. A non-transitory computer-readable medium having stored thereon instructions that, when executed by at least one data processor, result in operations including the method of any one of embodiments 1 to 28.

[0067] The subject matter described herein may be embodied in systems, devices, methods, and / or articles, depending on the desired configuration. The implementations set forth in the above description do not necessarily represent all implementations of the subject matter described herein. Rather, they are merely some examples consistent with aspects related to the described subject matter. While several variations have been detailed above, other modifications and additions are possible. In particular, additional features and / or variations may be provided in addition to those described herein. For example, the above-described embodiments may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several additional features disclosed above. Additionally, the logic flow depicted in the accompanying figures and / or described herein does not necessarily require the particular order shown or sequential order to achieve desirable results. Other implementations may be within the scope of the following claims.

Claims

1. 1. A computer-implemented method comprising: receiving an image of the biological sample; applying a cell classification model to identify one or more cell types present in the biological sample based at least on the image of the biological sample, the cell classification model being trained to distinguish between a plurality of cell types including a first cell type whose likelihood of being a macrophage meets a threshold and a second cell type whose likelihood of being a macrophage does not meet the threshold; generating a compositional profile for the biological sample based at least on the one or more cell types identified in the biological sample; 11. A computer-implemented method comprising:

2. The method of claim 1 , wherein the first cell type is a macrophage.

3. 3. The method of claim 1 or claim 2, wherein the first cell type comprises foamy macrophages, intraalveolar macrophages, and pigmented interstitial macrophages.

4. The method of any one of claims 1 to 3, wherein the second cell type is a stromal cell.

5. The method of any one of claims 1 to 4, wherein the second cell type comprises fibroblasts and non-pigmented interstitial macrophages.

6. The method of any one of claims 1 to 5, wherein the plurality of cell types further comprises tumor cells, lymphocytes, plasma cells, endothelial cells, adipocytes, and neutrophils.

7. 7. The method of any one of claims 1 to 6, wherein the cell classification model comprises a first machine learning model trained to extract one or more features from the image of the biological sample, and wherein the cell classification model further comprises a second machine learning model trained to identify the one or more cell types present in the biological sample based at least on the one or more features extracted from the image of the biological sample.

8. 8. The method of any one of claims 1 to 7, wherein said cell classification model comprises an end-to-end machine learning model trained to identify the one or more cell types present in the biological sample.

9. The method of any one of claims 1 to 8, wherein the image is an image of the whole slide.

10. 10. The method of any one of claims 1 to 9, wherein the image is a hematoxylin and eosin (H&E) stained whole slide image or an immunohistochemistry (IHC) stained whole slide image.

11. The method of any one of claims 1 to 10, wherein the biological sample comprises one or more tissue fragments, free cells, and / or body fluids.

12. The method of any one of claims 1 to 11, wherein the biological sample comprises tumor tissue.

13. 13. The method of any one of claims 1 to 12, wherein the compositional profile of the biological sample is generated to comprise the density and / or spatial distribution of the one or more cell types present in the biological sample.

14. 14. The method of any one of claims 1 to 13, wherein the compositional profile of the biological sample is generated to include an indication of whether cells identified as one cell type are within a threshold distance of cells identified as another cell type.

15. 15. The method of any one of claims 1 to 14, wherein the compositional profile of the biological sample is generated to include an indication of whether cells identified as the second cell type are within a threshold distance of cells identified as lymphocytes.

16. 16. The method of any one of claims 1 to 15, wherein the compositional profile of the biological sample is generated to include the amount and / or relative proportion of the one or more cell types present in the biological sample.

17. 17. The method of any one of claims 1 to 16, wherein the compositional profile of the biological sample is generated to include a first indication of whether cells identified as one cell type are present in a tumor region of the biological sample.

18. 18. The method of claim 17, wherein the compositional profile of the biological sample is further generated to include a second indication of whether cells identified as another cell type are also present in the tumor region of the biological sample and / or in non-tumor regions of the biological sample.

19. 19. The method of any one of claims 1 to 18, wherein the compositional profile of the biological profile is generated to include a spatial distribution of the one or more cell types across a tumor region of the biological sample and / or a non-tumor region of the biological sample.

20. 20. The method of any one of claims 1 to 19, wherein the compositional profile of the biological sample is generated to include an indication of whether cells identified as the first cell type and / or the second cell type are present in a tumor region of the biological sample.

21. 21. The method of any one of claims 1 to 20, wherein the compositional profile of the biological sample is generated to include an indication of whether cells identified as the first cell type and / or the second cell type are present in non-tumorous regions of the biological sample.

22. training the cell classification model to identify the one or more cell types present in the biological sample based on at least training data; The method of any one of claims 1 to 21, further comprising:

23. generating the training data to include one or more images annotated with ground truth labels that identify at least one cell type present in each image; 23. The method of claim 22, further comprising:

24. 24. The method of claim 23, wherein generating the training data comprises assigning to the one or more images a first ground truth label identifying the first cell type based at least on a first expert annotation identifying one or more macrophages associated with a first confidence value that satisfies one or more thresholds.

25. 25. The method of claim 24, wherein generating the training data further comprises assigning to the one or more images a second ground truth label identifying the second cell type based at least on a second expert annotation identifying the one or more macrophages associated with a second confidence value that does not satisfy the one or more thresholds.

26. determining at least one of a disease diagnosis, disease progression, disease burden, and treatment response of a patient associated with the biological sample based at least on the compositional profile of the biological sample; The method of any one of claims 1 to 25, further comprising:

27. identifying a patient associated with the biological sample as a responder to a treatment or a non-responder to the treatment based at least on the composition profile of the biological sample; The method of any one of claims 1 to 26, further comprising:

28. determining, based at least on the composition profile of the biological sample, (i) a first likelihood associated with the biological sample that the patient will respond to a treatment, (ii) a second likelihood that the patient will relapse after the treatment, and / or (iii) a durability of the patient's response to the treatment; The method of any one of claims 1 to 27, further comprising:

29. 1. A system comprising: at least one data processor; at least one memory having stored thereon instructions which, when executed by said at least one data processor, result in operations comprising the method of any one of claims 1 to 28; A system comprising:

30. A non-transitory computer readable medium having stored thereon instructions which, when executed by at least one data processor, result in operations comprising the method of any one of claims 1 to 28.