Probabilistic feature identification for machine learning-enabled cell phenotyping
A machine learning-based phenotyping model with probabilistic outputs addresses variability in traditional and existing methods by accurately quantifying uncertainty in cell phenotyping, enhancing disease diagnosis and treatment response in heterogeneous diseases.
Patent Information
- Application Number
- JP2025524262
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-02
- Filing Date
- 2023-10-31
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional histological analysis techniques for identifying cellular phenotypes in microscopic images are prone to high inter- and intra-pathologist variability, and existing machine learning-based solutions are susceptible to similar variability due to unreliable expert annotation, leading to inaccuracies in determining cell phenotypes.
A machine learning-based phenotyping model generates probabilistic outputs to quantify uncertainty in identifying cell phenotypes, using biomarker identification models and phenotypic discrimination models to provide accurate quantification of cell features and phenotypes, reducing reliance on binary outputs.
The probabilistic approach enhances the accuracy of cell phenotyping by capturing uncertainty, thereby improving disease diagnosis, progression prediction, and treatment response determination in heterogeneous diseases like cancer.
Smart Images

Figure 2025537513000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 382,075, filed November 2, 2022, entitled "Probabilistic Identification of Features for Machine Learning-Enabled Cellular Phenotyping," the disclosure of which is incorporated herein by reference in its entirety.
[0002] Technical Field The subject matter described herein relates generally to digital and computational pathology, and more particularly to probabilistic approaches to identifying features for machine-learning-enabled determination of cellular phenotype. [Background technology]
[0003] background A cellular phenotype may refer to a unique combination of morphological and functional characteristics resulting from various cellular processes, including, for example, gene expression, protein expression, and / or the like. In some cases, the complex interplay between a cell's genome, epigenome, and local environment can result in a combination of observable characteristics collectively known as the cellular phenotype. Cellular phenotypes, including those of tumor cells, are typically attributed to genomic instability, but in recent years, the influence of epigenetics and the microenvironment has gained attention. Such non-genetic factors may further increase the inherent diversity and plasticity of tumor cells. At the tumor level, non-genetic factors may contribute to greater phenotypic heterogeneity, allowing tumor cells to evade immune responses and resist drug intervention. Summary of the Invention
[0004] overview Systems, methods, and articles, including computer program products, are provided for probabilistic identification of features for machine learning-enabled cellular phenotyping. In one aspect, a system for probabilistic identification of features for machine learning-enabled cellular phenotyping is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code that, when executed by the at least one processor, provides operations. The operations may include extracting a plurality of features for each cell in the population of cells from a first image depicting the population of cells; applying a biomarker identification model to determine, based at least on the plurality of features associated with each cell in the population of cells, whether the cell is associated with a plurality of biomarkers, is positive for the plurality of biomarkers, or is negative for the plurality of biomarkers; determining a set of probabilities for each cell in the population of cells based at least on the output of the biomarker identification model, the set of probabilities including, for each biomarker in the plurality of biomarkers, a probability that a corresponding cell is associated with the biomarker, is positive for the biomarker, or is negative for the biomarker; identifying a first subset of cells exhibiting a first phenotype based at least on the set of probabilities associated with each cell in the population of cells; and identifying a first set of features associated with the first subset of cells as indicative of a first probability that the cell is associated with the first phenotype, is positive for the first phenotype, or is negative for the first phenotype.
[0005] In another aspect, a method for probabilistic identification of features for cell phenotyping enabled by machine learning is provided. The method may include extracting a plurality of features for each cell in the population of cells from a first image depicting the population of cells, applying a biomarker identification model to determine whether the cell is associated with a plurality of biomarkers, is positive for the plurality of biomarkers, or is negative for the plurality of biomarkers based at least on the plurality of features associated with each cell in the population of cells, determining a set of probabilities for each cell in the population of cells based at least on the output of the biomarker identification model, the set of probabilities comprising, for each biomarker in the plurality of biomarkers, a probability that a corresponding cell is associated with the biomarker, is positive for the biomarker, or is negative for the biomarker, identifying a first subset of cells exhibiting a first phenotype based at least on the set of probabilities associated with each cell in the population of cells, and identifying the first set of features associated with the first subset of cells as indicative of a first probability that the cell is associated with the first phenotype, is positive for the first phenotype, or is negative for the first phenotype.
[0006] In another aspect, a computer program product is provided for machine learning-enabled probabilistic identification of features for cell phenotyping. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause an operation. The operations may include extracting a plurality of features for each cell in the population of cells from a first image depicting the population of cells; applying a biomarker identification model to determine, based at least on the plurality of features associated with each cell in the population of cells, whether the cell is associated with a plurality of biomarkers, is positive for the plurality of biomarkers, or is negative for the plurality of biomarkers; determining a set of probabilities for each cell in the population of cells based at least on the output of the biomarker identification model, the set of probabilities including, for each biomarker in the plurality of biomarkers, a probability that a corresponding cell is associated with the biomarker, is positive for the biomarker, or is negative for the biomarker; identifying a first subset of cells exhibiting a first phenotype based at least on the set of probabilities associated with each cell in the population of cells; and identifying a first set of features associated with the first subset of cells as indicative of a first probability that the cell is associated with the first phenotype, is positive for the first phenotype, or is negative for the first phenotype.
[0007] Implementations of the present subject matter may include, but are not limited to, methods according to the description provided herein, as well as articles comprising tangibly embodied machine-readable media operable to cause one or more machines (e.g., computers, etc.) to perform operations that implement one or more of the described features. Similarly, computer systems are described that may include one or more processors and one or more memories coupled to the one or more processors. The memory, which may include a non-transitory computer-readable or machine-readable storage medium, may include, encode, or store one or more programs that cause the one or more processors to perform one or more of the operations described herein. Computer-implemented methods consistent with one or more implementations of the present subject matter may be implemented by one or more data processors present in a single computing system or in multiple computing systems. Such multiple computing systems may be connected, e.g., to exchange data and / or commands or other instructions, etc., via one or more connections, including, for example, connections via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), direct connections between one or more of the multiple computing systems, etc.
[0008] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the presently disclosed subject matter are described for illustrative purposes in connection with identifying features for machine learning-enabled cell phenotyping, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure define the scope of the protected subject matter. [Brief explanation of the drawings]
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, serve to explain some of the principles associated with the disclosed embodiments.
[0010] [Figure 1] FIG. 1 depicts a system diagram showing an example of a digital pathology system, according to some exemplary embodiments.
[0011] [Figure 2] 1 depicts a flowchart illustrating an example of a process for machine learning enabled probabilistic feature identification for cell phenotyping, according to some exemplary embodiments.
[0012] [Figure 3] FIG. 1 depicts a schematic diagram showing an example of a workflow for machine learning enabled probabilistic identification of features for cell phenotyping, according to some exemplary embodiments.
[0013] [Figure 4] 1 depicts examples of features extracted from an image depicting a population of cells, according to some exemplary embodiments.
[0014] [Figure 5A] FIG. 1 depicts a schematic diagram showing an example of a process for training and validating a biomarker discrimination model, according to some exemplary embodiments.
[0015] [Figure 5B] FIG. 1 depicts a schematic diagram illustrating an example of a process for training and validating a phenotype discrimination model, according to some exemplary embodiments.
[0016] [Figure 6A] 1 depicts a visualization of an example of a reduced dimensionality representation of a biomarker probability dataset, according to some exemplary embodiments.
[0017] [Figure 6B]1 depicts another visualization of an example of a reduced dimensionality representation of a biomarker probability dataset, according to some exemplary embodiments.
[0018] [Figure 7] 1 depicts a block diagram of an example computing system, according to some illustrative embodiments.
[0019] Wherever practical, like reference numerals refer to like structures, features, or elements. DETAILED DESCRIPTION OF THE INVENTION
[0020] Detailed Description In highly heterogeneous diseases such as cancer, insight into the phenotype of cells forming diseased tissue and the surrounding microenvironment can be essential for accurate diagnosis of disease subtypes, prognosis of disease progression, and prediction of response to various treatments. For example, non-Hodgkin's lymphoma patients at high risk for disease progression with standard-of-care treatments (e.g., the combination immunochemotherapy R-CHOP (rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone)) can be identified by characterizing the immune microenvironment, including identification of lymph node-resident immune cells and infiltrating neoplastic cells. Nevertheless, traditional histological analysis techniques for identifying cellular phenotypes depicted in microscopic images (e.g., hematoxylin and eosin (H&E)-stained whole-slide images, multiplex immunofluorescence (MxIF)-stained whole-slide images, etc.) are prone to error due to high levels of inter- and intra-pathologist variability. On the other hand, existing machine learning-based solutions are also susceptible to inter- and intra-pathologist variability, at least because the training of machine learning-based cell phenotyping models relies on expert annotation of training samples, which is not sufficiently reliable due to the presence of inter- and intra-pathologist variability.
[0021] In some exemplary embodiments, a machine learning-based phenotyping model may be trained to determine the probability that one or more of the cells depicted in the image are associated with a particular phenotype, are positive for a particular phenotype, or are negative for a particular phenotype based on one or more features extracted from an image depicting a population of cells. The machine learning-based phenotyping model may be trained to generate a probabilistic output instead of a binary output to provide a more accurate quantification of the error (or uncertainty) present in determining that one or more cells depicted in the image are positive (or negative) for a particular phenotype. In some cases, if the probability that one or more of the cells depicted in the image are positive for a particular phenotype meets one or more thresholds, one or more downstream tasks, such as disease diagnosis, disease progression, disease burden, and / or treatment response determination, may be performed. Furthermore, in some cases, the same machine learning-based phenotyping model or one or more separate machine learning-based phenotyping models may be trained to determine the probability that one or more of the cells depicted in the image are positive for another phenotype.
[0022] In some exemplary embodiments, one or more features present in an image may be identified as indicating that a cell is associated with, positive for, or negative for a particular phenotype, or the probability that a cell is associated with, positive for, or negative for a particular phenotype. For example, in some cases, multiple features may be extracted from images depicting a population of cells, including, for example, hematoxylin and eosin (H&E)-stained whole slide images, multiplex immunofluorescence (MxIF)-stained whole slide images, and / or the like. These features may be collected across multiple channels. For example, in some cases, each channel may correspond to one or more of the emission wavelength of a fluorescent dye applied to the image, a metal ion collected by a mass cytometer, a nucleotide sequence identified by barcode hybridization, a nucleotide sequence identified by sequencing, and / or the like. The one or more features that indicate that a cell is associated with, positive for, or negative for a particular phenotype, or that indicate the probability of being associated with, positive for, or negative for a particular phenotype, may include a combination of features that distinguish one subset of cells from another subset of cells depicted in the image.
[0023] In some exemplary embodiments, one or more subsets of cells present in an image may be identified by applying a biomarker identification model, including, for example, a machine learning-based biomarker identification model. For example, in some cases, the biomarker identification model may be applied to determine whether a cell is associated with multiple biomarkers, is positive for multiple biomarkers, or is negative for multiple biomarkers based at least on features associated with each cell in a population of cells depicted in the image. The output of the biomarker identification model may include a set of probabilities, each of which is the probability that a cell is associated with a corresponding biomarker, is positive for the corresponding biomarker, or is negative for the corresponding biomarker. The biomarker identification model may generate a probabilistic output instead of a binary output to provide a more accurate quantification of the error (or uncertainty) present in determining that an individual cell depicted in the image is positive (or negative) for a particular biomarker. For example, the binary output may include either a first value (e.g., “1”) indicating that the cell is positive for the biomarker, or a second value (e.g., “0”) indicating that the cell is negative for the biomarker even though there is uncertainty as to whether the cell is positive (or negative) for the biomarker. A probabilistic output, such as the probability that a cell is positive (or negative) for a biomarker, may capture the uncertainty involved in determining that a cell is positive (or negative) for a biomarker.
[0024] In some exemplary embodiments, one or more subsets of cells may be identified based at least on a set of probabilities associated with each cell depicted in the image. For example, in some cases, each subset of cells may correspond to one or more clusters of cells present in a reduced-dimensional representation of the dataset that includes a set of probabilities associated with each cell. Features associated with each subset of cells may be identified as indicating that the cells are associated with the corresponding phenotype, positive for the corresponding phenotype, or negative for the corresponding phenotype, or as indicating a probability that the cells are associated with the corresponding phenotype, positive for the corresponding phenotype, or negative for the corresponding phenotype. For example, a first set of features associated with a first subset of cells may be identified as indicative of the cell being associated with a first phenotype, being positive for the first phenotype, or being negative for the first phenotype (or the probability that the cell is associated with the first phenotype, being positive for the first phenotype, or being negative for the first phenotype), while a second set of features associated with a second subset of cells may be identified as indicative of the cell being positive for a second phenotype (or the probability that the cell is associated with the second phenotype, being positive for the second phenotype, or being negative for the second phenotype). In some cases, a first phenotypic discrimination model may be trained to determine a first probability that a cell is associated with a first phenotype, is positive for the first phenotype, or is negative for the first phenotype based on a first set of features associated with a first subset of cells, while a second phenotypic discrimination model may be trained to determine a second probability that a cell is associated with a second phenotype, is positive for the second phenotype, or is negative for the second phenotype based on a second set of features associated with a second subset of cells.
[0025] FIG. 1 depicts a system diagram illustrating an example digital pathology system 100, according to some exemplary embodiments. Referring to FIG. 1 , digital pathology system 100 may include digital pathology platform 110, imaging system 120, and client device 130. As shown in FIG. 1 , digital pathology platform 110, imaging system 120, and client device 130 may be communicatively coupled via network 140. Network 140 may be a wired and / or wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc. Imaging system 120 may include one or more imaging devices, including, for example, a microscope, a digital camera, a whole slide scanner, a robotic microscope, etc. Client device 130 may be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc.
[0026] Referring again to FIG. 1 , the digital pathology platform 110 may include a feature extractor 112, a controller 114, a biomarker identification model 116, and one or more phenotype identification models 118. As shown in FIG. 1 , the feature extractor 112 may extract multiple features associated with each cell in a population of cells from a first image 115 depicting the population of cells. In some cases, the first image 115 may be a stained whole slide image (WSI), including, for example, a hematoxylin and eosin (H&E)-stained whole slide image, a multiplex immunofluorescence (MxIF)-stained whole slide image, and / or the like. Further, in some cases, the feature extractor 112 may collect multiple features across multiple channels for each cell in the population of cells depicted in the first image 115. For example, in some cases, each channel (e.g., each individual feature) may correspond to one or more of the emission wavelengths of a fluorescent dye applied to the first image 115. Alternatively and / or additionally, each channel (e.g., each individual feature) may correspond to a metal ion collected by a mass cytometer, a nucleotide sequence identified by barcode hybridization, a nucleotide sequence identified by sequencing, etc.
[0027] In some exemplary embodiments, the controller 114 may apply the biomarker identification model 116 to determine, based at least on features associated with each cell depicted in the first image 115, whether the cell is associated with each of the plurality of biomarkers, is positive for each of the plurality of biomarkers, or is negative for each of the plurality of biomarkers. Example biomarkers may include Pax5, CD68, CD3, CD8, Foxp3, CD335, and Ki67. In some cases, the biomarker identification model 116 may be implemented using one or more machine learning models, including, for example, gradient-boosted decision trees, random forests, naive Bayes classifiers, neural networks, k-means clustering models, logistic regression models, and / or the like. Furthermore, the output of the biomarker identification model 116 may include, for each cell in the population of cells depicted in the first image 115, a set of probabilities, each of which is the probability that the cell is associated with the corresponding biomarker, is positive for the corresponding biomarker, or is negative for the corresponding biomarker. For example, n biomarkers b1, b2, …, b n , the output of the biomarker identification model 116 is a set of n probabilities P(b1), P(b2), ..., P(b n That is, the output of biomarker identification model 116 may include, for each cell in the population of cells depicted in first image 115, a set of probabilities including a first probability that the cell is associated with a first biomarker, is positive for the first biomarker, or is negative for the first biomarker, and a second probability that the cell is associated with a second biomarker, is positive for the second biomarker, or is negative for the second biomarker.
[0028] In some exemplary embodiments, controller 114 may identify one or more features that indicate a cell is associated with a particular phenotype, is positive for a particular phenotype, or is negative for a particular phenotype, or that indicate a probability that a cell is associated with a particular phenotype, is positive for a particular phenotype, or is negative for a particular phenotype. For example, in some cases, controller 114 may identify one or more subsets of cells present in the population of cells depicted in first image 115 based at least on a set of probabilities associated with each cell depicted in first image 115. Each of the subsets of cells identified within the population of cells depicted in first image 115 may correspond to a distinct phenotype. Furthermore, the set of features associated with each subset of cells identified within the population of cells depicted in first image 115 may indicate whether the cell is associated with the corresponding phenotype, is positive for the corresponding phenotype, or is negative for the corresponding phenotype, or may indicate a probability that the cell is associated with the corresponding phenotype, is positive for the corresponding phenotype, or is negative for the corresponding phenotype.
[0029] In some exemplary embodiments, the controller 114 may identify one or more subsets of cells present in the population of cells depicted in the first image 115 by at least generating a reduced-dimensional representation of a data set that includes a set of probabilities associated with each cell in the population of cells. For example, in some cases, the controller 114 may generate a reduced-dimensional representation of the data set by applying techniques such as t-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), principal component analysis (PCA), linear discriminant analysis (LDA), machine learning models, and the like. In some cases, the reduced-dimensional representation of the data set may occupy a lower-dimensional space (e.g., a space defined by fewer features) than the data set itself. For example, a data set that includes a set of probabilities associated with each cell in a population of cells may occupy an n-dimensional space in which each dimension (or feature) corresponds to the probability that a cell is associated with one of n biomarkers, is positive with respect to that biomarker, or is negative with respect to that biomarker. The reduced-dimensional representation of the data set may occupy an m-dimensional space where m < n (or m ≪ n).
[0030] In some exemplary embodiments, the reduced-dimensional representation of the dataset may occupy a two-dimensional (or three-dimensional) space that provides a visualization of the connections between individual cells depicted in the first image 115. For example, in the reduced-dimensional representation of the dataset, individual cells may be spatially distributed according to their respective probabilities of being associated with, being positive for, or being negative for each of n biomarkers. Cells that are associated with, being positive for, or exhibiting similar probabilities of being negative for similar combinations of biomarkers may be closely located within the m-dimensional space occupied by the reduced-dimensional representation of the dataset. For example, in some cases, cells that have similar probabilities of being associated with, being positive for, or being negative for similar combinations of biomarkers may form one or more cell clusters within the m-dimensional space occupied by the reduced-dimensional representation of the dataset. The controller 114 may identify each subset of cells to include one or more of the cell clusters present within the m-dimensional space occupied by the reduced-dimensional representation of the dataset.
[0031] In some exemplary embodiments, each subset of cells identified within the population of cells depicted in first image 115 may correspond to a particular phenotype. In some cases, the phenotype of a cell may correspond to a transient cellular state exhibited by the cell. Alternatively and / or additionally, other examples of phenotypes may include tumor cells, macrophages, regulatory T cells, CD8+ T cells, B cells, and natural killer (NK) cells. Controller 114 may identify a feature set associated with each subset of cells as indicating that the cell is associated with the corresponding phenotype, is positive for the corresponding phenotype, or is negative for the corresponding phenotype, or indicating a probability that the cell is associated with the corresponding phenotype, is positive for the corresponding phenotype, or is negative for the corresponding phenotype. For example, a first feature set associated with a first subset of cells may be identified as indicative of the cell being associated with a first phenotype, being positive for the first phenotype, or being negative for the first phenotype (or the probability that the cell is associated with the first phenotype, being positive for the first phenotype, or being negative for the first phenotype), while a second feature set associated with a second subset of cells may be identified as indicative of the cell being positive for a second phenotype (or the probability that the cell is associated with the second phenotype, being positive for the second phenotype, or being negative for the second phenotype). Further, in some cases, controller 114 may generate training data 117 for training one or more phenotype identification models 118 based at least on the first feature set and the second feature set.
[0032] In some exemplary embodiments, controller 114 may train one or more phenotype discrimination models 118 to determine whether cells depicted in second image 119 are associated with a phenotype, are positive for a phenotype, or are negative for a phenotype, based at least on a set of features associated with the phenotype. For example, controller 114 may train a first phenotype discrimination model 118a to determine whether cells depicted in second image 119 are associated with a first phenotype, are positive for a first phenotype, or are negative for a first phenotype, based at least on a first set of features extracted from second image 119. Further, in some cases, controller 114 may train a second phenotype discrimination model 118b to determine whether cells depicted in second image 119 are associated with a second phenotype, are positive for a second phenotype, or are negative for a second phenotype, based at least on a second set of features extracted from second image 119.
[0033] In some cases, the one or more phenotype identification models 118 may be implemented as one or more machine learning models including, for example, gradient-boosted decision trees, random forests, naive Bayes classifiers, neural networks, k-means clustering models, logistic regression models, and / or the like. Furthermore, while FIG. 1 depicts the first phenotype identification model 118a and the second phenotype identification model 118b being implemented using two separate machine learning models, in some cases, a single machine learning model may implement more than one of the one or more phenotype identification models 118 (e.g., the first phenotype identification model 118a and the second phenotype identification model 118b). Furthermore, in some cases, a single machine learning model may implement the biomarker identification model 116 and one or more phenotype identification models 118.
[0034] As mentioned above, in some exemplary embodiments, the output of the biomarker identification model 116 for each cell depicted in the first image 115 may be, for example, a set of n probabilities P(b1), P(b2), ..., P(b n), where each probability P(b i ) is the biomarker that the cell corresponds to b i The biomarker identification model 116 determines whether the individual cells depicted in the first image 115 are positive for each biomarker b, or negative for this corresponding biomarker b, associated with the corresponding biomarker b. i In some cases, the output of one or more phenotypic discrimination models 118 may also include the probability that a cell is associated with the corresponding phenotype, is positive for the corresponding phenotype, or is negative for the corresponding phenotype. It should be understood that one or more phenotypic discrimination models 118 may be trained to generate a probabilistic output instead of a binary output to provide a more accurate quantification of the error (or uncertainty) present in determining that a cell is positive (or negative) for a particular phenotype.
[0035] Figure 2 is a flowchart illustrating an example of a process 200 for probabilistic feature identification for machine learning-enabled cellular phenotyping, according to some exemplary embodiments. A corresponding workflow 300 for machine learning-enabled probabilistic feature identification for cellular phenotyping is shown in Figure 3. With reference to Figures 1-3, process 200 (and corresponding workflow 300) may be performed by digital pathology platform 110, which may include, for example, one or more of feature extractor 112, controller 114, biomarker identification model 116, and one or more phenotype identification models 118.
[0036] At 202, feature extractor 112 may extract multiple features for each cell in the population of cells from a first image 115 depicting the population of cells. To further illustrate, FIG. 4 depicts example features extracted from first image 115. As shown in FIG. 4, in some exemplary embodiments, feature extractor 112 may extract multiple features from first image 115, including, for example, one or more geometric features, statistical features, textural features, and / or the like. Furthermore, FIG. 4 illustrates that feature extractor 112 may extract features across multiple channels. For example, in some cases, each channel (or feature) may correspond to one or more of the following: an emission wavelength of a fluorescent dye applied to the image, a metal ion collected by a mass cytometer, a nucleotide sequence identified by barcode hybridization, a nucleotide sequence identified by sequencing, and / or the like. In the example shown in FIG. 4, a total of 197 features were extracted for each cell depicted in first image 115. However, it should be understood that the feature extractor 112 may extract a different amount of features from the first image 115 than the example shown in FIG.
[0037] At 204, controller 114 may apply biomarker identification model 116 to determine, based at least on a plurality of features associated with each cell in the population of cells, whether the cell is associated with a plurality of biomarkers, is positive for the plurality of biomarkers, or is negative for the plurality of biomarkers. In some exemplary embodiments, controller 114 may apply biomarker identification model 116 to determine, for each cell in the population of cells depicted in first image 115, whether the cell is associated with each of the plurality of biomarkers, is positive for each of the plurality of biomarkers, or is negative for each of the plurality of biomarkers. For example, in some cases, biomarker identification model 116 may be applied to determine, based on features extracted from first image 115, whether each individual cell depicted in first image 115 is associated with n biomarkers b1, b2, ..., b nAs mentioned, examples of biomarkers can include Pax5, CD68, CD3, CD8, Foxp3, CD335, and Ki67.
[0038] In some exemplary embodiments, to train the biomarker identification model 116, the training data 117 may include annotated training samples that include, for each cell in the population of cells depicted in the first image 115, a plurality of features associated with the cell and a ground truth label corresponding to each biomarker exhibited by the cell. In some cases, the ground truth labels assigned to the cells depicted in the first image 115 may be binary values, such as a first value (e.g., 1) indicating that the cell is associated with, positive for, or negative for a particular biomarker, and a second value (e.g., 0) indicating that the cell is negative for the biomarker. To further illustrate, FIG. 5A depicts a schematic diagram illustrating an example of a process 500 for training and validating the biomarker identification model 116, according to some exemplary embodiments. In the example shown in FIG. 5A, the biomarker identification model 116 calculates, for each region of interest (ROI) in the first image 115 corresponding to a cell, n probabilities P(b1), P(b2), ..., P(b n ), where each probability P(b i ) is the biomarker that the cell corresponds to b i is the probability of being positive for this corresponding biomarker or negative for this corresponding biomarker, associated with
[0039] At 206, the controller 114 may determine a set of probabilities for each cell in the population of cells based at least on the output of the biomarker identification model 116. In some exemplary embodiments, the output of the biomarker identification model 116 may be a set of n probabilities P(b1), P(b2), ..., P(b n), where each probability P(b i ) is the biomarker that the cell corresponds to b i The biomarker identification model 116 determines whether the individual cells depicted in the first image 115 are positive for each biomarker b, or negative for this corresponding biomarker b, associated with the corresponding biomarker b. i It may be trained to produce a probabilistic output instead of a binary output to provide a more accurate quantification of the error (or uncertainty) present in a decision that is positive (or negative) for .
[0040] At 208, the controller 114 may identify a first subset of cells exhibiting a first phenotype and a second subset of cells exhibiting a second phenotype based at least on a set of probabilities associated with each cell in the population of cells. In some exemplary embodiments, the controller 114 may identify n probabilities P(b1), P(b2), ..., P(b n ) to identify a first subset of cells exhibiting a first phenotype and a second subset of cells exhibiting a second phenotype. For example, in some cases, the controller 114 may generate the reduced dimensional representation of the dataset by applying one or more of t-distributed stochastic neighborhood embedding (t-SNE), uniform manifold approximation and projection (UMAP), principal component analysis (PCA), linear discriminant analysis (LDA), applying a machine learning model, etc.
[0041] FIG. 6A shows n probabilities P(b1), P(b2), ..., P(b n6A depicts a visualization of an example of a reduced-dimensional representation of a dataset including a set of n biomarkers. A dataset may occupy an n-dimensional space, where each dimension (or feature) corresponds to a probability that a cell is associated with one of the n biomarkers, is positive for that biomarker, or is negative for that biomarker. The example reduced-dimensional representation of the dataset shown in FIG. 6A may occupy a two-dimensional space, where individual cells are spatially distributed according to their respective probabilities of being associated with, being positive for, or being negative for each of the n biomarkers. For example, the visualization of the reduced-dimensional representation of the dataset shown in FIG. 6A includes multiple clusters of cells, each populated by cells that have a similar probability of being associated with, being positive for, or being negative for a similar combination of biomarkers. FIG. 6B illustrates how the cells within each cell cluster are spatially distributed within image 115. Thus, controller 114 may, for example, identify first cell subset 600 and second cell subset 650, each including one or more clusters of cells. As shown in FIG. 6B, some subsets of cells, such as first cell subset 600, may include a single cell cluster, while some subsets of cells, such as second cell subset 650, may include multiple cell clusters. Furthermore, in some cases, controller 114 may identify first cell subset 600 as being associated with a first phenotype and second cell subset 650 as being associated with a second phenotype.
[0042] At 210, controller 114 may identify a first feature set associated with a first subset of cells as indicating a first probability that the cell is associated with a first phenotype, positive for the first phenotype, or negative for the first phenotype, and may identify a second feature set associated with a second subset of cells as indicating a second probability that the cell is associated with a second phenotype, positive for the second phenotype, or negative for the second phenotype. For example, in the example shown in FIGS. 6A-6B , a first feature set associated with first cell subset 600 may be identified as indicating a first probability that the cell is associated with the first phenotype, positive for the first phenotype, or negative for the first phenotype, and a second feature set associated with second cell subset 650 may be identified as indicating a second probability that the cell is associated with the second phenotype, positive for the second phenotype, or negative for the second phenotype.
[0043] At 212, the controller 114 may train a first phenotype discrimination model 118a to determine a first probability that a cell is associated with a first phenotype, is positive for the first phenotype, or is negative for the first phenotype based on a first feature set associated with the first subset of cells. To further explain, FIG. 5B depicts a schematic diagram showing an example of a process 550 for training and validating one or more phenotype discrimination models 118. For example, in some cases, the controller 114 may train the first phenotype discrimination model 118a to determine a first probability that a cell is associated with a first phenotype, is positive for the first phenotype, or is negative for the first phenotype based at least on the training data 117. The training data 117 in this case may include annotated training samples that include, for each cell in the first cell subset 600, a first feature set associated with the first phenotype and a ground truth label corresponding to the cell's first phenotype. For example, in some cases, the ground truth label assigned to a cell may be a binary value, such as a first value (e.g., 1) indicating that the cell is associated with, positive for, or negative for the first phenotype, and a second value (e.g., 0) indicating that the cell is negative for the first phenotype. Thus, the first phenotype discrimination model 118a may be trained to learn the link between the first feature set and the first phenotype such that the first phenotype discrimination model 118a may determine a first probability that a cell is associated with, positive for, or negative for the first phenotype based on whether the cell exhibits one or more of the features included in the first feature set.
[0044] At 214, the controller 114 may train a second phenotype discrimination model 118b to determine a second probability that a cell is associated with, positive for, or negative for a second phenotype based at least on a second set of features associated with the second subset of cells. In some exemplary embodiments, one or more phenotype discrimination models 118 may be phenotype-specific, such that the controller 114 may train a separate phenotype discrimination model 118 for each possible phenotype. Thus, in some cases, the controller 114 may further train a second phenotype discrimination model 118b based at least on the training data 117 to determine a second probability that a cell is associated with, positive for, or negative for a second phenotype. In this case, the training data 117 may include annotated training samples that include, for each cell in the second cell subset 650, a second feature set associated with a second phenotype and a ground truth label corresponding to the cell's second phenotype. In some cases, the ground truth label assigned to the cell may be a binary value, such as a first value (e.g., 1) indicating that the cell is associated with the second phenotype, positive for the second phenotype, or negative for the second phenotype, and a second value (e.g., 0) indicating that the cell is negative for the second phenotype. From the output when trained, the second phenotype discrimination model 118b may recognize a connection between the second feature set and the second phenotype, and thereby the second phenotype discrimination model 118b may determine a second probability that the cell is associated with the second phenotype, positive for the second phenotype, or negative for the second phenotype based on whether the cell exhibits one or more features included in the second feature set.
[0045] At 216, the controller 114 may apply the first phenotypic discrimination model 118a and / or the second phenotypic discrimination model 118b to determine a phenotype of one or more cells depicted in the second image 119. In some exemplary embodiments, the controller 114 may apply the first phenotypic discrimination model 118a to determine a first probability that one or more cells depicted in the second image 119 are associated with a first phenotype, are positive for the first phenotype, or are negative for the first phenotype, based at least on a first set of features extracted from the second image 119. Further, the controller 114 may apply the second phenotypic discrimination model 118b to determine a second probability that one or more cells depicted in the second image 119 are associated with a second phenotype, are positive for the second phenotype, or are negative for the second phenotype, based at least on a second set of features extracted from the second image 119.
[0046] In some exemplary embodiments, one or more downstream tasks, such as determining a disease diagnosis, disease progression, disease burden, and / or treatment response for a patient associated with the second image 119, may be performed based on a first probability that one or more cells depicted in the second image 119 are associated with a first phenotype, are positive for the first phenotype, or are negative for the first phenotype, and / or a second probability that one or more cells depicted in the second image 119 are associated with a second phenotype, are positive for the second phenotype, or are negative for the second phenotype. Alternatively, in some cases, one or more downstream tasks may be performed when the first probability that one or more cells are associated with the first phenotype, are positive for the first phenotype, or are negative for the first phenotype, and / or the second probability that one or more cells are associated with the second phenotype, are positive for the second phenotype, or are negative for the second phenotype, meet one or more thresholds. For example, if a first probability that one or more cells are associated with a first phenotype, are positive for the first phenotype, or are negative for the first phenotype, and / or a second probability that one or more cells are associated with a second phenotype, are positive for the second phenotype, or are negative for the second phenotype, exceeds one or more thresholds, the presence (or absence) of cells having the first phenotype and / or the second phenotype may be used to determine a disease diagnosis, disease progression, disease burden, and / or treatment response of a patient associated with the second image 119.
[0047] 7 depicts a block diagram of an example computing system 700, according to some exemplary embodiments. Referring to FIGS. 1 and 7, computing system 700 may be used to implement digital pathology platform 110, imaging system 120, client device 130, and / or any components therein.
[0048] 7, computing system 700 may include a processor 710, a memory 720, a storage device 730, and an input / output device 740. The processor 710, the memory 720, the storage device 730, and the input / output device 740 may be interconnected via a system bus 750. The processor 710 may process instructions for execution within the computing system 700. Such executed instructions may implement one or more components, such as, for example, the digital pathology platform 110, the imaging system 120, the client device 130, etc. In some exemplary embodiments, the processor 710 may be a single-threaded processor. Alternatively, the processor 710 may be a multi-threaded processor. The processor 710 may process instructions stored in the memory 720 and / or the storage device 730 to display graphical information for a user interface provided via the input / output device 740.
[0049] Memory 720 is a computer-readable medium, such as a volatile or non-volatile medium, that stores information within computing system 700. Memory 720 may store, for example, data structures representing a configuration object database. Storage device 730 may provide persistent storage for computing system 700. Storage device 730 may be a floppy disk drive, a hard disk drive, an optical disk drive, or a tape drive, or other suitable persistent storage means. Input / output device 740 provides input / output operations to computing system 700. In some exemplary embodiments, input / output device 740 includes a keyboard and / or a pointing device. In various embodiments, input / output device 740 includes a display unit for displaying a graphical user interface.
[0050] According to some exemplary embodiments, input / output devices 740 may provide input / output operations to network devices. For example, input / output devices 740 may include an Ethernet port or other networking port for communicating with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).
[0051] In some exemplary embodiments, computing system 700 may be used to execute various interactive computer software applications that may be used to organize, analyze, and / or store data in various formats. Alternatively, computing system 700 may be used to execute any type of software application. These applications may be used to perform various functions, such as planning functions (e.g., creating, managing, editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, etc. Applications may include various add-in functions or may be standalone computing products and / or functions. When activated within an application, functions may be used to generate a user interface that is provided via input / output devices 740. The user interface may be generated by computing system 700 and presented to a user (e.g., on a computer screen monitor, etc.).
[0052] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs, field programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0053] These computer programs, which may also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or in an assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as, for example, magnetic disks, optical disks, memories, and programmable logic devices (PLDs), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium may non-transitory store such machine instructions, such as, for example, a non-transitory solid-state memory, a magnetic hard drive, or any equivalent storage medium. Alternatively or additionally, a machine-readable medium may temporarily store such machine instructions, such as, for example, a processor cache or other random access memory associated with one or more physical processor cores.
[0054] To provide for user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having, for example, a display device, such as a cathode ray tube (CRT) or liquid crystal display (LCD) or light-emitting diode (LED) monitor, for displaying information to a user, and a keyboard and pointing device, such as a mouse or trackball, through which the user may provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software.
[0055] In the above description and in the claims, phrases such as "at least one of" or "one or more of" may appear followed by a conjunctive list of elements or features. The term "and / or" may also be used in listings of two or more elements or features. Unless otherwise implicitly or explicitly stated by the context in which it is used, such phrases are intended to refer to any of the listed elements or features individually, or any of the listed elements or features in combination with any of the other listed elements or features. For example, the phrases "at least one of A and B," "one or more of A and B," and "A and / or B" are intended to mean "A only, B only, or A and B together," respectively. A similar interpretation is intended for lists containing more than two items. For example, the phrases "at least one of A, B, and C;," "one or more of A, B, and C;," and "A, B, and / or C" are intended to mean "A alone, B alone, C alone, A and B, A and C, B and C, or A, B, and C," respectively. Use of the term "based on" above and in the claims means "based at least in part on," and implies that unrecited features or elements are also permitted.
[0056] Illustrative Embodiments Embodiments disclosed herein may include the following. 1. A computer-implemented method comprising: extracting a plurality of features for each cell in the population of cells from a first image depicting the population of cells; applying a biomarker discrimination model to determine whether a cell is associated with, positive for, or negative for a plurality of biomarkers based at least on a plurality of features associated with each cell in the population of cells; determining a set of probabilities for each cell in the population of cells based at least on the output of the biomarker discrimination model, and for each biomarker in the plurality of biomarkers, determining a set of probabilities comprising a probability that a corresponding cell is associated with the biomarker; identifying a first subset of cells exhibiting a first phenotype based at least on a set of probabilities associated with each cell in the population of cells; identifying a first set of features associated with a first subset of cells as indicative of a first probability that the cells are associated with a first phenotype; 10. A computer-implemented method comprising: 2. The method of embodiment 1, wherein the plurality of features includes one or more of geometric features, statistical features, and textural features. 3. The method of embodiment 1 or 2, wherein a plurality of features are collected across a plurality of channels, each channel of the plurality of channels corresponding to (i) an emission wavelength of a fluorescent dye applied to the first image, (ii) a metal ion collected by the mass cytometer, (iii) a nucleotide sequence identified by barcode hybridization, or (iv) a nucleotide sequence identified by sequencing. 4. The method of any one of embodiments 1 to 3, wherein each biomarker in the plurality of biomarkers corresponds to a protein of interest or an antigen comprising one or more carbohydrates, lipids, or nucleotides. 5. The method of any one of embodiments 1 to 4, wherein each biomarker in the plurality of biomarkers corresponds to a protein expressed by a population of cells. 6. Training a first phenotype discrimination model to determine a first probability that a cell is associated with a first phenotype based on a first set of features associated with a first subset of cells; 6. The method of any one of embodiments 1 to 5, further comprising: 7. applying a first phenotype discrimination model to determine a probability that one or more cells depicted in the second image are associated with the first phenotype based on a first set of features extracted from the second image; 7. The method of embodiment 6, further comprising: 8. Determining a disease diagnosis, disease progression, disease burden, and / or treatment response based at least on the probability that one or more cells depicted in the second image are associated with the first phenotype; 8. The method of embodiment 7, further comprising: 9. Identifying a second subset of cells exhibiting a second phenotype based at least on the set of probabilities associated with each cell in the population of cells; training a second phenotype discrimination model to determine a second probability that a cell is associated with a second phenotype based on a second set of features associated with a second subset of cells; 7. The method of embodiment 6, further comprising: 10. The method of any one of embodiments 1-9, wherein the set of probabilities for each cell in the population of cells comprises a first probability that the cell is associated with a first biomarker in the plurality of biomarkers and a second probability that the cell is associated with a second biomarker in the plurality of biomarkers. 11. The method of any one of embodiments 1 to 10, wherein the first subset of cells is identified by generating a reduced-dimensional representation of a dataset comprising a set of probabilities associated with each cell in the population of cells. 12. The method of embodiment 11, wherein the reduced-dimensional representation of the dataset is generated by applying t-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), principal component analysis (PCA), linear discriminant analysis (LDA), and / or a machine learning model. 13. The method of embodiment 11, wherein the reduced-dimensional representation of the dataset comprises a plurality of cell clusters, each cell cluster of the plurality of cell clusters comprising one or more cells having the same or similar phenotype. 14. The method of embodiment 11, wherein the first subset of cells comprises one or more cell clusters from the plurality of cell clusters. 15. The method of any one of embodiments 1 to 14, wherein the biomarker discrimination model and / or the first phenotype discrimination model comprises a gradient boosted decision tree, a random forest, a naive Bayes classifier, a neural network, a k-means clustering model, or a logistic regression model. 16. Training a biomarker identification model to determine, based at least on a plurality of features extracted from the first image, whether one or more cells depicted in the first image are associated with each biomarker in the plurality of biomarkers; 16. The method of any one of embodiments 1 to 15, further comprising: 17. For each cell in the population of cells, generating an annotated training sample including a plurality of features associated with the cell and a ground truth label corresponding to each biomarker exhibited by the cell; training a biomarker identification model based at least on the plurality of annotated training samples; 17. The method of embodiment 16, further comprising: 18. For each cell in the first subset of cells, generating an annotated training sample including a first feature set and a ground truth label corresponding to a first phenotype of the cell; training a first phenotype discrimination model based at least on the plurality of annotated training samples; 18. The method of any one of embodiments 1 to 17, further comprising: 19. The method of any one of embodiments 1 to 18, wherein the population of cells is part of a biological sample or a derivative of a biological sample. 20. The method of any one of embodiments 1-19, wherein the population of cells comprises a tissue fragment and / or a body fluid. 21. The method of any one of embodiments 1-20, wherein the first image comprises at least a portion of the entire slide image. 22. The method of any one of embodiments 1-21, wherein the plurality of biomarkers comprises Pax5, CD68, CD3, CD8, Foxp3, CD335 and / or Ki67. 23. The method of any one of embodiments 1 to 22, wherein the first phenotype is a transient cellular state exhibited by a first subset of cells. 24. The method of any one of embodiments 1 to 23, wherein the first phenotype is a tumor cell, a macrophage, a regulatory T cell, a CD8-positive T cell, a B cell, or a natural killer (NK) cell. 25. The method of any one of embodiments 1-24, wherein the first image is an entire slide image. 26. The method of any one of embodiments 1 to 25, wherein the first image is a hematoxylin and eosin (H&E) stained image or a multiplex immunofluorescence (MxIF) stained image. 27. at least one data processor; At least one memory storing instructions that, when executed by at least one data processor, result in operations including the method of any one of embodiments 1 to 26; A system comprising: 28. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, result in operations including the method described in any one of embodiments 1-26.
[0057] The subject matter described herein may be embodied in systems, devices, methods, and / or articles, depending on the desired configuration. The implementations set forth in the above description do not represent all implementations of the subject matter described herein. Instead, these implementations are merely some examples consistent with aspects associated with the described subject matter. While some variations have been described in detail above, other modifications or additions are possible. In particular, additional features and / or variations may be provided in addition to those described herein. For example, the implementations described above may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of certain additional features described above. Furthermore, the logic flow illustrated in the accompanying drawings and / or described herein does not necessarily require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims.
Claims
1. extracting a plurality of features for each cell in a population of cells from a first image depicting the population of cells; applying a biomarker discrimination model to determine whether the cell is associated with a plurality of biomarkers, is positive for a plurality of biomarkers, or is negative for a plurality of biomarkers based at least on the plurality of features associated with each cell in the population of cells; determining a set of probabilities for each cell in the population of cells based at least on an output of the biomarker discrimination model, the set of probabilities comprising, for each biomarker in the plurality of biomarkers, a probability that a corresponding cell is associated with the biomarker; identifying a first subset of cells exhibiting a first phenotype based at least on the set of probabilities associated with each cell in the population of cells; identifying a first set of features associated with a first subset of the cells as indicative of a first probability that the cells are associated with the first phenotype; 10. A computer-implemented method comprising:
2. The method of claim 1 , wherein the plurality of features comprises one or more geometric features, statistical features, and textural features.
3. 3. The method of claim 1 or 2, wherein the plurality of features are collected across a plurality of channels, each channel of the plurality of channels corresponding to (i) an emission wavelength of a fluorescent dye applied to the first image, (ii) a metal ion collected by a mass cytometer, (iii) a nucleotide sequence identified by barcode hybridization, or (iv) a nucleotide sequence identified by sequencing.
4. 4. The method of claim 1, wherein each biomarker in the plurality of biomarkers corresponds to a protein of interest or an antigen comprising one or more carbohydrates, lipids or nucleotides.
5. The method of any one of claims 1 to 4, wherein each biomarker in the plurality of biomarkers corresponds to a protein expressed by the population of cells.
6. training a first phenotype discrimination model to determine the first probability that the cells are associated with the first phenotype based on the first set of features associated with a first subset of the cells; The method of any one of claims 1 to 5, further comprising:
7. applying the first phenotype discrimination model to determine a probability that one or more cells depicted in the second image are associated with the first phenotype based on the first set of features extracted from the second image; The method of claim 6 further comprising:
8. determining a disease diagnosis, disease progression, disease burden, and / or treatment response based at least on the probability that the one or more cells depicted in the second image are associated with the first phenotype; The method of claim 7 further comprising:
9. identifying a second subset of cells exhibiting a second phenotype based at least on a set of probabilities associated with each cell in the population of cells; and training a second phenotype discrimination model to determine a second probability that the cells are associated with the second phenotype based on a second set of features associated with a second subset of the cells; The method of claim 6 further comprising:
10. 10. The method of any one of claims 1 to 9, wherein the set of probabilities for each cell in the population of cells comprises a first probability that the cell is associated with a first biomarker in the plurality of biomarkers and a second probability that the cell is associated with a second biomarker in the plurality of biomarkers.
11. 11. The method of any one of claims 1 to 10, wherein the first subset of cells is identified by generating a reduced dimensional representation of a dataset comprising the set of probabilities associated with each cell in the population of cells.
12. 12. The method of claim 11 , wherein the reduced-dimensional representation of the dataset is generated by applying t-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), principal component analysis (PCA), linear discriminant analysis (LDA), and / or a machine learning model.
13. 12. The method of claim 11, wherein the reduced dimensional representation of the dataset comprises a plurality of cell clusters, each cell cluster of the plurality of cell clusters comprising one or more cells having the same or similar phenotype.
14. 12. The method of claim 11, wherein the first subset of cells comprises one or more cell clusters from a plurality of cell clusters.
15. 15. The method of any one of claims 1 to 14, wherein the biomarker discrimination model and / or first phenotype discrimination model comprises a gradient boosted decision tree, a random forest, a naive Bayes classifier, a neural network, a k-means clustering model, or a logistic regression model.
16. training the biomarker identification model to determine whether one or more cells depicted in the first image are associated with each biomarker in the plurality of biomarkers based at least on the plurality of features extracted from the first image; The method of any one of claims 1 to 15, further comprising:
17. generating, for each cell in the population of cells, an annotated training sample comprising the plurality of features associated with the cell and a ground truth label corresponding to each biomarker exhibited by the cell; training the biomarker discrimination model based at least on a plurality of annotated training samples; 17. The method of claim 16, further comprising:
18. generating, for each cell in the first subset of cells, an annotated training sample comprising the first feature set and a ground truth label corresponding to the first phenotype of the cell; training a first phenotype discrimination model based at least on the plurality of annotated training samples; The method of any one of claims 1 to 17, further comprising:
19. The method of any one of claims 1 to 18, wherein the population of cells is part of a biological sample or a derivative of said biological sample.
20. The method of any one of claims 1 to 19, wherein the population of cells comprises a tissue fragment and / or a body fluid.
21. The method of any preceding claim, wherein the first image comprises at least a portion of an entire slide image.
22. 22. The method of any one of claims 1 to 21, wherein the plurality of biomarkers comprises Pax5, CD68, CD3, CD8, Foxp3, CD335 and / or Ki67.
23. 23. The method of any one of claims 1 to 22, wherein the first phenotype is a transient cellular state exhibited by the first subset of cells.
24. 24. The method of any one of claims 1 to 23, wherein the first phenotype is a tumor cell, a macrophage, a regulatory T cell, a CD8-positive T cell, a B cell, or a natural killer (NK) cell.
25. The method of any one of claims 1 to 24, wherein the first image is an entire slide image.
26. The method of any one of claims 1 to 25, wherein the first image is a hematoxylin and eosin (H&E) stained image or a multiplex immunofluorescence (MxIF) stained image.
27. at least one data processor; at least one memory storing instructions which, when executed by said at least one data processor, result in operations comprising the method of any one of claims 1 to 26; A system comprising:
28. A non-transitory computer readable medium storing instructions which, when executed by at least one data processor, result in operations comprising the method of any one of claims 1 to 26.