Probabilistic identification of features for machine learning supported cellular phenotypes
By applying machine learning models to generate probability outputs in cell images, quantifying the correlation between cells and biomarkers and phenotypes, the diagnosis inaccuracy caused by pathologist variability in the prior art is solved, and more accurate cell phenotype recognition is achieved.
Patent Information
- Application Number
- CN202380083012.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-02
- Filing Date
- 2023-10-31
- Publication Date
- 2025-07-22
AI Technical Summary
Existing machine learning-based cell phenotype recognition techniques are susceptible to variability among and within pathologists, resulting in inaccurate diagnosis and existing methods that are difficult to quantify cell uncertainty about biomarkers and phenotypes.
Using machine learning-based biomarker recognition model and phenotype recognition model, probability output is generated to quantify the correlation between cells and biomarkers and phenotypes. By extracting multiple features from the images and applying machine learning algorithms such as gradient enhancement decision trees, random forests, etc., the probability of cell subsets and phenotypes is identified.
Provides more accurate cell phenotype recognition, quantifies cell uncertainty about biomarkers and phenotypes, and improves diagnostic accuracy and consistency.
Smart Images

Figure CN120359551A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 382,075, filed on November 2, 2022, entitled "Probabilistic Identification of Features for Machine Learning Enabled Cellular Phenotyping", the disclosure of which is hereby incorporated by reference in its entirety. Technical Field
[0003] The subject matter described herein generally relates to digital and computational pathology, and more particularly to probabilistic methods for identifying features for the determination of machine - learning - enabled cellular phenotypes. Background Art
[0004] The phenotype of a cell can refer to a unique combination of morphological and functional characteristics produced by various cellular processes, including, for example, gene expression, protein expression, etc. In some cases, the complex interactions between the cell genome, epigenome, and the local environment may give rise to a range of observable features, collectively referred to as the cell phenotype. Although cell phenotypes (including those of tumor cells) are often attributed to genomic instability, there has recently been increasing attention paid to epigenetic and microenvironmental influences. Such non - genetic factors can further increase the intrinsic diversity and plasticity of tumor cells. At the tumor level, non - genetic factors may lead to greater phenotypic heterogeneity, enabling tumor cells to evade immune responses and resist drug interventions. Summary of the Invention
[0005] The present disclosure provides systems, methods, and articles of manufacture, including computer program products, for the probabilistic identification of features of machine learning-supported cell phenotypes. In one aspect, a system for the probabilistic identification of features of machine learning-supported cell phenotypes is provided. The system can include at least one processor and at least one memory. The at least one memory can include program code that, when executed by the at least one processor, provides operations. The operations can include: for each cell in a cell population, extracting a plurality of features from a first image depicting the cell population; applying a biomarker identification model to determine whether the cell is associated with, positive for, or negative for a plurality of biomarkers, based at least on the plurality of features associated with each cell in the cell population; determining, for each cell in the cell population, a probability set based at least on the output of the biomarker identification model, and for each of the plurality of biomarkers, the probability set includes the probability that the corresponding cell is associated with, positive for, or negative for the biomarker; identifying a first subset of cells exhibiting a first phenotype based at least on the probability set associated with each cell in the cell population; and identifying a first set of features associated with the first subset of cells as a first probability indicating that the cells are associated with, positive for, or negative for the first phenotype.
[0006] In another aspect, a method for the probabilistic identification of features of machine learning-supported cell phenotypes is provided. The method can include: for each cell in a cell population, extracting a plurality of features from a first image depicting the cell population; applying a biomarker identification model to determine whether the cell is associated with, positive for, or negative for a plurality of biomarkers, based at least on the plurality of features associated with each cell in the cell population; determining, for each cell in the cell population, a probability set based at least on the output of the biomarker identification model, and for each of the plurality of biomarkers, the probability set includes the probability that the corresponding cell is associated with, positive for, or negative for the biomarker; identifying a first subset of cells exhibiting a first phenotype based at least on the probability set associated with each cell in the cell population; and identifying a first set of features associated with the first subset of cells as a first probability indicating that the cells are associated with, positive for, or negative for the first phenotype.
[0007] In another aspect, a computer program product is provided for the probabilistic identification of features of a machine learning-supported cell phenotype. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations. The operations may include: for each cell in a cell population, extracting a plurality of features from a first image depicting the cell population; applying a biomarker identification model to determine whether the cell is associated with, positive for, or negative for a plurality of biomarkers based at least on the plurality of features associated with each cell in the cell population; determining, for each cell in the cell population, a probability set based at least on the output of the biomarker identification model, and for each of the plurality of biomarkers, the probability set includes the probability that the corresponding cell is associated with, positive for, or negative for the biomarker; identifying a first subset of cells exhibiting a first phenotype based at least on the probability set associated with each cell in the cell population; and identifying a first set of features associated with the first subset of cells as a first probability indicating that the cell is associated with, positive for, or negative for the first phenotype.
[0008] Specific implementations of the present subject matter may include, but are not limited to, methods consistent with the description provided herein and articles of manufacture including tangible embodied machine-readable media that are operable to cause one or more machines (e.g., computers, etc.) to cause operations implementing one or more of the described features. Similarly, a computer system may be described that may include one or more processors and one or more memories coupled to the one or more processors. The memory, which may include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. A computer-implemented method consistent with one or more implementations of the present subject matter may be implemented by one or more data processors present in a single computing system or multiple computing systems. Such multiple computing systems may be connected and may exchange data and / or commands or other instructions, etc., via one or more connections, including, for example, via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.) via a direct connection between one or more of the multiple computing systems, etc. to the connection.
[0009] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will become apparent from the description and drawings, and from the claims. While certain features of the presently disclosed subject matter are described for illustrative purposes in connection with the identification of features for machine learning supported cell phenotypes, it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed embodiments. In the drawings,
[0011] Figure 1 a system diagram depicting an example of a digital pathology system in accordance with some exemplary embodiments is shown;
[0012] Figure 2 a flowchart depicting an example of a process for probabilistic feature identification of machine learning supported cell phenotypes in accordance with some exemplary embodiments is shown;
[0013] Figure 3 a schematic diagram depicting an example of a workflow for probabilistic identification of features of machine learning supported cell phenotypes in accordance with some exemplary embodiments is shown;
[0014] Figure 4 examples of features extracted from an image depicting a cell population in accordance with some exemplary embodiments are shown;
[0015] Figure 5A a schematic diagram depicting an example of a process for training and validating a biomarker identification model in accordance with some exemplary embodiments is shown;
[0016] Figure 5B a schematic diagram depicting an example of a process for training and validating a phenotype identification model in accordance with some exemplary embodiments is shown;
[0017] Figure 6A a visualization of an example of a dimensionality reduced representation of a biomarker probability dataset in accordance with some exemplary embodiments is shown;
[0018] Figure 6B a further visualization of an example of a dimensionality reduced representation of a biomarker probability dataset in accordance with some exemplary embodiments is shown; and
[0019] Figure 7 a block diagram depicting an example of a computing system in accordance with some exemplary embodiments is shown.
[0020] In actual applications, like reference numerals denote like structures, features, or elements. Detailed implementation manners
[0021] In highly heterogeneous diseases such as cancer, insights into the cell phenotypes that form the diseased tissue and the surrounding microenvironment may be indispensable for accurately diagnosing disease subtypes, predicting disease progression, and predicting responses to various treatments. For example, patients with non-Hodgkin lymphoma at high risk of disease progression using standard-of-care treatments (e.g., combination immunochemotherapy R-CHOP (rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone)) can be identified by characterizing the immune microenvironment, which includes identifying lymph node-resident immune cells and infiltrating tumor cells. However, conventional histological analysis techniques for identifying cell phenotypes depicted in microscopic images (e.g., whole-slide images stained with hematoxylin and eosin (H&E), whole-slide images stained with multiplex immunofluorescence (MxIF), etc.) are error-prone due to the high level of variability between and within pathologists. At the same time, existing machine learning-based solutions are also vulnerable to variability between and within pathologists, at least because the training of machine learning-based cell phenotype models relies on expert annotations of training samples, but due to the existence of variability between and within pathologists, the expert annotations are not sufficiently reliable.
[0022] In some exemplary embodiments, a machine learning-based phenotype recognition model can be trained to determine the probability that one or more cells depicted in an image are associated with a specific phenotype, are positive for a specific phenotype, or are negative for a specific phenotype based on one or more features extracted from the image depicting the cell population. The machine learning-based phenotype recognition model can be trained to generate a probability output instead of a binary output in order to provide a more precise quantification of the error (or uncertainty) present in determining that one or more cells depicted in the image are positive (or negative) for a specific phenotype. In some cases, when the probability that one or more cells depicted in the image are positive for a specific phenotype meets one or more thresholds, one or more downstream tasks can be performed, such as determining disease diagnosis, disease progression, disease burden, and / or treatment response. Additionally, in some cases, the same machine learning-based phenotype recognition model or one or more independent machine learning-based phenotype recognition models can be trained to determine the probability that one or more cells depicted in the image are positive for another phenotype.
[0023] In some exemplary embodiments, one or more features present in an image can be identified as indicating that a cell is associated with a particular phenotype, positive for a particular phenotype, or negative for a particular phenotype, or the probability that a cell is associated with a particular phenotype, positive for a particular phenotype, or negative for a particular phenotype. For example, in some cases, multiple features can be extracted from an image depicting a cell population, which image includes, for example, a whole-slide image stained with hematoxylin and eosin (H&E), a whole-slide image stained with multiplex immunofluorescence (MxIF), etc. These features can be collected through multiple channels. For instance, in some cases, each channel can correspond to one or more of the emission wavelengths of fluorescent dyes applied to the image, metal ions collected by mass cytometry, nucleotide sequences identified by barcode hybridization, nucleotide sequences identified by sequencing, etc. One or more features indicating that a cell is associated with a particular phenotype, positive for a particular phenotype, or negative for a particular phenotype, or the probability that a cell is associated with a particular phenotype, positive for a particular phenotype, or negative for a particular phenotype can include a combination of features that distinguish one subset of cells from another subset of cells depicted in the image.
[0024] In some exemplary embodiments, one or more subsets of cells present in an image can be identified by applying a biomarker identification model that includes, for example, a machine learning-based biomarker identification model. For example, in some cases, a biomarker identification model can be applied to determine whether a cell is associated with multiple biomarkers, positive for multiple biomarkers, or negative for multiple biomarkers based at least on the features associated with each cell in the cell population depicted in the image. The output of the biomarker identification model can include a set of probabilities, each of which is the probability that a cell is associated with the corresponding biomarker, positive for the corresponding biomarker, or negative for the corresponding biomarker. The biomarker identification model can generate a probability output rather than a binary output in order to provide a more precise quantification of the error (or uncertainty) present in determining that an individual cell depicted in the image is positive (or negative) for a particular biomarker. For example, a binary output can include a first value (e.g., "1") indicating that a cell is positive for a biomarker or a second value (e.g., "0") indicating that a cell is negative for a biomarker, even if there is uncertainty as to whether the cell is positive (or negative) for the biomarker. A probability output, such as the probability that a cell is positive (or negative) for a biomarker, can capture the uncertainty included in determining that a cell is positive (or negative) for a biomarker.
[0025] In some exemplary embodiments, one or more subsets of cells can be identified based at least on a set of probabilities associated with each cell depicted in an image. For example, in some cases, each subset of cells can correspond to one or more cell clusters present in a reduced-dimensional representation of a data set that includes the set of probabilities associated with each cell. The features associated with each subset of cells can be identified as indicating that the cell is associated with a corresponding phenotype, positive for the corresponding phenotype, or negative for the corresponding phenotype, or the probability that the cell is associated with the corresponding phenotype, positive for the corresponding phenotype, or negative for the corresponding phenotype. By way of example, a first set of features associated with a first subset of cells can be identified as indicating that the cell is associated with a first phenotype, positive for the first phenotype, or negative for the first phenotype (or the probability that the cell is associated with the first phenotype, positive for the first phenotype, or negative for the first phenotype), while a second set of features associated with a second subset of cells can be identified as indicating that the cell is positive for a second phenotype (or the probability that the cell is associated with the second phenotype, positive for the second phenotype, or negative for the second phenotype). In some cases, a first phenotype recognition model can be trained to determine a first probability that a cell is associated with a first phenotype, positive for the first phenotype, or negative for the first phenotype based on the first set of features associated with the first subset of cells, while a second phenotype recognition model can be trained to determine a second probability that a cell is associated with a second phenotype, positive for the second phenotype, or negative for the second phenotype based on the second set of features associated with the second subset of cells.
[0026] Figure 1 FIG. depicts a system diagram showing an example of a digital pathology system 100 according to some exemplary embodiments. Referring Figure 1 to, the digital pathology system 100 can include a digital pathology platform 110, an imaging system 120, and a client device 130. As Figure 1 shown, the digital pathology platform 110, the imaging system 120, and the client device 130 can be communicatively coupled via a network 140. The network 140 can be a wired network and / or a wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc. The imaging system 120 can include one or more imaging devices (including, for example, a microscope, a digital camera, a whole slide scanner, an automated microscope, etc.). The client device 130 can be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc.
[0027] Referring again Figure 1 to, the digital pathology platform 110 can include a feature extractor 112, a controller 114, a biomarker recognition model 116, and one or more phenotype recognition models 118. As Figure 1As shown, the feature extractor 112 can extract a plurality of features associated with each cell in the cell population from the first image 115 depicting the cell population. In some cases, the first image 115 can be a stained whole slide image (WSI), including, for example, a hematoxylin and eosin (H&E) stained whole slide image, a multiplex immunofluorescence (MxIF) stained whole slide image, and the like. Additionally, in some cases, the feature extractor 112 can collect a plurality of features through a plurality of channels for each cell in the cell population depicted in the first image 115. For example, in some cases, each channel (e.g., each individual feature) can correspond to one or more of the emission wavelengths of the fluorescent dyes applied to the first image 115. Alternatively and / or additionally, each channel (e.g., each individual feature) can correspond to metal ions collected by mass cytometry, nucleotide sequences identified by barcode hybridization, nucleotide sequences identified by sequencing, and the like.
[0028] In some exemplary embodiments, the controller 114 can apply the biomarker recognition model 116 to determine whether each cell depicted in the first image 115 is associated with, positive for, or negative for each of a plurality of biomarkers, at least based on the features associated with each cell. Examples of biomarkers can include Pax5, CD68, CD3, CD8, Foxp3, CD335, and Ki67. In some cases, the biomarker recognition model 116 can be implemented using one or more machine learning models, including, for example, gradient boosting decision trees, random forests, naive Bayes classifiers, neural networks, k-means clustering models, logistic regression models, and the like. Additionally, for each cell in the cell population depicted in the first image 115, the output of the biomarker recognition model 116 can include a set of probabilities, each of which is the probability that the cell is associated with, positive for, or negative for the corresponding biomarker. For example, for a set of n biomarkers b1, b2, …, b n the output of the biomarker recognition model 116 for each cell in the cell population depicted in the first image 115 can include a set of n probabilities P(b1), P(b2), …, P(b n )). That is, for each cell in the cell population depicted in the first image 115, the output of the biomarker recognition model 116 can include a set of probabilities that includes a first probability that the cell is associated with, positive for, or negative for the first biomarker and a second probability that the cell is associated with, positive for, or negative for the second biomarker.
[0029] In some exemplary embodiments, the controller 114 may identify one or more features indicative that a cell is associated with a particular phenotype, positive for a particular phenotype, or negative for a particular phenotype, or the probability that a cell is associated with a particular phenotype, positive for a particular phenotype, or negative for a particular phenotype. For example, in some cases, the controller 114 may identify one or more cell subsets present in the cell population depicted in the first image 115 based at least on a set of probabilities associated with each cell depicted in the first image 115. Each of the cell subsets identified within the cell population depicted in the first image 115 may correspond to a separate phenotype. Additionally, the set of features associated with each cell subset identified within the cell population depicted in the first image 115 may indicate whether the cell is associated with the corresponding phenotype, positive for the corresponding phenotype, or negative for the corresponding phenotype, or the probability that the cell is associated with the corresponding phenotype, positive for the corresponding phenotype, or negative for the corresponding phenotype.
[0030] In some exemplary embodiments, the controller 114 may identify one or more cell subsets present in the cell population depicted in the first image 115 by at least generating a reduced-dimensional representation of a data set that includes a set of probabilities associated with each cell in the cell population. For example, in some cases, the controller 114 may generate a reduced-dimensional representation of the data set by applying t-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), principal component analysis (PCA), linear discriminant analysis (LDA), a machine learning model, and the like. In some cases, the reduced-dimensional representation of the data set may occupy a space of lower dimension than the data set itself (e.g., a space defined by fewer features). By way of example, a data set that includes a set of probabilities associated with each cell in the cell population may occupy an n-dimensional space, where each dimension (or feature) corresponds to the probability that a cell is associated with, positive for, or negative for one of an n-number of biomarkers. The reduced-dimensional representation of the data set may occupy an m-dimensional space, where m < n (or m << n).
[0031] In some exemplary embodiments, a dimensionality-reduced representation of a data set can occupy a two-dimensional (or three-dimensional) space that provides visualization of the connections between individual cells depicted in the first image 115. For example, in the dimensionality-reduced representation of the data set, individual cells can be spatially distributed according to their respective probabilities of being associated with, positive for, or negative for each of an n-number of biomarkers. Cells that exhibit similar probabilities of being associated with, positive for, or negative for a similar combination of biomarkers can be located close to one another in the m-dimensional space occupied by the dimensionality-reduced representation of the data set. For instance, in some cases, cells that have similar probabilities of being associated with, positive for, or negative for a similar combination of biomarkers can form one or more cell clusters in the m-dimensional space occupied by the dimensionality-reduced representation of the data set. The controller 114 can identify each subset of cells to include one or more cell clusters present in the m-dimensional space occupied by the dimensionality-reduced representation of the data set.
[0032] In some exemplary embodiments, each subset of cells identified within the cell population depicted in the first image 115 can correspond to a particular phenotype. In some cases, the phenotype of a cell can correspond to a transient cell state exhibited by the cell. Alternatively and / or additionally, other examples of phenotypes can include tumor cells, macrophages, regulatory T cells, CD8-positive T cells, B cells, and natural killer (NK) cells. The controller 114 can identify a set of features associated with each subset of cells as indicative of the cell being associated with, positive for, or negative for the corresponding phenotype, or the probability that the cell is associated with, positive for, or negative for the corresponding phenotype. For example, a first set of features associated with a first subset of cells can be identified as indicative of the cell being associated with, positive for, or negative for a first phenotype (or the probability that the cell is associated with, positive for, or negative for the first phenotype), while a second set of features associated with a second subset of cells can be identified as indicative of the cell being positive for a second phenotype (or the probability that the cell is associated with, positive for, or negative for the second phenotype). Additionally, in some cases, the controller 114 can generate training data 117 for training one or more phenotype recognition models 118 based at least on the first set of features and the second set of features.
[0033] In some exemplary embodiments, the controller 114 may train one or more phenotypic recognition models 118 to determine whether the cells depicted in the second image 119 are associated with, positive for, or negative for a phenotype, based at least on a set of features associated with the phenotype. For example, the controller 114 may train a first phenotypic recognition model 118a to determine whether the cells depicted in the second image 119 are associated with, positive for, or negative for a first phenotype, based at least on a first set of features extracted from the second image 119. Additionally, in some cases, the controller 114 may train a second phenotypic recognition model 118b to determine whether the cells depicted in the second image 119 are associated with, positive for, or negative for a second phenotype, based at least on a second set of features extracted from the second image 119.
[0034] In some cases, one or more phenotypic recognition models 118 may be implemented as one or more machine learning models, including, for example, gradient boosting decision trees, random forests, naive Bayes classifiers, neural networks, k-means clustering models, logistic regression models, and the like. Additionally, although Figure 1 the first phenotypic recognition model 118a and the second phenotypic recognition model 118b are depicted as being implemented using two separate machine learning models, in some cases, a single machine learning model may implement multiple ones of the one or more phenotypic recognition models 118 (e.g., the first phenotypic recognition model 118a and the second phenotypic recognition model 118b). Additionally, in some cases, a single machine learning model may also implement the biomarker recognition model 116 and one or more of the phenotypic recognition models 118.
[0035] As described above, in some exemplary embodiments, the output of the biomarker recognition model 116 for each cell depicted in the first image 115 may include a set of probabilities, such as, for example, a set of n probabilities P(b1), P(b2), …, P(b n ), where each probability P(b i ) is the probability that the cell is associated with, positive for, or negative for the corresponding biomarker b i . The biomarker recognition model 116 may be trained to generate a probability output rather than a binary output in order to provide a more precise quantification of the determination that an individual cell depicted in the first image 115 is positive for each biomarker b iThe error (or uncertainty) present in being positive (or negative). In some cases, the output of one or more phenotype recognition models 118 may also include the probability that a cell is associated with a corresponding phenotype, is positive for the corresponding phenotype, or is negative for the corresponding phenotype. It should be understood that one or more phenotype recognition models 118 can also be trained to generate a probability output instead of a binary output in order to provide a more precise quantification of the error (or uncertainty) present in determining that a cell is positive (or negative) for a particular phenotype.
[0036] Figure 2 FIG. shows a flowchart of an example of a process 200 for probability-based feature recognition of machine learning-supported cell phenotypes according to some exemplary embodiments. Figure 3 A corresponding workflow 300 for probability-based feature recognition of machine learning-supported cell phenotypes is shown. Referring Figures 1 to 3 to, the process 200 (and the corresponding workflow 300) can be performed by a digital pathology platform 110, including, for example, by one or more of a feature extractor 112, a controller 114, a biomarker recognition model 116, and one or more phenotype recognition models 118.
[0037] At 202, the feature extractor 112 can extract a plurality of features from a first image 115 depicting a cell population, for each cell in the cell population. For further illustration, Figure 4 FIG. shows an example of features extracted from the first image 115. As Figure 4 shown, in some exemplary embodiments, the feature extractor 112 can extract a plurality of features from the first image 115, including, for example, one or more geometric features, statistical features, texture features, etc. Additionally, Figure 4 FIG. shows that the feature extractor 112 can extract features through multiple channels. For example, in some cases, each channel (or feature) can correspond to one or more of the emission wavelengths of fluorescent dyes applied to the image, metal ions collected by mass cytometry, nucleotide sequences identified by barcode hybridization, nucleotide sequences identified by sequencing, etc. In the Figure 4 example shown, a total of 197 features are extracted for each cell depicted in the first image 115. However, it should be understood that the feature extractor 112 can extract a different number of features from the first image 115 than the Figure 4 example shown.
[0038] At 204, the controller 114 may apply the biomarker identification model 116 to determine whether a cell is associated with, positive for, or negative for multiple biomarkers based at least on a plurality of features associated with each cell in the cell population. In some exemplary embodiments, for each cell in the cell population depicted in the first image 115, the controller 114 may apply the biomarker identification model 116 to determine whether the cell is associated with, positive for, or negative for each of the multiple biomarkers. For example, in some cases, the biomarker identification model 116 may be applied to determine, based on features extracted from the first image 115, the probability that an individual cell depicted in the first image 115 is positive for a set of n biomarkers b1, b2, …, b n . As described above, examples of biomarkers may include Pax5, CD68, CD3, CD8, Foxp3, CD335, and Ki67.
[0039] In some exemplary embodiments, to train the biomarker identification model 116, for each cell in the cell population depicted in the first image 115, the training data 117 may include annotated training samples that include a plurality of features associated with the cell and ground truth labels corresponding to each biomarker exhibited by the cell. In some cases, the ground truth labels assigned to the cells depicted in the first image 115 may be binary values, such as, for example, a first value (e.g., 1) indicating that the cell is associated with, positive for, or negative for a particular biomarker and a second value (e.g., 0) indicating that the cell is negative for the biomarker. For further illustration, Figure 5A FIG. shows a schematic diagram illustrating an example of a process 500 for training and validating the biomarker identification model 116 according to some exemplary embodiments. In Figure 5A the example shown, for each target region (ROI) in the first image 115 corresponding to a cell, the biomarker identification model 116 may be trained to determine a set of n probabilities P(b1), P(b2), …, P(b n ), where each probability P(b i ) is the probability that the cell is associated with, positive for, or negative for the corresponding biomarker b i .
[0040] At 206, the controller 114 can determine a set of probabilities for each cell in the cell population based at least on the output of the biomarker identification model 116. In some exemplary embodiments, for each cell in the cell population depicted in the first image 115, the output of the biomarker identification model 116 can include a set of n probabilities P(b1), P(b2), …, P(b n ), where each probability P(b i ) is the probability that the cell is associated with, positive for, or negative for the corresponding biomarker b i . The biomarker identification model 116 can be trained to generate a probability output rather than a binary output in order to provide a more precise quantification of the error (or uncertainty) present in determining that an individual cell depicted in the first image 115 is positive (or negative) for each biomarker b i .
[0041] At 208, the controller 114 can identify a first subset of cells exhibiting a first phenotype and a second subset of cells exhibiting a second phenotype based at least on the set of probabilities associated with each cell in the cell population. In some exemplary embodiments, the controller 114 can identify the first subset of cells exhibiting the first phenotype and the second subset of cells exhibiting the second phenotype by at least generating a reduced-dimensional representation of the data set that includes the set of n probabilities P(b1), P(b2), …, P(b n ) associated with each cell in the cell population depicted in the image 115. For example, in some cases, the controller 114 can generate a reduced-dimensional representation of the data set by applying one or more of the following: applying t-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), principal component analysis (PCA), linear discriminant analysis (LDA), machine learning models, etc.
[0042] Figure 6A depicts a visualization of an example of a reduced-dimensional representation of a data set that includes the set of n probabilities P(b1), P(b2), …, P(b n ). Although the data set can occupy an n-dimensional space in which each dimension (or feature) corresponds to the probability that a cell is associated with, positive for, or negative for one of the n-number of biomarkers, the Figure 6A example of the reduced-dimensional representation of the data set shown can occupy a two-dimensional space in which individual cells are spatially distributed according to their respective probabilities of being associated with, positive for, or negative for each of the n-number of biomarkers. For example, Figure 6AVisualization of the dimensionality-reduced representation of the data set shown includes multiple cell clusters, where each cell cluster is populated by cells having a similar probability associated with a similar combination of biomarkers, positive for a similar combination of biomarkers, or negative for a similar combination of biomarkers. Figure 6B Shows how the cells in each cell cluster are spatially distributed in image 115. Thus, the controller 114 can identify, for example, a first cell subset 600 and a second cell subset 650, each containing one or more cell clusters. As Figure 6B shown, some cell subsets (such as the first cell subset 600) can include a single cell cluster, while some cell subsets (such as the second cell subset 650) can include multiple cell clusters. Additionally, in some cases, the controller 114 can identify the first cell subset 600 as being associated with a first phenotype and the second cell subset 650 as being associated with a second phenotype.
[0043] At 210, the controller 114 can identify a first set of features associated with the first cell subset as a first probability indicating that the cells are associated with the first phenotype, positive for the first phenotype, or negative for the first phenotype, and a second set of features associated with the second cell subset as a second probability indicating that the cells are associated with the second phenotype, positive for the second phenotype, or negative for the second phenotype. By way of example, in Figures 6A to 6B the example shown, the first set of features associated with the first cell subset 600 can be identified as a first probability indicating that the cells are associated with the first phenotype, positive for the first phenotype, or negative for the first phenotype, and the second set of features associated with the second cell subset 650 can be identified as a second probability indicating that the cells are associated with the second phenotype, positive for the second phenotype, or negative for the second phenotype.
[0044] At 212, the controller 114 can train a first phenotype recognition model 118a to determine, based on the first set of features associated with the first cell subset, a first probability that the cells are associated with the first phenotype, positive for the first phenotype, or negative for the first phenotype. For further illustration, Figure 5BFIG. 550 depicts an example of a process for training and validating one or more phenotypic recognition models 118. For example, in some cases, the controller 114 may train a first phenotypic recognition model 118a based at least on training data 117 to determine a first probability that a cell is associated with, positive for, or negative for a first phenotype. For each cell in the first cell subset 600, the training data 117 in this case may include an annotated training sample that includes a first set of features associated with the first phenotype and a ground truth label corresponding to the first phenotype of the cell. For example, in some cases, the ground truth label assigned to a cell may be a binary value, such as a first value (e.g., 1) indicating that the cell is associated with, positive for, or negative for the first phenotype and a second value (e.g., 0) indicating that the cell is negative for the first phenotype. Thus, the first phenotypic recognition model 118a may be trained to learn the association between the first set of features and the first phenotype such that the first phenotypic recognition model 118a can determine a first probability that a cell is associated with, positive for, or negative for the first phenotype based on whether the cell exhibits one or more features included in the first set of features.
[0045] At 214, the controller 114 may train a second phenotypic recognition model 118b to determine a second probability that a cell is associated with, positive for, or negative for a second phenotype based at least on a second set of features associated with a second cell subset. In some exemplary embodiments, one or more phenotypic recognition models 118 may be phenotype-specific such that the controller 114 may train a separate phenotypic recognition model 118 for each possible phenotype. Thus, in some cases, the controller 114 may further train the second phenotypic recognition model 118b based at least on the training data 117 to determine a second probability that a cell is associated with, positive for, or negative for the second phenotype. For each cell in the second cell subset 650, the training data 117 in this case may include an annotated training sample that includes a second set of features associated with the second phenotype and a ground truth label corresponding to the second phenotype of the cell. In some cases, the ground truth label assigned to a cell may be a binary value, such as for example, a first value (e.g., 1) indicating that the cell is associated with, positive for, or negative for the second phenotype and a second value (e.g., 0) indicating that the cell is negative for the second phenotype. As a result of the training output, the second phenotypic recognition model 118b may identify the association between the second set of features and the second phenotype such that the second phenotypic recognition model 118b can determine a second probability that a cell is associated with, positive for, or negative for the second phenotype based on whether the cell exhibits one or more features included in the second set of features.
[0046] At 216, the controller 114 may apply the first phenotype recognition model 118a and / or the second phenotype recognition model 118b to determine the phenotype of one or more cells depicted in the second image 119. In some exemplary embodiments, the controller 114 may apply the first phenotype recognition model 118a to determine a first probability that one or more cells depicted in the second image 119 are associated with, positive for, or negative for the first phenotype, based at least on a first set of features extracted from the second image 119. Additionally, the controller 114 may apply the second phenotype recognition model 118b to determine a second probability that one or more cells depicted in the second image 119 are associated with, positive for, or negative for the second phenotype, based at least on a second set of features extracted from the second image 119.
[0047] In some exemplary embodiments, one or more downstream tasks, such as determining a disease diagnosis, disease progression, disease burden, and / or treatment response of a patient associated with the second image 119, may be performed based on the first probability that one or more cells depicted in the second image 119 are associated with, positive for, or negative for the first phenotype and / or the second probability that one or more cells depicted in the second image 119 are associated with, positive for, or negative for the second phenotype. Alternatively, in some cases, one or more downstream tasks may be performed when the first probability that one or more cells are associated with, positive for, or negative for the first phenotype and / or the second probability that one or more cells are associated with, positive for, or negative for the second phenotype meet one or more thresholds. For example, in a case where the first probability that one or more cells are associated with, positive for, or negative for the first phenotype and / or the second probability that one or more cells are associated with, positive for, or negative for the second phenotype exceed one or more thresholds, the presence (or absence) of cells having the first phenotype and / or the second phenotype may be used to determine a disease diagnosis, disease progression, disease burden, and / or treatment response of a patient associated with the second image 119.
[0048] Figure 7 A block diagram depicting an example of a computing system 700 according to some exemplary embodiments is shown. Referring Figure 1 and Figure 7 , the computing system 700 may be used to implement the digital pathology platform 110, the imaging system 120, the client device 130, and / or any components thereof.
[0049] As Figure 7As shown, the computing system 700 may include a processor 710, a memory 720, a storage device 730, and an input / output device 740. The processor 710, the memory 720, the storage device 730, and the input / output device 740 may be interconnected via a system bus 750. The processor 710 is capable of processing instructions for execution within the computing system 700. Such executed instructions may implement one or more components of, for example, the digital pathology platform 110, the imaging system 120, the client device 130, and the like. In some exemplary embodiments, the processor 710 may be a single-threaded processor. Alternatively, the processor 710 may be a multi-threaded processor. The processor 710 is capable of processing instructions stored in the memory 720 and / or the storage device 730 to display graphical information for a user interface provided via the input / output device 740.
[0050] The memory 720 is a computer-readable medium for storing information within the computing system 700, such as a volatile or non-volatile computer-readable medium. For example, the memory 720 may store a data structure representing a configuration object database. The storage device 730 is capable of providing persistent storage for the computing system 700. The storage device 730 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device or other suitable persistent storage means. The input / output device 740 provides input / output operations for the computing system 700. In some exemplary embodiments, the input / output device 740 includes a keyboard and / or a pointing device. In various embodiments, the input / output device 740 includes a display unit for displaying a graphical user interface.
[0051] According to some exemplary embodiments, the input / output device 740 may provide input / output operations for a network device. For example, the input / output device 740 may include an Ethernet port or other networking port to communicate with one or more wired and / or wireless networks (e.g., local area network (LAN), wide area network (WAN), Internet).
[0052] In some exemplary embodiments, the computing system 700 may be used to execute various interactive computer software applications that may be used to organize, analyze, and / or store data in various formats. Alternatively, the computing system 700 may be used to execute any type of software application. These applications may be used to perform various functions, such as, for example, scheduling functions (e.g., generating, managing, editing spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functions, communication functions, and the like. The applications may include various additional features or may be stand-alone computing products and / or functions. After being activated within the application, the functions may be used to generate a user interface provided via the input / output device 740. The user interface may be generated by the computing system 700 and presented to the user (e.g., on a computer screen monitor, etc.).
[0053] One or more aspects or features of the subject matter described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor (which can be special purpose or general purpose and is coupled to receive data and instructions from, and to send data and instructions to, a storage system, at least one input device, and at least one output device). The programmable system or computing system can include a client and a server. Typically, the client and the server are remotely located from each other and generally interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and the client-server relationship between them.
[0054] These computer programs may also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer product, apparatus, and / or device (such as, for example, a disk, optical disk, memory, and programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor. The machine-readable medium can non-transitorily store such machine instructions (such as, for example, in non-transitory solid state memory or a magnetic hard disk drive or any equivalent storage medium). The machine-readable medium can alternatively or additionally store such machine instructions in a transitory manner (such as, for example, in a processor cache or other random access memory associated with one or more physical processor cores).
[0055] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device (such as, for example, a cathode ray tube (CRT), or a liquid crystal display (LCD), or a light emitting diode (LED) monitor for displaying information to the user) and a keyboard and a pointing device (such as, for example, a mouse or a trackball by which the user can provide input to the computer). Other kinds of devices can also be used to provide for interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback; and the input from the user can be received in any form, including sound, voice, or tactile input. Other possible input devices include a touch screen or other touch-sensitive devices, such as a single-point or multi-point resistive or capacitive trackpad, speech recognition hardware and software, an optical scanner, an optical indicator, a digital image capture device, and associated interpretation software, etc.
[0056] In the foregoing description and claims, phrases such as "at least one" or "one or more" may appear, followed by a list of elements or features. The term "and / or" may also appear in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, the phrase is intended to mean any of the elements or features listed individually, or any other element or feature in combination with any other recited element or feature. For example, the phrases "at least one of A and B"; "one or more of A and B"; "A and / or B" are each intended to mean "A alone, B alone, or A and B together". Similar interpretations apply to lists that include three or more items. For example, the phrases "at least one of A, B, and C"; "one or more of A, B, and C" and "A, B, and / or C" are each intended to mean "A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together". The use of the term "based on" in the foregoing and claims is intended to mean "at least in part based on", such that an unrecited feature or element is also permissible.
[0057] Exemplary embodiments
[0058] The embodiments disclosed herein can include:
[0059] 1. A computer-implemented method, comprising:
[0060] For each cell in a cell population, extracting a plurality of features from a first image depicting the cell population;
[0061] Apply a biomarker identification model to determine whether a cell is associated with, positive for, or negative for multiple biomarkers, based at least on a plurality of features associated with each cell in a cell population;
[0062] Determine a probability set for each cell in the cell population, based at least on the output of the biomarker identification model, and for each of the multiple biomarkers, the probability set includes the probability that the corresponding cell is associated with the biomarker;
[0063] Identify a first subset of cells exhibiting a first phenotype, based at least on the probability set associated with each cell in the cell population; and
[0064] Identify a first set of features associated with the first subset of cells as a first probability indicating that the cell is associated with the first phenotype.
[0065] 2. The method according to embodiment 1, wherein the plurality of features includes one or more geometric features, statistical features, and texture features.
[0066] 3. The method according to embodiment 1 or embodiment 2, wherein the plurality of features is collected through a plurality of channels, and each of the plurality of channels corresponds to (i) the emission wavelength of a fluorescent dye applied to a first image, (ii) a metal ion collected by mass cytometry, (iii) a nucleotide sequence identified by barcode hybridization, or (iv) a nucleotide sequence identified by sequencing.
[0067] 4. The method according to any one of embodiments 1 to 3, wherein each of the multiple biomarkers corresponds to a target protein or an antigen comprising one or more carbohydrates, lipids, or nucleotides.
[0068] 5. The method according to any one of embodiments 1 to 4, wherein each of the multiple biomarkers corresponds to a protein expressed by the cell population.
[0069] 6. The method according to any one of embodiments 1 to 5, further comprising:
[0070] Train a first phenotype identification model to determine a first probability that a cell is associated with the first phenotype, based on the first set of features associated with the first subset of cells.
[0071] 7. The method according to embodiment 6, further comprising:
[0072] Apply the first phenotype identification model to determine the probability that one or more cells depicted in a second image are associated with the first phenotype, based on the first set of features extracted from the second image.
[0073] 8. The method according to embodiment 7, further comprising:
[0074] Determining a disease diagnosis, disease progression, disease burden, and / or treatment response based at least on the probability that one or more cells depicted in the second image are associated with a first phenotype.
[0075] 9. The method according to embodiment 6, further comprising:
[0076] Identifying a second subset of cells exhibiting a second phenotype based at least on a set of probabilities associated with each cell in the cell population; and
[0077] Training a second phenotype recognition model to determine a second probability that a cell is associated with the second phenotype based on a second set of features associated with the second subset of cells.
[0078] 10. The method according to any one of embodiments 1 to 9, wherein the set of probabilities for each cell in the cell population includes a first probability that the cell is associated with a first biomarker among a plurality of biomarkers and a second probability that the cell is associated with a second biomarker among the plurality of biomarkers.
[0079] 11. The method according to any one of embodiments 1 to 10, wherein the first subset of cells is identified by generating a reduced-dimensional representation of a data set that includes the set of probabilities associated with each cell in the cell population.
[0080] 12. The method according to embodiment 11, wherein the reduced-dimensional representation of the data set is generated by applying t-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), principal component analysis (PCA), linear discriminant analysis (LDA), and / or a machine learning model.
[0081] 13. The method according to embodiment 11, wherein the reduced-dimensional representation of the data set includes a plurality of cell clusters, and wherein each cell cluster among the plurality of cell clusters includes one or more cells having the same or similar phenotype.
[0082] 14. The method according to embodiment 11, wherein the first subset of cells includes one or more cell clusters from the plurality of cell clusters.
[0083] 15. The method according to any one of embodiments 1 to 14, wherein the biomarker recognition model and / or the first phenotype recognition model includes a gradient boosting decision tree, a random forest, a naive Bayes classifier, a neural network, a k-means clustering model, or a logistic regression model.
[0084] 16. The method according to any one of embodiments 1 to 15, further comprising:
[0085] Train a biomarker recognition model to determine whether one or more cells depicted in a first image are associated with each of a plurality of biomarkers based at least on a plurality of features extracted from the first image.
[0086] 17. The method according to embodiment 16, further comprising:
[0087] For each cell in a cell population, generate an annotated training sample including a plurality of features associated with the cell and a ground truth label corresponding to each biomarker exhibited by the cell; and
[0088] Train the biomarker recognition model based at least on the plurality of annotated training samples.
[0089] 18. The method according to any one of embodiments 1 to 17, further comprising:
[0090] For each cell in a first subset of cells, generate an annotated training sample including a first set of features and a ground truth label corresponding to a first phenotype of the cell; and
[0091] Train a first phenotype recognition model based at least on the plurality of annotated training samples.
[0092] 19. The method according to any one of embodiments 1 to 18, wherein the cell population is part of a biological sample or a derivative of a biological sample.
[0093] 20. The method according to any one of embodiments 1 to 19, wherein the cell population comprises tissue fragments and / or body fluids.
[0094] 21. The method according to any one of embodiments 1 to 20, wherein the first image comprises at least a portion of a whole slide image.
[0095] 22. The method according to any one of embodiments 1 to 21, wherein the plurality of biomarkers comprises Pax5, CD68, CD3, CD8, Foxp3, CD335, and / or Ki67.
[0096] 23. The method according to any one of embodiments 1 to 22, wherein the first phenotype is a transient cell state exhibited by the first subset of cells.
[0097] 24. The method according to any one of embodiments 1 to 23, wherein the first phenotype is a tumor cell, a macrophage, a regulatory T cell, a CD8-positive T cell, a B cell, or a natural killer (NK) cell.
[0098] 25. The method according to any one of embodiments 1 to 24, wherein the first image is a whole slide image.
[0099] 26. The method according to any one of embodiments 1 to 25, wherein the first image is a hematoxylin and eosin (H&E) stained image or a multiplex immunofluorescence (MxIF) stained image.
[0100] 27. A system comprising:
[0101] at least one data processor; and
[0102] at least one memory storing instructions that, when executed by the at least one data processor, cause operations including the method according to any one of embodiments 1 to 26.
[0103] 28. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations including the method according to any one of embodiments 1 to 26.
[0104] Depending on the desired configuration, the subject matter described herein can be embodied in a system, apparatus, method, and / or article of manufacture. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Instead, they are only some examples consistent with aspects related to the subject matter described. Although some variations have been described in detail above, other modifications or additions are possible. In particular, other features and / or variations can be provided in addition to those features and / or variations set forth herein. For example, the above specific embodiments can be directed to various combinations and sub-combinations of the disclosed features and / or to combinations and sub-combinations of several further features disclosed above. Additionally, the logical flows depicted in the figures and / or described herein need not be in the particular order or sequential order shown to achieve the desired result. Other specific embodiments can be within the scope of the following claims.
Claims
1. A computer-implemented method, comprising: For each cell in a cell population, extracting a plurality of features from a first image depicting the cell population; Applying a biomarker identification model to determine whether the cell is associated with a plurality of biomarkers, positive for the plurality of biomarkers, or negative for the plurality of biomarkers, based at least on the plurality of features associated with each cell in the cell population; Determining a probability set for each cell in the cell population, based at least on the output of the biomarker identification model, and for each biomarker of the plurality of biomarkers, the probability set includes the probability that the corresponding cell is associated with the biomarker; Identifying a first subset of cells exhibiting a first phenotype, based at least on the probability set associated with each cell in the cell population; And Identifying a first set of features associated with the first subset of cells as a first probability indicating that the cells are associated with the first phenotype.
2. The method according to claim 1, wherein the plurality of features includes one or more geometric features, statistical features, and texture features.
3. The method according to claim 1 or claim 2, wherein the plurality of features are collected through a plurality of channels, and each of the plurality of channels corresponds to (i) the emission wavelength of a fluorescent dye applied to the first image, (ii) metal ions collected by mass cytometry, (iii) nucleotide sequences identified by barcode hybridization, or (iv) nucleotide sequences identified by sequencing.
4. The method according to any one of claims 1 to 3, wherein each biomarker of the plurality of biomarkers corresponds to a target protein or an antigen comprising one or more carbohydrates, lipids, or nucleotides.
5. The method according to any one of claims 1 to 4, wherein each biomarker of the plurality of biomarkers corresponds to a protein expressed by the cell population.
6. The method according to any one of claims 1 to 5, further comprising: Training a first phenotype identification model to determine the first probability that the cells are associated with the first phenotype, based on the first set of features associated with the first subset of cells.
7. The method according to claim 6, further comprising: Applying the first phenotype identification model to determine the probability that one or more cells depicted in a second image are associated with the first phenotype, based on the first set of features extracted from the second image.
8. The method according to claim 7, further comprising: Determining a disease diagnosis, disease progression, disease burden, and / or treatment response, based at least on the probability that one or more cells depicted in the second image are associated with the first phenotype.
9. The method according to claim 6, further comprising: Identifying a second subset of cells exhibiting a second phenotype, based at least on the probability set associated with each cell in the cell population; And Train a second phenotype recognition model to determine a second probability that a cell is associated with the second phenotype based on a second set of features associated with the second subset of cells.
10. The method according to any one of claims 1 to 9, wherein the set of probabilities for each cell in the cell population includes a first probability that the cell is associated with a first biomarker among the plurality of biomarkers and a second probability that the cell is associated with a second biomarker among the plurality of biomarkers.
11. The method according to any one of claims 1 to 10, wherein the first subset of cells is identified by generating a reduced-dimensional representation of a data set that includes the set of probabilities associated with each cell in the cell population.
12. The method according to claim 11, wherein the reduced-dimensional representation of the data set is generated by applying t-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), principal component analysis (PCA), linear discriminant analysis (LDA), and / or a machine learning model.
13. The method according to claim 11, wherein the reduced-dimensional representation of the data set includes a plurality of cell clusters, and wherein each cell cluster in the plurality of cell clusters includes one or more cells having the same or similar phenotype.
14. The method according to claim 11, wherein the first subset of cells includes one or more cell clusters from the plurality of cell clusters.
15. The method according to any one of claims 1 to 14, wherein the biomarker recognition model and / or the first phenotype recognition model includes a gradient boosting decision tree, a random forest, a naive Bayes classifier, a neural network, a k-means clustering model, or a logistic regression model.
16. The method according to any one of claims 1 to 15, further comprising: training the biomarker recognition model to determine whether one or more cells depicted in the first image are associated with each of the plurality of biomarkers based at least on the plurality of features extracted from the first image.
17. The method according to claim 16, further comprising: for each cell in the cell population, generating an annotated training sample that includes the plurality of features associated with the cell and a ground truth label corresponding to each biomarker exhibited by the cell; and training the biomarker recognition model based at least on the plurality of annotated training samples.
18. The method according to any one of claims 1 to 17, further comprising: for each cell in the first subset of cells, generating an annotated training sample that includes the first set of features and a ground truth label corresponding to the first phenotype of the cell; and training the first phenotype recognition model based at least on the plurality of annotated training samples.
19. The method according to any one of claims 1 to 18, wherein the cell population is part of a biological sample or a derivative of the biological sample.
20. The method according to any one of claims 1 to 19, wherein the cell population comprises tissue fragments and / or body fluids.
21. The method according to any one of claims 1 to 20, wherein the first image comprises at least a portion of a whole slide image.
22. The method according to any one of claims 1 to 21, wherein the plurality of biomarkers comprises Pax5, CD68, CD3, CD8, Foxp3, CD335, and / or Ki67.
23. The method according to any one of claims 1 to 22, wherein the first phenotype is a transient cell state exhibited by the first subset of cells.
24. The method according to any one of claims 1 to 23, wherein the first phenotype is a tumor cell, macrophage, regulatory T cell, CD8-positive T cell, B cell, or natural killer (NK) cell.
25. The method according to any one of claims 1 to 24, wherein the first image is a whole slide image.
26. The method according to any one of claims 1 to 25, wherein the first image is a hematoxylin and eosin (H&E) stained image or a multiplex immunofluorescence (MxIF) stained image.
27. A system, comprising: at least one data processor; and at least one memory storing instructions that, when executed by the at least one data processor, cause an operation comprising the method according to any one of claims 1 to 26.
28. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause an operation comprising the method according to any one of claims 1 to 26.