Methods and systems for multiplex analysis
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- VIEWSML TECHNOLOGIES INC
- Filing Date
- 2024-07-23
- Publication Date
- 2026-06-03
AI Technical Summary
Multiplexing techniques face challenges in designing and optimizing experiments due to signal interference and difficulty in distinguishing multiple parameters, leading to high costs and complex data analysis.
A computer-implemented method for multiplex image analysis of cells using trained machine learning models to classify or score cells based on predefined criteria, generating visualizations that distinguish each criterion, and performing flow cytometry analysis without physical assays.
The method enables efficient and comprehensive multiplex image analysis, reducing costs and improving data interpretation by distinguishing multiple parameters and generating actionable visualizations.
Smart Images

Figure IB2024000395_30012025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR MULTIPLEX ANALYSISCROSS-REFERENCE
[0001] This application claims the benefit of Provisional Application No. 63 / 528,496, filed July 24, 2023, which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Multiplexing may allow the analysis of multiple parameters or molecules (e.g., of cells or tissues) in a single experiment, thereby enabling highly efficient and comprehensive data collection and analysis.SUMMARY
[0003] Multiplexing may refer to a technique used in various fields, such as biology, engineering, telecommunications, and data science, to simultaneously analyze multiple samples (e.g., cell or tissue samples) or data points. It may be a powerful technique for analyzing multiple parameters simultaneously; however, it can also be challenging to design and optimize multiplexing experiments, as the signals from multiple parameters may interfere with each other or be difficult to distinguish.
[0004] Multiplexing allows for the simultaneous analysis of multiple targets, generating images that can provide a more comprehensive view of biological processes and interactions than may be possible with single-target analysis. However, large amounts of complex data including images can be challenging to analyze and interpret. Multiplexing may also encounter various challenges such as a high cost, for example, requiring about $500K in capital equipment purchases that allow for fluorescently labeled probes to multiplex.
[0005] Recognizing the above needs, the present disclosure provides improved methods and systems for performing multiplex image analysis of cells.
[0006] In an aspect, the present disclosure provides a computer-implemented method of performing multiplex image analysis of a plurality of cells, comprising: (a) obtaining an image of the plurality of cells; (b) obtaining a plurality of trained machine learning models, each of the plurality of trained machine learning models configured to analyze image data of cells to classify or score the cells based on a different pre-defined criterion among a plurality of predefined criteria; (c) for each of the plurality of trained machine learning models, processing the image of the plurality of cells to classify or score the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria; (d) identifying a region of interest in the image of the plurality of cells based at least in part on the classifying or thescoring in (c); and (e) generating, based at least in part on the region of interest identified in (d), a visualization of the plurality of cells that distinguishes each of the different pre-defined criteria.
[0007] In some embodiments, the image of the plurality of cells comprises a plurality of images. In some embodiments, the plurality of images comprise a time series of images.
[0008] In some embodiments, the image of the plurality of cells is a single image. In some embodiments, the method further comprises performing flow cytometry analysis based at least in part on analyzing the single image.
[0009] In some embodiments, the method further comprises analyzing the image using a technique comprising multiplexed imaging system, high parameter spatial multi-omic system, multiplexed ion beam imaging by time of flight (MIBI-TOF), imaging mass cytometry (IMC), matrix-assisted laser desorption ionization mass spectrometry imaging (MALDI-MSI), digital spatial analysis (DSP), or geospatial clustering analysis.
[0010] In some embodiments, the plurality of cells comprise primary cells, stem cells, immune cells, carcinoma cells, sarcoma cells, lymphoma cells, germ cell tumor cells, blastoma cells, bladder cancer cells, breast cancer cells, colon cancer cells, colorectal cancer cells, endocrine tumor cells, esophageal cancer cells, glioblastoma cells, Hodgkin lymphoma cells, lung cancer cells, melanoma cells, or prostate cancer cells. In some embodiments, the plurality of cells is obtained at least in part by one or more of: biopsy collection, surgical resection, xenograft, animal model, fine needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, neoplastic tissue sampling, malignant tissue sampling, diseased tissue sampling, and implanted tissue sampling.
[0011] In some embodiments, the plurality of pre-defined criteria comprises gene expression, cellular structure, cellular morphology, next-generation sequencing (NGS), proteomic features, biochemical features, genomic features, transcriptomic features, methylation features, metabolomic features, metagenomic features, phenomics features, phenotypic features, or signaling pathway features.
[0012] In some embodiments, the method further comprises processing the image of the plurality of cells to classify the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the method further comprises combining a plurality of classifications to generate a composite classification for the image.
[0013] In some embodiments, the method further comprises processing the image of the plurality of cells to score the plurality of cells based on a different pre-defined criterion amongthe plurality of pre-defined criteria. In some embodiments, the method further comprises combining a plurality of scores to generate a composite score for the image.
[0014] In some embodiments, the method further comprises providing the visualization of the plurality of cells using a plurality of different colors indicative of each of the different pre-defined criteria. In some embodiments, the method further comprises use of fluorescence. In some embodiments, the method is performed without use of fluorescence. In some embodiments, the method further comprises use of immunohistochemistry staining. In some embodiments, the multiplex image analysis is performed without performing physical assays on the plurality of cells. In some embodiments, the method is performed without use of immunohistochemistry staining.
[0015] In some embodiments, the output of the machine learning model is based on a pre-defined criterion of a biomarker of a plurality of biomarkers. In some embodiments, the trained machine learning models comprise a supervised machine learning model, an unsupervised machine learning model, a deep learning model, or a time-series machine learning model.
[0016] In some embodiments, the image of the plurality of cells comprises surrounding tissue of the plurality of cells.
[0017] In another aspect, the present disclosure provides a computer-implemented method of performing image analysis of a plurality of cells, comprising: (a) obtaining an image of the plurality of cells, wherein the image is obtained without performing physical assays on the plurality of cells; (b) obtaining a trained machine learning model configured to analyze image data of cells to classify or score the cells based on a plurality of pre-defined criteria; (c) processing the image of the plurality of cells to classify or score the plurality of cells based on the plurality of pre-defined criteria; (d) identifying a region of interest in the image of the plurality of cells based at least in part on the classifying or the scoring in (c); and (e) generating, based at least in part on the region of interest identified in (d), a visualization of the plurality of cells that distinguishes each of the different pre-defined criteria.
[0018] In some embodiments, the image of the plurality of cells comprises a plurality of images. In some embodiments, the plurality of images comprise a time series of images.
[0019] In some embodiments, the image of the plurality of cells is a single image. In some embodiments, the method further comprises performing flow cytometry analysis based at least in part on analyzing the single image.
[0020] In some embodiments, the method further comprises analyzing the image using a technique comprising multiplexed imaging system, high parameter spatial multi-omic system, multiplexed ion beam imaging by time of flight (MIBI-TOF), imaging mass cytometry (IMC),matrix-assisted laser desorption ionization mass spectrometry imaging (MALDI-MSI), digital spatial analysis (DSP), or geospatial clustering analysis.
[0021] In some embodiments, the plurality of cells comprise primary cells, stem cells, immune cells, carcinoma cells, sarcoma cells, lymphoma cells, germ cell tumor cells, blastoma cells, bladder cancer cells, breast cancer cells, colon cancer cells, colorectal cancer cells, endocrine tumor cells, esophageal cancer cells, glioblastoma cells, Hodgkin lymphoma cells, lung cancer cells, melanoma cells, or prostate cancer cells. In some embodiments, the plurality of cells is obtained at least in part by one or more of: biopsy collection, surgical resection, xenograft, animal model, fine needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, neoplastic tissue sampling, malignant tissue sampling, diseased tissue sampling, and implanted tissue sampling.
[0022] In some embodiments, the plurality of pre-defined criteria comprises gene expression, cellular structure, cellular morphology, next-generation sequencing (NGS), proteomic features, biochemical features, genomic features, transcriptomic features, methylation features, metabolomic features, metagenomic features, phenomics features, phenotypic features, or signaling pathway features.
[0023] In some embodiments, the method further comprises processing the image of the plurality of cells to classify the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the method further comprises combining a plurality of classifications to generate a composite classification for the image.
[0024] In some embodiments, the method further comprises processing the image of the plurality of cells to score the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the method further comprises combining a plurality of scores to generate a composite score for the image.
[0025] In some embodiments, the method further comprises providing the visualization of the plurality of cells using a plurality of different colors indicative of each of the different pre-defined criteria. In some embodiments, the image analysis is performed without use of fluorescence. In some embodiments, the image analysis is performed without use of immunohistochemistry staining.
[0026] In some embodiments, output of the machine learning model is based on a predefined criterion of a biomarker of a plurality of biomarkers. In some embodiments, the trained machine learning models comprise a supervised machine learning model, an unsupervised machine learning model, a deep learning model, or a time-series machine learning model.
[0027] In some embodiments, the image of the plurality of cells comprises surrounding tissue of the plurality of cells.
[0028] Another aspect of the present disclosure provides a computer-implemented system comprising one or more processors and a computer memory having machineexecutable instructions stored thereon, the machine-executable instructions cause the one or more processors to perform operations comprising: (a) obtaining an image of the plurality of cells; (b) obtaining a plurality of trained machine learning models, each of the plurality of trained machine learning models configured to analyze image data of cells to classify or score the cells based on a different pre-defined criterion among a plurality of pre-defined criteria; (b) obtaining a plurality of trained machine learning models, each of the plurality of trained machine learning models configured to analyze image data of cells to classify or score the cells based on a different pre-defined criterion among a plurality of pre-defined criteria; (d) identifying a region of interest in the image of the plurality of cells based at least in part on the classifying or the scoring in (c); and (e) generating, based at least in part on the region of interest identified in (d), a visualization of the plurality of cells that distinguishes each of the different pre-defined criteria.
[0029] Another aspect of the present disclosure provides a non-transitory computer readable storage medium encoded with instructions executable by one or more processors to cause the one or more processors to perform operations comprising: (a) obtaining an image of the plurality of cells; (b) obtaining a plurality of trained machine learning models, each of the plurality of trained machine learning models configured to analyze image data of cells to classify or score the cells based on a different pre-defined criterion among a plurality of predefined criteria; (c) for each of the plurality of trained machine learning models, processing the image of the plurality of cells to classify or score the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria; (d) identifying a region of interest in the image of the plurality of cells based at least in part on the classifying or the scoring in (c); and (e) generating, based at least in part on the region of interest identified in (d), a visualization of the plurality of cells that distinguishes each of the different pre-defined criteria.
[0030] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, only illustrative embodiments of the present disclosure are shown and described. The present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure.Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0031] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0033] FIG. 1 illustrates a computer system that is programmed or otherwise configured to implement methods provided herein.
[0034] FIG. 2A illustrates an example of virtual multiplexing of mostly negative staining in the lesional cells. FIGs. 2B-2D show images of physical immunohistochemistry of a sample for negative HER2 (FIG. 2B), negative estrogen receptor (ER) (FIG. 2C), and negative progesterone receptor (PR) (FIG. 2D) corresponding to the same region of FIG. 2A.
[0035] FIG. 3A illustrates an example of virtual multiplexing of a sample with HER2 positive breast cancer. HER2 positive breast cancer is highlighted by orange peri-nuclear staining in the circle area. A benign breast cancer lobule is also presented as highlighted with green nuclei as the arrow pointing. Occasional ER-positive lobular cells are also presented with pink / red expression. FIGs. 3B-3D show images of physical immunohistochemistry of a sample depicting the same regions of FIG. 3A.
[0036] FIG. 4A illustrates an example of virtual multiplexing of a sample for invasive breast cancer with predominantly co-expression of ER and PR as the nuclei are highlighted in red and HER2 negative. FIGs. 4B-4D show images of physical immunohistochemistry of a sample of negative HER2 (FIG. 4B) and positive ER (FIG. 4C) and positive PR (FIG. 4D) from the same regions of FIG. 4A.
[0037] FIG. 5 illustrates an example of a topographical map showing cell prediction densities for HER2, ER, and PR as they relate to one another.DETAILED DESCRIPTION
[0038] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0039] As used in the specification and claims, the singular form “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a nucleic acid” includes a plurality of nucleic acids, including mixtures thereof.
[0040] Reference throughout this specification to “some embodiments,” “further embodiments,” or “a particular embodiment,” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in some embodiments,” or “in further embodiments,” or “in a particular embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0041] The term "image”, as used herein, generally refers to data which may provide a visual representation or depiction (e.g., of cells or tissues derived from a biological sample). “Image” may refer to a single image or a plurality of images. Images can be generated using various imaging techniques, such as microscopy, histology, immunohistochemistry, or other methods for biological sample analysis. Images may serve as visual records or representations of the cellular or tissue structures, morphology, and characteristics present in the biological sample under investigation. Images may undergo further processing, analysis, or enhancement (e.g., using specialized software or techniques) to extract information, identify regions of interest (e.g., specific cellular components), quantify various features, or facilitate further scientific interpretation. Images may encompass various imaging approaches and applications employed in the study of cells, tissues, and biological systems.
[0042] The term “multiplexing”, as used herein, generally refers to a technique or method used to combine multiple signals or streams of data into a single channel or transmission medium. The purpose of multiplexing is to optimize the use of available resources and improve the efficiency of data transmission or processing.
[0043] The term “machine learning model”, as used herein, generally refers to a mathematical or computational algorithm that is configured to make predictions or decisions based on input data (e.g., based on learned patterns). It may be a key component of machine learning, a subfield of artificial intelligence (Al).
[0044] The term “biomarker”, as used herein, generally refers to a pre-defined characteristic that may be measured or analyzed as an indicator of biological processes, pathogenic processes, or responses to an exposure or intervention, including therapeutic interventions. Biomarkers can be found in various biological materials, such as blood, tissues, urine, saliva, or other bodily fluids. Biomarkers can provide valuable information about the presence, progression, or severity of a particular disease or condition. They can also be used to assess the effectiveness of a treatment or intervention, predict disease outcomes, or aid in the diagnosis and classification of diseases. Various types of biomarkers include molecular biomarkers, histologic biomarkers, radiographic biomarkers, physiologic characteristics biomarkers, diagnostic biomarkers, prognostic biomarkers, predictive biomarkers, or surrogate biomarkers.
[0045] Multiplexing may allow the analysis of multiple parameters or molecules (e.g., of cells or tissues) in a single experiment, thereby enabling highly efficient and comprehensive data collection. Multiplexing may refer to a technique used in various fields, such as biology, engineering, and data science, to simultaneously analyze multiple samples (e.g., cell or tissue samples) or data points. It is a powerful technique for analyzing multiple parameters simultaneously; however, it can also be challenging to design and optimize multiplexing experiments, as the signals from multiple parameters may interfere with each other or be difficult to distinguish. Multiplexing allows for the simultaneous analysis of multiple targets, generating images that can provide a more comprehensive view of biological processes and interactions than may be possible with single-target analysis. However, large amounts of complex data including images can be challenging to analyze and interpret. Multiplexing may also encounter various challenges such as a high cost, for example, requiring about $500K in capital equipment purchases that allow for fluorescently labeled probes to multiplex.
[0046] In one aspect, disclosed herein a computer-implemented method of performing multiplex image analysis of a plurality of cells, comprising: (a) obtaining an image of the plurality of cells; (b) obtaining a plurality of trained machine learning models, each of the plurality of trained machine learning models configured to analyze image data of cells to classify or score the cells based on a different pre-defined criterion among a plurality of predefined criteria; (c) for each of the plurality of trained machine learning models, processing the image of the plurality of cells to classify or score the plurality of cells based on a differentpre-defined criterion among the plurality of pre-defined criteria; (d) identifying a region of interest in the image of the plurality of cells based at least in part on the classifying or the scoring in (c); and (e) generating, based at least in part on the region of interest identified in (d), a visualization of the plurality of cells that distinguishes each of the different pre-defined criteria.
[0047] In some embodiments, the image of the plurality of cells comprises a plurality of images. In some embodiments, the plurality of images comprise a time series of images.
[0048] In some embodiments, the image of the plurality of cells is a single image. In some embodiments, the method further comprises performing flow cytometry analysis based at least in part on analyzing the single image.
[0049] In some embodiments, the method further comprises analyzing the image using a technique comprising multiplexed imaging system, high parameter spatial multi-omic system, multiplexed ion beam imaging by time of flight (MIBI-TOF), imaging mass cytometry (IMC), matrix-assisted laser desorption ionization mass spectrometry imaging (MALDI-MSI), digital spatial analysis (DSP), or geospatial clustering analysis.
[0050] In some embodiments, the plurality of cells comprise primary cells, stem cells, immune cells, carcinoma cells, sarcoma cells, lymphoma cells, germ cell tumor cells, blastoma cells, bladder cancer cells, breast cancer cells, colon cancer cells, colorectal cancer cells, endocrine tumor cells, esophageal cancer cells, glioblastoma cells, Hodgkin lymphoma cells, lung cancer cells, melanoma cells, or prostate cancer cells. In some embodiments, the plurality of cells is obtained at least in part by one or more of biopsy collection, surgical resection, xenograft, animal model, fine needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, neoplastic tissue sampling, malignant tissue sampling, diseased tissue sampling, and implanted tissue sampling.
[0051] In some embodiments, the plurality of pre-defined criteria comprises gene expression, cellular structure, cellular morphology, next-generation sequencing (NGS), proteomic features, biochemical features, genomic features, transcriptomic features, methylation features, metabolomic features, metagenomic features, phenomics features, phenotypic features, or signaling pathway features.
[0052] In some embodiments, the method further comprises processing the image of the plurality of cells to classify the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the method further comprises combining a plurality of classifications to generate a composite classification for the image.
[0053] In some embodiments, the method further comprises processing the image of the plurality of cells to score the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the method further comprises combining a plurality of scores to generate a composite score for the image.
[0054] In some embodiments, the method further comprises providing the visualization of the plurality of cells using a plurality of different colors indicative of each of the different pre-defined criteria. In some embodiments, the method further comprises use of fluorescence. In some embodiments, the method is performed without use of fluorescence. In some embodiments, the method further comprises use of immunohistochemistry staining. In some embodiments, the method is performed without use of immunohistochemistry staining.
[0055] In some embodiments, output of the machine learning model is based on a predefined criterion of a biomarker of a plurality of biomarkers. In some embodiments, the trained machine learning models comprise a supervised machine learning model, an unsupervised machine learning model, a deep learning model, or a time-series machine learning model.
[0056] In some embodiments, the image of the plurality of cells comprises surrounding tissue of the plurality of cells.
[0057] In another aspect, disclosed herein a computer-implemented method of performing image analysis of a plurality of cells, comprising: (a) obtaining an image of the plurality of cells, the image is obtained without use of fluorescence; (b) obtaining a trained machine learning model configured to analyze image data of cells to classify or score the cells based on a plurality of pre-defined criteria; (c) processing the image of the plurality of cells to classify or score the plurality of cells based on the plurality of pre-defined criteria; (d) identifying a region of interest in the image of the plurality of cells based at least in part on the classifying or the scoring in (c); and (e) generating, based at least in part on the region of interest identified in (d), a visualization of the plurality of cells that distinguishes each of the different pre-defined criteria.
[0058] In some embodiments, the image of the plurality of cells comprises a plurality of images. In some embodiments, the plurality of images comprise a time series of images. In some embodiments, the image of the plurality of cells is a single image. In some embodiments, the method further comprises performing flow cytometry analysis based at least in part on analyzing the single image.
[0059] In some embodiments, the method further comprises analyzing the image using a technique comprising multiplexed imaging system, high parameter spatial multi-omic system, multiplexed ion beam imaging by time of flight (MIBI-TOF), imaging mass cytometry (IMC),matrix-assisted laser desorption ionization mass spectrometry imaging (MALDI-MSI), digital spatial analysis (DSP), or geospatial clustering analysis.
[0060] In some embodiments, the plurality of cells comprise primary cells, stem cells, immune cells, carcinoma cells, sarcoma cells, lymphoma cells, germ cell tumor cells, blastoma cells, bladder cancer cells, breast cancer cells, colon cancer cells, colorectal cancer cells, endocrine tumor cells, esophageal cancer cells, glioblastoma cells, Hodgkin lymphoma cells, lung cancer cells, melanoma cells, or prostate cancer cells. In some embodiments, the plurality of cells is obtained at least in part by one or more of: biopsy collection, surgical resection, xenograft, animal model, fine needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, neoplastic tissue sampling, malignant tissue sampling, diseased tissue sampling, and implanted tissue sampling.
[0061] In some embodiments, the plurality of pre-defined criteria comprises gene expression, cellular structure, cellular morphology, next-generation sequencing (NGS), proteomic features, biochemical features, genomic features, transcriptomic features, methylation features, metabolomic features, metagenomic features, phenomics features, phenotypic features, or signaling pathway features.
[0062] In some embodiments, the method further comprises processing the image of the plurality of cells to classify the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the method further comprises combining a plurality of classifications to generate a composite classification for the image.
[0063] In some embodiments, the method further comprises processing the image of the plurality of cells to score the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the method further comprises combining a plurality of scores to generate a composite score for the image.
[0064] In some embodiments, the method further comprises providing the visualization of the plurality of cells using a plurality of different colors indicative of each of the different pre-defined criteria. In some embodiments, the method further comprises use of fluorescence. In some embodiments, the method is performed without use of immunohistochemistry staining.
[0065] In some embodiments, output of the machine learning model is based on a predefined criterion of a biomarker of a plurality of biomarkers. In some embodiments, the trained machine learning models comprise a supervised machine learning model, an unsupervised machine learning model, a deep learning model, or a time-series machine learning model.
[0066] In some embodiments, the image of the plurality of cells comprises surrounding tissue of the plurality of cells.
[0067] Another aspect of the present disclosure provides a computer-implemented system comprising one or more processors and a computer memory having machineexecutable instructions stored thereon, the machine-executable instructions cause the one or more processors to perform operations comprising: (a) obtaining an image of the plurality of cells; (b) obtaining a plurality of trained machine learning models, each of the plurality of trained machine learning models configured to analyze image data of cells to classify or score the cells based on a different pre-defined criterion among a plurality of pre-defined criteria; (b) obtaining a plurality of trained machine learning models, each of the plurality of trained machine learning models configured to analyze image data of cells to classify or score the cells based on a different pre-defined criterion among a plurality of pre-defined criteria; (d) identifying a region of interest in the image of the plurality of cells based at least in part on the classifying or the scoring in (c); and (e) generating, based at least in part on the region of interest identified in (d), a visualization of the plurality of cells that distinguishes each of the different pre-defined criteria.
[0068] Another aspect of the present disclosure provides a non-transitory computer readable storage medium encoded with instructions executable by one or more processors to cause the one or more processors to perform operations comprising: (a) obtaining an image of the plurality of cells; (b) obtaining a plurality of trained machine learning models, each of the plurality of trained machine learning models configured to analyze image data of cells to classify or score the cells based on a different pre-defined criterion among a plurality of predefined criteria; (c) for each of the plurality of trained machine learning models, processing the image of the plurality of cells to classify or score the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria; (d) identifying a region of interest in the image of the plurality of cells based at least in part on the classifying or the scoring in (c); and (e) generating, based at least in part on the region of interest identified in (d), a visualization of the plurality of cells that distinguishes each of the different pre-defined criteria.
[0069] Cell Images
[0070] In some embodiments, the methods, systems, and media disclosed herein may comprise obtaining an image of a plurality of cells. In some embodiments, the methods, systems, and media disclosed herein use an image that is a single image. In some embodiments, the methods, systems, and media disclosed herein may comprise obtaining a plurality of images of a plurality cells. In some embodiments, the methods, systems, and mediadisclosed herein may comprise obtaining a plurality of images comprise a time series of images. In some embodiments, the image of a plurality of cells comprises surrounding tissue of the plurality of cells.
[0071] In some embodiments, the methods, systems, and media disclosed herein may comprise image processing, image segmentation, and / or object detection process as encoded in an image processing, image segmenting, image classification or scoring, or object detection algorithm. The image processing procedure may filter, transform, scale, rotate, mirror, shear, combine, compress, segment, concatenate, extract features from, and / or smooth an image prior to downstream processing. In some embodiments, a plurality of images is combined to form an image quilt. The image quilt may be converted to a representation (e.g., a tensor) that is useful for downstream processing of image data. The image segmentation process may partition an image into one or more segments which contain a factor or region of interest. For example, an image segmentation algorithm may process digital histopathology slides to determine a region of tissue as opposed to a region of whitespace or an artifact. In some embodiments, the image segmentation algorithm may comprise a machine learning or artificial intelligence algorithm. In some embodiments, image segmentation may precede image processing. In some embodiments, image processing may precede image segmentation. In some embodiments, the cell location is identified, and that region is isolated as a single cell image, which is scored or classified. The object detection process may comprise detecting the presence or absence of a target object (e.g., a cell or cell part, such as a nucleus). In some embodiments, object detection may proceed image processing and / or image segmentation. For example, images which are found by an image detection algorithm to contain one or more objects of interest may be concatenated in a subsequent image processing step. Alternatively, or additionally, image processing may precede object detection and / or image segmentation. For example, raw image data may be processed (e.g., filtered) and the processed image data subjected to an object detection algorithm. Image data may be subject to multiple image processing, image segmentation, and / or object detection steps in any appropriate order. In an example, image data is optionally subjected to one or more image processing steps to improve image quality. The processed image is then subjected to an image segmentation algorithm to detect regions of interest (e.g., regions of tissue in a set of histopathology slides). The regions of interest are then subjected to an object detection algorithm (e.g., algorithm to detect nuclei in images of tissue) and regions found to possess at least one target object are concatenated to produce processed image data for downstream use.
[0072] In some embodiments, the methods, systems, and media disclosed herein may comprise use of fluorescence. In some embodiments, the methods, systems, and mediadisclosed herein may comprise methods performed without use of fluorescence. In some embodiments, the methods, systems, and media disclosed herein may comprise use of immunohistochemistry staining. In some embodiments, the methods, systems, and media disclosed herein may comprise methods performed without use of immunohistochemistry staining.
[0073] In some embodiments, the methods, systems, and media disclosed herein may comprise obtaining an image of a plurality of cells, cells comprise primary cells, stem cells, immune cells, carcinoma cells, sarcoma cells, lymphoma cells, germ cell tumor cells, blastoma cells, bladder cancer cells, breast cancer cells, colon cancer cells, colorectal cancer cells, endocrine tumor cells, esophageal cancer cells, glioblastoma cells, Hodgkin lymphoma cells, lung cancer cells, melanoma cells, or prostate cancer cells.
[0074] In some embodiments, the methods, systems, and media disclosed herein may comprise obtaining an image of a plurality of cells, cells comprise a variety of cells, including eukaryotic cells, prokaryotic cells, fungi cells, heart cells, lung cells, kidney cells, liver cells, pancreas cells, reproductive cells, stem cells, induced pluripotent stem cells, gastrointestinal cells, blood cells, cancer cells, bacterial cells, bacterial cells isolated from a human microbiome sample, and circulating cells in the human blood. In some embodiments, cells may comprise contents of a cell, such as, for example, the contents of a single cell or the contents of multiple cells.
[0075] In some embodiments, the methods, systems, and media disclosed herein may comprise obtaining an image of a plurality of cells, cells further comprise tumor or cancer cells. Non-limiting examples of tumors and associated cancers may include: acoustic neuroma, acute lymphoblastic leukemia, acute myeloid leukemia, adenocarcinoma, adrenocortical carcinoma, AIDS-related cancers, AIDS-related lymphoma, anal cancer, angiosarcoma, appendix cancer, astrocytoma, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancers, brain tumors, such as cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumors, visual pathway and hypothalamic glioma, breast cancer, bronchial adenomas, Burkitt lymphoma, carcinoma of unknown primary origin, central nervous system lymphoma, bronchogenic carcinoma, cerebellar astrocytoma, cervical cancer, childhood cancers, chondrosarcoma, chordoma, choriocarcinoma, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative disorders, colon cancer, colon carcinoma, craniopharyngioma, cutaneous T-cell lymphoma, cystadenocarcinoma, desmoplastic small round cell tumor, embryonal carcinoma, endocrine system carcinomas, endometrial cancer, endotheliosarcoma, ependymoma, epithelial carcinoma, esophageal cancer, Ewing’s sarcoma,fibrosarcoma, germ cell tumors, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, gastrointestinal system carcinomas, genitourinary system carcinomas, gliomas, hairy cell leukemia, head and neck cancer, heart cancer, hemangioblastoma, hepatocellular (liver) cancer, Hodgkin lymphoma, Hypopharyngeal cancer, intraocular melanoma, islet cell carcinoma, Kaposi sarcoma, kidney cancer, laryngeal cancer, leiomyosarcoma, lip and oral cavity cancer, liposarcoma, liver cancer, lung cancers, such as non-small cell and small cell lung cancer, lung carcinoma, lymphangiosarcoma, lymphangioendotheliosarcoma, lymphomas, leukemias, macroglobulinemia, malignant fibrous histiocytoma of bone / osteosarcoma, medulloblastoma, medullary carcinoma, melanomas, meningioma, mesothelioma, metastatic squamous neck cancer with occult primary, mouth cancer, multiple endocrine neoplasia syndrome, myelodysplastic syndromes, myeloid leukemia, myxosarcoma, nasal cavity and paranasal sinus cancer, nasopharyngeal carcinoma, neuroblastoma, non-Hodgkin lymphoma, non-small cell lung cancer, oligodendroma, oral cancer, oropharyngeal cancer, osteosarcoma / malignant fibrous histiocytoma of bone, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumor, pancreatic cancer, pancreatic cancer islet cell, papillary adenocarcinoma, papillary carcinoma, paranasal sinus and nasal cavity cancer, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pineal astrocytoma, pineal germinoma, pituitary adenoma, pleuropulmonary blastoma, plasma cell neoplasia, primary central nervous system lymphoma, prostate cancer, rectal cancer, renal cell carcinoma, renal pelvis and ureter transitional cell cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcomas, sebaceous gland carcinoma, seminoma, skin cancers, skin carcinoma merkel cell, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, stomach cancer, sweat gland carcinoma, synovioma, T-cell lymphoma, testicular tumor, throat cancer, thymoma, thymic carcinoma, thyroid cancer, trophoblastic tumor (gestational), cancers of unknown primary site, urethral cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, Wilms tumor, or combinations thereof. The tumors may be associated with various types of organs. Non-limiting examples of organs may include brain, breast, liver, lung, kidney, prostate, ovary, spleen, lymph node (including tonsil), thyroid, pancreas, heart, skeletal muscle, intestine, larynx, esophagus, stomach, or combinations thereof.
[0076] Images of individual frames from a time-lapse movie can be used to analyze and track cellular behavior, quantify changes in cell morphology, measure cell migration, or study the dynamics of specific cellular components such as organelles or proteins. In some embodiments, the methods, systems, and media disclosed herein may comprise acquiring a series of images over time. In some embodiments, the methods, systems, and media disclosedherein may comprise capturing images of cells over time by a time-lapse microscopy. In some embodiments, the time-lapse microscopy comprises Bright-Field Time-Lapse Microscopy, Phase-Contrast Time-Lapse Microscopy, Fluorescence Time-Lapse Microscopy, Confocal Time-Lapse Microscopy, or Spinning Disk Time-Lapse Microscopy.
[0077] Time-lapse images may be taken at one or more timepoints. In some embodiments, the methods, systems, and media disclosed herein may comprise acquiring images at a time interval.
[0078] In some embodiments, the time interval may comprise at least about 0.1 second, at least about 0.2 second, at least about 0.3 second, at least about 0.5 second, at least about 1 second, at least about 2 seconds, at least about 3 seconds, at least about 4 seconds, at least about 5 seconds, at least about 6 seconds, at least about 7 seconds, at least about 8 seconds, at least about 9 seconds, at least about 10 seconds, at least about 12 seconds, at least about 15 seconds, at least about 20 seconds, at least about 25 seconds, at least about 30 seconds, at least about 1 minutes, at least about 2 minutes, at least about 5 minutes, at least about 10 minutes, at least about 20 minutes, at least about 30 minutes, at least about 40 minutes, at least about 50 minutes, at least about 1 hour, at least about 1 day, at least about 1 week, at least about 2 weeks, at least about 3 weeks, or at least about 4 weeks or longer.
[0079] In some embodiments, the time interval may comprise at most about 4 weeks or longer, at most about 3 weeks, at most about 2 weeks, at most about 1 week, at most about 1 day, at most about 1 hour, at most about 50 minutes, at most about 40 minutes, at most about 30 minutes, at most about 20 minutes, at most about 10 minutes, at most about 5 minutes, at most about 2 minutes, at most about 1 minutes, at most about 30 seconds, at most about 25 seconds, at most about 20 seconds, at most about 15 seconds, at most about 12 seconds, at most about 10 seconds, at most about 9 seconds, at most about 8 seconds, at most about 7 seconds, at most about 6 seconds, at most about 5 seconds, at most about 4 seconds, at most about 3 seconds, at most about 2 seconds, at most about 1 second, at most about 0.5 second, at most about 0.3 second, at most about 0.2 second, at most about 0.1 second.
[0080] In some embodiments, the images are from more than two timepoints. In some embodiments, the images are from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more timepoints.
[0081] In some embodiments, the methods, systems, and media disclosed herein may comprise obtaining an image of a plurality of cells, cells are obtained at least in part by one or more of: biopsy collection, surgical resection, xenograft, animal model, fine needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, neoplastic tissue sampling, malignant tissue sampling, diseased tissue sampling, and implanted tissue sampling.
[0082] In some embodiments, the tissues may be associated with various types of organs. Non-limiting examples of organs may include brain, breast, liver, lung, kidney, prostate, ovary, spleen, lymph node (including tonsil), thyroid, pancreas, heart, skeletal muscle, intestine, larynx, esophagus, stomach, or combinations thereof. In some embodiments, the tissues are associated with a prostate of the subject. In the case of a biological sample comprising cells and / or tissue (e.g., a biopsy sample), the biological sample may be further analyzed or assayed. In some embodiments, the biopsy sample may be fixed, processed (e.g., dehydrated), embedded, frozen, stained, and / or examined under a microscope. In some embodiments, digital slides are generated from processed samples.
[0083] Biomarkers
[0084] In some embodiments, the methods, systems, and media disclosed herein may comprise use of a plurality of pre-defined criteria. For example, the plurality of pre-defined criteria may comprise gene expression, cellular structure, cellular morphology, next-generation sequencing (NGS), proteomic features, biochemical features, genomic features, transcriptomic features, methylation features, metabolomic features, metagenomic features, phenomics features, phenotypic features, or signaling pathway features. In some embodiments, the methods, systems, and media disclosed herein may comprise obtaining a plurality of trained machine learning models, each of the plurality of trained machine learning models configured to analyze image data of cells to classify or score the cells based on a different pre-defined criterion among a plurality of pre-defined criteria. In some embodiments, output of the machine learning model is based on a pre-defined criterion of a biomarker of a plurality of biomarkers. A biomarker can be, for example, a cell, a small molecule, a macromolecule, a protein, a glycoprotein, a carbohydrate, a sugar, a polypeptide, a nucleic acid (e.g., deoxyribonucleic acid (DNA), ribonucleic acid (RNA)), a cell-free nucleic acid (e.g., cf-DNA, cf-RNA), a lipid, a cellular component, or combinations thereof.
[0085] Gene expression profiles can serve as biomarkers to distinguish between healthy and diseased states, assess disease severity or progression, predict treatment responses, or identify specific subtypes or stages of diseases. In some embodiments, gene expression can be measured using techniques of RNA sequencing (RNA-seq), microarrays, or quantitative polymerase chain reaction (qPCR).
[0086] Cellular structure and cellular morphology can serve as biomarkers for diagnostic and prognostic information about diseases or cellular states. In some embodiments, cellular structure and cellular morphology can be measured and analyzed by cell staining dyes comprised of hematoxylin and eosin (H&E) staining, or Giemsa staining. In some embodiments, cellular structure and cellular morphology can be measured and analyzed byfluorescent dyes or probes targeting cellular structures comprising 4',6-diamidino-2- phenylindole (DAP I), phalloidin, ER-Tracker™, Lyso-Tracker™ or Mito-Tracker™ dyes. In some embodiments, cellular structure and cellular morphology can be measured and analyzed by specific antibodies used in immunocytochemistry, immunohistochemistry (IHC), or immunofluorescence. In some embodiments, cellular structure and cellular morphology can be measured and analyzed by cell membrane markers comprise lipid-binding dyes or fluorescently tagged lectins.
[0087] In some embodiments, the methods, systems, and media disclosed herein may comprise applying Next-generation sequencing (NGS) to analyze biomarkers and genetic information in a high-throughput manner. NGS enables the simultaneous sequencing of millions of DNA or RNA fragments, to obtain detailed information about the genetic material present in a biological sample. In some embodiments, NGS can be utilized to identify and characterize genomic biomarkers comprising analyzing genomic DNA, detecting genetic variations comprising single nucleotide polymorphisms (SNPs), insertions, deletions, or structural variants. In some embodiments, NGS can be utilized to analyze transcriptomic biomarkers for analysis of RNA molecule to provide insights into gene expression levels and alternative splicing patterns. In some embodiments, NGS can be utilized to analyze epigenetic modifications, including DNA methylation and histone modifications. In some embodiments, NGS can be utilized to analyze the composition and diversity of microbial communities present in a biological sample based on microbiome biomarkers. In some embodiments, NGS can be utilized for discovery and profiling of non-coding RNAs, such as microRNAs and long non-coding RNAs (Inc-RNAs).
[0088] Proteomic features can serve as biomarkers by providing valuable information about protein expression, post-translational modifications, and protein-protein interactions within biological systems. In some embodiments, biomarkers derived from proteomic analysis comprise protein expression levels, post-translational modifications, protein isoforms and splice variants, protein-protein interactions, secreted or exosome proteins, or drug response and resistance.
[0089] Biochemical features encompass various molecules and their properties, such as enzymes, proteins, metabolites, or small molecules. Changes in the levels or activities of these biomolecules can be indicative of specific diseases or physiological conditions. In some embodiments, biomarkers derived from biological features comprise enzyme activity, protein biomarkers, metabolites, hormones, lipids, oxidative stress markers, coagulation markers, or genetic biomarkers. Genomic features refer to genetic variations, mutations, or alterations in the DNA sequence of an individual. Certain genetic variations, such as single nucleotidepolymorphisms (SNPs) or structural variants, can be associated with disease susceptibility or drug response. In some embodiments, biomarkers derived from genomic features comprise single nucleotide polymorphisms, copy number variations, structural variants, gene expression profiling, mutations in specific genes, pharmacogenomics, or genomic signatures.
[0090] Transcriptomic features involve analyzing the expression levels of genes or RNA molecules, such as messenger RNA (mRNA) or non-coding RNA, in a particular sample. Differential gene expression patterns or alternative splicing events can serve as biomarkers for specific diseases or physiological conditions. In some embodiments, in-situ hybridization (ISH) can be used with specific nucleic acid probes that target RNA sequences of interest. In some embodiments, in-situ hybridization (ISH) can be used to provide valuable information about the expression patterns of specific genes, help identify cell types or regions of interest based on their RNA content. In some embodiments, in-situ hybridization (ISH) can be used to analyze single genes or a small set of genes in a spatial context, to allow visualization of gene expression patterns within intact tissue sections, providing information about the distribution and localization of specific transcripts. Transcriptomic biomarkers can be identified through techniques like RNA sequencing (RNA-seq) or microarray analysis. Transcriptomic features encompass biomarkers derived from the analysis of gene expression patterns, RNA molecules, and other transcriptomic data. In some embodiments, biomarkers derived from transcriptomic features comprise differential gene expression, alternative splicing patterns, fusion genes and chimeric transcripts, non-coding RNA biomarkers, gene expression signatures, or regulatory networks and pathways.
[0091] Metabolomic features involve the analysis of small molecules or metabolites present in a biological sample. Metabolites are sensitive to changes in cellular processes and can reflect alterations in metabolic pathways associated with diseases. Metabolomic biomarkers can be identified through techniques like mass spectrometry or nuclear magnetic resonance spectroscopy, enabling the detection of disease-specific metabolic signatures. In some embodiments, biomarkers derived from metabolomic features comprise metabolite profiling, disease specific metabolic signatures, metabolic pathway biomarkers, drug response biomarkers, environmental exposure markers, or microbiome related biomarkers.
[0092] Metagenomic features involve the analysis of genetic material derived from a microbial community present in a biological sample. Metagenomics allows the identification and characterization of microbial species and their functional potential. Changes in the composition or diversity of the microbiome can serve as biomarkers for various diseases or conditions, including gastrointestinal disorders, autoimmune diseases, or even mental health disorders. In some embodiments, biomarkers derived from metagenomic features comprisetaxonomic biomarkers, functional biomarkers, antibiotic resistance markers, viral biomarkers, community structure biomarkers, or functional potential biomarkers. Phenomics features involve the comprehensive analysis of phenotypic traits or characteristics of an organism, such as physical traits, behavior, or response to stimuli. These features can provide insights into disease susceptibility, treatment response, or prognosis. Phenotypic biomarkers can be identified through rigorous clinical observation, imaging techniques, or high-throughput phenotyping platforms. In some embodiments, biomarkers derived from phenomics features comprise physical characteristics, vital signs, cognitive function, behavioral biomarkers, metabolic biomarkers, or disease specific phenotypes. Signaling pathway features involve the analysis of molecular signaling cascades or networks involved in cellular processes. Dysregulation or perturbation of signaling pathways can be associated with diseases. In some embodiments, biomarkers derived from signaling pathway features comprise phosphorylation status of proteins, activation of transcription factors, expression of key signaling proteins, protein-protein interactions, expression or activity of downstream targets, small molecule metabolites, or gene expression signatures.
[0093] Trained algorithms
[0094] In some embodiments, the methods, systems, and media disclosed herein may comprise training machine learning models based on a different pre-defined criterion. In some embodiments, the methods, systems, and media disclosed herein may further comprise processing the image of the plurality of cells to classify the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the methods, systems, and media disclosed herein may comprise trained machine learning models comprise a supervised machine learning model, an unsupervised machine learning model, a deep learning model, or a time-series machine learning model. In some embodiments, the methods, systems, and media disclosed herein may further comprise combining a plurality of classifications to generate a composite classification for the image. In some embodiments, the methods, systems, and media disclosed herein may further comprise processing the image of the plurality of cells to score the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the methods, systems, and media disclosed herein may further comprises combining a plurality of scores to generate a composite score for the image. In some embodiments, the methods, systems, and media disclosed herein may comprise output of the machine learning model based on a pre-defined criterion of a biomarker of a plurality of biomarkers.
[0095] Methods and systems as disclosed herein are presenting outputs of multiple machine-learning models in a single output image. In some embodiments, machine learningmodels are trained based on a different pre-defined criterion among a plurality of pre-defined criteria. The trained algorithm may be configured to identify the criterion with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more than 99% for at least about 25, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, or more than about 500 independent samples.
[0096] The trained algorithm may comprise an unsupervised machine learning algorithm. The trained algorithm may comprise a supervised machine learning algorithm. The trained algorithm may comprise a deep learning algorithm. The trained algorithm may comprise a time-series machine learning algorithm. The trained algorithm may comprise a classification and regression tree (CART) algorithm. The supervised machine learning algorithm may comprise, for example, a Random Forest, a support vector machine (SVM), a neural network, or a deep learning algorithm. The trained algorithm may comprise a selfsupervised machine learning algorithm. The time-series machine learning algorithm may comprise autoregressive integrated moving average (ARIMA), recurrent neural networks (RNN), convolutional neural networks (CNN), Gaussian processes, long short-term memory networks, gated recurrent unit networks, Hidden Markov Models, or transformer-based models.
[0097] In some embodiments, a machine learning algorithm of a method or system as described herein utilizes one or more neural networks. In some case, a neural network is a type of computational system that can learn the relationships between an input dataset and a target dataset. A neural network may be a software representation of a human neural system (e.g., cognitive system), intended to capture “learning” and “generalization” abilities as used by a human. In some embodiments, the machine learning algorithm comprises a neural network comprising a CNN. Non-limiting examples of structural components of machine learning algorithms described herein include: CNNs, recurrent neural networks, dilated CNNs, fully- connected neural networks, deep generative models, and Boltzmann machines.
[0098] In some embodiments, a neural network comprises a series of layers termed “neurons.” In some embodiments, a neural network comprises an input layer, to which data is presented; one or more internal, and / or “hidden”, layers; and an output layer. A neuron may be connected to neurons in other layers via connections that have weights, which are parameters that control the strength of the connection. The number of neurons in each layer may be related to the complexity of the problem to be solved. The minimum number of neuronsrequired in a layer may be determined by the problem complexity, and the maximum number may be limited by the ability of the neural network to generalize. The input neurons may receive data being presented and then transmit that data to the first hidden layer through connections’ weights, which are modified during training. The first hidden layer may process the data and transmit its result to the next layer through a second set of weighted connections. Each subsequent layer may “pool” the results from the previous layers into more complex relationships. In addition, whereas conventional software programs require writing specific instructions to perform a function, neural networks are programmed by training them with a known sample set and allowing them to modify themselves during (and after) training so as to provide a desired output such as an output value. After training, when a neural network is presented with new input data, it is configured to generalize what was “learned” during training and apply what was learned from training to the new previously unseen input data in order to generate an output associated with that input.
[0099] In some embodiments, the neural network comprises artificial neural networks (ANNs). ANNs may be machine learning algorithms that may be trained to map an input dataset to an output dataset, where the ANN comprises an interconnected group of nodes organized into multiple layers of nodes. For example, the ANN architecture may comprise at least an input layer, one or more hidden layers, and an output layer. The ANN may comprise any total number of layers, and any number of hidden layers, where the hidden layers function as trainable feature extractors that allow mapping of a set of input data to an output value or set of output values. As used herein, a deep learning algorithm (such as a deep neural network (DNN)) is an ANN comprising a plurality of hidden layers, e.g., two or more hidden layers. Each layer of the neural network may comprise a number of nodes (or “neurons”). A node receives input that comes either directly from the input data or the output of nodes in previous layers, and performs a specific operation, e.g., a summation operation. A connection from an input to a node is associated with a weight (or weighting factor). The node may sum up the products of all pairs of inputs and their associated weights. The weighted sum may be offset with a bias. The output of a node or neuron may be gated using a threshold or activation function. The activation function may be a linear or non-linear function. The activation function may be, for example, a rectified linear unit (ReLU) activation function, a Leaky ReLU activation function, or other function such as a saturating hyperbolic tangent, identity, binary step, logistic, arctan, softsign, parametric rectified linear unit, exponential linear unit, softplus, bent identity, softexponential, sinusoid, sine, Gaussian, or sigmoid function, or any combination thereof.
[0100] The weighting factors, bias values, and threshold values, or other computational parameters of the neural network, may be “taught” or “learned” in a training phase using one or more sets of training data. For example, the parameters may be trained using the input data from a training dataset and a gradient descent or backward propagation method so that the output value(s) that the ANN computes are consistent with the examples included in the training dataset.
[0101] The number of nodes used in the input layer of the ANN or DNN may be at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, or greater. In other instances, the number of nodes used in the input layer may be at most about 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9000, 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10, or less. In some instances, the total number of layers used in the ANN or DNN (including input and output layers) may be at least about 3, 4, 5, 10, 15, 20, or greater. In other instances, the total number of layers may be at most about 20, 15, 10, 5, 4, 3, or less.
[0102] In some instances, the total number of learnable or trainable parameters, e.g., weighting factors, biases, or threshold values, used in the ANN or DNN may be at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, or greater. In other instances, the number of learnable parameters may be at most about 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9000, 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10, or less.
[0103] In some embodiments of a machine learning algorithm as described herein, a machine learning algorithm comprises a neural network such as a deep CNN. In some embodiments in which a CNN is used, the network is constructed with any number of convolutional layers, dilated layers or fully-connected layers. In some embodiments, the number of convolutional layers is between 1-10 and the dilated layers between 0-10. The total number of convolutional layers (including input and output layers) may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or greater, and the total number of dilated layers may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or greater. The total number of convolutional layers may be at most about 20, 15, 10, 5, 4, 3, or less, and the total number of dilated layers may be at most about 20, 15, 10, 5, 4, 3, or less. In some embodiments, the number of convolutional layers is between 1-10 and the fully-connected layers between 0-10. The total number of convolutional layers(including input and output layers) may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or greater, and the total number of fully-connected layers may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or greater. The total number of convolutional layers may be at most about 20, 15, 10, 5, 4, 3, 2, 1, or less, and the total number of fully-connected layers may be at most about 20, 15, 10, 5, 4, 3, 2, 1, or less.
[0104] In some embodiments, a machine learning algorithm comprises a neural network comprising a CNN, RNN, dilated CNN, fully-connected neural networks, deep generative models and / or deep restricted Boltzmann machines.
[0105] In some embodiments, a machine learning algorithm comprises one or more CNNs. The CNN may be deep and feedforward ANNs. The CNN may be applicable to analyzing visual imagery. The CNN may comprise an input, an output layer, and multiple hidden layers. The hidden layers of a CNN may comprise convolutional layers, pooling layers, fully-connected layers and normalization layers. The layers may be organized in 3 dimensions: width, height, and depth.
[0106] The convolutional layers may apply a convolution operation to the input and pass results of the convolution operation to the next layer. For processing images, the convolution operation may reduce the number of free parameters, allowing the network to be deeper with fewer parameters. In neural networks, each neuron may receive input from some number of locations in the previous layer. In a convolutional layer, neurons may receive input from only a restricted subarea of the previous layer. The convolutional layer's parameters may comprise a set of learnable filters (or kernels). The learnable filters may have a small receptive field and extend through the full depth of the input volume. During the forward pass, each filter may be convolved across the width and height of the input volume, compute the dot product between the entries of the filter and the input, and produce a two-dimensional activation map of that filter. As a result, the network may learn filters that activate when it detects some specific type of feature at some spatial position in the input.
[0107] In some embodiments, a machine learning algorithm comprises an RNN. RNNs are neural networks with cyclical connections that can encode and process sequential data. An RNN can include an input layer that is configured to receive a sequence of inputs. An RNN may additionally include one or more hidden recurrent layers that maintain a state. At each step, each hidden recurrent layer can compute an output and a next state for the layer. The next sate may depend on the previous state and the current input. The state may be maintained across steps and may capture dependencies in the input sequence.
[0108] An RNN can be a long short-term memory (LSTM) network. An LSTM network may be made of LSTM units. An LSTM unit may include of a cell, an input gate, an outputgate, and a forget gate. The cell may be responsible for keeping track of the dependencies between the elements in the input sequence. The input gate can control the extent to which a new value flows into the cell, the forget gate can control the extent to which a value remains in the cell, and the output gate can control the extent to which the value in the cell is used to compute the output activation of the LSTM unit.
[0109] Alternatively, an attention mechanism (e.g., a transformer) is applied to mimic human cognitive process of selectively focusing on relevant information while filtering out irrelevant details. Attention mechanisms may focus on, or “attend to,” certain input regions while ignoring others. This may increase model performance because certain input regions may be less relevant. At each step, an attention unit can compute a dot product of a context vector and the input at the step, among other operations. The output of the attention unit may define where the most relevant information in the input sequence is located.
[0110] In some embodiments, the pooling layers comprise global pooling layers. The global pooling layers may combine the outputs of neuron clusters at one layer into a single neuron in the next layer. For example, max pooling layers may use the maximum value from each of a cluster of neurons in the prior layer; and average pooling layers may use the average value from each of a cluster of neurons at the prior layer.
[0111] In some embodiments, the fully-connected layers connect every neuron in one layer to every neuron in another layer. In neural networks, each neuron may receive input from some number locations in the previous layer. In a fully-connected layer, each neuron may receive input from every element of the previous layer.
[0112] In some embodiments, the normalization layer is a batch normalization layer. The batch normalization layer may improve the performance and stability of neural networks. The batch normalization layer may provide any layer in a neural network with inputs that are zero mean / unit variance. The advantages of using batch normalization layer may include faster trained networks, higher learning rates, easier to initialize weights, more activation functions viable, and simpler process of creating deep networks.
[0113] The trained algorithm may be configured to accept a plurality of input variables and to produce one or more output values based on the plurality of input variables. The plurality of input variables may comprise one or more datasets indicative of pre-defined criteria. For example, an input variable may comprise features extracted from an image of a Hematoxylin and eosin (H&E) slide (e.g., prepared from formalin-fixed paraffin-embedded tissue). The plurality of input variables may also include other pre-defined criteria.
[0114] In some embodiments, the methods, systems, and media disclosed herein may comprise feature extraction for image processing. In some embodiments, feature extractioncomprises regions of interest (RO I), segmentation, texture features, automatic image processing or machine vision algorithms.
[0115] In some embodiments, the methods, systems, and media disclosed herein may comprise image processing, image segmentation, and / or object detection process as encoded in an image processing, image segmenting, or object detection algorithm. The image processing procedure may filter, transform, scale, rotate, mirror, shear, combine, compress, segment, concatenate, extract features from, and / or smooth an image prior to downstream processing. In some embodiments, a plurality of images is combined to form an image quilt. The image quilt may be converted to a representation (e.g., a tensor) that is useful for downstream processing of image data.
[0116] The image segmentation process may partition an image into one or more segments which contain a factor or region of interest. For example, an image segmentation algorithm may process digital histopathology slides to determine a region of tissue as opposed to a region of whitespace or an artifact. In some embodiments, the image segmentation algorithm may comprise a machine learning or artificial intelligence algorithm. In some embodiments, image segmentation may precede image processing.
[0117] In some embodiments, image processing may precede image segmentation. The object detection process may comprise detecting the presence or absence of a target object (e.g., a cell or cell part, such as a nucleus). In some embodiments, object detection may precede image processing and / or image segmentation. For example, images which are found by an image detection algorithm to contain one or more objects of interest may be concatenated in a subsequent image processing step.
[0118] Alternatively, or additionally, image processing may precede object detection and / or image segmentation. For example, raw image data may be processed (e.g., filtered) and the processed image data subjected to an object detection algorithm. Image data may be subject to multiple image processing, image segmentation, and / or object detection steps in any appropriate order. In an example, image data is optionally subjected to one or more image processing steps to improve image quality. The processed image is then subjected to an image segmentation algorithm to detect regions of interest (e.g., regions of tissue in a set of histopathology slides). The regions of interest are then subjected to an object detection algorithm (e.g., algorithm to detect nuclei in images of tissue) and regions found to possess at least one target object are concatenated to produce processed image data for downstream use.
[0119] A region of interest (ROI) selection can be based on various criteria. In some embodiments, ROI can be selected based on spatial coordinates, specific objects, features within the image, or user-defined boundaries. In some embodiments, ROI can be selectedbased on regions or structures within an image that correspond to individual cells or specific cellular features. In some embodiments, ROI selection nay be started by cell segmentation. Cell segmentation assigns each pixel or region in the image to its corresponding cell, creating distinct regions of interest for each cell. In some embodiments, specific morphological features can be measured within the defined ROIs. These features may include area, perimeter, diameter, circularity, eccentricity, solidity, aspect ratio, and others that describe the size, shape, and spatial characteristics of the cells. In some embodiments, ROIs can be applied to analyze subcellular structures and or regions of interest within the cells. In some embodiments, analysis can be applied for studying cell density, clustering, proximity to neighboring cells, or orientation by selecting ROI around cells of interest or specific regions within the tissue. In some further embodiments, ROI selection can be used for cell tracking purposes by defining ROIs around cells of interest at different time points.
[0120] In some embodiments, texture features can be selected on statistical or structural properties of the spatial distribution of pixel intensities within an image. Texture features provide information about the texture patterns, variations, or regularities present in an image of biological sample. Texture analysis can be useful in various applications, such as image classification, segmentation, and object recognition. In some embodiments, texture features comprise Gray-Level Co-occurrence Matrix (GLCM), Local Binary Patterns (LBP), Gabor Filters, Haralick features, Local Binary Patterns on Three Orthogonal Planes (LBP-TOP), Wavelet Transform, Histogram of Oriented Gradients (HOG), or Tamura Texture Features.
[0121] In some embodiments, image processing comprises automatic image processing applying algorithms and computational techniques to analyze, enhance, or extract information from images without requiring manual intervention. In some embodiments, automatic image processing includes image acquisition, preprocessing, image analysis and feature extraction, feature representation, classification and decision making, post-processing and visualization, automation and integration or combination of all.
[0122] In some embodiments, image processing comprises machine vision algorithms. In some embodiments, machine vision algorithms comprise image filtering, image segmentation, object detection and recognition, feature extraction, image registration, optical character recognition (OCR), tracking and motion analysis, 3D vision and depth estimation, or deep learning based methods.
[0123] The trained algorithm may comprise a classifier, such that each of the one or more output values comprises one of a fixed number of possible values (e.g., a linear classifier, a logistic regression classifier, etc.) indicating a classification of biomarkers by the classifier. The trained algorithm may comprise a binary classifier, such that each of the one ormore output values comprises one of two values (e.g., {0, 1 }, {positive, negative}, or {high- likelihood, low-likelihood}) indicating a classification of biomarkers by the classifier. The trained algorithm may be another type of classifier, such that each of the one or more output values comprises one of more than two values (e.g., {0, 1, 2}, {positive, negative, or indeterminate}, {high-likelihood, intermediate-likelihood, or low-likelihood}, or {high- intensity staining, intermediate-intensity staining, or low-intensity staining) indicating a classification of biomarkers by the classifier. The output values may comprise descriptive labels, numerical values, or a combination thereof. Some of the output values may comprise descriptive labels. Such descriptive labels may provide conditions of cells, comprising classification of cells for live or dead, classification of cells for malignant or benign, classification of cells for normal or abnormal, classification of cells for differentiated or undifferentiated, classification of cells for stem or non-stem, classification of cells for immunophenotype (e.g., T cells, B cells, macrophages), classification of cells for cell cycle phases (e.g. gap phase 1, synthesis phase, gap phase 2, or mitosis), classification of cells for subtype or lineage, classification of cells for metastatic or non-metastatic, or classification of cells for drug response. Such descriptive labels may provide an identification of biomarkers, and may comprise, for example, status of cellular structure and cellular morphology.
[0124] Some of the output values may comprise numerical values, such as binary, integer, or continuous values. Such binary output values may comprise, for example, {0, 1 }, {positive, negative}, {high-likelihood, low-likelihood}, or {high-intensity staining, low- intensity staining}. Such integer output values may comprise, for example, {0, 1, 2}. Such continuous output values may comprise, for example, cell area, cell diameter, cell perimeter, cell circularity, cell aspect ratio, cell eccentricity, cell solidity, cell intensity statistics, cell texture feature, or cell gradient or edge-based feature. Such continuous output values may comprise, for example, a probability value of at least 0 and no more than 1. Such continuous output values may comprise, for example, an un-normalized probability value of at least 0. Some numerical values may be mapped to descriptive labels, for example, by mapping 1 to “positive” and 0 to “negative.”
[0125] Systems and methods as described herein may use more than trained algorithm to determine an output (e.g., biomarkers). Systems and methods may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more trained algorithms. A trained algorithm of the plurality of trained algorithms may be trained on a particular type of data (e.g., image data or tabular data). Alternatively, a trained algorithm may be trained on more than one type of data. The inputs of one trained algorithm may comprise the outputs of one or more othertrained algorithms. Additionally, a trained algorithm may receive as its input the output of one or more trained algorithms.
[0126] Statistical analysis
[0127] In some embodiments, the methods, systems, and media disclosed herein further comprise applying statistical analysis on outputs from machine learning models. In some embodiments, statistical analysis includes evaluating the minimum values, the maximum values, the standard deviation, the variance, the mean, the median, the mode, the quartiles, the percentiles, or the sum of the outputs from machine learning models. In some embodiments, statistical analysis is performed on multiple model outputs in the region corresponding to the cell or associated tissue.
[0128] In some embodiments, statistical analysis is used to categorize or score each cell multiple times, once per model. In some embodiments, the results obtained from multiple models can further be combined mathematically to result in new additional categories. In some embodiments, a threshold is applied using the mean score from multiple machine learning models. Cells with features above a given threshold are placed into a given category. In some embodiments, the category with the maximum score may be used as the final category. In some embodiments, statistical analysis may also be used to infer a category which can be reported to a user such as standard deviation, variance, mean, mode, the minimum, the maximum or the quartiles. In some embodiments, cells with features comprising a combined score from multiple machine learning models above a threshold may be reported differently than other cells. In some embodiments, a new cell category can be generated by applying Boolean logic to multiple categorical model outputs. In some embodiments, cells that are positive for marker 1 and negative for marker 2 can be placed in a new, separate category.
[0129] In some embodiments, the methods, systems, and media disclosed herein may further comprise combining a plurality of classifications to generate a composite classification for the image. In some embodiments, the methods, systems, and media disclosed herein may further comprise processing the image of the plurality of cells to score the plurality of cells based on a different pre-defined criterion among the plurality of pre-defined criteria. In some embodiments, the methods, systems, and media disclosed herein may further comprises combining a plurality of scores to generate a composite score for the image. The output of a given model is processed so that the area of each cell is identified and separated from every other cell. For each cell, statistical analysis is then performed using that first model as well as the other models. The region in the image that corresponds to a given cell is then categorized, scored, or subjected to statistical analysis based on the output of the previously mentioned machine learning models.
[0130] Statistical analysis is performed on outputs from machine learning models to compare the performance and / or predictions of different models to assess their relative strengths and identify significant differences. In some embodiments, the methods, systems, and media disclosed herein further comprise defining performance metrics. In some embodiments, performance metrics are evaluated based on the problem domain for each model on accuracy, precision, recall, Fl score, area under the curve (AUC), mean squared error (MSE), or any other relevant features.
[0131] In some embodiments, the methods, systems, and media disclosed herein further comprise data splitting to ensure fair evaluation and avoid overfitting. In some embodiments, dataset are divided into training and testing sets. In some embodiments, data splitting adopts a cross-validation strategy to evaluate on unseen data.
[0132] In some embodiments, the methods, systems, and media disclosed herein further comprise performing statistical tests to verify dependency or independency of outputs from multiple models. In some embodiments, statistical tests to verify dependency or independency of outputs from multiple models comprises paired T-test, independent T-test, analysis of variance (ANOVA), Wilcoxon signed-rank test, or Kruskal-Wallis test.
[0133] In some embodiments, the methods, systems, and media disclosed herein further comprise interpreting results based on the p-value obtained from the statistical test. If the p- value is below a pre-determined significance level (e.g., 0.05), there is a statistically significant difference in performance between the models. If the p-value is above the significance level, there is not enough evidence to claim a significant difference. While p- values provide information about statistical significance, it is also essential to consider effect sizes and confidence intervals. In some embodiments, the methods, systems, and media disclosed herein further comprise evaluating effect size and confidence intervals. Effect sizes quantify the magnitude of the observed differences. Confidence intervals provide a range of plausible values for the true difference. In some embodiments, the methods, systems, and media disclosed herein further comprise performing validation on the results to provide additional insights and strengthen the reliability of the results. In some embodiments, validation comprises cross-validation, bootstrapping, or permutation testing.
[0134] The present disclosure provides methods and systems for presenting outputs of multiple machine-learning models in a single output image. Multiple statistical metrics from more than one model are then used to alter the color scheme of the cell or surrounding tissue in an output image so that the color reflects how the cell is scored, categorized, or otherwise reflects the statistical metric. These metrics are combined and represented using statistical diagrams. Statistical analysis can be carried out on the predicted metrics along with geospatialcluster analysis using the coordinates of each cell, which can also allow to cluster cells for interpretive analysis. The resulting diagrams can help select thresholds for the various metrics to evaluate different cell populations. In some embodiments, the aggregate of this data can be used to generate predictions about the entire H&E image such as prognosis, diagnostic categorization, molecular or proteomic expression, or response to therapy. Finally, thresholding or clustering can be applied to the resulting data and can be used to segment out a cell population, which can then be projected geospatially onto the original image or a modified version of the original image.
[0135] The present disclosure provides methods and systems for presenting outputs of multiple machine-learning models in a single output image. In some embodiments, the methods, systems, and media disclosed herein further comprise performing flow cytometry analysis based at least in part on analyzing the single image. In some embodiments, the methods, systems, and media disclosed herein further comprise analyzing the image using a technique comprising multiplexed imaging system, high parameter spatial multi-omic system, multiplexed ion beam imaging by time of flight (MIBI-TOF), imaging mass cytometry (IMC), matrix-assisted laser desorption ionization mass spectrometry imaging (MALDI-MSI), digital spatial analysis (DSP), or geospatial clustering analysis.
[0136] Imaging flow cytometry with multiplexing capabilities provides the advantages of high-throughput analysis combined with detailed imaging information. It allows for the simultaneous examination of multiple cellular parameters, providing a comprehensive understanding of cell populations and their heterogeneity. Flow cytometry is used for the analysis of individual cells in a suspension. It involves the labeling of cells with fluorescent markers, followed by their detection and quantification as they pass through a flow cytometer. The present disclosure provides methods and systems for presenting outputs of multiple machine-learning models in a single output multiplexing image. Flow cytometry is applied to simultaneous detecting multiple parameters or markers within the single image with outputs of multiple machine-learning models. Thus, imaging flow cytometry combines the principles of flow cytometry and microscopy to provide both high-throughput analysis and imaging capabilities.
[0137] The high-parameter spatial multi-omic system using multiplexing imaging may provide a comprehensive approach for studying complex biological systems, such as tissues or tumors, by simultaneously capturing multiple molecular features with spatial context. The present disclosure provides methods and systems for presenting outputs of multiple machinelearning models in a single output multiplexing image. Applying high-parameter spatial multi- omic system on multiplexing imaging with outputs of multiple machine-learning modelsenables the simultaneous detection of multiple molecular features within the context of tissue samples. This advanced system combines imaging, multiplexing, and omics techniques to generate comprehensive spatial and molecular information. The single image with outputs of multiple machine-learning models is subjected to image analysis algorithms specifically designed for high-parameter multiplexed data. These algorithms identify and separate individual signals corresponding to different molecular targets based on their fluorescence properties, spatial localization, and other characteristics. The algorithms assign spatial coordinates to each signal, creating a spatial map of the detected molecular features. The spatial map of molecular features is integrated with omic data, such as genomics, transcriptomics, or proteomics data. This integration can be achieved by overlaying the spatial map with omic data obtained from the same tissue sample.
[0138] Multiplexed ion beam imaging by time of flight (MIBI-TOF) with highdimensional imaging enables the simultaneous detection of multiple molecular targets with high resolution and sensitivity. MIBI-TOF provides valuable insights into the spatial organization of cellular components, signaling pathways, and microenvironmental interactions within tissues, facilitating a deeper understanding of complex biological processes and disease mechanisms. The multiplexing capability, for example antibody labels, allows for the assessment of complex cellular phenotypes and the exploration of interrelationships between different biomarkers. The present disclosure provides methods and systems for presenting outputs of multiple machine-learning models in a single output multiplexing image. The acquired data is processed and reconstructed into images that represent the spatial distribution of the biomarkers. Each pixel in the resulting images represents a specific m / z value corresponding to a labeled antibody.
[0139] Digital Spatial Analysis (DSP) utilizes multiplexing imaging data to extract spatial information and perform quantitative analysis of biomarkers or molecular features within tissues. DSP allows for the quantitative analysis and interpretation of multiple biomarkers or molecular features within their spatial context, providing insights into cellular interactions, tissue organization, disease heterogeneity, and potential diagnostic or therapeutic implications. The present disclosure provides methods and systems for presenting outputs of multiple machine-learning models in a single output multiplexing image. DSP involves the segmentation of tissue regions or individual cells within the single output multiplexed image. This segmentation process separates different regions or cell boundaries to isolate specific areas of interest for analysis. Following segmentation, features of interest are extracted, which can include intensity, texture, shape, or other quantitative measurements related to the biomarkers or molecular targets. DSP leverages the spatial information within the multiplexedimage to analyze the relationships and interactions between different biomarkers or molecular features. This analysis can involve quantifying co-localization, proximity, or spatial patterns of expression between different targets. Various statistical and computational methods are employed to assess the spatial relationships and identify significant patterns or correlations.
[0140] Geospatial clustering analysis provides a process for identifying clusters or groups of spatially related objects or features within a geographic or spatial dataset. Geospatial clustering analysis can also be applied to multiplexing images to identify spatially related patterns or clusters of features. The present disclosure provides methods and systems for presenting outputs of multiple machine-learning models in a single output multiplexing image. This allows for the identification and exploration of spatial patterns or clusters of features within complex biological samples. It also enables to discover and interpret spatial relationships between different biomarkers or molecular targets, which will facilitate the understanding of tissue organization, cellular interactions, disease heterogeneity, and potential biological implications. Geospatial clustering analysis starts with a single output multiplexing image. The image can represent different biomarkers, molecular targets, or other features that are labeled with distinct labels, such as fluorophores. Relevant features are extracted from the multiplexing images. These features can be derived from intensity values, texture information, or other quantitative measures associated with the labeled biomarkers or molecular targets. The feature extraction process aims to represent the characteristics of the objects or features in a numerical form. Geospatial clustering algorithms are applied to the extracted features to identify spatially related patterns or clusters. These clustering algorithms consider both the feature values and the spatial proximity of the objects or features within the single output multiplexing image. Examples of clustering algorithms commonly used in geospatial analysis include k-means clustering, hierarchical clustering, DBSCAN (Density -Based Spatial Clustering of Applications with Noise), or spatially-constrained clustering algorithms. The resulting clusters are evaluated and validated using appropriate metrics to assess the quality and significance of the clustering results. Various visualization techniques can be employed to represent the identified clusters within the single output multiplexing image. This can include color coding or overlays to highlight the clusters' spatial distribution and relationships with the labeled biomarkers or molecular targets. The identified clusters can be further analyzed and interpreted based on their spatial patterns, associations with specific biomarkers or molecular targets, or other relevant contextual information. This analysis may involve exploring the characteristics or functional implications of the clusters, investigating their relationship with clinical outcomes or experimental conditions, or integrating the clustering results with additional data sources for comprehensive insights.
[0141] Matrix- Assisted Laser Desorption Ionization Mass Spectrometry Imaging (MALDI-MSI) is an advanced technique used for spatially resolved analysis of molecules within tissues. MALDI-MSI allows for the multiplexed analysis of molecules within tissues, providing spatially resolved information about their distribution and abundance. The present disclosure provides methods and systems for presenting outputs of multiple machine-learning models in a single output multiplexing image. Then mass spectra represent the molecular composition of the analytes at each pixel location, providing a spatially resolved dataset. The acquired mass spectra are compared to reference databases based on outputs of multiple machine-learning models, allowing for the identification and annotation of the detected molecules. The identified molecules and their corresponding intensities can be spatially mapped onto the tissue section. By correlating the molecular information with the spatial location, a multiplexed image can be generated. Each pixel in the image represents the intensity or abundance of a specific molecule or ion at that particular location.
[0142] Computer systems
[0143] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 1 shows a computer system 101 that is programmed or otherwise configured to, for example, obtain an image of cells; obtain trained machine learning models processing images of cells to classify or score cells, identify a region of interest in cells, and generate a visualization of cells.
[0144] The computer system 101 can regulate various aspects of analysis, calculation, and generation of the present disclosure, such as, for example, obtaining an image of cells; obtaining trained machine learning models processing images of cells to classify or score cells, identifying a region of interest in cells, and generating a visualization of cells. The computer system 101 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.
[0145] The computer system 101 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 105, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 101 also includes memory or memory location 110 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 115 (e.g., hard disk), communication interface 120 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 125, such as cache, other memory, data storage and / or electronic display adapters. The memory 110, storage unit 115, interface 120 and peripheral devices 125 are in communication with the CPU 105 through a communication bus (solid lines), such as a motherboard. The storage unit 115 can be a data storage unit (or data repository) for storing data. The computersystem 101 can be operatively coupled to a computer network (“network”) 130 with the aid of the communication interface 120. The network 130 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet.
[0146] In some embodiments, the network 130 is a telecommunication and / or data network. The network 130 can include one or more computer servers, which can enable distributed computing, such as cloud computing. For example, one or more computer servers may enable cloud computing over the network 130 (“the cloud”) to perform various aspects of analysis, calculation, and generation of the present disclosure, such as, for example, obtaining an image of cells; obtaining trained machine learning models processing images of cells to classify or score cells, identifying a region of interest in cells, and generating a visualization of cells. Such cloud computing may be provided by cloud computing platforms such as, for example, Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM cloud. In some embodiments, the network 130, with the aid of the computer system 101, can implement a peer-to-peer network, which may enable devices coupled to the computer system 101 to behave as a client or a server.
[0147] The CPU 105 may comprise one or more computer processors and / or one or more graphics processing units (GPUs). The CPU 105 can execute a sequence of machine- readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 110. The instructions can be directed to the CPU 105, which can subsequently program or otherwise configure the CPU 105 to implement methods of the present disclosure. Examples of operations performed by the CPU 105 can include fetch, decode, execute, and writeback.
[0148] The CPU 105 can be part of a circuit, such as an integrated circuit. One or more other components of the system 101 can be included in the circuit. In some embodiments, the circuit is an application specific integrated circuit (ASIC).
[0149] The storage unit 115 can store files, such as drivers, libraries, and saved programs. The storage unit 115 can store user data, e.g., user preferences and user programs. In some embodiments, the computer system 101 can include one or more additional data storage units that are external to the computer system 101, such as located on a remote server that is in communication with the computer system 101 through an intranet or the Internet.
[0150] The computer system 101 can communicate with one or more remote computer systems through the network 130. For instance, the computer system 101 can communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab, Microsoft® Surface), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 101 via the network 130.
[0151] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 101, such as, for example, on the memory 110 or electronic storage unit 115. The machine executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 105. In some embodiments, the code can be retrieved from the storage unit 115 and stored on the memory 110 for ready access by the processor 105. In some situations, the electronic storage unit 115 can be precluded, and machine-executable instructions are stored on memory 110.
[0152] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.
[0153] Embodiments of the systems and methods provided herein, such as the computer system 101, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, or disk drives, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical, and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links, or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0154] Hence, a machine readable medium, such as computer-executable code, may take many forms, including a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0155] The computer system 101 can include or be in communication with an electronic display 135 that comprises a user interface (UI) 140 for providing, for example, (i) a visual display indicative of training and testing of a trained algorithm, (ii) a visual display of data indicative of a status of a pre-defined criteria, (iii) a quantitative measure of a status of a predefined criteria, or (iv) a visual display of outputs of machine-learning models. Examples of UIs include, without limitation, a graphical user interface (GUI) and web-based user interface.
[0156] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 205. The algorithm can, for example, obtain an image of cells; obtain trained machine learning models processing images of cells to classify or score cells, identify a region of interest in cells, and generate a visualization of cells.EXAMPLES
[0157] Example 1: Virtual Multiplexing: Triple negative breast cancer
[0158] Breast cancer photomicrographs were processed by 3 virtual IHC models: estrogen receptor (ER), progesterone receptor (PR), and HER2. The nuclei and peri-nuclear regions of cells were re-colored based on the predicted expression of the various virtualbiomarkers. A prediction of no staining expression was indicated by coloring the nuclei blue. A prediction of ER expression was indicated by coloring the nuclei pink. A prediction of PR expression was indicated by coloring the nuclei green.
[0159] FIG. 2A illustrates an example of virtual multiplexing of mostly negative staining in the lesional cells. FIGs. 2B-2D show images of physical immunohistochemistry of a sample for negative HER2 (FIG. 2B), negative estrogen receptor (ER) (FIG. 2C), and negative progesterone receptor (PR) (FIG. 2D) corresponding to the same region of FIG. 2A.
[0160] Example 2: Virtual Multiplexing: Her2 positive, ER / PR negative
[0161] Breast cancer photomicrographs were processed by 3 virtual IHC models: estrogen receptor (ER), progesterone receptor (PR), and HER2. The nuclei and peri-nuclear regions of cells were re-colored based on the predicted expression of the various virtual biomarkers. A prediction of HER2 expression was indicated by coloring the perinuclear region roughly corresponding to the cell cytoplasm as orange. For HER2 expression, the prediction of weak, moderate, and strong staining intensity is represented by the orange color intensity. A prediction of ER expression was indicated by coloring the nuclei pink. A prediction of PR expression was indicated by coloring the nuclei green.
[0162] FIG. 3A illustrates an example of virtual multiplexing of a sample with HER2 positive breast cancer. HER2 positive breast cancer is highlighted by orange peri-nuclear staining in the circle area. A benign breast cancer lobule is also presented as highlighted with green nuclei as the arrow pointing. Occasional ER-positive lobular cells are also presented with pink / red expression. FIGs. 3B-3D show images of physical immunohistochemistry of a sample depicting the same regions of FIG. 3A.
[0163] Example 3: Virtual Multiplexing: Co-localized ER / PR and Her2 negative
[0164] Breast cancer photomicrographs were processed by 3 virtual IHC models: estrogen receptor (ER), progesterone receptor (PR), and HER2. The nuclei and peri-nuclear regions of cells were re-colored based on the predicted expression of the various virtual biomarkers. A prediction of co-expression of both ER and PR was indicated by coloring the nuclei dark red.
[0165] FIG. 4A illustrates an example of virtual multiplexing of a sample for invasive breast cancer with predominantly co-expression of ER and PR as the nuclei are highlighted in red and HER2 negative. FIGs. 4B-4D show images of physical immunohistochemistry of a sample of negative HER2 (FIG. 4B) and positive ER (FIG. 4C) and positive PR (FIG. 4D) from the same regions of FIG. 4A.
[0166] Example 4: Virtual Flow of prediction of cell densities for ER, PR, and Her2
[0167] Breast cancer photomicrographs were processed by 3 virtual IHC models: estrogen receptor (ER), progesterone receptor (PR), and HER2. The regions corresponding to each cell nuclei were obtained by first applying a low pass filter to the background prediction channel and performing a connected components analysis on the resulting binary image.
[0168] For ER and PR virtual biomarkers in which the prediction was binary (positive or negative), the mean prediction score was obtained over the area corresponding to each cell.
[0169] For HER2 in which there were 4 possible cell categories (negative, weak, moderate, and strong), a composite score was calculated. The mean prediction scores for each category was obtained over the area corresponding to each cell. These were then multiplied by 0, 0.33, 0.67, and 1 for each category and then summed to produce the composite score.
[0170] The results were then displayed in a topographical map in order to highlight the relationship between the virtual biomarkers in the various cell populations as shown in FIG. 5.
[0171] Example 5: Automated image registration based on nuclei location
[0172] Registering IHC and H&E images can be challenging, even with multiple coarse and fine registration passes and using a patch-wise approach. The difficulty arises due to the different color spaces of the images, non-homogeneity of the changes in colors of the various features in the images, and tissue warping during the re-staining process. Not all features in an image require registration, and in some cases, the location of cell nuclei is the only relevant information needed for annotation. Using methods and systems of the present disclosure, automated image registration is performed based on nuclei location.
[0173] An ML model is utilized to identify the nuclei’s location in the H&E image, and employ either the same or a separate ML model to identify nuclei in the corresponding IHC, ISH, special stain, or other modality image. This generates two point clouds that can be registered more easily compared to arbitrary features obtained through off the shelf registration techniques. The registration characteristics obtained from the point clouds enables an option to warp the corresponding IHC or H&E images using deformable or non-deformable transformations. Alternatively, this approach can be used to annotate the H&E cells by directly referencing the corresponding cell nucleus location and image in the IHC / ISH / special stain or other relevant image without warping the image itself.
[0174] Example 6: Alternative annotation approaches
[0175] Using methods and systems of the present disclosure, as an alternative to manually registering images, a generative adversarial network (GAN) is trained that can create synthetic H&E images using IHC images as templates for the input. The IHC image is then converted to a synthetic H&E image, which is already registered. These images are then processed, e.g., as described elsewhere herein. Alternatively, a synthetic IHC image isgenerated from the H&E image. The resulting pair is then processed, e.g., as described elsewhere herein. This approach may result in reduced costs of procuring the dataset, which may outweigh any artifacts that may be introduced with the creation of synthetic H&E or IHC images.
[0176] Example 7: Alternative annotation approaches
[0177] Using methods and systems of the present disclosure, the annotations created either directly through the staining / re-staining method or the output of the virtual IHC model can be used as a surrogate for alternative annotations in a 2-step process.
[0178] For example, an IHC marker specific to a population of cancer cells can be used to first identify the location of each cancer cell within a set of images from a plurality of patient samples. The dataset can then be divided into two groups: lesions with a given molecular aberration and lesions without a given molecular aberration. As molecular aberrations in cancer are usually only present in the cancer cell population, the dataset can then be used train an Al model to differentiate between non-cancer cells, cancers cells with the molecular aberration, and cancer cells without the aberration. Because it performs this prediction on a per-cell basis, this approach allows for a considerable margin of error.
[0179] Example 8: Pharmaceutical and research studies
[0180] Using methods and systems of the present disclosure, pharmaceutical and research studies may comprise performing rapid, large scale analysis of numerous tissue slides. A business analytics dashboard can then be used wherein statistical analysis from individual cell staining from one or multiple biomarker is combined with the geospatial location of the cells within the lesion if needed, and is used for, as an example, stratifying patients in clinical trials, or to investigate and characterize which population of patients respond to treatments.
[0181] Multiple virtual biomarkers can also be displayed simultaneously using the multiplex functionality for the discovery of new intra-cellular interactions, differentiation of cell subpopulations, and improved prognostication. For example, tumor infiltrating lymphocytes can automatically be detected, quantified, and their geospatial relationship to tumor cells can be analyzed.
[0182] Example 9: Complex biomarker profiling
[0183] Many diseases and biological processes involve the dysregulation or interplay of multiple molecular targets. Using methods and systems of the present disclosure, virtual multiplexing approaches enable the assessment of complex biomarker profiles by examining the expression of multiple proteins simultaneously. This can provide a more comprehensive understanding of disease mechanisms, heterogeneity, and treatment responses.
[0184] Example 10: Co-expression analysis
[0185] Using methods and systems of the present disclosure, multiplexing IHC enables the analysis of co-expression patterns, by visualizing multiple markers within the same tissue section. This helps identify subpopulations of cells or disease-specific molecular signatures, leading to insights into cellular phenotypes, functional states, and potential therapeutic targets. Such analysis is vital for the diagnosis of, for example, hematolymphoid neoplasms.Pathologists currently rely on flow cytometry to achieve this goal.
[0186] Example 11: Pharmaceutical applications of virtual IHC
[0187] Using methods and systems of the present disclosure, various pharmaceutical applications of virtual IHC are performed, such as patient triaging, companion diagnostics, and sample enrichment.
[0188] For example, patient triaging may be performed. Several precision therapeutics use Next Generation Sequencing as the diagnostic to identify the right patient, which in itself is expensive on a per patient basis. Using methods and systems of the present disclosure upstream to this allows one to select down from the total patient pool, allowing better per patient justification of using NGS.
[0189] As another example, virtual biomarkers may be applied to develop companion diagnostics. A single or multiple markers simultaneously may be used to predict response to specific therapeutic interventions. Such efforts may be coordinated with a pharmaceutical company during drug development and drug testing.
[0190] As another example, sample enrichment may be performed. By more precisely characterizing the tissue samples based on objective virtual IHC data, potentially complementary and parallel to physical IHC biomarkers, allows one to better characterize study samples. This aims to reduce the risk early in drug development by obtaining a more granular description of the tumors.
[0191] Example 12: Virtual IHC markers for evaluating different disease types and fundamental biological sciences
[0192] Using methods and systems of the present disclosure, various virtual IHC markers may be used for evaluating different disease types and fundamental biological sciences.
[0193] For example, research on developmental biology may comprise analysis of one or more of the following virtual IHC markers: Oct4, Sox2, Nanog, Brachyury, Pax6, MyoD, Nestin, Ki67, Neurofilament, Beta III Tubulin, Desmin, Smooth muscle actin, and Vimentin.
[0194] As another example, neuroscientific studies may comprise analysis of one or more of the following virtual IHC markers: GFAP, NeuN, Synaptophysin, MAP2, TyrosineHydroxylase, GABA, 5-HT (Serotonin), DARPP-32, and Calbindin D.
[0195] As another example, immunological research may comprise analysis of one or more of the following virtual IHC markers: CD3, CD4, CD8, CD20, CD68, CD45, Foxp3, Granzyme B, Interferon-gamma, IL-10, and IL-17.
[0196] As another example, pharmacological research may comprise analysis of one or more of the following virtual IHC markers: Phospho-ERKl / 2, Phospho- Akt, Phospho-p38, Phospho- JNK, Cyclin DI, PCNA, Bcl-2, Cleaved caspase-3, VEGF, EGFR, PD-L1, HER2, PDGFR, CDK4, and PARP.
[0197] As another example, immuno-oncology studies may comprise analysis of one or more of the following virtual IHC markers: CD4, CD4, CD8, CD31, CD45, FoxP3 / Treg, PD- Ll, PD-1, CTLA4, F4 / 80, IBA1, and Ki67.
[0198] As another example, toxicology and drug safety assessment studies may comprise analysis of one or more of the following virtual IHC markers: CD3, CD4, CD8, CD20, CD45, CYP1A1, CYP2E1, GSTP1, Nrf2, F4 / 80, HO-1, Bax, p53, Cleaved caspase-9, Ki67, Albumin, KIM-1, AQP2, TGF-betal, Collagen IV, and Alpha-SMA.
[0199] As another example, investigation of infectious diseases may comprise analysis of one or more of the following virtual IHC markers: HIV p24, Influenza A / B nucleoprotein, CMV pp65, Hepatitis B surface antigen, Chlamydia trachomatis LPS, Treponema pallidum, Legionella pneumophila, Mycobacterium tuberculosis antigen, Toxoplasma gondii antigen, Candida albicans antigen, Pneumocystis jirovecii, HSV (Herpes Simplex Virus), EBV (Epstein-Barr Virus), HBV (Hepatitis B Virus), HCV (Hepatitis C Virus), SARS-CoV-2 (Severe Acute Respiratory Syndrome Coronavirus 2), Histoplasma capsulatum, and Aspergillus spp.
[0200] As another example, veterinary pathology studies may comprise analysis of one or more of the following virtual IHC markers: VimH, CD3, CD79a, CD117 (c-Kit), S100, Myosin, Desmin, SMA (Smooth Muscle Actin), GFAP, Ki67, Melan-A, CD31, CD34, E- cadherin, Pancytokeratin, CK7, CK20, and Calretinin.
[0201] As another example, studies of cancer may comprise analysis of one or more of the following virtual IHC markers: ER (Estrogen Receptor), PR (Progesterone Receptor), HER2 (Human Epidermal Growth Factor Receptor 2), Ki67, p53, Bcl-2, EGFR (Epidermal Growth Factor Receptor), CK7 (Cytokeratin 7), CK20 (Cytokeratin 20), PDL-1 (Programmed Death-Ligand 1), BRAF, ALK (Anaplastic Lymphoma Kinase), CDX2 (Caudal Type Homeobox 2), Melan-A, S100, CD45, CD20, CD3, CD5, CD30, CD15, CD99, WT1 (Wilms' Tumor 1), MSH2 (MutS Homolog 2), MSH6, MLH1 (MutL Homolog 1), PMS2, TTF-1(Thyroid Transcription Factor 1), Napsin A, GATA3, CAIX (Carbonic Anhydrase IX), and HMB-45.
[0202] As another example, studies of autoimmune diseases may comprise analysis of one or more of the following virtual IHC markers: ANA (Antinuclear Antibodies), ENA (Extractable Nuclear Antigens), dsDNA (Double-Stranded DNA), Ro (SS-A), La (SS-B), Scl- 70 (Anti-Topoisomerase I), Jo-1 (Histidyl-tRNA Synthetase), C3, C4, IgG, IgA, IgM, CD3, CD20, CD4, CD8, CD68, and CD31.
[0203] As another example, studies of gastrointestinal diseases may comprise analysis of one or more of the following virtual IHC markers: CDX2 (Caudal Type Homeobox 2), CK20 (Cytokeratin 20), CK7 (Cytokeratin 7), CD 10, CDH17 (Cadherin-17), Villin, Chromogranin A, Synaptophysin, SI 00, Mucin (e.g., MUC1, MUC2, MUC5AC), and CD34.
[0204] As another example, studies of hematological diseases may comprise analysis of one or more of the following virtual IHC markers: CD20, CD3, CD5, CD10, CD23, CD30, CD 15, CD99, CD45, CD79a, CD 138, PAX5 (Paired Box 5), Bcl-2, Bcl-6, Cyclin DI, MUM1 (Multiple Myeloma Oncogene 1), ALK (Anaplastic Lymphoma Kinase), TdT (Terminal Deoxynucleotidyl Transferase), CD117 (c-Kit), CD34, CD43, CD56, CD7,and CD79a,.
[0205] As another example, studies of respiratory diseases may comprise analysis of one or more of the following virtual IHC markers: TTF-1 (Thyroid Transcription Factor 1), Napsin A, P40, CK5, CK7 (Cytokeratin 7), Chromogranin A, Synaptophysin, CD56, PDL-1 (Programmed Death-Ligand 1), CDX2 (Caudal Type Homeobox 2), CD45, CD20, and CD3.
[0206] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A computer-implemented method of performing multiplex image analysis of a plurality of cells, comprising:(a) obtaining an image of said plurality of cells;(b) obtaining a plurality of trained machine learning models, each of said plurality of trained machine learning models configured to analyze image data of cells to classify or score said cells based on a different pre-defined criterion among a plurality of pre-defined criteria;(c) for each of said plurality of trained machine learning models, processing said image of said plurality of cells to classify or score said plurality of cells based on a different pre-defined criterion among said plurality of pre-defined criteria;(d) identifying a region of interest in said image of said plurality of cells based at least in part on said classifying or said scoring in (c); and(e) generating, based at least in part on said region of interest identified in (d), a visualization of said plurality of cells that distinguishes each of said different pre-defined criteria.
2. The method of claim 1, wherein said image of said plurality of cells comprises a plurality of images.
3. The method of claim 2, wherein said plurality of images comprise a time series of images.
4. The method of claim 1, wherein said image of said plurality of cells is a single image.
5. The method of claim 3, further comprising performing flow cytometry analysis based at least in part on analyzing said single image.
6. The method of claim 1, further comprising analyzing the image using a technique comprising multiplexed imaging system, high parameter spatial multi-omic system, multiplexed ion beam imaging by time of flight (MIBI-TOF), imaging mass cytometry (IMC), matrix-assisted laser desorption ionization mass spectrometry imaging (MALDI- MSI), digital spatial analysis (DSP), or geospatial clustering analysis.
7. The method of claim 1, wherein said plurality of cells comprise primary cells, stem cells, immune cells, carcinoma cells, sarcoma cells, lymphoma cells, germ cell tumor cells,blastoma cells, bladder cancer cells, breast cancer cells, colon cancer cells, colorectal cancer cells, endocrine tumor cells, esophageal cancer cells, glioblastoma cells, Hodgkin lymphoma cells, lung cancer cells, melanoma cells, or prostate cancer cells.
8. The method of claim 1, wherein said plurality of cells are obtained at least in part by one or more of: biopsy collection, surgical resection, xenograft, animal model, fine needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, neoplastic tissue sampling, malignant tissue sampling, diseased tissue sampling, and implanted tissue sampling.
9. The method of claim 1, wherein said plurality of pre-defined criteria comprise gene expression, cellular structure, cellular morphology, next-generation sequencing (NGS), proteomic features, biochemical features, genomic features, transcriptomic features, methylation features, metabolomic features, metagenomic features, phenomics features, phenotypic features, or signaling pathway features.
10. The method of claim 1, wherein (c) further comprises processing said image of said plurality of cells to classify said plurality of cells based on a different pre-defined criterion among said plurality of pre-defined criteria.
11. The method of claim 9, wherein (c) further comprises combining a plurality of classifications to generate a composite classification for said image.
12. The method of claim 1, wherein (c) further comprises processing said image of said plurality of cells to score said plurality of cells based on a different pre-defined criterion among said plurality of pre-defined criteria.
13. The method of claim 11, wherein (c) further comprises combining a plurality of scores to generate a composite score for said image.
14. The method of claim 1, wherein (d) further comprises providing said visualization of said plurality of cells using a plurality of different colors indicative of each of said different pre-defined criteria.
15. The method of claim 1, wherein the method comprises use of fluorescence.
16. The method of claim 1, wherein the method is performed without use of fluorescence.
17. The method of claim 1, wherein the method comprises use of immunohistochemistry staining.
18. The method of claim 1, wherein said multiplex image analysis is performed without performing physical assays on said plurality of cells.
19. The method of claim 18, wherein said multiplex image analysis is performed without use of immunohistochemistry staining.
20. The method of claim 1, wherein output of said machine learning model is based on a predefined criterion of a biomarker of a plurality of biomarkers.
21. The method of claim 1, wherein said trained machine learning models comprise a supervised machine learning model, an unsupervised machine learning model, a deep learning model, or a time-series machine learning model.
22. The method of claim 1, wherein said image of said plurality of cells comprises surrounding tissue of said plurality of cells.
23. A computer-implemented method of performing image analysis of a plurality of cells, comprising:(a) obtaining an image of said plurality of cells, wherein said image is obtained without performing physical assays on said plurality of cells;(b) obtaining a trained machine learning model configured to analyze image data of cells to classify or score said cells based on a plurality of pre-defined criteria;(c) processing said image of said plurality of cells to classify or score said plurality of cells based on said plurality of pre-defined criteria;(d) identifying a region of interest in said image of said plurality of cells based at least in part on said classifying or said scoring in (c); and(e) generating, based at least in part on said region of interest identified in (d), a visualization of said plurality of cells that distinguishes each of said different pre-defined criteria.
24. The method of claim 23, wherein said image of said plurality of cells comprises a plurality of images.
25. The method of claim 24, wherein said plurality of images comprise a time series of images.
26. The method of claim 23, wherein said image of said plurality of cells is a single image.
27. The method of claim 25, further comprising performing flow cytometry analysis based at least in part on analyzing said single image.
28. The method of claim 23, further comprising analyzing the image using a technique comprising multiplexed imaging system, high parameter spatial multi-omic system,multiplexed ion beam imaging by time of flight (MIBI-TOF), imaging mass cytometry (IMC), matrix-assisted laser desorption ionization mass spectrometry imaging (MALDI- MSI), digital spatial analysis (DSP), or geospatial clustering analysis.
29. The method of claim 23, wherein said plurality of cells comprise primary cells, stem cells, immune cells, carcinoma cells, sarcoma cells, lymphoma cells, germ cell tumor cells, blastoma cells, bladder cancer cells, breast cancer cells, colon cancer cells, colorectal cancer cells, endocrine tumor cells, esophageal cancer cells, glioblastoma cells, Hodgkin lymphoma cells, lung cancer cells, melanoma cells, or prostate cancer cells.
30. The method of claim 23, wherein said plurality of cells are obtained at least in part by one or more of: biopsy collection, surgical resection, xenograft, animal model, fine needle aspiration, peripheral blood collection, bone marrow biopsies, healthy tissue sampling, neoplastic tissue sampling, malignant tissue sampling, diseased tissue sampling, and implanted tissue.
31. The method of claim 23, wherein said plurality of pre-defined criteria comprise gene expression, cellular structure, cellular morphology, next-generation sequencing (NGS), proteomic features, biochemical features, genomic features, transcriptomic features, methylation features, metabolomic features, metagenomic features, phenomics features, phenotypic features, or signaling pathway features.
32. The method of claim 23, wherein (c) further comprises processing said image of said plurality of cells to classify said plurality of cells based on a different pre-defined criterion among said plurality of pre-defined criteria.
33. The method of claim 23, wherein (c) further comprises combining a plurality of classifications to generate a composite classification for said image.
34. The method of claim 23, wherein (c) further comprises processing said image of said plurality of cells to score said plurality of cells based on a different pre-defined criterion among said plurality of pre-defined criteria.
35. The method of claim 34, wherein (c) further comprises combining a plurality of scores to generate a composite score for said image.
36. The method of claim 23, wherein (d) further comprises providing said visualization of said plurality of cells using a plurality of different colors indicative of each of said different pre-defined criteria.
37. The method of claim 23, wherein said image analysis is performed without use of fluorescence.
38. The method of claim 23, wherein said image analysis is performed without use of immunohistochemistry staining.
39. The method of claim 23, wherein output of said machine learning model is based on a pre-defined criterion of a biomarker of a plurality of biomarkers.
40. The method of claim 23, wherein said trained machine learning models comprise a supervised machine learning model, an unsupervised machine learning model, a deep learning model, or a time-series machine learning model.
41. The method of claim 23, wherein said image of said plurality of cells comprises surrounding tissue of said plurality of cells.
42. A computer-implemented system comprising one or more processors and a computer memory having machine-executable instructions stored thereon, wherein said machineexecutable instructions cause said one or more processors to perform operations comprising:(a) obtaining an image of said plurality of cells;(b) obtaining a plurality of trained machine learning models, each of said plurality of trained machine learning models configured to analyze image data of cells to classify or score said cells based on a different pre-defined criterion among a plurality of pre-defined criteria;(c) for each of said plurality of trained machine learning models, processing said image of said plurality of cells to classify or score said plurality of cells based on a different pre-defined criterion among said plurality of pre-defined criteria;(d) identifying a region of interest in said image of said plurality of cells based at least in part on said classifying or said scoring in (c); and(e) generating, based at least in part on said region of interest identified in (d), a visualization of said plurality of cells that distinguishes each of said different pre-defined criteria.
43. A non-transitory computer-readable storage medium encoded with instructions executable by one or more processors to cause the one or more processors to perform operations comprising:(a) obtaining an image of said plurality of cells;(b) obtaining a plurality of trained machine learning models, each of said plurality of trained machine learning models configured to analyze image data of cells to classify or score said cells based on a different pre-defined criterion among a plurality of pre-defined criteria;(c) for each of said plurality of trained machine learning models, processing said image of said plurality of cells to classify or score said plurality of cells based on a different pre-defined criterion among said plurality of pre-defined criteria;(d) identifying a region of interest in said image of said plurality of cells based at least in part on said classifying or said scoring in (c); and(e) generating, based at least in part on said region of interest identified in (d), a visualization of said plurality of cells that distinguishes each of said different pre-defined criteria.