Methods and systems for multimodal subcellular segmentation
The novel cell segmentation method combines morphological markers and spatial transcriptomics data to improve cell boundary information, addressing inaccuracies in existing methods and enhancing the precision of spatial analyses by using photocleavable markers and readout density maps.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BRUKER SPACIAL BIOLOGY INC
- Filing Date
- 2024-06-12
- Publication Date
- 2026-07-29
AI Technical Summary
Existing cell segmentation methods are inaccurate and highly dependent on specific experimental conditions and tissue types, often failing to provide reliable cell membrane labeling due to variability in morphological marker expression, leading to missed or mis-segmented cells in spatial transcriptomics analyses.
A novel method for cell segmentation that utilizes a combination of morphological markers and spatial transcriptomics data, including photocleavable markers, to generate a readout density map, and performs segmentation based on this map and multiple morphological markers, using 3D scanned images in high dynamic range mode, with optional 2D reduction and deconvolution techniques.
Enhances cell boundary information, providing precise cell segmentation across various tissue types, improving the accuracy of spatial analyses such as spatial transcriptomics by leveraging additional morphological markers and readout density maps.
Smart Images

Figure 2026525377000001_ABST
Abstract
Description
cross reference
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 507,824, filed on 13 June 2023, which is incorporated herein by reference in its entirety, and this application claims priority under 35 U.S.C. 120. [Background technology]
[0002] Cell segmentation is a crucial step in many biological and medical analyses, providing a vital foundation for spatial analysis that leverages the spatial organization and characteristics of cells within tissue samples. Precise cell segmentation allows for the study of spatial relationships and patterns, aiding in the understanding of tissue composition, cell communication, or disease progression. Existing cell segmentation methods are often inaccurate and highly dependent on specific experimental conditions and tissue types. [Overview of the project]
[0003] Cell segmentation is a visual and image processing technique that divides an image or digital representation of a cell into individual regions or segments, each corresponding to a specific cell or cellular component. In spatial biology, cell segmentation enables several specific analyses that leverage the spatial configuration and characteristics of cells within a tissue sample, including spatial transcriptomics, spatial profiling, spatial clustering and spatial interaction analysis, spatial single-cell analysis, spatial co-expression analysis, spatial visualization, and data integration. However, cell boundary information is often noisy due to the variability of expression of morphological markers used. In some cases, expression levels may be missing or saturated, causing cells to be missed or mis-segmented. The lack of reliable cell membrane labeling is a common challenge for robust cell segmentation in current spatial transcriptomics analyses.
[0004] Existing cell segmentation methods often typically use a single-channel nuclear image or a small set of morphological stains between a single-channel nuclear image or a nuclear RGB image with up to three channels and membrane / cytoplasmic staining. This invention relates to a novel method for segmenting cells from tissue microscopy-derived images. Photocleavable markers can provide an infinite set of source-input channels through a process of repeatedly re-staining and removing markers in the tissue using UV irradiation and chemical washing. The additional number of morphological markers can dramatically increase the amount of information about cell boundaries presented in the tissue. This invention also provides a novel approach that utilizes spatial density generated from the readout density of a spatial transcriptomics assay as a morphological image of the cell body and performs cell segmentation on the generated readout density map.
[0005] In one embodiment, disclosed herein is a system comprising at least one processor and instructions executable by at least one processor, for providing a multimodal segmentation application including: (a) a software module configured to retrieve 3D scanned images of a biological sample in high dynamic range (HDR) mode, wherein the biological sample is labeled with one of a plurality of morphological markers; (b) a software module configured to reduce the 3D scanned images to 2D images, wherein an optimal focal region in the z slice is obtained; or utilizes the entire 3D stack volume; (c) a software module configured to retrieve a readout density map from a transcriptomics assay of a biological sample; and (d) a software module configured to perform intracellular segmentation on a biological sample based on a plurality of morphological markers or a readout density map. In one embodiment, the morphological marker includes a fluorescent dye, nuclear stain, fluorescently labeled antibody, immunohistochemistry (IHC) stain, photocleavable morphological marker, genetically encoded tag, magnetic resonance imaging (MRI) contrast agent, or nucleic acid probe. In one embodiment, the image is a microscope-derived image, including images from an optical microscope, electron microscope, or scanning probe microscope. In one embodiment, the image may include images from a spatial molecular imager. In one embodiment, the image may be subjected to bleaching correction. In one embodiment, the transcriptome assay may include a gene expression assay with a fluorescently labeled probe, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), cap analysis of gene expression, or single-cell RNA sequencing (scRNA-seq).
[0006] In some embodiments, the software module is configured to retrieve at least one, at least three, at least five, at least ten, at least fifteen, at least 20, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, at least 200, or more than 200 images of a biological sample. In some embodiments, the biological sample is obtained by at least a portion of one or more of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, neoplasm tissue sampling, malignant tissue sampling, diseased tissue sampling, and transplanted tissue sampling. In some embodiments, the biological sample includes cells or tissue. In one embodiment, cells include primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, cancer cells, or neoplasm cells. In one embodiment, intracellular segmentation includes nuclear segmentation, cytoplasmic segmentation, or extracellular segmentation. In one embodiment, intracellular segmentation includes training and / or applying machine learning algorithms.
[0007] In another embodiment, disclosed herein is a non-temporary computer-readable storage medium encoded with instructions executable by one or more processors to provide a multimodal segmentation application including: (a) a software module configured to retrieve a 3D scanned image of a biological specimen in high dynamic range (HDR) mode, wherein the biological specimen is labeled with one of a plurality of morphological markers. (b.1) A software module configured to optionally reduce 3D scan images to 2D images while preserving the optimal focal area from various z slices using techniques such as maximum intensity projection, focus stacking, and extended depth of field; (b.2) Alternatively, a software module configured to deconvolve 2D images to improve image sharpness and clarity while preserving the z position of those images in a 3D scan image stack; (b.3) Alternatively, a software module configured to directly use each deconvolved 2D slice forming a 3D volume stack for 3D cell segmentation; (c) A software module configured to retrieve readout density maps from transcriptome assays of biological samples; and (d) A software module configured to perform intracellular segmentation on a biological sample based on multiple morphological markers or readout density maps. In one embodiment, morphological markers include fluorescent dyes, nuclear stains, fluorescently labeled antibodies, immunohistochemistry (IHC) stains, photocleavable morphological markers, genetically encoded tags, magnetic resonance imaging (MRI) contrast agents, or nucleic acid probes. In one embodiment, the image is a microscope-derived image, including images from an optical microscope, electron microscope, or scanning probe microscope. In one embodiment, the image may include images from a spatial molecular imager. In one embodiment, the image is subjected to bleaching correction.In some embodiments, the transcriptome assay includes a gene expression assay using a fluorescently labeled probe, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), cap analysis of gene expression, or single-cell RNA sequencing (scRNA-seq). In some embodiments, the software module is configured to retrieve at least one, at least three, at least five, at least ten, at least fifteen, at least 20, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, or at least 200 or more images of the biological sample. In some embodiments, biological samples are obtained by at least a portion of one or more of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, tumor tissue sampling, malignant tissue sampling, diseased tissue sampling, and transplanted tissue sampling. In some embodiments, biological samples include cells or tissues. In some embodiments, cells include primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, cancer cells, or tumor cells. In some embodiments, intracellular segmentation includes nuclear segmentation, cytoplasmic segmentation, protrusion segmentation, or extracellular segmentation. In some embodiments, intracellular segmentation includes training a machine learning algorithm and / or applying a machine learning algorithm.
[0008] In another embodiment, disclosed herein is a computer-aided method comprising: (a) a computer retrieving 3D scanned images of a biological sample in high dynamic range (HDR) mode, where the biological sample is labeled with one of several morphological markers; (b.1) a software module configured to reduce the 3D scanned images to 2D images while maintaining the optimal focal area from various z-slice, using techniques such as maximum intensity projection, focus stacking, and extended depth of field, if necessary; (b.2) alternatively, a software module configured to deconvolve the 2D images while maintaining the z-positions of those images in the 3D scanned image stack to improve image sharpness and clarity; (c) a computer retrieving a readout density map from a transcriptome assay of the biological sample; and (d) performing image preprocessing and segmentation on the biological sample based on the multiple morphological markers or the readout density map. In some embodiments, the morphological marker includes a fluorescent dye, nuclear stain, fluorescently labeled antibody, immunohistochemistry (IHC) stain, photocleavable morphological marker, genetically encoded tag, magnetic resonance imaging (MRI) contrast agent, or nucleic acid probe.
[0009] In some embodiments, the images are microscope-derived images, including images from an optical microscope, electron microscope, or scanning probe microscope. In some embodiments, the images may include images from a spatial molecular imager. In some embodiments, the images are subjected to bleaching correction. In some embodiments, the transcriptome assay includes a gene expression assay with a fluorescently labeled probe, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), cap analysis of gene expression, or single-cell RNA sequencing (scRNA-seq). In some embodiments, the software module is configured to retrieve at least one image, at least three images, at least five images, at least ten images, at least fifteen images, at least 20 images, at least 30 images, at least 35 images, at least 40 images, at least 45 images, at least 50 images, at least 55 images, at least 60 images, at least 65 images, at least 70 images, at least 80 images, at least 90 images, at least 100 images, at least 120 images, at least 150 images, and at least 200 or more images of a biological sample. In some embodiments, biological samples are obtained by at least one of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, tumor tissue sampling, malignant tissue sampling, diseased tissue sampling, and transplanted tissue sampling. In some embodiments, biological samples include cells or tissues. In some embodiments, cells include primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, cancer cells, or tumor cells. In some embodiments, intracellular segmentation includes nuclear segmentation, cytoplasmic segmentation, or extracellular segmentation. In some embodiments, intracellular segmentation includes training a machine learning algorithm and / or applying a machine learning algorithm.
[0010] The systems, media, and methods disclosed herein include combining morphological staining with both proteomics and transcriptomics data to obtain rich information about cellular structures, including the nucleus, cytoplasm, and membrane regions. The systems, media, and methods disclosed herein include combining protein images expressing similar structures and / or cell types with multi-channel input morphological images. The systems, media, and methods disclosed herein include segmentation of extracellular processes. Various properties and statistics can be calculated on a cell-unit basis to enable researchers to study specific structures or regions of interest. These properties and statistics enable the establishment of downstream analyses using cell location and cell-unit fluorescence and shape properties generated during segmentation. It also serves as a baseline performance across multiple sample types and as a metric for evaluating ground-truth-free segmentation of new datasets. In some embodiments, the evaluation of cell segmentation performance is based on direct comparison. In some embodiments, the evaluation of cell segmentation performance is achieved by projecting the statistics of a new dataset onto an embedding space generated by the statistics of a reference dataset of similar sample types.
[0011] In some embodiments, image preprocessing is performed on the raw images before feeding them to at least one cell segmentation module. In some embodiments, image preprocessing is based on conventional image processing techniques such as denoising, deblurring, and image enhancement. In some embodiments, a variety of denoising algorithms can be performed, including filter-based (Wiener, Gaussian, median), total variation (TV) denoising, and deep learning-based approaches using wavelet transforms. In some embodiments, a variety of deblurring algorithms can be performed, including nearest neighbor deconvolution, Richardson-Lucy deconvolution, and total variation (TV) regularization. In some embodiments, a variety of image enhancement algorithms can be performed, including histogram equalization and contrast-limited adaptive histogram equalization (CLAHE), contrast stretching, gamma correction, and unsharp masking.
[0012] In some embodiments, cell segmentation is based on algorithms including Cellpose deep learning algorithms, adaptive thresholding, watershed algorithms, thresholding, region-based segmentation, active contours, convolutional neural networks (CNNs), level set methods, graph-based algorithms, fuzzy c-means clustering, deep watershed, image morphology, or Markov random fields. In some embodiments, nuclear segmentation, cytoplasmic segmentation, protrusion segmentation, and extracellular segmentation are based on the same algorithm. In some embodiments, nuclear segmentation, cytoplasmic segmentation, protrusion segmentation, and extracellular segmentation are based on different algorithms.
[0013] In various embodiments, image segmentation software based on machine learning (ML) algorithms can be applied to create cell boundaries from fluorescence images of protein assays. In some embodiments, the protein assay may include protein antibodies that bind to membrane proteins. In some embodiments, the ML algorithm applied to image segmentation may include semantic segmentation, instance segmentation, or a generative network for segmentation. In some embodiments, the image segmentation software may include ImageJ, CellProfiler, Cellpose, Ilastik, or QuPath, or any combination thereof.
[0014] In various embodiments, applications and use cases include, but are not limited to, the discovery and mapping of cell types and cellular states, phenotyping of the tissue microenvironment, differential expression of cell types based on spatial context, quantification of intracellular expression, and identification of spatially degraded biomarkers.
[0015] A patent or application file shall include at least one drawing made in color. A copy of this patent or patent application publication shall be provided by the Patent and Trademark Office upon request and payment of the required fees.
[0016] A better understanding of the features and advantages of this subject matter will be obtained by referring to the following detailed description which includes exemplary embodiments and accompanying drawings. [Brief explanation of the drawing]
[0017] [Figure 1] A non-specific example of a computing device is given. In this case, the device has one or more processors, memory, storage, and network interfaces. [Figure 2] This is a non-specific example of a web / mobile application delivery system. In this case, the system provides a browser-based and / or native mobile user interface. [Figure 3] Shows a non - limiting example of a cloud - based Web / mobile application provisioning system. In this case, the system includes elastically load - balanced and auto - scaling Web server and application server resources, and a synchronously replicated database. [Figure 4A] Shows a non - limiting example of a multimodal cell acquisition and segmentation system, particularly illustrating exemplary protein acquisition and 3D morphology sample collection. [Figure 4B] Shows exemplary RNA acquisition and detection. [Figure 4C] Shows an exemplary sub - cellular cell segmentation. [Figure 5] Shows a non - limiting example of a high - dynamic - range (HDR) imaging method. [Figure 6] Shows a non - limiting example of two - stage morphology scanning. [Figure 7] Shows a non - limiting example of an image affected by photobleaching from exposure to adjacent light. <* [Figure 8A] Shows non - limiting examples of scaling and normalization. [Figure 8B] Shows non - limiting examples of scaling and normalization. [Figure 8C] Shows non - limiting examples of scaling and normalization. [Figure 9] Shows a non - limiting example of chromatic aberration. [Figure 10] Shows a non - limiting example of chromatic aberration correction. [Figure 11] Shows a non - limiting example of spot density generation. [Figure 12] Shows a non - limiting example of generating density using a grid. [Figure 13A] Shows non - limiting examples of segmentation results based on morphological staining and spot read - out density, particularly showing non - limiting examples of several protein channels. [Figure 13B] Shows non - limiting examples of segmented channels. [Figure 14]A non-restrictive example of a diagram for multimodal cell segmentation is shown. [Figure 15] This demonstrates a non-limiting example of processing cell segmentation using protein and spot density inputs. [Figure 16] This shows a non-restrictive example of generating a new N-channel model. [Figure 17A] This provides non-restrictive examples of comparing 2-channel and 5-channel model outputs, and in particular, non-restrictive examples of various 2-channel model outputs. [Figure 17B] This shows non-restrictive examples of various 5-channel model outputs. [Figure 18A] This paper presents a non-limiting example of high-density neuronal cells in the human brain, particularly neuronal cells with nuclear markers. [Figure 18B] This shows neuronal cells with membrane markers. [Figure 18C] This shows astrocytes containing GFAP markers. [Figure 18D] This shows nuclei containing the DAPI marker. [Figure 19] This shows a non-restrictive example of extracellular process segmentation. [Figure 20A] This paper presents non-limiting examples of results from cell segmentation, particularly non-limiting examples of segmentation result overlays of cellular and extracellular processes in the human brain. [Figure 20B] A non-limiting example of a corresponding cell segmentation mask is shown. [Figure 20C] This shows a non-limiting example of detecting extracellular cell segmentation masks. [Figure 21] This shows a non-limiting example of cell segmentation output that provides intracellular and extracellular masks and statistics. [Figure 22] This shows a non-restrictive example of the objective lens cone angle for the blurring function. [Figure 23] This shows a non-restrictive example of a three-dimensional segmentation pipeline. [Figure 24]This section demonstrates image preprocessing techniques such as image sharpening and enhancement. [Figure 25] This shows a non-exclusive example of an image acquired using the High Definition Range (HDR) setting. [Figure 26] This shows a non-restrictive example of output generated by the nearest neighbor (NN) deconvolution method. [Figure 27] This shows a non-restrictive example of cell protrusion merging. [Figure 28] This shows a non-restrictive example of the result of cross-linking (IoU) merging of the binding between nuclear segmentation output and membrane-plus-nuclear segmentation. [Figure 29A] This paper presents a non-restrictive example of cross-analysis, particularly a non-restrictive example of IoU merging results in 2D image slices. [Figure 29B] This shows a non-limiting example of IoU merging results within a 3D volume. [Figure 29C] This shows a non-restrictive example of a single cell marked across all z-slices throughout the entire Z-stack. [Figure 30] This shows non-restrictive examples of cell segmentation in various tissues. [Modes for carrying out the invention]
[0018] In certain embodiments, a system comprising at least one processor and instructions executable by at least one processor is described herein to provide a multimodal segmentation application comprising: (a) a software module configured to retrieve a 3D scanned image of a biological sample in high dynamic range (HDR) mode, wherein the biological sample is labeled with one of a plurality of morphological markers; (b) a software module configured to reduce the 3D scanned image to a 2D image to obtain an optimal focal region within a z slice, or to utilize the entire 3D stack in which each 2Dz slice is deconvoluted to enhance features and reduce blur; (c) a software module configured to retrieve a readout density map from a transcriptome assay of a biological sample; and (d) a software module configured to perform intracellular segmentation on a biological sample based on a plurality of morphological markers or a readout density map. In some embodiments, the morphological markers may include one or more fluorescent dyes, nuclear stains, fluorescently labeled antibodies, immunohistochemistry (IHC) stains, photocleavable morphological markers, genetically encoded tags, magnetic resonance imaging (MRI) contrast agents, or nucleic acid probes, or any combination thereof.
[0019] In some embodiments, the images may be microscopic images. In some embodiments, the microscopic images may be from an optical microscope, electron microscope, or scanning probe microscope. In some embodiments, the images may be from a spatial molecular imager. In some embodiments, bleaching correction may be applied to the images. In some embodiments, the transcriptome assay may include one or more of the following: a gene expression assay using a fluorescently labeled probe, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), cap analysis of gene expression, or single-cell RNA sequencing (scRNA-seq), or any combination thereof. In some embodiments, the software module can be configured to retrieve at least one, at least three, at least five, at least ten, at least fifteen, at least 20, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, at least 200, or more than 200 images of a biological sample. In some embodiments, the biological sample can be obtained from at least a portion of one or more of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue collection, tumor tissue collection, malignant tissue collection, diseased tissue collection, or transplant tissue collection, or any combination thereof. In some embodiments, the biological sample can include cells or tissue. In one embodiment, the cells may include one or more of the following: primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, cancer cells, or tumor cells, or any combination thereof.In some embodiments, intracellular segmentation may include one or more nuclear segmentation, cytoplasmic segmentation, or extracellular segmentation, or any combination thereof. In some embodiments, intracellular segmentation may include training a machine learning algorithm, or applying a machine learning algorithm, or both.
[0020] Furthermore, in certain embodiments, non-temporary computer-readable storage media encoded with instructions executable by one or more processors are described herein, which can provide multimodal segmentation applications including: (a) a software module configured to retrieve 3D scanned images of a biological sample in high dynamic range (HDR) mode, wherein the biological sample is labeled with one of a plurality of morphological markers; (b.1) a software module configured to reduce the 3D scanned images to 2D images while maintaining an optimal focal area from various z-slice, using techniques such as maximum intensity projection, focus stacking, and extended depth of field, as needed; (b.2) alternatively, a software module configured to deconvolve the 2D images to improve image sharpness and clarity while maintaining the z-position of those images in a stack of 3D scanned images; (c) a software module configured to retrieve readout density maps from a transcriptome assay of a biological sample; and (d) a software module configured to perform intracellular segmentation on a biological sample based on a plurality of morphological markers or readout density maps. In some embodiments, the morphological marker may include one or more of the following: a fluorescent dye, nuclear stain, fluorescently labeled antibody, immunohistochemistry (IHC) stain, photocleavable morphological marker, genetically encoded tag, magnetic resonance imaging (MRI) contrast agent, or nucleic acid probe, or any combination thereof. In some embodiments, the image may be a microscopic image. In some embodiments, the microscopic image may include an image from an optical microscope, electron microscope, or scanning probe microscope, or any combination thereof. In some embodiments, the image may include an image from a spatial molecular imager. In some embodiments, bleaching correction may be applied to the image.In some embodiments, a transcriptome assay may include one or more of the following: a gene expression assay using a fluorescently labeled probe, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), cap analysis of gene expression, or single-cell RNA sequencing (scRNA-seq), or any combination thereof. In some embodiments, a software module may be configured to search for at least one image, at least three images, at least five images, at least ten images, at least fifteen images, at least 20 images, at least 30 images, at least 35 images, at least 40 images, at least 45 images, at least 50 images, at least 55 images, at least 60 images, at least 65 images, at least 70 images, at least 80 images, at least 90 images, at least 100 images, at least 120 images, at least 150 images, at least 200 images, or at least more than 200 images of a biological sample. In some embodiments, a biological sample can be obtained at least partially from one or more of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, tumor tissue sampling, malignant tissue sampling, diseased tissue sampling, and transplanted tissue sampling, or any combination thereof. In some embodiments, a biological sample can include cells or tissues. In some embodiments, cells can include primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, cancer cells, or tumor cells, or any combination thereof. In some embodiments, intracellular segmentation can include nuclear segmentation, cytoplasmic segmentation, or extracellular segmentation, or any combination thereof. In some embodiments, intracellular segmentation can include training a machine learning algorithm, applying a machine learning algorithm, or both.
[0021] In another embodiment also described herein, there are computer-aided methods that include: (a) a computer retrieving 3D scanned images of a biological sample in high dynamic range (HDR) mode, where the biological sample may be labeled with one of several morphological markers; (b.1) a software module optionally configured to reduce the 3D scanned images to 2D images while maintaining the optimal focal area from various z-slice using techniques such as maximum intensity projection, focus stacking, and extended depth of field; (b.2) alternatively, a software module which may be configured to deconvolve the 2D images while maintaining the z-positions of those images in a stack of 3D scanned images to improve image sharpness and clarity; (c) a computer retrieving a readout density map from a transcriptome assay of the biological sample; and (d) performing segmentation on the biological sample based on several morphological markers or the readout density map. In some embodiments, the morphological marker may include one or more of the following: a fluorescent dye, nuclear stain, fluorescently labeled antibody, immunohistochemistry (IHC) stain, photocleavable morphological marker, genetically encoded tag, magnetic resonance imaging (MRI) contrast agent, or nucleic acid probe, or any combination thereof. In some embodiments, the image may be a microscopic image. In some embodiments, the microscopic image may include an image from a light microscope, electron microscope, or scanning probe microscope, or any combination thereof. In some embodiments, the image may include an image from a spatial molecular imager. In some embodiments, bleaching correction may be applied to the image. In some embodiments, the transcriptome assay may include a gene expression assay comprising one or more fluorescently labeled probes, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), gene expression cap analysis, or single-cell RNA sequencing (scRNA-seq), or any combination thereof.In some embodiments, the software module can be configured to retrieve at least one, at least three, at least five, at least ten, at least fifteen, at least 20, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, at least 200, or more than 200 images of a biological sample. In some embodiments, the biological sample can be obtained at least partially from one or more of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue collection, tumor tissue collection, malignant tissue collection, diseased tissue collection, and transplant tissue collection, or any combination thereof. In some embodiments, the biological sample may include cells or tissues. In some embodiments, the cells may include one or more primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, cancer cells, or tumor cells, or any combination thereof. In some embodiments, intracellular segmentation may include one or more nuclear segmentation, cytoplasmic segmentation, or extracellular segmentation, or any combination thereof. In some embodiments, intracellular segmentation may include training a machine learning algorithm, applying a machine learning algorithm, or both.
[0022] definition Unless otherwise defined, all technical terms used herein have the same meaning as those generally understood by those skilled in the art to which this subject matter belongs.
[0023] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural references unless the context explicitly indicates otherwise. References to "or" herein are intended to include "and / or" unless otherwise stated.
[0024] Throughout this specification, any reference to “several embodiments,” “further embodiments,” or “specific embodiments” means that the particular features, structures, or characteristics described in relation to an embodiment are included in at least one embodiment. Therefore, the occurrences of the phrases “in several embodiments,” “further embodiments,” or “specific embodiments” in various places throughout this specification do not necessarily all refer to the same embodiment. Furthermore, specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0025] Multimodal cell acquisition and segmentation system In some embodiments, the systems, media, and methods disclosed herein may include combining morphological staining with both proteomics and transcriptomics data. In some embodiments, the systems, media, and methods disclosed herein may further include receiving information about cellular structures. In some embodiments, cellular structures may include the nucleus, cytoplasm, or membrane regions, or any combination thereof. In some embodiments, the systems, media, and methods disclosed herein may include combining protein images expressing similar structures with multichannel input morphological images. In some embodiments, the systems, media, and methods disclosed herein may include combining protein images expressing similar cell types with multichannel input morphological images. In some embodiments, the systems, media, and methods disclosed herein may include combining protein images expressing similar cell types and similar structures with multichannel input morphological images.In some embodiments, the multichannel input image may include at least 1-channel input morphology image, at least 3-channel input morphology image, at least 5-channel input morphology image, at least 10-channel input morphology image, at least 15-channel input morphology image, at least 20-channel input morphology image, at least 30-channel input morphology image, at least 35-channel input morphology image, at least 40-channel input morphology image, at least 45-channel input morphology image, at least 50-channel input morphology image, at least 55-channel input morphology image, at least 60-channel input morphology image, at least 65-channel input morphology image, at least 70-channel input morphology image, at least 80-channel input morphology image, at least 90-channel input morphology image, at least 100-channel input morphology image, at least 120-channel input morphology image, at least 150-channel input morphology image, or at least 200-channel input morphology image, or at least more than 200-channel input morphology images (including increments).
[0026] In some embodiments, the systems, media, and methods disclosed herein may include a multi-step process of segmentation. In some embodiments, the multi-step process of segmentation may include segmentation of nuclear regions. In some embodiments, the multi-step process of segmentation may include segmentation of cytoplasmic regions. In some embodiments, the multi-step process of segmentation may include segmentation regions of membranes. In some embodiments, the systems, media, and methods disclosed herein may include segmentation of extracellular objects for segmenting regions. In some embodiments, the segmented region may be a segmented region of an organ or tissue. In some embodiments, the segmented region may be a segmented region of the brain. In some embodiments, the segmented region may be distal to or cleaved relative to somatic cells, or both. In some embodiments, the systems, media, and methods disclosed herein may include combining images expressing similar structures with one or more input morphological images. In some embodiments, the systems, media, and methods disclosed herein may include combining images expressing similar cell types with one or more input morphological images. In some embodiments, the systems, media, and methods disclosed herein may include combining images expressing similar structures with one or more input morphological images, and combining cell types with one or more input morphological images. In some embodiments, the input morphological images may be multi-channel input morphological images. In some embodiments, the systems, media, and methods disclosed herein may include combining images expressing similar structures with a five-channel input morphological image. In some embodiments, the systems, media, and methods disclosed herein may include combining images expressing similar cell types with a five-channel input morphological image. In some embodiments, the systems, media, and methods disclosed herein may include combining images expressing similar structures with a five-channel input morphological image, and combining cell types with a five-channel input morphological image.
[0027] In some embodiments, the systems, media, and methods disclosed herein may include combining one or more protein images. In some embodiments, the number of protein images combined may be about 1, about 2, about 3, about 4, about 5, about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 120, about 140, about 160, about 180, about 200, or about 200 or more protein images. In some embodiments, the systems, media, and methods disclosed herein may include processing additional input channels of raw RNA spot readout density to compensate for signal loss regions from morphological markers.
[0028] In some embodiments, the systems, media, and methods disclosed herein may include segmentation of extracellular processes. In some embodiments, the systems, media, and methods disclosed herein may further include labeling one or more images having a unique ID for each cell. In some embodiments, the systems, media, and methods disclosed herein may further include labeling one or more images having a unique ID for each connected process. In some embodiments, the systems, media, and methods disclosed herein may further include labeling one or more images having a unique ID for each cell and each connected process. The systems, media, and methods disclosed herein may further include identifying “compartment labels” for different compartments of a region for each cell.
[0029] As a non-limiting example, a multimodal cell acquisition and segmentation system is shown in Figures 4A-4C. For example, protein acquisition and 3D morphology samples were collected as shown in Figure 4A. Protein samples and morphological markers, such as photosection markers, were incubated and scanned into 3D images in high dynamic range (HDR) mode. The markers were then removed, for example, by UV irradiation and chemical washing within the tissue. Repeated restaining and removal of the markers were performed, for example, by UV irradiation and chemical washing within the tissue. As an exemplary example, RNA acquisition and detection were performed as shown in Figure 4B. Next, for example, RNA reporter hybridization samples were scanned and RNA spots were detected. Then, for example, the markers were removed, for example, by UV irradiation and washing within the sample.
[0030] In some embodiments, the systems, media, and methods disclosed herein may include retrieving images from one or more tissue scans. In some embodiments, the one or more tissue scans may include one or more scans of one or more tissues stained with morphological markers. In some embodiments, the one or more tissue scans may include one or more scans of one or more tissues incubated with a protein reporter. In some embodiments, the one or more tissue scans may include one or more scans of one or more tissues incubated with an RNA reporter. In some embodiments, the systems, media, and methods disclosed herein may include removing one or more reporters from one or more tissue samples. In some embodiments, one or more reporters may be reduced in one or more tissue samples. In some embodiments, one or more reporters may be bleached, cut, or both in one or more tissue samples. In one embodiment, the reporter may be reduced in one or more tissue samples by applying one or more of the following: ultraviolet light, heat, proteases, endonucleases, nucleases, esterases, ribonucleases, RNase A, RNase T1, RNase H, disulfate-bonded reducing agents (e.g., dithiothreitol, tris(2-carboxyethyl)phosphine), salt buffers, alkalis, hydrogen bond destabilizing solvents (e.g., formamide, DMSO), or any combination thereof.
[0031] In some non-limiting examples, after processing images of one or more proteins using morphological and density measurements for RNA detection, subcellular cell segmentation was performed, for example, as shown in Figure 4C. As illustrated in Figure 4C, the cell segmentation process includes, for example, background subtraction, normalization, region-to-nucleus segmentation, region-to-cytoplasmic or region-to-membrane segmentation, or any combination thereof. In one embodiment, different compartments of a region in each cell may be identified by unique "compartment labels." In one embodiment, various properties and statistics may be calculated for each cell. In one embodiment, the various properties and statistics calculated for each cell may provide information about a particular structure or region of interest, or both.
[0032] In some embodiments, the systems, media, and methods disclosed herein may include retrieving images of biological specimens. In some embodiments, retrieved images of biological specimens may be labeled with morphological markers. In some embodiments, the morphological markers may include one or more of the following: fluorescent dyes, nuclear stains, fluorescently labeled antibodies, immunohistochemistry (IHC) stains, photocleavable morphological markers, genetically encoded tags, magnetic resonance imaging (MRI) contrast agents, or nucleic acid probes, or any combination thereof.
[0033] In some embodiments, the systems, media, and methods disclosed herein may include generating a readout density map using a transcriptome assay of a biological sample. In some embodiments, the density map may be a heatmap, choropleth map, kernel density map, dot density map, contour map, or proportional symbol map. In some embodiments, the heatmap may further include a graphic representation of the data where values are indicated using color. In some embodiments, the heatmap can be used to visualize the density or intensity of a particular phenomenon across a two-dimensional space. In some embodiments, the two-dimensional space may be a geographical region, an image, or a grid, or any combination thereof. In some embodiments, the density map may show the density or concentration of a particular attribute or event within a given region. In some embodiments, the density map may provide visual information regarding the distribution and intensity of data points or events across a spatial region.
[0034] In one embodiment, a transcriptome assay may include one or more of the following: a gene expression assay using a fluorescently labeled probe, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), cap analysis of gene expression, or single-cell RNA sequencing (scRNA-seq), or any combination thereof.
[0035] In some embodiments, the systems, media, and methods disclosed herein may include retrieving images of a biological sample. In some embodiments, a biological sample may be obtained at least partially using one or more of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, tumor tissue sampling, malignant tissue sampling, diseased tissue sampling, or transplanted tissue sampling, or any combination thereof.
[0036] In some embodiments, the biological sample may include cells, tissue, or both. In some embodiments, the cells may include one or more primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, neoplasm cells, or any combination thereof.
[0037] In some embodiments, the systems, media, and methods disclosed herein may include retrieving images of spatially degraded highplex gene expression data from tissue. In some embodiments, the systems, media, and methods disclosed herein may include retrieving images of spatially degraded highplex protein data from tissue. In some embodiments, the systems, media, and methods disclosed herein may include retrieving images of RNA assays. In some embodiments, the systems, media, and methods disclosed herein may include retrieving images of RNA assays for profiling an entire transcriptome. In some embodiments, the systems, media, and methods disclosed herein may include retrieving images of RNA assays for profiling an entire transcriptome from tissue on a single formalin-fixed paraffin-embedded (FFPE) sample. In some embodiments, the systems, media, and methods disclosed herein may include retrieving images of RNA assays for profiling an entire transcriptome from tissue on a fresh-frozen (FF) sample slide. In some embodiments, the systems, media, and methods disclosed herein may include retrieving images of protein assays for generating quantitative analysis, spatial analysis, or both of multiple proteins. In one embodiment, quantitative analysis, spatial analysis, or both of multiple proteins can be generated from a single FFPE or FF sample slide. In one embodiment, the FFPE or FF tissue section may be stained with a barcoded in-situ hybridization probe that binds to endogenous mRNA transcripts. In one embodiment, the user may select a region of interest (ROI) to profile. In one embodiment, each ROI segment may be further subdivided into illuminated regions (AOIs) based on tissue morphology. In one embodiment, the spatial molecular imaging device may photosection each AOI segment separately. In one embodiment, the spatial molecular imaging device may separately collect expression tags or barcodes for each AOI segment.In one embodiment, tags or barcodes may be used for downstream sequencing or data processing, or both.
[0038] Computing systems Referring to Figure 1, a block diagram is shown illustrating an exemplary machine including a computer system 100 (e.g., a processing or computing system) in which a set of instructions can be executed to cause a device to execute or perform any one or more embodiments and / or methodologies for static code scheduling of the present disclosure. The components herein are merely examples and do not limit the scope of use or functionality of any hardware, software, embedded logic components, or combinations of two or more such components that implement a particular embodiment.
[0039] The computer system 100 may include one or more processors 101, memory 103, and storage 108 that communicate with each other and with other components via a bus 140. The bus 140 may also link a display 132, one or more input devices 133 (e.g., keypads, keyboards, mice, styluses, etc.), one or more output devices 134, one or more storage devices 135, and various tangible storage media 136. All of these elements may interface with the bus 140 directly or via one or more interfaces or adapters. For example, various tangible storage media 136 may interface with the bus 140 via a storage media interface 126. The computer system 100 may have any suitable physical form, including, but is not limited to, one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile phones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0040] The computer system 100 includes one or more processors 101 that perform functions (e.g., a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), or a quantum processing unit (QPU)). The processors 101 optionally include cache memory units 102 for temporary local storage of instructions, data, or computer addresses. The processors 101 are configured to assist in the execution of computer-readable instructions. The computer system 100 can provide the functionality of the components shown in Figure 1 as a result of the processors 101 executing non-temporary processor-executable instructions embodied in one or more tangible computer-readable storage media such as memory 103, storage device 108, storage device 135, and / or storage medium 136. The computer-readable media can store software that implements a particular embodiment, and the processors 101 can execute the software. Memory 103 can read software from one or more other computer-readable media (e.g., mass storage devices 135, 136) or from one or more other sources via a suitable interface such as a network interface 120. The software may cause the processor 101 to execute one or more processes or one or more steps of one or more processes described or illustrated herein. Executing such processes or steps may include defining a data structure to be stored in memory 103 and modifying the data structure as directed by the software.
[0041] Memory 103 may include, but is not limited to, various components (e.g., machine-readable media), including, random access memory components (e.g., RAM 104) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM®), phase-change random access memory (PRAM), etc.), read-only memory components (e.g., ROM 105), and any combination thereof. ROM 105 can act to communicate data and instructions unidirectionally to the processor 101, and RAM 104 can act to communicate data and instructions bidirectionally to the processor 101. ROM 105 and RAM 104 may include any suitable tangible computer-readable media as described below. For example, a basic input / output system 106 (BIOS) containing basic routines that help transfer information between elements within the computer system 100, such as during startup, may be stored in memory 103.
[0042] The fixed storage device 108 is optionally connected bidirectionally to the processor 101 via a storage control unit 107. The fixed storage device 108 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. The storage device 108 can be used to store the operating system 109, executable files 110, data 111, applications 112, etc. The storage device 108 may also include an optical disc drive, a solid-state memory device (e.g., a flash-based system), or any combination of the above. Information in the storage device 108 can be incorporated as virtual memory in memory 103, where appropriate.
[0043] In one example, the storage device 135 can be detachably interfaced with a computer system 100 (for example, via an external port connector (not shown)) via a storage device interface 125. In particular, the storage device 135 and associated machine-readable media can provide non-volatile and / or volatile storage for machine-readable instructions, data structures, program modules, and / or other data for the computer system 100. In one example, the software may reside entirely or partially within the machine-readable media on the storage device 135. In another example, the software may reside entirely or partially within the processor 101.
[0044] Bus 140 connects a wide variety of subsystems. Here, a reference to a bus may include, where necessary, one or more digital signal lines that provide common functionality. Bus 140 may include, but is not limited to, several types of bus structures, including, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof using any of the various bus architectures. Examples of such architectures, but are not limited to, include the Industrial Standards Architecture (ISA) bus, the Extended ISA (EISA) bus, the Microchannel Architecture (MCA) bus, the Video Electronics Standards Institute Local Bus (VLB), the Peripheral Interconnect (PCI) bus, the PCI-Express (PCI-X) bus, the High-Speed Graphics Port (AGP) bus, the Hypertransport (HTX) bus, the Serial Advanced Technology Attachment (SATA) bus, and any combination thereof.
[0045] The computer system 100 may also include an input device 133. In one example, a user of the computer system 100 may input commands and / or other information to the computer system 100 via the input device 133. Examples of input devices 133 include, but are not limited to, alphanumeric input devices (e.g., keyboards), pointing devices (such as mice and touchpads), touchpads, touchscreens, multitouchscreens, joysticks, styluses, gamepads, audio input devices (e.g., microphones, voice response systems), optical scanners, video or still image capture devices (e.g., cameras), and any combination thereof. In some embodiments, the input device may be Kinect, Leap Motion, etc. The input device 133 may interface to the bus 140 via one of various input interfaces 123 (e.g., input interface 123) including, but not limited to, serial, parallel, game port, USB, FIREWIRE®, THUNDERBOLT®, or any combination thereof.
[0046] In certain embodiments, if computer system 100 is connected to network 130, computer system 100 can communicate with other devices connected to network 130, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, and cloud computing systems. Communication with computer system 100 can be transmitted via network interface 120. For example, network interface 120 can receive incoming communications (such as requests or responses from other devices) from network 130 in the form of one or more packets (such as Internet Protocol (IP) packets), and computer system 100 can store the incoming communications in memory 103 for processing. Similarly, computer system 100 can store outgoing communications (such as requests or responses to other devices) in memory 103 in the form of one or more packets and communicate from network interface 120 to network 130. Processor 101 can access these communication packets stored in memory 103 for processing.
[0047] Examples of network interface 120 include, but are not limited to, network interface cards, modems, and any combination thereof. Examples of network 130 or network segment 130 include, but are not limited to, distributed computing systems, cloud computing systems, wide area networks (WANs) (e.g., the internet, corporate networks), local area networks (LANs) (e.g., networks related to offices, buildings, campuses, or other relatively small geographical spaces), telephone networks, direct connections between two computing devices, peer-to-peer networks, and any combination thereof. Networks like network 130 can use wired and / or wireless communication modes. In general, any network topology can be used.
[0048] Information and data can be displayed via the display 132. Examples of the display 132 include, but are not limited to, cathode ray tubes (CRTs), liquid crystal displays (LCDs), thin-film transistor liquid crystal displays (TFT-LCDs), organic liquid crystal displays (OLEDs) such as passive-matrix OLEDs (PMLEDs) or active-matrix OLEDs (AMOLEDs), plasma displays, and any combination thereof. The display 132 can interface via the bus 140 to the processor 101, memory 103, and fixed storage device 108, as well as other devices such as input device 133. The display 132 is linked to the bus 140 via a video interface 122, and the transfer of data between the display 132 and the bus 140 can be controlled via graphics control 121. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD), such as a VR headset. In further embodiments, suitable VR headsets include, but are not limited to, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOYE VR, Zeiss VR One, Avegant Glyph, and Freefly VR headsets. In yet another embodiment, the display is a combination of devices such as those disclosed herein.
[0049] In addition to the display 132, the computer system 100 may include one or more other peripheral output devices 134, including but not limited to audio speakers, printers, storage devices, and any combination thereof. Such peripheral output devices may be connected to the bus 140 via an output interface 124. Examples of the output interface 124 include, but are not limited to, serial ports, parallel connections, USB ports, FIREWIRE® ports, THUNDERBOLT® ports, and any combination thereof.
[0050] Furthermore, or alternatively, the computer system 100 may provide functionality as a result of logic hardwired into or incorporated into the circuit. This circuit may operate in place of or with software to perform one or more processes or one or more steps of one or more processes described or illustrated herein. References to software in this disclosure may include logic, and references to logic may include software. Furthermore, references to computer-readable media may, as appropriate, include circuitry (such as an IC) that houses software for execution, circuitry that embodies logic for execution, or both. This disclosure includes any suitable combination of hardware, software, or both.
[0051] Those skilled in the art will understand that various exemplary logic blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly demonstrate this hardware and software compatibility, various exemplary components, blocks, modules, circuits, and steps are generally described above in relation to their functions.
[0052] Various exemplary logic blocks, modules, and circuits described in connection with embodiments disclosed herein may be implemented or run on general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or other such configurations.
[0053] Steps of methods or algorithms described in relation to embodiments disclosed herein can be implemented directly in hardware, in software modules executed by one or more processors, or in a combination of the two. The software modules may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and storage medium may reside as separate components within a user terminal.
[0054] As described herein, suitable computing devices include, in non-limiting examples, distributed computing systems, cloud computing platforms, server clusters, server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netpad computers, handheld computers, internet devices, mobile smartphones, and tablet computers. Those skilled in the art will also recognize that selected televisions, video players, and digital music players with any computer network connectivity are also suitable for use in the systems described herein. Suitable tablet computers include, in various embodiments, those having booklets, slates, and convertible configurations known to those skilled in the art.
[0055] In some embodiments, a computing device includes an operating system configured to execute executable instructions. The operating system is software, for example, including programs and data, that manages the device's hardware and provides services for running applications. Those skilled in the art will recognize that suitable server operating systems include, in non-limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux®, Apple®, Mac OS X Server®, Oracle®, Solaris®, Windows Server®, Novell®, and NetWare®. Those skilled in the art will also recognize that suitable personal computer operating systems include, in non-limiting examples, UNIX®-like operating systems such as Microsoft®, Windows®, Apple®, Mac OS X®, UNIX®, and GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those skilled in the art will also recognize that suitable mobile smartphone operating systems include, but are not limited to, Nokia®, Symbian®, OS, Apple®, iOS®, Research In Motion®, BlackBerry OS®, Google®, Android®, Microsoft®, Windows Phone®, OS, Microsoft®, Windows Mobile®, OS, Linux®, and Palm®, WebOS®.
[0056] Non-temporary computer-readable storage medium In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-temporary computer-readable storage media encoded with a program containing instructions that can optionally be executed by the operating system of a network computing device. In further embodiments, the computer-readable storage medium is a tangible component of the computing device. In yet another embodiment, the computer-readable storage medium is optionally removable from the computing device. In some embodiments, the computer-readable storage medium includes, in non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid-state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are encoded permanently, substantially permanently, semi-permanently, or non-temporarily on the medium.
[0057] Computer program In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program or use thereof. A computer program includes a sequence of instructions executable by one or more processors of the CPU of a computing device, written to perform a particular task. Computer-readable instructions may be implemented as program modules such as functions, objects, application programming interfaces (APIs), and computing data structures, which perform a particular task or implement a particular abstract data type. In light of the disclosures provided herein, those skilled in the art will recognize that computer programs may be written in various versions of various languages.
[0058] The functionality of computer-readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program includes one sequence of instructions. In some embodiments, a computer program includes multiple sequence of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from multiple locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plugins, extensions, add-ins, or add-ons, or a combination thereof.
[0059] Web application In some embodiments, a computer program may include a web application. Those skilled in the art will recognize, in light of the disclosures provided herein, that a web application may, in various embodiments, utilize one or more software frameworks and one or more database systems. In some embodiments, a web application is built on a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, a web application utilizes one or more database systems, including relational, non-relational, object-oriented, associative, XML, and document-oriented database systems. In further embodiments, suitable relational database systems include, in non-limiting examples, Microsoft® SQL Server, MySQL®, and Oracle®. Those skilled in the art will also recognize that, in various embodiments, a web application may be written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation-definition languages, client-side scripting languages, server-side coding languages, database query languages, or a combination thereof. In some embodiments, the web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or Extensible Markup Language (XML). In some embodiments, the web application can be written to some extent in a presentation-defining language such as Cascading Style Sheets (CSS). In some embodiments, the web application can be written to some extent in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Flash® ActionScript, JavaScript, or Silverlight®.In some embodiments, the web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel, Smalltalk, WebDNA®, or Groovy. In some embodiments, the web application can be written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, the web application integrates with an enterprise server product such as IBM® Lotus Domino®. In some embodiments, the web application includes a media player element. In various further embodiments, the media player element utilizes one or more of many suitable multimedia technologies, including, but not limited to, Adobe® Flash®, HTML5, Apple®, QuickTime®, Microsoft®, Silverlight®, Java™, and Unity®.
[0060] Referring to Figure 2, in a particular embodiment, the application delivery system includes one or more databases 200 accessed by a relational database management system (RDBMS) 210. Suitable RDBMSs include Firebird, MySQL®, PostgreSQL, SQLite, Oracle Database, Microsoft SQL Server, IBM DB2, IBM Informix, SAP Sybase, Teradata, etc. In this embodiment, the application delivery system further includes one or more application servers 220 (e.g., Java Server, .NET Server, PHP Server, etc.) and one or more web servers 230 (e.g., Apache, IIS, GWS, etc.). The web servers optionally expose one or more web services via an application programming interface (API) 240. Over a network such as the Internet, the system provides a browser-based and / or mobile-native user interface.
[0061] Referring to Figure 3, in a particular embodiment, the application delivery system has a distributed cloud-based architecture 300 and includes elastically load-balanced, auto-scaling web server resources 310 and application server resources 320, as well as a synchronously replicated database 330.
[0062] Standalone application In some embodiments, a computer program may include a standalone application, which is a program that runs as an independent computer process, not as a plug-in or as an add-on to an existing process. Those skilled in the art will recognize that standalone applications are often compiled. A compiler is a computer program that translates source code written in a programming language into binary object code, such as assembly language or machine code. Suitable compiled programming languages, in non-limiting examples, include C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB.NET, or combinations thereof. Compilation is often performed at least partially to create an executable program. In some embodiments, a computer program may include one or more executable compiled applications.
[0063] Software Module In some embodiments, the platforms, systems, media, and methods disclosed herein may include software, servers, and / or database modules, or uses thereof. With regard to the disclosures provided herein, software modules are created by techniques known to those skilled in the art, using machines, software, and languages known to those skilled in the art. Software modules disclosed herein are implemented in numerous ways. In various embodiments, a software module may include files, sections of code, programming objects, programming structures, distributed computing resources, cloud computing resources, or a combination thereof. Further various embodiments may include multiple files, multiple sections of code, multiple programming objects, multiple programming structures, multiple distributed computing resources, multiple cloud computing resources, or a combination thereof. In various embodiments, one or more software modules may include, in non-limiting examples, web applications, mobile applications, standalone applications, and distributed or cloud computing applications. In some embodiments, a software module resides within one computer program or application. In other embodiments, a software module resides within two or more computer programs or applications. In some embodiments, a software module is hosted on one machine. In other embodiments, a software module is hosted on one or more machines. In further embodiments, the software module is hosted on a distributed computing platform such as a cloud computing platform. In one embodiment, the software module is hosted on one or more machines in one location. In another embodiment, the software module is hosted on one or more machines in one or more locations.
[0064] database In some embodiments, the systems, media, and methods disclosed herein may include one or more databases or the use thereof. In consideration of the disclosures provided herein, those skilled in the art will recognize that many databases are suitable for storing and retrieving user information, examination information, slide information, field of view (FoV) information, flow cell information, image information, genomic information, transcriptome information, and proteomics information. In various embodiments, suitable databases include, in non-limiting examples, relational databases, non-relational databases, object-oriented databases, object databases, entity-relational model databases, associative databases, XML databases, document-oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL®, Oracle, DB2, Sybase, and MongoDB. In some embodiments, the database may be internet-based. In further embodiments, the database may be web-based. In yet further embodiments, the database may be cloud computing-based. In certain embodiments, the database may be a distributed database. In other embodiments, the database may be based on one or more local computer storage devices.
[0065] In some embodiments, the data stored in the database may include biological image data. In some embodiments, the biological image data may include, in some non-limiting examples, microscopic images (e.g., micrographs) of formalin-fixed paraffin-embedded (FFPE) cell samples, microscopic images (e.g., micrographs) of formalin-fixed paraffin-embedded (FFPE) tissue samples, microscopic images (e.g., micrographs) of fresh-frozen (FF) cell samples, or microscopic images (e.g., micrographs) of fresh-frozen (FF) tissue samples, or any combination thereof. In some embodiments, data from a single slide may be split into two datasets. In some cases, the two datasets may include an RNA assay dataset and a protein assay dataset. In some embodiments, data from a single slide may be combined into a single dataset. In some embodiments, the single dataset may include both RNA assay data and protein assay data. In some embodiments, the image data may be two-dimensional image data. In some embodiments, the image data may be three-dimensional image data. In some embodiments, the data may include, in non-limiting examples, “omics” data such as genomic data, proteomics data, metabolomics data, metagenomics data, phenomics data, or transcriptomics data, or any combination thereof. In some embodiments, omics data may be associated with image data. In some embodiments, omics data may be spatially associated with image data. In some cases, omics data may be spatially associated with image data in two dimensions, three dimensions, or both. In some embodiments, omics data may be associated with image data as metadata, or as an overlay on the image, or both.In some embodiments, the data may include, in some non-limiting examples, patient data, demographic data, diagnostic data, disease data, treatment data, laboratory data, or any combination thereof.
[0066] Image acquisition The systems, media, and methods disclosed herein may include retrieving images from one or more tissue scans of tissue stained with a morphological marker. In one embodiment, the tissue may be incubated with a protein reporter. In one embodiment, the tissue may be incubated with an RNA reporter. In one embodiment, the systems, media, and methods disclosed herein may include removing the reporter from the tissue. In one embodiment, the reporter may be bleached. In one embodiment, the reporter may be cleaved. In one embodiment, the reporter may be bleached, cleaved, or both by applying one or more of the following: ultraviolet light, heat, proteases, endonucleases, nucleases, esterases, ribonucleases, RNase A, RNase Tl, RNase H, disulfate-binding reducing agents (e.g., dithiothreitol, tris(2-carboxyethyl)phosphine), salt buffers, alkalis, hydrogen bond destabilizing solvents (e.g., formamide, DMSO), or any combination thereof.
[0067] In some embodiments, the systems, media, and methods disclosed herein may include image retrieval. In some embodiments, the images may include images derived from a microscope. In some embodiments, the images derived from a microscope may include images from an optical microscope, an electron microscope, or a scanning probe microscope, or any combination thereof. In some embodiments, the images may include images from a spatial molecular imaging device.
[0068] In one embodiment, a morphological signal is acquired. In one embodiment, an RNA-protein signal is acquired. In one embodiment, both a morphological signal and an RNA-protein signal are acquired. In one embodiment, both the morphological signal and the RNA-protein signal are acquired using High Dynamic Range (HDR) mode. In one embodiment, the signals are acquired using HDR to improve the dynamic range of the signals. As shown in Figure 5, Image 1 was obtained with a 0.0125-second exposure with detail in the bright areas in one non-limiting example, and Image 2 was obtained with a 0.1-second exposure with detail in the dark areas. In one embodiment, the resulting image can be generated by splitting the accumulator with a counter. As shown in a non-limiting example in Figure 5, the final intensity scale was equivalent to a 0.0125-second exposure, which is 1 / 8 of the nominal exposure time.
[0069] In some embodiments, the systems, media, and methods disclosed herein may include combining protein images expressing similar structures or cell types, or both, with multi-channel input morphological images. In some embodiments, the systems, media, and methods disclosed herein may include adding additional input channels. In some embodiments, the additional input channels may include bioRNA spot density input to compensate for missing signals from morphological markers. In some embodiments, a spatial molecular imager may be used to image the nucleus and DAPI for staining. In some embodiments, the spatial molecular imager may be a CosMx™ instrument. In some embodiments, the systems, media, and methods disclosed herein may include biological samples labeled with photocleavable morphological markers. In some embodiments, DAPI channel acquisition for one FOV may cause reporter cleavage or bleaching, or both, in adjacent FOVs. In some embodiments, the order of image or channel acquisition, or both, is determined. In some cases, the order of image or channel acquisition may be determined along the edges of the FOV, particularly in adjacent scans, to prevent visible and irrecoverable signal loss.
[0070] In some embodiments, the systems, media, and methods disclosed herein may include combining multi-channel input morphological images. In some embodiments, the morphological images may be used for different markers expressing the nucleus, cytoplasm, or cell membrane, or any combination thereof. In one non-limiting example, a five-channel morphological image used for segmentation may include blue (B), green (G), yellow (Y), red (R), and UV (U).
[0071] In some embodiments, the systems, media, and methods disclosed herein further include a multi-stage image acquisition process. In some embodiments, the multi-stage image acquisition process can be performed to avoid signal loss. In some embodiments, the systems, media, and methods disclosed herein may include a software module configured to reduce a 3D scanned image to a 2D image. In some embodiments, reducing a 3D scanned image to a 2D image can obtain an optimal focal region in the z slice. In some embodiments, the systems, media, and methods disclosed herein may include a software module configured to deconvolve a 2D image. In some embodiments, deconvolution of a 2D image can provide an optimal focal region for the image while maintaining the z position of the image in the 3D scanned image stack. In some embodiments, deconvolution of a 2D image can provide optimal clarity for the image while maintaining the z position of the image in the 3D scanned image stack. As shown in the non-limiting example of Figure 6, two-stage image acquisitions P01 and P02 were performed to avoid signal loss from DAPI acquisition using UV light. In step P01, the system acquired four channels—B, G, Y, and R—of a 3D high dynamic range (HDR) z-stack spanning the tissue for the selected FOV. For example, in step P02 in Figure 6, the system separated the DAPI from the other two morphological channels for registration with the previous step P01 into 3D coordinate space. After the POI step was completed, for example, the FOV location was revisited to acquire a new Z-stack with the DAPI and two additional morphological channels, e.g., B and Y. The B and Y morphological channels were necessary in one non-limiting example to register the P01 and P02 z-stacks in the same 3D coordinate space, as shown in the step register in Figure 6. In a non-limiting example, after the stacks were aligned, they were combined to produce a complete 5-channel Z-stack, as shown in step P99 in Figure 6. In some embodiments, once the morphological scan is performed, protein images can be acquired for all FOVs following RNA acquisition.In some embodiments, in each cycle, images can be registered in the same 3D coordinate space using reference markers that can be placed on the tissue. In some embodiments, the systems, media, and methods disclosed herein may include acquiring time-lapse images of a sample of interest using a multi-stage image acquisition process.
[0072] Image enhancement and bleaching correction In some embodiments, photocleavable markers are used to obtain an unlimited set of source input channels by a process of repeatedly re-staining and removing the markers using UV irradiation and chemical washing within the tissue. In some cases, repeated scanning of tissue regions introduces some photobleaching and / or attenuation at the edges of the FOV. In one non-limiting example, Figure 7 shows darkened bands on the right and bottom of the image that reduce the signaling and detectability of cells in these regions. In some cases, the closer the distance between FOVs and the greater the number of repeated acquisition cycles, the greater the photobleaching effect, as shown in the non-limiting example in Figure 7.
[0073] In some embodiments, the systems, media, and methods disclosed herein may further include normalizing the intensity of the attenuation region using the nearest unbleached portion of the image to correct for edge artifacts. In some embodiments, measurements can be taken where bleaching occurs by downsampling the image and dividing it into smaller blocks, as shown, for example, in Figure 8A. In some embodiments, a transition smoothing function can be applied during the scaling calculation, as shown, for example, in Figure 8B. In one non-limiting example, a scheme for normalization and scaling factor calculation is shown in Figure 8C. In some embodiments, the lower 25th percentile and upper 75th percentile values can be calculated in each block, as shown in Figure 8C. In some embodiments, the ratio of the lower 25th percentile and upper 75th percentile values can be calculated, as shown in Figure 8C. In some embodiments, the intensity attenuation at a given location can be normalized using the low intensity quantile ratio by applying a scaling factor, for example, shown in Figure 8B. In some embodiments, the high intensity values for each block can be determined using the high quantile ratio. In some embodiments, this abrupt drop in high-intensity quantiles between adjacent blocks can be used to determine the location where bleaching begins. In some embodiments, the systems, media, and methods disclosed herein may further include applying bleaching correction to morphological and protein image acquisition.
[0074] In some embodiments, images can be enhanced by conventional image enhancement techniques, including denoising, deblurring, or image sharpening, or any combination thereof. In some embodiments, one or more denoising algorithms can be applied to one or more images. One or more denoising algorithms include filter-based (Wiener, Gaussian, median), total variation (TV) denoising, and deep learning-based approaches using wavelet transforms. Deblurring algorithms include nearest neighbor deconvolution, Richardson-Lucy deconvolution, and total variation (TV) regularization, or any combination thereof. In some embodiments, image enhancement algorithms can include histogram equalization, contrast-limited adaptive histogram equalization (CLAHE), contrast stretching, gamma correction, unsharp masking, or any combination thereof. In some embodiments, nearest neighbor deconvolution can be used. In some cases, nearest neighbor deconvolution can include removing blur signals by utilizing adjacent z-planes. In some embodiments, nearest neighbor deconvolution does not require point spread function (PSF) data. In one embodiment, nearest neighbor deconvolution occurs at the previous and next z position I zprev and I znext Using the image from, corrections can be performed on the image at position z, I, as described below.
[0075] I decon =IA*(G(I zprev ,σ)+G(I znext ,σ)) In the formula, A is a multiplier coefficient of 0-1. G = Gaussian kernel, where σ is the optical numerical aperture (NA) and z(δ z It is determined by the sampling rate in ).
[0076] In some embodiments, the relationship between NA and the width of the blur function can be described as NA = n sin (a) using the objective lens cone angle, as shown in Figure 22, where n is the refractive index. Thus, in some embodiments, a = arcsin (NA / n). In some non-restrictive examples, using a water-based objective lens, n = 1.33333 and NA = 1.10, it can be derived that a = 55.59°. For example, using Figure 22, the Gaussian kernel width is σ = δ z It can be derived as tan(a).
[0077] RNA spot detection and density map generation In one embodiment, RNA molecules can be detected using single-molecule fluorescent barcodes that appear in the image as visible spots. In one embodiment, combinations of spots can correspond to specific RNA transcripts encoded in the chemical reagents used. In one embodiment, the systems, media, and methods disclosed herein may further include applying the location and density of raw spots to infer further cell morphology. In one embodiment, the systems, media, and methods disclosed herein may further include applying the location and density of raw spots to enhance further cell morphology. In one embodiment, each RNA spot can be detected from image slices acquired across tissue sections. In one embodiment, detection can be performed using a Gaussian-Laplace (LoG) bandpass filter. In one embodiment, the LoG bandpass filter can be designed to match the morphology of the reporter signal in each channel. In one embodiment, the precise location of the spots can be estimated using 2D parabolic fitting. In one embodiment, chromatic aberration of the optical system can be accounted for to ensure spatial alignment between each channel. In one embodiment, the lens does not focus all wavelengths (channels) to the same point. In one non-limiting example, Figure 9 shows the aberrations generated in the axial direction (center figure) and the lateral direction (right figure) compared to a perfect lens on the left. In some embodiments, the systems, media, and methods disclosed herein may further include applying a unique calibration that determines the z offset for each wavelength during acquisition in order to correct axial chromatic aberration.
[0078] In some embodiments, the systems, media, and methods disclosed herein may further include performing lateral chromatic aberration correction for each multichannel image. In some embodiments, performing lateral chromatic aberration correction can be done using a pre-calibrated transform. In some non-limiting examples, for example, as shown in Figure 10, there are high-confidence (high-signal) spot locations greater than 10M detected for each FOV resolved to one-tenth of a pixel (-12nm). For example, as shown in Figure 10, when high-density spots are plotted in image space, they coincide with cellular structures and can be used to infer and enhance cellular morphology. In some embodiments, the spot locations can be binned into lower-resolution pixel space. For example, as shown in Figure 10, this binning creates a detailed morphological image that can be successfully segmented using a cell segmentation pipeline. In some embodiments, the systems, media, and methods disclosed herein may further include applying a binned two-dimensional histogram of spot locations as an additional input segmentation channel. As shown in Figure 11, in non-limiting examples, raw spots were detected at a resolution of 0.1 pixels to form a detailed morphological density image. In some non-limiting examples, as shown in Table 1, the spots for each FOV were resolved to 1 / 10 of a pixel. In non-limiting examples, X and Y were decipixels with a range of [0.42560].
[0079] [Table 1]
[0080] For example, as shown in Figure 12, a 1330x1330 grid was superimposed on the range [0.42560]. For example, the total number of spots in each box of the grid was tallied to determine the grayscale value in the resulting spot density map image. For example, the density map image was resized to a 4256x4256 FOV dimension to reduce noise in the density map image. For example, the image was then used in a cell segmentation workflow as a standard morphology channel obtained by fluorescence staining.
[0081] Multimodal cell segmentation In some embodiments, the systems, media, and methods disclosed herein may include searching for five morphological channels. In some embodiments, searching for five morphological channels can express a greater description of the cell boundary than using only nuclear channels. In some embodiments, the systems, media, and methods disclosed herein may include searching for 64 or more different protein channels. In some embodiments, approximately 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100 different protein channels may be searched. For example, some protein channels are shown, including channels for histone staining, DAPI staining, and Soma rRNA staining. For example, some segmented channels are shown, including readout density heatmaps, segmentation based on all three morphological staining images, and segmentation based on density heatmaps and histone staining.
[0082] In some embodiments, the number of different protein channels to be searched may be indicated by the number of protein markers, scanning cycling parameters, additional transcription-based information using spot density maps, or any combination thereof, as shown in Figure 14. In some embodiments, the systems, media, and methods disclosed herein may include combining morphological and protein images having the same morphological type into a smaller set of input channels. In some embodiments, each protein may be annotated to describe its morphology. In some embodiments, the annotations may include cytoplasmic or nuclear expression, or both. In some embodiments, the annotations may describe different cell type expression. In some cases, cell type expression may include glial, astrocyte, and other neuronal cell types, or any combination thereof. In some embodiments, these annotations may be used to combine each protein image together based on cells having the same morphology or the same cell type, or both. In some embodiments, a projection method may be used to fold multiple protein images having the same morphology into a 5-channel image. In some embodiments, the projection method may be one or more of maximal intensity projection, principal component analysis, singular value decomposition, multi-reference alignment, density map averaging, or any combination thereof. In some embodiments, readout or spot density images, or both, may be added as additional channels. In some embodiments, all input channels may be background subtracted and normalized before the segmentation step. In some embodiments, cell segmentation may be one or more of nuclear segmentation, cytoplasmic segmentation, or extracellular segmentation, or any combination thereof. In some embodiments, the outputs from each cell segmentation may be combined as the final segmentation output. In some embodiments, per-cell and per-compartment statistics may be performed.
[0083] In some embodiments, image preprocessing to minimize image artifacts and improve image quality may be applied to the raw image before processing by the cell segmentation module. In some embodiments, the transcription spot density heatmap may be processed using image preprocessing techniques. In some embodiments, preprocessing can be used to achieve a better signal-to-noise ratio between intracellular and extracellular regions. An example of denoising preprocessing used for the spot density heatmap is shown in Figure 24, but is not limited to this. In some embodiments, the spot density heatmap can be used without further preprocessing. In some embodiments, the image preprocessing step can use a combination of conventional image processing techniques such as deconvolution, Gaussian blurring, normalization, feature extraction, Fourier transform, linear filtering, contrast enhancement, binarization, or any one or more of these. In some embodiments, the image preprocessing step can utilize machine learning methods, such as the NoiseSelf and celpose 3.0 image reconstruction models.
[0084] In some embodiments, nuclear channels can be segmented to provide intracellular segmentation. In some embodiments, nuclear channels can be segmented in addition to other channels, including membranes and cytoplasmic contents. In some embodiments, multi-stage segmentation produces intracellular results that define the separation of nuclear and cytoplasmic regions. In some embodiments, overlap between cells can be measured using an IoU (Issue of Unit) score. For example, in some embodiments, if the IoU score is low, the overlapping region can be assigned to a segment with a higher confidence signal. In some embodiments, nuclear channels can have a stronger signal. In some cases, low crossover pixels can be assigned as nuclear segments. In some embodiments, nuclear segments are not cytoplasmic. In some embodiments, if the crossover is high, the two segmentation results can be merged to combine both nuclear and cytoplasmic. In some embodiments, the segmentation system generates masks that define compartments. In some embodiments, a compartment can be defined as nuclear or cytoplasmic. In some embodiments, a compartment can be defined as the region of each cell detected in the image. In some embodiments, measurements and statistics can be performed for each cell by using these masks. In one embodiment, specific membranes and extracellular regions can provide biological information.
[0085] In some embodiments, an interlocking (IoU) based combination approach, such as shown in Figure 23, can be used to combine multiple segmentation results. In some embodiments, the IoU-based combination approach can be implemented in a stepwise manner. In some embodiments, implementing the IoU-based combination in a stepwise manner allows for the combination of multiple modalities of morphology in achieving the final result. In some embodiments, the IoU-based combination approach can combine the 2D cell segmentation results of individual 2D images with the 3D cell segmentation volumetric measurement results of the corresponding 3D scanned image stack. In some embodiments, correlation can include labeling each cell at all z-positions using a maximum IoU score for consistent mapping. In some embodiments, cell label continuity through z-slices can be determined using other measures, such as cell focus profiles across z-slices. In some embodiments, cells in a 3D stack can be labeled using analysis of cell centroid and shape changes across z-slices. In some embodiments, a 3D stack can be segmented directly using a machine learning model that accepts a 3D stack input. In some embodiments, extending this configuration from two dimensions to three dimensions, as shown in Figure 23, may include performing multimodal segmentation in two dimensions for each z-slice, followed by performing a three-dimensional cell label stitch to correlate the cell IDs within each z-slice and form three-dimensional cell labels. Figure 23 shows a non-limiting example of a three-dimensional segmentation pipeline.
[0086] In some embodiments, the localization of intracellular objects, such as the nucleus, can be mapped to the cell body contour using the maximum IoU score. In some embodiments, the localization of cells in different source images, such as images taken at different z-coordinates or time points, can be mapped to each other using the maximum IoU score. In some embodiments, localization points from different source images taken at different z-coordinates or time points that yield the maximum IoU score when matched can be mapped together to form a stacked image. In some embodiments, spot images can be registered to the same criteria as morphological staining images. In some embodiments, time-lapse images can be taken on the sample of interest. In some embodiments, an IoU-based combination approach can be applied to combine cell segmentation results of the same sample at different time points to track the behavior of the same cells over time.
[0087] In one embodiment, extracellular regions can be detected. In one embodiment, extracellular regions can be included as part of the segmentation results. In one embodiment, the extracellular region may include one or more structures in the extracellular space of, for example, the brain. In one embodiment, the extracellular structure may include one or more vesicles. In one embodiment, each connected component can be detected and masked. In one embodiment, masking can be performed using automated thresholding techniques. In one embodiment, automated thresholding techniques may include finding high-intensity regions within input morphological channels or using a neural network model trained to detect cell membranes or specific extracellular regions.
[0088] In one non-limiting example, as shown in Figure 15, all proteins expressing a microglial cell type were combined with a microglial morphology channel using an Ibal marker. As shown in Figure 15, the input image channel included morphology channels and spot density map channels for the same biological sample. In some embodiments, maximum intensity pixel values may be projected onto all input morphology channels. In some embodiments, readout or spot density imaging, or both, may be added as separate channels. In some embodiments, nuclear, cytoplasmic, or extracellular segmentation may be performed on each supply channel. In some embodiments, a cell segmentation algorithm may be selected based on one or more supply channels. In some embodiments, cell segmentation may be based on one or more algorithms. In some cases, one or more algorithms may include Cellpose deep learning algorithms, adaptive thresholding, watershade algorithms, thresholding, region-based segmentation, active contours, convolutional neural networks (CNNs), level set methods, graph-based algorithms, fuzzy c-means clustering, deep watershade, image morphology, or Markov random fields, or any combination thereof. In some embodiments, nuclear segmentation, cytoplasmic segmentation, and extracellular segmentation may be based on the same algorithm. In some embodiments, nuclear segmentation, cytoplasmic segmentation, and extracellular segmentation may be based on different algorithms. For example, as shown in Figure 15, the selected algorithm was used for a neural network model, and the model was generated using a ground truth dataset. In some embodiments, a model may be selected to fit a particular input morphology by comparing an image and a model style vector.In some embodiments, after neural network inference, the output can be generated by measuring the image properties within each cell or cell partition mask region. In some embodiments, the outputs from all models can be combined into a single final output. In some embodiments, cell-by-cell statistics, or both, are performed.
[0089] In some embodiments, the systems, media, and methods disclosed herein may include multiple input channels in a model. Figure 16 shows a non-limiting example for generating a new N-channel model that fits a new number of morphological channels. In some embodiments, the systems, media, and methods disclosed herein may include applying transfer learning techniques. In some embodiments, transfer learning techniques can be applied to reuse the weights of an existing trained deep learning model in a higher dimension. In some embodiments, new channels can be added by copying weights from the original two channels, as shown in the weight transfer block of Figure 16, where a predefined two-channel model is used as the starting model. In some embodiments, the new model may be trained on newly defined weights and a ground truth dataset. In some embodiments, a new N-channel model may then be generated. Various new two-channel and five-channel outputs are shown by Figures 17A–17B, where Figure 17A shows a non-limiting example of a new two-channel model, and Figure 17B shows a non-limiting example of a new five-channel model.
[0090] In some embodiments, separate nuclear-only segmentation and cytoplasmic segmentation can be combined. In some embodiments, the combination of nuclear and cytoplasmic segmentation can ensure segmentation output for tissues with a complete set of nuclear and membrane signals. In some embodiments, the combination of nuclear and cytoplasmic segmentation can ensure segmentation output for regions with weak membrane signals. In some embodiments, both nuclear and cytoplasmic segmentation can provide subcellular resolution and precision. As shown in Figures 18A–18D, which are non-limiting embodiments, nuclear expression images were selected as input for nuclear segmentation, and all morphological input channels, including the nucleus, were supplied for cytoplasmic segmentation. As shown in Figures 18A–18D, the results from nuclear and cytoplasmic segmentation were combined by analyzing overlap and crossover (IoU) between segmentation results. Various markers were applied to the cell segmentation results. For example, Figure 18A shows neurons to which the nuclear marker HistoneH3 was applied, Figure 18B shows neurons to which the rRNA membrane marker was applied, Figure 18C shows astrocytes to which the GFAP marker was applied, and Figure 18D shows nuclei to which the DAPI marker was applied.
[0091] In some embodiments, the systems, media, and methods disclosed herein may further include segmentation of extracellular processes as a third segmentation step. In some embodiments, the extracellular region may be a structure present in the image but without nuclear information. In some embodiments, the systems, media, and methods disclosed herein may further include tracing each of one or more neuronal processes using a filter. In some embodiments, the filter may include a Gaussian-Laplace (LoG) filter. In some embodiments, as shown in Figure 19, the filter may be followed by a concatenated component algorithm for separating the detected processes into individual components. In some embodiments, the LoG filter may extract thin neuronal processes from normalized brain-specific markers. In some embodiments, the brain-specific markers may include astrocytes or microglia, or both, input channels. In some embodiments, a threshold may be applied to the LoG output to create a mask. In some embodiments, the mask is then expanded to match the size of the neuron. In some embodiments, each process may be detected using standard thresholding techniques. In some cases, standard thresholding techniques may include Otsu, Moments, Li, Huang, and Bernsen local thresholding methods, or any combination thereof. In some embodiments, a connected component algorithm can be used to number each connected neuron. In some embodiments, a connected component algorithm can be used to number cells segmented from an image. In some embodiments, each protrusion can be detected using a machine learning model trained to detect cellular structures.
[0092] In one embodiment, a process that completely encapsulates the cell body can be merged with the nuclear segment and marked as a cell. In another embodiment, a process that does not have a connection to the cell body can be marked as an object, as shown in Figures 20A–20C. For example, Figure 20A shows an overlay of cells and extracellular processes masked as shown in Figure 20B to separate the extracellular segment, as shown in Figure 20C, for example. In one embodiment, the merging between cellular processes and the nucleus can be performed using a machine learning algorithm trained to perform the merge operation.
[0093] In some embodiments, each neuronal process is individually identifiable as one or more objects, so transcriptomics, cell typing, and other spatial data analyses can be performed in the same manner as single-cell analysis. In some embodiments, all three segmentation steps can generate single-cell labels using cellular measurements and statistics, as well as object-specific measurements. In some embodiments, the cell labels can be further divided into nuclear and cytoplasmic compartments. In one non-limiting example relating to brain tissue, each individual cell type was marked with a different channel and identified as shown in Figure 21.
[0094] Statistics per cell In one embodiment, the segmentation output may provide statistics for each detected cell. The statistics may include statistics describing fluorescence properties, spatial location, or features, or any combination thereof. In one embodiment, these properties may include mean, median, and maximum fluorescence intensity. In one embodiment, these properties may include various shape descriptors. In some cases, the shape descriptors may include circularity, compactness, aspect ratio, convexity, stereochemistry, eccentricity, and circumference.
[0095] These shape descriptors are described below. Aspect ratio = W bb / Hbb Here, W bb = the width of the bounding box, and H bb = the height of the bounding box.
[0096] In some embodiments, to limit the aspect ratio value to 0 - 1, the width can be assigned a shorter value of the bounding box size. In some embodiments, to limit the aspect ratio value to 0 - 1, the height can be assigned a longer value of the bounding box size.
[0097] Circularity = (4 * π * A) / (P conv * P conv ) Here, A is the area of the cell, and P conv is the convex perimeter to avoid concave irregularities.
[0098] Compactness = (4 * π * A) / (P * P) Here, A is the area of the cell, and P is the perimeter of a regular cell.
[0099] Convexity = P conv / P Eccentricity = L minor / L major Here, L minor is the length of the minor axis, and L major is the length of the major axis.
[0100] Solidity = \frac{A}{A_{conv}} Here, A is the area of the cell, and A conv is the convex area.
[0101] In one aspect, the measurement criteria can be calculated for all compartments including the nucleus and cytoplasmic compartments. In one aspect, these properties can be utilized to analyze cell sub - populations. In one aspect, other properties such as texture can also be incorporated.
[0102] In one embodiment, when a small section of a tissue sample is imaged, some cells may be segmented across different fields of view (FOV). In one embodiment, when used in the context of single-cell transcriptome analysis, these partial cell segments at FOV crossovers may create an incomplete mask, leading to inaccurate analysis. In one embodiment, segmentation output measurements can identify cells segmented at image boundaries and filter them out downstream.
[0103] Machine Learning In some embodiments, the systems, media, and methods disclosed herein may include intracellular segmentation further comprising one or more deep learning models. In some embodiments, one or more deep learning models may be trained using images of cells or tissues. In some embodiments, the images of cells or tissues may include microscopic images, morphological images, or fluorescence images, or any combination thereof. In some embodiments, the deep learning models may be trained to segment specific tissue types. In some embodiments, the deep learning models may be trained to restore image quality for better cell segmentation results.
[0104] In some embodiments, the systems, media, and methods disclosed herein may include training a machine learning model, applying a machine learning model, or both. In some embodiments, the machine learning model may perform dimensionality reduction. In some embodiments, dimensionality reduction may be performed by a nonlinear dimensionality reduction algorithm. In some embodiments, the nonlinear dimensionality reduction algorithm may include summon mapping, principal curves and manifolds, Laplace eigenmaps, isomaps, local linear embeddings, local tangent space alignment, maximum variance expansion, Gaussian process latent variable models, t-distribution stochastic neighborhood embeddings, relational perspectives, contagion maps, curve component analysis, curve distance analysis, heteromorphic dimensionality reduction, manifold alignment, diffusion maps, local multidimensional scaling, nonlinear PCA, data-driven high-dimensional scaling, manifold sculpting, RankVisu, topological constraint conformal embeddings, uniform manifold approximation or projection (UMAP), or any combination thereof. In some embodiments, the UMAP algorithm may apply a feedforward neural network to a subset of the data. In some embodiments, the UMAP algorithm may project manifold clustering onto the entire dataset. In some embodiments, the UMAP algorithm may be a feedforward algorithm. In some embodiments, the feedforward neural network may be trained to approximate an identity function. In some embodiments, approximating an identity function may involve mapping a vector of values to the same vector. In some embodiments, the feedforward neural network can be used for dimensionality reduction. In some embodiments, one of the hidden layers in the network may be restricted to contain only a small number of network units.
[0105] In some embodiments, the systems, media, and methods disclosed herein may include training a machine learning model, applying a machine learning model, or both. In some embodiments, the machine learning model may include an unsupervised machine learning model, a supervised machine learning model, a semi-supervised machine learning model, or a self-supervised machine learning model, or any combination thereof. In some embodiments, the supervised machine learning model may include, for example, a random forest, a support vector machine (SVM), a neural network, or a deep learning algorithm, or any combination thereof.
[0106] In some embodiments, the platforms, systems, media, and methods disclosed herein may include machine learning models that utilize one or more neural networks. In some embodiments, the neural network may learn relationships between input datasets and target datasets. In some embodiments, the neural network may be a software representation of the human nervous system. In some embodiments, the human nervous system may be a cognitive system. In some embodiments, the neural network captures the “learning” and “generalization” capabilities used by humans. In some embodiments, the machine learning algorithm may include neural networks, including CNNs. In some embodiments, the structural components of the machine learning algorithm may include one or more of CNNs, recurrent neural networks, augmented CNNs, fully connected neural networks, deep generative models, transformers, Boltzmann machines, or any combination thereof.
[0107] In some embodiments, a neural network may include a set of layers called “neurons.” In some embodiments, a neural network may include an input layer where data is presented; one or more internal and / or “hidden” layers; and an output layer. In some embodiments, neurons may be connected to neurons in other layers via connections that have weights, which are parameters that control the strength of the connections. In some embodiments, the number of neurons in each layer may relate to the complexity of the problem to be solved. In some embodiments, the minimum number of neurons required in a layer may be determined by the complexity of the problem, and the maximum number may be limited by the neural network’s ability to generalize. In some embodiments, an input neuron may receive presented data and then transmit that data to a first hidden layer via connection weights that are modified during training. The first hidden layer may process the data and transmit the results to the next layer via a second set of weighted connections. In some embodiments, each subsequent layer may “pool” the results from the previous layer into more complex relationships. In some embodiments, while conventional software programs require the writing of specific instructions to perform a function, neural networks are programmed by training them with a known set of samples, allowing them to modify themselves during (and after) training to provide desired outputs, such as output values. In some embodiments, when new input data is presented to the neural network after training, it is configured to generalize what it "learned" during training and apply what it has learned from training to new, previously unseen input data in order to generate an output relevant to that input.
[0108] In some embodiments, the neural network may include an artificial neural network (ANN). In some embodiments, the ANN may be a machine learning algorithm that can be trained to map input datasets to output datasets, and the ANN includes a group of interconnected nodes organized into multiple layers of nodes. For example, an ANN architecture may include at least an input layer, one or more hidden layers, and an output layer. In some embodiments, the ANN may include any total number of layers and any number of hidden layers, the hidden layers acting as trainable feature extractors that enable mapping a set of input data to output values or a set of output values. The deep learning algorithms used herein (such as deep neural networks (DNNs)) are ANNs including multiple hidden layers, e.g., two or more hidden layers. Each layer of the neural network may include a number of nodes (or "neurons"). In some embodiments, a node may receive input coming directly from either input data or the output of a node in the previous layer and perform a specific operation, e.g., an addition operation. In some embodiments, the connection from the input to the node may be associated with a weight or weight coefficient.
[0109] In some embodiments, a node can sum the products of all pairs of inputs and their associated weights. In some embodiments, the weighted sum can be offset by a bias. In some embodiments, the output of a node or neuron can be gated using a threshold or activation function. In some embodiments, the activation function may be a linear or nonlinear function. In some embodiments, the activation function may be, for example, a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or other functions such as saturated hyperbolic tangent, identity, binary step, logistic equation, arctan, soft sine, parametric rectified linear unit, exponential linear unit, soft plus, Bent's identity, soft exponential, sine, Gaussian, or sigmoid function, or any combination thereof.
[0110] In some embodiments, the weight coefficients, bias values, thresholds, or other computational parameters of a neural network may be “taught” or “learned” during the training phase using one or more sets of training data. For example, the parameters may be trained using input data from the training dataset and gradient descent or backpropagation so that the output values computed by the ANN match examples contained in the training dataset.
[0111] The number of nodes used in the input layer of an ANN or DNN, including its increments, may be at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, or more than 100,000.
[0112] In other examples, the number of nodes used in the input layer, including its increments, may be up to approximately 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9,000, 8,000, 7,000, 6,000, 5,000, 4,000, 3,000, 2,000, 1,000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10, or less than 10. In some examples, the total number of layers used in an ANN or DNN (including input and output layers), including its increments, may be at least approximately 3, 4, 5, 10, 15, 20, or more. In other examples, the total number of layers, including its increments, may be up to approximately 20, 15, 10, 5, 4, 3, or less.
[0113] In some examples, the total number of learnable or trainable parameters used in an ANN or DNN, such as weighting coefficients, biases, or thresholds, including increments, may be at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, or more than 100,000. In other examples, the number of learnable parameters, including their increments, may be up to approximately 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9,000, 8,000, 7,000, 6,000, 5,000, 4,000, 3,000, 2,000, 1,000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10, or less than 10.
[0114] In some embodiments, the platforms, systems, media, and methods disclosed herein include machine learning models and may include neural networks such as deep CNNs. In some embodiments where a CNN is used, the network can be constructed with any number of convolutional layers, extension layers, or fully connected layers. In some embodiments, the number of convolutional layers may be between 1 and 10. In some embodiments, the number of extension layers may be between 0 and 10. In some embodiments, the total number of convolutional layers (including input and output layers) may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more, and the total number of extension layers may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more. In some embodiments, the total number of convolutional layers may be at most about 20, 15, 10, 5, 4, 3, or less, and the total number of extension layers may be at most about 20, 15, 10, 5, 4, 3, or less. In some embodiments, the number of convolutional layers may be between 1 and 10, and the number of fully connected layers may be between 0 and 10. In some embodiments, the total number of convolutional layers (including input and output layers) may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more, and the total number of fully connected layers may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more. In some embodiments, the total number of convolutional layers may be at most about 20, 15, 10, 5, 4, 3, 2, 1, or less, and the total number of fully connected layers may be at most about 20, 15, 10, 5, 4, 3, 2, 1, or less.
[0115] In one embodiment, the machine learning algorithm may include neural networks such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), augmented CNNs, fully connected neural networks, deep generative models or deep restricted Boltzmann machines, or any combination thereof.
[0116] In some embodiments, a machine learning model may include one or more CNNs. In some embodiments, the CNN may be a deep and feedforward ANN. In some embodiments, the CNN may be applicable to the analysis of visual images. In some embodiments, the CNN may include an input layer, an output layer, and a number of hidden layers. In some embodiments, the hidden layers of the CNN may include a convolutional layer, a pooling layer, a fully connected layer, and a normalization layer. In some embodiments, the layers may be configured in three dimensions: width, height, and depth.
[0117] In some embodiments, a convolutional layer can apply a convolution operation to an input and pass the result of the convolution to the next layer. In some embodiments, to process images, the convolution operation can reduce the number of available parameters, allowing the network to be deeper with fewer parameters. In some embodiments of the neural network, each neuron can receive input from several locations in the previous layer. In some embodiments of the convolutional layer, a neuron can receive input from only a limited subregion of the previous layer. In some embodiments, the parameters of the convolutional layer can include a set of learnable filters (or kernels). In some cases, the learnable filters have a small receptive field and can extend across the entire depth of the input volume. In some cases, during the forward pass, each filter can be convolved across the width and height of the input volume, and the dot product between the filter's entry and the input can be computed to generate a two-dimensional activation map of that filter. In some embodiments, as a result, the network can learn filters that are activated when certain types of features are detected at certain spatial locations in the input.
[0118] In some embodiments, the machine learning model may include an RNN. An RNN is a neural network having circular connections that can encode and process sequential data. In some embodiments, the RNN may include an input layer configured to receive a sequence of inputs. In some embodiments, the RNN may further include one or more hidden recurrent layers that maintain a state. In some embodiments, at each step, each hidden recurrent layer may compute the layer's output and the next state. In some embodiments, the next state may depend on the previous state and the current input. In some embodiments, the state may be maintained between steps and may capture dependencies within the input sequence.
[0119] In some embodiments, the RNN may be a long-term short-term memory (LSTM) network. In some embodiments, the LSTM network may consist of LSTM units. In some embodiments, an LSTM unit may include a cell, an input gate, an output gate, and a forget gate. In some embodiments, the cell may be responsible for tracking dependencies between elements in an input sequence. In some embodiments, the input gate may control the extent to which new values flow into the cell, the forget gate may control the extent to which values remain in the cell, and the output gate may control the extent to which values in the cell are used to compute the output activation of the LSTM unit.
[0120] In some embodiments, the neural network may include an attention mechanism (such as a transformer). In some embodiments, the attention mechanism can focus on, or "attention to," a specific input region while ignoring other input regions. In some embodiments, this may improve model performance because the specific input region may be less relevant. In some embodiments, at each step, the attention unit may compute the inner product of the context vector and the input at the step, among other operations. In some embodiments, the output of the attention unit may define where the most relevant information in the input sequence is placed. In some embodiments, the attention mechanism may include a vision transformer. In some embodiments, the vision transformer may include a natural language model. In some cases, the vision transformer may include a masked autoencoder, a Swin transformer, a vector-quantized variational autoencoder, or another type of vision transformer. In some embodiments, the vision transformer may be combined with a generative adversarial network.
[0121] In some embodiments, the pooling layer may include a global pooling layer. In some embodiments, the global pooling layer may combine the outputs of neuron clusters in one layer with a single neuron in the next layer. For example, a max pooling layer may use the maximum value from each cluster of neurons in the previous layer. Similarly, an average pooling layer may use the average value from each cluster of neurons in the previous layer.
[0122] In some embodiments, a fully connected layer can connect all neurons in one layer to all neurons in another layer. In some neural networks, each neuron can receive input from several positions in the previous layer. In some cases, in a fully connected layer, each neuron can receive input from all elements in the previous layer.
[0123] In some embodiments, the normalization layer can be a batch normalization layer. In some embodiments, the batch normalization layer can improve the performance and stability of the neural network. In some embodiments, the batch normalization layer can provide any layer in the neural network with an input that has a mean / unit variance of 0. In some embodiments, the advantages of using a batch normalization layer may include faster trained networks, higher learning rates, easier weight initialization, more viable activation functions, and a simpler process for creating deep networks.
[0124] In some embodiments, the trained algorithm can be configured to accept multiple input variables and generate one or more output values based on the multiple input variables. In some embodiments, the trained algorithm can include a classifier such that each of the one or more output values can contain one of a fixed number of possible values. In some embodiments, the possible values can include a linear classifier, a logistic regression classifier. In some embodiments, the algorithm can represent the classification of a biological sample or subject by the classifier, or both. In some embodiments, the trained algorithm can include a binary classifier. In some embodiments, each of the one or more output values can contain one of two values. In some embodiments, the values can include {0, 1}, {positive, negative}, or {high risk, low risk}, representing the classification of a biological sample or subject by the classifier, or both. In some embodiments, the trained algorithm may include another type of classifier such that each of the one or more output values can contain one of more than two values. In some embodiments, the values can include {0, 1, 2}, {positive, negative, uncertain}, or {high risk, medium risk, low risk}, representing the classification of a biological sample or subject by the classifier, or both. In some embodiments, the output values may include descriptive labels, numerical values, or combinations thereof. In some embodiments, some of the output values may include descriptive labels. In some embodiments, some of the output values may include numerical values such as binary values, integer values, or continuous values. Such binary output values may include, for example, {0, 1}, {positive, negative}, or {high risk, low risk}. Such integer output values may include, for example, {0, 1, 2}. Such continuous output values may include, for example, probability values of at least 0 and less than or equal to 1. Such continuous output values may include, for example, unnormalized probability values of at least 0. Such continuous output values may indicate the prognosis of the subject for cancer-related categories.Some numerical values can be mapped to descriptive labels, for example, by mapping 1 to "positive" and 0 to "negative".
[0125] In some embodiments, the classification of a sample can assign an output value of “positive,” “negative,” “uncertain,” or 2 if the sample is not classified as “positive,” “negative,” 1, or 0. In some cases, two sets of cutoff values can be used to classify a sample into one of three possible output values. Examples of cutoff value sets include {1%, 99%}, {2%, 98%}, {5%, 95%}, {10%, 90%}, {15%, 85%}, {20%, 80%}, {25%, 75%}, {30%, 70%}, {35%, 65%}, {40%, 60%}, and {45%, 55%}. In some embodiments, a set of cutoff values can be used to classify a sample into one of n+1 possible output values, where n is any positive integer.
[0126] visualization In some embodiments, the systems, media, and methods disclosed herein may further include generating data summaries, results, or visualizations, or any combination thereof. In some embodiments, the visualization may include graphs, plots, and overlay points representing transcriptions or cell annotations on a tissue image. In some embodiments, the visualization may query its Seurat / tileDB object to retrieve relevant data from the object for display. In some further embodiments, the relevant data may include transcription locations and / or cell annotations. In some embodiments, the visualization may be displayed as transcription locations as points, cell annotations as points (e.g., cell types), box plots, violin plots, dot plots, or any combination thereof. In some embodiments, a visualization relating to cell segmentation statistics may be displayed as points of raw values or projected coordinates in an embedding space to evaluate a segmentation result of interest with respect to any alternative segmentation results of the same sample. In some embodiments, the visualization may form an overlay between segmentation boundaries, transcripts, and downstream data analysis outputs. In some cases, the downstream data analysis outputs include cell types. In some cases, the downstream data analysis output includes one or more of the following: cellular state, phenotype of the tissue microenvironment, differential expression of cell types based on spatial context, quantification of intracellular expression, and spatially degraded biomarker identification, or any combination thereof. In some embodiments, visualization may be obtained from different segmentation configurations or methods, or from segmentation results of a reference sample.
[0127] In some embodiments, visualization may be performed using a vision transformer. In some embodiments, the vision transformer may process an input image into a series of patches. In some cases, the vision transformer may serialize each patch into a vector. In some embodiments, the vision transformer may use matrix multiplication to match each vector to a smaller dimension. In some embodiments, a transformer encoder may process the vision transformer data. In some embodiments, the transformer may measure the relationship between one or more input tokens via the transformer encoder. In some embodiments, the visualization may be trained with a masked encoder, a Swin transformer, a vector-quantized variational autoencoder, or any combination thereof.
[0128] [Examples] The following illustrative examples represent, and are not intended to limit in any way, embodiments of the software applications, systems, and methods described herein.
[0129] [Example 1 - Chromatic aberration correction] Transverse chromatic aberration correction was performed on each multichannel image using a pre-calibrated transform. As shown in Figure 10, the raw image on the left was corrected for chromatic aberration and improved upwards. With the correction, the reference point where all channels are aligned is visibly observed in the corrected image on the right. The parameters for chromatic aberration correction are listed in the correction transform table.
[0130] [Example 2 - Spot Density Heatmap] As shown in Figures 13A and 13B, valid segmentation results were achieved, revealing cellular structures not shown in the original segmentation using morphological markers alone. The three columns in Figure 13A show morphological images of brain tissue with imaging channels including histone staining, DAPI staining, and Soma rRNA staining. The three columns in Figure 13B show a readout density heatmap, segmentation based on all three morphological staining images, and segmentation based on the density heatmap and histone staining. The readout density heatmap in column 1 of Figure 13B displays cell body contours with finer detail, with a higher signal-to-noise ratio compared to the rRNA Soma-stained morphological image in the third column of Figure 13A. As a result, the results of cell segmentation based on all three morphological staining images combined, and the readout density heatmap and histone staining combined, are shown in the second and third columns of Figure 13B, respectively.
[0131] As shown in Figure 24, image preprocessing techniques such as image sharpening and enhancement or machine learning denoising models were applied to the raw images to improve image quality. In one embodiment, Figure 24 shows a raw image of a readout density heatmap compared to the image reconstruction output after processing with the CellPose 3.0 denoising model.
[0132] [Example 3 - Comparison of 2-channel and 5-channel model outputs] Figure 17A shows the initial network output using the Cellpose algorithm. The gradient flow output was generated using a 2-channel model. Ground truth annotations were generated from segmentation results produced by an existing custom cell segmentation algorithm (a process called bootstrapping), based on the 2-channel model as the initial ground truth. The annotations were then reviewed and modified to ensure the quality of the training data. The new 5-channel network model was initialized with the parameters of the existing 2-channel model. This reduced the training burden and resulted in good initial segmentation performance. The network was trained and fine-tuned to optimize the parameters of all channels based on the annotated ground truth. The parameter set was finalized after satisfying a specified learning rate and minimizing a specific cost function. Figure 17B shows the output gradient flow results from the trained 5-channel model initialized from transfer learning. The trained and validated neural networks can be exported to Open Neural Network Exchange Format (ONNX) models for interoperability between different frameworks and deployments.
[0133] [Example 4: Image acquisition: HDR settings] Figure 25 shows an unrestricted example of an image acquired using a High Definition Range (HDR) setting. In the unrestricted example of Figure 25, the nuclear-stained image was acquired in two non-HDR modes, one with a higher exposure time setting and the other with a lower exposure time. In the unrestricted example shown in Figure 25, DAPI nuclear staining was used, excited with 385 nm UV and detected in the blue channel with emission at 512 nm. In the unrestricted example, the low-exposure scan was acquired at 3 ms with 3% power (A), followed by an acquisition at 24 ms with 4% power (B). As shown in Figure 25, some cells in B are saturated. As shown in Figure 25, (C) shows an image acquired in HDR using both exposure times combined to provide a larger dynamic range and simultaneously reduced intensity saturation. As shown in Figure 25, (A) is the image acquired at 3 ms (low). (B) is the image acquired at 24 ms (high). And (C) is an image acquired using HDR.
[0134] For example, the morphology channel describing the cell membrane used the acquisition settings shown in Table 2 below. Table 2 lists the channels / dyes used, as well as the Excite / Emission wavelengths and exposure times. For example, using 8x HDR, two images are acquired per channel, one with the specified exposure time and the other with 1 / 8 the exposure time of the first image.
[0135] [Table 2]
[0136] [Example 5: Image preprocessing: Nearest neighbor deconvolution] Figure 26 shows a non-limiting example of the output produced by the nearest neighbor (NN) deconvolution method. In Figure 26, for example, a Gaussian filter sigma value of 9.7 pixels was used based on the aforementioned estimation, and the image was sampled at 800 nm in each z acquisition. For example, a multiplier A = 0.75 was used to ensure that dark cells were still retained while removing most of the blurred signal. As shown in Figure 26, areas are circled to mark regions where cells previously blurred by out-of-focus signals (left panel) are recovered by nearest neighbor deconvolution, revealing clear boundaries between cells that were previously difficult to see (right panel).
[0137] [Example 6 - Cross-Issues (IoU) Analysis on Unions] For example, as shown in Figure 27, improved segmentation was achieved by integrating multiple segmentation steps from various modalities. For instance, in some embodiments, combining nuclear-only segmentation with cytoplasmic and nuclear segmentation can yield enhanced results. For example, certain cells may rely solely on nuclear staining and lack appropriate membrane staining, and vice versa. In some embodiments, this multimodality segmentation approach can be used to track brain cells. In some embodiments, a separate segmentation step may be required to identify elongated projections, which are then integrated with nuclear and cell body data. In some embodiments, the criteria for merging or segmenting are determined by calculations involving crossovers to joins and area measurements, typically with a minimum set of overlaps. Figure 27 shows a non-limiting example of merging cell projections.
[0138] Figure 28 shows a non-limiting example of IoU merging results between nuclear segmentation output and membrane-plus nuclear segmentation. The enclosed area in Figure 28 marks cases of cell merging. One example of merging in the upper right position of Figure 28 shows that the cell mask was detected only in the nuclear segmentation step (A) and not in the cytoplasmic / membrane segmentation step (B). The second example of merging in the lower left of Figure 28 shows that when the two masks (A) and (B) overlap, the masks are merged and retain the shape from both (A) and (B).
[0139] Similarly, as shown in the figure, this cross-analysis can be extended to correlate cells in different z-slices to perform three-dimensional cell segmentation. Figure 29A shows an unrestricted example of IoU merging results for correlating cellIds across multiple 2D image slices. Figure 29B shows an unrestricted example of IoU merging results in a 3D volume. Figure 29C shows an unrestricted example of a single cell marked in all z-slices to show that this cell is segmented across the entire Z-stack. This data was generated using the NanoString whole transcriptome (Wtx) pancreas public dataset.
[0140] [Example 7: Segmentation results across different organizational types] The multimodal segmentation system has been tested with multiple tissue samples, including those typically used in immuno-oncology studies, such as human breast, colon, lung, kidney, liver, and tonsils, in both normal and diseased states. As shown in Figure 30, it has been successfully segmented in human and mouse brain tissue, as well as other tissues such as ovaries, osteosarcoma, skin, and muscle.
[0141] [Example 8: Segmentation Mask Measurement Statistics Table] Using a segmentation mask, various cell measurements were generated, including intensity and shape measurements. Tables 3 and 4 below show examples of intensity and shape characteristics for a subset of the cell population, e.g., 20 out of 5309 cells. In some embodiments, hundreds of images can be obtained.
[0142] Strength and morphological properties of cell subsets
[0143] [Table 3]
[0144] Area and shape characteristics of cell subsets
[0145] [Table 4]
[0146] In some embodiments, in addition to position, intensity, and shape characteristics, this cell segmentation measurement may also include other characteristics such as cell fluorescence texture / entropy and neighbor / intercellular analysis. For example, to mark cells segmented across multiple FOVs, the calculation was performed by finding all cell IDs that intersect the FOV boundary. Furthermore, for these segmented cells, the ratio between these cells to the local mean was measured. In some embodiments, if only a portion of the cell segment is present in the current FOV, it may be further filtered during the analysis. Table 5 shows examples of segmented cell measurements. In Table 5, values close to 1 indicate that their size is similar to the mean and only a small portion of the segmented cells are located in adjacent FOVs, so those cells at the boundary may be retained.
[0147] Cell division ratio
[0148] [Table 5]
Claims
1. A computer implementation system that provides a multimodal segmentation application comprising a computing device including at least one processor and instructions executable by at least one processor, wherein the system includes the following: (a) A software module configured to search for 3D scanned images of a biological sample in high dynamic range (HDR) mode, wherein the biological sample is labeled with one of a plurality of morphological markers; (b) Optionally, a software module configured to reduce a 3D scanned image to a 2D image, thereby obtaining an optimal focal region in the z-slice; or alternatively, a software module configured to deconvolve each of the 2D images and use a stack of 3D scanned images of the 2D images to obtain an optimal focal region; (c) A software module configured to perform image preprocessing for correcting and enhancing images before segmentation, wherein the software module is further configured to perform intracellular segmentation on a biological sample based on multiple morphological markers; (d) Software modules configured to combine different imaging modalities, wherein the imaging modalities describe nuclear modalities, cytoplasmic modalities, and membrane modalities; (e) Software modules configured to perform cell segmentation in 2D, 3D, and time-lapse images; and A computer execution system comprising: (f) a software module configured to generate one or more measurements of cell-specific characteristics and images (c)–(e), including fluorescence intensity and shape information as part of the processing results.
2. The system according to claim 1, further comprising a software module configured to retrieve a readout density map from a transcriptome assay of a biological sample.
3. The system according to claim 2, wherein the imaging modality includes a readout density map or a fluorescence image, or both.
4. The system according to claim 1, wherein the morphological marker comprises a fluorescent dye, nuclear stain, fluorescently labeled antibody, immunohistochemistry (IHC) stain, photocleavable morphological marker, genetically encoded tag, magnetic resonance imaging (MRT) contrast agent, or nucleic acid probe, or any combination thereof.
5. The system according to claim 1, wherein the image is a microscope-derived image, including an image from an optical microscope, an electron microscope, or a scanning probe microscope.
6. The system according to claim 1, wherein the aforementioned image is applied to bleach correction and deconvolution.
7. The system according to claim 1, wherein the transcriptome assay comprises a gene expression assay using a fluorescently labeled probe, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), cap analysis of gene expression, or single-cell RNA sequencing (scRNA-seq), or any combination thereof.
8. The system according to claim 1, wherein the aforementioned image is applied to chromatic aberration correction.
9. The system according to claim 1, wherein the software module is configured to search for at least one image, at least three images, at least five images, at least ten images, at least fifteen images, at least twenty images, at least thirty images, at least thirty-five images, at least forty images, at least forty-five images, at least fifty images, at least fifty-five images, at least sixty images, at least sixty images, at least seventy images, at least eighty images, at least ninety images, at least one hundred images, at least one hundred twelve images, at least one fifteen images, and at least two hundred or more images of the biological sample.
10. The system according to claim 1, wherein one or more image results of the subcellular segmentation are combined before (f).
11. The system according to claim 1, wherein the biological sample is obtained as at least part of one or more of one of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, tumor tissue sampling, malignant tissue sampling, diseased tissue sampling, or transplanted tissue sampling, or any combination thereof.
12. The system according to claim 1, wherein the biological sample comprises cells or tissue.
13. The system according to claim 12, wherein the cells include primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, cancer cells, or tumor cells.
14. The system according to claim 1, wherein intracellular segmentation includes nuclear segmentation, cytoplasmic segmentation, or extracellular segmentation.
15. The system according to claim 1, wherein intracellular segmentation includes training a machine learning algorithm and / or applying a machine learning algorithm.
16. A non-temporary computer-readable storage medium coded with instructions executable by one or more processors, in order to provide a multimodal segmentation application comprising the following: (a) A software module configured to search for 3D (three-dimensional) scanned images of a biological sample in high dynamic range (HDR) mode, wherein the biological sample is labeled with one of a plurality of morphological markers; (b) A software module configured to reduce the three-dimensional scan image to a two-dimensional image, wherein the optimal focal region in the z slice can be obtained; or alternatively, a software module configured to deconvolve the two-dimensional image in order to obtain the optimal focal region while maintaining the z position of the two-dimensional image in the three-dimensional scan image stack of the two-dimensional image; and (c) A non-temporary computer-readable storage medium comprising a software module configured to perform intracellular segmentation on the biological sample based on the plurality of morphological markers.
17. The non-temporary computer-readable storage medium according to claim 16, further comprising a software module configured to retrieve a read-out density map from a transcriptome assay of the biological sample.
18. The non-temporary computer-readable storage medium according to claim 16, wherein the morphological marker comprises a fluorescent dye, nuclear stain, fluorescently labeled antibody, immunohistochemistry (IHC) stain, photocleavable morphological marker, genetically encoded tag, magnetic resonance imaging (MRI) contrast agent, or nucleic acid probe, or any combination thereof.
19. The non-temporary computer-readable storage medium according to claim 16, wherein the image is a microscope-derived image, including an image from an optical microscope, an electron microscope, or a scanning probe microscope.
20. The aforementioned image is a non-temporary computer-readable storage medium according to claim 16, which is applied to bleach correction.
21. The non-temporary computer-readable storage medium according to claim 16, wherein the transcriptome assay comprises a gene expression assay using a fluorescently labeled probe, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), cap analysis of gene expression, or single-cell RNA sequencing (scRNA-seq), or any combination thereof.
22. The non-temporary computer-readable storage medium according to claim 17, wherein the software module is configured to retrieve at least one image, at least three images, at least five images, at least ten images, at least fifteen images, at least 20 images, at least 30 images, at least 35 images, at least 40 images, at least 45 images, at least 50 images, at least 55 images, at least 60 images, at least 65 images, at least 70 images, at least 80 images, at least 90 images, at least 100 images, at least 120 images, at least 150 images, and at least 200 or more images of the biological sample.
23. The non-temporary computer-readable storage medium according to claim 16, wherein the biological sample is obtained by at least part of one or more of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue sampling, tumor tissue sampling, malignant tissue sampling, diseased tissue sampling, or transplanted tissue sampling, or any combination thereof.
24. A non-temporary computer-readable storage medium according to claim 16, wherein the biological sample comprises cells or tissue.
25. The non-temporary computer-readable storage medium according to claim 24, wherein the cells include primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, cancer cells, or tumor cells.
26. The non-temporary computer-readable storage medium according to claim 16, wherein intracellular segmentation includes nuclear segmentation, cytoplasmic segmentation, or extracellular segmentation.
27. A non-temporary computer-readable storage medium according to claim 16, wherein one or more image results of intracellular segmentation are combined.
28. The non-temporary computer-readable storage medium according to claim 16, wherein intracellular segmentation includes training a machine learning algorithm and / or applying a machine learning algorithm.
29. A method performed by a computer, (a) A step of searching for a three-dimensional scan image of a biological sample in high dynamic range (HDR) mode using a computer, wherein the biological sample is labeled with one of a plurality of morphological markers; (b) A step of reducing the three-dimensional scan image to a two-dimensional image, wherein the optimal focal region within the z slice is obtained; or a step of performing deconvolution while maintaining the z position of the two-dimensional image within a stack of the corresponding three-dimensional scan images, thereby generating a two-dimensional image for optimal contrast; and (c) A method comprising the step of performing segmentation on the biological sample based on the plurality of morphological markers.
30. The method according to claim 29, further comprising the step of retrieving a readout density map from a transcriptome assay of a biological sample.
31. The method according to claim 29, wherein the morphological marker comprises a fluorescent dye, nuclear stain, fluorescently labeled antibody, immunohistochemistry (IHC) stain, photocleavable morphological marker, genetically encoded tag, magnetic resonance imaging (MRI) contrast agent, or nucleic acid probe, or any combination thereof.
32. The method according to claim 29, wherein the image is a microscope-derived image, including an image from an optical microscope, an electron microscope, or a scanning probe microscope.
33. The method according to claim 29, wherein the aforementioned image is subjected to bleaching correction.
34. The method according to claim 29, wherein the transcriptome assay comprises a gene expression assay using a fluorescently labeled probe, RNA sequencing (RNA-seq), microarray analysis, reverse transcription polymerase chain reaction (RT-PCR), gene expression cap analysis, or single-cell RNA sequencing (scRNA-seq), or any combination thereof.
35. The method according to claim 29, wherein the software module is configured to search for at least one image, at least three images, at least five images, at least ten images, at least fifteen images, at least 20 images, at least 30 images, at least 35 images, at least 40 images, at least 45 images, at least 50 images, at least 55 images, at least 29 images, at least 65 images, at least 70 images, at least 80 images, at least 90 images, at least 100 images, at least 120 images, at least 150 images, and at least 200 or more images of the biological sample.
36. The method according to claim 29, wherein one or more image results of intracellular segmentation are combined.
37. The method according to claim 29, wherein the biological sample is obtained by at least part of one or more of the following: biopsy, surgical excision, xenograft, animal model, fine-needle aspiration, peripheral blood collection, bone marrow biopsy, healthy tissue collection, tumor tissue collection, malignant tissue collection, diseased tissue collection, or transplant tissue collection, or any combination thereof.
38. The method according to claim 29, wherein the biological sample comprises cells or tissue.
39. The method according to claim 38, wherein the cells include primary cells, stem cells, immune cells, cancer cells, sarcoma cells, lymphoma cells, melanoma cells, cancer cells, or tumor cells.
40. The method according to claim 29, wherein intracellular segmentation includes nuclear segmentation, cytoplasmic segmentation, or extracellular segmentation.
41. The method according to claim 29, wherein intracellular segmentation includes training and / or applying a machine learning algorithm.