In situ cell characterization method

Through contact with the co-regulatory genes in biological samples, the combination of emission signals is detected, and the problem of low RNA diffusion and capture efficiency in the prior art is solved, and efficient and robust cell spatial characterization of biological tissue samples is achieved.

CN120265787APending Publication Date: 2025-07-04AGENCY FOR SCI TECH & RES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380080491.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-29
Filing Date
2023-11-29
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, when analyzing biological tissue samples with high-dimensional spatial resolution, there are problems such as low RNA diffusion and capture efficiency, high nonspecific background noise, high resolution requirements and poor scalability, making it difficult to achieve spatial characterization of cells.

Method used

Multiple probes are used to contact ribonucleic acid (RNA) transcripts of multiple predefined genes in biological samples, combined with detectable markers and specific domains, detect emission signal combinations, and perform cell characterization by co-regulating genes.

Benefits of technology

It improves the intensity and scalability of signal detection, achieves robust and efficient spatial characterization of cells, reduces the requirements for experimental equipment and time, and is suitable for healthy and diseased tissue samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120265787A_ABST
    Figure CN120265787A_ABST
Patent Text Reader

Abstract

The present technology relates to a method and a kit for in situ characterization of cells in a biological sample. The method includes contacting the biological sample with a plurality of probes that bind ribonucleic acid (RNA) transcripts of a plurality of predetermined genes. In one embodiment, the plurality of predetermined genes are expressed in cells associated with cancer, particularly colorectal cancer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of Singapore Patent Application No. 10202260245V, filed on November 29, 2022, the entire content of which is hereby incorporated by reference, especially with respect to the drawings, legends and claims herein. Technical field

[0003] The present invention generally relates to the field of molecular and cell biology. In particular, the present invention relates to methods for cell characterization. Background art

[0004] High - dimensional spatially resolved analysis of intact biological tissue samples holds promise for transforming biomedical research and diagnosis. Recent advances in single - cell RNA sequencing (scRNA - seq) have made it possible to unbiasedly define cell types that reflect ontogeny, function, or anatomical location. However, high - throughput mapping of these cells within intact biological systems remains a technical challenge. Existing methods, such as spatial indexing combined with next - generation sequencing, enable spatial mapping of sequencing reads and in - situ reconstruction of cell types. However, sequencing - based spatial transcriptomics methods are limited by RNA diffusion and capture efficiency. Alternatively, cell types can be characterized via imaging - based spatial transcriptomics methods by targeting RNA using multiplexed single - molecule fluorescence in situ hybridization (FISH) or in - situ sequencing. These methods are highly quantitative and scalable to the whole transcriptome (about 10,000 genes), but drawbacks include higher non - specific background noise, limitations due to molecular crowding, and the need for high - resolution microscopes. Imaging - based spatial transcriptomics methods also become increasingly laborious for a large number of targets. Another approach to cell spatial mapping is multiplexed immunostaining or spatial proteomics. Although the increased protein copy number compared to RNA may improve detection robustness, antibody panels are more expensive, less flexible, and less scalable.

[0005] Accordingly, there is a need for a technology that can implement a simple, efficient, and scalable method for spatially characterizing cells in the context of normal tissue physiology or disease microenvironments. Additionally, other desired features and characteristics will become apparent from the following detailed description and the appended claims, taken in conjunction with the accompanying drawings referenced herein. Summary of the invention

[0006] In one aspect, the present disclosure relates to a method for in-situ characterization of cells in a biological sample, comprising: (a) contacting the biological sample with a plurality of probes that bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes, wherein each probe comprises: i) a detectable label; and ii) a domain that specifically binds to the ribonucleic acid transcript of one of the predetermined genes; wherein a signal is emitted when the probe binds to the ribonucleic acid transcript; (b) detecting a combination of emission signals from the plurality of probes or a plurality of emission signals; and (c) characterizing the cells based on the combination of emission signals or the plurality of emission signals.

[0007] In another aspect, the present disclosure relates to a method for determining the prognosis of a subject having cancer, comprising: (a) obtaining a sample from the subject; (b) characterizing one or more cancer cells in the sample using the method according to any one of claims 1 to 13 to determine the stage of the cancer; and (c) determining the prognosis based on the stage of the cancer.

[0008] In another aspect, the present disclosure relates to a kit for in-situ characterization of cells in a biological sample, comprising: a plurality of probes that bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes; wherein each probe comprises: i) a detectable label, and ii) a domain that specifically binds to the ribonucleic acid transcript of one of the predetermined genes; and instructions for use.

[0009] In another aspect, the present disclosure relates to a kit for in-situ characterization of colorectal cancer in a biological sample, comprising: a plurality of probes that bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes, wherein the plurality of predetermined genes are selected from the genes listed in Tables 6(6a)-(6d); wherein each probe comprises: i) a detectable label, and ii) a domain that specifically binds to the ribonucleic acid transcripts of the plurality of predetermined genes; and instructions for use. Description of the Drawings

[0010] Figure 1 A schematic overview of the in-situ hybridization (ISH) method for characterizing cells described herein is provided. The methods described herein can be used to precisely map cell types without disrupting tissue architecture. As described herein, the method is a sensitive, robust, and scalable in-situ hybridization (ISH)-based spatial transcriptomics method that uses multiple co-regulated genes to profile single cells. As used herein, co-regulated genes refer to genes that exhibit co-variation in terms of gene expression levels, i.e., co-varying genes. As Figure 1 shown in A, co-regulated genes are spatially co-localized in the same cells within a tissue, which allows for the design of hybridization probes to target a large number of genes to reliably detect cell populations of interest. Figure 1A provides a cell-by-gene count matrix derived from single-cell RNA sequencing (scRNA-seq). This matrix is used to cluster cell types, which are characterized by their unique gene expression profiles (e.g., grouping genes A-D into one cell cluster and genes E-I into a different cluster). Figure 1 B provides an illustrative description of identifying related gene groups based on reference scRNA-seq data. Genes that exhibit co-variation with each other in terms of expression levels are spatially co-localized in the same cells within a tissue. Based on the identified related gene groups, thousands of oligonucleotide probes targeting their transcripts are designed, which results in tens of thousands of detectable tags per cell (taking into account the number of genes, the number of transcript copies per cell, and the number of probes per transcript). By designing labeled oligonucleotide probes targeting a large number of co-regulated transcripts, the in situ hybridization cell characterization method described herein improves the intensity of signal detection. Figure 1 C shows the workflow for in situ hybridization-based expression profiling of cells in animal tissues (such as kidney and brain) by combining an oligonucleotide pool synthesized by array and continuous flow control technology. This method can be applied to healthy tissues or diseased tissues (e.g., normal tissues or cancer tissues). By combining repeated hybridization and washing cycles, the in situ hybridization method described herein for characterizing cells enables robust and scalable mapping of cell types in tissue samples. For example, a commonly used detectable signal is a fluorescent signal. A useful application of the in situ hybridization method can be fluorescence in situ hybridization for characterizing cell heterogeneity (referred to as "FISHnCHIPs" in some specific examples). Thus, as summarized herein, the present disclosure provides a robust in situ hybridization method for characterizing cells in a biological sample, which has amplified signal intensity and high scalability.

[0011] Figure 2 A comparison of the exemplary application of this method with conventional single-molecule RNA FISH (smFISH) in an exemplary mouse kidney tissue is provided. In the exemplary method ("FISHnCHIPs") shown in this figure, fluorescently labeled probes are designed for five selected cell types using the mouse kidney scRNA-seq dataset: renal macrophages, glomerular endothelial cells, loop of Henle (LOH) cells, collecting duct (CD) cells, and glomerular podocytes. Figure 2A provides a gene expression heatmap generated based on scRNA-seq reference data, which highlights five corresponding cell clusters representing each cell type. A suitable cut-off value is applied to the correlation coefficients calculated for the genes to determine the genes to be targeted for each cell type using FISHnCHIP. The heatmap shows the relative expression levels of 84 genes associated with the most differentially expressed (DE) genes in the five selected cell types, with a maximum of 300 cells sampled per cluster. Figure 2 B shows untreated smFISH images of mouse kidney tissue sections in the five selected cell types in the left and middle images and FISHnCHIPs images in the right image, which simultaneously label multiple co-regulated genes (14 to 23 genes, as Figure 2 shown in B) to detect the target cell type. For each cell type, the smFISH and FISHnCHIPs images are adjusted to the same camera intensity range. Nuclei staining is shown with DAPI. The scale bar is 3 μm. From Figure 2 the comparison between the smFISH image and the FISHnCHIPs image in B, a high co-localization between the two most co-regulated genes in each of these cell types is observed, confirming that the relevant genes from scRNA-seq are indeed spatially co-localized in the same cells. Figure 2 C shows FISHnCHIPs images of five different cell types of mouse kidney tissue. Panel (i) shows the FISHnCHIPs image of endothelial cells of mouse kidney tissue. Panel (ii) shows the FISHnCHIPs image of collecting duct cells of mouse kidney tissue. Panel (iii) shows the FISHnCHIPs image of podocytes of mouse kidney tissue. Panel (iv) shows the FISHnCHIPs image of loop of Henle cells of mouse kidney tissue. Panel (v) shows the FISHnCHIPs image of macrophages of mouse kidney tissue. Panel (vi) shows the DAPI image of the nuclei in the same mouse kidney tissue. Figure 2 The scale bar for all images in D is 25 μm. As Figure 2 shown, when using a combination of multiple genes to label the selected cell type, it is easier to detect cells compared to only labeling a single most differentially expressed (DE) gene. Although these 5 cell types only account for approximately 12% of the total kidney cell population (estimated based on scRNA-seq), the method shown in this figure reveals complex spatial details of the kidney tissue structure (such as the arrangement of podocytes within the highly perforated Bowman's capsule where they wrap around the glomerular endothelial cells at the Bowman's capsule). Therefore, Figure 2Examples of cell - centric strategies for in situ hybridization (ISH) methods for characterizing cells as described herein are provided, which amplify detectable signals based on multiple co - regulated genes corresponding to user - predefined known cell types (e.g., renal macrophages, glomerular endothelial cells, loop of Henle (LOH) cells, collecting duct (CD) cells, and glomerular podocytes).

[0012] Figure 3 Examples of cell - centric FISHnCHIPs signal readings quantification for five cell types in mouse kidney are provided in relation to Figure 2 the above. Figure 3 A shows box plots (solid boxes) of the ratio of the average fluorescence intensity per cell of FISHnCHIPs to single - molecule FISH (smFISH), which indicates the actual increase in measured fluorescence intensity; and shows the count ratio of 14 to 23 genes to the top DE gene (hollow box) based on scRNA - seq results, which indicates the predicted value of the increase in fluorescence intensity. The number of cells used for FISHnCHIPs calculations is as follows: collecting duct cells: 146; podocytes: 461; loop of Henle cells: 727; endothelial cells: 400; and macrophages: 341. The number of cells used for scRNA - seq calculations is as follows: collecting duct cells: 1825; podocytes: 77; loop of Henle cells: 1496; endothelial cells: 701; and macrophages: 216. The box plots show the median (mid - line), the first and third quartiles (box boundaries), and 1.5 times the inter - quartile range (whiskers). The horizontal line indicates the position where the fluorescence signal gain is 1. Compared with the conventional single - molecule FISH (smFISH) method, the FISHnCHIPs fluorescence intensity per cell increased by approximately 6 to 39 - fold (median of at least 146 cells) in all 5 cell types and was consistent with or exceeded the predicted signal increase. However, according to Figure 2 the scRNA - seq data shown in A, some selected genes of FISHnCHIPs may be expressed in off - target cell types. For example, Slc5a3 (Pearson correlation (r) with Slc12a1, a marker of loop of Henle (LOH), is 0.33) is also expressed in collecting duct (CD) cells. To estimate the crosstalk in FISHnCHIPs results, the Manders overlap coefficients for all five cell - type channels were calculated, which ranged from 0.001 to 0.09, indicating minimal crosstalk between these cell types. Figure 3 B provides a heat map showing the normalized average scRNA - seq counts of selected genes of FISHnCHIPs for all 5 cell types, which is a predictive signal crosstalk level. Figure 3Panel C shows the Manders overlap coefficients for all five cell type channels measured by FISHnCHIPs, indicating the actual measured signal crosstalk in the FISHnCHIPs imaging results. In Figure 3 B and Figure 3 C, the number of cells analyzed was the same. Thus, based on the quantitative comparison between the conventional smFISH method and the FISHnCHIPs method illustrated herein, this method shows a signal intensity increase of up to 39-fold. Further comparison with the predictive crosstalk based on scRNA-seq data shows that the FISHnCHIPs method illustrated herein shows minimal crosstalk between cell types and thus shows high specificity.

[0013] Figure 4 Provides computational predictions of signal gain and specificity for the cell-centered FISHnCHIPs method as Figure 2 shown. As Figure 4 shown in Panel A, the heatmap visually presents the scRNA-seq gene expression of the FISHnCHIPs gene set, which targets all previously annotated mouse kidney cell types, with a maximum of 300 cells sampled per cluster. Figure 4 Panel B provides the predicted signal gain (SG) and signal specificity ratio (SSR) based on scRNA-seq reference data, both expressed as a function of the number of genes used (ranked according to their Pearson correlation with the most differentially expressed gene). The signal gain (SG) is defined as the ratio of the sum of the counts of the FISHnCHIPs genes to the sum of the counts of the most highly DE gene, and the signal specificity ratio (SSR) is defined as the ratio of the sum of the counts of the FISHnCHIPs genes in the target cell type to the sum of the counts of the FISHnCHIPs genes in the most likely off-target cell type. When the SSR approaches unity, the fluorescence intensity of the cell type of interest should equal that of the off-target cell type, making them indistinguishable. A high signal gain (SG) indicates the expected signal amplification of FISHnCHIPs. As Figure 4 shown in Panel B, for 9 out of the 16 previously annotated cell types, the SSR exceeds 4, indicating high specificity for these cell types when a cell-centered strategy is adopted for FISHnCHIPs panel design. Figure 4C provides an overview of predicted signal crosstalk in the heatmap, which shows the normalized average scRNA-seq counts of FISHnCHIPs gene combinations for all renal cell types. Although the signal-to-noise ratio has been improved, the cell type specificity of these cell-centered FISHnCHIPs can be further enhanced. Given the predicted signal gain and specificity of the method (cell-centered strategy), it is shown that the method improves sensitivity with a minimal trade-off in specificity.

[0014] Figure 5 Alternative examples of the in situ hybridization (ISH) cell characterization method described herein are provided. As an alternative to the cell-centered strategy that requires user input of known cell type information, the gene-centered strategy utilizes related genes from gene expression program clusters (i.e., co-regulated genes in biological pathways). Figure 5 An exemplary gene-centered FISHnCHIPs profiling of 18 gene modules in the mouse cortex is shown. To reduce crosstalk, genes are clustered based on pathways and gene expression programs, which are known to exhibit coordinated expression variability in the mammalian genome at least without prior clustering of cell types. The gene-gene correlation matrix (instead of the gene-cell matrix) of the mouse visual cortex dataset is clustered. A total of 255 candidate genes are selected, which are highly correlated (Pearson correlation (r) > 0.7) with at least three genes. From the candidate pool, 18 gene modules that are significantly enriched for gene ontology (GO) are identified. Figure 5 A provides a gene-gene correlation heatmap (heatmap of pairwise Pearson correlation coefficients), which is clustered into 18 gene module (gene expression program) clusters according to the identification. Each module (on average containing 14 genes) is imaged sequentially in fresh frozen mouse brain tissue sections under an automated flow cytometry-coupled fluorescence microscopy system. For gene modules 1, 2, 3, and 18, exemplary FISHnCHIPs images of mouse brain tissue sections are stained. The scale bar for all images is 50 μm. Single cells in the images are segmented using DAPI staining, and a cell mask is applied after quality control to define 6180 cells. The average fluorescence intensity of each cell for each imaged module is quantified. Figure 5 B provides a heatmap showing the average fluorescence intensity of each cell. The cell-by-module intensity matrix is clustered using the Louvain algorithm to obtain 8 cell clusters. Then, the resulting cell clusters are targeted separately in the sample, and detectable markers are measured. Figure 5C shows spatial maps of the detected cells in images (i) to (viii), which are classified into cell types: glutamatergic neurons (i), GABAergic neurons (ii), astrocytes (iii), oligodendrocytes (iv), endothelial cells (v), microglia (vi), perivascular cells (vii), and leptomeningeal cells (viii). Figure 5 The scale bar in C is 500 μm. As Figure 5 shown in C, the 8 cell types exhibit different spatial organization patterns. To verify whether the identified cell types are consistent with existing methods, Figure 5 D shows in a scatter plot the frequencies of cell types detected by FISHnCHIPs versus the frequencies of cell types detected by the multiplex error-robust fluorescence in situ hybridization (MERFISH) method (Pearson correlation r = 0.97). The inset is a pie chart showing the proportion of each FISHnCHIPs cluster. FISHnCHIPs demonstrates a high degree of correlation and consistency with existing technology methods. Thus, Figure 5 an example of a gene-centered in situ hybridization (ISH) cell characterization method is provided, which effectively depicts tissue samples into 8 different cell types based on 18 gene expression programs, showing results consistent with existing methods.

[0015] Figure 6 Further details of the combinatorial design of the 18 gene expression programs are provided, as well as the results of clustering 8 cell types using gene-centered FISHnCHIPs in the mouse cortex as Figure 5 shown. Figure 6 A provides a uniform manifold approximation and projection (UMAP) rendering of the predicted clusters of single-cell RNA-seq (scRNA-seq) mock module-cell (meta-gene) expression, indicated by markers from the scRNA-seq reference dataset. As shown in the UMAP plot, approximately 8 cell types are clearly separated using the selected features. Figure 6 B predicts the conserved signal gain (cumulative conserved signal gain), which is defined as the ratio of the combined signal to the highest gene signal, as a function of the number of genes. As Figure 6 shown in C, the FISHnCHIPs signal is expected to be 1.2 to 22.3 times brighter than spectral analysis using individual marker genes. Figure 6 C provides a module-cell expression heatmap, which is grouped into 8 distinguishable cell types. Using the gene-centered in situ hybridization (ISH) cell characterization method, amplified signals can be obtained for each gene expression program.

[0016] Figure 7A schematic overview of an exemplary software pipeline for aligning, segmenting, and clustering cell types based on the acquired FISHnCHIPs imaging data is provided. In summary, the stepwise data processing includes the following: 1) The input to the image processing workflow includes DAPI, FISHnCHIPs, and background (after 55% formamide wash) images; 2) Image segmentation is preprocessed based on the DAPI image to generate cell masks; 3) The FISHnCHIPs images are registered and background subtracted; 4) A cell intensity matrix with a list of cytoplasmic centroids is generated using the cell masks; 5) The cell intensity matrix is clustered; 6) The output of the pipeline can be visually presented in a heatmap, UMAP, or spatial map. Further analysis can also be performed on the output generated from this pipeline, such as classification of spatial patterns and analysis of cell-cell interactions. The imaging results obtained from the in situ hybridization method described herein provide insights into cell types, cell-cell interactions, and cell spatial distribution within tissues. Further processing of the imaging data is available and can be designed accordingly based on the purpose of the experiment.

[0017] Figure 8 Scatter plots of cell type abundances between three different replicate datasets are provided, demonstrating the reliable reproducibility of mouse brain FISHnCHIPs cell type profiling data in technical replicates.

[0018] Figure 9 Another example of the in situ hybridization method described herein is provided, which is based on gene-centered FISHnCHIPs profiling of 20 gene expression programs in the mouse cortex. Related genes are identified based on a dimensionality reduction-based algorithm (consensus non-negative matrix factorization (NMF)) that infers co-gene expression in neurons, rather than Figure 5 the gene-gene correlation matrix shown. Gene-gene correlation analysis is performed on 20 previously annotated gene expression programs, resulting in FISHnCHIPs combinations that on average contain 16 genes per program. Twenty neuronal gene expression programs (including 14 identity programs (ExcL2, ExcL3, ……, and Sub) and 6 activity programs (Erp, LrpD, ……, and Syn)) are detected by the FISHnCHIPs method described herein, and Figure 9 the resulting images are shown in A. Figure 9A provides exemplary FISHnCHIPs images of mouse brain tissue sections stained for the programs ExcL2, ExcL5p3, ExcL6p1, ExcL6p2, IntSst, and IntPv out of the 20 programs used, while imaging an average of 16 co-related genes. The scale bar for all images is 500 μm. Identity programs appear to be more spatially local, while active programs are more ubiquitously expressed. Cluster analysis was performed on 2794 segmented single cells using the identity programs. Figure 9 B shows a heatmap of the average fluorescence intensity per cell for each imaging program. As Figure 9 visually presented by Uniform Manifold Approximation and Projection (UMAP) in C, the cell-program intensity matrix was further clustered using the Louvain algorithm, resulting in 11 cell type clusters, each labeled with program annotations (L2 / 3, L3 / 4, L4 / 5 …… and Sub). Figure 9 D provides a spatial map of the cells detected within the tissue, classified by their cell type into: L2 / 3 excitatory neurons (panel i), L3 / 4 excitatory neurons (panel ii), L4 / 5 excitatory neurons (panel iii), L5p1 excitatory neurons (panel iv), L5 / 6 excitatory neurons (panel v), L6p1 excitatory neurons (panel vi), IntPv inhibitory neurons (panel vii), IntSst inhibitory neurons (panel viii), IntNpy / CckVip inhibitory neurons (panel ix), hippocampus (panel x), and subiculum (panel xi). The scale bar for all images is 400 μm. The distribution of excitatory and inhibitory neurons along the cortical depth was further quantified. Quantification of the neuronal cell distribution reproduced previous findings of the layered structural organization of cells in the cortex. As Figure 9 shown in E, excitatory neurons are spatially organized into 6 distinct layers. According to Figure 9 F, inhibitory neurons also exhibit layer-specific localization, with Npy and CckVip more concentrated in the upper layers, while Sst and Pv-expressing neurons are clustered in the deeper layers. This example demonstrates that the present method can distinguish neuronal subtypes that stratify the typical laminar structure of the visual cortex. It also shows that the method for identifying gene modules (gene expression programs) is not limited to Figure 5 the gene-gene correlation matrix shown, but is also applicable to other methods for determining related genes.

[0019] Figure 10 provides an evaluation of the gene-centered FISHnCHIPs combination in the mouse visual cortex using a scRNA-seq reference dataset. As Figure 9 shown in Figure 10As shown in A, for all programs, the predicted conservative signal gain (cumulative conservative signal gain) (defined as the ratio of the combined signal to the highest gene signal, as a function of the number of genes) increased by 1.2 to 7.6-fold. Figure 10 B is a scRNA-seq expression heatmap of 20 gene expression programs. The heatmap visually presents the predicted signals of the 20 gene expression programs (rows normalized to the maximum value, which is the sum of the expression levels of co-regulated genes in the program). The heatmap provides an overview of the expression levels of the programs in different cell types (columns). As Figure 10 shown in B, the identity programs are expressed in a cell type-specific manner (high specificity), while the active programs are more generally expressed. Figure 10 C provides a uniform manifold approximation and projection (UMAP) illustration of 20 gene expression programs (labeled by reference cell type annotation). UMAP shows that cells from the same cell type cluster close to each other. For example, excitatory neurons are close together, while inhibitory interneurons and inhibitory neurons on the left side of the UMAP are well separated in the cluster. Figure 10 D provides a simulated scRNA-seq feature map of 14 identity programs. Similar to Figure 10 the heatmap of B, Figure 10 D visually presents the program expression according to Figure 10 the cell types plotted in C. The evaluation of the exemplary gene-centered in situ hybridization method described herein shows amplified signal intensity (sensitivity) while providing cell type specificity.

[0020] Figure 11 shows the formation of a gene expression gradient along the cortical depth of the mouse visual cortex as shown by Figure 9 the gene-centered FISHnCHIPs combined imaging in. Figure 11 A provides a heatmap of the FISHnCHIPs expression cell-program intensity matrix, where the cells are sorted by their distance from the cortical outer edge. As Figure 9 defined in D, the cortical depth distance of each cell type is calculated based on two white arcs. According to the heatmap, some programs show a gradual change in intensity along the cortical depth. Figure 11 B provides a uniform manifold approximation and projection (UMAP) illustration of the FISHnCHIPs feature maps of 14 identity programs. These results show that the excitatory programs (except ExcL6p1) change continuously with the distance from the cortical outer edge. The expression distributions of some programs partially overlap along the cortical depth, indicating that the spatial gene expression gradient may underlie continuous neuronal subtypes. As shown herein, this in situ hybridization method can be used to reveal potential structural patterns in tissue organization.

[0021] Figure 12Shows imaging of the mouse brain at lower magnification using the in situ hybridization method described herein. Figure 12 A provides an overview of six different objective lenses used according to their respective magnification (M), numerical aperture (N.A.), and predicted condenser power (F(epi)) specifications under epi-illumination configuration. As Figure 12 shown in B, for six different objective lenses, the average fluorescence intensity of Alexa594, Cy5, and IR800CW for each cell was measured. Objective lenses with higher magnification were able to collect signals with higher intensity, which was true for Alexa594, Cy5, and IR800CW. At the same magnification level, a water lens could obtain an image with higher signal intensity compared to an air lens. For the six different objective lenses (pictures a to f), Figure 12 an exemplary untreated FISHnCHIPs image (one field of view, FOV) of the mouse cortex is shown in C. In all three-color channels, signals above the background level could be detected in cells labeled with FISHnCHIPs even at the lowest magnification of 10x, indicating that this method significantly improved signal intensity compared to conventional methods. Figure 12 D provides quantification of the number of cells detected in each field of view (FOV) (n = 5 FOVs, error bar columns represent standard deviation). Because the field of view is wider, the number of imaged cells was approximately 40 times larger when using a 10x compared to a 60x objective lens. The average number of cells detected by each lens was: 10x air lens: 3130; 10x water lens: 3088; 20x air lens: 1003; 20x water lens: 1041; 40x: 261; 60x: 73. Due to the enhanced signal, cells labeled using the method described herein could be well detected at lower magnification, enabling a larger field of view and spectral analysis of more cells within the same time. To capture a larger number of cells, the 10x water objective lens was later used for Figure 13 data acquisition in.

[0022] Figure 13 Shows an exemplary gene-centered FISHnCHIPs spectral analysis of 53 gene modules in the mouse brain at a large field of view (FOV) (10x objective lens) of whole tissue sections. Compared to a 60x objective lens, this allowed coverage of an area 36 times larger within the same assay time (21 hours). Similar to previous analyses, as Figure 13 shown in A, unsupervised clustering of 54,834 cells is shown in the cell-module intensity matrix ( Figure 13A, left), which reveals 18 major cell types. As shown in the matrix, co-regulated gene modules were observed to co-localize in the same cells, and biologically relevant modules were tightly clustered in the expression space. Uniform Manifold Approximation and Projection (UMAP) plots of all cells are provided ( Figure 13 A, right), with the separate clusters marked accordingly. Figure 13 B provides individual spatial maps of 18 different cell clusters in a large field of view (FOV) in Panels a to r: neurons 1, 2, 3, 4, 5, 6, 7, and 8, astrocytes, vascular-associated cells, endothelial cells, ependymal cells, immature oligodendrocytes, mature oligodendrocytes 1 and 2, microglia, pericytes, and an unknown cell type. Scale bar is 1000 μm. Spectral analysis of cell types using the gene-centered in situ hybridization method of the present invention at low magnification demonstrated that the method described herein enhanced signal sensitivity and provided proof-of-concept for spectral analysis of cells within a large field of view (FOV) in tissue, covering both neuronal and non-neuronal cell types.

[0023] Figure 14 Simulations of gene-centered FISHnCHIPs combinations were provided, which utilized an exemplary unsorted scRNA-seq dataset to evaluate clustering accuracy with respect to a reference annotation. Figure 14 A provides a scRNA-seq gene-gene correlation heatmap of 674 feature genes from the Figure 13 mouse cortex library imaged in. Pairwise Pearson correlation coefficients of the feature genes were calculated. Based on the correlation coefficients, the Leiden algorithm was used to cluster the correlation matrix. Hierarchical clustering was used to further sub-cluster the resulting gene clusters into 53 gene modules with a signal gain (SG) of approximately 1.9 to 20.2. Figure 14 B to Figure 14 E provide UMAP plots of cells in the scRNA-seq dataset predicted according to different feature sets: Figure 14 B shows the prediction based on 1000 highly variable genes. Figure 14 C shows the prediction based on 2000 highly variable genes. Figure 14 D shows the prediction based on 3000 highly variable genes. Figure 14 E shows the prediction based on the Figure 13 53 modules shown in. Figure 14 F shows the adjusted Rand index (ARI) for clustering cells at a resolution of 0.1, using the Figure 14 B to Figure 14E was used as a feature and compared with markers from scRNA-seq datasets as the ground truth. The ARI score for the combination of 53 modules was 0.814, indicating that it could largely reproduce known brain cell types. As a comparison, the ARI score for 1000 highly variable genes (simulating a conventional assay that profiles 1000 genes separately) was only slightly higher, at 0.846. Thus, the simulation shows that the in situ hybridization method described herein provides amplified signal readings while maintaining comparable profiling specificity compared to conventional assays.

[0024] Figure 15 Exemplary normalized images of FISHnCHIPs profiling of 53 modules at a 10x objective are provided, which cover an area 36 times larger within the same assay time (21 hours). For example, in Figure 15 A, gene modules 39, 41, and 53 were imaged using Alexa 594. Figure 15 B shows representative images of gene modules 20, 33, and 36 using Cy5. Figure 15 C shows gene modules 1, 5, and 6 using IRDye 800CW. These images were taken at a 10x objective. The scale bar for all images is 1000 μm. The insets are magnified in the white box area with a scale bar of 100 μm. Despite capturing a large field of view (FOV), these exemplary images show strong signals with good resolution obtained using the method described herein, indicating an enhancement in imaging quality and efficiency of the present method.

[0025] Figure 16 The cell types identified by FISHnCHIPs were compared with the results of single-cell RNA sequencing (scRNA-seq). Figure 16 A provides a uniform manifold approximation and projection (UMAP) plot of frontal cortex cells, which is derived from the Harmony algorithm integration of scRNA-seq reference and FISHnCHIPs composite data. Figure 16 B provides a uniform manifold approximation and projection (UMAP) plot of scRNA-seq cells with cell type markers provided by Saunders et al. Figure 16 C shows the UMAP and markers of cells processed using Figure 13 the same FISHnCHIP method described above. The UMAP plot shows the correspondence between the cell types identified by the in situ hybridization method described herein and the scRNA-seq data.

[0026] Figure 17 Provided is Figure 13Sub-cluster analysis of FISHnCHIPs data for the 53 modules described above. Figure 17 A provides a FISHnCHIPs expression heatmap of the identified vascular-related cell subtypes. Figure 17 B provides a FISHnCHIPs spatial map of the identified vascular-related cell subtypes. Figure 17 C provides the Uniform Manifold Approximation and Projection (UMAP) of the identified vascular-related cell subtypes. FISHnCHIPs experimental data is used to identify various cell subtypes. For example, different localizations of vascular-related cell subtypes (such as CNN1+ smooth muscle cells, DCN+ fibroblasts, MRC1+ (also known as CD206) boundary-related macrophages almost exclusively located on the cortical surface, and GKN3+ arterial endothelial cells forming larger penetrating vascular structures) are observed. Therefore, the in situ hybridization method described herein not only provides a map of cell types, but also reveals fine subtype cells with different spatial distribution patterns.

[0027] Figure 18 The performance of the high-throughput FISHnCHIPs assay was further verified. By using two closely adjacent frozen sections to compare the frequency and spatial distribution of cell types observed under 10x and 60x objective lenses, a highly correlated cluster size (Pearson correlation, r = 0.95) between the 10x dataset and the 60x dataset was shown. Figure 18 A shows the experimental dataset generated under a 10x objective lens, including a figure showing all segmented cells (Picture a), filtered cells after excluding low-expression cells in the first quality control stage (Picture b), a spatial map of cells after Leiden clustering (Picture c), and a Uniform Manifold Approximation and Projection (UMAP) illustration of the clustering (Picture d). Figure 18 B shows the experimental dataset generated under a 60x objective lens, including a figure showing all segmented cells (Picture e), filtered cells after excluding low-expression cells in the first quality control stage (Picture f), a spatial map of cells after Leiden clustering (Picture g), and a Uniform Manifold Approximation and Projection (UMAP) illustration of the clustering (Picture h). Figure 18 A and Figure 18 B both have a scale bar of 500 μm. Figure 18 C provides a scatter plot of the number of cells in each cluster detected at 60x and 10x. The dashed line represents the x = y line. This comparison shows that although the throughput is increased at a lower magnification (such as 10x) compared to a higher magnification (such as 60x), the quality of the FISHnCHIPs data does not show a significant decline.

[0028] Figure 19Shows imaging of cancer-associated fibroblast (CAF) subtypes using the in situ hybridization method described herein. Imaging of two cancer-associated fibroblast (CAF) subtypes was performed on frozen biopsies of human colorectal cancer (CRC) tissue using the FISHnCHIPs method. Co-staining of epithelial cells (labeled by tumor marker genes) and immune cells (labeled by human leukocyte antigen (HLA) genes) in CRC tissue was performed using FISHnCHIPs. Figure 19 A provides exemplary images of cancer-associated fibroblast 1 (CAF-1), cancer-associated fibroblast 2 (CAF-2), colonic epithelial cells, and immune cells (HLA gene) in Panels a to d, respectively. The scale bar is 200 μm. Figure 19 B provides magnified regions of the white box inset in Composite Panel i in Panels ii to v, with a scale bar of 25 μm. Figure 19 B shows the centroids of the segmented cell masks of CAF-1 (vi), CAF-2 (vii), and immune cells (viii) in Panels vi to viii. The scale bar is 200 μm. Figure 19 B shows box plots of the number of immune cells within a 100-μm radius of CAF-1 (vi) and CAF-2 (vii) cells. The number of cells in the box plots is: CAF-1: 2946 cells; CAF-2: 2671 cells. The box plots show the median (midline), first and third quartiles (box limits), and 1.5 times the interquartile range (whiskers). p = 1.4×10 -72 , two-sided Mann-Whitney U test. As Figure 19 shown in B, different spatial organizations of the two CAF subtypes were observed. The CAF-2 subtype expressing genes related to muscle contraction seems to promote an immunosuppressive microenvironment, in which fewer immune cells were detected near CAF-2 compared to the CAF-1 subtype (0.74-fold, p = 1.4×10 -72 (two-sided Mann-Whitney U test)). The frequency of occurrence of immune cells near CAF-2 was 0.74-fold lower than that near CAF-1. As shown in this example, the in situ hybridization method described herein can not only characterize cells from healthy tissue samples but also cells from diseased tissue samples (such as cancer tissues). Additional insights related to pathological development can be revealed from the spatial organization information of specific cell types within tissue samples.

[0029] Figure 20 Provides an estimate of the signal gain (SG) of the Figure 19 human colorectal cancer (CRC) FISHnCHIPs panel in Figure 20A shows the scRNA-seq gene expression heatmap of the human colorectal cancer (CRC) FISHnCHIPs combination, which is based on the information previously published by Li, H. et al. (Li, H. et al. Reference component analysis of single-cell transcriptomes elucidates cellular heterogeneity in human colorectal tumors. Nat Genet 49, 708–718 (2017)). The reference scRNA-seq data can be downloaded from Gene Expression Omnibus: EGAS00001001945 / GSE81861. Figure 20 B shows the scRNA-seq gene expression heatmap of the human colorectal cancer (CRC) FISHnCHIPs combination, which is based on the more recent scRNA-seq dataset published by Pelka et al. (Pelka, K. et al. Spatially organized multicellular immune hubs in human colorectal cancer. Cell 4734–4752 (2021)). Figure 20 C provides the predicted conserved signal gain (SG) of the human colorectal cancer (CRC) FISHnCHIPs combination, which shows significant signal gain detected in all four cell types. The RNA quality of clinical samples is usually low, which limits the imaging quality of these samples. Using genes that exhibit co-varying expression levels in the methods described herein results in higher robustness and higher signal gain, which is beneficial for imaging of clinical samples.

[0030] Figure 21 Additional FISHnCHIPs technical replicates of human colorectal cancer (CRC) tissues were generated. Figure 21 A provides exemplary FISHnCHIPs images of CAF-1 subtype cells (Picture a), CAF-2 subtype cells (Picture b), colonic epithelial cells (Picture c), and immune cells (HLA gene) (Picture d). Figure 21 The scale bar for all images in A is 250 μm. Figure 21 B shows the combined FISHnCHIPs image of the four cell types in Picture i. The scale bar is 250 μm. Figure 21 B magnifies the white box in Picture i below Pictures ii to v, and the scale bar is shown as 50 μm. Figure 21B provided box plots showing the number of immune cells within a 100-μm radius of CAF-1 (vi) and CAF-2 (vii) cells. Consistent with previous findings, the frequency of immune cell presence near CAF-2 subtype cells was 0.51-fold lower than that near CAF-1 subtype cells. The number of cells quantified in the box plots was: CAF-1: 2,548 cells; CAF-2: 2,199 cells. The box plots show the median (midline), the first and third quartiles (box limits), and 1.5 times the interquartile range (whiskers). p = 8.5 × 10 -142 , two-sided Mann-Whitney U test. The consistency of the in situ hybridization imaging results in cancer tissues demonstrated the reproducibility of the method described herein.

[0031] Figure 22 Triple-color immunofluorescence (IF) staining of the immunomarkers CD68, the CAF-1 marker PDPN, LUM, PDGFA, and the CAF-2 markers aSMA, MMP2 on four frozen human colorectal cancer tissues was provided. All images were contrasted at the 1st to 99.9th percentiles of the maximum intensity in each channel. The scale bar for all images was 250 μm. The observed CAF-1 and CAF-2 patterns were consistent with immunofluorescence (IF) labeling, confirming the specificity and sensitivity of the method.

[0032] Figure 23 Dual-color single-molecule FISH (smFISH) staining of the CAF-1 markers DCN, MMP2 and the CAF-2 markers ACTA2, TAGLN at different concentrations on frozen human colorectal cancer tissues was provided. DCN and TAGLN were stained together, while MMP2 and ACTA2 were stained together on the same sample. Single-molecule FISH staining of the pan-fibroblast SPARC was included as a positive control. The scale bar for all images was 10 μm. Contrary to the strong signals detected in the FISHnCHIPs illustrated in Figure 21 , the smFISH staining for DCN or MMP2 (markers of CAF-1) and TAGLN or ACTA2 (markers of CAF-2) was weak, and the CAFs subtypes were difficult to distinguish from background noise. Therefore, the method described herein for labeling cell types based on multiple co-regulated genes is effective in signal amplification compared to conventional methods such as single-molecule FISH.

[0033] Figure 24 Summarized the software workflow for the combined design and evaluation of cell-centered and gene-centered strategies for the in situ hybridization method disclosed herein.

[0034] Definition

[0035] As used herein, the term "spatial transcriptomics" refers to molecular profiling methods that allow measurement of all gene activity (i.e., transcription) in a tissue and allow mapping of the location of the activity. Spatial transcriptomics includes methods for assigning cell types (identified by mRNA readout) to their locations in tissue sections. Commonly used methods in spatial transcriptomics include fluorescence in situ hybridization (FISH), in situ sequencing, in situ capture, and computational reconstruction.

[0036] As used herein, the term "hybridization" refers to the formation of a hybrid nucleic acid molecule with a complementary nucleotide sequence. Hybridization generally occurs between DNA and / or multiple RNAs in forms such as DNA:DNA, DNA:RNA, or RNA:RNA. The hybridization process can occur naturally in vivo (e.g., during DNA replication and transcription of DNA into RNA) or in vitro (such as during nucleic acid sequencing or polymerase chain reaction (PCR)).

[0037] As used herein, the term "in situ hybridization" or "ISH" refers to an established and highly sensitive molecular biology technique that can be used to detect the presence or location of nucleic acids in a preserved cell or tissue sample. The method is based on the complementary binding of a nucleotide probe to a specific target sequence of DNA or RNA. This technique can be further divided into two types according to the visual presentation method, namely: fluorescence in situ hybridization (FISH) or chromogenic in situ hybridization (CISH).

[0038] As used herein, the term "fluorescence in situ hybridization" or "FISH" refers to in situ hybridization visualized by a fluorescent signal. A typical fluorescence in situ hybridization experiment requires a fluorescent copy of the probe sequence or a modified probe sequence that can be fluorescently labeled later. The probe sequence is designed to be complementary to a specific target sequence. During hybridization, the probe and the target strand are separated into single strands, for example, by heating or chemically disrupting existing hydrogen bonds. Then, the strands separated from the probe and the target are allowed to reanneal through complementary regions, thereby forming new hydrogen bonds. After hybridization, the probe can be visualized, for example, using a fluorescence microscope. There are also other fluorescence in situ hybridization variants, such as multiplex FISH, spectral karyotyping, cross-species color banding, and comparative genomic hybridization that allows multicolor imaging of fluorescent signals. Single molecule FISH (smFISH) (also known as smRNAFISH or RNAFISH) can be used for imaging and quantification of individual RNA molecules. Multiplex error-robust FISH (MERFISH) can simultaneously measure the copy number and spatial distribution of a large number of RNA species in a single cell.

[0039] As used herein, the term "co-expression" or "co-expressed" is used to describe genes that are expressed within the same cell, which means that these genes are also expressed very close to each other spatially within a tissue.

[0040] As used herein, the term "co-regulated" or "co-regulatory" is used to describe genes that exhibit coordinated changes in terms of gene expression levels, i.e., co-varying genes.

[0041] As used herein, the terms "co-vary", "coordinate change", or "co-variance" refer to the consistency of changes in gene expression levels between two or more genes in terms of the direction of change (increase or decrease) and time. The term "co-vary" refers to a positive correlation between the expression levels of genes in a cell. For example, the expression levels of two or more genes can increase or decrease simultaneously. The magnitude of the change can also be coordinated. Correlation analysis is one way to identify co-regulated or co-expressed genes. The default measure of correlation is the Pearson correlation coefficient. Methods for calculating such a correlation coefficient are well-established in the art. In addition to the Pearson correlation coefficient, other potential methods for calculating correlation coefficients include mutual information, Spearman's rank correlation coefficient, and Euclidean distance calculations. As used herein, the term "gene expression level" refers to the copy number of RNA in a cell or the transcriptional level of RNA from a gene in a cell. The expression level of a gene within a cell is the combined result of its synthesis and degradation. In the context of the present invention, "co-regulated" genes typically exhibit co-variance in terms of expression levels. This is because, for eukaryotic transcription or RNA synthesis, co-regulated genes are likely to be co-transcribed and they may share common regulatory elements or mechanisms (such as transcription factors, enhancers, and repressors). For degradation, the RNA copy number may be co-regulated by post-transcriptional mechanisms (such as miRNAs).

[0042] As used herein, the term "cell-centered" refers to a strategy for applying the in situ hybridization method described herein. As an initial step, this method requires the user to input a list of marker genes that define a cell type. In the "cell-centered" strategy, the marker genes correspond to the cell type of interest defined by the user. This definition can be based on existing information (such as information published in the literature or previous experimental observations). For example, as Figure 2As shown, when designing a panel for in situ hybridization, five known cell types (renal macrophages, glomerular endothelial cells, loop of Henle (LOH) cells, collecting duct (CD) cells, and glomerular podocytes) are predefined. In addition to the "cell-centered" strategy, a different "gene-centered" strategy of this method can be adopted. As used herein, the term "gene-centered" in situ hybridization refers to a method in which the initial input is a set of thresholds / parameters to identify a set of genes whose expression levels co-vary, rather than a predefined gene that defines a specific cell type as defined by the user. These gene sets can be "gene expression programs" or "gene modules". Various data types (e.g., sequencing-based spatial transcriptomics, sorted and unsorted scRNA-seq data) can also be used as reference for the purposes of the methods described herein. The "gene-centered" strategy can be used to image multiple gene expression programs and the acquired signals can be further processed, e.g., by quality control (QC), normalization, and clustering, in order to characterize cells in a more unbiased manner. For example, since cell types can also be defined by the expression of multiple gene expression programs, by decoding the acquired "gene-centered" signals, one of ordinary skill in the art can classify the imaged cells into various cell types based on their expression profiles.

[0043] As used herein, the terms "gene module", "gene regulatory module", or "gene expression program" refer to multiple genes that exhibit coordinated changes in their expression profiles under a given set of conditions (such as the binding of the same set of transcription factors or co-factors). In the context of the methods described herein, multiple predefined genes exhibit co-variation in their expression levels within a cell. These genes are co-regulated biologically and can be, but are not limited to, markers of a specific cell type, differentially expressed genes of a specific cell type, markers of a gene expression program or gene regulatory module, or markers of a biological pathway. For example, the "muscle contraction program" refers to multiple genes related to the function of muscle contraction, while the "neuron program" refers to multiple genes related to neurons. Mechanisms such as the action of cis / trans regulatory sequences, the binding of non-coding RNAs, etc. can be used as "gene expression programs". "Gene expression programs" can be obtained from algorithmic techniques in the art that identify sets of genes with co-varying expression levels. For example, the clustering results of a gene-gene correlation matrix are "gene modules", which will be used as input for subsequent signal detection. Methods for obtaining "gene modules" or "gene expression programs" can include various unbiased methods established in the art.

[0044] As used herein, the term "biological pathway" includes a set of protein / compound-coding genes that interact continuously with each other to initiate a biological process or form a certain product. According to databases or literature, the number of genes in a "pathway" is usually less than that in a "module". For example, in the Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway annotation, "pathway" is at a lower level compared to "module". For example, biological pathways can be derived from co-gene expression changes through gene set enrichment analysis.

[0045] As used herein, the term "signal gain" or "SG" refers to the ratio of the sum of the counts of a predetermined target gene to the sum of the counts of the most differentially expressed gene. Signal gain quantifies the expected signal enhancement when using the in situ hybridization method described herein compared to conventional methods (such as single-gene FISH). The SG metric can be easily explained. For example, if the predicted SG is 10, it is predicted that the cells labeled by the in situ hybridization method will be 10 times brighter. In Figure 4 the kidney FISHnCHIPs experiments shown, 4 out of 5 cell types had a higher experimentally determined brightness than expected. The minimum threshold should be determined by the user according to the situation, taking into account the signal specificity ratio threshold.

[0046] As used herein, the term "signal specificity ratio" or "SSR" refers to the ratio of the sum of the counts of a predetermined target gene in the target cell type to the sum of the counts in the cell type most likely to be off-target. Signal specificity ratio quantifies the predicted "noise" when using the in situ hybridization method described herein compared to conventional methods (such as single-gene FISH). When the SSR is close to unity, the fluorescence intensity of the cell type of interest should be equal to that of the off-target cell type, making them indistinguishable. The SSR metric can be easily explained. For example, if the predicted SSR is 10, it is predicted that the target cells labeled by the in situ hybridization method will be 10 times brighter than the off-target cells. In Figure 4 the kidney FISHnCHIPs experiments shown, all 5 cell types had a lower experimentally determined background noise than expected. The minimum threshold should be determined by the user according to the situation, taking into account the SG threshold. It should be emphasized that "SSR" and "SG" are predictive and depend on the quality of the input dataset.

[0047] As used herein, the term "adjusted Rand index" or "ARI" refers to a term that measures the similarity between two data clusters. ARI is an opportunity-corrected version of the Rand index, which establishes a baseline by using the expected similarity of all pairwise comparisons between clusters specified by a random model. ARI can be used to quantify and compare clustering accuracy when using the in situ hybridization method described herein compared to conventional methods (such as single-gene FISH).

[0048] As used herein, the term "ground truth" refers to information that is known to be true or genuine as provided by direct observation or measurement (i.e., empirical evidence), rather than information provided by inference.

[0049] As used herein, the term "single-cell RNA sequencing" or "scRNA-seq" refers to state-of-the-art sequencing methods that allow the detection of the expression profiles of individual cells. Single-cell RNA sequencing reveals the heterogeneity and complexity of RNA transcripts within individual cells and reveals the composition of different cell types and functions in highly organized tissues / organs / organisms.

[0050] As used herein, the term "preprocessing" refers to the data preparation and operations performed on the original input dataset.

[0051] As used herein, in the context of selecting marker genes, the term "targeted" or "supervised" refers to the selection of one or more genes based on prior knowledge of the expression levels or biological specificities of reference genes or markers. For example, the cell-centered strategy of the methods described herein is a targeted approach. In a targeted approach, the user needs to consider genome-wide gene co-expression to ensure that the set of genes they select is specific to the target cell type. A targeted approach can be employed when an untargeted approach fails to produce specific markers or genes that match prior knowledge or existing experimental results.

[0052] As used herein, in the context of selecting co-expressed genes, the term "untargeted" or "unsupervised" refers to the selection of genes without prior knowledge of the expression levels or biological specificities of the genes. For example, the gene-centered strategy of the methods described herein is an untargeted approach. "Untargeted" or "unsupervised" selection of genes can allow for the clustering of cells based on the inherent similarity of expression patterns without relying on previously known markers or classes. Untargeted approaches are applicable to tissues or samples for which there is little or no prior literature. Additionally, untargeted approaches have the potential to reveal previously unknown cell types.

[0053] As used herein, the term "identity program" refers to a set of genes that together are responsible for determining the identity or a specific function of a particular cell type or tissue in an organism.

[0054] As used herein, the term "activity program" refers to a set of genes that are turned on or off in response to specific environmental cues or cellular signals.

[0055] As used herein, the term "detectable label" refers to a label that allows a labeled target to be distinguished from an unlabeled target, typically by detecting a visual signal from the label. Detectable labels can be proteins, nucleotides, or compounds. For example, commonly used detectable labels include, but are not limited to: fluorescent proteins, isotopes, and mass tags. Fluorescent protein labeling in combination with imaging techniques is widely used in biological research, which allows the detection of labeled targets in fixed or live samples. Visual presentation of fluorescent protein labeling typically requires excitation by light in a specific wavelength range (excitation wavelength range), which allows the emission of detectable light in a different wavelength range (emission wavelength range). Collecting signals in the emission wavelength range allows for the visual presentation of the fluorescent protein, thereby identifying the presence, location, and / or quantity of the labeled target.

[0056] As used herein, the term "combination of emission signals" refers to a collection of emission signals from multiple predetermined genes (having the same label or multiple identical tags or a similar label or multiple similar tags that emit the same type of signal), which can be detected simultaneously by methods known in the art. In the context of the present disclosure, a fluorescence microscope can be used to detect the combined emission signals from a set of predetermined genes (e.g., gene modules or gene expression programs) from the same fluorophore using a single set of excitation and emission wavelengths. The detected signal will be a combination of all the emission signals from each labeled gene in the set of predetermined genes, without distinguishing the signals from each individual gene.

[0057] As used herein, the term "multiple emission signals" refers to a collection of different signals emitted by multiple detectable labels. In the context of the present disclosure, multiple gene modules or gene expression programs can be detectably labeled, each gene module or gene expression program containing multiple predetermined genes. Each gene module or gene expression program can be labeled with a different type of label (such as a fluorophore), which allows for differentiation between different gene modules or gene expression programs when measuring emission signals. In a gene module or gene expression program, individual genes are labeled with the same label (such as a fluorophore). "Multiple emission signals" refers to the different signals emitted by the excitation labels from each gene module or gene expression program. Detailed Description

[0058] High-throughput spatial characterization of cells in intact biological samples has been a technical challenge. Existing methods are often inefficient, costly, and have poor scalability. To address these limitations, as described herein, the present disclosure provides an in situ hybridization (ISH) method for cell heterogeneity characterization that is capable of precisely delineating cell types without disrupting tissue architecture.

[0059] The following detailed description is merely exemplary in nature and is not intended to limit the present invention or the application and uses of the present invention. Additionally, there is no intention to be bound by any theory presented in the foregoing background of the invention or the following detailed description.

[0060] The present disclosure provides an in situ hybridization (ISH) method that simultaneously labels multiple genes in a specific cell type or molecular pathway, rather than a single gene, and measures the collective signal emitted from these multiple genes in each cell. By targeting multiple genes, each cell will have a large number of detectable labels (the product of the number of transcript copies per cell, the number of probes per transcript, and the number of target genes). Depending on the cell type or biological pathway of interest, the signal gain is greater than 1, 10, 100, or 1000-fold, making detection more robust and easier. Figure 1 An overview of the method described herein is shown. Instead of focusing on precisely determining the possible discrimination of a single gene, the present invention enhances the signal by increasing the signal of a set of predetermined genes that are related to each other by co-variation or co-variance in expression levels (e.g., due to the predetermined genes belonging to the same pathway). These predetermined genes can be detected simultaneously using the same detectable label (e.g., fluorophore), thereby amplifying the collected signal. Compared to conventional ISH methods that determine the attribution of each individual gene to the overall signal, the method of the present invention utilizes the sum of the signals obtained from different predetermined genes, which can improve the signal-to-noise ratio of the collected data.

[0061] The method described herein is applicable to any cell population with a known transcriptomic profile, thereby allowing investigation of cell states that cannot be accessed by antibody-based methods. The method also allows determination of the spatial location of enhanced cell signals in tissues or 3D cell clusters / structures without disrupting the tissue architecture, thereby providing insights into the spatial organization of cells within the tissue.

[0062] The in situ hybridization method described herein can be achieved through three main steps. A) Design a combination of predetermined genes or use an existing set of predetermined genes to be targeted; B) Label and image the genes; and C) Finally, collect the data and process the collected data. Based on the way the gene combination is designed, the in situ hybridization method can be further subdivided into two different strategies, namely: a cell-centered strategy and a gene-centered strategy.

[0063] The present disclosure provides examples of the cell-centered strategy and the gene-centered strategy of the in situ hybridization method. As Figure 2 exemplarily shown, a cell-centered FISH method was performed on five selected cell types in the mouse kidney. For example, Figure 5Gene-centric and gene-centric FISH methods based on an 18-gene module in mouse cortex are presented. Both strategies effectively delineate cell types in tissue samples, showing results consistent with existing methods. In addition, the methods described herein demonstrate improved signal intensity. In the cell-centric strategy, fluorescence intensity per cell increased approximately 6- to 39-fold in all 5 cell types, as shown in Figure 2. Figure 3 As shown in A. Figure 6 C, Signal gain in gene-centric strategies can be about 1.2 to 22.3 times brighter than profiling using individual marker genes. The workflows of these methods are briefly described below.

[0064] Cell-centered in situ hybridization (ISH) strategy

[0065] 1. Identify a set of genes by calculating the covariation of expression of other genes with reference cell type defining markers;

[0066] 2. Design ISH probes for this series of marker genes;

[0067] 3. Evaluate ISH probe combinations;

[0068] 4. exposing the cell sample to the probe and visually presenting the probe after exposure;

[0069] 5. quantifying the detectable signal obtained from the probe bound to its target; and

[0070] 6. Perform data analysis (e.g., clustering, cell-cell contacts / proximities, and tissue partitioning) and present graphical data of cell clusters / heat maps.

[0071] Gene-centered in situ hybridization (ISH) strategy

[0072] 1. Identify co-varying gene sets (e.g., gene expression programs, gene modules, or pathways of interest) from a reference dataset or a database of interest.

[0073] 2. Design ISH probes for these gene sets;

[0074] 3. Evaluate ISH probe combinations;

[0075] 4. exposing the cell sample to the probe and visually presenting the probe after exposure;

[0076] 5. quantifying the detectable signal obtained from the probe bound to its target; and

[0077] 6. Perform data analysis (e.g., clustering, cell-cell contacts / proximities, and tissue partitioning) and present graphical data of cell clusters / heat maps.

[0078] As described above, one feature of the present invention will be the use of in situ hybridization probes targeting individual gene sets or multiple gene sets (not individual genes) that will be labeled with the same label (such as a fluorophore, a readout probe, or a sequencing tag). Another feature of the present invention is to group genes based on the correlation between gene expression and cell type marker genes and to cluster the correlation matrix. Gene-gene correlation analysis is employed as an algorithmic approach for detecting the above-described gene sets across the entire transcriptome or for cell type marker genes. Another technical feature of the present disclosure is the sequential hybridization of multiple gene modules to allow for de novo reconstruction of cell types in tissues.

[0079] Compared with conventional methods, an improved in situ hybridization (ISH) method for cell heterogeneity characterization enhances signal sensitivity. In one example, compared with conventional in situ hybridization methods, the sensitivity can be increased by about 2 to 200 times (depending on the required "cell type resolution"). In another example, the sensitivity can be increased by about 20 to 200 times. In another example, the signal sensitivity can be enhanced by at least 2 times. In some examples, the signal sensitivity can be enhanced by at least about 5 times, at least about 10 times, at least about 20 times, at least about 30 times, at least about 40 times, at least about 50 times, at least about 60 times, at least about 70 times, at least about 80 times, at least about 90 times, or at least about 100 times. In some examples, the signal sensitivity can be enhanced by about 2 to 20 times, 20 to 100 times, about 50 to 100 times, or about 50 to 200 times. Contrary to existing marker gene selection strategies that minimize redundancy or use compressive sensing to improve the multiplexing efficiency of individual genes, the method described herein exploits the redundancy of correlated genes to improve sensitivity and robustness. For example, as Figure 3 shown in the box plot of A, the fluorescence signal gain per cell using the method described herein is about 6 to 39 times higher than that of conventional single molecule FISH. In addition, the method described herein reduces the requirements for experimental equipment, experimental costs, and assay time. Imaging with a large field of view (FOV) at a lower magnification can accelerate the imaging process while maintaining comparable imaging quality because there is a higher signal-to-noise ratio even at a lower magnification (10x), as Figure 13 exemplarily shown in. Using co-expressed genes, this in situ hybridization method is also robust in analyzing clinical tissues, which are typically characterized by low RNA amounts. In addition, optical crowding in small cells usually hinders the accurate decoding of highly expressed RNA transcripts, but the method disclosed herein allows for the simultaneous delineation of co-localized genes at the single cell level. Compared with conventional multiplex immunostaining methods, this method provides flexibility and throughput because it utilizes specially designed inexpensive oligonucleotide probes. In addition, the labeling of antibody combinations often requires individual optimization, but the detectable signals from the in situ hybridization method described herein are more consistent because of the higher probe hybridization efficiency across the entire transcriptome.

[0080] Thus, as described herein, the present disclosure provides a method for in-situ characterization of cells in a biological sample.

[0081] In one instance, the method includes: contacting a biological sample with a plurality of probes that bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes. In one instance, the methods described herein are in vitro methods. In another instance, the methods described herein are implemented on a biological sample obtained from a subject. The biological sample can be, but is not limited to, a tissue sample, a culture sample (such as an in vitro or ex vivo sample or an organoid), or a biopsy sample. The biological sample can be an unprocessed biological sample (fresh sample) or a processed biological sample (e.g., fixed, frozen, embedded, or tissue-cleared sample). In one instance, the biological sample is fixed or presented on an imaging slide, coverslip, or cell culture dish. In a specific instance, the biological sample can be formalin-fixed paraffin-embedded (FFPE) tissue, which typically has a lower RNA quality that can affect the labeled signal intensity. Compared to the conventional methods described above, signals from FFPE tissue samples can be easily detected using the methods described herein due to the (high) signal intensity. In some cases, the biological sample contains cells of the same tissue type. In some other cases, the biological sample contains different types of cells. For example, as Figure 13 shown, the methods described herein can be used to analyze an entire tissue section, which encompasses both neuronal and non-neuronal cell types simultaneously. In other cases, Figure 9 shows a cell type profiling performed in the mouse cortex that only encompasses neuronal cell types. Thus, the biological sample can contain a homogeneous or heterogeneous cell population. In some instances, the biological sample can contain healthy cells, diseased cells, or both. Figure 19 An example is provided of imaging cancer-associated fibroblast (CAF) subtypes from a frozen biopsy of human colorectal cancer (CRC) tissue using the in-situ hybridization method described herein. In one instance, the biological sample contains cells adhered to a solid matrix. In another instance, the biological sample is one of a plurality of samples within a tissue array or one of a plurality of samples on a coverslip.

[0082] In one instance, the probes described herein are probes made of nucleic acid. The nucleic acid probe can be ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In another instance, the probes described herein contain a nucleotide sequence. In another instance, the probe contains a domain that specifically binds to the ribonucleic acid transcript of one of the predetermined genes. The binding between the probe and the target RNA transcript can be hybridization, which is mediated by the formation of hydrogen bonds between complementary nucleotides.

[0083] In one instance, the selection of multiple predetermined genes is an unsupervised selection, a supervised selection, or a combination of both. Unsupervised methods are applicable to tissues or samples with little or no prior literature. Additionally, unsupervised methods have the potential to reveal previously unknown cell types. In cases where unsupervised methods do not yield specific markers or genes that match prior knowledge or existing experimental results, supervised methods can be employed. In supervised methods, the user needs to consider genome-wide gene co-expression to ensure that the set of genes they select is specific to the target cell type.

[0084] In one instance, the probes target multiple predetermined genes. The multiple predetermined genes include at least one gene and at least one other gene that exhibit co-variation in terms of expression levels. The methods described herein differ from conventional ISH methods (such as MERFISH, seqFISH, osmFISH, smFISH, or RNA scope) in that the methods described herein use probes to simultaneously hybridize to the transcripts of multiple co-regulated gene targets (regulatory modules / gene expression programs), while conventional methods only label a single target gene. The at least one predetermined gene and the at least one other predetermined gene can include, but are not limited to, markers of a specific cell type, differentially expressed genes of a specific cell type, markers of a gene expression program or gene regulatory module, markers of a biological pathway, or combinations thereof.

[0085] In another instance, the at least one other gene includes, but is not limited to, one or more input datasets, such as: bulk RNA sequencing, single-cell RNA sequencing, microarray datasets, chromatin accessibility sequencing, methylation sequencing, DNA-associated protein sequencing, spatial transcriptomics sequencing, multiplexed RNA fluorescence in situ hybridization, multiplexed immunohistochemistry, bioinformatics databases, or any user-defined dataset, or combinations thereof. In another instance, the bioinformatics database is selected from the group consisting of the Kyoto Encyclopedia of Genes and Genomes (KEGG), Panther, or the Database for Annotation, Visualization, and Integrated Discovery (DAVID), or Gene Ontology (GO), or combinations thereof. Additionally, prior knowledge regarding biochemical pathways, transcription factor motifs, chromatin accessibility, bulk gene expression, sequencing-based spatial transcriptomics, or cis-regulatory sequences can be incorporated as part of the input. The in situ hybridization method can be combined with split probes, tissue clearing, or amplification to further enhance the signal. The availability of scRNA-seq methods and integrated cell atlas reference datasets can facilitate the use of the methods described herein to map a broader range of cell types.

[0086] Based on the input dataset, a person skilled in the art will be able to use existing mathematical tools to calculate whether two genes are likely to exhibit co-variation (i.e., co-regulation) in terms of expression levels within a cell, for example, through clustering of genes in a gene-gene correlation matrix, dimensionality reduction analysis (non-negative matrix factorization (NMF)), differential expression gene analysis, or a combination thereof. Mathematical analyses such as Pearson coefficient, mutual information, Spearman correlation coefficient, Euclidean distance, non-negative matrix factorization, principal component analysis, Louvain or Leiden community detection algorithms, rank-based, centroid-based clustering algorithms, or non-parametric Wilcoxon rank sum test can be employed to perform correlation, clustering, and dimensionality reduction analyses.

[0087] In some instances, co-regulated genes are further evaluated to identify a plurality of predetermined genes. For example, the signal gain (SG) of co-regulated genes is calculated to predict the expected improvement in signal intensity when using the methods described herein compared to conventional ISH methods. The signal gain (SG) is the ratio of the sum of the signals of co-regulated genes to the signal of a single gene (such as a differentially expressed gene or the gene with the highest expression). In some instances, a plurality of predetermined genes are identified when the SG is higher than 1, 2, 5, 10, or 50. In another instance, the signal specificity ratio (SSR) of co-regulated genes is calculated to predict the (background) "noise" caused by off-target cell types in the signals generated when using the methods described herein compared to conventional ISH methods. The signal specificity ratio (SSR) is the ratio of the sum of the signals of co-regulated genes in target cells to the off-target cells or the cell cluster with the second highest expression. In some instances, a plurality of predetermined genes are identified when the SSR is higher than 2, 5, 10, or 50. Figure 4 B provides an exemplary figure showing the SG and SSR calculated in a cell-centered FISHnCHIP experiment using signal readings of 5 cell types in a mouse kidney.

[0088] In one instance, the probes described herein include a detectable label. In some instances, the detectable label can be detected directly. In other instances, the detectable label can be detected when contacted with one or more reagents (sandwich label). In some instances, the detectable label is included in a separate readout probe. In one instance, the detectable label is a fluorophore, a fluorescent protein, or a fluorescent dye. As described herein, the probe can emit a detectable signal when bound to a target ribonucleic acid transcript, which allows for the detection of the signal. For example, when the signal is a fluorophore, the signal can be detected by exciting the fluorophore near its excitation maximum and observing the fluorescence emission near its emission maximum. The resulting emission can be detected by an optical imaging instrument such as a fluorescence microscope. Commonly used fluorophore colors include, but are not limited to: a) near infrared; b) far red; c) red; d) yellow; e) green; f) cyan; and g) blue. Although some of the examples provided herein are based on fluorescence in situ hybridization (FISH), those skilled in the art will understand that the same improved in situ hybridization (ISH) methods are compatible with other detection methods and detectable labels such as chromophores, radioisotopes, and chromogens.

[0089] Fluorescently labeled readout probes can be designed for transcriptome analysis in the improved fluorescence in situ hybridization (FISH) methods described herein. The probes are labeled at the 5' or 3' end. Exemplary sequences of probe sequences and tags are listed in Table 1 below:

[0090] Table 1: FISHnCHIPs Readout Probes

[0091]

[0092]

[0093]

[0094] In another instance, the method includes: detecting a combination of emission signals from multiple probes or multiple emission signals. Detection of the combination of emission signals or multiple emission signals allows for amplification of the detectable signal (taking into account the number of genes, the number of transcript copies per cell, and the number of probes per transcript), which increases the signal sensitivity of the methods described herein by approximately 20 to 200-fold. In some instances, the level of the detected emission signal can be quantified and / or processed based on the purpose of the experiment.

[0095] In some instances of the methods described herein, the steps of contacting a biological sample with a plurality of probes that bind to RNA transcripts of a plurality of different predetermined genes and detecting a combined emission signal from the plurality of probes or a plurality of emission signals can be repeated one or more times. This step facilitates imaging of a collection of multiple genes targeted by the probes within the same tissue, thereby allowing for the simultaneous acquisition of multiple data sets.

[0096] In another instance, the method further includes: characterizing a cell based on the combined emission signal or the plurality of emission signals. A cell type can be defined by the expression profiles of a plurality of gene regulatory modules (or gene expression programs). In some cases, the characterization of the cell includes one or more of the following: mapping the location of the cell within the biological sample; identifying the interactions between the cell and one or more other cells; identifying the gene expression pattern of the cell within the biological sample and visually presenting the spatial transcriptome of the cell within the biological sample; and stratifying cancer subtypes to determine the severity of the cancer. Thus, the in situ hybridization method for cell heterogeneity characterization described herein can be used to capture signals from a plurality of gene regulatory modules (or gene expression programs) or even signals from the entire genome, and the resulting signals can be further processed to reveal cell types in a more unbiased manner. In another instance, the characterization of the cell includes: processing the input data set to improve data quality. Methods for processing experimental data obtained from in situ hybridization are known in the art. For example, the experimental data can be subjected to a preprocessing process (such as quality control (QC), normalization, and log / linear transformation). The preprocessed data can be further analyzed by methods such as correlation analysis, clustering analysis, dimensionality reduction analysis, or differential expression gene analysis.

[0097] Thus, as described herein, the present disclosure provides a method for in situ characterization of cells in a biological sample, including: contacting the biological sample with a plurality of probes that bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes, wherein each probe comprises a detectable label and a domain that specifically binds to an RNA transcript of one of the predetermined genes; wherein a signal is emitted when the probe binds to the RNA transcript; detecting a combined emission signal from the plurality of probes or a plurality of emission signals; and characterizing the cell based on the combined emission signal or the plurality of emission signals, wherein the plurality of predetermined genes includes at least one gene and at least one other gene that are co-regulated within the cell. The methods described herein improve the signal-to-noise ratio, reduce instrument requirements, and shorten experimental run times by grouping and labeling multiple co-regulated genes together. The methods described herein allow for the characterization of cells in a biological sample based on information regarding cell type, cell subtype, and cell spatial localization.

[0098] In another instance of the methods described herein, the plurality of predetermined genes are expressed in the kidney, brain, digestive tract, or a combination thereof.Figure 2 An example of cell - centered cell - type profiling performed in the mouse kidney is provided. In addition, Figure 5 exemplary experimental data of cell - type profiling performed in a mouse cerebral cortex sample are shown. Figure 19 Gene - centered cell - type profiling performed in a human colorectal tissue sample is presented. Although the exemplary data demonstrate the use of the methods described herein in the kidney, brain, and digestive tract, those skilled in the art should understand that the method can generally be applied to other organ or tissue types. In addition, the methods described herein can be applied to any biological sample containing cells and are not limited to the exemplified species including mice and humans.

[0099] In one example, multiple predetermined genes are expressed in the kidney, as Figures 2 to 4 shown. In another example, genes are specifically expressed in the loop of Henle cells, collecting duct cells, endothelial cells, podocyte cells, and macrophage cells of the kidney.

[0100] In one example, the multiple predetermined genes expressed in podocyte cells include the genes listed in Table 2(2a). In another example, the multiple predetermined genes expressed in endothelial cells include the genes listed in Table 2(2b). In another example, the multiple predetermined genes expressed in loop of Henle cells include the genes listed in Table 2(2c). In another example, the multiple predetermined genes expressed in collecting duct cells include the genes listed in Table 2(2d). In another example, the multiple predetermined genes expressed in macrophage cells include the genes listed in Table 2(2e).

[0101] Table 2: Figure 2 FISHnCHIPs of the mouse kidney library

[0102]

[0103]

[0104]

[0105]

[0106] In one example, multiple predetermined genes are expressed in neuronal tissue. In another example, the predetermined genes are expressed in the cerebral cortex. Figures 5 to 8 Exemplary gene - centered profiling of 18 gene modules in the mouse cortex is shown.

[0107] In another example, multiple predetermined genes are expressed in gene regulatory modules within the brain, where the gene regulatory modules are selected from M1, M2, M3, M4, M5, M6, M8, M9, M10, M11, M12, M13, M14, M15, M21, M22, M23, and M24. In another example, the multiple predetermined genes expressed in M1 include the genes listed in Table 3(3a). In another example, the multiple predetermined genes expressed in M2 include the genes listed in Table 3(3b). In another example, the multiple predetermined genes expressed in M3 include the genes listed in Table 3(3c). In another example, the multiple predetermined genes expressed in M4 include the genes listed in Table 3(3d). In another example, the multiple predetermined genes expressed in M5 include the genes listed in Table 3(3e). In another example, the multiple predetermined genes expressed in M6 include the genes listed in Table 3(3f). In another example, the multiple predetermined genes expressed in M8 include the genes listed in Table 3(3g). In another example, the multiple predetermined genes expressed in M9 include the genes listed in Table 3(3h). In another example, the multiple predetermined genes expressed in M10 include the genes listed in Table 3(3i). In another example, the multiple predetermined genes expressed in M11 include the genes listed in Table 3(3j). In another example, the multiple predetermined genes expressed in M12 include the genes listed in Table 3(3k). In another example, the multiple predetermined genes expressed in M13 include the genes listed in Table 3(3l). In another example, the multiple predetermined genes expressed in M14 include the genes listed in Table 3(3m). In another example, the multiple predetermined genes expressed in M15 include the genes listed in Table 3(3n). In another example, the multiple predetermined genes expressed in M21 include the genes listed in Table 3(3o). In another example, the multiple predetermined genes expressed in M22 include the genes listed in Table 3(3p). In another example, the multiple predetermined genes expressed in M23 include the genes listed in Table 3(3q). In another example, the multiple predetermined genes expressed in M24 include the genes listed in Table 3(3r).

[0108] Table 3: Figure 5 FISHnCHIPs of mouse cortex library

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116] In another example, such as Figure 9As shown, the present invention provides a gene-centered profiling using 20 gene expression programs in the mouse cortex. The gene-gene correlation analysis of the 20 gene expression programs was performed using the non-negative matrix factorization (NMF) algorithm. In one example, multiple predetermined genes are expressed in gene expression programs selected from Erp, ExcL2, ExcL3, ExcL4, ExcL5p1, ExcL5p2, ExcL5p3, ExcL6p1, ExcL6p2, Hip, IntCckVip, IntNpy, IntPv, IntSst, LrpD, LrpS, NS, Other, Sub, and Syn. In another example, the multiple predetermined genes expressed in Erp include the genes listed in Table 4 (4a). In another example, the multiple predetermined genes expressed in ExcL2 include the genes listed in Table 4 (4b). In another example, the multiple predetermined genes expressed in ExcL3 include the genes listed in Table 4 (4c). In another example, the multiple predetermined genes expressed in ExcL4 include the genes listed in Table 4 (4d). In another example, the multiple predetermined genes expressed in ExcL5p1 include the genes listed in Table 4 (4e). In another example, the multiple predetermined genes expressed in ExcL5p2 include the genes listed in Table 4 (4f). In another example, the multiple predetermined genes expressed in ExcL5p3 include the genes listed in Table 4 (4g). In another example, the multiple predetermined genes expressed in ExcL6p1 include the genes listed in Table 4 (4h). In another example, the multiple predetermined genes expressed in ExcL6p2 include the genes listed in Table 4 (4i). In another example, the multiple predetermined genes expressed in Hip include the genes listed in Table 4 (4j). In another example, the multiple predetermined genes expressed in IntCckVip include the genes listed in Table 4 (4k). In another example, the multiple predetermined genes expressed in IntNpy include the genes listed in Table 4 (4l). In another example, the multiple predetermined genes expressed in IntPv include the genes listed in Table 4 (4m). In another example, the multiple predetermined genes expressed in IntSst include the genes listed in Table 4 (4n). In another example, the multiple predetermined genes expressed in LrpD include the genes listed in Table 4 (4o). In another example, the multiple predetermined genes expressed in LrpS include the genes listed in Table 4 (4p). In another example, the multiple predetermined genes expressed in NS include the genes listed in Table 4 (4q). In another example, the multiple predetermined genes expressed in Other (characterized by high expression of non-coding RNA Meg3 and other genes related to cerebral ischemic injury) include the genes listed in Table 4 (4r).In another instance, the multiple predetermined genes expressed in Sub include the genes listed in Table 4(4s). In another instance, the multiple predetermined genes expressed in Syn include the genes listed in Table 4(4t).

[0117] Table 4: Figure 9 FISHnCHIPs of mouse cortex library

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126] In one instance, the multiple predetermined genes are expressed in the mouse brain, as Figures 13 to 18 shown. In one instance, the multiple predetermined genes are expressed in a gene module selected from any one of gene modules M1 to M53.

[0127] In one example, the plurality of predetermined genes expressed in the M1 gene module include the genes listed in Table 5(5a). In another example, the plurality of predetermined genes expressed in the M2 gene module include the genes listed in Table 5(5b). In another example, the plurality of predetermined genes expressed in the M3 gene module include the genes listed in Table 5(5c). In another example, the plurality of predetermined genes expressed in the M4 gene module include the genes listed in Table 5(5d). In another example, the plurality of predetermined genes expressed in the M5 gene module include the genes listed in Table 5(5e). In another example, the plurality of predetermined genes expressed in the M6 gene module include the genes listed in Table 5(5f). In another example, the plurality of predetermined genes expressed in the M7 gene module include the genes listed in Table 5(5g). In another example, the plurality of predetermined genes expressed in the M8 gene module include the genes listed in Table 5(5h). In another example, the plurality of predetermined genes expressed in the M9 gene module include the genes listed in Table 5(5i). In another example, the plurality of predetermined genes expressed in the M10 gene module include the genes listed in Table 5(5j). In another example, the plurality of predetermined genes expressed in the M11 gene module include the genes listed in Table 5(5k). In another example, the plurality of predetermined genes expressed in the M12 gene module include the genes listed in Table 5(5l). In another example, the plurality of predetermined genes expressed in the M13 gene module include the genes listed in Table 5(5m). In another example, the plurality of predetermined genes expressed in the M14 gene module include the genes listed in Table 5(5n). In another example, the plurality of predetermined genes expressed in the M15 gene module include the genes listed in Table 5(5o). In another example, the plurality of predetermined genes expressed in the M16 gene module include the genes listed in Table 5(5p). In another example, the plurality of predetermined genes expressed in the M17 gene module include the genes listed in Table 5(5q). In another example, the plurality of predetermined genes expressed in the M18 gene module include the genes listed in Table 5(5r). In another example, the plurality of predetermined genes expressed in the M19 gene module include the genes listed in Table 5(5s). In another example, the plurality of predetermined genes expressed in the M20 gene module include the genes listed in Table 5(5t). In another example, the plurality of predetermined genes expressed in the M21 gene module include the genes listed in Table 5(5u). In another example, the plurality of predetermined genes expressed in the M22 gene module include the genes listed in Table 5(5v). In another example, the plurality of predetermined genes expressed in the M23 gene module include the genes listed in Table 5(5w). In another example, the plurality of predetermined genes expressed in the M24 gene module include the genes listed in Table 5(5x).In another example, the multiple predetermined genes expressed in the M25 gene module include the genes listed in Table 5(5y). In another example, the multiple predetermined genes expressed in the M26 gene module include the genes listed in Table 5(5z). In another example, the multiple predetermined genes expressed in the M27 gene module include the genes listed in Table 5(5aa). In another example, the multiple predetermined genes expressed in the M28 gene module include the genes listed in Table 5(5ab). In another example, the multiple predetermined genes expressed in the M29 gene module include the genes listed in Table 5(5ac). In another example, the multiple predetermined genes expressed in the M30 gene module include the genes listed in Table 5(5ad). In another example, the multiple predetermined genes expressed in the M31 gene module include the genes listed in Table 5(5ae). In another example, the multiple predetermined genes expressed in the M32 gene module include the genes listed in Table 5(5af). In another example, the multiple predetermined genes expressed in the M33 gene module include the genes listed in Table 5(5ag). In another example, the multiple predetermined genes expressed in the M34 gene module include the genes listed in Table 5(5ah). In another example, the multiple predetermined genes expressed in the M35 gene module include the genes listed in Table 5(5ai). In another example, the multiple predetermined genes expressed in the M36 gene module include the genes listed in Table 5(5aj). In another example, the multiple predetermined genes expressed in the M37 gene module include the genes listed in Table 5(5ak). In another example, the multiple predetermined genes expressed in the M38 gene module include the genes listed in Table 5(5al). In another example, the multiple predetermined genes expressed in the M39 gene module include the genes listed in Table 5(5am). In another example, the multiple predetermined genes expressed in the M40 gene module include the genes listed in Table 5(5an). In another example, the multiple predetermined genes expressed in the M41 gene module include the genes listed in Table 5(5ao). In another example, the multiple predetermined genes expressed in the M42 gene module include the genes listed in Table 5(5ap). In another example, the multiple predetermined genes expressed in the M43 gene module include the genes listed in Table 5(5aq). In another example, the multiple predetermined genes expressed in the M44 gene module include the genes listed in Table 5(5ar). In another example, the multiple predetermined genes expressed in the M45 gene module include the genes listed in Table 5(5as). In another example, the multiple predetermined genes expressed in the M46 gene module include the genes listed in Table 5(5at). In another example, the multiple predetermined genes expressed in the M47 gene module include the genes listed in Table 5(5au).In another instance, the plurality of predetermined genes expressed in the M48 gene module include the genes listed in Table 5 (5av). In another instance, the plurality of predetermined genes expressed in the M49 gene module include the genes listed in Table 5 (5aw). In another instance, the plurality of predetermined genes expressed in the M50 gene module include the genes listed in Table 5 (5ax). In another instance, the plurality of predetermined genes expressed in the M51 gene module include the genes listed in Table 5 (5ay). In another instance, the plurality of predetermined genes expressed in the M52 gene module include the genes listed in Table 5 (5az). In another instance, the plurality of predetermined genes expressed in the M53 gene module include the genes listed in Table 5 (5ba).

[0128] Table 5: Figure 13 FISHnCHIPs of the mouse brain library

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146] In one instance, multiple predetermined genes are expressed in the digestive tract. In another instance, the predetermined genes are expressed in intestinal cells. In another instance, multiple predetermined genes are expressed in cells associated with colorectal cancer. In some instances, the cells can include, but are not limited to, epithelial cells, CAF-1 cells, immune cells, and CAF-2 cells. In another instance, the multiple predetermined genes expressed in epithelial cells include the genes listed in Table 6(6a). In another instance, the multiple predetermined genes expressed in CAF-1 cells include the genes listed in Table 6(6b). In another instance, the multiple predetermined genes expressed in immune cells include the genes listed in Table 6(6c). In another instance, the multiple predetermined genes expressed in CAF-2 cells include the genes listed in Table 6(6d). As Figure 19 shown in B, the method described herein identifies the different spatial organizations of two CAF subtypes, demonstrating the specificity and sensitivity of the ISH method for cell heterogeneity characterization.

[0147] Table 6: Figure 19 FISHnCHIPs of human colorectal cancer libraries

[0148]

[0149]

[0150]

[0151] Although Tables 2 to 6 provide exemplary gene combinations to be targeted in the kidney, brain, and digestive tract in the in situ hybridization methods described herein, it will be appreciated by those skilled in the art that the gene combinations were identified for experimental purposes. Thus, the methods described herein are not limited by the listed exemplary combinations. According to the methods described herein, alternative combinations can be obtained based on user-defined cell types (for cell-centric strategies) or selected gene expression programs (for gene-centric strategies).

[0152] The methods described herein are applicable to profiling cell types within a biological sample, identifying new cell types, and validating new cell types identified from scRNA-seq studies. For example, Figure 13 a large field of view (FOV) in situ hybridization using the gene-centric strategy described herein is provided. As Figure 13 shown in the UMAP (right) of A, an unknown cell cluster has been identified independently of other cell types.

[0153] Similar to conventional methods such as multiplexed single molecule FISH (smFISH), the in situ hybridization method can be used to quantify cell types, derive compartmentalization patterns, and analyze cell-cell interactions. For example, asFigure 11 As shown in A, the method described herein can reveal the spatial pattern of signal intensity. Figure 11 A shows the intensity gradient that occurs along the cortical depth within the mouse cerebral cortex for some gene expression programs. Figure 19 B shows the new cell-cell interactions between immune cells and cancer subtype cells cancer-associated fibroblast 1 (CAF-1) and cancer-associated fibroblast 2 (CAF-2), which were observed using the in situ hybridization method described herein. By grouping multiple genes and labeling them together to improve the signal-to-noise ratio, the method described herein provides robust and sensitive signal measurements at the cellular level. In addition, by combining the method described herein with multiplexed smFISH, transcriptomic information at both the cellular and transcriptional levels can be obtained simultaneously.

[0154] The sensitivity of the method described herein makes spatial transcriptomics simpler, faster, and less instrument costly, thereby improving the accessibility of spatial assays for a wider range of biomedical research. In addition to neuroscience and oncology, the method can also be used for other biological studies, such as understanding spatial gene coordination during embryonic development or defining the multicellular ecosystem of infectious pathogens. The method can be used for molecular histopathology of formalin-fixed paraffin-embedded (FFPE) tissues, where clinically actionable cell states can be diagnosed precisely and on a large scale. Thus, as described herein, the in situ hybridization method is a sensitive, robust, and scalable spatial transcriptomics method that depicts single cells in tissue samples.

[0155] In another aspect, the present disclosure provides a method for performing / providing a prognosis for a subject having cancer. The method includes: obtaining a sample from the subject. The sample can be, but is not limited to, a biopsy sample obtained from the subject or a tissue sample obtained from cancerous tissue. The method further includes: characterizing one or more cancer cells in the sample using the methods described herein to determine the stage of the cancer. Methods and criteria for determining the stage of cancer have been established in the art. For example, the TNM staging system is the most commonly used staging system by medical professionals. Typically, the TNM staging system includes three dimensions: T for describing the size of the tumor (T1 - T4); N for describing the presence of cancer in the lymph nodes (N0 - N3); and finally, M represents metastasis of the cancer (M0 or M1). Alternatively, in the case of a numerical staging system, the progression of cancer includes five stages, namely, stage 0: carcinoma in situ; stage I: early cancer; stages II and III: cancer has spread to nearby tissues; and stage IV: metastatic cancer. The different stages of cancer can be distinguished by depicting the gene expression of cells within each stage of tissue. Those skilled in the art will be able to determine the stage of cancer based on appropriate information from a biological sample (such as a biopsy sample) revealed by the method. In another example, the method includes: determining a prognosis based on the stage of the cancer.

[0156] In another aspect, the present invention provides a kit for in situ characterization of cells in a biological sample. The kit contains a plurality of probes that bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes described herein. In one example, each probe contains a detectable label. In another example, each probe contains a domain that specifically binds to the ribonucleic acid transcript of one of the predetermined genes described herein. In another example, the kit includes instructions for use.

[0157] In another example of the kit described herein, the plurality of predetermined genes include at least one co-regulated gene and at least one other gene, wherein the at least one gene and the at least one other gene are markers of a specific cell type, differentially expressed genes of a specific cell type, markers of a gene expression program or a gene regulatory module, markers of a biological pathway, or a combination thereof. In another example, the at least one other gene is selected from one or more input data sets. Those skilled in the art can select appropriate input data sets based on the experimental design, including but not limited to: bulk RNA sequencing, single-cell RNA sequencing, microarray data sets, chromatin accessibility sequencing, methylation sequencing, DNA-associated protein sequencing, spatial transcriptomics sequencing, multiplexed RNA fluorescence in situ hybridization, multiplexed immunohistochemistry, bioinformatics databases, or any user-defined data sets, or a combination thereof. In another example, the bioinformatics database used to obtain the set of predetermined genes is selected from the group consisting of the Kyoto Encyclopedia of Genes and Genomes (KEGG), Panther, the Database for Annotation, Visualization and Integrated Discovery (DAVID), the Gene Ontology (GO), or a combination thereof. In addition, prior knowledge about biochemical pathways, transcription factors, or cis-regulatory sequences can also be incorporated as part of the input. Based on the input data set of the predetermined genes, those skilled in the art can use existing mathematical tools to calculate whether two genes are likely to exhibit co-variation in terms of expression levels within a cell.

[0158] In one example of the kit described herein, the plurality of predetermined genes are expressed in the kidney, brain, or digestive tract. In another example, the plurality of predetermined genes are expressed in cancerous tissues. In another example, the plurality of predetermined genes are selected from the genes listed in Tables 2 (2a)-(2e), 3 (3a)-(3r), 4 (4a)-(4t), 5 (5a)-(5ba), and 6 (6a)-(6d).

[0159] In another aspect, the present disclosure provides a kit for in situ characterization of colorectal cancer. In one example, the kit contains a plurality of probes that bind to ribonucleic acid (RNA) transcripts of the plurality of predetermined genes described herein. In another example, the plurality of predetermined genes are selected from the genes listed in Table 6 (6a)-(6d). In another example, each of the plurality of probes contains a detectable label described herein. In another example, each of the plurality of probes contains a domain that specifically binds to the ribonucleic acid transcripts of the plurality of predetermined genes described herein. In another example, the kit further includes instructions for use.

[0160] The present disclosure has been described broadly and generically. Each narrower genus and subgenus grouping that is part of the general disclosure also forms part of the present invention. This includes a general description of the invention with the proviso or negative limitation that any subject matter is removed from the class, regardless of whether or not the removed material is specifically recited herein. Other embodiments are in the following claims and non-limiting examples. In addition, the terms and expressions used herein are used as terms of description and not of limitation, and the use of such terms and expressions is not intended to exclude any equivalent features of the features shown and described or portions thereof, but it should be recognized that various modifications are acceptable within the scope of the invention as claimed. Accordingly, it should be understood that although the present invention has been specifically disclosed by way of preferred embodiments and alternative features, those skilled in the art can make modifications and variations to the invention as embodied herein, and such modifications and variations should be considered to be within the scope of the present invention.

[0161] Experimental Section

[0162] Gene Combination Design and Evaluation Software

[0163] Figure 24Summarized in this paper is the software workflow for in situ hybridization panel design and evaluation. To target specific cell types, the cell-centric strategy of the in situ hybridization method described herein accepts user input of reference markers and cell markers or performs de novo clustering of cell types, and identifies one or more differentially expressed (DE) genes as one or more reference markers. The default measure of correlation is the Pearson correlation coefficient. Other possible measures include mutual information, Spearman rank correlation coefficient, and Euclidean distance. To explore gene expression activity without prior cell type clustering of scRNA-seq data, the gene-centric in situ hybridization method performs feature selection and / or dimensionality reduction (e.g., using non-negative matrix factorization (NMF)), and then performs cluster analysis on the gene-gene correlation matrix to identify gene modules. In the method based on eigen-gene modules, genes highly correlated (>min.corr) with the minimum number of genes (>min.genes) are used as nodes in a network that is constructed based on the gene-gene correlation matrix and partitioned using the Leiden algorithm. Gene partitions can be further sub-clustered using hierarchical clustering based on their log-transformed expression matrix. For the method based on dimensionality reduction, the non-negative matrix factorization (NMF) algorithm that identifies gene programs and their relative contributions can be used. The top N genes from each program are selected to construct the gene-gene correlation matrix. The clustering of the matrix can be refined by setting the correlation range. A hybrid in situ hybridization method is also designed, in which differentially expressed (DE) genes are used as features to construct the gene-gene correlation matrix to identify gene modules. The user is recommended to perform clustering in the gene-gene space to reduce crosstalk. The output gene combinations are evaluated by predicting signal gain and specificity and by simulating the expected cell module expression profiles and clusters. This application demonstrates cell-centric in situ hybridization for a mouse kidney library ( Figures 2 to 4 ), gene-centric in situ hybridization for a mouse cortex library ( Figures 5 to 11 ), and hybrid methods for a mouse brain ( Figures 12 to 18 ) and a human CRC library ( Figures 19 to 23 ).

[0164] The following paragraphs describe the in situ hybridization panel design and evaluation process in more detail:

[0165] Data preprocessing

[0166] Preprocess the scRNA-seq count matrix using the Seurat pipeline. First, quality control (QC) filters out empty droplets and doublets, i.e., cells that express too few or too many unique genes. After QC, three versions of the gene count matrix will be prepared for different downstream analyses: 1) Scale the total count of cells to a constant by dividing by the total count of the cell and multiplying by a scaling factor. The cell-scaled matrix will be used to predict the expected signal of in situ hybridization combinations; 2) Add pseudocounts to the cell-scaled matrix and apply a natural logarithm transformation. The log-transformed matrix will be used for differential gene analysis and gene-gene correlation analysis; 3) Apply a linear transformation to the gene expression vector such that the average gene expression of all cells is 0 and the variance of all cells is 1. The gene-scaled matrix will be used for dimensionality reduction and visual presentation of the heatmap of individual gene expression.

[0167] Combined evaluation

[0168] The in situ hybridization combination can be evaluated by the signal gain and signal specificity ratio:

[0169] Represent the in situ hybridization combination with n genes as P t for target cell type C t ={g1, g2, …, g i , …, g n};

[0170] The number of probes for the genes corresponds to K = {k1, k2, …, k i , …, k t}.

[0171] The predicted signal of a gene g t in cell type C i (denoted as signal(g i , C t )) is defined as the product of k i and the average expression of g t in cell type C i .

[0172] The signal of the combination P t in cell type C t (denoted as signal(P t , C t )) is the sum of the signals of all genes in the target cell type or module.

[0173] Denote g1 as the reference gene and g max as the gene with the maximum signal.

[0174] Define the general signal gain as That is, the ratio of the combined signal to the reference gene signal.

[0175] The conservative signal gain is defined as i.e., the ratio of the combined signal to the highest gene signal.

[0176] Crosstalk can be estimated by calculating the signal specificity ratio (defined as t i.e., the ratio of the combined signal in cell type C t to the combined signal in cell type C t ′) of the combination P between cell type C t and cell type C t ′.

[0177] The general signal specificity is defined as the ratio of the combined signal in the target cell type to the combined signals in all off-target cell types. The conservative signal specificity is defined as the ratio of the combined signal in the target cell type to the combined signal in the cell cluster with the highest predicted crosstalk. The general signal gain is used for cell-centered mouse kidney combinations, while the conservative signal gain is used for all other in situ hybridization combinations. The in situ hybridization combinations can be further evaluated by re-clustering the scRNA-seq dataset using the module-cell expression matrix. The module-cell expression matrix is calculated from the cell-scaled expression matrix by summing the cell counts of genes in the same group. When the module is regarded as a meta-gene, the module-expression matrix can serve as a meta-gene expression matrix. Therefore, conventional clustering methods for processing single-cell gene count matrices can be applied. The module-cell expression heatmap and dimensionality reduction visualization tools (such as UMAP or tSNE) can be used to simulate the reconstruction of cell types according to the in situ hybridization assay described herein.

[0178] Design of cell-centered mouse kidney combinations

[0179] Retrieve the scRNA-seq data and cell markers of mouse kidneys from the NCBI Gene Expression Omnibus (GEO) according to accession number GSE115746. Select the gene with the highest log fold change in the average expression between the targeted cluster and other clusters as the reference marker. Exclude cells with less than 200 or more than 3000 uniquely expressed genes. Exclude cells with mitochondrial genes greater than 50%. Exclude genes expressed in less than 10 cells. Then scale the cells to a sequence depth of 10,000 per cell and perform a log transformation with a pseudocount of 1. Scale the genes so that the average expression of all cells is 0 and the variance of all cells is 1. For each cluster, select genes that are related to the reference marker and have a Pearson correlation greater than 0.5. If there are less than 15 genes highly correlated with the reference, select the top 15 genes. For all clusters, delete genes that appear more than once. For glomerular endothelial cells, the top marker Plat is expressed in only 59.5% of glomerular endothelial cells and is also highly expressed in glomerular podocytes. Therefore, Emcn rather than Plat is used as the reference marker. For renal macrophages, both C1qa and C1qb are used as references. As Figure 2 shown, five cell types are used for imaging. However, all previously annotated cell types have been computationally evaluated, as Figure 4 shown.

[0180] Design of gene-centered mouse cortex assemblies

[0181] The scRNA-seq dataset of the mouse primary visual cortex (VISp) is used to compare with Figures 5 to 8Related mouse brain combinatorial design. First, scale the cells to 10,000, and then binarize the gene expression in the cells by the average expression of all genes in all cells. Genes expressed in less than 5 cells or more than 80% of the total number of cells are filtered out. Gene names starting with "Mt" or "Gm" followed by a number are removed. 330 genes highly correlated with at least 5 genes with a correlation greater than 0.7 are selected as candidate genes. A graph is created based on the 330×330 correlation matrix, and edges with lower correlation (less than 0.6) are removed. By performing Leiden partitioning on the graph with 330 candidate genes, 11 clusters are generated. Hierarchical clustering is performed on the Leiden clusters based on gene expression, so that the dendrogram of genes is cut into k sub-clusters: for larger clusters (>30 genes), k = 6; for medium-sized clusters (11 to 30 genes), k = 4; for smaller clusters (6 - 10 genes), k = 2; for very small clusters (<6 genes), k = 1. After removing sub-clusters with a single gene, 255 genes are distributed in 18 modules, and these genes are not found in the probe design transcriptome database (Hsp25-ps1 and Gstm2-ps1) or are associated with multiple IDs in our probe design transcriptome database (Schip1). The g:GOst is used to perform functional enrichment analysis on the combined genes (referred to as gene set enrichment analysis).

[0182] Mouse cortex combination based on dimensionality reduction

[0183] Non-negative matrix factorization (NMF) provides a low-rank approximation of the gene-cell matrix through the product of two non-negative matrices and is able to capture the structure of co-expressed genes in scRNA-seq data. The gene-contribution matrix of mouse visual cortex neurons was downloaded from Kotliar, D. et al. (Kotliar, D. et al. Identifying gene expression programs of cell-type identity and cellular activity with single-cell RNA-Seq. Elife 8, 1–26 (2019)). The 50 genes with the highest contributions are selected from 20 factors. Gene names starting with "Gm" followed by a number are removed. Clustering of the gene-gene correlation matrix generates one or more gene modules for each program. As Figures 9 to 11As shown, by comparing gene expression heatmaps and gene-gene correlation matrices, most genes with a Pearson correlation (r) higher than 0.3 showed expression across multiple programs and were markers associated with major cell types (such as the major cell type of all inhibitory neurons). Therefore, genes with r higher than 0.3 and lower than 0.02 were excluded. After further discarding genes for which no probes were found, 311 genes were distributed across 20 programs.

[0184] Mouse brain combination of 674 genes

[0185] Using subcluster markers provided by the mouse brain Drop-seq scRNA dataset, up to 50 differentially expressed (DE) genes were identified using the Wilcoxon rank sum test algorithm implemented in Seurat, with a difference of at least 0.25-fold for all subclusters. For each subcluster, genes with the lowest correlation with any DE gene were excluded until the minimum Pearson correlation matrix of the remaining genes was greater than 0.1. To further improve the quality of this combination, genes starting with "mt" and small modules with fewer than 5 genes were excluded, resulting in 53 gene modules containing 674 genes. To evaluate this combination, the scRNA-seq dataset was reclustered using 53 modules as features, and the adjusted Rand index was calculated using the "aricode" package in R. For further comparison, the scRNA-seq data was also reclustered using 1000, 2000, and 3000 highly variable genes as features to simulate single-gene-based multiplex FISH assays ( Figure 14 ).

[0186] Human colorectal cancer (CRC) combination

[0187] Previously, two cancer-associated fibroblast (CAF) subtypes were identified using scRNA-seq. These two subtypes have been further confirmed using a recent scRNA sequencing dataset ( Figure 20 ). Genes expressed in fewer than 5 cells or in more than 70% of the total cell count were filtered out. Gene names starting with "Rp", "Mt", or "Gm" followed by a number were excluded. Based on 125 selected marker genes, a graph was created according to the gene-gene correlation matrix, and edges with lower correlation (less than 0.7) were excluded. Leiden partitioning of the graph produced approximately 20 modules, and 4 modules highly expressed in both CAFs, epithelial cells, and immune cells were selected for demonstrating the in situ hybridization method described in this article.

[0188] In situ hybridization library design and probe sequences

[0189] For all genes, a 25-nucleotide target region was identified using a previously published algorithm (DeTomaso, D. & Yosef, N., 2021). Briefly, reference transcript sequences were downloaded from the GENCODE website (human v24 and mouse m4). A specificity table was calculated using a 15-nucleotide seed, and a 0.2 specificity cutoff was employed. Quadruple repeats (“AAAA”, “TTTT”, “GGGG”, and “CCCC”) were excluded from the possible target regions. The list of resulting readout probe sequences is shown in Table 1. Initially, a total of 56 readout probe sequences were generated, but B16, B48, and B55 were not used.

[0190] Probe Amplification and Preparation

[0191] The probe library (Genscript) was amplified according to a previously published protocol (Kuemmerle, L.B. et al. Probe set selection for targeted spatial transcriptomics. Bioarxiv (2022)). Briefly, the oligonucleotide library was first amplified by limited-cycle PCR using Phusion Hot Start Flex 2x Master Mix with an annealing temperature of 68 °C. During PCR, the T7 promoter sequence was introduced into the reverse primer. Further amplification was achieved by in vitro transcription overnight using a high-yield in vitro transcription kit (NEB, cat. no. E2050S). Then, the RNA template was reverse-transcribed using Maxima H-Reverse Transcriptase (Thermo Fisher, cat. no. EP0753) to generate a DNA-RNA hybrid. The RNA portion was then cleaved off by alkaline hydrolysis, leaving single-stranded DNA (ssDNA), which was then purified by magnetic bead purification and eluted in nuclease-free water (Ambion, cat. no. AM9930). The primers used for PCR were as follows:

[0192] Figure 2 for the mouse kidney library:

[0193] Forward primer: 5’-CTATGCGCTATCCCGGACGC-3’ (SEQ ID NO:53)

[0194] Reverse primer:

[0195] 5’-TAATACGACTCACTATAGGGTCGCATATCCGTACCGGC-3’

[0196] (SEQ ID NO:54)

[0197] Figure 5 Mouse cortical library:

[0198] Forward primer: 5’-CCGTTCAAGACTGCCGTGCTA-3’ (SEQ ID NO:55)

[0199] Reverse primer:

[0200] 5’-TAATACGACTCACTATAGGGCTAGGGAGCCTACAGGCTGC-3’ (SEQ ID NO:56)

[0201] Figure 9 Mouse cortical library:

[0202] Forward primer: 5’-TTGCGTTCGGTCTGAATGCG-3’ (SEQ ID NO:57) Reverse primer:

[0203] 5’-TAATACGACTCACTATAGGGACTCCTGCTCTTTGGGTCCG-3’ (SEQ ID NO:58)

[0204] Figure 13 Mouse brain library:

[0205] Forward primer: 5'-CGCCCTAATCTCCGCTTGGG’-3' (SEQ ID NO:59)

[0206] Reverse primer: 5'-

[0207] TAATACGACTCACTATAGGGGCTTCGACCGAGGGCGAAAT’-3'

[0208] (SEQ ID NO:60)

[0209] Figure 19 Human colorectal cancer library:

[0210] Forward primer: 5’-TGCCCGCCTTTCGTTACTCA-3’ (SEQ ID NO:61) Reverse primer:

[0211] 5’-TAATACGACTCACTATAGGGCGCAATCGTCGGCTAACGGT-3’ (SEQ ID NO:62)

[0212] Cover glass functionalization

[0213] Coverslip functionalization was performed as previously described by Goh, J.J.L. et al. (Goh, J.J.L. et al. Highly specific multiplexed RNA imaging in tissues with split-FISH. Nat Methods 17, 689–693 (2020)) and Lyubimova, A. et al. (Lyubimova, A. et al. Single-molecule mRNA detection and counting in mammalian tissue. Nat Protoc 8, 1743–58 (2013)). Briefly, coverslips (Warner Instruments, cat. no. 64–1500) were washed in 1 M KOH by gentle rocking for 1 h and rinsed three times with MilliQ water. The coverslips were rinsed with 100% methanol and then immersed in an aminosilane solution (methanol solution of 3% vol / vol (3-aminopropyl)triethoxysilane (Merck cat no. 440140) and 5% vol / vol acetic acid (Sigma, cat. no. 537020)) for 2 min at room temperature, then rinsed three times with MilliQ water and dried overnight at 47 °C in an oven. The functionalized coverslips were then used immediately or stored in a dry desiccated environment at room temperature for several weeks.

[0214] Preparation of mouse tissue samples

[0215] Eight-week-old C57BL / 6nTac female mice (InVivos) were used in this study. All animal care and experiments were conducted in accordance with the guidelines of the Institutional Animal Care and Use Committee (IACUC) of the Agency for Science, Technology and Research (A*STAR) in Singapore (IACUC#211580). The mice were euthanized, their kidneys and brains were rapidly collected, and these kidneys and brains were immediately frozen in optimal cutting temperature compound (Tissue-Tek O.C.T.; VWR, cat. no. 25608–930) and then stored at -80 °C. Fresh frozen samples were then sectioned directly onto functionalized coverslips at 7 μm using a cryostat. To compare 10x and 60x objective lenses ( Figure 18) Adjacent mouse sagittal brain sections were used. The sections were air-dried at room temperature for 5 minutes and then fixed in a 1×PBS solution containing 4% vol / vol paraformaldehyde for 15 minutes. After fixation, the samples were rinsed once with 1×PBS and immediately permeabilized in a 1×PBS solution containing 0.5% TritonX-100 at room temperature for 10 minutes, or permeabilized in 70% ethanol at 4 °C overnight, or stored at -80 °C. No sample size estimation was performed as the aim was to demonstrate the technique.

[0216] Preparation of human colorectal cancer tissue samples

[0217] As part of an ongoing study on colorectal cancer (CRC) approved by the SingHealth (2020 - 186) Institutional Review Board, sample collection was carried out according to ethical guidelines and patients provided written informed consent. To demonstrate the FISHnCHIPs technique, aliquots from tumor colon tissue that could not be individually identified (A*STAR IRB F-112) were used, which were collected immediately after resection, frozen on dry ice, and stored at -80 °C. Before sectioning, the tissue was embedded in optimal cutting temperature compound (Tissue-Tek O.C.T.; VWR, cat. no. 25608–930). Sections were obtained as described above and, after fixation, the samples were rinsed once with 1×PBS and then immediately permeabilized in 70% ethanol at 4 °C overnight. The sections were further permeabilized in a 1×PBS solution containing 0.5% TritonX-100 at room temperature for 15 minutes.

[0218] Sample staining

[0219] After permeabilization, the tissue samples were rinsed three times with 1×PBS and then rinsed with 2×SSC. The coding probes were diluted to a final concentration of 1-2 nM per probe in 20% or 30% hybridization buffer. 20% hybridization buffer consisted of 20% deionized formamide (AmbionTMCat: AM9342, AM9344) (vol / vol) in 2×SSC, 1 mg ml-1 yeast tRNA (Life Technologies, cat. no. 15401-011), and 10% dextran sulfate (Sigma, cat. no. D8906) (wt / vol). The samples were stained with the coding probes for 16 to 48 hours at 37 °C or 47 °C. After hybridization, the samples were washed twice in 20% formamide wash buffer containing 20% deionized formamide and 2×SSC, with each wash incubated at 37 °C or 47 °C for 15-30 minutes. Then the wash buffer was removed and the samples were washed twice with 2×SSC. The staining and washing conditions were optimized separately for each sample type. The samples were stained with DAPI (Sigma, cat. no. D9564) at a concentration of 1 μg / ml in 2×SSC for 10 minutes at room temperature. Then the samples were washed three times with 2×SSC and immediately imaged or stored in 2×SSC at 4 °C for less than 12 hours before imaging. For single molecule FISH of DCN, MMP2, TAGLN, ACTA2, and SPARC (Biosearch technologies), the probes were diluted in 10% hybridization buffer and the samples were stained overnight at 37 °C. Then the samples were washed twice with 10% formamide wash buffer, each wash for 15 minutes at 37 °C, then rinsed with 2×SSC, followed by imaging.

[0220] Imaging cycle

[0221] Use a flow chamber (Bioptechs, cat. no. FCS2) that can be fixed to the microscope stage to mount the sample. Readout probe hybridization is performed directly in the flow chamber by buffer exchange, which is controlled by a custom fluidics system controlled by a computer, as previously described by Chen, K. H. et al. (Chen, K. H., Boettiger, A. N., Moffitt, J. R., Wang, S. & Zhuang, X. Spatially resolved, highly multiplexed RNA profiling in single cells. Science 348, aaa6090 (2015)). All buffer solutions (about 1 ml per exchange) are flowed in within 1 minute. A 10 nM fluorescently labeled readout probe in 10% high-salt hybridization buffer is flowed into the flow chamber and incubated for 10 minutes at room temperature. The 10% high-salt hybridization buffer consists of 10% deionized formamide (vol / vol) and 10% dextran sulfate (Sigma, cat. no. D8906) (wt / vol) in 4×SSC. After hybridization, the sample is rinsed with 2×SSC and then flowed in 10% formamide wash buffer containing 0.1% Triton X-100. It is flowed with 2×SSC again before the imaging buffer. The imaging buffer consists of 2×SSC, 10% glucose, 50 mM Tris-HCl pH 8, 2 mM Trolox (Sigma, cat. no. 238813), 0.5 mg / ml glucose oxidase (Sigma, cat. no. G2133), and 40 μg / ml catalase (Sigma, cat. no. C30). To remove the fluorescent signal, the sample is washed with 55% formamide wash buffer containing 0.1% Triton X-100. This hybridization and washing cycle is repeated until all readout probes have been imaged.

[0222] Imaging device 1

[0223] Imaging is performed on the device described by Goh, J. J. L. et al. (ibid.). Briefly, the microscope is constructed around a Nikon Ti2-E body, a Marzhauser SCANplus IM 130 mm × 85 mm motorized X-Y stage, a Nikon CFI Plan Apo Lambda 60×1.4-n.a. oil-immersion objective, and an Andor Sona4.2B-11s CMOS camera. For the whole-slide imaging experiment ( Figure 6),The Nikon CFIPlan Apo 10×0.5-n.a. water immersion objective lens was used. The DAPI channel was excited by a Coherent Obis405 100-mW laser. MPB Communications fiber lasers were used as the illumination for Alexa594 (592 nm), Cy5 (647 nm), and IRDye 800CW (750 nm) respectively: 2RU-VFL-P-500-592-B1R (500 mW), 2RU-VFL-P-1000-647-B1R (1000 mW), and 2RU-VFL-P-500-750-B1R (500 mW). The Nikon Perfect Focus system was used to maintain focus during imaging, and in each imaging cycle, one Z position was imaged for each field of view. When imaging under the 10× water immersion objective lens, the Perfect Focus system was not used. Images were acquired at different exposure times (1 second, 500 milliseconds, and 1 second at 60× and 3 seconds, 3 seconds, and 5 seconds at 10× for Alexa594, Cy5, and IRDye 800CW respectively) to avoid camera saturation.

[0224] Imaging device 2

[0225] A custom microscope was constructed using a Nikon Ti2-E body, a Marzhauser SCANplus IM 130 mm × 85 mm motorized X-Y stage, and a pco.edge 4.2BI-USB Back Illuminated sCMOS camera. A custom fiber-coupled laser box from CNIlaser was used for illumination of DAPI (405 nm), Alexa Fluor 488 (488 nm), Alexa Fluor 594 (588 nm), Cy5 (637 nm), and IRDye 800CW (750 nm). Custom multi-wavelength filters (445 / 503 / 560 / 615 / 683 / 813 (Semrock) and 405 / 473 / 532 / 588 / 637 / 730 (Semrock)) were used. The following objective lenses were tested: Nikon CFIPlan Apo Lambda 10×0.45-n.a. air objective (MRD00105), Nikon CFIPlan Apo 10×0.5-n.a. water immersion objective (MRD71120), Nikon CFIPlan Fluor 20×0.75-n.a. water immersion objective (MRH07241), Nikon CFI S Plan Fluor ELWD 20×0.45-n.a. air objective (MRH08230), Nikon CFI Apo LWD Lambda S 40×1.15-n.a. water immersion objective (MRD77410), and Nikon CFIPlan Apo Lambda 60×1.4-n.a. oil immersion objective (MRD01605). At 40× and 60×, the Nikon Perfect Focus system was used to maintain focus. One Z position was imaged per field of view. The setup was used for objective lens comparison experiments and immunofluorescence imaging.

[0226] Immunofluorescence staining

[0227] Wash the tissues three times with 1×PBS at room temperature. Block for 1 hour with 1% BSA (NEB) and 0.1% Tween-20 in 1×PBS at room temperature. Stain the tissues overnight at 4°C with the following antibodies diluted in the blocking solution: anti-LUM (Abcam, ab168384; 1:75), anti-MMP2 (Abcam, ab37150; 1:200), anti-α-SMA (Abcam, ab7817; 1:600), and anti-PDGFA (Santa Cruz Biotechnology, sc-9974; 1:600). Detect PDPN using AF488-conjugated primary antibody (BioLegend, 337005; 1:75). Then perform secondary antibody staining for 1 hour at room temperature using anti-mouse AF594 (ThermoFisher, A11005; 1:1000) and anti-rabbit AF488 (ThermoFisher, A11008; 1:1000). Finally, stain the samples overnight at 4°C with anti-CD68 (Cell Signalling Technology, #79594; 1:50). After washing three times with 1×PBS, counterstain the tissues with DAPI (Sigma), and then mount (Vectashield, H-1700-10).

[0228] Image processing and data analysis

[0229] Create a custom pipeline( Figure 7) To align the images (DAPI image, FISHnCHIPs image, and background image), segment, and cluster cell types. First, nuclear masks are obtained by performing nuclear segmentation using the deep learning-based Cellpose algorithm (Stringer, C., Wang, T., Michaelos, M. & Pachitariu, M. Cellpose: a generalist algorithm for cellular segmentation. Nat Methods 18, 100–106 (2021)) or the watershed algorithm. The in situ hybridization images are registered to the DAPI image by phase correlation using the subpixel registration algorithm provided in the Scikit-Image package (van der Walt, S. et al. scikit-image: image processing in Python. PeerJ 2, e453 (2014)). Subsequently, after alignment (i.e., applying the same displacement), the background image is subtracted from the in situ hybridization images (images are taken after washing with 55% formamide, and these images are used to estimate the tissue autofluorescence background). The nuclear masks obtained from DAPI segmentation are enlarged to generate cell masks, and these cell masks are applied to all background-subtracted in situ hybridization images. An in situ hybridization intensity matrix is constructed for cell type clustering and subsequent analysis. After quality control and normalization, the Louvain algorithm is used to cluster the intensity matrix. The cell clusters are visually presented in heatmaps, dimensionality reduction plots, and clustering plots. The analysis pipeline can be downloaded as supplementary software.

[0230] Gain and crosstalk analysis of mouse kidneys

[0231] Nuclear segmentation and image alignment are performed as described above. Nuclear masks smaller than 3000 pixels are discarded. The nuclear masks are enlarged by 5 pixels to create cell masks. The images are normalized by dividing by the 99th percentile of the pixel intensity. A cell-channel intensity matrix is constructed by calculating the average fluorescence intensity of each cell using the cell masks. Since only five renal cell types are imaged in this experiment, cells with a normalized intensity lower than 0.5 are removed (only about 18.6% of the cells brightly labeled by the in situ hybridization method described herein are retained). Qualified cells with the highest normalized intensity in all channels are assigned to the corresponding cell types. As Figure 2As shown, the fluorescence signal gain of in situ hybridization was calculated by taking the ratio of the average FISHnCHIPs intensity to the average smFISH intensity in the same cells (the same cell mask was applied to the FISHnCHIPs image and the smFISH image because they were imaged sequentially on the same sample). The crosstalk of the in situ hybridization method was estimated by calculating the Manders overlap coefficient, which is a measure that quantifies the degree of colocalization of objects in a pair of images (originally developed for two-color confocal microscopy). It is the overlapping part between two channels:

[0232]

[0233] where t1 and t2 are the thresholds used to binarize the two channels C1 and C2, respectively.

[0234] Analysis of mouse cortex data of 18 modules

[0235] As Figure 5 shown, gene-centered in situ hybridization profiling was performed on 18 gene modules in the mouse cortex. Nucleus segmentation and image alignment were performed as described above. Nucleus masks smaller than 3000 pixels were discarded. The nucleus masks were enlarged by 15 pixels to create cell masks. The images were normalized to the 99th percentile of pixel intensity. A cell-module intensity matrix was constructed by taking the average intensity of the segmented cell masks. For quality control, cells with a total intensity below the 15th percentile were removed. The cell-module intensity matrix was used for clustering using the Seurat package. The modules were Z-scaled before calculating the principal components and dimensionality reduction projection. The Louvain clustering algorithm was used for clustering analysis. The cells were clustered at a resolution of 0.8 using the top 10 PCs with 20 nearest neighbors. Finally, the cell clusters were mapped back to the positions of the cell masks to reconstruct the spatial map.

[0236] Analysis of data of mouse cortical neuron subtypes

[0237] Nucleus segmentation and image alignment were performed as described above. Nucleus masks smaller than 3000 pixels were discarded. The nucleus masks were enlarged by 10 pixels to create cell masks. The images were normalized to the 99th percentile of pixel intensity. A cell-program intensity matrix was constructed by taking the average intensity of the cell masks. As Figure 9As shown, the image was cropped to include only the cortical region. For quality control, cells with a total intensity below the 20th percentile were removed. Clustering analysis was performed as described above, but at a higher resolution (1.2). Five of the 18 clusters (29.7% of the cells) contained cells with a weak or no neuronal expression signature, and these cells were then removed. Thus, 50.3% of all cells (defined by DAPI) were considered neurons. To quantify the cortical depth of neuronal cells, the edges of two circles with the same radius (R = 25500 pixels) were used to cover the region with excitatory neurons, as Figure 9 shown. The distance between the two centers was 10000 pixels. The normalized depth of a cell was defined as follows: the distance from the outer edge divided by the distance between the two centers. A cortical depth cell intensity heatmap ( Figure 11 ) was plotted by arranging the cells in increasing order of depth. The cell density along the cortical depth was estimated by applying kernel density estimation (KDE) with a 0.05 Gaussian kernel.

[0238] Analysis of large field-of-view mouse brain data for 53 modules

[0239] To generate Figure 13For the cell-module intensity matrix and cell positions, the nuclear images were normalized to the 99th percentile of pixel intensity and the same nuclear segmentation pipeline as described above was utilized. Each in situ hybridization image was registered to its corresponding DAPI image and the displacements were recorded. Displacements exceeding 50 pixels in any direction were discarded. The average displacement was then applied to all fields of view. To correct for illumination variations between fields of view, the 60th percentile of pixel intensity outside the cell mask was subtracted. Cells with low intensity (<0.2%) in all modules or high intensity (>98%) in all over 30 modules were removed. A cell graph based on the 15 nearest neighbors using the first 20 PCs was initially constructed. Leiden clustering was performed at a resolution of 2. 133 cells (0.25%) from 2 of the preliminary clusters were affected by autofluorescence of dust particles in the sample and were discarded from further analysis. 54,834 (97.3%) eligible cells were clustered at a lower resolution (0.6), resulting in 18 clusters or cell types. The vascular-related cell cluster and the inhibitory neuron cluster showed a finer structure in the UMAP and were further subclustered. To validate the cluster annotations, an integration analysis was performed using the Harmony algorithm between the in situ hybridization method described and scRNA-seq (Korsunsky, I. et al. Fast, sensitive and accurate integration of single-cell data with Harmony. Nat Methods 16, 1289–1296 (2019)) ( Figure 16 ). To ensure compatibility, the in situ hybridization data were cropped to the frontal cortex region. Additionally, as recommended by the Harmony authors, the scRNA-seq data were randomly subsampled to balance the cell numbers. Before integration, the scRNA-seq and in situ hybridization data were normalized and scaled. One of the clusters (2,773 cells or 5% of the cells) could not be annotated because they showed low-level expression and were spatially heterogeneous in all neuronal and non-neuronal modules. Based on the integration analysis, these cells were observed to be very close to the oligodendrocyte and excitatory neuron clusters. Based on this observation, the "unknown" cluster may be one or more true cell populations that cannot be resolved by the current probe set.

[0240] Proximity of cancer-associated fibroblasts (CAFs) to immune cells in human colorectal cancer (CRC) tissues

[0241] The fibroblasts and immune cells were segmented using the watershed segmentation algorithm provided in the Scikit-image package. For each cell type, the truncation threshold and opening threshold used for watershed segmentation were manually adjusted. Using the centroid of the segmented cell masks, the number of immune cells within a 100-μm radius of CAF-1 or CAF-2 cells was calculated. As Figure 19 shown, the number of immune cells closer to CAF-1 cells was significantly higher compared to CAF-2 cells (two-sided Mann-Whitney U test). This result was consistent with the visual inspection of cell positions ( Figure 19 and Figure 21 ).

[0242] Summary

[0243] In summary, the present invention demonstrates that the in situ hybridization method described herein can be used for robust imaging and characterization of cells in biological tissue samples with high sensitivity and high throughput, while reducing the requirements and costs of experimental instruments.

Claims

1. A method for in-situ characterization of cells in a biological sample, comprising: a. contacting the biological sample with a plurality of probes that bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes, wherein each probe comprises i) a detectable label, and ii) a domain that specifically binds to the ribonucleic acid transcript of one of the predetermined genes; wherein a signal is emitted when the probe binds to the ribonucleic acid transcript; b. detecting a combination of emission signals from the plurality of probes or a plurality of emission signals; and c. characterizing the cells based on the combination of emission signals or the plurality of emission signals.

2. The method according to claim 1, wherein steps a and b are repeated one or more times using a plurality of probes that bind to RNA transcripts of a plurality of different predetermined genes.

3. The method according to claim 1 or 2 further comprises the following steps: Before characterizing the cells, quantify the level of the emission signals detected in step b, process the signals, or perform both operations simultaneously.

4. The method according to claim 1, wherein the plurality of predetermined genes comprise at least one gene and at least one other gene, both of which exhibit co-variation in their expression levels, and both of which are: a) markers of a specific cell type; b) differentially expressed genes of a specific cell type; c) markers of a gene expression program or a gene regulatory module; d) markers of a biological pathway; or a combination thereof; wherein the at least one other gene is selected from one or more input data sets.

5. The method according to claim 4, wherein the input data set is bulk RNA sequencing, single-cell RNA sequencing, microarray data set, chromatin accessibility sequencing, methylation sequencing, DNA-associated protein sequencing, spatial transcriptomics sequencing, multiplexed RNA fluorescence in-situ hybridization, multiplexed immunohistochemistry, bioinformatics database, any user-defined data set, or a combination thereof.

6. The method according to any one of claims 1 to 5, wherein the selection of the plurality of predetermined genes is unsupervised selection, supervised selection, or a combination thereof.

7. The method according to any one of claims 4 to 6, wherein the co-variation in the expression levels of the at least one gene and the at least one other gene is determined by correlation analysis, clustering analysis, dimensionality reduction analysis, differentially expressed gene analysis, or a combination thereof of the input data set.

8. The method according to any one of claims 4 to 7, wherein genes that exhibit co-variation in their expression levels are further analyzed using signal gain (SG) or signal specificity ratio (SSR) to identify the plurality of predetermined genes.

9. The method according to any one of claims 1 to 8, wherein the domain of the probe is a ribonucleic acid (RNA) oligonucleotide that specifically binds to RNA.

10. The method according to any one of claims 1 to 9, wherein the biological sample comprises a homogeneous or heterogeneous cell population.

11. The method according to any one of claims 1 to 10, wherein the characterization of the cell comprises one or more of the following: mapping the position of the cell in the biological sample; identifying the interaction between the cell and one or more other cells; identifying the gene expression pattern of the cell or the biological sample and visually presenting the spatial transcriptome of the cell or the biological sample; stratifying cancer subtypes to determine the severity of the cancer.

12. The method according to any one of claims 1 to 11, further comprising: Preprocessing the input data set before performing correlation analysis, clustering analysis, dimensionality reduction analysis or differential expression gene analysis.

13. The method according to any one of claims 1 to 12, wherein the plurality of predetermined genes are expressed in cells associated with cancer.

14. A method for determining the prognosis of a subject with cancer, comprising: a. Obtaining a sample of the subject; b. Characterizing one or more cancer cells in the sample using the method according to any one of claims 1 to 13 to determine the stage of the cancer; and c. Determining the prognosis based on the stage of the cancer.

15. A kit for in-situ characterization of cells in a biological sample, comprising: A plurality of probes that bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes, wherein each probe comprises: i) A detectable label, and ii) A domain that specifically binds to the ribonucleic acid transcript of one of the predetermined genes; and instructions for use.

16. The kit according to claim 15, wherein the plurality of probes bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes, wherein the plurality of predetermined genes comprise at least one gene and at least one other gene, and the at least one gene and the at least one other gene exhibit co-variation in their expression levels, and both are: e) Markers of specific cell types; f) Differentially expressed genes of specific cell types; g) Markers of gene expression programs or gene regulatory modules; h) Markers of biological pathways; or a combination thereof; wherein the at least one other gene is selected from one or more input data sets.

17. The kit according to claim 15 or 16, wherein the plurality of predetermined genes are expressed in the kidney, brain, cancer, or a combination thereof.

18. A kit for in-situ characterization of colorectal cancer in a biological sample, comprising: A plurality of probes that bind to ribonucleic acid (RNA) transcripts of a plurality of predetermined genes, wherein the plurality of predetermined genes are selected from the genes listed in Tables 6(6a) to Table (6d); wherein each probe comprises: i) A detectable label, and ii) A domain that specifically binds to the ribonucleic acid transcripts of the plurality of predetermined genes; and instructions for use.