Co-mapping of transcriptional state and protein organization
Patent Information
- Application Number
- JP2023573328
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-28
- Filing Date
- 2022-05-27
- Publication Date
- 2025-06-03
AI Technical Summary
Existing methods for studying Alzheimer's disease (AD) are limited by their inability to integrate spatially resolved single-cell transcriptomics with protein detection in the same tissue section, masking cellular heterogeneity and failing to preserve spatial patterns, and they cannot effectively map gene expression and protein histology in the same tissue sample.
The development of STARmap Pro, a method that allows high-resolution spatial transcriptomics with simultaneous localization of specific proteins within the same tissue section, enabling the mapping of gene and protein expression at subcellular resolution using oligonucleotide probes and detection agents, such as antibodies, to identify disease-associated microglia and astrocyte populations in AD models.
Enables comprehensive molecular atlases of AD pathophysiology across multiple cell types, providing insights into disease progression and facilitating drug screening and treatment strategies by integrating spatial and temporal gene and protein expression patterns.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Related Applications This application claims priority under 35 USC § 119(e) to U.S. Provisional Application No. USSN 63 / 194,536, filed May 28, 2021, which is incorporated herein by reference. [Background technology]
[0002] 2. Background of the Invention Alzheimer's disease (AD) is a progressive neurodegenerative disease and the most common form of dementia in older adults (Masters et al., 2015). Extensive deposition of amyloid-β (Aβ) plaques and neurofibrillary tangles (hyperphosphorylated tau deposits), especially in the neocortex and hippocampus, is the neuropathological hallmark of AD (Braak and Braak, 1991; Hardy and Selkoe, 2002; Masters et al., 2015). In addition, AD pathology is also characterized by gliosis (reactive changes in microglia and astrocytes) and white matter abnormalities (Beach et al., 1989; Henstrindge et al., 2019; Butt et al., 2019). A key question in AD research is how morphological features correlate with cellular gene pathways that drive neurodegeneration. Genome-wide association studies (GWAS) have revealed genes associated with AD risk and contributed to elucidating the mechanisms of AD pathology, and it has been shown that the majority of AD risk genes are highly expressed in microglia (Pimenova et al., 2018; Cauwenberghe et al. 2016). Multiple bulk and scRNA-seq studies from AD mouse models and other neurodegenerative models have discovered a population of microglia with a unique transcriptional state, called DAM (Disease-associated microglia) (Bohlen et al., 2019; Hansen et al., 2018). In addition to DAM, astrocyte populations associated with AD pathology have also been characterized. Established analytical methods are at a disadvantage in revealing the molecular and cellular complexity of AD: bulk tissue analysis hides the heterogeneity of cell populations in the brain, and standard imaging methods can only visualize a few genes and proteins and identify only limited cell types. Recent application of single-cell RNA sequencing (scRNA-seq) to AD brain tissue has revealed substantial heterogeneous changes in gene expression in major brain cell types ( Grubman et al., 2019 ; Keren-Shou et al., 2017 ; Mathys et al., 2019 ).However, while scRNA-seq studies provide single-cell resolution, they fail to preserve spatial patterns. Also, single-cell preparations of any cell type cannot be easily isolated from the brain in an unbiased manner. To truly understand the extent and heterogeneity of diverse cellular responses to amyloid plaques, tau aggregation, cell death, and synapse loss, and to investigate the spatial relationships between the localized lesions and cellular responses mentioned above, fundamentally different technological platforms are required. Thus, methods that integrate spatially resolved single-cell transcriptomics and tissue histology are highly desired in AD research and will be useful for many other applications.
[0003] Many existing spatially resolved transcriptome techniques (e.g., spatial transcriptomics, STARmap, etc.) are not compatible with protein detection in the same tissue section (Stahl et al., 2016; Stuart and Satija, 2019; Wang et al., 2018). Plaque-inducible genes (PIGs) have been revealed using spatial transcriptomics by fluorescent staining of adjacent brain sections (Stahl et al., 2016; Stuart and Satija, 2019; Wang et al., 2018). However, the resolution is limited and only a small set of genes has been validated at cellular resolution. Furthermore, due to the relative thickness of each section, the adjacent section strategy is less precise and cannot be used to investigate the effect of tau tangles on gene expression in the same cells. Therefore, new methods are needed to map gene expression and protein histology in the same tissue sample, and such methods would be useful for Alzheimer's disease research and treatment. DISCLOSURE OF THEINVENTION
[0004] Summary of the Invention This disclosure describes a method for profiling gene and protein expression in the same cell. In particular, the development of a method / system, referred to herein as "STARmap Pro", is described in this disclosure. STARmap Pro allows high-resolution spatial transcriptomics to be performed simultaneously with the localization of specific proteins in the same tissue section. This method / system is useful for understanding, for example, the pathophysiology of AD across multiple cell types at subcellular resolution with a comprehensive molecular atlas (Figure 1A). This method / system is useful for studying, diagnosing, or treating any disease that involves alterations in gene and / or protein expression, such as cancer, and also for studying tissue development. STARmap Pro can also be used to characterize gene and protein expression associated with any disease, study the development of normal tissue, or study the effect of drugs on tissue (including screening drugs that have specific effects on tissue). For example, the present disclosure describes the use of an established mouse model of AD (TauPS2APP triple transgenic mouse) with both amyloid plaques and tauopathy that expresses mutant forms of hPresenilin 2 (PS2), hAPP, and hTAU and exhibits age-related brain amyloid deposition, tauopathy, gliosis, and cognitive impairment (Grueninger et al., 2010; Lee et al., 2021). By mapping a target list of 2,766 genes extracted from previous bulk and single-cell RNA-seq references and diverse AD-related databases, a spatial cell atlas of 8-month-old and 13-month-old TauPS2APP mice was created at subcellular resolution in the context of extracellular Aβ plaques and intracellular phosphorylated tau accumulation. Single-cell resolved transcriptome analysis identified disease-related gene pathways across various cell types in the cortical and hippocampal regions of the TauPS2APP model compared to control samples. By synthesizing spatial maps of diverse cell types and states at various disease stages, a comprehensive spatiotemporal model of AD disease progression is established and described herein.
[0005] In one aspect, the present disclosure provides methods and systems for mapping gene and protein expression (of one or more genes and proteins) within the same cell (i.e., at single-cell resolution, see, e.g., Figures 1A and 1B). Gene and protein expression can also be mapped in multiple cells at once, e.g., multiple cells present in a tissue sample. In the methods disclosed herein, the cells can be contacted with one or more pairs of oligonucleotide probes to amplify the nucleic acid of interest and generate one or more ligated amplicons. The cells can then be contacted with one or more detection agents (e.g., antibodies, or antibody fragments or variants), where each detection agent binds to a protein of interest. The one or more ligated amplicons and the one or more detection agents can then be embedded in a polymer matrix, and the one or more amplicons can be sequenced to determine the identity of the transcripts and their location within the polymer matrix. The location of the detection agent bound to one or more proteins of interest within the polymer matrix can also be determined by imaging (e.g., by confocal microscopy), allowing the location of transcripts and proteins of interest within a cell to be simultaneously mapped within the same sample (e.g., the same cell). The locations of the transcripts and proteins of interest can be used to identify individual cells, subcellular locations, and organelles based on the location and expression patterns of specific genes and proteins. This method can be useful, for example, to compare a cell (or multiple cells) from a diseased tissue sample and a healthy tissue sample. This method can also be useful in drug discovery (e.g., screening candidate agents with specific effects), studying the side effects of drugs, and diagnosing and treating diseases (e.g., Alzheimer's disease and cancer).
[0006] In some embodiments, the present disclosure provides a method for mapping gene and protein expression in a cell, the method comprising the steps of: a) contacting a cell with one or more pairs of oligonucleotide probes, where each pair of oligonucleotide probes comprises a first oligonucleotide probe (also referred to herein as a "padlock" probe) and a second oligonucleotide probe (also referred to herein as a "primer" probe), i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, where the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe; b) ligating together the 5' and 3' ends of a first oligonucleotide probe to generate a circular oligonucleotide; c) performing rolling circle amplification to amplify the circular oligonucleotide using the second oligonucleotide probe as a primer to generate one or more ligated amplicons; d) contacting the cells with one or more detection agents, where each detection agent binds to a protein of interest; e) embedding the one or more ligated amplicons and the one or more detection agents in a polymer matrix; f) contacting the one or more ligated amplicons embedded in the polymer matrix with a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe; and g) Imaging the one or more ligated amplicons embedded in the polymer matrix and the one or more detection agents embedded in the polymer matrix to determine the location of the nucleic acid of interest and the protein of interest within the cell, and optionally, to map gene and protein expression.
[0007] The methods and systems described herein may be useful for studying gene and protein expression in tissues (e.g., developing tissues), diagnosing and treating various diseases, and drug discovery. Thus, in another aspect, the disclosure provides a method for diagnosing a disease or disorder (e.g., Alzheimer's disease) in a subject. For example, the methods for profiling gene and protein expression described herein can be performed in cells from a sample taken from a subject (e.g., a subject believed to have or at risk of having a disease or disorder, or a healthy subject, or a subject believed to be healthy). Expression of various nucleic acids and proteins of interest in the cells can then be compared to expression of the same nucleic acids and proteins of interest in non-disease cells or cells from a non-disease tissue sample (e.g., a cell from a healthy individual, or a plurality of cells from a population of healthy individuals). Any change in expression of the nucleic acid of interest compared to expression in non-disease cells may indicate that the subject has a disease or disorder. Gene and protein expression in one or more non-disease cells can be profiled in parallel with expression in diseased cells as a control experiment. Gene and protein expression in one or more non-diseased cells may have been previously profiled, and expression in the diseased cells can be compared to this reference data for non-diseased cells.
[0008] In another aspect, the present disclosure provides a method for screening agents that can modulate gene and / or protein expression of a nucleic acid or protein of interest, or multiple nucleic acids and / or proteins of interest. For example, the methods for mapping gene and protein expression described herein can be carried out in cells in the presence of one or more candidate agents. The expression of various nucleic acids and / or proteins of interest in cells (e.g., normal cells or diseased cells) can then be compared with the expression of the same nucleic acids and / or proteins of interest in cells that have not been exposed to one or more candidate agents. Any change in the expression of the nucleic acid(s) and / or protein(s) of interest compared to the expression in cells that have not been exposed to the candidate agent(s) can indicate that the expression of the nucleic acid(s) and / or protein(s) of interest is modulated by the candidate agent(s).
[0009] In another aspect, the disclosure provides a method for treating a disease or disorder (e.g., Alzheimer's disease) in a subject. For example, the methods for profiling gene and protein expression described herein can be performed in a cell (or a plurality of cells, e.g., those making up a tissue) from a sample taken from a subject (e.g., a subject suspected of having or at risk of having a disease or disorder). Expression of various nucleic acids and / or proteins of interest in the cells can then be compared to expression of the same nucleic acids and / or proteins of interest in cells from a non-disease tissue sample. Treatment of the disease or disorder can then be administered to the subject if any changes in expression of the nucleic acids and / or proteins of interest are observed compared to expression in non-disease cells. Gene and protein expression in one or more non-disease cells can be profiled in parallel with expression in disease cells as a control experiment. Gene and protein expression in one or more non-disease cells may have been profiled previously, and expression in disease cells (or test cells suspected to be disease cells) can be compared to this reference data for non-disease cells.
[0010] In another aspect, the present disclosure provides a plurality of oligonucleotide probes including a first oligonucleotide probe (also referred to herein as a "padlock" probe) and a second oligonucleotide probe (also referred to herein as a "primer" probe), wherein: i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence; Here, the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe.
[0011] In another aspect, the disclosure provides kits (e.g., kits including any of the oligonucleotide probes disclosed herein). In some embodiments, the kits include a library of oligonucleotide probes described herein, each of which can be used to identify a particular nucleic acid of interest. In some embodiments, the kits further include a detection agent, or a library of detection agents, for detecting various proteins of interest. The kits described herein can also include any other reagents or components useful for carrying out the methods described herein, including, but not limited to, cells, ligases, polymerases, amine-modified nucleotides, primary antibodies, secondary antibodies, buffers, and / or reagents for making a polymer matrix (e.g., a polyacrylamide matrix).
[0012] Another aspect of the present disclosure provides a method for identifying spatial variation of cell types within at least one image (i.e., examining the variation in the relative spatial distribution of a particular cell type between multiple samples, e.g., comparing healthy and diseased tissue). In some embodiments, such a method comprises the following steps: receiving, for each of a plurality of cells in the at least one image, a spatial location of the cell in the at least one image; receiving, for each of a plurality of proteins in the at least one image, a spatial location of the protein within the image; determining, for a first protein of the plurality of proteins, a number of cells of the first cell type that are less than a threshold distance to the first protein, where the distance is determined based on spatial locations of at least some of the plurality of cells and spatial locations of at least some of the plurality of proteins; identifying a spatial variation of cells of the first cell type in the at least one image based on the number of cells of the first cell type; and Outputting a representation of the spatial variability of cells of the first cell type in at least one image. Such methods are useful, for example, for identifying cell types that are within a certain distance from a protein of interest that may be associated with disease. For example, as further described herein, the presence of a particular cell type at a certain distance from Aβ and / or tau inclusions may be associated with Alzheimer's disease.
[0013] In another aspect, the disclosure provides an apparatus comprising: at least one computer processor; and at least one non-transitory computer-readable storage medium encoded with a plurality of instructions that, when executed by the at least one computer processor, performs a method for identifying spatial variation of cell types in at least one image, the method comprising: receiving, for each of a plurality of cells in the at least one image, a spatial location of the cell in the at least one image; receiving, for each of a plurality of proteins in the at least one image, a spatial location of the protein within the image; determining, for a first protein of the plurality of proteins, a number of cells of the first cell type that are less than a threshold distance to the first protein, where the distance is determined based on spatial locations of at least some of the plurality of cells and spatial locations of at least some of the plurality of proteins; identifying a spatial variation of cells of the first cell type in the at least one image based on the number of cells of the first cell type; and outputting a representation of the spatial variance of cells of the first cell type within the at least one image; The apparatus includes:
[0014] In another aspect, the present disclosure provides at least one non-transitory computer readable storage medium encoded with a plurality of instructions that, when executed by at least one computer processor, performs a method for identifying spatial variation of cell types in at least one image, the method comprising: receiving, for each of a plurality of cells in the at least one image, a spatial location of the cell in the at least one image; receiving, for each of a plurality of proteins in the at least one image, a spatial location of the protein within the image; determining, for a first protein of the plurality of proteins, a number of cells of the first cell type that are less than a threshold distance to the first protein, where the distance is determined based on spatial locations of at least some of the plurality of cells and spatial locations of at least some of the plurality of proteins; identifying a spatial variation of cells of the first cell type in the at least one image based on the number of cells of the first cell type; and outputting a representation of the spatial variance of cells of the first cell type within the at least one image; The at least one non-transitory computer-readable storage medium includes:
[0015] It should be understood that the foregoing concepts, and additional concepts described below, can be arranged in any suitable combination, as the present disclosure is not limited thereto. Moreover, other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non-limiting embodiments, when considered in conjunction with the accompanying drawings. [Brief description of the drawings]
[0016] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the disclosure, which may be better understood by reference to one or more of these drawings in combination with the detailed description of specific aspects presented herein.
[0017] [Fig. 1A-1F]Figure 1A-1F show co-mapping of single-cell transcriptional states with amyloid-β and tau pathology in Alzheimer's disease using STARmap Pro. Simultaneous mapping of cell types, single-cell transcriptional states, and histopathology at 200 nm resolution is displayed. Figure 1A shows an overview of the STARmap Pro method; it is an integrated in situ method that can simultaneously map thousands of RNA species and protein disease markers at subcellular resolution (200 nm) within the same intact three-dimensional (3D) tissue. STARmap Pro was applied to characterize TauPS2APP transgenic mice (a mouse disease model of Alzheimer's disease (AD) pathology) with amyloid-β and tau pathology. An integrated analysis of single-cell transcriptional states and histopathology (spatial cell typing, pseudo-temporal trajectories, differential gene expression, and gene pathway analysis) was used to reveal key cell types, cell states, and gene programs associated with Alzheimer's disease progression. Figure 1B shows a schematic flow chart of the STARmap Pro method. Brain tissues were dissected and fixed, intracellular mRNAs were recognized by a pair of SNAIL (specific amplification of nucleic acids by intramolecular ligation) probes, and amine-modified cDNA amplicons were synthesized in situ by enzymatic ligation and rolling circle amplification. After labeling of protein targets with primary antibodies, the tissues bearing amine-modified cDNA amplicons, proteins, and primary antibodies were functionalized with acrylic acid N-hydroxysuccinimide ester (AA-NHS) and copolymerized with acrylamide to generate hydrogel-tissue hybrids that fixed the positions of biomolecules (e.g., amplicons, proteins, and antibodies) for in situ mapping. Each cDNA amplicon contains a gene-specific identifier sequence (labeled at the top of FIG. 1B), which is read out by in situ SEDAL (sequencing with error reduction by dynamic annealing and ligation) and then subjected to fluorescent protein staining (with secondary antibodies and the small molecule dye X-34) to visualize the protein signal.Figure 1C shows the expanded coding capacity of STARmap Pro. The SNAIL probe of STARmap Pro contains two 5-nt barcodes (labeled "Barcode A" and "Barcode B") with a theoretical coding capacity of 1 million (4^10). Figure 1D provides imaging results of cell nuclei, cDNA amplicons, and protein signals in a 13-month-old TauPS2APP mouse brain. The enlarged image shows p-Tau positive cells in the CA1 region of the hippocampus. Propidium iodide (PI) staining of cell nuclei, fluorescent DNA probe staining of all cDNA amplicons, X-34 staining of amyloid-β plaques, and immunofluorescent staining of p-Tau (AT8 primary antibody followed by fluorescent goat anti-mouse secondary antibody) are shown according to the provided legend. Figure 1E provides representative imaging results showing simultaneous mapping of cell nuclei, cDNA amplicons, and protein signals in a 13-month-old TauPS2APP mouse brain slice. 3D projection of raw confocal fluorescence images shows p-Tau positive cells in the CA1 region of the hippocampus (left). A close-up of the dashed area in the left panel is also provided (center), which shows the last cycle of histopathology imaging detecting both protein and cDNA amplicons. Propidium iodide (PI) staining of cell nuclei, fluorescent DNA probe staining of all cDNA amplicons, X-34 staining of amyloid-β plaques, and immunofluorescence staining of p-Tau (AT8 primary antibody followed by fluorescent goat anti-mouse secondary antibody) are shown according to the provided legend. Eight cycles of in situ RNA sequencing of the diagram in the middle panel are also shown (right). Fluorescence channels, each representing one round of in situ sequencing, are shown. Figure 1F shows an exemplary synthesis scheme for preparing DNA-tagged antibodies.
[0018] [Figure 2A-2L]Figures 2A-2L show top-level cell type classification and spatial analysis in brain slices of TauPS2AAPP and control mice. Figure 2A provides a Uniform Manifold Approximation (UMAP) plot visualizing the nonlinear dimensionality compression of the transcriptome profiles of 33,106 cells from four samples. The Leiden algorithm was used to identify well-connected cells in the low-dimensional representation of the transcriptome profiles as clusters. Thirteen cell types were defined by gene markers enriched in each cluster. As visualized in the UMAP of individual samples, the clusters highlighted by dashed rectangles (astrocytes, microglia, oligodendrocytes, and dentate gyrus) showed differential distribution of cell populations between TauPS2APP and control samples. Figure 2B shows the hierarchical classification of cell types. The plots provided show that both of the following were identified during the cell type classification process: 13 top-level clusters (cortical excitatory neurons (CTX-Ex, 8,687 cells), inhibitory neurons (In, 2,005 cells), CA1 excitatory neurons (CA1-Ex, 2,754 cells), CA2 excitatory neurons (CA2-Ex, 436 cells), CA3 excitatory neurons (CA3-Ex, 1,878 cells), dentate gyrus (DG, 4,377 cells), astrocytes (Astro, 2,884 cells), endothelial cells (Endo, 1,849 cells), microglia (Micro, 1,723 cells), oligodendrocytes (Oligo, 4,966 cells), oligodendrocyte progenitor cells (OPC, 549 cells), smooth muscle cells (SMC, 877 cells), lateral habenula neurons (LHb, 121 cells)), and 24 sub-level clusters. Gene expression profiles for each top-level cluster of interest were obtained and used again for sub-level clustering in the same manner. Figure 2C shows a cell atlas of the cortical and hippocampal regions of a TauPS2APP 13-month sample with Aβ and tau pathology. Top-level cell types and pathological signals are shaded as indicated in the legend in the upper right corner, and Aβ plaques and p-Tau protein are colored black and gray, respectively, as indicated in the provided legend.Imaging sections are manually separated into cortex and subcortex, and the borders are marked with black dashed lines. Scale bar, 100 μm. Enlarged sections: (I) is an enlarged section of the cortical region, with Aβ plaques shaded black in the center; (II-III) are enlarged sections of the subcortical region, with p-Tau protein shaded gray on the cells; scale bar, 10 μm. Figure 2D provides a schematic plot showing the strategy used to analyze the cell type composition around Aβ plaques. Considering the size of each plaque, five concentric circles of different radii (10, 20, 30, 40, and 50 μm) were generated, and the cell type composition was quantified at different distance intervals from the edge of each plaque. If cells resided at the border between two concentric circles, they were merged to prevent repeated counts. Scale bar, 50 μm. Figure 2E shows the cell type composition around Aβ plaques at different distance intervals for a 13-month sample of TauPS2APP. Stacked bar plots are provided showing the percentage of each top cell type in each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 µm) around the Aβ plaque. For each region, the overall (averaged) cell type composition is included as a reference. Figure 2F provides a Uniform Manifold Approximation (UMAP) plot visualizing the nonlinear dimensionality compression of the transcriptome profiles of 72,165 cells collected from coronal brain sections of 8- and 13-month-old TauPS2APP and control mice. The Leiden algorithm was used to identify well-connected cells in the low-dimensional representation of the transcriptome profiles as clusters. Thirteen major cell types were identified by gene markers enriched in each cluster. As visualized in the UMAP of individual samples, the clusters highlighted by dashed rectangles (astrocytes, microglia, oligodendrocytes, and dentate gyrus) showed differential distribution of cell type populations between TauPS2APP and control samples in the UMAP. FIG. 2G shows a hierarchical classification of cell types.The plots provided show that both of the following were identified based on representative gene markers during the cell type classification process: 13 top clusters (cortical excitatory neurons (CTX-Ex, 18,483 cells), inhibitory neurons (Inh, 4,163 cells), CA1 excitatory neurons (CA1-Ex, 3,225 cells), CA2 excitatory neurons (CA2-Ex, 1,439 cells), CA3 excitatory neurons (CA3-Ex, 3,22 5 cells), dentate gyrus (DG, 9,562 cells), astrocytes (Astro, 6,789 cells), endothelial cells (Endo, 4,168 cells), microglia (Micro, 3,732 cells), oligodendrocytes (Oligo, 11,265 cells), oligodendrocyte progenitor cells (OPC, 1,269 cells), smooth muscle cells (SMC, 2,397 cells), lateral habenula neurons (LHb, 204 cells), and 27 sub-level clusters. Gene expression profiles of each top-level cluster of interest were analyzed using Leiden clustering again, and sub-level clusters were identified in the same way. Figure 2H shows a representative spatial cell type atlas of cortical and hippocampal regions of a 13-month sample of TauPS2APP with Aβ and tau pathology. Top cell types and pathological signals are shaded as indicated in the legend in the upper right corner, and Aβ plaques and p-Tau protein are colored in black and gray, respectively, as indicated in the provided legend. Imaged sections are manually separated into cortex, corpus callosum (CC), and hippocampus, and the boundaries are marked with black dashed lines. Scale bar, 100 μm. Enlarged cross sections: (I) is an enlarged cross section of the cortical region, with Aβ plaques shaded black in the center and surrounded by various types of cells; (II) is an enlarged cross section of the hippocampal region, with p-Tau protein shaded gray on the cells; Scale bar, 10 μm. Figure 2I provides a schematic plot showing the strategy used to analyze the cell type composition surrounding Aβ plaques. Considering the size of each plaque, five concentric borders (10, 20, 30, 40, and 50 μm away from each plaque) were generated to quantify the cell type composition at different distance intervals from the edge of each plaque. If cells were present at the border between two stripes, they were merged to prevent repeated counts. Scale bar, 50 μm.Figure 2J shows a representative spatial distribution of cell type composition around Aβ plaques at different distance intervals for a 13-month TauPS2APP sample. Stacked bar plots are provided showing the density (cells per mm2) of each top cell type at each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 μm) around Aβ plaques. Each region includes the cell density of each major cell type as a reference for comparison. Figure 2K provides a schematic showing the method used for p-Tau signal quantification. Tissue sections were divided into a 20 μm × 20 μm grid. Shading represents the integrated intensity of p-Tau in each square and was used as an index to analyze the degree of colocalization of p-Tau with different cell types. Figure 2L shows the cell type composition analysis based on the 20 μm × 20 μm grid in a 13-month TauPS2APP sample, ranked by p-Tau density. Blocks divided by grid lines were ranked by the percentage of tau positive pixels and grouped into three bins: 0% (zero p-Tau), 1-50% (low p-Tau), and 51-100% (high p-Tau). The high p-Tau group was further divided into plaque positive and negative groups to further analyze the impact of plaques and tauopathy on cell type distribution. Stacked bar plots show the average number of cells per block for each major cell type.
[0019] [Fig. 3A-3S]Figures 3A-3S demonstrate the unique transcriptomic response and spatial composition of microglial populations under Tau+ plaque pathology. Spatiotemporal gene expression analysis of microglia in TauPS2APP and control samples is shown. Figure 3A provides the UMAP of microglial cell populations with subcluster annotation. The plot shows a low-dimensional representation of the transcriptomic profile of 1723 microglial cells identified from Figure 2A. Three sublevel clusters (Micro1 (n=779), Micro2 (n=415), and Micro3 (n=529)) were identified by the Leiden algorithm. The Micro3 cluster was annotated as a disease-associated microglia (DAM) population by its genetic markers in accordance with previous reports and significant enrichment in disease samples. UMAP plots of the sublevel cell clusters for each sample (control 8 months, TauPS2APP 8 months, control 13 months, TauPS2APP 13 months) were also included. Figure 3B shows the expression levels of representative markers across different subclusters (Z-scores per row). Figure 3C provides the UMAP of the microglial cell population with pseudotime trajectories. A plot showing a low-dimensional representation of the transcriptome profile of microglial cells generated by Monocle 3 is provided. The color map represents the pseudotime values. Corresponding trajectories are constructed and plotted on the low-dimensional embedding to show the biological progression of the cell population. Plots for each sample are also included. The trajectory start anchor was manually selected (based on the 8-month control sample). Figure 3D shows the pseudotime embedding of the microglial cell population with subcluster annotations. A plot showing the low-dimensional embedding used in the pseudotime calculation is provided with annotations of the microglial subclusters. Figure 3E shows the pseudotime distribution. The distribution of pseudotimes of the three microglial subpopulations and the distribution of microglial pseudotimes in the four samples are shown. Mann-Whitney-Wilcoxon test, ns: p>0.05, *p≦0.05, **p≦0.01, ***p≦0.001, ****p≦0.0001. Figure 3F provides a spatial map of microglial populations in a control 13-month sample. Scale bar, 100 μm.Figure 3G provides a spatial map of the microglial population of TauPS2APP 13-month samples. Scale bar, 100 μm. Two enlarged sections (I, II) are shown in the upper right corner. Scale bar, 10 μm. Figure 3H shows the cell density and cell type composition by region around the plaque. Stacked bar graphs showing the density (number per mm2) of each microglial subcluster in both the cortical and subcortical regions of each sample are provided (top). The shading of the bar graph corresponds to the cell type legend in Figure 3F. The plot shows a significant enrichment of DAMs in both 8-month and 13-month TauPS2APP samples. Stacked bar graphs showing the percentage of each microglial subcluster in each distance interval (0–10, 10–20, 20–30, 30–40, 40–50 μm) around the Aβ plaque are also provided (bottom). For each region, the overall cell type composition was included as a reference. Figure 3I shows the results of gene set enrichment analysis (GSEA) of differentially expressed genes (DEGs). Statistically significant (nominal p-value < 0.01) enrichment scores are shaded by the sign. Terms are filtered by term size 20-1000. Figure 3J provides a matrix plot showing the validated subset of DEGs in microglial populations from AD vs. controls comparison. A plot is provided showing row-wise scaled expression values of the top significantly changed DEGs (ranked by p-value) in the 2766 gene dataset and the 64 gene validation results. Figure 3K shows the subclustering of microglial cell populations. The plot shows a low-dimensional representation of the transcriptome profile of 3732 microglial cells identified from Figure 2A (top). Three sublevel clusters (Micro1 (n = 1924), Micro2 (n = 784), and Micro3 (n = 1024)) were identified by the Leiden algorithm. The Micro3 cluster was annotated as a disease-associated microglia (DAM) population by its genetic markers, in accordance with previous reports and significant enrichment in disease samples.Visualization of diffusion maps of sub-level cell clusters of the microglial cell population for each sample (Control 8 months, TauPS2APP 8 months, Control 13 months, TauPS2APP 13 months, n=2 for each condition) was also included. Figure 3L shows the expression levels of representative gene markers among different sub-clusters of cell populations. Dot plots are provided showing both the average gene expression (shading) of representative markers for each microglial subpopulation and the percentage of cells expressing them (dot size). Gene expression values were normalized for each column. Figure 3M provides diffusion maps of microglial cell populations together with pseudo-time trajectories. The plot shows a low-dimensional representation of the transcriptomic profile of microglial cells generated by Monocle 3 (top). The color map represents the pseudo-time values. Corresponding trajectories were constructed and plotted on the low-dimensional embedding to show the biological progression of the cell populations. Plots for each sample are also included. Trajectory start anchors were manually selected (based on the 8-month control sample). Cells from different samples are highlighted separately within the diffusion map embedding (bottom). Figure 3N shows the pseudotime embedding of the microglial cell population with subcluster annotation. A plot showing the low-dimensional embedding used in the pseudotime calculation is provided with the microglial subcluster annotation identified in Figure 3K and the trajectories identified in Figure 3M. Figure 3O provides a spatial map of the microglial population of a control 13-month sample. Scale bar, 100 μm. Figure 3P provides a spatial map of the microglial population of a TauPS2APP 13-month sample. Scale bar, 100 μm. Two enlarged sections (I, II) are highlighted within black boxes. Scale bar, 10 μm. Black dashed lines mark the border between the cortex, corpus callosum (cc), and hippocampus. Figure 3Q shows the cell density and cell type composition per region around the plaque. Box plots showing the density (number per mm2) of each microglial subcluster in both the cortical and hippocampal regions of control and TauPS2APP mice at the two time points are provided (top). Bar shading corresponds to the cell type legend in Figure 3P. The plots show significant enrichment of DAMs in both brain regions in TauPS2APP samples at 8 and 13 months.Stacked bar graphs showing the percentage of each microglial subcluster in each distance interval (0–10, 10–20, 20–30, 30–40, 40–50 μm) around Aβ plaques at 13 months are also provided (bottom). For each region, the overall cell density of each subpopulation in each area was included as a standard of comparison. Figure 3R provides a diffusion map showing the expression of four representative gene markers of microglial subtypes. Shading indicates the log10 (mean gene expression value) of genes in each cell. Figure 3S provides a matrix plot showing the Z-scores of microglial spatial DEGs across multiple distance intervals from the plaque (0–10, 10–20, 20–30, 30–40, 40+ μm).
[0020] [Fig. 4A-4T]Figures 4A-4T show the spatial composition of disease-associated populations and astrocytes under Tau+ plaque pathology. Spatiotemporal gene expression analysis of astrocytes in TauPS2APP and control samples is shown. Figure 4A provides the UMAP of astrocyte populations with annotation of subclusters. A plot showing a low-dimensional representation of the transcriptome profile of 2884 astrocytes identified from Figure 2A is provided. Three sublevel clusters (Astro1 (n=1068), Astro2 (n=1271), and Astro3 (n=545)) were identified by the Leiden algorithm. The Astro3 cluster was annotated as disease-associated astrocytes (DAA) by its genetic markers according to previous reports and significant enrichment in disease samples. UMAP plots of the sublevel cell clusters for each sample (control 8 months, TauPS2APP 8 months, control 13 months, TauPS2APP 13 months) were also included. Figure 4B shows the expression levels of representative markers across the different subclusters. Shading by row-wise Z-score. Figure 4C provides the UMAP of the astrocyte population with pseudotime trajectories. Plots showing a low-dimensional representation of the transcriptome profile of astrocytes generated by Monocle 3 are provided. The color map represents the pseudotime values. Corresponding trajectories were constructed and plotted on the low-dimensional embedding to show the biological progression of the population. Plots for each sample are also included. Figure 4D shows the pseudotime embedding of the astrocyte cell population with annotation of the subclusters. Plots showing the low-dimensional embedding used in the pseudotime calculation are provided with annotation of the astrocyte subclusters. Figure 4E shows the pseudotime distribution. The distribution of pseudotimes of the three subpopulations of astrocytes and the distribution of pseudotimes of astrocytes in the four samples are shown. Mann-Whitney-Wilcoxon test, ns: p>0.05, *p≤0.05, **p≤0.01, ***p≤0.001, ****p≤0.0001. Figure 4F provides a spatial map of the astrocyte population of a control 13-month sample. Scale bar, 100 μm. Figure 4G provides a spatial map of the astrocyte population of a TauPS2APP 13-month sample. Scale bar, 100 μm. Two enlarged sections (I, II) are shown in the upper right corner.Scale bar, 10 μm. Figure 4H shows cell density and cell type composition by region around the plaque. Stacked bar graphs showing the density (number per mm2) of each astrocyte subcluster in both the cortical and subcortical regions of each sample are provided (top). Bar graph shading corresponds to the cell type legend in Figure 4F. The plot shows significant enrichment of DAAs in the cortex, especially at 13 months. Stacked bar graphs showing the percentage of each astrocyte subcluster in each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 μm) around Aβ plaques are also provided (bottom). For each region, the overall cell type composition was included as a reference. Figure 4I shows the DEGs on pseudotime embedding. UMAP showing the expression of the four upregulated disease genes on the pseudotime embedding is shown, and the shaded scale of the raw counts was adjusted by log10 scale. Figure 4J shows the GSEA results of differentially expressed genes (DEGs). Terms are filtered by term size (20–1000) and nominal p-value (<0.01). Figure 4K provides a matrix plot showing the validated subset of DEGs of astrocyte populations obtained from AD vs. controls comparison. A plot is provided showing row-wise scaled expression values of the top significantly changed (ranked by p-value) DEGs in the 2766 gene dataset and validation results for 64 genes. Figure 4L shows the subclusters of astrocyte populations. A plot is provided showing the zoomed-in diffusion map visualization of the transcriptome profiles of 6,789 astrocytes identified from Figure 2A (top). Three sublevel clusters of astrocytes (Astro1 (n=2,547), Astro2 (n=3,278), and Astro3 (n=964)) were identified by the Leiden algorithm. The Astro3 cluster was annotated as disease-associated astrocytes (DAA) by its genetic markers in accordance with previous reports and significant enrichment in disease samples. Diffusion map embeddings of astrocyte sub-level cell clusters for each sample (control 8 months, TauPS2APP 8 months, control 13 months, TauPS2APP 13 months, n=2 for each condition) were also included.Figure 4M shows the expression levels of representative markers across the different subclusters. Dot plots showing both the average gene expression (shading) of representative markers for each astrocyte subpopulation and the percentage of cells expressing them (dot size) are provided. Gene expression values were normalized for each column. Figure 4N provides diffusion maps of astrocyte populations along with pseudotime trajectories. A diffusion map visualization of the transcriptomic profile of astrocytes generated by Monocle 3 is provided (top). The color map represents the pseudotime values. Corresponding trajectories were constructed and plotted on the low-dimensional embedding to show the biological progression of the populations. Plots for each sample are also included. Trajectory start anchors were manually selected based on the Astro1 population. Black arrows highlight branching points on the trajectories associated with disease-associated gene expression changes. Cells from different samples were highlighted separately within the diffusion map embedding (bottom). Figure 4O provides a diffusion map showing the different types of astrocytes identified in Figure 4L along with the trajectories identified in Figure 4N. Pseudotime embedding of the astrocyte cell population is shown with annotation of the subclusters. A plot showing the low-dimensional embedding used in the pseudotime calculation is provided with annotation of the astrocyte subclusters. Figure 4P provides a spatial map of the astrocyte population of the TauPS2APP 13-month sample. Scale bar, 100 μm. Figure 4Q provides a spatial map of the astrocyte population of the TauPS2APP 13-month sample. Scale bar, 100 μm. Two enlarged sections (I, II) are shown. Scale bar, 10 μm. The black dashed lines mark the border between the cortex, corpus callosum (cc), and hippocampus. Figure 4R shows the cell density and cell type composition by region around the plaque. Box plots showing the density (number per mm2) of each astrocyte subcluster in both the cortical and hippocampal regions of each sample (control and TauPS2APP mice) at two time points are provided (top). The shading of the bar graphs corresponds to the cell type legend in Figure 4Q. The plot shows significant enrichment of TauPS2APP13 in the cortex especially at 13 months.Stacked bar graphs showing the density (number per mm2) of astrocyte subclusters in each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 μm) around Aβ plaques at 13 months are also provided (bottom). In each region, the cell density of each subpopulation in each area was included as a standard of comparison. Figure 4S shows the DEGs on pseudotime embedding. Diffusion maps showing the expression of four representative gene markers of astrocyte subpopulations on pseudotime embedding are shown, and the shaded scale of raw counts was adjusted by log10 transformation. Figure 4T provides a matrix plot showing the Z-scores of spatial DEGs of astrocytes across multiple distance intervals (0–10, 10–20, 20–30, 30–40, 40+ μm) from plaques.
[0021] [Figure 5A-5V]Figures 5A-5V show the sub-level clustering and composition of oligodendrocytes and progenitor cells under Tau+ plaque pathology. Spatiotemporal gene expression analysis of oligodendrocytes and progenitor cells in TauPS2APP and control samples is shown. Figure 5A provides the UMAP of oligodendrocyte and progenitor cell (OPC) populations with annotation of sub-clusters. UMAP plots showing low-dimensional representation of transcriptome profiles of 4966 oligodendrocytes and 549 OPCs identified from Figure 2A are provided. Three sub-level clusters (Oligo1 (n=4,295), Oligo2 (n=181), Oligo3 (n=490)) and OPCs (n=549) were identified by the Leiden algorithm. The Oligo3 cluster was annotated as disease-enriched oligodendrocytes. UMAP plots of sub-level cell clusters for each sample (Control 8 months, TauPS2APP 8 months, Control 13 months, TauPS2APP 13 months) were also included. Figure 5B shows the expression levels of representative markers across different sub-clusters of oligodendrocytes (Z-scores per row). Figure 5C provides the UMAP of oligodendrocyte-associated cell populations with pseudo-time trajectories. Plots showing low-dimensional representations of transcriptome profiles of oligodendrocytes and OPCs generated by Monocle 3 are provided. Color maps represent pseudo-time values. Corresponding trajectories were constructed and plotted on the low-dimensional embedding to show the biological progression of the populations. Plots for each sample are also included. Figure 5D shows the pseudo-time embedding of oligodendrocyte-associated cell populations with annotations of sub-clusters. Plots showing the low-dimensional embedding used in the pseudo-time calculation are provided with sub-cluster annotations of oligodendrocytes and progenitors. Figure 5E shows the distribution of pseudotimes of the three subpopulations of oligodendrocytes and oligodendrocyte precursor cells, as well as the distribution of pseudotimes of those cells in the four samples (Mann-Whitney-Wilcoxon test, ns: p>0.05, *p≦0.05, **p≦0.01, ***p≦0.001, ****p≦0.0001).Figure 5F provides a spatial map of the oligodendrocyte-associated populations of a control 13-month sample. Scale bar, 100 μm. Figure 5G provides a spatial map of the oligodendrocyte-associated populations of a TauPS2APP 13-month sample. Scale bar, 100 μm. Two enlarged sections (I, II) are provided in the upper right corner. Scale bar, 10 μm. Figure 5H shows the cell density and cell type composition per region around the plaque. Stacked bar graphs are provided showing the density (number per mm2) of each oligodendrocyte subcluster and OPCs in both the cortical and subcortical regions of each sample. The shading of the bar graphs corresponds to the cell type legend in Figure 5F. Stacked bar graphs showing the percentage of each oligodendrocyte subcluster and OPCs in each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 μm) around the Aβ plaque are also provided (bottom). For each region, the overall cell type composition was included as a reference. Figure 5I shows the results of gene set enrichment analysis (GSEA) of differentially expressed genes (DEGs). Terms are filtered by term size 20–1000 and nominal p-value <0.01. Figure 5J shows the pseudo-time embedding of cells expressing marker genes. UMAPs showing the pseudo-time distribution of cells expressing representative markers are provided. Figure 5K provides a matrix plot showing the validated subset of DEGs of oligodendrocyte-related populations from AD vs. controls comparison. Plots showing scaled expression values by row of the top significantly changed DEGs (ranked by p-value) in the 2766 gene dataset and validation results for 64 genes are provided. Figure 5L shows subclusters of oligodendrocyte and progenitor cell (OPC) populations. A zoomed-in diffusion map visualization of 11,265 oligodendrocytes and 1,269 OPCs identified from Figure 2A is provided (top). Shown are three subclusters of oligodendrocytes identified by Leiden clustering: Oligo1 (n = 9,594), Oligo2 (n = 910), Oligo3 (n = 761), and OPC (n = 1,269).Diffusion map embeddings of oligodendrocyte and OPC subclusters for each sample (Control 8 months, TauPS2APP 8 months, Control 13 months, TauPS2APP 13 months) are also provided (bottom). Figure 5M shows the expression levels of representative markers across different subclusters of oligodendrocytes and OPCs. Dot plots are provided showing both the average gene expression (shading) of representative markers for each subpopulation and the percentage of cells expressing them (dot size). Gene expression values were normalized for each column. Figure 5N shows the diffusion map pseudotime trajectory visualization of oligodendrocyte-associated cell populations. Diffusion map visualization of pseudotime trajectories of oligodendrocytes and OPCs generated by Monocle 3 is provided (top). Color maps represent pseudotime values. Corresponding pseudotime trajectories were plotted on the diffusion map embedding. Trajectory start anchors were manually selected based on the OPC population. Black arrows highlight branch points on the trajectories associated with disease-associated gene expression changes. Cells from different samples were highlighted separately within the diffusion map embedding (bottom). Figure 5O provides diffusion maps showing the different types of oligodendrocytes and OPCs identified in Figure 5L, along with trajectories. Figure 5P provides spatial cell maps of the oligodendrocyte-associated populations of a control 13-month sample. Scale bar, 100 μm. Figure 5Q provides spatial cell maps of the oligodendrocyte-associated populations of a TauPS2APP 13-month sample. Scale bar, 100 μm. Two magnified sections (I, II) are provided with magnified areas highlighted in black boxes. Scale bar, 10 μm. Black dashed lines mark the boundaries between the cortex, corpus callosum (cc), hippocampus, and alveus. Figure 5R provides box plots showing the density (number per mm2) of each oligodendrocyte subtype and OPCs in the cortex, corpus callosum, and hippocampus regions of control and TauPS2APP mice at two time points. The plots show a significant enrichment of Oligo2 in the corpus callosum (cc) and hippocampus regions.Figure 5S provides stacked bar graphs showing the density (number per mm2) of each oligodendrocyte subcluster and OPCs in each distance interval (0–10, 10–20, 20–30, 30–40, 40–50 μm) around Aβ plaques at 13 months. The overall cell density of each subpopulation in each area was included as a standard of comparison (all). Figure 5T shows cell type composition analysis of oligodendrocyte lineages associated with p-Tau pathology. Tissue areas of TauPS2APP samples at 13 months were divided into a 20 μm × 20 μm grid and ranked by p-Tau density. Blocks divided by grid lines were ranked by the percentage of p-Tau positive pixels and grouped into three bins: zero (0%), low (1%–50%), and high (50%–100%). The high p-Tau bins are further divided into two groups based on the presence or absence of Aβ plaques. Stacked bar plots showing the average number of cells per block of each oligodendrocyte-associated subtype are provided. Figure 5U shows the cell density and subtype composition of oligodendrocytes and OPCs in the hippocampal alveolar region. Figure 5V provides a matrix plot showing the z-scores of spatial DEGs of oligodendrocytes across multiple distance intervals from the plaque (0–10, 10–20, 20–30, 30–40, 40+ µm).
[0022] [Fig. 6A-6U]Figure 6A-6U show differential gene expression analysis and spatial information of neurons in TauPS2APP and control samples. Figure 6A provides a spatial map of excitatory neuronal populations. Scale bar, 100 μm. A 13-month sample of TauPS2APP (total cell counts: CTX-Ex1: 2,292, CTX-Ex2: 2,766, CTX-Ex3: 1,184, CTX-Ex4: 2,445, CA1: 2,754, CA2: 436, CA3: 1,878, DG: 4,377) and a 1x magnified section of the CA1 region at the bottom. Scale bar, 10 μm. p-Tau protein signal is colored gray on the cells. Figure 6B provides a spatial map of inhibitory neuronal populations. Scale bar, 100 μm. TauPS2APP Cnr1:421, Lamp5:191, Pvalb:864, Sst:529. A 1× magnified section of the CA1 region from a 13-month sample is shown below. Scale bar, 10 μm. p-Tau protein signals are colored gray on the cells. Figure 6C shows the excitatory neuronal composition around plaques. Stacked bar graphs are provided showing the percentage of each subcluster of excitatory neuronal population in each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 μm) around Aβ plaques from the cortical and subcortical regions of a 13-month AD sample. For each region, the overall cell type composition was included as a reference. Figure 6D shows the inhibitory neuronal composition around plaques. Stacked bar graphs are provided showing the percentage of each subcluster of inhibitory neuron population in each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around Aβ plaques from cortical and subcortical regions of 13-month AD samples. For each region, the overall cell type composition is included as a reference. Figure 6E shows the cell type composition of p-Tau positive neurons. Stacked bar graphs are provided showing the composition of Tau positive excitatory and inhibitory neurons in each current subthreshold AD sample. Figure 6F shows the quantification of p-Tau signal around plaques.p-Tau+ pixels (intensity > threshold) were quantified in each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 μm) around Aβ plaques from cortical and subcortical regions of TauPS2APP 13-month samples. Values were normalized by ring area. Figure 6G provides UMAPs of dentate gyrus (DG) cell populations along with pseudotime trajectories. Plots are provided showing low-dimensional representations of transcriptome profiles of cell populations within the dentate gyrus region of each sample generated by Monocle 3. Color maps represent pseudotime values. Corresponding trajectories were constructed and plotted on the low-dimensional embedding to show the biological progression of the populations. Plots for each sample are also included. Figure 6H provides spatial cell maps shaded by pseudotime for the DG population. Scale bar, 100 μm. Figure 6I shows DEGs on the pseudotime embedding. From the DEG analysis of DG neurons on pseudotime embedding, a UMAP is provided showing the expression of four significantly changed genes. The shading scale of the raw counts was adjusted by log10 scale. Figure 6J provides a matrix plot showing the validated subsets of DEGs of excitatory and inhibitory neuronal populations from AD vs. controls comparison. A plot showing the row-wise scaled expression values of the top significantly changed DEGs (ranked by p-value) in the 2,766 gene dataset and the validation results of 64 genes is provided. Figure 6K shows the sub-clustering of cortical excitatory neuronal cell populations. Visualization of the UMAP showing four sub-clusters of cortical excitatory neuronal cells identified by Leiden clustering: CTX-Ex1 (n=4,972), CTX-Ex2 (n=6,049), CTX-Ex3 (n=4,158), and CTX-Ex4 (n=3,304). Figure 6L shows the expression levels of representative gene markers among different sub-clusters of cortical excitatory neuronal cells. Dot plots are provided showing both the average gene expression (shading) of representative markers for each subpopulation, and the percentage of cells expressing them (dot size). Gene expression values were normalized for each column. Figure 6M shows the subclustering of inhibitory neuronal cell populations.UMAP visualization showing six subclusters of inhibitory cells identified by Leiden clustering: Cnr1 (n=512), Lamp5 (n=573), Pvalb (n=1,392), Pvalb-Nog (n=410), Sst (n=959), and Vip (n=317). Figure 6N shows the expression levels of representative gene markers among different subclusters of inhibitory neuronal cells. Dot plots are provided showing both the average gene expression (shading) of representative markers for each subpopulation and the percentage of cells expressing them (dot size). Gene expression values were normalized for each column. Figure 6O provides spatial maps of Aβ plaques and p-Tau in excitatory neuronal populations in a 13-month sample of TauPS2APP (top). Scale bar, 100 μm. Total cell counts in the 13-month sample of TauPS2APP: CTX-Ex1: 2,292, CTX-Ex2: 2,766, CTX-Ex3: 1,184, CTX-Ex4: 2,445, CA1: 2,754, CA2: 436, CA3: 1,878, DG: 4,377 and a higher magnification view of a section of the CA1 region indicated in the black box is shown at the bottom. Scale bar, 10 μm. p-Tau protein signal is colored gray on the cells. Figure 6P provides a spatial map of Aβ plaques and p-Tau in the inhibitory neuronal population in the 13-month sample of TauPS2APP (top). Scale bar, 100 μm. Cnr1:421, Lamp5:191, Pvalb:864, Sst:529 in TauPS2APP. A high magnification view of a 13-month sample and a section of the CA1 region indicated in the black box is shown below. Scale bar, 10 μm. p-Tau protein signal is colored gray on the cells. Figure 6Q shows the excitatory neuron composition around plaques. Stacked bar graphs are provided showing the density (number per mm2) of each subcluster of inhibitory neuron populations from various brain regions in each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 μm) around Aβ plaques from cortical and subcortical regions of a 13-month sample of TauPS2APP in AD. In each region, the overall cell density of each cell type subpopulation was included as a reference comparison standard.Figure 6R shows the inhibitory neuron composition around plaques. Stacked bar graphs are provided showing the density (number per mm2) of each subcluster of inhibitory neuron populations from various brain regions at each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around Aβ plaques from cortical and subcortical regions of 13-month AD TauPS2APP samples. In each region, the overall cell density of each cell type subpopulation was included as a reference comparison standard. Figure 6S provides bar graphs showing the density (number per mm2) of each cortical excitatory and inhibitory neuron subtype in control and TauPS2APP mice. Figure 6T shows the quantification of p-Tau signal around Aβ plaques. p-Tau+ pixels (intensity > threshold) were quantified in each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 µm) around Aβ plaques from cortical and subcortical regions of 13-month samples of TauPS2APP. Values were normalized by ring area. Y-axis values were normalized by total p-Tau signal in each brain region. Figure 6U shows the cell type composition of p-Tau positive neurons. Stacked bar graphs show the composition of p-Tau positive excitatory and inhibitory neurons in each AD sample (control and TauPS2APP) under the current threshold at two time points, defined by the ratio of Tau positive pixels to the area of each cell.
[0023] [Figure 7A-7M]Figures 7A-7M show integrated pathway and spatial analysis of disease-associated cell types. Figure 7A provides a GSEA heatmap showing significant (nominal p-value < 0.05) biological process-related terms enriched in DEGs for each cell type of interest in AD 13-month samples. Terms are filtered by term size (20-1000). Tile shading represents normalized enrichment scores. Figure 7B provides a matrix plot showing upregulated genes (LogFC>0.1) in the plaque-proximal region of AD 13-month samples from cells within the 25 μm ring compared to cells outside the 25 μm ring (shaded by row-wise Z-score). Figure 7C provides a Venn diagram highlighting the overlap of plaque-induced genes (PIGs) in AD 13-month samples with PIGs in AD 8-month samples and previously reported PIGs in 18-month AppNL-GF mice. Figure 7D provides spatial histograms of disease-associated microglia (DAMs), disease-associated astrocytes (DAAs), oligodendrocytes, OPCs, and neuronal cell populations around Aβ plaques in 8-month (left) and 13-month (right) AD samples. In the histograms, cells were counted in 10 μm bins in 2D maximum projections from the edge of each plaque. Figure 7E provides schematics of DAMs, DAAs, oligodendrocytes, OPCs, microglia (excluding DAMs), astrocytes (excluding DAAs), and neuronal cell populations around Aβ plaques in the brains of 8-month AD mice (left) and 13-month mice. The number of cells in the schematics is an approximation of the calculated average cell number of each cell type within each ring. Figures 7F-7G provide matrix plots showing gene clustering results in each distance interval (0-10, 10-20, 20-30, 30-40, >40 μm) around Aβ plaques in 8-month samples of TauPS2APP (Figure 7F) and 13-month samples of TauPS2APP (Figure 7G), shaded by row-wise Z-score.Figures 7H-7I provide matrix plots showing plaque-induced genes (enrichment in intervals 0-40, adjusted p-values <0.01) in each distance interval (0-10, 10-20, 20-30, 30-40, >40 μm) around Aβ plaques in TauPS2APP 8-month samples (Figure 7H) and TauPS2APP 13-month samples (Figure 7I). Colored by Z-scores per row. Figure 7J shows a Venn diagram highlighting the overlap of SDEGs in TauPS2APP 8-month and 13-month samples with SDEGs in TauPS2APP and previously reported PIGs in 18-month AppNL-GF mice. Figure 7K shows significantly enriched GO terms for SDEGs in TauPS2APP 8-month and 13-month samples and previously reported PIGs. Figure 7L provides spatial histograms of Micro3 disease-associated microglia (DAMs), Astro3 disease-associated astrocytes (DAAs), Oligo2 / 3 oligodendrocytes, OPCs, and neuronal cell populations around Aβ plaques in 8-month (left) and 13-month (right) AD TauPS2APP samples. In the histograms, cells were counted in 10 μm bins in 2D maximum projections from the edge of each plaque. Figure 7M provides schematics of various cell types (e.g., DAMs, DAAs, oligodendrocytes, OPCs, microglia (except DAMs), astrocytes (except DAAs), and neuronal cell populations) around Aβ plaques and oligodendrocyte subtypes in the hippocampal alveolus of 8-month AD (left) and 13-month AD mouse brains. The number of cells in the schematics is an approximate ratio of the number of cells of each cell type in each ring.
[0024] [Figure 8A-8E]Figures 8A-8E show the development of the STARmap Pro method. Figure 8A shows the STARmap Pro procedure, where p-Tau primary antibody staining was performed after in situ hybridization and amplification of mRNA. The imaging results showed strong signals from both the cDNA amplicon and protein. Figure 8B shows an alternative procedure, where p-Tau primary antibody staining was performed before in situ hybridization and amplification of mRNA. The imaging results showed a much weaker signal from the cDNA amplicon, suggesting that RNA degradation was occurring during the antibody incubation and washing steps. Figure 8C provides a schematic illustrating the enhanced specificity of STARmap Pro compared to the previous STARmap method. In the original STARmap method design, the DNA probes of all genes share identical DNA sequences at the ligation junctions, so the padlock probes can be circularized and amplified nonspecifically, even when there are primer probes nearby. This can lead to false signals when using the previous STARmap method. STARmap Pro places an additional 5 nt barcode at the ligation site to prevent ligation of non-specifically bound probes, thus improving specificity. Figures 8D and 8E show the imaging results of STARmap Pro using a SNAIL probe with (Figure 8D) and without (Figure 8E) a barcode mismatch near the ligation site. DAPI staining of cell nuclei and fluorescent DNA probe staining of all cDNA amplicons are shown according to the provided legend.
[0025] [Figures 9A-9G]Figures 9A-9G show the top cell type classification results for all samples. Figure 9A provides stacked violin plots of representative gene markers aligned with each top cell type for the 2,766 gene dataset. Figure 9B provides gene expression heat maps of representative markers aligned with each top cell type for the 2,766 and 64 gene datasets. The 64 gene data successfully reproduced the top clustering results. Expression of each gene is Z-scored across all genes within each cell. Figure 9C provides a Uniform Manifold Approximation (UMAP) plot visualizing the nonlinear dimensionality reduction of the transcriptome profiles of 36,625 cells from four samples in the validation dataset. Plots are provided showing the following: 13 top clusters (cortical excitatory neurons (CTX-Ex, 8,640 cells), inhibitory neurons (In, 2,858 cells), CA1 excitatory neurons (CA1-Ex, 2,967 cells), CA2 excitatory neurons (CA2-Ex, 331 cells), CA3 excitatory neurons (CA3-Ex, 1751 cells), dentate gyrus (DG, 4,560 cells), astrocytes (Astro, 3,423 cells), endothelial cells (Endo, 1,661 cells), microglia (Micro, 2,183 cells), oligodendrocytes (Oligo, 5,268 cells), oligodendrocyte precursor cells (OPC, 711 cells), smooth muscle cells (SMC, 1,146 cells), and mixed unidentified cells (Mix, 1126 cells). Figures 9D and 9E show spatial atlases of top cell types in the cortical and hippocampal regions of four samples in the 2,766-gene dataset (Figure 9D) and the 64-gene dataset (Figure 9E). Scale bars, 100 µm. Figures 9F and 9G show the cell type composition around Aβ plaques at different distance intervals in the 2,766-gene dataset (F) and the 64-gene validation dataset (G) for both 8-month and 13-month samples. Stacked bar plots are provided showing the percentage of each top cell type in each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 µm) around Aβ plaques. For each region, the overall cell type composition was included as a reference.
[0026] [Figure 10A-10E] Figures 10A-10E show additional gene expression and spatial information of the microglial population. Figure 10A provides a spatial map of microglial subtypes in 8-month control and TauPS2APP samples. Scale bar, 100 μm. Two enlarged sections (I, II) are shown on the right. Scale bar, 10 μm. Figure 10B shows the cell type composition around Aβ plaques at different distance intervals in TauPS2APP samples at 8 months. Stacked bar plots are provided showing the percentage of each microglial subpopulation in each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around Aβ plaques. For each region, the overall cell type composition was included as a reference. Figure 10C provides a spatial map of microglia colored by pseudotime. Scale bar, 100 μm. Two enlarged sections (I, II) are shown on the right. Scale bar, 10 μm. Figure 10D shows the microglial pseudotime values relative to plaques. Box plots showing the distribution of microglial pseudotimes at each distance interval (0–10, 10–20, 20–30, 30–40, and 40–50 µm) around Aβ plaques are provided. Total microglial distribution was included as a reference. Box shading represents the median pseudotime. Figure 10E provides a volcano plot for microglial differential expression. The plot shows microglial gene expression across AD and control samples at 8 and 13 months of age (y-axis: -log adjusted p-values, x-axis: mean log fold change). Differentially expressed genes (adjusted p-values <0.05, absolute logFC values >0.1) have positive fold change values (upregulated) or negative fold change values (downregulated).
[0027] [Figures 11A-11E]Figures 11A-11E show additional gene expression and spatial information for the astrocyte populations. Figure 11A provides a spatial map of astrocyte subtypes for control and TauPS2APP samples at 8 months. Scale bar, 100 μm. Two enlarged sections (I, II) are shown on the right. Scale bar, 10 μm. Figure 11B shows the cell type composition around Aβ plaques at different distance intervals for the TauPS2APP 8 months sample. Stacked bar plots are provided showing the percentage of each astrocyte subpopulation in each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around the Aβ plaques. For each region, the overall cell type composition was included as a reference. Figure 11C provides a spatial map of astrocytes shaded by pseudotime for the astrocyte cell populations. Scale bar, 100 μm. Two enlarged sections (I, II) are shown on the right. Scale bar, 10 μm. Figure 11D shows pseudotimes relative to plaques. Box plots showing the distribution of astrocyte pseudotimes in each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around Aβ plaques are provided. The distribution of total astrocytes was included as a reference. Figure 11E provides a volcano plot showing differential gene expression in astrocytes. Plots showing astrocyte gene expression across 8-month and 13-month samples of AD and controls are provided (y-axis: -log adjusted p-value, x-axis: mean log fold change). Differentially expressed genes (adjusted p-value < 0.05, absolute value of logFC > 0.1) have positive fold change values (upregulated) or negative fold change values (downregulated).
[0028] [Fig. 12A-12F]Figures 12A-12F show additional gene expression and spatial information of oligodendrocyte and OPC cell populations. Figure 12A provides cell-resolved spatial maps for oligodendrocyte and OPC populations of both control and TauPS2APP 8-month samples. Scale bar, 100 μm. Two enlarged sections (I, II) are shown on the right. Scale bar, 10 μm. Figure 12B shows the cell type composition around Aβ plaques at different distance intervals for TauPS2APP 8-month samples. Stacked bar plots are provided showing the percentage of each oligodendrocyte subpopulation and OPCs in each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around Aβ plaques. For each region, the overall cell type composition was included as a reference. Figure 12C provides pseudotime shaded spatial maps for oligodendrocyte-associated cell populations. Scale bar, 100 μm. Two enlarged sections (I, II) are shown on the right. Scale bar, 10 μm. Figure 12D shows pseudotimes relative to plaques. Box plots are provided showing the distribution of pseudotimes of oligodendrocytes in each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around Aβ plaques. The distribution of total oligodendrocytes was included as a reference. Figure 12E shows pseudotimes relative to plaques. Box plots are provided showing the distribution of pseudotimes of OPCs in each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around Aβ plaques. The distribution of total OPCs was included as a reference. Figure 12E provides a volcano plot showing differential expression in oligodendrocytes. Plots are provided showing oligodendrocyte gene expression across 8-month and 13-month AD and control samples (y-axis: -log adjusted p-value, x-axis: mean log fold change). Differentially expressed genes (adjusted p-value < 0.05, absolute logFC > 0.1) have positive fold change values (upregulated) or negative fold change values (downregulated).
[0029] [Figures 13A-13G]Figures 13A-13G show additional gene expression and spatial information for the neuronal populations. Figure 13A provides spatial maps of the excitatory neuronal populations of four samples. Scale bar, 100 μm. Figure 13B provides spatial maps of the inhibitory neuronal populations of four samples. Scale bar, 100 μm. Figure 13C shows the cell density and cell type composition by region around excitatory neuronal plaques. Stacked bar graphs showing the density (number per mm2) of each excitatory neuronal subcluster in both cortical and subcortical regions of each sample are provided (top). The shading of the bar graphs corresponds to the cell type legend in Figure 13A. Stacked bar graphs showing the percentage of each excitatory neuronal subcluster in each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around Aβ plaques are also provided (bottom). For each region, the overall cell type composition was included as a reference. Figure 13D shows the cell density and cell type composition by region around inhibitory neuron plaques. Stacked bar graphs showing the density (number per mm2) of each inhibitory neuron subcluster in both cortical and subcortical regions of each sample are provided (top). Bar graph shading corresponds to the cell type legend in Figure 13B. Stacked bar graphs showing the percentage of each inhibitory neuron subcluster in each distance interval (0-10, 10-20, 20-30, 30-40, and 40-50 μm) around Aβ plaques are also provided (bottom). For each region, the overall cell type composition was included as a reference. Figures 13E-13G show enrichment of synaptic Gene Ontology terms using SynGO for excitatory neurons, inhibitory neurons, and DEGs from CA1 cells with tau pathology. Shading in sunburst plots represents enrichment-log10Q values at 1% FDR.
[0030] [Figures 14A-14E]Figures 14A-14E show pathway and spatial analysis of AD mouse brain. Figure 14A provides a gene ontology heat map showing biological process-related terms enriched for upregulated genes in each cell type of interest in AD 8-month samples. Terms are filtered by term size 20-1000. Tile shading represents enrichment-log10(FDR) values. Figure 14B provides functional enrichment maps generated using Cytoscape with EnrichmentMap and AutoAnnotate apps from differentially expressed genes (abs(LogFC)>0.1, Wilcox test p-value<0.05) in neuronal cells (excitatory, inhibitory neurons) and non-neuronal cells (microglia, astrocytes, and oligodendrocytes). Each circle represents a GO term and is shaded by cell type to illustrate the common and distinct contributions of DEGs from each cell cluster. Ellipses indicate constellations of GO terms clustered by the AutoAnnotate app. Figure 14C (left) provides a heatmap showing the average nearest neighbor distance between associated cell types and Aβ plaques in 13-month samples of AD. The average nearest neighbor distance is calculated as follows: 1) For all cells and plaques, calculate the Euclidean distance to each other object, find the closest cell to each cell type or plaque, and store that distance. 2) For all comparisons of interest (i.e., Micro vs. Plaque), calculate the average of that distance distribution. Figure 14C (right) shows the same distances as in Figure 14C (left), but using shuffled (randomized) cell type labels. Figure 14D (left) provides a heatmap showing the average nearest neighbor distance between associated cell types and Aβ plaques in 8-month samples of AD. The calculation is similar to that used in Figure 14C. Figure 14D (right) shows the same distances as in 14D (left), but using shuffled (randomized) cell type labels. FIG. 14E provides a matrix plot showing genes upregulated in the juxta-plaque region of AD 8-month samples from cells within the 25 μm ring compared to cells outside the 25 μm ring.
[0031] [Figure 15] FIG. 15 is a flow diagram of one embodiment of a method relating to identifying spatial variation of cell types within at least one image.
[0032] [Figure 16] FIG. 16 is a flow diagram of one embodiment of a method relating to identifying spatial variation of cells of a first cell type in at least one image.
[0033] [Figure 17] FIG. 17 is a flow diagram of an embodiment of a method relating to capturing at least one image using a camera.
[0034] [Figure 18] FIG. 18 is a flow diagram of one embodiment for spatially aligning spatial locations from multiple images.
[0035] [Figure 19] FIG. 19 is a flow diagram of one embodiment of a method directed to determining the number of cells of a cell type that are less than a threshold distance to a first protein.
[0036] [Figure 20] FIG. 20 is a flow diagram of one embodiment of a method relating to determining the cell type of a cell.
[0037] [Figure 21] FIG. 21 is a block diagram of a computer system in which various functions or methods may be implemented.
[0038] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which the present invention belongs.The following references provide those skilled in the art with the general definitions of many terms used in the present invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994);The Cambridge Dictionary of Science and Technology (Walker ed., 1988);The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991);And Hale & Marham, The Harper Collins Dictionary of Biology (1991).As used herein, the following terms have the meanings ascribed to them unless otherwise specified.
[0039] The terms "administer," "administering," and "administration" refer to implanting, absorbing, ingesting, injecting, inhaling, or otherwise introducing a treatment or therapeutic agent, or a treatment or therapeutic agent composition, into or onto a subject.
[0040] The term "amplicon" as used herein refers to a nucleic acid (e.g., DNA or RNA) that is the product of an amplification reaction (i.e., the generation of one or more copies of a gene fragment or target sequence) or a replication reaction. Amplicons can be formed artificially, for example, using PCR or other polymerization reactions. The term "concatenated amplicons" refers to multiple amplicons that are linked together to form a single nucleic acid molecule. Concatenated amplicons can be formed, for example, by rolling circle amplification (RCA), in which a circular oligonucleotide is amplified to generate multiple linear copies of the oligonucleotide as a single nucleic acid molecule that contains multiple linked amplicons.
[0041] "Antibody" refers to a glycoprotein that belongs to the immunoglobulin superfamily. The terms antibody and immunoglobulin are used interchangeably. With some exceptions, mammalian antibodies are typically made of a basic structural unit that has two large heavy chains and two small light chains each. There are several different types of antibody heavy chains, and several different types of antibodies that are grouped into different isotypes based on which heavy chains they have. Five different antibody isotypes are known in mammals (IgG, IgA, IgE, IgD, and IgM), which play different roles and help direct the appropriate immune response to each different type of foreign body that they encounter. In some embodiments, an antibody as used herein binds to a protein of interest (e.g., any protein of interest expressed in a cell). The term "antibody" as used herein also encompasses antibody fragments and nanobodies, as well as antibody variants and antibody fragments and nanobody variants.
[0042] The term "cancer" refers to a class of diseases characterized by the development of abnormal cells that grow uncontrollably and have the ability to invade and destroy normal body tissues. See, e.g., Stedman's Medical Dictionary, 25th ed.; Hensyl ed.; Williams & Wilkins: Philadelphia, 1990. Examples of cancer include, but are not limited to, acoustic neuroma; adenocarcinoma; adrenal cancer; angiosarcoma (e.g., lymphangiosarcoma, lymphangioendothelial sarcoma, angiosarcoma); appendix cancer; benign monoclonal gammopathy; biliary tract cancer (e.g., cholangiocarcinoma); bladder cancer; breast cancer (e.g., adenocarcinoma of the breast, papillary carcinoma of the breast, breast carcinoma, medullary carcinoma of the breast); brain cancer (e.g., meningioma, glioblastoma, glioma (e.g., astrocytoma, oligodendroglioma), medulloblastoma); bronchial cancer: carcinoid tumor tumors;cervical cancer (e.g., cervical adenocarcinoma);choriocarcinoma;chordoma;craniopharyngioma;colorectal cancer (e.g., colon carcinoma, rectal carcinoma, colorectal adenocarcinoma);connective tissue cancer;epithelial carcinoma;ependymoma;endothelial sarcoma (e.g., Kaposi's sarcoma, multiple idiopathic hemorrhagic sarcoma);endometrial cancer (e.g., uterine carcinoma, uterine sarcoma);esophageal cancer (e.g., esophageal adenocarcinoma, Barrett's adenocarcinoma);Ewing's sarcoma;eye cancer (e.g., intraocular melanoma, retinoblastoma);familial eosinophilia;gallbladder cancer;gastric cancer (e.g., gastric adenocarcinoma);gastrointestinal stromal tumors (GIST);germ cell cancer;head and neck cancer (e.g. head and neck squamous cell carcinoma, oral cavity cancer (e.g. oral squamous cell carcinoma), pharyngeal cancer (e.g. laryngeal cancer, pharyngeal cancer, nasopharyngeal cancer, oropharyngeal cancer));hematopoietic cancer (e.g. leukemia such as acute lymphoblastic leukemia (ALL) (e.g. B-cell ALL, T-cell ALL), acute myeloid leukemia (AML) (e.g. B-cell AML, T-cell AML), chronic myeloid leukemia (CML) (e.g. B-cell CML, T-cell CML), and chronic lymphocytic leukemia (CLL) (e.g. B-cell CLL, T-cell CLL));Lymphomas, such as Hodgkin's lymphoma (HL) (e.g. B-cell HL, T-cell HL) and non-Hodgkin's lymphoma (NHL) (e.g. B-cell NHL, e.g. diffuse large cell lymphoma (DLCL) (e.g. diffuse large B-cell lymphoma), follicular lymphoma, chronic lymphocytic leukemia / small lymphocytic lymphoma (CLL / SLL), mantle cell lymphoma (MCL), marginal zone B-cell lymphoma (e.g. mucosal intralymphoid tissue (MALT) lymphoma, nodal marginal zone B-cell lymphoma, splenic marginal zone B-cell lymphoma), primary mediastinal B-cell lymphoma, Burkitt's lymphoma, lymphoplasmacytic lymphoma, lympho-plasmocy ... lymphomas (i.e., Waldenström's macroglobulinemia), hairy cell leukemia (HCL), immunoblastic large cell lymphoma, precursor B-lymphoblastic lymphoma, and primary central nervous system (CNS) lymphomas; and T-cell NHLs, such as precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphomas (PTCLs) (e.g., cutaneous T-cell lymphoma (CTCL) (e.g., mycosis fungoides, Sézary syndrome), angioimmunoblastic T-cell lymphoma, extranodal natural killer T-cell lymphoma, enteropathy-type T-cell lymphoma, subcutaneous panniculitis-like T-cell lymphoma, and anaplastic large cell lymphoma). ; one or more of the above leukemia / lymphoma mixtures; and multiple myeloma (MM)); heavy chain disorders (e.g., alpha chain disorders, gamma chain disorders, mu chain disorders); hemangioblastoma; hypopharyngeal carcinoma; inflammatory myofibroblastic tumors; immune cell amyloidosis; kidney cancer (e.g., nephroblastoma, also known as Wilms tumor, renal cell carcinoma); liver cancer (e.g., hepatocellular carcinoma (HCC), malignant hepatoma); lung cancer (e.g., bronchogenic carcinoma, small cell lung cancer (SCLC), non-small cell lung cancer (NSCLC), adenocarcinoma of the lung); leiomyosarcoma (LMS); mastocytosis (e.g., systemic mastocytosis); muscle cancer; myelodysplastic syndromes (MDS) );mesothelioma;myeloproliferative disorders (MPDs) (e.g., polycythemia vera (PV), essential thrombocytosis (ET), myeloid metaplasia of unknown origin (AMM), also known as myelofibrosis (MF), chronic idiopathic myelofibrosis, chronic myeloid leukemia (CML), chronic neutrophilic leukemia (CNL), hypereosinophilic syndrome (HES));neuroblastoma;neurofibromas (e.g., neurofibromatosis (NF) type 1 or 2, schwannomatosis);neuroendocrine carcinomas (e.g., gastroenteropancreatic neuroendocrine tumors (GEP-NETs), carcinoid tumors);osteosarcoma (e.g., bone cancer);ovarian cancer (e.g., cystadenocarcinoma, ovarian embryonal carcinoma, ovarian adenocarcinoma);papillary adenocarcinoma;Pancreatic cancer (e.g., pancreatic adenocarcinoma, intraductal papillary mucinous neoplasm (IPMN), islet cell tumor); penile cancer (e.g., Paget's disease of the penis and scrotum); pinealoma; primitive neuroectodermal tumor (PNT); plasmacytoma; paraneoplastic neurological syndrome; intraepithelial neoplasia; prostate cancer (e.g., prostatic adenocarcinoma); rectal cancer; rhabdomyosarcoma; salivary gland cancer; skin cancer (e.g., squamous cell carcinoma (SCC), keratoacanthoma (KA), melanoma, basal cell carcinoma (BCC) ));small intestine cancer (e.g. appendix cancer);soft tissue sarcomas (e.g. malignant fibrous histiocytoma (MFH), liposarcoma, malignant peripheral nerve sheath tumor (MPNST), chondrosarcoma, fibrosarcoma, myosarcoma);sebaceous gland carcinoma;small intestine cancer;sweat gland carcinoma;synoviomas;testicular cancer (e.g. seminoma, testicular embryonal carcinoma);thyroid cancer (e.g. papillary carcinoma of the thyroid, papillary thyroid carcinoma (PTC), medullary thyroid carcinoma);urethral cancer;vaginal cancer;vulvar cancer (e.g. Paget's disease of the vulva). ;
[0043] A "cell" as used herein may be present in a population of cells (e.g., in a tissue, sample, biopsy, organ, or organoid). In some embodiments, the cell population is composed of multiple different cell types. The cells used in the methods of the present disclosure may be present in an organism, in a single cell type derived from an organism, or in a mixture of cell types. This includes naturally occurring cells and cell populations, genetically engineered cell lines, cells derived from transgenic animals, and the like. Virtually any cell type and size may be adapted for the methods and systems described herein. Suitable cells include bacterial cells, fungal cells, plant cells, and animal cells. In some embodiments, the cells are mammalian cells (e.g., complex cell populations such as naturally occurring tissues). In some embodiments, the cells are of human origin. In some embodiments, the cells are collected from a subject (e.g., a human) through a medical procedure such as a biopsy. Alternatively, the cells may be a cultured population (e.g., a culture derived from a complex population, or a culture derived from a single cell type in which the cells have differentiated into multiple lineages). The cells may also be provided in situ in a cellular tissue.
[0044] Cell types contemplated for use in the methods of the present disclosure include, but are not limited to, stem and progenitor cells (e.g., embryonic stem cells, hematopoietic stem cells, mesenchymal stem cells, neural crest cells, etc.), endothelial cells, muscle cells, cardiomyocytes, smooth and skeletal muscle cells, mesenchymal cells, epithelial cells, hematopoietic cells, lymphocytes such as T cells (e.g., Th1 T cells, Th2 T cells, Th0 T cells, cytotoxic T cells), and B cells (e.g., pre-B cells), monocytes, dendritic cells, neutrophils, macrophages, natural killer cells, mast cells, adipocytes, immune cells, neurons, hepatocytes, and cells involved in specific organs (e.g., thymus, endocrine glands, pancreas, brain, neurons, glia, astrocytes, dendritic cells, and genetically modified cells thereof). The cells may also be different types of transformed or neoplastic cells (e.g., cancers of different cellular origins, lymphomas of different cell types, etc.), or any type of cancerous cell (e.g., derived from any of the cancers disclosed herein). Cells of different origins (e.g., ectoderm, mesoderm, and endoderm) are also contemplated for use in the methods of the present disclosure. In some embodiments, the cells are microglia, astrocytes, oligodendrocytes, excitatory neurons, or inhibitory neurons. In some embodiments, cells of multiple cell types are present in the same sample.
[0045] As used herein, the term "detection agent" refers to any agent that can be used to detect the presence or location of a protein or peptide of interest. In some embodiments, the methods disclosed herein include contacting one or more cells with one or more detection agents. Each detection agent used in the methods disclosed herein binds to a protein or peptide of interest. Detection agents that can be used in the methods described herein include, but are not limited to, proteins, peptides, nucleic acids, and small molecules. In some embodiments, the detection agent is an antibody that binds to a protein of interest. In some embodiments, the detection agent includes an antibody fragment, an antibody variant, and a nanobody. In some embodiments, the detection agent includes an aptamer. In some embodiments, the detection agent includes a receptor or a fragment thereof. In some embodiments, the detection agent includes a small molecule dye (e.g., small molecule X-34).
[0046] As used herein, the term "gene" refers to a nucleic acid fragment that expresses a protein, including regulatory sequences preceding (5' non-coding sequences) and following (3' non-coding sequences) the coding sequence. A "native gene" refers to a gene as found in nature with its own regulatory sequences.
[0047] As used herein, "gene expression" refers to the process by which information from a gene is used to synthesize a gene product. Gene products include proteins and RNA transcripts (e.g., messenger RNA, transfer RNA, or small nuclear RNA). Gene expression includes transcription and translation. Transcription is the process by which a segment of DNA is transcribed into RNA by RNA polymerase. Translation is the process by which RNA is translated into peptides or proteins by ribosomes. As used herein, the term "genetic information" refers to one or more genes and / or one or more RNA transcripts (e.g., any number of genes and / or RNA transcripts).
[0048] "Neurodegenerative disease" refers to a type of neurological disease characterized by loss of nerve cells, including, but not limited to, Alzheimer's disease, Parkinson's disease, amyotrophic lateral sclerosis, tauopathies (including frontotemporal dementia), and Huntington's disease. In some embodiments, the neurodegenerative disease is Alzheimer's disease. The causes of Alzheimer's disease are poorly understood, but are often believed to involve a genetic basis. The disease is characterized by loss of neurons and synapses in the cerebral cortex, resulting in atrophy of the affected areas. Biochemically, Alzheimer's disease is characterized as a protein misfolding disease caused by the accumulation of plaques of abnormally folded amyloid beta and tau proteins in the brain. Symptoms of Alzheimer's disease include, but are not limited to, difficulty remembering recent events, language problems, disorientation, mood swings, decreased motivation, self-neglect, and behavioral problems. Eventually, bodily functions are gradually lost and Alzheimer's disease eventually leads to death. Treatment is currently aimed at treating the cognitive problems caused by the disease (e.g., with acetylcholinesterase inhibitors or NMDA receptor antagonists), psychosocial interventions (e.g., behaviorally or cognitively oriented approaches), and general care. Currently, there is no treatment that completely halts or reverses the progression of the disease.
[0049] The terms "polynucleotide", "nucleotide sequence", "nucleic acid", "nucleic acid molecule", "nucleic acid sequence", and "oligonucleotide" refer to a series of nucleotide bases (also called "nucleotides") in DNA and RNA, and mean any chain of two or more nucleotides. Polynucleotides can be chimeric mixtures or derivatives or modified versions thereof, and single-stranded or double-stranded. Oligonucleotides can be modified at the base moiety, sugar moiety, or phosphate backbone to improve, for example, the stability of the molecule, its hybridization parameters, and the like.
[0050] A "protein," "peptide," or "polypeptide" comprises a polymer of amino acid residues linked together by peptide bonds. The term refers to proteins, polypeptides, and peptides of any size, structure, or function. Typically, a protein is at least three amino acids in length. A protein may refer to an individual protein or a collection of proteins. The proteins of the invention preferably contain only natural amino acids, but non-natural amino acids (i.e., compounds that do not occur in nature but can be incorporated into a polypeptide chain) and / or amino acid analogs known in the art may alternatively be used. Also, one or more amino acids in a protein may be modified, for example, by adding a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a chemical such as a linker for conjugation or functionalization, or other modification. A protein may be a single molecule or a multi-molecular complex. A protein may be a fragment of a naturally occurring protein or peptide. A protein may be natural, recombinant, synthetic, or any combination thereof.
[0051] "Pseudotime," as used herein, refers to a method for modeling differential expression of genes in cells, which is further described in Van den Berge, K. et al. Trajectory-based differential expression analysis for single-cell sequencing data. Nature Communications 2020, 11, 1-13, incorporated herein by reference.
[0052] A "transcript" or "RNA transcript" is the product resulting from RNA polymerase-catalyzed transcription of a DNA sequence. If the RNA transcript is a complementary copy of a DNA sequence, it is called the primary transcript, or it is an RNA sequence derived from post-transcriptional processing of the primary transcript, called the mature RNA. "Messenger RNA (mRNA)" refers to RNA that has no introns and can be translated into a polypeptide by the cell. "cRNA" refers to complementary RNA transcribed from a recombinant cDNA template. "cDNA" refers to DNA that is complementary to and derived from an mRNA template.
[0053] The term "sample" or "biological sample" refers to any sample, including tissue samples (e.g., tissue sections, surgical biopsies, and needle biopsies of tissues); cell samples (e.g., cytological smears (e.g., Pap or blood smears) or samples of cells obtained by microdissection); or cell fractions, fragments, or organelles (e.g., obtained by lysing cells and separating their components by centrifugation, etc.). Other examples of biological samples include, but are not limited to, blood, serum, urine, semen, feces, cerebrospinal fluid, interstitial fluid, mucus, tears, sweat, pus, biopsy tissue (e.g., obtained by surgical or needle biopsy), nipple aspirate, milk, vaginal fluid, saliva, swabs (e.g., oral swabs), or any material containing biomolecules derived from a first biological sample. In some embodiments, the biological sample is a surgical biopsy taken from a subject, e.g., a biopsy of any tissue described herein. In some embodiments, the biological sample is a tumor biopsy (e.g., from a subject diagnosed with, suspected of, or believed to have cancer). In some embodiments, the sample is brain tissue. In some embodiments, the tissue is heart tissue. In some embodiments, the tissue is muscle tissue.
[0054] A "subject" to which administration is contemplated refers to a human (i.e., male or female of any age group, e.g., a pediatric subject (e.g., infant, child, or adolescent) or an adult subject (e.g., young adult, middle-aged adult, or elderly adult)), or a non-human animal. In some embodiments, the non-human animal is a mammal (e.g., a primate (e.g., a cynomolgus or rhesus monkey) or a mouse). The term "patient" refers to a subject in need of treatment for a disease. In some embodiments, the subject is a human. In some embodiments, the patient is a human. A human can be male or female at any stage of development. Subjects or patients "in need" of treatment for a disease or disorder include, but are not limited to, those who exhibit any risk factors or symptoms of a disease or disorder (e.g., Alzheimer's disease). In some embodiments, the subject is a non-human experimental animal (e.g., a mouse).
[0055] A "therapeutically effective amount" of a treatment or therapeutic agent is an amount sufficient to provide a therapeutic benefit in the treatment of a condition or to delay or minimize one or more symptoms associated with the condition. A therapeutically effective amount of a treatment or therapeutic agent refers to an amount of a treatment that, alone or in combination with other treatments, provides a therapeutic effect in the treatment of a condition. The term "therapeutically effective amount" can include an amount that improves overall treatment, reduces or avoids symptoms, signs, or causes of a condition, and / or enhances the therapeutic effect of another therapeutic agent.
[0056] As used herein, a "tissue" is a group of cells from the same origin and their extracellular matrix. Together, the cells perform a specific function. Multiple tissue types combine together to form an organ. The cells may be of different cell types. In some embodiments, the tissue is an epithelial tissue. Epithelial tissue is formed by cells that cover the surfaces of organs (e.g., the surface of the skin, the respiratory tract, the soft organs, the reproductive tract, and the lining of the digestive tract). Epithelial tissue performs protective functions and is also involved in secretion, excretion, and absorption. Examples of epithelial tissue include, but are not limited to, simple squamous epithelium, stratified squamous epithelium, simple cuboidal epithelium, transitional epithelium, pseudostratified epithelium, columnar epithelium, and glandular epithelium. In some embodiments, the tissue is a connective tissue. Connective tissue is a fibrous tissue that is composed of cells separated by non-living substances (e.g., extracellular matrix). Connective tissue gives organs their shape and holds them in place. Connective tissue includes fibrous connective tissue, skeletal connective tissue, and fluid connective tissue. Examples of connective tissues include, but are not limited to, blood, bone, tendons, ligaments, fat, and areolar tissue. In some embodiments, the tissue is muscle tissue. Muscle tissue is an active contractile tissue formed from muscle cells. Muscle tissue functions to generate force and cause movement. Muscle tissue includes smooth muscle (e.g., found in the lining of organs), skeletal muscle (e.g., typically attached to bone), and cardiac muscle (e.g., found in the heart, which contracts to pump blood throughout an organism). In some embodiments, the tissue is nervous tissue. Nervous tissue includes cells that comprise the central and peripheral nervous systems. Nervous tissue forms the brain, spinal cord, cranial nerves, and spinal nerves (e.g., motor neurons). In some embodiments, the tissue is brain tissue. In some embodiments, the tissue is placental tissue. In some embodiments, the tissue is cardiac tissue.
[0057] The terms "treatment", "treat" and "treating" refer to reversing, alleviating, delaying onset or inhibiting progression of a disease described herein (e.g., Alzheimer's disease). In some embodiments, treatment may be administered after one or more signs or symptoms of a disease have developed or been observed (e.g., prophylactically (as may be further described herein) or based on suspicion or risk of disease). In other embodiments, treatment may be administered in the absence of signs or symptoms of a disease. For example, treatment may be administered to a susceptible subject prior to the onset of symptoms (e.g., taking into account the subject's or the subject's family's history of the condition). Treatment may also be continued after symptoms have resolved, e.g., to delay or prevent recurrence. In some embodiments, treatment may be administered after observing changes in gene and / or protein expression of one or more nucleic acids and / or proteins of interest in a cell or tissue compared to a healthy cell or tissue using the methods disclosed herein.
[0058] The terms "tumor" and "neoplasm" as used herein refer to an abnormal mass of tissue whose growth exceeds and is out of step with that of normal tissue. Tumors are "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation (including morphology and function), rate of growth, local invasion, and metastasis. "Benign neoplasms" are generally well differentiated and characterized by slower growth than malignant neoplasms, and remain localized at the site of origin. Furthermore, benign neoplasms do not have the ability to infiltrate, invade, or metastasize to distant sites. Exemplary benign neoplasms include, but are not limited to, lipomas, chondromas, adenomas, acrochordons, senile hemangiomas, seborrheic keratosis, lentigines, and sebaceous hyperplasia. In some cases, certain "benign" tumors may later give rise to malignant neoplasms, which may result from additional genetic changes in a subpopulation of neoplastic cells of the tumor, and these tumors are referred to as "premalignant neoplasms." An exemplary pre-malignant neoplasm is a teratoma. In contrast, a "malignant neoplasm" is generally poorly differentiated (anaplastic) and has a characteristic rapid growth accompanied by progressive infiltration, invasion, and destruction of surrounding tissue. In addition, a malignant neoplasm generally has the ability to metastasize to a distant site. The terms "metastasis", "metastatic", or "metastasizing" refer to the spread or migration of cancer cells from a primary or original tumor to another organ or tissue, typically identifiable by the presence of a "secondary tumor" or "secondary cell mass" of the histotype of the primary or original tumor, rather than the histotype of the organ or tissue in which the secondary (metastatic) tumor is located. For example, prostate cancer that has migrated to bone is said to be metastatic prostate cancer, which includes cancerous prostate cancer cells growing in bone tissue. Detailed Description of Specific Embodiments
[0059] The aspects described herein are not limited to specific embodiments, systems, compositions, methods, or configurations, as such may, of course, vary, and the terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting, unless specifically defined herein.
[0060] The present disclosure provides methods and systems for mapping gene and protein expression in cells (i.e., simultaneously mapping gene and protein expression in the same cell). The present disclosure also provides methods for diagnosing a disease or disorder (e.g., Alzheimer's disease or cancer) in a subject. Methods for screening or testing candidate agents capable of modulating gene and / or protein expression are also provided by the present disclosure. The present disclosure also provides methods for treating a disease or disorder, such as a neurological disorder (e.g., Alzheimer's disease), in a subject in need thereof. The present disclosure also describes pairs of oligonucleotide probes that may be useful for carrying out the methods described herein, and kits that include any of the oligonucleotide probes described herein. Additionally, the present disclosure provides methods, devices, systems, and non-transitory computer-readable storage media for identifying spatial variations of cell types in at least one image. Methods for mapping gene and protein expression in cells - Patents.com
[0061] In one aspect, the present disclosure provides a method for mapping gene and protein expression in cells (see, e.g., Figures 1A and 1B). In the methods disclosed herein, cells can be contacted with one or more pairs of oligonucleotide probes, which are further described herein, and can be used to amplify the transcripts (e.g., by rolling circle amplification) to generate one or more ligated amplicons. The cells can then be contacted with one or more detection agents (e.g., antibodies), where each detection agent binds to a protein of interest in the cell. The one or more ligated amplicons and the one or more detection agents can then be embedded in a polymer matrix, and the one or more amplicons can be sequenced to determine the identity of the transcripts (e.g., through SEDAL sequencing (sequencing with error reduction by dynamic annealing and ligation), which is further described herein), and their location within the polymer matrix. The location of detection agents (e.g., antibodies) bound to one or more proteins of interest within the polymer matrix can also be determined by imaging (e.g., by confocal microscopy), allowing the location of transcripts and proteins of interest within the cell to be mapped. The locations of transcripts and proteins of interest can be used to identify individual cells, subcellular locations, and organelles. This method can be useful, for example, to compare a cell (or multiple cells) from diseased and healthy tissue samples.
[0062] In some embodiments, the present disclosure provides a method for mapping gene and protein expression in a cell, comprising the steps of: a) contacting a cell with one or more pairs of oligonucleotide probes, where each pair of oligonucleotide probes comprises a first oligonucleotide probe and a second oligonucleotide probe, i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, where the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe; b) ligating together the 5' and 3' ends of a first oligonucleotide probe to generate a circular oligonucleotide; c) performing rolling circle amplification to amplify the circular oligonucleotide using the second oligonucleotide probe as a primer to generate one or more ligated amplicons; d) contacting the cells with one or more detection agents, where each detection agent binds to a protein of interest; e) embedding the one or more ligated amplicons and the one or more detection agents in a polymer matrix; f) contacting the one or more ligated amplicons embedded in the polymer matrix with a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe; and g) Imaging one or more ligated amplicons embedded in the polymer matrix and one or more detection agents embedded in the polymer matrix to determine the location of the nucleic acid of interest and the protein of interest within the cell, and optionally map gene and protein expression. In some embodiments, the expression of one nucleic acid and protein of interest is mapped using the methods described herein. In some embodiments, any of the methods described herein can be used to map the expression of multiple nucleic acids of interest and proteins of interest within the same cell (or multiple cells, e.g., in a tissue sample).
[0063] The use of any type of cell in the methods disclosed herein is contemplated by the present disclosure (e.g., any cell type described herein). In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a cell from the nervous system. In some embodiments, the cell is a cancer cell. The present disclosure also contemplates performing the methods described herein on multiple cells simultaneously. In some embodiments, the method is performed on multiple cells of the same cell type. In some embodiments, the method is performed on multiple cells, including cells of different cell types. Cell types for which gene and protein expression can be mapped using the methods disclosed herein include, but are not limited to, stem cells, progenitor cells, neural cells, astrocytes, dendritic cells, endothelial cells, microglia, oligodendrocytes, muscle cells, cardiomyocytes, mesenchymal cells, epithelial cells, immune cells, and hepatocytes. In some embodiments, the cell is a microglia, astrocyte, oligodendrocyte, excitatory neuron, and / or inhibitory neuron. In some embodiments, the cell(s) are permeabilized cells (e.g., the cells are permeabilized prior to contacting with one or more pairs of oligonucleotide probes). In some embodiments, the cell(s) are present within an intact tissue (e.g., any of the tissue types described herein). In some embodiments, the intact tissue is a fixed tissue sample. In some embodiments, the intact tissue comprises multiple cell types (e.g., microglia, astrocytes, oligodendrocytes, excitatory neurons, and / or inhibitory neurons). In some embodiments, the tissue is heart tissue, lymph node tissue, liver tissue, muscle tissue, bone tissue, eye tissue, or ear tissue.
[0064] The nucleic acid or nucleic acids of interest whose gene expression is profiled in the methods described herein can be the transcription product expressed from the genomic DNA of the cell. In some embodiments, the nucleic acid of interest is DNA. In some embodiments, the nucleic acid of interest is RNA. In some embodiments, the nucleic acid of interest is mRNA. The methods described herein can be used to profile the gene expression in cells for one nucleic acid of interest at a time or for multiple nucleic acids of interest simultaneously. In some embodiments, the gene expression in one cell or multiple cells is simultaneously mapped for more than 100, more than 200, more than 500, more than 1000, more than 2000, more than 3000, more than 5000, more than 10,000, more than 15,000, more than 20,000, more than 25,000, or more than 30,000 nucleic acids of interest. In some embodiments, the gene expression in one cell or multiple cells is simultaneously profiled for up to 1 million nucleic acids of interest.
[0065] The methods disclosed herein also contemplate the use of a first oligonucleotide probe and a second oligonucleotide probe provided as a pair of oligonucleotide probes. The first oligonucleotide probe (also referred to herein as a "padlock" probe) used in the methods described herein comprises a first barcode sequence and a second barcode sequence, each of which is composed of a specific sequence of nucleotides. In some embodiments, the first barcode sequence of the first oligonucleotide probe is about 5 to about 15, about 6 to about 14, about 7 to about 13, about 8 to about 12, or about 9 to about 11 nucleotides in length. In some embodiments, the first barcode sequence of the first oligonucleotide probe is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides in length. In some embodiments, the first barcode sequence of the first oligonucleotide probe is 10 nucleotides in length. In some embodiments, the second barcode sequence of the first oligonucleotide probe is about 5 to about 15, about 6 to about 14, about 7 to about 13, about 8 to about 12, or about 9 to about 11 nucleotides in length. In some embodiments, the second barcode sequence of the first oligonucleotide probe is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides in length. In some embodiments, the second barcode sequence of the first oligonucleotide probe is 10 nucleotides in length. The second barcode sequence provides additional complementary sites between the first and second oligonucleotide probes, providing an advantage over previous oligonucleotide probe designs. The second barcode sequence of the first oligonucleotide probe can increase the specificity of detection of the nucleic acid of interest in the methods described herein (i.e., compared to when the method is performed using a first oligonucleotide probe that does not include a second barcode). The second barcode sequence of the first oligonucleotide probe can also play a role in reducing non-specific amplification of the nucleic acid of interest in the methods described herein.
[0066] The barcode of the oligonucleotide probe described herein can comprise a gene-specific sequence that is used to identify the nucleic acid of interest.The use of barcode in the oligonucleotide probe described herein is further described in, for example, International Patent Application Publication No. WO2019 / 199579, and Wang et al., Science 2018, 361, 380, both of which are incorporated herein by reference in their entirety.
[0067] The first oligonucleotide probe also comprises a portion that is complementary to the second oligonucleotide probe and a portion that is complementary to the nucleic acid of interest.In some embodiments, the portion of the first oligonucleotide probe that is complementary to the second oligonucleotide probe is divided between the 5'-end and the 3'-end of the first oligonucleotide probe.In some embodiments, the portion of the first oligonucleotide probe that is complementary to the nucleic acid of interest is about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides long. In some embodiments, the first oligonucleotide probe is about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, or about 40 nucleotides in length. In some embodiments, the first oligonucleotide probe has the following structure: 5'-[part complementary to the second probe]-[part complementary to the nucleic acid of interest]-[first barcode sequence]-[second barcode sequence]-3'; where ]-[ includes any linker (e.g., any nucleotide linker). In some embodiments, ]-[ represents a direct bond (i.e., a phosphodiester bond) between two portions of the first oligonucleotide probe.
[0068] The second oligonucleotide probe (also referred to herein as a "primer" probe) used in the methods disclosed herein comprises a barcode sequence comprised of a specific sequence of nucleotides. In some embodiments, the barcode sequence of the second oligonucleotide probe is about 5 to about 15, about 6 to about 14, about 7 to about 13, about 8 to about 12, or about 9 to about 11 nucleotides in length. In some embodiments, the barcode sequence of the second oligonucleotide probe is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides in length. In some embodiments, the barcode sequence of the second oligonucleotide probe is 10 nucleotides in length.
[0069] The second oligonucleotide probe also comprises a portion that is complementary to the nucleic acid of interest and a portion that is complementary to the portion of the first oligonucleotide probe.In some embodiments, the first and second oligonucleotide probes are complementary to and bind to different portions of the nucleic acid of interest.In some embodiments, the portion of the second oligonucleotide probe that is complementary to the nucleic acid of interest is about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides long. In some embodiments, the second oligonucleotide probe is about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, or about 40 nucleotides in length. In some embodiments, the second oligonucleotide probe has the following structure: 5'-[part complementary to the nucleic acid of interest]-[part complementary to the first probe]-[barcode sequence]-3'; where ]-[ comprises an optional nucleotide linker. In some embodiments, ]-[ represents a direct bond between two portions of the second oligonucleotide probe.
[0070] The methods disclosed herein also include the use of a third oligonucleotide probe. In some embodiments, the third oligonucleotide probe comprises a detectable label (i.e., any label that can be used to visualize the location of the third oligonucleotide probe, for example, through imaging). In some embodiments, the detectable label is fluorescent (e.g., a fluorophore). As described herein, the third oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe. In some embodiments, the second barcode sequence of the first oligonucleotide probe is a gene-specific sequence that is used to identify the nucleic acid of interest. In some embodiments, the step of contacting one or more ligated amplicons embedded in a polymer matrix with a third oligonucleotide probe is performed to identify the nucleic acid of interest. This method for identifying nucleic acids of interest is known as sequencing with error reduction by dynamic annealing and ligation (SEDAL sequencing) and is further described in Wang, X. et al., Three-dimensional intact-tissue sequencing of single-cell transcriptional states. Science 2018, 36, eaat5691 and International Patent Application Publication No. WO 2019 / 199579, each of which is incorporated herein by reference.
[0071] The third oligonucleotide probe used in the methods described herein (e.g., as used in SEDAL sequencing) can be read out using any suitable imaging technique known in the art. For example, in embodiments where the third oligonucleotide probe comprises a fluorophore, the fluorophore can be read out using imaging to identify the nucleic acid of interest. As described above, the third oligonucleotide probe comprises a sequence complementary to the second barcode sequence of the first oligonucleotide probe, which is used to detect the specific nucleic acid of interest. By imaging the location of the third oligonucleotide probe comprising the fluorophore, the location of the specific nucleic acid of interest within the sample can be determined. In some embodiments, the imaging step comprises fluorescent imaging. In some embodiments, the imaging step comprises confocal microscopy. In some embodiments, the imaging step comprises epifluorescence microscopy. In some embodiments, the location of the nucleic acid of interest and the protein of interest are determined in the same round of imaging. In some embodiments, the location of the nucleic acid of interest and the protein of interest are determined in another round of imaging. In some embodiments, two rounds of imaging are performed. In some embodiments, three rounds of imaging are performed. In some embodiments, four rounds of imaging are performed. In some embodiments, 5 or more rounds of imaging are performed.
[0072] In some embodiments, the methods disclosed herein include contacting the cells with one or more detection agents, where each detection agent binds to a protein of interest. In some embodiments, the detection agent is a protein, peptide, nucleic acid, or small molecule. In some embodiments, the detection agent is an antibody. In some embodiments, the detection agent is an antibody fragment, antibody variant, or nanobody. In some embodiments, the detection agent is an aptamer. In some embodiments, the detection agent is a receptor or a fragment thereof. The use of any antibody that binds to a protein of interest is contemplated by the present disclosure. Antibodies that may be used in the methods described herein also include antibody fragments, as well as variants of full length antibodies or antibody fragments. In some embodiments, the one or more detection agents are antibodies, each of which includes a unique detectable label (e.g., a fluorophore). In some embodiments, the method further includes contacting the one or more antibodies with a secondary detection agent. For example, each antibody that binds to a protein of interest can be contacted with a different secondary detection agent. In some embodiments, the secondary detection agent is a secondary antibody (i.e., an antibody that binds to one of the antibodies bound to the protein of interest). In some embodiments, the secondary antibody includes a detectable label. In some embodiments, the detectable label of the secondary antibody is a fluorophore.In some embodiments, one or more detection agents are antibodies that each bind to a protein of interest, and each antibody that binds to a protein of interest is conjugated to a unique oligonucleotide sequence.Then, one or more antibodies can be contacted with an oligonucleotide that is conjugated to a detectable label (e.g., a fluorophore) that is complementary to the oligonucleotide sequence that is conjugated to one or more antibodies, and the location of one or more antibodies can be visualized (e.g., by confocal microscopy or other means for detecting fluorescence).
[0073] In some embodiments, the detection agent that binds to the protein of interest is a small molecule dye. In some embodiments, the small molecule dye is X-34. For example, X-34 (an amyloid-specific fluorescent dye that is commonly used as a highly fluorescent marker for β-sheet structure) can be used to detect the presence of Aβ plaques in cells and determine pathological changes associated with Alzheimer's disease. The use of X-34 in detecting Aβ plaques is well known in the art and is described, for example, in Styren, SD et al. X-34, a fluorescent derivative of Congo red: a novel histochemical stain for Alzheimer's disease pathology. J. Histochem. Cytochem. 2000, 48(9), 1223-1232, which is incorporated herein by reference. In some embodiments, the step of contacting each of the one or more detection agents embedded in the polymer matrix with a secondary detection agent is performed after the step of performing rolling circle amplification to amplify the circular oligonucleotide. In some embodiments, the step of contacting the cell with one or more detection agents is performed before the step of embedding.
[0074] The methods provided herein can be used to map gene and protein expression in cells at subcellular resolution.For example, any of the methods provided herein can be carried out at a subcellular resolution of 200nm, 150nm, 100nm, 50nm, 40nm, 30nm, 20nm or 10nm.In some embodiments, any of the methods provided herein is carried out at a subcellular resolution of 200nm.
[0075] The use of various polymer matrices is contemplated in the present disclosure, and any polymer matrix that can embed one or more linked amplicons and detection agents is suitable for use in the methods described herein.In some embodiments, the polymer matrix is a hydrogel (i.e., a network of cross-linked polymers that are hydrophilic).In some embodiments, the hydrogel is a polyvinyl alcohol hydrogel, a polyethylene glycol hydrogel, a sodium polyacrylate hydrogel, an acrylic acid polymer hydrogel, or a polyacrylamide hydrogel.In some embodiments, the hydrogel is a polyacrylamide hydrogel.Such hydrogels can be prepared, for example, by incubating a sample in a buffer solution that includes acrylamide and bisacrylamide, removing the buffer solution, and incubating the sample in a polymerization mixture that includes, for example, ammonium persulfate and tetramethylethylenediamine.
[0076] In some embodiments, performing rolling circle amplification to amplify the circular oligonucleotides to generate one or more linked amplicons further comprises providing nucleotides modified with a reactive chemical group (e.g., 5-(3-aminoallyl)-dUTP). In some embodiments, the nucleotides modified with a reactive chemical group constitute about 5%, about 6%, about 7%, about 8%, about 9%, or about 10% of the nucleotides used in the amplification reaction. For example, performing rolling circle amplification to amplify the circular oligonucleotides to generate one or more linked amplicons may further comprise providing amine-modified nucleotides. During the amplification process, the amine-modified nucleotides are incorporated into one or more linked amplicons as they are generated. The resulting amplicons are functionalized with a primary amine, which can be further reacted with another compatible chemical moiety (e.g., N-hydroxysuccinimide) to facilitate embedding the linked amplicons into a polymer matrix. In some embodiments, the step of embedding the one or more linked amplicons in a polymer matrix comprises reacting amine-modified nucleotides of the one or more linked amplicons with an acrylic acid N-hydroxysuccinimide ester to copolymerize the one or more linked amplicons with the polymer matrix.
[0077] In an embodiment, the present disclosure provides a method for mapping gene expression in a cell, comprising the steps of: a) contacting a cell with one or more pairs of oligonucleotide probes, where each pair of oligonucleotide probes comprises a first oligonucleotide probe and a second oligonucleotide probe, i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, where the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe; b) ligating together the 5' and 3' ends of a first oligonucleotide probe to generate a circular oligonucleotide; c) performing rolling circle amplification to amplify the circular oligonucleotide using the second oligonucleotide probe as a primer to generate one or more ligated amplicons; d) embedding one or more ligated amplicons into a polymer matrix; e) contacting the one or more ligated amplicons embedded in the polymer matrix with a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe; and f) Imaging one or more ligated amplicons embedded in the polymer matrix to determine the location of the nucleic acid of interest within the cell and, optionally, to map gene and protein expression.
[0078] In an embodiment, the present disclosure provides a method for mapping gene and protein expression in a cell, comprising the steps of: a) contacting a cell with one or more pairs of oligonucleotide probes, where each pair of oligonucleotide probes comprises a first oligonucleotide probe and a second oligonucleotide probe, i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest and a portion complementary to the first oligonucleotide probe; b) ligating together the 5' and 3' ends of a first oligonucleotide probe to generate a circular oligonucleotide; c) performing rolling circle amplification to amplify the circular oligonucleotide using the second oligonucleotide probe as a primer to generate one or more ligated amplicons; d) contacting the cells with one or more detection agents, where each detection agent binds to a protein of interest; e) embedding the one or more ligated amplicons and the one or more detection agents in a polymer matrix; f) contacting the one or more ligated amplicons embedded in the polymer matrix with a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe; and g) Imaging the one or more ligated amplicons embedded in the polymer matrix and the one or more detection agents embedded in the polymer matrix to determine the location of the nucleic acid of interest and the protein of interest within the cell, and optionally, to map gene and protein expression. Methods for diagnosing a disease or disorder in a subject
[0079] In another aspect, the present disclosure provides a method for diagnosing a disease or disorder in a subject. For example, the method of profiling gene and protein expression described herein can be performed in one or more cells from a sample taken from a subject (e.g., a subject believed to have or at risk of having a disease or disorder, or a subject believed to be healthy or healthy). The expression of various nucleic acids and proteins of interest in the cells can then be compared to the expression of the same nucleic acids and proteins of interest in non-disease cells or cells from a non-disease tissue sample (e.g., a cell from a healthy individual, or a plurality of cells from a population of healthy individuals). Any change in the expression of a nucleic acid and / or protein of interest (or a plurality of nucleic acids and / or proteins of interest, e.g., a particular disease signature) compared to the expression in a non-disease cell indicates that the subject has a disease or disorder. Gene and protein expression in one or more non-disease cells can be profiled in parallel with the expression in a diseased cell as a control experiment. Gene and protein expression in one or more non-disease cells can also have been profiled previously, and the expression in a diseased cell can be compared to this reference data for a non-disease cell.
[0080] In some embodiments, the method for diagnosing a disease or disorder in a subject comprises the steps of: a) contacting a cell taken from the subject with one or more pairs of oligonucleotide probes, wherein each pair of oligonucleotide probes comprises a first oligonucleotide probe and a second oligonucleotide probe, i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, where the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe; b) ligating together the 5' and 3' ends of a first oligonucleotide probe to generate a circular oligonucleotide; c) performing rolling circle amplification to amplify the circular oligonucleotide using the second oligonucleotide probe as a primer to generate one or more ligated amplicons; d) contacting the cells with one or more detection agents, where each detection agent binds to a protein of interest; e) embedding the one or more ligated amplicons and the one or more detection agents in a polymer matrix; f) contacting the one or more ligated amplicons embedded in the polymer matrix with a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe; and g) imaging the one or more ligated amplicons embedded in the polymer matrix and the one or more detection agents embedded in the polymer matrix to determine the location of the nucleic acid of interest and the protein of interest within the cell, and optionally to map gene and protein expression; Here, alteration in expression of the nucleic acid of interest and / or protein of interest, relative to expression in one or more non-diseased cells, indicates that the subject has the disease or disorder.
[0081] In some embodiments, the gene and protein expression in one or more non-disease cells is profiled at the same time using the method disclosed herein as a control experiment.In some embodiments, the gene and protein expression data in one or more non-disease cells that is compared with the expression in disease cell comprises the reference data from the previous implementation of the method on one or more non-disease cells.
[0082] Any disease or disorder diagnosis is contemplated by the method described herein.In some embodiments, the disease or disorder is genetic disease, proliferation disease, inflammatory disease, autoimmune disease, liver disease, spleen disease, lung disease, blood disease, neurological disease, gastrointestinal (GI) tract disease, urogenital disease, infectious disease, musculoskeletal disease, endocrine disease, metabolic disorder, immune disorder, central nervous system (CNS) disorder, neurological disorder, ophthalmological disease, or cardiovascular disease.In some embodiments, the disease or disorder is neurodegenerative disease.In some embodiments, the disease or disorder is Alzheimer's disease.In some embodiments, the disease or disorder is cancer.
[0083] In some embodiments, the cell is present in a tissue. In some embodiments, the tissue is a tissue sample from a subject. In some embodiments, the subject is a non-human experimental animal (e.g., a mouse). In some embodiments, the subject is a farm animal. In some embodiments, the subject is a human. In some embodiments, the tissue sample comprises a fixed tissue sample. In some embodiments, the tissue sample is a biopsy (e.g., a bone, bone marrow, breast, gastrointestinal tract, lung, liver, pancreas, prostate, brain, nerve, kidney, endometrial, cervical, lymph node, muscle, or skin biopsy). In some embodiments, the biopsy is a tumor biopsy. In some embodiments, the tissue is brain tissue. In some embodiments, the tissue is from the central nervous system.
[0084] The use of any type of cell in the methods disclosed herein for diagnosing a disease or disorder in a subject is contemplated by the present disclosure. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a cancer cell. The present disclosure also contemplates performing the methods described herein on multiple cells simultaneously. In some embodiments, the method is performed on multiple cells of the same cell type. In some embodiments, the method is performed on multiple cells of different cell types. Cell types for which gene and protein expression can be mapped using the methods disclosed herein include, but are not limited to, stem cells, progenitor cells, neural cells, astrocytes, dendritic cells, endothelial cells, microglia, oligodendrocytes, muscle cells, cardiomyocytes, mesenchymal cells, epithelial cells, immune cells, and hepatocytes. In some embodiments, the cell is a microglia, astrocyte, oligodendrocyte, excitatory neuron, and / or inhibitory neuron.
[0085] A variety of nucleic acids and proteins of interest can be profiled using the methods disclosed herein. For example, the expression of any nucleic acid or protein known or believed to be associated with a disease or disorder (e.g., Alzheimer's disease) can be mapped using the methods disclosed herein and used to diagnose the disease or disorder. In some embodiments, the nucleic acid of interest is selected from the group consisting of Vsnl1, Snap25, Dnm1, Slc6a1, Aldoc, Bsg, Ctss, Plp1, Cst7, Ctsb, Apoe, Trem2, C1qa, P2ry12, Gfap, Vim, Aqp4, Clu, Plp1, Mbp, C4b, Ccnb2, Gpm6a, Ddit3, Dapk1, Myo5a, Tspan7, and Rhoc. In some embodiments, the protein of interest comprises an amyloid beta (Aβ) peptide. In some embodiments, the Aβ peptide is present in the form of Aβ plaques. In some embodiments, the protein of interest comprises Tau protein. In some embodiments, Tau protein exists in the form of inclusions (p-Tau) in cells. The presence of Aβ plaques and / or p-Tau can be used to diagnose, for example, neurodegenerative diseases (e.g., Alzheimer's disease) in subjects.
[0086] The use of various detection agents to detect one or more proteins of interest is contemplated by the present disclosure. In some embodiments, when the one or more proteins of interest include Aβ peptides, the Aβ peptides are detected using a small molecule detection agent (e.g., a small molecule fluorescent dye). In some embodiments, the small molecule detection agent is X-34, as described herein and in, for example, Styren, SD et al. X-34, a fluorescent derivative of Congo red: a novel histochemical stain for Alzheimer's disease pathology. J. Histochem. Cytochem. 2000, 48(9), 1223-1232, which is incorporated herein by reference. In some embodiments, when the proteins of interest include Tau protein or p-Tau, the Tau protein is detected using a p-Tau primary antibody. In some embodiments, the method further includes detecting the p-Tau primary antibody with a secondary antibody. In some embodiments, the secondary antibody is conjugated to a detectable label (e.g., a fluorophore).
[0087] Using the methods disclosed herein for diagnosing a disease or disorder in a subject, examining changes in expression of a nucleic acid of interest and / or a protein of interest compared to expression in one or more non-disease cells can indicate that the subject has a disease or disorder. In some embodiments, changes in expression of a nucleic acid of interest and / or a protein of interest are used to identify cell types in close proximity to plaques. For example, if a particular cell type is identified in close proximity to plaques, the subject may have or may be suspected of having Alzheimer's disease. In some embodiments, the plaques are Aβ plaques. In some embodiments, identification of disease-related microglial cell types in close proximity to plaques (e.g., including high expression of C1qa, Trem2, Cst7, Ctsb, and / or Apoe) indicates that the subject has or is at risk of having Alzheimer's disease. In some embodiments, identification of disease-related astrocytic cell types in close proximity to plaques (e.g., including high expression of Gfap, Vim, and / or Apoe) indicates that the subject has or is at risk of having Alzheimer's disease. In some embodiments, identification of oligodendrocyte precursor cell types proximate to plaques (e.g., including high expression of Cldn11, Klk6, Serpina3n, and / or C4b) indicates that the subject has or is at risk for having Alzheimer's disease. Cells proximate to plaques can include cells that are within about 10 μm, within about 20 μm, within about 30 μm, within about 40 μm, or within about 50 μm of the plaque.
[0088] The methods disclosed herein can also be used to map proteins modified with a post-translational modification of interest (e.g., phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation, lipidation, etc.). Studying the location of such modified proteins relative to specific cell types may provide a useful tool, for example, in cancer research, diagnosis, and treatment, since protein modifications such as phosphorylation play important roles during cancer development and progression. Methods for screening agents capable of modulating gene and / or protein expression
[0089] In another aspect, the present disclosure provides a method for screening agents capable of modulating gene and / or protein expression of a nucleic acid or protein of interest, or multiple nucleic acids and / or proteins of interest. For example, the methods for mapping gene and protein expression described herein can be performed in cells in the presence of one or more candidate agents. Expression of various nucleic acids and / or proteins of interest in cells (e.g., normal cells or diseased cells) can then be compared to expression of the same nucleic acids and / or proteins of interest in cells that have not been exposed to one or more candidate agents. Any change in expression of the nucleic acid(s) and / or protein(s) of interest compared to expression in cells that have not been exposed to the candidate agent(s) indicates that expression of the nucleic acid(s) and / or protein(s) of interest is modulated by the candidate agent(s). In some embodiments, certain signatures (e.g., expression of nucleic acids and proteins) known to be associated with disease treatment can be used to identify candidate agents that can modulate gene and / or protein expression in a desired manner. The methods described herein can also be used to identify drugs that have particular side effects, for example, by looking for specific nucleic acid and protein expression signatures when one or more cells are treated with a candidate agent or known drug.
[0090] In some embodiments, the present disclosure provides a method for screening for agents capable of modulating gene and / or protein expression, comprising the steps of: a) contacting a cell being treated or having been treated with a candidate agent with one or more pairs of oligonucleotide probes, wherein each pair of oligonucleotide probes comprises a first oligonucleotide probe and a second oligonucleotide probe, i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, where the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe; b) ligating together the 5' and 3' ends of a first oligonucleotide probe to generate a circular oligonucleotide; c) performing rolling circle amplification to amplify the circular oligonucleotide using the second oligonucleotide probe as a primer to generate one or more ligated amplicons; d) contacting the cells with one or more detection agents, where each detection agent binds to a protein of interest; e) embedding the one or more ligated amplicons and the one or more detection agents in a polymer matrix; f) contacting the one or more ligated amplicons embedded in the polymer matrix with a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe; and g) imaging the one or more ligated amplicons embedded in the polymer matrix and the one or more detection agents embedded in the polymer matrix to determine the location of the nucleic acid of interest and the protein of interest within the cell, and optionally to map gene and protein expression; Here, a change in expression of the nucleic acid of interest and / or protein of interest in the presence of the candidate agent compared to expression in the absence of the candidate agent indicates that the candidate agent modulates gene and / or protein expression.
[0091] In some embodiments, the candidate agent is a small molecule, a protein, a peptide, a nucleic acid, a lipid, or a carbohydrate. In some embodiments, the small molecule is an anti-cancer therapeutic agent. In some embodiments, the small molecule comprises a known drug. In some embodiments, the small molecule comprises an FDA-approved drug. In some embodiments, the protein is an antibody. In some embodiments, the protein is an antibody fragment or antibody variant. In some embodiments, the protein is a receptor. In some embodiments, the protein is a cytokine. In some embodiments, the nucleic acid is an mRNA, an antisense RNA, a miRNA, an siRNA, an RNA aptamer, a double-stranded RNA (dsRNA), a short hairpin RNA (shRNA), or an antisense oligonucleotide (ASO). Any candidate agent can be screened using the methods described herein. In particular, any candidate agent that is believed to be capable of modulating gene and / or protein expression can be screened using the methods described herein. In some embodiments, the modulation of gene and / or protein expression by a candidate agent is associated with reducing, alleviating, or eliminating symptoms of a disease or disorder, or preventing the onset or progression of a disease or disorder. In some embodiments, the disease or disorder regulated by the candidate agent is genetic disease, proliferative disease, inflammatory disease, autoimmune disease, liver disease, spleen disease, lung disease, blood disease, neurological disease, gastrointestinal (GI) tract disease, urogenital disease, infectious disease, musculoskeletal disease, endocrine disease, metabolic disorder, immune disorder, central nervous system (CNS) disorder, neurological disorder, ophthalmological disease, or cardiovascular disease.In some embodiments, the disease or disorder is Alzheimer's disease.In some embodiments, the disease or disorder is cancer. Methods for Treating a Disease or Disorder in a Subject
[0092] In another aspect, the present disclosure provides a method for treating a disease or disorder in a subject. For example, the method for profiling gene and protein expression described herein can be performed on cells from a sample taken from a subject (e.g., a subject who is suspected of having and at risk of having a disease or disorder). The expression of various nucleic acids and / or proteins of interest in the cells can then be compared with the expression of the same nucleic acids and / or proteins of interest in cells from a non-disease tissue sample. If any changes in the expression of the nucleic acids and / or proteins of interest are observed compared to the expression in non-disease cells, treatment for the disease or disorder can then be administered to the subject. Gene and protein expression in one or more non-disease cells can be profiled in parallel with the expression in disease cells as a control experiment. Gene and protein expression in one or more non-disease cells may have been profiled previously, and the expression in disease cells can be compared to this reference data for non-disease cells.
[0093] In some embodiments, the present disclosure provides a method for treating a disease or disorder in a subject, comprising the steps of: a) contacting a cell taken from the subject with one or more pairs of oligonucleotide probes, wherein each pair of oligonucleotide probes comprises a first oligonucleotide probe and a second oligonucleotide probe, i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, where the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe; b) ligating together the 5' and 3' ends of a first oligonucleotide probe to generate a circular oligonucleotide; c) performing rolling circle amplification to amplify the circular oligonucleotide using the second oligonucleotide probe as a primer to generate one or more ligated amplicons; d) contacting the cells with one or more detection agents, where each detection agent binds to a protein of interest; e) embedding the one or more ligated amplicons and the one or more detection agents in a polymer matrix; f) contacting the one or more ligated amplicons embedded in the polymer matrix with a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe; and g) imaging the one or more ligated amplicons embedded in the polymer matrix and the one or more detection agents embedded in the polymer matrix to determine the location of the nucleic acid of interest and the protein of interest within the cell, and optionally to map gene and protein expression; and h) treating the subject for a disease or disorder if altered expression of the nucleic acid of interest and / or protein of interest is observed compared to expression in one or more non-diseased cells.
[0094] In some embodiments, the gene and protein expression in one or more non-disease cells is profiled at the same time using the method disclosed herein as a control experiment.In some embodiments, the gene and protein expression data in one or more non-disease cells compared with the expression in disease cells comprises reference data from when the method is previously performed on non-disease cells.
[0095] Any suitable treatment for the disease or disorder can be administered to the subject. In some embodiments, the treatment includes administering a therapeutic agent. In some embodiments, the treatment includes surgery. In some embodiments, the treatment includes imaging. In some embodiments, the treatment includes performing an additional diagnostic method. In some embodiments, the treatment includes radiation therapy. In some embodiments, the therapeutic agent is a small molecule, a protein, a peptide, a nucleic acid, a lipid, or a carbohydrate. In some embodiments, the small molecule is an anti-cancer therapeutic agent. In some embodiments, the small molecule is a known drug. In some embodiments, the small molecule is an FDA-approved drug. In some embodiments, the small molecule comprises a library of compounds. In some embodiments, the protein is an antibody. In some embodiments, the protein is an antibody fragment. In some embodiments, the protein is an antibody variant. In some embodiments, the protein is a receptor. In some embodiments, the protein is a cytokine. In some embodiments, the nucleic acid is an mRNA, an antisense RNA, a miRNA, an siRNA, an RNA aptamer, a double-stranded RNA (dsRNA), a short hairpin RNA (shRNA), or an antisense oligonucleotide (ASO). In some embodiments, the nucleic acid is DNA. In some embodiments, the protein is an antibody. In some embodiments, the protein is an antibody fragment or antibody variant. In some embodiments, the protein is a receptor, or a fragment or variant thereof. In some embodiments, the protein is a cytokine. In some embodiments, the nucleic acid is an mRNA, an antisense RNA, an miRNA, an siRNA, an RNA aptamer, a double-stranded RNA (dsRNA), a short hairpin RNA (shRNA), or an antisense oligonucleotide (ASO).
[0096] Any disease or disorder is intended to be treated by the method described herein.In some embodiments, the disease or disorder is genetic disease, proliferation disease, inflammatory disease, autoimmune disease, liver disease, spleen disease, lung disease, blood disease, neurological disease, gastrointestinal (GI) tract disease, urogenital disease, infectious disease, musculoskeletal disease, endocrine disease, metabolic disorder, immune disorder, central nervous system (CNS) disorder, neurological disorder, ophthalmological disease, or cardiovascular disease.In some embodiments, the disease or disorder is neurodegenerative disease.In some embodiments, the disease or disorder is Alzheimer's disease.In some embodiments, the disease or disorder is cancer.
[0097] In some embodiments, the subject is a human. In some embodiments, the sample comprises a biological sample. In some embodiments, the sample comprises a tissue sample. In some embodiments, the tissue sample is a biopsy (e.g., a bone, bone marrow, breast, gastrointestinal tract, lung, liver, pancreas, prostate, brain, nerve, kidney, endometrial, cervical, lymph node, muscle, or skin biopsy). In some embodiments, the biopsy is a tumor biopsy. In some embodiments, the biopsy is a solid tumor biopsy. In some embodiments, the tissue sample is a brain tissue sample. In some embodiments, the tissue sample is a central nervous system tissue sample. Oligonucleotide Probes
[0098] The present disclosure also provides oligonucleotide probes for use in the methods and systems described herein.
[0099] In one aspect, the present disclosure provides a plurality of oligonucleotide probes including a first oligonucleotide probe (also referred to herein as a "padlock" probe) and a second oligonucleotide probe (also referred to herein as a "primer" probe), wherein: i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence; Here, the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe.
[0100] All oligonucleotide probes described herein may optionally have spacers or linkers of various nucleotide lengths between each of the recited components, or the components of the oligonucleotide probe may be directly linked to each other. All oligonucleotide probes described herein may contain standard nucleotides, or some of the standard nucleotides may be replaced with any modified nucleotide known in the art.
[0101] A first oligonucleotide probe (also referred to herein as a "padlock" probe) of the plurality of oligonucleotide probes described herein comprises a first barcode sequence and a second barcode sequence, each of which is composed of a particular sequence of nucleotides. In some embodiments, the first barcode sequence of the first oligonucleotide probe is about 5 to about 15, about 6 to about 14, about 7 to about 13, about 8 to about 12, or about 9 to about 11 nucleotides in length. In some embodiments, the first barcode sequence of the first oligonucleotide probe is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides in length. In some embodiments, the first barcode sequence of the first oligonucleotide probe is 10 nucleotides in length. In some embodiments, the second barcode sequence of the first oligonucleotide probe is about 5 to about 15, about 6 to about 14, about 7 to about 13, about 8 to about 12, or about 9 to about 11 nucleotides in length. In some embodiments, the second barcode sequence of the first oligonucleotide probe is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides long. In some embodiments, the second barcode sequence of the first oligonucleotide probe is 10 nucleotides long. The second barcode sequence provides additional complementary sites between the first and second oligonucleotide probes, providing an advantage over previous oligonucleotide probe designs. The second barcode sequence of the first oligonucleotide probe can increase the specificity of detection of the target nucleic acid using multiple oligonucleotide probes (e.g., using multiple oligonucleotide probes in any method disclosed herein). The second barcode sequence of the first oligonucleotide probe can also play a role in reducing non-specific amplification of the target nucleic acid.
[0102] The barcode of the oligonucleotide probe described herein can comprise a gene-specific sequence that is used to identify the nucleic acid of interest.The use of barcode in the oligonucleotide probe described herein is further described in, for example, International Patent Application Publication No. WO2019 / 199579, and Wang et al., Science 2018, 361, 380, both of which are incorporated herein by reference in their entirety.
[0103] The first oligonucleotide probe also comprises a portion that is complementary to the second oligonucleotide probe and a portion that is complementary to the nucleic acid of interest.In some embodiments, the portion of the first oligonucleotide probe that is complementary to the second oligonucleotide probe is divided between the 5'-end and the 3'-end of the first oligonucleotide probe.In some embodiments, the portion of the first oligonucleotide probe that is complementary to the nucleic acid of interest is about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides long. In some embodiments, the first oligonucleotide probe is about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, or about 40 nucleotides in length. In some embodiments, the first oligonucleotide probe has the following structure: 5'-[part complementary to the second probe]-[part complementary to the nucleic acid of interest]-[first barcode sequence]-[second barcode sequence]-3,' where ]-[ includes any linker (e.g., any nucleotide linker). In some embodiments, ]-[ represents a direct bond (i.e., a phosphodiester bond) between two portions of the first oligonucleotide probe.
[0104] The second oligonucleotide probe (also referred to herein as a "primer" probe) of the plurality of oligonucleotide probes disclosed herein comprises a barcode sequence comprised of a particular sequence of nucleotides. In some embodiments, the barcode sequence of the second oligonucleotide probe is about 5 to about 15, about 6 to about 14, about 7 to about 13, about 8 to about 12, or about 9 to about 11 nucleotides in length. In some embodiments, the barcode sequence of the second oligonucleotide probe is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides in length. In some embodiments, the barcode sequence of the second oligonucleotide probe is 10 nucleotides in length.
[0105] The second oligonucleotide probe also comprises a portion that is complementary to the nucleic acid of interest and a portion that is complementary to the portion of the first oligonucleotide probe.In some embodiments, the first and second oligonucleotide probes are complementary to and bind to different portions of the nucleic acid of interest.In some embodiments, the portion of the second oligonucleotide probe that is complementary to the nucleic acid of interest is about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides long. In some embodiments, the second oligonucleotide probe is about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, or about 40 nucleotides in length. In some embodiments, the second oligonucleotide probe has the following structure: 5'-[part complementary to the nucleic acid of interest]-[part complementary to the first probe]-[barcode sequence]-3'; where ]-[ includes any linker (e.g., any nucleotide linker). In some embodiments, ]-[ represents a direct bond (i.e., a phosphodiester bond) between the two portions of the second oligonucleotide probe.
[0106] In some embodiments, the plurality of oligonucleotide probes comprises a third oligonucleotide probe. In some embodiments, the third oligonucleotide probe comprises a detectable label. In some embodiments, the detectable label is a fluorophore. As described herein, the third oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe. In some embodiments, the second barcode sequence of the first oligonucleotide probe is a gene-specific sequence that is used to identify the nucleic acid of interest. In some embodiments, the third oligonucleotide probe is used to identify the nucleic acid of interest (e.g., through SEDAL sequencing). kit
[0107] The present disclosure also provides kits. In one aspect, the kits provided may include one or more oligonucleotide probes as described herein. In some embodiments, the kits may further include a container (e.g., a vial, an ampoule, a bottle, and / or a dispenser package, or other suitable container). In some embodiments, the kits include a plurality of oligonucleotide probes as described herein. In some embodiments, the kits further include one or more detection agents (e.g., a peptide, a protein such as an antibody, a nucleic acid such as an aptamer, or a small molecule such as a fluorescent dye), where each detection agent binds to a protein of interest. In some embodiments, the one or more detection agents include an anti-p-Tau antibody. In some embodiments, the kits further include a small molecule detection agent (e.g., X-34 as described herein). In some embodiments, the kits further include a third oligonucleotide probe as described herein. The third oligonucleotide probe may include a sequence complementary to the first barcode sequence of the first oligonucleotide probe. In some embodiments, the third oligonucleotide probe includes a detectable label (e.g., a fluorophore). In some embodiments, the kits include a library of two or more sets of oligonucleotide probes, where each set of oligonucleotide probes is used to identify a particular nucleic acid of interest. In some embodiments, the kit comprises a library of detection agents for detecting multiple proteins of interest. In some embodiments, the kit may further comprise other reagents for carrying out the methods disclosed herein, such as cells, ligases, polymerases, amine-modified nucleotides, primary antibodies, secondary antibodies, buffers, and / or reagents for making polymer matrices (e.g., polyacrylamide matrices). In some embodiments, the kit is useful for profiling gene and protein expression in cells. In some embodiments, the kit is useful for diagnosing disease in a subject (e.g., Alzheimer's disease). In some embodiments, the kit is useful for screening agents capable of modulating gene and / or protein expression.In some embodiments, the kits are useful for diagnosing a disease or disorder in a subject. In some embodiments, the kits are useful for treating a disease or disorder in a subject. In some embodiments, the kits described herein further comprise instructions for using the kit. Method, apparatus, and non-transitory computer-readable storage medium for identifying spatial variation of cell types within at least one image - Patents.com
[0108] In various aspects, the present disclosure provides methods for identifying spatial variation of cell types in one or more images (i.e., examining the variation in the relative spatial distribution of a particular cell type among a plurality of samples, e.g., comparing healthy tissue to diseased tissue). In some embodiments, the one or more images are acquired using any of the methods disclosed herein. FIG. 15 is a flow diagram of one embodiment of a method relating to identifying spatial variation of cell types in at least one image. The process flow 1500 may be performed by at least one computer processor. In some embodiments, there may be at least one non-transitory computer-readable storage medium encoded with a plurality of instructions that perform the process flow 1500 when executed by the at least one computer processor. The process flow 1500 includes steps 1502, 1504, 1506, 1508, and 1510. In step 1502, the at least one computer processor receives, for each of a plurality of cells in the at least one image, a spatial location of the cell in the at least one image. In step 1504, the at least one computer processor receives, for each of a plurality of proteins in the at least one image, a spatial location of the protein in the image. In step 1506, the at least one computer processor determines, for a first protein of the plurality of proteins, a number of cells of the first cell type that are less than a threshold distance to the first protein, where the distance is determined based on at least some of the spatial locations of the plurality of cells and at least some of the spatial locations of the plurality of proteins. In step 1508, the at least one computer processor identifies spatial variability of cells of the first cell type in the at least one image based on the number of cells of the first cell type. In step 1510, the at least one computer processor outputs a representation of spatial variability of cells of the first cell type in the at least one image.
[0109] FIG. 16 is a flow diagram of an embodiment of a method relating to identifying spatial variation of cells of a first cell type in at least one image (i.e., examining variation in relative spatial distribution of a particular cell type among a plurality of samples, e.g., healthy versus diseased tissue). The process flow 1600 may be implemented by at least one computer processor. In some embodiments, there may be at least one non-transitory computer-readable storage medium encoded with a plurality of instructions that, when executed by at least one computer processor, implements the process flow 1600. The process flow 1600 includes steps 1602, 1604, 1606, and 1608. In step 1602, the at least one computer processor determines a first percentage of cells associated with the first cell type for cells of the plurality of cells that are less than a threshold distance to the first protein. In step 1604, the at least one computer processor determines a second percentage of cells associated with the first cell type for cells of the plurality of cells in the at least one image. In step 1606, the at least one computer processor compares the first percentage to the second percentage to obtain a comparison result. In step 1608, the at least one computer processor uses the comparison result to identify spatial variation of cells of the first cell type in the at least one image.
[0110] FIG. 17 is a flow diagram of an embodiment of a method related to capturing at least one image using a camera. The process flow 1700 may be implemented by at least one computer processor. In some embodiments, there may be at least one non-transitory computer readable storage medium encoded with a plurality of instructions that, when executed by the at least one computer processor, implements the process flow 1700. The process flow 1700 includes steps 1702 and 1704. In step 1702, the at least one computer processor captures a first image of the plurality of images, the first image being used to determine the spatial location of the plurality of cells. In step 1704, the at least one computer processor captures a second image of the plurality of images, the second image being used to determine the spatial location of the plurality of proteins.
[0111] FIG. 18 is a flow diagram of an embodiment of spatially aligning spatial locations from a plurality of images. The process flow 1800 can be implemented by at least one computer processor. In some embodiments, there can be at least one non-transitory computer readable storage medium encoded with a plurality of instructions that implement the process flow 1800 when executed by at least one computer processor. The process flow 1800 includes steps 1802 and 1804. In step 1802, the at least one computer processor spatially aligns the spatial locations of the plurality of cells from the first image with the spatial locations of the plurality of proteins from the second image. In step 1804, the at least one computer processor determines the number of cells of the first cell type that are less than a threshold distance to the first protein based on the spatially aligned spatial locations.
[0112] 19 is a flow diagram of an embodiment of a method for determining the number of cells of a first cell type that are less than a threshold distance to a first protein. The process flow 1900 can be performed by at least one computer processor. In some embodiments, there can be at least one non-transitory computer readable storage medium encoded with a plurality of instructions that perform the process flow 1900 when executed by at least one computer processor. The process flow 1900 includes steps 1902 and 1904. In step 1902, the at least one computer processor determines the shortest distance from the cells of the first cell type to the nearest protein of the plurality of proteins. In step 1904, the at least one computer processor compares the shortest distance to a threshold distance.
[0113] 20 is a flow diagram of one embodiment of a method for associating a cell with a cell type. Process flow 2000 may be performed by at least one computer processor. In some embodiments, there may be at least one non-transitory computer readable storage medium encoded with a plurality of instructions that, when executed by at least one computer processor, perform process flow 2000. Process flow 2000 includes steps 2002 and 2004. In step 2002, at least one computer processor receives genetic information of the cell (e.g., information regarding gene and / or protein expression). In step 2004, at least one computer processor associates the cell with a cell type from a plurality of cell types based on the genetic information of the cell, where the plurality of cell types includes a first cell type.
[0114] An exemplary implementation of a computer system 2100 that may be used in connection with any of the aspects of the disclosure provided herein is shown in Figure 21. The computer system 2100 may include one or more computer processors 2110 and one or more articles of manufacture that include a non-transitory computer-readable storage medium, such as a memory 2120 and one or more non-volatile storage media 2130. The processor 2110 may control the writing of data to and reading of data from the memory 2120 and the non-volatile storage 2130 in any suitable manner. To perform any of the functions or methods described herein, such as methods associated with process flows 1500, 1600, 1700, 1800, 1900, 2000, or, for example, to identify spatial variation of a cell type in at least one image (i.e., examining variation in relative spatial distribution of a particular cell type among multiple samples, e.g., comparing healthy tissue to diseased tissue), identify spatial variation of cells of a first cell type in at least one image, capture at least one image using a camera, spatially align spatial locations from the multiple images, determine a number of cells of a first cell type that are less than a threshold distance to a first protein, or associate cells with a cell type, the processor 2110 can execute one or more computer processor executable instructions stored in one or more non-transitory computer readable storage media, such as, for example, memory 2120, which can function as a non-transitory computer readable storage medium that stores processor executable instructions for execution by the processor 2110. System for mapping gene and protein expression in cells
[0115] The present disclosure also provides a system for mapping gene and protein expression in a cell. In some embodiments, the system comprises: a) a cell; and b) one or more pairs of oligonucleotide probes comprising a first oligonucleotide probe and a second oligonucleotide probe, wherein: i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, wherein the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe.
[0116] In some embodiments, the system further comprises a microscope (e.g., a confocal microscope). In some embodiments, the system further comprises a computer. In some embodiments, the system further comprises software for performing microscopy and / or image analysis (e.g., using any of the image analysis methods described herein). In some embodiments, the system further comprises a ligase. In some embodiments, the system further comprises a polymerase. In some embodiments, the system further comprises amine-modified nucleotides. In some embodiments, the system further comprises reagents for making a polymer matrix (e.g., a polyacrylamide matrix). The cells in the systems of the present disclosure can be any of the cell types disclosed herein. In some embodiments, the system comprises a plurality of cells. In some embodiments, the cells are of different cell types. In some embodiments, the cells are present in a tissue. In some embodiments, the tissue is a tissue sample provided by or from the subject. In some embodiments, the subject is a human.
[0117] example Example 1: Development of the STARmap Pro method Amyloid-β plaques and neurofibrillary tau tangles are neuropathological hallmarks of Alzheimer's disease (AD), but the molecular events and cellular mechanisms underlying AD pathophysiology remain poorly understood in spatial and temporal dimensions. The STARmap Pro method was developed and applied to simultaneously detect the transcriptional status and protein disease markers (amyloid-β aggregates and tau pathology) of single cells in brain tissues of AD mouse models. Through pseudo-time trajectory construction and joint analysis of differential gene expression at subcellular resolution (200 nm), high-resolution spatial maps of cell types and states in AD pathology were constructed. Disease-associated microglial (DAM) cells formed an inner shell in direct contact with amyloid-β plaques from the early stages of disease progression, while disease-associated astrocytes (DAA) and oligodendrocyte precursor cells (OPCs) were found to be enriched in the outer shell surrounding amyloid-β plaques in the later stages of the disease. Hyperphosphorylated tau appeared mainly in excitatory neurons and axonal processes. Furthermore, disease-associated gene pathways were pinpointed and validated across different cell types, suggesting inflammatory and gliotic processes in glial cells, and a reduction in neurogenesis in the adult hippocampus.
[0118] Previous STARmap methods were not compatible with histological staining (immunostaining or small molecule staining) and were limited to the detection of 1024 genes. To overcome such limitations with STARmap Pro, we first streamlined the experimental protocol to incorporate antibody (AT8 antibody, detecting phosphorylated tau) and dye staining (X-34, detecting Aβ plaques) into the library preparation and in situ sequencing steps (Figure 1B, Figures 8A and 8B). Intracellular mRNA in brain tissue was detected by a pair of DNA probes (primer and padlock, Figure 1C) and enzymatically amplified as cDNA amplicons. Proteins were labeled with primary antibodies, and then both the cDNA amplicons and the primary antibodies were embedded in a hydrogel matrix. Each cDNA amplicon contains a gene-specific identifier (barcode), which is read out by in situ SEDAL sequencing (sequencing with error reduction by dynamic annealing and ligation) ( Wang et al., 2018 ) and subsequent fluorescent protein staining (secondary antibody staining and small molecule dye X-34 staining) to visualize the protein signal.
[0119] The barcodes encoding genes in the DNA probes were then extended from 5 nucleotides (nt) to 10 nt (10^6 coding capacity), sufficient to code for over 20,000 genes. Furthermore, an additional 5 nt barcode was strategically designed near the ligation site to increase the specificity of gene detection by reducing non-specific amplification of mismatched primer-padlock pairs (Figure 8C). The barcode design was then verified by incorporating mismatches near the ligation site. The results showed undetectable signals, suggesting that the STARmap Pro method has high specificity (Figures 8D and 8E).
[0120] To investigate how AD-related pathology, including amyloid deposits and hyperphosphorylated Tau, affects transcriptional responses at the cellular level, we performed eight rounds of in situ sequencing to map 2766 genes and imaging after one round of sequencing (Figure 1D) to localize Aβ plaques and phosphorylated Tau (p-Tau) in thin coronal sections of brain from TauPS2APP and control mice. Transgenic TauPS2APP mice express mutant forms of human amyloid precursor protein (APPK670N / M671L) and presenilin 2 (N141I), which produce high levels of Abeta (Aβ) peptides and cause plaque formation, in addition to expressing a mutant form of human MAPT (P301L) that becomes hyperphosphorylated and aggregates. Sections were analyzed at 8 months (an age when tau and Aβ pathology are well established and expanding) and 13 months (a more advanced disease stage with severe pathology and elevated neuroinflammatory activity) (Lee et al., 2021). We primarily investigated the cortex and hippocampus, two of the most susceptible regions in AD. The spatial protein signal of TauPS2APP showed that Aβ plaque signals were distributed in both cortical and hippocampal regions, and p-Tau immunoreactivity was mainly distributed in the CA1 region of the hippocampus (Figure 1D), in agreement with previous reports (Grueninger et al. 2010, Lee et al. Neuron 2021), suggesting the fidelity of protein detection in STARmap Pro. 19,932 cells from two TauPS2APP mice and 17,135 cells from two control mice (non-transgenic littermates) were mapped at subcellular resolution (voxel size 95 nm × 95 nm × 346 nm, Figure 1A). The 3D RNA reads were projected onto a 2D plane for cell segmentation and quality control filtering of single-cell transcriptional profiles (area is 1000 pixels or 9.025 µm). 2 After 30 min of sequencing (>68 RNA reads per cell; see also STAR methodology), the remaining 33,106 cells pooled from all four samples were subjected to downstream analysis. Example 2: Hierarchical cell classification and spatial analysis
[0121] To identify cell types from STARmap Pro data, we employed a hierarchical clustering strategy; in this strategy, top-level clustering helps to classify cells into common cell types shared across all samples, while sub-level clustering helps to further identify disease-associated subtypes. During top-level clustering, the Leiden algorithm was applied along with Uniform Manifold Approximation and Projection (UMAP) to a low-dimensional representation of all transcriptome profiles (McInnes et al., 2018; Traag et al., 2019). Thirteen major cell types were identified with clear annotations according to previously reported genetic markers and tissue morphology (Figure 2A, Figure 9A). For example, excitatory neurons were annotated by high expression levels of genes related to ion channels and synaptic signaling, such as Vsnl1, Snap25, and Dnm1. Inhibitory neurons were separated by enrichment of the gamma-aminobutyric acid (GABA) transporter Slc6a1. Corresponding clusters were also annotated using other non-neuronal cell type specific markers, e.g. Aldoc for astrocytes, Bsg for endothelial cells, Ctss for microglia, and Plp1 for oligodendrocytes (Figure 9A). UMAP plots of TauPS2APP samples showed a distinct distribution of astrocytes, microglia, oligodendrocytes, and dentate gyrus (DG) cells compared to control samples, suggesting disease-associated cell subtypes (Figure 2A). Heterogeneity within each major cell type was therefore explored, and 24 sub-level clusters were identified based on their transcriptomic signatures (Figure 2B).
[0122] Benefiting from the spatial information preserved at subcellular resolution, spatial cell type atlases were generated together with histopathological features in the cortical and hippocampal regions of the four samples (Figure 2C and Figure 9D). All spatial cell atlases of the different samples showed similar anatomical structures in the cortical and hippocampal regions, confirming the robustness and reliability of the top clustering results. Furthermore, several major cell types showed various spatial distributions near Aβ plaques in TauPS2APP samples, where p-Tau signals were evident in the hippocampal region at 13 months (Figure 2C). To further quantify the cell type composition in spatial relationship with Aβ plaques, the distance from the centroid of each cell to the edge of its nearest plaque was measured. Cells were then counted (within five 10 μm concentric circles) as a function of distance from Aβ plaques and grouped based on the annotation of major cell types (Figure 2D, 2E, and 9F). Of the 13 major cell types, microglia, astrocytes, and oligodendrocyte progenitor cells (OPCs) were enriched around plaques in TauPS2APP samples compared to the overall cell type composition. Microglia were the most predominant cell type within the 10 μm ring. OPCs and astrocytes showed moderate enrichment at distances of 20–30 μm. The remaining cell types were relatively depleted within the 10–20 μm ring.
[0123] Top-level cell clustering and spatial analysis revealed that microglia, astrocytes, OPCs, oligodendrocytes and neurons exhibited changes in transcriptional profiles, spatial locations, or both features. These cell types were therefore selected for detailed sub-level clustering analysis to pinpoint disease-associated cell subtypes and gene pathways. AD is an inherently progressive disease, with ongoing molecular and cellular fluctuations. However, clustering analysis can only identify different cell types and cannot describe the continuous transition of cell states. To capture the gradient of cell states with disease progression and determine the relationship between different subtypes, in the following sections, we used Monocle pseudotime analysis (Cao et al., 2019), a computational tool widely used to reconstruct cell differentiation trajectories, as a complement to the sub-cluster analysis. Example 3: Disease-associated microglia directly contact Aβ plaques from the early stage of disease onset
[0124] After confirming that microglia were enriched in the immediate vicinity of plaques (Figure 2E), heterogeneity within the microglial population was investigated by subclustering analysis of the transcriptome profile. Three subpopulations were identified, named Micro1 (n=779), Micro2 (n=415), and Micro3 (n=529), respectively (Figure 3A). While Micro1 and Micro2 subtypes were present in all four samples, the Micro3 population was largely expanded from 8 to 13 months in TauPS2APP samples and was nearly absent in control samples (Figure 3A). Furthermore, the Micro3 subcluster expressed high levels of C1qa, Trem2, Cst7, Ctsb, and Apoe, which are known markers of disease-associated microglia (DAMs) and associated with neurodegeneration (Friedman et al., 2018; Keren-Shouetal. et al., 2017) (Figure 3B). Given its strong association with disease models and concordance with known DAM genetic markers, Micro3 was annotated as a DAM.
[0125] The continuous slope of cell state transitions was then further reconstructed in microglia by pseudotime trajectory analysis. Microglia populations exhibited linear pseudotime trajectories that matched very well with the actual disease progression timeline. Microglia in control samples were enriched at the beginning of the trajectory, whereas microglia in TauPS2APP samples continued to vary along a continuous path from 8 to 13 months (Figure 3C-3E). The Micro3 subcluster was presented at later pseudotime points along the trajectory, whereas both Micro1 and Micro2 spread along the trajectory from early to intermediate pseudotime values (Figure 3D-3E). We then analyzed the spatial distribution of the three microglia subclusters. The following was observed: (i) all three subclusters were present in both the cortical and hippocampal regions of TauPS2APP samples (Figures 3F-3G); (ii) Micro3 was almost exclusively present in TauPS2APP mice but absent in control samples (Figures 3H, 10A); (iii) Micro2 showed higher density in the hippocampal region than in the cortical region in both TauPS2APP and control samples (Figure 3H). In cell type composition analysis around Aβ plaques, the following was observed: (i) at 8 months, 64.5% of all cells within the 10 μm ring around the plaque were microglia, and 73.3% of the microglia were Micro3 (Figure 10B); (ii) at 13 months, their numbers increased to 75.8% and 81.5%, respectively (Figure 3H, bottom). Furthermore, although Micro2 is not strictly disease specific, it was significantly enriched within the 10 μm ring around plaques compared to other (non-microglial) cell types (Figure 3H and Figure S10B), suggesting that Micro2 also responds to Aβ plaques.The spatial cell atlas of microglia was further plotted with pseudotime values, and the pseudotime distribution around plaques was calculated using a similar concentric ring quantification used for cell type analysis.The results showed the expected observation that microglial cells near the plaque (within 10 μm) had higher pseudotime values than cells farther away from the plaque, revealing the spatiotemporal trajectory of microglial activation as microglia migrate toward Aβ plaques (Figures 10C and 10D).
[0126] To obtain a more comprehensive understanding of the microglial response to AD pathology at the molecular level, we then compared microglial gene expression in TauPS2APP and control mice (Figure 10E). Most of the upregulated DEGs identified in AD samples were DAM gene markers involved in biological processes, such as cell activation (i.e., Cst7, Apoe, Cd9, and Clec7a), inflammatory response (i.e., Trem2, Ccl6, and Cd68), and control of antigen processing and presentation (i.e., Ctss and H2-k1). Meanwhile, downregulated genes in microglia in disease samples were related to regulation of morphogenesis of anatomical structures (i.e., Sparc, Numb, Cdc42), endocytosis (Calm1 and Bin1), and positive control of trafficking (i.e., P2ry12 and Glud1). We selected a subset of the most significant DEGs and cell type markers and validated their expression variations using an additional mouse set in a dataset of 64 genes (Figure 3I). Example 4: Disease-associated astrocytes appear at late stages near the plaque-DAM complex
[0127] Astrocytes were another non-neuronal cell population with significant differences between TauPS2APP and control samples. Sub-level clustering analysis of astrocytes identified three transcriptomically distinct subpopulations, Astro1 (n=1,068), Astro2 (n=1,271), and Astro3 (n545) (Figure 4A). All three subclusters were observed in all samples; however, Astro3 cells were almost absent in control mice, and the Astro3 population was greatly expanded from 8 to 13 months in TauPS2APP mice (Figure 4A). Previous studies have identified disease-associated astrocytes (DAA) with a characteristic upregulation of Gfap, and genetic markers for Astro3 were similar to those of DAA (Habibetal., 2020) (i.e., Gfap, Vim, and Apoe, Figure 4B). Therefore, the Astro3 cluster was annotated as DAA because its transcriptome profile is similar to DAA and is associated with AD models. Furthermore, Astro1 and Astro2 correspond to the previously reported low-Gfap and intermediate-Gfap cell populations (Habib et al., 2020).
[0128] Despite the apparent linear slope of many marker genes across Astro1-3 subtypes, branching paths were observed from pseudotime trajectory analysis of astrocyte populations (Figure 4C). By visualizing the annotation of subclusters along the pseudotime trajectory, the longest path (lower path) was identified as consistent with a slope from Astro1 (start point) and Astro2 (end point) with no obvious association with disease progression. In contrast, a shorter branch from the branch point (upper path) represented a transition from a non-disease state (mainly Astro2) to Astro3 (DAA) (Figure 4D).
[0129] Spatial cell maps of astrocytes further showed that Astro1 was located near cortical and hippocampal neurons, while Astro2 was enriched in corpus callosum and lacunos stratum corneum molecules (Figure 4F-4H and Figure Figure11A). 11A). Cell type analysis in relation to tissue pathology revealed that DAM interacts directly with plaques within 10 μm to form DAM-plaque complexes, while Astro3 (DAA) is enriched relatively far from plaques (20-40 μm) in both cortical and hippocampal regions. In the cortex, corpus callosum, and hippocampal regions, approximately 3.1%, 7.8%, or 10.1% of cells in the 20-40 μm range around plaques were DAA, respectively, in the 13-month TauPS2APP sample (Figure 4H). Furthermore, Astro2 was also enriched near plaques at 8 months, contributing 3-16% of all cells in the 20-40 μm range; however, Astro2 presence declined to 2-7% at 13 months (Figures 4H and and11B). 11B). The observed shift from Astro2 to Astro3 populations near plaques from 8 to 13 months, combined with the disease-associated transcriptome pseudo-time trajectory of Astro2 to Astro3 (Figure 4D, top path), suggests that Astro2 to Astro3 (DAA) conversion may be present near plaques with disease progression.
[0130] Most of the DEGs in astrocytes of TauPS2APP and control samples were related to glial differentiation (i.e., Gfap, Vim, Clu, and Stat3) and extracellular matrix (i.e., Ctsb and Bcan) (Figure 11E). By showing the gene expression levels of the top DEGs identified on the pseudotime embedding, the expression profiles of Gfap and Vim showed the best correlation with DAAs and resembled the molecular slope along the disease-related branch (Figure 4I, upper path). Enriched GO terms from the DEGs included negative control of neuronal projection development and positive control of glial cell proliferation, as well as intermediate filament organization (Figure 4J), suggesting an overall gliotic process in the AD disease model. Again, a subset of the most significant DEGs and cell type markers were selected and their expression variations were also observed in the 64-gene validation dataset (Figure 4K). Example 5: Oligodendrocyte subtypes and progenitor cells are enriched near the middle of the plaque
[0131] Sub-level clustering analysis was then performed on oligodendrocyte lineage cells, identifying four clusters (Figures 5A and 5B): Oligo1 (n=4,295), Oligo2 (n=181), Oligo3 (n=490), and OPC (n=549). The oligodendrocyte gene marker Plp1 showed relatively uniform expression across all clusters, whereas Klk6 and Cldn11 marked the Oligo2 and Oligo3 populations, respectively (Figure 5B). TauPS2APP disease model and control samples showed similar distribution of all four cell populations on UMAP embeddings (Figure 5A).
[0132] Pseudotemporal trajectory analysis of a mixed population of OPC and oligodendrocyte cells recapitulated the known differentiation path from OPC to mature oligodendrocytes (Figure 5C, lower trajectory). Similar to astrocytes, disease-associated trajectories branched off alongside the main path from OPC to oligodendrocytes (Figure 5C, upper branch). Oligodendrocytes from TauPS2APP samples were highly enriched around the disease-associated trajectory branch compared to 8-month TauPS2APP and control samples (Figure 5C). Furthermore, disease-associated trajectories included all three Oligo1-3 subtypes (Figure 5D). Overall, the pseudotemporal distribution of cells in 13-month disease samples had a significantly higher mean compared to both 13-month control and 8-month disease samples (Figure 5E).
[0133] From the spatial maps and cell density calculations (Figures 5F-5H, top, and Figure 12A), we observed that the cell density of Oligo2 and Oligo3 populations in TauPS2APP was 100-200% higher in disease samples at 13 months compared to control samples. Oligo2 and Oligo3 were therefore annotated as disease-enriched oligodendrocytes. Cell type distribution analysis near Aβ plaques revealed that OPCs were enriched within a 20-40 μm ring around the plaques in TauPS2APP mice at both 8 and 13 months (Figures 5H and 14B). Furthermore, Oligo2 and Oligo3 were also enriched at a distance of 20-40 μm from the plaques in TauPS2APP mice, especially in the subcortical regions.
[0134] Through DEG analysis in oligodendrocytes from the comparison of TauPS2APP and control mice, a group of genes (i.e., Cldn11, Klk6, Serpina3n, and C4b) that were strongly upregulated in 13-month-old TauPS2APP mice were identified and validated. In contrast, the fold changes of DEGs between 8-month-old TauPS2APP and control mice were not significant, and some DEGs showed inconsistent changes in later experimental validation (Figures 5K and 14F). GO term analysis indicated key biological processes of oligodendrocytes in cytokine production, regulation of neuronal death, and regulation of synaptic plasticity and myelination during disease progression. Further DEG analysis along the disease-associated pseudotime trajectory further confirmed the disease association of Klk6, Cldn11, and C4b genes in oligodendrocytes (Figure 5J). Example 6: Neuronal response to Aβ plaques and tau tangles
[0135] In addition to cellular changes in non-neuronal cells, transcriptomic responses in neurons are important for understanding the mechanisms of neurodegeneration. Sub-level clustering analysis of neuronal populations was performed, and eight excitatory neuron subclusters and four inhibitory neuron subclusters were identified. As visualized in the neuronal spatial cell maps (Figures 6A, 6B, 13A, and 13B), the four subtypes of cortical excitatory neurons correspond to different cortical layers (CTX-Ex1–4). The types of excitatory neurons in the hippocampal regions were the same as the major types (DG, CA1, CA2, and CA3), and no additional subclusters were identified within each major type. Of the four subtypes of inhibitory neurons, Pvalb and Sst neurons were enriched in the cortex, whereas Cnr1 and Lamp5 neurons were overrepresented in the hippocampus.
[0136] Next, we investigated the neuronal type composition and their transcriptomic profiles in relation to Aβ plaques. In general, due to the migration and expansion of non-neuronal cells (microglia, astrocytes, OPCs, oligodendrocytes) near plaques, the percentage of all neuronal cells was low near plaques and positively correlated with their distance to the plaques. However, in the hippocampal region, the neuronal population around plaques was mostly composed of DG cells, which is consistent with the observation that the majority of plaques appeared near the DG. A recent report showed that adult hippocampal neurogenesis (AHN) activity in the DG was rapidly reduced in human AD patients. Therefore, we tested whether Aβ plaques in the TauPS2APP mouse model also affected AHN in the DG by pseudotemporal trajectory analysis. At 8 months, plaques were almost absent near the DG, so the pseudotemporal distribution of DG cells in TauPSAPP and control mice was indistinguishable. However, at 13 months, with an increase in the number of plaques near the DG, the pseudotime trajectory bifurcated into two branches, corresponding to the DG populations of TauPS2APP and control samples, respectively. The molecular responses of neurons in the DG region were further studied as they are related to neurogenesis. According to the pseudotime analysis, the samples at 13 months followed two distinct paths. In the diseased samples at 13 months, genes such as Dapk1 were upregulated, which were also involved in regulating neuronal cell death (Figures 6I and 6J).
[0137] To investigate the changes induced by tau tangles, quantification of tau protein intensity in each cell was performed by calculating the ratio of the number of tau-positive pixels to the total pixel area of each cell. Because tau tangles in axons may be erroneously attributed to cells, a threshold was used to select tau-positive neurons in each sample. In this case, at 8 months, the majority of tau-positive excitatory neurons in the sample were in the CTX-Ex2 population, whereas at 13 months, most of them were found in the CA1 region. For inhibitory neurons, most tau-positive cells were from the Pvalb population, whereas at 13 months, most of them were from the Sst population (Figure 6E). Tau signals around plaques were also quantified, revealing that in cortical regions, tau signals were enriched near plaques, whereas no pattern was observed in subcortical regions (Figure 6F).
[0138] Finally, the combined effect of Aβ plaques and tau tangles on neurons was investigated by pooling all subtypes into two large categories, excitatory and inhibitory neurons, and analyzing the DEGs in neurons between TauPS2APP and control samples. Overall, the DEGs identified from the excitatory and inhibitory neuronal populations were highly consistent. GO term analysis showed that the identified DEGs were enriched in the biological processes of cell cycle regulation (Ccnb2), neuronal differentiation (Gpm6a), and regulation of neuronal death (Ddit3) (Figures 6G and 6H). Example 7: Integrated analysis of disease-related cells in AD pathology
[0139] The above mentioned analyses have focused on analyzing disease-related subtypes and DEGs within individual major cell types. To synthesize a comprehensive picture of AD gene pathways from multiple cell types, gene set enrichment analysis (GSEA) was performed using DEGs from four major cell types (microglia, astrocytes, oligodendrocytes, and neurons). As shown in the GO term enrichment heatmap (Figures 7A and 14A), at 13 months, most of the significantly enriched terms of non-neuronal cells in biological processes were glial differentiation and glial development, which recapitulate the emergence of disease-related populations, while those of neurons are involved in regulating synaptic signaling, cognition, and synaptic plasticity. Therefore, we collected all annotations enriched in the target cell types and grouped them based on their similarity. A GO enrichment map was generated, showing both cell type-specific annotations and shared annotations of major AD participants (Figure 14B). Cell type-specific terms associated with immune responses, inflammatory responses, and lysosomes were enriched in microglia as the populations directly interacted with sites of pathology, whereas terms associated with cell motility, morphogenesis, and differentiation were shared across multiple cell types.
[0140] In addition to the DEGs identified in the comparison between AD and control samples, spatial DEGs were also calculated using cells close to the plaque (within 25 μm) compared to cells far from the plaque (more than 25 μm away). Genes specifically upregulated in the area near the plaque were considered plaque-induced genes (PIGs). Sixteen PIGs were identified in the 8-month and 29 PIGs in the AD samples at 13 months (Figures 7B and 14G). All 16 PIGs in the 8-month AD samples were included in the PIGs in the 13-month AD samples (Figure 7C). PIGs enriched within 10 μm of the plaque were found to be mainly DAM marker genes such as Trem2, Cst7, Ctsb, Apoe, and Cd9 (Figures 7B and 14G). Vim, a DAA marker, was upregulated in the area 20–30 μm from the plaque in the 13-month AD samples (Figure 7B). This is consistent with previous findings that DAAs were enriched in the region 20–30 μm from plaques in 13-month AD samples (Figure 4H). The discovered PIGs were also compared with previously reported PIGs in 18-month AppNL-GF mice. The results showed that 58.9% (17 of 29) of PIGs in 13-month AD samples and 87.5% (14 of 16) of PIGs in 8-month AD samples overlapped with the reported PIGs (Figure 7C). The high overlap rate implied the accuracy of STARmap Pro in detecting spatially conserved RNA signals.
[0141] To gain a more comprehensive and quantitative understanding of the spatial relationship between each cell type and Aβ plaques, we calculated the average Euclidean distance from plaques, DAMs, DAAs, and neurons to their nearest neighbors (Figure 14C and 14E). Compared to the shuffled references (Figure 14D and 14F), we found that DAMs were the closest neighbors to amyloid plaques. Instead of plaques, DAAs had shorter distances to DAMs, suggesting that DAAs are in contact with the DAM-plaque complex. To further validate this finding, cells were next quantified at different distance intervals from each plaque and grouped based on their cell type identity (Figure 7D). In summary, DAMs were closest to plaques and were enriched in the region between 0 and 20 μm. DAAs and OPCs were relatively similar in terms of distance to plaques and distributed in the region between 20 and 40 μm. Oligodendrocytes and neurons had higher density in the outer region (Figure 7D). Based on this spatial information, we generated diagrams of the cell distribution patterns around plaques in the AD 8-month and AD 13-month samples (Figure 7E). In both stages, DAMs are in direct contact with the plaques to form an inner core. With disease progression, more DAAs appeared and were enriched in the second-most proximal plaque regions. The enrichment of OPCs in the regions near plaques was also shown in the diagram of the TauPS2APP 13-month sample. Marker genes or DEGs of these target cell types were also labeled in the diagram. Consideration
[0142] STARmap Pro was developed for in situ detection of RNA and protein signals in the same tissue section at single-cell resolution. Based on STARmap, its detection capability and specificity have been further improved, providing a new capability to simultaneously profile RNA and protein at single-cell resolution while preserving spatial information. STARmap Pro offers the opportunity to study biological systems in a more comprehensive manner and enables multimodal spatial gene expression analysis. Because proteins are the most common disease hallmarks, the method described here may be widely used in pathology studies. Using STARmap Pro to investigate gene expression changes near disease hallmarks will advance our understanding of disease pathogenesis. The original STARmap only detects RNA signals and cannot accurately represent protein abundance or detect protein modifications. In contrast, the immunostaining strategy of STARmap Pro can distinguish not only various protein types but also protein modifications, making it a useful tool for cancer research; because protein modifications such as phosphorylation play important roles in cancer development and progression.
[0143] STARmap Pro was applied to detect RNA and AD pathology protein signals in AD mouse models. Spatially secured single-cell RNA signals were used to generate spatial maps of cell types and reveal the cellular distribution patterns associated with Aβ plaques; the results showed that microglia were significantly enriched in the areas near plaques. Then, subclustering analysis of five target cell types (microglia, astrocytes, oligodendrocytes, excitatory neurons, and inhibitory neurons) was performed. The results provide a comprehensive spatial map that identifies the appearance of DAMs and DAAs and their enrichment behavior near Aβ plaque areas. This is consistent with previous studies, further confirming the detection accuracy of STARmap Pro. When analyzing the cellular distribution around plaques, it was found that DAMs were distributed in the closest area around Aβ plaques, and then DAAs were distributed in the outer area near DAMs. This implies that the distribution of DAMs may be directly affected by Aβ plaques, and then DAAs may be affected by DAMs. Indeed, a previous study revealed that reactive astrocytes, which highly express the DAA marker genes Gfap and Vim, are induced by activated microglia. Regarding the distribution pattern of cell types, we also found that OPCs are enriched in the area near the plaques. These OPCs near the plaques may differentiate into oligodendrocytes, which are involved in the pathology of AD.
[0144] Transcriptional signatures of pathology responsiveness in five major cell types were also provided. While most gene markers were cell type specific and may be involved in specific perturbations, some DEGs were found to be shared across cell types. For example, Gfap was upregulated in all major cell types. This was consistent with a previous mononuclear study that also showed high Gfap expression in Alzheimer's disease-specific subclusters of various cell types. Microglia and astrocytes also shared DEGs, such as CTSB. Pathway analysis also indicated that DEGs from different cell types were involved in similar biological processes. method mouse
[0145] All animal procedures followed animal care guidelines approved by the Genentech Institutional Animal Care and Use Committee (IACUC), and animal experiments were performed in accordance with IACUC policies and NIH guidelines. Mice used in STARmap Pro contain the human tau P301L mutant and PS2 N141I and APP swe (PS2APP homo ;P301L hemi ), and a non-transgenic control. Tissue collection and sample preparation for STARmap Pro
[0146] Animals were anesthetized with isoflurane and quickly decapitated. Brain tissue was removed, placed in OCT, frozen in liquid nitrogen, and stored at -80°C. For mouse brain tissue sections, brains were transferred to a cryostat (Leica CM1950) and cut as 20 μm slices in coronal sections. Brain slices were fixed with 4% PFA in 1× PBS buffer for 15 min at room temperature, permeabilized with cold methanol, and left at -80°C for 1 h. STARmap Pro detects spatial RNA and protein signals
[0147] Samples were removed from -80°C and washed at room temperature for 5 min and then with PBSTR buffer (0.1% Tween-20, 0.1 U / μL SUPERase·InRNase inhibitor in PBS). After washing, samples were incubated with 300 μl of 1× hybridization buffer (2× SSC, 10% formamide, 1% Tween-20, 0.1 mg / ml yeast tRNA, 20 mM RVC, 0.1 U / μL SUPERase·InRNase inhibitor, and pooled SNAIL probe at 1 nM per oligo) for 36 h in a humidified oven at 40°C with shaking and parafilm wrapping. Samples were washed twice with PBSTR and once with high salt wash buffer (4× SSC in PBSTR) at 37°C. Finally, samples were rinsed once with PBSTR at room temperature. The samples were then incubated with ligation mixture (1:10 dilution of T4 DNA ligase in 1× T4 DNA ligase buffer supplemented with 0.5 mg / ml BSA and 0.2 U / μL SUPERase·InRNase inhibitor) for 2 h at room temperature with gentle shaking. After ligation, the samples were washed twice with PBASR buffer and incubated with rolling circle amplification (RCA) mixture (1:10 dilution of Phi29 DNA polymerase in 1× Phi29 buffer supplemented with 250 μM dNTPs, 20 μM 5-(3-aminoallyl)-dUTP, 0.5 mg / ml BSA and 0.2 U / μL SUPERase·InRNase inhibitor) for 2 h at 30°C with gentle shaking. Samples were washed twice with PBST (0.1% Tween-20 in PBS) and blocked with blocking solution (5 mg / ml BSA in PBST) for 30 min at room temperature. Samples were incubated with p-Tau primary antibody (1:100 dilution in blocking solution) for 2 h at room temperature. Samples were washed three times for 5 min each with PBST. Samples were then treated with 20 mM acrylic acid NHS ester in PBST for 1 h and then rinsed once with PBST. Samples were incubated in monomer buffer (4% acrylamide, 0.2% bisacrylamide in 2×SSC) for 15 min at room temperature.The buffer was then aspirated and 35 polymerization mix (0.2% ammonium persulfate, 0.2% tetramethylethylenediamine in monomer buffer) was added to the center of the sample and immediately covered with a Gel Slick-coated coverslip. The polymerization reaction was carried out at room temperature for 1 h, followed by two washes with PBST for 5 min each. The sample was treated with dephosphorylation mix (1:100 dilution of shrimp alkaline phosphatase in 1× CutSmart buffer supplemented with 0.5 mg / ml BSA) for 1 h at 37°C, followed by three washes with PBST for 5 min each.
[0148] For in situ RNA sequencing, each cycle started with treating the samples with stripping buffer (60% formamide and 0.1% Triton-X-100 in water) twice for 10 min at room temperature, followed by washing with PBST three times for 5 min each. Samples were incubated with sequencing mixture (1:25 dilution of T4 DNA ligase in 1× T4 DNA ligase buffer supplemented with 0.5 mg / ml BSA, 10 μM reading probe and 5 μM fluorescent oligo) for at least 3 h at room temperature. Samples were washed three times for 10 min each with washing and imaging buffer (10% formamide in 2× SSC) and then immersed in washing and imaging buffer for imaging. Images were acquired using a Leica TCS SP8 confocal microscope. Eight cycles of imaging were performed to detect 2766 genes.
[0149] After in situ sequencing of RNA signals, samples were incubated in X-34 solution (10 μM X-34, 40% ethanol and 0.02 M NaOH in 1× PBS) for 10 min at room temperature. Samples were then washed 3 times with 1× PBS, incubated for 1 min in 80% EtOH, and then washed 3 times with PBS for 1 min each. Samples were then incubated with secondary antibody (1:80 dilution in blocking solution) for 12 h at room temperature. Samples were then washed 3 times for 5 min each with PBST. For cell segmentation purposes, propidium iodide (PI) staining was performed according to the manufacturer's instructions. Another round of imaging was performed to detect spatial protein signals. Thin-section STARmap Pro data processing
[0150] All image processing steps were implemented using MATLAB R2019b and related open source packages in Python 3.6 and applied according to Wang et al., 2018.
[0151] Image preprocessing: Multidimensional histogram matching was performed for each tile using the MATLAB function "imhistmatchn". The image of the first color channel from the first sequencing round was used as a reference to ensure uniform illumination and contrast levels.
[0152] Image registration: Image registration was applied according to Wang et al., 2018. Global image registration was performed using a 3D fast Fourier transform (FFT) to calculate the cross-correlation between two image volumes at all translational offsets. The position of the maximum correlation coefficient was identified and used to transform the image volumes to compensate for the offsets.
[0153] Spot calling: After registration, individual dots were identified separately in each color channel in the first round of sequencing. Dots approximately 6 pixels in diameter were identified by finding local maxima in 3D. After identification of each dot, the dominant color of that dot across all four channels was determined in each round within a 5x5x3 voxel volume surrounding the dot's location.
[0154] Barcode filtering: Dots were first filtered based on their quality score. The quality score quantified the extent to which each dot in each sequencing round arose from one color rather than a mixture of colors. The barcode codebook was converted to color space based on the expected color sequence after two-base coding of the barcode DNA sequence. Dot color sequences that passed the quality threshold and matching sequences in the codebook were retained and identified as the specific gene that the barcode represented; all other dots were rejected. High-quality dots in the codebook and the identity of the associated gene were kept for downstream analysis.
[0155] 2D cell segmentation: Nuclei were automatically identified by the StarDist 2D machine learning model (Schmidt et al., 2018) from the maximum intensity projection of the stitched DAPI channel after the final round of sequencing. Cell positions were then extracted from the segmented DAPI images. Cell bodies were represented by an overlay of the stitched Nissl stain and merged amplicon images. Finally, a marker-based watershed transformation was applied to segment thresholded cell bodies based on the combined thresholded cell body map and the positions of identified nuclei. Points that overlap each segmented cell region in 2D were then assigned to that cell and a per-cell gene expression matrix was calculated.
[0156] Cell type classification: A two-level clustering strategy was applied to identify both major and sub-level cell types in the dataset. The processing steps in this section were implemented using Scanpy v1.4.6 (Wolf et al., 2018) and other customized scripts in Python 3.6 and were applied according to Wang et al., 2018. After filtering, normalization, and scaling, Principal Component Analysis (PCA) was applied to reduce the dimensionality of the cell expression matrix. Based on the explained variance ratio, the top PCs were used to calculate a neighborhood graph of the observations. Then, the Leiden algorithm was used to identify well-connected cells as clusters in a low-dimensional representation of the transcriptome profile. Cells are displayed using Uniform Manifold Approximation and Projection (UMAP) and color-coded according to their cell type. Cells in each top-level cluster were then sub-clustered using PCA decomposition, followed by Leiden clustering to determine sub-level cell types. spatial analysis
[0157] Plaque segmentation: The spatial analysis starts with plaque segmentation. Using the “bwlabel” function from the EBImage package, plaques can be segmented from the binary image of the plaque channel. Then, the size and center of each plaque are calculated using the “computeFeature.moment” and “computeFeature.shape” functions. Finally, plaques with an area of 400 pixels (approximately 36.7 μm) are segmented. 2 ) of plaque is removed by the filter.
[0158] Cell distribution around plaques: Since the cell positions PC are obtained from the data pre-processing step, considering a sample with $n$ cells and $m$ plaques, for each cell:
number
number
[0159] The number of cells is then counted for all cell types that fall into the different ranges. The ranges are set from 0-10 μm (ring 1) to 40-50 μm (ring 5). To remove differences in total cell numbers, the statistics are normalized by calculating the percentage of each cell type within the ring. A graphical illustration of this analysis is shown in Figure 2D. For global statistics, the percentage of each cell type in the entire sample was calculated.
[0160] Calculating inter-type distances: In a sample with n cells and m plaques, they are treated as objects with coordinates P and type labels (plaques, DAAs, DAMs, etc.). There are t type labels, and the set of objects of each type is
number
number
[0161] Shuffled control analysis is performed using the same algorithm, but labels are randomly assigned (the number of cells / plaques is the same). Differential expression and pathway analysis
[0162] Differential Expression Analysis: Before performing DE analysis, the dataset is normalized by: 1) dividing the number of genes in a sample by the median *total counts per cell* of that sample and multiplying by a scale factor (defined as the average of the median *total counts per cell* of all samples); 2) performing a log2 transformation by adding a pseudocount of 1:
number
[0163] DE genes are identified by performing a Wilcoxon rank sum test between two groups of cells using the "FindMarkers" function in Seurat. For comparisons such as "disease vs. control", the two groups of cells naturally extract a certain type of cell from the TauPS2APP and control samples. For the "cells close to plaque vs. cells far from plaque" comparison, "cells close to plaque" are cells with a nearest plaque distance of less than 25 μm, and all other cells are defined as "cells far from plaque". For the "CA1 Tau+ vs. Tau-" comparison, Tau+ CA1 cells are filtered according to the ratio of the Tau signal area to the area of the cell body. The threshold of that ratio is set to 0.3.
[0164] To filter out some low-expressing genes, the minimum threshold for the percentage of cells in which gene expression was detected in either cell population was set to 0.1. The following thresholds were also applied to the generated gene list to filter out non-significant genes: absolute log fold change >0.1, p-value <0.05.
[0165] To visualize the DE results, we generated a volcano plot using the "EnhancedVolcanoplot" package. DE genes with logFC>0 are displayed in red, other genes in blue. Significant genes that failed to pass the LogFC threshold (p-value < 0.05) are displayed in green. All other non-significant genes are displayed in grey. Note that some genes with extremely high -log(P-value) or logFC are capped. Gene Ontology (GO) enrichment analysis:
[0166] The website g:Profiler was used to perform GO enrichment analysis of DE genes for each comparison: the list of DE genes between cells from disease and control samples, or tau positive and negative, or cells near and far from plaques (25 μm was used as threshold) are the input for GO analysis. The scope of the statistical domain is restricted to annotated genes. The significance threshold was determined using g:SCS, with the user threshold set to 0.05. To limit the size of the functional categories subjected to enrichment analysis, GO terms with fewer than 20 genes or more than 1000 genes were excluded. The results were downloaded in generic enrichment map (GEM) format to be used as input for further functional enrichment analysis.
[0167] Cytoscape (v3.8.2) with EnrichmentMap (v3.3.1) and AutoAnnotate (v1.3.3) apps were used to integrate and visualize GO enrichment results from DEGs (Ex / In / Astro / Micro / Oligo) of five major cell types. The lists of statistically significant GO and KEGG terms obtained in g:Profiler were imported into Cytoscape with the following parameters: node (GO / KEGG term) cutoffs were set at adjusted p-value < 0.05 and FDRq-value < 0.1; for edges (representing similarity between gene lists in each node), a similarity threshold of 0.375 was used. Each node is colored by cell type to illustrate the common and distinct contributions of DEGs from each cell cluster. The enrichment map was automatically annotated using AutoAnnotate and a three-word label for each cluster was created using the WordCould app.
[0168] The SynGO enrichment tool was used to further characterize the synaptic functions enriched in DEGs from excitatory neurons, inhibitory neurons, and CA1 cells with tau pathology. Brain-expressed genes are used as background gene lists. References
[0169] 1. Braak, H., and Braak, E. (1991). Neuropathological stageing of Alzheimer-related changes. Acta Neuropathologica. 82, 239-259.
[0170] 2. Busche, M.A., and Hyman, B.T. (2020). Synergy between amyloid-β and tau in Alzheimer’s disease. Nature Neuroscience. 23, 1183-1193.
[0171] 3. Cao, J., Spielmann, M., Qiu, X., Huang, X., Ibrahim, D.M., Hill, A.J., Zhang, F., Mundlos, S., Christiansen, L., Steemers, F.J., et al. (2019). The single-cell transcriptional landscape of mammalian organogenesis. Nature. 566, 496-502.
[0172] 4. Chen, W.-T., Lu, A., Craessaerts, K., Pavie, B., Sala Frigerio, C., Corthout, N., Qian, X., Lalakova, J., Kuhnemund, M., Voytyuk, I., et al. (2020). Spatial Transcriptomics and In Situ Sequencing to Study Alzheimer’s Disease. Cell. 182, 976-991.e19.
[0173] 5. Grubman, A., Chew, G., Ouyang, J.F., Sun, G., Choo, X.Y., McLean, C., Simmons, R.K., Buckberry, S., Vargas-Landin, D.B., Poppe, D., et al. (2019). A single-cell atlas of entorhinal cortex from individuals with Alzheimer’s disease reveals cell-type-specific gene expression regulation. Nat. Neurosci. 22, 2087-2097.
[0174] 6. Grueninger, F., Bohrmann, B., Czech, C., Ballard, T.M., Frey, J.R., Weidensteiner, C., von Kienlin, M., and Ozmen, L. (2010). Phosphorylation of Tau at S422 is enhanced by Aβ in TauPS2APP triple transgenic mice. Neurobiol. Dis. 37, 294-306.
[0175] 7. Habib, N., McCabe, C., Medina, S., Varshavsky, M., Kitsberg, D., Dvir-Szternfeld, R., Green, G., Dionne, D., Nguyen, L., Marshall, J.L., et al. (2020). Disease-associated astrocytes in Alzheimer’s disease and aging. Nat. Neurosci. 23, 701-706.
[0176] 8. Hardy, J., and Selkoe, D.J. (2002). The amyloid hypothesis of Alzheimer’s disease: progress and problems on the road to therapeutics. Science. 297, 353-356.
[0177] 9. Keren-Shaul, H., Spinrad, A., Weiner, A., Matcovitch-Natan, O., Dvir-Szternfeld, R., Ulland, T.K., David, E., Baruch, K., Lara-Astaiso, D., Toth, B., et al. (2017). A Unique Microglia Type Associated with Restricting Development of Alzheimer’s Disease. Cell. 169, 1276-1290.e17.
[0178] 10. Lau, S.-F., Cao, H., Fu, A.K.Y., and Ip, N.Y. (2020). Single-nucleus transcriptome analysis reveals dysregulation of angiogenic endothelial cells and neuroprotective glia in Alzheimer’s disease. Proceedings of the National Academy of Sciences. 117, 25800-25809.
[0179] 11. Lee, S.-H., Meilandt, W.J., Xie, L., Gandham, V.D., Ngu, H., Barck, K.H., Rezzonico, M.G., Imperio, J., Lalehzadeh, G., Huntley, M.A., et al. (2021). Trem2 restrains the enhancement of tau accumulation and neurodegeneration by β-amyloid pathology. Neuron. 109, 1283-1301.e6.
[0180] 12. Masters, C.L., Bateman, R., Blennow, K., Rowe, C.C., Sperling, R.A., and Cummings, J.L. (2015). Alzheimer’s disease. Nat Rev Dis Primers. 1, 15056.
[0181] 13. Mathys, H., Davila-Velderrain, J., Peng, Z., Gao, F., Mohammadi, S., Young, J.Z., Menon, M., He, L., Abdurrob, F., Jiang, X., et al. (2019). Single-cell transcriptomic analysis of Alzheimer’s disease. Nature. 570, 332-337.
[0182] 14. McInnes, L., Healy, J., Saul, N., and Grosberger, L. (2018). UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software. 3, 861.
[0183] 15. Stahl , PL , Salmen , F , Vickovic , S , Lundmark , A , Navarro , JF , Magnusson , J , Giacomello , S , Asp , M , Westholm , JO , Huss , M , et al. (2016). Visualization and analysis of gene expression in tissue sections by spatial transcriptomics. Science. 353, 78-8
[0184] 16. Stuart, T.; Integrative single-cell analysis. Nat. Rev. Fr. Genet. 20, 257-2
[0185] 17. Traag, VA, Waltman, L., and van Eck, NJ. From Leuven to Leiden: guaranteeing well-connected communities. Sci. Rep. 9 , 5233 .
[0186] [ PubMed ] 18. Wang, X., Allen, WE, Wright, MA, Sylwestrak, EL, Samusik, N., Vesuna, S., Evans, K., Liu, C., Ramakrishnan, C., Liu, J., et al. (2018). Three-dimensional intact-tissue sequencing of single-cell transcriptional states. Science. 361.
[0187] 19. Zhou, Y., Song, WM, Andhey, PS, Swain, A., Levy, T., Miller, KR, Poliani, PL, Cominelli, M., Grover, S., Gilfillan, S., et al. (2020). Human and mouse single-nucleus transcriptomics reveal TREM2-dependent and TREM2-independent cellular responses in Alzheimer's disease. Nat. Med. 26, 131-142. Incorporation by Reference
[0188] This application refers to various issued patents, published patent applications, scientific journal articles, and other publications, all of which are incorporated herein by reference.Details of one or more aspects of the invention are described herein.Other features, objects, and advantages of the invention will become apparent from the detailed description, drawings, examples, and claims. Equivalents and Scope
[0189] Articles such as "a," "an," and "the" can mean one or more than one, unless indicated to the contrary or clear from context. An embodiment or description including "or" between one or more members of a group is deemed to be satisfied when one, more than one, or all of the group members are present in, used in, or otherwise relevant to a given product or process, unless indicated to the contrary or clear from context. The invention includes embodiments in which exactly one member of a group is present in, used in, or otherwise relevant to a given product or process. The invention includes embodiments in which more than one, or all of the group members are present in, used in, or otherwise relevant to a given product or process.
[0190] Furthermore, the disclosure encompasses all variations, combinations, and permutations in which one or more limitations, elements, clauses, and descriptive terms from one or more of the enumerated claims are introduced into another claim. For example, any claim that is dependent on another claim can be amended to include one or more limitations found in other claims that are dependent on the same base claim. When elements are presented as lists, such as in Markus group format, each subgroup of elements is also disclosed, and any element(s) can be removed from the group. In general, when the invention or aspects of the invention are said to include certain elements and / or features, it is to be understood that certain aspects of the disclosure or aspects of the disclosure consist of or consist essentially of such elements and / or features. For simplicity, these aspects have not been specifically described in these exact terms herein. It should also be noted that the terms "comprise" and "contain" are intended to be open and allow for the inclusion of additional elements or steps. When ranges are specified, the endpoints are also included. Additionally, unless otherwise indicated or clear from the context and the understanding of one of ordinary skill in the art, values expressed as ranges can take any particular value or subrange within the ranges set forth in various aspects of the invention to one tenth of the unit of the lower limit of the range, unless the context clearly dictates otherwise.
[0191] This application refers to various issued patents, published patent applications, journal articles, and other publications, all of which are incorporated herein by reference. In the event of any inconsistency between any of the incorporated references and this specification, this specification shall control. Furthermore, certain aspects of the present invention that fall within the prior art may be expressly excluded from one or more aspects. Such aspects may be excluded even if the exclusion is not expressly set forth herein, since such aspects are deemed known to those of ordinary skill in the art. Any particular aspect of the present invention may be excluded from any aspect for any reason, whether or not related to the existence of prior art.
[0192] Those skilled in the art will recognize or be able to ascertain, using no more than routine experimentation, many equivalents to the specific embodiments described herein. The scope of the embodiments described herein is not intended to be limited to the above description, but rather as described in the accompanying embodiments. Those skilled in the art will appreciate that various changes and modifications to this description are possible without departing from the spirit or scope of the invention as defined in the following claims.
Claims
1. A method for mapping gene and protein expression in cells, comprising the following: a) contacting the cells with one or more pairs of oligonucleotide probes, wherein each pair of oligonucleotide probes comprises a first oligonucleotide probe and a second oligonucleotide probe, wherein i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, wherein the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe; b) ligating the 5' end and the 3' end of the first oligonucleotide probe together to generate a circular oligonucleotide; c) performing rolling circle amplification to amplify the circular oligonucleotide using the second oligonucleotide probe as a primer to generate one or more concatenated amplicons; d) contacting the cells with one or more detection agents, wherein each detection agent binds to the protein of interest; e) embedding the one or more concatenated amplicons and the one or more detection agents in a polymer matrix; f) contacting the one or more concatenated amplicons embedded in the polymer matrix with a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe; and g) imaging the one or more concatenated amplicons embedded in the polymer matrix and the one or more detection agents embedded in the polymer matrix to determine the positions of the nucleic acid of interest and the protein of interest within the cells The method as described above.
2. The method according to claim 1, wherein gene and protein expression are profiled in a plurality of cells.
3. The method according to claim 2, wherein the cells comprise a plurality of cell types.
4. The method according to claim 1, wherein the cells are present within intact tissue.
5. The method according to claim 1, wherein the nucleic acid of interest is RNA.
6. The method according to claim 1, wherein gene expression of more than 2000 nucleic acids of interest is mapped.
7. The method according to claim 1, wherein the barcode sequences on the first and second oligonucleotide probes are 5 to 15 nucleotides in length.
8. The first oligonucleotide probe has the structure: 5′-[portion complementary to the second probe]-[portion complementary to the nucleic acid of interest]-[first barcode sequence]-[second barcode sequence]-3′, The method according to claim 1, comprising:
9. The second oligonucleotide probe has the structure: 5′-[portion complementary to the nucleic acid of interest]-[portion complementary to the first probe]-[barcode sequence]-3′, The method according to claim 1, comprising:
10. The method according to claim 1, wherein the second barcode sequence of the first oligonucleotide probe enhances the specificity of detection of the nucleic acid of interest or reduces non-specific amplification.
11. The method according to claim 1, wherein the third oligonucleotide probe comprises a detectable label.
12. The method according to any one of claims 1 to 11, wherein one or more detection agents are antibodies each comprising a detectable label.
13. The method according to any one of claims 1 to 11, further comprising contacting one or more detection agents with a secondary detection agent.
14. The method according to any one of claims 1 to 11, wherein one or more detection agents are antibodies each binding to an oligonucleotide sequence and each binding to a protein of interest.
15. The method according to claim 14, further comprising contacting each of one or more antibodies that bind to the protein of interest with an oligonucleotide conjugated to a detectable label, wherein each oligonucleotide conjugated to a detectable label is complementary to the oligonucleotide sequence conjugated to one of the antibodies.
16. The method according to any one of claims 1 to 11, wherein the first barcode sequence of the first oligonucleotide probe is a gene-specific sequence used to identify the nucleic acid of interest.
17. The method according to any one of claims 1 to 11, wherein the method is performed with a resolution of 200 nm or less of a cell.
18. A method for mapping gene expression in a cell, comprising: a) contacting the cells with one or more pairs of oligonucleotide probes, wherein each pair of oligonucleotide probes comprises a first oligonucleotide probe and a second oligonucleotide probe, wherein i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, wherein the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe; b) ligating the 5' and 3' ends of the first oligonucleotide probe together to generate a circular oligonucleotide; c) performing rolling circle amplification to amplify the circular oligonucleotide using the second oligonucleotide probe as a primer to generate one or more concatenated amplicons; d) embedding the one or more concatenated amplicons in a polymer matrix; e) contacting the one or more concatenated amplicons embedded in the polymer matrix with a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe; and f) imaging the one or more concatenated amplicons embedded in the polymer matrix to determine the location of the nucleic acid of interest within the cell comprising the method.
19. A plurality of oligonucleotide probes comprising a first oligonucleotide probe and a second oligonucleotide probe, wherein i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, wherein the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe, said plurality of oligonucleotide probes.
20. The plurality of oligonucleotide probes according to claim 19, further comprising a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe.
21. A method for identifying spatial variations in cell types within at least one image, comprising: receiving, for each of a plurality of cells within at least one image, the spatial position of the cells within the at least one image; receiving, for each of a plurality of proteins within at least one image, the spatial position of the proteins within the image; determining, for a first protein of the plurality of proteins, the number of cells of a first cell type whose distance to the first protein is less than a threshold distance, where the distance is determined based on the spatial positions of at least some of the plurality of cells and the spatial positions of at least some of the plurality of proteins; identifying spatial variations in cells of the first cell type within at least one image based on the number of cells of the first cell type; and outputting a display of the spatial variations in cells of the first cell type within at least one image The method as described above.
22. a) cells; and b) one or more pairs of oligonucleotide probes comprising a first oligonucleotide probe and a second oligonucleotide probe A system comprising, wherein: i) the first oligonucleotide probe comprises a portion complementary to the second oligonucleotide probe, a portion complementary to the nucleic acid of interest, a first barcode sequence, and a second barcode sequence; and ii) the second oligonucleotide probe comprises a portion complementary to the nucleic acid of interest, a portion complementary to the first oligonucleotide probe, and a barcode sequence, where the barcode sequence of the second oligonucleotide probe is complementary to the second barcode sequence of the first oligonucleotide probe. The system as described above.
23. The system according to claim 22, further comprising a third oligonucleotide probe comprising a sequence complementary to the first barcode sequence of the first oligonucleotide probe.