Imaging-based pooled genetic screens with highly multiplexed phenotypes

Image-based pooled genetic screens using MERFISH and RCA-MERFISH overcome limitations in detecting genetic function and cellular phenotypes by enabling high-resolution, multiplexed detection of nucleic acids and proteins, preserving spatial and morphological information.

WO2025122729A1PCT designated stage expired Publication Date: 2025-06-12PRESIDENT & FELLOWS OF HARVARD COLLEGE +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058643
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-26
Filing Date
2024-12-05
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current methods for detecting genetic function and cellular phenotypes are limited by low throughput, inability to multiplex, and loss of spatial and morphological information.

Method used

The development of image-based techniques for pooled genetic screens that allow for the simultaneous detection of nucleic acids and proteins at high resolution, using methods such as multiplexed error-robust FISH (MERFISH) and rolling circle amplification (RCA-MERFISH).

Benefits of technology

These methods enable the assessment of diverse cellular phenotypes, including transcriptional state, subcellular morphology, and tissue organization, allowing for the mapping of gene function in organs under various physiological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058643_12062025_PF_FP_ABST
    Figure US2024058643_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure generally relates to certain image-based techniques for detecting nucleic acids and proteins in a sample. The methods provided include simultaneous detection of RNAs and proteins using RCA-MERFISH technology.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] IMAGING-BASED POOLED GENETIC SCREENS WITH HIGHLY MULTIPLEXED PHENOTYPES

[0002] RELATED APPLICATIONS

[0003] This application claims the benefit under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 63 / 606,973, filed December 6, 2023, U.S. Provisional Application No. 63 / 650,996, filed May 23, 2024, and U.S. Provisional Application No. 63 / 699,689, filed September 26, 2024, the entire contents of each of which are incorporated herein by reference.

[0004] REFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0005] The contents of the electronic sequence listing (H049870810WO00-SEQ-KUW.xml; Size: 91,624 bytes; and Date of Creation: November 27, 2024) is herein incorporated by reference in its entirety.

[0006] BACKGROUND

[0007] A major problem in biology is understanding and controlling how individual genes or combinations of genes function within cells and tissues to produce distinct cellular phenotypes. To address this, we have developed several new approaches to perturb gene function at a very large scale in cells in their native position in the tissue of a living animal and to measure complex cellular and tissue-level phenotypes associated with these perturbations. Single-molecule fluorescent in situ hybridization (smFISH) is a powerful method for detecting individual nucleic acid molecules in cells. A fundamental limitation of smFISH, however, is its low throughput, typically only a few genes at a time. This low throughput is due to a lack of distinguishable probes with which to label cells and the cost of producing large amounts of labeled probe required for high efficiency staining. Thus, improvements in detecting molecules are needed. Further, sequencing-based approaches do not retain spatial and morphological information such as subcellular structures and tissue organization, which are crucial aspects of cellular and tissue physiology. Furthermore, methods for simultaneous detection of proteins and nucleic acids at high resolution are needed.

[0008] SUMMARY The present disclosure generally relates to certain image-based techniques for performing large-scale pooled genetic screens in native tissue with rich, multimodal phenotypic readouts, as well as its application to map the function of hundreds of genes in organs under multiple physiological conditions. In some embodiments, provided are methods that assess the effects of many genetic perturbations on diverse cellular phenotypes, including transcriptional state, subcellular morphology, and tissue organization in an organ. In some embodiments, provided herein are methods for determining spatial positions of nucleic acids in a cell or a tissue. In some embodiments, the methods further provide determination of spatial positions of proteins in the cell or tissue. In some embodiments, multiplexed simultaneous spatial position determinations of nucleic acids, e.g., RNAs, and proteins in a cell can be accomplished using the methods and composition described herein. In some embodiments, methods are provided that use fixed-cell perturbation (using CRIPSR technology) and sequencing as well as imaging-based pooled genetic screening in heavily fixed tissue. Advantageously, the methods allow joint sequencing and imaging analysis of diverse perturbations in the same tissue in certain embodiments. In some aspects, the methods provided apply multiplexed protein and RNA imaging, using in situ enzymatic probe amplification followed by multiplexed error-robust FISH (MERFISH) to read out both endogenous RNAs and short barcodes for genotyping. For example, through an integrated analysis of morphological and transcriptional phenotypes, using the methods described herein, novel regulators of hepatocyte zonation were identified, mechanisms how proteostatic stress pathway activation can broadly affect the expression of secreted proteins were revealed, and mechanisms how diverse cellular pathways can produce convergent effects on steatosis were demonstrated.

[0009] Certain aspects are generally directed to: 1. In vivo mosaic screens with in situ phenotypes; 2. Barcode readouts; 3. RCA-MERFISH; 4. Sequential immunofluorescence imaging for multiplexed cellular phenotype measurements; 5. Tissue preparation protocols for combining immuno staining and MERFISH; 6. Enzymatic synthesis of probes for RCA- MERFISH and other RNA-templated probe ligation methods; 7. Enzymatic production of amplifiers for protein and RNA imaging; or 8. Self-supervised analysis methods for highly multiplexed structural imaging.

[0010] In some aspects, provided is a method comprising: exposing a sample comprising a plurality of cells to oligonucleotide-conjugated antibodies and nucleic acid probes; fixing the sample in a gel; and determining spatial positions of the oligonucleotides and the nucleic acid probes within the plurality cells of the sample the gel. In some embodiments, the method further comprises: exposing the sample to adapter oligonucleotides, wherein the adapter oligonucleotides hybridize with the oligonucleotides of the conjugated antibodies and nucleic acid probes present in the sample.

[0011] In some embodiments, the oligonucleotides, nucleic acid probes, and adapter oligonucleotides are modified with an agent for attaching the oligonucleotides, nucleic acid probes, and adapter oligonucleotides to the gel.

[0012] In some aspects, provided is a method comprising: exposing a sample to oligonucleo tide-conjugated antibodies; exposing RNA and the oligonucleotides of the oligonucleo tide-conjugated antibodies in the sample to an agent for attaching RNA and the oligonucleotides to a gel; fixing the sample in a gel; and determining spatial positions of the RNA and the oligonucleotides within the gel.

[0013] In some aspects, provided is a method comprising: exposing a sample to oligonucleo tide-conjugated antibodies; fixing the sample in a gel; and determining spatial positions of nucleic acids within the gel, wherein the nucleic acids include the oligonucleotides in the conjugates.

[0014] In some embodiments, the nucleic acids are modified with an agent for attaching the oligonucleotide to a gel. In some embodiments, the agent for attaching the oligonucleotide to the gel is:

[0015] In some embodiments, the gel is an acrylamide gel. In some embodiments, the method further comprises clearing proteins and / or lipids from the sample. In some embodiments, the method further comprises retrieving antigens in the sample prior to exposing the sample to oligonucleo tide-conjugated antibodies. In some embodiments, the method further comprises exposing the sample to a plurality of nucleic acid probes.

[0016] In some embodiments, the plurality of nucleic acid probes comprises nucleic acid probes with different sequences. In some embodiments, the plurality of nucleic acid probes target RNA species and oligonucleotides of oligonucleo tide-conjugated antibodies in the sample. In some aspects, provided is a method comprising: exposing nucleic acids in a sample to a padlock probe comprising a barcode flanked by a 5’ homology arm and a 3’ homology arm, wherein the 5’ and 3’ homology arms bind to the nucleic acid; circularizing the padlock probe; generating an amplicon by rolling circle amplification using the circularized padlock probe as a template, wherein the amplicon comprises a plurality of barcode complements; and determining the spatial positions of the plurality of barcode complements using MERFISH.

[0017] In some embodiments, the methods described herein further comprise: exposing nucleic acids in a sample to a padlock probe comprising a barcode flanked by a 5’ homology arm and a 3’ homology arm, wherein the 5’ and 3’ homology arms bind to the nucleic acid; circularizing the padlock probe; generating an amplicon by rolling circle amplification using the circularized padlock probe as a template, wherein the amplicon comprises a plurality of barcode complements; and determining the spatial positions of the plurality of barcode complements using MERFISH.

[0018] In some aspects, provided is a method comprising: exposing a sample to oligonucleo tide-conjugated antibodies; fixing the sample in a gel; and determining spatial positions of the oligonucleotides and nucleic acids within the sample.

[0019] In some aspects, provided is a method comprising: introducing, into a sample comprising a plurality of cells, a nucleic acid that perturbs expression of at least one gene in the plurality of cells; exposing the sample to oligonucleo tide-conjugated antibodies and a plurality of nucleic acid probes; fixing the sample in a gel; and determining spatial positions of the oligonucleotides and the nucleic acid probes within the plurality of cells in the sample.

[0020] In some embodiments, the method further comprises: exposing the sample to adapter oligonucleotides, wherein the adapter oligonucleotides hybridize with the oligonucleotides of the conjugated antibodies and the nucleic acid probes present in the sample.

[0021] In some embodiments, the nucleic acid that perturbs expression of at least one gene comprises a guide portion, a reporter portion, and an identification portion. In some embodiments, the nucleic acid that perturbs expression of at least one gene comprises an open reading frame (ORF) of the at least one gene. In some embodiments, the spatial positions of the oligonucleotides and nucleic acid probes are determined in neighboring cells of the plurality of cells of the sample.

[0022] In another aspect, the present disclosure encompasses methods of making one or more of the embodiments described herein. In still another aspect, the present disclosure encompasses methods of using one or more of the embodiments described herein. Other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non-limiting embodiments of the disclosure when considered in conjunction with the accompanying figures.

[0023] BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Non-limiting embodiments of the present disclosure will be described by way of example with reference to the accompanying figures, which are schematic and are not intended to be drawn to scale. In the figures, each identical or nearly identical component illustrated is typically represented by a single numeral. For purposes of clarity, not every component is labeled in every figure, nor is every component of each embodiment of the disclosure shown where illustration is not necessary to allow those of ordinary skill in the art to understand the disclosure.

[0025] FIGs. 1A-1F show the workflow and optimizations for imaging-based screening. FIG. 1A depicts an experimental procedure for tissue preservation by PFA fixation and cryoprotection. FIG. IB is a detailed experimental protocol for multimodal oligo-conjugated antibody staining and RCA-MERFISH in fixed tissue. FIG. 1C shows different readout modalities by imaging or sequencing. FIG. ID is a diagram of padlock probe design for RCA-MERFISH with a Hamming Weight (HW) 6 code. The readout sequences (marked as bit 1 to bit 6), the presence of which determined the MERFISH code, were directly encoded in the padlock probe. FIG. IE is a diagram of the RCA-MERFISH signal amplification process. Probes that have both arms hybridized to target RNA and adjacent to each other were ligated, and then the ligated probes were amplified through rolling circle amplification (RCA). After RCA, the individual bits in each amplicon were read out through fluorescent microscopy, over multiple rounds of staining with fluorescent readout probes, and then dehybridizing (‘stripping’) the probes off with a high formamide concentration wash. FIG. IF is an alternative experimental protocol for multimodal oligo-conjugated antibody staining and RCA-MERFISH in fixed tissue.

[0026] FIG. 2 shows a schematic of a process of reversible staining in multiplexed imaging.

[0027] FIG. 3A-3E show RCA-MERFISH padlock probe production. FIG. 3A shows the RCA-MERFISH padlock probe design for pooled oligonucleotide synthesis prior to amplification and probe library synthesis. FIG. 3B shows the testing of different Type IIS restriction enzymes on single restriction digested, PAGE-purified probe against GFP made through phosphoramidite synthesis. FIG. 3C shows the RCA amplicons of digested RCA- MERFISH probe against GFP in U-2 OS cells expressing either GFP or mCherry. There were many more GFP amplicons in GFP-expressing cells than there were in mCherry-expressing cells, indicating the specificity of RCA-MERFISH. FIG. 3D shows the probe library preparation protocol using RCA followed by restriction digestion. Femtomolar pools of oligos synthesized in arrays were amplified by PCR with phosphorylated primers containing an Nt.BbvCI site. The PCR amplicons were then circularized and nicked with Nt.BbvCI. The nick site was used to initiate rolling circle amplification, and the RCA product was then digested with BccI and BciVI after annealing on primers. FIG. 3E shows a 4% agarose gel electrophoresis of digested single probes and concatenated probes produced through RCA synthesis.

[0028] FIGs. 4A-4J show the optimization of the RCA-MERFISH protocol. FIG. 4A shows Phi29 used in RCA degrades ssDNA smFISH probes. Left: smFISH signal in U-2 OS cells without Phi29 treatment at 37°C for 1 hr. Center: smFISH signal in U-2 OS cells pre-treated with Phi29 at 37°C for 1 hour before FISH staining. Right: smFISH signal in U-2 OS cells treated with Phi29 at 37°C for 1 hour after FISH staining. FIG. 4B shows dextran sulfate inclusion in hybridization buffer inhibits Phi29 enzymatic activity. Left: amplified RCA- MERFISH padlock probe against GFP in U-2 OS cells not expressing GFP, detected by readout probes complementary to readout sequences on the padlock. Center: amplified RCA- MERFISH probe against GFP in U-2 OS cells expressing GFP, with dextran sulfate in the hybridization buffer, detected by readout probes complementary to readout sequences on the padlock. Right: amplified RCA-MERFISH probe again GFP in U-2 OS cells expressing GFP, without dextran sulfate in the hybridization buffer, detected by readout probes complementary to readout sequences on the padlock. FIG. 4C shows the quantification of effect of Phi29 on ssDNA smFISH probes from (A), in number of spots / cell. FIG. 4D shows the quantification of effect of Phi29 on ssDNA smFISH probes from (A), in spot intensity. FIG. 4E shows the optimization of alternative crowding agents to dextran sulfate. Multiple additives to hybridization mixture, staining U-2 OS cells expressing either GFP or mCherry with a single probe against GFP, in terms of number of spots per cell, distinguishing GFP+ (signal) and mCherry+ (background) cells. Peg8k = Poly(ethylene glycol) average mol wt 8,000, Peg35k = Poly(ethylene glycol) average mol wt 35,000, Dex = unsulfonated dextran; all were added to the hybridization solution so the final w / v was at the indicated percent. FIG. 4F shows the quantification of increase in efficiency of different additives to hybridization mix from (Fig. 4E), relative to control (no additive to hybridization mix). FIG. 4G shows the optimization of RCA-MERFISH. Fresh-frozen and PFA-fixed tissue was tested with staining RNA after gel embedding (+ Post-gel stain) or before gel embedding (- Post- gel stain), with or without the addition of methacrylic acid NHS ester (MAA) along with MelphaX, with or without antigen retrieval, and varying the digestion temperature from 37°C to 60°C. Counts are number of amplicons per cell across the first two bits, detected in the Cy7 and Cy5 fluorescence channel, from a 120 gene RCA-MERFISH library. FIG. 4H shows the optimization of RCA-MERFISH across different decrosslinking conditions, quantified by amplicon counts per cell for a single bit, using an RCA-MERFISH library at low concentration (~0.1 nM / probe), hence the lower counts per cell than (Fig. 5G) which used ~10X higher probe concentration. The conditions tested were (1) no decrosslinking, (2) decrosslinking at 47° for 30 minutes, (3) decrosslinking at 60° for 30 minutes, (4) decrosslinking at 70° for 30 minutes, and (5) decrosslinking at 85° for 15 minutes. All decrosslinking was conducted in TE pH 9. FIG. 41 shows the optimization of immunofluorescence across different decrosslinking conditions, measuring total intensity per field of view for a Tomm20 antibody. The decrosslinking conditions were the same as in Fig. 5H. FIG. 4J shows the correlation of 209 gene RCA-MERFISH (after optimization) with bulk RNA-seq from the liver.

[0029] FIGs. 5A-5D show data processing pipeline and deep learning model architecture. FIG. 5A is a diagram of the data processing pipeline. Each panel of barcodes, endogenous RNA or morphological data (protein or RNA) that was measured was collected back-to-back in the same experiment and then processed in parallel. The RCA-MERFISH data was processed by first registering to common fiducials across multiple rounds, then decoding the identity of individual molecules. The molecules were then filtered using machine learning on features of molecules (mean intensity, size, variance, difference between mean on- and off-bit intensity), to obtain a final 5% false positive rate that decode to a blank barcode. In parallel, the polyA and Na+ / K+ ATPase channels of the morphological data were used to segment cells, which were then merged to eliminate duplicates of the same cells segmented in multiple fields of view. The cell segmentations were used to assign molecules to individual cells for quantification, and then the morphological channels were used with the segmentation mask to export the final images and per-gene quantification of expression for each cell. FIG. 5B is a diagram of final the annotated data matrix combining all features. FIG. 5C is a high-level diagram of VQ-VAE network across all channels. An input image was put into an artificial neural network that attempts to reconstruct the same image after passing the image through a low-dimensional bottleneck. In this case, two separate representations were created (top- and bottom-level) that attempted to capture different scales of features in the image. The bottom representation was formed first through one network, then further compressed to form a top representation with a separate autoencoder. The two representations were then concatenated and passed through a final network to reconstruct the original image. FIG. 5D shows the detail of the VQ-VAE network for each individual channel, trained simultaneously. A separate embedding was created for each morphological channel at the same time, using a mean squared error (MSE) loss to determine the accuracy of reconstruction. For each morphological channel, the embedding was used in a classification task to predict the identity of the protein or RNA being represented. The embeddings for each morphological channel were then concatenated and used to predict higher-level information about the type or state of each cell. When added to the overall training loss, these auxiliary predictive tasks were intended to constrain the representations that were formed by the VQ-VAE networks to capture salient features that discriminate different morphological channels and cell types or states.

[0030] FIG. 6A shows a schematic of a lentiviral vector used for dual-mode mosaic screens. The vector was derived from a CROP-seq vector and included an mU6 promoter driving sgRNA expression and a hepatocyte promoter driving expression of a mTurquoise transcript that contained a perturbation- specific barcode sequence in the 3’ UTR. FIG. 6B shows a construct of sgRNA expressed under a Pol III promoter and a barcode expressed in the 3’ UTR of a transcript containing a fluorescent protein ORF under a Pol II promoter.

[0031] FIG. 7 shows regulated pathways revealed through co-clustering of gene expression and perturbations using methods described herein.

[0032] FIG. 8A-8G show a multimodal investigation of the regulation of hepatocyte zonation. FIG. 8A shows kernel density estimate plots showing the distribution of zonation gene expression in single cells with control sgRNAs, sgRNAs targeting Ctnnbl, and sgRNAs targeting APC. The single cell zonation scores reflect the expression of periportal marker genes such as Cyp2f2 and Hal and pericentral marker genes such Glul and Cyp2el, with periportal expression contributing positively and pericentral expression contributing negatively to the score. The scores were defined relative to the mean expression in cells with control sgRNAs. FIG. 8B is a ranking of perturbed genes by their average impact on the zonal gene expression score. FIG. 8C shows a heatmap summarizing the categories of genes whose perturbation has a large impact on zonated gene expression. Here, the periportal and pericentral expression scores are shown separately. FIG. 8D is a heat map of the perturbation-perturbation correlation of pseudobulk transcriptional changes associated with the indicated sgRNAs. The gray shades represent the Pearson correlation coefficient of pseudobulk phenotypes between indicated perturbations in the Perturb-seq dataset. FIG. 8E is a diagram of data-driven zonal segmentation process. The proportion of each of the different hepatocyte subtypes was calculated in 50 pm x 50 pm bins. The bins were then grouped into two zones based on the local cell type distribution and the enrichment of cells with each perturbation in the two zones was quantified. FIG. 8F is a representative tissue section showing cell types from the RCA-MERFISH experiment (left) and the resulting spatial segmentation of periportal and pericentral zones (right). FIG. 8G is a barplot of the fraction of cells in periportal and pericentral zones (as defined above), for the indicated genetic perturbations. The white line represents the fraction of cells with control sgRNAs. No perturbations had significantly altered spatial zonation.

[0033] FIGs. 9A-9F show heterogeneity in subcellular morphology by cell type and state. FIG. 9A shows images of a single field of view imaged with a full subcellular morphology panel, comprising 4 abundant RNA species and 14 proteins. Proteins were labeled with oligoconjugated antibodies and imaged together with RNAs through sequential rounds of multifluorescence FISH. FIG. 9B are zoom-in images of a subset of subcellular morphologies. FIG. 9C shows a diagram of deep learning autoencoder to reduce dimensionality and featurize subcellular morphology images. Protein morphologies were reduced from images to a 512 dimensional embedding, using a VQ-VAE model. FIG. 9D shows a confusion matrix representing hepatocyte transcriptional subtype classification accuracy on held out test data, predicted using the transcriptomic data obtained from MERFISH imaging of 209 genes. FIG. 9E shows a confusion matrix representing hepatocyte transcriptional subtype classification accuracy on held out test data, predicted using the morphological feature embeddings obtained from imaging the 14 proteins and 4 abundant RNAs. FIG. 9F shows a heat map showing mutual information between each morphological channel at hepatocyte subtype identity (Hepl vs Hep6). Channels with higher KL divergence values had larger differences in the embeddings between Hepl and Hep6 cells.

[0034] FIG. 10A-10F show multimodal in vivo screening with strong phenotypes. FIG. 10A shows an example exploration of a spatial phenotype in the imaging dataset (right and zoom: cells with no called sgRNA are light gray). FIG. 10B shows an example exploration of a transcriptional phenotype in the Perturb-seq dataset. A UMAP representing individual cells was generated from the single cell transcriptomes of cells with sgRNAs targeting Hnf4a and from a random sub-sampling of cells with control sgRNAs. The UMAP is shaded by sgRNA identity (left) or by Apoal expression (right). FIG. 10C shows a heat map representation of pseudobulk transcriptional changes (measured by sequencing, left) and staining protein and RNA level changes (measured by imaging, right) associated with each sgRNA, relative to cells with control sgRNAs. The gray shades represent perturbed gene-level log2-fold RNA expression changes in cells with perturbed genes versus negative control cells (transcriptional sequencing data), or Z-normalized protein / RNA expression change in cells with perturbed genes relative to negative control cells (staining intensity imaging data) and are clipped for visual emphasis. FIG. 10D shows a heat map of the perturbation-perturbation correlation of RNA and protein changes associated with active sgRNAs, with zooms. The gray shades represent Pearson correlation of perturbed gene-level pseudobulk phenotypes in the sequencing dataset (below diagonal) or the imaging dataset (above diagonal). The genetic perturbations are ordered by single linkage hierarchical clustering of joint phenotype vectors that represent both transcriptional changes measured by sequencing and protein and RNA changes measured by imaging. FIG. 10E shows minimal distortion embedding where each dot represents an mRNA expressed in hepatocytes. mRNAs that are co-varying in expression across the genetic perturbations are placed in proximity. FIG. 10F is a heat map of the correlation between the expression levels of indicated proteins / RNAs across perturbations, in the imaging dataset. The proteins / RNAs are ordered by single linkage hierarchical clustering of the correlation matrix.

[0035] Figs. 11A-11M show the identifying genetic drivers of liver physiology. FIG. 110A shows a ranking of perturbed genes by their impact on a literature-defined set of lipid and cholesterol biosynthesis genes in the Perturb-seq experiment. The genes included in the lipid score include Hmgcsl, Sqle, and Fasn and the score reflects mean log2-fold change across this panel, relative to negative control cells. Light gray reflects FDR-corrected significance (corrected p < 0.05, Benjamini-Yekutieli correction). FIG. 11B shows a scatterplot showing the impact of perturbed genes on literature-defined sets of Unfolded Protein Response (UPR) and Integrated Stress Response (ISR) genes. The UPR score included genes such as Hspa5 and Herpudl; the ISR score included genes such as Atf4 and Ddit4. The scores were calculated relative to cells with control sgRNAs. FIG. 11C shows a ranking of perturbed genes by their impact on anti-GAPDH intensity in the imaging experiment, relative to mean intensity in cells with control sgRNAs. Light gray reflects FDR-corrected significance (corrected p < 0.05, Benjamini-Yekutieli correction). FIG. 11D shows a ranking of perturbed genes by their impact on Albumin mRNA FISH intensity in the imaging experiment, relative to mean intensity in cells with control sgRNAs. Light gray reflects FDR-corrected significance (corrected p < 0.05, Benjamini-Yekutieli correction). FIG. HE shows a ranking of perturbed genes by their impact on pre-rRNA FISH intensity in the imaging experiment, relative to mean intensity in cells with control sgRNAs. Light gray reflects FDR-corrected significance (corrected p < 0.05, Benjamini-Yekutieli correction). FIG. 11F shows a ranking of perturbed genes by their impact on anti-Phospho S6 ribosomal protein intensity in the imaging experiment, relative to mean intensity in cells with control sgRNAs. Light gray reflects FDR-corrected significance (corrected p < 0.05, Benjamini-Yekutieli correction). FIG. 11G shows a Leiden-clustered UMAP representation of CathB embeddings from the imaging experiment. Every point represents an individual cell. Sampled cells from each cluster are shown as insets. FIG. 11H shows a bar plot of enrichment in each of the clusters in Npcl knockout cells, relative to cells with control sgRNA. FIG. Ill shows a sampling of cells with control sgRNAs or Npcl sgRNAs showing fluorescence micrographs the anti- CathB channel alongside polyA FISH. FIG. 11J shows a schematic of the diet + genetic perturbation Perturb-seq experiment. Two diet conditions were tested, including food access (ad lib) and deprived (fasted) conditions. FIG. 11K shows a scatterplot comparing the energy distance vs control cells for each knockout in the ad lib and fasted Perturb-seq datasets. The two distances were correlated (Pearson’s R = 0.77). FIG. 11L shows a heatmap representing the number of differentially expressed genes between the indicated samples. Genes that were differentially expressed between the indicated conditions were defined with Benjamini-Hochberg-corrected, Mann-Whitney corrected p < 0.05. Cells were subsampled such that each of the comparisons included the same number of cells. FIG. 11M shows a heatmap representing Pearson correlations of pseudobulk transcriptional responses between the indicated knockouts, in the indicated condition. The knockout- specific transcriptional responses were calculated relative to cells with control sgRNAs from the same mouse. The phenotypes of these knockouts were more correlated in fasted tissue (mean Pearson’s R = 0.91, vs 0.19 in an ad libitum animal).

[0036] FIGs. 12A-12J show a multimodal investigation of stress response and steatosis. FIG. 12A shows a ranking of perturbed genes by their impact on anti-Calreticulin intensity in the imaging experiment, relative to mean intensity in cells with control sgRNAs. Light gray reflects FDR-corrected significance (corrected p < 0.05, Benjamini-Yekutieli correction). FIG. 12B shows a volcano plot showing intensity change and significance of the imaged protein and RNA channels between cells with sgRNA targeting Selll vs cells with control sgRNAs. Dash line indicated corrected p value = 0.05. FIG. 12C shows an unbiased sampling of cells with control sgRNAs and Selll sgRNAs, showing the anti-Calreticulin channel (top) and the Albumin mRNA FISH channel (bottom) alongside polyA FISH. FIG. 12D shows heatmaps showing log2-fold change in the expression of UPR genes in the sequencing experiment. FIG. 12E shows a stacked bar plot representing the proportion of the measured pseudobulk transcriptome composed of the ten most abundant secretory mRNAs, ranked according to their abundance in control cells. FIG. 13F shows a violin plot representing the fraction of each measured single-cell transcriptome that is composed of mRNAs encoding the ten abundant secretory mRNAs, from cells with sgRNAs targeting Selll and cells with control sgRNA. FIG. 13G shows a ranking of perturbed genes by their impact on anti-Perilipin intensity in the imaging experiment, relative to mean intensity in cells with control sgRNAs. Light gray reflects FDR-corrected significance (corrected p < 0.05, Benjamini-Yekutieli correction). FIG. 12H shows an unbiased sampling of cells with control sgRNAs and sgRNAs targeting the indicated gene, showing the anti-Perilipin channel alongside polyA FISH. FIG. 121 shows a single-link hierarchically ordered heatmap showing log2 fold change in the expression of lipid biosynthesis genes and Integrated Stress Response genes, for the indicated genetic perturbations. FIG. 12J shows a diagram illustrating convergent mechanisms that all caused lipid droplet accumulation.

[0037] DETAILED DESCRIPTION

[0038] Provided are methods and compositions for determining spatial positions of nucleic acids and proteins in a sample, e.g., a cell or a tissue. In some embodiments, the methods and compositions comprise simultaneously and / or sequentially labeling nucleic acids and / or proteins in a sample and simultaneously and / or sequentially determining spatial positions of nucleic acids and / or proteins.

[0039] The methods and compositions provided herein can be used to test the function of individual genes, e.g., on the state of individual cells in a particular tissue in a living organism and / or on the overall function and organization of a tissue. In some embodiments, methods and compositions are provided to assess spatial positions of nucleic acids and proteins in a cell or tissue following a perturbation delivered to the cell or tissue. In some embodiments, the methods provided comprise performing large-scale pooled genetic screens in native tissue with rich, multimodal phenotypic readouts. In some embodiments, the methods comprise mapping the function of hundreds of genes in organs under multiple physiological conditions. In some embodiments, provided are methods that assess the effects of many genetic perturbations on diverse cellular phenotypes, including transcriptional state, subcellular morphology, and tissue organization in an organ. In some embodiments, methods provided use fixed-cell perturbation (e.g., using CRIPSR technology) and sequencing as well as imaging-based pooled genetic screening in heavily fixed tissue. Advantageously, the methods enable joint sequencing and imaging analysis of diverse perturbations in the same tissue.

[0040] In some embodiments, the methods provided comprise applying multiplexed protein and RNA imaging, using in situ enzymatic probe amplification and performing multiplexed error-robust FISH (MERFISH) to read out both endogenous RNAs and short barcodes for genotyping. In some embodiments, methods provided herein comprise measuring multiple distinct dimensions of cellular states to capture a comprehensive picture of gene function.

[0041] In some embodiments, methods comprise delivering genome-wide in vivo CRISPR libraries to an organ, rapidly fix the organ and capturing multiple aspects of cellular state while simultaneously identifying perturbed genes using single-cell sequencing and / or imaging-based approaches described herein. In some embodiments, methods comprise singlecell sequencing of the perturbed cells in the organ and further comprise quantification of proteins and mRNA and assessing subcellular morphology using imaging. Advantageously, by linking cellular properties measured through different modalities via their shared perturbation identity, integrated multimodal analyses can be performed to examine the multifarious effects of the same perturbation on different aspects of cellular function.

[0042] In some cases, the perturbations may be delivered randomly to individual cells in tissues such as brain, liver, heart, or muscle, for example, by viruses such as lentivirus or AAV, with either a low or a high multiplicity of infection (MOI). The throughput of the technology allows such screens at a genome- wide scale in certain embodiments (e.g., knocking out all genes in the genome, or overexpress all proteins). In some embodiments, the perturbing viruses may contain a CRISPR sgRNA for gene knockout and / or an open reading frame (ORF) for overexpression. However, a variety of technologies that allow perturbations can be used, and some embodiments described herein may be agnostic to the specific type of perturbation being used. Non-limiting examples of potentially suitable genome engineering technologies include RNAi, TALENs, CRISPR KO, CRISPRi, CRISPRa, base editing, prime editing, ORF overexpression, and other technologies that allow pooled genetic perturbation screening.

[0043] In some embodiments, the methods comprise measuring cellular phenotypes through both sequencing and imaging in a single sample, the methods comprising using modified readout modality strategies and protocols such as is described herein. In some embodiments, the methods comprise contacting a cell, tissue and / or organ with sgRNA and, optionally, an endonuclease for targeted gene perturbations. In some embodiments, the methods comprise hybridizing cells with a transcriptome- wide library of split probes against different RNA species for transcriptome- wide detection and performing microfluidic encapsulation of dissociated single cells. In some embodiments, the methods comprise hybridizing pairs of split probes that hybridize adjacent to each other and ligating the split probes in a droplet. In some embodiments, the ligated split probes comprise a cellular barcode for sequencing and expression quantification. In some embodiments, the methods further comprise hybridizing custom probes directed to sgRNAs.

[0044] In one set of embodiments, a MERFISH-style in situ readout of the identity of the perturbation in each cell is provided, e.g., in its native position in fixed tissue. In some embodiments, a method comprises determining the phenotype of a targeted cell and its neighbors , for example, with multiplexed error-robust fluorescence in situ hybridization (MERFISH)- style spatial transcriptomics, with multiplexed immunofluorescence, and / or other techniques. See, e.g., Int. Pat. Apl. Pub. No. WO 2016 / 018960, incorporated herein by reference in its entirety.

[0045] In some embodiments, the methods provided comprise a series of tissue fixation, embedding, clearing, labeling and imaging steps. In some embodiments, the methods described herein are used in tissue samples fixed in a manner that is standard for high resolution immunofluorescence, in a single sample preparation and, e.g., on an automated microscope. In some embodiments, the methods comprise amplifying RNA FISH signals, both in order to increase the throughput of data collection and to allow the readout of short barcode sequences associated with the sgRNAs. In some embodiments, the methods comprise producing stable samples are stored and imaged at a later time, such that many samples are prepared in parallel and then imaged sequentially.

[0046] In some embodiments, the methods comprise performing rolling circle amplification (RCA) and performing multiplexed error-robust fluorescence in situ hybridization (MERFISH). In some embodiments, the methods comprise hybridizing padlock probes that target both endogenous mRNA and barcode RNA (e.g., sgRNA), ligating the padlock probes upon binding to mRNA, and amplifying the ligated padlock probes in situ. In some embodiments, each padlock probe serves as an encoding probe and comprises multiple readout sequences that together provide a MERFISH code for the target molecule. In some aspects, a method comprises detecting said readout sequences over multiple rounds of hybridization by readout probes complementary to the readout sequences.

[0047] In some embodiments, the methods provided comprise using heavily fixed, e.g., cross-linked tissues used for immunofluorescence imaging of proteins. In some embodiments, the methods comprise embedding the RNA of the heavily fixed tissue in a hydrogel. In some embodiments, the methods comprise clearing the tissue. In some embodiments, the methods comprise embedding of RCA-amplicons in a polyacrylamide gel.

[0048] In some embodiments, the methods provided comprise providing a heavily fixed, e.g., cross-linked tissue, embedding the RNA of the heavily fixed tissue in a hydrogel, clearing the tissue, hybridizing padlock probes that target both endogenous mRNA and, optionally, barcode RNA (e.g., sgRNA), ligating the padlock probes upon binding to mRNA, and amplifying the ligated padlock probes in situ using RCA, and embedding the RCA-amplicons in a polyacrylamide gel. In some embodiments, the methods provided further comprise multiplex immunolabeling proteins with oligonucleotide-conjugated antibodies. In some embodiments, the methods further comprise sequentially hybridizing complementary readout probes and detecting oligonucleotide-conjugated antibodies with the complementary readout probes. In some embodiments, the methods provided further comprise sequentially performing FISH and detecting abundant cellular RNAs through sequential FISH.

[0049] In some embodiments, the methods provided comprise perfusing a tissue, cryopreserving the tissue, cutting the tissue into thin sections, decrosslinking the tissue, and staining the tissue with a pool of oligonucleotide-conjugated antibodies targeting different subcellular organelles, membrane, and proteins.

[0050] In some embodiments, the methods comprise treating the tissue with an alkylating agent containing an acrylamide moiety (e.g., MelphaX) to modify all RNA and DNA molecules in the tissue.

[0051] In some embodiments, the methods comprise embedding the sample in a thin acrylamide gel, such that the RNA and DNA, including cellular RNAs, sgRNAs with barcodes, and antibody-associated DNA oligonucleotides are covalently linked to the acrylamide gel. In some embodiments, the methods comprise digesting the proteins and washing away lipids to both clear the tissue and make the remaining RNA and DNA accessible to enzymes and probes, and hybridizing a library of padlock probes targeting both endogenous RNA and perturbation barcodes. In some embodiments, the methods comprise washing the sample stringently. In some embodiments, the methods comprise ligating the RNA-hybridized padlock probes and performing rolling circle amplification. In some embodiments, the methods comprise decrosslinking. In some embodiments, the methods comprise optimizing the decrosslinking conditions. In some embodiments, the methods comprise modifying oligonucleotides. In some embodiments, the methods comprise optimizing additives. In some embodiments, the methods comprise optimizing digestion conditions to identify optimal conditions that allow for amplification while being compatible with immunostaining of proteins with oligonucleo tide-conjugated antibodies in heavily fixed tissue.

[0052] In some embodiments, the methods comprise providing multiple padlock probes which each target RNA species by tiling the transcript, wherein each padlock probe comprises two binding arms that hybridize adjacent to each other. In some embodiments, the method further comprises ligating the adjacent binding arms, and wherein each padlock probe comprises 4-6 readout sequence for encoding the target RNA. In some embodiments, the padlock probes are relatively long (-140-170 nt) and must be full-length in order to ligate properly.

[0053] In some embodiments, multiplexed error-robust fluorescence in situ hybridization (MERFISH)- style spatial transcriptomics are combined with multiplexed immunofluorescence. In some embodiments, a method comprises using machine learning techniques (including, e.g., CNNs, RNNs, image transformers, etc.) to generate meaningful embeddings of the morphologies in the multiplexed imaging data and / or to interpret both cell autonomous and cell-non-autonomous phenotypes.

[0054] In some embodiments, a method comprises assessing the distribution of cells / local cell neighborhoods that received a non-targeting control virus in comparison to the distribution of cells / local cell neighborhoods that received each of the perturbing viruses. In some embodiments, a method comprises screening in vivo with a spatially resolved, in situ readout. In some cases, imaging-based screening may include lower cost per measurement, allowing for much larger-scale measurements than is typically feasible using sequencing. In some embodiments, a method comprises determining, qualitative and / or quantitative, both transcriptomic and protein-based phenotypes, where the latter are measured using high- resolution imaging of antibody staining. In some embodiments, a method comprises antibody staining performed using oligonucleotide-conjugated antibodies. In some cases, a method comprises determining both cell-autonomous and cell-non-autonomous phenotypes. In some cases, a method comprises resolving phenotypes that vary as a function of spatial position with a tissue. The ability to image proteins and other subcellular features in addition to measuring gene expression is an advantageous feature of the methods described herein. In some embodiments, a method comprises measuring changes in subcellular morphology that are indicative of different healthy (e.g., mitochondrial, ER, lysosome structure, structural protein organization, etc.) and diseased states of cells (e.g. lipid droplet accumulation, formation of autophagosomes, etc.). In some embodiments, a method comprises imaging and / or measuring cellular phenotypes based on RNA (e.g. nuclear speckles), and / or synthetic reporters of cellular state (e.g. GFP expression in response to cellular stimulation).

[0055] In some embodiments, systems and methods such as those described herein comprise the use of directly compatible other imaging-based phenotyping modalities that are commonly performed in vivo, such as intravital live cell structural (e.g., synapses, organelle positions,) and time-lapse functional imaging (e.g. calcium signaling, membrane electrical potential, protein trafficking, cell migration, cell division, morphological changes). In some embodiments, systems and methods such as those described herein comprise the use in an in vitro context - for example, in iPS-derived cell types or cultured cancer cells -combined with live cell imaging or other techniques.

[0056] In some embodiments, the methods provided comprise mapping transcriptional gradients in cells. In some embodiments, the methods comprise imaging and scRNA- sequencing to map the spatial organization of distinct cell types and cell states in an organ.

[0057] To achieve a mosaic distribution of cells, in some embodiments methods provided herein use viral delivery of pools of genetic perturbation reagents. Examples include adeno associated virus (AAV) and lentivirus (LV) to deliver pooled perturbations. AAV is a small, replication-deficient, non-enveloped single-stranded DNA virus from the Parvovirus family that does not integrate in the host genome, but instead expresses episomally in the nucleus. AAVs can be administered systemically or injected locally into animals. Different serotypes of AAVs have different tropisms for specific organs and cell types; for example, AAV9 can cross the blood-brain barrier to infect brain cells and AAV8 is enriched for targeting the liver. LV is an enveloped RNA virus that integrates into the host genome after reverse transcription. LV can be delivered to infect cells throughout the liver via systemic delivery into the bloodstream, particularly in neonatal animals. Those of ordinary skill in the art will be aware of methods of delivering AAVs or LVs.

[0058] Lor both AAV and LV delivery, in some embodiments, provided are methods comprising molecular cloning techniques to generate a plasmid library of different genetic perturbations (e.g. CRISPR guides with associated DNA barcodes), and, e.g., propagate and expand the plasmid library in bacteria. In some embodiments, a method comprises producing viral libraries by transfecting mammalian cell culture with the perturbation plasmid library as well as virus- specific helper plasmids. In some embodiments, the method further comprises after transfection, isolating LV from the cell culture supernatant and concentrating by ultracentrifugation; isolating AAV from the cell pellet and purifying by ultracentrifugation. In some embodiments, a method comprises introducing the perturbations with viral vectors, and / or with other approaches for pooled in vivo screens, such as hydrodynamic injection of DNA, LNP mRNA delivery, VLP mRNA delivery, or fully transgenic approaches such as iMAP.

[0059] In some embodiments, a method comprises administering the resulting viral library to animals at a sufficient dose, for example, having a desired multiplicity-of-infection (MOI; i.e., the number of expressed viral particles per infected cell). For example, the desired MOI is less than 1, e.g., so that most cells have at most one gene perturbed. In some embodiments, a method comprises using an MOI less than 0.9, less than 0.7, less than 0.5, less than 0.3, less than 0.1, or less. However, in some embodiments, a method comprises using a MOI greater than 1. For instance, a method comprises using a MOI greater than 1.1, greater than 1.5, greater than 2, greater than 3, greater than 5, greater than 7, or more. In some embodiments, a method comprises using a MOI greater than lin various applications, such as for combinatorial genetic screens that examine epistatic interactions between genes, for cellular reprogramming experiments that require overexpressing multiple transcription factors in the same cell. Because the identities of all perturbations in a cell are measured at single-cell resolution in the methods provided herein, the methods comprise the use of higher MOI for more efficient genetic screens.

[0060] In one set of embodiments, systems and methods such as those described herein comprise performing screens in the liver, for example, using both AAV and lentivirus to deliver CRISPR perturbations or ORF overexpression of dominant negative stress response genes. In some embodiments, methods comprise using other perturbation techniques, including any of those described herein. In some embodiments, a method determines, e.g., assesses, measures, or quantifies various phenotypes. For example, in some embodiments, a method determines phenotypes involved in metabolism and organismal metabolic state changes (e.g. high fat diet, fasting, refeeding).

[0061] In some embodiments, methods provided herein comprise using high-throughput genetic screens in a native tissue context, e.g., without dissociating individual cells. In some embodiments, a method comprises using a screen and determining a phenotype that is an arbitrary pattern of gene expression or protein localization. Unlike a cultured cellular system, a method provided herein comprises determining gene function in cells within their native physiological context. In some embodiments, a method comprises determining, assessing, measuring, or quantifying a variety of developmental or disease contexts allowing insights into the genes controlling these processes. In some embodiments, methods provided herein comprise performing lower cost pooled in vivo genetic screens. In some embodiments, methods provided comprise systematically testing gene function using larger scale experiments, rather than testing the function of only at most a few dozen genes.

[0062] In some embodiments, methods comprise studying gene function in a real physiological context within an intact animal. In some aspects, the methods provided have implications for target discovery, disease modeling, basic biology, or other applications. For example, in some embodiments, methods are provided that identify genes that sensitize cells to the effects of particular disease models and determine potential new targets for drugs. In some embodiments, methods are provided comprising identifying genetic perturbations that change the state of cells in a disease context to a healthier phenotype and identifying new genetic therapies. In some embodiments, a method comprises systematically measuring the effects of perturbing thousands of different genes in different cell types, and providing datasets for machine learning models that and predicting effects of different genetic interventions without a need to perform experiments. In some aspects, the methods provided herein further comprise performing therapeutic design or cellular engineering.

[0063] Barcode readout

[0064] In some embodiments, methods are provided for determining the genotype of cells, the methods comprising introducing a relatively short DNA “barcode” indicating the identity of the genetic perturbation into a cell and determining or “reading out.” The terms “read out,” or “reading out,” as used herein, refer to the process of removing information from a device, e.g., a computer or signal detector, and displaying the information in an understandable form, e.g., an image, a schematic or a graph. In some embodiments, a method comprises using techniques such as Rolling Circle Amplification-MERFISH (RCA-MERFISH) (described below) for a high degree of amplification, e.g., to readout the identity of barcodes in situ in tissue (e.g., 180-mer barcodes, or barcodes of other lengths). Although some of the examples described herein use RCA-MERFISH, the present disclosure is not so limited, and other versions of MERFISH may be used, e.g., branched tree amplification (capable of detecting sequencing > 200 nt in length).

[0065] In order to express these barcodes, in some embodiments, various molecular biology methods are provided to clone libraries of genetic constructs that both express a CRISPR sgRNA or ORF and a DNA barcode. For example, in one embodiment, for CRISPR guides, a method comprises synthesizing the sgRNA and barcode together as a single oligonucleotide, and performing a two-step cloning process such that the sgRNA is expressed under a Pol III promoter and the guide is expressed in the 3’ UTR of a transcript containing a fluorescent protein ORF under a Pol II promoter (FIG. 1). This allows for highly scalable synthesis of CRISPR guide libraries that are compatible with RCA-MERFISH based readout.

[0066] RCA-MERFISH.

[0067] In certain aspects, provided are methods comprising an amplified version of MERFISH spatial transcriptomics with certain advantages for screening, called RCA- MERFISH. The term “RCA-MERFISH,” refers to a process that combines aspects for spatial transcriptomics such as MERFISH with rolling circle amplification. In some aspects, methods provided using RCA-MERFISH comprise detecting highly multiplexed in situ RNA for both endogenous RNA and barcodes associated with perturbations. In some embodiments, the methods comprise using an amplified signal that is amenable to fast imaging speeds, and has specificity to the requirement for ligation of adjacent ends of a probe, and using massively multiplexed detection by MERFISH. In some embodiments, the methods comprise using read out through fast and inexpensive probe hybridization and stripping.

[0068] In one embodiment, a method comprising the use of RCA-MERFISH comprises generating 5’ phosphorylated DNA padlock probes from oligonucleotide pools by pooled enzymatic synthesis. In some embodiments, the method further comprises fixing the sample, generating an RNA-retaining hydrogel, and hybridizing the padlock probes to the endogenous RNA transcripts in cells or tissues and / or barcode RNA transcripts associated with genetic perturbations. In some embodiments, a method comprises circularizing the probes in a template-dependent manner by an RNA-templated DNA ligase and amplifying using a strand-displacing polymerase. In some embodiments, a method further comprises reading out the identity of each amplified molecule on a microscope using MERFISH through sequential rounds of hybridization of short, fluorescently labeled oligonucleotide readout probes followed by probe stripping. In some embodiments, a method comprises varying various aspects of the chemistry for this approach, including the probe concentration, ligation and amplification conditions, readout conditions, hybridization buffer additives, etc. However, in some aspects, the method comprises using relatively efficient detection of both nucleic acids, e.g., RNAs and proteins in a sample. In some embodiments, the method comprises performing a sequence of steps in sample preparation and labeling for multiplexed detection by RCA-MERFISH of both RNAs and proteins in a sample. Several features distinguish RCA-MERFISH from previous approaches for imagingbased spatial transcriptomics.

[0069] In some embodiments, a method comprises generating padlock probes using enzymatic chemistry. In some cases, a method comprises chip- synthesizing an oligo pool template, followed by enzymatic amplification as compared to direct chemical synthesis of these probes. Advantageously, a method comprising the use of a chip- synthesized oligo pool template makes it easier to obtain large probe-sets and measure many genes.

[0070] In some embodiments, methods provided comprise using longer RCA probes due to the high efficiency and low truncation rate of enzymatic probe synthesis vs chemical probe synthesis. In some embodiments, a method comprises direct encoding of the probe identity as an array of readout bit annealing sequences directly in the RCA probe, simplifying, and accelerating readout protocols.

[0071] In some embodiments, a method comprises using robust and compact MERFISH codes for the labeling ligation specificity and amplification that is obtained through enzymatic approaches. Using MERFISH decoding, in some embodiments, a method comprises measuring only a subset of amplicons on every round of imaging, and in some embodiments, the method comprises controlling the sparsity of this subset by the number of readout “bits” used in the codebook.

[0072] Sequential immunofluorescence imaging for multiplexed cellular phenotype measurements In order to capture both RNA and protein phenotypes, certain systems and methods provided herein comprise employing multiplexed, sequential immunofluorescence and RCA- MERFISH imaging. In some embodiments, the systems and methods comprise using multiplexed, simultaneous immunofluorescence and RCA-MERFISH imaging. In some embodiments, a method provided herein comprises using simultaneous imaging, the method comprising using oligonucleo tide-conjugated antibodies that bind, e.g., cellular proteins and primary nucleic acid probes, e.g., padlock probes that hybridize to, e.g., cellular RNAs. In some embodiments, a method comprises using secondary probes that bind simultaneously to the oligonucleotides conjugated to the antibodies and the primary probes hybridized to the RNAs. In some embodiments, a method comprises using primary oligonucleotide-conjugated antibodies that are stained as a pool, and then reading out through sequential rounds of imaging alongside measurements of RNA via MERFISH. In some embodiments, a method comprises using oligonucleotides conjugated to antibodies, using secondary nucleic acid probes described herein that bind the antibody conjugated oligonucleotides, and reading out in a sample at the same time as reading out secondary nucleic acid probes bound to primary probes, e.g., padlock probes present in the sample. For example, in some embodiments, a method comprises binding oligonucleotide-conjugated antibodies to proteins present in a sample and hybridizing padlock probes to endogenous RNA transcripts in a sample, e.g., a cell or tissue and / or a barcode RNA transcript associated with a genetic perturbation. In some embodiments, a method comprises amplifying padlock probes, e.g., by rolling circle amplification (RCA). In some embodiments, a method comprises adding secondary probes to the sample to hybridize to the oligonucleotides conjugated to the antibodies and to the amplified padlock probes. In some embodiments, a method comprises reading out the identity of each nucleic acid amplified from a padlock probe and some or all of the proteins bound by an oligonucleotide-conjugated antibody on a microscope using MERFISH through sequential rounds of hybridization of short, fluorescently labeled oligonucleotide readout probes, e.g., bound to the secondary probes followed by probe stripping. In some embodiments, a method comprises using the sequence of steps in the sample preparation and labeling procedure for successful detection of both nucleic acids and proteins in a sample.

[0073] In one embodiment, a method comprises using Phi29 enzyme for RCA. In some cases, the Phi29 enzyme used for the RCA can degrade single stranded DNA, e.g., oligonucleotides conjugated to antibodies. Therefore, in some embodiments, a method comprises reducing or preventing degradation, and / or allowing multimodal measurements of RNAs and proteins. In one embodiment, a method comprises labeling an antibody with an oligonucleotide that has one or more 3’ phosphorothioate linkages, a chemical modification to DNA that blocks the exonuclease activity of Phi29, and using a blocked 3’ hydroxyl to prevent the extension of these barcodes. In some embodiments, a method comprises modifying the 5’ end of the oligonucleotide with an acrydite moiety, e.g., having a structure (A) (below):

[0074] (A).

[0075] In some embodiments, the acrydite moiety is Acrydite™. In some embodiments, a method comprises chemically attaching an acrydite moiety to an acrylamide or acrylate gel during polymerization, thereby maintaining the position of the oligonucleotide during proteolytic clearing of the sample, that is the position of the oligonucleotide conjugated to an antibody bound to a protein in a specific position, e.g., within a cell. In another embodiment, a method comprises hybridizing oligonucleotide amplifiers (either a single amplifier or branched chain amplifiers) to a stained and post-fixed oligonucleotide-conjugated antibody. These amplifiers may be 5’ acrydite-modified and 3’ phosphoro thio ate-modified and blocked in certain embodiments, which may allow these amplifiers to themselves be crosslinked into the gel. These amplifiers may or may not be modified with MelphaX.

[0076] In another embodiment, a method comprises hybridizing short blocking probes to the single- stranded DNA (ssDNA) oligonucleotides on the antibodies. In some embodiments, the method comprises double- stranding the oligonucleotides attached to the antibodies with the short probes, preventing degradation. In some cases, the method comprises 3’ phosphorothioate-modifying and blocking the short probes. In some embodiments, the method further comprises stripping off the short probes prior to imaging.

[0077] Certain embodiments are directed to tissue preparation protocols for combined immuno staining of oligonucleotide-conjugated antibodies and MERFISH. Some embodiments are directed to the capture of both high-dimensional RNA and protein phenotypes, along with short RNA barcodes that give information about the genotype (e.g. genetic perturbations) of each measured cell.

[0078] In one set of embodiments, a method comprises using enzymatic reactions (ligation, rolling circle amplification, etc.) in situ in an acrylamide gel. Advantageously, the methods provided herein in some embodiments allow use of enzymatic reactions without interfering with the labeling and readout of the RNA and protein phenotypes. In some aspects, a method comprises several components that differ from the original MERFISH protocol: (1) Using an optimized antigen retrieval + antibody staining protocol that allows obtaining good immuno staining results in heavily crosslinked tissue, (2) Pre-labeling of RNA using the chemical reagent MelphaX, which chemically modified RNA to add an acrydite moiety, e.g., Acrydite™ allowing the RNA to be covalently attached to an acrylamide gel and survive harsh protein digestion and detergent conditions. (3) Staining of adapters onto the oligonucleotide-conjugated antibodies to prevent degradation by Phi29 enzyme and allow for signal amplification and gel anchoring, as described above, (3) Using acrylamide gelation of RNA and proteins and digestion before labeling of RNAs with oligonucleotide probes, (4) Staining of gel-embedded RNA with probes for RCA-MERFISH, (5) Performing in gel ligation and RCA amplification of probes, (6) Performing in gel readout using short fluorescently-conjugated oligonucleotides. Exemplary work-flow diagrams are shown in FIGs. 1A, IB, and IF.

[0079] Sample Preparation

[0080] A sample may be a cell, or a tissue. In some embodiments, the tissue may arise from a living animal, for example, a human or a non-human animal. For example, the animal may be a mammal, an invertebrate, a fish, an amphibian, a reptile, a bird, etc. Non-limiting examples of mammals include a monkey, cow, sheep, goat, horse, rabbit, pig, mouse, rat, dog, or cat. In some cases, the sample may be a relatively thick sample, e.g., samples having a thickness of at least 10 micrometers, at least 20 micrometers, at least 30 micrometers, at least 40 micrometers, at least 50 micrometers, at least 60 micrometers, at least 70 micrometers, at least 80 micrometers, at least 90 micrometers, at least 100 micrometers, at least 110 micrometers, at least 125 micrometers, at least 150 micrometers, at least 175 micrometers, at least 200 micrometers, at least 225 micrometers, at least 250 micrometers, at least 300 micrometers, at least 350 micrometers, at least 400 micrometers, at least 450 micrometers, at least 500 micrometers, etc.

[0081] For example, in some embodiments, a method provided herein comprises fixing a sample prior to introducing a nucleic acid probe and / or an oligonucleotide-conjugated antibody. Techniques for fixing samples comprising cells or tissues are known to those of ordinary skill in the art. As non-limiting examples, a sample may be fixed using chemicals such as formaldehyde, paraformaldehyde, glutaraldehyde, ethanol, methanol, acetone, acetic acid, or the like. In one embodiment, a sample may be fixed using Hepes-glutamic acid buffer-mediated organic solvent (HOPE). In some embodiments, a method comprises performing an antigen retrieval step on a fixed or unfixed sample prior to adding an oligonucleotide-conjugated antibody and / or a nucleic acid probe. Techniques for antigen retrieval are known to those of ordinary skill in the art. As non-limiting examples, enzymatic / proteolytic, high pH buffer and / or heat-induced antigen retrieval can be used.

[0082] In some embodiments, a method comprises using an antibody in a method described herein including any antibody suitable to bind a target on a surface of a cell in a sample, a target inside a cell in a sample, or a target between cells of a sample. Antibodies useful in a method described herein include antibodies that bind proteins, lipids, carbohydrates and / or nucleic acids in a sample.

[0083] In some embodiments, a method comprises using antibodies conjugated to oligonucleotides in the methods described herein. Methods for conjugating oligonucleotides, e.g., single- stranded DNA oligonucleotides to antibodies are known to a person of ordinary skill in the art and can be used by such person based on the descriptions provided herein.

[0084] In some embodiments, a method comprises exposing a sample, e.g., a cell and / or tissue, to an oligonucleo tide-conjugated antibody. In some embodiments, a method comprises exposing an unfixed sample, e.g., a cell and / or tissue, to an oligonucleotide-conjugated antibody. In some embodiments, a method comprises exposing a fixed sample, e.g., a cell and / or tissue, to an oligonucleotide-conjugated antibody.

[0085] In some embodiments, a method comprises fixing a sample after the sample has been exposed to an oligonucleotide-conjugated antibody. In some embodiments, a method comprises fixing a sample in a gel after the sample has been exposed to an oligonucleotide- conjugated antibody. In some embodiments, techniques for fixing as described above can be used.

[0086] In some embodiments, a method comprises exposing a sample to an agent for attaching a nucleic acid and / or oligonucleotide to a gel. In some embodiments, a method comprises modifying an oligonucleotide of an oligonucleotide-conjugated antibody with an agent for attaching the oligonucleotide to a gel. In some embodiments, a method comprises modifying a nucleic acid present in a sample, e.g., a RNA, with an agent for attaching the nucleic acid, e.g., RNA to a gel. In some embodiments, an agent for attaching a nucleic acid to a gel is a molecule of structure A (below):

[0087] In some embodiments, the agent for attaching a nucleic acid and / or oligonucletoide to a gel is Acrydite™. In some embodiments, a sample may be modified with MelphaX (see below):

[0088] MelphaX

[0089] In some embodiments, a method comprises exposing a sample to an oligonucleotide- conjugated antibody and further exposing the sample to the agent for attaching a nucleic acid and / or oligonucleotide to a gel, e.g., a molecule of structure A. In some embodiments, a method comprises exposing a sample to a nucleic acid probe and further exposing the sample to an agent for attaching a nucleic acid and / or oligonucleotide to a gel, e.g., a molecule of structure A.

[0090] In some embodiments, a method comprises exposing a sample to an adaptor oligonucleotide. In some embodiments, an adaptor oligonucleotide can be an oligonucleotide amplifier (either a single amplifier or branched chain amplifier) that hybridizes to an oligonucleotide conjugated to an antibody. In some embodiments, a method comprises modifying an adaptor oligonucleotide with an agent for attaching the adaptor oligonucleotide to a gel. For example, the adaptor oligonucleotide may be modified with a molecule of structure A. In some embodiments, a method comprises using an adaptor oligonucleotide that is 5’ modified with a molecule of structure A and / or 3’ phosphorothioate-modified and blocked. In some embodiment, a method comprises modifying an adaptor oligonucleotide with MelphaX. In some embodiment, a method comprises not modifying an adaptor oligonucleotide with MelphaX. In some embodiments, a method comprises not stripping off the adaptor oligonucleotide prior to imaging. In some embodiments, a method comprises attaching an adaptor oligonucleotide 5’ modified with a molecule of structure A to a gel. In some embodiments, a method comprises binding or hybridizing an adaptor oligonucleotide with a nucleic acid probe described herein for determination, e.g., of a spatial position of an oligonucleotide conjugated to an antibody. In some embodiments, a method comprises hybridizing a short blocking probe to an oligonucleotide conjugated to an antibody. The short probe makes the oligonucleotide attached to the antibody double- stranded and prevents the degradation of the oligonucleotide. In some embodiments, a method comprises using short probes that are 3’ phosphorothioate-modified and blocked. In some embodiments, the short probes are stripped off prior to imaging.

[0091] In some embodiments, a method comprises fixing a sample in a gel. In some embodiments, a method comprises fixing a sample by embedding the sample in a gel. In some embodiments, a gel is an acrylamide gel. In some embodiments, a method comprises attaching a nucleic acid, an oligonucleotide of an oligonucleotide-conjugated antibody, an adaptor oligonucleotide and / or a nucleic acid probe present in the sample to the gel by embedding the sample in a gel.

[0092] In some embodiments, a method comprises fixing a spatial position of a nucleic acid and / or oligonucleotide conjugated to an antibody in a sample in a gel by an interaction of a molecule of structure A in the nucleic acid and / or oligonucleotide in the sample with the gel.

[0093] In some embodiments, a method comprises attaching a molecule of structure A to an end of a nucleic acid, antibody-conjugated oligonucleotide, adaptor oligonucleotide, and / or nucleic acid probe present in the sample. In some embodiments, a method comprises attaching a molecule of structure A to a 5’ end of a nucleic acid, antibody-conjugated oligonucleotide, adaptor oligonucleotide, and / or nucleic acid probe present in the sample. In some embodiments, a method comprises attaching a molecule of structure A to a 3’ end of a nucleic acid, antibody-conjugated oligonucleotide, adaptor oligonucleotide, and / or nucleic acid probe present in the sample. In some embodiments, a method comprises attaching a molecule of structure A to a non-end portion, e.g., a middle portion of a nucleic acid, antibody-conjugated oligonucleotide, adaptor oligonucleotide, and / or nucleic acid probe present in the sample.

[0094] In some aspects, a method comprises digesting a gel-embedded (e.g., a gel-fixed) sample to remove proteins. In some embodiments, a method comprises digesting a gel- embedded sample with a proteolytic enzyme. In some embodiments, a method comprises further treating a gel-embedded sample with a non-ionic detergent.

[0095] In some embodiments, a method comprises exposing a sample prepared as described above to a variety of nucleic acid probes to determine one or more nucleic acids and / or one or more oligonucleotide conjugated to an antibody within a cell of the sample. The probes may comprise nucleic acids (or entities that can hybridize to a nucleic acid, e.g., specifically) such as DNA, RNA, LNA (locked nucleic acids), PNA (peptide nucleic acids), or combinations thereof. In some embodiments, additional components may also be present within the nucleic acid probes, e.g., as discussed below. Any suitable method may be used to introduce nucleic acid probes into a cell. For example, a method comprises adding a nucleic acid probe to a sample by flowing a fluid containing the nucleic acid probe around the cell. In some embodiments, a method comprises permeabilizing the cell such that the nucleic acid probe is introduced into the cell by flowing a fluid containing the nucleic acid probe around the cell. In some embodiments, a method comprises sufficiently permeabilizing the cell as part of a fixation process; in other embodiments, a method comprises permeabilizing the cell by exposure to certain chemicals such as ethanol, methanol, Triton, or the like. In addition, in some embodiments, a method comprises using techniques such as electroporation or microinjection to introduce a nucleic acid probe into a cell or other sample.

[0096] In some embodiments, a method comprises adding a nucleic acid probe to the sample after the sample has been embedded in a gel. In some embodiments, a method comprises adding a nucleic acid probe to the sample after an oligonucleo tide-conjugated antibody has been added to the sample and the sample has been embedded in a gel. In some embodiments, a method comprises adding a nucleic acid probe to a sample prior to exposing the sample to an oligonucleotide-conjugated antibody. In some embodiments, a method comprises adding a nucleic acid probe to a sample prior to embedding the sample in a gel.

[0097] In some embodiments, a method is performed in the following order of steps: first an oligonucleotide-conjugated antibody is added to the sample. Next, the sample is embedded in the gel. Then, a nucleic acid probe is added to the sample.

[0098] Nucleic Acid Probes

[0099] In some embodiments, a method provided herein comprises exposing nucleic acids in a sample to a padlock probe comprising a barcode flanked by a 5’ homology arm and a 3’ homology arm, wherein the 5’ and 3’ homology arms bind to the nucleic acid; circularizing the padlock probe; generating an amplicon by rolling circle amplification using the circularized padlock probe as a template, wherein the amplicon comprises a plurality of barcode complements; and determining the spatial positions of the plurality of barcode complements using MERFISH.

[0100] In some embodiments, a probe, e.g., a padlock probe may comprise any of a variety of entities that can hybridize to a nucleic acid, typically by Watson-Crick base pairing, such as DNA, RNA, LNA, PNA, etc., depending on the application. The nucleic acid probe, e.g., the padlock probe typically contains a target sequence that is able to bind to at least a portion of a target nucleic acid, in some cases specifically.

[0101] A “target nucleic acid,” as used herein, refers to both a nucleic acid present in a cell, e.g., a RNA and an oligonucleotide sequence of an oligonucleotide conjugated to an antibody present in a cell. A “target,” as used herein, refers to any molecule bound by a nucleic acid probe, or an antibody as described herein.

[0102] In some embodiments, a method comprises introducing into a cell or other system a nucleic probe, e.g., a padlock probe that binds to a specific target nucleic acid (e.g., an mRNA, or other nucleic acids as discussed herein) and / or an oligonucleotide of an oligonucleo tide-conjugated antibody and / or an adaptor oligonucleotide described herein. In some embodiments, a method comprises introducing a nucleic acid probe, e.g., a padlock probe that binds only to a target nucleic acid that is, e.g., a RNA present in a cell and not to an oligonucleotide of an oligonucleo tide-conjugated antibody. In some embodiments, a method comprises introducing a nucleic acid probe, e.g., a padlock probe that binds only to an oligonucleotide of an oligonucleo tide-conjugated antibody and not to a target nucleic acid that is, e.g., a RNA present in a cell. In some embodiments, a method comprises introducing a nucleic acid probe, e.g., a padlock probe that binds to a target nucleic acid that is, e.g., a RNA present in a cell and an oligonucleotide of an oligonucleotide-conjugated antibody. In some embodiments, a method comprises using a target nucleic acid, e.g., a RNA that encodes a protein to which the oligonucleotide-conjugated antibody binds. In some embodiments, a method comprises using a nucleic acid probe, e.g., a padlock probe that binds both an RNA that encodes a protein and an oligonucleotide of an oligonucleotide-conjugated antibody that binds that same protein to further amplify a detection signal, e.g., for a protein expressed at low levels in a cell or tissue.

[0103] In some embodiments, a method comprises determining, e.g., assessing or measuring a spatial position of a nucleic acid probe, by using a padlock probe that comprises a signaling entity (e.g., as discussed below), and / or by using a secondary nucleic acid probe able to bind to the nucleic acid probe (i.e., to a primary nucleic acid probe, e.g., a padlock probe). The determination of such nucleic acid probe is discussed in detail below.

[0104] In some embodiments, a method comprises applying more than one type of (primary) nucleic acid probe, e.g., more than one type of padlock probe to a sample, e.g., simultaneously. In some embodiments, a method comprises applying more than one type of (primary) nucleic acid probe, e.g., padlock probe to a sample, e.g., sequentially. For example, a method comprises applying at least 2, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, or at least 30,000 distinguishable nucleic acid probes, e.g., padlock probes to a sample, e.g., simultaneously or sequentially.

[0105] In some embodiments, provided is a method of preparing a nucleic acid probe, the method comprising positioning a target sequence of a nucleic acid probe, e.g., a padlock probe anywhere within the nucleic acid probe (or primary nucleic acid probe, e.g., the padlock probe). A “target sequence of a nucleic acid probe,” as used herein, refers to the nucleotides of a nucleic acid probe that bind or hybridize to the target nucleic acid sequence. In some embodiments, a target sequence of a nucleic acid probe contains a region that is substantially complementary to a portion of a target nucleic acid. In some embodiments, the portion is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to a portion of a target nucleic acid. In some embodiments, the target sequence is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. In some embodiments, the target sequence is no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, be no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides in length. In some embodiments, a target sequence comprises combinations of any of these characteristics, e.g., the target sequence has a length of between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, between 10 and 300 nucleotides, etc. Typically, complementarity is determined on the basis of Watson-Crick nucleotide base pairing.

[0106] In some embodiments, a method comprises determining a target sequence of a (primary) nucleic acid probe with reference to a target nucleic acid suspected of being present within a cell or other sample. For example, a method comprises determining a target nucleic acid to a protein using the protein’s sequence, by determining the nucleic acids that are expressed to form the protein. In some embodiments, a method comprises using only a portion of the nucleic acids encoding the protein, e.g., having the lengths as discussed above. In addition, in some embodiments, a method comprises using more than one target sequence to identify a particular target nucleic acid. For instance, a method comprises using multiple nucleic acid probes, sequentially and / or simultaneously, that can bind to or hybridize to different regions of the same target nucleic acid. Hybridization typically refers to an annealing process by which complementary single- stranded nucleic acids associate through Watson-Crick nucleotide base pairing (e.g., hydrogen bonding, guanine-cytosine and adenine-thymine) to form double- stranded nucleic acid.

[0107] In some embodiments, a nucleic acid probe, such as a primary nucleic acid probe, also comprises one or more “read” sequences. However, it should be understood that read sequences are not necessary in all cases. In some embodiments, the nucleic acid probe as described herein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more, 20 or more, 32 or more, 40 or more, 50 or more, 64 or more, 75 or more, 100 or more, 128 or more read sequences. In some embodiments, a method comprises positioning the read sequences anywhere within a nucleic acid probe. If more than one read sequence is present, a method comprises positioning the read sequences next to each other, and / or interspersing the read sequences with other sequences.

[0108] The read sequences, if present, may be of any length. In some embodiments, if more than one read sequence is used, the read sequences may independently have the same or different lengths. For instance, the read sequence may be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. In some cases, the read sequence may be no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, be no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides in length. In some embodiments, a method comprises using combinations of any of these characteristics, e.g., the read sequence have a length of between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, or between 10 and 300 nucleotides, etc.

[0109] In some embodiments, the read sequences are arbitrary or random. In certain cases, provided are methods comprising selecting the read sequences so as to reduce or minimize homology with other components of the cell or other sample, e.g., such that the read sequences do not themselves bind to or hybridize with other nucleic acids suspected of being within the cell or other sample. In some cases, the homology may be less than 10%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1%. In some cases, there may be a homology of less than 20 base pairs (bp), less than 18 bp, less than 15 bp, less than 14 bp, less than 13 bp, less than 12 bp, less than 11 bp, or less than 10 bp. In some cases, the base pairs are sequential.

[0110] In some embodiments, a population of nucleic acid probes may contain a certain number of read sequences. In some embodiments, a method comprises using a certain number of read sequences where the number of read sequences is less than the number of targets of the nucleic acid probes in a sample. Those of ordinary skill in the art will be aware that if there is one signaling entity and n read sequences, then in general In- 1 different nucleic acid targets may be uniquely identified. However, not all possible combinations need be used. For instance, in some embodiments, a population of nucleic acid probes targets 12 different nucleic acid sequences yet contains no more than 8 read sequences. As another example, a population of nucleic acids targets 140 different nucleic acid species yet contains no more than 16 read sequences.

[0111] In some embodiments, a method comprises separately identifying different nucleic acid sequence targets by using different combinations of read sequences within each nucleic acid probe. For instance, each nucleic acid probe may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, etc. or more read sequences. In some cases, a population of nucleic acid probes may each contain the same number of read sequences, although in other cases, there may be different numbers of read sequences present on the various probes. More than 20 are also possible in some embodiments. In addition, in some embodiments, a population of nucleic acid probes has, in total, 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 20 or more, 24 or more, 32 or more, 40 or more, 50 or more, 60 or more, 64 or more, 100 or more, 128 or more, etc. of possible read sequences present, although some or all of the nucleic acid probes may each contain more than one read sequence, as discussed herein. In addition, in some embodiments, the population of nucleic acid probes has no more than 100, no more than 80, no more than 64, no more than 60, no more than 50, no more than 40, no more than 32, no more than 24, no more than 20, no more than 16, no more than 15, no more than 14, no more than 13, no more than 12, no more than 11, no more than 10, no more than 9, no more than 8, no more than 7, no more than 6, no more than 5, no more than 4, no more than 3, or no more than two read sequences present. In some embodiments, a method comprises using combinations of any of these are also possible, e.g., a population of nucleic acid probes may comprise between 10 and 15 read sequences in total.

[0112] As a non-limiting example, a first nucleic acid probe may contain a first target sequence, a first read sequence, and a second read sequence, while a second, different nucleic acid probe may contain a second target sequence, the same first read sequence, but a third read sequence instead of the second read sequence. Such probes may thereby be distinguished by determining the various read sequences present or associated with a given probe or spatial location in a sample, as described herein. In some embodiments, the nucleic acid probe contains a signaling entity. It should be understood that signaling entities are not required in all cases. For example, in some embodiments, a method comprises determining, e.g., assessing or measuring the nucleic acid probe using secondary nucleic acid probes. Examples of signaling entities that can be used are discussed in more detail below.

[0113] Other components may also be present within a nucleic acid probe as well. For example, in some embodiments, one or more primer sequences are present, e.g., to allow for enzymatic amplification of probes. Those of ordinary skill in the art will be aware of primer sequences suitable for applications such as amplification (e.g., using PCR or other suitable techniques). Many such primer sequences are available commercially. Other examples of sequences that may be present within a primary nucleic acid probe include, but are not limited to promoter sequences, operons, identification sequences, nonsense sequences, or the like.

[0114] In methods provided herein, after introduction of the nucleic acid probes into a cell or other sample, the method comprises directly determining the nucleic acid probes by determining, e.g., assessing or measuring signaling entities (if present), and / or the nucleic acid probes by using one or more secondary nucleic acid probes, in accordance with certain embodiments described herein. In some embodiments, the determination is spatial, e.g., in two or three dimensions. In some embodiments, the determination is quantitative, e.g., the method comprises determining the amount or concentration of a primary nucleic acid probe (and of a target nucleic acid). Additionally, the secondary probes may comprise any of a variety of entities able to hybridize a nucleic acid, e.g., DNA, RNA, LNA, and / or PNA, etc., depending on the application.

[0115] In some embodiments, a secondary nucleic acid probe contains a recognition sequence able to bind to or hybridize with a read sequence of a primary nucleic acid probe. In some embodiments, the binding is specific, or the binding is such that a recognition sequence preferentially binds to or hybridizes with only one of the read sequences that are present. In some embodiments, a secondary nucleic acid probe contains one or more signaling entities. If more than one secondary nucleic acid probe is used, the signaling entities may be the same or different.

[0116] In some embodiments, the recognition sequences are of any length, and multiple recognition sequences are of the same or different lengths. If more than one recognition sequence is used, the recognition sequences independently have the same or different lengths. For instance, the recognition sequence is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, or at least 50 nucleotides in length. In some cases, the recognition sequence may be no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides in length.

[0117] Combinations of any of these are also possible, e.g., the recognition sequence has a length of between 10 and 30, between 20 and 40, or between 25 and 35 nucleotides, etc. In one embodiment, the recognition sequence is of the same length as the read sequence. In addition, in some cases, the recognition sequence is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% complementary to a read sequence of the primary nucleic acid probe.

[0118] In some embodiments, a nucleic acid probe is a member of more than one group or pool. Members of a nucleic acid pool contain features in addition to target sequences, read sequences, and or signaling entities that allow them to be distinguished from other groups. In some embodiments, these features are short nucleic acid sequences that are used for the amplification, production, or separation of these sequences. The nucleic acid probes of each group are applied to a sample, e.g., sequentially, as described herein.

[0119] Signaling Entities

[0120] In some embodiments, a method comprises determining signaling entities within a sample, e.g., spatially, using a variety of techniques. In some embodiments, the signaling entities is fluorescent, and techniques for determining fluorescence within a sample, such as fluorescence microscopy or confocal microscopy, is used to spatially identify the positions of signaling entities within, e.g., a cell in the sample. In some embodiments, the positions of entities within the sample are determined in two or even three dimensions.

[0121] In some embodiments, the spatial positions of the signaling entities (and thus, nucleic acid probes that the entities may be associated with) are determined at relatively high resolutions. The term “relatively high resolution,” as used herein, refers to a spatial resolution of better than about 100 micrometers. For instance, in some embodiments, the positions are determined at spatial resolutions of better than about 100 micrometers, better than about 30 micrometers, better than about 10 micrometers, better than about 3 micrometers, better than about 1 micrometer, better than about 800 nm, better than about 600 nm, better than about 500 nm, better than about 400 nm, better than about 300 nm, better than about 200 nm, better than about 100 nm, better than about 90 nm, better than about 80 nm, better than about 70 nm, better than about 60 nm, better than about 50 nm, better than about 40 nm, better than about 30 nm, better than about 20 nm, or better than about 10 nm, etc.

[0122] There are a variety of techniques to determine or image the spatial positions of signaling entities optically, e.g., using fluorescence microscopy. In some embodiments, the spatial positions is determined at super resolutions, or at resolutions better than the wavelength of light. Non-limiting examples include STORM (stochastic optical reconstruction microscopy), STED (stimulated emission depletion microscopy), NSOM (Near-field Scanning Optical Microscopy), 4Pi microscopy, SIM (Structured Illumination Microscopy), SMI (Spatially Modulated Illumination) microscopy, RESOLFT (Reversible Saturable Optically Linear Fluorescence Transition Microscopy), GSD (Ground State Depletion Microscopy), SSIM (Saturated Structured-Illumination Microscopy), SPDM (Spectral Precision Distance Microscopy), Photo-Activated Localization Microscopy (PALM), Fluorescence Photoactivation Localization Microscopy (FPALM), LIMON (3D Light Microscopical Nanosizing Microscopy), Super-resolution optical fluctuation imaging (SOFI), or the like. See, e.g., U.S. Pat. No. 7,838,302, issued November 23, 2010, entitled “Sub-Diffraction Limit Image Resolution and Other Imaging Techniques,” by Zhuang, et al.; U.S. Pat. No. 8,564,792, issued October 22, 2013, entitled “Sub-diffraction Limit Image Resolution in Three Dimensions,” by Zhuang, et al.; or Int. Pat. Apl. Pub. No. WO 2013 / 090360, published June 20, 2013, entitled “High Resolution Dual-Objective Microscopy,” by Zhuang, et al., each incorporated herein by reference in their entireties.

[0123] In addition, the signaling entity is inactivated in some cases. For example, in some embodiments, a first secondary nucleic acid probe containing a signaling entity is applied to a sample that can recognize a first read sequence, then the first secondary nucleic acid probe is inactivated before a second secondary nucleic acid probe is applied to the sample. In some embodiments, if multiple signaling entities are used, the same or different techniques may be used to inactivate the signaling entities, and some or all of the multiple signaling entities may be inactivated, e.g., sequentially or simultaneously.

[0124] In some embodiments, a method comprises inactivating a signaling entity, e.g., by removing the signaling entity (e.g., from the sample, or from the nucleic acid probe, etc.), and / or by chemically altering the signaling entity in some fashion, e.g., by photobleaching the signaling entity, bleaching or chemically altering the structure of the signaling entity, etc. For instance, in some embodiments, a method comprises inactivating a fluorescent signaling entity by chemical or optical techniques such as oxidation, photobleaching, chemically bleaching, stringent washing or enzymatic digestion or reaction by exposure to an enzyme, dissociating the signaling entity from other components (e.g., a probe), chemical reaction of the signaling entity (e.g., to a reactant able to alter the structure of the signaling entity) or the like.

[0125] In some embodiments, various nucleic acid probes (including primary and / or secondary nucleic acid probes) include one or more signaling entities. In some embodiments, if more than one nucleic acid probe is used, the signaling entities may are the same or different. In some embodiments, a signaling entity is any entity able to emit light. For instance, in some embodiments, the signaling entity is fluorescent. In some embodiments, the signaling entity is phosphorescent, radioactive, absorptive, etc. In some embodiments, the signaling entity is any entity that can be determined within a sample at relatively high resolutions, e.g., at resolutions better than the wavelength of visible light. In some embodiments, the signaling entity is, for example, a dye, a small molecule, a peptide or protein, or the like. In some embodiments , the signaling entity is a single molecule. In some embodiments, if multiple secondary nucleic acid probes are used, the nucleic acid probes comprise the same or different signaling entities.

[0126] Non-limiting examples of signaling entities include fluorescent entities (fluorophores) or phosphorescent entities, for example, cyanine dyes (e.g., Cy2, Cy3, Cy3B, Cy5, Cy5.5, Cy7, etc.), Alexa Fluor dyes, Atto dyes, photoswitchable dyes, photoactivatable dyes, fluorescent dyes, metal nanoparticles, semiconductor nanoparticles or “quantum dots”, fluorescent proteins such as GFP (Green Fluorescent Protein), or photoactivatable fluorescent proteins, such as PAGFP, PSCFP, PSCFP2, Dendra, Dendra2, EosFP, tdEos, mEos2, mEos3, PAmCherry, PAtagRFP, mMaple, mMaple2, and mMaple3. Other suitable signaling entities are known to those of ordinary skill in the art. See, e.g., U.S. Pat. No. 7,838,302 or U.S. Pat. Apl. Ser. No. 61 / 979,436, each incorporated herein by reference in its entirety.

[0127] As used herein, the term “light” generally refers to electromagnetic radiation, having any suitable wavelength (or equivalently, frequency). For instance, in some embodiments, the light may include wavelengths in the optical or visual range (for example, having a wavelength of between about 400 nm and about 700 nm, i.e., “visible light”), infrared wavelengths (for example, having a wavelength of between about 300 micrometers and 700 nm), ultraviolet wavelengths (for example, having a wavelength of between about 400 nm and about 10 nm), or the like. In some embodiments, more than one signaling entity may be used, i.e., signaling entities that are chemically different or distinct, for example, structurally. However, in some embodiments, the signaling entities may be chemically identical or at least substantially chemically identical. In some embodiments, the signaling entity is “switchable,” i.e., the signaling entity can be switched between two or more states, at least one of which emits light having a desired wavelength. In the other state(s), the signaling entity may emit no light, or emit light at a different wavelength. For instance, a signaling entity may be “activated” to a first state able to produce light having a desired wavelength, and “deactivated” to a second state not able to emit light of the same wavelength. A signaling entity is “photoactivatable” if it can be activated by incident light of a suitable wavelength. As a non-limiting example, Cy5, can be switched between a fluorescent and a dark state in a controlled and reversible manner by light of different wavelengths, i.e., 633 nm (or 642nm, 647nm, 656 nm) red light can switch or deactivate Cy5 to a stable dark state, while 405 nm green light can switch or activate the Cy5 back to the fluorescent state. In some embodiments, the signaling entity is reversibly switched between the two or more states, e.g., upon exposure to the proper stimuli. For example, a first stimuli (e.g., a first wavelength of light) is used to activate the switchable signaling entity, while a second stimuli (e.g., a second wavelength of light) is used to deactivate the switchable signaling entity, for instance, to a non-emitting state. Any suitable method may be used to activate the signaling entity. For example, in some embodiments, incident light of a suitable wavelength is used to activate the signaling entity to emit light, i.e., the signaling entity is “photoswitchable.” Thus, the photoswitchable signaling entity can be switched between different light-emitting or non-emitting states by incident light, e.g., of different wavelengths. The light may be monochromatic (e.g., produced using a laser) or polychromatic. In some embodiments, a method comprises activating the signaling entity upon stimulation by electric field and / or magnetic field. In some embodiments, a method comprises activating the signaling entity upon exposure to a suitable chemical environment, e.g., by adjusting the pH, or inducing a reversible chemical reaction involving the entity, etc. Similarly, a method comprises any suitable method to deactivate the signaling entity, and the methods of activating and deactivating the signaling entity need not be the same. For instance, in some embodiments, a method comprises deactivating the signaling entity upon exposure to incident light of a suitable wavelength or deactivating the signaling entity by waiting a sufficient time.

[0128] Typically, a “switchable” signaling entity can be identified by one of ordinary skill in the art by determining conditions under which a signaling entity in a first state can emit light when exposed to an excitation wavelength, switching the signaling entity from the first state to the second state, e.g., upon exposure to light of a switching wavelength, then showing that the signaling entity, while in the second state can no longer emit light (or emits light at a much reduced intensity) when exposed to the excitation wavelength.

[0129] In some embodiments, a method comprises using the light to activate a switchable signaling entity from an external source, e.g., a light source such as a laser light source, another light-emitting entity proximate the switchable entity, etc. In some embodiments, a method comprises using a second, light emitting signaling entity, e.g., a fluorescent signaling entity. In certain embodiments, the second, light-emitting signaling entity is itself also a switchable signaling entity.

[0130] In some embodiments, the switchable signaling entity includes a first, light-emitting portion (e.g., a fluorophore), and a second portion that activates or “switches” the first portion. For example, upon exposure to light, the second portion of the switchable signaling entity activates the first portion, causing the first portion to emit light. Examples of activator portions include, but are not limited to, Alexa Fluor 405 (Invitrogen), Alexa Fluor 488 (Invitrogen), Cy2 (GE Healthcare), Cy3 (GE Healthcare), Cy3B (GE Healthcare), Cy3.5 (GE Healthcare), or other suitable dyes. Examples of light-emitting portions include, but are not limited to, Cy5, Cy5.5 (GE Healthcare), Cy7 (GE Healthcare), Alexa Fluor 647 (Invitrogen), Alexa Fluor 680 (Invitrogen), Alexa Fluor 700 (Invitrogen), Alexa Fluor 750 (Invitrogen), Alexa Fluor 790 (Invitrogen), DiD, DiR, YOYO-3 (Invitrogen), YO-PRO-3 (Invitrogen), TOT-3 (Invitrogen), TO-PRO-3 (Invitrogen) or other suitable dyes. These may linked together, e.g., covalently, for example, directly, or through a linker, e.g., forming compounds such as, but not limited to, Cy5-Alexa Fluor 405, Cy5-Alexa Fluor 488, Cy5-Cy2, Cy5-Cy3, Cy5-Cy3.5, Cy5.5-Alexa Fluor 405, Cy5.5-Alexa Fluor 488, Cy5.5-Cy2, Cy5.5-Cy3, Cy5.5- Cy3.5, Cy7-Alexa Fluor 405, Cy7-Alexa Fluor 488, Cy7-Cy2, Cy7-Cy3, Cy7-Cy3.5, Alexa Fluor 647- Alexa Fluor 405, Alexa Fluor 647-Alexa Fluor 488, Alexa Fluor 647-Cy2, Alexa Fluor 647-Cy3, Alexa Fluor 647-Cy3.5, Alexa Fluor 750- Alexa Fluor 405, Alexa Fluor 750- Alexa Fluor 488, Alexa Fluor 750-Cy2, Alexa Fluor 750-Cy3, or Alexa Fluor 750-Cy3.5. Those of ordinary skill in the art will be aware of the structures of these and other compounds, many of which are available commercially. In some embodiments, the portions are linked via a covalent bond, or by a linker, such as those described in detail below. Other light-emitting or activator portions include portions having two quaternized nitrogen atoms joined by a polymethine chain, where each nitrogen is independently part of a heteroaromatic moiety, such as pyrrole, imidazole, thiazole, pyridine, quinoine, indole, benzothiazole, etc., or part of a nonaromatic amine. In some cases, there may be 5, 6, 7, 8, 9, or more carbon atoms between the two nitrogen atoms. In some embodiments, the light-emitting portion and the activator portions, when isolated from each other, each are fluorophores, i.e., entities that can emit light of a certain, emission wavelength when exposed to a stimulus, for example, an excitation wavelength. However, when a switchable signaling entity is formed that comprises the first fluorophore and the second fluorophore, the first fluorophore forms a first, light-emitting portion and the second fluorophore forms an activator portion that activates or “switches” the first portion in response to a stimulus. For example, in some embodiments, the switchable signaling entity comprises a first fluorophore directly bonded to the second fluorophore, or the first and second portion are connected via a linker or a common portion. Whether a pair of lightemitting portion and activator portion produces a suitable switchable signaling entity can be tested by methods known to those of ordinary skills in the art. For example, in some embodiments, light of various wavelength is used to stimulate the pair and emission light from the light-emitting portion is measured to determine whether the pair makes a suitable switch.

[0131] As a non-limiting example, Cy3 and Cy5 are linked together to form such a signaling entity. In this example, Cy3 is an activator portion that is able to activate Cy5, the lightemission portion. Thus, light at or near the absorption maximum (e.g., near 532 nm light for Cy3) of the activation or second portion of the signaling entity causes that portion to activate the first, light-emitting portion, thereby causing the first portion to emit light (e.g., near 647 nm for Cy5). See, e.g., U.S. Pat. No. 7,838,302, incorporated herein by reference in its entirety. In some cases, the first, light-emitting portion is subsequently deactivated by any suitable technique (e.g., by directing 647 nm red light to the Cy5 portion of the molecule).

[0132] Other non-limiting examples of potentially suitable activator portions include 1,5 IAEDANS, 1,8-ANS, 4-Methylumbelliferone, 5-carboxy-2,7-dichlorofluorescein, 5- Carboxyfluorescein (5-FAM), 5-Carboxynapthofluorescein, 5-Carboxytetramethylrhodamine (5-TAMRA), 5-FAM (5-Carboxyfluorescein), 5-HAT (Hydroxy Tryptamine), 5-Hydroxy Tryptamine (HAT), 5-ROX (carboxy -X -rhodamine), 5-TAMRA (5- Carboxytetramethylrhodamine), 6-Carboxyrhodamine 6G, 6-CR 6G, 6-JOE, 7-Amino-4- methylcoumarin, 7- Aminoactinomycin D (7- A AD), 7-Hydroxy-4-methylcoumarin, 9-Amino- 6-chloro-2-methoxyacridine, ABQ, Acid Fuchsin, ACMA (9-Amino-6-chloro-2- methoxy acridine), Acridine Orange, Acridine Red, Acridine Yellow, Acriflavin, Acriflavin Feulgen SITSA, Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 430, Alexa Fluor 488, Alexa Fluor 500, Alexa Fluor 514, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 610, Alexa Fluor 633, Alexa Fluor 635, Alizarin Complexon, Alizarin Red, AMC, AMCA-S, AMCA (Aminomethylcoumarin), AMCA-X, Aminoactinomycin D, Aminocoumarin, Aminomethylcoumarin (AMCA), Anilin Blue, Anthrocyl stearate, APTRA-BTC, APTS, Astrazon Brilliant Red 4G, Astrazon Orange R, Astrazon Red 6B, Astrazon Yellow 7 GLL, Atabrine, ATTO 390, ATTO 425, ATTO 465, ATTO 488, ATTO 495, ATTO 520, ATTO 532, ATTO 550, ATTO 565, ATTO 590, ATTO 594, ATTO 610, ATTO 61 IX, ATTO 620, ATTO 633, ATTO 635, ATTO 647, ATTO 647N, ATTO 655, ATTO 680, ATTO 700, ATTO 725, ATTO 740, ATTO-TAG CBQCA, ATTOTAG FQ, Auramine, Aurophosphine G, Aurophosphine, BAO 9 (Bisaminophenyloxadiazole), BCECF (high pH), BCECF (low pH), Berberine Sulphate, Bimane, Bisbenzamide, Bisbenzimide (Hoechst), bis-BTC, Blancophor FFG, Blancophor SV, BOBO -1, BOBO -3, Bodipy 492 / 515, Bodipy 493 / 503, Bodipy 500 / 510, Bodipy 505 / 515, Bodipy 530 / 550, Bodipy 542 / 563, Bodipy 558 / 568, Bodipy 564 / 570, Bodipy 576 / 589, Bodipy 581 / 591, Bodipy 630 / 650-X, Bodipy 650 / 665-X, Bodipy 665 / 676, Bodipy Fl, Bodipy FL ATP, Bodipy Fl-Ceramide, Bodipy R6G, Bodipy TMR, Bodipy TMR-X conjugate, Bodipy TMR-X, SE, Bodipy TR, Bodipy TR ATP, Bodipy TR-X SE, BO-PRO -1, BO-PRO -3, Brilliant Sulphoflavin FF, BTC, BTC-5N, Calcein, Calcein Blue, Calcium Crimson, Calcium Green, Calcium Green- 1 Ca2+ Dye, Calcium Green-2 Ca2+, Calcium Green-5N Ca2+, Calcium Green-C18 Ca2+, Calcium Orange, Calcofluor White, Carboxy -X-rhodamine (5-ROX), Cascade Blue, Cascade Yellow, Catecholamine, CCF2 (GeneBlazer), CFDA, Chromomycin A, Chromomycin A, CL-NERF, CMFDA, Coumarin Phalloidin, CPM Methylcoumarin, CTC, CTC Formazan, Cy2, Cy3.1 8, Cy3.5, Cy3, Cy5.1 8, cyclic AMP Fluorosensor (FiCRhR), Dabcyl, Dansyl, Dansyl Amine, Dansyl Cadaverine, Dansyl Chloride, Dansyl DHPE, Dansyl fluoride, DAPI, Dapoxyl, Dapoxyl 2, Dapoxyl 3' DCFDA, DCFH (Dichlorodihydrofluorescein Diacetate), DDAO, DHR (Dihydorhodamine 123), Di-4- ANEPPS, Di-8-ANEPPS (non-ratio), DiA (4-Di-16-ASP), Dichlorodihydrofluorescein Diacetate (DCFH), DiD - Lipophilic Tracer, DiD (DilCl 8(5)), DIDS, Dihydorhodamine 123 (DHR), Dil (DiIC18(3)), Dinitrophenol, DiO (DiOC18(3)), DiR, DiR (DiIC18(7)), DM- NERF (high pH), DNP, Dopamine, DTAF, DY-630-NHS, DY-635-NHS, DyLight 405, DyLight 488, DyLight 549, DyLight 633, DyLight 649, DyLight 680, DyLight 800, ELF 97, Eosin, Erythrosin, Erythrosin ITC, Ethidium Bromide, Ethidium homodimer -1 (EthD-1), Euchrysin, EukoLight, Europium (III) chloride, Fast Blue, FDA, Feulgen (Pararosaniline), FIF (Formaldehyd Induced Fluorescence), FITC, Flazo Orange, Fluo-3, Fluo-4, Fluorescein (FITC), Fluorescein Diacetate, Fluoro-Emerald, Fluoro-Gold (Hydroxy stilb amidine), FluorRuby, FluorX, FM 1-43, FM 4-46, Fura Red (high pH), Fura Red / Fluo-3, Fura-2, Fura- 2 / BCECF, Genacryl Brilliant Red B, Genacryl Brilliant Yellow 10GF, Genacryl Pink 3G, Genacryl Yellow 5GF, GeneBlazer (CCF2), Gloxalic Acid, Granular blue, Haematoporphyrin, Hoechst 33258, Hoechst 33342, Hoechst 34580, HPTS, Hydroxycoumarin, Hydroxy stilbamidine (FluoroGold), Hydroxytryptamine, Indo-1, high calcium, Indo-1, low calcium, Indodicarbocyanine (DiD), Indotricarbocyanine (DiR), Intrawhite Cf, JC-1, JO-JO-1, JO-PRO- 1, LaserPro, Laurodan, LDS 751 (DNA), LDS 751 (RNA), Leucophor PAF, Leucophor SF, Leucophor WS, Lissamine Rhodamine, Lissamine Rhodamine B, Calcein / Ethidium homodimer, LOLO-1, LO-PRO-1, Lucifer Yellow, Lyso Tracker Blue, Lyso Tracker Blue-White, Lyso Tracker Green, Lyso Tracker Red, Lyso Tracker Yellow, Ly soSensor Blue, Ly soSensor Green, Ly soSensor Yellow / Blue, Mag Green, Magdala Red (Phloxin B), Mag-Lura Red, Mag-Lura-2, Mag-Lura-5, Mag-Indo-1, Magnesium Green, Magnesium Orange, Malachite Green, Marina Blue, Maxiion Brilliant Llavin 10 GPP, Maxiion Brilliant Plavin 8 GPP, Merocyanin, Methoxycoumarin, Mitotracker Green PM, Mitotracker Orange, Mitotracker Red, Mitramycin, Monobromobimane, Monobromobimane (mBBr-GSH), Monochlorobimane, MPS (Methyl Green Pyronine Stilbene), NBD, NBD Amine, Nile Red, Nitrobenzoxadidole, Noradrenaline, Nuclear Past Red, Nuclear Yellow, Nylosan Brilliant lavin E8G, Oregon Green, Oregon Green 488-X, Oregon Green, Oregon Green 488, Oregon Green 500, Oregon Green 514, Pacific Blue, Pararosaniline (Peulgen), PBPI, Phloxin B (Magdala Red), Phorwite AR, Phorwite BKL, Phorwite Rev, Phorwite RPA, Phosphine 3R, PKH26 (Sigma), PKH67, PMIA, Pontochrome Blue Black, POPO-1, POPO-3, PO-PRO-1, PO-PRO-3, Primuline, Procion Yellow, Propidium lodid (PI), PyMPO, Pyrene, Pyronine, Pyronine B, Pyrozal Brilliant Plavin 7GP, QSY 7, Quinacrine Mustard, Resorufin, RH 414, Rhod-2, Rhodamine, Rhodamine 110, Rhodamine 123, Rhodamine 5 GLD, Rhodamine 6G, Rhodamine B, Rhodamine B 200, Rhodamine B extra, Rhodamine BB, Rhodamine BG, Rhodamine Green, Rhodamine Phallicidine, Rhodamine Phalloidine, Rhodamine Red, Rhodamine WT, Rose Bengal, S65A, S65C, S65L, S65T, SBPI, Serotonin, Sevron Brilliant Red 2B, Sevron Brilliant Red 4G, Sevron Brilliant Red B, Sevron Orange, Sevron Yellow L, SITS, SITS (Primuline), SITS (Stilbene Isothiosulphonic Acid), SNAPL calcein, SNAPL-1, SNAPL-2, SNARL calcein, SNARL 1, Sodium Green, SpectrumAqua, SpectrumGreen, SpectrumOrange, Spectrum Red, SPQ (6-methoxy-N-(3-sulfopropyl)quinolinium), Stilbene, Sulphorhodamine B can C, Sulphorhodamine Extra, SYTO 11, SYTO 12, SYTO 13, SYTO 14, SYTO 15, SYTO 16, SYTO 17, SYTO 18, SYTO 20, SYTO 21, SYTO 22, SYTO 23, SYTO 24, SYTO 25, SYTO 40, SYTO 41, SYTO 42, SYTO 43, SYTO 44, SYTO 45, SYTO 59, SYTO 60, SYTO 61, SYTO 62, SYTO 63, SYTO 64, SYTO 80, SYTO 81, SYTO 82, SYTO 83, SYTO 84, SYTO 85, SYTOX Blue, SYTOX Green, SYTOX Orange, Tetracycline, Tetramethylrhodamine (TAMRA), Texas Red, Texas Red-X conjugate, Thiadicarbocyanine (DiSC3), Thiazine Red R, Thiazole Orange, Thioflavin 5, Thioflavin S, Thioflavin TCN, Thiolyte, Thiozole Orange, Tinopol CBS (Calcofluor White), TMR, TO-PRO-1, TO-PRO-3, TO-PRO-5, TOTO-1, TOTO-3, TRITC (tetramethylrodamine isothiocyanate), True Blue, TruRed, Ultralite, Uranine B, Uvitex SFC, WW 781, X-Rhodamine, XRITC, Xylene Orange, Y66F, Y66H, Y66W, YO-PRO-1, YO-PRO-3, YOYO-1, YOYO-3, SYBR Green, Thiazole orange (interchelating dyes), or combinations thereof.

[0133] Codes

[0134] In some embodiments, provided is a method comprising converting a pattern of binding or hybridization of secondary nucleic acid probes to primary nucleic acid probes into a “codeword.” For example, the codewords may be “101” and “110,” where a value of 1 represents binding and a value of 0 represents no binding. The codewords may also have longer lengths in other embodiments. In some embodiments, a codeword is directly related to a specific target nucleic acid sequence of the primary nucleic acid probe. Accordingly, methods are provided comprising matching different primary nucleic acid probes to certain codewords, and using the codewords to identify the different targets of the primary nucleic acid probes based on the binding patterns of the secondary probes, even if in some cases, there is overlap in the read sequences of different secondary probes. However, if no binding occurs, the codeword would be “000” in a sample.

[0135] In some embodiments, a method comprises dividing a plurality of different primary probes containing codewords into as many separate pools as there are positions in the codewords, such that each primary probe pool corresponds to a certain value in a certain position of the codewords (e.g., a “one” in the first position as in “1001”).

[0136] In some embodiments, a method comprises assigning the codewords using an errordetection system or an error-correcting system, such as a Hamming system, a Golay code, or an extended Hamming system (or a SECDED system, i.e., single error correction, double error detection). Generally, such systems can be used to identify where errors have occurred, and in some cases, such systems can also be used to correct the errors and determine what the correct codeword should have been. For example, a method comprises detecting a codeword such as 001 as invalid and correcting using such a system to 101, e.g., if 001 is not previously assigned to a different target sequence. Accordingly, in some embodiments, a method comprises determining a codeword and, once the codeword is determined, comparing the codeword to the known nucleic acid probe codewords. In some embodiments, if a match is found, the method further comprises identifying or determining the nucleic acid probe. If no match is found, then an error in the reading of the codeword is identified. In some embodiments, a method comprises applying error correction to determine the correct codeword, and thus the correct nucleic acid probe. In some embodiments, a method comprises selecting the codewords such that, assuming that there is only one error present, only one possible correct codeword is available, and thus, only one of the nucleic acid probes is possible. In some embodiments, this may also be generalized to larger codeword spacings or Hamming distances; for instance, the codewords may be selected such that if two, three, or four errors are present (or more in some cases), only one possible correct codeword is available, and thus, only one of the nucleic acid probes is possible. A variety of different error-correcting codes can be used, some of which have previously been developed for use within the computer industry.

[0137] It should also be understood that all possible codewords in a code need not be used in some cases. For example, in some embodiments, codewords that are not used serve as negative controls. Similarly, in some embodiments, some codewords are left out because they are more prone to errors in measurement than other codewords. For example, in some embodiments, reading a codeword with more values of ‘1’ might be more error-prone that reading a codeword with fewer values of ‘1.’

[0138] Spatial Position Determination

[0139] In some embodiments, a method provided herein comprises determining the spatial positions of the nucleic acid probes within a cell or other sample. In some embodiments, a method comprises determining the spatial positions of oligonucleotides conjugated to antibodies and of nucleic acid probes hybridized to cellular RNAs. In some embodiments, a method comprises determining spatial positions of antibody-conjugated oligonucleotides and nucleic acids at resolutions better than the wavelength of visible light.

[0140] In some embodiments, the nucleic acids are mRNAs, or other nucleic acids described herein. In some embodiments, a method comprises determining spatial positions of a relatively large numbers of nucleic acids and / or oligonucleotide labeled proteins by using combinatorial approaches using a relatively small number of different labels on the nucleic acid probes. Thus, for example, a method comprises using a relatively small number of experiments to determine spatial positions of a relatively large number of nucleic acids and / or oligonucleotide labeled proteins in a sample due, e.g., to simultaneous binding of the nucleic acid probes to different nucleic acids and / or oligonucleotide labeled proteins in the sample.

[0141] In some embodiments, a method comprises applying a plurality of primary nucleic acid probes to the cell (or other sample), which plurality of primary nucleic acid probes bind or hybridize to nucleic acids and / or oligonucleo tide-conjugated antibodies bound to proteins suspected of being present within the cell. Afterwards, the method comprises adding sequentially secondary nucleic acid probes that can bind to or otherwise interact with some of the primary nucleic acids and determining the spatial positions of the secondary nucleic acid probes, e.g., using imaging techniques such as STORM (stochastic optical reconstruction microscopy) or other imaging techniques. After imaging, the method comprises inactivating or removing the secondary nucleic acid probes and adding different secondary nucleic acid probes to the sample. This may be repeated multiple times with multiple different secondary nucleic acid probes. In some embodiments, a method comprises using the pattern of binding of the various secondary nucleic acid probes to determine the spatial positions of primary nucleic acid probes at locations within the cell or other sample, to determine the spatial position of mRNA or other nucleic acids and / or oligonucleotide labeled proteins (by way of oligonucleo tide-conjugated antibodies described herein) that are present in the sample (e.g., a cell of the sample). It should be understood that the above description is an example of some embodiments of the invention, and that primary and secondary nucleic acid probes are not necessary in all embodiments. For example, in some embodiments, a method comprises using a series of primary nucleic acid probes containing signaling entities to determine spatial positions of nucleic acids and / or oligonucleotide-conjugated antibodies within a cell or other sample, without necessarily requiring secondary nucleic acid probes.

[0142] For example, a first spatial location within the cell or other sample may exhibit binding of a first secondary probe and a third secondary probe, but not the binding of a second or a fourth secondary probe, while a second spatial location may exhibit a different pattern of binding of various secondary probes. In some embodiments, a method comprises determining the spatial positions of the primary nucleic acid probe that the secondary probes are able to bind to or hybridize with by considering the pattern of binding of various secondary probes.

[0143] Samples

[0144] A sample includes a cell culture, a suspension of cells, a biological tissue, a biopsy, an organ, an organism, or the like. The sample may also be cell-free but nevertheless contain nucleic acid and / or oligonucleotide-conjugated antibodies. If a sample contains a cell, the cell may be a human cell, or any other suitable cell, e.g., a mammalian cell, a fish cell, an insect cell, a plant cell, or the like. More than one cell may be present in a sample.

[0145] In some embodiments, the nucleic acids and / or oligonucleotides conjugated to antibodies, the spatial locations of which are determined by the methods described herein, are, e.g., DNA, RNA, or other nucleic acids that are present within a cell (or other sample). In some embodiments, the nucleic acids and / or proteins bound by oligonucleotide-conjugated antibodies are endogenous to the cell or added to the cell. For instance, a nucleic acid may be viral, or artificially created. In some embodiments, a nucleic acid and / or protein bound by an oligonucleotide-conjugated antibody is expressed by the cell. In some embodiments, the nucleic acid is RNA. In some embodiments, the RNA is coding and / or non-coding RNA. Non-limiting examples of RNA that may be studied within a cell include mRNA, siRNA, rRNA, miRNA, tRNA, IncRNA, snoRNAs, snRNAs, exRNAs, piRNAs, or the like.

[0146] In some cases, a method comprises determining spatial positions of a significant portion of the nucleic acids and / or proteins bound by oligonucleotide-conjugated antibodies within a cell. For example, in some embodiments, a method comprises determining enough of the RNA present within a cell so as to produce a partial or complete transcriptome of the cell including spatial positions of a partial or complete transcriptome of the cell. In some embodiments, a method comprises determining spatial positions of at least 4 types of mRNAs within a cell. In some embodiments, a method comprises determining spatial positions of at least 3, at least 4, at least 7, at least 8, at least 12, at least 14, at least 15, at least 16, at least 22, at least 30, at least 31, at least 32, at least 50, at least 63, at least 64, at least 72, at least 75, at least 100, at least 127, at least 128, at least 140, at least 255, at least 256, at least 500, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 types of mRNAs within a cell. In some embodiments, a method comprises determining enough of the proteins present within a cell so as to produce a partial or complete proteome of the cell including spatial positions of a partial or complete proteome of the cell

[0147] It should be understood that the transcriptome generally encompasses all RNA molecules produced within a cell, not just mRNA. Thus, for instance, the transcriptome also includes rRNA, tRNA, etc. In some embodiments, a method comprises determining at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the transcriptome of a cell may be determined. In some embodiments, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the proteome of a cell.

[0148] In some aspects, provided are methods for enzymatic synthesis of probes for RCA- MERFISH and other RNA-templated probe ligation methods. In some embodiments, certain systems and methods comprise synthesizing probes that are ligated in situ either as DNA:DNA hybrids or RNA:DNA hybrids. In some embodiments, a method comprises synthesizing RNA:DNA hybrids, using the RNA-templated PBCV-1 DNA ligase (brand name SplintR from New England Biolabs). In one embodiment, the synthesis involves (1) performing PCR from an oligonucleotide pool, (2) transcribing with RNA polymerase, (3) reverse transcribing and degrading the RNA, (4) hybridizing short oligos to the remaining ssDNA to generate dsDNA templates for restriction enzymes, and (5) using type IIS restriction enzymes to generate the final padlock probes with exact 5’ and 3’ homology arms. In some embodiments, a method comprises using techniques such as annealing and digestion in other embodiments, e.g., with use of type IIS restriction enzymes to generate padlock probes.

[0149] In certain embodiments, provided are method that comprise using RCA instead of using in vitro transcription - reverse transcription to amplify ssDNA. In some embodiments, the methods comprise: (1) performing PCR from an oligonucleotide pool with phosphorylated primers, (2) circularizing phosphorylated PCR products, (3) combining enzymatic nicking and exonuclease degradation of one strand of the circularized DNA, (4) annealing of a primer to the circularized ssDNA, (5) amplifying by rolling circle amplification under conditions with single- strand binding protein that primarily generate ssDNA, (6) digesting using either type IIS restriction enzyme digestion or DNAzyme (selfcleaving DNA enzyme) cleavage of the resulting long ssDNA product to generate a short probe for hybridization.

[0150] In certain embodiments, a method comprises generating complex libraries of probes for both in situ RCA-based approaches and for hybridization-based single cell sequencing. For example, in one embodiment, a method comprises synthesizing a probe to reduce the price of droplet-based and split-pool-based single-cell RNA-seq approaches, in addition to the in situ approaches.

[0151] In some embodiments, a method comprises performing enzymatic probe synthesis from oligonucleotide libraries synthesized in array for split probe, RNA-templated- ligation- based single cell sequencing. In some embodiments, a method comprises performing enzymatic probe synthesis from oligonucleotide libraries synthesized in array to generate probes for RCA-MERFISH.

[0152] In some embodiments, a method comprises using a probe-based single cell sequencing assay comprising a droplet-based method. In some embodiments, a method comprises using a probe-based single cell sequencing assay that comprises split-pool combinatorial indexing. In some embodiments, methods provided herein comprise providing libraries of enzymatically-synthesized split probes that target all mRNAs ("genome-wide"). In some embodiments, methods provided herein comprise providing libraries of enzymatically-synthesized split probes that target a subset of genes. In some embodiments, methods provided herein comprise providing libraries of enzymatically-synthesized split probes that target sgRNAs. In some embodiments, methods provided herein comprise providing libraries of enzymatically-synthesized split probes that target barcodes.

[0153] Padlock Probe Production

[0154] Further provided are methods comprising producing padlock probes described herein from low-cost femtomolar pools of long oligos synthesized in array. In some embodiments, the methods comprise performing limited cycle PCR. In some embodiments, the methods comprise amplifying the initial array-synthesized oligo pool by RCA. In some embodiments, the methods comprise liberating the full length phosphorylated probes. In some embodiments, the methods comprise digesting samples using Type IIS restriction enzymes.

[0155] In some embodiments, a method comprises synthesizing a probe. In some embodiments, a method comprises amplifying a pooled oligonucleotide library. In some embodiments, a method comprises synthesizing ssDNA. In some embodiments, a method comprises digesting the synthesized ssDNA to create single stranded split probes with defined 5' and / or 3' ends for RNA-templated ligation. In some embodiments, a method comprises initially amplifying by PCR an original oligonucleotide library. In some embodiments, a method comprises appending by PCR a probe barcode useful for sample multiplexing. In some embodiments, a method comprises introducing by PCR a T7 promoter to the amplified DNA. In some embodiments, a method comprises, T7 -transcribing the PCR- amplified material. In some embodiments, a method comprises reverse transcribing the T7- transcribed material to generate ssDNA. In some embodiments, a method comprises circularizing and rolling-circle amplifying PCR product to generate ssDNA. In some embodiments, a method comprises cloning the PCR product into a phage genome. In some embodiments, a method comprises producing ssDNA from the phage / phagemid. In some embodiments, a method comprises digesting the ssDNA by restriction enzymes to generate ssDNA fragments with the proper 5' and / or 3' ends. In some embodiments, a method comprises generating ssDNA hairpin sequences that create cut sites for the restriction enzymes. In some embodiments, a method comprises hybridizing the ssDNA to oligonucleotide probes to create cut sites for the restriction enzymes. In some embodiments, a method comprises digesting the ssDNA with DNAzyme to create probes with the proper 5' and / or 3' ends. In some embodiments, a method further comprises activating the DNAzymes by metal ions.

[0156] In certain embodiments, methods provided herein are generally directed to enzymatic production of amplifiers for protein and RNA imaging. Because MERFISH and multiplexed immunofluorescence with oligonucleotide-conjugated antibodies can be relatively dim, necessitating high laser power and long exposure times to achieve sufficient signal; additionally, the dim signal may mean that, in some tissues, background autofluorescence can overwhelm the fluorescence from the MERFISH or IF stains, methods provided comprise using amplifiers. The term “amplifier”, as used herein, refers to a DNA sequence with one copy of a sequence that binds to the bit and many copies of a distinct sequence to increase the signal intensity.

[0157] In certain embodiments, methods are provided that allow for producing oligonucleotides in pools with enzymatic approaches. In some embodiments, a method comprises starting with a cheaper, un-purified oligonucleotide amplifier pool, selecting full- length products using a special PCR process that avoids intramolecular recombination, and then producing single-stranded DNA amplifiers with T7 transcription and reverse transcription, with circularization and RCA, or the like.

[0158] Additionally, in some embodiments, a method comprises enzymatically producing probes by incorporating nucleotide analogues such as aminoallyl-dUTP or vinyl-dUTP. In some embodiments, a method comprises covalently incorporating amplifiers into a polyacrylamide hydrogel, which increases the storage stability of the samples, and / or allows reversible stripping out readout probes. Additionally, a method comprises modifying the probes to protect the amplifiers against exonucleolytic degradation, such as by Phi29 during rolling circle amplification.

[0159] One non-limiting example of synthesis using the PCR-T7-RT version, with modification with a gel-reaction agent, is shown in FIG. 2. In certain embodiments, provided are methods generally directed to novel selfsupervised analysis methods for multiplexed structural imaging. Interpreting imaging-based cellular phenotypes is currently a challenging problem. Where single-cell RNA transcriptomes can be represented as high-dimensional vectors to compute correlations or perform regressions across perturbations, other imaging-based phenotypes such as those measured by multiplexed immuno staining are more challenging for automated analysis and interpretation. Individual cells may vary in their orientation, shape, intracellular structure, dynamics, cell-cell interactions and communications, and many other attributes that are irrelevant to the effect of the genetic perturbation, but which would confound a naive analysis of e.g. protein intensity or localization.

[0160] In some embodiments, methods provided are generally directed to a low-dimensional representation of cellular protein state. In some embodiments, methods comprise using techniques that use self-supervised deep learning methods (e.g., variational autoencoders [VAEs] or visual transformers [ViTs], etc.) to reduce the dimensionality of imaging data for phenotypic interpretation. In some embodiments, methods comprise using an auxiliary loss function that discriminates between different perturbations or experimental conditions on a per-cell basis, in order to generate a representation of the features of each cell that is optimized to maximally discriminate different phenotypes. In some embodiments, methods comprise using a multi-level model that models each protein or other stains in a cell individually, and / or a joint model that models all of the proteins and stains together (e.g., holistically).

[0161] The following are incorporated herein by reference in their entireties: Int. Pat. Apl. Pub. Nos. WO 2016 / 018960, WO 2016 / 018963, WO 2018 / 089445, WO 2018 / 218150, WO 2018 / 089438, WO 2020 / 123742, WO 2020 / 214885, WO 2021 / 102122, and WO 2021 / 138078.

[0162] In some embodiments, the methods provided comprise providing a plurality of padlock probes comprising at least three different signaling entities, ligating the padlock probes upon binding to mRNA, amplifying the ligated padlock probes in situ, and detecting readout sequences over multiple rounds of hybridization. In some embodiments, the methods further comprise binding a plurality of oligonucleotide-conjugated antibodies to a plurality of proteins, which proteins, optionally, are subcellular organelle proteins, lipid droplet proteins, signaling pathway proteins. In some embodiments, the methods further comprise hybridizing a plurality of padlock probes to cellular RNAs, optionally RNAs that exhibit specific cellular pattems of cellular localization. In some embodiments, the methods further comprise sequentially imaging the plurality of proteins and RNAs.

[0163] In some embodiments, the methods comprise processing the image data, the processing comprising reducing each image channel in each cell from a high-resolution z- stack to a 128x128 pixel matrix. In some embodiments, the methods comprise further reducing the dimensionality of the protein and / or RNA stain images to a 512-dimensional vector. In some embodiments, the methods further comprise using low-dimensional embedding to discriminate different transcriptionally-defined cell types and / or physiological conditions.

[0164] In some embodiments, the methods comprise providing a 512-dimensional embedding for each protein and / or RNA across all cells and visualizing individual cellular features in the embeddings. In some embodiments, the methods further comprise conducting unsupervised clustering of the embeddings to identify cellular subtypes. In some embodiments, the methods comprise training a classifier to predict cellular subtypes using transcriptional profiles and / or 512-dimensional feature embeddings.

[0165] In some embodiments, the methods comprise obtaining a tissue from a subject. In some embodiments, the methods further comprise obtaining a tissue from a subject that has undergone a cellular stress. In some embodiments, the methods comprise contacting a cell or tissue with a plurality of sgRNAs and a Cas nuclease to knock-out one or more target genes. In some embodiments, the methods comprise contacting a cell or tissue with a nucleic acid comprising a sgRNAs comprising a barcode. In some embodiments, the methods comprise contacting a cell or tissue with a nucleic acids comprising a plurality of sgRNAs each comprising a barcode. In some embodiments, the methods comprise using a plurality of sgRNAs comprising a barcode and further comprise using intervening regions between each sgRNA and barcode pair to reduce barcode swapping. In some embodiments, the methods further comprise contacting a cell or tissue with a plurality of padlock probes targeting the sgRNA barcode sequences. In some embodiments, the methods further comprise sequentially imaging the plurality of padlock probes targeting the sgRNA barcode sequences. In some embodiments, the methods further comprise providing a plurality of padlock probes targeting cellular mRNAs and / or a plurality of padlock probes targeting antibody-conjugated oligonucleotides, amplifying the plurality of padlock probes in situ, and detecting readout sequences over multiple rounds of hybridization.

[0166] In some embodiments, the methods comprise dissociating a fixed tissue comprising padlock probes to cellular RNAs, antibody-conjugated oligonucleotides, and padlock probes to sgRNAs, and enriching sgRNA-containing cells using cell sorting methods, e.g., flow cytometry cell sorting. In some embodiments, the methods further comprise amplifying the plurality of padlock probes in situ and detecting readout sequences over multiple rounds of hybridization. In some embodiments, the methods comprise single-cell RNA sequencing. In some embodiments, the methods further comprise applying an autoencoder model to generate lower-dimensional embeddings for an image obtained following padlock probe hybridization and performing an unbiased clustering on the embedding. In some embodiments, the methods further comprise identifying sgRNAs that shift a representation of cells between different clusters. In some embodiments, the methods comprise using an energy distance metric to quantify an impact of a sgRNA on a distribution of global transcriptional states. The following examples are intended to illustrate certain embodiments of the present disclosure, but do not exemplify the full scope of the disclosure.

[0167] EXAMPLE 1

[0168] Methods and Materials

[0169] Fixed Cell scRNA-seq Library Preparation and Sequencing scRNA-seq was conducted with the Single Cell Gene Expression Flex platform (lOx Genomics). Spike-in probes for the sgRNAs were included to a final concentration of 2 nM. Cells were counted on a Countess II (ThermoFisher Scientific). Four sample barcodes were used and the cells were recovered across one (diet experiment) or eight (Perturb-seq experiment) of microfluidic channels. mRNA libraries were prepared according to the manufacturer’s instructions. sgRNA libraries were prepared with the Fixed RNA Feature Barcode Kit according to the manufacturer’s instructions. Libraries were sequenced using a NovaSeq 6000 (Illumina). scRNA-seq Alignment and Calling

[0170] The scRNA-seq mRNA data was aligned with CellRanger (lOx Genomics). The sgRNA reads were aligned with a custom pipeline using the cell barcodes produced by CellRanger. Briefly, bowtie2 (flags -very-sensitive —local) was used to align the reads to the sgRNA probe library. SgRNA reads were then selected with a cell barcode that was shared with a cellranger-called cell and the number of UMIs identified for each sgRNA in each cell. Several perturbation calling approaches were tested including mixed model calling and identifying outliers in a Poisson distribution or zero-inflated Poisson distribution and the best performance across sgRNAs was found with thresholding, choosing thresholds empirically to maximize on target knockdown of perturbation targets with apparent NMD. All cells with exactly one sgRNA were identified over the chosen threshold and the rest were excluded from all analyses.

[0171] Antibody Labeling

[0172] Antibodies were obtained in >50 ug quantities and labeled with bifunctional 5' Acrydite - bit sequence - 3’ DBCO oligos (IDT) by enzymatic modification and click chemistry (SiteClick Antibody Azido Modification Kit, ThermoFisher) according to the manufacturer’s instructions. Antibody-oligo conjugates were concentrated in PBS by ultrafiltration with a lOOkDa membrane (Millipore), which also removed residual unconjugated oligonucleotides. Antibody-oligonucleotide conjugates were aliquoted into tube strips and snap frozen in liquid nitrogen.

[0173] Probe Synthesis

[0174] Enzymatic probe synthesis from oligonucleotide libraries synthesized in array (such as by Twist, Agilent, etc) were used for split probe, RNA-templated-ligation-based single cell sequencing assays, and to generate probes for RCA-MERFISH. These probe-based single cell sequencing assays were used in the form of droplet-based methods (similar to lOx FLEX or a probe-based adaptation of Fluent Biosciences technology) or for split-pool combinatorial indexing. In some implementations, these libraries of enzymatically- synthesized split probes were used to target all mRNAs ("genome-wide"). In some implementations, these libraries were used to target a subset of genes. In some implementations, these libraries were used to target sgRNAs. In some implementations, these libraries were used to target barcodes.

[0175] The probe synthesis involved the amplification of the pooled oligonucleotide libraries and then synthesis of ssDNA, followed by digestion to create single stranded split probes with defined 5' and / or 3' ends for RNA-templated ligation. In some implementations, the original oligonucleotide libraries were initially amplified by PCR. In some implementations, the PCR appended a probe barcode at this step for sample multiplexing. In some implementations, the PCR introduced a T7 promoter to the amplified DNA. In some implementations, the PCR- amplified material was then be TV-transcribed and then reverse transcribed to generate ssDNA. In some implementations, the PCR product was circularized and rolling-circle amplification was used to generate ssDNA. In some implementations, the PCR product was cloned into a phage genome and phage / phagemid was used to produce ssDNA. In some implementations, the ssDNA was digested by restriction enzymes to generate ssDNA fragments with the proper 5' and / or 3' ends. In some implementations, the ssDNA had hairpin sequences that created cut sites for the restriction enzymes. In some implementations, the ssDNA hybridized to oligonucleotide probes to create cut sites for the restriction enzymes. In some implementations, the ssDNA had DNAzyme sites that enabled the cleavage of the ssDNA to create probes with the proper 5' and / or 3' ends. In some implementations, the DNAzymes were activated by metal ions.

[0176] RCA-MERFISH Readout Probe Synthesis

[0177] Amine-modified 15mer oligonucleotides were obtained from IDT (standard desalting). Oligos were resuspended to 300 pm in 112 mM sodium bicarbonate solution (ThermoFisher). 300 mM Sulfo-Cy3-NHS ester, Sulfo-Cy5-NHS ester, and Sulfo-Cy7-NHS ester (Lumiprobe) solutions were made in dry DMSO (Sigma Aldrich). The appropriate dye was added to each oligonucleotide to a final concentration of 10 mM and the dyes were allowed to react for 24 hours in the dark at room temperature. Sodium acetate pH 5.5 was added to a final concentration of 500 mM and then ice cold ethanol was added to a final concentration of 80%. The oligonucleotides were incubated at -20° C for >24 hours, pelleted by centrifugation at > 18,000g at 4° C for >20 minutes, washed 3x with ice cold 80% ethanol, dried in a vacufuge, and resuspended in TE pH 8 to a final concentration of 100 pm. Labeled oligonucleotides were stored at 4° C until use.

[0178] RCA-MERFISH Encoding Probe Design and Construct

[0179] For the 205 endogenous genes, a 21 -bit, Hamming Weight 4, Hamming Distance 4 codebook was used for MERFISH imaging. For the 456 perturbation barcodes, an 18-bit, Hamming Weight 6, Hamming Distance 4 codebook was used for MERFISH imaging. For both the endogenous genes and perturbation barcodes, individual genes / barcodes were randomly assigned to codewords in the codebook. For each gene / barcode, a total of 8 (endogenous genes) encoding probes (shifted by various nucleotide lengths) or 3 (barcodes) encoding probes (shifted by 60 nucleotides) were designed with a single codeword targeting the gene or barcode mRNA sequence. Following the guidelines for MERFISH probe design as previously described (Allen, W.E., Cell 186, 194-208.el8, 2023), 60 mer regions were selected that could be split into two 30 mers, where each half had GC content between 30 and 70%, melting temperature Tm within 60-80°C, isoform specificity index between 0.7 and 1, gene specificity index between 0.75 and 1, and no homology longer than 15 nt to rRNAs or tRNAs. Pairs of adjacent probes that had a ligation junction with a G or C at the donor (5’ phosphorylated) end of the probe were excluded. The two halves of the probes were then split, and between them was added an RCA primer sequence and the reverse complements of the 4 (endogenous) or 6 (barcode) readout sequences that encoded the identity of that gene. PCR handles with BciVI (left hand side) and Bed (right hand side) restriction sites were then appended to either end of the probe. The readout sequences on the encoding probes were detected with dye-labeled readout probes with complementary sequences in order to decode the gene or barcode.

[0180] RCA-MERFISH Encoding Probe Synthesis

[0181] RCA-MERFISH encoding probe libraries were first synthesized at femtomolar scale in a pool by Twist Biosciences. Probe libraries were then amplified by limited-cycle PCR (Phusion Polymerase, New England Biolabs), purified using SPRI beads (Beckman Coulter), and then blunt-end ligated using T4 DNA to circularize. Circularized DNA molecules were nicked using Nt.BbvCI (NEB), and then RCA amplified overnight at 30°C using Phi29 DNA polymerase from the nick site with T4 gene 32 single-stranded binding protein (SSB) (New England Biolabs). RCA amplified DNA was ethanol precipitated and resuspended in CutSmart buffer. Oligonucleotides with degenerate ends containing restriction enzyme sites were annealed to the RCA product, and the mixture was digested overnight with BccI and BciVI (New England Biolabs). The final libraries were then purified using magnetic beads, eluting in Tris-EDTA (TE) (ThermoFisher) buffer to a final concentration of -10 nM / probe. The resulting library was stored at -20°C until use.

[0182] RCA-MERFISH Sample Preparation

[0183] Sample preparation occurred over several days. Blocking, antibody staining, MelphaX modification, probe hybridization, ligation, and RCA were conducted with the coverslips inverted onto small volumes of reaction mixture over parafilm, whereas decrosslinking, washing, and digestion were conducted with the coverslips upright in >5 ml of solution in 60 mm tissue culture dishes. Silanized, PDL-coated coverslips were prepared according to the method of Moffitt, J.R., et al., Proc. Natl. Acad. Sci. U. S. A. 113, 14456-14461.

[0184] 10 pm sections of liver were cut onto silanized, PDL-coated coverslips, warmed to room temperature for 15 minutes, and attached to the surface by 15 minutes of postfixation in PBS + 4% paraformaldehyde. The sections were washed 3x with PBS and decrosslinked at 60° in TE pH 9 (Genemed) for an hour. The sections were then washed with PBS.

[0185] The sections were blocked at RT for 20 minutes in blocking buffer (lx PBS, 10 mg / ml BSA (UltraPure, ThermoFisher), 0.3% Triton X-100, 0.5 mg / ml sheared salmon sperm DNA (ThermoFisher), and 0.1 U / pL SUPERase ln RNase Inhibitor (ThermoFisher). It was ensured that the blocking buffer did not contain dextran sulfate, as even trace amounts inhibited later ligation and / or RCA.

[0186] The antibody-oligonucleotide conjugates were pooled and diluted into blocking buffer. The sections were then stained overnight at 4° with this mix. Each coverslip was stained with -150 ng of each oligonucleotide- antibody, diluted into 150 ul of blocking buffer. The sections were then washed 3x with PBS and then incubated with PBS at RT for 15 minutes, then postfixed for 5 minutes in PBS + 4% paraformaldehyde. The sections were then washed 3x with PBS and fixed with 1.5 mM BS(PEG)9 in PBS for 20 minutes, inactivated in PBS + 100 mM Tris pH 8 for 5 minutes, washed 3x in PBS, and stored in tightly sealed tissue culture dishes at 4° in 70% ethanol for 24 hours to one month.

[0187] After antibody staining, the samples were then modified with MelphaX that was diluted 1:10 in 20 mM MOPS pH7.7 at 37°C for 1 hr. The samples were washed three times with PBS, then embedded in an acrylamide gel (4% v / v 19:1 acrylamide:bis-acrylamide (BioRad), 300 mM NaCl, 60 mM Tris pH 8, 0.2% v / v tetramethylethylenediamine [TEMED], and 0.2% w / v ammonium persulfate [APS]) for 1.5 hrs at room temperature. The samples were then digested at 42°C for 48 hours in digestion buffer (2% v / v sodium dodecyl sulfate [SDS] (Thermo Fisher), 1% v / v proteinase K (New England Biolabs), 50 mM Tris pH 8 (Ambion), 300 mM NaCl (Ambion), 0.25% Triton X-100 (Sigma), 0.5 mM ethylenediaminetetraacetic acid [EDTA] (Ambion)).

[0188] After digestion, samples were washed three times with PBS + 0.1% Triton X-100 to remove residual SDS. Samples were then hybridized with a mixture containing 2xSSC, 30% v / v formamide (Ambion), 1% v / v murine RNase inhibitor (New England Biolabs), 0.1% w / v yeast tRNA, 5% w / v PEG35000 (Sigma), 1 uM of a polyA probe, a library of probes for imaging pre-rRNA, mtRNA, and Albumin at 1 nM / probe, and each RCA-MERFISH encoding probe library at 1 nM / probe. The polyA probe had a mixture of DNA and LNA nucleotides ( / 5Acryd / TTGAGTGGATGGAGTGTAATT+TT+TT+TT+TT+TT+TT+TT+TT+TT+T) (SEQ ID NO: 1) where T+ is a locked nucleic acid and / 5Acryd / is a 5’ acrydite modification. The samples were then hybridized for 36-48 hours at 37°C in a humidified chamber.

[0189] After hybridization, the samples were washed twice for 30 min in 2xSSC, 30% v / v formamide at 47°C. The samples were then washed three times with PBS + 0.1% v / v Tween 20, and once briefly with preligation buffer (50 mM Tris-HCl pH8, 10 mM MgC12). The RCA-MERFISH probes were ligated with 10% v / v (~1 uM) SplintR ligase (New England Biolabs), lx SplintR ligase buffer, 1% v / v murine RNAse inhibitor, and 100 nM RCA primer (TCTTCACCCGGGGCAGCTGAA*G*T (SEQ ID NO: 2), where * is a phosphorothioate bond) at 37°C for 1 hr. Samples were then washed three times with PBS + 0.1% Tween 20. Next, samples underwent rolling circle amplification using lx Phi29 Buffer (Lucigen), 10% v / v Phi29 enzyme (Lucigen), 0.2 mg / ml BSA, 250 pM dNTP (New England Biolabs), 25 pM aminoallyl-dUTP (ThermoFisher), and 1% v / v murine RNase inhibitor for 2 hrs at 37°C. Finally, the samples were washed three times in IxPBS, and then crosslinked for 30 min with IxPBS + 1 mM BS(PEG)9 at room temperature, before a final quick wash with IxPBS. Samples were stored in IxPBS with 1% v / v murine RNAse inhibitor at 4°C for up to one week.

[0190] Multiplexed. RNA and protein imaging by RCA-MERFISH and sequential hybridization

[0191] RCA-MERFISH samples were imaged on a custom epifluorescent microscope with automated fluidics, as previously described for MERFISH imaging (Allen, W.E., et al. Cell 186, 194-208.el8). Briefly, samples were mounted in a flow cell (Bioptechs) with a 0.75- mm-thick flow gasket on a Nikon epifluorescence microscope.

[0192] For each round of hybridization, mixtures of Cy7-, Cy5-, Cy3-labeled readout probes for each triplet of bits to be read out were diluted to a final concentration of 10 nM / probe in 5 mL of 2xSSC, 10% formamide, 0.1% Triton X-100. The samples were stained for 15 min, then washed with 2xSSC, 10% formamide, 0.1% Triton X-100. Finally, imaging buffer was flowed into the chamber. The imaging buffer contained 2xSSC, 10% w / v glucose (Sigma), 60 mM Tris-HCl pH8.0, ~0.5 mg / mL glucose oxidase (Sigma), 0.05 mg / mL catalase (Sigma), 50 pM trolox quinone (generated by UV irradiation of 6-hydroxy-2, 5,7,8- tetramethylchroman-2-carboxylic acid (Sigma)), 0.2% v / v murine RNAse inhibitor, and 0.1% v / v of Hoechst 33342 dye (ThermoFisher).

[0193] After the readouts were hybridized and imaging buffer added, the samples were imaged with a high-magnification, high-numerical aperture objective. For wildtype animals, a 60X 1.4NA oil immersion objective was used, with a pixel size of 108 nm / pixel. For genetically perturbed experiments, a 40X 1.3NA oil immersion was used, with a pixel size of 162 nm / pixel. Each field of view (FOV) was imaged with a 10-plane z stack with 1.5 pm spacing between adjacent z planes, where each z-plane was imaged in the 750-nm, 650-nm, 560-nm, and 405-nm channels. After each round of imaging, the readout probes were stripped off using 2xSSC, 80% formamide stripping buffer for 10 min, followed by two washes of readout buffer, and one wash of 2xSSC. The different panels (antibodies and structural RNAs, endogenous RNA, and barcode RNA) were imaged back-to-back on the same tissue sections, where protein (labeled with oligonucleotide-conjugated antibodies) and structural RNA were first imaged by with sequential rounds of multi-fluorescence FISH, followed by endogenous RNAs and barcode RNAs in two separate RCA-MERFISH runs. Each experiment took 36-48 hours depending on the number of fields of view, and whether the perturbation barcode library was imaged in addition to the endogenous RNA and protein panels.

[0194] Animals

[0195] Male and female wildtype and B6;129-Gt(ROSA)26Sortml(CAG-cas9*,- EGFP)Fezh / J in a C57BL / 6J background were used in this study. Mice were obtained from Jackson labs or bred at Harvard or MIT. Mice were maintained on a 12 hr light / 12 hr dark cycle (14:200 to 02:00 dark period) at a temperature of 22 ± 1 °C, a humidity of 30-70%, with ad libitum access to food and water unless otherwise noted. Mice were fed ad libitum with a diet containing 60% kcal from fat (Research Diets, Inc. D12492i) or standard chow.

[0196] Cell lines

[0197] HEK 293T / 17 (ATCC CRL- 11268) cells were cultured in DMEM supplemented with glutamax, HEPES, 10% fetal bovine serum, 100 units / mL penicillin, and 100 mg / mL streptomycin (ThermoFisher Scientific). K562-CRISPRi cells (Gilbert et al, 2014) were cultured in RPMI supplemented with glutamine, HEPES, 10% fetal bovine serum, 100 units / mL penicillin, and 100 mg / mL streptomycin (ThermoFisher Scientific). AML 12 cells (ATCC CRL-22540 were cultured in DMEM / F12 medium supplemented with 10% fetal bovine serum, 10 mg / mL insulin, 5.5 mg / mL transferrin, 5 ng / mL selenium, and 40 ng / mL dexamethasone (ThermoFisher Scientific).

[0198] Lentiviral Construct Cloning

[0199] The parental mosaic screening vector (pVVl) was generated from a modified CROP- seq vector (pBA950; addgene 122239 123) and from a lentiviral vector used for CRISPR screens in the liver (pLentiCRISPRv2-Stuffer-HepmTurquoise2; addgene 192826 27) using standard methods. First, the EFla-BFP was replaced by mTurquoise driven by a hepatocytespecific promoter. Then, the parental mU6 promoter was replaced by an mU6 promoter that did not contain a loxP site 124, yielding pVVl (FIG. 7).

[0200] CRISPR Guide Library Cloning

[0201] The parental pVVl vector was digested with BstXI and BamHI and purified by gel electrophoresis. Library elements containing both sgRNAs and their associated barcodes were ordered as eblocks and pooled before cloning. The library elements were synthesized with three different sgRNA constant regions, which decreases recombination between the sgRNA and the barcode during lentiviral packaging. The library elements were also digested with BstXI and BamHI and purified by gel electrophoresis. The digested library was ligated into digested pVV 1. The ligation was purified with a silica column (Zymo) and electroporated into MegaX DH10B T1 R electrocompetent cells (ThermoFisher Scientific) according to the manufacturer’s protocol. The cells were recovered and then directly introduced into liquid culture and maxiprepped the next day. Samples were plated and used to confirm >1000x library coverage during cloning. Colonies were sequenced to confirm cloning fidelity.

[0202] Lentiviral Preparation

[0203] The lentivirus was generated according to standard methods by the transfection of 293T / 17 cells with Fugene HD (Promega), our library vector, psPAX2, and pMD2.G. ViralBoost (Alstem) was used according to the manufacturer’s instructions. Supernatant was harvested, filtered through a 0.45 pm PES membrane, and a lOx concentration was conducted with PEG and NaCl (Lenti-X Concentrator; Takara) according to the manufacturer’s instructions. The concentrate was then pelleted by centrifugation in a swinging bucket centrifuge (25,000g, 2 hours, 4°, Beckman Coulter SW 32 Ti) and resuspended in cold PBS + 4% glucose, leading to an additional lOOx concentration. Lentivirus was flash frozen and titered approximately through the transduction of AML 12 cells. sgRNA probes for Fixed Cell scRNA-seq

[0204] Probes for the sgRNAs were obtained from IDT as opools. The left-hand side probes targeted the sgRNA constant region; three variants of this probe were included in each hybridization, targeting each of the three constant regions included in the sgRNA library. The left-hand side probes also had a sample barcode sequence. The right hand side probes were 5’ phosphorylated and targeted the protospacer sequence directly. The barcode spike in probes contained the appropriate TruSeq sequences such that they could be amplified separately from the transcriptome probes and sequenced independently. A variable number of N bases was included to ensure base diversity during sequencing.

[0205] Mosaic Liver Preparation

[0206] The lentivirus was thawed on ice and up to 50ul was injected into the temporal vein of postnatal day one mice 27,125. Approximately 1 x 107titer units of lentivirus were injected per animal. The mice were allowed to grow to adulthood (>P30) and Cas9 induced through the retro-orbital injection of AAV8 with Cre driven by a hepatocyte promoter (Addgene 107787-AAV8; ~5 x 1011genome copies per animal). The mice were maintained for ten days, or alternatively maintained for nine days and then fasted for 16 hours. The mice were anesthetized and fixed by perfusion of PBS + 4% paraformaldehyde. Livers were removed and fixation was continued for 3 hours in PBS + 4% paraformaldehyde, followed by ~13 hours in PBS + 4% paraformaldehyde + 30% sucrose. Lixed livers were then frozen in Optimal Cutting Temperature compound (OCT) and stored at -80°. Mosaic Liver Dissociation for Fixed Cell scRNA-seq

[0207] I OOp m sections of liver were generated on a cryostat and stored at -80° for later processing. The sections were washed with cold 0.5x PBS to remove residual OCT and then resuspended in warm RPMI + 1 mg / ml Liberase Th. The material was transferred to a gentleMACS C tube and dissociated on a gentleMACS Octo Dissociator with heaters (Miltenyi) with the program 37°C_FFPE_1. The cells were strained through a 30pm preseparation filters (Miltenyi) and singlets (diet experiment) or GFP+, mTurqiouse+ singlets (Perturb-seq experiment) were isolated by FACS (ARIA II, BD) and maintained at 4° in 0.5x PBS until scRNA-seq.Dato Processing Pipeline

[0208] RCA-MERFISH gene expression was processed using a modified version of the MERlin pipeline, as previously described (Allen, W.E., Cell 186, 194-208. el8, 2023), with the addition of a machine learning filtering step that used XGBoost to train a classifier to discriminate incorrectly decoded molecules that were assigned to blank barcodes and putatively correctly decoded molecules that were assigned to coding barcodes on a subset of the data. This classifier was then applied to the remainder of the data, and only molecules that were classified as coding molecules with an adaptive 5% false-positive threshold were exported for assignment to individual cells. Each library (endogenous RNA or barcodes) was decoded separately.

[0209] Cell segmentation was accomplished using a custom Cellpose model that was trained on images of polyA staining for cytoplasm + nucleus, and Na+ / K+ ATPase staining for membrane. A separate model was trained for data collected using the 60X and 40X objectives and applied to the respective datasets. Each cell was assigned a unique identifier and its position, shape, and z extent were recorded. The molecules that were exported from each field of view were then assigned to cells based on overlap between molecules and Cellpose- created mask for each cell in three dimensional space. The total number of molecules of each type for each cell was summed, and the result exported as an AnnData objective.

[0210] After decoding and cell assignment, all RCA-MERFISH datasets were concatenated into a single dataset. This was done separately for the wildtype physiological perturbation data and the genetically perturbed data. Cells were filtered to remove all cells with less than 25 or greater than 1500 molecules per cell, and genes expressed in <3 cells were removed. Each cell’s gene expression values were scaled such that the sum of gene expression values per cell added to 10,000, then was log-transformed. The area and number of molecules per cell were regressed out using linear regression, and the residuals were z-scored. The first 20 principal components were then computed and the data were integrated with log-transformed, z-scored Flex data using Harmony, using the top 40 principal components computed based on the subset of genes measured in RCA-MERFISH. The data were then jointly clustered using Leiden clustering with resolution = 0.4. Clusters with fewer than 10 differentially expressed genes were then greedily merged, and the final set of clusters were manually annotated based on marker gene expression. For perturbation data, the individual guide was called for each cell when there were > 3 molecules per cell for a given barcode.

[0211] The labels assigned to clusters in the integrated Flex and RCA-MERFISH data from wildtype animals were transferred to the perturbed data through integration and label transfer. After pre-processing to filter out cells, normalize, log transform, and z-score the data as described above, the two datasets were integrated using Harmony. A K- nearest neighbors classifier was then trained to predict the clusters annotations of the wildtype RCA-MERFISH data from the top 20 principal components of the data, with K=10. This classifier was then applied to the perturbed RCA-MERFISH data to predict cluster identity.

[0212] For morphological imaging, each channel used for morphological imaging was adaptively contrast adjusted. A single z-plane in the middle of each segmented cell was selected, and the image of each cell in each channel was cropped out of the larger field of view, in a 256 x 256 square, using the segmentation for that cell as a mask. The background around each cell outside of the mask within that 256 x 256 square was set to zero. The imaging data for each cell was then associated with the unique identifier for that cell, for integration with RCA-MERFISH data for endogenous RNA and perturbation barcodes.

[0213] Deep autoencoder model

[0214] A deep autoencoder model was developed to extract biologically relevant features from single-cell multiplexed protein and abundant RNA images based on the second generation of the vector quantized variational autoencoder (VQ-VAE) (Razavi, A., et al. Generating Diverse High-Fidelity Images with VQ-VAE-2. arXiv [cs.LG], 2019; accessible at doi.org / 10.48550 / ARXIV.1906.00446). The model takes advantage of the power of VQ- VAE to learn meaningful representations using self- supervised training. Inspired by the cytoself model (Kobayashi, H., et al., Self-Supervised Deep Learning Encodes High- Resolution Features of Protein Subcellular Localization. bioRxiv, 2021.03.29.437595; accessible at doi.org / 10.1101 / 2021.03.29.437595), auxiliary classification tasks were used to guide the model to focus on features of biological importance. Each single-fluorescence protein image of a particular protein / RNA channel was concatenated to a fiducial channel to create a two-color image. The polyA FISH staining was used as the fiducial channel to provide information of the relative location of the stained proteins / RNAs to the nuclei. From the two-fluorescence image, the model used a ResNet with two residual blocks to generate a bottom level representation as described in the second generation VQ-VAE design. The bottom level representation was then used to generate a top level representation with another ResNet with two residual blocks. Both bottom and top representations were converted to discrete representations by vector quantization (Razavi A., et al). The quantized bottom and top representations were concatenated and decoded by a ResNet with two residual blocks to regenerate the input image. The concatenated representation vectors were fed into a multilayer perceptron (MLP) classifier with one hidden layer to predict the identity of the protein / RNA channels. The hidden representation of the MLP classifier had 512 dimensions, which was used as the representation vector of the input image. The representation vectors for different protein / RNA channels of one cell were concatenated to generate the morphological representation of a cell. The cell representation was fed into another MLP classifier to predict the transcriptionally defined cell type and / or the diet condition of the cell.

[0215] The loss function of the autoencoder model contained 4 terms, namely, the latent loss, the image reconstruction loss, the protein / RNA classification loss and the cell-type or dietcondition classification loss. The same latent loss definition as the VQ-VAE model (Razavi A., et al) was used, which measured the difference between the latent representations before and after quantization. The image reconstruction loss was defined as the mean squared error of the reconstructed image compared to the input image. The protein / RNA and cell-type / diet- condition classification losses were defined as the cross entropy loss of classification.

[0216] Training the deep autoencoder

[0217] Images of single-cells were cropped out from the fields of view as square boxes. All the pixels outside the target cells were set to zero using the cell segmentation masks, such that each cropped image only contained one cell. The cropped single-cell images were rescaled to 128x128 pixels and used as input for the autoencoder training. The autoencoder model was optimized with stochastic gradient descent using the Adam optimizer until the loss function converged.

[0218] A VQ-VAE model was trained to embed cells under different diet conditions using both the protein classification auxiliary task and the cell classification auxiliary task that predicted transcriptionally defined cell types and diet conditions. Lor embedding the cells under CRISPR perturbations, because only hepatocytes were included, the autoencoder model was trained using only the protein classification auxiliary task.

[0219] Tissue zone segmentation The zonal segmentation of tissue was accomplished by computing for each replicate a 2D histograms at 50 pm resolution of the number of Hepl+Hep2 or Hep5+Hep6 cells in each bin. These histograms were then blurred with a Gaussian filter with sigma = 0.5 and normalized by dividing by the maximum across all bins. For each bin, the zone was determined as whether the normalized Hepl+Hep2 or Hep5+Hep6 count was greater for that bin.

[0220] QUANTIFICATION AND STATISTICAL ANALYSIS scRNA-seq Analyses

[0221] Energy Distance Calculation and Permutation Testing

[0222] Energy distances and permutation tests were calculated according to the method of Peidli, S., et al., scPerturb: harmonized single-cell perturbation data. Nat. Methods 21, 531— 540, 2024. The first 20 PCs were used generated from the tp 10k- normalized, loglp- transformed data. The Holm-Sidak multiple testing correction was used for the permutation testing. For the fasted vs ad libitum comparisons, the transcriptional states of perturbed cells were compared to those of cells with control sgRNAs in the same mouse in the same condition.

[0223] Z-scoring relative to control

[0224] For some purposes, z-scored transcriptional data were analyzed. The z-scores were calculated by tplOk-normalizing the data, identifying the mean and standard deviation for each mRNA in control cells, and then used these values to identify the z-score for each gene in each cell, relative to control cells. This calculation was performed separately for each barcode in each GEM group and then concatenated the cells to form a final z-normalized expression matrix. This transformation emphasizes changes in mRNAs with low variance in control cells and may decrease batch effects between barcodes and GEM groups.

[0225] Pseudobulk Correlation Calculations

[0226] Correlations calculated between pseudobulk transcriptional responses are either Pearson correlations calculated from mean log-transformed transcriptional responses, with the expression in controls subtracted to a Perturbation-associated phenotype, or correlations of mean Z-scored transcriptional changes. Significant differences between the two were not observed. To decrease noise, only genes with high expression in control cells were included (highest 250 or highest 1000) in the calculations. In Figure 1 IM, sgRNAs targeting genes are shown that have two sgRNAs that have substantial transcriptional phenotypes.

[0227] Differential Expression Benjamini-Hochberg-corrected Mann- Whitney testing was used to identify genes whose expression is significantly affected by perturbations (corrected p < 0.05). tplOk- normalized, loglp-transformed data was used for these calculations. The heat map showing changes in expression (Figure 12D) represents log2-fold changes of pseudobulk expression versus cells with control perturbations.

[0228] Scoring and significance testing

[0229] The scores in Figures 11 A and 1 IB were calculated using gene sets derived from a previous large-scale perturbation experiment in cell culture (Replogle, J.M., Cell 185, 2559- 2575.e28, 2022). The scores were calculated as the pseudobulk z-scored change relative to control cells, averaged for all genes in the gene set, for all cells with each perturbation. The scores in Figures 8 A, 8B, and 8C were calculated with a literature-curated set of zonation markers and reflect the sum of pseudobulk or single cell z-scored changes relative to control cells. The periportal markers were Cyp2f2, Hal, Hsdl7bl3, Sds, Ctsc, Aldhlbl, and Pckl and the pericentral markers are Cyp4al4, Cyp2d9, Gstm3, Cyp4al0, Mupl7, Slcla2, Slc22al, Cypla2, Aldhlal, Cyp2a5, Gulo, Cyp2c37, Lect2, Cyp2el, Oat, Glul. Periportal expression contributed positively to the score and pericentral contributed negatively to the score, or alternatively they were shown separately. Significance was calculated relative to negative control cells by Mann- Whitney with Benjamini-Yekutieli correction for multiple testing (corrected p < 0.05).

[0230] Clustering and genotype-phenotype mapping

[0231] The perturbation heatmap was generated from the hierarchical clustering of a joint vector including Pearson correlations of pseudobulk log-transformed transcriptional responses measured by sequencing and pseudobulk z-scored staining intensity changes measured by imaging (Euclidean distance metric, UPGMA algorithm). The representation of gene co-regulation in Figure 10E was generated from correlations between z-scored pseudobulk expression levels of pairs of genes across perturbation (the transpose of the transcriptional component of the data used to generate the correlations in Figure 10D). The position of spots derived from a two-dimensional minimal distortion embedding of these correlations that tried to place co-regulated genes in proximity. The gray shades of spots derived from a density-based clustering of a separate twenty-dimensional minimal distortion embedding of co-expression, calculated according to the method used for mRNAs in Replogle, J.M., et al. Cell 185, 2559-2575.e28, 2022.

[0232] Intensity Analysis

[0233] Z-scoring relative to control When analyzing intensity, each protein / RNA channel was z-scored relative to the mean and standard deviation of that channel in control cells (cells with control sgRNAs), from that same imaging sample. The cells were then concatenated from the various imaging samples to form a final z-normalized intensity matrix. This transformation emphasized changes in protein / RNA intensities with low variance in control cells and decreased batch effects between imaging samples.

[0234] Pseudobulk Correlation Calculations

[0235] Correlations calculated between pseudobulk protein / RNA intensity responses were Pearson correlations calculated from mean z-scored intensity responses. Specifically, perturbations targeting genes that have two sgRNAs that are significant in a corrected energy distance permutation test were evaluated.

[0236] Number of Protein / RNA Channels Exhibiting Differentially Intense Signals

[0237] To quantify the number of differentially intense protein / RNA channels, Benjamini- Hochberg-corrected Mann-Whitney testing was used (corrected p < 0.05) on the z-scored intensity data.

[0238] Intensity Scoring and significance testing

[0239] The scores in Figures 11C, 11D, HE, 11F, 12A, 12B, and 12G were calculated as the mean z-scored change relative control cells for all cells with each perturbation. Significance was calculated relative to negative control cells by Mann-Whitney with Benjamini-Yekutieli correction for multiple testing, corrected p < 0.05.

[0240] EXAMPLE 2

[0241] This example describes the development of an approach for performing large-scale pooled genetic screens in native tissue with rich, multimodal phenotypic readouts, as well as its application to map the function of hundreds of genes in the mouse liver under multiple physiological conditions. The effects of many genetic perturbations on diverse cellular phenotypes, including transcriptional state, subcellular morphology, and tissue organization, in the liver were determined. Developing this approach required solving multiple technical challenges towards multimodal, in vivo phenotype mapping through both imaging and sequencing. Specifically, new methods for fixed-cell Perturb-seq as well as imaging-based pooled genetic screening in heavily fixed tissue were developed, enabling joint sequencing and imaging analysis of diverse perturbations in the same tissue. For the latter, multiplexed protein and RNA imaging were applied, using in situ enzymatic probe amplification followed by multiplexed error-robust FISH (MERFISH) to read out both endogenous RNAs and short barcodes for genotyping. Through an integrated analysis of morphological and transcriptional phenotypes, novel regulators of hepatocyte zonation were identified and revealed how proteostatic stress pathway activation could broadly affect the expression of secreted proteins, and showed how diverse cellular pathways could produce convergent effects on steatosis. Beyond enabling new ways of interrogating the genetic basis of complex cellular and organismal physiology, this approach provided crucial training data for emerging machine leaming / AI efforts to create predictive models of “virtual” cells.

[0242] Pooled in vivo genetic screens through imaging and sequencing

[0243] An approach was designed to profile the effects of a large number of genetic perturbations at subcellular resolution in different cell types in the tissue of a living mouse under defined physiological conditions. For each perturbation, multiple distinct dimensions of cellular state were measured to capture a comprehensive picture of gene function. Pioneering work established the liver as an important system for performing large-scale in vivo pooled genetic screens, due to the effectiveness of viral delivery to hepatocytes in homeostasis, disease, and regeneration. The previous studies focused on cellular survival / proliferation as a phenotype through barcode sequencing. Whereas, this approach was designed to explore complex, multimodal phenotypes at a far wider range of gene functions.

[0244] To this end, an approach was developed where hepatocytes in a transgenic Cas9 mouse were infected in a mosaic with a pool of lentiviruses, each delivering a different CRISPR guide RNA (). After perfusion with paraformaldehyde to rapidly fix cellular phenotypes, perturbed cells were interrogated using sequencing- or imaging-based approaches to capture multiple aspects of cellular state while simultaneously identifying perturbed genes. From single-cell sequencing, the effects of perturbations on genome-scale transcriptional profiles were measured to obtain insights into regulators of cell state in distinct cell type. From imaging, a set of phenotypes was measured, including the subcellular morphology and quantity of proteins and mRNAs, as well as the cellular organization of the tissue. By linking cellular properties measured through different modalities via their shared perturbation identity, integrated multimodal analyses was then performed to examine the multifarious effects of the same perturbation on different aspects of cellular function. For sequencing-based phenotyping, the 10X Genomics Flex technology was used. This technology uses microfluidic encapsulation of dissociated single cells that have been hybridized with a transcriptome- wide library of split probes against different RNA species for transcriptome-wide detection). Pairs of split probes that hybridized adjacent to each other were then ligated in the droplet and given a cellular barcode for sequencing and expression quantification. This approach was extended to also measure CRISPR-based genetic perturbations through the inclusion of custom probes that directly target the sgRNA.

[0245] An imaging approach for genotyping and phenotyping intact tissue sections was created. This approach, had several outcomes: 1) Created a multimodal protocol for the readout of both RNA and protein from tissue fixed in a manner that was standard for high resolution immunofluorescence, in a single sample preparation on an automated microscope. 2) Amplified RNA FISH signals, both in order to increase the throughput of data collection and to allow the readout of short barcode sequences associated with the sgRNAs. 3) Produced stable samples that could be stored and imaged later, such that many samples could be prepared in parallel and then imaged sequentially.

[0246] RCA-MERFISH was selected for spatial transcriptomics and proteomics analyses. MERFISH was used to read the molecular identity and for signal amplification prior to MERFISH detection, rolling circle amplification (RCA) was used. Specifically, both endogenous mRNA and barcode RNA (for sgRNA identification) with padlock probes that can be ligated upon binding to target RNA and amplified in situ (FIGs. 1A-1E) were targeted. Each padlock probe served as an encoding probe and contained multiple readout sequences that together provided a MERFISH code for the target molecule, which could be detected over multiple rounds of hybridization by readout probes complementary to the readout sequences (FIG. 1A-1E). Heavy fixation was preferred for immunofluorescence imaging of proteins, therefore, to overcome limitations on performing in situ enzymatic reactions for RCA in heavily crosslinked tissue, a hydrogel-based RNA-embedding and tissue clearing approach and embedding of RCA-amplicons in a polyacrylamide gel was adopted (FIG. 1A, IB). Finally, this approach was combined with multiplexed immunolabeling of proteins with oligo-conjugated antibodies, which were then detected with complementary readout probes using sequential rounds of hybridization. Several abundant cellular RNAs were detected through sequential FISH.

[0247] The final protocol used a series tissue fixation, embedding, clearing, labeling and imaging steps (FIG. 1A-1E, and 2). Briefly, after perfusion and cryopreservation, thin sections of tissue were cut, decrosslinked, and stained with a pool of oligonucleotide- conjugated antibodies targeting different subcellular organelles, membrane, and proteins. After antibody staining, all RNA and DNA were modified in the tissue with an alkylating agent containing an acrylamide moiety (MelphaX), and the sample was embedded in a thin acrylamide gel, such that the RNA and DNA, including cellular RNAs, sgRNAs with barcodes, and antibody-associated DNA oligonucleotides were covalently linked to the gel. The proteins were then digested and the lipids were washed away to both clear the tissue and make the remaining RNA and DNA accessible to enzymes and probes. After clearing, a library of padlock probes targeting both endogenous RNA and perturbation barcodes was hybridized. After stringent washing, the RNA-hybridized padlock probes were ligated and rolling circle amplification was performed. The final protocol was the result of extensive optimization of decrosslinking conditions, oligo modifications, additives, and digestion conditions to identify optimal conditions that allowed for amplification while being compatible with immunostaining of proteins with oligonucleotide-conjugated antibodies in heavily fixed tissue (FIG. 4A-4J; STAR Methods).

[0248] For the RCA-MERFISH detection, each RNA species was targeted by eight padlock probes tiling the transcript and each padlock probe comprised two binding arms that hybridized adjacent to each other to be ligated, and 4-6 readout sequences for encoding the target RNA (FIG ID). These probes are relatively long (-140-170 nt) and may be full-length in order to ligate properly. A method to produce these probes was developed from low-cost femtomolar pools of long oligos synthesized in array (FIG. 3). After testing multiple different options, including a previously used reverse transcription followed by in vitro transcription method (Chen, K.H., et al., Science 348, aaa6090, 2015), an approach that used limited cycle PCR and then RCA to amplify the initial array-synthesized oligo pool was utilized, followed by Type IIS restriction enzyme digestion to liberate the full length, phosphorylated probes (FIG. 3A-3E).

[0249] EXAMPLE 3

[0250] Mapping hepatocyte transcriptional gradients through RCA-MERFISH

[0251] The imaging approach was applied in combination with scRNA-seq to map the spatial organization of distinct cell types and cell states in the mouse liver. The liver is a classic system for studying cell biology and has allowed discoveries over the past century ranging from the isolation of different organelles to uncovering mechanisms of mTor signaling. The liver provided a test system, as the principal cell type - hepatocytes - have a well-defined molecular composition and spatial organization, and are readily infected with virus for perturbation delivery (see, FIG. 6A, 6B). At the anatomical level, the liver has long been understood to have distinct cellular ‘zones’ radially organized with respect to the central and portal veins (FIG. 8). Periportal and pericentral hepatocytes have distinct gene expression and correspondingly distinct physiological functions. For example, periportal hepatocytes specialize in oxidative metabolism, gluconeogenesis, and the urea cycle, whereas pericentral hepatocytes specialize in glycolysis and lipogenesis. This periportal to pericentral distinction was repeated in multiple units around different central veins in a crystalline manner throughout the organ (data not shown).

[0252] After optimization, the multimodal approach allowed the endogenous gene expression to be measured while imaging multiple subcellular structures simultaneously. In total, across 7 rounds of 3 fluorescent dye imaging, with a total of 21 bits (Table 1), 209 genes were measured (see example in Table 2) using a Hamming Weight 4 (HW4), Hamming Distance 4 (MHD4) codebook, which resolved hundreds of amplicons per cell. This gene panel was primarily focused on hepatocytes but also included markers for other cell types. In addition to combinatorial RNA imaging with RCA-MERFISH, 14 protein labelings were sequentially imaged with various subcellular organelles, lipid droplets, and signaling pathways and 4 abundant RNA species that exhibited specific patterns of subcellular localization (FIG. 9A, 9B). One of the protein targets, the membrane marker Na / K+ ATPase, was used for cell segmentation.

[0253] Unsupervised clustering revealed a number of hepatocyte subtypes and nonhepatocyte cell types. Different cell types were classified as well as subtypes of hepatocytes using an integrated analysis of RCA-MERFISH and scRNA-seq data. The analysis focused on hepatocytes. A periportal and pericentral gene expression score was computed based on summing the expression of known marker genes such as G6pc, Aldhlbl , and Pckl and it was determined that the hepatocyte clusters fell onto a zonated continuum from Hepl to Hep6, with Hepl being the most pericentral and Hep6 being the most periportal (). Visualizing these zonation scores and hepatocyte subtypes across the tissue revealed the expected patchwork localization of periportal and pericentral cells (), whereas non-hepatocytes had less obvious spatial arrangement. In addition, one subtype of hepatocytes (Hep3) expressing multiple interferon signaling markers that were localized in small patches was identified (data not shown; markers included lsgl5, Statl , and Irj7). The expression of the full transcriptome was imputed for each cell measured with RCA-MERFISH from the integration with scRNA-seq data and a pattern of zonated gene expression was observed for imputed genes. Finally, it was determined that a large fraction of the transcriptome in hepatocytes varied along the periportal to pericentral axis. EXAMPLE 4

[0254] Heterogeneity in subcellular morphology among hepatocytes

[0255] To capture diverse aspects of cellular state through high-resolution microscopy, a panel of 18 different protein and RNA targets that covered a variety of subcellular organelles, morphological features such as plasma membrane, and phosphorylated states of proteins was devised. Targeted proteins included both organellar and structural markers such as GAPDH (cytoplasm), Calreticulin (endoplasmic reticulum), Perilipin (lipid droplets), Na+ / K+ ATPase (plasma membrane), CathB (lysosome), and Tomm20 (mitochondria), as well as cell statedependent markers such as the mTOR marker phospo-S6 ribosomal protein (pS6RP) (Table 3). In addition to proteins, a few RNA species too abundant to be included in MERFISH imaging, including the highly abundant Albumin mRNA, pre-rRNA in the nucleolus, mitochondrial RNA (mtRNA), as well as total polyadenylated RNA (FIG. 9A, 9B) were imaged.

[0256] To analyze this highly multiplexed cellular imaging data, self- supervised learning methods for dimensionality reduction were used. As part of a data processing pipeline that was developed (FIG. 5A, 5B), each image channel (i.e. each protein or abundant RNA target) in each cell was reduced from a high-resolution z- stack to a 128x128 pixel matrix and trained a VQ-VAE network to minimize the dimensionality of that protein or RNA stain further, to a 512-dimensional vector (FIG. 5C, 5D and 9C). The dimensionality reduction was constrained by an auxiliary task that sought to use the low-dimensional embedding to discriminate different transcriptionally-defined cell types or physiological conditions, in order to find an embedding that reflected true biological variation in the sample. These embeddings were visualized with UMAP and it was found that embeddings representing images of morphologically similar targets were indeed grouped together (FIG. 4D); this was also observed through quantification of mutual information between protein / RNA targets (FIG. 5E). Further analysis of these embeddings revealed that individual features within the embedding often reflected interpretable patterns (data not shown).

[0257] Quantifying the variations in subcellular morphology with single-cell resolution at the scale of entire tissues was a challenging task. The machine-leaming-based feature analysis provided a means to overcome this challenge. Each 512-dimension embedding for a protein / RNA across all cells was visualized for individual features. It was determined that many protein / RNA features exhibited zonal spatial patterns similar to those observed in hepatocyte gene expression. For some proteins or abundant RNAs, such zonated patterns were observed, such as Albumin mRNA and Perilipin. For others zonated profiles to a greater or lesser extent (Rab7, CathB) were observed. The zonated variation in subcellular morphology of these proteins may reflect differences in hepatic metabolism across different liver zones.

[0258] To compare the relative information contained in protein embeddings versus transcriptome in more detail, the relative information that was contained in the single-cell transcriptional profiles versus the morphological feature embeddings in the same cells was quantified. A classifier was trained to predict the subtype identity of individual hepatocytes - from Hepl to Hep6 - on held out cells, using either the 209 genes measured through RCA- MERFISH or the concatenated 18 x 512 dimensional feature embeddings. It was determined that the transcriptional profiles discriminated all subtypes of hepatocytes (FIG. 9D), whereas the morphological feature embeddings only distinguished the most periportal and pericentral subtypes Hepl and Hep6 (FIG. 9E). By contrast, the other subtypes were less clearly distinguished, suggesting that the transcriptionally-defined intermediate subtypes may exhibit a range of subcellular morphologies.

[0259] To explore the differences in subcellular morphology between the most periportal and pericentral hepatocyte subtypes, the morphological feature embeddings that best discriminated between Hepl and Hep6 cells were identified (FIG. 9F). Albumin mRNA and Perilipin were the most dissimilar between these two hepatocyte subtypes, consistent with the previous understanding of zonated Albumin expression and lipid storage and metabolism. The anti-Perilipin embeddings were visualized with UMAP and it was found that the embeddings representing cells with different hepatocyte subtype identity were separated; an unsupervised clustering of the embeddings was conducted and it was found that periportal Hep6 cells were most enriched in anti-Perilipin embedding cluster 2, which had very few lipid droplets, whereas pericentral Hepl cells were most enriched in anti-Perilipin embedding cluster 6, which had many lipid droplets (data not shown).

[0260] To understand the features of liver cellular state that respond to changing physiological conditions and metabolic stress, mice were fast overnight or fed a high-fat diet (HFD) for a month and tissue samples were fixed for analysis by scRNA-seq and imaging. These physiological perturbations caused large changes in gene expression, especially in hepatocytes. The morphological images were embedded with the VQ-VAE model to identify image channels whose embeddings best discriminated between these two physiological conditions. Phospho-S6 ribosomal protein (pS6RP) staining separated hepatocytes between the fasted and ad libitum (ad lib) samples best, consistent with S6 ribosomal protein phosphorylation by mTOR in response to nutritional status (data not shown). Perilipin differed in cells from the ad lib and HFD samples (data not shown). Further, it was observed that stains such as anti-calreticulin and anti-pS6RP were also different between the ad lib and HFD samples. This may in part reflect negative space caused by the exclusion of these stains by large lipid droplets, in addition to potential changes in the abundance of these targets caused by the altered dietary conditions.

[0261] In total, this data demonstrated that the multiplexed protein and RNA imaging, combined with deep learning analysis, was able to resolve significant heterogeneity in subcellular morphology, which could be further linked to distinctions in transcriptionally- defined cell types and the animal’s physiological state. These experiments created the first integrated transcriptional and subcellular morphological map of the liver, while validating the ability of protein and RNA readouts to provide complementary information on liver anatomy and physiology.

[0262] EXAMPLE 5

[0263] Large-scale in vivo multimodal screening in CRISPR mosaic livers

[0264] Having established an integrated sequencing and imaging methodology for studying the molecular and cellular organization of tissue at multiple scales, this approach was combined with large-scale pooled in vivo CRISPR screening to map the effect of genetic knockout perturbations on genome- wide transcriptional state as well as the intensity and morphology of the imaged proteins and RNAs. A total of 202 genes were targeted, each with two independent sgRNAs, and included 50 negative control sgRNAs. The genes chosen for perturbation were curated from the literature and were involved in diverse aspects of hepatocyte and liver physiology, including metabolism, intercellular and intracellular signaling, transcription and translation, and protein trafficking and secretion. Genes selected for high or specific expression in hepatocytes were also targeted, including genes with uncertain functions in hepatocytes.

[0265] Barcoded genetic perturbations that could be read out through either imaging or sequencing into the liver through lentiviral delivery were introduced. Building on the CROP- seq lentiviral vector design to mitigate the effects of recombination during lentivirus production, a vector that expressed a fluorescent protein (mTurquoise) as well as a 185-mer barcode associated with a sgRNA for imaging from a hepatocyte-specific Pol II promoter while simultaneously expressing the sgRNA from a Pol III promoter was engineered (FIG. 8A). Each sgRNA and its associated 185-mer barcode were synthesized on a single DNA fragment and then cloned into the vector in a pool. An intervening region composed of one of three alternative sgRNA constant regions was introduced between the barcode and sgRNA to reduce the effects of barcode swapping through recombination of the constant intervening sequence.

[0266] This vector was used to introduce a mosaic of genetic perturbations into the liver of Cas9 transgenic mice. High-titer pools of CROP-seq lentivirus were produced and the virus was introduced systemically into transgenic Cre-inducible Cas9 mice at postnatal day 1 (Pl) through injection into the temporal vein. A relatively low MOI (10-30%) was used, such that most infected cells had a single perturbation. Once the mice reached adulthood (>P30) and Cas9 expression was induced in the hepatocytes with systemically delivered AAV containing the Cre recombinase driven by a hepatocyte- specific TBG promoter. After waiting 10 days to allow Cas9 to express and perturb the targeted gene, the mice were perfused and the fixed livers were cryopreserved for analysis. Lentivirus introduced at Pl expressed mTurquoise in mosaic distribution throughout the liver of adult mice, whereas EGFP expressed with Cas9 was uniformly expressed in presumptive hepatocytes following Cre recombination (data not shown). The effects of these perturbations on hepatocytes were mapped using a combination of imaging and sequencing.

[0267] For imaging-based phenotype and genotype measurement, the RNA and morphological imaging approach described above was extended with the inclusion of padlock probes targeting the perturbation-specific barcode sequences, which were orthogonal to the mouse genome. These perturbation- specific probes carried the readout sequences that encoded the perturbations (e.g., the sgRNA, but this could be extended to other genetically- encoded perturbations) with a HW6, HD4 MERFISH code, which were imaged by MERFISH alongside the endogenous mRNA imaging and sequential imaging of proteins and abundant RNAs on the same automated microscope. Gene expression was determined in 7 rounds using 209 mRNA species, morphology imaging was conducted in 8 rounds using 14 proteins and 4 RNAs, and barcode imaging was conducted in 7 rounds using 450 sgRNAs.

[0268] The sgRNA-associated barcodes were expressed only within a subset of cells. After establishing a count threshold to confidently identify cells with each perturbation, it was determined that most cells did not contain any perturbation and that most perturbed cells had only a single perturbation, consistent with the low infection MOI and approximately Poisson- distributed transduction of the lentivirus-accessible cells in the liver (1 call - 85.3%; 2 calls 11.6%; 3 calls 2.1%, and >3 calls 1%). A relatively homogeneous infection was observed throughout the tissue and within lobules. Altogether, approximately -79,000 cells perturbed with single sgRNAs were analyzed. In parallel, to measure genome-wide transcriptional responses to the genetic perturbations, a fixed-cell Perturb-seq method was developed. In addition to compatibility with imaging, it was reasoned that a fixed cell protocol could overcome many of the challenges that have limited the scale and practicality of in vivo Perturb-seq studies, such as sample degradation during tissue dissociation and FACS enrichment, especially when processing many tissues in parallel or recovering millions of cells.

[0269] Dissociation of fixed tissue from the same livers that were imaged yielded intact cells with morphologies similar to those observed in tissue. Despite heavy fixation, the cells could be efficiently dissociated with high yield, obviating the need for nuclear RNA sequencing or complex enzymatic dissociation approaches. FACS was used to enrich for sgRNA-transduced (mTurquoise+) and Cas9-active (GFP+) cells and then fixed-cell single-cell RNA-seq was performed using the hybridization-based 10X Flex approach (see FIG. 1). To determine the identity of expressed sgRNA alongside the mRNA transcriptome, a pool of custom probes that directly targeted the protospacers of the sgRNA library (right hand side probe) and the constant region of the sgRNA (left hand side probe) was spiked in. In most cells, one or two sgRNA probes had much higher UMI counts than other sgRNAs, most of which gave zero UMI counts. This distribution was used to establish a threshold and confidently identify the sgRNAs present in each cell. Most perturbed cells had a single sgRNA above threshold, with only a small fraction of cells harboring two or more distinct sgRNAs (1 call 85.7%; 2 calls 12.4%; 3 calls 1.7%; >3 calls 0.3%), again consistent with the low MOI. Altogether, approximately -55,000 cells perturbed with single sgRNAs were analyzed.

[0270] EXAMPLE 6

[0271] Validation of perturbation calling

[0272] The accuracy of the sgRNA calling was confirmed using Perturb-seq and RCA- MERFISH by inspecting for depletion of sgRNA targets. Fortuitously, sgRNAs targeting Albumin (Alb), an abundant secreted protein whose mRNA provides -10% of the mRNA UMIs in hepatocytes, caused strong depletion of Albumin mRNA relative to control sgRNAs, presumably due to nonsense-mediated decay (NMD) of targeted mRNAs. With Perturb-seq, Albumin mRNA depletion directly was quantified and a minimal overlap was observed between the distributions of Albumin expression in cells expressing sgRNAs against Albumin and those expressing control sgRNA, indicating a low false positive rate of sgRNA calling. A fixed-cell Perturb-seq experiment in cultured CRISPRi K562 cells, which allowed the measurement of on-target knockdown without relying on NMD (median 87% knockdown across targets), was conducted and confirmed the accuracy and specificity of sgRNA calling with hybridization-based scRNA-seq.

[0273] Because probes against Albumin mRNA were also included in the imaging-based screens, raw FISH signal was quantified for Albumin mRNA and a lower intensity in cells expressing sgRNAs against Albumin as compared to cells expressing control sgRNA was observed. Likewise, the immunofluorescence imaging showed that anti-Gapdh signals were lower in cells expressing sgRNAs against Gapdh than in cells expressing control sgRNA (data not shown). These results suggested that perturbation readout by RCA-MERFISH also had a low false positive rate.

[0274] Next, the broader phenotypic consequences of each genetic perturbation was measured by Perturb-seq and imaging. As a first-pass analysis of the imaging data, the mean intensity of each 14 proteins or 4 abundant RNA targets in each cell was quantified.

[0275] An energy-distance permutation test with stringent multiple testing correction was used to determine whether perturbations caused a difference in the distribution of (1) global transcriptional states measured by Perturb-seq and (2) overall morphological states measured by imaging, relative to cells with control sgRNAs. Such global analyses are crucial, as experiments perturbing hundreds of genes and measuring tens of thousands of phenotypes across a multitude of cells can be prone to statistical false discovery. Although the library was designed to be enriched for targets with functions in liver biology, it was not known a priori which sgRNAs could cause phenotype changes in the assays. In the Perturb-seq experiment, 109 / 406 targeting sgRNAs had a statistically significant impact on the distribution of transcriptional states (measured as the first 20 PCs of mRNA expression); at that significance threshold (multiple-testing corrected p value < 0.05), 0 / 50 control sgRNAs had a significant transcriptional impact. In the imaging experiment, 84 / 406 targeting sgRNAs and 3 / 50 control sgRNAs had a significant impact on the distribution of morphological state. Furthermore, pairs of sgRNAs targeting the same gene were compared. Pseudo-bulk sgRNA-phenotype maps were generated by groups of cells with identical sgRNAs and it was found that pairs of sgRNAs targeting the same gene caused both highly correlated transcriptional phenotypes and highly correlated morphological phenotypes, relative to random pairs of non-targeting sgRNAs. Together, these results supported the accuracy of the sgRNA calling and specificity of the perturbations caused by these sgRNAs.

[0276] Perturbations that had the largest impact on the distribution of global phenotypic states, as reflected in energy distance from cells with control sgRNAs were inspected. In the Perturb-seq experiment, many genes with functions essential for cellular growth and survival, such as Atp2a2, Aars, and Sf3b6 caused large impacts to global transcriptional state upon acute knockout. The knockout of signaling genes such as Vhl also had large transcriptional impacts. Many of the same genes also had large impacts on the global immunofluorescence state (hypergeometric p < 10'13). Overall, a positive correlation between the magnitude of the knockout phenotypes in the Perturb-seq and imaging assays was observed (data not shown, Pearson’s R = 0.5). Even though, since the Perturb-seq and imaging experiments measured different phenotypes — genome-wide transcription in the former and 14 proteins and 4 abundant RNAs in the latter — a high correlation between these two measurements was not anticipated.

[0277] Next, the differentially expressed genes associated with each sgRNA were identified. Although most sgRNAs caused differential expression in a relatively small number genes, comparable to the effects of control sgRNA, some targeting sgRNAs caused hundreds to thousands of genes to be differentially expressed in the Perturb-seq measurements, consistent with the earlier observation that -25% of targeting sgRNAs had a significant impact on the transcriptional state. The protein and abundant RNA channels that were significantly impacted by each perturbation in the imaging measurements were also identified. Most active targeting perturbations impacted a relatively small subset of imaged targets, but some perturbations such as the essential nuclear genes Sf3b6, Sbnol , Polrla, and Kin had a significant impact on many channels. These results reflected the pleiotropic consequences of disrupted transcription and RNA processing.

[0278] EXAMPLE 7

[0279] Multimodal in vivo screening with strong phenotypes

[0280] The imaging and sequencing experiments resulted in >100,000 total perturbed cells with either 18-plex imaging phenotypes (subcellular morphology and cell locations in tissue), or genome- wide transcriptional phenotypes. The spatial organization of perturbed cells in tissue was visualized and small clusters of cells sharing the same perturbation were observed, distributed throughout the various zones that form liver tissue (FIG. 10A); these clusters presumably reflected local proliferation of single infected hepatocytes during postnatal growth after lentiviral transduction. The Perturb-seq transcriptional phenotypes with dimensionality reduction were visualized. For example, a UMAP projection of transcriptional phenotypes clearly separated cells with control sgRNAs from cells with sgRNAs targeting the core hepatocyte transcription factor Hnf4a (FIG. 10B). The Hnf4a knockout was associated with lower expression of Apolipoprotein Al, consistent with the role of Hnf4a in promoting the expression of abundant secreted factors in hepatocytes. Unbiased clustering of the multimodal genotype-phenotype map was performed to identify perturbations that caused similar multimodal phenotypes measured by Perturb-seq and multiplexed imaging (FIG. s and 10D). The perturbation cluster map derived from joint analysis of both modalities recovered many interpretable blocks of perturbed gene function, including a large block of ribosomal proteins and ribosome biogenesis genes, a block of nuclear mRNA processing genes, and a block of genes whose disruption activates the integrated stress response (FIG. 10D). Broadly, the sequencing and imaging data identified related but complementary perturbation-perturbation relationships (Pearson’s R = 0.42 for correlation-of-correlations) .

[0281] The mRNA and protein phenotypes that co-varied across perturbations were identified. The mRNA co-expression data from sequencing was observed by constructing a minimum distortion embedding that placed genes with correlated expression nearby (FIG. 10E). Clusters of co-regulated genes with similar biological functions, such as a group of genes involved in the urea cycle and the transport and metabolism of basic amino acids were identified. Imaging stains whose intensity co-varied across perturbations (FIG. 10F), such as the mitochondria proteins Tomm20 and Tomm70, the lysosomal protein mannose 6- phosphate receptors (M6PRs), and the ER chaperone calreticulin were also identified.

[0282] EXAMPLE 8

[0283] Identifying genetic regulators of liver physiology

[0284] Targeted analyses were conducted to discover key regulators of diverse aspects of in vivo liver physiology. With the Perturb-seq data, the drivers of gene expression modules with crucial functions in hepatocytes were identified. For example, genetic perturbations were ranked by their impact on a core program of lipid biosynthesis genes defined by a previous large-scale Perturb-seq study (FIG. 11A). Knockout of the lipid biosynthesis inhibitor Insigl had the strongest positive effect on lipid biosynthesis gene expression; knockout of the Ldl receptor (Ldl ) and Srebfl also had strong positive impacts, possibly due to compensatory expression.

[0285] Perturbations that activated transcriptional stress response pathways such as the integrated stress response (ISR) or the unfolded protein response (UPR), in hepatocytes in an intact liver were identified (FIG. 11B). These two pathways were activated by the knockouts of separate sets of genes, suggesting that the stress- specific and selective transcriptional responses that were observed in cell culture were also a feature of hepatocyte physiology. For example, knockout of aa-tRNA synthetase genes such as Nars and Aars activated the ISR, as was previously observed in a genome- wide Perturb-seq experiment in cultured cells. Knockout of translation initiation factors such as Eif2a (Eif2sl ) and Eif2b (Eif2b4) also activated the ISR. These results were consistent with established mechanisms of ISR activation by uncharged tRNAs and delayed translation initiation. Knockout of ER import and quality control genes such as Sec61al, Dnajb9, and Selll activated the unfolded protein response. Knockout of the UPR effectors Xbpl and Ire I (Eml) caused a decrease in UPR gene expression, suggesting some basal level of UPR activation in hepatocytes in ad libitum mice. Further, knockout of the ER calcium transport gene Atp2a2 caused strong activation of both the UPR and the ISR, suggesting that calcium homeostasis played a central role in regulating both ER function and broader cellular stress responses in hepatocytes.

[0286] For first-pass exploration of the imaging data, genetic perturbations were ranked by their impact on the intensity of each immunofluorescence or RNA FISH stain and perturbations that had a significant impact on the stain intensity were identified. As a validation, it was observed that sgRNAs targeting Gapdh had the strongest negative impact on anti-Gapdh staining intensity (FIG. 11C) and sgRNAs targeting Albumin had the strongest negative impact on Albumin mRNA FISH staining intensity (FIG. 11D). Targeting the RNA polymerase II machinery (Polr2l, Gpnl) or the core hepatocyte transcription factor Hnfla, which directly binds the Albumin promoter, also had a strong negative impact on Albumin mRNA FISH, as did sgRNAs against the ER associated degradation (ERAD) factor Selll.

[0287] The screens correctly identified genetic regulators of many highly essential processes, suggesting that the ten days that passed between AAV-Cre infection and tissue perfusion captured the acute impacts of perturbations. For example, targeting the RNA polymerase I machinery (Pair I ) and Pol I transcription factors (Tafia, Rrn3, Ublf) had the strongest negative impact on pre-rRNA FISH intensity (FIG. HE). Unexpectedly, targeting the nuclear speckle component Sf3b6 caused an increase in pre-rRNA FISH intensity, suggesting a possible interplay between subnuclear structures (e.g. nuclear speckle and nucleolus) in hepatocytes in vivo.

[0288] The screens also identified the in vivo regulators of cellular signaling pathways that play important roles in liver physiology. Phosphorylation of ribosomal protein S6 is a core function of the mTor signaling pathway. Knockout of Mtor and the potential mTor complex co-chaperone Cdc37 caused the strongest decrease in anti-phospho S6 ribosomal protein intensity (FIG. HF); knockout of the mTor inhibitors Tscl and Tsc2 and the growth inhibitor Pten caused the strongest increase in anti-phospho S6 ribosomal protein intensity. The imaging-based screening was uniquely powerful for probing genetic regulators of protein modification, which could not be measured by sequencing. EXAMPLE 9

[0289] Deep learning on the imaging-based perturbation data

[0290] In addition to modification of signaling proteins, the imaging was also uniquely capable of detecting morphological changes beyond simple assessment of protein expression levels. Deep-learning-based approaches were used to explore more complex morphological phenotypes in the perturbation-imaging data. An autoencoder model was used to generate a lower-dimensional embedding for each image and then unbiased clustering was performed on the embeddings. These embeddings and the clustering with UMAP were illustrated, using the imaging data of the lysosomal protein CathB as an example (FIG. 11G). The perturbations that shifted the representation of cells between the different clusters were then identified. For example, knockout of the lysosomal cholesterol transport gene Npcl caused enrichment of CathB morphologies in clusters F and G, which had many distinct anti-CathB punctae, and depletion from clusters A and E, which had very few anti-CathB punctae (FIGs. 11H and 111). Npcl knockout also caused a significant increase in anti-CathB intensity (mean increase = 0.53 AUs, FDR-corrected p < 10'34). This was consistent with the accumulation of dysfunctional lysosomes that occurred when lysosomal cholesterol export was blocked, such as in Niemann-Pick disease type C with NPC1 or NPC2 mutations in humans.

[0291] EXAMPLE 10

[0292] Investigating the impact of perturbations in altered physiological states

[0293] A key ability of performing screens in vivo rather than in vitro is to identify perturbations that are sensitive to changes in organismal physiology and homeostasis. A mouse was fasted from the lenti-sgRNA infected cohort for 16 hours prior to perfusion and tissue preparation and a parallel Perturb-seq experiment was conducted on both the fasted animal and ad lib animal (FIG. 20J). An energy distance metric was used to quantify the impact of each gene knockout on the distribution of global transcriptional states. The impact of each knockout was measured relative to cells with control sgRNAs in the same animal. The overall magnitude of transcriptional phenotypic changes caused by each gene knockout was similar between fasted and ad libitum mice for most genes (FIG. 11K, Pearson’s R = 0.77). However, some genes involved in autophagy and the lysosome Atp6v0c, Atp6apl, Lamlor2) and the broader endomembrane system Dnm2, Arfrpl, Jtb, ZwIO) had especially strong phenotypes in fasted animals in comparison to ad libitum animals. Consistent with the energy distance analysis, knockout of these genes caused many more genes to be differentially expressed in fasted mice than in ad libitum mice, with Dnm2 knockout shown as an example (FIG. 11L). This may have been due to a higher dependence on lysosomal function during fasting; the failure of autophagy, and / or endocytic function during fasting may cause severe cellular distress.

[0294] The transcriptional phenotypes associated with knockout of these lysosome and endomembrane genes were examined in ad libitum and fasted mice. The transcriptional responses associated with these knockouts were similar in a fasted animal (mean Pearson’s R = 0.91, versus 0.19 in an ad libitum animal; Fig. 20M), suggesting that these knockouts caused a convergent phenotype in some physiological states. More broadly this illustrated how the modular structure of in vivo genotype-phenotype maps can be re-wired by environmental conditions.

[0295] EXAMPLE 11

[0296] Three case studies exploring liver physiology with multimodal screening

[0297] Hepatocyte Zonation

[0298] Zonated gene expression in hepatocytes contributes to the spatial division of liver function. This zonation is believed to be maintained by gradients of morphogens, oxygen, nutrients, and hormones, but the requirements for these different processes to maintain zonal gene expression in a hepatocyte in an adult mouse are unclear. The multimodal screening was used to shed light on mechanisms underlying hepatocyte zonal identity maintenance and dynamics.

[0299] Wnt signaling is an established driver of hepatocyte zonation. Wnt ligand secretion from central vein endothelial cells promotes pericentral gene expression in hepatocytes; low Wnt in the periportal region contributes to the periportal gene expression program. All cells in the Perturb-seq dataset were scored according to zonated gene expression and the distribution of zonal expression in cells with knockouts of central Wnt effectors and inhibitors was assessed (FIG. 8A; see EXAMPLE 1, Tissue zone segmentation). Knockout of the Wnt transducer Ctnnbl (P-catenin) caused a major decrease in the fraction of cells with pericentral gene expression, consistent with the role of Wnt signaling in pericentral expression. Conversely, knockout of the Wnt inhibitor Ape caused a decrease in the population of hepatocytes with periportal gene expression.

[0300] All gene knockouts in the experiment were then assessed to identify the other drivers of zonal gene expression (FIG. 8B). The periportal and pericentral expression were also assessed as two separate scores, since the knockout of some core essential genes caused a general loss of zonal identity rather than a shift in one direction or the other (FIG. 8C). Ape knockout caused the strongest increase in pericentral expression, as well as a substantial decrease in periportal expression. Knockout of Vhl, which mediated the degradation of hypoxia-inducible factors (HIFs) under normoxic conditions, caused a major decrease in periportal expression and also increased pericentral expression, consistent with the idea that oxygen tension directly regulates zonal gene expression. Loss of Vhl activity in the oxygenrich periportal region may cause the specific decrease in periportal expression. Knockout of Pten and the gene Zfp830, which was previously studied in the context of hepatocyte zonation, also caused a shift toward a more pericentral-like expression state. Knockout of the R-spondin receptor Lgr4 caused a strong increase in periportal gene expression, phenocopying Ctnnbl. The knockout of protein kinase, a subunit Prkarla, also caused a substantial increase in periportal gene expression. Knockout of the heparan sulfate synthesis gene Hs6stl and the proteoglycan synthesis enzyme B4galt7 also caused an increase in periportal expression and a decrease in pericentral expression, possibly by impacting the extracellular matrix in a manner that influences the interactions of signaling molecules such as Wnt ligands with their receptors.

[0301] Although one- and two-dimensional zonal expression scores described above were useful for exploration and interpretation of knockout effects on zonation, they did not necessarily reflect the full complexity of transcriptional phenotypes. Interestingly, when comparing the transcriptome- wide phenotype caused by these zonation-modulating knockouts, it was found that the knockout of some pro-pericentral factors, such as Hs6stl and B4galt7, closely phenocopied Lgr4 and Ctnnbl knockout (FIG. 8D), suggesting that their functions in hepatocytes are intertwined with Wnt signaling. However, knockout of the pro- pericentral factor Prkarla caused a globally distinct transcriptional response despite its convergent impact on zonation-associated gene expression. Knockout of the apparent pro- periportal factors Ape, Vhl, Pten, and Zfp830 all caused globally distinct transcriptional phenotypes despite their convergent impacts on zonal gene expression. These results emphasized the complexity of the relationship between cell state and metabolism in hepatocytes.

[0302] The initial lentiviral transduction appeared to target cells in all zones of the lobule; cells with control sgRNAs exhibited a full range of zonated gene expression (FIG. 8A). In Perturb-seq, the endpoint distribution of gene expression was measured, but what happened during the time course of the perturbation or where each cell was in the tissue, prior to dissociation was not apparent from Perturb-seq. The shift in endpoint zonal gene expression that was observed may have been due to (1) in situ trans-differentiation without cell migration to an expression-appropriate zone, (2) trans-differentiation with migration, or (3) the death of cells in one zone and / or the proliferation of cells in the other zone. To distinguish between these scenarios, the perturbation-imaging data was examined. The tissue was automatically segmented into two spatial zones (pericentral and periportal segmentations) and assessed to see whether the knockouts that impacted zonal gene expression also impacted the positions of cells with regard to the segmented pericentral and periportal zones (FIG. 8E and 8F). The fraction of cells physically located in the pericentral and periportal zones did not change with any of the gene knockouts that significantly impacted zonal expression, even with potent knockouts such as Ape and Ctnnbl (FIG. 8G). These results suggested that, on the time scale of the experiment and with the statistical power, cells with zonation-impacting knockouts exhibited gene expression changes due to trans-differentiation without migration or changes to survival or proliferation.

[0303] EXAMPLE 12

[0304] Stress Response

[0305] One major function of hepatocytes is to secrete abundant plasma proteins, so special attention was directed to perturbations that impacted the endoplasmic reticulum. It was observed that the knockout of the gene Selll, an ER-associated degradation (ERAD) cofactor that plays a crucial role in ER protein quality control, caused a strong increase in Calreticulin protein level (FIG. 12A and 12C). This suggested that the transcriptional upregulation of UPR targets that were observed upon Selll knockout in the Perturb-seq experiment (FIG. 11B) led to a corresponding increase in the ER protein chaperone capacity. Selll knockout also caused a decrease in Albumin mRNA level but had relatively small impacts on the other proteins and abundant RNAs that were imaged (FIG. 12B and 12C).

[0306] Leveraging the paired imaging and sequencing data, the genome- wide transcriptional changes associated with Selll knockout in hepatocytes was assessed. The Selll knockout caused the upregulation of many other ER chaperone mRNAs, such as Hspa5 and Dnajb9, as well as Calreticulin itself (FIG. 12D). Selll knockout decreased Albumin mRNA by 50% (Fig. 22E), as observed from our imaging data, but it also caused the downregulation of other mRNAs that encoded for a range of abundant secreted plasma proteins, such as Apolipoprotein Al (Apoal; 38% decrease) and the GC Vitamin D Binding Protein (Gc, 36% decrease) (FIG. 12E). Altogether, the mRNAs downregulated by Selll knockout reflected a -40% decrease in the total quantity of abundant secreted protein mRNAs (FIG. 12F). These results suggested that a major function of the unfolded protein response in hepatocytes in vivo was to decrease the burden of secreted protein mRNAs and therefore of nascent proteins on the ER translocon and folding machineries. This decrease in mRNA abundance could occur indirectly through transcriptional downregulation or directly through post-transcriptional stress response pathways such as Regulated Irela-Dependent Decay (RIDD), where it was found that ER stress sensor Irel can directly cleave mRNAs of secreted proteins on the ER surface to lower the flux of proteins through the ER.

[0307] EXAMPLE 13

[0308] Lipid. Accumulation

[0309] Hepatic steatosis, characterized by the accumulation of lipids within hepatocytes, is linked to metabolic syndrome and can progress to serious diseases such as metabolic dysfunction-associated steatohepatitis (MASH) and fibrosis. The imaging screen was used to identify in vivo drivers of steatosis. Acute knockout of the lipid biosynthesis inhibitor Insig 1, the growth inhibitor Pten, the integrated stress response inhibitor Eif2a (Eif2sl and the alanine-tRNA synthetase Aars all caused steatosis, as measured by enlargement and increase in signal from lipid droplets, marked by anti-Perilipin (FIG. 12G and 12H).

[0310] The Perturb-seq data from the same mice was then assessed to identify the transcriptome-wide changes associated with each of these four perturbations (FIG. 121). Consistent with its function as an inhibitor of lipid biosynthesis gene transcription, Insigl knockout caused the upregulation of mRNAs for cholesterol biosynthesis genes such as Hmgcr and Fdps and fatty acid biosynthesis genes such as Acaca and Fasn. On the other hand, knockout of Aars and Eif2sl drove upregulation of ISR target genes such as Ddit4 and Atf5. Intriguingly, Aars and Eif2sl knockout caused a decrease in most lipid biosynthesis mRNAs. Finally, knockout of Pten caused yet another distinct transcriptional response that could not be summarized either as the ISR or lipid biosynthesis.

[0311] Through integrated analysis of Perturb-seq and imaging, it was determined that knockouts of these four genes drove three separate transcriptional responses but converged on the steatotic phenotype. Considering these results, it was possible that the knockouts caused steatosis through three separate mechanisms — (1) through the transcriptional activation of lipid biosynthesis genes in the case of Insigl knockout; (2) through the apparent sequestration of free lipids into lipid droplets in the case of Eif2sHAars knockout, possibly as a component of the ISR that may serve to protect the cytoplasm and other organelles in stressed cells from biophysically disruptive lipid molecules, and (3) a distinct P / cn-associatcd mechanism, possibly including uptake from plasma, some synthesis increase, sequestration, and / or the repurposing of organellar stores (FIG. 12J).

[0312] EXAMPLE 14

[0313] Provided below are tables with Examples from an RCA MERFISH mRNA library (Table 2), a Liver Perturbation Library (Table 6), and an RCA MERFISH Perturbation Barcode Library (Table 7). Table 2

[0314] Table 6

[0315] Table 7

[0316] Also provided below are Readout Bits (Table 1), a Stain Panel (Table 3), Examples of

[0317] Abundant RNA Probes (Table 4), RCA MERFISH Imaging Experiment Rounds WT Animal (Table 5), RCA MERFISH Imaging Experiment Rounds (Table 8) and Examples of Perturb

[0318] Seq sgRNA Probes (Table 9).

[0319] Table 1 Table 3

[0320] Table 4 Table 5

[0321] Table 8

[0322] Table 9

[0323] Probe Pool ID Probe Sequence sgRNA Constant Region Probe Pool gtgactggagttcagacgtgtgctcttccgatctactttaggGCTATGCTGTTTCCAGCTTAGCTCT BC001 (SEQ ID NO: 80) sgRNA Constant Region Probe Pool gtgactggagttcagacgtgtgctcttccgatctactttaggNGCTATGCTGTTTCCAGCTTAGCTCT BC001 (SEQ ID NO: 81) sgRNA Constant Region Probe Pool gtgactggagttcagacgtgtgctcttccgatctactttaggNNGCTATGCTGTTTCCAGCTTAGCTCT BC001 (SEQ ID NO: 82) sgRNA Constant Region Probe Pool gtgactggagttcagacgtgtgctcttccgatctactttaggNNNGCTATGCTGTTTCCAGCTTAGCTCT BC001 (SEQ ID NO: 83) sgRNA Constant Region Probe Pool gtgactggagttcagacgtgtgctcttccgatctNactttaggGCTATGCTGTTTCCAGCTTAGCTCT BC001 (SEQ ID NO: 84)

[0324] EXAMPLE 15 DISCUSSION

[0325] Massively parallel in vivo genetic screens provide a powerful approach to dissect regulators of cellular and tissue physiology in their native context. As described herein, through extensive technical development and integration of multiple single cell profiling methods, a general and scalable approach for multimodal, pooled genetic screens with Perturb-seq- and multiplexed-imaging-based phenotyping was established. This multimodal in vivo approach provided access and interrogated physiological processes that were previously difficult to study by other means. Screens in hepatocytes in the mouse liver were conducted and these screens demonstrated how the resulting data allowed causal dissection of multifarious aspects of liver biology, from hepatocyte zonation to steatosis to the signaling of nutritional status. The data also provided a rich resource for future studies of liver biology and development of computational methods. In particular, the multimodal imaging data proved useful for developing new analytical and machine learning tools for dimensionality reduction, feature extraction, and cross-modal analysis. More generally, this approach allowed the creation of large-scale training datasets for machine learning and Al efforts to build models of “virtual” cells, driving both new discoveries and new forms of understanding.

[0326] The combination of transcriptomic profiling and subcellular morphology imaging allowed new analyses of cellular heterogeneity within tissue, under both homeostatic and perturbed physiological conditions. Tissue-scale spatial analysis of protein feature embeddings revealed zonal heterogeneity in several imaging targets, beyond the expected Albumin mRNA and Perilipin. Cross modal analysis of subcellular protein and abundant RNA morphology across transcriptionally-defined hepatocyte subtypes revealed a continuum of morphological types between different liver zones. More generally, the multimodal combination of spatial transcriptomics to define cell types and states with morphological features that linked to the extensive histopathology literature allowed studies of both healthy and diseased tissue.

[0327] The screening results demonstrated the value of multimodal phenotyping that combined both multiplexed protein and RNA imaging and transcriptome- wide mRNA sequencing to study genetic regulators of cellular state in vivo. Applying this approach to cells in native tissue provided insights into three core aspects of liver physiology in a single experiment: hepatocyte zonation, dynamic stress responses, and lipid droplet accumulation.

[0328] Zonation: Hepatocytes in different zones of the liver express different genes and have different functions. The regulation of hepatocyte zonation has been challenging to study in vitro, as hepatocytes rapidly lose their zonal identity when isolated and introduced into cell culture. Even complex organ-on-a-chip and organoid models provided imperfect recapitulations of tissue architecture. As described herein, the regulation of zonal gene expression in the mouse liver was directly probed. In addition to confirming the known roles of oxygen sensing and Wnt signaling in regulating zonal identity, it was determined that even brief (<10 day) disruption of these pathways rewired hepatocyte zonal gene expression in adult mice, apparently in a cell autonomous manner. Several other novel candidate regulators of zonation such as Hs6stl and B4galt7 were identified that had an equal impact on zonation relative to perturbing the Wnt pathway, which illustrated the value of systematic exploration. These factors were involved in the modification of secreted and extracellular matrix proteins and pointed toward a unique capacity of in vivo screens to explore causal relationships between the extracellular microenvironment and intercellular signaling. By analyzing the spatial location of the genetically perturbed cells, it was determined that upon deletion of these key regulatory genes, the cells changed their transcriptional zonal states without physical relocation across pericentral and periportal zones in the tissue. The methods and materials described herein were used with genome-scale perturbation libraries and multiplex perturbation vectors to comprehensively classify zonation regulators, and to identify epistatic relationships between the associated signaling pathways.

[0329] Stress response: Hepatocytes are secretory cells that must produce Albumin, among other proteins, in massive quantities. As such, stress response pathways that monitor protein folding in the endoplasmic reticulum, such as unfolded protein response (UPR), play an important role in maintaining hepatocyte function. Notably the UPR is activated in several liver diseases, including viral hepatitis, alcohol-associated liver disease, and MASH. The secretion of proteins such as Albumin drops dramatically when the liver is dissociated and hepatocytes are introduced into an in vitro setting, indicating the importance of in vivo models for the investigation of hepatocyte secretion. In the imaging -based pooled genetic screening described herein, it was determined that the knockout of Sei 11, an important ubiquitin ligase adaptor protein involved in ER-associated degradation, caused massive upregulation of the ER quality control protein Calreticulin, and downregulation of Albumin RNA. In subsequent analysis of the Perturb-seq data, it was found that Selll knockout caused both transcriptional upregulation of UPR target genes and the 40% transcriptional downregulation of abundant secreted proteins. These results indicated that a major function of the UPR in hepatocytes in the liver is to downregulate secretory protein genes, thereby reducing the burden on the ER folding machinery, either transcriptionally or through post- transcriptional mechanisms such as RIDD.

[0330] Lipid droplets: Lipid droplets play a crucial role in hepatocyte energy storage and lipid metabolism. Accumulation of lipid droplets is key to the pathology of metabolic dysfunction-associated steatotic liver disease (MASLD), whereby progress from steatosis to steatohepatitis is marked by distinct transcriptional states in humans. As an example of interrogating in vivo physiology, it was determined that genetic knockouts targeting Insigl, Eif2sHAars and Pten all caused a convergent morphological phenotype with the dramatic accumulation of lipid droplets, but entirely distinct transcriptional responses. While the highly interpretable imaging results indicated that the function of these genes impinged on lipid homeostasis, their distinct effects on the transcriptional states indicated they likely induced lipid accumulation through distinct mechanisms. The paired imaging and sequencing approach was necessary to reach this conclusion. Imaging lipid droplet accumulation on its own would reveal the ultimate cellular consequences of the perturbations; measuring the transcriptional responses by Perturb-seq would show the complexity and distinctiveness of each perturbation, but mask their common functional impact on the cell. The experiments described herein demonstrated the power of multimodal functional interrogation of diverse genes in vivo at the scale of hundreds of genes. The described approach is highly scalable in terms of the number of phenotypes, genotypes, and perturbed cells measured. In terms of phenotype, the number of genes imaged using MERFISH, as well as the number of proteins assayed through multiplexed immunofluorescence, can be increased to obtain a more unbiased picture of the phenotypic states of cells. In terms of genotype, the number of perturbations can be increased, ultimately to the genome- wide scale, enabling truly unbiased and comprehensive mapping of gene function.

[0331] While several embodiments of the present disclosure have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the present disclosure. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the teachings of the present disclosure is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the disclosure described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, the disclosure may be practiced otherwise than as specifically described and claimed. The present disclosure is directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.

[0332] In cases where the present specification and a document incorporated by reference include conflicting and / or inconsistent disclosure, the present specification shall control. If two or more documents incorporated by reference include conflicting and / or inconsistent disclosure with respect to each other, then the document having the later effective date shall control. All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0333] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”

[0334] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0335] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e. “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.”

[0336] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.

[0337] When the word “about” is used herein in reference to a number, it should be understood that still another embodiment of the disclosure includes that number not modified by the presence of the word “about.”

[0338] It should also be understood that, unless clearly indicated to the contrary, in any methods claimed herein that include more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited.

[0339] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.

Claims

CLAIMSWhat is claimed is:

1. A method, comprising: exposing a sample comprising a plurality of cells to oligonucleotide- conjugated antibodies and nucleic acid probes; fixing the sample in a gel; and determining spatial positions of the oligonucleotides and the nucleic acid probes within the plurality cells of the sample the gel.

2. The method of claim 1, further comprising: exposing the sample to adapter oligonucleotides, wherein the adapter oligonucleotides hybridize with the oligonucleotides of the conjugated antibodies and nucleic acid probes present in the sample.

3. The method of claim 1 or 2, wherein the oligonucleotides, nucleic acid probes, and adapter oligonucleotides are modified with an agent for attaching the oligonucleotides, nucleic acid probes, and adapter oligonucleotides to the gel.

4. A method, comprising: exposing a sample to oligonucleo tide-conjugated antibodies; exposing RNA and the oligonucleotides of the oligonucleo tide-conjugated antibodies in the sample to an agent for attaching RNA and the oligonucleotides to a gel; fixing the sample in a gel; and determining spatial positions of the RNA and the oligonucleotides within the gel.

5. A method, comprising: exposing a sample to oligonucleo tide-conjugated antibodies; fixing the sample in a gel; and determining spatial positions of nucleic acids within the gel, wherein the nucleic acids include the oligonucleotides in the conjugates.13416186.

26. The method of claim 5, wherein the nucleic acids are modified with an agent for attaching the oligonucleotide to a gel.

7. The method of claim 6, wherein the agent for attaching the oligonucleotide to the gel is:

8. The method of any one of claims 1-7, wherein the gel is an acrylamide gel.

9. The method of any one of claims 1-8, further comprising clearing proteins and / or lipids from the sample.

10. The method of any one of claims 1-9, further comprising retrieving antigens in the sample prior to exposing the sample to oligonucleotide-conjugated antibodies.

11. The method of claim 4 or 5, further comprising exposing the sample to a plurality of nucleic acid probes.

12. The method of any one of claims 1-11, wherein the plurality of nucleic acid probes comprises nucleic acid probes with different sequences.

13. The method of any one of claims 1-12, wherein the plurality of nucleic acid probes target RNA species and oligonucleotides of oligonucleotide-conjugated antibodies in the sample.

14. A method, comprising: exposing nucleic acids in a sample to a padlock probe comprising a barcode flanked by a 5’ homology arm and a 3’ homology arm, wherein the 5’ and 3’ homology arms bind to the nucleic acid; circularizing the padlock probe;13416186.2generating an amplicon by rolling circle amplification using the circularized padlock probe as a template, wherein the amplicon comprises a plurality of barcode complements; and determining the spatial positions of the plurality of barcode complements using MERFISH.

15. The method of any one of claims 1-13, further comprising: exposing nucleic acids in a sample to a padlock probe comprising a barcode flanked by a 5’ homology arm and a 3’ homology arm, wherein the 5’ and 3’ homology arms bind to the nucleic acid; circularizing the padlock probe; generating an amplicon by rolling circle amplification using the circularized padlock probe as a template, wherein the amplicon comprises a plurality of barcode complements; and determining the spatial positions of the plurality of barcode complements using MERFISH.

16. A method, comprising: exposing a sample to oligonucleo tide-conjugated antibodies; fixing the sample in a gel; and determining spatial positions of the oligonucleotides and nucleic acids within the sample.

17. A method, comprising: introducing, into a sample comprising a plurality of cells, a nucleic acid that perturbs expression of at least one gene in the plurality of cells; exposing the sample to oligonucleo tide-conjugated antibodies and a plurality of nucleic acid probes; fixing the sample in a gel; and determining spatial positions of the oligonucleotides and the nucleic acid probes within the plurality of cells in the sample.

18. The method of claim 17, further comprising:13416186.2exposing the sample to adapter oligonucleotides, wherein the adapter oligonucleotides hybridize with the oligonucleotides of the conjugated antibodies and the nucleic acid probes present in the sample.

19. The method of claim 17 or 18, wherein the nucleic acid that perturbs expression of at least one gene comprises a guide portion, a reporter portion, and an identification portion.

20. The method of claim 17 or 18, wherein the nucleic acid that perturbs expression of at least one gene comprises an open reading frame (ORF) of the at least one gene.

21. The method of any one of claims 1-20, wherein the spatial positions of the oligonucleotides and nucleic acid probes are determined in neighboring cells of the plurality of cells of the sample.

22. A method comprising: providing a library of enzymatically-synthesized split probes that target all mRNAs in a cell.

23. A method comprising: providing a library of enzymatically-synthesized split probes that target a subset of mRNAs present in a cell.

24. A method comprising: providing a library of enzymatically-synthesized split probes that target sgRNAs in a cell.

25. A method comprising: providing a library of enzymatically-synthesized split probes that target barcodes present in a cell.

26. The method of any one of claims 22-25, further comprising amplifying by PCR oligonucleotides of a pooled library of oligonucleotides.13416186.

227. The method of claim 26, further comprising adding by PCR a probe barcode to the amplified oligonucleotides.

28. The method of claim 27, further comprising introducing by PCR a T7 promoter to the amplified oligonucleotides.

29. The method of claim 28, further comprising T7-transcribing the amplified oligonucleotides.

30. The method of claim 29, further comprising reverse transcribing the amplified, T7- transcribed oligonucleotides to generate ssDNA.

31. The method of claim 30, further comprising rolling-circle amplifying and circularizing the amplified, T7-transcribed oligonucleotides.

32. The method of any one of claims 27-29, further comprising cloning the amplified, oligonucleotides into a phage / phagemid genome.

33. The method of claim 32, further comprising producing ssDNA from the phage / phagemid.

34. The method of claim 33, further comprising digesting the ssDNA by restriction enzymes.

35. The method of claim 34, further comprising hybridizing the ssDNA to oligonucleotide probes to create cut sites for the restriction enzymes.

36. The method of claim 33, further comprising digesting the ssDNA with DNAzyme.

37. The method of claim 36, further comprising activating the DNAzyme with a metal ion.13416186.2

Citation Information

Patent Citations

  • Methods for making transcription products

    US20120156679A1

  • Novel constructs and screening methods

    US20180187184A1

  • Methods for Attaching Cellular Constituents to a Matrix

    US20190127786A1

  • Multiplexed imaging using merfish, expansion microscopy, and related technologies

    US20190276881A1

  • Nucleic acid probes

    US20230083623A1