Barcoding of nuclei for multiplex screening of cells

Vectors encoding fusion proteins with nuclear localization and RNA-binding domains facilitate barcoding of cell nuclei, enhancing high-throughput transcriptome profiling and genetic screening by combining snRNA-seq with scRNA-seq, addressing the challenges of cell damage and stress-induced transcriptome alterations.

WO2026035851A1PCT designated stage Publication Date: 2026-02-12RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/040914
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-08
Filing Date
2025-08-06
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing methods for high-throughput transcriptome profiling and genetic screening of cells, particularly for tissues and cell types like neurons, kidney, heart, liver, and adipocytes, face challenges such as cell damage during dissociation and stress-induced transcriptome alterations, making single nucleus RNA sequencing (snRNA-seq) a preferable approach.

Method used

Vectors are developed that encode fusion proteins with nuclear localization and RNA-binding domains, generating barcoded RNA transcripts that are translocated to the nucleus, allowing for barcoding of cell nuclei, which can be used in combination with scRNA-seq for comprehensive transcriptome profiling.

Benefits of technology

The solution enables improved barcoding of cell nuclei for high-throughput transcriptome profiling and genetic screening, providing more comprehensive cell type identification and reducing transcriptome perturbation.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Vectors, cell lines comprising the vectors, recombinant virions produced from the vectors, and methods of using the vectors for barcoding nuclei of cells are provided. Vector libraries that collectively encode a plurality of genetic elements of interest for screening and methods of screening cells having nuclei barcoded according to the subject methods are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

BARCODING OF NUCLEI FOR MULTIPLEX SCREENING OF CELLSCROSS REFERENCE TO APPLICATIONS

[0001] Pursuant to 35 U.S.C. § 119(e), this application claims priority to the filing date of U.S. Provisional Application No. 63 / 681,047, filed August 8, 2024, the disclosure of which is incorporated herein by reference.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This invention was made with government support under MH130700, and MH121268 awarded by the National Institutes of Health. The government has certain rights in the invention.INCORPORATION BY REFERENCE OF A SEQUENCE LISTING

[0003] A Sequence Listing is provided herewith as a Sequence Listing XML file, “UCSF- 806PRV”, created on July 25, 2024, and having a size of 5,241 bytes. The contents of the Sequence Listing XML file are incorporated by reference herein in their entirety.BACKGROUND OF THE INVENTION

[0004] Single-cell RNA sequencing (scRNA-seq) allows high-throughput transcriptome profiling and investigation of cellular heterogeneity (Kolodziejczyk et al. (2015) Mol. Cell. 58:610-620). However, some cells such as neurons in the brain are easily damaged during tissue dissociation, which is required for scRNA-seq, and some tissues and cell types are difficult to dissociate including the kidney, heart, liver, adipocytes, and myofibers (Lake et al. (2016) Science 352:1586-1590; Andrews et al. (2022) Hepatol. Commun. 6:821-40, Petrany et al. (2020) Nat. Commun. 11 :6374, Sun et al. (2020) Nature 587:98-102, Tucker et al. (2020) Circulation 142:466- 482, Wu et al. (2019) J. Am. Soc. Nephrol. 30:23-32). For such tissues and cell types, single nucleus RNA sequencing (snRNA-seq) may be a preferable approach. In addition, enzymatic dissociation, which is often used for single cell scRNA-seq may induce a stress response that alters the transcriptome (Wu et al., supra Denisenko et al. (2020) Genome Biol. 21:130), whereas snRNA-seq has the advantage that it does not perturb the transcriptome to the same extent.Therefore, the use of snRNA-seq data may improve transcriptome profiling of cells and tissues that arc less amenable to scRNA-scq, and the combination of snRNA-scq with scRNA scq should provide more comprehensive transcriptome profiling for cell type identification than either technique alone.

[0005] There remains a need in the art for improved methods of barcoding cell nuclei for high- throughput transcriptome profiling and genetic screening of cells.SUMMARY OF THE INVENTION

[0006] Vectors, cell lines comprising the vectors, recombinant virions produced from the vectors, and methods of using the vectors for barcoding nuclei of cells are provided. Vector libraries that collectively encode a plurality of genetic elements of interest for screening and methods of screening cells having nuclei barcoded according to the subject methods are also provided.

[0007] In one aspect, a vector is provided, the vector comprising: a first expression cassette comprising a first promoter operably linked to a first nucleotide sequence encoding a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain, wherein expression of the first nucleotide sequence results in production of the fusion protein in a cell; and a second expression cassette comprising a second promoter operably linked to a second nucleotide sequence encoding an RNA recognition sequence for the RNA binding domain and a barcode, wherein transcription of the second nucleotide sequence generates a barcoded RNA transcript comprising the RNA recognition sequence and the barcode, wherein binding of the RNA binding domain to the RNA recognition sequence results in formation of a complex between the fusion protein and the barcoded RNA transcript, wherein the complex is translocated to the nucleus of the cell.

[0008] In certain embodiments, the vector is a viral vector. In some embodiments, the viral vector is an adenovirus associated virus (AAV) vector, a lentivirus vector, or a G-deleted rabies vims (RVdG) vector. In embodiments in which the viral vector is an adenovirus associated virus (AAV) vector, the AAV vector further comprises a 5 ’-inverted terminal repeat (ITR) and a 3’- ITR, wherein the first expression cassette and the second expression cassette are positioned between the 5 ’-ITR and the 3 ’-ITR.

[0009] In certain embodiments, the barcode is located in a 3’ untranslated region (3’ UTR) of the barcoded RNA transcript. In some embodiments, the barcode is transcribed by a separate U6 promoter and DNA polymerase III.

[0010] In certain embodiments, the nuclear localization domain comprises a nuclear localization sequence or a nuclear-localized protein. In some embodiments, the nuclear localization sequence is a histone H2B nuclear localization sequence. In some embodiments, the nuclear localization domain comprises a Klarsicht-ANC-l-Syne homology (KASH) domain. In some embodiments, the nuclear- localized protein is a histone H2B protein or a Klarsicht ANC-1 Syne Homology (KASH) domain-containing protein.

[0011] In some embodiments, the RNA-binding domain comprises a plurality of RNA binding proteins.

[0012] In some embodiments, the RNA recognition sequence comprises a plurality of binding sites for the RNA-binding protein.

[0013] In certain embodiments, the RNA binding protein comprises an MS2 coat protein (MCP) and the RNA recognition sequence comprises an MS2 RNA aptamer, wherein the MCP binds to the MS2 RNA aptamer. In some embodiments, the RNA recognition sequence comprises a plurality of MS2 RNA aptamers. In some embodiments, the RNA recognition sequence comprises at least 3, at least 4, at least 5, or at least 6 MS2 RNA aptamers. In some embodiments, the RNA recognition sequence comprises 1 to 30 MS2 RNA aptamers, including any number of MS2 RNA aptamers in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 MS2 RNA aptamers. In some embodiments, the RNA-binding domain comprises a plurality of RNA binding proteins. In some embodiments, the RNA-binding domain comprises multiple repeats of MCP, such as 1 to 6 MCPs, including any number of MCP in this range such as 1, 2, 3, 4, 5, or 6 MCPs.

[0014] In certain embodiments, the RNA binding protein comprises a lambda N peptide and the RNA recognition sequence comprises a box B sequence, wherein the lambda N peptide binds to the box B sequence. In some embodiments, the RNA recognition sequence comprises a plurality of box B sequences. In some embodiments, the RNA recognition sequence comprises at least 3, at least 4, at least 5, or at least 6 box B sequences. In some embodiments, the RNA recognition sequence comprises 1 to 30 box B sequences, including any number of box B sequences in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 box B sequences.In some embodiments, the RNA binding protein comprises multiple repeats of the lambda N peptide, such as 1 to 6 lambda N peptides, including any number of lambda N peptides in this range such as 1, 2, 3, 4, 5, or 6 lambda N peptides.

[0015] In certain embodiments, the RNA binding protein comprises a bacteriophage PP7 coat protein (PCP) and the RNA recognition sequence comprises a PCP binding site, wherein the PCP binds to the PCP binding site. In some embodiments, the RNA recognition sequence comprises a plurality of PCP binding sites. In some embodiments, the RNA recognition sequence comprises at least 3, at least 4, at least 5, or at least 6 PCP binding sites. In some embodiments, the RNA recognition sequence comprises 1 to 30 PCP binding sites, including any number of PCP binding sites in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 PCP binding sites. In some embodiments, the RNA binding protein comprises multiple repeats of the PCP, such as 1 to 6 PCPs, including any number of PCPs in this range such as 1, 2, 3, 4, 5, or 6 PCPs.

[0016] In certain embodiments, the RNA binding protein comprises a BglG transcriptional antiterminator, and the RNA recognition sequence comprises a BglG binding site, wherein the BglG transcriptional antiterminator binds to the BglG binding site. In some embodiments, the RNA recognition sequence comprises a plurality of BglG binding sites. In some embodiments, the RNA recognition sequence comprises at least 3, at least 4, at least 5, or at least 6 BglG binding sites. In some embodiments, the RNA recognition sequence comprises 1 to 30 BglG binding sites, including any number of BglG binding sites in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 BglG binding sites. In some embodiments, the RNA binding protein comprises multiple repeats of the BglG transcriptional antiterminator, such as 1 to 6 BglG transcriptional antiterminators, including any number of the BglG transcriptional antiterminators in this range such as 1, 2, 3, 4, 5, or 6 BglG transcriptional anti terminators.

[0017] In certain embodiments, the RNA binding protein comprises spliceosomal protein U1 A (U1A), and the RNA recognition sequence comprises a U1A binding site, wherein the U1A binds to the U1A binding site. In some embodiments, the RNA recognition sequence comprises a plurality of U1A binding sites. In some embodiments, the RNA recognition sequence comprises at least 3, at least 4, at least 5, or at least 6 U1A binding sites. In some embodiments, the RNA recognition sequence comprises 1 to 30 U1A binding sites, including any number of U1A binding sites in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 PCPbinding sites. In some embodiments, the RNA binding protein comprises multiple repeats of the U1A, such as 1 to 6 UlAs, including any number of UlAs in this range such as 1, 2, 3, 4, 5, or 6 UlAs.

[0018] In certain embodiments, the vector further comprises a polyadenylation site.

[0019] In certain embodiments, the vector further comprises a pseudoknot RNA motif sequence, wherein the pseudoknot RNA motif sequence is positioned between the barcode and the polyadenylation site. In some embodiments, the pseudoknot RNA motif sequence is an evopreql pseudoknot RNA motif sequence.

[0020] In certain embodiments, the second nucleotide further comprises a pair of RNA circulation sites flanking the RNA recognition sequence and the barcode in the barcoded RNA transcript, wherein the barcoded RNA transcript forms a circularized RNA transcript in a cell.

[0021] In certain embodiments, the second nucleotide sequence further encodes a pair of ribozymes comprising a first ribozyme and a second ribozyme, wherein the first ribozyme and the second ribozyme flank the RNA recognition sequence and the barcode in the barcoded RNA transcript, wherein the first ribozyme and the second ribozyme undergo autocatalytic cleavage to generate a 5-end comprising a hydroxyl group and a 3’ -end comprising a 2’, 3’ -cyclic phosphate group, wherein the 5 ’-end and the 3 ’-end are ligated by an endogenous RtcB RNA ligase in a cell resulting in circularization of the barcoded RNA transcript. In some embodiments, the pair of ribozymes are Twister ribozymes.

[0022] In certain embodiments, the first promoter and the second promoter are constitutive or inducible.

[0023] In certain embodiments, the first promoter is a hybrid cytomegalovirus (CMV) / chicken P-actin promoter (CAG) promoter.

[0024] In certain embodiments, the second promoter is a U6 promoter.

[0025] In certain embodiments, the fusion protein further comprises a selectable marker. In some embodiments, the selectable marker is a fluorescent protein. Exemplary fluorescent proteins include, without limitation, green fluorescent proteins, red fluorescent proteins, blue fluorescent proteins, yellow fluorescent proteins, orange fluorescent proteins, and violet fluorescent proteins.

[0026] In certain embodiments, the vector further comprises a pair of restriction sites comprising a first restriction site and a second restriction site, wherein the first restriction site and the second restriction site flank the second expression cassette.

[0027] In certain embodiments, the vector further comprises a multiple cloning site (MCS) for insertion of an expressible sequence into the vector.

[0028] In certain embodiments, the vector further comprises an expressible sequence, wherein the barcode identifies the expressible sequence. In some embodiments, the expressible sequence encodes a protein or an RNA. In some embodiments, the protein is a therapeutic protein or a genome-editing enzyme. In some embodiments, the protein is a variant. In some embodiments, the RNA is a messenger RNA (mRNA) or a non-coding RNA such as, but not limited to, a microRNA (miRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a small nuclear RNA (snRNA), a piwi-interacting RNA (piRNA), a small nucleolar RNA (snoRNA), or a long non-coding RNA (IncRNA).

[0029] In certain embodiments, the RNA is an aptamer or a guide RNA.

[0030] In certain embodiments, the second nucleotide sequence further comprises a sequence that is complementary to an oligonucleotide capture probe.

[0031] In another aspect, a vector is provided, the vector comprising an expression cassette comprising a promoter operably linked to a nucleotide sequence encoding a nuclear retention element connected to a barcode, wherein transcription of the nucleotide sequence in a cell generates a barcoded RNA transcript comprising the nuclear retention element connected to the barcode, wherein the barcoded RNA transcript is translocated to the nucleus of the cell.

[0032] In certain embodiments, the nuclear retention element is a nuclear retention element of a long non-coding RNA such as, but not limited to, MEG3, XIST, MALAT1, or SIRLOIN. The barcoded RNA transcript may comprise an entire long non-coding RNA or a fragment thereof comprising a nuclear retention element, as long as the barcoded RNA transcript is translocated to the nucleus of the cell. In some embodiments, the nuclear retention element comprises the nucleotide sequence of RCCTCCC, wherein R is an A or a G.

[0033] In certain embodiments, the second nucleotide further comprises a pair of RNA circulation sites flanking the RNA recognition sequence and the barcode in the barcoded RNA transcript, wherein the barcoded RNA transcript forms a circularized RNA transcript in a cell.

[0034] In certain embodiments, the nucleotide sequence encoding the nuclear retention element connected to the barcode further encodes a pair of ribozymes comprising a first ribozyme and a second ribozyme, wherein the first ribozyme and the second ribozyme flank the nuclear retention element and the barcode in the barcoded RNA transcript, wherein the first ribozyme andthe second ribozyme undergo autocatalytic cleavage to generate a 5-end comprising a hydroxyl group and a 3’-cnd comprising a 2’,3’-cyclic phosphate group, wherein the 5’-cnd and the 3’-cnd are ligated by an endogenous RtcB RNA ligase in a cell resulting in circularization of the barcoded RNA transcript. In some embodiments, the pair of ribozymes are Twister ribozymes.

[0035] In certain embodiments, the vector is a viral vector. In some embodiments, the viral vector is an adenovirus associated virus (AAV) vector, lentivirus vector, or a G-deleted rabies virus (RVdG) vector. In embodiments in which the viral vector is an adenovirus associated virus (AAV) vector, the AAV vector further comprises a 5 ’-inverted terminal repeat (ITR) and a 3’- ITR, wherein the expression cassette is positioned between the 5 ’-ITR and the 3’-ITR.

[0036] In certain embodiments, the vector further comprises a polyadenylation site.

[0037] In certain embodiments, the vector further comprises a pseudoknot RNA motif sequence, wherein the pseudoknot RNA motif sequence is positioned between the barcode and the polyadenylation site. In some embodiments, the pseudoknot RNA motif sequence is an evopreql pseudoknot RNA motif sequence.

[0038] In certain embodiments, the promoter is constitutive or inducible.

[0039] In certain embodiments, the promoter is a U6 promoter.

[0040] In certain embodiments, the nucleotide sequence further encodes a selectable marker.In some embodiments, the selectable marker is a fluorescent protein. Exemplary fluorescent proteins include, without limitation, green fluorescent proteins, red fluorescent proteins, blue fluorescent proteins, yellow fluorescent proteins, orange fluorescent proteins, and violet fluorescent proteins.

[0041] In certain embodiments, the vector further comprises a pair of restriction sites comprising a first restriction site and a second restriction site, wherein the first restriction site and the second restriction site flank the expression cassette.

[0042] In certain embodiments, the vector further comprises a multiple cloning site (MCS) for insertion of an expressible sequence into the vector.

[0043] In certain embodiments, the vector further comprises an expressible sequence, wherein the barcode identifies the expressible sequence.

[0044] In certain embodiments, the vector further comprises a sequence that is complementary to an oligonucleotide capture probe.

[0045] In another aspect, a cell transfected with a vector, described herein, is provided. In some embodiments, the cell is a mammalian cell. In some embodiments, the mammalian cell is a human cell. In some embodiments, the cell is a neuron.

[0046] In another aspect, a method of barcoding nuclei in a population of cells is provided, the method comprising; providing a plurality of vectors described herein, wherein the barcode in each vector of the plurality comprises a unique identifier sequence; and transfecting the population of cells with the plurality of vectors to generate cells having barcoded nuclei.

[0047] In certain embodiments, the transfecting is performed in vitro, ex vivo, or in vivo.

[0048] In certain embodiments, the population of cells is in a tissue or an organ.

[0049] In another aspect, a method for multiplex screening of a population of cells is provided, the method comprising: providing a plurality of vectors described herein, wherein the barcode in each vector of the plurality comprises a unique identifier sequence; transfecting the population of cells with the plurality of vectors to generate cells having barcoded nuclei; measuring morphological or functional characteristics of the cells having the barcoded nuclei; isolating the barcoded nuclei; and identifying the barcode in each barcoded nucleus.

[0050] In certain embodiments, the population of cells is in a tissue or an organ. In some embodiments, the tissue is nervous tissue.

[0051] In certain embodiments, the population of cells is in a culture.

[0052] In certain embodiments, the population of cells comprises neurons or glial cells or a combination thereof.

[0053] In certain embodiments, the method further comprises optogenetically modifying the neurons.

[0054] In certain embodiments, measuring the morphological or functional characteristics comprises performing gene expression profiling, microscopy, calcium imaging, an electrophysiology measurement, functional neuroimaging, a migration assay, an axonal growth and pathfinding assay, retrograde or monosynaptic tracing, a phagocytosis assay, an enzymatic assay, a cell receptor assay, an ion channel assay, a signal transduction assay, or a cell secretion assay.

[0055] In certain embodiments, the gene expression profiling comprises performing microarray analysis, RNA sequencing (e.g., single nucleus RNA sequencing, single cell RNA sequencing), or quantitative polymerase chain reaction.

[0056] In certain embodiments, the microscopy is confocal microscopy, atomic force microscopy, super-resolution microscopy, light-sheet microscopy, two-photon microscopy, or fluorescence microscopy.

[0057] In certain embodiments, the electrophysiology measurement is patch clamping, electroencephalography (EEG), or magnetoencephalography (MEG).

[0058] In certain embodiments, the functional neuroimaging is functional magnetic resonance imaging (fMRI), positron emission tomography (PET), functional near-infrared spectroscopy (fNIRS), single-photon emission computed tomography (SPECT), or functional ultrasound imaging (fUS).

[0059] In certain embodiments, the morphological or functional characteristics are measured in the population of cells in a tissue of a live subject in vivo or in culture in vitro.

[0060] In certain embodiments, the subject is a nonhuman animal.

[0061] In certain embodiments, the method further comprises removing the tissue from the subject after measuring the morphological or functional characteristics.

[0062] In certain embodiments, isolating the barcoded nuclei comprises performing fluorescence activated nuclei sorting (FANS).

[0063] In certain embodiments, the method further comprises isolating the barcoded RNA transcript from each barcoded nucleus. In some embodiments, isolating the barcoded RNA transcript comprises hybridization of the barcoded RNA transcript to a capture probe comprising a sequence complementary to a region of the barcoded RNA transcript. In some embodiments, isolating the barcoded RNA transcript comprises hybridization of the barcoded RNA transcript to a capture probe comprising a poly-T sequence complementary to the poly-A tail of the barcoded RNA transcript. In some embodiments, the capture probe is immobilized on a solid support. In some embodiments, the solid support is a magnetic bead or polystyrene bead.

[0064] In certain embodiments, identifying the barcode comprises sequencing the barcode sequence. In some embodiments, the sequencing comprises performing single nucleus RNA sequencing (snRNA-seq).

[0065] In certain embodiments, the method further comprises amplifying the barcode sequence prior to said sequencing. In some embodiments, amplifying comprises performing reverse transcription polymerase chain reaction or isothermal amplification.

[0066] In certain embodiments, the method further comprises reverse transcribing the barcoded RNA transcript to generate a complementary DNA (cDNA) copy of the barcoded RNA transcript. In some embodiments, the method further comprises adding an adapter to the 5’ end and the 3’ end of the cDNA copy of the barcoded RNA transcript.

[0067] In certain embodiments, each vector of the plurality further comprises a different expressible sequence, wherein the barcode in each vector identifies the expressible sequence.

[0068] In certain embodiments, the expressible sequence encodes a protein or an RNA. In some embodiments, the protein is a therapeutic protein or a genome-editing enzyme. In some embodiments, the therapeutic protein is a hormone, a cytokine, a chemokine, a growth factor, an enzyme, or an antibody. In some embodiments, the protein or RNA is a variant. In some embodiments, the genome-editing enzyme is a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) nuclease, a meganuclease, a zinc-finger nuclease (ZFN), or a transcription activator-like effector nuclease (TALEN). In some embodiments, the Cas nuclease is Cas9 or Cas 12a. In some embodiments, the RNA is a messenger RNA (mRNA) or a non-coding RNA. In some embodiments, the non-coding RNA is a microRNA (miRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a small nuclear RNA (snRNA), a piwi-interacting RNA (piRNA), a small nucleolar RNA (snoRNA), or a long non-coding RNA (IncRNA). In some embodiments, the RNA is an aptamer or a guide RNA. In some embodiments, the protein is a fluorescent protein or a bioluminescent protein. In some embodiments, the method further comprises imaging the fluorescent protein or the bioluminescent protein, wherein a location of a cell expressing the fluorescent protein or the bioluminescent protein is determined from the imaging. In some embodiments, the method further comprises mapping the location of the cell expressing the fluorescent protein or the bioluminescent protein onto a reference image of the tissue.

[0069] In another aspect, a vector system is provided, the vector system comprising a vector described herein and a helper virus vector. In some embodiments in which the viral vector is an AAV vector, the helper virus vector encodes Ela and Elb or E2a and E4. In some embodiments in which the viral vector is a lentivirus vector, the helper virus vector encodes Gag, Pol, Rev, and VSV-G. In some embodiments in which the viral vector is a G-deleted rabies vims (RVdG) vector, the helper virus vector encodes a glycoprotein (G) gene.

[0070] In certain embodiments, the vector system further comprises a vector encoding one or more capsid proteins. In some embodiments in which the viral vector is an AAV vector, the one or more capsid proteins comprise VP1, VP2, and VP3. In some embodiments in which the viral vector is an AAV vector, the vector system further comprises a vector encoding an AAV Rep protein.

[0071] In another aspect, a cell transfected with a vector system, described herein, is provided. In some embodiments, the cell is a mammalian cell. In some embodiments, the mammalian cell is a human cell. In some embodiments, the cell is a neuron.

[0072] In another aspect, a vector library comprising a plurality of vectors, as described herein, wherein the plurality of vectors collectively comprise a plurality of expressible sequences, and wherein the barcode in each vector identifies the expressible sequence in each vector.

[0073] In certain embodiments the plurality of expressible sequences collectively encode different therapeutic proteins, protein variants, antibodies, mRNAs, shRNAs, siRNAs, aptamers, or guide RNAs for RNA-guided nucleases (e.g., Cas9 for gene editing by CRISPR system).DETAILED DESCRIPTION

[0074] Vectors, cell lines comprising the vectors, recombinant virions produced from the vectors, and methods of using the vectors for barcoding nuclei of cells are provided. Vector libraries that collectively encode a plurality of genetic elements of interest for screening and methods of screening cells having nuclei barcoded according to the subject methods are also provided.

[0075] Before the present vectors, cell lines, recombinant virions, vector libraries, and methods of using the vectors for barcoding nuclei and screening cells are described, it is to be understood that this invention is not limited to particular methods or compositions described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.

[0076] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that statedrange is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where cither, neither or both limits are included in the smaller ranges is also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0077] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, some potential and preferred methods and materials are now described. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. It is understood that the present disclosure supersedes any disclosure of an incorporated publication to the extent there is a contradiction.

[0078] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible.

[0079] It must be noted that as used herein and in the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a cell" includes a plurality of such cells, and reference to "the vector" includes reference to one or more vectors and equivalents thereof, such as viral vectors, plasmids, constructs, and the like, known to those skilled in the ail, and so forth.

[0080] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.Definitions

[0081] The term "about", particularly in reference to a given quantity, is meant to encompass deviations of plus or minus five percent.

[0082] "AAV" is an abbreviation for adeno-associated virus, and may be used to refer to the vims itself or derivatives thereof. The term covers all subtypes and both naturally occurring and recombinant forms, except where required otherwise.

[0083] By "recombinant virus" is meant a virus that has been genetically altered, e.g., by the addition or insertion of a heterologous nucleic acid construct into the particle.

[0084] The abbreviation "rAAV" refers to recombinant adeno-associated virus, also referred to as a recombinant AAV vector (or "rAAV vector"). The term “AAV” includes any AAV serotype, such as, but not limited to, AAV type 1 (AAV-1), AAV type 2 (AAV-2), AAV type 3 (AAV-3), AAV type 4 (AAV-4), AAV type 5 (AAV-5), AAV type 6 (AAV-6), AAV type 7 (AAV-7), AAV type 8 (AAV-8), AAV type 9 (AAV-9), AAV type 10 (AAV- 10), avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and ovine AAV. “Primate AAV” refers to AAV isolated from a primate, “non-primate AAV” refers to AAV isolated from a non-primate mammal, “bovine AAV” refers to AAV isolated from a bovine mammal (e.g., a cow), etc.

[0085] An "rAAV vector" as used herein refers to an AAV vector comprising a polynucleotide sequence not of AAV origin (i.e., a polynucleotide heterologous to AAV), typically a sequence of interest for introducing into a target cell. In general, the heterologous polynucleotide is flanked by at least one, and generally by two AAV inverted terminal repeat sequences (ITRs). The term rAAV vector encompasses both rAAV vector particles and rAAV vector plasmids.

[0086] An "AAV virus" or "AAV viral particle" or "rAAV vector particle" refers to a viral particle composed of at least one AAV capsid protein (typically by all of the capsid proteins of a wild-type AAV) and an encapsidated polynucleotide rAAV vector. If the particle comprises a heterologous polynucleotide (i.e., a polynucleotide other than a wild-type AAV genome, such as a transgene to be delivered to a mammalian cell), it is typically referred to as an "rAAV vector particle" or simply an "rAAV vector". Thus, production of a rAAV particle necessarily includes production of a rAAV vector, as such a vector contained within an rAAV particle.

[0087] "Packaging" refers to a series of intracellular events that result in the assembly and encapsidation of a viral particle.

[0088] AAV "rep" and "cap" genes refer to polynucleotide sequences encoding replication and cncapsidation proteins of adcno-associatcd virus. AAV rep and cap arc referred to herein as AAV "packaging genes."

[0089] By "AAV rep coding region" is meant the art-recognized region of the AAV genome which encodes the replication proteins of the virus which are required to replicate the viral genome and to insert the viral genome into a host genome during latent infection. The term also includes functional homologues thereof such as the human herpesvirus 6 (HHV-6) rep gene which is also known to mediate AAV-2 DNA replication (Thomson et al. (1994) Virology 204, 304-311). For a further description of the AAV rep coding region, see, e.g., Muzyczka, N. (1992) Current Topics in Microbiol, and Immunol. 158, 97-129; Kotin, R. M. (1994) Human Gene Therapy 5, 793-801. The rep coding region, as used herein, can be derived from any viral serotype, such as those described above. The region need not include all of the wild-type genes but may be altered, e.g., by the insertion, deletion or substitution of nucleotides, so long as the rep genes present provide for sufficient integration functions when expressed in a suitable recipient cell.

[0090] By " AAV cap coding region" is meant the art-recognized region of the AAV genome which encodes the coat proteins of the vims which are required for packaging the viral genome. For a further description of the cap coding region, see, e.g., Muzyczka, N. (1992) Current Topics in Microbiol, and Immunol. 158, 97-129; Kotin, R. M. (1994) Human Gene Therapy 5, 793-801. The AAV cap coding region, as used herein, can be derived from any AAV serotype, as described above. The region need not include all of the wild-type cap genes but may be altered, e.g., by the insertion, deletion or substitution of nucleotides, so long as the genes provide for sufficient packaging functions when present in a host cell along with an AAV vector.

[0091] By " adeno-associated virus inverted terminal repeats" or "AAV ITRs” is meant the art- recognized regions found at each end of the AAV genome which function together in cis as origins of DNA replication and as packaging signals for the viral genome. AAV ITRs, together with the AAV rep coding region, provide for the efficient excision and rescue from, and integration of a nucleotide sequence interposed between two flanking ITRs into a mammalian cell genome. The nucleotide sequences of AAV ITR regions are known. See, e.g., Kotin, R. M. (1994) Human Gene Therapy 5, 793-801; Berns, K. I. "Parvoviridae and their Replication" in Fundamental Virology, 2d ed., (B. N. Fields and D. M. Knipe, eds.) for the AAV-2 sequence. As used herein, an "AAV ITR" need not have the wild-type nucleotide sequence depicted in the previously cited references,but may be altered, e.g., by the insertion, deletion or substitution of nucleotides. Additionally, the AAV ITR may be derived from any of several AAV serotypes, including without limitation, AAV- 1, AAV-2, AAV-3, AAV-4, AAV-5, AAVX7, etc. Furthermore, 5' and 3' ITRs which flank a selected nucleotide sequence in an AAV vector need not necessarily be identical or derived from the same AAV serotype or isolate, so long as they function as intended, i.e., to allow for excision and rescue of the sequence of interest from a host cell genome or vector, and to allow integration of the heterologous sequence into the recipient cell genome when AAV Rep gene products are present in the cell.

[0092] A "helper virus" for AAV refers to a virus that allows AAV (e.g., wild-type AAV) to be replicated and packaged by a mammalian cell. A variety of such helper viruses for AAV are known in the art, including adenoviruses, herpesviruses and poxviruses such as vaccinia. The adenoviruses encompass a number of different subgroups, although Adenovirus type 5 of subgroup C is most commonly used. Numerous adenoviruses of human, non-human mammalian and avian origin are known and available from depositories such as the ATCC. Viruses of the herpes family include, for example, herpes simplex viruses (HSV) and Epstein-Barr viruses (EBV), as well as cytomegaloviruses (CMV) and pseudorabies viruses (PRV); which are also available from depositories such as ATCC.

[0093] "Helper virus function(s)" refers to function(s) encoded in a helper virus genome which allow viral AAV replication and packaging (in conjunction with other requirements for replication and packaging described herein). As described herein, "helper virus function" may be provided in a number of ways, including by providing helper virus or providing, for example, polynucleotide sequences encoding the requisite function(s) to a producer cell in trans.

[0094] "RVdG " is an abbreviation for a G-deleted rabies virus and may be used to refer to the rabies virus itself or derivatives thereof. The term covers all G-deleted rabies virus subtypes. RVdG is an enveloped, non-segmented negative-stranded RNA virus, genetically modified to lack the envelope glycoprotein gene (G gene). The glycoprotein is not required for the transcription or replication of the genome within infected cells but is required for RVdG to spread from cell to cell. Thus, the viral core of RVdG can proliferate within initially infected cells, but the inability to synthesize the glycoprotein prevents its progeny from infecting other cells.

[0095] RVdG infects neurons retrogradely, infecting neurons through axon terminals and spreading between synaptically coupled neurons exclusively in the retrograde direction. Thisfeature of the virus makes RVdG useful for retrograde synaptic tracing (see, e.g., Wickersham et al. (2007) Nat Methods. 4(l):47-49, Wickcrsham ct al. (2007) Neuron. 53(5):639-647; herein incorporated by reference in their entireties). In some embodiments, the G gene is replaced with a gene encoding a fluorescent protein to allow fluorescence imaging of cells infected with the virus. In some embodiments, the G gene is provided in trans to allow the virus to spread from cell to cell.

[0096] An "infectious1' virus or viral particle is one that comprises a polynucleotide component which is capable of delivering into a cell for which the viral species is tropic. The term does not necessarily imply any replication capacity of the virus. As used herein, an “infectious” virus or viral particle is one that can access a target cell, can infect a target cell, and can express a heterologous nucleic acid in a target cell. Thus, “infectivity” refers to the ability of a viral particle to access a target cell, infect a target cell, and express a heterologous nucleic acid in a target cell. Infectivity can refer to in vitro infectivity or in vivo infectivity. Assays for counting infectious viral particles are described elsewhere in this disclosure and in the art. Viral infectivity can be expressed as the ratio of infectious viral particles to total viral particles. Total viral particles can be expressed as the number of viral genome (vg) copies. The ability of a viral particle to express a heterologous nucleic acid in a cell can be referred to as “transduction.” The ability of a viral particle to express a heterologous nucleic acid in a cell can be assayed using a number of techniques, including assessment of a marker gene, such as a green fluorescent protein (GFP) assay (e.g., where the virus comprises a nucleotide sequence encoding GFP), where GFP is produced in a cell infected with the viral particle and is detected and / or measured; or the measurement of a produced protein, for example by an enzyme-linked immunosorbent assay (ELISA). Viral infectivity can be expressed as the ratio of infectious viral particles to total viral particles. Methods of determining the ratio of infectious viral particle to total viral particle are known in the art. See, e.g., Grainger et al. (2005) Mol. Ther. 11 :S337 (describing a TCID50 infectious titer assay); and Zolotukhin et al. (1999) Gene Ther. 6:973.

[0097] A "replication-competent" virus (e.g., a replication-competent AAV) refers to a phenotypically wild-type virus that is infectious, and is also capable of being replicated in an infected cell (i.e., in the presence of a helper virus or helper virus functions). In the case of AAV, replication competence generally requires the presence of functional AAV packaging genes.

[0098] A "barcode" refers to one or more nucleotide sequences that are used to identify a nucleic acid, vector, or cell with which the barcode is associated. Barcodes can be 3-1000 or morenucleotides in length, preferably 10-250 nucleotides in length, and more preferably 10-30 nucleotides in length, including any length within these ranges, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length. Barcodes may be used, for example, to identify a single cell, subpopulation of cells, colony, or sample from which a nucleic acid or vector originated. Barcodes may also be used to identify the position (i.e., positional barcode) of a cell, colony, or sample from which a nucleic acid or vector originated, such as the position of a colony in a cellular’ array, the position of a well in a multi-well plate, or the position of a tube, flask, or other container in a rack. In particular, a barcode may be used to identify a genetically modified cell from which a nucleic acid or vector originated. In some embodiments, a barcode is used to identify a genetic element or expressible sequence from a vector in a cell. Furthermore, multiple barcodes can be used in combination to identify different features of a nucleic acid, vector, or cell. For example, positional barcoding (e.g., to identify the position of a cell, colony, culture, or sample in an array, multi-well plate, or rack) can be combined with barcodes identifying expressible sequences (e.g., encoding therapeutic proteins, shRNAs, siRNAs, aptamers, guide-RNAs, etc.) from a vector.

[0099] The term "polynucleotide" refers to a polymeric form of nucleotides of any length, including deoxyribonucleotides or ribonucleotides, or analogs thereof. A polynucleotide may comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs, and may be interrupted by non-nucleotide components. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The term polynucleotide, as used herein, refers interchangeably to double- and single-stranded molecules. Unless otherwise specified or required, any embodiment of the invention described herein that is a polynucleotide encompasses both the double-stranded form and each of two complementary single-stranded forms known or predicted to make up the double-stranded form.

[0100] A polynucleotide or polypeptide has a certain percent "sequence identity" to another polynucleotide or polypeptide, meaning that, when aligned, that percentage of bases or amino acids are the same when comparing the two sequences. Sequence similarity can be determined in a number of different manners. To determine sequence identity, sequences can be aligned using the methods and computer programs, including BLAST, available over the world wide web at ncbi.nlm.nih.gov / BLAST / . Another alignment algorithm is FASTA, available in the GeneticsComputing Group (GCG) package, from Madison, Wisconsin, USA, a wholly owned subsidiary of Oxford Molecular Group, Inc. Other techniques for alignment arc described in Methods in Enzymology, vol. 266: Computer Methods for Macromolecular Sequence Analysis (1996), ed. Doolittle, Academic Press, Inc., a division of Harcourt Brace & Co., San Diego, California, USA. Of particular interest are alignment programs that permit gaps in the sequence. The Smith- Waterman is one type of algorithm that permits gaps in sequence alignments. See Meth. Mol. Biol. 70: 173-187 (1997). Also, the GAP program using the Needleman and Wunsch alignment method can be utilized to align sequences. See J. Mol. Biol. 48: 443-453 (1970)

[0101] Of interest is the BestFit program using the local homology algorithm of Smith Waterman (Advances in Applied Mathematics 2: 482-489 (1981) to determine sequence identity. The gap generation penalty will generally range from 1 to 5, usually 2 to 4 and in some cases will be 3. The gap extension penalty will generally range from about 0.01 to 0.20 and in many instances will be 0.10. The program has default parameters determined by the sequences inputted to be compared. Preferably, the sequence identity is determined using the default parameters determined by the program. This program is available also from Genetics Computing Group (GCG) package, from Madison, Wisconsin, USA.

[0102] Another program of interest is the FastDB algorithm. FastDB is described in Current Methods in Sequence Comparison and Analysis, Macromolecule Sequencing and Synthesis, Selected Methods and Applications, pp. 127-149, 1988, Alan R. Liss, Inc. Percent sequence identity is calculated by FastDB based upon the following parameters:Mismatch Penalty: 1.00;Gap Penalty: 1.00;Gap Size Penalty: 0.33; andJoining Penalty: 30.0.

[0103] A "gene" refers to a polynucleotide containing at least one open reading frame that is capable of encoding a particular protein after being transcribed and translated.

[0104] " Gene transfer" or "gene delivery" refers to methods or systems for inserting foreign DNA into host cells. Gene transfer can result in transient expression of non-integrated transferred DNA, extrachromosomal replication and expression of transferred replicons (e.g., episomes), or integration of transferred genetic material into the genomic DNA of host cells.

[0105] The term "host cell" denotes, for example, microorganisms, yeast cells, insect cells, and mammalian cells, that can be, or have been, used as recipients of a vector or vector system described herein, or other transfer DNA. The term includes the progeny of the original cell which has been transfected. Thus, a "host cell" as used herein generally refers to a cell which has been transfected with an exogenous DNA sequence. It is understood that the progeny of a single parental cell may not necessarily be completely identical in morphology or in genomic or total DNA complement to the original parent, due to natural, accidental, or deliberate mutation.

[0106] As used herein, the term "cell line" refers to a population of cells capable of continuous or prolonged growth and division in vitro. Often, cell lines are clonal populations derived from a single progenitor cell. It is further known in the art that spontaneous or induced changes can occur in karyotype during storage or transfer of such clonal populations. Therefore, cells derived from the cell line referred to may not be precisely identical to the ancestral cells or cultures, and the cell line referred to includes such variants.

[0107] By ' ’vector" is meant any genetic element, such as a plasmid, phage, transposon, cosmid, chromosome, artificial chromosome, virus, virion, etc., which is capable of replication when associated with the proper control elements and which can transfer gene sequences between cells. Thus, the term includes cloning and expression vehicles, as well as viral vectors.

[0108] " Recombinant," as applied to a polynucleotide means that the polynucleotide is the product of various combinations of cloning, restriction or ligation steps, and other procedures that result in a construct that is distinct from a polynucleotide found in nature. A recombinant virus is a viral particle comprising a recombinant polynucleotide. The terms respectively include replicates of the original polynucleotide construct and progeny of the original virus construct.

[0109] A "control element" or "control sequence" is a nucleotide sequence involved in an interaction of molecules that contributes to the functional regulation of a polynucleotide, including replication, duplication, transcription, splicing, translation, or degradation of the polynucleotide. The regulation may affect the frequency, speed, or specificity of the process, and may be enhancing or inhibitory in nature. Control elements known in the art include, for example, transcriptional regulatory sequences such as promoters and enhancers. A promoter is a DNA region capable under certain conditions of binding RNA polymerase and initiating transcription of a coding region usually located downstream (in the 3' direction) from the promoter.

[0110] "Operatively linked" or "operably linked" refers to a juxtaposition of genetic elements, wherein the elements arc in a relationship permitting them to operate in the expected manner. For instance, a promoter is operatively linked to a coding region if the promoter helps initiate transcription of the coding sequence. There may be intervening residues between the promoter and coding region so long as this functional relationship is maintained.

[0111] The term “expressible sequence” refers to a polynucleotide which is operably linked to a promoter element such that the promoter element is able to cause transcriptional expression of the expressible sequence. An expressible sequence is typically linked downstream, on the 3'-end of the promoter element(s) in order to achieve transcriptional expression. The result of this transcriptional expression is the production of an RNA macromolecule. The expressed RNA molecule may encode a protein and may thus be subsequently translated by the appropriate cellular machinery to produce a polypeptide / protein molecule. In some embodiments, the expressible sequence encodes a therapeutic protein such as, but not limited to, a hormone, a cytokine, a chemokine, a growth factor, an enzyme, or an antibody. Alternately, the RNA molecule may be an antisense RNA or other non-coding RNA molecule such as a microRNA, shRNA, or siRNA, which is capable of modulating the expression of specific genes in a cell, as is known in the art. In some embodiments, the RNA is a guide RNA for an RNA-guided nuclease (e.g., Cas9 for gene editing by CRISPR system). In some embodiments, the RNA is an aptamer.

[0112] "Expression cassette" or "expression construct" refers to an assembly which is capable of directing the expression of the sequence(s) or gene(s) of interest. An expression cassette generally includes control elements, as described above, such as a promoter which is operably linked to (so as to direct transcription of) the sequence(s) or gene(s) of interest, and often includes a polyadenylation sequence as well. Within certain embodiments of the invention, the expression cassette described herein may be contained within a plasmid construct. In addition to the components of the expression cassette, the plasmid construct may also include, one or more selectable markers, a signal which allows the plasmid construct to exist as single stranded DNA (e.g., a M13 origin of replication), at least one multiple cloning site, and a "mammalian" origin of replication (e.g., a SV40 or adenovirus origin of replication).

[0113] An "expression vector" is a vector comprising a region which encodes a polypeptide or RNA transcript of interest, and is used for effecting the expression of the protein or RNA transcript in an intended target cell. An expression vector also comprises control elements operatively linkedto the encoding region to facilitate expression of the protein or RNA transcript in the target cell. The combination of control elements and a gene or genes to which they arc operably linked for expression is sometimes referred to as an "expression cassette," a large number of which are known and available in the art or can be readily constructed from components that are available in the art.

[0114] "Heterologous" means derived from a genotypically distinct entity from that of the rest of the entity to which it is being compared. For example, a polynucleotide introduced by genetic engineering techniques into a plasmid or vector derived from a different species is a heterologous polynucleotide. A promoter removed from its native coding sequence and operatively linked to a coding sequence with which it is not naturally found linked is a heterologous promoter. Thus, for example, an rAAV that includes a heterologous nucleic acid encoding a heterologous gene product is an rAAV that includes a nucleic acid not normally included in a naturally occurring, wild-type AAV, and the encoded heterologous gene product is a gene product not normally encoded by a naturally occurring, wild-type AAV. As another example, a variant AAV capsid protein that comprises a heterologous peptide inserted into the GH loop of the capsid protein is a variant AAV capsid protein that includes an insertion of a peptide not normally included in a naturally occurring, wild-type AAV.

[0115] The terms “genetic alteration” and “genetic modification” (and grammatical variants thereof) are used interchangeably herein to refer to a process wherein a genetic element (e.g., a polynucleotide) is introduced into a cell other than by mitosis or meiosis. The element may be heterologous to the cell, or it may be an additional copy or improved version of an element already present in the cell. Genetic alteration may be effected, for example, by transfecting a cell with a recombinant plasmid or other polynucleotide through any process known in the art, such as electroporation, calcium phosphate precipitation, or contacting with a polynucleotide-liposome complex. Genetic alteration may also be effected, for example, by transduction or infection with a DNA or RNA virus or viral vector. Generally, the genetic element is introduced into a chromosome or mini-chromosome in the cell; but any alteration that changes the phenotype and / or genotype of the cell and its progeny is included in this term.

[0116] The term "transfection" is used to refer to the uptake of foreign DNA by a cell. A cell has been "transfected" when exogenous DNA has been introduced inside the cell membrane. A number of transfection techniques are generally known in the art. See, e.g., Graham et al. (1973) Virology, 52:456; Green and Sambrook Molecular Cloning: A Laboratory Manual (Cold SpringHarbor Laboratory Press, 4thedition, 2012); Current Protocols in Molecular Biology (Ausubel ed., John Wiley & Sons, 1995); Davis ct al. (1995) Basic Methods in Molecular Biology, 2ndedition, McGraw-Hill; and Chu et al. (1981) Gene 13:197. Such techniques can be used to introduce one or more exogenous DNA moieties into suitable host cells. The term refers to both stable and transient uptake of the genetic material, and includes uptake of peptide- or antibody-linked DNAs.

[0117] A cell is said to be "stably" altered, transduced, genetically modified, or transformed with a genetic sequence if the sequence is available to perform its function during extended culture of the cell in vitro. Generally, such a cell is "heritably" altered (genetically modified) in that a genetic alteration is introduced which is also inheritable by progeny of the altered cell.

[0118] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The terms also encompass an amino acid polymer that has been modified; for example, disulfide bond formation, glycosylation, lipidation, phosphorylation, or conjugation with a labeling component. Polypeptides such as therapeutic polypeptides, when discussed in the context of delivering a gene product to a mammalian subject, and compositions therefor, refer to the respective intact polypeptide, or any fragment or genetically engineered derivative thereof, which retains the desired biochemical function of the intact protein. Similarly, references to nucleic acids encoding anti-angiogenic polypeptides, nucleic acids encoding neuroprotective polypeptides, and other such nucleic acids for use in delivery of a gene product to a mammalian subject (which may be referred to as "transgenes" to be delivered to a recipient cell), include polynucleotides encoding the intact polypeptide or any fragment or genetically engineered derivative possessing the desired biochemical function.

[0119] The terms "fusion protein," "fusion polypeptide," or "fusion peptide" as used herein refer to a fusion comprising a nuclear localization domain in combination with an RNA-binding domain as part of a single continuous chain of amino acids, which chain does not occur in nature. The nuclear localization domain and the RNA-binding domain may be connected directly to each other by peptide bonds or may be separated by intervening amino acid sequences. The fusion protein may also contain other sequences such as a selectable marker or detectable label (e.g., fluorescent or bioluminescent protein).

[0120] By "fragment" is intended a molecule consisting of only a part of the intact full-length sequence and structure. The fragment can include a C-terminal deletion an N- terminal deletion, and / or an internal deletion of the polypeptide. Active fragments of a particular protein orpolypeptide will generally include at least about 5-14 contiguous amino acid residues of the full length molecule, but may include at least about 15-25 contiguous amino acid residues of the full- length molecule, and can include at least about 20-50 or more contiguous amino acid residues of the full-length molecule, or any integer between 5 amino acids and the full-length sequence, provided that the fragment in question retains biological activity.

[0121] “Isolated” refers to an entity of interest that is in an environment different from that in which it may naturally occur. “Isolated” is meant to include entities that are within samples that are substantially enriched for the entity of interest and / or in which the entity of interest is partially or substantially purified.

[0122] An "isolated" plasmid, nucleic acid, vector, virus, virion, host cell, nucleus, or other substance refers to a preparation of the substance devoid of at least some of the other components that may also be present where the substance or a similar substance naturally occurs or is initially prepared from. Thus, for example, an isolated substance may be prepared by using a purification technique to enrich it from a source mixture. Enrichment can be measured on an absolute basis, such as weight per volume of solution, or it can be measured in relation to a second, potentially interfering substance present in the source mixture. Increasing enrichments of the embodiments of this invention are increasingly more isolated. An isolated plasmid, nucleic acid, vector, virus, host cell, nucleus, or other substance is in some cases purified, e.g., from about 80% to about 90% pure, at least about 90% pure, at least about 95% pure, at least about 98% pure, or at least about 99%, or more, pure.

[0123] "Substantially purified" generally refers to isolation of a substance (e.g., compound, nucleus, cell, vector, nucleic acid, or protein) such that the substance comprises the majority percent of the sample in which it resides. Typically, in a sample, a substantially purified component comprises 50%, preferably 80%-85%, more preferably 90-95% of the sample. Techniques for purifying entities of interest are well-known in the art and include, for example, ion-exchange chromatography, affinity chromatography, differential centrifugation and sedimentation according to density, and flow cytometry.

[0124] The terms “individual”, “subject”, and “patient”, are used interchangeably herein and refer to any mammalian subject, including human and non-human mammals such as non-human primates, including chimpanzees and other apes and monkey species; laboratory animals such as mice, rats, rabbits, hamsters, guinea pigs, and chinchillas; domestic animals such as dogs and cats;farm animals such as sheep, goats, pigs, horses and cows. In some cases, the methods of the invention find use in experimental animals, in veterinary application, and in the development of animal models for disease, including, but not limited to, rodents including mice, rats, and hamsters; primates, and transgenic animals.

[0125] The term "animal" is used herein to include all vertebrate and invertebrate animals, except humans. The term also includes animals at all stages of development.

[0126] The terms "hybridize" and "hybridization" refer to the formation of complexes between nucleotide sequences which are sufficiently complementary to form complexes via Watson-Crick base pairing.

[0127] The term "homologous region" refers to a region of a nucleic acid with homology to another nucleic acid region. Thus, whether a "homologous region" is present in a nucleic acid molecule is determined with reference to another nucleic acid region in the same or a different molecule. Further, since a nucleic acid is often double-stranded, the term "homologous, region," as used herein, refers to the ability of nucleic acid molecules to hybridize to each other. For example, a single-stranded nucleic acid molecule can have two homologous regions which are capable of hybridizing to each other. Thus, the term "homologous region" includes nucleic acid segments with complementary sequences. Homologous regions may vary in length, but will typically be between 4 and 500 nucleotides (e.g., from about 4 to about 40, from about 40 to about 80, from about 80 to about 120, from about 120 to about 160, from about 160 to about 200, from about 200 to about 240, from about 240 to about 280, from about 280 to about 320, from about 320 to about 360, from about 360 to about 400, from about 400 to about 440, etc.).

[0128] As used herein, the terms "complementary" or "complementarity" refers to polynucleotides that are able to form base pairs with one another. Base pairs are typically formed by hydrogen bonds between nucleotide units in an anti-parallel orientation between polynucleotide strands. Complementary polynucleotide strands can base pair in a Watson-Crick manner (e.g., A to T, A to U, C to G), or in any other manner that allows for the formation of duplexes. As persons skilled in the ail are aware, when using RNA as opposed to DNA, uracil (U) rather than thymine (T) is the base that is considered to be complementary to adenosine. However, when a uracil is denoted in the context of the present invention, the ability to substitute a thymine is implied, unless otherwise stated. "Complementarity" may exist between two RNA strands, two DNA strands, or between an RNA strand and a DNA strand. It is generally understood that two or morepolynucleotides may be "complementary" and able to form a duplex despite having less than perfect or less than 100% complementarity. Two sequences arc "perfectly complementary" or "100% complementary" if at least a contiguous portion of each polynucleotide sequence, comprising a region of complementarity, perfectly base pairs with the other polynucleotide without any mismatches or interruptions within such region. Two or more sequences are considered "perfectly complementary" or " 100% complementary" even if either or both polynucleotides contain additional non-complementary sequences as long as the contiguous region of complementarity within each polynucleotide is able to perfectly hybridize with the other. "Less than perfect" complementarity refers to situations where less than all of the contiguous nucleotides within such region of complementarity are able to base pair with each other. Determining the percentage of complementarity between two polynucleotide sequences is a matter of ordinary skill in the art.

[0129] As used herein, the term “capture oligonucleotide” refers to an oligonucleotide that contains a nucleic acid sequence complementary to a nucleic acid sequence present in the target nucleic acid analyte such that the capture oligonucleotide can “capture” the target nucleic acid. One or more capture oligonucleotides can be used in order to capture the target analyte. The polynucleotide regions of a capture oligonucleotide may be composed of DNA, and / or RNA, and / or synthetic nucleotide analogs. By “capture” is meant that the analyte can be separated from other components of the sample by virtue of the binding of the capture molecule to the analyte. Typically, the capture molecule is associated with a solid support, either directly or indirectly.

[0130] It will be appreciated that the hybridizing sequences need not have perfect complementarity to provide stable hybrids. In many situations, stable hybrids will form where fewer than about 10% of the bases are mismatches, ignoring loops of four or more nucleotides. Accordingly, as used herein the term “complementary” refers to an oligonucleotide that forms a stable duplex with its “complement” under assay conditions, generally where there is about 90% or greater homology.

[0131] The terms “hybridize” and “hybridization” refer to the formation of complexes between nucleotide sequences which are sufficiently complementary to form complexes via Watson-Crick base pairing.Vectors for Barcoding Nuclei of Cells

[0132] Vectors arc provided for barcoding nuclei of cells in a cell population. The nucleus of an individual cell is barcoded by expression of a vector-encoded RNA transcript carrying a barcode that is translocated to the nucleus. Nuclei can be subsequently isolated from cells, and the barcodes can be identified, for example, by sequencing. The vectors described herein are especially useful for high-throughput screening of genetically modified cells in pooled assays and transcriptome profiling by single nucleus RNA sequencing.

[0133] In some embodiments, the vector used for nuclear barcoding comprises: a) a first expression cassette comprising a first promoter operably linked to a first nucleotide sequence encoding a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain, wherein expression of the first nucleotide sequence results in production of the fusion protein in a cell; and b) a second expression cassette comprising a second promoter operably linked to a second nucleotide sequence encoding an RNA recognition sequence for the RNA binding domain and a barcode, wherein transcription of the second nucleotide sequence generates a barcoded RNA transcript comprising the RNA recognition sequence and the barcode, wherein binding of the RNA binding domain to the RNA recognition sequence results in formation of a complex between the fusion protein and the barcoded RNA transcript, wherein the complex is translocated to the nucleus of the cell. The nuclear localization domain may comprise a nuclear localization sequence or a nuclear-localized protein.

[0134] In some embodiments, the nuclear localization domain comprises a nuclear localization sequence of H2B such as QKKGGKKRK (SEQ ID NO:3) or KKRKRS (SEQ ID NO:4). The full- length H2B protein or a fragment thereof containing a H2B nuclear localization sequence may be included in the fusion protein such that the complex of the fusion protein and the barcoded RNA transcript is translocated to the nucleus. In some embodiments, a fragment of H2B comprising at least amino acid residues 1-52 of H2B is included in the fusion protein. For a description of putative H2B nuclear localization domains, see, e.g., Musinova et al. (2011) Biochim Biophys Acta 1813(l):27-38, Mosammaparast et al. (2001) J. Cell Biol. 153(2):251-262; herein incorporated by reference in their entireties.

[0135] In some embodiments, the nuclear localization domain is a Klarsicht-ANC-l-Syne homology (KASH) nuclear envelope localization domain. The KASH domain comprises approximately 60 amino acids, including a hydrophobic transmembrane region of about 20 aminoacids spanning the outer nuclear membrane and a 30-35-residue C-terminal region that lies between the inner and the outer nuclear membranes in the perinuclear space. In some embodiments, the nuclear localization domain of the fusion protein comprises a KASH domain comprising the amino acid sequence of SEQ ID NO: 5, or an amino acid sequence having at least about 80-100% sequence identity to the amino acid sequence of SEQ ID NO: 5, including any percent identity within this range, such as 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto, wherein the complex of the fusion protein and the barcoded RNA transcript is translocated to the nucleus. The KASH domain or a full-length KASH-domain containing protein or a fragment thereof containing the KASH domain may be included in the fusion protein such that the complex of the fusion protein and the barcoded RNA transcript is translocated to the nucleus. Exemplary mammalian KASH- domain containing proteins include nesprins such as nesprin-1, nesprin-2, nesprin-3, and nesprin- 4. For a further description of the KASH domain, see, e.g., Starr (2011) Curr. Biol. 21(11):R414- 415, Rajgor et al. (2013) Expert Rev. Mol. Med. 15:e5, and InterPro database Pfam entry PF10541 (ebi.ac.uk / interpro / entry / pfam / PF10541 / ); herein incorporated by reference in their entireties.

[0136] In certain embodiments, the RNA binding protein comprises an MS2 bacteriophage coat protein (MCP), and the RNA recognition sequence comprises an MS2 RNA aptamer, wherein the MCP binds to the MS2 RNA aptamer. The RNA recognition sequence may comprise a plurality of MS2 RNA aptamers. In some embodiments, the RNA recognition sequence comprises at least3, at least 4, at least 5, or at least 6 MS2 RNA aptamers. In some embodiments, the RNA recognition sequence comprises 1 to 30 MS2 RNA aptamers, including any number of MS2 RNA aptamers in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 MS2 RNA aptamers. In some embodiments, the RNA-binding domain comprises a plurality of RNA binding proteins. In some embodiments, the RNA-binding domain comprises multiple repeats of MCP, such as 1 to 6 MCPs, including any number of MCP in this range such as 1, 2, 3,4, 5, or 6 MCPs. For a description of MCP and its MS2 binding site, see, e.g., Tutucci et al. (2018) Nat Methods 15(1):81-89, Pichon et al. (2020) Methods Mol. Biol. 2166:121-144, Johansson et al. (1997) Seminars in Virology. 8 (3): 176-185; herein incorporated by reference in their entireties.

[0137] In certain embodiments, the RNA binding protein comprises a lambda N peptide, and the RNA recognition sequence comprises a box B sequence, wherein the lambda N peptide binds to the box B sequence. The RNA recognition sequence may comprise a plurality of box Bsequences. In some embodiments, the RNA recognition sequence comprises at least 3, at least 4, at least 5, or at least 6 box B sequences. In some embodiments, the RNA recognition sequence comprises 1 to 30 box B sequences, including any number of box B sequences in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 box B sequences. In some embodiments, the RNA binding domain comprises multiple repeats of the lambda N peptide, such as 1 to 6 lambda N peptides, including any number of lambda N peptides in this range such as 1, 2, 3, 4, 5, or 6 lambda N peptides. For a description of the lambda N peptide and its box B binding site, see, e.g., Baron-Benhamou et al. (2004) Methods Mol. Biol. 257:135-54, Lange et al. (2008) Traffic 9(8): 1256- 1267; herein incorporated by reference in their entireties.

[0138] In certain embodiments, the RNA binding protein comprises a bacteriophage PP7 coat protein (PCP), and the RNA recognition sequence comprises a PCP binding site, wherein the PCP binds to the PCP binding site. The RNA recognition sequence may comprise a plurality of PCP binding sites. In some embodiments, the RNA recognition sequence comprises at least 3, at least 4, at least 5, or at least 6 PCP binding sites. In some embodiments, the RNA recognition sequence comprises 1 to 30 PCP binding sites, including any number of PCP binding sites in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 PCP binding sites. In some embodiments, the RNA binding domain comprises multiple repeats of the PCP, such as 1 to 6 PCPs, including any number of PCPs in this range such as 1 , 2, 3, 4, 5, or 6 PCPs. For a description of the PCP and its binding site, see, e.g., Chao et al. (2008) Nat. Struct. Mol. Biol. 15(1): 103-105, Das et al. (2018) Sci. Adv. 4(6):eaar3448, Heinrich et al. (2017) RNA 23(2): 134- 141 ; herein incorporated by reference in their entireties.

[0139] In certain embodiments, the RNA binding protein comprises a BglG transcriptional antiterminator, and the RNA recognition sequence comprises a BglG binding site, wherein the BglG transcriptional antiterminator binds to the BglG binding site. The RNA recognition sequence may comprise a plurality of BglG binding sites. In some embodiments, the RNA recognition sequence comprises at least 3, at least 4, at least 5, or at least 6 BglG binding sites. In some embodiments, the RNA recognition sequence comprises 1 to 30 BglG binding sites, including any number of BglG binding sites in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 BglG binding sites. In some embodiments, the RNA binding domain comprises multiple repeats of the BglG transcriptional antiterminator, such as 1 to 6 BglG antiterminators, including any number of BglG antiterminators in this range such as 1, 2, 3, 4, 5,or 6 BglG antiterminators. For a description of the BglG transcriptional antiterminator and its binding site, see, c.g., Chen ct al. (2009) Proc. Natl. Acad. Sei. USA 106(32): 13535- 13540, Pena et al. (2022) Methods Mol. Biol. 2457:411-426; herein incorporated by reference in their entireties.

[0140] In certain embodiments, the RNA binding protein comprises U1A, and the RNA recognition sequence comprises a U1A binding site, wherein the U1A binds to the U1A binding site. The RNA recognition sequence may comprise a plurality of U1A binding sites. In some embodiments, the RNA recognition sequence comprises at least 3, at least 4, at least 5, or at least 6 U1A binding sites. In some embodiments, the RNA recognition sequence comprises 1 to 30 U1A binding sites, including any number of U1A binding sites in this range such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30 U1A binding sites. In some embodiments, the RNA binding domain comprises multiple repeats of the U1A, such as 1 to 6 U1 As, including any number of UlAs in this range such as 1, 2, 3, 4, 5, or 6 UlAs. For a description of the U1A and its binding site, see, e.g., Chung et al. (2011) Methods Mol. Biol. 714:221-235, Brodsky et al. (2002) Methods 26(2): 151155, Takizawa et al. (2000) Proc. Natl. Acad. Sci. USA 97(10):5273-5278; herein incorporated by reference in their entireties.

[0141] In certain embodiments, the second nucleotide further comprises a pair of RNA circulation sites flanking the RNA recognition sequence and the barcode in the barcoded RNA transcript, wherein the barcoded RNA transcript forms a circularized RNA transcript in a cell. Circularization of the barcoded RNA transcript eliminates the free 5’ and 3’ ends, which prevents exonuclease degradation and increases the stability of the barcoded RNA transcript in a host cell. In some embodiments, the second nucleotide sequence further encodes a pair of ribozymes comprising a first ribozyme and a second ribozyme, wherein the first ribozyme and the second ribozyme flank the RNA recognition sequence and the barcode in the barcoded RNA transcript, wherein the first ribozyme and the second ribozyme undergo autocatalytic cleavage to generate a 5-end comprising a hydroxyl group and a 3’-end comprising a 2’,3’-cyclic phosphate group, wherein the 5 ’-end and the 3 ’-end are ligated by an endogenous RtcB RNA ligase in a cell resulting in circularization of the barcoded RNA transcript. In some embodiments, the pair of ribozymes a e Twister ribozymes. For a description of RNA circularization using Twister ribozymes in a Tornado (Twister-optimized RNA for durable overexpression) expression system, see, e.g., Litke et al. (2019) Nat. Biotechnol. 37(6):667-675; herein incorporated by reference in its entirety.

[0142] In other embodiments, the vector used for nuclear barcoding comprises an expression cassette comprising a promoter operably linked to a nucleotide sequence encoding a nuclear retention element connected to a barcode, wherein transcription of the nucleotide sequence in a cell generates a barcoded RNA transcript comprising the nuclear retention element connected to the barcode, wherein the barcoded RNA transcript is translocated to the nucleus of the cell.

[0143] The nuclear retention element may include, without limitation, an exonic repeat, a C- rich motif, a U1 motif, a short interspersed nuclear element (SINE), an Alu-like element, and the like. Exemplary nuclear retention elements include those from long non-coding RNAs such as, but not limited to, MEG3, XIST, MALAT1, and SIRLOIN. Either a nuclear retention element of a long non-coding RNA, an entire long non-coding RNA, or a fragment thereof containing the nuclear retention element may be included in the barcoded RNA transcript as long as the barcoded RNA transcript is translocated to the nucleus. In some embodiments, the nuclear retention element comprises the nucleotide sequence, RCCTCCC, wherein R is an A or a G. For a further description of nuclear’ retention elements in long non-coding RNAs, see, e.g., Hasenson et al. (2022) Cells 11(12):1942, Azam et al. (2019) RNA Biol. 16(8): 1001-1009, Tong et al. (2021) RNA Biol. 18(12);2073-2086, Guo et al. (2020) Trends Biochem Sei. 45(ll):947-960, Palazzo et al. (2018) Front Genet. 9:440; Cohen et al. (2007) Chromosomal 16(4):373-83, Yin et al. (2020) Nature 580(7801): 147-150, Nguyen et al. (2020) Nucleic Acids Res 48(5):2621-2642, Lubelsky et al. (2021) EMBO J. 40(12):el06357, Lubelsky et al. (2018) Nature 555:107-111, Agostini et al. (2018) EMBO J. 37(6):e99123; herein incorporated by reference in their entireties.

[0144] In certain embodiments, the nucleotide sequence encoding the nuclear retention element connected to the barcode further encodes a pair of ribozymes comprising a first ribozyme and a second ribozyme, wherein the first ribozyme and the second ribozyme flank the nuclear retention element and the barcode in the barcoded RNA transcript, wherein the first ribozyme and the second ribozyme undergo autocatalytic cleavage to generate a 5-end comprising a hydroxyl group and a 3’-end comprising a 2’,3’-cyclic phosphate group, wherein the 5’-end and the 3’-end are ligated by an endogenous RtcB RNA ligase in a cell resulting in circularization of the barcoded RNA transcript. In some embodiments, the pair of ribozymes are Twister ribozymes.

[0145] Any suitable vector may be used for delivery of an expression cassette (e.g., encoding a fusion protein comprising a nuclear- localization domain connected to an RNA-binding domain and / or a barcoded RNA transcript) to a cell. A number of viral based systems have been developedfor gene transfer into mammalian cells. Suitable vectors include viral vectors based on adeno- associatcd virus (AAV) (sec, e.g., Li ct al. (2020) Nat Rev Genet. 21(4):255-272, Balakrishnan ct al. (2014) Curr. Gene Ther. 14(2):86-100, Ali et al., Hum Gene Ther 9:81 86, 1998, Flannery et al., PNAS 94:6916 6921, 1997; Bennett et al., Invest Opthalmol Vis Sci 38:2857 2863, 1997; Jomary et al., Gene Ther 4:683 690, 1997, Rolling et al., Hum Gene Ther 10:641 648, 1999; Ali et al., Hum Mol Genet 5:591 594, 1996; Srivastava in WO 93 / 09239, Samulski et al., J. Vir. (1989) 63:38223828; Mendelson et al., Virol. (1988) 166:154165; and Flotte et al., PNAS (1993) 90:1061310617); rabies virus (e.g., Wickersham et al. (2007) Nat Methods. 4(l):47-49, Wickersham et al. (2007) Neuron. 53(5):639-647, Hagendorf et al. (2015) Cold Spring Harb Protoc. 2015(12):pdb.prot089417, Suzuki et al. (2020) Front Neural Circuits 13:77, Osakada et al. (2013) Nat. Protoc. 8(8): 1583-601); adenovirus (see, e.g., Li et al., Invest Opthalmol Vis Sci 35:2543 2549, 1994; Borras et al., Gene Ther 6:515 524, 1999; Li and Davidson, PNAS 92:7700 7704, 1995; Sakamoto et al., H Gene Ther 5:1088 1097, 1999; WO 94 / 12649, WO 93 / 03769; WO 93 / 19191; WO 94 / 28938; WO 95 / 11984 and WO 95 / 00655); SV40; herpes simplex virus (e.g., Artusi et al. (2018) Diseases 14;6(3):74, Lachmann (2004) Int. J. Exp. Pathol. 85(4): 177-90); human immunodeficiency virus (see, e.g., Miyoshi et al., PNAS 94:10319 23, 1997; Takahashi et al., J Virol 73:78127816, 1999); vaccinia virus (e.g., Xie et al. (2022) Vaccine 40(49):7022-7031, Yang et al. (2018) I Cancer Res Clin Oncol. 144(12):2433-2440, Mackett et al. (1986) J Gen Virol. 67 ( Pt 10):2067-82); poliovirus (e.g., Girard et al. (193) Biologicals 21(4):371-7); anellovirus (Prince et al. (2024) biorxiv.org / content / 10.1101 / 2024.03.27.586964vl); and retroviruses, including lentivirus (e.g., Cockrell et al. (2007) Mol. Biotechnol. 36(3): 184-204, Wang et al. (2021) Sci China Life Sci. 64(11): 1842-1857, Lever et al. (1999) Biochem Soc Trans. 27(6):841- 7), y- retro virus such as murine leukemia virus and feline leukemia virus, an avian retrovirus such as spleen necrosis virus, and vectors derived from retroviruses such as Rous Sarcoma Virus, Harvey Sarcoma Virus, avian leukosis virus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus; and the like. See also, e.g., Warnock et al. (2011) Methods Mol. Biol. 737:1-25; Walther et al. (2000) Drugs 60(2):249-271; and Lundstrom (2003) Trends Biotechnol. 21(3): 117-122; herein incorporated by reference.

[0146] In some embodiments, an adeno-associated virus (AAV) vector is used for delivery of the expression cassettes (e.g., expression cassette encoding a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain and / or an expression cassette encodinga barcoded RNA transcript). The components of the AAV DNA genome consists of two open reading frames, Rep and Cap, flanked by two 145 base inverted terminal repeats (ITRs). The Rep and Cap regions are translated to produce multiple distinct proteins including Rep78, Rep68, Rep52, and Rep40 and the capsid proteins VP1, VP2, and VP3 required for production of rAAV virions. In addition, AAV requires a helper plasmid containing genes from a helper virus such as adenovirus, including Ela, Elb, E4, E2a, and VA genes for AAV replication. Expression cassettes (e.g., encoding fusion protein and / or barcoded RNA transcript) can be inserted between the 5’- ITR and 3’-ITR. The structural (cap) and packaging (rep) proteins can be delivered in trans. Various AAV vector systems are available, and AAV vectors can be readily constructed using techniques well known in the art. See, e.g., U.S. Pat. Nos. 5,173,414 and 5,139,941; International Publication Nos. WO 92 / 01070 (published 23 January 1992) and WO 93 / 03769 (published 4 March 1993); Lebkowski et al., Molec. Cell. Biol. (1988) 8:3988-3996; Vincent et al., Vaccines 90 (1990) (Cold Spring Harbor Laboratory Press); Carter, B. J. Current Opinion in Biotechnology (1992) 3:533-539; Muzyczka, N. Current Topics in Microbiol, and Immunol. (1992) 158:97-129; Kotin, R. M. Human Gene Therapy (1994) 5:793-801; Shelling and Smith, Gene Therapy (1994) 1:165-169; Zhou et al., J. Exp. Med. (1994) 179:1867-1875; Li et al. (2020) Nat Rev Genet. 21(4):255-272; and Balakrishnan et al. (2014) Curr. Gene Ther. 14(2):86-100; herein incorporated by reference.

[0147] In some embodiments, a G-deleted rabies virus (RVdG) is used for delivery of the expression cassettes (e.g., expression cassette encoding a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain and / or an expression cassette encoding a barcoded RNA transcript). RVdG is an enveloped, non-segmented negative-stranded RNA virus, genetically modified to lack the envelope glycoprotein gene (G gene). The glycoprotein is not required for the transcription or replication of the genome within infected cells but is required for RVdG to spread from cell to cell. Thus, the viral core of RVdG can proliferate within infected cells, but the inability to synthesize the glycoprotein prevents its progeny from infecting other cells. RVdG vector can be used to infect neurons to allow synaptic tracing. RVdG infects neurons through their axon terminals and spreads between synaptically coupled neurons exclusively in the retrograde direction. This feature of the virus makes RVdG useful for retrograde synaptic tracing (see, e.g., Wickersham et al. (2007) Nat Methods. 4(l):47-49, Wickersham et al. (2007) Neuron. 53(5):639-647, Hagendorf et al. (2015) Cold Spring Harb Protoc. 2015(12):pdb.prot089417,Suzuki et al. (2020) Front Neural Circuits 13:77, Osakada et al. (2013) Nat. Protoc. 8(8): 1583- 601; herein incorporated by reference in their entireties). In some embodiments, the G gene is replaced with a gene encoding a fluorescent protein to allow fluorescence imaging of cells infected with the virus. In some embodiments, the G gene is provided in trans to allow the virus to spread from cell to cell. The G gene can be inserted suitably for use in a research or therapeutic applications or gene therapy.

[0148] Other vectors may also be used for delivery of the expression cassettes such as retroviral vectors. Commonly used retroviral vectors are “replication defective”, i.e., unable to produce viral proteins required for productive infection. Rather, replication of the vector requires growth in a packaging cell line. To generate viral particles comprising nucleic acids of interest, the retroviral nucleic acids comprising the nucleic acid are packaged into viral capsids by a packaging cell line. Different packaging cell lines provide a different envelope protein (ecotropic, amphotropic or xenotropic) to be incorporated into the capsid, this envelope protein determining the specificity of the viral particle for the cells (ecotropic for murine and rat; amphotropic for most mammalian cell types including human, dog and mouse; and xenotropic for most mammalian cell types except murine cells). The appropriate packaging cell line may be used to ensure that the cells are targeted by the packaged viral particles. Methods of introducing expression vectors into packaging cell lines and of collecting the viral particles that are generated by the packaging lines are well known in the ail (see, e.g., Kafri et al. (2004) Methods Mol Biol. 246:367-390, herein incorporated by reference).

[0149] Lentiviruses belong to the genus of retroviruses, including human immunodeficiency virus, and are enveloped viruses comprising a nucleocapsid containing two copies of singlestranded positive-sense RNA. The genomic RNA undergoes reverse transcription catalyzed by a viral reverse transcriptase to produce a DNA copy of the RNA genome. Lentivirus vectors can infect dividing or nondividing cells and DNA produced by reverse transcription can integrate into the genome of a host cell to provide stable transgene expression.

[0150] Commonly used lentiviral vectors a e based on the genome of human immunodeficiency virus type 1 (HIV-1), but other lentiviral vectors derived from human immunodeficiency virus type 2 (HIV-2), simian immunodeficiency virus, and nonprimate lentiviruses such as equine infectious anemia virus, feline immunodeficiency vims, and bovine immunodeficiency vims can also be used for gene delivery. Due to safety concerns becauseof the pathogenic nature of HIV in humans, lentivirus vectors are generally “replication defective” and arc genetically modified to prevent disease. In third-generation lentiviral vector systems (or higher), virus production is split across four or more plasmids. Current third generation lentiviral vectors encode only three of the nine HIV-1 proteins (group specific antigen (Gag), polymerase (Pol), and regulator of expression of virion proteins (Rev)), which are expressed from separate plasmids to avoid recombination-mediated generation of a replication-competent virus. The envelope protein may be replaced by the more stable VSV-G protein, which allows entry into more cell types. For a further description of lentivirus vectors, see, e.g., Ghaleh et al. (2020) Biomed. Pharmacother. 128:110276, Shearer et al. (2015) Genes Cells 20(1): 1-10, White et al. (2017) Hum Gene Ther Methods 28(4):163-176, Berkhout (2017) Mol Ther. 25(8): 1741- 1743, Kafri et al. (2001) Curr. Opin. Mol. Ther. 3(4):316-26, Cockrell et al. (2007) Mol. Biotechnol. 36(3): 184-204, Wang et al. (2021) Sei China Life Sei. 64(11): 1842- 1857, Lever et al. (1999) Biochem Soc Trans. 27(6):841-847.

[0151] A number of adenovirus vectors have also been described. Unlike retroviruses which integrate into the host genome, adenoviruses persist extrachromosomally thus minimizing the risks associated with insertional mutagenesis (Haj-Ahmad and Graham, J. Virol. (1986) 57:267-274; Bett et al., J. Virol. (1993) 67:5911-5921; Mittereder et al., Human Gene Therapy (1994) 5:717- 729; Seth et al., J. Virol. (1994) 68:933-940; Barr et al., Gene Therapy (1994) 1:51-58; Berkner, K. L. BioTechniques (1988) 6:616-629; and Rich et al., Human Gene Therapy (1993) 4:461-476). Molecular conjugate vectors, such as the adenovirus chimeric vectors described in Michael et al., J. Biol. Chem. (1993) 268:6866-6869 and Wagner et al., Proc. Natl. Acad. Sci. USA (1992) 89:6099-6103, can also be used for gene delivery.

[0152] Another vector system useful for delivering expression cassettes is the enterically administered recombinant poxvirus vaccines described by Small, Jr., P. A., et al. (U.S. Pat. No. 5,676,950, issued Oct. 14, 1997, herein incorporated by reference). Alternatively, avipoxviruses, such as the fowlpox and canarypox viruses, can also be used to deliver the expression cassettes. The use of an avipox vector is suitable for human and other mammalian species since members of the avipox genus can only productively replicate in susceptible avian species and therefore are not infective in mammalian cells. Methods for producing recombinant avipoxviruses are known in the art and employ genetic recombination. See, e.g., WO 91 / 12882; WO 89 / 03429; and WO 92 / 03545.

[0153] Members of the Alphavirus genus, such as, but not limited to, vectors derived from the Sindbis virus (SIN), Scmliki Forest virus (SFV), and Venezuelan Equine Encephalitis virus (VEE), can be used as viral vectors for delivering the expression cassettes. For a description of Sindbis- virus derived vectors useful for the practice of the instant methods, see, Dubensky et al. (1996) J. Virol. 70:508-519; and International Publication Nos. WO 95 / 07995, WO 96 / 17072; as well as Dubensky, Jr., T. W., et al., U.S. Pat. No. 5,843,723, issued Dec. 1, 1998, and Dubensky, Jr., T. W., U.S. Patent No. 5,789,245, issued Aug. 4, 1998, both herein incorporated by reference. Particularly preferred are chimeric alphavirus vectors comprised of sequences derived from Sindbis virus and Venezuelan equine encephalitis virus. See, e.g., Perri et al. (2003) J. Virol. 77: 10394-10403 and International Publication Nos. WO 02 / 099035, WO 02 / 080982, WO 01 / 81609, and WO 00 / 61772; herein incorporated by reference in their entireties.

[0154] A vaccinia-based infection / transfection system can be conveniently used to provide for inducible, transient expression of the expression cassettes in a host cell. In this system, cells are first infected in vitro with a vaccinia virus recombinant that encodes the bacteriophage T7 RNA polymerase. This polymerase only transcribes templates bearing T7 promoters. Following infection, cells are transfected with the polynucleotide of interest, driven by a T7 promoter. The polymerase expressed in the cytoplasm from the vaccinia virus recombinant transcribes the transfected DNA into RNA which may be translated into protein by the host translational machinery. The method provides for high level, transient, cytoplasmic production of large quantities of RNA and / or its translation products. See, e.g., Elroy-Stein and Moss, Proc. Natl. Acad. Sci. USA (1990) 87:6743-6747; Fuerst et al., Proc. Natl. Acad. Sci. USA (1986) 83:8122-8126.

[0155] In addition, anelloviral vectors can be used for gene delivery. An anellovector, based on a virus of the Betatorquevirus genus, has been developed (see, e.g., Prince et al. (2024) (biorxiv.org / content / 10.1101 / 2024.03.27.586964vl). The vector comprises a self-amplifying trans -complementation of a universal recombinant anellovector (SATURN) system, which relies on a self-replicating plasmid to provide viral proteins in trans that drive replication and capsiddependent packaging of vector genomes. The SATURN system uses Cre-lox-based recombination to generate single unit-sized circular genomes inside a MOLT-4 production cell line. Capsid protein-dependent particles that encapsidate single stranded DNA vector genomes can be produced using the SATURN system.

[0156] A vector may be provided directly to a target host cell, for example, by contacting the host cell with the vector such that the vector is taken up by the cells. Methods of transfecting cells are well known in the art, and include, without limitation, electroporation, calcium chloride transfection, microinjection, and lipofection. For viral vector delivery, cells can be contacted with viral particles comprising viral expression vectors.

[0157] Nucleic acids encoding, for example, a nuclear localization domain or nuclear retention element, an RNA-binding domain, an RNA recognition sequence for the RNA binding domain, or a barcode a can be inserted into an expression vector to create an expression cassette capable of producing a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain, a barcoded RNA transcript comprising an RNA recognition sequence and a barcode, or a barcoded RNA transcript comprising a nuclear retention element connected to a barcode in a suitable host cell. The ability of constructs to produce a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain, a barcoded RNA transcript comprising an RNA recognition sequence and a barcode, or a barcoded RNA transcript comprising a nuclear retention element connected to a barcode can be empirically determined.

[0158] Expression cassettes typically include control elements operably linked to a coding sequence, which allow for the expression of a gene in vivo in the subject species. Depending on the host / vector system utilized, any of a number of suitable transcription and translation control elements, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc. may be used in the expression vector.

[0159] Promoters can be used to drive expression by an RNA polymerase (e.g., pol I, pol II, pol III). Suitable promoters can be derived from viruses (i.e., viral promoters) or an organism, including prokaryotic or eukaryotic organisms. Exemplary promoters include, but are not limited to the SV40 early promoter, mouse mammary tumor virus long terminal repeat (LTR) promoter, adenovirus major late promoter (Ad MLP), herpes simplex virus (HSV) promoter, cytomegalovirus (CMV) promoter such as the CMV immediate early promoter region (CM VIE), hybrid cytomegalovirus (CMV) / chicken -actin promoter (CAG) promoter, Rous sarcoma virus (RSV) promoter, human U6 small nuclear promoter (U6) (Miyagishi et al., Nature Biotechnology 20, 497-500 (2002)), enhanced U6 promoter (e.g., Xia et al., Nucleic Acids Res. 2003 Sep. 1; 31(17)), and human Hl promoter (Hl), and the like.

[0160] The promoter can be a constitutively active promoter (i.e., a promoter that is constitutively in an activc / “ON” state) or an inducible promoter (i.e., a promoter whose state, active / “ON” or inactive / “OFF” is controlled by an external stimulus, e.g., the presence of a particular temperature, compound, or protein). In some cases, a promoter is a spatially restricted promoter (e.g., tissue- specific promoter or cell type-specific promoter controlled by a transcriptional control element, enhancer, etc.). In some cases, a promoter is a temporally restricted promoter (i.e., the promoter is in the “ON” state or “OFF” state during specific stages of embryonic development or during specific stages of a biological process).

[0161] Inducible promoters suitable for use include any inducible promoter described herein or known to one of ordinary skill in the art. Examples of inducible promoters include, without limitation, chemically / biochemically-regulated and physically-regulated promoters such as alcohol-regulated promoters, tetracycline-regulated promoters (e.g., anhydrotetracycline (aTc)- responsive promoters and other tetracycline-responsive promoter systems, which include a tetracycline repressor protein (tetR), a tetracycline operator sequence (tetO) and a tetracycline transactivator fusion protein (tTA)), steroid-regulated promoters (e.g., promoters based on the rat glucocorticoid receptor, human estrogen receptor, moth ecdysone receptors, and promoters from the steroid / retinoid / thyroid receptor superfamily), metal-regulated promoters (e.g., promoters derived from metallothionein (proteins that bind and sequester metal ions) genes from yeast, mouse and human), pathogenesis-regulated promoters (e.g., induced by salicylic acid, ethylene or benzothiadiazole (BTH)), temperature / heat-inducible promoters (e.g., heat shock promoters), and light-regulated promoters (e.g., light responsive promoters from plant cells).

[0162] In some cases, the promoter is a spatially restricted promoter (i.e., cell type-specific promoter, tissue-specific promoter, organ-specific, etc.) such that in a multi-cellular organism, the promoter is active (i.e., “ON”) in a subset of specific cells. Spatially restricted promoters may be regulated by enhancers, transcriptional control elements, control sequences, etc. Any convenient spatially restricted promoter may be used as long as the promoter is functional in the targeted host cell (e.g., eukaryotic cell). In some cases, the promoter is a tissue-specific promoter. In some cases, the promoter is a cell type-specific promoter. In some cases, the transcriptional control element (e.g., the promoter) is functional in a targeted cell type or targeted cell population. For example, in some cases, the transcriptional control element can be functional in a muscle cell (e.g., a cardiac muscle cell (cardiomyocyte), a skeletal muscle cell (skeletal myofiber), or a smooth muscle cell),a neuron, a retinal cell, a T cell, a B cell, a hematopoietic stem cell, a liver cell, a lung cell, or other targeted cell. In some cases, the transcriptional control clement is functional in a postmitotic cell or non-dividing cell such as, but not limited to, a neuron, a cardiomyocyte, a skeletal muscle myofiber, a retinal ganglion cell, a cochlear hair cell, an osteocyte, or an adipocyte.

[0163] In some cases, the promoter is a reversible promoter. Suitable reversible promoters, including reversible inducible promoters are known in the art. Such reversible promoters may be isolated and derived from any of a variety of organisms. Modification of reversible promoters derived from a first organism for use in a second (different) organism is well known in the ail. Such reversible promoters, and systems based on such reversible promoters but also comprising additional control proteins, include, but are not limited to, alcohol regulated promoters (e.g., alcohol dehydrogenase I (alcA) gene promoter, promoters responsive to alcohol transactivator proteins (AlcR), etc.), tetracycline regulated promoters, (e.g., promoter systems including TetActivators, TetON, TetOFF, etc.), steroid regulated promoters (e.g., rat glucocorticoid receptor promoter systems, human estrogen receptor promoter systems, retinoid promoter systems, thyroid promoter systems, ecdysone promoter systems, mifepristone promoter systems, etc.), metal regulated promoters (e.g., metallothionein promoter systems, etc.), pathogenesis-related regulated promoters (e.g., salicylic acid regulated promoters, ethylene regulated promoters, benzothiadiazole regulated promoters, etc.), temperature regulated promoters (e.g., heat shock inducible promoters (e.g., HSP-70, HSP-90, soybean heat shock promoter, etc.), light regulated promoters, synthetic inducible promoters, and the like. A suitable promoter can include elements that are responsive to transactivation, e.g., hypoxia response elements, Gal4 response elements, lac repressor response element, and small molecule control systems such as tetracycline-regulated systems and the RU- 486 system (see, e.g., Gossen & Bujard, 1992, Proc. Natl. Acad. Sei. USA, 89:5547; Oligino et al., 1998, Gene Ther., 5:491-496; Wang et al., 1997, Gene Ther., 4:432-441; Neering et al., 1996, Blood, 88:1147-55; and Rendahl et al., 1998, Nat. Biotechnol., 16:757-761).

[0164] For illustration purposes, examples of spatially restricted promoters include, but are not limited to, neuron- specific promoters, cardiomyocyte-specific promoters, skeletal muscle-specific promoters, smooth muscle-specific promoters, photoreceptor- specific promoters, retinal ganglion cell-specific promoters, adipocyte- specific promoters, etc.

[0165] In some embodiments, the promoter is a neuron- specific promoter. Examples of neuron- specific promoters include, but are not limited to, a neuron-specific enolase (NSE)promoter (see, e.g., EMBL HSENO2, X51956; see also, e.g., U.S. Pat. No. 6,649,81 1 , U.S. Pat. No. 5,387,742); an aromatic amino acid decarboxylase (AADC) promoter; a neuro filament promoter (see, e.g., GenBank HUMNFL, L04147); a synapsin promoter (see, e.g., GenBank HUMSYNIB, M55301); a thy-1 promoter (see, e.g., Chen et al. (1987) Cell 51:7-19; and Llewellyn et al. (2010) Nat. Med. 16:1161); a serotonin receptor promoter (see, e.g., GenBank S62283); a tyrosine hydroxylase promoter (TH) (see, e.g., Nucl. Acids. Res. 15:2363-2384 (1987) and Neuron 6:583-594 (1991)); a GnRH promoter (see, e.g., Radovick et al., Proc. Natl. Acad. Sci. USA 88:3402-3406 (1991)); an L7 promoter (see, e.g., Oberdick et al., Science 248:223-226 (1990)); a DNMT promoter (see, e.g., Bartge et al., Proc. Natl. Acad. Sci. USA 85:3648-3652 (1988)); an enkephalin promoter (see, e.g., Comb et al., EMBO J. 17:3793-3805 (1988)); a myelin basic protein (MBP) promoter; a CMV enhancer / platelet-derived growth factor- .beta, promoter (see, e.g., Liu et al. (2620) Gene Therapy 11:52-60); a motor neuron- specific gene Hb9 promoter (see, e.g., U.S. Pat. No. 7,632,679; and Lee et al. (2620) Development 131:3295-3306); an alpha subunit of Ca2+-calmodulin-dependent protein kinase II (CaMKII) promoter (see, e.g., Mayford et al. (1996) Proc. Natl. Acad. Sci. USA 93:13250), and a retinal ganglion cell Nefh promoter (see, e.g., Hanlon et al. (2017) Front Neurosci. 11:521). Other suitable promoters include elongation factor (EF) 1 and dopamine transporter (DAT) promoters, and the like.

[0166] In some embodiments, the promoter is a cardiomyocyte- specific promoter. Examples of cardiomyocyte-specific promoters include, but are not limited to, a cardiac muscle-specific alpha myosin heavy chain (MHC) gene promoter (see, e.g., Gulick et al. (1991) J. Biol. Chem. 266:9180-9185, Aikawa et al. (2002) J. Biol. Chem. 277(21): 18979-18985). a ventricle-specific cardiac myosin light chain 2 (MLC-2v) promoter (see, e.g., Boecker et al. (2004) Mol. Imaging 3(2):69-75, Griscelli et al. (1997) C R Acad. Sci. Ill 320(2): 103-12), a cardiac troponin T (cTNT) promoter (see, e.g., Ai et al. (2018) Cell Physiol. Biochem. 48(5): 1894-1900), a troponin 2 (TNNT2) promoter (see, e.g., Fiedorowicz et al. (2020) Sci. Rep.10(1): 1895), an alpha cardiac actin (ACTC) promoter (see, e.g., Fiedorowicz et al., supra), and a cardiac ankyrin repeat protein gene (Carp / Ankrdl) promoter (see, e.g., Briegel et al. (2005) Development 132(14):3305-16).

[0167] In some embodiments, the promoter is a skeletal muscle-specific promoter. Examples of skeletal muscle-specific promoters include, but are not limited to, a skeletal muscle a-actin promoter, creatine kinase promoter, desmin promoter, troponin promoter, myosin light chain promoter, myosin heavy chain promoter, dystrophin promoter, and Pitx3 promoter (see, e.g., (see,e.g., Skopenkova et al. (2021) Acta Naturae 13(1): 47-58, Coulon et al. (2007) J. Biol. Chem. 282(45):33192-33200, Sartorelli et al. (1993) Circ. Res. 72(5):925-931).

[0168] In some embodiments, cell subtype-specific expression is achieved by using a recombination system, e.g., Cre-Lox recombination, Flp-FRT recombination, etc. Cell typespecific expression of genes using recombination has been described in, e.g., Fenno et al., Nat Methods, 2014 July; 11(7):763; Gompf et al., Front. Behav. Neurosci. 2015 Jul. 2;9:152, and McCarthy et al. (2012) Skelet. Muscle. 2(1):8; which are herein incorporated by reference.

[0169] Typically, transcription termination and polyadenylation sequences will also be present, located 3' to the translation stop codon. Preferably, a sequence for optimization of initiation of translation, located 5' to the coding sequence, is also present. Examples of transcription terminator / polyadenylation signals include those derived from SV40, as described in Sambrook et al., supra, as well as a bovine growth hormone terminator sequence.

[0170] Enhancer elements may also be used herein to increase expression levels of the mammalian constructs. Examples include the SV40 early gene enhancer, as described in Dijkema et al., EMPO J. (1985) 4:761, the enhancer / promoter derived from the long terminal repeat (LTR) of the Rous Sarcoma Virus, as described in Gorman et al., Proc. Natl. Acad. Sci. USA (1982b) 79:6777 and elements derived from human CMV, as described in Boshart et al., Cell (1985) 41:521, such as elements included in the CMV intron A sequence.

[0171] Additionally, 5'- UTR sequences can be placed adjacent to the coding sequence in order to enhance expression of the same. Such sequences may include UTRs comprising an internal ribosome entry site (IRES). Inclusion of an IRES permits the translation of one or more open reading frames from a vector. For example, a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain can be co-expressed with a therapeutic protein from a multicistronic vector including an IRES element. The IRES element attracts a eukaryotic ribosomal translation initiation complex and promotes translation initiation. See, e.g., Kaufman et al., Nuc. Acids Res. (1991) 19:4485-4490; Gurtu et al., Biochem. Biophys. Res. Comm. (1996) 229:295-298; Rees et al., BioTechniques (1996) 20:102-110; Kobayashi et al., BioTechniques (1996) 21:399-402: and Mosser et al., BioTechniques (1997) 22 150-161. A multitude of IRES sequences are known and include sequences derived from a wide variety of viruses, such as from leader sequences of picornaviruses such as the encephalomyocarditis virus (EMCV) UTR (Jang et al. J. Virol. (1989) 63:1651-1660), the polio leader sequence, the hepatitis A virus leader, thehepatitis C virus IRES, human rhinovirus type 2 IRES (Dobrikova et ah, Proc. Natl. Acad. Sci. (2003) 100(25): 15125- 15130), an IRES element from the foot and mouth disease virus (Ramcsh et al., Nucl. Acid Res. (1996) 24:2697-2700), a giardiavirus IRES (Garlapati et al., J. Biol. Chem. (2004) 279(5):3389-3397), and the like. A variety of nonviral IRES sequences will also find use herein, including, but not limited to IRES sequences from yeast, as well as the human angiotensin II type 1 receptor IRES (Martin et al., Mol. Cell Endocrinol. (2003) 212:51-61), fibroblast growth factor IRESs (FGF-1 IRES and FGF-2 IRES, Martineau et al. (2004) Mol. Cell. Biol. 24(17):7622- 7635), vascular endothelial growth factor IRES (Baranick et al. (2008) Proc. Natl. Acad. Sci. U.S.A. 105(12):4733-4738, Stein et al. (1998) Mol. Cell. Biol. 18(6):3112-3119, Bert et al. (2006) RNA 12(6): 1074-1083), and insulin-like growth factor 2 IRES (Pedersen et al. (2002) Biochem. J. 363(Pt l):37-44). These elements are readily commercially available in plasmids sold, e.g., by Clontech (Mountain View, CA), Invivogen (San Diego, CA), Addgene (Cambridge, MA) and GeneCopoeia (Rockville, MD). See also IRESite: The database of experimentally verified IRES structures (iresite.org). An IRES sequence may be included in a vector, for example, to express multiple protein products in combination.

[0172] Alternatively, a polynucleotide encoding a viral T2A peptide can be used to allow production of multiple protein products (e.g., fusion protein and therapeutic protein) from a single vector. 2A linker peptides are inserted between the coding sequences in the multicistronic construct. The 2A peptide, which is self-cleaving, allows co-expressed proteins from the multicistronic construct to be produced at equimolar levels. 2A peptides from various viruses may be used, including, but not limited to 2A peptides derived from the foot-and-mouth disease virus, equine rhinitis A virus, Thosea asigna virus and porcine teschovirus-1. See, e.g., Kim et al. (2011) PLoS One 6(4):el8556, Trichas et al. (2008) BMC Biol. 6:40, Provost et al. (2007) Genesis 45(10):625-629, Furler et al. (2001) Gene Ther. 8(11):864-873; herein incorporated by reference in their entireties.

[0173] In certain embodiments, the barcode is located in a 3’ untranslated region (3’ UTR) of a barcoded RNA transcript. In some embodiments, the barcode is transcribed by a separate U6 promoter and DNA polymerase III.

[0174] In certain embodiments, cells containing a construct comprising the expression cassettes are identified in vitro or in vivo by including a selection marker in the construct. Selection markers confer an identifiable change to the cell permitting positive selection of cells having theconstruct. For example, fluorescent markers such as, but not limited to, green fluorescent protein (GFP), enhanced green fluorescent protein (EGFP), turboGFP, yellow fluorescent protein, enhanced yellow fluorescent protein, blue fluorescent protein, cyan fluorescent protein, cerulean, oScarlet, mCherry, mOrange, mPlum, Venus, mVenus, tdTomato, Dendra2, dsRed, YPet, and phycoerythrin; bioluminescent markers such as but not limited to, luciferase (e.g., bacterial, firefly, click beetle and the like), luciferin, aequorin and the like; enzyme systems having visually detectable signals such as, but are not limited to, galactosidases, glucorimidases, phosphatases, peroxidases, cholinesterases and the like; cell surface markers; expression of a reporter gene (e.g., GFP, GUS, lacZ, CAT); or drug selection markers such as genes that confer resistance to neomycin, puromycin, hygromycin, DHFR, GPT, zeocin, or histidinol may be used to identify cells. Alternatively, enzymes such as herpes simplex virus thymidine kinase (tk) or chloramphenicol acetyltransferase (CAT) may be employed. Any selectable marker may be used as long as it is capable of being expressed in the cell to allow identification of cells containing the construct. Further examples of selectable markers are well known to one of skill in the art.

[0175] In certain embodiments, a selection marker expression cassette encodes two or more selection markers. Selection markers may be used in combination, for example, a cell surface marker may be used with a fluorescent marker, or a drug resistance gene may be used with a suicide gene. In certain embodiments, the selection marker expression cassette is multicistronic to allow expression of multiple selection markers in combination. The multicistronic vector may include an IRES or viral 2A peptide to allow expression of more than one selection marker from a single vector.

[0176] In certain embodiments, a suicide marker is included as a negative selection marker to facilitate negative selection of cells. Suicide genes can be used to selectively kill cells by inducing apoptosis or converting a nontoxic drug to a toxic compound in genetically modified cells. Examples include suicide genes encoding thymidine kinases, cytosine deaminases, intracellular antibodies, telomerases, caspases, and DNases. In certain embodiments, a suicide gene is used in combination with one or more other selection markers, such as those described above for use in positive selection of cells. In addition, a suicide gene may be used in cells containing constructs expressing ZIP7 and / or Rpnll, for example, to improve their safety by allowing their destruction at will. See, e.g., Jones et al. (2014) Front. Pharmocol. 5:254, Mitsui et al. (2017) Mol. Ther.Methods Clin. Dev. 5:51-58, Greco et al. (2015) Front. Pharmacol. 6:95; herein incorporated by reference.

[0177] In certain embodiments, a selection marker such as a fluorescent or bioluminescent protein is included in the fusion protein comprising the nuclear localization domain connected to the RNA-binding domain. In some embodiments, a fluorescent protein is positioned between the nuclear localization domain and the RNA-binding domain in the fusion protein.

[0178] Once complete, vectors comprising the expression cassettes can be administered to a cell, population of cells, a tissue, organ, or subject using standard gene delivery protocols. Methods for gene delivery are known in the ail. See, e.g., U.S. Pat. Nos. 5,399,346, 5,580,859, 5,589,466. Genes can be delivered either directly to a subject or, alternatively, delivered ex vivo, to cells derived from the subject and the cells reimplanted in the subject. For example, methods for the ex vivo delivery and reimplantation of transformed cells into a subject are known in the art and can include, e.g., dextran-mediated transfection, calcium phosphate precipitation, polybrene mediated transfection, lipofectamine and LT-1 mediated transfection, protoplast fusion, electroporation, encapsulation of the polynucleotide(s) in liposomes, and direct microinjection of the DNA into nuclei.

[0179] Direct delivery of synthetic expression cassette compositions in vivo will generally be accomplished with a viral vector, as described above, by injection using either a conventional syringe, needless devices such as Bioject™ or a gene gun, such as the Accell™ gene delivery system (PowderMed Ltd, Oxford, England).Production of Recombinant Virions

[0180] The present disclosure further provides host cells comprising the vectors described herein. A subject host cell can be an isolated cell, e.g., a cell in in vitro culture or a cell in an organism, organ, or tissue. A subject host cell is useful for producing recombinant virions, as described below. If a subject host cell is used to produce recombinant virions, it is referred to as a “packaging cell.” In some cases, a subject host cell is stably genetically modified with a vector. In other cases, a subject host cell is transiently genetically modified with a vector. The vectors described herein can be used in a variety of host cells for recombinant virion production.

[0181] In some embodiments, a vector system is provided, comprising a helper virus vector in addition to the vector for production of barcoded RNA transcripts that are translocated to thenucleus of a cell, as described herein. For example, for production of AAV virions, a helper virus vector encoding El a and Elb or E2a and E4 may be included in the vector system. In addition, the vector system may comprise expression cassettes encoding the AAV Rep proteins and the capsid proteins, VP1, VP2, and VP3. For production of RVdG virions, a vector encoding a glycoprotein (G) gene may be included in the vector system to allow the virions to spread from cell to cell or excluded to prevent such spreading. For production of lentivirus virions, one or more vectors encoding Gag, Pol, Rev, and VSV-G may be included in the vector system.

[0182] Suitable host cells that are transfected with a vector system are rendered capable of producing recombinant virions. Vectors of a vector system can be introduced into a host cell, either simultaneously or serially, using established transfection techniques, including, but not limited to, electroporation, calcium phosphate precipitation, liposome-mediated transfection, lipid nanoparticle (LNP) -mediated transfection, and the like. In some embodiments, vectors for producing recombinant virions are introduced into a host cell, and a vector comprising an expressible sequence encoding a gene product of interest is introduced later when production of the gene product of interest is desired.

[0183] A subject host cell is generated by introducing a vector or vector system into any of a variety of cells. In certain embodiments, the cell is present in a population of cells. In certain embodiments, the population of cells includes a plurality of cell types. For example, a population of cells from nervous tissue may include excitatory neurons, inhibitory neurons, and non-neuronal cells. Cells can be in an organism, organ, or tissue. The cells may include a single cell type derived from an organism or can be a mixture of cell types. Included are naturally occurring cells and cell populations, genetically engineered cell lines, cells derived from transgenic animals, etc. The cells may be of any cell type or size. Suitable cells include fungal, plant, and animal cells. In some embodiments, the cells are mammalian cells and may contain, for example, complex cell populations such as are naturally occurring in tissues or organs, for example, blood, liver, pancreas, lung, kidney, stomach, intestine, bladder, heart, brain, muscle, nervous tissue, bone marrow, skin, and the like. Tissue may be intact or disrupted into a monodisperse suspension. Alternatively, the cells may be a cultured population, e.g. a culture derived from a complex population, a culture derived from a single cell type where the cells have differentiated into multiple lineages, or where the cells are responding differentially to a stimulus, and the like.

[0184] Cell types that can find use in the subject methods include stem and progenitor cells, c.g. embryonic stem cells, hematopoietic stem cells, mesenchymal stem cells, neural crest cells, etc., endothelial cells, muscle cells, myocardial, smooth and skeletal muscle cells, mesenchymal cells, epithelial cells; hematopoietic cells, such as lymphocytes, including T-cells, such as Thl T cells, Th2 T cells, ThO T cells, cytotoxic T cells; B cells, pre- B cells, etc.; monocytes; dendritic cells; neutrophils; and macrophages; natural killer cells; mast cells, etc.; adipocytes, cells involved with particular organs, such as thymus, endocrine glands, pancreas, or brain, such as neurons, glia, astrocytes, dendrocytes, etc. and genetically modified cells thereof. Hematopoietic cells may be associated with inflammatory processes, autoimmune diseases, etc., endothelial cells, smooth muscle cells, myocardial cells, etc. may be associated with cardiovascular diseases; almost any type of cell may be associated with neoplasias, such as sarcomas, carcinomas and lymphomas; liver diseases with hepatic cells; kidney diseases with kidney cells; etc.

[0185] The cells may also be transformed or neoplastic cells of different types, e.g. carcinomas of different cell origins, lymphomas of different cell types, etc. The American Type Culture Collection (Manassas, VA) has collected and makes available over 4,000 cell lines from over 150 different species, over 950 cancer cell lines including 700 human cancer cell lines. The National Cancer Institute has compiled clinical, biochemical and molecular data from a large panel of human tumor cell lines, these are available from ATCC or the NCI (Phelps et al. (1996) Journal of Cellular Biochemistry Supplement 24:32-91). Included are different cell lines derived spontaneously, or selected for desired growth or response characteristics from an individual cell line; and may include multiple cell lines derived from a similar tumor type but from distinct patients or sites.

[0186] Cells may be non-adherent, e.g. blood cells including monocytes, T cells, B-cells; tumor cells, etc., or adherent cells, e.g. epithelial cells, endothelial cells, neural cells, etc. In order to profile adherent cells, they may be dissociated from the substrate that they are adhered to, and from other cells.

[0187] Such cells can be acquired from an individual using, e.g., a draw, a lavage, a wash, surgical dissection, biopsy, etc., from a variety of tissues, e.g., blood, marrow, a solid tissue (e.g., a solid tumor), ascites, by a variety of techniques that are known in the art. Cells may be obtained from fixed or unfixed, fresh or frozen, whole or disaggregated samples. Disaggregation of tissue may occur either mechanically or enzymatically using known techniques.

[0188] Suitable mammalian cells include, but are not limited to, primary cells and cell lines, where suitable cell lines include, but arc not limited to, 293 cells, COS cells, HcLa cells, Vcro cells, 3T3 mouse fibroblasts, C3H10T1 / 2 fibroblasts, baby hamster kidney (BHK) fibroblasts, Chinese hamster ovarian (CHO) cells, and the like. Non-limiting examples of suitable host cells include, e.g., HeLa cells (e.g., American Type Culture Collection (ATCC) No. CCL-2), CHO cells (e.g., ATCC Nos. CRL9618, CCL61, CRL9096), 293 cells (e.g., ATCC No. CRL-1573), Vero cells, NIH 3T3 cells (e.g., ATCC No. CRL-1658), Huh-7 cells, BHK cells (e.g., ATCC No. CCL10), PC12 cells (ATCC No. CRL1721), COS cells, COS-7 cells (ATCC No. CRL1651), RATI cells, mouse L cells (ATCC No. CCLI.3), human embryonic kidney (HEK) cells (ATCC No. CRL1573), HEK293T cells (ATCC CRL-3216), HLHepG2 cells, and the like. A subject host cell can also be made using a baculovirus to infect insect cells such as Sf9 cells, which produce AAV (see, e.g., U.S. Patent Nos. 7,271,002 and 8,945,918).Multiplex Screening

[0189] Cells having their nuclei barcoded with vector-encoded RNA transcripts carrying barcodes can be pooled and tested simultaneously in multiplexed assays. The subject methods allow cells with different expressible sequences or genetic modifications to be screened simultaneously in multiplexed assays. Cells can be tested individually or in large pools. Multiple different perturbation types can be screened together. For example, gene knockdowns can be screened simultaneously with synthetic gene knockins, or gene knockouts or knockins of portions of genes. Different types of genetic manipulations can be screened across categories (testing knockouts, knockins, knockdowns, overexpression, underexpression, expression of exogenous genes, etc.) simultaneously.

[0190] In some embodiments, multiplex screening comprises transfection of a population of cells with a plurality of vectors collectively comprising a plurality of expressible sequences, wherein each vector has a barcode identifying the expressible sequence in the vector. For example, a vector comprising an expression cassette encoding a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain and / or a barcoded RNA transcript may further comprise an expression cassette comprising an expressible sequence encoding a gene product of interest. The expressible sequence of interest is placed under the control of a promoter so that the sequence of interest is transcribed into RNA in a host cell. Expressible sequences ofinterest may include without limitation, genes encoding therapeutic proteins, antibodies, enzymes, hormones, cytokines, growth factors, mRNAs, shRNAs, siRNAs, aptamers, site-specific endonucleases, or guide RNAs for RNA-guided nucleases (e.g., Cas9 for gene editing by CRISPR system).

[0191] In some cases, a construct comprises an expressible sequence encoding both a heterologous nucleic acid gene product and a heterologous polypeptide gene product. If the gene product is an RNA, in some cases, the RNA gene product encodes a polypeptide, and in other cases, the RNA gene product does not encode a polypeptide. In some cases, the expressible sequence encodes a single heterologous gene product. In other cases, the expressible sequence encodes two or more heterologous gene products. If the expressible sequence encodes two heterologous gene products, in some cases, the nucleotide sequences encoding the two heterologous gene products are operably linked to the same promoter. If the expressible sequence encodes two heterologous gene products, in some cases, the nucleotide sequences encoding the two heterologous gene products are operably linked to two different promoters. In some cases, a construct comprises an expressible sequence encoding three heterologous gene products. If the expressible sequence encodes three heterologous gene products, in some cases, the nucleotide sequences encoding the three heterologous gene products are operably linked to the same promoter. If the expressible sequence encodes three heterologous gene products, in some cases, the nucleotide sequences encoding the three heterologous gene products are operably linked to two or three different promoters. In some cases, a construct of the present disclosure comprises two or more expressible sequences, each comprising a nucleotide sequence encoding a heterologous gene product.

[0192] In some embodiments, the expressible sequence encodes a polypeptide of interest. The polypeptide of interest may be any type of protein / peptide including, without limitation, an enzyme, an extracellular matrix protein, a receptor, transporter, ion channel, or other membrane protein, a hormone, a growth factor, a neuropeptide, a cytokine, a chemokine, an antibody, or a cytoskeletal protein; or a fragment thereof, or a biologically active domain of interest. In some cases, the gene product is a therapeutic polypeptide, e.g., a polypeptide that provides clinical benefit.

[0193] In some embodiments, the expressible sequence is a polynucleotide encoding a RNA interference (RNAi) nucleic acid or regulatory RNA of interest such as, but not limited to, amicroRNA (miRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a small nuclear RNA (snRNA), a long non-coding RNA (IncRNA), an antisense nucleic acid, and the like. The nucleotide sequence encoding the RNAi nucleic acid or regulatory RNA may be operably linked to a promoter to allow production of the RNAi nucleic acid or regulatory RNA by transcription in a suitable host cell.

[0194] In some embodiments, the gene product is an aptamer. In some cases, the aptamer is a therapeutic aptamer. For example, the aptamer may function as an antagonist by blocking interactions at a disease-associated target (e.g., receptor-ligand interactions). Alternatively, an aptamer can serve as an agonist for activating the function of a target receptor. Exemplary aptamers of interest include aptamers against growth factor receptors and growth factors such as aptamers that bind to epidermal growth factor receptor (see, e.g., Wang et al. (2014) Biochem. Biophys. Res. Commun. 453(4):681-5), transforming growth factor-beta type III receptor (see, e.g., Ohuchi et al. (2006) Biochimie 88(7):897-904.), vascular endothelial growth factor (VEGF) (see, e.g., Ng et al. (2006) Nat. Rev. Drug Discovery 5:123; and Lee et al. (2005) Proc. Natl. Acad. Sci. USA 102:18902) or platelet-derived growth factor (PDGF), e.g., E10030 (see, e.g., Ni and Hui (2009) Ophthalmologica 223:401; and Akiyama et al. (2006) J. Cell Physiol. 207:407).

[0195] In some embodiments, the expressible sequence encodes a sequence- specific endonuclease for use in genome editing. The sequence specific endonuclease can be used to create a double-stranded break at a specific site in the genome. The double stranded breaks can then be repaired by non-homologous end joining (NHEJ), microhomology-mediated end joining (MMEJ), or homology-directed repair (HDR) pathways. Desired genome edits can be introduced into the genome using donor DNA to repair double-strand breaks by homologous recombination. Various sequence-specific endonucleases can be used in genome editing for creation of double-strand breaks in DNA, including, without limitation, engineered zinc-finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, and clustered regularly interspaced short palindromic repeats (CRISPR) Cas endonucleases such as Cas9. See, e.g., Targeted Genome Editing Using Site-Specific Nucleases: ZFNs, TALENs, and the CRISPR / Cas9 System (T. Yamamoto ed., Springer, 2015); Genome Editing: The Next Step in Gene Therapy (Advances in Experimental Medicine and Biology, T. Cathomen, M. Hirsch, and M. Porteus eds., Springer, 2016); Aachen Press Genome Editing (CreateSpace Independent Publishing Platform, 2015); herein incorporated by reference. Precise control over the timing of production of thegenome editing enzyme can be achieved by using an inducible promoter to allow turning on and off of expression as desired.

[0196] In some cases, a gene product of interest is a site- specific endonuclease that provides for site-specific knock-down of gene function, e.g., where the endonuclease knocks out an allele associated with a disease. For example, in a case where a dominant allele encodes a defective copy of a gene, and the wild-type gene provides for normal function, a site- specific endonuclease can be targeted to the defective allele and knock out the defective allele. In some cases, a site-specific endonuclease is an RNA-guided endonuclease.

[0197] A site-specific nuclease can also be used to stimulate homologous recombination with a donor DNA that encodes a functional copy of the protein encoded by the defective allele. Thus, e.g., a vector can be used to deliver a site-specific endonuclease that knocks out a defective allele and also be used to deliver a functional copy of the defective allele, resulting in repair of the defective allele, thereby providing for production of a functional gene product.

[0198] In some cases, the gene product is an RNA-guided endonuclease. In some cases, the gene product is an RNA comprising a nucleotide sequence encoding an RNA-guided endonuclease. In some cases, the gene product is a guide RNA, e.g., a single-guide RNA. In some cases, the gene products are: 1) a guide RNA; and 2) an RNA-guided endonuclease. The guide RNA can comprise: a) a protein-binding region that binds to the RNA-guided endonuclease; and b) a region that binds to a target nucleic acid. An RNA-guided endonuclease is also referred to herein as a “genome editing nuclease.”

[0199] Examples of RNA-guided endonucleases are CRISPR / Cas endonucleases (e.g., class 2 CRISPR / Cas endonucleases such as a type II, type V, or type VI CRISPR / Cas endonucleases). A suitable genome editing nuclease is a CRISPR / Cas endonuclease (e.g., a class 2 CRISPR / Cas endonuclease such as a type II, type V, or type VI CRISPR / Cas endonuclease). In some cases, a suitable RNA-guided endonuclease is a class 2 CRISPR / Cas endonuclease. In some cases, a suitable RNA-guided endonuclease is a class 2 type II CRISPR / Cas endonuclease (e.g., a Cas9 protein). In some cases, a genome targeting composition includes a class 2 type V CRISPR / Cas endonuclease (e.g., a Cpfl protein, a C2cl protein, or a C2c3 protein). In some cases, a suitable RNA-guided endonuclease is a class 2 type VI CRISPR / Cas endonuclease (e.g., a C2c2 protein; also referred to as a “Casl3a” protein). Also suitable for use is a CasX protein. Also suitable for use is a CasY protein.

[0200] In some cases, the genome-editing endonuclease is a Type II CRISPR / Cas endonuclease. In some cases, the genome-editing endonuclease is a Cas9 polypeptide. The Cas9 protein is guided to a target site (e.g., stabilized at a target site) within a target nucleic acid sequence (e.g., a chromosomal sequence or an extrachromosomal sequence, e.g., an episomal sequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.) by virtue of its association with the protein-binding segment of the Cas9 guide RNA. In some cases, the Cas9 polypeptide used in a composition or method of the present disclosure is a Staphylococcus aureus Cas9 (saCas9) polypeptide. In some cases, a suitable Cas9 polypeptide is a high-fidelity (HF) Cas9 polypeptide. Kleinstiver et al. (2016) Nature 529:490. In some cases, a suitable Cas9 polypeptide exhibits altered PAM specificity. See, e.g., Kleinstiver et al. (2015) Nature 523:481. In some cases, the genome-editing endonuclease is a type V CRISPR / Cas endonuclease. In some cases, a type V CRISPR / Cas endonuclease is a Cpfl protein. In some cases, the genome-editing endonuclease is a CasX or a CasY polypeptide. CasX and CasY polypeptides are described in Burstein et al. (2017) Nature 542:237.

[0201] In some cases, a genome editing nuclease is a fusion protein that is fused to a heterologous polypeptide (also referred to as a “fusion partner”). In some cases, a genome editing nuclease is fused to an amino acid sequence (a fusion partner) that provides for subcellular localization, i.e., the fusion partner is a subcellular localization sequence (e.g., one or more nuclear localization signals (NLSs) for targeting to the nucleus, two or more NLSs, three or more NLSs, etc.).

[0202] Also suitable for use is an RNA-guided endonuclease with reduced enzymatic activity. Such an RNA-guided endonuclease is referred to as a “dead” RNA-guided endonuclease; for example, a Cas9 polypeptide that comprises certain amino acid substitutions such that it exhibits substantially no endonuclease activity, but such that it still binds to a target nucleic acid when complexed with a guide RNA, is referred to as a “dead” Cas9 or “dCas9.” In some cases, a “dead” Cas9 protein has a reduced ability to cleave both the complementary and the non-complementary strands of a double stranded target nucleic acid. For example, a “nuclease defective” Cas9 lacks a functioning RuvC domain (i.e., does not cleave the non-complementary strand of a double stranded target DNA) and lacks a functioning HNH domain (i.e., does not cleave the complementary strand of a double stranded target DNA). Such a Cas9 protein has a reduced ability to cleave a target nucleic acid (e.g., a single stranded or double stranded target nucleic acid) but retains the abilityto bind a target nucleic acid. A Cas9 protein that cannot cleave target nucleic acid (e.g., due to one or more mutations, c.g., in the catalytic domains of the RuvC and HNH domains) is referred to as a “nuclease defective Cas9”, “dead Cas9” or simply “dCas9.” Other residues can be mutated to achieve the above effects (i.e. inactivate one or the other nuclease portions).

[0203] In some cases, the genome-editing endonuclease is an RNA-guided endonuclease (and its corresponding guide RNA) known as Cas9-synergistic activation mediator (Cas9-SAM). The RNA-guided endonuclease (e.g., Cas9) of the Cas9-SAM system is a “dead” Cas9 fused to a transcriptional activation domain (wherein suitable transcriptional activation domains include, e.g., VP64, p65, MyoDl, HSF1, RTA, and SET7 / 9) or a transcriptional repressor domain (where suitable transcriptional repressor domains include, e.g., a KRAB domain, a NuE domain, an NcoR domain, a SID domain, and a SID4X domain). The guide RNA of the Cas9-SAM system comprises a loop that binds an adapter protein fused to a transcriptional activator domain (e.g., VP64, p65, MyoDl, HSF1, RTA, or SET7 / 9) or a transcriptional repressor domain (e.g., a KRAB domain, a NuE domain, an NcoR domain, a SID domain, or a SID4X domain). For example, in some cases, the guide RNA is a single-guide RNA comprising an MS2 RNA aptamer inserted into one or two loops of the sgRNA; the dCas9 is a fusion polypeptide comprising dCas9 fused to VP64; and the adaptor / functional protein is a fusion polypeptide comprising: i) MS2; ii) p65; and iii) HSF1. See, e.g., U.S. Patent Publication No. 2016 / 0355797.

[0204] Also suitable for use is a chimeric polypeptide comprising: a) a dead RNA-guided endonuclease; and b) a heterologous fusion polypeptide. Examples of suitable heterologous fusion polypeptides include a polypeptide having, e.g., methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, DNA integration activity, or nucleic acid binding activity.

[0205] A nucleic acid that binds to a class 2 CRISPR / Cas endonuclease (e.g., a Cas9 protein; a type V or type VI CRISPR / Cas protein; a Cpf 1 protein; etc.) and targets the complex to a specific location within a target nucleic acid is referred to herein as a “guide RNA” or “CRISPR / Cas guide nucleic acid” or “CRISPR / Cas guide RNA.” A guide RNA provides target specificity to the complex (the RNP complex) by including a targeting segment, which includes a guide sequence (also referred to herein as a targeting sequence), which is a nucleotide sequence that is complementary to a sequence of a target nucleic acid.

[0206] In some cases, a guide RNA includes two separate nucleic acid molecules: an “activator” and a “targctcr” and is referred to herein as a “dual guide RNA”, a “double-molecule guide RNA”, a “two-molecule guide RNA”, or a “dgRNA.” In some cases, the guide RNA is one molecule (e.g., for some class 2 CRISPR / Cas proteins, the corresponding guide RNA is a single molecule; and in some cases, an activator and targeter are covalently linked to one another, e.g., via intervening nucleotides), and the guide RNA is referred to as a “single guide RNA”, a “singlemolecule guide RNA,” a “one-molecule guide RNA”, or simply “sgRNA.”

[0207] Where the gene product is an RNA-guided endonuclease, or is both an RNA-guided endonuclease and a guide RNA, the gene product can modify a target nucleic acid. In some cases, e.g., where a target nucleic acid comprises a deleterious mutation in a defective allele (e.g., a deleterious mutation in a neural cell target nucleic acid), the RNA-guided endonuclease / guide RNA complex, together with a donor nucleic acid comprising a nucleotide sequence that corrects the deleterious mutation (e.g., a donor nucleic acid comprising a nucleotide sequence that encodes a functional copy of the protein encoded by the defective allele), can be used to correct the deleterious mutation, e.g., via homology-directed repair (HDR).

[0208] In some cases, the gene products are an RNA-guided endonuclease and 2 separate sgRNAs, where the 2 separate sgRNAs provide for deletion of a target nucleic acid via non- homologous end joining (NHEJ).

[0209] In some cases, the gene products are: i) an RNA-guided endonuclease; and ii) one guide RNA. In some cases, the guide RNA is a single-molecule (or “single guide”) guide RNA (an “sgRNA”). In some cases, the guide RNA is a dual-molecule (or “dual-guide”) guide RNA (“dgRNA”).

[0210] In some cases, the gene products are: i) an RNA-guided endonuclease; and ii) 2 separate sgRNAs, where the 2 separate sgRNAs provide for deletion of a target nucleic acid via non- homologous end joining (NHEJ). In some cases, the guide RNAs are sgRNAs. In some cases, the guide RNAs are dgRNAs.

[0211] In some cases, the gene products are: i) a Cpfl polypeptide; and ii) a guide RNA precursor; in these cases, the precursor can be cleaved by the Cpfl polypeptide to generate 2 or more guide RNAs.

[0212] In certain embodiments, detecting the effects of expression of different expressible sequences or sequence- specific genetic perturbations in a plurality of host cells comprisesdetecting a change in cell morphology, cell growth, cell proliferation, gene expression, biological activity of a protein, or any combination thereof in the plurality of host cells compared to an unmodified host cell.

[0213] In certain embodiments, a plurality of host cells are contacted with a test agent to determine the effects of expression of different expressible sequences or different sequencespecific genetic perturbations in the plurality of host cells on activity of the test agent. Cells comprising constructs, described herein, can be subjected to a plurality of candidate agents or other therapeutic intervention. Candidate agents encompass numerous chemical classes, e.g., small organic compounds having a molecular weight of more than 50 daltons and less than about 10,000 daltons, less than about 5,000 daltons, or less than about 2,500 daltons. Test agents can comprise functional groups necessary for structural interaction with proteins, e.g., hydrogen bonding, and can include at least an amine, carbonyl, hydroxyl or carboxyl group, or at least two of the functional chemical groups. The test agents can comprise cyclical carbon or heterocyclic structures and / or aromatic or polyaromatic structures substituted with one or more of the above functional groups. Test agents are also found among biomolecules including peptides, peptide fragments, receptor fragments, co-receptor fragments, saccharides, fatty acids, steroids, purines, pyrimidines, derivatives, structural analogs or combinations thereof.

[0214] Test agents are obtained from a wide variety of sources including libraries of synthetic or natural compounds. For example, numerous means are available for random and directed synthesis of a wide variety of organic compounds and biomolecules, including expression of randomized oligonucleotides and oligopeptides. Alternatively, libraries of natural compounds in the form of bacterial, fungal, plant and animal extracts are available or readily produced. Additionally, natural or synthetically produced libraries and compounds are readily modified through conventional chemical, physical and biochemical means, and may be used to produce combinatorial libraries. Known pharmacological agents may be subjected to directed or random chemical modifications, such as acylation, alkylation, esterification, amidification, etc. to produce structural analogs. Moreover, screening may be directed to known pharmacologically active compounds and chemical analogs thereof, or to new agents with unknown properties such as those created through rational drug design.

[0215] In some embodiments, test agents are synthetic compounds. A number of techniques are available for the random and directed synthesis of a wide variety of organic compounds andbiomolecules, including expression of randomized oligonucleotides. See for example WO 94 / 24314, hereby expressly incorporated by reference, which discusses methods for generating new compounds, including random chemistry methods as well as enzymatic methods.

[0216] In another embodiment, the test agents are provided as libraries of natural compounds in the form of bacterial, fungal, plant and animal extracts that are available or readily produced. Additionally, natural or synthetically produced libraries and compounds are readily modified through conventional chemical, physical and biochemical means. Known pharmacological agents may be subjected to directed or random chemical modifications, including enzymatic modifications, to produce structural analogs.

[0217] In some embodiments, the test agents are organic moieties. In this embodiment, test agents are synthesized from a series of substrates that can be chemically modified. “Chemically modified” herein includes traditional chemical reactions as well as enzymatic reactions. These substrates generally include, but are not limited to, alkyl groups (including alkanes, alkenes, alkynes and heteroalkyl), aryl groups (including arenes and heteroaryl), alcohols, ethers, amines, aldehydes, ketones, acids, esters, amides, cyclic compounds, heterocyclic compounds (including purines, pyrimidines, benzodiazepins, beta-lactams, tetracylines, cephalosporins, and carbohydrates), steroids (including estrogens, androgens, cortisone, ecodysone, etc.), alkaloids (including ergots, vinca, curare, pyrollizdine, and mitomycines), organometallic compounds, hetero-atom bearing compounds, amino acids, and nucleosides. Chemical (including enzymatic) reactions may be done on the moieties to form new substrates or candidate agents which can then be tested using the present invention.

[0218] In some embodiments a test agent is assessed for any cytotoxic activity it may exhibit toward a living eukaryotic cell, using well-known assays, such as trypan blue dye exclusion, an MTT (3-(4,5-dimethylthiazol-2-yl)-2,5-diphenyl-2 H-tetrazolium bromide) assay, and the like. Agents that do not exhibit significant cytotoxic activity are considered candidate agents.

[0219] Cells comprising constructs, described herein, may be tested in culture with one or a panel of cellular' environments, where the cellular environment includes one or more of: exposure to a candidate agent of interest, contact with other cells, electrical stimulation, alterations in ionicity, alterations in temperature, contact with pro-inflammatory or anti-inflammatory agents, contact with infectious agents, e.g. bacterial, viral, fungal, or parasitic infectious agents, and the like, and where cells may vary in types of genetic modifications, in prior exposure to anenvironment of interest, in the dose of an agent that is provided, etc. Usually at least one control is included, for example, a negative control and / or a positive control. Culture of genetically modified cells is typically performed in a sterile environment, for example, at 37 °C. in an incubator containing a humidified 92-95% air / 5-8% CO2 atmosphere. Cell culture may be carried out in nutrient mixtures containing undefined biological fluids such as fetal calf serum, or media which is fully defined and serum free. The effect of the altering the environment may be assessed by monitoring multiple output parameters, including morphological, functional, and genetic changes.

[0220] The agents are conveniently added in solution, or readily soluble form, to the medium used in culturing the genetically modified cells. The agents may be added in a flow-through system, as a stream, intermittent or continuous, or alternatively, adding a bolus of the compound, singly or incrementally, to an otherwise static solution. In a flow-through system, two fluids are used, where one is a physiologically neutral solution, and the other is the same solution with the test compound added. The first fluid is passed over the cells, followed by the second. In a single solution method, a bolus of the test compound is added to the volume of medium surrounding the cells. The overall concentrations of the components of the culture medium should not change significantly with the addition of the bolus, or between the two solutions in a flow through method.

[0221] Preferred agent formulations do not include additional components, such as preservatives, that may have a significant effect on the overall formulation. Thus, preferred formulations consist essentially of a biologically active compound and a physiologically acceptable carrier, e.g., water, ethanol, DMSO, etc. However, if a compound is liquid without a solvent, the formulation may consist essentially of the compound itself.

[0222] A plurality of assays may be run in parallel with different agent concentrations to obtain a differential response to the various concentrations. As known in the art, determining the effective concentration of an agent typically uses a range of concentrations resulting from 1:10, or other log scale, dilutions. The concentrations may be further refined with a second series of dilutions, if necessary. Typically, one of these concentrations serves as a negative control, i.e., at zero concentration or below the level of detection of the agent or at or below the concentration of agent that does not give a detectable change in the phenotype.

[0223] Any suitable method known in the art may be used for analyzing cells. Microscopy techniques including, without limitation, fluorescence microscopy, confocal microscopy, two- photon microscopy, multi-photon microscopy, light-field microscopy, expansion microscopy, andlight sheet microscopy may be used, for example, to detect morphological changes. Additionally, immunofluorescence may be used to detect changes in localization of antigens and expression of surface markers on cells. Patch-clamping can be used to detect electrophysiological changes (e.g., particularly in excitable cells such as neurons and muscle cells). Microarray analysis, single cell RNA sequencing, single nucleus RNA sequencing, and / or proteomic profiling techniques may be used to detect changes in gene expression and / or the distribution of proteins in cells. Biochemical assays may be used to assess changes in activities of particular proteins in cells.

[0224] Various methods can be utilized for quantifying the presence of selected parameters, For measuring the amount of a target analyte that is present, a convenient method is to label a molecule with a detectable moiety, which may be fluorescent, luminescent, radioactive, enzymatically active, etc., particularly a molecule specific for binding to the target analyte with high affinity. Fluorescent moieties are readily available for labeling virtually any biomolecule, structure, or cell type. Immunofluorescent moieties can be directed to bind not only to specific proteins but also specific conformations, cleavage products, or site modifications like phosphorylation. Individual peptides and proteins can be engineered to fluoresce, e.g., by expressing them as green fluorescent protein chimeras inside cells (for a review see Jones et al. (1999) Trends Biotechnol. 17(12):477-81). Cells can be genetically modified to provide fusions of an antibody to a fluorescent or bioluminescent protein.

[0225] Depending upon the label chosen, parameters may be measured using immunoassay techniques such as a radioimmunoassay (RIA) or enzyme linked immuno sorbance assay (ELISA), homogeneous enzyme immunoassays, and related non-enzymatic techniques. These techniques utilize specific antibodies as reporter molecules, which are particularly useful due to their high degree of specificity for attaching to a single molecular target. U.S. Pat. No. 4,568,649 describes ligand detection systems, which employ scintillation counting. These techniques are particularly useful for protein or modified protein parameters or epitopes, or carbohydrate determinants. Cell readouts for proteins and other cell determinants can be obtained using fluorescent or otherwise tagged reporter molecules. Cell based ELISA or related non-enzymatic or fluorescence-based methods enable measurement of cell surface parameters and secreted parameters. Capture ELISA and related non-enzymatic methods usually employ two specific antibodies or reporter molecules and arc useful for measuring parameters in solution. Flow cytometry methods are useful for measuring cell surface and intracellular parameters, as well as shape change and granularity andfor analyses of beads used as antibody- or probe-linked reagents. Readouts from such assays may be the mean fluorescence associated with individual fluorescent antibody-detected cell surface molecules or cytokines, or the average fluorescence intensity, the median fluorescence intensity, the variance in fluorescence intensity, or some relationship among these.

[0226] Quantitative readouts of parameters may include baseline measurements in the absence of agents or a pre-defined genetic control condition and test measurements in the presence of a single or multiple agents or a genetic test condition. Furthermore, quantitative readouts of parameters may include long-term recordings and may therefore be used as a function of time (change of parameter value). Readouts may be acquired either spontaneously or in response to stimulation or perturbation of the cells. The quantitative readouts of parameters may further include a single determined value, the mean or median values of parallel, subsequent or replicate measurements, the variance of the measurements, various normalizations, the cross -correlation between parallel measurements, etc. and every statistic used to a calculate a meaningful and informative factor.Single Nucleus RNA Sequencing

[0227] The subject methods are compatible with single nucleus RNA sequencing, which can be used, for example, for high-throughput transcriptomic profiling and identification of cell types of individual cells in a tissue (see, e.g., Kim et al. (2023) J. Pathol. Transl Med. 57(l):52-59, Deleersnijder et al. (2021) J. Am. Soc. Nephrol. 32(8): 1838- 1852, Guo et al. (2023) Int. J. Mol. Sci. 24(18): 13744, Wang et al. (2023) Curr. Protoc. 3(1 l):e919; herein incorporated by reference in their entireties). Single-cell RNA sequencing typically involves isolation of nuclei from cells, lysis of nuclei, RNA capture, reverse transcription (conversion of RNA into cDNA), cDNA amplification, and sequencing of cDNA amplicons. In some embodiments, a common sequencing primer binding site (e.g., near the barcode sequence) and a constant 5’-adapter and constant 3’- adapter are included in constructs to facilitate high-throughput amplification and sequencing of barcodes.

[0228] In order to isolate the nucleus from a cell, the plasma membrane is lysed without disrupting the nuclear membrane. Both chemical and mechanical methods may be used for cell lysis. Typically, the plasma membrane is first treated with nonionic detergents followed by mechanical disruption of the plasma membrane with a dounce homogenizer. Exemplary nonionicdetergents that can be used include, without limitation, 0.1 % TRITON X-100, 1 % NP40, 0.1 % IGEPAL, or 0.03% TWEEN 20. Bovine scrum albumin and RNase inhibitors may be included throughout the isolation process to prevent RNA degradation. Nuclei can be collected by filtration and further purified by sucrose gradient centrifugation or fluorescence activated nuclei sorting (FANS). Isolation of nuclei from brain tissue may require additional steps to remove myelin debris such as centrifugation with an iodixanol gradient (OptoPrep, San Diego, CA, USA) or passage through a myelin removal column (Miltenyi, Bergisch Gladbach, Germany). After isolation, the morphology of nuclei should be checked by microscopy to confirm their integrity before further processing. For a further description of single nucleus RNA sequencing and methods of isolating nuclei, see, Example 1 and Habib et al. (2017) Nat Methods 14:955-958, Lake et al. (2016) Science 352:1586-1590, Slyper et al. (2020) Nat Med. 26:792-802, Waag et al. (2023) Curr. Protoc. 3(l l):e919, and Kim et al. (2023) J. Pathol. Transl. Med. 57(1): 52-59; herein incorporated by reference in their entireties.

[0229] In certain embodiments, the barcoded RNA transcript is separated from non- homologous nucleic acids in the nucleus using capture oligonucleotides immobilized on a solid support. Such capture oligonucleotides contain nucleic acid sequences that are complementary to a nucleic acid sequence present in the barcoded RNA transcript such that the capture oligonucleotide can “capture” the barcoded RNA transcript. In some embodiments, the capture oligonucleotide comprises a poly T homopolymer chain to capture the RNA transcript by its poly A tail.

[0230] In certain embodiments, capture oligonucleotides are used to bind barcoded RNA transcripts either prior to or after amplification by primers and / or prior to sequencing. A nuclear lysate from an isolated nucleus is contacted with a solid support in association with capture oligonucleotides. The capture oligonucleotides may be associated with the solid support, for example, by covalent binding of the capture moiety to the solid support, by affinity association, hydrogen binding, or nonspecific association.

[0231] The capture oligonucleotides can include from about 5 to about 500 nucleotides, preferably about 10 to about 100 nucleotides, or more preferably about 10 to about 60 nucleotides, or any integer within these ranges, such as a sequence including 18, 19, 20, 21, 22, 23, 24, 25, 26 . . . 35 . . . 40, etc. nucleotides of a sequence that is complementary to a region of the barcodedRNA transcript. The capture oligonucleotide may also be phosphorylated at the 3' end in order to prevent extension of the capture oligonucleotide.

[0232] The capture oligonucleotide may be attached to the solid support in a variety of manners. For example, the oligonucleotide may be attached to the solid support by attachment of the 3' or 5' terminal end to the solid support. Alternatively, the capture oligonucleotide is attached to the solid support by a linker which serves to distance the probe from the solid support. The linker is usually at least 10-50 atoms in length, more preferably, at least 15-30 atoms in length. The required length of the linker will depend on the particular solid support used. For example, a six atom linker is generally sufficient when high cross-linked polystyrene is used as the solid support.

[0233] The solid support may take many forms including, for example, nitrocellulose reduced to particulate form and retrievable upon passing the sample medium containing the support through a sieve; nitrocellulose or the materials impregnated with magnetic particles or the like, allowing the nitrocellulose to migrate within the sample medium upon the application of a magnetic field; beads or particles which may be filtered or exhibit electromagnetic properties; and polystyrene beads which partition to the surface of an aqueous medium. Examples of types of solid supports for immobilization of the oligonucleotide probe include controlled pore glass, glass plates, polystyrene, avidin-coated polystyrene beads, cellulose, nylon, acrylamide gel and activated dextran.

[0234] In one embodiment, the solid support comprises magnetic beads. The magnetic beads may contain primary amine functional groups, which facilitate covalent binding or association of the capture oligonucleotides to the magnetic support particles. Alternatively, the magnetic beads have immobilized thereon homopolymers, such as poly T or poly A sequences. The homopolymers on the solid support will generally be complementary to any homopolymer on the capture oligonucleotide to allow attachment of the capture oligonucleotide to the solid support by hybridization. The use of a solid support with magnetic beads allows for a one-pot method of isolation, amplification and detection as the solid support can be separated from the sample by magnetic means.

[0235] The magnetic beads or particles can be produced using standard techniques or obtained from commercial sources. In general, the particles or beads may be comprised of magnetic particles, although they can also include other magnetic metal or metal oxides, whether in impure,alloy, or composite form, as long as they have a reactive surface and exhibit an ability to react to a magnetic field. Other materials that may be used individually or in combination with iron include, but are not limited to, cobalt, nickel, and silicon. A magnetic bead suitable for use with the present invention includes magnetic beads containing poly dT groups marketed under the trade name Sera- Mag magnetic oligonucleotide beads by Seradyn, Indianapolis, Ind.

[0236] In certain embodiments, the capture oligonucleotides are combined with a nuclear lysate under conditions suitable for hybridization with barcoded RNA transcripts prior to immobilization of the capture oligonucleotides on a solid support. The capture oligonucleotide- target nucleic acid complexes formed are then bound to the solid support. In other embodiments, a solid support with associated capture oligonucleotides is brought into contact with a nuclear lysate under hybridizing conditions. The immobilized capture oligonucleotides hybridize to the barcoded RNA transcripts present in the nuclear lysate. Typically, hybridization of capture oligonucleotides to the barcoded RNA transcript can be accomplished in approximately 15 minutes, but may take as long as 3 to 48 hours.

[0237] The solid support is then separated from the sample, for example, by filtering, centrifugation, passing through a column, or by magnetic means. The solid support may be washed to remove unbound contaminants and transferred to a suitable container (e.g., a microtiter plate). As will be appreciated by one of skill in the art, the method of separation will depend on the type of solid support selected. Since the barcoded RNA transcripts are hybridized to the capture oligonucleotides immobilized on the solid support, the RNA transcripts are thereby separated from the impurities in the nuclear lysate. In some cases, extraneous nucleic acids, proteins, carbohydrates, lipids, cellular debris, and other impurities may still be bound to the support, although at much lower concentrations than initially found in the nuclear lysate. Those skilled in the art will recognize that some undesirable materials can be removed by washing the support with a washing medium. The separation of a solid support from a sample preferably removes at least about 70%, more preferably about 90% and, most preferably, at least about 95% or more of the non-target nucleic acids present in the sample.

[0238] Any high-throughput technique for sequencing can be used to sequence the barcodes of the barcoded RNA transcripts isolated from the nuclei. Barcodes in RNA transcripts may include cell tracing sequences, unique molecular identifier sequences, vector feature identifiersequences, and / or positional barcodes. Tn some embodiments, a barcode identifies an expressible sequence or genome modification module introduced into a cell by a vector.

[0239] DNA sequencing techniques include dideoxy sequencing reactions (Sanger method) using labeled terminators or primers and gel separation in slab or capillary, sequencing by synthesis using reversibly terminated labeled nucleotides, pyrosequencing, 454 sequencing, sequencing by synthesis using allele specific hybridization to a library of labeled clones followed by ligation, real time monitoring of the incorporation of labeled nucleotides during a polymerization step, polony sequencing, SOLID sequencing, and the like.

[0240] Certain high-throughput methods of sequencing comprise a step in which individual molecules are spatially isolated on a solid surface where they are sequenced in parallel. Such solid surfaces may include nonporous surfaces (such as in Solexa sequencing, e.g. Bentley et al, Nature, 456: 53-59 (2008) or Complete Genomics sequencing, e.g. Drmanac et al, Science, 327: 78-81 (2010)), arrays of wells, which may include bead- or particle-bound templates (such as with 454, e.g. Margulies et al, Nature, 437: 376-380 (2005) or Ion Torrent sequencing, e.g., U.S. patent publication 2010 / 0137143 or 2010 / 0304982), micromachined membranes (such as with SMRT sequencing, e.g. Eid et al, Science, 323: 133-138 (2009)), or bead arrays (as with SOLiD sequencing or polony sequencing, e.g. Kim et al, Science, 316: 1481-1414 (2007)). Such methods may comprise amplifying the isolated molecules either before or after they are spatially isolated on a solid surface. Prior amplification may comprise emulsion-based amplification, such as emulsion PCR or rolling circle amplification.

[0241] Of particular interest is sequencing on the Illumina MiSeq, NextSeq, and HiSeq platforms, which use reversible-terminator sequencing by synthesis technology (see, e.g., Shen et al. (2012) BMC Bioinformatics 13:160; Junemann et al. (2013) Nat. Biotechnol. 31(4):294-296; Glenn (2011) Mol. Ecol. Resour. l l(5):759-769; Thudi et al. (2012) Brief Funct. Genomics 11(1):3-11; herein incorporated by reference); the Oxford Nanopore Technologies Inc. MinlON, GridlON, and PromethlON nanopore sequencing platforms, which can be used to determine the sequences of DNA or RNA by monitoring changes in electrical current as nucleic acids are passed through a protein nanopore (see, e.g., Lu et al. (2016) Genomics Proteomics Bioinformatics 14(5):265-279, Petersen et al. (2019) J. Clin. Microbiol. 58(l):e01315- 19, Kono et al. (2019) Dev Growth Differ. 61(5):316-326, Deamer et al. (2016) Nat. Biotechnol. 34(5):518-24, Madoui et al. (2015) BMC Genomics 16:327, Szalay et al. (2015) Nat. Biotechnol 33, 1087-1091; hereinincorporated by reference); the PacBIO Single Molecule, Real-Time (SMRT) sequencing platforms, including the Sequel, HiFi, and RS II sequencing platforms (sec, c.g., Ardui ct al. (2018) Nucleic Acids Res. 46(5):2159-2168, An et al. (2018) Genes (Basel) 9(1):43), Nakano et al. (2017) Hum Cell. 30(3) : 149- 161 ; herein incorporated by reference), the Omniome sequencing by binding (SBB®) short-read sequencing platform using high fidelity plasmonic nanohole arrays (see, e.g., Cetin et al. (2018) ACS Sens. 3(3):561-568; herein incorporated by reference), the Gynapsys compact DNA sequencer, which uses metal oxide semiconductor (CMOS) sequencing chips for electronic data detection and sequencing by synthesis (SBS) chemistry, the Singular Genomics G4 benchtop sequencing platform, which uses SBS chemistry, and the Element Biosciences AVITI™ benchtop sequencer, which uses a modified form of SBS chemistry that reduces reagent usage.Amplification of Barcodes

[0242] Any primer-dependent amplification method known in the art may be used for amplification of barcodes or a portion of an RNA transcript comprising its barcode. Nucleic acid amplification methods include, without limitation, polymerase chain reaction (PCR), rolling circle amplification, and isothermal amplification methods such as recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase-dependent amplification (HDA), nicking enzyme amplification reaction (NEAR), ligase chain reaction (LGR), nucleic acid sequence based amplification (NASBA), transcription-mediated amplification (TMA), Q-beta amplification, and the like.

[0243] In some embodiments, polymerase chain reaction (PCR)-based techniques are used to amplify a construct or barcode region. PCR is a technique for amplifying a desired target nucleic acid sequence contained in a nucleic acid molecule or mixture of molecules. In PCR, a pair of primers is employed in excess to hybridize to the complementary strands of the target nucleic acid. The primers are each extended by a polymerase using the target nucleic acid as a template. The extension products become target sequences themselves after dissociation from the original target strand. New primers are then hybridized and extended by a polymerase, and the cycle is repeated to geometrically increase the number of target sequence molecules. The PCR method for amplifying target nucleic acid sequences in a sample is well known in the art and has been described in, e.g., Innis et al. (eds.) PCR Protocols (Academic Press, NY 1990); Taylor (1991) Polymerase chain reaction: basic principles and automation, in PCR: A PracticalApproach, McPherson et al. (eds.) IRL Press, Oxford; Saiki et al. (1986) Nature 324:163; as well as in U.S. Pat. Nos. 4,683,195, 4,683,202 and 4,889,818, all incorporated herein by reference in their entireties.

[0244] In particular, PCR uses relatively short oligonucleotide primers which flank the target nucleotide sequence to be amplified, oriented such that their 3' ends face each other, each primer extending toward the other. The polynucleotide sample is extracted and denatured, preferably by heat, and hybridized with first and second primers that are present in molar excess. Polymerization is catalyzed in the presence of the four deoxyribonucleotide triphosphates (dNTPs — dATP, dGTP, dCTP and dTTP) using a primer- and template-dependent polynucleotide polymerizing agent, such as any enzyme capable of producing primer extension products, for example, E. coli DNA polymerase I, Klenow fragment of DNA polymerase I, T4 DNA polymerase, thermostable DNA polymerases isolated from Thermits aquaticus (Taq), available from a variety of sources (for example, Perkin Elmer), Thermus thermophilus polymerase (United States Biochemicals), Bacillus stereothermophilus polymerase (Bio-Rad), Thermococcus litoralis polymerase (“Vent” polymerase, New England Biolabs), Pyrococcus species GB-D polymerase (“Deep Vent” polymerase, New England Biolabs), Pyrococcus woesei polymerase (Pwo polymerase, Sigma- Aldrich) or Pyrococcus furiosus polymerase (Pfu polymerase from Promega Corporation). This results in two “long products” which contain the respective primers at their 5' ends covalently linked to the newly synthesized complements of the original strands. The reaction mixture is then returned to polymerizing conditions, e.g., by lowering the temperature, inactivating a denaturing agent, or adding more polymerase, and a second cycle is initiated. The second cycle provides the two original strands, the two long products from the first cycle, two new long products replicated from the original strands, and two “short products” replicated from the long products. The short products have the sequence of the target sequence with a primer at each end. On each additional cycle, an additional two long products are produced, and a number of short products equal to the number of long and short products remaining at the end of the previous cycle. Thus, the number of short products containing the target sequence grows exponentially with each cycle. Preferably, PCR is carried out with a commercially available thermal cycler, e.g., Perkin Elmer.

[0245] RNA transcripts may be amplified by reverse transcribing the RNA into cDNA, and then performing PCR (RT-PCR), as described above. Alternatively, a single enzyme may be used for both steps as described in U.S. Pat. No. 5,322,770, incorporated herein by reference in itsentirety. RNA may also be reverse transcribed into cDNA, followed by asymmetric gap ligase chain reaction (RT-AGLCR) as described by Marshall ct al. (1994) PCR Meth. App. 4:80-84.

[0246] PCR primers should be of sufficient length to provide for hybridization to complementary template DNA under annealing conditions. The primers will generally be at least 6 bp in length, including but not limited to e.g., at least 10 bp in length, at least 15 bp in length, at least 16 bp in length, at least 17 bp in length, at least 18 bp in length, at least 19 bp in length, at least 20 bp in length, at least 21 bp in length, at least 22 bp in length, at least 23 bp in length, at least 24 bp in length, at least 25 bp in length, at least 26 bp in length, at least 27 bp in length, at least 28 bp in length, at least 29 bp in length, at least 30 bp in length, and may be as long as 60 bp in length or longer, where the length of the primers will generally range from 18 to 50 bp in length, including but not limited to, e.g., from about 20 to 35 bp in length. In some instances, the template DNA may be contacted with a single primer or a set of two primers (forward and reverse primers), depending on whether primer extension, linear or exponential amplification of the template DNA is desired. Methods of PCR that may be employed in the subject methods include but are not limited to those described in U.S. Pat. Nos.: 4,683,202; 4,683,195; 4,800,159; 4,965,188 and 5,512,462, the disclosures of which are herein incorporated by reference.

[0247] Alternatively, a polymerase that preferentially uses dUTP rather than dTTP can be used to perform PCR. Such polymerases include archaeal family B DNA polymerases such as Nanoarchaeum equitans B DNA polymerase, which can utilize deaminated bases such as uracil and hypoxanthine and performs PCR with higher fidelity than Thermits aquaticus (Taq) DNA polymerase (e.g., as described in Choi et al. (2008) Appl. Environ. Microbiol. 74(21): 6563-6569; herein incorporated by reference). In addition, engineered polymerases such as Q5U Hot Start High-Fidelity DNA Polymerase from New England Biolabs (Ipswich, MA) and Phusion U DNA polymerase from Thermo Fisher Scientific (Waltham, MA), which contain a mutation in the nucleotide-binding pocket that enables these polymerases to amplify templates containing uracil and inosine bases, may be used to perform PCR with dUTP. The use of polymerases that utilize UTP is useful for preventing carryover contamination in different PCR runs. The uracil-containing amplicon products of such polymerases can be digested by a uracil-DNA glycosylase to remove residual products from previous PCR amplifications and suppress template contamination between runs.

[0248] In addition, one or more PCR additives or enhancing agents may be included to improve the yield of the amplification reaction, for example, by reducing secondary structure in a nucleic acid or mispriming events. Such additives or enhancing agents include, but are not limited to, dimethyl sulfoxide (DMSO), N,N,N-trimethylglycine (betaine), formamide, glycerol, nonionic detergents (e.g., Triton X-100, Tween 20, and Nonidet P-40 (NP-40)), 7-deaza-2'-deoxyguanosine, bovine serum albumin, T4 gene 32 protein, polyethylene glycol, 1,2-propanediol, and tetramethylammonium chloride.

[0249] A PCR reaction will generally be carried out by cycling the reaction mixture between appropriate temperatures for annealing, elongation / extension, and denaturation for specific times. Such temperature and times will vary and will depend on the particular components of the reaction including, e.g., the polymerase and the primers as well as the expected length of the resulting PCR product. In some instances, e.g., where nested or two-step PCR are employed the cycling-reaction may be carried out in stages, e.g., cycling according to a first stage having a particular cycling program or using particular temperature(s) and subsequently cycling according to a second stage having a particular cycling program or using particular temperature(s).

[0250] Multistep PCR processes may or may not include that addition of one or more reagents following the initiation of amplification. For example, in some instances, amplification may be initiated by elongation with the use of a polymerase and, following an initial phase of the reaction, additional reagent(s) (e.g., one or more additional primers, additional enzymes, etc.) may be added to the reaction to facilitate a second phase of the reaction. In some instances, amplification may be initiated with a first primer or a first set of primers and, following an initial phase of the reaction, additional reagent(s) (e.g., one or more additional primers, additional enzymes, etc.) may be added to the reaction to facilitate a second phase of the reaction. In certain embodiments, the initial phase of amplification may be referred to as “preamplification”.

[0251] In particular, the subject methods are applicable to digital PCR techniques. For digital PCR, a sample containing nucleic acids is separated into a large number of partitions before performing PCR. Partitioning can be achieved in a variety of ways known in the ail, for example, by use of micro well plates, capillaries, emulsions, arrays of miniaturized chambers or nucleic acid binding surfaces. Separation of the sample may involve distributing any suitable portion including up to the entire sample among the partitions. Each partition includes a fluid volume that is isolated from the fluid volumes of other partitions. The partitions may be isolated from one another by afluid phase, such as a continuous phase of an emulsion, hy a solid phase, such as at least one wall of a container, or a combination thereof. In certain embodiments, the partitions may comprise droplets disposed in a continuous phase, such that the droplets and the continuous phase collectively form an emulsion.

[0252] The partitions may be formed by any suitable procedure, in any suitable manner, and with any suitable properties. For example, the partitions may be formed with a fluid dispenser, such as a pipette, with a droplet generator, by agitation of the sample (e.g., shaking, stirring, sonication, etc.), and the like. Accordingly, the partitions may be formed serially, in parallel, or in batch. The partitions may have any suitable volume or volumes. The partitions may be of substantially uniform volume or may have different volumes. Exemplary partitions having substantially the same volume are monodisperse droplets. Exemplary volumes for the partitions include an average volume of less than about 100, 10 or 1 mL, less than about 100, 10, or 1 nL, or less than about 100, 10, or 1 pL, among others.

[0253] After separation of the sample, PCR is carried out in the partitions. The partitions, when formed, may be competent for performance of one or more reactions in the partitions. Alternatively, one or more reagents may be added to the partitions after they are formed to render them competent for reaction. The reagents may be added by any suitable mechanism, such as a fluid dispenser, fusion of droplets, or the like.

[0254] In some embodiments, nucleic acids are amplified by emulsion PCR to compartmentalize the amplification reactions of individual DNA molecules. An aqueous PCR mixture with forward and reverse primers is mixed with an oil to create the emulsion. Preferably, each droplet of water in the oil emulsion contains one bead and one molecule of template DNA, such that individual molecules are amplified in separate emulsion droplets. After amplification, the emulsion is broken, e.g., using isopropanol and detergent with vortexing. In some embodiments, the gene fragment library and the sequencing library are bound to magnetic beads or superparamagnetic beads prior to amplification, wherein amplification and breaking of the emulsion is followed by magnetic separation of the beads. For a description of emulsion PCR, see, e.g., Kanagal-Shamanna et al. (2016) Methods Mol Biol. 1392:33-42, Zhu et al. (2012) Anal Bioanal Chem. 403(8):2127-43, Zhang et al. (2020) Lab Chip 20(13):2328-2333, Siu et al. (2021) Taianta 221:121593, Zheng et al. (2011) Nat. Protoc. 6(9):1367-1376, and Kojima et al. (2015) Methods Mol. Biol. 2015;1347:87-100; herein incorporated by reference.

[0255] After PCR amplification, nucleic acids can be quantified by counting the partitions that contain PCR amplicons. Partitioning of the sample allows quantification of the number of different molecules by assuming that the population of molecules follows a Poisson distribution. For a description of digital PCR methods, see, e.g., Hindson et al. (2011) Anal. Chem. 83(22):8604- 8610; Pohl and Shih (2004) Expert Rev. Mol. Diagn. 4(l):41-47; Pekin et al. (2011) Lab Chip 11 (13): 2156-2166; Pinheiro et al. (2012) Anal. Chem. 84 (2): 1003-1011; Day et al. (2013) Methods 59(1) : 101 - 107; herein incorporated by reference in their entireties.

[0256] In some instances, amplification may be carried out under isothermal conditions, e.g., by means of isothermal amplification. Methods of isothermal amplification generally make use of enzymatic means of separating DNA strands to facilitate amplification at constant temperature, such as, e.g., strand-displacing polymerase or a helicase, thus negating the need for thermocycling to denature DNA. Any convenient and appropriate means of isothermal amplification may be employed in the subject methods including but are not limited to: recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase-dependent amplification (HDA), nicking enzyme amplification reaction (NEAR), ligase chain reaction (LGR), nucleic acid sequence based amplification (NASBA), transcription-mediated amplification (TMA), Q-beta amplification, and the like.

[0257] RPA combines isothermal recombinase-mediated primer targeting with stranddisplacement DNA synthesis (Piepenburg et al. (2006) PLOS Biology. 4 (7): e204; herein incorporated by reference). The technique uses two primers together with a recombinase, a singlestranded DNA-binding protein, and a strand-displacing polymerase for amplification. Unlike PCR, heat is not required for melting of the DNA strands. Instead, a recombinase-primer complex is used for localized strand exchange to place oligonucleotide primers at homologous sequences of the DNA template. The single-stranded DNA-binding protein binds to the displaced template strand to prevent the primers from being ejected by branch migration. Dissociation of the recombinase leaves the 3 '-end of the primer accessible to the strand displacing DNA polymerase (e.g., the large fragment of Bacillus subtilis Pol I), which catalyzes primer extension. Cyclic repetition of this process results in exponential amplification.

[0258] LAMP generally utilizes a plurality of primers, e.g., 4-6 primers, which may recognize a plurality of distinct regions, e.g., 6-8 distinct regions, of target DNA. Synthesis is generally initiated by a strand-displacing DNA polymerase with two of the primers forming loop structuresto facilitate subsequent rounds of amplification. LAMP is rapid and sensitive. In addition, the magnesium pyrophosphate produced during the LAMP amplification reaction may, in some instances, be visualized without the use of specialized equipment, e.g., by eye.

[0259] SDA generally involves the use of a strand-displacing DNA polymerase (e.g., Bst DNA polymerase, Large (Klenow) Fragment polymerase, Klenow Fragment (3'-5' exo-), and the like) to initiate at nicks created by a strand-limited restriction endonuclease or nicking enzyme at a site contained in a primer. In SDA, the nicking site is generally regenerated with each polymerase displacement step, resulting in exponential amplification.

[0260] HDA generally employs: a helicase which unwinds double-stranded DNA unwinding to separate strands; primers, e.g., two primers, that may anneal to the unwound DNA; and a stranddisplacing DNA polymerase for extension.

[0261] NEAR generally involves a strand-displacing DNA polymerase that initiates elongation at nicks, e.g., created by a nicking enzyme. NEAR is rapid and sensitive, quickly producing many short nucleic acids from a target sequence.

[0262] Nucleic acid sequence-based amplification (NASBA) is an isothermal RNA-specific amplification method that does not require thermal cycling instrumentation. RNA is initially reverse transcribed such that the single- stranded RNA target is copied into a double- stranded DNA molecule that serves as a template for RNA transcription. Detection of the amplified RNA is typically accomplished either by electrochemiluminescence or in real-time, for example, with fluorescently labeled molecular beacon probes. See, e.g., Lau et al. (2006) Dev. Biol. (Basel) 126:7-15; and Deiman et al. (2002) Mol. Biotechnol. 20(2): 163-179.

[0263] The Ligase Chain Reaction (LCR) is an alternate method for nucleic acid amplification. In LCR, probe pairs are used which include two primary (first and second) and two secondary (third and fourth) probes, all of which are employed in molar excess to the target. The first probe hybridizes to a first segment of the target strand, and the second probe hybridizes to a second segment of the target strand, the first and second segments being contiguous so that the primary probes abut one another in 5' phosphate-3' hydroxyl relationship, and so that a ligase can covalently fuse or ligate the two probes into a fused product. In addition, a third (secondary) probe can hybridize to a portion of the first probe and a fourth (secondary) probe can hybridize to a portion of the second probe in a similar abutting fashion. If the target is initially double stranded, the secondary probes also will hybridize to the target complement in the first instance. Once the ligatedstrand of primary probes is separated from the target strand, it will hybridize with the third and fourth probes which can be ligated to form a complementary, secondary ligated product. It is important to realize that the ligated products are functionally equivalent to either the target or its complement. By repeated cycles of hybridization and ligation, amplification of the target sequence is achieved. This technique is described more completely in EPA 320,308 to K. Backman published Jun. 16, 1989 and EPA 439,182 to K. Backman et al., published Jul. 31, 1991, both of which are incorporated herein by reference.

[0264] Other known methods for amplification of nucleic acids include, but are not limited to, self-sustained sequence replication (3SR) described by Guatelli et al., Proc. Natl. Acad. Sci. f / SA (1990) 87:1874-1878 and J. Compton, Nature (1991) 350:91-92 (1991); Q-beta amplification; strand displacement amplification (as described in Walker et al., Clin. Chem. (1996) 42:9-13 and EPA 684,315; target mediated amplification, as described in International Publication No. WO 93 / 22461, and the TaqMan™ assay.

[0265] In some instances, entire amplification methods may be combined, or aspects of various amplification methods may be recombined to generate a hybrid amplification method. For example, in some instances, aspects of PCR may be used, e.g., to generate the initial template or amplicon or first round or rounds of amplification, and an isothermal amplification method may be subsequently employed for further amplification. In some instances, an isothermal amplification method or aspects of an isothermal amplification method may be employed, followed by PCR for further amplification of the product of the isothermal amplification reaction. In some instances, a sample may be preamplified using a first method of amplification and may be further processed, including e.g., further amplified or analyzed, using a second method of amplification. As a nonlimiting example, a sample may be preamplified by PCR and further analyzed by qPCR. In some instances, the method further comprises monitoring the amplification of a target DNA molecule such as is performed in, e.g., real-time PCR, also referred to herein as quantitative PCR (qPCR).

[0266] The fluorogenic 5' nuclease assay is conveniently performed using, for example, AmpliTaq Gold™ DNA polymerase, which has endogenous 5' nuclease activity, to digest an internal oligonucleotide probe labeled with both a fluorescent reporter dye and a quencher (see, Holland et al., Proc. Natl. Acad. Sci. USA (1991) 88:7276-7280; and Lee et al., Nucl. Acids Res. (1993) 21:3761-3766). Assay results are detected by measuring changes in fluorescence that occur during the amplification cycle as the fluorescent probe is digested, uncoupling the dye andquencher labels and causing an increase in the fluorescent signal that is proportional to the amplification of target nucleic acid.

[0267] The amplification products can be detected in solution or using solid supports. In this method, the TaqMan™ probe is designed to hybridize to a target sequence within the desired PCR product. The 5' end of the TaqMan™ probe contains a fluorescent reporter dye. The 3' end of the probe is blocked to prevent probe extension and contains a dye that will quench the fluorescence of the 5' fluorophore. During subsequent amplification, the 5' fluorescent label is cleaved off if a polymerase with 5' exonuclease activity is present in the reaction. Excision of the 5' fluorophore results in an increase in fluorescence that can be detected. For a detailed description of the TaqMan™ assay, reagents and conditions for use therein, see, e.g., Holland et al., Proc. Natl. Acad. Sci, U.S.A. (1991) 88:7276-7280; U.S. Pat. Nos. 5,538,848, 5,723,591, and 5,876,930, all incorporated herein by reference in their entireties.

[0268] TMA is an isothermal, autocatalytic nucleic acid target amplification system that can provide more than a billion RNA copies of a target sequence, and thus provides a method of identifying target nucleic acid sequences present in very small amounts in a sample. For a detailed description of TMA assay methods, see, e.g., Hill (2001) Expert Rev. Mol. Diagn. 1:445-55; WO 89 / 1050; WO 88 / 10315; EPO Publication No. 408,295; EPO Application No. 8811394-8.9; WO91 / 02818; U.S. Pat. Nos. 5,399,491, 6,686,156, and 5,556,771, all incorporated herein by reference in their entireties.

[0269] Suitable DNA polymerases include reverse transcriptases, such as avian myeloblastosis virus (AMV) reverse transcriptase (available from, e.g., Seikagaku America, Inc.) and Moloney murine leukemia virus (MMLV) reverse transcriptase (available from, e.g., Bethesda Research Laboratories).Adapters

[0270] Adapter oligonucleotides comprising known sequences can be added to the 5' and 3' ends of cDNA copies of the barcoded RNA transcript to facilitate amplification and / or sequencing. Adapters can be designed with a primer binding site (e.g., near the barcode sequence) having a sequence suitable for hybridizing to primers for primer-dependent amplification and / or sequencing, Adapters may also include sites that allow cDNA to attach to a solid support. To facilitate multiplexing, the adapters can be barcoded.

[0271] Adapter oligonucleotides can be ligated to the ends of a nucleic acid using a ligase. Any suitable ligase may be used, including, without limitation, phage ligases such as T4 or T7 DNA ligase, archaeal ligases, or bacterial ligases. The adapters that are ligated to either end of a nucleic acid may be the same or different. Adapters may be attached to the ends of DNA by either blunt end ligation or sticky end ligation. Ligation of the ends of two DNA molecules involves formation of a phosphodiester bond between the 3'-hydroxyl group at the 3’ end of one DNA molecule with the 5'-phosphoryl group at the 5’ end of another DNA molecule. The ends of DNA molecules may be prepared for ligation by blunting of the DNA ends and phosphorylation of the 5’ end. Blunting involves removing a single-stranded overhang (e.g., which may have been created by a restriction enzyme) by adding nucleotides to the complementary strand using the overhang as a template for polymerization, or removing the overhang using an exonuclease. DNA ends may be "blunted" to allow non-compatible ends to be joined by ligation. DNA polymerases, such as the Klenow fragment of DNA polymerase I and T4 DNA polymerase can be used to fill in nucleotides or chew back a 3’ overhang. A nuclease, such as Mung Bean Nuclease may be used for removal of a 5' overhang. In some cases, adapters are designed with a poly T overhang, which allows an adapter to be ligated to the 3 '-ends of DNA having a poly A-o verhang. For example, the ends of nucleic acids may be blunted and phosphorylated at the 5’ ends by treating the DNA with T4 polynucleotide kinase, T4 DNA polymerase, and the Klenow large fragment. A poly A tail can be added to the 3' ends of nucleic acids using either Taq polymerase or the Klenow large fragment.

[0272] Solid phase amplification of polynucleotides is typically performed by first ligating known adapter sequences to each end of a target nucleic acid. The double-stranded nucleic acid is then denatured to form a single- stranded template molecule that is immobilized on a solid support (e.g., the surface of a flow-cell for the Illumina platform or beads for the Ion Torrent platform). The adapter sequence on the 3' end of the template is hybridized to an extension primer, and amplification is performed by extending the primer. In certain aspects, a sequencing platform adapter construct includes one or more nucleic acid domains selected from: a domain (e.g., a "capture site" or "capture sequence") that specifically binds to a surface-attached sequencing platform oligonucleotide (e.g., the P5 or P7 oligonucleotides attached to the surface of a flow cell in an Illumina® sequencing system); a sequencing primer binding domain (e.g., a domain to which the Read 1 or Read 2 primers of the Illumina® platform may bind); a barcode domain (e.g., a domain that uniquely identifies the sample source of a nucleic acid being sequenced to enablesample multiplexing by marking every molecule from a given sample with a specific barcode or "tag"); a barcode sequencing primer binding domain (a domain to which a primer used for sequencing a barcode binds); or any combination of such domains.

[0273] The adapters chosen for the preparation of a sequencing library should be compatible with the sequencing system to be used. Polynucleotides are incorporated into the sequencing library by ligation to sequencing adapters containing specific sequences designed to work with the sequencing platform (i.e., sequencing platform adapter domain). A sequencing platform adapter domain, when present in an adapter, may include one or more nucleic acid domains of any length and sequence suitable for the sequencing platform of interest. In some embodiments, the nucleic acid domains are from 4 to 200 nucleotides in length. For example, the nucleic acid domains may be from 4 to 100 nucleotides in length, such as from 6 to 75, from 8 to 50, or from 10 to 40 nucleotides in length. According to certain embodiments, the sequencing platform adapter construct includes a nucleic acid domain that is from 2 to 8 nucleotides in length, such as from 9 to 15, from 16 to 22, from 23 to 29, or from 30 to 36 nucleotides in length.

[0274] The nucleic acid domains may have a length and sequence that enables a polynucleotide (e.g., an oligonucleotide) employed by the sequencing platform of interest to specifically bind to the nucleic acid domain, e.g., for solid phase amplification and / or sequencing. Example nucleic acid domains include the P5, P7, Read 1 primer, and Read 2 primer domains employed on the Illumina®-based sequencing platforms. Other example nucleic acid domains include the A adapter and Pl adapter domains employed on the Ion Torrent™-based sequencing platforms. The nucleotide sequences of nucleic acid domains useful for sequencing on a sequencing platform of interest may vary and / or change over time. Adapter sequences are typically provided by the manufacturer of the sequencing platform (e.g., in technical documents provided with the sequencing system and / or available on the manufacturer's website). Based on such information, the sequence of any sequencing platform adapter domains, amplification primers, and / or the like, may be designed to include all or a portion of one or more nucleic acid domains in a configuration that enables sequencing the nucleic acid insert (corresponding to the assembled nucleic acid product) on the platform of interest.

[0275] Any suitable type of adapter may be ligated to the assembled nucleic acids, including, without limitation, linear adapters, Y adapters, stubby adapters, and hairpin adapters. In some embodiments, two different linear adapters are ligated to the 5’ and 3’ ends of every member of alibrary. A disadvantage of using linear double stranded adapters is that if two different adapters (referred to as A and B) arc ligated to double-stranded DNA in a library, a mixture of ligation products is generated including about 50% incorrectly ligated products with the same adapter on both ends (A-insert-A, B-insert-B), and only 50% having correctly ligated adapters (e.g., A-insert- B or B-insert-A). The DNA inserts having A-A and B-B adapters are unable to be amplified by PCR.

[0276] Y-adapters have a Y-shaped structure with a single- stranded 5’ arm, a single stranded 3’ arm, and a double-stranded stem. The arms of the Y have noncomplementary sequences, and the stem is double- stranded with complementary strands of DNA that hybridize to each other. Y- adapters are attached to the assembled nucleic acids by ligating the double-stranded stem of the Y-adapters to the 5’ and 3’ ends of the assembled nucleic acids. The use of Y adapters allows the addition of different adapter sequences to the 5' and 3' ends of a DNA library simultaneously. An advantage of using Y adapters rather than linear adapters is that the Y adapters provide an efficient method of attaching two different adapter sequences on the 5’ and 3’ ends of every assembled nucleic acid in a library with a single adapter. After ligation of a Y adapter to both DNA ends, the library can be amplified using primers having sequences that are complementary to the sequences of the 5’ and 3’ arms of the Y-adapters. The use of Y adapters also allows both DNA strands to be sequenced. Exemplary Y-adapters compatible with the Ion Torrent sequencing platform are described by Forth et al. (BioTechniques (2019) 67 :229-237 ; herein incorporated by reference. Y- adapters compatible with the Illumina sequencing platform are also available. The xGen Stubby adapter from Integrated DNA Technologies, Inc. (Coralville, Iowa) is a short Y adapter that can be ligated to fragments with A overhangs and used with proprietary unique dual indexing (UDI) primer pairs.

[0277] In some embodiments, a hairpin adapter is used. Hairpin adapters generally comprise a double-stranded "stem" region and a single stranded "loop" region. In some embodiments, the hairpin adapter comprises one strand (i.e., one continuous strand) capable of adopting a hairpin structure, wherein the hairpin adapter comprises a self-complementary palindromic region that forms the stem and a non-complementary region that forms the loop of the hairpin adapter. In addition, hairpin adapters may comprise various components of adapters as described herein, including, without limitation, amplification priming sites, barcode sequences, and specific sequencing platform adapter domain sequences (e.g., P5 and P7 or A and Pl adapter sequences).Hairpin adapters may further comprise one or more cleavage sites capable of being cleaved under cleavage conditions. In some embodiments, a cleavage site is located in the loop region. Cleavage at a cleavage site generates two separate strands from the hairpin adapter. In some embodiments, cleavage at a cleavage site in the loop region generates a partially double stranded adapter with a double stranded stem and two unpaired strands forming a "Y" structure. Cleavage sites may comprise, for example, uracil and / or deoxyuridine bases, which may be cleaved, for example, using DNA glycosylases, endonucleases, RNAses, and the like and combinations thereof. In some embodiments, the loop of a hairpin adapter contains a uracil residue and can be cleaved using uracil DNA glycosylase and endonuclease VIII. An advantage of using hairpin adapters is that such adapters allow contiguous sequencing of both strands of a double-stranded DNA molecule by covalently attaching one strand to the other. Hairpin adapters also help to minimize the formation of adaptor-dimers during adaptor ligation.

[0278] Hairpin adapters are commercially available from New England Biolabs (Ipswich, MA) such as the NEBNext® adaptor, which has sequences that are compatible with Illumina sequencing platforms. The Oxford Nanopore MinlON sequencing platform uses a Y-adapter and a hairpin adapter. The Y-adapter is ligated to the one end of a double-stranded DNA molecule, which provides for attachment of the DNA molecule to the sequencing nanopore. The hairpin-like adapter is ligated to the other end of the double-stranded DNA molecule to allow sequencing of both strands of the DNA molecule in series. Exemplary Oxford Nanopore Y adapters are described by Karamitros et al. (2015) Nucleic Acids Res. 43(22): e!52; herein incorporated by reference.Kits

[0279] Also provided are kits comprising a vector for barcoding nuclei, a recombinant virion comprising such a vector, a vector system, or a cell line for producing a recombinant virion comprising the vector / vector system, as described herein. In some embodiments, a vector is provided with cells (e.g., already transfected with a vector / vector system or separately). Other agents may also be included in the kit such as transfection agents, suitable media for culturing cells, buffers, antibiotics, agents for inducing production of virions and expression of an expressible sequence encoding a product of interest (e.g., tetracyclin, doxycycline, tamoxifen), and the like.

[0280] In certain embodiments, the vector included in the kit is a viral vector. In some embodiments, the viral vector is an adenovirus associated virus (AAV) vector, lentivirus vector, or a G-deleted rabies virus (RVdG) vector.

[0281] In addition to the above components, the subject kits may further include (in certain embodiments) instructions for practicing the subject methods. In some embodiments, instructions for using the vectors, vector systems, virions, or cell lines to barcode nuclei of cells are provided in the kits. These instructions may be present in the subject kits in a variety of forms, one or more of which may be present in the kit. One form in which these instructions may be present is as printed information on a suitable medium or substrate, e.g., a piece or pieces of paper on which the information is printed, in the packaging of the kit, in a package insert, and the like. Yet another form of these instructions is a computer readable medium, e.g., diskette, compact disk (CD), DVD, flash drive, SD drive, and the like, on which the information has been recorded. Yet another form of these instructions that may be present is a website address which may be used via the internet to access the information at a removed site.EXAMPLES OF NON-LIMITING ASPECTS OF THE DISCLOSURE

[0282] Aspects, including embodiments, of the present subject matter described above may be beneficial alone or in combination, with one or more other aspects or embodiments. Without limiting the foregoing description, certain non-limiting aspects of the disclosure numbered 1-136 are provided below. As will be apparent to those of skill in the art upon reading this disclosure, each of the individually numbered aspects may be used or combined with any of the preceding or following individually numbered aspects. This is intended to provide support for all such combinations of aspects and is not limited to combinations of aspects explicitly provided below:1. A vector comprising: a first expression cassette comprising a first promoter operably linked to a first nucleotide sequence encoding a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain, wherein expression of the first nucleotide sequence results in production of the fusion protein in a cell; and a second expression cassette comprising a second promoter operably linked to a second nucleotide sequence encoding an RNA recognition sequence for the RNA binding domain and abarcode, wherein transcription of the second nucleotide sequence generates a barcoded RNA transcript comprising the RNA recognition sequence and the barcode, wherein binding of the RNA binding domain to the RNA recognition sequence results in formation of a complex between the fusion protein and the barcoded RNA transcript, wherein the complex is translocated to the nucleus of the cell.2. The vector of aspect 1, wherein the vector is a viral vector.3. The vector of aspect 2, wherein the viral vector is an adenovirus associated virus (AAV) vector, a lentivirus vector, or a G-deleted rabies vims (RVdG) vector.4. The vector of aspect 3, wherein the AAV vector further comprises a 5’-inverted terminal repeat (ITR) and a 3’-ITR, wherein the first expression cassette and the second expression cassette are positioned between the 5’-ITR and the 3’-ITR.5. The vector of any one of aspects 1-4, wherein the barcode is located in a 3’ untranslated region (3’ UTR) of the barcoded RNA transcript.6. The vector of any one of aspects 1-4, wherein the barcode is transcribed by a separate U6 promoter and DNA polymerase III.7. The vector of any one of aspects 1-6, wherein the nuclear localization domain comprises a nuclear localization sequence or a nuclear-localized protein.8. The vector of aspect 7, wherein the nuclear localization sequence is a histone H2B nuclear localization sequence.9. The vector of any one of aspects 1-6, wherein the nuclear localization domain comprises a Klarsicht-ANC-l-Syne homology (KASH) domain.10. The vector of any one of aspect 7-9, wherein the nuclear-localized protein is a histone H2B protein or a Klarsicht ANC-1 Sync Homology (KASH) domain-containing protein.11. The vector of any one of aspects 1-10, wherein the RNA-binding domain comprises one or more RNA-binding proteins.12. The vector of any one of aspects 1-11, wherein the RNA binding domain comprises an MS2 coat protein (MCP), and the RNA recognition sequence comprises an MS2 RNA aptamer, wherein the MCP binds to the MS2 RNA aptamer.13. The vector of aspect 12, wherein the RNA-binding domain comprises 1 to 6 MCPs, and the RNA recognition sequence comprises 1 to 30 MS2 RNA aptamers.14. The vector of any one of aspects 1-11, wherein the RNA binding domain comprises a lambda N peptide, and the RNA recognition sequence comprises a box B sequence, wherein the lambda N peptide binds to the box B sequence.15. The vector of aspect 14, wherein the RNA-binding domain comprises 1 to 6 lambda N peptides, and the RNA recognition sequence comprises 1 to 30 box B sequences.16. The vector of any one of aspects 1-11, wherein the RNA binding domain comprises a bacteriophage PP7 coat protein (PCP), and the RNA recognition sequence comprises a PCP binding site, wherein the PCP binds to the PCP binding site.17. The vector of aspect 16, wherein the RNA-binding domain comprises 1 to 6 PCPs, and the RNA recognition sequence comprises 1 to 30 PCP binding sites.18. The vector of any one of aspects 1-11, wherein the RNA binding domain comprises a BglG transcriptional antiterminator, and the RNA recognition sequence comprises a BglG binding site, wherein the BglG transcriptional antiterminator binds to the BglG binding site.19. The vector of aspect 18, wherein the RNA-binding domain comprises 1 to 6 BglG transcriptional antiterminators, and the RNA recognition sequence comprises 1 to 30 BglG binding sites.20. The vector of any one of aspects 1-11, wherein the RNA binding domain comprises spliceosomal protein U1A (U1A), and the RNA recognition sequence comprises a U1A binding site, wherein the U1A binds to the U1A binding site.21. The vector of aspect 20, wherein the RNA-binding domain comprises 1 to 6 UlAs, and the RNA recognition sequence comprises 1 to 30 U1A binding sites.22. The vector of any one of aspects 1-21, further comprising a polyadenylation site.23. The vector of aspect 22, further comprising a pseudoknot RNA motif sequence, wherein the pseudoknot RNA motif sequence is positioned between the barcode and the polyadenylation site.24. The vector of aspect 23, wherein the pseudoknot RNA motif sequence is an evopreql pseudoknot RNA motif sequence.25. The vector of any one of aspects 1-24, wherein the second nucleotide sequence further encodes a pair of ribozymes comprising a first ribozyme and a second ribozyme, wherein the first ribozyme and the second ribozyme flank the RNA recognition sequence and the barcode in the barcoded RNA transcript, wherein the first ribozyme and the second ribozyme undergo autocatalytic cleavage to generate a 5 ’-end comprising a hydroxyl group and a 3 ’-end comprising a 2’,3’-cyclic phosphate group, wherein the 5’-end and the 3’-end are ligated by an endogenous RtcB RNA ligase in a cell resulting in circularization of the barcoded RNA transcript.26. The vector of aspect 25, wherein the pair of ribozymes are Twister ribozymes.27. The vector of any one of aspects 1 -26, wherein the first promoter and the second promoter arc constitutive or inducible.28. The vector of any one of aspects 1-27, wherein the first promoter is a hybrid cytomegalovirus (CMV) / chicken 0-actin promoter (CAG) promoter.29. The vector of any one of aspects 1-28, wherein the second promoter is a U6 promoter.30. The vector of any one of aspects 1-29, wherein the fusion protein further comprises a selectable marker.31. The vector of aspect 30, wherein the selectable marker is a fluorescent protein.32. The vector of aspect 31, wherein the fluorescent protein is a green fluorescent protein, a red fluorescent protein, a blue fluorescent protein, a yellow fluorescent protein, an orange fluorescent protein, or a violet fluorescent protein.33. The vector of any one of aspects 1-32, further comprising a pair of restriction sites comprising a first restriction site and a second restriction site, wherein the first restriction site and the second restriction site flank the second expression cassette.34. The vector of any one of aspects 1-33, further comprising a multiple cloning site (MCS) for insertion of an expressible sequence into the vector.35. The vector of any one of aspects 1-34, further comprising an expressible sequence, wherein the barcode identifies the expressible sequence.36. The vector of aspect 35, wherein the expressible sequence encodes a protein or anRNA.37. The vector of aspect 36, wherein the protein is a therapeutic protein or a genomeediting enzyme.38. The vector of aspect 36 or 37, wherein the protein is a variant.39. The method of aspect 37, wherein the genome-editing enzyme is a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) nuclease, a meganuclease, a zinc-finger nuclease (ZFN), or a transcription activator-like effector nuclease (TALEN).40. The method of aspect 39, wherein the Cas nuclease is Cas9 or Cas 12a.41. The vector of aspect 36, wherein the RNA is a messenger RNA (mRNA) or a non-coding RNA.42. The vector of aspect 41, wherein the non-coding RNA is a microRNA (miRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a small nuclear RNA (snRNA), a piwi-interacting RNA (piRNA), a small nucleolar RNA (snoRNA), or a long noncoding RNA (IncRNA).43. The vector of aspect 36, wherein the RNA is a guide RNA or an aptamer.44. The vector of any one of aspects 1-43, wherein the second nucleotide sequence further comprises a sequence that is complementary to an oligonucleotide capture probe.45. A vector comprising an expression cassette comprising a promoter operably linked to a nucleotide sequence encoding a nuclear retention element connected to a barcode, wherein transcription of the nucleotide sequence in a cell generates a barcoded RNA transcript comprising the nuclear retention element connected to the barcode, wherein the barcoded RNA transcript is translocated to the nucleus of the cell.46. The vector of aspect 45, wherein the nuclear retention element is a nuclear retention element of a long non-coding RNA.47. The vector of aspect 46, wherein the long non-coding RNA is MEG3, XIST, MALAT1, or SIRLOIN.48. The vector of aspect 46 or 47, wherein the nuclear retention element comprises the entire long non-coding RNA or a fragment thereof, wherein the barcoded RNA transcript is translocated to the nucleus of the cell.49. The vector of aspect 46, wherein the nuclear retention element comprises the nucleotide sequence of RCCTCCC, wherein R is an A or a G.50. The vector of any one of aspects 45-48, wherein the vector is a viral vector.51. The vector of aspect 50, wherein the viral vector is an adenovirus associated virus (AAV) vector, a lentivirus vector, or a G-deleted rabies virus (RVdG) vector.52. The vector of aspect 51, wherein the AAV vector further comprises a 5 ’-inverted terminal repeat (ITR) and a 3’-ITR, wherein the expression cassette is positioned between the 5’- ITR and the 3 ’-ITR.53. The vector of any one of aspects 45-52, wherein the barcode is located in a 3’ untranslated region (3’ UTR) of the barcoded RNA transcript.54. The vector of any one of aspects 45-52, wherein the barcode is transcribed by a separate U6 promoter and DNA polymerase III.55. The vector of any one of aspects 45-54, further comprising a polyadenylation site.56. The vector of aspect 55, further comprising a pseudoknot RNA motif sequence, wherein the pscudoknot RNA motif sequence is positioned between the barcode and the polyadenylation site.57. The vector of aspect 56, wherein the pseudoknot RNA motif sequence is an evopreql pseudoknot RNA motif sequence.58. The vector of any one of aspects 45-57, wherein the nucleotide sequence encoding the nuclear retention element connected to the barcode further encodes a pair of ribozymes comprising a first ribozyme and a second ribozyme, wherein the first ribozyme and the second ribozyme flank the nuclear retention element and the barcode in the barcoded RNA transcript, wherein the first ribozyme and the second ribozyme undergo autocatalytic cleavage to generate a 5 ’-end comprising a hydroxyl group and a 3 ’-end comprising a 2’, 3 ’-cyclic phosphate group, wherein the 5 ’-end and the 3 ’-end are ligated by an endogenous RtcB RNA ligase in a cell resulting in circularization of the barcoded RNA transcript.59. The vector of aspect 58, wherein the pair of ribozymes are Twister ribozymes.60. The vector of any one of aspects 45-59, wherein the promoter is constitutive or inducible.61. The vector of any one of aspects 45-60, wherein the promoter is a U6 promoter.62. The vector of any one of aspects 45-61, wherein the nucleotide sequence further encodes a selectable marker.63. The vector of aspect 62, wherein the selectable marker is a fluorescent protein.64. The vector of aspect 63, wherein the fluorescent protein is a green fluorescent protein, a red fluorescent protein, a blue fluorescent protein, a yellow fluorescent protein, an orange fluorescent protein, or a violet fluorescent protein.65. The vector of any one of aspects 45-64, further comprising a pair of restriction sites comprising a first restriction site and a second restriction site, wherein the first restriction site and the second restriction site flank the expression cassette.66. The vector of any one of aspects 45-65, further comprising a multiple cloning site (MCS) for insertion of an expressible sequence into the vector.67. The vector of any one of aspects 45-66, further comprising an expressible sequence, wherein the barcode identifies the expressible sequence.68. The vector of aspect 67, wherein the expressible sequence encodes a protein or an RNA.69. The vector of aspect 68, wherein the protein is a therapeutic protein or a genomeediting enzyme.70. The vector of aspect 68 or 69, wherein the protein is a variant.71. The method of aspect 69, wherein the genome-editing enzyme is a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) nuclease, a meganuclease, a zinc-finger nuclease (ZFN), or a transcription activator-like effector nuclease (TALEN).72. The method of aspect 71, wherein the Cas nuclease is Cas9 or Casl2a.73. The vector of aspect 68, wherein the RNA is a messenger RNA (mRNA) or a non-coding RNA.74. The vector of aspect 73, wherein the non-coding RNA is a microRNA (miRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a small nuclear RNA(snRNA), a piwi-interacting RNA (piRNA), a small nucleolar RNA (snoRNA), or a long noncoding RNA (IncRNA).75. The vector of aspect 68, wherein the RNA is a guide RNA or an aptamer.76. The vector of any one of aspects 45-75, wherein the nucleotide sequence further comprises a sequence that is complementary to an oligonucleotide capture probe.77. A cell transfected with the vector of any one of aspects 1-76.78. The cell of aspect 77, wherein the cell is a mammalian cell.79. The cell of aspect 78, wherein the mammalian cell is a human cell.80. The cell of aspect 78 or 79, wherein the cell is a neuron.81. A method of barcoding nuclei in a population of cells, the method comprising: providing a plurality of vectors according to any one of aspects 1-76, wherein the barcode in each vector of the plurality comprises a unique identifier sequence; and transfecting the population of cells with the plurality of vectors to generate cells having barcoded nuclei.82. The method of aspect 81, wherein said transfecting is performed in vitro, ex vivo, or in vivo.83. The method of aspect 81 or 82, wherein the population of cells is in a tissue or an organ.84. A method of multiplex screening of a population of cells, the method comprising: providing a plurality of vectors according to any one of aspects 1-76, wherein the barcode in each vector of the plurality comprises a unique identifier sequence;transfecting the population of cells with the plurality of vectors to generate cells having barcoded nuclei; measuring morphological or functional characteristics of the cells having the barcoded nuclei; isolating the barcoded nuclei; and identifying the barcode in each barcoded nucleus.85. The method of aspect 84, wherein the population of cells is in a tissue or an organ.86. The method of aspect 85, wherein the tissue is nervous tissue.87. The method of aspect 84, wherein the population of cells is in a culture.88. The method of any one of aspects 84-87, wherein the population of cells comprises neurons or glial cells or a combination thereof.89. The method of aspect 88, further comprising optogenetically modifying the neurons.90. The method of any one of aspects 84-89, wherein said measuring the morphological or functional characteristics comprises performing gene expression profiling, microscopy, calcium imaging, an electrophysiology measurement, functional neuroimaging, a migration assay, an axonal growth and pathfinding assay, monosynaptic tracing, a phagocytosis assay, an enzymatic assay, a cell receptor assay, an ion channel assay, a signal transduction assay, or a cell secretion assay.91. The method of aspect 90, wherein the gene expression profiling comprises performing microarray analysis, RNA sequencing, or quantitative polymerase chain reaction.92. The method of aspect 91, wherein the RNA sequencing comprises single nucleus RNA sequencing, single cell RNA sequencing, or a combination thereof.93. The method of aspect 90, wherein the microscopy is confocal microscopy, atomic force microscopy, super-resolution microscopy, light-sheet microscopy, two-photon microscopy, or fluorescence microscopy.94. The method of aspect 90 wherein the electrophysiology measurement is patch clamping, electroencephalography (EEG), or magnetoencephalography (MEG).95. The method of aspect 90, wherein the functional neuroimaging is functional magnetic resonance imaging (fMRI), positron emission tomography (PET), functional nearinfrared spectroscopy (fNIRS), single-photon emission computed tomography (SPECT), or functional ultrasound imaging (fUS).96. The method of any one of aspects 84-95, wherein the morphological or functional characteristics are measured in the population of cells in a tissue of a live subject in vivo or in culture in vitro.97. The method of aspect 96, wherein the live subject is a nonhuman animal.98. The method of aspect 96 or 97, further comprising removing the tissue from the subject after said measuring the morphological or functional characteristics.99. The method of any one of aspects 84-98, wherein said isolating the barcoded nuclei comprises performing fluorescence activated nuclei sorting (FANS).100. The method of aspect 99, further comprising isolating the barcoded RNA transcript from each barcoded nucleus.101. The method of aspect 100, wherein said isolating the barcoded RNA transcript comprises hybridization of the barcoded RNA transcript to a capture probe comprising a sequence complementary to a region of the barcoded RNA transcript.102. The method of aspect 100, wherein said isolating the barcoded RNA transcript comprises hybridization of the barcoded RNA transcript to a capture probe comprising a poly-T sequence complementary to the poly-A tail of the barcoded RNA transcript.103. The method of aspect 101 or 102, wherein the capture probe is immobilized on a solid support.104. The method of aspect 103, wherein the solid support is a magnetic bead or polystyrene bead.105. The method of any one of aspects 84-104, wherein said identifying comprises sequencing the barcode sequence.106. The method of aspect 105, wherein said sequencing comprises performing single nucleus RNA sequencing (snRNA-seq).107. The method of any one of aspects 84-106, further comprising amplifying the barcode sequence prior to said sequencing.108. The method of aspect 107, wherein said amplifying comprises performing reverse transcription polymerase chain reaction or isothermal amplification.109. The method of any one of aspects 84-108, further comprising reverse transcribing the barcoded RNA transcript to generate a complementary DNA (cDNA) copy of the barcoded RNA transcript.110. The method of aspect 109, further comprising adding adapters to the 5’ end and the 3’ end of the cDNA copy of the barcoded RNA transcript.111. The method of any one of aspects 84-110, where each vector of the plurality further comprises a different expressible sequence, wherein the barcode in each vector identifies the expressible sequence.112. The method of aspect 111, wherein the expressible sequence encodes a protein or an RNA.113. The method of aspect 112, wherein the protein is a therapeutic protein or a genome-editing enzyme.114. The method of aspect 123, wherein the therapeutic protein is a hormone, a cytokine, a chemokine, a growth factor, an enzyme, or an antibody.115. The method of any one of aspects 112-114, wherein the protein or RNA is a variant.116. The method of aspect 115, wherein the genome-editing enzyme is a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) nuclease, a meganuclease, a zinc-finger nuclease (ZFN), or a transcription activator-like effector nuclease (TALEN).117. The method of aspect 116, wherein the Cas nuclease is Cas9 or Casl2a.118. The method of aspect 112, wherein the RNA is a messenger RNA (mRNA) or a non-coding RNA.119. The method of aspect 118, wherein the non-coding RNA is a microRNA (miRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a small nuclear RNA (snRNA), a piwi-interacting RNA (piRNA), a small nucleolar RNA (snoRNA), or a long non-coding RNA (IncRNA).120. The method of aspect 1 12, wherein the RNA is a guide RNA or an aptamer.121. The method of aspect 112, wherein the protein is a fluorescent protein or a bioluminescent protein.122. The method of aspect 121, further comprising imaging the fluorescent protein or the bioluminescent protein, wherein a location of a cell expressing the fluorescent protein or the bioluminescent protein is determined from the imaging.123. The method of aspect 122, further comprising mapping the location of the cell expressing the fluorescent protein or the bioluminescent protein onto a reference image of the tissue.124. A vector system comprising the vector of any one of aspects 1-76 and a helper virus vector.125. The vector system of aspect 124, wherein the helper virus vector encodes Ela and Elb or E2a and E4.126. The vector system of aspect 124, wherein the helper virus vector encodes a glycoprotein (G) gene.127. The vector system of any one of aspects 124-126, further comprising a vector encoding one or more capsid proteins.128. The vector system of aspect 127, wherein the one or more capsid proteins comprise VP1, VP2, and VP3.129. The vector system of any one of aspects 124-128, further comprising a vector encoding an AAV Rep protein.130. The vector system of aspect 124, wherein the helper virus vector encodes Gag, Pol, Rev, and VSV-G.131. A cell transfected with the vector system of any one of aspects 124-130.132. The cell of aspect 131, wherein the cell is a mammalian cell.133. The cell of aspect 132, wherein the mammalian cell is a human cell.134. The cell of any one of aspects 131-133, wherein the cell is a neuron.135. A vector library comprising a plurality of vectors according to any one of aspects 1-76, wherein the plurality of vectors collectively comprise a plurality of expressible sequences, and wherein the barcode in each vector identifies the expressible sequence in the vector.136. The vector library of aspect 135, wherein the plurality of expressible sequences collectively encode different proteins, protein variants, antibodies, mRNAs, shRNAs, siRNAs, aptamers, or guide RNAs for RNA-guided nucleases.EXAMPLES

[0283] The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the disclosed subject matter, and are not intended to limit the scope of what the inventors regard as their invention nor are they intended to represent that the experiments below are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperature, etc.) but some experimental errors and deviations should be accounted for.Example 1: Plasmid Construction for Nuclei-Capture Compatible Feature Barcodes

[0284] The AAV transgene plasmids were based on an AAV vector from the Schaffer Lab of UC Berkeley. The hybrid CMV / chicken -actin promoter (CAG) promoter driving eGFP expression inserted into the AAV vector downstream of the 5’ ITR was derived from the STICRlentiviral plasmid described by Ryan Delgado in Delgado et al. (2022). The transgene H2B-GFP- MCP was constructed via Gibson assembly of PCR fragments for H2B-GFP, and MCP. The enzyme mix used for Gibson assembly was NEBuilder® HiFi DNA Assembly Master Mix (NEB) and the mix used for PCR was Q5® High-Fidelity 2X Master Mix (NEB). H2B-GFP was amplified via PCR from a plasmid in the Nowakowski Lab and MCP was amplified via PCR from a plasmid gifted from the DeRisi Lab of UCSF. T2A and P2A linker sequences were included in the 3’ overhangs of the PCR primers for H2B-GFP and GFP-MCP respectively. The WPRE and bGH polyadenylation signal were also derived from the STICR plasmid (Delgado et al. 2022) and inserted via Gibson assembly downstream of H2B-GFP-MCP. The U6 promoter and MCS barcode array were synthesized and ordered from IDT and inserted via Gibson assembly upstream of the 3’ AAV ITR. Six MS2-aptamer repeats were amplified via PCR from Addgene plasmid # 84561 and inserted via Gibson assembly downstream of the U6 promoter and upstream of the 3’ barcode array. The sequences for aptamer evopreql and polyT termination-tract were included in the 3’ overhangs of PCR primers and inserted downstream of the lOx Capture Sequence 1 for the U6- driven barcode. To construct the 5 ’ -capture-compatible barcoded AAV vector, a barcode array gBlock was designed in Snapgene with the following elements constructed in 5’ to 3’ order: Apal restriction site, U6 promoter, barcode array, Cas9 Scaffold (sequence given by lOx in 5’ capture kit V3), Capture Sequence 1, six MS2-aptamer repeats, Bglll restriction site. The dsDNA fragment above was synthesized by IDT. The 3’ capture-compatible AAV vector described in the previous paragraph was digested with Apal and Bglll along with the IDT 5’ gBlock fragment and the fragments were ligated to construct the 5 ’-compatible vector. To facilitate multiplexed infections of different AAV valiants, the AAV transgene vectors were digested with BbsI and 36 nucleotide barcodes flanked by compatible sticky-ends were ligated into the AAV transgene plasmids.

[0285] The lentivirus plasmids were derived from the STICR plasmid (Delgado et al. 2022) and inserted via Gibson assembly downstream of H2B-GFP-MCP. The transgene H2B-GFP-MCP was constructed via Gibson assembly of PCR fragments for H2B-GFP, and MCP. The enzyme mix used for Gibson assembly was NEBuilder® HiFi DNA Assembly Master Mix (NEB) and the mix used for PCR was Q5® High-Fidelity 2X Master Mix (NEB). H2B-GFP was amplified via PCR from a plasmid in the Nowakowski Lab and MCP was amplified via PCR from a plasmid gifted from the DeRisi Lab of UCSF. The U6 promoter, tornado RNA circularization sites, and MCS barcode array were synthesized and ordered from IDT and inserted via Gibson assembly.The sequences for aptamer evopreql and polyT termination-tract were included in the 3’ overhangs of PCR primers and inserted downstream of the lOx Capture Sequence 1 for the U6- driven barcode.

[0286] The G-deleted rabies virus (RVdG) is a genetically modified non-segmented negative- stranded RNA virus, engineered to lack the glycoprotein gene (G gene). This modification requires the provision of the glycoprotein in trans for the virus to spread from cell to cell, enabling retrograde or monosynaptic tracing. Recent work has enabled simultaneous readout of connectivity and molecularly-defined cell type identity after monosynaptic tracing with RVdG by inserting a high complexity molecular barcode library into one of the 3’ UTRs of the RVdG genome. However, given that rabies virus replicates and undergoes transcription in the cytoplasm, the inability to read barcodes from nuclei necessitates whole cell dissociations, a process that can compromise neuron viability and integrity, thereby affecting the accuracy and reliability of connectivity mapping. To address this issue, we engineered a barcoded RVdG to be compatible with single nuclei RNA sequencing (snRNA-seq). To prepare a RVdG genome plasmid for barcoding and nuclear capture, the RabV CVS-N2c(deltaG)-mCherry plasmid (Addgene #73464) was digested with Smal (NEB, R0141) and Nhell-HF (NEB, R3131). Gibson assembly using NEB HiFi Assembly master mix was used to replace the gene encoding mCherry with a synthesized double-stranded gene fragment (IDT) encoding a fusion protein consisting of the MS2 coat protein MCP, the fluorescent protein oScarlet, and either the KASH or H2B nuclear localization signal. The 3’ UTR of these gene fragments carried six MS2 hairpin sequences followed by a barcode cloning site represented by two SacII restriction enzyme sites. A high-complexity barcode library was subsequently cloned into the resulting RVdG plasmids. Plasmids were digested with SacII (NEB, R0157) overnight and purified using the Zymo DNA Clean and Concentrator-5 kit (Zymo, D4004). The insert containing the barcode library was generated by performing PCR on the STICR barcoded lentiviral plasmid library (Addgene 180483), with overhangs matching the SacII digested plasmid. Gibson assembly (16 reactions) was performed using NEB HiFi Assembly master mix, in which a single reaction contains 250 ng of digested backbone and 1:2 molar excess of PCR- amplified insert. The gibson assembly reactions were incubated at 50°C for one hour and DNA was purified using phenol-chloroform extraction and ethanol precipitation. DNA from gibson assembly reactions was electroporated into MegaX DH10B T1R electrocompetent bacteria (500 ng of DNA per 50 pL of bacteria per electroporation).

[0287] Additional constructs were similarly cloned that facilitate nuclear import and / or retention, including: (1) the lambda N pcptidc / BoxB system derived from bacteriophage lambda, in which the lambda N peptide is fused to H2B or KASH nuclear localization signal to bind BoxB hairpins flanking feature barcodes, (2) the PP7 / PCP system derived from PP7 bacteriophage in which the PP7 coat protein (PCP) is fused to H2B nuclear localization signal to bind PCP binding sites flanking feature barcodes, and (3) natural nuclei retention elements identified in long noncoding RNAs including from MEG3, XIST, MALAT1, and the SIRLOIN consensus sequence, which were cloned flanking feature barcodes.AAV production

[0288] HEK293T cells were seeded in a 15 -cm dish (Corning) at a density of -60% confluence. 24 hours later, a triple-transfection using PEI (3:1 ratio of PEI volume to DNA mass) was performed with the rep-cap pXX5 plasmid, pHelper plasmid, and barcoded AAV transgene plasmid. 72 hours after transfection, the cells were scraped from the dish and lysed in buffer (0.15M NaCl, 50 mM Tris-HCl, 0.05% Tween). The crude lysate was further purified using lodixanol gradients of 15%, 25%, 40%, and 60% lodixanol (Optiprep, Sigma) and subsequent centrifugation in Amicon Ultra- 15 tubes (MilliporeSigma) previously washed with PBS + Tween (Sigma- Aldrich) .Lentivirus production

[0289] HEK293 cells were first transfected with the barcode library along with lentiviral helper plasmids psPAX2 (Addgene, #12260) and pMD2.G (Addgene, #12259) using Lipofectamine™ 3000 Transfection Reagent (Thermo Fisher, L3OOOOO8). HEK293 media were changed 12 h after transfection and replaced with 30 ml low-FBS cell culture media (DMEM (Thermo Fisher, #11965092), 2% FBS (Hyclone, # SH30071.03), 1% Penicillin-Streptomycin (Thermo Fisher, # 15070063)). After an additional 48 h, media were collected, concentrated with an ultracentrifuge, and then resuspended in 50-100 pl of sterile PBS.EnvA-pseudotyped RVdG production

[0290] Tissue culture plates were coated with 10 pg / mL poly-D-lysine diluted in PBS for 1 hour at room temperature, and rinsed twice with PBS. HEK-GT cells (Sumser et al. (2022) Elifel l :e79848) were dissociated using 0.05% Trypsin-EDTA (Thermo, 25200056) and 20 million cells were plated into each PEI-coatcd 15-cm dish in DMEM (Thermo, 41966) containing 10% Heat-inactivated FBS (HI-FBS, Cytiva, SH30071.03HI) and 0.2% primocin (Invivogen, ant-pm- 1). The next day, when the cells had reached -90% confluence, each 15 cm dish was transfected with 55.1 pg of barcoded pSADdeltaG-dTomato, 27.1 of SADB19-N (Addgene #32630), 13.9 pg of SADB19-P (Addgene #32631), 13.9 pg of SADB19-L (Addgene #32632), and 16.6 pg of CAGGS-T7opt (Addgene #65974) using Lipofectamine 3000. The day following transfection, cells were passaged at a 1.5:1 ratio based on surface area. Media was replaced every other day until day of collection. When -40-50% of cells were producing rabies virus based on dTomato expression, viral supernatant was collected and filtered through a 0.45 pm syringe filter (Corning 431220). Supernatant was concentrated using PEG precipitation with 1 volume of PEG concentrator for every 3 volumes of supernatant (PEG concentrator: 40% (w / v) PEG 8000 (Promega, V3011) in lx PBS (Thermo, 10010023) supplemented with 1.2 M NaCl (Fisher, S271- 500), as per recommendations of MD Anderson Functional Genomics Core. The next day, the viral pellet was resuspended using DMEM / F12-HEPES supplemented with 10% FBS, 2% primocin, and 8 pg / mL of polybrene (Millipore Sigma TR-1003-G) at 1:100 of the original media volume.

[0291] Per every mL of concentrated G-RVdG virus, 1 million BHK-EnvA cells (Salk Institute Viral Core) were transduced by spin-infection (800xg, 1 h, room temperature, using a swinging bucket rotor). Subsequently, BHK-EnvA cells were plated into 15 cm dishes at a density of 5 million cells / dish in DMEM containing 10% FBS, 2% primocin, and 8 pg / mL of polybrene. The following day, the media was replaced with fresh media without polybrene. One or two days later, when the plate reached confluence, the cells were passaged once more using 0.25% trypsin-EDTA for 10 minutes at 37°C, in order to deactivate any residual G-pseudotyped virus. Cells were spun down to remove the trypsin-EDTA and passaged at a 1:4 dilution into fresh 15 cm dishes. Cells were given a media change the next day. One or two days later, when plates had reached confluence, viral supernatant was collected every other day for a total of 3-4 collections, depending on the health of the cells. On the day of collection, viral supernatant was filtered through a 0.45 pm syringe filter and concentrated by ultracentrifuge (Beckman Coulter) through a 4 mL sucrose cushion using the SW32Ti rotor (20,300 RPM, 2 h, 4C). After ultracentrifugation, the viral pellet was resuspended in 50 pL of HBSS while gently rocking overnight at 4C. Virus was aliquoted, frozen at -80C, and titered by infecting TVA800 HEK293T cells (Salk Institute Viral Core) usinga serial dilution of virus, and measuring the percentage of dTomato-positive cells two days later by flow cytometry.Viral transduction of organotypic slice cultures of primary cortical tissue

[0292] To generate organotypic slice cultures, cortical tissue was embedded in 3% low- melting point agarose and sliced at a thickness of 300um using a Leica VT1600S vibratome. Slicing was performed in artificial cerebrospinal fluid (ACSF) containing 125 mM NaCl, 2.5 mM KC1, 1 mM MgCh, 2 mM CaCl2, 1.25 mM NaH2PO4, 25 mM NaHCCh, and 25 mM D-(+)- glucose, which had been previously bubbled with carbogen (95% 02 / 5% CO2). Slices were cultured at the air-liquid interface on cell culture inserts (Millipore, #PICM0RG50) in media containing 60% Basal Medium Eagle, 32% Hanks Buffered Saline Solution, 5% heat-inactivated fetal bovine serum, 1% glucose, 1% N2 and 0.2% primocin.

[0293] Transduction with AAV was performed by directly adding approximately 3 pL of 1012- 1013genomic copies (gc) / mL virus directly to the surface of the slice, one day after plating. Slices were cultured for up to 7 days after transduction to confirm fluorescence before proceeding with nuclear isolation. Transduction with lentivirus was performed by directly adding virus directly to the surface of the slice, one day after plating. Slices were cultured for up to 7 days after transduction to confirm fluorescence before proceeding with nuclear isolation. Transduction with EnvA-pseudotyped RVdG requires a prior infection with a virus expressing TVA, which is the cognate receptor for EnvA, and the rabies G, which is required for trans synaptic spread. The day of plating, slices were transduced with 3 pL of “helper” AAVs carrying these transgenes (AAV- syn-TVA66t-mStayGold2-Cre, IxlO11gc / mL; AAV-CAG-FLEX-N2cG, 2xl012gc / mL). Seven days later, slices were infected with EnvA-RVdG and cultured for an additional five days before proceeding with nuclear isolation.Nuclei extraction

[0294] Tissue dounce and pestles were prepared by sequentially washing with EtOH, RNAse Zap and RNAse-free water (3 washes), and then precooled on ice. A tissue dounce was filled with 2 ml of ice-cold nucleus isolation buffer (EZ lysis buffer Sigma-Aldrich, #NUC101-lKT) and 1% Kollidon VA64. Pieces of tissue were directly placed into the buffer in the dounce. The tissue was mechanically disrupted using 10-20 strokes with pestle A, followed by 120 strokes with pestle B.The homogenized solution was transferred to a chilled protein low-binding tube (Eppendorf, #0030122216). The douncc was then rinsed with 2 ml of ice-cold nucleus isolation buffer, and the rinsate was transferred to the same tube. The solution was gently mixed, incubated on ice for 5 minutes, then centrifuged at 500g for 5 minutes at 4°C. The supernatant was discarded, and the pellet was resuspended in 4 ml of EZ lysis buffer by gentle pipetting, incubated for 5 minutes, and centrifuged again at 500g for 5 minutes at 4°C. The pellet was then resuspended in 0.5 ml of nucleus suspension buffer (0.1% BSA in lx PBS with 0.2 U / pl RNase inhibitor) and filtered through a 30 pm cell strainer into a new FACS tube. Hoechst (BioRad, #1351304) was added to nuclei at 1:100 dilution. Nuclei carrying feature barcodes were isolated based on their expression of H2B-GFP-MCP and Hoechst using fluorescence activated nuclei sorting (FANS) using an 85 um nozzle on a BD Aria Fusion into 20uL of HBSS + 0.1% BSA.Library preparation of snRNA-seq transcriptome and feature barcodes

[0295] Preparation of snRNA-seq libraries was carried out using a Chromium Single Cell 3' GEM, Library & Gel Bead Kit v3.1 (lOx Genomics, #PN-1000268) and 3’ Feature Barcode Kit (lOx Genomics, #PN- 1000262), or Chromium GEM-X Single Cell 3' Kit v4 (lOx Genomics, # PN-1000691) and 3’ Feature Barcode Kit (lOx Genomics, #PN- 1000262), or using PIPseq T2 3’ Single Cell RNA Kit v4PLUS (FBS-SCR-T2-8-V4.05-FLU). Transcriptome libraries were prepared according to the manufacturer’s protocol. For preparing feature barcode libraries, the cDNA amplification step following reverse transcription was modified to include a spike-in primer targeting the 5-prime sequence immediately proximal to feature barcodes at a final concentration of 0.25 pM. The sequence for spike-in primer was CCTTGGCACCCGAGAATTCCA (SEQ ID NO:1) for AAV feature barcodes, AACCATGCCGAGTGCGG (SEQ ID NO:2) for lentivirus feature barcodes, and TAGCAAACTGGGGCACAAGC (SEQ ID NO:3) for RVdG feature barcodes. AAV feature barcodes, which are expressed from U6 promoter and captured using the 10X Genomics capture sequence, were recovered from the supernatant fraction of the 0.6X SPRIselect cleanup after cDNA amplification. RVdG feature barcodes, which are embedded in the 3’ UTR of the fourth open reading frame of the rabies genome and is captured by the polyA tail, were recovered from the pellet / transcriptome library fraction of the SPRIselect cleanup. Sequencing-ready barcode libraries were prepared by performing a one-step PCR reaction using Q5 2x HotStart Mastermix. PCR with primers carrying Illumina P7 and P5 handles and Truseqread adapters was performed with 25% of transferred supernatant cleanup (AAV feature barcodes) or transcriptomc library fraction (RVdG feature barcodes) as a DNA template at a final concentration of 0.5 pM. Amplification reactions were performed as follows: (1) 98 °C for 30 s; (2) 98 °C for 10 s, 62 °C for 20 s, 72C for 10 s; repeated for 6-15 cycles depending on the input amount; (3) 72 °C 2 min. The PCR product was purified by using 0.6X-0.8X double-sided SPRI selection and was ready for sequencing. For 5’ recovery strategy of AAV feature barcodes, Preparation of snRNA-seq libraries was carried out using a Chromium GEM-X Single Cell 5' Kit v3 (lOx Genomics, # PN- 1000695) and Chromium GEM-X Single Cell 5' Feature Barcode Kit v3 (lOx Genomics, #PN1000703). Transcriptome libraries were prepared according to the manufacturer’s protocol. Feature barcodes were recovered by following the manufacturer’s protocol in step 6 for CRISPR screening library construction.

[0296] The above examples are provided to illustrate the invention but not to limit its scope. Other variants of the invention will be readily apparent to one of ordinary skill in the art and are encompassed by the appended claims. All publications, accessions, references, databases, and patents cited herein are hereby incorporated by reference for all purposes.

Claims

CLAIMSWhat is claimed is:

1. A vector comprising: a first expression cassette comprising a first promoter operably linked to a first nucleotide sequence encoding a fusion protein comprising a nuclear localization domain connected to an RNA-binding domain, wherein expression of the first nucleotide sequence results in production of the fusion protein in a cell; and a second expression cassette comprising a second promoter operably linked to a second nucleotide sequence encoding an RNA recognition sequence for the RNA binding domain and a barcode, wherein transcription of the second nucleotide sequence generates a barcoded RNA transcript comprising the RNA recognition sequence and the barcode, wherein binding of the RNA binding domain to the RNA recognition sequence results in formation of a complex between the fusion protein and the barcoded RNA transcript, wherein the complex is translocated to the nucleus of the cell.

2. The vector of claim 1, wherein the vector is a viral vector.

3. The vector of claim 2, wherein the viral vector is an adenovirus associated virus (AAV) vector, a lentivirus vector, or a G-deleted rabies vims (RVdG) vector.

4. The vector of claim 3, wherein the AAV vector further comprises a 5’-inverted terminal repeat (ITR) and a 3’-ITR, wherein the first expression cassette and the second expression cassette are positioned between the 5’-ITR and the 3’-ITR.

5. The vector of any one of claims 1-4, wherein the barcode is located in a 3’ untranslated region (3’ UTR) of the barcoded RNA transcript.

6. The vector of any one of claims 1-4, wherein the barcode is transcribed by a separate U6 promoter and DNA polymerase III.

7. The vector of any one of claims 1-6, wherein the nuclear localization domain comprises a nuclear localization sequence or a nuclear-localized protein.

8. The vector of claim 7, wherein the nuclear localization sequence is a histone H2B nuclear localization sequence.

9. The vector of any one of claims 1-6, wherein the nuclear localization domain comprises a Klarsicht-ANC-l-Syne homology (KASH) domain.

10. The vector of any one of claim 7-9, wherein the nuclear-localized protein is a histone H2B protein or a Klarsicht ANC-1 Syne Homology (KASH) domain-containing protein.

11. The vector of any one of claims 1-10, wherein the RNA-binding domain comprises one or more RNA-binding proteins.

12. The vector of any one of claims 1-11, wherein the RNA binding domain comprises an MS2 coat protein (MCP), and the RNA recognition sequence comprises an MS2 RNA aptamer, wherein the MCP binds to the MS2 RNA aptamer.

13. The vector of claim 12, wherein the RNA-binding domain comprises 1 to 6 MCPs, and the RNA recognition sequence comprises 1 to 30 MS2 RNA aptamers.

14. The vector of any one of claims 1-11, wherein the RNA binding domain comprises a lambda N peptide, and the RNA recognition sequence comprises a box B sequence, wherein the lambda N peptide binds to the box B sequence.

15. The vector of claim 14, wherein the RNA-binding domain comprises 1 to 6 lambda N peptides, and the RNA recognition sequence comprises 1 to 30 box B sequences.

16. The vector of any one of claims 1- 11 , wherein the RNA binding domain comprises a bacteriophage PP7 coat protein (PCP), and the RNA recognition sequence comprises a PCP binding site, wherein the PCP binds to the PCP binding site.

17. The vector of claim 16, wherein the RNA-binding domain comprises 1 to 6 PCPs, and the RNA recognition sequence comprises 1 to 30 PCP binding sites.

18. The vector of any one of claims 1-11, wherein the RNA binding domain comprises a BglG transcriptional antiterminator, and the RNA recognition sequence comprises a BglG binding site, wherein the BglG transcriptional antiterminator binds to the BglG binding site.

19. The vector of claim 18, wherein the RNA-binding domain comprises 1 to 6 BglG transcriptional anti terminators, and the RNA recognition sequence comprises 1 to 30 BglG binding sites.

20. The vector of any one of claims 1-11, wherein the RNA binding domain comprises spliceosomal protein U1A (U1A), and the RNA recognition sequence comprises a U1A binding site, wherein the U1A binds to the U1A binding site.

21. The vector of claim 20, wherein the RNA-binding domain comprises 1 to 6 UlAs, and the RNA recognition sequence comprises 1 to 30 U1 A binding sites.

22. The vector of any one of claims 1-21, further comprising a polyadenylation site.

23. The vector of claim 22, further comprising a pseudoknot RNA motif sequence, wherein the pseudoknot RNA motif sequence is positioned between the barcode and the polyadenylation site.

24. The vector of claim 23, wherein the pseudoknot RNA motif sequence is an evopreql pseudoknot RNA motif sequence.

25. The vector of any one of claims 1-24, wherein the second nucleotide sequence further encodes a pair of ribozymes comprising a first ribozyme and a second ribozyme, wherein the first ribozyme and the second ribozyme flank the RNA recognition sequence and the barcode in the barcoded RNA transcript, wherein the first ribozyme and the second ribozyme undergo autocatalytic cleavage to generate a 5 ’-end comprising a hydroxyl group and a 3 ’-end comprising a 2’,3’-cyclic phosphate group, wherein the 5’-end and the 3’-end are ligated by an endogenous RtcB RNA ligase in a cell resulting in circularization of the barcoded RNA transcript.

26. The vector of claim 25, wherein the pair of ribozymes are Twister ribozymes.

27. The vector of any one of claims 1-26, wherein the first promoter and the second promoter are constitutive or inducible.

28. The vector of any one of claims 1-27, wherein the first promoter is a hybrid cytomegalovirus (CMV) / chicken 0-actin promoter (CAG) promoter.

29. The vector of any one of claims 1-28, wherein the second promoter is a U6 promoter.

30. The vector of any one of claims 1-29, wherein the fusion protein further comprises a selectable marker.

31. The vector of claim 30, wherein the selectable marker is a fluorescent protein.

32. The vector of claim 31, wherein the fluorescent protein is a green fluorescent protein, a red fluorescent protein, a blue fluorescent protein, a yellow fluorescent protein, an orange fluorescent protein, or a violet fluorescent protein.

33. The vector of any one of claims 1 -32, further comprising a pair of restriction sites comprising a first restriction site and a second restriction site, wherein the first restriction site and the second restriction site flank the second expression cassette.

34. The vector of any one of claims 1-33, further comprising a multiple cloning site (MCS) for insertion of an expressible sequence into the vector.

35. The vector of any one of claims 1-34, further comprising an expressible sequence, wherein the barcode identifies the expressible sequence.

36. The vector of claim 35, wherein the expressible sequence encodes a protein or an RNA.

37. The vector of claim 36, wherein the protein is a therapeutic protein or a genomeediting enzyme.

38. The vector of claim 36 or 37, wherein the protein is a variant.

39. The method of claim 37, wherein the genome-editing enzyme is a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) nuclease, a meganuclease, a zine-finger nuclease (ZFN), or a transcription activator-like effector nuclease (TALEN).

40. The method of claim 39, wherein the Cas nuclease is Cas9 or Casl2a.

41. The vector of claim 36, wherein the RNA is a messenger RNA (mRNA) or a noncoding RNA.

42. The vector of claim 41, wherein the non-coding RNA is a microRNA (miRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a small nuclear RNA (snRNA),a piwi-interacting RNA (piRNA), a small nucleolar RNA (snoRNA), or a long non-coding RNA (IncRNA).

43. The vector of claim 36, wherein the RNA is a guide RNA or an aptamer.

44. The vector of any one of claims 1-43, wherein the second nucleotide sequence further comprises a sequence that is complementary to an oligonucleotide capture probe.

45. A vector comprising an expression cassette comprising a promoter operably linked to a nucleotide sequence encoding a nuclear retention element connected to a barcode, wherein transcription of the nucleotide sequence in a cell generates a barcoded RNA transcript comprising the nuclear retention element connected to the barcode, wherein the barcoded RNA transcript is translocated to the nucleus of the cell.

46. The vector of claim 45, wherein the nuclear retention element is a nuclear retention element of a long non-coding RNA.

47. The vector of claim 46, wherein the long non-coding RNA is MEG3, XIST, MALAT1, or SIRLOIN.

48. The vector of claim 46 or 47, wherein the nuclear retention element comprises the entire long non-coding RNA or a fragment thereof, wherein the barcoded RNA transcript is translocated to the nucleus of the cell.

49. The vector of claim 46, wherein the nuclear retention element comprises the nucleotide sequence of RCCTCCC, wherein R is an A or a G.

50. The vector of any one of claims 45-48, wherein the vector is a viral vector.

51. The vector of claim 50, wherein the viral vector is an adenovirus associated virus (AAV) vector, a lentivirus vector, or a G-deleted rabies virus (RVdG) vector.

52. The vector of claim 51, wherein the AAV vector further comprises a 5’-invcrtcd terminal repeat (ITR) and a 3’-ITR, wherein the expression cassette is positioned between the 5’- ITR and the 3 ’-ITR.

53. The vector of any one of claims 45-52, wherein the barcode is located in a 3’ untranslated region (3’ UTR) of the barcoded RNA transcript.

54. The vector of any one of claims 45-52, wherein the barcode is transcribed by a separate U6 promoter and DNA polymerase III.

55. The vector of any one of claims 45-54, further comprising a polyadenylation site.

56. The vector of claim 55, further comprising a pseudoknot RNA motif sequence, wherein the pseudoknot RNA motif sequence is positioned between the barcode and the polyadenylation site.

57. The vector of claim 56, wherein the pseudoknot RNA motif sequence is an evopreql pseudoknot RNA motif sequence.

58. The vector of any one of claims 45-57, wherein the nucleotide sequence encoding the nuclear retention element connected to the barcode further encodes a pair of ribozymes comprising a first ribozyme and a second ribozyme, wherein the first ribozyme and the second ribozyme flank the nuclear retention element and the barcode in the barcoded RNA transcript, wherein the first ribozyme and the second ribozyme undergo autocatalytic cleavage to generate a 5’-end comprising a hydroxyl group and a 3’-end comprising a 2’,3’-cyclic phosphate group, wherein the 5 ’-end and the 3 ’-end are ligated by an endogenous RtcB RNA ligase in a cell resulting in circularization of the barcoded RNA transcript.

59. The vector of claim 58, wherein the pair of ribozymes are Twister ribozymes.

60. The vector of any one of claims 45-59, wherein the promoter is constitutive or inducible.

61. The vector of any one of claims 45-60, wherein the promoter is a U6 promoter.

62. The vector of any one of claims 45-61, wherein the nucleotide sequence further encodes a selectable marker.

63. The vector of claim 62, wherein the selectable marker is a fluorescent protein.

64. The vector of claim 63, wherein the fluorescent protein is a green fluorescent protein, a red fluorescent protein, a blue fluorescent protein, a yellow fluorescent protein, an orange fluorescent protein, or a violet fluorescent protein.

65. The vector of any one of claims 45-64, further comprising a pair of restriction sites comprising a first restriction site and a second restriction site, wherein the first restriction site and the second restriction site flank the expression cassette.

66. The vector of any one of claims 45-65, further comprising a multiple cloning site (MCS) for insertion of an expressible sequence into the vector.

67. The vector of any one of claims 45-66, further comprising an expressible sequence, wherein the barcode identifies the expressible sequence.

68. The vector of claim 67, wherein the expressible sequence encodes a protein or an RNA.

69. The vector of claim 68, wherein the protein is a therapeutic protein or a genomeediting enzyme.

70. The vector of claim 68 or 69, wherein the protein is a variant.

71. The method of claim 69, wherein the genome-editing enzyme is a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) nuclease, a meganuclease, a zine-finger nuclease (ZFN), or a transcription activator-like effector nuclease (TALEN).

72. The method of claim 71, wherein the Cas nuclease is Cas9 or Casl2a.

73. The vector of claim 68, wherein the RNA is a messenger RNA (mRNA) or a noncoding RNA.

74. The vector of claim 73, wherein the non-coding RNA is a microRNA (miRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a small nuclear RNA (snRNA), a piwi-interacting RNA (piRNA), a small nucleolar RNA (snoRNA), or a long non-coding RNA (IncRNA).

75. The vector of claim 68, wherein the RNA is a guide RNA or an aptamer.

76. The vector of any one of claims 45-75, wherein the nucleotide sequence further comprises a sequence that is complementary to an oligonucleotide capture probe.

77. A cell transfected with the vector of any one of claims 1-76.

78. The cell of claim 77, wherein the cell is a mammalian cell.

79. The cell of claim 78, wherein the mammalian cell is a human cell.

80. The cell of claim 78 or 79, wherein the cell is a neuron.

81. A method of barcoding nuclei in a population of cells, the method comprising:providing a plurality of vectors according to any one of claims 1-76, wherein the barcode in each vector of the plurality comprises a unique identifier sequence; and transfecting the population of cells with the plurality of vectors to generate cells having barcoded nuclei.

82. The method of claim 81, wherein said transfecting is performed in vitro, ex vivo, or in vivo.

83. The method of claim 81 or 82, wherein the population of cells is in a tissue or an organ.

84. A method of multiplex screening of a population of cells, the method comprising: providing a plurality of vectors according to any one of claims 1-76, wherein the barcode in each vector of the plurality comprises a unique identifier sequence; transfecting the population of cells with the plurality of vectors to generate cells having barcoded nuclei; measuring morphological or functional characteristics of the cells having the barcoded nuclei; isolating the barcoded nuclei; and identifying the barcode in each barcoded nucleus.

85. The method of claim 84, wherein the population of cells is in a tissue or an organ.

86. The method of claim 85, wherein the tissue is nervous tissue.

87. The method of claim 84, wherein the population of cells is in a culture.

88. The method of any one of claims 84-87, wherein the population of cells comprises neurons or glial cells or a combination thereof.

89. The method of claim 88, further comprising optogenetically modifying the neurons.

90. The method of any one of claims 84-89, wherein said measuring the morphological or functional characteristics comprises performing gene expression profiling, microscopy, calcium imaging, an electrophysiology measurement, functional neuroimaging, a migration assay, an axonal growth and pathfinding assay, monosynaptic tracing, a phagocytosis assay, an enzymatic assay, a cell receptor assay, an ion channel assay, a signal transduction assay, or a cell secretion assay.

91. The method of claim 90, wherein the gene expression profiling comprises performing microarray analysis, RNA sequencing, or quantitative polymerase chain reaction.

92. The method of claim 91, wherein the RNA sequencing comprises single nucleus RNA sequencing, single cell RNA sequencing, or a combination thereof.

93. The method of claim 90, wherein the microscopy is confocal microscopy, atomic force microscopy, super-resolution microscopy, light-sheet microscopy, two-photon microscopy, or fluorescence microscopy.

94. The method of claim 90 wherein the electrophysiology measurement is patch clamping, electroencephalography (EEG), or magnetoencephalography (MEG).

95. The method of claim 90, wherein the functional neuroimaging is functional magnetic resonance imaging (fMRI), positron emission tomography (PET), functional nearinfrared spectroscopy (fNIRS), single-photon emission computed tomography (SPECT), or functional ultrasound imaging (fUS).

96. The method of any one of claims 84-95, wherein the morphological or functional characteristics are measured in the population of cells in a tissue of a live subject in vivo or in culture in vitro.

97. The method of claim 96, wherein the live subject is a nonhuman animal.

98. The method of claim 96 or 97, further comprising removing the tissue from the subject after said measuring the morphological or functional characteristics.

99. The method of any one of claims 84-98, wherein said isolating the barcoded nuclei comprises performing fluorescence activated nuclei sorting (FANS).

100. The method of claim 99, further comprising isolating the barcoded RNA transcript from each barcoded nucleus.

101. The method of claim 100, wherein said isolating the barcoded RNA transcript comprises hybridization of the barcoded RNA transcript to a capture probe comprising a sequence complementary to a region of the barcoded RNA transcript.

102. The method of claim 100, wherein said isolating the barcoded RNA transcript comprises hybridization of the barcoded RNA transcript to a capture probe comprising a poly-T sequence complementary to the poly-A tail of the barcoded RNA transcript.

103. The method of claim 101 or 102, wherein the capture probe is immobilized on a solid support.

104. The method of claim 103, wherein the solid support is a magnetic bead or polystyrene bead.

105. The method of any one of claims 84-104, wherein said identifying comprises sequencing the barcode sequence.

106. The method of claim 105, wherein said sequencing comprises performing single nucleus RNA sequencing (snRNA-seq).

107. The method of any one of claims 84-106, further comprising amplifying the barcode sequence prior to said sequencing.

108. The method of claim 107, wherein said amplifying comprises performing reverse transcription polymerase chain reaction or isothermal amplification.

109. The method of any one of claims 84-108, further comprising reverse transcribing the barcoded RNA transcript to generate a complementary DNA (cDNA) copy of the barcoded RNA transcript.

110. The method of claim 109, further comprising adding adapters to the 5’ end and the 3’ end of the cDNA copy of the barcoded RNA transcript.

111. The method of any one of claims 84- 110, where each vector of the plurality further comprises a different expressible sequence, wherein the barcode in each vector identifies the expressible sequence.

112. The method of claim 111, wherein the expressible sequence encodes a protein or an RNA.

113. The method of claim 112, wherein the protein is a therapeutic protein or a genome-editing enzyme.

114. The method of claim 123, wherein the therapeutic protein is a hormone, a cytokine, a chemokine, a growth factor, an enzyme, or an antibody.

115. The method of any one of claims 112-114, wherein the protein or RNA is a variant.

116. The method of claim 115, wherein the genome-editing enzyme is a clustered regularly interspaced short palindromic repeats (CRISPR)-associatcd (Cas) nuclease, a meganuclease, a zinc-finger nuclease (ZFN), or a transcription activator-like effector nuclease (TALEN).

117. The method of claim 116, wherein the Cas nuclease is Cas9 or Cas 12a.

118. The method of claim 112, wherein the RNA is a messenger RNA (mRNA) or a non-coding RNA.

119. The method of claim 118, wherein the non-coding RNA is a microRNA (miRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a small nuclear RNA (snRNA), a piwi-interacting RNA (piRNA), a small nucleolar RNA (snoRNA), or a long non-coding RNA (IncRNA).

120. The method of claim 112, wherein the RNA is a guide RNA or an aptamer.

121. The method of claim 112, wherein the protein is a fluorescent protein or a bioluminescent protein.

122. The method of claim 121, further comprising imaging the fluorescent protein or the bioluminescent protein, wherein a location of a cell expressing the fluorescent protein or the bioluminescent protein is determined from the imaging.

123. The method of claim 122, further comprising mapping the location of the cell expressing the fluorescent protein or the bioluminescent protein onto a reference image of the tissue.

124. A vector system comprising the vector of any one of claims 1-76 and a helper vims vector.Il l1 5. The vector system of claim 124, wherein the helper virus vector encodes E 1 a andElb or E2a and E4.

126. The vector system of claim 124, wherein the helper virus vector encodes a glycoprotein (G) gene.

127. The vector system of any one of claims 124-126, further comprising a vector encoding one or more capsid proteins.

128. The vector system of claim 127, wherein the one or more capsid proteins comprise VP1, VP2, and VP3.

129. The vector system of any one of claims 124-128, further comprising a vector encoding an AAV Rep protein.

130. The vector system of claim 124, wherein the helper virus vector encodes Gag, Pol, Rev, and VSV-G.

131. A cell transfected with the vector system of any one of claims 124-130.

132. The cell of claim 131, wherein the cell is a mammalian cell.

133. The cell of claim 132, wherein the mammalian cell is a human cell.

134. The cell of any one of claims 131-133, wherein the cell is a neuron.

135. A vector library comprising a plurality of vectors according to any one of claims 1-76, wherein the plurality of vectors collectively comprise a plurality of expressible sequences, and wherein the barcode in each vector identifies the expressible sequence in the vector.

136. The vector library of claim 135, wherein the plurality of expressible sequences collectively encode different proteins, protein variants, antibodies, mRNAs, shRNAs, siRNAs, aptamers, or guide RNAs for RNA-guided nucleases.