High-content and high-resolution in vivo screen for analyzing gene functions

By employing AAV vectors and CRISPR-Cas9 technology to introduce genetic perturbations into specific cell types in vivo, followed by single-cell RNA sequencing, the challenges of high-content phenotypic screens are addressed, achieving efficient and precise analysis of gene functions in vivo.

WO2025054224A9PCT designated stage expired Publication Date: 2025-06-05THE SCRIPPS RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/045238
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-08
Filing Date
2024-09-05
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current methods for high-content phenotypic screens in vivo are challenging due to difficulties in scalably labeling, perturbing, and isolating cells from various cell types, and in robustly capturing and deconvoluting each cell's perturbation identity in sparse single-cell omics data.

Method used

The use of AAV vectors encoding genetic perturbations, combined with a CRISPR-Cas9 expressing transgenic system, to introduce genetic perturbations into specific cell types in vivo, followed by single-cell RNA sequencing to determine the perturbed gene and correlate it with the observed phenotype.

Benefits of technology

This approach enables high-content and high-resolution analysis of gene functions in vivo, overcoming previous limitations by achieving rapid and robust transgene expression, efficient gRNA capture, and scalable genetic perturbation screens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024045238_05062025_PF_FP_ABST
    Figure US2024045238_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides high-content and high-resolution in vivo screens for analyzing functions of a plurality of genes. In the functional screens of the invention, genetic perturbations are delivered to a CRISPR-expressing transgenic system by specific AAV vectors that are fully compatible with Perturb-seq platforms. The screens enable functional genomics analyses across diverse tissues, cell types, and model organisms in vivo, with high-throughput single-cell readout.
Need to check novelty before this filing date? Find Prior Art

Description

PATENT Attorney Docket No.: 2219.1PC High-Content And High-Resolution In Vivo Screen For Analyzing Gene Functions CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The subject patent application claims the benefit of priority to U.S. Provisional Patent Application No.63 / 581,297 (filed September 8, 2023; now pending). The full disclosure of the priority application is incorporated herein by reference in its entirety and for all purposes. STATEMENT CONCERNING GOVERNMENT SUPPORT

[0002] This invention was made with government support under HG012819 awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND OF THE INVENTION

[0003] Human genetics studies have identified risk genes confidently associated with disease and disorder. In particular, major international initiatives in identifying genetic contribution of neurodevelopmental disorders have come to fruition, including mapping variants implicated in autism spectrum disorder and neurodevelopmental delay. Collectively, hundreds of genes have been identified to contribute to neurodevelopmental disorder risks, with additional genes and loci involved in schizophrenia and other psychiatric conditions. Even still, little is known about the underlying mechanisms leading to the disease manifestation. One of the key aims in functional genomics is to define the cell type and molecular specificity of gene actions: among all the heterogenous cell types, including the central and peripheral nervous systems, what are the most vulnerable cell types and molecular networks that are affected by each risk gene? Genetic technologies, including programmable perturbations enabled by CRISPR (clustered regularly interspaced short palindromic repeats), have opened new avenues to experimentally test these risk genes at scale. In vivo, pooled genetic screens in mammals have been applied to several systems and contexts to readout gRNA abundance as a proxy for cellular proliferation or depletion. However, changes of gRNA abundance cannot fully capture all the molecular consequences and cellular phenotypes.

[0004] To understand gene function with both depth and precision, scalable technologies like CRISPR-Cas9 gene editing and single-cell genomics were combined to give rise toPerturb-seq and CROP-seq, using transcriptomic or epigenomic readouts. The prevailing applications have been within the in vitro settings in immortalized cell lines. These in vitro cultures, though controlled, often do not mirror the complex intercellular signaling and dynamic physiology in vivo where disorders take root and manifest. In vivo Perturb-seq adapted this system to developing embryos and characterized disorder risk genes and transcription factors in carcinogenesis, across the developmental period when the disorder manifests. See, e.g., Elena et al., bioRxiv 2023.01.30.525356; and Jin et al., Science 2020; 370, eaaz6063. Gene expression analysis revealed convergent molecular networks and specific neuronal or glial subtypes impacted by risk genes within the context of a developing brain. Nevertheless, conducting high-content phenotypic screens in vivo remains challenging due to at least two hurdles: (1) the need to scalably label, perturb, and isolate enough cells from various cell types in vivo, which is harder than most in vitro counterparts, and (2) by multiplexing the screen and mixing perturbation agents (gRNAs), we face the challenge to then robustly capture and deconvolute each cell’s perturbation identity in the sparse single- cell omics data. To compound these challenges, the prevalent use of lentiviral vectors in most Perturb-seq applications, known for their limited in vivo penetration, transduction, and thermostability, hampers systemic screens in hard-to-reach tissues such as the adult central or peripheral nervous systems. See, e.g., Higashikawa and Chang, Virology 2001; 280: 124-131.

[0005] Thus, there is an unmet need for high-content phenotypic screens for functional genomics studies in vivo. The present invention addresses this and other needs in the art. SUMMARY OF THE INVENTION

[0006] In one aspect, the invention provides methods for analyzing functions of a plurality of target genes in one or more specific cell types in vivo. The methods entail (1) introducing a library of AAV vectors encoding genetic perturbations for the plurality of target genes into a CRISPR-Cas9 expressing transgenic system, (2) identifying one or more cells of a specific cell type from the transgenic system that express a genetic perturbation and display a specific phenotype, and (3) determining (a) the genetic perturbation encoded by the AAV vector introduced into each of the one or more cells and (b) the corresponding perturbed gene, thereby correlating the corresponding perturbed gene with the phenotype displayed by each of the one or more cells. In some embodiments, the employed transgenic system is a developing embryo of a Cas9-expressing transgenic animal, and the library of AAV vectors are vectors of serotype AAV-SCH9 or AAV2.NN. In some of these embodiments, theemployed library of AAV vectors are AAV-pRep2-SCH9 vectors or AAV-pRep2-SCH9 (repeat 136bp) vectors. In various embodiments, the library of AAV vectors are administered to the animal embryo around a development stage that is equivalent to mouse embryonic day 11.5 (E11.5), 12.5 (E12.5), 13.5 (E13.5), 14.5 (E14.5), or 15.5 (E15.5). In some embodiments, the library of AAV vectors is administered in utero into the brain (lateral ventricles) of the embryo.

[0007] In some other embodiments, the employed transgenic system is a postnatal or adult Cas9-expressing transgenic animal, and the library of AAV vectors are vectors of serotype AAV-PHP.eB or AAV.PHP.S. In some of these embodiments, the library of AAV vectors is administered to the animal at an age that is equivalent to mouse age of from about postnatal day 10 (P10) to about 18 months old. In some embodiments, the library of AAV vectors are administered retroorbitally to the animal.

[0008] In some methods of the invention, the specific cell type to be analyzed is newborn neuron or progenitor cell thereof obtained from a brain tissue. In some of these methods, the brain tissue is from neocortex, olfactory bulb, striatum, hippocampus, thalamus, or cerebellum. In some methods, the target genes to be analyzed are known or suspected to be associated with a neurodevelopmental disorder.

[0009] In some embodiments, each of the genetic perturbations to be introduced into the specific cell types in vivo contains one or more guide RNAs (gRNAs) for introducing genomic changes via CRISPR-Cas9 gene editing in the coding region of a target gene. In some of these embodiments, the one or more gRNAs each target a sequence at the 5’ end of a target gene’s coding region. In some methods, the AAV vectors further encode a transposable element that is flanked with the gRNAs, and the library of AAV vectors are introduced into the transgenic system in the presence of a transposase that recognizes the transposable element and inserts it into the host genome. In some of these methods, the transposase is encoded by a second AAV vector co-introduced with the library of AAV vectors into the transgenic system. In some embodiments, the employed transposase is Transposase hypPB.

[0010] Some methods of the invention can additionally entail enriching the identified one or more cells that express one or more genetic perturbations and display a specific phenotype. In some methods, the one or more gRNAs’ coding sequence is operatively linked to a reporter gene in the AAV vectors. In some of these methods, the one or more cells expressing a genetic perturbation is identified by detecting expression of the reporter gene. In some methods, the genetic perturbation encoded by the AAV vector introduced into each of the oneor more cells is determined via capturing the one or more gRNAs via single-cell RNA sequencing (scRNA-seq). In some methods, the one or more gRNAs are functional for gene editing and can be directly captured in scRNA-seq.

[0011] Some methods of the invention utilize AAV-SCH9 or AAV2.NN vectors. In some of these methods, the library of AAV vectors is administered to the animal embryo around a development stage that is equivalent to mouse embryonic day 11.5 (E11.5), 12.5 (E12.5), 13.5 (E13.5), 14.5 (E14.5), or 15.5 (E15.5). Some methods of the invention utilize AAV- PHP.eB or AAV.PHP.S vectors. In some of these methods, the library of AAV vectors are administered to the animal retroorbitally in an age from postnatal day 10 (P10) to 18 months old.

[0012] In a related aspect, the invention provides methods for analyzing functions of a plurality of target genes in one or more specific cell types in vivo. These methods involve: (1) co-introducing into a developing embryo of a Cas9-expressing transgenic animal (i) an AAV vector encoding a transposase, and (ii) a library of AAV vectors each encoding (a) one or more guide RNAs (gRNAs) for one of the target genes and (b) a transposable element that is recognized by the transposase and that is flanked with the gRNAs; wherein the AAV vectors are of serotype AAV-SCH9 or AAV2.NN; (2) identifying one or more cells of a specific cell type from the embryo that express the gRNAs and display a specific phenotype, and (3) determining the gRNAs encoded by each of the AAV vectors introduced into the one or more cells and the corresponding perturbed gene; thereby correlating the specific phenotype with the function of each of the perturbed target genes. In some of these methods, the employed transposase is hypPB (Hyperactive piggybac).

[0013] A further understanding of the nature and advantages of the present invention may be realized by reference to the remaining portions of the specification and claims. DESCRIPTION OF THE DRAWINGS

[0014] Figure 1. Barcoded AAV serotype screen in vivo identified AAV-SCH9 efficiently targeting developing brains. (A) Schematics of AAV library administration in utero at E13.5 followed by Fluorescence-activated cell sorting (FACS) cell enrichment and barcode analysis with next generation sequencing. (B-C) Immunofluorescence analysis of brain sections two days after AAV library administration: (B) co-stained with markers of newborn projection neurons (TBR1 and CTIP2) and intermediate progenitors (TBR2), indorsal cortical laminar and ganglionic eminence (GE); (C) quantification of percentage of GFP+cells co-expressing neuronal and progenitor markers including TBR1, TBR2 and CTIP2. (D) Heatmap of proportion of 86 AAV serotype abundance in AAV library, 48 hours post transduction in HT22 cells and in embryonic mouse brain; each row represents the abundance of an AAV serotype. (E) Principal component analysis of the percentage abundance of AAV library, HT22 cells and mouse brain 24- or 48-hours post transduction. (F) AAV-SCH9 percentage abundance in the initial AAV library (input for transductions), HT22 cells and mouse brain 24- or 48-hours post transduction. Error bars indicated standard error of the mean. (G) Volcano plots of AAV serotype changes in mouse brain (left) or HT22 cells (right) 48 hours post transduction compared to the initial AAV library. Scale bars indicate 250μm (left in B), 50μm (right in B and in C).

[0015] Figure 2. AAV-SCH9 labeled newborn neurons spreading across brain regions in vivo across developmental stages. (A) Schematics of a secondary AAV serotype screen: 14 AAV serotypes were barcoded and introduced in pool in utero at E13.5, followed by scRNA- seq 48 hours later. (B) Uniform Manifold Approximation and Projection (UMAP) visualization of 11 major cell populations identified (left) from sorted (GFP+) and unsorted cells (right); cell types include: upper and deep layer projection neurons (ULPN, DLPN), migrating neurons (Mig. neurons), apical progenitors (Api. prog), intermediate progenitors (IP), interneurons derived from the medial ganglionic eminence (IN-MGE), interneurons derived from the non-medial ganglionic eminence (IN-non-MGE), Cajal-Retzius cells (CR), fibroblast (Fibro), mural cells (Mural), and microglia (Mg). (C) AAV serotype barcode expression in each cell type. Each row represents an AAV barcode, and each AAV serotype is associated with a distinct set of 3 barcodes (BCa, BCb and BCc). Each column represents barcode expression in a cell, arranged by cell types. (D) Immunofluorescence analysis of AAV-SCH9-GFP-KASH E14.5 transduced brain sections co-stained with markers including CTIP2, TBR1, and TBR2 as well as the percentage of marker co-expression in the GFP+cells. Boxes on top panels indicate chosen fields of view in the bottom panels; arrows indicate representative cells with marker co-localizations. (E) AAV-SCH9-GFP-KASH was administered at two different time points (E13.5 or E17.5) and the cortical sections were analyzed 48 hours later to quantify GFP+cells distribution across laminar layers, which were divided evenly into six bins. (F) AAV-SCH9-GFP-KASH or lentiviral reporter (GFP) administered at E13.5 resulted in diverse brain region labeling at P7. Scale bars indicate 500μm (left in D and top in F) or 50μm (right in D, E, and bottom in F).

[0016] Figure 3. hypPB transposon enhanced and stabilized expression in embryonic and adult brains and peripheral nervous systems. (A) Schematics of the molecular design to enhance transgene expression and prevent loss due to cell division and differentiation. Shaded box indicates transposon flanked by the inverted repeats (IR). (B) Timelapse imaging showed co-transfection of hypPB increased the expression of the GFP transgene in vitro. Each dot represents analysis from a chosen field of view from a well. (C-E) Transposon stabilized expression in vivo across embryonic brain, adult brain, and adult dorsal root ganglion using three targeting vectors: AAV-SCH9, AAV9-PHP.eB and AAV9-PHP.S. (F) AAV-SCH9 labeled more cortical neurons in the presence of hypPB; Y axis indicated the number of GFP-expressing cells per mm2(n=3 animals / condition). Asterisks indicate P- value<0.0001 with Welch’s t-test. (G) AAV9-PHP.eB labeled cortical neurons with increased expression intensity with hypPB: average fluorescence intensity (X axis) across cortical layers (Y axis), from ventricular zone (0) to pia (1) (n=3 animals / condition). (H) Whole genome sequencing of AAV-SCH9 transduced cells in vivo showed the genomic regions of integration events. Left: each line indicates a unique hybrid read between mouse genome and transposon, illustrated as a wheel plot. Right: percentage of integration events occurred in intergenic and coding regions in the genome. (I) Example reads that captured the junction of mouse genome and transposon, indicating the integration sites. Top: schematic illustration of the transposon sequence (SEQ ID NOs:27 and 37). Bottom: example reads aligned to the transposon sequence (in orange boxes) and to the mouse genome with their chromosome numbers (Left portion: SEQ ID NOs:28-36 respectively; Right portion: SEQ ID NOs:38-46 respectively). Scale bars indicate 100μm (in B, right in C, right in D and in E) or 500μm (left in C and left in D).

[0017] Figure 4. In vivo Perturb-seq identified cell type-specific changes across perturbations of transcription factors. (A) Schematics of the screen design. Shaded box indicates transposon flanked by the inverted repeats (IR); NT and ST indicate non-targeting and safe-targeting controls, respectively. (B) UMAP plot of filtered cells with a single perturbation, with each cell colored by annotated cell type. (C) UMAP plot of filtered cells, with each cell colored by gRNA identity estimated by DemuxEM with down-sampling (top) and batch / channel (bottom). (D) Top: schematics of 5’ and 3’ scRNA-seq capture mechanism. Bottom: comparing gRNA capture rate by measuring percentage of cells in each cell type with gRNA UMI number greater than 5 in 5’ and 3’ scRNA-seq, with cell types on the X axis, percentage of cells assigned gRNA identity on the Y axis. (E) Bar plot showingpercentage of cells assigned to one or more gRNA across major cell types, color represents gRNA identity assignment. (F) Percentage of insertion / deletion reads in Foxg1 gRNAs targeting loci by Foxg1 perturbation comparing to Non-Targeting control 2 (NT2) controls, extracted from scRNA-seq data. (G) Heatmap showing cell type proportion changes by each perturbation. The rings highlight FDR adjusted P-value<0.05. (H) Dot plot of number of differentially expressed genes (DEG) and cell number (size of dots) across cell type- perturbation combinations. The boxes highlight the robust changes of ≥10 DEGs supported by at least 2 gRNAs. (I) Volcano plots of cell-type specific effects on differential expressed genes of Foxg1 perturbation in Layer 6 CT and Layer 5 IT excitatory neurons. The dots label significantly altered DEGs; the rings label the highlighted DEGs. DETAILED DESCRIPTION OF THE INVENTION I. Overview

[0018] Thousands of disease risk genes identified through human genetics far outstrip our current capacity to systematically study their functions in physiological contexts. Two substantial hurdles impede advancement of in vivo genetic screens: a lack of tools to access large number of cells within live tissues, and the capability to introduce robust enough expressions to be detected in pooled assays such as single-cell omics. AAV presents a promising strategy for delivering genetic perturbations to a wide range of cell types in vivo. However, the conventional AAV expression tends to be transient, at a relatively low level, with subsequently dilutions through cell divisions, which together poses a challenge to accurately recover the perturbation identity of sparsely labeled cells in pooled assays like Perturb-seq. See, e.g., Lang et al., 2019; Nat. Commun.10, 3415. Most critically, the slow onset of AAV-mediated expression – often taking days or even weeks – is suboptimal for studying dynamic gene functions in fast-evolving cellular contexts such as neurodevelopment. For example, peak AAV-transgene expression commonly occurs after more than seven days, a timespan that encompasses the entire window for corticogenesis. These limitations underscore the pressing need for an AAV system that delivers rapid and robust transgene expression, to allow the expanded capabilities of massively parallel Perturb- seq in vivo.

[0019] The present invention provides AAV-based high-content in vivo screen platforms that overcome these limitations. The generalizable platforms for in vivo Perturb-seq are able to target diverse tissue types and cell types with high content phenotypic readout with singlecell resolution. The invention is derived in part from studies undertaken by the inventors to develop an AAV-based platform for massively parallel in vivo Perturb-seq to target a broad spectrum of tissues and cell types with gene expression-based characterization at single cell resolution. As detailed herein, through a barcoded screen of 86 phylogenetically diverse AAVs in vivo, the inventors identified specific AAV serotypes including AAV-SCH9, which enable swift and robust transgene delivery in newborn neurons and progenitors within 48 hours post transduction (versus 2-3 weeks by the currently available methods). The identified AAV vectors were further synergized with a transposon system to ensure rapid and sustained expression of gRNAs in both target cells and their daughter cells, facilitating efficient gRNA captures in the single-cell gene expression-based analysis, with >82% cells recovered with perturbation identities. Through in utero perturbation screens, the inventors additionally uncovered the cell-type specific impact of perturbations on neuronal subtype proportions. Differential expression analysis further elucidated cell-type differential transcriptomic changes: Foxg1 predominantly affects Layer 5 intratelencephalic and Layer 6 corticothalamic neurons, and its loss-of-function, strongly associated with neurodevelopmental delay, de- represses transcriptional networks and leads to hybrid cell fates and misexpression of axon guidance and synaptogenesis molecules.

[0020] Previously, in vivo Perturb-seq was limited by the number of cells that can be labeled and harvested from tissues, and the ability to capture gRNA reliably. Importantly, the AAV-based in vivo Perturb-seq platforms as exemplified herein achieved 10-fold higher labeling in embryonic brains (two batches for 50,075 cells) and can label >6% of brain cells, surpassing the lentiviral efficacy of <0.1% as reported in Jin et al. (Science 2020; 370, eaaz6063), permitting in vivo analysis of over 30,000 cells in a single trial. In contrast, conventional methods employing lentiviral vectors typically can only analyze a few thousand cells in vivo from one batch, e.g., 17 batches 46,770 cells as described in Jin et al., 2020. Thus, the current invention represents significant improvement and technological advantages over other known in vivo Perturb-seq screens that rely on other viral vectors, e.g., the lentiviral vector based screen as described in Jin et al., 2020; and US Patent Application Publication NO.2021 / 0172017A1. Compatible with various perturbation techniques (CRISPRa / i) and phenotypic measurements (single-cell or spatial multiomics), the Perturb- seq platforms described herein provides a flexible approach to interrogate gene function across diverse cell types in vivo, translating gene variants to their causal functions.

[0021] It is noted that, unless otherwise specified, this invention is not limited to the particular methodology, protocols, and reagents described as these may vary. Unless otherwise indicated, the practice of the present invention employs conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry and immunology, which are within the skill of the art. Such techniques are explained fully in the literature. For example, exemplary methods are described in the following references, Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press (3rded., 2001); Brent et al., Current Protocols in Molecular Biology, John Wiley & Sons, Inc. (ringbou ed., 2003); Freshney, Culture of Animal Cells: A Manual of Basic Technique, Wiley-Liss, Inc. (4thed., 2000); and Weissbach & Weissbach, Methods for Plant Molecular Biology, Academic Press, NY, Section VIII, pp.421-463, 1988. In addition, the following sections provide more detailed guidance for practicing the invention. II. Definitions

[0022] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. The following references provide one of skill with a general definition of many of the terms used in this invention: Academic Press Dictionary of Science and Technology, Morris (Ed.), Academic Press (1sted., 1992); Oxford Dictionary of Biochemistry and Molecular Biology, Smith et al. (Eds.), Oxford University Press (revised ed., 2000); Encyclopaedic Dictionary of Chemistry, Kumar (Ed.), Anmol Publications Pvt. Ltd. (2002); Dictionary of Microbiology and Molecular Biology, Singleton et al. (Eds.), John Wiley & Sons (3rded., 2002); Dictionary of Chemistry, Hunt (Ed.), Routledge (1sted., 1999); Dictionary of Pharmaceutical Medicine, Nahler (Ed.), Springer-Verlag Telos (1994); Dictionary of Organic Chemistry, Kumar and Anandand (Eds.), Anmol Publications Pvt. Ltd. (2002); and A Dictionary of Biology (Oxford Paperback Reference), Martin and Hine (Eds.), Oxford University Press (4thed., 2000). Further clarifications of some of these terms as they apply specifically to this invention are provided herein.

[0023] As used herein, the singular forms "a", "an", and "the" include plural reference unless the context clearly dictates otherwise. Thus, for example, reference to "a cell" includes a plurality of such cells, reference to "a protein" includes one or more proteins and equivalents thereof known to those skilled in the art, and so forth.

[0024] As used herein, AAV refers to adeno-associated virus, and may be used to encompass the naturally occurring wild-type virus itself or derivatives thereof. An AAV vector is a small single-stranded DNA viral vector containing icosahedral protein capsids. The term covers all subtypes, serotypes and pseudotypes, and both naturally occurring and recombinant forms, except where required otherwise. Pseudotyped AAV refers to an AAV that contains capsid proteins from one serotype and a viral genome including 5'-3' ITRs of a second serotype. The abbreviation "rAAV" refers to recombinant adeno-associated viral particle or a recombinant AAV vector (or "rAAV vector"). An "AAV virus" or "AAV viral particle" refers to a viral particle composed of at least one AAV capsid protein (preferably by all of the capsid proteins of a wild-type AAV) and an encapsidated polynucleotide. If the particle comprises a heterologous polynucleotide (i.e., a polynucleotide other than a wild-type AAV genome such as a transgene to be delivered to a mammalian cell), it is typically referred to as "rAAV".

[0025] Cas9 CRISPR genome editing system requires the Cas9 DNase obtained from Streptococcus pyogenes and a guide RNA (gRNA) for targeting the Cas9 DNase activity to complementary genomic sequences. For Cas9 CRISPR systems, the gRNA is made up of two parts: crispr RNA (crRNA), a 17-20 nucleotide sequence complementary to the target DNA, and a trans-activating crispr RNA (tracrRNA), which serves as a binding scaffold for the Cas nuclease. In addition to the sequence complementarity, the cleavage site in the target DNA is also preceded by a short protospacer-adjacent motif (PAM). In bacteria, Cas9 relies on RNase III to excise crRNAs from a CRISPR array.

[0026] Protospacer adjacent motif (PAM) is a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. PAM is a component of the invading virus or plasmid, but is not a component of the bacterial CRISPR locus. Cas9 will not successfully bind to or cleave the target DNA sequence if it is not followed by the PAM sequence. PAM is an essential targeting component (not found in bacterial genome) which distinguishes bacterial self from non-self DNA, thereby preventing the CRISPR locus from being targeted and destroyed by nuclease. The canonical PAM is the sequence 5'-NGG-3' where "N" is any nucleobase followed by two guanine ("G") nucleobases. Guide RNAs (gRNAs) can transport Cas9 to anywhere in the genome for gene editing, but no editing can occur at any site other than one at which Cas9 recognizes PAM.

[0027] The canonical PAM is associated with the Cas9 nuclease of Streptococcus pyogenes (designated SpCas9), whereas different PAMs are associated with the Cas9 proteins of the bacteria Neisseria meningitidis, Treponema denticola, and Streptococcus thermophilus. 5'-NGA-3' can be a highly efficient non-canonical PAM for human cells, but efficiency varies with genome location. Attempts have been made to engineer Cas9s to recognize different PAMs to improve ability of CRISPR-Cas9 to do gene editing at any desired genome location. Cas9 of Francisella novicida recognizes the canonical PAM sequence 5'-NGG-3', but has been engineered to recognize the PAM 5'-YG-3' (where "Y" is a pyrimidine), thus adding to the range of possible Cas9 targets.

[0028] Protospacers are spacer sequences in CRISPR loci in a bacterium that were inserted into a CRISPR locus by invading viral or plasmid DNA. On subsequent invasion, Cas9 nuclease attaches to tracrRNA:crRNA which guides Cas9 to the invading protospacer sequence. But Cas9 will not cleave the protospacer sequence unless there is an adjacent PAM sequence. The spacer in the bacterial CRISPR loci will not contain a PAM sequence, and will thus not be cut by the nuclease. But the protospacer in the invading virus or plasmid will contain the PAM sequence, and will thus be cleaved by the Cas9 nuclease. For editing genes, guide RNAs (gRNAs) are synthesized to perform the function of the tracrRNA:crRNA complex in recognizing gene sequences having a PAM sequence at the 3'-end.

[0029] A “host cell” or “target cell” refers to a living cell into which a heterologous polynucleotide sequence is to be or has been introduced. The living cell includes both a cultured cell and a cell within a living organism. Means for introducing the heterologous polynucleotide sequence into the cell are well known, e.g., transfection, electroporation, calcium phosphate precipitation, microinjection, transformation, viral infection, and / or the like. Often, the heterologous polynucleotide sequence to be introduced into the cell is a replicable expression vector or cloning vector. In some embodiments, host cells can be engineered to incorporate a desired gene on its chromosome or in its genome. Many host cells that can be employed in the practice of the present invention (e.g., CHO cells) serve as hosts are well known in the art. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press (3rded., 2001); and Brent et al., Current Protocols in Molecular Biology, John Wiley & Sons, Inc. (ringbou ed., 2003). In some preferred embodiments, the host cell is a mammalian cell.

[0030] The term “operably linked” or “operably associated” refers to functional linkage between genetic elements that are joined in a manner that enables them to carry out theirnormal functions. For example, a gene is operably linked to a promoter when its transcription is under the control of the promoter and the transcript produced is correctly translated into the protein normally encoded by the gene. Similarly, a gRNA-encoding sequence is operably linked to a reporter gene if the gRNA is co-produced with a transcript encoded by the reporter gene.

[0031] A “substantially identical” nucleic acid or amino acid sequence refers to a polynucleotide or amino acid sequence which comprises a sequence that has at least 75%, 80% or 90% sequence identity to a reference sequence as measured by one of the well-known programs described herein (e.g., BLAST) using standard parameters. The sequence identity is preferably at least 95%, more preferably at least 98%, and most preferably at least 99%. In some embodiments, the subject sequence is of about the same length as compared to the reference sequence, i.e., consisting of about the same number of contiguous amino acid residues (for polypeptide sequences) or nucleotide residues (for polynucleotide sequences). Polynucleotide sequences are no less substantially identical if they are composed of RNA or DNA, despite the chemical differences between RNA and DNA, and the presence of uracil in RNA instead of thymidine in DNA.

[0032] Sequence identity can be readily determined with various methods known in the art. For example, the BLASTN program (for nucleotide sequences) uses as defaults a wordlength (W) of 11, an expectation (E) of 10, M=5, N=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1989)). Percentage of sequence identity is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity.

[0033] As used herein, “complementary” or “complement” refers to a nucleotide or nucleotide sequence that hybridizes to a given nucleotide or nucleotide sequence. For instance, for DNA, the nucleotide A is complementary to T and vice versa, and the nucleotideC is complementary to G and vice versa. For instance, in RNA, the nucleotide A is complementary to the nucleotide U and vice versa, and the nucleotide C is complementary to the nucleotide G and vice versa.

[0034] “Paired” and “unpaired” refer to Watson-Crick base pairs, in either the context of DNA, RNA or a DNA-RNA hybrid. A sequence is said to be paired with its reverse complement.

[0035] A cell has been “transformed” or “transfected” by exogenous or heterologous polynucleotide (or “a transgene” or “a target gene” as used interchangeably herein) when such polynucleotide has been introduced inside the cell. The transforming polynucleotide may or may not be integrated (covalently linked) into the genome of the cell. In prokaryotes, yeast, and mammalian cells for example, the transforming polynucleotide may be maintained on an episomal element such as a plasmid. With respect to eukaryotic cells, a stably transformed cell is one in which the transforming polynucleotide has become integrated into a chromosome so that it is inherited by daughter cells through chromosome replication. This stability is demonstrated by the ability of the eukaryotic cell to establish cell lines or clones comprised of a population of daughter cells containing the transforming polynucleotide. A “clone” is a population of cells derived from a single cell or common ancestor by mitosis. A “cell line” is a clone of a primary cell that is capable of stable growth in vitro for many generations.

[0036] Transposons (TEs) (transposable elements or jumping genes) refer to DNA sequences that can move and integrate to different locations within the genome. There are two main classes of TEs, class I TEs (or retrotransposons) that move through an RNA intermediate via a “copy and paste” mechanism of transposition, and class II TEs that move through a DNA intermediate a “cut and paste” mechanism of transposition. DNA transposons can move in the DNA of an organism via a single-or double-stranded DNA intermediate. DNA transposons have been found in both prokaryotic and eukaryotic organisms. They can make up a significant portion of an organism's genome, particularly in eukaryotes. In prokaryotes, TE's can facilitate the horizontal transfer of antibiotic resistance or other genes associated with virulence. DNA transposons can be autonomous or non-autonomous. Autonomous transposons encode their own transposase enzyme that facilitates the jumping of the gene while non-autonomous transposons require the transposase activity of another transposable element. DNA transposons are delineated by flanking terminal repeats that markthe location that the transposase excises the DNA. These DNA elements then re-integrate at a different location within the genome.

[0037] A "vector" or “construct" is a non-naturally occurring nucleic acid with or without a carrier that can be introduced into a cell, or has been introduced into a cell. Vectors that have been introduced into a cell include transfected plasmids and integrated DNA molecules, including those resulting from retroviral integration, integration of an AAV vector, and integration by homologous recombination. Vectors capable of directing the expression of heterologous polynucleotide or transgene sequences encoding for one or more polypeptides are referred to as "expression vectors" or "expression constructs". The cloned transgene sequence or open reading frame (ORF) is usually placed under the control of (i.e., operably linked to) certain regulatory sequences such as promoters, enhancers and polynucleotide switch sequences. III. AAV vector based Perturb-seq platforms for analyzing gene functions in vivo

[0038] The present invention provides in vivo Perturb-seq platforms to analyze target genes in diverse tissue types and cell types with high content phenotypic readout and single cell resolution. The in vivo systems of the invention utilize a library AAV vectors of specific serotypes for delivering genetic perturbations to modulate one or more target genes of interest in specific cell or tissue types, e.g., central and peripheral nervous system, in an animal embryo or an adult or aged animal. Adeno-associated virus (AAV) is a small, nonenveloped virus that was adapted for use as a gene transfer vehicle. AAV vectors refer to recombinant adeno-associated viruses that are derived from nonpathogenic parvoviruses. They evoke essentially no cellular immune response, and produce transgene expression lasting months in most systems. Like adenovirus, adeno-associated virus vectors also have the capability to infect replicating and nonreplicating cells and are believed to be nonpathogenic to humans. Delivery of heterologous polynucleotide sequences via recombinant AAV can provide for safe, unobtrusive and sustained expression (> 2 years) of high levels of protein therapeutics.

[0039] As detailed in the Examples herein, the specific AAV vector serotypes suitable for the invention were identified through a barcoded screen of phylogenetically diverse AAVs in vivo. Serotypes of AAV vectors refer to classifications of AAV variants based on their molecular composition and their targeting specificity in tissues and cell types. To date, there are at least 11 number of AAV serotypes, including AAV1, 2, 4, 5, 6, 7, 8, 9 (see, e.g., Issa et al., Cells 2023; 12(5):785). General characteristics including structural information of well-known AAV serotypes are provided in, e.g., Srivastava, Curr Opin Virol.2016; 21: 75–80; Issa et al., Cells 2023; 12:785; Chen et al., Human Gene Therapy.2020 ; 31: 440-447; Pillay et al., J. Virol.2017; 91: e00391-17; Naso et al., BioDrugs.2017; 31: 317–334; Wang et al., Nat. Rev. Drug Discov.2019; 18: 358–378; and Wu et al., Mol. Ther.2006; 14: 316-327. As demonstrated herein, specific AAV serotypes and variants that are suitable for the practice of the present invention include AAV-SCH9, AAV2.NN, AAV9-PHP.eB and AAV9-PHP.S. AAV-SCH9 is an AAV serotype that is a hybrid of AAV2, 8, and 9 and originally made to transduce adult neural stem cells. Detailed structural information of AAV-SCH9 is provided in, e.g., Ojala et al., Mol. Ther.2018; 26: 304-319. Specific AAV-SCH9 and AAV2.NN vectors can be obtained from commercial vendors, e.g., Addgene (Watertown, MA). Similarly, other well-known AAV vectors including AAV9-PHP.eB and AAV9-PHP.S exemplified herein can also be obtained commercially, e.g., from vendors such as Addgene (Watertown, MA).

[0040] By utilizing the specific AAV serotypes, the in vivo Perturb-Seq platforms described herein provide a flexible and generalizable strategy to efficiently target selective cell types in vivo, and perform scalable genetic perturbation screens with reliable phenotypic readout using high-throughput single-cell omic methods. Compared to conventional approaches, in vivo Perturb-seq screens based on these AAV vectors can efficiently and economically target and perturb a far greater number of cells in vivo within a single experiment. In some embodiments, the employed AAV vectors are intended to deliver genetic perturbations to target genes in tissues in animal embryos. In some of these embodiments, the employed AAV vector is AAV-Rep2-SCH9. In some other embodiments, the employed AAV vector is AAV2.NN. In some embodiments, the employed AAV vectors are intended to deliver genetic perturbations to target genes in tissues in adult or aged animals, e.g., adult brains and peripheral nerve systems. In some of these embodiments, the employed AAV vector is AAV-PHP.eB. In some other embodiments, the employed AAV vector is AAV9-PHP.S. Detailed structural information of these specific AAV vectors are known in the art. See, e.g., Ojala et al., Mol. Ther., 2018; 26:304-319; and Chan et al., Nat. Neurosci., 2017; 20:1172-1179.

[0041] As demonstrated herein, the AAV vectors can be administered to a transgenic animal system (e.g., an embryo or an adult animal) that expresses a CRISSPR-Cas gene editing system. In some embodiments, the library of AAV vectors is administered to a developing embryo of a CRISSPR-Cas9 transgenic mouse. Some of these embodiments canemploy the AAV-Rep2-SCH9 vector or the AAV2.NN. In some embodiments, these vectors can be administered in utero to the brain (e.g., lateral ventricle) at any time that is equivalent to a mouse embryo developing stage between E11.5 and E19.5. In some of these embodiments, the vectors are administered to the embryo during the E11.5-15.5 period or the E13.5-17.5 period. In some other embodiments, the library of AAV vectors is administered to a developed adult animal of a CRISSPR-Cas9 transgenic mouse. Some of these embodiments can employ the AAV-PHP.eB vector or the AAV9-PHP.S vector. In various embodiments, the AAV vectors are administered to the animal at an age that is equivalent to mouse age of from about postnatal day 10 (P10) to about 18 months old. In some of these embodiments, the vectors are administered to the adult mice via retroorbital injection, as exemplified herein.

[0042] In some preferred embodiments, the genetic perturbations encoded by each of the AAV vectors contain gRNAs that can guide the target genes for modulation (e.g., cleavage) by a CRISPR gene editing system. At least one gRNA targeting a gene of interest (“target gene”) is encoded by the AAV vectors. In some embodiments, two or more gRNAs designed for targeting each gene of interest are included in the library of AAV vectors. The two or more gRNAs targeting one gene can be encoded by one AAV vector. Alternatively, the multiple gRNAs targeting the same gene can be encoded by different AAV vectors in the library. In some embodiments, the animal or animal embryo employed in the screen is derived from a transgenic non-human animal that is engineered to express the CRISPR editing system. In some preferred embodiments, the gene editing system expressed by the animal is CRISPR is CRISPR-Cas9. In some embodiments, each of the gRNAs can be put under the control of a different promoter, including, e.g., human U6 polymerase III promoter as exemplified herein. In some embodiments, a reporter gene such as GFP as exemplified herein can be operably linked to the gRNA-coding sequences in the AAV vectors. The reporter gene allows enrichment of cells expressing the corresponding gRNAs after the AAV vectors are introduced into the animal embryos or developed animals. Unless otherwise described herein, delivery of the vectors into the CRSIPR-expressing transgenic system (e.g., a mouse embryo), expression and detection of the genetic perturbations (e.g., gRNAs for specific target genes), identification (e.g., by reporter gene expression) and enrichment of cells (perturbed cells) expressing the genetic perturbations, and analyses of phenotype- correlating genes in the perturbed cells (e.g., single-cell transcriptomic analysis) can all be performed with techniques specifically exemplified herein and / or well known in the art. See, e.g., US Patent Application Publication No.2021 / 0172017.

[0043] In some embodiments, the AAV vectors delivering genetic perturbations used in the in vivo screen platforms can be further coupled with a transposon system. As described in detail below, this is to ensure rapid, cell-type specific, and reliable expression of gRNAs in target cells as well as their daughter cells for highly efficient gRNA captures in the single-cell transcriptomic analysis. Due to incorporation of a transposon system, the AAV-based in vivo screen platforms of the invention can efficiently and economically target and perturb genes in both embryonic and adult contexts. This represents a significant improvement over genetic screens utilizing lentiviral or any other viral vectors, e.g., the lentiviral vector-based system as described in US Patent Application Publication No.2021 / 0172017.

[0044] In order to construct AAV vectors encoding genetic perturbations in the practice of the invention, the coding sequence for a genetic perturbation (e.g., gRNA-coding sequence) and any other operably linked sequences (e.g., a reporter gene) are often inserted into the viral genome in the place of certain viral sequences to produce a viral construct that is replication-defective. Methods for producing AAV vectors are well-known in the art. AAV vectors have also been used in many reported studies for gene therapy in research and clinical environment. See, e.g., Kaplitt et al., Lancet 369: 2097–105, 2007; Daya et al., Clin Microbiol Rev.21(4): 583–593, 2008; Strobel et al., Am. J. Resp. Cell Mol. Biol.53: 291– 302, 2015; and Kotterman et al., Nat. Rev. Genet.15:445–451, 2014. In some embodiments, the AAV viral vectors and related reagents (e.g., packaging cell lines) suitable for the invention can be obtained commercially. In some of these embodiments, the specific AAV vectors for practicing the invention can be based on the pAAV-MCS construct that is available from Agilent Technologies (Santa Clara, CA). IV. Transgenic animal systems expressing CRISPR functionality

[0045] In some preferred embodiments, the genetic perturbations delivered by the AAV vectors in the in vivo Perturb-seq platforms of the invention include gRNAs to guide target genes for genetic modulation. Typically, a transgenic animal system expressing a CRISPR gene editing functionality is used to express the genetic perturbations and examine their effect on target genes. As exemplified herein, the transgenic animal system can express the CRISSPR-Cas9 gene editing functionality. In these embodiments, the genetic perturbations encoded by the AAV vectors can be gRNAs that are designed to subject one or more target genes to the activities of the CRISPPR-Cas system in the transgenic animal. The library of AAV vectors to be introduced into the transgenic system encode at least one gRNA that isdesigned for each of target genes. In some embodiments, two or more gRNAs are encoded by the library of AAV vectors for each of the target genes. In some of these embodiments, 3 or 4 gRNAs are encoded by the library of AAV vectors for each of the target genes. In some embodiments, the multiple gRNAs designed for a specific target gene are encoded by multiple AAV vectors. In some other embodiments, the multiple gRNAs designed for a specific target gene are encoded by one AAV vector.

[0046] Any transgenic systems with CRISPPR-Cas functionality can be employed in the practice of the invention. In general, these include any transgenic animals or embryos at various development stages that are engineered to express a Cas enzyme and optionally other components required for the CRISPR functionality. The transgenic systems can constitutively or inducibly expresses the Cas enzyme and / or the other components in some or all of their cells. In various embodiments, the Cas enzyme expressed in the transgenic systems is a Cas Type I, II, III, IV, or V protein. In some embodiments, the Cas enzyme (e.g., Cas9) is expressed by the transgenic systems while the other components can be separately provided, e.g., by another vector to be introduced into the transgenic systems (embryos or developed animals). A more detailed description of the various CRISPR-Cas gene editing systems that may be incorporated in the transgenic systems of the invention is provided in, e.g., US Patent Application Publication No.2021 / 0172017.

[0047] In some preferred embodiments as exemplified herein, the genetic perturbations encoded by the vectors are gRNAs that are designed to subject the target genes to CRISPPR- Cas9 gene editing in the transgenic system. In these embodiments, each AAV vector encodes one or more gRNAs that can direct a target gene to the enzymatic activity of Cas9. Methods well-known in the art for genome-scale screening of perturbations in single cells using the CRISPR-Cas9 gene editing system can be readily employed and modified as necessary in the practice of the invention. See, e.g., Dixit et al., “Perturb-Seq: Dissecting Molecular Circuits with Scalable Single-Cell RNA Profiling of Pooled Genetic Screens” 2016, Cell 167, 1853- 1866; Adamson et al., “A Multiplexed Single-Cell CRISPR Screening Platform Enables Systematic Dissection of the Unfolded Protein Response” 2016, Cell 167, 1867-1882; Feldman et al., Lentiviral co-packaging mitigates the effects of intermolecular recombination and multiple integrations in pooled genetic screens, bioRxiv 262121, doi: doi.org / 10.1101 / 262121; Datlinger, et al., 2017, Pooled CRISPR screening with single-cell transcriptome readout. Nature Methods. Vol.14 No.3 DOI: 10.1038 / nmeth.4177; Hill et al., On the design of CRISPR-based single cell molecular screens, Nat Methods.2018 April;15(4): 271-274; Replogle, et al., “Combinatorial single-cell CRISPR screens by direct guide RNA capture and targeted sequencing” Nat Biotechnol (2020). doi.org / 10.1038 / s41587-020- 0470-y; and PCT publication WO / 2017 / 075294). The transgenic systems described in any of these reports may be employed and adapted for use in the practice of the invention. In some preferred embodiments, transgenic mice expressing Cas9 can be used in the practice of the invention. As exemplified herein, Cas9-expressing transgenic mice can be readily obtained from commercial vendors, e.g., the Jackson Laboratory (Bar Harbor, ME). V. Transposon system co-delivered with AAV vectors

[0048] In some embodiments, the AAV-based in vivo Perturb-seq platforms of the invention can additionally incorporate a transposon system. As exemplified herein, co- introduction of AAV vectors encoding gRNAs for perturbing targeting genes and a transposon system can substantially boost gRNA expression level and enable gRNA capture with sparse scRNA-seq readout. Co-expression of the transposon element effectively enhances and stabilizes expression of the gRNAs in embryonic and adult brains and peripheral nerve systems. It ensures rapid, cell-type specific, and reliable expression of gRNAs in target cells as well as their daughter cells for highly efficient gRNA captures in the single-cell transcriptomic analysis.

[0049] Various transposon systems can be incorporated into the AAV-based in vivo screen platforms of the invention. In some embodiments, the employed transposon system is a DNA transposon. Integration of a DNA transposon system with the genetic perturbation delivering AAV vectors can be performed in accordance with the specific protocols exemplified or methods well known in the art. Many DNA transposon systems well known in the art can also be employed and modified for use in the practice of the present invention. These include a wide range of individual DNA transposons from different families and donor organisms that have been characterized in detail in the art, e.g., Tol2 originating from medeka fish, Sleeping Beauty - synthetic sequences derived from transposons found in the white cloud minnow, Atlantic salmon and rainbow trout, and piggyBac isolated from the cabbage looper moth. The piggyBac DNA transposon encodes a transposase (iPB) that is highly active in mammalian cells and integrates at the conserved TTAA tetranucleotide sequence. There is also a murine codon-optimized PB transposase (mPB) that was 20 times more efficient than iPB in mouse embryonic stem cells. Further, a hyperactive form of PB transposase (hypPB) with several amino acid substitutions on mPB is known in the art. This DNA transposonsystem is able to increase relative integration frequency by nine times over mPB in mouse embryonic stem cells. Detailed technical information of this transposon system is provided in, e.g., Yusa et al., Proc Natl Acad Sci USA.2011; 108:1531–1536. In addition to the PB transposase and derivatives, other transposon systems well-known in the art can also be employed in the practice of the invention. These include, e.g., the Sleeping Beauty transposon system, hAT transposons, the MuDR superfamily of transposons, the Tc1 / mariner transposon family, transposons utilizing the IS30 transposase such as Tn2700 and Tn2702, and miniature inverted-repeat transposable elements from Bombyx mori and Rhodnius prolixus. See, e.g., Querques et al., Nat. Biotechnol.2019; 37: 1502–1512; Rubin et al., Genetics.2001; 158: 949–957; Raizada et al., Plant J.2001; 25:79-91; Dupeyron et al., Mobile DNA 2020; 11: 21; Szabó et al., J. Bacteriol.2010; 192: 3414-23; and Zhang et al., Genome Biol Evol.2013; 5: 2020–2031.

[0050] These well-known DNA transposons are composed of a transposase gene and flanking inverted terminal repeats (ITRs). The enzyme transposase recognizes specific short target sequences, called directed repeats (DRs) located in the ITRs. Upon binding, the transposase cuts out the transposon sequence from the surrounding genomic DNA of the host cell. The formed complex consisting of the mobilized transposon DNA fragment and the still bound transposases is now able to change its position to a new location in the cell genome. The transposases open the genomic DNA backbone at the new locus and insert the transposon fragment. The ligation of the open DNA ends is mediated by cellular key factors of the non- homologous end joining pathway (NHEJ) within the double strand break (DSB) repair system. Employment of any of these transposon systems in the practice of the invention can be readily carried out in accordance with the teachings provided in the art. See, e.g., Fraser et al., Insect Mol Biol.1996; 5:141–151; Ivics et al., Insect Mol Biol.1996; 5:141–151; Kawakami et al., Gene.1998; 225:17–22; Cadiñanos et al., Nucleic Acids Res.2007; 35:e87; Muñoz-López and García-Pérez, Curr Genomics.2010; 11:115–128; Grabundzija et al., Mol Ther.2010; 18:1200–1209; Mátés et al., Curr Genomics.2010; 11:115–128; Huang et al., Mol Ther.2010; 18:1803–1813; Yusa et al., Proc Natl Acad Sci USA.2011; 108:1531–1536; and Gogol-Döring et al., Mol Ther.2016; 24:592–606.

[0051] In the practice of the invention, the transposon system can be introduced into the transgenic animal system in the same AAV vectors that encode the genetic perturbations. Alternatively, a dual vector system can be used. In these embodiments, the coding sequence of the genetic perturbation (e.g., gRNA) is flanked with a transposon element in a first AAVvector, and a separate vector (e.g., a second AAV vector) encodes a transposase that is able to recognize the transposon sequence. In some embodiments of the invention, the employed DNA transposon system includes transposase hypPB (e.g., hyperactive piggyBac) and a transposon sequence recognized by the transposase (e.g., inverted terminal repeat sequences). The PiggyBac (PB) transposon is a highly useful transposon for genetic engineering of a wide variety of species. It can efficiently transpose between vectors and chromosomes via a "cut and paste" mechanism. During transposition, the PB transposase recognizes transposon- specific inverted terminal repeat sequences (ITRs) located on both ends of the transposon vector and efficiently moves the contents from the original sites and integrates them into TTAA chromosomal sites. The PiggyBac transposon system enables genes of interest between the two ITRs (e.g., gRNA-encoding sequence) in the PB vector (e.g., AAV vectors exemplified herein) to be easily mobilized into target genomes. As demonstrated herein, incorporation of this transposon system into the AAV-based Perturb-seq platform is able to overcome the challenge of gRNA dilution and significantly improve genetic perturbation and labeling efficiency in vivo. VI. Correlating perturbed genes with cellular phenotypes or functions

[0052] Once a library of AAV vectors encoding genetic perturbations (e.g., gRNAs) specific for one or more target genes are introduced into the transgenic animal or animal embryo, cells expressing the genetic perturbations can be identified and isolated from the live animal. The identified cells can be further enriched, before being subject to examination for their biological and cellular properties including phenotypic characteristics. In various embodiments, examination of the perturbed cells for modified or altered biological and cellular properties can include measuring genomic, genetic, proteomic, epigenetic and / or phenotypic differences in single cells. The specific genetic perturbations expressed in individual cells (e.g., gRNAs targeting specific genes) can then be correlated with the biological properties (e.g., phenotypic changes) revealed from the examination for the different cells. In some embodiments, phenotypic changes resulting from the genetic perturbations in the identified cells relate to changes in gene or protein expression. In some embodiments, a plurality of target genes can be perturbed in single cells, and gene expressions in the perturbed cells are then analyzed (e.g., via scRNA-seq). These methods allow identification of networks of genes that are disrupted due to perturbation of a single or signature target gene. Understanding the network of genes effected by a perturbation mayallow for a gene to be linked to a specific pathway that may be targeted to modulate the signature gene. Thus, in certain embodiments, perturb-seq is used to discover novel gene and drug targets to allow treatment of various diseases in which the target genes are involved.

[0053] Various tools are available to examine cells expressing the genetic perturbations and exhibiting changes in certain gene or protein expressions. In some embodiments, a sequencing analysis can be used to identify cell types and modulated genes in the perturbed cells. In some of these embodiments, analysis of phenotypic changes due to the introduced genetic perturbations (e.g., editing of a specific target gene) in the identified cells and correlation of a perturbation with an observed phenotypic change in single cells can be performed via single-cell RNA sequencing (scRNA-seq). In some of these embodiments, the gRNAs encoded by the AAV vectors are functional for gene editing and also can be directly captured in the scRNA-seq technology. This is to facilitate correlation of an observed phenotype with a specific perturbed gene in the identified cells. In some embodiments, the scRNA-seq analysis aims to detect gRNAs in the identified cells by sequencing transcripts expressed from the AAV vectors. In some embodiments, a given gRNA and a corresponding specific barcode are co-expressed from the same AAV vector, and detection of the gRNA is realized by detection of the specific barcode via RNA sequencing.

[0054] In some embodiments, phenotypic analyses of the identified cells entail determining cell types and corresponding perturbations (e.g., gRNA sequences encoded by the AAV vectors) via single cell RNA (scRNA) sequencing of the identified and optionally further enriched perturbed cell population. scRNA sequencing of a population of cells can be readily performed in accordance with experimental protocols that are routinely practiced in the art. See, e.g., Picelli et al., Nat. Protoc.2014; 9: 171-181; Ramskold et al., Nat. Biotechnol.2012; 30: 777-782; Hashimshony et al., Cell Rep.2012; 2: 666-673; Kalisky et al., Ann. Rev. Genetics 2011; 45: 431-445; Kalisky et al., Nat. Methods 2011; 8: 311-314; Islam et al., Genome Res.2011; 21: 1160-7; Tang et al., Nat. Protoc.2010; 5: 516-535; and Tang et al., Nat. Methods 2009; 6: 377-382. In some preferred embodiments, phenotypic analyses of the identified and enriched perturbed cells involve high-throughput scRNA sequencing. This can be performed with the specific techniques exemplified herein and / or experimental procedures well-known in the art. See, e.g., Rosenberg et al., Science 2018; 360: 176-182; Vitak et al., Nat. Methods 2017; 14: 302-308; Cao et al., Science 2017; 357: 661-667; Gierahn et al., Nat. Methods 2017; 14: 395-398; Macosko et al., Cell 2015; 161: 1202-1214; Klein et al., Cell 2015; 161: 1187-1201; Zheng et al., Nat. Biotechnol.2016, 34:303-311; Zheng et al., Nat. Commun.2017; 8: 14049; Zilionis et al., Nat Protoc.2017; 12: 44-73; and PCT publications WO2016 / 040476, WO2016168584, and WO2014210353. In some embodiments, phenotypic analyses of the identified and enriched perturbed cells involve can be accomplished by single nucleus RNA sequencing. This can be performed using standard protocols that have been reported in the literature. See, e.g., Drokhlyanskyetal., Cell 2020; 182: 1606–1622; Swiech et al., Nat. Biotechnol.2014; 33: 102-106; Habib et al., Science 2016; 353: 925-928; Habib et al., Nat. Methods.2017; 14: 955-958; and PCT publications WO2017 / 164936, WO2019 / 094984 and WO2020 / 077236.

[0055] Some embodiments of the invention utilize a CRISPPR-Cas9 gene editing system to perturb target genes in a CRISPPR-expressing transgenic system described herein. This can be performed with the transgenic system (e.g., Cas9 transgenic mice) and specific experimental protocols exemplified herein. Additional guidance for performing genome-scale screening of perturbations in single cells using CRISPR-Cas9 have been provided in the art. See, e.g., Dixit et al., “Perturb-Seq: Dissecting Molecular Circuits with Scalable Single-Cell RNA Profiling of Pooled Genetic Screens” 2016, Cell 167, 1853-1866; Adamson et al., “A Multiplexed Single-Cell CRISPR Screening Platform Enables Systematic Dissection of the Unfolded Protein Response” 2016, Cell 167, 1867-1882; Feldman et al., Lentiviral co- packaging mitigates the effects of intermolecular recombination and multiple integrations in pooled genetic screens, bioRxiv 262121, doi: doi.org / 10.1101 / 262121; Datlinger, et al., 2017, Pooled CRISPR screening with single-cell transcriptome readout. Nature Methods. Vol.14 No.3 DOI: 10.1038 / nmeth.4177; Hill et al., On the design of CRISPR-based single cell molecular screens, Nat Methods.2018 April; 15(4): 271-274; Replogle, et al., “Combinatorial single-cell CRISPR screens by direct guide RNA capture and targeted sequencing” Nat Biotechnol (2020). doi.org / 10.1038 / s41587-020-0470-y; International publication serial number WO / 2017 / 075294); and US Patent Application Publication No.2021 / 0172017. Any of the methods described in the art can be adopted and modified in the practice of the invention. EXAMPLES

[0056] The following examples are provided to further illustrate the invention but not to limit its scope. Other variants of the invention will be readily apparent to one of ordinary skill in the art and are encompassed by the appended claims.Example 1. Barcoded AAV screens identify serotypes efficiently expressed in the developing brain

[0057] Existing Perturb-seq systems relies on lentiviral vectors, which have relatively low packaging yield and limited tissue penetration in vivo. AAV vectors are advantageous for in vivo studies due to their high production yields and capsid engineering potential. However, most AAVs take weeks to reach prime expression; and there is no documented AAV to efficiently transduce newborn neurons and progenitors in the developing brain, which requires a faster expression onset. To identify an AAV serotype targeting developing brain in vivo, we constructed a barcoded library of 86 phylogenetically diverse AAV serotypes; each expresses a green fluorescent protein (GFP) with a unique DNA barcode upstream of the polyadenylation initiation sites (Fig.1, Panel A). This library consists of engineered serotypes that were reported or published previously, including AAV1, AAV2, AAV3B, AAV4, AAV5, AAV6.2, AAV7, AAV8, AAV9, AAV10, AAV 12 and AAV13; several of these serotypes are commonly used for in vivo studies with minimal toxicity as previously reported (Table 1). Table 1. AAV serotypes lists and barcodes in the 86-AAV and 14-AAV libraries.

[0058] To identify a fast-acting serotype in developing mouse brains (C57BL / 6J strain), we administered this library of AAV vectors in utero into the lateral ventricles at embryonic day 13.5 (E13.5) (0.5-1.5μL, 0.5-1.5e10viral genome / embryo). At 48 hours, we observed many GFP+cells in the cortex and ganglionic eminence (Fig.1, Panels B-C). Throughimmunofluorescence marker labeling, we observed that 56% of GFP+cells were newborn neurons (TBR1+) that our AAV library could directly or indirectly transduce. We found 3% of GFP+cells express the intermediate progenitor (IP) marker TBR2, which indicates that either the AAVs were less potent at transducing IPs, or the GFP was quickly diluted as IPs differentiate. At 24 hours, a minimal GFP signal was detected from the brain tissue with immunofluorescent amplification, suggesting that most AAV vectors took 24 to 48 hours to express the transgene in vivo to detectable levels. In parallel to this in vivo analysis, we conducted the screen in vitro with a mouse hippocampal neuronal cell line HT22 and could detect GFP expressions in vitro at 24 hours post transduction and it persisted at 48 hours. Example 2. AAV-SCH9 rapidly transduces developing brains within 48 hours

[0059] To identify the AAV serotypes that are enriched in GFP+cells, we purified the cells from the neocortex (in vivo) and HT22 cells (in vitro) at 24- and 48-hours post- transduction and quantified the barcode abundance using next-generation sequencing (Fig.1, Panel D). Compared to the initial distribution in the AAV library, 48 hours post-transduction, both in vivo and in vitro, we detected significant shifts in the barcode distribution. We found that, 48 hours post-transduction, several AAV serotypes were enriched in vitro (AAV-P1529, AAV2-NN, AAV2-P1583 and AAV-DJ), several serotypes were enriched in vivo (AAV1, AAV2-P1558, AAV-Hu48.2 and AAV2-7M8), and several were enriched in both (AAV- SCH9, AAV-SCH9repeat136bp, AAV2-P1576, AAV2-P1579 and AAV2-P1596) (Fig.1, Panel D; Table 1). This likely reflects a shared pattern of AAV tropism for newborn neurons and progenitors in vivo and in vitro, as well as a distinct pattern of enrichment in vivo.

[0060] To further examine the pattern of AAV barcode distribution across time points and across different contexts, we performed a principal component analysis of the samples (Fig.1, Panel E). PC1 explained 63% variance of the data and separated in vitro data (24 and 48h), in vivo data (48hr), from in vivo data (24h) and the initial AAV library. PC2 (10% variance) separated in vivo (48h) data from all other samples. Consistent with the in vivo tissue analysis, it takes 48h to observe the change of barcode distribution in vivo, while the impact is quicker and more consistent at 24h and 48h in vitro.

[0061] We observed a 11.5-fold increase in the proportion of barcode expression in our top hit, AAV-SCH9: it constituted 2.3% of the initial AAV library and 48 hours after in vivo transduction, it constituted 27.4% of the enriched population. Differential expression analysis revealed that AAV-SCH9 was the most significant hit from the in vivo experiment, followedby a variant of the same serotype (AAV-SCH9-repeat136bp). Both AAV-SCH9 and its variant showed significant enrichment in vitro 24 hours post transduction that persisted at 48 hours (Fig.1, Panel G). We also quantified the relative expression dynamics of several top hits, including AAV-SCH9, by comparing the abundance of the barcodes in 24 and 48-hour post-transduction (Fig.1, Panel F).

[0062] Critically, compared to AAV-SCH9, several widely used neuron-targeting AAVs required longer time to reach prime expression. Within the initial 48 hours, AAV9-PHP.eB and AAV-DJ showed only minimal expression in vivo, with 0.2-fold and 1.1-fold changes relative to their compositions in the initial AAV library, respectively (Table 1; and Fig.1, Panel D). In contrast, AAV-SCH9 demonstrated a 11.5-fold increase in expression abundance, making it an optimal choice for perturbing and studying gene function in a dynamic developmental context when cells are actively differentiating and maturing. Example 3. Diverse cell type tropisms of AAV serotypes in vivo with single-cell resolution

[0063] We identified several serotypes, including AAV-SCH9, that exhibit efficient transduction of the developing brain and neuronal cell lines by bulk measurement, but their precise cell type-specificity were still unclear. To further characterize the tropism of the AAV hits, we selected the top 14 serotype hits into a secondary, new library for validation with single cell resolution (Table 1). For these experiments, each AAV was designed to express a GFP reporter with a set of 3 unique barcodes upstream of the polyadenylation initiation sites and then pooled them with equal titer (Fig.2, Panel A). We administered this 14-AAV library into the lateral ventricle of E13.5 mouse embryo brain and collected cortical cells at E15.5 to perform droplet-based single-cell RNA-seq (scRNA-seq) with both the sorted (GFP+) population and unsorted population (Fig.2, Panels A-B). We assigned the AAV serotype identities by using the barcodes, captured in a dial-out PCR library (Fig.2, Panel C).

[0064] After quality control, we retained a total of 14,835 neocortical cells for further analysis (7,630 cells from the FACS-enriched GFP+population, and 7,205 cells without enrichment) (Fig.2, Panel B). We partitioned the cells into major cell types and annotated them based on known marker gene expressions (Table 3) (Di Bella et al., 2021; La Manno et al., 2021; Tasic et al., 2018). These cells were clustered into 11 cell types including upper and deep layer projection neurons (ULPN, DLPN), migrating neurons (Mig. neurons), apical progenitors (Api. prog), intermediate progenitors (IP), interneurons derived from the medial ganglionic eminence (IN-MGE), interneurons derived from the non-medial ganglioniceminence (IN-non-MGE), Cajal-Retzius cells (CR), fibroblast (Fibro), mural cells (Mural), and microglia (Mg).

[0065] From inspecting the relative abundance of the cell types in GFP+and unsorted populations, most cell types were conserved. Two populations, apical and intermediate progenitors, showed decreased representation in the GFP+population, reflecting either a reduced capacity of the AAV to transduce dividing cells, or the cellular division diluting the transgene products, or both. From inspecting the individual barcodes across cell populations, we found that four serotypes were highly efficient in transducing in vivo: AAV-SCH9 (BC1), AAV2-NN (BC2), AAV-SCH9-G160D (BC4) and AAV2-7M8 (BC10) (Fig.2, Panel C). Within the AAV-SCH9 transduced populations, we detected 12.0% upper layer projection neurons, 56.5% deep layer projection neurons, 10.1% MGE-derived interneurons, 7.9% migrating neurons, 2.8% apical progenitors, and 1.3% intermediate progenitors (Tables 2-3). Table 2. Summaries of scRNA-seq experiments (detailed information about each batch of Perturb-seq)Table 3. Cell type classification and differential expressed genes in 14-AAV library E16.5 scRNA-seq dataExample 4. In situ characterizations of AAV-SCH9 reveal its fast-acting dynamics and broad neuronal labeling in the developing brain

[0066] Using bulk sequencing and scRNA-seq, we identified and validated AAV-SCH9 as an effective vector to rapidly (< 48 hours) transduce embryonic cortical tissues in vivo. Previously, AAV-SCH9 has been reported to target subventricular adult neural stem cells (Ojala et al., 2018). To further characterize its brain-wide action and tropism, we administered AAV-SCH9 expressing a nuclear membrane-anchored fluorophore (GFP- KASH) in utero into the lateral ventricles at embryonic day 14.5 (E14.5). After 48 hours, we observed many GFP+cells in the neocortex, ganglionic eminence, and around the lateral ventricle, similar to the pattern from the 86-AAV library (Fig.2, Panel D; and Fig.1, Panels B-C).

[0067] We co-stained the section for markers of deep layer excitatory projection neurons (CTIP2 and TBR1) and intermediate progenitors (TBR2). We found on average 37% of GFP+cells express CTIP2 and 73% express TBR1, indicating that AAV-SCH9 can induce broad expression in newborn neurons. For GFP+cells, 11% expressed the intermediate progenitor marker TBR2, indicating either the transduction was less efficient in progenitors and / or the expression is quickly diluted as IPs divide. The relative proportion of cell types measured byimmunohistochemistry and scRNA-seq data generally agree, as expected. These data confirmed the in vivo performance of AAV-SCH9 to label neurons and intermediate progenitors within 48 hours.

[0068] AAV-SCH9 may directly transduce neurons, or first transduce progenitors which then differentiate into neurons. To identify its tropism and differentiate these two possibilities, we performed in utero transduction of AAV-SCH9-GFP-KASH at two embryonic ages when two distinct populations of neurons (deep versus upper layers) are born. If the AAV only labels progenitors, we expected to observe enrichment of GFP expression in different layers, targeting different classes of newly born neurons on different injection day (only deep layers for E13 injection, and only upper layers for E17 injection, respectively). Brains were harvested after 48 hours and co-stained with DLPN marker CTIP2 and neural progenitor marker PAX6 to distinguish the laminar layers. We divided the cortical layers evenly into 6 bins: the E13.5 transduced cells were enriched in bin 2 (46%) and bin 3 (23%) with CTIP2 expression, indicating their deep layer cortical sub-cerebral projection neuron identities (Arlotta et al., 2005; Chen et al., 2008). By contrast, the E17.5-transduced cells were broadly distributed in the later-born ULPNs in bin 1 (22%) and Layer 5 DLPNs in bin 2 (31%), supporting that AAV-SCH9 could label newborn neurons. (Fig.2, Panel E). Interestingly, we observed a higher level of GFP+cell labeling in the ventricular and subventricular zone in the E17.5- than E13.5-transduced sample (18% versus 9% in bin 5, 15% versus 5% in bin 6), supporting the progenitor-labeling capacity of AAV-SCH9 and consistent with the slower cycling and differentiation capacity in the late neurogenesis stage (Greig et al., 2013). These data, together with the scRNA-seq results, collectively supports that AAV-SCH9 transduces both newborn neurons and progenitors in embryonic mouse brain with a fast onset of expression.

[0069] One of the advantages of AAVs is their excellent tissue penetration and labeling in vivo. Besides the neocortex, AAV-SCH9 can transduce additional regions, including olfactory bulb, striatum, hippocampus, thalamus, and cerebellum, with labeling density suitable for future cross-region Perturb-seq studies (Fig.2, Panel F). By contrast, we observed limited expression with lentiviral transduction (using optimal, >9×109U / mL high- titer vector): there were much fewer GFP+cells in most regions, especially cerebellum, thalamus, and striatum (Fig.2, Panel F), possibly due to limited tissue penetration in vivo. Altogether, this scale-up Perturb-seq holds greater potential for its ability to access abundant cells across brain regions, allowing a thorough comparison of the brain region- and cell type-specific perturbation effects of disease risk genes in vivo. Furthermore, AAV-SCH9 transduction surpasses conventional lentiviral vectors in labeling neocortex and other brain regions, significantly streamlining the sampling process for scRNA-seq and reducing the associated costs and labors. Example 5. Transposase hypPB stabilizes and enhances transgene expression

[0070] In contrast to lentiviral vectors, most AAVs remain episomal in host cells due to the lack of integrase (Deyle and Russell, 2009). Therefore, when performing Perturb-seq using AAV vectors in the rapidly dividing cellular context in vivo, the gRNA and fluorophore expression from AAV will be diluted, and eventually lost, upon cell division and growth. To overcome the challenge of gRNA dilution and improve genetic perturbation and labeling efficiency in vivo, we designed a dual-vector system including a hypPB (hyperactive piggyBac) transposase and a transposon with inverted repeat flanking the gRNA and fluorophore (Moudgil et al., 2020) (Fig.3, Panel A). Upon transduction, hypPB can integrate the transgene into the nuclear genome for inheritance in the daughter cells, allowing consistent expression rather than transient, episomal expression (Fig.3, Panel A).

[0071] We first tested the design in vitro and transfected HT22 cells followed by time- lapse imaging. In the presence of hypPB, we detected more GFP+cells (Fig.3, Panel B) and enhanced GFP expression levels with faster onset; the expression was resistant to cell passages over the course of six days. hypPB allowed faster expression onset within a few hours post-transfection which would be highly advantageous for labeling and perturbing cells during the dynamic developmental process in vivo. Example 6. hypPB improves transgene expression in vivo through embryonic transduction

[0072] To evaluate whether hypPB can enhance gene expression in vivo, we first administered AAV-SCH9 containing a transposon, with or without an AAV-SCH9-hypPB in utero at E14.5 and performed immunofluorescence analysis at P7. We detected hypPB (HA- tagged) expression in many brain regions including cortical laminar layers, similar to the expected AAV-SCH9 tropism (Fig.3, Panel C). hypPB expression did not introduce overt toxicity in development (Fig.3, Panel C) as we measured the brain sizes and gliosis markers GFAP and IBA1 expression changes in the presence of hypPB.

[0073] We quantified the number of GFP+cells in the somatosensory neocortex from brains administered with AAV-transposon, with and without hypPB, in utero. Co-transduction of hypPB increased the labeling efficiency: 2.2-fold more cells were GFP+(Fig. 3, Panel F). Consistently, we divided the cortical laminar layers evenly into 4 bins and detected 3.0-fold more GFP+cells with hypPB co-expression in the upper layers (15 versus 5 cells / mm2, corresponding to bin 1), 1.8-fold in Layer 5 (108 versus 59 cells / mm2in bin 3), and 3.2-fold in Layer 6 (123 versus 39 cells / mm2in bin 4). This result is consistent with our in vitro study that hypPB co-expression leads to retention of expression in dividing and differentiating cell lineages. Example 7. hypPB enhances transgene expression in adult central and peripheral nervous system

[0074] The utility of applying Perturb-seq in vivo have been demonstrated in the past in embryonic brains (Dvoretskova et al., 2023; Jin et al., 2020). However, the need for a comprehensive in vivo screen platform for the adult central and peripheral nervous systems – especially pertinent to neurodegenerative diseases – still exists. We hypothesized that hypPB could amplify AAV transgene expression in postmitotic neurons in adult central and peripheral nervous systems, where AAV transgene levels have been observed to be relatively low compared to genomic expression (Lang et al., 2019). To test this hypothesis, we delivered AAV vectors to postmitotic neurons in adult mice using two variants that cross the blood-brain barriers: AAV9-PHP.eB for neurons and glia in the brain, and AAV9-PHP.S for peripheral neurons, both through non-invasive retro-orbital administrations (Chan et al., 2017).

[0075] In adult brain, hypPB generally increased GFP expression levels (Fig.3, Panel D). In the somatosensory cortex, average fluorescence intensity was 1.7-fold higher across laminar layers with hypPB (Fig.3, Panel G). Furthermore, by extracting nuclei from cortex and performing flow cytometry analysis, we found that the AAV9-PHP.eB transposon system can label 6.5-7.0% of total nuclei in the cortex, 4.0-fold higher than without hypPB. This indicates hypPB can enhance transgene expression in adult, post-mitotic cells. Notably, different from the AAV-SCH9 delivery, elevated levels of gliosis markers IBA1 and GFAP were detected in the retro-orbital AAV9-PHP.eB-hypPB transduction conditions. This is likely due to the neuroinflammation response by the systemic viral administration method and the amount of vector used (Perez et al., 2020) as well as hypPB genome integration, indicating the importance of carefully evaluating AAV administration to minimize toxicity concerns.

[0076] To test the performance of this system in the adult peripheral nervous system including the dorsal root ganglion, we performed retro-orbital injection of the AAV9-PHP.S gRNA-tdTomato construct, with or without the AAV-hypPB (Fig.3, Panel E). Similarly, the presence of hypPB also increased the reporter expression levels and labeling efficiency (Fig. 3, Panel E). Altogether, hypPB transposon can further enhance transgene expression in adult nervous systems, combined with other AAV9 variants. This opens doors to study genetic perturbations and gene functions in adult tissues, especially in the context of aging or degeneration of central and peripheral nervous systems, beyond the capability of conventional lentivirus-based genetic screens. Example 8. hypPB transposon integrates into the host genome with random insertions

[0077] Genome integration of the transgenes, via lentiviral vectors or transposons, could trigger unwanted changes in the genome and cellular activities. For instance, integrations in coding regions could profoundly influence function of nearby genes. To characterize the transposon integration preferences in neurons in vivo, we extracted the mouse genome from over 170,000 tdTomato+nuclei from brains co-administered with AAV-SCH9-hypPB and transposons in utero. We then performed 60x paired-end whole-genome sequencing to identify hybrid reads that capture the junctions between the mouse genome and transposon, which provides evidence of transposon integration sites (Fig.3, Panels H-I). We identified 46 such insertion events, several of which were supported by multiple reads, distributed across the mouse genome (Table 4). Overall, most of the insertion events were in the intergenic regions (41%). We observed a slight preference for transcription start sites amongst insertions within the promoter regions (Fig.3, Panels H-I), consistent with the literature on hypPB specificity (Chen et al., 2020). Integration events occurred at the expected TTAA flanking sites, further demonstrating the true insertion performance of hypPB in vivo. These analysis showed that hypPB integration sites are largely random across the genome. The perturbation effects in each cell could be confounded by the hypPB-induced genome integration, which also exists for lentiviral vector-based screens. Since the events are largely random and in the intergenic regions, sampling large number of cells for each perturbation group could minimize this bias to allow a robust and well-powered analysis. Table 4. Whole genome sequencing analysis of transposon insertion eventsExample 9. AAV-hypPB permits gRNA capture with sparse scRNA-seq readout

[0078] One challenge of in vivo CRISPR screens is to retain high gRNA expression for efficient gene editing as well as gRNA recovery in the sparse scRNA-seq (Kalamakis and Platt, 2023). We next tested if the transposase system also could enhance gRNA expression levels and found that in vitro hypPB co-transfection led to a 4.2-fold increase in the gRNA expression. Interestingly, the presence of Cas9 stabilized gRNA expression by 5.7-fold, likely by forming ribonucleoprotein complexes to prevent gRNA degradation (Hendel et al., 2015). With the co-expression of hypPB and Cas9, gRNA expression level was increased by 10.8- fold, likely through additive mechanisms, which could improve the gene editing efficiency as well as gRNA detection in the scRNA-seq readout.

[0079] With the fast-acting AAV-SCH9 and the hypPB system to enhance the expression, we moved on to characterize the gRNA identity recovery rate of this Perturb-seq platform by two established scRNA-seq strategies. Embryonically transduced cortical cells were postnatally dissociated, purified and followed by scRNA-seq, with two methods to capture gene expression and gRNA identities: 3’ scRNA-seq, which uses on-beads barcoded oligo-dT and gRNA primers to label and amplify the 3’ ends, or 5’ scRNA-seq, which uses in-solution oligo-dT and gRNA primers to capture the transcript from 5’ end (Replogle et al., 2020). The 5’ method benefits from a higher concentration of gRNA-specific primers in solution (rather than on the beads) during reverse transcription, which is expected to give rise to a higher gRNA capture rate.

[0080] We observed that, with similar sequencing depths, the 5’ and 3’ scRNA-seq gene expression analysis resulted in similar cell clustering and data quality (Table 5). However, with 3’ scRNA-seq, we detected only 0.1% of reads assigned to gRNA from the library, whereas the 5’ scRNA-seq gave rise to 52.9% of reads assigned to gRNA. This is consistent with reliable assignment of the gRNA to the cell barcodes and high detection of gRNA levels from 5’ scRNA-seq. Moreover, in the 5’ scRNA-seq we can observe a threshold to separate cells with high gRNA expression from those with low gRNA expression, likely separating the true expression from the ambient, spurious, or background low-level expression. Using in vivo AAV Perturb-seq, we reliably detected gRNA identity to 63-83% of cells by 5’ scRNA- seq, which was about 7-10-fold higher than those by 3’ chemistry at 7%-14% (Fig.4, Panel D; and Table 5). Table 5. Cell type classification of 5’ and 3’ scRNA-seq data (parameters used in Seurat for cell type clustering)Example 10. Efficient cell labeling, high gRNA recovery, and validated CRISPR effects in vivo

[0081] We applied the AAV-based Perturb-seq system to a proof-of-principle in vivo screen targeting transcription factors with roles in brain development: Foxg1, Nr2f1 (COUP- TFI), Tbr1, and Tcf4. Haploinsufficiency of these genes is associated with neurodevelopmental disorders and diseases, including Bosch-Boonstra-Schaaf optic atrophy syndrome (Nr2f1), Pitt-Hopkins Syndrome (Tcf4), FOXG1 syndrome (Foxg1) and TBR1 syndrome (Tbr1). These transcription factors play critical roles in cortical neuronal differentiation and progenitor maintenance, regional patterning, neuronal migration, and circuit assembly through the regulation of networks of other transcription factors (Chen et al., 2021; Greig et al., 2013; Hou et al., 2020). Most of their expression spans from embryonic to postnatal stages: Tcf4 and Nr2f1 are broadly expressed in most cell types, whereas Foxg1 and Tbr1 expressions are restricted to subclasses of neurons including deep layer projection neurons and immature neurons (Di Bella et al., 2021; La Manno et al., 2021).

[0082] We designed four gRNAs for each gene by targeting the early 5’ coding exons and tested them in vitro to select the three best-performing gRNAs to be included in the pool (Table 6). We pooled them equally along with four control gRNAs including non-targeting (NT) control and safe-targeting (ST) control gRNAs (Morgens et al., 2017). AAV-SCH9 vectors expressing hypPB and pooled gRNAs were administered into embryonic lateral ventricles at E14.5 (Fig.4, Panel A). At P7, the isocortex and hippocampal formation was micro-dissected and dissociated; the BFP-expressing cells were enriched and processed for droplet-based 5’ scRNA-seq with direct capture of gRNA (Replogle et al., 2020). In this experiment, we intentionally controlled the AAV injection titer to achieve <2% of cells transduced in the neocortex to limit multiple perturbation events, although this ratio can be further increased if testing combinations of perturbation is desired. Notably, with this intentional dilution, our system already achieves the collection of >10-fold more BFP+perturbed cells than the conventional lentiviral labeling, which significantly streamlines the experiment. Table 6. gRNA sequences

[0083] Across five replicates (10x Chromium channels) with a total of 11 animals from two litters, we obtained a total of 50,075 cells profiled after the primary quality control(Table 2). Remarkably, this is much more efficient than our previous lentiviral-based work (to collect 46,770 cells over 17 litters and 163 animals) (Jin et al., 2020), demonstrating the prospective scalability of this new platform. All the cells were classified into six general cell types and 21 subclusters, annotated by direct comparison with public scRNA-seq data (Yao et al., 2021).

[0084] To accurately assign the perturbation identities to cells, we evaluated multiple methods. We found that the most reliable approach involved first down-sampling the gRNA counts in each cell to minimize biases due to count differences, followed by DemuxEM, a tool designed for demultiplexing hashing barcodes (Gaublomme et al., 2019). We identified 16,067 cells with a single gRNA perturbation assigned, 16,373 cells with multiple gRNAs assigned, and 17,635 cells with no gRNA assigned, most of which were glia. Glial populations were associated with low gRNA recovery and low BFP detection (expressed from the AAV transgene), consistent with our previous characterization of AAV-SCH9 tropism (Fig.2). Overall, we successfully assigned 57-77% of cells to perturbations in our data from the non-glia cell types, with a median of 152 UMI gRNA detected per cell and 9,687 UMI endogenous transcripts per cell in excitatory neurons (Fig.4, Panel E). We further filtered low-quality cells with low UMI or a high percentage of intronic reads, which likely indicates they were cytoplasmic, nuclear debris, or a similar population (La Manno et al., 2021).

[0085] Next, we narrowed our downstream analysis to 11,471 high-quality cells, each with only a single gRNA perturbation. We retained 14 annotated cell types, following the classification taxonomy of isocortex and hippocampal formation previously described (Yao et al., 2021). This included ten clusters of excitatory glutamatergic neurons including the L2-4 upper layer, L5 / 6 IT (intratelencephalic), Car3, PT (pyramidal tract), NP (near-projecting), CT (corticothalamic), and L6b; three clusters of GABAergic inhibitory neurons (Id2, Sst, and Lhx6+Sst-), and Cajal-Retzius cells (Fig.4, Panel B) from a total of 18 subclusters. We did not observe any apparent batch effects on cell type across the five channels in UMAP space (Fig.4, Panel C). Each gRNA had 318-1,065 cells assigned to it, giving us 1,609-2,527 cells perturbed for each gene and 3,316 control cells.

[0086] To evaluate the CRISPR loss-of-function and target gene dosage in vivo, we extracted the endogenous transcript reads from reads that are associated with cells receiving a gRNA perturbation in the 5’ scRNA-seq libraries. We detected a substantial proportion of reads harboring insertion or deletion events within each gRNA-targeting region (Fig.4, PanelF). For example, gRNA1 of Foxg1 induced several distinct 1-base pair insertions in 39% of the reads, frameshift mutations leading to premature transcriptional termination downstream. This is likely an underestimation of the perturbation effects, as nonsense mediated decay should degrade much of the mutated mRNA. We speculated that different gRNAs, targeting the same gene, may yield variable phenotypic outcomes due to their differential efficacies in gene editing, introducing potential phenotypic heterogeneity. We also observed that wild-type transcripts were detectable in our data, suggesting that in vivo Perturb-seq could allow the analysis of both full knockout and heterozygous loss-of-function effects across different cells collected. Example 11. Differential cell type proportion changes by transcription factor perturbation in vivo

[0087] To analyze the cell type proportion changes resulting from each perturbation, we performed statistical tests considering the proportion differences across batches and gRNA identities, compared to a non-targeting control group (NT-2) (Phipson et al., 2022) (Fig.4, Panel G). We decided to perform this analysis at the gRNA level, rather than the target-gene level, due to the known variability of gRNA performance (Fig.4, Panel F). We also included other control gRNA groups, compared to this chosen control group NT-2, to test the robustness of the analysis, as we expected little to no effect in these control-to-control comparisons.

[0088] Different gRNAs that are targeting the same gene often showed a similar trend of changes, even if they did not all reach the same level of statistical significance (Fig.4, Panel G). Notably, upper layer, deep layer excitatory neurons, and inhibitory neurons all showed changes of proportion by perturbations of Foxg1, Tbr1 and Tcf4. Amongst them, Tcf4 perturbation was associated with a 2.2-fold decrease in Sst interneuron, consistent with Tcf4’s high expression in both excitatory and inhibitory lineages in development.

[0089] As the most profound effect, Foxg1 loss-of-function was associated with a significant 9.9-fold reduction of L6 IT neurons, accompanied by a 2-fold increase of upper layer projection neurons (Fig.4, Panel G). But this perturbation did not significantly impact the proportion of L6 CT neurons which reside in the same laminar layer as L6 IT neurons. This result highlights the cell type-dependent effects of Foxg1 in governing Layer 6 neuronal cell fate, with different effects in two neuronal classes within the same layer and sharing the same developmental lineage.

[0090] Intriguingly, we observed that Tbr1 perturbation is associated with a reduction in the proportion of deep layer excitatory neurons including L6 CT (1.9-fold) and L6 IT (3.6- fold). This is consistent with the known role of Tbr1 in maintaining L6 neuronal identity as previously reported (Bedogni et al., 2010; Fazel Darbandi et al., 2018), and our analysis provides a refined annotation of its cell type-specific effects. Moreover, Tbr1 perturbation led to a 3.1-fold increase in L5 NP excitatory neurons, a distinct effect from its role in L6. Example 12. Cellular context-dependent perturbation impact on deep layer glutamatergic neurons

[0091] To further explore the impact of each perturbation at the molecular level, we performed differential expression (DE) analysis for each perturbation (gRNA) across cell types. For each cell type, cells containing the same gRNA were compared to those with a control gRNA (NT-2). We subset the data to include all the perturbation-cell type pairs with >50 cells, filtered out lowly expressed genes, and performed DE analysis with edgeR (McCarthy et al., 2012) with caution on false-positive results (Soneson and Robinson, 2018).

[0092] First, we examined the number of significant DE genes across cell type- perturbation combinations (Fig.4, Panel H). Control groups, when compared to the NT-2 control, showed very few to none (0-1) significant DE genes, as expected. Four perturbation- cell cluster combinations showed markedly altered gene expression (≥10 significant DE genes supported by at least two gRNAs): Foxg1 in L5 IT and L6 CT neurons; Nr2f1 in L6 CT neurons; and Tbr1 in L6 CT neurons. Some of these changes aligned closely with established roles described in the literature: Foxg1 in enforcing L6 CT neuronal identity (Liu et al., 2022b); Tbr1 loss-of-function causing defects in L6 CT neurons that acquire both L5- and L6-like identity and electrophysiological properties (Bedogni et al., 2010; Fazel Darbandi et al., 2018).

[0093] In late neurogenesis, Foxg1 is critical for cell fate specification across neuronal cell types through regulations of transcription factor networks to maintain L6 neuronal identity (Liu et al., 2022b). Indeed, we found Foxg1 loss-of-function affected several signaling pathways in L6 CT including expression changes of adhesion and synaptic molecules (Stxbp6, Grin3a), axon guidance molecules (Nrp1, Sema5b), and transcription factors (Bhlhe22, Lhx2) (Fig.4, Panel I). We identified upregulations of neuronal markers from nearby layers that should not be expressed in L6, including L4 spiny stellate neuron marker Rorb4 (4-fold increase) and L5 subcerebral / corticospinal projection neuron markersBcl11b and Tcerg1l (1.5-fold and 34-fold increase) in these L6 CT perturbed neurons (Fig.4, Panel I). These data suggested that, in L6 CT postmitotic neurons, Foxg1 maintains the identity by actively suppressing other transcription factors and alternative cell fates; and its loss of function de-represses this network.

[0094] Additionally, we detected a sub-cluster of L6 CT neurons, which was predominantly comprised of cells carrying Foxg1 perturbations (Fig.4, Panel B). This subcluster expresses some cell type- or layer-specific markers (Tcerg1l, Lhx2), as well as markers unique to this cluster (Nkd1), suggesting this is not simply a cluster of doublets or empty droplets and is indeed likely to be cells with hybrid fates. The proportion analysis showed that Foxg1 perturbation significantly increased the production of cluster 7, a strong effect that stood out as statistically significant across all three independent gRNAs (2-24-fold increase, P-adj<0.05). The misregulation of transcription factor expression as well as the emergence of a hybrid neuronal subcluster, altogether, provided a high-resolution characterization of the altered cell fate and their underlying misregulation of transcription factors when Foxg1 failed to maintain L6 CT identity.

[0095] Importantly, we uncovered distinct patterns of gene regulation of Foxg1 across cell types: the same perturbation induced different DE genes in different cell types and had little overlap. Amongst all the Foxg1-DE genes in the four deep layer neuron cell types: the only shared target is Stxbp6 (Syntaxin binding protein 6) whose expression is upregulated upon the perturbation in L5 PT, L5 IT, L5 PT, and L6 CT neurons (Fig.4, Panel I). However, we found that there were more distinct, than shared, regulatory networks in different deep layer excitatory neuron types. Most of the Foxg1-knockout induced changes of transcription factors are highly specific to the cell types. For example, the DE genes observed in L6 CT are largely absent in L5 IT, L5 PT or L5 NP cell types (Fig.4, Panel I). Additionally, it is known that Foxg1-Lhx2 interactions are crucial for cortical hem formation; loss of Foxg1 results in a reduction of Lhx2 expression in progenitors during early embryogenesis at E9.5 (Chou and Tole, 2019). However, in perinatal and early postnatal stage, the time window of our analysis, Lhx2 expression was substantially increased in the postmitotic L6 CT and L5 NP neurons following Foxg1 perturbation (20 and 74-fold, respectively), and not significantly changed in L5 IT or L5 PT neurons. Since our perturbation agent AAV was administered at E14.5 in this experiment, which is days after L6 neurons were born, the effect is equivalent to a conditional knockout in the postmitotic neurons, rather than knockout in the progenitors and their previously reported effects in E9.5. These discoveries revealed Foxg1’s pleotropiceffects on gene regulation that is spatiotemporally dynamic – with distinct molecular consequences in different cell types and developmental time windows.

[0096] Taken together, our analysis uncovered the cell type- and developmental stage- specific role of Foxg1 in maintaining L6 CT neuronal properties, through actively repressing alternative cell fates in postmitotic neurons. Foxg1 loss-of-function led to the de-repression of several transcription factors which ultimately leads to hybrid cell identities as well as altered capacity for synaptogenesis and circuit assembly, distinct from its other known roles in the progenitors. These analyses, with high temporal, spatial, and cell type specificity, altogether, reveal key molecular pathways through which Foxg1 orchestrates the cell fate determination and maturation of diverse neuronal cell types. Our data, altogether, demonstrated the potential and massively parallelizable capacity of the in vivo Perturb-seq platform to expand the scale and depth of genetic analysis of the highly cell type-specific regulatory networks from intact tissues. Example 13. Some exemplified experimental protocols and materials

[0097] Resources of some materials used in the studies are listed in Table 7. Some experimental models and subject details exemplified herein are described below.

[0098] C57BL / 6J, Cas9, and CD-1 mice:

[0099] All animal experiments were performed according to protocols approved by the Institutional Animal Care and Use Committees (IACUC) of The Scripps Research Institute. E15 to P9 mice of varying sex and weight were used in the scRNA-seq experiments and mice ranging from E15 to adult were used in the immunohistochemistry experiments. All mice were kept in standard conditions (a 12-h light / dark cycle with ad libitum access to food and water).

[0100] HT22 and HEK293FT cell lines:

[0101] Mammalian cell culture experiments were performed in the HT-22 mouse hippocampal neuronal cell line (Millipore Sigma, #SCC129) or HEK293FT cell line (Thermo Fisher Scientific, #R70007) grown in DMEM (Thermo Fisher Scientific, #11965092) with 25mM high glucose, 1mM sodium pyruvate and 4mM L-Glutamine (Thermo Fisher Scientific, #11995073), additionally supplemented with 1× penicillin–streptomycin (Thermo Fisher Scientific, #15140122), and 5-10% fetal bovine serum (Thermo Fisher Scientific, #16000069). HT-22 cells were maintained at confluency below 80% and HEK293FT cells were maintained at confluency below 90%.

[0102] Method details:

[0103] Mammalian cell culture and time lapse imaging: All transfections were performed with PEI (Polysciences, #24765-1) in 96-well plates unless otherwise noted. Cells were plated at approximately 10,000 cells per well 16–20 hrs before transfection to ensure 50-60% confluency at the time of transfection. For each well on the plate, transfection plasmids were combined with OptiMEM I Reduced Serum Medium (Thermo Fisher, #31985070) with PEI to a total of 20 µl. This solution was added directly to the media dropwise.

[0104] Cells were transfected and incubated in In Cell6000 Analyzer (GE Healthcare) at 37°C and 5% CO2 and imaged with 10x air objective. Images were collected every 4 hours for 20 hours. Image file names were blinded, and cell numbers were counted; fluorescence was analyzed with a custom script that is available on GitHub.

[0105] RT-qPCR: HEK293FT cells were washed once with PBS, followed by trypsinization and resuspension in PBS supplemented with 0.04% BSA (NEB, #B9000S). Cells were FACS purified at 4°C and collected in Trizol (Thermo Fisher Scientific, #10296010) with 5,000-50,000 cells per sample. RNA was extracted using Zymo Direct-zol RNA MiniPrep isolation kit (Zymo Research, #R2052). RT-qPCR was performed with Maxima H Minus Reverse Transcriptase (Thermo Fisher Scientific, #EP0753) and Power SYBR Green PCR Master Mix (Applied Biosystems, #4367659) with an equal mixture of the two RT primers: oligo (dT)18 primer (Thermo Fisher Scientific, #SO131) and direct capture primer 5’-ttgctaggaccggccttaaagc-3’ (SEQ ID NO:1). The 2-ΔΔCT method was used for the analysis of qPCR data and normalized to GAPDH expression.

[0106] AAV vector construction and production: Viral vectors and plasmids were constructed as previously reported (Jin et al., 2020). The backbone plasmid contains the human U6 promoter to express one gRNA, and the EF1α promoter to express a fluorescent protein conjugated to the nuclear membrane localized domain KASH. Cloning of the vectors was done individually and confirmed by Sanger sequencing. The gRNA designs were defined using the online tool at benchling.com and the full sequences of the gRNAs used in this work are listed in Table 6. AAV production and titration was performed by the viral vector core facility at Sanford Burnham Prebys and UCI Center for Neural Circuit Mapping viral core.

[0107] AAV administration: AAV (0.5-1.5 µL per embryo) was administered in utero to the lateral ventricle at E13.5-17.5 in CD1, C57BL / 6J or Cas9 transgenic mice (Jax#026179) (Platt et al., 2014) for immunohistochemistry analysis and scRNA-seq. Adult mice wereinjected retro-orbitally with AAV (50-100 µL and ~1-4e11viral genome per animal) and perfused for immunohistochemistry experiments and nuclei flow cytometry.

[0108] AAV library barcode extraction: At 24- or 48-hours post transduction, HT-22 cells and mouse primary cortical cells were purified with FACS. Genomic DNA was extracted from approximately 3,000 purified cells by QuickExtract DNA Extraction Solution (Lucigen, #NC0302740) following the manufacture’s protocol. The AAV serotype library was lysed by DNase I digestion and Proteinase K digestion. PCRs with genomic DNA were performed with NEBNext High-Fidelity 2X PCR Master Mix (New England BioLabs, #M0541L) with the following primers: 5’-ctttccctacacgacgctcttccgatct-gacgagtcggatctcccttt-3’ (SEQ ID NO:2), and 5’-gactggagttcagacgtgtgctcttccgatct-gcgatgcaatttcctcattt-3’ (SEQ ID NO:3). Amplicons were amplified to include adaptors and sequenced on iSeq 1000 or MiSeq platforms (> 2 million reads per sample). BCL files were converted to FASTQ files using bcl2fastq (Illumina).

[0109] Immunofluorescent staining of brain sections and whole-mount DRGs: Embryonic brains were directly harvested after decapitation and frozen immediately on dry ice in OCT. Continuous sets of 15-20μm tissue sections were prepared on a cryostat, followed by the fixation for 15 min with 4% paraformaldehyde in PBS on ice. Postnatal pups and adult mice were anesthetized and transcardially perfused with ice-cold PBS followed by ice-cold 4% paraformaldehyde in PBS. Dissected brains were postfixed overnight in 4% paraformaldehyde at 4 °C. Postnatal and adult brains were embedded in 2% agar and 60- 100μm tissue sections were collected on a vibratome.

[0110] The slides with tissue sections mounted were washed 4 times with PBS with 0.3% TritonX-100 and incubated with blocking media (10% donkey serum (Sigma Aldrich, #S30- 100ML) in 0.3% TritonX-100 with PBS) for 2 hrs at room temperature, then incubated with primary antibodies in the blocking media overnight at 4°C. Slides were washed with PBS with 0.3% TritonX-1004 times. Secondary antibodies were applied at 1:1000 dilution in blocking media and incubated for 2 hrs at room temperature. Slides were then washed 4 times with PBS with 0.3% TritonX-100 and incubated with DAPI for 10 mins before mounting with Antifade Mounting Medium (Vector Laboratories, #H-1700-10). All images were taken using a Nikon AX Confocal Microscope with a 10x air or 20x air objective.

[0111] The primary antibodies and dilutions were: Chicken anti-GFP antibody (ab16901, 1:500; Millipore), Rabbit anti-GFP antibody (A-11122, 1:500; Invitrogen), Rabbit anti-RFP (600-401-379, 1:500; Rockland), Rabbit anti-Tbr1 (ab31940, 1:500: Abcam), Rabbit anti-Tbr2 (ab183991, 1:500: Abcam), Rat anti-Ctip2 (ab18465, 1:1000, Abcam), Rabbit anti-Pax6 (Cat#901302, 1:500, BioLegend), Rabbit anti-HA tag (5017, 1:500; Cell Signaling), Rat anti- HA tag (11867423001, 1:500, Roche), Chicken anti-GFAP (ab4674, 1:500, Abcam) and Goat anti-Iba1 (ab5076, 1:500, Abcam).

[0112] Mouse dorsal root ganglia (DRGs) were extracted and post-fixed in 4% PFA for 1 hr before being washed with PBS. DRGs were mounted onto silicone isolators and mounted using EasyIndex.

[0113] Nuclei isolation and FACS-enrichment for genomic analysis: For the whole genome sequencing, >170,000 nuclei were sorted into DNA / RNA Shield (Zymo Research, Cat# R1100-250) and processed for DNA purification using a Quick DNA Microprep Kit (Zymo Research, #D3020) to isolate >500 ng genomic DNA. Genomic DNA was sequenced to 60x coverage with paired-end (150 base pair-150 base pair) Illumina sequencing (Novogene).

[0114] Tissue Dissociation and FACS-enrichment for genomic analysis: Tissue dissociation was performed with the Papain Dissociation kit (Worthington, #LK003150) in a modification of a previously described protocol (Jin et al., 2020). Briefly, young mice were anesthetized then disinfected with 70% ethanol and decapitated. The brains were quickly extracted and gently dabbed with a PBS-soaked Kimwipe (Kimberly-Clark) to remove the meninge and fibroblasts. Cortices were micro-dissected in ice-cold dissection medium (Hibernate A medium (Thermo Fisher Scientific, #A1247501) with B27 supplement (Thermo Fisher Scientific, #17504044) and Trehalose (Sigma Aldrich, Cat# T9531) under a dissecting microscope. Microdissected cortices were transferred into papain solution with DNase in a cell culture dish and cut into small pieces with a razor blade. The dish was then placed onto a digital rocker in a cell culture incubator for 30 mins with rocking speed at 30 rpm at 37°C. The digested tissues were collected into a 15 mL tube and triturated with a 10 mL low bind plastic pipette 20 times and the cell suspension was carefully transferred to a new 15 mL tube.2.7 mL of EBSS, 3 mL of reconstituted Worthington inhibitor solution, and DNase solution were added to the 15 mL tube and mixed gently. Cells were pelleted by centrifugation at 300 g for 5 mins at 4°C, followed by washing with 8ml cold dissection medium at 200g for 5 min at 4°C. Cells were resuspended in 0.5 mL ice-cold dissection medium with 10% fetal bovine serum (FBS) (Thermo Fisher Scientific, #16000069) and live / dead cell dyes, and subjected to FACS purification.

[0115] After collection, the cells were immediately centrifuged and resuspended in ice- cold PBS with 0.04% BSA (NEB, #B9000S). Each 10x scRNA-seq library was prepared by combining the FACS sorted cells from 1-2 litters (5-8 animals) of E15 or P7-9 animals harvested on the same day (Table 2). We performed the dissociation, FACS purification and resuspension within 3 hrs while keeping the cells on ice to prevent necrosis.

[0116] The scRNA-seq libraries were constructed using the Chromium Next GEM Single Cell 3' Solution v3.1 kit with Feature Barcode Technology or the Chromium Next GEM Single Cell 5’ Solution v2 kit with Feature Barcode Technology (10x Genomics) following the manufacturer’s protocol. The gene expression library was sequenced with NextSeq500 high-output 75-cycle kits (Illumina) with sequencing saturation to ensure greater than 20,000 reads coverage per cell (R1: 26 base pair, R2: 46 base pair). The CRISPR gRNA screening library was sequenced with Illumina iSeq100300-cycle (R1: 151 base pair, R2: 151 base pair) and Nextseq500 mid-output 150-cycle kits (Illumina) (R1: 73 base pair, R2: 74 base pair).

[0117] AAV barcode enrichment from scRNA-seq library: Following whole transcriptome amplification (WTA) in the 10x Chromium library construction, a fraction of the WTA product was used to amplify AAV serotype barcodes as well as cell barcodes using a dial-out PCR strategy. Briefly, 10 ng of WTA libraries were amplified for 11 cycles of PCR with AAVlib-dialout-NGS1: gactggagttcagacgtgtgctcttccgatctgacgagtcggatctcccttt (SEQ ID NO:4), and Read1-F: ctacacgacgctcttccgatct (SEQ ID NO:5), followed by a 1X SPRI bead cleanup. The sample was amplified another 24 cycles with 8 bp indexed PCR primers and gel purified. The final dial-out library was sequenced along with transcriptome library with NextSeq500 high-output 75-cycle kits (Illumina) flow cell (R1: 26 base pair, R2: 46 base pair).

[0118] Quantification and Statistical Analysis: All images were analyzed with ImageJ (NIH), Photoshop (Adobe) and Illustrator (Adobe). Cells were counted manually from blinded files using ImageJ CellCount function.

[0119] AAV barcode analysis in the primary serotype screen: FASTQ files of Illumina libraries were mapped to AAV barcodes using a custom script. Briefly, FASTQ that began with the correct initial primer sequence were kept. Of these sequences, the barcode sequence following the initial primer sequence was compared to our list of AAV barcodes and assigned to matching AAV barcodes with Levenshtein distance less than 2.

[0120] The barcode counts matrix was analyzed using DESeq2 v1.40.2 (Love et al., 2014). The DESeqDataSetFromMatrix command was used to create the DESeq object with the library counts as the reference level. Results were tabulated for each condition using alpha=0.05 as the threshold for significance. Volcano plots were produced using the R package EnhancedVolcano v1.18.0 and heatmaps with pheatmap v1.0.12.

[0121] Transposon integration site analysis: A custom reference genome was created by appending the transposon reporter plasmid to the mm39 mouse genome as an additional chromosome. FASTQ files were aligned to the custom genome using bwa mem (v0.7.17) (Li and Durbin, 2010) using the SP5M flags. The resulting files were filtered for reads that aligned to the piggybac plasmid and their pairs using SAMtools™ v1.15.1. After checking the distribution along the plasmid of these reads, they were further filtered down to reads aligning to the ends of the insert (position 1653-2053 and 6096-6496) within the plasmid. The filtered files were then parsed using Pairtools™ v0.3.0 and default settings. They were then filtered for junctions (pair_type of UU, UR, or RU with one end aligning to the mouse genome) that were then validated manually.

[0122] To annotate the integration sites, the annotatePeak function from the R package ChIPseeker™ (v1.36.0) (Yu et al., 2015) was used with the TSS set as + / - 3,000 bp. The R package Circlize v0.4.15 was used to create a Circos Plot to visualize integration sites.

[0123] scRNA-seq data processing: Starting with raw Illumina BCL files of transcriptome libraries were used to generate FASTQ files using the default parameters by “cellranger mkfastq” command (from Cell Ranger package v7.1.0) (Zheng et al., 2017). The gRNA library or AAV barcode dial-out library were demultiplexed using bcl2fastq. The “cellranger count” command was used to align the transcriptome reads to the mouse genome reference mm10 (GENCODE vM23 / Ensembl 98) and generate a gene expression count matrix, using expect-cells = 9,000. The AAV barcodes or gRNA reads were quantified at the single cell level with the feature-ref flag in Cell Ranger.

[0124] Cell type classification and cell identity annotation: For the AAV serotype secondary screen and comparing 5’ vs 3’ scRNA-seq, filtered count matrices from Cell Ranger were loaded into R v4.3.0 with the Read10X command from Seurat v4.3.0.9003 (Hao et al., 2021) and loaded into Seurat object with CreateSeuratObject, filtering out cells with <500 genes or mitochondrial count percent > 25%. The data were log normalized and the 2,000 most variable features were selected by FindVariableFeatures. In each analysis, two conditions (sort vs unsort, 5’ vs 3’) were integrated by the variable features agreed acrossdatasets by IntegrateData. The integrated data was scaled by ScaleData and PCA was performed by RunPCA. A UMAP was generated with RunUMAP and clustering was performed by FindNeighbors and FindClusters (with default parameters, except for dims = 1:25, resolution = 0.3 for the sort vs unsort, or 0.2 for 5’ vs 3’). Clusters were assigned to cell types based on the known markers (Di Bella et al., 2021; La Manno et al., 2021; Tasic et al., 2018). In the AAV serotype screen, non-cortical cell clusters were removed, while cell clusters with low gRNA mapping were removed in the 5p vs 3p analysis. The data were re- clustered after removal (dims = 1:28, resolution = 0.3 for the AAV serotype screen library or 0.2 for 5’ vs 3’). The AAV barcode or gRNA identity was assigned for cells with >5 UMI detected.

[0125] Perturb-seq data processing: cell type classification, cell identity annotation, and perturbation identity annotation: The filtered count matrices from Cell Ranger were loaded into R v4.0.3 with the Read10X command from Seurat v4.0.0 (Hao et al., 2021) and loaded into a Seurat object with CreateSeuratObject, filtering out cells with <500 genes. The data was normalized with the NormalizeData command, with normalization.method="LogNormalize" and scale.factor=10,000 and variable genes were selected with FindVariableFeatures. Doublet scores were calculated with scds v1.6.0 (Bais and Kostka, 2020) separately for each 10x channel. The UMI count data for the variable genes were extracted and used as input to scGBM v0.1.0 (Phillip and Jeffrey, 2023) with M=20 and subset=30000, resulting in a 20-dimensional embedding of the count data that was added to the Seurat object. A UMAP was calculated on this reduction using RunUMAP, and clustering was performed on this reduction with FindNeighbors and FindClusters with otherwise default settings. The UMI count matrix for the gRNAs produced by Cell Ranger was also added to this Seurat object as an additional assay.

[0126] Quality control (QC) metrics for each channel were calculated with CellLevel_QC tool (Github) and loaded into the metadata for the Seurat object. The % mitochondrial reads was also calculated for each cell. Azimuth v0.3.2 was used to produce an initial annotation of the data using a single cell reference from the Allen brain atlas (Yao et al., 2021). Clusters with high percent intronic reads or high doublet scores were removed, as were cells with >20% intronic reads or >10% mitochondrial reads. Clusters were then labeled with cell type using the cell type labels from Azimuth and by comparing DE genes from our dataset to DE genes from the Allen brain atlas dataset (DE genes between clusters were calculated with presto v1.0.0).

[0127] We next built a Nextflow pipeline (DSL v2) (Di Tommaso et al., 2017) that took the Cell Ranger output, extracted the guide RNA data into a csv file, assigned cells to guides with DemuxEM (v0.1.5) (Gaublomme et al., 2019), and extracted the resulting perturbation labels in the form of a tsv file (where the creation of the tsv and csv was performed using PegasusIO v0.2.11 (Li et al., 2020) for loading the data and pandas v1.2.4 for saving the data (pandas-development-team, 2023)). These tables were then loaded into the Seurat objects metadata. We also built a version of the pipeline that allowed for down sampling before running DemuxEM (the down sampling approached mentioned in the main text). In particular, for each 10x channel and each gRNA, if there were more than 50,000 UMIs assigned to that gRNA, the pipeline down-sampled the gRNA UMI counts for that gRNA using the binomial function from numpy.random (Harris et al., 2020), where the proportion used for down sampling was equal to 50,000 divided by the total number of UMI coming from the gRNA (resulting in an expected value of 50,000 UMIs per gRNA per channel after the down sampling). Similarly, we built a Nextflow pipeline to extract reads mapping to the BFP or GFP sequence. The pipeline takes in the BAM file produced by Cell Ranger, extracts unmapped reads with SAMtools v1.8 (using the command “samtools view -f 4 -b”) (Li et al., 2009), extracts UMIs and CBCs with UMI-tools v1.0.1 (Smith et al., 2017), maps these reads to the sequences for GFP and BFP (including 5’ and 3’ UTR regions) with Minimap2 v2.11 (using the arguments -ax sr) (Li, 2018), and transforms them into a BAM file with SAM-tools view -b. Unmapped reads were discarded and the name of the contig mapped to (GFP or BFP) for the remaining reads mapped to was added to the BAM file as an additional tag (XT tag) with awk. The resulting BAM file was sorted, indexed, and the number of UMIs mapping to GFP and BFP for each cell was extracted with command umi_tools count --per- gene --gene-tag=XT --per-cell. These counts were also loaded into the Seurat object.

[0128] In order to accurately assign the perturbation identities to cells, we compared 3 methods: the built-in gRNA assignments from Cell Ranger, the tool for demultiplexing the hashing barcodes known as DemuxEM (Gaublomme et al., 2019), and down sampling gRNA counts followed by DemuxEM. We found that all 3 methods performed well, though we saw evidence showing that the standard DemuxEM approach led to biases in cell type proportion due to large differences in the number of gRNA UMIs recovered per guide. For example, we observed a strong correlation between the average number of gRNA UMIs recovered from a given gRNA and the proportion of Cajal Retzius cells among cells assigned to that gRNA (Spearman correlation=0.84, P-value=2.934e-05); this effect disappeared after the downsampling (Spearman correlation=-0.03, P-value=0.89). Cell Ranger did show evidence of this bias but did report a larger than expected number of doublets, particularly in the Cajal Retzius cluster (with 71% of high-quality Cajal Retzius cells being assigned as doublets in Cell Ranger, versus 60% in the down sampling approach and 52% in the standard DemuxEM approach). We used the DemuxEM with down sampling for the full analysis.

[0129] Non-neuronal cells (cells not labeled as Excitatory, Inhibitory, or Cajal-Retzius neurons) were removed from the Seurat object, as were cells with <3% intronic reads. Excitatory and Inhibitory cells with <3000 genes were filtered out, as were Cajal-Retzius cells with <2,000 genes. Finally, all cells that were not assigned to exactly one guide by DemuxEM were removed as well. This Seurat object was used for downstream analysis.

[0130] Detecting insertions and deletions in targeted regions: For each targeted gene, the reads overlapping it in each 10x channel were extracted from the Cell Ranger BAM files with SAM-tools view and combined into one BAM file with SAM-tools concat, which was sorted and indexed with SAM-tools. The sinto filterbarcode command (Stuart et al., 2021) was used to split the BAM file into one BAM file for each perturbation, consisting of cells in the final analysis assigned to each gRNA (excluding cell barcodes that occur in multiple 10x channels). The resulting BAM files were indexed. This resulted in one BAM file for each targeted gene and each gRNA, consisting of reads in cells assigned to that gRNA / perturbation identity that overlap the targeted gene.

[0131] For each targeted region, we counted the number of reads overlapping that region for each perturbation, as well as the number of reads with an indel (insertion or deletion) in that region. This was performed with a Pysam (Github) based python script. The script loaded each read one by one from the BAM file. For each position in each read it used get_reference_positions to: 1) check if that position was in the range required 2) if it was in the range, checked if that location was an insertion (excluding insertions at the beginning / end of reads to avoid soft clipped regions) 3) if it was in the range and did not contain an insertion, checked if that location was a deletion (if there was a difference of >1 bp between the mapping position of the current base and the previous base). The results were then tabulated into a pandas data frame and returned before being saved then loaded into R for downstream analysis.

[0132] Perturbation-associated analysis: Having assigned a gRNA and a cell type to all the cells, we looked at cell type proportion changes across gRNAs. To do so, we selected a control group (Non-Targeting control 2) and compared each of the perturbation / controlgroups containing other gRNAs to this group. Statistics for these pairwise composition comparisons were computed using the propeller.ttest function from speckle (R package v0.99.7) (Phipson et al., 2022). Here, the batch (10x channel) was additionally considered as another fixed effect to the linear model. Cell types and clusters with less than 200 cells overall (Excit_Car3 and Inhib_Id2 for the cell type-level comparisons and clusters 5, 25 and 26 for the cluster-level comparisons) were excluded from this analysis. Results are collected and visualized together using ComplexHeatmap (R package v2.14.0) (Gu et al., 2016). For this visualization, effect sizes were capped at the absolute value of 2.

[0133] We next looked for genes that were differentially expressed in each gRNA and within each cell type. Using control (gRNA Non-Targeting control 2), we compared each gRNA group population to the control. Statistics were computed with a pseudo-bulk approach using edgeR (R package v3.40.2) (McCarthy et al., 2012). Lowly expressed genes (<10 supporting reads and <.1 normalized expression) were identified within each comparison and excluded from the DE test. Additionally, we applied surrogate variable analysis (sva; R package v3.46.0) (Leek and Storey, 2007) to capture potential batch- or technical artifact-related variable from the data and add it to the model testing for DE (n.sv=1). Computed P-values were adjusted using the Benjamini-Hochberg method. Table 7. List of resources of some materials used in the studies_

[0134] Some cited references: Arlotta, P., Molyneaux, B.J., Chen, J., Inoue, J., Kominami, R., and Macklis, J.D. (2005). Neuronal subtype-specific genes that control corticospinal motor neuron development in vivo. Neuron 45, 207-221.10.1016 / j.neuron.2004.12.036. Bais, A.S., and Kostka, D. (2020). scds: computational annotation of doublets in single-cell RNA sequencing data. Bioinformatics 36, 1150-1158.10.1093 / bioinformatics / btz698. Bedogni, F., Hodge, R.D., Elsen, G.E., Nelson, B.R., Daza, R.A., Beyer, R.P., Bammler, T.K., Rubenstein, J.L., and Hevner, R.F. (2010). Tbr1 regulates regional and laminar identity of postmitotic neurons in developing neocortex. Proc Natl Acad Sci U S A 107, 13129- 13134.10.1073 / pnas.1002285107. Bock, C., Datlinger, P., Chardon, F., Coelho, M.A., Dong, M.B., Lawson, K.A., Lu, T., Maroc, L., Norman, T.M., Song, B., et al. (2022). High-content CRISPR screening. Nat Rev Methods Primers 2.10.1038 / s43586-022-00098-7. Boulting, G.L., Durresi, E., Ataman, B., Sherman, M.A., Mei, K., Harmin, D.A., Carter, A.C., Hochbaum, D.R., Granger, A.J., Engreitz, J.M., et al. (2021). Activity-dependent regulome of human GABAergic neurons reveals new patterns of gene regulation and neurological disease heritability. Nat Neurosci 24, 437-448.10.1038 / s41593-020-00786-1. Chan, K.Y., Jang, M.J., Yoo, B.B., Greenbaum, A., Ravi, N., Wu, W.L., Sanchez-Guardado, L., Lois, C., Mazmanian, S.K., Deverman, B.E., and Gradinaru, V. (2017). Engineered AAVs for efficient noninvasive gene delivery to the central and peripheral nervous systems. Nat Neurosci 20, 1172-1179.10.1038 / nn.4593. Chen, B., Wang, S.S., Hattox, A.M., Rayburn, H., Nelson, S.B., and McConnell, S.K. (2008). The Fezf2-Ctip2 genetic pathway regulates the fate choice of subcortical projection neurons in the developing cerebral cortex. Proc Natl Acad Sci U S A 105, 11382-11387. 10.1073 / pnas.0804918105. Chen, H.Y., Bohlen, J.F., and Maher, B.J. (2021). Molecular and Cellular Function of Transcription Factor 4 in Pitt-Hopkins Syndrome. Dev Neurosci 43, 159-167. 10.1159 / 000516666. Chen, L.F., Lin, Y.T., Gallegos, D.A., Hazlett, M.F., Gomez-Schiavon, M., Yang, M.G., Kalmeta, B., Zhou, A.S., Holtzman, L., Gersbach, C.A., et al. (2019). Enhancer Histone Acetylation Modulates Transcriptional Bursting Dynamics of Neuronal Activity-Inducible Genes. Cell Rep 26, 1174-1188 e1175.10.1016 / j.celrep.2019.01.032.Chen, Q., Luo, W., Veach, R.A., Hickman, A.B., Wilson, M.H., and Dyda, F. (2020). Structural basis of seamless excision and specific targeting by piggyBac transposase. Nat Commun 11, 3446.10.1038 / s41467-020-17128-1. Chen, S., Sanjana, N.E., Zheng, K., Shalem, O., Lee, K., Shi, X., Scott, D.A., Song, J., Pan, J.Q., Weissleder, R., et al. (2015). Genome-wide CRISPR screen in a mouse model of tumor growth and metastasis. Cell 160, 1246-1260.10.1016 / j.cell.2015.02.038. Chen, X., Ravindra Kumar, S., Adams, C.D., Yang, D., Wang, T., Wolfe, D.A., Arokiaraj, C.M., Ngo, V., Campos, L.J., Griffiths, J.A., et al. (2022). Engineered AAVs for non- invasive gene delivery to rodent and non-human primate nervous systems. Neuron 110, 2242- 2257 e2246.10.1016 / j.neuron.2022.05.003. Cheroni, C., Caporale, N., and Testa, G. (2020). Autism spectrum disorder at the crossroad between genes and environment: contributions, convergences, and interactions in ASD developmental pathophysiology. Mol Autism 11, 69.10.1186 / s13229-020-00370-1. Chong, L., Jonas Simon, F., Catarina, M.-C., Thomas, R.B., Marlene, S., Ábel, V., Angela Maria, P., Christopher, E., Ulrich, E., Gregor, K., et al. (2022). Single-cell brain organoid screening identifies developmental defects in autism. bioRxiv, 2022.2009.2015.508118. 10.1101 / 2022.09.15.508118. Chou, S.J., and Tole, S. (2019). Lhx2, an evolutionarily conserved, multifunctional regulator of forebrain development. Brain Res 1705, 1-14.10.1016 / j.brainres.2018.02.046. Chow, R.D., Guzman, C.D., Wang, G., Schmidt, F., Youngblood, M.W., Ye, L., Errami, Y., Dong, M.B., Martinez, M.A., Zhang, S., et al. (2017). AAV-mediated direct in vivo CRISPR screen identifies functional suppressors in glioblastoma. Nat Neurosci 20, 1329-1341. 10.1038 / nn.4620. Chuapoco, M.R., Flytzanis, N.C., Goeden, N., Christopher Octeau, J., Roxas, K.M., Chan, K.Y., Scherrer, J., Winchester, J., Blackburn, R.J., Campos, L.J., et al. (2023). Adeno- associated viral vectors for functional intravenous gene transfer throughout the non-human primate brain. Nat Nanotechnol.10.1038 / s41565-023-01419-x. Deyle, D.R., and Russell, D.W. (2009). Adeno-associated virus vector integration. Curr Opin Mol Ther 11, 442-447. Dhainaut, M., Rose, S.A., Akturk, G., Wroblewska, A., Nielsen, S.R., Park, E.S., Buckup, M., Roudko, V., Pia, L., Sweeney, R., et al. (2022). Spatial CRISPR genomics identifies regulators of the tumor microenvironment. Cell 185, 1223-1239 e1220. 10.1016 / j.cell.2022.02.015.Di Bella, D.J., Habibi, E., Stickels, R.R., Scalia, G., Brown, J., Yadollahpour, P., Yang, S.M., Abbate, C., Biancalani, T., Macosko, E.Z., et al. (2021). Molecular logic of cellular diversification in the mouse cerebral cortex. Nature 595, 554-559.10.1038 / s41586-021- 03670-5. Di Tommaso, P., Chatzou, M., Floden, E.W., Barja, P.P., Palumbo, E., and Notredame, C. (2017). Nextflow enables reproducible computational workflows. Nat Biotechnol 35, 316- 319.10.1038 / nbt.3820. Dvoretskova, E., May, C.H., Volker, K., Ilaria, V., Daniel, D.L., Irene, D., Chao, F., Miguel, T., Juliane, W., and Christian, M. (2023). Spatial enhancer activation determines inhibitory neuron identity. bioRxiv, 2023.2001.2030.525356.10.1101 / 2023.01.30.525356. Fazel Darbandi, S., Robinson Schwartz, S.E., Qi, Q., Catta-Preta, R., Pai, E.L., Mandell, J.D., Everitt, A., Rubin, A., Krasnoff, R.A., Katzman, S., et al. (2018). Neonatal Tbr1 Dosage Controls Cortical Layer 6 Connectivity. Neuron 100, 831-845 e837. 10.1016 / j.neuron.2018.09.027. Feldman, D., Singh, A., Schmid-Burgk, J.L., Carlson, R.J., Mezger, A., Garrity, A.J., Zhang, F., and Blainey, P.C. (2019). Optical Pooled Screens in Human Cells. Cell 179, 787-799 e717.10.1016 / j.cell.2019.09.016. Fleck, J.S., Jansen, S.M.J., Wollny, D., Zenk, F., Seimiya, M., Jain, A., Okamoto, R., Santel, M., He, Z., Camp, J.G., and Treutlein, B. (2022). Inferring and perturbing cell fate regulomes in human brain organoids. Nature.10.1038 / s41586-022-05279-8. Fu, J.M., Satterstrom, F.K., Peng, M., Brand, H., Collins, R.L., Dong, S., Wamsley, B., Klei, L., Wang, L., Hao, S.P., et al. (2022). Rare coding variation provides insight into the genetic architecture and phenotypic context of autism. Nat Genet.10.1038 / s41588-022-01104-0. Gaublomme, J.T., Li, B., McCabe, C., Knecht, A., Yang, Y., Drokhlyansky, E., Van Wittenberghe, N., Waldman, J., Dionne, D., Nguyen, L., et al. (2019). Nuclei multiplexing with barcoded antibodies for single-nucleus genomics. Nat Commun 10, 2907. 10.1038 / s41467-019-10756-2. Gilbert, L.A., Horlbeck, M.A., Adamson, B., Villalta, J.E., Chen, Y., Whitehead, E.H., Guimaraes, C., Panning, B., Ploegh, H.L., Bassik, M.C., et al. (2014). Genome-Scale CRISPR-Mediated Control of Gene Repression and Activation. Cell 159, 647-661. 10.1016 / j.cell.2014.09.029.Greig, L.C., Woodworth, M.B., Galazo, M.J., Padmanabhan, H., and Macklis, J.D. (2013). Molecular logic of neocortical projection neuron specification, development and diversity. Nat Rev Neurosci 14, 755-769.10.1038 / nrn3586. Gu, Z., Eils, R., and Schlesner, M. (2016). Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinformatics 32, 2847-2849. 10.1093 / bioinformatics / btw313. Hao, Y., Hao, S., Andersen-Nissen, E., Mauck, W.M., 3rd, Zheng, S., Butler, A., Lee, M.J., Wilk, A.J., Darby, C., Zager, M., et al. (2021). Integrated analysis of multimodal single-cell data. Cell 184, 3573-3587 e3529.10.1016 / j.cell.2021.04.048. Harris, C.R., Millman, K.J., van der Walt, S.J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N.J., et al. (2020). Array programming with NumPy. Nature 585, 357-362.10.1038 / s41586-020-2649-2. Hayashi, H., Kubo, Y., Izumida, M., and Matsuyama, T. (2020). Efficient viral delivery of Cas9 into human safe harbor. Sci Rep 10, 21474.10.1038 / s41598-020-78450-8. Hendel, A., Bak, R.O., Clark, J.T., Kennedy, A.B., Ryan, D.E., Roy, S., Steinfeld, I., Lunstad, B.D., Kaiser, R.J., Wilkens, A.B., et al. (2015). Chemically modified guide RNAs enhance CRISPR-Cas genome editing in human primary cells. Nat Biotechnol 33, 985-989. 10.1038 / nbt.3290. Higashikawa, F., and Chang, L. (2001). Kinetic analyses of stability of simple and complex retroviral vectors. Virology 280, 124-131.10.1006 / viro.2000.0743. Hou, P.S., hAilin, D.O., Vogel, T., and Hanashima, C. (2020). Transcription and Beyond: Delineating FOXG1 Function in Cortical Development and Disorders. Front Cell Neurosci 14, 35.10.3389 / fncel.2020.00035. Jang, M.J., Coughlin, G.M., Jackson, C.R., Chen, X., Chuapoco, M.R., Vendemiatti, J.L., Wang, A.Z., and Gradinaru, V. (2023). Spatial transcriptomics for profiling the tropism of viral vectors in tissues. Nat Biotechnol.10.1038 / s41587-022-01648-w. Jin, X., Simmons, S.K., Guo, A., Shetty, A.S., Ko, M., Nguyen, L., Jokhi, V., Robinson, E., Oyler, P., Curry, N., et al. (2020). In vivo Perturb-Seq reveals neuronal and glial abnormalities associated with autism risk genes. Science 370.10.1126 / science.aaz6063. Johnston, S., Parylak, S.L., Kim, S., Mac, N., Lim, C., Gallina, I., Bloyd, C., Newberry, A., Saavedra, C.D., Novak, O., et al. (2021). AAV ablates neurogenesis in the adult murine hippocampus. Elife 10.10.7554 / eLife.59291.Joung, J., Ma, S., Tay, T., Geiger-Schuller, K.R., Kirchgatterer, P.C., Verdine, V.K., Guo, B., Arias-Garcia, M.A., Allen, W.E., Singh, A., et al. (2023). A transcription factor atlas of directed differentiation. Cell 186, 209-229 e226.10.1016 / j.cell.2022.11.026. Kalamakis, G., and Platt, R.J. (2023). CRISPR for neuroscientists. Neuron. 10.1016 / j.neuron.2023.04.021. Kaplanis, J., Samocha, K.E., Wiel, L., Zhang, Z., Arvai, K.J., Eberhardt, R.Y., Gallone, G., Lelieveld, S.H., Martin, H.C., McRae, J.F., et al. (2020). Evidence for 28 genetic disorders discovered by combining healthcare and research data. Nature 586, 757-762. 10.1038 / s41586-020-2832-5. Keys, H.R., and Knouse, K.A. (2022). Genome-scale CRISPR screening in a single mouse liver. Cell Genom 2.10.1016 / j.xgen.2022.100217. Klingler, E., Francis, F., Jabaudon, D., and Cappello, S. (2021). Mapping the molecular and cellular complexity of cortical malformations. Science 371.10.1126 / science.aba4517. Kuzmin, D.A., Shutova, M.V., Johnston, N.R., Smith, O.P., Fedorin, V.V., Kukushkin, Y.S., van der Loo, J.C.M., and Johnstone, E.C. (2021). The clinical landscape for AAV gene therapies. Nat Rev Drug Discov 20, 173-174.10.1038 / d41573-021-00017-7. La Manno, G., Siletti, K., Furlan, A., Gyllborg, D., Vinsland, E., Mossi Albiach, A., Mattsson Langseth, C., Khven, I., Lederer, A.R., Dratva, L.M., et al. (2021). Molecular architecture of the developing mouse brain. Nature 596, 92-96.10.1038 / s41586-021-03775-x. Lang, J.F., Toulmin, S.A., Brida, K.L., Eisenlohr, L.C., and Davidson, B.L. (2019). Standard screening methods underreport AAV-mediated transduction and gene editing. Nat Commun 10, 3415.10.1038 / s41467-019-11321-7. Leek, J.T., and Storey, J.D. (2007). Capturing heterogeneity in gene expression studies by surrogate variable analysis. PLoS Genet 3, 1724-1735.10.1371 / journal.pgen.0030161. Li, B., Gould, J., Yang, Y., Sarkizova, S., Tabaka, M., Ashenberg, O., Rosen, Y., Slyper, M., Kowalczyk, M.S., Villani, A.C., et al. (2020). Cumulus provides cloud-based data analysis for large-scale single-cell and single-nucleus RNA-seq. Nat Methods 17, 793-798. 10.1038 / s41592-020-0905-x. Li, H. (2018). Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094-3100.10.1093 / bioinformatics / bty191. Li, H., and Durbin, R. (2010). Fast and accurate long-read alignment with Burrows-Wheeler transform. Bioinformatics 26, 589-595.10.1093 / bioinformatics / btp698.Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N., Marth, G., Abecasis, G., Durbin, R., and Genome Project Data Processing, S. (2009). The Sequence Alignment / Map format and SAMtools. Bioinformatics 25, 2078-2079. 10.1093 / bioinformatics / btp352. Liu, G., Lin, Q., Jin, S., and Gao, C. (2022a). The CRISPR-Cas toolbox and gene editing technologies. Mol Cell 82, 333-347.10.1016 / j.molcel.2021.12.002. Liu, J., Yang, M., Su, M., Liu, B., Zhou, K., Sun, C., Ba, R., Yu, B., Zhang, B., Zhang, Z., et al. (2022b). FOXG1 sequentially orchestrates subtype specification of postmitotic cortical projection neurons. Sci Adv 8, eabh3568.10.1126 / sciadv.abh3568. Love, M.I., Huber, W., and Anders, S. (2014). Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 15, 550.10.1186 / s13059-014-0550- 8. McCarthy, D.J., Chen, Y., and Smyth, G.K. (2012). Differential expression analysis of multifactor RNA-Seq experiments with respect to biological variation. Nucleic Acids Res 40, 4288-4297.10.1093 / nar / gks042. Mimitou, E.P., Cheng, A., Montalbano, A., Hao, S., Stoeckius, M., Legut, M., Roush, T., Herrera, A., Papalexi, E., Ouyang, Z., et al. (2019). Multiplexed detection of proteins, transcriptomes, clonotypes and CRISPR perturbations in single cells. Nat Methods 16, 409- 412.10.1038 / s41592-019-0392-0. Morgens, D.W., Wainberg, M., Boyle, E.A., Ursu, O., Araya, C.L., Tsui, C.K., Haney, M.S., Hess, G.T., Han, K., Jeng, E.E., et al. (2017). Genome-scale measurement of off-target activity using Cas9 toxicity in high-throughput screens. Nat Commun 8, 15178. 10.1038 / ncomms15178. Moudgil, A., Wilkinson, M.N., Chen, X., He, J., Cammack, A.J., Vasek, M.J., Lagunas, T., Jr., Qi, Z., Lalli, M.A., Guo, C., et al. (2020). Self-Reporting Transposons Enable Simultaneous Readout of Gene Expression and Transcription Factor Binding in Single Cells. Cell 182, 992-1008 e1021.10.1016 / j.cell.2020.06.037. Nam, A.S., Kim, K.T., Chaligne, R., Izzo, F., Ang, C., Taylor, J., Myers, R.M., Abu-Zeinah, G., Brand, R., Omans, N.D., et al. (2019). Somatic mutations and cell identity linked by Genotyping of Transcriptomes. Nature 571, 355-360.10.1038 / s41586-019-1367-0. Nunez, J.K., Chen, J., Pommier, G.C., Cogan, J.Z., Replogle, J.M., Adriaens, C., Ramadoss, G.N., Shi, Q., Hung, K.L., Samelson, A.J., et al. (2021). Genome-wide programmabletranscriptional memory by CRISPR-based epigenome editing. Cell 184, 2503-2519 e2517. 10.1016 / j.cell.2021.03.025. Ojala, D.S., Sun, S., Santiago-Ortiz, J.L., Shapiro, M.G., Romero, P.A., and Schaffer, D.V. (2018). In Vivo Selection of a Computationally Designed SCHEMA AAV Library Yields a Novel Variant for Infection of Adult Neural Stem Cells in the SVZ. Mol Ther 26, 304-319. 10.1016 / j.ymthe.2017.09.006. Pablo Perez-Pinera, D.G.O., Matthew T. Brown, David Yudovich, Charles A. Gersbach (2011). Targeted Gene Integration into a Mouse Safe Harbor Locus Directed by Engineered Zinc Finger Nucleases. Molecular Therapy 19, S163. https: / / doi.org / 10.1016 / S1525- 0016(16)36994-5. pandas-development-team (2023). pandas-dev / pandas: Pandas (v2.1.0rc0). Zenodo (doi.org / 10.5281 / zenodo.8239932). Perez, B.A., Shutterly, A., Chan, Y.K., Byrne, B.J., and Corti, M. (2020). Management of Neuroinflammatory Responses to AAV-Mediated Gene Therapies for Neurodegenerative Diseases. Brain Sci 10.10.3390 / brainsci10020119. Phillip, B.N., and Jeffrey, W.M. (2023). Model-based dimensionality reduction for single-cell RNA-seq using generalized bilinear models. bioRxiv, 2023.2004.2021.537881. 10.1101 / 2023.04.21.537881. Phipson, B., Sim, C.B., Porrello, E.R., Hewitt, A.W., Powell, J., and Oshlack, A. (2022). propeller: testing for differences in cell type proportions in single cell data. Bioinformatics 38, 4720-4726.10.1093 / bioinformatics / btac582. Platt, R.J., Chen, S., Zhou, Y., Yim, M.J., Swiech, L., Kempton, H.R., Dahlman, J.E., Parnas, O., Eisenhaure, T.M., Jovanovic, M., et al. (2014). CRISPR-Cas9 knockin mice for genome editing and cancer modeling. Cell 159, 440-455.10.1016 / j.cell.2014.09.014. Przybyla, L., and Gilbert, L.A. (2022). A new era in functional genomics screens. Nat Rev Genet 23, 89-103.10.1038 / s41576-021-00409-w. Querques, I., Mades, A., Zuliani, C., Miskey, C., Alb, M., Grueso, E., Machwirth, M., Rausch, T., Einsele, H., Ivics, Z., et al. (2019). A highly soluble Sleeping Beauty transposase improves control of gene insertion. Nat Biotechnol 37, 1502-1512.10.1038 / s41587-019- 0291-z. Replogle, J.M., Norman, T.M., Xu, A., Hussmann, J.A., Chen, J., Cogan, J.Z., Meer, E.J., Terry, J.M., Riordan, D.P., Srinivas, N., et al. (2020). Combinatorial single-cell CRISPRscreens by direct guide RNA capture and targeted sequencing. Nat Biotechnol 38, 954-961. 10.1038 / s41587-020-0470-y. Rubin, A.J., Parker, K.R., Satpathy, A.T., Qi, Y., Wu, B., Ong, A.J., Mumbach, M.R., Ji, A.L., Kim, D.S., Cho, S.W., et al. (2019). Coupled Single-Cell CRISPR Screening and Epigenomic Profiling Reveals Causal Gene Regulatory Networks. Cell 176, 361-376 e317. 10.1016 / j.cell.2018.11.022. Samulski, R.J., and Muzyczka, N. (2014). AAV-Mediated Gene Therapy for Research and Therapeutic Purposes. Annu Rev Virol 1, 427-451.10.1146 / annurev-virology-031413- 085355. Sanchez-Priego, C., Hu, R., Boshans, L.L., Lalli, M., Janas, J.A., Williams, S.E., Dong, Z., and Yang, N. (2022). Mapping cis-regulatory elements in human neurons links psychiatric disease heritability and activity-regulated transcriptional programs. Cell Rep 39, 110877. 10.1016 / j.celrep.2022.110877. Schraivogel, D., Steinmetz, L.M., and Parts, L. (2023). Pooled Genome-Scale CRISPR Screens in Single Cells. Annu Rev Genet.10.1146 / annurev-genet-072920-013842. Smith, T., Heger, A., and Sudbery, I. (2017). UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy. Genome Res 27, 491-499. 10.1101 / gr.209601.116. Soneson, C., and Robinson, M.D. (2018). Bias, robustness and scalability in single-cell differential expression analysis. Nat Methods 15, 255-261.10.1038 / nmeth.4612. Stuart, T., Srivastava, A., Madad, S., Lareau, C.A., and Satija, R. (2021). Single-cell chromatin state analysis with Signac. Nat Methods 18, 1333-1341.10.1038 / s41592-021- 01282-5. Tabebordbar, M., Lagerborg, K.A., Stanton, A., King, E.M., Ye, S., Tellez, L., Krunnfusz, A., Tavakoli, S., Widrick, J.J., Messemer, K.A., et al. (2021). Directed evolution of a family of AAV capsid variants enabling potent muscle-directed gene delivery across species. Cell 184, 4919-4938 e4922.10.1016 / j.cell.2021.08.028. Tasic, B., Yao, Z., Graybuck, L.T., Smith, K.A., Nguyen, T.N., Bertagnolli, D., Goldy, J., Garren, E., Economo, M.N., Viswanathan, S., et al. (2018). Shared and distinct transcriptomic cell types across neocortical areas. Nature 563, 72-78.10.1038 / s41586-018- 0654-5. Tian, F., Cheng, Y., Zhou, S., Wang, Q., Monavarfeshani, A., Gao, K., Jiang, W., Kawaguchi, R., Wang, Q., Tang, M., et al. (2023). Core transcription programs controllinginjury-induced neurodegeneration of retinal ganglion cells. Neuron 111, 444. 10.1016 / j.neuron.2023.01.008. Townsley, K.G., Brennand, K.J., and Huckins, L.M. (2020). Massively parallel techniques for cataloguing the regulome of the human brain. Nat Neurosci 23, 1509-1521. 10.1038 / s41593-020-00740-1. Tyson, J.R., Chloe, M.K., Bhek, M., Robin, W.Y., Dena, S.L., David, W.M., Tsui, C.K., Amy, L., Michael, C.B., and Anne, B. (2021). In vitro and in vivo CRISPR-Cas9 screens reveal drivers of aging in neural stem cells of the brain. bioRxiv, 2021.2011.2023.469762. 10.1101 / 2021.11.23.469762. Wang, G., Chow, R.D., Ye, L., Guzman, C.D., Dai, X., Dong, M.B., Zhang, F., Sharp, P.A., Platt, R.J., and Chen, S. (2018). Mapping a functional cancer genome atlas of tumor suppressors in mouse liver using AAV-CRISPR-mediated direct in vivo screening. Sci Adv 4, eaao5508.10.1126 / sciadv.aao5508. Wertz, M.H., Mitchem, M.R., Pineda, S.S., Hachigian, L.J., Lee, H., Lau, V., Powers, A., Kulicke, R., Madan, G.K., Colic, M., et al. (2020). Genome-wide In Vivo CNS Screening Identifies Genes that Modify CNS Neuronal Survival and mHTT Toxicity. Neuron 106, 76- 89 e78.10.1016 / j.neuron.2020.01.004. West, A.E., and Greenberg, M.E. (2011). Neuronal activity-regulated gene transcription in synapse development and cognitive function. Cold Spring Harb Perspect Biol 3. 10.1101 / cshperspect.a005744. Wong, F.K., Bercsenyi, K., Sreenivasan, V., Portales, A., Fernandez-Otero, M., and Marin, O. (2018). Pyramidal cell regulation of interneuron survival sculpts cortical networks. Nature 557, 668-673.10.1038 / s41586-018-0139-6. Xie, S., Cooley, A., Armendariz, D., Zhou, P., and Hon, G.C. (2018). Frequent sgRNA- barcode recombination in single-cell perturbation assays. PLoS One 13, e0198635. 10.1371 / journal.pone.0198635. Xiong, W., Wu, D.M., Xue, Y., Wang, S.K., Chung, M.J., Ji, X., Rana, P., Zhao, S.R., Mai, S., and Cepko, C.L. (2019). AAV cis-regulatory sequences are correlated with ocular toxicity. Proc Natl Acad Sci U S A 116, 5785-5794.10.1073 / pnas.1821000116. Yao, Z., Liu, H., Xie, F., Fischer, S., Adkins, R.S., Aldridge, A.I., Ament, S.A., Bartlett, A., Behrens, M.M., Van den Berge, K., et al. (2021). A transcriptomic and epigenomic cell atlas of the mouse primary motor cortex. Nature 598, 103-110.10.1038 / s41586-021-03500-8.Ye, L., Lam, S.Z., Yang, L., Suzuki, K., Zou, Y., Lin, Q., Zhang, Y., Clark, P., Peng, L., and Chen, S. (2023). Therapeutic immune cell engineering with an mRNA : AAV- Sleeping Beauty composite system. bioRxiv.10.1101 / 2023.03.14.532651. Ye, L., Park, J.J., Dong, M.B., Yang, Q., Chow, R.D., Peng, L., Du, Y., Guo, J., Dai, X., Wang, G., et al. (2019). In vivo CRISPR screening in CD8 T cells with AAV-Sleeping Beauty hybrid vectors identifies membrane targets for improving immunotherapy for glioblastoma. Nat Biotechnol 37, 1302-1313.10.1038 / s41587-019-0246-4. Yim, Y.Y., Teague, C.D., and Nestler, E.J. (2020). In vivo locus-specific editing of the neuroepigenome. Nat Rev Neurosci 21, 471-484.10.1038 / s41583-020-0334-y. Yu, G., Wang, L.G., and He, Q.Y. (2015). ChIPseeker: an R / Bioconductor package for ChIP peak annotation, comparison and visualization. Bioinformatics 31, 2382-2383. 10.1093 / bioinformatics / btv145. Zheng, G.X., Terry, J.M., Belgrader, P., Ryvkin, P., Bent, Z.W., Wilson, R., Ziraldo, S.B., Wheeler, T.D., McDermott, G.P., Zhu, J., et al. (2017). Massively parallel digital transcriptional profiling of single cells. Nat Commun 8, 14049.10.1038 / ncomms14049. ***

[0135] Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be readily apparent to one of ordinary skill in the art in light of the teachings of this invention that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.

[0136] All publications, databases, GenBank sequences, patents, and patent applications cited in this specification are herein incorporated by reference as if each was specifically and individually indicated to be incorporated by reference.

Claims

WHAT IS CLAIMED IS:

1. A method for analyzing functions of a plurality of target genes in one or more specific cell types in vivo, comprising (1) introducing a library of AAV vectors encoding genetic perturbations for the plurality of target genes into a CRISPR-Cas9 expressing transgenic system, (2) identifying one or more cells of a specific cell type from the transgenic system that express a genetic perturbation and display a specific phenotype, and (3) determining (a) the genetic perturbation encoded by the AAV vector introduced into each of the one or more cells and (b) the corresponding perturbed gene, thereby correlating the corresponding perturbed gene with the phenotype displayed by each of the one or more cells.

2. The method of claim 1, wherein the transgenic system is a developing embryo of a Cas9-expressing transgenic animal, and wherein the library of AAV vectors comprise vectors of serotype AAV-SCH9 or AAV2.NN.

3. The method of claim 2, wherein the library of AAV vectors comprise AAV- pRep2-SCH9 vectors or AAV-pRep2-SCH9 (repeat 136bp) vectors.

4. The method of claim 2, wherein the library of AAV vectors are administered to the animal embryo around a development stage that is equivalent to mouse embryonic day 11.5 (E11.5), 12.5 (E12.5), 13.5 (E13.5), 14.5 (E14.5), or 15.5 (E15.5).

5. The method of claim 2, wherein the library of AAV vectors are administered in utero into the brain of the embryo.

6. The method of claim 1, wherein the transgenic system is a postnatal or adult Cas9-expressing transgenic animal, and wherein the library of AAV vectors comprise vectors of serotype AAV-PHP.eB or AAV.PHP.S.

7. The method of claim 6, wherein the library of AAV vectors are administered to the animal at an age that is equivalent to mouse age of from about postnatal day 10 (P10) to about 18 months old.

8. The method of claim 6, wherein the library of AAV vectors are administered retroorbitally to the animal.

9. The method of claim 1, wherein the specific cell type is newborn neuron or progenitor cell thereof obtained from a brain tissue.

10. The method of claim 9, wherein the brain tissue is from neocortex, olfactory bulb, striatum, hippocampus, thalamus, or cerebellum.

11. The method of claim 9, wherein the target genes are known or suspected to be associated with a neurodevelopmental disorder.

12. The method of claim 1, wherein each of the genetic perturbations comprises one or more guide RNAs (gRNAs) for introducing genomic changes via CRISPR-Cas9 gene editing in the coding region of a target gene.

13. The method of claim 12, wherein the one or more gRNAs each target a sequence at the 5’ end of a target gene’s coding region.

14. The method of claim 12, wherein the AAV vectors further encode a transposable element that is flanked with the gRNAs, and wherein the library of AAV vectors is introduced into the transgenic system in the presence of a transposase that recognizes the transposable element and inserts it into the host genome.

15. The method of claim 14, wherein the transposase is encoded by a second AAV vector co-introduced with the library of AAV vectors into the transgenic system.

16. The method of claim 14, wherein the transposase is Transposase Hyperactive piggybac (hypPB).

17. The method of claim 1, further comprising enriching the identified one or more cells that express one or more genetic perturbations and display a specific phenotype.

18. The method of claim 12, wherein the one or more gRNAs’ coding sequence is operatively linked to a reporter gene in the AAV vectors.

19. The method of claim 18, wherein the one or more cells expressing a genetic perturbation is identified by detecting expression of the reporter gene.

20. The method of claim 18, wherein the genetic perturbation encoded by the AAV vector introduced into each of the one or more cells is determined via capturing the one or more gRNAs via single-cell RNA sequencing.

21. The method of claim 18, wherein the one or more gRNAs are functional for gene editing and can be directly captured in scRNA-seq.

22. A method for analyzing functions of a plurality of target genes in one or more specific cell types in vivo, comprising: (1) co-introducing into a developing embryo of a Cas9-expressing transgenic animal (i) an AAV vector encoding a transposase, and (ii) a library of AAV vectors each encoding (a) one or more guide RNAs (gRNAs) for one of the target genes and (b) a transposable element that is recognized by the transposase and that is flanked with the gRNAs; wherein the AAV vectors are of serotype AAV-SCH9 or AAV2.NN. (2) identifying one or more cells of a specific cell type from the embryo that express the gRNAs and display a specific phenotype, and (3) determining the gRNAs encoded by each of the AAV vectors introduced into the one or more cells and the corresponding perturbed gene; thereby correlating the specific phenotype with the function of each of the perturbed target genes.

23. The method of claim 22, wherein the transposase is hypPB.