High content and high resolution in vivo screening for analysis of gene function

By combining AAV vectors with transposase systems, the challenge of in vivo screening for high-content phenotypes was overcome, enabling rapid and robust gene editing and expression in multiple cell types, improving the ability to capture single-cell omics data, and significantly enhancing in vivo screening efficiency.

CN122122305APending Publication Date: 2026-05-29THE SCRIPPS RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE SCRIPPS RES INST
Filing Date
2024-09-05
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to perform high-content phenotypic screening in vivo, particularly in effectively labeling, perturbing, and isolating cells across multiple cell types. Furthermore, it is difficult to robustly capture and deconvolve the perturbed identity of each cell by reusing screening agents and mixing perturbing agents, especially in hard-to-reach tissues such as the adult central or peripheral nervous system.

Method used

By combining AAV vector libraries with transposase systems, genetic perturbations are delivered in developing embryos or postnatal animals using AAV-SCH9 or AAV2.NN vectors, and in adult animals using AAV-PHP.eB or AAV.PHP.S vectors. gRNA is captured by single-cell RNA sequencing, enabling rapid and robust gene editing and expression.

Benefits of technology

It enables rapid and robust delivery of transgene expression in vivo, increases the number of labeled cells, enhances the ability to capture single-cell omics data, and enables the analysis of more than 30,000 cells in a single experiment, significantly improving the efficiency and accuracy of in vivo screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122122305A_ABST
    Figure CN122122305A_ABST
Patent Text Reader

Abstract

The present invention provides high-content and high-resolution in vivo screening for the analysis of the function of multiple genes. In the functional screening of the present invention, genetic perturbations are delivered to CRISPR-expressing transgenic systems by specific AAV vectors fully compatible with the Perturb-seq platform. This screening enables in vivo functional genomics analysis in a variety of tissues, cell types, and model organisms with high-throughput single-cell readout.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This patent application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 581,297 (filed September 8, 2023; currently pending). The entire disclosure of that priority application is incorporated herein by reference in its entirety and for all purposes.

[0003] Statement regarding government support

[0004] This invention was completed with government support under license number HG012819 granted by the National Institutes of Health (NIH). The government owns certain rights to this invention. Background Technology

[0005] Human genetic research has identified risk genes clearly associated with diseases and symptoms. In particular, major international initiatives to determine the genetic contribution of neurodevelopmental disorders have yielded results, including information on variations associated with autism spectrum disorder and neurodevelopmental delay. In general, hundreds of genes have been identified as contributing to the risk of neurodevelopmental disorders, and additional genes and loci are involved in schizophrenia and other mental illnesses. Nevertheless, little is known about the underlying mechanisms leading to disease manifestations. One of the key goals of functional genomics is to define the cell type and molecular specificity of gene action: among all heterogeneous cell types, including the central and peripheral nervous systems, what are the most vulnerable cell types and molecular networks affected by each risk gene? Genetic techniques, including programmable perturbations enabled by CRISPR (clustered regularly interspaced short palindromic repeats), have opened new avenues for large-scale experimental testing of these risk genes. In vivo, combined genetic screening in mammals has been applied to several systems and environments to read out gRNA abundance as an indicator of cell proliferation or exhaustion. However, changes in gRNA abundance cannot fully capture all molecular consequences and cellular phenotypes.

[0006] To gain a deep and precise understanding of gene function, scalable technologies such as CRISPR-Cas9 gene editing and single-cell genomics are combined to produce Pertub-seq and CROP-seq, which are readouts using transcriptomics or epigenomics. Existing applications have been developed in in vitro settings in immortalized cell lines. While these in vitro cultures are controlled, they often fail to reflect the complex in vivo intercellular signaling and dynamic physiology of disease rooting and manifestation. In vivo Pertub-seq adapts the system to the developing embryo and characterizes disease risk genes and transcription factors in carcinogenesis at developmental stages where the disease manifests. See, for example, Elena et al., bioRxiv 2023.01.30.525356; and Jin et al., Science 2020; 370, eaaz6063. Gene expression analysis reveals convergent molecular networks and specific neuronal or glial cell subtypes influenced by risk genes in the developing brain environment. However, high-content phenotypic screening in vivo remains challenging due to at least two obstacles: (1) the need to scalably label, perturb, and isolate sufficient cells from multiple cell types in vivo, which is more difficult than most in vitro correspondences, and (2) the challenge of robustly capturing and unconvolving the perturbation identity of each cell in sparse single-cell omics data by reusing screening and mixing perturbing agents (gRNAs). Adding to these challenges is the existing use of lentiviral vectors in most Pertub-seq applications (known for their limited in vivo penetration, transduction, and thermostability) which hinders systematic screening in hard-to-reach tissues such as the adult central or peripheral nervous system. See, for example, Higashikawa and Chang, Virology 2001;280: 124-131.

[0007] Therefore, there is an unmet need for high-content phenotypic screening for in vivo functional genomics studies. This invention addresses this and other needs in the field. Summary of the Invention

[0008] In one aspect, the present invention provides a method for in vivo analysis of the function of multiple target genes in one or more specific cell types. The method requires (1) introducing a library of AAV vectors encoding genetic perturbations of multiple target genes into a transgenic system expressing CRISPR-Cas9, (2) identifying one or more cells from the transgenic system expressing the genetic perturbations and exhibiting a specific phenotype, and (3) determining (a) the genetic perturbation encoded by the AAV vector introduced into each of the one or more cells and (b) the corresponding perturbed gene, thereby associating the corresponding perturbed gene with the phenotype displayed by each of the one or more cells. In some embodiments, the transgenic system used is a developing embryo of a transgenic animal expressing Cas9, and the AAV vector library is a vector with serotypes AAV-SCH9 or AAV2.NN. In some of these embodiments, the AAV vector library used is an AAV-pRep2-SCH9 vector or an AAV-pRep2-SCH9 (136 bp repeat) vector. In several embodiments, the AAV vector library is administered to animal embryos at developmental stages corresponding to approximately day 11.5 (E11.5), 12.5 (E12.5), 13.5 (E13.5), 14.5 (E14.5), or 15.5 (E15.5) of mouse embryonic development. In some embodiments, the AAV vector library is administered intrauterinely to the embryonic brain (lateral ventricle).

[0009] In other embodiments, the transgenic system used is a newborn or adult transgenic animal expressing Cas9, and the AAV vector library is a vector with serotypes AAV-PHP.eB or AAV.PHP.S. In some of these embodiments, the AAV vector library is administered to the animals at an age corresponding to approximately 10 days after birth (P10) to approximately 18 months of age. In some embodiments, the AAV vector library is administered retro-orbitally to the animals.

[0010] In some methods of this invention, the specific cell type to be analyzed is a newly generated neuron or its progenitor cell derived from brain tissue. In some of these methods, the brain tissue is derived from the neocortex, olfactory bulb, striatum, hippocampus, thalamus, or cerebellum. In some methods, the target gene to be analyzed is known or suspected to be associated with neurodevelopmental disorders.

[0011] In some embodiments, each of the genetic perturbations to be introduced into a specific cell type in vivo contains one or more guide RNAs (gRNAs) for introducing genomic alterations in the coding region of a target gene via CRISPR-Cas9 gene editing. In some of these embodiments, the one or more gRNAs each target a sequence at the 5' end of the coding region of the target gene. In some methods, the AAV vector also encodes a transposon element flanked by the gRNA, and a library of the AAV vector is introduced into the transgenic system in the presence of a transposase that recognizes the transposon element and inserts it into the host genome. In some of these methods, the transposase is encoded by a second AAV vector introduced into the transgenic system along with the library of the AAV vector. In some embodiments, the transposase used is the transposase hypoPB.

[0012] Some methods of the present invention may additionally require enrichment of one or more cells identified as expressing one or more genetic perturbations and exhibiting a specific phenotype. In some methods, the coding sequences of one or more gRNAs are operatively linked to a reporter gene in an AAV vector. In some of these methods, one or more cells expressing a genetic perturbation are identified by detecting the expression of the reporter gene. In some methods, the genetic perturbation encoded by the AAV vector introduced into each of one or more cells is determined by capturing one or more gRNAs using single-cell RNA sequencing (scRNA-seq). In some methods, one or more gRNAs are functional for gene editing and can be captured directly in scRNA-seq.

[0013] Some methods of the present invention utilize AAV-SCH9 or AAV2.NN vectors. In some of these methods, the AAV vector library is administered to animal embryos at developmental stages corresponding to approximately day 11.5 (E11.5), 12.5 (E12.5), 13.5 (E13.5), 14.5 (E14.5), or 15.5 (E15.5) of mouse embryonic development. Some methods of the present invention utilize AAV-PHP.eB or AAV.PHP.S vectors. In some of these methods, the AAV vector library is administered to animals from day 10 (P10) to 18 months of age.

[0014] In one related aspect, the present invention provides methods for in vivo analysis of the function of multiple target genes in one or more specific cell types. These methods include: (1) co-introducing a library of (i) an AAV vector encoding a transposase and (ii) AAV vectors each encoding (a) one or more guide RNAs (gRNAs) targeting one of the target genes and (b) transposable elements recognized by the transposase and flanked by the gRNAs into a developing embryo of a transgenic animal expressing Cas9; wherein the AAV vector has a serotype AAV-SCH9 or AAV2.NN;

[0015] (2) Identify one or more cells from the embryo that express the gRNA and exhibit a specific phenotype of a specific cell type.

[0016] (3) Identify the gRNA encoded by each of the AAV vectors introduced into one or more cells and the corresponding perturbed gene; thereby associating a specific phenotype with the function of each of the perturbed target genes.

[0017] In some of these methods, the transposase used is hypoPB (highly active piggybac).

[0018] A further understanding of the nature and advantages of the invention can be achieved by referring to the remainder of the specification and the claims. Attached Figure Description

[0019] Figure 1 In vivo barcoded AAV serotype screening identified AAV-SCH9 as an effective target for the developing brain. (A) Schematic diagram of intrauterine administration of the AAV library at E13.5, followed by fluorescence-activated cell sorting (FACS) cell enrichment and barcode analysis using next-generation sequencing. (B to C) Immunofluorescence analysis of brain slices two days after AAV library administration: (B) Co-staining with markers of newly generated projection neurons (TBR1 and CTIP2) and intermediate progenitor cell marker (TBR2) in the dorsal cortex and ganglion iceminence (GE); (C) GFP co-expressing neuronal and progenitor cell markers (including TBR1, TBR2, and CTIP2). +Quantification of cell percentages. (D) Heatmap of the abundance proportions of 86 AAV serotypes in the AAV library, HT22 cells 48 hours after transduction, and embryonic mouse brain; each row represents the abundance of AAV serotypes. (E) Principal component analysis of the percentage abundance of AAV libraries, HT22 cells 24 or 48 hours after transduction, and mouse brain. (F) Percentage abundance of AAV-SCH9 in the initial AAV library (input for transduction), HT22 cells 24 or 48 hours after transduction, and mouse brain. Error bars indicate the standard error of the mean. (G) Volcano plot of changes in AAV serotypes in mouse brain (left panel) or HT22 cells (right panel) 48 hours after transduction compared to the initial AAV library. Scale bars indicate 250 μm (left panel in B) and 50 μm (right panel in B and C).

[0020] Figure 2 AAV-SCH9-labeled newborn neurons are distributed throughout brain regions at different developmental stages. (A) Schematic diagram of secondary AAV serotype screening: 14 AAV serotypes were barcoded and introduced intrauterine as a serotype at E13.5, followed by scRNA-seq 48 hours later. (B) From sorting (GFP) + Visualization of 11 major cell populations identified from unsorted cells (right panel) and unsorted cells (left panel) using uniform manifold approximation and projection (UMAP); cell types include: upper and deep projection neurons (ULPN, DLPN), migrating neurons (Mig. neurons), apical progenitor cells (Api. prog), intermediate progenitor cells (IP), interneurons originating from the medial ganglion eminence (IN-MGE), interneurons originating from the non-medial ganglion eminence (IN-non-MGE), Cajal-Retzius cells (CR), fibroblasts (Fibro), mural cells, and microglia (Mg). (C) AAV serotype barcode expression in each cell type. Each row represents one AAV barcode, and each AAV serotype is associated with a unique set of 3 barcodes (BCa, BCb, and BCC). Each column represents the barcode expression in one cell, arranged by cell type. (D) Immunofluorescence analysis of brain slices transduced with AAV-SCH9-GFP-KASH E14.5 and co-stained with markers including CTIP2, TBR1, and TBR2, and GFP. +Percentage of marker co-expression in cells. The box in the top plot indicates a selected field of view in the bottom plot; the arrow indicates representative cells with marker co-localization. (E) AAV-SCH9-GFP-KASH was administered at two different time points (E13.5 or E17.5), and cortical sections were analyzed 48 hours later to quantify GFP in the laminar layer. + Cell distribution, which is evenly divided into six bins. (F) AAV-SCH9-GFP-KASH or lentiviral reporter (GFP) administered at E13.5 resulted in different brain region markings at P7. Scale bars indicate 500 μm (left panel in D and top panel in F) or 50 μm (right panels in D and E and bottom panel in F).

[0021] Figure 3 The hypPB transposon enhanced and stabilized expression in the embryonic and adult brain, as well as the peripheral nervous system. (A) Schematic diagram of the molecular design that enhances transgene expression and prevents loss due to cell division and differentiation. Shaded boxes indicate transposons with inverted repeats (IR) flanking them. (B) Time-lapse imaging shows that co-transfection with hypPB enhances in vitro expression of the GFP transgene. Each point represents an analysis from a selected field of view from one well. (C to E) Using three targeting vectors: AAV-SCH9, AAV9-PHP.eB, and AAV9-PHP.S, the transposon was stably expressed in vivo in the embryonic brain, adult brain, and adult dorsal root ganglia. (F) In the presence of hypPB, AAV-SCH9 labeled more cortical neurons; the Y-axis represents the number of cells expressing GFP / mm. 2(n=3 animals / condition). An asterisk indicates a p-value <0.0001 for the Welch t-test. (G) Cortical neurons expressing AAV9-PHP.eB-labeled neurons with increased intensity under hypoPB conditions: mean fluorescence intensity (X-axis) in the cortical layer (Y-axis) from the ventricular region (0) to the pia mater (1) (n=3 animals / condition). (H) Whole-genome sequencing of AAV-SCH9 transduced cells in vivo showing genomic regions of integration events. Left panel: Each line represents a unique hybridization read between the mouse genome and transposons, shown as a wheel diagram. Right panel: Percentage of integration events occurring between genes and coding regions in the genome. (I) Example reads capturing the connection between the mouse genome and transposons, indicating integration sites. Top panel: Schematic diagram of transposon sequences (SEQ ID NO: 27 and 37). Bottom image: Example reads aligned with transposon sequences (orange box) and the mouse genome, and chromosome numbers of the mouse genome (left side: SEQ ID NO: 28-36; right side: SEQ ID NO: 38-46). Scale bars indicate 100 μm (right images in B, C, D, and E) or 500 μm (left images in C and D).

[0022] Figure 4In vivo Pertub-seq identified cell type-specific changes in transcription factor perturbations. (A) Schematic diagram of the screening design. Shaded boxes indicate transposons with inverted repeats (IR) flanking them; NT and ST represent the non-targeted control and the safe-targeted control, respectively. (B) UMAP plot of filtered cells with a single perturbation, where each cell is colored according to the annotated cell type. (C) UMAP plot of filtered cells, where each cell is colored according to the gRNA identity estimated by DemuxEM (downsampling (top) and batch / channel (bottom)). (D) Top: Schematic diagram of the 5' and 3' scRNA-seq capture mechanism. Bottom: Comparison of gRNA capture rates by measuring the percentage of cells with a gRNA UMI greater than 5 in each cell type in 5' and 3' scRNA-seq, where the X-axis represents cell type and the Y-axis represents the percentage of cells assigned gRNA identities. (E) Bar chart showing the percentage of cells assigned to one or more gRNAs in the major cell types, with colors representing gRNA identity assignments. (F) Percentage of insertion / deletion reads in the Foxg1 gRNA target loci perturbed by scRNA-seq data compared to the non-targeted control 2 (NT2). (G) Heatmap showing the change in cell type proportion by perturbed. Circles highlight FDR-corrected P-values ​​<0.05. (H) Dot plot of the number of differentially expressed genes (DEGs) and cell number (dot size) in cell type-perturbed combinations. Boxes highlight those supported by at least 2 gRNAs. Robust changes in 10 DEGs. (I) Volcano plot of the cell type-specific effects of Foxg1 perturbation on differentially expressed genes in excitatory neurons of layer 6 CT and layer 5 IT. Dot markers indicate DEGs with significant alterations; circular markers highlight DEGs. Detailed Implementation

[0023] I. Overview

[0024] The thousands of disease risk genes identified through human genetics far exceed our current ability to systematically study their function in the physiological environment. Two substantial barriers hinder progress in in vivo genetic screening: the lack of tools to obtain large numbers of cells within living tissue, and the ability to introduce sufficiently robust expression for detection in pooled assays (e.g., single-cell omics). AAV offers a promising strategy for delivering genetic perturbations in vivo to a wide range of cell types. However, conventional AAV expression tends to be transient, at relatively low levels, and subsequently diluted by cell division, which together poses a challenge to accurately recovering the perturbation identity of sparsely labeled cells in pooled assays (e.g., Perturb-seq). See, for example, Lang et al., 2019; Nat. Commun. 10, 3415. Most critically, the slow onset of AAV-mediated expression—often taking days or even weeks—is suboptimal for studying dynamic gene function in rapidly evolving cellular environments (e.g., neural development). For example, peak AAV-transgenic expression typically occurs more than seven days later, a time span that covers the entire window of cortical development. These limitations underscore the urgent need for AAV systems that deliver rapid and robust transgenic expression to allow for the scaling up of large-scale parallel perturb-seq capabilities in vivo.

[0025] This invention provides an AAV-based high-content in vivo screening platform that overcomes these limitations. The universal platform for in vivo perturb-seq can target diverse tissue and cell types at single-cell resolution with high content phenotypic readout. This invention is partly derived from research conducted by the inventors to develop an AAV-based platform for massively parallel in vivo perturb-seq, enabling the targeting of a broad range of tissue and cell types at single-cell resolution with gene expression-based characterization. As detailed herein, through barcoding screening of 86 phylogenetically distinct AAVs in vivo, the inventors identified specific AAV serotypes, including AAV-SCH9, capable of rapid and robust transgene delivery in newborn neurons and progenitor cells within 48 hours post-transduction (compared to 2–3 weeks using currently available methods). The identified AAV vector further synergizes with transposon systems to ensure rapid and sustained gRNA expression in both target and daughter cells, facilitating efficient gRNA capture in single-cell gene expression-based analyses, where >82% of cells recovered their perturbation identity. Through intrauterine perturbation screening, the inventors further revealed the cell type-specific effects of perturbation on the proportion of neuronal subtypes. Differential expression analysis further elucidated the cell type-differential transcriptomic changes: Foxg1 mainly affects neurons in the 5th layer of the teleencephalon and neurons in the 6th layer of the corticothalamus, and it leads to the loss of function of the disinhibitory transcriptional network closely related to neurodevelopmental delay, resulting in hybrid cell fate and misexpression of axonal guidance and synaptic molecules.

[0026] Previously, in vivo perturb-seq was limited by the number of cells that could be labeled and harvested from tissues, as well as the ability to reliably capture gRNA. Importantly, the AAV-based in vivo perturb-seq platform illustrated in this paper achieved 10-fold higher labeling in embryonic brain (50,075 cells in two batches) and could label >6% of brain cells, exceeding the <0.1% lentiviral potency reported by Jin et al. (Science 2020; 370, eaaz6063), allowing for in vivo analysis of more than 30,000 cells in a single assay. In contrast, conventional methods using lentiviral vectors typically only allow for in vivo analysis of thousands of cells from a single batch, such as 46,770 cells in 17 batches, as described in Jin et al., 2020. Therefore, this invention represents a significant improvement and technical advantage over other known in vivo perturb-seq screening methods that rely on other viral vectors, such as those described in Jin et al., 2020; and lentiviral vector-based screening as described in U.S. Patent Application Publication No. 2021 / 0172017A1. Compatible with a variety of perturbation techniques (CRISPRa / i) and phenotypic measurements (single-cell or spatial multi-omics), the perturb-seq platform described herein provides a flexible approach to studying gene function in different cell types in vivo, translating gene variants into their causal functions.

[0027] It should be noted that, unless otherwise stated, the present invention is not limited to the specific methods, procedures, and reagents described, as these can vary. Unless otherwise indicated, the practice of the present invention utilizes conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology, all of which are within the scope of the art. Such techniques are well illustrated in the literature. Exemplary methods are described, for example, in the following references: Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press (3rd edition, 2001); Brent et al., Current Protocols in Molecular Biology, John Wiley & Sons, Inc. (ringbou editor, 2003); Freshney, Culture of Animal Cells: A Manual of Basic Technique, Wiley-Liss, Inc. (4th edition, 2000); and Weissbach & Weissbach, Methods for Plant Molecular Biology, Academic Press, NY, Part VIII, pp. 421-463, 1988. In addition, the following sections provide more detailed guidance for practicing this invention.

[0028] II. definition

[0029] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The following references provide general definitions for many of the terms used in this invention: Academic Press Dictionary of Science and Technology, Morris (ed.), Academic Press (1st edition, 1992); Oxford Dictionary of Biochemistry and Molecular Biology, Smith et al. (ed.), Oxford University Press (revised edition, 2000); Encyclopaedic Dictionary of Chemistry, Kumar (ed.), Anmol Publications Pvt. Ltd. (2002); Dictionary of Microbiology and Molecular Biology, Singleton et al. (ed.), John Wiley & Sons (3rd edition, 2002); Dictionary of Chemistry, Hunt (ed.), Routledge (1st edition, 1999); Dictionary of Pharmaceutical Medicine, Nahler (ed.), Springer-Verlag Telos (1994); Dictionary of Organic Chemistry, Kumar and Anandand (ed.), Anmol Publications Pvt. Ltd. (2002); and A Dictionary of Biology (Oxford Paperback Reference), Martin and Hine (eds.), Oxford University Press (4th edition, 2000). Further clarification of these terms is provided herein when some of them are specifically applied to this invention.

[0030] Unless the context clearly specifies otherwise, nouns used herein without quantifiers include plural referents. Thus, for example, reference to “cell” includes multiple such cells, reference to “protein” includes one or more proteins and their equivalents known to those skilled in the art, and so on.

[0031] As used herein, AAV refers to adeno-associated virus and may be used to encompass naturally occurring wild-type viruses themselves or their derivatives. An AAV vector is a small, single-stranded DNA viral vector containing an icosahedral protein capsid. Unless otherwise required, the term encompasses all subtypes, serotypes, and pseudotypes, as well as both naturally occurring and recombinant forms. A pseudotype AAV is an AAV containing a capsid protein from one serotype and a viral genome containing a 5'-3' ITR from a second serotype. The abbreviation "rAAV" refers to recombinant adeno-associated virus particles or recombinant AAV vectors (or "rAAV vectors"). "AAV virus" or "AAV virus particle" refers to a viral particle composed of at least one AAV capsid protein (preferably all capsid proteins of wild-type AAV) and capsid-encapsulated polynucleotides. If the particle contains heterologous polynucleotides (i.e., polynucleotides other than the wild-type AAV genome, such as transgenes to be delivered into mammalian cells), it is generally referred to as "rAAV".

[0032] The Cas9 CRISPR genome editing system requires a Cas9 DNase obtained from *Streptococcus pyogenes* and a guide RNA (gRNA) to target the Cas9 DNase activity to a complementary genomic sequence. For the Cas9 CRISPR system, the gRNA consists of two parts: crrRNA (crRNA), a 17-20 nucleotide sequence complementary to the target DNA; and trans-activating crRNA (tracrRNA), which acts as a binding scaffold for the Cas nuclease. In addition to sequence complementarity, a short protospacer-adjacent motif (PAM) precedes the cleavage site in the target DNA. In bacteria, Cas9 relies on RNase III to excise the crRNA from the CRISPR array.

[0033] The protospacer adjacent motif (PAM) is a 2- to 6-base-pair DNA sequence that follows the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. The PAM is a component of invading viruses or plasmids, but not a component of the bacterial CRISPR locus. If the target DNA sequence is not followed by a PAM sequence, Cas9 will not successfully bind to or cleave the target DNA sequence. The PAM is an essential targeting component (not found in the bacterial genome) that distinguishes the bacterial's own DNA from non-self DNA, thus preventing the CRISPR locus from being targeted and destroyed by nucleases. A typical PAM is the sequence 5'-NGG-3', where "N" is any nucleobase followed by two guanine ("G") nucleotides. Guide RNA (gRNA) can transport Cas9 anywhere in the genome for gene editing, but editing cannot occur at any site other than where Cas9 recognizes the PAM.

[0034] Typical PAMs are associated with the Cas9 nuclease of Streptococcus pyogenes (named SpCas9), while different PAMs are associated with the Cas9 proteins of bacteria such as Neisseria meningitidis, Treponema denticola, and Streptococcus thermophilus. 5'-NGA-3' can be a highly efficient atypical PAM for human cells, but its efficiency varies with genomic location. Attempts have been made to engineer Cas9 to recognize different PAMs to improve CRISPR-Cas9's ability to perform gene editing at any desired genomic location. The Cas9 of the novel culprit *Francisella novicida* recognizes the typical PAM sequence 5'-NGG-3', but has been engineered to recognize PAM 5'-YG-3' (where "Y" stands for pyrimidine), thus increasing the range of possible Cas9 targets.

[0035] The protospacer is a spacer sequence within a bacterial CRISPR locus that is inserted into the CRISPR locus by invading viral or plasmid DNA. Upon subsequent invasion, the Cas9 nuclease attaches to tracrRNA:crRNA, which guides Cas9 to the invading protospacer sequence. However, Cas9 does not cleave the protospacer sequence unless an adjacent PAM sequence is present. Spacers within bacterial CRISPR loci do not contain PAM sequences and therefore will not be cleaved by nucleases. However, protospacers in invading viruses or plasmids will contain PAM sequences and will therefore be cleaved by the Cas9 nuclease. To edit genes, guide RNA (gRNA) is synthesized to perform the function of the tracrRNA:crRNA complex in recognizing gene sequences with a PAM sequence at the 3' end.

[0036] "Host cell" or "target cell" refers to a living cell in which a heterologous polynucleotide sequence is to be introduced or has already been introduced. Living cells include both cultured cells and cells in a living organism. Methods for introducing heterologous polynucleotide sequences into cells are well known, such as transfection, electroporation, calcium phosphate precipitation, microinjection, transformation, viral infection, etc. Typically, the heterologous polynucleotide sequence to be introduced into the cell is a reproducible expression vector or cloning vector. In some embodiments, the host cell may be engineered to incorporate the desired gene onto its chromosome or into its genome. Many host cells (e.g., CHO cells) that can be used in the practice of this invention are well known in the art as hosts. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press (3rd edition, 2001); and Brent et al., Current Protocols in Molecular Biology, John Wiley & Sons, Inc. (ringbou edition, 2003). In some preferred embodiments, the host cell is a mammalian cell.

[0037] The term "operably linked" or "operably associated" refers to a functional connection between genetic elements in a manner that enables them to perform their normal functions. For example, a gene is operably linked to a promoter when transcription of a gene is under the control of a promoter and the resulting transcript is correctly translated into a protein normally encoded by that gene. Similarly, if a gRNA is co-produced with a transcript encoded by a reporter gene, the sequence encoding the gRNA is operably linked to the reporter gene.

[0038] A “substantially identical” nucleic acid or amino acid sequence refers to a polynucleotide or amino acid sequence that contains at least 75%, 80%, or 90% sequence identity with a reference sequence as measured using standard parameters by one of the known procedures described herein (e.g., BLAST). Sequence identity is preferably at least 95%, more preferably at least 98%, and most preferably at least 99%. In some embodiments, the target sequence has approximately the same length as the reference sequence, i.e., it consists of approximately the same number of consecutive amino acid residues (for polypeptide sequences) or nucleotide residues (for polynucleotide sequences). If the polynucleotide sequence is composed of RNA or DNA, they are also substantially identical, despite chemical differences between RNA and DNA, and the presence of uracil in RNA instead of thymidine in DNA.

[0039] Sequence identity can be readily determined using a variety of methods known in the art. For example, the BLASTN program (for nucleotide sequences) defaults to a word length (W) of 11, an expected value (E) of 10, M=5, N=-4, and compares two strands. For amino acid sequences, the BLASTP program defaults to a word length (W) of 3, an expected value (E) of 10, and a BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1989)). The percentage of sequence identity is determined by comparing two best-aligned sequences in a comparison window, where a portion of the polynucleotide sequence in the comparison window may contain additions or deletions (i.e., vacancies) compared to a reference sequence (which does not contain additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions in both sequences where the same nucleic acid base or amino acid residue appears to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity.

[0040] As used in this article, “complementary” or “complementary” refers to a nucleotide or nucleotide sequence that hybridizes with a given nucleotide or nucleotide sequence. For example, in DNA, nucleotide A is complementary to T and vice versa, and nucleotide C is complementary to G and vice versa. For example, in RNA, nucleotide A is complementary to nucleotide U and vice versa, and nucleotide C is complementary to nucleotide G and vice versa.

[0041] In the context of DNA, RNA, or DNA-RNA hybrids, "paired" and "unpaired" refer to Watson-Crick base pairs. A sequence is said to be paired with its inverse complementary pair.

[0042] When a foreign or heterologous polynucleotide is introduced into a cell, the cell has been "transformed" or "transfected" by that foreign or heterologous polynucleotide (or, as used interchangeably here, "transgenic" or "target gene"). The transforming polynucleotide may or may not be integrated (covalently linked) into the cell's genome. For example, in prokaryotes, yeast, and mammalian cells, the transforming polynucleotide can be maintained on addendum elements such as plasmids. For eukaryotic cells, stably transformed cells are those in which the transforming polynucleotide has been integrated into the chromosome, allowing it to be inherited by daughter cells through chromosome replication. This stability is demonstrated by the eukaryotic cell's ability to establish cell lines or clones containing daughter cells containing the transforming polynucleotide. A "clone" is a population of cells derived from a single cell or common progenitor cell through mitosis. A "cell line" is a clone of a primary cell that can stably grow for multiple generations in vitro.

[0043] Transposons (TEs) (transposable elements or jumping genes) are DNA sequences that can move and integrate into different locations within the genome. There are two main classes of TEs: class I TEs (or retrotransposons), which move via RNA intermediates through a "copy and paste" mechanism of transposition, and class II TEs, which move via DNA intermediates through a "cut and paste" mechanism of transposition. DNA transposons can move within an organism's DNA via single-stranded or double-stranded DNA intermediates. DNA transposons are found in both prokaryotes and eukaryotes. They can form an important part of an organism's genome, especially in eukaryotes. In prokaryotes, TEs can promote the horizontal transfer of genes associated with antibiotic resistance or other virulence. DNA transposons can be autonomous or non-autonomous. Autonomous transposons encode their own transposases that promote gene jumping, while non-autonomous transposons require the transposase activity of another transposable element. DNA transposons are delineated by flanking repeats that mark the location where the transposase will cut the DNA. These DNA elements then reintegrate into different locations within the genome.

[0044] A "vector" or "construct" is a non-naturally occurring nucleic acid, with or without a vector, that can be introduced into or has been introduced into a cell. Vectors already introduced into cells contain transfected plasmids and integrated DNA molecules, including those resulting from retroviral integration, AAV vector integration, and integration via homologous recombination. Vectors capable of guiding the expression of heterologous polynucleotide or transgenic sequences encoding one or more polypeptides are called "expression vectors" or "expression constructs." Cloned transgenic sequences or open reading frames (ORFs) are typically placed under the control (i.e., operatively linked) of certain regulatory sequences, such as promoters, enhancers, and polynucleotide switch sequences.

[0045] III. AAV-based Perturb-seq platform for in vivo gene function analysis

[0046] This invention provides an in vivo perturb-seq platform for analyzing target genes in different tissue and cell types with high intrinsic phenotypic readout and single-cell resolution. The in vivo system of this invention utilizes a serotype-specific library of AAV vectors to deliver genetic perturbations to regulate one or more target genes in specific cell or tissue types (e.g., central and peripheral nervous systems) in animal embryos or adult or aged animals. Adeno-associated virus (AAV) is a small, non-enveloped virus suitable for use as a gene transfer vector. AAV vectors refer to recombinant adeno-associated viruses derived from non-pathogenic parvoviruses. They substantially do not elicit a cellular immune response and produce transgenic expression lasting for months in most systems. Like adenoviruses, AAV vectors are capable of infecting both replicating and non-replicating cells and are considered non-pathogenic to humans. Delivery of heterologous polynucleotide sequences via recombinant AAV can provide safe, unobtrusive, and sustained expression (>2 years) of high-level protein therapeutics.

[0047] As detailed in the embodiments herein, specific AAV vector serotypes suitable for the present invention are identified by barcode screening of phylogenetically different AAVs in vivo. AAV vector serotypes refer to classifications of AAV variants based on their molecular composition and their targeting specificity in tissue and cell types. To date, at least 11 AAV serotypes exist, including AAV1, 2, 4, 5, 6, 7, 8, and 9 (see, for example, Issa et al., Cells 2023; 12(5):785). General features, including structural information of known AAV serotypes, are provided in, for example, Srivastava, Curr Opin Virol. 2016;21: 75–80; Issa et al., Cells 2023; 12:785; Chen et al., Human Gene Therapy. 2020; 31:440-447; Pillay et al., J. Virol. 2017; 91: e00391-17; Naso et al., BioDrugs. 2017; 31:317–334; Wang et al., Nat. Rev. Drug Discov. 2019; 18: 358–378; and Wu et al., Mol. Ther.2006; 14: 316-327. As indicated herein, specific AAV serotypes and variants suitable for the practice of this invention include AAV-SCH9, AAV2.NN, AAV9-PHP.eB, and AAV9-PHP.S. AAV-SCH9 is such an AAV serotype that is a hybrid of AAV2, 8, and 9, and was originally prepared for transduction of adult neural stem cells. Detailed structural information for AAV-SCH9 is provided, for example, in Ojala et al., Mol. Ther. 2018; 26: 304-319. Specific AAV-SCH9 and AAV2.NN vectors are available from commercial suppliers such as Addgene (Watertown, MA). Similarly, other known AAV vectors (including AAV9-PHP.eB and AAV9-PHP.S exemplified herein) are also commercially available, for example, from suppliers such as Addgene (Watertown, MA).

[0048] By utilizing specific AAV serotypes, the in vivo perturb-seq platform described herein provides a flexible and scalable strategy for effectively targeting selective cell types in vivo and performing scalable genetic perturbation screening with reliable phenotypic readouts using high-throughput single-cell omics methods. Compared to conventional methods, in vivo perturb-seq screening based on these AAV vectors can effectively and economically target and perturb a much larger number of cells in vivo within a single experiment. In some embodiments, the AAV vectors used are designed to deliver genetic perturbations to target genes in tissues of animal embryos. In some of these embodiments, the AAV vector used is AAV-Rep2-SCH9. In other embodiments, the AAV vector used is AAV2.NN. In some embodiments, the AAV vectors used are designed to deliver genetic perturbations to target genes in tissues of adult or aged animals (e.g., adult brain and peripheral nervous system). In some of these embodiments, the AAV vector used is AAV-PHP.eB. In other embodiments, the AAV vector used is AAV9-PHP.S. Detailed structural information about these specific AAV vectors is known in the art. See, for example, Ojala et al., Mol. Ther., 2018; 26:304-319; and Chan et al., Nat. Neurosci., 2017; 20:1172-1179.

[0049] As shown herein, AAV vectors can be administered to transgenic animal systems (e.g., embryos or adult animals) expressing the CRISSPR-Cas gene editing system. In some embodiments, a library of the AAV vector is administered to developing embryos of CRISSPR-Cas9 transgenic mice. Some of these embodiments may use the AAV-Rep2-SCH9 vector or AAV2.NN. In some embodiments, these vectors may be administered intrauterine to the brain (e.g., the lateral ventricle) at any time corresponding to mouse embryonic developmental stages E11.5 to E19.5. In some of these embodiments, the vector is administered to the embryo during the time period E11.5 to 15.5 or E13.5 to 17.5. In other embodiments, a library of the AAV vector is administered to mature adult animals of CRISSPR-Cas9 transgenic mice. Some of these embodiments may use the AAV-PHP.eB vector or AAV9-PHP.S vector. In several embodiments, the AAV vector is administered to the animal at an age corresponding to the age of the mouse from approximately day 10 after birth (P10) to approximately 18 months of age. In some of these implementation schemes, the vector is administered to adult mice via retroorbital injection, as illustrated herein.

[0050] In some preferred embodiments, the genetic perturbation encoded by each of the AAV vectors comprises a gRNA that directs the regulation (e.g., cleavage) of a target gene via a CRISPR gene editing system. At least one gRNA targeting the target gene (“target gene”) is encoded by an AAV vector. In some embodiments, the AAV vector library contains two or more gRNAs designed to target each of the target genes. Two or more gRNAs targeting a gene may be encoded by a single AAV vector. Alternatively, multiple gRNAs targeting the same gene may be encoded by different AAV vectors in the library. In some embodiments, the animal or animal embryo used in the screening is derived from a transgenic nonhuman animal modified to express a CRISPR editing system. In some preferred embodiments, the gene editing system expressed by the animal is CRISPR-Cas9. In some embodiments, each of the gRNAs may be placed under the control of a different promoter, including, for example, the human U6 polymerase III promoter illustrated herein. In some embodiments, a reporter gene, such as GFP illustrated herein, may be operatively linked to the sequence encoding the gRNA in the AAV vector. After the AAV vector is introduced into an animal embryo or mature animal, the reporter gene allows for the enrichment of cells expressing the corresponding gRNA. Unless otherwise described herein, the delivery of the vector into a transgenic system expressing CRSIPR (e.g., a mouse embryo), the expression and detection of genetic perturbations (e.g., gRNAs targeting specific gene segments), the identification (e.g., by reporter gene expression) and enrichment of cells expressing genetic perturbations (perturbed cells), and the analysis of phenotype-related genes in perturbed cells (e.g., single-cell transcriptome analysis) can all be performed using techniques specifically illustrated herein and / or those known in the art. See, for example, U.S. Patent Application Publication No. 2021 / 0172017.

[0051] In some embodiments, the AAV vector delivering genetic perturbations used in the in vivo screening platform may also be coupled to a transposon system. As detailed below, this is to ensure rapid, cell-type-specific, and reliable expression of gRNA in target cells and their daughter cells for efficient gRNA capture in single-cell transcriptome analysis. Due to the incorporation of the transposon system, the AAV-based in vivo screening platform of the present invention can effectively and economically target and perturb genes in both embryos and adults, as well as in the environment. This represents a significant improvement over genetic screening using lentiviruses or any other viral vectors (e.g., lentiviral vector-based systems as described in U.S. Patent Application Publication No. 2021 / 0172017).

[0052] To construct AAV vectors encoding genetic perturbations in the practice of this invention, the coding sequence for the genetic perturbation (e.g., a sequence encoding gRNA) and any other operatively linked sequence (e.g., a reporter gene) are typically inserted into the viral genome in place of certain viral sequences to produce a replication-defective viral construct. Methods for generating AAV vectors are well known in the art. AAV vectors have also been used in numerous reported studies for gene therapy in both research and clinical settings. See, for example, Kaplitt et al., Lancet 369: 2097–105, 2007; Daya et al., Clin Microbiol Rev. 21(4): 583–593, 2008; Strobel et al., Am. J. Resp. Cell Mol. Biol. 53: 291–302, 2015; and Kotterman et al., Nat. Rev. Genet. 15:445–451, 2014. In some embodiments, the AAV viral vectors and related reagents (e.g., packaging cell lines) suitable for use in this invention are commercially available. In some of these embodiments, the specific AAV vector used to practice this invention may be based on the pAAV-MCS construct available from Agilent Technologies (Santa Clara, CA).

[0053] IV. Transgenic animal systems expressing CRISPR function

[0054] In some preferred embodiments, the genetic perturbation delivered by the AAV vector in the in vivo Pertub-seq platform of the present invention comprises a gRNA that guides the genetic regulation of the target gene. Typically, transgenic animal systems expressing CRISPR gene-editing function are used to express genetic perturbations and examine their effects on target genes. As illustrated herein, transgenic animal systems may express CRISPR-Cas9 gene-editing function. In these embodiments, the genetic perturbation encoded by the AAV vector may be a gRNA designed to subject one or more target genes to the activity of the CRISPPR-Cas system in the transgenic animal. The library of the AAV vector to be introduced into the transgenic system encodes at least one gRNA designed for each of the target genes. In some embodiments, two or more gRNAs are encoded by a library of AAV vectors targeting each of the target genes. In some of these embodiments, three or four gRNAs are encoded by a library of AAV vectors targeting each of the target genes. In some embodiments, multiple gRNAs designed for a specific target gene are encoded by multiple AAV vectors. In other embodiments, multiple gRNAs designed for a specific target gene are encoded by a single AAV vector.

[0055] Any transgenic system with CRISPPR-Cas function can be used in the practice of this invention. Generally, these include any transgenic animal or embryo at multiple developmental stages that has been engineered to express the Cas enzyme and, optionally, other components required for CRISPR function. Transgenic systems may constitutively or inducibly express the Cas enzyme and / or other components in some or all of their cells. In many embodiments, the Cas enzyme expressed in the transgenic system is a Cas type I, II, III, IV, or V protein. In some embodiments, the Cas enzyme (e.g., Cas9) is expressed by the transgenic system, while other components may be provided separately, for example, by another vector to be introduced into the transgenic system (embryo or mature animal). A more detailed description of various CRISPR-Cas gene-editing systems that can be incorporated into the transgenic systems of this invention is provided, for example, in U.S. Patent Application Publication No. 2021 / 0172017.

[0056] In some preferred embodiments illustrated herein, the genetic perturbation encoded by the vector is a gRNA designed to affect the target gene in the transgenic system through CRISPPR-Cas9 gene editing. In these embodiments, each AAV vector encodes one or more gRNAs that can direct the target gene to the enzymatic activity of Cas9. Methods known in the art for genome-scale screening of perturbations in single cells using a CRISPR-Cas9 gene editing system can be readily used and modified as needed in the practice of this invention.See, e.g., Dixit et al., “Perturb-Seq: DissectingMolecular Circuits with Scalable Single-Cell RNA Profiling of Pooled GeneticScreens” 2016, Cell 167, 1853-1866; Adamson et al., “A Multiplexed Single-CellCRISPR Screening Platform Enables Systematic Dissection of the UnfoldedProtein Response” 2016, Cell 167, 1867-1882; Feldman et al., Lentiviral co-packaging mitigates the effects of intermolecular recombination and multiple integrations in pooled genetic screens, bioRxiv 262121, doi: doi.org / 10.1101 / 262121; Datlinger, et al., 2017, Pooled CRISPR screening with single-celltranscriptome readout. Nature Methods. Volume 14 Issue 3 DOI: 10.1038 / nmeth.4177; Hill et al., On the design of CRISPR-based single cell molecular screens, NatMethods. 2018 April; 15(4): 271-274; Replogle et al., “Combinatorial single-cellCRISPR screens by direct guide RNA capture and targeted sequencing” NatBiotechnol (2020). Doi.org / 10.1038 / s41587-020-0470-y; and PCT Publication WO / 2017 / 075294). The transgenic systems described in any of these reports may be used and are suitable for use in the practice of this invention. In some preferred embodiments, transgenic mice expressing Cas9 may be used in the practice of this invention.As illustrated in this article, transgenic mice expressing Cas9 can be readily obtained from commercial suppliers such as Jackson Laboratories (Bar Harbor, ME).

[0057] V. Transposon subsystem co-delivered with AAV vector

[0058] In some embodiments, the AAV-based in vivo pertub-seq platform of the present invention can be further incorporated with a transposon system. As illustrated herein, the co-introduction of an AAV vector encoding gRNA for perturbing target genes and a transposon system can significantly enhance gRNA expression levels and enable gRNA capture using sparse scRNA-seq readout. Co-expression of transposon elements effectively enhances and stabilizes gRNA expression in the embryonic and adult brain, as well as the peripheral nervous system. It ensures rapid, cell-type-specific, and reliable gRNA expression in target cells and their daughter cells for efficient gRNA capture in single-cell transcriptome analysis.

[0059] Multiple transposon systems can be incorporated into the AAV-based in vivo screening platform of this invention. In some embodiments, the transposon system used is a DNA transposon. The integration of the DNA transposon system with the AAV vector delivering the genetic perturbation can be performed according to the specific scheme illustrated or methods known in the art. Many DNA transposon systems known in the art can also be used and modified for the practice of this invention. These include a wide range of individual DNA transposons from different families and donor organisms that have been characterized in detail in the art, such as Tol2 from medeka fish, the synthetic sequence Sleeping Beauty from transposons present in whitecloud minnow, Atlantic salmon, and rainbow trout, and piggyBac isolated from cabbage looper moth. The piggyBac DNA transposon encodes a highly active transposase (iPB) in mammalian cells and is integrated into a conserved TTAA tetranucleotide sequence. A mouse codon-optimized PB transposase (mPB) also exists, which is 20 times more efficient than iPB in mouse embryonic stem cells. Furthermore, a highly active form of PB transposase (hypPB) with several amino acid substitutions on mPB is known in the art. In mouse embryonic stem cells, this DNA transposon system is able to increase the relative integration frequency by nine times relative to mPB. Detailed technical information on this transposon system is provided, for example, in Yusa et al., Proc Natl Acad Sci USA. 2011; 108:1531–1536. Other transposon systems known in the art, besides PB transposases and derivatives, can also be used in the practice of this invention. These include, for example, the Sleeping Beauty transposon system, the hAT transposon, the MuDR superfamily of transposons, the Tc1 / mariner transposon family, transposons utilizing the IS30 transposase such as Tn2700 and Tn2702, and miniature anti-repeating transposon elements from the silkworm (Bombyx mori) and the long red cone bug (Rhodnius prolixus).See, for example, Querques et al., Nat. Biotechnol. 2019; 37: 1502–1512; Rubin et al., Genetics. 2001; 158: 949–957; Raizada et al., Plant J. 2001; 25:79-91; Dupeyron et al., Mobile DNA 2020; 11: 21; Szabó et al., J. Bacteriol. 2010; 192: 3414-23; and Zhang et al., Genome Biol Evol. 2013; 5: 2020–2031.

[0060] These well-known DNA transposons consist of a transposase gene and flanking inverted terminal repeats (ITRs). The transposase recognizes a specific short target sequence located within the ITR, called a directed repeat (DR). After binding, the transposase cleaves the transposon sequence from the surrounding genomic DNA of the host cell. The complex formed by the mobilized transposon DNA fragment and the still-bound transposase is now able to reposition it to a new location within the cellular genome. The transposase opens the genomic DNA backbone at the new locus and inserts the transposon fragment. The ligation of the open DNA ends is mediated by cellular key factors of the non-homologous end ligation pathway (NHEJ) within the double-strand break (DSB) repair system. The application of any of these transposon systems in the practice of this invention can be readily made based on the teachings provided in the art. See, for example, Fraser et al., Insect Mol Biol. 1996; 5:141–151; Ivics et al., Insect Mol Biol. 1996; 5:141–151; Kawakami et al., Gene. 1998; 225:17–22; Cadiñanos et al., Nucleic Acids Res. 2007; 35:e87; Muñoz-López and García-Pérez, Curr Genomics. 2010; 11:115–128; Grabundzija et al., Mol Ther. 2010; 18:1200–1209; Mátés et al., Curr Genomics. 2010; 11:115–128; Huang et al., Mol Ther. 2010; 18:1803–1813; Yusa et al., Proc Natl Acad Sci USA. 2011; 108:1531–1536; and Gogol-Döring et al., Mol Ther. 2016; 24:592–606.

[0061] In practice, the transposon system can be introduced into a transgenic animal system within the same AAV vector encoding the genetic perturbation. Alternatively, a dual-vector system can be used. In these embodiments, the coding sequence for the genetic perturbation (e.g., gRNA) is flanked by transposon elements in a first AAV vector, and a separate vector (e.g., a second AAV vector) encodes a transposase capable of recognizing the transposon sequence. In some embodiments of the invention, the DNA transposon system used comprises the transposase hypPB (e.g., highly active piggyBac) and a transposon sequence (e.g., an inverted terminal repeat sequence) recognized by the transposase. The PiggyBac (PB) transposon is a very useful transposon for genetic engineering in a wide variety of species. It can efficiently transposon between a vector and a chromosome via a "cut and paste" mechanism. During transposition, the PB transposase recognizes the transposon-specific inverted terminal repeat (ITR) sequences located at both ends of the transposon vector and efficiently moves the contents from the original site and integrates them into the TTAA chromosomal site. The PiggyBac transposon system enables the easy transfer of a target gene (e.g., a gRNA-encoding sequence) between two ITRs in a PB vector (e.g., an AAV vector exemplified herein) to the target genome. As demonstrated herein, incorporating this transposon system into an AAV-based perturb-seq platform overcomes the challenge of gRNA dilution and significantly improves in vivo genetic perturbation and labeling efficiency.

[0062] VI. Associated genes with cell phenotype or function

[0063] Once a library of AAV vectors encoding genetic perturbations (e.g., gRNAs) specific to one or more target genes is introduced into a transgenic animal or animal embryo, cells expressing the genetic perturbations can be identified and isolated from the living organism. The identified cells can be further enriched before undergoing examination of biological and cellular characteristics, including phenotypic traits. In several embodiments, examining the modified or altered biological and cellular characteristics of the perturbated cells may include measuring genomic, genetic, proteomic, epigenetic, and / or phenotypic differences in single cells. A specific genetic perturbation expressed in a single cell (e.g., gRNA targeting a specific gene) can then be correlated with biological characteristics (e.g., phenotypic changes) revealed from examinations of different cells. In some embodiments, phenotypic changes caused by genetic perturbations in the identified cells involve changes in gene or protein expression. In some embodiments, multiple target genes can be perturbed in a single cell, and gene expression in the perturbated cells can then be analyzed (e.g., via scRNA-seq). These methods allow for the identification of gene networks disrupted due to perturbations of single or characteristic target genes. Understanding perturbation-affected gene networks allows genes to be associated with specific pathways, which can then be targeted to regulate characteristic genes. Therefore, in some implementations, perturb-seq is used to discover novel gene and drug targets, enabling the treatment of a variety of diseases involving target genes.

[0064] Various tools can be used to examine cells that express genetic perturbations and exhibit changes in the expression of certain genes or proteins. In some embodiments, sequencing analysis can be used to identify cell types and regulated genes in the perturbated cells. In some of these embodiments, phenotypic changes caused by genetic perturbations introduced into the identified cells (e.g., editing of a specific target gene) and the correlation between the perturbations and the phenotypic changes observed in single cells can be analyzed using single-cell RNA sequencing (scRNA-seq). In some of these embodiments, the gRNA encoded by the AAV vector is functional for gene editing and can also be directly captured in scRNA-seq technology. This is to facilitate the association of observed phenotypes with specific perturbated genes in the identified cells. In some embodiments, scRNA-seq analysis aims to detect gRNAs in the identified cells by sequencing transcripts expressed from the AAV vector. In some embodiments, a given gRNA and a corresponding specific barcode are co-expressed from the same AAV vector, and the detection of the gRNA is achieved by detecting the specific barcode via RNA sequencing.

[0065] In some implementations, phenotypic analysis of the identified cells requires determining the cell type and corresponding perturbation (e.g., the gRNA sequence encoded by an AAV vector) by single-cell RNA (scRNA) sequencing of the identified and optionally further enriched perturbed cell population. scRNA sequencing of the cell population can be readily performed according to experimental protocols of conventional practice in the art. See, for example, Picelli et al., Nat. Protoc. 2014; 9: 171-181; Ramskold et al., Nat. Biotechnol. 2012; 30: 777-782; Hashimshony et al., Cell Rep. 2012; 2: 666-673; Kalisky et al., Ann. Rev. Genetics 2011; 45: 431-445; Kalisky et al., Nat. Methods 2011; 8: 311-314; Islam et al., Genome Res. 2011; 21: 1160-7; Tang et al., Nat. Protoc. 2010; 5: 516-535; and Tang et al., Nat. Methods 2009; 6: 377-382. In some preferred embodiments, phenotypic analysis of identified and enriched perturbed cells involves high-throughput scRNA sequencing. This can be performed using the specific techniques illustrated herein and / or experimental procedures known in the art. See, for example, Rosenberg et al., Science 2018; 360:176-182; Vitak et al., Nat. Methods 2017; 14: 302-308; Cao et al., Science 2017; 357:661-667; Gierahn et al., Nat. Methods 2017; 14: 395-398; Macosko et al., Cell 2015; 161: 1202-1214; Klein et al., Cell 2015; 161: 1187-1201; Zheng et al., Nat. Biotechnol. 2016, 34: 303-311; Zheng et al., Nat. Commun. 2017; 8: 14049; Zilionis et al., NatProtoc. 2017; 12: 44-73; and PCT Publications WO2016 / 040476, WO2016168584, and WO2014210353. In some embodiments, the phenotypic analysis of the identified and enriched perturbed cells involved can be performed by single-nuclear RNA sequencing. This can be done using standard protocols already reported in the literature.See, for example, Drokhlyansky et al., Cell 2020; 182: 1606–1622; Swiech et al., Nat. Biotechnol. 2014; 33: 102-106; Habib et al., Science 2016; 353: 925-928; Habib et al., Nat. Methods. 2017; 14: 955-958; and PCT publications WO2017 / 164936, WO2019 / 094984 and WO2020 / 077236.

[0066] Some embodiments of the present invention utilize the CRISPPR-Cas9 gene editing system to perturb target genes in transgenic systems expressing CRISPPR as described herein. This can be done using transgenic systems exemplified herein (e.g., Cas9 transgenic mice) and specific experimental protocols. Further guidance has been provided in the art for genome-scale screening using perturbations in single cells using CRISPR-Cas9.See, e.g., Dixit et al., “Perturb-Seq: DissectingMolecular Circuits with Scalable Single-Cell RNA Profiling of Pooled GeneticScreens” 2016, Cell 167, 1853-1866; Adamson et al., “A Multiplexed Single-CellCRISPR Screening Platform Enables Systematic Dissection of the UnfoldedProtein Response” 2016, Cell 167, 1867-1882; Feldman et al., Lentiviral co-packaging mitigates the effects of intermolecular recombination and multiple integrations in pooled genetic screens, bioRxiv 262121, doi: doi.org / 10.1101 / 262121; Datlinger, et al., 2017, Pooled CRISPR screening with single-celltranscriptome readout. Nature Methods. Volume 14 Issue 3 DOI: 10.1038 / nmeth.4177; Hill et al., On the design of CRISPR-based single cell molecular screens, NatMethods. 2018 April; 15(4): 271-274; Replogle et al., “Combinatorial single-cellCRISPR screens by direct guide RNA capture and targeted sequencing” NatBiotechnol (2020). doi.org / 10.1038 / s41587-020-0470-y; International Publication Serial No. WO / 2017 / 075294); and U.S. Patent Application Publication No. 2021 / 0172017. In the practice of this invention, any method described in the art may be employed and modified.

[0067] Example

[0068] The following embodiments are provided to further illustrate the invention, but do not limit its scope. Other variations of the invention will be apparent to those skilled in the art and are covered by the appended claims.

[0069] Example 1. Barcode-based AAV screening identifies serotypes that are effectively expressed in the developing brain.

[0070] Existing perturb-seq systems rely on lentiviral vectors, which have relatively low packaging yields and limited in vivo tissue penetration. AAV vectors are advantageous for in vivo studies due to their high production yields and capsid modification potential. However, most AAVs require several weeks to reach optimal expression; and no documented AAVs can effectively transduce newborn neurons and progenitor cells in the developing brain, requiring faster expression initiation. To identify AAV serotypes targeting the developing brain in vivo, we constructed a barcoded library of 86 phylogenetically distinct AAV serotypes; each expressing green fluorescent protein (GFP) and possessing a unique DNA barcode upstream of the polyadenylation initiation site (PDE). Figure 1 (Panel A). This library consists of previously reported or publicly available modified serotypes, including AAV1, AAV2, AAV3B, AAV4, AAV5, AAV6.2, AAV7, AAV8, AAV9, AAV10, AAV12, and AAV13; several of these serotypes are commonly used in in vivo studies and have minimal toxicity, as previously reported (Table 1).

[0071] Table 1. List and barcodes of AAV serotypes in the 86-AAV and 14-AAV libraries.

[0072]

[0073] To identify the rapid-acting serotype in the developing mouse brain (C57BL / 6J strain), we administered a library of this AAV vector intrauterinely into the lateral ventricle on embryonic day 13.5 (E13.5) (0.5 to 1.5 μL, 0.5 to 1.5 e). 10 (Viral genome / embryo). At 48 hours, we observed numerous GFPs in the cortex and ganglion ridges. + cell( Figure 1 (Sub-figures B to C). Using immunofluorescence markers, we observed 56% GFP. + Cells are newborn neurons (TBR1) that our AAV library can directly or indirectly transduce. + We found 3% GFP. +Cells expressed the intermediate progenitor (IP) marker TBR2, suggesting low efficacy of AAV transduction of IP, or rapid dilution of GFP during IP differentiation. Minimal GFP signal was detected in brain tissue using immunofluorescence amplification at 24 hours, indicating that most AAV vectors require 24 to 48 hours to express transgenes at detectable levels in vivo. Parallel to this in vivo analysis, we performed in vitro screening using the mouse hippocampal neuronal cell line HT22, and were able to detect GFP expression in vitro 24 hours after transduction, persisting until 48 hours.

[0074] Example 2. AAV-SCH9 rapidly transduces the developing brain within 48 hours.

[0075] To identify in GFP + AAV serotypes enriched in cells were purified from neocortical (in vivo) and HT22 cells (in vitro) at 24 and 48 hours post-transduction, and barcode abundance was quantified using next-generation sequencing. Figure 1 (Subgraph D). Compared to the initial distribution in the AAV library, we detected significant changes in the barcode distribution both in vivo and in vitro at 48 hours post-transduction. We found that at 48 hours post-transduction, several AAV serotypes were enriched in vitro (AAV-P1529, AAV2-NN, AAV2-P1583, and AAV-DJ), several serotypes were enriched in vivo (AAV1, AAV2-P1558, AAV-Hu48.2, and AAV2-7M8), and several serotypes were enriched in both (AAV-SCH9, AAV-SCH9 repeat 136bp, AAV2-P1576, AAV2-P1579, and AAV2-P1596). Figure 1 (See Figure D; Table 1). This may reflect a common pattern of AAV tropism for newborn neurons and progenitors in vivo and in vitro, as well as different patterns of enrichment in vivo.

[0076] To further examine the AAV barcode distribution pattern at different time points and under different environments, we performed principal component analysis on the samples. Figure 1 (Subplot E). PC1 explains 63% of the data variance and separates in vitro data (24 and 48 hours), in vivo data (48 hours), and in vivo data (24 hours) from the initial AAV library. PC2 (10% variance) separates the in vivo (48 hours) data from all other samples. Consistent with in vivo tissue analysis, changes in barcode distribution were observed in vivo after 48 hours, while the effects were observed more rapidly and consistently in vitro at 24 and 48 hours.

[0077] We observed an 11.5-fold increase in barcode expression in our top hit AAV-SCH9: it constituted 2.3% of the initial AAV library and 27.4% of the enriched clusters at 48 hours post-transduction in vivo. Differential expression analysis showed that AAV-SCH9 was the most significant hit from the in vivo experiment, followed by the serotype variant (AAV-SCH9-repeater 136bp). Both AAV-SCH9 and its variant showed significant enrichment in vitro at 24 hours post-transduction, which persisted at 48 hours. Figure 1 (Subgraph G). We also quantified the relative expression kinetics of several optimal hits, including AAV-SCH9, by comparing the abundance of barcodes at 24 and 48 hours after transduction. Figure 1 Subgraph F).

[0078] Crucially, several widely used AAVs targeting neurons require significantly longer to reach optimal expression compared to AAV-SCH9. Within the first 48 hours, AAV9-PHP.eB and AAV-DJ showed only minimal in vivo expression, with changes of 0.2-fold and 1.1-fold, respectively, relative to their composition in the initial AAV libraries (Table 1; and). Figure 1 (Figure D). In contrast, AAV-SCH9 showed an 11.5-fold increase in expression abundance, making it the best choice for perturbing and studying gene function in a dynamic developmental environment when cells are actively differentiating and maturing.

[0079] Example 3. AAV serotypes and their tropism in different cell types in vivo at single-cell resolution

[0080] We identified several serotypes, including AAV-SCH9, through batch measurements, which demonstrated efficient transduction in developing brain and neuronal cell lines; however, their precise cell type specificity remains unclear. To further characterize AAV hit tropism, we selected the top 14 serotype hits for validation in a new secondary library at single-cell resolution (Table 1). For these experiments, each AAV was programmed to express a GFP reporter with a set of three unique barcodes upstream of the polyadenylation initiation site, and these were subsequently merged at equal titers. Figure 2 (Figure A). We administered this 14-AAV library to the lateral ventricle of the brain of E13.5 mouse embryos and collected cortical cells at E15.5 for sorting (GFP). + Both the sorted and unsorted populations were subjected to droplet-based single-cell RNA-seq (scRNA-seq). Figure 2 (Subfigures A through B). We assigned AAV serotype identities using barcodes captured in the dial-out PCR library. Figure 2 (Figure C).

[0081] Following quality control, we retained a total of 14,835 neocortical cells for further analysis (7,630 cells from FACS-enriched GFP). + (Population, and 7,205 cells not enriched) Figure 2 (Figure B). We classified the cells into major cell types and annotated them based on the expression of known marker genes (Table 3) (Di Bella et al., 2021; LaManno et al., 2021; Tasic et al., 2018). These cells were clustered into 11 cell types, including upper and deep projection neurons (ULPN, DLPN), migrating neurons (Mig. neurons), apical progenitor cells (Api. prog), intermediate progenitor cells (IP), interneurons originating from the medial ganglion eminence (IN-MGE), interneurons originating from the non-medial ganglion eminence (IN-non-MGE), Cajal-Retzius cells (CR), fibroblasts (Fibro), mural cells, and microglia (Mg).

[0082] By checking GFP + The relative abundance of cell types in the unsorted population showed that most cell types were conserved. Two populations, apical progenitors and intermediate progenitors, were observed in GFP... + The reduced representativeness observed in the population reflected either a decreased ability of AAV to transduce dividing cells, or cell division diluting the transgene product, or both. By examining individual barcodes within the cell population, we identified four serotypes that were highly efficient at transduction in vivo: AAV-SCH9 (BC1), AAV2-NN (BC2), AAV-SCH9-G160D (BC4), and AAV2-7M8 (BC10). Figure 2 (Figure C). In the AAV-SCH9 transduced population, we detected 12.0% of upper projection neurons, 56.5% of deep projection neurons, 10.1% of MGE-derived interneurons, 7.9% of migrating neurons, 2.8% of apical progenitors and 1.3% of intermediate progenitors (Tables 2 to 3).

[0083] Table 2. Summary of scRNA-SEQ experiments (details of each batch of Perturb-seq)

[0084]

[0085] Table 3. Cell type classification and differentially expressed genes in 14-AAV library E16.5 scRNA-seq data

[0086]

[0087] Example 4. In situ characterization of AAV-SCH9 revealed its rapid dynamics and broad-spectrum effects in the developing brain. Pan-neuronal labeling

[0088] Using batch sequencing and scRNA-seq, we identified and validated AAV-SCH9 as an efficient vector for rapid (<48 hours) in vivo transduction of embryonic cortical tissue. Previously, AAV-SCH9 targeting subventricular adult neural stem cells has been reported (Ojala et al., 2018). To further characterize its whole-brain effects and tropism, we administered AAV-SCH9 expressing nuclear membrane anchoring fluorophore (GFP-KASH) intrauterine into the lateral ventricle at embryonic day 14.5 (E14.5). Forty-eight hours later, we observed abundant GFP in the neocortex, ganglion eminence, and around the lateral ventricle. + Cells, similar to the pattern from the 86-AAV library ( Figure 2 Subgraph D; and Figure 1 (Subgraphs B to C).

[0089] We co-stained the sections with markers of deep excitatory projection neurons (CTIP2 and TBR1) and intermediate progenitor cells (TBR2). We found an average of 37% GFP. + Cells expressed CTIP2 and 73% expressed TBR1, indicating that AAV-SCH9 can induce widespread expression in newly formed neurons. For GFP... + Cells, 11% expressing the intermediate progenitor marker TBR2, indicate lower transduction efficiency in progenitor cells and / or rapid dilution of expression with IP division. As expected, the relative proportions of cell types measured by immunohistochemistry and scRNA-seq data were generally consistent. These data confirm the in vivo performance of AAV-SCH9 in labeling neurons and intermediate progenitor cells within 48 hours.

[0090] AAV-SCH9 can directly transduce neurons, or first transduce progenitor cells, which then differentiate into neurons. To identify its tropism and differentiate between these two possibilities, we performed intrauterine transduction of AAV-SCH9-GFP-KASH at two embryonic ages when two distinct neuronal populations (deep and upper layers) were generated. If AAV only labels progenitor cells, we expected to observe enrichment of GFP expression in different layers, targeting different classes of newborn neurons on different injection days (deep layers only injected at E13, and upper layers only injected at E17, respectively). Brains were harvested after 48 hours and co-stained with the DLPN marker CTIP2 and the neural progenitor marker PAX6 to differentiate laminar layers. We divided the cortex into six blocks: cells transduced at E13.5 were enriched in blocks 2 (46%) and 3 (23%), exhibiting CTIP2 expression, indicating their deep cortical subcortical projection neuronal identity (Arlotta et al., 2005; Chen et al., 2008). In contrast, E17.5 transduced cells were widely distributed in the later-developing ULPN in block 1 (22%) and in layer 5 DLPN in block 2 (31%), supporting the ability of AAV-SCH9 to label newly generated neurons. Figure 2 (Figure E). Interestingly, we observed GFP in the ventricles and subventricular regions of the E17.5 transduced samples. + The levels of cellular markers were higher in samples transduced with E13.5 (18% vs. 9% in block 5 and 15% vs. 5% in block 6), supporting AAV-SCH9's progenitor cell marker capacity and consistent with its slower circulation and differentiation capacity in late neurogenesis (Greig et al., 2013). These data, along with scRNA-seq results, collectively support the rapid onset of AAV-SCH9 expression in both nascent neurons and progenitor cells transduced in the embryonic mouse brain.

[0091] One of the advantages of AAV is its excellent in vivo tissue penetration and labeling. In addition to the neocortex, AAV-SCH9 can transduce other regions, including the olfactory bulb, striatum, hippocampus, thalamus, and cerebellum, and the labeling density is suitable for future cross-regional perturb-seq studies. Figure 2 (Subplot F). In contrast, we observed limited expression in lentiviral transduction (using optimal >9 × 10⁻⁶). 9 (High-titer vector, U / mL): In most regions, especially the cerebellum, thalamus, and striatum, GFP... + Much fewer cells ( Figure 2(Figure F) This may be due to limited tissue penetration in vivo. In summary, this expanded perturb-seq has greater potential because it can acquire a large number of cells across brain regions, allowing for comprehensive comparison of brain region and cell type-specific perturbation effects of disease risk genes in vivo. Furthermore, AAV-SCH9 transduction outperforms conventional lentiviral vectors in labeling the neocortex and other brain regions, significantly simplifies the scRNA-seq sampling process, and reduces associated costs and labor.

[0092] Example 5. The transposase hypoPB stabilizes and enhances transgene expression.

[0093] Compared to lentiviral vectors, most AAVs remain free in host cells due to the lack of integrase (Deyle and Russell, 2009). Therefore, when using AAV vectors for perturb-seq in vivo in a rapidly dividing cellular environment, the expression of gRNA and fluorophore from AAVs will be diluted and eventually lost during cell division and growth. To overcome the challenge of gRNA dilution and improve in vivo genetic perturbation and labeling efficiency, we designed a dual-vector system containing a hypPB (highly active piggyBac) transposon and a transposon with inverted repeats flanking the gRNA and fluorophore (Moudgil et al., 2020). Figure 3 (Figure A). After transduction, hypoPB can integrate the transgene into the nuclear genome for inheritance in daughter cells, allowing consistent expression rather than transient, free expression. Figure 3 (Subgraph A).

[0094] We first tested the design in vitro and transfected HT22 cells, followed by time-lapse imaging. In the presence of hypPB, we detected more GFP. + cell( Figure 3 (Figure B) and enhanced GFP expression levels, with a faster onset; expression is resistant to cell passaging over a six-day period. hypPB allows for faster expression initiation within hours of transfection, which is highly advantageous for in vivo labeling and perturbing of cells during dynamic development.

[0095] Example 6. hypPB improves in vivo transgene expression through embryo transduction.

[0096] To evaluate whether hypoPB could enhance gene expression in vivo, we first administered transposon-containing AAV-SCH9 to E14.5 uterus with or without AAV-SCH9-hypPB, and performed immunofluorescence analysis at P7. We detected hypoPB (HA-labeled) expression in many brain regions, including the lamellar cortex, similar to the expected AAV-SCH9 tropism. Figure 3(Figure C). hypPB expression did not introduce significant toxicity during development ( Figure 3 (Figure C) because we measured changes in brain size and the expression of gliosis markers GFAP and IBA1 in the presence of hypoPB.

[0097] We quantified GFP in the somatosensory neocortex of brains administered AAV-transposons in utero, with or without hypoPB. + Cell number. Co-transduction of hypPB improved labeling efficiency: more than 2.2 times the number of cells were GFP-positive. + ( Figure 3 (Sub-figure F). Consistently, we divided the cortical laminar layer into four blocks and detected more than 3.0 times more GFP with hypPB co-expression in the upper layer. + Cells (15 cells / mm) 2 Relative to 5 cells / mm 2 (corresponding to block 1), 1.8 times was detected in layer 5 (108 cells / mm in block 3). 2 Relative to 59 cells / mm 2 ), and detected 3.2 times in layer 6 (123 cells / mm in block 4). 2 Relative to 39 cells / mm 2 This result is consistent with our in vitro studies, which showed that co-expression of hypPB leads to retention of expression in both dividing and differentiating cell lineages.

[0098] Example 7. hypPB enhances transgene expression in the adult central and peripheral nervous systems.

[0099] The efficacy of in vivo application of perturb-seq has been demonstrated in the embryonic brain in the past (Dvoretskova et al., 2023; Kim et al., 2020). However, there remains a need for a comprehensive in vivo screening platform for the adult central and peripheral nervous systems, particularly those associated with neurodegenerative diseases. We hypothesize that hypoPB can amplify AAV transgene expression in postmitotic neurons of the adult central and peripheral nervous systems, where relatively low levels of AAV transgene expression have been observed compared to genomic expression (Lang et al., 2019). To test this hypothesis, we delivered AAV vectors to postmitotic neurons of adult mice using two blood-brain barrier-crossing variants: AAV9-PHP.eB targeting neurons and glial cells in the brain, and AAV9-PHP.S targeting peripheral neurons, both administered non-invasively via the retroorbital route (Chan et al., 2017).

[0100] In the adult brain, hypPB typically increases GFP expression levels ( Figure 3(Figure D). In the somatosensory cortex, the average fluorescence intensity of the lamellar layer was 1.7 times higher in the presence of hypoPB (sub-figure D). Figure 3 (Figure G). Furthermore, by extracting nuclei from the cortex and performing flow cytometry analysis, we found that the AAV9-PHP.eB transposon system could label 6.5% to 7.0% of the total nuclei in the cortex, which is 4.0 times higher than in the absence of hypPB. This indicates that hypPB can enhance transgene expression in postmitotic cells in adults. Notably, unlike AAV-SCH9 delivery, elevated levels of gliosis markers IBA1 and GFAP were detected under retroorbital AAV9-PHP.eB-hypPB transduction conditions. This may be due to the systemic viral administration method and the amount of vector used (Perez et al., 2020) as well as the neuroinflammatory response induced by hypPB genome integration, highlighting the importance of carefully evaluating AAV administration to minimize toxicity issues.

[0101] To test the performance of this system in the adult peripheral nervous system, including the dorsal root ganglion, we performed retroorbital injections of the AAV9-PHP.S gRNA-tdTomato construct in the presence or absence of AAV-hypPB. Figure 3 Subgraph E). Similarly, the presence of hypPB also improves the report expression level and labeling efficiency ( , subgraph E). Figure 3 (Figure E). In summary, in combination with other AAV9 variants, the hypoPB transposon can further enhance transgene expression in the adult nervous system. This goes beyond the capabilities of conventional lentivirus-based genetic screening, providing opportunities to study genetic perturbations and gene function in adult tissues, especially in cases of aging or degeneration of the central and peripheral nervous systems.

[0102] Example 8. hypPB transposons integrate into the host genome via random insertion.

[0103] Transgenic integration via lentiviral vectors or transposons can induce undesirable changes in the genome and cellular activity. For example, integration in coding regions can significantly affect the function of nearby genes. To characterize transposon integration preferences in neurons in vivo, we analyzed over 170,000 tdTomato transposons from brains co-administered with AAV-SCH9-hypPB and transposons in utero. + The mouse genome was extracted from the nucleus. We then performed 60x paired-end whole-genome sequencing to identify hybrid reads that captured the connection between the mouse genome and transposons, providing evidence of transposon integration sites. Figure 3(Subgraphs H to I). We identified 46 such insertion events, several of which were supported by multiple reads and distributed across the mouse genome (Table 4). Overall, most insertion events occurred in intergenic regions (41%). We observed a slight preference for transcription start sites in insertions within promoter regions (…). Figure 3 (Figures H to I), consistent with literature on hypPB specificity (Chen et al., 2020). Integration events occurred at the expected TTAA flanking sites, further demonstrating the true in vivo insertion performance of hypPB. These analyses show that hypPB integration sites are largely randomized across the genome. The perturbation effects in each cell can be confounded by hypPB-induced genomic integration, a phenomenon also present in lentiviral vector-based screening. Because the events are largely randomized and located in intergenic regions, sampling a large number of cells for each perturbation group can minimize this bias, allowing for robust and efficient analyses.

[0104] Table 4. Whole-genome sequencing analysis of transposon insertion events

[0105]

[0106] Example 9. AAV-hypPB allows for the readout of captured gRNA using sparse scRNA-seq.

[0107] One challenge in in vivo CRISPR screening is maintaining high gRNA expression in sparse scRNA-seq for efficient gene editing and gRNA recovery (Kalamakis and Platt, 2023). We then tested whether transposase systems could also enhance gRNA expression levels and found that in vitro hypoPB co-transfection resulted in a 4.2-fold increase in gRNA expression. Interestingly, the presence of Cas9 stabilized gRNA expression by 5.7-fold, possibly through the formation of a ribonucleoprotein complex to prevent gRNA degradation (Hendel et al., 2015). With co-expression of hypoPB and Cas9, gRNA expression levels increased to 10.8-fold, likely through an additive mechanism, which could improve gene editing efficiency and gRNA detection in scRNA-seq readouts.

[0108] Using the fast-acting AAV-SCH9 and hypPB systems to enhance expression, we continued to characterize the gRNA identity recovery rate of this Perturb-seq platform using two established scRNA-seq strategies. Embryo-transduced cortical cells were dissociated, purified, and subsequently subjected to scRNA-seq after birth, capturing gene expression and gRNA identity using two approaches: 3' scRNA-seq, which uses bead-barcoded oligo-dT and gRNA primers to label and amplify the 3' ends; or 5' scRNA-seq, which uses in-solution oligo-dT and gRNA primers to capture transcripts from the 5' ends (Replogle et al., 2020). During reverse transcription, the 5' approach benefits from a higher concentration of gRNA-specific primers in solution (rather than on beads), which is expected to result in a higher gRNA capture rate.

[0109] We observed that at similar sequencing depths, 5' and 3' scRNA-seq gene expression analyses produced similar cell clustering and data quality (Table 5). However, for 3' scRNA-seq, we detected only 0.1% of reads allocated to gRNA from the library, while 5' scRNA-seq produced 52.9% of reads allocated to gRNA. This is consistent with the reliable allocation of gRNA to cell barcodes and the high detection rate of gRNA levels from 5' scRNA-seq. Furthermore, in 5' scRNA-seq, we observed a threshold that distinguishes cells with high gRNA expression from those with low gRNA expression, potentially differentiating true expression from environmental, spurious, or low-level background expression. Using in vivo AAV Perturb-seq, we reliably detected 63% to 83% of cell gRNA identity via 5' scRNA-seq, which is approximately 7 to 10 times higher than the 7% to 14% achieved via 3' chemistry. Figure 4 (Subgraph D; and Table 5).

[0110] Table 5. Cell type classification of 5' and 3' scRNA-seq data (parameters for cell type clustering in Seurat)

[0111]

[0112] Example 10. Effective in vivo cell labeling, high gRNA recovery, and validated CRISPR effects

[0113] We applied an AAV-based perturb-seq system to conduct a proof-of-principle in vivo screening of transcription factors that play a role in brain development: Foxg1, Nr2f1 (COUP-TFI), Tbr1, and Tcf4. Haploid deficiency of these genes is associated with neurodevelopmental disorders and diseases, including Bosch-Boonstra-Schaaf optic atrophy syndrome (Nr2f1), Pitt-Hopkins syndrome (Tcf4), FOXG1 syndrome (Foxg1), and TBR1 syndrome (Tbr1). These transcription factors play key roles in cortical neuronal differentiation and progenitor cell maintenance, regional pattern formation, neuronal migration, and circuit assembly by regulating networks of other transcription factors (Chen et al., 2021; Greig et al., 2013; Hou et al., 2020). Most of their expression spans from embryonic to postnatal stages: Tcf4 and Nr2f1 are widely expressed in most cell types, while Foxg1 and Tbr1 expression is limited to neuronal subclasses, including deep projection neurons and immature neurons (Di Bella et al., 2021; La Manno et al., 2021).

[0114] We designed four gRNAs for each gene by targeting the early 5' coding exon and tested them in vitro to select the three best-performing gRNAs for inclusion in the confluence (Table 6). We confluenced them on average with four control gRNAs (including a non-targeted (NT) control and a safe-targeted (ST) control gRNA) (Morgens et al., 2017). At E14.5, the AAV-SCH9 vector expressing hypPB and the confluenced gRNAs was administered to the lateral ventricle of the embryo. Figure 4 (Figure A). At P7, isomorphic cortical and hippocampal structures were microdissected and dissected; cells expressing BFP were enriched and subjected to droplet-based 5' scRNA-seq to directly capture gRNA (Replogle et al., 2020). In this experiment, we intentionally controlled the AAV injection titer to achieve <2% of cells transduced in the neocortex to limit multiple perturbation events; however, this ratio could be further increased if a combination of perturbations is desired. Notably, with this intentional dilution, our system has achieved more than 10-fold higher BFP levels than conventional lentiviral labeling. + The collection of disturbed cells significantly simplified the experiment.

[0115] Table 6. gRNA Sequences

[0116]

[0117] In five replicates (10x Chromium channels) from a total of 11 animals across two litters, we obtained a total of 50,075 cells for profiling analysis after primary quality control (Table 2). Notably, this is significantly more efficient than our previous lentivirus-based work (collecting 46,770 cells from 17 litters and 163 animals) (Jin et al., 2020), demonstrating the potential scalability of this new platform. All cells were classified into six general cell types and 21 subclusters, annotated by direct comparison with publicly available scRNA-seq data (Yao et al., 2021).

[0118] To accurately assign perturbation identities to cells, we evaluated several methods. We found that the most reliable method involved first downsampling the gRNA count in each cell to minimize bias due to count differences, followed by DemuxEM, a tool designed for demultiplexing hash barcodes (Gaublomme et al., 2019). We identified 16,067 cells assigned a single gRNA perturbation, 16,373 cells assigned multiple gRNAs, and 17,635 cells assigned no gRNAs, the majority of which were glial cells. The glial population was associated with low gRNA recovery and low BFP detection (from AAV transgene expression), consistent with our previous characterization of AAV-SCH9 tropism. Figure 2 Overall, we successfully assigned 57% to 77% of the cells to perturbations from non-glial cell types, with a median of 152 UMI gRNAs / cell and 9,687 UMI endogenous transcripts / cell detected in excitatory neurons. Figure 4 (Figure E). We further filtered for low-quality cells with low UMI or high percentage of intron reads, which may indicate that they are cytoplasmic, nuclear fragments or similar populations (La Manno et al., 2021).

[0119] Next, we narrowed down the downstream analysis to 11,471 high-quality cells, each with only a single gRNA perturbation. We retained 14 annotated cell types, following the previously described classification of homomorphic cortical and hippocampal structures (Yao et al., 2021). This included ten excitatory glutamatergic neuronal clusters from a total of 18 subclusters, including upper L2-4, L5 / 6IT (telebrain), Car3, PT (pyramidal tract), NP (proximal projection), CT (corticothalamus), and L6b; three GABAergic inhibitory neuronal clusters (Id2, Sst, and Lhx6+Sst-) and Cajal-Retzius cells (…). Figure 4(Subfigure B). We did not observe any significant batch effect for cell type in the five channels of the UMAP space ( Figure 4 (Figure C). Each gRNA was allocated to 318 to 1,065 cells, resulting in 1,609 to 2,527 perturbed cells and 3,316 control cells per gene.

[0120] To evaluate CRISPR loss of function and target gene dosage in vivo, we extracted endogenous transcript reads from 5' scRNA-seq libraries associated with cells receiving gRNA perturbations. We detected a significant proportion of reads carrying insertion or deletion events within each gRNA targeting region. Figure 4 (Figure F). For example, Foxg1 gRNA1 induced several distinct 1-base pair insertions in 39% of the reads, and frameshift mutations led to premature transcriptional termination downstream. This may be an underestimation of the perturbation effect, as nonsense-mediated decay should degrade most of the mutated mRNA. We hypothesize that different gRNAs targeting the same gene may produce variable phenotypic outcomes due to their varying efficiencies in gene editing, thus introducing potential phenotypic heterogeneity. We also observed that wild-type transcripts were detectable in our data, suggesting that in vivo perturb-seq can allow for the analysis of both complete knockout and loss-of-function heterozygous effects in different collected cells.

[0121] Example 11. Changes in the proportion of differentially expressed cell types caused by in vivo transcription factor perturbations

[0122] To analyze the changes in cell type proportions caused by each perturbation, we performed statistical tests that considered the differences in proportions across batches and gRNA identities compared to the non-targeted control group (NT-2) (Phipson et al., 2022). Figure 4 (Figure G). Since the variability of gRNA performance is known, we decided to perform this analysis at the gRNA level rather than the target gene level. Figure 4 (Figure F). We also included other control gRNA groups compared to the selected control group NT-2 to test the robustness of the analysis, as we expected little effect in comparisons between these controls.

[0123] Different gRNAs targeting the same gene often show similar trends, even if they do not all reach the same level of statistical significance. Figure 4 (Figure G). Notably, the proportions of excitatory and inhibitory neurons in the upper and deep layers all showed changes due to perturbations of Foxg1, Tbr1, and Tcf4. Among them, Tcf4 perturbation was associated with a 2.2-fold reduction in Sst interneurons, consistent with the high expression of Tcf4 in both the excitatory and inhibitory lineages during development.

[0124] As the most significant effect, Foxg1 function loss was associated with a 9.9-fold reduction in L6 IT neurons, accompanied by a 2-fold increase in upper-layer projection neurons. Figure 4 (Figure G). However, this perturbation did not significantly affect the proportion of L6 CT neurons residing in the same layer as L6 IT neurons. This result highlights the cell type-dependent role of Foxg1 in controlling neuronal cell fate in layer 6, with different roles in the two neuronal classes within the same layer and sharing the same developmental lineage.

[0125] Interestingly, we observed that Tbr1 perturbation was associated with a decrease in the proportion of deep excitatory neurons, including L6 CT (1.9-fold) and L6 IT (3.6-fold). This is consistent with previously reported known roles of Tbr1 in maintaining L6 neuronal identity (Bedogni et al., 2010; Fazel Darbandi et al., 2018), and our analysis provides a fine annotation of its cell-type-specific role. Furthermore, Tbr1 perturbation resulted in a 3.1-fold increase in L5 NP excitatory neurons, a role distinct from its function in L6.

[0126] Example 12. Effects of cellular environment-dependent perturbations on deep glutamatergic neurons

[0127] To further explore the effects of each perturbation at the molecular level, we performed differential expression (DE) analysis for each perturbation (gRNA) across cell types. For each cell type, cells containing the same gRNA were compared with cells containing the control gRNA (NT-2). We subsetted the data to include all perturbation cell type pairs with >50 cells, filtered out low-expression genes, and performed DE analysis using edgeR (McCarthy et al., 2012), while being cautious of false positive results (Soneson and Robinson, 2018).

[0128] First, we examined the number of significant DE genes in the cell type-perturbation combination ( Figure 4 (Figure H). As expected, the control group showed very little to no (0 to 1) significant DE gene expression compared to the NT-2 control. The four perturbation-cell cluster combinations showed significantly altered gene expression (Figure H). Ten notable DE genes supported by at least two gRNAs: Foxg1 in L5 IT and L6 CT neurons; Nr2f1 in L6 CT neurons; and Tbr1 in L6 CT neurons. Some of these changes are closely associated with established roles described in the literature: Foxg1 enhances L6 CT neuron identity (Liu et al., 2022b); loss of Tbr1 function leads to defects in L6 CT neurons, acquiring both L5 and L6-like identity and electrophysiological properties (Bedogni et al., 2010; Fazel Darbandi et al., 2018).

[0129] In late neurogenesis, Foxg1 plays a crucial role in maintaining L6 neuronal identity across neuronal cell types by regulating the transcription factor network (Liu et al., 2022b). Indeed, we found that Foxg1 loss of function affects several signaling pathways in L6 CT, including changes in the expression of adhesion and synaptic molecules (Stxbp6, Grin3a), axonal guidance molecules (Nrp1, Sema5b), and transcription factors (Bhlhe22, Lhx2). Figure 4 (Figure I). We identified upregulation of neuronal markers from nearby layers that should not be expressed in L6, including the L4 spinous stellate neuron marker Rorb4 (4-fold increase) and the L5 inferior / cortical spinal projection neuron markers Bcl11b and Tcerg1l (1.5-fold and 34-fold increases, respectively) in these L6 CT-perturbed neurons. Figure 4 (Figure I). These data suggest that in L6 CT postmitotic neurons, Foxg1 maintains its identity by actively inhibiting other transcription factors and alternative cell fates; and its loss of function causes the network to de-inhibit.

[0130] In addition, we detected subclusters of L6 CT neurons, which were mainly composed of cells carrying Foxg1 perturbations. Figure 4 (Figure B). This subcluster expresses several cell type or layer-specific markers (Tcerg1l, Lhx2), as well as a cluster-specific marker (Nkd1), indicating that it is not merely a cluster of doublet or empty droplets, but may indeed be a cell with hybridization fate. Proportional analysis showed that Foxg1 perturbation significantly increased the production of cluster 7, a strong effect that was statistically significant among all three independent gRNAs (2 to 24-fold increase, P-adj < 0.05). When Foxg1 failed to maintain L6 CT identity, the misregulation of transcription factor expression and the emergence of hybrid neuron subclusters together provided a high-resolution characterization of altered cell fate and its potential misregulation of transcription factors.

[0131] Importantly, we discovered distinct patterns of Foxg1 gene regulation across different cell types: the same perturbation induces different DE genes in different cell types with minimal overlap. Among all Foxg1-DE genes in four deep neuronal cell types, the only common target is Stxbp6 (synaptic fusion protein-binding protein 6), whose expression is upregulated after perturbation in L5PT, L5IT, L5PT, and L6CT neurons. Figure 4 (Figure I). However, we found more distinct rather than shared regulatory networks across different deep excitatory neuronal types. Most Foxg1 knockout-induced transcriptional factor changes were highly cell type specific. For example, the DE gene observed in L6 CT was largely absent in L5 IT, L5PT, or L5 NP cell types (Figure I). Figure 4 (Figure I). Furthermore, the Foxg1-Lhx2 interaction is known to be crucial for cortical hem formation; Foxg1 deficiency leads to decreased Lhx2 expression in progenitor cells during early embryogenesis at E9.5 (Chou and Tole, 2019). However, in the perinatal and early postnatal period, the time window we analyzed, Lhx2 expression was significantly increased (20-fold and 74-fold, respectively) in postmitotic L6 CT and L5 NP neurons after Foxg1 perturbation, while no significant changes were observed in L5IT or L5 PT neurons. Since our perturbator AAV was administered at E14.5 in this experiment, a few days after L6 neuron birth, this effect corresponds to a conditional knockout in postmitotic neurons, rather than knockout in progenitor cells and its previously reported effect at E9.5. These findings reveal the pleiotropic role of Foxg1 in gene regulation, which is spatiotemporally dynamic and has different molecular outcomes across different cell types and developmental time windows.

[0132] In summary, our analyses reveal a cell-type and developmental stage-specific role for Foxg1 in maintaining L6 CT neuronal characteristics by actively inhibiting alternative cell fates in postmitotic neurons. Foxg1 loss of function leads to the de-repression of several transcription factors, ultimately resulting in altered hybrid cell identity and the ability to synapse and assemble circuits, distinct from its other known roles in progenitor cells. These analyses, highly temporal, spatial, and cell-type specific, collectively reveal key molecular pathways by which Foxg1 coordinates cell fate determination and maturation across different neuronal cell types. In conclusion, our data demonstrate the potential and large-scale parallelizability of the in vivo Perturb-seq platform to extend the scale and depth of genetic analyses of highly cell-type-specific regulatory networks from intact tissues.

[0133] Example 13. Some exemplary experimental schemes and materials

[0134] The resources for some of the materials used in this study are listed in Table 7. Details of some of the experimental models and objects illustrated in this paper are described below.

[0135] C57BL / 6J, Cas9, and CD-1 mice:

[0136] All animal experiments were conducted according to protocols approved by the Institutional Animal Care and Use Committees (IACUC) of the Scripps Research Institute. E15 to P9 mice of different sexes and weights were used in scRNA-seq experiments, while mice ranging from E15 to adult size were used in immunohistochemistry experiments. All mice were kept under standard conditions (12-hour light / dark cycles with free access to food and water).

[0137] HT22 and HEK293FT cell lines:

[0138] Mammalian cell culture experiments were conducted using HT-22 mouse hippocampal neuronal cell lines (Millipore Sigma, #SCC129) or HEK293FT cell lines (Thermo Fisher Scientific, #R70007) cultured in DMEM (Thermo Fisher Scientific, #11965092) containing 25 mM high glucose, 1 mM sodium pyruvate, and 4 mM L-glutamine (Thermo Fisher Scientific, #11995073), supplemented with 1× penicillin-streptomycin (Thermo Fisher Scientific, #15140122) and 5% to 10% fetal bovine serum (Thermo Fisher Scientific, #16000069). HT-22 cells were maintained at less than 80% confluence, and HEK293FT cells were maintained at less than 90% confluence.

[0139] Method details:

[0140] Mammalian cell culture and time-lapse imaging: Unless otherwise specified, all transfections were performed in 96-well plates using PEI (Polysciences, #24765-1). Cells were seeded at approximately 10,000 cells / well 16–20 hours prior to transfection to ensure 50–60% confluence at transfection. For each well on the plate, the transfection plasmid was combined with OptiMEM I reduced serum medium (Thermo Fisher, #31985070) containing PEI to a total of 20 μl. This solution was added dropwise directly to the medium.

[0141] Cells were transfected and incubated in an In Cell 6000 analyzer (GE Healthcare) at 37°C and 5% CO2, and imaged using a 10x air objective. Images were acquired every 4 hours for 20 hours. Image filenames were blinded, and cell counts were performed; fluorescence was analyzed using a custom script available on GitHub.

[0142] RT-qPCR: HEK293FT cells were washed once with PBS, then trypsinized and resuspended in PBS supplemented with 0.04% BSA (NEB, #B9000S). Cells were purified by FACS at 4°C and collected in Trizol (Thermo Fisher Scientific, #10296010) at 5,000 to 50,000 cells / sample. RNA was extracted using the Zymo Direct-zol RNA MiniPrep Isolation Kit (Zymo Research, #R2052). RT-qPCR was performed using Maxima H Minus reverse transcriptase (Thermo Fisher Scientific, #EP0753) and Power SYBR Green PCR premix (MasterMix) (Applied Biosystems, #4367659), along with a homogeneous mixture of two RT primers: oligo(dT)18 primers (Thermo Fisher Scientific, #SO131) and direct capture primers. The 2-ΔΔCT method was used for the analysis of qPCR data and normalized relative to GAPDH expression.

[0143] AAV Vector Construction and Generation: Viral vectors and plasmids were constructed as previously reported (Jin et al., 2020). The backbone plasmid contained a human U6 promoter to express a gRNA and an EF1α promoter to express a fluorescent protein conjugated to the nuclear membrane localization domain KASH. Vector cloning was performed individually and confirmed by Sanger sequencing. The gRNA design was defined using online tools at benchling.com, and the complete sequences of the gRNAs used in this work are listed in Table 6. AAV generation and titration were performed at the Viral Vector Core Facility at Sanford Burnham Prebys and the viral core at the UCI Center for Neural Circuit Mapping.

[0144] AAV administration: In CD1, C57BL / 6J, or Cas9 transgenic mice (Jax#026179) (Platt et al., 2014), AAV (0.5 to 1.5 μL / embryo) was administered intrauterinely to the lateral ventricle at E13.5 to 17.5 for immunohistochemical analysis and scRNA-seq. Adult mice were injected retroorbitally with AAV (50 to 100 μL and approximately 1 to 4 e-). 11 (One viral genome / animal), and perfused for immunohistochemical experiments and nuclear flow cytometry.

[0145] AAV library barcode extraction: HT-22 cells and mouse primary cortical cells were purified using FACS 24 or 48 hours after transduction. Genomic DNA was extracted from approximately 3,000 purified cells using QuickExtract DNA Extraction Solution (Lucigen, #NC0302740) according to the manufacturer's protocol. The AAV serotype library was lysed by DNase I and proteinase K digestion. PCR was performed on the genomic DNA using NEBNext High-Fidelity 2X PCR premix (New England BioLabs, #M0541L) and the following primers: The amplicons were amplified to include the adaptor and sequenced on an iSeq 1000 or MiSeq platform (>2 million reads / sample). The BCL files were converted to FASTQ files using bcl2fastq (Illumina).

[0146] Immunofluorescence staining of brain sections and whole-mount DRG: Embryonic brains were harvested directly after decapitation and immediately frozen on dry ice in an OCT chamber. Serial sets of 15-20 μm tissue sections were prepared on a cryostat and then fixed on ice with 4% paraformaldehyde in PBS for 15 minutes. Newborn and adult mice were anesthetized and perfused cardiacally with ice-cold PBS, followed by perfusion with ice-cold 4% paraformaldehyde in PBS. Dissected brains were post-fixed overnight in 4% paraformaldehyde at 4°C. Newborn and adult brains were embedded in 2% agar, and 60-100 μm tissue sections were collected using a vibratory microtome.

[0147] Slides containing mounted tissue sections were washed four times with PBS containing 0.3% Triton X-100 and incubated for 2 hours at room temperature with blocking medium (10% donkey serum in 0.3% Triton X-100 containing PBS (Sigma Aldrich, #S30-100ML)). Then, they were incubated overnight at 4°C with the primary antibody in the blocking medium. Slides were washed four times with PBS containing 0.3% Triton X-100. Secondary antibody was applied to the blocking medium at a 1:1000 dilution and incubated for 2 hours at room temperature. Slides were then washed four times with PBS containing 0.3% Triton X-100 and incubated with DAPI for 10 minutes, followed by mounting with antifluorescence quenching mounting medium (Vector Laboratories, #H-1700-10). All images were taken using a Nikon AX confocal microscope with either a 10x or 20x air objective.

[0148] The primary antibodies and diluents were: chicken anti-GFP antibody (ab16901, 1:500; Millipore), rabbit anti-GFP antibody (A-11122, 1:500; Invitrogen), rabbit anti-RFP (600-401-379, 1:500; Rockland), rabbit anti-Tbr1 (ab31940, 1:500; Abcam), rabbit anti-Tbr2 (ab183991, 1:500; Abcam), and rat anti-Ctip2 (a b18465, 1:1000, Abcam), rabbit anti-Pax6 (catalog number 901302, 1:500, BioLegend), rabbit anti-HA tag (5017, 1:500; CellSignaling), rat anti-HA tag (11867423001, 1:500, Roche), chicken anti-GFAP (ab4674, 1:500, Abcam) and goat anti-Iba1 (ab5076, 1:500, Abcam).

[0149] Mouse dorsal root ganglia (DRGs) were extracted and post-fixed in 4% PFA for 1 hour, then washed with PBS. The DRGs were then mounted onto silicone isolators using EasyIndex.

[0150] Nuclear isolation and FACS enrichment for genomic analysis: For whole-genome sequencing, >170,000 nuclei were sorted into a DNA / RNA Shield (Zymo Research, catalog number R1100-250) and treated with a Quick DNA Microprep kit (Zymo Research, catalog number D3020) for DNA purification to isolate >500 ng of genomic DNA. The genomic DNA was sequenced to 60x coverage using paired-end (150-150 base pairs) Illumina sequencing (Novogene).

[0151] Tissue dissociation and FACS enrichment for genomic analysis: Tissue dissociation was performed using a modified version of the previously described protocol (Jin et al., 2020) of the papain dissociation kit (Worthington, #LK003150). Briefly, young mice were anesthetized, then sterilized with 70% ethanol and decapitated. The brain was rapidly extracted and gently wiped with PBS-soaked Kimwipe (Kimberly-Clark) to remove meninges and fibroblasts. The cortex was microdissected under a dissecting microscope in ice-cold dissecting medium (Hibernate A medium (Thermo Fisher Scientific, #A1247501)) containing B27 supplement (Thermo Fisher Scientific, #17504044) and trehalose (Sigma Aldrich, catalog number T9531). The microdissected cortex was transferred to a papain solution containing DNase in a cell culture dish and cut into small pieces with a blade. The dish was then placed on a digital shaker in a cell culture incubator and shaken at 30 rpm for 30 minutes at 37°C. The digested tissue was collected into 15 mL tubes and ground 20 times with a 10 mL low-binding plastic pipette. The cell suspension was carefully transferred to a new 15 mL tube. 2.7 mL of EBSS, 3 mL of reconstituted Worthington inhibitor solution, and DNase solution were added to the 15 mL tubes and gently mixed. The cells were pelleted by centrifugation at 300 g for 5 minutes at 4°C, followed by centrifugation at 200 g for 5 minutes at 4°C. The cells were washed with 8 ml of cold dissection medium for 5 minutes. The cells were then resuspended in 0.5 ml of ice-cold dissection medium containing 10% fetal bovine serum (FBS) (Thermo Fisher Scientific, #16000069) and live / dead cell dye, and purified by FACS.

[0152] Immediately after collection, cells were centrifuged and resuspended in ice-cold PBS containing 0.04% BSA (NEB, #B9000S). Each 10x scRNA-seq library was prepared by combining FACS-sorted cells from 1 to 2 litters (5 to 8 animals) of E15 or P7 to 9 animals harvested on the same day (Table 2). Dissociation, FACS purification, and resuspending were performed over 3 hours while keeping cells on ice to prevent necrosis.

[0153] Following the manufacturer's instructions, construct scRNA-seq libraries using either the Chromium Next GEM Single Cell 3' Solution v3.1 kit with Feature Barcode Technology or the Chromium Next GEM Single Cell 5' Solution v2 kit (10x Genomics) with Feature Barcode Technology. Sequencing the gene expression libraries was performed using the NextSeq500 High Output 75 Cycle Kit (Illumina) at a sequencing saturation of greater than 20,000 reads per cell (R1: 26 base pairs, R2: 46 base pairs). CRISPR gRNA screening libraries were sequenced using the Illumina iSeq100 300-cycle kit (R1: 151 base pairs, R2: 151 base pairs) and the Nextseq500 150-cycle kit (Illumina) (R1: 73 base pairs, R2: 74 base pairs).

[0154] AAV barcode enrichment from scRNA-seq libraries: Following whole transcriptome amplification (WTA) in 10x Chromium library construction, a portion of the WTA product was used to amplify AAV serotype and cellular barcodes using a pull-out PCR strategy. In short, using AAVlib-pull-out-NSG1: A 10 ng WTA library was amplified, followed by 11 cycles of PCR and then cleaned with 1X SPRI beads. The samples were amplified for another 24 cycles using 8 bp indexed PCR primers and then purified by gel electrophoresis. The final extracted library was sequenced along with the transcriptome library using the NextSeq500 High Output 75 Cycle Kit (Illumina) flow cell (R1: 26 base pairs, R2: 46 base pairs).

[0155] Quantitative and statistical analysis: All images were analyzed using ImageJ (NIH), Photoshop (Adobe), and Illustrator (Adobe). Cells were manually counted from blinded files using the ImageJ CellCount function.

[0156] AAV barcode analysis in primary serotype screening: A custom script is used to map FASTQ files from the Illumina library to AAV barcodes. In short, FASTQs starting with the correct initial primer sequence are retained. Among these sequences, the barcode sequence following the initial primer sequence is compared to our list of AAV barcodes and assigned to matching AAV barcodes with a Levenshtein distance less than 2.

[0157] The barcode count matrix was analyzed using DESeq2 v1.40.2 (Love et al., 2014). The DESeqDataSetFromMatrix command was used to create the DESeq object, with the library counts used as the reference level. The results for each condition were tabulated using α=0.05 as the significance threshold. Volcano plots were generated using the R package EnhancedVolcano v1.18.0, and heatmaps were generated using pheatmap v1.0.12.

[0158] Transposon integration site analysis: A custom reference genome was created by attaching a transposon reporter plasmid as an additional chromosome to the mm39 mouse genome. Using the SP5M flag, the FASTQ file was aligned to the custom genome using bwa mem (v0.7.17) (Li and Durbin, 2010). Reads aligned to the piggybac plasmid and their pairings in the resulting files were filtered using SAMtools™ v1.15.1. After examining the distribution of these reads along the plasmid, they were further filtered to reduce them to reads aligned to the ends of the plasmid inserts (positions 1653–2053 and 6096–6496). The filtered files were then parsed using Pairtools™ v0.3.0 with default settings. Their connection points (pair type UU, UR, or RU, with one end aligned to the mouse genome) were then filtered, followed by manual verification.

[0159] To annotate integration sites, the `annotatePeak` function from the R package `ChIPseeker™` (v1.36.0) (Yu et al., 2015) was used, with the TSS set to + / - 3,000 bp. The R package `Circlize v0.4.15` was used to create a Circos plot to visualize the integration sites.

[0160] scRNA-seq data processing: Starting with the raw Illumina BCL file of the transcriptome library, a FASTQ file was generated using the "cellranger mkfastq" command (from Cell Ranger package v7.1.0) (Zheng et al., 2017) with default parameters. bcl2fastq was used to demultiplex the gRNA library or AAV barcode-extracted library. The "cellrangercount" command was used to align transcriptome reads to the mouse genome reference mm10 (GENCODE vM23 / Ensembl 98), and a gene expression count matrix was generated using the expected cell count (expect-cells) = 9,000. AAV barcodes or gRNA reads were quantified at the single-cell level using characteristic reference markers in Cell Ranger.

[0161] Cell type classification and cell identity annotation: For secondary screening of AAV serotypes and comparison of 5' vs 3' scRNA-seq, the filtered count matrix from Cell Ranger was loaded into R v4.3.0 using the Read10X command from Seurat v4.3.0.9003 (Hao et al., 2021), and then loaded into a Seurat object using CreateSeuratObject, filtering out cells with <500 genes or mitochondrial count percentage >25%. The data were log-normalized, and the 2,000 most variable features were selected using FindVariableFeatures. In each analysis, the two conditions (sorted vs. unsorted, 5' vs. 3') were integrated using IntegrateData based on consistent variable features across datasets. The integrated data were scaled using ScaleData, and PCA was performed using RunPCA. UMAPs were generated using RunUMAP and clustered using FindNeighbors and FindClusters (using default parameters, except that the resolution was 0.3 for sorted vs. unsorted, or 0.2 for 5' vs. 3', except that dims=1:25). Clusters were assigned to cell types based on known markers (Di Bella et al., 2021; La Manno et al., 2021; Tasic et al., 2018). In AAV serotyping, non-cortical cell clusters were removed, while in 5p vs. 3p analysis, cell clusters with low gRNA mapping were removed. The data were re-clustered after removal (dims=1:28, resolution 0.3 for AAV serotyping library, or 0.2 for 5' vs. 3'). Cells with a detection >5 UMI were assigned either an AAV barcode or gRNA identity.

[0162] Perturb-seq Data Processing: Cell Type Classification, Cell Identity Annotation, and Perturbation Identity Annotation: The filtered count matrix from the Cell Ranger was loaded into Rv4.0.3 using the Read10X command from Seurat v4.0.0 (Hao et al., 2021), and then loaded into a Seurat object using CreateSeuratObject to filter out cells with <500 genes. The data was normalized using the NormalizeData command, with normalization.method="LogNormalize" and scale.factor=10,000, and variable genes were selected using FindVariableFeatures. Two-cell scores for each 10x channel were calculated using scds v1.6.0 (Bais and Kostka, 2020). UMI count data for variable genes were extracted and used as input to scGBM v0.1.0 (Phillip and Jeffrey, 2023), where M=20 and subset=30000, generating 20-dimensional embeddings of the count data. These 20-dimensional embeddings were added to the Seurat object. The UMAP was calculated based on this dimensionality reduction using RunUMAP, and clustering of the dimensionality reduction was performed using FindNeighbors and FindClusters, with other settings remaining at default. The UMI count matrix of gRNAs generated by Cell Ranger was also added to the Seurat object as an additional assay.

[0163] The quality control (QC) metric for each channel was calculated using the CellLevel_QC tool (Github) and loaded into the metadata of the Seurat object. The mitochondrial read percentage for each cell was also calculated. Initial annotation of the data was generated using single-cell references from the Allen Brain Atlas (Yao et al., 2021) with Azimuth v0.3.2. Clusters with high percentages of intron reads or high double-cell scores, as well as cells with >20% intron reads or >10% mitochondrial reads, were removed. Clusters were then labeled as follows: using cell type labels from Azimuth and by comparing DE genes from our dataset with DE genes from the Allen Brain Atlas dataset (DE genes between clusters were calculated using Presto v1.0.0).

[0164] Next, we constructed a Nextflow pipeline (DSL v2) (Di Tommaso et al., 2017) that uses CellRanger output to extract guide RNA data into CSV files, assigns cells to guides using DemuxEM (v0.1.5) (Gaublomme et al., 2019), and extracts the resulting perturbation markers as TSV files (where PegasusIO v0.2.11 (Li et al., 2020) is used for loading data and pandas v1.2.4 (pandas development team, 2023) is used for saving data to create the TSV and CSV files). These tables are then loaded into Seurat object metadata. We also constructed a version of this pipeline that allows downsampling (the downsampling method mentioned in the text) before running DemuxEM. Specifically, for each 10x channel and each gRNA, if more than 50,000 UMIs are assigned to that gRNA, the pipeline downsamples the gRNA UMI count for that gRNA using a binomial function from numpy.random (Harris et al., 2020), where the proportion used for downsampling is equal to 50,000 divided by the total number of UMIs from the gRNA (producing the expected value of 50,000 UMIs / gRNA / channel after downsampling). Similarly, we constructed a Nextflow pipeline to extract reads mapped to BFP or GFP sequences. This pipeline accepts a BAM file generated by Cell Ranger, extracts unmapped reads using SAMtools v1.8 (using the command "samtoolsview-f 4-b") (Li et al., 2009), extracts UMI and CBCs using UMI-tools v1.0.1 (Smith et al., 2017), maps these reads to GFP and BFP sequences (including 5' and 3' UTR regions) using Minimap2 v2.11 (using the parameter -ax sr) (Li, 2018) (Li, 2018), and converts them into a BAM file using SAM-tools view-b. Unmapped reads are discarded, and the contig names (GFP or BFP) mapped to the remaining mapped reads are added to the BAM file as additional tags (XT tags) using awk. The resulting BAM files are sorted and indexed, and the number of UMIs mapped to GFP and BFP in each cell is extracted using the command `umi_tools count --per-gene --gene-tag=XT --per-cell`. These counts are then loaded into the Seurat object.

[0165] To accurately assign perturbation identities to cells, we compared three methods: built-in gRNA allocation from the Cell Ranger; a tool for demultiplexing hash barcodes called DemuxEM (Gaublomme et al., 2019); and DemuxEM after downsampling the gRNA count. We found all three methods performed well, although the evidence we observed suggests that the standard DemuxEM method leads to biases in cell type proportions due to significant differences in the number of recovered gRNA UMIs / guidelines. For example, we observed a strong correlation between the average number of gRNA UMIs recovered from a given gRNA and the proportion of Cajal Retzius cells assigned to that gRNA (Spearman correlation = 0.84, P = 2.934e-05); this effect disappeared after downsampling (Spearman correlation = -0.03, P = 0.89). Cell Ranger does show evidence of this bias, but it does report a higher number of double cells than expected, particularly in the Cajal Retzius cluster (where 71% of high-quality Cajal Retzius cells were assigned as double cells in Cell Ranger, compared to 60% in the downsampling method and 52% in the standard DemuxEM method). We performed a full analysis using DemuxEM versus downsampling.

[0166] Non-neuronal cells (cells not labeled as excitatory, inhibitory, or Cajal-Retzius neurons) and cells with <3% intron reads were removed from the Seurat object. Excitatory and inhibitory cells with <3000 genes and Cajal-Retzius cells with <2000 genes were filtered out. Finally, all cells not assigned exactly one guide by DemuxEM were also removed. This Seurat object was used for downstream analysis.

[0167] Detection of insertions and deletions in target regions: For each target gene, reads overlapping with that target gene in each 10x channel were extracted from the CellRanger BAM file using SAM-tools view and combined into a single BAM file using SAM-toolsconcat. Sorting and indexing were then performed using SAM-tools. The BAM file was split into a single BAM file for each perturbation using the sinto filterbarcode command (Stuart et al., 2021), consisting of cells assigned to each gRNA in the final analysis (excluding cell barcodes appearing in multiple 10x channels). The resulting BAM files were indexed. This produced a single BAM file for each target gene and each gRNA, consisting of reads overlapping with the target gene in cells assigned to the gRNA / perturbation identity.

[0168] For each target region, we counted the number of reads that overlapped with each perturbation and the number of reads with insertions or deletions (insertions or deletions) within that region. This was done using a Pysam-based Python script (Github). The script loaded each read one by one from a BAM file. For each position within each read, it used `get_reference_positions` to: 1) check if the position was within the desired range, 2) if it was within the range, check if the position was an insertion (excluding insertions at the beginning / end of the read to avoid soft-shear regions), and 3) if it was within the range and did not contain an insertion, check if the position was a deletion (if there was a difference of >1 bp between the mapping positions of the current and previous bases). The results were then tabulated into a pandas dataframe, returned, saved, and loaded into R for downstream analysis.

[0169] Perturbation-related analysis: After assigning gRNAs and cell types to all cells, we observed changes in cell type ratios among gRNAs. For this purpose, we selected a control group (non-targeted control 2) and compared each of the perturbation / control groups containing other gRNAs to this group. Statistics for these pairwise compositional comparisons were calculated using the propeller.ttest function from speckle (R package v0.99.7) (Phipson et al., 2022). Here, batches (10x channels) were treated as another fixed effect in the linear model. Cell types and clusters with fewer than 200 cells in total (Excit_Car3 and Inhib_Id2 for cell type-level comparisons and clusters 5, 25, and 26 for cluster-level comparisons) were excluded from this analysis. Results were collected and visualized together using ComplexHeatmap (R package v2.14.0) (Gu et al., 2016). For this visualization, effect sizes were limited to an absolute value of 2.

[0170] Next, we identified differentially expressed genes for each gRNA and for each cell type. Using controls (gRNA non-targeted control 2), we compared each gRNA group population to the control. Statistics were calculated using the pseudo-batch method with edgeR (R package v3.40.2) (McCarthy et al., 2012). Low-expression genes (<10 supporting reads and <.1 normalized expression) were identified in each comparison and excluded from the DE test. Additionally, we applied alternative variable analysis (sva; R package v3.46.0) (Leek and Storey, 2007) to capture potential batch or technical artifact-related variables from the data and added them to the model test for DE (n.sv=1). Calculated p-values ​​were corrected using the Benjamini-Hochberg method.

[0171] Table 7. List of resources for some materials used in the study

[0172]

[0173] Some cited references

[0174]

[0175]

[0176]

[0177]

[0178]

[0179]

[0180]

[0181]

[0182]

[0183]

[0184]

[0185] Although the foregoing invention has been described in detail by way of example and embodiments for the purpose of clarity, it will be apparent to those skilled in the art that certain changes and modifications may be made to the invention without departing from the spirit or scope of the appended claims, based on the teachings of the invention.

[0186] All publications, databases, GenBank sequences, patents and patent applications cited in this specification are incorporated herein by reference as if each were expressly and individually indicated to be incorporated by reference.

Claims

1. A method for analyzing the function of multiple target genes in one or more specific cell types in vivo, comprising (1) introducing a library of an AAV vector encoding a genetic perturbation of the multiple target genes into a transgenic system expressing CRISPR-Cas9, (2) identifying one or more cells of a specific cell type expressing the genetic perturbation and exhibiting a specific phenotype from the transgenic system, and (3) determining (a) the genetic perturbation encoded by the AAV vector introduced into each of the one or more cells and (b) the corresponding perturbed gene, thereby associating the corresponding perturbed gene with the phenotype exhibited by each of the one or more cells.

2. The method of claim 1, wherein the transgenic system is a developing embryo of a transgenic animal expressing Cas9, and wherein the library of the AAV vector contains a vector with serotype AAV-SCH9 or AAV2.NN.

3. The method of claim 2, wherein the AAV vector library comprises an AAV-pRep2-SCH9 vector or an AAV-pRep2-SCH9 (136bp repeat) vector.

4. The method of claim 2, wherein the AAV vector library is administered to the animal embryo at approximately the developmental stage corresponding to mouse embryonic day 11.5 (E11.5), 12.5 (E12.5), 13.5 (E13.5), 14.5 (E14.5), or 15.5 (E15.5).

5. The method of claim 2, wherein the library of the AAV vector is administered intrauterinely to the brain of the embryo.

6. The method of claim 1, wherein the transgenic system is a newborn or adult transgenic animal expressing Cas9, and wherein the library of the AAV vector contains a vector with serotype AAV-PHP.eB or AAV.PHP.S.

7. The method of claim 6, wherein the library of the AAV vector is administered to the animal at an age corresponding to the age of the mouse from about 10 days after birth (P10) to about 18 months of age.

8. The method of claim 6, wherein the library of the AAV vector is administered retro-orbitally to an animal.

9. The method of claim 1, wherein the specific cell type is a newly generated neuron or its progenitor cell obtained from brain tissue.

10. The method of claim 9, wherein the brain tissue is derived from the neocortex, olfactory bulb, striatum, hippocampus, thalamus, or cerebellum.

11. The method of claim 9, wherein the target gene is known or suspected to be associated with a neurodevelopmental disorder.

12. The method of claim 1, wherein each of the genetic perturbations comprises one or more guide RNAs (gRNAs) for introducing genomic alterations in the coding region of a target gene via CRISPR-Cas9 gene editing.

13. The method of claim 12, wherein the one or more gRNAs each target a sequence at the 5' end of the coding region of a target gene.

14. The method of claim 12, wherein the AAV vector further encodes a transposon element flanked by the gRNA, and wherein a library of the AAV vector is introduced into the transgenic system in the presence of a transposase, the transposase recognizing the transposon element and inserting it into the host genome.

15. The method of claim 14, wherein the transposase is encoded by a second AAV vector introduced into the transgenic system together with a library of the AAV vector.

16. The method of claim 14, wherein the transposase is a highly active transposase piggybac (hypPB).

17. The method of claim 1, further comprising enriching one or more cells identified as expressing one or more genetic perturbations and exhibiting a particular phenotype.

18. The method of claim 12, wherein the coding sequence of one or more gRNAs is operatively linked to a reporter gene in an AAV vector.

19. The method of claim 18, wherein one or more cells expressing a genetic perturbation are identified by detecting the expression of the reporter gene.

20. The method of claim 18, wherein the genetic perturbation encoded by the AAV vector introduced into each of the one or more cells is determined by capturing the one or more gRNAs using single-cell RNA sequencing.

21. The method of claim 18, wherein the one or more gRNAs are functional for gene editing and can be directly captured in scRNA-seq.

22. A method for analyzing the function of multiple target genes in one or more specific cell types in vivo, comprising: (1) A library of AAV vectors encoding (i) a transposase and (ii) AAV vectors each encoding (a) one or more guide RNAs (gRNAs) targeting one of the target genes and (b) transposable elements recognized by the transposase and flanked by the gRNAs are co-introduced into the developing embryo of a transgenic animal expressing Cas9; wherein the AAV vectors have serotypes AAV-SCH9 or AAV2.NN (2) Identify from the embryo one or more cells of a specific cell type expressing the gRNA and exhibiting a specific phenotype, and (3) Identify the gRNA encoded by each of the AAV vectors introduced into the one or more cells and the corresponding perturbed gene; thereby associating the specific phenotype with the function of each of the perturbed target genes.

23. The method of claim 22, wherein the transposase is hypoPB.