AAV evolution at single cell resolution using SPLIT-SEQ
The problem of AAV transduction detection at the single-cell level is solved by delivering barcoded RNA expression constructs in AAV vectors, combining with RNA polymerase III promoter and modified AAV capsid proteins, and achieving accurate identification of cell types and high-throughput analysis.
Patent Information
- Application Number
- CN202380073433.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-19
- Filing Date
- 2023-09-19
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to simultaneously detect transduction of AAV capsids at the single-cell level and provide cellular identity information, especially detection of mRNA sequences and delivered AAV capsid DNA or expressed RNA.
Barcoded RNA expression constructs are provided, delivered to cells through AAV vectors, express mRNA sequences with marker barcodes, and bind to the RNA polymerase III promoter, using a population of recombinant adeno-associated viral vectors, containing modified AAV capsid proteins targeting peptides, to achieve the identification of cell tropism.
Accurate identification of AAV transduction at the single-cell level, enabling the identification of transduction of specific cell types, provides a high-throughput method to detect and analyze the cellular tropism of AAV.
Smart Images

Figure BDA0005361777460000131 
Figure BDA0005361777460000141 
Figure BDA0005361777460000351
Abstract
Description
Citation of Related Applications
[0001] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 407,826, filed on September 19, 2022, the entire content of which is incorporated herein by reference. Citation of Sequence Listing
[0002] This application contains a Sequence Listing XML that has been submitted electronically and is hereby incorporated by reference in its entirety. The Sequence Listing XML, created on September 19, 2023, is named CHOPP0057WO_ST26.xml and is 28,221 bytes in size. Background Art 1. Technical Field
[0003] The present invention generally relates to the fields of molecular biology, virology, and medicine. More specifically, it relates to compositions and methods for determining the cellular tropism of AAV capsid proteins with targeting peptides. 2. Description of Related Art
[0004] One of the major challenges in detecting the transduction of barcoded AAV capsids at the single-cell level is detecting the mRNA sequences that provide information about cell identity and simultaneously detecting both the delivered AAV capsid DNA and the expressed RNA. Methods are needed that allow the identification of the specific cell types transduced by a given barcoded AAV capsid. Summary of the Invention
[0005] Barcoded RNA expression constructs are provided herein that, when packaged into AAVs and delivered to cells, express mRNA sequences having an identifying barcode that is distinct from the modified region of the capsid DNA sequence. In one embodiment, a recombinant adeno-associated virus (rAAV) vector is provided that contains an expression cassette encoding a barcode sequence operably linked to an RNA polymerase III promoter. In one embodiment, a population of recombinant adeno-associated virus (rAAV) vectors, each rAAV vector containing an expression cassette encoding a barcode sequence operably linked to an RNA polymerase III promoter. The population of vectors can contain a ratio of barcode sequences to vectors of 1:1, 1:5, 1:10, 1:50, 1:100, 1:500, or at least 1:1000. In one embodiment, a population of recombinant adeno-associated virus (rAAV) vectors is provided herein, wherein each rAAV vector independently contains (i) a modified adeno-associated virus (AAV) Cap gene that encodes a modified AAV capsid protein containing a targeting peptide, and (ii) an expression cassette that encodes a barcode sequence operably linked to an RNA polymerase III promoter, wherein each targeting peptide and each barcode are uniquely paired.
[0006] The barcode sequence can be at least 9, at least 12, at least 15, at least 18, or at least 21 nucleotides in length. The barcode sequence can be 9 to 21 nucleotides in length, 12 to 21 nucleotides in length, 15 to 21 nucleotides in length, 9 to 18 nucleotides in length, 9 to 15 nucleotides in length, or 9 to 12 nucleotides in length. The barcode sequence can be 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 nucleotides in length. The barcode sequence can be flanked by sequences capable of hybridizing to and activating a padlock probe. Each barcode sequence independently contains (NNNT) n sequence.
[0007] The RNA polymerase III promoter can be a type III RNA polymerase III promoter. The RNA polymerase promoter can be a U6 snRNA gene promoter, an H1 RNA gene promoter, or a 7SK gene promoter.
[0008] The rAAV vector can further contain a reverse transcription primer binding site located 3' of the barcode sequence and an enrichment primer binding site located 5' of the barcode sequence. The expression cassette can contain a sequence identical to SEQ ID NO:7, a sequence having at least 90% identity or at least 95% identity to SEQ ID NO:7.
[0009] The rAAV vector may also contain a modified adeno-associated virus (AAV) Cap gene, which encodes a modified AAV capsid protein containing a targeting peptide. The modified AAV capsid protein may be a modified AAV1 capsid protein, a modified AAV2 capsid protein, or a modified AAV9 capsid protein. The length of the targeting peptide may be three to ten amino acids. The length of the targeting peptide may be 3, 4, 5, 6, 7, 8, 9, or 10 amino acids.
[0010] If the modified AAV capsid protein is derived from the AAV1 capsid protein (see SEQ ID NO:1), the targeting peptide may be inserted after the 590th residue of the AAV1 capsid protein. The targeting peptide may be flanked by linker sequences, where the linker sequences on each side of the targeting peptide are two or three amino acids in length. The linker sequences may be SSA on the N-terminal side of the targeting peptide and AS on the C-terminal side of the targeting peptide. The modified AAV1 capsid protein may have the same sequence as SEQ ID NO:4, a sequence having at least 90% identity or at least 95% identity with SEQ ID NO:4.
[0011] If the modified AAV capsid protein is derived from the AAV2 capsid protein (see SEQ ID NO:2), the targeting peptide may be inserted after the 587th residue of the AAV2 capsid protein. The targeting peptide may be flanked by linker sequences, where the linker sequences on each side of the targeting peptide are two or three amino acids in length. The linker sequences may be AAA on the N-terminal side of the targeting peptide and AA on the C-terminal side of the targeting peptide. The modified AAV2 capsid protein may have the same sequence as SEQ ID NO:5, a sequence having at least 90% identity or at least 95% identity with SEQ ID NO:5.
[0012] If the modified AAV capsid protein is derived from the AAV9 capsid protein (see SEQ ID NO:3), the targeting peptide may be inserted after the 588th residue of the AAV9 capsid protein. The targeting peptide may be flanked by linker sequences, where the linker sequences on each side of the targeting peptide are two or three amino acids in length. The linker sequences may be AAA on the N-terminal side of the targeting peptide and AS on the C-terminal side of the targeting peptide. The modified AAV9 capsid protein may have the same sequence as SEQ ID NO:6, a sequence having at least 90% identity or at least 95% identity with SEQ ID NO:6.
[0013] An rAAV vector population can be provided, wherein the population contains a plurality of capsid protein targeting peptides, and each capsid protein targeting peptide is paired with more than one barcode sequence. An rAAV vector population can be provided, wherein the population contains a plurality of capsid protein targeting peptides, and all rAAVs having the same barcode sequence also have the same capsid protein targeting peptide. In other words, a plurality of RNAbc sequences represent a single AAV peptide insertion sequence. This feature enables the use of "randomer" sequences during plasmid production, making the methods provided herein more high-throughput because it does not require individual cloning of each RNAbc-AAV peptide insertion combination.
[0014] Cells containing the rAAV vectors of the embodiments of the present invention are also provided herein. The cells can be mammalian cells. The cells can be human cells. The cells can be in vitro or in vivo.
[0015] Library preparation techniques are also provided herein, which are designed to barcode and recover both mRNA and AAV-derived RNA simultaneously. In one embodiment, a method for determining the cell tropism of a recombinant adeno-associated virus (rAAV) having a modified AAV capsid protein containing a targeting peptide is provided, the method comprising: (i) contacting a plurality of cell types with the modified rAAV vector according to any one of the embodiments of the present invention; (ii) identifying the cells transduced by the modified rAAV vector in the presence of a barcode sequence; and (iii) detecting the transcriptome expressed in each transduced cell on a cell-by-cell basis to determine the cell tropism of the modified rAAV. In one embodiment, a method for determining the cell tropism of a recombinant adeno-associated virus (rAAV) having a modified AAV capsid protein containing a targeting peptide is provided, the method comprising: (i) contacting a plurality of cell types with the rAAV vector population provided herein; (ii) detecting both the transcriptome expressed and the rAAV transducing each cell on a cell-by-cell basis; and (iii) determining which cell types are transduced by which modified rAAV vectors to determine the cell tropism of the modified rAAV.
[0016] The contact in (i) can be carried out in vitro or in vivo.
[0017] Detecting the expressed transcriptome and rAAV in (ii) may include: (a) isolating, fixing, and permeabilizing the nuclei of the cells contacted in (i); (b) dividing the nuclei into a plurality of first aliquots; (c) reverse transcribing the cellular RNA molecules expressed in the nuclei using primers comprising a poly(T) sequence to form complementary DNA (cDNA) molecules, and reverse transcribing the rAAV RNA molecules expressed in the nuclei using primers comprising a sequence sufficient to hybridize to and reverse transcribe the barcode sequence within the expression cassette to form AAV amplicons; (d) labeling the cDNA molecules and AAV amplicons with a first 5' barcode, wherein the first 5' barcode of the primers in each first aliquot is unique such that the cDNA molecules and AAV amplicons from the nuclei of each aliquot can be identified compared to the cDNA molecules and AAV amplicons from the nuclei of all other aliquots; (e) pooling the plurality of first aliquots; (f) dividing the pooled plurality of first aliquots into a plurality of second aliquots; (g) ligating a second 5' barcode to the 5' ends of the cDNA molecules and AAV amplicons to form doubly barcoded cDNA molecules and AAV amplicons, wherein the second 5' barcode in each second aliquot is unique; (h) pooling the plurality of second aliquots; (i) dividing the pooled plurality of first aliquots into a plurality of third aliquots; (j) ligating a third 5' barcode to the 5' ends of the cDNA molecules and AAV amplicons to form triply barcoded cDNA molecules and AAV amplicons, wherein the third 5' barcode in each third aliquot is unique; (k) pooling the plurality of third aliquots; (l) lysing the nuclei to release the cDNA molecules and AAV amplicons from within the nuclei to form a lysate; and (m) sequencing the cDNA molecules and AAV amplicons to detect both the expressed transcriptome and the rAAV transducing each cell.
[0018] The cDNA molecules and AAV amplicons can be labeled with the first 5' barcode during reverse transcription, wherein the reverse transcription primers comprise the first 5' barcode. The nuclei can be fixed and permeabilized at a temperature below about 8°C, below about 7°C, below about 6°C, below about 5°C, below about 4°C, below about 3°C, below about 2°C, or below about 1°C. Most of the triply barcoded cDNA molecules and AAV molecules from a single nucleus may contain the same series of barcodes. Most of the triply barcoded cDNA molecules and AAV molecules from a single nucleus may have a unique series of barcodes compared to the triply barcoded cDNA molecules and AAV molecules from other nuclei. The cell type can be determined based on the expressed transcriptome.
[0019] Sequencing cDNA molecules and AAV amplicons includes preparing a sequencing library, which may include: (i) adding a common adapter sequence to the 3' ends of the cDNA molecules and AAV amplicons; (ii) performing full-length cDNA and AAV amplicon amplification; (iii) fragmenting the amplified full-length cDNA and AAV amplicons; (iv) performing end repair and A-tailing on the fragmented cDNA and AAV amplicons; (v) ligating an adapter to the 5' ends of the end-repaired and A-tailed cDNA and AAV amplicons; and (vi) performing sample indexing PCR to add sequencing adapters and dual indices to the adapter-ligated cDNA and AAV amplicons.
[0020] Sequencing AAV amplicons includes preparing a sequencing library enriched for AAV amplicons, which may include: (i) adding a common adapter sequence to the 3' ends of the cDNA molecules and AAV amplicons; (ii) performing full-length amplification of the cDNA and AAV amplicons; (iii) performing enrichment amplification of the AAV amplicons with a forward primer that hybridizes upstream of the AAV barcode in the AAV amplicon, wherein the forward primer has a 5' phosphate; (iv) performing A-tailing on the AAV amplicon with a 5' phosphate; (v) ligating an adapter to the 5' ends of the A-tailed AAV amplicons; and (vi) performing sample indexing PCR to add sequencing adapters and dual indices to the adapter-ligated AAV amplicons.
[0021] The common adapter sequence can be added to the 3' ends of the cDNA molecules and AAV amplicons by template switching.
[0022] Sequencing can be paired-end sequencing, amplicon sequencing, single-cell RNA sequencing, or in situ sequencing.
[0023] Other objects, features, and advantages of the present invention will become apparent from the following detailed description. However, it should be understood that although the detailed description and specific examples indicate some preferred embodiments of the present invention, they are given by way of illustration only, since various changes and modifications within the spirit and scope of the present invention will become apparent to those skilled in the art from this detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The following drawings form a part of this specification and are included to further illustrate certain aspects of the present invention. The present invention can be better understood by referring to one or more of these drawings in combination with the detailed description of the specific embodiments given herein.
[0025] Figures 1A to 1F. Design of AAV cargo containing dual barcodes. (A) Schematic diagram showing each element in the barcoded expression construct (SEQ ID NO:7) component of the AAV cargo. (B) Cartoon schematic showing the layout of the construct shown in Figure 1A within the packaged AAV genome and relative to the AAV Cap gene sequence that has been modified to contain a peptide insertion. (C) Gel image showing RNAbc sequence amplification after reverse transcription using primers pr749 and 750. (D) Schematic highlighting the presence of two different DNA barcodes, namely RNAbc and the peptide inserted into the Cap sequence. (E and F) Sanger sequencing across each insert confirmed the successful creation of this dual-barcode construct. In Figure 1E , the five sequences from top to bottom are SEQ ID NOs: 15 to 19 respectively. In Figure 1F , the top sequence is SEQ ID NO: 20, and the four sequences at the bottom are all SEQ ID NO: 21.
[0026] Figures 2A to 2B . Adaptation of Split-Pool Ligation-based whole-Transcriptome Sequencing (SPLiT-seq) for AAV.RNAbc detection. (A) Schematic diagram showing the procedural steps of single-nucleus combinatorial barcoding. (B) Schematic diagram showing the procedural steps of next-generation library preparation for whole-transcriptome and AAV.RNAbc amplicon sequencing.
[0027] Figures 3A to 3D. Single-cell RNA-Seq results showing detection of the RNA barcode (RNAbc) expressed after transfection of HEK 293 cells. (A) UMAP unbiased clustering of SPLiT-Seq barcoded single cells from an experiment in which HEK293 cells were transfected with plasmids containing AAV.RNAbc or AAV.noBarcode. (B) Unique UMI counts obtained from Illumina sequencing reads after amplification and library preparation. (C) UMAP unbiased clustering of SPLiT-Seq barcoded single cells from cells treated only with AAV.RNAbc. Heatmap shading represents the log UMI counts derived from AAV.RNAbc amplicons. (D) UMAP unbiased clustering of SPLiT-Seq barcoded single cells from cells treated only with AAV.eGFP. Heatmap shading represents the log UMI counts derived from AAV.eGFP amplicons.
[0028] Figure 4ATo 4D. In vivo applications of SPLiT-Seq. (A) Schematic showing the procedural steps. (B) Seurat cell annotation containing cDNA expression and AAV transduction information. (C) Identification of transduced single cells. (D) Assessment of AAV transduction status within the target single cell type.
[0029] Figures 5A to 5C. Tools for analyzing transduction performance. (A) Transduction performance by tissue. (B) Transduction performance by cell type. (C) Transduction performance visualized spatially using UMAP unbiased clustering. Detailed Description
[0030] Barcoded RNA expression constructs are provided herein that, when packaged into AAVs and delivered to cells, express mRNA sequences with identifying barcodes that are distinct from modified regions of the capsid DNA sequences. Library preparation techniques are also provided herein that are designed to barcode and recover both mRNA and AAV-derived RNA simultaneously. Finally, a custom software pipeline is provided that integrates the publicly available SPLiT-Seq demultiplexing pipeline with a pipeline for counting AAV amplicon sequences. I. Adeno-Associated Virus (AAV) Vectors
[0031] Adeno-associated virus (AAV) is a small non-pathogenic virus of the family Parvoviridae. To date, many serologically distinct AAVs have been identified and more than a dozen have been isolated from humans or primates. AAV differs from other members of the family in that it is replication-dependent on a helper virus.
[0032] The AAV genome can exist in an episomal state without integrating into the host cell genome; has a broad host range; transduces both dividing and non-dividing cells in vitro and in vivo and maintains high-level expression of the transduced gene. AAV viral particles are heat-stable; resistant to solvents, detergents, pH changes, and temperature; and can be column-purified and / or concentrated on a CsCl gradient or by other means. The AAV genome contains single-stranded deoxyribonucleic acid (ssDNA) that is either sense or antisense. The approximately 4.7 kb genome of AAV consists of a segment of single-stranded DNA that is either positive or negative strand. The ends of the genome are short inverted terminal repeats (ITRs) that can fold into hairpin structures and serve as origins of viral DNA replication.
[0033] An AAV "genome" refers to a recombinant nucleic acid sequence that is ultimately packaged or encapsulated to form an AAV particle. An AAV particle typically contains an AAV genome packaged with AAV capsid proteins. In the case of constructing or preparing a recombinant vector using a recombinant plasmid, the AAV vector genome does not contain a "plasmid" portion that does not correspond to the vector genome sequence of the recombinant plasmid. This non-vector genome portion of the recombinant plasmid is referred to as the "plasmid backbone", which is important for the cloning and amplification of the plasmid (processes required for plasmid propagation and production), but is not itself packaged or encapsulated into virus particles. Thus, an AAV vector "genome" refers to the nucleic acid packaged or encapsulated by AAV capsid proteins.
[0034] An AAV virion (particle) is a non-enveloped icosahedral particle approximately 25 nm in diameter that contains an AAV capsid. The AAV particle exhibits icosahedral symmetry and contains three related capsid proteins, VP1, VP2, and VP3, which interact to form the capsid. The genomes of most natural AAVs typically contain two open reading frames (ORFs), sometimes referred to as the left ORF and the right ORF. The right ORF typically encodes the capsid proteins VP1, VP2, and VP3. These proteins typically exist in a ratio of 1:1:10, respectively, but can exist in different ratios and all originate from the right-hand ORF. The VP1, VP2, and VP3 capsid proteins differ from one another by using alternative splicing and unusual start codons. Deletion analysis has shown that removal or alteration of VP1 translated from alternative splicing information results in a reduced yield of infectious particles. Mutations within the VP3 coding region result in the inability to produce any single-stranded progeny DNA or infectious particles. In certain embodiments, the genome of an AAV particle encodes one, two, or all three of the VP1, VP2, and VP3 polypeptides.
[0035] The left ORF typically encodes the non-structural Rep proteins, Rep 40, Rep 52, Rep 68, and Rep 78, which are involved in the regulation of replication and transcription in addition to the production of single-stranded progeny genomes. Two of the Rep proteins are associated with the preferential integration of the AAV genome into the region of the q arm of human chromosome 19. Rep68 / 78 has been shown to have NTP-binding activity as well as DNA and RNA helicase activity. Some of the Rep proteins have a nuclear localization signal as well as several potential phosphorylation sites. In certain embodiments, the genome of an AAV (e.g., rAAV) encodes some or all of the Rep proteins. In certain embodiments, the genome of an AAV (e.g., rAAV) does not encode the Rep proteins. In certain embodiments, one or more of the Rep proteins can be delivered in trans and thus are not included in the AAV particle containing the nucleic acid encoding the polypeptide.
[0036] The ends of the AAV genome contain short inverted terminal repeats (ITRs) that have the potential to fold into a T-shaped hairpin structure, which serves as the origin of viral DNA replication. Thus, the AAV genome contains one or more ITR sequences (e.g., ITR sequence pairs) that flank the single-stranded viral DNA genome. The length of the ITR sequences is typically about 145 bases each. Within the ITR region, two elements that are considered to be the functional core of the ITR have been described, the GAGC repeat motif and the terminal resolution site (trs). When the ITR is in a linear or hairpin conformation, the repeat motif shows binding to Rep. This binding is thought to localize Rep68 / 78 to the trs for cleavage, which occurs in a site-specific and strand-specific manner. In addition to their role in replication, these two elements also appear to be central to viral integration. Contained within the chromosome 19 integration locus is a Rep binding site with an adjacent trs. These elements have been shown to be functional and are required for locus-specific integration.
[0037] The term "recombinant" as a modifier of a vector (e.g., recombinant virus, such as a lentivirus or parvovirus (e.g., AAV) vector) and of a sequence (e.g., recombinant nucleic acid sequence and polypeptide) means that the composition has been manipulated (i.e., engineered) in a manner that does not normally occur in nature. A specific example of a recombinant vector (e.g., AAV, retroviral, or lentiviral vector) is the insertion of a nucleic acid sequence that is not normally present in the wild-type viral genome into the viral genome. An example of a recombinant nucleic acid sequence is a nucleic acid (e.g., gene) encoding an inhibitory RNA cloned into a vector that has or does not have the 5', 3', and / or intron regions normally associated with that gene in the viral genome. Although the term "recombinant" is not always used herein to refer to vectors (e.g., viral vectors) and sequences (e.g., polynucleotides), "recombinant" forms including nucleic acid sequences, polynucleotides, transgenes, etc. are specifically included, notwithstanding any such omission.
[0038] The recombinant viral "vector" is derived from the wild-type genome of a virus by using molecular methods to remove a portion of the wild-type genome from the virus and replace it with a non-native nucleic acid (e.g., a nucleic acid sequence). Generally speaking, for example, in the case of AAV, one or two inverted terminal repeat (ITR) sequences of the AAV genome are retained in the recombinant AAV vector. The "recombinant" viral vector (e.g., rAAV) is different from the viral (e.g., AAV) genome because a portion of the viral genome has been replaced with a non-native sequence directed against viral genomic nucleic acid (e.g., nucleic acid encoding a transactivator or nucleic acid encoding an inhibitory RNA or nucleic acid encoding a therapeutic protein). Thus, the incorporation of such a non-native nucleic acid sequence defines the viral vector as a "recombinant" vector, which in the case of AAV can be referred to as an "rAAV vector".
[0039] In certain embodiments, AAV (e.g., rAAV) contains two ITRs. In certain embodiments, AAV (e.g., rAAV) contains a pair of ITRs. In certain embodiments, AAV (e.g., rAAV) contains a pair of ITRs that flank at least a nucleic acid sequence encoding a polypeptide having function or activity (i.e., at each 5' and 3' end).
[0040] AAV vectors (e.g., rAAV vectors) can be packaged and are referred to herein as "AAV particles", which are used for subsequent ex vivo, in vitro or in vivo infection (transduction) of cells. When the recombinant AAV vector is encapsulated or packaged into an AAV particle, the particle can also be referred to as an "rAAV particle". In certain embodiments, the AAV particle is an rAAV particle. rAAV particles generally contain an rAAV vector or a portion thereof. rAAV particles can be one or more rAAV particles (e.g., multiple AAV particles). rAAV particles generally contain proteins (e.g., capsid proteins) that encapsulate or package the rAAV vector genome. It should be noted that a reference to an rAAV vector can also be used to refer to an rAAV particle.
[0041] Any suitable AAV particle (e.g., rAAV particle) can be used in the methods or uses herein. The rAAV particle and / or the genome contained therein can be derived from any suitable AAV serotype or strain. The rAAV particle and / or the genome contained therein can be derived from two or more AAV serotypes or strains. Thus, rAAV can contain proteins and / or nucleic acids or portions thereof of any AAV serotype or strain, wherein the AAV particle is suitable for infection and / or transduction of mammalian cells. Some non-limiting examples of AAV serotypes include AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV-rh74, AAV-rh10, and AAV-2i8.
[0042] In certain embodiments, the plurality of rAAV particles comprise particles of the same strain or serotype (or subgroup or variant), or comprise particles derived from the same strain or serotype (or subgroup or variant). In certain embodiments, the plurality of rAAV particles comprise a mixture of two or more different rAAV particles (e.g., different serotypes and / or strains).
[0043] As used herein, the term "serotype" refers to a distinction used to denote an AAV that has a capsid that is serologically distinct from other AAV serotypes. Serological distinctiveness is determined based on the lack of cross-reactivity between the antibodies of one AAV and the antibodies of another AAV. Such differences in cross-reactivity are generally due to differences in the capsid protein sequences / epitopes (e.g., due to differences in the VP1, VP2, and / or VP3 sequences of the AAV serotype). Although there is a possibility that AAV variants (including capsid variants) may not be serologically distinguishable from a reference AAV or other AAV serotypes, they differ by at least one nucleotide or amino acid residue compared to the reference or other AAV serotypes.
[0044] In certain embodiments, the rAAV vector genome based on a first serotype corresponds to the serotype of one or more capsid proteins that package the vector. For example, the serotype of one or more AAV nucleic acids (e.g., ITRs) that contain the AAV vector genome corresponds to the serotype of the capsid that contains the rAAV particle.
[0045] In certain embodiments, the rAAV vector genome can be based on an AAV (e.g., AAV2) serotype genome that is different from the serotype of one or more AAV capsid proteins that package the vector. For example, the rAAV vector genome can contain nucleic acids (e.g., ITRs) from AAV2, while at least one or more of the three capsid proteins are derived from a different serotype, such as the AAV1, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, Rh10, Rh74, or AAV-2i8 serotype or variants thereof.
[0046] In certain embodiments, the rAAV particles or their vector genomes related to a reference serotype have a polynucleotide, polypeptide, or a subsequence thereof that comprises or consists of a sequence comprising: having at least 60% or more (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc.) identity to the polynucleotide, polypeptide, or subsequence of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, Rh10, Rh74, or AAV-2i8 particles. In some specific embodiments, the rAAV particles or their vector genomes related to a reference serotype have a capsid or ITR sequence that comprises or consists of a sequence comprising: having at least 60% or more (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc.) identity to the capsid or ITR sequence of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, Rh10, Rh74, or AAV-2i8 serotype.
[0047] In certain embodiments, the methods herein include using, administering, or delivering rAAV1, rAAV2, rAAV3, rAAV4, rAAV5, rAAV6, rAAV7, rAAV8, rAAV9, rAAV10, rAAV11, rAAV12, rRh10, rRh74, or rAAV-2i8 particles.
[0048] In certain embodiments, the methods herein include using, administering, or delivering rAAV2 particles. In certain embodiments, the rAAV2 particles comprise an AAV2 capsid. In certain embodiments, the rAAV2 particles comprise one or more capsid proteins (e.g., VP1, VP2, and / or VP3) that have at least 60%, 65%, 70%, 75%, or higher identity to the corresponding capsid proteins of a native or wild-type AAV2 particle, such as 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identity. In certain embodiments, the rAAV2 particles comprise VP1, VP2, and VP3 capsid proteins that have at least 75% or higher identity to the corresponding capsid proteins of a native or wild-type AAV2 particle, such as 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identity. In certain embodiments, the rAAV2 particles are variants of a native or wild-type AAV2 particle. In some aspects, compared to the capsid proteins of a native or wild-type AAV2 particle, one or more capsid proteins of the AAV2 variant have 1, 2, 3, 4, 5, 5 to 10, 10 to 15, 15 to 20, or more amino acid substitutions.
[0049] In certain embodiments, the rAAV9 particles comprise an AAV9 capsid. In certain embodiments, the rAAV9 particles comprise one or more capsid proteins (e.g., VP1, VP2, and / or VP3), which have at least 60%, 65%, 70%, 75% or higher identity with the corresponding capsid proteins of native or wild-type AAV9 particles, such as 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identity. In certain embodiments, the rAAV9 particles comprise VP1, VP2, and VP3 capsid proteins, which have at least 75% or higher identity with the corresponding capsid proteins of native or wild-type AAV9 particles, such as 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identity. In certain embodiments, the rAAV9 particles are variants of native or wild-type AAV9 particles. In some aspects, compared to the capsid proteins of native or wild-type AAV9 particles, one or more capsid proteins of the AAV9 variant have 1, 2, 3, 4, 5, 5 to 10, 10 to 15, 15 to 20 or more amino acid substitutions.
[0050] In some embodiments, the rAAV comprises a modified capsid, wherein the modified capsid comprises a targeting peptide. In certain embodiments, the AAV is AAV1, AAV2, or AAV9. Exemplary wild-type reference AAV1 capsid protein sequences are provided in SEQ ID NO:1. Exemplary wild-type reference AAV2 capsid protein sequences are provided in SEQ ID NO:2. Exemplary wild-type reference AAV9 capsid protein sequences are provided in SEQ ID NO:3. In certain aspects, the targeting peptide is inserted at position 590 of the AAV1 capsid, position 587 of the AAV2 capsid, or position 588 of the AAV9 capsid. Exemplary modified AAV1 capsid protein sequences are provided in SEQ ID NO:4, which shows the insertion of a targeting peptide SSAX7AS after position 590, where the front SSA and the tail AS are linker sequences and X7 represents the targeting peptide. Exemplary modified AAV2 capsid protein sequences are provided in SEQ ID NO:5, which shows the insertion of a targeting peptide AAAX7AA after position 587, where the front AAA and the tail AA are linker sequences and X7 represents the targeting peptide. Exemplary modified AAV9 capsid protein sequences are provided in SEQ ID NO:6, which shows the insertion of a targeting peptide AAAX7AS after position 588, where the front AAA and the tail AS are linker sequences and X7 represents the targeting peptide. Table 1. AAV Capsid Sequences
[0051] In certain embodiments, the rAAV particle comprises one or two ITRs (e.g., an ITR pair), and the ITR (e.g., the ITR pair) has at least 75% or higher identity with the corresponding ITRs of native or wild-type AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV-rh74, AAV-rh10, or AAV-2i8, such as 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identity, as long as it retains one or more desired ITR functions (e.g., the ability to form a hairpin, which allows DNA replication; integrating AAV DNA into the host cell genome; and / or packaging, if desired).
[0052] In certain embodiments, the rAAV2 particle comprises one or two ITRs (e.g., an ITR pair), and the ITR (e.g., the ITR pair) has at least 75% or higher identity with the corresponding ITR of a native or wild-type AAV2 particle, such as 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identity, provided that it retains one or more desired ITR functions (e.g., the ability to form a hairpin, which allows DNA replication; integration of AAV DNA into the host cell genome; and / or packaging, if desired).
[0053] In certain embodiments, the rAAV9 particle comprises one or two ITRs (e.g., an ITR pair), and the ITR (e.g., the ITR pair) has at least 75% or higher identity with the corresponding ITR of a native or wild-type AAV2 particle, such as 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identity, provided that it retains one or more desired ITR functions (e.g., the ability to form a hairpin, which allows DNA replication; integration of AAV DNA into the host cell genome; and / or packaging, if desired).
[0054] The rAAV particle may comprise an ITR having any suitable number of "GAGC" repeats. In certain embodiments, the ITR of the AAV2 particle comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more "GAGC" repeats. In certain embodiments, the rAAV2 particle comprises an ITR containing three "GAGC" repeats. In certain embodiments, the rAAV2 particle comprises an ITR having fewer than four "GAGC" repeats. In certain embodiments, the rAAV2 particle comprises an ITR having more than four "GAGC" repeats. In certain embodiments, the ITR of the rAAV2 particle comprises a Rep binding site, wherein the fourth nucleotide in the first two "GAGC" repeats is C instead of T.
[0055] Exemplary suitable lengths of DNA that can be incorporated into the rAAV vector for packaging / encapsulation into rAAV particles can be about 5 kilobases (kb) or less. In some specific embodiments, the length of the DNA is less than about 5 kb, less than about 4.5 kb, less than about 4 kb, less than about 3.5 kb, less than about 3 kb, or less than about 2.5 kb.
[0056] rAAV vectors containing nucleic acid sequences that direct the expression of RNAi or polypeptides can be generated using suitable recombinant techniques known in the art (e.g., see Sambrook et al., 1989). Recombinant AAV vectors are typically packaged into transducing AAV particles and propagated using an AAV viral packaging system. Transducing AAV particles are capable of binding to and entering mammalian cells and subsequently delivering a nucleic acid cargo (e.g., a heterologous gene) to the nucleus of the cell. Thus, intact rAAV particles with transducing ability are configured to transduce mammalian cells. rAAV particles configured to transduce mammalian cells generally do not have the ability to replicate and require additional protein machinery for self-replication. Thus, rAAV particles configured to transduce mammalian cells are engineered to bind to and enter mammalian cells and deliver nucleic acids to the cell, where the nucleic acid for delivery is typically located between AAV ITR pairs in the rAAV genome.
[0057] Suitable host cells for generating transducing AAV particles include, but are not limited to, microorganisms, yeast cells, insect cells, and mammalian cells that can or have been used as recipients for heterologous rAAV vectors. Cells from the stable human cell line HEK293 (readily available, for example, from the American Type Culture Collection under accession number ATCC CRL1573) can be used. In certain embodiments, a modified human embryonic kidney cell line (e.g., HEK293) that has been transformed with a type 5 adenovirus DNA fragment and expresses the adenovirus E1a and E1b genes is used to generate recombinant AAV particles. The modified HEK293 cell line is readily transfected and provides a particularly convenient platform for generating rAAV particles therein. Methods for generating high-titer AAV particles capable of transducing mammalian cells are known in the art. For example, AAV particles can be prepared as shown in Wright, 2008 and Wright, 2009.
[0058] In certain embodiments, AAV helper functions are introduced into a host cell by transfecting the host cell with an AAV helper construct, either before or simultaneously with transfection of an AAV expression vector. Thus, AAV helper constructs are sometimes used to provide at least transient expression of the AAV rep and / or cap genes to complement missing AAV functions necessary for productive AAV transduction. AAV helper constructs typically lack AAV ITRs and are neither capable of replicating nor packaging themselves. These constructs can be in the form of plasmids, phages, transposons, cosmids, viruses, or virions. Many AAV helper constructs have been described, such as the commonly used plasmids pAAV / Ad and pIM29+45 that encode both Rep and Cap expression products. Many other vectors encoding Rep and / or Cap expression products are known. II. Methods for Determining Modified AAV Cell Tropism
[0059] Methods for uniquely labeling or barcoding molecules within one or more nuclei are provided herein. It is understood that the embodiments generally described herein are exemplary. The following more detailed description of various embodiments is not intended to limit the scope of the disclosure, but merely represents various embodiments. Additionally, those skilled in the art may change the order of steps or actions of the methods disclosed herein without departing from the scope of the disclosure. In other words, unless the correct operation of an embodiment requires a specific order of steps or actions, the order or use of specific steps or actions may be modified.
[0060] The term "binding" as used extensively throughout this disclosure refers to any form of connecting or coupling two or more components, entities, or objects. For example, two or more components may bind to each other through chemical bonds, covalent bonds, ionic bonds, hydrogen bonds, electrostatic forces, Watson-Crick hybridization, etc.
[0061] One aspect of the present disclosure relates to methods of tagging nucleic acids. In some embodiments, the method may include tagging nucleic acids in a first nucleus. The method may include: (a) generating complementary DNA (cDNA) from cellular RNA and / or AAV.RNAbc amplicons within a plurality of nuclei by reverse transcribing RNA using a reverse transcription primer comprising a 5' overhang sequence; (b) dividing the plurality of nuclei into a number (n) of aliquots; (c) providing a plurality of barcode tags to each of the n aliquots, wherein each tagging sequence of the plurality of barcode tags provided to a given aliquot is the same, and wherein different tagging sequences are provided to each of the n aliquots; (d) binding at least one cDNA and / or AAV.RNAbc amplicon in each of the n aliquots to the barcode tags; (e) pooling the n aliquots; and (f) repeating steps (b), (c), (d), and (e) with the pooled aliquots. In some aspects, two different reverse transcription primers are used, wherein the first reverse transcription primer comprises a poly(A) hybridization sequence (i.e., a poly(T) sequence), and wherein the second reverse transcription primer comprises a sequence capable of hybridizing downstream of the barcode sequence to RNA expressed from an AAV barcoded expression construct (i.e., an AAV.RNAbc transcript).
[0062] In certain embodiments, each barcode tag may comprise a first strand comprising a 3' hybridization sequence extending from the 3' end of the tagging sequence and a 5' hybridization sequence extending from the 5' end of the tagging sequence. Each barcode tag may further comprise a second strand comprising an overhang sequence. The overhang sequence may comprise: (i) a first portion complementary to at least one of the 5' hybridization sequence and the 5' overhang sequence, and (ii) a second portion complementary to the 3' hybridization sequence. In some embodiments, the barcode tag (e.g., the final nucleic acid tag) may comprise a capture agent, such as but not limited to 5'-biotin. A cDNA or AAV.RNAbc amplicon labeled with a barcode tag comprising 5'-biotin may allow or permit the cDNA or AAV.RNAbc amplicon to be linked or coupled to streptavidin-coated magnetic beads. In other embodiments, a plurality of beads may be coated with a capture strand (i.e., a nucleic acid sequence) configured to hybridize to the final sequence overhang of the barcode tag. In other embodiments, cDNA or AAV.RNAbc amplicon molecules may be purified or isolated by using a commercially available kit (e.g., the RNEASY TM kit).
[0063] In multiple embodiments, step (f) (i.e., steps (b), (c), (d), and (e)) can be repeated a number of times sufficient to generate a unique series of marker sequences for the cDNA and AAV.RNAbc amplicons in the first nucleus. In other words, step (f) can be repeated multiple times such that the cDNA and AAV.RNAbc amplicons in the first nucleus can have a first unique series of marker sequences, the cDNA and AAV.RNAbc amplicons in the second nucleus can have a second unique series of marker sequences, the cDNA and AAV.RNAbc amplicons in the third nucleus can have a third unique series of marker sequences, and so on. The methods of the present disclosure can provide for barcoding cDNA and AAV.RNAbc amplicon sequences from individual nuclei, where the unique barcode can identify or assist in identifying the cell from which the cDNA and AAV.RNAbc amplicons are derived. In other words, a portion, majority, or substantially all of the cDNA and AAV.RNAbc amplicons from a single cell can have the same barcode, and the barcode can not be repeated in the cDNA or AAV.RNAbc amplicons derived from one or more other cells in the sample (e.g., from a second cell, a third cell, a fourth cell, etc.).
[0064] In some embodiments, the barcoded cDNA and AAV.RNAbc amplicons can be mixed together and sequenced (e.g., using NGS) such that data on RNA expression and AAV transduction can be collected at the single cell level. For example, certain embodiments of the methods of the present disclosure can be used to evaluate, analyze, or study the cell tropism of a modified AAV (i.e., the specific cell type or types selectively or specifically targeted by any given modified AAV capsid).
[0065] As described above, aliquots or pools of nuclei can be separated into different reaction vessels or containers, and a first set of barcode tags can be added to multiple cDNA transcripts and AAV.RNAbc transcripts. Vessels or containers may also be referred to herein as receptacles, samples, and wells. Thus, the terms vessel, container, receptacle, sample, and well may be used interchangeably herein. The aliquots of nuclei can then be regrouped, mixed, and separated again, and a second set of barcode tags can be added to the first set of barcode tags. In multiple embodiments, the same barcode tag can be added to more than one aliquot of nuclei in a single or given round of tagging. However, after repeated rounds of separation, tagging, and repooling, the cDNA and AAV.RNAbc amplicons of each nucleus can be associated with a unique combination or sequence of barcode tags that identify the individual nucleus. In some embodiments, the nuclei in a single sample can be divided into multiple different reaction containers. For example, the number of reaction containers can include four 1.5 ml microcentrifuge tubes, multiple wells of a 96-well plate, multiple wells of a 384-well plate, or other suitable numbers and types of reaction containers.
[0066] In certain embodiments, step (f) (i.e., steps (b), (c), (d), and (e)) can be repeated multiple times, where the number of times is selected from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, etc. In certain other embodiments, step (f) can be repeated a sufficient number of times such that the cDNA and AAV.RNAbc amplicons of each nucleus are likely to be associated with a unique sequence of barcode tags. The number of times can be selected to provide a likelihood greater than 50%, greater than 90%, greater than 95%, greater than 99%, or some other probability that the cDNA and AAV.RNAbc amplicons in each nucleus are associated with a unique sequence of barcode tags. In other embodiments, step (f) can be repeated some other suitable number of times.
[0067] In some embodiments, a method of labeling nucleic acids in a first nucleus may include fixing a plurality of nuclei prior to step (a). For example, the components of the nuclei may be fixed or crosslinked such that the components are immobilized or maintained in place. Formaldehyde in phosphate buffered saline (PBS) may be used to fix the plurality of nuclei. The plurality of nuclei may be fixed in, for example, about 1% to 4% formaldehyde in PBS. In a plurality of embodiments, methanol (e.g., 100% methanol) may be used to fix the plurality of nuclei at about -20 °C or about 25 °C. In a plurality of other embodiments, methanol (e.g., 100% methanol) may be used to fix the plurality of nuclei at about -20 °C to about 25 °C. In a plurality of other embodiments, ethanol (e.g., about 70% to 100% ethanol) may be used to fix the plurality of nuclei at about -20 °C or at room temperature. In a plurality of other embodiments, ethanol (e.g., about 70% to 100% ethanol) may be used to fix the plurality of nuclei at about -20 °C to room temperature. In a plurality of other embodiments, acetic acid may be used to fix the plurality of nuclei, for example, at about -20 °C. In a plurality of other embodiments, acetone may be used to fix the plurality of nuclei, for example, at about -20 °C. Other suitable methods of fixing the plurality of nuclei are also within the scope of the present disclosure.
[0068] In certain embodiments, a method of labeling nucleic acids in a first nucleus may include permeabilizing a plurality of nuclei prior to step (a). For example, pores or openings may be formed in the nuclear membranes of the plurality of nuclei. TRITON TM X-100 may be added to the plurality of nuclei, followed optionally by the addition of HCl to form one or more pores. For example, about 0.2% TRITON TM X-100 may be added to the plurality of nuclei, followed by the addition of about 0.1 N HCl. In certain other embodiments, ethanol (e.g., about 70% ethanol), methanol (e.g., about 100% methanol), Tween 20 (e.g., about 0.2% Tween 20), and / or NP-40 (e.g., about 0.1% NP-40) may be used to permeabilize the plurality of nuclei. In a plurality of embodiments, a method of labeling nucleic acids in a first nucleus may include fixing and permeabilizing a plurality of nuclei prior to step (a).
[0069] In some embodiments, a method of tagging nucleic acids in a first nucleus may include ligating at least two barcode tags that bind to cDNA and / or AAV.RNAbc amplicons. The ligation may be performed before or after a lysis and / or nucleic acid purification step. The ligation may include covalently linking a 5' phosphate sequence on the barcode tag to an adjacent strand or the 3' end of an adjacent barcode tag such that the individual tags form a continuous or substantially continuous barcode sequence that binds to the 3' end of the cDNA sequence. In various embodiments, a double-stranded DNA or RNA ligase may be used with an additional adapter strand that is configured to hold the barcode tag and adjacent nucleic acid together in a "nick" duplex conformation. The double-stranded DNA or RNA ligase may then be used to seal the "nick". In various other embodiments, a single-stranded DNA or RNA ligase may be used without an additional adapter. In certain embodiments, the ligation may be performed within multiple nuclei.
[0070] In certain other embodiments, the method may include lysing multiple nuclei (i.e., breaking down the nuclear structure) to release cDNA and / or AAV.RNAbc amplicons from within the multiple nuclei, e.g., after step (f). In some embodiments, the multiple nuclei may be lysed in a lysis solution (e.g., 10 mM Tris-HCl (pH 7.9), 50 mM EDTA (pH 7.9), 0.2 M NaCl, 2.2% SDS, 0.5 mg / ml ANTI-RNase (protein ribonuclease inhibitor; ) and 1000 mg / ml proteinase K ), e.g., by shaking (e.g., vigorously) at about 55 °C for about 1 to 3 hours. In other embodiments, the multiple nuclei may be lysed using sonication and / or by passing through an 18 to 25 gauge syringe needle at least once. In other embodiments, the multiple nuclei may be lysed by heating to about 70 °C to 90 °C. For example, the multiple nuclei may be lysed by heating to about 70 °C to 90 °C for about one or more hours. The cDNA and / or AAV.RNAbc amplicons may then be isolated from the lysed nuclei. In some embodiments, RNase H may be added to the cDNA and AAV.RNAbc amplicons to remove RNA. The method may also include ligating at least two of the barcode tags that bind to the released cDNA and AAV.RNAbc amplicons. In other embodiments, a method of tagging nucleic acids in a first cell may include ligating at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, etc. of the barcode tags that bind to the cDNA and AAV.RNAbc amplicons.
[0071] In multiple embodiments, methods of tagging nucleic acids in a first nucleus can include removing one or more unbound barcode labels (e.g., washing the multiple nuclei). For example, the method can include removing a portion, a majority, or substantially all of the unbound barcode labels. Unbound barcode labels can be removed so that further rounds of the disclosed methods are not contaminated by one or more unbound barcode labels from a previous round of a given method. In some embodiments, unbound barcode labels can be removed by centrifugation. For example, the multiple nuclei can be centrifuged so that a nuclear pellet forms at the bottom of a centrifuge tube. The supernatant (i.e., the liquid containing the unbound barcode labels) can be removed from the centrifuged nuclei. The nuclei can then be resuspended in a buffer (e.g., fresh buffer that does not contain or substantially does not contain unbound barcode labels). In another example, the multiple nuclei can be coupled or linked to magnetic beads that are coated with an antibody configured to bind to the nuclear membrane. The multiple nuclei can then be pelleted using a magnet to attract them to one side of a reaction vessel.
[0072] As described above, the multiple nuclei can be re - combined, and the method can be repeated any number of times, adding more barcode labels to the cDNA and AAV.RNAbc amplicons to generate a unique set of barcode labels that can serve to identify cDNA and AAV.RNAbc amplicons derived from the same cell. As more and more rounds are added, the number of paths the nuclei can take increases, and thus the number of possible unique barcode label sequences that can be created also increases. Given sufficient rounds and division, the number of possible barcodes will be far higher than the number of nuclei, resulting in each nucleus potentially having a unique barcode. For example, if the distribution is carried out in a 96 - well plate, after 4 distributions, there will be 96 4 = 84,934,656 possible barcodes.
[0073] In some embodiments, the cDNA reverse transcription primer can be configured to reverse - transcribe all or substantially all of the RNA in a cell (e.g., a random hexamer with a 5' overhang). In other embodiments, the cDNA reverse transcription primer can be configured to reverse - transcribe RNA having a poly(A) tail (e.g., a poly(dT) primer with a 5' overhang, such as a dT(15) primer). In other embodiments, the cDNA reverse transcription primer can be configured to reverse - transcribe a predetermined RNA (e.g., a transcript - specific primer). For example, the cDNA reverse transcription primer can be configured to barcode specific transcripts so that fewer transcripts in each cell can be profiled, but so that each transcript can be profiled in a greater number of cells.
[0074] In some embodiments, the AAV.RNAbc reverse transcription primer can be configured to reverse transcribe RNA expressed by the AAV barcoded expression construct (i.e., the AAV.RNAbc transcript). For example, the AAV.RNAbc reverse transcription primer can be configured to hybridize with the AAV.RNAbs transcript downstream of the barcode sequence.
[0075] Reverse transcription can be performed or carried out on multiple nuclei. In certain embodiments, reverse transcription can be performed on multiple fixed and / or permeabilized nuclei. In some embodiments, variants of M-MuLV reverse transcriptase can be used for reverse transcription. Any suitable reverse transcription method is within the scope of the present disclosure. For example, the reverse transcription mixture can contain a reverse transcription primer with a 5' overhang, and the reverse transcription primer can be configured to initiate reverse transcription and / or serve as a binding sequence for the barcode label. In other embodiments, a portion of the reverse transcription primer configured to bind to RNA and / or initiate reverse transcription can include one or more of the following: random hexamers, heptamers, octamers, nonamers, decamers, poly(T) extensions of nucleotides, and / or one or more gene-specific primers.
[0076] Another aspect of the present disclosure relates to a method for uniquely tagging multiple intranuclear RNA molecules. The method may include: (a) fixing and permeabilizing a first plurality of nuclei prior to step (b), wherein the first plurality of nuclei may be fixed and permeabilized at a temperature below about 8 °C; (b) reverse transcribing RNA molecules in the first plurality of cells to form complementary DNA (cDNA) molecules and AAV.RNAbc amplicons within the first plurality of nuclei, wherein reverse transcribing the RNA molecules includes coupling a primer to the RNA molecules, and the primer includes at least one of a poly(T) sequence or a sequence capable of hybridizing downstream of a barcode sequence to RNA expressed by an AAV barcoded expression construct (i.e., an AAV.RNAbc transcript); (c) dividing the first plurality of nuclei containing the cDNA molecules and AAV.RNAbc amplicons into at least two primary aliquots, the at least two primary aliquots including a first primary aliquot and a second primary aliquot; (d) providing primary barcode tags to the at least two primary aliquots, wherein the primary barcode tag provided to the first primary aliquot is different from the primary barcode tag provided to the second primary aliquot; (e) coupling the cDNA molecules and AAV.RNAbc amplicons within each of the at least two primary aliquots to the provided primary barcode tags; (f) combining the at least two primary aliquots; (g) dividing the combined primary aliquots into at least two secondary aliquots, the at least two secondary aliquots including a first secondary aliquot and a second secondary aliquot; (h) providing secondary barcode tags to the at least two secondary aliquots, wherein the secondary barcode tag provided to the first secondary aliquot is different from the secondary barcode tag provided to the second secondary aliquot; (i) coupling the cDNA molecules and AAV.RNAbc amplicons within each of the at least two secondary aliquots to the provided secondary barcode tags; (j) repeating steps (f), (g), (h), and (i) for subsequent aliquots, wherein the final barcode tag includes a capture agent; (k) combining the final aliquots; (I) lysing the first plurality of nuclei to release the cDNA molecules and AAV.RNAbc amplicons from within the first plurality of nuclei to form a lysate; and / or (m) adding a protease inhibitor and / or a binding agent to the lysate such that the cDNA molecules and AAV.RNAbc amplicons bind to the binding agent.
[0077] The method may further include dividing the combined final aliquot into at least two final aliquots, the at least two final aliquots including a first final aliquot and a second final aliquot. In some embodiments, the first plurality of nuclei may be fixed and permeabilized at a temperature below about 8 °C, below about 7 °C, below about 6 °C, below about 5 °C, at about 4 °C, below about 4 °C, below about 3 °C, below about 2 °C, below about 1 °C, or at another suitable temperature. In certain embodiments, the method may include splitting nuclei. For example, after the last or final round of barcoding (by ligation), the nuclei may be combined prior to lysis and subsequently divided into different lysate aliquots. Each lysate aliquot may include a predetermined number of nuclei.
[0078] Referring to, for example, step (m), the protease inhibitor may include phenylmethanesulfonyl fluoride (PMSF), 4-(2-aminoethyl)benzenesulfonyl fluoride hydrochloride (AEBSF), combinations thereof, and / or another suitable protease inhibitor. Referring to, for example, steps (j), (k), (l), and / or (m), the capture agent may include biotin or another suitable capture agent. Additionally, the binder may comprise avidin (e.g., streptavidin) or another suitable binder.
[0079] In certain embodiments, methods for uniquely labeling multiple intronic RNA molecules may further include (e.g., after step (m)): (n) performing template switching on the cDNA molecules and AAV.RNAbc amplicons bound to the binder using a template switch oligonucleotide; (o) amplifying the cDNA molecules and AAV.RNAbc amplicons to form a solution of amplified cDNA molecules and AAV.RNAbc amplicons; and / or (p) introducing a solid phase reversible immobilization (SPRI) bead solution into the solution of amplified cDNA molecules and AAV.RNAbc amplicons to remove polynucleotides less than about 200 base pairs, less than about 175 base pairs, or less than about 150 base pairs (see DeAngelis, M.M., et al. Nucleic Acids Research (1995) 23(22):4742). In other words, the cDNA molecules and AAV.RNAbc amplicons can bind to streptavidin beads within the lysate. Template switching can be performed on the cDNA molecules and AAV.RNAbc amplicons linked to the beads, for example, to add adaptors to the 3' ends of the cDNA molecules and AAV.RNAbc amplicons. The cDNA molecules and AAV.RNAbc amplicons can then be PCR amplified, followed by the addition of SPRI beads to remove polynucleotides less than about 200 base pairs. The ratio of the SPRI bead solution to the solution of amplified cDNA molecules can be from about 0.9:1 to about 0.7:1, from about 0.875:1 to about 0.775:1, from about 0.85:1 to about 0.75:1, from about 0.825:1 to about 0.725:1, from about 0.8:1, or other suitable ratios. In addition, the SPRI bead solution can contain from about 1 M to 4 M NaCl, from about 2 M to 3 M NaCl, from about 2.25 M to 2.75 M NaCl, about 2.5 M NaCl, or another suitable amount of NaCl. The SPRI bead solution can also contain from about 15% w / v to 25% w / v polyethylene glycol (PEG), where the molecular weight of the PEG is from about 7,000 g / mol to 9,000 g / mol (PEG 8000). In multiple embodiments, the SPRI bead solution can contain from about 17% w / v to 23% w / v PEG 8000, from about 18% w / v to 22% w / v PEG 8000, from about 19% w / v to 21% w / v PEG 8000, about 20% w / v PEG 8000, or another suitable % w / v of PEG 8000.
[0080] Methods for uniquely tagging multiple intranuclear RNA molecules can also include adding a common adapter sequence to the 3' ends of the released cDNA molecules and AAV.RNAbc amplicons. For each of the cDNA molecules and AAV.RNAbc amplicons (i.e., within a given experiment), the common adapter sequence can be the same or substantially the same adapter sequence. The addition of the common adapter can be performed or carried out in a solution containing up to about 10% w / v PEG, where the PEG has a molecular weight of about 7,000 g / mol to 9,000 g / mol. In certain embodiments, the common adapter sequence can be added to the 3' ends of the released cDNA molecules and AAV.RNAbc amplicons by template switching. (See Picelli, S, et al. Nature Methods 10, 1096-1098 (2013)).
[0081] Step (j) can be repeated a number of times sufficient to generate a unique barcode tag series for the nucleic acids in a single nucleus. For example, the number of times can be selected from: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, and 100.
[0082] In multiple embodiments, the primer of step (b) can also include a first specific barcode. In other words, the first barcode added to the cDNA molecules and AAV.RNAbc amplicons in a particular container, mixture, reaction, receptacle, sample, well, or vessel can be predetermined (e.g., specific to a given container, mixture, reaction, receptacle, sample, well, or vessel). For example, 96 different well-specific RT primer sets can be used (e.g., in a 96-well plate). Thus, if there are 96 samples or aliquots, each sample or aliquot can obtain a unique well-specific barcode.
[0083] In multiple embodiments, each of the barcode tags can include a first strand, where the first strand includes (i) a barcode sequence having a 3' end and a 5' end, and (ii) a 3' hybridization sequence and a 5' hybridization sequence flanking the 3' end and 5' end of the barcode sequence, respectively. Each of the barcode tags can also include a second strand, where the second strand includes (i) a first portion complementary to at least one of the 5' hybridization sequence and the adapter sequence, and (ii) a second portion complementary to the 3' hybridization sequence.
[0084] Methods for uniquely tagging RNA molecules within multiple nuclei may also include ligating at least two (or more) of the barcode tags that are bound to cDNA molecules and AAV.RNAbc amplicons. Ligation may be performed within the first plurality of nuclei.
[0085] The method may also include removing unbound barcode tags. In some embodiments, the method may include ligating at least two of the barcode tags that are bound to the released cDNA molecules and AAV.RNAbc amplicons. The cDNA molecules and AAV.RNAbc amplicons from a single nucleus to which most nucleic acid tags are bound may contain the same series of bound barcode tags.
[0086] cDNA molecules may be formed or generated in aliquots (e.g., reaction mixtures). The concentration of the first reverse transcription primer in the aliquot may be from about 0.5 μM to about 10 μΜ, from about 1 μM to about 7 μM, from about 1.5 μM to about 4 μM, from about 2 μM to about 3 μM, about 2.5 μM, or another suitable concentration. The concentration of the second reverse transcription primer in the aliquot may be from about 0.5 μM to about 10 μM, from about 1 μM to about 7 μM, from about 1.5 μM to about 4 μM, from about 2 μM to about 3 μM, about 2.5 μM, or another suitable concentration. III. Sequencing Library Preparation
[0087] In multiple embodiments, sequencing may be performed on a variety of sequencing platforms that require the preparation of a sequencing library. In the case of whole transcriptome sequencing, preparation typically involves fragmenting the cDNA (by sonication, nebulization, or shearing), followed by cDNA repair and end polishing (blunt end or overhang), and platform-specific adapter ligation. In one embodiment, the methods described herein may utilize next generation sequencing (NGS) techniques, which allow multiple samples to be sequenced individually as genomic molecules (i.e., singleplex sequencing) or as pooled samples that include indexed genomic molecules in a single sequencing run (e.g., multiplex sequencing). These methods can generate readouts of up to billions of DNA sequences. In multiple embodiments, the sequences of genomic nucleic acids and / or indexed genomic nucleic acids may be determined using, for example, the next generation sequencing (NGS) techniques described herein. In multiple embodiments, a large amount of sequence data obtained using NGS may be analyzed using one or more processors.
[0088] The preparation of a sequencing library for whole transcriptome sequencing is facilitated by fragmenting a polynucleotide (e.g., cDNA) to obtain polynucleotides within a desired size range.
[0089] Paired-end reads can be used in the sequencing methods and systems disclosed herein. The fragment or insert length is longer than the read length and is sometimes longer than the sum of two read lengths.
[0090] In some illustrative embodiments, the sample nucleic acid is obtained as cDNA, which is fragmented into fragments longer than about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 base pairs, to which NGS methods can be readily applied. In some embodiments, paired-end reads are obtained from inserts of about 100 to 5000 bp. In some embodiments, the insert is about 100 to 1000 bp long. These are sometimes implemented as conventional short-insert paired-end reads. In some embodiments, the insert is about 1000 to 5000 bp long.
[0091] Fragmentation can be achieved by any of a variety of methods known to those skilled in the art. For example, fragmentation can be achieved by mechanical means (including but not limited to nebulization, sonication, and hydroshear) or enzymatically. However, mechanical fragmentation typically cleaves the DNA backbone at C-O, P-O, and C-C bonds, resulting in a heterogeneous mixture of blunt ends and 3' and 5' overhangs with broken C-O, P-O, and / C-C bonds (see, e.g., Alnemri and Liwack, J Biol Chem 265:17323-17333
[1990] ; Richards and Boyer, J Mol Biol 11:327-240
[1965] ), which may need to be repaired because they may lack the 5'-phosphate necessary for subsequent enzymatic reactions (e.g., ligation of sequencing adapters), which is necessary for preparing DNA for sequencing.
[0092] Generally, DNA fragments are converted to blunt-end DNA with 5'-phosphate and 3'-hydroxyl. Standard protocols, for example, those used for sequencing on the Illumina platform as described in the Figure 2B example workflow, direct the user to perform end repair of the sample DNA; purify the end-repaired product, then adenylate or dA-tail the 3' end; and purify the dA-tailed product, then perform the adapter ligation step of library preparation.
[0093] Multiple embodiments of the sequence library preparation methods described herein eliminate the need to perform one or more of the steps typically specified in standard protocols to obtain a modified DNA product that can be sequenced by NGS. For example, for enriched AAV.RNAbc amplicon sequencing, fragmentation is not performed. Instead, PCR enrichment is performed using a forward primer specific for the AAV.RNAbc amplicon, and the forward primer has a 5' phosphate, thus eliminating the need to perform end repair. IV. Sequencing Methods
[0094] The methods and devices described herein can employ next-generation sequencing (NGS) technologies that permit large-scale parallel sequencing. In certain embodiments, clonally amplified DNA templates or single DNA molecules are sequenced in a large-scale parallel manner within a flow cell (e.g., as described in Volkerding et al. Clin Chem 55:641-658
[2009] ; Metzker M Nat Rev 11:31-46
[2010] ). NGS sequencing technologies include, but are not limited to, pyrosequencing, sequencing by synthesis with reversible dye terminators, sequencing by ligation of oligonucleotide probes, and ion semiconductor sequencing. DNA from individual samples can be sequenced separately (i.e., singleplex sequencing), or DNA from multiple samples can be pooled and sequenced as indexed genomic molecules in a single sequencing run (i.e., multiplex sequencing) to generate up to hundreds of millions of DNA sequence reads. Some examples of sequencing technologies that can be used to obtain sequence information according to the methods of the invention are further described herein.
[0095] Some sequencing technologies are commercially available, such as the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, Calif.) and the sequencing-by-synthesis platforms from 454 Life Sciences (Bradford, Conn.), Illumina / Solexa (Hayward, Calif.), and Helicos Biosciences (Cambridge, Mass.), and the sequencing-by-ligation platform from Applied Biosystems (Foster City, Calif.), as described below. In addition to single molecule sequencing using the sequencing by synthesis of Helicos Biosciences, other single molecule sequencing technologies include, but are not limited to, the SMRT TM technology of Pacific Biosciences, the ION TORRENT TM technology, and nanopore sequencing, such as that developed by Oxford Nanopore Technologies.
[0096] Although the automated Sanger method is considered a "first-generation" technology, Sanger sequencing, including automated Sanger sequencing, can also be used in the methods described herein. Additional suitable sequencing methods include, but are not limited to, nucleic acid imaging techniques such as atomic force microscopy (AFM) or transmission electron microscopy (TEM). Illustrative sequencing techniques are described in more detail below.
[0097] In some embodiments, the disclosed methods involve obtaining sequence information of nucleic acids in a test sample by performing massively parallel sequencing of millions of DNA fragments using Illumina's sequencing-by-synthesis and reversible terminator-based sequencing chemistries (e.g., as described in Bentley et al., Nature 6:53-59
[2009] ). The template DNA can be cDNA. In some embodiments, cDNA from isolated cells is used as a template and is fragmented into lengths of several hundred base pairs. In other embodiments, AAV.RNAbc amplicons are prepared by PCR amplification and do not require fragmentation. If desired, the template DNA is end-repaired to generate 5'-phosphorylated blunt ends. The polymerase activity of the Klenow fragment is used to add a single A base to the 3' end of the blunt-end phosphorylated DNA fragment. This addition prepares the DNA fragment for ligation to an oligonucleotide adapter that has a single T-base overhang at its 3' end to enhance ligation efficiency. The adapter oligonucleotide is complementary to the flow cell anchor oligonucleotide. Under conditions of limited dilution, the adapter-modified single-stranded template DNA is added to the flow cell and immobilized by hybridization to the anchor oligonucleotide. The ligated DNA fragments are extended and bridge amplified to generate an ultra-high density sequencing flow cell with hundreds of millions of clusters, each cluster containing approximately 1000 copies of the same template. In one embodiment, the adapter-ligated DNA is amplified by PCR before subjecting it to cluster amplification. In some applications, the template is sequenced using a robust four-color DNA sequencing-by-synthesis technique that employs reversible terminators with removable fluorescent dyes. Laser excitation and total internal reflection optics are used to achieve high-sensitivity fluorescence detection. Short sequence reads of approximately dozens to hundreds of base pairs are aligned to a reference genome, and specially developed data analysis pipeline software is used to identify the unique mapping of the short sequence reads relative to the reference genome. After completion of the first read, the template can be regenerated in situ so that a second read can be performed from the other end of the fragment. Thus, single-end or paired-end sequencing of DNA fragments can be used.
[0098] Multiple embodiments of the present disclosure may use synthetic sequencing that allows paired-end sequencing. In some embodiments, Illumina's synthetic sequencing platform involves fragment clustering. Clustering is a process in which each fragment molecule is isothermally amplified. In some embodiments, as in the examples described herein, the fragment has two different adapters ligated to the two ends of the fragment, which allow the fragment to hybridize to two different oligonucleotides on the surface of the flow cell lane. The fragment also contains two index sequences located at the two ends of the fragment or ligated to the two index sequences located at the two ends of the fragment, which provide tags for identifying different samples in multiplex sequencing. In some sequencing platforms, the fragment to be sequenced from both ends is also referred to as an insert.
[0099] In some embodiments, the flow cell used for clustering in the Illumina platform is a slide with lanes. Each lane is a glass channel coated with a lawn of two types of oligonucleotides (e.g., P5 and P7' oligonucleotides). Hybridization is enabled by the first of the two types of oligonucleotides on the surface. This oligonucleotide is complementary to the first adapter on one end of the fragment. The polymerase generates a complementary strand of the hybridized fragment. The double-stranded molecule is denatured, and the initial template strand is washed away. The remaining strands parallel to many other remaining strands are clonally amplified by bridge amplification.
[0100] In bridge amplification and other sequencing methods involving clustering, the strand folds over, and the second adapter region on the second end of the strand hybridizes to the second type of oligonucleotide on the surface of the flow cell. The polymerase generates a complementary strand, forming a double-stranded bridge molecule. The double-stranded molecule is denatured, producing two single-stranded molecules tethered to the flow cell by two different oligonucleotides. Then the process is continuously repeated and occurs simultaneously in millions of clusters, enabling the clonal amplification of all fragments. After bridge amplification, the reverse strand is cleaved and washed away, leaving only the forward strand. The 3' end is blocked to prevent unwanted priming.
[0101] After clustering, sequencing begins by extending the first sequencing primer to generate the first readout. In each cycle, fluorescently labeled nucleotides compete for addition to the growing strand. Based on the template sequence, only one nucleotide is incorporated. After adding each nucleotide, the cluster is excited by a light source and emits a characteristic fluorescent signal. The number of cycles determines the length of the readout. The emission wavelength and signal intensity determine the base call. For a given cluster, all identical strands are read out simultaneously. Hundreds of millions of clusters are sequenced in a massively parallel manner. When the first readout is complete, the readout product is washed away.
[0102] In the next step of the protocol involving two indexing primers, the index 1 primer is introduced and hybridized to the index 1 region on the template. The index region provides an identification of the fragment, which can be used to demultiplex samples during the multiplex sequencing process. The generation of the index 1 readout is similar to the first readout. After completion of the index 1 readout, the readout product is washed away, and the 3' end of the strand is deprotected. Then the template strand is flipped and bound to the second oligonucleotide on the flow cell. The index 2 sequence is read out in the same manner as index 1. Then, upon completion of this step, the index 2 readout product is washed away.
[0103] After reading the two indexes, read 2 starts by using polymerase to initiate the extension of the second flow cell oligonucleotide, forming a double-stranded bridge. The double-stranded DNA is denatured, and the 3' end is blocked. The initial forward strand is cleaved and washed away, leaving the reverse strand. Read 2 starts by introducing the read 2 sequencing primer. Similar to read 1, the sequencing steps are repeated until the desired length is achieved. The read 2 product is washed away. The entire process generates millions of readouts, representing all the fragments. Based on the unique indexes introduced during sample preparation, the sequences from the pooled sample library are separated. For each sample, the readouts of similar segments of base calling are locally clustered. The forward and reverse readouts are paired to generate a contiguous sequence. V. Kit
[0104] Another aspect of the present disclosure relates to a kit for labeling nucleic acids in at least a first cell. In some embodiments, the kit may comprise at least two reverse transcription primers comprising 5' overhang sequences. The kit may comprise at least one poly(T)-containing reverse transcription primer. The kit may comprise at least one AAV.RNAbc transcript-specific reverse transcription primer.
[0105] The kit may further comprise a plurality of first barcode tags. Each first barcode tag may comprise a first strand. The first strand may comprise a 3' hybridization sequence extending from the 3' end of the first marker sequence and a 5' hybridization sequence extending from the 5' end of the first marker sequence. Each first barcode tag may further comprise a second strand. The second strand may comprise an overhang sequence, wherein the overhang sequence may comprise (i) a first portion complementary to at least one of the 5' hybridization sequence and the 5' overhang sequence of the reverse transcription primer, and (ii) a second portion complementary to the 3' hybridization sequence.
[0106] The kit may also include a plurality of second barcode labels. Each second barcode label may include a first strand. The first strand may include a 3' hybridization sequence extending from the 3' end of the second marker sequence and a 5' hybridization sequence extending from the 5' end of the second marker sequence. Each second barcode label may also include a second strand. The second strand may include an overhang sequence, where the overhang sequence may include (i) a first portion complementary to at least one of the 5' hybridization sequence and the 5' overhang sequence of the reverse transcription primer, and (ii) a second portion complementary to the 3' hybridization sequence. In some embodiments, the first marker sequence may be different from the second marker sequence.
[0107] In some embodiments, the kit may also include one or more additional pluralities of barcode labels. Each barcode label in the one or more additional pluralities of barcode labels may include a first strand. The first strand may include a 3' hybridization sequence extending from the 3' end of the marker sequence and a 5' hybridization sequence extending from the 5' end of the marker sequence. Each barcode label in the one or more additional pluralities of barcode labels may also include a second strand. The second strand may include an overhang sequence, where the overhang sequence includes (i) a first portion complementary to at least one of the 5' hybridization sequence and the 5' overhang sequence of the reverse transcription primer, and (ii) a second portion complementary to the 3' hybridization sequence. In some embodiments, within each given additional plurality of barcode labels, the marker sequences may be different.
[0108] In a plurality of embodiments, the kit may also include at least one of reverse transcriptase, fixative, permeabilizing agent, ligating agent, and / or lysing agent. VI. Definitions
[0109] The terms "polynucleotide", "nucleic acid", and "transgene" are used interchangeably herein to refer to all forms of nucleic acids, oligonucleotides, including deoxyribonucleic acid (DNA) and ribonucleic acid (RNA) and their polymers. Polynucleotides include genomic DNA, cDNA, and antisense DNA, as well as spliced or unspliced mRNA, rRNA, tRNA, and inhibitory DNA or RNA (RNAi, e.g., small or short hairpin (sh) RNA, microRNA (miRNA), small or short interfering (si) RNA, trans-spliced RNA, or antisense RNA). Polynucleotides may include naturally occurring, synthetic, and intentionally modified or altered polynucleotides (e.g., variant nucleic acids). Polynucleotides may be single-stranded, double-stranded, or triple-stranded, linear or circular, and may have any suitable length. When discussing polynucleotides, the sequence or structure of a particular polynucleotide may be described herein according to the convention of providing the sequence in the 5' to 3' direction.
[0110] Nucleic acids encoding polypeptides generally contain an open reading frame encoding the polypeptide. Unless otherwise indicated, a particular nucleic acid sequence also contains degenerate codon substitutions.
[0111] The nucleic acid may contain one or more expression control or regulatory elements operably linked to the open reading frame, wherein the one or more regulatory elements are configured to direct the transcription and translation of the polypeptide encoded by the open reading frame in mammalian cells. Some non-limiting examples of expression control / regulatory elements include transcription initiation sequences (e.g., promoters, enhancers, TATA boxes, etc.), translation initiation sequences, mRNA stability sequences, poly A sequences, secretion sequences, etc. The expression control / regulatory elements can be obtained from the genomes of any suitable organisms.
[0112] A "promoter" refers to a nucleotide sequence that is generally located upstream (5') of the coding sequence and directs and / or controls the expression of the coding sequence by providing recognition for RNA polymerase and other factors required for correct transcription. Pol II promoters include minimal promoters, which are short DNA sequences containing a TATA box and optionally other sequences for specifying the transcription start site, to which regulatory elements are added to control expression. Type 1 pol III promoters contain three cis-acting sequence elements downstream of the transcription start site: a) a 5' sequence element (Block A); b) an intermediate sequence element (Block I); c) a 3' sequence element (Block C). Type 2 pol III promoters contain two essential cis-acting sequence elements downstream of the transcription start site: a) an A box (5' sequence element); and b) a B box (3' sequence element). Type 3 pol III promoters contain several cis-acting promoter elements upstream of the transcription start site, such as the conventional TATA box, proximal sequence element (PSE), and distal sequence element (DSE).
[0113] An "enhancer" is a DNA sequence that can stimulate transcriptional activity and can be an intrinsic element of a promoter or a heterologous element that enhances the expression level or tissue specificity. It can function in either orientation (5' -> 3' or 3' -> 5') and can act even when located upstream or downstream of the promoter.
[0114] The promoter and / or enhancer can be entirely derived from a natural gene, or be composed of different elements derived from different elements present in nature, or even contain synthetic DNA segments. A promoter or enhancer can contain DNA sequences involved in the binding of protein factors that regulate / control the efficiency of transcription initiation in response to stimuli, physiological, or developmental conditions.
[0115] Some non-limiting examples of promoters include the SV40 early promoter, the mouse mammary tumor virus LTR promoter; the adenovirus major late promoter (Ad MLP); the herpes simplex virus (HSV) promoter, the cytomegalovirus (CMV) promoter such as the CMV immediate early promoter region (CMVIE), the rous sarcoma virus (RSV) promoter, the pol II promoter, the pol III promoter, synthetic promoters, chimeric promoters, etc. Additionally, sequences derived from non-viral genes such as the murine metallothionein gene have also been found to be useful herein. Some exemplary constitutive promoters include promoters for the following genes encoding certain constitutive or "housekeeping" functions: hypoxanthine phosphoribosyl transferase (HPRT), dihydrofolate reductase (DHFR), adenosine deaminase, phosphoglycerol kinase (PGK), pyruvate kinase, phosphoglyceromutase, actin promoter, U6, and other constitutive promoters known to those of skill in the art. Additionally, many viral promoters function constitutively in eukaryotic cells. These include: the early and late promoters of SV40; the long terminal repeat (LTR) of Moloney leukemia virus and other retroviruses; and the thymidine kinase promoter of herpes simplex virus, etc. Additionally, sequences derived from intronic miRNA promoters, such as the miR107, miR206, miR208b, miR548f-2, miR569, miR590, miR566, and miR128 promoters have also been found to be useful herein (see, e.g., Monteys et al., 2010). Thus, any of the constitutive promoters mentioned above can be used to control the transcription of heterologous gene inserts.
[0116] As used herein, a "transgene" conveniently refers to a nucleic acid sequence / polynucleotide that is intended or has been introduced into a cell or organism. Transgenes include any nucleic acid, such as a gene encoding a barcode, and are typically heterologous relative to the naturally occurring AAV genomic sequence.
[0117] The term "transduction" refers to the introduction of a nucleic acid sequence into a cell or host organism by means of a vector (e.g., a viral particle). Thus, the introduction of a transgene into a cell by a viral particle can be referred to as "transduction" of the cell. The transgene may or may not be integrated into the genomic nucleic acid of the transduced cell. If the introduced transgene is integrated into the nucleic acid (genomic DNA) of the recipient cell or organism, it can be stably maintained in that cell or organism and further passed on to the progeny cells or organisms of the recipient cell or organism or inherited by them. Finally, the introduced transgene may be extrachromosomal in the recipient cell or host organism or may be present only transiently. Thus, a "transduced cell" is a cell into which a transgene has been introduced by transduction. Thus, a "transduced" cell is a cell or its progeny into which a transgene has been introduced. Transduced cells can proliferate, transcribe the transgene, and express the encoded inhibitory RNA or protein. For gene therapy uses and methods, the transduced cells can be in a mammal.
[0118] A nucleic acid / transgene is "operably linked" when it is placed in a functional relationship with another nucleic acid sequence. The nucleic acid / transgene encoding a barcode or the nucleic acid directing the expression of a polypeptide may contain an inducible promoter or a tissue-specific promoter for controlling the transcription of the encoded polypeptide. A nucleic acid operably linked to an expression control element may also be referred to as an expression cassette.
[0119] As used herein, the terms "modified" or "variant" and their grammatical variations mean that a nucleic acid, polypeptide, or a subsequence thereof deviates from a reference sequence. Thus, a modified sequence and a variant sequence may have expression, activity, or function that is substantially the same as, higher than, or lower than that of the reference sequence, but at least retain some activity or function of the reference sequence. A particular type of variant is a mutant protein, which refers to a protein encoded by a gene having a mutation such as a missense or nonsense mutation.
[0120] A "nucleic acid" or "polynucleotide" variant refers to a modified sequence that has been genetically altered compared to the wild type. The sequence can be genetically modified without changing the encoded protein sequence. Alternatively, the sequence can be genetically modified to encode a variant protein. A nucleic acid or polynucleotide variant can also refer to a combinatorial sequence that has been codon-modified to encode a protein that still retains at least partial sequence identity with a reference sequence (e.g., a wild-type protein sequence) and has also been codon-modified to encode a variant protein. For example, some codons of such a nucleic acid variant will be changed without changing the amino acids of the protein encoded by it, and some codons of the nucleic acid variant will be changed, which in turn changes the amino acids of the protein encoded by it.
[0121] The terms "protein" and "polypeptide" are used interchangeably herein. The "polypeptides" encoded by "nucleic acids" or "polynucleotides" or "transgenes" disclosed herein include partial or full-length native sequences, such as naturally occurring wild-type and functional polymorphic proteins, their functional subsequences (fragments), and sequence variants thereof, provided that the polypeptide retains a certain degree of function or activity. Thus, in the methods and uses of the present invention, such polypeptides encoded by nucleic acid sequences need not be identical to the endogenous protein that is defective or has insufficient, absent, or non-existent activity, function, or expression in the mammal being treated.
[0122] Some non-limiting examples of modifications include one or more nucleotide or amino acid substitutions (e.g., about 1 to about 3, about 3 to about 5, about 5 to about 10, about 10 to about 15, about 15 to about 20, about 20 to about 25, about 25 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 100, about 100 to about 150, about 150 to about 200, about 200 to about 250, about 250 to about 500, about 500 to about 750, about 750 to about 1000 or more nucleotides or residues).
[0123] An example of an amino acid modification is a conservative amino acid substitution or deletion. In some specific embodiments, the modified sequence or variant sequence retains at least part of the function or activity of the unmodified sequence (e.g., wild-type sequence).
[0124] Another example of an amino acid modification is the introduction of a targeting peptide into the capsid protein of a viral particle. Peptides that target the central nervous system (e.g., different brain regions) have been identified for recombinant viral vectors.
[0125] A "variant" of a molecule is a sequence that is substantially similar to the sequence of the native molecule. For nucleotide sequences, variants include those sequences that, due to the degeneracy of the genetic code, encode the same amino acid sequence as the native protein. For example, naturally occurring allelic variants such as these can be determined using molecular biology techniques such as, for example, polymerase chain reaction (PCR) and hybridization techniques. Variant nucleotide sequences also include nucleotide sequences of synthetic origin, such as those produced by site-directed mutagenesis (which encode the native protein), and those that encode polypeptides having amino acid substitutions. In general, nucleotide sequence variants of the present invention will have at least 40%, 50%, 60% to 70%, such as 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78% to 79%, typically at least 80%, such as 81% to 84%, at least 85%, such as 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97% to 98% sequence identity with the native (endogenous) nucleotide sequence. In certain embodiments, the variant is biologically functional (i.e., retains 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or 100% of the activity or function of the wild type).
[0126] "Conservative variation" of a particular nucleic acid sequence refers to those nucleic acid sequences that encode the same or substantially the same amino acid sequence. Due to the degeneracy of the genetic code, a large number of functionally identical nucleic acids encode any given polypeptide. For example, the codons CGT, CGC, CGA, CGG, AGA, and AGG all encode the amino acid arginine. Thus, at each position where arginine is specified by a codon, that codon can be changed to any of the corresponding codons described without changing the encoded protein. Such nucleic acid variations are "silent variations" and are a type of "conservative modification variation". Unless otherwise indicated, each nucleic acid sequence described herein that encodes a polypeptide also describes every possible silent variation. One of ordinary skill in the art will recognize that each codon in a nucleic acid (except ATG, which is typically the only codon for methionine) can be modified by standard techniques to produce a functionally identical molecule. Thus, every "silent variation" of a nucleic acid encoding a polypeptide is implicit in each described sequence.
[0127] "Substantial identity" of a polynucleotide sequence means that the polynucleotide comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78% or 79%, or at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88% or 89%, or at least 90%, 91%, 92%, 93% or 94%, or even at least 95%, 96%, 97%, 98% or 99% sequence identity compared to a reference sequence using one of the alignment programs with standard parameters. Those skilled in the art will recognize that these values can be appropriately adjusted to determine the corresponding identity of the proteins encoded by two nucleotide sequences by considering codon degeneracy, amino acid similarity, reading frame positioning, etc. Substantial identity of amino acid sequences for these purposes generally means at least 70%, at least 80%, 90% or even at least 95% sequence identity.
[0128] In the case of polypeptides, the term "substantial identity" means that within a specified comparison window, the polypeptide comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, or 79%, or 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88% or 89%, or at least 90%, 91%, 92%, 93% or 94%, or even 95%, 96%, 97%, 98% or 99% sequence identity to a reference sequence. An indication that two polypeptide sequences are identical is that one polypeptide is immunoreactive with an antibody raised against the second polypeptide. Thus, for example, a polypeptide is identical to a second polypeptide where the two peptides differ only by conservative substitutions.
[0129] As used herein, "substantially free of" with respect to a specified component is used to mean that the specified component has not been purposefully formulated into the composition and / or is present only as a contaminant or in trace amounts. Thus, the total amount of the specified component due to any accidental contamination of the composition is far less than 0.05%, preferably less than 0.01%. Most preferably, the composition is one in which the amount of the specified component cannot be detected by standard analytical methods.
[0130] As used herein, a noun not qualified by a quantifier in the specification may mean one or more. As used herein in the claims, when used in conjunction with the word "comprising", a noun not qualified by a quantifier may mean one or more than one.
[0131] Unless expressly indicated to refer to alternatives only or that the alternatives are mutually exclusive, the use of the term "or" in a claim is used to mean "and / or", but the present disclosure supports definitions that refer to alternatives only and to "and / or". As used herein, "another" can mean at least a second or more.
[0132] Throughout this application, the term "about" is used to indicate that a value includes the inherent error variations of the device used to determine the value, the inherent variations of the method used to determine the value, the variations that exist between the subjects of study, or values within 10% of the value. VII. Examples
[0133] The following examples are included to illustrate some preferred embodiments of the invention. Those skilled in the art will understand that the techniques disclosed in the following examples represent techniques discovered by the inventors to function well in the practice of the invention and thus can be considered to constitute preferred modes for its practice. However, those skilled in the art will understand from the present disclosure that many changes can be made to the specific embodiments disclosed without departing from the spirit and scope of the invention and still obtain the same or similar results. Example 1 - Design of a Dual-Barcode Containing AAV Cargo
[0134] One of the major challenges in detecting transduction of barcoded AAV capsids at the single-cell level is detecting the mRNA sequences that provide information about cell identity and simultaneously detecting both the delivered AAV capsid DNA or expressed RNA. To overcome this challenge, the inventors engineered an AAV-delivered expression construct that uses the human U6 promoter to drive robust expression of the barcode sequence (Figures 1A to B). The barcode sequences designed by the inventors can be detected by a variety of methods, including amplicon sequencing, single-cell RNA sequencing, and in situ sequencing. The construct shown in Figure 1A was packaged into the AAV genome, as shown in Figures 1B and D. As an example, the complete DNA sequence from such a construct is included in SEQ ID NO:14, where nucleotides 1 to 141 are the AAV-1 ITR, nucleotides 148 to 404 are the human U6 promoter, nucleotides 187 to 207 are the binding site for Pr766, nucleotide 404 is the human U6 transcription start site, nucleotides 413 to 434 are the binding site for the BC0108 scramble primer, nucleotides 441 to 456 are the 3’ padlock sequence, nucleotides 457 to 464 are the 9nt RNA barcode, nucleotides 466 to 488 are the 5’ padlock sequence, nucleotides 495 to 516 are the binding site for Split-seq Pr Rev, nucleotides 525 to 705 are the Pr40 promoter, nucleotides 1028 to 3274 are the coding sequence for the AAV1 capsid, nucleotides 2807 to 2827 are the coding sequence for the (NNK)7 peptide, nucleotides 3117 to 3139 are the binding site for Pr521 Rev, and nucleotides 3408 to 3548 are the AAV2 ITR. Additionally, the AAV Cap gene sequence has been modified to include a peptide insertion. Importantly, each modified Cap sequence is paired with a single RNA barcode “RNAbc”. These pairings are resolved by long-read sequencing that captures both the RNAbc and the Cap insertion sequence. Figure 1C provides an example of successful amplification of only the RNAbc sequence after reverse transcription using primers pr749 and 750, in contrast to non-specific amplification observed with other primer sets. Sanger sequencing across each insert confirmed the successful creation of this dual-barcode construct ( Figure 1E and F). Table 2. Primer sequences SEQ ID NO: 14 Example 2 - Adaptation of Split - Pool Ligation - based Transcriptome Sequencing (SPLiT - seq) for AAV.RNAbc Detection.
[0135] As Figure 2A shown, during three rounds of barcoding, fixed and permeabilized nuclei were randomly assigned to each well of a 96 - well plate. Each well of the first - round barcoding plate corresponded to two barcoding primers: an oligo(dT)RT primer and an AAV - specific RT primer to capture AAV.RNAbc transcripts. During the first round, both poly(A) and AAV.RNAbc transcripts from the same nucleus were reverse - transcribed and labeled with the same first - round barcode. All nuclei from the same well would receive the same first - round barcode, allowing sample information to be encoded by the first - round well position. After reverse - transcription, in the second and third rounds, all nuclei were pooled and randomly re - assigned. The barcoding in the second and third rounds consisted of ligation reactions to add additional single - nucleus barcodes. The third - round barcoding also added Unique Molecular Identifiers (UMIs). After three rounds of barcoding in the 96 - well plate, there could be 884,736 nucleus - barcode combinations (96×96×96). All nuclei were pooled and split into sub - libraries of <10,000 nuclei, followed by lysis, de - crosslinking, and streptavidin - bead - based cDNA isolation.
[0136] Figure 2B Next - generation library preparation for whole - transcriptome and AAV.RNAbc amplicon sequencing is shown. A template - switching reaction was performed to add a 5′ common sequence for full - length cDNA amplification. After amplification, the library was split for sequencing of the whole - transcriptome and AAV.RNAbc amplicons from the same nucleus. The library for whole - transcriptome sequencing was fragmented, followed by end - repair, A - tailing, and adapter ligation (see Figure 2B Section I). A final PCR was performed to add Illumina adapters and dual indices. Paired - end Illumina sequencing was performed. Read 1 contained both mRNA and AAV.RNAbc sequence information and Read 2 corresponded to the single - nucleus barcode for downstream demultiplexing. To enrich AAV.RNAbc sequences, a second PCR - based amplification was performed upstream of the AAV barcode with a 5′ primer sequence (see Figure 2BPart II). In addition to enrichment, this step also controls the size and starting position of the AAV.RNAbc amplicon. The forward primer is also modified to add a phosphate to the PCR product, allowing subsequent ligation of the Illumina sequencing adapters. Then, poly(A) tailing and adapter ligation are performed, followed by a final PCR to add the Illumina adapters and sample indices. Paired-end Illumina sequencing is performed. Here, read 1 output corresponds to the AAV.RNAbc amplicon sequence and read 2 corresponds to a single nuclear barcode for downstream demultiplexing. Given that the cDNA libraries from the same nucleus are split after amplification and processed simultaneously by two library preparations, the single nuclear barcode is used to identify the cell type identity of the cells expressing the AAV.RNAbc transcript. Example 3 - Detection of Expressed RNA Barcodes (RNAbc) after Transfection of HEK 293 Cells
[0137] HEK 293 cells were transfected with plasmids containing AAV.RNAbc or AAV.barcode-free. Uniform manifold approximation and projection (UMAP) unbiased clustering of SPLiT-Seq barcoded single cells is shown in Figure 3A. After performing amplification and library preparation dedicated to the RNAbc amplicon, unique UMI counts were obtained from the Illumina sequencing reads (Figure 3B). As expected, the vast majority of the counts belong to single cells that received the AAV.RNAbc construct. A small number of counts were found to originate from cells treated with AAV.barcode-free, likely representing binucleated cells. This is consistent with the 0.1% doublet rate of SPLiT-Seq. The Seurat single cell object is a subset that contains only cells that received the AAV.RNAbc treatment (Figure 3C) or only cells that received the AAV.eGFP construct (Figure 3D). When the Seurat single cell object is a subset that includes only cells that received the AAV.eGFP construct, a background level of counts originating from the AAV.RNAbc amplicon was observed. Further filtering in a non-single culture background will help remove this background. Example 4 - In Vivo Detection of 67 AAV1 Capsid Variants in 67 out of 1430 of 6739 Nuclei after Delivery of an AAV Variant Library to the Mouse Brain
[0138] A library of 67 AAV1 capsid variants was delivered to the mouse brain by intra-striatal and intra-thalamic injection. After incubation, the brain tissue was recovered, and the hippocampal, thalamic, striatal, and cortical regions were microdissected and subjected to single nuclear isolation ( Figure 4A) Fixed and permeabilized nuclei were processed using the dual Poly(A) and -AAV RNAbc reverse transcription barcoding method disclosed herein. SPLiT-Seq-based barcoding was then performed to apply single-cell barcodes to mRNAs and AAV transcripts contained within the permeabilized nuclei. After adding three single-cell barcodes, the nuclei were split into two pools and lysed. In parallel, the barcoded mRNAs and AAV RNAs were then amplified and Illumina indices and sequencing adapters were added to facilitate sequencing on an Illumina NovaSeq 6000.
[0139] The resulting fastq files were processed using a custom bioinformatics pipeline that integrated the RNAbc sequences with AAV peptide insertion information obtained from long-read sequencing of the input capsid library. This allowed the detected RNA-bcs to be converted into counts of AAV peptide insertions. A count table of cDNA counts / gene and RNA-bc counts / gene was generated and read into Seurat for downstream single-cell analysis and capsid transduction quantification.
[0140] Using single-cell barcodes that match between the cDNA and AAV-derived datasets, the inventors were able to create a single Seurat object containing both cDNA expression and AAV transduction information (Figure 4B). This enabled the inventors to determine which single cells had been transduced (Figure 4C). Then, by observing cell types associated with the disease, such as Drd1- and Drd2-positive medium spiny neurons (MSNs), the inventors were able to evaluate the AAV transduction status within the single-cell types of interest (Figure 4D).
[0141] Using this dataset and future datasets generated using this technology, the inventors were able to examine the tissues used in the study to identify which capsids perform best in the tissue of interest (Figure 5A) and the single-cell types of interest (Figure 5B). This information can also be represented spatially using UMAP unbiased clustering plots (Figure 5C). These tools enable the identification of capsid variants with performance characteristics suitable for disease treatment. ***
[0142] According to the present disclosure, all of the methods disclosed and claimed herein can be made and executed without undue experimentation. Although the compositions and methods of the present invention have been described in accordance with some preferred embodiments, it will be apparent to those skilled in the art that changes can be made in the methods described herein and in the steps or the order of the steps of the methods without departing from the concept, spirit, and scope of the present invention. More specifically, it will be apparent that certain chemically and physiologically related substances can be substituted for the substances described herein, while the same or similar results will be obtained. All such similar substitutions and modifications that are apparent to those skilled in the art are considered to be within the spirit, scope, and concept of the present invention as defined by the appended claims. References The following references, to the extent that they provide exemplary procedures or other details supplementary to those described herein, are specifically incorporated herein by reference. Rosenberg et al., “SPLiT-seq reveals cell types and lineages in the developing brain and spinal cord,” Science, 360:176-182, 2018.
Claims
1. A population of recombinant adeno-associated virus (rAAV) vectors, wherein each rAAV vector independently comprises: (i) a modified adeno-associated virus (AAV) Cap gene encoding a modified AAV capsid protein containing a targeting peptide, and (ii) an expression cassette encoding a barcode sequence operably linked to an RNA polymerase III promoter, wherein the targeting peptide and the barcode in each rAAV vector are uniquely paired.
2. The population of rAAV vectors according to claim 1, wherein the barcode sequence is 9 to 20 nucleotides in length.
3. The population of rAAV vectors according to claim 1 or 2, wherein the barcode sequence is (NNNT) n 。 4. The population of rAAV vectors according to any one of claims 1 to 3, wherein the barcode sequence is flanked by sequences capable of hybridizing to and activating a padlock probe.
5. The population of rAAV vectors according to any one of claims 1 to 4, wherein the RNA polymerase III promoter is a type III RNA polymerase III promoter.
6. The population of rAAV vectors according to claim 8, wherein the RNA polymerase promoter is a U6 snRNA gene promoter, an H1 RNA gene promoter, or a 7SK gene promoter.
7. The population of rAAV vectors according to any one of claims 1 to 6, further comprising a reverse transcription primer binding site located 3' of the barcode sequence and an enrichment primer binding site located 5' of the barcode sequence.
8. The population of rAAV vectors according to any one of claims 1 to 7, wherein the expression cassette comprises a sequence having at least 90% identity to SEQ ID NO:
7.
9. The population of rAAV vectors according to any one of claims 1 to 8, wherein the modified AAV capsid protein is a modified AAV1 capsid protein, a modified AAV2 capsid protein, or a modified AAV9 capsid protein.
10. The population of rAAV vectors according to claim 9, wherein the modified AAV capsid protein is derived from the AAV1 capsid protein (see SEQ ID NO:1), and wherein the targeting peptide is inserted after the 590th residue of the AAV1 capsid protein.
11. The population of rAAV vectors according to claim 10, wherein the targeting peptide is flanked by linker sequences, and wherein the linker sequences on each side of the targeting peptide are two or three amino acids in length.
12. The population of rAAV vectors according to claim 11, wherein the linker sequences are SSA on the N-terminal side of the targeting peptide and AS on the C-terminal side of the targeting peptide.
13. The group of rAAV vectors according to any one of claims 10 to 12, wherein the modified AAV1 capsid protein has a sequence having at least 95% identity with SEQ ID NO:
4.
14. The group of rAAV vectors according to claim 9, wherein the modified AAV capsid protein is derived from the AAV2 capsid protein (see SEQ ID NO: 2), and wherein the targeting peptide is inserted after the 587th residue of the AAV2 capsid protein.
15. The group of rAAV vectors according to claim 14, wherein the targeting peptide is flanked by linker sequences, and wherein the linker sequences on each side of the targeting peptide are two or three amino acids in length.
16. The group of rAAV vectors according to claim 15, wherein the linker sequences are AAA on the N-terminal side of the targeting peptide and AA on the C-terminal side of the targeting peptide.
17. The group of rAAV vectors according to any one of claims 14 to 16, wherein the modified AAV2 capsid protein has a sequence having at least 95% identity with SEQ ID NO:
5.
18. The group of rAAV vectors according to claim 9, wherein the modified AAV capsid protein is derived from the AAV9 capsid protein (see SEQ ID NO: 3), and wherein the targeting peptide is inserted after the 588th residue of the AAV9 capsid protein.
19. The group of rAAV vectors according to claim 18, wherein the targeting peptide is flanked by linker sequences, and wherein the linker sequences on each side of the targeting peptide are two or three amino acids in length.
20. The group of rAAV vectors according to claim 19, wherein the linker sequences are AAA on the N-terminal side of the targeting peptide and AS on the C-terminal side of the targeting peptide.
21. The group of rAAV vectors according to any one of claims 18 to 20, wherein the modified AAV9 capsid protein has a sequence having at least 95% identity with SEQ ID NO:
6.
22. The group of rAAV vectors according to any one of claims 8 to 21, wherein the targeting peptide is 3 to 10 amino acids in length.
23. The group of rAAV vectors according to claim 22, wherein the targeting peptide is 7 amino acids in length.
24. The group of rAAV vectors according to any one of claims 8 to 23, wherein the group comprises multiple capsid protein targeting peptides, and wherein each capsid protein targeting peptide is paired with more than one barcode sequence.
25. An rAAV vector population according to any one of claims 8 to 24, wherein the population comprises a plurality of capsid protein targeting peptides, and wherein all rAAVs having the same barcode sequence also have the same capsid protein targeting peptide.
26. A cell population comprising the rAAV vector population according to any one of claims 1 to 25.
27. The cell population according to claim 26, wherein the cells are mammalian cells.
28. The cell population according to claim 26, wherein the cells are human cells.
29. The cell population according to claim 26, wherein the cells are in vitro.
30. The cell population according to claim 26, wherein the cells are in vivo.
31. A method for determining the cell tropism of a recombinant adeno-associated virus (rAAV) having a modified AAV capsid protein comprising a targeting peptide, the method comprising: (i) Contacting a plurality of cell types with the modified rAAV vector according to any one of claims 6 to 25; (ii) Identifying the cells transduced by the modified rAAV vector based on the presence of the barcode sequence; And (iii) Detecting the transcriptome expressed by each transduced cell on a cell-by-cell basis to determine the cell tropism of the modified rAAV.
32. A method for determining the cell tropism of a recombinant adeno-associated virus (rAAV) having a modified AAV capsid protein comprising a targeting peptide, the method comprising: (i) Contacting a plurality of cell types with a population of rAAV vectors according to any one of claims 6 to 25; (ii) Detecting both the expressed transcriptome and the rAAV transducing each cell on a cell-by-cell basis; And (iii) Determining which cell types are transduced by which modified rAAV vectors to determine the cell tropism of the modified rAAV.
33. The method according to claim 32, wherein the contact in (i) is performed in vitro.
34. The method according to claim 32, wherein the contact in (i) is performed in vivo.
35. The method according to any one of claims 32 to 34, wherein detecting the expressed transcriptome in (ii) and the rAAV comprises: (a) Isolating, fixing, and permeabilizing the nuclei of the cells contacted in (i); (b) Dividing the nuclei into a plurality of first aliquots; (c) Reverse transcribing the cellular RNA molecules expressed in the nuclei using primers containing a poly(T) sequence to form complementary DNA (cDNA) molecules, and reverse transcribing the rAAV RNA molecules expressed in the nuclei using primers containing a sequence sufficient to hybridize to and reverse transcribe the barcode sequence within the expression cassette to form AAV amplicons; (d) Labeling the cDNA molecules and AAV amplicons with a first 5' barcode, wherein the first 5' barcode of the primers in each first aliquot is unique, such that the cDNA molecules and AAV amplicons from the nuclei of each aliquot can be identified compared to the cDNA molecules and AAV amplicons from the nuclei of all other aliquots; (e) Combining the plurality of first aliquots; (f) Dividing the combined plurality of first aliquots into a plurality of second aliquots; (g) Ligating a second 5' barcode to the 5' ends of the cDNA molecules and the AAV amplicons to form doubly barcoded cDNA molecules and AAV amplicons, wherein the second 5' barcode in each second aliquot is unique; (h) Combining the plurality of second aliquots; (i) Dividing the combined plurality of first aliquots into a plurality of third aliquots; (j) Ligating a third 5' barcode to the 5' ends of the cDNA molecules and the AAV amplicons to form triply barcoded cDNA molecules and AAV amplicons, wherein the third 5' barcode in each third aliquot is unique; (k) Combining the plurality of third aliquots; (l) Lysing the nuclei to release the cDNA molecules and the AAV amplicons from within the nuclei to form a lysate; And (m) Sequencing the cDNA molecules and the AAV amplicons to detect both the expressed transcriptome and the rAAV transducing each cell.
36. The method according to claim 35, wherein the cDNA molecules and AAV amplicons are labeled with the first 5' barcode while performing the reverse transcription, and the reverse transcription primer comprises the first 5' barcode.
37. The method according to claim 35 or 36, wherein the nuclei are fixed and permeabilized at a temperature below about 8°C, below about 7°C, below about 6°C, below about 5°C, below about 4°C, below about 3°C, below about 2°C or below about 1°C.
38. The method according to any one of claims 35 to 37, wherein most of the triply barcoded cDNA molecules and AAV molecules from a single nucleus comprise the same series of barcodes.
39. The method according to claim 38, wherein most of the triply barcoded cDNA molecules and AAV molecules from a single nucleus have a unique series of barcodes compared to the triply barcoded cDNA molecules and AAV molecules from other nuclei.
40. The method according to any one of claims 32 to 39, wherein the cell type is determined based on the expressed transcriptome.
41. The method according to any one of claims 32 to 40, wherein sequencing the cDNA molecules and the AAV amplicons comprises preparing a sequencing library, and preparing the sequencing library comprises: (i) Adding a common adapter sequence to the 3' ends of the cDNA molecules and AAV amplicons; (ii) Amplifying the full-length cDNA and AAV amplicons; (iii) Fragmenting the amplified full-length cDNA and AAV amplicons; (iv) Perform end repair and A-tailing on fragmented cDNA and AAV amplicons; (v) Ligate adapters to the 5'-ends of the end-repaired and A-tailed cDNA and AAV amplicons; and (vi) Perform sample indexing PCR to add sequencing adapters and dual indices to the adapter-ligated cDNA and AAV amplicons.
42. The method according to any one of claims 32 to 41, wherein sequencing the AAV amplicons comprises preparing a sequencing library that enriches the AAV amplicons, and preparing the sequencing library comprises: (i) Add a common adapter sequence to the 3'-ends of the cDNA molecules and AAV amplicons; (ii) Perform full-length amplification of the cDNA and AAV amplicons; (iii) Perform enrichment amplification of the AAV amplicons with a forward primer that hybridizes upstream of the AAV barcode in the AAV amplicon, wherein the forward primer has a 5'-phosphate; (iv) Perform A-tailing on the AAV amplicon with a 5'-phosphate; (v) Ligate adapters to the 5'-ends of the A-tailed AAV amplicons; and (vi) Perform sample indexing PCR to add sequencing adapters and dual indices to the adapter-ligated AAV amplicons.
43. The method according to claim 41 or 42, wherein the common adapter sequence is added to the 3' ends of the cDNA molecules and AAV amplicons by template switching.
44. The method according to any one of claims 35 to 43, wherein the sequencing is paired-end sequencing, amplicon sequencing, single-cell RNA sequencing or in situ sequencing.