AAV evolution at single-cell resolution using SPLiT-seq
By introducing a modified capsid protein with a targeting peptide and a barcode sequence connected to an RNA polymerase III promoter into AAV, the difficulty of AAV capsid transduction detection at the single-cell level was solved, and efficient identification and transduction detection of cell types were achieved.
Patent Information
- Application Number
- JP2025516100
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-19
- Filing Date
- 2023-09-19
- Publication Date
- 2025-10-01
AI Technical Summary
Existing technologies make it difficult to simultaneously detect the mRNA sequence transduced by AAV capsid and the delivered AAV capsid DNA or expressed RNA at the single-cell level, and are unable to effectively identify specific cell types transduced by barcoded AAV capsid.
Provided is a recombinant adeno-associated virus (rAAV) vector containing a modified AAV capsid protein with a targeting peptide, comprising an expression cassette with a barcode sequence operably linked to an RNA polymerase III promoter. By packaging in AAV and delivering it into cells, combined with specific library preparation technology and sequencing methods, cell type identification and AAV transduction detection at the single-cell level can be achieved.
It achieves efficient detection of AAV capsid transduction at the single-cell level, can identify and distinguish the transduction conditions of different cell types, and improves the accuracy and efficiency of cell type identification.
Smart Images

Figure 2025532627000006 
Figure 2025532627000007 
Figure 2025532627000008
Abstract
Description
[Technical Field]
[0001] REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 407,826, filed September 19, 2022, the entire contents of which are incorporated herein by reference.
[0002] Sequence Listing Reference This application contains a Sequence Listing XML, which has been submitted electronically and is incorporated herein by reference in its entirety. The Sequence Listing XML, created on September 19, 2023, is named CHOPP0057WO_ST26.xml and is 28,221 bytes in size.
[0003] background 1. Field The present invention relates generally to the fields of molecular biology, virology, and medicine, and more particularly to compositions and methods for determining the cell tropism of AAV capsid proteins bearing targeting peptides. [Background technology]
[0004] 2. Description of Related Technical Fields One of the main challenges in detecting barcoded AAV capsid transduction at the single-cell level is both detecting mRNA sequences that provide information about cell identity and simultaneously detecting the delivered AAV capsid DNA or expressed RNA. Methods that allow for the identification of specific cell types transduced by a given barcoded AAV capsid are needed. Summary of the Invention
[0005] overview Provided herein is a barcoded RNA expression construct, which is packaged in AAV and delivered to cells, expressing an mRNA sequence with a different identification barcode from the modified region of the capsid DNA sequence.In one embodiment, a recombinant adeno-associated virus (rAAV) vector is provided, which comprises an expression cassette encoding a barcode sequence that is operably linked to an RNA polymerase III promoter.In one embodiment, a population of recombinant adeno-associated virus (rAAV) vectors, each rAAV vector comprises an expression cassette encoding a barcode sequence that is operably linked to an RNA polymerase III promoter.The population of vectors can comprise a barcode sequence to vector ratio of 1:1, 1:5, 1:10, 1:50, 1:100, 1:500, or at least 1:1000. In one aspect, provided herein is a population of recombinant adeno-associated virus (rAAV) vectors, wherein each rAAV vector independently comprises (i) a modified adeno-associated virus (AAV) Cap gene encoding a modified AAV capsid protein comprising a targeting peptide, and (ii) an expression cassette encoding a barcode sequence operably linked to an RNA polymerase III promoter, wherein each targeting peptide and each barcode is uniquely paired.
[0006] The barcode sequence can be at least 9, at least 12, at least 15, at least 18, or at least 21 nucleotides in length. The barcode sequence can be 9-21 nucleotides in length, 12-21 nucleotides in length, 15-21 nucleotides in length, 9-18 nucleotides in length, 9-15 nucleotides in length, or 9-12 nucleotides in length. The barcode sequence can be 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 nucleotides in length. The barcode sequence can be flanked by sequences that can hybridize to and activate a padlock probe. Each barcode sequence, independently, is (NNNT) n Contains arrays.
[0007] The RNA polymerase III promoter can be a type III RNA polymerase III promoter. The RNA polymerase promoter can be a U6 snRNA gene promoter, an H1 RNA gene promoter, or a 7SK gene promoter.
[0008] The rAAV vector may further comprise a reverse transcription primer binding site located 3' of the barcode sequence and an enrichment primer binding site located 5' of the barcode sequence. The expression cassette may comprise a sequence identical, at least 90% identical, or at least 95% identical to SEQ ID NO:7.
[0009] The rAAV vector may further comprise a modified adeno-associated virus (AAV) Cap gene encoding a modified AAV capsid protein containing a targeting peptide. The modified AAV capsid protein may be a modified AAV1 capsid protein, a modified AAV2 capsid protein, or a modified AAV9 capsid protein. The targeting peptide may be 3 to 10 amino acids in length. The targeting peptide may be 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in length.
[0010] When the modified AAV capsid protein is derived from the AAV1 capsid protein (see SEQ ID NO: 1), the targeting peptide can be inserted after residue 590 of the AAV1 capsid protein. The targeting peptide can be flanked by linker sequences, each of which is 2 or 3 amino acids long. The linker sequence can be SSA at the N-terminus of the targeting peptide and AS at the C-terminus of the targeting peptide. The modified AAV1 capsid protein can have a sequence identical, at least 90% identical, or at least 95% identical to SEQ ID NO: 4.
[0011] When the modified AAV capsid protein is derived from the AAV2 capsid protein (see SEQ ID NO: 2), the targeting peptide can be inserted after residue 587 of the AAV2 capsid protein. The targeting peptide can be flanked by linker sequences, which are 2 or 3 amino acids long on each side of the targeting peptide. The linker sequence can be AAA on the N-terminal side of the targeting peptide and AA on the C-terminal side of the targeting peptide. The modified AAV2 capsid protein can have a sequence identical, at least 90% identical, or at least 95% identical to SEQ ID NO: 5.
[0012] When the modified AAV capsid protein is derived from the AAV9 capsid protein (see SEQ ID NO: 3), the targeting peptide can be inserted after residue 588 of the AAV9 capsid protein. The targeting peptide can be flanked by linker sequences, which are 2 or 3 amino acids long on each side of the targeting peptide. The linker sequence can be AAA on the N-terminal side of the targeting peptide and AS on the C-terminal side of the targeting peptide. The modified AAV9 capsid protein can have a sequence identical, at least 90% identical, or at least 95% identical to SEQ ID NO: 6.
[0013] A population of rAAV vectors may be provided that contain multiple capsid protein targeting peptides, where each capsid protein targeting peptide is paired with more than one barcode sequence.A population of rAAV vectors may be provided that contain multiple capsid protein targeting peptides, where all rAAV vectors with the same barcode sequence also contain the same capsid protein targeting peptide.In other words, multiple RNAbc sequences represent a single AAV peptide insertion sequence.This feature allows the use of "randomizer" sequences during plasmid generation, and makes the method provided herein more high-throughput because it is not necessary to individually clone each RNAbc-AAV peptide insertion combination.
[0014] Also provided herein are cells comprising the rAAV vectors of the present invention. The cells can be mammalian cells. The cells can be human cells. The cells can be in vitro or in vivo.
[0015] Also provided herein is a library preparation technique designed to simultaneously barcode and recover both mRNA and AAV-derived RNA. In one embodiment, provided herein is a method for determining the cellular tropism of a recombinant adeno-associated virus (rAAV) having a modified AAV capsid protein containing a targeting peptide, the method comprising: (i) contacting various cell types with a modified rAAV vector according to any one of the aspects of the present invention; (ii) identifying cells transduced by the modified rAAV vector based on the presence of a barcode sequence; and (iii) detecting the expressed transcriptome of each transduced cell on a cell-by-cell basis, thereby determining the cellular tropism of the modified rAAV. In one aspect, provided herein is a method for determining the cellular tropism of a recombinant adeno-associated virus (rAAV) having a modified AAV capsid protein comprising a targeting peptide, the method comprising: (i) contacting various cell types with a population of rAAV vectors provided herein; (ii) detecting, on a cell-by-cell basis, both the expressed transcriptome and the rAAV transduced into each cell; and (iii) determining which cell types have been transduced by which modified rAAV vectors, thereby determining the cellular tropism of the modified rAAV.
[0016] The contacting step in (i) can be carried out in vitro or in vivo.
[0017] The step of detecting the expressed transcriptome and rAAV in (ii) includes: (a) isolating, fixing, and permeabilizing the nuclei of the cells contacted in (i); (b) dividing the nuclei into a plurality of first aliquots; (c) reverse transcribing the cellular RNA molecules expressed in the nuclei using a primer containing a poly(T) sequence to form complementary DNA (cDNA) molecules, and (d) reverse transcribing the cellular RNA molecules expressed in the nuclei using a primer containing a sequence sufficient to hybridize to and reverse transcribe a barcode sequence in the expression cassette to form an AAV amplicon. (d) labeling the cDNA molecules and AAV amplicons with first 5' barcodes, wherein the first 5' barcode for the primer in each first aliquot is unique such that the cDNA molecules and AAV amplicons from each aliquot of nuclei can be distinguished relative to the cDNA molecules and AAV amplicons from all other aliquots of nuclei; (e) combining the plurality of first aliquots; (f) dividing the combined plurality of first aliquots into a plurality of second aliquots; (g) ligating second 5' barcodes to the 5' ends of the cDNA molecules and AAV amplicons to form dual-barcoded cDNA molecules and AAV amplicons, wherein the second 5' barcode in each second aliquot is unique. (h) combining the plurality of second aliquots; (i) dividing the combined plurality of first aliquots into a plurality of third aliquots; (j) ligating third 5' barcodes to the 5' ends of the cDNA molecules and AAV amplicons to form triple-barcoded cDNA molecules and AAV amplicons, wherein the third 5' barcode in each third aliquot is unique; (k) combining the plurality of third aliquots; (l) lysing the nuclei to release the cDNA molecules and AAV amplicons from within the nuclei to form a lysate; and (m) sequencing the cDNA molecules and AAV amplicons, thereby detecting both the expressed transcriptome and the rAAV transduced into each cell.
[0018] The cDNA molecules and AAV amplicons may be labeled with a first 5' barcode simultaneously with reverse transcription, and the reverse transcription primer may include the first 5' barcode. Nuclei may be fixed and permeabilized at temperatures below about 8°C, below about 7°C, below about 6°C, below about 5°C, below about 4°C, below about 3°C, below about 2°C, or below about 1°C. The majority of triple-barcoded cDNA molecules and AAV molecules from a single nucleus may contain the same set of barcodes. The majority of triple-barcoded cDNA molecules and AAV molecules from a single nucleus may have a unique set of barcodes compared to triple-barcoded cDNA molecules and AAV molecules from other nuclei. Cell types may be determined based on the expressed transcriptome.
[0019] Sequencing the cDNA molecules and AAV amplicons includes preparing a sequencing library, which may include (i) adding a common adapter sequence to the 3' ends of the cDNA molecules and AAV amplicons; (ii) amplifying the full-length cDNA and AAV amplicons; (iii) fragmenting the amplified full-length cDNA and AAV amplicons; (iv) repairing the ends of the fragmented cDNA and AAV amplicons and adding A-tails; (v) ligating adapters to the 5' ends of the end-repaired and A-tailed cDNA and AAV amplicons; and (vi) performing sample index PCR to add sequencing adapters and double indexes to the adapter-ligated cDNA and AAV amplicons.
[0020] Sequencing the AAV amplicons can include preparing an AAV amplicon-enriched sequencing library, which can include (i) adding a common adapter sequence to the 3' ends of the cDNA molecules and the AAV amplicons; (ii) performing full-length amplification of the cDNA and the AAV amplicons; (iii) performing AAV amplicon-enriched amplification with a forward primer that hybridizes to the AAV amplicon upstream of the AAV barcode, wherein the forward primer has a 5' phosphate; (iv) adding an A-tail to the AAV amplicons that have a 5' phosphate; (v) ligating an adapter to the 5' end of the A-tailed AAV amplicon; and (vi) performing sample index PCR to add sequencing adapters and double indexes to the adapter-ligated AAV amplicons.
[0021] A common adapter sequence can be added to the 3' ends of the cDNA molecules and AAV amplicons by template switching.
[0022] The sequencing can be paired-end sequencing, amplicon sequencing, single-cell RNA sequencing, or in situ sequencing.
[0023] Other objects, features, and advantages of the present invention will become apparent from the following detailed description. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description. [Brief explanation of the drawings]
[0024] The accompanying drawings form part of the present specification and are included to further demonstrate certain aspects of the present invention. The invention may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[0025] [Figure 1-1] Figures 1A-1F. Design of a double-barcode containing AAV cargo. (A) Schematic diagram illustrating each element in the barcoded expression construct (SEQ ID NO: 7) component of the AAV cargo. (B) Cartoon diagram illustrating the layout of the construct shown in Figure 1A within the packaged AAV genome and comparing it to the AAV Cap gene sequence, which has been modified to contain a peptide insert. (C) Gel image showing amplification of the RNAbc sequence after reverse transcription using primers pr749 and 750. (D) Schematic diagram highlighting the peptide insertion into the two different DNA barcodes, RNAbc and Cap, present. (E and F) Sanger sequencing spanning each insert confirms the successful creation of this double-barcoded construct. In Figure 1E, the five sequences from top to bottom are SEQ ID NOs: 15-19, respectively. In Figure 1F, the top sequence is SEQ ID NO: 20, and the bottom four sequences are all SEQ ID NO: 21. [Figure 1-2] Please refer to the description of Figure 1-1. [Figure 1-3] Please refer to the description of Figure 1-1. [Figure 1-4] Please refer to the description of Figure 1-1. [Figure 2A]Figures 2A-2B. Adaptation of Split-Pool Ligation-based whole-Transcriptome Sequencing (SPLiT-seq) for AAV.RNAbc detection. (A) Schematic showing the procedural steps of single-nucleus combinatorial barcoding. (B) Schematic showing the procedural steps for next-generation library preparation for whole-transcriptome and AAV.RNAbc amplicon sequencing. [Figure 2B] See legend to Figure 2A. [Figure 3] Figures 3A-3D. Single-cell RNA-Seq results showing detection of expressed RNA barcode (RNAbc) after transfection of HEK293 cells. (A) UMAP unbiased clustering of SPLiT-Seq barcoded single cells from an experiment in which HEK293 cells were transfected with plasmids containing either AAV.RNAbc or AAV.noBarcode. (B) Unique UMI counts obtained from Illumina sequencing reads after amplification and library preparation. (C) UMAP unbiased clustering of SPLiT-Seq barcoded single cells showing only cells that received AAV.RNAbc treatment. Heatmap shading indicates log UMI counts originating from the AAV.RNAbc amplicon. (D) UMAP unbiased clustering of SPLiT-Seq barcoded single cells showing only cells that received AAV.eGFP treatment. Heatmap shading indicates log UMI counts originating from the AAV.eGFP amplicon. [Figure 4-1] Figures 4A-4D. In vivo application of SPLit-Seq. (A) Schematic showing procedural steps. (B) Seurat cell annotation containing cDNA expression and AAV transduction information. (C) Identification of transduced single cells. (D) Assessment of AAV transduction status within a single cell type of interest. [Figure 4-2] See the description of Figure 4-1. [Figure 5]Figures 5A-5C. Tools for analyzing transduction performance. (A) Transduction performance by tissue. (B) Transduction performance by cell type. (C) Transduction performance spatially visualized using UMAP unbiased clustering. DETAILED DESCRIPTION OF THE INVENTION
[0026] Detailed Description Provided herein is a barcoded RNA expression construct, which is packaged in AAV and delivered to cells, expresses an mRNA sequence with a different identification barcode from the modified region of capsid DNA sequence.Also provided herein is a library preparation technique designed to simultaneously barcode and recover both mRNA and AAV-derived RNA.Finally, provided is a custom software pipeline that integrates the publicly available SPLiT-Seq demultiplexing pipeline with the pipeline for counting AAV amplicon sequences.
[0027] I. Adeno-associated virus (AAV) vectors Adeno-associated virus (AAV) is a small, non-pathogenic virus of the Parvoviridae family. To date, numerous serologically distinct AAVs have been identified, with more than 12 species isolated from humans or primates. AAV differs from other members of this family by its dependence on helper viruses for replication.
[0028] The AAV genome exists extrachromosomally without integrating into the host cell genome, has a broad host range, and can transduce both dividing and non-dividing cells in vitro and in vivo, maintaining high levels of transduced gene expression. AAV viral particles are thermostable, resistant to solvents, detergents, pH changes, and temperature, and can be column-purified and / or concentrated on a CsCl gradient or by other means. The AAV genome contains single-stranded deoxyribonucleic acid (ssDNA) of either the positive or negative strand. The approximately 4.7 kb AAV genome consists of a single segment of single-stranded DNA of either positive or negative polarity. The ends of the genome are short inverted terminal repeats (ITRs) that can fold into hairpin structures and serve as origins of viral DNA replication.
[0029] AAV "genome" refers to the recombinant nucleic acid sequence that is ultimately packaged or encapsulated to form AAV particles. AAV particles often contain an AAV genome packaged with AAV capsid proteins. When a recombinant plasmid is used to construct or produce a recombinant vector, the AAV vector genome does not contain any portion of the "plasmid" that does not correspond to the vector genome sequence of the recombinant plasmid. This non-vector genome portion of the recombinant plasmid is called the "plasmid backbone," which is important for the cloning and amplification of the plasmid, a process required for the propagation and production of the plasmid, but is not itself packaged or encapsulated into the viral particle. Therefore, AAV vector "genome" refers to the nucleic acid that is packaged or encapsulated by the AAV capsid proteins.
[0030] AAV virions (particles) are non-enveloped icosahedral particles approximately 25 nm in diameter that contain the AAV capsid. AAV particles have icosahedral symmetry, consisting of three related capsid proteins, VP1, VP2, and VP3, which interact with each other to form the capsid. Most native AAV genomes often contain two open reading frames (ORFs), sometimes referred to as the left and right ORFs. The right ORF often encodes the capsid proteins VP1, VP2, and VP3. These proteins are often found in a 1:1:10 ratio, respectively, but may occur in various ratios, all derived from the right ORF. The VP1, VP2, and VP3 capsid proteins differ from each other due to alternative splicing and aberrant start codon usage. Deletion analysis has shown that removing or altering VP1, which is translated from the alternatively spliced message, results in reduced yields of infectious particles. Mutations within the VP3 coding region result in the inability to produce any single-stranded progeny DNA or infectious particles. In certain embodiments, the genome of an AAV particle encodes one, two, or all three VP1, VP2, and VP3 polypeptides.
[0031] The left ORF often encodes nonstructural Rep proteins, Rep40, Rep52, Rep68, and Rep78, which are involved in regulating replication and transcription as well as producing single-stranded progeny genomes. Two of the Rep proteins are associated with preferential integration of the AAV genome into the q arm region of human chromosome 19. Rep68 / 78 has been shown to have NTP binding activity, as well as DNA and RNA helicase activity. Some Rep proteins have nuclear localization signals and several potential phosphorylation sites. In certain embodiments, the genome of an AAV (e.g., rAAV) encodes some or all of the Rep proteins. In certain embodiments, the genome of an AAV (e.g., rAAV) does not encode a Rep protein. In certain embodiments, one or more of the Rep proteins can be delivered in trans and are therefore not included in AAV particles containing nucleic acids encoding polypeptides.
[0032] The ends of the AAV genome contain short inverted terminal repeats (ITRs) that can potentially fold into T-shaped hairpin structures that function as origins of viral DNA replication. Thus, the AAV genome contains one or more (e.g., paired) ITR sequences flanking the single-stranded viral DNA genome. The ITR sequences are often approximately 145 bases long each. Within the ITR region, two elements thought to be central to ITR function have been described: a GAGC repeat motif and a terminal resolution site (trs). The repeat motif has been shown to bind Rep when the ITR is in either a linear or hairpin conformation. This binding is thought to position Rep68 / 78 for cleavage at the trs in a site- and strand-specific manner. In addition to their role in replication, these two elements appear to be central to viral integration. The integration locus on chromosome 19 contains a Rep binding site with adjacent trs. These elements have been shown to be functional and necessary for locus-specific integration.
[0033] The term "recombinant," as a modifier of vectors such as recombinant viral vectors, e.g., lentiviral or parvoviral (e.g., AAV) vectors, and as a modifier of sequences such as recombinant nucleic acid sequences and polypeptides, means that the composition has been manipulated (i.e., engineered) in a manner that generally does not occur in nature. A particular example of a recombinant vector, such as an AAV vector, retroviral vector, or lentiviral vector, would be where a nucleic acid sequence not normally present in the wild-type viral genome has been inserted into the viral genome. An example of a recombinant nucleic acid sequence would be where a nucleic acid (e.g., a gene) encodes an inhibitory RNA cloned into the vector with or without the 5', 3', and / or intronic regions with which the gene is normally associated in the viral genome. Although the term "recombinant" is not always used herein with reference to vectors, such as viral vectors, and sequences such as polynucleotides, "recombinant" forms, including nucleic acid sequences, polynucleotides, transgenes, and the like, are expressly included, despite any such omission.
[0034] Recombinant viral "vector" is derived from the wild-type genome of a virus by using molecular methods to remove part of the wild-type genome from the virus and replace it with non-native nucleic acid, such as a nucleic acid sequence.Typically, for example, for AAV, one or both inverted terminal repeat (ITR) sequences of the AAV genome are retained in the recombinant AAV vector.A "recombinant" viral vector (e.g., rAAV) is distinguished from a viral (e.g., AAV) genome because part of the viral genome is replaced with a non-native sequence with respect to the viral genome nucleic acid, such as the nucleic acid encoding a transactivator, or the nucleic acid encoding an inhibitory RNA, or the nucleic acid encoding a therapeutic protein.Therefore, by incorporating such a non-native nucleic acid sequence, the viral vector is defined as a "recombinant" vector, and in the case of AAV, it can be called a "rAAV vector."
[0035] In certain embodiments, an AAV (e.g., rAAV) comprises two ITRs. In certain embodiments, an AAV (e.g., rAAV) comprises a pair of ITRs. In certain embodiments, an AAV (e.g., rAAV) comprises a pair of ITRs that are adjacent to (i.e., at the 5' and 3' ends, respectively) a nucleic acid sequence encoding at least a polypeptide having a function or activity.
[0036] AAV vectors (e.g., rAAV vectors) can be packaged for subsequent infection (transduction) of cells ex vivo, in vitro, or in vivo, and are referred to herein as "AAV particles." When a recombinant AAV vector is encapsulated or packaged in an AAV particle, the particle can also be referred to as an "rAAV particle." In certain embodiments, the AAV particle is an rAAV particle. The rAAV particle often comprises an rAAV vector or a portion thereof. The rAAV particle can be one or more rAAV particles (e.g., multiple AAV particles). The rAAV particle typically comprises a protein (e.g., a capsid protein) that encapsulates or packages the rAAV vector genome. It is noted that the term "rAAV vector" can also be used to refer to an rAAV particle.
[0037] Any suitable AAV particle (e.g., rAAV particle) can be used in the methods or uses herein. The rAAV particle, and / or the genome contained therein, can be derived from any suitable serotype or strain of AAV. The rAAV particle, and / or the genome contained therein can be derived from two or more serotypes or strains of AAV. Thus, the rAAV can comprise the protein and / or nucleic acid, or a portion thereof, of any serotype or strain of AAV, where the AAV particle is suitable for infecting and / or transducing mammalian cells. Non-limiting examples of AAV serotypes include AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV-rh74, AAV-rh10, and AAV-2i8.
[0038] In certain embodiments, the plurality of rAAV particles comprises particles of or derived from the same strain or serotype (or subgroup or variant). In certain embodiments, the plurality of rAAV particles comprises a mixture of two or more different rAAV particles (e.g., of different serotypes and / or strains).
[0039] As used herein, the term "serotype" refers to an AAV having a capsid that is serologically distinct from other AAV serotypes. Serological uniqueness is determined based on the lack of cross-reactivity between antibodies against one AAV compared to antibodies against another AAV. Such differences in cross-reactivity are usually due to differences in capsid protein sequences / antigenic determinants (e.g., differences in the sequences of VP1, VP2, and / or VP3 of AAV serotypes). AAV variants, including capsid variants, may not be serologically distinct from a reference AAV or other AAV serotypes, but they differ in at least one nucleotide or amino acid residue compared to the reference or other AAV serotypes.
[0040] In certain embodiments, an rAAV vector based on a first serotype genome corresponds to one or more serotypes of the capsid proteins that package the vector. For example, the serotype of one or more AAV nucleic acids (e.g., ITRs) that comprise the AAV vector genome corresponds to the serotype of the capsid that comprises the rAAV particle.
[0041] In certain embodiments, the rAAV vector genome can be based on a genome of an AAV (e.g., AAV2) serotype that is different from one or more serotypes of the AAV capsid proteins that package the vector. For example, the rAAV vector genome can include nucleic acid (e.g., ITR) from AAV2, while at least one or more of the three capsid proteins are derived from a different serotype, such as AAV1, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, Rh10, Rh74, or AAV-2i8 serotype, or a variant thereof.
[0042] In certain embodiments, an rAAV particle or its vector genome related to a reference serotype has a polynucleotide, polypeptide, or subsequence thereof that comprises or consists of a sequence that is at least 60% or more (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc.) identical to a polynucleotide, polypeptide, or subsequence of an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, Rh10, Rh74, or AAV-2i8 particle. In certain embodiments, an rAAV particle or its vector genome related to a reference serotype has a capsid or ITR sequence that comprises or consists of at least 60% or more (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc.) identical sequence to the capsid or ITR sequence of an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, Rh10, Rh74, or AAV-2i8 serotype.
[0043] In certain embodiments, the methods herein include the use, administration, or delivery of rAAV1, rAAV2, rAAV3, rAAV4, rAAV5, rAAV6, rAAV7, rAAV8, rAAV9, rAAV10, rAAV11, rAAV12, rRh10, rRh74, or rAAV-2i8 particles.
[0044] In certain embodiments, the methods herein include the use, administration, or delivery of rAAV2 particles. In certain embodiments, the rAAV2 particles include an AAV2 capsid. In certain embodiments, the rAAV2 particles include one or more capsid proteins (e.g., VP1, VP2, and / or VP3) that are at least 60%, 65%, 70%, 75% or more identical to the corresponding capsid proteins of native or wild-type AAV2 particles, for example, 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identical. In certain embodiments, the rAAV2 particles comprise VP1, VP2, and VP3 capsid proteins that are at least 75% or more identical to the corresponding capsid proteins of native or wild-type AAV2 particles, e.g., 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identical. In certain embodiments, the rAAV2 particles are variants of native or wild-type AAV2 particles. In some aspects, one or more capsid proteins of the AAV2 variant have 1, 2, 3, 4, 5, 5-10, 10-15, 15-20, or more amino acid substitutions compared to the capsid proteins of a native or wild-type AAV2 particle.
[0045] In certain embodiments, the rAAV9 particles comprise an AAV9 capsid. In certain embodiments, the rAAV9 particles comprise one or more capsid proteins (e.g., VP1, VP2, and / or VP3) that are at least 60%, 65%, 70%, 75%, or more identical to the corresponding capsid protein of a native or wild-type AAV9 particle, e.g., 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identical. In certain embodiments, the rAAV9 particles comprise VP1, VP2, and VP3 capsid proteins that are at least 75% or more identical to the corresponding capsid proteins of native or wild-type AAV9 particles, e.g., 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identical. In certain embodiments, the rAAV9 particles are variants of native or wild-type AAV9 particles. In some aspects, one or more capsid proteins of the AAV9 variant have 1, 2, 3, 4, 5, 5-10, 10-15, 15-20, or more amino acid substitutions compared to the capsid proteins of a native or wild-type AAV9 particle.
[0046] In some embodiments, the rAAV comprises a modified capsid, wherein the modified capsid comprises a targeting peptide. In certain embodiments, the AAV is AAV1, AAV2, or AAV9. An exemplary wild-type reference AAV1 capsid protein sequence is provided in SEQ ID NO: 1. An exemplary wild-type reference AAV2 capsid protein sequence is provided in SEQ ID NO: 2. An exemplary wild-type reference AAV9 capsid protein sequence is provided in SEQ ID NO: 3. In certain aspects, the targeting peptide is inserted at position 590 of the AAV1 capsid, position 587 of the AAV2 capsid, or position 588 of the AAV9 capsid. An exemplary modified AAV1 capsid protein sequence is provided in SEQ ID NO: 4, which shows a targeting peptide insert after position 590 as SSAX7AS, where the leading SSA and following AS are linker sequences and X7 represents the targeting peptide. An exemplary modified AAV2 capsid protein sequence is provided in SEQ ID NO: 5, which shows a targeting peptide insert after position 587 as AAAX7AA, where the leading AAA and following AA are linker sequences and X7 represents the targeting peptide. An exemplary modified AAV9 capsid protein sequence is provided in SEQ ID NO: 6, which shows a targeting peptide insert after position 588 as AAAX7AS, where the leading AAA and following AS are linker sequences and X7 represents the targeting peptide.
[0047] Table 1. AAV capsid sequence TIFF2025532627000001.tif171149TIFF2025532627000002.tif207149
[0048] In certain embodiments, the rAAV particles contain one or two ITRs (e.g., a pair of ITRs) that are at least 75% or more identical to the corresponding ITRs of native or wild-type AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV-rh74, AAV-rh10, or AAV-2i8, e.g., 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identical, such that they fulfill one or more desired ITR functions (e.g., the ability to form a hairpin that allows DNA replication; AAV This includes, so long as the vector maintains integration of the DNA into the host cell genome; and / or packaging, if desired.
[0049] In certain embodiments, the rAAV2 particles comprise one or two ITRs (e.g., a pair of ITRs) that are at least 75% or more identical to the corresponding ITRs of a native or wild-type AAV2 particle, e.g., 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identical, so long as they retain one or more desired ITR functions (e.g., the ability to form a hairpin to allow DNA replication; integration of AAV DNA into the host cell genome; and / or packaging, if desired).
[0050] In certain embodiments, the rAAV9 particles comprise one or two ITRs (e.g., pairs of ITRs) that are at least 75% or more identical to the corresponding ITRs of native or wild-type AAV2 particles, e.g., 80%, 85%, 85%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, etc., up to 100% identical, so long as they retain one or more desired ITR functions (e.g., the ability to form a hairpin to enable DNA replication; integration of AAV DNA into the host cell genome; and / or packaging, if desired).
[0051] rAAV particles can comprise ITRs with any suitable number of "GAGC" repeats. In certain embodiments, the ITRs of AAV2 particles comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more "GAGC" repeats. In certain embodiments, the rAAV2 particles comprise ITRs with three "GAGC" repeats. In certain embodiments, the rAAV2 particles comprise ITRs with fewer than four "GAGC" repeats. In certain embodiments, the rAAV2 particles comprise ITRs with more than four "GAGC" repeats. In certain embodiments, the ITRs of rAAV2 particles comprise a Rep binding site in which the fourth nucleotide in the first two "GAGC" repeats is C rather than T.
[0052] An exemplary suitable length of DNA that can be incorporated into an rAAV vector for packaging / encapsidation into an rAAV particle can be about 5 kilobases (kb) or less. In certain embodiments, the length of the DNA is less than about 5 kb, less than about 4.5 kb, less than about 4 kb, less than about 3.5 kb, less than about 3 kb, or less than about 2.5 kb.
[0053] rAAV vectors containing nucleic acid sequences directing the expression of RNAi or polypeptides can be produced using suitable recombinant techniques known in the art (see, for example, Sambrook et al., 1989). Recombinant AAV vectors are typically packaged into transducible AAV particles and propagated using AAV viral packaging systems. Transducible AAV particles can bind to and enter mammalian cells, and then deliver nucleic acid cargo (e.g., heterologous genes) to the nucleus of the cells. Therefore, intact transducible rAAV particles are configured to transduce mammalian cells. rAAV particles configured to transduce mammalian cells are often not replicative and require additional protein machinery for self-replication. Therefore, rAAV particles configured to transduce mammalian cells are engineered to bind to and enter mammalian cells and deliver nucleic acids to the cells, where the nucleic acid for delivery is often located between a pair of AAV ITRs in the rAAV genome.
[0054] Suitable host cells for producing transducible AAV particles include, but are not limited to, microorganisms, yeast cells, insect cells, and mammalian cells that can be or have been used as recipients of heterologous rAAV vectors. Cells derived from the stable human cell line HEK293 (e.g., readily available through the American Type Culture Collection under accession number ATCC CRL1573) can be used. In certain embodiments, modified human embryonic kidney cell lines (e.g., HEK293) transformed with adenovirus type 5 DNA fragments and expressing adenovirus E1a and E1b genes can be used to produce recombinant AAV particles. The modified HEK293 cell line is easily transfected, providing a particularly convenient platform for producing rAAV particles. Methods for producing high-titer AAV particles capable of transducing mammalian cells are known in the art. For example, AAV particles can be produced as described in Wright, 2008 and Wright, 2009.
[0055] In certain embodiments, AAV helper functions are introduced into host cells by transfecting them with an AAV helper construct either before or simultaneously with transfection of the AAV expression vector. Thus, AAV helper constructs are sometimes used to provide at least transient expression of the AAV rep and / or cap genes to complement missing AAV functions necessary for productive AAV transduction. AAV helper constructs often lack AAV ITRs and are unable to replicate or package themselves. These constructs can be in the form of plasmids, phages, transposons, cosmids, viruses, or virions. Numerous AAV helper constructs have been described, such as the commonly used plasmids pAAV / Ad and pIM29+45, which encode both Rep and Cap expression products. Numerous other vectors encoding Rep and / or Cap expression products are known.
[0056] II. Methods for Determining Modified AAV Cell Tropism Provided herein is a method for uniquely labeling or barcoding molecules in a nucleus or multiple nuclei.It will be readily understood that the embodiments generally described herein are exemplary.The following more detailed description of various embodiments is not intended to limit the scope of the present disclosure, but merely represents various embodiments.Furthermore, the order of steps or actions of the methods disclosed herein may be changed by those skilled in the art without departing from the scope of the present disclosure.In other words, unless a specific order of steps or actions is required for the proper implementation of the embodiment, the order or use of specific steps or actions may be modified.
[0057] The term "bond" is used broadly throughout this disclosure to refer to any form of attachment or coupling of two or more components, entities, or objects. For example, two or more components may be bound to one another via chemical bonds, covalent bonds, ionic bonds, hydrogen bonds, electrostatic forces, Watson-Crick hybridization, etc.
[0058] One aspect of the present disclosure relates to a method for labeling nucleic acids. In some embodiments, the method may include labeling nucleic acids in a first nucleus. The method may include: (a) generating complementary DNA (cDNA) from cellular RNA and / or AAV.RNAbc amplicons in a plurality of nuclei by reverse transcribing the RNA using a reverse transcription primer containing a 5'-overhanging sequence; (b) dividing the plurality of nuclei into a number (n) of aliquots; (c) providing a plurality of barcode tags to each of the n aliquots, wherein each label sequence of the plurality of barcode tags provided in a given aliquot is the same, and a different label sequence is provided in each of the n aliquots; (d) binding at least one of the cDNAs and / or AAV.RNAbc amplicons in each of the n aliquots to the barcode tag; (e) combining the n aliquots; and (f) repeating steps (b), (c), (d), and (e) with the combined aliquots. In some aspects, two different reverse transcription primers are used, where the first of the reverse transcription primers contains a poly(A) hybridizing sequence (i.e., a poly(T) sequence) and the second of the reverse transcription primers contains a sequence capable of hybridizing to RNA expressed from the AAV barcoding expression construct downstream of the barcode sequence (i.e., an AAV.RNAbc transcript).
[0059] In certain embodiments, each barcode tag may include a first strand including a 3' hybridization sequence extending from the 3' end of the target sequence and a 5' hybridization sequence extending from the 5' end of the target sequence. Each barcode tag may also include a second strand including an overhanging sequence. The overhanging sequence may include (i) a first portion complementary to at least one of the 5' hybridization sequence and the 5' overhanging sequence, and (ii) a second portion complementary to the 3' hybridization sequence. In some embodiments, the barcode tag (e.g., the final nucleic acid tag) may include a capture agent, such as, but not limited to, 5' biotin. A cDNA or AAV.RNAbc amplicon labeled with a barcode tag including 5' biotin may allow or permit attachment or coupling of the cDNA or AAV.RNAbc amplicon to streptavidin-coated magnetic beads. In some other embodiments, multiple beads may be coated with a capture strand (i.e., a nucleic acid sequence) configured to hybridize to the final sequence overhang of the barcode tag. In yet some other embodiments, the cDNA or AAV.RNAbc amplicon molecules may be prepared using commercially available kits (e.g., RNEASY (商標) The protein can be purified or isolated by use of a kit.
[0060] In various embodiments, step (f) (i.e., steps (b), (c), (d), and (e)) can be repeated a sufficient number of times to generate a unique set of target sequences for the cDNA and AAV.RNAbc amplicon in a first nucleus. Stated another way, step (f) can be repeated a number of times such that the cDNA and AAV.RNAbc amplicon in a first nucleus can have a first unique set of target sequences, the cDNA and AAV.RNAbc amplicon in a second nucleus can have a second unique set of target sequences, the cDNA and AAV.RNAbc amplicon in a third nucleus can have a third unique set of target sequences, and so on. The disclosed methods can provide for labeling of cDNA and AAV.RNAbc amplicon sequences from a single nucleus with unique barcodes, which can identify or aid in identifying the cell from which the cDNA and AAV.RNAbc amplicon originated. In other words, some, most, or substantially all of the cDNA and AAV.RNAbc amplicons from a single cell can have the same barcode, and that barcode may not be repeated in cDNA or AAV.RNAbc amplicons originating from one or more other cells in the sample (e.g., a second cell, a third cell, a fourth cell, etc.).
[0061] In some embodiments, the barcoded cDNA and AAV.RNAbc amplicons can be mixed together and sequenced (e.g., using NGS) so that data can be gathered regarding RNA expression and AAV transduction at the single-cell level. For example, certain embodiments of the disclosed methods can be useful for assessing, analyzing, or studying the cellular tropism of modified AAV (i.e., the particular cell type that any given modified AAV capsid selectively or specifically targets).
[0062] As discussed above, aliquots or groups of nuclei can be separated into different reaction vessels or containers, and a first set of barcode tags can be added to multiple cDNA transcripts and AAV.RNAbc transcripts. Vessels or containers can also be referred to herein as receptacles, samples, and wells. Thus, the terms receptacle, container, receptacle, sample, and well can be used interchangeably herein. The aliquots of nuclei can then be regrouped, mixed, and separated again, and a second set of barcode tags can be added to the first set of barcode tags. In various embodiments, the same barcode tag can be added to more than one aliquot of nuclei in a single or given round of labeling. However, after repeated rounds of separation, tagging, and repooling, each nucleus' cDNA and AAV.RNAbc amplicon can be bound to a unique combination or sequence of barcode tags that identifies a single nucleus. In some embodiments, nuclei in a single sample can be separated into multiple different reaction vessels. For example, the multiple reaction vessels can include four 1.5 ml microcentrifuge tubes, multiple wells of a 96-well plate, multiple wells of a 384-well plate, or another suitable number and type of reaction vessels.
[0063] In certain embodiments, step (f) (i.e., steps (b), (c), (d), and (e)) may be repeated a number of times selected from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 times, etc. In certain other embodiments, step (f) may be repeated a sufficient number of times so that each nuclear cDNA and AAV.RNAbc amplicon is likely to bind to a unique sequence in the barcode tag. The number of times can be selected to provide a greater than 50%, greater than 90%, greater than 95%, greater than 99%, or some other probability that the cDNA and AAV.RNAbc amplicon in each nucleus will bind to a unique sequence in the barcode tag. In yet other embodiments, step (f) can be repeated some other suitable number of times.
[0064] In some embodiments, the method for labeling nucleic acids in a first nucleus can include fixing the nuclei prior to step (a). For example, components of the nuclei can be fixed or crosslinked so that the components are immobilized or held in place. The nuclei can be fixed using formaldehyde in phosphate-buffered saline (PBS). The nuclei can be fixed, for example, in about 1-4% formaldehyde in PBS. In various embodiments, the nuclei can be fixed using methanol (e.g., 100% methanol) at about -20°C or about 25°C. In various other embodiments, the nuclei can be fixed using methanol (e.g., 100% methanol) at between about -20°C and about 25°C. In still various other embodiments, the nuclei can be fixed using ethanol (e.g., about 70-100% ethanol) at about -20°C or at room temperature. In still various other embodiments, the nuclei can be fixed using ethanol (e.g., about 70-100% ethanol) at between about -20°C and room temperature. In still various other embodiments, the nuclei may be fixed using acetic acid, e.g., at about −20° C. In still various other embodiments, the nuclei may be fixed using acetone, e.g., at about −20° C. Other suitable methods of fixing the nuclei are also within the scope of this disclosure.
[0065] In certain embodiments, the method for labeling nucleic acids in a first nucleus may include permeabilizing the plurality of nuclei prior to step (a). For example, holes or openings may be formed in the nuclear membrane of the plurality of nuclei. To form one or more holes, TRITON (商標) X-100 can be added to multiple nuclei, followed by the optional addition of HCl. For example, about 0.2% TRITON (商標)X-100 may be added to the plurality of nuclei, followed by the addition of about 0.1 N HCl. In certain other embodiments, the plurality of nuclei may be permeabilized with ethanol (e.g., about 70% ethanol), methanol (e.g., about 100% methanol), Tween 20 (e.g., about 0.2% Tween 20), and / or NP-40 (e.g., about 0.1% NP-40). In various embodiments, the method of labeling nucleic acids in a first nuclei may include fixing and permeabilizing the plurality of nuclei prior to step (a).
[0066] In some embodiments, the method for labeling nucleic acids in a first nucleus can include ligating at least two of the barcode tags attached to the cDNA and / or AAV.RNAbc amplicon. Ligation can be performed before or after the lysis and / or nucleic acid purification steps. Ligation can include covalently linking the 5' phosphate sequence on the barcode tag to the 3' end of an adjacent strand or barcode tag, so that each tag forms a continuous or substantially continuous barcode sequence attached to the 3' end of the cDNA sequence. In various embodiments, a double-stranded DNA or RNA ligase can be used with an additional linker strand configured to hold the barcode tag together with the adjacent nucleic acid in a "nicked" double-stranded conformation. A double-stranded DNA or RNA ligase can then be used to seal the "nicks." In various other embodiments, a single-stranded DNA or RNA ligase can be used without an additional linker. In certain embodiments, ligation can be performed in multiple nuclei.
[0067] In certain other embodiments, the method may include, e.g., after step (f), lysing the plurality of nuclei (i.e., disassembling the nuclear structure) to release the cDNA and / or AAV.RNAbc amplicons from within the plurality of nuclei. In some embodiments, the plurality of nuclei are lysed in a lysis solution (e.g., 10 mM Tris-HCl (pH 7.9), 50 mM EDTA (pH 7.9), 0.2 M NaCl, 2.2% SDS, 0.5 mg / ml ANTI-RNase (protein ribonuclease inhibitor; AMBION)). (登録商標) ), and 1000 mg / ml proteinase K (AMBION (登録商標) )) for about 1-3 hours at about 55°C with shaking (e.g., vigorous shaking). In some other embodiments, the nuclei can be lysed using sonication and / or by passing them through an 18-25 gauge syringe needle at least once. In still other embodiments, the nuclei can be lysed by heating to about 70-90°C. For example, the nuclei can be lysed by heating to about 70-90°C for about 1 hour or more. The cDNA and / or AAV.RNAbc amplicon can then be isolated from the lysed nuclei. In some embodiments, RNase H can be added to the cDNA and AAV.RNAbc amplicon to remove RNA. The method can further include ligating at least two of the barcode tags attached to the released cDNA and AAV.RNAbc amplicon. In some other embodiments, the method of labeling nucleic acids in a first cell may include ligating at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, etc. of the barcode tags attached to the cDNA and AAV.RNAbc amplicon.
[0068] In various embodiments, the method of labeling nucleic acids in first nuclei can include removing one or more unbound barcode tags (e.g., washing the plurality of nuclei). For example, the method can include removing a portion, most, or substantially all of the unbound barcode tags. The unbound barcode tags can be removed so that further rounds of the disclosed method are not contaminated with one or more unbound barcode tags from previous rounds of a given method. In some embodiments, the unbound barcode tags can be removed via centrifugation. For example, the plurality of nuclei can be centrifuged so that a pellet of nuclei forms at the bottom of the centrifuge tube. The supernatant (i.e., the liquid containing unbound barcode tags) can be removed from the centrifuged nuclei. The nuclei can then be resuspended in a buffer solution (e.g., fresh buffer solution free of or substantially free of unbound barcode tags). In another example, the plurality of nuclei can be coupled or linked to magnetic beads coated with an antibody configured to bind to the nuclear membrane. The plurality of nuclei can then be pulled to one side of the reaction vessel using a magnet and pelleted.
[0069] As discussed above, multiple nuclei can be repooled, and the method can be repeated any number of times to add more barcode tags to the cDNA and AAV.RNAbc amplicons, creating unique sets of barcode tags that can serve to identify the cDNA and AAV.RNAbc amplicons as originating from the same cell. As more rounds are added, the number of paths a nucleus can take increases, and consequently, the number of possible unique barcode tag sequences that can be created also increases. Given enough rounds and splits, the number of possible barcodes will far exceed the number of nuclei, resulting in each nucleus having a high likelihood of having a unique barcode. For example, if splitting is performed in a 96-well plate, after four splits, 96 4 = 84,934,656 possible barcodes.
[0070] In some embodiments, the cDNA reverse transcription primer can be configured to reverse transcribe all or substantially all RNA in a cell (e.g., random hexamers with 5' overhangs). In some other embodiments, the cDNA reverse transcription primer can be configured to reverse transcribe RNA with a poly(A) tail (e.g., a poly(dT) primer, such as a dT(15) primer with a 5' overhang). In yet some other embodiments, the cDNA reverse transcription primer can be configured to reverse transcribe a predetermined RNA (e.g., a transcript-specific primer). For example, the cDNA reverse transcription primer can be configured to barcode a specific transcript so that fewer transcripts can be profiled per cell, but each transcript can be profiled across a larger number of cells.
[0071] In some embodiments, the AAV.RNAbc reverse transcription primer can be configured to reverse transcribe RNA expressed from an AAV barcoding expression construct (i.e., an AAV.RNAbc transcript). For example, the AAV.RNAbc reverse transcription primer can be configured to hybridize to an AAV.RNAbc transcript downstream of the barcode sequence.
[0072] Reverse transcription can be performed or carried out on multiple nuclei. In certain embodiments, reverse transcription can be performed on fixed and / or permeabilized multiple nuclei. In some embodiments, a variant of M-MuLV reverse transcriptase can be used in reverse transcription. Any suitable method of reverse transcription is within the scope of this disclosure. For example, the reverse transcription mix can include a reverse transcription primer that includes a 5' overhang, and the reverse transcription primer can be configured to initiate reverse transcription and / or act as a binding sequence for a barcode tag. In some other embodiments, the portion of the reverse transcription primer that is configured to bind to RNA and / or initiate reverse transcription can include one or more of a random hexamer, a septamer, an octamer, a nonamer, a decamer, a poly(T) stretch of nucleotides, and / or one or more gene-specific primers.
[0073] Another aspect of the present disclosure relates to a method for uniquely labeling RNA molecules in multiple nuclei. The method may include: (a) fixing and permeabilizing the first plurality of nuclei prior to step (b), wherein the first plurality of nuclei may be fixed and permeabilized at less than about 8°C; (b) reverse transcribing RNA molecules in the first plurality of cells to form complementary DNA (cDNA) molecules and AAV.RNAbc amplicons in the first plurality of nuclei, wherein reverse transcribing the RNA molecules comprises coupling a primer to the RNA molecules, wherein the primer comprises at least one of a poly(T) sequence or a sequence capable of hybridizing to RNA expressed from an AAV barcoding expression construct downstream of a barcode sequence (i.e., an AAV.RNAbc transcript); (c) dividing the first plurality of nuclei comprising the cDNA molecules and AAV.RNAbc amplicons into at least two primary aliquots, wherein the at least two primary aliquots comprise a first primary aliquot and a second primary aliquot; (e) coupling the cDNA molecules and AAV.RNAbc amplicons in each of the at least two primary aliquots with the provided primary barcode tags; (f) combining the at least two primary aliquots; (g) dividing the combined primary aliquots into at least two secondary aliquots, the at least two secondary aliquots comprising a first secondary aliquot and a second secondary aliquot; (h) providing secondary barcode tags to the at least two secondary aliquots, the secondary barcode tag provided in the first secondary aliquot being different from the secondary barcode tag provided in the second secondary aliquot; (i) coupling the cDNA molecules and AAV.RNAbc amplicons in each of the at least two secondary aliquots with the provided primary barcode tags;(j) repeating steps (f), (g), (h), and (i) with subsequent aliquots, wherein the final barcode tag comprises a capture agent; (k) combining the final aliquots; (l) lysing the first plurality of nuclei to release the cDNA molecules and AAV.RNAbc amplicons from within the first plurality of nuclei to form a lysate; and / or (m) adding a protease inhibitor and / or a binding agent to the lysate, such that the cDNA molecules and AAV.RNAbc amplicons bind to the binding agent.
[0074] The method may further comprise dividing the combined final aliquot into at least two final aliquots, wherein the at least two final aliquots comprise a first final aliquot and a second final aliquot. In some embodiments, the first plurality of nuclei may be fixed and permeabilized at a temperature below about 8°C, below about 7°C, below about 6°C, below about 5°C, below about 4°C, below about 4°C, below about 3°C, below about 2°C, below about 1°C, or another suitable temperature. In certain embodiments, the method may comprise dividing the nuclei. For example, after the last or final round of barcoding (via ligation), the nuclei may be pooled before lysis, and then the nuclei may be divided into different lysate aliquots. Each lysate aliquot may contain a predetermined number of nuclei.
[0075] For example, with respect to step (m), the protease inhibitor may include phenylmethanesulfonyl fluoride (PMSF), 4-(2-aminoethyl)benzenesulfonyl fluoride hydrochloride (AEBSF), a combination thereof, and / or another suitable protease inhibitor. For example, with respect to steps (j), (k), (l), and / or (m), the capture agent may include biotin or another suitable capture agent. Furthermore, the binding agent may include avidin (e.g., streptavidin) or another suitable binding agent.
[0076] In certain embodiments, the method for uniquely labeling RNA molecules in multiple nuclei can further include (e.g., after step (m)): (n) performing template switching of the cDNA molecules and AAV.RNAbc amplicons bound to the binder using a template switch oligonucleotide; (o) amplifying the cDNA molecules and AAV.RNAbc amplicons to form an amplified cDNA molecule and AAV.RNAbc amplicon solution; and / or (p) introducing a solid-phase reversible immobilization (SPRI) bead solution into the amplified cDNA molecule and AAV.RNAbc amplicon solution to remove polynucleotides less than about 200 base pairs, less than about 175 base pairs, or less than about 150 base pairs (see DeAngelis, MM, et al. Nucleic Acids Research (1995) 23(22):4742). In other words, the cDNA molecules and AAV.RNAbc amplicons can bind to streptavidin beads within the lysate. Template switching of the bead-attached cDNA molecules and AAV.RNAbc amplicons can be performed, for example, to add adapters to the 3' ends of the cDNA molecules and AAV.RNAbc amplicons. PCR amplification of the cDNA molecules and AAV.RNAbc amplicons can then be performed, followed by the addition of SPRI beads to remove polynucleotides less than about 200 base pairs. The ratio of SPRI bead solution to amplified cDNA molecule solution can be between about 0.9:1 and about 0.7:1, between about 0.875:1 and about 0.775:1, between about 0.85:1 and about 0.75:1, between about 0.825:1 and about 0.725:1, about 0.8:1, or another suitable ratio. Additionally, the SPRI bead solution may contain between about 1 M and 4 M NaCl, between about 2 M and 3 M NaCl, between about 2.25 M and 2.75 M NaCl, about 2.5 M NaCl, or another suitable amount of NaCl. The SPRI bead solution may also contain between about 15% and 25% w / v polyethylene glycol (PEG), where the molecular weight of PEG is between about 7,000 g / mol and 9,000 g / mol (PEG 8000).In various embodiments, the SPRI bead solution may contain between about 17% w / v and 23% w / v PEG 8000, between about 18% w / v and 22% w / v PEG 8000, between about 19% w / v and 21% w / v PEG 8000, about 20% w / v PEG 8000, or another suitable % w / v PEG 8000.
[0077] The method for uniquely labeling RNA molecules in multiple nuclei may further include adding a common adapter sequence to the 3' end of the released cDNA molecules and AAV.RNAbc amplicons. The common adapter sequence can be the same or substantially the same for each cDNA molecule and AAV.RNAbc amplicon (i.e., within a given experiment). Addition of the common adapter may be performed or carried out in a solution containing up to about 10% w / v PEG, the molecular weight of which is between about 7,000 g / mol and 9,000 g / mol. In certain embodiments, the common adapter sequence can be added to the 3' end of the cDNA molecules and AAV.RNAbc amplicons released by template switching (see Picelli, S, et al. Nature Methods 10, 1096-1098 (2013)).
[0078] Step (j) can be repeated a sufficient number of times to generate a series of unique barcode tags for the nucleic acids in a single nucleus. For example, the number of times can be selected from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, and 100 times.
[0079] In various embodiments, the primer in step (b) can further comprise a first specific barcode.In other words, the first barcode that is added to the cDNA molecules and AAV.RNAbc amplicons in a specific container, mixture, reaction product, container, sample, well, or vessel can be predetermined (e.g., specific to a given container, mixture, reaction product, container, sample, well, or vessel).For example, 96 sets of different well-specific RT primers can be used (e.g., in a 96-well plate).Therefore, if there are 96 samples or aliquots, each sample or aliquot can obtain a unique well-specific barcode.
[0080] In various embodiments, each of the barcode tags may comprise a first strand, the first strand comprising (i) a barcode sequence comprising a 3' end and a 5' end, and (ii) a 3' hybridization sequence and a 5' hybridization sequence adjacent to the 3' end and the 5' end of the barcode sequence, respectively. Each of the barcode tags may also comprise a second strand, the second strand comprising (i) a first portion complementary to at least one of the 5' hybridization sequence and the adapter sequence, and (ii) a second portion complementary to the 3' hybridization sequence.
[0081] The method for uniquely labeling RNA molecules in a plurality of nuclei may further include ligating at least two (or more) of the barcode tags attached to the cDNA molecule and the AAV.RNAbc amplicon. The ligation may be performed within a first plurality of nuclei.
[0082] The method may further comprise removing unbound barcode tags. In some embodiments, the method may comprise ligating at least two of the barcode tags bound to the released cDNA molecules and AAV.RNAbc amplicons. Most of the cDNA molecules and AAV.RNAbc amplicons bound to nucleic acid tags from a single nucleus may contain the same set of bound barcode tags.
[0083] The cDNA molecules can be formed or generated in aliquots (e.g., reaction mixtures). The concentration of the first reverse transcription primer in the aliquots can be between about 0.5 μM and about 10 μM, between about 1 μM and about 7 μM, between about 1.5 μM and about 4 μM, between about 2 μM and about 3 μM, about 2.5 μM, or another suitable concentration. The concentration of the second reverse transcription primer in the aliquots can be between about 0.5 μM and about 10 μM, between about 1 μM and about 7 μM, between about 1.5 μM and about 4 μM, between about 2 μM and about 3 μM, about 2.5 μM, or another suitable concentration.
[0084] III. Sequencing library preparation In various embodiments, sequencing can be performed on various sequencing platforms, requiring the preparation of a sequencing library. For whole-transcriptome sequencing, preparation typically involves fragmenting the cDNA (sonication, spraying, or shearing), followed by cDNA repair and end-polishing (blunt-end or A-overhang), and platform-specific adapter ligation. In one embodiment, the methods described herein can utilize next-generation sequencing (NGS) technologies, which allow multiple samples to be sequenced individually as genomic molecules (i.e., singleplex sequencing) or as a pooled sample containing indexed genomic molecules (e.g., multiplex sequencing) in a single sequencing run. These methods can generate up to billions of DNA sequence reads. In various embodiments, the sequences of genomic nucleic acids and / or indexed genomic nucleic acids can be determined, for example, using next-generation sequencing (NGS) technologies described herein. In various embodiments, the analysis of the enormous amount of sequence data obtained using NGS can be performed using one or more processors.
[0085] Preparation of sequencing libraries for whole transcriptome sequencing is facilitated by fragmentation of large polynucleotides (e.g., cDNA) to obtain polynucleotides in the desired size range.
[0086] Paired-end reads can be used in the sequencing methods and systems disclosed herein, where the length of the fragment or insert is longer than the length of the read, and sometimes longer than the sum of the lengths of the two reads.
[0087] In some illustrative embodiments, the sample nucleic acid is obtained as cDNA, which is subjected to fragmentation into fragments longer than approximately 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 base pairs, which can be easily amenable to NGS methods. In some embodiments, paired-end reads are obtained from inserts of approximately 100 to 5000 bp. In some embodiments, the inserts are approximately 100 to 1000 bp in length. These are sometimes implemented as regular short-insert paired-end reads. In some embodiments, the inserts are approximately 1000 to 5000 bp in length.
[0088] Fragmentation can be achieved by any of a number of methods known to those skilled in the art.For example, fragmentation can be achieved by mechanical means, including but not limited to spraying, ultrasonication, and hydroshearing, or by enzymatic means.However, mechanical fragmentation typically cuts DNA backbone at CO, PO, and CC bond, resulting in a heterogeneous mixture of blunt ends and 3'-protruding ends and 5'-protruding ends with broken CO, PO, and / or CC bond (see, for example, Alnemri and Liwack, J Biol. Chem 265:17323-17333
[1990] ; Richards and Boyer, J Mol Biol 11:327-240
[1965] ), which may need to be repaired because they may lack the 5'-phosphate required for subsequent enzymatic reactions, such as the ligation of sequencing adaptors, that are required to prepare DNA for sequencing.
[0089] Typically, DNA fragments are converted into blunt-end DNA with 5'-phosphate and 3'-hydroxyl.Standard protocols, such as the sequencing protocol using the Illumina platform as described in the exemplary workflow in Figure 2B, instruct users to end-repair sample DNA, purify the end-repaired product before 3'-end adenylation or dA tailing, and purify the dA tailing product before the adapter ligation step of library preparation.
[0090] Various embodiments of the method for preparing sequence libraries described herein avoid the need to perform one or more of the steps typically required by standard protocols to obtain modified DNA products that can be sequenced by NGS.For example, for sequencing enriched AAV.RNAbc amplicons, no fragmentation is performed.Rather, PCR enrichment is performed using a forward primer that is specific to AAV.RNAbc amplicons, and this forward primer has a 5' phosphate, thereby eliminating the need for end repair.
[0091] IV. Sequencing Methods Methods and devices described herein can adopt next-generation sequencing technology (NGS), which enables massively parallel sequencing.In certain embodiments, clone-amplified DNA template or single DNA molecule is sequenced in a flow cell in a massively parallel manner (for example, as described in Volkerding et al. Clin Chem 55:641-658
[2009] ; Metzker M Nature Rev 11:31-46
[2010] ).The sequencing technology of NGS includes but is not limited to pyrosequencing, reversible dye terminator sequencing by synthesis, oligonucleotide probe ligation sequencing and ion semiconductor sequencing. In a single sequencing run, DNA from each sample can be sequenced individually (i.e., singleplex sequencing) to generate up to hundreds of millions of DNA sequence reads, or DNA from multiple samples can be pooled and sequenced as an indexed genome molecule (i.e., multiplex sequencing). Examples of sequencing technologies that can be used to obtain sequence information according to the present methods are further described herein.
[0092] As described below, several sequencing technologies are commercially available, such as the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, Calif.), sequencing-by-synthesis platforms from 454 Life Sciences (Bradford, Conn.), Illumina / Solexa (Hayward, Calif.), and Helicos Biosciences (Cambridge, Mass.), and the sequencing-by-ligation platform from Applied Biosystems (Foster City, Calif.). In addition to single molecule sequencing performed using Helicos Biosciences' sequencing-by-synthesis platform, other single molecule sequencing technologies include Pacific Biosciences' SMRT (商標) Technology, ION TORRENT (商標) techniques, and nanopore sequencing, for example, as developed by Oxford Nanopore Technologies.
[0093] Although automated Sanger sequencing is considered to be " first generation " technology, Sanger sequencing, including automated Sanger sequencing, can also be employed in the method described herein.Other suitable sequencing methods include but are not limited to nucleic acid imaging technology, such as atomic force microscope (AFM) or transmission electron microscope (TEM).Illustrative sequencing technology will be described in more detail below.
[0094] In some embodiments, the disclosed methods involve obtaining nucleic acid sequence information in a test sample by massively parallel sequencing of millions of DNA fragments using Illumina sequencing-by-synthesis and reversible terminator-based sequencing chemistry (e.g., as described in Bentley et al., Nature 6:53-59
[2009] ). The template DNA can be cDNA. In some embodiments, cDNA from isolated cells is used as a template and fragmented to lengths of several hundred base pairs. In other embodiments, AAV.RNAbc amplicons are prepared by PCR amplification, and fragmentation is not required. If necessary, the template DNA is end-repaired to generate 5'-phosphorylated blunt ends. The polymerase activity of Klenow fragment is used to add a single A base to the 3' end of the blunt-phosphorylated DNA fragment. This addition prepares the DNA fragments for ligation to oligonucleotide adapters with a single T base overhang at the 3' end to increase ligation efficiency. The adapter oligonucleotide is complementary to the flow cell anchor oligo. Under limiting dilution conditions, adapter-modified single-stranded template DNA is added to a flow cell and immobilized by hybridization to the anchor oligo. The attached DNA fragments are extended and bridge-amplified to create an ultra-high-density sequencing flow cell with hundreds of millions of clusters, each containing approximately 1,000 copies of the same template. In one embodiment, the adapter-ligated DNA is amplified using PCR before being subjected to cluster amplification. In some applications, the templates are sequenced using robust four-color DNA sequencing by synthesis, which employs reversible terminators with removable fluorescent dyes. High-sensitivity fluorescence detection is achieved using laser excitation and total internal reflection optics. Specially developed data analysis pipeline software aligns short sequence reads, approximately 10 to several hundred base pairs, to a reference genome to identify unique mappings of the short sequence reads to the reference genome. After completion of the first read, the template can be regenerated in situ to allow a second read from the opposite end of the fragment.Thus, either single-end or paired-end sequencing of DNA fragments can be used.
[0095] Various embodiments of the present disclosure may use sequencing by synthesis, which allows paired-end sequencing. In some embodiments, Illumina's sequencing by synthesis platform involves fragment clustering. Clustering is a process in which each fragment molecule is isothermally amplified. In some embodiments, as in the example described herein, fragments have two different adapters attached to their two ends, and the adapters allow the fragments to hybridize with two different oligos on the surface of a flow cell lane. The fragments further include or are connected to two index sequences at their two ends, which provide labels for identifying different samples in multiplex sequencing. In some sequencing platforms, the fragments to be sequenced from both ends are also called inserts.
[0096] In some implementations, the flow cell for clustering in the Illumina platform is a glass slide with lanes. Each lane is a glass channel coated with a lawn of two types of oligos (e.g., P5 oligos and P7' oligos). Hybridization is enabled by the first of the two types of oligos on the surface. This oligo is complementary to the first adapter at one end of the fragment. A polymerase creates a complementary strand of the hybridized fragment. The double-stranded molecule is denatured, and the original template strand is washed away. The remaining strand is clonally amplified by bridge application in parallel with many other remaining strands.
[0097] In other sequencing methods involving bridge amplification and clustering, the strands are folded over, and a second adapter region at the second end of the strand hybridizes to a second type of oligo on the surface of the flow cell. A polymerase generates a complementary strand, forming a double-stranded bridge molecule. This double-stranded molecule is denatured, resulting in two single-stranded molecules tethered to the flow cell through two different oligos. The process is then repeated multiple times, resulting in millions of clusters simultaneously, resulting in clonal amplification of all fragments. After bridge amplification, the reverse strand is cleaved and washed away, leaving only the forward strand. The 3' end is blocked to prevent unwanted priming.
[0098] After clustering, sequencing begins by extending the first sequencing primer to generate the first read. At each cycle, fluorescently tagged nucleotides compete for addition to the growing strand. Only one is incorporated based on the sequence of the template. After each nucleotide is added, the cluster is excited by a light source, causing it to emit a characteristic fluorescent signal. The number of cycles determines the length of the read. The emission wavelength and signal intensity determine the base call. For a given cluster, all identical strands are read simultaneously. Hundreds of millions of clusters are sequenced in a massively parallel manner. Upon completion of the first read, the read product is washed away.
[0099] In the next step of the protocol, which involves two index primers, an index 1 primer is introduced and hybridized to the index 1 region on the template. The index region provides fragment identification, which is useful for demultiplexing samples in multiplex sequencing processes. The index 1 read is generated in the same manner as the first read. After the index 1 read is completed, the read product is washed away and the 3' end of the strand is deprotected. The template strand then folds over and binds to the second oligo on the flow cell. The index 2 sequence is read in the same manner as index 1. The index 2 read product is then washed away upon completion of the process.
[0100] After reading the two indices, Read 2 begins by using a polymerase to extend the second flow cell oligo to form a double-stranded bridge. This double-stranded DNA is denatured and the 3' end is blocked. The original forward strand is cleaved and washed away, leaving the reverse strand. Read 2 begins with the introduction of a Read 2 sequencing primer. As with Read 1, the sequencing process is repeated until the desired length is achieved. The Read 2 product is washed away. This entire process generates millions of reads, representing all fragments. Sequences from the pooled sample library are separated based on unique indices introduced during sample preparation. For each sample, reads with similar stretches of base calls are locally clustered. Forward and reverse reads are paired to create a contiguous sequence.
[0101] V. Kit Another aspect of the present disclosure relates to a kit for labeling nucleic acids in at least a first cell. In some embodiments, the kit may include at least two reverse transcription primers containing 5'-overhanging sequences. The kit may include at least one reverse transcription primer containing poly(T). The kit may include at least one AAV.RNAbc transcript-specific reverse transcription primer.
[0102] The kit may also include a plurality of first barcode tags. Each first barcode tag may include a first strand. The first strand may include a 3' hybridization sequence extending from the 3' end of the first target sequence and a 5' hybridization sequence extending from the 5' end of the first target sequence. Each first barcode tag may further include a second strand. The second strand may include an overhanging sequence, and the overhanging sequence may include (i) a first portion complementary to at least one of the 5' hybridization sequence and the 5' overhanging sequence of the reverse transcription primer, and (ii) a second portion complementary to the 3' hybridization sequence.
[0103] The kit may further include a plurality of second barcode tags. Each second barcode tag may include a first strand. The first strand may include a 3' hybridization sequence extending from the 3' end of the second label sequence and a 5' hybridization sequence extending from the 5' end of the second label sequence. Each second barcode tag may further include a second strand. The second strand may include an overhanging sequence, and the overhanging sequence may include (i) a first portion complementary to at least one of the 5' hybridization sequence and the 5' overhanging sequence of the reverse transcription primer, and (ii) a second portion complementary to the 3' hybridization sequence. In some embodiments, the first label sequence may be different from the second label sequence.
[0104] In some embodiments, the kit may also include one or more additional barcode tags. Each barcode tag of the one or more additional barcode tags may include a first strand. The first strand may include a 3' hybridization sequence extending from the 3' end of the target sequence and a 5' hybridization sequence extending from the 5' end of the target sequence. Each barcode tag of the one or more additional barcode tags may also include a second strand. The second strand may include an overhanging sequence, the overhanging sequence including (i) a first portion complementary to at least one of the 5' hybridization sequence and the 5' overhanging sequence of the reverse transcription primer, and (ii) a second portion complementary to the 3' hybridization sequence. In some embodiments, the label sequence may be different in each of the given additional barcode tags.
[0105] In various embodiments, the kit may further comprise at least one of a reverse transcriptase, a fixation agent, a permeabilization agent, a ligation agent, and / or a lysis agent.
[0106] VI. Definition The terms "polynucleotide," "nucleic acid," and "transgene" are used interchangeably herein to refer to all forms of nucleic acids, oligonucleotides, including deoxyribonucleic acid (DNA) and ribonucleic acid (RNA) and polymers thereof. Polynucleotides include genomic DNA, cDNA, and antisense DNA, as well as spliced or unspliced mRNA, rRNA, tRNA, and inhibitory DNA or RNA (RNAi, e.g., small or short hairpin (sh)RNA, microRNA (miRNA), small or short interfering (si)RNA, trans-splicing RNA, or antisense RNA). Polynucleotides can include naturally occurring, synthetic, and intentionally modified or altered polynucleotides (e.g., variant nucleic acids). Polynucleotides can be single-stranded, double-stranded, or triplexed, linear or circular, and of any suitable length. When discussing polynucleotides, the sequence or structure of a particular polynucleotide may be described herein according to the convention of providing the sequence in the 5' to 3' direction.
[0107] A nucleic acid that encodes a polypeptide often includes an open reading frame that encodes the polypeptide. Unless otherwise indicated, a particular nucleic acid sequence also includes degenerate codon substitutions.
[0108] The nucleic acid can include one or more expression control or regulatory elements operably linked to the open reading frame, wherein the one or more regulatory elements are configured to direct the transcription and translation of the polypeptide encoded by the open reading frame in mammalian cells. Non-limiting examples of expression control / regulatory elements include transcription initiation sequences (e.g., promoters, enhancers, TATA boxes, etc.), translation initiation sequences, mRNA stability sequences, polyA sequences, secretory sequences, etc. Expression control / regulatory elements can be obtained from the genome of any suitable organism.
[0109] "Promoter" refers to a nucleotide sequence, usually upstream (5') of a coding sequence, that directs and / or controls expression of the coding sequence by providing recognition for RNA polymerase and other factors required for proper transcription. Pol II promoters contain a minimal promoter, a short DNA sequence consisting of a TATA box and, optionally, other sequences that serve to specify the site of transcription initiation, to which regulatory elements are added for control of expression. Type 1 Pol III promoters contain three cis-acting sequence elements downstream of the transcription start site: a) a 5' sequence element (A block); b) an intermediate sequence element (I block); and c) a 3' sequence element (C block). Type 2 Pol III promoters contain two essential cis-acting sequence elements downstream of the transcription start site: a) an A box (5' sequence element); and b) a B box (3' sequence element). Type 3 Pol III promoters contain several cis-acting promoter elements upstream of the transcription start site, such as the traditional TATA box, a proximal sequence element (PSE), and a distal sequence element (DSE).
[0110] An "enhancer" is a DNA sequence that can stimulate transcriptional activity and can be a native element of a promoter or a heterologous element that enhances the level or tissue specificity of expression. It can operate in either orientation (5'→3' or 3'→5') and can function when placed either upstream or downstream of a promoter.
[0111] A promoter and / or enhancer may be derived entirely from a native gene, or may be composed of various elements derived from various elements found in nature, or may even be composed of synthetic DNA segments. A promoter or enhancer may contain DNA sequences that are involved in the binding of protein factors that regulate / control the effectiveness of transcription initiation in response to stimuli, physiological or developmental conditions.
[0112] Non-limiting examples of promoters include the SV40 early promoter, mouse mammary tumor virus LTR promoter, adenovirus major late promoter (Ad MLP), herpes simplex virus (HSV) promoter, cytomegalovirus (CMV) promoters such as the CMV immediate early promoter region (CMVIE), Rous sarcoma virus (RSV) promoter, pol II promoter, pol III promoter, synthetic promoters, hybrid promoters, and the like. In addition, sequences derived from non-viral genes, such as the mouse metallothionein gene, also find use herein. Exemplary constitutive promoters include promoters of the following genes encoding certain constitutive or "housekeeping" functions: hypoxanthine phosphoribosyltransferase (HPRT), dihydrofolate reductase (DHFR), adenosine deaminase, phosphoglycerol kinase (PGK), pyruvate kinase, phosphoglycerol mutase, actin promoter, U6, and other constitutive promoters known to those skilled in the art. In addition, many viral promoters function constitutively in eukaryotic cells. These include, among others, the early and late promoters of SV40, the long terminal repeat (LTR) of Moloney leukemia virus and other retroviruses, and the thymidine kinase promoter of herpes simplex virus.In addition, the sequences derived from intronic miRNA promoters, such as miR107, miR206, miR208b, miR548f-2, miR569, miR590, miR566 and miR128 promoters, are also useful herein (see, for example, Monteys et al., 2010).Therefore, any of the constitutive promoters referred to above can be used to control the transcription of heterologous gene inserts.
[0113] "Transgene" is used herein to conveniently refer to a nucleic acid sequence / polynucleotide that is intended to be or has been introduced into a cell or organism. Transgenes include any nucleic acid, such as a gene encoding a barcode, and are generally heterologous with respect to the naturally occurring AAV genome sequence.
[0114] The term "transduce" refers to the introduction of a nucleic acid sequence into a cell or host organism by a vector (e.g., a viral particle). Thus, the introduction of a transgene into a cell by a viral particle can be referred to as "transduction" of the cell. The transgene may or may not be integrated into the genomic nucleic acid of the transduced cell. If the introduced transgene is integrated into the nucleic acid (genomic DNA) of a recipient cell or organism, it can be stably maintained in the cell or organism and further transmitted or inherited by the cells or organisms of the recipient cell's or organism's descendants. Finally, the introduced transgene can exist extrachromosomally or only transiently in the recipient cell or host organism. Thus, a "transduced cell" is a cell into which a transgene has been introduced by transduction. Thus, a "transduced" cell is a cell into which a transgene has been introduced or its progeny. The transduced cell can proliferate, transcribe the transgene, and express the encoded inhibitory RNA or protein. For gene therapy applications and methods, the transduced cells can be in a mammal.
[0115] A nucleic acid / transgene is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. A nucleic acid / transgene encoding a barcode or directing expression of a polypeptide may include an inducible promoter or a tissue-specific promoter to control transcription of the encoded polypeptide. A nucleic acid operably linked to an expression control element may also be referred to as an expression cassette.
[0116] As used herein, the term "modify" or "variant" and its grammatical variations refer to nucleic acid, polypeptide, or its subsequence deviating from the reference sequence.Thus, modified and variant sequences can have substantially the same, greater, or less expression, activity, or function as the reference sequence, but retain at least partial activity or function of the reference sequence.A particular type of variant is mutant protein, which refers to the protein encoded by a gene that has mutation, for example, missense mutation or nonsense mutation.
[0117] A "nucleic acid" variant or a "polynucleotide" variant refers to a modified sequence that is genetically altered compared to the wild type. The sequence may be genetically altered without changing the encoded protein sequence. Alternatively, the sequence may be genetically altered to encode a variant protein. A nucleic acid variant or a polynucleotide variant may also refer to a combination sequence that has been codon-altered to encode a protein that still retains at least partial sequence identity with a reference sequence, such as a wild-type protein sequence, and that has also been codon-altered to encode a variant protein. For example, some codons in such nucleic acid variants are altered without changing the amino acids of the protein encoded thereby, and some codons in nucleic acid variants are altered to subsequently change the amino acids of the protein encoded thereby.
[0118] The terms "protein" and "polypeptide" are used interchangeably herein. A "polypeptide" encoded by a "nucleic acid" or "polynucleotide" or "transgene" disclosed herein includes partial or full-length native sequences, as well as naturally occurring wild-type and functional polymorphic proteins, functional subsequences (fragments) thereof, and sequence variants thereof, so long as the polypeptide retains some function or activity. Thus, in the methods and uses of the present invention, such polypeptides encoded by nucleic acid sequences are not required to be identical to endogenous proteins that are defective or whose activity, function, or expression is insufficient, missing, or absent in the treated mammal.
[0119] Non-limiting examples of modifications include substitution of one or more nucleotides or amino acids (e.g., about 1 to about 3, about 3 to about 5, about 5 to about 10, about 10 to about 15, about 15 to about 20, about 20 to about 25, about 25 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 100, about 100 to about 150, about 150 to about 200, about 200 to about 250, about 250 to about 500, about 500 to about 750, about 750 to about 1000, or more nucleotides or residues).
[0120] Examples of amino acid modifications are conservative amino acid substitutions or deletions. In certain embodiments, the modified or variant sequence retains at least some of the function or activity of the unmodified sequence (e.g., the wild-type sequence).
[0121] Another example of an amino acid modification is a targeting peptide introduced into the capsid protein of the viral particle. Peptides have been identified that target recombinant viral vectors to the central nervous system, for example, to different brain regions.
[0122] A "variant" of a molecule is a sequence that is substantially similar to the sequence of the native molecule. For nucleotide sequences, variants include sequences that, due to the degeneracy of the genetic code, encode the same amino acid sequence as the native protein. Naturally occurring allelic variants such as these can be identified using molecular biology techniques, such as polymerase chain reaction (PCR) and hybridization techniques. Variant nucleotide sequences also include synthetically derived nucleotide sequences, such as those that encode native proteins, for example, those generated by site-directed mutagenesis, and those that encode polypeptides with amino acid substitutions. Generally, the nucleotide sequence variants of the present invention will have at least 40%, 50%, 60%, to 70%, e.g., 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, to 79%, typically at least 80%, e.g., 81% to 84%, at least 85%, e.g., 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, to 98% sequence identity to the native (endogenous) nucleotide sequence. In certain embodiments, the variant is biologically functional (i.e., retains 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% of the wild-type activity or function).
[0123] "Conservative variations" of a particular nucleic acid sequence refer to nucleic acid sequences that encode identical or essentially identical amino acid sequences. Due to the degeneracy of the genetic code, a large number of functionally identical nucleic acids encode any given polypeptide. For example, the codons CGT, CGC, CGA, CGG, AGA, and AGG all encode the amino acid arginine. Thus, at every position where arginine is specified by a codon, the codon can be altered to any of the corresponding codons described without altering the encoded protein. Such nucleic acid variations are "silent variations," which are one species of "conservatively modified variations." Every nucleic acid sequence described herein that encodes a polypeptide also describes all possible silent variations, unless otherwise specified. Those skilled in the art will recognize that each codon in a nucleic acid (except ATG, which is usually the only codon for methionine) can be altered using standard techniques to produce a functionally identical molecule. Thus, each "silent variation" of a nucleic acid that encodes a polypeptide is implicit in each described sequence.
[0124] The term "substantial identity" of a polynucleotide sequence means that the polynucleotide comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, or 79%, or at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, or 89%, or at least 90%, 91%, 92%, 93%, or 94%, or even at least 95%, 96%, 97%, 98%, or 99% sequence identity compared to a reference sequence using one of the alignment programs described with standard parameters. Those skilled in the art will recognize that these values can be appropriately adjusted to determine the corresponding identity of proteins encoded by two nucleotide sequences by taking into account codon degeneracy, amino acid similarity, reading frame arrangement, etc. Substantial amino acid sequence identity for these purposes typically means at least 70%, at least 80%, 90%, or even at least 95% sequence identity.
[0125] The term "substantial identity" in the context of a polypeptide refers to a polypeptide comprising a sequence that has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, or 79%, or 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, or 89%, or at least 90%, 91%, 92%, 93%, or 94%, or even 95%, 96%, 97%, 98%, or 99% sequence identity with a reference sequence over a specified comparison window. An indication that two polypeptide sequences are identical is that one polypeptide is immunologically reactive with an antibody raised against the second polypeptide. Thus, a polypeptide is identical to a second polypeptide, for example, if the two peptides differ only by conservative substitutions.
[0126] As used herein, "essentially free" of a specified component means that none of the specified components are intentionally formulated in the composition and / or are present only as contaminants or in trace amounts.Therefore, the total amount of the specified components resulting from any unintentional contamination of the composition is well below 0.05%, preferably below 0.01%.Most preferred is a composition in which no amount of the specified components can be detected by standard analytical methods.
[0127] As used herein in the specification, "a" or "an" may mean one or more. As used herein in the claims, when used in conjunction with the word "comprising," the words "a" or "an" may mean one or more than one.
[0128] Use of the term "or" in the claims is used to mean "and / or," unless expressly indicated to refer to alternatives only or the alternatives are not mutually exclusive, although the present disclosure supports a definition that refers to alternatives only and "and / or." As used herein, "another" can mean at least a second or more.
[0129] Throughout this application, the term "about" is used to indicate that a value includes the inherent variation of error for the device, the inherent variation in the method employed to determine the value, the variation that exists among study subjects, or a value that is within 10% of the stated value. [Example]
[0130] VII. Working Examples The following examples are included to demonstrate preferred embodiments of the invention. It should be recognized by those of skill in the art that the techniques disclosed in the examples that follow represent techniques discovered by the inventors to function well in the practice of the invention, and therefore can be considered to constitute preferred modes for its practice. However, those of skill in the art should, in light of the present disclosure, recognize that many changes can be made in the specific embodiments that are disclosed and still obtain like or similar results without departing from the spirit and scope of the invention.
[0131] Example 1 – Design of a double barcode containing AAV cargo One of the major challenges in detecting barcoded AAV capsid transduction at the single-cell level is both detecting the mRNA sequence, which provides information about cell identity, and simultaneously detecting the delivered AAV capsid DNA or expressed RNA. To overcome this challenge, we engineered an AAV-delivered expression construct that uses the human U6 promoter to drive robust expression of the barcode sequence (Figure 1A-B). We designed the barcode sequence to be detectable by multiple methodologies, including amplicon sequencing, single-cell RNA sequencing, and in situ sequencing. The construct shown in Figure 1A is packaged into the AAV genome, as shown in Figure 1B and D. As an example, the complete DNA sequence from such a construct is contained in SEQ ID NO: 14, where nucleotides 1-141 are the AAV-1 ITRs, nucleotides 148-404 are the human U6 promoter, nucleotides 187-207 are the binding site for Pr766, nucleotide 404 is the human U6 transcription start site, nucleotides 413-434 are the binding site for the BC0108 scrambled primer, nucleotides 441-456 are the 3' padlock sequence, nucleotides 457-464 are a 9 nt RNA barcode, nucleotides 466-488 are the 5' padlock sequence, and nucleotides 495-516 are the Split-seq Pr The AAV1 Cap sequence is the Rev binding site, nucleotides 525-705 are the Pr40 promoter, nucleotides 1028-3274 are the coding sequence for the AAV1 capsid, nucleotides 2807-2827 are the coding sequence for the (NNK)7 peptide, nucleotides 3117-3139 are the Pr521 Rev binding site, and nucleotides 3408-3548 are the AAV2 ITRs. Additionally, the AAV Cap gene sequence has been modified to contain peptide inserts. Importantly, each modified Cap sequence is paired with a single RNA barcode, "RNAbc." These pairings can be resolved by long-read sequencing, which captures both the RNAbc and Cap insert sequences.Figure 1C provides an example of successful exclusive RNAbc sequence amplification after reverse transcription using primers pr749 and 750, in contrast to the nonspecific amplification seen with other primer sets. Sanger sequencing across each insert confirmed the successful creation of this dual-barcoded construct (Figure 1E and F).
[0132] (Table 2) Primer sequences TIFF2025532627000003.tif33150
[0133] TIFF2025532627000004.tif226146
[0134] Example 2 - Adaptation of split-pool ligation-based whole transcriptome sequencing (SPLiT-seq) for AAV.RNAbc detection As shown in Figure 2A, during three rounds of barcoding, fixed and permeabilized nuclei are randomly distributed into each well of a 96-well plate. Each well of the round 1 barcoding plate corresponds to two barcoding primers: an oligo(dT) RT primer and an AAV-specific RT primer for capturing AAV.RNAbc transcripts. During round 1, both poly(A) transcripts and AAV.RNAbc transcripts from the same nucleus are reverse transcribed and labeled with the same round 1 barcode. All nuclei from the same well receive the same round 1 barcode, allowing for encoding of sample information via the round 1 well location. After reverse transcription, all nuclei are pooled and randomly redistributed in rounds 2 and 3. Rounds 2 and 3 of barcoding consist of ligation reactions to add additional single-nucleus barcodes. Round 3 of barcoding also adds a unique molecular identifier (UMI). After three rounds of barcoding in a 96-well plate, 884,736 possible nuclear barcode combinations are possible (96 x 96 x 96). Prior to lysis, decrosslinking, and streptavidin bead-based cDNA isolation, all nuclei are pooled and divided into sublibraries of less than 10,000 nuclei.
[0135] Figure 2B shows next-generation library preparation for whole transcriptome and AAV.RNAbc amplicon sequencing. A template-switching reaction adds a 5' consensus sequence for full-length cDNA amplification. After amplification, the library is split for whole transcriptome and AAV.RNAbc amplicon sequencing from the same nucleus. The library for whole transcriptome sequencing is fragmented before undergoing end repair, A-tail addition, and adapter ligation (see Figure 2B, part I). A final PCR is performed to add Illumina adapters and dual indexes. Paired-end Illumina sequencing is then performed. Read 1 contains sequence information for both mRNA and AAV.RNAbc, while Read 2 corresponds to a single nucleus barcode for downstream demultiplexing. To enrich for AAV.RNAbc sequences, a second PCR-based amplification is performed with a 5' primer sequence upstream of the AAV barcode (see Figure 2B, part II). In addition to enrichment, this step controls the size and start position of the AAV.RNAbc amplicon. The forward primer is also modified to add a phosphate to the PCR product, allowing for subsequent ligation of Illumina sequencing adapters. A-tailing and adapter ligation are then performed before the final PCR to add Illumina adapters and sample indexes. Paired-end Illumina sequencing is then performed. Here, Read 1 output corresponds to the AAV.RNAbc amplicon sequence, and Read 2 corresponds to a single nucleus barcode for downstream demultiplexing. Single nucleus barcodes are used to identify the cell type identity of cells expressing AAV.RNAbc transcripts, allowing for cDNA libraries from the same nucleus to be split after amplification and processed simultaneously through both library preparations.
[0136] Example 3 - Detection of RNA barcodes (RNAbc) expressed after transfection of HEK293 cells HEK293 cells were transfected with plasmids containing either AAV.RNAbc or AAV.noBarcode. Unbiased clustering of SPLiT-Seq barcoded single cells using uniform manifold approximation and projection (UMAP) is shown in Figure 3A. Unique UMI counts were obtained from Illumina sequencing reads after amplification and library preparation exclusive to the RNAbc amplicon (Figure 3B). As expected, the majority of counts belonged to single cells that received the AAV.RNAbc construct. A small minority of counts were found to originate from cells treated with AAV.noBarcode, likely representing doublet nuclei. This is consistent with the 0.1% doublet rate of SPLiT-Seq. Seurat single-cell objects were subset to include only cells that received AAV.RNAbc treatment (Figure 3C) or only cells that received the AAV.eGFP construct (Figure 3D). When the Seurat single-cell objects were subset to include only cells that received the AAV.eGFP construct, a background level of counts was observed originating from the AAV.RNAbc amplicon. Further filtering in non-monoculture contexts helps to remove this background.
[0137] Example 4 - In vivo detection of 67 AAV1 capsid variants in 1430 of 6739 nuclei after delivery of an AAV variant library into the mouse brain A library of 67 AAV1 capsid variants was delivered into mouse brains via intrastriatal and intrathalamic injections. After incubation, brain tissue was collected, and the hippocampus, thalamus, striatum, and cortex were microdissected for single-nucleus isolation (Figure 4A). Fixed and permeabilized nuclei were processed using the double poly(A) and -AAV.RNAbc reverse transcription barcoding method described herein. SPLiT-Seq-based barcoding was then performed to apply single-cell barcodes to the mRNA and AAV transcripts contained within the permeabilized nuclei. After adding three single-cell barcodes, the nuclei were separated into two pools and lysed. In parallel, the barcoded mRNA and AAV.RNA were then amplified and tagged with Illumina indexes and sequencing adapters to facilitate sequencing on an Illumina NovaSeq 6000.
[0138] The resulting fastq files were processed using a custom bioinformatics pipeline that integrated the RNAbc sequences with AAV-peptide-insert information obtained from long-read sequencing of the input capsid library. This enabled the conversion of detected RNAbc to AAV peptide insert counts. Count tables of cDNA counts per gene and RNAbc counts per gene were generated and run in Seurat for downstream single-cell analysis and capsid transduction quantification.
[0139] Using matching single-cell barcodes between the cDNA and AAV-derived datasets, we were able to create a single Seurat object containing both cDNA expression and AAV transduction information (Figure 4B). This allowed us to determine which single cells were transduced (Figure 4C). Then, by examining disease-associated cell types, such as Drd1- and Drd2-positive medium spiny neurons (MSNs), we were able to assess the AAV transduction status within a single cell type of interest (Figure 4D).
[0140] This dataset, and future datasets generated using this technology, allows us to examine the tissues utilized in this study to identify which capsids perform best within the tissue of interest (Figure 5A) and within individual cell types of interest (Figure 5B). This information can also be represented spatially using a UMAP unbiased clustering plot (Figure 5C). These tools allow for the identification of capsid variants with performance characteristics applicable to disease treatment.
[0141] All of the methods disclosed and claimed herein can be made and executed without undue experimentation in light of this disclosure. Although the compositions and methods of the present invention have been described in terms of preferred embodiments, it will be apparent to those skilled in the art that variations can be applied to the methods described herein and to the steps or order of steps of the methods without departing from the concept, spirit, and scope of the invention. More specifically, it will be apparent that certain agents that are both chemically and physiologically related may be substituted for the agents described herein while still achieving the same or similar results. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope, and concept of the invention as defined by the appended claims.
[0142] References The following references, to the extent that they provide exemplary procedural or other details supplementary to those set forth herein, are specifically incorporated herein by reference. TIFF2025532627000005.tif11146
Claims
1. A population of recombinant adeno-associated virus (rAAV) vectors, each rAAV vector independently comprising (i) a modified adeno-associated virus (AAV) Cap gene encoding a modified AAV capsid protein comprising a targeting peptide, and (ii) an expression cassette encoding a barcode sequence operably linked to an RNA polymerase III promoter, wherein the targeting peptide and the barcode in each rAAV vector are uniquely paired.
2. 2. The population of rAAV vectors of claim 1, wherein the barcode sequences are 9 to 20 nucleotides in length.
3. The barcode sequence is (NNNT) n 3. The population of rAAV vectors of claim 1 or 2, wherein
4. 4. The population of rAAV vectors of any one of claims 1 to 3, wherein the barcode sequence is flanked by sequences that can hybridize to and activate a padlock probe.
5. The population of rAAV vectors of any one of claims 1 to 4, wherein the RNA polymerase III promoter is a type III RNA polymerase III promoter.
6. The population of rAAV vectors of claim 8, wherein the RNA polymerase promoter is a U6 snRNA gene promoter, an H1 RNA gene promoter, or a 7SK gene promoter.
7. 7. The population of rAAV vectors of any one of claims 1 to 6, further comprising a reverse transcription primer binding site located 3' of the barcode sequence and an enrichment primer binding site located 5' of the barcode sequence.
8. The population of rAAV vectors of any one of claims 1 to 7, wherein the expression cassette comprises a sequence that is at least 90% identical to SEQ ID NO:
7.
9. The population of rAAV vectors of any one of claims 1 to 8, wherein the modified AAV capsid protein is a modified AAV1 capsid protein, a modified AAV2 capsid protein, or a modified AAV9 capsid protein.
10. 10. The population of rAAV vectors of claim 9, wherein the modified AAV capsid protein is derived from the AAV1 capsid protein (see SEQ ID NO: 1) and the targeting peptide is inserted after residue 590 of the AAV1 capsid protein.
11. 11. The population of rAAV vectors of claim 10, wherein the targeting peptide is flanked by linker sequences, the linker sequences on each side of the targeting peptide being 2 or 3 amino acids in length.
12. 12. The population of rAAV vectors of claim 11, wherein the linker sequence is SSA at the N-terminus of the targeting peptide and AS at the C-terminus of the targeting peptide.
13. The population of rAAV vectors of any one of claims 10 to 12, wherein the modified AAV1 capsid protein has a sequence that is at least 95% identical to SEQ ID NO:
4.
14. 10. The population of rAAV vectors of claim 9, wherein the modified AAV capsid protein is derived from the AAV2 capsid protein (see SEQ ID NO: 2) and the targeting peptide is inserted after residue 587 of the AAV2 capsid protein.
15. 15. The population of rAAV vectors of claim 14, wherein the targeting peptide is flanked by linker sequences, the linker sequences on each side of the targeting peptide being 2 or 3 amino acids in length.
16. 16. The population of rAAV vectors of claim 15, wherein the linker sequence is AAA at the N-terminus of the targeting peptide and AA at the C-terminus of the targeting peptide.
17. The population of rAAV vectors of any one of claims 14 to 16, wherein the modified AAV2 capsid protein has a sequence that is at least 95% identical to SEQ ID NO:
5.
18. 10. The population of rAAV vectors of claim 9, wherein the modified AAV capsid protein is derived from the AAV9 capsid protein (see SEQ ID NO: 3) and the targeting peptide is inserted after residue 588 of the AAV9 capsid protein.
19. 19. The population of rAAV vectors of claim 18, wherein the targeting peptide is flanked by linker sequences, the linker sequences on each side of the targeting peptide being 2 or 3 amino acids in length.
20. 20. The population of rAAV vectors of claim 19, wherein the linker sequence is AAA at the N-terminus of the targeting peptide and AS at the C-terminus of the targeting peptide.
21. 21. The population of rAAV vectors of any one of claims 18 to 20, wherein the modified AAV9 capsid protein has a sequence that is at least 95% identical to SEQ ID NO:
6.
22. 22. The population of rAAV vectors of any one of claims 8 to 21, wherein the targeting peptide is 3 to 10 amino acids in length.
23. 23. The population of rAAV vectors of claim 22, wherein the targeting peptide is 7 amino acids in length.
24. 24. The population of rAAV vectors of any one of claims 8-23, wherein the population comprises a plurality of capsid protein targeting peptides, each capsid protein targeting peptide paired with more than one barcode sequence.
25. 25. The population of rAAV vectors of any one of claims 8 to 24, wherein the population comprises multiple capsid protein targeting peptides, and all rAAVs having the same barcode sequence also have the same capsid protein targeting peptide.
26. A population of cells comprising a population of rAAV vectors of any one of claims 1 to 25.
27. 27. The population of cells of claim 26, wherein the cells are mammalian cells.
28. 27. The population of cells of claim 26, wherein the cells are human cells.
29. 27. The population of cells of claim 26, wherein the cells are in vitro.
30. 27. The population of cells of claim 26, wherein the cells are in vivo.
31. A method for determining the cell tropism of a recombinant adeno-associated virus (rAAV) having a modified AAV capsid protein containing a targeting peptide, comprising: (i) contacting various cell types with the modified rAAV vector of any one of claims 6 to 25; (ii) identifying cells transduced by the modified rAAV vector based on the presence of a barcode sequence; and (iii) detecting the expressed transcriptome of each transduced cell on a cell-by-cell basis, thereby determining the cellular tropism of the modified rAAV.
32. A method for determining the cell tropism of a recombinant adeno-associated virus (rAAV) having a modified AAV capsid protein containing a targeting peptide, comprising: (i) contacting a variety of cell types with a population of rAAV vectors of any one of claims 6 to 25; (ii) detecting, on a cell-by-cell basis, both the expressed transcriptome and the rAAV transduced into each cell; and (iii) determining which cell types are transduced by which modified rAAV vectors, thereby determining the cell tropism of the modified rAAV.
33. 33. The method of claim 32, wherein the contacting step in (i) is performed in vitro.
34. 33. The method of claim 32, wherein the contacting step in (i) is performed in vivo.
35. (ii) detecting the expressed transcriptome and the rAAV, (a) isolating, fixing, and permeabilizing the nuclei of the cells contacted in (i); (b) dividing the nuclei into a plurality of first aliquots; (c) reverse transcribing the cellular RNA molecules expressed in the nucleus using a primer containing a poly(T) sequence to form a complementary DNA (cDNA) molecule; and reverse transcribing the rAAV RNA molecules expressed in the nuclei using a primer containing a sequence sufficient to hybridize to and reverse transcribe a barcode sequence within the expression cassette to form an AAV amplicon; (d) labeling the cDNA molecules and AAV amplicons with a first 5' barcode, wherein the first 5' barcode for the primer in each first aliquot is unique such that the cDNA molecules and AAV amplicons from the nuclei of each aliquot can be distinguished relative to the cDNA molecules and AAV amplicons from the nuclei of all other aliquots; (e) combining said plurality of first aliquots; (f) dividing the combined first aliquots into a plurality of second aliquots; (g) ligating a second 5' barcode to the 5' ends of the cDNA molecule and the AAV amplicon to form a dual-barcoded cDNA molecule and AAV amplicon, wherein the second 5' barcode in each second aliquot is unique; (h) combining said plurality of second aliquots; (i) dividing the combined first aliquots into a plurality of third aliquots; (j) ligating a third 5' barcode to the 5' ends of the cDNA molecules and the AAV amplicons to form triple-barcoded cDNA molecules and AAV amplicons, wherein the third 5' barcode in each third aliquot is unique; (k) combining said plurality of third aliquots; (l) lysing the nuclei to release the cDNA molecules and the AAV amplicons from within the nuclei to form a lysate; and (m) sequencing the cDNA molecules and the AAV amplicons, thereby detecting both the expressed transcriptome and the rAAV that transduced each cell.
35. The method of any one of claims 32 to 34, comprising:
36. 36. The method of claim 35, wherein the cDNA molecules and AAV amplicons are labeled with a first 5' barcode simultaneously with reverse transcription, and the reverse transcription primer comprises the first 5' barcode.
37. 37. The method of claim 35 or 36, wherein the nuclei are fixed and permeabilized at less than about 8°C, less than about 7°C, less than about 6°C, less than about 5°C, less than about 4°C, less than about 3°C, less than about 2°C, or less than about 1°C.
38. 38. The method of any one of claims 35-37, wherein a majority of said triple-barcoded cDNA molecules and AAV molecules from a single nucleus comprise the same set of barcodes.
39. 39. The method of claim 38, wherein a majority of the triple-barcoded cDNA and AAV molecules from a single nucleus have a unique set of barcodes compared to triple-barcoded cDNA and AAV molecules from other nuclei.
40. 40. The method of any one of claims 32 to 39, wherein the cell type is determined based on the expressed transcriptome.
41. Sequencing the cDNA molecules and the AAV amplicons includes preparing a sequencing library, and preparing the sequencing library includes: (i) adding a common adapter sequence to the 3' end of the cDNA molecule and the AAV amplicon; (ii) performing amplification of full-length cDNA and AAV amplicons; (iii) fragmenting the amplified full-length cDNA and AAV amplicons; (iv) repairing and A-tailing the ends of the fragmented cDNA and AAV amplicons; (v) ligating adapters to the 5' ends of the end-repaired and A-tailed cDNA and AAV amplicons; and (vi) performing sample index PCR to add sequencing adapters and double indexes to the adapter-ligated cDNA and AAV amplicons; 41. The method of any one of claims 32 to 40, comprising:
42. Sequencing the AAV amplicons comprises preparing a sequencing library enriched for the AAV amplicons, and preparing the sequencing library comprises: (i) adding a common adapter sequence to the 3' end of the cDNA molecule and the AAV amplicon; (ii) performing full-length amplification of cDNA and AAV amplicons; (iii) performing an AAV amplicon enrichment amplification with a forward primer that hybridizes to the AAV amplicon upstream of the AAV barcode, wherein the forward primer has a 5' phosphate; (iv) adding an A-tail to the AAV amplicon having a 5' phosphate; (v) ligating an adapter to the 5' end of the A-tailed AAV amplicon; and (vi) performing sample index PCR to add sequencing adapters and double indexes to the adapter-ligated AAV amplicons; 42. The method of any one of claims 32 to 41, comprising:
43. 43. The method of claim 41 or 42, wherein the common adapter sequence is added to the 3' end of the cDNA molecule and the AAV amplicon by template switching.
44. 44. The method of any one of claims 35 to 43, wherein the sequencing is paired-end sequencing, amplicon sequencing, single-cell RNA sequencing, or in situ sequencing.