Using base modifications to characterize viral genomes
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2026-03-18
AI Technical Summary
The prior art has difficulty in efficiently characterizing and packing single-stranded DNA AAV viral particles, especially lacking detailed analysis on the distinction between sense and antisense sequences.
Complementary synthetic DNA strands are synthesized by introducing modified bases (such as 5-methylcytosine) on single-stranded DNA of virus particles to form double-stranded DNA products and analyzed on a nanopore sequencing platform using long-read sequence technology to determine the specificity of DNA strands.
High-resolution quality control of single-strand DNA in viral particles is achieved, which can accurately distinguish sense and antisense sequences, and improve the efficiency and safety of virus packaging and delivery.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to methods for characterizing viral genomes and packaging (eg, AAV genomes) using modified bases. [Background technology]
[0002] Adeno-associated virus (AAV), a non-enveloped single-stranded DNA virus, has emerged as an attractive class of therapeutic agent to deliver genetic material to host cells for gene therapy due to its ability to transduce a wide range of species and tissues in vivo, low risk of immunotoxicity, and mild innate and adaptive immune responses. Recombinant AAV technology depends on proper genome packaging, which can be both sense and antisense conformation. A standard method for evaluating the DNA content packaged in AAV particles is still unknown.
[0003] Existing methods allow characterization of structural anomalies in AAV packaging of self-complementary AAV variants, but do not provide detailed analysis of the single-stranded AAV genome to account for its strand specificity. The complex nature of viral vectors such as AAV requires improved methods to enable product testing and characterization. Thus, a method is needed to determine the specificity of the single-stranded DNA of viral particles. Summary of the Invention
[0004] The present disclosure is directed to a method for differentiating single-stranded DNA strands that preserve strand-specific information in samples of viral particles (e.g., AAV) using a combination of modified base incorporation and long-read sequencing approaches. In an exemplary embodiment, modified bases such as 5-methylcytosine (5mC) are used to label single-stranded DNA detected by a nanopore sequencer platform, allowing identification of the specificity of the sequenced strand. Selective labeling of single-stranded DNA of viral genomes and its subsequent analysis uniquely enables quantitative molecular-level quality control of viral gene therapy and adds additional resolution to the analysis of single-stranded viral constructs.
[0005] In one aspect, the disclosure provides a method for assaying a sample of viral particles comprising a single-stranded DNA genome for strand specificity comprising: (a) synthesizing a synthetic DNA strand complementary to a natural DNA strand of the sample of viral particles to obtain a double-stranded DNA product, wherein the synthetic DNA strand is synthesized by incorporating a modified base; (b) purifying the synthesized double-stranded DNA product; (c) sequencing the purified double-stranded DNA product on a sequencer, wherein the sequencer identifies a sequence of nucleotides within one sequenced strand of the double-stranded DNA product; and (d) determining the specificity of the natural DNA strand by identifying the specificity of the sequenced strand and detecting whether the sequenced strand contains a modified base, wherein the presence of a modified base indicates that the sequenced strand is a synthetic DNA strand of the double-stranded DNA product.
[0006] In some embodiments, the sample of viral particles comprises adeno-associated viral (AAV) particles. In some cases, the AAV particles are of the serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV-DJ, AAV-DJ / 8, AAV-Rh10, AAV-retro, AAV-PHP.B, AAV8-PHP.eB, or AAV-PHP.S.
[0007] In some embodiments, the modified base comprises 5-methylcytosine, 5-hydroxymethylcytosine, or N6-methyldeoxyadenosine. In some embodiments, the modified base is 5-methylcytosine.
[0008] In some embodiments, the synthetic DNA strand is methylated by incorporation of a modified base.
[0009] In some embodiments, the double-stranded DNA product is purified using magnetic beads.
[0010] In some embodiments, the sequencer is a nanopore sequencer.
[0011] In some embodiments, the nanopore sequencer allows for the entry of either the native DNA strand of the purified double-stranded DNA product or a selectively labeled synthetic DNA strand for sequencing.
[0012] In some embodiments, the natural DNA strands or selectively labeled synthetic DNA strands are sequenced using a long-read sequencing approach.
[0013] In some embodiments, modified bases in synthetic DNA strands are detected by changes in current flow through the nanopore sequencer. In some cases, detection of the modified bases distinguishes the synthetic DNA strand from a natural DNA strand that enters the nanopore sequencer.
[0014] In some embodiments, the method further includes determining the percentage of packaged sense and antisense strands of the sample of viral particles based on detection of the modified bases on the synthetic DNA strands.
[0015] In one aspect, the disclosure provides a method for distinguishing between natural and synthetic DNA strands in a sample of adeno-associated virus (AAV) particles comprising a single-stranded DNA genome, the method comprising: (a) synthesizing a synthetic DNA strand complementary to a natural DNA strand to obtain a double-stranded DNA product, the synthetic DNA strand comprising a modified base that selectively labels the synthetic DNA strand; (b) purifying the double-stranded DNA product; (c) sequencing the purified double-stranded DNA product on a nanopore sequencer; and (d) determining the specificity of the natural DNA strand by identifying the specificity of the sequenced strand and detecting whether the sequenced strand contains the modified base of the synthetic DNA strand during nanopore sequencing by detecting a change in current flow through the nanopore sequencer, wherein the presence of the modified base indicates that the sequenced strand is the synthetic DNA strand of the double-stranded DNA product.
[0016] In some embodiments, the modified base comprises 5-methylcytosine, 5-hydroxymethylcytosine, or N6-methyldeoxyadenosine. In some embodiments, the modified base is 5-methylcytosine.
[0017] In some embodiments, the selective labeling results in methylation of the synthetic DNA strand.
[0018] In some embodiments, the natural DNA strands or selectively labeled synthetic DNA strands are sequenced using a long-read sequencing approach.
[0019] In some embodiments, the double-stranded DNA product is purified using magnetic beads.
[0020] In some embodiments, the AAV particles are of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV-DJ, AAV-DJ / 8, AAV-Rh10, AAV-retro, AAV-PHP.B, AAV8-PHP.eB, or AAV-PHP.S.
[0021] In some embodiments, the method further comprises determining the orientation of the sequenced strand based on detection of the modified bases of the synthetic DNA strand.
[0022] In some embodiments, the method further includes identifying the percentage of each of the sense and antisense strands packaged in the sample of adeno-associated virus (AAV) particles based on the determination of the orientation of the sequenced strands.
[0023] In some embodiments, the viral particle comprises a foreign gene in its single-stranded DNA genome. In some cases, the foreign gene is a therapeutic gene.
[0024] In some embodiments, the method further comprises determining the complete sequence of the sequenced strand.
[0025] In some embodiments, the method further comprises identifying a complementary nucleotide sequence that corresponds to the complete sequence when the sequenced strand is a synthetic DNA strand of a double-stranded DNA product.
[0026] In any of the various embodiments described above or discussed herein, the sequencing of the nucleic acid identifies between about 100 and about 5000 nucleotides in the sequenced strand. In some cases, the sequencing identifies between 100 and 1000 nucleotides in the sequenced strand. In some cases, the sequencing identifies between 100 and 500 nucleotides in the sequenced strand.
[0027] In any of the various embodiments described above or discussed herein, the sequencing of the nucleic acid is performed without fragmentation of the sequenced strand.
[0028] In various embodiments, any of the features or components of the embodiments described above or discussed herein may be combined, and such combinations are encompassed within the scope of the present disclosure. Any specific value described above or discussed herein may be combined with another associated value described above or discussed herein to recite a range having values representing the upper and lower limits of the range, and such ranges are encompassed within the scope of the present disclosure.
[0029] Other embodiments will be apparent from consideration of the detailed description that follows. [Brief description of the drawings]
[0030] [Figure 1] FIG. 1 illustrates an AAV capsid carrying a single-stranded DNA genome containing a therapeutic gene or gene of interest (GOI).
[0031] [Diagram 2] FIG. 2 illustrates a schematic of an exemplary method for assaying AAV particles for strand specificity according to an embodiment of the present disclosure.
[0032] [Diagram 3] 3 illustrates a process of second strand synthesis according to an embodiment of the present disclosure, showing the incorporation of modified bases and selective labeling of the synthesized second strand. "CH3" refers to methylation.
[0033] [Figure 4] FIG. 4 illustrates the purification of second strand synthesis products using magnetic beads.
[0034] [Diagram 5] FIG. 5 illustrates sequencing of second strand synthesis products on a nanopore sequencer and determining the sequence of the sequenced strand based on changes in current flow through the nanopore sequencer.
[0035] [Figure 6]6A and 6B illustrate the chemical structures of the unmodified base cytosine and the modified base 5-methylcytosine, respectively.
[0036] 6C illustrates sequencing of selectively labeled versus unlabeled DNA strands via a nanopore sequencer, where detection of altered current flow indicates the presence of the modified base 5-methylcytosine. "Me" or "CH3" refers to methylation. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0037] Before describing the present invention, it is to be understood that the present invention is not limited to the particular methods and experimental conditions described, as such methods and conditions may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention is limited only by the appended claims.
[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.As used herein, the term "about" when used in relation to a specific referenced numerical value means that the value may vary by 1% or less from the referenced value.For example, as used herein, the expression "about 100" includes 99 and 101, and all values therebetween (e.g., 99.1, 99.2, 99.3, 99.4, etc.).
[0039] As used herein, the terms "include", "includes", and "including" are intended to be open-ended and are understood to mean "comprise", "comprises", and "comprising", respectively.
[0040] Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are now described. All patents, applications, and non-patent publications mentioned herein are incorporated by reference in their entirety. Selected Abbreviations
[0041] AAV: adeno-associated virus
[0042] ssDNA: Single-stranded DNA
[0043] dsDNA: double stranded DNA
[0044] GOI: Gene of interest
[0045] PCR: polymerase chain reaction
[0046] 5mC: 5-methylcytosine
[0047] dNTP: deoxyribonucleotide triphosphate
[0048] dATP: deoxyadenosine triphosphate
[0049] dGTP: deoxyguanosine triphosphate
[0050] dTTP: deoxythymidine triphosphate
[0051] dCTP: deoxycytidine triphosphate
[0052] NFW: Nuclease-free water definition
[0053] "Adeno-associated virus" or "AAV" is a non-pathogenic parvovirus with a single-stranded DNA, non-enveloped genome of approximately 4.7 kb, and an icosahedral conformation. AAV was first discovered in 1965 as a contaminant of adenovirus preparations. AAV belongs to the genus Dependovirus and family Parvoviridae, and requires helper functions from either herpesvirus or adenovirus for replication. In the absence of helper virus, AAV can set up latency by integrating into human chromosome 19 at position 19q13.4. The AAV genome consists of two open reading frames (ORFs), one for each of the two AAV genes Rep and Cap. The AAV DNA termini have 145 bp inverted terminal repeats (ITRs), and the 125 terminal bases are palindromic, leading to a characteristic T-shaped hairpin structure.
[0054] As used herein, the term "sample" refers to a mixture of viral particles (e.g., AAV particles) that contain at least one viral capsid component encapsulating a single-stranded DNA genome to be manipulated according to the methods of the invention, including, for example, selective labeling and sequencing.
[0055] The term "nucleic acid" as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, the term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The backbone of a nucleic acid can include sugars and phosphate groups (as typically found in RNA or DNA), or modified or substituted sugar or phosphate groups.
[0056] "Recombinant viral particle" refers to a viral particle that includes one or more foreign genes or heterologous sequences (eg, nucleic acid sequences not of viral origin) that may be adjacent to at least one viral nucleotide sequence.
[0057] "Recombinant AAV particle" refers to an adeno-associated virus particle that includes one or more heterologous sequences (e.g., nucleic acid sequences not of AAV origin), which may be flanked by at least one, e.g., two, AAV inverted terminal repeats (ITRs). Such rAAV particles can be replicated and packaged when present in a host cell that is infected with a suitable helper virus (or expresses suitable helper functions) and expresses the AAV rep and cap gene products (i.e., AAV Rep and Cap proteins).
[0058] "Virus particle" refers to a viral particle that is composed of at least one viral capsid protein and an encapsidated viral genome.
[0059] "Heterologous" or "foreign" means derived from a genotypically distinct entity from the rest of the entity to which it is compared or to which it is introduced or incorporated. For example, a nucleic acid that is introduced into a different cell type by genetic engineering techniques is a heterologous nucleic acid (which, when expressed, can encode a heterologous polypeptide). Similarly, a cellular sequence (e.g., a gene or portion thereof) that is incorporated into a viral particle is a heterologous or foreign nucleotide sequence with respect to the viral particle.
[0060] The term "therapeutic gene" refers to a genetically modified gene that produces a therapeutic effect (eg, by encoding a protein of interest) or the treatment of a disease by repairing or rearranging defective genetic material.
[0061] "Inverted terminal repeat" or "ITR" sequences are relatively short sequences found at the ends of the viral genome and are inverted. The "AAV inverted terminal repeat (ITR)" sequence is an approximately 145 nucleotide sequence present at both ends of the single-stranded AAV genome.
[0062] The term "isolated" as used herein refers to a biological component (such as a nucleic acid, peptide, protein, lipid, viral particle, or metabolite) that has been substantially separated, produced, or purified from other biological components in the cells of an organism in which the component naturally occurs or is transgenically expressed.
[0063] As used herein, a "vector" refers to a recombinant plasmid or virus containing a nucleic acid that is delivered into a host cell either in vitro or in vivo.
[0064] The term "corresponding" is a relative term indicating similarity in location, purpose, or structure.
[0065] The term "read" in the context of sequencing refers to the nucleic acid sequence of a collection of nucleotides obtained after the sequencing process is completed, which is ultimately the sequence of a section of a complete nucleic acid sequence. A "read" is a series of nucleotide base call values derived from the raw signal.
[0066] The term "long-read sequencing" refers to a DNA sequencing technique that can determine the sequence of nucleotides of a long sequence of DNA from about 100 base pairs to about 1,000,000 base pairs or more at a time (the upper limit may depend only on the genome being sequenced), thereby eliminating the need to fragment and amplify the DNA that is typically required with other DNA sequencing techniques. In one example, the AAV genome (which is about 5000 base pairs) can be sequenced in its entirety without any fragmentation.
[0067] As used herein, "amplification" refers to the production of multiple copies of a segment of DNA or RNA. Amplification is usually induced by the polymerase chain reaction.
[0068] As used herein, "PCR" refers to the polymerase chain reaction, a molecular biology technique used to amplify single copies of DNA or RNA segments, generating thousands to millions of copies of a particular DNA or RNA sequence. PCR is commonly used to amplify the number of copies of DNA or RNA segments for cloning or other analytical procedures.
[0069] The term "nanopore sequencing" refers to the sequencing of a nucleic acid molecule resulting from the change in electrical current flow with each base as the nucleic acid molecule passes through a nanopore.
[0070] The term "strand specificity" refers to the plus or sense and minus or antisense strands of the viral genome. The plus or sense strand is the coding strand of the gene and the minus or antisense strand is the non-coding strand. overview
[0071] The present disclosure provides a method for the synthesis and labeling of synthetic DNA strands by incorporation of modified bases, and the strand specificity of the ssDNA genome of a sample of viral particles (e.g., AAV particles) is rapidly identified by sequencing the labeled synthetic DNA strand. The method utilizes a long-read sequencing-based approach on a nanopore sequencer to identify and quantify the strand specificity of the viral genome. Quantitative characterization to identify the proportion of each of the packaged sense and antisense strands in a sample of AAV particles is necessary to ensure the quality and consistency of the product, thereby greatly streamlining the quality control process. Methods for identifying and quantifying strand specificity of viral single-stranded DNA genomes
[0072] Aspects of the present disclosure are directed to methods for identifying and quantifying the strand specificity of ssDNA genomes in samples of viral particles (e.g., recombinant AAV particles) using a combination of modified base incorporation and long-read sequencing approaches.
[0073] In some cases, the method includes: (a) synthesizing a synthetic DNA strand complementary to a natural DNA strand of the sample of viral particles to obtain a double-stranded DNA product, wherein the synthetic DNA strand is synthesized by incorporating a modified base; (b) purifying the synthesized double-stranded DNA product; (c) sequencing the purified double-stranded DNA product on a sequencer, wherein the sequencer identifies a sequence of nucleotides within one sequenced strand of the double-stranded DNA product; and (d) determining the specificity of the natural DNA strand by identifying the specificity of the sequenced strand and detecting whether the sequenced strand contains a modified base, wherein the presence of a modified base indicates that the sequenced strand is the synthetic DNA strand of the double-stranded DNA product.
[0074] An adeno-associated virus particle 100 is illustrated with its single-stranded DNA genome in Figure 1. In the example shown in Figure 1, the ssDNA genome of AAV is highly symmetric with palindromic elements. In addition, the ssDNA genome of AAV has a GC content of about 70% and has inverted terminal repeats. In some examples, the single-stranded DNA genome of the recombinant AAV can contain a therapeutic gene or gene of interest (GOI), for example, for gene therapy purposes.
[0075] In the methods disclosed herein, a method 200 for assaying AAV particles for strand specificity is illustrated by the schematic diagram shown in FIG. 2. In the exemplary method shown in FIG. 2, ssDNA genomes 204 are extracted or isolated from a sample 202 of AAV particles. In one example, the isolated ssDNA genomes 204 may be prepared by dissolving nucleocapsids and releasing viral ssDNA using a phenol-chloroform extraction method, while in another example, the isolated ssDNA genomes 204 may be prepared by an alkaline lysis extraction method. The isolated ssDNA genomes 204 are then subjected to synthesis of a synthetic DNA strand or synthesis of a second strand by incorporation of modified bases to obtain selectively labeled double-stranded DNA 206. The labeled double-stranded DNA 206 is purified from the second strand synthesis reaction. In one example, purification of the double-stranded DNA 206 may be performed using magnetic beads. The purified double-stranded DNA 206 is then sequenced on a nanopore sequencer 208, where the sequencer identifies the sequence of nucleotides within one sequenced strand of the double-stranded DNA product. More detailed information regarding second strand synthesis, purification of labeled double-stranded DNA, and nanopore sequencing is provided in subsequent figures.
[0076] In various embodiments of the methods discussed herein, an example of second strand synthesis 300 is illustrated in FIG. 3. A natural ssDNA strand 302 (which may be similar to the ssDNA genome 204 in FIG. 2) may be used as a template to synthesize a synthetic DNA strand complementary to the natural ssDNA strand 302. Second strand synthesis is catalyzed by Klenow fragment in the presence of nucleotide bases and modified random hexamers. The second strand synthesis reaction utilizes modified bases to selectively label the synthetic DNA strand. In various embodiments of the method, the modified bases include 5-methylcytosine, 5-hydroxymethylcytosine, or N6-methyldeoxyadenosine. In some cases, the modified base is 5-methylcytosine. Second strand synthesis results in the formation of a double-stranded DNA product 310 in which the synthetic DNA strand is methylated (CH3), as shown in FIG. 3.
[0077] The labeled double-stranded DNA product 310 is then purified from the second strand synthesis reaction. An exemplary purification method 400 is illustrated in FIG. 4, showing the purification of the labeled double-stranded DNA product 310 using magnetic beads 410. The labeled dsDNA product 310 is conjugated to the magnetic beads 410 and separated from the second strand synthesis reaction using a magnet. The labeled dsDNA product 310 is then washed and eluted from the magnetic beads 410.
[0078] In embodiments of the methods discussed herein, the purified double-stranded DNA product is finally subjected to sequencing. As shown in FIG. 5, sequencing 500 of the product of second strand synthesis is performed on a nanopore sequencer 502 using long read technology that provides single reads of about 100 base pairs to 1000 kb or more. As shown, the nanopore 504 of the nanopore sequencer 502 identifies the sequence of nucleotides within one sequenced strand of the double-stranded DNA product. Identification of the nucleotide sequence within the sequenced strand is based on the ability of each nucleotide base to uniquely change the current flow while passing through the nanopore 504. A plot 506 showing each base changing the current flow over time is illustrated in FIG. 5. In each example, the double-stranded DNA product does not need to be fragmented for sequencing, thereby allowing reads up to the length of the entire viral (e.g., AAV) genome. In various embodiments, the number of nucleotides sequenced without fragmentation is about 100 to about 1,000,000. In some cases, the number of nucleotides sequenced without fragmentation is from about 100 to about 10,000. In some cases, the number of nucleotides sequenced without fragmentation is from about 100 to about 9000, up to about 8000, up to about 7000, up to about 6000, or up to about 5000 (e.g., nucleotides in the sequenced strand).
[0079] The resulting sequence data is analyzed for the presence of modified bases in the sequenced strand of the dsDNA product. The chemical structures of the unmodified nucleobase cytosine and the modified nucleobase 5-methylcytosine with an additional methyl group (CH3) are illustrated in Figures 6A and 6B, respectively. Both modified and unmodified nucleobases alter the current flow through the nanopore differently. As shown in Figure 6C, detection of a modified base during sequencing indicates that the sequenced strand is a synthetic DNA strand of the double-stranded DNA product, since the synthetic DNA strand is modified (e.g., methylated). However, the absence of a modified base during sequencing indicates that the sequenced strand is a natural DNA strand of the double-stranded DNA product, since the natural DNA strand is not methylated. A plot 610 showing the change in current flow over time corresponding to the methylated (modified) strand and the unmethylated (modified) strand is illustrated in Figure 6C.
[0080] The method discussed herein includes determining the specificity of the natural DNA strand by identifying the specificity of the sequenced strand. The specificity of the natural DNA strand of the AAV genome can be determined by identifying the complementary nucleotide sequence corresponding to the complete sequence when the sequenced strand is a synthetic DNA strand of a double-stranded DNA product. Furthermore, the proportion of each of the packaged sense strand and antisense strand in a sample of AAV particles can be determined based on the detection of modified bases in the synthetic DNA strand. Virus particles
[0081] In certain embodiments, the viral particles are AAV particles, and the disclosed methods can be used to determine the specificity of ssDNA genomes and identify the proportion of each of the sense and antisense strands packaged in a sample of AAV particles.The AAV particles can be recombinant AAV (rAAV) particles.The rAAV particles include an AAV vector that encodes a heterologous transgene or a heterologous nucleic acid molecule.
[0082] In certain aspects, the AAV particles comprise an AAV1 capsid, an AAV2 capsid, an AAV3 capsid, an AAV4 capsid, an AAV5 capsid, an AAV6 capsid, an AAV7 capsid, an AAV8 capsid, an AAVrh8 capsid, an AAV9 capsid, an AAV10 capsid, an AAV11 capsid, an AAV12 capsid, or variants thereof. In certain aspects, the AAV particles are particles of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV-DJ, AAV-DJ / 8, AAV-Rh10, AAV-retro, AAV-PHP.B, AAV8-PHP.eB, or AAV-PHP.S. In some embodiments, the AAV particles are of serotype AAV1 or AAV8.
[0083] While AAV has been a model viral particle for this disclosure, it is contemplated that the disclosed methods may be applied to characterize a variety of viruses, including viral families, subfamilies, and genera. The disclosed methods may be used, for example, in characterizing viral particles, to monitor or detect the relative abundance of each of the packaged sense and antisense strands in a composition of viral particles during production, purification, or storage of such compositions.
[0084] In an exemplary embodiment, the viral particle belongs to a virus family selected from the group consisting of Parvoviridae.
[0085] In certain aspects, the viral particle belongs to a viral genus selected from the group consisting of ambidensovirus, brevidensovirus, hepandensovirus, iteradensovirus, pentyldensovirus, andoparvovirus, abeparvovirus, bocaparvovirus, copiparvovirus, dependoparvovirus, erythroparvovirus, protoparvovirus, and tetraparvovirus.
[0086] In some embodiments, the viral particle (e.g., AAV particle) comprises a heterologous nucleic acid molecule or a foreign gene (e.g., a therapeutic gene or a gene of interest). In some embodiments, the heterologous nucleic acid molecule is operably linked to a promoter. Exemplary promoters include, but are not limited to, the cytomegalovirus (CMV) immediate early promoter, the RSV LTR, the MoMLV LTR, the phosphoglycerate kinase-1 (PGK) promoter, the simian virus 40 (SV40) promoter and the CK6 promoter, the transthyretin promoter (TTR), the TK promoter, the tetracycline responsive promoter (TRE), the HBV promoter, the hAAT promoter, the LSP promoter, the chimeric liver-specific promoter (LSP), the E2F promoter, the telomerase (hTERT) promoter, the cytomegalovirus enhancer / chicken beta-actin / rabbit beta-globin promoter, and the elongation factor 1-alpha promoter (EF1-alpha) promoter. In some aspects, the promoter comprises a human beta-glucuronidase promoter or a cytomegalovirus enhancer linked to a chicken beta-actin (CBA) promoter. The promoter can be a constitutive, inducible, or repressible promoter. In some aspects, the present invention provides a recombinant vector comprising a nucleic acid encoding a heterologous transgene of the present disclosure operably linked to a CBA promoter. In some cases, a native promoter for the transgene or a fragment thereof is used. A native promoter can be used when it is desired that the expression of the transgene should mimic the native expression. A native promoter can be used when the expression of the transgene must be regulated temporally or developmentally, or in a tissue-specific manner, or in response to a specific transcriptional stimulus. In further aspects, other native expression control elements, such as enhancer elements, polyadenylation sites, or Kozak consensus sequences, can also be used to mimic the native expression. EXAMPLES
[0087] The following examples are provided to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the methods and compositions of the present invention, and are not intended to limit the scope of what the inventor regards as his invention. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should be accounted for. Unless otherwise indicated, parts are parts by weight, molecular weight is average molecular weight, temperature is degrees Celsius, and pressure is at or near atmospheric pressure. Example 1: Characterization of the viral ssDNA genome of AAV particles
[0088] AAV samples of different serotypes were prepared in-house. Total nucleic acid (ssDNA genome) was extracted from AAV cultured in human hosts. Nucleic acid extracts were subjected to second strand synthesis, which incorporates modified bases to yield labeled dsDNA products. dsDNA products were purified from the second strand synthesis reaction using Agencourt AMPure XP beads (Beckman Coulter). Sequencing of the purified dsDNA products was performed on a GridION nanopore sequencer (Oxford Nanopore Technologies) and data were analyzed using bioinformatics analysis. Chemicals and Reagents
[0089] All chemicals and reagents were obtained from MilliporeSigma (Burlington, MA, USA) unless otherwise stated. AAV samples and their nucleic acid extracts were prepared in-house (Regeneron Pharmaceuticals Inc., Tarrytown, NY, USA). Agencourt AMPure XP beads were obtained from Beckman Coulter (Indianapolis, IN, USA). Sigma-Aldrich 200 proof molecular grade ethanol was obtained from MilliporeSigma. Oxford Nanopore Technology Flow Cell Priming Kit and Oxford Nanopore Technology Rapid Sequencing Kit were purchased from Oxford Nanopore Technologies (Oxford, UK). 1.5 ml DNA / LoBind tubes were obtained from VWR (Atlanta, GA, USA). NEB Individual dNTPs and NEB Modified 5mC dCTP were obtained from Thermo Fisher Scientific (Waltham, MA, USA). Modified random hexamers were purchased from Integrated DNA Technologies (Coralville, IA, USA). Second strand synthesis experiment
[0090] Second strand synthesis reactions were performed on AAV total nucleic acid extracts. A dNTP master mix was prepared with a working concentration of 10 mM each of dATP, dTTP, dGTP, and dCTP (5mC). The reaction mixture was prepared in nuclease-free water. To the reaction mixture, 1 μg of total nucleic acid extract, 10 mM dNTP master mix, and 60 μM modified random hexamers (primers) were added. The reaction mixture was heated to 95° C. for 3 minutes and then allowed to slowly cool at room temperature for 10 minutes. 2 U / μL Klenow enzyme was then added to the reaction mixture and the reaction mixture was incubated at 30° C. for 1 hour. The second strand synthesis reaction was stopped by incubating the reaction mixture at 75° C. for 10 minutes. Purification of second strand synthesis products
[0091] The second strand synthesis reaction was placed in a 1.5 mL DNA / RNA LoBind tube prior to the purification assay. Purification of the dsDNA product from the second strand synthesis reaction was performed using Agencourt AMPure XP beads (magnetic beads). First, the dsDNA product was conjugated to magnetic beads by adding resuspended AMPure XP beads to the second strand synthesis reaction in a 1:1 ratio. The mixture of magnetic beads and second strand synthesis reaction was incubated for 5 minutes on a mixer set at medium agitation and then briefly centrifuged. The tube containing the mixture of magnetic beads and second strand synthesis reaction was then placed on a magnetic stand for 3 minutes until the magnetic beads abutted against the tube wall. The supernatant was discarded without disturbing the magnetic beads. The first ethanol wash was performed by resuspending the magnetic beads conjugated to dsDNA in 80% ethanol. The tube was then again placed on the magnetic stand for 3 minutes until the magnetic beads abutted against the tube wall and the supernatant was discarded without disturbing the magnetic beads. A second ethanol wash with 80% ethanol was performed on the magnetic beads conjugated to dsDNA in a similar manner. Finally, the dsDNA was eluted from the magnetic beads by resuspending the beads in nuclease-free water with gentle stirring or agitation for 5 minutes. The tube containing the NFW and magnetic beads was briefly centrifuged and placed on a magnetic stand for 5 minutes until the beads formed a pellet against the tube wall and the aqueous solution containing the amplicons appeared clear. The aqueous solution (eluate) was then removed and transferred to a 0.2 mL PCR tube for storage, and the magnetic beads were discarded. Nanopore sequencing of purified second-strand synthesis products
[0092] Sequencing of the purified dsDNA products was performed on a GridION nanopore sequencer (Oxford Nanopore Technologies, Oxford, UK). The GridION sequencer and flow cell were prepared for sequencing by gently placing the flow cell into the designated GridION slot. The number of active pores for sequencing was verified to be 1000 or more. The flow cell was primed using a Flow Cell Priming Kit (Oxford Nanopore Technologies, Oxford, UK) to check for fluid flow without introducing any air bubbles. After keeping the flow cell undisturbed for at least 15 minutes, the SpotON sample port cover was slowly lifted to allow access to the SpotON sample port. The purified dsDNA products were then prepared for sequencing to form the sequencing template library. Briefly, 1000 ng of purified dsDNA product was fragmented to an average size of approximately 800 base pairs, end-repaired, and ligated to sequencing adapters using the Rapid Sequencing Kit (Oxford Nanopore Technologies, Oxford, UK). This sequencing template library was mixed with the flow cell loading mix, and the entire volume was added dropwise to the flow cell of the sequencer through the SpotON sample port, ensuring that no air bubbles were present. The SpotON sample port cover was slowly replaced, the flow cell priming port was closed, and sequencing was started on the nanopore sequencer using standard protocols. Sequencing was automatically terminated at the end of 48 hours. Data analysis
[0093] Once sequencing was completed, the FASTQ files were transferred to the AWS cloud infrastructure platform for bioinformatics analysis. The resulting sequence data was analyzed for the presence of modified bases within the sequenced strand of the dsDNA product. Metrics for nanopore-based sequencing included absolute current flow at a given time, duration of a given current flow measurement, duration that a current measurement indicates that the detection region does not contain a nucleobase, duration between current measurements that indicate a particular base, magnitude of change from a baseline current measurement, and magnitude of change from the previous consecutive current measurement. Results and Discussion
[0094] An exemplary method (schematic diagram shown in FIG. 2) was utilized to assay AAV particles for strand specificity. The method performed second strand synthesis of isolated ssDNA genomes of AAV particles (FIG. 3), magnetic bead-assisted purification of second strand synthesis products (FIG. 4), and nanopore sequencing of second strand synthesis products over 48 hours (FIG. 5). After sequencing, the specificity of the isolated single-stranded AAV genomes was determined by identifying the specificity of the sequenced strand (FIGS. 6A, 6B, and 6C).
[0095] In the second strand synthesis reaction, a synthetic DNA strand was synthesized that was complementary to the natural strand of the isolated ssDNA genome of AAV. Modified bases (5mC-dCTP) were incorporated into the second strand synthesis reaction, thereby methylating the synthetic DNA strand and resulting in selective labeling of the synthetic DNA strand. Thus, this reaction results in a dsDNA product in which the synthetic DNA strand is labeled. Labeling of the synthetic second strand allows for greater resolution of the positive-sense strand versus the negative-sense strand. After second strand synthesis, the labeled dsDNA product was purified from the second strand synthesis reaction using magnetic beads. The labeled dsDNA product was first conjugated to magnetic beads and then eluted from the magnetic beads after separation from the remaining second strand synthesis reaction.
[0096] Nanopore sequencing of the purified dsDNA product utilized long-read sequencing technology, which provides single reads of approximately 100 base pairs to 1000 kb or more. In some embodiments, the purified double-stranded DNA product does not need to be fragmented for sequencing, which allows reads up to the full length of the viral (e.g., AAV) genome. The long reads are particularly useful for determining where the short reads map, which aids in the assembly of metagenomic samples into larger contigs and whole genome assembly. The nanopore sequencer identifies the sequence of nucleotides within one sequenced strand of the double-stranded DNA product (as shown in FIG. 6C).
[0097] Single molecule sequencing reactions using nanopore sensors are centered on characterizing the change in current through the nanopore itself, for example, the rate at which the current fluctuation indicates that the "next" base in the template is within the detection region of the pore, as well as the magnitude of the current fluctuation, as well as all aspects of the noise measured during the fluctuation. The current flow is uniquely altered by each base as the template passes through the nanopore (Figure 5). In addition, modified nucleobases in the template are also detected by changes in the current flow through the nanopore, in contrast to their unmodified counterparts (Figure 6C). Furthermore, when a protein (e.g., a helicase or polymerase) is used to pass the template through the nanopore, the interaction of the protein with the template prior to entry into the pore can affect the kinetics of sequence detection of the template portion within the detection region of the pore. In this way, upstream modified nucleobases can affect the detection of unmodified bases within the nanopore.
[0098] The resulting sequence data was analyzed for the presence of modified bases (5mC-dCTP) in the sequenced strand of the dsDNA product. Reads can be grouped based on the similarity of the motif with the modified base. For example, reads with a methylation / modification pattern were grouped into a first group, and reads without a methylation / modification pattern were grouped into a second group. Detection of a modified base or 5mC indicates that the sequenced strand is a synthetic DNA strand of the double-stranded DNA product, while the absence of a modified base or 5mC indicates that the sequenced strand is a natural DNA strand of the double-stranded DNA product (e.g., Figures 6A, 6B, and 6C).
[0099] The specificity of the isolated single-stranded AAV genomes was determined by identifying the complementary nucleotide sequence corresponding to the complete sequence when the sequenced strand was the synthetic DNA strand of the double-stranded DNA product. Furthermore, the percentage of each of the packaged sense and antisense strands in a sample of AAV particles was determined based on detection of modified bases in the synthetic DNA strand.
[0100] The present invention is not to be limited in scope by the specific embodiments described herein. Indeed, various modifications of the invention in addition to those described herein will become apparent to those skilled in the art from the foregoing description. Such modifications are intended to be within the scope of the appended claims. *******
Claims
1. A method for assaying a sample of viral particles containing a single-stranded DNA genome for strand specificity, (a) To obtain a double-stranded DNA product by synthesizing a synthetic DNA strand complementary to the natural DNA strand of the virus particle sample, wherein the synthetic DNA strand is synthesized by incorporating modified bases, (b) Purify the synthesized double-stranded DNA product, (c) Sequence the purified double-stranded DNA product on a sequencer, wherein the sequencer identifies the sequence of nucleotides within one of the sequenced strands of the double-stranded DNA product. (d) A method comprising determining the specificity of the native DNA strand by identifying the specificity of the sequenced strand, and detecting whether the sequenced strand contains the modified base, wherein the presence of the modified base indicates that the sequenced strand is the synthetic DNA strand of the double-stranded DNA product.
2. The method according to claim 1, wherein the sample of the virus particles includes adeno-associated virus (AAV) particles.
3. The method according to claim 2, wherein the AAV particles are particles of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV-DJ, AAV-DJ / 8, AAV-Rh10, AAV-retro, AAV-PHP.B, AAV8-PHP.eB, or AAV-PHP.S.
4. The method according to claim 1, wherein the modified base comprises 5-methylcytosine, 5-hydroxymethylcytosine, or N6-methyldeoxyadenosine.
5. The method according to claim 4, wherein the modified base is 5-methylcytosine.
6. The method according to claim 4, wherein the synthetic DNA strand is methylated by the incorporation of the modified base.
7. The method according to claim 1, wherein the double-stranded DNA product is purified using magnetic beads.
8. The method according to claim 1, wherein the sequencer is a nanopore sequencer.
9. The method according to claim 8, wherein the nanopore sequencer allows either the natural DNA strand or the selectively labeled synthetic DNA strand of the purified double-stranded DNA product to enter for sequencing.
10. The method according to claim 1, wherein the natural DNA strand or the selectively labeled synthetic DNA strand is sequenced using a long-read sequencing approach.
11. The method according to claim 1, wherein the modified bases of the synthetic DNA strand are detected by a change in the flow of current passing through the nanopore sequencer.
12. The method according to any one of claims 8 to 11, wherein the synthetic DNA strand is distinguished from the natural DNA strand that enters the nanopore sequencer by detecting the modified base.
13. The method according to any one of claims 1 to 11, further comprising determining the proportion of packaged sense strands and antisense strands in a sample of virus particles based on the detection of the modified bases of the synthetic DNA strand.
14. A method for distinguishing between native DNA strands and synthetic DNA strands in a sample of adeno-associated virus (AAV) particles containing a single-stranded DNA genome, (a) To obtain a double-stranded DNA product by synthesizing a synthetic DNA strand complementary to the natural DNA strand, wherein the synthetic DNA strand contains a modified base that selectively labels the synthetic DNA strand, (b) Purifying the double-stranded DNA product, (c) Sequence the purified double-stranded DNA product on a nanopore sequencer, (d) A method comprising determining the specificity of the native DNA strand by identifying the specificity of the sequenced strand, and detecting whether the sequenced strand contains the modified base of the synthetic DNA strand during nanopore sequencing by detecting a change in the flow of current through the nanopore sequencer, wherein the presence of the modified base indicates that the sequenced strand is the synthetic DNA strand of the double-stranded DNA product.
15. The method according to claim 14, wherein the modified base comprises 5-methylcytosine, 5-hydroxymethylcytosine, or N6-methyldeoxyadenosine.
16. The method according to claim 15, wherein the modified base is 5-methylcytosine.
17. The method according to claim 15, wherein the synthetic DNA strand is methylated by the selective labeling.
18. The method according to claim 14, wherein the natural DNA strand or the selectively labeled synthetic DNA strand is sequenced using a long-read sequencing approach.
19. The method according to claim 14, wherein the double-stranded DNA product is purified using magnetic beads.
20. The method according to claim 14, wherein the AAV particles are particles of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV-DJ, AAV-DJ / 8, AAV-Rh10, AAV-retro, AAV-PHP.B, AAV8-PHP.eB, or AAV-PHP.S.
21. The method according to any one of claims 14 to 20, further comprising determining the orientation of the sequenced strand based on the detection of the modified bases of the synthetic DNA strand.
22. The method according to claim 21, further comprising determining the proportion of packaged sense strands and antisense strands in a sample of adeno-associated virus (AAV) particles based on the determination of the orientation of the sequenced strands.
23. The method according to any one of claims 1 to 11 and 14 to 22, wherein the virus particle contains an exogenous gene in the single-stranded DNA genome.
24. The method according to claim 23, wherein the foreign gene is a therapeutic gene.
25. The method according to any one of claims 1 to 11 and 14 to 22, further comprising determining the complete sequence of the sequenced strand.
26. The method according to claim 25, further comprising identifying a complementary nucleotide sequence corresponding to the complete sequence, if the sequenced strand is the synthetic DNA strand of the double-stranded DNA product.
27. The method according to any one of claims 1 to 11 and 14 to 22, wherein the sequencing identifies about 100 to about 5000 nucleotides in the sequenced chain.
28. The method according to claim 27, wherein the sequencing identifies 100 to 1000 nucleotides in the sequenced chain.
29. The method according to claim 28, wherein the sequencing identifies 100 to 500 nucleotides in the sequenced chain.
30. The method according to any one of claims 1 to 11 and 14 to 22, wherein the sequencing is performed without fragmentation of the sequenced strand.