Compositions of DNA molecules encoding amylo-alpha-1,6-glucosidase, 4-alpha-glucanotransferase, methods of making same, and methods of use thereof

JP2024517427A5Pending Publication Date: 2025-05-20アンジャリウム バイオサイエンシズ エージー
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023564174
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-04-20
Filing Date
2022-04-19
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

Current gene therapy methods using viral vectors, such as AAV, are limited by transgene size, immunogenicity, and the need for repeated administration, which is hindered by neutralizing antibodies, and there is a lack of effective treatments for glycogen storage diseases like GSDIII due to these limitations.

Method used

Development of biocompatible carriers, such as hybridosomes or lipid nanoparticles, containing DNA molecules encoding human GDE or its catalytically active fragments, which can be administered in multiple doses to treat diseases associated with decreased GDE activity, using non-viral vectors that avoid viral components.

Benefits of technology

The method achieves sustained expression of GDE proteins, reducing glycogen accumulation and associated symptoms in GSDIII patients, with improved safety and efficacy by avoiding immune responses and viral vector limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are double-stranded DNA molecules comprising an inverted repeat, an expression cassette, and one or more restriction sites for a nicking endonuclease, methods of use thereof, and methods of making thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Priority) This application claims the benefit of priority to U.S. Serial No. 63 / 177,016, filed April 20, 2021, which is incorporated herein by reference in its entirety.

[0002] (Reference to electronically submitted sequence listing) This application incorporates by reference the Sequence Listing submitted herewith as a text file entitled "14497-008-228_Sequence_Listing.txt", created on April 19, 2022, and having a size of 167,403 bytes.

[0003] (1. Field) Provided herein are double-stranded DNA molecules encoding amylo-α-1,6-glucosidase, 4-α-glucanotransferase, methods of use thereof, and methods of making same, as well as methods of treating glycogen storage disorders. [Background technology]

[0004] (2.Background) Gene therapy aims to treat or prevent disease by introducing genes into target cells. By providing a transcription cassette with an active gene product (sometimes called a transgene), gene therapy can improve clinical outcomes. This is because the gene product can result in the gain of a beneficial functional effect, the loss of a harmful functional effect, or another result, such as an oncolytic effect in patients with cancer. The delivery and expression of the corrective gene in a patient's target cells can be achieved by many methods, including non-viral delivery (e.g., via liposomes) or viral delivery, including the use of engineered viruses and viral gene delivery vectors. Among the available virus-derived vectors, also known as viral particles (e.g., recombinant retroviruses, recombinant lentiviruses, recombinant adenoviruses, etc.), the AAV system has become popular as a versatile vector in gene therapy.

[0005] However, the use of viral particles as gene delivery vectors has several significant drawbacks. One significant drawback is the reliance on viral life cycles and viral proteins to package transcription cassettes into viral particles. As a result, the use of viral vectors is limited by the size of the transgene (e.g., a protein-coding capacity of less than 150,000 Da for AAV) or the need for specific viral sequences (e.g., Rep-binding elements) that can destabilize the expression cassette to ensure efficient replication and packaging. Therefore, two or more viral particles may be required to deliver large transgenes (e.g., transgenes encoding proteins greater than 150,000 Da or transgenes longer than approximately 4.7 Kb). The use of more than one AAV construct may increase the risk of reactivation of the AAV genome.

[0006] A second drawback is that the viral particles used in gene therapy are often derived from wild-type viruses to which a portion of the population has been exposed during their lifetime. These patients are known to harbor neutralizing antibodies that can compromise the efficacy of gene therapy, as further described in Snyder, Richard O., and Philippe Moullier, "Adeno-associated virus: methods and protocols," Totowa, NJ: Humana Press, 2011. Print...For the remaining seronegative patients, the viral vector capsid is often immunogenic, preventing re-administration of viral vector therapy to these patients if the initial dose is insufficient or if the therapy becomes ineffective over time.

[0007] Thus, there remains an unmet need for non-viral-based gene therapies as alternatives to viral particles, particularly those that deliver large transgenes. There also remains an unmet need for methods to produce these capsid-free vectors in host cells without the coexistence of plasmids or DNA sequences encoding viral replication machinery (e.g., AAV Rep genes), because these viral proteins or the viral DNA sequences encoding them may contaminate the isolated DNA of capsid-free viral vectors.

[0008] Furthermore, there remains a significant unmet need for recombinant DNA vectors with improved manufacturability and / or expression properties, as well as for non-immunogenic gene delivery vectors that allow for repeated administration without loss of efficacy due to, for example, neutralizing antibodies.

[0009] Disorders associated with defective or absent amylo-α-1,6-glucosidase, 4-α-glucanotransferase (GDE) function, such as glycogen storage diseases (GSDIII types A-C), cause defects in glycogen metabolism. Specifically, the debranching activity of GDE is impaired, leading to glycogen accumulation in various tissues, with the liver being the most affected. Due to metabolic defects, patients suffer from low blood sugar (hypoglycemia), enlarged liver (hepatomegaly), excessive amounts of blood fat (hyperlipidemia), elevated blood liver enzyme levels, chronic liver disease (cirrhosis), liver failure, growth retardation, short stature, benign tumors (adenomas), hypertrophic cardiomyopathy, cardiac dysfunction, congestive heart failure, skeletal myopathy, and / or insufficient muscle tone (hypotonia). Currently, disease management is limited to dietary therapy to prevent severe ketotic hypoglycemia at a very young age. Strict dietary restrictions must be initiated as soon as possible after birth and continued for at least 15 years, if not lifelong. Furthermore, the majority of GSDIII patients experience long-term symptoms. Despite recent successes with adeno-associated virus (AAV)-based gene replacement for metabolic diseases, current limitations of AAV-mediated gene transfer, such as gene size, remain a challenge for successful gene therapy in GSDIII (Louisa Jauze et al., Human Gene Therapy; October 2019, pp. 1263-1273). Furthermore, transgene loss over time has been observed in AAV gene therapy for the liver, likely due to the pathological condition of the hepatocytes being treated.

[0010] Despite significant advances in our understanding of the molecular biology and diagnosis of GSDIII, little progress has been made in developing new treatments for this disorder. There remains a significant unmet need for long-lasting disease-modifying therapies in GSDIII. Current therapies primarily aim to maintain normoglycemia in the short term, which requires strict dietary restrictions, and nonadherence can lead to seizures and, in extreme cases, coma. Furthermore, the need to prevent long-term damage to tissues such as the liver (including severe fibrosis) and muscle remains unaddressed. There are no approved gene therapies for GSDIII, and conventional AAV-based therapies cannot accommodate large transgenes, and pre-existing antibodies prevent their use in 25% to 40% of patients. Other viral gene therapy vectors that can accommodate large transgenes pose challenges in determining dose levels, as they can only be administered once and the resulting GDE expression levels may not be high enough to be effective or may be supernormal.

[0011] Thus, there is a need in the art for techniques that allow for the expression of therapeutic GDE proteins in cells, tissues, or subjects for the treatment of GDSIII. Summary of the Invention

[0012] (3. Overview) Provided herein is a method for treating a disease associated with decreased activity of amylo-α-1,6-glucosidase, 4-α-glucanotransferase (GDE) in a human patient, comprising administering to the patient a biocompatible carrier (hybridosome) or lipid nanoparticle, wherein the hybridosome or lipid nanoparticle contains a DNA molecule comprising an expression cassette containing an introduced gene encoding a human GDE or a catalytically active fragment thereof.

[0013] Provided herein is a method for treating a disease associated with decreased activity of amylo-α-1,6-glucosidase, 4-α-glucanotransferase (GDE) in a human patient, comprising administering to the patient a DNA molecule comprising an expression cassette containing a transgene encoding a human GDE or a catalytically active fragment thereof, wherein the DNA molecule is contained within a single delivery vector.

[0014] Provided herein is a method for treating a disease associated with decreased activity of a GDE in a human patient, the method comprising the steps of: (i) administering to the patient a first dose of a DNA molecule comprising an expression cassette containing a transgene encoding a human GDE or a catalytically active fragment thereof; and (ii) administering to the patient a second dose of the DNA molecule.

[0015] In one embodiment, the first dose of the DNA molecule is administered to the patient at least 3 months, at least 4 months, at least 5 months, at least 6 months, at least 7 months, at least 8 months, at least 9 months, at least 10 months, or at least 11 months before the second dose of the DNA molecule.

[0016] In one embodiment, the first dose of the DNA molecule is administered to the patient at least 1 year, at least 2 years, at least 3 years, at least 4 years, at least 5 years, at least 10 years, at least 15 years, or at least 20 years before the second dose of the DNA molecule.

[0017] In one embodiment, said first dose of said double-stranded DNA molecule and said second dose of said DNA molecule contain the same amount of said DNA molecule.

[0018] In one embodiment, said first dose of said DNA molecule and said second dose of said DNA molecule contain different amounts of said DNA molecule.

[0019] In one embodiment, the method further comprises administering one or more additional doses of the DNA molecule.

[0020] In one embodiment, the DNA molecule is administered once a week, once every two weeks, or once a month.

[0021] In one embodiment, the DNA molecule is administered to the patient about every 6 months, about every 12 months, about every 18 months, about every 2 years, about every 3 years, about every 5 years, about every 10 years, about every 15 years, or about every 20 years.

[0022] In one embodiment, the DNA molecule is administered to the patient for the life of the patient.

[0023] In one embodiment, the patient is an adult patient.

[0024] In one embodiment, the patient is a pediatric patient.

[0025] In one embodiment, the patient is a pediatric patient when the first dose of the DNA molecule is administered.

[0026] In one embodiment, the pediatric patient is an infant.

[0027] In one embodiment, the pediatric patient is about 1 year old, about 2 years old, about 3 years old, about 4 years old, about 5 years old, about 6 years old, about 7 years old, about 8 years old, about 9 years old, about 10 years old, about 11 years old, about 12 years old, about 13 years old, about 14 years old, about 15 years old, about 16 years old, about 17 years old, or about 18 years old.

[0028] In one embodiment, the disease is glycogen storage disease (GDS) type III (GSDIII).

[0029] In one embodiment, the disease is GSDIIIa, GSDIIIb, GSDIIIc, and GSDIIId.

[0030] In one embodiment, the transgene comprises a sequence that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO:174, 175, 178, or 179.

[0031] In one embodiment, the method results in improvement of one or more of the following clinical symptoms of GSDIII: fasting intolerance, exercise intolerance, failure to thrive, myopathy, muscle weakness, and hepatomegaly.

[0032] In one embodiment, the method results in about a 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100% reduction in the number of hypoglycemic episodes per year in the patient.

[0033] In one embodiment, the method results in about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100% improvement in liver function in the patient as determined by liver function tests.

[0034] In one embodiment, the method results in about a 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100% reduction in the number of hyperlipidemic episodes per year in the patient.

[0035] In one embodiment, the method results in greater than about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 95% clinical improvement as measured by one or more of the following metabolic markers: glucose, lactate, ketones, creatine phosphokinase, uric acid, lipids, or ketones.

[0036] In one embodiment, the method results in a clinical improvement of greater than about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 95% as measured by urinary glucose tetrasaccharide (Glc4) levels in the patient.

[0037] In one embodiment, the method results in GDE protein activity that is about 1-10%, about 10-20%, about 20-30%, about 30-40%, about 40-50%, about 50-60%, about 60-70%, about 70-80%, or about 80-90% of the biological activity level of the native GDE protein.

[0038] In one embodiment, said DNA molecule is detectable in liver cells of said patient by quantitative real-time PCR.

[0039] In one embodiment, the method results in a greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% reduction in limit dextrin accumulation in a biological sample (e.g., a liver sample) from the patient.

[0040] In one embodiment, said DNA molecule is detectable in muscle tissue of said patient by quantitative real-time PCR.

[0041] In one embodiment, the method results in a greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% reduction in limit dextrin accumulation in a biological sample (e.g., a muscle sample) from the patient.

[0042] Provided herein are oligonucleotides in the 5' to 3' direction of the top strand: (a) a first inverted repeat, wherein first and second restriction sites for a nicking endonuclease are located on opposite strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a top strand 5' overhang that includes the first inverted repeat; (b) an expression cassette containing a transgene encoding a human GDE or a catalytically active fragment thereof; and (c) a second inverted repeat, wherein third and fourth restriction sites for a nicking endonuclease are located on opposite strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a top strand 3' overhang that includes the second inverted repeat. is a double-stranded DNA molecule comprising:

[0043] Provided herein are oligonucleotides in the 5' to 3' direction of the top strand: (a) a first inverted repeat, wherein first and second restriction sites for a nicking endonuclease are located on opposite strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a bottom strand 3' overhang that includes the first inverted repeat; (b) an expression cassette containing a transgene encoding a human GDE or a catalytically active fragment thereof; and (c) a second inverted repeat, wherein third and fourth restriction sites for a nicking endonuclease are located on opposite strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a bottom strand 5' overhang that includes the second inverted repeat. is a double-stranded DNA molecule comprising:

[0044] Provided herein are oligonucleotides in the 5' to 3' direction of the top strand: (a) a first inverted repeat, wherein first and second restriction sites for a nicking endonuclease are located on opposite strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a top strand 5' overhang that includes the first inverted repeat; (b) an expression cassette containing a transgene encoding a human GDE or a catalytically active fragment thereof; and (c) a second inverted repeat, wherein third and fourth restriction sites for a nicking endonuclease are located on opposite strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a bottom strand 5' overhang that includes the second inverted repeat. is a double-stranded DNA molecule comprising:

[0045] Provided herein are oligonucleotides in the 5' to 3' direction of the top strand: (a) a first inverted repeat, wherein first and second restriction sites for a nicking endonuclease are located on opposite strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a bottom strand 3' overhang that includes the first inverted repeat; (b) an expression cassette containing a transgene encoding a human GDE or a catalytically active fragment thereof; and (c) a second inverted repeat, wherein third and fourth restriction sites for a nicking endonuclease are located on opposite strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a top strand 3' overhang that includes the second inverted repeat. is a double-stranded DNA molecule comprising:

[0046] In one embodiment, the DNA molecules provided herein are isolated DNA molecules.

[0047] In one embodiment, the first, second, third, and fourth restriction sites for a nicking endonuclease of the DNA molecules provided herein are all restriction sites for the same nicking endonuclease.

[0048] In one embodiment, the first and second inverted repeats of the DNA molecules provided herein are identical.

[0049] In one embodiment, the first and / or second inverted repeat of the DNA molecule provided herein is an ITR of a parvovirus.

[0050] In one embodiment, the first and / or second inverted repeat of the DNA molecules provided herein is a modified ITR of a parvovirus.

[0051] In one embodiment, the parvovirus is dependoparvovirus, bocaparvovirus, erythroparvovirus, protoparvovirus, or tetraparvovirus.

[0052] In one embodiment, the nucleotide sequence of the modified ITR of the DNA molecule provided herein is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, or at least 99% identical to the ITR of said parvovirus.

[0053] In one embodiment, the ITRs of the DNA molecules provided herein contain viral replication-associated protein binding sequences ("RABS").

[0054] In one embodiment, the RABS comprises a Rep binding sequence.

[0055] In one embodiment, the RABS comprises an NS1 binding sequence.

[0056] In one embodiment, the ITRs of the DNA molecules provided herein do not contain RABS.

[0057] In one embodiment, the transgene comprises the sequence of SEQ ID NO: 174, 175, 178, or 179.

[0058] In one embodiment, the DNA molecule provided herein is: (a) the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the DNA molecule wherein the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat.

[0059] In one embodiment, the DNA molecule provided herein is: (a) the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the DNA molecule wherein the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat.

[0060] In some embodiments, the DNA molecules provided herein are: (a) the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the DNA molecule wherein the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat.

[0061] In some embodiments, the DNA molecules provided herein are: (a) the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the DNA molecule wherein the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat.

[0062] In one embodiment, the nick is within the inverted repeat.

[0063] In one embodiment, the nick is outside of the inverted repeat.

[0064] In one embodiment, the DNA molecule is a plasmid.

[0065] In one embodiment, the plasmid further comprises a bacterial origin of replication.

[0066] In one embodiment, the plasmid further comprises a restriction enzyme site in a region 5' to the first inverted repeat and 3' to the second inverted repeat, wherein the restriction enzyme site is not present in any of the first inverted repeat, the second inverted repeat, or the region between the first and second inverted repeats.

[0067] In one embodiment, cleavage with said restriction enzyme results in single-stranded overhangs that do not anneal at a detectable level under conditions suitable for annealing of said first inverted repeat and / or second inverted repeat.

[0068] In one embodiment, the plasmid further comprises fifth and sixth restriction sites for a nicking endonuclease in the region 5' to the first inverted repeat and 3' to the second inverted repeat, wherein the fifth and sixth restriction sites for a nicking endonuclease are: (a) on opposing chains; and (b) creating a break in the double-stranded DNA molecule such that the single-stranded overhangs of the break do not undergo detectable inter- or intramolecular annealing under conditions suitable for annealing of the first and / or second inverted repeats.

[0069] In one embodiment, the fifth and sixth nicks are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides apart.

[0070] In one embodiment, said first, second, third, fourth, fifth and sixth restriction sites for a nicking endonuclease are all target sequences for the same nicking endonuclease.

[0071] In one embodiment, the nicking endonuclease recognizing the first, second, third, and / or fourth restriction site for the nicking endonuclease is Nt. BsmAI; Nt. BtsCI; N. ALwl; N. BstNBI; N. BspD6I; Nb. Mva1269I; Nb. BsrDI; Nt. BtsI; Nt. BsaI; Nt. Bpu10I; Nt. BsmBI; Nb. BbvCI; Nt. BbvCI; or Nt. BspQI.

[0072] In one embodiment, the nicking endonuclease recognizing the fifth and sixth restriction sites for the nicking endonuclease is Nt. BsmAI; Nt. BtsCI; N. ALwl; N. BstNBI; N. BspD6I; Nb. Mva1269I; Nb. BsrDI; Nt. BtsI; Nt. BsaI; Nt. Bpu10I; Nt. BsmBI; Nb. BbvCI; Nt. BbvCI; or Nt. BspQI.

[0073] In one embodiment, the nicking endonuclease recognizing said first, second, third, and / or fourth restriction site for the nicking endonuclease is a programmable nicking endonuclease.

[0074] In one embodiment, the nicking endonuclease that recognizes said fifth and sixth restriction sites for the nicking endonuclease is a programmable nicking endonuclease.

[0075] In one embodiment, the nicking endonuclease is a Cas nuclease.

[0076] In one embodiment, the expression cassette further comprises a promoter operably linked to the transcription unit.

[0077] In one embodiment, the transcription unit comprises an open reading frame.

[0078] In one embodiment, the expression cassette further comprises a post-transcriptional regulatory element.

[0079] In one embodiment, the expression cassette further comprises polyadenylation and termination signals.

[0080] In one embodiment, the size of the expression cassette is at least 4 kb, at least 4.5 kb, at least 5 kb, at least 5.5 kb, at least 6 kb, at least 6.5 kb, at least 7 kb, at least 7.5 kb, at least 8 kb, at least 8.5 kb, at least 9 kb, at least 9.5 kb, or at least 10 kb.

[0081] Provided herein is a kit for expressing human GDE in vivo, the kit comprising 0.1 to 500 mg of a DNA molecule provided herein and a device for administering the DNA molecule.

[0082] In one embodiment, the device is a syringe needle.

[0083] Provided herein are compositions comprising one or more of the DNA molecules provided herein and a pharmaceutically acceptable carrier.

[0084] In one embodiment, the carrier comprises a transfection reagent, a nanoparticle, a hybridosome, or a liposome.

[0085] In one embodiment, the compositions provided herein are used in medical therapy.

[0086] In one embodiment, the compositions provided herein are used to prepare or produce a medicament for ameliorating, preventing, delaying the onset of, or treating a disease or disorder associated with decreased activity of GDE in a subject in need thereof. [Brief explanation of the drawings]

[0087] (4. Brief description of the drawings) [Figure 1] FIG. 1 depicts various exemplary hairpin structures and their structural elements.

[0088] [Figure 2]Figures 2A-2C show linear interaction plots showing exemplary strand conformations and intramolecular forces within the overhangs, as well as intermolecular forces between the strands, and Figure 2C shows the expected annealed structure of Figures 2A and 2B.

[0089] [Figure 3] 3A-3C depict various exemplary configurations of the hairpin and the location of various restriction sites, as well as restriction sites for type II nicking endonucleases in the primary stem of the hairpin.

[0090] [Figure 4] FIG. 4 depicts the structures of various exemplary hairpins and structural elements of human mitochondrial DNA OriL and OriL-derived ITRs.

[0091] [Figure 5] FIG. 5 depicts the structure of an exemplary aptamer and aptamer ITR hairpin.

[0092] [Figure 6] FIG. 6A shows an exemplary structure of a circular plasmid from which the DNA product for expression of the GDE proteins disclosed herein may be generated after carrying out the method steps as described in Example 1.

[0093] Figure 6B shows an exemplary structure of a hairpin-ended DNA molecule for expression of the GDE protein disclosed herein. In this embodiment, the exemplary hairpin-ended DNA comprises an expression cassette containing a PGK promoter, an open reading frame (ORF) encoding the GDE transgene, and a BGH poly(A) tail. The expression cassette is flanked by two single-stranded terminal hairpins. Figure 6C shows a visualization of the DNA product from Construct 1 after performing the method steps described in Example 1.

[0094] [Figure 7]7A shows a further exemplary structure of a plasmid from which a DNA product for expression of a GDE protein disclosed herein results after performing the method steps as described in Example 1. In this embodiment, 12 (6 pairs) restriction sites for nicking endonucleases (e.g., restriction sites for nicking endonucleases as described in Sections 5.3.4 and 5.4.2) in the region 5' to the first inverted repeat and 3' to the second inverted repeat.

[0095] Figure 7B shows an exemplary structure of a hairpin-ended DNA molecule for expression of the GDE protein disclosed herein. In this embodiment, the exemplary hairpin-ended DNA comprises an expression cassette containing a promoter, an open reading frame (ORF) encoding the GDE transgene, a WPRE regulatory element, and a poly(A) tail. The expression cassette is flanked by two single-stranded terminal hairpins. Unique restriction endonuclease recognition sites were also introduced at specific sites between each component, facilitating the introduction of new genetic components within the construct.

[0096] [Figure 8] Figures 8A and 8B show GDE protein activity in cells transfected with hairpin-ended DNA molecules encoding GDE.

[0097] [Figure 9] Figure 9A shows the glycogen content, calculated as glucose, over time in lysates of glucose-starved GSDIII patient-derived fibroblasts treated with hairpin-ended DNA molecules encoding GDE or GFP, and Figure 9B shows the glycogen content, calculated as glucose, over time in lysates of glucose-starved wild-type GDE-expressing fibroblasts treated with hairpin-ended DNA molecules encoding GDE or GFP.

[0098] [Figure 10]Figures 10A-10C show luciferase expression in dividing and non-dividing cells as described in Section 6.3. Figure 10A shows the time course of luciferase expression in non-dividing cells transfected with equimolar amounts of hairpin-ended DNA molecules encoding secreted luciferase encapsulated in LNPs or hybridosomes. Figure 10B shows luciferase expression in non-dividing cells after transfection with equimolar amounts of hairpin-ended DNA molecules and a complete circular plasmid encapsulated in hybridosomes, each encoding the same expression cassette for secreted luciferase. Figure 10C shows luciferase expression in dividing cells after transfection with equimolar amounts of hairpin-ended DNA molecules and a complete circular plasmid encapsulated in hybridosomes, each encoding the same expression cassette for secreted luciferase. Luciferase activity peaks at day 2 in dividing cells, while expression continues for 4 weeks in non-dividing cells. In non-dividing cells, in direct comparison, luciferase expression from the intact circular plasmid decreases over time.

[0099] [Figure 11] FIG. 11 shows a sequence alignment of ITRs from AAV1, highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0100] [Figure 12] FIG. 12 shows a sequence alignment of ITRs from AAV2 highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0101] [Figure 13] FIG. 13 shows a sequence alignment of ITRs from AAV3, highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0102] [Figure 14]FIG. 14 depicts a sequence alignment of the ITRs from AAV4 left, highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0103] [Figure 15] FIG. 15 depicts a sequence alignment of the ITRs from AAV4 right, highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0104] [Figure 16] FIG. 16 shows a sequence alignment of ITRs from AAV5, highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0105] [Figure 17] FIG. 17 shows a sequence alignment of ITRs from AAV7 highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites. DETAILED DESCRIPTION OF THE INVENTION

[0106] (5. Detailed Description) Provided herein are methods and compositions for treating a disease or disorder associated with decreased abundance or function of amylo-α-1,6-glucosidase, 4-α-glucanotransferase (GDE) in a subject. In some embodiments, the disease associated with decreased abundance or function of GDE is glycogen storage disease type III (GSDIII). Such compositions comprise hairpin-ended DNA molecules comprising one or more nucleic acids encoding a GDE therapeutic protein or fragments thereof. In one embodiment, a composition described herein comprises a hairpin-ended DNA molecule comprising one nucleic acid encoding a GDE therapeutic protein or fragments thereof. In one embodiment, a composition described herein comprises a hairpin-ended DNA molecule comprising two, three, four, or more nucleic acids encoding a GDE therapeutic protein or fragments thereof. Also provided herein are hairpin-ended DNA molecules for expression of the GDE proteins described herein, comprising one or more nucleic acids encoding the GDE proteins. Also provided herein are methods for producing the hairpin-ended DNA molecules described herein. Also provided herein are methods of treating GSDIII using the hairpin-end DNA and related pharmaceutical compositions provided herein. More particularly, provided herein are methods of treating GSDIII comprising administering to a subject in need thereof the hairpin-end DNA described herein.

[0107] Provided herein are methods for generating hairpin-ended DNA molecules. Also provided herein are methods for using hairpin-ended DNA molecules, such as using hairpin-ended DNA molecules in gene therapy. Various methods for generating hairpin-ended DNA molecules are further described in Section 5.2 below. Various methods for using hairpin-ended DNA molecules are described in Section 5.8 below. Hairpin-ended DNA generated by these methods is shown in Section 5.5 below, and contains hairpinned inverted repeats at both ends, each of which is further described below, and an expression cassette. In some embodiments, the hairpin-ended DNA also contains one or two nicks, as further described in Section 5.5 below. Hairpins, hairpinned inverted repeats, and hairpinned ends are described in Section 5.5, below; inverted repeats that form hairpinned ends are described in Section 5.4.1, below; nicks, nicking endonucleases, and restriction sites for nicking endonucleases are described in Sections 5.4.2 and 5.5, below; expression cassettes are described in Sections 5.4.3 and 5.5, below; and functional properties of hairpin-ended DNA molecules are described in Section 5.6, below. Thus, the present disclosure provides hairpin-ended DNA molecules, methods for making them, and methods of use therefor, with any combination or permutation of the elements provided herein.

[0108] Also provided herein are parent DNA molecules for use in methods of generating the hairpin-ended DNA molecules, the parent DNA molecules comprising two inverted repeats, each of which comprises two or more restriction sites for nicking endonucleases, as further described below, and an expression cassette. The restriction sites for the nicking endonucleases are positioned such that upon nicking and denaturation with the nicking endonucleases, a single-stranded overhang containing the inverted repeat sequence is generated, which then folds upon annealing to form a hairpin (each step as described in Section 5.2). The inverted repeat is described in Section 5.4.1 below; the nick, nicking endonucleases, and restriction sites for the nicking endonucleases are described in Section 5.4.2 below; and the expression cassette is described in Section 5.4.3 below. Thus, the present disclosure provides parent DNA molecules for use in the methods of generation with any combination or permutation of the elements provided herein.

[0109] (5.1 definition) As used herein, the term "isolated" when used in reference to a DNA molecule is intended to mean that the referenced DNA molecule is free from at least one component found in its native, natural, or synthetic environment. This term includes DNA molecules that have been removed from some or all of the other components found in its native, natural, or synthetic environment. A component of a DNA molecule's native, natural, or synthetic environment includes anything in that environment that is required, used, or otherwise plays a role in the replication and maintenance of the DNA molecule in that environment. Components of a DNA molecule's native, natural, or synthetic environment also include, for example, cells, tissue debris, organelles, proteins, peptides, amino acids, lipids, polysaccharides, nucleic acids other than the referenced DNA molecule, salts, nutrients for cell culture, and / or chemicals used in DNA synthesis. A DNA molecule of the present disclosure can be partially, completely, or substantially free from all or any other component of the native, natural, or synthetic environment from which the DNA molecule is isolated, synthetically produced, naturally produced, or recombinantly produced. Examples of isolated DNA molecules include partially pure and substantially pure DNA molecules.

[0110] As used herein, the term "delivery vehicle" refers to a substance that can be used, with or without the agent to be delivered, to administer or deliver one or more agents to cells, tissues, or subjects, particularly human subjects. A delivery vehicle can preferentially deliver an agent to a specific subset or type of cell. The selective or preferential delivery achieved by a delivery vehicle can be achieved by the properties of the vehicle or by a moiety conjugated to, bound to, or contained in the delivery vehicle. This moiety specifically or preferentially binds to a specific subset of cells. A delivery vehicle can also increase the in vivo half-life of the agent to be delivered, the efficiency of delivery of the agent compared to delivery without the delivery vehicle, and / or the bioavailability of the agent to be delivered. Non-limiting examples of delivery vehicles are hydridosomes, liposomes, lipid nanoparticles, polymersomes, mixtures of natural / synthetic lipids, membrane or lipid extracts, exosomes, virus particles, proteins or protein complexes, peptides, and / or polysaccharides.

[0111] As used herein, the term "subject" refers to a human or any non-human animal (e.g., a mouse, rat, rabbit, dog, cat, cow, pig, sheep, horse, or primate). Human includes prenatal and postnatal forms. In many embodiments, the subject is a human. A subject may also be a patient, meaning a human who visits a healthcare provider seeking diagnosis or treatment of a disease. The term "subject" is used interchangeably herein with "individual" or "patient." A subject may be suffering from a disease or disorder or may be susceptible to a disease or disorder, but may or may not exhibit symptoms of the disease or disorder. In exemplary embodiments, the subject of the present disclosure is a subject with reduced activity of amylo-α-1,6-glucosidase, 4-α-glucanotransferase (AGL) (e.g., due to decreased concentration, decreased abundance, and / or decreased function). In further exemplary embodiments, the subject is a human.

[0112] As used herein, the term "and / or" in phrases such as "A and / or B" is intended to include both A and B; A or B; A alone; and B alone. Similarly, the term "and / or" in phrases such as "A, B, and / or C" is intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A alone; B alone; and C alone.

[0113] 5.2 Hairpin-Ended DNA Molecules and Methods for Producing Hairpin-Ended DNA Molecules The methods and compositions described herein refer to compositions and methods for delivering GDE nucleic acid sequences encoding human GDE proteins to a subject in need thereof for the treatment of GSDIII.

[0114] In some embodiments, a polynucleotide molecule for expressing human amylo-α-1,6-glucosidase, 4-α-glucanotransferase (collectively or individually referred to herein as "AGL" or "GDE") or a fragment thereof having GDE activity.

[0115] In some embodiments, the hairpin-ended DNA molecules of the present disclosure can be used in methods for ameliorating, preventing, or treating one or more of GSDIIIa, GSDIIIb, GSDIIIc, and GSDIIId (collectively or individually referred to herein as "GSDIII" or "Glycogen Storage Disease Type III").

[0116] The disease or disorder treated herein (e.g., GSDIIIa, GSDIIIb, GSDIIIc, or GSDIIId) may be associated with low blood sugar (hypoglycemia), enlarged liver (hepatomegaly), excessive amounts of fat in the blood (hyperlipidemia), elevated blood liver enzyme levels, chronic liver disease (cirrhosis), liver failure, growth retardation, short stature, benign tumors (adenomas), hypertrophic cardiomyopathy, cardiac dysfunction, congestive heart failure, skeletal myopathy, and / or insufficient muscle tone (hypotonia).

[0117] As will be understood by one of ordinary skill in the art, GSDIII may be referred to by any number of aliases in the art, including, but not limited to, AGL deficiency, Cori disease, Cori's disease, debrancher deficiency, Forbes disease, glycogen debranching enzyme deficiency, GSDIII, or limit dextrinosis. Accordingly, GSDIII may be used interchangeably with any of these aliases in the specification, examples, figures, and claims.

[0118] In a further aspect, provided herein is a method for generating and preparing a hairpin-ended DNA molecule for expressing human amylo-α-1,6-glucosidase, 4-α-glucanotransferase (AGL). In one aspect, provided herein is a method for preparing a hairpin-ended DNA molecule, comprising: a. culturing a host cell containing a DNA molecule described in Section 5.4 under conditions that result in amplification of the DNA molecule; b. releasing the DNA molecule from the host cell; c. incubating the DNA molecule with one or more nicking endonucleases that recognize the four restriction sites and produce four nicks; d. denaturing, thereby generating a DNA fragment containing the expression cassette and flanked by the two single-stranded DNA overhangs; and e. allowing the single-stranded DNA overhangs to intramolecularly anneal, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d.

[0119] 5.3 Methods for generating hairpin-ended DNA molecules In one aspect, provided herein is a method for preparing a hairpin-ended DNA molecule, the method comprising: a. culturing a host cell containing a DNA molecule described in Section 5.4 under conditions that result in amplification of the DNA molecule; b. releasing the DNA molecule from the host cell; c. incubating the DNA molecule with one or more nicking endonucleases that recognize the four restriction sites and produce four nicks; d. denaturing, thereby generating a DNA fragment that contains the expression cassette and is flanked by the two single-stranded DNA overhangs; and e. intramolecularly annealing the single-stranded DNA overhangs, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d.

[0120] In another aspect, provided herein is a method for preparing hairpin-ended DNA, the method comprising: a. culturing a host cell containing the plasmid of 5.4.6 under conditions that result in amplification of the plasmid; b. releasing the plasmid from the host cell; c. incubating the DNA molecule with one or more nicking endonucleases that recognize the four restriction sites and produce four nicks; d. denaturing, thereby generating a DNA fragment containing the expression cassette and flanked by the two single-stranded DNA overhangs; e. allowing the single-stranded DNA overhangs to intramolecularly anneal, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d; f. incubating the plasmid or the fragment resulting from step d with the restriction enzyme, thereby cleaving the plasmid or the fragment of the plasmid; and g. incubating the fragment of the plasmid with an exonuclease, thereby digesting the fragment of the plasmid except for the fragment resulting from step e.

[0121] In a further aspect, provided herein is a method for preparing hairpin-ended DNA, comprising: a. 25. The method of claim 24, comprising culturing a host cell containing the plasmid of claim 24 under conditions that result in amplification of the plasmid; b. releasing the plasmid from the host cell; c. incubating the DNA molecule with one or more nicking endonucleases that recognize the first, second, third, and fourth restriction sites to produce four nicks; d. denaturing, thereby generating a DNA fragment that contains the expression cassette and is flanked by the two single-stranded DNA overhangs; e. allowing the single-stranded DNA overhangs to intramolecularly anneal, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d; f. incubating the plasmid or the fragment resulting from step d with one or more nicking endonucleases that recognize the fifth and sixth restriction sites to produce cleavages in the double-stranded DNA molecule; and g. incubating the fragment of the plasmid with an exonuclease, thereby digesting the fragment of the plasmid except for the fragment resulting from step e.

[0122] In one aspect, provided herein is a method for preparing a hairpin-ended DNA molecule, the method comprising: a. culturing a host cell containing a DNA molecule described in Section 5.4 under conditions that result in amplification of the DNA molecule; b. releasing the DNA molecule from the host cell; c. incubating the DNA molecule with one or more programmable nicking enzymes that recognize the four target sites for guide nucleic acid and produce four nicks; d. denaturing, thereby generating a DNA fragment that contains the expression cassette and is flanked by the two single-stranded DNA overhangs; and e. intramolecularly annealing the single-stranded DNA overhangs, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d.

[0123] In another aspect, provided herein is a method for preparing hairpin-ended DNA, the method comprising: a. culturing a host cell containing a plasmid described in 5.4.6 under conditions that result in amplification of the plasmid; b. releasing the plasmid from the host cell; c. incubating the DNA molecule with one or more programmable nicking enzymes that recognize the four target sites for a guide nucleic acid and produce four nicks; d. denaturing, thereby generating a DNA fragment containing the expression cassette and flanked by the two single-stranded DNA overhangs; e. allowing the single-stranded DNA overhangs to intramolecularly anneal, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d; f. incubating the plasmid or the fragment resulting from step d with the restriction enzyme, thereby cleaving the plasmid or the fragment of the plasmid; and g. incubating the fragment of the plasmid with an exonuclease, thereby digesting the fragment of the plasmid except for the fragment resulting from step e.

[0124] In a further aspect, provided herein is a method for preparing hairpin-ended DNA, comprising: a. 25. The method of claim 24, comprising culturing a host cell containing the plasmid of claim 24 under conditions that result in amplification of the plasmid; b. releasing the plasmid from the host cell; c. incubating the DNA molecule with one or more programmable nicking enzymes that recognize the first, second, third, and fourth target sites for a guide nucleic acid and result in four nicks; d. denaturing, thereby generating a DNA fragment that contains the expression cassette and is flanked by the two single-stranded DNA overhangs; e. allowing the single-stranded DNA overhangs to intramolecularly anneal, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d; f. incubating the plasmid or the fragment resulting from step d with a programmable nicking enzyme that recognizes the fifth and sixth target sites for a guide nucleic acid and result in a cleavage in the double-stranded DNA molecule; and g. incubating the fragment of the plasmid with an exonuclease, thereby digesting the fragment of the plasmid except for the fragment resulting from step e. In another embodiment, step f of this paragraph can be replaced by step f: incubating the plasmid or the fragment resulting from step d with one or more nicking endonucleases that recognize the two restriction sites and result in cuts in the double-stranded DNA molecule.

[0125] In one embodiment, a DNA molecule (described in Section 5.4) containing an expression cassette flanked by inverted repeats can be provided by culturing host cells containing the DNA molecule or plasmid and releasing the DNA molecule or plasmid from the host cells, as provided in steps a and b in the previous paragraph. Alternatively, such DNA molecules can be synthesized in a cell-free system or a combination of a cell-free system and a host cell-based system. For example, chemical synthesis of DNA fragments and plasmids of various sizes and sequences is known and widely used in the art; fragments can be chemically synthesized and then ligated or recombined in host cells by any means known in the art. In another embodiment, the DNA molecule or plasmid can be provided by in vitro replication. Various methods can be used for in vitro replication, including amplification by polymerase chain reaction (PCR). PCR methods for replicating DNA fragments or plasmids of various sizes are well known and widely used in the art, for example, as described in Michael Green and Joseph Sambrook's "Molecular Cloning: A Laboratory Manual," 4th Edition, ISBN 978-1-936113-42-2 (2012), the entire contents of which are incorporated herein by reference. In some embodiments, steps a and b can be replaced by preparing DNA molecules by chemical synthesis or PCR. In another embodiment, steps a, b, c, and d can be replaced by preparing DNA molecules by chemical synthesis.

[0126] The order of the method steps is listed in the method for exemplary purposes. In certain embodiments, the method steps are performed in the order in which they appear in the claims. In some embodiments, the method steps can be performed in an order other than the order in which they appear in the claims. Specifically, in some embodiments, the method steps for generating hairpin-ended DNA molecules can be performed in the order in which they appear in the claims, or in the order in which they are listed alphabetically in claims a through e or a through g. Alternatively, the method steps for generating hairpin-ended DNA molecules can be performed in a manner other than the order in which they appear in the claims. In one embodiment, if the host cell naturally expresses, is engineered to express, or otherwise contains one or more nicking endonucleases, step c (incubating the DNA molecule with one or more nicking endonucleases that recognize the four restriction sites and produce four nicks) can be performed before step b (releasing the plasmid from the host cell). In another embodiment, step f (incubating the plasmid or the fragment resulting from step d with the restriction enzyme, or incubating the plasmid or the fragment resulting from step d with one or more nicking endonucleases) can be performed before step d (performing denaturation and thereby generating a DNA fragment comprising the expression cassette and flanked by the two single-stranded DNA overhangs) or before step c (incubating the DNA molecule with one or more nicking endonucleases). Furthermore, one or more steps can be combined into a single step that performs all the functions of the separate steps. In one embodiment, step a (culturing the host cell) can be combined with step c (incubating the DNA molecule with one or more nicking endonucleases) if the host cell naturally expresses, has been engineered to express, or otherwise contains one or more nicking endonucleases.In another embodiment, step f (incubating the plasmid or the fragment resulting from step d with said restriction enzyme or incubating the plasmid or the fragment resulting from step d with one or more nicking endonucleases) can be combined with step c (incubating the DNA molecule with one or more nicking endonucleases) by incubating with the nicking endonucleases or restriction enzymes described in steps f and c. Thus, the present disclosure provides that the steps can be performed in various combinations and permutations according to the state of the art.

[0127] Additional steps can be added to the methods provided herein before all of the steps, after all of the steps, or between any of the steps. In one embodiment, the methods provided herein further comprise step h: repairing the nick with a ligase to generate a circular DNA. In another embodiment, step h: repairing the nick with a ligase to generate a circular DNA is performed after all of the other steps of the methods described herein.

[0128] As described further below in Sections 5.4.1 and 5.5, the hairpin formed at the end of a DNA molecule is determined by the nature of the overhang between the restriction sites for the nicking endonuclease. Thus, by designing the nature, including sequence and structural properties, of the overhang between the restriction sites for the nicking endonuclease in accordance with Sections 5.4.1 and 5.5, the present method can be used to generate one, two, or more hairpinned ends. In one embodiment, the method produces hairpin-ended DNA comprising one hairpin end. In another embodiment, the method produces hairpin-ended DNA consisting of one hairpin end. In yet another embodiment, the method produces hairpin-ended DNA comprising two hairpin ends. In a further embodiment, the method produces hairpin-ended DNA consisting of two hairpin ends.

[0129] The methods provided herein can be used to produce DNA molecules comprising an artificial sequence, a natural DNA sequence, or a sequence having both a natural DNA sequence and an artificial sequence. In one embodiment, the method produces a hairpin-end DNA molecule comprising an artificial sequence. In another embodiment, the method produces a hairpin-end DNA molecule comprising a natural sequence. In yet another embodiment, the method produces a hairpin-end DNA molecule comprising both a natural sequence and an artificial sequence. In one embodiment, the method produces a hairpin-end DNA molecule comprising a viral inverted terminal repeat (ITR). In a further embodiment, the method produces a hairpin-end DNA molecule comprising a viral genome. In some embodiments, the viral genome is an engineered viral genome comprising one or more non-viral genes in an expression cassette. In certain embodiments, the viral genome is an engineered viral genome in which one or more viral genes have been knocked out. In some specific embodiments, the viral genome is an engineered viral genome in which a replication protein (Rep) gene, a capsid (Cap) gene, or both a Rep gene and a Cap gene have been knocked out. In another embodiment, the viral genome is a parvovirus genome. In yet another embodiment, the parvovirus is dependoparvovirus, bocaparvovirus, erythroparvovirus, protoparvovirus, or tetraparvovirus.

[0130] The steps performed in the various methods provided herein are described in further detail below. Host cell and host cell culturing embodiments are described in Section 5.3.1; Denaturing DNA molecules is described in Section 5.3.3; Annealing embodiments are described in Section 5.3.5; Incubating DNA molecules with a nicking endonuclease or restriction enzyme is described in Section 5.3.4; Incubating with an exonuclease is described in Section 5.3.6; and Ligating embodiments are described in Section 5.3.7. Accordingly, the present disclosure provides methods that include permutations and combinations of various embodiments of the steps described herein.

[0131] 5.3.1 Host Cells and Culturing of the Host Cells The present disclosure provides that a variety of host cells can be cultured to amplify the DNA molecules. Host cells for use in the methods provided herein can be eukaryotic host cells, prokaryotic host cells, or any transformable organism capable of replicating or amplifying recombinant DNA molecules. In some embodiments, the host cells can be microbial host cells. In further embodiments, the host cells can be host microbial cells selected from bacteria, yeast, fungi, or any of a variety of other microbial cells applicable to replicating or amplifying DNA molecules. Bacterial host cells include Escherichia coli, Klebsiella oxytoca, Anaerobiospirillum succiniciproducens, Actinobacillus succinogenes, Mannheimia succiniciproducens, Rhizobium etli, Bacillus subtilis, Corynebacterium glutamicum, Gluconobacter oxydans, Zymomonas mobilis, and Lactococcus lactis. lactis, Lactobacillus plantarum, Streptomyces coelicolor, Clostridium acetobutylicum, Pseudomonas fluorescens, and Pseudomonas putida.Yeast or fungal host cells can be of any species selected from Saccharomyces cerevisiae, Schizosaccharomyces pombe, Kluyveromyces lactis, Kluyveromyces marxianus, Aspergillus terreus, Aspergillus niger, Pichia pastoris, Rhizopus arrhizus, Rhizobus oryzae, etc. Escherichia coli is a particularly useful host cell because it is a well-characterized microbial cell and has been widely used in molecular cloning. Other particularly useful host cells include yeast, such as Saccharomyces cerevisiae. It will be appreciated that any suitable microbial host cell known in the art can be used to amplify the DNA molecules.

[0132] Similarly, eukaryotic host cells for use in the methods provided herein can be any eukaryotic cell capable of replicating or amplifying recombinant DNA molecules known and used in the art. In some embodiments, host cells for use in the methods provided herein can be mammalian host cells. In further embodiments, host cells can be human or non-human mammalian host cells. In other embodiments, host cells can be insect host cells. Some commonly used non-human mammalian host cells include CHO, mouse myeloma cell lines (e.g., NS0, SP2 / 0), rat myeloma cell lines (e.g., YB2 / 0), and BHK. Some commonly used human host cells include HEK293 and its derivatives, HT-1080, PER.C6, and Huh-7. In some embodiments, the host cell is selected from the group consisting of HeLa, NIH3T3, Jurkat, HEK293, COS, CHO, Saos, SF9, SF21, High 5, NS0, SP2 / 0, PC12, YB2 / 0, BHK, HT-1080, PER.C6, and Huh-7.

[0133] Host cells can be cultured as each host cell is known and cultured in the art. Culture conditions and media for different host cells can vary, as known and practiced in the art. For example, bacterial or other microbial host cells can be cultured at 37°C with an agitation speed of up to 300 rpm, with or without forced aeration. Some insect host cells can generally be optimally cultured at 25-30°C, without agitation or with an agitation speed of up to 150 rpm, with or without forced aeration. Some mammalian host cells can be optimally cultured at 37°C with an agitation speed of up to 150 rpm, with or without forced aeration. Furthermore, conditions for culturing various host cells can be determined by examining the growth curves of the host cells under various conditions, as known and practiced in the art. Culture media and conditions for several widely used host cells are described in Michael Green and Joseph Sambrook, "Molecular Cloning: A Laboratory Manual," 4th Edition, ISBN 978-1-936113-42-2 (2012), which is incorporated herein by reference in its entirety.

[0134] 5.3.2 Releasing DNA Molecules from Host Cells DNA molecules can be released from host cells in various ways known and practiced in the art. For example, DNA molecules can be released by disrupting host cells physically, mechanically, enzymatically, chemically, or by a combination of physical, mechanical, enzymatic, and chemical actions. In some embodiments, DNA molecules can be released from host cells by exposing the cells to a solution of a cell lysis reagent. The cell lysis reagent includes a detergent such as Triton, SDS, Tween, NP-40, and / or CHAPS. In another embodiment, DNA molecules can be released from host cells by exposing the host cells to a difference in osmolality, for example, by exposing the host cells to a hypotonic solution. In another embodiment, DNA molecules can be released from host cells by exposing the host cells to a solution with high or low pH. In some embodiments, DNA molecules can be released from host cells by treating the host cells with an enzyme, for example, by treating the host cells with lysozyme. In some further embodiments, the DNA molecule can be released from the host cell by exposing the host cell to any combination of detergent, osmolarity pressure, high or low pH, and / or enzymes (e.g., lysozyme).

[0135] Alternatively, DNA molecules can be released from host cells by exerting physical force on the host cells. In one embodiment, DNA molecules can be released from host cells by directly applying force to the host cells, for example, using a Waring blender or polytron. A Waring blender uses high-speed rotating blades to disrupt cells, and a polytron draws tissue into a long shaft containing rotating blades. In another embodiment, DNA molecules can be released from host cells by applying shear stress or force to the host cells. Various homogenizers can be used to force host cells through a narrow space, thereby shearing the cell membrane. In some embodiments, DNA molecules can be released from host cells by liquid-based homogenization. In a specific embodiment, DNA molecules can be released from host cells using a Dounce homogenizer. In another specific embodiment, DNA molecules can be released from host cells using a Potter-Elvehjem homogenizer. In yet another specific embodiment, DNA molecules can be released from host cells using a French press. Other physical forces that can release DNA molecules from host cells include manual grinding, e.g., using a mortar and pestle, in which the host cells are often frozen, e.g., in liquid nitrogen, and then ground with a mortar and pestle, during which the tensile strength of the cellulose and other polysaccharides in the cell walls breaks the host cells.

[0136] Furthermore, DNA molecules can be released from host cells by subjecting the cells to a freeze-thaw cycle. In some embodiments, a suspension of host cells is frozen and then thawed over several such freeze-thaw cycles. In some embodiments, DNA molecules can be released from host cells by subjecting the host cells to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 freeze-thaw cycles.

[0137] The above-described methods for releasing DNA molecules from host cells are not mutually exclusive, and thus the present disclosure provides that DNA molecules can be released from host cells by any combination of the DNA release methods provided in this Section 5.3.2.

[0138] 5.3.3 Denaturing DNA Molecules DNA molecules can be denatured in a variety of ways known and practiced in the art. The process of denaturing DNA molecules can separate DNA molecules from double-stranded DNA (dsDNA) to form single-stranded DNA (ssDNA). To separate the two DNA strands, the temperature can be increased until the DNA unwinds, weakening the hydrogen bonds holding the two strands together and eventually separating them. The process of separating double-stranded DNA into single strands is known as DNA denaturation or DNA denaturing.

[0139] In some embodiments, denaturing a DNA molecule can separate the two DNA strands of one or more segments of the dsDNA molecule while maintaining other segment(s) of the DNA molecule as dsDNA. In some further embodiments, denaturing a DNA molecule can separate the dsDNA into ssDNA in a segment between a first and second restriction site for a nicking endonuclease on the top and bottom strands of the DNA while maintaining other portions of the DNA molecule (e.g., a DNA molecule described in Section 5.4) as dsDNA, thereby generating an overhang between the first and second restriction sites. In certain embodiments, denaturing a DNA molecule can separate the dsDNA into ssDNA in a segment between a third and fourth restriction site for a nicking endonuclease on the top and bottom strands of the DNA while maintaining other portions of the DNA molecule (e.g., a DNA molecule described in Section 5.4) as dsDNA, thereby generating an overhang between the third and fourth restriction sites. In another embodiment, denaturing a DNA molecule can separate dsDNA into ssDNA at a segment between the first and second restriction sites and a segment between the third and fourth restriction sites for a nicking endonuclease on the top and bottom strands of the DNA, while maintaining other portions of the DNA molecule (e.g., a DNA molecule described in Section 5.4) as dsDNA, thereby (1) degrading the DNA molecule into two daughter DNA molecules, and (2) generating one overhang between the first and second restriction sites and one overhang between the third and fourth restriction sites. In one embodiment, the overhang between the first and second restriction sites for a nicking endonuclease can be a top-strand 5' overhang. In another embodiment, the overhang between the first and second restriction sites for a nicking endonuclease can be a bottom-strand 3' overhang.In yet another embodiment, the overhang between the third and fourth restriction sites for the nicking endonuclease can be a top strand 3' overhang. In a further embodiment, the overhang between the third and fourth restriction sites for the nicking endonuclease can be a bottom strand 5' overhang. In some embodiments, denaturing the DNA molecules can separate the DNA molecules in any combination of the embodiments provided herein.

[0140] The overhangs may vary in length depending on the distance between each restriction site for the nicking endonuclease. In one embodiment, the overhangs may be identical in length and / or sequence. In another embodiment, the overhangs may differ in length and / or sequence. In some embodiments, the top strand 5' overhang is at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides.In another embodiment, the top strand 5' overhang has a length of about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, It can be about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, or more nucleotides.In certain embodiments, the bottom strand 3' overhang has a length of at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59 at least 9, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides.In further embodiments, the bottom strand 3' overhang has a length of about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59 , about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, or more nucleotides.In yet another embodiment, the top strand 3' overhang has a length of at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides.In another embodiment, the top strand 3' overhang has a length of about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, It can be about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, or more nucleotides.In some embodiments, the bottom strand 5' overhang has a length of at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides.In another embodiment, the bottom strand 5' overhang has a length of about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, It can be about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, or more nucleotides.

[0141] As is known and practiced in the art, DNA molecules can be denatured by heat, by changing the pH in the DNA molecule's environment, by increasing the salt concentration, or by any combination of these and other known means. The present disclosure provides that the DNA molecule can be denatured in the above manner by using denaturing conditions that selectively separate the dsDNA into ssDNA in the segment between the first and second restriction sites and / or the segment between the third and fourth restriction sites on the top and bottom strands of the DNA, while maintaining the remaining portions of the DNA molecule as dsDNA. Such selective separation of dsDNA into ssDNA can be achieved by controlling the denaturing conditions and / or the time the DNA molecule is exposed to the denaturing conditions. In one embodiment, the DNA molecules are denatured at a temperature of at least 70°C, at least 71°C, at least 72°C, at least 73°C, at least 74°C, at least 75°C, at least 76°C, at least 77°C, at least 78°C, at least 79°C, at least 80°C, at least 81°C, at least 82°C, at least 83°C, at least 84°C, at least 85°C, at least 86°C, at least 87°C, at least 88°C, at least 89°C, at least 90°C, at least 91°C, at least 92°C, at least 93°C, at least 94°C, or at least 95°C. In another embodiment, the DNA molecules are denatured at a temperature of about 70°C, about 71°C, about 72°C, about 73°C, about 74°C, about 75°C, about 76°C, about 77°C, about 78°C, about 79°C, about 80°C, about 81°C, about 82°C, about 83°C, about 84°C, about 85°C, about 86°C, about 87°C, about 88°C, about 89°C, about 90°C, about 91°C, about 92°C, about 93°C, about 94°C, or about 95°C. In one particular embodiment, the DNA molecules are denatured at a temperature of about 90°C.

[0142] In addition to heat denaturation, some or all of the DNA molecules provided herein can undergo a denaturation process by adding various chemical agents, such as guanidine, formamide, sodium salicylate, dimethyl sulfoxide, propylene glycol, and urea. These chemical denaturants lower the melting temperature by competing with existing nitrogen base pairs for hydrogen bond donors and acceptors, allowing isothermal denaturation. In some embodiments, chemical agents can induce denaturation at room temperature. In some specific embodiments, alkaline agents (e.g., NaOH) can be used to denature DNA by changing the pH and removing protons that contribute to hydrogen bonds. In another embodiment, chemically denaturing the DNA molecules provided herein can be a gentler treatment with respect to DNA stability than heat-induced denaturation. In another embodiment, chemically denaturing and renaturing the DNA molecules provided herein (e.g., by changing the pH) can be faster than by heating. In some embodiments, the DNA of the present disclosure can be replicated and nicked in bacteria and simultaneously denatured within the period of release from the bacteria (e.g., alkaline lysis process).

[0143] In one embodiment, the DNA molecule is denatured at a pH of at least 10, at least 10.1, at least 10.2, at least 10.3, at least 10.4, at least 10.5, at least 10.6, at least 10.7, at least 10.8, at least 10.9, at least 11, at least 11.1, at least 11.2, at least 11.3, at least 11.4, at least 11.5, at least 11.6, at least 11.7, at least 11.8, at least 11.9, at least 12, at least 12.1, at least 12.2, at least 12.3, at least 12.4, at least 12.5, at least 13, at least 13.5, or at least 14. In another embodiment, the DNA molecule is denatured at a pH of about 10, about 10.1, about 10.2, about 10.3, about 10.4, about 10.5, about 10.6, about 10.7, about 10.8, about 10.9, about 11, about 11.1, about 11.2, about 11.3, about 11.4, about 11.5, about 11.6, about 11.7, about 11.8, about 11.9, about 12, about 12.1, about 12.2, about 12.3, about 12.4, about 12.5, about 13, about 13.5, or about 14. In yet another embodiment, the DNA molecule is denatured at a salt concentration of at least 1 M, at least 1.5 M, at least 2 M, at least 2.5 M, at least 3 M, at least 3.5 M, or at least 4 M salt. In further embodiments, the DNA molecules are denatured at a salt concentration of about 1 M, about 1.5 M, about 2 M, about 2.5 M, about 3 M, about 3.5 M, or about 4 M salt. In certain embodiments, the DNA molecules are exposed to denaturing conditions for at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 minutes. In other embodiments, the DNA molecules are exposed to denaturing conditions for about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 minutes.In some embodiments, the DNA molecules can be denatured by any combination of denaturing conditions and denaturing periods provided herein.

[0144] Denaturation conditions can be determined for a method step that selectively denatures a segment between the first and second restriction sites and a segment between the third and fourth restriction sites on the top and bottom strands of a DNA molecule while maintaining the remaining portions of the DNA molecule as dsDNA. Such selective denaturation conditions can be determined depending on the properties of the DNA segment to be selectively denatured. The stability of a DNA double helix is ​​correlated with the length and percentage G / C content of the DNA segment. The present disclosure provides that the selective denaturation conditions can be determined by the sequence of the DNA segment to be selectively denatured or the sequence of the resulting overhang. For example, the temperature for selective denaturation can be roughly determined as Tm = 2°C × number of AT pairs + 4°C × number of GC pairs for the DNA sequence to be selectively denatured. Other, more rigorous calculations of Tm are also known and used in the art, such as those described in, for example, Freier SM et al., Proc Natl Acad Sci, 83, 9373-9377 (1986); Breslauer KJ et al., Proc Natl Acad Sci, 83, 3746-3750 (1986); Panjkovich, A. and Melo, F., Bioinformatics 21:711-722 (2005); Panjkovich, A. et al., Nucleic Acids Res 33:W570-W572 (2005), all of which are incorporated herein by reference in their entireties.

[0145] The overhang can comprise a variety of DNA sequences. In one embodiment, the overhang comprises an inverted repeat. In another embodiment, the overhang comprises a viral inverted repeat. In yet another embodiment, the overhang comprises or consists of any embodiment of the sequences described in Sections 5.4.1, 5.4.2, 5.4.3, and 5.5. In a further embodiment, the overhang comprises or consists of any one of the sequences described in Sections 5.4.1 and 5.5.

[0146] 5.3.4 Incubating DNA molecules with one or more nicking endonucleases or restriction enzymes The present disclosure provides one or more method steps for incubating a DNA molecule with one or more nicking endonucleases or restriction enzymes, as described in Sections 3 and 5.2. Without being bound by theory, a nicking endonuclease recognizes a restriction site within the DNA molecule and cleaves only on one strand of the dsDNA (e.g., by hydrolyzing a phosphodiester bond in one DNA strand) at a site either inside or outside the restriction site, thereby generating a nick in the dsDNA. Conversely, a restriction enzyme recognizes a restriction site and cleaves both strands of the dsDNA, thereby cleaving the DNA molecule at or near a specific restriction site.

[0147] In various embodiments of the compositions and methods provided herein, the nicking endonuclease may be methylation-dependent, methylation-sensitive, or methylation-insensitive. Various nicking endonucleases known and practiced in the art are provided herein. In some embodiments, the nicking endonuclease for the compositions and methods provided herein may be a naturally occurring nicking endonuclease that is not 5-methylcytosine-dependent, such as Nb.Bsml, Nb.BbvCI, Nb.BsrDI, Nb.Btsl, Nt.BbvCI, Nt.Alwl, Nt.CviPII, Nt.BsmAI, Nt.Alwl, and Nt.BstNBI. Nicking endonucleases for the compositions and methods provided herein can also be engineered from Type IIs restriction enzymes (e.g., Alwl, BpulOI, BbvCI, Bsal, BsmBI, BsmAI, Bsml, BspOJ, mlyl, Mval269l, and Sapl, etc.), and methods for making nicking endonucleases are described in the literature, e.g., U.S. Pat. No. 7,081,358; U.S. Pat. No. 7,011,966; U.S. Pat. No. 7,943,303; U.S. Pat. No. 7,820,424; WO201804514, all of which are incorporated by reference in their entireties.

[0148] Alternatively, instead of a nicking endonuclease, a programmable nicking enzyme can be used in the compositions and methods provided herein. Such programmable nicking enzymes include, for example, Cas9 or its functional equivalents (such as Pyrococcus furiosus Argonaute (PfAgo) or Cpfl). Cas9 contains two catalytic domains, RuvC and HNH. Inactivation of one of these domains results in a programmable nicking enzyme that can replace the nicking endonuclease for the methods and compositions provided herein. In Cas9, the RuvC domain can be inactivated by an amino acid substitution at position D10 (e.g., D10A), and the HNH domain can be inactivated by an amino acid substitution at H840 (e.g., H840A) or at positions corresponding to these amino acids in other Cas9-equivalent proteins. Such programmable nicking enzymes can be Argonaute or Type II CRISPR / Cas endonucleases that contain two components: a nicking enzyme (e.g., D10A Cas9 nicking enzyme or a variant or ortholog thereof) that cleaves the target DNA, and a guide nucleic acid, e.g., a guide DNA or RNA (gDNA or gRNA), that targets or programs the nicking enzyme to a specific site within the target DNA (see, e.g., Hsu et al., Nature Biotechnology 2013 31: 827-832, incorporated herein by reference in its entirety). Programmable nicking enzymes can also be generated by fusing a site-specific DNA-binding domain (targeting domain), such as the DNA-binding domain of a DNA-binding protein (e.g., a restriction endonuclease, a transcription factor, a zinc finger, or another domain that binds to DNA at a non-random location), to the nicking endonuclease so that it acts at a specific, non-random site.As is apparent from the foregoing, programmable cleavage by a programmable nicking enzyme results from a targeting domain within or fused to the nicking enzyme, or a guide molecule (gDNA or gRNA) that directs the nicking enzyme to a specific, non-random site, a site that can be programmed by varying the targeting domain or guide molecule. Such programmable nicking enzymes are described in the literature, e.g., US 7,081,358 and WO2010021692A, which are incorporated herein by reference in their entireties.

[0149] Suitable guide nucleic acid (e.g., gDNA or gRNA) sequences and target sites suitable for guide nucleic acids are known and widely available in the art. A guide nucleic acid (e.g., gDNA or gRNA) is a specific nucleic acid (e.g., gDNA or gRNA) sequence that recognizes a target DNA region of interest and directs a programmable nicking enzyme (e.g., Cas nuclease) thereto for editing. A guide nucleic acid (e.g., gDNA or gRNA) often consists of two parts: a targeting nucleic acid, which is a 15-20 nucleotide sequence complementary to the target DNA, and a scaffold nucleic acid, which serves as a binding scaffold for the programmable nicking enzyme (e.g., Cas nuclease). A target site suitable for a guide nucleic acid must have two elements: a sequence complementary to the targeting nucleic acid in the programmable nicking enzyme and an adjacent protospacer adjacent motif (PAM). The PAM serves as a binding signal for the programmable nicking enzyme (e.g., Cas nuclease). A variety of PAMs are known, characterized, and utilized in the art, as described, for example, in Daniel Gleditzsch et al., RNA Biol. 16(4): 504-517 (April 2019); Ryan T. Leenay et al., Mol Cell. 62(1): 137-147 (April 7, 2016), both of which are incorporated by reference in their entireties. Exemplary gRNA and gDNA sequences targeting the main stem sequence of the AAV2 ITR include those listed in Table 1. Table 1: Exemplary nicking endonucleases and their corresponding restriction sites [Table 1]

[0150] A variety of nicking endonucleases known and used in the art can be used in the methods provided herein. An exemplary list of nicking endonucleases and corresponding restriction sites for some of the nicking endonucleases provided as embodiments for nicking endonucleases for use in the methods is described in The Restriction Enzyme Database (known in the art as REBASE), available at www.rebase.neb.com / cgi-bin / azlist?nick, which is incorporated herein by reference in its entirety. In one embodiment, the nicking endonucleases that recognize the first, second, third, and / or fourth restriction sites are all directed against the same nicking endonuclease target sequence. In another embodiment, the first, second, third, and fourth restriction sites for nicking endonucleases are target sequences for two different nicking endonucleases, including all possible combinations for assigning the four sites to two different nicking endonuclease target sequences (e.g., the first restriction site for the first nicking endonuclease and the remaining restriction site for the second nicking endonuclease, the first and second restriction sites for the first nicking endonuclease and the remaining restriction site for the second nicking endonuclease, etc.). In yet another embodiment, the first, second, third, and fourth restriction sites for nicking endonucleases are target sequences for three different nicking endonucleases, including all possible combinations for assigning the four sites to three different endonuclease target sequences. In a further embodiment, the first, second, third, and fourth restriction sites for nicking endonucleases are target sequences for four different nicking endonucleases. In some embodiments, the nicking endonuclease can be any one selected from those listed in Table 2. Table 2: Exemplary nicking endonucleases and their corresponding restriction sites [Table 2]

[0151] Conditions under which various nicking endonucleases cleave one strand of dsDNA, including temperature, salt concentration, pH, buffering agents, the presence or absence of certain detergents, and incubation periods to achieve a desired percentage of nicked DNA molecules, are known for the various nicking endonucleases presented herein. These conditions are readily available from the websites or catalogs of various suppliers of nicking endonucleases, such as New England BioLabs. The present disclosure provides that the step of incubating the DNA molecules with one or more nicking endonucleases is carried out according to incubation conditions known and practiced in the art.

[0152] Various restriction enzymes known and used in the art can be used in the methods provided herein. An exemplary list of restriction enzymes and their corresponding restriction sites provided as embodiments of restriction enzymes for use in the methods is described in the New England Biolabs catalog, available at neb.com / products / restriction-endonucleases and incorporated herein by reference in its entirety. Conditions under which various restriction enzymes cleave dsDNA are known for the various restriction enzymes provided herein, including temperature, salt concentration, pH, buffering reagents, the presence or absence of certain detergents, and incubation periods to achieve a desired percentage of nicked DNA molecules. These conditions are readily available from various suppliers of restriction enzymes, such as the websites or catalogs of New England BioLabs. The present disclosure provides that the step of incubating the DNA molecules with the restriction enzyme is performed according to incubation conditions known and practiced in the art.

[0153] (5.3.5 Annealing) The annealing step in the methods provided herein is performed to selectively anneal the ssDNA overhangs intramolecularly, thereby generating hairpinned inverted repeats on one end of the DNA fragments (e.g., those in Sections 5.4 and 5.5) obtained in the denaturing step (Section 5.3.3) described above. In certain embodiments, the annealing step in the methods provided herein is performed to selectively anneal the ssDNA overhangs intramolecularly, thereby generating hairpinned inverted repeats on both ends of the DNA fragments (e.g., those in Sections 5.4 and 5.5) obtained in the denaturing step (Section 5.3.3) described above. Without being bound or otherwise limited by theory, such selective intramolecular annealing of ssDNA overhangs is achieved because intramolecular complementary sequences within the ssDNA overhangs make intramolecular annealing of the ssDNA overhangs thermodynamically and / or kinetically favorable compared to intermolecular annealing of the ssDNA overhangs.

[0154] Without being bound or otherwise limited by theory, it is recognized that certain lengths and / or sequences of overhangs can make intramolecular annealing of ssDNA overhangs thermodynamically and / or kinetically favorable compared to intermolecular annealing of the ssDNA overhangs. For example, linear interaction plots showing the intramolecular forces within the overhang and the intermolecular forces between strands and the resulting structure are shown in Figures 2A-2C. The thermodynamics and kinetics of annealing of ssDNA overhangs are determined by, among other factors, enthalpy (ΔH) and entropy (ΔS). The inventors have recognized that the entropy loss in intramolecular annealing is less than that in intramolecular annealing because the loss of freedom of motion from a free ssDNA overhang to an intramolecularly annealed overhang is less than the loss of freedom of motion from a free ssDNA overhang to an intermolecularly annealed overhang. On the other hand, because the number of complementary nucleotide pairs in an intramolecularly annealed overhang is smaller than the number of complementary nucleotide pairs in an intermolecularly annealed overhang (and therefore, fewer Watson-Crick and Hoogsteen hydrogen bonds), the enthalpy increase in intramolecular annealing may be smaller than the enthalpy increase in intramolecular annealing. The present disclosure provides that ssDNA overhangs can be designed to have a specific length, number of complementary nucleotide pairs, and percentage of GC and AT pairs such that the free energy increase (ΔG = ΔH - TΔS) for intramolecular annealing of the overhang is greater than that for intermolecular annealing, thereby making intramolecular annealing thermodynamically favorable compared to intermolecular annealing. The inventors further recognize that the reaction rate of intramolecular annealing of ssDNA overhangs may be higher than that of intermolecular annealing because nucleotides within an ssDNA overhang have a higher probability of contacting each other during molecular motion than they do with nucleotides in another ssDNA overhang.The present disclosure provides that the superior kinetics of intramolecular annealing of ssDNA overhangs can result in the formation of intramolecularly annealed overhangs in preference to intermolecularly annealed overhangs, even though intramolecular annealing is thermodynamically unfavorable compared to intermolecular annealing.

[0155] The annealing step can be carried out at a variety of temperatures that favor intramolecular annealing over intermolecular annealing. In one embodiment, the ssDNA overhangs are at least 15°C, at least 16°C, at least 17°C, at least 18°C, at least 19°C, at least 20°C, at least 21°C, at least 22°C, at least 23°C, at least 24°C, at least 25°C, at least 26°C, at least 27°C, at least 28°C, at least 29°C, at least 30°C, at least 31°C, at least 32°C, at least 33°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 43°C, at least 44°C, at least 45°C, at least 46°C, at least 47°C, at least 48°C, at least 49°C, at least 50°C, at least 51°C, at least 52°C, at least 53°C, at least 54°C, at least 55°C, at least 56°C, at least 57°C, at least 58°C, at least 59°C, at least 60°C, at least 61°C, at least 62°C, at least 63°C, at least 64°C, at least 65°C, at least 66°C, at least 67°C, at least 68°C, at least 69°C, at least 70°C, at least 71°C, at least 72°C or at least 56°C, at least 57°C, at least 58°C, at least 59°C, or at least 60°C. In another embodiment, the ssDNA overhang is at about 15°C, about 16°C, about 17°C, about 18°C, about 19°C, about 20°C, about 21°C, about 22°C, about 23°C, about 24°C, about 25°C, about 26°C, about 27°C, about 28°C, about 29°C, about 30°C, about 31°C, about 32°C, about 33°C, about 34°C, about 35°C, about 36°C, The annealing is carried out at a temperature of about 37°C, about 38°C, about 39°C, about 40°C, about 41°C, about 42°C, about 43°C, about 44°C, about 45°C, about 46°C, about 47°C, about 48°C, about 49°C, about 50°C, about 51°C, about 52°C, about 53°C, about 54°C, about 55°C, about 56°C, about 57°C, about 58°C, about 59°C, or about 60°C. In one specific embodiment, the ssDNA overhangs are annealed at a temperature of at least 25°C. In another specific embodiment, the ssDNA overhangs are annealed at a temperature of about 25°C. In yet another specific embodiment, the ssDNA overhangs are annealed at room temperature.

[0156] Furthermore, the annealing step can be carried out for various periods of time that favor intramolecular annealing over intermolecular annealing. In certain embodiments, the ssDNA overhangs are annealed for at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, or at least 40 minutes. In another embodiment, the ssDNA overhang is annealed for about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, or about 40 minutes. In a specific embodiment, the ssDNA overhang is annealed for at least 20 minutes. In another specific embodiment, the ssDNA overhang is annealed for about 20 minutes.

[0157] In some embodiments, annealing can be achieved by lowering the temperature below the calculated melting temperature of the sense and antisense sequence pair.The melting temperature depends on the specific nucleotide base content and the properties of the solution used, such as salt concentration.The melting temperature of any combination of sequence and solution can be easily calculated as known and practiced in the art.

[0158] In some embodiments, annealing can be achieved isothermally by reducing the amount of denaturing chemicals to allow interaction between the sense and antisense sequence pairs.The minimum concentration of denaturing chemicals required to denature a DNA sequence may depend on the specific nucleotide base content and the properties of the solution used, such as temperature or salt concentration.The concentration of chemical denaturants that do not cause denaturation for any combination of sequence and solution can be easily identified as known and practiced in the art.The concentration of chemical denaturants can also be easily changed as known and practiced in the art.For example, the amount of urea can be reduced by dialysis or tangential flow filtration, or the pH can be changed by adding acid or base.

[0159] The annealing temperature and annealing time for intramolecular annealing correlate with the length of the ssDNA overhang, the number of complementary nucleotide pairs, and the percentage of GC and AT pairs, as well as the sequence (arrangement of complementary nucleotide pairs) of the ssDNA overhang. In certain embodiments, the ssDNA overhangs provided in the methods provided herein comprise any number of nucleotides in length, as described in Section 5.3.3. In certain embodiments, the ssDNA overhangs provided in the methods provided herein comprise at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, or at least 50 intramolecularly complementary nucleotide pairs. In some embodiments, the ssDNA overhangs provided in the methods provided herein comprise about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, or about 50 intramolecular complementary nucleotide pairs.In some embodiments, the ssDNA overhangs provided in the methods provided herein comprise at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, or at least 90% GC pairs of intramolecular complementary nucleotide pairs. In certain embodiments, the ssDNA overhangs provided in the methods provided herein contain about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, or about 90% GC pairs among the intramolecular complementary nucleotide pairs.

[0160] Furthermore, the inventors recognize that the concentration of DNA molecules, which correlates with the concentration of overhangs, can affect the equilibrium and kinetics of intramolecular and intermolecular annealing of overhangs. Without being bound or otherwise limited by theory, if the concentration of overhangs is too high, the probability of intermolecular contacts between overhangs increases, which in turn reduces the kinetic dominance of intramolecular contacts over intermolecular contacts seen at such low concentrations.

[0161] As mentioned above, in some embodiments, intramolecular interactions may occur at a faster rate, and intermolecular interactions occur at a slower rate.In some embodiments, base pair interactions involving three or more molecules (e.g., three different strands) occur at the slowest rate.In some embodiments, the kinetic rate of intramolecular interactions to intermolecular interactions is governed by the concentration of each molecule.In some embodiments, the lower the concentration of DNA strands, the faster the kinetics of intramolecular interactions or the greater the intramolecular force.

[0162] Viewed individually, the absolute free energy of forming each complementary domain of the IR or ITR can be different, resulting in regions of the IR or ITR that can locally fold earlier as the strands transition from the denatured to the annealed state. The presence of locally folded domains (e.g., central hairpins or branched hairpins, such as in the AAV2 ITRs described elsewhere in this section (Section 5.4.1) and in Section 5.5) can reduce the amount of bases available for pairing with the other strand, thus reducing the likelihood of intermolecular annealing or hybridization and shifting the equilibrium from intermolecular annealing to intramolecular annealing or ITR formation.

[0163] Thus, the present disclosure provides that the annealing step can be carried out at various concentrations that favor intramolecular annealing over intermolecular annealing. In some embodiments, the ssDNA overhangs are 1 or less, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, 7 or less, 8 or less, 9 or less, 10 or less, 11 or less, 12 or less, 13 or less, 14 or less, 15 or less, 16 or less, 17 or less, 18 or less, 19 or less, 20 or less, 21 or less, 22 or less, 23 or less, 24 or less, 25 or less, 26 or less, 27 or less, 28 or less, 29 or less, 30 or less, 31 or less, 32 or less, 33 or less, 34 or less, 35 or less, 36 or less, 37 or less, 38 or less, 39 or less, 40 or less, 41 or less, 42 or less, 43 or less, 44 or less, 45 or less, 46 or less, 47 or less, 48 ​​or less, 49 or less, 50 or less, 55 or less, 60 or less for said DNA molecule. , 65 or less, 70 or less, 75 or less, 80 or less, 85 or less, 90 or less, 95 or less, 100 or less, 110 or less, 120 or less, 130 or less, 140 or less, 150 or less, 160 or less, 170 or less, 180 or less, 190 or less, 200 or less, 210 or less, 220 or less, 230 or less, 240 or less, 250 or less, 260 or less, 270 or less, 280 or less, 290 or less, 300 or less, 325 or less, 350 or less, 375 or less, 400 or less, 425 or less, 450 or less, 475 or less, 500 or less, 550 or less, Annealed at ng / μl concentrations of 600 or less, 650 or less, 700 or less, 750 or less, 800 or less, 850 or less, 900 or less, 950 or less, 1000 or less.In certain embodiments, the ssDNA overhang is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 55, about 60, about 65 , about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 260, about 270, about 280, about 290, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, or about 1000 ng / μl.

[0164] Similarly, the present disclosure provides that the annealing step can be carried out at various molar concentrations that favor intramolecular annealing over intermolecular annealing. In some embodiments, the ssDNA overhangs are 1 or less, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, 7 or less, 8 or less, 9 or less, 10 or less, 11 or less, 12 or less, 13 or less, 14 or less, 15 or less, 16 or less, 17 or less, 18 or less, 19 or less, 20 or less, 21 or less, 22 or less, 23 or less, 24 or less, 25 or less, 26 or less, 27 or less, 28 or less, 29 or less, 30 or less, 31 or less, 32 or less, 33 or less, 34 or less, 35 or less, 36 or less, 37 or less, 38 or less, 39 or less, 40 or less, 41 or less, 42 or less, 43 or less, 44 or less, 45 or less, 46 or less, 47 or less, 48 ​​or less, 49 or less, 50 or less, 55 or less, 60 or less, 61 or less, 62 or less, 63 or less, 64 or less, 65 or less, 66 or less, 67 or less, 68 or less, 69 or less, 70 or less, 71 or less, 72 or less, 73 or less, 74 or less, 75 or less, 76 or less, 77 or less, 78 or less, 79 or less, 80 or less, 81 or less, 82 or less, 83 or less, 84 or less, 85 or less, 86 or less, 87 or less, 88 or less, 89 or less, 90 or less, 91 or less, 92 or less, 93 Below, 65 or less, 70 or less, 75 or less, 80 or less, 85 or less, 90 or less, 95 or less, 100 or less, 110 or less, 120 or less, 130 or less, 140 or less, 150 or less Lower, 160 or less, 170 or less, 180 or less, 190 or less, 200 or less, 210 or less, 220 or less, 230 or less, 240 or less, 250 or less, 260 or less, 270 or less Lower, 280 or less, 290 or less, 300 or less, 325 or less, 350 or less, 375 or less, 400 or less, 425 or less, 450 or less, 475 or less, 500 or less, 550 or less Bottom, annealed at nM concentrations below 600, below 650, below 700, below 750, below 800, below 850, below 900, below 950, and below 1000.In certain embodiments, the ssDNA overhang is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 55, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108 about 5, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 260, about 270, about 280, about 290, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, or about 1000 nM. In some further embodiments, the ssDNA overhang is annealed at a concentration of 1 or less, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, 7 or less, 8 or less, 9 or less, 10 or less, 11 or less, 12 or less, 13 or less, 14 or less, 15 or less, 16 or less, 17 or less, 18 or less, 19 or less, or 20 or less μM. In yet another embodiment, the ssDNA overhang is annealed at a concentration of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 μM. In a specific embodiment, the ssDNA overhang is annealed at a concentration of about 10 nM for the DNA molecule. In another specific embodiment, the ssDNA overhang is annealed at a concentration of about 20 nM for the DNA molecule. In yet another specific embodiment, the ssDNA overhangs are annealed to said DNA molecule at a concentration of about 30 nM. In an even more specific embodiment, the ssDNA overhangs are annealed to said DNA molecule at a concentration of about 40 nM.In yet another specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 50 nM. In another specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 60 nM. In a specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 10 ng / μl. In another specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 20 ng / μl. In yet another specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 30 ng / μl. In an even more specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 40 ng / μl. In a specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 50 ng / μl. In another specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 60ng / μl.In yet another specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 70ng / μl.In a specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 80ng / μl.In another specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 90ng / μl.In yet another specific embodiment, the ssDNA overhang is annealed to the DNA molecule at a concentration of about 100ng / μl.

[0165] In some embodiments, the ssDNA overhang provided in the methods provided herein comprises any sequence listed in Table 3. Table 3: Sequences of ssDNA overhangs and corresponding structures after annealing [Table 3]

[0166] In some embodiments, the structure of a DNA molecule provided herein is identical after 2, 3, 4, 5, 10, or 20 cycles of denaturation / renaturation (e.g., denaturation as described in Section 5.3.3 and reannealing as described in this section (Section 5.3.5)). DNA structure can be described by an ensemble of structures at or near an energy minimum. In certain embodiments, the ensemble DNA structure is identical after 2, 3, 4, 5, 10, or 20 cycles of denaturation / renaturation. In one embodiment, the folded hairpin structure formed from an ITR or IR provided herein is identical after 2, 3, 4, 5, 10, or 20 cycles of denaturation / renaturation. In another embodiment, the ensemble structure of folded hairpins is identical after 2, 3, 4, 5, 10, or 20 cycles of denaturation / renaturation.

[0167] 5.3.6 Incubating with Exonucleases The present disclosure provides for the step of incubating with an exonuclease as described in Section 3. Exonucleases cleave nucleotides from the ends (exo) of DNA molecules. Exonucleases can cleave nucleotides in the 5' to 3' direction, the 3' to 5' direction, or in both directions. In certain embodiments, exonucleases for use in the methods provided herein cleave nucleotides without sequence specificity. In some embodiments, exonucleases for use in the methods provided herein digest DNA fragments containing ends generated by one or more nicking endonucleases that recognize and cleave the fifth and sixth restriction sites or by restriction enzymes that cleave the plasmid or fragments of the plasmid provided in Section 3.

[0168] Various exonucleases known and used in the art can be used in the methods provided herein. An exemplary list of exonucleases provided as embodiments of restriction enzymes for use in the methods is described in the New England Biolabs catalog, available at neb.com / products / dna-modifying-enzymes-and-cloning-technologies / nucleases and incorporated herein by reference in its entirety. Conditions under which various exonucleases digest DNA molecules are known for the various exonucleases provided herein, including temperature, salt concentration, pH, buffering reagents, the presence or absence of certain detergents, and incubation periods to achieve a desired percentage of digestion. These conditions are readily available from various suppliers of restriction enzymes, such as the websites or catalogs of New England BioLabs. The present disclosure provides that the step of incubating the DNA molecule with the restriction enzyme is performed according to incubation conditions known and practiced in the art.

[0169] The exonuclease incubation step selectively digests DNA molecules with one or more termini while leaving hairpin-ended DNA molecules intact. As is apparent from the descriptions in Sections 5.3.5 and 5.5, hairpin-ended DNA molecules contain zero, one, two, or more nicks. In some embodiments, the exonuclease for use in the methods provided herein can be an exonuclease that selectively digests DNA molecules with one or more termini while leaving circular ssDNA / dsDNA molecules or DNA molecules containing one or more nicks but no termini intact. In one embodiment, the exonuclease for use in the methods provided herein can be exonuclease V (RecBCD). In one embodiment, the exonuclease for use in the methods provided herein can be exonuclease VIII or truncated exonuclease VIII. Exonuclease V (RecBCD), Exonuclease VIII, and Truncated Exonuclease VIII have the selectivities described in this paragraph. Other suitable exonucleases are known and used in the art, as described on the websites and catalogs of various suppliers of exonucleases, such as New England BioLabs, and are presented herein.

[0170] In some embodiments, after exonuclease treatment, the DNA molecule of the present disclosure is substantially free of prokaryotic backbone sequences. In some embodiments, backbone refers to the plasmid sequence that is not part of the sequence encompassing the expression cassette between the two ITRs. In some embodiments, backbone refers to the vector sequence that is not part of the sequence encompassing the expression cassette between the two ITRs. In some embodiments, the isolated DNA molecule of the present disclosure is 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% free of the prokaryotic backbone sequence of the original plasmid.

[0171] (5.3.7 Repairing nicks with ligase) The present disclosure provides a step of repairing the nick using a ligase described in Section 3. DNA ligase catalyzes the joining of two ends of DNA molecules by forming one or more new covalent bonds. For example, the commonly used T4 DNA ligase catalyzes the formation of a phosphodiester bond between juxtaposed 5' phosphate and 3' hydroxyl ends in DNA. The formation of a new covalent bond catalyzed by a ligase to join two DNA molecules is called "ligation." In certain embodiments, a DNA ligase for use in the methods provided herein ligates nucleotides without sequence specificity. In some embodiments, a DNA ligase for use in the methods provided herein ligates two ends at one nick in a DNA molecule described in Section 5.5, thereby repairing the one nick. In some embodiments, a DNA ligase for use in the methods provided herein ligates each pair of two ends at two nicks in a DNA molecule described in Section 5.5, thereby repairing the two nicks. In some embodiments, the DNA ligase for use in the methods provided herein ligates each pair of two ends of every nick in a DNA molecule described in Section 5.5, thereby repairing every nick in the DNA molecule. When the DNA molecule described in Section 5.5 forms a circular DNA after every nick in the DNA molecule described in Section 5.5 is repaired. As described in Section 5.5, in some embodiments, the DNA molecule described in Section 5.5 consists of two nicks. In one embodiment, the DNA molecule described in Section 5.5 contains two nicks. In another embodiment, the DNA molecule described in Section 5.5 consists of one nick. In yet another embodiment, the DNA molecule described in Section 5.5 contains one nick.

[0172] The present disclosure provides that the step of repairing the nick with a ligase is carried out according to incubation conditions known and practiced in the art.

[0173] 5.4 DNA Molecules Used in the Method The DNA molecules provided herein can be DNA molecules in their natural environment or isolated DNA molecules. In some embodiments, the DNA molecules are DNA molecules in their natural environment. In some embodiments, the DNA molecules are isolated DNA molecules. In one embodiment, the isolated DNA molecule is at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, The DNA molecules can be at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% pure.In another embodiment, the isolated DNA molecule is about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54% , about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% pure DNA molecules. Other embodiments of the isolated DNA molecules provided herein in terms of purity are further described in Section 5.4.8, which can be combined in any suitable combination with the embodiments provided in this paragraph.

[0174] Because DNA molecules can be entirely engineered (e.g., synthetically or recombinantly produced), the DNA molecules provided herein, including those in Section 3 and this Section 5.4, may lack certain sequences or characteristics, as further described in Section 5.4.5.

[0175] (5.4.1 Backward Iteration) The ITRs or IRs provided in Section 3 and this section (Section 5.4.1) can, for example, be subjected to the method steps described in Sections 3, 5.3.3, 5.3.4, and 5.3.5 to form hairpinned ITRs in hairpin-ended DNA molecules provided in Section 5.5. Thus, in some embodiments, the ITRs or IRs provided in Section 3 and this section (Section 5.4.1) can include any of the embodiments of the IRs or ITRs provided in Section 3 and Section 5.5 and additional embodiments provided in this section (Section 5.4.1), in any combination.

[0176] "Inverted repeat" or "IR" refers to a single-stranded nucleic acid sequence that contains a palindromic region. This palindromic region includes a sequence of nucleotides and its reverse complement, i.e., a "palindromic sequence," as further described below, on the same strand. Under denatured conditions, meaning conditions in which hydrophobic stacking attractions between bases are nullified, the IR nucleic acid sequence exists in a random coil state (e.g., at high temperature, in the presence of chemicals, at high pH, ​​etc.). As conditions become more physiological, the IR can fold to form a secondary structure in which the outermost regions are held together non-covalently by base pairing. In some embodiments, the IR can be an ITR. In certain embodiments, the IR includes an ITR.

[0177] "Inverted terminal repeat," "terminal repeat," "TR," or "ITR" refers to an inverted repeat region at or near the end of a single-stranded DNA molecule, or an inverted repeat in or within a single-stranded overhang of a dsDNA molecule. An ITR can fold back on itself as a result of a palindromic sequence in the ITR. In one embodiment, an ITR is at or near one end of a ssDNA. In another embodiment, an ITR is at or near one end of a dsDNA. In yet another embodiment, two ITRs are at or near each of the two respective ends of a ssDNA. In a further embodiment, two ITRs are at or near each of the two respective ends of a dsDNA. In some embodiments, the non-ITR portion of a ssDNA or dsDNA is heterologous to the ITR. In certain embodiments, the non-ITR portion of a ssDNA or dsDNA is homologous to the ITR. Under denaturing conditions, meaning that the hydrophobic stacking attraction between bases is nullified, nucleic acid sequences containing ITRs exist in a random coil state (e.g., at high temperatures, in the presence of chemicals, at high pH, ​​etc.). In some embodiments, as conditions become more suitable for annealing, as described in Section 5.3.5, the ITRs can fold back on themselves to form a structure held together non-covalently by base pairing, while the heterologous non-ITR portion of the dsDNA molecule remains intact, or the heterologous non-ITR portion of the ssDNA molecule can hybridize to a second ssDNA molecule containing the reverse complement of the heterologous DNA molecule. The resulting complex of the two hybridized DNA strands contains three distinct regions: a first folded, single-stranded ITR is covalently linked to a double-stranded DNA region, which is in turn covalently linked to a second folded, single-stranded ITR. In certain embodiments, the ITR sequences can begin at one of the restriction sites for a nicking endonuclease described in Sections 3, 5.3.4, and 5.4.2 and end at the last base before the dsDNA.In one embodiment, in contrast to linear double-stranded DNA molecules, the ITRs located at the 5' and 3' ends of the top and bottom strands at both ends of a DNA molecule can fold inward and face each other (e.g., 3' to 5', 5' to 3', or vice versa), thus not exposing free 5' or 3' ends on either side of the nucleic acid duplex. When an ITR folds on itself, in some embodiments, the dsDNA within the folded ITR can be immediately adjacent to the dsDNA in the non-ITR portion of the DNA molecule, generating a nick adjacent to the dsDNA, or in another embodiment, the dsDNA within the folded ITR can be separated from the dsDNA in the non-ITR portion of the DNA molecule by one or more nucleotides, generating a "ssDNA gap" adjacent to the dsDNA. Two ITRs located on either side of a non-ITR DNA sequence are referred to as an "ITR pair." In some embodiments, when the ITR adopts a folded state, it becomes resistant to exonuclease digestion (eg, exonuclease V), for example, for a period of more than 1 hour at 37°C.

[0178] The interface between the terminal bases of the ITR, which has folded into its secondary structure, and the terminal bases of the DNA hybridized duplex can be further stabilized by stacking interactions (e.g., coaxial stacking) between base pairs located on either side of the nick or ssDNA gap; these interactions are sequence-dependent. In the case of nick-like structures, an equilibrium between two conformations may exist, where one conformation is very close to that of an intact double helix, in which stacking between base pairs located on either side of the nick is preserved, and the other conformation corresponds to a complete loss of stacking at the nick site, thus inducing a kink in the DNA. Nicked molecules are known to migrate somewhat slower than intact molecules of the same size during polyacrylamide and agarose gel electrophoresis. In some cases, this retardation is enhanced at higher temperatures. It is believed that the rapid equilibration between the stacked / linear and unstacked / bent conformations of the nicks directly affects the mobility of DNA molecules during gel electrophoresis, resulting in the specific retardation characteristic of nicked DNA molecules.

[0179] Without being bound by theory, it is believed that cellular proteins can recognize parallel 5' and 3' ends as double-stranded breaks and can bind to and process them, which can have deleterious effects on the fate of the DNA within the cell. Thus, ITRs can prevent premature, unwanted degradation of expression cassettes with ITRs, such as those provided in Sections 3 and 5.5 and this section (Section 5.4.1), at one or both of their two ends.

[0180] By placing first and second restriction sites for a nicking endonuclease on opposite strands and near the inverted repeat, and subsequently separating the top strand from the bottom strand of the inverted repeat, the resulting overhang can fold back on itself to form a double-stranded end containing at least one restriction site for a nicking endonuclease. In some embodiments, the folded ITR resembles the secondary structure conformation of a viral ITR. In one embodiment, the ITRs are located at both the 5'-end and 3'-end of the bottom strand (e.g., the left ITR and the right ITR). In another embodiment, the ITRs are located at both the 5'-end and 3'-end of the top strand. In yet another embodiment, one ITR is located at the 5'-end of the top strand, and the other ITR is located at the opposite end of the bottom strand (e.g., the left ITR is at the 5'-end on the top strand, and the right ITR is at the 5'-end of the bottom strand). In yet another embodiment, one ITR is located at the 3' end of the top strand and the other ITR is located at the 3' end of the bottom strand.

[0181] In some embodiments, the present disclosure provides DNA molecules comprising palindromic sequences. A "palindromic sequence" or "palindrome" is a self-complementary DNA sequence that can fold back to form a section of dsDNA in the self-complementary region under conditions that favor intramolecular annealing. In some embodiments, a palindromic sequence comprises a contiguous polynucleotide section that is identical when read in the forward direction to when read in the reverse direction on the complementary strand. In one embodiment, a palindromic sequence comprises a polynucleotide section that is identical when read in the forward direction to when read in the reverse direction on the complementary strand, interrupted by one or more non-palindromic polynucleotide sections. In another embodiment, a palindromic sequence comprises a stretch of polynucleotide that, when read in the forward direction, is 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to when read in the reverse direction on the complementary strand. In yet another embodiment, a palindromic sequence, when read in the forward direction, has a length that is 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151%, 152%, 153%, 154%, 155%, 156%, , 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, and which are interrupted by one or more non-palindromic polynucleotide segments.The ssDNA encoding one or more palindromic sequences can fold back on itself to form double-stranded base pairs with secondary structures (e.g., hairpin loops, or three-way junctions).

[0182] For example, under appropriate conditions, such as those described in Sections 5.3.3, 5.3.4, and 5.3.5, an IR or ITR provided in this section (Section 5.4.1) can fold to form a hairpin structure as described in this section (Section 5.4.1) and Section 5.5, containing a stem, a main stem, a loop, a turning point, a bulge, a bifurcation, a bifurcation loop, an internal loop, and / or any combination or permutation of the structural features described in Section 5.5.

[0183] In one embodiment, an IR or ITR for use in the methods and compositions provided herein comprises one or more palindromic sequences. In some embodiments, an IR or ITR described herein comprises a palindromic sequence or domain capable of forming a branched hairpin structure in addition to forming a main stem domain. In some embodiments, an IR or ITR comprises a palindromic sequence capable of forming any number of branched hairpins. In specific embodiments, an IR or ITR comprises a palindromic sequence capable of forming 1 to 30 branched hairpins, or any subrange of 1 to 30 branched hairpins. In some particular embodiments, an IR or ITR comprises a palindromic sequence capable of forming 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 branched hairpins. In some embodiments, the IR or ITR comprises a sequence capable of forming two branched hairpin structures resulting in a three-way junction domain (T-shaped). In some embodiments, the IR or ITR comprises a sequence capable of forming a three-branched hairpin structure resulting in a four-way junction domain (or cruciform structure). In some embodiments, the IR or ITR comprises a sequence capable of forming a non-T-shaped hairpin structure, e.g., a U-shaped hairpin structure. In some embodiments, the IR or ITR comprises a sequence capable of forming an interrupted U-shaped hairpin structure comprising a series of bulges and base pair mismatches. In some embodiments, the branched hairpins all have stems and / or loops of the same length. In some embodiments, one branched hairpin is smaller (e.g., truncated) than the other branched hairpin. Some exemplary embodiments of hairpin structures and structural elements of hairpin structures are shown in Figure 1.

[0184] "Hairpin closing base pair" refers to the first base pair following an unpaired loop sequence. Certain stem-loop sequences have a preferred closing base pair (e.g., GC in the AAV2 ITR). In one embodiment, the stem-loop sequence contains a GC pair as the closing base pair. In another embodiment, the stem-loop sequence contains a CG pair as the closing base pair.

[0185] "ITR closing base pair" refers to the first and last nucleotides that form a base pair within a folded ITR. The terminal base pair is typically the pair of nucleotides in the main stem domain that is closest to the non-ITR sequence of the DNA molecule (e.g., an expression cassette). The ITR closing base pair can be any type of base pair (e.g., CG, AT, GC, or TA). In one embodiment, the ITR closing base pair is a GC base pair. In another embodiment, the ITR closing base pair is an AT base pair. In yet another embodiment, the ITR closing base pair is a CG base pair. In a further embodiment, the ITR closing base pair is a TA base pair.

[0186] The present disclosure provides that DNA secondary structure can be computationally predicted as known and practiced in the art. DNA secondary structure can be represented in several ways: squiggle plot, graph representation, dot-bracket notation, circular plot, arc diagram, mountain plot, dot plot, etc. In a circular plot, the backbone is represented by a circle and the base pairs are represented by arcs inside the circle. In an arc diagram, the DNA backbone is depicted as a straight line and the nucleotides of each base pair are connected by arcs. Both circular and arc plots allow for identification of similarities and differences in secondary structure.

[0187] One of the many methods for predicting DNA secondary structure uses nearest-neighbor models, which minimize the total free energy associated with the DNA structure. The minimum free energy is estimated by summing the individual energy contributions from base-pair stacking, hairpins, bulges, internal loops, and multi-branch loops. The energy contributions of these elements are sequence- and length-dependent and have been experimentally determined. The separation of a sequence into stem-loops and substems can be illustrated, for example, by displaying the structure as a graph plot. In a linear interaction plot, each residue is represented on the horizontal axis, and semi-elliptical lines connect paired bases (e.g., Figures 2A and 2B).

[0188] In some embodiments, the ITRs improve the long-term survival of nucleic acid molecules in the nucleus of a cell. In some embodiments, the ITRs improve the permanent survival of nucleic acid molecules in the nucleus of a cell (e.g., throughout the lifespan of the cell). In some embodiments, the ITRs increase the stability of nucleic acid molecules in the nucleus of a cell. In some embodiments, the ITRs inhibit or prevent degradation of nucleic acid molecules in the nucleus of a cell.

[0189] In one embodiment, the IR or ITR can comprise any viral ITR, hi another embodiment, the IR or ITR can comprise a synthetic palindromic sequence capable of forming a palindromic hairpin structure that does not expose the 5' or 3' ends at the outermost vertices or turnover points of the repeats.

[0190] In some embodiments, the single-stranded ITR sequence extending from one nucleotide of the ITR closing base pair to the other nucleotide of the ITR closing base pair has a Gibbs free energy of unfolding (ΔG) in the range of −10 kcal / mol to −100 kcal / mol under physiological conditions. In one embodiment, the Gibbs free energy of unfolding (ΔG) (kcal / mol) referred to in the preceding sentence is -10 or less (meaning ≦-10, including, for example, -20, -30, etc.), -11 or less, -12 or less, -13 or less, -14 or less, -15 or less, -16 or less, -17 or less, -18 or less, -19 or less, -20 or less, -21 or less, -22 or less, -23 or less, -24 or less, -25 or less, -26 or less, -27 or less, -28 or less, -29 or less, -30 or less, -31 or less, -32 or less, -33 or less, -34 or less, -35 or less, -36 or less, -37 or less, -38 or less, -39 or less, -40 or less, -41 or less, -42 or less, -43 or less, -44 or less, -45 or less, -46 or less, -47 or less, -48 or less, 49 or less, -50 or less, -51 or less, -52 or less, -53 or less, -54 or less, -55 or less, -56 or less, -57 or less, -58 or less, -59 or less, -60 or less, -61 or less, -6 2 or less, -63 or less, -64 or less, -65 or less, -66 or less, -67 or less, -68 or less, -69 or less, -70 or less, -71 or less, -72 or less, -73 or less, -74 or less, -75 or less, -76 or less, -77 or less, -78 or less, -79 or less, -80 or less, -81 or less, -82 or less, -83 or less, -84 or less, -85 or less, -86 or less, -87 or less, -88 or less, -89 or less, -90 or less, -91 or less, -92 or less, -93 or less, -94 or less, -95 or less, -96 or less, -97 or less, -98 or less, -99 or less, or -100 or less.In another embodiment, the Gibbs free energy of unfolding (ΔG) (kcal / mol) referred to in the preceding sentence is about -10 (meaning ≦-10, including, for example, -20, -30, etc.), about -11, about -12, about -13, about -14, about -15, about -16, about -17, about -18, about -19, about -20, about -21, about -22, about -23, about -24, about -25, about -26, about -27, about -28, about -29, about -30, about -31, about -32, about -33, about -34, about -35, about -36, about -37, about -38, about -39, about -40, about -41, about -42, about -43, about -44, about -45, about -46, or about -47. , about -48, about -49, about -50, about -51, about -52, about -53, about -54, about -55, about -56, about -57, about -58, about -59, about -60, about -61, about -62, about -63, about -64, about -65, about -66, about -67, about -68, about -69, about -70, about -71, about -72, about -73, about -74, about −75, about −76, about −77, about −78, about −79, about −80, about −81, about −82, about −83, about −84, about −85, about −86, about −87, about −88, about −89, about −90, about −91, about −92, about −93, about −94, about −95, about −96, about −97, about −98, about −99, or about −100. In some embodiments, the ITR sequence extending from one nucleotide of the ITR closing base pair to the other nucleotide of the ITR closing base pair has a Gibbs free energy of unfolding (ΔG) under physiological conditions in the range of −26 kcal / mol to −95 kcal / mol. In some embodiments, the ITR sequence extending from one nucleotide of the ITR closing base pair to the other nucleotide of the ITR closing base pair contributes all of the Gibbs free energy of unfolding (ΔG) for that ITR sequence under physiological conditions.

[0191] In some embodiments, in the folded state, the single-stranded IR or ITR has an overall Watson-Crick self-complementarity of about 50% to 98%. In one embodiment, in the folded state, the single-stranded IR or ITR has an overall Watson-Crick self-complementarity of about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 100%, about 101%, about 102%, about 103%, about 104%, about 105%, about 106%, about 107%, about 108%, about 109%, about 110%, about 111%, about 112%, about 113%, about 114%, about 115%, about 116%, about 117%, about 118%, about 119%, about 120%, about 121%, about 122%, about 123%, about 124%, about 125%, about 126%, about 127%, about 128%, about 12 3%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% overall Watson-Crick self-complementarity. In another embodiment, in the folded state, the single chain IR or ITR is at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 1 have an overall Watson-Crick self-complementarity of at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.In some embodiments, in the folded state, the IRs or ITRs have an overall Watson-Crick complementarity of about 60% to 98%.

[0192] In some embodiments, the single-chain IR or ITR has a total GC content of about 60-95%. In certain embodiments, the single-chain IR or ITR has a total GC content of at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%. In another embodiment, the single-chain IR or ITR has a total GC content of about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, or about 95%. In some embodiments, the single-chain IR has a total GC content of about 60-91%.

[0193] Table 4 lists the folding free energies, GC content, percentage complementarity, and lengths of exemplary ITRs, and Table 5 lists the sequences of the ITRs in Table 4. Table 4: Folding free energy, GC content, percent complementarity, and length of exemplary ITRs. [Table 4] (Table 5: ITR sequences in Table 4) [Table 5]

[0194] The DNA molecules for use in the methods and compositions provided herein can contain IRs or ITRs of various origins. In one embodiment, the IRs or ITRs in the DNA molecule are viral ITRs. "Viral ITRs" include any viral terminal repeat or synthetic sequence containing at least one minimally essential replication origin and a region containing a palindromic hairpin structure. In one embodiment, the viral ITRs are derived from the Parvoviridae family. Viral ITRs from the Parvoviridae family contain a "minimum essential replication origin" that includes a viral replication-associated protein-binding sequence ("RABS"), which refers to a DNA sequence to which viral DNA replication-associated protein ("RAP") and its isoforms, encoded by the Parvoviridae genes Rep and NS1, can bind. In some embodiments, RABS includes a Rep-binding sequence ("RBS") (also referred to as an RBE (Rep-binding element)), which refers to a nucleotide sequence that includes both a nucleotide sequence recognized by a Rep protein (directed at replicating viral nucleic acid molecules) and a site of specific interaction between the Rep protein and the nucleotide sequence. In another embodiment, viral ITRs from the Parvoviridae family contain RABSs that include an NS1-binding element ("NSBE") to which the replication-associated viral protein NS1 can bind. In some embodiments, the viral ITRs are from the Parvoviridae family and contain terminal separation sites ("TRS") that allow the viral DNA replication-associated protein NS1 or Rep to nick the sequence in the TRS by endonucleolytic cleavage. In yet another embodiment, the viral ITRs contain at least one RBS or NSBE and at least one TRS. In the context of producing a virus or a recombinant Rep-based viral genome, the ITRs mediate replication and viral packaging. As unexpectedly discovered by the inventors and provided herein, double-stranded linear DNA vectors having ITRs similar to viral ITRs can be produced without the need for Rep or NS1 proteins, and consequently, DNA replication is independent of RABS or TRS sequences.Thus, RABS and TRS can optionally be encoded within the nucleotide sequences disclosed herein, but are not required, providing flexibility in ITR design. In one embodiment, the ITRs for the methods and compositions provided herein do not comprise a RABS. In another embodiment, the ITRs for the methods and compositions provided herein do not comprise a RBS. In another embodiment, the ITRs for the methods and compositions provided herein do not comprise a NSBE. In yet another embodiment, the ITRs for the methods and compositions provided herein do not comprise a TRS. In a further embodiment, the ITRs for the methods and compositions provided herein do not comprise either a RABS or a TRS. In a further embodiment, the ITRs for the methods and compositions provided herein comprise a RBS, a TRS, or both a RBS and a TRS. In a further embodiment, the ITRs for the methods and compositions provided herein comprise a NBSE, a TRS, or both a NBSE and a TRS.

[0195] An "ITR pair" refers to two ITRs within a single DNA molecule. In some embodiments, both ITRs in an ITR pair are derived from a wild-type viral ITR (e.g., an AAV2 ITR) that has reverse-complementary sequences throughout its entire length. An ITR can be considered a wild-type sequence even if it has one or more nucleotides that deviate from the standard, naturally occurring sequence, as long as the changes do not affect the nature of the sequence or the overall three-dimensional structure. The present disclosure provides that, in some embodiments, the insertion, deletion, or substitution of one or more nucleotides can create a restriction site for a nicking endonuclease without altering the overall three-dimensional structure of the viral ITR. In some embodiments, the deviating nucleotides represent a conservative sequence change. In certain embodiments, the sequences of the ITRs provided herein can have at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a reference sequence (e.g., as determined using BLAST with default settings) and contain restriction sites for a nicking endonuclease such that the 3D structure has the same shape in geometric space. In other embodiments, the sequences of the ITRs provided herein can have about 95%, about 96%, about 97%, about 98%, or about 99% sequence identity to a reference sequence (e.g., as determined using BLAST with default settings) and contain restriction sites for a nicking endonuclease such that the 3D structure has the same shape in geometric space.

[0196] In some embodiments, the DNA molecules for use in the methods and compositions provided herein comprise a set of wt-ITRs. In specific embodiments, the DNA molecules for use in the methods and compositions provided herein comprise a set of wt-ITRs selected from the group set forth in Table 6. Table 6 shows the AAV serotype 1 (AAV1), AAV serotype 2 (AAV2), AAV serotype 3 (AAV3), AAV serotype 4 (AAV4), AAV serotype 5 (AAV5), AAV serotype 6 (AAV6), AAV serotype 7 (AAV7), AAV serotype 8 (AAV8), AAV serotype 9 (AAV9), AAV serotype 10 (AAV10), AAV serotype 11 (AAV11), or AAV serotype 12 (AAV12); AAVrh8, AAVrhlO, AAV-DJ, and AAV-DJ8 genomes (e.g., NCBI: NC 002077; NC 001401; NC001729; NC001829; NC006152; NC 006260; NC Exemplary ITRs from the same or different serotypes, or from other parvoviruses, are shown, including ITRs from mammals (avian AAV (AAAV), bovine AAV (BAAV), canine, equine, and ovine AAV), B19 parvovirus (GenBank Accession No. NC 000883), minute virus of mice (MVM) (GenBank Accession No. NC 001510); goose: goose parvovirus (GenBank Accession No. NC 001701); snake: snake parvovirus 1 (GenBank Accession No. NC 006148). Table 6: Exemplary ITR Sequences [Table 6] TIFF2024517427000008.tif248170TIFF2024517427000009.tif57170

[0197] In some embodiments, the DNA molecules for use in the methods and compositions provided herein comprise all or part of a parvovirus genome. Parvovirus genomes are linear and 3.9 to 6.3 kb in size, with the coding region flanked by either different (heterotelomeric, e.g., HBoV) or identical (homotelomeric, e.g., AAV2) terminal repeats that can fold into hairpin-like structures. In one embodiment, the DNA molecules for use in the methods and compositions provided herein comprise two different ITRs at the two ends of the DNA molecule. In another embodiment, the DNA molecules for use in the methods and compositions provided herein comprise two identical ITRs at the two ends of the DNA molecule. In yet another embodiment, the DNA molecules for use in the methods and compositions provided herein comprise two different ITRs corresponding to the two HBoV ITRs at the two ends of the DNA molecule. In a further embodiment, the DNA molecules for use in the methods and compositions provided herein comprise two identical ITRs corresponding to the AAV2 ITRs at the two ends of the DNA molecule.

[0198] In certain embodiments, the ITRs in the DNA molecules provided herein can be AAV ITRs. In another embodiment, the ITRs can be non-AAV ITRs. In one embodiment, the ITRs in the DNA molecules provided herein can be derived from AAV ITRs or non-AAV TRs. In some specific embodiments, the ITRs can be derived from any one of the Parvoviridae family, including parvoviruses and dependoviruses (e.g., canine parvovirus, bovine parvovirus, mouse parvovirus, porcine parvovirus, and human parvovirus B-19). In another specific embodiment, the ITRs can be derived from the SV40 hairpin that serves as the origin of SV40 replication. Viruses in the Parvoviridae family consist of two subfamilies: the Parvovirinae, which infect vertebrates, and the Densovirinae, which infect invertebrates. Thus, in one embodiment, the ITRs can be derived from any one of the Parvovirinae subfamily. In another embodiment, the ITRs can be derived from any one of the Densovirinae subfamily.

[0199] Compared with the T-shaped AAV ITRs, human erythrovirus B19 has ITRs that fold into long linear duplexes with a small number of unpaired nucleotides and terminate in imperfect palindromes that can generate a series of small but highly conserved mismatch bulges. In some embodiments, the ITRs of any parvovirus can be used as ITRs (e.g., wild-type or modified ITRs) for the DNA molecules provided herein, or the ITRs of any parvovirus can serve as template ITRs for modification and subsequent incorporation into the DNA molecules provided herein. In some specific embodiments, the parvovirus from which the ITRs of the DNA molecules are derived is Dependovirus, Erythroparvovirus, or Bocaparvovirus. In another specific embodiment, the ITRs of the DNA molecules provided herein are derived from AAV, B19, or HBoV. In certain embodiments, the serotype of the AAV ITRs selected for the DNA molecules provided herein can be based on the tissue tropism of the serotype. AAV2 has broad tissue tropism, AAV1 preferentially targets neurons and skeletal muscle, and AAV5 preferentially targets neurons, retinal pigment epithelial cells, and photoreceptor cells. AAV6 preferentially targets skeletal muscle and lung. AAV8 preferentially targets liver, skeletal muscle, heart, and pancreatic tissue. AAV9 preferentially targets liver, skeletal, and lung tissue. In one embodiment, the ITRs or modified ITRs of the DNA molecules provided herein are based on the AAV2 ITRs. In one embodiment, the ITRs or modified ITRs of the DNA molecules provided herein are based on the AAV1 ITRs. In one embodiment, the ITRs or modified ITRs of the DNA molecules provided herein are based on the AAV5 ITRs. In one embodiment, the ITRs or modified ITRs of the DNA molecules provided herein are based on the AAV6 ITRs. In one embodiment, the ITRs or modified ITRs of the DNA molecules provided herein are based on the AAV8 ITRs. In one embodiment, the ITRs or modified ITRs of the DNA molecules provided herein are based on the AAV9 ITRs.

[0200] In one embodiment, the DNA molecule for use in the methods and compositions provided herein comprises one or more non-AAV ITRs. In further embodiments, such non-AAV ITRs can be derived from hairpin sequences found in mammalian genomes. In one particular embodiment, such non-AAV ITRs include the OriL hairpin sequence, which adopts a stem-loop structure and is involved in initiating DNA synthesis of mitochondrial DNA. [ka] (See Fuste et al., Molecular Cell, 37, 67-78, January 15, 2010, which is incorporated herein by reference in its entirety.) In another specific embodiment, a DNA molecule for use in the methods and compositions provided herein comprises an ITR derived from an OriL sequence that is mirrored to form a T-junction with two self-complementary palindromic regions and a 12-nucleotide loop at both vertices of the hairpin. In one embodiment, a DNA molecule for use in the methods and compositions provided herein comprises an ITR derived from an OriL sequence that maintains the OriL hairpin loop and the subsequent unpaired bulge and GC-rich stem. Some exemplary embodiments of ITRs derived from mitochondrial OriL are shown in Figure 2.

[0201] In one embodiment, the DNA molecule for use in the methods and compositions provided herein comprises one or more non-AAV ITRs derived from aptamers. Similar to viral ITRs, aptamers are composed of ssDNA that fold to form a three-dimensional structure and have the ability to recognize biological targets with high affinity and specificity. DNA aptamers can be generated by systematic evolution of ligands by exponential enrichment (SELEX). For example, some aptamers have been shown to target the nuclei of human cells (see Shen et al., ACS Sens. 2019, 4, 6, 1612-1618, incorporated herein by reference in its entirety). In one embodiment, the DNA molecule for use in the methods and compositions provided herein comprises a nuclear-targeting aptamer ITR or a derivative thereof, wherein the aptamer specifically binds to a nuclear protein. In some embodiments, the aptamer ITR folds to form a secondary structure that can include hairpins and internal loops, as well as bulges and stem regions. Some exemplary embodiments of aptamers or ITRs derived therefrom are shown in Figures 3A-3C.

[0202] In some specific embodiments, the DNA molecules for use in the methods and compositions provided herein comprise one or more AAV2 ITRs, human erythrovirus B19 ITRs, goose parvovirus ITRs, and / or derivatives thereof, in any combination. In another specific embodiment, the DNA molecules for use in the methods and compositions provided herein comprise two ITRs selected from AAV2 ITRs, human erythrovirus B19 ITRs, goose parvovirus ITRs, and derivatives thereof, in any combination. In some specific embodiments, the DNA molecules for use in the methods and compositions provided herein comprise one or more AAV2 ITRs, human erythrovirus B19 ITRs, goose parvovirus ITRs, and / or derivatives thereof, in any combination, wherein the palindromic regions of those ITRs remain functional regardless of whether the 5' and 3' ITR orientations relative to the expression cassette are direct, opposite, or any conceivable combination (as described in WO2019143885, the entire contents of which are incorporated herein by reference).

[0203] In some embodiments, the modified IR or ITR in the DNA molecules provided herein is a synthetic IR sequence that includes a restriction site for an endonuclease, such as 5'-GAGTC-3', in addition to various palindromic sequences that allow for hairpin secondary structure formation as described in this section (Section 5.4.1).

[0204] In certain embodiments, the IRs or ITRs in the DNA molecules provided herein can be IRs or ITRs that vary in sequence homology to the IR or ITR sequences described in this section (Section 5.4.1). In other embodiments, the IRs or ITRs in the DNA molecules provided herein can be IRs or ITRs that vary in sequence homology to known IR or ITR sequences of the various ITR origins described in this section (Section 5.4.1) (e.g., viral ITRs, mitochondrial ITRs, artificial or synthetic ITRs such as aptamers, etc.). In one embodiment, such homology provided in this paragraph can be at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology. In another embodiment, such homology provided in this paragraph can be about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% homology.

[0205] In some embodiments, the IRs or ITRs within the DNA molecules provided herein can comprise any one or more of the features described in this section (Section 5.4.1), in various permutations and combinations.

[0206] 5.4.2 Restriction Enzymes, Nicking Endonucleases, and Their Respective Restriction Sites; Programmable Nicking Enzymes and Their Target Sites Various embodiments of nicking endonucleases, restriction enzymes, and / or restriction sites thereof, as described in Section 5.3.4, are provided for the DNA molecules provided herein. In some embodiments, the first, second, third, and fourth restriction sites for nicking endonucleases provided for a DNA molecule as described in Section 3 and this section (Section 5.4) can all be target sequences for the same nicking endonuclease. In some embodiments, the first, second, third, and fourth restriction sites for nicking endonucleases provided for a DNA molecule as described in Section 3 and this section (Section 5.4) can be target sequences for four different nicking endonucleases. In another embodiment, the first, second, third, and fourth restriction sites for nicking endonucleases are target sequences for two different nicking endonucleases, including all possible combinations for assigning the four sites to two different nicking endonuclease target sequences (e.g., the first restriction site for the first nicking endonuclease and the remaining restriction site for the second nicking endonuclease, the first and second restriction sites for the first nicking endonuclease and the remaining restriction site for the second nicking endonuclease, etc.). In one embodiment, the first, second, third, and fourth restriction sites for nicking endonucleases are target sequences for three different nicking endonucleases, including all possible combinations for assigning the four sites to three different nicking endonuclease target sequences. In some embodiments, the nicking endonuclease and the restriction site for the nicking endonuclease can be any one selected from those described in Section 5.3.4, including Table 2. In further embodiments, each of the first, second, third, and fourth restriction sites for a nicking endonuclease can be a site for any nicking endonuclease selected from those described in Section 5.3.4, including Table 2.

[0207] Tables 7-16 show exemplary modified AAV ITR sequences containing two antiparallel recognition sites for the same nicking endonuclease, grouped by nicking endonuclease species. The modified ITR sequences and corresponding alignments of wild-type AAV1, AAV2, AAV3, AAV4 left, AAV4 right, AAV5, and AAV7 are shown in Figures 11-17. Table 7: Exemplary AAV-derived ITRs with internal antiparallel recognition sites for the nicking endonuclease Nb.BvCI [Table 7] Table 8: ITRs from exemplary AAVs containing antiparallel recognition sites for the nicking endonuclease Nb.BsmI [Table 8] Table 9: ITRs from exemplary AAVs containing antiparallel recognition sites for the nicking endonuclease Nb.BsrDI [Table 9] Table 10: ITRs from exemplary AAVs containing antiparallel recognition sites for the nicking endonuclease Nb.BssSi [Table 10] Table 11: ITRs from exemplary AAVs containing antiparallel recognition sites for the nicking endonuclease Nb.BtsI [Table 11] Table 12: ITRs from exemplary AAVs with internal antiparallel recognition sites for the nicking endonuclease Nt.AlwI [Table 12] TIFF2024517427000022.tif229170 (Table 13: ITRs from exemplary AAVs containing antiparallel recognition sites for the nicking endonuclease Nt.BbvCI) [Table 13] Table 14: ITRs from exemplary AAVs containing antiparallel recognition sites for the nicking endonuclease Nt.BsmAI [Table 14] TIFF2024517427000026.tif158170 (Table 15: Exemplary AAV-derived ITRs containing antiparallel recognition sites for the nicking endonuclease Nt.BspQI) [Table 15] Table 16: ITRs from exemplary AAVs containing antiparallel recognition sites for the nicking endonuclease Nt.BstNBI [Table 16] TIFF2024517427000030.tif114170 (Table 17: Reverse complement of nicking enzyme target) [Table 17] TIFF2024517427000032.tif242170TIFF2024517427000033.tif244170TIFF2024517427000034.tif245170TIFF2024517427000035.tif244170TIFF2024517427000036.tif244170TIFF2024517427000037.tif242170TIFF2024517427000038.tif248170TIFF2024517427000039.tif249170TIFF2024517427000040.tif241170TIFF2024517427000041.tif26170

[0208] The first, second, third and fourth restriction sites for the nicking endonuclease can be arranged in a variety of configurations.In some embodiments, the first and second restriction sites for the nicking endonuclease are at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64 , at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, At least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, or at least 200 nucleotides apart.In another embodiment, the first and second restriction sites for the nicking endonuclease are about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64 , about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, or about 200 nucleotides apart.

[0209] Similarly, in certain embodiments, the third and fourth restriction sites for nicking endonucleases are at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64 , at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, At least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, or at least 200 nucleotides apart.In further embodiments, the third and fourth restriction sites for the nicking endonuclease are about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about 103, about 104, about 105, about 10 4, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, or about 200 nucleotides apart.

[0210] The present disclosure provides that the overhangs described in Sections 3, 5.2 (including 5.3.3), and 5.4 (including 5.4.1) can be the result of nicking at the first and second restriction sites by a nicking endonuclease and the denaturation described in Sections 3 and 5.2 (including 5.3.3). Thus, in some embodiments, the overhangs resulting from nicking at the first and second restriction sites can be the same length (in number of nucleotides) that separates the first and second restriction sites as described in the previous paragraph of this section (Section 5.4.2). Because a nicking endonuclease can cleave DNA inside or outside the restriction site for the nicking endonuclease, in certain embodiments, the overhang resulting from nicking at the first and second restriction sites can be at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 nucleotides longer or shorter than the length separating the first and second restriction sites. In another embodiment, the overhang resulting from nicking at the first and second restriction sites can be about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides longer or shorter than the length separating the first and second restriction sites.

[0211] Similarly, the present disclosure provides that the overhangs described in Sections 3, 5.2 (including 5.3.3), and 5.4 (including 5.4.1) can be the result of nicking at the third and fourth restriction sites by a nicking endonuclease and the denaturation described in Sections 3 and 5.2 (including 5.3.3). Thus, in some embodiments, the overhangs resulting from nicking at the third and fourth restriction sites can be the same length (in number of nucleotides) as the separation of the third and fourth restriction sites as described in the previous paragraph of this section (Section 5.4.2). Because a nicking endonuclease can cleave DNA inside or outside the restriction site for the nicking endonuclease, in certain embodiments, the overhang resulting from nicking at the third and fourth restriction sites can be at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 nucleotides longer or shorter than the length separating the third and fourth restriction sites. In another embodiment, the overhang resulting from nicking at the third and fourth restriction sites can be about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides longer or shorter than the length separating the third and fourth restriction sites.

[0212] As is apparent from the descriptions in Sections 3 and 5.5 and this section (Section 5.4), the DNA molecules provided herein comprise an expression cassette. In some embodiments, the expression cassette is located between a first and a second restriction site for a nicking endonuclease(s) at one end and a third and a fourth restriction site for a nicking endonuclease(s) at the other end. In another embodiment, the expression cassette is located within a dsDNA segment of a DNA molecule produced by performing steps a through d of the methods described in Sections 3 and 5.2, including the denaturation step that provides two DNA overhangs as described in Section 5.3.3. In certain embodiments, the first, second, third, and fourth restriction sites for a nicking endonuclease are positioned such that the length of the dsDNA segment described in this paragraph is at least 0.2 kb, at least 0.3 kb, at least 0.4 kb, at least 0.5 kb, at least 0.6 kb, at least 0.7 kb, at least 0.8 kb, at least 0.9 kb, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 2.5 kb, at least 3 kb, at least 3.5 kb, at least 4 kb, at least 4.5 kb, at least 5 kb, at least 5.5 kb, at least 6 kb, at least 6.5 kb, at least 7 kb, at least 7.5 kb, at least 8 kb, at least 8.5 kb, at least 9 kb, at least 9.5 kb, or at least 10 kb. In another embodiment, the first, second, third, and fourth restriction sites for a nicking endonuclease are positioned such that the length of the dsDNA segment described in this paragraph is about 0.2 kb, about 0.3 kb, about 0.4 kb, about 0.5 kb, about 0.6 kb, about 0.7 kb, about 0.8 kb, about 0.9 kb, about 1 kb, about 1.5 kb, about 2 kb, about 2.5 kb, about 3 kb, about 3.5 kb, about 4 kb, about 4.5 kb, about 5 kb, about 5.5 kb, about 6 kb, about 6.5 kb, about 7 kb, about 7.5 kb, about 8 kb, about 8.5 kb, about 9 kb, about 9.5 kb, or about 10 kb.

[0213] As described in Section 5.3.4, incubation with a nicking endonuclease results in a first nick corresponding to a first restriction site for the nicking endonuclease, a second nick corresponding to a second restriction site for the nicking endonuclease, a third nick corresponding to a third restriction site for the nicking endonuclease, and / or a fourth nick corresponding to a fourth restriction site for the nicking endonuclease. The present disclosure provides that the first, second, third, and / or fourth nick can be at various positions relative to the inverted repeat. In one embodiment, the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat. In another embodiment, the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat. In yet another embodiment, the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat.In further embodiments, the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat. In one embodiment, the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat. In another embodiment, the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat. In yet another embodiment, the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat.In further embodiments, the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides from the 3' nucleotide of the ITR-closing base pair of the second inverted repeat. In some embodiments, any or any combination of the first, second, third, and fourth nicks is inside the inverted repeat. In certain embodiments, any or any combination of the first, second, third, and fourth nicks is outside the inverted repeat. In some additional embodiments, the first, second, third, and fourth nicks can have any relative positions between themselves, between any of them and the inverted repeat, and / or between any of them and the expression cassette described in this section (Section 5.4.2), in any combination or permutation. In some further embodiments, the first, second, third, and fourth restriction sites for nicking endonucleases can have any relative positions between themselves, between any of them and the inverted repeat, and / or between any of them and the expression cassette described in this section (Section 5.4.2), in any combination or permutation.

[0214] 5.4.3 Expression cassettes encoding GDEs The DNA molecules provided herein can comprise an expression cassette (see also Sections 3, 5.4, and 5.5). An "expression cassette" is a nucleic acid molecule or portion of a nucleic acid molecule that contains sequences or other information that direct a cellular apparatus to produce RNA or protein. In some embodiments, an expression cassette comprises a promoter sequence. In certain embodiments, an expression cassette comprises a transcription unit. In some further embodiments, an expression cassette comprises a promoter operably linked to a transcription unit. In one embodiment, a transcription unit comprises an open reading frame (ORF). Embodiments of ORFs for use with the methods and compositions provided herein are further described in the final paragraph of this section (Section 5.4.3). An expression cassette can further comprise features that direct a cellular apparatus to produce RNA and protein. In one embodiment, an expression cassette comprises post-transcriptional regulatory elements. In another embodiment, the expression cassette further comprises a polyadenylation and / or termination signal. In yet other embodiments, the expression cassette comprises regulatory elements known and used in the art to modulate (enhance, inhibit, and / or turn on / off expression of an ORF). Such regulatory elements include, for example, the 5' untranslated region (UTR), the 3'-UTR, or both the 5' and 3' UTRs. In some further embodiments, the expression cassette comprises any one or more features provided in this section (Section 5.4.3), in any combination or permutation.

[0215] An expression cassette can contain a protein-coding sequence in its ORF (sense strand). Alternatively, the expression cassette can contain a complementary sequence (antisense strand) of the ORF encoding the protein, as well as regulatory elements and / or other signals that direct the cellular machinery to produce the sense strand DNA / RNA and the corresponding protein. In some embodiments, the expression cassette contains a protein sequence without introns. In other embodiments, the expression cassette contains a protein sequence with introns that are removed upon transcription and splicing. An expression cassette can also contain a variable number of ORFs or transcription units. In one embodiment, the expression cassette contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 ORFs. In another embodiment, the expression cassette comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 transcription units.

[0216] The human AGL gene encodes a 1532 amino acid protein (SEQ ID NO: 1; accession number P35573) with a molecular weight of approximately 174.8 kDa. The AGL gene is located on chromosome 1 at site 1p21.2. AGL is a multifunctional enzyme that functions in glycogen degradation as a 1,4-α-D-glucan: 1,4-α-D-glucan-4-α-D-glycosyltransferase and amylo-1,6-glucosidase, and is also referred to as glycogen debranching enzyme (GDE), glycogen debranching enzyme, amylo-α-1,6-glucosidase, 4-α-glucanotransferase, EC: 2.4.1.25, EC: 3.2.1.33. The consensus human AGL coding sequence can be found under NCBI accession number NM_000028.2 and is SEQ ID NO: 1.

[0217] Those skilled in the art will understand that GDE therapeutic proteins include all splice variants and orthologs of GDE proteins. Essentially, any version of a GDE therapeutic protein or fragment thereof (e.g., functional fragment) can be encoded by and expressed from the DNA vectors described herein. GDE therapeutic proteins include intact molecules and fragments thereof (e.g., functional fragments). In some embodiments, a GDE therapeutic protein can be a functional truncated version as outlined in WO2020030661A1.

[0218] In some embodiments, the hairpin DNA molecule for expressing GDE protein provides advantages over traditional AAV vectors because there is no size restriction on the heterologous nucleic acid sequence encoding desired protein.Therefore, even full-length GDE4599nt protein can be expressed from a single DNA vector.Therefore, the DNA vector described herein can be used to express therapeutic GDE protein in subjects who need it, such as subjects with glycogen storage disease. Table 18: Exemplary transgenes [Table 18] TIFF2024517427000043.tif248170TIFF2024517427000044.tif248170TIFF202 4517427000045.tif249170TIFF2024517427000046.tif248170TIFF20245174270 00047.tif248170TIFF2024517427000048.tif249170TIFF2024517427000049.t if248170TIFF2024517427000050.tif248170TIFF2024517427000051.tif115170

[0219] In one aspect, a codon-optimized engineered nucleic acid sequence encoding a human GDE is provided. In one embodiment, an engineered human GDE cDNA is provided herein (as SEQ ID NO: 175), which was designed to maximize translation compared to the native GDE sequence (SEQ ID NO: 174). Preferably, the codon-optimized GDE coding sequence has less than about 80% identity, preferably about 75% or less identity, to the full-length native GDE coding sequence (SEQ ID NO: 174). In one embodiment, the codon-optimized GDE coding sequence has about 75% identity to the native GDE coding sequence of SEQ ID NO: 174. In one embodiment, the codon-optimized GDE coding sequence is characterized by an improved translation rate after delivery compared to the native GDE. In one embodiment, the codon-optimized GDE coding sequence shares less than about 99%, less than 98%, less than 97%, less than 96%, less than 95%, less than 94%, less than 93%, less than 92%, less than 91%, less than 90%, less than 89%, less than 88%, less than 87%, less than 86%, less than 85%, less than 84%, less than 83%, less than 82%, less than 81%, less than 80%, less than 79%, less than 78%, less than 77%, less than 76%, less than 75%, less than 74%, less than 73%, less than 72%, less than 71%, less than 70%, less than 69%, less than 68%, less than 67%, less than 66%, less than 65%, less than 64%, less than 63%, less than 62%, less than 61% or less identity to the full-length native GDE coding sequence of SEQ ID NO:174. In one embodiment, the codon-optimized nucleic acid sequence is a variant of SEQ ID NO: 175. In another embodiment, the codon-optimized nucleic acid sequence is a sequence that shares about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61% or more identity to SEQ ID NO: 175. In one embodiment, the codon-optimized nucleic acid sequence is SEQ ID NO: 175. In another embodiment, the nucleic acid sequence is codon-optimized for expression in humans.In another embodiment, a different GDE coding sequence is selected.

[0220] In one aspect, a CpG-minimized engineered nucleic acid sequence encoding a human GDE is provided. In one embodiment, an engineered human GDE cDNA is provided herein (as SEQ ID NO: 179), which was designed to minimize CpG motifs compared to the native GDE sequence (SEQ ID NO: 174). Preferably, the CpG-minimized GDE coding sequence has less than about 90% identity, preferably about 85% identity or less, to the full-length native GDE coding sequence (SEQ ID NO: 174). In one embodiment, the CpG-minimized GDE coding sequence has about 81% identity to the native GDE coding sequence of SEQ ID NO: 174. In one embodiment, the CpG-minimized GDE coding sequence is characterized by reduced activation of a host immune response after delivery into a host cell compared to the native GDE sequence. In one embodiment, the CpG-minimized GDE coding sequence shares less than about 99%, less than 98%, less than 97%, less than 96%, less than 95%, less than 94%, less than 93%, less than 92%, less than 91%, less than 90%, less than 89%, less than 88%, less than 87%, less than 86%, less than 85%, less than 84%, less than 83%, less than 82%, less than 81%, less than 80%, less than 79%, less than 78%, less than 77%, less than 76%, less than 75%, less than 74%, less than 73%, less than 72%, less than 71%, less than 70%, less than 69%, less than 68%, less than 67%, less than 66%, less than 65%, less than 64%, less than 63%, less than 62%, less than 61% or less identity to the full-length native GDE coding sequence of SEQ ID NO:174. In one embodiment, the CpG-minimized nucleic acid sequence is a variant of SEQ ID NO: 179. In another embodiment, the CpG-minimized nucleic acid sequence has a sequence that shares about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61% or more identity to SEQ ID NO: 179. In one embodiment, the CpG-minimized nucleic acid sequence is SEQ ID NO: 179.

[0221] In some embodiments, the hairpin-ended DNA molecules described herein encode a fusion protein comprising a full-length, fragment, or portion of a GDE protein fused to another sequence (e.g., an N- or C-terminal fusion). In some embodiments, the N- or C-terminal sequence is a signal sequence or a cell-targeting sequence.

[0222] In specific embodiments, the expression cassette comprises a GDE transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 174. In specific embodiments, the expression cassette comprises a GDE transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 175. In specific embodiments, the expression cassette comprises a GDE transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 179. In specific embodiments, the expression cassette comprises a GDE transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 178. In specific embodiments, the expression cassette comprises a GDE transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 179.

[0223] In a specific embodiment, the expression cassette comprises a GDE transgene identical to the sequence set forth in SEQ ID NO: 174. In a specific embodiment, the expression cassette comprises a GDE transgene identical to the sequence set forth in SEQ ID NO: 175. In a specific embodiment, the expression cassette comprises a GDE transgene identical to the sequence set forth in SEQ ID NO: 179. In a specific embodiment, the expression cassette comprises a GDE transgene identical to the sequence set forth in SEQ ID NO: 178. In a specific embodiment, the expression cassette comprises a GDE transgene identical to the sequence set forth in SEQ ID NO: 179.

[0224] The terms "percent (%) identity," "sequence identity," "percent sequence identity," or "percent identical," in the context of nucleic acid sequences encoding a GDE, refer to the number of residues in two sequences that are identical when aligned. The length of sequence identity comparison can span the entire length of a genome, the entire length of a gene coding sequence, or a fragment of at least about 500-5000 nucleotides, where appropriate. However, identity over smaller fragments, e.g., at least about 9 nucleotides, usually at least about 20-24 nucleotides, at least about 28-32 nucleotides, at least about 36 or more nucleotides, can also be desirable.

[0225] Percent identity can be readily determined for amino acid sequences over the full length of a protein, polypeptide, about 32 amino acids, about 330 amino acids, or fragment peptides thereof, or for the corresponding nucleic acid sequence coding sequence. Suitable amino acid fragments can be at least about 8 amino acids in length and up to about 700 amino acids in length. Generally, when referring to "identity," "homology," or "similarity" between two different sequences, the "identity," "homology," or "similarity" is determined with respect to "aligned" sequences. "Aligned" sequences or "alignment" refers to multiple nucleic acid or protein (amino acid) sequences that often contain modifications in terms of missing or additional bases or amino acids compared to a reference sequence.

[0226] Identity can be determined by preparing an alignment of sequences and using various algorithms and / or computer programs known in the art or commercially available (e.g., BLAST, ExPASy; ClustalO; FASTA; e.g., using the Needleman-Wunsch algorithm, Smith-Waterman algorithm). Alignment can be performed using any of a variety of publicly or commercially available sequence alignment programs. Sequence alignment programs, such as the "Clustal Omega" and "Clustal X" programs, are available for amino acid sequences. Generally, any of these programs are used with default settings, although those skilled in the art can change these settings as needed. Alternatively, those skilled in the art can use other algorithms or computer programs that provide at least a similar level of identity or alignment to that provided by the referenced algorithms and programs. See, e.g., J.D. Thomson et al., Nucl. Acids. Res., "A comprehensive comparison of multiple sequence alignments," 27(13):2682-2690 (1999). Numerous sequence alignment programs are also available for nucleic acid sequences. Examples of such programs include "Clustal Omega," "Clustal W," "CAP Sequence Assembly," "BLAST," "MAP," and "MEME," which are accessible via web servers on the Internet.

[0227] Codon-optimized coding regions can be designed in a variety of different ways. This optimization can be performed using methods available online (e.g., GeneArt), published methods, or companies that provide codon optimization services, such as DNA2.0 (Menlo Park, CA). Preferably, the entire length of the product open reading frame (ORF) is modified. However, in some embodiments, only a fragment of the ORF may be altered. By using one of these methods, frequencies can be applied to any polypeptide sequence to produce a nucleic acid fragment of a codon-optimized coding region that encodes the polypeptide. Several options are available for actually making the changes to the codons or for synthesizing the codon-optimized coding regions designed as described herein. Such modifications or synthesis can be performed using standard and routine molecular biology procedures well known to those skilled in the art.

[0228] The GDE expression cassette can be located at any base pair distance from either the 5' and / or 3' ITR closing pair (described in Section 5.4.1) suitable to allow or maintain efficient transcription of the expression cassette in the host cell. In some embodiments, the distance between the expression cassette and the 5' ITR and the distance between the expression cassette and the 3' ITR closing pair are the same. In some embodiments, the distance between the expression cassette and the 5' ITR and the distance between the expression cassette and the 3' ITR closing pair are not the same.In some embodiments, the distance between the expression cassette and / or the 3' ITR closing pair and the distance between the expression cassette and the 5' ITR closing pair is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least at least 195, at least 200, at least 205, at least 210, at least 215, at least 220, at least 225, at least 230, at least 235, at least 240, at least 245, at least 250, at least 255, at least 260, at least 265, at least 270, at least 275, at least 280, at least 285, at least 290, at least 295, at least 300, at least 305, at least 310, at least 315, at least 320, at least 325, at least 330, at least 335, at least 340, at least 345, at least 350, at least 355, at least 360, at least 365, at least 370, at least 375, at least 380, at least 385, at least 390, at least 395, or at least 400 nucleotides.In some embodiments, the distance between the expression cassette and the 3' ITR closing pair and / or the distance between the expression cassette and the 5' ITR closing pair is about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 200, about 210, about 215, about 220, about 225, about 230, about 235, about 240, about 245, about 250, about 255, about 260, about 265, about 270, about 275, about 280, about 285, about 290, about 300, about 310, about 315, about 320, about 325, about 330, about 335, about 340, about 345, about 350, about 355, about 360, about 365, about 370, about 375, about 380, about 385, about 390, about 400, about 410, about 420, about 430, about 440, about 450, about 460, about 470, about about 190, about 195, about 200, about 205, about 210, about 215, about 220, about 225, about 230, about 235, about 240, about 245, about 250, about 255, about 260, about 265, about 270, about 275, about 280, about 285, about 290, about 295, about 300, about 305, about 310, about 315, about 320, about 325, about 330, about 335, about 340, about 345, about 350, about 355, about 360, about 365, about 370, about 375, about 380, about 385, about 390, about 395 or about 400 nucleotides.

[0229] By "engineered nucleic acid sequence," it is meant that the nucleic acid sequence encoding the GDE protein described herein is incorporated into and positioned on any suitable genetic element, e.g., naked DNA, phage, transposon, cosmid, episome, etc., that transfers the GDE sequence it carries to a host cell, e.g., to generate a non-viral delivery system (e.g., RNA-based system, naked DNA, etc.), or to generate a viral vector within a packaging host cell, and / or for delivery to a host cell in a subject. In one embodiment, the genetic element is a circular plasmid. Methods used to generate such engineered constructs are known to those skilled in nucleic acid manipulation and include genetic engineering, recombinant engineering, and synthetic techniques. See, e.g., Green and Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, NY (2012).

[0230] In one embodiment, the nucleic acid sequence encoding the GDE further comprises a nucleic acid encoding a tag polypeptide covalently linked thereto. The tag polypeptide may be selected from known "epitope tags," such as, but not limited to, myc tag polypeptide, glutathione-S-transferase tag polypeptide, luciferase protein tag polypeptide, green fluorescent protein tag polypeptide, myc-pyruvate kinase tag polypeptide, His6 tag polypeptide, influenza virus hemagglutinin tag polypeptide, flag tag polypeptide, and maltose binding protein tag polypeptide. In some embodiments, a hairpin-end vector expressing a GDE protein linked to a reporter polypeptide may be used for diagnostic purposes, to determine efficacy, or as a marker of activity of the hairpin-end vector in a subject to which it is administered.

[0231] 5.4.4 Hairpin-Ended DNA Molecules Encoding GDEs As is apparent from the above description, the hairpin-ended DNA molecules for expressing human amylo-α-1,6-glucosidase, 4-α-glucanotransferase provided herein contain an expression cassette. An "expression cassette" is a nucleic acid molecule or portion of a nucleic acid molecule that contains sequences or other information that direct the cellular machinery to produce RNA and protein. An expression cassette may contain a transcription unit or open reading frame (ORF) encoding a GDE protein or a fragment thereof. In some embodiments, the expression cassette contains a promoter sequence. In some further embodiments, the expression cassette contains a promoter operably linked to the transcription unit. The expression cassette may further contain features that direct the cellular machinery to produce RNA and protein. In one embodiment, the expression cassette contains post-transcriptional control elements. In another embodiment, the expression cassette further contains polyadenylation and / or termination signals. In yet another embodiment, the expression cassette contains regulatory elements known and used in the art to regulate (enhance, inhibit, and / or turn on / off expression of an ORF). Such regulatory elements include, for example, the 5'-untranslated region (UTR), the 3'-UTR, or both the 5'UTR and the 3'UTR. In some further embodiments, the expression cassette comprises any one or more features provided in this section (Section 5.4.3), in any combination or permutation.

[0232] An expression cassette may contain a protein-coding sequence in its ORF (sense strand). Alternatively, the expression cassette may contain a complementary sequence (antisense) of the protein-encoding ORF and regulatory components and / or other signals that direct the cellular machinery to produce the sense strand DNA / RNA and the corresponding protein. In some embodiments, the expression cassette contains a GDE protein sequence without an intron. In other embodiments, the expression cassette contains a GDE protein sequence with an intron that is removed during transcription and splicing. An expression cassette may also contain a variable number of ORFs or transcription units. In one embodiment, the expression cassette contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 ORFs. In another embodiment, the expression cassette comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 transcription units.

[0233] An expression cassette can also contain one or more transcriptional regulatory elements, one or more post-transcriptional regulatory elements, or both one or more transcriptional regulatory elements and one or more post-transcriptional regulatory elements. Such regulatory elements are any sequences that enable, contribute to, or regulate the functional regulation of a nucleic acid molecule, including replication, duplication, transcription, splicing, translation, stability, and / or transport of the nucleic acid or one of its derivatives (e.g., mRNA) into a host cell or organism. Such regulatory elements include, but are not limited to, promoters, enhancers, polyadenylation signals, translation stop codons, ribosome binding elements, transcription terminators, selectable markers, origins of replication, etc.

[0234] In some embodiments, the expression cassette includes an enhancer. Any enhancer sequence known to those skilled in the art in light of the present disclosure can be used. In some embodiments, the enhancer sequence can be human actin, human myosin, human hemoglobin, human muscle creatine, or a viral enhancer such as one of CMV, HA, RSV, or EBV. In specific embodiments, the enhancer can be a woodchuck HBV posttranscriptional regulator (WPRE), an intron / exon sequence derived from human apolipoprotein A1 precursor (ApoAI), the untranslated R-U5 domain of the human T-cell leukemia virus type 1 (HTLV-1) long terminal repeat (LTR), a splicing enhancer, a synthetic rabbit β-globin intron, an AAV P5 promoter, or any combination thereof.

[0235] As described above, the expression cassette can include a promoter that controls the expression of a protein of interest. A promoter includes any nucleotide sequence that initiates transcription of an operably linked nucleotide sequence. Promoters can be constitutive, inducible, or repressible. Promoters can be derived from viral, bacterial, fungal, plant, insect, and animal sources. Promoters can be homologous (e.g., derived from the same genetic source) or heterologous (e.g., derived from a different genetic source). In some embodiments, the promoter can be a promoter derived from simian virus 40 (SV40), a mouse mammary tumor virus (MMTV) promoter, a human immunodeficiency virus (HIV) promoter, such as the bovine immunodeficiency virus (BIV) long terminal repeat (LTR) promoter, a Moloney virus promoter, an avian leukosis virus (ALV) promoter, a cytomegalovirus (CMV) promoter, such as the CMV immediate early promoter (CMV-IE), an Epstein-Barr virus (EBV) promoter, or a Rous sarcoma virus (RSV) promoter. In another embodiment, the promoter can be a promoter derived from a human gene, such as human actin, human myosin, human hemoglobin, human muscle creatine, or human metallothionein. In a further embodiment, the promoter can be a tissue-specific promoter, such as a natural or synthetic muscle- or skin-specific promoter, that drives expression in cells or tissues in which expression of the GDE is desired, such as in cells or tissues in which expression of the GDE is desired in GDE-deficient patients.

[0236] In certain embodiments, the promoter is a muscle-specific promoter. Non-limiting examples of muscle-specific promoters include the muscle creatine kinase (MCK) promoter. Non-limiting examples of suitable muscle creatine kinase promoters include the human muscle creatine kinase promoter and the truncated mouse muscle creatine kinase (tMCK) promoter (Wang B et al., "Construction and analysis of compact muscle-selective promoters for AAV vectors." Gene Ther. 2008 Nov;15(22):1489-99) (Representative GenBank accession number: AF188002). Human muscle creatine kinase has gene ID number 1158 (Representative GenBank accession number: NC 000019.9). Other examples of muscle-specific promoters include the synthetic promoter C5.12 (spC5.12, or also referred to herein as "C5.12"), such as the spC5.12 or spC5.12 promoter (disclosed in Wang et al., Gene Therapy 15:1489-1499 (2008)), the MHCK7 promoter (Salva et al., Mol Ther. 2007 Feb;15(2):320-9), myosin light chain (MLC) promoters, such as MLC2 (Gene ID No. 4633; representative GenBank Accession No. NG 007554.1); myosin heavy chain (MHC) promoters, such as α-MHC (Gene ID No. 4624; representative GenBank Accession No. NG 023444.1); and the desmin promoter (Gene ID No. 1674; representative GenBank Accession No. NG 023444.1). 008043.1); cardiac troponin C promoter (Gene ID No. 7134; representative GenBank accession number NG_008963.1); troponin I promoter (Gene ID Nos. 7135, 7136, and 7137; representative GenBank accession numbers NG_016649.1, NG_011621.1, and NG_007866.2); myoD gene family promoter (Weintraub et al., Science, 251, 761 (1991); Gene ID No. 4654; representative GenBank accession number NM_002478); α-actin promoter (Gene ID Nos. 58, 59, and 70; representative GenBank accession numbers NG_006672.1, NG_011541.1, and NG_007866.2). 007553.1,); β-actin promoter (Gene ID number 60; representative GenBank accession number NG 007992.1); γ-actin promoter (Gene ID numbers 71 and 72; representative GenBank accession number NG 011433.1 and NM 001199893); the muscle-specific promoter within intron 1 of the eye form of Pitx3 (Gene ID No. 5309) (Coulon et al.; this muscle-selective promoter corresponds to residues 11219-11527 of representative GenBank accession number NG 008147); and promoters described in U.S. Patent Publication US 2003 / 0157064, and the CK6 promoter (Wang et al., 2008 doi: 10.1038 / gt.2008.104). In another specific embodiment, the muscle-specific promoter is the E-Syn promoter described in Wang et al., Gene Therapy 15:1489-1499 (2008), which comprises a combination of an MCK-inducible enhancer and an spC5.12 promoter. In certain embodiments of the present disclosure, the muscle-specific promoter is selected from the group consisting of spC5.12 promoter, MHCK7 promoter, E-syn promoter, muscle creatine kinase myosin light chain (MLC) promoter, myosin heavy chain (MHC) promoter, cardiac troponin C promoter, troponin I promoter, myoD gene family promoter, alpha actin promoter, beta actin promoter, gamma actin promoter, the muscle-specific promoter in intron 1 of the ophthalmic form of Pitx3, CK6 promoter, CK8 promoter, and Acta promoter. In certain embodiments, the muscle-specific promoter is selected from the group consisting of spC5.12, desmin, and MCK promoters. In further embodiments, the muscle-specific promoter is selected from the group consisting of spC5.12 and MCK promoters. In certain embodiments, the muscle-specific promoter is the spC5.12 promoter.

[0237] In certain embodiments, the promoter is a liver-specific promoter. Non-limiting examples of liver-specific promoters include the alpha-1 antitrypsin promoter (hAAT), transthyretin promoter, albumin promoter, thyroxine-binding globulin (TBG) promoter, LSP promoter (thyroid hormone-binding globulin promoter sequence, two copies of alpha-microglobulin / bikunin enhancer sequence, and leader sequence - Ill, CR et al. (1997). "Optimization of the human factor VIII complementary DNA expression plasmid for gene therapy of hemophilia A." Blood Coag. Fibrinol. 8: S23-S30), and the like. Other useful liver-specific promoters can be found, for example, in the Liver Specific Gene Promoter Database, Cold Spring Harbor Laboratory of Gene Therapy. Those listed in the Laboratory Compilation (http: / / rulai.cshl.edu / LSPD / ) are known in the art. A suitable liver-specific promoter in the context of the present disclosure is the hAAT promoter. In another specific embodiment, the promoter is a neuron-specific promoter. Non-limiting examples of neuron-specific promoters will be apparent to those skilled in the art, and include, but are not limited to, the synapsin-1 (Syn) promoter, neuron-specific enolase (NSE) promoter (Andersen et al., Cell. Mol. Neurobiol., 13:503-15 (1993)), neurofilament light chain gene promoter (Piccioli et al., Proc. Natl. Acad. Sci. USA, 88:5611-5 (1991)), and neuron-specific vgf gene promoter (Piccioli et al., Neuron, 15:373-84 (1995)).In certain embodiments, the neuron-specific promoter is a Syn promoter. Other neuron-specific promoters include, but are not limited to, synapsin-2 promoter, tyrosine hydroxylase promoter, dopamine b-hydroxylase promoter, hypoxanthine phosphoribosyltransferase promoter, low-affinity NGF receptor promoter, and choline acetyltransferase promoter (Bejanin et al., 1992; Carroll et al., 1995; Chin and Greengard, 1994; Foss-Petter et al., 1990; Harrington et al., 1987; Mercer et al., 1991; Patei et al., 1986). Representative promoters specific to motor neurons include, but are not limited to, the promoter of calcitonin gene-related peptide (CGRP), a known motor neuron-derived factor. Other promoters functional in motor neurons include the promoters of choline acetyltransferase (ChAT), neuron-specific enolase (NSE), synapsin, and Hb9. Other neuron-specific promoters useful in the present disclosure include, but are not limited to: GFAP (astrocytes), Calbindin 2 (interneurons), Mnxl (motomeurons), Nestin (neurons), Parvalbumin, Somatostation, and Plpl (oligodendrocytes and Schwann cells). In another specific embodiment, the promoter is a ubiquitous promoter.Representative ubiquitous promoters include the cytomegalovirus enhancer / chicken β-actin (CAG) promoter, the cytomegalovirus enhancer / promoter (CMV) (optionally with a CMV enhancer) [see, for example, Boshart et al., Cell, 41:521-530 (1985)], the PGK promoter, the SV40 early promoter, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with a RSV enhancer), the dihydrofolate reductase promoter, the b-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1 alpha promoter. The promoter may also be an endogenous promoter, such as an albumin promoter or a GDE promoter. In certain embodiments, the promoter is linked to an enhancer sequence, such as a cis-regulatory module (CRM) or an artificial enhancer sequence. CRMs useful in practicing the present disclosure include those described in Rincon et al., Mol Ther. 2015 Jan;23(1):43-52; Chuah et al., Mol Ther. 2014 Sep;22(9):1605-13; or Nair et al., Blood. 2014 May 15;123(20):3195-9. Other regulatory elements that can enhance muscle-specific expression of genes, particularly cardiac and / or skeletal muscle expression, are those disclosed in WO2015110449. Particular examples of nucleic acid regulatory elements comprising artificial sequences include regulatory elements obtained by rearranging transcription factor binding sites (TFBSs) present within the sequences disclosed in WO2015110449. Such rearrangements can involve changing the order of the TFBSs and / or changing the position of one or more TFBSs relative to other TFBSs and / or changing the copy number of one or more of the TFBSs.For example, nucleic acid regulatory elements for enhancing muscle-specific gene expression, particularly cardiac and skeletal muscle-specific gene expression, can include binding sites for E2A, HNH1, NF1, C / EBP, LRF, MyoD, and SREBP; or binding sites for E2A, NF1, p53, C / EBP, LRF, and SREBP; or binding sites for E2A, HNH The protein may comprise binding sites for E2A, HNF3a, HNF3b, NF1, C / EBP, LRF, MyoD, and SREBP; or binding sites for E2A, HNF3a, NF1, CEBP, LRF, MyoD, and SREBP; or binding sites for E2A, HNF3a, NF1, CEBP, LRF, MyoD, and SREBP; or binding sites for HNF4, NF1, RSRFC4, C / EBP, LRF, and MyoD, or binding sites for NF1, PPAR, p53, C / EBP, LRF, and MyoD. For example, a nucleic acid regulatory element for enhancing muscle-specific gene expression, particularly skeletal muscle-specific gene expression, can also include binding sites for E2A, NF1, SRFC, p53, C / EBP, LRF, and MyoD; or binding sites for E2A, NF1, C / EBP, LRF, MyoD, and SREBP; or binding sites for E2A, HNF3a, C / EBP, LRF, MyoD, SEREBP, and Tall b; or binding sites for E2A, SRF, p53, C / EBP, LRF, MyoD, and SREBP; or binding sites for HNF4, NF1, RSRFC4, C / EBP, LRF, and SREBP; or binding sites for E2A, HNF3a, HNF3b, NF1, SRF, C / EBP, LRF, MyoD, and SREBP; or binding sites for E2A, CEBP, and MyoD. In other examples, these nucleic acid regulatory elements contain at least two copies, e.g., two, three, four, or more copies, of one or more of the aforementioned TFBSs. In particular, other regulatory elements that can enhance liver-specific expression of genes are those disclosed in WO2009130208. Table 19: Exemplary Regulatory Elements [Table 19] TIFF2024517427000053.tif248170TIFF2024517427000054.tif133170

[0238] In some embodiments, the expression cassette can include a polyadenylation signal, a termination signal, or both a polyadenylation signal and a termination signal. Any polyadenylation signal known to those of skill in the art in light of the present disclosure can be used. In some embodiments, the polyadenylation signal can be an SV40 polyadenylation signal, an AAV2 polyadenylation signal (bp 4411-4466, NC_001401), a polyadenylation signal derived from the herpes simplex virus thymidine kinase gene, an LTR polyadenylation signal, a bovine growth hormone (bGH) polyadenylation signal, a human growth hormone (hGH) polyadenylation signal, or a human β-globin polyadenylation signal.

[0239] In some embodiments, the expression cassette can have various sizes to accommodate one or more ORFs of various lengths. In some embodiments, the size of the expression cassette is at least 4.5 kb, at least 5 kb, at least 5.5 kb, at least 6 kb, at least 6.5 kb, at least 7 kb, at least 7.5 kb, at least 8 kb, at least 8.5 kb, at least 9 kb, at least 9.5 kb, at least 10 kb, at least 15 kb, at least 20 kb, at least 25 kb, at least 30 kb, at least 35 kb, at least 40 kb, at least 45 kb, at least 50 kb, at least 55 kb, at least 60 kb, at least 65 kb, at least 70 kb, at least 75 kb, or at least 80 kb. In a specific embodiment, the expression cassette is at least 4.5 kb. In another specific embodiment, the expression cassette is at least 4.6 kb. In yet another specific embodiment, the expression cassette is at least 4.7 kb. In a more specific embodiment, the expression cassette is at least 4.8 kb. In one particular embodiment, the expression cassette is at least 4.9 kb. About 4.5 kb, about 5 kb, about 5.5 kb, about 6 kb, about 6.5 kb, about 7 kb, about 7.5 kb, about 8 kb, about 8.5 kb, about 9 kb, about 9.5 kb, about 10 kb, about 15 kb, about 20 kb, about 25 kb, about 30 kb, about 35 kb, about 40 kb, about 45 kb, about 50 kb, about 55 kb, about 60 kb, about 65 kb, about 70 kb, about 75 kb, or about 80 kb. In one particular embodiment, the expression cassette is about 4.5 kb. In another specific embodiment, the expression cassette is about 4.6 kb. In yet another specific embodiment, the expression cassette is about 4.7 kb. In a more specific embodiment, the expression cassette is about 4.8 kb. In one particular embodiment, the expression cassette is about 4.9 kb. In another specific embodiment, the expression cassette is about 5 kb. The expression cassette can also contain a variable number of genes of interest ("transgenes").In one embodiment, the expression cassette comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 transgenes. In some specific embodiments, the expression cassette comprises one transgene. In some embodiments, the transgene is a recombinant gene. In some further embodiments, the transgene comprises a cDNA sequence (e.g., no introns in the transgene).

[0240] In some embodiments, the DNA molecules provided herein do not have the size limitations of encapsidated AAV vectors, thus allowing for the delivery of large expression cassettes and providing efficient transgene delivery. In certain embodiments, the DNA molecules provided herein contain expression cassettes that are the same size as or larger than any naturally occurring AAV genome.

[0241] The expression cassette can occupy various positions relative to the inverted repeat. In some embodiments, the expression cassette is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides apart.In some embodiments, the expression cassette is at least 0.2 kb, at least 0.3 kb, at least 0.4 kb, at least 0.5 kb, at least 0.6, at least kb, at least 0.7 kb, at least 0.8 kb, at least 0.9 kb, at least 1 kb, at least 1.5 kb, or at least 2 kb away from the inverted repeat. In another embodiment, the expression cassette comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102 about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99 or about 100 nucleotides apart. In further embodiments, the expression cassette is about 0.2 kb, about 0.3 kb, about 0.4 kb, about 0.5 kb, about 0.6 kb, about 0.7 kb, about 0.8 kb, about 0.9 kb, about 1 kb, about 1.5 kb, or about 2 kb away from the inverted repeat. In one embodiment, the inverted repeat of this paragraph is the first inverted repeat described in Sections 3 and 5.4 (including 5.4.1). In another embodiment, the inverted repeat of this paragraph is the second inverted repeat described in Sections 3 and 5.4 (including 5.4.1). In yet another embodiment, the inverted repeat of this paragraph is both the first and second inverted repeats described in Sections 3 and 5.4 (including 5.4.1).

[0242] In one aspect, provided herein is a therapeutic antibody comprising, in the 5' to 3' direction of the sense strand: i) a first inverted repeat (e.g., as described in Section 5.4.1), wherein first and second restriction sites for a nicking endonuclease are located on opposite strands near the first inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the first inverted repeat, nicking results in a sense strand 5' overhang that includes the first inverted repeat; a sense expression cassette encoding a GDE protein; and iii) a double-stranded DNA molecule comprising a second inverted repeat (e.g., as described in Section 5.4.1), wherein third and fourth restriction sites for a nicking endonuclease are located on opposite strands near the second inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the top strand from the antisense strand of the second inverted repeat, nicking results in a sense strand 3' overhang that includes the second inverted repeat.

[0243] In another aspect, provided herein is a nucleic acid sequence comprising, in the 5' to 3' direction of the sense strand: i) a first inverted repeat (e.g., as described in Section 5.4.1), where first and second restriction sites for a nicking endonuclease are located on opposite strands near the first inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the first inverted repeat, nicking results in an antisense strand 3' overhang that includes the first inverted repeat. ii) a sense expression cassette encoding a therapeutic GDE protein; and iii) a second inverted repeat (e.g., as described in Section 5.4.1), wherein third and fourth restriction sites for a nicking endonuclease are located on opposite strands near the second inverted repeat, such that upon separation of the sense from the antisense of the second inverted repeat, nicking results in an antisense strand 5' overhang that includes the second inverted repeat.

[0244] In yet another aspect, provided herein is a nucleic acid sequence comprising, in the 5' to 3' direction of the sense strand: i) a first inverted repeat (e.g., as described in Section 5.4.1), wherein first and second restriction sites for a nicking endonuclease are located on opposite strands near the first inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the first inverted repeat, nicking results in a sense strand 5' overhang that includes the first inverted repeat; and iii) a double-stranded DNA molecule comprising a second inverted repeat (e.g., as described in Section 5.4.1), wherein third and fourth restriction sites for a nicking endonuclease are located on opposite strands near the second inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the second inverted repeat, nicking results in an antisense-strand 5' overhang that includes the second inverted repeat.

[0245] In a further aspect, provided herein is a nucleic acid sequence comprising, in the 5' to 3' direction of the sense strand: i) a first inverted repeat (e.g., as described in Section 5.4.1), wherein first and second restriction sites for a nicking endonuclease are located on opposite strands near the first inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the first inverted repeat, nicking results in an antisense strand 3' overhang that includes the first inverted repeat; and iii) a double-stranded DNA molecule comprising a second inverted repeat (e.g., as described in Section 5.4.1), wherein third and fourth restriction sites for a nicking endonuclease are located on opposite strands near the second inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2 or shown in Figures 2B and 2C), such that upon separation of the sense strand from the antisense strand of the second inverted repeat, nicking results in a sense strand 3' overhang that includes the second inverted repeat.

[0246] In one aspect, provided herein is a nucleic acid sequence comprising, in the 5' to 3' direction of the sense strand: i) a first inverted repeat (e.g., as described in Section 5.4.1), wherein first and second target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposite strands near the first inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the first inverted repeat, nicking by the programmable nicking enzyme results in a sense strand 5' overhang that includes the first inverted repeat; a double-stranded DNA molecule comprising: a sense expression cassette encoding a therapeutic GDE protein; and iii) a second inverted repeat (e.g., as described in Section 5.4.1), wherein third and fourth target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposite strands near the second inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense from the antisense strand of the second inverted repeat, nicking by the programmable nicking enzyme results in a sense strand 3' overhang that includes the second inverted repeat.

[0247] In another aspect, provided herein is a nucleic acid sequence comprising, in the 5' to 3' direction of the sense strand: i) a first inverted repeat (e.g., as described in Section 5.4.1), wherein first and second target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposite strands near the first inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the first inverted repeat, nicking by the programmable nicking enzyme results in an antisense strand 3' overhang that includes the first inverted repeat; and iii) a double-stranded DNA molecule comprising a second inverted repeat (e.g., as described in Section 5.4.1), wherein third and fourth target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposite strands near the second inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the second inverted repeat, nicking by the programmable nicking enzyme results in an antisense-strand 5' overhang that includes the second inverted repeat.

[0248] In yet another aspect, provided herein is a nucleic acid sequence comprising, in the 5' to 3' direction of the sense strand: i) a first inverted repeat (e.g., as described in Section 5.4.1), wherein first and second target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposite strands near the first inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the first inverted repeat, nicking by the programmable nicking enzyme results in a sense strand 5' overhang that includes the first inverted repeat; and iii) a double-stranded DNA molecule comprising a second inverted repeat (e.g., as described in Section 5.4.1), wherein third and fourth target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposite strands near the second inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the second inverted repeat, nicking by the programmable nicking enzyme results in an antisense-strand 5' overhang that includes the second inverted repeat.

[0249] In a further aspect, provided herein is a nucleic acid sequence comprising, in the 5' to 3' direction of the sense strand: i) a first inverted repeat (e.g., as described in Section 5.4.1), wherein first and second target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposite strands near the first inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the sense strand from the antisense strand of the first inverted repeat, nicking by the programmable nicking enzyme results in an antisense strand 3' overhang that includes the first inverted repeat; and iii) a double-stranded DNA molecule comprising a second inverted repeat (e.g., as described in Section 5.4.1), wherein third and fourth target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposite strands near the second inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2 or shown in Figures 2B and 2C), such that upon separation of the sense strand from the antisense strand of the second inverted repeat, nicking by the programmable nicking enzyme results in a sense strand 3' overhang that includes the second inverted repeat. In one embodiment, the first, second, third, and fourth target sites for the programmable nicking enzyme in this and the three immediately preceding paragraphs are all the same. In another embodiment, three of the first, second, third, and fourth target sites for the programmable nicking enzyme in this and the three immediately preceding paragraphs are the same. In yet another embodiment, two of the first, second, third, and fourth target sites for the programmable nicking enzyme in this paragraph and the three immediately preceding paragraphs are the same. In a further embodiment, the first, second, third, and fourth target sites for the programmable nicking enzyme in this paragraph and the three immediately preceding paragraphs are all different.

[0250] An expression cassette can also include one or more transcriptional regulatory elements, one or more post-transcriptional regulatory elements, or both one or more transcriptional regulatory elements and one or more post-transcriptional regulatory elements. Such regulatory elements are any sequences that enable, contribute to, or regulate the functional regulation of a nucleic acid molecule, including replication, duplication, transcription, splicing, translation, stability, and / or transport of the nucleic acid or one of its derivatives (e.g., mRNA) into a host cell or organism. Such regulatory elements include, but are not limited to, promoters, enhancers, polyadenylation signals, translation stop codons, ribosome binding elements, transcription terminators, selectable markers, and / or origins of replication.

[0251] Expression cassettes can be of various sizes to accommodate one or more ORFs of various lengths. In some embodiments, the size of the expression cassette is at least 0.2 kb, at least 0.3 kb, at least 0.4 kb, at least 0.5 kb, at least 0.6 kb, at least 0.7 kb, at least 0.8 kb, at least 0.9 kb, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 2.5 kb, at least 3 kb, at least 3.5 kb, at least 4 kb, at least 4.5 kb, at least 5 kb, at least 5.5 kb, at least 6 kb, The expression cassette may be at least 6.5 kb, at least 7 kb, at least 7.5 kb, at least 8 kb, at least 8.5 kb, at least 9 kb, at least 9.5 kb, at least 10 kb, at least 15 kb, at least 20 kb, at least 25 kb, at least 30 kb, at least 35 kb, at least 40 kb, at least 45 kb, at least 50 kb, at least 55 kb, at least 60 kb, at least 65 kb, at least 70 kb, at least 75 kb, or at least 80 kb. In a specific embodiment, the expression cassette is at least 4.5 kb. In another specific embodiment, the expression cassette is at least 4.6 kb. In yet another specific embodiment, the expression cassette is at least 4.7 kb. In an even more specific embodiment, the expression cassette is at least 4.8 kb. In a specific embodiment, the expression cassette is at least 4.9 kb. In another specific embodiment, the expression cassette is at least 5 kb.In another embodiment, the size of the expression cassette is about 0.2 kb, about 0.3 kb, about 0.4 kb, about 0.5 kb, about 0.6 kb, about 0.7 kb, about 0.8 kb, about 0.9 kb, about 1 kb, about 1.5 kb, about 2 kb, about 2.5 kb, about 3 kb, about 3.5 kb, about 4 kb, about 4.5 kb, about 5 kb, about 5.5 kb, The length of the expression cassette is about 6 kb, about 6.5 kb, about 7 kb, about 7.5 kb, about 8 kb, about 8.5 kb, about 9 kb, about 9.5 kb, about 10 kb, about 15 kb, about 20 kb, about 25 kb, about 30 kb, about 35 kb, about 40 kb, about 45 kb, about 50 kb, about 55 kb, about 60 kb, about 65 kb, about 70 kb, about 75 kb, or about 80 kb. In a specific embodiment, the expression cassette is about 4.5 kb. In another specific embodiment, the expression cassette is about 4.6 kb. In yet another specific embodiment, the expression cassette is about 4.7 kb. In an even more specific embodiment, the expression cassette is about 4.8 kb. In a specific embodiment, the expression cassette is about 4.9 kb. In another specific embodiment, the expression cassette is about 5 kb. The expression cassette can also contain a variable number of genes of interest ("transgenes"). In embodiments, the expression cassette contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 transgenes. In some specific embodiments, the expression cassette contains one transgene. In some embodiments, the transgene is a recombinant gene. In some further embodiments, the transgene comprises a cDNA sequence (e.g., no introns in the transgene).

[0252] Furthermore, an expression cassette can comprise at least 4,000 nucleotides, at least 5,000 nucleotides, at least 10,000 nucleotides, at least 20,000 nucleotides, at least 30,000 nucleotides, at least 40,000 nucleotides, or at least 50,000 nucleotides. In some embodiments, an expression cassette can comprise any range of nucleotides, from about 4,000 to about 10,000 nucleotides, from about 10,000 to about 50,000 nucleotides, or greater than 50,000 nucleotides. In some embodiments, an expression cassette can comprise a transgene ranging in length from about 500 to about 50,000 nucleotides. In some embodiments, an expression cassette can comprise a transgene ranging in length from about 500 to about 75,000 nucleotides. In some embodiments, an expression cassette can comprise a transgene ranging in length from about 500 to about 10,000 nucleotides. In some embodiments, an expression cassette can comprise a transgene ranging in length from about 1,000 to about 10,000 nucleotides. In some embodiments, the expression cassette can contain a transgene ranging in length from about 500 to about 5,000 nucleotides. In some embodiments, the DNA molecules provided herein do not have the size limitations of encapsidated AAV vectors, thus allowing for the delivery of large expression cassettes and providing efficient transgene delivery. In certain embodiments, the DNA molecules provided herein contain expression cassettes that are the same size as or larger than any naturally occurring AAV genome.

[0253] The expression cassette can occupy various positions relative to the inverted repeat. In some embodiments, the expression cassette is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides apart.In some embodiments, the expression cassette is at least 0.2 kb, at least 0.3 kb, at least 0.4 kb, at least 0.5 kb, at least 0.6, at least kb, at least 0.7 kb, at least 0.8 kb, at least 0.9 kb, at least 1 kb, at least 1.5 kb, or at least 2 kb away from the inverted repeat. In another embodiment, the expression cassette comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102 about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99 or about 100 nucleotides apart. In further embodiments, the expression cassette is about 0.2 kb, about 0.3 kb, about 0.4 kb, about 0.5 kb, about 0.6 kb, about 0.7 kb, about 0.8 kb, about 0.9 kb, about 1 kb, about 1.5 kb, or about 2 kb away from the inverted repeat. In one embodiment, the inverted repeat of this paragraph is the first inverted repeat described in Sections 3 and 5.4 (including 5.4.1). In another embodiment, the inverted repeat of this paragraph is the second inverted repeat described in Sections 3 and 5.4 (including 5.4.1). In yet another embodiment, the inverted repeat of this paragraph is both the first and second inverted repeats described in Sections 3 and 5.4 (including 5.4.1).

[0254] Various embodiments described in this section (Section 5.4.3) using a nicking endonuclease and / or a restriction site for the nicking endonuclease are further provided where the nicking endonuclease is replaced with a programmable nicking enzyme and the restriction site is replaced with a target site for the programmable nicking enzyme. Programmable nicking enzymes and their target sites for purposes of this paragraph and this section (Section 5.4.3) are provided in Section 5.3.4.

[0255] 5.4.5 Viral DNA Sequence Characteristics Not Present in the DNA Molecules Provided Herein As further described in Sections 3, 5.2, 5.4.1, 5.4.2, 5.4.3, 5.4.6, 5.4.7, and 5.5, the provided DNA molecules can be produced either synthetically or recombinantly, with or without specific sequence elements or features. Accordingly, certain appropriate and desired sequence features or elements can be included in, or excluded from, the DNA molecules provided herein. Corresponding methods for making such DNA molecules, with or without sequence features or elements, are also provided herein, as described by applying the methods of 5.2 with the DNA molecules of 5.4, thereby producing the various DNA molecules described in 5.5.

[0256] As described in Sections 3, 5.4.1, 5.6, and 6, such DNA sequence elements or features that can be excluded from the DNA molecules provided herein may be viral replication-associated protein binding sequences ("RABS"), which refer to DNA sequences to which viral DNA replication-associated proteins and their isoforms encoded by the Parvoviridae genes Rep and NS1 can bind. RABS refers to a nucleotide sequence that contains both a nucleotide sequence recognized by the Rep or NS1 protein (for replication of viral nucleic acid molecules) and a site of specific interaction between the Rep or NS1 protein and the nucleotide sequence. RABS can be a sequence of 5 to 300 nucleotides.In some embodiments of the DNA molecules provided herein, including those provided in this Section 5.4.5, the RABS is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, The sequence may be at least 200, at least 205, at least 210, at least 215, at least 220, at least 225, at least 230, at least 235, at least 240, at least 245, at least 250, at least 255, at least 260, at least 265, at least 270, at least 275, at least 280, at least 285, at least 290, at least 295, at least 300, at least 305, at least 310, at least 315, at least 320, at least 325, at least 330, at least 335, at least 340, at least 345, at least 350, at least 355, at least 360, at least 365, at least 370, at least 375, at least 380, at least 385, at least 390, at least 395, or at least 400 nucleotides.In some other embodiments, the RABS is about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, about 200, about 205, about 210, about 215, about 220, about 225, about 230, about 235, about 240, about 245, about 250, about 255, about 260, about 265, about 270, about 275, about 280, about 285, about 290, about 300, about 310, about 320, about 330, about 340, about 350, about 355, about 360, about 365, about 370, about 375, about 380, about 385, about 390, about 395, about 400, about 410, about 420, about 430, about 440, about 450, about 460, about 470, about 480, about 490, about 510, about 520, about 530, about 540 The sequence may be 0, about 215, about 220, about 225, about 230, about 235, about 240, about 245, about 250, about 255, about 260, about 265, about 270, about 275, about 280, about 285, about 290, about 295, about 300, about 305, about 310, about 315, about 320, about 325, about 330, about 335, about 340, about 345, about 350, about 355, about 360, about 365, about 370, about 375, about 380, about 385, about 390, about 395 or about 400 nucleotides. In some further embodiments, any embodiment of a DNA molecule lacking RABS described in this paragraph can be combined with any method or DNA molecule provided herein, including those provided in Sections 3, 5.2, 5.4, 5.5, and 6.

[0257] Alternatively, the DNA molecules provided herein, including those in Sections 3, 5.2, 5.4, 5.5, and 6, may lack a functional RABS by functionally inactivating a RABS sequence present in the DNA molecule using a mutation, insertion, deletion (including partial deletion or truncation) such that the RABS can no longer serve as a Rep protein or NS1 protein recognition and / or binding site. Thus, in some embodiments of the DNA molecules provided herein, including those in Sections 3, 5.2, 5.4, 5.5, and 6, the DNA molecule comprises a functionally inactivated RABS. Such functional inactivation can be assessed by measuring and comparing binding between a Rep or NS1 protein and a DNA molecule comprising a functionally inactivated RABS to binding between the Rep or NS1 protein and a reference molecule comprising a wild-type (wt) RBS or NSBE sequence (e.g., an identical DNA molecule except for having a wt RBS or wt NSBE sequence). Such binding can be determined by any binding assay known and used in the field of molecular biology, for example, a chromatin immunoprecipitation (ChIP) assay, a DNA electrophoretic mobility shift assay (EMSA), a DNA pull-down assay, or a microplate capture and detection assay, as further described in Matthew J. Guille and G. Geoff Kneale, Molecular Biotechnology 8:35-52 (1997); Bipasha Dey et al., Mol Cell Biochem. 2012 Jun;365(1-2):279-99, both of which are incorporated herein by reference in their entireties.In one embodiment, the binding between the RAP and a functionally inactivated RABS in a DNA molecule is at most 0.001%, at most 0.01%, at most 0.1%, at most 1%, at most 1.5%, at most 2%, at most 2.5%, at most 3%, at most 3.5, at most 4%, at most 4.5%, at most 5%, at most 5.5%, at most 6%, at most 6.5%, at most 7%, at most 7.5%, at most 8%, at most 8.5%, at most 9%, at most 9.5%, or at most 10%, compared to the binding between the RAP and a wild-type RBS or NSBE in a reference DNA molecule (e.g., an identical DNA molecule except for having a wild-type RBS or NSBE sequence). In another embodiment, the binding between the RAP and the functionally inactivated RABS in the DNA molecule is about 0.001%, about 0.01%, about 0.1%, about 1%, about 1.5%, about 2%, about 2.5%, about 3%, about 3.5, about 4%, about 4.5%, about 5%, about 5.5%, about 6%, about 6.5%, about 7%, about 7.5%, about 8%, about 8.5%, about 9%, about 9.5%, or about 10% compared to the binding between the RAP and the wild-type RABS in a reference DNA molecule (e.g., an identical DNA molecule except for having a wt RBS or NSBE sequence). In yet another embodiment, the binding between the RAP and a functionally inactivated RABS in a DNA molecule is 0.001%, 0.01%, 0.1%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10% compared to the binding between the RAP and a wild-type RABS in a reference DNA molecule (e.g., an identical DNA molecule except for having a wt RBS or NSBE sequence).

[0258] Additionally, the DNA molecules provided herein, including those in Sections 3, 5.2, 5.4, 5.5, and 6, may lack a functional RAP or viral capsid coding sequence by functionally inactivating the Rep protein, NS1, or viral capsid coding sequence present in the DNA molecule using a mutation, insertion, deletion (including partial deletion or truncation) such that the RAP or viral capsid coding sequence is no longer capable of functionally expressing the Rep protein, NS1 protein, or viral capsid protein. Such functionally inactivating mutations, insertions, or deletions can be achieved, for example, by using mutations, insertions, and / or deletions that shift the open reading frame of the Rep protein or viral capsid coding sequence, by using mutations, insertions, and / or deletions that remove the start codon, by using mutations, insertions, and / or deletions that remove the promoter or transcription start point, by using mutations, insertions, and / or deletions that remove RNA polymerase binding sites, by using mutations, insertions, and / or deletions that remove ribosome recognition or binding sites, or by other means known and used in the art.

[0259] In one embodiment, the DNA molecule comprises an RBS that has been inactivated by a mutation. In one embodiment, the DNA molecule comprises an RBS that has been inactivated by a mutation of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 10, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in the RBS. In another embodiment, the DNA molecule comprises an RBS that is inactivated by mutation of 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 10%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleotides in the RBS. In further embodiments, the DNA molecule comprises an RBS that has been inactivated by a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 10, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in the RBS. In yet another embodiment, the DNA molecule comprises an RBS that has been inactivated by deletion of 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 10%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleotides at the RBS. In some embodiments, the deletion in the preceding sentence is an internal deletion, a deletion from the 5' end, or a deletion from the 3' end. In some embodiments, the deletion in this paragraph can be any combination of internal deletions, deletions from the 5' end, and / or deletions from the 3' end. In some embodiments, the DNA molecule comprises an RBS that has been inactivated by deletion of the entire RBS sequence.In some additional embodiments, the DNA molecule comprises an RBS that has been inactivated by a partial deletion of the RBS sequence.

[0260] In one embodiment, the DNA molecule comprises an NSBE that has been inactivated by a mutation. In one embodiment, the DNA molecule comprises an NSBE that has been inactivated by a mutation of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 10, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in the NSBE. In another embodiment, the DNA molecule comprises an NSBE that has been inactivated by mutation of 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 10%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleotides in the NSBE. In further embodiments, the DNA molecule comprises an NSBE that has been inactivated by a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 10, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in the NSBE. In yet another embodiment, the DNA molecule comprises an NSBE that has been inactivated by deletion of 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 10%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleotides in the NSBE. In some embodiments, the deletion in the preceding sentence is an internal deletion, a deletion from the 5' end, or a deletion from the 3' end. In some embodiments, the deletion in this paragraph can be any combination of internal deletion, deletion from the 5' end, and / or deletion from the 3' end. In certain embodiments, the DNA molecule comprises an NSBE that has been inactivated by deletion of the entire NSBE sequence.In some additional embodiments, the DNA molecule comprises an NSBE that has been inactivated by a partial deletion of the NSBE sequence.

[0261] Similarly, elements or features of a DNA sequence can be included in or excluded from any particular region of the DNA molecules provided herein (e.g., Sections 5.4 and 5.5) or any particular region of a DNA molecule used in the methods provided herein (e.g., Section 5.2). In one embodiment, the DNA molecule lacks a Rep protein coding sequence. In one embodiment, the DNA molecule lacks an NS1 protein coding sequence. In another embodiment, the DNA molecule lacks a viral capsid protein coding sequence. In some embodiments, the expression cassette lacks a Rep protein coding sequence. In some embodiments, the expression cassette lacks an NS1 protein coding sequence. In certain embodiments, the expression cassette lacks a viral capsid protein coding sequence. In a further embodiment, the DNA molecule lacks a RABS. In yet another embodiment, the first inverted repeat lacks a RABS. In one embodiment, the second inverted repeat lacks a RABS. In another embodiment, the DNA sequence between the ITR-closing base pair of the first inverted repeat and the ITR-closing base pair of the second inverted repeat lacks a RABS. In one embodiment, the DNA molecule comprises a functionally inactivated Rep protein coding sequence. In one embodiment, the DNA molecule comprises a functionally inactivated NS1 protein coding sequence. In another embodiment, the DNA molecule comprises a functionally inactivated viral capsid protein coding sequence. In some embodiments, the expression cassette comprises a functionally inactivated Rep protein coding sequence. In some embodiments, the expression cassette comprises a functionally inactivated NS1 protein coding sequence. In some embodiments, the expression cassette comprises a functionally inactivated viral capsid protein coding sequence. In a further embodiment, the DNA molecule comprises a functionally inactivated RABS. In yet another embodiment, the first inverted repeat comprises a functionally inactivated RABS. In one embodiment, the second inverted repeat comprises a functionally inactivated RABS.In another embodiment, the DNA sequence between the ITR closing base pair of said first inverted repeat and the ITR closing base pair of said second inverted repeat comprises a functionally inactivated RABS.

[0262] Furthermore, DNA sequence elements or features can be functionally inactivated from any specific region of the DNA molecules provided herein (e.g., Sections 5.4 and 5.5) or any combination of specific regions of the DNA molecules used in the methods provided herein (e.g., Section 5.2). In one embodiment, the first inverted repeat comprises a functionally inactivated RABS, and the second inverted repeat comprises a functionally inactivated RABS. In another embodiment, the first inverted repeat comprises a functionally inactivated RABS, and the DNA sequence between the ITR-closing base pair of the first inverted repeat and the ITR-closing base pair of the second inverted repeat comprises a functionally inactivated RABS. In a further embodiment, the second inverted repeat comprises a functionally inactivated RABS, and the DNA sequence between the ITR-closing base pair of the first inverted repeat and the ITR-closing base pair of the second inverted repeat comprises a functionally inactivated RABS. In yet another embodiment, the first inverted repeat comprises a functionally inactivated RABS, the second inverted repeat comprises a functionally inactivated RBS, and the DNA sequence between the ITR closing base pair of said first inverted repeat and the ITR closing base pair of said second inverted repeat comprises a functionally inactivated RABS.

[0263] As described in Sections 3, 5.4.1, 5.6, and 6, one such DNA sequence element or feature that can be excluded from the DNA molecules provided herein can be a terminal separation site ("TRS"). A TRS refers to a nucleotide sequence within an inverted repeat of a DNA molecule that contains a nucleotide sequence recognized by RAP (for replication of a viral nucleic acid molecule), a site of specific interaction between the RAP and the nucleotide sequence, and a site of specific cleavage by the endonuclease activity of the RAP protein. The nucleotide sequence of the conserved site of specific cleavage by the endonuclease activity of a RAP protein can be determined by DNA nicking assays known and used in the field of molecular biology, such as gel electrophoresis, fluorophore-based in vitro nicking assays, and radioactive in vitro nicking assays, as further described in Xu P et al., 2019. Antimicrob Agents Chemother 63:e01879-18; US20190203229A, both of which are incorporated herein by reference in their entireties. In some embodiments, a TRS can be a nucleotide sequence within an inverted repeat of a DNA molecule, comprising a nucleotide sequence recognized by a Rep protein (responsible for replicating viral nucleic acid molecules), a site of specific interaction between the Rep protein and the nucleotide sequence, and a site of specific cleavage by the endonuclease activity of the Rep protein. In one embodiment, a TRS can be a nucleotide sequence within an inverted repeat of a DNA molecule, comprising a nucleotide sequence recognized by an NS1 protein (responsible for replicating viral nucleic acid molecules), a site of specific interaction between the NS1 protein and the nucleotide sequence, and a site of specific cleavage by the endonuclease activity of the NS1 protein. A TRS can be a sequence of 5 to 300 nucleotides.In some embodiments of the methods provided herein, such as those provided in this Section 5.4.5, the TRS is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, at least The sequence may be at least 200, at least 205, at least 210, at least 215, at least 220, at least 225, at least 230, at least 235, at least 240, at least 245, at least 250, at least 255, at least 260, at least 265, at least 270, at least 275, at least 280, at least 285, at least 290, at least 295, at least 300, at least 305, at least 310, at least 315, at least 320, at least 325, at least 330, at least 335, at least 340, at least 345, at least 350, at least 355, at least 360, at least 365, at least 370, at least 375, at least 380, at least 385, at least 390, at least 395, or at least 400 nucleotides.In some other embodiments, the TRS is about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, about 200, about 205, about 210 about 215, about 220, about 225, about 230, about 235, about 240, about 245, about 250, about 255, about 260, about 265, about 270, about 275, about 280, about 285, about 290, about 295, about 300, about 305, about 310, about 315, about 320, about 325, about 330, about 335, about 340, about 345, about 350, about 355, about 360, about 365, about 370, about 375, about 380, about 385, about 390, about 395, or about 400 nucleotides. In some further embodiments, any embodiment of a TRS described in this paragraph can be combined with any method or DNA molecule provided herein, such as those provided in Sections 3, 5.2, 5.4, 5.5, and 6.

[0264] Alternatively, the DNA molecules provided herein, including those in Sections 3, 5.2, 5.4, 5.5, and 6, may lack a functional TRS by functionally inactivating a TRS sequence present in the DNA molecule using mutation, insertion, deletion (including partial deletion or truncation) such that the TRS can no longer serve as a recognition and / or binding site for RAP (i.e., Rep and NS1). Thus, in some embodiments of the DNA molecules provided herein, including those in Sections 3, 5.2, 5.4, 5.5, and 6, the DNA molecule comprises a functionally inactivated TRS. Such functional inactivation can be assessed by measuring and comparing binding between RAP (i.e., Rep and NS1) and a DNA molecule comprising a functionally inactivated TRS to binding between RAP and a reference molecule comprising a wild-type (wt) TRS sequence (e.g., an identical DNA molecule except for having the wt TRS sequence). Such binding can be determined by any binding assay known and used in the field of molecular biology, for example, a chromatin immunoprecipitation (ChIP) assay, a DNA electrophoretic mobility shift assay (EMSA), a DNA pull-down assay, or a microplate capture and detection assay, as further described in Matthew J. Guille and G. Geoff Kneale, Molecular Biotechnology 8:35-52 (1997); Bipasha Dey et al., Mol Cell Biochem. 2012 Jun;365(1-2):279-99, both of which are incorporated herein by reference in their entireties.In one embodiment, the binding between the RAP (i.e., Rep and NS1) and the functionally inactivated TRS in a DNA molecule is at most 0.001%, at most 0.01%, at most 0.1%, at most 1%, at most 1.5%, at most 2%, at most 2.5%, at most 3%, at most 3.5, at most 4%, at most 4.5%, at most 5%, at most 5.5%, at most 6%, at most 6.5%, at most 7%, at most 7.5%, at most 8%, at most 8.5%, at most 9%, at most 9.5%, or at most 10%, compared to the binding between the RAP (i.e., Rep and NS1) and the wild-type TRS in a reference DNA molecule (e.g., an identical DNA molecule except for having the wt TRS sequence). In another embodiment, the binding between the RAP (i.e., Rep and NS1) and the functionally inactivated TRS in a DNA molecule is about 0.001%, about 0.01%, about 0.1%, about 1%, about 1.5%, about 2%, about 2.5%, about 3%, about 3.5, about 4%, about 4.5%, about 5%, about 5.5%, about 6%, about 6.5%, about 7%, about 7.5%, about 8%, about 8.5%, about 9%, about 9.5%, or about 10% compared to the binding between the RAP (i.e., Rep and NS1) and the wild-type TRS in a reference DNA molecule (e.g., an identical DNA molecule except for having the wt TRS sequence). In yet another embodiment, the binding between the RAP (i.e., Rep and NS1) and the functionally inactivated TRS in a DNA molecule is 0.001%, 0.01%, 0.1%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10% compared to the binding between the RAP (i.e., Rep and NS1) and the wild-type TRS in a reference DNA molecule (e.g., an identical DNA molecule except for having the wt TRS sequence).

[0265] In one embodiment, the DNA molecule comprises a TRS that has been inactivated by a mutation. In one embodiment, the DNA molecule comprises a TRS that has been inactivated by a mutation of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 10, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in the TRS. In another embodiment, the DNA molecule comprises a TRS that is inactivated by mutation of 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 10%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleotides in the TRS. In further embodiments, the DNA molecule comprises a TRS that has been inactivated by a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 10, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in the TRS. In yet another embodiment, the DNA molecule comprises a TRS that has been inactivated by deletion of 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 10%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleotides in the TRS. In some embodiments, the deletion in the preceding sentence is an internal deletion, a deletion from the 5' end, or a deletion from the 3' end. In some embodiments, the deletion in this paragraph can be any combination of internal deletions, deletions from the 5' end, and / or deletions from the 3' end. In certain embodiments, the DNA molecule comprises a TRS that has been inactivated by deletion of the entire TRS sequence.In some additional embodiments, the DNA molecule comprises a TRS that has been inactivated by a partial deletion of the TRS sequence.

[0266] Similarly, DNA sequence elements or features can be included in or excluded from any particular region of the DNA molecules provided herein (e.g., Sections 5.4 and 5.5) or any particular region of the DNA molecules used in the methods provided herein (e.g., Section 5.2). In one embodiment, the DNA molecule lacks a TRS. In yet another embodiment, the first inverted repeat lacks a TRS. In another embodiment, the second inverted repeat lacks a TRS. In a further embodiment, the first inverted repeat lacks a TRS and the second inverted repeat lacks a TRS.

[0267] Alternatively, a TRS sequence element or feature can be functionally inactivated from any particular region of a DNA molecule provided herein (e.g., Sections 5.4 and 5.5) or any particular region of a DNA molecule used in a method provided herein (e.g., Section 5.2). In one embodiment, the DNA molecule comprises a functionally inactivated TRS. In yet another embodiment, the first inverted repeat comprises a functionally inactivated TRS. In another embodiment, the second inverted repeat comprises a functionally inactivated TRS. In a further embodiment, the first inverted repeat comprises a functionally inactivated TRS and the second inverted repeat comprises a functionally inactivated TRS.

[0268] In some specific embodiments, the RBS that is eliminated or functionally inactivated in the DNA molecules provided herein can be any, any combination of any number, or all of the RBS sequences listed in Table 20. Table 20: Exemplary RAPs [Table 20]

[0269] In one particular embodiment, the DNA molecule lacks the coding sequence for any one, any combination of any number, or all of the RAPs listed in the table in the previous paragraph. In another specific embodiment, the DNA molecule comprises a functionally inactivated sequence encoding any one, any combination of any number, or all of the RAPs listed in the table in the previous paragraph.

[0270] In another specific embodiment, the TRS that is eliminated or functionally inactivated in the DNA molecules provided herein can be any, or any combination of any number of, or all of the TRS sequences listed in Table 21. Table 21: Exemplary RAPs [Table 21]

[0271] Because the methods provided herein do not require a viral replication step and because the DNA molecules provided herein do not need to be produced or replicated in the viral life cycle, the present disclosure provides, and readers of this disclosure will understand, that the DNA molecules provided herein can lack various DNA sequences or features, including those provided in this section (Section 5.4.5). DNA molecules lacking RABS and / or TRS as provided in this Section 5.4.5 and DNA molecules containing functionally inactivated RABS and / or functionally inactivated TRS offer at least a significant advantage over DNA molecules containing such RABS and / or TRS sequences in that they have no or significantly reduced risk of mobilization or replication when administered to a patient. Mobilization risk or mobilization risk refers to the risk that a replication-deficient DNA molecule will revert to replicating or producing viral particles in a host to which the DNA molecule is administered. Such mobilization risk may result from the presence of viral proteins (e.g., Rep protein, NS1 protein, or viral capsid protein) expressed by viruses infecting the same host as the host to which the DNA molecule is administered. Mobilization risk poses significant safety concerns regarding the use of replication-deficient viral genomes as gene therapy vectors, as described, for example, in Liujiang Song, Hum Gene Ther, 2020 Oct;31(19-20):1054-1067, incorporated herein by reference in its entirety. DNA molecules lacking such RBS and / or TRS will not have a binding site for viral Rep protein to initiate replication, even if another helper virus is present in the same host to provide Rep protein.

[0272] Thus, in some embodiments of the DNA molecules provided herein, including those in this Section 5.4.5, DNA molecules that do not comprise a RABS and / or do not comprise a TRS have a lower risk of mobilization after administration to a subject or patient, compared to DNA molecules that comprise a RABS and / or a TRS. In certain embodiments of the DNA molecules provided herein, including those in this Section 5.4.5, DNA molecules that comprise a functionally inactivated RABS and / or a functionally inactivated TRS have a lower risk of mobilization after administration to a subject or patient, compared to DNA molecules that comprise a RABS and / or a TRS. Such a reduced risk of mobilization can be determined as (Pm-Po) / Pm, where Pm is the number of viral particles produced from a control DNA molecule that comprises an RBS when a RAP is present (e.g., due to infection with any virus that contains a RAP in the same host or engineered expression of a RAP); Po is the number of viral particles produced from a DNA molecule that lacks or comprises a functionally inactivated RABS provided herein, under comparable conditions in the same host as used for the control DNA molecule. Alternatively, such a reduction in mobilization risk can be determined as (Pm-Po) / Pm, where Pm is the number of viral particles produced from a control DNA molecule containing a TRS when RAP is present (e.g., due to infection with any virus containing a Rep protein in the same host or engineered expression of a Rep protein); Po is the number of viral particles produced from a DNA molecule lacking a TRS or containing a functionally inactivated TRS provided herein under similar conditions in the same host as used for the control DNA molecule.Furthermore, such a reduction in mobilization risk can be determined as (Pm-Po) / Pm, where Pm is the number of viral particles produced from a control DNA molecule containing a RABS and containing a TRS when a RABS is present (e.g., due to infection with any virus containing a Rep protein, an NS1 protein, or engineered expression of a Rep protein in the same host); Po is the number of viral particles produced from a DNA molecule (i) lacking a RABS or containing a functionally inactivated RABS, and (ii) lacking a TRS or containing a functionally inactivated TRS, as provided herein, under comparable conditions in the same host as used for the control DNA molecule. As described in Liujiang Song, Hum Gene Ther, 2020 Oct;31(19-20):1054-1067, incorporated herein by reference in its entirety, the host used to determine the number of particles produced can be a cell, an animal (e.g., a mouse, hamster, rat, dog, rabbit, guinea pig, and other suitable mammals), or a human. The present disclosure further provides, and those skilled in the art will understand upon reading this disclosure, that Pm and Po, each as described in this paragraph, can also be used to determine absolute or relative levels of mobilization. Briefly, in such assays, a DNA molecule is transduced into host cells (e.g., HEK293 cells) by transfecting the host cells or infecting them with viral particles containing the DNA molecule. The host cells are further transfected with Rep protein, NS1 protein, or co-infected with another virus (e.g., wild-type virus) expressing Rep protein or NS1 protein. The host cells are then cultured to produce and release viral particles. After 48 to 72 hours (e.g., 65 hours) of culture, virions are then harvested by collecting both the host cells and the culture medium.Viral particle titers (surrogates for Pm and Po) can be determined by probe-based quantitative PCR (qPCR) analysis after benzonase treatment to remove non-encapsidated DNA, as described in Song et al., Cytotherapy 2013;15:986-998, which is incorporated by reference in its entirety. An exemplary implementation of such an assay is provided in Liujiang Song, Hum Gene Ther 2020 Oct;31(19-20):1054-1067, which is incorporated by reference in its entirety.

[0273] Based on the determination of the reduced mobilization risk and mobilization risk level, in some embodiments of the DNA molecules provided herein, including those in this Section 5.4.5, the mobilization risk of the DNA molecule, when administered to a host, is 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 78%, 79 ... 6%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, 50%, 49%, 48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, or 20% lower.In certain embodiments, the mobilization risk of the DNA molecule when administered to a host is at least 99%, at least 98%, at least 97%, at least 96%, at least 95%, at least 94%, at least 93%, at least 92%, at least 91%, at least 90%, at least 89%, at least 88%, at least 87%, at least 86%, at least 85%, at least 84%, at least 83%, at least 82%, at least 81%, at least 80%, at least 79%, at least 78%, at least 77%, at least 76%, at least 75%, at least 74%, at least 73%, at least 72%, at least 71%, at least 70%, at least 69%, at least 68%, at least 67%, at least 66%, at least 65%, at least 64%, at least 63%, at least 62%, at least 61%, at least 60%, at least 59%, at least 58%, at least 57%, at least 56%, at least 55%, at least 54%, at least 53%, at least 52%, at least 51%, at least 50%, at least 49%, at least 48%, at least 47%, at least 46%, at least 45%, at least 44%, at least 43%, at least 42%, at least 41%, at least 40%, at least 39%, at least 38%, at least 37%, at least 36%, at least 35%, at least 34%, at least 33%, at least 32%, at least 31%, at least 30%, at least 29%, at least 28%, at least 27%, at least 26%, at least 25%, at least 24%, at least 23%, at least 22%, at least 21%, or at least 20% lower.In another embodiment, the risk of mobilization of the DNA molecule when administered to a host is about 100%, about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 79%, about 89%, about 90%, about 91 ... 66%, about 65%, about 64%, about 63%, about 62%, about 61%, about 60%, about 59%, about 58%, about 57%, about 56%, about 55%, about 54%, about 53%, about 52%, about 51%, about 50%, about 49%, about 48%, about 47%, about 46%, about 45%, about 44%, about 43%, about 42%, about 41%, about 40%, about 39%, about 38%, about 37%, about 36%, about 35%, about 34%, about 33%, about 32%, about 31%, about 30%, about 29%, about 28%, about 27%, about 26%, about 25%, about 24%, about 23%, about 22%, about 21% or about 20% lower.

[0274] Alternatively, in one embodiment, the DNA molecules provided herein, including those in this Section 5.4.5, do not result in detectable mobilization (e.g., based on measurements of Po as provided in this Section 5.4.5). In another embodiment, the DNA molecules provided herein, including those in this Section 5.4.5, result in mobilization of 0.0001% or less, 0.001% or less, 0.01% or less, 0.1% or less, 1% or less, 1.5% or less, 2% or less, 2.5% or less, 3% or less, 3.5% or less, 4% or less, 4.5% or less, 5% or less, 5.5% or less, 6% or less, 6.5% or less, 7% or less, 7.5% or less, 8% or less, 8.5% or less, 9% or less, 9.5% or less, or 10% or less of the mobilization resultant by a reference DNA molecule (e.g., an identical DNA molecule except for having a wild-type RABS and / or a wild-type TRS sequence). In further embodiments, the DNA molecules provided herein, including those in this Section 5.4.5, provide about 0.0001%, about 0.001%, about 0.01%, about 0.1%, about 1%, about 1.5%, about 2%, about 2.5%, about 3%, about 3.5, about 4%, about 4.5%, about 5%, about 5.5%, about 6%, about 6.5%, about 7%, about 7.5%, about 8%, about 8.5%, about 9%, about 9.5%, or about 10% of the mobilization provided by a reference DNA molecule (e.g., an identical DNA molecule except for having a wild-type RABS and / or a wild-type TRS sequence). In yet another embodiment, the DNA molecules provided herein, including those in this Section 5.4.5, mobilize 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10% of the mobilization mediated by a reference DNA molecule (e.g., an identical DNA molecule except for having a wild-type RABS and / or a wild-type TRS sequence). Such percentages of mobilization can be determined by using Pm and Po determined as further described in the previous paragraph (including the previous two paragraphs).

[0275] As will be apparent from the description in this Section 5.4.5, the DNA sequences or features excluded in the DNA molecules provided herein can be combined in any manner with any of the methods provided herein (e.g., Sections 3, 5.2, and 6), any of the DNA molecules provided herein (e.g., Sections 3, 5.4, and 6), and any of the hairpin-ended DNA molecules provided herein (e.g., Sections 3, 5.5, and 6), and contribute to the functional properties of the DNA molecules provided herein (e.g., Sections 3, 5.6, and 6).

[0276] (5.4.6 Plasmids and other vectors) The present disclosure provides that DNA molecules can be in various forms. In one embodiment, the DNA molecules provided for the methods and compositions herein are vectors. A vector is a nucleic acid molecule that can be replicated and / or expressed in a host cell. Any vector known to those skilled in the art is provided herein. In some embodiments, the vector can be a plasmid, a viral vector, a cosmid, and an artificial chromosome (e.g., a bacterial artificial chromosome or a yeast artificial chromosome). In a particular embodiment, the vector is a plasmid. As will be apparent from the description, when a DNA molecule is in the form of a vector (including a plasmid), the vector will have all the characteristics described herein for a DNA molecule, including those described in Section 3 and this section (Section 5.4).

[0277] In some embodiments, the vectors provided in this section (Section 5.4.6) can be used to produce the DNA molecules provided in Sections 3 and 5.5, e.g., by performing the steps of the methods provided in Section 5.2. Thus, the vectors provided in this section (Section 5.4.6) (1) comprise features of the DNA molecules provided in Sections 3 and 5.5, including IRs or ITRs capable of forming a hairpin as described in Sections 5.4.1 and 5.5, an expression cassette as described in 5.4.3, and restriction sites for nicking endonucleases or restriction enzymes as described in Sections 5.4.2, 5.3.4, and 5.4.7, and / or (2) lack RABS and / or TRS sequences as described in Section 5.4.5. Thus, the present disclosure provides that the vectors provided in this section (Section 5.4.6) can include any combination of (1) IRs or ITRs capable of forming a hairpin as described in Sections 5.4.1 and 5.5, an expression cassette as described in 5.4.3, restriction sites for nicking endonucleases or restriction enzymes as described in Sections 5.4.2, 5.3.4, and 5.4.7, and additional feature embodiments for vectors provided in this section (Section 5.4.6), and / or (2) lack RABS and / or TRS sequences as described in Section 5.4.5. In some embodiments, vectors can be constructed using known techniques that provide, as operably linked components, at least the following, in the direction of transcription: (1) 5' ITR sequence; (2) an expression cassette including cis-regulatory elements, e.g., promoters, inducible promoters, regulatory switches, enhancers, etc.; and (3) 3' IR sequence. In some embodiments, the expression cassette is flanked by ITRs and includes cloning sites for introducing exogenous sequences.

[0278] Specifically, in one embodiment, the DNA molecule is a plasmid. Plasmids are widely known and used in the art as vectors for replicating or expressing the DNA molecules contained therein. Plasmids often refer to double-stranded and / or circular DNA molecules that can autonomously replicate in suitable host cells. The plasmids provided for the methods and compositions described herein are available from various suppliers and / or include commercially available plasmids for use in well-known host cells (including both prokaryotic and eukaryotic host cells), such as those described in Michael Green and Joseph Sambrook's Molecular Cloning: A Laboratory Manual, 4th Edition, ISBN 978-1-936113-42-2 (2012), the entire contents of which are incorporated herein by reference.

[0279] The plasmids described in this section (Section 5.4.6) can further include other features. In some embodiments, the plasmid further comprises a restriction enzyme site (e.g., a restriction enzyme site as described in Sections 5.3.4 and 5.4.2) in the region 5' to the first inverted repeat and 3' to the second inverted repeat, wherein the restriction enzyme site is not present in any of the first inverted repeat, the second inverted repeat, or the region between the first and second inverted repeats. In certain embodiments, cleavage with a restriction enzyme at a restriction site described in this paragraph results in single-stranded overhangs that do not anneal at a detectable level under conditions suitable for annealing of the first inverted repeat and / or the second inverted repeat (e.g., conditions as described in Section 5.3.5). In some other embodiments, the plasmid further comprises an open reading frame encoding a restriction enzyme that recognizes and cleaves a restriction site described in this paragraph. In certain embodiments, the restriction enzyme site and corresponding restriction enzyme can be any one of the restriction enzyme sites and corresponding restriction enzymes described in Sections 5.3.4 and 5.4.2. In further embodiments, expression of the restriction enzyme described in this paragraph is under the control of a promoter. In some embodiments, the promoter described in this paragraph can be any promoter described above in Section 5.4.3. In another embodiment, the promoter described is an inducible promoter. In certain embodiments, the inducible promoter is a chemically inducible promoter.In further embodiments, the inducible promoter is any one selected from the group consisting of the tetracycline ON (Tet-On) promoter, the negative inducible pLac promoter, alcA, amyB, bli-3, bphA, catR, cbhl, cre1, exylA, gas, glaA, gla1, mir1, niiA, qa-2, Smxyl, tcu-1, thiA, vvd, xyl1, xyl1, xylP, xyn1, and ZeaR, as described in Janina Kluge et al., Applied Microbiology and Biotechnology, 102: 6357-6372 (2018), the entire contents of which are incorporated herein by reference.

[0280] Similarly, in certain embodiments, the plasmid can further comprise fifth and sixth restriction sites for a nicking endonuclease in the region 5' to the first inverted repeat and 3' to the second inverted repeat (e.g., restriction sites for a nicking endonuclease as described in Sections 5.3.4 and 5.4.2), wherein the fifth and sixth restriction sites for a nicking endonuclease: a.) are on opposite strands; and b.) cause the cleavage to occur within the double-stranded DNA molecule such that the single-stranded overhangs of the cleavage do not inter- or intramolecularly anneal at a detectable level under conditions suitable for annealing of the first inverted repeat and / or the second inverted repeat (e.g., conditions as described in Section 5.3.5). As is apparent from the description in Section 5.3.4, incubation with a nicking endonuclease results in a fifth nick corresponding to a fifth restriction site for the nicking endonuclease and a sixth nick corresponding to a sixth restriction site for the nicking endonuclease. The present disclosure provides that the fifth and sixth nicks can be at various relative positions between them. In one embodiment, the fifth and sixth nicks are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides apart. In some embodiments, the ssDNA overhang resulting from the fifth and sixth nicks has a lower melting temperature than the ssDNA overhangs described in Sections 5.3.3 and 5.4.2 because the ssDNA overhang between the fifth and sixth nicks does not detectably inter- or intramolecularly anneal under conditions suitable for annealing of the first inverted repeat and / or the second inverted repeat. In certain embodiments, the ssDNA overhang resulting from the fifth and sixth nicks is shorter than the ssDNA overhangs described in Sections 5.3.3 and 5.4.2. In other embodiments, the ssDNA overhang resulting from the fifth and sixth nicks has a lower GC percentage than the ssDNA overhangs described in Sections 5.3.3 and 5.4.2.In some particular embodiments, the ssDNA overhangs resulting from the fifth and sixth nicks are 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length.

[0281] In certain embodiments, the plasmid can further comprise 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or more restriction sites for nicking endonucleases (e.g., restriction sites for nicking endonucleases as described in Sections 5.3.4 and 5.4.2) in the region 5' to the first inverted repeat and 3' to the second inverted repeat, where the additional restriction sites for nicking endonucleases: a.) are on opposite strands; and b.) cause the cleavage to occur within the double-stranded DNA molecule such that the single-stranded overhangs of the cleavage do not inter- or intramolecularly anneal at a detectable level under conditions suitable for annealing of the first inverted repeat and / or the second inverted repeat (e.g., conditions as described in Section 5.3.5). The present disclosure provides that the nicks in the region 5' to the first inverted repeat and 3' to the second inverted repeat can have various relative positions therebetween. In one embodiment, the nicks in the region 5' to the first inverted repeat and 3' to the second inverted repeat are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides apart. In some embodiments, the ssDNA overhangs resulting from nicks in the region 5' to the first inverted repeat and 3' to the second inverted repeat have a lower melting temperature than the ssDNA overhangs described in Sections 5.3.3 and 5.4.2 because the ssDNA overhangs between the nicks in the region 5' to the first inverted repeat and 3' to the second inverted repeat do not undergo detectable inter- or intramolecular annealing under conditions suitable for annealing of the first inverted repeat and / or the second inverted repeat. In certain embodiments, the ssDNA overhangs resulting from nicks in the region 5' to the first inverted repeat and 3' to the second inverted repeat are shorter than the ssDNA overhangs described in Sections 5.3.3 and 5.4.2.In another embodiment, the ssDNA overhang resulting from the nick in the region 5' to the first inverted repeat and 3' to the second inverted repeat has a lower percentage GC content than the ssDNA overhangs described in Sections 5.3.3 and 5.4.2. In some particular embodiments, the ssDNA overhang resulting from the nick in the region 5' to the first inverted repeat and 3' to the second inverted repeat is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length.

[0282] As described above in Sections 5.3.4 and 5.4.2, in various embodiments, the first, second, third, and fourth restriction sites for nicking endonucleases can be target sequences for the same or different nicking endonucleases. Similarly, in certain embodiments, the fifth and sixth restriction sites for nicking endonucleases can be target sequences for the same or different nicking endonucleases. In some embodiments, the first, second, third, fourth, fifth, and sixth restriction sites for nicking endonucleases provided on a DNA molecule as described in Sections 3 and 5.3.4 and this Section 5.4 can all be target sequences for the same nicking endonuclease. Alternatively, in another embodiment, the first, second, third, fourth, fifth, and sixth numbered restriction sites for nicking endonucleases are target sequences for two different nicking endonucleases, including all possible combinations for assigning the six sites to two different nicking endonuclease target sequences (e.g., the first restriction site to the first nicking endonuclease and the remainder to the second nicking endonuclease, the first and second restriction sites to the first nicking endonuclease and the remainder to the second nicking endonuclease, etc.). Furthermore, in certain embodiments, the first, second, third, fourth, fifth, and sixth restriction sites for nicking endonucleases are target sequences for three different nicking endonucleases, including all possible combinations for assigning the six sites to three different nicking endonuclease target sequences. Furthermore, in some embodiments, the first, second, third, fourth, fifth, and sixth restriction sites for nicking endonucleases are target sequences for four different nicking endonucleases, including all possible combinations of assigning the six sites to four different nicking endonuclease target sequences.Furthermore, in some embodiments, the first, second, third, fourth, fifth, and sixth restriction sites for nicking endonucleases are target sequences for five different nicking endonucleases, including all possible combinations of assigning the six sites to five different nicking endonuclease target sequences. Furthermore, in some embodiments, the first, second, third, fourth, fifth, and sixth restriction sites for nicking endonucleases are target sequences for six different nicking endonucleases.

[0283] In some embodiments, one or more of the nicking endonuclease sites described in the preceding paragraph is a target sequence for an endogenous nicking endonuclease. In some particular embodiments, the plasmid further comprises an ORF encoding a nicking endonuclease that recognizes one or more of the first, second, third, fourth, fifth, and sixth restriction sites for the nicking endonuclease described in this section (Section 5.4.6), including the preceding paragraph. In one particular embodiment, the plasmid further comprises two ORFs encoding two nicking endonucleases that recognize two or more of the first, second, third, fourth, fifth, and sixth restriction sites for the nicking endonuclease described in this section (Section 5.4.6), including the preceding paragraph. In another specific embodiment, the plasmid further comprises three ORFs encoding three nicking endonucleases that recognize three or more of the first, second, third, fourth, fifth, and sixth restriction sites for a nicking endonuclease described in this section (Section 5.4.6), including the preceding paragraph. In yet another specific embodiment, the plasmid further comprises four ORFs encoding four nicking endonucleases that recognize four or more of the first, second, third, fourth, fifth, and sixth restriction sites for a nicking endonuclease described in this section (Section 5.4.6), including the preceding paragraph. In an even more specific embodiment, the plasmid further comprises five ORFs encoding five nicking endonucleases that recognize five or more of the first, second, third, fourth, fifth, and sixth restriction sites for a nicking endonuclease described in this section (Section 5.4.6), including the preceding paragraph. In one particular embodiment, the plasmid further comprises six ORFs encoding six nicking endonucleases, each recognizing a first, second, third, fourth, fifth, and sixth restriction site for a nicking endonuclease described in this section (Section 5.4.6), including the preceding paragraph. In certain embodiments, expression of one or more of the nicking endonucleases described in this paragraph is under the control of a promoter.In some embodiments, expression of one or more nicking endonucleases described in this paragraph is under the control of an inducible promoter. In some specific embodiments, the inducible promoter can be any of the inducible promoters described above in this section (Section 5.4.6).

[0284] In some embodiments, the nicking endonuclease recognizing the first, second, third, and / or fourth restriction site for a nicking endonuclease can be any one of those described in Sections 3, 5.3.4, and 5.4.2. In certain embodiments, the nicking endonuclease recognizing the first, second, third, and / or fourth restriction site for a nicking endonuclease is Nt. BsmAI; Nt. BtsCI; N. ALwl; N. BstNBI; N. BspD6I; Nb. Mva1269I; Nb. BsrDI; Nt. BtsI; Nt. BsaI; Nt. Bpu10I; Nt. BsmBI; Nb. BbvCI; Nt. BbvCI; or Nt. BspQI. In some embodiments, the nicking endonuclease recognizing the fifth and sixth restriction sites for a nicking endonuclease can be any one of those described in Sections 3, 5.3.4, and 5.4.2. In certain embodiments, the nicking endonuclease recognizing the fifth and sixth restriction sites for a nicking endonuclease is Nt. BsmAI; Nt. BtsCI; N. ALwl; N. BstNBI; N. BspD6I; Nb. Mva1269I; Nb. BsrDI; Nt. BtsI; Nt. BsaI; Nt. Bpu10I; Nt. BsmBI; Nb. BbvCI; Nt. BbvCI; or Nt. BspQI.

[0285] In some embodiments, the DNA molecules for the methods and compositions provided herein (e.g., as provided in Section 3 and this section (Section 5.4)) can be linear, non-circular DNA molecules.

[0286] In some embodiments, the vectors for the methods and compositions provided herein comprise any one or more of the features described in this section (Section 5.4.6), in various permutations and combinations. In certain embodiments, the plasmids for the methods and compositions provided herein comprise any one or more of the features described in this section (Section 5.4.6), in various permutations and combinations.

[0287] Various embodiments described in this section (Section Section 5.4.6) using nicking endonucleases and / or restriction sites for nicking endonucleases are further provided where the nicking endonucleases are replaced with programmable nicking enzymes and the restriction sites are replaced with target sites for the programmable nicking enzymes. Programmable nicking enzymes and their target sites for purposes of this paragraph and this section (Section 5.4.3) are provided in Section 5.3.4.

[0288] 5.4.7 DNA molecules with fewer than four restriction sites for nicking endonucleases and fewer than four target sites for programmable nicking enzymes In an additional aspect, provided herein is a nucleic acid sequence comprising, in the 5' to 3' direction of the top strand: i) a first inverted repeat (e.g., as described in Section 5.4.1), a first restriction site for a nicking endonuclease and a first restriction site for a restriction enzyme located at opposite ends and near the first inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking and cleavage with the restriction enzyme results in a top strand 5' overhang that includes the first inverted repeat; ii) an expression cassette (e.g., as described in Section 5.4.3); and iii) a second inverted repeat (e.g., as described in Section 5.4.1). a double-stranded DNA molecule comprising a second inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2), wherein a second restriction site for a nicking endonuclease and a second restriction site for a restriction enzyme are positioned at op...

Claims

1. 1. A method for preparing a hairpin-ended DNA molecule, comprising: (a) providing a double-stranded DNA molecule, comprising in a 5' to 3' direction of a top strand: (i) a first inverted repeat, (1) such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a top strand 5' overhang that includes the first inverted repeat; or (2) such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a bottom strand 3' overhang that includes the first inverted repeat; a first inverted repeat, wherein first and second restriction sites for one or more nicking endonucleases are located on opposite strands near the first inverted repeat; (ii) an expression cassette comprising a transgene encoding human glycogen debranching enzyme (GDE) or a catalytically active fragment thereof; and (iii) a second inverted repeat, (1) such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a top strand 3' overhang that includes the second inverted repeat; or (2) such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a bottom strand 5' overhang that includes the second inverted repeat; providing said double stranded DNA molecule comprising a second inverted repeat, wherein third and fourth restriction sites for one or more nicking endonucleases are located on opposing strands proximal to said second inverted repeat; (b) incubating said double-stranded DNA molecule with one or more nicking endonucleases that recognize four of said restriction sites and result in four nicks; (c) performing denaturation, thereby generating a DNA fragment containing the expression cassette and flanked by two single-stranded DNA overhangs; and (d) intramolecularly annealing the single-stranded DNA overhangs, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step (c). The method comprising:

2. The double-stranded DNA molecule of step (a) is (a) culturing a host cell containing the double-stranded DNA molecule under conditions that result in amplification of the double-stranded DNA molecule; and (b) releasing the double-stranded DNA molecule from the host cell; 2. The method of claim 1, wherein the host cell is a bacterial host cell.

3. The method of claim 1, wherein the double-stranded DNA molecule of step (a) is provided by in vitro replication.

4. 2. The method of claim 1, wherein said first, second, third, and fourth restriction sites for said one or more nicking endonucleases are all restriction sites for the same nicking endonuclease.

5. The double-stranded DNA molecule of step (a) (a) the first and / or second inverted repeat is a modified inverted terminal repeat (ITR) of a parvovirus, and / or (b) the first and / or second inverted repeat is a modified ITR of a parvovirus having a nucleotide sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, or at least 99% identical to the modified ITR of the parvovirus; 2. The method of claim 1, wherein the parvovirus is dependoparvovirus, bocaparvovirus, erythroparvovirus, protoparvovirus, or tetraparvovirus.

6. The double-stranded DNA molecule of step (a) (a) a first said nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second said nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat.

2. The method according to claim 1 .

7. The double-stranded DNA molecule of step (a) (a) a first said nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (c) a third said nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat; 2. The method according to claim 1 .

8. The double-stranded DNA molecule of step (a) (a) a first said nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (c) a third said nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat; 2. The method according to claim 1 .

9. The double-stranded DNA molecule of step (a) (a) a first said nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat; 2. The method according to claim 1 .

10. the double-stranded DNA molecule of step (a) is a plasmid, the plasmid being circular; (i) the plasmid further comprises a restriction enzyme site in a region 5' to the first inverted repeat and 3' to the second inverted repeat, wherein the restriction enzyme site is not present in any of the first inverted repeat, the second inverted repeat, and the region between the first and second inverted repeats; and / or The method of claim 1, wherein (ii) cleavage with a restriction enzyme results in single-stranded overhangs that do not anneal at a detectable level under conditions suitable for annealing of the first inverted repeat and / or the second inverted repeat.

11. The method of claim 1, wherein the expression cassette of the double-stranded DNA molecule of step (a) comprises one or more open reading frames (ORFs).

12. The method of claim 1, wherein the expression cassette comprises a promoter operably linked to a transcription unit, the transcription unit comprises an open reading frame (ORF), and the transcription unit further comprises post-transcriptional regulatory elements or polyadenylation and termination signals.

13. the size of the expression cassette of the double stranded DNA molecule of step (a) is at least 4 kb, at least 4.5 kb, at least 5 kb, at least 5.5 kb, at least 6 kb, at least 6.5 kb, at least 7 kb, at least 7.5 kb, at least 8 kb, at least 8.5 kb, at least 9 kb, at least 9.5 kb, or at least 10 kb; and / or 2. The method of claim 1, wherein the expression cassette comprises a transgene ranging from about 500 to about 50,000 nucleotides in length, from about 500 to about 75,000 nucleotides in length, from about 500 to about 10,000 nucleotides in length, from about 1000 to about 10,000 nucleotides in length, or from about 500 to about 5,000 nucleotides in length.

14. 2. The method of claim 1, wherein each of the first and second inverted repeats of the double-stranded DNA molecule of step (a) lacks a functional viral replication-associated protein binding sequence (RABS) or lacks a functional terminal separation site (TRS).

15. 11. The method of claim 10, (e) incubating the plasmid or the fragment resulting from step (c) with the restriction enzyme, thereby cleaving the plasmid or the fragment of the plasmid; and (f) incubating the fragment of the plasmid with an exonuclease, thereby digesting the fragment of the plasmid except for the fragment resulting from step (d). The method further comprises:

16. 11. The method of claim 10, (e) incubating the plasmid or the resulting fragment of step (c) with one or more nicking endonucleases that recognize the fifth and sixth restriction sites resulting in cleavage in the double-stranded DNA molecule; and (f) incubating the fragment of the plasmid with an exonuclease, thereby digesting the fragment of the plasmid except for the fragment resulting from step (d). The method further comprises:

17. 2. The method of claim 1, further comprising repairing nicks in the fragment resulting from step (b) using a ligase to generate circular DNA.

18. the double stranded DNA molecule of step (a), wherein the single stranded overhang of the double stranded DNA molecule comprises a palindromic sequence which, when read in a forward direction, comprises a stretch of polynucleotide that is 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to when read in a reverse direction on the complementary strand; or (a) the first and second restriction sites for a nicking endonuclease are at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least at least 96, at least 97, at least 98, at least 99, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, or at least 200 nucleotides apart; and / or (b) the third and fourth restriction sites for a nicking endonuclease are at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 9 at least 5, at least 96, at least 97, at least 98, at least 99, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, or at least 200 nucleotides apart.

2. The method according to claim 1 .

19. 2. The method of claim 1, wherein either or both of the hairpinned inverted repeats of step (d) comprise a stem, a main stem, a loop, a turn point, a bulge, a branch, a branched loop, an internal loop, and / or any combination thereof.

20. The length of the top strand 5' overhang, the bottom strand 3' overhang, the top strand 3' overhang, and / or the bottom strand 5' overhang can, independently, be at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 2. The method of claim 1, wherein the nucleic acid sequence is at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides in length.