Compositions of DNA molecules encoding factor VIII, methods of making same, and methods of use thereof - Patents.com

JP2025504404A5Pending Publication Date: 2026-03-27アンジャリウム バイオサイエンシズ エージー
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-01-13
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing viral vectors have packaging restrictions, immune response and immune tolerance problems in gene therapy, making it difficult to effectively express large transgenes, especially Factor VIII protein for treating hemophilia A, and frequent injections lead to inconvenience.

Method used

Using non-viral vector DNA molecules, a double-stranded DNA molecule is prepared by designing an expression cassette containing reverse repeat sequences, and a programmatic cleavage nuclease is used to prepare double-stranded DNA molecules to achieve stable expression of factor VIII protein, avoid immune responses, and reduce injection frequency.

Benefits of technology

The stable expression of factor VIII protein is achieved, the frequency of injection is reduced, the therapeutic effect of hemophilia A is improved, and the risk of immune response is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are double-stranded DNA molecules comprising an inverted repeat, an expression cassette, and one or more restriction sites for a nicking endonuclease, methods of use thereof, and methods of making same.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (Priority) This application claims the benefit of priority to U.S. Application No. 63 / 299,500, filed January 14, 2022, which is incorporated by reference herein in its entirety.

[0002] (REFERENCE TO ELECTRONICALLY SUBMITTED SEQUENCE LISTING) This application contains a computer readable sequence listing submitted herewith in XML file format, the entire contents of which are incorporated herein by reference in their entirety. The sequence listing XML file submitted herewith is named "14497-009-228_SequenceListing.xml", was created on January 12, 2023, and is 677,006 bytes in size.

[0003] 1. Field Provided herein are DNA molecules encoding coagulation factor VIII, methods of use thereof, and methods of making same. Also provided are methods of treating bleeding disorders. [Background technology]

[0004] 2.Background Gene therapy aims to introduce genes into target cells to treat or prevent disease. By providing a transcription cassette (sometimes called a transgene) with an active gene product, gene therapy can improve clinical outcomes. This is because the gene product can result in the gain of a beneficial functional effect, the loss of a harmful functional effect, or another outcome, which may have an oncolytic effect, for example, in patients with cancer. Delivery and expression of the corrective gene in the target cells of the patient can be achieved through a number of methods, such as non-viral delivery (e.g., by liposomes) or viral delivery methods, including the use of engineered viruses and viral gene delivery vectors. Of the available virus-derived vectors, also known as viral particles (e.g., recombinant retroviruses, recombinant lentiviruses, recombinant adenoviruses, etc.), the AAV system has become popular as a versatile vector in gene therapy.

[0005] However, there are some significant drawbacks to using viral particles as gene delivery vectors. One of the significant drawbacks is the reliance on viral life cycle and viral proteins to package the transcription cassette into the viral particle. As a result, the use of viral vectors is limited in terms of the size of the transgene (e.g., protein coding capacity of less than 150,000 Da for AAV) or the need for the presence of certain viral sequences (e.g., Rep binding elements) that may destabilize the expression cassette to ensure efficient replication and packaging. Thus, more than one viral particle may be required to deliver a large transgene (e.g., a transgene encoding a protein larger than 150,000 Da or a transgene longer than about 4.7 kb). The use of more than one AAV construct may increase the risk of reactivation of the AAV genome. In addition, the use of viral Rep or nonstructural protein 1 binding elements may increase the risk of vector mobilization in patients.

[0006] A second drawback is that the viral particles used in gene therapy are often derived from wild-type viruses to which a portion of the population has been exposed during their lifetime. These patients are known to have neutralizing antibodies that may in turn impair the efficacy of gene therapy, as further described in Snyder, Richard O., and Philippe Moullier. Adeno-associated viruses: methods and protocols. Totowa, NJ: Humana Press, 2011. For the remaining seronegative patients, the capsid of the viral vector is often immunogenic, preventing re-administration of the patient to viral vector therapy if the initial dose is insufficient or the therapy becomes ineffective over time.

[0007] Therefore, there is still an unmet need for non-viral gene therapy as an alternative to viral particles, especially for therapy that delivers large transgenes.There is also a need for DNA vectors that provide higher stability in cell nuclei and allow extended expression compared to circular plasmid DNA.In addition, there is still an unmet need for a method to produce these DNA vectors without the coexistence of plasmids or DNA sequences that code for viral replication machinery (e.g., AAV Rep genes).This is because these viral proteins, or the viral DNA sequences that code for them, may be contaminated with the isolated DNA of DNA vectors.

[0008] Additionally, there remains a significant unmet need for recombinant DNA vectors with improved manufacturability and / or expression characteristics, as well as DNA-based vectors that do not induce anti-viral (e.g., viral capsid, Toll-like receptor activation, etc.) immune responses and allow for repeated administration without loss of efficacy (e.g., due to neutralizing antibodies) or loss of transgene-expressing cells.

[0009] Disorders associated with impaired function or deficiency of the coagulation factor VIII (FVIII), including hemophilia A, result in blood clotting disorders. Because joints are the most frequently involved anatomical site, patients are subject to injury due to increased risk of bleeding, which can be life-threatening depending on where the bleeding occurs. Although all joints can be involved, joint bleeding usually occurs in the large synovial joints (e.g., knees, ankles, elbows) and gradually leads to severe disabling arthropathy. Currently, disease management involves frequent intravenous injections of recombinant FVIII protein, with high frequency due to its short half-life. Enzyme replacement must be initiated as soon as possible after birth and must be continued for at least 15 years, if not lifelong. Furthermore, most patients with hemophilia A develop long-term pathology. Despite recent successes with adeno-associated virus (AAV)-based gene replacement for metabolic diseases, current limitations of AAV-mediated gene transfer, including gene size, remain a challenge for successful gene therapy in hemophilia A (Leebeek and Miesbach, Gene Therapy for Hemophilia: a review on clinical benefit, limitations and remaining issues, Blood, 2021). Furthermore, loss of transgenes over time has been observed in liver-directed AAV gene therapy, likely due to the pathological state of the treated hepatocytes.

[0010] Despite great advances in understanding molecular biology and diagnosing hemophilia A, little progress has been made in developing new treatments for the disorder. There remains a large unmet need for sustained disease-modifying therapy in hemophilia A. Classical treatment of hemophilia A is through replacement therapy targeted at restoring factor VIII activity. Replacement therapy to treat hemophilia A involves restoring factor VIII activity to 1-5% of normal levels to prevent spontaneous bleeding. There are plasma-derived and recombinant factor VIII products available to prevent the occurrence of bleeding episodes by treating them on demand or prophylactically. Based on the half-life of these products, treatment regimens require frequent intravenous administration. Such frequent administration is painful and inconvenient. Furthermore, the need to prevent long-term damage to joints and chronic pain remains unaddressed. There is no approved gene therapy for hemophilia A, and regular AAV-based therapy cannot accommodate large wild-type transgenes and cannot be used for 25%-40% of patients due to pre-existing antibodies. Other viral gene therapy vectors that can accommodate large transgenes pose the challenge that they can only be administered once, the resulting Factor VIII (FVIII) expression levels may not be high enough to be effective, or dose levels may not be titrated above normal.

[0011] Thus, there is a need in the art for techniques that allow for the expression of therapeutic FVIII proteins in cells, tissues, or subjects in need of treatment for hemophilia A. Summary of the Invention

[0012] 3. Overview In one aspect, provided herein is a method for treating a disease associated with decreased activity of coagulation factor VIII in a human patient, the method comprising administering to the patient a biocompatible carrier (hybridosome) or lipid nanoparticle, wherein the hybridosome or lipid nanoparticle comprises a DNA molecule comprising an expression cassette comprising a transgene encoding human FVIII or a catalytically active fragment thereof.

[0013] Provided herein is a method for treating a disease associated with decreased activity of coagulation factor VIII in a human patient, the method comprising administering to the patient a DNA molecule comprising an expression cassette including a transgene encoding human coagulation factor VIII or a catalytically active fragment thereof, wherein the DNA molecule is contained within a single delivery vector.

[0014] Described herein is a method for treating a disease associated with reduced activity of FVIII in a human patient, the method comprising: (i) administering to the patient a first dose of a DNA molecule comprising an expression cassette including an introduced gene encoding human FVIII or a catalytically active fragment thereof; and (ii) administering to the patient a second dose of the DNA molecule.

[0015] In one embodiment, the first dose of the DNA molecule is administered to the patient at least 3 months, at least 4 months, at least 5 months, at least 6 months, at least 7 months, at least 8 months, at least 9 months, at least 10 months, or at least 11 months before the second dose of the DNA molecule.

[0016] In one embodiment, the first dose of the DNA molecule is administered to the patient at least 1 year, at least 2 years, at least 3 years, at least 4 years, at least 5 years, at least 10 years, at least 15 years, or at least 20 years before the second dose of the DNA molecule.

[0017] In one embodiment, the first dose of double stranded DNA molecules and the second dose of DNA molecules contain the same amount of DNA molecules.

[0018] In one embodiment, the first dose of DNA molecules and the second dose of DNA molecules contain different amounts of DNA molecules.

[0019] In one embodiment, the method further comprises administering one or more additional doses of the DNA molecule.

[0020] In one embodiment, the DNA molecule is administered once every week, once every other week, or once every month.

[0021] In one embodiment, the DNA molecule is administered to the patient about every 6 months, about every 12 months, about every 18 months, about every 2 years, about every 3 years, about every 5 years, about every 10 years, about every 15 years, or about every 20 years.

[0022] In one embodiment, the DNA molecule is administered to the patient for the life of the patient.

[0023] In one embodiment, the patient is an adult patient.

[0024] In one embodiment, the patient is a pediatric patient.

[0025] In one embodiment, the patient is a pediatric patient when the first dose of the DNA molecule is administered.

[0026] In one embodiment, the pediatric patient is an infant.

[0027] In one embodiment, the pediatric patient is about 1 year old, about 2 years old, about 3 years old, about 4 years old, about 5 years old, about 6 years old, about 7 years old, about 8 years old, about 9 years old, about 10 years old, about 11 years old, about 12 years old, about 13 years old, about 14 years old, about 15 years old, about 16 years old, about 17 years old, or about 18 years old.

[0028] In one embodiment, the disease is hemophilia A.

[0029] In one embodiment, the transgene comprises a sequence that is at least 60%, at least 70%, at least 80% or at least 90% identical to the sequence set forth in SEQ ID NO: 174, 175, 176, 177, 178, 179, 180, 181, 379, 380, 381, 383, 384, 385, 387, 388, 389, 391, 392, 393, 395, 396, 397, 399, 400, 401, 403, 404, 405, 407, 408, or 409.

[0030] In one embodiment, the method results in amelioration of one or more of the clinical symptoms of hemophilia A: excessive annual bleeding rate, hemophilic arthropathy, and irreversible joint damage.

[0031] In one embodiment, the method results in a reduction in the number of bleeding episodes in a patient by about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100% per year.

[0032] In one embodiment, the method results in an improvement in blood clotting cascade function of about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100% as determined by a coagulation function test.

[0033] In one embodiment, the method results in a reduction in the number of joint bleeds in a patient by about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100% per year.

[0034] In one embodiment, the method results in greater than about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 95% clinical improvement as measured by one or more of the following coagulation markers: prothrombin time test, partial thromboplastin time, and clotting factor tests.

[0035] In one embodiment, the method results in greater than about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 95% clinical improvement as measured by the level of FVIII in the patient's plasma.

[0036] In one embodiment, the method results in a FVIII protein activity that is about 1-10%, about 10-20%, about 20-30%, about 30-40%, about 40-50%, about 50-60%, about 60-70%, about 70-80%, or about 80-90% of the biological activity level of the native FVIII protein.

[0037] In one embodiment, the DNA molecule is detectable in the patient's liver cells by quantitative real-time PCR.

[0038] Provided herein is a 5' to 3' direction of the top strand: (a) a first inverted repeat, where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a top strand 5' overhang that includes the first inverted repeat; (b) an expression cassette comprising a transgene encoding human FVIII or a catalytically active fragment thereof; (c) a second inverted repeat, where a third and a fourth restriction site for a nicking endonuclease are positioned on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a top strand 3' overhang that includes the second inverted repeat.

[0039] Provided herein is a 5' to 3' direction of the top strand: (a) a first inverted repeat, where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a bottom strand 3' overhang that includes the first inverted repeat; (b) an expression cassette comprising a transgene encoding human FVIII or a catalytically active fragment thereof; (c) a second inverted repeat, where a third and a fourth restriction site for a nicking endonuclease are positioned on opposing strands near the second inverted repeat such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a bottom strand 5' overhang that includes the second inverted repeat.

[0040] Provided herein is a 5' to 3' direction of the top strand: (a) a first inverted repeat, where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a top strand 5' overhang that includes the first inverted repeat; (b) an expression cassette comprising a transgene encoding human FVIII or a catalytically active fragment thereof; (c) a second inverted repeat, where a third and a fourth restriction site for a nicking endonuclease are positioned on opposing strands near the second inverted repeat such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a bottom strand 5' overhang that includes the second inverted repeat.

[0041] Provided herein is a 5' to 3' direction of the top strand: (a) a first inverted repeat, where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a bottom strand 3' overhang that includes the first inverted repeat; (b) an expression cassette comprising a transgene encoding human FVIII or a catalytically active fragment thereof; (c) a second inverted repeat, where a third and a fourth restriction site for a nicking endonuclease are positioned on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a top strand 3' overhang that includes the second inverted repeat.

[0042] In one embodiment, the DNA molecules provided herein are isolated DNA molecules.

[0043] In one embodiment, the first, second, third, and fourth restriction sites for a nicking endonuclease of a DNA molecule provided herein are all restriction sites for the same nicking endonuclease.

[0044] In one embodiment, the first and second inverted repeats of the DNA molecules provided herein are the same.

[0045] In one embodiment, the first and / or second inverted repeat of the DNA molecule provided herein is an ITR of a parvovirus.

[0046] In one embodiment, the first and / or second inverted repeat of the DNA molecule provided herein is a modified ITR of a parvovirus.

[0047] In one embodiment, the parvovirus is a dependoparvovirus, a bocaparvovirus, an erythroparvovirus, a protoparvovirus, or a tetraparvovirus.

[0048] In one embodiment, the nucleotide sequence of the modified ITR of the DNA molecule provided herein is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, or at least 99% identical to an ITR of a parvovirus.

[0049] In one embodiment, the ITRs of the DNA molecules provided herein contain viral replication-associated protein binding sequences ("RABS").

[0050] In one embodiment, the RABS comprises a Rep binding sequence.

[0051] In one embodiment, the RABS comprises an NS1 binding sequence.

[0052] In one embodiment, the ITRs of the DNA molecules provided herein do not contain a RABS.

[0053] In one embodiment, the transgene comprises the sequence of SEQ ID NO: 174, 175, 176, 177, 178, 179, 180, 181, 379, 380, 381, 383, 384, 385, 387, 388, 389, 391, 392, 393, 395, 396, 397, 399, 400, 401, 403, 404, 405, 407, 408, or 409.

[0054] In one embodiment, the DNA molecule provided herein comprises: (a) the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat.

[0055] In one embodiment, the DNA molecule provided herein comprises: (a) the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the 3' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat.

[0056] In some embodiments, the DNA molecule provided herein comprises: (a) the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the 3' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat.

[0057] In some embodiments, the DNA molecule provided herein comprises: (a) the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the first inverted repeat; (b) the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the first inverted repeat; (c) the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 5' nucleotide of the ITR closing base pair of the second inverted repeat; and / or (d) the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides of the 3' nucleotide of the ITR closing base pair of the second inverted repeat.

[0058] In one embodiment, the nick is within the inverted repeat.

[0059] In one embodiment, the nick is outside the inverted repeat.

[0060] In one embodiment, the DNA molecule is a plasmid.

[0061] In one embodiment, the DNA molecule is a linear DNA molecule.

[0062] In one embodiment, the plasmid further comprises a bacterial origin of replication.

[0063] In one embodiment, the plasmid further comprises a restriction enzyme site in a region 5' to the first inverted repeat and 3' to the second inverted repeat, wherein the restriction enzyme site is not present in any of the first inverted repeat, the second inverted repeat, or the region between the first and second inverted repeats.

[0064] In one embodiment, cleavage with a restriction enzyme results in single-stranded overhangs that do not detectably anneal under conditions that favor annealing of the first inverted repeat and / or the second inverted repeat.

[0065] In one embodiment, the plasmid further comprises a fifth and sixth restriction site for a nicking endonuclease in the region 5' to the first inverted repeat and 3' to the second inverted repeat, the fifth and sixth restriction sites for the nicking endonuclease being (a) are on opposing chains, and (b) creating a break in the double-stranded DNA molecule such that the single-stranded overhangs of the break do not detectably anneal intermolecularly or intramolecularly under conditions that favor annealing of the first and / or second inverted repeats.

[0066] In one embodiment, the fifth and sixth nicks are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides apart.

[0067] In one embodiment, the first, second, third, fourth, fifth, and sixth restriction sites for a nicking endonuclease are all target sequences for the same nicking endonuclease.

[0068] In one embodiment, the nicking endonuclease recognizing the first, second, third, and / or fourth restriction site for the nicking endonuclease is Nt.BsmAI, Nt.BtsCI, N.ALwl, N.BstNBI, N.BspD6I, Nb.Mva1269I, Nb.BsrDI, Nt.BtsI, Nt.BsaI, Nt.Bpu10I, Nt.BsmBI, Nb.BbvCI, Nt.BbvCI, or Nt.BspQI.

[0069] In one embodiment, the nicking endonuclease recognizing the fifth and sixth restriction sites for the nicking endonuclease is Nt.BsmAI, Nt.BtsCI, N.ALwl, N.BstNBI, N.BspD6I, Nb.Mva1269I, Nb.BsrDI, Nt.BtsI, Nt.BsaI, Nt.Bpu10I, Nt.BsmBI, Nb.BbvCI, Nt.BbvCI, or Nt.BspQI.

[0070] In one embodiment, the nicking endonuclease that recognizes the first, second, third, and / or fourth restriction site for the nicking endonuclease is a programmable nicking endonuclease.

[0071] In one embodiment, the nicking endonuclease that recognizes the fifth and sixth restriction sites for the nicking endonuclease is a programmable nicking endonuclease.

[0072] In one embodiment, the nicking endonuclease is a Cas nuclease.

[0073] In one embodiment, the expression cassette further comprises a promoter operably linked to the transcription unit.

[0074] In one embodiment, a transcription unit comprises an open reading frame.

[0075] In one embodiment, the expression cassette further comprises a post-transcriptional regulatory element.

[0076] In one embodiment, the expression cassette further comprises a polyadenylation signal and a termination signal.

[0077] In one embodiment, the size of the expression cassette is at least 4 kb, at least 4.5 kb, at least 5 kb, at least 5.5 kb, at least 6 kb, at least 6.5 kb, at least 7 kb, at least 7.5 kb, at least 8 kb, at least 8.5 kb, at least 9 kb, at least 9.5 kb, or at least 10 kb.

[0078] Provided herein is a kit for expressing human FVIII in vivo, the kit comprising 0.1 to 500 mg of a DNA molecule provided herein and a device for administering the DNA molecule.

[0079] In one embodiment, the device is a syringe needle.

[0080] Provided herein are compositions comprising one or more of the DNA molecules provided herein and a pharma- ceutically acceptable carrier.

[0081] In one embodiment, the carrier comprises a transfection reagent, a nanoparticle, a hybridosome, a lipid nanoparticle, or a liposome.

[0082] In one embodiment, the compositions provided herein are used in medical therapy.

[0083] In one embodiment, the compositions provided herein are used to prepare or manufacture a medicament for alleviating, preventing, delaying the onset of, or treating a disease or disorder associated with reduced activity of FVIII in a subject in need thereof. [Brief description of the drawings]

[0084] 4. Brief description of the drawings [Figure 1] 1 depicts various exemplary hairpin structures and hairpin structural elements.

[0085] [Diagram 2] A and B represent linear interaction plots showing exemplary strand conformations and intramolecular forces within the overhangs, as well as intermolecular forces between the strands, and C represents the expected annealed structure of Figures 2A and 2B.

[0086] [Diagram 3] Various exemplary configurations of the hairpin are depicted, as well as the location of various restriction sites, as well as a restriction site for a type II nicking endonuclease in the primary stem of the hairpin.

[0087] [Figure 4]1 depicts the structures of various exemplary hairpins and structural elements of human mitochondrial DNA OriL and OriL-derived ITRs.

[0088] [Diagram 5] 1 depicts the structure of an exemplary aptamer and an aptamer ITR hairpin.

[0089] [Figure 6] 1 depicts construct 1 and a visualization of the DNA products derived from construct 1 after carrying out the method steps described in Example 1.

[0090] [Figure 7] 1 shows a visualization of the DNA products derived from construct 2 and construct 1 after carrying out the method steps described in Example 1.

[0091] [Figure 8] AC represent multiple renaturation / denaturation cycles as described in Example 2.

[0092] [Figure 9] A and B represent the isothermal denaturation of construct 1 described in Example 3.

[0093] [Figure 10] Figure 1 shows the expression levels of luciferase from different amounts of DNA vector. Cells were transfected with various concentrations of DNA vector containing either hybridosomes or lipid nanoparticles. Luciferase activity was determined 48 hours after transfection.

[0094] [Figure 11]Figure 6 represents luciferase expression in dividing and non-dividing cells as described in section 6.5 (Example 5 Expression in dividing and non-dividing cells). A and B represent expression of non-secreted Turboluc (construct 1) in dividing cells (A) and non-dividing cells (B). For non-secreted Turboluc (construct 1), luciferase activity in dividing cells peaks on day 2, while expression continues to increase in non-dividing cells. C and D represent expression of secreted Turboluc (construct 2) in non-dividing cells (C) and dividing cells (D). For secreted Turboluc (construct 2), luciferase activity peaks on day 2 in dividing cells, while expression increases in non-dividing cells and remains stable for 9 days thereafter. As a direct comparison, an equimolar amount of the complete circular plasmid encoding construct 2 was also transfected and, as shown in C and D, generally lower luciferase activities were recorded, indicating improved nuclear delivery of purified construct 2 with folded ITRs.

[0095] [Figure 12] A sequence alignment of the ITRs from AAV1 is shown highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0096] [Figure 13] A sequence alignment of the ITRs from AAV2 is shown highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0097] [Figure 14] A sequence alignment of the ITRs from AAV3 is shown highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0098] [Figure 15] A sequence alignment of the ITRs from AAV4 left is shown highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0099] [Figure 16] A sequence alignment of the ITRs from AAV4 right is shown highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0100] [Figure 17] A sequence alignment of the ITRs from AAV5 is shown highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0101] [Figure 18] A sequence alignment of the ITRs from AAV7 left is shown highlighting sequence modifications that create recognition sites for various nicking endonuclease recognition sites.

[0102] [Figure 19] A and B represent agarose gels showing successful ligation of DNA constructs and corresponding luciferase expression in non-dividing hepatocytes transfected with hybridosomes encapsulating the ligated and unligated constructs, and the parental plasmid, respectively.

[0103] [Figure 20] 1 depicts the time course of luciferase expression by non-dividing cells transfected with equimolar amounts of hairpin-ended DNA molecules encoding secreted luciferase encapsulated within LNPs or hybridosomes.

[0104] [Figure 21] 13 depicts the percentage of RFP-positive color-switched HEK293 cells after 72 h transfection with hairpin-ended DNA encoding Cre recombinase delivered by lipid nanoparticles, hybridosomes, and jetprime as described in Example 9.

[0105] [Figure 22] A and B represent agarose gels showing successful formation of hairpin-ended DNA from plasmids containing right and left ITRs with wild-type AAV RBE compared to mutants in which the RBE was replaced with the corresponding sequences shown in the figure. Luciferase expression in non-dividing hepatocytes transfected with the corresponding ITR sequences is shown in B.

[0106] [Diagram 23] A and B show a further exemplary cloning method (FIG. 24A) and a map of the resulting plasmid (FIG. 23B) from which hairpinned inverted repeat DNA molecules as disclosed herein can be prepared by carrying out the method steps as described in Example 11. In this example, six restriction sites for nicking endonucleases are located in the region 5' to the left ITR and 3' to the right ITR.

[0107] [Figure 24] A and B represent and visualize the products of nicking, denaturation / annealing and exonuclease digestion starting from the plasmid depicted in FIG. 23B on an agarose gel and the luminescence readout of the above-mentioned products subjected to transfection.

[0108] [Diagram 25] A and B represent and visualize the products of nicking, denaturation / annealing, and exonuclease digestion of the FVIII-encoding construct described in Example 12 on an agarose gel. For the DNA construct encoding truncated FVIII, the agarose gel (A) shows the nicked plasmid in lane 2, the denatured / renatured DNA product in lane 3, the digestion-resistant vector in lane 4, and the purified product in lane 5. For the DNA construct encoding full-length FVIII, the agarose gel (B) shows the nicked plasmid in lane 2, the denatured / renatured DNA product in lane 3, a single band of the digestion-resistant vector in lane 4, and the purified product in lane 5.

[0109] [Figure 26] 1 shows the levels of FVIII activity after transfection of truncated FVIII and full-length FVIII of Example 13.

[0110] [Figure 27] FIG. 1 shows the levels of FVIII activity following transfection of CpG-free ITR constructs encoding truncated FVIII variants with partial B domain deletion and truncated FVIII variants with partial B domain / linker a3 domain deletion in Huh-7 cells as described in Example 14.

[0111] [Figure 28] 1 shows the results of FVIII expression from transfected hairpin-end DNA molecules encoding different FVIII variants and codon optimization and the effect on supernatant FVIII concentration (IU / ml) as described in Example 17.

[0112] [Figure 29] 13 depicts the results of an in vivo study of activated partial thromboplastin clotting time on day 3 following administration of FVIII encoding hairpin-ended DNA molecules formulated into LNPs, as described in Example 18. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0113] 5. Detailed Description Provided herein are methods and compositions for treating a disease or disorder associated with reduced presence or function of coagulation factor VIII (FVIII) in a subject. In some embodiments, the disease associated with reduced presence or function of FVIII is hemophilia A (hemophilia A). Such compositions include hairpin-ended DNA molecules comprising one or more nucleic acids encoding a therapeutic FVIII protein or fragments thereof. In one embodiment, the compositions described herein include hairpin-ended DNA molecules comprising one nucleic acid encoding a therapeutic FVIII protein or fragments thereof. In one embodiment, the compositions described herein include hairpin-ended DNA molecules comprising two, three, four, or more nucleic acids encoding a therapeutic FVIII protein or fragments thereof. Also provided herein are hairpin-ended DNA molecules for expression of a FVIII protein described herein comprising one or more nucleic acids encoding a FVIII protein. Also provided herein are methods of producing the hairpin-ended DNA molecules described herein. Also provided herein are methods of treating hemophilia A using the hairpin-ended DNA and related pharmaceutical compositions provided herein. More specifically, provided herein is a method of treating hemophilia A, comprising administering to a subject in need thereof a hairpin-ended DNA described herein.

[0114] Provided herein are methods of making hairpin-ended DNA molecules. Also provided herein are methods of using hairpin-ended DNA molecules, including, for example, using hairpin-ended DNA molecules in gene therapy. Various methods of making hairpin-ended DNA molecules are further described below in Section 5.2. Various methods of using hairpin-ended DNA molecules are described below in Section 5.8. Hairpin-ended DNA made by these methods is shown below in Section 5.5, which includes two terminal hairpinned inverted repeats, each of which is further described below, and an expression cassette. In some embodiments, the hairpin-ended DNA also includes one or two nicks as further provided below in Section 5.5. Hairpins, hairpinned inverted repeats, and hairpinned ends are described below in Section 5.5, inverted repeats that form hairpinned ends are described below in Section 5.4.1, nicks, nicking endonucleases, and restriction sites for nicking endonucleases are described below in Sections 5.4.2 and 5.5, expression cassettes are described below in Sections 5.4.3 and 5.5, and functional properties of hairpin end DNA molecules are described below in Section 5.6. Thus, the present disclosure provides hairpin end DNA molecules, methods of making same, and methods of use therefor, with any combination or permutation of the components provided herein.

[0115] Also provided herein are parent DNA molecules for use in methods of making hairpin-ended DNA molecules, the parent DNA molecules comprising two inverted repeats, two or more restriction sites for a nicking endonuclease, and an expression cassette, each as further described below. The restriction sites for the nicking endonuclease are positioned such that upon nicking and denaturation with the nicking endonuclease, a single-stranded overhang having the inverted repeat sequence is generated, which then folds upon annealing to form a hairpin (each step as described in Section 5.2). The inverted repeat is described in Section 5.4.1 below, the nick, the nicking endonuclease, and the restriction sites for the nicking endonuclease are described in Section 5.4.2 below, and the expression cassette is described in Section 5.4.3 below. Thus, the present disclosure provides parent DNA molecules for use in the methods of making with any combination or permutation of the components provided herein.

[0116] 5.1 Definition As used herein, the term "isolated" when used in relation to a DNA molecule is intended to mean that the referenced DNA molecule is free of at least one component found in its native, natural, or synthetic environment. This term includes a DNA molecule that has been removed from some or all other components found in its native, natural, or synthetic environment. A component of a DNA molecule's native, natural, or synthetic environment includes anything in the native, natural, or synthetic environment that is required, used, or otherwise plays a role in the replication and maintenance of the DNA molecule in that environment. Components of a DNA molecule's native, natural, or synthetic environment also include, for example, cells, debris, organelles, proteins, peptides, amino acids, lipids, polysaccharides, nucleic acids other than the referenced DNA molecule, salts, nutrients for cell culture, and / or chemicals used in DNA synthesis. A DNA molecule of the present disclosure can be partially, completely, or substantially free of all of these components, or any other components of its native, natural, or synthetic environment that are isolated, synthetically produced, naturally produced, or recombinantly produced. Examples of isolated DNA molecules include partially pure and substantially pure DNA molecules.

[0117] As used herein, the term "delivery vehicle" refers to a substance that can be used to administer or deliver one or more agents to cells, tissues, or subjects, particularly human subjects, that contain or do not contain the agent(s) to be delivered. A delivery vehicle can preferentially deliver agent(s) to a particular subset or type of cells. The selective or preferential delivery achieved by a delivery vehicle can be achieved by the nature of the vehicle, or by a moiety conjugated to, associated with, or contained in the delivery vehicle, which specifically or preferentially binds to a particular subset of cells. A delivery vehicle can also increase the in vivo half-life of the agent to be delivered, the efficiency of delivery of the agent compared to delivery without a delivery vehicle, and / or the bioavailability of the agent to be delivered. Non-limiting examples of delivery vehicles are hybridosomes, liposomes, lipid nanoparticles, polymersomes, mixtures of natural / synthetic lipids, membranes or lipid extracts, exosomes, viral particles, proteins or protein complexes, peptides, and / or polysaccharides.

[0118] As used herein, the term "subject" refers to a human or any non-human animal (e.g., mouse, rat, rabbit, dog, cat, cow, pig, sheep, horse, or primate). Human includes prenatal and postnatal forms. In many embodiments, the subject is a human. A subject may be a patient, which refers to a human presenting to a health care provider for diagnosis or treatment of a disease. The term "subject" is used interchangeably herein with "individual" or "patient." A subject is afflicted with or susceptible to a disease or disorder, but may or may not exhibit symptoms of the disease or disorder. In an exemplary embodiment, the subject of the present disclosure is a subject having reduced activity (e.g., resulting from reduced concentration, presence, and / or function) of coagulation factor VIII (FVIII). In a further exemplary embodiment, the subject is a human.

[0119] As used herein, the term "therapeutic protein" refers to any polypeptide known in the art that, when expressed in a subject for the treatment of a disease or disorder associated with reduced presence or function of FVIII in a subject (e.g., hemophilia A), results in a significant, measurable change in expression of a hemophilia A biomarker, or a reduction in symptoms associated with a given disease.

[0120] In some embodiments, the therapeutic protein comprises a protein selected from a clotting factor, a functional fragment thereof, or a combination thereof. As used herein, the term "clotting factor" refers to a naturally occurring or recombinantly produced protein, or a fragment or analog thereof, that prevents or shortens the duration of a bleeding episode in a subject. In other words, it refers to a protein that has a pro-clotting activity, such as a protein involved in the conversion of fibrinogen to a mesh of insoluble fibrin that causes blood to clot or form a clot. As used herein, "clotting factor" includes an activated clotting factor, its zymogen, or an activatable clotting factor. An "activatable clotting factor" is a clotting factor in an inactive form (e.g., its zymogen form) that can be converted to an active form.

[0121] As used herein, the term "and / or" in a phrase such as "A and / or B" is intended to include both A and B, A or B, A (alone), and B (alone). Similarly, the term "and / or" in a phrase such as "A, B, and / or C" is intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone), and C (alone).

[0122] 5.2 Hairpin-end DNA molecules and methods for making hairpin-end DNA molecules The methods and compositions described herein involve compositions and methods for delivering a FVIII nucleic acid sequence encoding a human FVIII protein to a subject in need thereof for the treatment of hemophilia A.

[0123] In some embodiments, the polynucleotide molecules provided herein express human FVIII having anti-hemophilia factor VIII (collectively or individually referred to herein as "FVIII," "F8," or "Factor VIII"), or a fragment thereof.

[0124] In some embodiments, the hairpin-ended DNA molecules of the present disclosure can be used in a method for alleviating, preventing, or treating hemophilia A in a subject in need thereof.

[0125] The disease or disorder (e.g., hemophilia A) treated herein may be associated with spontaneous bleeding and excessive bleeding after trauma. Over time, repeated bleeding in muscles and joints, often beginning in early childhood, causes hemophilic arthropathy and irreversible joint damage. This damage is progressive and can severely limit joint mobility, leading to muscle atrophy and chronic pain. As will be understood by those skilled in the art, hemophilia A may be referred to by any number of alternative names in the art, including, but not limited to, FVIII deficiency, bleeding tendency disease, or classical hemophilia. Thus, hemophilia A may be used interchangeably with any of these alternative names in the specification, examples, figures, and claims.

[0126] In a further aspect, provided herein is a method for making a hairpin-ended DNA molecule for expressing human coagulation factor VIII (FVIII) and / or a functional fragment thereof. In one aspect, provided herein is a method for preparing a hairpin-ended DNA molecule, the method comprising: a. amplifying a DNA molecule; b. incubating the DNA molecule with one or more nicking endonucleases that recognize four restriction sites and produce at least four nicks; c. denaturing, thereby generating a DNA fragment that contains an expression cassette and is flanked by two single-stranded DNA overhangs; and d. annealing the single-stranded DNA overhangs intramolecularly, thereby generating hairpinned inverted repeats at both ends of the DNA fragment obtained from step c.

[0127] 5.3 Methods for making hairpin-ended DNA molecules In one aspect, provided herein is a method for preparing a hairpin-ended DNA molecule comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. culturing a host cell comprising the DNA molecule described in section 5.4 under conditions resulting in amplification of the DNA molecule; b. releasing the DNA molecule from the host cell; c. incubating the DNA molecule with one or more nicking endonucleases that recognize four restriction sites for the nicking endonuclease that results in four nicks; d. denaturing, thereby generating a DNA fragment comprising an expression cassette encoding FVIII or a functional fragment thereof and flanked by two single-stranded DNA overhangs; and e. intramolecularly annealing the single-stranded DNA overhangs, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d.

[0128] In another aspect, provided herein is a method for preparing a hairpin-ended DNA comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. culturing a host cell comprising the plasmid of section 5.4.6 under conditions resulting in amplification of the plasmid; b. releasing the plasmid from the host cell; c. incubating the DNA molecule with one or more nicking endonucleases that recognize four restriction sites resulting in four nicks; d. denaturing, thereby generating a DNA fragment comprising the expression cassette encoding FVIII or a functional fragment thereof and flanked by two single-stranded DNA overhangs; e. intramolecularly annealing the single-stranded DNA overhangs, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d; f. incubating the plasmid or fragment resulting from step d with a restriction enzyme, thereby cleaving the plasmid or fragment of the plasmid; and g. incubating the fragment of the plasmid with an exonuclease, thereby digesting the fragment of the plasmid except for the fragment resulting from step e.

[0129] In a further aspect, provided herein is a method for preparing a hairpin-ended DNA comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. culturing a host cell comprising the plasmid of section 5.4 under conditions resulting in amplification of the plasmid; b. releasing the plasmid from the host cell; c. incubating the DNA molecule with one or more nicking endonucleases that recognize a first, second, third, and fourth restriction site resulting in four nicks; d. denaturing, thereby producing a hairpin-ended DNA comprising an expression cassette encoding FVIII (or a functional fragment thereof); generating a DNA fragment flanked by two single-stranded DNA overhangs; e. intramolecularly annealing the single-stranded DNA overhangs, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d; f. incubating the plasmid or fragment resulting from step d with one or more nicking endonucleases that recognize the fifth and sixth restriction sites resulting in cleavage within the double-stranded DNA molecule; and g. incubating the plasmid fragment with an exonuclease, thereby digesting the plasmid fragment except for the fragment resulting from step e.

[0130] In one aspect, provided herein is a method for preparing a hairpin-ended DNA molecule comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. culturing a host cell comprising the DNA molecule described in section 5.4 under conditions resulting in amplification of the DNA molecule; b. releasing the DNA molecule from the host cell; c. incubating the DNA molecule with one or more programmable nicking enzymes that recognize four target sites for the guide nucleic acid resulting in four nicks; d. denaturing, thereby generating a DNA fragment comprising an expression cassette encoding FVIII (or a functional fragment thereof) and flanked by two single-stranded DNA overhangs; and e. intramolecularly annealing the single-stranded DNA overhangs, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d.

[0131] In another aspect, provided herein is a method for preparing a hairpin-ended DNA molecule comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. culturing a host cell comprising the plasmid of section 5.4.6 under conditions resulting in amplification of the plasmid; b. releasing the plasmid from the host cell; c. incubating the DNA molecule with one or more programmable nicking enzymes that recognize four target sites for the guide nucleic acid resulting in four nicks; d. denaturing, thereby generating a DNA fragment comprising the expression cassette encoding FVIII or a functional fragment thereof and flanked by two single-stranded DNA overhangs; e. intramolecularly annealing the single-stranded DNA overhangs, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d; f. incubating the plasmid or fragment resulting from step d with a restriction enzyme, thereby cleaving the plasmid or fragment of the plasmid; and g. incubating the fragment of the plasmid with an endonuclease, thereby digesting the fragment of the plasmid except for the fragment resulting from step e.

[0132] In a further aspect, provided herein is a method for preparing a hairpin-ended DNA comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. culturing a host cell comprising the plasmid of section 5.4 under conditions resulting in amplification of the plasmid; b. releasing the plasmid from the host cell; c. incubating the DNA molecule with one or more programmable nicking enzymes that recognize first, second, third, and fourth target sites for the guide nucleic acid resulting in four nicks; d. denaturing, thereby producing a hairpin-ended DNA comprising an expression cassette encoding FVIII or a functional fragment thereof; generating a DNA fragment flanked by two single-stranded DNA overhangs; e. intramolecularly annealing the single-stranded DNA overhangs, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step d; f. incubating the plasmid or fragment resulting from step d with a programmable nicking enzyme that recognizes target sites for the fifth and sixth guide nucleic acid, resulting in cleavage within the double-stranded DNA molecule; g. incubating the plasmid fragment with an exonuclease, thereby digesting the plasmid fragment except for the fragment resulting from step e. In another embodiment, step f of this paragraph can be replaced with step f: incubating the plasmid or fragment resulting from step d with one or more nicking endonucleases that recognize two restriction sites, resulting in cleavage within the double-stranded DNA molecule.

[0133] In a particular embodiment, a DNA molecule comprising an expression cassette encoding FVIII flanked by inverted repeats (as described in section 5.4) can be provided by culturing a host cell containing the DNA molecule or plasmid and releasing the DNA molecule or plasmid from the host cell, as provided in steps a and b of the previous paragraph. Alternatively, such a DNA molecule can be synthesized in a cell-free system or a combination of a cell-free system and a host cell-based system. For example, chemical synthesis of DNA fragments and plasmids of various sizes and sequences is known and widely used in the art. The fragments can be chemically synthesized and then ligated or recombined in the host cell by any means known in the art. In another embodiment, the DNA molecule or plasmid can be provided by in vitro replication. Various methods can be used for in vitro replication, including amplification by polymerase chain reaction (PCR). PCR methods for replicating DNA fragments or plasmids of various sizes are well known and widely used in the art, for example, as described in Molecular Cloning: A Laboratory Manual, 4th Edition, by Michael Green and Joseph Sambrook, ISBN 978-1-936113-42-2 (2012) (incorporated herein by reference in its entirety). In some embodiments, the method of in vitro replication may be isothermal DNA amplification. In some embodiments, steps a and b can be replaced by providing DNA molecules by chemical synthesis or PCR. In another embodiment, steps a, b, c, and d can be replaced by providing DNA molecules by chemical synthesis.

[0134] In one embodiment, the method provided herein can be used to prepare a hairpin-ended DNA comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. providing a double-stranded DNA molecule as described in Section 5.4; b. incubating the DNA molecule with at least one nicking enzyme under conditions resulting in nicking of the double-stranded DNA molecule, thereby generating at least two stoichiometric DNA fragments; c. denaturing the DNA fragments; d. annealing the DNA fragments, whereby at least one DNA fragment comprises the expression cassette and a single-stranded DNA overhang capable of annealing intramolecularly, thereby generating hairpinned inverted repeats at both ends of said DNA; and e. annealing the DNA fragment to at least one of the nicking enzymes. and incubating the resulting mixture with an exonuclease of 0.1 to 100 μg / ml, thereby digesting the stoichiometric DNA fragments of step b, except for the hairpin end fragments comprising the expression cassette obtained from step d. In specific embodiments, step b of the method of this paragraph produces at least two, at least three, at least four, at least five, at least six, or more stoichiometric fragments. In further embodiments, step b of the method of this paragraph produces at least two stoichiometric DNA fragments, the DNA fragments comprising the expression cassette being stoichiometrically equivalent to the DNA molecule provided in step a. In some embodiments, the digestion-resistant hairpin end fragments comprising the expression cassette obtained from step e of this paragraph may be approximately stoichiometrically equivalent compared to the DNA molecule provided in step a.

[0135] In some embodiments, a hairpin-ended DNA molecule comprising an expression cassette encoding FVIII (or a functional fragment thereof) can be prepared using the methods provided herein, comprising: a. providing a double-stranded DNA molecule as described in section 5.4; b. incubating the DNA molecule with at least one nicking enzyme under conditions resulting in nicking of the double-stranded DNA molecule, thereby generating at least two stoichiometric DNA fragments; c. denaturing the DNA fragments into single-stranded DNA; d. annealing the sense and antisense strands of the DNA fragment comprising the expression cassette, whereby the sense and / or antisense strands comprise single-stranded DNA overhangs capable of annealing intramolecularly, thereby generating hairpinned inverted repeats at both ends of the DNA fragment comprising the expression cassette; and e. incubating the DNA fragments with at least one exonuclease, thereby digesting the stoichiometric DNA fragments of step b, except for the hairpin-ended fragments comprising the expression cassette obtained from step d.

[0136] In some embodiments, the methods provided herein can be used to prepare a hairpin-ended DNA molecule comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. providing a double-stranded DNA molecule as described in section 5.4; b. incubating the DNA molecule with at least one nicking enzyme under conditions resulting in nicking of the double-stranded DNA molecule; c. denaturing the double-stranded DNA, thereby generating at least two stoichiometric DNA fragments; and d. annealing the DNA fragments, whereby at least one DNA fragment comprises the expression cassette and a single-stranded DNA overhang capable of annealing intramolecularly, thereby generating hairpinned inverted repeats at both ends of the DNA fragment.

[0137] In a further aspect, a hairpin-ended DNA molecule encoding FVIII (or a functional fragment thereof) can be prepared using the methods provided herein, comprising at least one pot (e.g., a container, vessel, well, tube, plate, or other receptacle) containing a double-stranded DNA molecule as described in Section 5.4 in an aqueous buffer, to which (i) a nicking enzyme, (ii) a denaturing agent (e.g., a base), (iii) an annealing agent (e.g., an acid), and (iv) an exonuclease are sequentially added. The ability to perform the methods provided herein as one-pot reactions would provide at least a further advantage in that the process for producing hairpin-ended DNA molecules could be completed without the need to purify any intermediates, contaminants (e.g., enzymes) or DNA digestion by-products (i.e., nucleotides, oligos or single-stranded DNA fragments) between steps (i)-(iv) of the method, thereby providing a method that is favorable in terms of cost and risk of production failure (e.g., by minimizing purification losses, reducing required starting materials, tighter control of process variables, etc.).

[0138] In a further embodiment, the methods provided herein can be used to produce hairpin-ended DNA molecules encoding FVIII (or functional fragments thereof), comprising: a. providing a pot (e.g., a container, vessel, well, tube, plate, or other receptacle) containing a double-stranded DNA molecule described in Section 5.4 and at least one nicking enzyme under conditions that result in nicking of the double-stranded DNA molecule; b. denaturing and annealing the DNA (e.g., by changing temperature, pH, or buffer composition); and c. adding an exonuclease without the need to purify an intermediate (e.g., between steps a and c). In a specific embodiment, in the method of this paragraph, the pot contains in step a at least one species of double stranded DNA molecule (e.g., a plasmid or derivative thereof), at least one species of a nicking enzyme and an aqueous buffer, and in step c at least one species of hairpin ended DNA, at least one species of a nicking enzyme, an exonuclease and an aqueous buffer containing at least one species of DNA digestion product (e.g., dNMPs, dinucleotides and / or short oligos).

[0139] In a further embodiment, the method provided herein can be used to produce a hairpin-ended DNA molecule encoding FVIII (or a functional fragment thereof), comprising: a. providing a double-stranded DNA molecule as described in section 5.4 and at least one nicking enzyme in at least one pot (e.g., a container, vessel, well, tube, plate, or other receptacle) under conditions that result in nicking of the double-stranded DNA molecule; b. denaturing the DNA molecule, thereby generating a DNA fragment comprising the expression cassette and flanked by two single-stranded DNA overhangs; c. annealing the single-stranded DNA overhangs intramolecularly, thereby generating hairpinned inverted repeats at both ends of the DNA fragment resulting from step b; and d. adding an exonuclease to the pot, thereby digesting the DNA fragments of the DNA molecule in step b, except for the fragment resulting from step c.

[0140] In some embodiments, a hairpin-ended DNA molecule comprising an expression cassette encoding FVIII (or a functional fragment thereof) can be prepared using the methods provided herein, comprising: a. providing a double-stranded DNA molecule as described in section 5.4 and at least one nicking enzyme in at least one pot (e.g., a container, vessel, well, tube, plate, or other receptacle) under conditions that result in nicking of the double-stranded DNA molecule, thereby generating at least two stoichiometric DNA fragments; b. denaturing the DNA fragments; c. annealing the DNA fragments, whereby at least one DNA fragment comprises a single-stranded DNA overhang capable of intramolecularly annealing, thereby generating hairpinned inverted repeats at both ends of the DNA fragment comprising the expression cassette; and d. adding at least one exonuclease to the pot, thereby digesting the DNA fragments of step a, except for the hairpin-ended DNA fragments comprising the expression cassette obtained from step c. In further embodiments, in step a of the method of this paragraph, at least 2, at least 3, at least 4, at least 5, at least 6, or more stoichiometric fragments are generated.

[0141] In some embodiments, the methods provided herein can be used to prepare hairpin-ended DNA molecules comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. providing a double-stranded DNA molecule as described in section 5.4 in at least one pot (e.g., a container, vessel, well, tube, plate, or other receptacle); b. adding at least one nicking enzyme to the pot under conditions that result in nicking of the double-stranded DNA molecule, thereby generating at least two stoichiometric DNA fragments; c. denaturing the DNA fragments; d. annealing the DNA fragments, thereby generating hairpinned inverted repeats at both ends of the DNA fragment comprising the expression cassette, and e. adding at least one exonuclease to the pot, thereby digesting the stoichiometric DNA fragments of step b, except for the hairpin-ended fragments comprising the expression cassette obtained from step d. In a specific embodiment, the pot of the method of this paragraph comprises one species of DNA molecule in step a, and comprises at least 2, at least 3, at least 4, at least 5, at least 6, or more stoichiometric fragments compared to the DNA molecule of step a in step b, where the fragments containing the expression cassette are stoichiometrically equivalent to the DNA molecule provided in step a. In a specific embodiment, the pot of the method of this paragraph in step b comprises at least 2, at least 3, at least 4, at least 5, at least 6, or more stoichiometric fragments compared to the DNA molecule of step a, where the fragments containing the expression cassette are stoichiometrically equivalent to the DNA molecule provided in step a. In a non-limiting example, the pot of the method of this paragraph in step b comprises three stoichiometric DNA fragments, such as (i) two fragments lacking an expression cassette and one fragment containing an expression cassette, where the fragments containing the expression cassette are stoichiometrically equivalent to the DNA molecule provided in step a.In further specific embodiments, the pot in step e of the method of this paragraph comprises a digestion-resistant hairpin-end DNA molecule in an amount that is stoichiometrically equivalent to the DNA molecule provided in step a. In specific embodiments, the pot in step e of the method of this paragraph comprises at most an amount of digestion-resistant hairpin-end DNA molecule that is approximately stoichiometrically equivalent to the DNA molecule provided in step a. In some embodiments, the pot in step e of the method of this paragraph comprises at most an amount of digestion-resistant hairpin-end DNA molecule that is approximately stoichiometrically equivalent to the DNA molecule provided in step a, where the total mass of the DNA molecule is reduced approximately by the ratio of the nucleotides present in the hairpin-end DNA molecule comprising the expression cassette divided by the nucleotides present in the DNA molecule provided in step a. In specific embodiments, the pot in step e of the method of this paragraph comprises at most an amount of digestion-resistant hairpin-end DNA molecule that is approximately stoichiometrically equivalent to the expression cassette of the DNA molecule provided in step a.

[0142] In some embodiments, the methods provided herein can be used to prepare hairpin-ended DNA molecules comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. culturing a host cell comprising the DNA molecule described in Section 5.4 under conditions that result in amplification of the DNA molecule; b. releasing the DNA molecule from the host cell; c. adding the DNA molecule to at least one pot (e.g., a container, vessel, well, tube, plate, or other receptacle); and d. nicking at least one double-stranded DNA molecule under conditions that result in nicking of the double-stranded DNA molecule. a nicking enzyme to the pot, thereby generating at least two stoichiometric DNA fragments; e. denaturing the DNA fragments; f. annealing the DNA fragments, whereby at least one DNA fragment contains the expression cassette and a single-stranded DNA overhang capable of annealing intramolecularly, thereby generating hairpinned inverted repeats at both ends of the DNA fragment; and g. adding at least one exonuclease to the pot, whereby the stoichiometric DNA fragments of step f are digested except for the hairpin end fragments containing the expression cassette obtained from step f. In a specific embodiment, the pot in step g of the method of this paragraph contains an approximately stoichiometrically equivalent amount of digestion-resistant hairpin end DNA molecules compared to the DNA molecules provided in step c.

[0143] In some embodiments, the methods provided herein can be used to prepare hairpin-ended DNA molecules comprising an expression cassette encoding FVIII (or a functional fragment thereof), the method comprising: a. culturing a host cell comprising a plasmid as described in Section 5.4 under conditions that result in amplification of the plasmid; b. releasing the plasmid from the host cell; c. adding the plasmid to at least one pot (e.g., a container, vessel, well, tube, plate, or other receptacle); and d. isolating at least one of the plasmids under conditions that result in nicking of the plasmid. adding a nicking enzyme to the pot, thereby generating at least two stoichiometric DNA fragments; e. denaturing the DNA fragments; f. annealing the DNA fragments, thereby generating hairpinned inverted repeats at both ends of the DNA fragment, whereby at least one DNA fragment contains the expression cassette and a single-stranded DNA overhang capable of annealing intramolecularly; g. adding at least one exonuclease to the pot, thereby digesting the stoichiometric DNA fragments of step f, except for the hairpin-end fragments containing the expression cassette obtained from step f. In a specific embodiment, the pot in step g of the method of this paragraph contains an approximately stoichiometrically equivalent amount of digestion-resistant hairpin-end DNA molecules compared to the plasmid provided in step c.

[0144] The order of the steps of the method is listed in the method for exemplary purposes. In certain embodiments, the steps of the method are performed in the order in which they appear as described herein. In some embodiments, the steps of the method can be performed in a different order than the order in which they appear as described herein. Specifically, in some embodiments, the steps of the method of making a hairpin end DNA molecule can be performed in the order in which they appear, or in the order listed alphabetically as described herein from a to e, or from a to g. Alternatively, the steps of the method of making a hairpin end DNA molecule can be performed other than in the order in which they appear as described herein. In one embodiment, step c (incubating the DNA molecule with one or more nicking endonucleases that recognize four restriction sites resulting in four nicks) can be performed before step b (releasing the plasmid from the host cell) when the host cell naturally expresses, is engineered to express, or otherwise contains one or more nicking endonucleases. In another embodiment, step f (incubating the plasmid or fragment resulting from step d with a restriction enzyme, or incubating the plasmid or fragment resulting from step d with one or more nicking endonucleases) can be performed before step d (denaturing and thereby creating a DNA fragment that contains an expression cassette encoding FVIII (or a functional fragment thereof) and is flanked by two single-stranded DNA overhangs) or before step c (incubating the DNA molecule with one or more nicking endonucleases). In addition, one or more steps can be combined into one step that performs all the functions of the separate steps. In certain embodiments, step a (culturing the host cell) can be combined with step c (incubating the DNA molecule with one or more nicking endonucleases) when the host cell naturally expresses, is engineered to express, or otherwise contains one or more nicking endonucleases.In other embodiments, step f (incubating the plasmid or fragment resulting from step d with a restriction enzyme or incubating the plasmid or fragment resulting from step d with one or more nicking endonucleases) can be combined with step c (incubating the DNA molecule with one or more nicking endonucleases) by incubating with the nicking endonucleases or restriction enzymes recited together in steps f and c. Thus, the present disclosure provides that the steps can be performed in various combinations and permutations according to the state of the art.

[0145] Additional steps can be added to the methods provided herein before all steps of the method, after all steps of the method, or between any of the steps of the method. In one embodiment, the method provided herein further comprises step h (repairing the nicks with ligase to form circular DNA). In another embodiment, step h of repairing the nicks with ligase to form circular DNA is performed after all other steps of the methods described herein.

[0146] As described further below in Sections 5.4.1 and 5.5, the hairpin formed at the end of the DNA molecule is determined by the nature of the overhang between the restriction sites for the nicking endonuclease. Thus, by designing the nature, including the sequence and structural nature, of the overhang between the restriction sites for the nicking endonuclease in accordance with Sections 5.4.1 and 5.5, the method can be used to produce one, two, or more hairpinned ends. In one embodiment, the method produces hairpin end DNA comprising one hairpin end. In another embodiment, the method produces hairpin end DNA consisting of one hairpin end. In yet another embodiment, the method produces hairpin end DNA comprising two hairpin ends. In a further embodiment, the method produces hairpin end DNA consisting of two hairpin ends.

[0147] The methods provided herein can be used to produce DNA molecules that include artificial sequences, natural DNA sequences, or sequences with both natural DNA sequences and artificial sequences. In one embodiment, the method produces hairpin-ended DNA molecules that include artificial sequences. In another embodiment, the method produces hairpin-ended DNA molecules that include natural sequences. In yet another embodiment, the method produces hairpin-ended DNA molecules that include both natural and artificial sequences. In a particular embodiment, the method produces hairpin-ended DNA molecules that include a viral inverted terminal repeat (ITR). In yet another embodiment, the method produces hairpin-ended DNA molecules that include a hairpinned inverted repeat that lacks a RABS. In yet another embodiment, the method produces hairpin-ended DNA molecules that include two hairpinned inverted repeats, both hairpinned inverted repeats lacking a RABS. In another embodiment, the method produces hairpin-ended DNA molecules that include two hairpinned inverted repeats, both hairpinned inverted repeats lacking a TRS. In a further embodiment, the method produces a hairpin-ended DNA molecule comprising two hairpinned inverted repeats, both of which lack a RABS and a TRS. In another embodiment, the method produces a hairpin-ended DNA molecule comprising two hairpinned inverted repeats, both of which lack promoter activity (e.g., P5 promoter activity) and transcription activity (e.g., transcription start site [TSS]), a hairpin-ended DNA molecule. In another embodiment, the method produces a hairpin-ended DNA molecule comprising two hairpinned inverted repeats, both of which lack a RABS, promoter activity (e.g., P5 promoter activity), transcription activity (e.g., transcription start site [TSS]), and a TRS. In yet another embodiment, the method produces a hairpin-ended DNA molecule comprising two hairpinned inverted repeats, both of which lack a RABS. In yet another embodiment, the method produces a hairpin-ended DNA molecule containing one or two ITRs that lacks RAPS.In further embodiments, the method produces hairpin-ended DNA molecules comprising a viral genome. In some embodiments, the viral genome is an engineered viral genome that includes one or more non-viral genes in an expression cassette. In certain embodiments, the viral genome is an engineered viral genome in which one or more viral genes have been knocked out. In some specific embodiments, the viral genome is an engineered viral genome in which the replication-associated protein ("RAP" i.e., Rep or NS1) gene, the capsid (Cap) gene, or both the RAP gene and the Cap gene have been knocked out. In other embodiments, the viral genome is a parvovirus genome. In yet other embodiments, the parvovirus is a dependoparvovirus, a bocaparvovirus, an erythroparvovirus, a protoparvovirus, or a tetraparvovirus. In one embodiment, the parvovirus is an adeno-associated virus (e.g., AAV1, AAV2, AAV4, AAV5, AAV6, AAV7, AAV8, or AAV9).

[0148] The steps performed in the various methods provided herein are described in further detail below: host cells and host cell culture embodiments are described in Section 5.3.1, releasing DNA molecules from host cells embodiments are described in Section 5.3.2, denaturing DNA molecules embodiments are described in Section 5.3.3, annealing embodiments are described in Section 5.3.5, incubating DNA molecules with a nicking endonuclease or restriction enzyme embodiments are described in Section 5.3.4, incubating with an exonuclease embodiments are described in Section 5.3.6, and ligating embodiments are described in Section 5.3.7. Thus, the present disclosure provides methods that include permutations and combinations of the various embodiments of the steps described herein.

[0149] 5.3.1 Host Cells and Host Cell CultivationThe present disclosure provides that various host cells can be cultured to amplify DNA molecules. The host cell for use in the methods provided herein can be a eukaryotic host cell, a prokaryotic host cell, or any transformable organism capable of replicating or amplifying recombinant DNA molecules. In some embodiments, the host cell can be a microbial host cell. In further embodiments, the host cell can be a host microbial cell selected from bacteria, yeast, fungi, or any of a variety of other microbial cells that can be applied to replicating or amplifying DNA molecules. Bacterial host cells include Escherichia coli, Klebsiella oxytoca, Anaerobiospirillum succiniciproducens, Actinobacillus succinogenes, Mannheimia succiniciproducens, Rhizobium etli, Bacillus subtilis, Corynebacterium glutamicum, Gluconobacter oxydans, Zymomonas mobilis, Lactococcus lactis, and others. lactis, Lactobacillus plantarum, Streptomyces coelicolor, Clostridium acetobutylicum, Pseudomonas fluorescens, and Pseudomonas putida.The yeast or fungal host cell can be of any species selected from Saccharomyces cerevisiae, Schizosaccharomyces pombe, Kluyveromyces lactis, Kluyveromyces marxianus, Aspergillus terreus, Aspergillus niger, Pichia pastoris, Rhizopus arrhizus, Rhizobus oryzae, etc. E. coli is a particularly useful host cell because it is a well-characterized microbial cell and has been widely used in molecular cloning. Other particularly useful host cells include yeast, such as Saccharomyces cerevisiae. It will be appreciated that any suitable microbial host cell known in the art can be used to amplify the DNA molecules.

[0150] Similarly, eukaryotic host cells for use in the methods provided herein can be any eukaryotic cell capable of replicating or amplifying recombinant DNA molecules as known and used in the art. In some embodiments, host cells for use in the methods provided herein can be mammalian host cells. In further embodiments, the host cells can be human or non-human mammalian host cells. In other embodiments, the host cells can be insect host cells. Some commonly used non-human mammalian host cells include CHO, mouse myeloma cell lines (e.g., NS0, SP2 / 0), rat myeloma cell lines (e.g., YB2 / 0), and BHK. Some commonly used human host cells include HEK293 and its derivatives, HT-1080, PER.C6, and Huh-7. In certain embodiments, the host cell is selected from the group consisting of HeLa, NIH3T3, Jurkat, HEK293, COS, CHO, Saos, SF9, SF21, High 5, NS0, SP2 / 0, PC12, YB2 / 0, BHK, HT-1080, PER.C6, and Huh-7.

[0151] The host cells can be cultured as each host cell is known and cultured in the art. Culture conditions and culture media for different host cells can vary as known and practiced in the art. For example, bacterial or other microbial host cells can be cultured at 37°C, with an agitation speed of up to 300 rpm, with or without forced aeration. Some insect host cells can be optimally cultured at generally 25-30°C, without agitation or with an agitation speed of up to 150 rpm, with or without forced aeration. Some mammalian host cells can be optimally cultured at 37°C, without agitation or with an agitation speed of up to 150 rpm, with or without forced aeration. In addition, conditions for culturing various host cells can be determined by investigating the growth curves of the host cells under various conditions, as known and practiced in the art. Culture media and conditions for several widely used host cells are described in Molecular Cloning: A Laboratory Manual, 4th Edition, by Michael Green and Joseph Sambrook, ISBN 978-1-936113-42-2 (2012), which is incorporated herein by reference in its entirety.

[0152] 5.3.2 Releasing DNA molecules from the host cell The DNA molecules can be released from the host cells by various methods as known and practiced in the art. For example, the DNA molecules can be released by degrading the host cells physically, mechanically, enzymatically, chemically, or by a combination of physical, mechanical, enzymatic, and chemical action. In some embodiments, the DNA molecules can be released from the host cells by exposing the cells to a solution of a cell lysis reagent. The cell lysis reagent includes detergents such as Triton, SDS, Tween, NP-40, and / or CHAPS. In other embodiments, the DNA molecules can be released from the host cells by exposing the host cells to a difference in osmolality, for example, exposing the host cells to a hypotonic solution. In other embodiments, the DNA molecules can be released from the host cells by exposing the host cells to a solution of high or low pH. In certain embodiments, the DNA molecules can be released from the host cells by subjecting the host cells to an enzymatic treatment, for example, treatment with lysozyme. In some further embodiments, the DNA molecule can be released from the host cell by exposing the host cell to any combination of detergent, osmolarity pressure, high or low pH, and / or enzymes (e.g., lysozyme).

[0153] Alternatively, the DNA molecules can be released from the host cells by exerting a physical force on the host cells. In one embodiment, the DNA molecules can be released from the host cells by applying force directly to the host cells, for example, with a Waring blender and a Polytron. The Waring blender utilizes high speed rotating blades to break up the cells, while the Polytron draws in tissue through a long shaft containing the rotating blades. In another embodiment, the DNA molecules can be released from the host cells by applying shear stress or force to the host cells. Various homogenizers can be used to force the host cells through a narrow space, thereby shearing the cell membrane. In some embodiments, the DNA molecules can be released from the host cells by liquid-based homogenization. In one specific embodiment, the DNA molecules can be released from the host cells using a Dounce homogenizer. In another specific embodiment, the DNA molecules can be released from the host cells using a Potter-Elvehjem homogenizer. In yet another specific embodiment, the DNA molecules can be released from the host cells using a French press. Other physical forces that release DNA molecules from host cells include manual grinding, e.g., with a mortar and pestle, in which the host cells are often frozen, e.g., in liquid nitrogen, and then ground with a mortar and pestle, during which the tensile strength of the cellulose and other polysaccharides in the cell walls breaks down the host cells.

[0154] In addition, the DNA molecules can be released from the host cells by subjecting the cells to freeze and thaw cycles. In some embodiments, the host cell suspension is frozen for several such freeze and thaw cycles and then thawed. In some embodiments, the DNA molecules can be released from the host cells by subjecting the host cells to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 freeze and thaw cycles.

[0155] The above-mentioned methods for releasing DNA molecules from a host cell are not mutually exclusive. Thus, the present disclosure provides that DNA molecules can be released from a host cell by any combination of the DNA release methods provided in this Section 5.3.2.

[0156] 5.3.3 Denaturing DNA molecules DNA molecules can be denatured in a variety of ways known and practiced in the art. The process of denaturing DNA molecules can separate DNA molecules from double-stranded DNA (dsDNA) to single-stranded DNA (ssDNA). To separate the two DNA strands, the temperature can be increased until the DNA unwinds, the hydrogen bonds that hold the two strands together weaken, and finally, they split. The process of splitting double-stranded DNA into single strands is known as DNA denaturation or DNA denaturing.

[0157] In some embodiments, denaturing the DNA molecule can separate the two DNA strands of one or more segments of the dsDNA molecule while keeping the other segment(s) of the DNA molecule as dsDNA. In some embodiments, denaturing the DNA molecule can separate all DNA strands of one or more segments of the dsDNA molecule into ssDNA strands. In some further embodiments, denaturing the DNA molecule can separate the dsDNA in a segment between a first and a second restriction site for a nicking endonuclease on the top and bottom strands of the DNA into ssDNA while keeping the other portions of the DNA molecule (e.g., a DNA molecule described in Section 5.4) as dsDNA, thereby generating an overhang between the first and second restriction sites. In certain embodiments, denaturing the DNA molecule can separate dsDNA into ssDNA in a segment between a third and a fourth restriction site for a nicking endonuclease on the top and bottom strands of the DNA while keeping other portions of the DNA molecule (e.g., a DNA molecule described in Section 5.4) as dsDNA, thereby generating an overhang between the third and the fourth restriction sites. In another embodiment, denaturing the DNA molecule can separate dsDNA into ssDNA in a segment between a first and a second restriction site and a segment between a third and a fourth restriction site for a nicking endonuclease on the top and bottom strands of the DNA while keeping other portions of the DNA molecule (e.g., a DNA molecule described in Section 5.4) as dsDNA, thereby (1) decomposing the DNA molecule into two daughter DNA molecules, and (2) generating one overhang between the first and the second restriction site and one overhang between the third and the fourth restriction site. In one embodiment, the overhang between the first and second restriction sites for the nicking endonuclease can be a top strand 5' overhang.In another embodiment, the overhang between the first and second restriction sites for the nicking endonuclease can be a bottom strand 3' overhang. In yet another embodiment, the overhang between the third and fourth restriction sites for the nicking endonuclease can be a top strand 3' overhang. In a further embodiment, the overhang between the third and fourth restriction sites for the nicking endonuclease can be a bottom strand 5' overhang. In some embodiments, denaturing the DNA molecules can separate the DNA molecules in any combination of the embodiments provided herein.

[0158] The overhangs may vary in length depending on the distance between each restriction site for the nicking endonuclease. In one embodiment, the overhangs may be identical in length and / or sequence. In another embodiment, the overhangs may vary in length and / or sequence. In some embodiments, the top strand 5' overhangs are at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 9, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides.In another embodiment, the top strand 5' overhang has a length of about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108, about 109, about 110, about 111, about 112, about 113, about 114, about 115 about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, or more nucleotides.In certain embodiments, the bottom strand 3' overhang has a length of at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59 , at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides.In further embodiments, the bottom strand 3' overhang has a length of about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about It can be about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100 or more nucleotides.In yet another embodiment, the top strand 3' overhang has a length of at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, 9, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides.In other embodiments, the top strand 3' overhang has a length of about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108, about 109, about 110, about 111, about 112, about 113, about 114, about 115 about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, or more nucleotides.In some embodiments, the bottom strand 5' overhang has a length of at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, 9, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 nucleotides.In other embodiments, the bottom strand 5' overhang has a length of about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108, about 109, about 110, about 111, about 112, about 113, about 114, about 115 about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, or more nucleotides.

[0159] As known and practiced in the art, DNA molecules can be denatured by heat, by changing the pH in the environment of the DNA molecules, by increasing the salt concentration, or by any combination of these and other known means. The present disclosure provides that the DNA molecules can be denatured in the present method by using denaturation conditions that selectively separate dsDNA into ssDNA in the segments between the first and second restriction sites and / or the segments between the third and fourth restriction sites on the top and bottom strands of the DNA, optionally while keeping other portions of the DNA molecule as dsDNA. In some embodiments, the denaturation completely separates dsDNA into ssDNA. Such selective separation of dsDNA into ssDNA can be achieved by controlling the denaturation conditions and / or the time that the DNA molecules are exposed to the denaturation conditions. In one embodiment, the DNA molecule is denatured at a temperature of at least 70°C, at least 71°C, at least 72°C, at least 73°C, at least 74°C, at least 75°C, at least 76°C, at least 77°C, at least 78°C, at least 79°C, at least 80°C, at least 81°C, at least 82°C, at least 83°C, at least 84°C, at least 85°C, at least 86°C, at least 87°C, at least 88°C, at least 89°C, at least 90°C, at least 91°C, at least 92°C, at least 93°C, at least 94°C, or at least 95°C. In another embodiment, the DNA molecules are denatured at a temperature of about 70° C., about 71° C., about 72° C., about 73° C., about 74° C., about 75° C., about 76° C., about 77° C., about 78° C., about 79° C., about 80° C., about 81° C., about 82° C., about 83° C., about 84° C., about 85° C., about 86° C., about 87° C., about 88° C., about 89° C., about 90° C., about 91° C., about 92° C., about 93° C., about 94° C., or about 95° C. In a specific embodiment, the DNA molecules are denatured at a temperature of about 90° C.

[0160] In addition to heat denaturation, some or all of the DNA molecules provided herein can undergo a denaturation process by the addition of various chemical agents, such as guanidine, formamide, sodium salicylate, dimethyl sulfoxide, propylene glycol, and urea. These chemical denaturants reduce the melting temperature by competing with existing nitrogen base pairs for hydrogen bond donors and acceptors, allowing isothermal denaturation. In some embodiments, the chemical agents can induce denaturation at room temperature. In some specific embodiments, an alkaline agent (e.g., NaOH) can be used to denature the DNA by changing the pH and removing protons that contribute to hydrogen bonds. In other embodiments, chemically denaturing the DNA molecules provided herein may be a gentler treatment in terms of DNA stability compared to heat-induced denaturation. In other embodiments, chemically denaturing and renature (e.g., changing the pH) the DNA molecules provided herein can be made faster by heating. In some embodiments, the DNA of the present disclosure can be replicated in bacteria, nicked, and simultaneously denatured during release from the bacteria (e.g., alkaline lysis step).

[0161] In one embodiment, the DNA molecule is denatured at a pH of at least 10, at least 10.1, at least 10.2, at least 10.3, at least 10.4, at least 10.5, at least 10.6, at least 10.7, at least 10.8, at least 10.9, at least 11, at least 11.1, at least 11.2, at least 11.3, at least 11.4, at least 11.5, at least 11.6, at least 11.7, at least 11.8, at least 11.9, at least 12, at least 12.1, at least 12.2, at least 12.3, at least 12.4, at least 12.5, at least 13, at least 13.5, or at least 14. In another embodiment, the DNA molecules are denatured at a pH of about 10, about 10.1, about 10.2, about 10.3, about 10.4, about 10.5, about 10.6, about 10.7, about 10.8, about 10.9, about 11, about 11.1, about 11.2, about 11.3, about 11.4, about 11.5, about 11.6, about 11.7, about 11.8, about 11.9, about 12, about 12.1, about 12.2, about 12.3, about 12.4, about 12.5, about 13, about 13.5, or about 14. In yet another embodiment, the DNA molecules are denatured at a salt concentration of at least 1 M, at least 1.5 M, at least 2 M, at least 2.5 M, at least 3 M, at least 3.5 M, or at least 4 M salt. In further embodiments, the DNA molecules are denatured at a salt concentration of about 1M, about 1.5M, about 2M, about 2.5M, about 3M, about 3.5M, or about 4M salt. In certain embodiments, the DNA molecules are exposed to denaturing conditions for at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 minutes. In other embodiments, the DNA molecules are exposed to denaturing conditions for about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 minutes.In some embodiments, the DNA molecules can be denatured by any combination of denaturing conditions and denaturation periods provided herein.

[0162] Denaturation conditions can be determined for the steps of the method of selectively denaturing the segments between the first and second restriction sites and the segments between the third and fourth restriction sites on the top and bottom strands of DNA while keeping the rest of the DNA molecule as dsDNA. Such selective denaturation conditions can be determined according to the nature of the DNA segments to be selectively denatured. The stability of the DNA double helix correlates with the length and the percentage of G / C content of the DNA segment. The present disclosure provides that the selective denaturation conditions can be determined by the sequence of the DNA segment to be selectively denatured or the sequence of the resulting overhang. For example, the temperature for selective denaturation can be roughly determined as Tm=2°C×number of AT pairs+4°C×number of GC pairs for the DNA sequence to be selectively denatured. Other, more accurate calculations of Tm are also known and used in the art, as described, for example, in Freier SM, et a., Proc Natl Acad Sci, 83, 9373-9377 (1986); Breslauer KJ, et al., Proc Natl Acad Sci, 83, 3746-3750 (1986); Panjkovich, A. and Melo, F. Bioinformatics 21:711-722 (2005); Panjkovich, A., et al. Nucleic Acids Res 33:W570-W572 (2005), all of which are incorporated by reference in their entirety.

[0163] The overhang may comprise a variety of DNA sequences. In one embodiment, the overhang comprises an inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the inverted repeat). In another embodiment, the overhang comprises a viral inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the viral inverted repeat). In yet another embodiment, the overhang comprises or consists of any of the embodiments of the sequences described in Sections 5.4.1, 5.4.2, 5.4.3, and 5.5. In a further embodiment, the overhang comprises or consists of any one of the sequences described in Sections 5.4.1 and 5.5. In some embodiments, the overhang does not include one or more of the viral replication-associated sequences (e.g., RABS, RBE, or TRS) described in Section 5.4.5. In some embodiments, the overhang does not include one or more of the transcription activity-associated sequences (e.g., TSS or CpG motifs) described in Section 5.4.5.

[0164] 5.3.4 Incubating DNA molecules with one or more nicking endonucleases or restriction enzymes The present disclosure provides one or more method steps for incubating a DNA molecule with one or more nicking endonucleases or restriction enzymes, as described in Section 5.4.2. Without being bound by theory, a nicking endonuclease recognizes a restriction site for the nicking endonuclease in a DNA molecule and cuts (e.g., hydrolyzes a phosphodiester bond in one DNA strand) only on one strand of the dsDNA at a site either inside or outside the restriction site for the nicking endonuclease, thereby generating a nick in the dsDNA. On the other hand, a restriction enzyme recognizes a restriction site for the restriction enzyme and cuts both strands of the dsDNA, thereby cutting the DNA molecule at or near a particular restriction site.

[0165] In various embodiments of the compositions and methods provided herein, the nicking endonuclease may be methylation-dependent, methylation-sensitive, or methylation-insensitive. A variety of nicking endonucleases known and used in the art are provided herein. In some embodiments, the nicking endonuclease for the compositions and methods provided herein may be a naturally occurring nicking endonuclease that is not 5-methylcytosine-dependent, including Nb.Bsml, Nb.BbvCI, Nb.BsrDI, Nb.Btsl, Nt.BbvCI, Nt.Alwl, Nt.CviPII, Nt.BsmAI, Nt.Alwl, and Nt.BstNBI. Nicking endonucleases for the compositions and methods provided herein can also be engineered from Type IIs restriction enzymes (e.g., Alwl, BpulOI, BbvCI, Bsal, BsmBI, BsmAI, Bsml, BspOJ, Mlyl, Mval269l, and Sapl, etc.), and methods of making nicking endonucleases can be found in references such as, for example, US 7,081,358, US 7,011,966, US 7,943,303, US 7,820,424, WO2018 / 04514, all of which are incorporated by reference herein in their entireties.

[0166] Alternatively, instead of a nicking endonuclease, a programmable nicking enzyme can be used in the compositions and methods provided herein. Such programmable nicking enzymes include, for example, Cas9 or functional equivalents thereof, such as Pyrococcus furiosus Argonaute (PfAgo) or Cpfl. Cas9 contains two catalytic domains, RuvC and HNH. Inactivating one of these domains generates a programmable nicking enzyme that can replace the nicking endonuclease for the methods and compositions provided herein. In Cas9, the RuvC domain can be inactivated by an amino acid substitution at position D10 (e.g., D10A), and the HNH domain can be inactivated by an amino acid substitution at position H840 (e.g., H840A), or at positions corresponding to these amino acids in other Cas9-equivalent proteins. Such programmable nicking enzymes can be two components: a nicking enzyme (e.g., D10A Cas9 nicking enzyme or a variant or ortholog thereof) that cleaves the target DNA, and an Argonaute or type II CRISPR / Cas endonuclease that includes a guide nucleic acid, e.g., a guide DNA or RNA (gDNA or gRNA), that targets or programs the nicking enzyme to a specific site within the target DNA (see, e.g., Hsu, et al., Nature Biotechnology 2013 31:827-832, which is incorporated herein by reference in its entirety). Programmable nicking enzymes can also be made by fusing a site-specific DNA-binding domain (targeting domain), such as the DNA-binding domain of a DNA-binding protein (e.g., a restriction endonuclease, a transcription factor, a zinc finger, or another domain that binds to DNA at a non-random location), to the nicking endonuclease so that it acts on a specific, non-random site.As is evident from the above, programmable cleavage by a programmable nicking enzyme results from a targeting domain within or fused to the nicking enzyme, or a guide molecule (gDNA or gRNA) that directs the nicking enzyme to a specific, non-random site that is a programmable site by varying the targeting domain or guide molecule. Such programmable nicking enzymes can be found in references, e.g., US 7,081,358 and WO 2010 / 021692 A, which are incorporated herein by reference in their entirety.

[0167] Suitable guide nucleic acid (e.g., gDNA or gRNA) sequences and target sites suitable for guide nucleic acids are known in the art and are widely available. Guide nucleic acids (e.g., gDNA or gRNA) are specific nucleic acid (e.g., gDNA or gRNA) sequences that recognize a target DNA region of interest and direct a programmable nicking enzyme (e.g., Cas nuclease) thereto for editing. Guide nucleic acids (e.g., gDNA or gRNA) are often composed of two parts: a targeting nucleic acid, which is a 15-20 nucleotide sequence complementary to the target DNA, and a scaffold nucleic acid that serves as a binding scaffold for the programmable nicking enzyme (e.g., Cas nuclease). A target site suitable for a guide nucleic acid must have two components: a sequence complementary to the targeting nucleic acid in the programmable nicking enzyme, and an adjacent protospacer adjacent motif (PAM). The PAM serves as a binding signal for the programmable nicking enzyme (e.g., Cas nuclease). A variety of PAMs are known, characterized, and utilized in the art, as discussed, for example, in Daniel Gleditzsch et al., RNA Biol. 16(4):504-517 (April 2019); Ryan T. Leenay et al., Mol Cell. 62(1):137-147 (Apr 7, 2016), both of which are incorporated by reference in their entireties. Exemplary gRNA and gDNA sequences targeting the main stem sequence of the AAV2 ITR include those listed in Table 1. [Table 1]

[0168] A variety of nicking endonucleases known and used in the art can be used in the methods provided herein. An exemplary list of nicking endonucleases provided as embodiments of nicking endonucleases for use in the methods and corresponding restriction sites for some of the nicking endonucleases are listed in The Restriction Enzyme Database (known in the art as REBASE), available at www.rebase.neb.com / cgi-bin / azlist?nick, and incorporated herein by reference in its entirety. In one embodiment, the nicking endonucleases that recognize the first, second, third, and / or fourth restriction sites are all for the same nicking endonuclease target sequence. In another embodiment, the first, second, third, and fourth restriction sites for the nicking endonucleases are target sequences for two different nicking endonucleases, including all possible combinations that arrange four sites for two different nicking endonuclease target sequences (e.g., a first restriction site for a first nicking endonuclease and a remaining site for a second nicking endonuclease, a first and second restriction site for a first nicking endonuclease and a remaining site for a second nicking endonuclease, etc.). In yet another embodiment, the first, second, third, and fourth restriction sites for the nicking endonuclease are target sequences for three different nicking endonucleases, including all possible combinations that arrange four sites for three different nicking endonuclease target sequences. In a further embodiment, the first, second, third, and fourth restriction sites for the nicking endonuclease are target sequences for four different nicking endonucleases. In some embodiments, the nicking endonuclease can be any one selected from those listed in Table 2. [Table 2]

[0169] Conditions under which various nicking endonucleases cleave one strand of dsDNA are known for various nicking endonucleases presented herein, including temperature, salt concentration, pH, buffering reagents, the presence or absence of specific detergents, and incubation periods to achieve a desired percentage of nicked DNA molecules. These conditions are readily available from the websites or catalogs of various suppliers of nicking endonucleases, for example, from New England BioLabs. The present disclosure provides that the step of incubating DNA molecules with one or more nicking endonucleases is performed according to incubation conditions known and practiced in the art. In some embodiments, the step of incubating DNA molecules with one or more nicking endonucleases is according to incubation conditions optimized by methods known in the art.

[0170] Various restriction enzymes known and used in the art can be used in the methods provided herein. An exemplary list of restriction enzymes provided as embodiments of the restriction enzymes for use in the methods, and the corresponding restriction sites for the restriction enzymes, are described in the New England Biolabs catalog available at neb.com / products / restriction-endonucleases, which is incorporated herein by reference in its entirety. Conditions for various restriction enzymes to cleave dsDNA are known for the various restriction enzymes provided herein, including temperature, salt concentration, pH, buffering reagents, the presence or absence of specific detergents, and incubation periods to achieve a desired percentage of nicked DNA molecules. These conditions are readily available from various suppliers of restriction enzymes, such as the websites or catalogs of New England BioLabs. The present disclosure provides that the step of incubating DNA molecules with a restriction enzyme is performed according to incubation conditions known and practiced in the art.

[0171] 5.3.5 Annealing The annealing step in the methods provided herein is performed to selectively anneal the ssDNA overhangs intramolecularly, thereby generating hairpinned inverted repeats at one end of the DNA fragments (e.g., those in Sections 5.4 and 5.5) obtained in the denaturing step (Section 5.3.3) described above. In certain embodiments, the annealing step in the methods provided herein is performed to selectively anneal the ssDNA overhangs intramolecularly, thereby generating hairpinned inverted repeats at the two ends of the DNA fragments (e.g., those in Sections 5.4 and 5.5) obtained in the denaturing step (Section 5.3.3) described above. Without being bound or otherwise limited by theory, such selective intramolecular annealing of the ssDNA overhangs is achieved because the intramolecular complementary sequences within the ssDNA overhangs make intramolecular annealing of the ssDNA overhangs thermodynamically and / or kinetically favorable compared to intermolecular annealing of the ssDNA overhangs.

[0172] Without wishing to be bound or otherwise limited by theory, it is recognized that certain lengths and / or sequences of the overhangs can make intramolecular annealing of ssDNA overhangs thermodynamically and / or kinetically favorable compared to intermolecular annealing of ssDNA overhangs. For example, linear interaction plots showing the intramolecular forces within the overhangs and the intermolecular forces between the strands and the resulting structures are depicted in Figures 2A-C. The thermodynamics and kinetics of annealing of ssDNA overhangs are determined by enthalpy (ΔH) and entropy (ΔS), among other factors. The inventors recognize that the entropy reduction in intramolecular annealing is less than the entropy reduction in intramolecular annealing because the loss of freedom of motion from the free ssDNA overhang to the intramolecularly annealed overhang is less than the reduction of freedom of motion from the free ssDNA overhang to the intermolecularly annealed overhang. On the other hand, the enthalpy increase in intramolecular annealing may be less than that in intramolecular annealing because the number of complementary nucleotide pairs in the intramolecularly annealed overhang is less than that in the intermolecularly annealed overhang (hence, fewer Watson-Crick and Hoogsteen type hydrogen bonds). The present disclosure provides that ssDNA overhangs can be designed to have a specific length, number of complementary nucleotide pairs, and percentage of GC and AT pairs such that the free energy increase (ΔG=ΔH-TΔS) of intramolecular annealing of the overhang is greater than that of intermolecular annealing, thereby making intramolecular annealing thermodynamically favored compared to intermolecular annealing. The inventors further recognize that the reaction rate of intramolecular annealing of ssDNA overhangs may be higher than that of intermolecular annealing because nucleotides in an ssDNA overhang have a higher probability of contacting each other during molecular motion than the probability of contacting nucleotides of another ssDNA overhang.The present disclosure provides that even when intramolecular annealing is thermodynamically unfavorable compared to intermolecular annealing, the superior kinetics of intramolecular annealing of ssDNA overhangs can result in the formation of intramolecularly annealed overhangs in preference to intermolecularly annealed overhangs.

[0173] The annealing step can be carried out at a variety of temperatures that favor intramolecular annealing over intermolecular annealing. In one embodiment, the ssDNA overhangs are at least 15°C, at least 16°C, at least 17°C, at least 18°C, at least 19°C, at least 20°C, at least 21°C, at least 22°C, at least 23°C, at least 24°C, at least 25°C, at least 26°C, at least 27°C, at least 28°C, at least 29°C, at least 30°C, at least 31°C, at least 32°C, at least 33°C, at least 34°C, at least 35°C, at least 36°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 43°C, at least 44°C, at least 45°C, at least 46°C, at least 47°C, at least 48°C, at least 49°C, at least 50°C, at least 51°C, at least 52°C, at least 53°C, at least 54°C, at least 55°C, at least 56°C, at least 57°C, at least 58°C, at least 59°C, at least 60°C, at least 61°C, at least 62°C, at least 63°C, at least 64°C, at least 65°C, at least 66°C, at least 67°C, at least 68°C, at least 69°C, at least 70°C, at least 71°C, at least 72°C, at least 73°C 7°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 43°C, at least 44°C, at least 45°C, at least 46°C, at least 47°C, at least 48°C, at least 49°C, at least 50°C, at least 51°C, at least 52°C, at least 53°C, at least 54°C, at least 55°C, at least 56°C, at least 57°C, at least 58°C, at least 59°C, or at least 60°C. In another embodiment, ssDNA annealing is performed at about 15°C, about 16°C, about 17°C, about 18°C, about 19°C, about 20°C, about 21°C, about 22°C, about 23°C, about 24°C, about 25°C, about 26°C, about 27°C, about 28°C, about 29°C, about 30°C, about 31°C, about 32°C, about 33°C, about 34°C, about 35°C, about 36°C, about 37°C, about 38°C, about 39°C, about 40°C, about 41°C, about 42°C, about 43°C, about 44°C, about 45°C, about 46°C, about 47°C, about 48°C, about The ssDNA overhangs are annealed at a temperature of about 7° C., about 38° C., about 39° C., about 40° C., about 41° C., about 42° C., about 43° C., about 44° C., about 45° C., about 46° C., about 47° C., about 48° C., about 49° C., about 50° C., about 51° C., about 52° C., about 53° C., about 54° C., about 55° C., about 56° C., about 57° C., about 58° C., about 59° C., or about 60° C. In a specific embodiment, the ssDNA overhangs are annealed at a temperature of at least 25° C. In another specific embodiment, the ssDNA overhangs are annealed at a temperature of about 25° C. In yet another specific embodiment, the ssDNA overhangs are annealed at room temperature.

[0174] In addition, the annealing step can be performed for various times that favor intramolecular annealing over intermolecular annealing. In certain embodiments, the ssDNA overhangs are annealed for at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, or at least 40 minutes. In other embodiments, the ssDNA overhang is annealed for about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, or about 40 minutes. In a specific embodiment, the ssDNA overhang is annealed for at least 20 minutes. In another specific embodiment, the ssDNA overhang is annealed for about 20 minutes.

[0175] In some embodiments, annealing can be achieved by lowering the temperature below the calculated melting temperature of the sense and antisense sequence pair. The melting temperature depends on the specific nucleotide base content and the properties of the solution used, such as salt concentration. The melting temperature of any given sequence and solution combination is easily calculated as known and practiced in the art.

[0176] In some embodiments, annealing can be achieved isothermally by reducing the amount of denaturing chemical agent to allow interaction between sense and antisense sequence pairs. The minimum concentration of denaturing chemical agent required to denature DNA sequence may depend on the specific nucleotide base content and the properties of the solution used, such as temperature or salt concentration. The concentration of chemical denaturant that does not cause denaturation for any given sequence and solution combination is easily identified as known and practiced in the art. The concentration of chemical denaturant can also be easily modified as known and practiced in the art. For example, the amount of urea can be reduced by dialysis or tangential flow filtration, or the pH can be changed by adding acid or base.

[0177] The annealing temperature and annealing time for intramolecular annealing are correlated with the length of the ssDNA overhang, the number of complementary nucleotide pairs, and the percentage of GC and AT pairs, and the sequence (arrangement of complementary nucleotide pairs) of the ssDNA overhang. In certain embodiments, the ssDNA overhangs provided in the methods provided herein comprise any number of nucleotides in length, as described in Section 5.3.3. In certain embodiments, the ssDNA overhangs provided in the methods provided herein comprise at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, or at least 50 intramolecular complementary nucleotide pairs. In some embodiments, the ssDNA overhangs provided in the methods provided herein comprise about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, or about 50 intramolecular complementary nucleotide pairs.In some embodiments, the ssDNA overhangs provided in the methods provided herein comprise at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, or at least 90% GC pairs among intramolecular complementary nucleotide pairs. In certain embodiments, the ssDNA overhangs provided in the methods provided herein comprise about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, or about 90% GC pairs among the intramolecular complementary nucleotide pairs.

[0178] In addition, the inventors recognize that the concentration of DNA molecules, which correlates with the concentration of overhangs, can affect the equilibrium and kinetics of intra- and intermolecular annealing of overhangs. Without being bound or otherwise limited by theory, if the concentration of overhangs is too high, the probability of intermolecular contacts between overhangs increases, which reduces the kinetic dominance of intramolecular contacts over intermolecular contacts seen at such low concentrations.

[0179] As mentioned above, in some embodiments, intramolecular interactions may occur at a faster rate, and intermolecular interactions occur at a slower rate. In some embodiments, base pair interactions involving three or more molecules (e.g., three different strands) occur at the slowest rate. In some embodiments, the kinetic rate of intramolecular interactions to intermolecular interactions is governed by the concentration of each molecule. In some embodiments, the lower the concentration of DNA strands, the faster the kinetics of intramolecular interactions or the greater the intramolecular forces.

[0180] Taken individually, the absolute free energy of forming each complementary domain of the IR or ITR can be different, resulting in regions of the IR or ITR that may locally fold earlier as the strands transition from a denatured to an annealed state. The presence of locally folded domains (e.g., a central hairpin or branched hairpins as in the AAV2 ITR as described elsewhere in this section (Section 5.4.1) and in Section 5.5) can reduce the amount of bases available for pairing with the other strand, thus reducing the likelihood of intermolecular annealing or hybridization and shifting the equilibrium from intermolecular annealing to intramolecular annealing or ITR formation.

[0181] Thus, the present disclosure provides that the annealing step can be performed at various concentrations that favor intramolecular annealing over intermolecular annealing. In some embodiments, the ssDNA overhangs are 1 or less, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, 7 or less, 8 or less, 9 or less, 10 or less, 11 or less, 12 or less, 13 or less, 14 or less, 15 or less, 16 or less, 17 or less, 18 or less, 19 or less, 20 or less, 21 or less, 22 or less, 23 or less, 24 or less, 25 or less, 26 or less, 27 or less, 28 or less, 29 or less, 30 or less, 31 or less, 32 or less, 33 or less, 34 or less, 35 or less, 36 or less, 37 or less, 38 or less, 39 or less, 40 or less, 41 or less, 42 or less, 43 or less, 44 or less, 45 or less, 46 or less, 47 or less, 48 ​​or less, 49 or less, 50 or less, 55 or less, 60 or less, 61 or less, 62 or less, 63 or less, 64 or less, 65 or less, 66 or less, 67 or less, 68 or less, 69 or less, 70 or less, 71 or less, 72 or less, 73 or less, 74 or less, 75 or less, 76 or less, 77 or less, 78 or less, 79 or less, 79 or less, 80 or less, 81 or less, 82 or less, 83 or less, 84 or less, 85 or less, 86 or less, 87 or less, 88 or less, 89 or less, 80 or less, 85 or less, 86 5 or less, 70 or less, 75 or less, 80 or less, 85 or less, 90 or less, 95 or less, 100 or less, 110 or less, 120 or less, 130 or less, 140 or less, 150 or less, 1 60 or less, 170 or less, 180 or less, 190 or less, 200 or less, 210 or less, 220 or less, 230 or less, 240 or less, 250 or less, 260 or less, 270 or less, 2 Annealed at ng / μl concentrations of 80 or less, 290 or less, 300 or less, 325 or less, 350 or less, 375 or less, 400 or less, 425 or less, 450 or less, 475 or less, 500 or less, 550 or less, 600 or less, 650 or less, 700 or less, 750 or less, 800 or less, 850 or less, 900 or less, 950 or less, 1000 or less.In certain embodiments, the ssDNA overhang is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 55, about 60, about 65, about The nucleic acid is annealed at a concentration of about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 260, about 270, about 280, about 290, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, or about 1000 ng / μl.

[0182] Similarly, the present disclosure provides that the annealing step can be performed at various molar concentrations that favor intramolecular annealing over intermolecular annealing. In some embodiments, the ssDNA overhangs are 1 or less, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, 7 or less, 8 or less, 9 or less, 10 or less, 11 or less, 12 or less, 13 or less, 14 or less, 15 or less, 16 or less, 17 or less, 18 or less, 19 or less, 20 or less, 21 or less, 22 or less, 23 or less, 24 or less, 25 or less, 26 or less, 27 or less, 28 or less, 29 or less, 30 or less, 31 or less, 32 or less, 33 or less, 34 or less, 35 or less, 36 or less, 37 or less, 38 or less, 39 or less, 40 or less, 41 or less, 42 or less, 43 or less, 44 or less, 45 or less, 46 or less, 47 or less, 48 ​​or less, 49 or less, 50 or less, 55 or less, 60 or less , 65 or less, 70 or less, 75 or less, 80 or less, 85 or less, 90 or less, 95 or less, 100 or less, 110 or less, 120 or less, 130 or less, 140 or less, 150 or less , 160 or less, 170 or less, 180 or less, 190 or less, 200 or less, 210 or less, 220 or less, 230 or less, 240 or less, 250 or less, 260 or less, 270 or less , 280 or less, 290 or less, 300 or less, 325 or less, 350 or less, 375 or less, 400 or less, 425 or less, 450 or less, 475 or less, 500 or less, 550 or less, 600 or less, 650 or less, 700 or less, 750 or less, 800 or less, 850 or less, 900 or less, 950 or less, 1000 or less nM concentrations.In certain embodiments, the ssDNA overhang is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 55, about 60, about 65 , about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 260, about 270, about 280, about 290, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000 nM. In some further embodiments, the ssDNA overhang is annealed at a concentration of 1 or less, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, 7 or less, 8 or less, 9 or less, 10 or less, 11 or less, 12 or less, 13 or less, 14 or less, 15 or less, 16 or less, 17 or less, 18 or less, 19 or less, or 20 or less μM. In yet another embodiment, the ssDNA overhang is annealed at a concentration of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 μM. In a specific embodiment, the ssDNA overhang is annealed at a concentration of about 10 nM for the DNA molecule. In another specific embodiment, the ssDNA overhang is annealed at a concentration of about 20 nM for the DNA molecule. In yet another specific embodiment, the ssDNA overhang is annealed at a concentration of about 30 nM for the DNA molecule.In a further specific embodiment, the ssDNA overhang is annealed at a concentration of about 40 nM for the DNA molecule.In yet another specific embodiment, the ssDNA overhang is annealed at a concentration of about 50 nM for the DNA molecule.In another specific embodiment, the ssDNA overhang is annealed at a concentration of about 60 nM for the DNA molecule. In one specific embodiment, the ssDNA overhang is annealed at a concentration of about 10 ng / μl for the DNA molecule. In another specific embodiment, the ssDNA overhang is annealed at a concentration of about 20 ng / μl for the DNA molecule. In yet another specific embodiment, the ssDNA overhang is annealed at a concentration of about 30 ng / μl for the DNA molecule. In yet another specific embodiment, the ssDNA overhang is annealed at a concentration of about 40 ng / μl for the DNA molecule. In one specific embodiment, the ssDNA overhang is annealed at a concentration of about 50 ng / μl for the DNA molecule. In another specific embodiment, the ssDNA overhang is annealed at a concentration of about 60 ng / μl for the DNA molecule. In yet another specific embodiment, the ssDNA overhang is annealed at a concentration of about 70 ng / μl for the DNA molecule. In one specific embodiment, the ssDNA overhang is annealed at a concentration of about 80ng / μl for the DNA molecule. In another specific embodiment, the ssDNA overhang is annealed at a concentration of about 90ng / μl for the DNA molecule. In yet another specific embodiment, the ssDNA overhang is annealed at a concentration of about 100ng / μl for the DNA molecule.

[0183] In some embodiments, the ssDNA overhang provided in the methods provided herein comprises any sequence listed in Table 3. [Table 3]

[0184] In some embodiments, the structure of the DNA molecules provided herein is the same after 2, 3, 4, 5, 10, or 20 cycles of denaturation / renaturation (e.g., denaturation as described in Section 5.3.3 and reannealing as described in this section (Section 5.3.5)). The DNA structure can be described by an ensemble of structures at or near an energy minimum. In certain embodiments, the ensemble DNA structure is the same after 2, 3, 4, 5, 10, or 20 cycles of denaturation / renaturation. In one embodiment, the folded hairpin structure formed from the ITR or IR provided herein is the same after 2, 3, 4, 5, 10, or 20 cycles of denaturation / renaturation. In another embodiment, the ensemble structure of folded hairpins is the same after 2, 3, 4, 5, 10, or 20 cycles of denaturation / renaturation.

[0185] 5.3.6 Incubation with exonuclease The present disclosure provides for incubating with an exonuclease as described in Section 3. An exonuclease cleaves nucleotides from the ends (exo) of a DNA molecule. An exonuclease can cleave nucleotides along the 5' to 3' direction, along the 3' to 5' direction, or along both directions. In certain embodiments, an exonuclease for use in the methods provided herein cleaves nucleotides without sequence specificity. In some embodiments, an exonuclease for use in the methods provided herein digests DNA fragments containing ends generated by one or more nicking endonucleases that recognize and cleave the fifth and sixth restriction sites, or by a restriction enzyme that cleaves a plasmid or a fragment of a plasmid, as provided in Section 5.4.6.

[0186] Various exonucleases known and used in the art can be used in the methods provided herein. An exemplary list of exonucleases provided as embodiments of restriction enzymes for use in the methods is available at neb.com / products / dna-modifying-enzymes-and-cloning-technologies / nucleases and is described in the New England Biolabs catalog, which is incorporated herein by reference in its entirety. Conditions under which various exonucleases digest DNA molecules are known for various exonucleases provided herein, including temperature, salt concentration, pH, buffering reagents, the presence or absence of specific detergents, and incubation periods to achieve a desired digestion rate. These conditions are readily available from various suppliers of restriction enzymes, such as the website or catalog of New England BioLabs. The present disclosure provides that the step of incubating DNA molecules with restriction enzymes is performed according to incubation conditions known and practiced in the art.

[0187] The step of incubating the exonuclease selectively digests DNA molecules with one or more ends while leaving hairpin-ended DNA molecules intact. As is evident from the description in Sections 5.3.5 and 5.5, hairpin-ended DNA molecules contain zero, one, two, or more nicks. In some embodiments, the exonuclease for use in the methods provided herein can be an exonuclease that selectively digests DNA molecules with one or more ends while leaving circular ssDNA / dsDNA molecules, or DNA molecules that contain one or more nicks but no ends, intact. In one embodiment, the exonuclease for use in the methods provided herein can be exonuclease V (RecBCD). In one embodiment, the exonuclease for use in the methods provided herein can be exonuclease VIII or truncated exonuclease VIII. Exonuclease V (RecBCD), Exonuclease VIII, and Truncated Exonuclease VIII include the selectivities described in this paragraph. In certain embodiments, the exonuclease for use in the methods provided herein can be an exonuclease that selectively digests linear segments of DNA molecules starting from one or more nicks, but cannot proceed to digest folded hairpins, terminating digestion at the hairpin and leaving ssDNA behind. In certain embodiments, the exonuclease for use to initiate one or more nicks and / or double-stranded breaks can be T7 exonuclease. Other suitable exonucleases are also known and used in the art and are provided herein, for example, as described on the websites or catalogs of various suppliers of exonucleases, including New England BioLabs.

[0188] In some embodiments, after exonuclease treatment, the DNA molecule of the present disclosure is substantially free of prokaryotic backbone sequences. In some embodiments, backbone refers to plasmid sequences that are not part of the sequence encompassing the expression cassette between the two ITRs. In some embodiments, backbone refers to vector sequences that are not part of the sequence encompassing the expression cassette between the two ITRs. In some embodiments, the isolated DNA molecule of the present disclosure is 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% free of prokaryotic backbone sequences of the parent plasmid.

[0189] 5.3.7 Repairing nicks with ligase The present disclosure provides an optional step of repairing the nick using a ligase as described in Section 3. A DNA ligase catalyzes the connection of two ends of a DNA molecule by forming one or more new covalent bonds. For example, the commonly used T4 DNA ligase catalyzes the formation of a phosphodiester bond between juxtaposed 5' phosphate and 3' hydroxyl ends in DNA. The formation of a new covalent bond catalyzed by a ligase to connect two DNA molecules is called "ligation". In certain embodiments, a DNA ligase for use in the methods provided herein ligates nucleotides without sequence specificity. In some embodiments, a DNA ligase for use in the methods provided herein ligates two ends at one nick of a DNA molecule as described in Section 5.5, thereby repairing the one nick described above. In some embodiments, a DNA ligase for use in the methods provided herein ligates each pair of two ends at two nicks of a DNA molecule as described in Section 5.5, thereby repairing the two nicks described above. In some embodiments, the DNA ligase for use in the methods provided herein ligates each pair of two ends at every nick in the DNA molecule described in Section 5.5, thereby repairing every nick in the DNA molecule. If the DNA molecule described in Section 5.5 forms a circular DNA, then every nick in the DNA molecule described in Section 5.5 is repaired. As described in Section 5.5, in some embodiments, the DNA molecule described in Section 5.5 consists of two nicks. In certain embodiments, the DNA molecule described in Section 5.5 contains two nicks. In other embodiments, the DNA molecule described in Section 5.5 consists of one nick. In yet other embodiments, the DNA molecule described in Section 5.5 contains one nick.

[0190] In some embodiments, the step of repairing the nick with a ligase can be carried out according to incubation conditions known and practiced in the art.

[0191] Various ligases known and used in the art can be used in the methods provided herein. An exemplary list of ligases provided as embodiments of ligases for use in the methods is available at neb.com / products / dna-modifying-enzymes-and-cloning-technologies / dna-ligases / dna-ligases and is described in the New England Biolabs catalog, which is incorporated herein by reference in its entirety. Conditions under which various ligases digest DNA molecules are known for various ligases provided herein, including temperature, salt concentration, pH, buffering reagents, the presence or absence of specific detergents, and incubation periods to achieve the desired digestion rate. These conditions are readily available from various suppliers of restriction enzymes, such as the websites or catalogs of New England BioLabs. Ligation conditions also correlate with the freedom of movement of the two DNA ends to be ligated. For example, two DNA ends can be brought close to each other by annealing both ends to a common DNA strand, or ligation can be enhanced by increasing the probability that two DNA ends will be close to each other. In one embodiment, the steps of the method provided in this section (Section 5.3.7) repair the nick with a ligase to generate a circular DNA, where the two DNAs at either nick of the DNA molecule described in Section 5.5 are annealed to a common DNA strand. In some embodiments, the step of repairing the nick with a ligase is performed according to incubation conditions known and practiced in the art.

[0192] In certain embodiments, the methods provided in this section 5.2 can be used to produce the hairpin-ended DNA molecules described herein at large scale, high yield, and / or high purity. In certain embodiments, large scale, high yield, and / or high purity can be achieved in a single reaction vessel. In certain embodiments, large scale is at least 1 mg, 10 mg, 100 mg, 1 g, 10 g, 100 g, 1 kg, or at least 10 kg. In certain embodiments, high yield is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, or at least 99% yield (comparing the number of plasmid copies used as input to the number of hairpin-ended DNA molecules as product). In certain embodiments, high purity is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, or at least 99% purity of the hairpin-ended DNA molecules as product as a result of the methods provided herein.

[0193] 5.4 DNA molecules used in the method The present disclosure provides various aspects and embodiments of DNA molecules for use in the methods provided herein described in Section 5.3 above. In one aspect, provided herein is a DNA molecule comprising, in the 5' to 3' direction of the top strand, i) a first inverted repeat (e.g., as described in Section 5.4.1), where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a top strand 5' overhang that includes the first inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat); and ii) an expression cassette ( and iii) a second inverted repeat (e.g., as described in Section 5.4.1), where a third and a fourth restriction site for a nicking endonuclease are located on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a top strand 3' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the top strand 5' overhang comprises the first inverted repeat. In certain embodiments, the top strand 3' overhang comprises the second inverted repeat. In certain embodiments, the top strand 5' overhang comprises a first inverted repeat and the top strand 3' overhang comprises a second inverted repeat.

[0194] In another aspect, provided herein is a nucleic acid sequence encoding a nucleic acid sequence that includes, in a 5' to 3' direction of the top strand, i) a first inverted repeat (e.g., as described in Section 5.4.1), where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a bottom strand 3' overhang that includes the first inverted repeat, or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat); and ii) an expression cassette (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2). 5.4.3), and iii) a second inverted repeat (e.g., as described in Section 5.4.1), where a third and a fourth restriction site for a nicking endonuclease are disposed on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a bottom strand 5' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the bottom strand 5' overhang comprises the first inverted repeat. In certain embodiments, the bottom strand 5' overhang comprises the second inverted repeat. In certain embodiments, the bottom strand 3' overhang comprises a first inverted repeat and the bottom strand 5' overhang comprises a second inverted repeat.

[0195] In yet another aspect, provided herein is an expression cassette that includes, in a 5' to 3' direction of the top strand, i) a first inverted repeat (e.g., as described in Section 5.4.1), where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a top strand 5' overhang that includes the first inverted repeat, or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat); and ii) an expression cassette. and iii) a second inverted repeat (e.g., as described in Sections 5.4.1) where a third and a fourth restriction site for a nicking endonuclease are located on opposing strands near the second inverted repeat such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a bottom strand 5' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the top strand 5' overhang comprises the first inverted repeat. In certain embodiments, the bottom strand 5' overhang comprises the second inverted repeat. In certain embodiments, the top strand 5' overhang comprises a first inverted repeat and the bottom strand 5' overhang comprises a second inverted repeat.

[0196] In a further aspect, provided herein is a nucleic acid sequence encoding a nucleic acid sequence comprising, in a 5' to 3' direction of the top strand, i) a first inverted repeat (e.g., as described in Section 5.4.1), where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a bottom strand 3' overhang that includes the first inverted repeat, or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat); and ii) an expression cassette (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2). iii) a second inverted repeat (e.g., as described in Section 5.4.1), where a third and a fourth restriction site for a nicking endonuclease are disposed on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a top strand 3' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the bottom strand 3' overhang comprises the first inverted repeat. In certain embodiments, the top strand 3' overhang comprises the second inverted repeat. In certain embodiments, the bottom strand 3' overhang comprises a first inverted repeat and the top strand 3' overhang comprises a second inverted repeat.

[0197] In one aspect, provided herein is a nucleic acid sequence comprising, in a 5' to 3' direction of a top strand, i) a first inverted repeat (e.g., as described in Section 5.4.1), where first and second target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking by the programmable nicking enzyme results in a top strand 5' overhang that comprises the first inverted repeat, or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat); and ii) an expression cassette ( and iii) a second inverted repeat (e.g., as described in Sections 5.4.1), where third and fourth target sites for a guide nucleic acid for the programmable nicking enzyme are located on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking by the programmable nicking enzyme results in a top strand 3' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the top strand 5' overhang comprises the first inverted repeat. In certain embodiments, the top strand 3' overhang comprises the second inverted repeat. In certain embodiments, the top strand 5' overhang comprises a first inverted repeat and the top strand 3' overhang comprises a second inverted repeat.

[0198] In another aspect, provided herein is a nucleic acid sequence comprising, in a 5' to 3' direction of the top strand, i) a first inverted repeat (e.g., as described in Section 5.4.1), where first and second target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking by the programmable nicking enzyme results in a bottom strand 3' overhang that comprises the first inverted repeat, or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat); and ii) an expression cassette (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2). 5.4.3), and iii) a second inverted repeat (e.g., as described in Section 5.4.1), where third and fourth target sites for a guide nucleic acid for the programmable nicking enzyme are located on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking by the programmable nicking enzyme results in a bottom strand 5' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the bottom strand 3' overhang comprises the first inverted repeat. In certain embodiments, the bottom strand 5' overhang comprises a second inverted repeat. In certain embodiments, the bottom strand 3' overhang comprises a first inverted repeat and the bottom strand 5' overhang comprises a second inverted repeat.

[0199] In yet another aspect, provided herein is an expression cassette comprising, in a 5' to 3' direction of the top strand, i) a first inverted repeat (e.g., as described in Section 5.4.1), where first and second target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking by the programmable nicking enzyme results in a top strand 5' overhang that includes the first inverted repeat, or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat). and iii) a second inverted repeat (e.g., as described in Sections 5.4.1), where third and fourth target sites for a guide nucleic acid for the programmable nicking enzyme are located on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking by the programmable nicking enzyme results in a bottom strand 5' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the top strand 5' overhang comprises the first inverted repeat. In certain embodiments, the bottom strand 5' overhang comprises the second inverted repeat. In certain embodiments, the top strand 5' overhang comprises a first inverted repeat and the bottom strand 5' overhang comprises a second inverted repeat.

[0200] In a further aspect, provided herein is a nucleic acid sequence comprising, in a 5' to 3' direction of the top strand, i) a first inverted repeat (e.g., as described in Section 5.4.1), where first and second target sites for a guide nucleic acid for a programmable nicking enzyme are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking by the programmable nicking enzyme results in a bottom strand 3' overhang that includes the first inverted repeat, or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat); and ii) an expression cassette (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2). iii) a second inverted repeat (e.g., as described in Section 5.4.1), where third and fourth target sites for a guide nucleic acid for the programmable nicking enzyme are located on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking by the programmable nicking enzyme results in a top strand 3' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the bottom strand 3' overhang comprises the first inverted repeat. In certain embodiments, the top strand 3' overhang comprises a second inverted repeat. In certain embodiments, the bottom strand 3' overhang comprises a first inverted repeat and the top strand 3' overhang comprises a second inverted repeat.

[0201] The DNA molecules provided herein include various features or have various embodiments, as described in Section 3 and the above paragraphs of this section (Section 5.4), which are further described in various subsections below: embodiments for inverted repeats including a first inverted repeat and / or a second inverted repeat are described in Section 5.4.1, embodiments for restriction enzymes, nicking endonucleases and their respective restriction sites are described in Sections 5.4.2 and 5.3.4, embodiments for programmable nicking enzymes and their target sites are described in Section 5.3.4, embodiments for expression cassettes are described in Section 5.4.3, embodiments for plasmids and vectors are described in Section 5.4.6, and embodiments for DNA molecules containing less than four restriction sites for nicking endonucleases are described in Section 5.4.7. Thus, the present disclosure provides DNA molecules that include any permutation and combination of the various embodiments of DNA molecules and embodiments of the features of DNA molecules described herein. In further embodiments, the arrangement of the ITRs, expression cassette, restriction sites for the nicking endonuclease or restriction enzyme, and the programmable nicking enzyme and its target site can be any of the arrangements described in Sections 5.3.3, 5.3.4, 5.3.5, 5.4.1, 5.4.2, 5.4.3 5.4.7, and 5.5.

[0202] In one aspect, provided herein is a vector comprising, in a 5' to 3' direction of the top strand, i) a first viral replication-deficient inverted repeat (e.g., as described in Sections 5.4.1 and 5.4.5), where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a top strand 5' overhang that includes the first inverted repeat, or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat); and ii) an expression cassette. and iii) a second viral replication-deficient inverted repeat (e.g., as described in Sections 5.4.1 and 5.4.5), where third and fourth restriction sites for a nicking endonuclease are located on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a top strand 3' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the top strand 5' overhang comprises the first viral replication-deficient inverted repeat. In certain embodiments, the top strand 3' overhang comprises a second viral replication-deficient inverted repeat. In certain embodiments, the top strand 5' overhang comprises a first viral replication-deficient inverted repeat and the top strand 3' overhang comprises a second viral replication-deficient inverted repeat.

[0203] In another aspect, provided herein is a vector comprising, in a 5' to 3' direction on a top strand, i) a first viral replication-deficient inverted repeat (e.g., as described in Sections 5.4.1 and 5.4.5), where a first and a second restriction site for a nicking endonuclease are located on opposing strands near the first inverted repeat such that nicking is effected as described in Sections 5.3.3, 5.3.4, and 5.4.2; ii) an expression cassette (e.g., as described in Section 5.4.3); and iii) a second viral replication-deficient inverted repeat (e.g., as described in Sections 5.4.1 and 5.4.5). and a second viral replication-deficient inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2, or depicted in Figures 2A and 2C), where third and fourth restriction sites for a nicking endonuclease are located on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a bottom strand 5' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the bottom strand 3' overhang comprises the first viral replication-deficient inverted repeat. In certain embodiments, the bottom strand 5' overhang comprises the second viral replication-deficient inverted repeat. In certain embodiments, the bottom strand 3' overhang comprises a first viral replication-deficient inverted repeat and the bottom strand 5' overhang comprises a second viral replication-deficient inverted repeat.

[0204] In yet another aspect, provided herein is a vector comprising, in a 5' to 3' direction of the top strand, i) a first viral replication-deficient inverted repeat (e.g., as described in Sections 5.4.1 and 5.4.5), where first and second restriction sites for a nicking endonuclease are located on opposing strands near the first inverted repeat, such that upon separation of the top strand from the bottom strand of the first inverted repeat, nicking results in a top strand 5' overhang that includes the first inverted repeat, or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the first inverted repeat); and ii) an expression cassette. and iii) a second viral replication-deficient inverted repeat (e.g., as described in Sections 5.4.1 and 5.4.5), where third and fourth restriction sites for a nicking endonuclease are located on opposing strands near the second inverted repeat, such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a bottom strand 5' overhang that comprises the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat). In certain embodiments, the top strand 5' overhang comprises the first viral replication-deficient inverted repeat. In certain embodiments, the bottom strand 5' overhang comprises a second viral replication-deficient inverted repeat. In certain embodiments, the top strand 5' overhang comprises a first viral replication-deficient inverted repeat and the bottom strand 5' overhang comprises a second viral replication-deficient inverted repeat.

[0205] In a further aspect, provided herein is a vector comprising, in a 5' to 3' direction of the top strand, i) a first viral replication-deficient inverted repeat (e.g., as described in Sections 5.4.1 and 5.4.5), where a first and a second restriction site for a nicking endonuclease are located on opposing strands near the first inverted repeat such that nicking results in the bottom strand as described in Sections 5.3.3, 5.3.4, and 5.4.2; ii) an expression cassette (e.g., as described in Section 5.4.3); and iii) a second viral replication-deficient inverted repeat (e.g., as described in Sections 5.4.1 and 5.4.5). and a second viral replication-deficient inverted repeat (e.g., as described in Sections 5.3.3, 5.3.4, and 5.4.2, or depicted in Figures 2B and 2C), where third and fourth restriction sites for a nicking endonuclease are located on opposing strands near the second inverted repeat or a fragment thereof (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the second inverted repeat), such that upon separation of the top strand from the bottom strand of the second inverted repeat, nicking results in a top strand 3' overhang that includes the second inverted repeat. In certain embodiments, the bottom strand 3' overhang includes the first viral replication-deficient inverted repeat. In certain embodiments, the top strand 3' overhang includes the second viral replication-deficient inverted repeat. In certain embodiments, the bottom strand 3' overhang comprises a first viral replication-deficient inverted repeat and the top strand 3' overhang comprises a second viral replication-deficient inverted repeat.

[0206] The DNA molecule provided herein can be a DNA molecule in its natural environment or an isolated DNA molecule.In certain embodiments, the DNA molecule is a DNA molecule in its natural environment.In some embodiments, the DNA molecule is an isolated DNA molecule. In one embodiment, the isolated DNA molecule is at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 9 %, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% pure DNA molecules.In another embodiment, the isolated DNA molecule is about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 100%, about 101%, about 102%, about 103%, about 104%, about 105%, about 106%, about 107%, about 10 The DNA molecules may be about 4%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% pure. Other embodiments of the isolated DNA molecules provided herein in terms of purity are further described in Section 5.4.8, which can be combined in any suitable combination with the embodiments provided in this paragraph.

[0207] Because DNA molecules can be entirely engineered (e.g., synthetically or recombinantly produced), the DNA molecules provided herein, including those in Section 3 and herein in Section 5.4, may lack certain sequences or characteristics, as further described in Section 5.4.5.

[0208] 5.4.1 Backward Iteration The ITRs or IRs provided in Section 3 and this section (Section 5.4.1) can, for example, upon performing the method steps described in Sections 3, 5.3.3, 5.3.4, and 5.3.5, form hairpinned ITRs in hairpin-ended DNA molecules provided in Section 5.5. Thus, in some embodiments, the ITRs or IRs provided in Section 3 and this section (Section 5.4.1) can include any of the IR or ITR embodiments provided in Section 3 and Section 5.5 and additional embodiments provided in this section (Section 5.4.1), in any combination.

[0209] Most DNA in cells contains two strands joined by Watson-Crick base pairing that segregates most functional groups and limits structural and functional diversity. Although a desirable property for a molecule whose function is to store genetic information, single-stranded viruses have evolved to take advantage of the intramolecular interactions of linear single-stranded DNA (ssDNA) to form secondary structures that add another layer of functional complexity. One of the major contributing factors is that these ssDNA viral genomes consist of only one strand that folds back on itself to form a hairpin.

[0210] The secondary structure of a single-stranded DNA molecule can represent the pattern of complementary base pairing that occurs between the constituent nucleotides based on the initial DNA sequence. The sequence, represented as a sequence of four letters (one for each nucleotide type), is a single strand of nucleotides that are generally believed to form different secondary structures with the lowest free energy governed by thermodynamic interactions.

[0211] "Inverted repeat" or "IR" refers to a single stranded nucleic acid sequence that includes a palindromic region. This palindromic region includes a sequence of nucleotides and its reverse complement, i.e., a "palindromic sequence" as further described below, on the same strand as further described below. Under denatured conditions, meaning conditions in which hydrophobic stacking attractions between bases are nullified, the IR nucleic acid sequence exists in a random coil state (e.g., at high temperature, in the presence of chemical agents, at high pH, ​​etc.). As conditions become more physiological, the IR described above can fold into a secondary structure in which the outermost regions are held together non-covalently by base pairing. In some embodiments, the IR can be an ITR. In certain embodiments, the IR includes an ITR. In some embodiments, the IR can be a hairpinned inverted repeat. In certain embodiments, the inverted repeat, when folded on itself, includes at least one hairpin loop (also known as a stem loop) that results in an unpaired loop of single stranded DNA being formed when the DNA strand folds and base pairs with another section of the same strand. In certain embodiments, the inverted repeat may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 such hairpin loop structures.

[0212] "Inverted terminal repeat", "terminal repeat", "TR", or "ITR" refers to an inverted repeat region at or near the end of a single-stranded DNA molecule, or an inverted repeat at or within a single-stranded overhang of a dsDNA molecule. The ITR can fold on itself as a result of a palindromic sequence in the ITR. In one embodiment, the ITR is at or near one end of the ssDNA. In another embodiment, the ITR is at or near one end of the dsDNA. In yet another embodiment, each of the two ITRs is at or near two respective ends of the ssDNA. In a further embodiment, each of the two ITRs is at or near two respective ends of the dsDNA. In some embodiments, the non-ITR portion of the ssDNA or dsDNA is heterologous to the ITR. In certain embodiments, the non-ITR portion of the ssDNA or dsDNA is homologous to the ITR. In denatured conditions, meaning conditions in which hydrophobic stacking attractions between bases are nullified, nucleic acid sequences containing ITRs exist in a random coil state (e.g., at high temperature, in the presence of chemical agents, at high pH, ​​etc.). In some embodiments, as conditions become more suitable for annealing as described in Section 5.3.5, the ITRs can fold on themselves to form a structure that is held together non-covalently by base pairing while leaving the heterologous non-ITR portion of the dsDNA intact, or the heterologous non-ITR portion of the ssDNA molecule can hybridize with a second ssDNA molecule that contains the inverted complementary sequence of the heterologous DNA molecule. The resulting complex of the two hybridized DNA strands contains three distinct regions, namely, a first folded single-stranded ITR covalently linked to a double-stranded DNA region that is, in turn, covalently linked to a second folded single-stranded ITR. In certain embodiments, the ITR sequence can start at one of the restriction sites for a nicking endonuclease described in Sections 3, 5.3.4, and 5.4.2, and end at the last base before the dsDNA.In one embodiment, in contrast to linear double-stranded DNA molecules, the ITRs present at the 5' and 3' ends of the top and bottom strands at either end of the DNA molecule can fold inward and face each other (e.g., 3' to 5', 5' to 3', or vice versa), thus not exposing free 5' or 3' ends on either side of the nucleic acid duplex. When an ITR folds on itself, in some embodiments, the dsDNA in the folded ITR can be immediately adjacent to the dsDNA in the non-ITR portion of the DNA molecule, generating a nick adjacent to the dsDNA, or in other embodiments, the dsDNA in the folded ITR can be one or more nucleotides away from the dsDNA in the non-ITR portion of the DNA molecule, generating a "ssDNA gap" adjacent to the dsDNA. Two ITRs located in a non-ITR DNA sequence are referred to as an "ITR pair." In some embodiments, when the ITR is in its presumed folded state, it is resistant to exonuclease digestion (eg, exonuclease V), for example, for 1 hour or more at 37°C.

[0213] The interface between the terminal bases of the ITR that folds to its secondary structure and the terminal bases of the DNA hybridized duplex can be further stabilized by stacking interactions (e.g., coaxial stacking) between the base pairs adjacent to the nick or ssDNA gap, and these interactions are sequence-dependent. In the case of structures similar to a nick, there may be an equilibrium between two conformations, where the first conformation is very close to that of an intact double helix in which the stacking between the base pairs located on either side of the nick is preserved, and the other conformation corresponds to a complete loss of stacking at the nick site, thus inducing a twist in the DNA. Nicked molecules are known to migrate somewhat slower than intact molecules of the same size during polyacrylamide and agarose gel electrophoresis. In some cases, this retardation is enhanced at higher temperatures. It is believed that the rapid equilibration between the stacked / linear and unstacked / bent conformations of the nicks results in the specific retardation characteristic of nicked DNA molecules, which directly affects the mobility of the DNA molecules during gel electrophoresis.

[0214] Without wishing to be bound by theory, it is believed that cellular proteins can recognize, bind to and process parallel 5' and 3' ends as double-stranded breaks, which can have deleterious effects on the fate of the DNA within the cell. Thus, ITRs can prevent premature and unwanted degradation of expression cassettes with ITRs as provided in Sections 3 and 5.5, as well as in this section (Section 5.4.1), at one or both of their two ends.

[0215] By placing the first and second restriction sites for a nicking endonuclease on opposite strands and near the inverted repeat, and subsequent separation of the top strand from the bottom strand of the inverted repeat, the resulting overhang can fold on itself to form a double-stranded end containing at least one restriction site for a nicking endonuclease. In some embodiments, the folded ITR resembles the secondary structure conformation of a viral ITR. In one embodiment, the ITRs are located at both the 5' and 3' ends of the bottom strand (e.g., left and right ITRs). In another embodiment, the ITRs are located at both the 5' and 3' ends of the top strand. In yet another embodiment, one ITR is located at the 5' end of the top strand and the other ITR is located at the opposite ends of the bottom strand (e.g., left ITR at the 5' end on the top strand and right ITR at the 5' end of the bottom strand). In yet another embodiment, one ITR is located at the 3' end of the top strand and the other ITR is located at the 3' end of the bottom strand.

[0216] In some aspects, the present disclosure provides DNA molecules that include palindromic sequences. A "palindromic sequence" or "palindrome" is a self-complementary DNA sequence that can fold back to form a stretch of dsDNA in the self-complementary region under conditions that favor intramolecular annealing. In some embodiments, a palindromic sequence comprises a stretch of contiguous polynucleotides that is identical when read forward to when read backward on a complementary strand. In one embodiment, a palindromic sequence comprises a stretch of polynucleotides that is identical when read forward to when read backward on a complementary strand, interrupted by one or more stretches of non-palindromic polynucleotides. In another embodiment, a palindromic sequence comprises a stretch of polynucleotide that when read in the forward direction is 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to when read in the reverse direction on the complementary strand. In yet another embodiment, the palindrome sequence, when read in the forward direction, has a length that is 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77% greater than when read in the reverse direction on the complementary strand. , 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the first or second polynucleotides, wherein such stretches are interrupted by one or more stretches of non-palindromic polynucleotides.In some specific embodiments, provided herein are double-stranded DNA molecules having first and second restriction sites for a nicking endonuclease on opposing strands of the double-stranded DNA, wherein after nicking and separating the top strand from the bottom strand of the inverted repeat, the resulting inverted repeat has a length between the first and second restriction sites as described above that is 50%, 51%, 52%, 53% or more shorter when read in a forward direction than when read in a reverse direction on the complementary strand. , 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the ssDNA encoding one or more palindromic sequences. The ssDNA encoding one or more palindromic sequences can fold back on itself to form double-stranded base pairs that include secondary structures (e.g., hairpin loops, or three-way junctions).

[0217] For example, under appropriate conditions, such as those described in Sections 5.3.3, 5.3.4, and 5.3.5, an IR or ITR provided in this section (Section 5.4.1) can fold to form a hairpin structure as described in this section (Section 5.4.1) and Section 5.5, containing a stem, a main stem, a loop, a turning point, a bulge, a branch, a branched loop, an internal loop, and / or any combination or permutation of the structural features described in Section 5.5.

[0218] In one embodiment, an IR or ITR for the methods and compositions provided herein comprises one or more palindromic sequences. In some embodiments, an IR or ITR described herein comprises a palindromic sequence or domain that can form a branched hairpin structure in addition to forming a main stem domain. In some embodiments, an IR or ITR comprises a palindromic sequence that can form any number of branched hairpins. In specific embodiments, an IR or ITR comprises a palindromic sequence that can form 1-30 branched hairpins, or any subrange of 1-30 branched hairpins. In some specific embodiments, an IR or ITR comprises a palindromic sequence that can form 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 branched hairpins. In some embodiments, the IR or ITR comprises a sequence capable of forming two branched hairpin structures resulting in a three-way junction domain (T-shaped). In some embodiments, the IR or ITR comprises a sequence capable of forming three branched hairpin structures resulting in a four-way junction domain (or a cruciform structure). In some embodiments, the IR or ITR comprises a sequence capable of forming a non-T-shaped hairpin structure, e.g., a U-shaped hairpin structure. In some embodiments, the IR or ITR comprises a sequence capable of forming an interrupted U-shaped hairpin structure that includes a series of bulges and base pair mismatches. In some embodiments, the branched hairpins all have stems and / or loops of the same length. In some embodiments, one branched hairpin is smaller (e.g., truncated) than the other branched hairpin. Some exemplary embodiments of hairpin structures and structural elements of hairpin structures are depicted in FIG. 1.

[0219] "Hairpin closing base pair" refers to the first base pair following the mismatched loop sequence. Certain stem-loop sequences have a preferred closing base pair (e.g., GC in the AAV2 ITR). In one embodiment, the stem-loop sequence contains a GC pair as the closing base pair. In another embodiment, the stem-loop sequence contains a CG pair as the closing base pair.

[0220] "ITR closing base pair" refers to the first and last nucleotides that form a base pair within a folded ITR. The terminal base pair is usually the pair of nucleotides in the main stem domain that is closest to the non-ITR sequence of the DNA molecule (e.g., an expression cassette). The ITR closing base pair can be any type of base pair (e.g., CG, AT, GC, or TA). In one embodiment, the ITR closing base pair is a GC base pair. In another embodiment, the ITR closing base pair is an AT base pair. In yet another embodiment, the ITR closing base pair is a CG base pair. In a further embodiment, the ITR closing base pair is a TA base pair.

[0221] The present disclosure provides that DNA secondary structures can be computationally predicted as known and practiced in the art. DNA secondary structures can be represented in several ways, such as Squiggle plots, graph representations, dot-bracket notations, circular plots, arc diagrams, mountain plots, dot plots, etc. In circular plots, the backbone is represented by a circle and the base pairs are represented by arcs inside the circle. In arc diagrams, the DNA backbone is depicted as a straight line and the nucleotides of each base pair are connected by arcs. Both circular and arc plots allow identification of similarities and differences in secondary structures.

[0222] One of the many methods for DNA secondary structure prediction uses the nearest neighbor model, which minimizes the total free energy associated with the DNA structure. The minimum free energy is estimated by summing the individual energy contributions from base pair stacking, hairpins, bulges, internal loops, and multi-branched loops. The energy contributions of these elements are sequence and length dependent and have been experimentally determined. The separation of the sequence into stem-loops and sub-stems can be represented, for example, by showing the structure as a graph plot. In a linear interaction plot, each residue is represented on the horizontal axis and a semi-elliptical line connects the bases that are paired with each other (e.g., Figures 2A and B).

[0223] In some embodiments, the ITRs promote long-term persistence of the nucleic acid molecule in the nucleus of the cell. In some embodiments, the ITRs promote permanent persistence of the nucleic acid molecule in the nucleus of the cell (e.g., for the entire lifespan of the cell). In some embodiments, the ITRs promote stability of the nucleic acid molecule in the nucleus of the cell. In some embodiments, the ITRs inhibit or prevent degradation of the nucleic acid molecule in the nucleus of the cell.

[0224] In certain embodiments, the IR or ITR may comprise any viral ITR, hi other embodiments, the IR or ITR may comprise a synthetic palindromic sequence capable of forming a palindromic hairpin structure that does not expose the 5' or 3' ends at the outermost apex or turn point of the repeat.

[0225] In some embodiments, a single-stranded ITR sequence extending from one nucleotide of the ITR closing base pair to the other nucleotide of the ITR closing base pair has a Gibbs free energy of unfolding (ΔG) in the range of −10 kcal / mol to −100 kcal / mol under physiological conditions. In one embodiment, the Gibbs free energy of unfolding (ΔG) (kcal / mol) referred to in the preceding sentence is -10 or less (meaning ≦-10, including, for example, -20, -30, etc.), -11 or less, -12 or less, -13 or less, -14 or less, -15 or less, -16 or less, -17 or less, -18 or less, -19 or less, -20 or less, -21 or less, -22 or less, -23 or less, -24 or less, -25 or less, -26 or less, -27 or less, -28 or less, -29 or less, -30 or less, -31 or less, -32 or less, -33 or less, -34 or less, -35 or less, -36 or less, -37 or less, -38 or less, -39 or less, -40 or less, -41 or less, -42 or less, -43 or less, -44 or less, -45 or less, -46 or less, -47 or less, -48 or less, -49 or less. Below, -50 or less, -51 or less, -52 or less, -53 or less, -54 or less, -55 or less, -56 or less, -57 or less, -58 or less, -59 or less, -60 or less, -61 or less, -62 or less Lower, -63 or less, -64 or less, -65 or less, -66 or less, -67 or less, -68 or less, -69 or less, -70 or less, -71 or less, -72 or less, -73 or less, -74 or less, -75 or less , -76 or less, -77 or less, -78 or less, -79 or less, -80 or less, -81 or less, -82 or less, -83 or less, -84 or less, -85 or less, -86 or less, -87 or less, -88 or less, -89 or less, -90 or less, -91 or less, -92 or less, -93 or less, -94 or less, -95 or less, -96 or less, -97 or less, -98 or less, -99 or less, or -100 or less.In another embodiment, the Gibbs free energy of unfolding (ΔG) (kcal / mol) referred to in the preceding sentence is about -10 (meaning ≦-10, including, for example, -20, -30, etc.), about -11, about -12, about -13, about -14, about -15, about -16, about -17, about -18, about -19, about -20, about -21, about -22, about -23, about -24, about -25, about -26, about -27, about -28, about -29, about -30, about -31, about -32, about -33, about -34, about -35, about -36, about -37, about -38, about -39, about -40, about -41, about -42, about -43, about -44, about -45, about -46, about -47, about -48, about -49, about -50, about -51, about -52, about -53, about -54, about -55, about -56, about -57, about -58, about -59, about -60, about -61, about -62, about -63, about -64, about -65, about -66, about -67, about -68, about -69, about -70, about -71, about -72, about -73, about -74, about -75, about -76, about -77, about -78, about -79, about -80, about -81, about -82, about -83, about -84, about -85, about -86, about -87, about -88, about -89, about -90, about -91, about -92, about -93, about -94, about -95, about -96, about -97, about -98, about -99, or about -100. In some embodiments, the ITR sequence extending from one nucleotide of the ITR closing base pair to the other nucleotide of the ITR closing base pair has a Gibbs free energy of unfolding (ΔG) in the range of -26 kcal / mol to -95 kcal / mol under physiological conditions. In some embodiments, the ITR sequence extending from one nucleotide of the ITR closing base pair to the other nucleotide of the ITR closing base pair contributes all of the Gibbs free energy of unfolding (ΔG) for the ITR sequence under physiological conditions.

[0226] In some embodiments, in the folded state, the single-stranded IR or ITR has an overall Watson-Crick self-complementarity of about 50% to 98%. In one embodiment, in the folded state, the single-stranded IR or ITR has an overall Watson-Crick self-complementarity of about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 100%, about 101%, about 102%, about 103%, about 104%, about 105%, about 106%, about 107%, about 108%, about 109%, about 110%, about 111%, about 112%, about 113%, about 114%, about 115%, about 116%, about 117%, about 118%, about 119%, about 120%, about 121%, about 122%, about 123%, about 124%, about 125%, about 126%, about 127%, about 128%, about 12 %, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% overall Watson-Crick self-complementarity. In another embodiment, in the folded state, the single-chain IR or ITR is at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% overall Watson-Crick self complementarity.In some embodiments, in the folded state, the IR or ITR has an overall Watson-Crick complementarity of about 60%-98%. In some embodiments, the single-stranded IR or ITR has an overall GC content of about 60%-95%. In certain embodiments, the single-stranded IR or ITR has a total GC content of at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%. In another embodiment, the single-stranded IR or ITR has a total GC content of about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, or about 95%. In some embodiments, the single-stranded IR has a total GC content of about 60-91%.

[0227] Table 4 lists the folding free energies, GC content, percent complementarity, and lengths of exemplary ITRs, and Table 5 lists the sequences of the ITRs in Table 4. [Table 4] [Table 5]

[0228] The DNA molecules for the methods and compositions provided herein can include IRs or ITRs of various origins. In one embodiment, the IRs or ITRs in the DNA molecule are viral ITRs. "Viral ITRs" include any viral terminal repeat or synthetic sequence that includes at least one minimally essential replication origin and a region that includes a palindromic hairpin structure. In one embodiment, the viral ITRs are from Parvoviridae. In another embodiment, the viral ITRs from Parvoviridae include a minimally essential replication origin that includes at least one viral replication-associated protein binding sequence ("RABS"), which refers to a DNA sequence to which viral DNA replication-associated protein ("RAP") and its isoforms encoded by Rep and / or NS1 of Parvoviridae can bind. In some embodiments, the RABS is a Rep binding sequence ("RBS"). In some embodiments, the RABS includes a Rep binding sequence ("RBS"), and Rep can bind to two elements within the ITR. It can bind to a nucleotide sequence within the stem structure of the ITR (i.e., a nucleotide sequence recognized by the Rep protein for replication of the viral nucleic acid molecule). Such an RBS is also referred to as an RBE (Rep binding element). Rep can also bind to a nucleotide sequence forming a small palindrome containing a single tip of an internal hairpin within the ITR, thereby stabilizing the binding between Rep and the ITR. Such an RBS is also referred to as an RBE'. In another embodiment, the viral ITR from Parvoviridae contains a RABS that includes an NS1 binding element ("NSBE") to which the replication-associated viral protein NS1 can bind. In another embodiment, the RABS is an NS1 binding element ("NSBE") to which the replication-associated viral protein NS1 can bind. In some embodiments, the viral ITR is from Parvoviridae and includes a terminal separation site ("TRS") that allows the viral DNA replication-associated protein NS1 and / or Rep to endonucleolytically nick a sequence in the TRS.In yet another embodiment, the viral ITRs comprise at least one RBS or NSBE and at least one TRS. In the context of the production of viral or recombinant RAP (i.e., Rep or NS1)-based viral genomes, the ITRs mediate replication and viral packaging. As unexpectedly discovered by the present inventors and provided herein, double-stranded linear DNA vectors having ITRs similar to viral ITRs can be produced without the need for Rep or NS1 proteins, and thus DNA replication is independent of RABS or TRS sequences. Thus, RABS and TRS can be optionally encoded within the nucleotide sequences disclosed herein, but are not essential, providing flexibility regarding the design of the ITRs. In one embodiment, the ITRs for the methods and compositions provided herein do not comprise at least one RABS (e.g., one RABS, two RABS, or more than two RABS). In another embodiment, the ITRs for the methods and compositions provided herein do not comprise any RABS. In another embodiment, the ITRs for the methods and compositions provided herein do not comprise at least one RBS. In another embodiment, the ITR for the methods and compositions provided herein does not include any RBS. In another embodiment, the ITR for the methods and compositions provided herein does not include an RBE. In another embodiment, the ITR for the methods and compositions provided herein does not include an RBE'. In another embodiment, the ITR for the methods and compositions provided herein does not include an RBE and an RBE'. In another embodiment, the ITR for the methods and compositions provided herein does not include an NSBE. In yet another embodiment, the ITR for the methods and compositions provided herein does not include a TRS. In a further embodiment, the ITR for the methods and compositions provided herein does not include at least one RABS (e.g., one RABS, two RABS, or more than two RABS) and does not include a TRS. In a further embodiment, the ITR for the methods and compositions provided herein does not include any RABS and does not include a TRS.In further embodiments, an ITR for the methods and compositions provided herein comprises an RBS (i.e., an RBE and / or an RBE'), a TRS, or both an RBS (i.e., an RBE and / or an RBE') and a TRS. In further embodiments, an ITR for the methods and compositions provided herein comprises an NBSE, a TRS, or both an NBSE and a TRS.

[0229] "ITR pair" refers to two ITRs in a single DNA molecule. In some embodiments, both ITRs in an ITR pair are derived from a wild-type viral ITR (e.g., AAV2 ITR) with reverse-complementary sequences throughout their entire length. An ITR can be considered a wild-type sequence even if it has one or more nucleotides that deviate from the standard naturally occurring sequence, as long as the changes do not affect the nature of the sequence and the overall three-dimensional structure. The present disclosure provides that, in some embodiments, the insertion, deletion, or substitution of one or more nucleotides can create a restriction site for a nicking endonuclease without changing the overall three-dimensional structure of the viral ITR. In some aspects, the deviating nucleotides represent conservative sequence changes. In certain embodiments, the sequences of the ITRs provided herein can have at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the standard sequence (e.g., as measured using BLAST with default settings) and have a restriction site for a nicking endonuclease such that the 3D structure has the same shape in geometric space. In other embodiments, the sequences of the ITRs provided herein can have about 95%, about 96%, about 97%, about 98%, or about 99% sequence identity to the standard sequence (e.g., as measured using BLAST with default settings) and have a restriction site for a nicking endonuclease such that the 3D structure has the same shape in geometric space.

[0230] In some embodiments, the DNA molecules for the methods and compositions provided herein comprise a pair of wt-ITRs. In certain specific embodiments, the DNA molecules for the methods and compositions provided herein comprise a pair of wt-ITRs selected from the group shown in Table 6. Table 6 shows the AAV serotype 1 (AAV1), AAV serotype 2 (AAV2), AAV serotype 3 (AAV3), AAV serotype 4 (AAV4), AAV serotype 5 (AAV5), AAV serotype 6 (AAV6), AAV serotype 7 (AAV7), AAV serotype 8 (AAV8), AAV serotype 9 (AAV9), AAV serotype 10 (AAV10), AAV serotype 11 (AAV11), or AAV serotype 12 (AAV12); AAVrh8, AAVrhlO, AAV-DJ, and AAV-DJ8 genomes (see, e.g., NCBI: NC 002077; NC 001401; NC001729; NC001829; NC006152; NC 006260; NC Exemplary ITRs are shown from the same or different serotypes, or from different parvoviruses, including ITRs from warm-blooded animals (avian AAV (AAAV), bovine AAV (BAAV), canine, equine, and ovine AAV), ITRs from B19 parvovirus (GenBank Accession No. NC 000883), minute virus of mice (MVM) (GenBank Accession No. NC 001510); goose: goose parvovirus (GenBank Accession No. NC 001701); snake: snake parvovirus 1 (GenBank Accession No. NC 006148). [Table 6] TIFF2025504404000008.tif150165

[0231] In some embodiments, the DNA molecules for the methods and compositions provided herein comprise all or part of the genome of a parvovirus. Parvovirus genomes are linear and 3.9-6.3 kb in size, with the coding region flanked by terminal repeats that can fold into hairpin-like structures, either different (heterotelomeric, e.g., HBoV) or identical (homotelomeric, e.g., AAV2). In one embodiment, the DNA molecules for the methods and compositions provided herein comprise two different ITRs at the two ends of the DNA molecule. In another embodiment, the DNA molecules for the methods and compositions provided herein comprise two identical ITRs at the two ends of the DNA molecule. In yet another embodiment, the DNA molecules for the methods and compositions provided herein comprise two different ITRs corresponding to the two HBoV ITRs at the two ends of the DNA molecule. In a further embodiment, the DNA molecules for the methods and compositions provided herein comprise two identical ITRs corresponding to the AAV2 ITRs at the two ends of the DNA molecule.

[0232] In certain embodiments, the ITRs in the DNA molecules provided herein can be AAV ITRs. In other embodiments, the ITRs can be non-AAV ITRs. In one embodiment, the ITRs in the DNA molecules provided herein can be derived from AAV ITRs or non-AAV TRs. In some specific embodiments, the ITRs can be derived from any one of the Parvoviridae family, which includes parvoviruses and dependoviruses (e.g., canine parvovirus, bovine parvovirus, mouse parvovirus, porcine parvovirus, human parvovirus B-19). In other specific embodiments, the ITRs can be derived from the SV40 hairpin that serves as the origin of SV40 replication. The Parvoviridae family of viruses consists of two subfamilies: Parvovirinae, which infect vertebrates, and Densovirinae, which infect invertebrates. Thus, in one embodiment, the ITRs can be derived from any one of the Parvovirinae subfamilies. In another embodiment, the ITRs can be derived from any one of the Densovirinae subfamilies.

[0233] Compared with the T-shaped AAV ITR, human erythrovirus B19 has an ITR that ends in an imperfect palindrome that can fold into a long linear duplex with a few unpaired nucleotides and generate a series of small but highly conserved mismatch bulges. In some embodiments, any parvovirus ITR can be used as an ITR (e.g., wild-type or modified ITR) for the DNA molecules provided herein, or can act as a template ITR for modification and then incorporation into the DNA molecules provided herein. In some specific embodiments, the parvovirus from which the ITR of the DNA molecule originates is a dependovirus, an erythroparvovirus, or a bocaparvovirus. In other specific embodiments, the ITR of the DNA molecule provided herein is derived from AAV, B19, or HBoV. In one specific embodiment, the ITRs of the DNA molecules provided herein are derived from the HBoV genome (Accession No. JQ923422) nucleotides 19-122 (TCTTGGAATCCAATATGTCTGCCGGCTCAGTCATGCCTGCGCTGCGCGCAGCGCGCTGCGCGCGCGCATGATCTAATCGCCGGCAGACATATTGGATTCCAAGA SEQ ID NO: 183). In another specific embodiment, the ITRs of the DNA molecules provided herein are derived from the B19 genome (Accession No. AY386330) nucleotides 129-237 (GGGTTGGCTCTGGGCCAGCTTGCTTGGGGTTGCCTTGACACTAAGACAAGCGGCGCGCCGCTTGATCTTAGTGGCACGTCAACCCCAAGCGCTGGCCCAGAGCCAACCC SEQ ID NO: 184). In certain embodiments, the serotype of the AAV ITRs selected for the DNA molecules provided herein can be based on the tissue tropism of the serotype. AAV2 has broad tissue tropism, AAV1 preferentially targets neurons and skeletal muscle, AAV5 preferentially targets neurons, retinal pigment epithelial cells, and photoreceptor cells, and AAV6 preferentially targets skeletal muscle and lung.AAV8 preferentially targets liver, skeletal muscle, heart, and pancreatic tissues. AAV9 preferentially targets liver, skeletal, and lung tissues. In one embodiment, the ITR or modified ITR of the DNA molecule provided herein is based on the AAV2 ITR. In one embodiment, the ITR or modified ITR of the DNA molecule provided herein is based on the AAV1 ITR. In one embodiment, the ITR or modified ITR of the DNA molecule provided herein is based on the AAV5 ITR. In one embodiment, the ITR or modified ITR of the DNA molecule provided herein is based on the AAV6 ITR. In one embodiment, the ITR or modified ITR of the DNA molecule provided herein is based on the AAV8 ITR. In one embodiment, the ITR or modified ITR of the DNA molecule provided herein is based on the AAV9 ITR.

[0234] In one embodiment, the DNA molecules for the methods and compositions provided herein comprise one or more non-AAV ITRs. In further embodiments, such non-AAV ITRs may be derived from hairpin sequences found in mammalian genomes. In a specific embodiment, such non-AAV ITRs may be derived from hairpin sequences found in mitochondrial genomes, including the OriL hairpin sequence (SEQ ID NO: 30: 5'GAAGAGGGCGGCGGCCCTTTTTTCCGCCCTCTTCGGGGCCGTCCAAACTT'3), which accommodates a stem-loop structure and is involved in the initiation of DNA synthesis of mitochondrial DNA (see Fuste et al., Molecular Cell, 37, 67-78, January 15, 2010, which is incorporated herein by reference in its entirety). In another specific embodiment, the DNA molecules for the methods and compositions provided herein comprise an ITR derived from the OriL sequence that is mirrored to form a T-junction with two self-complementary palindromic regions and a 12-nucleotide loop at either apex of the hairpin. In one embodiment, the DNA molecules for the methods and compositions provided herein comprise ITRs derived from an OriL sequence that maintains the OriL hairpin loop followed by an unpaired bulge and a GC-rich stem. Some exemplary embodiments of ITRs derived from mitochondrial OriL are depicted in FIG.

[0235] In one embodiment, the DNA molecule for the methods and compositions provided herein comprises one or more non-AAV ITRs derived from aptamers. Similar to viral ITRs, aptamers are composed of ssDNA that fold to form a three-dimensional structure and have the ability to recognize biological targets with high affinity and high specificity. DNA aptamers can be generated by Systematic Evolution of Ligands by Exponential Enrichment (SELEX). For example, it has already been shown that some aptamers can target the nuclei of human cells (see Shen et al ACS Sens. 2019, 4, 6, 1612-1618, which is incorporated herein by reference in its entirety). In one embodiment, the DNA molecule for the methods and compositions provided herein comprises a nuclear targeting aptamer ITR or a derivative thereof, where the aptamer specifically binds to a nuclear protein. In some embodiments, the aptamer ITR folds to form a secondary structure that can include hairpins and internal loops, as well as bulges and stem regions, and the like. Some exemplary embodiments of derived aptamers or ITRs are depicted in Figure 5. In a specific embodiment, the aptamer comprises the sequence ATCCGGCTTTAAACGGGCAACTGCGTCTCATTCACGTTAGAGACTACAACCGTCGGAT (SEQ ID NO:346).

[0236] In some specific embodiments, the DNA molecules for the methods and compositions provided herein comprise one or more AAV2 ITRs, human erythrovirus B19 ITRs, goose parvovirus ITRs, and / or derivatives thereof in any combination. In other specific embodiments, the DNA molecules for the methods and compositions provided herein comprise two ITRs selected from AAV2 ITRs, human erythrovirus B19 ITRs, goose parvovirus ITRs, and derivatives thereof in any combination. In some specific embodiments, the DNA molecules for the methods and compositions provided herein comprise one or more AAV2 ITRs, human erythrovirus B19 ITRs, goose parvovirus ITRs, and / or derivatives thereof in any combination, where the ITRs remain functional regardless of whether the palindromic regions of those ITRs are in direct, opposite, or any possible combination of 5' and 3' ITR orientation relative to the expression cassette (as described in WO2019 / 143885, which is incorporated herein by reference in its entirety).

[0237] In some embodiments, the modified IR or ITR in the DNA molecules provided herein is a synthetic IR sequence that contains a restriction site for an endonuclease, such as 5'-GAGTC-3', in addition to various palindromic sequences that allow for hairpin secondary structure formation as described in this section (Section 5.4.1).

[0238] In certain embodiments, the IR or ITR in the DNA molecules provided herein may be an IR or ITR with varying sequence homology with the IR or ITR sequences described in this section (Section 5.4.1). In other embodiments, the IR or ITR in the DNA molecules provided herein may be an IR or ITR with varying sequence homology with known IR or ITR sequences of the origin of the various ITRs described in this section (Section 5.4.1) (e.g., viral ITRs, mitochondrial ITRs, artificial or synthetic ITRs such as aptamers, etc.). In one embodiment, such homology as provided in this paragraph may be at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology. In another embodiment, such homology as provided in this paragraph can be about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% homology.

[0239] In certain embodiments, the right and / or left IR or ITR is a synthetic IR or ITR. In certain embodiments, the synthetic IR or ITR comprises the nucleotide sequence of SEQ ID NO: 417 or 418. In certain embodiments, the synthetic IR or ITR comprises a nucleotide sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 417 or 418.

[0240] In some embodiments, the IRs or ITRs in the DNA molecules provided herein can comprise any one or more of the features described in this section (Section 5.4.1), in various permutations and combinations.

[0241] 5.4.2 Restriction enzymes, nicking endonucleases and their respective restriction sites, programmable nicking enzymes and their target sites Various embodiments of nicking endonucleases, restriction enzymes, and / or their restriction sites as described in Section 5.3.4 are provided for the DNA molecules provided herein. In some embodiments, the first, second, third, and fourth restriction sites for a nicking endonuclease provided for a DNA molecule as described in Section 3 and this section (Section 5.4) can all be target sequences for the same nicking endonuclease. In some embodiments, the first, second, third, and fourth restriction sites for a nicking endonuclease provided for a DNA molecule as described in Section 3 and this section (Section 5.4) can be target sequences for four different nicking endonucleases. In other embodiments, the first, second, third, and fourth restriction sites for the nicking endonucleases are target sequences for two different nicking endonucleases, including all possible combinations for allocating four sites to two different nicking endonuclease target sequences (e.g., the first restriction site for the first nicking endonuclease and the remaining restriction site for the second nicking endonuclease, the first and second restriction sites for the first nicking endonuclease and the remaining restriction site for the second nicking endonuclease, etc.). In certain embodiments, the first, second, third, and fourth restriction sites for the nicking endonucleases are target sequences for three different nicking endonucleases, including all possible combinations for allocating four sites to three different nicking endonuclease target sequences. In some embodiments, the nicking endonuclease and the restriction site for the nicking endonuclease can be any one selected from those described in Section 5.3.4, including Table 2. In further embodiments, each of the first, second, third, and fourth restriction sites for the nicking endonuclease can be a site for any nicking endonuclease selected from those described in Section 5.3.4, including Table 2.

[0242] Tables 7-16 show exemplary modified AAV ITR sequences with two antiparallel recognition sites for the same nicking endonuclease, grouped by nicking endonuclease species. The modified sequences of the ITRs and the corresponding alignments of wild-type AAV1, AAV2, AAV3, AAV4 left, AAV4 right, AAV5, and AAV7 are shown in Figures 12-18. [Table 7] TIFF2025504404000010.tif158165 [Table 8] TIFF2025504404000012.tif151165 [Table 9] TIFF2025504404000014.tif151165 [Table 10] TIFF2025504404000016.tif151165 [Table 11] TIFF2025504404000018.tif151165 [Table 12] TIFF2025504404000020.tif151165 [Table 13] TIFF2025504404000022.tif151165 [Table 14] TIFF2025504404000024.tif145165 [Table 15] TIFF2025504404000026.tif139165

Table 16

Table 17

[0243] The first, second, third and fourth restriction sites for the nicking endonuclease can be arranged in a variety of configurations.In some embodiments the first and second restriction sites for a nicking endonuclease are at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, or at least 200 nucleotides apart.In other embodiments, the first and second restriction sites for the nicking endonuclease are about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, or about 200 or more nucleotides apart.In other embodiments, the nucleotide sequence between the first and second restriction sites for the nicking endonuclease is from about 10 to about 500 nucleotides, e.g., from about 10 to about 250, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62 , about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about about 96, about 97, about 98, about 99, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, or about 200 nucleotides apart.

[0244] Similarly, in certain embodiments, the third and fourth restriction sites for the nicking endonuclease are at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64 , at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, or at least 200 nucleotides apart.In a further embodiment, the third and fourth restriction sites for the nicking endonuclease are about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 105, about 106, about 107, about 108, about 109, about 1 4, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, or about 200 nucleotides apart.

[0245] The present disclosure provides that the overhangs described in Sections 3, 5.2 (including 5.3.3), and 5.4 (including 5.4.1) may be the result of nicking at the first and second restriction sites with a nicking endonuclease and denaturation as described in Sections 3, and 5.2 (including 5.3.3). Thus, in some embodiments, the overhang resulting from nicking at the first and second restriction sites may be the same length (in number of nucleotides) that separates the first and second restriction sites as described in the previous paragraph of this section (Section 5.4.2). Because a nicking endonuclease can cleave DNA inside or outside of a restriction site for the nicking endonuclease, in certain embodiments the overhang resulting from nicking at the first and second restriction sites can be at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 nucleotides longer or shorter than the length separating the first and second restriction sites. In other embodiments, the overhang resulting from nicking at the first and second restriction sites may be about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides longer or shorter than the length separating the first and second restriction sites.

[0246] Similarly, the present disclosure provides that the overhangs described in Sections 3, 5.2 (including 5.3.3), and 5.4 (including 5.4.1) may be the result of nicking at the third and fourth restriction sites by a nicking endonuclease and denaturation as described in Sections 3 and 5.2 (including 5.3.3). Thus, in some embodiments, the overhangs resulting from nicking at the third and fourth restriction sites may be the same length (in number of nucleotides) that separates the third and fourth restriction sites as described in the previous paragraph of this section (Section 5.4.2). Because a nicking endonuclease can cleave DNA inside or outside of the restriction site for the nicking endonuclease, in certain embodiments the overhang resulting from nicking at the third and fourth restriction sites can be at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 nucleotides longer or shorter than the length separating the third and fourth restriction sites. In other embodiments, the overhang resulting from nicking at the third and fourth restriction sites may be about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides longer or shorter than the length separating the third and fourth restriction sites.

[0247] As will be apparent from the description in Sections 3 and 5.5 and this section (Section 5.4), the DNA molecules provided herein include an expression cassette. In some embodiments, the expression cassette is located between a first and a second restriction site for a nicking endonuclease(s) at one end and a third and a fourth restriction site for a nicking endonuclease(s) at the other end. In other embodiments, the expression cassette is located within a dsDNA segment of a DNA molecule produced by performing steps a-d of the method described in Sections 3 and 5.2, including the denaturation step that provides two ssDNA overhangs described in Section 5.3.3. In certain embodiments, the first, second, third, and fourth restriction sites for a nicking endonuclease are positioned such that the length of the dsDNA segment described in this paragraph is at least 0.2 kb, at least 0.3 kb, at least 0.4 kb, at least 0.5 kb, at least 0.6, at least kb, at least 0.7 kb, at least 0.8 kb, at least 0.9 kb, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 2.5 kb, at least 3 kb, at least 3.5 kb, at least 4 kb, at least 4.5 kb, at least 5 kb, at least 5.5 kb, at least 6 kb, at least 6.5 kb, at least 7 kb, at least 7.5 kb, at least 8 kb, at least 8.5 kb, at least 9 kb, at least 9.5 kb, or at least 10 kb. In other embodiments, the first, second, third, and fourth restriction sites for a nicking endonuclease are positioned such that the length of the dsDNA segment described in this paragraph is about 0.2 kb, about 0.3 kb, about 0.4 kb, about 0.5 kb, about 0.6, about kb, about 0.7 kb, about 0.8 kb, about 0.9 kb, about 1 kb, about 1.5 kb, about 2 kb, about 2.5 kb, about 3 kb, about 3.5 kb, about 4 kb, about 4.5 kb, about 5 kb, about 5.5 kb, about 6 kb, about 6.5 kb, about 7 kb, about 7.5 kb, about 8 kb, about 8.5 kb, about 9 kb, about 9.5 kb, or about 10 kb.

[0248] As described in Section 5.3.4, incubation with a nicking endonuclease results in a first nick corresponding to a first restriction site for the nicking endonuclease, a second nick corresponding to a second restriction site for the nicking endonuclease, a third nick corresponding to a third restriction site for the nicking endonuclease, and / or a fourth nick corresponding to a fourth restriction site for the nicking endonuclease. The present disclosure provides that the first, second, third, and / or fourth nick can be at various positions relative to the inverted repeat. In one embodiment, the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides from the 5' nucleotide of the ITR closing base pair of the first inverted repeat. In another embodiment, the first nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides from the 3' nucleotide of the ITR closing base pair of the first inverted repeat. In yet another embodiment, the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides from the 5' nucleotide of the ITR closing base pair of the first inverted repeat.In further embodiments, the second nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides from the 3' nucleotide of the ITR closing base pair of the first inverted repeat. In one embodiment, the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides from the 5' nucleotide of the ITR closing base pair of the second inverted repeat. In another embodiment, the third nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides from the 3' nucleotide of the ITR closing base pair of the second inverted repeat. In yet another embodiment, the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides from the 5' nucleotide of the ITR closing base pair of the second inverted repeat.In further embodiments, the fourth nick is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides from the 3' nucleotide of the ITR closing base pair of the second inverted repeat. In some embodiments, any or any combination of the first, second, third, and fourth nicks are inside the inverted repeat. In certain embodiments, any or any combination of the first, second, third, and fourth nicks are outside the inverted repeat. In some additional embodiments, the first, second, third, and fourth nicks may have any relative positions between themselves, between any of them and the inverted repeat, and / or between any of them and the expression cassette described in this section (Section 5.4.2), in any combination or permutation. In some further embodiments, the first, second, third, and fourth restriction sites for nicking endonucleases may have any relative positions between themselves, between any of them and the inverted repeat, and / or between any of them and the expression cassette described in this section (Section 5.4.2), in any combination or permutation.

[0249] 5.4.3 Expression Cassettes Encoding FVIII The DNA molecules provided herein include at least one expression cassette. An "expression cassette" is a nucleic acid molecule or a portion of a nucleic acid molecule that contains sequences or other information that direct the cellular machinery to make RNA and proteins. In some embodiments, the expression cassette includes a promoter sequence. In certain embodiments, the expression cassette includes a transcription unit. In some further embodiments, the expression cassette includes a promoter operably linked to the transcription unit. In one embodiment, the transcription unit includes an open reading frame (ORF). An embodiment of an ORF used with the methods and compositions provided herein is further described in the final paragraph of this section (Section 5.4.3). The expression cassette can further include features that direct the cellular machinery to produce RNA and proteins. In one embodiment, the expression cassette includes a post-transcriptional regulator. In another embodiment, the expression cassette further includes a polyadenylation and / or termination signal. In yet another embodiment, the expression cassette includes regulatory elements known and used in the art to regulate (enhance, inhibit, and / or turn on / off expression of the ORF). Such regulatory elements include, for example, the 5' untranslated region (UTR), the 3'-UTR, or both the 5'UTR and the 3'UTR. In some further embodiments, the expression cassette comprises any one or more of the features provided in this section (Section 5.4.3), in any combination or permutation.

[0250] An expression cassette can contain a protein coding sequence in its ORF (sense strand). Alternatively, an expression cassette can contain a complementary sequence (antisense strand) of the ORF that codes for a protein, as well as regulatory components and / or other signals to direct the cellular machinery to produce the sense strand DNA / RNA and the corresponding protein. In some embodiments, an expression cassette contains a protein sequence without introns. In other embodiments, an expression cassette contains a protein sequence with introns that are removed upon transcription and splicing. An expression cassette can also contain a variable number of ORFs or transcription units. In one embodiment, an expression cassette contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 ORFs. In another embodiment, the expression cassette comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 transcription units.

[0251] The human F8 gene encodes a 2351 amino acid protein (SEQ ID NO: 359, Accession No. P00451) with a molecular weight of approximately 174.8 kDa. The F8 gene is located on chromosome X. The consensus human F8 coding sequence can be found in NCBI Accession No. NM_000132 and translated to SEQ ID NO: 359.

[0252] Upon expression, factor VIII is secreted into the blood and circulates in an inactive form. The inactive form typically binds to Willebrand factor, which stabilizes it. Following a triggering event, such as injury, factor VIII becomes activated. The activated protein then interacts with coagulation factor IX, which proteolytically activates factor X, triggering the coagulation pathway and leading to clot formation.

[0253] Factor VIII consists of six domains, namely A1-A2-B-A3-C1-C2, and three acidic linker regions a1, a2, and a3 (referred to herein as "linker a1", "linker a2" and "linker a3"). Factor VIII has 19 consensus sites for N-linked glycosylation. FVIII is divided into a heavy chain (A1-a1-A2-a2-B) and a light chain (a3-A3-C1-C2). Factor VIII is expressed and circulates in an inactive form as a heterodimeric complex consisting of the A1-A2-B and A3-C1-C2 domains. Factor VIII is activated by proteolytic cleavage by thrombin. After thrombin cleavage, factor VIII forms a heterotrimeric complex consisting of the A1, A2, and A3-C1-C2 domains and undergoes a conformational change that allows it to bind to factor IXa and activate factor X. Following activation, factor VIIIa can undergo further proteolysis and / or the individual components of the heterotrimeric complex can dissociate from one another, thereby inactivating factor VIIIa. It has been shown that the B domain is not required for factor VIII cofactor activity.

[0254] Those skilled in the art will understand that FVIII therapeutic proteins (also referred to herein as therapeutic FVIII proteins) include all splice variants and orthologues of FVIII proteins. Essentially any aspect of a FVIII therapeutic protein or a fragment thereof (e.g., functional fragment) can be encoded by and expressed in or from a hairpin end DNA vector described herein. FVIII therapeutic proteins include intact molecules and fragments thereof (e.g., functional). In some embodiments, a FVIII therapeutic protein or a fragment thereof is modified compared to wild-type FVIII. Examples of modified FVIII therapeutic proteins include, but are not limited to, those explicitly described herein.

[0255] The term "variant factor VIII (FVIII)" refers to a genetically modified modified FVIII compared to unmodified wild-type FVIII (e.g., SEQ ID NO: 359) or truncated FVIII. Such variants can be referred to as "nucleic acid variants encoding factor VIII (FVIII)". A specific example of a variant is a CpG-reduced nucleic acid encoding FVIII or a functional fragment thereof. The term "variant" need not appear in each instance of reference to a CpG-reduced nucleic acid encoding FVIII. Similarly, terms such as "CpG-reduced nucleic acid" may omit the term "variant", but reference to "CpG-reduced nucleic acid" is intended to include variants at the genetic level.

[0256] FVIII constructs with reduced CpG content can exhibit improvements compared to wild-type FVIII or functional fragments thereof that do not have reduced CpG content, and these improvements can be observed even in the absence of modifications to the nucleic acid that would result in a change in the primary amino acid sequence of the encoded FVIII protein.

[0257] A "functional fragment(s)" of FVIII provided herein includes modified FVIII fragments in which the modified protein has amino acid changes compared to wild-type FVIII but retains some functionality of the native full-length protein, increases protein activity compared to the native full-length protein, and / or improves protein functionality compared to the native full-length protein. For example, in certain embodiments, the CpG-reduced nucleic acid encoding FVIII or a truncated FVIII protein comprises a B domain deletion as described herein, and the expressed protein retains clotting function. In another embodiment, the CpG-reduced nucleic acid encoding FVIII or a truncated FVIII protein comprises a B domain and / or linker a3 deletion as described herein, and the expressed protein retains clotting function. In certain embodiments, the variant truncated FVIII may retain a portion of the B domain. Thus, in certain embodiments, the truncated FVIII comprises a portion of the B domain.

[0258] In some embodiments, the hairpin DNA molecule for expressing FVIII protein provides advantages over traditional AAV vectors, because there is no size restriction for the heterologous nucleic acid sequence encoding desired protein.Therefore, even full-length FVIII protein can be expressed from a single DNA molecule.Therefore, the DNA molecule described herein can be used to express therapeutic FVIII protein in subjects who need it, such as subjects with hemophilia A. [Table 18] TIFF2025504404000045.tif214165TIFF2025504404000046.tif210165TIFF2025504404000047.tif215165TIFF2025504404000048.tif210165TIFF2025504404000049.tif214165TIFF2025504404000050.tif209165TIFF2025504404000051.tif206165TIFF2025504404000052.tif209165TIFF2025504404000053.tif209165TIFF2025504404000054.tif205165TIFF2025504404000055.tif206165TIFF2025504404000056.tif209165TIFF2025504404000057.tif209165TIFF2025504404000058.tif209165TIFF2025504404000059.tif206165TIFF2025504404000060.tif209165TIFF2025504404000061.tif211165TIFF2025504404000062.tif209165TIFF2025504404000063.tif211165TIFF2025504404000064.tif214165TIFF2025504404000065.tif214165TIFF2025504404000066.tif206165TIFF2025504404000067.tif205165TIFF2025504404000068.tif197165TIFF2025504404000069.tif208165TIFF2025504404000070.tif209165TIFF2025504404000071.tif121165

[0259] In one aspect, a codon-optimized engineered nucleic acid sequence is provided that encodes human FVIII. In a particular embodiment, an engineered human FVIII cDNA is provided herein (as SEQ ID NO: 175), which is designed to remove selected nicking enzyme recognition sites compared to the native FVIII sequence (SEQ ID NO: 174). Preferably, the codon-engineered FVIII coding sequence has less than about 80% identity to the full-length native FVIII coding sequence, or less. In one embodiment, the codon-optimized FVIII coding sequence has about 75% identity to the native FVIII coding sequence of SEQ ID NO: 174. In one embodiment, the engineered FVIII coding sequence is characterized by an improved translation rate compared to the native FVIII after delivery. In one embodiment, the engineered FVIII coding sequence shares less than about 99%, less than 98%, less than 97%, less than 96%, less than 95%, less than 94%, less than 93%, less than 92%, less than 91%, less than 90%, less than 89%, less than 88%, less than 87%, less than 86%, less than 85%, less than 84%, less than 83%, less than 82%, less than 81%, less than 80%, less than 79%, less than 78%, less than 77%, less than 76%, less than 75%, less than 74%, less than 73%, less than 72%, less than 71%, less than 70%, less than 69%, less than 68%, less than 67%, less than 66%, less than 65%, less than 64%, less than 63%, less than 62%, less than 61%, or less than identity to the full-length native FVIII coding sequence of SEQ ID NO:174. In one embodiment, the codon-optimized nucleic acid sequence is a variant of SEQ ID NO: 175. In another embodiment, the codon-optimized nucleic acid sequence is a sequence that shares about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61% or more identity with SEQ ID NO: 175. In one embodiment, the codon-optimized nucleic acid sequence is SEQ ID NO: 175. In another embodiment, the nucleic acid sequence is codon-optimized for expression in humans.In other embodiments, a different FVIII coding sequence is selected.

[0260] Another example of nucleic acid modification provided herein is CpG reduction.In certain embodiments, the CpG-reduced nucleic acid encoding FVIII, such as human FVIII protein, or its functional fragment has 10 or less, 5 or less, or 5 or less CpG compared to the wild-type sequence encoding human FVIII factor.

[0261] A further example of nucleic acid modification provided herein is the reduction of a restriction site for a selected nicking endonuclease. In certain embodiments, a restriction site-reduced nucleic acid encoding FVIII, such as a human FVIII protein, or a functional fragment thereof, has five or fewer restriction sites for a nicking endonuclease compared to a wild-type sequence encoding human FVIII factor. In certain embodiments, a restriction site-reduced nucleic acid encoding FVIII, or a functional fragment thereof, has no restriction sites for a nicking endonuclease compared to a wild-type sequence encoding human FVIII factor. In some embodiments, the expression cassette comprises a nucleic acid modification to reduce CpG and restriction sites for a selected nicking endonuclease.

[0262] In one aspect, a CpG-minimized engineered nucleic acid sequence is provided that encodes human FVIII that translates to the F328S amino acid substitution. In a particular embodiment, an engineered human FVIII cDNA that translates to the F328S mutation is provided herein (as SEQ ID NO: 176), which is also designed to remove selected nicking enzyme recognition sites and minimize CpG motifs compared to the native FVIII sequence. Preferably, the CpG-minimized FVIII coding sequence has less than about 90%, preferably about 85% or less identity to the full-length native FVIII coding sequence. In one embodiment, the CpG-minimized FVIII coding sequence has about 81% identity to the native FVIII coding sequence of SEQ ID NO: 174. In one embodiment, the CpG-minimized FVIII coding sequence is characterized by reduced activation to host immune response compared to the native FVIII sequence after delivery to a host cell. In one embodiment, the CpG-minimized FVIII coding sequence shares less than about 99%, less than 98%, less than 97%, less than 96%, less than 95%, less than 94%, less than 93%, less than 92%, less than 91%, less than 90%, less than 89%, less than 88%, less than 87%, less than 86%, less than 85%, less than 84%, less than 83%, less than 82%, less than 81%, less than 80%, less than 79%, less than 78%, less than 77%, less than 76%, less than 75%, less than 74%, less than 73%, less than 72%, less than 71%, less than 70%, less than 69%, less than 68%, less than 67%, less than 66%, less than 65%, less than 64%, less than 63%, less than 62%, less than 61%, or less than identity to the full-length native FVIII coding sequence of SEQ ID NO:174. In one embodiment, the CpG-minimized nucleic acid sequence is a variant of SEQ ID NO: 176. In another embodiment, the CpG-minimized nucleic acid sequence is a sequence that shares about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61% or greater identity to SEQ ID NO:176.In one embodiment, the CpG-minimized nucleic acid sequence is SEQ ID NO:176.

[0263] In some embodiments, the hairpin-end DNA molecule described herein encodes a fusion protein that includes a full length, fragment, or portion of a FVIII protein fused to another sequence (e.g., N- or C-terminal fusion). In some embodiments, the N- or C-terminal sequence is a signal sequence or a cell targeting sequence. In further embodiments, the hairpin-end DNA molecule encodes a FVIII protein with a complete or partial B domain deletion, which may include the remaining portion of the B domain linked to the A2 and A3 domains as in the wild-type amino acid sequence, or may include one or more spacer amino acids to replace part or all of the deleted domain. In further embodiments, the hairpin-end DNA molecule encodes a FVIII protein with a complete or partial B domain deletion and a partial linker a3 deletion, which may include the remaining portion of the B domain linked to the A2 domain and a partial linker a3 as in the wild-type amino acid sequence, or may include one or more spacer amino acids to replace part or all of the deletion in the domain or linker, as described herein.

[0264] In certain embodiments, the truncated linker a3 has a deletion of the N-terminal portion of the linker a3 domain, and the number of amino acids deleted from the N-terminal portion of the linker a3 domain ranges from about 1 to about 10, about 1 to about 20, about 1 to about 30, or about 1 to about 40 amino acids. In specific embodiments, the truncated linker a3 has a deletion of 9 N-terminal amino acids of the linker a3. In one embodiment, the truncated FVIII variant comprising a truncated N-terminus of the linker a3 is characterized by a higher procoagulant activity compared to native FVIII after expression by a host cell.

[0265] In some embodiments, the truncated B domain is an N-terminal portion of the B domain, and the number of amino acids in the truncated N-terminal portion of the B domain ranges from about 20 to about 300, about 25 to about 300, about 29 to about 300, about 29 to about 269, about 29 to about 250, about 30 to about 300, about 30 to about 250, about 50 to about 300, about 50 to about 250, about 100 to about 300, about 100 to about 500, about 100 to about 250, about 150 to about 300, about 150 to about 250, about 200 to about 300, or about 250 to about 300 amino acids. In some embodiments, the hairpin end DNA molecule comprises a truncated version of a peptide engineered to replace the wild-type factor VIII B domain or the B domain of a factor VIII polypeptide. As used herein, the factor VIII linker, according to some embodiments, is located between the C-terminus of the factor VIII heavy chain and the N-terminus of the factor VIII light chain in the factor VIII variant polypeptide. Non-limiting examples of B domain replacement linkers are disclosed in US Patent Application Publication Nos. 2013 / 024960, 2015 / 0071883 and 2015 / 0158930, and PCT Publication Nos. WO2014 / 064277 and WO2014 / 127215, the disclosures of which are incorporated herein by reference in their entirety for all purposes.

[0266] In a specific embodiment, a hairpin-ended DNA molecule is provided that includes a nucleic acid sequence encoding a truncated human FVIII protein comprising a partial B domain and a truncated linker a3 (B-NA3), wherein the truncated protein encoded by the DNA molecule is characterized by a deletion of amino acids 985-1677 compared to the native FVIII sequence. In one embodiment, the B-NA3 truncated FVIII protein coding sequence has a deletion of amino acids 985-1677 relative to the full-length native FVIII coding sequence of SEQ ID NO:174, which is less than about 99%, less than about 98%, less than about 97%, less than about 96%, less than about 95%, less than about 94%, less than about 93%, less than about 92%, less than about 91%, less than about 90%, less than about 89%, less than about 88%, less than about 87%, less than about 86%, less than about 85%, less than about 84%, or less than about 98%. , less than about 83%, less than about 82%, less than about 81%, less than about 80%, less than about 79%, less than about 78%, less than about 77%, less than about 76%, less than about 75%, less than about 74%, less than about 73%, less than about 72%, less than about 71%, less than about 70%, less than about 69%, less than about 68%, less than about 67%, less than about 66%, less than about 65%, less than about 64%, less than about 63%, less than about 62%, or less than about 61% identity. In one embodiment, the B-NA3 truncated FVIII protein sequence is a variant of SEQ ID NO:358. In another embodiment, the B-NA3 truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity to SEQ ID NO:358. In one embodiment, the B-NA3 truncated FVIII protein sequence is SEQ ID NO:358.

[0267] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 354. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:354. In another embodiment, the truncated FVIII encoding protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 354. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO: 354.

[0268] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, wherein the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 355. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:355. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 355. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:355.

[0269] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 356. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:356. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 356. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:356.

[0270] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 357. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:357. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 357. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:357.

[0271] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 360. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:360. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 360. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:360.

[0272] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 378. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:378. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 378. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:378.

[0273] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 382. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:382. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 382. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:382.

[0274] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 386. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:386. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 386. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:386.

[0275] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 390. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:390. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 390. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO: 390.

[0276] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 394. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:394. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 394. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:394.

[0277] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 398. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:398. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 398. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:398.

[0278] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 402. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:402. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 402. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:402.

[0279] In a specific embodiment, the hairpin end DNA molecule comprises a nucleic acid sequence encoding a truncated human FVIII protein, and the truncated protein encoded by the DNA molecule has the amino acid sequence of SEQ ID NO: 406. In one embodiment, the truncated FVIII protein sequence is a variant of SEQ ID NO:406. In another embodiment, the truncated FVIII protein sequence has a sequence that shares about 99%, about 98%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 89%, about 88%, about 87%, about 86%, about 85%, about 84%, about 83%, about 82%, about 81%, about 80%, about 79%, about 78%, about 77%, about 76%, about 75%, about 74%, about 73%, about 72%, about 71%, about 70%, about 69%, about 68%, about 67%, about 66%, about 65%, about 64%, about 63%, about 62%, about 61%, or greater than 61% identity with SEQ ID NO: 406. In one embodiment, the truncated FVIII protein sequence is SEQ ID NO:406.

[0280] In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 174. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 175. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 176. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 177. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 178. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 179. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 180. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 379. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 380. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 381.In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 383. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 384. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 385. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 387. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 388. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 389. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 391. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 392. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 393. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 395.In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 396. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 397. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 399. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 400. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 401. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 403. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 404. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 405. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 407. In a specific embodiment, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO: 408.In specific embodiments, the expression cassette comprises a FVIII transgene that is at least 60%, at least 70%, at least 80%, or at least 90% identical to the sequence set forth in SEQ ID NO:409.

[0281] In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 174. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 175. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 176. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 177. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 178. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 179. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 180. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 379. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 380. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 381. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 383. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 384. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 385. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 387. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 388. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 389. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 391. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 392.In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 393. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 395. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 396. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 397. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 399. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 400. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 401. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 403. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 404. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 405. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 407. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 408. In a specific embodiment, the expression cassette comprises a FVIII transgene identical to the sequence set forth in SEQ ID NO: 409.

[0282] The terms "percent identity (%)", "sequence identity", "percent sequence identity", or "percent identical" in the context of FVIII-encoding nucleic acid sequences refer to residues in two sequences that are identical when aligned for correspondence. The length of sequence identity comparison may be desired over the entire length of a genome, over the entire length of a gene coding sequence, or over a fragment of at least about 500-5000 nucleotides. However, identity over smaller fragments, e.g., at least about 9 nucleotides, usually at least about 20-24 nucleotides, at least about 28-32 nucleotides, at least about 36 or more nucleotides, may also be desired.

[0283] Percent identity can be readily determined for amino acid sequences spanning the entire length of a protein, polypeptide, about 32 amino acids, about 330 amino acids, or peptide fragments thereof or the corresponding nucleic acid sequence coding sequence. Suitable amino acid fragments may be at least about 8 amino acids in length and may be up to about 700 amino acids in length. In general, when referring to "identity", "homology", or "similarity" between two different sequences, the "identity", "homology", or "similarity" is determined with reference to "aligned" sequences. An "aligned" sequence or "alignment" refers to multiple nucleic acid or protein (amino acid) sequences, often including corrections of deleted or additional bases or amino acids compared to a reference sequence.

[0284] Identity can be determined by preparing an alignment of sequences and by using various algorithms and / or computer programs known in the art or commercially available [e.g., BLAST, ExPASy, ClustalO, FASTA, e.g., using the Needleman-Wunsch algorithm, Smith-Waterman algorithm]. Alignment is performed using any of a variety of publicly or commercially available multiple sequence alignment programs. Sequence alignment programs are available for amino acid sequences, e.g., the "Clustal Omega" and "Clustal X" programs. Generally, any of these programs are used with default settings, but those skilled in the art can change these settings as needed. Alternatively, those skilled in the art can utilize another algorithm or computer program that provides at least the level of identity or alignment provided by the referenced algorithms and programs. See, e.g., JD Thomson et al, Nucl. Acids. Res., "A comprehensive comparison of multiple sequence alignments", 27(13):2682-2690 (1999). For nucleic acid sequences, several sequence alignment programs are also available, examples of such programs include "Clustal Omega", "Clustal W", "CAP Sequence Assembly", "BLAST", "MAP", and "MEME", which are accessible via web servers on the Internet.

[0285] Codon-optimized coding regions can be designed by a variety of different methods. This optimization may be performed using methods available online (e.g., GeneArt), publicly available methods, or companies that provide codon optimization services, such as DNA2.0 (Menlo Park, CA). Suitably, the entire length of the open reading frame (ORF) for the product is modified. However, in some embodiments, only a fragment of the ORF may be altered. By using one of these methods, frequencies can be applied to any given polypeptide sequence to generate a nucleic acid fragment of a codon-optimized coding region that encodes a polypeptide. Several options are available for making the actual changes to the codons or for synthesizing the codon-optimized coding region designed as described herein. Such modifications or synthesis can be performed using standard and routine molecular biology operations well known to those skilled in the art.

[0286] The FVIII expression cassette may be located at any suitable distance in base pairs from either the 5' and / or 3' ITR closing pair to enable or maintain efficient transcription of said expression cassette in the host cell (as described in section 5.4.1). In some embodiments, the distance between the expression cassette and the 5' ITR and the distance between the expression cassette and the 3' ITR closing pair are identical. In some embodiments, the distance between the expression cassette and the 5' ITR and the distance between the expression cassette and the 3' ITR closing pair are not identical.In some embodiments the distance between the expression cassette and / or the 3' ITR closing pair, and the distance between the expression cassette and the 5' ITR closing pair is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least at least 195, at least 200, at least 205, at least 210, at least 215, at least 220, at least 225, at least 230, at least 235, at least 240, at least 245, at least 250, at least 255, at least 260, at least 265, at least 270, at least 275, at least 280, at least 285, at least 290, at least 295, at least 300, at least 305, at least 310, at least 315, at least 320, at least 325, at least 330, at least 335, at least 340, at least 345, at least 350, at least 355, at least 360, at least 365, at least 370, at least 375, at least 380, at least 385, at least 390, at least 395, or at least 400 nucleotides.In some embodiments, the distance between the expression cassette and the 3' ITR closing pair and / or the distance between the expression cassette and the 5' ITR closing pair is about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, about 200, about 205, about 210, about 215, about 220, about 225, about 230, about 235, about 240, about 245, about 250, about 255, about 260, about 265, about 270, about 275, about 280, about 285, about 290, about 295, about 300, about 305, about 310, about 315, about 320, about 325, about 330, about 335, about 340, about 345, about 350, about 355, about 360, about 365, about 370, about 375, about 380, about 385, about 390, about 395, or about 400 nucleotides.

[0287] By "engineered nucleic acid sequence" is meant that the nucleic acid sequence encoding the FVIII protein described herein is assembled and placed in any suitable genetic element, e.g., naked DNA, phage, transposon, cosmid, episome, etc., which transfers the FVIII sequence carried thereon to a host cell, e.g., to generate a non-viral delivery system (e.g., RNA-based system, naked DNA, etc.), or to generate a viral vector in a packaging host cell, and / or for delivery to a host cell in a subject. In one embodiment, the genetic element is a circular plasmid. Methods used to generate such engineered constructs are known to those skilled in nucleic acid manipulation and include genetic engineering, recombinant engineering, and synthetic techniques. See, e.g., Green and Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, NY (2012).

[0288] In one embodiment, the nucleic acid sequence encoding FVIII further comprises a nucleic acid encoding a tag polypeptide covalently linked thereto. The tag polypeptide may be selected from known "epitope tags" including, but not limited to, myc tag polypeptide, glutathione-S-transferase tag polypeptide, luciferase protein tag polypeptide, green fluorescent protein tag polypeptide, myc-pyruvate kinase tag polypeptide, His6 tag polypeptide, influenza virus hemagglutinin tag polypeptide, flag tag polypeptide, and maltose binding protein tag polypeptide. In some aspects, hairpin end vectors expressing FVIII protein linked to a reporter polypeptide may be used for diagnostic purposes and to determine efficacy in subjects to whom they are administered, or as a marker of activity of the hairpin end vector.

[0289] In yet another embodiment, the expression cassette comprises regulatory elements known and used in the art to regulate (enhance, inhibit, and / or turn on / off expression of an ORF). Such regulatory elements include, for example, the 5' untranslated region (UTR), the 3'-UTR, or both the 5'UTR and the 3'UTR. In some further embodiments, the expression cassette comprises any one or more features provided in this Section 5.4.3 in any combination or permutation.

[0290] An untranslated region is a nucleic acid segment of a polynucleotide that is not translated before the start codon (5'UTR) and after the stop codon (3'UTR). In some embodiments, a polynucleotide of the invention comprising an open reading frame (ORF) encoding a Factor VIII polypeptide further comprises a UTR (e.g., a 5'UTR or a functional fragment thereof, a 3'UTR or a functional fragment thereof, or a combination thereof). Further exemplary translational regulatory activities provided by components, structures, elements, motifs, and / or specific sequences comprising a polynucleotide include, but are not limited to, mRNA stabilization or destabilization (Baker & Parker (2004) Curr Opin Cell Biol 16(3):293-299), translational activation (Villalba et al., (2011) Curr Opin Genet Dev 21(4):452-457), and translational repression (Blumer et al., (2002) Mech Dev 110(1-2):97-112). Studies have shown that naturally occurring cis-acting RNA elements can confer their respective functions when used to modify by incorporation into heterologous polynucleotides (Goldberg-Cohen et al., (2002) J Biol Chem 277(16):13635-13640). The UTRs may be homologous or heterologous to the coding region within the polynucleotide. In some embodiments, the UTRs are homologous to the ORF encoding the Factor VIII polypeptide. In some embodiments, the UTRs are heterologous to the ORF encoding the Factor VIII polypeptide. In some embodiments, the polynucleotide comprises two or more 5'UTRs or functional fragments thereof, each having the same or different nucleotide sequence. In some embodiments, the polynucleotide comprises two or more 3'UTRs or functional fragments thereof, each having the same or different nucleotide sequence. In some embodiments, the 5'UTR and 3'UTR may be heterologous. In some embodiments, the 5'UTR may be from a different species than the 3'UTR. In some embodiments, the 3'UTR may be from a different species than the 5'UTR.The UTRs or portions thereof can be positioned in the same orientation as in the transcript from which they were selected, or the orientation or position can be changed. Thus, the 5' and / or 3' UTRs can be inverted, shortened, extended, or combined with one or more other 5' or 3' UTRs.

[0291] The 5'UTR or 3'UTR of the hairpin-ended DNA molecule of the present disclosure may, in some embodiments, contain a sequence that destabilizes the mRNA transcript obtained by cellular processes. In certain embodiments, the mRNA transcript of the hairpin-ended DNA molecule may be destabilized in response to increased endogenous endoplasmic reticulum (ER) stress by cleavage of the mRNA (e.g., by riboendonuclease). In specific embodiments, ER stress activates a riboendonuclease, which then cleaves the mRNA transcript at a specific site or motif. Since ectopic expression of a polypeptide (e.g., FVIII) can induce ER stress, it may be advantageous to engineer in a sequence that contains an ER stress-dependent riboendonuclease recognition motif. In some embodiments, the recognition motif is folded into a loop and a stem. In a specific embodiment, the stem comprises four nucleotides that form an energetically stable stem, as well as a loop consisting of an approximately 7 nt consensus transcribed RNA sequence of CNGCAGN (whereby N can be A, G, T, or C, SEQ ID NO: 333).

[0292] In certain embodiments, mouse FVIII, hemoglobin alpha locus (e.g., hemoglobin subunit alpha 1), hemoglobin subunit alpha 2, or hemoglobin beta locus (e.g., hemoglobin subunit beta) 5'UTR and / or 3'UTR sequences can be used with the methods and compositions of the invention. In certain embodiments, rabbit FVIII, hemoglobin alpha locus (e.g., hemoglobin subunit alpha 1), hemoglobin subunit alpha 2, or hemoglobin beta locus (e.g., hemoglobin subunit beta) 5'UTR and / or 3'UTR sequences can be used with the methods and compositions of the invention. In certain embodiments, human FVIII, hemoglobin alpha locus (e.g., hemoglobin subunit alpha 1), hemoglobin subunit alpha 2, or hemoglobin beta locus (e.g., hemoglobin subunit beta) 5'UTR and / or 3'UTR sequences can be used with the methods and compositions of the invention. In certain embodiments, the 5'UTR and / or 3'UTR sequences of the present disclosure comprise a nucleotide sequence that is at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 333 adjacent to the 4 nt stem. In certain embodiments, an expression cassette open reading frame may comprise a nucleotide sequence that is at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 333 adjacent to the 4 nt stem.

[0293] An expression cassette can contain a protein coding sequence in its ORF (sense strand). Alternatively, an expression cassette can contain a complementary sequence (antisense strand) of the ORF that codes for a protein, as well as regulatory components and / or other signals to direct the cellular machinery to produce the sense strand DNA / RNA and the corresponding protein. In some embodiments, the expression cassette contains a FVIII protein sequence without an intron. In other embodiments, the expression cassette contains a FVIII protein sequence with an intron that is removed upon transcription and splicing. An expression cassette can also contain a variable number of ORFs or transcription units. In one embodiment, the expression cassette contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 ORFs. In another embodiment, the expression cassette comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 transcription units.

[0294] An expression cassette can also include one or more transcriptional regulators, one or more post-transcriptional regulators, or both one or more transcriptional regulators and one or more post-transcriptional regulators. Such regulatory elements are any sequences that allow for, contribute to, or regulate the functional regulation of a nucleic acid molecule, including replication, duplication, transcription, splicing, translation, stability, and / or transport of a nucleic acid or one of its derivatives (e.g., mRNA) into a host cell or organism. Such regulatory elements include, but are not limited to, promoters, enhancers, polyadenylation signals, translation stop codons, ribosome binding elements, transcription terminators, selection markers, origins of replication, and the like.

[0295] In some embodiments, the expression cassette comprises an enhancer. Any enhancer sequence known to those skilled in the art in view of the present disclosure can be used. In some embodiments, the enhancer is a liver-specific enhancer. In some embodiments, the enhancer can be an enhancer sequence such as human actin gene, human myosin gene, human hemoglobin gene, human muscle creatine gene, transthyretin gene, alpha-1-antitrypsin gene, albumin gene, apolipoprotein E gene, FVIII gene, or a viral enhancer, such as one from CMV, HA, RSV, or EBV. In a specific embodiment, the enhancer can be the woodchuck HBV post-transcriptional regulatory element (WPRE), an intron / exon sequence derived from human apolipoprotein A1 precursor (ApoAI), the untranslated R-U5 domain of the long terminal repeat (LTR) of human T-cell leukemia virus type 1 (HTLV-1), a splicing enhancer, a synthetic rabbit β-globin intron, the P5 promoter of AAV, or any combination thereof. In certain embodiments, the enhancer comprises a nucleotide sequence that is at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 416. In certain embodiments, the enhancer comprises the nucleotide sequence of SEQ ID NO:416.

[0296] As described above, the expression cassette can include a promoter that controls the expression of a protein of interest. A promoter includes any nucleotide sequence that initiates transcription of an operably linked nucleotide sequence. A promoter can be constitutive, inducible, or repressible. A promoter can be derived from sources including viral, bacterial, fungal, plant, insect, and animal. A promoter can be a homologous promoter (e.g., derived from the same genetic source) or a heterologous promoter (e.g., derived from a different genetic source). In some embodiments, a promoter can be a promoter from simian virus 40 (SV40), a mouse mammary tumor virus (MMTV) promoter, a human immunodeficiency virus (HIV) promoter, such as a bovine immunodeficiency virus (BIV) long terminal repeat (LTR) promoter, a Moloney virus promoter, an avian leukosis virus (ALV) promoter, a cytomegalovirus (CMV) promoter, such as a CMV immediate early promoter (CMV-IE), an Epstein-Barr virus (EBV) promoter, or a Rous sarcoma virus (RSV) promoter. In other embodiments, the promoter may be a promoter from a human gene, such as human actin, human myosin, human hemoglobin, human muscle creatine, or human metallothionein. In further embodiments, the promoter may be a tissue-specific promoter, such as a muscle- or skin-specific promoter, natural or synthetic, to promote expression in cells or tissues in which expression of FVIII is desired, for example, in FVIII-deficient patients. In certain embodiments, the inducible promoter is a liver-specific promoter. In certain embodiments, the promoter comprises a nucleotide sequence at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 340, 341, 342, 347, 410, 411, 412, 413, 414, or 415.In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 340. In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 341. In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 342. In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 347. In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 410. In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 411. In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 412. In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 413. In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 414. In certain embodiments, the promoter comprises the nucleotide sequence of SEQ ID NO: 415. Exemplary regulatory elements can be found in Table 19.

[0297] In certain embodiments, the promoter is a muscle-specific promoter. Non-limiting examples of muscle-specific promoters include muscle creatine kinase (MCK) promoter. Non-limiting examples of suitable muscle creatine kinase promoters are human muscle creatine kinase promoter and truncated mouse muscle creatine kinase [(tMCK) promoter] (Wang B et al, Construction and analysis of compact muscle-selective promoters for AAV vectors. Gene Ther. 2008 Nov; 15 (22): 1489-99) (representative GenBank accession number AF188002). Human muscle creatine kinase has gene 1D number 1158 (representative GenBank accession number NC 000019.9). Other examples of muscle-specific promoters include the synthetic promoter C5.12 (spC5.12, or referred to herein as "C5.12"), such as the spC5.12 or spC5.12 promoter (disclosed in Wang et al., Gene Therapy volume 15, pages 1489-1499 (2008)), the MHCK7 promoter (Salva et al. Mol Ther. 2007 Feb;15(2):320-9), myosin light chain (MLC) promoters, such as MLC2 (Gene 1D No. 4633, representative GenBank Accession No. NG 007554.1), myosin heavy chain (MHC) promoters, such as alpha-MHC (Gene 1D No. 4624, representative GenBank Accession No. NG 023444.1), the desmin promoter (Gene 1D No. 1674, representative GenBank Accession No. NG 023444.1), the MHCK7 promoter (Salva et al. Mol Ther. 2007 Feb;15(2):320-9), the myosin light chain (MLC) promoters, such as MLC2 (Gene 1D No. 4633, representative GenBank Accession No. NG 007554.1), the myosin heavy chain (MHC) promoters, such as alpha-MHC (Gene 1D No. 4624, representative GenBank Accession No. NG 023444.1), the desmin promoter (Gene 1D No. 1674, representative GenBank Accession No. NG 023444.1), the MHCK7 promoter (Salva et al. Mol Ther. 2007 Feb;15(2):320-9), the MHCK7 promoter (Salva et al. Mol Ther. 2007 Feb;15( 008043.1), cardiac troponin C promoter (Gene ID no. 7134, representative GenBank accession no. NG 008963.1), troponin I promoter (Gene ID nos. 7135, 7136, and 7137, representative GenBank accession nos. NG 016649.1, NG 011621.1, and NG_007866.2,), myoD gene family promoter (Weintraub et al., Science, 251, 761 (1991), gene ID number 4654, representative GenBank accession number NM 002478), alpha actin promoter (gene ID numbers 58, 59, and 70, representative GenBank accession numbers NG 006672.1, NG 011541.1, and NG 007553.1), beta actin promoter (gene ID number 60, representative GenBank accession number NG 007992.1), gamma actin promoter (gene ID numbers 71 and 72, representative GenBank accession numbers NG 011433.1 and NM 001199893), muscle-specific promoter remaining within intron 1 of eye morphology of Pitx3 (gene ID number 5309) (Coulon et al., representative GenBank accession number NG 002478). 008147), and promoters described in U.S. Patent Publication No. 2003 / 0157064, and the CK6 promoter (Wang et al 2008 doi:10.1038 / gt.2008.104). In another specific embodiment, the muscle-specific promoter is the E-Syn promoter described in Wang et al., Gene Therapy volume 15, pages 1489-1499 (2008), which includes a combination of the MCK-derived enhancer and the spC5.12 promoter. In certain embodiments of the present disclosure, the muscle-specific promoter is selected from the group consisting of spC5.12 promoter, MHCK7 promoter, E-syn promoter, muscle creatine kinase myosin light chain (MLC) promoter, myosin heavy chain (MHC) promoter, cardiac troponin C promoter, troponin I promoter, myoD gene family promoter, alpha actin promoter, beta actin promoter, gamma actin promoter, the muscle-specific promoter present within intron 1 of the eye form of Pitx3, CK6 promoter, CK8 promoter, and Actal promoter. In certain embodiments, the muscle-specific promoter is spC5.In a further embodiment, the muscle-specific promoter is selected from the group consisting of spC5.12, desmin an...

Claims

[Claim 1] The novel products, methods and processes substantially as herein described.