Nucleic acid molecules and their use for non-viral gene therapy

By designing nucleic acid molecules containing specific ITRs and heterologous polynucleotide sequences, the problems of limited viral packaging capabilities and immune response in gene therapy are solved, and durable and safe gene expression in vivo are achieved.

CN119955796APending Publication Date: 2025-05-09BIOVILA DIVI THERAPEUTICS INC
View PDF 134 Cites 0 Cited by

Patent Information

Application Number
CN202411809059.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-08-09
Filing Date
2019-08-09
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing AAV vectors have limited viral packaging capabilities and problems in gene therapy, affecting their lasting expression and safety in vivo.

Method used

A nucleic acid molecule is designed, comprising a first reverse terminal repeat (ITR) and a second ITR flanked by a gene cassette containing a heterologous polynucleotide sequence, the first ITR and/or the second ITR at least 75% matched to a particular nucleotide sequence to optimize gene expression and avoid immune responses.

Benefits of technology

Through this nucleic acid molecule, effective and persistent expression of target sequences is achieved in vitro and in vivo environments, reducing the limitations of immune response and virus packaging capabilities, and improving the safety and efficiency of gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119955796A_ABST
    Figure CN119955796A_ABST
Patent Text Reader

Abstract

The invention relates to a nucleic acid molecule and its use for non-viral gene therapy, and specifically provides a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette encoding a target sequence. In some embodiments, the target sequence encodes a miRNA and / or a therapeutic protein. In certain embodiments, the therapeutic proteins comprise blood coagulation factors, growth factors, hormones, cytokines, antibodies, fragments thereof, and combinations thereof. In some embodiments, the first ITR and / or the second ITR are / is an ITR of a non-adeno-associated virus (AAV). The disclosure also provides a method of treating a liver metabolic disorder in a subject comprising administering to the subject a nucleic acid molecule or a polypeptide encoded thereby.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention application is a divisional application based on the invention patent application with application date of August 9, 2019, application number 201980065714.1 (international application number PCT / US2019 / 045957) and name “Nucleic Acid Molecules and Their Use for Non-Viral Gene Therapy”.

[0002] Related applications

[0003] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 716,826, filed on August 9, 2018, the entire disclosure of which is incorporated herein by reference.

[0004] Reference to electronically submitted sequence listing

[0005] The contents of the Sequence Listing submitted electronically as an ASCII text file (Name: SA9-465PC_SL_ST25.txt; Size: 460,648 bytes; Creation Date: August 8, 2019) are incorporated herein by reference in their entirety. Background Art

[0006] Gene therapy provides the possibility of a lasting means of treating a variety of diseases. In the past, many gene therapy treatments usually relied on the use of viruses. There are a variety of viral agents that can be selected for this purpose, each with different characteristics, which makes them more suitable or less suitable for gene therapy. Zhou et al., Adv Drug Deliv Rev.106(Pt A):3-26,2016. However, the undesirable properties of some viral vectors (including their immunogenicity spectrum or their tendency to cause cancer) have led to clinical safety issues and, until recently, limited their clinical use in certain applications (e.g., vaccines and oncolytic strategies). Cotter et al., Front Biosci.10:1098-105(2005).

[0007] Adeno-associated virus (AAV) is one of the most commonly studied gene therapy vectors. AAV is a protein coat that surrounds and protects a small single-stranded DNA genome of approximately 4.8 kilobases (kb). Naso et al., BioDrugs, 31(4):317-334, 2017. AAV belongs to the parvovirus family and relies on co-infection with other viruses (primarily adenoviruses) to replicate. Same as above. Its single-stranded genome contains three genes: Rep (replication), Cap (capsid), and aap (assembly). Same as above. These coding sequences are flanked by inverted terminal repeats (ITRs) required for genome replication and packaging. Same as above. The two cis-acting AAV ITRs are approximately 145 nucleotides in length and have an interrupted palindromic sequence that can fold into a T-shaped hairpin structure that acts as a primer during the initiation of DNA replication.

[0008] However, the use of conventional AAV as a gene delivery vector has certain disadvantages. One major disadvantage is related to the limited viral packaging capacity of AAV's approximately 4.5 kb heterologous DNA. (Dong et al., Hum Gene Ther. 7 (17): 2101-12, 1996). In addition, the administration of AAV vectors can induce human immune responses. Although it has been shown that the immunogenicity of AAV is lower than that of some other viruses (i.e., adenovirus), the capsid protein can trigger multiple components of the human immune system. See Naso et al., 2017. AAV is a common virus in the population, and most people have been exposed to AAV, so most people have already developed an immune response to the specific variants they have previously been exposed to. This pre-existing adaptive response may include neutralizing antibodies (NAb) and T cells, which can reduce the clinical efficacy of subsequent reinfection with AAV and / or the elimination of cells that have been transduced, which may make patients with pre-existing anti-AAV immunity ineligible for AVV-based gene therapy treatment. In addition, there is evidence that the T-shaped hairpin loop of AAV ITR is susceptible to inhibition by host cell proteins / protein complexes that bind to the T-shaped hairpin structure of AAV ITR. See, for example, Zhou et al., Scientific Reports 7:5432 (July 14, 2017).

[0009] Therefore, there is a need in the art for efficient and durable expression of target sequences (e.g., therapeutic proteins and / or miRNAs) in both in vitro and in vivo settings while avoiding some of the unintended consequences and limitations of current AAV vector technology. Summary of the Invention

[0010] In certain aspects, a nucleic acid molecule is provided, comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette, the gene cassette comprising a heterologous polynucleotide sequence, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0011] In certain exemplary embodiments, the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 180, and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 181. In certain exemplary embodiments, the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 183, and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 184. In certain exemplary embodiments, the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 185, and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 186. In certain exemplary embodiments, the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 187, and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 188.

[0012] In certain exemplary embodiments, the first ITR and / or the second ITR consists of the nucleotide sequence set forth in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187, or 188. In certain exemplary embodiments, the first ITR and the second ITR are reverse complements of each other.

[0013] In certain exemplary embodiments, the nucleic acid molecule further comprises a promoter. In certain exemplary embodiments, the activator is a tissue-specific promoter. In certain exemplary embodiments, the promoter drives the expression of the heterologous polynucleotide sequence in an organ selected from the group consisting of muscle, central nervous system (CNS), eyes, liver, heart, kidney, pancreas, lung, skin, bladder, urinary tract, or any combination thereof. In certain exemplary embodiments, the promoter drives the expression of the heterologous polynucleotide sequence in the following cells: hepatocytes, endothelial cells, cardiomyocytes, skeletal muscle cells, sinusoidal cells, afferent neurons, efferent neurons, interneurons, glial cells, astrocytes, oligodendrocytes, microglia, ependymal cells, lung epithelial cells, Schwann cells, satellite cells, photoreceptor cells, retinal ganglion cells, or any combination thereof. In certain exemplary embodiments, the promoter is positioned at the 5' of the heterologous polynucleotide sequence. In certain exemplary embodiments, the promoter is selected from mouse thyroxine promoter (mTTR), endogenous human factor VIII promoter (F8), human alpha-1-antitrypsin promoter (hAAT), human albumin minimal promoter, mouse albumin promoter, triple tetraproline (TTP) promoter, CASI promoter, CAG promoter, cytomegalovirus (CMV) promoter, alpha 1-antitrypsin (AAT), muscle creatine kinase (MCK), myosin heavy chain alpha (alphaMHC), myoglobin (MB), desmin (DES), SPc5-12, 2R5Sc5-12, dMCK, tMCK and phosphoglycerate kinase (PGK) promoter.

[0014] In certain exemplary embodiments, the heterologous polynucleotide sequence further comprises an intron sequence. In certain exemplary embodiments, the intron sequence is located 5' of the heterologous polynucleotide sequence. In certain exemplary embodiments, the intron sequence is located 3' of the promoter. In certain exemplary embodiments, the intron sequence comprises a synthetic intron sequence. In certain exemplary embodiments, the intron sequence comprises SEQ ID NO: 115 or 192.

[0015] In certain exemplary embodiments, the gene cassette further comprises a post-transcriptional regulatory element. In certain exemplary embodiments, the post-transcriptional regulatory element is located 3' to the heterologous polynucleotide sequence. In certain exemplary embodiments, the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, or any combination thereof. In certain exemplary embodiments, the microRNA binding site comprises a binding site for miR142-3p.

[0016] In certain exemplary embodiments, the gene cassette further comprises a 3'UTR poly(A) tail sequence. In certain exemplary embodiments, the 3'UTR poly(A) tail sequence is selected from bGH poly(A), actin poly(A), hemoglobin poly(A), and any combination thereof. In certain exemplary embodiments, the 3'UTR poly(A) tail sequence comprises bGH poly(A).

[0017] In certain exemplary embodiments, the gene cassette further comprises an enhancer sequence. In certain exemplary embodiments, the enhancer sequence is positioned between the first ITR and the second ITR.

[0018] In certain exemplary embodiments, the nucleic acid molecule comprises, from 5' to 3', a first ITR, a gene cassette, and a second ITR; wherein the gene cassette comprises a tissue-specific promoter sequence, an intron sequence, a heterologous polynucleotide sequence, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence. In certain exemplary embodiments, the gene cassette comprises, from 5' to 3', a tissue-specific promoter sequence, an intron sequence, a heterologous polynucleotide sequence, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence. In certain exemplary embodiments, the tissue-specific promoter sequence comprises a TTT promoter; the intron is a synthetic intron; the post-transcriptional regulatory element comprises a WPRE; and the 3'UTR poly(A) tail sequence comprises bGHpA.

[0019] In certain exemplary embodiments, the gene cassette comprises a single-stranded nucleic acid. In certain exemplary embodiments, the gene cassette comprises a double-stranded nucleic acid.

[0020] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a coagulation factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or any combination thereof.

[0021] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a growth factor selected from the group consisting of adrenomedullin (AM), angiopoietin (Ang), autotaxin, bone morphogenetic protein (BMP) (e.g., BMP2, BMP4, BMP5, BMP7), ciliary neurotrophic factor family members (e.g., ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6)), colony stimulating factors (e.g., macrophage colony stimulating factor (m-CSF), granulocyte colony stimulating factor (G-CSF), granulocyte macrophage colony stimulating factor (GM-CSF)), epidermal growth factor (EGF), ephrins (e.g., ephrin A1), , ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1, ephrin B2, ephrin B3), erythropoietin (EPO), fibroblast growth factor (FGF) (e.g., FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, FGF23), fetal bovine somatotropin (FBS), GDNF family members (e.g., glial cell line-derived neurotrophic factor growth factor (GDNF), neurotrophic factor (NGF), persephin, artemin), growth differentiation factor-9 (GDF9), hepatocyte growth factor (HGF), hepatoma-derived growth factor (HDGF), insulin, insulin-like growth factor (e.g., insulin-like growth factor-1 (IGF-1) or IGF-2, interleukins (IL) (e.g., IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7), keratinocyte growth factor (KGF), migration stimulating factor (MSF), macrophage stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)), myostatin ( GDF-8), neuregulins (e.g., neuregulin 1 (NRG1), NRG2, NRG3, NRG4), neurotrophins (e.g., brain-derived neurotrophic factor (BDNF), nerve growth factor (NGF), neurotrophin-3 (NT-3), NT-4, placental growth factor (PGF), platelet-derived growth factor (PDGF), renalase (RNLS), T-cell growth factor (TCGF), thrombopoietin (TPO), transforming growth factors (e.g., transforming growth factor alpha (TGF-α), TGF-β, tumor necrosis factor-α (TNF-α), and vascular endothelial growth factor (VEGF), and any combination thereof.

[0022] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a hormone.

[0023] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a cytokine.

[0024] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes an antibody or fragment thereof.

[0025] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a gene selected from the group consisting of X-linked dystrophin, MTM1 (myotubulein), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (mitochondrial ataxia protein (frataxin)), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), and any combination thereof.

[0026] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a microRNA (miRNA). In certain exemplary embodiments, the miRNA downregulates the expression of a target gene selected from the group consisting of SOD1, HTT, RHO, and any combination thereof.

[0027] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a coagulation factor selected from the group consisting of factor I (FI), factor II (FII), factor III (FIII), factor IV (FVI), factor V (FV), factor VI (FVI), factor VII (FVII), factor VIII (FVIII), factor IX (FIX), factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FVIII), von Willebrand factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, α2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), and any combination thereof.

[0028] In certain exemplary embodiments, the coagulation factor is FVIII. In certain exemplary embodiments, the FVIII comprises full-length mature FVIII. In certain exemplary embodiments, the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 106.

[0029] In certain exemplary embodiments, FVIII comprises an A1 domain, an A2 domain, an A3 domain, a C1 domain, a C2 domain, and part of or no B domain. In certain exemplary embodiments, FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 109.

[0030] In certain exemplary embodiments, the coagulation factor comprises a heterologous portion. In certain exemplary embodiments, the heterologous portion is selected from the C-terminal peptide (CTP) of the β subunit of albumin or its fragment, immunoglobulin Fc region, human chorionic gonadotropin, PAS sequence, HAP sequence, transferrins or its fragment, albumin binding moiety, its derivative or its any combination. In certain exemplary embodiments, the heterologous portion is connected to the N-terminal or C-terminal of FVIII or inserts between two amino acids of FVIII. In certain exemplary embodiments, the heterologous portion is inserted between two amino acids at one or more insertion sites selected from the insertion sites listed in Table 4.

[0031] In certain exemplary embodiments, FVIII further comprises an A1 domain, an A2 domain, a C1 domain, a C2 domain, an optional B domain, and a heterologous portion, wherein the heterologous portion is inserted immediately downstream of amino acid 745 corresponding to mature FVIII (SEQ ID NO: 106).

[0032] In certain exemplary embodiments, FVIII further comprises an FcRn binding partner. In certain exemplary embodiments, the FcRn binding partner comprises the Fc region of an immunoglobulin constant domain.

[0033] In certain exemplary embodiments, the nucleic acid sequence encoding FVIII is codon-optimized. In certain exemplary embodiments, the nucleic acid sequence encoding FVIII is codon-optimized for expression in humans.

[0034] In certain exemplary embodiments, the nucleic acid sequence encoding FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO: 107.

[0035] In certain exemplary embodiments, the nucleic acid sequence encoding FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO:71.

[0036] In certain exemplary embodiments, the heterologous polynucleotide sequence is codon-optimized. In certain exemplary embodiments, the heterologous polynucleotide sequence is codon-optimized for expression in humans.

[0037] In certain exemplary embodiments, the nucleic acid molecule is formulated with a delivery agent. In certain exemplary embodiments, the delivery agent comprises a lipid nanoparticle. In certain exemplary embodiments, the delivery agent is selected from liposomes, non-lipid polymeric molecules, and endosomes, and any combination thereof.

[0038] In certain exemplary embodiments, the nucleic acid molecule is formulated for intravenous, transdermal, intradermal, subcutaneous, pulmonary or oral delivery, or any combination thereof. In certain exemplary embodiments, the nucleic acid molecule is formulated for intravenous delivery.

[0039] In certain aspects, a vector comprising a nucleic acid molecule as described herein is provided.

[0040] In certain aspects, a host cell comprising a nucleic acid molecule as described herein is provided.

[0041] In certain aspects, pharmaceutical compositions are provided, comprising a nucleic acid molecule or vector as described herein and a pharmaceutically acceptable excipient.

[0042] In certain aspects, pharmaceutical compositions are provided, comprising a host cell as described herein and a pharmaceutically acceptable excipient.

[0043] In certain aspects, a kit is provided comprising a nucleic acid molecule as described herein and instructions for administering the nucleic acid molecule to a subject in need thereof.

[0044] In certain aspects, baculovirus systems are provided for producing nucleic acid molecules as described herein.

[0045] In certain exemplary embodiments, nucleic acid molecules as described herein are produced in insect cells.

[0046] In certain aspects, nanoparticle delivery systems comprising a nucleic acid molecule as described herein are provided.

[0047] In certain aspects, methods of producing a polypeptide are provided, comprising culturing a host cell as described herein under suitable conditions and recovering the polypeptide.

[0048] In certain aspects, a method of producing a polypeptide having coagulation activity is provided, comprising culturing a host cell as described herein under suitable conditions and recovering the polypeptide having coagulation activity.

[0049] In certain aspects, methods are provided for expressing a heterologous polynucleotide sequence in a subject in need thereof, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, or a pharmaceutical composition as described herein.

[0050] In certain aspects, methods are provided for expressing a coagulation factor in a subject in need thereof, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, a polypeptide as described herein, or a pharmaceutical composition as described herein.

[0051] In certain aspects, provided are methods of treating a disease or condition in a subject in need thereof, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, or a pharmaceutical composition as described herein.

[0052] In certain aspects, methods of treating a subject having a coagulation factor deficiency are provided, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, a polypeptide as described herein, or a pharmaceutical composition as described herein.

[0053] In certain aspects, methods are provided for treating a coagulation factor deficiency in a subject in need thereof, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, a polypeptide as described herein, or a pharmaceutical composition as described herein.

[0054] In certain exemplary embodiments, the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonary, or any combination thereof. In certain exemplary embodiments, the nucleic acid molecule is administered intravenously.

[0055] In certain exemplary embodiments, the method further comprises administering a second agent to the subject.

[0056] In certain exemplary embodiments, the subject is a mammal. In certain exemplary embodiments, the subject is a human.

[0057] In certain exemplary embodiments, administration of the nucleic acid molecule to a subject results in increased FVIII activity relative to the FVIII activity in the subject prior to administration, wherein the FVIII activity is increased by at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 11-fold, at least about 12-fold, at least about 13-fold, at least about 14-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, or at least about 100-fold.

[0058] In certain exemplary embodiments, the subject has a bleeding disorder. In certain exemplary embodiments, the bleeding disorder is hemophilia. In certain exemplary embodiments, the bleeding disorder is hemophilia A.

[0059] In certain aspects, methods are provided for treating a bleeding disorder in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding a coagulation factor, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0060] In certain aspects, methods are provided for treating hemophilia A in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding Factor VIII (FVIII), wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0061] In certain aspects, methods are provided for treating a liver metabolic disorder in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding a liver-related metabolic enzyme that is deficient in the subject, wherein the first ITR and / or the second ITR is an ITR of a non-adeno-associated virus (non-AAV).

[0062] In certain exemplary embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0063] In certain aspects, methods are provided for treating a liver metabolic disorder in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding a liver-associated metabolic enzyme that is deficient in the subject, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0064] In certain exemplary embodiments, the gene cassette comprises a single-stranded nucleic acid. In certain exemplary embodiments, the gene cassette comprises a double-stranded nucleic acid.

[0065] In certain exemplary embodiments, the liver metabolic disorder is selected from phenylketonuria (PKU), a urea cycle disease, a lysosomal storage disorder, and a glycogen storage disease. In certain exemplary embodiments, the liver metabolic disorder is phenylketonuria (PKU).

[0066] In certain exemplary embodiments, the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonary, or any combination thereof. In certain exemplary embodiments, the nucleic acid molecule is administered intravenously.

[0067] In certain exemplary embodiments, the method further comprises administering a second agent to the subject.

[0068] In certain exemplary embodiments, the subject is a mammal. In certain exemplary embodiments, the subject is a human.

[0069] In certain aspects, methods are provided for treating phenylketonuria (PKU) in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding phenylalanine hydroxylase, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0070] In certain exemplary embodiments, the gene cassette comprises a single-stranded nucleic acid. In certain exemplary embodiments, the gene cassette comprises a double-stranded nucleic acid.

[0071] In certain exemplary embodiments, the nucleic acid molecule is formulated with a delivery agent. In certain exemplary embodiments, the delivery agent comprises lipid nanoparticles.

[0072] In certain aspects, methods are provided for cloning nucleic acid molecules comprising inserting a nucleic acid molecule capable of forming a complex secondary structure into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex.

[0073] In certain exemplary embodiments, the disruption in the SbcCD complex comprises a gene disruption in the SbcC gene and / or the SbcD gene. In certain exemplary embodiments, the disruption in the SbcCD complex comprises a gene disruption in the SbcC gene. In certain exemplary embodiments, the disruption in the SbcCD complex comprises a gene disruption in the SbcD gene.

[0074] In certain exemplary embodiments, the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first and / or second ITR is a non-adeno-associated virus (non-AAV) ITR.

[0075] In certain exemplary embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0076] In certain exemplary embodiments, the nucleic acid molecule further comprises a gene cassette, wherein the gene cassette is flanked by a first ITR and a second ITR.

[0077] In certain exemplary embodiments, the gene cassette comprises a heterologous polynucleotide sequence.

[0078] In certain exemplary embodiments, the suitable vector is a low copy vector. In certain exemplary embodiments, the suitable vector is pBR322.

[0079] In certain exemplary embodiments, the bacterial host strain is unable to resolve the cruciform DNA structure.

[0080] In certain exemplary embodiments, the bacterial host strain is PMC103, which comprises the genotypes sbcC, recD, mcrA, and ΔmcrBCF. In certain exemplary embodiments, the bacterial host strain is PMC107, which comprises the genotypes recBC, recJ, sbcBC, mcrA, and ΔmcrBCF. In certain exemplary embodiments, the bacterial host strain is SURE, which comprises the genotypes recB, recJ, sbcC, mcrA, and ΔmcrBCF, umuC, and uvrC.

[0081] In certain aspects, a method of cloning a nucleic acid molecule is provided, comprising inserting a nucleic acid molecule capable of forming a complex secondary structure into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex, wherein the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0082] Specifically, the present invention includes but is not limited to the following:

[0083] 1. A nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette, wherein the gene cassette comprises a heterologous polynucleotide sequence, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0084] 2. The nucleic acid molecule of claim 1 , wherein the first ITR comprises the nucleotide sequence described in SEQ ID NO: 180, and the second ITR comprises the nucleotide sequence described in SEQ ID NO: 181.

[0085] 3. The nucleic acid molecule of claim 1 , wherein the first ITR comprises the nucleotide sequence described in SEQ ID NO: 183, and the second ITR comprises the nucleotide sequence described in SEQ ID NO: 184.

[0086] 4. The nucleic acid molecule of claim 1 , wherein the first ITR comprises the nucleotide sequence described in SEQ ID NO: 185, and the second ITR comprises the nucleotide sequence described in SEQ ID NO: 186.

[0087] 5. The nucleic acid molecule of claim 1 , wherein the first ITR comprises the nucleotide sequence described in SEQ ID NO: 187, and the second ITR comprises the nucleotide sequence described in SEQ ID NO: 188.

[0088] 6. The nucleic acid molecule of claim 1 , wherein the first ITR and / or the second ITR consists of the nucleotide sequence described in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188.

[0089] 7. The nucleic acid molecule of item 1, wherein the first ITR and the second ITR are reverse complements of each other.

[0090] 8. The nucleic acid molecule according to any one of items 1 to 7, further comprising a promoter.

[0091] 9. The nucleic acid molecule of claim 8, wherein the promoter is a tissue-specific promoter.

[0092] 10. A nucleic acid molecule as described in item 8 or 9, wherein the promoter drives the expression of the heterologous polynucleotide sequence in an organ selected from the group consisting of muscle, central nervous system (CNS), eye, liver, heart, kidney, pancreas, lung, skin, bladder, urinary tract or any combination thereof.

[0093] 11. A nucleic acid molecule as described in any one of items 8 to 10, wherein the promoter drives expression of the heterologous polynucleotide sequence in the following cells: hepatocytes, endothelial cells, cardiomyocytes, skeletal muscle cells, sinusoidal cells, afferent neurons, efferent neurons, interneurons, glial cells, astrocytes, oligodendrocytes, microglia, ependymal cells, lung epithelial cells, Schwann cells, satellite cells, photoreceptor cells, retinal ganglion cells or any combination thereof.

[0094] 12. The nucleic acid molecule of any one of items 8 to 11, wherein the promoter is located 5' to the heterologous polynucleotide sequence.

[0095] 13. A nucleic acid molecule as described in any one of items 8 to 12, wherein the promoter is selected from the following group: mouse thyroxine promoter (mTTR), endogenous human factor VIII promoter (F8), human α-1-antitrypsin promoter (hAAT), human albumin minimal promoter, mouse albumin promoter, triple tetraproline (TTP) promoter, CASI promoter, CAG promoter, cytomegalovirus (CMV) promoter, α1-antitrypsin (AAT), muscle creatine kinase (MCK), myosin heavy chain α (αMHC), myoglobin (MB), desmin (DES), SPc5-12, 2R5Sc5-12, dMCK, tMCK and phosphoglycerate kinase (PGK) promoter.

[0096] 14. The nucleic acid molecule of any one of items 1 to 13, wherein the heterologous polynucleotide sequence further comprises an intron sequence.

[0097] 15. The nucleic acid molecule of claim 14, wherein the intron sequence is located 5' to the heterologous polynucleotide sequence.

[0098] 16. The nucleic acid molecule of item 14 or 15, wherein the intron sequence is located 3' to the promoter.

[0099] 17. The nucleic acid molecule of any one of items 14 to 16, wherein the intron sequence comprises a synthetic intron sequence.

[0100] 18. The nucleic acid molecule of any one of items 14 to 17, wherein the intron sequence comprises SEQ ID NO: 115 or 192.

[0101] 19. The nucleic acid molecule of any one of items 1 to 18, wherein the gene cassette further comprises a post-transcriptional regulatory element.

[0102] 20. The nucleic acid molecule of claim 19, wherein the post-transcriptional regulatory element is located 3' to the heterologous polynucleotide sequence.

[0103] 21. A nucleic acid molecule as described in item 19 or 20, wherein the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence or any combination thereof.

[0104] 22. The nucleic acid molecule of claim 21 , wherein the microRNA binding site comprises a binding site for miR142-3p.

[0105] 23. The nucleic acid molecule of any one of items 1 to 22, wherein the gene cassette further comprises a 3'UTR poly(A) tail sequence.

[0106] 24. The nucleic acid molecule of item 23, wherein the 3'UTR poly(A) tail sequence is selected from the group consisting of bGH poly(A), actin poly(A), hemoglobin poly(A), and any combination thereof.

[0107] 25. The nucleic acid molecule of item 23 or 24, wherein the 3'UTR poly(A) tail sequence comprises bGH poly(A).

[0108] 26. The nucleic acid molecule of any one of items 1 to 25, wherein the gene cassette further comprises an enhancer sequence.

[0109] 27. The nucleic acid molecule of item 26, wherein the enhancer sequence is positioned between the first ITR and the second ITR.

[0110] 28. A nucleic acid molecule as described in any one of items 1 to 27, wherein the nucleic acid molecule comprises from 5' to 3': the first ITR, the gene cassette and the second ITR; wherein the gene cassette comprises a tissue-specific promoter sequence, an intron sequence, a heterologous polynucleotide sequence, a post-transcriptional regulatory element and a 3'UTR poly (A) tail sequence.

[0111] 29. The nucleic acid molecule of item 28, wherein the gene cassette comprises from 5' to 3': a tissue-specific promoter sequence, an intron sequence, the heterologous polynucleotide sequence, a post-transcriptional regulatory element and a 3'UTR poly(A) tail sequence.

[0112] 30. The nucleic acid molecule of item 28 or 29, wherein:

[0113] The tissue-specific promoter sequence comprises a TTT promoter;

[0114] The intron is a synthetic intron;

[0115] The post-transcriptional regulatory element comprises WPRE; and

[0116] The 3'UTR poly (A) tail sequence comprises bGHpA.

[0117] 31. The nucleic acid molecule of any one of items 1 to 30, wherein the gene cassette comprises a single-stranded nucleic acid.

[0118] 32. The nucleic acid molecule of any one of items 1 to 30, wherein the gene cassette comprises a double-stranded nucleic acid.

[0119] 33. The nucleic acid molecule of any one of items 1 to 32, wherein the heterologous polynucleotide sequence encodes a coagulation factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or any combination thereof.

[0120] 34. A nucleic acid molecule as described in claim 33, wherein the heterologous polynucleotide sequence encodes a growth factor selected from the group consisting of adrenomedullin (AM), angiogenin (Ang), autocrine motility factor, bone morphogenetic protein (BMP) (e.g., BMP2, BMP4, BMP5, BMP7), ciliary neurotrophic factor family members (e.g., ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6)), colony stimulating factors (e.g., macrophage colony stimulating factor (m-CSF), granulocyte colony stimulating factor (G-CSF), granulocyte macrophage colony stimulating factor (GM-CSF)), epidermal growth factor (EGF), liver glycosides proteins (e.g., ephrin A1, ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1, ephrin B2, ephrin B3), erythropoietin (EPO), fibroblast growth factor (FGF) (e.g., FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, FGF23), fetal bovine somatotropin (FBS), GDNF family members (e.g., glial cell line-derived neurotrophic factor (GDNF), neurotrophin, proteolytic protein, atermin), growth differentiation factor-9 (GDF9), hepatocyte growth factor (HGF), hepatoma-derived growth factor (HDGF), insulin, insulin-like growth factor (e.g., insulin-like growth factor-1 (IGF-1) or IGF-2, interleukins (IL) (e.g., IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7), keratinocyte growth factor (KGF), migration stimulating factor (MSF), macrophage stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)), myostatin

[00135] The present invention also includes but is not limited to: neuregulin (GDF-8), neuregulin (e.g., neuregulin 1 (NRG1), NRG2, NRG3, NRG4), neurotrophins (e.g., brain-derived neurotrophic factor (BDNF), nerve growth factor (NGF), neurotrophin-3 (NT-3), NT-4, placental growth factor (PGF), platelet-derived growth factor (PDGF), renalase (RNLS), T-cell growth factor (TCGF), thrombopoietin (TPO), transforming growth factors (e.g., transforming growth factor alpha (TGF-α), TGF-β, tumor necrosis factor-α (TNF-α) and vascular endothelial growth factor (VEGF), and any combination thereof.

[0121] 35. A nucleic acid molecule as described in item 33, wherein the heterologous polynucleotide sequence encodes a hormone.

[0122] 36. A nucleic acid molecule as described in item 33, wherein the heterologous polynucleotide sequence encodes a cytokine.

[0123] 37. A nucleic acid molecule as described in item 33, wherein the heterologous polynucleotide sequence encodes an antibody or a fragment thereof.

[0124] 38. A nucleic acid molecule as described in any one of items 1 to 32, wherein the heterologous polynucleotide sequence encodes a gene selected from the group consisting of: X-linked dystrophin, MTM1 (myotubulein), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (mitochondrial ataxia protein), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase) and any combination thereof.

[0125] 39. The nucleic acid molecule of any one of items 1 to 32, wherein the heterologous polynucleotide sequence encodes a microRNA (miRNA).

[0126] 40. The nucleic acid molecule of claim 39, wherein the miRNA downregulates the expression of a target gene selected from the group consisting of SOD1, HTT, RHO, and any combination thereof.

[0127] 41. A nucleic acid molecule as described in any one of items 1 to 32, wherein the heterologous polynucleotide sequence encodes a coagulation factor selected from the group consisting of factor I (FI), factor II (FII), factor III (FIII), factor IV (FVI), factor V (FV), factor VI (FVI), factor VII (FVII), factor VIII (FVIII), factor IX (FIX), factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FVIII), von Willebrand factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, α2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2) and any combination thereof.

[0128] 42. A nucleic acid molecule as described in item 41, wherein the coagulation factor is FVIII.

[0129] 43. A nucleic acid molecule as described in claim 42, wherein the FVIII comprises full-length mature FVIII.

[0130] 44. A nucleic acid molecule as described in claim 43, wherein the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the amino acid sequence having SEQ ID NO: 106.

[0131] 45. The nucleic acid molecule of claim 42, wherein the FVIII comprises an A1 domain, an A2 domain, an A3 domain, a C1 domain, a C2 domain, and a partial domain or no B domain.

[0132] 46. ​​A nucleic acid molecule as described in claim 45, wherein the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the amino acid sequence of SEQ ID NO:109.

[0133] 47. A nucleic acid molecule as described in any one of items 41 to 46, wherein the coagulation factor comprises a heterologous part.

[0134] 48. A nucleic acid molecule as described in claim 47, wherein the heterologous part is selected from the following group: albumin or a fragment thereof, an immunoglobulin Fc region, a C-terminal peptide (CTP) of the β subunit of human chorionic gonadotropin, a PAS sequence, a HAP sequence, transferrin or a fragment thereof, an albumin binding part, a derivative thereof or any combination thereof.

[0135] 49. A nucleic acid molecule according to claim 47 or 48, wherein the heterologous portion is linked to the N-terminus or C-terminus of the FVIII or is inserted between two amino acids in the FVIII.

[0136] 50. A nucleic acid molecule as described in item 49, wherein the heterologous part is inserted between two amino acids at one or more insertion sites selected from the insertion sites listed in Table 4.

[0137] 51. The nucleic acid molecule of any one of items 42 to 50, wherein the FVIII further comprises an A1 domain, an A2 domain, a C1 domain, a C2 domain, an optional B domain, and a heterologous portion, wherein the heterologous portion is inserted immediately downstream of amino acid 745 corresponding to mature FVIII (SEQ ID NO: 106).

[0138] 52. The nucleic acid molecule of any one of items 49 to 51, wherein the FVIII further comprises an FcRn binding partner.

[0139] 53. A nucleic acid molecule as described in item 52, wherein the FcRn binding partner comprises the Fc region of an immunoglobulin constant domain.

[0140] 54. The nucleic acid molecule of any one of items 42 to 53, wherein the nucleic acid sequence encoding the FVIII is codon-optimized.

[0141] 55. The nucleic acid molecule of any one of items 42 to 54, wherein the nucleic acid sequence encoding the FVIII is codon-optimized for expression in humans.

[0142] 56. A nucleic acid molecule as described in claim 55, wherein the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or about 100% identical to the nucleotide sequence of SEQ ID NO: 107.

[0143] 57. A nucleic acid molecule as described in claim 55, wherein the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or about 100% identical to the nucleotide sequence of SEQ ID NO:71.

[0144] 58. A nucleic acid molecule as described in any one of items 1 to 53, wherein the heterologous polynucleotide sequence is codon optimized.

[0145] 59. A nucleic acid molecule as described in item 58, wherein the heterologous polynucleotide sequence is codon-optimized for expression in humans.

[0146] 60. The nucleic acid molecule of any one of items 1 to 59, wherein the nucleic acid molecule is formulated with a delivery agent.

[0147] 61. A nucleic acid molecule as described in claim 60, wherein the delivery agent comprises lipid nanoparticles.

[0148] 62. The nucleic acid molecule of claim 60, wherein the delivery agent is selected from the group consisting of liposomes, non-lipid polymeric molecules, and endosomes, and any combination thereof.

[0149] 63. The nucleic acid molecule of any one of items 1 to 62, wherein the nucleic acid molecule is formulated for intravenous, transdermal, intradermal, subcutaneous, pulmonary or oral delivery, or any combination thereof.

[0150] 64. A nucleic acid molecule as described in item 63, wherein the nucleic acid molecule is formulated for intravenous delivery.

[0151] 65. A vector comprising the nucleic acid molecule of any one of items 1 to 59.

[0152] 66. A host cell comprising the nucleic acid molecule of any one of items 1 to 59.

[0153] 67. A pharmaceutical composition comprising the nucleic acid of any one of items 1 to 59 or the vector of item 65, and a pharmaceutically acceptable excipient.

[0154] 68. A pharmaceutical composition comprising the host cell of claim 66 and a pharmaceutically acceptable excipient.

[0155] 69. A kit comprising the nucleic acid molecule of any one of items 1 to 59 and instructions for administering the nucleic acid molecule to a subject in need thereof.

[0156] 70. A baculovirus system for producing the nucleic acid molecule of any one of items 1 to 59.

[0157] 71. The baculovirus system of claim 70, wherein the nucleic acid molecule of any one of claims 1 to 59 is produced in insect cells.

[0158] 72. A nanoparticle delivery system comprising the nucleic acid molecule of any one of items 1 to 59.

[0159] 73. A method for producing a polypeptide, which comprises culturing the host cell of item 66 under suitable conditions and recovering the polypeptide.

[0160] 75. A method for producing a polypeptide with coagulation activity, comprising: culturing the host cell of item 66 under suitable conditions and recovering the polypeptide with coagulation activity.

[0161] 76. A method of expressing a heterologous polynucleotide sequence in a subject in need thereof, comprising administering to the subject a nucleic acid molecule of any one of items 1 to 59, a vector of item 65 or a pharmaceutical composition of item 67.

[0162] 77. A method of expressing a coagulation factor in a subject in need thereof, comprising administering to the subject a nucleic acid molecule of any one of items 41 to 57, a vector of item 65, a polypeptide of item 75 or a pharmaceutical composition of item 67.

[0163] 78. A method of treating a disease or condition in a subject in need thereof, comprising administering to the subject a nucleic acid molecule of any one of items 1 to 59, a vector of item 65, or a pharmaceutical composition of item 67.

[0164] 79. A method for treating a subject suffering from a coagulation factor deficiency, comprising administering to the subject a nucleic acid molecule of any one of items 41 to 57, a vector of item 65, a polypeptide of item 75 or a pharmaceutical composition of item 67.

[0165] 80. A method of treating a coagulation factor deficiency in a subject in need thereof, comprising administering to the subject a nucleic acid molecule of any one of items 41 to 57, a vector of item 65, a polypeptide of item 75 or a pharmaceutical composition of item 67.

[0166] 81. The method of claim 79 or 80, wherein the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonary, or any combination thereof.

[0167] 82. The method of claim 81 , wherein the nucleic acid molecule is administered intravenously.

[0168] 83. The method of any one of items 79 to 82, further comprising administering a second agent to the subject.

[0169] 84. The method of any one of items 79 to 83, wherein the subject is a mammal.

[0170] 85. The method of any one of items 79 to 84, wherein the subject is a human.

[0171] 86. The method of any one of items 79 to 85, wherein administration of the nucleic acid molecule to the subject results in increased FVIII activity relative to the FVIII activity in the subject before the administration, wherein the increase in FVIII activity is at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 11-fold, at least about 12-fold, at least about 13-fold, at least about 14-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, or at least about 100-fold.

[0172] 87. The method of any one of items 79 to 86, wherein the subject has a bleeding disorder.

[0173] 88. The method of claim 87, wherein the bleeding disorder is hemophilia.

[0174] 89. The method of claim 87 or 88, wherein the bleeding disorder is hemophilia A.

[0175] 90. A method for treating a bleeding disorder in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding a coagulation factor, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0176] 91. A method of treating hemophilia A in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding Factor VIII (FVIII), wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0177] 92. A method for treating a liver metabolic disorder in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding a liver-related metabolic enzyme that is deficient in the subject, wherein the first ITR and / or the second ITR is an ITR of a non-adeno-associated virus (non-AAV).

[0178] 93. A method as described in claim 92, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence described in SEQ ID NO:180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0179] 94. A method of treating a liver metabolic disorder in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding a liver-associated metabolic enzyme that is deficient in the subject, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0180] 95. The method of any one of items 92 to 94, wherein the gene cassette comprises a single-stranded nucleic acid.

[0181] 96. The method of any one of items 92 to 94, wherein the gene cassette comprises a double-stranded nucleic acid.

[0182] 97. The method of any one of items 92 to 96, wherein the liver metabolic disorder is selected from the group consisting of phenylketonuria (PKU), urea cycle diseases, lysosomal storage disorders, and glycogen storage diseases.

[0183] 98. The method of claim 97, wherein the liver metabolic disorder is phenylketonuria (PKU).

[0184] 99. The method of any one of items 92 to 98, wherein the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonary, or any combination thereof.

[0185] 100. The method of claim 99, wherein the nucleic acid molecule is administered intravenously.

[0186] 101. The method of any one of items 92 to 100, further comprising administering a second agent to the subject.

[0187] 102. The method of any one of items 92 to 101, wherein the subject is a mammal.

[0188] 103. The method of any one of items 92 to 102, wherein the subject is a human.

[0189] 104. A method of treating phenylketonuria (PKU) in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence encoding phenylalanine hydroxylase, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0190] 105. The method of claim 104, wherein the gene cassette comprises a single-stranded nucleic acid.

[0191] 106. The method of claim 104, wherein the gene cassette comprises a double-stranded nucleic acid.

[0192] 107. The method of any one of items 104 to 106, wherein the nucleic acid molecule is formulated with a delivery agent.

[0193] 108. A method as described in claim 107, wherein the delivery agent comprises lipid nanoparticles.

[0194] 109. A method of cloning a nucleic acid molecule comprising inserting a nucleic acid molecule capable of forming a complex secondary structure into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex.

[0195] 110. A method as described in claim 109, wherein the disruption in the SbcCD complex comprises a gene disruption in the SbcC gene and / or the SbcD gene.

[0196] 111. The method of claim 109 or 110, wherein the disruption in the SbcCD complex comprises a gene disruption in the SbcC gene.

[0197] 112. The method of claim 109 or 110, wherein the disruption in the SbcCD complex comprises a gene disruption in the SbcD gene.

[0198] 113. A method as described in any one of items 109 to 112, wherein the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first and / or second ITR is a non-adeno-associated virus (non-AAV) ITR.

[0199] 114. A method as described in any one of items 109 to 113, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence described in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0200] 115. The method of any one of items 109 to 114, wherein the nucleic acid molecule further comprises a gene cassette, wherein the gene cassette is flanked by the first and second ITRs.

[0201] 116. A method as described in item 115, wherein the gene cassette comprises a heterologous polynucleotide sequence.

[0202] 117. A method as described in any one of items 109 to 116, wherein the suitable vector is a low copy vector.

[0203] 118. A method as described in any one of items 109 to 116, wherein the suitable vector is pBR322.

[0204] 119. A method as described in any one of items 109 to 118, wherein the bacterial host strain is unable to resolve cruciform DNA structures.

[0205] 120. The method of any one of items 109 to 118, wherein the bacterial host strain is PMC103 comprising the genotypes sbcC, recD, mcrA, ΔmcrBCF.

[0206] 121. A method as described in any one of items 109 to 118, wherein the bacterial host strain is PMC107, which comprises the genotypes recBC, recJ, sbcBC, mcrA, ΔmcrBCF.

[0207] 122. The method of any one of items 109 to 118, wherein the bacterial host strain is SURE, comprising the genotypes recB, recJ, sbcC, mcrA, ΔmcrBCF, umuC, uvrC.

[0208] 123. A method for cloning a nucleic acid molecule, comprising inserting a nucleic acid molecule capable of forming a complex secondary structure into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex, wherein the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0209] Figure 1A-1B Schematic diagram of a single-chain coagulation factor (e.g., FVIII) expression cassette. Shown are the positions of the 5' ITR from non-AAV (with a hairpin loop at the end of the ssDNA structure), the 3' ITR from non-AAV (with a hairpin loop), a promoter sequence (e.g., TTPp or CAGp), and a transgene sequence (e.g., a FVIII co6 XTEN sequence with XTEN144 inserted within the B domain). The exemplary expression cassette also shows other possible components, such as intron sequences, WPREmut sequences, and bGHpA sequences.

[0210] Figure 1C-1F It is used to prepare a single-chain coagulation factor expression cassette (such as Figure 1A-1B Schematic diagram of a plasmid containing the cassette shown in ), wherein the ITRs of the cassette are derived from AAV2 ( Figure 1C )、B19( Figure 1D )、GPV( Figure 1E ), or the wild-type B19 ITR sequence ( Figure 1F ). With PvuII (at the PvuII site) ( Figure 1C ) or LguI (at the LguI site) ( Figure 1D-1F ) digested a plasmid construct containing the ssFVIII expression cassette as shown here to precisely release the sequences containing the ITRs and the expression cassette. The double-stranded DNA was heat denatured at 95°C to generate ssDNA, and then incubated at 4°C to allow the ITR structure to form.

[0211] Figure 2A is a phylogenetic tree illustrating the relationships among the various parvoviridae family members. B19, AAV-2, and GPV are marked with outlined boxes.

[0212] Figure 2B Schematic diagram of various boxes including hairpin structures.

[0213] Figure 3A and Figure 3B It is the ITR of B19, GPV and AAV2 ( Figure 3A ) and the ITR of B19 and GPV ( Figure 3B Grey shading shows homology.

[0214] Figures 4A-4C We show that single-chain FVIII-AAV naked DNA (ssAAV-FVIII; Figure 1C )、ssDNA-B19 FVIII( Figure 1D ) or ssDNA-GPV FVIII ( Figure 1E ) after FVIII plasma activity. Figure 4C ), 20 μg / mouse ( Figure 4A and Figure 4B ), 10 μg / mouse ( Figure 4A 、 Figure 4B and Figure 4C ) or 5 μg / mouse ( Figure 4A FVIII activity (as a percentage of normal physiological levels in humans) was measured in plasma samples of mice treated with a single HDI of ssDNA containing 5 μg / mouse of plasmid DNA at 24 hours, 3 days, 2 weeks, 3 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, and 6 months. Figure 4A 、 Figure 4B and Figure 4C ).

[0215] Figure 5 The results show that after a single hydrodynamic injection of equimolar amounts of single-stranded naked DNA (ssAAV-FVIII, Figure 1A FVIII activity in the plasma of hemophilia A mice was assessed after the expression of double-stranded AAV-FVIII DNA (dsDNA) containing ITR sequences, double-stranded FVIII DNA without ITR sequences (dsDNA without ITRs), or circularized double-stranded FVIII DNA without ITRs or bacterial sequences (minicircles). dsDNA was generated by PCR amplification of the AAV-FVIII plasmid ( Figure 1C ) was generated by enzymatic cleavage without heat denaturation. The ITR-free dsDNA was generated by cleaving the AAV-FVIII plasmid ( Figure 1C ) were enzymatically cleaved and subsequently purified. Minicircle DNA was generated by ligating dsDNA without ITR DNA at the AflII site. Mouse plasma was collected over 3 or 4 months and FVIII was determined by chromogenic activity assay.

[0216] Figure 6 The results show that the hydrodynamic injection of 30 μg single-stranded naked FVIII-DNA ( Figure 1A 、 Figure 1D-1F Figure 3. FVIII activity in plasma of hemophilia A mice after 4 weeks of treatment. Plasma was collected weekly for 7 weeks and FVIII activity was determined by chromogenic assay. After 35 days (depicted as black arrows), 30 μg was administered again by hydrodynamic injection to mice receiving FVIII-B19d135 and FVIII-GPVd162 ssDNA.

[0217] Figure 7A Figure 1 is a schematic diagram of a single-chain murine phenylalanine hydroxylase (e.g., PAH) expression cassette. The locations of the non-AAV-derived 5' ITR (with a hairpin loop at the end of the ssDNA structure), the non-AAV-derived 3' ITR (with a hairpin loop), the promoter sequence (e.g., CAGp), and the transgene sequence (e.g., 3xFLAG_mPAH sequence) are shown. The exemplary expression cassette also shows other possible components, such as the WPREmut sequence and the bGHpA sequence.

[0218] Figure 7B-7D Figure 2 shows plasma concentrations of phenylalanine (Phe) in phenylketonuria (PKU) mice before (day 0) and after a single administration of single-stranded DNA containing murine PAH cDNA and non-AAV ITRB19d135 or GPVd162 by hydrodynamic injection. Plasma was collected on days 3, 7, 14, 28, 42, and 56 after ssDNA administration. Residual phenylalanine levels are shown as concentrations in μg / ml ( Figure 7B-7C ) or as a percentage before application ( Figure 7D ). The horizontal line depicts the baseline Phe level before administration.

[0219] Figure 7E Shown are Western immunoblots of liver lysates from PKU mice treated with ssDNA containing the murine PAH transgene and either the B19d135 or GPVd165 ITR. Livers were harvested on day 81 post-treatment, and protein lysates were extracted. Each well represents a single animal. FLAG-tagged murine PAH protein was detected using the M2 anti-FLAG antibody, and a GAPDH loading control was included for comparison.

[0220] Figure 8A-8B showed that the encapsulated FVIII-AAV DNA ( Figure 1A-1C FVIII activity levels in Huh7 cell supernatants after transduction with lipid nanoparticles containing CAGp promoter ( Figure 1B) were encapsulated with an amine to phosphate (NP) ratio equal to 3 and applied to Huh7 cells at different concentrations determined by picogreen assay ( Figure 8A ). In the TTPp promoter ( Figure 1A ) plasmids, double-stranded linear (ds) and single-stranded (ss) AAV-FVIII were also encapsulated in lipid nanoparticles at an NP ratio equal to 2 and used to transduce Huh7 cells at different DNA concentrations ( Figure 8B ). FVIII was measured by chromogenic activity assay compared to human FACT plasma standards. DETAILED DESCRIPTION

[0221] The present disclosure describes a plasmid-like nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette (e.g., encoding a target sequence (also referred to herein as a heterologous polynucleotide sequence), such as a therapeutic protein or miRNA), wherein the first ITR and / or the second ITR is an ITR of a non-adeno-associated virus (e.g., the first ITR and / or the second ITR is from a non-AAV). In some embodiments, the gene cassette encodes a therapeutic protein, e.g., the target sequence encodes a therapeutic protein. In some embodiments, the therapeutic protein comprises a protein selected from the group consisting of a coagulation factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or a combination thereof. In some embodiments, the gene cassette encodes X-linked dystrophin, MTM1 (myotubulin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (mitochondrial ataxia protein), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof.

[0222] In some embodiments, the therapeutic protein comprises a coagulation factor. In a specific embodiment, the therapeutic protein comprises a FVIII or FIX protein.

[0223] In some embodiments, the gene cassette encodes a miRNA. In certain embodiments, the miRNA downregulates the expression of a target gene selected from the group consisting of SOD1, HTT, RHO, or any combination thereof.

[0224] In certain embodiments, the non-AAV is selected from members of the virus family Parvoviridae and any combination thereof. The present disclosure also relates to a method for expressing a therapeutic protein (e.g., a coagulation factor (e.g., FVIII)) in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette (e.g., encoding a therapeutic protein or miRNA), wherein the first ITR and / or the second ITR are ITRs of a non-adeno-associated virus (non-AAV). In certain embodiments, the present disclosure describes an isolated nucleic acid molecule comprising a nucleotide sequence having sequence homology to a nucleotide sequence selected from SEQ ID NOs: 113 and 120.

[0225] In certain embodiments, the present disclosure provides nucleic acid molecules comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette comprising a heterologous polynucleotide sequence, wherein the first and / or second ITR is derived from parvovirus B19 or goose parvovirus (GPV).

[0226] Exemplary constructs of the present disclosure are illustrated in the accompanying drawings and sequence listing.In order to provide a clear understanding of the specification and claims, the following definitions are provided below.

[0227] I. Definition

[0228] It should be noted that the term "a" or "an" entity refers to one or more of the entity: for example, "a nucleotide sequence" should be understood to represent one or more nucleotide sequences. Similarly, "a therapeutic protein" and "a miRNA" should be understood to represent one or more therapeutic proteins and one or more miRNAs, respectively. Therefore, the terms "a" (or "an"), "one or more" and "at least one" are used interchangeably herein.

[0229] The term "about" is used herein to mean approximately, roughly, roughly, or around. When the term "about" is used in conjunction with a numerical range, it modifies that range by extending the boundaries above and below the stated values. Generally, the term "about" is used herein to modify a numerical value above or below the stated value by 10% above or below.

[0230] Also as used herein, "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when construed in the alternative ("or").

[0231] "Nucleic acid," "nucleic acid molecule," "nucleotide," "sequence of one or more nucleotides," and "polynucleotide" are used interchangeably and refer to a phosphate polymeric form of ribonucleosides (adenosine, guanosine, uridine, or cytidine; "RNA molecule") or deoxyribonucleosides (deoxyadenosine, deoxyguanosine, deoxythymidine, or deoxycytidine; "DNA molecule"), or any of their phosphate analogs, such as phosphorothioates and thioesters, in single-stranded form or in a double-stranded helix. A single-stranded nucleic acid sequence refers to single-stranded DNA (ssDNA) or single-stranded RNA (ssRNA). Double-stranded DNA-DNA, DNA-RNA, and RNA-RNA helices are possible. The term nucleic acid molecule, particularly DNA or RNA molecule, refers solely to the primary and secondary structure of the molecule and does not limit it to any particular tertiary form. Thus, this term includes double-stranded DNA found, inter alia, in linear or circular DNA molecules (e.g., restriction fragments), plasmids, supercoiled DNA, and chromosomes. When discussing the structure of a particular double-stranded DNA molecule, the sequence may be described herein according to conventional practice, with the sequence given in a 5' to 3' direction along the non-transcribed strand of the DNA (i.e., the strand with a sequence homologous to the mRNA). A "recombinant DNA molecule" is a DNA molecule that has been subjected to molecular biological manipulation. DNA includes, but is not limited to, cDNA, genomic DNA, plasmid DNA, synthetic DNA, and semisynthetic DNA. A "nucleic acid composition" of the present disclosure comprises one or more nucleic acids as described herein.

[0232] As used herein, "inverted terminal repeat" (or "ITR") refers to a nucleic acid subsequence located at the 5' end or 3' end of a single-stranded nucleic acid sequence, which comprises a group of nucleotides (the initial sequence) followed by its reverse complement, i.e., a palindrome, downstream. The inserted nucleotide sequence between the initial sequence and the reverse complement can have any length, including zero. In one embodiment, the ITRs useful in the present disclosure comprise one or more "palindromes." ITRs can have any number of functions. In some embodiments, the ITRs described herein form a hairpin structure. In some embodiments, the ITRs form a T-shaped hairpin structure. In some embodiments, the ITRs form a non-T-shaped hairpin structure, such as a U-shaped hairpin structure. In some embodiments, the ITRs promote the long-term survival of nucleic acid molecules in the cell nucleus of a cell. In some embodiments, the ITRs promote the permanent survival of nucleic acid molecules in the cell nucleus of a cell (e.g., for the entire lifespan of the cell). In some embodiments, the ITRs promote the stability of nucleic acid molecules in the cell nucleus of a cell. In some embodiments, the ITRs promote the retention of nucleic acid molecules in the cell nucleus of a cell. In some embodiments, the ITRs promote the persistence of nucleic acid molecules in the cell nucleus of a cell. In some embodiments, the ITRs inhibit or prevent the degradation of nucleic acid molecules in the cell nucleus of a cell.

[0233] In one embodiment, the original sequence and / or the reverse complement comprises about 2-600 nucleotides, about 2-550 nucleotides, about 2-500 nucleotides, about 2-450 nucleotides, about 2-400 nucleotides, about 2-350 nucleotides, about 2-300 nucleotides, or about 2-250 nucleotides. In some embodiments, the original sequence and / or the reverse complement comprises about 5-600 nucleotides, about 10-600 nucleotides, about 15-600 nucleotides, about 20-600 nucleotides, about 25-600 nucleotides, about 30-600 nucleotides, about 35-600 nucleotides, about 40-600 nucleotides, about 45-600 nucleotides, about 50-600 nucleotides, about 60-600 nucleotides, About 70-600 nucleotides, about 80-600 nucleotides, about 90-600 nucleotides, about 100-600 nucleotides, about 150-600 nucleotides, about 200-600 nucleotides, about 300-600 nucleotides, about 350-600 nucleotides, about 400-600 nucleotides, about 450-600 nucleotides, about 500-600 nucleotides, or about 550-600 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 5-550 nucleotides, about 5 to 500 nucleotides, about 5-450 nucleotides, about 5 to 400 nucleotides, about 5-350 nucleotides, about 5 to 300 nucleotides, or about 5-250 nucleotides. In certain embodiments, the initial sequence and / or reverse complement comprises about 10-550 nucleotides, about 15-500 nucleotides, about 20-450 nucleotides, about 25-400 nucleotides, about 30-350 nucleotides, about 35-300 nucleotides, or about 40-250 nucleotides. In certain embodiments, the initial sequence and / or reverse complement comprises about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, or about 600 nucleotides. In a specific embodiment, the initial sequence and / or reverse complement comprises about 400 nucleotides.

[0234] In other embodiments, the original sequence and / or the reverse complement comprises about 2-200 nucleotides, about 5-200 nucleotides, about 10-200 nucleotides, about 20-200 nucleotides, about 30-200 nucleotides, about 40-200 nucleotides, about 50-200 nucleotides, about 60-200 nucleotides, about 70-200 nucleotides, about 80-200 nucleotides, about 90-200 nucleotides, about 100-200 nucleotides, about 125-200 nucleotides, about 150-200 nucleotides, or about 175-200 nucleotides. In other embodiments, the original sequence and / or the reverse complement comprises about 2-150 nucleotides, about 5-150 nucleotides, about 10-150 nucleotides, about 20-150 nucleotides, about 30-150 nucleotides, about 40-150 nucleotides, about 50-150 nucleotides, about 75-150 nucleotides, about 100-150 nucleotides, or about 125-150 nucleotides. In other embodiments, the original sequence and / or the reverse complement comprises about 2-100 nucleotides, about 5-100 nucleotides, about 10-100 nucleotides, about 20-100 nucleotides, about 30-100 nucleotides, about 40-100 nucleotides, about 50-100 nucleotides, or about 75-100 nucleotides. In other embodiments, the original sequence and / or the reverse complement comprises about 2-50 nucleotides, about 10-50 nucleotides, about 20-50 nucleotides, about 30-50 nucleotides, about 40-50 nucleotides, about 3-30 nucleotides, about 4-20 nucleotides, or about 5-10 nucleotides. In another embodiment, the original sequence and / or the reverse complement consists of 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides. In other embodiments, the intervening nucleotides between the original sequence and the reverse complement are (e.g., consist of) 0 nucleotides, 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides.

[0235] Therefore, " ITR " as used herein can fold back on itself and form a double-stranded segment. For example, when folded to form a double helix, the sequence GATCXXXXGATC comprises the initial sequence of GATC and its complement (3'CTAG5'). In some embodiments, the ITR comprises a continuous palindromic sequence (e.g., GATCGATC) between the initial sequence and the reverse complement. In some embodiments, the ITR comprises an interrupted palindromic sequence (e.g., GATCXXXXGATC) between the initial sequence and the reverse complement. In some embodiments, the complementary parts of the continuous or interrupted palindromic sequences interact with each other to form a "hairpin loop" structure. As used herein, when at least two complementary sequence bases on a single-stranded nucleotide molecule are paired to form a double-stranded portion, a "hairpin loop" structure is produced. In some embodiments, only a portion of the ITR forms a hairpin loop. In other embodiments, the entire ITR forms a hairpin loop.

[0236] In the present disclosure, at least one ITR is an ITR of a non-adenovirus-associated virus (non-AAV). In certain embodiments, the ITR is an ITR of a non-AAV member of the Parvoviridae family of the virus family. In some embodiments, the ITR is an ITR of a non-AAV member of the genus Dependovirus or the genus Erythrovirus. In specific embodiments, the ITR is an ITR of the following viruses: goose parvovirus (GPV), Muscovy duck parvovirus (MDPV), or Erythrovirus parvovirus B19 (also known as parvovirus B19, primate erythroparvovirus 1, B19 virus, and erythrovirus). In certain embodiments, one of the two ITRs is an ITR of AAV. In other embodiments, one of the two ITRs in the construct is an ITR of an AAV serotype selected from the following: serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and any combination thereof. In a specific embodiment, the ITR is derived from AAV serotype 2, such as an ITR of AAV serotype 2.

[0237] In certain aspects of the present disclosure, the nucleic acid molecule comprises two ITRs, a 5'ITR and a 3'ITR, wherein the 5'ITR is located at the 5' end of the nucleic acid molecule and the 3'ITR is located at the 3' end of the nucleic acid molecule. The 5'ITR and the 3'ITR can be derived from the same virus or different viruses. In certain embodiments, the 5'ITR is derived from AAV and the 3'ITR is not derived from an AAV virus (e.g., non-AAV). In certain embodiments, the 3'ITR is derived from AAV and the 5'ITR is not derived from an AAV virus (e.g., non-AAV). In other embodiments, the 5'ITR is not derived from an AAV virus (e.g., non-AAV), and the 3'ITR is derived from the same or different non-AAV viruses.

[0238] As used herein, the term "parvovirus" encompasses the family Parvoviridae, including but not limited to autonomously replicating parvoviruses and Dependinoviruses. Autonomous parvoviruses include, for example, members of the genera Bocavirus, Dependinovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstyldensovirus.

[0239] Exemplary autonomous parvoviruses include, but are not limited to, porcine parvovirus, minute virus of mice, canine parvovirus, mink enteritis virus, bovine parvovirus, chicken parvovirus, feline panleukopenia virus, feline parvovirus, goose parvovirus, H1 parvovirus, Muscovy duck parvovirus, snake parvovirus, and B19 virus. Other autonomous parvoviruses are known to those skilled in the art. See, for example, FIELDS et al. VIROLOGY, Volume 2, Chapter 69 (4th Edition, Lippincott-Raven Publishers).

[0240] As used herein, the term "non-AAV" encompasses nucleic acids, proteins, and viruses from the Parvoviridae family, excluding any adeno-associated virus (AAV) of the Parvoviridae family. "Non-AAV" includes, but is not limited to, autonomously replicating members of the genera Bocavirus, Dependovirus, Erythrovirus, Aleutianvirus, Parvovirus, Densovirus, Repetitivevirus, Contravirus, Avian Parvovirus, Ruminant Parvovirus, Proparvovirus, Quadrangular Parvovirus, Ambisense Densovirus, Brevovirus, Hepatopancreatic Densovirus, and Prawn Densovirus.

[0241] As used herein, the term "adeno-associated virus" (AAV) includes, but is not limited to, AAV type 1, AAV type 2, AAV type 3 (including types 3A and 3B), AAV type 4, AAV type 5, AAV type 6, AAV type 7, AAV type 8, AAV type 9, AAV type 10, AAV type 11, AAV type 12, AAV type 13, snake AAV, avian AAV, bovine AAV, canine AAV, equine AAV, ovine AAV, caprine AAV, shrimp AAV, those AAV serotypes and clades disclosed by Gao et al. (J. Virol. 78:6381 (2004)) and Moris et al. (Virol. 33:375 (2004)), and any other AAV currently known or discovered in the future. See, for example, FIELDS et al. VIROLOGY, Vol. 2, Chapter 69 (4th ed., Lippincott-Raven Publishers).

[0242] As used herein, the term "derived from" refers to a component that is separated from a specified molecule or organism or prepared using a specified molecule or organism, or information (e.g., amino acid or nucleic acid sequence) is from a specified molecule or organism. For example, a nucleic acid sequence (e.g., ITR) derived from a second nucleic acid sequence (e.g., ITR) may include a nucleotide sequence that is identical or substantially similar to the nucleotide sequence of the second nucleic acid sequence. In the case of nucleotides or polypeptides, the derived material can be obtained by, for example, naturally occurring mutagenesis, artificial directed mutagenesis, or artificial random mutagenesis. The mutagenesis used to derive nucleotides or polypeptides can be intentionally directed or intentionally random, or a mixture of each. Mutagenesis of nucleotides or polypeptides to produce different nucleotides or polypeptides derived from the first nucleotide or polypeptide can be a random event (e.g., caused by polymerase distortion), and the identification of the derived nucleotides or polypeptides can be performed by appropriate screening methods (e.g., as discussed herein). Mutagenesis of polypeptides generally requires manipulation of the polynucleotide encoding the polypeptide. In some embodiments, a nucleotide or amino acid sequence derived from a second nucleotide or amino acid sequence has at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, In some embodiments, the ITR is at least 90% identical to the non-AAV ITR (or AAV ITR), wherein the non-AAV ITR retains the functional properties of the non-AAV ITR (or AAV ITR). In some embodiments, the ITR derived from a non-AAV (or AAV) ITR is at least 80% identical to the non-AAV ITR (or, respectively, AAV ITR), wherein the non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or, respectively, AAV ITR).In some embodiments, the ITR derived from a non-AAV (or AAV) ITR is at least 70% identical to the non-AAV ITR (or, respectively, AAV ITR), wherein the non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or, respectively, AAV ITR). In some embodiments, the ITR derived from a non-AAV (or AAV) ITR is at least 60% identical to the non-AAV ITR (or, respectively, AAV ITR), wherein the non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or, respectively, AAV ITR). In some embodiments, the ITR derived from a non-AAV (or AAV) ITR is at least 50% identical to the non-AAV ITR (or, respectively, AAV ITR), wherein the non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or, respectively, AAV ITR).

[0243] In certain embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR. In some embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR, wherein the fragment comprises at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about nucleotides, at least about 100 nucleotides, at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides, or at least about 600 nucleotides; wherein the ITR derived from a non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or, respectively, AAV ITR). In certain embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR, wherein the fragment comprises at least about 129 nucleotides, and wherein the ITR derived from a non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or, respectively, AAV ITR). In certain embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR, wherein the fragment comprises at least about 102 nucleotides, and wherein the ITR derived from a non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or, respectively, AAV ITR).

[0244] In some embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR, wherein the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the length of the non-AAV (or AAV) ITR.

[0245] In certain embodiments, a nucleotide or amino acid sequence derived from a second nucleotide or amino acid sequence, when correctly aligned, has at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, In some embodiments, the ITRs of the non-AAV (or AAV) ITRs are at least 90% identical to the homologous portions of the non-AAV ITRs (or, respectively, AAV ITRs) when correctly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, when correctly aligned, the ITR derived from a non-AAV (or AAV) ITR is at least 80% identical to the homologous portion of the non-AAV ITR (or, respectively, AAV ITR), wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, when correctly aligned, the ITR derived from a non-AAV (or AAV) ITR is at least 70% identical to the homologous portion of the non-AAV ITR (or, respectively, AAV ITR), wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, when correctly aligned, the ITR derived from a non-AAV (or AAV) ITR is at least 60% identical to the homologous portion of the non-AAV ITR (or, respectively, AAV ITR), wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, an ITR derived from a non-AAV (or AAV) ITR is at least 50% identical to a homologous portion of a non-AAV ITR (or, respectively, an AAV ITR) when correctly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence.

[0246] A "capsid-free" or "capsid-less" vector or nucleic acid molecule refers to a vector construct that does not have a capsid. In some embodiments, the capsid-free vector or nucleic acid molecule does not contain a sequence encoding, for example, an AAV Rep protein.

[0247] As used herein, "coding region" or "coding sequence" is a portion of a polynucleotide that is composed of codons that are translatable into amino acids. Although "stop codons" (TAG, TGA, or TAA) are not typically translated into amino acids, they are considered to be part of the coding region, but any flanking sequences (e.g., promoters, ribosome binding sites, transcription terminators, introns, etc.) are not part of the coding region. The boundaries of the coding region are typically determined by the start codon at the 5' end (encoding the amino terminus of the resulting polypeptide) and the translation stop codon at the 3' end (encoding the carboxyl terminus of the resulting polypeptide). Two or more coding regions may be present in a single polynucleotide construct (e.g., on a single vector), or in separate polynucleotide constructs (e.g., on separate (different) vectors). The result is then that a single vector may contain only a single coding region, or may comprise two or more coding regions.

[0248] Certain proteins secreted by mammalian cells are associated with a secretory signal peptide that is cleaved from the mature protein once export of the growing protein chain across the rough endoplasmic reticulum has been initiated. Those skilled in the art will appreciate that signal peptides are typically fused to the N-terminus of a polypeptide and are cleaved from the intact or "full-length" polypeptide to produce the secreted or "mature" form of the polypeptide. In certain embodiments, a native signal peptide or a functional derivative of such a sequence that retains the ability to direct the secretion of a polypeptide is operably associated with the polypeptide. Alternatively, a heterologous mammalian signal peptide (e.g., human tissue plasminogen activator (TPA) or mouse β-glucuronidase signal peptide) or a functional derivative thereof may be used.

[0249] The term "downstream" refers to a nucleotide sequence located 3' of a reference nucleotide sequence. In certain embodiments, a downstream nucleotide sequence refers to a sequence following the start of transcription. For example, the translation start codon of a gene is located downstream of the transcription start site.

[0250] The term "upstream" refers to a nucleotide sequence located 5' of a reference nucleotide sequence. In certain embodiments, an upstream nucleotide sequence refers to a sequence located 5' of a coding region or transcription start site. For example, most promoters are located upstream of the transcription start site.

[0251] As used herein, the term "gene regulatory region" or "regulatory region" refers to a nucleotide sequence that is located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding region and that affects the transcription, RNA processing, stability, or translation of the associated coding region. A regulatory region can include a promoter, a translation leader sequence, introns, a polyadenylation recognition sequence, an RNA processing site, an effector binding site, or a stem-loop structure. If the coding region is intended to be expressed in eukaryotic cells, a polyadenylation signal and transcription termination sequence will typically be located 3' to the coding sequence.

[0252] A polynucleotide encoding a product (e.g., a miRNA or a gene product (e.g., a polypeptide, such as a therapeutic protein)) can include a promoter and / or other expression (e.g., transcription or translation) control components operably associated with one or more coding regions. In operably associated, the coding region of a gene product (e.g., a polypeptide) is associated with one or more regulatory regions in a manner that places the expression of the gene product under the influence or control of the regulatory regions. For example, a coding region and a promoter are "operably associated" if induction of promoter function results in transcription of an mRNA encoding the gene product encoded by the coding region, and if the nature of the connection between the promoter and the coding region does not interfere with the ability of the promoter to direct expression of the gene product or the ability of the DNA template to be transcribed. Other expression control components besides promoters (e.g., enhancers, operators, repressors, and transcription termination signals) can also be operably associated with a coding region to direct expression of a gene product.

[0253] "Transcription control sequence" refers to a DNA regulatory sequence that provides for the expression of a coding sequence in a host cell, such as a promoter, enhancer, terminator, and the like. A variety of transcription control regions are known to those skilled in the art. These include, but are not limited to, those that act in vertebrate cells, such as, but not limited to, promoter and enhancer segments from cytomegalovirus (immediate early promoter, associated with intron-A), simian virus 40 (early promoter), and retroviruses (such as Rous sarcoma virus). Other transcription control regions include those derived from vertebrate genes, such as actin, heat shock protein, bovine growth hormone, and rabbit β-globin, as well as other sequences capable of controlling gene expression in eukaryotic cells. Other suitable transcription control regions include tissue-specific promoters and enhancers and lymphokine-inducible promoters (e.g., promoters inducible by interferons or interleukins).

[0254] Similarly, a variety of translation control components are known to those skilled in the art. These include, but are not limited to, ribosome binding sites, translation initiation and termination codons, and components derived from picornaviruses (particularly internal ribosome entry sites, or IRES, also known as CITE sequences).

[0255] As used herein, the term "expression" refers to the process by which a polynucleotide produces a gene product (e.g., RNA or polypeptide). It includes but is not limited to the transcription of a polynucleotide into messenger RNA (mRNA), transfer RNA (tRNA), small hairpin RNA (shRNA), small interfering RNA (siRNA) or any other RNA product, and the translation of mRNA into a polypeptide. Expression produces a "gene product." As used herein, a gene product can be a nucleic acid (e.g., messenger RNA produced by gene transcription) or a polypeptide translated from a transcript. Gene products as described herein also include nucleic acids with post-transcriptional modifications (e.g., polyadenylation or splicing), or polypeptides with post-translational modifications (e.g., methylation, glycosylation, lipid addition, association with other protein subunits, or proteolytic cleavage). As used herein, the term "yield" refers to the amount of a polypeptide produced by gene expression.

[0256] "Vector" refers to any medium for cloning and / or transferring nucleic acids into a host cell. A vector can be a replicon to which another nucleic acid segment can be connected to achieve replication of the connected segment. "Replicon" refers to any genetic component (e.g., plasmid, bacteriophage, cosmid, chromosome, virus) that acts as an autonomous replicating unit in vivo (i.e., capable of replicating under its own control). The term "vector" includes a medium for introducing nucleic acids into cells in vitro, in vitro, or in vivo. A large number of vectors are known and used in the art, including, for example, plasmids, modified eukaryotic viruses, or modified bacterial viruses. Inserting a polynucleotide into a suitable vector can be accomplished by connecting the appropriate polynucleotide fragments to a selected vector with complementary sticky ends.

[0257] The vector can be engineered to encode a selectable marker or reporter that provides selection or identification of cells incorporated into the vector. The expression of the selectable marker or reporter allows identification and / or selection of host cells incorporated into and expressing other coding regions contained in the vector. Examples of selectable marker genes known in the art and used include: genes that provide resistance to ampicillin, streptomycin, gentamicin, kanamycin, hygromycin, bialaphos herbicides, sulfonamides, etc.; and genes used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentanyl transferase genes, etc. The example of reporters known in the art and used include: luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), β-galactosidase (LacZ), β-glucuronidase (Gus), etc. Selectable markers can also be considered as reporters.

[0258] As used herein, the term "host cell" refers to, for example, microorganisms, yeast cells, insect cells, and mammalian cells that can be or have been used as recipients of ssDNA or vectors. The term includes the progeny of an initial cell that has been transduced. Thus, as used herein, "host cell" generally refers to a cell that has been transduced with an exogenous DNA sequence. It should be understood that the morphology or genomic or total DNA complement of the progeny of a single parent cell may not necessarily be identical to that of the initial parent due to natural, accidental, or deliberate mutations. In some embodiments, the host cell may be an in vitro host cell.

[0259] The term "selectable marker" refers to an identification factor (typically an antibiotic or chemical resistance gene) that can be selected based on the effect of the marker gene (i.e., resistance to antibiotics, resistance to herbicides, colorimetric markers, enzymes, fluorescent markers, etc.), wherein the effect is used to track the inheritance of a nucleic acid of interest and / or to identify cells or organisms that have inherited the nucleic acid of interest. Examples of selectable marker genes known and used in the art include: genes that provide resistance to ampicillin, streptomycin, gentamicin, kanamycin, hygromycin, bialaphos herbicides, sulfonamides, etc.; and genes used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentanyl transferase genes, etc.

[0260] The term "reporter gene" refers to a nucleic acid encoding an identification factor that can be identified based on the effect of a reporter gene, wherein the effect is used to track the heritability of a nucleic acid of interest, identify cells or organisms that have inherited a nucleic acid of interest, and / or measure gene expression induction or transcription. Examples of reporter genes known in the art and used include: luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), beta-galactosidase (LacZ), beta-glucuronidase (Gus), etc. Selectable marker genes can also be considered as reporter genes.

[0261] "Promoter" and "promoter sequence" are used interchangeably and refer to a DNA sequence that can control the expression of a coding sequence or functional RNA. Typically, the coding sequence is located 3' to the promoter sequence. A promoter can be derived in its entirety from a natural gene, or be composed of different components derived from different promoters found in nature, or even comprise synthetic DNA segments. Those skilled in the art will appreciate that different promoters can direct gene expression in different tissues or cell types, or at different developmental stages, or in response to different environmental or physiological conditions. Promoters that cause a gene to be expressed in most cell types at most times are generally referred to as "constitutive promoters." Promoters that cause a gene to be expressed in a specific cell type are generally referred to as "cell-specific promoters" or "tissue-specific promoters." Promoters that cause a gene to be expressed at a specific stage of development or cell differentiation are generally referred to as "development-specific promoters" or "cell differentiation-specific promoters." Promoters that are induced and cause gene expression after exposure or treatment of cells with promoter-inducing agents, biomolecules, chemicals, ligands, light, etc. are generally referred to as "inducible promoters" or "regulatable promoters." It is also recognized that because in most cases the exact boundaries of regulatory sequences have not been completely defined, DNA fragments of different lengths can have identical promoter activity.

[0262] The promoter sequence is usually bounded at its 3' terminus by the transcription start site and extends upstream (5' direction) to include the minimum number of bases or modules necessary to initiate transcription at a level detectable above background. Within the promoter sequence will be found the transcription start site (conveniently defined, for example, by mapping with nuclease S1) and protein binding domains (consensus sequences) responsible for the binding of RNA polymerase.

[0263] In certain embodiments, nucleic acid molecules include tissue-specific promoters. In certain embodiments, tissue-specific promoters drive the expression of therapeutic proteins (e.g., coagulation factors) in the liver (e.g., in hepatocytes and / or endothelial cells). In specific embodiments, promoters are selected from mouse thyroxine promoter (mTTR), endogenous human factor VIII promoter (F8), human alpha-1-antitrypsin promoter (hAAT), human albumin minimal promoter, mouse albumin promoter, triple tetraproline (TTP) promoter, CASI promoter, CAG promoter, cytomegalovirus (CMV) promoter, phosphoglycerate kinase (PGK) promoter, and any combination thereof. In some embodiments, the promoter is selected from a liver-specific promoter (e.g., α1-antitrypsin (AAT)), a muscle-specific promoter (e.g., muscle creatine kinase (MCK), myosin heavy chain α (αMHC), myoglobin (MB), and desmin (DES)), a synthetic promoter (e.g., SPc5-12, 2R5Sc5-12, dMCK, and tMCK), and any combination thereof. In a specific embodiment, the promoter comprises a TTP promoter.

[0264] The term "restriction endonuclease" is used interchangeably with "restriction enzyme" and refers to an enzyme that binds and cuts within a specific nucleotide sequence within double-stranded DNA.

[0265] The term "plasmid" refers to an extrachromosomal component that usually carries genes that are not part of the central metabolism of the cell and is usually in the form of a circular double-stranded DNA molecule. Such components can be linear, circular, or supercoiled autonomously replicating sequences, genome-integrating sequences, phage, or nucleotide sequences derived from single-stranded or double-stranded DNA or RNA from any source, in which multiple nucleotide sequences have been joined or recombined into a unique construct capable of introducing a promoter fragment and DNA sequence for a selected gene product, as well as appropriate 3' non-translated sequences, into the cell.

[0266] Eukaryotic viral vectors that can be used include, but are not limited to, adenoviral vectors, retroviral vectors, adeno-associated viral vectors, poxvirus (e.g., vaccinia virus vectors), baculovirus vectors, or herpesvirus vectors. Non-viral vectors include plasmids, liposomes, charged lipids (cytofectins), DNA-protein complexes, and biopolymers.

[0267] "Cloning vector" refers to a "replicon," which is a unit length of continuously replicating nucleic acid and which contains an origin of replication, such as a plasmid, phage, or cosmid, to which another nucleic acid segment can be attached to achieve replication of the attached segment. Certain cloning vectors are capable of replicating in one cell type (e.g., bacteria) and expressing in another cell type (e.g., eukaryotic cells). Cloning vectors typically contain one or more sequences that can be used to select cells containing the vector and / or one or more multiple cloning sites for inserting nucleic acid sequences of interest.

[0268] The term "expression vector" refers to a vehicle designed to enable expression of an inserted nucleic acid sequence after insertion into a host cell. The inserted nucleic acid sequence is placed in operable association with regulatory regions as described above.

[0269] The vector is introduced into the host cell by methods well known in the art, such as transfection, electroporation, microinjection, transduction, cell fusion, DEAE dextran, calcium phosphate precipitation, lipofection (lysosomal fusion), use of a gene gun, or a DNA vector transporter. As used herein, "culture," "to culture," and "culturing" mean incubating cells or maintaining cells alive under in vitro conditions that allow cell growth or division. As used herein, "cultured cells" means cells propagated in vitro.

[0270] As used herein, the term "polypeptide" is intended to encompass the singular "polypeptide" as well as the plural "polypeptides" and refers to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also known as peptide bonds). The term "polypeptide" refers to any one or more chains of two or more amino acids and does not refer to the specific length of the product. Thus, peptides, dipeptides, tripeptides, oligopeptides, "proteins," "amino acid chains," or any other terms used to refer to one or more chains of two or more amino acids are included within the definition of "polypeptide," and the term "polypeptide" can replace any of these terms or be used interchangeably therewith. The term "polypeptide" is also intended to refer to products of post-expression modifications of the polypeptide, including but not limited to glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting / blocking groups, proteolytic cleavage, or modification by non-naturally occurring amino acids. Polypeptides can be derived from natural biological sources or produced by recombinant technology, but are not necessarily translated from a specified nucleic acid sequence. They can be produced in any manner, including by chemical synthesis.

[0271] The term "amino acid" includes alanine (Ala or A); arginine (Arg or R); asparagine (Asn or N); aspartic acid (Asp or D); cysteine ​​(Cys or C); glutamine (Gln or Q); glutamic acid (Glu or E); glycine (Gly or G); histidine (His or H); isoleucine (Ile or I); leucine (Leu or L); lysine (Lys or K); methionine (Met or M); phenylalanine (Phe or F); proline (Pro or P); serine (Ser or S); threonine (Thr or T); tryptophan (Trp or W); tyrosine (Tyr or Y); and valine (Val or V). Non-traditional amino acids are also within the scope of the present disclosure and include norleucine, ornithine, norvaline, homoserine and other amino acid residue analogs, such as those described in Ellman et al. Meth. Enzym. 202: 301-336 (1991). To generate such non-naturally occurring amino acid residues, the procedures of Noren et al. Science 244: 182 (1989) and Ellman et al. (supra) can be used. Briefly, these procedures involve chemical activation of inhibitor tRNA with non-naturally occurring amino acid residues, followed by in vitro transcription and translation of RNA. The introduction of non-traditional amino acids can also be achieved using peptide chemistry known in the art. As used herein, the term "polar amino acid" includes amino acids having zero net charge but having non-zero partial charges at different parts of their side chains (e.g., M, F, W, S, Y, N, Q, C). These amino acids can participate in hydrophobic interactions and electrostatic interactions. As used herein, the term "charged amino acid" includes amino acids with a non-zero net charge on their side chains (e.g., R, K, H, E, D). These amino acids can participate in hydrophobic interactions and electrostatic interactions.

[0272] The present disclosure also includes fragments or variants of polypeptides and any combination thereof. The term "fragment" or "variant" when referring to a polypeptide binding domain or binding molecule of the present disclosure includes any polypeptide that retains at least some of the properties of a reference polypeptide (e.g., FcRn binding affinity to an FcRn binding domain or Fc variant, coagulation activity to a FVIII variant, or FVIII binding activity to a VWF fragment). In addition to the specific antibody fragments discussed elsewhere herein, polypeptide fragments include proteolytic fragments and deletion fragments, but do not include naturally occurring full-length polypeptides (or mature polypeptides). Variants of the polypeptide binding domain or binding molecule of the present disclosure include fragments as described above, as well as polypeptides having altered amino acid sequences due to amino acid substitutions, deletions, or insertions. Variants may be naturally occurring or non-naturally occurring. Non-naturally occurring variants can be produced using mutagenesis techniques known in the art. Variant polypeptides may contain conservative or non-conservative amino acid substitutions, deletions, or additions.

[0273] " conservative amino acid substitution " is the substitution of amino acid residues with amino acid residues having similar side chains. The industry has defined families of amino acid residues with similar side chains, including basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), non-polar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), β-branched side chains (e.g., threonine, valine, isoleucine) and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Therefore, if the amino acid in the polypeptide is substituted with another amino acid from the same side chain family, the substitution is considered conservative. In another embodiment, an amino acid string can be conservatively substituted with a string that is structurally similar but different in order and / or composition of side chain family members.

[0274] As known in the art, the term "percent identity" refers to the relationship between two or more polypeptide sequences or two or more polynucleotide sequences, as determined by comparing the sequences. In the art, "identity" also means the degree of sequence relatedness between polypeptide or polynucleotide sequences, as determined by the match between strings of such sequences, as the case may be. "Identity" can be readily calculated by known methods, including but not limited to those described in Computational Molecular Biology (Lesk, AM, ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, DW, ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, AM and Griffin, HG, eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Stockton Press, New York (1991). Optimal methods for determining identity are designed to give the best match between the sequences tested. Methods for determining identity are codified in publicly available computer programs. Sequence alignments and percent identity calculations can be performed using sequence analysis software, such as the Megalign program of the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, Wisconsin), the GCG program suite (Wisconsin Package Version 9.0, Genetics Computer Group (GCG), Madison, Wisconsin), BLASTP, BLASTN, BLASTX (Altschul et al., J. Mol. Biol. 215:403 (1990)), and DNASTAR (DNASTAR, Inc. 1228 S. Park St. Madison, Wisconsin 53715 USA). In the context of this application, it will be understood that if sequence analysis software is used for analysis, unless otherwise specified, the results of the analysis will be based on the "default values" of the referenced program. "Default values," as used herein, will mean any set of values ​​or parameters that are initially loaded with the software upon first initialization.For the purpose of determining the percent identity between a therapeutic protein (e.g., coagulation factor) sequence of the present disclosure and a reference sequence, only the nucleotides corresponding to the nucleotides in the therapeutic protein (e.g., coagulation factor) sequence of the present disclosure in the reference sequence are used to calculate the percent identity. For example, when comparing a full-length FVIII nucleotide sequence containing a B domain and an optimized B domain deletion (BDD) FVIII nucleotide sequence of the present disclosure, the comparison portion including A1, A2, A3, C1, and C2 domains is used to calculate the percent identity. Nucleotides in the portion encoding the B domain of the full-length FVIII sequence (which will produce a large " gap " in the comparison) will not be counted as mispairing. Additionally, when determining the percent identity between an optimized BDD FVIII sequence of the present disclosure, or a specified portion thereof (e.g., nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3) and a reference sequence, the percent identity will be calculated by aligning and dividing the number of matching nucleotides by the total number of nucleotides in the entire sequence of the optimized BDD-FVIII sequence, or a specified portion thereof, as described herein.

[0275] As used herein, nucleotides corresponding to nucleotides in a particular sequence of the present disclosure are identified by aligning the sequence of the present disclosure to maximize identity with a reference sequence. The numbering used to identify equivalent amino acids in a reference sequence is based on the numbering used to identify the corresponding amino acid in the sequence of the present disclosure.

[0276] A "fusion" or "chimeric" protein comprises a first amino acid sequence linked to a second amino acid sequence to which the first amino acid sequence is not naturally linked in nature. Amino acid sequences that are normally found in separate proteins can be brought together in a fusion polypeptide, or amino acid sequences that are normally found in the same protein can be placed in a fusion polypeptide in a new arrangement (e.g., a fusion of a Factor VIII domain with an Ig Fc domain of the present disclosure). Fusion proteins are produced, for example, by chemical synthesis or by generating and translating a polynucleotide that encodes the peptide regions in the desired relationship. A chimeric protein can also comprise a second amino acid sequence associated with a first amino acid sequence by a covalent, non-peptide bond or a non-covalent bond.

[0277] As used herein, the term "insertion site" refers to the following position in a polypeptide or a fragment, variant or derivative thereof, which is immediately upstream of the position where a heterologous portion can be inserted. An "insertion site" is designated as a number, which is the number of the amino acid in the reference sequence. For example, an "insertion site" in FVIII refers to the number of the amino acid sequence corresponding to the insertion site in mature native FVIII (SEQ ID NO: 15), which is immediately adjacent to the N-terminus of the insertion position. For example, the phrase "a3 comprises a heterologous portion at the insertion site corresponding to amino acid 1656 of SEQ ID NO: 15" indicates that the heterologous portion is located between two amino acids corresponding to amino acid 1656 and amino acid 1657 of SEQ ID NO: 15.

[0278] As used herein, the phrase "immediately downstream of an amino acid" refers to a position immediately adjacent to the terminal carboxyl group of an amino acid. Similarly, the phrase "immediately upstream of an amino acid" refers to a position immediately adjacent to the terminal amine group of an amino acid.

[0279] As used herein, the terms "inserted," "is inserted," "inserted into," or grammatically related terms refer to the position of a heterologous moiety in a polypeptide (e.g., a coagulation factor) relative to an analogous position in a parent polypeptide. For example, in certain embodiments, "inserted" and the like refer to the position of a heterologous moiety in a recombinant FVIII polypeptide relative to an analogous position in native mature human FVIII. As used herein, the terms refer to characteristics of a polypeptide and do not indicate, suggest, or imply any method or process for preparing the polypeptide.

[0280] As used herein, the term "half-life" refers to the biological half-life of a particular polypeptide in vivo. Half-life can be expressed as the time required for half of the amount administered to a subject to be cleared from the animal's circulation and / or other tissues. When the clearance curve for a given polypeptide is constructed as a function of time, the curve is typically biphasic, with a rapid α phase and a longer β phase. The α phase typically represents the equilibrium between the intravascular and extravascular spaces of the administered Fc polypeptide and depends in part on the size of the polypeptide. The β phase typically represents the catabolism of the polypeptide in the intravascular space. In some embodiments, therapeutic proteins (e.g., coagulation factors, such as FVIII) and chimeric proteins comprising the therapeutic proteins are monophasic and therefore do not have an α phase, but only have a single β phase. Therefore, in certain embodiments, the term half-life as used herein refers to the half-life of a polypeptide in the β phase.

[0281] As used herein, the term "connection" refers to the covalent or non-covalent bonding of a first amino acid sequence or nucleotide sequence to a second amino acid sequence or nucleotide sequence, respectively. The first amino acid or nucleotide sequence can be directly bonded or juxtaposed with the second amino acid or nucleotide sequence, or alternatively, an insertion sequence can covalently bond the first sequence to the second sequence. The term "connection" not only means that the first amino acid sequence and the second amino acid sequence are fused at the C-terminus or N-terminus, but also includes inserting a complete first amino acid sequence (or second amino acid sequence) between any two amino acids in the second amino acid sequence (or respectively, the first amino acid sequence). In one embodiment, the first amino acid sequence can be connected to the second amino acid sequence by a peptide bond or a joint. The first nucleotide sequence can be connected to the second nucleotide sequence by a phosphodiester bond or a joint. The joint can be a peptide or polypeptide (for a polypeptide chain) or a nucleotide or nucleotide chain (for a nucleotide chain) or any chemical moiety (for both polypeptide and polynucleotide chains). The term "connection" is also indicated by a hyphen (-).

[0282] As used herein, hemostasis means stopping or slowing bleeding or hemorrhage; or stopping or slowing the flow of blood through a blood vessel or body part.

[0283] As used herein, hemostatic disorders refer to inherited or acquired disorders characterized by a tendency to bleed spontaneously or due to trauma due to impaired or inability to form fibrin clots. Examples of such disorders include hemophilia. The three main forms are hemophilia A (factor VIII deficiency), hemophilia B (factor IX deficiency or "Kreissmann's disease"), and hemophilia C (factor XI deficiency, mild bleeding tendency). Other hemostatic disorders include, for example, von Willebrand's disease, factor XI deficiency (PTA deficiency), factor XII deficiency, fibrinogen, prothrombin, factor V, factor VII, factor X, or factor XIII deficiency or structural abnormalities, GPIb deficiency, or giant platelet syndrome. The vWF receptor GPIb may be defective and lead to a lack of primary clot formation (primary hemostasis) and increased bleeding tendency, as well as Glanzman and Naegeli's thrombasthenia (Glanzmann thrombasthenia). In liver failure (acute and chronic forms), the liver's coagulation factors are not produced enough; this may increase the risk of bleeding.

[0284] The isolated nucleic acid molecules, isolated polypeptides, or vectors comprising the isolated nucleic acid molecules of the present disclosure can be used prophylactically. As used herein, the term "prophylactic treatment" refers to the administration of a molecule prior to an episode of bleeding. In one embodiment, the subject in need of a general hemostatic agent is undergoing or about to undergo surgery. The polynucleotides, polypeptides, or vectors of the present disclosure can be administered as a prophylactic before or after surgery. The polynucleotides, polypeptides, or vectors of the present disclosure can be administered during or after surgery to control an acute bleeding episode. Surgeries can include, but are not limited to, liver transplantation, hepatectomy, dental surgery, or stem cell transplantation.

[0285] The isolated nucleic acid molecules, isolated polypeptides or vectors of the present disclosure are also used for on-demand treatment. The term "on-demand treatment" refers to administering an isolated nucleic acid molecule, isolated polypeptide or vector in response to the symptoms of a bleeding episode or prior to an activity that may cause bleeding. In one aspect, on-demand treatment can be administered to a subject at the onset of bleeding (e.g., after an injury) or in anticipation of bleeding (e.g., before surgery). In another aspect, on-demand treatment can be administered prior to an activity that increases the risk of bleeding (e.g., contact sports).

[0286] As used herein, the term "acute hemorrhage" refers to an episode of bleeding regardless of the underlying cause. For example, the subject may have trauma, uremia, an inherited bleeding disorder (e.g., Factor VII deficiency), a platelet disorder, or resistance due to the development of antibodies to a coagulation factor.

[0287] As used herein, treat, treatment, and treating refer to, for example, a decrease in the severity of a disease or condition; a decrease in the duration of the disease course; an improvement in one or more symptoms associated with the disease or condition; providing a beneficial effect to a subject having the disease or condition, but not necessarily curing the disease or condition; or preventing one or more symptoms associated with the disease or condition. In one embodiment, the term "treating" or "treatment" means maintaining, for example, a FVIII trough level of at least about 1 IU / dL, 2 IU / dL, 3 IU / dL, 4 IU / dL, 5 IU / dL, 6 IU / dL, 7 IU / dL, 8 IU / dL, 9 IU / dL, 10 IU / dL, 11 IU / dL, 12 IU / dL, 13 IU / dL, 14 IU / dL, 15 IU / dL, 16 IU / dL, 17 IU / dL, 18 IU / dL, 19 IU / dL, etc., in a subject by administering an isolated nucleic acid molecule, isolated polypeptide, or vector of the present disclosure. dL, 20IU / dL, 25IU / dL, 30IU / dL, 35IU / dL, 40IU / dL, 45IU / dL, 50IU / dL, 55IU / dL, 60IU / dL, 65IU / dL, 70IU / dL, 75IU / dL, 80IU / dL, 85IU / dL, 90IU / dL, 95IU / dL, 100IU / dL, 105IU / dL, 110IU / dL, 115IU / dL, 120IU / dL, 125IU / dL, 130IU / dL, 135IU / dL, 140IU / dL, 145IU / dL or 150IU / dL.In another embodiment, treating or treatment means maintaining a FVIII trough level of about 1 to about 150 IU / dL, about 1 to about 125 IU / dL, about 1 to about 100 IU / dL, about 1 to about 90 IU / dL, about 1 to about 85 IU / dL, about 1 to about 80 IU / dL, about 1 to about 75 IU / dL, about 1 to about 70 IU / dL, about 1 to about 65 IU / dL, about 1 to about 60 IU / dL, about 1 to about 55 IU / dL, about 1 to about 50 IU / dL, about 1 to about 45 IU / dL, about 1 to about 40 IU / dL / dL, about 1-about 35 IU / dL, about 1-about 30 IU / dL, about 1-about 25 IU / dL, about 25-about 125 IU / dL, about 50-about 100 IU / dL, about 50-about 75 IU / dL, about 75-about 100 IU / dL, about 1-about 20 IU / dL, about 2-about 20 IU / dL, about 3-about 20 IU / dL, about 4-about 20 IU / dL, about 5-about 20 IU / dL, about 6-about 20 IU / dL, about 7-about 20 IU / dL, about 8-about 20 IU / dL, about 9-about 20 IU / dL, or about 10-about 20 IU / dL. Treatment or treating of a disease or condition can also include maintaining FVIII activity in a subject at a level that is equivalent to at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 105%, 110%, 115%, 120%, 125%, 130%, 135%, 140%, 145%, or 150% of the FVIII activity in a non-hemophilic subject. The trough level required for treatment can be measured by one or more known methods and can be adjusted (increased or decreased) for each individual.

[0288] As used herein, "administering" means administering a pharmaceutically acceptable nucleic acid molecule of the present disclosure, a polypeptide expressed from the nucleic acid molecule, or a vector comprising the nucleic acid molecule to a subject via a pharmaceutically acceptable route. The route of administration can be intravenous, such as intravenous injection and intravenous infusion. Other routes of administration include, for example, subcutaneous, intramuscular, oral, nasal, and pulmonary administration. Nucleic acid molecules, polypeptides, and vectors can be administered as part of a pharmaceutical composition comprising at least one excipient.

[0289] As used herein, the term "pharmaceutically acceptable" refers to molecular entities and compositions that are physiologically tolerable and do not typically produce toxic or allergic reactions or similar adverse reactions (such as stomach upset, dizziness, etc.) when administered to humans. Optionally, as used herein, the term "pharmaceutically acceptable" means approved by a regulatory agency of the Federal or a state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeia for use in animals, and more particularly in humans.

[0290] As used herein, the phrase "subject in need thereof" includes subjects (e.g., mammalian subjects) who would benefit from the administration of a nucleic acid molecule, polypeptide, or vector of the present disclosure, e.g., to improve hemostasis. In one embodiment, subjects include, but are not limited to, individuals with hemophilia. In another embodiment, subjects include, but are not limited to, individuals who have developed inhibitors against a therapeutic protein (e.g., a coagulation factor, such as FVIII) and therefore require a bypass therapy. The subject may be an adult or a minor (e.g., less than 12 years of age).

[0291] As used herein, the term "therapeutic protein" refers to any polypeptide known in the art that can be administered to a subject. In some embodiments, the therapeutic protein comprises a protein selected from the group consisting of a coagulation factor, a growth factor, an antibody, a functional fragment thereof, or a combination thereof. As used herein, the term "coagulation factor" refers to a naturally occurring or recombinantly produced molecule or its analog that prevents bleeding episodes in a subject or shortens the duration of a bleeding episode. In other words, it means a molecule with procoagulant activity (i.e., responsible for converting fibrinogen into a network of insoluble fibrin, thereby causing blood coagulation or coagulation). As used herein, "coagulation factor" includes activated coagulation factors, their zymogens, or activatable coagulation factors. "Activatable coagulation factor" is a coagulation factor in an inactive form (e.g., in its zymogen form) that can be converted into an active form. The term "coagulation factor" includes, but is not limited to, factor I (FI), factor II (FII), factor V (FV), FVII, FVIII, FIX, factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FXIII), von Willebrand factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, α2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), zymogens thereof, activated forms thereof, or any combination thereof.

[0292] Coagulant activity as used herein means the ability to participate in the cascade of biochemical reactions that culminates in the formation of a fibrin clot and / or reduces the severity, duration or frequency of a bleeding disorder or bleeding episodes.

[0293] As used herein, "growth factors" include any growth factor known in the art, including cytokines and hormones. In some embodiments, the growth factor is selected from adrenomedullin (AM), angiogenin (Ang), autotaxin, bone morphogenetic protein (BMP) (e.g., BMP2, BMP4, BMP5, BMP7), ciliary neurotrophic factor family members (e.g., ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6)), colony stimulating factors (e.g., macrophage colony stimulating factor (m-CSF), granulocyte colony stimulating factor (G-CSF), granulocyte macrophage colony stimulating factor (GM-CSF)), epidermal growth factor (EGF), ephrins (e.g., ephrin A1, ephrin B2), ephrin C1, ephrin C2, ephrin C3, ephrin D1, ephrin E2, ephrin F3, ephrin F4, ephrin F5, ephrin B7), ephrin F6, ephrin F7, ephrin F7, ephrin F8, ephrin F9, ephrin F1, ephrin F2, ephrin F3, ephrin F4, ephrin F5, ephrin F7), ephrin F3, ephrin F2, ephrin F3, ephrin F4, ephrin F6, ephrin F7, ephrin F8, ephrin F1, ephrin F2, ephrin F3, ephrin F3, ephrin F4, ephrin F1, ephrin F2, ephrin F3, ephrin F3 ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1, ephrin B2, ephrin B3), erythropoietin (EPO), fibroblast growth factor (FGF) (e.g., FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, FGF23), fetal bovine somatotropin (FBS), GDNF family members (e.g., glial cell line-derived neurotrophic factor (GDNF), neurotrophic factor, proteolytic protein, atremin protein), growth differentiation factor-9 (GDF9), hepatocyte growth factor (HGF), hepatoma-derived growth factor (HDGF), insulin, insulin-like growth factor (e.g., insulin-like growth factor-1 (IGF-1) or IGF-2, interleukins (IL) (e.g., IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7), keratinocyte growth factor (KGF), migration stimulating factor (MSF), macrophage stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)), myostatin Protein (GDF-8), neuregulins (e.g., neuregulin 1 (NRG1), NRG2, NRG3, NRG4), neurotrophins (e.g., brain-derived neurotrophic factor (BDNF), nerve growth factor (NGF), neurotrophin-3 (NT-3), NT-4, placental growth factor (PGF), platelet-derived growth factor (PDGF), renalase (RNLS), T-cell growth factor (TCGF), thrombopoietin (TPO), transforming growth factors (e.g., transforming growth factor alpha (TGF-α), TGF-β, tumor necrosis factor-α (TNF-α), and vascular endothelial growth factor (VEGF).

[0294] In some embodiments, the therapeutic protein is encoded by a gene selected from the group consisting of X-linked dystrophin, MTM1 (myotubulin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (mitochondrial ataxia), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof.

[0295] As used herein, the term "heterologous" or "exogenous" refers to a molecule not normally found in a given context (e.g., in a cell or in a polypeptide). For example, an exogenous or heterologous molecule can be introduced into a cell and only be present after manipulation of the cell, for example, by transfection or other forms of genetic engineering, or a heterologous amino acid sequence can be present in a protein in which it is not naturally present.

[0296] As used herein, the term "heterologous nucleotide sequence" refers to a nucleotide sequence that does not naturally occur with a given polynucleotide sequence. In one embodiment, the heterologous nucleotide sequence encodes a polypeptide that can prolong the half-life of a therapeutic protein (e.g., a coagulation factor, such as FVIII). In another embodiment, the heterologous nucleotide sequence encodes a polypeptide that increases the hydrodynamic radius of a therapeutic protein (e.g., a coagulation factor, such as FVIII). In other embodiments, the heterologous nucleotide sequence encodes a polypeptide that improves one or more pharmacokinetic properties of a therapeutic protein but does not significantly affect its biological activity or function (e.g., procoagulant activity). In some embodiments, the therapeutic protein is connected or linked to the polypeptide encoded by the heterologous nucleotide sequence via a linker. Non-limiting examples of polypeptide portions encoded by heterologous nucleotide sequences include, inter alia, immunoglobulin constant regions or portions thereof, albumin or fragments thereof, albumin binding moieties, transferrin, the PAS polypeptide of U.S. Patent Application No. 20100292130, HAP sequences, transferrin or fragments thereof, the C-terminal peptide (CTP) of the beta subunit of human chorionic gonadotropin, albumin binding small molecules, XTEN sequences, FcRn binding moieties (e.g., complete Fc regions or portions thereof that bind to FcRn), single chain Fc regions (ScFc regions, e.g., as described in US 2008 / 0260738, WO 2008 / 012543 or WO 2008 / 1439545), polyglycine linkers, polyserine linkers, peptides and short polypeptides of 6-40 amino acids (having a degree of secondary structure varying from less than 50% to greater than 50%) of two types of amino acids selected from glycine (G), alanine (A), serine (S), threonine (T), glutamate (E) and proline (P), or two or more combinations thereof. In some embodiments, the polypeptide encoded by the heterologous nucleotide sequence is linked to a non-polypeptide moiety. Non-limiting examples of non-polypeptide moieties include polyethylene glycol (PEG), albumin binding small molecules, polysialic acid, hydroxyethyl starch (HES), derivatives thereof, or any combination thereof.

[0297] As used herein, the term "Fc region" is defined as the portion of a polypeptide that corresponds to the Fc region of a native Ig, i.e., as formed by the dimeric association of the corresponding Fc domains of its two heavy chains. A native Fc region forms a homodimer with another Fc region. In contrast, the term "genetically fused Fc region" or "single-chain Fc region" (scFc region), as used herein, refers to a synthetic dimeric Fc region that includes the Fc domains genetically linked within a single polypeptide chain (i.e., encoded in a single contiguous gene sequence).

[0298] In one embodiment, the "Fc region" refers to the portion of a single Ig heavy chain that begins at the hinge region immediately upstream of the papain cleavage site (i.e., residue 216 in IgG, with the first residue of the heavy chain constant region being 114) and ends at the C-terminus of the antibody. Thus, a complete Fc domain comprises at least the hinge domain, the CH2 domain, and the CH3 domain.

[0299] Depending on the Ig isotype, the Fc region of the Ig constant region can include CH2, CH3 and CH4 domains and a hinge region. Chimeric proteins comprising the Fc region of Ig confer several desirable properties on the chimeric protein, including increased stability, increased serum half-life (see Capon et al., 1989, Nature 337:525) and binding to Fc receptors such as neonatal Fc receptors (FcRn) (U.S. Patent Nos. 6,086,875, 6,485,726, 6,030,613; WO 03 / 077834; US2003-0235536A1), which are incorporated herein by reference in their entirety.

[0300] A "reference nucleotide sequence," when used herein as a comparison to a nucleotide sequence of the present disclosure, is a polynucleotide sequence that is substantially identical to a nucleotide sequence of the present disclosure, but the reference sequence has not been optimized. For example, a reference nucleotide sequence of a nucleic acid molecule consisting of the codon-optimized BDD FVIII of SEQ ID NO: 1 and a heterologous nucleotide sequence encoding a single-chain Fc region linked at its 3' end to SEQ ID NO: 1 is a nucleic acid molecule consisting of the original (or "parent") BDD FVIII of SEQ ID NO: 16 and the same heterologous nucleotide sequence encoding a single-chain Fc region linked at its 3' end to SEQ ID NO: 16.

[0301] As used herein, the term "optimization" with respect to a nucleotide sequence refers to a polynucleotide sequence encoding a polypeptide, wherein the polynucleotide sequence has been mutated to enhance the properties of the polynucleotide sequence. In some embodiments, optimization is performed to increase transcription levels, increase translation levels, increase steady-state mRNA levels, increase or decrease the binding of regulatory proteins (such as universal transcription factors), increase or decrease splicing, or increase the yield of the polypeptide produced by the polynucleotide sequence. Examples of changes that can be performed to optimize a polynucleotide sequence include codon optimization, G / C content optimization, removal of repetitive sequences, removal of AT-rich components, removal of hidden splice sites, removal of cis-acting components that repress transcription or translation, addition or removal of poly-T or poly-A sequences, addition of sequences that enhance transcription (such as Kozak consensus sequences) around the transcription start site, removal of sequences that can form stem-loop structures, removal of destabilizing sequences, and two or more combinations thereof.

[0302] II. Nucleic Acid Molecules

[0303] The present disclosure relates to a plasmid-like capsid-free nucleic acid molecule encoding a target sequence, wherein the target sequence encodes a therapeutic protein or a gene that can regulate the expression of a target protein (e.g., miRNA). Capsid is the protein shell of a virus that encapsulates the genetic material of the virus. Known capsids assist the function of virions by protecting the viral genome, delivering the genome to the host, and interacting with the host. However, especially when used in gene therapy, viral capsids can be factors that limit the packaging capacity of the vector and / or induce an immune response.

[0304] AAV vectors have emerged as a more common type of gene therapy vector. However, the presence of the capsid limits the utility of AAV vectors in gene therapy. Specifically, the capsid itself can limit the size of the transgene contained in the vector to less than 4.5 kb. Even before the addition of regulatory components, many therapeutic proteins used in gene therapy can easily exceed this size.

[0305] Furthermore, the proteins that make up the capsid may act as antigens that can be targeted by the subject's immune system. AAV is very common in the general population, and most people have been exposed to AAV in their lifetime. Therefore, most potential gene therapy recipients are likely to have already developed an immune response to AAV and are therefore more likely to reject the therapy.

[0306] Some aspects of the present disclosure are intended to overcome these defects of AAV vectors. Specifically, some aspects of the present disclosure relate to nucleic acid molecules, which comprise the first ITR, the second ITR and a gene cassette (e.g., encoding a therapeutic protein and / or miRNA). In certain embodiments, the first ITR and the second ITR are flanked by a gene cassette comprising a heterologous polynucleotide sequence. In certain embodiments, the nucleic acid molecule does not comprise genes encoding capsid proteins, replication proteins and / or assembly proteins. In certain embodiments, the gene cassette encodes a therapeutic protein. In certain embodiments, the therapeutic protein comprises a coagulation factor. In certain embodiments, the gene cassette encodes a miRNA. In certain embodiments, the gene cassette is positioned between the first ITR and the second ITR. In certain embodiments, the nucleic acid molecule further comprises one or more non-coding regions. In certain embodiments, the one or more non-coding regions comprise a promoter sequence, an intron, a post-transcriptional regulatory element, a 3'UTR poly (A) sequence or any combination thereof.

[0307] In one embodiment, the gene cassette is a single-stranded nucleic acid. In another embodiment, the gene cassette is a double-stranded nucleic acid.

[0308] In one embodiment, the nucleic acid molecule comprises:

[0309] (a) a first ITR that is an ITR of a non-AAV family member of the Parvoviridae family (e.g., a B19 or GPV ITR);

[0310] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0311] (c) introns, such as synthetic introns;

[0312] (d) nucleotides encoding miRNA or therapeutic proteins (e.g., coagulation factors);

[0313] (e) post-transcriptional regulatory elements, such as WPRE;

[0314] (f) 3′UTR poly(A) tail sequence, such as bGHpA;

[0315] (g) A second ITR that is an ITR of a non-AAV family member of the Parvoviridae family (e.g., a B19 or GPV ITR).

[0316] In one embodiment, the nucleic acid molecule comprises:

[0317] (a) The first ITR of an ITR of a non-AAV family member of the Parvoviridae family;

[0318] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0319] (c) introns, such as synthetic introns;

[0320] (d) nucleotides encoding a miRNA, wherein the miRNA downregulates the expression of a target gene selected from the group consisting of SOD1, HTT, RHO, and any combination thereof;

[0321] (e) post-transcriptional regulatory elements, such as WPRE;

[0322] (f) 3′UTR poly(A) tail sequence, such as bGHpA;

[0323] (g) Second ITR that is an ITR of a non-AAV family member of the Parvoviridae family.

[0324] In one embodiment, the nucleic acid molecule comprises:

[0325] (a) The first ITR of an ITR of a non-AAV family member of the Parvoviridae family;

[0326] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0327] (c) introns, such as synthetic introns;

[0328] (d) nucleotides encoding X-linked dystrophin, MTM1 (myotubulein), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (mitochondrial ataxia protein), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof;

[0329] (e) post-transcriptional regulatory elements, such as WPRE;

[0330] (f) 3′UTR poly(A) tail sequence, such as bGHpA;

[0331] (g) Second ITR that is an ITR of a non-AAV family member of the Parvoviridae family.

[0332] In one embodiment, the nucleic acid molecule comprises:

[0333] (a) a first ITR that is an ITR of an AAV (e.g., an AAV serotype 2 genome);

[0334] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0335] (c) introns, such as synthetic introns;

[0336] (d) a nucleotide encoding FVIII; wherein the nucleotide has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a nucleotide sequence selected from SEQ ID NOs: 1-14 or SEQ ID NO: 71, wherein the FVIII encoded by the nucleotide retains FVIII activity;

[0337] (e) post-transcriptional regulatory elements, such as WPRE;

[0338] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and

[0339] (g) A second ITR that is an ITR of an AAV (e.g., an AAV serotype 2 genome).

[0340] In one embodiment, the nucleic acid molecule comprises:

[0341] (a) a first ITR that is an ITR of an AAV (e.g., an AAV serotype 2 genome);

[0342] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0343] (c) introns, such as synthetic introns;

[0344] (d) nucleotides encoding miRNA, wherein the miRNA downregulates the expression of a target gene, such as SOD1, HTT, RHO, and any combination thereof;

[0345] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and

[0346] (g) A second ITR that is an ITR of an AAV (e.g., an AAV serotype 2 genome).

[0347] In one embodiment, the nucleic acid molecule comprises:

[0348] (a) a first ITR that is an ITR of an AAV (e.g., an AAV serotype 2 genome);

[0349] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0350] (c) introns, such as synthetic introns;

[0351] (d) nucleotides encoding X-linked dystrophin, MTM1 (myotubulein), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (mitochondrial ataxia protein), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof;

[0352] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and

[0353] (g) A second ITR that is an ITR of an AAV (e.g., an AAV serotype 2 genome).

[0354] In another embodiment, the nucleic acid molecule comprises:

[0355] (a) first ITR;

[0356] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0357] (c) introns, such as synthetic introns;

[0358] (d) nucleotides encoding miRNA or therapeutic proteins (e.g., coagulation factors);

[0359] (e) post-transcriptional regulatory elements, such as WPRE;

[0360] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and

[0361] (g) Second ITR,

[0362] wherein one of the first ITR or the second ITR is an ITR of a non-AAV family member of the Parvoviridae family, and the other ITR is an ITR of AAV (eg, an AAV serotype 2 genome).

[0363] In another embodiment, the nucleic acid molecule comprises:

[0364] (a) a 5'ITR having the AAV2 5'ITR sequence set forth in SEQ ID NO: 111;

[0365] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0366] (c) introns, such as synthetic introns;

[0367] (d) a heterologous polynucleotide sequence encoding FVIII (e.g., FVIIIco6XTEN);

[0368] (e) post-transcriptional regulatory elements, such as WPRE;

[0369] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0370] (g) 3' ITR with the AAV2 3' ITR sequence set forth in SEQ ID NO:124.

[0371] In another embodiment, the nucleic acid molecule comprises:

[0372] (a) a 5'ITR having the AAV2 5'ITR sequence set forth in SEQ ID NO: 111;

[0373] (b) tissue-specific promoter sequences, such as the CAG promoter;

[0374] (c) introns, such as synthetic introns;

[0375] (d) a heterologous polynucleotide sequence encoding FVIII (e.g., FVIIIco6XTEN);

[0376] (e) post-transcriptional regulatory elements, such as WPRE;

[0377] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0378] (g) 3' ITR with the AAV2 3' ITR sequence set forth in SEQ ID NO:193.

[0379] In another embodiment, the nucleic acid molecule comprises:

[0380] (a) first ITR;

[0381] (b) tissue-specific promoter sequence, TTP activator;

[0382] (c) introns, such as synthetic introns;

[0383] (d) nucleotides encoding miRNA or therapeutic proteins (e.g., coagulation factors);

[0384] (e) post-transcriptional regulatory elements, such as WPRE;

[0385] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and

[0386] (g) Second ITR,

[0387] The first ITR is a synthetic ITR, the second ITR is a synthetic ITR, or both the first ITR and the second ITR are synthetic ITRs.

[0388] In another embodiment, the nucleic acid molecule comprises:

[0389] (a) First B19 ITR;

[0390] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0391] (c) introns, such as synthetic introns;

[0392] (d) a heterologous polynucleotide sequence encoding a therapeutic protein selected from the group consisting of a coagulation factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, and a combination thereof;

[0393] (e) post-transcriptional regulatory elements, such as WPRE;

[0394] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0395] (g) Second B19 ITR.

[0396] In another embodiment, the nucleic acid molecule comprises:

[0397] (a) First GPV ITR;

[0398] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0399] (c) introns, such as synthetic introns;

[0400] (d) a heterologous polynucleotide sequence encoding a therapeutic protein selected from the group consisting of a coagulation factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, and a combination thereof;

[0401] (e) post-transcriptional regulatory elements, such as WPRE;

[0402] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0403] (g) Second GPV ITR.

[0404] In another embodiment, the nucleic acid molecule comprises:

[0405] (a) First B19 ITR;

[0406] (b) ubiquitous promoter sequences, such as the CAG promoter;

[0407] (c) introns, such as synthetic introns;

[0408] (d) a heterologous polynucleotide sequence encoding a therapeutic protein selected from the group consisting of a coagulation factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, and a combination thereof;

[0409] (e) post-transcriptional regulatory elements, such as WPRE;

[0410] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0411] (g) Second B19 ITR.

[0412] In another embodiment, the nucleic acid molecule comprises:

[0413] (a) First GPV ITR;

[0414] (b) ubiquitous promoter sequences, such as the CAG promoter;

[0415] (c) introns, such as synthetic introns;

[0416] (d) a heterologous polynucleotide sequence encoding a therapeutic protein selected from the group consisting of a coagulation factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, and a combination thereof;

[0417] (e) post-transcriptional regulatory elements, such as WPRE;

[0418] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0419] (g) Second GPV ITR.

[0420] In another embodiment, the nucleic acid molecule comprises:

[0421] (a) First B19 ITR;

[0422] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0423] (c) introns, such as synthetic introns;

[0424] (d) a heterologous polynucleotide sequence encoding phenylalanine hydroxylase (PAH);

[0425] (e) post-transcriptional regulatory elements, such as WPRE;

[0426] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0427] (g) Second B19 ITR.

[0428] In another embodiment, the nucleic acid molecule comprises:

[0429] (a) First GPV ITR;

[0430] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0431] (c) introns, such as synthetic introns;

[0432] (d) a heterologous polynucleotide sequence encoding phenylalanine hydroxylase (PAH);

[0433] (e) post-transcriptional regulatory elements, such as WPRE;

[0434] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0435] (g) Second GPV ITR.

[0436] In another embodiment, the nucleic acid molecule comprises:

[0437] (a) a 5'ITR having the B19d135 5'ITR sequence described in SEQ ID NO: 180;

[0438] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0439] (c) introns, such as synthetic introns;

[0440] (d) a heterologous polynucleotide sequence encoding FVIII (e.g., FVIIIco6XTEN);

[0441] (e) post-transcriptional regulatory elements, such as WPRE;

[0442] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0443] (g) A 3'ITR having the B19d135 3'ITR sequence set forth in SEQ ID NO:181.

[0444] In another embodiment, the nucleic acid molecule comprises:

[0445] (a) a 5'ITR having the GPVd162 5'ITR sequence described in SEQ ID NO: 183;

[0446] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0447] (c) introns, such as synthetic introns;

[0448] (d) a heterologous polynucleotide sequence encoding FVIII (e.g., FVIIIco6XTEN);

[0449] (e) post-transcriptional regulatory elements, such as WPRE;

[0450] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0451] (g) 3'ITR with the GPVd162 3'ITR sequence depicted in SEQ ID NO:184.

[0452] In another embodiment, the nucleic acid molecule comprises:

[0453] (a) a 5'ITR having the full-length B19 5'ITR sequence set forth in SEQ ID NO: 185;

[0454] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0455] (c) introns, such as synthetic introns;

[0456] (d) a heterologous polynucleotide sequence encoding FVIII (e.g., FVIIIco6XTEN);

[0457] (e) post-transcriptional regulatory elements, such as WPRE;

[0458] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0459] (g) 3'ITR with the full-length B19 3'ITR sequence set forth in SEQ ID NO:186.

[0460] In another embodiment, the nucleic acid molecule comprises:

[0461] (a) a 5'ITR having the full-length GPV 5'ITR sequence set forth in SEQ ID NO: 187;

[0462] (b) tissue-specific promoter sequences, such as the TTP promoter;

[0463] (c) introns, such as synthetic introns;

[0464] (d) a heterologous polynucleotide sequence encoding FVIII (e.g., FVIIIco6XTEN);

[0465] (e) post-transcriptional regulatory elements, such as WPRE;

[0466] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0467] (g) 3'ITR with the full-length GPV 3'ITR sequence set forth in SEQ ID NO:188.

[0468] In another embodiment, the nucleic acid molecule comprises:

[0469] (a) a 5'ITR having the B19d135 5'ITR sequence described in SEQ ID NO: 180;

[0470] (b) tissue-specific promoter sequences, such as the CAG promoter;

[0471] (c) introns, such as synthetic introns;

[0472] (d) a heterologous polynucleotide sequence encoding PAH;

[0473] (e) post-transcriptional regulatory elements, such as WPRE;

[0474] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0475] (g) A 3'ITR having the B19d135 3'ITR sequence set forth in SEQ ID NO:181.

[0476] In another embodiment, the nucleic acid molecule comprises:

[0477] (a) a 5'ITR having the GPVd162 5'ITR sequence described in SEQ ID NO: 183;

[0478] (b) tissue-specific promoter sequences, such as the CAG promoter;

[0479] (c) introns, such as synthetic introns;

[0480] (d) a heterologous polynucleotide sequence encoding PAH;

[0481] (e) post-transcriptional regulatory elements, such as WPRE;

[0482] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0483] (g) 3'ITR with the GPVd162 3'ITR sequence depicted in SEQ ID NO:184.

[0484] In another embodiment, the nucleic acid molecule comprises:

[0485] (a) a 5'ITR having the full-length B19 5'ITR sequence set forth in SEQ ID NO: 185;

[0486] (b) tissue-specific promoter sequences, such as the CAG promoter;

[0487] (c) introns, such as synthetic introns;

[0488] (d) a heterologous polynucleotide sequence encoding PAH;

[0489] (e) post-transcriptional regulatory elements, such as WPRE;

[0490] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0491] (g) 3'ITR with the full-length B19 3'ITR sequence set forth in SEQ ID NO:186.

[0492] In another embodiment, the nucleic acid molecule comprises:

[0493] (a) a 5'ITR having the full-length GPV 5'ITR sequence set forth in SEQ ID NO: 187;

[0494] (b) tissue-specific promoter sequences, such as the CAG promoter;

[0495] (c) introns, such as synthetic introns;

[0496] (d) a heterologous polynucleotide sequence encoding PAH;

[0497] (e) post-transcriptional regulatory elements, such as WPRE;

[0498] (f) a 3'UTR poly(A) tail sequence, such as bGHpA; and / or

[0499] (g) 3'ITR with the full-length GPV 3'ITR sequence set forth in SEQ ID NO:188.

[0500] A. Inverted terminal repeats

[0501] Certain aspects of the present disclosure relate to nucleic acid molecules comprising a first ITR (e.g., a 5' ITR) and a second ITR (e.g., a 3' ITR). Generally, ITRs are involved in DNA replication and rescue or excision of parvoviruses (e.g., AAV) from prokaryotic plasmids (Samulski et al., 1983, 1987; Senapathy et al., 1984; Gottlieb and Muzyczka, 1988). Additionally, ITRs appear to be the minimal sequences required for AAV proviral integration and packaging of AAV DNA into virions (McLaughlin et al., 1988; Samulski et al., 1989). These components are essential for efficient manipulation of the parvovirus genome. It is hypothesized that the minimal defining components essential for ITR function are a Rep binding site (e.g., RBS; for AAV2, GCGCGCTCGCTCGCTC (SEQ ID NO: 104)) and a terminal melting site (e.g., TRS; for AAV2, AGTTGG (SEQ ID NO: 105)) plus a variable palindromic sequence that allows hairpin formation. Palindromic nucleotide regions typically act together in cis as the origin of DNA replication and as a packaging signal for the virus. Complementary sequences in the ITR fold into a hairpin structure during DNA replication. In some embodiments, the ITR folds into a hairpin T-shaped structure. In other embodiments, the ITR folds into a non-T-shaped hairpin structure, for example, a U-shaped hairpin structure. Data suggest that the T-shaped hairpin structure of the AAV ITR can inhibit the expression of transgenes flanked by the ITR. See, e.g., Zhou et al., Scientific Reports 7: 5432 (July 14, 2017). By utilizing ITRs that do not form a T-shaped hairpin structure, this form of inhibition can be avoided. Thus, in certain aspects, polynucleotides comprising non-AAV ITRs have improved transgene expression compared to polynucleotides comprising AAV ITRs that form T-shaped hairpins.

[0502] In some embodiments, the ITR comprises a naturally occurring ITR, e.g., an ITR comprises all or a portion of a parvoviral ITR. In some embodiments, the ITR comprises a synthetic sequence. In one embodiment, the first ITR or the second ITR comprises a synthetic sequence. In another embodiment, the first ITR and the second ITR each comprise a synthetic sequence. In some embodiments, the first ITR or the second ITR comprises a naturally occurring sequence. In another embodiment, the first ITR and the second ITR each comprise a naturally occurring sequence.

[0503] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR (e.g., a truncated ITR). In some embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides, or at least about 600 nucleotides; wherein the ITR retains the functional properties of a naturally occurring ITR. In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 129 nucleotides; wherein the ITR retains the functional properties of a naturally occurring ITR. In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 102 nucleotides; wherein the ITR retains the functional properties of the naturally occurring ITR.

[0504] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR, wherein the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the length of the naturally occurring ITR; wherein the fragment retains the functional properties of the naturally occurring ITR.

[0505] In certain embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 90% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of a naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when correctly aligned, has at least 80% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of a naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when correctly aligned, has at least 70% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of a naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when correctly aligned, has at least 60% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of a naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when correctly aligned, has at least 50% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of a naturally occurring ITR.

[0506] In some embodiments, the ITR comprises an ITR from an AAV genome. In some embodiments, the ITR is an ITR from an AAV genome selected from the group consisting of: AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, and any combination thereof. In a specific embodiment, the ITR is an ITR from an AAV2 genome. In another embodiment, the ITR is a synthetic sequence genetically engineered to include at its 5' and 3' ends an ITR derived from one or more AAV genomes.

[0507] In some embodiments, the ITR is not derived from an AAV genome. In some embodiments, the ITR is a non-AAV ITR. In some embodiments, the ITR is an ITR from a non-AAV genome selected from, but not limited to, the following viral family Parvoviridae: Bocavirus, Dependentovirus, Erythrovirus, Aleutianvirus, Parvovirus, Densovirus, Repetitive Virus, Contravirus, Avian Parvovirus, Ruminant Parvovirus, Protoparvovirus, Quadparvovirus, Amphivirus, Brevovirus, Hepatopancreatic Densovirus, Shrimp Densovirus, and any combination thereof. In certain embodiments, the ITR is derived from Erythrovirus Parvovirus B19 (human virus). In another embodiment, the ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, such as MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, such as MDPV strain YY. In some embodiments, the ITR is derived from porcine parvovirus, such as porcine parvovirus U44978. In some embodiments, the ITR is derived from minute virus of mice, such as minute virus of mice U34256. In some embodiments, the ITR is derived from canine parvovirus, such as canine parvovirus M19296. In some embodiments, the ITR is derived from mink enteritis virus, such as mink enteritis virus D00765. In some embodiments, the ITR is derived from Dependoparvovirus. In one embodiment, the Dependoparvovirus is a Dependovirus goose parvovirus (GPV) strain. In a specific embodiment, the GPV strain is attenuated, such as GPV strain 82-0321V. In another specific embodiment, the GPV strain is pathogenic, such as GPV strain B.

[0508] The first ITR and the second ITR of the nucleic acid molecule can be derived from the same genome (e.g., derived from the genome of the same virus), or derived from different genomes (e.g., derived from genomes of two or more different viral genomes). In certain embodiments, the first ITR and the second ITR are derived from the same AAV genome. In a specific embodiment, the two ITRs present in the nucleic acid molecule of the present invention are identical, and in particular can be AAV2 ITRs. In other embodiments, the first ITR is derived from the AAV genome and the second ITR is not derived from the AAV genome (e.g., non-AAV genome). In other embodiments, the first ITR is not derived from the AAV genome (e.g., non-AAV genome) and the second ITR is derived from the AAV genome. In still other embodiments, both the first ITR and the second ITR are not derived from the AAV genome (e.g., non-AAV genome). In a specific embodiment, the first ITR and the second ITR are identical.

[0509] In some embodiments, the first ITR is derived from an AAV genome and the second ITR is derived from a genome selected from the group consisting of: Bocavirus, Dependovirus, Erythrovirus, Aleutianvirus, Parvovirus, Densovirus, Retauvirus, Contravirus, Avian Parvovirus, Ruminant Parvovirus, Proparvovirus, Quadparvovirus, Ambisense Densovirus, Brevdensovirus, Hepatopancreatic Densovirus, Shrimp Densovirus, and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome and the first ITR is derived from a genome selected from the group consisting of: Bocavirus, Dependovirus, Erythrovirus, Aleutianvirus, Parvovirus, Densovirus, Retauvirus, Contravirus, Avian Parvovirus, Ruminant Parvovirus, Proparvovirus, Quadparvovirus, Ambisense Densovirus, Brevdensovirus, Hepatopancreatic Densovirus, Shrimp Densovirus, and any combination thereof. In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of: Bocavirus, Dependovirus, Erythrovirus, Aleutianvirus, Parvovirus, Densovirus, Repetitivevirus, Contravirus, Avian Parvovirus, Ruminant Parvovirus, Proparvovirus, Quadparvovirus, Ambisense Densovirus, Brevovirus, Hepatopancreatic Densovirus, Shrimp Densovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from the same genome. In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of: Bocavirus, Dependovirus, Erythrovirus, Aleutianvirus, Parvovirus, Densovirus, Repetitivevirus, Contravirus, Avian Parvovirus, Ruminant Parvovirus, Proparvovirus, Quadparvovirus, Ambisense Densovirus, Brevovirus, Hepatopancreatic Densovirus, Shrimp Densovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from different genomes.

[0510] In some embodiments, the first ITR is derived from the AAV genome and the second ITR is derived from Erythrovirus Parvovirus B19 (a human virus). In other embodiments, the second ITR is derived from the AAV genome and the first ITR is derived from Erythrovirus Parvovirus B19 (a human virus).

[0511] In certain embodiments, the first and / or second ITR comprises or consists of all or a portion of an ITR derived from B 19. In some embodiments, the first and / or second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171, wherein the first and / or second ITR retains the functional properties of the B 19 ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0512] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 167, 168, 169, 170, and 171. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 167. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 168. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 169. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 170. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 171.

[0513] Table 1. Sample parvovirus ITR sequences.

[0514]

[0515]

[0516]

[0517] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 167. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 168. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 168. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 169, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167, and retains the functional properties of the B19 ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 169, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0518] In certain embodiments, the first and / or second ITR comprises or consists of all or a portion of an ITR derived from B19. In some embodiments, the second ITR is the reverse complement of the first ITR. In some embodiments, the first ITR is the reverse complement of the second ITR. In some embodiments, the first and / or second ITR comprises or consists of a nucleotide sequence, or a functional derivative thereof, that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 180, 181, 185, and 186. In some embodiments, the functional derivative retains the functional properties of the B19 ITR from which it is derived. In some embodiments, the first and / or second ITR comprises or consists of a nucleotide sequence or a functional derivative thereof that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 180, 181, 185, and 186. In some embodiments, the functional derivative is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0519] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 180. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 180. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 181. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 181. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 185. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 185. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 186. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 186.

[0520] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 180, 181, 185, and 186. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 180. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 181. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 185. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 186. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 180. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 181. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 185. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 186.

[0521] In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 180, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 181. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 181, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 180. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 185, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 186. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 186, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 185.

[0522] In some embodiments, the first ITR is derived from the AAV genome and the second ITR is derived from GPV. In other embodiments, the second ITR is derived from the AAV genome and the first ITR is derived from GPV.

[0523] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 172. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 173. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 173. In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from GPV. In some embodiments, the first and / or second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 172, 173, 174, 175, and 176, wherein the first and / or second ITR retains the functional properties of the GPV ITR from which it is derived. In some embodiments, the first and / or second ITR comprises or consists of all or a portion of an ITR derived from GPV. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 172, 173, 174, 175, and 176, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 172, 173, 174, 175, and 176. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 172. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 173. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 174. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 175. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 176.

[0524] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 174, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR retains the functional properties of the GPV ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 174, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0525] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 176, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR retains the functional properties of the GPV ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 176, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0526] In certain embodiments, the first and / or second ITR comprises or consists of all or a portion of an ITR derived from GPV. In some embodiments, the second ITR is the reverse complement of the first ITR. In some embodiments, the first ITR is the reverse complement of the second ITR. In some embodiments, the first and / or second ITR comprises or consists of a nucleotide sequence, or a functional derivative thereof, that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 183, 184, 187, and 188. In some embodiments, the functional derivative retains the functional properties of the GPV ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence or a functional derivative thereof that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 183, 184, 187, and 188. In some embodiments, the functional derivative is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0527] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 183. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 183. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 184. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 184. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 187. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 187. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 188. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 188.

[0528] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 183, 184, 187, and 188. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 183. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 184. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 187. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 188. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 183. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 184. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 187. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 188.

[0529] In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 183, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 184. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 184, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 183. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 187, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 188. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 188, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 187.

[0530] In certain embodiments, one of the first ITR or the second ITR comprises or consists of all or a portion of an ITR derived from AAV2. In some embodiments, the first ITR or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 177 or 178, wherein the first ITR and / or the second ITR retains the functional properties of the AAV2 ITR from which it is derived. In some embodiments, the first or second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 177 or 178, wherein the first and / or second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin. In some embodiments, the first and / or second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 177 or 178. In some embodiments, the first and / or second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 177. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 178.

[0531] In some embodiments, the first ITR is derived from the AAV genome and the second ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In other embodiments, the second ITR is derived from the AAV genome and the first ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, such as MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, such as MDPV strain YY.

[0532] In some embodiments, the first ITR is derived from the AAV genome and the second ITR is derived from a Dependaviridae virus. In some embodiments, the second ITR is derived from the AAV genome and the first ITR is derived from a Dependaviridae virus. In other embodiments, the first ITR is derived from the AAV genome and the second ITR is derived from a Dependaviridae virus goose parvovirus (GPV) strain. In other embodiments, the second ITR is derived from the AAV genome and the first ITR is derived from a Dependaviridae virus GPV strain. In certain embodiments, the GPV strain is attenuated, such as GPV strain 82-0321V. In other embodiments, the GPV strain is pathogenic, such as GPV strain B.

[0533] In certain embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a genome selected from the group consisting of: porcine parvovirus, such as porcine parvovirus strain U44978; minute virus of mice, such as minute virus of mice strain U34256; canine parvovirus, such as canine parvovirus strain M19296; mink enteritis virus, such as mink enteritis virus strain D00765; and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a genome selected from the group consisting of: porcine parvovirus, such as porcine parvovirus strain U44978; minute virus of mice, such as minute virus of mice strain U34256; canine parvovirus, such as canine parvovirus strain M19296; mink enteritis virus, such as mink enteritis virus strain D00765; and any combination thereof.

[0534] In another specific embodiment, the ITR is a synthetic sequence that has been genetically engineered to include an ITR at its 5' and 3' ends that is not derived from an AAV genome. In another specific embodiment, the ITR is a synthetic sequence that has been genetically engineered to include an ITR at its 5' and 3' ends that is derived from one or more non-AAV genomes. The two ITRs present in the nucleic acid molecule of the present invention may be the same or different non-AAV genomes. In particular, the ITRs may be derived from the same non-AAV genome. In a specific embodiment, the two ITRs present in the nucleic acid molecule of the present invention are identical and may in particular be AAV2 ITRs.

[0535] In some embodiments, the ITR sequence comprises one or more palindromic sequences. The palindromic sequences of the ITRs disclosed herein include, but are not limited to, natural palindromic sequences (i.e., sequences found in nature), synthetic sequences (i.e., sequences not found in nature) (such as pseudo palindromic sequences), and combinations or modified forms thereof. A "pseudo palindromic sequence" is a palindromic DNA sequence including an imperfect palindromic sequence that shares less than 80% (including less than 70%, 60%, 50%, 40%, 30%, 20%, 10% or 5% or no) nucleic acid sequence identity with a sequence in a natural AAV or non-AAV palindromic sequence that forms a secondary structure. Natural palindromic sequences can be obtained or derived from any genome disclosed herein. Synthetic palindromic sequences can be based on any genome disclosed herein.

[0536] The palindrome sequence can be continuous or interrupted. In some embodiments, the palindrome sequence is interrupted, wherein the palindrome sequence comprises an insertion of a second sequence. In some embodiments, the second sequence comprises a promoter, an enhancer, an integration site for an integrase (e.g., a site for Cre or Flp recombinase), an open reading frame for a gene product, or a combination thereof.

[0537] In some embodiments, the ITRs form a hairpin loop structure. In one embodiment, the first ITR forms a hairpin structure. In another embodiment, the second ITR forms a hairpin structure. In still another embodiment, both the first ITR and the second ITR form a hairpin structure. In some embodiments, the first ITR and / or the second ITR do not form a T-shaped hairpin structure. In certain embodiments, the first ITR and / or the second ITR form a non-T-shaped hairpin structure. In some embodiments, the non-T-shaped hairpin structure comprises a U-shaped hairpin structure.

[0538] In some embodiments, the ITR in the nucleic acid molecules described herein can be a transcriptionally activated ITR. A transcriptionally activated ITR can comprise all or part of a wild-type ITR that has been transcriptionally activated by including at least one transcriptionally activated element. Various types of transcriptionally activated elements are suitable for use in this context. In some embodiments, the transcriptionally activated element is a constitutive transcriptionally activated element. Constitutive transcriptionally activated elements provide a sustained level of gene transcription and are preferred when sustained expression of a transgene is desired. In other embodiments, the transcriptionally activated element is an inducible transcriptionally activated element. Inducible transcriptionally activated elements typically exhibit low activity in the absence of an inducer (or inducing conditions) and are upregulated in the presence of an inducer (or when switched to inducing conditions). Inducible transcriptionally activated elements may be preferred when expression is desired only at certain times or in certain locations, or when expression levels need to be gradually increased using an inducer. Transcriptionally activated elements can also be tissue-specific; that is, they exhibit activity only in certain tissues or cell types.

[0539] Transcriptional activity elements can be incorporated into ITRs in a variety of ways. In certain embodiments, transcriptional activity elements are incorporated into the 5′ or 3′ of any part of an ITR. In other embodiments, the transcriptional activity element of a transcriptionally activated ITR is located between two ITR sequences. If a transcriptional activity element comprises two or more components that must be spaced apart, those components can alternate with parts of the ITR. In certain embodiments, the hairpin structure of an ITR is missing and replaced with an inverted repeat of a transcriptional component. This latter arrangement will produce a hairpin of the missing part in the emulation structure. Multiple tandem transcriptional activity elements may also be present in the transcriptionally activated ITR, and these components may be adjacent or spaced apart. In addition, a protein binding site (e.g., a Rep binding site) may be introduced into the transcriptional activity element of a transcriptionally activated ITR. A transcriptional activity element may comprise any sequence that enables controlled transcription of DNA to form RNA by RNA polymerase, and may comprise, for example, a transcriptional activity element as defined below.

[0540] The ITR of transcriptional activation provides both transcriptional activation and ITR function to nucleic acid molecules with relatively limited nucleotide sequence length, which effectively maximizes the length of the transgenic transgene that can be carried and expressed from the nucleic acid molecule. Incorporating transcriptional activation elements into the ITR can be accomplished in a variety of ways. Comparison of the sequence requirements of the ITR sequence and the transcriptional activation elements can provide an understanding of the mode of the assembly in the coding ITR. For example, transcriptional activity can be added to the ITR by introducing specific changes in the ITR sequence of the functional component of the replication transcriptional activation element. There are multiple technologies in the industry that can effectively add, delete and / or change the specific nucleotide sequence of a specific site (see, for example, Deng and Nickoloff (1992) Anal. Biochem. 200: 81-88). Another way to produce the ITR of transcriptional activation relates to introducing restriction sites at the desired position in the ITR. In addition, methods known in the art can be used to incorporate multiple transcriptional activation elements into the ITR of transcriptional activation.

[0541] By way of illustration, a transcriptionally activated ITR can be generated by including one or more transcriptionally active elements such as a TATA box, a GC box, a CCAAT box, an Sp1 site, an Inr region, a CRE (cAMP regulatory element) site, an ATF-1 / CRE site, an APBβ box, an APBα box, a CArG box, a CCAC box, or any other component known in the art to be involved in transcription.

[0542] Aspects of the present disclosure provide methods for cloning nucleic acid molecules as described herein, comprising inserting a nucleic acid molecule capable of forming a complex secondary structure into a suitable vector, and introducing the resulting vector into a suitable bacterial host strain. As known in the art, the complex secondary structure of nucleic acids (e.g., long palindromes) may be unstable and difficult to clone in bacterial host strains. For example, nucleic acid molecules comprising the first ITR and the second ITR (e.g., non-AAV parvovirus ITRs, such as B19 or GPV ITRs) of the present disclosure may be difficult to clone using conventional methods. Long DNA palindromic sequences inhibit DNA replication and are unstable in the genomes of Escherichia coli (E. coli), Bacillus (Bacillus), Streptococcus (Steptococcus), Streptomyces (Streptomyces), Saccharomyces cerevisiae (S. cerevisiae), mice, and humans. These effects are due to the formation of hairpin or cruciform structures by intrachain base pairing. In E. coli, inhibition of DNA replication can be significantly overcome in SbcC or SbcD mutants. SbcD is the nuclease subunit, and SbcC is the ATPase subunit of the SbcCD complex. The Escherichia coli SbcCD complex is an exonuclease complex responsible for preventing the replication of long palindromic sequences. The SbcCD complex is a nuclear complex with ATP-dependent double-stranded DNA exonuclease activity and ATP-independent single-stranded DNA endonuclease activity. SbcCD can recognize DNA palindromes and disrupt replication forks by attacking the resulting hairpin structure.

[0543] In certain embodiments, the suitable bacterial host strain is unable to split the cruciform DNA structure. In certain embodiments, the suitable bacterial host strain comprises a disruption in the SbcCD complex. In some embodiments, the disruption in the SbcCD complex comprises a gene disruption in the SbcC gene and / or the SbcD gene. In certain embodiments, the disruption in the SbcCD complex comprises a gene disruption in the SbcC gene. Various bacterial host strains comprising a gene disruption in the SbcC gene are known in the art. For example, but not limited to, bacterial host strain PMC103 comprises the genotypes sbcC, recD, mcrA, ΔmcrBCF; bacterial host strain PMC107 comprises the genotypes recBC, recJ, sbcBC, mcrA, ΔmcrBCF; and bacterial host strain SURE comprises the genotypes recB, recJ, sbcC, mcrA, ΔmcrBCF, umuC, uvrC. Thus, in some embodiments, the method for cloning a nucleic acid molecule described herein comprises inserting a nucleic acid molecule capable of forming a complex secondary structure into a suitable vector, and introducing the resulting vector into a host strain PMC103, PMC107, or SURE. In certain embodiments, the method for cloning a nucleic acid molecule described herein comprises inserting a nucleic acid molecule capable of forming a complex secondary structure into a suitable vector, and introducing the resulting vector into a host strain PMC103.

[0544] Suitable vectors are known in the art and are described elsewhere herein. In certain embodiments, a suitable vector for use in the cloning methods of the present disclosure is a low copy vector. In certain embodiments, a suitable vector for use in the cloning methods of the present disclosure is pBR322.

[0545] Accordingly, the present disclosure provides a method for cloning a nucleic acid molecule, comprising inserting a nucleic acid molecule capable of forming a complex secondary structure into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex, wherein the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence depicted in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0546] B. Therapeutic Proteins

[0547] Certain aspects of the present disclosure relate to nucleic acid molecules comprising a first ITR, a second ITR, and a gene cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein. In some embodiments, the gene cassette encodes one therapeutic protein. In some embodiments, the gene cassette encodes more than one therapeutic protein. In some embodiments, the gene cassette encodes two or more copies of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more variants of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more different therapeutic proteins.

[0548] Certain embodiments of the present disclosure relate to nucleic acid molecules comprising the first ITR, the second ITR and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a coagulation factor. In certain embodiments, the coagulation factor is selected from FI, FII, FIII, FIV, FV, FVI, FVII, FVIII, FIX, FX, FXI, FXII, FXIII, VWF, prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, α2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), any zymogen thereof, any active form thereof and any combination thereof. In one embodiment, the coagulation factor comprises FVIII or its variant or fragment. In another embodiment, the coagulation factor comprises FIX or its variant or fragment. In another embodiment, the coagulation factor comprises FVII or its variant or fragment. In another embodiment, the coagulation factor comprises VWF or a variant or fragment thereof.

[0549] 1. Coagulation factors

[0550] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR and a gene cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein, wherein the therapeutic protein comprises a factor VIII polypeptide. Unless otherwise specified, "factor VIII" as used herein is abbreviated as "FVIII" throughout this application to refer to a functional FVIII polypeptide having its normal role in coagulation. Thus, the term FVIII includes functional variant polypeptides. "FVIII protein" can be used interchangeably with FVIII polypeptides (or proteins) or FVIII. Examples of FVIII functions include, but are not limited to, the ability to activate coagulation, the ability to act as a cofactor for factor IX, or in Ca 2+The ability to form a tenase complex with factor IX in the presence of phospholipids, which then converts factor X into the activated form, Xa. The FVIII protein can be human, porcine, canine, rat, or murine FVIII protein. In addition, comparisons between FVIII from humans and other species have identified conserved residues that may be required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US 6,251,632). Full-length polypeptide and polynucleotide sequences are known, as are many functional fragments, mutants, and modified forms. Multiple FVIII amino acid and nucleotide sequences are disclosed, for example, in the following documents: U.S. Publication Nos. 2015 / 0158929A1, 2014 / 0308280A1, and 2014 / 0370035A1, and International Publication No. WO 2015 / 106052 A1. FVIII polypeptides include, for example, full-length FVIII, full-length FVIII lacking a Met at the N-terminus, mature FVIII (lacking a signal sequence), mature FVIII with an additional Met at the N-terminus, and / or FVIII with a complete or partial deletion of the B domain. FVIII variants include partial or complete deletions of the B domain.

[0551] a. FVIII and polynucleotide sequences encoding FVIII proteins

[0552] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR and a gene cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein, wherein the therapeutic protein comprises a factor VIII polypeptide. Unless otherwise specified, "factor VIII" as used herein is abbreviated as "FVIII" throughout this application to refer to a functional FVIII polypeptide having its normal role in coagulation. Thus, the term FVIII includes functional variant polypeptides. "FVIII protein" can be used interchangeably with FVIII polypeptides (or proteins) or FVIII. Examples of FVIII functions include, but are not limited to, the ability to activate coagulation, the ability to act as a cofactor for factor IX, or in Ca 2+The ability to form a factor X enzyme complex with factor IX in the presence of phospholipids, which then converts factor X to the activated form, Xa. The FVIII protein can be human, porcine, canine, rat, or murine FVIII protein. In addition, comparisons between FVIII from humans and other species have identified conserved residues that may be required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US 6,251,632). Full-length polypeptide and polynucleotide sequences are known, as are many functional fragments, mutants, and modified forms. Multiple FVIII amino acid and nucleotide sequences are disclosed, for example, in the following documents: U.S. Publication Nos. 2015 / 0158929A1, 2014 / 0308280A1, and 2014 / 0370035A1, and International Publication No. WO 2015 / 106052 A1. FVIII polypeptides include, for example, full-length FVIII, full-length FVIII lacking a Met at the N-terminus, mature FVIII (lacking a signal sequence), mature FVIII with an additional Met at the N-terminus, and / or FVIII with a complete or partial deletion of the B domain. FVIII variants include partial or complete deletions of the B domain.

[0553] The FVIII portion in the chimeric protein used herein has FVIII activity. FVIII activity can be measured by any known method in the industry. A variety of tests can be used to assess the function of the coagulation system: activated partial thromboplastin time (aPTT) test, chromogenic assay, ROTEM assay, prothrombin time (PT) test (also used to determine INR), fibrinogen test (usually performed by Claus method), platelet count, platelet function test (usually performed by PFA-100), TCT, bleeding time, mixed test (if the patient's plasma is mixed with normal plasma, whether the abnormality is corrected), coagulation factor determination, antiphospholipid antibodies, D-dimer, genetic test (e.g., Factor V Leiden, prothrombin mutation G20210A), diluted Russell's viper venom time (dRVVT), miscellaneous platelet function tests, thromboelastometry (TEG or Sonoclot), thromboelastometry ( For example ) or euglobulin lysis time (ELT).

[0554] The aPTT test is an indicator of the efficacy of both the "intrinsic" (also known as the contact activation pathway) and the common coagulation pathway. This test is generally used to measure the coagulation activity of commercially available recombinant coagulation factors (e.g., FVIII). It is used in conjunction with the prothrombin time (PT), which measures the extrinsic pathway.

[0555] ROTEM analysis provides information on the overall dynamics of hemostasis: clotting time, clot formation, clot stability, and lysis. The different parameters in thromboelastometry depend on the activity of the plasma coagulation system, platelet function, fibrinolysis, or a number of factors that influence these interactions. This analysis can provide a comprehensive understanding of secondary hemostasis.

[0556] The chromogenic assay mechanism is based on the principle of the blood coagulation cascade, in which activated FVIII accelerates the conversion of factor X to factor Xa in the presence of activated factor IX, phospholipids, and calcium ions. Factor Xa activity is assessed by hydrolysis of a p-nitroaniline (pNA) substrate specific for factor Xa. The initial release rate of p-nitroaniline, measured at 405 nM, is directly proportional to the factor Xa activity and, therefore, to the FVIII activity in the sample.

[0557] The chromogenic assay is recommended by the FVIII and Factor IX Subcommittee of the Scientific Standardization Committee (SSC) of the International Society on Thrombosis and Haemostasis (ISTH). The chromogenic assay has also been the reference method for specifying the potency of FVIII concentrates in the European Pharmacopoeia since 1994. Thus, in one embodiment, a chimeric polypeptide comprising FVIII has a similar potency to a chimeric polypeptide comprising mature FVIII or BDD FVIII (e.g., or ) equivalent FVIII activity.

[0558] In another embodiment, the Factor Xa production rate of a chimeric protein comprising FVIII of the present disclosure is comparable to that of a chimeric protein comprising mature FVIII or BDD FVIII (e.g., or )quite.

[0559] To activate Factor X to Factor Xa, activated Factor IX (Factor IXa) is activated in the presence of Ca 2+ In the presence of phospholipids, membrane phospholipids and FVIII cofactor, an arginine-isoleucine bond in factor X is hydrolyzed to form factor Xa. Thus, the interaction of FVIII with factor IX plays a key role in the coagulation pathway. In certain embodiments, a chimeric polypeptide comprising FVIII can be used in conjunction with a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ) interacts with Factor IXa at a comparable rate.

[0560] In addition, FVIII binds to von Willebrand factor while being inactive in the circulation. FVIII rapidly degrades when not bound to VWF and is released from VWF by the action of thrombin. In some embodiments, a chimeric polypeptide comprising FVIII is conjugated to a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ) binds to von Willebrand factor at comparable levels.

[0561] FVIII can be inactivated by activated protein C in the presence of calcium and phospholipids. Activated protein C cuts the FVIII heavy chain after arginine 336 in the A1 domain, thereby destroying the factor X substrate interaction site, and cuts after arginine 562 in the A2 domain, thereby enhancing the dissociation of the A2 domain and destroying the interaction site with factor IXa. This cleavage also bisects the A2 domain (43 kDa) and generates A2-N (18 kDa) and A2-C (25 kDa) domains. Thus, activated protein C can catalyze multiple cleavage sites in the heavy chain. In one embodiment, activated protein C is used to bind to a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ) levels that inactivate the chimeric polypeptide comprising FVIII.

[0562] In other embodiments, the chimeric protein comprising FVIII has a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or In a specific embodiment, in the HemA mouse tail vein transection model, the chimeric polypeptide comprising FVIII can be used to produce FVIII at a rate comparable to that of a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ) protected HemA mice at a comparable level.

[0563] As used herein, the "B domain" of FVIII is identical to the B domain known in the art, which is defined by the internal amino acid sequence identity of mature human FVIII and the proteolytic cleavage site of thrombin (e.g., residues Ser741-Arg1648). Other human FVIII domains are defined relative to mature human FVIII by the following amino acid residues: A1, residues Ala1-Arg372 of mature FVIII; A2, residues Ser373-Arg740; A3, residues Ser1690-Ile2032; C1, residues Arg2033-Asn2172; C2, residues Ser2173-Tyr2332. Unless otherwise indicated, the sequence residue numbers used herein correspond to the FVIII sequence without signal peptide sequence (19 amino acids) without reference to any SEQ ID numbering. The A3-C1-C2 sequence (also referred to as the FVIII heavy chain) includes residues Ser1690-Tyr2332. The remaining sequence (residues Glu1649-Arg1689) is generally referred to as the FVIII light chain activation peptide. The locations of the boundaries of all domains (including the B domain) of porcine, mouse, and canine FVIII are also known in the art. In one embodiment, the B domain of FVIII is deleted ("B domain deleted FVIII" or "BDD FVIII"). An example of BDD FVIII is (Recombinant BD FVIII). In a specific embodiment, the B-domain deleted FVIII variant comprises a deletion of amino acid residues 746 to 1648 of mature FVIII.

[0564] "B-domain deleted FVIII" can have a complete or partial deletion as disclosed in U.S. Patent Nos. 6,316,226, 6,346,513, 7,041,635, 5,789,203, 6,060,447, 5,595,886, 6,228,620, 5,972,885, 6,048,720, 5,543,502, 5,610,278, 5,171,844, 5,112,950, 4,868,112, and 6,458,563 and International Publication No. WO 2015106052 A1 (PCT / US2015 / 010738). In some embodiments, the B-domain deleted FVIII sequence used in the methods of the present disclosure comprises any one of the deletions disclosed in column 4, line 4 to column 5, line 28 of U.S. Pat. No. 6,316,226 (also in U.S. Pat. No. 6,346,513) and Examples 1-5. In another embodiment, the B-domain deleted Factor VIII is S743 / Q1638 B-domain deleted Factor VIII (SQ BDD FVIII) (e.g., Factor VIII having a deletion of amino acids 744 to amino acids 1637, e.g., Factor VIII having amino acids 1-743 and amino acids 1638-2332 of mature FVIII). In some embodiments, the B-domain deleted FVIII used in the methods of the present disclosure has a deletion disclosed in column 2, lines 26-51 of U.S. Pat. No. 5,789,203 (also in U.S. Pat. No. 6,060,447, U.S. Pat. No. 5,595,886, and U.S. Pat. No. 6,228,620) and Examples 5-8. In some embodiments, the B-domain deleted Factor VIII has a deletion as described in U.S. Pat. No. 5,972,885 at column 1, line 25 to column 2, line 40; U.S. Pat. No. 6,048,720 at column 6, lines 1-22 and Example 1; U.S. Pat. No. 5,543,502 at column 2, lines 17-46; U.S. Pat. No. 5,171,844 at column 4, line 22 to column 5, line 36; U.S. Pat. No. 5,112, 950, col. 2, lines 55-68, Figure 2, and Example 1; U.S. Pat. No. 4,868,112, col. 2, line 2 to col. 19, line 21, and Table 2; U.S. Pat. No. 7,041,635, col. 2, line 1 to col. 3, line 19, col. 3, line 40 to col. 4, line 67, col. 7, line 43 to col. 8, line 26, and col. 11, line 5 to col. 13, line 39; or U.S. Pat. No. 6,458,563, col. 4, lines 25-53. In some embodiments, the B-domain-deleted FVIII has a deletion of most of the B domain, but still contains the amino-terminal sequence of the B domain necessary for proteolytic processing of the primary translation product into two polypeptide chains in vivo, as disclosed in WO 91 / 09122.In some embodiments, a B-domain deleted FVIII is constructed with a deletion of amino acids 747-1638, i.e., a virtually complete deletion of the B domain. Hoeben RC et al. J. Biol. Chem. 265(13):7318-7323 (1990). A B-domain deleted Factor VIII may also contain a deletion of amino acids 771-1666 or amino acids 868-1562 of FVIII. Meulien P. et al. Protein Eng. 2(4):301-6 (1988). Other B domain deletions that are part of the present invention include: amino acids 982 to 1562 or 760 to 1639 (Toole et al., Proc. Natl. Acad. Sci. USA (1986) 83, 5939-5942), 797 to 1562 (Eaton et al., Biochemistry (1986) 25:8343-8347), 741 to 1646 (Kaufman et al. (PCT Published Application No. WO 87 / 04187)), 747-1560 (Sarver et al., DNA (1987) 6:553-564), 741 to 1648 (Pasek (PCT Application No. 88 / 00831)), or 816 to 1598 or 741 to 1648 (Lagner (Behring Inst. Mitt. (1988) 82:16-25, EP 2004005)). 295597)). In a specific embodiment, the B domain deleted FVIII comprises a deletion of amino acid residues 746 to 1648 of mature FVIII. In another embodiment, the B domain deleted FVIII comprises a deletion of amino acid residues 745 to 1648 of mature FVIII. In some embodiments, BDD FVIII comprises a single chain FVIII (also referred to as rVIII - Single Chain and) containing a deletion corresponding to amino acids 765 to 1652 of mature full-length FVIII. ). See U.S. Patent No. 7,041,635.

[0565] In other embodiments, BDD FVIII includes a FVIII polypeptide comprising a fragment of the B domain that retains one or more N-linked glycosylation sites, such as residues 757, 784, 828, 900, 963, or optionally 943 of the amino acid sequence corresponding to the full-length FVIII sequence. Examples of B domain fragments include 226 amino acids or 163 amino acids of the B domain as disclosed in Miao, HZ et al., Blood 103(a):3412-3419 (2004); Kasuda, A et al., J. Thromb. Haemost. 6:1352-1359 (2008) and Pipe, SW et al., J. Thromb. Haemost. 9:2235-2242 (2011) (i.e., retaining the first 226 amino acids or 163 amino acids of the B domain). In still other embodiments, BDD FVIII further comprises a point mutation at residue 309 (from Phe to Ser) to improve expression of the BDD FVIII protein. See Miao, HZ et al., Blood 103(a):3412-3419 (2004). In still other embodiments, BDD FVIII includes a FVIII polypeptide containing a portion of the B domain, but without one or more furin cleavage sites (e.g., Arg1313 and Arg 1648). See Pipe, SW et al., J. Thromb. Haemost. 9:2235-2242 (2011). In some embodiments, BDD FVIII comprises a single-chain FVIII (also referred to as rVIII-SingleChain and rVIII-SingleChain) containing a deletion in amino acids 765 to 1652 corresponding to mature full-length FVIII. ). See U.S. Patent No. 7,041,635. Each of the foregoing deletions can be made in any FVIII sequence.

[0566] Numerous functional FVIII variants are known, as discussed above and below. Additionally, hundreds of non-functional mutations in FVIII have been identified in hemophilia patients, and it has been determined that the impact of these mutations on FVIII function is more due to its position within the three-dimensional structure of FVIII rather than the properties of substitution (Cutler et al., Hum. Mutat. 19: 274-8 (2002)), which is incorporated herein by reference in its entirety. Additionally, comparisons between FVIII from humans and other species have identified conserved residues (Cameron et al., Thromb. Haemost. 79: 317-22 (1998) that may be required for function; US 6,251,632), which is incorporated herein by reference in its entirety.

[0567] In some embodiments, the FVIII polypeptide comprises a FVIII variant or fragment thereof, wherein the FVIII variant or fragment thereof has FVIII activity. In some embodiments, the gene cassette encodes a full-length FVIII polypeptide. In other embodiments, the gene cassette encodes a B-domain deleted (BDD) FVIII polypeptide, wherein all or a portion of the B domain of FVIII is deleted. In a specific embodiment, the gene cassette encodes a polypeptide comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to SEQ ID NO: 106, 107, 109, 110, 111 or 112. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 17, or a fragment thereof. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 106, or a fragment thereof. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 107. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 109, or a fragment thereof. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 16. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 109.

[0568] In some embodiments, the gene cassette of the present disclosure encodes a FVIII polypeptide or fragment thereof comprising a signal peptide. In other embodiments, the gene cassette encodes a FVIII polypeptide lacking a signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.

[0569] In some embodiments, the gene cassette comprises a nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence is codon optimized. In certain embodiments, the gene cassette comprises a nucleotide sequence disclosed in International Application No. PCT / US2017 / 015879, which is incorporated herein by reference in its entirety. In some embodiments, the gene cassette comprises a nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence is codon optimized. In certain embodiments, the gene cassette comprises a nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence is codon optimized. In certain embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to a nucleotide sequence selected from SEQ ID NOs: 1-14. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 71. In some embodiments, the gene-cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19.

[0570] i. Codon-optimized nucleotide sequence encoding FVIII polypeptide

[0571] In some embodiments, the nucleic acid molecules of the present disclosure comprise a first ITR, a second ITR, and a gene cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein, wherein the first ITR and the second ITR are derived from an AAV genome, and wherein the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide. In some embodiments, the codon-optimized nucleotide sequence encodes a full-length FVIII polypeptide. In other embodiments, the codon-optimized nucleotide sequence encodes a B-domain deleted (BDD) FVIII polypeptide, wherein all or a portion of the B domain of FVIII is deleted. In a specific embodiment, the codon-optimized nucleotide sequence encodes a polypeptide or fragment thereof comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 17. In one embodiment, the codon-optimized nucleotide sequence encodes a polypeptide or fragment thereof having the amino acid sequence of SEQ ID NO: 17.

[0572] In some embodiments, the codon-optimized nucleotide sequence encodes a FVIII polypeptide or fragment thereof comprising a signal peptide. In other embodiments, the codon-optimized sequence encodes a FVIII polypeptide lacking a signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.

[0573] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 58 to 1791 of SEQ ID NO: 3 or (ii) nucleotides 58 to 1791 of SEQ ID NO: 4; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In a specific embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 1791 of SEQ ID NO: 3. In another embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 4. In other embodiments, the first nucleotide sequence comprises nucleotides 58-1791 of SEQ ID NO: 3 or nucleotides 58-1791 of SEQ ID NO: 4.

[0574] In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1-1791 of SEQ ID NO: 3 or (ii) nucleotides 1-1791 of SEQ ID NO: 4; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1-1791 of SEQ ID NO: 3 or nucleotides 1-1791 of SEQ ID NO: 4. In another embodiment, the second nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 3 or nucleotides 1792-4374 of SEQ ID NO: 4. In a specific embodiment, the second nucleotide sequence comprises nucleotides 1792-4374 of SEQ ID NO: 3 or nucleotides 1792-4374 of SEQ ID NO: 4. In yet another embodiment, the second nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:3 or nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:4 (i.e., nucleotides 1792-4374 of SEQ ID NO:3 or nucleotides 1792-4374 of SEQ ID NO:4 without the nucleotides encoding the B domain or a B domain fragment). In a specific embodiment, the second nucleotide sequence comprises nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 3 or nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 1792-4374 of SEQ ID NO: 3 or nucleotides 1792-4374 of SEQ ID NO: 4 without nucleotides encoding a B domain or a B domain fragment).

[0575] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5 or (ii) nucleotides 1792-4374 of SEQ ID NO: 6; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 5. In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 6. In a specific embodiment, the second nucleic acid sequence comprises nucleotides 1792-4374 of SEQ ID NO: 5 or nucleotides 1792-4374 of SEQ ID NO: 6. In some embodiments, the first nucleic acid sequence listed above that is linked to the second nucleic acid sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 5 or nucleotides 58-1791 of SEQ ID NO: 6. In other embodiments, the first nucleic acid sequence listed above that is linked to the second nucleic acid sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1-1791 of SEQ ID NO:5 or nucleotides 1-1791 of SEQ ID NO:6.

[0576] In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without nucleotides encoding the B domain or a B domain fragment) or (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without nucleotides encoding the B domain or a B domain fragment); and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without nucleotides encoding the B domain or a B domain fragment). In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without nucleotides encoding the B domain or a B domain fragment). In a specific embodiment, the second nucleic acid sequence comprises nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 or nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 or nucleotides 1792-4374 of SEQ ID NO: 6 without nucleotides encoding the B domain or a B domain fragment). In some embodiments, the first nucleic acid sequence listed above, linked to the second nucleic acid sequence, has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 5 or nucleotides 58-1791 of SEQ ID NO: 6.In other embodiments, the first nucleic acid sequence listed above that is linked to the second nucleic acid sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1-1791 of SEQ ID NO:5 or nucleotides 1-1791 of SEQ ID NO:6.

[0577] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 1, (ii) nucleotides 58-1791 of SEQ ID NO: 2, (iii) nucleotides 58-1791 of SEQ ID NO: 70, or (iv) nucleotides 58-1791 of SEQ ID NO: 71; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In other embodiments, the first nucleotide sequence comprises nucleotides 58-1791 of SEQ ID NO: 1, nucleotides 58-1791 of SEQ ID NO: 2, (iii) nucleotides 58-1791 of SEQ ID NO: 70, or (iv) nucleotides 58-1791 of SEQ ID NO: 71.

[0578] In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1-1791 of SEQ ID NO: 1, (ii) nucleotides 1-1791 of SEQ ID NO: 2, (iii) nucleotides 1-1791 of SEQ ID NO: 70, or (iv) nucleotides 1-1791 of SEQ ID NO: 71; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1-1791 of SEQ ID NO: 1, nucleotides 1-1791 of SEQ ID NO: 2, (iii) nucleotides 1-1791 of SEQ ID NO: 70, or (iv) nucleotides 1-1791 of SEQ ID NO: 71. In another embodiment, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 1, nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In a specific embodiment, the second nucleotide sequence linked to the first nucleotide sequence comprises (i) nucleotides 1792-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In other embodiments, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71.In one embodiment, the second nucleotide sequence comprises (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71.

[0579] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In a specific embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In some embodiments, the codon-optimized sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence is identical to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1792-4374 of SEQ ID NO: 1, nucleotides 1792-4374 of SEQ ID NO: 2, nucleotides 1792-4374 of SEQ ID NO: 70, or nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71, without nucleotides encoding a B domain or a B domain fragment. NO:71) having at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity; and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity.In one embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1792-4374 of SEQ ID NO: 1, nucleotides 1792-4374 of SEQ ID NO: 2, nucleotides 1792-4374 of SEQ ID NO: 70, or nucleotides 1792-4374 of SEQ ID NO: 71, without nucleotides encoding a B domain or a B domain fragment).

[0580] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 1. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 2277 and 2320 to 4374 of SEQ ID NO: 1 (i.e., nucleotides 58 to 4374 of SEQ ID NO: 1 excluding nucleotides encoding the B domain or a B domain fragment). In other embodiments, the nucleic acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 1. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 1-4374 of SEQ ID NO: 1 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 1.

[0581] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 2. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 2. In other embodiments, the nucleic acid sequence has at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 2. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 2 (i.e., nucleotides 58-4374 of SEQ ID NO: 2 without nucleotides encoding the B domain or a B domain fragment), or nucleotides 58 to 4374 of SEQ ID NO: 2. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 2 (i.e., nucleotides 1-4374 of SEQ ID NO: 2 without nucleotides encoding the B domain or a B domain fragment), or nucleotides 1 to 4374 of SEQ ID NO: 2.

[0582] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 70. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 2277 and 2320 to 4374 of SEQ ID NO: 70 (i.e., nucleotides 58 to 4374 of SEQ ID NO: 70 without nucleotides encoding the B domain or a B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 70. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 58-4374 of SEQ ID NO: 70 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 70. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 1-4374 of SEQ ID NO: 70 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 70.

[0583] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 71. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 2277 and 2320 to 4374 of SEQ ID NO: 71 (i.e., nucleotides 58 to 4374 of SEQ ID NO: 71 excluding nucleotides encoding the B domain or a B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 71. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 71 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 71. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1-4374 of SEQ ID NO: 71 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 71.

[0584] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 3. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 2277 and 2320 to 4374 of SEQ ID NO: 3 (i.e., nucleotides 58 to 4374 of SEQ ID NO: 3 excluding nucleotides encoding the B domain or a B domain fragment). In certain embodiments, the nucleic acid sequence has at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 3. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 58-4374 of SEQ ID NO: 3 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 3. In still other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 1-4374 of SEQ ID NO: 3 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 3.

[0585] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 4. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 4 (i.e., nucleotides 58 to 4374 of SEQ ID NO: 4 excluding nucleotides encoding the B domain or a B domain fragment). In other embodiments, the nucleic acid sequence has at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 4. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 58-4374 of SEQ ID NO: 4 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 4. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 1-4374 of SEQ ID NO: 4 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 4.

[0586] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 5. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 2277 and 2320 to 4374 of SEQ ID NO: 5 (i.e., nucleotides 58 to 4374 of SEQ ID NO: 5 excluding nucleotides encoding the B domain or a B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 5. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 58-4374 of SEQ ID NO: 5 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 5. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1-4374 of SEQ ID NO: 5 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 5.

[0587] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 6. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 2277 and 2320 to 4374 of SEQ ID NO: 6 (i.e., nucleotides 58 to 4374 of SEQ ID NO: 6 excluding nucleotides encoding the B domain or a B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 6. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 58-4374 of SEQ ID NO: 6 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 6. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1-4374 of SEQ ID NO: 6 without nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 6.

[0588] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleic acid sequence encoding a signal peptide. In certain embodiments, the signal peptide is a FVIII signal peptide. In some embodiments, the nucleic acid sequence encoding the signal peptide is codon-optimized. In a specific embodiment, the nucleic acid sequence encoding the signal peptide has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to (i) nucleotides 1 to 57 of SEQ ID NO: 1; (ii) nucleotides 1 to 57 of SEQ ID NO: 2; (iii) nucleotides 1 to 57 of SEQ ID NO: 3; (iv) nucleotides 1 to 57 of SEQ ID NO: 4; (v) nucleotides 1 to 57 of SEQ ID NO: 5; (vi) nucleotides 1 to 57 of SEQ ID NO: 6; (vii) nucleotides 1 to 57 of SEQ ID NO: 70; (viii) nucleotides 1 to 57 of SEQ ID NO: 71; or (ix) nucleotides 1 to 57 of SEQ ID NO: 68.

[0589] SEQ ID NOs: 1-6, 70, and 71 are optimized versions of the starting or "parental" or "wild-type" FVIII nucleotide sequence, SEQ ID NO: 16. SEQ ID NO: 16 encodes a human FVIII with a B-domain deletion. Although SEQ ID NOs: 1-6, 70, and 71 are derived from a specific B-domain deleted form of FVIII (SEQ ID NO: 16), it should be understood that the present disclosure also encompasses optimized versions of nucleic acids encoding other forms of FVIII. For example, other forms of FVIII can include full-length FVIII, other B-domain deletions of FVIII (described herein), or other FVIII fragments that retain FVIII activity.

[0590] In one embodiment, the gene cassette comprises a FVIII construct comprising a polynucleotide sequence as listed in Tables 2A-2F. In one embodiment, the gene cassette comprises a FVIII construct comprising a polynucleotide sequence as described in Table 2A. In one embodiment, the gene cassette comprises a FVIII construct comprising a polynucleotide sequence as described in Table 2B. In one embodiment, the gene cassette comprises a FVIII construct comprising a polynucleotide sequence as described in Table 2C. In one embodiment, the gene cassette comprises a FVIII construct comprising a polynucleotide sequence as described in Table 2D. In one embodiment, the gene cassette comprises a FVIII construct comprising a polynucleotide sequence as described in Table 2E. In one embodiment, the gene cassette comprises a FVIII construct comprising a polynucleotide sequence as described in Table 2F.

[0591] In certain embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 179, 182, 189, or 194. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 179. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 182. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 189. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 194. In some embodiments, the isolated nucleic acid molecule retains the ability to express a functional FVIII protein.

[0592] Table 2A: Exemplary AAV-FVIII constructs (nucleotides 1-6526; SEQ ID NO: 110)

[0593]

[0594]

[0595]

[0596]

[0597] Table 2B: Exemplary B19-FVIII constructs with B19d135 ITR (nucleotides 1-6762; SEQ ID NO: 179)

[0598]

[0599]

[0600]

[0601]

[0602]

[0603]

[0604]

[0605] Table 2C: Exemplary GPV-FVIII constructs with GPVd162 ITR (nucleotides 1-6830; SEQ ID NO: 182)

[0606]

[0607]

[0608]

[0609]

[0610]

[0611]

[0612] Table 2D: Exemplary B19-FVIII constructs with full-length B19 ITR (nucleotides 1-7032; SEQ ID NO: 189)

[0613]

[0614]

[0615]

[0616]

[0617]

[0618]

[0619]

[0620] Table 2E: Exemplary AAV-FVIII constructs (nucleotides 1-6824; SEQ ID NO: 190)

[0621]

[0622]

[0623]

[0624]

[0625]

[0626]

[0627] Table 2F: Exemplary GPV-FVIII constructs with full-length GPV ITR (nucleotides 1-7154; SEQ ID NO: 194)

[0628]

[0629]

[0630]

[0631]

[0632]

[0633]

[0634]

[0635] In one embodiment, the gene cassette comprises a phenylalanine hydroxylase (PAH) construct comprising a polynucleotide sequence as listed in Tables 10A and 10B. In one embodiment, the gene cassette comprises a PAH construct comprising a polynucleotide sequence as described in Table 10A. In one embodiment, the gene cassette comprises a PAH construct comprising a polynucleotide sequence as described in Table 10B.

[0636] In certain embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 197 or 198. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 197. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 198. In some embodiments, the isolated nucleic acid molecule retains the ability to express a functional phenylalanine hydroxylase.

[0637] A. Codon Adaptation Index

[0638] In one embodiment, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the human codon adaptation index of the codon-optimized nucleotide sequence is increased relative to SEQ ID NO: 16. For example, the codon-optimized nucleotide sequence can have a human codon adaptation index of at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%), at least about 0.91 (91%), at least about 0.92 (92%), at least about 0.93 (93%), at least about 0.94 (94%), at least about 0.95 (95%), at least about 0.96 (96%), at least about 0.97 (97%), at least about 0.98 (98%), at least about 0.99 (99%), at least about 0.91 (99%), at least about 0.92 (99%), at least about 0.93 (99%), at least about 0.94 (99%), at least about 0.95 (99%), at least about 0.96 (99%), at least about 0.97 (99%), at least about 0.98 (99%), at least about In some embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index of at least about 0.88 (88%). In other embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index of at least about 0.91 (91%). In other embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index of at least about 0.91 (97%). In other embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index of at least about 0.91 (97%).

[0639] In a specific embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence is identical to (i) nucleotides 58 to 1791 of SEQ ID NO: 3; (ii) nucleotides 1 to 1791 of SEQ ID NO: 3; (iii) nucleotides 58 to 1791 of SEQ ID NO: 4; or (iv) nucleotides 58 to 1791 of SEQ ID NO: 5. NO:4 has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO:16. In some embodiments, the nucleotide sequence has a human codon adaptation index of at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%), or at least about 0.91 (91%). In a specific embodiment, the nucleotide sequence has a human codon adaptation index of at least about 0.88 (88%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about 0.91 (91%).

[0640] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 or (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the human codon adaptation index of the nucleotide sequence is relative to that of SEQ ID NO: NO: 16 increases. In some embodiments, the nucleotide sequence has a human codon adaptation index of at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), or at least about 0.88 (88%). In a specific embodiment, the nucleotide sequence has a human codon adaptation index of at least about 0.83 (83%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about 0.88 (88%).

[0641] In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 without nucleotides encoding the B domain or a B domain fragment); and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO: 16. In some embodiments, the nucleotide sequence has a human codon adaptation index of at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), or at least about 0.88 (88%). In a specific embodiment, the nucleotide sequence has a human codon adaptation index of at least about 0.75 (75%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about 0.83 (83%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about 0.88 (88%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about 0.91 (91%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about 0.97 (97%).

[0642] In some embodiments, the codon-optimized nucleotide sequences encoding FVIII polypeptides of the present disclosure have an increased optimal codon frequency (FOP) relative to SEQ ID NO: 16. In certain embodiments, the FOP of the codon-optimized nucleotide sequence encoding a FVIII polypeptide is at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 64, at least about 65, at least about 70, at least about 75, at least about 79, at least about 80, at least about 85, or at least about 90.

[0643] In other embodiments, the codon-optimized nucleotide sequences encoding the FVIII polypeptides of the present disclosure have an increased relative synonymous codon usage (RCSU) relative to SEQ ID NO: 16. In some embodiments, the RCSU of the isolated nucleic acid molecule is greater than 1.5. In other embodiments, the RCSU of the isolated nucleic acid molecule is greater than 2.0. In certain embodiments, the RCSU of the isolated nucleic acid molecule is at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2.0, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, or at least about 2.7.

[0644] In still other embodiments, the codon-optimized nucleotide sequences encoding FVIII polypeptides of the present disclosure have a reduced significant codon number relative to SEQ ID NO: 16. In some embodiments, the codon-optimized nucleotide sequences encoding FVIII polypeptides have a significant codon number of less than about 50, less than about 45, less than about 40, less than about 35, less than about 30, or less than about 25. In a specific embodiment, the isolated nucleic acid molecule has a significant codon number of about 40, about 35, about 30, about 25, or about 20.

[0645] BG / C content optimization

[0646] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains a higher percentage of G / C nucleotides as compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, or at least about 60%.

[0647] In a specific embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58 to 1791 of SEQ ID NO: 3; (ii) nucleotides 1 to 1791 of SEQ ID NO: 3; (iii) nucleotides 58 to 1791 of SEQ ID NO: 4; or (iv) nucleotides 1 to 1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a G / C content of at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, or at least about 58%. In a specific embodiment, the nucleotide sequence encoding a polypeptide having FVIII activity has a G / C content of at least about 58%.

[0648] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence is identical to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without nucleotides encoding the B domain or a B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without nucleotides encoding the B domain or a B domain fragment). NO:6) having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains a greater percentage of G / C nucleotides than the percentage of G / C nucleotides in SEQ ID NO:16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, or at least about 57%. In a specific embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 52%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 55%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 57%.

[0649] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, or 71 without nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71; In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 45%. In a specific embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 52%. In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 55%. In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 57%. In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 58%. In still another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 60%.

[0650] "G / C content" (or guanine-cytosine content) or "percentage of G / C nucleotides" refers to the percentage of nitrogenous bases in a DNA molecule that are guanine or cytosine. G / C content can be calculated using the following formula:

[0651]

[0652] The G / C content of human genes is highly heterogeneous, with some genes having a G / C content as low as 20% and others having a G / C content as high as 95%. Generally, genes rich in G / C have higher expression. In fact, it has been shown that increasing the G / C content of a gene can lead to increased gene expression, primarily due to increased transcription and higher steady-state mRNA levels. See Kudla et al., PLoS Biol., 4(6):e180 (2006).

[0653] C. Matrix attachment region-like sequence

[0654] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a MARS / ARS sequence.

[0655] In a specific embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58 to 1791 of SEQ ID NO: 3; (ii) nucleotides 1 to 1791 of SEQ ID NO: 3; (iii) nucleotides 58 to 1791 of SEQ ID NO: 4; or (iv) nucleotides 1 to 1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence has at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: NO: 16 contains fewer MARS / ARS sequences. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains no more than 6, no more than 5, no more than 4, no more than 3, or no more than 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains no more than 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a MARS / ARS sequence.

[0656] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence is identical to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without nucleotides encoding a B domain or a B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without nucleotides encoding a B domain or a B domain fragment). NO:6 (nucleotides 1792-4374) having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains no more than 6, no more than 5, no more than 4, no more than 3, or no more than 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains no more than 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a MARS / ARS sequence.

[0657] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises the same sequence as (i) nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71, or (ii) nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 (i.e., SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 without nucleotides encoding the B domain or B domain fragment). NO: 1, 2, 3, 4, 5, 6, 70 or 71) has at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity; and wherein the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains no more than 6, no more than 5, no more than 4, no more than 3 or no more than 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains no more than 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a MARS / ARS sequence.

[0658] AT-rich components have been identified in the human FVIII nucleotide sequence that share sequence similarity with the autonomously replicating sequence (ARS) and matrix attachment region (MAR) of Saccharomyces cerevisiae. (Fallux et al., Mol. Cell. Biol. 16:4264-4272 (1996)). One of these components has been shown to bind to nuclear factor in vitro and repress the expression of the chloramphenicol acetyltransferase (CAT) reporter gene. Ibid. It has been hypothesized that these sequences may contribute to transcriptional repression of the human FVIII gene. Therefore, in one embodiment, all MAR / ARS sequences are eliminated in the codon-optimized nucleotide sequence encoding the FVIII polypeptide of the present disclosure. In the parental FVIII sequence (SEQ ID NO: 16), there are 4 MAR / ARS ATATTT sequences (SEQ ID NO: 21) and 3 MAR / ARS AAATAT sequences (SEQ ID NO: 22). All of these sites were mutated to disrupt the MAR / ARS sequence in the optimized FVIII sequences (SEQ ID NOs: 1-6). The positions of each of these components in the optimized sequences and the sequences of the corresponding nucleotides are shown in Table 3 below.

[0659] Table 3: Overview of changes in repressor components

[0660]

[0661]

[0662] D. Destabilizing sequence

[0663] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains no more than 9, no more than 8, no more than 7, no more than 6, or no more than 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains no more than 4, no more than 3, no more than 2, or no more than 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a destabilizing element.

[0664] In a specific embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58 to 1791 of SEQ ID NO: 3; (ii) nucleotides 1 to 1791 of SEQ ID NO: 3; (iii) nucleotides 58 to 1791 of SEQ ID NO: 4; or (iv) nucleotides 1 to 1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence has at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: NO: 16 contains fewer destabilizing elements. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains no more than 9, no more than 8, no more than 7, no more than 6, or no more than 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains no more than 4, no more than 3, no more than 2, or no more than 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a destabilizing element.

[0665] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence is identical to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without nucleotides encoding a B domain or a B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without nucleotides encoding a B domain or a B domain fragment). In some embodiments, the nucleotide sequence encoding the polypeptide having FVIII activity comprises at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 6 (nucleotides 1792-4374); wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains at most 9, at most 8, at most 7, at most 6, or at most 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a destabilizing element.

[0666] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises (i) nucleotides 58-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, excluding nucleotides encoding a B domain or a B domain fragment). In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least ab...

Claims

1. A nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanked by a gene cassette, wherein the gene cassette comprises a heterologous polynucleotide sequence, wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence described in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

2. The nucleic acid molecule of claim 1, wherein the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 180, and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO:

181.

3. The nucleic acid molecule of claim 1, wherein the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 183, and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO:

184.

4. The nucleic acid molecule of claim 1, wherein the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 185, and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO:

186.

5. The nucleic acid molecule of claim 1, wherein the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 187, and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO:

188.

6. The nucleic acid molecule of claim 1, wherein the first ITR and / or the second ITR consists of the nucleotide sequence described in SEQ ID NO: 180, 181, 183, 184, 185, 186, 187 or 188.

7. The nucleic acid molecule of claim 1, wherein the first ITR and the second ITR are reverse complements of each other.

8. The nucleic acid molecule according to any one of claims 1 to 7, further comprising a promoter.

9. The nucleic acid molecule of claim 8, wherein the promoter is a tissue-specific promoter.

10. The nucleic acid molecule of claim 8 or 9, wherein the promoter drives expression of the heterologous polynucleotide sequence in an organ selected from the group consisting of muscle, central nervous system (CNS), eye, liver, heart, kidney, pancreas, lung, skin, bladder, urinary tract, or any combination thereof.

Citation Information

Patent Citations

  • Factor VIII:C-like molecule with a coagulant activity

    EP0295597A2

  • Lentiviral vectors encoding clotting factors for gene therapy

    EP1395293A1

  • Serum albumin binding moieties

    US20030069395A1

  • Central airway administration for systemic delivery of therapeutics

    US20030235536A1

  • Fc Variants Having Increased Affinity for FcyRIIb

    US20070231329A1