NUCLEIC ACID MOLECULES AND THEIR USES FOR NON-VIRAL GENE THERAPY.

MX431897BActive Publication Date: 2026-02-25BIOVERATIV THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
MX2021001599
Authority / Receiving Office
MX · MX
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-08-09
Filing Date
2021-02-09
Publication Date
2026-02-25
Estimated Expiration
2039-08-09

AI Technical Summary

Technical Problem

Conventional adeno-associated virus (AAV) gene therapy vectors face limitations due to their limited viral packaging capacity and immunogenic properties, which can trigger immune responses and reduce clinical efficacy, especially in patients with pre-existing immunity.

Method used

A nucleic acid molecule comprising a first and second inverted terminal repeat (ITR) flanking a genetic cassette with a heterologous polynucleotide sequence, where the ITRs are at least 75% to 100% identical to specific sequences, is used to enhance gene expression and persistence, avoiding the limitations of AAV vectors.

Benefits of technology

The solution enables efficient and persistent expression of therapeutic proteins or miRNAs, overcoming the packaging capacity and immunogenicity issues of AAV vectors, and achieving significant increases in clotting factor activity, such as factor VIII, in treated subjects.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This disclosure provides nucleic acid molecules comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette encoding a target sequence. In some embodiments, the target sequence encodes a microRNA and / or a therapeutic protein. In certain embodiments, the therapeutic protein comprises a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, and a combination thereof. In some embodiments, the first ITR and / or the second ITR is an ITR of a non-adeno-associated virus (AAV). This disclosure also provides methods of treating a metabolic liver disorder in a subject comprising administering to the subject the nucleic acid molecule or a polypeptide so encoded.
Need to check novelty before this filing date? Find Prior Art

Description

NUCLEIC ACID MOLECULES AND USES THEREOF FOR NON-VIRAL GENE THERAPYRELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application Serial No.62 / 716,826, filed August 9, 2018, the entire disclosure of which is hereby incorporated herein by reference.REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY

[0002] The content of the electronically submitted sequence listing in ASCII text file(Name: SA9-465PC_SL_ST25.txt; Size: 460,648 bytes; and Date of Creation: August 8, 2019) is incorporated herein by reference in its entirety.BACKGROUND OF THE DISCLOSURE

[0003] Gene therapy offers the potential for a lasting means of treating a variety of diseases. In the past, many gene therapy treatments typically relied on the use of viruses. There are numerous viral agents that could be selected for this purpose, each with distinct properties that would make them more or less suitable for gene therapy. Zhou et al., Adv Drug Deliv Rev. 106(Pt A):3-26, 2016. However, the undesired properties of some viral vectors, including their immunogenic profiles ortheir propensity to cause cancer, have resulted in clinical safety concerns and, until recently, limited their clinical use to certain applications, for example, vaccines and oncolytic strategies. Cotter et al., Front Biosci. 10:1098-105 (2005).

[0004] Adeno-associated virus (AAV) is one of the most commonly investigated gene therapy vectors. AAV is a protein shell surrounding and protecting a small, single-stranded DNA genome of approximately 4.8 kilobases (kb). Naso et al., BioDrugs, 31 (4): 317-334, 2017. AAV belongs to the parvovirus family and is dependent on co-infection with other viruses, mainly adenoviruses, in order to replicate. Id. Its single-stranded genome contains three genes, Rep (Replication), Cap (Capsid), and aap (Assembly). Id. These coding sequences are flanked by inverted terminal repeats (ITRs) that are required for genome replication and packaging. Id. The two cis-acting AAV ITRs are approximately 145 nucleotides in length with interrupted palindromic sequences that can fold into T shaped hairpin structures that function as primers during initiation of DNA replication.

[0005] The use of conventional AAV as a gene delivery vector has certain drawbacks, however. One of the major drawbacks is associated with the AAV's limited viral packaging capacity of about 4.5 kb of heterologous DNA. (Dong et al., Hum Gene Ther. 7(17): 2101-12, 1996). In addition, administration of AAV vectors can induce an immune response in humans.Although AAV has been shown to be less immunogenic than some other viruses (i.e. adenovirus), the capsid proteins can trigger various components of the human immune system. See Naso et al., 2017. AAV is a common virus in the human population, and most people have been exposed to AAV, accordingly most people have already developed an immune response against the particular variants to which they had previously been exposed. This pre-existing adaptive response can include neutralizing antibodies (NAbs) and T cells that could diminish the clinical efficacy of subsequent re-infections with AAV and / or the elimination of cells that have been transduced, which may disqualify patients with pre-existing anti-AAV immunity to AW based gene therapy treatment. Furthermore, evidence suggests that the T-shaped hairpin loops of AAV ITRs are susceptible to inhibition by host cell proteins / protein complexes that bind the T- shaped hairpin structures of AAV ITRs. See, e.g., Zhou et al., Scientific Reports 7:5432 (July 14, 2017).

[0006] Thus, there exists a need in the art to efficiently and persistently express target sequences, e.g., therapeutic proteins and / or miRNAs, in in vitro and in vivo settings, while avoiding some of the unintended consequences and limitations of existing AAV vector technology.SUMMARY OF THE DISCLOSURE

[0007] In certain aspects, a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof, is provided

[0008] In certain exemplary embodiments, the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 180 and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 181. In certain exemplary embodiments, the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 183 and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 184. In certain exemplary embodiments, the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 185 and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 186. In certain exemplary embodiments, the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 187 and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 188.

[0009] In certain exemplary embodiments, the first ITR and / or the second ITR consists of a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188. In certain exemplary embodiments, the first ITR and the second ITR are reverse complements of each other.

[0010] In certain exemplary embodiments, the nucleic acid molecule further comprises a promoter. In certain exemplary embodiments, the promoter is a tissue-specific promoter. In certain exemplary embodiments, the promoter drives expression of the heterologous polynucleotide sequence in an organ selected from the muscle, central nervous system (CNS), ocular, liver, heart, kidney, pancreas, lungs, skin, bladder, urinary tract, or any combination thereof. In certain exemplary embodiments, the promoter drives expression of the heterologous polynucleotide sequence in hepatocytes, endothelial cells, cardiac muscle cells, skeletal muscle cells, sinusoidal cells, afferent neurons, efferent neurons, interneurons, glial cells, astrocytes, oligodendrocytes, microglia, ependymal cells, lung epithelial cells, Schwann cells, satellite cells, photoreceptor cells, retinal ganglion cells, or any combination thereof. In certain exemplary embodiments, the promoter is positioned 5' to the heterologous polynucleotide sequence. In certain exemplary embodiments, the promoter is selected from the group consisting of a mouse thyretin promoter (mTTR), an endogenous human factor VIII promoter (F8), a human alpha-1- antitrypsin promoter (hAAT), a human albumin minimal promoter, a mouse albumin promoter, a tristetraprolin (TTP) promoter, a CASI promoter, a CAG promoter, a cytomegalovirus (CMV) promoter, a1 -antitrypsin (AAT), muscle creatine kinase (MCK), myosin heavy chain alpha (aMHC), myoglobin (MB), desmin (DES), SPc5-12, 2R5Sc5-12, dMCK, tMCK, and a phosphoglycerate kinase (PGK) promoter.

[0011] In certain exemplary embodiments, the heterologous polynucleotide sequence further comprises an intronic sequence. In certain exemplary embodiments, the intronic sequence is positioned 5' to the heterologous polynucleotide sequence. In certain exemplary embodiments, the intronic sequence is positioned 3' to the promoter. In certain exemplary embodiments, the intronic sequence comprises a synthetic intronic sequence. In certain exemplary embodiments, the intronic sequence comprises SEQ ID NO: 1 15 or 192.

[0012] In certain exemplary embodiments, the genetic cassette further comprises a post- transcriptional regulatory element. In certain exemplary embodiments, the post-transcriptional regulatory element is positioned 3' to the heterologous polynucleotide sequence. In certain exemplary embodiments, the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, or any combination thereof. In certain exemplary embodiments, the microRNA binding site comprises a binding site to miR142-3p.

[0013] In certain exemplary embodiments, the genetic cassette further comprises a3'UTR poly(A) tail sequence. In certain exemplary embodiments, the 3'UTR poly(A) tail sequence is selected from the group consisting of bGH poly(A), actin poly(A), hemoglobin poly(A), and any combination thereof. In certain exemplary embodiments, the 3'UTR poly(A) tail sequence comprises bGH poly(A).

[0014] In certain exemplary embodiments, the genetic cassette further comprises an enhancer sequence. In certain exemplary embodiments, the enhancer sequence is positioned between the first ITR and the second ITR.

[0015] In certain exemplary embodiments, the nucleic acid molecule comprises from 5’ to 3’: the first ITR, the genetic cassette, and the second ITR; wherein the genetic cassette comprises a tissue-specific promoter sequence, an intronic sequence, the heterologous polynucleotide sequence, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence. In certain exemplary embodiments, the genetic cassette comprises from 5’ to 3’: a tissue-specific promoter sequence, an intronic sequence, the heterologous polynucleotide sequence, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence. In certain exemplary embodiments, the tissue specific promoter sequence comprises a TTT promoter; the intron is a synthetic intron; the post-transcriptional regulatory element comprises WPRE; andthe 3'UTR poly(A) tail sequence comprises bGHpA.

[0016] In certain exemplary embodiments, the genetic cassette comprises a single stranded nucleic acid. In certain exemplary embodiments, the genetic cassette comprises a double stranded nucleic acid.

[0017] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or any combination thereof.

[0018] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a growth factor selected from the group consisting of adrenomedullin (AM), angiopoietin (Ang), autocrine motility factor, a bone morphogenetic protein (BMP) (e.g. BMP2, BMP4, BMP5, BMP7), a ciliary neurotrophic factor family member (e.g., ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6)), a colony-stimulating factor (e.g., macrophage colony-stimulating factor (m-CSF), granulocyte colony-stimulating factor (G-CSF), granulocyte macrophage colony-stimulating factor (GM-CSF)), an epidermal growth factor (EGF), an ephrin (e.g., ephrin A1 , ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1 , ephrin B2, ephrin B3), erythropoietin (EPO), a fibroblast growth factor (FGF) (e.g. , FGF1 , FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF1 1 , FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21 , FGF22, FGF23), foetal bovine somatotrophin (FBS), a GDNFfamily member ( e.g . , glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, artemin), growth differentiation factor-9 (GDF9), hepatocyte growth factor (HGF), hepatoma- derived growth factor (HDGF), insulin, an insulin-like growth factors (e.g., insulin-like growth factor-1 (IGF- 1 ) or IGF-2, an interleukin (IL) (e.g., IL-1 , IL-2, IL-3, IL-4, IL-5, IL-6, IL-7), keratinocyte growth factor (KGF), migration-stimulating factor (MSF), macrophage-stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)), myostatin (GDF-8), a neuregulin (e.g., neuregulin 1 (NRG1), NRG2, NRG3, NRG4), a neurotrophin (e.g., brain-derived neurotrophic factor (BDNF), nerve growth factor (NGF), a neurotrophin-3 (NT-3), NT-4, placental growth factor (PGF), platelet-derived growth factor (PDGF), renalase (RNLS), T-cell growth factor (TCGF), thrombopoietin (TPO), a transforming growth factor (e.g., transforming growth factor alpha (TGF-a), TGF-b, tumor necrosis factor-alpha (TNF-a), and vascular endothelial growth factor (VEGF), and any combination thereof.

[0019] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a hormone.

[0020] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a cytokine.

[0021] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes an antibody or a fragment thereof.

[0022] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a gene selected from dystrophin X-linked, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1 , FXN (frataxin), GUCY2D, RS1 , CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha- glucosidase), and any combination thereof.

[0023] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a microRNA (miRNA). In certain exemplary embodiments, the miRNA down regulates the expression of a target gene selected from SOD1 , HTT, RHO, and any combination thereof

[0024] In certain exemplary embodiments, the heterologous polynucleotide sequence encodes a clotting factor selected from the group consisting of factor I (FI), factor II (Fll), factor III (Fill), factor IV (FVI), factor V (FV), factor VI (FVI), factor VII (FVII), factor VIII (FVIII), factor IX (FIX), factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FVIII), Von Willebrand factor (VWF), prekallikrein, high-molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, Protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator(tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), and any combination thereof.

[0025] In certain exemplary embodiments, the clotting factor is FVIII. In certain exemplary embodiments, the FVIII comprises full-length mature FVIII. In certain exemplary embodiments, the FVIII comprises an amino acid sequence at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to an amino acid sequence having SEQ ID NO: 106.

[0026] In certain exemplary embodiments, the FVIII comprises A1 domain, A2 domain,A3 domain, C1 domain, C2 domain, and a partial or no B domain. In certain exemplary embodiments, the FVIII comprises an amino acid sequence at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO:109.

[0027] In certain exemplary embodiments, the clotting factor comprises a heterologous moiety. In certain exemplary embodiments, the heterologous moiety is selected from the group consisting of albumin or a fragment thereof, an immunoglobulin Fc region, the C-terminal peptide (CTP) of the b subunit of human chorionic gonadotropin, a PAS sequence, a HAP sequence, a transferrin or a fragment thereof, an albumin-binding moiety, a derivative thereof, or any combination thereof. In certain exemplary embodiments, the heterologous moiety is linked to the N-terminus or the C-terminus of the FVIII or inserted between two amino acids in the FVIII. In certain exemplary embodiments, the heterologous moiety is inserted between two amino acids at one or more insertion site selected from the insertion sites listed in Table 4.

[0028] In certain exemplary embodiments, the FVIII further comprises A1 domain, A2 domain, C1 domain, C2 domain, an optional B domain, and a heterologous moiety, wherein the heterologous moiety is inserted immediately downstream of amino acid 745 corresponding to mature FVIII (SEQ ID NO:106).

[0029] In certain exemplary embodiments, the FVIII further comprises an FcRn binding partner. In certain exemplary embodiments, the FcRn binding partner comprises an Fc region of an immunoglobulin constant domain.

[0030] In certain exemplary embodiments, the nucleic acid sequence encoding the FVIII is codon optimized. In certain exemplary embodiments, the nucleic acid sequence encoding the FVIII is codon optimized for expression in a human.

[0031] In certain exemplary embodiments, the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%,at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to a nucleotide sequence of SEQ ID NO: 107.

[0032] In certain exemplary embodiments, the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO: 71.

[0033] In certain exemplary embodiments, the heterologous polynucleotide sequence is codon optimized. In certain exemplary embodiments, the heterologous polynucleotide sequence is codon optimized for expression in a human.

[0034] In certain exemplary embodiments, the nucleic acid molecule is formulated with a delivery agent. In certain exemplary embodiments, the delivery agent comprises a lipid nanoparticle. In certain exemplary embodiments, the delivery agent is selected from the group consisting of liposomes, non-lipid polymeric molecules, and endosomes, and any combination thereof.

[0035] In certain exemplary embodiments, the nucleic acid molecule is formulated for intravenous, transdermal, intradermal, subcutaneous, pulmonary, or oral delivery, or any combination thereof. In certain exemplary embodiments, the nucleic acid molecule is formulated for intravenous delivery.

[0036] In certain aspects, a vector comprising a nucleic acid molecule as described herein, is provided.

[0037] In certain aspects, a host cell comprising a nucleic acid molecule as described herein, is provided.

[0038] In certain aspects, a pharmaceutical composition comprising a nucleic acid molecule or a vector as described herein, and a pharmaceutically acceptable excipient, is provided.

[0039] In certain aspects, a pharmaceutical composition comprising a host cell as described herein, and a pharmaceutically acceptable excipient, is provided.

[0040] In certain aspects, a kit, comprising a nucleic acid molecule as described herein, and instructions for administering the nucleic acid molecule to a subject in need thereof, is provided.

[0041] In certain aspects, a baculovirus system for production of a nucleic acid molecule as described herein, is provided.

[0042] In certain exemplary embodiments, a nucleic acid molecule as described herein, is produced in insect cells.

[0043] In certain aspects, a nanoparticle delivery system comprising a nucleic acid molecule as described herein, is provided.

[0044] In certain aspects, a method of producing a polypeptide, comprising culturing a host cell as described herein under suitable conditions and recovering the polypeptide, is provided.

[0045] In certain aspects, a method of producing a polypeptide with clotting activity, comprising: culturing a host cell as described herein under suitable conditions and recovering the polypeptide with clotting activity, is provided.

[0046] In certain aspects, a method of expressing a heterologous polynucleotide sequence in a subject in need thereof, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, or a pharmaceutical composition as described herein, is provided.

[0047] In certain aspects, a method of expressing a clotting factor in a subject in need thereof, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, a polypeptide as described herein, or a pharmaceutical composition as described herein, is provided.

[0048] In certain aspects, a method of treating a disease or disorder in a subject in need thereof, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, or a pharmaceutical composition as described herein, is provided.

[0049] In certain aspects, a method of treating a subject having a clotting factor deficiency, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, a polypeptide as described herein, or a pharmaceutical composition as described herein, is provided.

[0050] In certain aspects, a method of treating a clotting factor deficiency in a subject in need thereof, comprising administering to the subject a nucleic acid molecule as described herein, a vector as described herein, a polypeptide as described herein, or a pharmaceutical composition as described herein, is provided.

[0051] In certain exemplary embodiments, the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonarily, or any combination thereof. In certain exemplary embodiments, the nucleic acid molecule is administered intravenously.

[0052] In certain exemplary embodiments, the method further comprising administering to the subject a second agent.

[0053] In certain exemplary embodiments, the subject is a mammal. In certain exemplary embodiments, the subject is a human.

[0054] In certain exemplary embodiments, the administration of the nucleic acid molecule to the subject results in an increased FVIII activity, relative to a FVIII activity in the subject prior to the administration, wherein the FVIII activity is increased by at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 1 1-fold, at least about 12-fold, at least about 13-fold, at least about 14-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, or at least about 100-fold.

[0055] In certain exemplary embodiments, the subject has a bleeding disorder. In certain exemplary embodiments, the bleeding disorder is a hemophilia. In certain exemplary embodiments, the bleeding disorder is hemophilia A.

[0056] In certain aspects, a method of treating a bleeding disorder in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding a clotting factor, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof, is provided.

[0057] In certain aspects, a method of treating hemophilia A in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding factor VIII (FVIII), wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof, is provided

[0058] In certain aspects, a method of treating a metabolic disorder of the liver in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding a liver-associated metabolic enzyme that is deficient in the subject, wherein the first ITR and / or second ITR are an ITR of a non-adeno- associated virus (non-AAV), is provided.

[0059] In certain exemplary embodiments, the the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0060] In certain aspects, a method of treating a metabolic disorder of the liver in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding a liver-associated metabolic enzyme that is deficient in the subject, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof, is provided.

[0061] In certain exemplary embodiments, the genetic cassette comprises a single stranded nucleic acid. In certain exemplary embodiments, the genetic cassette comprises a double stranded nucleic acid.

[0062] In certain exemplary embodiments, the metabolic disorder of the liver is selected from the group consisting of phenylketonuria (PKU), a urea cycle disease, a lysosomal storage disorder, and a glycogen storage disease. In certain exemplary embodiments, the metabolic disorder of the liver is phenylketonuria (PKU).

[0063] In certain exemplary embodiments, the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonarily, or any combination thereof. In certain exemplary embodiments, the nucleic acid molecule is administered intravenously.

[0064] In certain exemplary embodiments, the method further comprising administering to the subject a second agent.

[0065] In certain exemplary embodiments, the subject is a mammal. In certain exemplary embodiments, the subject is a human.

[0066] In certain aspects, a method of treating phenylketonuria (PKU) in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding phenylalanine hydroxylase, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, atleast about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof, is provided.

[0067] In certain exemplary embodiments, the genetic cassette comprises a single stranded nucleic acid. In certain exemplary embodiments, the genetic cassette comprises a double stranded nucleic acid.

[0068] In certain exemplary embodiments, the nucleic acid molecule is formulated with a delivery agent. In certain exemplary embodiments, the delivery agent comprises a lipid nanoparticle.

[0069] In certain aspects, a method of cloning a nucleic acid molecule, comprising inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex, is provided

[0070] In certain exemplary embodiments, the the disruption in the SbcCD complex comprises a genetic disruption in the SbcC gene and / or SbcD gene. In certain exemplary embodiments, the disruption in the SbcCD complex comprises a genetic disruption in the SbcC gene. In certain exemplary embodiments, the disruption in the SbcCD complex comprises a genetic disruption in the SbcD gene.

[0071] In certain exemplary embodiments, the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first and / or second ITR is a non- adeno-associated virus (non-AAV) ITR.

[0072] In certain exemplary embodiments, the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

[0073] In certain exemplary embodiments, the nucleic acid molecule further comprises a genetic cassette, wherein the genetic cassette is flanked by the first ITR and second ITR.

[0074] In certain exemplary embodiments, the genetic cassette comprises a heterologous polynucleotide sequence.

[0075] In certain exemplary embodiments, the uitable vector is a low copy vector. In certain exemplary embodiments, the suitable vector is pBR322.

[0076] In certain exemplary embodiments, the bacterial host strain is incapable of resolving cruciform DNA structures.

[0077] In certain exemplary embodiments, the bacterial host strain is PMC103, comprising the genotype sbcC, recD, mcrA, AmcrBCF. In certain exemplary embodiments, thebacterial host strain is PMC107, comprising the genotype recBC, recJ, sbcBC, mcrA, AmcrBCF. In certain exemplary embodiments, the bacterial host strain is SURE, comprising the genotype recB, recJ, sbcC, mcrA, AmcrBCF, umuC, uvrC.

[0078] In certain aspects, a method of cloning a nucleic acid molecule, comprising inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex, wherein the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof, is providedBRIEF DESCRIPTION OF THE DRAWINGS

[0079] FIG. 1A-1 B are schematic representations of a single strand clotting factor (e.g.,FVIII) expression cassette. The locations of 5' ITR from a non-AAV (with hairpin loop at the end of the ssDNA structure), 3' ITR from a non-AAV (with hairpin loop), a promotor sequence (e.g., TTPp or CAGp), and a transgene sequence, e.g., FVIIIco6XTEN sequence with an XTEN144 inserted within the B domain are shown. The exemplary expression cassettes also show additional possible elements, e.g., an intron sequence, WPREmut sequence, and bGHpA sequence.

[0080] FIGs. 1C-1 F are schematic representations of plasmids used to prepare single strand clotting factor expression cassettes, such as the cassette shown in FIG. 1A-1 B, wherein the ITRs of the cassette are derived from AAV2 (FIG. 1 C), B19 (FIG. 1 D), GPV (FIG. 1 E), or are the wildtype B19 ITR sequence (FIG. 1 F). A plasmid construct comprising an ssFVIII expression cassette as shown here was digested with Pvull (at Pvull sites) (FIG. 1 C) or Lgul (at Lgul sites) (FIGs. 1 D-1 F) to precisely release the sequence comprising the ITRs and expression cassette. The double stranded DNA was heat denatured at 95°C to produce ssDNA and then incubated at 4°C to allow for ITR structure formation.

[0081] FIG. 2A is a phylogenetic tree illustrating that relationships between various parvovirus family members. B19, AAV-2, and GPV are marked by outlined boxes.

[0082] FIG. 2B is a schematic drawing of the various cassettes, including the hairpin structures.

[0083] FIGs. 3A and 3B are alignments of the ITRs of B19, GPV, and AAV2 (FIG. 3A) and B19 and GPV (FIG. 3B). Gray shading shows homology.

[0084] FIGs. 4A-4C show FVIII plasma activity following single-stranded FVIII-AAV naked DNA (ssAAV-FVIII; FIG. 1 C), SSDNA-B19 FVIII (FIG. 1 D), or ssDNA-GPV FVIII (FIG. 1 E) administration via hydrodynamic injection (HDI) in Hem A mice. FVIII Activity was measured (as a percentage of normal physiological levels in humans) in plasma samples at 24 hours, 3 days, 2 weeks, 3 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, and 6 months in mice treated with a single HDI of ssDNA at 50 pg / mouse (FIG. 4C), 20 pg / mouse (FIGs. 4A and 4B), 10 pg / mouse (FIGs. 4A, 4B, and 4C), or 5 pg / mouse (FIG. 4A). An HDI of 5 pg / mouse of plasmid DNA was given as a control (FIGs. 4A, 4B, and 4C).

[0085] FIG. 5 shows FVIII activity in hemophilia A mouse plasma following a single hydrodynamic injection of equal molar amounts of single- stranded naked DNA (ssAAV-FVIII, FIG. 1A), double-stranded AAV-FVIII DNA containing the ITR sequence (dsDNA), double-stranded FVIII DNA without the ITR sequence (dsDNA No ITR), or circularized double-stranded FVIII DNA without ITR or bacterial sequences (minicircle). dsDNA was generated by enzyme cleavage of the AAV-FVIII plasmid (FIG. 2C) with Pvull but not heat denatured. dsDNA No ITR was generated by enzyme cleavage of the AAV-FVIII plasmid (FIG. 2C) with Aflll and subsequently purified. Minicircle DNA was generated by ligation of the dsDNA No ITR DNA at Aflll sites. Mouse plasma was collected over 3 months or 4 months and FVIII was determined by chromogenic activity assay.

[0086] FIG. 6 shows FVIII activity in hemophilia A mouse plasma following a hydrodynamic injection of 30 pg of single-stranded naked FVIII-DNA (FIG. 1A, FIGs. 1 D-1 F). Plasma was collected weekly for 7 weeks and FVIII activity was determined by chromogenic assay. After 35 days (depicted as black arrow), mice receiving FVIII-B19d135 and FVIII- GPVd162 ssDNA were re-administered 30 pg via hydrodynamic injection.

[0087] FIG. 7A is a schematic representations of a single strand murine phenylalanine hydroxylase (e.g., PAH) expression cassette. The locations of 5' ITR from a non-AAV (with hairpin loop at the end of the ssDNA structure), 3' ITR from a non-AAV (with hairpin loop), a promotor sequence (e.g., CAGp), and a transgene sequence, e.g., 3xFLAG_mPAH sequence are shown. The exemplary expression cassettes also show additional possible elements, e.g., WPREmut sequence, and bGHpA sequence.

[0088] FIGs. 7B-7D show plasma concentrations of phenylalanine (Phe) in phenylketonuria (PKU) mice before (day 0) and after single administration of single-stranded DNA containing the murine PAH cDNA and non-AAV ITRs B19d135 or GPVd162 via hydrodynamic injection. Plasma was collected at days 3, 7, 14, 28, 42, and 56 following ssDNA administration. Residual phenylalanine levels are shown as concentration in pg / ml (FIGs. 7B-7C)or as percent prior to administration (FIG.7D). The horizontal line depicts baseline Phe levels prior to administration.

[0089] FIG. 7E shows a Western immunoblot of liver lysates from PKU mice treated with ssDNA containing the murine PAH transgene and either B19d135 or GPVd165 ITRs. Livers were collected at day 81 post treatment and protein lysates were extracted. Each well represents a single animal. The FLAG-tagged murine PAH protein was detected using the M2 anti-FLAG antibody and a GAPDH loading control was included for comparison.

[0090] FIGs. 8A-B show FVIII activity levels in Huh7 cell supernatant following transduction with FVIII-AAV DNA (FIGs. 1A-1 C) encapsulated lipid nanoparticles. Plasmid FVIII- AAV under the CAGp promoter (FIG. 1 B) was encapsulated at three amine-to-phosphate (NP) ratios and applied to Huh7 cells at various concentrations determined by picogreen assay (FIG 8A). Plasmid, double stranded linear (ds), and single-stranded (ss) AAV-FVIII under the TTPp promoter (FIG 1A) was also encapsulated in lipid nanoparticles at two NP ratios and used to transduce Huh7 cells at various DNA concentrations (FIG. 8B). FVIII was measured by chromogenic activity assay compared to a human FACT plasma standard.DETAILED DESCRIPTION OF THE DISCLOSURE

[0091] The present disclosure describes plasmid-like nucleic acid molecules comprising a first inverted terminal repeat (ITR), a second ITR, and a genetic cassette, e.g., encoding a target sequence (also referred to herein as a heterologous polynucleotide sequence), e.g., a therapeutic protein or a miRNA, wherein the first ITR and / or the second ITR are an ITR of a non- adeno-associated virus (e.g., the first ITR and / or the second ITR are from a non-AAV). In some embodiments, the genetic cassette encodes a therapeutic protein, e.g., the target sequence encodes a therapeutic protein. In some embodiments, the therapeutic protein comprises a protein selected from a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or a combination thereof. In some embodiments, the genetic cassette encodes dystrophin X-linked, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1 , FXN (frataxin), GUCY2D, RS1 , CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof.

[0092] In some embodiments, the therapeutic protein comprises a clotting factor. In one particular embodiment, the therapeutic protein comprises a FVIII or a FIX protein.

[0093] In some embodiments, the genetic cassette encodes a miRNA. In certain embodiments, the miRNA down regulates the expression of a target gene selected from SOD1 , HTT, RHO, or any combination thereof.

[0094] In certain embodiments, the non-AAV is selected from the group consisting of a member of the viral family Parvoviridae and any combination thereof. The present disclosure is further directed to methods of expressing a therapeutic protein, e.g., a clotting factor, e.g., a FVIII, in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a genetic cassette, e.g. , encoding a therapeutic protein or an miRNA, wherein the first ITR and / or the second ITR are an ITR of a non-adeno-associated virus (non-AAV). In certain embodiments, the disclosure describes an isolated nucleic acid molecule comprising a nucleotide sequence, which has sequence homology to a nucleotide sequence selected from SEQ ID NOs: 1 13 and 120.

[0095] In certain embodiments, the present disclosure provides nucleic acid molecules comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence, wherein the first and / or second ITR is derived from parvovirus B19 or goose parvovirus (GPV).

[0096] Exemplary constructs of the disclosure are illustrated in the accompanying figures and sequence listing. In order to provide a clear understanding of the specification and claims, the following definitions are provided below.I. Definitions

[0097] It is to be noted that the term "a" or "an" entity refers to one or more of that entity: for example, "a nucleotide sequence" is understood to represent one or more nucleotide sequences. Similarly, "a therapeutic protein" and "a miRNA" is understood to represent one or more therapeutic protein and one or more miRNA, respectively. As such, the terms "a" (or "an"), "one or more," and "at least one" can be used interchangeably herein.

[0098] The term "about" is used herein to mean approximately, roughly, around, or in the regions of. When the term "about" is used in conjunction with a numerical range, it modifies that range by extending the boundaries above and below the numerical values set forth. In general, the term "about" is used herein to modify a numerical value above and below the stated value by a variance of 10 percent, up or down (higher or lower).

[0099] Also as used herein, "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").

[0100] "Nucleic acids," "nucleic acid molecules," "nucleotides," "nucleotide(s) sequence," and "polynucleotide" are used interchangeably and refer to the phosphate ester polymeric form of ribonucleosides (adenosine, guanosine, uridine or cytidine; "RNA molecules") or deoxyribonucleosides (deoxyadenosine, deoxyguanosine, deoxythymidine, or deoxycytidine; "DNA molecules"), or any phosphoester analogs thereof, such as phosphorothioates andthioesters, in either single stranded form, or a double-stranded helix. Single stranded nucleic acid sequences refer to single-stranded DNA (ssDNA) or single-stranded RNA (ssRNA). Double stranded DNA-DNA, DNA-RNA and RNA-RNA helices are possible. The term nucleic acid molecule, and in particular DNA or RNA molecule, refers only to the primary and secondary structure of the molecule, and does not limit it to any particular tertiary forms. Thus, this term includes double-stranded DNA found, inter alia, in linear or circular DNA molecules (e.g., restriction fragments), plasmids, supercoiled DNA and chromosomes. In discussing the structure of particular double-stranded DNA molecules, sequences can be described herein according to the normal convention of giving only the sequence in the 5’ to 3’ direction along the non- transcribed strand of DNA (i.e., the strand having a sequence homologous to the mRNA). A "recombinant DNA molecule" is a DNA molecule that has undergone a molecular biological manipulation. DNA includes, but is not limited to, cDNA, genomic DNA, plasmid DNA, synthetic DNA, and semi-synthetic DNA. A "nucleic acid composition" of the disclosure comprises one or more nucleic acids as described herein.

[0101] As used herein, an "inverted terminal repeat" (or "ITR") refers to a nucleic acid subsequence located at either the 5' or 3' end of a single stranded nucleic acid sequence, which comprises a set of nucleotides (initial sequence) followed downstream by its reverse complement, i.e., palindromic sequence. The intervening sequence of nucleotides between the initial sequence and the reverse complement can be any length including zero. In one embodiment, the ITR useful for the present disclosure comprises one or more "palindromic sequences." An ITR can have any number of functions. In some embodiments, an ITR described herein forms a hairpin structure. In some embodiments, the ITR forms a T-shaped hairpin structure. In some embodiments, the ITR forms a non-T-shaped hairpin structure, e.g., a U- shaped hairpin structure. In some embodiments, the ITR promotes the long-term survival of the nucleic acid molecule in the nucleus of a cell. In some embodiments, the ITR promotes the permanent survival of the nucleic acid molecule in the nucleus of a cell (e.g., for the entire lifespan of the cell). In some embodiments, the ITR promotes the stability of the nucleic acid molecule in the nucleus of a cell. In some embodiments, the ITR promotes the retention of the nucleic acid molecule in the nucleus of a cell. In some embodiments, the ITR promotes the persistence of the nucleic acid molecule in the nucleus of a cell. In some embodiments, the ITR inhibits or prevents the degradation of the nucleic acid molecule in the nucleus of a cell.

[0102] In one embodiment, the initial sequence and / or the reverse complement comprise about 2-600 nucleotides, about 2-550 nucleotides, about 2-500 nucleotides, about 2-450 nucleotides, about 2-400 nucleotides, about 2-350 nucleotides, about 2-300 nucleotides, or about 2-250 nucleotides. In some embodiments, the initial sequence and / or the reverse complementcomprise about 5-600 nucleotides, about 10-600 nucleotides, about 15-600 nucleotides, about 20-600 nucleotides, about 25-600 nucleotides, about 30-600 nucleotides, about 35-600 nucleotides, about 40-600 nucleotides, about 45-600 nucleotides, about 50-600 nucleotides, about 60-600 nucleotides, about 70-600 nucleotides, about 80-600 nucleotides, about 90-600 nucleotides, about 100-600 nucleotides, about 150-600 nucleotides, about 200-600 nucleotides, about 300-600 nucleotides, about 350-600 nucleotides, about 400-600 nucleotides, about 450- 600 nucleotides, about 500-600 nucleotides, or about 550-600 nucleotides. In some embodiments, the initial sequence and / or the reverse complement comprise about 5-550 nucleotides, about 5 to 500 nucleotides, about 5-450 nucleotides, about 5 to 400 nucleotides, about 5-350 nucleotides, about 5 to 300 nucleotides, or about 5-250 nucleotides. In some embodiments, the initial sequence and / or the reverse complement comprise about 10-550 nucleotides, about 15-500 nucleotides, about 20-450 nucleotides, about 25-400 nucleotides, about 30-350 nucleotides, about 35-300 nucleotides, or about 40-250 nucleotides. In certain embodiments, the initial sequence and / or the reverse complement comprise about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, or about 600 nucleotides. In particular embodiments, the initial sequence and / or the reverse complement comprise about 400 nucleotides.

[0103] In other embodiments, the initial sequence and / or the reverse complement comprise about 2-200 nucleotides, about 5-200 nucleotides, about 10-200 nucleotides, about 20- 200 nucleotides, about 30-200 nucleotides, about 40-200 nucleotides, about 50-200 nucleotides, about 60-200 nucleotides, about 70-200 nucleotides, about 80-200 nucleotides, about 90-200 nucleotides, about 100-200 nucleotides, about 125-200 nucleotides, about 150-200 nucleotides, or about 175-200 nucleotides. In other embodiments, the initial sequence and / or the reverse complement comprise about 2-150 nucleotides, about 5-150 nucleotides, about 10-150 nucleotides, about 20-150 nucleotides, about 30-150 nucleotides, about 40-150 nucleotides, about 50-150 nucleotides, about 75-150 nucleotides, about 100-150 nucleotides, or about 125- I SO nucleotides. In other embodiments, the initial sequence and / or the reverse complement comprise about 2-100 nucleotides, about 5-100 nucleotides, about 10-100 nucleotides, about 20- 100 nucleotides, about 30-100 nucleotides, about 40-100 nucleotides, about 50-100 nucleotides, or about 75-100 nucleotides. In other embodiments, the initial sequence and / or the reverse complement comprise about 2-50 nucleotides, about 10-50 nucleotides, about 20-50 nucleotides, about 30-50 nucleotides, about 40-50 nucleotides, about 3-30 nucleotides, about 4-20nucleotides, or about 5-10 nucleotides. In another embodiment, the initial sequence and / or the reverse complement consist of two nucleotides, three nucleotides, four nucleotides, five nucleotides, six nucleotides, seven nucleotides, eight nucleotides, nine nucleotides, ten nucleotides, 1 1 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides. In other embodiments, an intervening nucleotide between the initial sequence and the reverse complement is (e.g., consists of) 0 nucleotide, 1 nucleotide, two nucleotides, three nucleotides, four nucleotides, five nucleotides, six nucleotides, seven nucleotides, eight nucleotides, nine nucleotides, 10 nucleotides, 1 1 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides.

[0104] Therefore, an "ITR" as used herein can fold back on itself and form a double stranded segment. For example, the sequence GATCXXXXGATC comprises an initial sequence of GATC and its complement (3'CTAG5') when folded to form a double helix. In some embodiments, the ITR comprises a continuous palindromic sequence (e.g., GATCGATC) between the initial sequence and the reverse complement. In some embodiments, the ITR comprises an interrupted palindromic sequence (e.g., GATCXXXXGATC) between the initial sequence and the reverse complement. In some embodiments, the complementary sections of the continuous or interrupted palindromic sequence interact with each other to form a "hairpin loop" structure. As used herein, a "hairpin loop" structure results when at least two complimentary sequences on a single-stranded nucleotide molecule base-pair to form a double stranded section. In some embodiments, only a portion of the ITR forms a hairpin loop. In other embodiments, the entire ITR forms a hairpin loop.

[0105] In the present disclosure, at least one ITR is an ITR of a non-adenovirus associated virus (non-AAV). In certain embodiments, the ITR is an ITR of a non-AAV member of the viral family Parvoviridae. In some embodiments, the ITR is an ITR of a non-AAV member of the genus Dependovirus or the genus Erythrovirus. In particular embodiments, the ITR is an ITR of a goose parvovirus (GPV), a Muscovy duck parvovirus (MDPV), or an erythrovirus parvovirus B19 (also known as parvovirus B19, primate erythroparvovirus 1 , B19 virus, and erythrovirus). In certain embodiments, one ITR of two ITRs is an ITR of an AAV. In other embodiments, one ITR of two ITRs in the construct is an ITR of an AAV serotype selected from serotype 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 1 1 and any combination thereof. In one particular embodiment, the ITR is derived from AAV serotype 2, e.g., an ITR of AAV serotype 2.

[0106] In certain aspects of the present disclosure, the nucleic acid molecule comprises two ITRs, a 5' ITR and a 3' ITR, wherein the 5' ITR is located at the 5' terminus of the nucleic acid molecule, and the 3' ITR is located at the 3' terminus of the nucleic acid molecule. The 5' ITR andthe 3' ITR can be derived from the same virus or different viruses. In certain embodiments, the 5' ITR is derived from an AAV and the 3' ITR is not derived from an AAV virus ( e.g ., a non-AAV). In some embodiments, the 3' ITR is derived from an AAV and the 5' ITR is not derived from an AAV virus (e.g., a non-AAV). In other embodiments, the 5' ITR is not derived from an AAV virus (e.g. , a non-AAV), and the 3' ITR is derived from the same or a different non-AAV virus.

[0107] The term "parvovirus" as used herein encompasses the family Parvoviridae, including but not limited to autonomously-replicating parvoviruses and Dependoviruses. The autonomous parvoviruses include, for example, members of the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstyldensovirus.

[0108] Exemplary autonomous parvoviruses include, but are not limited to, porcine parvovirus, mice minute virus, canine parvovirus, mink entertitus virus, bovine parvovirus, chicken parvovirus, feline panleukopenia virus, feline parvovirus, goose parvovirus, H1 parvovirus, muscovy duck parvovirus, snake parvovirus, and B19 virus. Other autonomous parvoviruses are known to those skilled in the art. See, e.g., FIELDS et al. VIROLOGY, volume 2, chapter 69 (4th ed., Lippincott-Raven Publishers).

[0109] The term "non-AAV" as used herein encompasses nucleic acids, proteins, and viruses from the family Parvoviridae excluding any adeno-associated viruses (AAV) of the Parvoviridae family. "Non-AAV" includes but is not limited to autonomously-replicating members of the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstyldensovirus.

[0110] As used herein, the term "adeno-associated virus" (AAV), includes but is not limited to, AAV type 1 , AAV type 2, AAV type 3 (including types 3A and 3B), AAV type 4, AAV type 5, AAV type 6, AAV type 7, AAV type 8, AAV type 9, AAV type 10, AAV type 11 , AAV type 12, AAV type 13, snake AAV, avian AAV, bovine AAV, canine AAV, equine AAV, ovine AAV, goat AAV, shrimp AAV, those AAV serotypes and clades disclosed by Gao et al. (J. Virol. 78:6381 (2004)) and Moris et al. (Virol. 33:375 (2004)), and any other AAV now known or later discovered. See, e.g., FIELDS et al. VIROLOGY, volume 2, chapter69 (4th ed., Lippincott-Raven Publishers).

[0111] The term“derived from,” as used herein, refers to a component that is isolated from or made using a specified molecule or organism, or information (e.g., amino acid or nucleic acid sequence) from the specified molecule or organism. For example, a nucleic acid sequence (e.g., ITR) that is derived from a second nucleic acid sequence (e.g., ITR) can include a nucleotide sequence that is identical or substantially similar to the nucleotide sequence of thesecond nucleic acid sequence. In the case of nucleotides or polypeptides, the derived species can be obtained by, for example, naturally occurring mutagenesis, artificial directed mutagenesis or artificial random mutagenesis. The mutagenesis used to derive nucleotides or polypeptides can be intentionally directed or intentionally random, or a mixture of each. The mutagenesis of a nucleotide or polypeptide to create a different nucleotide or polypeptide derived from the first can be a random event (e.g., caused by polymerase infidelity) and the identification of the derived nucleotide or polypeptide can be made by appropriate screening methods, e.g., as discussed herein. Mutagenesis of a polypeptide typically entails manipulation of the polynucleotide that encodes the polypeptide. In some embodiments, a nucleotide or amino acid sequence that is derived from a second nucleotide or amino acid sequence has a sequence identity of at least 50%, at least 51 %, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least57%, at least 58%, at least 59%, at least 60%, at least 61 %, at least 62%, at least 63%, at least64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least71 %, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least78%, at least 79%, at least 80%, at least 81 %, at least 82%, at least 83%, at least 84%, at least85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least99%, or 100% to the second nucleotide or amino acid sequence, respectively, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In other embodiments, an ITR derived from an ITR of a non-AAV (or AAV) is at least 90% identical to the non-AAV ITR (or AAV ITR, respectively), wherein the non-AAV (or AAV) ITR retains a functional property of the non-AAV ITR (or AAV ITR, respectively). In some embodiments, an ITR derived from an ITR of a non-AAV (or AAV) is at least 80% identical to the non-AAV ITR (or AAV ITR, respectively), wherein the non-AAV (or AAV) ITR retains a functional property of the non-AAV ITR (or AAV ITR, respectively). In some embodiments, an ITR derived from an ITR of a non-AAV (or AAV) is at least 70% identical to the non-AAV ITR (or AAV ITR, respectively), wherein the non-AAV (or AAV) ITR retains a functional property of the non-AAV ITR (or AAV ITR, respectively). In some embodiments, an ITR derived from an ITR of a non-AAV (or AAV) is at least 60% identical to the non-AAV ITR (or AAV ITR, respectively), wherein the non-AAV (or AAV) ITR retains a functional property of the non-AAV ITR (or AAV ITR, respectively). In some embodiments, an ITR derived from an ITR of a non-AAV (or AAV) is at least 50% identical to the non-AAV ITR (or AAV ITR, respectively), wherein the non-AAV (or AAV) ITR retains a functional property of the non-AAV ITR (or AAV ITR, respectively).

[0112] In certain embodiments, an ITR derived from an ITR of a non-AAV (or AAV) comprises or consists of a fragment of the ITR of the non-AAV (or AAV). In some embodiments,the ITR derived from an ITR of a non-AAV (or AAV) comprises or consists of a fragment of the ITR of the non-AAV (or AAV), wherein the fragment comprises at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides, or at least about 600 nucleotides; wherein the ITR derived from an ITR of a non-AAV (or AAV) retains a functional property of the non-AAV ITR (or AAV ITR, respectively). In certain embodiments, the ITR derived from an ITR of a non-AAV (or AAV) comprises or consists of a fragment of the ITR of the non-AAV (or AAV), wherein the fragment comprises at least about 129 nucleotides, and wherein the ITR derived from an ITR of a non-AAV (or AAV) retains a functional property of the non-AAV ITR (or AAV ITR, respectively). In certain embodiments, the ITR derived from an ITR of a non-AAV (or AAV) comprises or consists of a fragment of the ITR of the non-AAV (or AAV), wherein the fragment comprises at least about 102 nucleotides, and wherein the ITR derived from an ITR of a non-AAV (or AAV) retains a functional property of the non-AAV ITR (or AAV ITR, respectively).

[0113] In some embodiments, the ITR derived from an ITR of a non-AAV (or AAV) comprises or consists of a fragment of the ITR of the non-AAV (or AAV), wherein the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the length of the ITR of the non-AAV (or AAV).

[0114] In certain embodiments, a nucleotide or amino acid sequence that is derived from a second nucleotide or amino acid sequence has a sequence identity of at least 50%, at least 51 %, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least58%, at least 59%, at least 60%, at least 61 %, at least 62%, at least 63%, at least 64%, at least65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71 %, at least72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least79%, at least 80%, at least 81 %, at least 82%, at least 83%, at least 84%, at least 85%, at least86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 92%, at least93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% to a homologous portion of the second nucleotide or amino acid sequence, respectively, when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In other embodiments, an ITR derived from an ITR of a non-AAV (or AAV) is at least 90% identical to a homologous portion of the non-AAV ITR (or AAV ITR, respectively), when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, an ITR derived from an ITR of a non-AAV (or AAV) is at least 80% identical to a homologous portion of the non-AAV ITR (or AAV ITR, respectively), when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, an ITR derived from an ITR of a non- AAV (or AAV) is at least 70% identical to a homologous portion of the non-AAV ITR (or AAV ITR, respectively), when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, an ITR derived from an ITR of a non-AAV (or AAV) is at least 60% identical to a homologous portion of the non-AAV ITR (or AAV ITR, respectively), when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, an ITR derived from an ITR of a non-AAV (or AAV) is at least 50% identical to a homologous portion of the non-AAV ITR (or AAV ITR, respectively), when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence.

[0115] A "capsid-free" or "capsid-less" vector or nucleic acid molecule refers to a vector construct free from a capsid. In some embodiments, the capsid-less vector or nucleic acid molecule does not contain sequences encoding, e.g. , an AAV Rep protein.

[0116] As used herein, a "coding region" or "coding sequence" is a portion of polynucleotide which consists of codons translatable into amino acids. Although a "stop codon" (TAG, TGA, or TAA) is typically not translated into an amino acid, it can be considered to be part of a coding region, but any flanking sequences, for example promoters, ribosome binding sites, transcriptional terminators, introns, and the like, are not part of a coding region. The boundaries of a coding region are typically determined by a start codon at the 5’ terminus, encoding theamino terminus of the resultant polypeptide, and a translation stop codon at the 3’ terminus, encoding the carboxyl terminus of the resulting polypeptide. Two or more coding regions can be present in a single polynucleotide construct, e.g., on a single vector, or in separate polynucleotide constructs, e.g., on separate (different) vectors. It follows, then, that a single vector can contain just a single coding region, or comprise two or more coding regions.

[0117] Certain proteins secreted by mammalian cells are associated with a secretory signal peptide which is cleaved from the mature protein once export of the growing protein chain across the rough endoplasmic reticulum has been initiated. Those of ordinary skill in the art are aware that signal peptides are generally fused to the N-terminus of the polypeptide, and are cleaved from the complete or "full-length" polypeptide to produce a secreted or "mature" form of the polypeptide. In certain embodiments, a native signal peptide or a functional derivative of that sequence that retains the ability to direct the secretion of the polypeptide that is operably associated with it. Alternatively, a heterologous mammalian signal peptide, e.g. , a human tissue plasminogen activator (TPA) or mouse b-glucuronidase signal peptide, or a functional derivative thereof, can be used.

[0118] The term "downstream" refers to a nucleotide sequence that is located 3’ to a reference nucleotide sequence. In certain embodiments, downstream nucleotide sequences relate to sequences that follow the starting point of transcription. For example, the translation initiation codon of a gene is located downstream of the start site of transcription.

[0119] The term "upstream" refers to a nucleotide sequence that is located 5’ to a reference nucleotide sequence. In certain embodiments, upstream nucleotide sequences relate to sequences that are located on the 5’ side of a coding region or starting point of transcription. For example, most promoters are located upstream of the start site of transcription.

[0120] As used herein, the term "gene regulatory region" or "regulatory region" refers to nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding region, and which influence the transcription, RNA processing, stability, or translation of the associated coding region. Regulatory regions can include promoters, translation leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites, or stem-loop structures. If a coding region is intended for expression in a eukaryotic cell, a polyadenylation signal and transcription termination sequence will usually be located 3’ to the coding sequence.

[0121] A polynucleotide which encodes a product, e.g. , a miRNA or a gene product (e.g. , a polypeptide such as a therapeutic protein), can include a promoter and / or other expression (e.g., transcription or translation) control elements operably associated with one or more coding regions. In an operable association a coding region for a gene product, e.g., a polypeptide, isassociated with one or more regulatory regions in such a way as to place expression of the gene product under the influence or control of the regulatory region(s). For example, a coding region and a promoter are "operably associated" if induction of promoter function results in the transcription of mRNA encoding the gene product encoded by the coding region, and if the nature of the linkage between the promoter and the coding region does not interfere with the ability of the promoter to direct the expression of the gene product or interfere with the ability of the DNA template to be transcribed. Other expression control elements, besides a promoter, for example enhancers, operators, repressors, and transcription termination signals, can also be operably associated with a coding region to direct gene product expression.

[0122] "Transcriptional control sequences" refer to DNA regulatory sequences, such as promoters, enhancers, terminators, and the like, that provide for the expression of a coding sequence in a host cell. A variety of transcription control regions are known to those skilled in the art. These include, without limitation, transcription control regions which function in vertebrate cells, such as, but not limited to, promoter and enhancer segments from cytomegaloviruses (the immediate early promoter, in conjunction with intron-A), simian virus 40 (the early promoter), and retroviruses (such as Rous sarcoma virus). Other transcription control regions include those derived from vertebrate genes such as actin, heat shock protein, bovine growth hormone and rabbit b-globin, as well as other sequences capable of controlling gene expression in eukaryotic cells. Additional suitable transcription control regions include tissue-specific promoters and enhancers as well as lymphokine-inducible promoters (e.g., promoters inducible by interferons or interleukins).

[0123] Similarly, a variety of translation control elements are known to those of ordinary skill in the art. These include, but are not limited to ribosome binding sites, translation initiation and termination codons, and elements derived from picornaviruses (particularly an internal ribosome entry site, or IRES, also referred to as a CITE sequence).

[0124] The term "expression" as used herein refers to a process by which a polynucleotide produces a gene product, for example, an RNA or a polypeptide. It includes without limitation transcription of the polynucleotide into messenger RNA (mRNA), transfer RNA (tRNA), small hairpin RNA (shRNA), small interfering RNA (siRNA) or any other RNA product, and the translation of an mRNA into a polypeptide. Expression produces a "gene product." As used herein, a gene product can be either a nucleic acid, e.g., a messenger RNA produced by transcription of a gene, or a polypeptide which is translated from a transcript. Gene products described herein further include nucleic acids with post transcriptional modifications, e.g. , polyadenylation or splicing, or polypeptides with post translational modifications, e.g. , methylation, glycosylation, the addition of lipids, association with other protein subunits, orproteolytic cleavage. The term "yield," as used herein, refers to the amount of a polypeptide produced by the expression of a gene.

[0125] A "vector" refers to any vehicle for the cloning of and / or transfer of a nucleic acid into a host cell. A vector can be a replicon to which another nucleic acid segment can be attached so as to bring about the replication of the attached segment. A "replicon" refers to any genetic element (e.g., plasmid, phage, cosmid, chromosome, virus) that functions as an autonomous unit of replication in vivo, i.e., capable of replication under its own control. The term "vector" includes vehicles for introducing the nucleic acid into a cell in vitro, ex vivo or in vivo. A large number of vectors are known and used in the art including, for example, plasmids, modified eukaryotic viruses, or modified bacterial viruses. Insertion of a polynucleotide into a suitable vector can be accomplished by ligating the appropriate polynucleotide fragments into a chosen vector that has complementary cohesive termini.

[0126] Vectors can be engineered to encode selectable markers or reporters that provide for the selection or identification of cells that have incorporated the vector. Expression of selectable markers or reporters allows identification and / or selection of host cells that incorporate and express other coding regions contained on the vector. Examples of selectable marker genes known and used in the art include: genes providing resistance to ampicillin, streptomycin, gentamycin, kanamycin, hygromycin, bialaphos herbicide, sulfonamide, and the like; and genes that are used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentanyl transferase gene, and the like. Examples of reporters known and used in the art include: luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), b-galactosidase (LacZ), b-glucuronidase (Gus), and the like. Selectable markers can also be considered to be reporters.

[0127] The term "host cell" as used herein refers to, for example microorganisms, yeast cells, insect cells, and mammalian cells, that can be, or have been, used as recipients of ssDNA or vectors. The term includes the progeny of the original cell which has been transduced. Thus, a "host cell" as used herein generally refers to a cell which has been transduced with an exogenous DNA sequence. It is understood that the progeny of a single parental cell may not necessarily be completely identical in morphology or in genomic or total DNA complement to the original parent, due to natural, accidental, or deliberate mutation. In some embodiments, the host cell can be an in vitro host cell.

[0128] The term "selectable marker" refers to an identifying factor, usually an antibiotic or chemical resistance gene, that is able to be selected for based upon the marker gene’s effect, i.e., resistance to an antibiotic, resistance to a herbicide, colorimetric markers, enzymes, fluorescent markers, and the like, wherein the effect is used to track the inheritance of a nucleicacid of interest and / or to identify a cell or organism that has inherited the nucleic acid of interest. Examples of selectable marker genes known and used in the art include: genes providing resistance to ampicillin, streptomycin, gentamycin, kanamycin, hygromycin, bialaphos herbicide, sulfonamide, and the like; and genes that are used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentanyl transferase gene, and the like.

[0129] The term "reporter gene" refers to a nucleic acid encoding an identifying factor that is able to be identified based upon the reporter gene’s effect, wherein the effect is used to track the inheritance of a nucleic acid of interest, to identify a cell or organism that has inherited the nucleic acid of interest, and / or to measure gene expression induction or transcription. Examples of reporter genes known and used in the art include: luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), b-galactosidase (LacZ), b- glucuronidase (Gus), and the like. Selectable marker genes can also be considered reporter genes.

[0130] "Promoter" and "promoter sequence" are used interchangeably and refer to a DNA sequence capable of controlling the expression of a coding sequence or functional RNA. In general, a coding sequence is located 3' to a promoter sequence. Promoters can be derived in their entirety from a native gene, or be composed of different elements derived from different promoters found in nature, or even comprise synthetic DNA segments. It is understood by those skilled in the art that different promoters can direct the expression of a gene in different tissues or cell types, or at different stages of development, or in response to different environmental or physiological conditions. Promoters that cause a gene to be expressed in most cell types at most times are commonly referred to as "constitutive promoters." Promoters that cause a gene to be expressed in a specific cell type are commonly referred to as "cell-specific promoters" or "tissue- specific promoters." Promoters that cause a gene to be expressed at a specific stage of development or cell differentiation are commonly referred to as "developmentally-specific promoters" or "cell differentiation-specific promoters." Promoters that are induced and cause a gene to be expressed following exposure or treatment of the cell with an agent, biological molecule, chemical, ligand, light, or the like that induces the promoter are commonly referred to as "inducible promoters" or "regulatable promoters." It is further recognized that since in most cases the exact boundaries of regulatory sequences have not been completely defined, DNA fragments of different lengths can have identical promoter activity.

[0131] The promoter sequence is typically bounded at its 3’ terminus by the transcription initiation site and extends upstream (5’ direction) to include the minimum number of bases or elements necessary to initiate transcription at levels detectable above background. Within the promoter sequence will be found a transcription initiation site (conveniently defined for example,by mapping with nuclease S1), as well as protein binding domains (consensus sequences) responsible for the binding of RNA polymerase.

[0132] In some embodiments, the nucleic acid molecule comprises a tissue specific promoter. In certain embodiments, the tissue specific promoter drives expression of the therapeutic protein, e.g., the clotting factor, in the liver, e.g. , in hepatocytes and / or endothelial cells. In particular, embodiments, the promoter is selected from the group consisting of a mouse thyretin promoter (mTTR), an endogenous human factor VIII promoter (F8), a human alpha-1- antitrypsin promoter (hAAT), a human albumin minimal promoter, a mouse albumin promoter, a tristetraprolin (TTP) promoter, a CASI promoter, a CAG promoter, a cytomegalovirus (CMV) promoter, a phosphoglycerate kinase (PGK) promoter and any combination thereof. In some embodiments, the promoter is selected from a liver specific promoter (e.g., a1 -antitrypsin (AAT)), a muscle specific promoter (e.g., muscle creatine kinase (MCK), myosin heavy chain alpha (aMHC), myoglobin (MB), and desmin (DES)), a synthetic promoter (e.g., SPc5-12, 2R5Sc5-12, dMCK, and tMCK) and any combination thereof. In one particular embodiment, the promoter comprises a TTP promoter.

[0133] The terms "restriction endonuclease" and "restriction enzyme" are used interchangeably and refer to an enzyme that binds and cuts within a specific nucleotide sequence within double stranded DNA.

[0134] The term "plasmid" refers to an extra-chromosomal element often carrying a gene that is not part of the central metabolism of the cell, and usually in the form of circular double- stranded DNA molecules. Such elements can be autonomously replicating sequences, genome integrating sequences, phage or nucleotide sequences, linear, circular, or supercoiled, of a single- or double-stranded DNA or RNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construct, which is capable of introducing a promoter fragment and DNA sequence for a selected gene product along with appropriate 3' untranslated sequence into a cell.

[0135] Eukaryotic viral vectors that can be used include, but are not limited to, adenovirus vectors, retrovirus vectors, adeno-associated virus vectors, poxvirus, e.g. , vaccinia virus vectors, baculovirus vectors, or herpesvirus vectors. Non-viral vectors include plasmids, liposomes, electrically charged lipids (cytofectins), DNA-protein complexes, and biopolymers.

[0136] A "cloning vector" refers to a "replicon," which is a unit length of a nucleic acid that replicates sequentially and which comprises an origin of replication, such as a plasmid, phage or cosmid, to which another nucleic acid segment can be attached so as to bring about the replication of the attached segment. Certain cloning vectors are capable of replication in one cell type, e.g. , bacteria and expression in another, e.g., eukaryotic cells. Cloning vectors typicallycomprise one or more sequences that can be used for selection of cells comprising the vector and / or one or more multiple cloning sites for insertion of nucleic acid sequences of interest.

[0137] The term "expression vector" refers to a vehicle designed to enable the expression of an inserted nucleic acid sequence following insertion into a host cell. The inserted nucleic acid sequence is placed in operable association with regulatory regions as described above.

[0138] Vectors are introduced into host cells by methods well known in the art, e.g. , transfection, electroporation, microinjection, transduction, cell fusion, DEAE dextran, calcium phosphate precipitation, lipofection (lysosome fusion), use of a gene gun, or a DNA vector transporter. "Culture," "to culture" and "culturing," as used herein, means to incubate cells under in vitro conditions that allow for cell growth or division or to maintain cells in a living state. "Cultured cells," as used herein, means cells that are propagated in vitro.

[0139] As used herein, the term "polypeptide" is intended to encompass a singular"polypeptide" as well as plural "polypeptides," and refers to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also known as peptide bonds). The term "polypeptide" refers to any chain or chains of two or more amino acids, and does not refer to a specific length of the product. Thus, peptides, dipeptides, tripeptides, oligopeptides, "protein," "amino acid chain," or any other term used to refer to a chain or chains of two or more amino acids, are included within the definition of "polypeptide," and the term "polypeptide" can be used instead of, or interchangeably with any of these terms. The term "polypeptide" is also intended to refer to the products of post-expression modifications of the polypeptide, including without limitation glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting / blocking groups, proteolytic cleavage, or modification by non-naturally occurring amino acids. A polypeptide can be derived from a natural biological source or produced recombinant technology, but is not necessarily translated from a designated nucleic acid sequence. It can be generated in any manner, including by chemical synthesis.

[0140] The term "amino acid" includes alanine (Ala or A); arginine (Arg or R); asparagine(Asn or N); aspartic acid (Asp or D); cysteine (Cys or C); glutamine (Gin or Q); glutamic acid (Glu or E); glycine (Gly or G); histidine (His or H); isoleucine (lie or I): leucine (Leu or L); lysine (Lys or K); methionine (Met or M); phenylalanine (Phe or F); proline (Pro or P); serine (Ser or S); threonine (Thr or T); tryptophan (Trp or W); tyrosine (Tyr or Y); and valine (Val or V). Non- traditional amino acids are also within the scope of the disclosure and include norleucine, ornithine, norvaline, homoserine, and other amino acid residue analogues such as those described in Ellman et al. Meth. Enzym. 202:301-336 (1991). To generate such non-naturally occurring amino acid residues, the procedures of Noren etal. Science 244: 182 (1989) and Ellman et al., supra, can be used. Briefly, these procedures involve chemically activating a suppressortRNA with a non-naturally occurring amino acid residue followed by in vitro transcription and translation of the RNA. Introduction of the non-traditional amino acid can also be achieved using peptide chemistries known in the art. As used herein, the term "polar amino acid" includes amino acids that have net zero charge, but have non-zero partial charges in different portions of their side chains (e.g., M, F, W, S, Y, N, Q, C). These amino acids can participate in hydrophobic interactions and electrostatic interactions. As used herein, the term "charged amino acid" includes amino acids that can have non-zero net charge on their side chains (e.g., R, K, H, E, D). These amino acids can participate in hydrophobic interactions and electrostatic interactions.

[0141] Also included in the present disclosure are fragments or variants of polypeptides, and any combination thereof. The term "fragment" or "variant" when referring to polypeptide binding domains or binding molecules of the present disclosure include any polypeptides which retain at least some of the properties (e.g., FcRn binding affinity for an FcRn binding domain or Fc variant, coagulation activity for an FVIII variant, or FVIII binding activity for the VWF fragment) of the reference polypeptide. Fragments of polypeptides include proteolytic fragments, as well as deletion fragments, in addition to specific antibody fragments discussed elsewhere herein, but do not include the naturally occurring full-length polypeptide (or mature polypeptide). Variants of polypeptide binding domains or binding molecules of the present disclosure include fragments as described above, and also polypeptides with altered amino acid sequences due to amino acid substitutions, deletions, or insertions. Variants can be naturally or non-naturally occurring. Non- naturally occurring variants can be produced using art-known mutagenesis techniques. Variant polypeptides can comprise conservative or non-conservative amino acid substitutions, deletions or additions.

[0142] A "conservative amino acid substitution" is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art, including basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g. , alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), beta-branched side chains (e.g., threonine, valine, isoleucine) and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Thus, if an amino acid in a polypeptide is replaced with another amino acid from the same side chain family, the substitution is considered to be conservative. In another embodiment, a string of amino acids can be conservatively replaced with a structurally similar string that differs in order and / or composition of side chain family members.

[0143] The term "percent identity" as known in the art, is a relationship between two or more polypeptide sequences or two or more polynucleotide sequences, as determined by comparing the sequences. In the art, "identity" also means the degree of sequence relatedness between polypeptide or polynucleotide sequences, as the case can be, as determined by the match between strings of such sequences. "Identity" can be readily calculated by known methods, including but not limited to those described in: Computational Molecular Biology (Lesk, A. M., ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, D. W., ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, A. M., and Griffin, H. G., eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Stockton Press, New York (1991). Preferred methods to determine identity are designed to give the best match between the sequences tested. Methods to determine identity are codified in publicly available computer programs. Sequence alignments and percent identity calculations can be performed using sequence analysis software such as the Megalign program of the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, Wl), the GCG suite of programs (Wisconsin Package Version 9.0, Genetics Computer Group (GCG), Madison, Wl), BLASTP, BLASTN, BLASTX (Altschul et al., J. Mol. Biol. 215:403 (1990)), and DNASTAR (DNASTAR, Inc. 1228 S. Park St. Madison, Wl 53715 USA). Within the context of this application it will be understood that where sequence analysis software is used for analysis, that the results of the analysis will be based on the "default values" of the program referenced, unless otherwise specified. As used herein "default values" will mean any set of values or parameters which originally load with the software when first initialized. For the purposes of determining percent identity between a therapeutic protein, e.g. , a clotting factor, sequence of the disclosure and a reference sequence, only nucleotides in the reference sequence corresponding to nucleotides in the therapeutic protein, e.g., the clotting factor, sequence of the disclosure are used to calculate percent identity. For example, when comparing a full length FVIII nucleotide sequence containing the B domain to an optimized B domain deleted (BDD) FVIII nucleotide sequence of the disclosure, the portion of the alignment including the A1 , A2, A3, C1 , and C2 domain will be used to calculate percent identity. The nucleotides in the portion of the full length FVIII sequence encoding the B domain (which will result in a large "gap" in the alignment) will not be counted as a mismatch. In addition, in determining percent identity between an optimized BDD FVIII sequence of the disclosure, or a designated portion thereof (e.g., nucleotides 58-2277 and 2320-4374 of SEQ ID NO:3), and a reference sequence, percent identity will be calculated by aligning dividing the number ofmatched nucleotides by the total number of nucleotides in the complete sequence of the optimized BDD-FVIII sequence, or a designated portion thereof, as recited herein.

[0144] As used herein, nucleotides corresponding to nucleotides in a particular sequence of the disclosure are identified by alignment of the sequence of the disclosure to maximize the identity to a reference sequence. The number used to identify an equivalent amino acid in a reference sequence is based on the number used to identify the corresponding amino acid in the sequence of the disclosure.

[0145] A "fusion" or "chimeric" protein comprises a first amino acid sequence linked to a second amino acid sequence with which it is not naturally linked in nature. The amino acid sequences which normally exist in separate proteins can be brought together in the fusion polypeptide, or the amino acid sequences which normally exist in the same protein can be placed in a new arrangement in the fusion polypeptide, e.g., fusion of a Factor VIII domain of the disclosure with an Ig Fc domain. A fusion protein is created, for example, by chemical synthesis, or by creating and translating a polynucleotide in which the peptide regions are encoded in the desired relationship. A chimeric protein can further comprises a second amino acid sequence associated with the first amino acid sequence by a covalent, non-peptide bond or a non-covalent bond.

[0146] As used herein, the term "insertion site" refers to a position in a polypeptide, or fragment, variant, or derivative thereof, which is immediately upstream of the position at which a heterologous moiety can be inserted. An "insertion site" is specified as a number, the number being the number of the amino acid in a reference sequence. For example, an "insertion site" in FVIII refers to the number of the amino acid sequence in mature native FVIII (SEQ ID NO: 15) to which the insertion site corresponds, which is immediately N-terminal to the position of the insertion. For example, the phrase "a3 comprises a heterologous moiety at an insertion site which corresponds to amino acid 1656 of SEQ ID NO: 15" indicates that the heterologous moiety is located between two amino acids corresponding to amino acid 1656 and amino acid 1657 of SEQ ID NO: 15.

[0147] The phrase "immediately downstream of an amino acid" as used herein refers to position right next to the terminal carboxyl group of the amino acid. Similarly, the phrase "immediately upstream of an amino acid" refers to the position right next to the terminal amine group of the amino acid.

[0148] The terms "inserted," "is inserted," "inserted into" or grammatically related terms, as used herein refer to the position of a heterologous moiety in a polypeptide, e.g., a clotting factor, relative to the analogous position in the parental polypeptide. For example, in certain embodiment, "inserted" and the like refer to the position of a heterologous moiety in arecombinant FVIII polypeptide, relative to the analogous position in native mature human FVIII. As used herein the terms refer to the characteristics of the polypeptide, and do not indicate, imply or infer any methods or process by which the polypeptide was made.

[0149] As used herein, the term "half-life" refers to a biological half-life of a particular polypeptide in vivo. Half-life can be represented by the time required for half the quantity administered to a subject to be cleared from the circulation and / or other tissues in the animal. When a clearance curve of a given polypeptide is constructed as a function of time, the curve is usually biphasic with a rapid a-phase and longer b-phase. The a-phase typically represents an equilibration of the administered Fc polypeptide between the intra- and extra-vascular space and is, in part, determined by the size of the polypeptide. The b-phase typically represents the catabolism of the polypeptide in the intravascular space. In some embodiments, the therapeutic protein, e.g. , the clotting factor, e.g., FVIII, and chimeric proteins comprising the same are monophasic, and thus do not have an alpha phase, but just the single beta phase. Therefore, in certain embodiments, the term half-life as used herein refers to the half-life of the polypeptide in the b-phase.

[0150] The term "linked" as used herein refers to a first amino acid sequence or nucleotide sequence covalently or non-covalently joined to a second amino acid sequence or nucleotide sequence, respectively. The first amino acid or nucleotide sequence can be directly joined or juxtaposed to the second amino acid or nucleotide sequence or alternatively an intervening sequence can covalently join the first sequence to the second sequence. The term "linked" means not only a fusion of a first amino acid sequence to a second amino acid sequence at the C-terminus or the N-terminus, but also includes insertion of the whole first amino acid sequence (or the second amino acid sequence) into any two amino acids in the second amino acid sequence (or the first amino acid sequence, respectively). In one embodiment, the first amino acid sequence can be linked to a second amino acid sequence by a peptide bond or a linker. The first nucleotide sequence can be linked to a second nucleotide sequence by a phosphodiester bond or a linker. The linker can be a peptide or a polypeptide (for polypeptide chains) or a nucleotide or a nucleotide chain (for nucleotide chains) or any chemical moiety (for both polypeptide and polynucleotide chains). The term "linked" is also indicated by a hyphen (-).

[0151] Hemostasis, as used herein, means the stopping or slowing of bleeding or hemorrhage; or the stopping or slowing of blood flow through a blood vessel or body part.

[0152] Hemostatic disorder, as used herein, means a genetically inherited or acquired condition characterized by a tendency to hemorrhage, either spontaneously or as a result of trauma, due to an impaired ability or inability to form a fibrin clot. Examples of such disorders include the hemophilias. The three main forms are hemophilia A (factor VIII deficiency),hemophilia B (factor IX deficiency or "Christmas disease") and hemophilia C (factor XI deficiency, mild bleeding tendency). Other hemostatic disorders include, e.g., von Willebrand disease, Factor XI deficiency (PTA deficiency), Factor XII deficiency, deficiencies or structural abnormalities in fibrinogen, prothrombin, Factor V, Factor VII, Factor X or factor XIII, Bernard-Soulier syndrome, which is a defect or deficiency in GPIb. GPIb, the receptor for vWF, can be defective and lead to lack of primary clot formation (primary hemostasis) and increased bleeding tendency), and thrombasthenia of Glanzman and Naegeli (Glanzmann thrombasthenia). In liver failure (acute and chronic forms), there is insufficient production of coagulation factors by the liver; this can increase bleeding risk.

[0153] The isolated nucleic acid molecules, isolated polypeptides, or vectors comprising the isolated nucleic acid molecule of the disclosure can be used prophylactically. As used herein the term "prophylactic treatment" refers to the administration of a molecule prior to a bleeding episode. In one embodiment, the subject in need of a general hemostatic agent is undergoing, or is about to undergo, surgery. A polynucleotide, polypeptide, or vector of the disclosure can be administered prior to or after surgery as a prophylactic. The polynucleotide, polypeptide, or vector of the disclosure can be administered during or after surgery to control an acute bleeding episode. The surgery can include, but is not limited to, liver transplantation, liver resection, dental procedures, or stem cell transplantation.

[0154] The isolated nucleic acid molecules, isolated polypeptides, or vectors of the disclosure are also used for on-demand treatment. The term "on-demand treatment" refers to the administration of an isolated nucleic acid molecule, isolated polypeptide, or vector in response to symptoms of a bleeding episode or before an activity that can cause bleeding. In one aspect, the on-demand treatment can be given to a subject when bleeding starts, such as after an injury, or when bleeding is expected, such as before surgery. In another aspect, the on-demand treatment can be given prior to activities that increase the risk of bleeding, such as contact sports.

[0155] As used herein the term "acute bleeding" refers to a bleeding episode regardless of the underlying cause. For example, a subject can have trauma, uremia, a hereditary bleeding disorder (e.g., factor VII deficiency) a platelet disorder, or resistance owing to the development of antibodies to clotting factors.

[0156] Treat, treatment, treating, as used herein refers to, e.g., the reduction in severity of a disease or condition; the reduction in the duration of a disease course; the amelioration of one or more symptoms associated with a disease or condition; the provision of beneficial effects to a subject with a disease or condition, without necessarily curing the disease or condition, or the prophylaxis of one or more symptoms associated with a disease or condition. In one embodiment, the term "treating" or "treatment" means maintaining, e.g., a FVIII trough level atleast about 1 lU / dL, 2 lU / dL, 3 lU / dL, 4 lU / dL, 5 lU / dL, 6 lU / dL, 7 lU / dL, 8 lU / dL, 9 lU / dL, 10 lU / dL, 1 1 lU / dL, 12 lU / dL, 13 lU / dL, 14 lU / dL, 15 lU / dL, 16 lU / dL, 17 lU / dL, 18 lU / dL, 19 lU / dL, 20 lU / dL, 25 lU / dL, 30 lU / dL, 35 lU / dL, 40 lU / dL, 45 lU / dL, 50 lU / dL, 55 lU / dL, 60 lU / dL, 65 lU / dL, 70 lU / dL, 75 lU / dL, 80 lU / dL, 85 lU / dL, 90 lU / dL, 95 lU / dL, 100 lU / dL, 105 lU / dL, 1 10 lU / dL, 1 15 lU / dL, 120 lU / dL, 125 lU / dL, 130 lU / dL, 135 lU / dL, 140 lU / dL, 145 lU / dL, or 150 lU / dL in a subject by administering an isolated nucleic acid molecule, isolated polypeptide or vector of the disclosure. In another embodiment, treating or treatment means maintaining a FVIII trough level between about 1 and about 150 lU / dL, about 1 and about 125 lU / dL, about 1 and about 100 lU / dL, about 1 and about 90 lU / dL, about 1 and about 85 lU / dL, about 1 and about 80 lU / dL, about 1 and about 75 lU / dL, about 1 and about 70 lU / dL, about 1 and about 65 lU / dL, about 1 and about 60 lU / dL, about 1 and about 55 lU / dL, about 1 and about 50 lU / dL, about 1 and about 45 lU / dL, about 1 and about 40 lU / dL, about 1 and about 35 lU / dL, about 1 and about 30 lU / dL, about 1 and about 25 lU / dL, about 25 and about 125 lU / dL, about 50 and about 100 lU / dL, about 50 and about 75 lU / dL, about 75 and about 100 lU / dL, about 1 and about 20 lU / dL, about 2 and about 20 lU / dL, about 3 and about 20 lU / dL, about 4 and about 20 lU / dL, about 5 and about 20 lU / dL, about 6 and about 20 lU / dL, about 7 and about 20 lU / dL, about 8 and about 20 lU / dL, about 9 and about 20 lU / dL, or about 10 and about 20 lU / dL. Treatment or treating of a disease or condition can also include maintaining FVIII activity in a subject at a level comparable to at least about 1 %, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 1 1 %, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 105%, 1 10%, 115%, 120%, 125%, 130%, 135%, 140%, 145%, or 150% of the FVIII activity in a non-hemophiliac subject. The minimum trough level required for treatment can be measured by one or more known methods and can be adjusted (increased or decreased) for each person.

[0157] "Administering," as used herein, means to give a pharmaceutically acceptable nucleic acid molecule, polypeptide expressed therefrom, or vector comprising the nucleic acid molecule of the disclosure to a subject via a pharmaceutically acceptable route. Routes of administration can be intravenous, e.g., intravenous injection and intravenous infusion. Additional routes of administration include, e.g., subcutaneous, intramuscular, oral, nasal, and pulmonary administration. The nucleic acid molecules, polypeptides, and vectors can be administered as part of a pharmaceutical composition comprising at least one excipient.

[0158] The term "pharmaceutically acceptable" as used herein refers to molecular entities and compositions that are physiologically tolerable and do not typically produce toxicity or an allergic or similar untoward reaction, such as gastric upset, dizziness and the like, when administered to a human. Optionally, as used herein, the term "pharmaceutically acceptable"means approved by a regulatory agency of the federal or a state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeia for use in animals, and more particularly in humans.

[0159] As used herein, the phrase "subject in need thereof" includes subjects, such as mammalian subjects, that would benefit from administration of a nucleic acid molecule, polypeptide, or vector of the disclosure, e.g., to improve hemostasis. In one embodiment, the subjects include, but are not limited to, individuals with hemophilia. In another embodiment, the subjects include, but are not limited to, individuals who have developed an inhibitor to the therapeutic protein, e.g., the clotting factor, e.g., FVIII, and thus are in need of a bypass therapy. The subject can be an adult or a minor (e.g., under 12 years old).

[0160] As used herein, the term "therapeutic protein" refers to any polypeptide known in the art that can be administered to a subject. In some embodiments, the therapeutic protein comprises a protein selected from a clotting factor, a growth factor, an antibody, a functional fragment thereof, or a combination thereof. As used herein, the term "clotting factor," refers to molecules, or analogs thereof, naturally occurring or recombinantly produced which prevent or decrease the duration of a bleeding episode in a subject. In other words, it means molecules having pro-clotting activity, i.e., are responsible for the conversion of fibrinogen into a mesh of insoluble fibrin causing the blood to coagulate or clot. "Clotting factor" as used herein includes an activated clotting factor, its zymogen, or an activatable clotting factor. An "activatable clotting factor" is a clotting factor in an inactive form (e.g., in its zymogen form) that is capable of being converted to an active form. The term "clotting factor" includes but is not limited to factor I (FI), factor II (Fll), factor V (FV), FVII, FVIII, FIX, factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FXIII), Von Willebrand factor (VWF), prekallikrein, high-molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, Protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator(tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), zymogens thereof, activated forms thereof, or any combination thereof.

[0161] Clotting activity, as used herein, means the ability to participate in a cascade of biochemical reactions that culminates in the formation of a fibrin clot and / or reduces the severity, duration or frequency of hemorrhage or bleeding episode.

[0162] A "growth factor," as used herein, includes any growth factor known in the art including cytokines and hormones. In some embodiments, the growth factor is selected from adrenomedullin (AM), angiopoietin (Ang), autocrine motility factor, a bone morphogenetic protein (BMP) (e.g. BMP2, BMP4, BMP5, BMP7), a ciliary neurotrophic factor family member (e.g., ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6)), a colony-stimulating factor (e.g., macrophage colony-stimulating factor (m-CSF), granulocyte colony- stimulating factor (G-CSF), granulocyte macrophage colony-stimulating factor (GM-CSF)), an epidermal growth factor (EGF), an ephrin (e.g., ephrin A1 , ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1 , ephrin B2, ephrin B3), erythropoietin (EPO), a fibroblast growth factor (FGF) (e.g. , FGF1 , FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF1 1 , FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21 , FGF22, FGF23), foetal bovine somatotrophin (FBS), a GDNF family member (e.g., glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, artemin), growth differentiation factor-9 (GDF9), hepatocyte growth factor (HGF), hepatoma-derived growth factor (HDGF), insulin, an insulin-like growth factors (e.g., insulin-like growth factor-1 (IGF-1) or IGF-2, an interleukin (IL) (e.g., IL-1 , IL-2, IL- 3, IL-4, IL-5, IL-6, IL-7), keratinocyte growth factor (KGF), migration-stimulating factor (MSF), macrophage-stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)), myostatin (GDF-8), a neuregulin (e.g., neuregulin 1 (NRG1), NRG2, NRG3, NRG4), a neurotrophin (e.g., brain-derived neurotrophic factor (BDNF), nerve growth factor (NGF), a neurotrophin-3 (NT-3), NT-4, placental growth factor (PGF), platelet-derived growth factor (PDGF), renalase (RNLS), T-cell growth factor (TCGF), thrombopoietin (TPO), a transforming growth factor (e.g., transforming growth factor alpha (TGF-a), TGF-b, tumor necrosis factor-alpha (TNF-a), and vascular endothelial growth factor (VEGF).

[0163] In some embodiments, the therapeutic protein is encoded by a gene selected from dystrophin X-linked, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1 , FXN (frataxin), GUCY2D, RS1 , CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof.

[0164] As used herein the terms "heterologous" or "exogenous" refer to such molecules that are not normally found in a given context, e.g., in a cell or in a polypeptide. For example, an exogenous or heterologous molecule can be introduced into a cell and are only present after manipulation of the cell, e.g., by transfection or other forms of genetic engineering or a heterologous amino acid sequence can be present in a protein in which it is not naturally found.

[0165] As used herein, the term "heterologous nucleotide sequence" refers to a nucleotide sequence that does not naturally occur with a given polynucleotide sequence. In one embodiment, the heterologous nucleotide sequence encodes a polypeptide capable of extending the half-life of the therapeutic protein, e.g., the clotting factor, e.g., FVI II. In another embodiment, the heterologous nucleotide sequence encodes a polypeptide that increases the hydrodynamic radius of the therapeutic protein, e.g. , the clotting factor, e.g., FVI 11. In other embodiments, the heterologous nucleotide sequence encodes a polypeptide that improves one or morepharmacokinetic properties of the therapeutic protein without significantly affecting its biological activity or function (e.g., a procoagulant activity). In some embodiments, the therapeutic protein is linked or connected to the polypeptide encoded by the heterologous nucleotide sequence by a linker. Non-limiting examples of polypeptide moieties encoded by heterologous nucleotide sequences include an immunoglobulin constant region or a portion thereof, albumin or a fragment thereof, an albumin-binding moiety, a transferrin, the PAS polypeptides of U.S. Pat Application No. 20100292130, a HAP sequence, transferrin or a fragment thereof, the C-terminal peptide (CTP) of the b subunit of human chorionic gonadotropin, albumin-binding small molecule, an XTEN sequence, FcRn binding moieties (e.g., complete Fc regions or portions thereof which bind to FcRn), single chain Fc regions (ScFc regions, e.g., as described in US 2008 / 0260738, WO 2008 / 012543, or WO 2008 / 1439545), polyglycine linkers, polyserine linkers, peptides and short polypeptides of 6-40 amino acids of two types of amino acids selected from glycine (G), alanine (A), serine (S), threonine (T), glutamate (E) and proline (P) with varying degrees of secondary structure from less than 50% to greater than 50%, amongst others, or two or more combinations thereof. In some embodiments, the polypeptide encoded by the heterologous nucleotide sequence is linked to a non-polypeptide moiety. Non-limiting examples of the non-polypeptide moieties include polyethylene glycol (PEG), albumin-binding small molecules, polysialic acid, hydroxyethyl starch (HES), a derivative thereof, or any combinations thereof.

[0166] As used herein, the term "Fc region" is defined as the portion of a polypeptide which corresponds to the Fc region of native Ig, .e., as formed by the dimeric association of the respective Fc domains of its two heavy chains. A native Fc region forms a homodimer with another Fc region. In contrast, the term "genetically-fused Fc region" or "single-chain Fc region" (scFc region), as used herein, refers to a synthetic dimeric Fc region comprised of Fc domains genetically linked within a single polypeptide chain ( .e., encoded in a single contiguous genetic sequence).

[0167] In one embodiment, the "Fc region" refers to the portion of a single Ig heavy chain beginning in the hinge region just upstream of the papain cleavage site ( .e., residue 216 in IgG, taking the first residue of heavy chain constant region to be 114) and ending at the C-terminus of the antibody. Accordingly, a complete Fc domain comprises at least a hinge domain, a CH2 domain, and a CH3 domain.

[0168] The Fc region of an Ig constant region, depending on the Ig isotype can include the CH2, CH3, and CH4 domains, as well as the hinge region. Chimeric proteins comprising an Fc region of an Ig bestow several desirable properties on a chimeric protein including increased stability, increased serum half-life (see Capon et a!., 1989, Nature 337:525) as well as binding to Fc receptors such as the neonatal Fc receptor (FcRn) (U.S. Pat. Nos. 6,086,875, 6,485,726,6,030,613; WO 03 / 077834; US2003-0235536A1), which are incorporated herein by reference in their entireties.

[0169] A "reference nucleotide sequence," when used herein as a comparison to a nucleotide sequence of the disclosure, is a polynucleotide sequence essentially identical to the nucleotide sequence of the disclosure except that sequence is not optimized. For example, the reference nucleotide sequence for a nucleic acid molecule consisting of the codon optimized BDD FVIII of SEQ ID NO: 1 and a heterologous nucleotide sequence that encodes a single chain Fc region linked to SEQ ID NO: 1 at its 3' end is a nucleic acid molecule consisting of the original (or "parent") BDD FVIII of SEQ ID NO: 16 and the identical heterologous nucleotide sequence that encodes a single chain Fc region linked to SEQ ID NO: 16 at its 3' end.

[0170] As used herein, the term "optimized," with regard to nucleotide sequences, refers to a polynucleotide sequence that encodes a polypeptide, wherein the polynucleotide sequence has been mutated to enhance a property of that polynucleotide sequence. In some embodiments, the optimization is done to increase transcription levels, increase translation levels, increase steady-state mRNA levels, increase or decrease the binding of regulatory proteins such as general transcription factors, increase or decrease splicing, or increase the yield of the polypeptide produced by the polynucleotide sequence. Examples of changes that can be made to a polynucleotide sequence to optimize it include codon optimization, G / C content optimization, removal of repeat sequences, removal of AT rich elements, removal of cryptic splice sites, removal of cis-acting elements that repress transcription or translation, adding or removing poly- T or poly-A sequences, adding sequences around the transcription start site that enhance transcription, such as Kozak consensus sequences, removal of sequences that could form stem loop structures, removal of destabilizing sequences, and two or more combinations thereof.II. Nucleic Acid Molecules

[0171] The present disclosure is directed to a plasmid-like, capsid free, nucleic acid molecule that encodes a target sequence, wherein the target sequence encodes a therapeutic protein or a gene that can modulate expression of a target protein, e.g., a miRNA. A capsid, the protein shell of a virus, encloses the genetic material of the virus. Capsids are known to aid the functions of the virion by protecting the viral genome, delivering the genome to a host, and interacting with the host. Nonetheless, the viral capsids may be a factor in limiting the packaging capacity of the vectors and / or inducing immune responses, especially when used in gene therapy.

[0172] AAV vectors have emerged as one of the more common types of gene therapy vectors. However, the presence of the capsid limits the utility of an AAV vector in gene therapy. In particular, the capsid itself can limit the size of the transgene that is included in the vector toas low as less than 4.5 kb. Various therapeutic proteins that may be useful in a gene therapy can easily exceed this size even before regulatory elements are added.

[0173] Furthermore, proteins that make up the capsid can serve as antigens that can be targeted by a subject’s immune system. AAV is very common in the general population, with most people having been exposed to an AAV throughout their lives. As a result, most potential gene therapy recipients have likely already developed an immune response to an AAV, and thus are more likely to reject the therapy.

[0174] Certain aspects of the present disclosure aim to overcome these deficiencies ofAAV vectors. In particular, certain aspects of the present disclosure are directed to a nucleic acid molecule, comprising a first ITR, a second ITR, and a genetic cassette, e.g., encoding a therapeutic protein and / or a miRNA. In some embodiments, the first ITR and second ITR flank a genetic cassette comprising a heterologous polynucleotide sequence. In some embodiments, the nucleic acid molecule does not comprise a gene encoding a capsid protein, a replication protein, and / or an assembly protein. In some embodiments, the genetic cassette encodes a therapeutic protein. In some embodiments, the therapeutic protein comprises a clotting factor. In some embodiments, the genetic cassette encodes a miRNA. In certain embodiments, the genetic cassette is positioned between the first ITR and the second ITR. In some embodiments, the nucleic acid molecule further comprises one or more noncoding region. In certain embodiments, the one or more non-coding region comprises a promoter sequence, an intron, a post- transcriptional regulatory element, a 3'UTR poly(A) sequence, or any combination thereof.

[0175] In one embodiment, the genetic cassette is a single stranded nucleic acid. In another embodiment, the genetic cassette is a double stranded nucleic acid.

[0176] In one embodiment, the nucleic acid molecule comprises:(a) a first ITR that is an ITR of a non-AAV family member of Parvoviridae (e.g., a B19 or GPV ITR);(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a nucleotide encoding a miRNA or a therapeutic protein, e.g., a clotting factor;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA;(g) a second ITR that is an ITR of a non-AAV family member of Parvoviridae (e.g., a B19 or GPV ITR).

[0177] In one embodiment, the nucleic acid molecule comprises:(a) a first ITR that is an ITR of a non-AAV family member of Parvoviridae·,(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a nucleotide encoding a miRNA, wherein the miRNA down regulates the expression of a target gene selected from SOD1 , HTT, RHO, and any combination thereof;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA;(g) a second ITR that is an ITR of a non-AAV family member of Parvoviridae

[0178] In one embodiment, the nucleic acid molecule comprises:(a) a first ITR that is an ITR of a non-AAV family member of Parvoviridae·,(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a nucleotide encoding dystrophin X-linked, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1 , FXN (frataxin), GUCY2D, RS1 , CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha- glucosidase), or any combination thereof;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA;(g) a second ITR that is an ITR of a non-AAV family member of Parvoviridae

[0179] In one embodiment, the nucleic acid molecule comprises:(a) a first ITR that is an ITR of an AAV, e.g., an AAV serotype 2 genome;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a nucleotide encoding FVIII; wherein the nucleotide has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a nucleotide sequence selected from SEQ ID NOs: 1 -14 or SEQ ID NO: 71 , wherein the FVIII encoded by the nucleotide retains a FVIII activity;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and(g) a second ITR that is an ITR of an AAV, e.g., an AAV serotype 2 genome.

[0180] In one embodiment, the nucleic acid molecule comprises:(a) a first ITR that is an ITR of an AAV, e.g., an AAV serotype 2 genome;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a nucleotide encoding a miRNA, wherein the miRNA down regulates the expression of a target gene, e.g., SOD1 , HTT, RHO, and any combination thereof;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and(g) a second ITR that is an ITR of an AAV, e.g., an AAV serotype 2 genome.

[0181] In one embodiment, the nucleic acid molecule comprises:(a) a first ITR that is an ITR of an AAV, e.g., an AAV serotype 2 genome;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a nucleotide encoding dystrophin X-linked, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1 , FXN (frataxin), GUCY2D, RS1 , CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha- glucosidase), or any combination thereof;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and(g) a second ITR that is an ITR of an AAV, e.g., an AAV serotype 2 genome.

[0182] In another embodiment, the nucleic acid molecule comprises:(a) a first ITR;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a nucleotide encoding a miRNA or a therapeutic protein, e.g., clotting factor;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and(g) a second ITR,wherein one of the first ITR or the second ITR is an ITR of a non-AAV family member of Parvoviridae and the other ITR is an ITR of an AAV, e.g., an AAV serotype 2 genome.

[0183] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the AAV2 5’ ITR sequence set forth in SEQ ID NO: 1 1 1 ;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding FVIII, e.g., FVIIIco6XTEN;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the AAV2 3’ ITR sequence set forth in SEQ ID NO: 124.

[0184] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the AAV2 5’ ITR sequence set forth in SEQ ID NO: 1 1 1 ;(b) a tissue specific promoter sequence, e.g., CAG promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding FVIII, e.g., FVIIIco6XTEN;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the AAV2 3’ ITR sequence set forth in SEQ ID NO: 193.

[0185] In another embodiment, the nucleic acid molecule comprises:(a) a first ITR;(b) a tissue specific promoter sequence, TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a nucleotide encoding a miRNA or a therapeutic protein, e.g., clotting factor;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and(g) a second ITR,wherein the first ITR is a synthetic ITR, the second ITR is a synthetic ITR, or both the first ITR and the second ITR are synthetic ITRs.

[0186] In another embodiment, the nucleic acid molecule comprises:(a) a first B19 ITR;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding therapeutic protein selected from the group consisting of a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, and a combination thereof;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a second B19 ITR.

[0187] In another embodiment, the nucleic acid molecule comprises:(a) a first GPV ITR;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding therapeutic protein selected from the group consisting of a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, and a combination thereof;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a second GPV ITR.

[0188] In another embodiment, the nucleic acid molecule comprises:(a) a first B19 ITR;(b) a ubiquitous promoter sequence, e.g., CAG promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding therapeutic protein selected from the group consisting of a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, and a combination thereof;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a second B19 ITR.

[0189] In another embodiment, the nucleic acid molecule comprises:(a) a first GPV ITR;(b) a ubiquitous promoter sequence, e.g., CAG promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding therapeutic protein selected from the group consisting of a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, and a combination thereof;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a second GPV ITR.

[0190] In another embodiment, the nucleic acid molecule comprises:(a) a first B19 ITR;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding phenylalanine hydroxylase (PAH);(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a second B19 ITR.

[0191] In another embodiment, the nucleic acid molecule comprises:(a) a first GPV ITR;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding phenylalanine hydroxylase (PAH);(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a second GPV ITR.

[0192] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the B19d135 5’ ITR sequence set forth in SEQ ID NO: 180;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding FVIII, e.g., FVIIIco6XTEN;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the B19d135 3’ ITR sequence set forth in SEQ ID NO: 181.

[0193] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the GPVd162 5’ ITR sequence set forth in SEQ ID NO: 183;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding FVIII, e.g., FVIIIco6XTEN;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the GPVd162 3’ ITR sequence set forth in SEQ ID NO: 184.

[0194] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the full length B19 5’ ITR sequence set forth in SEQ ID NO: 185;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding FVIII, e.g., FVIIIco6XTEN;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the full length B19 3’ ITR sequence set forth in SEQ ID NO: 186.

[0195] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the full length GPV 5’ ITR sequence set forth in SEQ ID NO: 187;(b) a tissue specific promoter sequence, e.g., TTP promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding FVIII, e.g., FVIIIco6XTEN;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the full length GPV 3’ ITR sequence set forth in SEQ ID NO: 188.

[0196] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the B19d135 5’ ITR sequence set forth in SEQ ID NO: 180;(b) a tissue specific promoter sequence, e.g., CAG promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding PAH;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the B19d135 3’ ITR sequence set forth in SEQ ID NO: 181.

[0197] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the GPVd162 5’ ITR sequence set forth in SEQ ID NO: 183;(b) a tissue specific promoter sequence, e.g., CAG promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding PAH;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the GPVd162 3’ ITR sequence set forth in SEQ ID NO: 184.

[0198] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the full length B19 5’ ITR sequence set forth in SEQ ID NO: 185;(b) a tissue specific promoter sequence, e.g., CAG promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding PAH;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the full length B19 3’ ITR sequence set forth in SEQ ID NO: 186.

[0199] In another embodiment, the nucleic acid molecule comprises:(a) a 5’ ITR bearing the full length GPV 5’ ITR sequence set forth in SEQ ID NO: 187;(b) a tissue specific promoter sequence, e.g., CAG promoter;(c) an intron, e.g., a synthetic intron;(d) a heterologous polynucleotide sequence encoding PAH;(e) a post-transcriptional regulatory element, e.g., WPRE;(f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and / or(g) a 3’ ITR bearing the full length GPV 3’ ITR sequence set forth in SEQ ID NO: 188.A. Inverted Terminal Repeats

[0200] Certain aspects of the present disclosure are directed to a nucleic acid molecule comprising a first ITR, e.g. , a 5' ITR, and second ITR, e.g., a 3' ITR. Typically, ITRs are involved in parvovirus (e.g., AAV) DNA replication and rescue, or excision, from prokaryotic plasmids (Samulski et al., 1983, 1987; Senapathy et al., 1984; Gottlieb and Muzyczka, 1988). In addition, ITRs appear to be the minimum sequences required for AAV proviral integration and for packaging of AAV DNA into virions (McLaughlin et al., 1988; Samulski et al., 1989). Theseelements are essential for efficient multiplication of a parvovirus genome. It is hypothesized that the minimal defining elements indispensable for ITR function are a Rep-binding site (e.g., RBS; GCGCGCTCGCTCGCTC (SEQ ID NO: 104) for AAV2) and a terminal resolution site (e.g. , TRS; AGTTGG (SEQ ID NO: 105) for AAV2) plus a variable palindromic sequence allowing for hairpin formation. Palindromic nucleotide regions normally function together in cis as origins of DNA replication and as packaging signals for the virus. Complimentary sequences in the ITRs fold into a hairpin structure during DNA replication. In some embodiments, the ITRs fold into a hairpin T- shaped structure. In other embodiments, the ITRs fold into non-T-shaped hairpin structures, e.g., into a U-shaped hairpin structure. Data suggests that the T-shaped hairpin structures of AAV ITRs may inhibit the expression of a transgene flanked by the ITRs. See, e.g., Zhou et al., Scientific Reports 7:5432 (July 14, 2017). By utilizing an ITR that does not form T-shaped hairpin structures, this form of inhibition may be avoided. Therefore, in certain aspects, a polynucleotide comprising a non-AAV ITR has an improved transgene expression compared to a polynucleotide comprising an AAV ITR that forms a T-shaped hairpin.

[0201] In some embodiments, the ITR comprises a naturally occurring ITR, e.g. the ITR comprises all or a portion of a parvovirus ITR. In some embodiments, the ITR comprises a synthetic sequence. In one embodiment, the first ITR or the second ITR comprises a synthetic sequence. In another embodiment, each of the first ITR and the second ITR comprises a synthetic sequence. In some embodiments, the first ITR or the second ITR comprises a naturally occurring sequence. In another embodiment, each of the first ITR and the second ITR comprises a naturally occurring sequence.

[0202] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR, e.g., a truncated ITR. In some embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides, or at least about 600 nucleotides; wherein the ITR retains a functional property of the naturally occurring ITR. In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 129 nucleotides; wherein the ITR retains a functional property of the naturally occurring ITR. In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 102 nucleotides; wherein the ITR retains a functional property of the naturally occurring ITR.

[0203] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR, wherein the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the length of the naturally occurring ITR; wherein the fragment retains a functional property of the naturally occurring ITR.

[0204] In certain embodiments, the ITR comprises or consists of a sequence that has a sequence identity of at least 50%, at least 51 %, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61 %, at least62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least69%, at least 70%, at least 71 %, at least 72%, at least 73%, at least 74%, at least 75%, at least76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81 %, at least 82%, at least83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least97%, at least 98%, at least 99%, or 100% to a homologous portion of a naturally occurring ITR, when properly aligned; wherein the ITR retains a functional property of the naturally occurring ITR. In other embodiments, the ITR comprises or consists of a sequence that has a sequence identity of at least 90% to a homologous portion of a naturally occurring ITR, when properly aligned; wherein the ITR retains a functional property of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that has a sequence identity of at least 80% to a homologous portion of a naturally occurring ITR, when properly aligned; wherein the ITR retains a functional property of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that has a sequence identity of at least 70% to a homologous portion of a naturally occurring ITR, when properly aligned; wherein the ITR retains a functional property of the naturally occurring ITR. In some embodiments, the ITR comprises orconsists of a sequence that has a sequence identity of at least 60% to a homologous portion of a naturally occurring ITR, when properly aligned; wherein the ITR retains a functional property of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that has a sequence identity of at least 50% to a homologous portion of a naturally occurring ITR, when properly aligned; wherein the ITR retains a functional property of the naturally occurring ITR.

[0205] In some embodiments, the ITR comprises an ITR from an AAV genome. In some embodiments, the ITR is an ITR of an AAV genome selected from AAV1 , AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10 AAV11 , and any combination thereof. In a particular embodiment, the ITR is an ITR of the AAV2 genome. In another embodiment, the ITR is a synthetic sequence genetically engineered to include at its 5' and 3' ends ITRs derived from one or more of AAV genomes.

[0206] In some embodiments, the ITR is not derived from an AAV genome. In some embodiments, the ITR is an ITR of a non-AAV. In some embodiments, the ITR is an ITR of a non- AAV genome from the viral family Parvoviridae selected from, but not limited to, the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus and any combination thereof. In certain embodiments, the ITR is derived from erythrovirus parvovirus B19 (human virus). In another embodiment, the ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, e.g. , MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, e.g., MDPV strain YY. In some embodiments, the ITR is derived from a porcine parvovirus, e.g. , porcine parvovirus U44978. In some embodiments, the ITR is derived from a mice minute virus, e.g., mice minute virus U34256. In some embodiments, the ITR is derived from a canine parvovirus, e.g. , canine parvovirus M19296. In some embodiments, the ITR is derived from a mink enteritis virus, e.g., mink enteritis virus D00765. In some embodiments, the ITR is derived from a Dependoparvovirus. In one embodiment, the Dependoparvovirus is a Dependovirus Goose parvovirus (GPV) strain. In a specific embodiment, the GPV strain is attenuated, e.g., GPV strain 82-0321 V. In another specific embodiment, the GPV strain is pathogenic, e.g. , GPV strain B.

[0207] The first ITR and the second ITR of the nucleic acid molecule can be derived from the same genome, e.g., from the genome of the same virus, or from different genomes, e.g., from the genomes of two or more different virus genomes. In certain embodiments, the first ITR and the second ITR are derived from the same AAV genome. In a specific embodiment, the two ITRs present in the nucleic acid molecule of the invention are the same, and can in particular be AAV2ITRs. In other embodiments, the first ITR is derived from an AAV genome and the second ITR is not derived from an AAV genome (e.g. , a non-AAV genome). In other embodiments, the first ITR is not derived from an AAV genome (e.g., a non-AAV genome) and the second ITR is derived from an AAV genome. In still other embodiments, both the first ITR and the second ITR are not derived from an AAV genome (e.g., a non-AAV genome). In one particular embodiment, the first ITR and the second ITR are identical.

[0208] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a genome selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus,Brevidensovirus, Hepandensovirus, Penstyldensovirus and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a genome selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus,Penstyldensovirus, and any combination thereof. In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus,Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from the same genome. In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from the different genomes.

[0209] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from erythrovirus parvovirus B19 (human virus). In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from erythrovirus parvovirus B19 (human virus).

[0210] In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from B19. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, atleast about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171 , wherein the first ITR and / or the second ITR retains a functional property of the B19 ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171 , wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0211] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 167. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 168. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 169. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 170. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 171.Table 1. Sample Parvovirus ITR Sequences.

[0212] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 167. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 168. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 168. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 169, and wherein the nucleotide sequence is a at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167, retains a functional property of the B19 ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 169, and wherein the nucleotide sequence is a at least about50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0213] In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from B19. In some embodiments, the second ITR is a reverse complement of the first ITR. In some embodiments, the first ITR is a reverse complement of the second ITR. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 180, 181 , 185, and 186, or a functional derivative thereof. In some embodiments, the functional derivative retains a functional property of the B19 ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 180, 181 , 185, and 186, or a functional derivative thereof. In some embodiments, the functional derivative is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0214] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 180. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 180. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO:181. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 181. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 185. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 185. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 186. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 186.

[0215] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 180, 181 , 185, and 186. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 180. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 181. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 185. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 186. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 180. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 181. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 185. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 186.

[0216] In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO:180, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 181. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 181 , and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 180. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO:185, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 186. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 186, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 185.

[0217] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from GPV. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from GPV.

[0218] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 172. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 173. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 173. In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from GPV. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 172, 173, 174, 175, and 176, wherein the first ITR and / or the second ITR retains a functional property of the GPV ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from GPV. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 172, 173, 174, 175, and 176, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 172, 173, 174, 175, and 176. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 172. In some embodiments, the first ITR and / or thesecond ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 173. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 174. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 175. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 176.

[0219] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 174, and wherein the nucleotide sequence is a at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR retains a functional property of the GPV ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 174, and wherein the nucleotide sequence is a at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0220] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 176, and wherein the nucleotide sequence is a at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR retains a functional property of the GPV ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 176, and wherein the nucleotide sequence is a at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0221] In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from GPV. In some embodiments, the second ITR is a reverse complement of the first ITR. In some embodiments, the first ITR is a reverse complement of the second ITR. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 183, 184, 187 and 188, or a functional derivative thereof. In some embodiments, the functional derivative retains a functional property of the GPV ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs:183, 184, 187 and 188, or a functional derivative thereof. In some embodiments, the functional derivative is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0222] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 183. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 183. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO:184. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 184. Incertain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 187. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 187. In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 188. In certain embodiments, the first ITR and / or the second ITR consists of SEQ ID NO: 188.

[0223] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 183, 184, 187 and 188. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 183. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 184. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 187. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 188. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 183. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 184. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 187. In some embodiments, the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 188.

[0224] In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO:183, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 184. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 184, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 183. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO:187, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 188. In some embodiments, the first ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 188, and the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 187.

[0225] In certain embodiments, one of the first ITR or the second ITR comprises or consists of all or a portion of an ITR derived from AAV2. In some embodiments, the first ITR or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NOs: 177 or 178, wherein the first ITR and / or the second ITR retains a functional property of the AAV2 ITR from which it is derived. In some embodiments, the first ITR or the second ITR comprises or consists of a nucleotide sequence at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NOs: 177 or 178, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NOs: 177 or 178. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 177. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 178.

[0226] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, e.g., MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, e.g., MDPV strain YY.

[0227] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a Dependoparvovirus. In some embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a Dependoparvovirus. In other embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a Dependovirus goose parvovirus (GPV) strain. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a Dependovirus GPV strain. In certain embodiments, the GPV strain is attenuated, e.g., GPV strain 82-0321V. In other embodiments, the GPV strain is pathogenic, e.g., GPV strain B.

[0228] In certain embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a genome selected from the group consisting of porcine parvovirus,e.g., porcine parvovirus strain U44978; mice minute virus, e.g., mice minute virus strain U34256; canine parvovirus, e.g., canine parvovirus strain M19296; mink enteritis virus, e.g., mink enteritis virus strain D00765; and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a genome selected from the group consisting of porcine parvovirus, e.g., porcine parvovirus strain U44978; mice minute virus, e.g., mice minute virus strain U34256; canine parvovirus, e.g., canine parvovirus strain M 19296; mink enteritis virus, e.g., mink enteritis virus strain D00765; and any combination thereof.

[0229] In another particular embodiment, the ITR is a synthetic sequence genetically engineered to include at its 5' and 3' ends ITRs not derived from an AAV genome. In another particular embodiment, the ITR is a synthetic sequence genetically engineered to include at its 5' and 3' ends ITRs derived from one or more of non-AAV genomes. The two ITRs present in the nucleic acid molecule of the invention can be the same or different non-AAV genomes. In particular, the ITRs can be derived from the same non-AAV genome. In a specific embodiment, the two ITRs present in the nucleic acid molecule of the invention are the same, and can in particular be AAV2 ITRs.

[0230] In some embodiments, the ITR sequence comprises one or more palindromic sequence. A palindromic sequence of an ITR disclosed herein includes, but is not limited to, native palindromic sequences (i.e., sequences found in nature), synthetic sequences (i.e., sequences not found in nature), such as pseudo palindromic sequences, and combinations or modified forms thereof. A "pseudo palindromic sequence" is a palindromic DNA sequence, including an imperfect palindromic sequence, which shares less than 80% including less than 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5%, or no, nucleic acid sequence identity to sequences in native AAV or non-AAV palindromic sequence which form a secondary structure. The native palindromic sequences can be obtained or derived from any genome disclosed herein. The synthetic palindromic sequence can be based on any genome disclosed herein.

[0231] The palindromic sequence can be continuous or interrupted. In some embodiments, the palindromic sequence is interrupted, wherein the palindromic sequence comprises an insertion of a second sequence. In some embodiments, the second sequence comprises a promoter, an enhancer, an integration site for an integrase (e.g., sites for Cre or Flp recombinase), an open reading frame for a gene product, or a combination thereof.

[0232] In some embodiments, the ITRs form hairpin loop structures. In one embodiment, the first ITR forms a hairpin structure. In another embodiment, the second ITR forms a hairpin structure. Still in another embodiment, both the first ITR and the second ITR form hairpin structures. In some embodiments, the first ITR and / or the second ITR does not form a T-shaped hairpin structure. In certain embodiments, the first ITR and / or the second ITR forms a non-T-shaped hairpin structure. In some embodiments, the non-T-shaped hairpin structure comprises a U-shaped hairpin structure.

[0233] In some embodiments, an ITR in a nucleic acid molecule described herein may be a transcriptionally activated ITR. A transcriptionally-activated ITR can comprise all or a portion of a wild-type ITR that has been transcriptionally activated by inclusion of at least one transcriptionally active element. Various types of transcriptionally active elements are suitable for use in this context. In some embodiments, the transcriptionally active element is a constitutive transcriptionally active element. Constitutive transcriptionally active elements provide an ongoing level of gene transcription, and are preferred when it is desired that the transgene be expressed on an ongoing basis. In other embodiments, the transcriptionally active element is an inducible transcriptionally active element. Inducible transcriptionally active elements generally exhibit low activity in the absence of an inducer (or inducing condition), and are up-regulated in the presence of the inducer (or switch to an inducing condition). Inducible transcriptionally active elements may be preferred when expression is desired only at certain times or at certain locations, or when it is desirable to titrate the level of expression using an inducing agent. Transcriptionally active elements can also be tissue-specific; that is, they exhibit activity only in certain tissues or cell types.

[0234] Transcriptionally active elements, can be incorporated into an ITR in a variety of ways. In some embodiments, a transcriptionally active element is incorporated 5' to any portion of an ITR or 3' to any portion of an ITR. In other embodiments, a transcriptionally active element of a transcriptionally-activated ITR lies between two ITR sequences. If the transcriptionally active element comprises two or more elements which must be spaced apart, those elements may alternate with portions of the ITR. In some embodiments, a hairpin structure of an ITR is deleted and replaced with inverted repeats of a transcriptional element. This latter arrangement would create a hairpin mimicking the deleted portion in structure. Multiple tandem transcriptionally active elements can also be present in a transcriptionally-activated ITR, and these may be adjacent or spaced apart. In addition, protein binding sites (e.g., Rep binding sites) can be introduced into transcriptionally active elements of the transcriptionally-activated ITRs. A transcriptionally active element can comprise any sequence enabling the controlled transcription of DNA by RNA polymerase to form RNA, and can comprise, for example, a transcriptionally active element, as defined below.

[0235] Transcriptionally-activated ITRs provide both transcriptional activation and ITR functions to the nucleic acid molecule in a relatively limited nucleotide sequence length which effectively maximizes the length of a transgene which can be carried and expressed from the nucleic acid molecule. Incorporation of a transcriptionally active element into an ITR can beaccomplished in a variety of ways. A comparison of the ITR sequence and the sequence requirements of the transcriptionally active element can provide insight into ways to encode the element within an ITR. For example, transcriptional activity can be added to an ITR through the introduction of specific changes in the ITR sequence that replicates the functional elements of the transcriptionally active element. A number of techniques exist in the art to efficiently add, delete, and / or change particular nucleotide sequences at specific sites (see, for example, Deng and Nickoloff (1992) Anal. Biochem. 200:81-88). Anotherway to create transcriptionally-activated ITRs involves the introduction of a restriction site at a desired location in the ITR. In addition, multiple transcriptionally activate elements can be incorporated into a transcriptionally-activated ITR, using methods known in the art.

[0236] By way of illustration, transcriptionally-activated ITRs can be generated by inclusion of one or more transcriptionally active elements such as: TATA box, GC box, CCAAT box, Sp1 site, Inr region, CRE (cAMP regulatory element) site, ATF-1 / CRE site, ARBb box, APBa box, CArG box, CCAC box, or any other element involved in transcription as known in the art.

[0237] Aspects of the present disclosure provide a method of cloning a nucleic acid molecule described herein, comprising inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector, and introducing the resulting vector into a suitable bacterial host strain. As known in the art, complex secondary structures (e.g., long palindromic regions) of nucleic acids may be unstable and difficult to clone in bacterial host strains. For example, nucleic acid molecules comprising a first ITR and a second ITR (e.g., non-AAV parvoviral ITRs, e.g., B19 or GPV ITRs) of the present disclosure may be difficult to clone using conventional methodologies. Long DNA plindromes inhibit DNA replication and are unstable in the genomes of E. coli, Bacillus, Steptococcus, Streptomyces, S. cerevisiae, mice, and humans. These effects result from the formation of hairpin or cruciform structures by intrastrand base pairing. In E. coli the inhibition of DNA replication can be significantly overcome in SbcC or SbcD mutants. SbcD is the nuclease subunit, and SbcC is the ATPase subunit of the SbcCD complex. The E. coli SbcCD complex is an exonuclease complex responsible for preventing the replication of long palindromes. The SbcCD complex is a nuclear with ATP-dependent double-stranded DNA exonuclease activity and ATP-independent single-stranded DNA endonuclease activity. SbcCD may recognize DNA plaindromes and collapse replication forks by attacking hairpin structures that arise.

[0238] In certain embodiments, a suitable bacterial host strain is incapable of resolving cruciform DNA structures. In certain embodiments, a suitable bacterial host strain comprises a disruption in the SbcCD complex. In some embodiments, the disruption in the SbcCD complex comprises a genetic disruption in the SbcC gene and / or SbcD gene. In certain embodiments,the disruption in the SbcCD complex comprises a genetic disruption in the SbcC gene. Various bacterial host strains that comprise a genetic disruption in the SbcC gene are known in the art. For example, without limitation, the bacterial host strain PMC103 comprises the genotype sbcC, recD, mcrA, AmcrBCF, the bacterial host strain PMC107 comprises the genotype recBC, recJ, sbcBC, mcrA, AmcrBCF, and the bacterial host strain SURE comprises the genotype recB, recJ, sbcC, mcrA, AmcrBCF, umuC, uvrC. Accordingly, in some embodiments a method of cloning a nucleic acid molecule described herein comprises inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector, and introducing the resulting vector into host strain PMC103, PMC107, or SURE. In certain embodiments, the method of cloning a nucleic acid molecule described herein comprises inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector, and introducing the resulting vector into host strain PMC103.

[0239] Suitable vectors are known in the art and described elsewhere herein. In certain embodiments, a suitable vector for use in a cloning methodology of the present disclosure is a low copy vector. In certain embodiments, a suitable vector for use in a cloning methodology of the present disclosure is pBR322.

[0240] Accordingly, the present disclosure provides a method of cloning a nucleic acid molecule, comprising inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex, wherein the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.B. Therapeutic Proteins

[0241] Certain aspects of the present disclosure are directed to a nucleic acid molecule comprising a first ITR, a second ITR, and a genetic cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein. In some embodiments, the genetic cassette encodes one therapeutic protein. In some embodiments, the genetic cassette encodes more than one therapeutic protein. In some embodiments, the genetic cassette encodes two or more copies of the same therapeutic protein. In some embodiments, the genetic cassette encodes two or more variants of the same therapeutic protein. In some embodiments, the genetic cassette encodes two or more different therapeutic proteins.

[0242] Certain embodiments of the present disclosure are directed to a nucleic acid molecule comprising a first ITR, a second ITR, and a genetic cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a clotting factor. In some embodiments, the clotting factor is selected from the group consisting of FI, FI I, Fill, FIV, FV, FVI, FVII, FVIII, FIX, FX, FXI, FXII, FXIII), VWF, prekallikrein, high-molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, Protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator(tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), any zymogen thereof, any active form thereof, and any combination thereof. In one embodiment, the clotting factor comprises FVIII or a variant or fragment thereof. In another embodiment, the clotting factor comprises FIX or a variant or fragment thereof. In another embodiment, the clotting factor comprises FVII or a variant or fragment thereof. In another embodiment, the clotting factor comprises VWF or a variant or fragment thereof.1. Clotting Factors

[0243] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a genetic cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein, wherein the therapeutic protein comprises a factor VIII polypeptide. "Factor VIII," abbreviated throughout the instant application as "FVIII," as used herein, means functional FVIII polypeptide in its normal role in coagulation, unless otherwise specified. Thus, the term FVIII includes variant polypeptides that are functional. "A FVIII protein" is used interchangeably with FVIII polypeptide (or protein) or FVIII. Examples of the FVIII functions include, but are not limited to, an ability to activate coagulation, an ability to act as a cofactor for factor IX, or an ability to form a tenase complex with factor IX in the presence of Ca2+and phospholipids, which then converts Factor X to the activated form Xa. The FVIII protein can be the human, porcine, canine, rat, or murine FVIII protein. In addition, comparisons between FVIII from humans and other species have identified conserved residues that are likely to be required for function (Cameron et a!., Thromb. Haemost. 79:317-22 (1998); US 6,251 ,632). The full length polypeptide and polynucleotide sequences are known, as are many functional fragments, mutants and modified versions. Various FVIII amino acid and nucleotide sequences are disclosed in, e.g. , US Publication Nos. 2015 / 0158929 A1 , 2014 / 0308280 A1 , and 2014 / 0370035 A1 and International Publication No. WO 2015 / 106052 A1. FVIII polypeptides include, e.g., full-length FVIII, full-length FVIII minus Met at the N-terminus, mature FVIII (minus the signal sequence), mature FVIII with an additional Met at the N-terminus, and / or FVIII with a full or partial deletion of the B domain. FVIII variants include B domain deletions, whether partial or full deletionsa. FVIII and Polynucleotide Sequences Encoding the FVIII Protein

[0244] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a genetic cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein, wherein the therapeutic protein comprises a factor VIII polypeptide. "Factor VIII," abbreviated throughout the instant application as "FVIII," as used herein, means functional FVIII polypeptide in its normal role in coagulation, unless otherwise specified. Thus, the term FVIII includes variant polypeptides that are functional. "A FVIII protein" is used interchangeably with FVIII polypeptide (or protein) or FVIII. Examples of the FVIII functions include, but are not limited to, an ability to activate coagulation, an ability to act as a cofactor for factor IX, or an ability to form a tenase complex with factor IX in the presence of Ca2+and phospholipids, which then converts Factor X to the activated form Xa. The FVIII protein can be the human, porcine, canine, rat, or murine FVIII protein. In addition, comparisons between FVIII from humans and other species have identified conserved residues that are likely to be required for function (Cameron et a!., Thromb. Haemost. 79:317-22 (1998); US 6,251 ,632). The full-length polypeptide and polynucleotide sequences are known, as are many functional fragments, mutants and modified versions. Various FVIII amino acid and nucleotide sequences are disclosed in, e.g. , US Publication Nos. 2015 / 0158929 A1 , 2014 / 0308280 A1 , and 2014 / 0370035 A1 and International Publication No. WO 2015 / 106052 A1. FVIII polypeptides include, e.g., full-length FVIII, full-length FVIII minus Met at the N-terminus, mature FVIII (minus the signal sequence), mature FVIII with an additional Met at the N-terminus, and / or FVIII with a full or partial deletion of the B domain. FVIII variants include B domain deletions, whether partial or full deletions.

[0245] The FVIII portion in the chimeric protein used herein has FVIII activity. FVIII activity can be measured by any known methods in the art. A number of tests are available to assess the function of the coagulation system: activated partial thromboplastin time (aPTT) test, chromogenic assay, ROTEM assay, prothrombin time (PT) test (also used to determine INR), fibrinogen testing (often by the Clauss method), platelet count, platelet function testing (often by PFA-100), TCT, bleeding time, mixing test (whether an abnormality corrects if the patient's plasma is mixed with normal plasma), coagulation factor assays, antiphospholipid antibodies, D- dimer, genetic tests (e.g., factor V Leiden, prothrombin mutation G20210A), dilute Russell's viper venom time (dRWT), miscellaneous platelet function tests, thromboelastography (TEG or Sonoclot), thromboelastometry (TEM®, e.g., ROTEM®), or euglobulin lysis time (ELT).

[0246] The aPTT test is a performance indicator measuring the efficacy of both the "intrinsic" (also referred to the contact activation pathway) and the common coagulation pathways. This test is commonly used to measure clotting activity of commercially available recombinant clotting factors, e.g. , FVIII. It is used in conjunction with prothrombin time (PT), which measures the extrinsic pathway.

[0247] ROTEM analysis provides information on the whole kinetics of haemostasis: clotting time, clot formation, clot stability and lysis. The different parameters in thromboelastometry are dependent on the activity of the plasmatic coagulation system, platelet function, fibrinolysis, or many factors which influence these interactions. This assay can provide a complete view of secondary haemostasis.

[0248] The chromogenic assay mechanism is based on the principles of the blood coagulation cascade, where activated FVIII accelerates the conversion of Factor X into Factor Xa in the presence of activated Factor IX, phospholipids and calcium ions. The Factor Xa activity is assessed by hydrolysis of a p-nitroanilide (pNA) substrate specific to Factor Xa. The initial rate of release of p-nitroaniline measured at 405 nM is directly proportional to the Factor Xa activity and thus to the FVIII activity in the sample.

[0249] The chromogenic assay is recommended by the FVIII and Factor IX Subcommittee of the Scientific and Standardization Committee (SSC) of the International Society on Thrombosis and Hemostatsis (ISTH). Since 1994, the chromogenic assay has also been the reference method of the European Pharmacopoeia for the assignment of FVIII concentrate potency. Thus, in one embodiment, the chimeric polypeptide comprising FVIII has FVIII activity comparable to a chimeric polypeptide comprising mature FVIII or a BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).

[0250] In another embodiment, the chimeric protein comprising FVIII of this disclosure has a Factor Xa generation rate comparable to a chimeric protein comprising mature FVIII or a BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).

[0251] In order to activate Factor X to Factor Xa, activated Factor IX (Factor IXa) hydrolyzes one arginine-isoleucine bond in Factor X to form Factor Xa in the presence of Ca2+, membrane phospholipids, and a FVIII cofactor. Therefore, the interaction of FVIII with Factor IX is critical in coagulation pathway. In certain embodiments, the chimeric polypeptide comprising FVIII can interact with Factor IXa at a rate comparable to a chimeric polypeptide comprising mature FVIII sequence or a BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).

[0252] In addition, FVIII is bound to von Willebrand Factor while inactive in circulation. FVIII degrades rapidly when not bound to VWF and is released from VWF by the action of thrombin. In some embodiments, the chimeric polypeptide comprising FVIII binds to von Willebrand Factor at a level comparable to a chimeric polypeptide comprising mature FVIII sequence or a BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).

[0253] FVIII can be inactivated by activated protein C in the presence of calcium and phospholipids. Activated protein C cleaves FVIII heavy chain after Arginine 336 in the A1 domain, which disrupts a Factor X substrate interaction site, and cleaves after Arginine 562 in the A2domain, which enhances the dissociation of the A2 domain as well as disrupts an interaction site with the Factor IXa. This cleavage also bisects the A2 domain (43 kDa) and generates A2-N (18 kDa) and A2-C (25 kDa) domains. Thus, activated protein C can catalyze multiple cleavage sites in the heavy chain. In one embodiment, the chimeric polypeptide comprising FVIII is inactivated by activated Protein C at a level comparable to a chimeric polypeptide comprising mature FVIII sequence or a BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).

[0254] In other embodiments, the chimeric protein comprising FVIII has FVIII activity in vivo comparable to a chimeric polypeptide comprising mature FVIII sequence or a BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®). In a particular embodiment, the chimeric polypeptide comprising FVIII is capable of protecting a HemA mouse at a level comparable to a chimeric polypeptide comprising mature FVIII sequence or a BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®) in a HemA mouse tail vein transection model.

[0255] A "B domain" of FVIII, as used herein, is the same as the B domain known in the art that is defined by internal amino acid sequence identity and sites of proteolytic cleavage by thrombin, e.g. , residues Ser741-Arg1648 of mature human FVIII. The other human FVIII domains are defined by the following amino acid residues, relative to mature human FVIII: A1 , residues Ala1-Arg372; A2, residues Ser373-Arg740; A3, residues Ser1690-lle2032; C1 , residues Arg2033-Asn2172; C2, residues Ser2173-Tyr2332 of mature FVIII. The sequence residue numbers used herein without referring to any SEQ ID Numbers correspond to the FVIII sequence without the signal peptide sequence (19 amino acids) unless otherwise indicated. The A3-C1-C2 sequence, also known as the FVIII heavy chain, includes residues Ser1690-Tyr2332. The remaining sequence, residues Glu1649-Arg1689, is usually referred to as the FVIII light chain activation peptide. The locations of the boundaries for all of the domains, including the B domains, for porcine, mouse and canine FVIII are also known in the art. In one embodiment, the B domain of FVIII is deleted ("B-domain-deleted FVIII" or "BDD FVIII"). An example of a BDD FVIII is REFACTO®(recombinant BDD FVIII). In one particular embodiment the B domain deleted FVIII variant comprises a deletion of amino acid residues 746 to 1648 of mature FVIII.

[0256] A "B-domain-deleted FVIII" may have the full or partial deletions disclosed in U.S. Pat. Nos. 6,316,226, 6,346,513, 7,041 ,635, 5,789,203, 6,060,447, 5,595,886, 6,228,620, 5,972,885, 6,048,720, 5,543,502, 5,610,278, 5,171 ,844, 5,1 12,950, 4,868,1 12, and 6,458,563 and Int'l Publ. No. WO 2015106052 A1 (PCT / US2015 / 010738). In some embodiments, a B-domain-deleted FVIII sequence used in the methods of the present disclosure comprises any one of the deletions disclosed at col. 4, line 4 to col. 5, line 28 and Examples 1-5 of U.S. Pat. No. 6,316,226 (also in US 6,346,513). In another embodiment, a B-domain deleted Factor VIII is the S743 / Q1638 B- domain deleted Factor VIII (SQ BDD FVIII) (e.g. , Factor VIII having a deletion from amino acid744 to amino acid 1637, e.g., Factor VIII having amino acids 1-743 and amino acids 1638-2332 of mature FVIII). In some embodiments, a B-domain-deleted FVIII used in the methods of the present disclosure has a deletion disclosed at col. 2, lines 26-51 and examples 5-8 of U.S. Patent No. 5,789,203 (also US 6,060,447, US 5,595,886, and US 6,228,620). In some embodiments, a B-domain-deleted Factor VIII has a deletion described in col. 1 , lines 25 to col. 2, line 40 of US Patent No. 5,972,885; col. 6, lines 1-22 and example 1 of U.S. Patent no. 6,048,720; col. 2, lines 17-46 of U.S. Patent No. 5,543,502; col. 4, line 22 to col. 5, line 36 of U.S. Patent no. 5,171 ,844; col. 2, lines 55-68, figure 2, and example 1 of U.S. Patent No. 5, 1 12,950; col. 2, line 2 to col. 19, line 21 and table 2 of U.S. Patent No. 4,868, 112; col. 2, line 1 to col. 3, line 19, col. 3, line 40 to col. 4, line 67, col. 7, line 43 to col. 8, line 26, and col. 1 1 , line 5 to col. 13, line 39 of U.S. Patent no. 7,041 ,635; or col. 4, lines 25-53, of U.S. Patent No. 6,458,563. In some embodiments, a B- domain-deleted FVIII has a deletion of most of the B domain, but still contains amino-terminal sequences of the B domain that are essential for in vivo proteolytic processing of the primary translation product into two polypeptide chain, as disclosed in WO 91 / 09122. In some embodiments, a B-domain-deleted FVIII is constructed with a deletion of amino acids 747-1638, i.e. , virtually a complete deletion of the B domain. Hoeben R.C., et al. J. Biol. Chem. 265 (13): 7318-7323 (1990). A B-domain-deleted Factor VIII may also contain a deletion of amino acids 771-1666 or amino acids 868-1562 of FVIII. Meulien P., et al. Protein Eng. 2(4): 301-6 (1988). Additional B domain deletions that are part of the invention include: deletion of amino acids 982 through 1562 or 760 through 1639 (Toole et al., Proc. Natl. Acad. Sci. U.S. A. (1986) 83, 5939- 5942)), 797 through 1562 (Eaton, et al. Biochemistry (1986) 25:8343-8347)), 741 through 1646 (Kaufman (PCT published application No. WO 87 / 04187)), 747-1560 (Sarver, et al., DNA (1987) 6:553-564)), 741 through 1648 (Pasek (PCT application No.88 / 00831)), or 816 through 1598 or 741 through 1648 (Lagner (Behring Inst. Mitt. (1988) No 82:16-25, EP 295597)). In one particular embodiment, the B-domain-deleted FVIII comprises a deletion of amino acid residues 746 to 1648 of mature FVIII. In another embodiment, the B-domain-deleted FVIII comprises a deletion of amino acid residues 745 to 1648 of mature FVIII. In some embodiments, the BDD FVIII comprises single chain FVIII that contains a deletion in amino acids 765 to 1652 corresponding to the mature full length FVIII (also known as rVIII-SingleChain and AFSTYLA®). See US Patent No. 7,041 ,635.

[0257] In other embodiments, BDD FVIII includes a FVIII polypeptide containing fragments of the B-domain that retain one or more N-linked glycosylation sites, e.g., residues 757, 784, 828, 900, 963, or optionally 943, which correspond to the amino acid sequence of the full-length FVIII sequence. Examples of the B-domain fragments include 226 amino acids or 163 amino acids of the B-domain as disclosed in Miao, H.Z., et al., Blood 103(a): 3412-3419 (2004), Kasuda, A, etal., J. Thromb. Haemost. 6: 1352-1359 (2008), and Pipe, S.W., et al., J. Thromb. Haemost. 9: 2235-2242 (2011) (i.e., the first 226 amino acids or 163 amino acids of the B domain are retained). In still other embodiments, BDD FVIII further comprises a point mutation at residue 309 (from Phe to Ser) to improve expression of the BDD FVIII protein. See Miao, H.Z., et al., Blood 103(a): 3412-3419 (2004). In still other embodiments, the BDD FVIII includes a FVIII polypeptide containing a portion of the B-domain, but not containing one or more furin cleavage sites (e.g., Arg1313 and Arg 1648). See Pipe, S.W., et al., J. Thromb. Haemost. 9: 2235-2242 (201 1). In some embodiments, the BDD FVIII comprises single chain FVIII that contains a deletion in amino acids 765 to 1652 corresponding to the mature full length FVIII (also known as rVIII-SingleChain and AFSTYLA®). See US Patent No. 7,041 ,635. Each of the foregoing deletions may be made in any FVIII sequence.

[0258] A great many functional FVIII variants are known, as is discussed above and below. In addition, hundreds of nonfunctional mutations in FVIII have been identified in hemophilia patients, and it has been determined that the effect of these mutations on FVIII function is due more to where they lie within the 3-dimensional structure of FVIII than on the nature of the substitution (Cutler et al., Hum. Mutat. 19\ 274-8 (2002)), incorporated herein by reference in its entirety. In addition, comparisons between FVIII from humans and other species have identified conserved residues that are likely to be required for function (Cameron et al. , Thromb. Haemost. 79:317-22 (1998); US 6,251 ,632), incorporated herein by reference in its entirety.

[0259] In some embodiments, the FVIII polypeptide comprises a FVIII variant or fragment thereof, wherein the FVIII variant or the fragment thereof has a FVIII activity. In some embodiments, the genetic cassette encodes a full-length FVIII polypeptide. In other embodiments, the genetic cassette encodes a B domain-deleted (BDD) FVIII polypeptide, wherein all or a portion of the B domain of FVIII is deleted. In one particular embodiment, the genetic cassette encodes a polypeptide comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NOs: 106, 107, 109, 1 10, 1 1 1 , or 1 12. In some embodiments, the genetic cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 17 or a fragment thereof. In some embodiments, the genetic cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 106 or a fragment thereof. In some embodiments, the genetic cassette comprises a nucleotide sequence which has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 107. In some embodiments, the genetic cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 109 or a fragment thereof.In some embodiments, the genetic cassette comprises a nucleotide sequence which has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 16. In some embodiments, the genetic cassette comprises a nucleotide sequence which has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 109.

[0260] In some embodiments, the genetic cassette of the disclosure encodes a FVIII polypeptide comprising a signal peptide or a fragment thereof. In other embodiments, the genetic cassette encodes a FVIII polypeptide which lacks a signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.

[0261] In some embodiments, the genetic cassette comprises a nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence is codon optimized. In certain embodiments, the genetic cassette comprises a nucleotide sequence which is disclosed in International Application No. PCT / US2017 / 015879, which is incorporated by reference in its entirety. In some embodiments, the genetic cassette comprises a nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence is codon optimized. In certain embodiments, the genetic cassette comprises a nucleotide sequence which has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a nucleotide sequence selected from SEQ ID NOs: 1-14. In some embodiments, the genetic cassette comprises a nucleotide sequence which has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 71. In some embodiments, the genetic cassette comprises a nucleotide sequence which has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19.i. Codon Optimized Nucleotide Sequences Encoding FVIII Polypeptides

[0262] In some embodiments, a nucleic acid molecule of the present disclosure comprises a first ITR, a second ITR, and a genetic cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein, wherein the first ITR and the second ITR are derived from an AAV genome, and wherein the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide. In some embodiments, the codon optimized nucleotide sequence encodes a full-length FVIII polypeptide. In other embodiments, the codon optimized nucleotide sequence encodes a B domain-deleted (BDD) FVIII polypeptide, wherein all or a portion of the B domain of FVIII is deleted. In one particular embodiment, the codon optimized nucleotide sequence encodes a polypeptide comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, atleast about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 17 or a fragment thereof. In one embodiment, the codon optimized nucleotide sequence encodes a polypeptide having the amino acid sequence of SEQ ID NO: 17 or a fragment thereof.

[0263] In some embodiments, the codon optimized nucleotide sequence encodes a FVIII polypeptide comprising a signal peptide or a fragment thereof. In other embodiments, the codon optimized sequence encodes a FVIII polypeptide which lacks a signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.

[0264] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3 or (ii) nucleotides 58-1791 of SEQ ID NO: 4; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In one particular embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 3. In another embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 4. In other embodiments, the first nucleotide sequence comprises nucleotides 58-1791 of SEQ ID NO: 3 or nucleotides 58-1791 of SEQ ID NO: 4.

[0265] In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1-1791 of SEQ ID NO: 3 or (ii) nucleotides 1-1791 of SEQ ID NO: 4; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1- 1791 of SEQ ID NO: 3 or nucleotides 1-1791 of SEQ ID NO: 4. In another embodiment, the second nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 3 or 1792-4374 of SEQ ID NO: 4. In one particular embodiment, the second nucleotide sequence comprises nucleotides 1792-4374 of SEQ ID NO: 3 or 1792-4374 of SEQ ID NO: 4. In still another embodiment, the second nucleotide sequence has at least 60%,at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 3 or 1792-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 1792-4374 of SEQ ID NO: 3 or 1792-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment). In one particular embodiment, the second nucleotide sequence comprises nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 3 or 1792-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 1792-4374 of SEQ ID NO: 3 or 1792-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment).

[0266] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5 or (ii) 1792-4374 of SEQ ID NO: 6; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 5. In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 6. In one particular embodiment, the second nucleic acid sequence comprises nucleotides 1792-4374 of SEQ ID NO: 5 or 1792-4374 of SEQ ID NO: 6. In some embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 5 or nucleotides 58-1791 of SEQ ID NO: 6. In other embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1-1791 of SEQ ID NO: 5 or nucleotides 1-1791 of SEQ ID NO: 6.

[0267] In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e.,nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment) or (ii) 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792- 2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment). In one particular embodiment, the second nucleic acid sequence comprises nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 or 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 or 1792-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment). In some embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 5 or nucleotides 58-1791 of SEQ ID NO: 6. In other embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1-1791 of SEQ ID NO: 5 or nucleotides 1- 1791 of SEQ ID NO: 6.

[0268] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 1 , (ii) nucleotides 58-1791 of SEQ ID NO: 2, (iii) nucleotides 58-1791 of SEQ ID NO: 70, or (iv) nucleotides 58-1791 of SEQ ID NO: 71 ; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In other embodiments, the first nucleotide sequence comprises nucleotides 58-1791 of SEQ ID NO: 1 , nucleotides 58-1791 of SEQ ID NO: 2, (iii) nucleotides 58-1791 of SEQ ID NO: 70, or (iv) nucleotides 58-1791 of SEQ ID NO: 71.

[0269] In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1-1791 of SEQ ID NO: 1 , (ii) nucleotides 1-1791 of SEQ ID NO: 2, (iii) nucleotides 1-1791 of SEQ ID NO: 70, or (iv) nucleotides 1-1791 of SEQ ID NO: 71 ; and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1-1791 of SEQ ID NO: 1 , nucleotides 1-1791 of SEQ ID NO: 2, (iii) nucleotides 1-1791 of SEQ ID NO: 70, or (iv) nucleotides 1-1791 of SEQ ID NO: 71. In another embodiment, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 1 , 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In one particular embodiment, the second nucleotide sequence linked to the first nucleotide sequence comprises (i) nucleotides 1792-4374 of SEQ ID NO: 1 , (ii) nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792- 4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In other embodiments, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1 , (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320- 4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71. In one embodiment, the second nucleotide sequence comprises (i) nucleotides 1792-2277 and 2320- 4374 of SEQ ID NO: 1 , (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71.

[0270] In another embodiment, the codon optimized nucleotide sequence encoding aFVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 1 , (ii) nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71 ; and wherein the N-terminal portion and the C-terminal portion together have a FVIIIpolypeptide activity. In one particular embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792-4374 of SEQ ID NO: 1 , (ii) nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In some embodiments, the codon optimized sequence encoding a FVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1 , (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792- 2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1792-4374 of SEQ ID NO: 1 , nucleotides 1792-4374 of SEQ ID NO: 2, nucleotides 1792-4374 of SEQ ID NO: 70, or nucleotides 1792- 4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity. In one embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792- 2277 and 2320-4374 of SEQ ID NO: 1 , (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1792-4374 of SEQ ID NO: 1 , nucleotides 1792-4374 of SEQ ID NO: 2, nucleotides 1792-4374 of SEQ ID NO: 70, or nucleotides 1792- 4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment).

[0271] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide with FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 1. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 1. In still other embodiments, the nucleotidesequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 1- 4374 of SEQ ID NO: 1 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 1.

[0272] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide with FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 2. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 2. In other embodiments, the nucleic acid sequence has at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 2. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 2 (i.e., nucleotides 58-4374 of SEQ ID NO: 2 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 2. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 2 (i.e., nucleotides 1- 4374 of SEQ ID NO: 2 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 2.

[0273] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide with FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 70. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 58-4374 of SEQ ID NO: 70 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 70. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 58-4374 of SEQ ID NO: 70 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 70. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 1-4374 of SEQID NO: 70 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 70.

[0274] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide with FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 71. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 71. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 71. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 71.

[0275] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide with FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 3. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 58-4374 of SEQ ID NO: 3 without the nucleotides encoding the B domain or B domain fragment). In certain embodiments, the nucleic acid sequence has at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 3. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 58-4374 of SEQ ID NO: 3 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 3. In still otherembodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 1-4374 of SEQ ID NO: 3 without the nucleotides encoding the B domain or B domain fragment)or nucleotides 1 to 4374 of SEQ ID NO: 3.

[0276] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide with FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 4. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 58- 4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 4. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 58-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 4. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 1-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 4.

[0277] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide with FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 5. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 58-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 5. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 58-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 5. In stillother embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 5.

[0278] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide with FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 6. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 58-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 6. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 58-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 6. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 6.

[0279] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleic acid sequence encoding a signal peptide. In certain embodiments, the signal peptide is a FVIII signal peptide. In some embodiments, the nucleic acid sequence encoding a signal peptide is codon optimized. In one particular embodiment, the nucleic acid sequence encoding a signal peptide has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to (i) nucleotides 1 to 57 of SEQ ID NO: 1 ; (ii) nucleotides 1 to 57 of SEQ ID NO: 2; (iii) nucleotides 1 to 57 of SEQ ID NO: 3; (iv) nucleotides 1 to 57 of SEQ ID NO: 4; (v) nucleotides 1 to 57 of SEQ ID NO: 5; (vi) nucleotides 1 to 57 of SEQ ID NO: 6; (vii) nucleotides 1 to 57 of SEQ ID NO: 70; (viii) nucleotides 1 to 57 of SEQ ID NO: 71 ; or (ix) nucleotides 1 to 57 of SEQ ID NO: 68.

[0280] SEQ ID NOs: 1-6, 70, and 71 are optimized versions of SEQ ID NO: 16, the starting or "parental" or "wild-type" FVIII nucleotide sequence. SEQ ID NO: 16 encodes a B domain-deleted human FVIII. While SEQ ID NOs: 1-6, 70, and 71 are derived from a specific Bdomain-deleted form of FVIII (SEQ ID NO: 16), it is to be understood that the present disclosure also includes optimized versions of nucleic acids encoding other versions of FVIII. For example, other versions of FVIII can include full length FVIII, other B-domain deletions of FVIII (described herein), or other fragments of FVIII that retain FVIII activity.

[0281] In one embodiment, the genetic cassette comprises a FVIII construct, which includes a polynucleotide sequence as listed in Tables 2A-2F. In one embodiment, the genetic cassette comprises a FVIII construct, which includes a polynucleotide sequence set forth in Table 2A. In one embodiment, the genetic cassette comprises a FVIII construct, which includes a polynucleotide sequence set forth in Table 2B. In one embodiment, the genetic cassette comprises a FVIII construct, which includes a polynucleotide sequence set forth in Table 2C. In one embodiment, the genetic cassette comprises a FVIII construct, which includes a polynucleotide sequence set forth in Table 2D. In one embodiment, the genetic cassette comprises a FVIII construct, which includes a polynucleotide sequence set forth in Table 2E. In one embodiment, the genetic cassette comprises a FVIII construct, which includes a polynucleotide sequence set forth in Table 2F

[0282] In certain embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 179, 182, 189, or 194. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 179. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 182. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 189. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%,at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 194. In some embodiments, the isolated nucleic acid molecule retains the ability to express a functional FVIII protein.Table 2A: Example AAV-FVIII construct (nucleotides 1-6526; SEQ ID NO: 1 10)Table 2B: Example B19-FVIII construct bearing B19d135 ITRs (nucleotides 1-6762; SEQ ID NO: 179)Table 2C: Example GPV-FVIII construct bearing GPVd162 ITRs (nucleotides 1-6830; SEQ ID NO: 182)Table 2D: Example B19-FVIII construct bearing full length B19 ITRs (nucleotides 1-7032; SEQ ID NO: 189)Table 2E: Example AAV-FVIII construct (nucleotides 1-6824; SEQ ID NO: 190)Table 2F: Example GPV-FVIII construct bearing full length GPV ITRs (nucleotides 1-7154; SEQ ID NO: 194)

[0283] In one embodiment, the genetic cassette comprises a phenylalanine hydroxylase(PAH) construct, which includes a polynucleotide sequence as listed in Tables 10A and 10B. In one embodiment, the genetic cassette comprises a PAH construct, which includes a polynucleotide sequence set forth in Table 10A. In one embodiment, the genetic cassette comprises a PAH construct, which includes a polynucleotide sequence set forth in Table 10B.

[0284] In certain embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 197 or 198. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 197. In some embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 198. In some embodiments, the isolated nucleic acid molecule retains the ability to express a functional phenylalanine hydroxylase.A. Codon Adaptation Index

[0285] In one embodiment, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the human codon adaptation index of the codon optimized nucleotide sequence is increased relative to SEQ ID NO: 16. For example, the codon optimized nucleotide sequence can have a human codon adaptation index that is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81 %), at least about 0.82(82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%), at least about 0.91 (91 %), at least about 0.92 (92%), at least about 0.93 (93%), at least about 0.94 (94%), at least about 0.95 (95%), at least about 0.96 (96%), at least about 0.97 (97%), at least about 0.98 (98%), or at least about 0.99 (99%). In some embodiments, the codon optimized nucleotide sequence has a human codon adaptation index that is at least about .88 (88%). In other embodiments, the codon optimized nucleotide sequence has a human codon adaptation index that is at least about .91 (91 %). In other embodiments, the codon optimized nucleotide sequence has a human codon adaptation index that is at least about .91 (97%).

[0286] In one particular embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO: 16. In some embodiments, the nucleotide sequence has a human codon adaptation index that is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81 %), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%), or at least about .91 (91 %). In one particular embodiment, the nucleotide sequence has a human codon adaptation index that is at least about .88 (88%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about .91 (91 %).

[0287] In another embodiment, the codon optimized nucleotide sequence encoding aFVIII polypeptide comprises a nucleotide sequence which comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-2277 and 2320- 4374 of SEQ ID NO: 5 or (ii) 1792-2277 and 2320-4374 of SEQ ID NO: 6; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO: 16. In some embodiments, the nucleotide sequence has a human codon adaptation index that is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81 %), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), or at least about 0.88 (88%). In one particular embodiment, the nucleotide sequence has a human codon adaptation index that is at least about .83 (83%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about .88 (88%).

[0288] In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide with FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to nucleotides 58-2277 and 2320- 4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO: 16. In some embodiments, the nucleotide sequence has a human codon adaptation index that is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81 %), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), or at least about 0.88 (88%). In one particular embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.75 (75%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.83 (83%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.88 (88%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.91 (91 %). In anotherembodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.97 (97%).

[0289] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide of the present disclosure has an increased frequency of optimal codons (FOP) relative to SEQ ID NO: 16. In certain embodiments, the FOP of the codon optimized nucleotide sequence encoding a FVIII polypeptide is at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 64, at least about 65, at least about 70, at least about 75, at least about 79, at least about 80, at least about 85, or at least about 90.

[0290] In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide of the present disclosure has an increased relative synonymous codon usage (RCSU) relative to SEQ ID NO: 16. In some embodiments, the RCSU of the isolated nucleic acid molecule is greater than 1.5. In other embodiments, the RCSU of the isolated nucleic acid molecule is greater than 2.0. In certain embodiments, the RCSU of the isolated nucleic acid molecule is at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2.0, at least about 2.1 , at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, or at least about 2.7.

[0291] In still other embodiments, the codon optimized nucleotide sequence encoding aFVIII polypeptide of the present disclosure has a decreased effective number of codons relative to SEQ ID NO: 16. In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide has an effective number of codons of less than about 50, less than about 45, less than about 40, less than about 35, less than about 30, or less than about 25. In one particular embodiment, the isolated nucleic acid molecule has an effective number of codons of about 40, about 35, about 30, about 25, or about 20.B. G / C Content Optimization

[0292] In some embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51 %, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, or at least about 60%.

[0293] In one particular embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIIIpolypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises has a G / C content that is at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51 %, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, or at least about 58%. In one particular embodiment, the nucleotide sequence that encodes a polypeptide with FVIII activity has a G / C content that is at least about 58%.

[0294] In another embodiment, the codon optimized nucleotide sequence encoding aFVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment), or (iv) 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51 %, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, or at least about 57%. In one particular embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptidehas a G / C content that is at least about 52%. In another embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 55%. In another embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 57%.

[0295] In other embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 or (ii) nucleotides 58- 2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 45%. In one particular embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 52%. In another embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 55%. In another embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 57%. In another embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 58%. In still another embodiment, the n codon optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 60%.

[0296] "G / C content" (or guanine-cytosine content), or "percentage of G / C nucleotides," refers to the percentage of nitrogenous bases in a DNA molecule that are either guanine or cytosine. G / C content can be calculated using the following formula:

[0297] Human genes are highly heterogeneous in their G / C content, with some genes having a G / C content as low as 20%, and other genes having a G / C content as high as 95%. In general, G / C rich genes are more highly expressed. In fact, it has been demonstrated that increasing the G / C content of a gene can lead to increased expression of the gene, due mostly to an increase in transcription and higher steady state mRNA levels. See Kudla et a!., PLoS Biol., 4(6): e180 (2006).C. Matrix Attachment Region-Like Sequences

[0298] In some embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a MARS / ARS sequence.

[0299] In one particular embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the nucleotide sequence that encodes a polypeptide with FVIII activity contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a MARS / ARS sequence.

[0300] In another embodiment, the codon optimized nucleotide sequence encoding aFVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792- 4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a MARS / ARS sequence.

[0301] In other embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, or 71 or (ii) nucleotides 58-2277 and 2320-4374 of SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, or 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a MARS / ARS sequence.

[0302] AT-rich elements in the human FVIII nucleotide sequence that share sequence similarity with Saccharomyces cerevisiae autonomously replicating sequences (ARSs) and nuclear-matrix attachment regions (MARs) have been identified. (Fallux et a!., Mol. Cell. Biol. 16:4264-4272 (1996). One of these elements has been demonstrated to bind nuclear factors in vitro and to repress the expression of a chloramphenicol acetyltransferase (CAT) reporter gene. Id. It has been hypothesized that these sequences can contribute to the transcriptional repression of the human FVIII gene. Thus, in one embodiment, all MAR / ARS sequences are abolished in the codon optimized nucleotide sequence encoding a FVIII polypeptide of the present disclosure. There are four MAR / ARS ATATTT sequences (SEQ ID NO: 21) and three MAR / ARS AAATATsequences (SEQ ID NO: 22) in the parental FVIII sequence (SEQ ID NO: 16). All of these sites were mutated to destroy the MAR / ARS sequences in the optimized FVIII sequences (SEQ ID NOs: 1-6). The location of each of these elements, and the sequence of the corresponding nucleotides in the optimized sequences are shown in Table 3, below.Table 3: Summary of Changes to Repressive ElementsD. Destabilizing Sequences

[0303] In some embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 destabilizing elements. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 destabilizing elements. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a destabilizing element.

[0304] In one particular embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4;wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the nucleotide sequence that encodes a polypeptide with FVIII activity contains at most 9, at most 8, at most 7, at most 6, or at most 5 destabilizing elements. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 destabilizing elements. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a destabilizing element.

[0305] In another embodiment, the codon optimized nucleotide sequence encoding aFVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792- 4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the nucleotide sequence that encodes a polypeptide with FVIII activity contains at most 9, at most 8, at most 7, at most 6, or at most 5 destabilizing elements. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 destabilizing elements. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a destabilizing element.

[0306] In other embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 destabilizing elements. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 destabilizing elements. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a destabilizing element.

[0307] There are ten destabilizing elements in the parental FVIII sequence (SEQ ID NO:16): six ATTTA sequences (SEQ ID NO: 23) and four TAAAT sequences (SEQ ID NO: 24). In one embodiment, sequences of these sites were mutated to destroy the destabilizing elements in optimized FVIII SEQ ID NOs: 1-6, 70, and 71. The location of each of these elements, and the sequence of the corresponding nucleotides in the optimized sequences are shown in Table 3.E. Potential Promoter Binding Sites

[0308] In some embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding sites. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a potential promoter binding site.

[0309] In one particular embodiment, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptideactivity; and wherein the codon optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding sites. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a potential promoter binding site.

[0310] In another embodiment, the codon optimized nucleotide sequence encoding aFVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792- 4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding sites. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a potential promoter binding site.

[0311] In other embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 ofSEQ ID NO: 1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding sites. In yet other embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide does not contain a potential promoter binding site.

[0312] TATA boxes are regulatory sequences often found in the promoter regions of eukaryotes. They serve as the binding site of TATA binding protein (TBP), a general transcription factor. TATA boxes usually comprise the sequence TATAA (SEQ ID NO: 28) or a close variant. TATA boxes within a coding sequence, however, can inhibit the translation of full-length protein. There are ten potential promoter binding sequences in the wild type BDD FVIII sequence (SEQ ID NO: 16): five TATAA sequences (SEQ ID NO: 28) and five TTATA sequences (SEQ ID NO: 29). In some embodiments, at least 1 , at least 2, at least 3, or at least 4 of the promoter binding sites are abolished in the FVIII genes of the present disclosure. In some embodiments, at least 5 of the promoter binding sites are abolished in the FVIII genes of the present disclosure. In other embodiments, at least 6, at least 7, or at least 8 of the promoter binding sites are abolished in the FVIII genes of the present disclosure. In one embodiment, at least 9 of the promoter binging sites are abolished in the FVIII genes of the present disclosure. In one particular embodiment, all promoter binding sites are abolished in the FVIII genes of the present disclosure. The location of each potential promoter binding site and the sequence of the corresponding nucleotides in the optimized sequences are shown in Table 3.F. Other Cis Acting Negative Regulatory Elements

[0313] In addition to the MAR / ARS sequences, destabilizing elements, and potential promoter sites described above, several additional potentially inhibitory sequences can be identified in the wild type BDD FVIII sequence (SEQ ID NO: 16). Two AU rich sequence elements (AREs) can be identified (ATTTTATT (SEQ ID NOs: 30); and ATTTTTAA (SEQ ID NO: 31), along with a poly-A site (AAAAAAA; SEQ ID NO: 26), a poly-T site (TTTTTT ; SEQ ID NO: 25), and a splice site (GGTGAT; SEQ ID NO: 27) in the non-optimized BDD FVIII sequence. One or more of these elements can be removed from the optimized FVIII sequences. The location of each of these sites and the sequence of the corresponding nucleotides in the optimized sequences are shown in Table 3.

[0314] In certain embodiments, the codon optimized nucleotide sequence encoding aFVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence does not contain one or more exacting negative regulatory elements, for example, a splice site, a poly-T sequence, a poly-A sequence, an ARE sequence, or any combinations thereof.

[0315] In another embodiment, the codon optimized nucleotide sequence encoding aFVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792- 4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence does not contain one or more ex acting negative regulatory elements, for example, a splice site, a poly-T sequence, a poly-A sequence, an ARE sequence, or any combinations thereof.

[0316] In other embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, for example, a splice site, a poly-T sequence, a poly-A sequence, an ARE sequence, or any combinations thereof.

[0317] In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence does not contain the splice site GGTGAT (SEQ ID NO: 27). In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence does not contain a poly-T sequence (SEQ ID NO: 25). In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N- terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at leastabout 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence does not contain a poly-A sequence (SEQ ID NO: 26). In some embodiments, the codon optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N- terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have a FVIII polypeptide activity; and wherein the codon optimized nucleotide sequence does not contain an ARE element (SEQ ID NO: 30 or SEQ ID NO: 31).

[0318] In some embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon optimized nucleotide sequence does not contain the splice site GGTGAT (SEQ ID NO: 27). In some embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO:1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon optimized nucleotide sequence does not contain a poly-T sequence (SEQ ID NO: 25). In some embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58- 4374 of SEQ ID NO: 1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon optimized nucleotide sequence does not contain a poly-A sequence (SEQ ID NO: 26). In some embodiments, the genetic cassette comprises a codon optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1 , 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 , 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon optimized nucleotide sequence does not contain an ARE element (SEQ ID NO: 30 or SEQ ID NO: 31).

[0319] In other embodiments, an optimized FVIII sequence of the disclosure does not comprise one or more of antiviral motifs, stem-loop structures, and repeat sequences.

[0320] In still other embodiments, the nucleotides surrounding the transcription start site are changed to a kozak consensus sequence (GCCGCCACCATGC (SEQ ID NO: 32), wherein the underlined nucleotides are the start codon). In other embodiments, restriction sites can be added or removed to facilitate the cloning process.b. FIX and Polynucleotide Sequences Encoding the FIX Protein

[0321] In some embodiments, the nucleic acid molecule comprises a first ITR, a secondITR, and a genetic cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein, wherein the therapeutic protein comprises a FIX polypeptide. In some embodiments, the FIX polypeptide comprises FIX or a variant or fragment thereof, wherein the FIX or the variant or fragment thereof has a FIX activity.

[0322] Human FIX is a serine protease that is an important component of the intrinsic pathway of the blood coagulation cascade. "Factor IX" or "FIX," as used herein, refers to a coagulation factor protein and species and sequence variants thereof, and includes, but is not limited to, the 461 single-chain amino acid sequence of human FIX precursor polypeptide ("prepro"), the 415 single-chain amino acid sequence of mature human FIX (SEQ ID NO: 125), and the R338L FIX (Padua) variant (SEQ ID NO: 126). FIX includes any form of FIX molecule with the typical characteristics of blood coagulation FIX. As used herein "Factor IX" and "FIX" are intended to encompass polypeptides that comprise the domains Gla (region containing y- carboxyglutamic acid residues), EGF1 and EGF2 (regions containing sequences homologous to human epidermal growth factor), activation peptide ("AP," formed by residues R136-R180 of the mature FIX), and the C-terminal protease domain ("Pro"), or synonyms of these domains known in the art, or can be a truncated fragment or a sequence variant that retains at least a portion of the biological activity of the native protein. FIX or sequence variants have been cloned, as described in U.S. Patent Nos. 4,770,999 and 7,700,734, and cDNA coding for human FIX has been isolated, characterized, and cloned into expression vectors (see, for example, Choo et al., Nature 299:178-180 (1982); Fair et al., Blood 64:194-204 (1984); and Kurachi et al., Proc. Natl. Acad. Sci., U.S.A. 79:6461-6464 (1982)). One particular variant of FIX, the R338L FIX (Padua) variant (SEQ ID NO: 2), characterized by Simioni et al, 2009, comprises a gain-of-function mutation, which correlates with a nearly 8-fold increase in the activity of the Padua variant relative to native FIX (Table 4). FIX variants can also include any FIX polypeptide having one or more conservative amino acid substitutions, which do not affect the FIX activity of the FIX polypeptide. In some embodiments, the FIX variant comprises rFIX-albumin fused by a cleavable linker, e.g., IDELVION®. See US 7,939,632, incorporated herein by reference in its entirety.Table 4: Example FIX Sequences* Grey shading = signal peptide; underline = XTEN sequence; bold = Fc.** SEQ ID NO: 67 of US Patent No. 9,856,468, which is incorporated by reference herein in its entirety.

[0323] The FIX polypeptide is 55 kDa, synthesized as a prepropolypetide chain (SEQ IDNO: 125) composed of three regions: a signal peptide of 28 amino acids (amino acids 1 to 28 of SEQ ID NO: 127), a propeptide of 18 amino acids (amino acids 29 to 46), which is required for gamma-carboxylation of glutamic acid residues, and a mature Factor IX of 415 amino acids (SEQ ID NO: 125 or 126). The propeptide is an 18-amino acid residue sequence N-terminal to the gamma-carboxyglutamate domain. The propeptide binds vitamin K-dependent gamma carboxylase and then is cleaved from the precursor polypeptide of FIX by an endogenous protease, most likely PACE (paired basic amino acid cleaving enzyme), also known as furin or PCSK3. Without the gamma carboxylation, the Gla domain is unable to bind calcium to assume the correct conformation necessary to anchor the protein to negatively charged phospholipid surfaces, thereby rendering Factor IX nonfunctional. Even if it is carboxylated, the Gla domain also depends on cleavage of the propeptide for proper function, since retained propeptide interferes with conformational changes of the Gla domain necessary for optimal binding to calcium and phospholipid. In humans, the resulting mature Factor IX is secreted by liver cells into the blood stream as an inactive zymogen, a single chain protein of 415 amino acid residues that contains approximately 17% carbohydrate by weight (Schmidt, A. E., et al. (2003) Trends Cardiovasc Med, 13: 39).

[0324] The mature FIX is composed of several domains that in an N- to C-terminus configuration are: a GLA domain, an EGF1 domain, an EGF2 domain, an activation peptide (AP) domain, and a protease (or catalytic) domain. A short linker connects the EGF2 domain with the AP domain. FIX contains two activation peptides formed by R145-A146 and R180-V181 , respectively. Following activation, the single-chain FIX becomes a 2-chain molecule, in which the two chains are linked by a disulfide bond. Clotting factors can be engineered by replacing their activation peptides resulting in altered activation specificity. In mammals, mature FIX must be activated by activated Factor XI to yield Factor IXa. The protease domain provides, upon activation of FIX to FIXa, the catalytic activity of FIX. Activated Factor VIII (FVIIIa) is the specific cofactor for the full expression of FIXa activity.

[0325] In certain embodiments, a FIX polypeptide comprises an Thr148 allelic form of plasma derived FIX and has structural and functional characteristics similar to endogenous FIX.

[0326] Many functional FIX variants are known in the art. International publication numberWO 02 / 040544 A3 discloses mutants that exhibit increased resistance to inhibition by heparin at page 4, lines 9-30 and page 15, lines 6-31. International publication number WO 03 / 020764 A2 discloses FIX mutants with reduced T cell immunogenicity in Tables 2 and 3 (on pages 14-24), and at page 12, lines 1-27. International publication number WO 2007 / 149406 A2 discloses functional mutant FIX molecules that exhibit increased protein stability, increased in vivo and invitro half-life, and increased resistance to proteases at page 4, line 1 to page 19, line 1 1. WO 2007 / 149406 A2 also discloses chimeric and other variant FIX molecules at page 19, line 12 to page 20, line 9. International publication number WO 08 / 1 18507 A2 discloses FIX mutants that exhibit increased clotting activity at page 5, line 14 to page 6, line 5. International publication number WO 09 / 051717 A2 discloses FIX mutants having an increased number of N-linked and / or O-linked glycosylation sites, which results in an increased half-life and / or recovery at page 9, line 1 1 to page 20, line 2. International publication number WO 09 / 137254 A2 also discloses Factor IX mutants with increased numbers of glycosylation sites at page 2, paragraph

[0006] to page 5, paragraph

[0011] and page 16, paragraph

[0044] to page 24, paragraph

[0057] International publication number WO 09 / 130198 A2 discloses functional mutant FIX molecules that have an increased number of glycosylation sites, which result in an increased half-life, at page 4, line 26 to page 12, line 6. International publication number WO 09 / 140015 A2 discloses functional FIX mutants that an increased number of Cys residues, which can be used for polymer (e.g., PEG) conjugation, at page 1 1 , paragraph

[0043] to page 13, paragraph

[0053] The FIX polypeptides described in International Application No. PCT / US201 1 / 043569 filed July 1 1 , 201 1 and published as WO 2012 / 006624 on January 12, 2012 are also incorporated herein by reference in its entirety. In some embodiments, the FIX polypeptide comprises a FIX polypeptide fused to an albumin, e.g., FIX-albumin. In certain embodiments, the FIX polypeptide is IDELVION® or rlX-FP.

[0327] In addition, hundreds of non-functional mutations in FIX have been identified in hemophilia subjects, many of which are disclosed in Table 6, at pages 11-14 of International publication number WO 09 / 137254 A2. Such non-functional mutations are not included in the invention, but provide additional guidance for which mutations are more or less likely to result in a functional FIX polypeptide.

[0328] In one embodiment, the FIX polypeptide (or Factor IX portion of a fusion polypeptide) comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 1 or 2 (amino acids 1 to 415 of SEQ ID NO: 125 or 126), or alternatively, with a propeptide sequence, or with a propeptide and signal sequence (full length FIX). In another embodiment, the FIX polypeptide comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 2.

[0329] FIX coagulant activity is expressed as International Unit(s) (IU). One IU of FIX activity corresponds approximately to the quantity of FIX in one milliliter of normal human plasma. Several assays are available for measuring FIX activity, including the one stage clotting assay (activated partial thromboplastin time; aPTT), thrombin generation time (TGA) and rotationalthromboelastometry (ROTEM®). The invention contemplates sequences that have homology to FIX sequences, sequence fragments that are natural, such as from humans, non-human primates, mammals (including domestic animals), and non-natural sequence variants which retain at least a portion of the biologic activity or biological function of FIX and / or that are useful for preventing, treating, mediating, or ameliorating a coagulation factor-related disease, deficiency, disorder or condition (e.g., bleeding episodes related to trauma, surgery, of deficiency of a coagulation factor). Sequences with homology to human FIX can be found by standard homology searching techniques, such as NCBI BLAST.

[0330] In certain embodiments, the FIX sequence is codon-optimized. Examples of codon-optimized FIX sequences include, but are not limited to, SEQ ID NOs: 1 and 54-58 of International Publication No. WO 2016 / 0041 13 A1 , which is incorporated by reference herein in its entirety.c. FVII and Polynucleotide Sequences Encoding the FVII Protein

[0331] In some embodiments, the nucleic acid molecule comprises a first ITR, a secondITR, and a genetic cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein, wherein the therapeutic protein comprises a Factor VII polypeptide. In some embodiments, the FVII polypeptide comprises FVII or a variant or fragment thereof, wherein the variant or fragment thereof has a FVII activity.

[0332] "Factor VII" ("FVII," or "F7;" also referred to as Factor 7, coagulation factor VII, serum factor VII, serum prothrombin conversion accelerator, SPCA, proconvertin and eptacog alpha) is a serine protease that is part of the coagulation cascade. In one embodiment, the clotting factor in the nucleic acid described herein is FVII. Recombinant activated Factor VII ("FVII") has become widely used for the treatment of major bleeding, such as that which occurs in patients having hemophilia A or B, deficiency of coagulation Factor XI, FVII, defective platelet function, thrombocytopenia, or von Willebrand's disease.

[0333] Recombinant activated FVII (rFVIIa; NOVOSEVEN®) is used to treat bleeding episodes in (i) hemophilia patients with neutralizing antibodies against FVIII or FIX (inhibitors), (ii) patients with FVII deficiency, or (iii) patients with hemophilia A or B with inhibitors undergoing surgical procedures. However, NOVOSEVEN®displays poor efficacy. Repeated doses of FVIIa at high concentration are often required to control a bleed, due to its low affinity for activated platelets, short half-life, and poor enzymatic activity in the absence of tissue factor. Accordingly, there is an unmet medical need for better treatment and prevention options for hemophilia patients with FVIII and FIX inhibitors and / or with FVII deficiency.

[0334] In one embodiment, the genetic cassette encodes a mature form of FVII or a variant thereof. FVII includes a Gla domain, two EGF domains (EGF-1 and EGF-2), and a serineprotease domain (or peptidase S1 domain) that is highly conserved among all members of the peptidase S1 family of serine proteases, such as for example with chymotrypsin. FVII occurs as a single chain zymogen (i.e., activatable FVII) and a fully activated two-chain form.C. Growth Factors

[0335] In some embodiments, the nucleic acid molecule comprises a first ITR, a secondITR, and a genetic cassette encoding a target sequence, wherein the target sequence encodes a therapeutic protein, and wherein the therapeutic protein comprises a growth factor. The growth factor can be selected from any growth factor known in the art. In some embodiments, the growth factor is a hormone. In other embodiments, the growth factor is a cytokine. In some embodiments, the growth factor is a chemokine.

[0336] In some embodiments, the growth factor is adrenomedullin (AM). In some embodiments, the growth factor is angiopoietin (Ang). In some embodiments, the growth factor is autocrine motility factor. In some embodiments, the growth factor is a Bone morphogenetic protein (BMP). In some embodiments, the BMP is selects from BMP2, BMP4, BMP5, and BMP7. In some embodiments, the growth factor is a ciliary neurotrophic factor family member. In some embodiments, the ciliary neurotrophic factor family member is selected from ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6). In some embodiments, the growth factor is a colony-stimulating factor. In some embodiments, the colony-stimulating factor is selected from macrophage colony-stimulating factor (m-CSF), granulocyte colony-stimulating factor (G-CSF), and granulocyte macrophage colony-stimulating factor (GM-CSF). In some embodiments, the growth factor is an epidermal growth factor (EGF). In some embodiments, the growth factor is an ephrin. In some embodiments, the ephrin is selected from ephrin A1 , ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1 , ephrin B2, and ephrin B3. In some embodiments, the growth factor is erythropoietin (EPO). In some embodiments, the growth factor is a fibroblast growth factor (FGF). In some embodiments, the FGF is selected from FGF1 , FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF1 1 , FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21 , FGF22, and FGF23. In some embodiments, the growth factor is foetal bovine somatotrophin (FBS). In some embodiments, the growth factor is a GDNF family member. In some embodiments, the GDNF family member is selected from glial cell line- derived neurotrophic factor (GDNF), neurturin, persephin, and artemin. In some embodiments, the growth factor is growth differentiation factor-9 (GDF9). In some embodiments, the growth factor is hepatocyte growth factor (HGF). In some embodiments, the growth factor is hepatoma- derived growth factor (HDGF). In some embodiments, the growth factor is insulin. In some embodiments, the growth factor is an insulin-like growth factor. In some embodiments, the insulinlike growth factor is insulin-like growth factor-1 (IGF-1) or IGF-2. In some embodiments, thegrowth factor is an interleukin (IL). In some embodiments, the IL is selected from IL-1 , IL-2, IL-3, IL-4, IL-5, IL-6, and IL-7. In some embodiments, the growth factor is keratinocyte growth factor (KGF). In some embodiments, the growth factor is migration-stimulating factor (MSF). In some embodiments, the growth factor is macrophage-stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)),. In some embodiments, the growth factor is myostatin (GDF-8). In some embodiments, the growth factor is a ...

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

2. The nucleic acid molecule of claim 1 , wherein the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 180 and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 181.

3. The nucleic acid molecule of claim 1 , wherein the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 183 and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 184.

4. The nucleic acid molecule of claim 1 , wherein the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 185 and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 186.

5. The nucleic acid molecule of claim 1 , wherein the first ITR comprises the nucleotide sequence set forth in SEQ ID NO: 187 and the second ITR comprises the nucleotide sequence set forth in SEQ ID NO: 188.

6. The nucleic acid molecule of claim 1 , wherein the first ITR and / or the second ITR consists of a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188.

7. The nucleic acid molecule of claim 1 , wherein the first ITR and the second ITR are reverse complements of each other.

8. The nucleic acid molecule of any one of claims 1 to 7, further comprising a promoter.

9. The nucleic acid molecule of claim 8, wherein the promoter is a tissue-specific promoter.

10. The nucleic acid molecule of claim 8 or 9, wherein the promoter drives expression of the heterologous polynucleotide sequence in an organ selected from the muscle, central nervous system (CNS), ocular, liver, heart, kidney, pancreas, lungs, skin, bladder, urinary tract, or any combination thereof.1 1. The nucleic acid molecule of any one of claims 8 to 10, wherein the promoter drives expression of the heterologous polynucleotide sequence in hepatocytes, endothelial cells, cardiac muscle cells, skeletal muscle cells, sinusoidal cells, afferent neurons, efferent neurons, interneurons, glial cells, astrocytes, oligodendrocytes, microglia, ependymal cells, lung epithelial cells, Schwann cells, satellite cells, photoreceptor cells, retinal ganglion cells, or any combination thereof.

12. The nucleic acid molecule of any one of claims 8 to 1 1 , wherein the promoter is positioned 5' to the heterologous polynucleotide sequence.

13. The nucleic acid molecule of any one of claims 8 to 12, wherein the promoter is selected from the group consisting of a mouse thyretin promoter (mTTR), an endogenous human factor VIII promoter (F8), a human alpha-1-antitrypsin promoter (hAAT), a human albumin minimal promoter, a mouse albumin promoter, a tristetraprolin (TTP) promoter, a CASI promoter, a CAG promoter, a cytomegalovirus (CMV) promoter, a1 -antitrypsin (AAT), muscle creatine kinase (MCK), myosin heavy chain alpha (aMHC), myoglobin (MB), desmin (DES), SPc5-12, 2R5Sc5- 12, dMCK, tMCK, and a phosphoglycerate kinase (PGK) promoter.

14. The nucleic acid molecule of any one of claims 1 to 13, wherein the heterologous polynucleotide sequence further comprises an intronic sequence.

15. The nucleic acid molecule of claim 14, wherein the intronic sequence is positioned 5' to the heterologous polynucleotide sequence.

16. The nucleic acid molecule of claim 14 or 15, wherein the intronic sequence is positioned 3' to the promoter.

17. The nucleic acid molecule of any one of claims 14 to 16, wherein the intronic sequence comprises a synthetic intronic sequence.

18. The nucleic acid molecule of any one of claims 14 to 17, wherein the intronic sequence comprises SEQ ID NO: 1 15 or 192.

19. The nucleic acid molecule of any one of claims 1 to 18, wherein the genetic cassette further comprises a post-transcriptional regulatory element.

20. The nucleic acid molecule of claim 19, wherein the post-transcriptional regulatory element is positioned 3' to the heterologous polynucleotide sequence.

21. The nucleic acid molecule of claim 19 or 20, wherein the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, or any combination thereof.

22. The nucleic acid molecule of claim 21 , wherein the microRNA binding site comprises a binding site to miR142-3p.

23. The nucleic acid molecule of any one of claims 1 to 22, wherein the genetic cassette further comprises a 3'UTR poly(A) tail sequence.

24. The nucleic acid molecule of claim 23, wherein the 3'UTR poly(A) tail sequence is selected from the group consisting of bGH poly(A), actin poly(A), hemoglobin poly(A), and any combination thereof.

25. The nucleic acid molecule of claim 23 or 24, wherein the 3'UTR poly(A) tail sequence comprises bGH poly(A).

26. The nucleic acid molecule of any one of claims 1 to 25, wherein the genetic cassette further comprises an enhancer sequence.

27. The nucleic acid molecule of claim 26, wherein the enhancer sequence is positioned between the first ITR and the second ITR.

28. The nucleic acid molecule of any one of claims 1 to 27, wherein the nucleic acid molecule comprises from 5’ to 3’: the first ITR, the genetic cassette, and the second ITR; wherein the genetic cassette comprises a tissue-specific promoter sequence, an intronic sequence, the heterologous polynucleotide sequence, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence.

29. The nucleic acid molecule of claim 28, wherein the genetic cassette comprises from 5’ to 3’: a tissue-specific promoter sequence, an intronic sequence, the heterologous polynucleotide sequence, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence.

30. The nucleic acid molecule of claim 28 or 29, wherein:(a) the tissue specific promoter sequence comprises a TTT promoter;(b) the intron is a synthetic intron;(c) the post-transcriptional regulatory element comprises WPRE; and(d) the 3'UTR poly(A) tail sequence comprises bGHpA.

31. The nucleic acid molecule of any one of claims 1 to 30, wherein the genetic cassette comprises a single stranded nucleic acid.

32. The nucleic acid molecule of any one of claims 1 to 30, wherein the genetic cassette comprises a double stranded nucleic acid.

33. The nucleic acid molecule of any one of claims 1 to 32, wherein the heterologous polynucleotide sequence encodes a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or any combination thereof.

34. The nucleic acid molecule of claim 33, wherein the heterologous polynucleotide sequence encodes a growth factor selected from the group consisting of adrenomedullin (AM), angiopoietin (Ang), autocrine motility factor, a bone morphogenetic protein (BMP) (e.g. BMP2, BMP4, BMP5, BMP7), a ciliary neurotrophic factor family member (e.g., ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6)), a colony-stimulating factor (e.g., macrophage colony-stimulating factor (m-CSF), granulocyte colony-stimulating factor (G-CSF), granulocyte macrophage colony-stimulating factor (GM-CSF)), an epidermal growth factor (EGF), an ephrin (e.g., ephrin A1 , ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1 , ephrin B2, ephrin B3),erythropoietin (EPO), a fibroblast growth factor (FGF) (e.g. , FGF1 , FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF1 1 , FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21 , FGF22, FGF23), foetal bovine somatotrophin (FBS), a GDNF family member (e.g. , glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, artemin), growth differentiation factor-9 (GDF9), hepatocyte growth factor (HGF), hepatoma- derived growth factor (HDGF), insulin, an insulin-like growth factors (e.g., insulin-like growth factor-1 (IGF- 1 ) or IGF-2, an interleukin (IL) (e.g., IL-1 , IL-2, IL-3, IL-4, IL-5, IL-6, IL-7), keratinocyte growth factor (KGF), migration-stimulating factor (MSF), macrophage-stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)), myostatin (GDF-8), a neuregulin (e.g., neuregulin 1 (NRG1), NRG2, NRG3, NRG4), a neurotrophin (e.g., brain-derived neurotrophic factor (BDNF), nerve growth factor (NGF), a neurotrophin-3 (NT-3), NT-4, placental growth factor (PGF), platelet-derived growth factor (PDGF), renalase (RNLS), T-cell growth factor (TCGF), thrombopoietin (TPO), a transforming growth factor (e.g., transforming growth factor alpha (TGF-a), TGF-b, tumor necrosis factor-alpha (TNF-a), and vascular endothelial growth factor (VEGF), and any combination thereof.

35. The nucleic acid molecule of claim 33, wherein the heterologous polynucleotide sequence encodes a hormone.

36. The nucleic acid molecule of claim 33, wherein the heterologous polynucleotide sequence encodes a cytokine.

37. The nucleic acid molecule of claim 33, wherein the heterologous polynucleotide sequence encodes an antibody or a fragment thereof.

38. The nucleic acid molecule of any one of claims 1 to 32, wherein the heterologous polynucleotide sequence encodes a gene selected from dystrophin X-linked, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1 , FXN (frataxin), GUCY2D, RS1 , CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), and any combination thereof.

39. The nucleic acid molecule of any one of claims 1 to 32, wherein the heterologous polynucleotide sequence encodes a microRNA (miRNA).

40. The nucleic acid molecule of claim 39, wherein the miRNA down regulates the expression of a target gene selected from SOD1 , HTT, RHO, and any combination thereof41. The nucleic acid molecule of any one of claims 1 to 32, wherein the heterologous polynucleotide sequence encodes a clotting factor selected from the group consisting of factor I (FI), factor II (Fll), factor III (Fill), factor IV (FVI), factor V (FV), factor VI (FVI), factor VII (FVII), factor VIII (FVIII), factor IX (FIX), factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FVIII), Von Willebrand factor (VWF), prekallikrein, high-molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, Protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator(tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), and any combination thereof.

42. The nucleic acid molecule of claim 41 , wherein the clotting factor is FVIII.

43. The nucleic acid molecule of claim 42, wherein the FVIII comprises full-length mature FVIII.

44. The nucleic acid molecule of claim 43, wherein the FVIII comprises an amino acid sequence at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to an amino acid sequence having SEQ ID NO: 106.

45. The nucleic acid molecule of claim 42, wherein the FVIII comprises A1 domain, A2 domain, A3 domain, C1 domain, C2 domain, and a partial or no B domain.

46. The nucleic acid molecule of claim 45, wherein the FVIII comprises an amino acid sequence at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 109.

47. The nucleic acid molecule of any one of claims 41 to 46, wherein the clotting factor comprises a heterologous moiety.

48. The nucleic acid molecule of claim 47, wherein the heterologous moiety is selected from the group consisting of albumin or a fragment thereof, an immunoglobulin Fc region, the C- terminal peptide (CTP) of the b subunit of human chorionic gonadotropin, a PAS sequence, a HAP sequence, a transferrin or a fragment thereof, an albumin-binding moiety, a derivative thereof, or any combination thereof.

49. The nucleic acid molecule of claim 47 or 48, wherein the heterologous moiety is linked to the N-terminus or the C-terminus of the FVIII or inserted between two amino acids in the F VI 11.

50. The nucleic acid molecule of claim 49, wherein the heterologous moiety is inserted between two amino acids at one or more insertion site selected from the insertion sites listed in Table 4.

51. The nucleic acid molecule of any one of claims 42 to 50, wherein the FVIII further comprises A1 domain, A2 domain, C1 domain, C2 domain, an optional B domain, and a heterologous moiety, wherein the heterologous moiety is inserted immediately downstream of amino acid 745 corresponding to mature FVIII (SEQ ID NO: 106).

52. The nucleic acid molecule of any one of claims 49 to 51 , wherein the FVIII further comprises an FcRn binding partner.

53. The nucleic acid molecule of claim 52, wherein the FcRn binding partner comprises an Fc region of an immunoglobulin constant domain.

54. The nucleic acid molecule of any one of claims 42 to 53, wherein the nucleic acid sequence encoding the FVIII is codon optimized.

55. The nucleic acid molecule of any one of claims 42 to 54, wherein the nucleic acid sequence encoding the FVIII is codon optimized for expression in a human.

56. The nucleic acid molecule of claim 55, wherein the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to a nucleotide sequence of SEQ ID NO: 107.

57. The nucleic acid molecule of claim 55, wherein the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO: 71.

58. The nucleic acid molecule of any one of claims 1 to 53, wherein the heterologous polynucleotide sequence is codon optimized.

59. The nucleic acid molecule of claim 58, wherein the heterologous polynucleotide sequence is codon optimized for expression in a human.

60. The nucleic acid molecule of any one of claims 1 to 59, wherein the nucleic acid molecule is formulated with a delivery agent.

61. The nucleic acid molecule of claim 60, wherein the delivery agent comprises a lipid nanoparticle.

62. The nucleic acid molecule of claim 60, wherein the delivery agent is selected from the group consisting of liposomes, non-lipid polymeric molecules, and endosomes, and any combination thereof.

63. The nucleic acid molecule of any one of claims 1 to 62, wherein the nucleic acid molecule is formulated for intravenous, transdermal, intradermal, subcutaneous, pulmonary, or oral delivery, or any combination thereof.

64. The nucleic acid molecule of claim 63, wherein the nucleic acid molecule is formulated for intravenous delivery.

65. A vector comprising the nucleic acid molecule of any one of claims 1 to 59.

66. A host cell comprising the nucleic acid molecule of any one of claims 1 to 59.

67. A pharmaceutical composition comprising the nucleic acid of any one of claims 1 to 59 or the vector of claim 65, and a pharmaceutically acceptable excipient.

68. A pharmaceutical composition comprising the host cell of claim 66 and a pharmaceutically acceptable excipient.

69. A kit, comprising the nucleic acid molecule of any one of claims 1 to 59 and instructions for administering the nucleic acid molecule to a subject in need thereof.

70. A baculovirus system for production of the nucleic acid molecule of any one of claims 1 to 59.

71. The baculovirus system of claim 70, wherein the nucleic acid molecule of any one of claims 1 to 59 is produced in insect cells.

72. A nanoparticle delivery system comprising the nucleic acid molecule of any one of claims 1 to 59.

73. A method of producing a polypeptide, comprising culturing the host cell of claim 66 under suitable conditions and recovering the polypeptide.

75. A method of producing a polypeptide with clotting activity, comprising: culturing a host cell of claim 66 under suitable conditions and recovering the polypeptide with clotting activity.

76. A method of expressing a heterologous polynucleotide sequence in a subject in need thereof, comprising administering to the subject the nucleic acid molecule of any one of claims 1 to 59, the vector of claim 65, or the pharmaceutical composition of claim 67.

77. A method of expressing a clotting factor in a subject in need thereof, comprising administering to the subject the nucleic acid molecule of any one of claims 41 to 57, the vector of claim 65, the polypeptide of claim 75, or the pharmaceutical composition of claim 67.

78. A method of treating a disease or disorder in a subject in need thereof, comprising administering to the subject the nucleic acid molecule of any one of claims 1 to 59, the vector of claim 65, or the pharmaceutical composition of claim 67.

79. A method of treating a subject having a clotting factor deficiency, comprising administering to the subject the nucleic acid molecule of any one of claims 41 to 57, the vector of claim 65, the polypeptide of claim 75, or the pharmaceutical composition of claim 67.

80. A method of treating a clotting factor deficiency in a subject in need thereof, comprising administering to the subject the nucleic acid molecule of any one of claims 41 to 57, the vector of claim 65, the polypeptide of claim 75, or the pharmaceutical composition of claim 67.

81. The method of claim 79 or 80, wherein the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonarily, or any combination thereof.

82. The method of claim 81 , wherein the nucleic acid molecule is administered intravenously.

83. The method of any one of claims 79 to 82, further comprising administering to the subject a second agent.

84. The method of any one of claims 79 to 83, wherein the subject is a mammal.

85. The method of any one of claims 79 to 84, wherein the subject is a human.

86. The method of any one of claims 79 to 85, wherein the administration of the nucleic acid molecule to the subject results in an increased FVIII activity, relative to a FVIII activity in the subject prior to the administration, wherein the FVIII activity is increased by at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 1 1- fold, at least about 12-fold, at least about 13-fold, at least about 14-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, or at least about 100-fold.

87. The method of any one of claims 79 to 86, wherein the subject has a bleeding disorder.

88. The method of claim 87, wherein the bleeding disorder is a hemophilia.

89. The method of claim 87 or 88, wherein the bleeding disorder is hemophilia A.

90. A method of treating a bleeding disorder in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding a clotting factor, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

91. A method of treating hemophilia A in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding factor VIII (FVIII), wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

92. A method of treating a metabolic disorder of the liver in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding a liver-associated metabolic enzyme that is deficient in the subject, wherein the first ITR and / or second ITR are an ITR of a non-adeno-associated virus (non- AAV).

93. The method of claim 92, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

94. A method of treating a metabolic disorder of the liver in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first invertedterminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding a liver-associated metabolic enzyme that is deficient in the subject, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

95. The method of any one of claims 92 to 94, wherein the genetic cassette comprises a single stranded nucleic acid.

96. The method of any one of claims 92 to 94, wherein the genetic cassette comprises a double stranded nucleic acid.

97. The method of any one of claims 92 to 96, wherein the metabolic disorder of the liver is selected from the group consisting of phenylketonuria (PKU), a urea cycle disease, a lysosomal storage disorder, and a glycogen storage disease.

98. The method of claim 97, wherein the metabolic disorder of the liver is phenylketonuria (PKU).

99. The method of any one of claims 92 to 98, wherein the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonarily, or any combination thereof.

100. The method of claim 99, wherein the nucleic acid molecule is administered intravenously.

101. The method of any one of claims 92 to 100, further comprising administering to the subject a second agent.

102. The method of any one of claims 92 to 101 , wherein the subject is a mammal.

103. The method of any one of claims 92 to 102, wherein the subject is a human.

104. A method of treating phenylketonuria (PKU) in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a genetic cassette comprising a heterologous polynucleotide sequence encoding phenylalanine hydroxylase, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.

105. The method of claim 104, wherein the genetic cassette comprises a single stranded nucleic acid.

106. The method of claim 104, wherein the genetic cassette comprises a double stranded nucleic acid.

107. The method of any one of claims 104 to 106, wherein the nucleic acid molecule is formulated with a delivery agent.

108. The method of claim 107, wherein the delivery agent comprises a lipid nanoparticle.

109. A method of cloning a nucleic acid molecule, comprising inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex.1 10. The method of claim 109, wherein the disruption in the SbcCD complex comprises a genetic disruption in the SbcC gene and / or SbcD gene.1 1 1. The method of claim 109 or 1 10, wherein the disruption in the SbcCD complex comprises a genetic disruption in the SbcC gene.1 12. The method of claim 109 or 1 10, wherein the disruption in the SbcCD complex comprises a genetic disruption in the SbcD gene.1 13. The method of any one of claims 109 to 1 12, wherein the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first and / or second ITR is a non-adeno-associated virus (non-AAV) ITR.1 14. The method of any one of claims 109 to 1 13, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.1 15. The method of any one of claims 109 to 1 14, wherein the nucleic acid molecule further comprises a genetic cassette, wherein the genetic cassette is flanked by the first ITR and second ITR.1 16. The method of claim 1 15, wherein the genetic cassette comprises a heterologous polynucleotide sequence.1 17. The method of any one of claims 109 to 1 16, wherein the suitable vector is a low copy vector.1 18. The method of any one of claims 109 to 116, wherein the suitable vector is pBR322.1 19. The method of any one of claims 109 to 1 18, wherein the bacterial host strain is incapable of resolving cruciform DNA structures.

120. The method of any one of claims 109 to 118, wherein the bacterial host strain is PMC103, comprising the genotype sbcC, recD, mcrA, AmcrBCF.

121. The method of any one of claims 109 to 118, wherein the bacterial host strain is PMC107, comprising the genotype recBC, recJ, sbcBC, mcrA, AmcrBCF.

122. The method of any one of claims 109 to 1 18, wherein the bacterial host strain is SURE, comprising the genotype recB, recJ, sbcC, mcrA, AmcrBCF, umuC, uvrC.

123. A method of cloning a nucleic acid molecule, comprising inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector, and introducing the resulting vector into a bacterial host strain comprising a disruption in the SbcCD complex, wherein the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, wherein the first ITR and / or second ITR comprises a nucleotide sequence at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence set forth in SEQ ID NO: 180, 181 , 183, 184, 185, 186, 187 or 188, or a functional derivative thereof.