Nucleic acid molecule and use thereof

By using nucleic acid molecules with non-AAV inverted terminal repeats and gene cassettes, the problems of viral packaging capacity and immune response of AAV vectors were solved, enabling sustained expression and efficient therapeutic protein expression in vivo, especially coagulation factors.

CN120905307APending Publication Date: 2025-11-07BIOVILA DIVI THERAPEUTICS INC
View PDF 132 Cites 0 Cited by

Patent Information

Application Number
CN202510759063.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-08-09
Filing Date
2018-08-09
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing adeno-associated virus (AAV) vectors in gene therapy suffer from limited viral packaging capacity, high immunogenicity, and antibody-dependent inhibition of repeated treatments, making it difficult to effectively and sustainably express target sequences such as therapeutic proteins and miRNAs.

Method used

Nucleic acid molecules containing non-AAV inverted terminal repeats (ITRs) and gene cassettes encoding therapeutic proteins such as coagulation factors are used, combined with tissue-specific promoters and post-transcriptional regulatory elements, and delivered to target cells using lipid nanoparticles to achieve sustained expression.

Benefits of technology

It improved the expression level of therapeutic proteins, avoided the immune response and repeated treatment obstacles of AAV vectors, and achieved sustained and effective expression in vivo.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The present invention relates to nucleic acid molecules and uses thereof, in particular to nucleic acid molecules comprising a first reverse terminal repeat (ITR), a second ITR, and a gene cassette encoding miRNA and / or a therapeutic protein. In certain embodiments, the therapeutic protein comprises a blood coagulation factor, e.g., a FVIII polypeptide, a FIX polypeptide, or a fragment thereof. In some embodiments, the first ITR and / or the second ITR are / is an ITR of a non-adeno-associated virus (AAV). The invention also provides a method of treating a hemorrhagic condition, such as hemophilia, comprising administering to a subject the nucleic acid molecule or the polypeptide encoded thereby.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application patent application with the application date of August 9, 2018, the application number of 201880065324.X (the international application number of PCT / US2018 / 046110), and the name of "Nucleic Acid Molecules and Uses Thereof".

[0002] Incorporation by Reference of Electronically Submitted Sequence Listing

[0003] The content of the electronically submitted sequence listing in ASCII text file (Name: 4159_493PC01_ST25; Size: 434,561 bytes; and Date of Creation: August 9, 2018) is incorporated herein by reference in its entirety. BACKGROUND

[0004] Gene therapy offers a way to treat a variety of diseases in a persistent manner. In the past, gene therapy has generally relied on the use of viruses. A number of viral agents are available for this purpose, each with unique properties that will make it more or less suitable for gene therapy. Zhou et al., Adv Drug Deliv Rev. 106(Pt A):3-26, 2016. However, the undesirable properties of some viral vectors, including their immunogenicity or propensity to cause cancer, have led to clinical safety concerns, and until recently, their current use in the clinic has been limited to certain applications, such as vaccines and oncolytic strategies. Cotter et al., Front Biosci. 10: 1098-105 (2005).

[0005] Adeno-associated virus (AAV) is one of the most commonly studied gene therapy vectors. AAV is a protein shell that surrounds and protects a small, single-stranded DNA genome of about 4.8 kilobases (kb). Naso et al., BioDrugs, 31(4):317-334, 2017. AAV belongs to the Parvoviridae family and depends on co-infection with other viruses, primarily adenovirus, in order to replicate. Id. Its single-stranded genome contains three genes: Rep (replication), Cap (capsid), and aap (assembly). Id. These coding sequences are flanked by inverted terminal repeats (ITRs) required for genome replication and packaging. Id. The two AAV ITRs, which are approximately 145 nucleotides in length, are interrupted by palindromic sequences that can fold into T-shaped hairpin structures that act as primers during the initiation of DNA replication.

[0006] However, the use of conventional AAV as a gene delivery vehicle has resulted in several drawbacks. One of the main drawbacks relates to the limited viral packaging capacity of AAV of about 4.5 kb of heterologous DNA. (Dong et al., Hum Gene Ther. 7(17):2101-12, 1996). In addition, the administration of AAV vectors can induce an immune response in humans. Although AAV has been shown to be less immunogenic than some other viruses (i.e., adenovirus), the capsid proteins can trigger multiple components of the human immune system. See Naso et al., 2017. AAV is a common virus in the human population, and most people have been exposed to AAV, thus, most people have developed an immune response against the particular variant to which they were previously exposed. This pre-existing adaptive response can include NAbs and T cells that can reduce the clinical effect of subsequent re-infection with AAV and / or eliminate cells that have been transduced, which makes patients with pre-existing anti-AAV immunity ineligible for AAV-based gene therapy treatments. Whether administered locally or systemically, the virus will be seen as a foreign protein, and thus, the adaptive immune system will attempt to eliminate it. In addition, anti-AAV neutralizing antibodies induced by AAV treatment hinder repeat treatment with AAV when the first AAV treatment fails to achieve therapeutic levels of efficacy. Furthermore, evidence suggests that the T-shaped hairpin loop of the AAV ITR is susceptible to inhibition by host cell protein / protein complexes that bind the T-shaped hairpin structure of the AAV ITR. See, e.g., Zhou et al., Scientific Reports 7:5432 (July 14, 2017).

[0007] Accordingly, there is a need in the art for effective and sustained expression of target sequences, e.g., therapeutic proteins and / or miRNAs, in both in vitro and in vivo settings, while avoiding some of the unintended consequences and limitations of existing AAV vector technology. SUMMARY

[0008] One aspect of the disclosure relates to a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette encoding a therapeutic protein; wherein the first ITR and / or the second ITR is a non-adenovirus-associated virus (non-AAV) ITR, wherein the gene cassette is positioned between the first ITR and the second ITR, and wherein the therapeutic protein comprises a coagulation factor. In some embodiments, the non- AAV is selected from the group consisting of members of the Parvoviridae family and any combination thereof. In some embodiments, the first ITR is a non- AAV ITR and the second ITR is an adeno-associated virus (AAV) ITR, or wherein the first ITR is an AAV ITR and the second ITR is a non- AAV ITR. In other embodiments, the first ITR and the second ITR are non- AAV ITRs. In some embodiments, the first ITR and the second ITR are identical. In some embodiments, the first ITR and / or the second ITR comprises a palindromic sequence that is uninterrupted. In other embodiments, the first ITR and / or the second ITR comprises a palindromic sequence that is interrupted.

[0009] In some embodiments, the non-AAV in the nucleic acid molecule is a member of the Parvoviridae family of viruses. In some embodiments, the member of the Parvoviridae family of viruses is selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof. In some embodiments, the member of the Parvoviridae family of viruses is Erythrovirus parvovirus B19 (human virus). In some embodiments, the member of the Parvoviridae family of viruses is a Muscovy duck parvovirus (MDPV) strain. In some embodiments, the MDPV strain is the attenuated FZ91-30. In some embodiments, the MDPV strain is the pathogenic YY. In yet other embodiments, the Dependoparvovirus is a Goose parvovirus (GPV) strain. In some embodiments, the GPV strain is the attenuated 82-0321V. In some embodiments, the GPV strain is the pathogenic B. In some embodiments, the member of the Parvoviridae family of viruses is selected from the group consisting of porcine parvovirus (U44978), mouse minute virus (U34256), canine parvovirus (M19296), mink enteritis virus (D00765), and any combination thereof.

[0010] In some embodiments, the nucleic acid molecule further comprises a promoter. In some embodiments, the promoter is a tissue-specific promoter. In some embodiments, the promoter drives expression of the therapeutic protein in a hepatocyte, an endothelial cell, a muscle cell, a sinusoidal cell, or any combination thereof. In some embodiments, the promoter is located 5' of the nucleic acid sequence encoding the coagulation factor. In some embodiments, the promoter is selected from the group consisting of a mouse thyroxine promoter (mTTR), an endogenous human factor VIII promoter (F8), a human alpha-1-antitrypsin promoter (hAAT), a human albumin minimal promoter, a mouse albumin promoter, a Tristetraprolin (TTP) promoter, a CASI promoter, a CAG promoter, a cytomegalovirus (CMV) promoter, alpha 1-antitrypsin (AAT), muscle creatine kinase (MCK), myosin heavy chain alpha (aMHC), myoglobin (MB), desmin (DES), SPc5-12, 2R5Sc5-12, dMCK, tMCK, phosphogly cerate kinase (PGK) promoter, and any combination thereof. In some embodiments, the promoter comprises a TTP promoter.

[0011] In some embodiments, the nucleic acid molecule further comprises an intron sequence. In some embodiments, the intron sequence is located 5' of the nucleic acid sequence encoding the coagulation factor. In some embodiments, the intron sequence is located 3' of the promoter. In some embodiments, wherein the intron sequence comprises a synthetic intron sequence. In some embodiments, the intron sequence comprises SEQ ID NO: 115.

[0012] In some embodiments, the nucleic acid molecule comprises a post-transcriptional regulatory element. In some embodiments, the post-transcriptional regulatory element is located 3' of the nucleic acid sequence encoding the coagulation factor. In some embodiments, the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, or any combination thereof. In some embodiments, the microRNA binding site comprises a binding site for miR142-3p.

[0013] In some embodiments, the nucleic acid molecule comprises a 3' UTR poly(A) tail sequence. In some embodiments, the 3' UTR poly(A) tail sequence is selected from the group consisting of a bGH poly(A), an actin poly(A), a hemoglobin poly(A), and any combination thereof. In some embodiments, the 3' UTR poly(A) tail sequence comprises a bGH poly(A).

[0014] In some embodiments, the nucleic acid molecule comprises an enhancer sequence. In some embodiments, the enhancer sequence is located between the first ITR and the second ITR.

[0015] In some embodiments, the nucleic acid molecule comprises, in this order: a first ITR, a gene cassette, and a second ITR; wherein the gene cassette comprises a tissue-specific promoter sequence, an intron sequence, a nucleic acid sequence encoding a therapeutic protein (e.g., a coagulation factor) or a miRNA, a post-transcriptional regulatory element, and a 3’ UTR poly(A) tail sequence.

[0016] In some embodiments, the nucleic acid molecule comprises, in this order: a tissue-specific promoter sequence, an intron sequence, a nucleic acid sequence encoding a FVIII polypeptide, a post-transcriptional regulatory element, and a 3’ UTR poly(A) tail sequence.

[0017] In one embodiment, the nucleic acid molecule comprises:

[0018] (a) a first ITR that is an ITR of a non-AAV family member of the Parvoviridae family;

[0019] (b) a tissue-specific promoter sequence, e.g., a TTP promoter;

[0020] (c) an intron, e.g., a synthetic intron;

[0021] (d) a nucleotide sequence encoding a therapeutic protein (e.g., a coagulation factor) or a miRNA;

[0022] (e) a post-transcriptional regulatory element, e.g., a WPRE;

[0023] (f) a 3’ UTR poly(A) tail sequence, e.g., a bGHpA; and

[0024] (g) a second ITR that is an ITR of a non-AAV family member of the Parvoviridae family.

[0025] In some embodiments, the nucleic acid molecule comprises a single-stranded nucleic acid. In some embodiments, the gene cassette comprises a double-stranded nucleic acid.

[0026] In some embodiments, the nucleic acid molecule comprises a gene encoding a therapeutic protein (e.g., a coagulation factor), wherein the coagulation factor is expressed by a hepatocyte, an endothelial cell, a muscle cell, a sinusoidal cell, or any combination thereof. In some embodiments, the coagulation factor comprises Factor I (FI), Factor II (FII), Factor V (FV), Factor VII (FVII), Factor VIII (FVIII), Factor IX (FIX), Factor X (FX), Factor XI (FXI), Factor XII (FXII), Factor XIII (FVIII), von Willebrand Factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor- 1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), or any combination thereof. In some embodiments, the coagulation factor is FVIII. In some embodiments, the FVIII comprises full-length mature FVIII. In some embodiments, the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 106. In some embodiments, the FVIII comprises an Al domain, an A2 domain, an A3 domain, a Cl domain, a C2 domain, and a partial B domain or does not comprise a B domain. In some embodiments, the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 109.

[0027] In some embodiments, the coagulation factor comprises a heterologous moiety. In some embodiments, the heterologous moiety is selected from the group consisting of an albumin or fragment thereof, an immunoglobulin Fc region, a C-terminal peptide of the beta subunit of human chorionic gonadotropin (CTP), a PAS sequence, a HAP sequence, a transferrin or fragment thereof, an albumin binding moiety, a derivative thereof, and any combination thereof. In some embodiments, the heterologous moiety is linked to the N-terminus or C-terminus of FVIII or is inserted between two amino acids in the FVIII. In some embodiments, the heterologous moiety is inserted between two amino acids at one or more insertion sites selected from the insertion sites listed in Table 5. In some embodiments, the FVIII further comprises an Al domain, an A2 domain, a Cl domain, a C2 domain, an optional B domain, and a heterologous moiety, wherein the heterologous moiety is inserted immediately downstream of an amino acid corresponding to amino acid 745 of mature FVIII (SEQ ID NO: 106).

[0028] In some embodiments, the FVIII further comprises an FcRn binding partner. In some embodiments, the FcRn binding partner comprises an Fc region of an immunoglobulin constant domain. In some embodiments, the nucleic acid sequence encoding the FVIII is codon optimized. In some embodiments, the nucleic acid sequence encoding the FVIII is codon optimized for expression in humans.

[0029] In some embodiments, the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO: 107. In some embodiments, the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO: 71.

[0030] In some embodiments, the nucleic acid molecule is formulated with a delivery agent. In some embodiments, the delivery agent comprises one or more lipid nanoparticles. In some embodiments, the delivery agent is selected from the group consisting of a liposome, a non-lipid polymeric molecule, an endosome, and any combination thereof.

[0031] In some embodiments, the nucleic acid molecule is formulated for intravenous, transdermal, intradermal, subcutaneous, pulmonary, or oral delivery, or any combination thereof. In some embodiments, the nucleic acid molecule is formulated for intravenous delivery.

[0032] In some embodiments, provided herein is a vector comprising a nucleic acid molecule as described throughout the disclosure. In some embodiments, provided herein is a polypeptide encoded by a nucleic acid molecule as described throughout the disclosure. In some embodiments, provided is a host cell comprising a nucleic acid molecule as described throughout the disclosure.

[0033] In some embodiments, provided herein is a pharmaceutical composition comprising (a) a nucleic acid as described herein, a vector comprising a nucleic acid molecule as described throughout the disclosure, a polypeptide encoded by a nucleic acid molecule as described throughout the disclosure, or a host cell comprising a nucleic acid molecule as described throughout the disclosure; (b) an LNP; and (c) a pharmaceutically acceptable excipient. In some embodiments, provided herein is a kit comprising a nucleic acid molecule as described throughout the disclosure and instructions for administering the nucleic acid molecule to a subject in need thereof. In some embodiments, provided herein is a baculovirus system for producing a nucleic acid molecule described herein. In some embodiments, provided herein is a baculovirus in which a nucleic acid molecule disclosed herein is produced in an insect cell.

[0034] In some embodiments, provided herein is a nanoparticle delivery system for an expression construct, wherein the expression construct comprises a nucleic acid molecule described herein.

[0035] Also disclosed herein is a method of producing a polypeptide having coagulation activity, comprising: culturing a host cell disclosed herein under suitable conditions and recovering the polypeptide having coagulation activity. In some embodiments, disclosed herein is a method of expressing a coagulation factor in a subject in need thereof, comprising administering to the subject a nucleic acid molecule disclosed herein, a vector disclosed herein, a polypeptide disclosed herein, or a pharmaceutical composition disclosed herein. In some embodiments, disclosed herein is a method of treating a subject having a coagulation factor deficiency, comprising administering to the subject a nucleic acid molecule disclosed herein, a vector disclosed herein, a polypeptide disclosed herein, or a pharmaceutical composition disclosed herein. In some embodiments, the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonarily, or any combination thereof. In some embodiments, the nucleic acid molecule is administered intravenously. In some embodiments, the method further comprises administering to the subject a second agent. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.

[0036] In some embodiments, administration of the nucleic acid molecule to the subject results in increased FVIII activity relative to FVIII activity in the subject prior to the administration, wherein the FVIII activity is increased by at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 11-fold, at least about 12-fold, at least about 13-fold, at least about 14-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, or at least about 100-fold.

[0037] In some embodiments, the subject has a bleeding disorder. In some embodiments, the bleeding disorder is hemophilia. In some embodiments, the bleeding disorder is hemophilia A.

[0038] In particular, the present application includes, but is not limited to, the following:

[0039] 1. A nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette;

[0040] wherein the first ITR and / or the second ITR is a non-adenovirus-associated virus (non- AAV) ITR, wherein the gene cassette is located between the first ITR and the second ITR, and wherein the gene cassette encodes a therapeutic protein, an miRNA, or both a therapeutic protein and an miRNA.

[0041] 2. The nucleic acid molecule of item 1, wherein the therapeutic protein comprises a coagulation factor.

[0042] 3. The nucleic acid molecule of item 1 or item 2, wherein the non- AAV is selected from a member of the Parvoviridae family of viruses.

[0043] 4. The nucleic acid molecule of any one of items 1 to 3, wherein the first ITR and the second ITR are non- AAV ITRs.

[0044] 5. The nucleic acid molecule of item 3 or item 4, wherein the member of the family Parvoviridae is selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, Muscovy duck parvovirus (MDPV) strain, porcine parvovirus (U44978), mice minute virus (U34256), canine parvovirus (M19296), and mink enteritis virus (D00765).

[0045] 6. The nucleic acid molecule of any one of items 3 to 5, wherein the member of the family Parvoviridae is Erythrovirus B19 (human virus).

[0046] 7. The nucleic acid molecule of any one of items 3 to 5, wherein the family Parvoviridae is Dependovirus Goose parvovirus (GPV) strain.

[0047] 8. The nucleic acid molecule of any one of items 1 to 7, further comprising a tissue-specific promoter.

[0048] 9. The nucleic acid molecule of item 8, wherein the promoter drives expression of the therapeutic protein in hepatocytes, endothelial cells, muscle cells, crypt cells, or any combination thereof.

[0049] 10. The nucleic acid molecule of item 8 or item 9, wherein the promoter is selected from the group consisting of a mouse thyroxine promoter (mTTR), an endogenous human Factor VIII promoter (F8), a human alpha- 1 -antitrypsin promoter (hAAT), a human albumin minimal promoter, a mouse albumin promoter, a Tristetraprolin (TTP) promoter, a CASI promoter, a CAG promoter, a cytomegalovirus (CMV) promoter, an alpha 1 -antitrypsin (AAT), a muscle creatine kinase (MCK), a myosin heavy chain alpha (aMHC), a myoglobin (MB), a desmin (DES), a SPc5-12, a 2R5Sc5-12, a dMCK, a tMCK, and a phosphogly cerate kinase (PGK) promoter.

[0050] 11. The nucleic acid molecule of any one of items 1 to 10, wherein the nucleotide sequence further comprises:

[0051] (a) an intron sequence,

[0052] (b) a post-transcriptional regulatory element,

[0053] (c) a 3' UTR poly(A) tail sequence,

[0054] (d) an enhancer sequence, or

[0055] (e) any combination of (a)-(d).

[0056] 12. The nucleic acid molecule of item 11, wherein:

[0057] (a) the intron sequence is located 5' of the nucleic acid sequence encoding the coagulation factor;

[0058] (b) the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, or any combination thereof;

[0059] (c) the 3' UTR poly(A) tail sequence is selected from the group consisting of a bGH poly(A), an actin poly(A), a hemoglobin poly(A), and any combination thereof; or

[0060] (d) any combination of (a)-(c).

[0061] 13. The nucleic acid molecule of item 12, wherein:

[0062] (a) the intron sequence comprises SEQ ID NO: 115;

[0063] (b) the microRNA binding site comprises a binding site for miR142-3p;

[0064] (c) the 3' UTR poly(A) tail sequence comprises bGH poly(A); or

[0065] (d) any combination of (a)-(c).

[0066] 14. The nucleic acid molecule of any one of clauses 1-13, wherein the nucleic acid molecule comprises, in this order:

[0067] (a) the first ITR is an ITR of a non-AAV family member of the Parvoviridae family;

[0068] (b) a tissue-specific promoter sequence comprising a TTP promoter;

[0069] (c) an intron that is a synthetic intron;

[0070] (d) a nucleotide sequence encoding a blood clotting factor;

[0071] (e) a post-transcriptional regulatory element comprising a WPRE;

[0072] (f) a 3' UTR poly(A) tail sequence comprising bGH pA; and

[0073] (g) the second ITR is an ITR of a non-AAV family member of the Parvoviridae family.

[0074] 15. The nucleic acid molecule of any one of clauses 2-14, wherein the blood clotting factor comprises Factor I (FI), Factor II (FII), Factor V (FV), Factor VII (FVII), Factor VIII (FVIII), Factor IX (FIX), Factor X (FX), Factor XI (FXI), Factor XII (FXII), Factor XIII (FVIII), von Willebrand Factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor- 1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), or any combination thereof.

[0075] 16. The nucleic acid molecule of item 15, wherein the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to an amino acid sequence as set forth in a sequence selected from the group consisting of SEQ ID NO: 71, 106, 107, and 109.

[0076] 17. The nucleic acid molecule of any one of items 1 to 16, wherein the coagulation factor comprises a heterologous moiety selected from the group consisting of albumin or a fragment thereof, an immunoglobulin Fc region, a C-terminal peptide of the beta subunit of human chorionic gonadotropin (CTP), a PAS sequence, a HAP sequence, a transferrin or a fragment thereof, an albumin binding moiety, a derivative thereof, or any combination thereof.

[0077] 18. The nucleic acid molecule of any one of items 15 to 17, wherein the FVIII further comprises an FcRn binding partner.

[0078] 19. The nucleic acid molecule of any one of items 1 to 18, wherein the nucleic acid molecule is formulated with a delivery agent comprising a lipid nanoparticle.

[0079] 20. A pharmaceutical composition comprising the nucleic acid molecule of any one of items 1 to 19 and a pharmaceutically acceptable carrier.

[0080] 21. A method of expressing a coagulation factor in a subject in need thereof, comprising administering to the subject the nucleic acid molecule of any one of items 1 to 19 or the pharmaceutical composition of item 20. BRIEF DESCRIPTION OF DRAWINGS

[0081] Figure 1A is a schematic of a single-stranded coagulation factor (e.g., FVIII) expression cassette. The positions of the 5’ ITR from non-AAV (with hairpin loop at the end of the ssDNA structure), the 3’ ITR from non-AAV (with hairpin loop), the promoter sequence (e.g., TTPp), and the transgene sequence (e.g., FVIII co6XTEN sequence with XTEN144 inserted within the B domain) are shown. The exemplary expression cassette also shows additional possible elements, e.g., an intron sequence, a WPREmut sequence, and a bGHpA sequence.

[0082] Figures 1B-1D is a schematic of a plasmid used to make a single-stranded coagulation factor expression cassette (e.g., the cassette shown in Figure 1A ). The ITRs of the cassette are derived from AAV2 Figure 1B ), B19 Figure 1C ), or GPV Figure 1D). The plasmid construct containing the ssFVIII expression cassette as shown was digested with PvuII (at the PvuII site) Figure 1B ) or with LguI (at the LguI site) Figure 1C and Figure 1D ) to release the viral genome. Double stranded DNA was heated to 95°C to generate ssDNA, then incubated at 4°C to allow ITR structure formation.

[0083] Figure 2A is a phylogenetic tree illustrating the relationship between members of the various parvovirus families. B19, AAV-2 and GPV are marked with open boxes.

[0084] Figure 2B is a schematic of various cassettes, including hairpins.

[0085] Figure 3A and Figure 3B is an alignment of the ITRs of B19, GPV and AAV2 Figure 3A ) and the ITRs of B19 and GPV Figure 3B ). Gray shading shows homology.

[0086] Figures 4A-4C shows FVIII plasma activity after administration of single stranded FVIII-AAV naked DNA (ssAAV-FVIII; Figure 4A ), ssDNA-B19 FVIII Figure 4B ) or ssDNA-GPV FVIII Figure 4C ) via hydrodynamic injection (HDI) in Hem A mice. FVIII activity (as percent of control) was measured in plasma samples at 24 hours, 3 days, 2 weeks, 3 weeks, 1 month, 2 months, 3 months and 4 months in mice treated with a single HDI of ssDNA at 50 μg / mouse Figure 4C ), 20 μg / mouse Figure 4A and Figure 4B ), 10 μg / mouse Figure 4A and Figure 4C ) or 5 μg / mouse Figure 4A . HDI of plasmid DNA given at 5 μg / mouse served as a control Figures 4A-4C ). DETAILED DESCRIPTION

[0087] The disclosure describes plasmid-like nucleic acid molecules comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette (e.g., a gene cassette encoding a therapeutic protein or an miRNA), wherein the first ITR and / or the second ITR is a non-adenovirus-associated virus ITR (e.g., the first ITR and / or the second ITR is from a non-AAV). In some embodiments, the gene cassette encodes a therapeutic protein. In some embodiments, the therapeutic protein comprises a protein selected from a blood clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or a combination thereof. In some embodiments, the gene cassette encodes a dystrophin X-linked, MTM1 (muscle myosin), tyrosine hydroxylase, AADC, cyclase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid a-glucosidase), or any combination thereof.

[0088] In some embodiments, the therapeutic protein comprises a blood clotting factor. In a particular embodiment, the therapeutic protein comprises a FVIII or FIX protein.

[0089] In some embodiments, the gene cassette encodes an miRNA. In certain embodiments, the miRNA downregulates expression of a target gene selected from SOD1, HTT, RHO, or any combination thereof.

[0090] In certain embodiments, the non- AAV is selected from a member of the Parvoviridae family of viruses, and any combination thereof. The disclosure also relates to methods of expressing a therapeutic protein (e.g., a blood clotting factor, e.g., FVIII) in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette (e.g., a gene cassette encoding a therapeutic protein or an miRNA), wherein the first ITR and / or the second ITR is a non-adenovirus-associated virus (non- AAV) ITR. In certain embodiments, the disclosure describes an isolated nucleic acid molecule comprising a nucleotide sequence having sequence homology to a nucleotide sequence selected from SEQ ID NOs: 113 and 120.

[0091] Exemplary constructs of the disclosure are shown in the accompanying drawings and sequence listing. To provide a clear understanding of the specification and claims, the following definitions are provided below.

[0092] I. Definitions

[0093] It should be noted that the terms "a," "an," and "the" refer to one or more of something unless otherwise indicated. For example, "a nucleotide sequence" should be understood to refer to one or more nucleotide sequences. Similarly, "a therapeutic protein" and "a miRNA" should be understood to refer to one or more therapeutic proteins and one or more miRNAs, respectively. Likewise, the terms "a," "an," and "one or more" are used interchangeably in this document.

[0094] The term "about" is used herein to refer to approximately, roughly, around, or in a region thereof. When the term "about" is used in conjunction with a numerical range, it modifies that range by extending the upper and lower boundaries of the indicated numerical values. Generally, the term "about" is used herein to modify a numerical value above or below the specified value by a variation of magnitude of higher or lower than 10%.

[0095] Also as used herein, "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").

[0096] "Nucleic acid," "nucleic acid molecule," "nucleotide," "nucleotide sequence," and "polynucleotide" are used interchangeably and refer to a polymeric form of either ribonucleotides (adenine, guanine, uracil, or cytosine "RNA molecules") or deoxyribonucleotides (deoxyadenine, deoxyguanine, deoxythymidine, or deoxycytosine "DNA molecules"), or any phosphoester analogs thereof, such as phosphorothioates and phosphorodithioates, in either single stranded form, or as double-stranded helices. Single-stranded nucleic acid sequences refer to either single-stranded DNA (ssDNA) or single-stranded RNA (ssRNA). Double-stranded DNA-DNA, DNA-RNA, and RNA-RNA helices are possible. The term nucleic acid molecule, particularly DNA or RNA molecule, refers only to the primary and secondary structure of the molecule, and does not limit it to any particular tertiary forms. Thus, this term includes double-stranded DNA found, inter alia, in linear or circular DNA molecules (e.g., restriction fragments), plasmids, supercoiled DNA, and chromosomes. When discussing the structure of particular double-stranded DNA molecules, the sequence given is in the 5' to 3' direction according to the conventional convention for sequences along the non-transcribed strand, i.e., the strand having a sequence homologous to mRNA. A "recombinant DNA molecule" is a DNA molecule that has been subjected to a molecular biological manipulation. DNA includes, but is not limited to, cDNA, genomic DNA, plasmid DNA, synthetic DNA, and semi-synthetic DNA. The "nucleic acid compositions" of the present disclosure comprise one or more nucleic acids as described herein.

[0097] As used herein, a "inverted terminal repeat" (or "ITR") refers to a nucleic acid subsequence located at the 5' or 3' end of a single-stranded nucleic acid sequence that comprises a set of nucleotides (the initial sequence) that is followed downstream by an inverted complement, i.e., a palindromic sequence. The intercalating sequence of nucleotides between the initial sequence and the inverted complement can be of any length, including zero. In one embodiment, an ITR useful in the present disclosure comprises one or more "palindromic sequences." An ITR can have a number of functions. In some embodiments, an ITR described herein forms a hairpin structure. In some embodiments, an ITR forms a T-shaped hairpin structure. In some embodiments, an ITR forms a non-T-shaped hairpin structure, e.g., a U-shaped hairpin structure. In some embodiments, an ITR promotes long-term survival of a nucleic acid molecule in the nucleus of a cell. In some embodiments, an ITR promotes permanent survival of a nucleic acid molecule in the nucleus of a cell (e.g., for the entire lifetime of the cell). In some embodiments, an ITR promotes stability of a nucleic acid molecule in the nucleus of a cell. In some embodiments, an ITR promotes retention of a nucleic acid molecule in the nucleus of a cell. In some embodiments, an ITR promotes persistence of a nucleic acid molecule in the nucleus of a cell. In some embodiments, an ITR inhibits or prevents degradation of a nucleic acid molecule in the nucleus of a cell.

[0098] In one embodiment, the initial sequence and / or reverse complement comprises about 2-600 nucleotides, about 2-550 nucleotides, about 2-500 nucleotides, about 2-450 nucleotides, about 2-400 nucleotides, about 2-350 nucleotides, about 2-300 nucleotides, or about 2-250 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 5-600 nucleotides, about 10-600 nucleotides, about 15-600 nucleotides, about 20-600 nucleotides, about 25-600 nucleotides, about 30-600 nucleotides, about 35-600 nucleotides, about 40-600 nucleotides, about 45-600 nucleotides, about 50-600 nucleotides, about 60-600 nucleotides, about 70-600 nucleotides, about 80-600 nucleotides, about 90-600 nucleotides, about 100-600 nucleotides, about 150-600 nucleotides, about 200-600 nucleotides, about 300-600 nucleotides, about 350-600 nucleotides, about 400-600 nucleotides, about 450-600 nucleotides, about 500-600 nucleotides, or about 550-600 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 5-550 nucleotides, about 5 to 500 nucleotides, about 5-450 nucleotides, about 5 to 400 nucleotides, about 5-350 nucleotides, about 5 to 300 nucleotides, or about 5-250 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 10-550 nucleotides, about 15-500 nucleotides, about 20-450 nucleotides, about 25-400 nucleotides, about 30-350 nucleotides, about 35-300 nucleotides, or about 40-250 nucleotides. In certain embodiments, the initial sequence and / or reverse complement comprises about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, or about 600 nucleotides. In particular embodiments, the initial sequence and / or reverse complement comprises about 400 nucleotides.

[0099] In other embodiments, the initial sequence and / or reverse complement comprises about 2-200 nucleotides, about 5-200 nucleotides, about 10-200 nucleotides, about 20-200 nucleotides, about 30-200 nucleotides, about 40-200 nucleotides, about 50-200 nucleotides, about 60-200 nucleotides, about 70-200 nucleotides, about 80-200 nucleotides, about 90-200 nucleotides, about 100-200 nucleotides, about 125-200 nucleotides, about 150-200 nucleotides, or about 175-200 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2-150 nucleotides, about 5-150 nucleotides, about 10-150 nucleotides, about 20-150 nucleotides, about 30-150 nucleotides, about 40-150 nucleotides, about 50-150 nucleotides, about 75-150 nucleotides, about 100-150 nucleotides, or about 125-150 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2-100 nucleotides, about 5-100 nucleotides, about 10-100 nucleotides, about 20-100 nucleotides, about 30-100 nucleotides, about 40-100 nucleotides, about 50-100 nucleotides, or about 75-100 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2-50 nucleotides, about 10-50 nucleotides, about 20-50 nucleotides, about 30-50 nucleotides, about 40-50 nucleotides, about 3-30 nucleotides, about 4-20 nucleotides, or about 5-10 nucleotides. In another embodiment, the initial sequence and / or reverse complement consists of two nucleotides, three nucleotides, four nucleotides, five nucleotides, six nucleotides, seven nucleotides, eight nucleotides, nine nucleotides, ten nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides. In other embodiments, the intervening nucleotides between the initial sequence and the reverse complement are (e.g., consist of) 0 nucleotides, 1 nucleotide, two nucleotides, three nucleotides, four nucleotides, five nucleotides, six nucleotides, seven nucleotides, eight nucleotides, nine nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides.

[0100] Accordingly, as used herein, an "ITR" can fold upon itself and form a double-stranded segment. For example, the sequence GATCXXXXGATC comprises the initial sequence of GATC and its complement (3'CTAG5') when folded to form a duplex. In some embodiments, an ITR comprises a contiguous palindromic sequence (e.g., GATCGATC) between the initial sequence and the reverse complement. In some embodiments, an ITR comprises an interrupted palindromic sequence (e.g., GATCXXXXGATC) between the initial sequence and the reverse complement. In some embodiments, the complementary portions of the contiguous or interrupted palindromic sequence interact with each other to form a "hairpin loop" structure. As used herein, a "hairpin loop" structure results when at least two complementary sequences on base pairs of a single-stranded nucleotide molecule form a double-stranded portion. In some embodiments, only a portion of an ITR forms a hairpin loop. In other embodiments, the entire ITR forms a hairpin loop.

[0101] In the present disclosure, at least one ITR is an ITR of a non-adenovirus-associated virus (non-AAV). In certain embodiments, the ITR is an ITR of a non-AAV member of the Parvoviridae family of viruses. In some embodiments, the ITR is an ITR of a non-AAV member of the Dependovirus or Erythrovirus genera. In particular embodiments, the ITR is an ITR of a goose parvovirus (GPV), muscovy duck parvovirus (MDPV), or Erythrovirus B19 (also known as parvovirus B19, primate erythrovirus 1, B19 virus, and Erythrovirus). In certain embodiments, one of the two ITRs is an ITR of an AAV. In other embodiments, one of the two ITRs in the construct is an ITR selected from an AAV serotype of serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and any combination thereof. In a particular embodiment, the ITR is derived from an AAV serotype 2, e.g., an ITR of an AAV serotype 2.

[0102] In certain aspects of the present disclosure, the nucleic acid molecule comprises two ITRs, a 5' ITR and a 3' ITR, wherein the 5' ITR is located at the 5' end of the nucleic acid molecule and the 3' ITR is located at the 3' end of the nucleic acid molecule. The 5' ITR and the 3' ITR can be derived from the same virus or different viruses. In certain embodiments, the 5' ITR is derived from an AAV and the 3' ITR is not derived from an AAV virus (e.g., a non- AAV). In some embodiments, the 3' ITR is derived from an AAV and the 5' ITR is not derived from an AAV virus (e.g., a non- AAV). In other embodiments, the 5' ITR is not derived from an AAV virus (e.g., a non- AAV) and the 3' ITR is derived from the same or a different non- AAV virus.

[0103] The term "parvovirus" as used herein encompasses the Parvoviridae family, including but not limited to the autonomous replication parvovirus genus and the dependovirus genus. Autonomous parvoviruses include, for example, members of the Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus genera.

[0104] Exemplary autonomous parvoviruses include, but are not limited to, porcine parvovirus, mouse parvovirus, canine parvovirus, mink enterovirus, bovine parvovirus, chicken parvovirus, feline panleukopenia virus, feline parvovirus, goose parvovirus, H1 parvovirus, Muscovy duck parvovirus, snake parvovirus, and B19 virus. Other autonomous parvoviruses are known to those of skill in the art. See, e.g., VIROLOGY, Vol. 2, Chapter 69 (4th Ed., Lippincott-Raven Publishers).

[0105] The term "non-AAV" as used herein encompasses nucleic acids, proteins, and viruses from the Parvoviridae family other than adeno-associated viruses (AAVs). "Non-AAV" includes, but is not limited to, autonomous replication members of the Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus genera.

[0106] As used herein, the term "adeno-associated virus" (AAV), includes, but is not limited to, AAV type 1, AAV type 2, AAV type 3 (including types 3 A and 3B), AAV type 4, AAV type 5, AAV type 6, AAV type 7, AAV type 8, AAV type 9, AAV type 10, AAV type 11, AAV type 12, AAV type 13, snake AAV, bird AAV, bovine AAV, canine AAV, equine AAV, ovine AAV, caprine AAV, shrimp AAV (these AAV serotypes and clades are disclosed by Gao et al. (J. Virol. 78:6381 (2004)) and Moris et al. (Virol. 33:375 (2004)), and any other AAV now known or later discovered. See, e.g., FIELDS et al. VIROLOGY, vol. 2, chap. 69 (4th ed., Lippincott-Raven Publishers).

[0107] As used herein, the term "derived from" refers to a component that is isolated from, or uses, or is made from information (e.g., an amino acid or nucleic acid sequence) from, a specified molecule or organism. For example, a nucleic acid sequence (e.g., an ITR) derived from a second nucleic acid sequence (e.g., an ITR) can include a nucleotide sequence that is identical or substantially similar to the nucleotide sequence of the second nucleic acid sequence. Derivatives can be obtained in the case of nucleotides or polypeptides, for example, by naturally occurring mutagenesis, artificial directed mutagenesis, or artificial random mutagenesis. Mutagenesis of nucleotides or polypeptides to produce a different nucleotide or polypeptide derived from a first nucleotide or polypeptide can be intentionally directed or intentionally random, or a mixture of each. Mutagenizing a nucleotide or polypeptide to produce a different nucleotide or polypeptide derived from a first nucleotide or polypeptide can be a random event (e.g., caused by polymerase infidelity), and the identification of the derived nucleotide or polypeptide can be by appropriate screening methods, e.g., as described herein. Mutagenesis of a polypeptide typically requires manipulation of the polynucleotide encoding the polypeptide. In some embodiments, a nucleotide or amino acid sequence derived from a second nucleotide or amino acid sequence has at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the second nucleotide or amino acid, respectively, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In other embodiments, an ITR derived from a non-AAV (or AAV) ITR has at least 90% identity to the non-AAV ITR (or AAV ITR), respectively, wherein the non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or AAV ITR), respectively. In some embodiments, an ITR derived from a non-AAV (or AAV) ITR has at least 80% identity to the non-AAV ITR (or AAV ITR), respectively, wherein the non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or AAV ITR), respectively.In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 70% identical to the non-AAV ITR (or AAV ITR), respectively, wherein the non-AAV (or AAV) ITRs retain the functional properties of the non-AAV ITR (or AAV ITR), respectively. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 60% identical to the non-AAV ITR (or AAV ITR), respectively, wherein the non-AAV (or AAV) ITRs retain the functional properties of the non-AAV ITR (or AAV ITR), respectively. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 50% identical to the non-AAV ITR (or AAV ITR), respectively, wherein the non-AAV (or AAV) ITRs retain the functional properties of the non-AAV ITR (or AAV ITR), respectively.

[0108] In certain embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR. In some embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR, wherein the fragment comprises at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides, or at least about 600 nucleotides; wherein the ITR derived from a non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or AAV ITR), respectively. In certain embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR, wherein the fragment comprises at least about 129 nucleotides, and wherein the ITR derived from a non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or AAV ITR), respectively. In certain embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR, wherein the fragment comprises at least about 102 nucleotides, and wherein the ITR derived from a non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or AAV ITR), respectively.

[0109] In some embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of a non-AAV (or AAV) ITR, wherein the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the length of the non-AAV (or AAV) ITR.

[0110] In certain embodiments, the nucleotide or amino acid sequence derived from the second nucleotide or amino acid sequence is at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical, in correct alignment, to the homologous portion of the second nucleotide or amino acid sequence, respectively, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In other embodiments, the ITR derived from a non-AAV (or AAV) ITR is at least 90% identical, in correct alignment, to the homologous portion of the non-AAV ITR (or AAV ITR), respectively, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, the ITR derived from a non-AAV (or AAV) ITR is at least 80% identical, in correct alignment, to the homologous portion of the non-AAV ITR (or AAV ITR), respectively, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, the ITR derived from a non-AAV (or AAV) ITR is at least 70% identical, in correct alignment, to the homologous portion of the non-AAV ITR (or AAV ITR), respectively, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, the ITR derived from a non-AAV (or AAV) ITR is at least 60% identical, in correct alignment, to the homologous portion of the non-AAV ITR (or AAV ITR), respectively, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, the ITR derived from a non-AAV (or AAV) ITR is at least 50% identical, in correct alignment, to the homologous portion of the non-AAV ITR (or AAV ITR), respectively, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence.

[0111] A "capsid-free" or "capsid-less" vector or nucleic acid molecule refers to a vector construct that is free of a capsid. In some embodiments, a capsid-less vector or nucleic acid molecule does not contain a sequence encoding, for example, an AAV Rep protein.

[0112] As used herein, a "coding region" or "coding sequence" is a portion of a polynucleotide that consists of codons translated into amino acids. Although a "stop codon" (TAG, TGA, or TAA) is generally not translated into an amino acid, it can be considered part of a coding region, but any flanking sequences (e.g., promoters, ribosome binding sites, transcription terminators, introns, etc.) are not part of the coding region. The boundaries of a coding region are typically determined by a start codon at the 5' terminus and a translation stop codon at the 3' terminus, which together with the sequences in between them, encode a polypeptide. Two or more coding regions can be present in a single polynucleotide construct, e.g., on a single vector, or in separate polynucleotide constructs, e.g., on separate (different) vectors. Thus, a single vector can comprise only a single coding region, or two or more coding regions.

[0113] Certain proteins secreted by mammalian cells are associated with a secretory signal peptide that is cleaved from the mature protein once the growing protein chain has been exported through the rough endoplasmic reticulum. Those of ordinary skill in the art will appreciate that signal peptides are typically fused to the N-terminus of a polypeptide and cleaved from the intact or "full-length" polypeptide to yield the secreted or "mature" form of the polypeptide. In certain embodiments, a native signal peptide or a functional derivative of a sequence that retains the ability to direct secretion of a polypeptide is operatively associated therewith. Alternatively, a heterologous mammalian signal peptide (e.g., the human tissue plasminogen activator (TPA) or mouse beta-glucuronidase signal peptide) or a functional derivative thereof can be used.

[0114] The term "downstream" refers to a nucleotide sequence that is located 3' of a reference nucleotide sequence. In certain embodiments, a downstream nucleotide sequence relates to sequences after the start of transcription. For example, the translation initiation codon of a gene is located downstream of the start site of transcription.

[0115] The term "upstream" refers to a nucleotide sequence that is located 5' of a reference nucleotide sequence. In certain embodiments, an upstream nucleotide sequence relates to sequences located 5' of a coding region or the start of transcription. For example, most promoters are located upstream of the start site of transcription.

[0116] As used herein, the term "genetic control region" or "control region" refers to a nucleotide sequence located upstream of (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding region and which influences the transcription, RNA processing, stability, translation, of an associated coding region. Control regions can include promoters, translation leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites, or stem-and-loop structures. If the coding region is intended for expression in a eukaryotic cell, a polyadenylation signal and transcription termination sequence will usually be located 3' to the coding sequence.

[0117] A polynucleotide encoding a product (e.g., an miRNA or a gene product (e.g., a polypeptide such as a therapeutic protein)) can include a promoter and / or other expression (e.g., transcriptional or translational) control elements operably associated with one or more coding regions. In operable association, a coding region for a gene product (e.g., a polypeptide) is associated with one or more control regions such that expression of the gene product is under the influence or control of the control regions. For example, a coding region and a promoter are "operably associated" if induction of promoter function results in transcription of mRNA encoding the gene product encoded by that coding region, and if the nature of the linkage between the promoter and the coding region does not interfere with the ability of the promoter to direct expression of the gene product or interfere with the ability of the DNA template to be transcribed. In addition to a promoter, other expression control elements, such as enhancers, operators, repressors, and transcription termination signals, can also be operably associated with a coding region to direct gene product expression.

[0118] A "transcription control sequence" refers to DNA control sequences, such as a promoter, enhancer, terminators, and the like, that provide for the expression of a coding sequence in a host cell. A variety of transcription control regions are known to those of skill in the art. These include, without limitation, transcription control regions operable in vertebrate cells, such as, but not limited to, promoters and enhancer segments from cytomegalovirus (the immediate early promoter in conjunction with intron A), simian virus 40 (early promoter), and retroviruses (e.g., Rous sarcoma virus). Other transcription control regions include those derived from vertebrate genes (e.g., actin, heat shock

[0119] Similarly, a variety of translation control elements are known to those of ordinary skill in the art. These include, without limitation, ribosome binding sites, translation initiation and termination codons, and elements derived from picornaviruses (in particular, the internal ribosome entry site or IRES, also known as the CITE sequence).

[0120] The term "expression" as used herein refers to the process by which a polynucleotide produces a gene product (e.g., RNA or polypeptide). It includes, but is not limited to, transcription of a polynucleotide into messenger RNA (mRNA), transfer RNA (tRNA), small hairpin RNA (shRNA). small interfering RNA (siRNA), or any other RNA product, and translation of mRNA into a polypeptide. Expression results in a "gene product." As used herein, a gene product can be a nucleic acid, such as messenger RNA produced by transcription of a gene, or can be a polypeptide translated from a transcript. Gene products described herein also include nucleic acids with post-transcriptional modifications (e.g., polyadenylation or splicing), or polypeptides with post-translational modifications (e.g., methylation, glycosylation, addition of lipids, association with other protein subunits, or proteolytic cleavage). As used herein, the term "yield" refers to the amount of polypeptide produced by expression of a gene.

[0121] A "vector" refers to any vehicle used to clone and / or transfer a nucleic acid into a host cell. A vector can be a replicon, wherein another nucleic acid segment can be attached to the replicon, causing replication of the attached segment. A "replicon" refers to any genetic element, such as a plasmid, a bacteriophage, a cosmid, a chromosome, a virus, which is capable of replication either in its entirety or as a fragment. The term "vector" includes vehicles used to introduce nucleic acids into cells in vitro, ex vivo, or in vivo. A number of vectors are known and used in the art, which include, for example, plasmids, modified eukaryotic viruses, or modified bacterial viruses. Insertion of a polynucleotide into a suitable vector can be achieved by ligating the appropriate polynucleotide fragment into a suitable vector that has complementary cohesive ends.

[0122] Vectors can be engineered to encode a selectable marker or reporter molecule to provide for selection or identification of cells that have incorporated the vector. Expression of a selectable marker or reporter molecule allows for the identification and / or selection of host cells that have incorporated and are expressing other coding regions contained on the vector. Examples of selectable marker genes known and used in the art include: genes that provide resistance to ampicillin, streptomycin, gentamycin, kanamycin, hygromycin, bialaphos herbicide, sulfonamides, and the like; and genes that serve as phenotypic markers, i.e., anthocyanin regulatory genes, isopentyl transferase genes, and the like. Examples of reporter molecules known and used in the art include: luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), beta-galactosidase (LacZ), beta-glucuronidase (Gus), and the like. Selectable markers can also be considered reporter molecules.

[0123] The term "host cell" as used herein refers to, for example, microorganisms, yeast cells, insect cells, and mammalian cells, which can or have been used as recipients of ssDNA or vectors. The term includes the progeny of the original transduced cell. Thus, "host cell" as used herein generally refers to a cell that has been transduced with an exogenous DNA sequence. It is understood that the progeny of a single parental cell can not necessarily be completely identical in morphology or genome or total DNA complement as the original parent, due to natural, accidental, or deliberate mutation. In some embodiments, the host cell can be an in vitro host cell.

[0124] The term "selection marker" refers to an identifying factor that is capable of being selected based on the action of the marker gene (i.e., resistance to an antibiotic, resistance to a herbicide, colorimetric marker, enzyme, fluorescent marker, etc.), typically an antibiotic or chemical resistance gene, where the action is used to track the inheritance of a nucleic acid of interest and / or to identify cells or organisms that have inherited a nucleic acid of interest. Examples of selection marker genes known and used in the art include: genes that provide resistance to ampicillin, streptomycin, gentamycin, kanamycin, hygromycin, bialaphos herbicide, sulfonamides, etc.; and genes that serve as phenotypic markers, i.e., anthocyanin regulatory genes, isopentyl transferase genes, etc.

[0125] The term "reporter gene" refers to a nucleic acid that encodes an identifying factor that is capable of being identified based on the action of the reporter gene, where the action is used to track the inheritance of a nucleic acid of interest, to identify cells or organisms that have inherited a nucleic acid of interest, and / or to measure gene expression induction or transcription. Examples of reporter genes known and used in the art include: luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), beta-galactosidase (LacZ), beta-glucuronidase (Gus), etc. Selection marker genes can also be considered reporter genes.

[0126] "Promoter" and "promoter sequence" are used interchangeably and refer to a DNA sequence capable of controlling the expression of a coding sequence or functional RNA. Generally, a coding sequence is located 3' to a promoter sequence. Promoters can be derived in their entirety from natural genes, or be composed of different elements derived from different promoters, or even comprise synthetic DNA segments. It is understood by those skilled in the art that different promoters can direct the expression of a gene in different tissues or cell types, or at different stages of development, or in response to different environmental or physiological conditions. A promoter that causes a gene to be expressed in most cell types at most times is often referred to as a "constitutive promoter". A promoter that causes a gene to be expressed in specific cell types is often referred to as a "cell-specific promoter" or "tissue-specific promoter". A promoter that causes a gene to be expressed at a particular stage of development or cell differentiation is often referred to as a "developmental-specific promoter" or "cell differentiation-specific promoter". A promoter that is induced and causes a gene to be expressed upon exposure of a cell to, or treatment of a cell with, a pharmacological agent, biological molecule, chemical, ligand, light, etc. is often referred to as an "inducible promoter" or "regulatable promoter". It is further recognized that different lengths of DNA fragments can have the same promoter activity, since in most cases the exact boundaries of regulatory sequences have not been completely defined.

[0127] A promoter sequence is usually defined by a transcription initiation site at its 3' end, and extends upstream (5' direction) therefrom to include the minimum sequence of bases or elements required to initiate transcription at levels detectable above background. The transcription initiation site will be found upstream of the start of translation, and will be conveniently defined by, for example, nuclease SI mapping. A protein binding domain (consensus sequence) responsible for binding RNA polymerase will be found in the promoter sequence.

[0128] In some embodiments, the nucleic acid molecule comprises a tissue-specific promoter. In certain embodiments, the tissue-specific promoter drives expression of the therapeutic protein (e.g., a coagulation factor) in the liver (e.g., in hepatocytes and / or endothelial cells). In particular embodiments, the promoter is selected from the group consisting of a mouse thyroxin promoter (mTTR), an endogenous human factor VIII promoter (F8), a human alpha-1-antitrypsin promoter (hAAT), a human albumin minimal promoter, a mouse albumin promoter, a Tristetraprolin (TTP) promoter, a CASI promoter, a CAG promoter, a cytomegalovirus (CMV) promoter, a phosphoglycerate kinase (PGK) promoter, and any combination thereof. In some embodiments, the promoter is selected from the group consisting of a liver-specific promoter (e.g., alpha 1-antitrypsin (AAT)), a muscle-specific promoter (e.g., muscle creatine kinase (MCK), myosin heavy chain alpha (aMHC), myoglobin (MB), and desmin (DES)), a synthetic promoter (e.g., SPc5-12, 2R5Sc5-12, dMCK, and tMCK), and any combination thereof. In a particular embodiment, the promoter comprises a TTP promoter.

[0129] The terms "restriction endonuclease" and "restriction enzyme" are used interchangeably and refer to an enzyme that binds and cleaves within specific nucleotide sequences within double-stranded DNA.

[0130] The term "plasmid" refers to an extra-chromosomal element typically carrying genes that are not part of the central metabolism of the cell, and is usually in the form of a circular double-stranded DNA molecule. Such elements can be linear, circular or supercoiled, single- or double-stranded DNA or RNA of any source that are autonomously replicating sequences, genome integrating sequences, phasmids, or nucleotide sequences, many of which have been connected or recombined with other nucleotide sequences to bring about the unique property of the construct.

[0131] Eukaryotic viral vectors that can be used include, but are not limited to, adenoviral vectors, retroviral vectors, adeno-associated viral vectors, poxvirus (e.g., vaccinia virus) vectors, baculovirus vectors, or herpesvirus vectors. Non-viral vectors include plasmids, liposomes, cytofectins, DNA-protein complexes, and biopolymers.

[0132] A "cloning vector" refers to a "replicon" which is a unit length of nucleic acid that replicates sequentially and which contains an origin of replication, such as a plasmid, bacteriophage or cosmid, onto which another nucleic acid segment can be ligated so as to bring about replication of the attached segment. Certain cloning vectors are capable of replicating in one cell type (e.g., bacteria) and expressing in another cell type (e.g., eukaryotic cells). Cloning vectors typically contain one or more sequences that can be used to select for cells that contain the vector and / or one or more multiple cloning sites for insertion of nucleic acid sequences of interest.

[0133] The term "expression vector" refers to a vehicle that is designed to express an inserted nucleic acid sequence upon insertion into a host cell. The inserted nucleic acid sequence is operably associated with regulatory regions as described above.

[0134] Vectors are introduced into host cells by methods well known in the art, such as by transfection, electroporation, microinjection, transduction, cell fusion, DEAE dextran, calcium phosphate precipitation, lipofection (lysolipid fusion), use of a gene gun or DNA vector transporter. As used herein, "culturing," "to be cultured," and "being cultured" mean incubating a cell under in vitro conditions that allow the cell to grow or divide or maintain the cell in a viable state. As used herein, "cultured cells" means cells that have been propagated in vitro.

[0135] As used herein, the term "polypeptide" is intended to encompass both singular "polypeptide" as well as plural "polypeptides" and refers to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also referred to as peptide bonds). The term "polypeptide" refers to any one or more chains of two or more amino acids and does not refer to a particular length of the product. Thus, the definition of "polypeptide" includes a peptide, dipeptide, tripeptide, oligopeptide, "protein," "amino acid chain," or any other term used to refer to a chain of two or more amino acids. The term "polypeptide" can be used interchangeably with any of these terms. The term "polypeptide" is also intended to refer to post-expression modifications of the polypeptide, such as glycosylation, acetylation, phosphorylation, amidation, derivatization with known protecting / protecting groups, proteolytic processing, or modification by non-naturally occurring amino acids. A polypeptide can be derived from a natural biological source or produced by recombinant techniques, but is not necessarily translated from a specified nucleic acid sequence. It can be generated in any manner, including by chemical synthesis.

[0136] The term "amino acid" includes alanine (Ala or A); arginine (Arg or R); asparagine (Asn or N); aspartic acid (Asp or D); cysteine (Cys or C); glutamine (Gin or Q); glutamic acid (Glu or E); glycine (Gly or G); histidine (His or H); isoleucine (lie or I); leucine (Leu or L); lysine (Lys or K); methionine (Met or M); phenylalanine (Phe or F); proline (Pro or P); serine (Ser or S); threonine (Thr or T); tryptophan (Trp or W); tyrosine (Tyr or Y); and valine (Val or V). Non-traditional amino acids are also within the scope of the present disclosure and include norleucine, ornithine, norvaline, homoserine, and other amino acid residue analogs such as those described in Ellman et al. Meth. Enzym. 202:301-336 (1991). To generate such non-naturally occurring amino acid residues, the procedures of Noren et al. Science 244: 182 (1989) and Ellman et al. can be used. Briefly, these procedures involve chemically activating a suppressor tRNA with a non-naturally occurring amino acid residue and then transcribing and translating the RNA in vitro. Introduction of non-traditional amino acids can also be achieved using peptide chemistry known in the art. As used herein, the term "polar amino acid" includes amino acids that have a net zero charge but have non-zero partial charges at different parts of their side chains (e.g., M, F, W, S, Y, N, Q, C). These amino acids can participate in both hydrophobic and electrostatic interactions. As used herein, the term "charged amino acid" includes amino acids that can have a non-zero net charge on their side chains (e.g., R, K, H, E, D). These amino acids can participate in both hydrophobic and electrostatic interactions.

[0137] Also included in the present disclosure are fragments or variants of polypeptides, and any combinations thereof. When referring to a polypeptide binding domain or binding molecule of the present disclosure, the term "fragment" or "variant" includes polypeptides that retain at least some property of the referenced polypeptide (e.g., FcRn binding affinity of an FcRn binding domain or Fc variant, coagulation activity of an FVIII variant, or FVIII binding activity of a VWF fragment). In addition to the specific antibody fragments discussed elsewhere herein, fragments of polypeptides also include proteolytic fragments as well as deletion fragments, but not the naturally occurring full-length polypeptide (or mature polypeptide). Variants of a polypeptide binding domain or binding molecule of the present disclosure include fragments as described above, as well as polypeptides having an altered amino acid sequence due to amino acid substitution, deletion, or insertion. Variants can be natural or non-natural. Non-naturally occurring variants can be generated using mutagenesis techniques known in the art. Variant polypeptides can comprise conservative or non-conservative amino acid substitutions, deletions, or additions.

[0138] A "conservative amino acid substitution" is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art, including basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), beta-branched side chains (e.g., threonine, valine, isoleucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Thus, if an amino acid in a polypeptide is replaced with another amino acid from the same side chain family, then the substitution is considered to be conservative. In another embodiment, a string of amino acids can be conservatively substituted with a string of structurally similar, but sequentially and / or compositionally different, members of the side chain family.

[0139] The term "percent identity," as known in the art, is a relationship between two or more polypeptide sequences or two or more polynucleotide sequences, as determined by comparing the sequences. In the art, "identity" also refers to the degree of sequence correlation between polypeptide or polynucleotide sequences, as the case can be, as determined by the matching of strings of such sequences. "Identity" can be readily calculated by known methods, including, but not limited to, those described in: Computational Molecular Biology (Lesk, A. M., ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, D. W., ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, A. M. and Griffin, H. G., eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Stockton Press, New York (1991). Preferred methods to determine identity are designed to give the best match between sequences tested. Methods to determine identity have been incorporated into public available computer programs. Sequence alignment and percent identity calculations can be performed using sequence analysis software such as the Megalign program of the LASERGENE bioinformatics computing suite (DNASTAR, Madison, WI), the GCG program suite (Wisconsin Package Version 9.0, Genetics Computer Group (GCG), Madison, WI), BLASTP, BLASTN, BLASTX (Altschul et al., J. Mol. Biol. 215:403 (1990)), and DNASTAR (DNASTAR Inc., 1228 S. Park St. Madison, WI 53715 USA). In the context of the present application, it will be understood that where sequence analysis software is used for analysis, the results of the analysis will be based on the "default values" of the program referenced, unless otherwise specified. As used herein, "default values" refer to any set of values or parameters initially loaded with the software when first initialized.For purposes of determining percent identity between a therapeutic protein (e.g., a coagulation factor) sequence of the disclosure and a reference sequence, only the nucleotides in the reference sequence that correspond to nucleotides in the therapeutic protein (e.g., a coagulation factor) sequence of the disclosure are used to calculate percent identity. For example, when comparing a full-length FVIII nucleotide sequence containing a B domain to an optimized B domain-deleted (BDD) FVIII nucleotide sequence of the disclosure, the partial alignment including the Al, A2, A3, Cl, and C2 domains will be used to calculate percent identity. Nucleotides in the portion of the full-length FVIII sequence encoding the B domain that would result in a large "gap" in the alignment will not be counted as mismatches. Furthermore, in determining percent identity between an optimized BDD FVIII sequence of the disclosure or a specified portion thereof (e.g., nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3) and a reference sequence, percent identity will be calculated by dividing the number of matched nucleotides by the total number of nucleotides in the entire sequence of the optimized BDD-FVIII sequence or the specified portion thereof, as described herein.

[0140] As used herein, the nucleotides corresponding to the nucleotides in a particular sequence of the invention are identified by aligning the sequence of the invention to maximize identity to the reference sequence. The numbers used to identify the equivalent amino acids in the reference sequence are based on the numbers used to identify the corresponding amino acids in the sequences of the disclosure.

[0141] A "fusion" or "chimeric" protein comprises a first amino acid sequence linked to a second amino acid sequence that is not naturally linked thereto. Amino acid sequences that normally exist in separate proteins can be brought together in a fusion polypeptide, or amino acid sequences that normally exist in the same protein can be placed in a new arrangement in a fusion polypeptide, such as the fusion of a Factor VIII domain of the invention to an Ig Fc domain. Fusion proteins are produced, for example, by chemical synthesis or by production and translation of a polynucleotide encoding the peptide regions in the desired relationship. Chimeric proteins can further comprise a second amino acid sequence associated with the first amino acid sequence by covalent, non-peptide bonds, or non-covalent bonds.

[0142] As used herein, the term "insertion site" refers to a position in a polypeptide or fragment, variant, or derivative thereof that is immediately upstream of a position into which a heterologous moiety can be inserted. The "insertion site" is designated by a number that is the number of amino acids in the reference sequence. For example, an "insertion site" in FVIII refers to the number of the amino acid sequence in mature native FVIII (SEQ ID NO: 15) to which the insertion site corresponds that is immediately N-terminal to the insertion position. For example, the phrase "a3 comprises a heterologous moiety at an insertion site corresponding to amino acid 1656 of SEQ ID NO: 15" indicates that the heterologous moiety is located between the two amino acids corresponding to amino acid 1656 and amino acid 1657 of SEQ ID NO: 15.

[0143] As used herein, the phrase "immediately downstream of an amino acid" refers to the position immediately to the terminal carboxyl of an amino acid. Similarly, the phrase "immediately upstream of an amino acid" refers to the position immediately to the terminal amino of an amino acid.

[0144] As used herein, the term "inserted," "is inserted," "inserted into," or grammatically related terms refers to the position of a heterologous moiety in a polypeptide (e.g., a coagulation factor) relative to the analogous position in a parent polypeptide. For example, in certain embodiments, "inserted" and the like refers to the position of a heterologous moiety in a recombinant FVIII polypeptide relative to the analogous position in native mature human FVIII. As used herein, the term refers to a characteristic of a polypeptide and does not indicate, suggest, or infer any method or process by which the polypeptide was made.

[0145] As used herein, the term "half-life" refers to the biological half-life of a particular polypeptide in vivo. Half-life can be represented by the time required for a half of the amount administered to a subject to be cleared from circulation and / or other tissues in an animal. When a clearance curve for a given polypeptide is constructed as a function of time, the curve is typically biphasic, with a rapid alpha phase and a longer beta phase. The alpha phase typically represents the equilibration of the administered Fc polypeptide between intravascular and extravascular spaces, and is determined in part by the size of the polypeptide. The beta phase typically represents the metabolism of the polypeptide in the intravascular space. In some embodiments, therapeutic proteins (e.g., coagulation factors, e.g., FVIII) and chimeric proteins comprising the proteins are monophasic, and thus do not have an alpha phase, but only a single beta phase. Thus, in certain embodiments, the term half-life as used herein refers to the half-life of a polypeptide in the beta phase.

[0146] The term "linked" as used herein refers to a first amino acid sequence or nucleotide sequence that is covalently or non-covalently linked to a second amino acid sequence or nucleotide sequence, respectively. The first amino acid or nucleotide sequence can be directly linked or juxtaposed to the second amino acid or nucleotide sequence, or alternatively, an intervening sequence can covalently link the first sequence to the second sequence. The term "linked" is meant not only to refer to fusing a first amino acid sequence to a second amino acid sequence at the C-terminus or N-terminus, but also includes inserting all or a portion of the first amino acid sequence (or second amino acid sequence) into two amino acids in the second amino acid sequence (or first amino acid sequence), respectively. In one embodiment, the first amino acid sequence can be linked to the second amino acid sequence by a peptide bond or a linker. The first nucleotide sequence can be linked to the second nucleotide sequence by a phosphodiester bond or a linker. The linker can be a peptide or polypeptide (for polypeptide chains) or a nucleotide or nucleotide chain (for nucleotide chains) or any chemical moiety (for both polypeptide and polynucleotide chains). The term "linked" is also indicated by a hyphen (-).

[0147] Hemostasis, as used herein, means stopping or slowing bleeding or hemorrhage; or preventing or slowing the flow of blood through a blood vessel or body part.

[0148] Hemostatic disorder, as used herein, means a genetic or acquired condition characterized by a tendency to bleed spontaneously or due to trauma as a result of impaired or inability to form a fibrin clot. Examples of such disorders include hemophilia. The three major forms are hemophilia A (factor VIII deficiency), hemophilia B (factor IX deficiency or "Christmas disease") and hemophilia C (factor XI deficiency, mild bleeding tendency). Other hemostatic disorders include, for example, von Willebrand disease, factor XI deficiency (PTA deficiency), factor XII deficiency, deficiencies or structural abnormalities of fibrinogen, prothrombin, factor V, factor VII, factor X or factor XIII, Bernard-Soulier syndrome, which is a defect or deficiency in GPIb. GPIb, the receptor for vWF, can be defective and result in a deficiency in primary clot formation (primary hemostasis) and increased bleeding tendency, and Glanzman and Naegeli's thrombasthenia (Glanzmann thrombasthenia). In liver failure (both acute and chronic forms), the liver produces insufficient clotting factors; this increases the risk of bleeding.

[0149] The isolated nucleic acid molecules, isolated polypeptides, or vectors comprising the isolated nucleic acid molecules of the disclosure can be used prophylactically. The term "prophylactic treatment" as used herein refers to administration of the molecules prior to the onset of bleeding. In one embodiment, the subject in need of a general hemostatic agent is undergoing or is about to undergo surgery. The polynucleotides, polypeptides, or vectors of the disclosure can be administered as a prophylactic medication prior to or after surgery. The polynucleotides, polypeptides, or vectors of the disclosure can be administered during or after surgery to control acute bleeding episodes. Surgery can include, but is not limited to, liver transplantation, liver resection, dental surgery, or stem cell transplantation.

[0150] The isolated nucleic acid molecules, isolated polypeptides, or vectors of the disclosure are also used for on-demand treatment. The term "on-demand treatment" refers to administration of the isolated nucleic acid molecules, isolated polypeptides, or vectors in response to symptoms of a bleeding episode or prior to an activity that can cause bleeding. In one aspect, on-demand treatment can be given to a subject when bleeding begins, such as after an injury, or when bleeding is anticipated, such as prior to surgery. In another aspect, on-demand treatment can be given prior to an activity that increases the risk of bleeding, such as contact sports.

[0151] The term "acute bleeding" as used herein refers to an episode of bleeding, regardless of its underlying cause. For example, a subject can have a trauma, uremia, a genetic bleeding disorder (e.g., Factor VII deficiency), a platelet disorder, or resistance due to the production of antibodies against coagulation factors.

[0152] Treat, treatment, treating, as used herein, refers to, for example, a decrease in severity of a disease or condition; a reduction in length of time of a course of a disease or condition; an improvement in one or more symptoms associated with a disease or condition; providing a beneficial effect to a subject having a disease or condition without necessarily curing the disease or condition or preventing the one or more symptoms associated with the disease or condition. In one embodiment, the term "treatment" means maintaining a trough level of FVIII, for example, at least about 1 IU / dL, 2 IU / dL, 3 IU / dL, 4 IU / dL, 5 IU / dL, 6 IU / dL, 7 IU / dL, 8 IU / dL, 9 IU / dL, 10 IU / dL, 11 IU / dL, 12 IU / dL, 13 IU / dL, 14 IU / dL, 15 IU / dL, 16 IU / dL, 17 IU / dL, 18 IU / dL, 19 IU / dL, 20 IU / dL, 25 IU / dL, 30 IU / dL, 35 IU / dL, 40 IU / dL, 45 IU / dL, 50 IU / dL, 55 IU / dL, 60 IU / dL, 65 IU / dL, 70 IU / dL, 75 IU / dL, 80 IU / dL, 85 IU / dL, 90 IU / dL, 95 IU / dL, 100 IU / dL, 105 IU / dL, 110 IU / dL, 115 IU / dL, 120 IU / dL, 125 IU / dL, 130 IU / dL, 135 IU / dL, 140 IU / dL, 145 IU / dL, or 150 IU / dL in a subject by administering an isolated nucleic acid molecule, an isolated polypeptide, or a vector of the disclosure. In another embodiment, treatment means maintaining a trough level of FVIII at about 1 to about 150 IU / dL, about 1 to about 125 IU / dL, about 1 to about 100 IU / dL, about 1 to about 90 IU / dL, about 1 to about 85 IU / dL, about 1 to about 80 IU / dL, about 1 to about 75 IU / dL, about 1 to about 70 IU / dL, about 1 to about 65 IU / dL, about 1 to about 60 IU / dL, about 1 to about 55 IU / dL, about 1 to about 50 IU / dL, about 1 to about 45 IU / dL, about 1 to about 40 IU / dL, about 1 to about 35 IU / dL, about 1 to about 30 IU / dL, about 1 to about 25 IU / dL, about 25 to about 125 IU / dL, about 50 to about 100 IU / dL, about 50 to about 75 IU / dL, about 75 to about 100 IU / dL, about 1 to about 20 IU / dL, about 2 to about 20 IU / dL, about 3 to about 20 IU / dL, about 4 to about 20 IU / dL, about 5 to about 20 IU / dL, about 6 to about 20 IU / dL, about 7 to about 20 IU / dL, about 8 to about 20 IU / dL, about 9 to about 20 IU / dL, or about 10 to about 20 IU / dL.Treatment of a disease or condition can also include maintaining FVIII activity in the subject at a level that is at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 105%, 110%, 115%, 120%, 125%, 130%, 135%, 140%, 145%, or 150% of the level of FVIII activity in a non-hemophilia subject. The minimum trough level required for treatment can be measured by one or more known methods and can be adjusted (increased or decreased) for each person.

[0153] As used herein, "administering" means giving a pharmaceutically acceptable nucleic acid molecule, a polypeptide expressed therefrom, or a vector comprising a nucleic acid molecule of the disclosure to a subject via a pharmaceutically acceptable route. The route of administration can be intravenous, e.g., intravenous injection and intravenous infusion. Additional routes of administration include, e.g., subcutaneous, intramuscular, oral, nasal, and pulmonary administration. The nucleic acid molecules, polypeptides, and vectors can be administered as part of a pharmaceutical composition comprising at least one excipient.

[0154] The term "pharmaceutically acceptable" as used herein refers to molecular entities and compositions that are physiologically tolerable and do not typically produce toxicity or allergic or similar untoward reaction (such as gastric discomfort, dizziness, and the like) when administered to a human. Optionally, as used herein, the term "pharmaceutically acceptable" means approved by a regulatory agency of the Federal or a state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeia for use in animals, and more particularly in humans.

[0155] As used herein, the phrase "a subject in need thereof includes a subject, such as a mammalian subject, who would benefit from administration of a nucleic acid molecule, polypeptide, or vector of the disclosure, e.g., to improve hemostasis. In one embodiment, the subject includes, but is not limited to, an individual with hemophilia. In another embodiment, the subject includes, but is not limited to, an individual who has developed an inhibitor to a therapeutic protein (e.g., a clotting factor, e.g., FVIII) and thus requires bypass therapy. The subject can be an adult or a minor (e.g., under 12 years of age).

[0156] As used herein, the term "therapeutic protein" refers to any polypeptide known in the art to be administered to a subject. In some embodiments, the therapeutic protein comprises a protein selected from a blood clotting factor, a growth factor, an antibody, a functional fragment thereof, or a combination thereof. As used herein, the term "blood clotting factor" refers to a molecule naturally occurring or produced recombinantly or an analog thereof that prevents or reduces the duration of a bleeding episode in a subject. In other words, it means a molecule with pro-coagulant activity, i.e., responsible for the conversion of fibrinogen into an insoluble fibrin network, leading to blood coagulation or clotting. As used herein, "clotting factor" includes an activated clotting factor, a zymogen thereof, or an activatable clotting factor. An "activatable clotting factor" is a clotting factor in a non-activated form (e.g., a zymogen form thereof) that can be converted into an activated form. The term "blood clotting factor" includes, but is not limited to, Factor I (FI), Factor II (FII), Factor V (FV), FVII, FVIII, FIX, Factor X (FX), Factor XI (FXI), Factor XII (FXII), Factor XIII (FXIII), von Willebrand Factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor- 1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), a zymogen thereof, an activated form thereof, or any combination thereof.

[0157] As used herein, blood clotting activity means the ability to participate in a cascade of biochemical reactions that peak the formation of a fibrin clot and / or reduce the severity, duration, or frequency of bleeding or bleeding episodes.

[0158] As used herein, "growth factors" include any growth factor known in the art, which includes growth factors and hormones.In some embodiments, the growth factor is selected from the group consisting of adrenomedullin (AM), angiogenin (Ang), autocrine motility factor, bone morphogenetic protein (BMP) (e.g., BMP2, BMP4, BMP5, BMP7), ciliary neurotrophic factor family member (e.g., ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6)), colony stimulating factor (e.g., macrophage colony stimulating factor (m-CSF), granulocyte colony stimulating factor (G-CSF), granulocyte macrophage colony stimulating factor (GM-CSF)), epidermal growth factor (EGF), ephrin (e.g., ephrin Al, ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin Bl, ephrin B2, ephrin B3), erythropoietin (EPO), fibroblast growth factor (FGF) (e.g., FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, FGF23), foetal bovine somatotrophin (FBS), GDNF family member (e.g., glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, artemin), growth differentiation factor-9 (GDF9), hepatocyte growth factor (HGF), hepatoma-derived growth factor (HDGF), insulin, insulin-like growth factor (e.g., insulin-like growth factor-1 (IGF-1) or IGF-2, interleukin (IL) (e.g., IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7), keratinocyte growth factor (KGF), migration-stimulating factor (MSF), macrophage-stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)), myostatin (GDF-8), neuregulin (e.g., neuregulin 1 (NRG1), NRG2, NRG3, NRG4), neurotrophin (e.g., brain-derived neurotrophic factor (BDNF), nerve growth factor (NGF), neurotrophin-3 (NT-3), NT-4, placenta growth factor (PGF), platelet-derived growth factor (PDGF), rensin (RNLS), T-cell growth factor (TCGF), thrombopoietin (TPO), transforming growth factor (e.g., transforming growth factor alpha (TGF-a), TGF-b, tumor necrosis factor-alpha (TNF-a), and vascular endothelial growth factor (VEGF).

[0159] In some embodiments, the therapeutic protein is encoded by a gene selected from the group consisting of dystrophin X-linked, MTM1 (muscle myosin), tyrosine hydroxylase, AADC, cyclase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid a-glucosidase), or any combination thereof.

[0160] The term "heterologous" or "foreign" as used herein refers to a class of molecules that is not normally found in a given context, e.g., in a cell or polypeptide. For example, a foreign or heterologous molecule can be introduced into a cell and only exists after manipulation of the cell (e.g., by transfection or other forms of genetic engineering), or a heterologous amino acid sequence can exist in a protein that is not naturally found.

[0161] As used herein, the term "heterologous nucleotide sequence" refers to a nucleotide sequence that is not naturally occurring in a given polynucleotide sequence. In one embodiment, the heterologous nucleotide sequence encodes a polypeptide that is capable of extending the half-life of a therapeutic protein (e.g., a coagulation factor, e.g., FVIII). In another embodiment, the heterologous nucleotide sequence encodes a polypeptide that increases the hydrodynamic radius of a therapeutic protein (e.g., a coagulation factor, e.g., FVIII). In other embodiments, the heterologous nucleotide sequence encodes a polypeptide that improves one or more pharmacokinetic properties of a therapeutic protein without significantly affecting its biological activity or function (e.g., procoagulant activity). In some embodiments, the therapeutic protein is linked or connected to the polypeptide encoded by the heterologous nucleotide sequence via a linker. Non-limiting examples of polypeptide moieties encoded by the heterologous nucleotide sequence include an immunoglobulin constant region or portion thereof, albumin or a fragment thereof, an albumin binding moiety, transferrin, a PAS polypeptide of U.S. Patent Application No. 20100292130, a HAP sequence, a transferrin or a fragment thereof, a C-terminal peptide (CTP) of the beta subunit of human chorionic gonadotropin, an albumin binding small molecule, an XTEN sequence, an FcRn binding moiety (e.g., an intact Fc region or a portion thereof that binds FcRn), a single chain Fc region (ScFc region, e.g., as described in US 2008 / 0260738, WO 2008 / 012543, or WO 2008 / 1439545), a polyglycine linker, a polyserine linker, a peptide of 6-40 amino acids of two types of amino acids selected from glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), and proline (P), and short polypeptides whose secondary structure varies to a degree of less than 50% to more than 50%, and the like, or a combination of two or more thereof. In some embodiments, the polypeptide encoded by the heterologous nucleotide sequence is linked to a non-polypeptide moiety. Non-limiting examples of non-polypeptide moieties include polyethylene glycol (PEG), an albumin binding small molecule, polysialic acid, hydroxyethyl starch (HES), a derivative thereof, or any combination thereof.

[0162] As used herein, the term "Fc region" is defined as the polypeptide portion of a native Ig that corresponds to that of an Fc region of a native Ig, i.e., the portion formed by the dimeric association of the Fc domains of each of the two heavy chains. A native Fc region forms a homodimer with another Fc region. Conversely, the term "genetically fused Fc region" or "single chain Fc region" (scFc region) as used herein refers to a synthetic dimeric Fc region composed of Fc domains that are genetically linked within a single polypeptide chain (i.e., a polypeptide chain encoded in a single contiguous genetic sequence).

[0163] In one embodiment, "Fc region" refers to the portion of a single Ig heavy chain that begins in the upper hinge region immediately upstream of the papain cleavage site (i.e., residue 216 in IgG, with the first residue of the heavy chain constant region being 114) and ends at the C-terminus of the antibody. Thus, a complete Fc domain comprises at least a hinge domain, a CH2 domain, and a CH3 domain.

[0164] The Fc region of an Ig constant region can include CH2, CH3, and CH4 domains, as well as a hinge region, depending on the Ig isotype. Chimeric proteins comprising the Fc region of an Ig confer several desirable properties on the chimeric protein, including increased stability, extended serum half-life (see Capon et al., 1989, Nature 337:525), and binding to Fc receptors (e.g., neonatal Fc receptor (FcRn)) (U.S. Patent Nos. 6,086,875, 6,485,726, 6,030,613; WO 03 / 077834; US2003-0235536A1), which are incorporated herein by reference in their entireties.

[0165] A "reference nucleotide sequence," when used herein as a comparison to a nucleotide sequence of the disclosure, is a polynucleotide sequence that is substantially identical to a nucleotide sequence of the disclosure, except that the sequence has not been optimized. For example, a reference nucleotide sequence consisting of the codon-optimized BDD FVIII of SEQ ID NO: 1 and a heterologous nucleotide sequence encoding a single-chain Fc region that is linked to the 3' end of SEQ ID NO: 1 is a nucleic acid molecule consisting of the original (or "parental") BDD FVIII of SEQ ID NO: 16 and the same heterologous nucleotide sequence encoding a single-chain Fc region that is linked to the 3' end of SEQ ID NO: 16.

[0166] As used herein, the term "optimized" with respect to a nucleotide sequence means a polynucleotide sequence that encodes a polypeptide, wherein the polynucleotide sequence has been mutated to enhance a property of the polynucleotide sequence. In some embodiments, optimization is performed to increase transcription levels, to increase translation levels, to increase steady-state mRNA levels, to increase or decrease binding of regulatory proteins (such as general transcription factors), to increase or decrease splicing, or to increase production of a polypeptide produced by the polynucleotide sequence. Examples of changes that can be made to a polynucleotide sequence include codon optimization, G / C content optimization, removal of repetitive sequences, removal of AT-rich elements, removal of cryptic splice sites, removal of cis-acting elements that inhibit transcription or translation, addition or removal of polyT or polyA sequences, addition of sequences around the transcription start site that enhance transcription (such as Kozak consensus sequences), removal of sequences that can form stem-loop structures, removal of destabilizing sequences, and combinations of two or more thereof.

[0167] II. Nucleic Acid Molecules

[0168] The present disclosure relates to plasmid-like, capsid-free nucleic acid molecules encoding therapeutic proteins or genes that can modulate expression of target proteins. The capsid (protein shell of a virus) encloses the genetic material of a virus. The capsid is known to assist the function of the virion by protecting the viral genome, delivering the genome to the host, and interacting with the host. Nonetheless, the viral capsid can be a factor limiting the packaging capacity of the vector and / or inducing an immune response, especially when it is used in gene therapy.

[0169] AAV vectors have become one of the more common types of gene therapy vectors. However, the presence of the capsid limits the utility of AAV vectors in gene therapy. In particular, the capsid itself can limit the size of the transgene contained in the vector to as low as less than 4.5 kb. Even before the addition of regulatory elements, various therapeutic proteins available for gene therapy can easily exceed this size.

[0170] Furthermore, the proteins that make up the capsid can serve as antigens that can be targeted by the immune system of a subject. AAV is very prevalent in the general population, with most people having been exposed to AAV in their lifetime. Thus, most potential recipients of gene therapy can have already developed an immune response to AAV, and thus are more likely to reject the therapy.

[0171] Certain aspects of the present disclosure are intended to overcome these deficiencies of AAV vectors. In particular, certain aspects of the present disclosure relate to a nucleic acid molecule comprising a first ITR, a second ITR, and a gene cassette, e.g., that encodes a therapeutic protein and / or an miRNA. In some embodiments, the nucleic acid molecule does not comprise genes encoding capsid proteins, replication proteins, and / or assembly proteins. In some embodiments, the gene cassette encodes a therapeutic protein. In some embodiments, the therapeutic protein comprises a clotting factor. In some embodiments, the gene cassette encodes an miRNA. In certain embodiments, the gene cassette is located between the first ITR and the second ITR. In some embodiments, the nucleic acid molecule further comprises one or more non-coding regions. In certain embodiments, the one or more non-coding regions comprise a promoter sequence, an intron, a post-transcriptional regulatory element, a 3’ UTR poly(A) sequence, or any combination thereof.

[0172] In one embodiment, the gene cassette is a single-stranded nucleic acid. In another embodiment, the gene cassette is a double-stranded nucleic acid.

[0173] In one embodiment, the nucleic acid molecule comprises:

[0174] (a) a first ITR that is an ITR of a non-AAV family member of the Parvoviridae family;

[0175] (b) a tissue-specific promoter sequence, e.g., a TTP promoter;

[0176] (c) an intron, e.g., a synthetic intron;

[0177] (d) a nucleotide encoding a miRNA or a therapeutic protein, e.g., a blood clotting factor;

[0178] (e) a post-transcriptional regulatory element, e.g., a WPRE;

[0179] (f) a 3’ UTR poly(A) tail sequence, e.g., bGHpA;

[0180] (g) a second ITR that is an ITR of a non-AAV family member of the Parvoviridae family.

[0181] In one embodiment, the nucleic acid molecule comprises:

[0182] (a) a first ITR that is an ITR of a non-AAV family member of the Parvoviridae family;

[0183] (b) a tissue-specific promoter sequence, e.g., a TTP promoter;

[0184] (c) an intron, e.g., a synthetic intron;

[0185] (d) a nucleotide encoding a miRNA, wherein the miRNA downregulates expression of a target gene selected from the group consisting of SOD1, HTT, RHO, and any combination thereof;

[0186] (e) a post-transcriptional regulatory element, e.g., a WPRE;

[0187] (f) a 3’ UTR poly(A) tail sequence, e.g., bGHpA;

[0188] (g) a second ITR that is an ITR of a non-AAV family member of the Parvoviridae family.

[0189] In one embodiment, the nucleic acid molecule comprises:

[0190] (a) a first ITR that is an ITR of a non-AAV family member of the Parvoviridae family;

[0191] (b) a tissue-specific promoter sequence, e.g., a TTP promoter;

[0192] (c) an intron, e.g., a synthetic intron;

[0193] (d) nucleotides encoding dystrophin X-linked, MTM1 (muscle myosin), tyrosine hydroxylase, AADC, cyclase-hydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid a-glucosidase), or any combination thereof;

[0194] (e) a post-transcriptional regulatory element, e.g., WPRE;

[0195] (f) a 3’ UTR poly(A) tail sequence, e.g., bGHpA;

[0196] (g) a second ITR that is an ITR of a non-AAV family member of the Parvoviridae family.

[0197] In one embodiment, the nucleic acid molecule comprises:

[0198] (a) a first ITR that is an ITR of AAV (e.g., an AAV serotype 2 genome);

[0199] (b) a tissue-specific promoter sequence, e.g., a TTP promoter;

[0200] (c) an intron, e.g., a synthetic intron;

[0201] (d) nucleotides encoding FVIII; wherein the nucleotides have at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 1-14 or SEQ ID NO: 71, wherein the FVIII encoded by the nucleotides retains FVIII activity;

[0202] (e) a post-transcriptional regulatory element, e.g., WPRE;

[0203] (f) a 3’ UTR poly(A) tail sequence, e.g., bGHpA; and

[0204] (g) a second ITR that is an ITR of AAV (e.g., an AAV serotype 2 genome).

[0205] In one embodiment, the nucleic acid molecule comprises:

[0206] (a) a first ITR that is an ITR of AAV (e.g., an AAV serotype 2 genome);

[0207] (b) a tissue-specific promoter sequence, e.g., a TTP promoter;

[0208] (c) an intron, e.g., a synthetic intron;

[0209] (d) a nucleotide encoding a miRNA, wherein the miRNA downregulates expression of a target gene (e.g., SOD1, HTT, RHO, and any combination thereof);

[0210] (f) a 3’ UTR poly(A) tail sequence, e.g., bGH pA; and

[0211] (g) a second ITR that is an ITR of an AAV (e.g., an AAV serotype 2 genome).

[0212] In one embodiment, the nucleic acid molecule comprises:

[0213] (a) a first ITR that is an ITR of an AAV (e.g., an AAV serotype 2 genome);

[0214] (b) a tissue-specific promoter sequence, e.g., a TTP promoter;

[0215] (c) an intron, e.g., a synthetic intron;

[0216] (d) a nucleotide encoding a dystrophin X-linked, MTM1 (muscle myosin), tyrosine hydroxylase, AADC, cyclase-hydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid a-glucosidase), or any combination thereof;

[0217] (f) a 3’ UTR poly(A) tail sequence, e.g., bGH pA; and

[0218] (g) a second ITR that is an ITR of an AAV (e.g., an AAV serotype 2 genome).

[0219] In another embodiment, the nucleic acid molecule comprises:

[0220] (a) a first ITR;

[0221] (b) a tissue-specific promoter sequence, e.g., a TTP promoter;

[0222] (c) an intron, e.g., a synthetic intron;

[0223] (d) a nucleotide encoding a miRNA or a therapeutic protein (e.g., a clotting factor);

[0224] (e) a post-transcriptional regulatory element, e.g., a WPRE;

[0225] (f) a 3' UTR poly(A) tail sequence, e.g., bGH pA; and

[0226] (g) a second ITR,

[0227] wherein one of the first ITR or the second ITR is an ITR of a non-AAV family member of the Parvoviridae family and the other ITR is an ITR of an AAV (e.g., an AAV serotype 2 genome).

[0228] In another embodiment, the nucleic acid molecule comprises:

[0229] (a) a first ITR;

[0230] (b) a tissue-specific promoter sequence, a TTP promoter;

[0231] (c) an intron, e.g., a synthetic intron;

[0232] (d) a nucleotide encoding a miRNA or a therapeutic protein (e.g., a coagulation factor);

[0233] (e) a post-transcriptional regulatory element, e.g., a WPRE;

[0234] (f) a 3' UTR poly(A) tail sequence, e.g., bGH pA; and

[0235] (g) a second ITR,

[0236] wherein the first ITR is a synthetic ITR, the second ITR is a synthetic ITR, or both the first ITR and the second ITR are synthetic ITRs.

[0237] A. inverted terminal repeat

[0238] Certain aspects of the disclosure relate to a nucleic acid molecule comprising a first ITR, e.g., a 5' ITR, and a second ITR, e.g., a 3' ITR. Generally, ITRs are involved in parvovirus (e.g., AAV) DNA replication and rescue or excision from a prokaryotic plasmid (Samulski et al., 1983, 1987; Senapathy et al., 1984; Gottlieb and Muzyczka, 1988). In addition, ITRs appear to be the minimal sequences required for AAV provirus integration and packaging of AAV DNA into virions (McLaughlin et al., 1988; Samulski et al., 1989). These elements are essential for efficient amplification of the parvovirus genome. The minimal defined elements that are hypothesized to be essential for ITR function are the Rep binding site (e.g., RBS; GCGCGCTCGCTCGCTC (SEQ ID NO: 104) for AAV2) and the terminal resolution site (e.g., TRS; AGTTGG (SEQ ID NO: 105) for AAV2), plus a variable palindromic sequence that allows the formation of a hairpin. The palindromic nucleotide region generally functions in cis together as an origin of DNA replication and a packaging signal for the virus. During DNA replication, complementary sequences in the ITR fold into a hairpin structure. In other embodiments, the ITR folds into a non-T-shaped hairpin structure, e.g., a U-shaped hairpin structure. Data suggest that the T-shaped hairpin structure of AAV ITRs can inhibit expression of a transgene flanked by the ITRs. See, e.g., Zhou et al., Scientific Reports 7:5432 (July 4, 2017). By utilizing an ITR that does not form a T-shaped hairpin structure, this form of inhibition can be avoided. Thus, in certain aspects, a polynucleotide comprising a non-AAV ITR has improved transgene expression compared to a polynucleotide comprising an AAV ITR that forms a T-shaped hairpin.

[0239] In some embodiments, the ITR comprises a naturally occurring ITR, e.g., the ITR comprises all portions of a parvovirus ITR. In some embodiments, the ITR comprises a synthetic sequence. In one embodiment, the first ITR or the second ITR comprises a synthetic sequence. In another embodiment, each of the first ITR and the second ITR comprises a synthetic sequence. In some embodiments, the first ITR or the second ITR comprises a naturally occurring sequence. In another embodiment, each of the first ITR and the second ITR comprises a naturally occurring sequence.

[0240] In some embodiments, an ITR comprises or consists of a portion of a naturally occurring ITR (e.g., a truncated ITR). In some embodiments, an ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides, or at least about 600 nucleotides; wherein the ITR retains a functional property of the naturally occurring ITR. In certain embodiments, an ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 129 nucleotides; wherein the ITR retains a functional property of the naturally occurring ITR. In certain embodiments, an ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 102 nucleotides; wherein the ITR retains a functional property of the naturally occurring ITR.

[0241] In some embodiments, an ITR comprises or consists of a portion of a naturally occurring ITR, wherein the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the length of the naturally occurring ITR; wherein the fragment retains a functional property of the naturally occurring ITR.

[0242] In certain embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a cognate portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In other embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 90% sequence identity to a cognate portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 80% sequence identity to a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 70% sequence identity to a cognate portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 60% sequence identity to a cognate portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 50% sequence identity to a cognate portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR.

[0243] In some embodiments, the ITR comprises an ITR from an AAV genome. In some embodiments, the ITR is an ITR of an AAV genome selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, and any combination thereof. In a particular embodiment, the ITR is an ITR of an AAV2 genome. In another embodiment, the ITR is a synthetic sequence genetically engineered to include at the 5' and 3' ends ITRs derived from one or more AAV genomes.

[0244] In some embodiments, the ITR is not derived from an AAV genome. In some embodiments, the ITR is a non-AAV ITR. In some embodiments, the ITR is a non-AAV genome from the family Parvoviridae selected from, but not limited to, the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof. In certain embodiments, the ITR is derived from the Erythrovirus parvovirus B19 (human virus). In another embodiment, the ITR is derived from the Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, for example, the MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, for example, the MDPV strain YY. In some embodiments, the ITR is derived from a porcine parvovirus, for example, porcine parvovirus U44978. In some embodiments, the ITR is derived from a mouse minute virus, for example, mouse minute virus U34256. In some embodiments, the ITR is derived from a canine parvovirus, for example, canine parvovirus M19296. In some embodiments, the ITR is derived from a mink enteritis virus, for example, mink enteritis virus D00765. In some embodiments, the ITR is derived from a dependovirus. In one embodiment, the dependovirus is a strain of the genus Dependovirus, goose parvovirus (GPV). In a specific embodiment, the GPV strain is attenuated, for example, the GPV strain 82-0321V. In another specific embodiment, the GPV strain is pathogenic, for example, the GPV strain B.

[0245] The first and second ITRs of the nucleic acid molecule can be derived from the same genome (e.g., derived from the genome of the same virus), or from different genomes, e.g., from the genomes of two or more different viral genomes. In certain embodiments, the first and second ITRs are derived from the same AAV genome. In particular embodiments, the two ITRs present in the nucleic acid molecule of the application are identical, and can be, in particular, AAV2 ITRs. In other embodiments, the first ITR is derived from an AAV genome and the second ITR is not derived from an AAV genome (e.g., a non-AAV genome). In other embodiments, the first ITR is not derived from an AAV genome (e.g., a non-AAV genome) and the second ITR is derived from an AAV genome. In yet other embodiments, both the first and second ITRs are not derived from an AAV genome (e.g., a non-AAV genome). In one particular embodiment, the first and second ITRs are identical.

[0246] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a genome selected from the group consisting of a Bocavirinae, a Dependovirus, an Erythrovirus, a Mink Aleutian virus, a Parvovirus, a Densovirinae, a Reovirinae, a Contavirinae, an Aveparvovirus, a Copiparvovirus, a ProParvovirus, a Tet Parvovirus, a Bipartite Densovirus, a Brevidensovirus, a Hepan Densovirus, a Prawn Densovirus, and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a genome selected from the group consisting of a Bocavirinae, a Dependovirus, an Erythrovirus, a Mink Aleutian virus, a Parvovirus, a Densovirinae, a Reovirinae, a Contavirinae, an Aveparvovirus, a Copiparvovirus, a ProParvovirus, a Tet Parvovirus, a Bipartite Densovirus, a Brevidensovirus, a Hepan Densovirus, a Prawn Densovirus, and any combination thereof. In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of a Bocavirinae, a Dependovirus, an Erythrovirus, a Mink Aleutian virus, a Parvovirus, a Densovirinae, a Reovirinae, a Contavirinae, an Aveparvovirus, a Copiparvovirus, a ProParvovirus, a Tet Parvovirus, a Bipartite Densovirus, a Brevidensovirus, a Hepan Densovirus, a Prawn Densovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from the same genome. In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of a Bocavirinae, a Dependovirus, an Erythrovirus, a Mink Aleutian virus, a Parvovirus, a Densovirinae, a Reovirinae, a Contavirinae, an Aveparvovirus, a Copiparvovirus, a ProParvovirus, a Tet Parvovirus, a Bipartite Densovirus, a Brevidensovirus, a Hepan Densovirus, a Prawn Densovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from different genomes.

[0247] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from an Erythrovirus Parvovirus B19 (human virus). In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from an Erythrovirus Parvovirus B19 (human virus).

[0248] In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from B19. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 167, 168, 169, 170, and 171, wherein the first ITR and / or the second ITR retains the functional properties of the B19 ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 167, 168, 169, 170, and 171, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0249] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 167, 168, 169, 170, and 171. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 167. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 168. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 169. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 170. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 171.

[0250] Table 1. Sample parvovirus ITR sequences.

[0251]

[0252]

[0253]

[0254] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 169, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167, which nucleotide sequence is derived from a B19 ITR, retaining the functional properties of a B19 ITR. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 169, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0255] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a GPV. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a GPV.

[0256] In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from GPV. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from the group consisting of the nucleotide sequences set forth in SEQ ID NOs: 172, 173, 174, 175, and 176, wherein the first ITR and / or the second ITR retains the functional properties of a GPV ITR from which the first ITR and / or the second ITR is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from GPV. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from the group consisting of the nucleotide sequences set forth in SEQ ID NOs: 172, 173, 174, 175, and 176, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 172, 173, 174, 175, and 176. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 172. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 173. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 174. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 175. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 176.

[0257] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 174, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR retains the functional properties of the GPV ITR from which the first ITR and / or the second ITR is derived. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 174, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0258] In certain embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 176, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR retains the functional properties of the GPV ITR, which the first ITR and / or the second ITR is derived from. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 176, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0259] In certain embodiments, one of the first ITR or the second ITR comprises or consists of all or a portion of an ITR derived from AAV2. In some embodiments, the first ITR or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 177 or 178, wherein the first ITR and / or the second ITR retains the functional properties of an AAV2 ITR, which the first ITR and / or the second ITR is derived from. In some embodiments, the first ITR or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 177 or 178, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 177 or 178. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 177. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 178.

[0260] In some embodiments, the first ITR is derived from an AAV genome and the second ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In other embodiments, the second ITR is derived from an AAV genome and the first ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, e.g., MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, e.g., MDPV strain YY.

[0261] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a Dependovirus. In some embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a Dependovirus. In other embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a Dependovirus Goose Parvovirus (GPV) strain. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a Dependovirus GPV strain. In certain embodiments, the GPV strain is attenuated, e.g., GPV strain 82-0321V. In other embodiments, the GPV strain is pathogenic, e.g., GPV strain B.

[0262] In certain embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a genome selected from the group consisting of a porcine parvovirus, e.g., porcine parvovirus strain U44978; a mouse minivirus, e.g., mouse minivirus strain U34256; a canine parvovirus, e.g., canine parvovirus strain M19296; a mink enteritis virus, e.g., mink enteritis virus strain D00765; and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a genome selected from the group consisting of a porcine parvovirus, e.g., porcine parvovirus strain U44978; a mouse minivirus, e.g., mouse minivirus strain U34256; a canine parvovirus, e.g., canine parvovirus strain M19296; a mink enteritis virus, e.g., mink enteritis virus strain D00765; and any combination thereof.

[0263] In another particular embodiment, the ITRs are genetically engineered to include synthetic sequences at their 5' and 3' ends that are not derived from an ITR of an AAV genome. In another particular embodiment, the ITRs are genetically engineered to include synthetic sequences at their 5' and 3' ends that are derived from one or more ITRs of non-AAV genomes. The two ITRs present in the nucleic acid molecule of the application can be the same or different non-AAV genome. In particular, the ITRs can be derived from the same non-AAV genome. In specific embodiments, the two ITRs present in the nucleic acid molecule of the application are the same, and can in particular be AAV2 ITRs.

[0264] In some embodiments, the ITR sequence comprises one or more palindromic sequences. The palindromic sequences of the ITRs disclosed herein include, but are not limited to, natural palindromic sequences (i.e., sequences found in nature), synthetic sequences (i.e., sequences not found in nature), such as pseudo-palindromic sequences, and combinations or modified versions thereof. A "pseudo-palindromic sequence" is a palindromic DNA sequence, including imperfect palindromic sequences, that shares less than 80%, including less than 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5%, or no nucleic acid sequence identity to sequences in naturally occurring AAV or non-AAV palindromic sequences that form secondary structures. Natural palindromic sequences can be obtained or derived from any of the genomes disclosed herein. Synthetic palindromic sequences can be based on any of the genomes disclosed herein.

[0265] The palindromic sequence can be continuous or interrupted. In some embodiments, the palindromic sequence is interrupted, wherein the palindromic sequence comprises an insertion of a second sequence. In some embodiments, the second sequence comprises a promoter, an enhancer, an integration site for an integrase (e.g., a site for a Cre or Flp recombinase), an open reading frame for a gene product, or a combination thereof.

[0266] In some embodiments, the ITR forms a hairpin loop structure. In one embodiment, the first ITR forms a hairpin structure. In another embodiment, the second ITR forms a hairpin structure. In still another embodiment, both the first ITR and the second ITR form a hairpin structure. In some embodiments, the first ITR and / or the second ITR does not form a T-shaped hairpin structure. In certain embodiments, the first ITR and / or the second ITR forms a non-T-shaped hairpin structure. In some embodiments, the non-T-shaped hairpin structure comprises a U-shaped hairpin structure.

[0267] In some embodiments, the ITRs of the nucleic acid molecules described herein can be transcriptionally activated ITRs. The transcriptionally activated ITRs can comprise all or a portion of a wild-type ITR that has been transcriptionally activated by the inclusion of at least one transcriptionally active element. A variety of types of transcriptionally active elements are suitable for use in this context. In some embodiments, the transcriptionally active element is a constitutively transcriptionally active element. Constitutively transcriptionally active elements provide a constant level of gene transcription, and are preferred when expression of a transgene on a constant basis is desired. In other embodiments, the transcriptionally activating element is an inducible transcriptionally activating element. Inducible transcriptionally activating elements generally exhibit low activity in the absence of an inducer (or inducing conditions), and are upregulated in the presence of an inducer (or switching to inducing conditions). Inducible transcriptionally active elements can be preferred when expression is only desired at certain times or in certain locations, or when the use of an inducer to titrate expression levels is desired. The transcriptionally activating elements can also be tissue specific; that is, they exhibit activity only in certain tissues or cell types.

[0268] The transcriptional activation element can be incorporated into the ITR in a variety of ways. In some embodiments, the transcriptional activation element is incorporated 5' to any portion of the ITR or 3' to any portion of the ITR. In other embodiments, the transcriptional activation element of the transcriptionally activated ITR is positioned between two ITR sequences. If the transcriptional activation element comprises two or more elements that must be spaced apart, these elements can be interposed with a portion of the ITR. In some embodiments, the hairpin structure of the ITR is deleted and replaced with an inverted repeat of the transcriptional element. The latter arrangement will produce a hairpin that mimics the missing portion of the structure. Multiple transcriptionally active elements in tandem can also be present in the transcriptionally activated ITR, and they can be adjacent or spaced apart. In addition, a protein binding site (e.g., a Rep binding site) can be introduced into the transcriptional activation element of the transcriptionally activated ITR. The transcriptional activation element can comprise any sequence that is capable of controlling the transcription of DNA by an RNA polymerase to form RNA, and can comprise, for example, a transcriptionally active element as defined below.

[0269] The transcriptionally activated ITR provides the nucleic acid molecule with transcriptional activation and ITR functions in a relatively limited nucleotide sequence length, which effectively maximizes the length of the transgene that can be carried and expressed from the nucleic acid molecule. Incorporation of a transcriptionally active element into the ITR can be accomplished in a variety of ways. Comparison of ITR sequences and sequence requirements for the transcriptional activation element can provide an understanding of the manner in which the element is encoded within the ITR. For example, transcriptional activity can be added to the ITR by introducing specific changes in the ITR sequence that replicate functional elements of the transcriptional activation element. There are many techniques in the art that effectively add, delete, and / or alter specific nucleotide sequences at particular sites (see, e.g., Deng and Nickoloff (1992) Anal. Biochem. 200:81-88). Another way of creating a transcriptionally activated ITR involves introducing restriction sites at desired locations in the ITR. In addition, multiple transcriptional activation elements can be incorporated into the transcriptionally activated ITR using methods known in the art.

[0270] B. Therapeutic Proteins

[0271] Certain aspects of the present disclosure relate to a nucleic acid molecule comprising a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein. In some embodiments, the gene cassette encodes one therapeutic protein. In some embodiments, the gene cassette encodes more than one therapeutic protein. In some embodiments, the gene cassette encodes two or more copies of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more variants of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more different therapeutic proteins.

[0272] Certain embodiments of the present disclosure relate to a nucleic acid molecule comprising a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a coagulation factor. In some embodiments, the coagulation factor is selected from the group consisting of FI, FII, FIII, FIV, FV, FVI, FVII, FVIII, FIX, FX, FXI, FXII, FXIII), VWF, prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor- 1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), any proenzyme thereof, any activated form thereof, and any combination thereof. In one embodiment, the coagulation factor comprises FVIII or a variant or fragment thereof. In another embodiment, the coagulation factor comprises FIX or a variant or fragment thereof. In another embodiment, the coagulation factor comprises FVII or a variant or fragment thereof. In another embodiment, the coagulation factor comprises VWF or a variant or fragment thereof.

[0273] 1. A coagulation factor

[0274] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a Factor VIII polypeptide. As used herein, "Factor VIII" abbreviated as "FVIII" throughout this application, means a functional FVIII polypeptide that normally functions in coagulation, unless otherwise indicated. Thus, the term FVIII includes functional variant polypeptides. "FVIII protein" is used interchangeably with FVIII polypeptide (or protein) or FVIII. Examples of FVIII function include, but are not limited to, the ability to activate coagulation, the ability to act as a cofactor for Factor IX, or the ability to bind von Willebrand Factor (VWF). 2+and phospholipids, which then converts factor X to the activated form Xa. The FVIII protein can be a human, porcine, canine, rat, or mouse FVIII protein. In addition, comparisons between FVIII from humans and other species have identified conserved residues that can be required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US 6,251,632). Full-length polypeptide and polynucleotide sequences are known, as are many functional fragments, mutants, and modified variants. Various FVIII amino acid and nucleotide sequences are disclosed in, e.g., U.S. Pub. Nos. 2015 / 0158929 Al, 2014 / 0308280 Al, and 2014 / 0370035 Al, and International Pub. No. WO 2015 / 106052 Al. FVIII polypeptides include, e.g., full-length FVIII, full-length FVIII minus Met at the N-terminus, mature FVIII (minus signal sequence), mature FVIII with an additional Met at the N-terminus, and / or FVIII with all or part of the B domain deleted. FVIII variants include B domain deletions, whether partial or complete.

[0275] a. FVIII and polynucleotide sequences encoding FVIII proteins

[0276] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a Factor VIII polypeptide. Unless otherwise specified, as used herein “Factor VIII” abbreviated as “FVIII” throughout this application means a functional FVIII polypeptide that normally functions in coagulation. Thus, the term FVIII includes functional variant polypeptides. “FVIII protein” is used interchangeably with FVIII polypeptide (or protein) or FVIII. Examples of FVIII functions include, but are not limited to, the ability to activate coagulation, the ability to act as a cofactor for Factor IX, or the ability to bind to von Willebrand Factor (vWF). In some embodiments, the FVIII protein is a human FVIII protein. In some embodiments, the FVIII protein is a porcine FVIII protein. In some embodiments, the FVIII protein is a canine FVIII protein. In some embodiments, the FVIII protein is a rat FVIII protein. In some embodiments, the FVIII protein is a mouse FVIII protein. 2+and phospholipids to form a tenase complex with Factor IX, which then converts Factor X to activated form Xa. The FVIII protein can be human, porcine, canine, rat, or mouse FVIII protein. In addition, comparisons between FVIII from humans and other species have identified conserved residues that can be required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US 6,251,632). Full-length polypeptide and polynucleotide sequences are known, as are many functional fragments, mutants, and modified variants. Various FVIII amino acid and nucleotide sequences are disclosed in, e.g., U.S. Pub. Nos. 2015 / 0158929 Al, 2014 / 0308280 Al, and 2014 / 0370035 Al and International Pub. No. WO 2015 / 106052 Al. FVIII polypeptides include, e.g., full-length FVIII, full-length FVIII minus Met at the N-terminus, mature FVIII (minus signal sequence), mature FVIII with an additional Met at the N-terminus, and / or FVIII with all or part of the B domain deleted. FVIII variants include B domain deletions, whether partial or complete.

[0277] The FVIII portion in the chimeric proteins used herein has FVIII activity. FVIII activity can be measured by any known method in the art. A number of tests are available to assess the function of the coagulation system: activated partial thromboplastin time (aPTT) test, chromogenic assay, ROTEM assay, prothrombin time (PT) test (also used to determine INR), fibrinogen test (usually by the Clauss method), platelet count, platelet function test (usually by PFA-100), TCT, bleeding time, mixing test (whether a patient's plasma can be corrected by mixing with normal plasma), clotting factor assays, anti-phospholipid antibodies, D-dimer, genetic tests (e.g., Factor V Leiden, Prothrombin mutation G20210A), dilute Russell's viper venom time (dRVVT), other platelet function tests, thromboelastography (TEG or Sonoclot), thromboelastometry (TEg®) For example, ) or euglobulin lysis time (ELT).

[0278] The aPTT test is a performance indicator that measures the efficacy of the "intrinsic" (also called contact activation pathway) and common coagulation pathways. This test is commonly used to measure the coagulation activity of commercially available recombinant clotting factors (e.g., FVIII). It is used in conjunction with prothrombin time (PT), which measures the extrinsic pathway.

[0279] ROTEM analysis provides overall kinetic information on hemostasis: clotting time, clot formation, clot stability, and lysis. Different parameters in thrombelastography depend on the activity of the plasma clotting system, platelet function, fibrinolysis, or many factors that influence these interactions. This assay can provide a complete view of secondary hemostasis.

[0280] The chromogenic assay mechanism is based on the principle of the coagulation cascade, in which activated FVIII accelerates the conversion of factor X to factor Xa in the presence of activated factor IX, phospholipids, and calcium ions. Factor Xa activity is assessed by hydrolysis of a p-nitroaniline (pNA) substrate specific for factor Xa. The initial rate of p-nitroaniline release measured at 405 nM is directly proportional to Xa factor activity, and thus to FVIII activity in the sample.

[0281] The chromogenic assay is recommended by the Subcommittee on Factor VIII and IX of the Scientific and Standardization Committee (SSC) of the International Society on Thrombosis and Haemostasis (ISTH). Since 1994, the chromogenic assay has been the reference method for the potency of FVIII concentrates in the European Pharmacopoeia monograph. Thus, in one embodiment, a chimeric polypeptide comprising FVIII has a FVIII activity comparable to a chimeric polypeptide comprising mature FVIII or BDD FVIII (e.g., or ).

[0282] In another embodiment, a chimeric protein comprising FVIII of the present disclosure has a rate of factor Xa generation comparable to a chimeric protein comprising mature FVIII or BDD FVIII (e.g., or ).

[0283] To activate factor X to factor Xa, activated factor IX (factor IXa) hydrolyzes an arginine-isoleucine bond in factor X to form factor Xa in the presence of Ca 2+ , membrane phospholipids, and FVIII cofactors. Thus, the interaction of FVIII with factor IX is critical in the coagulation pathway. In certain embodiments, a chimeric polypeptide comprising FVIII can interact with factor IXa at a rate comparable to a chimeric polypeptide comprising mature FVIII sequence or BDD FVIII (e.g., or ).

[0284] In addition, FVIII binds to von Willebrand Factor (VWF), but is inactive in the circulation. When not bound to VWF, FVIII is rapidly degraded and released from VWF by the action of thrombin. In some embodiments, a chimeric polypeptide comprising FVIII binds to von Willebrand Factor at a rate comparable to a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ).

[0285] FVIII can be inactivated by activated protein C in the presence of calcium and phospholipids. Activated protein C cleaves the FVIII heavy chain after arginine 336 in the Al domain, which disrupts the factor X substrate interaction site, and after arginine 562 in the A2 domain, which enhances dissociation of the A2 domain and disrupts the interaction site with IXa factor. This cleavage also bisects the A2 domain (43 kDa) and generates A2-N (18 kDa) and A2-C (25 kDa) domains. Thus, activated protein C can catalyze multiple cleavage sites in the heavy chain. In one embodiment, a chimeric polypeptide comprising FVIII is inactivated by activated protein C at a level comparable to a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ).

[0286] In other embodiments, a chimeric polypeptide comprising FVIII has an in vivo activity of FVIII comparable to a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ). In a particular embodiment, a chimeric polypeptide comprising FVIII is capable of protecting a HemA mouse in a HemA mouse tail transection model at a level comparable to a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ).

[0287] As used herein, the "B domain" of FVIII is the same as the B domain known in the art, defined by internal amino acid sequence identity and a proteolytic cleavage site for thrombin, e.g., residues Ser741-Arg1648 of mature human FVIII. Other human FVIII domains are defined by the following amino acid residues, relative to mature human FVIII: Al, residues Ala1-Arg372 of mature FVIII; A2, residues Ser373-Arg740; A3, residues Ser1690-Ile2032; Cl, residues Arg2033-Asn2172; C2, residues Ser2173-Tyr2332. Unless otherwise indicated, sequence residue numbering used herein corresponds to FVIII sequence without signal peptide sequence (19 amino acids). The A3-Cl-C2 sequence, also referred to as the FVIII heavy chain, includes residues Ser1690-Tyr2332. The remaining sequence, residues Glu1649-Arg1689, is commonly referred to as the FVIII light chain activation peptide. The boundary positions for all domains (including the B domain) of porcine, mouse, and canine FVIII are also known in the art. In one embodiment, the B domain of FVIII is deleted ("B domain deleted FVIII" or "BDD FVIII"). An example of BDD FVIII is (rBDD FVIII). In a particular embodiment, the B domain deleted FVIII variant comprises a deletion of amino acid residues 746 to 1648 of mature FVIII.

[0288] A "B-domain deleted FVIII" can have all or part of the deletions disclosed in U.S. Patent Nos. 6,316,226, 6,346,513, 7,041,635, 5,789,203, 6,060,447, 5,595,886, 6,228,620, 5,972,885, 6,048,720, 5,543,502, 5,610,278, 5,171,844, 5,112,950, 4,868,112, and 6,458,563, and International Patent No. WO 2015106052 Al (PCT / US2015 / 010738). In some embodiments, the B-domain deleted FVIII sequence used in the methods of the present disclosure comprises any one of the deletions disclosed in Column 4, line 4 to Column 5, line 28 of U.S. Patent No. 6,316,226, and Examples 1-5 (also disclosed in US 6,346,513). In another embodiment, the B-domain deleted Factor VIII is a S743 / Q1638 B-domain deleted Factor VIII (SQ BDD FVIII) (e.g., a Factor VIII having a deletion from amino acid 744 to amino acid 1637, e.g., a Factor VIII having amino acids 1-743 and amino acids 1638-2332 of mature FVIII). In some embodiments, the B-domain deleted FVIII used in the methods of the present disclosure has the deletions disclosed in Column 2, lines 26-52 and Examples 5-8 of U.S. Patent No. 5,789,203 (also disclosed in US 6,060,447, US 5,595,886, and US 6,228,620). In some embodiments, the B-domain deleted Factor VIII has the deletions disclosed in Column 1, line 25 to Column 2, line 40 of U.S. Patent No. 5,972,885; Column 6, lines 1-22 and Example 1 of U.S. Patent No. 6,048,720; Column 2, lines 17-46 of U.S. Patent No. 5,543,502; Column 4, lines 22 to Column 5, lines 36 of U.S. Patent No. 5,171,844; Column 2, lines 55-68, Figure 2, and Example 1 of U.S. Patent No. 5,112,950; Column 2, lines 2 to Column 19, lines 21 and Table 2 of U.S. Patent No. 4,868,112; Column 2, lines 1 to Column 3, lines 19, Column 3, lines 40 to Column 4, lines 67, Column 7, lines 43 to Column 8, lines 26, and Column 11, lines 5 to Column 13, lines 39 of U.S. Patent No. 7,041,635; or Column 4, lines 25-53 of U.S. Patent No. 6,458,563. In some embodiments, the B-domain deleted FVIII has a deletion of most of the B-domain, but still contains the amino-terminal sequence of the B-domain, which is critical for in vivo proteolytic processing of the primary translation product into two polypeptide chains, as disclosed in WO 91 / 09122.In some embodiments, the B-domain deleted FVIII construct has a deletion of amino acids 747-1638, i.e., a complete deletion of the B-domain. Hoeben R.C., et al. J. Biol. Chem. 265(13):7318-7323 (1990). B-domain deleted Factor VIII can also contain a deletion of amino acids 771-1666 or amino acids 868-1562 of FVIII. Meulien P., et al. Protein Eng. 2(4):301-6 (1988). Additional B-domain deletions that are part of the present application include: a deletion of amino acids 982 to 1562 or 760 to 1639 (Toole et al. Proc. Natl. Acad. Sci. U.S.A. (1986) 83, 5939-5942), a deletion of 797 to 1562 (Eaton, et al. Biochemistry (1986) 25:8343-8347), a deletion of 741 to 1646 (Kaufman (PCT International Publication No. WO 87 / 04187), a deletion of 747-1560 (Sarver, et al. DNA (1987) 6:553-564), a deletion of 741 to 1648 (Pasek (PCT Application No. 88 / 00831), or a deletion of 816 to 1598 or a deletion of 741 to 1648 (Lagner (Behring Inst. Mitt. (1988) No 82:16-25, EP 295597). In a particular embodiment, the B-domain deleted FVIII comprises a deletion of amino acid residues 746 to 1648 of mature FVIII. In another embodiment, the B-domain deleted FVIII comprises a deletion of amino acid residues 745 to 1648 of mature FVIII. In some embodiments, the BDD FVIII comprises a single chain FVIII that contains a deletion corresponding to amino acids 765 to 1652 of mature full-length FVIII (also known as rVIII-SNG®. See U.S. Patent No. 7,041,635. )·

[0289] In other embodiments, the BDD FVIII includes a FVIII polypeptide that contains a fragment of the B domain that retains one or more N-linked glycosylation sites, e.g., residues 757, 784, 828, 900, 963, or optionally 943, which correspond to the amino acid sequence of full-length FVIII sequence. Examples of B domain fragments include the 226 amino acids or 163 amino acids of the B domain disclosed in Miao, H. Z., et al., Blood 103(a):3412-3419 (2004), Kasuda, A, et al., J. Thromb. Haemost. 6:1352-1359 (2008), and Pipe, S. W., et al., J. Thromb. Haemost. 9:2235-2242 (2011) (i.e., retaining the first 226 amino acids or 163 amino acids of the B domain). In still other embodiments, the BDD FVIII further comprises a point mutation at residue 309 (Phe to Ser) to improve expression of the BDD FVIII protein. See Miao, H. Z., et al., Blood 103(a):3412-3419 (2004). In still other embodiments, the BDD FVIII includes a FVIII polypeptide that contains a portion of the B domain but does not contain one or more furin cleavage sites (e.g., Arg1313 and Arg1648). See Pipe, S. W., et al., J. Thromb. Haemost. 9:2235-2242 (2011). In some embodiments, the BDD FVIII comprises a single chain FVIII that contains a deletion corresponding to amino acids 765 to 1652 of mature full-length FVIII (also known as rVIII-SingleChain and ) See U.S. Patent No. 7,041,635. Each of the foregoing deletions can be made in any FVIII sequence.

[0290] As discussed above and below, many functional FVIII variants are known. In addition, hundreds of non-functional mutations of FVIII have been identified in hemophilia patients, and it has been determined that the effect of these mutations on FVIII function is more due to their location within the 3-dimensional structure of FVIII than to the nature of the substitution (Cutler et al., Hum. Mutat. 19:274-8 (2002)), which is incorporated by reference in its entirety. In addition, comparisons between FVIII from humans and other species have identified conserved residues that are likely to be required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US 6,251,632), which are incorporated by reference in their entirety.

[0291] In some embodiments, the FVIII polypeptide comprises a FVIII variant or fragment thereof, wherein the FVIII variant or fragment thereof fragment has FVIII activity. In some embodiments, the gene cassette encodes a full-length FVIII polypeptide. In other embodiments, the gene cassette encodes a B-domain deleted (BDD) FVIII polypeptide, wherein all or a portion of the B domain of FVIII is deleted. In one particular embodiment, the gene cassette encodes a polypeptide comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 106, 107, 109, 110, 111, or 112. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 17 or a fragment thereof. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 106 or a fragment thereof. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 107. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 109 or a fragment thereof. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 16. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 109.

[0292] In some embodiments, the gene cassette of the present disclosure encodes a FVIII polypeptide comprising a signal peptide or fragment thereof. In other embodiments, the gene cassette encodes a FVIII polypeptide lacking a signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.

[0293] In some embodiments, the gene cassette comprises a nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence is codon-optimized. In certain embodiments, the gene cassette comprises the nucleotide sequence of International Application No. PCT / US2017 / 015879, which is incorporated by reference herein in its entirety. In some embodiments, the gene cassette comprises a nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence is codon-optimized. In certain embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 1-14. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 71. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19.

[0294] i. a codon-optimized nucleotide sequence encoding a FVIII polypeptide

[0295] In some embodiments, the nucleic acid molecule of the present disclosure comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the first ITR and the second ITR are derived from an AAV genome, and wherein the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide. In some embodiments, the codon-optimized nucleotide sequence encodes a full-length FVIII polypeptide. In other embodiments, the codon-optimized nucleotide sequence encodes a B-domain deleted (BDD) FVIII polypeptide, wherein all or a portion of the B domain of FVIII is deleted. In one particular embodiment, the codon-optimized nucleotide sequence encodes a polypeptide comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 17, or a fragment thereof. In one embodiment, the codon-optimized nucleotide sequence encodes a polypeptide having the amino acid sequence of SEQ ID NO: 17, or a fragment thereof.

[0296] In some embodiments, the codon-optimized nucleotide sequence encodes a FVIII polypeptide comprising a signal peptide, or a fragment thereof. In other embodiments, the codon-optimized sequence encodes a FVIII polypeptide lacking a signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.

[0297] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3 or (ii) nucleotides 58-1791 of SEQ ID NO: 4; and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In a particular embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 3. In another embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 4. In other embodiments, the first nucleotide sequence comprises nucleotides 58-1791 of SEQ ID NO: 3 or nucleotides 58-1791 of SEQ ID NO: 4.

[0298] In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1-1791 of SEQ ID NO: 3 or (ii) nucleotides 1-1791 of SEQ ID NO: 4; and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1-1791 of SEQ ID NO: 3 or nucleotides 1-1791 of SEQ ID NO: 4. In another embodiment, the second nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 3 or nucleotides 1792-4374 of SEQ ID NO: 4. In a particular embodiment, the second nucleotide sequence comprises nucleotides 1792-4374 of SEQ ID NO: 3 or nucleotides 1792-4374 of SEQ ID NO: 4. In yet another embodiment, the second nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 3 or nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 1792-4374 of SEQ ID NO: 3 or nucleotides 1792-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment). In a particular embodiment, the second nucleotide sequence comprises nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 3 or nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 1792-4374 of SEQ ID NO: 3 or nucleotides 1792-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment).

[0299] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5 or (ii) 1792-4374 of SEQ ID NO: 6; and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 5. In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 6. In a particular embodiment, the second nucleic acid sequence comprises nucleotides 1792-4374 of SEQ ID NO: 5 or 1792-4374 of SEQ ID NO: 6. In some embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 5 or nucleotides 58-1791 of SEQ ID NO: 6. In other embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1-1791 of SEQ ID NO: 5 or nucleotides 1-1791 of SEQ ID NO: 6.

[0300] In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5, which is devoid of nucleotides encoding the B domain or B domain fragment) or (ii) 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6, which is devoid of nucleotides encoding the B domain or B domain fragment); and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5, which is devoid of nucleotides encoding the B domain or B domain fragment). In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6, which is devoid of nucleotides encoding the B domain or B domain fragment). In a particular embodiment, the second nucleic acid sequence comprises nucleotides 1792-2277 of SEQ ID NO: 5 and 2320-4374 of SEQ ID NO: 6 or 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 or nucleotides 1792-4374 of SEQ ID NO: 6, which is devoid of nucleotides encoding the B domain or B domain fragment). In some embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 5 or nucleotides 58-1791 of SEQ ID NO: 6.In other embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, less 98%, or at least 99% sequence identity to nucleotides 1-1791 of SEQ ID NO: 5 or nucleotides 1-1791 of SEQ ID NO: 6.

[0301] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 1, (ii) nucleotides 58-1791 of SEQ ID NO: 2, (iii) nucleotides 58-1791 of SEQ ID NO: 70, or (iv) nucleotides 58-1791 of SEQ ID NO: 71; and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In other embodiments, the first nucleotide sequence comprises nucleotides 58-1791 of SEQ ID NO: 1, nucleotides 58-1791 of SEQ ID NO: 2, (iii) nucleotides 58-1791 of SEQ ID NO: 70, or (iv) nucleotides 58-1791 of SEQ ID NO: 71.

[0302] In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1-1791 of SEQ ID NO: 1, (ii) nucleotides 1-1791 of SEQ ID NO: 2, (iii) nucleotides 1-1791 of SEQ ID NO: 70, or (iv) nucleotides 1-1791 of SEQ ID NO: 71; and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1-1791 of SEQ ID NO: 1, nucleotides 1-1791 of SEQ ID NO: 2, (iii) nucleotides 1-1791 of SEQ ID NO: 70, or (iv) nucleotides 1-1791 of SEQ ID NO: 71. In another embodiment, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 1, nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In a particular embodiment, the second nucleotide sequence linked to the first nucleotide sequence comprises (i) nucleotides 1792-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In other embodiments, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71.In one embodiment, the second nucleotide sequence comprises (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71.

[0303] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71; and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In a particular embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In some embodiments, the codon-optimized sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1792-4374 of SEQ ID NO: 1, nucleotides 1792-4374 of SEQ ID NO: 2, nucleotides 1792-4374 of SEQ ID NO: 70, or nucleotides 1792-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity.In one embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1792-4374 of SEQ ID NO: 1, nucleotides 1792-4374 of SEQ ID NO: 2, nucleotides 1792-4374 of SEQ ID NO: 70, or nucleotides 1792-4374 of SEQ ID NO: 71, which are devoid of nucleotides encoding a B domain or a B domain fragment).

[0304] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 1. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, which are devoid of nucleotides encoding a B domain or a B domain fragment). In other embodiments, the nucleic acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, which are devoid of nucleotides encoding a B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 1. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 1-4374 of SEQ ID NO: 1, which are devoid of nucleotides encoding a B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 1.

[0305] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 2. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 2. In other embodiments, the nucleic acid sequence has at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 2. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 2 (i.e., nucleotides 58-4374 of SEQ ID NO: 2 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 2. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 2 (i.e., nucleotides 1-4374 of SEQ ID NO: 2 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 2.

[0306] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 70. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 58-4374 of SEQ ID NO: 70 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 70. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 58-4374 of SEQ ID NO: 70 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 70. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 1-4374 of SEQ ID NO: 70 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 70.

[0307] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 71. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 71. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 71. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 71.

[0308] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 3. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 58-4374 of SEQ ID NO: 3 without the nucleotides encoding the B domain or B domain fragment). In certain embodiments, the nucleic acid sequence has at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 3. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 58-4374 of SEQ ID NO: 3 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 3. In yet other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 1-4374 of SEQ ID NO: 3 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 3.

[0309] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 4. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 58-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 4. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 58-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 4. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 1-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 4.

[0310] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 5. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 58-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 5. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 58-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 5. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 5.

[0311] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 6. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 58-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 6. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 58-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO: 6. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO: 6.

[0312] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleic acid sequence encoding a signal peptide. In certain embodiments, the signal peptide is a FVIII signal peptide. In some embodiments, the nucleic acid sequence encoding the signal peptide is codon-optimized. In one particular embodiment, the nucleic acid sequence encoding the signal peptide has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to (i) nucleotides 1 to 57 of SEQ ID NO: 1; (ii) nucleotides 1 to 57 of SEQ ID NO: 2; (iii) nucleotides 1 to 57 of SEQ ID NO: 3; (iv) nucleotides 1 to 57 of SEQ ID NO: 4; (v) nucleotides 1 to 57 of SEQ ID NO: 5; (vi) nucleotides 1 to 57 of SEQ ID NO: 6; (vii) nucleotides 1 to 57 of SEQ ID NO: 70; (viii) nucleotides 1 to 57 of SEQ ID NO: 71; or (ix) nucleotides 1 to 57 of SEQ ID NO: 68.

[0313] SEQ ID NOs: 1-6, 70, and 71 are optimized versions of SEQ ID NO: 16, which is a starting or “parental” or “wild-type” FVIII nucleotide sequence. SEQ ID NO: 16 encodes a human FVIII with a B-domain deletion. While SEQ ID NOs: 1-6, 70, and 71 are derived from a specific B-domain deleted version of FVIII (SEQ ID NO: 16), it is understood that the disclosure also includes optimized versions of nucleic acids encoding other versions of FVIII. For example, other versions of FVIII can include full-length FVIII, other B-domain deleted FVIII (described herein), or other fragments of FVIII that retain FVIII activity.

[0314] In one embodiment, the gene cassette comprises a FVIII construct comprising the polynucleotide sequence set forth in Table 2A-2C (6526 nucleotides). In one embodiment, the gene cassette comprises a FVIII construct comprising the polynucleotide sequence set forth in Table 2B (6526 nucleotides).

[0315] In certain embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 179 or 182, wherein the isolated nucleic acid molecule retains the ability to express a functional FVIII protein.

[0316] Table 2A: Example FVIII constructs (nucleotides 1-6526; SEQ ID NO: 110)

[0317]

[0318]

[0319]

[0320]

[0321] Table 2B: Example B19-FVIII construct (nucleotides 1-6762; SEQ ID NO: 179)

[0322]

[0323]

[0324]

[0325]

[0326]

[0327]

[0328] Table 2C: Example GPV-FVIII construct (nucleotides 1-6830; SEQ ID NO: 182)

[0329]

[0330]

[0331]

[0332]

[0333]

[0334]

[0335]

[0336] A. Codon adaptation index

[0337] In one embodiment, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the human codon adaptation index of the codon-optimized nucleotide sequence is increased relative to SEQ ID NO: 16. For example, the codon-optimized nucleotide sequence can have a human codon adaptation index that is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%), at least about 0.91 (91%), at least about 0.92 (92%), at least about 0.93 (93%), at least about 0.94 (94%), at least about 0.95 (95%), at least about 0.96 (96%), at least about 0.97 (97%), at least about 0.98 (98%), or at least about 0.99 (99%). In some embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index that is at least about.88 (88%). In other embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index that is at least about.91 (91%). In other embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index that is at least about.91 (97%).

[0338] In a particular embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO: 16. In some embodiments, the human codon adaptation index of the nucleotide sequence is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%), or at least about.91 (91%). In a particular embodiment, the human codon adaptation index of the nucleotide sequence is at least about.88 (88%). In another embodiment, the human codon adaptation index of the nucleotide sequence is at least about.91 (91%).

[0339] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 or (ii) 1792-2277 and 2320-4374 of SEQ ID NO: 6; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO: 16. In some embodiments, the human codon adaptation index of the nucleotide sequence is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), or at least about 0.88 (88%). In a particular embodiment, the human codon adaptation index of the nucleotide sequence is at least about.83 (83%). In another embodiment, the human codon adaptation index of the nucleotide sequence is at least about.88 (88%).

[0340] In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71, without the nucleotides encoding the B domain or B domain fragment); and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO: 16. In some embodiments, the human codon adaptation index of the nucleotide sequence is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), or at least about 0.88 (88%). In a particular embodiment, the human codon adaptation index of the nucleotide sequence is at least about.75 (75%). In another embodiment, the human codon adaptation index of the nucleotide sequence is at least about.83 (83%). In another embodiment, the human codon adaptation index of the nucleotide sequence is at least about.88 (88%). In another embodiment, the human codon adaptation index of the nucleotide sequence is at least about.91 (91%). In another embodiment, the human codon adaptation index of the nucleotide sequence is at least about.97 (97%).

[0341] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide of the present disclosure has an increased frequency of optimal codons (FOP) relative to SEQ ID NO: 16. In certain embodiments, the FOP of the codon-optimized nucleotide sequence encoding a FVIII polypeptide is at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 64, at least about 65, at least about 70, at least about 75, at least about 79, at least about 80, at least about 85, or at least about 90.

[0342] In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide of the present disclosure has an increased relative synonymous codon usage (RCSU) relative to SEQ ID NO: 16. In some embodiments, the isolated nucleic acid molecule has an RCSU of greater than 1.5. In other embodiments, the isolated nucleic acid molecule has an RCSU of greater than 2.0. In certain embodiments, the isolated nucleic acid molecule has an RCSU of at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2.0, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, or at least about 2.7.

[0343] In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide of the present disclosure has a decreased number of effective codons relative to SEQ ID NO: 16. In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a number of effective codons of less than about 50, less than about 45, less than about 40, less than about 35, less than about 30, or less than about 25. In a particular embodiment, the isolated nucleic acid molecule has a number of effective codons of about 40, about 35, about 30, about 25, or about 20.

[0344] B. G / C content optimization

[0345] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, or at least about 60%.

[0346] In a particular embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, or at least about 58%. In a particular embodiment, the nucleotide sequence encoding a polypeptide having FVIII activity has a G / C content of at least about 58%.

[0347] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment), or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, or at least about 57%. In a particular embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 52%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 55%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 57%.

[0348] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 or (ii) nucleotides 58-2277 and 2320-4374 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment) of an amino acid sequence selected from the group consisting of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71; and wherein the nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 45%. In a particular embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 52%. In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 55%. In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 57%. In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 58%. In yet other embodiments, the n codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content of at least about 60%.

[0349] The "G / C content" (or guanine-cytosine content), or "percentage of G / C nucleotides," refers to the percentage of nitrogenous bases (i.e., guanine or cytosine) in a DNA molecule. The G / C content can be calculated using the following formula:

[0350]

[0351] The G / C content of human genes is highly heterogeneous, with some genes having a G / C content as low as 20% and other genes having a G / C content as high as 95%. Generally, G / C-rich genes are expressed more highly. Indeed, it has been demonstrated that increasing the G / C content of a gene can result in increased expression of the gene, primarily due to increased transcription and higher steady-state mRNA levels. See Kudla et al., PLoS Biol., 4(6):el80 (2006).

[0352] C. Matrix Attachment Region-Like Sequences

[0353] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a MARS / ARS sequence.

[0354] In one particular embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a MARS / ARS sequence.

[0355] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains no MARS / ARS sequences.

[0356] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71, or (ii) nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains no MARS / ARS sequences.

[0357] An AT-rich element in the human FVIII nucleotide sequence has been identified that shares sequence similarity with the autonomously replicating sequence (ARS) and nuclear matrix attachment regions (MARs) of S. cerevisiae. (Fallux et al., Mol. Cell. Biol. 16:4264-4272 (1996). One of these elements has been shown to bind nuclear factors in vitro and repress expression of a chloramphenicol acetyltransferase (CAT) reporter gene. (Fallux et al., Mol. Cell. Biol. 16:4264-4272 (1996). It has been hypothesized that these sequences can contribute to transcriptional repression of the human FVIII gene. Thus, in one embodiment, all MAR / ARS sequences are abrogated in the codon-optimized nucleotide sequence encoding a FVIII polypeptide of the present disclosure. There are four MAR / ARS ATATTT sequences (SEQ ID NO: 21) and three MAR / ARS AAATAT sequences (SEQ ID NO: 22) in the parental FVIII sequence (SEQ ID NO: 16). All of these sites were mutated to disrupt the MAR / ARS sequences in the optimized FVIII sequences (SEQ ID NO: 1-6). The location of each of these elements, as well as the sequence of the corresponding nucleotides in the optimized sequences, are shown in Table 3 below.

[0358] Table 3: Summary of destabilizing element changes

[0359]

[0360]

[0361] D. Destabilizing sequences

[0362] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a destabilizing element.

[0363] In one particular embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains at most 9, at most 8, at most 7, at most 6, or at most 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a destabilizing element.

[0364] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the nucleotide sequence encoding the polypeptide having FVIII activity contains at most 9, at most 8, at most 7, at most 6, or at most 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a destabilizing element.

[0365] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains no destabilizing elements.

[0366] There are 10 destabilizing elements in the parental FVIII sequence (SEQ ID NO: 16): 6 ATTTA sequences (SEQ ID NO: 23) and 4 TAAAT sequences (SEQ ID NO: 24). In one embodiment, the sequences at these sites are mutated to disrupt the destabilizing elements in the optimized FVIII SEQ ID NOs: 1-6, 70, and 71. The location of each of these elements, as well as the sequence of the corresponding nucleotides in the optimized sequences, is shown in Table 3.

[0367] E. Potential Promoter Binding Sites

[0368] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding site. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a potential promoter binding site.

[0369] In one particular embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding site. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a potential promoter binding site.

[0370] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding site. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a potential promoter binding site.

[0371] In other embodiments, the gene cassette encoding the FVIII polypeptide comprises a codon-optimized nucleotide sequence, wherein the nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 that do not encode the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding site. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a potential promoter binding site.

[0372] TATA boxes are regulatory sequences commonly found in the promoter region of eukaryotic cells. They serve as binding sites for TATA binding protein (TBP), a general transcription factor. TATA boxes typically comprise the sequence TATAA (SEQ ID NO: 28) or close variants. However, TATA boxes within coding sequences can inhibit translation of the full-length protein. There are 10 potential promoter binding sequences in the wild-type BDD FVIII sequence (SEQ ID NO: 16): 5 TATAA sequences (SEQ ID NO: 28) and 5 TTATA sequences (SEQ ID NO: 29). In some embodiments, at least 1, at least 2, at least 3, or at least 4 of the promoter binding sites are abolished in the FVIII genes of the present disclosure. In some embodiments, at least 5 of the promoter binding sites are abolished in the FVIII genes of the present disclosure. In other embodiments, at least 6, at least 7, or at least 8 of the promoter binding sites are abolished in the FVIII genes of the present disclosure. In one embodiment, at least 9 of the promoter binding sites are abolished in the FVIII genes of the present disclosure. In one particular embodiment, all of the promoter binding sites are abolished in the FVIII genes of the present disclosure. The location of each potential promoter binding site and the sequence of the corresponding nucleotides in the optimized sequence are shown in Table 3.

[0373] F. Other cis-acting negative regulatory elements

[0374] In addition to the MAR / ARS sequences, destabilizing elements, and potential promoter sites described above, several other potential inhibitory sequences can be identified in the wild-type BDD FVIII sequence (SEQ ID NO: 16). Two AU-rich sequence elements (AREs) can be identified in the non-optimized BDD FVIII sequence (ATTTTATT (SEQ ID NO: 30); and ATTTTTAA (SEQ ID NO: 31), along with a polyA site (AAAAAAA; SEQ ID NO: 26), a polyT site (TTTTTT; SEQ ID NO: 25), and a splice site (GGTGAT; SEQ ID NO: 27). One or more of these elements can be deleted from the optimized FVIII sequence. The location of each of these sites and the sequence of the corresponding nucleotides in the optimized sequence are shown in Table 3.

[0375] In certain embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, e.g., a splice site, a poly-T sequence, a poly-A sequence, an ARE sequence, or any combination thereof.

[0376] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 5; (ii) nucleotides 1792-4374 of SEQ ID NO: 6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 5 (i.e., nucleotides 1792-4374 of SEQ ID NO: 5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment); wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, e.g., a splice site, a poly-T sequence, a poly-A sequence, an ARE sequence, or any combination thereof.

[0377] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or B domain fragment); and wherein the codon-optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, e.g., a splice site, a poly-T sequence, a poly-A sequence, an ARE sequence, or any combination thereof.

[0378] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain the splice site GGTGAT (SEQ ID NO: 27). In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain a poly-T sequence (SEQ ID NO: 25).In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain a polyA sequence (SEQ ID NO: 26). In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO: 3; (ii) nucleotides 1-1791 of SEQ ID NO: 3; (iii) nucleotides 58-1791 of SEQ ID NO: 4; or (iv) nucleotides 1-1791 of SEQ ID NO: 4; and wherein the codon-optimized nucleotide sequence does not contain an ARE element (SEQ ID NO: 30 or SEQ ID NO: 31).

[0379] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence does not contain the splice site GGTGAT (SEQ ID NO: 27). In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 without the nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence does not have a poly-T sequence (SEQ ID NO: 25).In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 that do not have nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence does not contain a polyA sequence (SEQ ID NO: 26). In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 that do not have nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence does not contain an ARE element (SEQ ID NO: 30 or SEQ ID NO: 31).

[0380] In other embodiments, the optimized FVIII sequences of the present disclosure do not comprise one or more of an antiviral motif, a stem loop structure, and a repeat sequence.

[0381] In yet other embodiments, the nucleotides around the transcription start site are changed to a kozak consensus sequence (GCCGCCACC ATG C (SEQ ID NO: 32), where the underlined nucleotides are the start codon). In other embodiments, restriction sites can be added or removed to facilitate cloning processes.

[0382] b. FIX and polynucleotide sequences encoding FIX proteins

[0383] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a FIX polypeptide. In some embodiments, the FIX polypeptide comprises FIX or a variant or fragment thereof, wherein the FIX or variant or fragment thereof has FIX activity.

[0384] Human FIX is a serine protease that is an important component of the intrinsic pathway of the coagulation cascade. As used herein, “Factor IX” or “FIX” refers to the coagulation factor protein and its species and sequence variants, and includes, but is not limited to, the 461 single chain amino acid sequence of the human FIX precursor polypeptide (“prepro”), the 415 single chain amino acid sequence of mature human FIX (SEQ ID NO: 125), and the R338L FIX (Padua) variant (SEQ ID NO: 126). FIX includes any form of the FIX molecule having the typical characteristics of coagulation FIX. As used herein, “Factor IX” and “FIX” are intended to encompass a polypeptide comprising the domain Gla (region containing gamma carboxyglutamic acid residues), EGF1 and EGF2 (regions containing sequences homologous to human epidermal growth factor), the activation peptide (“AP”, which is formed from residues R136-R180 of mature FIX), and the C-terminal protease domain (“Pro”), or synonyms for these domains known in the art, or can be a truncated fragment or sequence variant that retains at least a portion of the biological activity of the native protein. FIX or sequence variants have been cloned, as described in U.S. Patent Nos. 4,770,999 and 7,700,734, and the cDNA encoding human FIX has been isolated, characterized, and cloned into expression vectors (see, e.g., Choo et al., Nature 299:178-180 (1982); Fair et al., Blood 64:194-204 (1984); and Kurachi et al., Proc. Natl. Acad. Sci., U.S.A. 79:6461-6464 (1982)). One particular variant of FIX, the R338L FIX (Padua) variant (SEQ ID NO: 2) (which is characterized by Simioni et al., 2009) comprises a gain-of-function mutation that is associated with an approximately 8-fold increase in activity of the Padua variant relative to native FIX (Table 4). FIX variants can also include any FIX polypeptide having one or more conservative amino acid substitutions that do not affect the FIX activity of the FIX polypeptide. In some embodiments, the FIX variant comprises rFIX-albumin fused by a cleavable linker, e.g., See US 7,939,632, which is incorporated by reference herein in its entirety.

[0385] Table 4: Example FIX Sequences

[0386]

[0387]

[0388]

[0389]

[0390] Gray shading = signal peptide; underlined = XTEN sequence; bold = Fc.

[0391] SEQ ID NO: 67 in U.S. Patent No. 9,856,468, which is incorporated by reference in its entirety.

[0392] FIX polypeptide is 55 kDa, which is synthesized as a prepropolypetide chain consisting of three regions: a 28 amino acid signal peptide (amino acids 1 to 28 of SEQ ID NO: 127), an 18 amino acid propeptide (amino acids 29 to 46) that requires gamma carboxylation of glutamic acid residues, and a 415 amino acid mature Factor IX (SEQ ID NO: 125 or 126). The propeptide is an 18 amino acid residue sequence at the N-terminus of the gamma-carboxyglutamate domain. This propeptide binds the vitamin K-dependent gamma carboxylase, and then is cleaved from the prepolypeptide of FIX by an endogenous protease, most likely PACE (paired basic amino acid cleavage enzyme), also known as furin or PCSK3. Without gamma carboxylation, the Gla domain cannot bind calcium, and cannot adopt the correct conformation necessary to anchor the protein to the negatively charged phospholipid surface, rendering Factor IX nonfunctional. Even if the Gla domain is carboxylated, it relies on cleavage of the propeptide for proper function, because the retained propeptide interferes with the conformational changes of the Gla domain necessary for optimal binding to calcium and phospholipids. In humans, the resulting mature Factor IX is secreted by hepatocytes into the bloodstream in an inactive zymogen form that is a 415 amino acid residue single chain protein containing approximately 17% carbohydrate by weight (Schmidt, A.E., et al. (2003) Trends Cardiovasc Med, 13:39).

[0393] Mature FIX consists of several domains in the N-terminal to C-terminal configuration: a GLA domain, an EGF1 domain, an EGF2 domain, an activation peptide (AP) domain, and a protease domain (or catalytic domain). A short linker joins the EGF2 domain to the AP domain. FIX contains two activation peptides formed by R145-A146 and R180-V181, respectively. Upon activation, single-chain FIX becomes a 2-chain molecule, with the two chains linked by a disulfide bond. Coagulation factors can be engineered by replacing their activation peptides, resulting in altered activation specificity. In mammals, mature FIX must be activated by activated factor XI to produce factor IXa. Upon activation of FIX to FIXa, the protease domain provides the catalytic activity of FIX. Activated factor VIII (FVIIIa) is a fully expressed specific cofactor of FIXa activity.

[0394] In certain embodiments, the FIX polypeptide comprises a Thr148 allelic form of plasma-derived FIX and has structural and functional characteristics similar to endogenous FIX.

[0395] A number of functional FIX variants are known in the art. International Publication No. WO 02 / 040544 A3 discloses mutants that exhibit increased resistance to inhibition by heparin on page 4, lines 9-30 and page 15, lines 6-31. International Publication No. WO 03 / 020764 A2 discloses FIX mutants with reduced T-cell immunogenicity in Tables 2 and 3 (pages 14-24) and page 12, lines 1-27. International Publication No. WO 2007 / 149406 A2 discloses functional mutant FIX molecules that exhibit increased protein stability, increased in vivo and in vitro half-life, and increased resistance to proteases on page 4, line 1 to page 19, line 11. WO 2007 / 149406 A2 further discloses chimeric and other variant FIX molecules on page 19, line 12 to page 20, line 9. International Publication No. WO 08 / 118507 A2 discloses FIX mutants that exhibit increased coagulation activity on page 5, line 14 to page 6, line 5. International Publication No. WO 09 / 051717 A2 discloses FIX mutants with an increased number of N-linked and / or O-linked glycosylation sites that result in increased half-life and / or recovery on page 9, line 11 to page 20, line 2. International Publication No. WO 09 / 137254 A2 further discloses Factor IX mutants with an increased number of glycosylation sites on page 2, paragraph

[006] to page 5, paragraph

[011] and page 16, paragraph

[044] to page 24, paragraph

[057] . International Publication No. WO 09 / 130198 A2 discloses functional mutant FIX molecules with an increased number of glycosylation sites that result in increased half-life on page 4, line 26 to page 12, line 6. International Publication No. WO 09 / 140015 A2 discloses functional FIX mutants with an increased number of Cys residues that can be used for polymerase (e.g., PEG) conjugation on page 11, paragraph

[0043] to page 13, paragraph

[0053] . The FIX polypeptides described in International Application No. PCT / US2011 / 043569, filed July 11, 2011 and published as WO 2012 / 006624 on January 12, 2012, are also incorporated by reference in their entirety. In some embodiments, the FIX polypeptide comprises a FIX polypeptide fused to albumin, e.g., FIX-albumin. In certain embodiments, the FIX polypeptide is or rIX-FP.

[0396] In addition, hundreds of non-functional mutations in FIX have been identified in hemophilia subjects, many of which are disclosed in Table 6 on pages 11-14 of International Publication No. WO 09 / 137254 A2. Such non-functional mutations are not included in the present application, but provide additional guidance as to which mutations are more or less likely to result in a functional FIX polypeptide.

[0397] In one embodiment, the FIX polypeptide (or Factor IX portion of the fusion polypeptide) comprises an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 1 or 2 (amino acids 1 to 415 of SEQ ID NO: 125 or 126), or alternatively, to a sequence having a propeptide sequence, or a sequence having a propeptide and signal sequence (full-length FIX). In another embodiment, the FIX polypeptide comprises an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 2.

[0398] FIX coagulant activity is expressed in International Units (IU). One IU of FIX activity corresponds approximately to the amount of FIX in one milliliter of normal human plasma. Several assays can be used to measure FIX activity, including one-stage clotting assays (activated partial thromboplastin time; aPTT), thrombin generation time (TGA), and rotational thrombelastometry (ROTEM) The present application contemplates sequences having homology to FIX sequences, natural sequence fragments (e.g., from humans, non-human primates, mammals (including domesticated animals)), and non-natural sequence variants that retain at least a portion of the biological activity or biological function of FIX and / or are useful in preventing, treating, mediating, or ameliorating a disease, deficiency, disorder, or condition associated with a clotting factor (e.g., a bleeding episode associated with trauma, surgery, a clotting factor deficiency). Sequences having homology to human FIX can be found by standard homology searching techniques (e.g., NCBIBLAST).

[0399] In certain embodiments, the FIX sequence is codon-optimized. Examples of codon-optimized FIX sequences include, but are not limited to, SEQ ID NOs: 1 and 54-58 in International Publication No. WO 2016 / 004113 Al, which is incorporated by reference herein in its entirety.

[0400] c. FVII and polynucleotide sequences encoding FVII proteins

[0401] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a Factor VII polypeptide. In some embodiments, the Factor VII polypeptide comprises FVII or a variant or fragment thereof, wherein the variant or fragment thereof has FVII activity.

[0402] “Factor VII” (“FVII” or “F7”; also known as Factor 7, Coagulation Factor VII, Serum Factor VII, Serum Thromboplastin Co-Factor, SPCA, Prethrombin, and eptacog alfa) is a serine protease that is part of the coagulation cascade. In one embodiment, the coagulation factor in the nucleic acid described herein is FVII. Recombinant activated Factor VII (“FVII”) has been used extensively to treat major bleeding, such as in patients with hemophilia A or B, Factor XI, FVII deficiency, platelet function defects, thrombocytopenia, or von Willebrand disease.

[0403] Recombinant activated FVII (rFVIIa; is used to treat bleeding episodes in (i) hemophilia patients with neutralizing antibodies (inhibitors) against FVIII or FIX, (ii) patients with FVII deficiency, or (iii) hemophilia A or B patients with inhibitors undergoing surgical treatment. However, showed poor efficacy. Due to low affinity for activated platelets, short half-life, and poor enzyme activity in the absence of tissue factor, high concentrations of repeated doses of FVIIa are often required to control bleeding. Thus, the medical need for better treatment and prophylactic options for hemophilia patients with FVIII and FIX inhibitors and / or with FVII deficiency has not been met.

[0404] In one embodiment, the gene cassette encodes a mature form of FVII or a variant thereof. FVII includes a Gla domain, two EGF domains (EGF-1 and EGF-2), and a serine protease domain (or peptidase S1 domain) that is highly conserved among all members of the peptidase S1 family of serine proteases (such as, for example, chymotrypsin). FVII occurs as a single-chain zymogen (i.e., activatable FVII) and as a fully activated two-chain form.

[0405] C. Growth Factors

[0406] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a growth factor. The growth factor can be selected from any growth factor known in the art. In some embodiments, the growth factor is a hormone. In other embodiments, the growth factor is a cytokine. In some embodiments, the growth factor is a chemokine.

[0407] In some embodiments, the growth factor is adrenomedullin (AM). In some embodiments, the growth factor is angiogenin (Ang). In some embodiments, the growth factor is autocrine motility factor. In some embodiments, the growth factor is bone morphogenetic protein (BMP). In some embodiments, the BMP is selected from the group consisting of BMP2, BMP4, BMP5, and BMP7. In some embodiments, the growth factor is a member of the ciliary neurotrophic factor family. In some embodiments, the member of the ciliary neurotrophic factor family is selected from the group consisting of ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6). In some embodiments, the growth factor is a colony stimulating factor. In some embodiments, the colony stimulating factor is selected from the group consisting of macrophage colony stimulating factor (m-CSF), granulocyte colony stimulating factor (G-CSF), and granulocyte macrophage colony stimulating factor (GM-CSF). In some embodiments, the growth factor is epidermal growth factor (EGF). In some embodiments, the growth factor is ephrin. In some embodiments, the ephrin is selected from the group consisting of ephrin Al, ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin Bl, ephrin B2, and ephrin B3. In some embodiments, the growth factor is erythropoietin (EPO). In some embodiments, the growth factor is fibroblast growth factor (FGF). In some embodiments, the FGF is selected from the group consisting of FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, and FGF23. In some embodiments, the growth factor is fetal bovine somatotropin (FBS). In some embodiments, the growth factor is a member of the GDNF family. In some embodiments, the member of the GDNF family is selected from the group consisting of glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, and Artemin. In some embodiments, the growth factor is growth differentiation factor-9 (GDF9). In some embodiments, the growth factor is hepatocyte growth factor (HGF). In some embodiments, the growth factor is hepatoma-derived growth factor (HDGF). In some embodiments, the growth factor is insulin. In some embodiments, the growth factor is an insulin-like growth factor. In some embodiments, the insulin-like growth factor is insulin-like growth factor-1 (IGF-1) or IGF-2. In some embodiments, the growth factor is an interleukin (IL). In some embodiments, the IL is selected from the group consisting of IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, and IL-7.In some embodiments, the growth factor is keratinocyte growth factor (KGF). In some embodiments, the growth factor is migration-stimulating factor (MSF). In some embodiments, the growth factor is macrophage stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)). In some embodiments, the growth factor is myostatin (GDF-8). In some embodiments, the growth factor is neuregulin. In some embodiments, the neuregulin is selected from neuregulin 1 (NRG1), NRG2, NRG3, and NRG4. In some embodiments, the growth factor is a neurotrophin. In some embodiments, the growth factor is brain-derived neurotrophic factor (BDNF). In some embodiments, the growth factor is nerve growth factor (NGF). In some embodiments, the NGF is neurotrophin 3 (NT-3) or NT-4. In some embodiments, the growth factor is placental growth factor (PGF). In some embodiments, the growth factor is platelet-derived growth factor (PDGF). In some embodiments, the growth factor is Renalase (RNLS). In some embodiments, the growth factor is T-cell growth factor (TCGF). In some embodiments, the growth factor is thrombopoietin (TPO). In some embodiments, the growth factor is a transforming growth factor. In some embodiments, the transforming growth factor is transforming growth factor alpha (TGF-a) or TGF-b. In some embodiments, the growth factor is tumor necrosis factor-alpha (TNF-a). In some embodiments, the growth factor is vascular endothelial growth factor (VEGF).

[0408] D. MicroRNAs (miRNAs)

[0409] MicroRNAs (miRNAs) are small non-coding RNA molecules (approximately 18-22 nucleotides) that negatively regulate gene expression by inhibiting translation or inducing messenger RNA (mRNA) degradation. Since their discovery, miRNAs have been implicated in a variety of cellular processes, including apoptosis, differentiation, and cell proliferation, and they have been shown to play a critical role in oncogenic processes. The ability of miRNAs to regulate gene expression makes in vivo expression of miRNAs a valuable tool in gene therapy.

[0410] Certain aspects of the present disclosure relate to a plasmid-like nucleic acid molecule comprising a first ITR, a second ITR, and a gene cassette encoding a miRNA, wherein the first ITR and / or the second ITR is a non-adenovirus-associated virus ITR (e.g., the first ITR and / or the second ITR is from a non-AAV). The miRNA can be any miRNA known in the art. In some embodiments, the miRNA downregulates expression of a target gene. In certain embodiments, the target gene is selected from SOD1, HTT, RHO, or any combination thereof.

[0411] In some embodiments, the gene cassette encodes one miRNA. In some embodiments, the gene cassette encodes more than one miRNA. In some embodiments, the gene cassette encodes two or more different miRNAs. In some embodiments, the gene cassette encodes two or more copies of the same miRNA. In some embodiments, the gene cassette encodes two or more variants of the same therapeutic protein. In certain embodiments, the gene cassette encodes one or more miRNAs and one or more therapeutic proteins.

[0412] In some embodiments, the miRNA is a naturally occurring miRNA. In some embodiments, the miRNA is an engineered miRNA. In some embodiments, the miRNA is an artificial miRNA. In certain embodiments, the miRNA comprises a miHTT engineered miRNA disclosed by Evers et al., Molecular Therapy 26(9): 1-15 (June 2018 epub ahead of print). In certain embodiments, the miRNA comprises a miR SOD1 artificial miRNA disclosed by Dirren et al., Annals of Clinical and Translational Neurology 2(2): 167-84 (February 2015). In certain embodiments, the miRNA comprises miR-708, which targets RHO (see Behrman et al., JCB 192(6): 919-27 (2011).

[0413] In some embodiments, the miRNA upregulates expression of a gene by downregulating expression of a suppressor of the gene. In some embodiments, the suppressor is native, e.g., wild type, inhibitor. In some embodiments, the suppressor is produced by a mutated, heterologous, and / or misexpressed gene.

[0414] E. Heterologous Moieties

[0415] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises at least one heterologous moiety. In some embodiments, the heterologous moiety is fused to the N- or C-terminus of the therapeutic protein. In other embodiments, the heterologous moiety is inserted between two amino acids within the therapeutic protein.

[0416] In some embodiments, the therapeutic protein comprises a FVIII polypeptide and a heterologous moiety inserted between two amino acids within the FVIII polypeptide. In some embodiments, the heterologous moiety is inserted within the FVIII polypeptide at one or more insertion sites selected from Table 5. In some embodiments, the heterologous amino acid sequence can be inserted within the coagulation factor polypeptide encoded by the nucleic acid molecules of the present disclosure at any of the sites disclosed in International Publication No. WO 2013 / 123457 Al, WO 2015 / 106052 Al, or U.S. Publication No. 2015 / 0158929 Al. In a particular embodiment, the therapeutic protein comprises FVIII and a heterologous moiety, wherein the heterologous moiety is inserted within FVIII immediately downstream of amino acid 745 relative to mature FVIII. In a particular embodiment, the therapeutic protein comprises FVIII and XTEN, wherein the XTEN is inserted within FVIII immediately downstream of amino acid 745 relative to mature FVIII. In a particular embodiment, the FVIII comprises a deletion of amino acids 746-1646 (corresponding to mature human FVIII (SEQ ID NO: 15)), and the heterologous moiety is inserted immediately downstream of amino acid 745 (corresponding to mature human FVIII (SEQ ID NO: 15)).

[0417] Table 5: FVIII heterologous moiety insertion sites

[0418]

[0419]

[0420] In some embodiments, the therapeutic protein comprises a FIX polypeptide and a heterologous moiety inserted between two amino acids within the FIX polypeptide. In some embodiments, the heterologous moiety is inserted within the FIX polypeptide at one or more insertion sites selected from Table 5. In some embodiments, the heterologous amino acid sequence can be inserted within the coagulation factor polypeptide encoded by the nucleic acid molecules of the present disclosure at any of the sites disclosed in International Application No. PCT / US2017 / 015879, which is incorporated by reference in its entirety herein. In a particular embodiment, the therapeutic protein comprises a FIX polypeptide and a heterologous moiety, wherein the heterologous moiety is inserted within the FIX polypeptide immediately downstream of amino acid 166 relative to mature FIX. In a particular embodiment, the therapeutic protein comprises a FIX polypeptide and XTEN, wherein the XTEN is inserted within FIX immediately downstream of amino acid 166 relative to mature FVIII.

[0421] Table 6: FIX heterologous moiety insertion sites

[0422]

[0423]

[0424] In other embodiments, the therapeutic protein of the present disclosure further comprises 2, 3, 4, 5, 6, 7, or 8 heterologous nucleotide sequences. In some embodiments, all of the heterologous moieties are the same. In some embodiments, at least one heterologous moiety is different from the other heterologous moieties. In some embodiments, the present disclosure can comprise 2, 3, 4, 5, 6, or 7 more heterologous moieties in a tandem format.

[0425] In some embodiments, the heterologous moiety increases the half-life of the therapeutic protein (which is a "half-life extender").

[0426] In some embodiments, the heterologous moiety is a peptide or polypeptide that has unstructured or structured features associated with half-life extension in vivo when incorporated into a protein of the present disclosure. Non-limiting examples include albumin, albumin fragments, Fc fragments of immunoglobulins, C-terminal peptide (CTP) of the beta subunit of human chorionic gonadotropin, HAP sequences, XTEN sequences, transferrin or fragments thereof, PAS polypeptides, polyglycine linkers, polyserine linkers, albumin binding moieties, or any fragments, derivatives, variants, or combinations of these polypeptides. In a particular embodiment, the heterologous amino acid sequence is an immunoglobulin constant region or portion thereof, transferrin, albumin, or a PAS sequence. In some aspects, the heterologous moiety includes von Willebrand factor or fragments thereof. In other related aspects, the heterologous moiety can include attachment sites (e.g., cysteine amino acids) for non-polypeptide moieties such as polyethylene glycol (PEG), hydroxyethyl starch (HES), polysialic acid, or derivatives, variants, or combinations of these elements. In some aspects, the heterologous moiety comprises a cysteine amino acid that serves as an attachment point for a non-polypeptide moiety such as polyethylene glycol (PEG), hydroxyethyl starch (HES), polysialic acid, or derivatives, variants, or combinations of these elements.

[0427] In a particular embodiment, the first heterologous moiety is a half-life extension molecule that is known in the art, and the second heterologous moiety is a half-life extension molecule that is known in the art. In certain embodiments, the first heterologous moiety (e.g., a first Fc moiety) and the second heterologous moiety (e.g., a second Fc moiety) associate with each other to form a dimer. In one embodiment, the second heterologous moiety is a second Fc moiety, wherein the second Fc moiety is linked or associated with the first heterologous moiety (e.g., a first Fc moiety). For example, the second heterologous moiety (e.g., a second Fc moiety) is linked to the first heterologous moiety (e.g., a first Fc moiety) by a linker, or is associated with the first heterologous moiety by a non-covalent bond.

[0428] In some embodiments, the heterologous moiety is a polypeptide comprising or consisting essentially of or consisting of at least about 10, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, at least about 2000, at least about 2500, at least about 3000, or at least about 4000 amino acids. In other embodiments, the heterologous moiety is a polypeptide comprising or consisting essentially of or consisting of about 100 to about 200 amino acids, about 200 to about 300 amino acids, about 300 to about 400 amino acids, about 400 to about 500 amino acids, about 500 to about 600 amino acids, about 600 to about 700 amino acids, about 700 to about 800 amino acids, about 800 to about 900 amino acids, or about 900 to about 1000 amino acids.

[0429] In certain embodiments, the heterologous moiety improves one or more pharmacokinetic properties of the therapeutic protein without significantly affecting its biological activity or function.

[0430] In certain embodiments, the heterologous moiety increases the in vivo and / or in vitro half-life of the therapeutic protein of the disclosure. In other embodiments, the heterologous moiety facilitates visualization or localization of the therapeutic protein of the disclosure or a fragment thereof (e.g., a fragment comprising the heterologous moiety following proteolytic cleavage of the FVIII protein). Visualization and / or localization of the therapeutic protein of the disclosure or a fragment thereof can be in vivo, in vitro, ex vivo, or a combination thereof.

[0431] In other embodiments, the heterologous moiety increases the stability of the therapeutic protein or fragment thereof of the present disclosure (e.g., a fragment comprising a heterologous moiety following proteolytic cleavage of the therapeutic protein (e.g., a coagulation factor)). As used herein, the term "stability" refers to art-recognized measures of maintenance of one or more physical properties of the therapeutic protein in response to environmental conditions (e.g., increased or decreased temperature). In certain aspects, the physical property can be maintenance of the covalent structure of the therapeutic protein (e.g., absence of proteolytic cleavage, unwanted oxidation, or deamidation). In other aspects, the physical property can also be the presence of the therapeutic protein in a correctly folded state (e.g., absence of soluble or insoluble aggregates or precipitates). In one aspect, the stability of the therapeutic protein is measured by determining biophysical properties of the therapeutic protein, e.g., thermal stability, pH unfolding curve, stable removal of glycosylation, solubility, biochemical function (e.g., ability to bind to a protein, receptor, or ligand), and / or combinations thereof. In another aspect, biochemical function is evidenced by binding affinity of the interaction. In one aspect, a measure of protein stability is thermal stability, i.e., resistance to thermal challenge. Stability can be measured using methods known in the art, such as, HPLC (high performance liquid chromatography), SEC (size exclusion chromatography), DLS (dynamic light scattering), and the like. Methods of measuring thermal stability include, but are not limited to, differential scanning calorimetry (DSC), differential scanning fluorimetry (DSF), circular dichroism (CD), and thermal challenge assays.

[0432] In certain aspects, the therapeutic protein encoded by the nucleic acid molecule of the present disclosure comprises at least one half-life extension moiety, i.e., a heterologous moiety that increases the in vivo half-life of the therapeutic protein relative to the in vivo half-life of the corresponding therapeutic protein lacking such heterologous moiety. The in vivo half-life of the therapeutic protein can be determined by any method known to one of skill in the art, e.g., activity assays (e.g., a chromogenic assay or a primary clotting aPTT assay, where the therapeutic protein comprises a FVIII polypeptide), ELISA, etc.

[0433] In some embodiments, the presence of the one or more half-life extension moieties results in an increase in the half-life of the therapeutic protein compared to the half-life of the corresponding protein lacking such one or more half-life extension moieties. The half-life of the therapeutic protein comprising the half-life extension moiety is at least about 1.5-fold, at least about 2-fold, at least about 2.5-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 11-fold, or at least about 12-fold longer than the in vivo half-life of the corresponding therapeutic protein lacking such half-life extension moiety.

[0434] In one embodiment, the half-life of the therapeutic protein comprising a half-life extender is about 1.5-fold to about 20-fold, about 1.5-fold to about 15-fold, or about 1.5-fold to about 10-fold longer than the in vivo half-life of the corresponding protein lacking such a half-life extender. In another embodiment, the half-life of the therapeutic protein comprising a half-life extender is extended by about 2-fold to about 10-fold, about 2-fold to about 9-fold, about 2-fold to about 8-fold, about 2-fold to about 7-fold, about 2-fold to about 6-fold, about 2-fold to about 5-fold, about 2-fold to about 4-fold, about 2-fold to about 3-fold, about 2.5-fold to about 10-fold, about 2.5-fold to about 9-fold, about 2.5-fold to about 8-fold, about 2.5-fold to about 7-fold, about 2.5-fold to about 6-fold, about 2.5-fold to about 5-fold, about 2.5-fold to about 4-fold, about 2.5-fold to about 3-fold, about 3-fold to about 10-fold, about 3-fold to about 9-fold, about 3-fold to about 8-fold, about 3-fold to about 7-fold, about 3-fold to about 6-fold, about 3-fold to about 5-fold, about 3-fold to about 4-fold, about 4-fold to about 6-fold, about 5-fold to about 7-fold, or about 6-fold to about 8-fold, compared to the in vivo half-life of the corresponding protein lacking such a half-life extender.

[0435] In other embodiments, the half-life of the therapeutic protein comprising a half-life extender is at least about 17 hours, at least about 18 hours, at least about 19 hours, at least about 20 hours, at least about 21 hours, at least about 22 hours, at least about 23 hours, at least about 24 hours, at least about 25 hours, at least about 26 hours, at least about 27 hours, at least about 28 hours, at least about 29 hours, at least about 30 hours, at least about 31 hours, at least about 32 hours, at least about 33 hours, at least about 34 hours, at least about 35 hours, at least about 36 hours, at least about 48 hours, at least about 60 hours, at least about 72 hours, at least about 84 hours, at least about 96 hours, or at least about 108 hours.

[0436] In still other embodiments, the half-life of the therapeutic protein comprising a half-life extender is about 15 hours to about two weeks, about 16 hours to about one week, about 17 hours to about one week, about 18 hours to about one week, about 19 hours to about one week, about 20 hours to about one week, about 21 hours to about one week, about 22 hours to about one week, about 23 hours to about one week, about 24 hours to about one week, about 36 hours to about one week, about 48 hours to about one week, about 60 hours to about one week, about 24 hours to about 6 days, about 24 hours to about five days, about 24 hours to about four days, about 24 hours to about three days, or about 24 hours to about two days.

[0437] In some embodiments, the mean half-life of the therapeutic protein comprising a half-life extender per subject is about 15 hours, about 16 hours, about 17 hours, about 18 hours, about 19 hours, about 20 hours, about 21 hours, about 22 hours, about 23 hours, about 24 hours (1 day), about 25 hours, about 26 hours, about 27 hours, about 28 hours, about 29 hours, about 30 hours, about 31 hours, about 32 hours, about 33 hours, about 34 hours, about 35 hours, about 36 hours, about 40 hours, about 44 hours, about 48 hours (2 days), about 54 hours, about 60 hours, about 72 hours (3 days), about 84 hours, about 96 hours (4 days), about 108 hours, about 120 hours (5 days), about six days, about seven days (one week), about eight days, about nine days, about 10 days, about 11 days, about 12 days, about 13 days, or about 14 days.

[0438] One or more half-life extenders can be fused to the C-terminus or N-terminus of the therapeutic protein, or inserted within the therapeutic protein.

[0439] 1. Immunoglobulin constant region or portion thereof

[0440] In another aspect, the heterologous portion comprises one or more immunoglobulin constant regions or portions thereof (e.g., an Fc region). In one embodiment, the isolated nucleic acid molecule of the present disclosure further comprises a heterologous nucleic acid sequence encoding an immunoglobulin constant region or portion thereof. In some embodiments, the immunoglobulin constant region or portion thereof is an Fc region.

[0441] An immunoglobulin constant region is composed of domains denoted as CH (constant heavy) domains (CHI, CH2, etc.). Depending on the isotype (i.e., IgG, IgM, IgA IgD, or IgE), the constant region can be composed of three or four CH domains. Some isotype (e.g., IgG) constant regions also contain a hinge region. See Janeway et al. 2001, Immunobiology, Garland Publishing, N.Y., N.Y.

[0442] The immunoglobulin constant region or portion thereof of the present disclosure can be obtained from a number of different sources. In one embodiment, the immunoglobulin constant region or portion thereof is derived from a human immunoglobulin. However, it will be appreciated that the immunoglobulin constant region or portion thereof can be derived from an immunoglobulin of another mammalian species, including, for example, a rodent (e.g., mouse, rat, rabbit, guinea pig) or non-human primate (e.g., chimpanzee, macaque) species. Moreover, the immunoglobulin constant region or portion thereof can be derived from any immunoglobulin class, including IgM, IgG, IgD, IgA, and IgE, and any immunoglobulin isotype, including IgGl, IgG2, IgG3, and IgG4. In one embodiment, a human isotype IgGl is used.

[0443] A variety of immunoglobulin constant region gene sequences (e.g., human constant region gene sequences) are available in publicly available depositories. Constant domain sequences can be selected that have a particular effector function (or lack a particular effector function) or have a particular modification that reduces immunogenicity. The sequences of many antibodies and antibody-encoding genes have been published and the sequences of suitable Ig constant regions (e.g., hinge, CH2, and / or CH3 sequences or portions thereof) can be derived from these using art-recognized techniques. The genetic material obtained using any of the foregoing methods can then be altered or synthesized to obtain a polypeptide of the present disclosure. It will be further recognized that the scope of the present disclosure encompasses alleles, variants, and mutations of constant region DNA sequences.

[0444] Sequences of immunoglobulin constant regions or portions thereof can be cloned, for example, using polymerase chain reaction and primers selected to amplify the domain of interest. To clone sequences of immunoglobulin constant regions or portions thereof from antibodies, mRNA can be isolated from hybridomas, spleen, or lymphocytes, reverse transcribed into DNA, and then the antibody genes amplified by PCR. PCR amplification methods are described in U.S. Patent Nos. 4,683,195; 4,683,202; 4,800,159; 4,965,188; and, for example, "PCR Protocols: A Guide to Methods and Applications" Innis et al., eds., Academic Press, San Diego, CA (1990); Ho et al. 1989. Gene 77:51; Horton et al. 1993. Methods Enzymol. 217:270). PCR can be initiated by consensus constant region primers or by more specific primers based on published heavy and light chain DNA and amino acid sequences. PCR can also be used to isolate DNA clones encoding antibody light and heavy chains. In this case, libraries can be screened by consensus primers or by larger homologous probes (e.g., mouse constant region probes). Numerous primer sets suitable for amplifying antibody genes are known in the art (e.g., 5' primers based on N-terminal sequence of purified antibodies (Benhar and Pastan. 1994. Protein Engineering 7: 1509); rapid amplification of cDNA ends (Ruberti, F. et al. 1994. J. Immunol. Methods 173:33); antibody leader sequences (Larrick et al. 1989 Biochem. Biophys. Res. Commun. 160: 1250). Cloning of antibody sequences is further described in Newman et al., U.S. Patent No. 5,658,570, filed January 25, 1995, which is incorporated herein by reference.

[0445] The immunoglobulin constant region used herein can include all domains and hinge regions or portions thereof. In one embodiment, the immunoglobulin constant region or portion thereof comprises a CH2 domain, a CH3 domain, and a hinge region, i.e., an Fc region or FcRn binding partner.

[0446] As used herein, the term "Fc region" is defined as the polypeptide portion of the Fc region of a native Ig that corresponds to the dimeric association of the Fc domains of each of the two heavy chains. A native Fc region forms a homodimer with another Fc region. In contrast, the term "genetically fused Fc region" or "single chain Fc region" (scFc region) refers to a synthetic dimeric Fc region composed of Fc domains that are genetically linked within a single polypeptide chain (i.e., encoded by a single contiguous genetic sequence). See International Publication No. WO 2012 / 006635, incorporated herein by reference in its entirety.

[0447] In one embodiment, "Fc region" refers to the portion of a single Ig heavy chain that begins immediately upstream of the hinge region (i.e., residue 216 in IgG, with the first residue of the heavy chain constant region being 114) and ends at the C-terminus of the antibody. Thus, a complete Fc region comprises at least a hinge domain, a CH2 domain, and a CH3 domain.

[0448] An immunoglobulin constant region or portion thereof can be an FcRn binding partner. FcRn is active in adult epithelial tissues and is expressed in the lumen of the intestine, lung airway, nasal surface, vaginal surface, colon and rectal surface (U.S. Patent No. 6,485,726). An FcRn binding partner is a portion of an immunoglobulin that binds FcRn.

[0449] The FcRn receptor has been isolated from several mammalian species, including humans. The sequences of human FcRn, monkey FcRn, rat FcRn, and mouse FcRn are known (Story et al. 1994, J. Exp. Med. 180:2377). The FcRn receptor binds IgG (but not other immunoglobulin classes such as IgA, IgM, IgD, and IgE) at a relatively low pH, actively transports IgG across cells in the lumen to serosal direction, and then releases IgG at a relatively higher pH found in tissue fluid. It is expressed in adult epithelial tissues (U.S. Patent Nos. 6,485,726, 6,030,613, 6,086,875; WO 03 / 077834; US2003-0235536A1), including lung and intestinal epithelium (Israel et al. 1997, Immunology 92:69) kidney proximal tubule epithelium (Kobayashi et al. 2002, Am. J. Physiol. Renal Physiol. 282:F358) as well as nasal epithelium, vaginal surface, and biliary tree surface.

[0450] FcRn binding partners useful in the present disclosure encompass molecules that can be specifically bound by the FcRn receptor, including intact IgG, Fc fragments of IgG, and other fragments that include the complete binding region of the FcRn receptor. The region of the Fc portion of IgG that binds the FcRn receptor has been described based on x-ray crystallography (Burmeister et al. 1994, Nature 372:379). The primary contact surface of Fc with FcRn is near the junction of the CH2 and CH3 domains. The Fc-FcRn contacts are all within a single Ig heavy chain. FcRn binding partners include intact IgG, Fc fragments of IgG, and other fragments of IgG that include the complete binding region of the FcRn. The primary contact sites include amino acid residues 248, 250-257, 272, 285, 288, 290-291, 308-311, and 314 of the CH2 domain, and amino acid residues 385-387, 428, and 433-436 of the CH3 domain. References to the amino acid numbering of immunoglobulins or immunoglobulin fragments or regions are all based on Kabat et al. 1991, Sequences of Proteins of Immunological Interest, U.S. Department of Public Health, Bethesda, Md.

[0451] FcRn can efficiently transport Fc regions or FcRn binding partners that bind to FcRn across the epithelial barrier, thus providing a non-invasive way to systemically administer a desired therapeutic molecule. In addition, fusion proteins comprising Fc regions or FcRn binding partners are endocytosed by cells expressing FcRn. Rather than being tagged for degradation, however, these fusion proteins are recycled back into the circulation, thus increasing the in vivo half-life of these proteins. In certain embodiments, the portion of an immunoglobulin constant region is an Fc region or FcRn binding partner, which is typically associated with another Fc region or another FcRn binding partner to form a dimer or higher order multimer via disulfide bonds or other non-specific interactions.

[0452] Two FcRn receptors can bind a single Fc molecule. Crystallographic data indicates that each FcRn molecule binds a single polypeptide of the Fc homodimer. In one embodiment, linking an FcRn binding partner (e.g., an Fc fragment of IgG) to a biologically active molecule provides a means to deliver the biologically active molecule orally, buccally, sublingually, rectally, vaginally, as an aerosol for administration by nasal or pulmonary routes, or via the ocular route. In another embodiment, a coagulation factor protein can be administered invasively, e.g., subcutaneously, intravenously.

[0453] The FcRn binding partner region is a molecule or portion thereof that can be specifically bound by the FcRn receptor and thus actively transported by the FcRn receptor of the Fc region. Specific binding refers to two molecules that form a relatively stable complex under physiological conditions. Specific binding is characterized by a high affinity and low to moderate capacity, which distinguishes it from non-specific binding, which typically has a low affinity and moderate to high capacity. In general, binding is considered specific if the affinity constant, KA, is higher than 10 6 M -1 or higher than 10 8 M -1 If desired, non-specific binding can be reduced by altering the conditions of binding without substantially affecting specific binding. Suitable binding conditions, such as the concentration of the molecules, the ionic strength of the solution, the temperature allowed for binding, the time, the concentration of blocking agents (e.g., serum albumin, milk casein), etc., can be optimized by the skilled artisan using routine techniques.

[0454] In certain embodiments, the therapeutic protein encoded by the nucleic acid molecule of the present disclosure comprises one or more truncated Fc regions that are still sufficient to confer the binding properties of the Fc region to Fc receptors (FcRs). For example, the portion of the Fc region that binds FcRn (i.e., the FcRn binding portion) comprises about amino acids 282-438 from IgG1 (the primary contact sites are amino acids 248, 250-257, 272, 285, 288, 290-291, 308-311, and 314 of the CH2 domain and amino acid residues 385-387, 428, and 433-436 of the CH3 domain). Thus, the Fc region of the present disclosure can comprise or consist of the FcRn binding portion. The FcRn binding portion can be derived from the heavy chain of any isotype, including IgG1, IgG2, IgG3, and IgG4. In one embodiment, the FcRn binding portion is from an antibody of human isotype IgG1. In another embodiment, the FcRn binding portion is from an antibody of human isotype IgG4.

[0455] The Fc region can be obtained from a number of different sources. In one embodiment, the Fc region of the polypeptide is derived from a human immunoglobulin. However, it is understood that the Fc portion can be derived from an immunoglobulin of another mammalian species, including, for example, a rodent (e.g., mouse, rat, rabbit, guinea pig) or non-human primate (e.g., chimpanzee, macaque) species. Also, the polypeptide of the Fc domain or portion thereof can be derived from any immunoglobulin class (including IgM, IgG, IgD, IgA, and IgE) and immunoglobulin isotype (including IgG1, IgG2, IgG3, and IgG4). In another embodiment, the human isotype IgG1 is used.

[0456] In certain embodiments, the Fc variant confers an alteration (e.g., an increase or decrease) in at least one effector function imparted by an Fc portion comprising the wild-type Fc domain (e.g., the ability of the Fc region to bind an Fc receptor (e.g., FcyRI, FcyRII, or FcyRIII) or a complement protein (e.g., Clq), or to trigger antibody-dependent cellular cytotoxicity (ADCC), phagocytosis, or complement-dependent cellular cytotoxicity (CDCC)). In other embodiments, the Fc variant provides an engineered cysteine residue.

[0457] The Fc region of this disclosure may use Fc variants recognized in the art that are known to cause changes (e.g., enhancement or reduction) in effector function and / or FcR or FcRn binding. Specifically, the Fc region of this disclosure may include, for example, changes (e.g., substitutions) at one or more amino acid positions disclosed below: International PCT disclosures WO88 / 07089A1, WO96 / 14339A1, WO98 / 05787A1, WO98 / 23289A1, WO99 / 51642A1, WO99 / 58572A1, WO00 / 09560A2, WO00 / 32767A1, WO00 / 42072A2, WO02 / 44215A2, WO0 2 / 060919A2, WO03 / 074569A2, WO04 / 016750A2, WO04 / 029207A2, WO04 / 035752A2, WO04 / 063351A2, WO04 / 074455A 2. WO04 / 099249A2, WO05 / 040217A2, WO04 / 044859, WO05 / 070963A1, WO05 / 077981A2, WO05 / 092925A2, WO05 / 12378 0A2, WO06 / 019447A1, WO06 / 047350A2 and WO06 / 085967A2; US Patent Publications US2007 / 0231329, US2007 / 0231329, US2007 / 0237765, US2007 / 0237766, US2007 / 0237767, US2007 / 0243188, US20070248603, US20070286859, US20080057056; or U.S. Patents 5,648,260; 5,739,277; 5,834,250; 5,869,046; 6,096,871; 6,121,022; 6,194,551; 6,242,195; 6,277,375; 6,528,624; 6,538,124; 6,737,056; 6,821,505; 6,998,253; 7,083,784; 7,404,956 and 7,317,091, each of which is incorporated herein by reference. In one embodiment, specific changes may be made at one or more of the disclosed amino acid positions (e.g., specific substitutions of one or more amino acids disclosed in the art). In another embodiment, different changes may be made at one or more of the disclosed amino acid positions (e.g., different substitutions of one or more amino acid positions disclosed in the art).

[0458] The Fc region of IgG or FcRn binding partner can be modified according to accepted procedures such as site-directed mutagenesis, etc. to produce a modified IgG or Fc fragment or portion thereof that will be bound by FcRn. Such modifications include modifications away from the FcRn contact site as well as modifications within the contact site that preserve or even enhance binding to FcRn. For example, the following single amino acid residues in human IgGl Fc (Fcγl) can be substituted without significantly decreasing the binding affinity of Fc for FcRn: P238A, S239A, K246A, K248A, D249A, M252A, T256A, E258A, T260A, D265A, S267A, H268A, E269A, D270A, E272A, L274A, N276A, Y278A, D280A, V282A, E283A, H285A, N286A, T289A, K290A, R292A, E293A, E294A, Q295A, Y296F, N297A, S298A, Y300F, R301A, V303A, V305A, T307A, L309A, Q311A, D312A, N315A, K317A, E318A, K320A, K322A, S324A, K326A, A327Q, P329A, A330Q, P331A, E333A, K334A, T335A, S337A, K338A, K340A, Q342A, R344A, E345A, Q347A, R355A, E356A, M358A, T359A, K360A, N361A, Q362A, Y373A, S375A, D376A, A378Q, E380A, E382A, S383A, N384A, Q386A, E388A, N389A, N390A, Y391F, K392A, L398A, S400A, D401A, D413A, K414A, R416A, Q418A, Q419A, N421A, V422A, S424A, E430A, N434A, T437A, Q438A, K439A, S440A, S444A, and K447A, where, for example, P238A represents the wild-type proline substituted with alanine at position number 238. For example, one particular embodiment incorporates the N297A mutation, which removes the highly conserved N-glycosylation site. Other amino acids can be substituted for the wild-type amino acid at the positions specified above, in addition to alanine. Mutations can be introduced one at a time into the Fc, resulting in more than one hundred Fc regions that differ from the native Fc. In addition, combinations of two, three, or more of these single mutations can be introduced together, resulting in many hundreds more Fc regions.

[0459] Certain mutations described above can confer new functionality to the Fc region or Fc Rn binding partner. For example, one embodiment incorporates N297A, which removes a highly conserved N-glycosylation site. The effect of this mutation is to reduce immunogenicity, thereby enhancing the circulating half-life of the Fc region, and to render the Fc region incapable of binding FcyRI, FcyRIIA, FcyRIIB, and FcyRIIIA, without impairing affinity for FcRn (Routledge et al. 1995, Transplantation 60:847; Friend et al. 1999, Transplantation 68:1632; Shields et al. 1995, J. Biol. Chem. 276:6591). As yet another example of new functionality resulting from the mutations described above, in some cases, affinity for FcRn can be increased beyond that of the wild type. This increased affinity can reflect an increased "on" rate, a decreased "off rate, or both an increased "on" rate and a decreased "off rate. Examples of mutations that are believed to confer increased affinity for FcRn include, but are not limited to, T256A, T307A, E380A, and dN434A (Shields et al. 2001, J. Biol. Chem. 276:6591).

[0460] In addition, at least three human Fcy receptors appear to recognize binding sites on IgG in the lower hinge region, generally amino acids 234-237. Thus, another example of new functionality and potentially reduced immunogenicity can result from mutations in this region, such as by replacing the amino acids 233-236 of human IgGl "ELLG" (SEQ ID NO: 45) with the corresponding sequence of IgG2 "PVA" (with one amino acid deletion). It has been shown that FcyRI, FcyRII, and FcyRIII, which mediate a variety of effector functions, will not bind IgGl when such mutations have been introduced. Ward and Ghetie 1995, Therapeutic Immunology 2:77 and Armour et al. 1999, Eur. J. Immunol. 29:2613.

[0461] In another embodiment, the immunoglobulin constant region or portion thereof comprises an amino acid sequence in the hinge region or portion thereof that forms one or more disulfide bonds with a second immunoglobulin constant region or portion thereof. The second immunoglobulin constant region or portion thereof can be linked to a second polypeptide that binds the therapeutic protein and the second polypeptide together. In some embodiments, the second polypeptide is an enhancer moiety. As used herein, the term "enhancer moiety" refers to a molecule, fragment or polypeptide composition that is capable of enhancing the activity of a therapeutic protein. The enhancer moiety can be a cofactor, e.g., where the therapeutic protein is a coagulation factor, soluble tissue factor (sTF) or thrombopoeitin. Thus, upon activation of the coagulation factor, the enhancer moiety can be used to enhance the coagulation factor activity.

[0462] In certain embodiments, the therapeutic protein encoded by the nucleic acid molecule of the present disclosure comprises an amino acid substitution to an immunoglobulin constant region or portion thereof (e.g., an Fc variant) that alters the antigen-independent effector function of the Ig constant region, in particular the circulating half-life of the protein.

[0463] 2. scFc Region

[0464] In another aspect, the heterologous moiety comprises a scFC (single chain Fc) region. In one embodiment, the isolated nucleic acid molecule of the present disclosure further comprises a heterologous nucleic acid sequence encoding a scFc region. The scFc region comprises at least two immunoglobulin constant regions or portions thereof (e.g., Fc portions or domains (e.g., 2, 3, 4, 5, 6 or more Fc portions or domains)) within the same linear polypeptide chain that are capable of folding (e.g., intra- or inter-molecularly) to form a functional scFc region that is linked by an Fc peptide linker. For example, in one embodiment, the polypeptide of the present disclosure is capable of binding to at least one Fc receptor (e.g., FcRn, FcyR receptor (e.g., FcyRIII) or a complement protein (e.g., Clq) via its scFc in order to improve or trigger immune effector functions (e.g., antibody-dependent cellular cytotoxicity (ADCC), phagocytosis or complement-dependent cellular cytotoxicity (CDCC) and / or improve manufacturability.

[0465] 3. CTP

[0466] In another aspect, the heterologous moiety comprises a C-terminal peptide (CTP) of the beta subunit of human chorionic gonadotropin, or a fragment, variant or derivative thereof. Inhibition of one or more CTP peptides inserted into a recombinant protein increases the in vivo half-life of the protein. See, e.g., U.S. Patent No. 5,712,122, which is incorporated by referen...

Claims

1. A nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette; wherein the first ITR and / or the second ITR is a non-adenovirus-associated virus (non- AAV) ITR, wherein the gene cassette is located between the first ITR and the second ITR, and wherein the gene cassette encodes a therapeutic protein, an miRNA, or both a therapeutic protein and an miRNA.

2. The nucleic acid molecule of claim 1, wherein the therapeutic protein comprises a blood clotting factor.

3. The nucleic acid molecule of claim 1, wherein the non- AAV is selected from a member of the Parvoviridae family of viruses.

4. The nucleic acid molecule of claim 1, wherein the first ITR and the second ITR are non- AAV ITRs.

5. The nucleic acid molecule of claim 3, wherein the member of the Parvoviridae family of viruses is a Erythrovirus Parvovirus B19 (human virus) or a Dependovirus Goose Parvovirus (GPV) strain.

6. The nucleic acid molecule of claim 1, further comprising a tissue-specific promoter.

7. The nucleic acid molecule of claim 6, wherein the promoter drives expression of the therapeutic protein in a hepatocyte, an endothelial cell, a muscle cell, a sinusoidal cell, or any combination thereof.

8. The nucleic acid molecule of any one of claims 1-7, wherein the nucleotide sequence further comprises: (a) an intron sequence, (b) a post-transcriptional regulatory element, (c) a 3' UTR poly(A) tail sequence, (d) an enhancer sequence, or (e) any combination of (a)-(d).

9. The nucleic acid molecule of claim 2, wherein the blood clotting factor comprises Factor I (FI), Factor II (FII), Factor V (FV), Factor VII (FVII), Factor VIII (FVIII), Factor IX (FIX), Factor X (FX), Factor XI (FXI), Factor XII (FXII), Factor XIII (FXIII), von Willebrand Factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), or any combination thereof.

10. A pharmaceutical composition comprising the nucleic acid molecule of any one of claims 1-9 and a pharmaceutically acceptable carrier.

Citation Information

Patent Citations

  • Lentiviral vectors encoding clotting factors for gene therapy

    EP1395293A1

  • Serum albumin binding moieties

    US20030069395A1

  • Central airway administration for systemic delivery of therapeutics

    US20030235536A1

  • Fc Variants Having Increased Affinity for FcyRIIb

    US20070231329A1

  • Fc Variants Having Increased Affinity for FcyRl

    US20070237765A1