Nucleic acid molecules and their uses
By designing nucleic acid molecules containing non-AAV ITRs and gene cassettes encoding therapeutic proteins, the problems of AAV vector packaging ability and immune response in gene therapy are solved, and the continuous expression effect in vitro and in vivo is achieved.
Patent Information
- Application Number
- CN201880065324.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-08-09
- Filing Date
- 2018-08-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-03-27
AI Technical Summary
Existing AAV vectors have limited viral packaging capabilities and induce immune responses in gene therapy, making it difficult to achieve continuous expression of target sequences in vitro and in vivo environments.
A nucleic acid molecule is designed to include a first reverse terminal repeat (ITR) and a second ITR of non-AAV, and a gene cassette encoding a therapeutic protein, such as a coagulation factor. This nucleic acid molecule does not contain a gene encoding a capsid protein, and the effective expression of therapeutic proteins is achieved by utilizing non-AAV ITR and tissue-specific promoters.
Through this technical means, continuous and effective target sequence expression can be achieved in vitro and in vivo environments, avoiding the restriction of the immune response of AAV vectors, and is suitable for the treatment of a variety of diseases.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
[0001] Reference to Electronically Submitted Sequence Listing
[0002] The content of the sequence listing electronically submitted as an ASCII text file (Name: 4159_493PC01_ST25; Size: 434,561 bytes; and Creation Date: August 9, 2018) is incorporated herein by reference in its entirety. BACKGROUND OF THE INVENTION
[0004] Gene therapy offers a way to treat a variety of diseases durably. In the past, gene therapy has generally relied on the use of viruses. For this purpose, many viral agents can be selected, each having unique properties that will make it more or less suitable for gene therapy. Zhou et al., Adv Drug Deliv Rev. 106(Pt A):3-26, 2016. However, the undesired properties of some viral vectors, including their immunogenicity or tendency to cause cancer, have led to clinical safety concerns, and until recently, their current clinical applications have been limited to certain applications such as vaccines and oncolytic strategies. Cotter et al., Front Biosci. 10:1098-105(2005).
[0005] Adeno-associated virus (AAV) is one of the most commonly studied gene therapy vectors. AAV is a protein shell that surrounds and protects a small single-stranded DNA genome of approximately 4.8 kilobases (kb). Naso et al., BioDrugs, 31(4):317-334, 2017. AAV belongs to the family Parvoviridae and depends on co-infection with other viruses (primarily adenoviruses) in order to replicate. Ibid. Its single-stranded genome contains three genes: Rep (replication), Cap (capsid), and aap (assembly). Ibid. Flanking these coding sequences are inverted terminal repeats (ITRs) required for genome replication and packaging. Ibid. The length of the two cis-acting AAV ITRs is approximately 145 nucleotides, which are interrupted by palindromic sequences that can fold into T-shaped hairpin structures that act as primers during the initiation of DNA replication.
[0006] However, the use of conventional AAV as a gene delivery vector has led to several drawbacks. One of the major drawbacks is related to the limited viral packaging capacity of AAV for heterologous DNA of approximately 4.5 kb. (Dong et al., Hum Gene Ther. 7(17):2101-12, 1996). In addition, the administration of AAV vectors can induce an immune response in humans. Although the immunogenicity of AAV has been shown to be lower than that of some other viruses (i.e., adenovirus), the capsid proteins can trigger multiple components of the human immune system. See Naso et al., 2017. AAV is a common virus in the human population, and most people have been exposed to AAV, and thus, most people have developed an immune response against the specific variant to which they were previously exposed. This pre-existing adaptive response can include NAbs and T cells that can reduce the clinical efficacy of subsequent reinfection with AAV and / or eliminate the transduced cells, rendering patients with pre-existing anti-AAV immunity ineligible for AAV-based gene therapy treatment. Whether administered locally or systemically, the virus will be recognized as a foreign protein, and thus, the adaptive immune system will attempt to eliminate it. In addition, anti-AAV neutralizing antibodies induced by AAV treatment impede repeated treatment with AAV when the first AAV treatment fails to reach a therapeutic efficacy level. Furthermore, evidence suggests that the T-shaped hairpin loop of the AAV ITR is susceptible to inhibition by host cell proteins / protein complexes that bind to the T-shaped hairpin structure of the AAV ITR. See, e.g., Zhou et al., Scientific Reports 7:5432 (July 14, 2017).
[0007] Accordingly, there is a need in the art for effective and sustained expression of target sequences, e.g., therapeutic proteins and / or miRNAs, in in vitro and in vivo settings while avoiding some of the unforeseen consequences and limitations of existing AAV vector technologies. SUMMARY OF THE INVENTION
[0009] One aspect of the present disclosure relates to a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette encoding a therapeutic protein; wherein the first ITR and / or the second ITR is an ITR of a non-adeno-associated virus (non-AAV), wherein the gene cassette is located between the first ITR and the second ITR, and wherein the therapeutic protein comprises a clotting factor. In some embodiments, the non-AAV is selected from members of the virus family Parvoviridae and any combination thereof. In some embodiments, the first ITR is an ITR of a non-AAV and the second ITR is an ITR of an adeno-associated virus (AAV), or wherein the first ITR is an ITR of an AAV and the second ITR is an ITR of a non-AAV. In other embodiments, the first ITR and the second ITR are ITRs of a non-AAV. In some embodiments, the first ITR and the second ITR are the same. In some embodiments, the first ITR and / or the second ITR comprises a palindromic sequence that is uninterrupted. In other embodiments, the first ITR and / or the second ITR comprises a palindromic sequence that is interrupted.
[0010] In some embodiments, the non-AAV in the nucleic acid molecule is a member of the viral family Parvoviridae. In some embodiments, the member of the viral family Parvoviridae is selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof. In some embodiments, the member of the viral family Parvoviridae is parvovirus B19 (human virus) of the genus Erythrovirus. In some embodiments, the member of the viral family Parvoviridae is Muscovy duck parvovirus (MDPV) strain. In some embodiments, the MDPV strain is the attenuated FZ91-30. In some embodiments, the MDPV strain is the pathogenic YY. In still other embodiments, Dependoparvovirus is a strain of goose parvovirus (GPV) of the genus Dependovirus. In some embodiments, the GPV strain is the attenuated 82-0321V. In some embodiments, the GPV strain is the pathogenic B. In some embodiments, the member of the viral family Parvoviridae is selected from the group consisting of porcine parvovirus (U44978), minute virus of mice (U34256), canine parvovirus (M19296), mink enteritis virus (D00765), and any combination thereof.
[0011] In some embodiments, the nucleic acid molecule further comprises a promoter. In some embodiments, the promoter is a tissue-specific promoter. In some embodiments, the promoter drives the expression of the therapeutic protein in hepatocytes, endothelial cells, muscle cells, sinusoidal cells, or any combination thereof. In some embodiments, the promoter is located 5' of the nucleic acid sequence encoding the clotting factor. In some embodiments, the promoter is selected from the group consisting of the murine thyroxine promoter (mTTR), the endogenous human factor VIII promoter (F8), the human α-1-antitrypsin promoter (hAAT), the human albumin minimal promoter, the murine albumin promoter, the Tristetraprolin (TTP) promoter, the CASI promoter, the CAG promoter, the cytomegalovirus (CMV) promoter, α1-antitrypsin (AAT), muscle creatine kinase (MCK), myosin heavy chain α (αMHC), myoglobin (MB), desmin (DES), SPc5-12, 2R5Sc5-12, dMCK, tMCK, phosphoglycerate kinase (PGK) promoter, and any combination thereof. In some embodiments, the promoter comprises the TTP promoter.
[0012] In some embodiments, the nucleic acid molecule further comprises an intron sequence. In some embodiments, the intron sequence is located 5' of the nucleic acid sequence encoding the clotting factor. In some embodiments, the intron sequence is located 3' of the promoter. In some embodiments, the intron sequence comprises a synthetic intron sequence. In some embodiments, the intron sequence comprises SEQ ID NO:115.
[0013] In some embodiments, the nucleic acid molecule comprises a post-transcriptional regulatory element. In some embodiments, the post-transcriptional regulatory element is located 3' of the nucleic acid sequence encoding the clotting factor. In some embodiments, the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, or any combination thereof. In some embodiments, the microRNA binding site comprises a binding site for miR142-3p.
[0014] In some embodiments, the nucleic acid molecule comprises a 3'UTR poly(A) tail sequence. In some embodiments, the 3'UTR poly(A) tail sequence is selected from the group consisting of bGH poly(A), actin poly(A), hemoglobin poly(A), and any combination thereof. In some embodiments, the 3'UTR poly(A) tail sequence comprises bGH poly(A).
[0015] In some embodiments, the nucleic acid molecule comprises an enhancer sequence. In some embodiments, the enhancer sequence is located between the first ITR and the second ITR.
[0016] In some embodiments, the nucleic acid molecule comprises, in this order: a first ITR, a gene cassette, and a second ITR; wherein the gene cassette comprises a tissue-specific promoter sequence, an intron sequence, a nucleic acid sequence encoding a therapeutic protein (e.g., a clotting factor) or miRNA, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence.
[0017] In some embodiments, the nucleic acid molecule comprises, in this order: a tissue-specific promoter sequence, an intron sequence, a nucleic acid sequence encoding an FVIII polypeptide, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence.
[0018] In one embodiment, the nucleic acid molecule comprises:
[0019] (a) A first ITR, which is an ITR of a member of the non-AAV family of the Parvoviridae;
[0020] (b) A tissue-specific promoter sequence, e.g., the TTP promoter;
[0021] (c) An intron, e.g., a synthetic intron;
[0022] (d) A nucleotide sequence encoding a therapeutic protein (e.g., a clotting factor) or miRNA;
[0023] (e) A post-transcriptional regulatory element, e.g., WPRE;
[0024] (f) A 3'UTR poly(A) tail sequence, e.g., bGHpA; and
[0025] (g) A second ITR, which is an ITR of a member of the non-AAV family of the Parvoviridae.
[0026] In some embodiments, the nucleic acid molecule comprises single-stranded nucleic acid. In some embodiments, the gene cassette comprises double-stranded nucleic acid.
[0027] In some embodiments, the nucleic acid molecule comprises a gene encoding a therapeutic protein (e.g., a clotting factor), wherein the clotting factor is expressed by hepatocytes, endothelial cells, muscle cells, sinusoidal cells, or any combination thereof. In some embodiments, the clotting factor comprises factor I (FI), factor II (FII), factor V (FV), factor VII (FVII), factor VIII (FVIII), factor IX (FIX), factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FVIII), von Willebrand factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, α2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), or any combination thereof. In some embodiments, the clotting factor is FVIII. In some embodiments, the FVIII comprises full-length mature FVIII. In some embodiments, the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 106. In some embodiments, the FVIII comprises an A1 domain, an A2 domain, an A3 domain, a C1 domain, a C2 domain, and a partial B domain or does not comprise a B domain. In some embodiments, the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 109.
[0028] In some embodiments, the clotting factor comprises a heterologous moiety. In some embodiments, the heterologous moiety is selected from the group consisting of albumin or a fragment thereof, the Fc region of an immunoglobulin, the C-terminal peptide (CTP) of the β subunit of human chorionic gonadotropin, a PAS sequence, a HAP sequence, transferrin or a fragment thereof, an albumin-binding moiety, derivatives thereof, and any combination thereof. In some embodiments, the heterologous moiety is linked to the N-terminus or C-terminus of FVIII or inserted between two amino acids in the FVIII. In some embodiments, the heterologous moiety is inserted between two amino acids at one or more insertion sites selected from the insertion sites listed in Table 5. In some embodiments, the FVIII further comprises an A1 domain, an A2 domain, a C1 domain, a C2 domain, an optional B domain, and a heterologous moiety, wherein the heterologous moiety is inserted immediately downstream of the amino acid corresponding to amino acid 745 of mature FVIII (SEQ ID NO: 106).
[0029] In some embodiments, the FVIII further comprises an FcRn-binding moiety. In some embodiments, the FcRn-binding moiety comprises the Fc region of an immunoglobulin constant domain. In some embodiments, the nucleic acid sequence encoding the FVIII is codon-optimized. In some embodiments, the nucleic acid sequence encoding the FVIII is codon-optimized for expression in humans.
[0030] In some embodiments, the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO: 107. In some embodiments, the nucleic acid sequence encoding the FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO: 71.
[0031] In some embodiments, the nucleic acid molecule is formulated with a delivery agent. In some embodiments, the delivery agent comprises one or more lipid nanoparticles. In some embodiments, the delivery agent is selected from the group consisting of liposomes, non-lipid polymeric molecules, endosomes, and any combination thereof.
[0032] In some embodiments, the nucleic acid molecule is formulated for intravenous, transdermal, intradermal, subcutaneous, pulmonary, or oral delivery, or any combination thereof. In some embodiments, the nucleic acid molecule is formulated for intravenous delivery.
[0033] In some embodiments, provided herein is a vector comprising a nucleic acid molecule as described throughout this disclosure. In some embodiments, provided herein is a polypeptide encoded by a nucleic acid molecule as described throughout this disclosure. In some embodiments, a host cell is provided that comprises a nucleic acid molecule as described throughout this disclosure.
[0034] In some embodiments, provided herein is a pharmaceutical composition comprising (a) a nucleic acid as described herein, a vector comprising a nucleic acid molecule as described throughout this disclosure, a polypeptide encoded by a nucleic acid molecule as described throughout this disclosure, or a host cell comprising a nucleic acid molecule as described throughout this disclosure; (b) an LNP; and (c) a pharmaceutically acceptable excipient. In some embodiments, provided herein is a kit comprising a nucleic acid molecule as described throughout this disclosure and instructions for administering the nucleic acid molecule to a subject in need thereof. In some embodiments, provided herein is a baculovirus system for producing a nucleic acid molecule as described herein. In some embodiments, provided herein is a baculovirus in which the nucleic acid molecule disclosed herein is produced in insect cells.
[0035] In some embodiments, provided herein is a nanoparticle delivery system for an expression construct, wherein the expression construct comprises a nucleic acid molecule as described herein.
[0036] Also disclosed herein is a method for producing a polypeptide having blood coagulation activity, comprising: culturing a host cell disclosed herein under suitable conditions and recovering the polypeptide having blood coagulation activity. In some embodiments, disclosed herein is a method for expressing a blood coagulation factor in a subject in need thereof, comprising administering to the subject a nucleic acid molecule disclosed herein, a vector disclosed herein, a polypeptide disclosed herein, or a pharmaceutical composition disclosed herein. In some embodiments, disclosed herein is a method for treating a subject suffering from a blood coagulation factor deficiency, comprising administering to the subject a nucleic acid molecule disclosed herein, a vector disclosed herein, a polypeptide disclosed herein, or a pharmaceutical composition disclosed herein. In some embodiments, the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonary, or any combination thereof. In some embodiments, the nucleic acid molecule is administered intravenously. In some embodiments, the method further comprises administering a second agent to the subject. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.
[0037] In some embodiments, administration of the nucleic acid molecule to the subject results in an increase in FVIII activity relative to the FVIII activity in the subject prior to the administration, wherein the increase in FVIII activity is at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 11-fold, at least about 12-fold, at least about 13-fold, at least about 14-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold or at least about 100-fold.
[0038] In some embodiments, the subject has a bleeding disorder. In some embodiments, the bleeding disorder is hemophilia. In some embodiments, the bleeding disorder is hemophilia A. Brief Description of the Drawings
[0040] Figure 1A is a schematic diagram of a single-stranded coagulation factor (e.g., FVIII) expression cassette. The positions of the 5' ITR from a non-AAV (with a hairpin loop at the end of the ssDNA structure), the 3' ITR from a non-AAV (with a hairpin loop), a promoter sequence (e.g., TTPp), and a transgene sequence (e.g., the FVIIIco6XTEN sequence with XTEN144 inserted within the B domain) are shown. The exemplary expression cassette also shows additional possible elements, e.g., an intron sequence, a WPREmut sequence, and a bGHpA sequence.
[0041] Figures 1B - 1D is for the preparation of a plasmid of a single-stranded coagulation factor expression cassette (such as the cassette shown in Figure 1A ), wherein the ITR of the cassette is derived from AAV2 ( Figure 1B ), B19 ( Figure 1C ), or GPV ( Figure 1D ). The plasmid construct containing the ssFVIII expression cassette as shown is digested with PvuII (at the PvuII site) ( Figure 1B ) or digested with LguI (at the LguI site) ( Figure 1C and 1D ) to release the viral genome. The double-stranded DNA is heated to 95 °C to generate ssDNA, and then incubated at 4 °C to allow the ITR structure to form.
[0042] Figure 2A is a phylogenetic tree illustrating the relationships among various parvovirus family members. B19, AAV-2, and GPV are marked with hollow boxes.
[0043] Figure 2BSchematic diagrams of various boxes (including hairpin structures).
[0044] Figure 3A and 3B are alignments of the ITRs of B19, GPV, and AAV2 ( Figure 3A ) and the ITRs of B19 and GPV ( Figure 3B ). Gray shading indicates homology.
[0045] Figures 4A - 4C Shows FVIII plasma activity after administration of single-stranded FVIII-AAV naked DNA (ssAAV-FVIII; Figure 4A ), ssDNA-B19 FVIII ( Figure 4B ), or ssDNA-GPV FVIII ( Figure 4C ) via hydrodynamic injection (HDI) in Hem A mice. FVIII activity (as a percentage of the control) was measured in plasma samples at 24 hours, 3 days, 2 weeks, 3 weeks, 1 month, 2 months, 3 months, and 4 months in mice treated with a single HDI of ssDNA at 50 μg / mouse ( Figure 4C ), 20 μg / mouse ( Figure 4A and 4B ), 10 μg / mouse ( Figure 4A and 4C ), or 5 μg / mouse ( Figure 4A ). HDI of plasmid DNA at 5 μg / mouse was given as a control ( Figures 4A - 4C ). DETAILED DESCRIPTION OF THE INVENTION
[0047] The present disclosure describes plasmid-like nucleic acid molecules that comprise a first inverted terminal repeat (ITR), a second ITR, and a gene cassette (e.g., a gene cassette encoding a therapeutic protein or miRNA), wherein the first ITR and / or the second ITR is a non-adeno-associated virus ITR (e.g., the first ITR and / or the second ITR is from a non-AAV). In some embodiments, the gene cassette encodes a therapeutic protein. In some embodiments, the therapeutic protein comprises a protein selected from the group consisting of clotting factors, growth factors, hormones, cytokines, antibodies, fragments thereof, or combinations thereof. In some embodiments, the gene cassette encodes dystrophin X-linked type, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acidic α-glucosidase), or any combination thereof.
[0048] In some embodiments, the therapeutic protein comprises a clotting factor. In a particular embodiment, the therapeutic protein comprises an FVIII or FIX protein.
[0049] In some embodiments, the gene cassette encodes an miRNA. In certain embodiments, the miRNA downregulates the expression of a target gene selected from SOD1, HTT, RHO, or any combination thereof.
[0050] In certain embodiments, the non-AAV is selected from members of the parvovirus family of viruses and any combination thereof. The present disclosure also relates to a method of expressing a therapeutic protein (e.g., a clotting factor, e.g., FVIII) in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette (e.g., a gene cassette encoding a therapeutic protein or an miRNA), wherein the first ITR and / or the second ITR is a non-adeno-associated virus (non-AAV) ITR. In certain embodiments, the present disclosure describes an isolated nucleic acid molecule comprising a nucleotide sequence that has sequence homology to a nucleotide sequence selected from SEQ ID NO:113 and 120.
[0051] Exemplary constructs of the present disclosure are shown in the drawings and sequence listing. To provide a clear understanding of the specification and claims, the following definitions are provided.
[0052] I. Definitions
[0053] It should be noted that the term "a" or "an" entity refers to one or more of that entity: for example, "nucleotide sequence" should be understood to represent one or more nucleotide sequences. Similarly, "therapeutic protein" and "miRNA" should be understood to represent one or more therapeutic proteins and one or more miRNAs, respectively. Likewise, the terms "a", "one or more", and "at least one" may be used interchangeably herein.
[0054] The term "about" is used herein to mean approximately, roughly, around, or in the region of. When the term "about" is used in conjunction with a numerical range, it modifies that range by extending the upper and lower boundaries of the indicated numerical values. Generally, the term "about" is used herein to modify a numerical value that is above or below a specified value by a variation magnitude of 10% higher or lower (higher or lower).
[0055] Also as used herein, "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the absence of combinations when interpreted in the alternative ("or").
[0056] "Nucleic acid", "nucleic acid molecule", "nucleotide", "nucleotide sequence", and "polynucleotide" are used interchangeably and refer to the phosphoester polymeric forms of ribonucleosides (adenosine, guanosine, uridine, or cytidine; "RNA molecule") or deoxyribonucleosides (deoxyadenosine, deoxyguanosine, deoxythymidine, or deoxycytidine; "DNA molecule"), or any phosphoester analogs thereof, such as phosphorothioates and thioesters in single-stranded or double-stranded helical forms. A single-stranded nucleic acid sequence refers to single-stranded DNA (ssDNA) or single-stranded RNA (ssRNA). Double-stranded DNA-DNA, DNA-RNA, and RNA-RNA helices are possible. The term nucleic acid molecule, particularly a DNA or RNA molecule, refers only to the primary and secondary structures of the molecule and is not limited to any particular tertiary form. Thus, this term includes double-stranded DNA found particularly in linear or circular DNA molecules (e.g., restriction fragments), plasmids, supercoiled DNA, and chromosomes. When discussing the structure of a particular double-stranded DNA molecule, sequences may be described herein according to the conventional convention of giving only the sequence in the 5' to 3' direction along the non-transcribed strand of the DNA (i.e., the strand having a sequence homologous to the mRNA). A "recombinant DNA molecule" is a DNA molecule that has undergone molecular biology manipulations. DNA includes, but is not limited to, cDNA, genomic DNA, plasmid DNA, synthetic DNA, and semi-synthetic DNA. The "nucleic acid composition" of the present disclosure comprises one or more nucleic acids as described herein.
[0057] As used herein, "inverted terminal repeat" (or "ITR") refers to a nucleic acid subsequence located at the 5' or 3' end of a single-stranded nucleic acid sequence that comprises a set of nucleotides (initial sequence) that is downstream of and reverse complementary to, i.e., palindromic to, itself. The inter-nucleotide spacer sequence between the initial sequence and the reverse complement can be of any length, including zero. In one embodiment, the ITRs useful in the present disclosure comprise one or more "palindromic sequences". ITRs can have a number of functions. In some embodiments, the ITRs described herein form hairpin structures. In some embodiments, the ITRs form T-shaped hairpin structures. In some embodiments, the ITRs form non-T-shaped hairpin structures, e.g., U-shaped hairpin structures. In some embodiments, the ITRs promote the long-term survival of nucleic acid molecules in the cell nucleus. In some embodiments, the ITRs promote the permanent survival of nucleic acid molecules in the cell nucleus (e.g., throughout the life of the cell). In some embodiments, the ITRs promote the stability of nucleic acid molecules in the cell nucleus. In some embodiments, the retention of nucleic acid molecules in the cell nucleus is promoted. In some embodiments, the ITRs promote the persistence of nucleic acid molecules in the cell nucleus. In some embodiments, the ITRs inhibit or prevent the degradation of nucleic acid molecules in the cell nucleus.
[0058] In one embodiment, the initial sequence and / or reverse complement comprises from about 2 to 600 nucleotides, from about 2 to 550 nucleotides, from about 2 to 500 nucleotides, from about 2 to 450 nucleotides, from about 2 to 400 nucleotides, from about 2 to 350 nucleotides, from about 2 to 300 nucleotides, or from about 2 to 250 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises from about 5 to 600 nucleotides, from about 10 to 600 nucleotides, from about 15 to 600 nucleotides, from about 20 to 600 nucleotides, from about 25 to 600 nucleotides, from about 30 to 600 nucleotides, from about 35 to 600 nucleotides, from about 40 to 600 nucleotides, from about 45 to 600 nucleotides, from about 50 to 600 nucleotides, from about 60 to 600 nucleotides, from about 70 to 600 nucleotides, from about 80 to 600 nucleotides, from about 90 to 600 nucleotides, from about 100 to 600 nucleotides, from about 150 to 600 nucleotides, from about 200 to 600 nucleotides, from about 300 to 600 nucleotides, from about 350 to 600 nucleotides, from about 400 to 600 nucleotides, from about 450 to 600 nucleotides, from about 500 to 600 nucleotides, or from about 550 to 600 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises from about 5 to 550 nucleotides, from about 5 to 500 nucleotides, from about 5 to 450 nucleotides, from about 5 to 400 nucleotides, from about 5 to 350 nucleotides, from about 5 to 300 nucleotides, or from about 5 to 250 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises from about 10 to 550 nucleotides, from about 15 to 500 nucleotides, from about 20 to 450 nucleotides, from about 25 to 400 nucleotides, from about 30 to 350 nucleotides, from about 35 to 300 nucleotides, or from about 40 to 250 nucleotides. In certain embodiments, the initial sequence and / or reverse complement comprises about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, or about 600 nucleotides. In a particular embodiment, the initial sequence and / or reverse complement comprises about 400 nucleotides.
[0059] In other embodiments, the initial sequence and / or reverse complement comprises about 2 - 200 nucleotides, about 5 - 200 nucleotides, about 10 - 200 nucleotides, about 20 - 200 nucleotides, about 30 - 200 nucleotides, about 40 - 200 nucleotides, about 50 - 200 nucleotides, about 60 - 200 nucleotides, about 70 - 200 nucleotides, about 80 - 200 nucleotides, about 90 - 200 nucleotides, about 100 - 200 nucleotides, about 125 - 200 nucleotides, about 150 - 200 nucleotides, or about 175 - 200 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2 - 150 nucleotides, about 5 - 150 nucleotides, about 10 - 150 nucleotides, about 20 - 150 nucleotides, about 30 - 150 nucleotides, about 40 - 150 nucleotides, about 50 - 150 nucleotides, about 75 - 150 nucleotides, about 100 - 150 nucleotides, or about 125 - 150 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2 - 100 nucleotides, about 5 - 100 nucleotides, about 10 - 100 nucleotides, about 20 - 100 nucleotides, about 30 - 100 nucleotides, about 40 - 100 nucleotides, about 50 - 100 nucleotides, or about 75 - 100 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2 - 50 nucleotides, about 10 - 50 nucleotides, about 20 - 50 nucleotides, about 30 - 50 nucleotides, about 40 - 50 nucleotides, about 3 - 30 nucleotides, about 4 - 20 nucleotides, or about 5 - 10 nucleotides. In another embodiment, the initial sequence and / or reverse complement consists of two nucleotides, three nucleotides, four nucleotides, five nucleotides, six nucleotides, seven nucleotides, eight nucleotides, nine nucleotides, ten nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides. In other embodiments, the intervening nucleotides between the initial sequence and the reverse complement are (e.g., consist of) 0 nucleotides, 1 nucleotide, two nucleotides, three nucleotides, four nucleotides, five nucleotides, six nucleotides, seven nucleotides, eight nucleotides, nine nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides.
[0060] Thus, an "ITR" as used herein can fold upon itself and form a double-stranded segment. For example, the sequence GATCXXXXGATC contains the initial sequence of GATC and its complement (3'CTAG5') when folded to form a double helix. In some embodiments, the ITR contains a continuous palindromic sequence (e.g., GATCGATC) between the initial sequence and the reverse complement. In some embodiments, the ITR contains a discontinuous palindromic sequence (e.g., GATCXXXXGATC) between the initial sequence and the reverse complement. In some embodiments, the complementary portions of the continuous or discontinuous palindromic sequences interact with each other to form a "hairpin loop" structure. As used herein, a "hairpin loop" structure is produced when at least two complementary sequences on a single-stranded nucleotide molecule base pair to form a double-stranded portion. In some embodiments, only a portion of the ITR forms a hairpin loop. In other embodiments, the entire ITR forms a hairpin loop.
[0061] In the present disclosure, at least one ITR is an ITR of a non-adeno-associated virus (non-AAV). In certain embodiments, the ITR is an ITR of a non-AAV member of the viral family Parvoviridae. In some embodiments, the ITR is an ITR of a non-AAV member of the genus Dependovirus or Erythrovirus. In a particular embodiment, the ITR is an ITR of goose parvovirus (GPV), Muscovy duck parvovirus (MDPV), or parvovirus B19 of the genus Erythrovirus (also known as parvovirus B19, primate erythroparvovirus 1, B19 virus, and Erythrovirus). In certain embodiments, one of the two ITRs is an ITR of AAV. In other embodiments, one of the two ITRs in the construct is an ITR of an AAV serotype selected from serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and any combination thereof. In a particular embodiment, the ITR is derived from AAV serotype 2, e.g., the ITR of AAV serotype 2.
[0062] In certain aspects of the present disclosure, a nucleic acid molecule contains two ITRs, a 5' ITR and a 3' ITR, where the 5' ITR is located at the 5' end of the nucleic acid molecule and the 3' ITR is located at the 3' end of the nucleic acid molecule. The 5' ITR and the 3' ITR can be derived from the same virus or different viruses. In certain embodiments, the 5' ITR is derived from AAV and the 3' ITR is not derived from an AAV virus (e.g., non-AAV). In some embodiments, the 3' ITR is derived from AAV and the 5' ITR is not derived from an AAV virus (e.g., non-AAV). In other embodiments, the 5' ITR is not derived from an AAV virus (e.g., non-AAV), and the 3' ITR is derived from the same or a different non-AAV virus.
[0063] As used herein, the term "parvovirus" encompasses the family Parvoviridae, including but not limited to the autonomous parvoviruses and dependoviruses. Autonomous parvoviruses include, for example, members of the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstyldensovirus.
[0064] Exemplary autonomous parvoviruses include, but are not limited to, porcine parvovirus, murine parvovirus, canine parvovirus, mink enterovirus, bovine parvovirus, chicken parvovirus, feline panleukopeniavirus, feline parvovirus, goose parvovirus, H1 parvovirus, Muscovy duck parvovirus, snake parvovirus, and B19 virus. Other autonomous parvoviruses are known to those of skill in the art. See, e.g., VIROLOGY, Volume 2, Chapter 69 (4th Edition, Lippincott-Raven Publishers).
[0065] As used herein, the term "non-AAV" encompasses nucleic acids, proteins, and viruses from the Parvoviridae family other than any adeno-associated virus (AAV) of the Parvoviridae family. "Non-AAV" includes but is not limited to autonomous replicating members of the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstyldensovirus.
[0066] As used herein, the term "adeno-associated virus" (AAV) includes, but is not limited to, AAV serotype 1, AAV serotype 2, AAV serotype 3 (including subtypes 3A and 3B), AAV serotype 4, AAV serotype 5, AAV serotype 6, AAV serotype 7, AAV serotype 8, AAV serotype 9, AAV serotype 10, AAV serotype 11, AAV serotype 12, AAV serotype 13, snake AAV, avian AAV, bovine AAV, canine AAV, equine AAV, ovine AAV, caprine AAV, shrimp AAV (these AAV serotypes and clades are disclosed by Gao et al. (J. Virol. 78:6381 (2004)) and Moris et al. (Virol. 33:375 (2004))) and any other AAV now known or later discovered. See, e.g., FIELDS et al VIROLOGY, Volume 2, Chapter 69 (4th Edition, Lippincott-Raven Publishers).
[0067] As used herein, the term "derived from" means a component isolated from or made using a specified molecule or organism, or information (such as an amino acid or nucleic acid sequence) from a specified molecule or organism. For example, a nucleic acid sequence (such as an ITR) derived from a second nucleic acid sequence (such as an ITR) may include a nucleotide sequence that is identical or substantially similar to the nucleotide sequence of the second nucleic acid sequence. In the case of a nucleotide or polypeptide, derivative species can be obtained, for example, by naturally occurring mutagenesis, artificial directed mutagenesis, or artificial random mutagenesis. The mutagenesis used to derive a nucleotide or polypeptide can be intentionally directed or intentionally random, or a mixture of each case. Mutagenizing a nucleotide or polypeptide to produce a different nucleotide or polypeptide derived from a first nucleotide or polypeptide can be a random event (such as caused by polymerase infidelity), and the identification of the derived nucleotide or polypeptide can be carried out by appropriate screening methods, such as the methods described herein. Mutagenesis of a polypeptide generally requires manipulation of the polynucleotide encoding the polypeptide. In some embodiments, a nucleotide or amino acid sequence derived from a second nucleotide or amino acid sequence has at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity with the second nucleotide or amino acid, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In other embodiments, an ITR derived from a non-AAV (or AAV) has at least 90% identity with a non-AAV ITR (or AAV ITR), wherein the non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or AAV ITR). In some embodiments, an ITR derived from a non-AAV (or AAV) ITR has at least 80% identity with a non-AAV ITR (or AAV ITR), wherein the non-AAV (or AAV) ITR retains the functional properties of the non-AAV ITR (or AAV ITR).In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 70% identical to non-AAV ITRs (or AAV ITRs), respectively, wherein the non-AAV (or AAV) ITRs retain the functional characteristics of the non-AAV ITRs (or AAV ITRs), respectively. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 60% identical to non-AAV ITRs (or AAV ITRs), respectively, wherein the non-AAV (or AAV) ITRs retain the functional characteristics of the non-AAV ITRs (or AAV ITRs), respectively. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 50% identical to non-AAV ITRs (or AAV ITRs), respectively, wherein the non-AAV (or AAV) ITRs retain the functional characteristics of the non-AAV ITRs (or AAV ITRs), respectively.
[0068] In certain embodiments, the ITRs derived from non-AAV (or AAV) ITRs comprise or consist of fragments of non-AAV (or AAV) ITRs. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs comprise or consist of fragments of non-AAV (or AAV) ITRs, wherein the fragments comprise at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides or at least about 600 nucleotides; wherein the ITRs derived from non-AAV (or AAV) ITRs respectively retain the functional properties of non-AAV ITRs (or AAV ITRs). In certain embodiments, the ITRs derived from non-AAV (or AAV) ITRs comprise or consist of fragments of non-AAV (or AAV) ITRs, wherein the fragments comprise at least about 129 nucleotides, and wherein the ITRs derived from non-AAV (or AAV) ITRs respectively retain the functional properties of non-AAV ITRs (or AAV ITRs). In certain embodiments, the ITRs derived from non-AAV (or AAV) ITRs comprise or consist of fragments of non-AAV (or AAV) ITRs, wherein the fragments comprise at least about 102 nucleotides, and wherein the ITRs derived from non-AAV (or AAV) ITRs respectively retain the functional properties of non-AAV ITRs (or AAV ITRs).
[0069] In some embodiments, an ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of the non-AAV (or AAV) ITR, wherein the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the length of the non-AAV (or AAV) ITR.
[0070] In certain embodiments, when properly aligned, the nucleotide or amino acid sequence derived from the second nucleotide or amino acid sequence has at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity with the homologous portion of the second nucleotide or amino acid sequence, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In other embodiments, when properly aligned, the ITRs derived from non-AAV (or AAV) ITRs are at least 90% identical to the homologous portions of non-AAV ITRs (or AAV ITRs), wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, when properly aligned, the ITRs derived from non-AAV (or AAV) ITRs are at least 80% identical to the homologous portions of non-AAV ITRs (or AAV ITRs), wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, when properly aligned, the ITRs derived from non-AAV (or AAV) ITRs are at least 70% identical to the homologous portions of non-AAV ITRs (or AAV ITRs), wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 60% identical to the homologous portions of non-AAV ITRs (or AAV ITRs), wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 50% identical to the homologous portions of non-AAV ITRs (or AAV ITRs), wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence.
[0071] A "capsid-free" or "capsid-less" vector or nucleic acid molecule refers to a capsid-free vector construct. In some embodiments, the capsid-less vector or nucleic acid molecule does not contain a sequence encoding, for example, an AAV Rep protein.
[0072] As used herein, a "coding region" or "coding sequence" is a portion of a polynucleotide that consists of codons that are translated into amino acids. Although "stop codons" (TAG, TGA or TAA) are generally not translated into amino acids, they can be considered part of the coding region, but any flanking sequences (e.g., promoters, ribosome binding sites, transcription terminators, introns, etc.) are not part of the coding region. The boundaries of the coding region are generally determined by a start codon at the 5' end that encodes the amino terminus of the resulting polypeptide and a translation stop codon at the 3' end that encodes the carboxyl terminus of the resulting polypeptide. Two or more coding regions can be present in a single polynucleotide construct, e.g., on a single vector, or in separate polynucleotide constructs, e.g., on separate (different) vectors. Thus, a single vector can contain only a single coding region or contain two or more coding regions.
[0073] Certain proteins secreted by mammalian cells are associated with a secretory signal peptide that is cleaved from the mature protein once the growing protein chain begins to be exported through the rough endoplasmic reticulum. Those of ordinary skill in the art know that signal peptides are generally fused to the N-terminus of a polypeptide and are cleaved from the intact or "full-length" polypeptide to produce the secreted or "mature" form of the polypeptide. In certain embodiments, a native signal peptide or a functional derivative of a sequence that retains the ability to direct polypeptide secretion can be operably associated therewith. Alternatively, a heterologous mammalian signal peptide (e.g., human tissue plasminogen activator (TPA) or mouse β-glucuronidase signal peptide) or a functional derivative thereof can be used.
[0074] The term "downstream" refers to a nucleotide sequence located 3' to a reference nucleotide sequence. In certain embodiments, the downstream nucleotide sequence refers to a sequence after the transcription start point. For example, the translation start codon of a gene is located downstream of the transcription start site.
[0075] The term "upstream" refers to a nucleotide sequence located 5' to a reference nucleotide sequence. In certain embodiments, the upstream nucleotide sequence refers to a sequence located 5' to a coding region or the transcription start point. For example, most promoters are located upstream of the transcription start site.
[0076] As used herein, the term "gene regulatory region" or "regulatory region" refers to nucleotide sequences located upstream (5' non-coding sequence) of the coding region, within the coding region, or downstream (3' non-coding sequence) of the coding region that affect the transcription, RNA processing, stability, and translation of the associated coding region. Regulatory regions can include promoters, translational leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites, or stem-loop structures. If the coding region is intended to be expressed in a eukaryotic cell, the polyadenylation signal and transcription termination sequence will typically be located 3' of the coding sequence.
[0077] A polynucleotide encoding a product (e.g., miRNA or a gene product (e.g., a polypeptide such as a therapeutic protein)) can include a promoter and / or other expression (e.g., transcriptional or translational) control elements operably associated with one or more coding regions. In an operable association, the coding region of a gene product (e.g., a polypeptide) is associated with one or more regulatory regions such that the expression of the gene product is under the influence or control of the regulatory region. For example, if induction of promoter function results in transcription of an mRNA encoding the gene product encoded by the coding region, and if the nature of the linkage between the promoter and the coding region does not interfere with the ability of the promoter to direct the expression of the gene product or interfere with the ability of the DNA template to be transcribed, then the coding region and the promoter are "operably associated". In addition to promoters, other expression control elements, such as enhancers, operators, repressors, and transcription termination signals, can also be operably associated with the coding region to direct the expression of the gene product.
[0078] "Transcription control sequences" refer to DNA regulatory sequences that provide for the expression of a coding sequence in a host cell, such as promoters, enhancers, terminators, etc. A variety of transcription control regions are known to those of ordinary skill in the art. These transcription control regions include, but are not limited to, transcription control regions that function in vertebrate cells, such as, but not limited to, promoters and enhancer segments from cytomegalovirus (immediate early promoter, associated with intron A), simian virus 40 (early promoter), and retroviruses (such as Rous sarcoma virus). Other transcription control regions include those derived from vertebrate genes (such as actin, heat shock protein, bovine growth hormone, and rabbit β-globin) and other sequences capable of controlling gene expression in eukaryotic cells. Additionally suitable transcription control regions include tissue-specific promoters and enhancers, as well as lymphokine-inducible promoters (e.g., promoters that can be induced by interferons or interleukins).
[0079] Similarly, a variety of translation control elements are known to those of ordinary skill in the art. These translation control elements include, but are not limited to, ribosome binding sites, translation initiation and termination codons, and elements derived from picornaviruses (particularly internal ribosome entry sites or IRESs, also known as CITE sequences).
[0080] As used herein, the term "expression" refers to the process by which a polynucleotide gives rise to a gene product (e.g., RNA or polypeptide). It includes, but is not limited to, transcription of a polynucleotide into messenger RNA (mRNA), transfer RNA (tRNA), small hairpin RNA (shRNA), small interfering RNA (siRNA), or any other RNA product, and translation of the mRNA into a polypeptide. Expression gives rise to a "gene product". As used herein, a gene product can be a nucleic acid, such as messenger RNA produced by transcription of a gene, or can be a polypeptide translated from a transcript. Gene products described herein also include nucleic acids having post-transcriptional modifications (e.g., polyadenylation or splicing), or polypeptides having post-translational modifications (e.g., methylation, glycosylation, addition of lipids, association with other protein subunits, or proteolytic cleavage). As used herein, the term "yield" refers to the amount of polypeptide produced by expression of a gene.
[0081] "Vector" refers to any vehicle used for cloning and / or transferring nucleic acids into a host cell. A vector can be a replicon to which another nucleic acid segment can be attached so as to bring about the replication of the attached segment. A "replicon" is any genetic element (e.g., plasmid, phage, cosmid, chromosome, virus) that functions as an autonomous unit of replication in vivo, i.e., that is capable of replication under its own control. The term "vector" includes agents for introducing nucleic acids into cells in vitro, ex vivo, or in vivo. A large number of vectors are known and used in the art, including, for example, plasmids, modified eukaryotic viruses, or modified bacterial viruses. Insertion of a polynucleotide into a suitable vector can be accomplished by ligating the appropriate polynucleotide fragment into the selected vector having complementary sticky ends.
[0082] Vectors can be engineered to encode a selectable marker or reporter molecule to provide selection or identification of cells that have incorporated the vector. Expression of the selectable marker or reporter molecule permits identification and / or selection of host cells that have incorporated and express other coding regions contained on the vector. Examples of selectable marker genes known and used in the art include: genes that confer resistance to ampicillin, streptomycin, gentamicin, kanamycin, hygromycin, bialaphos herbicide, sulfonamides, etc.; and genes that are used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentyl transferase genes, etc. Examples of reporter molecules known and used in the art include: luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), β-galactosidase (LacZ), β-glucuronidase (Gus), etc. Selectable markers can also be considered reporter molecules.
[0083] As used herein, the term "host cell" refers to, for example, microbial, yeast, insect, and mammalian cells that can or have been used as recipients of ssDNA or a vector. The term includes progeny of the original cells that have been transduced. Thus, a "host cell" as used herein generally refers to a cell that has been transduced with an exogenous DNA sequence. It should be understood that the progeny of a single parental cell may not be identical in morphology or genomic or total DNA complement to the original parent due to natural, accidental, or deliberate mutations. In some embodiments, the host cell can be an in vitro host cell.
[0084] The term "selectable marker" refers to an identifying factor that can be selected based on the action of a marker gene (i.e., resistance to an antibiotic, resistance to a herbicide, colorimetric marker, enzyme, fluorescent marker, etc.), typically an antibiotic or chemical resistance gene, wherein the action is used to trace the inheritance of a nucleic acid of interest and / or to identify cells or organisms that have inherited the nucleic acid of interest. Examples of selectable marker genes known and used in the art include: genes that confer resistance to ampicillin, streptomycin, gentamicin, kanamycin, hygromycin, bialaphos herbicide, sulfonamide, etc.; and genes that are used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentenyl transferase genes, etc.
[0085] The term "reporter gene" refers to a nucleic acid encoding an identifying factor that can be identified based on the action of the reporter gene, wherein the action is used to trace the inheritance of a nucleic acid of interest, to identify cells or organisms that have inherited the nucleic acid of interest, and / or to measure gene expression induction or transcription. Examples of reporter genes known and used in the art include: luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), β-galactosidase (LacZ), β-glucuronidase (Gus), etc. A selectable marker gene can also be considered a reporter gene.
[0086] "Promoter" and "promoter sequence" are used interchangeably and refer to a DNA sequence capable of controlling the expression of a coding sequence or functional RNA. Typically, the coding sequence is located 3' to the promoter sequence. A promoter may be derived intact from a native gene, or be composed of different elements derived from different promoters found in nature, or even contain synthetic DNA segments. Those skilled in the art will understand that different promoters may direct gene expression in different tissues or cell types, or at different developmental stages, or in response to different environmental or physiological conditions. A promoter that causes a gene to be expressed in most cell types most of the time is often referred to as a "constitutive promoter". A promoter that causes a gene to be expressed in a specific cell type is often referred to as a "cell-specific promoter" or "tissue-specific promoter". A promoter that causes a gene to be expressed at a specific stage of development or cell differentiation is often referred to as a "development-specific promoter" or "cell differentiation-specific promoter". A promoter that induces and causes gene expression after a cell is exposed to or treated with an agent, biomolecule, chemical, ligand, light, etc. that induces the promoter is often referred to as an "inducible promoter" or "regulatable promoter". It is further recognized that, since the exact boundaries of regulatory sequences have not been fully determined in most cases, DNA fragments of different lengths may have the same promoter activity.
[0087] A promoter sequence is typically delimited at its 3' end by the transcription start site and extends upstream (in the 5' direction) to include the minimal bases or elements required to initiate transcription at a detectable level above background. The transcription start site (e.g., conveniently defined by nuclease S1 mapping) will be found within the promoter sequence, as well as the protein-binding domain (consensus sequence) responsible for binding RNA polymerase.
[0088] In some embodiments, the nucleic acid molecule comprises a tissue-specific promoter. In certain embodiments, the tissue-specific promoter drives the expression of a therapeutic protein (e.g., a clotting factor) in the liver (e.g., in hepatocytes and / or endothelial cells). In specific embodiments, the promoter is selected from the mouse thyroxine promoter (mTTR), the endogenous human factor VIII promoter (F8), the human α-1-antitrypsin promoter (hAAT), the human albumin minimal promoter, the mouse albumin promoter, the Tristetraprolin (TTP) promoter, the CASI promoter, the CAG promoter, the cytomegalovirus (CMV) promoter, the phosphoglycerate kinase (PGK) promoter, and any combination thereof. In some embodiments, the promoter is selected from liver-specific promoters (e.g., α1-antitrypsin (AAT)), muscle-specific promoters (e.g., muscle creatine kinase (MCK), myosin heavy chain α (αMHC), myoglobin (MB), and desmin (DES)), synthetic promoters (e.g., SPc5-12, 2R5Sc5-12, dMCK, and tMCK), and any combination thereof. In a specific embodiment, the promoter comprises the TTP promoter.
[0089] The terms "restriction endonuclease" and "restriction enzyme" are used interchangeably and refer to enzymes that bind and cut within a specific nucleotide sequence in double-stranded DNA.
[0090] The term "plasmid" refers to an extrachromosomal element that usually carries genes not belonging to the central metabolism of the cell and is usually in the form of a circular double-stranded DNA molecule. Such elements can be linear, circular, or supercoiled autonomously replicating sequences, genomic integration sequences, phages, or nucleotide sequences of single-stranded or double-stranded DNA or RNA derived from any source, many of which have been ligated or recombined into unique constructs that are capable of introducing a promoter fragment and a DNA sequence of a selected gene product, along with appropriate 3' untranslated sequences, into a cell.
[0091] Eukaryotic viral vectors that can be used include, but are not limited to, adenoviral vectors, retroviral vectors, adeno-associated viral vectors, poxviral (e.g., vaccinia virus) vectors, baculoviral vectors, or herpesviral vectors. Non-viral vectors include plasmids, liposomes, charged lipids (cytofectin), DNA-protein complexes, and biopolymers.
[0092] "Cloning vector" refers to a "replicon", which is a unit length of nucleic acid that replicates sequentially and which contains an origin of replication, such as a plasmid, phage or cosmid, to which another nucleic acid segment can be attached so as to bring about the replication of the attached segment. Some cloning vectors are capable of replicating in one cell type (e.g., bacteria) and expressing in another cell type (e.g., eukaryotic cells). Cloning vectors typically contain one or more sequences that can be used to select cells containing the vector and / or one or more multiple cloning sites for insertion of nucleic acid sequences of interest.
[0093] The term "expression vector" refers to a vehicle designed to be capable of expressing an inserted nucleic acid sequence after insertion into a host cell. The inserted nucleic acid sequence is operably associated with regulatory regions as described above.
[0094] Vectors are introduced into host cells by methods well known in the art, such as by transfection, electroporation, microinjection, transduction, cell fusion, DEAE-dextran, calcium phosphate precipitation, lipofection (lysosome fusion), use of a gene gun or a DNA vector transporter. As used herein, "culturing", "to be cultured" and "being cultured" mean incubating cells under in vitro conditions that permit the cells to grow or divide or to maintain the cells in a viable state. As used herein, "cultured cells" mean cells propagated in vitro.
[0095] As used herein, the term "polypeptide" is intended to cover both the singular "polypeptide" and the plural "polypeptides", and refers to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also known as peptide bonds). The term "polypeptide" refers to any one or more chains of two or more amino acids and does not refer to a specific length of the product. Thus, included in the definition of "polypeptide" are peptides, dipeptides, tripeptides, oligopeptides, "proteins", "amino acid chains" or any other term used to refer to one or more chains of two or more amino acids. The term "polypeptide" can be used in place of, or be used interchangeably with, any of these terms. The term "polypeptide" is also intended to refer to products of post-expression modification of polypeptides, which include but are not limited to glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting / blocking groups, proteolytic cleavage or modification by non-naturally occurring amino acids. Polypeptides can be derived from natural biological sources or produced by recombinant techniques, but are not necessarily translated from a designated nucleic acid sequence. It can be produced in any manner, including by chemical synthesis.
[0096] The term "amino acid" includes alanine (Ala or A); arginine (Arg or R); asparagine (Asn or N); aspartic acid (Asp or D); cysteine (Cys or C); glutamine (Gln or Q); glutamic acid (Glu or E); glycine (Gly or G); histidine (His or H); isoleucine (Ile or I); leucine (Leu or L); lysine (Lys or K); methionine (Met or M); phenylalanine (Phe or F); proline (Pro or P); serine (Ser or S); threonine (Thr or T); tryptophan (Trp or W); tyrosine (Tyr or Y); and valine (Val or V). Non-conventional amino acids are also within the scope of the present disclosure and include norleucine, ornithine, norvaline, homoserine, and other amino acid residue analogs, such as those described in Ellman et al., Meth. Enzym. 202:301-336 (1991). To generate such non-naturally occurring amino acid residues, the procedures of Noren et al., Science 244:182 (1989) and Ellman et al. can be used. Briefly, these procedures involve chemically activating suppressor tRNA with non-naturally occurring amino acid residues and then transcribing and translating the RNA in vitro. The introduction of non-conventional amino acids can also be achieved using peptide chemistry known in the art. As used herein, the term "polar amino acid" includes amino acids that have a net zero charge but have non-zero partial charges in different parts of their side chains (e.g., M, F, W, S, Y, N, Q, C). These amino acids can participate in hydrophobic interactions and electrostatic interactions. As used herein, the term "charged amino acid" includes amino acids that can have a non-zero net charge on their side chains (e.g., R, K, H, E, D). These amino acids can participate in hydrophobic interactions and electrostatic interactions.
[0097] Also included in the present disclosure are fragments or variants of polypeptides, and any combination thereof. When referring to the polypeptide binding domains or binding molecules of the present disclosure, the terms "fragment" or "variant" include those that retain at least some properties of the reference polypeptide (e.g., the FcRn binding affinity for the FcRn binding domain or Fc variant, the coagulation activity of the FVIII variant, or the FVIII binding activity of the VWF fragment). In addition to the specific antibody fragments discussed elsewhere herein, fragments of polypeptides also include proteolytic fragments and deletion fragments, but do not include naturally occurring full-length polypeptides (or mature polypeptides). Variants of the polypeptide binding domains or binding molecules of the present disclosure include the fragments described above, as well as polypeptides having an altered amino acid sequence due to amino acid substitutions, deletions, or insertions. Variants can be natural or non-natural. Non-naturally occurring variants can be generated using mutagenesis techniques known in the art. Variant polypeptides can contain conservative or non-conservative amino acid substitutions, deletions, or additions.
[0098] "Conservative amino acid substitution" is an amino acid in which an amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art and include basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), β-branched side chains (e.g., threonine, valine, isoleucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Thus, if an amino acid in a polypeptide is replaced with another amino acid from the same side chain family, the substitution is considered conservative. In another embodiment, a string of amino acids may be conservatively replaced with a structurally similar string having a different order and / or composition of side chain family members.
[0099] As is known in the art, the term "percent identity" is a relationship between two or more polypeptide sequences or two or more polynucleotide sequences, as determined by comparing the sequences. In the art, "identity" also refers to the degree of sequence relatedness between polypeptide or polynucleotide sequences, as the case may be, as determined by the match between strings of such sequences. "Identity" can be readily calculated by known methods, including but not limited to those described in the following: Computational Molecular Biology (Lesk, A.M. ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, D.W. ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, A.M. and Griffin, H.G. eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G. ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J. eds.) Stockton Press, New York (1991). Preferred methods for determining identity are designed to give the best match between the sequences being tested. Methods for determining identity have been incorporated into publicly available computer programs. Sequence alignments and percent identity calculations can be performed using sequence analysis software such as the Megalign program of the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, WI), the GCG program suite (Wisconsin Package version 9.0, Genetics Computer Group (GCG), Madison, WI), BLASTP, BLASTN, BLASTX (Altschul et al., J. Mol. Biol. 215:403 (1990)) and DNASTAR (DNASTAR Inc., 1228 S. Park St. Madison, WI 53715 USA). In the context of the present application, it will be understood that in cases where sequence analysis software is used for analysis, unless otherwise stated, the results of the analysis will be based on the "default values" of the program cited. As used herein, "default values" refers to any set of values or parameters that are initially loaded with the software upon first initialization.For the purpose of determining the percent identity between the sequence of a therapeutic protein (e.g., a clotting factor) of the present disclosure and a reference sequence, only the nucleotides in the reference sequence that correspond to the nucleotides in the therapeutic protein (e.g., a clotting factor) of the present disclosure are used to calculate the percent identity. For example, when the full-length FVIII nucleotide sequence containing the B domain is compared to the optimized B domain-deleted (BDD) FVIII nucleotide sequence of the present disclosure, a partial alignment including the A1, A2, A3, C1, and C2 domains will be used to calculate the percent identity. Nucleotides in the portion of the full-length FVIII sequence encoding the B domain (which would result in large "gaps" in the alignment) will not be counted as mismatches. In addition, when determining the percent identity between the optimized BDD FVIII sequence of the present disclosure or a specified portion thereof (e.g., nucleotides 58-2277 and 2320-4374 of SEQ ID NO:3) and a reference sequence, the percent identity will be calculated by dividing the number of matching nucleotides by the total number of nucleotides in the complete sequence of the optimized BDD-FVIII sequence or its specified portion, as described herein.
[0100] As used herein, nucleotides corresponding to nucleotides in a particular sequence of the present invention are identified by aligning the sequences of the present invention to maximize identity with a reference sequence. The numbers used to identify equivalent amino acids in the reference sequence are based on the numbers used to identify the corresponding amino acids in the sequences of the present disclosure.
[0101] A "fusion" or "chimeric" protein comprises a first amino acid sequence linked to a second amino acid sequence, where the second amino acid sequence is not naturally linked thereto. Amino acid sequences that are normally present in separate proteins can be placed together in a fusion polypeptide, or amino acid sequences that are normally present in the same protein can be placed in a new arrangement in a fusion polypeptide, such as the fusion of a factor VIII domain of the present invention with an Ig Fc domain. Fusion proteins are produced, for example, by chemical synthesis or by generating and translating a polynucleotide encoding the peptide regions in the desired relationship. A chimeric protein can further comprise a second amino acid sequence associated with the first amino acid sequence by covalent, non-peptide, or non-covalent bonds.
[0102] As used herein, the term "insertion site" refers to the position in a polypeptide or a fragment, variant, or derivative thereof that is immediately upstream of the position where a heterologous moiety can be inserted. The "insertion site" is designated by a number that is the number of the amino acid in the reference sequence. For example, the "insertion site" in FVIII refers to the number of the amino acid sequence in mature native FVIII (SEQ ID NO:15) corresponding to the insertion site, which is immediately N-terminal to the insertion position. For example, the phrase "a3 contains a heterologous moiety at the insertion site corresponding to amino acid 1656 of SEQ ID NO:15" means that the heterologous moiety is located between the two amino acids corresponding to amino acid 1656 and amino acid 1657 of SEQ ID NO:15.
[0103] As used herein, the phrase "downstream of an amino acid" refers to the position immediately adjacent to the terminal carboxyl group of the amino acid. Similarly, the phrase "upstream of an amino acid" refers to the position immediately adjacent to the terminal amino group of the amino acid.
[0104] As used herein, the terms "inserted", "is inserted", "inserted into...", or grammatically related terms refer to the position of a heterologous moiety in a polypeptide (e.g., a clotting factor) relative to a similar position in the parental polypeptide. For example, in certain embodiments, "inserted" and the like refer to the position of a heterologous moiety in a recombinant FVIII polypeptide relative to a similar position in native mature human FVIII. As used herein, the term refers to a characteristic of the polypeptide and does not indicate, imply, or infer any method or process for preparing the polypeptide.
[0105] As used herein, the term "half-life" refers to the biological half-life of a particular polypeptide in vivo. The half-life can be represented by the time required to clear half of the amount administered to a subject from the circulation and / or other tissues in an animal. When the clearance rate curve of a given polypeptide is constructed as a function of time, the curve is typically biphasic, consisting of a rapid α-phase and a longer β-phase. The α-phase generally represents the equilibrium between the administered Fc polypeptide in the intravascular and extravascular spaces and is partially determined by the size of the polypeptide. The β-phase generally represents the metabolism of the polypeptide in the intravascular space. In some embodiments, therapeutic proteins (e.g., clotting factors, e.g., FVIII) and chimeric proteins containing such proteins are monophasic and thus do not have an α-phase but only a single β-phase. Thus, in certain embodiments, the term half-life as used herein refers to the half-life of the polypeptide in the β-phase.
[0106] As used herein, the term "linked" refers to a first amino acid sequence or nucleotide sequence that is covalently or non-covalently linked to a second amino acid sequence or nucleotide sequence, respectively. The first amino acid or nucleotide sequence can be directly linked or juxtaposed to the second amino acid or nucleotide sequence, or alternatively, an intervening sequence can covalently link the first sequence to the second sequence. The term "linked" not only means fusing the first amino acid sequence to the second amino acid sequence at the C-terminus or N-terminus, but also includes inserting the entire first amino acid sequence (or second amino acid sequence) between two amino acids of the second amino acid sequence (or first amino acid sequence), respectively. In one embodiment, the first amino acid sequence can be linked to the second amino acid sequence by a peptide bond or a linker. The first nucleotide sequence can be linked to the second nucleotide sequence by a phosphodiester bond or a linker. The linker can be a peptide or polypeptide (for polypeptide chains) or a nucleotide or nucleotide chain (for nucleotide chains) or any chemical moiety (for both polypeptide and polynucleotide chains). The term "linked" is also denoted by a hyphen (-).
[0107] Hemostasis, as used herein, means to stop or slow bleeding or hemorrhage; or to prevent or slow blood flow through a blood vessel or body part.
[0108] A hemostatic disorder, as used herein, means a genetic or acquired condition characterized by a tendency to bleed spontaneously or due to trauma due to impaired ability to form fibrin clots or inability to form fibrin clots. Examples of such disorders include hemophilia. The three main forms are hemophilia A (factor VIII deficiency), hemophilia B (factor IX deficiency or "Christmas disease"), and hemophilia C (factor XI deficiency, mild bleeding tendency). Other hemostatic disorders include, for example, von Willebrand disease, factor XI deficiency (PTA deficiency), factor XII deficiency, deficiency or structural abnormality of fibrinogen, prothrombin, factor V, factor VII, factor X, or factor XIII, Bernard-Soulier syndrome, which is a defect or deficiency of GPIb. GPIb, the receptor for vWF, can be defective and result in a lack of primary clot formation (primary hemostasis) and an increased bleeding tendency), and Glanzmann and Naegeli thrombasthenia (Glanzmann thrombasthenia). In liver failure (acute and chronic forms), the liver produces insufficient amounts of clotting factors; this increases the risk of bleeding.
[0109] The isolated nucleic acid molecules, isolated polypeptides, or vectors comprising the isolated nucleic acid molecules of the present disclosure can be used prophylactically. As used herein, the term "prophylactic treatment" refers to the administration of the molecules prior to the onset of bleeding. In one embodiment, a subject in need of a general hemostatic agent is undergoing or about to undergo surgery. The polynucleotides, polypeptides, or vectors of the present disclosure can be administered before or after surgery as a prophylactic agent. The polynucleotides, polypeptides, or vectors of the present disclosure can be administered during or after surgery to control an acute bleeding episode. Surgery can include, but is not limited to, liver transplantation, hepatectomy, dental surgery, or stem cell transplantation.
[0110] The isolated nucleic acid molecules, isolated polypeptides, or vectors of the present disclosure are also used for treatment on demand. The term "treatment on demand" refers to the administration of the isolated nucleic acid molecules, isolated polypeptides, or vectors in response to the symptoms of a bleeding episode or prior to an activity that can cause bleeding. In one aspect, treatment on demand can be administered to a subject when bleeding begins, such as after an injury, or when bleeding is anticipated, such as before surgery. In another aspect, treatment on demand can be administered prior to an activity that increases the risk of bleeding, such as contact sports.
[0111] As used herein, the term "acute bleeding" refers to the onset of bleeding, regardless of its underlying cause. For example, a subject can have trauma, uremia, a hereditary bleeding disorder (e.g., factor VII deficiency), a platelet disorder, or resistance due to the production of antibodies against a clotting factor.
[0112] "Treatment" (treat, treatment, treating), as used herein, refers to, for example, a reduction in the severity of a disease or condition; a decrease in the duration of the course of the disease; an improvement in one or more symptoms associated with the disease or condition; and providing a beneficial effect to a subject having the disease or condition without necessarily curing the disease or condition or preventing one or more symptoms associated with the disease or condition. In one embodiment, the term "treatment" means maintaining in a subject, by administration of an isolated nucleic acid molecule, an isolated polypeptide, or a vector of the present disclosure, for example, an FVIII trough level of at least about 1 IU / dL, 2 IU / dL, 3 IU / dL, 4 IU / dL, 5 IU / dL, 6 IU / dL, 7 IU / dL, 8 IU / dL, 9 IU / dL, 10 IU / dL, 11 IU / dL, 12 IU / dL, 13 IU / dL, 14 IU / dL, 15 IU / dL, 16 IU / dL, 17 IU / dL, 18 IU / dL, 19 IU / dL, 20 IU / dL, 25 IU / dL, 30 IU / dL, 35 IU / dL, 40 IU / dL, 45 IU / dL, 50 IU / dL, 55 IU / dL, 60 IU / dL, 65 IU / dL, 70 IU / dL, 75 IU / dL, 80 IU / dL, 85 IU / dL, 90 IU / dL, 95 IU / dL, 100 IU / dL, 105 IU / dL, 110 IU / dL, 115 IU / dL, 120 IU / dL, 125 IU / dL, 130 IU / dL, 135 IU / dL, 140 IU / dL, 145 IU / dL, or 150 IU / dL. In another embodiment, treatment means maintaining the FVIII trough level at about 1 - about 150 IU / dL, about 1 - about 125 IU / dL, about 1 - about 100 IU / dL, about 1 - about 90 IU / dL, about 1 - about 85 IU / dL, about 1 - about 80 IU / dL, about 1 - about 75 IU / dL, about 1 - about 70 IU / dL, about 1 - about 65 IU / dL, about 1 - about 60 IU / dL, about 1 - about 55 IU / dL, about 1 - about 50 IU / dL, about 1 - about 45 IU / dL, about 1 - about 40 IU / dL, about 1 - about 35 IU / dL, about 1 - about 30 IU / dL, about 1 - about 25 IU / dL, about 25 - about 125 IU / dL, about 50 - about 100 IU / dL, about 50 - about 75 IU / dL, about 75 - about 100 IU / dL, about 1 - about 20 IU / dL, about 2 - about 20 IU / dL, about 3 - about 20 IU / dL, about 4 - about 20 IU / dL, about 5 - about 20 IU / dL, about 6 - about 20 IU / dL, about 7 - about 20 IU / dL, about 8 - about 20 IU / dL, about 9 - about 20 IU / dL, or about 10 - about 20 IU / dL.Treatment of a disease or condition may also include maintaining FVIII activity in a subject at a level equivalent to at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 105%, 110%, 115%, 120%, 125%, 130%, 135%, 140%, 145% or 150% of FVIII activity in a non-hemophilic subject. The minimum trough level required for treatment can be measured by one or more known methods and can be adjusted (increased or decreased) for each individual.
[0113] As used herein, "administering" means giving a pharmaceutically acceptable nucleic acid molecule, a polypeptide expressed therefrom, or a vector comprising a nucleic acid molecule of the present disclosure to a subject via a pharmaceutically acceptable route. The route of administration can be intravenous, e.g., intravenous injection and intravenous infusion. Additional routes of administration include, e.g., subcutaneous, intramuscular, oral, nasal, and pulmonary administration. The nucleic acid molecule, polypeptide, and vector can be administered as part of a pharmaceutical composition comprising at least one excipient.
[0114] As used herein, the term "pharmaceutically acceptable" refers to molecular entities and compositions that are physiologically tolerable and that typically do not produce toxicity, allergic reactions, or similar adverse reactions (such as gastric upset, dizziness, etc.) when administered to a human. Optionally, as used herein, the term "pharmaceutically acceptable" refers to those approved by a regulatory agency of the federal or state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeias for animals, and more particularly for humans.
[0115] As used herein, the phrase "subject in need thereof" includes a subject who would benefit from administration of a nucleic acid molecule, polypeptide, or vector of the present disclosure, such as a mammalian subject, e.g., to improve hemostasis. In one embodiment, the subject includes, but is not limited to, an individual with hemophilia. In another embodiment, the subject includes, but is not limited to, an individual who has developed an inhibitor to a therapeutic protein (e.g., a clotting factor, e.g., FVIII) and thus requires bypass therapy. The subject can be an adult or a minor (e.g., under 12 years old).
[0116] As used herein, the term "therapeutic protein" refers to any polypeptide known in the art that is administered to a subject. In some embodiments, the therapeutic protein comprises a protein selected from a coagulation factor, a growth factor, an antibody, a functional fragment thereof, or a combination thereof. As used herein, the term "coagulation factor" refers to a naturally occurring or recombinantly produced molecule or an analogue thereof that prevents or reduces the duration of bleeding episodes in a subject. In other words, it refers to a molecule having procoagulant activity, i.e., responsible for converting fibrinogen into a meshwork of insoluble fibrin, thereby causing the blood to coagulate or clot. As used herein, "coagulation factors" include activated coagulation factors, their zymogens, or coagulation factors that can be activated. An "activatable coagulation factor" is a coagulation factor that can be converted into an activated form (e.g., its zymogen form) from a non-activated form. The term "coagulation factor" includes, but is not limited to, factor I (FI), factor II (FII), factor V (FV), FVII, FVIII, FIX, factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FXIII), von Willebrand factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, α2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), their zymogens, their activated forms, or any combination thereof.
[0117] As used herein, procoagulant activity means the ability to participate in a cascade of biochemical reactions that culminates in the formation of a fibrin clot and / or reduces the severity, duration, or frequency of bleeding or bleeding episodes.
[0118] "Growth factor" as used herein includes any growth factor known in the art, which includes growth factors and hormones.In some embodiments, the growth factor is selected from adrenomedullin (AM), angiopoietin (Ang), autocrine motility factor, bone morphogenetic protein (BMP) (e.g., BMP2, BMP4, BMP5, BMP7), ciliary neurotrophic factor family members (e.g., ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6)), colony stimulating factor (e.g., macrophage colony stimulating factor (m-CSF), granulocyte colony stimulating factor (G-CSF), granulocyte macrophage colony stimulating factor (GM-CSF)), epidermal growth factor (EGF), ephrin (e.g., ephrin A1, ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1, ephrin B2, ephrin B3), erythropoietin (EPO), fibroblast growth factor (FGF) (e.g., FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, FGF23), foetal bovine somatotrophin (FBS), GDNF family members (e.g., glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, artemin), growth differentiation factor-9 (GDF9), hepatocyte growth factor (HGF), hepatoma-derived growth factor (HDGF), insulin, insulin-like growth factor (e.g., insulin-like growth factor-1 (IGF-1) or IGF-2), interleukin (IL) (e.g., IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7), keratinocyte growth factor (KGF), migration stimulating factor (MSF), macrophage-stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)), myostatin (GDF-8), neuregulin (e.g., neuregulin 1 (NRG1), NRG2, NRG3, NRG4), neurotrophin (e.g., brain-derived neurotrophic factor (BDNF), nerve growth factor (NGF), neurotrophin-3 (NT-3), NT-4, placental growth factor (PGF), platelet-derived growth factor (PDGF), renin (RNLS), T cell growth factor (TCGF), thrombopoietin (TPO), transforming growth factor (e.g., transforming growth factor α (TGF-α), TGF-β, tumor necrosis factor-α (TNF-α) and vascular endothelial growth factor (VEGF).
[0119] In some embodiments, the therapeutic protein is encoded by a gene selected from the group consisting of: dystrophin X-linked, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof.
[0120] As used herein, the term "heterologous" or "exogenous" refers to a class of molecules that are not normally found in a given context, such as in a cell or polypeptide. For example, an exogenous or heterologous molecule can be introduced into a cell and only exists after the cell has been manipulated (e.g., by transfection or other forms of genetic engineering), or a heterologous amino acid sequence can be present in a protein where it is not naturally found.
[0121] As used herein, the term "heterologous nucleotide sequence" refers to a nucleotide sequence that is not naturally occurring in a given polynucleotide sequence. In one embodiment, the heterologous nucleotide sequence encodes a polypeptide capable of extending the half-life of a therapeutic protein (e.g., a clotting factor, e.g., FVIII). In another embodiment, the heterologous nucleotide sequence encodes a polypeptide that increases the hydrodynamic radius of a therapeutic protein (e.g., a clotting factor, e.g., FVIII). In other embodiments, the heterologous nucleotide sequence encodes a polypeptide that improves one or more pharmacokinetic properties of a therapeutic protein without significantly affecting its biological activity or function (e.g., procoagulant activity). In some embodiments, the therapeutic protein is linked or connected to the polypeptide encoded by the heterologous nucleotide sequence via a linker. Non-limiting examples of the polypeptide portion encoded by the heterologous nucleotide sequence include immunoglobulin constant regions or portions thereof, albumin or fragments thereof, albumin-binding moieties, transferrin, the PAS polypeptide of U.S. Patent Application No. 20100292130, HAP sequences, transferrin or fragments thereof, the C-terminal peptide (CTP) of the β subunit of human chorionic gonadotropin, albumin-binding small molecules, XTEN sequences, FcRn-binding moieties (e.g., the intact Fc region or portions thereof that bind FcRn), single-chain Fc regions (ScFc regions, e.g., as described in US 2008 / 0260738, WO2008 / 012543, or WO 2008 / 1439545), polyglycine linkers, polyserine linkers, peptides and short polypeptides of 6 - 40 amino acids from two types of amino acids selected from glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), and proline (P) (the degree of change in secondary structure of which is less than 50% to greater than 50%), and the like, or combinations of two or more thereof. In some embodiments, the polypeptide encoded by the heterologous nucleotide sequence is linked to a non-polypeptide portion. Non-limiting examples of the non-polypeptide portion include polyethylene glycol (PEG), albumin-binding small molecules, polysialic acid, hydroxyethyl starch (HES), derivatives thereof, or any combination thereof.
[0122] As used herein, the term "Fc region" is defined as the polypeptide portion corresponding to the Fc region of a native Ig, i.e., the portion formed by the dimerization association of the respective Fc domains of its two heavy chains. The native Fc region forms a homodimer with another Fc region. In contrast, the term "gene-fused Fc region" or "single-chain Fc region" (scFc region) as used herein refers to a synthetic dimeric Fc region composed of Fc domains gene-linked within a single polypeptide chain (i.e., a polypeptide chain encoded by a single continuous gene sequence).
[0123] In one embodiment, "Fc region" refers to the portion of a single Ig heavy chain that begins at the hinge region immediately upstream of the papain cleavage site (i.e., residue 216 in IgG, with the first residue of the heavy chain constant region designated as 114) and terminates at the C-terminus of the antibody. Thus, the complete Fc domain includes at least the hinge domain, CH2 domain, and CH3 domain.
[0124] Depending on the Ig isotype, the Fc region of the Ig constant region may include CH2, CH3, and CH4 domains, as well as the hinge region. Chimeric proteins containing the Fc region of an Ig confer several desirable properties on the chimeric protein, including increased stability, extended serum half-life (see Capon et al., 1989, Nature 337:525), and binding to Fc receptors (e.g., the neonatal Fc receptor (FcRn)) (U.S. Patent Nos. 6,086,875, 6,485,726, 6,030,613; WO 03 / 077834; US2003-0235536A1), which are incorporated herein by reference in their entirety.
[0125] When used herein as a comparison to the nucleotide sequences of the present disclosure, a "reference nucleotide sequence" is a polynucleotide sequence that is substantially identical to the nucleotide sequences of the present disclosure, except that the sequence has not been optimized. For example, the reference nucleotide sequence of a nucleic acid molecule consisting of the codon-optimized BDD FVIII of SEQ ID NO:1 and a heterologous nucleotide sequence encoding a single-chain Fc region linked to SEQ ID NO:1 at its 3' end is a nucleic acid molecule consisting of the original (or "parental") BDD FVIII of SEQ ID NO:16 and the same heterologous nucleotide sequence encoding a single-chain Fc region linked to SEQ ID NO:16 at its 3' end.
[0126] As used herein, with respect to a nucleotide sequence, the term "optimized" refers to a polynucleotide sequence encoding a polypeptide, wherein the polynucleotide sequence has been mutated to enhance the properties of the polynucleotide sequence. In some embodiments, the optimization is performed to increase transcription levels, increase translation levels, increase steady-state mRNA levels, increase or decrease the binding of regulatory proteins (such as general transcription factors), increase or decrease splicing, or increase the yield of the polypeptide produced from the polynucleotide sequence. Examples of changes that can be made to optimize a polynucleotide sequence include codon optimization, G / C content optimization, removal of repetitive sequences, removal of AT-rich elements, removal of cryptic splice sites, removal of cis-acting elements that inhibit transcription or translation, addition or removal of poly-T or poly-A sequences, addition of sequences that enhance transcription (such as Kozak consensus sequences) around the transcription start site, removal of sequences that can form stem-loop structures, removal of destabilizing sequences, and combinations of two or more thereof.
[0127] II. Nucleic Acid Molecules
[0128] The present disclosure relates to plasmid-like, capsidless nucleic acid molecules encoding therapeutic proteins or genes that can regulate the expression of target proteins. The capsid (the protein shell of a virus) surrounds the genetic material of the virus. It is known that the capsid aids the function of the virion by protecting the viral genome, delivering the genome to the host, and interacting with the host. Nevertheless, the viral capsid can be a factor that limits the packaging capacity of the vector and / or induces an immune response, especially when it is used in gene therapy.
[0129] AAV vectors have become one of the more common types of gene therapy vectors. However, the presence of the capsid limits the utility of AAV vectors in gene therapy. In particular, the capsid itself can limit the size of the transgene contained in the vector to as low as less than 4.5 kb. Even before adding regulatory elements, various therapeutic proteins available for gene therapy can easily exceed this size.
[0130] In addition, the proteins that make up the capsid can act as antigens that can be targeted by the subject's immune system. AAV is very common in the general population, and most people have been exposed to AAV during their lives. Therefore, most potential gene therapy recipients are likely to have already developed an immune response to AAV and are thus more likely to reject the therapy.
[0131] Certain aspects of the present disclosure are intended to overcome these deficiencies of AAV vectors. In particular, certain aspects of the present disclosure relate to a nucleic acid molecule that includes a first ITR, a second ITR, and a gene cassette, e.g., that encodes a therapeutic protein and / or miRNA. In some embodiments, the nucleic acid molecule does not include genes encoding capsid proteins, replication proteins, and / or assembly proteins. In some embodiments, the gene cassette encodes a therapeutic protein. In some embodiments, the therapeutic protein includes a clotting factor. In some embodiments, the gene cassette encodes miRNA. In certain embodiments, the gene cassette is located between the first ITR and the second ITR. In some embodiments, the nucleic acid molecule further includes one or more non-coding regions. In certain embodiments, the one or more non-coding regions include a promoter sequence, an intron, a post-transcriptional regulatory element, a 3'UTR poly(A) sequence, or any combination thereof.
[0132] In one embodiment, the gene cassette is a single-stranded nucleic acid. In another embodiment, the gene cassette is a double-stranded nucleic acid.
[0133] In one embodiment, the nucleic acid molecule includes:
[0134] (a) a first ITR that is an ITR of a non-AAV family member of the Parvoviridae;
[0135] (b) a tissue-specific promoter sequence, e.g., the TTP promoter;
[0136] (c) Introns, e.g., synthetic introns;
[0137] (d) Nucleotides encoding miRNA or a therapeutic protein (e.g., a clotting factor);
[0138] (e) Post-transcriptional regulatory elements, e.g., WPRE;
[0139] (f) 3'UTR poly(A) tail sequence, e.g., bGHpA;
[0140] (g) A second ITR that is an ITR of a member of the non-AAV family of the Parvoviridae.
[0141] In one embodiment, the nucleic acid molecule comprises:
[0142] (a) A first ITR that is an ITR of a member of the non-AAV family of the Parvoviridae;
[0143] (b) A tissue-specific promoter sequence, e.g., the TTP promoter;
[0144] (c) Introns, e.g., synthetic introns;
[0145] (d) Nucleotides encoding miRNA, wherein the miRNA downregulates the expression of target genes selected from SOD1, HTT, RHO, and any combination thereof;
[0146] (e) Post-transcriptional regulatory elements, e.g., WPRE;
[0147] (f) 3'UTR poly(A) tail sequence, e.g., bGHpA;
[0148] (g) A second ITR that is an ITR of a member of the non-AAV family of the Parvoviridae.
[0149] In one embodiment, the nucleic acid molecule comprises:
[0150] (a) A first ITR that is an ITR of a member of the non-AAV family of the Parvoviridae;
[0151] (b) A tissue-specific promoter sequence, e.g., the TTP promoter;
[0152] (c) Introns, e.g., synthetic introns;
[0153] (d) Nucleotides encoding dystrophin X-linked, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof;
[0154] (e) Post-transcriptional regulatory elements, e.g., WPRE;
[0155] (f) 3'UTR poly(A) tail sequence, e.g., bGHpA;
[0156] (g) A second ITR, which is an ITR of a non-AAV family member of the Parvoviridae family.
[0157] In one embodiment, the nucleic acid molecule comprises:
[0158] (a) A first ITR, which is an ITR of AAV (e.g., AAV serotype 2 genome);
[0159] (b) A tissue-specific promoter sequence, e.g., the TTP promoter;
[0160] (c) An intron, e.g., a synthetic intron;
[0161] (d) Nucleotides encoding FVIII; wherein the nucleotides have at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with a nucleotide sequence selected from SEQ ID NO: 1-14 or SEQ ID NO: 71, and wherein the FVIII encoded by the nucleotides retains FVIII activity;
[0162] (e) Post-transcriptional regulatory elements, e.g., WPRE;
[0163] (f) 3'UTR poly(A) tail sequence, e.g., bGHpA; and
[0164] (g) A second ITR, which is an ITR of AAV (e.g., AAV serotype 2 genome).
[0165] In one embodiment, the nucleic acid molecule comprises:
[0166] (a) A first ITR, which is an ITR of AAV (e.g., AAV serotype 2 genome);
[0167] (b) A tissue-specific promoter sequence, e.g., the TTP promoter;
[0168] (c) Introns, e.g., synthetic introns;
[0169] (d) Nucleotides encoding miRNA, wherein the miRNA downregulates the expression of target genes (e.g., SOD1, HTT, RHO, and any combination thereof);
[0170] (f) 3'UTR poly(A) tail sequence, e.g., bGHpA; and
[0171] (g) A second ITR, which is an ITR of AAV (e.g., AAV serotype 2 genome).
[0172] In one embodiment, the nucleic acid molecule comprises:
[0173] (a) A first ITR, which is an ITR of AAV (e.g., AAV serotype 2 genome);
[0174] (b) A tissue-specific promoter sequence, e.g., the TTP promoter;
[0175] (c) Introns, e.g., synthetic introns;
[0176] (d) Nucleotides encoding dystrophin X-linked type, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acidic α-glucosidase), or any combination thereof;
[0177] (f) 3'UTR poly(A) tail sequence, e.g., bGHpA; and
[0178] (g) A second ITR, which is an ITR of AAV (e.g., AAV serotype 2 genome).
[0179] In another embodiment, the nucleic acid molecule comprises:
[0180] (a) A first ITR;
[0181] (b) A tissue-specific promoter sequence, e.g., the TTP promoter;
[0182] (c) Introns, e.g., synthetic introns;
[0183] (d) Nucleotides encoding miRNA or a therapeutic protein (e.g., a clotting factor);
[0184] (e) A post-transcriptional regulatory element, e.g., WPRE;
[0185] (f) A 3' UTR poly(A) tail sequence, e.g., bGHpA; and
[0186] (g) A second ITR,
[0187] wherein one of the first ITR or the second ITR is an ITR of a non-AAV family member of the Parvoviridae, and the other ITR is an ITR of AAV (e.g., an AAV serotype 2 genome).
[0188] In another embodiment, the nucleic acid molecule comprises:
[0189] (a) A first ITR;
[0190] (b) A tissue-specific promoter sequence, the TTP promoter;
[0191] (c) An intron, e.g., a synthetic intron;
[0192] (d) Nucleotides encoding an miRNA or a therapeutic protein (e.g., a clotting factor);
[0193] (e) A post-transcriptional regulatory element, e.g., WPRE;
[0194] (f) A 3' UTR poly(A) tail sequence, e.g., bGHpA; and
[0195] (g) A second ITR,
[0196] wherein the first ITR is a synthetic ITR, the second ITR is a synthetic ITR, or both the first ITR and the second ITR are synthetic ITRs.
[0197] A. Inverted terminal repeat
[0198] Certain aspects of the present disclosure relate to a nucleic acid molecule comprising a first ITR, e.g., a 5' ITR, and a second ITR, e.g., a 3' ITR. Generally, ITRs are involved in parvovirus (e.g., AAV) DNA replication and rescue or excision from prokaryotic plasmids (Samulski et al., 1983, 1987; Senapathy et al., 1984; Gottlieb and Muzyczka, 1988). In addition, ITRs appear to be the minimal sequences required for AAV proviral integration and packaging of AAV DNA into viral particles (McLaughlin et al., 1988; Samulski et al., 1989). These elements are essential for efficient amplification of the parvovirus genome. The minimal defined elements that are hypothesized to be indispensable for ITR function are the Rep binding site (e.g., RBS; GCGCGCTCGCTCGCTC (SEQ ID NO:104) for AAV2) and the terminal resolution site (e.g., TRS; AGTTGG (SEQ ID NO:105) for AAV2), plus a variable palindromic sequence that permits hairpin formation. The palindromic nucleotide region generally functions together in cis as the origin of DNA replication and the packaging signal for the virus. During DNA replication, the complementary sequences in the ITR fold into a hairpin structure. In other embodiments, the ITR folds into a non-T-shaped hairpin structure, e.g., folds into a U-shaped hairpin structure. Data indicate that the T-shaped hairpin structure of the AAV ITR can inhibit the expression of transgenes flanked by the ITR. See, e.g., Zhou et al., Scientific Reports 7:5432 (July 4, 2017). By using an ITR that does not form a T-shaped hairpin structure, this form of inhibition can be avoided. Thus, in certain aspects, a polynucleotide comprising a non-AAV ITR has improved transgene expression compared to a polynucleotide comprising an AAV ITR that forms a T-shaped hairpin.
[0199] In some embodiments, the ITR comprises a naturally occurring ITR, e.g., the ITR comprises the entire portion of a parvovirus ITR. In some embodiments, the ITR comprises a synthetic sequence. In one embodiment, the first ITR or the second ITR comprises a synthetic sequence. In another embodiment, each of the first ITR and the second ITR comprises a synthetic sequence. In some embodiments, the first ITR or the second ITR comprises a naturally occurring sequence. In another embodiment, each of the first ITR and the second ITR comprises a naturally occurring sequence.
[0200] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR (e.g., a truncated ITR). In some embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides or at least about 600 nucleotides; wherein the ITR retains the functional properties of the naturally occurring ITR. In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 129 nucleotides; wherein the ITR retains the functional properties of the naturally occurring ITR. In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 102 nucleotides; wherein the ITR retains the functional properties of the naturally occurring ITR.
[0201] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR, wherein the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% of the length of the naturally occurring ITR; wherein the fragment retains the functional properties of the naturally occurring ITR.
[0202] In some embodiments, when properly aligned, the ITR comprises or consists of a sequence having at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In other embodiments, when properly aligned, the ITR comprises or consists of a sequence having at least 90% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, when properly aligned, the ITR comprises or consists of a sequence having at least 80% sequence identity to a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, when properly aligned, the ITR comprises or consists of a sequence having at least 70% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, when properly aligned, the ITR comprises or consists of a sequence having at least 60% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, when properly aligned, the ITR comprises or consists of a sequence having at least 50% sequence identity to a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR.
[0203] In some embodiments, the ITR comprises an ITR from an AAV genome. In some embodiments, the ITR is an ITR of an AAV genome selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, and any combination thereof. In a particular embodiment, the ITR is an ITR of the AAV2 genome. In another embodiment, the ITR is a synthetic sequence genetically engineered to include ITRs derived from one or more AAV genomes at the 5' and 3' ends.
[0204] In some embodiments, the ITR is not derived from an AAV genome. In some embodiments, the ITR is a non-AAV ITR. In some embodiments, the ITR is an ITR of a non-AAV genome from the viral family Parvoviridae, which viral family Parvoviridae is selected from, but not limited to, the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof. In certain embodiments, the ITR is derived from the parvovirus B19 (human virus) of the genus Erythrovirus. In another embodiment, the ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, e.g., the MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, e.g., the MDPV strain YY. In some embodiments, the ITR is derived from porcine parvovirus, e.g., porcine parvovirus U44978. In some embodiments, the ITR is derived from minute virus of mice, e.g., minute virus of mice U34256. In some embodiments, the ITR is derived from canine parvovirus, e.g., canine parvovirus M19296. In some embodiments, the ITR is derived from mink enteritis virus, e.g., mink enteritis virus D00765. In some embodiments, the ITR is derived from a dependoparvovirus. In one embodiment, the dependoparvovirus is a goose parvovirus (GPV) strain of the genus Dependovirus. In a specific embodiment, the GPV strain is attenuated, e.g., the GPV strain 82-0321V. In another specific embodiment, the GPV strain is pathogenic, e.g., the GPV strain B.
[0205] The first ITR and the second ITR of the nucleic acid molecule can be derived from the same genome (e.g., derived from the genome of the same virus), or from different genomes, e.g., genomes derived from the genomes of two or more different viruses. In certain embodiments, the first ITR and the second ITR are derived from the same AAV genome. In a specific embodiment, the two ITRs present in the nucleic acid molecule of the present invention are identical and can in particular be AAV2 ITRs. In other embodiments, the first ITR is derived from an AAV genome and the second ITR is not derived from an AAV genome (e.g., a non-AAV genome). In other embodiments, the first ITR is not derived from an AAV genome (e.g., a non-AAV genome) and the second ITR is derived from an AAV genome. In yet another other embodiment, neither the first ITR nor the second ITR is derived from an AAV genome (e.g., a non-AAV genome). In one particular embodiment, the first ITR and the second ITR are identical.
[0206] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a genome selected from the group consisting of Bocaparvovirus, Dependoparvovirus, Erythroparvovirus, Aleutian mink disease virus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penaeus densovirus, and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a genome selected from the group consisting of Bocaparvovirus, Dependoparvovirus, Erythroparvovirus, Aleutian mink disease virus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penaeus densovirus, and any combination thereof. In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of Bocaparvovirus, Dependoparvovirus, Erythroparvovirus, Aleutian mink disease virus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penaeus densovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from the same genome. In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of Bocaparvovirus, Dependoparvovirus, Erythroparvovirus, Aleutian mink disease virus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penaeus densovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from different genomes.
[0207] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from Erythroparvovirus B19 (human virus). In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from Erythroparvovirus B19 (human virus).
[0208] In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a portion of an ITR derived from the ITR of B19. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171, wherein the first ITR and / or the second ITR retains the functional properties of the B19 ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.
[0209] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 167. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 168. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 169. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 170. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 171.
[0210] Table 1. Parvovirus ITR sequences of samples.
[0211]
[0212]
[0213]
[0214] In certain embodiments, the first ITR and / or the second ITR comprise a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence shown in SEQ ID NO:169, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence shown in SEQ ID NO:167, retaining the functional properties of the B19 ITR, and the nucleotide sequence is derived from the B19 ITR. In some embodiments, the first ITR and / or the second ITR comprise a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence shown in SEQ ID NO:169, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence shown in SEQ ID NO:167, wherein the first ITR and / or the second ITR are capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.
[0215] In some embodiments, the first ITR is derived from an AAV genome and the second ITR is derived from a GPV. In other embodiments, the second ITR is derived from an AAV genome and the first ITR is derived from a GPV.
[0216] In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or part of an ITR derived from the ITR of GPV. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to a nucleotide sequence selected from the nucleotide sequences shown in SEQ ID NO: 172, 173, 174, 175, and 176, wherein the first ITR and / or the second ITR retains the functional properties of the GPV ITR, and the first ITR and / or the second ITR is derived from the GPV ITR. In some embodiments, the first ITR and / or the second ITR comprises or consists of all or part of an ITR derived from the ITR of GPV. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to a nucleotide sequence selected from SEQ ID NO: 172, 173, 174, 175, and 176, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NO: 172, 173, 174, 175, and 176. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence shown in SEQ ID NO: 172. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence shown in SEQ ID NO: 173. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence shown in SEQ ID NO: 174. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence shown in SEQ ID NO: 175. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence shown in SEQ ID NO: 176.
[0217] In certain embodiments, the first ITR and / or the second ITR comprise a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence shown in SEQ ID NO: 174, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence shown in SEQ ID NO: 172, wherein the first ITR and / or the second ITR retain the functional properties of the GPV ITR, and the first ITR and / or the second ITR are derived from the GPV ITR. In some embodiments, the first ITR and / or the second ITR comprise a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence shown in SEQ ID NO: 174, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence shown in SEQ ID NO: 172, wherein the first ITR and / or the second ITR are capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.
[0218] In certain embodiments, the first ITR and / or the second ITR comprise a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence shown in SEQ ID NO: 176, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence shown in SEQ ID NO: 172, wherein the first ITR and / or the second ITR retain the functional characteristics of the GPV ITR, and the first ITR and / or the second ITR is derived from the GPV ITR. In some embodiments, the first ITR and / or the second ITR comprise a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence shown in SEQ ID NO: 176, and wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence shown in SEQ ID NO: 172, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.
[0219] In certain embodiments, one of the first ITR or the second ITR comprises or consists of all or part of an ITR derived from AAV2. In some embodiments, the first ITR or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence shown in SEQ ID NO:177 or 178, wherein the first ITR and / or the second ITR retains the functional properties of the AAV2 ITR, and the first ITR and / or the second ITR is derived from the AAV2 ITR. In some embodiments, the first ITR or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence shown in SEQ ID NO:177 or 178, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence shown in SEQ ID NO:177 or 178. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence shown in SEQ ID NO:177. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence shown in SEQ ID NO:178.
[0220] In some embodiments, the first ITR is derived from the AAV genome and the second ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In other embodiments, the second ITR is derived from the AAV genome and the first ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, e.g., the MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, e.g., the MDPV strain YY.
[0221] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from the genus Dependoparvovirus. In some embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from the genus Dependoparvovirus. In other embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a Dependovirus goose parvovirus (GPV) strain. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a Dependovirus GPV strain. In certain embodiments, the GPV strain is attenuated, e.g., GPV strain 82 - 0321V. In other embodiments, the GPV strain is pathogenic, e.g., GPV strain B.
[0222] In certain embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a genome selected from porcine parvovirus, e.g., porcine parvovirus strain U44978; minute virus of mice, e.g., minute virus of mice strain U34256; canine parvovirus, e.g., canine parvovirus strain M19296; mink enteritis virus, e.g., mink enteritis virus strain D00765; and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a genome selected from porcine parvovirus, e.g., porcine parvovirus strain U44978; minute virus of mice, e.g., minute virus of mice strain U34256; canine parvovirus, e.g., canine parvovirus strain M19296; mink enteritis virus, e.g., mink enteritis virus strain D00765; and any combination thereof.
[0223] In another specific embodiment, the ITR is a synthetic sequence genetically engineered to include at its 5' and 3' ends ITRs not derived from an AAV genome. In another specific embodiment, the ITR is a synthetic sequence genetically engineered to include at its 5' and 3' ends ITRs derived from one or more non - AAV genomes. The two ITRs present in the nucleic acid molecule of the present invention can be the same or different non - AAV genomes. In particular, the ITRs can be derived from the same non - AAV genome. In a specific embodiment, the two ITRs present in the nucleic acid molecule of the present invention are the same and can in particular be AAV2 ITRs.
[0224] In some embodiments, the ITR sequence comprises one or more palindromic sequences. The palindromic sequences of the ITRs disclosed herein include, but are not limited to, natural palindromic sequences (i.e., sequences found in nature), synthetic sequences (i.e., sequences not found in nature), such as pseudo-palindromic sequences, and combinations or modified forms thereof. A "pseudo-palindromic sequence" is a palindromic DNA sequence, including imperfect palindromic sequences, that shares less than 80%, including less than 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5%, or no nucleic acid sequence identity with the sequences in natural AAV or non-AAV palindromic sequences that form secondary structures. Natural palindromic sequences can be obtained from or derived from any genome disclosed herein. Synthetic palindromic sequences can be based on any genome disclosed herein.
[0225] The palindromic sequence can be continuous or interrupted. In some embodiments, the palindromic sequence is interrupted, where the palindromic sequence contains an insertion of a second sequence. In some embodiments, the second sequence comprises a promoter, an enhancer, an integration site for an integrase (e.g., a site for Cre or Flp recombinase), an open reading frame of a gene product, or a combination thereof.
[0226] In some embodiments, the ITRs form hairpin loop structures. In one embodiment, the first ITR forms a hairpin structure. In another embodiment, the second ITR forms a hairpin structure. Still in another embodiment, both the first ITR and the second ITR form hairpin structures. In some embodiments, the first ITR and / or the second ITR do not form a T-shaped hairpin structure. In certain embodiments, the first ITR and / or the second ITR form non-T-shaped hairpin structures. In some embodiments, the non-T-shaped hairpin structure comprises a U-shaped hairpin structure.
[0227] In some embodiments, the ITRs of the nucleic acid molecules described herein can be transcriptionally activated ITRs. The transcriptionally activated ITRs can comprise all or part of a wild-type ITR that has been transcriptionally activated by including at least one transcriptional activity element. Multiple types of transcriptional activity elements are suitable in this context. In some embodiments, the transcriptional activity element is a constitutive transcriptional activity element. Constitutive transcriptional activity elements provide a continuous level of gene transcription and are preferred when continuous expression of a transgene is required. In other embodiments, the transcriptional activation element is an inducible transcriptional activation element. Inducible transcriptional activation elements generally exhibit low activity in the absence of an inducer (or inducing conditions) and are upregulated in the presence of an inducer (or upon switching to inducing conditions). Inducible transcriptional activity elements can be preferred when expression is only needed at certain times or certain locations, or when titration of the expression level with an inducer is desired. The transcriptional activation elements can also be tissue-specific; i.e., they exhibit activity only in certain tissues or cell types.
[0228] Transcriptional activation elements can be incorporated into the ITRs in a variety of ways. In some embodiments, the transcriptional activation element is incorporated 5′ to any part of the ITR or 3′ to any part of the ITR. In other embodiments, the transcriptional activation element of the transcriptionally activated ITR is located between two ITR sequences. If the transcriptional activation element comprises two or more elements that must be spaced apart, these elements can be interspersed with a portion of the ITR. In some embodiments, the hairpin structure of the ITR is deleted and replaced with an inverted repeat of the transcriptional element. The latter arrangement will generate a hairpin that mimics the deleted portion in the structure. Multiple tandem transcriptional activation elements can also be present in the transcriptionally activated ITR, and they can be adjacent or spaced apart. In addition, protein binding sites (e.g., Rep binding sites) can be introduced into the transcriptional activation elements of the transcriptionally activated ITR. The transcriptional activation element can comprise any sequence capable of controlling the transcription of DNA by RNA polymerase to form RNA, and can comprise, for example, transcriptional activation elements as defined below.
[0229] The transcriptionally activated ITR provides transcriptional activation and ITR function to a nucleic acid molecule with a relatively limited nucleotide sequence length, which effectively maximizes the length of the transgene that can be carried and expressed from the nucleic acid molecule. Incorporating transcriptional activation elements into the ITR can be accomplished in a variety of ways. Comparison of ITR sequences and sequence requirements of transcriptional activation elements can provide an understanding of the manner in which the element is encoded within the ITR. For example, transcriptional activity can be added to the ITR by introducing specific changes into the ITR sequence that replicate the functional elements of the transcriptional activation element. Many techniques exist in the art for effectively adding, deleting, and / or altering specific nucleotide sequences at specific loci (see, e.g., Deng and Nickoloff (1992) Anal. Biochem. 200:81-88). Another way to create a transcriptionally activated ITR involves introducing restriction sites at desired positions within the ITR. In addition, multiple transcriptional activation elements can be incorporated into the transcriptionally activated ITR using methods known in the art.
[0230] B. Therapeutic Proteins
[0231] Certain aspects of the present disclosure relate to a nucleic acid molecule comprising a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein. In some embodiments, the gene cassette encodes a therapeutic protein. In some embodiments, the gene cassette encodes more than one therapeutic protein. In some embodiments, the gene cassette encodes two or more copies of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more variants of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more different therapeutic proteins.
[0232] Certain embodiments of the present disclosure relate to a nucleic acid molecule comprising a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a coagulation factor. In some embodiments, the coagulation factor is selected from the group consisting of: FI, FII, FIII, FIV, FV, FVI, FVII, FVIII, FIX, FX, FXI, FXII, FXIII), VWF, prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, α2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), any of their zymogens, any of their activated forms, and any combination thereof. In one embodiment, the coagulation factor comprises FVIII or a variant or fragment thereof. In another embodiment, the coagulation factor comprises FIX or a variant or fragment thereof. In another embodiment, the coagulation factor comprises FVII or a variant or fragment thereof. In another embodiment, the coagulation factor comprises VWF or a variant or fragment thereof.
[0233] 1. Coagulation factor
[0234] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a factor VIII polypeptide. Unless otherwise indicated, "factor VIII" abbreviated as "FVIII" throughout this application means a functional FVIII polypeptide that functions normally in blood coagulation. Thus, the term FVIII includes functional variant polypeptides. "FVIII protein" is used interchangeably with FVIII polypeptide (or protein) or FVIII. Examples of FVIII function include, but are not limited to, the ability to activate blood coagulation, the ability to act as a cofactor for factor IX, or in Ca 2+The ability to form a Tenase complex with factor IX in the presence of phospholipids, and then the Tenase complex converts factor X into the activated form Xa. The FVIII protein can be human, porcine, canine, rat, or murine FVIII protein. In addition, comparisons between human and other species' FVIII have identified conserved residues that may be required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US 6,251,632). Full-length polypeptide and polynucleotide sequences are known, as well as many functional fragments, mutants, and modified variants. Multiple FVIII amino acid and nucleotide sequences are disclosed in, for example, US Publication Nos. 2015 / 0158929 A1, 2014 / 0308280 A1, and 2014 / 0370035 A1, and International Publication No. WO 2015 / 106052 A1. FVIII polypeptides include, for example, full-length FVIII, full-length FVIII with Met removed from the N-terminus, mature FVIII (minus the signal sequence), mature FVIII with an additional Met at the N-terminus, and / or FVIII with all or part of the B domain deleted. FVIII variants include B domain deletions, whether partial or complete.
[0235] a. FVIII and polynucleotide sequences encoding the FVIII protein
[0236] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a factor VIII polypeptide. Unless otherwise indicated, "factor VIII," abbreviated as "FVIII" throughout this application, refers to a functional FVIII polypeptide that functions normally in blood coagulation. Thus, the term FVIII includes functional variant polypeptides. "FVIII protein" is used interchangeably with FVIII polypeptide (or protein) or FVIII. Examples of FVIII function include, but are not limited to, the ability to activate blood coagulation, the ability to act as a cofactor for factor IX, or in the presence of Ca 2+The ability to form a tenase complex with factor IX in the presence of phospholipids, and then the tenase complex converts factor X into the activated form Xa. The FVIII protein can be human, porcine, canine, rat or mouse FVIII protein. In addition, comparisons between FVIII from humans and other species have identified conserved residues that may be required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US 6,251,632). Full-length polypeptide and polynucleotide sequences are known, as well as many functional fragments, mutants and modified variants. Multiple FVIII amino acid and nucleotide sequences are disclosed in, for example, US Publication Nos. 2015 / 0158929 A1, 2014 / 0308280 A1 and 2014 / 0370035 A1 and International Publication No. WO 2015 / 106052 A1. FVIII polypeptides include, for example, full-length FVIII, full-length FVIII minus Met at the N-terminus, mature FVIII (minus the signal sequence), mature FVIII with an additional Met at the N-terminus, and / or FVIII with all or part of the B domain deleted. FVIII variants include B domain deletions, whether partial or complete.
[0237] The FVIII portion in the chimeric proteins used herein has FVIII activity. FVIII activity can be measured by any known method in the art. Many types of tests can be used to assess the function of the coagulation system: activated partial thromboplastin time (aPTT) test, chromogenic assay, ROTEM assay, prothrombin time (PT) test (also used to determine INR), fibrinogen test (usually by the Clauss method), platelet count, platelet function test (usually by PFA-100), TCT, bleeding time, mixing test (whether the patient's plasma can be corrected for the abnormality if mixed with normal plasma), coagulation factor assay, antiphospholipid antibody, D-dimer, genetic test (e.g., factor V Leiden, prothrombin mutation G20210A), dilute Russell's viper venom time (dRVVT), other platelet function tests, thromboelastography (TEG or Sonoclot), thromboelastometry ( For example, ) or euglobulin lysis time (ELT).
[0238] The aPTT test is a performance indicator that measures the efficacy of the "intrinsic" (also known as the contact activation pathway) and the common coagulation pathway. This test is commonly used to measure the coagulation activity of commercially available recombinant coagulation factors (e.g., FVIII). It is used in conjunction with the prothrombin time (PT), which measures the extrinsic pathway.
[0239] ROTEM analysis provides information on the overall dynamics of hemostasis: clotting time, clot formation, clot stability, and lysis. Different parameters in thromboelastometry depend on the activity of the plasma clotting system, platelet function, fibrinolysis, or many factors that affect these interactions. This assay provides a complete view of secondary hemostasis.
[0240] The chromogenic assay mechanism is based on the principle of the coagulation cascade, in which activated FVIII accelerates the conversion of factor X to factor Xa in the presence of activated factor IX, phospholipids, and calcium ions. Factor Xa activity is assessed by hydrolysis of a paranitroaniline (pNA) substrate specific for factor Xa. The initial release rate of paranitroaniline measured at 405 nM is proportional to factor Xa activity and thus to FVIII activity in the sample.
[0241] The chromogenic assay is recommended by the Factor VIII and IX Subcommittee of the Scientific and Standardization Committee (SSC) of the International Society on Thrombosis and Haemostasis (ISTH). Since 1994, the chromogenic assay has been the reference method for the European Pharmacopoeia monograph on the potency of FVIII concentrates. Thus, in one embodiment, a chimeric polypeptide comprising FVIII has FVIII activity comparable to that of a chimeric polypeptide comprising mature FVIII or BDD FVIII (e.g., or ).
[0242] In another embodiment, a chimeric protein comprising FVIII of the present disclosure has a factor Xa generation rate comparable to that of a chimeric protein comprising mature FVIII or BDD FVIII (e.g., or ).
[0243] For the activation of factor X to factor Xa, activated factor IX (factor IXa) hydrolyzes an arginine-isoleucine bond in factor X in the presence of Ca 2+ , membrane phospholipids, and the FVIII cofactor to form factor Xa. Thus, the interaction of FVIII with factor IX is crucial in the coagulation pathway. In certain embodiments, a chimeric polypeptide comprising FVIII can interact with factor IXa at a rate comparable to that of a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ).
[0244] In addition, FVIII binds to von Willebrand factor (VWF) but is inactive in circulation. When not bound to VWF, FVIII is rapidly degraded and released from VWF by the action of thrombin. In some embodiments, the chimeric polypeptide comprising FVIII binds to von Willebrand factor at a rate comparable to that of a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ).
[0245] FVIII can be inactivated by activated protein C in the presence of calcium and phospholipids. Activated protein C cleaves the FVIII heavy chain after arginine 336 in the A1 domain, which disrupts the factor X substrate interaction site, and after arginine 562 in the A2 domain, which enhances the dissociation of the A2 domain and disrupts the interaction site with factor IXa. This cleavage also splits the A2 domain (43 kDa) into two parts, generating the A2-N (18 kDa) and A2-C (25 kDa) domains. Thus, activated protein C can catalyze multiple cleavage sites in the heavy chain. In one embodiment, the chimeric polypeptide comprising FVIII is inactivated by activated protein C at a level comparable to that of a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ).
[0246] In other embodiments, the chimeric protein comprising FVIII has an in vivo FVIII activity comparable to that of a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ). In certain embodiments, in the HemA mouse tail vein transection model, the chimeric polypeptide comprising FVIII is able to protect HemA mice at a level comparable to that of a chimeric polypeptide comprising a mature FVIII sequence or BDD FVIII (e.g., or ).
[0247] As used herein, the "B domain" of FVIII is the same as that known in the art, which is defined by internal amino acid sequence identity and proteolytic cleavage sites for thrombin, such as residues Ser741 - Arg1648 of mature human FVIII. Other human FVIII domains are defined by the following amino acid residues, relative to mature human FVIII: A1 of mature FVIII, residues Ala1 - Arg372; A2, residues Ser373 - Arg740; A3, residues Ser1690 - Ile2032; C1, residues Arg2033 - Asn2172; C2, residues Ser2173 - Tyr2332. Unless otherwise indicated, the sequence residue numbers used herein without reference to any SEQ ID number correspond to the FVIII sequence without the signal peptide sequence (19 amino acids). The A3 - C1 - C2 sequence, also referred to as the FVIII heavy chain, includes residues Ser1690 - Tyr2332. The remaining sequence, residues Glu1649 - Arg1689, is commonly referred to as the FVIII light chain activation peptide. The boundary positions of all domains (including the B domain) of porcine, murine, and canine FVIII are also known in the art. In one embodiment, the B domain of FVIII is deleted ("B domain - deleted FVIII" or "BDD FVIII"). An example of BDD FVIII is (recombinant BDD FVIII). In a particular embodiment, the B domain - deleted FVIII variant comprises a deletion of amino acid residues 746 to 1648 of mature FVIII.
[0248] "FVIII with B-domain deletion" may have all or part of the deletions disclosed in U.S. Patent Nos. 6,316,226, 6,346,513, 7,041,635, 5,789,203, 6,060,447, 5,595,886, 6,228,620, 5,972,885, 6,048,720, 5,543,502, 5,610,278, 5,171,844, 5,112,950, 4,868,112, and 6,458,563, as well as International Patent No. WO 2015106052 A1 (PCT / US2015 / 010738). In some embodiments, the FVIII with B-domain deletion sequence used in the methods of the present disclosure comprises any of the deletions disclosed in column 4, line 4 to column 5, line 28 and Examples 1-5 of U.S. Patent No. 6,316,226 (also disclosed in US 6,346,513). In another embodiment, the factor VIII with B-domain deletion is S743 / Q1638 B-domain deleted factor VIII (SQ BDD FVIII) (e.g., factor VIII having a deletion from amino acid 744 to amino acid 1637, e.g., factor VIII having amino acids 1-743 and amino acids 1638-2332 of mature FVIII). In some embodiments, the FVIII with B-domain deletion used in the methods of the present disclosure has the deletions disclosed in column 2, lines 26-52 and Examples 5-8 of U.S. Patent No. 5,789,203 (also disclosed in US 6,060,447, US 5,595,886, and US 6,228,620). In some embodiments, the factor VIII with B-domain deletion has the deletions disclosed in column 1, line 25 to column 2, line 40 of U.S. Patent No. 5,972,885; column 6, lines 1-22 and Example 1 of U.S. Patent No. 6,048,720; column 2, lines 17-46 of U.S. Patent No. 5,543,502; column 4, line 22 to column 5, line 36 of U.S. Patent No. 5,171,844; column 2, lines 55-68, Figure 2, and Example 1 of U.S. Patent No. 5,112,950; column 2, line 2 to column 19, line 21 and Table 2 of U.S. Patent No. 4,868,112; column 2, line 1 to column 3, line 19, column 3, line 40 to column 4, line 67, column 7, line 43 to column 8, line 26, and column 11, line 5 to column 13, line 39 of U.S. Patent No. 7,041,635; or column 4, lines 25-53 of U.S. Patent No. 6,458,563. In some embodiments, the FVIII with B-domain deletion has a deletion of most of the B-domain but still contains the amino-terminal sequence of the B-domain, which is crucial for the proteolytic processing of the primary translation product in vivo into two polypeptide chains, as disclosed in WO 91 / 09122.In some embodiments, the FVIII construct lacking the B domain has a deletion of amino acids 747 - 1638, i.e., a complete deletion of the B domain in effect. Hoeben R.C., et al. J. Biol. Chem. 265(13):7318 - 7323(1990). The factor VIII lacking the B domain may also contain a deletion of amino acids 771–1666 or amino acids 868 - 1562 of FVIII. Meulien P., et al. Protein Eng. 2(4):301 - 6(1988). Additional B domain deletions that are part of the present invention include: deletions of amino acids 982 to 1562 or 760 to 1639 (Toole et al., Proc. Natl. Acad. Sci. U.S.A. (1986) 83,5939 - 5942)), 797 to 1562 (Eaton, et al. Biochemistry (1986) 25:8343 - 8347)), 741 to 1646 (Kaufman (PCT International Publication No. WO 87 / 04187)), 747 - 1560 (Sarver, et al., DNA (1987) 6:553 - 564)), 741 to 1648 (Pasek (PCT Application No. 88 / 00831)) or 816 to 1598 or 741 to 1648 (Lagner (Behring Inst. Mitt. (1988) No 82:16 - 25, EP 295597)). In a particular embodiment, the FVIII lacking the B domain contains a deletion of amino acid residues 746 to 1648 of mature FVIII. In another embodiment, the FVIII lacking the B domain contains a deletion of amino acid residues 745 to 1648 of mature FVIII. In some embodiments, BDD FVIII contains single-chain FVIII that contains a deletion corresponding to amino acids 765 to 1652 of mature full-length FVIII (also referred to as rVIII single-chain and. ). See U.S. Patent No. 7,041,635.
[0249] In other embodiments, the BDD FVIII comprises an FVIII polypeptide containing a fragment of the B domain that retains one or more N-linked glycosylation sites, e.g., residues 757, 784, 828, 900, 963 or optionally 943, which correspond to the amino acid sequence of the full-length FVIII sequence. Examples of B domain fragments include those disclosed in Miao, H.Z., et al., Blood 103(a):3412-3419 (2004), Kasuda, A, et al., J. Thromb. Haemost. 6:1352-1359 (2008), and Pipe, S.W., et al., J. Thromb. Haemost. 9:2235-2242 (2011), which are 226 amino acids or 163 amino acids of the B domain (i.e., retaining the first 226 amino acids or 163 amino acids of the B domain). In still other embodiments, the BDD FVIII further comprises a point mutation at residue 309 (Phe mutated to Ser) to improve the expression of the BDD FVIII protein. See Miao, H.Z., et al., Blood 103(a):3412-3419 (2004). In still other embodiments, the BDD FVIII comprises an FVIII polypeptide that contains a partial B domain but does not contain one or more furin cleavage sites (e.g., Arg1313 and Arg1648). See Pipe, S.W., et al., J. Thromb. Haemost. 9:2235-2242 (2011). In some embodiments, the BDD FVIII comprises a single-chain FVIII that contains a deletion corresponding to amino acids 765 to 1652 of mature full-length FVIII (also referred to as rVIII single-chain and ). See U.S. Patent No. 7,041,635. Each of the foregoing deletions can be made in any FVIII sequence.
[0250] As discussed above and below, many functional FVIII variants are known. In addition, hundreds of non-functional mutations of FVIII have been identified in hemophilia patients, and it has been determined that the effect of these mutations on FVIII function is due more to their location within the 3-dimensional structure of FVIII rather than the nature of the substitution (Cutler et al., Hum. Mutat. 19:274-8 (2002)), which is incorporated herein by reference in its entirety. In addition, comparison between FVIII from humans and other species has identified conserved residues that may be required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US 6,251,632), which is incorporated herein by reference in its entirety.
[0251] In some embodiments, the FVIII polypeptide comprises an FVIII variant or a fragment thereof, wherein the FVIII variant or the fragment thereof has FVIII activity. In some embodiments, the gene cassette encodes a full-length FVIII polypeptide. In other embodiments, the gene cassette encodes a B domain-deleted (BDD) FVIII polypeptide, wherein all or part of the B domain of FVIII is deleted. In a particular embodiment, the gene cassette encodes a polypeptide comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to SEQ ID NO: 106, 107, 109, 110, 111 or 112. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 17 or a fragment thereof. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 106 or a fragment thereof. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 107. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 109 or a fragment thereof. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 16. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 109.
[0252] In some embodiments, the gene cassette of the present disclosure encodes an FVIII polypeptide comprising a signal peptide or a fragment thereof. In other embodiments, the gene cassette encodes an FVIII polypeptide lacking a signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.
[0253] In some embodiments, the gene cassette comprises a nucleotide sequence encoding an FVIII polypeptide, wherein the nucleotide sequence is codon-optimized. In certain embodiments, the gene cassette comprises the nucleotide sequence of International Application No. PCT / US2017 / 015879, which is incorporated herein by reference in its entirety. In some embodiments, the gene cassette comprises a nucleotide sequence encoding an FVIII polypeptide, wherein the nucleotide sequence is codon-optimized. In certain embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to a nucleotide sequence selected from SEQ ID NOs: 1-14. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 71. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 19.
[0254] i. A codon-optimized nucleotide sequence encoding an FVIII polypeptide
[0255] In some embodiments, the nucleic acid molecule of the present disclosure comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the first ITR and the second ITR are derived from an AAV genome, and wherein the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide. In some embodiments, the codon-optimized nucleotide sequence encodes a full-length FVIII polypeptide. In other embodiments, the codon-optimized nucleotide sequence encodes a B-domain deleted (BDD) FVIII polypeptide, wherein all or part of the B domain of FVIII is deleted. In a particular embodiment, the codon-optimized nucleotide sequence encodes a polypeptide comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to SEQ ID NO: 17 or a fragment thereof. In one embodiment, the codon-optimized nucleotide sequence encodes a polypeptide having the amino acid sequence of SEQ ID NO: 17 or a fragment thereof.
[0256] In some embodiments, the codon-optimized nucleotide sequence encodes an FVIII polypeptide comprising a signal peptide or a fragment thereof. In other embodiments, the codon-optimized sequence encodes an FVIII polypeptide lacking a signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.
[0257] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with (i) nucleotides 58-1791 of SEQ ID NO:3 or (ii) nucleotides 58-1791 of SEQ ID NO:4; and wherein the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In a particular embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 58-1791 of SEQ ID NO:3. In another embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 58-1791 of SEQ ID NO:4. In other embodiments, the first nucleotide sequence comprises nucleotides 58-1791 of SEQ ID NO:3 or nucleotides 58-1791 of SEQ ID NO:4.
[0258] In other embodiments, a codon-optimized nucleotide sequence encoding an FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with (i) nucleotides 1-1791 of SEQ ID NO:3 or (ii) nucleotides 1-1791 of SEQ ID NO:4; and wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1-1791 of SEQ ID NO:3 or nucleotides 1-1791 of SEQ ID NO:4. In another embodiment, the second nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 1792-4374 of SEQ ID NO:3 or 1792-4374 of SEQ ID NO:4. In a particular embodiment, the second nucleotide sequence comprises nucleotides 1792-4374 of SEQ ID NO:3 or 1792-4374 of SEQ ID NO:4. In yet another embodiment, the second nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:3 or 1792-2277 and 2320-4374 of SEQ ID NO:4 (i.e., nucleotides 1792-4374 of SEQ ID NO:3 or 1792-4374 of SEQ ID NO:4 that do not have nucleotides encoding the B domain or a B domain fragment). In a particular embodiment, the second nucleotide sequence comprises nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:3 or 1792-2277 and 2320-4374 of SEQ ID NO:4 (i.e., nucleotides 1792-4374 of SEQ ID NO:3 or 1792-4374 of SEQ ID NO:4 that do not have nucleotides encoding the B domain or a B domain fragment).
[0259] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with (i) nucleotides 1792-4374 of SEQ ID NO:5 or (ii) 1792-4374 of SEQ ID NO:6; and wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 1792-4374 of SEQ ID NO:5. In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 1792-4374 of SEQ ID NO:6. In a particular embodiment, the second nucleic acid sequence comprises nucleotides 1792-4374 of SEQ ID NO:5 or 1792-4374 of SEQ ID NO:6. In some embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 58-1791 of SEQ ID NO:5 or nucleotides 58-1791 of SEQ ID NO:6. In other embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 1-1791 of SEQ ID NO:5 or nucleotides 1-1791 of SEQ ID NO:6.
[0260] In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1792-4374 of SEQ ID NO:5, which do not encode the B domain or a B domain fragment) or (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:6, which do not encode the B domain or a B domain fragment); and wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1792-4374 of SEQ ID NO:5, which do not encode the B domain or a B domain fragment). In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:6, which do not encode the B domain or a B domain fragment). In a particular embodiment, the second nucleic acid sequence comprises nucleotides 1792-2277 of SEQ ID NO:5 and 2320-4374 or 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:5 or 1792-4374 of SEQ ID NO:6, which do not encode the B domain or a B domain fragment). In some embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 58-1791 of SEQ ID NO:5 or nucleotides 58-1791 of SEQ ID NO:6.In other embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 1-1791 of SEQ ID NO:5 or nucleotides 1-1791 of SEQ ID NO:6.
[0261] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with (i) nucleotides 58-1791 of SEQ ID NO:1, (ii) nucleotides 58-1791 of SEQ ID NO:2, (iii) nucleotides 58-1791 of SEQ ID NO:70, or (iv) nucleotides 58-1791 of SEQ ID NO:71; and wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity. In other embodiments, the first nucleotide sequence comprises nucleotides 58-1791 of SEQ ID NO:1, nucleotides 58-1791 of SEQ ID NO:2, (iii) nucleotides 58-1791 of SEQ ID NO:70, or (iv) nucleotides 58-1791 of SEQ ID NO:71.
[0262] In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with (i) nucleotides 1-1791 of SEQ ID NO:1, (ii) nucleotides 1-1791 of SEQ ID NO:2, (iii) nucleotides 1-1791 of SEQ ID NO:70, or (iv) nucleotides 1-1791 of SEQ ID NO:71; and wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1-1791 of SEQ ID NO:1, nucleotides 1-1791 of SEQ ID NO:2, (iii) nucleotides 1-1791 of SEQ ID NO:70, or (iv) nucleotides 1-1791 of SEQ ID NO:71. In another embodiment, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with nucleotides 1792-4374 of SEQ ID NO:1, 1792-4374 of SEQ ID NO:2, (iii) nucleotides 1792-4374 of SEQ ID NO:70 or (iv) nucleotides 1792-4374 of SEQ ID NO:71. In a particular embodiment, the second nucleotide sequence linked to the first nucleotide sequence comprises (i) nucleotides 1792-4374 of SEQ ID NO:1, (ii) nucleotides 1792-4374 of SEQ ID NO:2, (iii) nucleotides 1792-4374 of SEQ ID NO:70 or (iv) nucleotides 1792-4374 of SEQ ID NO:71. In other embodiments, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:70 or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:71.In one embodiment, the second nucleotide sequence comprises (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:71.
[0263] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with (i) nucleotides 1792-4374 of SEQ ID NO:1, (ii) nucleotides 1792-4374 of SEQ ID NO:2, (iii) nucleotides 1792-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-4374 of SEQ ID NO:71; and wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity. In a particular embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792-4374 of SEQ ID NO:1, (ii) nucleotides 1792-4374 of SEQ ID NO:2, (iii) nucleotides 1792-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-4374 of SEQ ID NO:71. In some embodiments, the codon-optimized sequence encoding the FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:71 (i.e., nucleotides 1792-4374 of SEQ ID NO:1, nucleotides 1792-4374 of SEQ ID NO:2, nucleotides 1792-4374 of SEQ ID NO:70 or nucleotides 1792-4374 of SEQ ID NO:71, which do not have nucleotides encoding the B domain or a B domain fragment); and wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity.In one embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:71 (i.e., nucleotides 1792-4374 of SEQ ID NO:1, nucleotides 1792-4374 of SEQ ID NO:2, nucleotides 1792-4374 of SEQ ID NO:70, or nucleotides 1792-4374 of SEQ ID NO:71, which do not have nucleotides encoding the B domain or a B domain fragment).
[0264] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with nucleotides 58 to 4374 of SEQ ID NO:1. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with nucleotides 58-2277 and 2320-4374 of SEQ ID NO:1 (i.e., nucleotides 58-4374 of SEQ ID NO:1, which do not have nucleotides encoding the B domain or a B domain fragment). In other embodiments, the nucleic acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO:1. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:1 (i.e., nucleotides 58-4374 of SEQ ID NO:1, which do not have nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO:1. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO:1 (i.e., nucleotides 1-4374 of SEQ ID NO:1, which do not have nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO:1.
[0265] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO:2. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO:2. In other embodiments, the nucleic acid sequence has at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO:2. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:2 (i.e., nucleotides 58-4374 of SEQ ID NO:2 that do not encode the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO:2. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO:2 (i.e., nucleotides 1-4374 of SEQ ID NO:2 that do not encode the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO:2.
[0266] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO:70. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO:70 (i.e., nucleotides 58-4374 of SEQ ID NO:70 that do not have nucleotides encoding the B domain or a B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:70. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:70 (i.e., nucleotides 58-4374 of SEQ ID NO:70 that do not have nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO:70. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO:70 (i.e., nucleotides 1-4374 of SEQ ID NO:70 that do not have nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO:70.
[0267] In some embodiments, the codon-optimized nucleotide sequence encoding an FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO:71. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO:71 (i.e., nucleotides 58-4374 of SEQ ID NO:71 that do not have nucleotides encoding the B domain or a B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:71. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:71 (i.e., nucleotides 58-4374 of SEQ ID NO:71 that do not have nucleotides encoding the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO:71. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO:71 (i.e., nucleotides 1-4374 of SEQ ID NO:71 that do not have nucleotides encoding the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO:71.
[0268] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO:3. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO:3 (i.e., nucleotides 58-4374 of SEQ ID NO:3 that do not encode the B domain or a B domain fragment). In certain embodiments, the nucleic acid sequence has at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO:3. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:3 (i.e., nucleotides 58-4374 of SEQ ID NO:3 that do not encode the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO:3. In still other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:3 (i.e., nucleotides 1-4374 of SEQ ID NO:3 that do not encode the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO:3.
[0269] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO:4. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO:4 (i.e., nucleotides 58-4374 of SEQ ID NO:4 that do not encode the B domain or a B domain fragment). In other embodiments, the nucleic acid sequence has at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO:4. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:4 (i.e., nucleotides 58-4374 of SEQ ID NO:4 that do not encode the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO:4. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO:4 (i.e., nucleotides 1-4374 of SEQ ID NO:4 that do not encode the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO:4.
[0270] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO:5. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 58-4374 of SEQ ID NO:5 that do not encode the B domain or a B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO:5. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 58-4374 of SEQ ID NO:5 that do not encode the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO:5. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1-4374 of SEQ ID NO:5 that do not encode the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO:5.
[0271] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO:6. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 58-4374 of SEQ ID NO:6 that do not encode the B domain or a B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO:6. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 58-4374 of SEQ ID NO:6 that do not encode the B domain or a B domain fragment) or nucleotides 58 to 4374 of SEQ ID NO:6. In still other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1-4374 of SEQ ID NO:6 that do not encode the B domain or a B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO:6.
[0272] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleic acid sequence encoding a signal peptide. In certain embodiments, the signal peptide is the FVIII signal peptide. In some embodiments, the nucleic acid sequence encoding the signal peptide is codon-optimized. In a particular embodiment, the nucleic acid sequence encoding the signal peptide has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity with (i) nucleotides 1 to 57 of SEQ ID NO:1; (ii) nucleotides 1 to 57 of SEQ ID NO:2; (iii) nucleotides 1 to 57 of SEQ ID NO:3; (iv) nucleotides 1 to 57 of SEQ ID NO:4; (v) nucleotides 1 to 57 of SEQ ID NO:5; (vi) nucleotides 1 to 57 of SEQ ID NO:6; (vii) nucleotides 1 to 57 of SEQ ID NO:70; (viii) nucleotides 1 to 57 of SEQ ID NO:71; or (ix) nucleotides 1 to 57 of SEQ ID NO:68.
[0273] SEQ ID NOs: 1-6, 70, and 71 are optimized versions of SEQ ID NO:16, which is the starting or "parental" or "wild-type" FVIII nucleotide sequence. SEQ ID NO:16 encodes B-domain-deleted human FVIII. Although SEQ ID NOs: 1-6, 70, and 71 are derived from a specific B-domain-deleted form of FVIII (SEQ ID NO:16), it should be understood that the present disclosure also includes optimized versions of nucleic acids encoding other versions of FVIII. For example, other versions of FVIII may include full-length FVIII, other B-domain-deleted FVIIIs (described herein), or other fragments of FVIII that retain FVIII activity.
[0274] In one embodiment, the gene cassette comprises an FVIII construct that includes the polynucleotide sequence (6526 nucleotides) listed in Tables 2A-2C. In one embodiment, the gene cassette comprises an FVIII construct that includes the polynucleotide sequence (6526 nucleotides) listed in Table 2B.
[0275] In certain embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or about 100% sequence identity with the nucleotide sequence of SEQ ID NO:179 or 182, wherein the isolated nucleic acid molecule retains the ability to express a functional FVIII protein.
[0276] Table 2A: Example FVIII construct (nucleotides 1 - 6526; SEQ ID NO: 110)
[0277]
[0278]
[0279]
[0280]
[0281] Table 2B: Example B19 - FVIII construct (nucleotides 1 - 6762; SEQ ID NO: 179)
[0282]
[0283]
[0284]
[0285]
[0286]
[0287]
[0288] Table 2C: Example GPV - FVIII construct (nucleotides 1 - 6830; SEQ ID NO: 182)
[0289]
[0290]
[0291]
[0292]
[0293]
[0294]
[0295]
[0296] A. Codon Adaptation Index
[0297] In one embodiment, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the human codon adaptation index of the codon-optimized nucleotide sequence is increased relative to SEQ ID NO:16. For example, the codon-optimized nucleotide sequence can have a human codon adaptation index that is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%), at least about 0.91 (91%), at least about 0.92 (92%), at least about 0.93 (93%), at least about 0.94 (94%), at least about 0.95 (95%), at least about 0.96 (96%), at least about 0.97 (97%), at least about 0.98 (98%), or at least about 0.99 (99%). In some embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index of at least about.88 (88%). In other embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index of at least about.91 (91%). In other embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index of at least about.91 (97%).
[0298] In a particular embodiment, a codon-optimized nucleotide sequence encoding an FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) nucleotides 1-1791 of SEQ ID NO:3; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO:16. In some embodiments, the nucleotide sequence has a human codon adaptation index of at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%) or at least about.91 (91%). In a particular embodiment, the nucleotide sequence has a human codon adaptation index of at least about.88 (88%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about.91 (91%).
[0299] In another embodiment, a codon-optimized nucleotide sequence encoding an FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 or (ii) 1792-2277 and 2320-4374 of SEQ ID NO:6; wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO:16. In some embodiments, the nucleotide sequence has a human codon adaptation index of at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%) or at least about 0.88 (88%). In a particular embodiment, the nucleotide sequence has a human codon adaptation index of at least about.83 (83%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about.88 (88%).
[0300] In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of the amino acid sequences selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, and 71 nucleotides 58-2277 and 2320-4374 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71, which do not have nucleotides encoding the B domain or a B domain fragment) having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; and wherein the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO: 16. In some embodiments, the nucleotide sequence has a human codon adaptation index of at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), or at least about 0.88 (88%). In a particular embodiment, the nucleotide sequence has a human codon adaptation index of at least about.75 (75%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about.83 (83%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about.88 (88%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about.91 (91%). In another embodiment, the nucleotide sequence has a human codon adaptation index of at least about.97 (97%).
[0301] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide of the present disclosure has an increased frequency of optimal codons (FOP) relative to SEQ ID NO: 16. In certain embodiments, the FOP of the codon-optimized nucleotide sequence encoding the FVIII polypeptide is at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 64, at least about 65, at least about 70, at least about 75, at least about 79, at least about 80, at least about 85, or at least about 90.
[0302] In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide of the present disclosure has an increased relative codon usage (RCSU) relative to SEQ ID NO:16. In some embodiments, the RCSU of the isolated nucleic acid molecule is greater than 1.5. In other embodiments, the RCSU of the isolated nucleic acid molecule is greater than 2.0. In certain embodiments, the RCSU of the isolated nucleic acid molecule is at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2.0, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, or at least about 2.7.
[0303] In still other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide of the present disclosure has a reduced effective number of codons relative to SEQ ID NO:16. In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has an effective number of codons less than about 50, less than about 45, less than about 40, less than about 35, less than about 30, or less than about 25. In a particular embodiment, the isolated nucleic acid molecule has an effective number of codons of about 40, about 35, about 30, about 25, or about 20.
[0304] B. G / C Content Optimization
[0305] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding the FVIII polypeptide, wherein the codon-optimized nucleotide sequence has a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO:16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, or at least about 60%.
[0306] In a particular embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity with (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) nucleotides 1-1791 of SEQ ID NO:3; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO:16. In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57% or at least about 58%. In a particular embodiment, the nucleotide sequence encoding the polypeptide having FVIII activity has a G / C content of at least about 58%.
[0307] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity with (i) nucleotides 1792-4374 of SEQ ID NO:5; (ii) nucleotides 1792-4374 of SEQ ID NO:6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1792-4374 of SEQ ID NO:5 that do not contain nucleotides encoding the B domain or a B domain fragment), or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:6 that do not contain nucleotides encoding the B domain or a B domain fragment); wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO:16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56% or at least about 57%. In a particular embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 52%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 55%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 57%.
[0308] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-4374 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71 that do not contain nucleotides encoding the B domain or a B domain fragment); and wherein the nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 45%. In a particular embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 52%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 55%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 57%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content of at least about 58%. In still other embodiments, the n codon-optimized nucleotide sequences encoding the FVIII polypeptide have a G / C content of at least about 60%.
[0309] The "G / C content" (or guanine-cytosine content), or "percentage of G / C nucleotides" refers to the percentage of nitrogenous bases (i.e., guanine or cytosine) in a DNA molecule. The G / C content can be calculated using the following formula:
[0310]
[0311] The G / C content of human genes is highly heterogeneous, with some genes having a G / C content as low as 20% and other genes having a G / C content as high as 95%. Generally, genes rich in G / C are expressed at higher levels. Indeed, it has been shown that increasing the G / C content of a gene can lead to increased gene expression, which is mainly due to increased transcription and higher steady-state mRNA levels. See Kudla et al., PLoS Biol., 4(6):e180 (2006).
[0312] C. Matrix attachment region-like sequences
[0313] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO:16. In other embodiments, the codon-optimized nucleotide sequence encoding an FVIII polypeptide contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding an FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding an FVIII polypeptide does not contain a MARS / ARS sequence.
[0314] In one particular embodiment, the codon-optimized nucleotide sequence encoding an FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) nucleotides 1-1791 of SEQ ID NO:3; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO:16. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains at most 6, at most 5, at most 4, at most 3, or at most 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding an FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding an FVIII polypeptide does not contain a MARS / ARS sequence.
[0315] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity with (i) nucleotides 1792-4374 of SEQ ID NO:5; (ii) nucleotides 1792-4374 of SEQ ID NO:6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1792-4374 of SEQ ID NO:5 that do not contain nucleotides encoding the B domain or a B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:6 that do not contain nucleotides encoding the B domain or a B domain fragment); wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO:16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 6, at most 5, at most 4, at most 3 or at most 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 1 MARS / ARS sequence. In yet another other embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a MARS / ARS sequence.
[0316] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71, or (ii) nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71 that do not have nucleotides encoding the B domain or a fragment of the B domain); and wherein the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence polypeptide encoding FVIII contains at most 6, at most 5, at most 4, at most 3 or at most 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding FVIII contains at most 1 MARS / ARS sequence. In yet another other embodiment, the codon-optimized nucleotide sequence encoding FVIII does not contain a MARS / ARS sequence.
[0317] AT-rich elements have been identified in the human FVIII nucleotide sequence that share sequence similarity with the Saccharomyces cerevisiae autonomous replication sequence (ARS) and matrix attachment regions (MAR). (Fallux et al., Mol. Cell. Biol. 16:4264-4272 (1996). One of these elements has been shown to bind nuclear factors in vitro and inhibit the expression of the chloramphenicol acetyltransferase (CAT) reporter gene. (Fallux et al., Mol. Cell. Biol. 16:4264-4272 (1996). It has been hypothesized that these sequences may promote transcriptional repression of the human FVIII gene. Thus, in one embodiment, all MAR / ARS sequences are abrogated in the codon-optimized nucleotide sequence encoding the FVIII polypeptide of the present disclosure. There are 4 MAR / ARS ATATTT sequences (SEQ ID NO: 21) and 3 MAR / ARS AAATAT sequences (SEQ ID NO: 22) present in the parental FVIII sequence (SEQ ID NO: 16). All of these sites were mutated to disrupt the MAR / ARS sequences in the optimized FVIII sequences (SEQ ID NO: 1-6). The position of each of these elements, and the sequence of the corresponding nucleotides in the optimized sequences are shown in Table 3 below.
[0318] Table 3: Summary of Inhibitory Element Variations
[0319]
[0320] D. Destabilizing Sequences
[0321] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO:16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains up to 9, up to 8, up to 7, up to 6, or up to 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains up to 4, up to 3, up to 2, or up to 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain any destabilizing elements.
[0322] In a particular embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58 - 1791 of SEQ ID NO:3; (ii) nucleotides 1 - 1791 of SEQ ID NO:3; (iii) nucleotides 58 - 1791 of SEQ ID NO:4; or (iv) nucleotides 1 - 1791 of SEQ ID NO:4; wherein the N-terminal portion and the C-terminal portion both have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO:16. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains up to 9, up to 8, up to 7, up to 6, or up to 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains up to 4, up to 3, up to 2, or up to 1 destabilizing element. In yet another other embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain any destabilizing elements.
[0323] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO:5; (ii) nucleotides 1792-4374 of SEQ ID NO:6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1792-4374 of SEQ ID NO:5 that do not contain nucleotides encoding the B domain or a B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:6 that do not contain nucleotides encoding the B domain or a B domain fragment); wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO:16. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains at most 9, at most 8, at most 7, at most 6 or at most 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 4, at most 3, at most 2 or at most 1 destabilizing element. In yet another other embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a destabilizing element.
[0324] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 that do not contain nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6 or at most 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 4, at most 3, at most 2 or at most 1 destabilizing element. In yet another other embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a destabilizing element.
[0325] There are 10 destabilizing elements in the parental FVIII sequence (SEQ ID NO: 16): 6 ATTTA sequences (SEQ ID NO: 23) and 4 TAAAT sequences (SEQ ID NO: 24). In one embodiment, the sequences at these sites are mutated to disrupt the destabilizing elements in the optimized FVIII SEQ ID NOs: 1-6, 70 and 71. The position of each of these elements, and the sequence of the corresponding nucleotides in the optimized sequences, are shown in Table 3.
[0326] E. Potential promoter binding sites
[0327] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO:16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding sites. In yet another other embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain potential promoter binding sites.
[0328] In a particular embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) nucleotides 1-1791 of SEQ ID NO:3; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO:16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding sites. In yet another other embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain potential promoter binding sites.
[0329] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity with (i) nucleotides 1792-4374 of SEQ ID NO:5; (ii) nucleotides 1792-4374 of SEQ ID NO:6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1792-4374 of SEQ ID NO:5 that do not contain nucleotides encoding the B domain or a B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:6 that do not contain nucleotides encoding the B domain or a B domain fragment); wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO:16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6 or at most 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 4, at most 3, at most 2 or at most 1 potential promoter binding site. In yet another other embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a potential promoter binding site.
[0330] In other embodiments, the gene cassette encoding the FVIII polypeptide comprises a codon-optimized nucleotide sequence, wherein the nucleotide sequence comprises nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, and 71, or nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, and 71 that do not contain nucleotides encoding the B domain or a B domain fragment), and has at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the nucleic acid sequence; and wherein the codon-optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 9, at most 8, at most 7, at most 6, or at most 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide contains at most 4, at most 3, at most 2, or at most 1 potential promoter binding site. In yet another other embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a potential promoter binding site.
[0331] The TATA box is a regulatory sequence commonly found in the promoter region of eukaryotic cells. It serves as a binding site for the TATA-binding protein (TBP), a general transcription factor. The TATA box typically contains the sequence TATAA (SEQ ID NO:28) or a close variant. However, a TATA box within a coding sequence can inhibit translation of the full-length protein. There are 10 potential promoter-binding sequences in the wild-type BDD FVIII sequence (SEQ ID NO:16): 5 TATAA sequences (SEQ ID NO:28) and 5 TTATA sequences (SEQ ID NO:29). In some embodiments, at least 1, at least 2, at least 3, or at least 4 of the promoter-binding sites are abolished in the FVIII gene of the present disclosure. In some embodiments, at least 5 of the promoter-binding sites are abolished in the FVIII gene of the present disclosure. In other embodiments, at least 6, at least 7, or at least 8 of the promoter-binding sites are abolished in the FVIII gene of the present disclosure. In one embodiment, at least 9 of the promoter-binding sites are abolished in the FVIII gene of the present disclosure. In a particular embodiment, all of the promoter-binding sites are abolished in the FVIII gene of the present disclosure. The location of each potential promoter-binding site and the sequence of the corresponding nucleotides in the optimized sequence are shown in Table 3.
[0332] F. Other cis-acting negative regulatory elements
[0333] In addition to the MAR / ARS sequences, destabilizing elements, and potential promoter sites described above, several other potential inhibitory sequences can be identified in the wild-type BDD FVIII sequence (SEQ ID NO:16). Two AU-rich sequence elements (AREs) (ATTTTATT (SEQ ID NO:30); and ATTTTTAA (SEQ ID NO:31)) can be identified in the non-optimized BDD FVIII sequence, along with a polyA site (AAAAAAA; SEQ ID NO:26), a polyT site (TTTTTT; SEQ ID NO:25), and a splice site (GGTGAT; SEQ ID NO:27). One or more of these elements can be deleted from the optimized FVIII sequence. The location of each of these sites and the sequence of the corresponding nucleotides in the optimized sequence are shown in Table 3.
[0334] In certain embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) nucleotides 1-1791 of SEQ ID NO:3; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, e.g., splice sites, poly-T sequences, poly-A sequences, ARE sequences, or any combination thereof.
[0335] In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO:5; (ii) nucleotides 1792-4374 of SEQ ID NO:6; (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1792-4374 of SEQ ID NO:5 that do not have nucleotides encoding the B domain or a B domain fragment); or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:6 that do not have nucleotides encoding the B domain or a B domain fragment); wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, e.g., splice sites, poly-T sequences, poly-A sequences, ARE sequences, or any combination thereof.
[0336] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71 that do not contain nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, e.g., splice sites, poly T sequences, poly A sequences, ARE sequences or any combination thereof.
[0337] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity with (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) nucleotides 1-1791 of SEQ ID NO:3; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain the splice site GGTGAT (SEQ ID NO:27). In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity with (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) SEQ ID NO:3 nucleotides 1-1791; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain a poly-T sequence (SEQ ID NO:25).In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity with (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) nucleotides 1-1791 of SEQ ID NO:3; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; wherein the N-terminal portion and the C-terminal portion simultaneously have FVIII polypeptide activity; and wherein the codon-optimized nucleotide sequence does not contain a polyA sequence (SEQ ID NO:26). In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a first nucleic acid sequence encoding the N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding the C-terminal portion of the FVIII polypeptide; wherein the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity with (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) nucleotides 1-1791 of SEQ ID NO:3; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; and wherein the codon-optimized nucleotide sequence does not contain an ARE element (SEQ ID NO:30 or SEQ ID NO:31).
[0338] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71 that do not contain nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence does not contain the splice site GGTGAT (SEQ ID NO: 27). In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71 that do not contain nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence does not have a poly-T sequence (SEQ ID NO: 25).In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71 that do not contain nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence does not contain a polyA sequence (SEQ ID NO: 26). In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding an FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence identity to (i) nucleotides 58-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 or (ii) nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70 or 71 that do not contain nucleotides encoding the B domain or a B domain fragment); and wherein the codon-optimized nucleotide sequence does not contain an ARE element (SEQ ID NO: 30 or SEQ ID NO: 31).
[0339] In other embodiments, the optimized FVIII sequences of the present disclosure do not contain one or more of antiviral motifs, stem-loop structures, and repeats.
[0340] In still other embodiments, the nucleotides around the transcription start site are changed to the kozak consensus sequence (GCCGCCACCATGC (SEQ ID NO: 32), where the underlined nucleotide is the start codon). In other embodiments, restriction sites can be added or removed to facilitate the cloning process.
[0341] b. FIX and polynucleotide sequences encoding FIX proteins
[0342] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a FIX polypeptide. In some embodiments, the FIX polypeptide comprises FIX or a variant or fragment thereof, wherein FIX or the variant or fragment thereof has FIX activity.
[0343] Human FIX is a serine protease and an important component of the intrinsic pathway of the blood coagulation cascade. As used herein, "Factor IX" or "FIX" refers to the blood coagulation factor protein and its species and sequence variants, and includes, but is not limited to, the 461 single-chain amino acid sequence of the human FIX precursor polypeptide ("prepro"), the 415 single-chain amino acid sequence of mature human FIX (SEQ ID NO: 125), and the R338L FIX (Padua) variant (SEQ ID NO: 126). FIX includes any form of FIX molecule having the typical characteristics of blood coagulation FIX. As used herein, "Factor IX" and "FIX" are intended to encompass polypeptides comprising the domain Gla (region containing γ-carboxyglutamic acid residues), EGF1 and EGF2 (regions containing sequences homologous to human epidermal growth factor), the activation peptide ("AP", which is formed by residues R136-R180 of mature FIX), and the C-terminal protease domain ("Pro"), or synonyms of these domains known in the art, or may be truncated fragments or sequence variants that retain at least a portion of the biological activity of the native protein. FIX or sequence variants have been cloned, as described in U.S. Patent Nos. 4,770,999 and 7,700,734, and the cDNA encoding human FIX has been isolated, characterized, and cloned into expression vectors (see, e.g., Choo et al., Nature 299:178-180 (1982); Fair et al., Blood 64:194-204 (1984); and Kurachi et al., Proc. Natl. Acad. Sci., U.S.A. 79:6461-6464 (1982)). A particular variant of FIX, the R338L FIX (Padua) variant (SEQ ID NO: 2) (which was characterized by Simioni et al., 2009) contains a gain-of-function mutation that is associated with an almost 8-fold increase in activity of the Padua variant relative to native FIX (Table 4). FIX variants can also include any FIX polypeptide having one or more conservative amino acid substitutions that do not affect the FIX activity of the FIX polypeptide. In some embodiments, the FIX variant comprises rFIX-albumin fused by a cleavable linker, e.g., See US 7,939,632, which is incorporated herein by reference in its entirety.
[0344] Table 4: Example FIX Sequences
[0345]
[0346]
[0347]
[0348]
[0349] *Grey shading = signal peptide; Underline = XTEN sequence; Bold = Fc.
[0350] **SEQ ID NO:67 in U.S. Patent No. 9,856,468, which patent is incorporated herein by reference in its entirety.
[0351] The FIX polypeptide is 55 kDa and is synthesized as a prepropolypeptide chain (SEQ ID NO:125) consisting of the following three regions: a 28 - amino acid signal peptide (amino acids 1 to 28 of SEQ ID NO:127), an 18 - amino acid propeptide (amino acids 29 to 46), which requires γ - carboxylation of glutamate residues, and a 415 - amino acid mature factor IX (SEQ ID NO:125 or 126). The propeptide is an 18 - amino acid residue sequence at the N - terminus of the γ - carboxyglutamate domain. The propeptide binds to the vitamin K - dependent γ - carboxylase and is then cleaved from the FIX precursor polypeptide by an endogenous protease, most likely PACE (paired basic amino acid cleaving enzyme), also known as furin or PCSK3. In the absence of γ - carboxylation, the Gla domain cannot bind calcium and cannot assume the correct conformation necessary to anchor the protein to the negatively charged phospholipid surface, rendering factor IX non - functional. Even when the Gla domain is carboxylated, it depends on the cleavage of the propeptide for proper function because the retained propeptide interferes with the conformational changes in the Gla domain necessary for optimal binding to calcium and phospholipids. In humans, the resulting mature factor IX is secreted by hepatocytes into the bloodstream in the form of an inactive zymogen, which is a single - chain protein of 415 amino acid residues and contains approximately 17% carbohydrate by weight (Schmidt, A.E., et al. (2003) Trends Cardiovasc Med, 13:39).
[0352] Mature FIX consists of several domains in an N - to - C configuration: GLA domain, EGF1 domain, EGF2 domain, activation peptide (AP) domain, and protease domain (or catalytic domain). A short linker joins the EGF2 domain to the AP domain. FIX contains two activation peptides formed by R145 - A146 and R180 - V181, respectively. Upon activation, single - chain FIX becomes a two - chain molecule, where the two chains are linked by a disulfide bond. The coagulation factor can be engineered by replacing its activation peptide, thus generating altered activation specificity. In mammals, mature FIX must be activated by activated factor XI to produce factor IXa. After activation of FIX to FIXa, the protease domain provides the catalytic activity of FIX. Activated factor VIII (FVIIIa) is the fully expressed specific cofactor for FIXa activity.
[0353] In certain embodiments, the FIX polypeptide comprises the Thr148 allelic form of plasma - derived FIX and has structural and functional characteristics similar to endogenous FIX.
[0354] Many functional FIX variants are known in the art. International Publication No. WO 02 / 040544 A3 discloses mutants that exhibit enhanced resistance to the inhibitory effects of heparin on pages 4, lines 9-30 and 15, lines 6-31. International Publication No. WO 03 / 020764 A2 discloses FIX mutants with reduced T cell immunogenicity in Tables 2 and 3 (pages 14-24) and 12, lines 1-27. International Publication No. WO 2007 / 149406 A2 discloses functional mutant FIX molecules that exhibit increased protein stability, increased in vivo and in vitro half-lives, and increased resistance to proteases on pages 4, line 1 to 19, line 11. WO 2007 / 149406 A2 also discloses chimeric and other variant FIX molecules on pages 19, line 12 to 20, line 9. International Publication No. WO 08 / 118507 A2 discloses FIX mutants that exhibit increased coagulation activity on pages 5, line 14 to 6, line 5. International Publication No. WO 09 / 051717 A2 discloses FIX mutants with an increased number of N-linked and / or O-linked glycosylation sites (which result in increased half-life and / or recovery) on pages 9, line 11 to 20, line 2. International Publication No. WO 09 / 137254 A2 also discloses factor IX mutants with an increased number of glycosylation sites in paragraphs
[006] on page 2 to
[011] on page 5 and paragraphs
[044] on page 16 to
[057] on page 24. International Publication No. WO 09 / 130198 A2 discloses functional mutant FIX molecules with an increased number of glycosylation sites (which result in increased half-life) on pages 4, line 26 to 12, line 6. International Publication No. WO 09 / 140015 A2 discloses functional FIX mutants with an increased number of Cys residues (which can be used for conjugation to a polymer (e.g., PEG)) in paragraphs
[0043] on page 11 to
[0053] on page 13. The FIX polypeptides described in International Application No. PCT / US2011 / 043569, filed on July 11, 2011 and published as WO 2012 / 006624 on January 12, 2012, are also incorporated herein by reference in their entirety. In some embodiments, the FIX polypeptide comprises a FIX polypeptide fused to albumin, e.g., FIX-albumin. In certain embodiments, the FIX polypeptide is or rIX-FP.
[0355] In addition, hundreds of non-functional mutations in FIX have been identified in hemophilia subjects, many of which are disclosed in Table 6 on pages 11-14 of International Publication No. WO 09 / 137254 A2. Such non-functional mutations are not included in the present invention, but provide additional guidance on which mutations are more or less likely to result in a functional FIX polypeptide.
[0356] In one embodiment, the FI polypeptide (or the factor IX portion of the fusion polypeptide) comprises an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identical to the sequence shown in SEQ ID NO: 1 or 2 (amino acids 1 to 415 of SEQ ID NO: 125 or 126), or alternatively, to a sequence having a propeptide sequence, or a sequence having a propeptide and a signal sequence (full-length FIX). In another embodiment, the FIX polypeptide comprises an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identical to the sequence shown in SEQ ID NO: 2.
[0357] FIX procoagulant activity is expressed in international units (IU). 1 IU of FIX activity approximately corresponds to the amount of FIX in one milliliter of normal human plasma. Several assays can be used to measure FIX activity, including one-stage clotting assays (activated partial thromboplastin time; aPTT), thrombin generation assays (TGA), and rotational thromboelastometry. The present invention contemplates sequences that are homologous to the FIX sequence, natural sequence fragments (such as from humans, non-human primates, mammals (including livestock)), and sequences of non-natural sequence variants that retain at least a portion of the biological activity or biological function of FIX and / or can be used to prevent, treat, mediate, or ameliorate a disease, defect, disorder, or condition related to a blood coagulation factor (e.g., bleeding episodes related to trauma, surgery, blood coagulation factor deficiencies). Sequences homologous to human FIX can be found by standard homology search techniques (such as NCBI BLAST).
[0358] In certain embodiments, the FIX sequence is codon-optimized. Examples of codon-optimized FIX sequences include, but are not limited to, SEQ ID NOs: 1 and 54-58 in International Publication No. WO 2016 / 004113 A1, which is incorporated herein by reference in its entirety.
[0359] c. FVII and polynucleotide sequences encoding the FVII protein
[0360] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a factor VII polypeptide. In some embodiments, the FVII polypeptide comprises FVII or a variant or fragment thereof, wherein the variant or fragment thereof has FVII activity.
[0361] "Factor VII" ("FVII" or "F7"; also known as factor 7, coagulation factor VII, serum factor VII, serum prothrombin conversion accelerator, SPCA, proconvertin, and eptacog α) is a serine protease that is part of the coagulation cascade. In one embodiment, the coagulation factor in the nucleic acids described herein is FVII. Recombinant activated factor VII ("FVII") has been widely used to treat major bleeding, such as that which occurs in patients with hemophilia A or B, factor XI, FVII deficiency, platelet function defects, thrombocytopenia, or von Willebrand disease.
[0362] Recombinant activated FVII (rFVIIa; ) is used to treat bleeding episodes in patients who: (i) have neutralizing antibodies (inhibitors) against FVIII or FIX, (ii) have FVII deficiency, or (iii) are undergoing surgical treatment and have inhibitors of hemophilia A or B. However, it has shown poor efficacy. Due to its low affinity for activated platelets, short half-life, and poor enzymatic activity in the absence of tissue factor, high concentrations of repeated doses of FVIIa are typically required to control bleeding. Thus, there is an unmet medical need for better treatment and prevention options for hemophilia patients with FVIII and FIX inhibitors and / or with FVII deficiency.
[0363] In one embodiment, the gene cassette encodes the mature form of FVII or a variant thereof. FVII comprises a Gla domain, two EGF domains (EGF-1 and EGF-2), and a serine protease domain (or peptidase S1 domain), which is highly conserved among all members of the peptidase S1 family of serine proteases (such as, for example, chymotrypsin). FVII occurs as a single-chain zymogen (i.e., activatable FVII) and a fully activated two-chain form.
[0364] C. Growth Factor
[0365] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a growth factor. The growth factor can be selected from any growth factor known in the art. In some embodiments, the growth factor is a hormone. In other embodiments, the growth factor is a cytokine. In some embodiments, the growth factor is a chemokine.
[0366] In some embodiments, the growth factor is adrenomedullin (AM). In some embodiments, the growth factor is angiopoietin (Ang). In some embodiments, the growth factor is autocrine motility factor. In some embodiments, the growth factor is bone morphogenetic protein (BMP). In some embodiments, the BMP is selected from BMP2, BMP4, BMP5, and BMP7. In some embodiments, the growth factor is a ciliary neurotrophic factor family member. In some embodiments, the ciliary neurotrophic factor family member is selected from ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), and interleukin-6 (IL-6). In some embodiments, the growth factor is a colony-stimulating factor. In some embodiments, the colony-stimulating factor is selected from macrophage colony-stimulating factor (m-CSF), granulocyte colony-stimulating factor (G-CSF), and granulocyte-macrophage colony-stimulating factor (GM-CSF). In some embodiments, the growth factor is epidermal growth factor (EGF). In some embodiments, the growth factor is ephrin. In some embodiments, the ephrin is selected from ephrin A1, ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1, ephrin B2, and ephrin B3. In some embodiments, the growth factor is erythropoietin (EPO). In some embodiments, the growth factor is fibroblast growth factor (FGF). In some embodiments, the FGF is selected from FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, and FGF23. In some embodiments, the growth factor is fetal bovine serum (FBS). In some embodiments, the growth factor is a GDNF family member. In some embodiments, the GDNF family member is selected from glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, and Artemin. In some embodiments, the growth factor is growth differentiation factor-9 (GDF9). In some embodiments, the growth factor is hepatocyte growth factor (HGF). In some embodiments, the growth factor is hepatoma-derived growth factor (HDGF). In some embodiments, the growth factor is insulin. In some embodiments, the growth factor is an insulin-like growth factor. In some embodiments, the insulin-like growth factor is insulin-like growth factor-1 (IGF-1) or IGF-2. In some embodiments, the growth factor is interleukin (IL). In some embodiments, the IL is selected from IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, and IL-7.In some embodiments, the growth factor is keratinocyte growth factor (KGF). In some embodiments, the growth factor is motility stimulating factor (MSF). In some embodiments, the growth factor is macrophage stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)). In some embodiments, the growth factor is myostatin (GDF-8). In some embodiments, the growth factor is neuregulin. In some embodiments, neuregulin is selected from neuregulin 1 (NRG1), NRG2, NRG3, and NRG4. In some embodiments, the growth factor is neurotrophin. In some embodiments, the growth factor is brain-derived neurotrophic factor (BDNF). In some embodiments, the growth factor is nerve growth factor (NGF). In some embodiments, NGF is neurotrophin 3 (NT-3) or NT-4. In some embodiments, the growth factor is placental growth factor (PGF). In some embodiments, the growth factor is platelet-derived growth factor (PDGF). In some embodiments, the growth factor is Renalase (RNLS). In some embodiments, the growth factor is T cell growth factor (TCGF). In some embodiments, the growth factor is thrombopoietin (TPO). In some embodiments, the growth factor is transforming growth factor. In some embodiments, the transforming growth factor is transforming growth factor α (TGF-α) or TGF-β. In some embodiments, the growth factor is tumor necrosis factor-α (TNF-α). In some embodiments, the growth factor is vascular endothelial growth factor (VEGF).
[0367] D. MicroRNA (miRNA)
[0368] MicroRNA (miRNA) are small non-coding RNA molecules (about 18 - 22 nucleotides) that negatively regulate gene expression by inhibiting translation or inducing messenger RNA (mRNA) degradation. Since their discovery, miRNAs have been involved in a variety of cellular processes, including apoptosis, differentiation, and cell proliferation, and they have been shown to play a key role in carcinogenesis. The ability of miRNAs to regulate gene expression makes in vivo expression of miRNAs a valuable tool in gene therapy.
[0369] Certain aspects of the present disclosure relate to plasmid-like nucleic acid molecules comprising a first ITR, a second ITR, and a gene cassette encoding an miRNA, wherein the first ITR and / or the second ITR is a non-adeno-associated virus ITR (e.g., the first ITR and / or the second ITR is from a non-AAV). The miRNA can be any miRNA known in the art. In some embodiments, the miRNA downregulates the expression of a target gene. In certain embodiments, the target gene is selected from SOD1, HTT, RHO, or any combination thereof.
[0370] In some embodiments, the gene cassette encodes an miRNA. In some embodiments, the gene cassette encodes more than one miRNA. In some embodiments, the gene cassette encodes two or more different miRNAs. In some embodiments, the gene cassette encodes two or more copies of the same miRNA. In some embodiments, the gene cassette encodes two or more variants of the same therapeutic protein. In certain embodiments, the gene cassette encodes one or more miRNAs and one or more therapeutic proteins.
[0371] In some embodiments, the miRNA is a naturally occurring miRNA. In some embodiments, the miRNA is an engineered miRNA. In some embodiments, the miRNA is an artificial miRNA. In certain embodiments, the miRNA comprises the engineered miRHTT miRNA disclosed by Evers et al., Molecular Therapy 26(9):1-15 (epub ahead of print June 2018). In certain embodiments, the miRNA comprises the artificial miR SOD1 miRNA disclosed by Dirren et al., Annals of Clinical and Translational Neurology 2(2):167-84 (February 2015). In certain embodiments, the miRNA comprises miR-708, which targets RHO (see Behrman et al., JCB 192(6):919-27 (2011).
[0372] In some embodiments, the miRNA upregulates the expression of a gene by downregulating the expression of an inhibitor of the gene. In some embodiments, the inhibitor is natural, e.g., wild-type, inhibitor. In some embodiments, the inhibitor is produced by a mutated, heterologous, and / or misexpressed gene.
[0373] E. Heterologous moiety
[0374] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises at least one heterologous moiety. In some embodiments, the heterologous moiety is fused to the N-terminus or C-terminus of the therapeutic protein. In other embodiments, the heterologous moiety is inserted between two amino acids within the therapeutic protein.
[0375] In some embodiments, the therapeutic protein comprises an FVIII polypeptide and a heterologous moiety that is inserted between two amino acids within the FVIII polypeptide. In some embodiments, the heterologous moiety is inserted within the FVIII polypeptide at one or more insertion sites selected from Table 5. In some embodiments, the heterologous amino acid sequence can be inserted within the coagulation factor polypeptide encoded by the nucleic acid molecule of the present disclosure at any site disclosed in International Publication No. WO2013 / 123457 A1, WO 2015 / 106052 A1, or U.S. Publication No. 2015 / 0158929 A1. In a particular embodiment, the therapeutic protein comprises FVIII and a heterologous moiety, wherein the heterologous moiety is inserted within FVIII immediately downstream of amino acid 745 relative to mature FVIII. In a particular embodiment, the therapeutic protein comprises FVIII and XTEN, wherein XTEN is inserted within FVIII immediately downstream of amino acid 745 relative to mature FVIII. In a particular embodiment, FVIII comprises a deletion of amino acids 746 - 1646 (corresponding to mature human FVIII (SEQ ID NO:15)), and the heterologous moiety is inserted immediately downstream of amino acid 745 (corresponding to mature human FVIII (SEQ ID NO:15)).
[0376] Table 5: FVIII Heterologous Moiety Insertion Sites
[0377] Insertion site Domain Insertion site Domain Insertion site Domain 3 A1 375 A2 1749 A3 18 A1 378 A2 1796 A3 22 A1 399 A2 1802 A3 26 A1 403 A2 1827 A3 40 A1 409 A2 1861 A3 60 A1 416 A2 1896 A3 65 A1 442 A2 1900 A3 81 A1 487 A2 1904 A3 116 A1 490 A2 1905 A3 119 A1 494 A2 1910 A3 130 A1 500 A2 1937 A3 188 A1 518 A2 2019 A3 211 A1 599 A2 2068 C1 216 A1 603 A2 2111 C1 220 A1 713 A2 2120 C1 224 A1 745 B 2171 C2 230 A1 1656 a3 region 2188 C2 333 A1 1711 A3 2227 C2 336 A1 1720 A3 2332 CT 339 A1 1725 A3
[0378] In some embodiments, the therapeutic protein comprises a FIX polypeptide and a heterologous moiety that is inserted between two amino acids within the FIX polypeptide. In some embodiments, the heterologous moiety is inserted within the FIX polypeptide at one or more insertion sites selected from Table 5. In some embodiments, the heterologous amino acid sequence can be inserted within the coagulation factor polypeptide encoded by the nucleic acid molecule of the present disclosure at any site disclosed in International Application No. PCT / US2017 / 015879, which is incorporated herein by reference in its entirety. In a particular embodiment, the therapeutic protein comprises a FIX polypeptide and a heterologous moiety, wherein the heterologous moiety is inserted within the FIX polypeptide immediately downstream of amino acid 166 relative to mature FIX. In a particular embodiment, the therapeutic protein comprises a FIX polypeptide and XTEN, wherein XTEN is inserted within FIX immediately downstream of amino acid 166 relative to mature FVIII.
[0379] Table 6: FIX Heterologous Moiety Insertion Sites
[0380] Insertion site Domain Insertion site Domain Insertion site Domain 52 EGF1 149 AP 257 Catalytic 59 EGF1 162 AP 265 Catalytic 66 EGF1 166 AP 277 Catalytic 80 EGF1 174 AP 283 Catalytic 85 EGF2 188 Catalytic 292 Catalytic 89 EGF2 202 Catalytic 316 Catalytic 103 EGF2 224 Catalytic 341 Catalytic 105 EGF2 226 Catalytic 354 Catalytic 113 EGF2 228 Catalytic 392 Catalytic 129 Linker 230 Catalytic 403 Catalytic 142 Linker 240 Catalytic 413 Catalytic
[0381] In other embodiments, the therapeutic proteins of the present disclosure further comprise 2, 3, 4, 5, 6, 7, or 8 heterologous nucleotide sequences. In some embodiments, all of the heterologous moieties are the same. In some embodiments, at least one heterologous moiety is different from the other heterologous moieties. In some embodiments, the present disclosure can comprise 2, 3, 4, 5, 6, or more than 7 heterologous moieties in tandem.
[0382] In some embodiments, the heterologous moiety increases the half-life of the therapeutic protein (which is a "half-life extender").
[0383] In some embodiments, the heterologous moiety is a peptide or polypeptide that has unstructured or structured features associated with increased in vivo half-life when incorporated into the proteins of the present disclosure. Non-limiting examples include albumin, albumin fragments, the Fc fragment of an immunoglobulin, the C-terminal peptide (CTP) of the β subunit of human chorionic gonadotropin, HAP sequences, XTEN sequences, transferrin or fragments thereof, PAS polypeptides, polyglycine linkers, polyserine linkers, albumin-binding moieties, or any fragment, derivative, variant, or combination of these polypeptides. In a specific embodiment, the heterologous amino acid sequence is an immunoglobulin constant region or a portion thereof, transferrin, albumin, or a PAS sequence. In some aspects, the heterologous moiety includes von Willebrand factor or a fragment thereof. In other related aspects, the heterologous moiety can include an attachment site (e.g., a cysteine amino acid) for a non-polypeptide moiety (such as polyethylene glycol (PEG), hydroxyethyl starch (HES), polysialic acid, or a derivative, variant, or combination of these elements). In some aspects, the heterologous moiety contains a cysteine amino acid that serves as an attachment point for a non-polypeptide moiety (such as polyethylene glycol (PEG), hydroxyethyl starch (HES), polysialic acid, or a derivative, variant, or combination of these elements).
[0384] In a specific embodiment, the first heterologous moiety is a half-life extending molecule that is known in the art, and the second heterologous moiety is a half-life extending molecule that is known in the art. In certain embodiments, the first heterologous moiety (e.g., the first Fc portion) and the second heterologous moiety (e.g., the second Fc portion) associate with each other to form a dimer. In one embodiment, the second heterologous moiety is a second Fc portion, wherein the second Fc portion is linked or associated with the first heterologous moiety (e.g., the first Fc portion). For example, the second heterologous moiety (e.g., the second Fc portion) is linked to the first heterologous moiety (e.g., the first Fc portion) by a linker, or associates with the first heterologous moiety by a non-covalent bond.
[0385] In some embodiments, the heterologous moiety is a polypeptide comprising at least about 10, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, at least about 2000, at least about 2500, at least about 3000, or at least about 4000 amino acids or consisting essentially of or consisting of the same. In other embodiments, the heterologous moiety is a polypeptide comprising from about 100 to about 200 amino acids, from about 200 to about 300 amino acids, from about 300 to about 400 amino acids, from about 400 to about 500 amino acids, from about 500 to about 600 amino acids, from about 600 to about 700 amino acids, from about 700 to about 800 amino acids, from about 800 to about 900 amino acids, or from about 900 to about 1000 amino acids or consisting essentially of or consisting of the same.
[0386] In certain embodiments, the heterologous moiety improves one or more pharmacokinetic properties of the therapeutic protein without significantly affecting its biological activity or function.
[0387] In certain embodiments, the heterologous moiety increases the in vivo and / or in vitro half-life of the therapeutic protein of the present disclosure. In other embodiments, the heterologous moiety promotes visualization or localization of the therapeutic protein of the present disclosure or a fragment thereof (e.g., a fragment comprising the heterologous moiety after proteolytic cleavage of the FVIII protein). Visualization and / or localization of the therapeutic protein of the present disclosure or a fragment thereof can be in vivo, in vitro, ex vivo, or a combination thereof.
[0388] In other embodiments, the heterologous moiety increases the stability of the therapeutic protein of the present disclosure or a fragment thereof (e.g., a fragment comprising a heterologous moiety after proteolytic cleavage of a therapeutic protein (e.g., a clotting factor)). As used herein, the term "stability" refers to the well-recognized measure in the art of maintaining one or more physical properties of a therapeutic protein in response to environmental conditions (e.g., elevated or reduced temperature). In certain aspects, the physical property can be the maintenance of the covalent structure of the therapeutic protein (e.g., absence of proteolytic cleavage, unwanted oxidation or deamidation). In other aspects, the physical property can also be the presence of the therapeutic protein in the correctly folded state (e.g., absence of soluble or insoluble aggregates or precipitates). In one aspect, the stability of a therapeutic protein is measured by determining the biophysical properties of the therapeutic protein, such as thermal stability, pH unfolding curve, stable removal of glycosylation, solubility, biochemical function (e.g., the ability to bind to a protein, receptor or ligand), etc., and / or combinations thereof. In another aspect, the biochemical function is demonstrated by the binding affinity of the interaction. In one aspect, a measure of protein stability is thermal stability, i.e., resistance to heat challenge. Methods known in the art (such as HPLC (high performance liquid chromatography), SEC (size exclusion chromatography), DLS (dynamic light scattering), etc.) can be used to measure stability. Methods for measuring thermal stability include, but are not limited to, differential scanning calorimetry (DSC), differential scanning fluorimetry (DSF), circular dichroism (CD), and heat challenge assays.
[0389] In certain aspects, the therapeutic protein encoded by the nucleic acid molecule of the present disclosure comprises at least one half-life extender, i.e., a heterologous moiety that increases the in vivo half-life of the therapeutic protein relative to the in vivo half-life of the corresponding therapeutic protein lacking such a heterologous moiety. The in vivo half-life of a therapeutic protein can be determined by any method known to those skilled in the art, such as activity assays (e.g., chromogenic assays or one-stage clotting aPTT assays, where the therapeutic protein comprises an FVIII polypeptide), ELISA, etc.
[0390] In some embodiments, the presence of one or more half-life extenders results in an increase in the half-life of the therapeutic protein compared to the half-life of the corresponding protein lacking such one or more half-life extenders. The half-life of a therapeutic protein comprising a half-life extender is at least about 1.5-fold, at least about 2-fold, at least about 2.5-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 11-fold, or at least about 12-fold longer than the in vivo half-life of the corresponding therapeutic protein lacking such a half-life extender.
[0391] In one embodiment, the half-life of a therapeutic protein comprising a half-life extender is about 1.5 to about 20 times, about 1.5 to about 15 times, or about 1.5 to about 10 times as long as the in vivo half-life of the corresponding protein lacking such a half-life extender. In another embodiment, the half-life of a therapeutic protein comprising a half-life extender is extended by about 2 to about 10 times, about 2 to about 9 times, about 2 to about 8 times, about 2 to about 7 times, about 2 to about 6 times, about 2 to about 5 times, about 2 to about 4 times, about 2 to about 3 times, about 2.5 to about 10 times, about 2.5 to about 9 times, about 2.5 to about 8 times, about 2.5 to about 7 times, about 2.5 to about 6 times, about 2.5 to about 5 times, about 2.5 to about 4 times, about 2.5 to about 3 times, about 3 to about 10 times, about 3 to about 9 times, about 3 to about 8 times, about 3 to about 7 times, about 3 to about 6 times, about 3 to about 5 times, about 3 to about 4 times, about 4 to about 6 times, about 5 to about 7 times, or about 6 to about 8 times compared to the in vivo half-life of the corresponding protein lacking such a half-life extender.
[0392] In other embodiments, the half-life of a therapeutic protein comprising a half-life extender is at least about 17 hours, at least about 18 hours, at least about 19 hours, at least about 20 hours, at least about 21 hours, at least about 22 hours, at least about 23 hours, at least about 24 hours, at least about 25 hours, at least about 26 hours, at least about 27 hours, at least about 28 hours, at least about 29 hours, at least about 30 hours, at least about 31 hours, at least about 32 hours, at least about 33 hours, at least about 34 hours, at least about 35 hours, at least about 36 hours, at least about 48 hours, at least about 60 hours, at least about 72 hours, at least about 84 hours, at least about 96 hours, or at least about 108 hours.
[0393] In still other embodiments, the half-life of a therapeutic protein comprising a half-life extender is from about 15 hours to about two weeks, from about 16 hours to about one week, from about 17 hours to about one week, from about 18 hours to about one week, from about 19 hours to about one week, from about 20 hours to about one week, from about 21 hours to about one week, from about 22 hours to about one week, from about 23 hours to about one week, from about 24 hours to about one week, from about 36 hours to about one week, from about 48 hours to about one week, from about 60 hours to about one week, from about 24 hours to about 6 days, from about 24 hours to about five days, from about 24 hours to about four days, from about 24 hours to about three days, or from about 24 hours to about two days.
[0394] In some embodiments, the average half-life of the therapeutic protein comprising the half-life extender for each subject is about 15 hours, about 16 hours, about 17 hours, about 18 hours, about 19 hours, about 20 hours, about 21 hours, about 22 hours, about 23 hours, about 24 hours (1 day), about 25 hours, about 26 hours, about 27 hours, about 28 hours, about 29 hours, about 30 hours, about 31 hours, about 32 hours, about 33 hours, about 34 hours, about 35 hours, about 36 hours, about 40 hours, about 44 hours, about 48 hours (2 days), about 54 hours, about 60 hours, about 72 hours (3 days), about 84 hours, about 96 hours (4 days), about 108 hours, about 120 hours (5 days), about six days, about seven days (one week), about eight days, about nine days, about 10 days, about 11 days, about 12 days, about 13 days, or about 14 days.
[0395] One or more half-life extenders may be fused to the C-terminus or N-terminus of the therapeutic protein, or inserted within the therapeutic protein.
[0396] 1. An immunoglobulin constant region or a portion thereof
[0397] In another aspect, the heterologous moiety comprises one or more immunoglobulin constant regions or portions thereof (e.g., the Fc region). In one embodiment, the isolated nucleic acid molecule of the present disclosure further comprises a heterologous nucleic acid sequence encoding an immunoglobulin constant region or a portion thereof. In some embodiments, the immunoglobulin constant region or a portion thereof is the Fc region.
[0398] The immunoglobulin constant region is composed of domains designated as CH (constant heavy) domains (CH1, CH2, etc.). Depending on the isotype (i.e., IgG, IgM, IgA, IgD, or IgE), the constant region may be composed of three or four CH domains. Some isotypes (e.g., IgG) constant regions also contain a hinge region. See Janeway et al 2001, Immunobiology, Garland Publishing, N.Y., N.Y.
[0399] The immunoglobulin constant regions or portions thereof of the present disclosure can be obtained from many different sources. In one embodiment, the immunoglobulin constant region or portion thereof is derived from a human immunoglobulin. However, it should be understood that the immunoglobulin constant region or a portion thereof can be derived from the immunoglobulin of another mammalian species, said another mammalian species including, for example, rodents (e.g., mouse, rat, rabbit, guinea pig) or non-human primate (e.g., chimpanzee, macaque) species. In addition, the immunoglobulin constant region or portion thereof can be derived from any immunoglobulin class, including IgM, IgG, IgD, IgA, and IgE, and any immunoglobulin isotype, including IgG1, IgG2, IgG3, and IgG4. In one embodiment, the human isotype IgG1 is used.
[0400] Multiple immunoglobulin constant region gene sequences (e.g., human constant region gene sequences) can be obtained in a publicly available stored form. Constant region domain sequences with specific effector functions (or lacking specific effector functions) or with specific modifications that reduce immunogenicity can be selected. The sequences of many antibodies and antibody-encoding genes have been published, and the sequences of suitable Ig constant regions (e.g., hinge, CH2, and / or CH3 sequences or portions thereof) can be derived from these sequences using techniques well recognized in the art. The genetic material obtained using any of the foregoing methods can then be altered or synthesized to obtain the polypeptides of the present disclosure. It will be further recognized that the scope of the present disclosure encompasses alleles, variants, and mutations of the constant region DNA sequences.
[0401] The sequence of an immunoglobulin constant region or a portion thereof can be cloned, for example, using polymerase chain reaction and primers that select to amplify the domain of interest. To clone the sequence of an immunoglobulin constant region or a portion thereof from an antibody, mRNA can be isolated from a hybridoma, spleen, or lymphocytes, reverse transcribed into DNA, and then the antibody gene amplified by PCR. PCR amplification methods are described in U.S. Patent Nos. 4,683,195; 4,683,202; 4,800,159; 4,965,188; and, for example, "PCR Protocols: A Guide to Methods and Applications" edited by Innis et al., Academic Press, San Diego, CA (1990); Ho et al. 1989. Gene 77:51; Horton et al. 1993. Methods Enzymol. 217:270). PCR can be initiated by consensus constant region primers or by more specific primers based on published heavy and light chain DNA and amino acid sequences. PCR can also be used to isolate DNA clones encoding antibody light and heavy chains. In this case, a library can be screened with consensus primers or larger homologous probes (such as, murine constant region probes). Many primer sets suitable for amplifying antibody genes are known in the art (e.g., 5' primers based on the N-terminal sequence of a purified antibody (Benhar and Pastan. 1994. Protein Engineering 7:1509); rapid amplification of cDNA ends (Ruberti, F. et al. 1994. J. Immunol. Methods 173:33); antibody leader sequences (Larrick et al. 1989 Biochem. Biophys. Res. Commun. 160:1250). Cloning of antibody sequences is further described in Newman et al., U.S. Patent No. 5,658,570, filed January 25, 1995, which is incorporated herein by reference.
[0402] The immunoglobulin constant region used herein can include all domains and the hinge region or a portion thereof. In one embodiment, the immunoglobulin constant region or a portion thereof comprises the CH2 domain, the CH3 domain, and the hinge region, i.e., the Fc region or the FcRn binding partner.
[0403] As used herein, the term "Fc region" is defined as the polypeptide portion corresponding to the Fc region of native Ig, i.e., the portion formed by the dimeric association of the respective Fc domains of its two heavy chains. Native Fc regions form homodimers with another Fc region. In contrast, the term "gene-fused Fc region" or "single-chain Fc region" (scFc region) refers to a synthetic dimeric Fc region consisting of Fc domains that are gene-linked within a single polypeptide chain (i.e., encoded by a single contiguous gene sequence). See International Publication No. WO 2012 / 006635, which is incorporated herein by reference in its entirety.
[0404] In one embodiment, "Fc region" refers to the portion of a single Ig heavy chain that begins at the hinge region immediately upstream of the papain cleavage site (i.e., residue 216 in IgG, the first residue of the heavy chain constant region being 114) and ends at the C-terminus of the antibody. Thus, a complete Fc region includes at least the hinge domain, CH2 domain, and CH3 domain.
[0405] The immunoglobulin constant region or a portion thereof can be an FcRn binding partner. FcRn is active in adult epithelial tissues and is expressed in the intestinal lumen, lung airways, nasal surface, vaginal surface, and the surfaces of the colon and rectum (U.S. Patent No. 6,485,726). An FcRn binding partner is a portion of an immunoglobulin that binds FcRn.
[0406] FcRn receptors have been isolated from several mammalian species, including humans. The sequences of human FcRn, monkey FcRn, rat FcRn, and mouse FcRn are known (Story et al. 1994, J. Exp. Med. 180:2377). The FcRn receptor binds IgG at relatively low pH (but not other immunoglobulin classes such as IgA, IgM, IgD, and IgE), actively transcellularly transports IgG from the lumen to the serosal side, and then releases IgG at the relatively high pH found in tissue fluid. It is expressed in adult epithelial tissues (U.S. Patent Nos. 6,485,726, 6,030,613, 6,086,875; WO 03 / 077834; US2003-0235536A1), including lung and intestinal epithelium (Israel et al. 1997, Immunology 92:69), renal proximal tubular epithelium (Kobayashi et al. 2002, Am. J. Physiol. Renal Physiol. 282:F358), and nasal epithelium, vaginal surface, and the surface of the biliary tree.
[0407] FcRn-binding ligands useful in the present disclosure encompass molecules that can be specifically bound by the FcRn receptor, which includes intact IgG, the Fc fragment of IgG, and other fragments that include the intact binding region of the FcRn receptor. The region of the Fc portion of IgG that binds the FcRn receptor has been described based on X-ray crystallography (Burmeister et al 1994, Nature 372:379). The major contact surface of Fc with FcRn is near the junction of the CH2 and CH3 domains. The Fc-FcRn contacts are all within a single Ig heavy chain. FcRn-binding ligands include intact IgG, the Fc fragment of IgG, and other fragments of IgG that include the intact binding region of FcRn. The major contact sites include amino acid residues 248, 250-257, 272, 285, 288, 290-291, 308-311, and 314 in the CH2 domain, and amino acid residues 385-387, 428, and 433-436 in the CH3 domain. References to the amino acid numbering of immunoglobulins or immunoglobulin fragments or regions are all based on Kabat et al 1991, Sequences of Proteins of Immunological Interest, U.S. Department of Public Health, Bethesda, Md.
[0408] The Fc region or FcRn-binding ligand bound to FcRn can be effectively transported across the epithelial barrier by FcRn, thus providing a non-invasive way to systemically administer a desired therapeutic molecule. In addition, fusion proteins containing an Fc region or FcRn-binding ligand are endocytosed by cells expressing FcRn. However, these fusion proteins are not labeled for degradation but are recycled back into the circulation, thus increasing the in vivo half-life of these proteins. In certain embodiments, a portion of the immunoglobulin constant region is an Fc region or FcRn-binding ligand that typically associates via disulfide bonds or other non-specific interactions with another Fc region or another FcRn-binding ligand to form a dimer or higher-order multimer.
[0409] Two FcRn receptors can bind a single Fc molecule. Crystallographic data indicate that each FcRn molecule binds a single polypeptide of an Fc homodimer. In one embodiment, linking an FcRn-binding ligand (e.g., the Fc fragment of IgG) to a bioactive molecule provides a way to deliver the bioactive molecule orally, buccally, sublingually, rectally, vaginally, as an aerosol for nasal or pulmonary administration, or via the ocular route. In another embodiment, a coagulation factor protein can be administered invasively, e.g., subcutaneously, intravenously.
[0410] The FcRn-binding moiety is a molecule or a portion thereof that can be specifically bound by the FcRn receptor and thus actively transported by the FcRn receptor of the Fc region. Specific binding refers to two molecules that form a relatively stable complex under physiological conditions. Specific binding is characterized by high affinity and low to medium avidity, which is distinct from non-specific binding that typically has low affinity and medium to high avidity. Generally, when the affinity constant KA is higher than 10 6 M -1 or higher than 10 8 M -1 , the binding is considered to be specific. If desired, non-specific binding can be reduced by changing the binding conditions with little effect on specific binding. Suitable binding conditions such as the concentration of the molecules, the ionic strength of the solution, the temperature allowed for binding, the time, the concentration of blocking agents (e.g., serum albumin, milk casein), etc., can be optimized by those skilled in the art using conventional techniques.
[0411] In certain embodiments, the therapeutic protein encoded by the nucleic acid molecule of the present disclosure comprises one or more truncated Fc regions that are still sufficient to confer the binding properties of the Fc receptor (FcR) to the Fc region. For example, the portion of the Fc region that binds FcRn (i.e., the FcRn-binding portion) comprises approximately amino acids 282-438 (EU numbering) from IgG1 (the major contact sites are amino acids 248, 250-257, 272, 285, 288, 290-291, 308-311, and 314 in the CH2 domain and amino acid residues 385-387, 428, and 433-436 in the CH3 domain). Thus, the Fc region of the present disclosure can comprise or consist of the FcRn-binding portion. The FcRn-binding portion can be derived from the heavy chain of any isotype (including IgG1, IgG2, IgG3, and IgG4). In one embodiment, the FcRn-binding portion of an antibody from human isotype IgG1 is used. In another embodiment, the FcRn-binding portion of an antibody from human isotype IgG4 is used.
[0412] The Fc region can be obtained from many different sources. In one embodiment, the Fc region of the polypeptide is derived from human immunoglobulin. However, it should be understood that the Fc portion can be derived from the immunoglobulin of another mammalian species, including for example rodents (e.g., mouse, rat, rabbit, guinea pig) or non-human primate (e.g., chimpanzee, macaque) species. Moreover, the polypeptide of the Fc domain or a portion thereof can be derived from any immunoglobulin class (including IgM, IgG, IgD, IgA, and IgE) as well as immunoglobulin isotypes (including IgG1, IgG2, IgG3, and IgG4). In another embodiment, human isotype IgG1 is used.
[0413] In certain embodiments, the Fc variant confers an alteration of at least one effector function conferred by the Fc portion comprising the wild-type Fc domain (e.g., an ability to enhance or reduce binding of the Fc region to an Fc receptor (e.g., FcγRI, FcγRII, or FcγRIII) or a complement protein (e.g., C1q), or to trigger antibody-dependent cell cytotoxicity (ADCC), phagocytosis, or complement-dependent cytotoxicity (CDCC)). In other embodiments, the Fc variant provides engineered cysteine residues.
[0414] The Fc region of the present disclosure may utilize Fc variants that are well-known in the art and are known to cause alterations (e.g., enhancements or reductions) in effector function and / or FcR or FcRn binding. Specifically, the Fc region of the present disclosure may include, for example, alterations (e.g., substitutions) at one or more of the following disclosed amino acid positions: International PCT publications WO88 / 07089A1, WO96 / 14339A1, WO98 / 05787A1, WO98 / 23289A1, WO99 / 51642A1, WO99 / 58572A1, WO00 / 09560A2, WO00 / 32767A1, WO00 / 42072A2, WO02 / 44215A2, WO02 / 060919A2, WO03 / 074569A2, WO04 / 016750A2, WO04 / 029207A2, WO04 / 035752A2, WO04 / 063351A2, WO04 / 074455A2, WO04 / 099249A2, WO05 / 040217A2, WO04 / 044859, WO05 / 070963A1, WO05 / 077981A2, WO05 / 092925A2, WO05 / 123780A2, WO06 / 019447A1, WO06 / 047350A2, and WO06 / 085967A2; U.S. Patent Publication Nos. US2007 / 0231329, US2007 / 0231329, US2007 / 0237765, US2007 / 0237766, US2007 / 0237767, US2007 / 0243188, US20070248603, US20070286859, US20080057056; or U.S. Patents 5,648,260; 5,739,277; 5,834,250; 5,869,046; 6,096,871; 6,121,022; 6,194,551; 6,242,195; 6,277,375; 6,528,624; 6,538,124; 6,737,056; 6,821,505; 6,998,253; 7,083,784; 7,404,956, and 7,317,091, each of which is incorporated herein by reference. In one embodiment, specific alterations (e.g., specific substitutions of one or more amino acids disclosed in the art) may be made at one or more of the disclosed amino acid positions. In another embodiment, different alterations (e.g., different substitutions at one or more of the disclosed amino acid positions) may be made at one or more of the disclosed amino acid positions.
[0415] The Fc region of IgG or the FcRn binding partner can be modified according to well-known procedures such as site-directed mutagenesis to produce a modified IgG or Fc fragment or a portion thereof that will be bound by FcRn. Such modifications include modifications away from the FcRn contact site and modifications within the contact site that maintain or even enhance binding to FcRn. For example, the following single amino acid residues in human IgG1 Fc (Fcγ1) can be substituted without significantly reducing the binding affinity of Fc for FcRn: P238A, S239A, K246A, K248A, D249A, M252A, T256A, E258A, T260A, D265A, S267A, H268A, E269A, D270A, E272A, L274A, N276A, Y278A, D280A, V282A, E283A, H285A, N286A, T289A, K290A, R292A, E293A, E294A, Q295A, Y296F, N297A, S298A, Y300F, R301A, V303A, V305A, T307A, L309A, Q311A, D312A, N315A, K317A, E318A, K320A, K322A, S324A, K326A, A327Q, P329A, A330Q, P331A, E333A, K334A, T335A, S337A, K338A, K340A, Q342A, R344A, E345A, Q347A, R355A, E356A, M358A, T359A, K360A, N361A, Q362A, Y373A, S375A, D376A, A378Q, E380A, E382A, S383A, N384A, Q386A, E388A, N389A, N390A, Y391F, K392A, L398A, S400A, D401A, D413A, K414A, R416A, Q418A, Q419A, N421A, V422A, S424A, E430A, N434A, T437A, Q438A, K439A, S440A, S444A and K447A, where, for example, P238A represents the wild-type proline substituted by alanine at position number 238. For example, a particular embodiment incorporates the N297A mutation, which removes a highly conserved N-glycosylation site. In addition to alanine, other amino acids can substitute for the wild-type amino acid at the positions specified above. The mutations can be introduced into Fc one by one, resulting in more than a hundred Fc regions that are different from the native Fc. In addition, combinations of two, three or more of these single mutations can be introduced together, resulting in hundreds more Fc regions.
[0416] Some of the above mutations can confer new functions on the Fc region or FcRn binding partners. For example, one embodiment incorporates N297A, which removes a highly conserved N-glycosylation site. The effect of this mutation is to reduce immunogenicity, thereby enhancing the circulatory half-life of the Fc region and rendering the Fc region unable to bind FcγRI, FcγRIIA, FcγRIIB, and FcγRIIIA, without compromising the affinity for FcRn (Routledge et al. 1995, Transplantation 60:847; Friend et al. 1999, Transplantation 68:1632; Shields et al. 1995, J. Biol. Chem. 276:6591). As another example of a new function resulting from the above mutations, in some cases, the affinity for FcRn can be increased to exceed that of the wild type. This increased affinity can reflect an increased "opening" rate, a decreased "closing" rate, or both an increased "opening" rate and a decreased "closing" rate. Examples of mutations that are thought to confer increased affinity for FcRn include, but are not limited to, T256A, T307A, E380A, and dN434A (Shields et al. 2001, J. Biol. Chem. 276:6591).
[0417] In addition, at least three human Fcγ receptors appear to recognize binding sites on IgG in the lower hinge region, typically amino acids 234-237. Thus, another example of a new function and potentially reduced immunogenicity can be generated by mutations in this region, for example, by replacing amino acids 233-236 (SEQ ID NO:45) of human IgG1 "ELLG" with the corresponding sequence of IgG2 "PVA" (with one amino acid deletion). It has been shown that when such mutations have been introduced, FcγRI, FcγRII, and FcγRIII, which mediate multiple effector functions, will not bind IgG1. Ward and Ghetie 1995, Therapeutic Immunology 2:77 and Armour et al. 1999, Eur. J. Immunol. 29:2613.
[0418] In another embodiment, the immunoglobulin constant region or a portion thereof comprises an amino acid sequence in the hinge region or a portion thereof that forms one or more disulfide bonds with a second immunoglobulin constant region or a portion thereof. The second immunoglobulin constant region or a portion thereof may be linked to a second polypeptide, which binds the therapeutic protein and the second polypeptide together. In some embodiments, the second polypeptide is an enhancer moiety. As used herein, the term "enhancer moiety" refers to a molecule, fragment thereof, or polypeptide composition that is capable of enhancing the activity of a therapeutic protein. The enhancer moiety may be a cofactor, such as where the therapeutic protein is a coagulation factor, soluble tissue factor (sTF), or procoagulant peptide. Thus, the enhancer moiety may be used to enhance the activity of the coagulation factor after activation.
[0419] In certain embodiments, the therapeutic protein encoded by the nucleic acid molecules of the present disclosure comprises an amino acid substitution in an immunoglobulin constant region or a portion thereof (e.g., an Fc variant) that alters the antigen-independent effector function of the Ig constant region, particularly the circulating half-life of the protein.
[0420] 2.scFc region
[0421] In another aspect, the heterologous moiety comprises an scFC (single-chain Fc) region. In one embodiment, the isolated nucleic acid molecule of the present disclosure further comprises a heterologous nucleic acid sequence encoding an scFc region. The scFc region comprises at least two immunoglobulin constant regions or portions thereof (e.g., Fc portions or domains (e.g., 2, 3, 4, 5, 6, or more Fc portions or domains)) within the same linear polypeptide chain that are capable of folding (e.g., intramolecularly or intermolecularly) to form a functional scFc region that is linked by an Fc peptide linker. For example, in one embodiment, the polypeptides of the present disclosure are capable of binding to at least one Fc receptor (e.g., FcRn, FcγR receptor (e.g., FcγRIII), or complement protein (e.g., C1q)) via their scFc in order to improve or trigger immune effector functions (e.g., antibody-dependent cell cytotoxicity (ADCC), phagocytosis, or complement-dependent cytotoxicity (CDCC)) and / or improve manufacturability.
[0422] 3.CTP
[0423] In another aspect, the heterologous moiety comprises a C-terminal peptide (CTP) of the beta subunit of human chorionic gonadotropin, or a fragment, variant, or derivative thereof. Inhibiting the insertion of one or more CTP peptides into a recombinant protein increases the in vivo half-life of this protein. See, e.g., U.S. Patent No. 5,712,122, which is incorporated herein by reference in its entirety.
[0424] Exemplary CTP peptides include DPRFQDSSSSKAPPPSLPSPSRLPGPSDTPIL (SEQ ID NO:33) or SSSSKAPPPSLPSPSRLPGPSDTPILPQ (SEQ ID NO:34). See, e.g., U.S. Application Publication No. US 2009 / 0087411 A1, which is incorporated by reference.
[0425] 4. XTEN Sequences
[0426] In some embodiments, the heterologous moiety comprises one or more XTEN sequences, fragments, variants, or derivatives thereof. As used herein, an "XTEN sequence" refers to an extended-length polypeptide having a non-naturally occurring, substantially non-repetitive sequence (which consists primarily of small hydrophilic amino acids) that has a low or no secondary or tertiary structure under physiological conditions. As a heterologous moiety, XTEN can be used as a half-life extension moiety. In addition, XTEN can provide desirable properties, including but not limited to enhanced pharmacokinetic parameters and solubility characteristics.
[0427] Incorporating a heterologous moiety comprising an XTEN sequence into the proteins of the present disclosure can confer one or more of the following advantageous properties to the protein: conformational flexibility, enhanced water solubility, high protease resistance, low immunogenicity, low binding to mammalian receptors, or increased hydrodynamic (or Stokes) radius.
[0428] In certain aspects, the XTEN sequence can increase pharmacokinetic properties such as a longer in vivo half-life or increased area under the curve (AUC), such that the proteins of the present invention remain in the body and have procoagulant activity for a longer period of time compared to the same protein without the XTEN heterologous moiety.
[0429] In some embodiments, the XTEN sequences useful in the present disclosure are peptides or polypeptides having greater than about 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1200, 1400, 1600, 1800, or 2000 amino acid residues. In certain embodiments, XTEN is a peptide or polypeptide having greater than about 20 to about 3000 amino acid residues, greater than 30 to about 2500 residues, greater than 40 to about 2000 residues, greater than 50 to about 1500 residues, greater than 60 to about 1000 residues, greater than 70 to about 900 residues, greater than 80 to about 800 residues, greater than 90 to about 700 residues, greater than 100 to about 600 residues, greater than 110 to about 500 residues, or greater than 120 to about 400 residues. In one particular embodiment, XTEN comprises an amino acid sequence that is longer than 42 amino acids and shorter than 144 amino acids.
[0430] The XTEN sequences of the present disclosure may comprise one or more sequence motifs that are 5 to 14 (e.g., 9 to 14) amino acid residues or an amino acid sequence that is at least 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence motif, wherein the motif comprises, consists essentially of, or consists of 4 to 6 types of amino acids (e.g., 5 amino acids) selected from glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), and proline (P). See US 2010-0239554 A1.
[0431] In some embodiments, XTEN comprises non-overlapping sequence motifs, wherein about 80% or at least about 85% or at least about 90% or about 91% or about 92% or about 93% or about 94% or about 95% or about 96% or about 97% or about 98% or about 99% or about 100% of the sequence consists of non-overlapping sequences of multiple units selected from a single motif family selected from Table 7, thereby generating a family sequence. As used herein, "family" means that XTEN has motifs selected only from a single motif classification from Table 7; i.e., AD, AE, AF, AG, AM, AQ, BC or BD XTEN, and means that any other amino acids in XTEN that are not from the family motifs are selected to achieve desired properties, such as allowing coding nucleotides to incorporate restriction sites, incorporating cleavage sequences or achieving better linkage to therapeutic proteins. In some embodiments of the XTEN family, the XTEN sequence comprises non-overlapping sequence motifs of multiple units of the AD motif family, or the AE motif family, or the AF motif family, or the AG motif family, or the AM motif family, or the AQ motif family, or the BC family, or the BD family, and the resulting XTEN exhibits the homology ranges described above. In other embodiments, XTEN comprises multiple units of motif sequences from two or more motif families in Table 7. These sequences can be selected to achieve desired physical / chemical properties, including properties such as net charge, hydrophilicity, lack of secondary structure or lack of repeatability conferred by the amino acid composition of the motif, which will be described more fully below. In the embodiments described above in this paragraph, the motifs incorporated into XTEN can be selected and assembled using the methods described herein to obtain an XTEN of about 36 to about 3000 amino acid residues.
[0432] Table 7. XTEN sequence motifs and motif families of 12 amino acids
[0433]
[0434]
[0435] * denotes individual motif sequences which, when used together in various arrangements, give rise to a "family sequence"
[0436] Examples of XTEN sequences that can be used as heterologous moieties in the therapeutic proteins of the present disclosure are disclosed, for example, in U.S. Patent Publication Nos. 2010 / 0239554 A1, 2010 / 0323956 A1, 2011 / 0046060 A1, 2011 / 0046061 A1, 2011 / 0077199 A1 or 2011 / 0172146 A1 or Internati...
Claims
1. A baculovirus system comprising a nucleic acid molecule of a first inverted terminal repeat (ITR), a second ITR, and a gene cassette; wherein the first ITR and the second ITR are ITRs of a non-adeno-associated virus (non-AAV), wherein the non-AAV ITR is composed of the sequence SEQ ID NO: 168 or composed of the sequence SEQ ID NO: 173, wherein the gene cassette is located between the first ITR and the second ITR, and wherein the gene cassette encodes a therapeutic protein; wherein the system is capsidless.
2. The baculovirus system according to claim 1, wherein the therapeutic protein comprises a coagulation factor.
3. The baculovirus system according to claim 1 or 2, further comprising a tissue-specific promoter.
4. The baculovirus system according to claim 3, wherein the promoter drives the expression of the therapeutic protein in hepatocytes, endothelial cells, muscle cells, sinusoidal cells, or any combination thereof.
5. The baculovirus system according to claim 3, wherein the promoter is selected from the group consisting of: mouse thyroxine (mTTR) promoter, endogenous human factor VIII (F8) promoter, human albumin minimal promoter, mouse albumin promoter, Tristetraprolin (TTP) promoter, CASI promoter, CAG promoter, cytomegalovirus (CMV) promoter, α1-antitrypsin (AAT) promoter, muscle creatine kinase (MCK) promoter, myosin heavy chain α (αMHC) promoter, myoglobin (MB) promoter, desmin (DES) promoter, SPc5-12 promoter, 2R5Sc5-12 promoter, and phosphoglycerate kinase (PGK) promoter.
6. The baculovirus system according to claim 3, wherein the promoter is selected from the group consisting of: human α-1-antitrypsin (hAAT) promoter, double muscle creatine kinase (dMCK) promoter, and triple muscle creatine kinase (tMCK) promoter.
7. The baculovirus system according to claim 1 or 2, wherein the nucleic acid molecule further comprises: (a) an intron sequence, (b) a post-transcriptional regulatory element, (c) a 3'UTR poly(A) tail sequence, (d) an enhancer sequence, or (e) any combination of (a)-(d).
8. The baculovirus system according to claim 7, wherein: (a) the intron sequence is located at the 5' end of the nucleic acid molecule encoding the therapeutic protein; (b) the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, or any combination thereof; (c) the 3'UTR poly(A) tail sequence is selected from the group consisting of: bGH poly(A), actin poly(A), hemoglobin poly(A), and any combination thereof; or (d) any combination of (a)-(c).
9. The baculovirus system according to claim 8, wherein: (a) The intron sequence comprises SEQ ID NO: 115; (b) The microRNA binding site comprises a binding site for miR142-3p; (c) The 3'UTR poly(A) tail sequence comprises bGH poly(A); or (d) Any combination of (a)-(c).
10. The baculovirus system according to claim 1 or 2, wherein the nucleic acid molecule comprises, in this order: (a) The first ITR; (b) A tissue-specific promoter sequence, which comprises the TTP promoter; (c) An intron, which is a synthetic intron; (d) A nucleotide sequence encoding a coagulation factor; (e) A post-transcriptional regulatory element, which comprises WPRE; (f) A 3'UTR poly(A) tail sequence, which comprises bGHpA; and (g) The second ITR.
11. The baculovirus system according to claim 2, wherein the coagulation factor comprises factor I (FI), factor II (FII), factor V (FV), factor VII (FVII), factor VIII (FVIII), factor IX (FIX), factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FVIII), von Willebrand factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, α2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), or any combination thereof.
12. The baculovirus system according to claim 11, wherein the FVIII comprises an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identical to the amino acid sequence shown in a sequence selected from SEQ ID NOs: 106 and 109.
13. The baculovirus system according to claim 2, wherein the coagulation factor comprises a heterologous moiety selected from the group consisting of albumin or a fragment thereof, the Fc region of an immunoglobulin, the C-terminal peptide (CTP) of the β subunit of human chorionic gonadotropin, a PAS sequence, a HAP sequence, transferrin or a fragment thereof, an albumin-binding moiety, or any combination thereof.
14. The baculovirus system according to claim 11, wherein the FVIII further comprises an FcRn-binding ligand.
15. The baculovirus system according to claim 1 or 2, wherein the nucleic acid molecule is formulated with a delivery agent comprising a lipid nanoparticle.
16. A pharmaceutical composition comprising the baculovirus system according to any one of claims 1-15 and a pharmaceutically acceptable carrier.
17. Use of the baculovirus system according to any one of claims 1-15 or the pharmaceutical composition according to claim 16 for the preparation of a drug for expressing a blood coagulation factor.
Citation Information
Patent Citations
Lentiviral vectors encoding clotting factors for gene therapy
EP1395293A1
Serum albumin binding moieties
US20030069395A1
Central airway administration for systemic delivery of therapeutics
US20030235536A1
Fc Variants Having Increased Affinity for FcyRIIb
US20070231329A1
Fc Variants Having Increased Affinity for FcyRl
US20070237765A1