Nucleic acid molecules and uses thereof
A nucleic acid molecule with non-AAV ITRs and a gene cassette addresses AAV's limitations, enhancing therapeutic protein expression and stability, and reducing immune responses, thus improving treatment efficacy.
Patent Information
- Application Number
- JP2025085515
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-08-09
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-26
AI Technical Summary
Existing AAV gene delivery vectors face limitations such as limited viral packaging capacity and immune responses, leading to inefficiencies in therapeutic protein expression and ineligibility for patients with pre-existing anti-AAV immunity.
A nucleic acid molecule comprising non-AAV inverted terminal repeats (ITRs) and a gene cassette encoding a therapeutic protein, such as FVIII, with tissue-specific promoters and regulatory elements, to enhance expression and stability.
The solution achieves sustained and efficient expression of therapeutic proteins, overcoming AAV's limitations by increasing packaging capacity and reducing immune responses, thereby improving therapeutic efficacy.
Smart Images

Figure 2025124729000029 
Figure 2025124729000030 
Figure 2025124729000031
Abstract
Description
[Technical Field]
[0001] Reference to an electronically submitted sequence listing The contents of the electronically submitted Sequence Listing as an ASCII text file (Name: 4159_493PC01_ST25; Size: 434,561 bytes; Created: August 9, 2018) are hereby incorporated by reference in their entirety. [Background technology]
[0002] Gene therapy offers a sustainable means of treating a variety of diseases. In the past, gene therapy has generally relied on the use of viruses. There are numerous viral agents that can be selected for this purpose, each with distinct properties that may make them more or less suitable for gene therapy (Non-Patent Document 1). However, the undesirable properties of some viral vectors, including their immunogenic profile or their tendency to cause cancer, have raised clinical safety concerns and, until recently, limited their current use in the clinic to certain applications, such as vaccines and oncolytic strategies (Non-Patent Document 2).
[0003] Adeno-associated virus (AAV) is one of the most commonly investigated gene therapy vectors. AAV is a protein shell that surrounds and protects a small, single-stranded DNA genome of approximately 4.8 kilobases (kb) (Non-Patent Document 3). AAV belongs to the Parvoviridae family and depends on coinfection with other viruses, primarily adenoviruses, for replication. Id. Its single-stranded genome contains three genes: Rep (replication), Cap (capsid), and aap (assembly). Id. These coding sequences are flanked by inverted terminal repeats (ITRs) required for genome replication and packaging. Id. The two cis-acting AAV ITRs are approximately 145 nucleotides in length and contain interrupted palindromic sequences that can fold into T-shaped hairpin structures that function as primers during the initiation of DNA replication.
[0004] However, the use of conventional AAV as a gene delivery vector has brought about some drawbacks. One of the major drawbacks is related to AAV's limited viral packaging capacity of approximately 4.5 kb of heterologous DNA (Non-Patent Document 4). Furthermore, administration of AAV vectors can induce immune responses in humans. Although AAV has been shown to be less immunogenic than some other viruses (i.e., adenovirus), its capsid protein can elicit various components of the human immune system (see Non-Patent Document 3). AAV is a common virus in the human population, and most people have been exposed to AAV. Therefore, most people have already mounted an immune response against the specific variants to which they were previously exposed. This pre-existing adaptive response may include NAb and T cells, which may attenuate the clinical efficacy of subsequent reinfection with AAV, and / or the elimination of transduced cells, which makes patients with pre-existing anti-AAV immunity ineligible for AAV-based gene therapy treatment. Whether administered locally or systemically, the virus is seen as a foreign protein, and therefore the adaptive immune system attempts to eliminate it. Furthermore, anti-AAV neutralizing antibodies induced by AAV therapy interfere with repeated AAV therapy when therapeutic levels of efficacy are not achieved by the first AAV therapy. Furthermore, evidence suggests that the T-shaped hairpin loop of the AAV ITR is susceptible to inhibition by host cell proteins / protein complexes that bind to the T-shaped hairpin structure of the AAV ITR. See, for example, Non-Patent Document 5. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Zhou et al., Adv Drug Deliv Rev.106(Pt A):3-26, 2016 [Non-patent document 2] Cotter et al., Front Biosci. 10:1098-105 (2005) [Non-patent document 3] Naso et al., BioDrugs, 31(4):317-334, 2017 [Non-patent document 4] Dong et al., Hum Gene Ther. 7(17):2101-12, 1996. [Non-Patent Document 5] Zhou et al., Scientific Reports 7:5432 (July 14, 2017) Summary of the Invention [Problem to be solved by the invention]
[0006] Thus, there is a need in the art for efficient and sustained expression of target sequences, e.g., therapeutic proteins and / or miRNAs, in in vitro and in vivo settings while avoiding some of the unintended consequences and limitations of existing AAV vector technology. [Means for solving the problem]
[0007] One aspect of the present disclosure relates to a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette encoding a therapeutic protein; the first ITR and / or the second ITR are ITRs of a non-adeno-associated virus (non-AAV), the gene cassette is disposed between the first ITR and the second ITR, and the therapeutic protein comprises a coagulation factor. In some embodiments, the non-AAV is selected from the group consisting of members of the Parvoviridae family of viruses and any combination thereof. In some embodiments, the first ITR is a non-AAV ITR and the second ITR is an adeno-associated virus (AAV) ITR, or the first ITR is an AAV ITR and the second ITR is a non-AAV ITR. In other embodiments, the first ITR and the second ITR are non-AAV ITRs. In some embodiments, the first ITR and the second ITR are identical. In some embodiments, the first and / or second ITRs comprise an uninterrupted palindromic sequence, while in other embodiments, the first and / or second ITRs comprise an interrupted palindromic sequence.
[0008] In some embodiments, the non-AAV in the nucleic acid molecule is a member of the Parvoviridae family of viruses. In some embodiments, the member of the Parvoviridae family of viruses is selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidesovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof. In some embodiments, the member of the Parvoviridae family of viruses is erythrovirus parvovirus B19 (a human virus). In some embodiments, the member of the Parvoviridae family of viruses is Muscovy duck parvovirus. In some embodiments, the Dependoparvovirus is a Dependovirus goose parvovirus (GPV) strain. In some embodiments, the MDPV strain is attenuated FZ91-30. In some embodiments, the MDPV strain is pathogenic YY. Further, in some embodiments, the Dependoparvovirus is a Dependovirus goose parvovirus (GPV) strain. In some embodiments, the GPV strain is attenuated 82-0321V. In some embodiments, the GPV strain is pathogenic B. In some embodiments, the member of the Parvoviridae family of viruses is selected from the group consisting of porcine parvovirus (U44978), minute virus of mice (U34256), canine parvovirus (M19296), mink enteritis virus (D00765), and any combination thereof.
[0009] In some embodiments, the nucleic acid molecule further comprises a promoter. In some embodiments, the promoter is a tissue-specific promoter. In some embodiments, the promoter promotes expression of the therapeutic protein in hepatocytes, endothelial cells, muscle cells, sinusoidal cells, or any combination thereof. In some embodiments, the promoter is located 5' to the nucleic acid sequence encoding the coagulation factor. In some embodiments, the promoter is selected from the group consisting of mouse thyretin promoter (mTTR), endogenous human factor VIII promoter (F8), human alpha-1-antitrypsin promoter (hAAT), human albumin minimal promoter, mouse albumin promoter, tristetraprolin (TTP) promoter, CASI promoter, CAG promoter, cytomegalovirus (CMV) promoter, alpha 1-antitrypsin (AAT), muscle creatine kinase (MCK), myosin heavy chain alpha (αMHC), myoglobin (MB), desmin (DES), SPc5-12, 2R5Sc5-12, dMCK, tMCK, phosphoglycerate kinase (PGK) promoter, and any combination thereof. In some embodiments, the promoter comprises a TTP promoter.
[0010] In some embodiments, the nucleic acid molecule further comprises an intron sequence. In some embodiments, the intron sequence is located 5' to the nucleic acid sequence encoding the clotting factor. In some embodiments, the intron sequence is located 3' to the promoter. In some embodiments, the intron sequence comprises a synthetic intron sequence. In some embodiments, the intron sequence comprises SEQ ID NO: 115.
[0011] In some embodiments, the nucleic acid molecule comprises a post-transcriptional regulatory element. In some embodiments, the post-transcriptional regulatory element is located 3' to the nucleic acid sequence encoding the coagulation factor. In some embodiments, the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, or any combination thereof. In some embodiments, the microRNA binding site comprises a binding site for miR142-3p.
[0012] In some embodiments, the nucleic acid molecule comprises a 3'UTR poly(A) tail sequence. In some embodiments, the 3'UTR poly(A) tail sequence is selected from the group consisting of bGH poly(A), actin poly(A), hemoglobin poly(A), and any combination thereof. In some embodiments, the 3'UTR poly(A) tail sequence comprises bGH poly(A).
[0013] In some embodiments, the nucleic acid molecule comprises an enhancer sequence, hi some embodiments, the enhancer sequence is located between the first and second ITRs.
[0014] In some embodiments, the nucleic acid molecule comprises, in that order, a first ITR, a gene cassette, and a second ITR; wherein the gene cassette comprises a tissue-specific promoter sequence, an intron sequence, a nucleic acid sequence encoding a therapeutic protein, e.g., a clotting factor or miRNA, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence.
[0015] In some embodiments, the nucleic acid molecule comprises, in that order, a tissue-specific promoter sequence, an intron sequence, a nucleic acid sequence encoding a FVIII polypeptide, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence.
[0016] In one embodiment, the nucleic acid molecule is: (a) The first ITR, which is an ITR of a non-AAV family member of the Parvoviridae family; (b) tissue-specific promoter sequences, e.g., the TTP promoter; (c) an intron, e.g., a synthetic intron; (d) a nucleotide sequence encoding a therapeutic protein, e.g., a clotting factor or miRNA; (e) posttranscriptional regulatory elements, e.g., WPRE; (f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and (g) a second ITR that is an ITR of a non-AAV family member of the Parvoviridae family; Includes.
[0017] In some embodiments, the nucleic acid molecule comprises a single-stranded nucleic acid. In some embodiments, the gene cassette comprises a double-stranded nucleic acid.
[0018] In some embodiments, the nucleic acid molecule comprises a gene encoding a therapeutic protein, such as a clotting factor, where the clotting factor is expressed by hepatocytes, endothelial cells, myocytes, sinusoidal cells, or any combination thereof. In some embodiments, the coagulation factor comprises factor I (FI), factor II (FII), factor V (FV), factor VII (FVII), factor VIII (FVIII), factor IX (FIX), factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FXIII), von Willebrand factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), or any combination thereof. In some embodiments, the coagulation factor is FVIII. In some embodiments, the FVIII comprises full-length mature FVIII. In some embodiments, the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 106. In some embodiments, the FVIII comprises the A1 domain, the A2 domain, the A3 domain, the C1 domain, the C2 domain, and a partial B domain, or does not comprise the B domain. In some embodiments, the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 109.
[0019] In some embodiments, the coagulation factor comprises a heterologous moiety. In some embodiments, the heterologous moiety is selected from the group consisting of albumin or a fragment thereof, an immunoglobulin Fc region, the C-terminal peptide (CTP) of the β subunit of human chorionic gonadotropin, a PAS sequence, an HAP sequence, transferrin or a fragment thereof, an albumin-binding moiety, a derivative thereof, and any combination thereof. In some embodiments, the heterologous moiety is attached to the N-terminus or C-terminus of FVIII. The heterologous moiety is linked to or inserted between two amino acids of FVIII. In some embodiments, the heterologous moiety is inserted between two amino acids at one or more insertion sites selected from the insertion sites listed in Table 5. In some embodiments, the FVIII further comprises the A1 domain, the A2 domain, the C1 domain, the C2 domain, optionally the B domain, and the heterologous moiety, wherein the heterologous moiety is inserted immediately downstream of amino acid 745, corresponding to mature FVIII (SEQ ID NO: 106).
[0020] In some embodiments, FVIII further comprises an FcRn binding partner. In some embodiments, the FcRn binding partner comprises an Fc region of an immunoglobulin constant domain. In some embodiments, the nucleic acid sequence encoding FVIII is codon-optimized. In some embodiments, the nucleic acid sequence encoding FVIII is codon-optimized for expression in humans.
[0021] In some embodiments, the nucleic acid sequence encoding FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO: 107. In some embodiments, the nucleic acid sequence encoding FVIII comprises a nucleotide sequence that is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the nucleotide sequence of SEQ ID NO: 71.
[0022] In some embodiments, the nucleic acid molecule is formulated with a delivery agent. In some embodiments, the delivery agent comprises one or more lipid nanoparticles. In some embodiments, the delivery agent is selected from the group consisting of liposomes, non-lipid polymer molecules, endosomes, and any combination thereof.
[0023] In some embodiments, the nucleic acid molecule is formulated for intravenous, transdermal, intradermal, subcutaneous, pulmonary, or oral delivery, or any combination thereof, hi some embodiments, the nucleic acid molecule is formulated for intravenous delivery.
[0024] In some embodiments, provided herein are vectors comprising the nucleic acid molecules described throughout this disclosure. In some embodiments, provided herein are polypeptides encoded by the nucleic acid molecules described throughout this disclosure. In some embodiments, provided herein are host cells comprising the nucleic acid molecules described throughout this disclosure.
[0025] In some embodiments, provided herein is a pharmaceutical composition comprising: (a) a nucleic acid described herein, a vector comprising a nucleic acid molecule described throughout this disclosure, a polypeptide encoded by a nucleic acid molecule described throughout this disclosure, or a host cell comprising a nucleic acid molecule described throughout this disclosure; (b) an LNP; and (c) a pharmaceutically acceptable excipient. In some embodiments, provided herein is a kit comprising a nucleic acid molecule described throughout this disclosure and instructions for administering the nucleic acid molecule to a subject in need thereof. In some embodiments, provided herein is a baculovirus system for producing the nucleic acid molecules described herein. In some embodiments, provided herein is a baculovirus system in which the nucleic acid molecules disclosed herein are produced in insect cells.
[0026] In some embodiments, provided herein is a nanoparticle delivery system for an expression construct, wherein the expression construct comprises a nucleic acid molecule described herein.
[0027] Also disclosed herein is a method for producing a polypeptide having coagulation activity, the method comprising culturing a host cell disclosed herein under suitable conditions and recovering the polypeptide having coagulation activity. In some embodiments, disclosed herein is a method for expressing a coagulation factor in a subject in need thereof, the method comprising administering to the subject a nucleic acid molecule disclosed herein, a vector disclosed herein, a polypeptide disclosed herein, or a pharmaceutical composition disclosed herein. In some embodiments, disclosed herein is a method for treating a subject having a coagulation factor deficiency, the method comprising administering to the subject a nucleic acid molecule disclosed herein, a vector disclosed herein, a polypeptide disclosed herein, or a pharmaceutical composition disclosed herein. In some embodiments, the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonary, or any combination thereof. In some embodiments, the nucleic acid molecule is administered intravenously. In some embodiments, the method further comprises administering a second agent to the subject. In some embodiments, the subject is a mammal. In some embodiments, the subject is human.
[0028] In some embodiments, administration of a nucleic acid molecule to a subject results in increased FVIII activity compared to the FVIII activity in the subject prior to administration, wherein the FVIII activity is increased by at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 11-fold, at least about 12-fold, at least about 13-fold, at least about 14-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, or at least about 100-fold.
[0029] In some embodiments, the subject has a bleeding disorder. In some embodiments, the bleeding disorder is hemophilia. In some embodiments, the bleeding disorder is hemophilia A. [Brief explanation of the drawings]
[0030] [Figure 1-1]Figure 1A is a schematic diagram of a single-chain coagulation factor (e.g., FVIII) expression cassette. The locations of the non-AAV-derived 5' ITR (with a hairpin loop at the end of the ssDNA structure), the non-AAV-derived 3' ITR (with a hairpin loop), the promoter sequence (e.g., TTPp), and the transgene sequence, e.g., the FVIIIco6XTEN sequence with XTEN144 inserted within the B domain, are shown. The exemplary expression cassette also shows additional possible elements, such as intron sequences, WPREmut sequences, and bGHpA sequences. Figures 1B-1D are schematic diagrams of plasmids used to prepare single-chain coagulation factor expression cassettes, such as the cassette shown in Figure 1A, where the ITRs of the cassette are derived from AAV2 (Figure 1B), B19 (Figure 1C), or GPV (Figure 1D). As shown herein, the plasmid construct containing the ssFVIII expression cassette was digested with PvuII (at the PvuII site) (FIG. 1B) or LguI (at the LguI site) (FIGS. 1C and 1D) to release the viral genome. The double-stranded DNA was heated to 95°C to produce ssDNA, which was then incubated at 4°C to allow the ITR structures to form. [Figure 1-2] Continued from Figure 1-2. [Figure 1-3] Continued from Figure 1-3. [Figure 2A] Figure 2A is a phylogenetic tree showing the relationships among various Parvoviridae members, with B19, AAV-2, and GPV marked with open boxes. [Figure 2B] FIG. 2B is a schematic diagram of various cassettes, including hairpin structures. [Figure 3] Figures 3A and 3B show alignments of the ITRs of B19, GPV, and AAV2 (Figure 3A) and B19 and GPV (Figure 3B). Gray shading indicates homology. [Figure 4-1]Figures 4A-4C show FVIII plasma activity after administration of single-stranded FVIII-AAV naked DNA (ssAAV-FVIII; Figure 2A), ssDNA-B19 FVIII (Figure 2B), or ssDNA-GPV FVIII (Figure 2C) by hydrodynamic injection (HDI) in Hem A mice. FVIII activity was measured in plasma samples (as a percentage of control) at 24 hours, 3 days, 2 weeks, 3 weeks, 1 month, 2 months, 3 months, and 4 months in mice treated with a single HDI of ssDNA at 50 μg / mouse (Figure 2C), 20 μg / mouse (Figures 2A and 2B), 10 μg / mouse (Figures 2A and 2C), or 5 μg / mouse (Figure 2A). An HDI of 5 μg / mouse of plasmid DNA was given as a control (Figures 2A-2C). [Figure 4-2] Continued from Figure 4-1. DETAILED DESCRIPTION OF THE INVENTION
[0031] The present disclosure describes a plasmid-like nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette encoding, for example, a therapeutic protein or miRNA, wherein the first ITR and / or the second ITR are non-adeno-associated viral ITRs (e.g., the first ITR and / or the second ITR are derived from a non-AAV). In some embodiments, the gene cassette encodes a therapeutic protein. In some embodiments, the therapeutic protein comprises a protein selected from a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or a combination thereof. In some embodiments, the gene cassette encodes X-linked dystrophin, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof.
[0032] In some embodiments, the therapeutic protein comprises a clotting factor, hi one particular embodiment, the therapeutic protein comprises a FVIII or FIX protein.
[0033] In some embodiments, the gene cassette encodes an miRNA. In certain embodiments, the miRNA downregulates the expression of a target gene selected from SOD1, HTT, RHO, or any combination thereof.
[0034] In certain embodiments, the non-AAV is selected from the group consisting of members of the Parvoviridae family of viruses and any combination thereof. The present disclosure further relates to a method for expressing a therapeutic protein, such as a coagulation factor, such as FVIII, in a subject in need thereof, comprising administering to the subject a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette encoding, for example, a therapeutic protein or miRNA, wherein the first ITR and / or the second ITR are ITRs of a non-adeno-associated virus (non-AAV). In certain embodiments, the present disclosure describes an isolated nucleic acid molecule comprising a nucleotide sequence having sequence homology to a nucleotide sequence selected from SEQ ID NOs: 113 and 120.
[0035] Exemplary constructs of the present disclosure are illustrated in the accompanying figures and sequence listing. In order to provide a clear understanding of the specification and claims, the following definitions are provided below.
[0036] I. Definition It is noted that the terms "a" or "an" entity refer to one or more of that entity: for example, a "nucleotide sequence" refers to one or more nucleotides. "A" (or "an"), "one or more," and "at least one" can be used interchangeably herein.
[0037] The term "about" is used herein to mean approximately, roughly, around, or approximately. When the term "about" is used in conjunction with a numerical range, it modifies that range by extending the boundaries above and below the numerical values set forth. In general, the term "about" is used herein to modify numerical values above and below the stated value by a variance of 10% above and below (high and low).
[0038] Similarly, as used herein, "and / or" refers to and includes all possible combinations of one or more of the associated listed items, as well as the lack of a combination ("or") when interpreted as alternatives.
[0039] The terms "nucleic acid," "nucleic acid molecule," "nucleotide," "nucleotide sequence(s)," and "polynucleotide" are used interchangeably and refer to the phosphate polymeric form of ribonucleosides (adenosine, guanosine, uridine, or cytidine; "RNA molecule") or deoxyribonucleosides (deoxyadenosine, deoxyguanosine, deoxythymidine, or deoxycytidine; "DNA molecule"), or any of their phosphate analogs, such as phosphorothioates and thioesters, in single-stranded or double-stranded helical form. A single-stranded nucleic acid sequence refers to single-stranded DNA (ssDNA) or single-stranded RNA (ssRNA). Double-stranded DNA-DNA, DNA-RNA, and RNA-RNA helices are possible. The term nucleic acid molecule, particularly DNA or RNA molecule, refers only to the primary and secondary structure of the molecule and does not limit it to any particular tertiary form. Thus, the term includes double-stranded DNA found, inter alia, in linear or circular DNA molecules (e.g., restriction fragments), plasmids, supercoiled DNA, and chromosomes. In discussing the structure of a particular double-stranded DNA molecule, the sequence may be described herein by the usual convention of providing only the sequence in the 5' to 3' direction along the non-transcribed strand of DNA (i.e., the strand having sequence homology to mRNA). A "recombinant DNA molecule" is a DNA molecule that has undergone molecular biological manipulation. DNA includes, but is not limited to, cDNA, genomic DNA, plasmid DNA, synthetic DNA, and semi-synthetic DNA. A "nucleic acid composition" of the present disclosure comprises one or more nucleic acids described herein.
[0040] As used herein, "inverted terminal repeat" (or "ITR") refers to a nucleic acid subsequence located at either the 5' or 3' end of a single-stranded nucleic acid sequence that includes a set of nucleotides (an initial sequence) followed downstream by its reverse complement, i.e., a palindromic sequence. The intervening sequence of nucleotides between the initial sequence and the reverse complement can be of any length, including zero. In one embodiment, an ITR useful for the present disclosure includes one or more "palindromic sequences." An ITR can have any number of functions. In some embodiments, the ITRs described herein form a hairpin structure. In some embodiments, the ITRs form a T-shaped hairpin structure. In some embodiments, the ITRs form a non-T-shaped hairpin structure, e.g., a U-shaped hairpin structure. In some embodiments, the ITRs promote long-term survival of a nucleic acid molecule in the nucleus of a cell. In some embodiments, the ITRs promote permanent survival of a nucleic acid molecule in the nucleus of a cell (e.g., for the entire lifespan of the cell). In some embodiments, the ITRs promote stability of a nucleic acid molecule in the nucleus of a cell. In some embodiments, the ITRs promote the retention of nucleic acid molecules in the nucleus of a cell. In some embodiments, the ITRs promote the persistence of nucleic acid molecules in the nucleus of a cell. In some embodiments, the ITRs inhibit or prevent the degradation of nucleic acid molecules in the nucleus of a cell.
[0041] In one embodiment, the initial sequence and / or reverse complement comprises from about 2 to 600 nucleotides, from about 2 to 550 nucleotides, from about 2 to 500 nucleotides, from about 2 to 450 nucleotides, from about 2 to 400 nucleotides, from about 2 to 350 nucleotides, from about 2 to 300 nucleotides, or from about 2 to 250 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 5 to 600 nucleotides, about 10 to 600 nucleotides, about 15 to 600 nucleotides, about 20 to 600 nucleotides, about 25 to 600 nucleotides, about 30 to 600 nucleotides, about 35 to 600 nucleotides, about 40 to 600 nucleotides, about 45 to 600 nucleotides, about 50 to 600 nucleotides, about 60 to 600 nucleotides, about 70 to 600 nucleotides, about 80 to 600 nucleotides, about 90 to 600 nucleotides, about 100 to 600 nucleotides, about 150 to 600 nucleotides, about 200 to 600 nucleotides, about 300 to 600 nucleotides, about 350 to 600 nucleotides, about 400 to 600 nucleotides, about 450 to 600 nucleotides, about 500 to 600 nucleotides, or about 550 to 600 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 5 to 550 nucleotides, about 5 to 500 nucleotides, about 5 to 450 nucleotides, about 5 to 400 nucleotides, about 5 to 350 nucleotides, about 5 to 300 nucleotides, or about 5 to 250 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 10 to 550 nucleotides, about 15 to 500 nucleotides, about 20 to 450 nucleotides, about 25 to 400 nucleotides, about 30 to 350 nucleotides, about 35 to 300 nucleotides, or about 40 to 250 nucleotides. In certain embodiments, the initial sequence and / or reverse complement comprises about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, or about 600 nucleotides. In certain embodiments, the initial sequence and / or reverse complement comprises about 400 nucleotides.
[0042] In other embodiments, the initial sequence and / or reverse complement comprises about 2 to 200 nucleotides, about 5 to 200 nucleotides, about 10 to 200 nucleotides, about 20 to 200 nucleotides, about 30 to 200 nucleotides, about 40 to 200 nucleotides, about 50 to 200 nucleotides, about 60 to 200 nucleotides, about 70 to 200 nucleotides, about 80 to 200 nucleotides, about 90 to 200 nucleotides, about 100 to 200 nucleotides, about 125 to 200 nucleotides, about 150 to 200 nucleotides, or about 175 to 200 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2 to 150 nucleotides, about 5 to 150 nucleotides, about 10 to 150 nucleotides, about 20 to 150 nucleotides, about 30 to 150 nucleotides, about 40 to 150 nucleotides, about 50 to 150 nucleotides, about 75 to 150 nucleotides, about 100 to 150 nucleotides, or about 125 to 150 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2 to 100 nucleotides, about 5 to 100 nucleotides, about 10 to 100 nucleotides, about 20 to 100 nucleotides, about 30 to 100 nucleotides, about 40 to 100 nucleotides, about 50 to 100 nucleotides, or about 75 to 100 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2-50 nucleotides, about 10-50 nucleotides, about 20-50 nucleotides, about 30-50 nucleotides, about 40-50 nucleotides, about 3-30 nucleotides, about 4-20 nucleotides, or about 5-10 nucleotides. In another embodiment, the initial sequence and / or reverse complement consists of 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides. In another embodiment, the intervening nucleotides between the initial sequence and the reverse complement are 0 nucleotides. , 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides.
[0043] Thus, as used herein, an "ITR" can fold back on itself to form a double-stranded segment. For example, the sequence GATCXXXXGATC, when folded to form a double helix, contains the initial sequence of GATC and its complement (3'CTAG5'). In some embodiments, an ITR contains a continuous palindromic sequence (e.g., GATCGATC) between the initial sequence and its reverse complement. In some embodiments, an ITR contains an interrupted palindromic sequence (e.g., GATCXXXXGATC) between the initial sequence and its reverse complement. In some embodiments, the complementary portions of the continuous or interrupted palindromic sequence interact with each other to form a "hairpin loop" structure. As used herein, a "hairpin loop" structure results when at least two complementary sequences on a single-stranded nucleotide molecule base pair to form a double-stranded portion. In some embodiments, only a portion of the ITR forms a hairpin loop. In other embodiments, the entire ITR forms a hairpin loop.
[0044] In the present disclosure, at least one ITR is an ITR of a non-adenovirus-associated virus (non-AAV). In certain embodiments, the ITR is an ITR of a non-AAV member of the Parvoviridae family of viruses. In some embodiments, the ITR is an ITR of a non-AAV member of the Dependovirus or Erythrovirus genus. In certain embodiments, the ITR is an ITR of goose parvovirus (GPV), Muscovy duck parvovirus (MDPV), or erythrovirus parvovirus B19 (also known as parvovirus B19, primate erythroparvovirus 1, B19 virus, and erythrovirus). In certain embodiments, one ITR of the two ITRs is an AAV ITR. In other embodiments, one of the two ITRs in the construct is an ITR of an AAV serotype selected from serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and any combination thereof. In a particular embodiment, the ITR is derived from AAV serotype 2, e.g., from an ITR of AAV serotype 2.
[0045] In certain aspects of the present disclosure, the nucleic acid molecule comprises two ITRs, 5'ITR and 3'ITR, wherein the 5'ITR is located at the 5' end of the nucleic acid molecule and the 3'ITR is located at the 3' end of the nucleic acid molecule. The 5'ITR and the 3'ITR can be derived from the same virus or different viruses. In certain embodiments, the 5'ITR is derived from AAV, and the 3'ITR is not derived from an AAV virus (e.g., non-AAV). In some embodiments, the 3'ITR is derived from AAV, and the 5'ITR is not derived from an AAV virus (e.g., non-AAV). In other embodiments, the 5'ITR is not derived from an AAV virus (e.g., non-AAV), and the 3'ITR is derived from the same or different non-AAV virus.
[0046] As used herein, the term "parvovirus" encompasses the family Parvoviridae, which includes, but is not limited to, autonomous parvoviruses and dependoviruses. Autonomous parvoviruses include, for example, Bocaviruses, Dependoviruses, Erythroviruses, Amdoviruses, Parvoviruses, Densoviruses, Iteraviruses, Contraviruses, and Avepparvoviruses. Parvoviruses include members of the genera Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstyldensovirus.
[0047] Exemplary autonomous parvoviruses include, but are not limited to, porcine parvovirus, minute virus of mice, canine parvovirus, mink enteritis virus, bovine parvovirus, chicken parvovirus, feline panleukopenia virus, feline parvovirus, goose parvovirus, H1 parvovirus, Muscovy duck parvovirus, snake parvovirus, and B19 virus. Other autonomous parvoviruses are known to those skilled in the art. See, for example, FIELDS et al., VIROLOGY, Vol. 2, Chapter 69 (4th Edition, Lippincott-Raven Publishers).
[0048] As used herein, the term "non-AAV" encompasses nucleic acids, proteins and viruses from the Parvoviridae family, excluding any adeno-associated viruses (AAV) of the Parvoviridae family. "Non-AAV" includes, but is not limited to, autonomously replicating members of the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstyldensovirus.
[0049] As used herein, the term "adeno-associated virus" (AAV) includes, but is not limited to, AAV types 1, 2, 3 (including 3A and 3B), 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, snake AAV, avian AAV, bovine AAV, canine AAV, equine AAV, ovine AAV, caprine AAV, shrimp AAV, the AAV serotypes and clades disclosed by Gao et al. (J. Virol. 78:6381 (2004)) and Morris et al. (Virol. 33:375 (2004)), and any other AAV now known or later discovered. See, for example, FIELDS et al., VIROLOGY, Vol. 2, Chapter 69 (4th ed., Lippincott-Raven Publishers).
[0050] As used herein, the term "derived from" refers to a component that is isolated from a specified molecule or organism or that is made using information (e.g., amino acid or nucleic acid sequence) from a specified molecule or organism. For example, a nucleic acid sequence (e.g., ITR) derived from a second nucleic acid sequence (e.g., ITR) can contain a nucleotide sequence that is identical or substantially similar to the nucleotide sequence of the second nucleic acid sequence. In the case of nucleotides or polypeptides, the derived species can be obtained by, for example, naturally occurring mutagenesis, artificially directed mutagenesis, or artificially random mutagenesis. The mutagenesis used to derive a nucleotide or polypeptide can be deliberately directed, deliberately random, or each Mutagenesis of a nucleotide or polypeptide that creates a different nucleotide or polypeptide from the first can be a random event (e.g., caused by polymerase infidelity), and identification of the derived nucleotide or polypeptide can be performed by a suitable screening method, e.g., as discussed herein. Mutagenesis of a polypeptide generally requires manipulation of a polynucleotide that encodes the polypeptide. In some embodiments, a nucleotide or amino acid sequence derived from a second nucleotide or amino acid sequence is at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, or at least 76% identical to the second nucleotide or amino acid sequence, respectively. %, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In other embodiments, the ITRs derived from a non-AAV (or AAV) ITRs are at least 90% identical to the non-AAV ITRs (or AAV ITRs, respectively), wherein the non-AAV (or AAV) ITRs are non-AAV ITRs that retain a functional property of the ITR (or, respectively, AAV ITR). In some embodiments, an ITR derived from a non-AAV (or AAV) ITR is at least 80% identical to the non-AAV ITR (or, respectively, AAV ITR), wherein the non-AAV (or, AAV) ITR retains a functional property of the non-AAV ITR (or, respectively, AAV ITR). In some embodiments, an ITR derived from a non-AAV (or, AAV) ITR is at least 70% identical to the non-AAV ITR (or, respectively, AAV ITR), wherein the non-AAV (or, AAV) ITR retains a functional property of the non-AAV ITR (or, respectively, AAV ITR). In some embodiments, an ITR derived from a non-AAV (or, AAV) ITR is at least 60% identical to the non-AAV ITR (or, respectively, AAV ITR), wherein the non-AAV (or, AAV) ITR retains a functional property of the non-AAV ITR (or, respectively, AAV ITR). In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 50% identical to the non-AAV ITRs (or, respectively, AAV ITRs), wherein the non-AAV (or, respectively, AAV) ITRs retain functional properties of the non-AAV ITRs (or, respectively, AAV ITRs).
[0051] In certain embodiments, the ITRs derived from non-AAV (or AAV) ITRs comprise or consist of a fragment of the non-AAV (or AAV) ITR. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs comprise or consist of a fragment of the non-AAV (or AAV) ITR, wherein the fragment is at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides, or at least about 600 nucleotides; wherein the ITRs derived from non-AAV (or AAV) ITRs retain the functional properties of the non-AAV ITRs (or AAV ITRs, respectively). In certain embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of the non-AAV (or AAV) ITR, wherein the fragment comprises at least about 129 nucleotides, and wherein the ITR derived from the non-AAV (or AAV) ITR retains a functional property of the non-AAV ITR (or, respectively, the AAV ITR). In certain embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of the non-AAV (or AAV) ITR, wherein the fragment comprises at least about 102 nucleotides, and wherein the ITR derived from the non-AAV (or AAV) ITR retains a functional property of the non-AAV ITR (or, respectively, the AAV ITR).
[0052] In some embodiments, the ITR derived from a non-AAV (or AAV) ITR comprises or consists of a fragment of the non-AAV (or AAV) ITR, wherein the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the length of the non-AAV (or AAV) ITR.
[0053] In certain embodiments, a nucleotide or amino acid sequence derived from a second nucleotide or amino acid sequence, when properly aligned, has an identity with a homologous portion of the second nucleotide or amino acid sequence of at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123 %, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In other embodiments, the ITRs derived from a non-AAV (or AAV) ITRs, when properly aligned, have at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. ITR) in which the first nucleotide or The ITRs derived from non-AAV (or AAV) ITRs are at least 80% identical to the homologous portion of the non-AAV ITR (or, respectively, AAV ITR) when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 70% identical to the homologous portion of the non-AAV ITR (or, respectively, AAV ITR) when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 60% identical to the homologous portion of the non-AAV ITR (or, respectively, AAV ITR) when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence. In some embodiments, the ITRs derived from non-AAV (or AAV) ITRs are at least 50% identical to the homologous portion of the non-AAV ITR (or AAV ITR, respectively) when properly aligned, wherein the first nucleotide or amino acid sequence retains the biological activity of the second nucleotide or amino acid sequence.
[0054] A "capsid-free" or "capsid-less" vector or nucleic acid molecule refers to a vector construct that does not contain a capsid. In some embodiments, the capsid-less vector or nucleic acid molecule does not contain sequences encoding, for example, AAV Rep proteins.
[0055] As used herein, a "coding region" or "coding sequence" is a portion of a polynucleotide consisting of codons translatable into amino acids. Although a "stop codon" (TAG, TGA, or TAA) is not generally translated into an amino acid, it can be considered part of the coding region, but any adjacent sequences, such as promoters, ribosome binding sites, transcription terminators, introns, etc., are not part of the coding region either. The boundaries of a coding region are generally determined by a start codon at the 5' end, which encodes the amino terminus of the resulting polypeptide, and a translation stop codon at the 3' end, which encodes the carboxyl terminus of the resulting polypeptide. Two or more coding regions can be present in a single polynucleotide construct, e.g., on a single vector, or in separate polynucleotide constructs, e.g., on separate (different) vectors. Consequently, a single vector can contain only a single coding region, or can include two or more coding regions.
[0056] Certain proteins secreted by mammalian cells are associated with a secretory signal peptide that is cleaved from the mature protein upon initiation of translocation of the growing protein chain across the rough endoplasmic reticulum. Those skilled in the art will recognize that signal peptides are commonly fused to the N-terminus of a polypeptide and are cleaved from the complete or "full-length" polypeptide to generate the secreted or "mature" form of the polypeptide. In certain embodiments, the native signal peptide or a functional derivative of that sequence retains the ability to direct the secretion of a polypeptide operably associated therewith. Alternatively, a heterologous mammalian signal peptide, such as human tissue plasminogen activator (TPA) or mouse β-glucuronidase signal peptide, or a functional derivative thereof, can be used.
[0057] The term "downstream" refers to the nucleotide sequence located 3' of the reference nucleotide sequence. In certain embodiments, the downstream nucleotide sequence relates to the sequence following the start of transcription. For example, the translation start codon of a gene is located downstream of the start site of transcription.
[0058] The term "upstream" refers to a nucleotide sequence located 5' of a reference nucleotide sequence. In certain embodiments, an upstream nucleotide sequence is located 5' of a coding region or the start of transcription. It refers to sequences located 5' to the promoter. For example, most promoters are located upstream from the start site of transcription.
[0059] As used herein, the term "gene regulatory region" or "regulatory region" refers to a nucleotide sequence located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding region that influences the transcription, RNA processing, stability, or translation of the associated coding region. Regulatory regions can include promoters, translation leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites, or stem-loop structures. If the coding region is intended for expression in a eukaryotic cell, a polyadenylation signal and transcription termination sequence will usually be located 3' to the coding sequence.
[0060] A polynucleotide encoding a product, e.g., an miRNA or gene product (e.g., a polypeptide such as a therapeutic protein), can include a promoter and / or other expression (e.g., transcriptional or translational) control elements operably associated with one or more coding regions. In operably associated, a coding region for a gene product, e.g., a polypeptide, is associated with one or more regulatory regions in such a way that expression of the gene product is under the influence or control of the regulatory region(s). For example, a coding region and a promoter are "operably associated" if induction of promoter function results in transcription of an mRNA encoding the gene product encoded by the coding region, and if the nature of the linkage between the promoter and the coding region does not interfere with the promoter's ability to direct expression of the gene product or the ability of the DNA template to be transcribed. Other expression control elements besides promoters, e.g., enhancers, operators, repressors, and transcription termination signals, can also be operably associated with a coding region to direct gene product expression.
[0061] "Transcription control sequence" refers to a DNA regulatory sequence, such as a promoter, enhancer, terminator, etc., that enables expression of a coding sequence in a host cell. Various transcription control regions are known to those skilled in the art. These include, but are not limited to, transcription control regions that function in vertebrate cells, such as promoter and enhancer segments from cytomegalovirus (immediate-early promoter with intron-A), simian virus 40 (early promoter), and retroviruses (such as Rous sarcoma virus). Other transcription control regions include those derived from vertebrate genes, such as actin, heat shock protein, bovine growth hormone, and rabbit β-globin, as well as other sequences capable of controlling gene expression in eukaryotic cells. Additional suitable transcription control regions include tissue-specific promoters and enhancers, and lymphokine-inducible promoters (e.g., promoters inducible by interferon or interleukin).
[0062] Similarly, a variety of translation control elements are known to those skilled in the art, including, but not limited to, ribosome binding sites, translation initiation and termination codons, and elements derived from picornaviruses (particularly internal ribosome entry sites or IRES, also called CITE sequences).
[0063] As used herein, the term "expression" refers to the process by which a polynucleotide produces a gene product, e.g., an RNA or a polypeptide. It includes, but is not limited to, transcription of a polynucleotide into messenger RNA (mRNA), transfer RNA (tRNA), small hairpin RNA (shRNA), small interfering RNA (siRNA), or any other RNA product, and translation of mRNA into a polypeptide. Expression produces a "gene product." As used herein, a gene product may be a nucleic acid, e.g., a messenger RNA produced by transcription of a gene, or a polypeptide translated from a transcript. Gene products as described herein further include nucleic acids that have post-transcriptional modifications, such as polyadenylation or splicing, or polypeptides that have post-translational modifications, such as methylation, glycosylation, lipid addition, conjugation to other protein subunits, or proteolytic cleavage. As used herein, the term "yield" refers to the amount of polypeptide produced by expression of a gene.
[0064] A "vector" refers to any vehicle for the cloning and / or transfer of a nucleic acid into a host cell. A vector may be a replicon to which another nucleic acid segment may be attached so as to bring about the replication of the attached segment. A "replicon" is an in It refers to any genetic element (e.g., a plasmid, a phage, a cosmid, a chromosome, a virus) that functions as an autonomous unit of replication in vivo, i.e., capable of replication under its own control. The term "vector" includes a vehicle for introducing nucleic acids into cells in vitro, ex vivo, or in vivo. Numerous vectors are known and used in the art, including, for example, plasmids, modified eukaryotic viruses, or modified bacterial viruses. Insertion of a polynucleotide into a suitable vector can be accomplished by ligating an appropriate polynucleotide fragment into a selected vector with complementary cohesive termini.
[0065] Vectors can be engineered to encode selectable markers or reporters that allow for the selection or identification of cells that have incorporated the vector. Expression of the selectable marker or reporter allows for the identification and / or selection of host cells that incorporate and express other coding regions contained on the vector. Examples of selectable marker genes known and used in the art include: genes that provide resistance to ampicillin, streptomycin, gentamicin, kanamycin, hygromycin, bialaphos herbicides, sulfonamides, etc.; and genes used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentanyl transferase genes, etc. Examples of reporters known and used in the art include: luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), β-galactosidase (LacZ), β-glucuronidase (Gus), etc. Selectable markers can also be considered reporters.
[0066] The term "host cell" as used herein refers to, for example, microorganisms, yeast cells, insect cells, and mammalian cells that can be or have been used as recipients of ssDNA or vectors. The term includes progeny of the original cell that has been transduced. Thus, as used herein, "host cell" generally refers to a cell that has been transduced with an exogenous DNA sequence. It is understood that the progeny of a single parent cell may not necessarily be completely identical in morphology or in genomic or total DNA complement to the original parent due to natural, accidental, or deliberate mutation. In some embodiments, the host cell may be an in vitro host cell.
[0067] The term "selectable marker" refers to an identifying agent, usually an antibiotic or chemical resistance gene, that can be selected for based on the function of the marker gene, i.e., antibiotic resistance, herbicide resistance, colorimetric marker, enzyme, fluorescent marker, etc., where this function is used to track the inheritance of a nucleic acid of interest and / or to identify cells or organisms that have inherited the nucleic acid of interest. Examples of selectable marker genes known and used in the art include: genes that provide resistance to ampicillin, streptomycin, gentamicin, kanamycin, hygromycin, bialaphos herbicides, sulfonamides, etc.; and genes used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentanyl transferase genes, etc.
[0068] The term "reporter gene" refers to a nucleic acid encoding an identifying factor that can be identified based on the action of the reporter gene, where this action is used to track the inheritance of the nucleic acid of interest, to identify cells or organisms that have inherited the nucleic acid of interest, and / or to measure the induction or transcription of gene expression. Examples of reporter genes known and used in the art include luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), β-galactosidase (LacZ), β-glucuronidase (Gus), etc. Selectable marker genes can also be considered reporter genes.
[0069] The terms "promoter" and "promoter sequence" are used interchangeably and refer to a DNA sequence capable of controlling the expression of a coding sequence or functional RNA. Generally, the coding sequence is located 3' from the promoter sequence. Promoters can be derived entirely from a native gene, composed of different elements derived from different promoters found in nature, or can include synthetic DNA segments. Those skilled in the art will appreciate that different promoters can direct the expression of a gene in different tissues or cell types, or at different stages of development, or in response to different environmental or physiological conditions. A promoter that causes a gene to be expressed in most cell types most of the time is generally referred to as a "constitutive promoter." A promoter that causes a gene to be expressed in a specific cell type is generally referred to as a "cell-specific promoter" or "tissue-specific promoter." A promoter that causes a gene to be expressed at a specific stage of development or cell differentiation is generally referred to as a "development-specific promoter" or "cell differentiation-specific promoter." A promoter that is induced to cause a gene to be expressed following exposure or treatment of cells with a promoter-inducing drug, biomolecule, chemical, ligand, light, etc. is generally referred to as an "inducible promoter" or "regulatable promoter." It is further recognized that since in most cases the exact boundaries of regulatory sequences have not been completely defined, DNA fragments of different lengths may have identical promoter activity.
[0070] A promoter sequence generally contains a transcription initiation site bounded at its 3' end and extends upstream (5') to include the minimum number of bases or elements necessary to initiate transcription at levels detectable above background. Within the promoter sequence will be found a transcription initiation site (conveniently defined, for example, by mapping with nuclease S1), as well as protein binding domains (consensus sequences) responsible for the binding of RNA polymerase.
[0071] In some embodiments, the nucleic acid molecule comprises a tissue-specific promoter. In certain embodiments, the tissue-specific promoter promotes the expression of therapeutic proteins, such as coagulation factors, in the liver, for example, in hepatocytes and / or endothelial cells. In certain embodiments, the promoter is selected from the group consisting of mouse thyretin promoter (mTTR), endogenous human factor VIII promoter (F8), human alpha-1-antitrypsin promoter (hAAT), human albumin minimal promoter, mouse albumin promoter, tristetraprolin (TTP) promoter, CASI promoter, CAG promoter, cytomegalovirus (CMV) promoter, phosphoglycerate kinase (PGK) promoter, and any combination thereof. In some embodiments, the promoter is selected from a liver-specific promoter (e.g., α1-antitrypsin (AAT)), a muscle-specific promoter (e.g., muscle creatine kinase (MCK), myosin heavy chain alpha (αMHC), myoglobin (MB), and desmin (DES)), a synthetic promoter (e.g., SPc5-12, 2R5Sc5-12, dMCK, and tMCK), and any combination thereof. In one particular embodiment, the promoter comprises a TTP promoter.
[0072] The terms "restriction endonucleases" and "restriction enzymes" are used interchangeably and refer to enzymes that bind and cut at specific nucleotide sequences in double-stranded DNA.
[0073] The term "plasmid" refers to an extrachromosomal element that often carries genes that are not part of the central metabolism of a cell and that are usually in the form of circular double-stranded DNA molecules. Such elements can be linear, circular, or supercoiled, autonomously replicating, genomic integrating, phage, or nucleotide sequences of single- or double-stranded DNA or RNA derived from any source, in which several nucleotide sequences are linked or recombined into a specific construct that is capable of introducing into a cell a promoter fragment and DNA sequence for a selected gene product, along with appropriate 3' untranslated sequences.
[0074] Eukaryotic viral vectors that can be used include, but are not limited to, adenovirus vectors, retrovirus vectors, adeno-associated virus vectors, poxviruses such as vaccinia virus vectors, baculovirus vectors, or herpes virus vectors. Non-viral vectors include plasmids, liposomes, electrically charged lipids (cytofectins), DNA-protein complexes, and biopolymers.
[0075] "Cloning vector" refers to a "replicon," a unit-length nucleic acid, such as a plasmid, phage, or cosmid, that replicates sequentially and contains an origin of replication, to which another nucleic acid segment can be attached so as to bring about replication of the attached segment. Certain cloning vectors are capable of replication in one cell type, e.g., bacteria, and expression in another cell type, e.g., eukaryotic cells. Cloning vectors generally contain one or more sequences that can be used for selection of cells that contain the vector, and / or one or more multiple cloning sites for insertion of nucleic acid sequences of interest.
[0076] The term "expression vector" refers to a vehicle designed to enable expression of an inserted nucleic acid sequence after insertion into a host cell, the inserted nucleic acid sequence being placed in operable relationship with the regulatory regions described above.
[0077] Vectors are introduced into host cells by methods well known in the art, such as transfection, electroporation, microinjection, transduction, cell fusion, DEAE-dextran, calcium phosphate precipitation, lipofection (lysosome fusion), using gene guns, or DNA vector transporters.As used herein, "culture", "cultivating" and "culturing" refer to incubating cells under in vitro conditions that allow cell growth or division, or maintaining cells in a viable state.As used herein, "cultured cells" refers to cells that are grown in vitro.
[0078] As used herein, the term "polypeptide" encompasses the singular "polypeptide" as well as the plural "polypeptides," and refers to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also known as peptide bonds). The term "polypeptide" refers to any chain or chains of two or more amino acids, and does not refer to a specific length of the product. Thus, peptides, dipeptides, tripeptides, oligopeptides, "proteins," "amino acid chains," or any other term used to refer to a chain or chains of two or more amino acids are included within the definition of "polypeptide," and the term "polypeptide" can be used in place of or interchangeably with any of these terms. The term "polypeptide" refers to any molecule that is glycosylated, acetylated, phosphorylated, amidated, or otherwise modified, as known in the art. "A polypeptide also refers to the product of post-expression modification of a polypeptide, including, but not limited to, derivatization with protecting / blocking groups, proteolytic cleavage, or modification with non-naturally occurring amino acids. A polypeptide can be derived from a natural biological source or produced by recombinant technology, but is not necessarily translated from a designated nucleic acid sequence. It can be produced by any method, including chemical synthesis.
[0079] The term "amino acid" includes alanine (Ala or A); arginine (Arg or R); asparagine (Asn or N); aspartic acid (Asp or D); cysteine (Cys or C); glutamine (Gln or Q); glutamic acid (Glu or E); glycine (Gly or G); histidine (His or H); isoleucine (Ile or I); leucine (Leu or L); lysine (Lys or K); methionine (Met or M); phenylalanine (Phe or F); proline (Pro or P); serine (Ser or S); threonine (Thr or T); tryptophan (Trp or W); tyrosine (Tyr or Y); and valine (Val or V). Non-traditional amino acids are also within the scope of the present disclosure and include norleucine, ornithine, norvaline, homoserine, and other amino acid residue analogs such as those described in Ellman et al., Meth. Enzym. 202:301-336 (1991). To generate such non-naturally occurring amino acid residues, the techniques of Noren et al., Science 244:182 (1989) and Ellman et al., supra, can be used. Briefly, these techniques involve chemically activating a suppressor tRNA with the non-naturally occurring amino acid residue, followed by in vitro transcription and translation of the RNA. Introduction of non-traditional amino acids can also be achieved using peptide chemistry known in the art. As used herein, the term "polar amino acid" includes amino acids that have a zero total charge but non-zero partial charges in different portions of their side chains (e.g., M, F, W, S, Y, N, Q, C). These amino acids can participate in hydrophobic and electrostatic interactions. As used herein, the term "charged amino acid" includes amino acids that can have a non-zero net charge on their side chains (e.g., R, K, H, E, D). These amino acids can participate in hydrophobic and electrostatic interactions.
[0080] Polypeptide fragments or variants, and any combinations thereof, are also included in the present disclosure. The term "fragment" or "variant," when referring to a polypeptide binding domain or binding molecule of the present disclosure, includes any polypeptide that retains at least some of the properties of the reference polypeptide (e.g., FcRn-binding affinity for an FcRn-binding domain or Fc variant, clotting activity of an FVIII variant, or FVIII-binding activity to a VWF fragment). Polypeptide fragments include proteolytic fragments and deletion fragments, as well as specific antibody fragments discussed elsewhere herein, but do not include naturally occurring full-length polypeptides (or mature polypeptides). Variants of the polypeptide binding domain or binding molecule of the present disclosure also include the above-mentioned fragments, as well as polypeptides with altered amino acid sequences due to amino acid substitutions, deletions, or insertions. Variants may be naturally occurring or non-naturally occurring. Non-naturally occurring variants can be generated using mutagenesis techniques known in the art. Variant polypeptides can include conservative or non-conservative amino acid substitutions, deletions, or additions.
[0081] A "conservative amino acid substitution" is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain, such as a basic side chain (e.g., lysine, arginine, histidine), an acidic side chain (e.g., aspartic acid, glutamic acid), an uncharged polar side chain (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), a polar side chain (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), a beta-branched side chain (e.g., threonine, valine, isoleucine), or a non-polar side chain (e.g., threonine, valine, isoleucine). Families of amino acid residues with similar side chains, including amino acids (e.g., leucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine), have been defined in the art. Thus, when an amino acid in a polypeptide is replaced with another amino acid from the same side chain family, the substitution is considered conservative. In another embodiment, a string of amino acids can be conservatively replaced with a structurally similar string that differs in the order and / or composition of the side chain family members.
[0082] As known in the art, the term "percent identity" refers to a relationship between two or more polypeptide sequences or two or more polynucleotide sequences, as determined by comparing the sequences. In the art, "identity" also means the degree of sequence relatedness between polypeptide or polynucleotide sequences, as the case may be, as determined by the match between strings of such sequences. "Identity" can be easily calculated by known methods, including but not limited to those described in the following: Computational Molecular Biology (Lesk, AM, ed.), Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, DW, ed.), Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, AM and Griffin, HG, eds.), Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.), Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.), Stockton Press, New York (1991). Preferred methods for determining identity are designed to give the best match between the sequences tested. Methods for determining identity are codified in publicly available computer programs.Sequence alignments and percent identity calculations can be performed using sequence analysis software such as the Megalign program of the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, WI), the GCG program suite (Wisconsin Package version 9.0, Genetics Computer Group (GCG), Madison, WI), BLASTP, BLASTN, BLASTX (Altschul et al., J. Mol. Biol. 215:403 (1990)), and DNASTAR (DNASTAR, Inc. 1228 S. Park St. Madison, WI 53715 USA). Within the context of this application, when sequence analysis software is used for analysis, it is understood that the results of the analysis will be based on the "default values" of the referenced program unless otherwise specified. As used herein, "default values" refers to any set of values or parameters that initially load with the software when first initialized. For purposes of determining percent identity between a therapeutic protein, e.g., a clotting factor, sequence of the present disclosure and a reference sequence, only nucleotides of the reference sequence that correspond to nucleotides in the therapeutic protein, e.g., a clotting factor, sequence of the present disclosure are used to calculate percent identity. For example, when comparing a full-length FVIII nucleotide sequence containing the B domain with an optimized B-domain deleted (BDD) FVIII nucleotide sequence of the present disclosure, the portion of the alignment including the A1, A2, A3, C1, and C2 domains is used to calculate percent identity. Nucleotides in the portion of the full-length FVIII sequence encoding the B domain (which results in a large "gap" in the alignment) are not counted as mismatches. Furthermore, in determining percent identity between an optimized BDD FVIII sequence of the present disclosure, or a designated portion thereof (e.g., nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3), and a reference sequence, the percent identity is calculated by dividing the number of matched nucleotides by the total number of nucleotides in the entire sequence of the optimized BDD-FVIII sequence or a designated portion thereof listed herein. will be done.
[0083] As used herein, nucleotides corresponding to nucleotides in a particular sequence of the present disclosure are identified by aligning the sequences of the present disclosure to maximize identity with the reference sequence. The number used to identify an equivalent amino acid in the reference sequence is based on the number used to identify the corresponding amino acid in the sequence of the present disclosure.
[0084] A "fusion" or "chimeric" protein comprises a first amino acid sequence linked to a second amino acid sequence to which it is not originally naturally linked. Amino acid sequences normally present in separate proteins can be brought together in a fusion polypeptide, or amino acid sequences normally present in the same protein can be placed in a new configuration in a fusion polypeptide, e.g., a fusion of the Factor VIII domain with the Ig Fc domain of the present disclosure. Fusion proteins are made, for example, by chemical synthesis, or by creating and translating a polynucleotide in which the peptide regions are encoded in the desired relationship. Chimeric proteins can further comprise a second amino acid sequence associated with the first amino acid sequence by a covalent, non-peptide, or non-covalent bond.
[0085] As used herein, the term "insertion site" refers to a position in a polypeptide or a fragment, variant, or derivative thereof immediately upstream of a position at which a heterologous moiety can be inserted. An "insertion site" is defined by a number, which is the number of an amino acid in a reference sequence. For example, an "insertion site" in FVIII refers to the number of the amino acid sequence in mature, native FVIII (SEQ ID NO: 15) to which the insertion site corresponds, immediately N-terminal to the insertion position. For example, the phrase "a3 contains a heterologous moiety at an insertion site corresponding to amino acid 1656 of SEQ ID NO: 15" indicates that the heterologous moiety is located between the two amino acids corresponding to amino acids 1656 and 1657 of SEQ ID NO: 15.
[0086] As used herein, the phrase "immediately downstream of an amino acid" refers to a position immediately adjacent to the terminal carboxyl group of an amino acid. Similarly, the phrase "immediately upstream of an amino acid" refers to a position immediately adjacent to the terminal amine group of an amino acid.
[0087] As used herein, the terms "inserted," "inserted," "inserted into," or grammatically related terms refer to the position of a heterologous moiety in a polypeptide, e.g., a clotting factor, relative to the analogous position in a parent polypeptide. For example, in certain embodiments, "inserted," or the like, refers to the position of a heterologous moiety in a recombinant FVIII polypeptide relative to the analogous position in native, mature human FVIII. As used herein, these terms refer to characteristics of the polypeptide and do not indicate, imply, or refer to any method or process by which the polypeptide was made.
[0088] As used herein, the term "half-life" refers to the biological half-life of a particular polypeptide in vivo. Half-life can be expressed as the time required for half of an administered dose to be cleared from the animal's circulation and / or other tissues. When the elimination curve of a given polypeptide is constructed as a function of time, the curve is usually biphasic, with a rapid α-phase and a longer β-phase. The α-phase generally represents the equilibration of an administered Fc polypeptide between the intravascular and extravascular spaces and is determined, in part, by the size of the polypeptide. The β-phase generally represents the catabolism of the polypeptide in the intravascular space. In some embodiments, therapeutic proteins, such as coagulation factors, e.g., FVIII, and chimeric proteins containing same, are monophasic, i.e., they have no alpha phase but only a single beta phase. Thus, in certain embodiments, the term half-life as used herein refers to the half-life of a polypeptide in the β-phase.
[0089] As used herein, the term "linked" refers to a first amino acid sequence or nucleotide sequence that is covalently or non-covalently linked to a second amino acid sequence or nucleotide sequence, respectively. The first amino acid or nucleotide sequence can be directly linked or juxtaposed to the second amino acid or nucleotide sequence, or the first sequence can be covalently linked to the second sequence via an intervening sequence. The term "linked" not only refers to the fusion of the first amino acid sequence to the second amino acid sequence at the C-terminus or N-terminus, but also includes the insertion of any two amino acids of the second amino acid sequence (or of the first amino acid sequence, respectively) throughout the first amino acid sequence (or second amino acid sequence). In one embodiment, the first amino acid sequence can be linked to the second amino acid sequence by a peptide bond or a linker. The first nucleotide sequence can be linked to the second nucleotide sequence by a phosphodiester bond or a linker. The linker can be a peptide or polypeptide (in the case of a polypeptide chain), a nucleotide or nucleotide chain (in the case of a nucleotide chain), or any chemical moiety (in the case of a polypeptide and a polynucleotide chain). The term "coupled" can also be indicated by a hyphen (-).
[0090] As used herein, hemostasis means the stopping or slowing of bleeding or hemorrhage; or the stopping or slowing of blood flow through a blood vessel or body part.
[0091] Hemostatic disorder, as used herein, refers to a genetically inherited or acquired condition characterized by a tendency to bleed, either spontaneously or as a result of trauma, due to an impaired or defective ability to form fibrin clots. Examples of such disorders include hemophilia. The three major forms are hemophilia A (factor VIII deficiency), hemophilia B (factor IX deficiency or "Christmas disease"), and hemophilia C (factor XI deficiency, mild bleeding tendency). Other hemostatic disorders include, for example, von Willebrand disease, factor XI deficiency (PTA deficiency), factor XII deficiency, deficiencies or structural abnormalities of fibrinogen, prothrombin, factor V, factor VII, factor X, or factor XIII, and Bernard-Soulier syndrome, which is a defect or deficiency of GPIb. The receptor for vWF, GPIb, is defective, which can lead to loss of primary clot formation (primary hemostasis) and an increased tendency to bleed, as well as Glanzmann and Naegeli thrombasthenia (Glanzmann thrombasthenia).In liver failure (acute and chronic forms), the liver produces insufficient clotting factors, which can increase the risk of bleeding.
[0092] The isolated nucleic acid molecule, isolated polypeptide, or vector comprising the isolated nucleic acid molecule of the present disclosure can be used prophylactically. As used herein, the term "prophylactic treatment" refers to the administration of the molecule before a bleeding episode. In one embodiment, the subject requiring a systemic hemostatic agent is undergoing or about to undergo surgery. The polynucleotide, polypeptide, or vector of the present disclosure can be administered before or after surgery as a prophylactic agent. The polynucleotide, polypeptide, or vector of the present disclosure can be administered during or after surgery to control acute bleeding episodes. Surgery can include, but is not limited to, liver transplantation, liver resection, dental procedures, or stem cell transplantation.
[0093] The isolated nucleic acid molecules, isolated polypeptides, or vectors of the present disclosure can also be used for on-demand therapy. The term "on-demand therapy" refers to the administration of an isolated nucleic acid molecule, isolated polypeptide, or vector in response to symptoms of a bleeding episode or before an activity that may cause bleeding. In one aspect, on-demand therapy can be administered to a subject when bleeding begins, such as after an injury, or when bleeding is expected, such as before surgery. In another aspect, on-demand therapy can be administered before an activity that increases the risk of bleeding, such as contact sports.
[0094] As used herein, the term "acute bleeding" refers to a bleeding episode regardless of the underlying cause. For example, the subject may have trauma, uremia, a hereditary bleeding disorder (e.g., factor VII deficiency), a platelet disorder, or tolerance due to the development of antibodies against clotting factors.
[0095] As used herein, treat, therapy, or treatment refers to, for example, reducing the severity of a disease or condition; reducing the duration of a disease; ameliorating one or more symptoms associated with a disease or condition; or providing a beneficial effect to a subject having a disease or condition without necessarily curing the disease or condition or preventing one or more symptoms associated with the disease or condition. In one embodiment, the term "treating" or "treatment" refers to the administration of an isolated nucleic acid molecule, isolated polypeptide, or vector of the present disclosure to achieve, for example, at least about 1 IU / dL, 2 IU / dL, 3 IU / dL, 4 IU / dL, 5 IU / dL, 6 IU / dL, 7 IU / dL, 8 IU / dL, 9 IU / dL, 10 IU / dL, 11 IU / dL, 12 IU / dL, 13 IU / dL, 14 IU / dL, 15 IU / dL, 16 IU / dL, 17 IU / dL, 18 IU / dL, 19 IU / dL, 20 IU / dL, 25 IU / dL, 26 IU / dL, 27 IU / dL, 28 IU / dL, 29 IU / dL, 30 IU / dL, 31 IU / dL, 32 IU / dL, 33 IU / dL, 34 IU / dL, 35 IU / dL, 36 IU / dL, 37 IU / dL, 38 IU / dL, 39 IU / dL, 40 IU / dL, 41 IU / dL, 42 IU / dL, 43 IU / dL, 44 IU / dL, 45 IU / dL, 46 IU / dL, 47 IU / dL, 48 IU / dL, 49 IU / dL, 50 IU / dL, 51 IU / dL, 52 IU / dL, 53 IU / dL, 54 IU / dL, 55 IU / dL, 56 IU / dL, 57 IU / dL, This means maintaining a FVIII trough level of 100 IU / dL, 30 IU / dL, 35 IU / dL, 40 IU / dL, 45 IU / dL, 50 IU / dL, 55 IU / dL, 60 IU / dL, 65 IU / dL, 70 IU / dL, 75 IU / dL, 80 IU / dL, 85 IU / dL, 90 IU / dL, 95 IU / dL, 100 IU / dL, 105 IU / dL, 110 IU / dL, 115 IU / dL, 120 IU / dL, 125 IU / dL, 130 IU / dL, 135 IU / dL, 140 IU / dL, 145 IU / dL or 150 IU / dL.In another embodiment, treating or treatment includes increasing FVIII trough levels to about 1 to about 150 IU / dL, about 1 to about 125 IU / dL, about 1 to about 100 IU / dL, about 1 to about 90 IU / dL, about 1 to about 85 IU / dL, about 1 to about 80 IU / dL, about 1 to about 75 IU / dL, about 1 to about 70 IU / dL, about 1 to about 65 IU / dL, about 1 to about 60 IU / dL, about 1 to about 55 IU / dL, about 1 to about 50 IU / dL, about 1 to about 45 IU / dL, about 1 to about 40 IU / dL, about 1 to about 35 IU / dL, This means maintaining between about 1 to about 30 IU / dL, about 1 to about 25 IU / dL, about 25 to about 125 IU / dL, about 50 to about 100 IU / dL, about 50 to about 75 IU / dL, about 75 to about 100 IU / dL, about 1 to about 20 IU / dL, about 2 to about 20 IU / dL, about 3 to about 20 IU / dL, about 4 to about 20 IU / dL, about 5 to about 20 IU / dL, about 6 to about 20 IU / dL, about 7 to about 20 IU / dL, about 8 to about 20 IU / dL, about 9 to about 20 IU / dL, or about 10 to about 20 IU / dL. Treating or treating a disease or condition can also include maintaining FVIII activity in a subject at a level equivalent to at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 105%, 110%, 115%, 120%, 125%, 130%, 135%, 140%, 145% or 150% of the FVIII activity in a non-hemophilic subject. The minimum trough level required for treatment can be determined by one or more known methods and can be adjusted (increased or decreased) for each individual.
[0096] As used herein, "administering" means providing a subject with a pharmaceutically acceptable nucleic acid molecule, a polypeptide expressed therefrom, or a vector comprising a nucleic acid molecule of the present disclosure through a pharmaceutically acceptable route. The administration route can be intravenous, for example, intravenous injection and intravenous infusion. Additional administration routes include, for example, subcutaneous, intramuscular, oral, nasal, and pulmonary administration. The nucleic acid molecule, polypeptide, and vector can be administered as part of a pharmaceutical composition comprising at least one excipient.
[0097] As used herein, the term "pharmaceutically acceptable" means a compound that is physiologically acceptable when administered to a human. "Pharmaceutically acceptable" refers to molecular entities and compositions that are tolerated by the body and generally do not cause toxicity or allergic or similar untoward reactions, such as stomach upset, dizziness, etc. As sometimes used herein, the term "pharmaceutically acceptable" means approved by a federal or state government regulatory agency or listed in the United States Pharmacopeia or other recognized pharmacopoeia for use in animals, and more particularly in humans.
[0098] As used herein, the phrase "subject in need thereof" includes subjects, such as mammalian subjects, who would benefit from administration of a nucleic acid molecule, polypeptide, or vector of the present disclosure, e.g., to improve hemostasis. In one embodiment, the subject includes, but is not limited to, an individual with hemophilia. In another embodiment, the subject includes, but is not limited to, an individual who has developed an inhibitor to a therapeutic protein, e.g., a clotting factor, e.g., FVIII, and therefore requires bypass therapy. The subject may be an adult or a minor (e.g., under 12 years of age).
[0099] As used herein, the term "therapeutic protein" refers to any polypeptide known in the art that can be administered to a subject. In some embodiments, the therapeutic protein comprises a protein selected from a clotting factor, a growth factor, an antibody, a functional fragment thereof, or a combination thereof. As used herein, the term "clotting factor" refers to a naturally occurring or recombinantly produced molecule or analog thereof that prevents or shortens the duration of bleeding episodes in a subject. In other words, it refers to a molecule that has procoagulant activity, i.e., is responsible for the conversion of fibrinogen into an insoluble fibrin mesh that causes blood to clot or coagulate. As used herein, "clotting factor" includes activated clotting factors, their zymogens, or activatable clotting factors. An "activatable clotting factor" is a clotting factor in an inactive form (e.g., in its zymogen form) that can be converted to an activated form. The term "clotting factor" includes, but is not limited to, Factor I (FI), Factor II (FII), Factor V (FV), FVII, FVIII, FIX, Factor X (FX), Factor XI (FXI), Factor XII (FXII), Factor XIII (FXIII), von Willebrand Factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), a zymogen thereof, an activated form thereof, or any combination thereof.
[0100] As used herein, coagulation activity means the ability to complete the formation of a fibrin clot and / or participate in a cascade of biochemical reactions that reduce the severity, duration, or frequency of hemorrhage or bleeding episodes.
[0101] As used herein, "growth factor" includes any growth factor known in the art, including cytokines and hormones. In some embodiments, the growth factor is selected from the group consisting of adrenomedullin (AM), angiopoietin (Ang), autocrine motility factors, bone morphogenetic proteins (BMPs) (e.g., BMP2, BMP4, BMP5, BMP7), ciliary neurotrophic factor family members (e.g., ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), interleukin-6 (IL-6)), colony-stimulating factors (e.g., macrophage colony-stimulating factor (MCSF)), and the like. Macrophage colony-stimulating factor (m-CSF), granulocyte colony-stimulating factor (G-CSF), granulocyte-macrophage colony-stimulating factor (GM-CSF), epidermal growth factor (EGF), ephrins (e.g., ephrin A1, ephrin A2, ephrin A3, ephrin A4, ephrin A5, ephrin B1, ephrin B2, ephrin B3), erythropoietin (EPO), fibroblast growth factors (FGFs) (e.g., FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19 ... FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, FGF23), fetal bovine growth hormone (FBS), GDNF family members (e.g., glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, artemin), growth differentiation factor-9 (GDF9), hepatocyte growth factor (HGF), hepatoma-derived growth factor (HDGF), insulin, insulin-like growth factors (e.g., insulin-like growth factor-1 (IGF-1) or IGF-2), interleukins (IL) (e.g., IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7), keratinocyte growth factor ( KGF), migration stimulating factor (MSF), macrophage stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)), myostatin (GDF-8), neuregulins (e.g., neuregulin 1 (NRG1), NRG2, NRG3, NRG4), neurotrophins (e.g., brain-derived neurotrophic factor (BDNF), nerve growth factor (NGF), neurotrophin-3 (NT-3), NT-4, placental growth factor (PGF), platelet-derived growth factor (PDGF), renalase (RNLS), T-cell growth factor (TCGF), thrombopoietin (TPO), transforming growth factors (e.g., transforming growth factor alpha (TGF-α), TGF-β, tumor necrosis factor-alpha (TNF-α), and vascular endothelial growth factor (VEGF).
[0102] In some embodiments, the therapeutic protein is encoded by a gene selected from X-linked dystrophin, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof.
[0103] As used herein, the terms "heterologous" or "exogenous" refer to a molecule that is not normally found in a given context, e.g., a cell or polypeptide. An exogenous or heterologous molecule can be introduced into a cell and is present only after manipulation of the cell, e.g., by transfection or other form of genetic engineering, or a heterologous amino acid sequence can be present in a protein in which it is not found in nature.
[0104] As used herein, the term "heterologous nucleotide sequence" refers to a nucleotide sequence that does not naturally occur with a given polynucleotide sequence. In one embodiment, the heterologous nucleotide sequence encodes a polypeptide that can extend the half-life of a therapeutic protein, e.g., a clotting factor, e.g., FVIII. In another embodiment, the heterologous nucleotide sequence encodes a polypeptide that increases the hydrodynamic radius of a therapeutic protein, e.g., a clotting factor, e.g., FVIII. In other embodiments, the heterologous nucleotide sequence encodes a polypeptide that improves one or more pharmacokinetic properties of a therapeutic protein without significantly affecting its biological activity or function (e.g., procoagulant activity). In some embodiments, the therapeutic protein is linked or connected to the polypeptide encoded by the heterologous nucleotide sequence by a linker. Non-limiting examples of polypeptide moieties encoded by heterologous nucleotide sequences include immunoglobulin constant regions or portions thereof, albumin or fragments thereof, albumin-binding moieties, transferrin, the PAS polypeptide of U.S. Patent Application No. 20100292130, HAP sequences, transferrin or fragments thereof, the C-terminal peptide of the beta subunit of human chorionic gonadotropin (CTP), albumin-binding small molecules, XTEN sequences, FcRn-binding moieties (e.g., complete Fc regions or portions thereof that bind to FcRn), single-chain Fc regions (ScFc regions, e.g., those described in US2008 / 0260738, WO2008 / 012543 or WO2008 / 1439545), polyglycine linkers, polyserine linkers, peptides, and 50% or more of the above-mentioned polypeptides, among others. The heterologous nucleotide sequence may comprise a short polypeptide of 6-40 amino acids of two types of amino acids selected from glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), and proline (P), or a combination of two or more thereof, with a degree of secondary structure varying from less than 50% to more than 50%. In some embodiments, the polypeptide encoded by the heterologous nucleotide sequence is linked to a non-polypeptide moiety. Non-limiting examples of non-polypeptide moieties include polyethylene glycol (PEG), albumin-binding small molecules, polysialic acid, hydroxyethyl starch (HES), derivatives thereof, or any combination thereof.
[0105] As used herein, the term "Fc region" is defined as the portion of a polypeptide corresponding to the Fc region of a native Ig, i.e., formed by the dimeric association of the Fc domains of each of its two heavy chains. A native Fc region forms a homodimer with another Fc region. In contrast, as used herein, the term "genetically fused Fc region" or "single-chain Fc region" (scFc region) refers to a synthetic dimeric Fc region composed of Fc domains genetically linked (i.e., encoded by a single contiguous gene sequence) in a single polypeptide chain.
[0106] In one embodiment, "Fc region" refers to that portion of a single Ig heavy chain beginning at the hinge region just upstream of the papain cleavage site (i.e., residue 216 of IgG, where the first residue of the heavy chain constant region is 114) and ending at the C-terminus of the antibody. Thus, a complete Fc domain includes at least the hinge, CH2, and CH3 domains.
[0107] The Fc region of an Ig constant region can comprise CH2, CH3, and CH4 domains, as well as a hinge region, depending on the Ig isotype. Chimeric proteins comprising the Fc region of an Ig confer several desirable properties to the chimeric protein, including increased stability, increased serum half-life (see Capon et al., 1989, Nature 337:525), and binding to Fc receptors such as the neonatal Fc receptor (FcRn) (U.S. Patent Nos. 6,086,875, 6,485,726, 6,030,613; WO03 / 077834; US2003-0235536A1), which are incorporated herein by reference in their entireties.
[0108] As used herein in comparison with a nucleotide sequence of the present disclosure, a "reference nucleotide sequence" is a polynucleotide sequence that is substantially identical to the nucleotide sequence of the present disclosure, except that the sequence is not optimized. For example, the codon-optimized BDD of SEQ ID NO: 1 The reference nucleotide sequence for a nucleic acid molecule consisting of FVIII and a heterologous nucleotide sequence encoding a single-chain Fc region linked at its 3' end to SEQ ID NO: 1 is a nucleic acid molecule consisting of the original (or "parent") BDD FVIII of SEQ ID NO: 16 and the same heterologous nucleotide sequence encoding a single-chain Fc region linked at its 3' end to SEQ ID NO: 16.
[0109] As used herein, the term "optimized" with respect to a nucleotide sequence refers to a polynucleotide sequence that encodes a polypeptide, wherein the polynucleotide sequence has been mutated to enhance a property of the polynucleotide sequence. In some embodiments, optimization is performed to increase transcription levels, increase translation levels, increase steady-state mRNA levels, increase or decrease binding of regulatory proteins such as general transcription factors, increase or decrease splicing, or increase the yield of a polypeptide produced by the polynucleotide sequence. Examples of modifications that can be made to a polynucleotide sequence to optimize it include codon optimization, G / C content optimization, removal of repetitive sequences, removal of AT-rich elements, removal of cryptic splice sites, removal of cis-acting elements that repress transcription or translation, adding or removing poly-T or poly-A sequences, and adding or removing transcription-enhancing sequences such as Kozak consensus sequences around the transcription start site. These include adding sequences that may form stem-loop structures, removing sequences that may form destabilizing sequences, removing destabilizing sequences, and combinations of two or more thereof.
[0110] II. Nucleic acid molecules The present disclosure relates to a plasmid-like, capsid-free nucleic acid molecule encoding a therapeutic protein or gene capable of modulating the expression of a target protein. The capsid, the protein shell of a virus, encloses the viral genetic material. The capsid is known to aid in the function of the virion by protecting the viral genome, delivering the genome to the host, and interacting with the host. Nevertheless, the viral capsid may limit the packaging capacity of the vector and / or be a factor in inducing an immune response, especially when used in gene therapy.
[0111] AAV vector has emerged as one of the more common types of gene therapy vector.However, the existence of capsid limits the usefulness of AAV vector in gene therapy.In particular, capsid itself can limit the size of the transgene contained in vector to less than 4.5kb.Various therapeutic proteins that may be useful in gene therapy can easily exceed this size even before adding regulatory elements.
[0112] Furthermore, the proteins that make up the capsid can serve as antigens that can be targeted by a subject's immune system. AAV is very common in the general population, and most people have been exposed to AAV during their lifetime. As a result, most potential gene therapy recipients may already have mounted an immune response to AAV and are therefore more likely to reject the therapy.
[0113] Certain aspects of the present disclosure aim to overcome these deficiencies of AAV vectors. Specifically, certain aspects of the present disclosure relate to nucleic acid molecules comprising a first ITR, a second ITR, and a gene cassette encoding, for example, a therapeutic protein and / or an miRNA. In some embodiments, the nucleic acid molecule does not comprise a gene encoding a capsid protein, a replication protein, or an assembly protein. In some embodiments, the gene cassette encodes a therapeutic protein. In some embodiments, the therapeutic protein comprises a coagulation factor. In some embodiments, the gene cassette encodes an miRNA. In certain embodiments, the gene cassette is disposed between the first ITR and the second ITR. In some embodiments, the nucleic acid molecule further comprises one or more non-coding regions. In certain embodiments, the one or more non-coding regions comprise a promoter sequence, an intron, a post-transcriptional regulatory element, a 3'UTR poly(A) sequence, or any combination thereof.
[0114] In one embodiment, the gene cassette is a single-stranded nucleic acid. In another embodiment, the gene cassette is a double-stranded nucleic acid.
[0115] In one embodiment, the nucleic acid molecule comprises: (a) The first ITR, which is an ITR of a non-AAV family member of the Parvoviridae family; (b) tissue-specific promoter sequences, e.g., the TTP promoter; (c) introns, e.g., synthetic introns; (d) nucleotides encoding miRNAs or therapeutic proteins, e.g., clotting factors; (e) post-transcriptional regulatory elements, e.g., WPREs; (f) 3′UTR poly(A) tail sequence, e.g., bGHpA; (g) Non-AAV family members of the Parvoviridae family The second ITR being an ITR.
[0116] In one embodiment, the nucleic acid molecule comprises: (a) The first ITR, which is an ITR of a non-AAV family member of the Parvoviridae family; (b) tissue-specific promoter sequences, e.g., the TTP promoter; (c) introns, e.g., synthetic introns; (d) nucleotides encoding a miRNA, wherein the miRNA downregulates expression of a target gene selected from SOD1, HTT, RHO, or any combination thereof; (e) posttranscriptional regulatory elements, e.g., WPRE; (f) 3′UTR poly(A) tail sequence, e.g., bGHpA; (g) The second ITR is an ITR of a non-AAV family member of the Parvoviridae family.
[0117] In one embodiment, the nucleic acid molecule comprises: (a) The first ITR, which is an ITR of a non-AAV family member of the Parvoviridae family; (b) tissue-specific promoter sequences, e.g., the TTP promoter; (c) introns, e.g., synthetic introns; (d) nucleotides encoding X-linked dystrophin, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof; (e) posttranscriptional regulatory elements, e.g., WPRE; (f) 3′UTR poly(A) tail sequence, e.g., bGHpA; (g) The second ITR is an ITR of a non-AAV family member of the Parvoviridae family.
[0118] In one embodiment, the nucleic acid molecule comprises: (a) the first ITR, which is an ITR of an AAV, e.g., AAV serotype 2 genome; (b) tissue-specific promoter sequences, e.g., the TTP promoter; (c) introns, e.g., synthetic introns; (d) nucleotides encoding FVIII; wherein the nucleotides have at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a nucleotide sequence selected from SEQ ID NOs: 1-14 or SEQ ID NO: 71, and the FVIII encoded by the nucleotides retains FVIII activity; (e) posttranscriptional regulatory elements, e.g., WPRE; (f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and (g) A second ITR that is an ITR of an AAV, e.g., AAV serotype 2 genome.
[0119] In one embodiment, the nucleic acid molecule comprises: (a) the first ITR, which is an ITR of an AAV, e.g., AAV serotype 2 genome; (b) tissue-specific promoter sequences, e.g., the TTP promoter; (c) introns, e.g., synthetic introns; (d) nucleotides encoding miRNAs, wherein the miRNAs downregulate the expression of target genes, e.g., SOD1, HTT, RHO, and any combination thereof; (f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and (g) A second ITR that is an ITR of an AAV, e.g., AAV serotype 2 genome.
[0120] In one embodiment, the nucleic acid molecule comprises: (a) the first ITR, which is an ITR of an AAV, e.g., AAV serotype 2 genome; (b) tissue-specific promoter sequences, e.g., the TTP promoter; (c) introns, e.g., synthetic introns; (d) nucleotides encoding X-linked dystrophin, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha-glucosidase), or any combination thereof; (f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and (g) A second ITR that is an ITR of an AAV, e.g., AAV serotype 2 genome.
[0121] In another embodiment, the nucleic acid molecule comprises: (a) First ITR; (b) tissue-specific promoter sequences, e.g., the TTP promoter; (c) introns, e.g., synthetic introns; (d) nucleotides encoding miRNAs or therapeutic proteins, e.g., clotting factors; (e) post-transcriptional regulatory elements, e.g., WPREs; (f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and (g) Second ITR; wherein one of the first or second ITRs is an ITR of a non-AAV family member of the Parvoviridae family, and the other ITR is an ITR of an AAV, eg, AAV serotype 2 genome.
[0122] In one embodiment, the nucleic acid molecule comprises: (a) First ITR; (b) tissue-specific promoter sequence, TTP promoter; (c) introns, e.g., synthetic introns; (d) nucleotides encoding miRNAs or therapeutic proteins, e.g., clotting factors; (e) post-transcriptional regulatory elements, e.g., WPREs; (f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and (g) Second ITR; Here, the first ITR is a synthetic ITR, the second ITR is a synthetic ITR, or both the first ITR and the second ITR are synthetic ITRs.
[0123] A. Inverted terminal repeats Certain embodiments of the present disclosure relate to nucleic acid molecules comprising a first ITR, e.g., a 5'ITR, and a second ITR, e.g., a 3'ITR. Generally, ITRs are involved in the replication and release, or excision, of parvovirus (e.g., AAV) DNA from prokaryotic plasmids (Samulski et al., 1983, 1987; Senapathy et al., 1984; Gottlieb and Muzyczka, 1988). Furthermore, ITRs appear to be the minimal sequences required for AAV proviral incorporation and packaging of AAV DNA into virions (McLaughlin et al., 1988; Samulski et al., 1989). These elements are essential for efficient propagation of the parvovirus genome. The minimal defined elements essential for ITR function are a Rep binding site (e.g., RBS; for AAV2, GCGCGCTCGCTCGCTC (SEQ ID NO: 104)) and a terminal resolution site (e.g., TRS; for AAV2, AGTTGG (SEQ ID NO: 105)), which is hypothesized to be a variable palindromic sequence that further allows hairpin formation. Palindromic nucleotide regions usually function together in cis as an origin of DNA replication and as a packaging signal for the virus. Complementary sequences within the ITRs fold into a hairpin structure during DNA replication. In some embodiments, the ITRs fold into a hairpin T-shaped structure. In other embodiments, the ITRs fold into a non-T-shaped hairpin structure, e.g., a U-shaped hairpin structure. Data suggests that the T-shaped hairpin structure of AAV ITRs can suppress expression of transgenes adjacent to the ITRs. See, e.g., Zhou et al., Scientific Reports 7:5432 (July 14, 2017). This form of suppression can be avoided by utilizing ITRs that do not form T-shaped hairpin structures. Thus, in certain aspects, polynucleotides containing non-AAV ITRs have improved transgene expression compared to polynucleotides containing AAV ITRs that form T-shaped hairpins.
[0124] In some embodiments, the ITR comprises a naturally occurring ITR, for example, the ITR comprises all or a portion of a parvovirus ITR. In some embodiments, the ITR comprises a synthetic sequence. In one embodiment, the first ITR or the second ITR comprises a synthetic sequence. In another embodiment, each of the first ITR and the second ITR comprises a synthetic sequence. In some embodiments, the first ITR or the second ITR comprises a naturally occurring sequence. In another embodiment, each of the first ITR and the second ITR comprises a naturally occurring sequence.
[0125] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR, e.g., a truncated ITR. In some embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment is at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, or at least about 50 nucleotides. and / or at least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides, or at least about 600 nucleotides; wherein the ITR retains a functional characteristic of a naturally occurring ITR. In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 129 nucleotides and wherein the ITR retains a functional characteristic of a naturally occurring ITR.In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, wherein the fragment comprises at least about 102 nucleotides, and wherein the ITR retains the functional properties of the naturally occurring ITR.
[0126] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR, where the fragment is at least about 5%, at least about 10% of the length of the naturally occurring ITR. , at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%; wherein the fragment retains a functional property of the naturally occurring ITR.
[0127] In certain embodiments, the ITR, when properly aligned, has at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 127%, at least In another embodiment, the ITR comprises or consists of a sequence having at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In another embodiment, the ITR comprises or consists of a sequence having at least 90% sequence identity with the homologous portion of the naturally occurring ITR when properly aligned; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 80% sequence identity with the homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 70% sequence identity with the homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 60% sequence identity with the homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR.In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 50% sequence identity with a homologous portion of a naturally occurring ITR; wherein the ITR retains the functional properties of the naturally occurring ITR.
[0128] In some embodiments, ITR comprises ITR from AAV genome.In some embodiments, ITR is ITR of AAV genome selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11 and any combination thereof.In certain embodiments, ITR is ITR of AAV2 genome.In another embodiment, ITR is a synthetic sequence that is genetically engineered to comprise ITR from one or more AAV genomes at its 5' and 3' ends.
[0129] In some embodiments, the ITRs are not derived from an AAV genome. In some embodiments, the ITRs are non-AAV ITRs. In some embodiments, the ITRs are derived from viruses including, but not limited to, Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, etc. The ITRs are non-AAV genomic ITRs from the Parvoviridae family of viruses selected from the group consisting of Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidesovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof. In certain embodiments, the ITRs are derived from Erythrovirus parvovirus B19 (a human virus). In another embodiment, the ITRs are derived from a Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, e.g., MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, e.g., MDPV strain YY. In some embodiments, the ITRs are derived from a porcine parvovirus, e.g., porcine parvovirus U44978. In some embodiments, the ITRs are derived from minute virus of mice, e.g., minute virus of mice U34256. In some embodiments, the ITRs are derived from canine parvovirus, e.g., canine parvovirus M19296. In some embodiments, the ITRs are derived from mink enteritis virus, e.g., mink enteritis virus D00765. In some embodiments, the ITRs are derived from Dependoparvovirus. In one embodiment, the Dependoparvovirus is a Dependovirus goose parvovirus (GPV) strain. In a specific embodiment, the GPV strain is attenuated, e.g., GPV strain 82-0321V. In another specific embodiment, the GPV strain is pathogenic, e.g., GPV strain B.
[0130] The first ITR and the second ITR of a nucleic acid molecule can be derived from the same genome, for example, the genome of the same virus, or from different genomes, for example, the genomes of two or more different virus genomes. In certain embodiments, the first ITR and the second ITR are derived from the same AAV genome. In specific embodiments, the two ITRs present in the nucleic acid molecule of the present invention are the same, and may in particular be AAV2 ITRs. In other embodiments, the first ITR is derived from an AAV genome, and the second ITR is not derived from an AAV genome (e.g., a non-AAV genome). In other embodiments, the first ITR is not derived from an AAV genome (e.g., a non-AAV genome), and the second ITR is derived from an AAV genome. In yet other embodiments, both the first ITR and the second ITR are not derived from an AAV genome (e.g., a non-AAV genome). In a specific embodiment, the first ITR and the second ITR are identical.
[0131] In some embodiments, the first ITR is derived from an AAV genome and the second ITR is derived from a genome selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome and the first ITR is derived from a virus such as Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, The virus is derived from a genome selected from the group consisting of Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof. In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from the same genome.In other embodiments, the first ITR and the second ITR are derived from a genome selected from the group consisting of Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstyldensovirus, and any combination thereof, wherein the first ITR and the second ITR are derived from different genomes.
[0132] In some embodiments, the first ITR is derived from an AAV genome and the second ITR is derived from erythrovirus parvovirus B19 (a human virus). In other embodiments, the second ITR is derived from an AAV genome and the first ITR is derived from erythrovirus parvovirus B19 (a human virus).
[0133] In certain embodiments, the first and / or second ITR comprises or consists of all or a portion of an ITR from B 19. In some embodiments, the first and / or second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171, wherein the first and / or second ITR retains a functional property of the B 19 ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 100%, at least about 110%, at least about 120%, at least about 130%, at least about 140%, at least about 150%, at least about 160%, at least about 170%, at least about 180%, at least about 190%, at least about 200%, at least about 210%, at least about 220%, at least about 230%, at least about 240%, at least about 250%, at least about 260%, at least about 270%, at least about 280%, at least about 290%, at least about 300%, at least about 310%, at least about 320%, at least about 330%, at least about 340%, at least about 350%, at least about 360%, at least about 370%, at least about 380%, at least about 390%, at least about 400%, at least about 410%, at least about 420%, at least about 430%, at least about 440%, at least about 450%, at least about 460%, at least about 470%, at least about 480%, at least about 490%, at least about 500%, at least about 510%, at least about 520%, at least about 530%, at least about 540%, at least about 550%, at least about 560%, at least about 570%, at least about 580%, at least about 590%, at least about In certain embodiments, the first ITR and / or the second ITR comprise or consist of a nucleotide sequence that is at least about 99% or 100% identical to the first ITR and / or the second ITR, and wherein the first ITR and / or the second ITR are capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.
[0134] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 167, 168, 169, 170, and 171. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 167. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 168. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 169. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 170. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 171.
[0135] [Table 1] [Table 2]
[0136] In certain embodiments, the first ITR and / or second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 169, wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167 and retains the functional properties of the B19 ITR from which it is derived. In some embodiments, the first and / or second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 169, wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 167, and wherein the first and / or second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.
[0137] In some embodiments, the first ITR is derived from the AAV genome and the second ITR is derived from the GP In other embodiments, the second ITR is derived from an AAV genome and the first ITR is derived from GPV.
[0138] In certain embodiments, the first and / or second ITRs comprise or consist of all or a portion of an ITR derived from GPV. In some embodiments, the first and / or second ITRs comprise or consist of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 172, 173, 174, 175, and 176, wherein the first and / or second ITRs retain the functional properties of the GPV ITRs from which they are derived. In some embodiments, the first and / or second ITRs comprise or consist of all or a portion of an ITR derived from GPV. In some embodiments, the first and / or second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 172, 173, 174, 175, and 176, wherein the first and / or second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin. In some embodiments, the first and / or second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 172, 173, 174, 175, and 176. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 172. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 173.In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 174. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 175. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 176.
[0139] In certain embodiments, the first ITR and / or second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 174, wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, and wherein the first ITR and / or second ITR retains a functional property of the GPV ITR from which it is derived. In some embodiments, the first and / or second ITR comprises a nucleotide sequence comprising the minimal nucleotide sequence set forth in SEQ ID NO: 174, wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, and wherein the first and / or second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.
[0140] In certain embodiments, the first ITR and / or second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 176, wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, and wherein the first ITR and / or second ITR retains the functional properties of the GPV ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises a nucleotide sequence, wherein the nucleotide sequence comprises the minimal nucleotide sequence set forth in SEQ ID NO: 176, wherein the nucleotide sequence is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 172, and wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.
[0141] In certain embodiments, one of the first or second ITRs comprises or consists of all or a portion of an ITR from AAV2. In some embodiments, the first or second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 177 or 178, wherein the first and / or second ITR retains the functional properties of the AAV2 ITR from which it is derived. In some embodiments, the first ITR or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NO: 177 or 178, wherein the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 177 or 178. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 177. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO:178.
[0142] In some embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a Muscovy duck parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is attenuated, for example, MDPV strain FZ91-30. In other embodiments, the MDPV strain is pathogenic, for example, MDPV strain YY.
[0143] In some embodiments, the first ITR is derived from an AAV genome and the second ITR is derived from Dependoparvovirus. In some embodiments, the second ITR is derived from an AAV genome and the first ITR is derived from Dependoparvovirus. In other embodiments, the first ITR is derived from an AAV genome and the second ITR is derived from a Dependovirus goose parvovirus (GPV) strain. In other embodiments, the second ITR is derived from an AAV genome and the first ITR is derived from a Dependovirus GPV strain. In certain embodiments, the GPV strain is attenuated, e.g., GPV strain 82-0321V. In other embodiments, the GPV strain is pathogenic, e.g., GPV strain B.
[0144] In certain embodiments, the first ITR is derived from an AAV genome, and the second ITR is derived from a genome selected from the group consisting of porcine parvovirus, e.g., porcine parvovirus strain U44978; minute virus of mice, e.g., minute virus of mice strain U34256; canine parvovirus, e.g., canine parvovirus strain M19296; mink enteritis virus, e.g., mink enteritis virus strain D00765; and any combination thereof. In other embodiments, the second ITR is derived from an AAV genome, and the first ITR is derived from a genome selected from the group consisting of porcine parvovirus, e.g., porcine parvovirus strain U44978; minute virus of mice, e.g., minute virus of mice strain U34256; canine parvovirus, e.g., canine parvovirus strain M19296; mink enteritis virus, e.g., mink enteritis virus strain D00765; and any combination thereof.
[0145] In another specific embodiment, the ITR is a synthetic sequence engineered to contain ITRs at its 5' and 3' ends that are not derived from the AAV genome. In another embodiment, the ITR is a synthetic sequence engineered to contain ITRs at its 5' and 3' ends that are derived from one or more non-AAV genomes. The two ITRs present in the nucleic acid molecule of the present invention may be the same or different non-AAV genomes. In particular, the ITRs may be derived from the same non-AAV genome. In a specific embodiment, the two ITRs present in the nucleic acid molecule of the present invention may be the same, and in particular, may be AAV2 ITRs.
[0146] In some embodiments, the ITR sequence comprises one or more palindromic sequences. The palindromic sequences of the ITRs disclosed herein include, but are not limited to, naturally occurring palindromic sequences (i.e., sequences found in nature), synthetic sequences (i.e., sequences not found in nature), such as pseudopalindromic sequences, and combinations or modifications thereof. A "pseudopalindromic sequence" is a palindromic DNA sequence containing an imperfect palindromic sequence that shares less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%, or less than 80%, including no identity, with a sequence in a naturally occurring AAV or non-AAV palindromic sequence that forms a secondary structure. Naturally occurring palindromic sequences can be obtained or derived from any genome disclosed herein. Synthetic palindromic sequences can be based on any genome disclosed herein.
[0147] The palindromic sequence may be continuous or interrupted. In some embodiments, the palindromic sequence is interrupted, where the palindromic sequence comprises the insertion of a second sequence. In some embodiments, the second sequence comprises a promoter, an enhancer, an integration site for an integrase (e.g., a site for Cre or Flp recombinase), an open reading frame for a gene product, or a combination thereof.
[0148] In some embodiments, the ITRs form a hairpin loop structure. In one embodiment, the first ITR forms a hairpin structure. In another embodiment, the second ITR forms a hairpin structure. In yet another embodiment, both the first ITR and the second ITR form a hairpin structure. In some embodiments, the first ITR and / or the second ITR do not form a T-shaped hairpin structure. In certain embodiments, the first ITR and / or the second ITR form a non-T-shaped hairpin structure. In some embodiments, a non-T-shaped hairpin structure is formed. The shaped hairpin structure includes a U-shaped hairpin structure.
[0149] In some embodiments, the ITRs in the nucleic acid molecules described herein may be transcriptionally activating ITRs. The transcriptionally activating ITRs may comprise all or part of a wild-type ITR that has been transcriptionally activated by the inclusion of at least one transcriptionally active element. Various transcriptionally active elements are suitable for use in this context. In some embodiments, the transcriptionally active element is a constitutive transcriptionally active element. Constitutive transcriptionally active elements provide a sustained level of gene transcription and are preferred when continuous expression of the transgene is desired. In other embodiments, the transcriptionally active element is an inducible transcriptionally active element. Inducible transcriptionally active elements generally exhibit low activity in the absence of an inducer (or inducing conditions) and are upregulated in the presence of an inducer (or when inducing conditions are switched on). Inducible transcriptionally active elements may be preferred when expression is desired only at certain times or in certain locations, or when it is desirable to titrate expression levels using an inducer. Transcriptionally active elements may also be tissue-specific; i.e., they are active only in certain tissues or cell types.
[0150] Transcriptionally active elements can be incorporated into ITRs in various ways. In some embodiments, transcriptionally active elements are incorporated 5' to any portion of the ITR or 3' to any portion of the ITR. In other embodiments, the transcriptionally active element of a transcriptionally activating ITR is located between two ITR sequences. When a transcriptionally active element contains two or more elements that must be spaced apart, the elements can alternate between portions of the ITR. In some embodiments, the hairpin structure of an ITR is deleted and replaced with an inverted repeat of a transcription element. This latter configuration would form a hairpin that mimics the deleted portion of the structure. Multiple tandem transcriptionally active elements can be present in a transcriptionally activating ITR, and these can be adjacent or separated. Additionally, protein binding sites (e.g., Rep binding sites) can be introduced into the transcriptionally active element of a transcriptionally activating ITR. The transcriptionally active element can include any sequence that allows for controlled transcription of DNA by RNA polymerase to form RNA, and can include, for example, the transcriptionally active elements defined below.
[0151] Transcriptionally activating ITRs provide both transcriptional activation and ITR functions to nucleic acid molecules of relatively limited nucleotide sequence length, effectively maximizing the length of the transgene that can be carried and expressed from the nucleic acid molecule. Incorporation of transcriptionally activating elements into ITRs can be achieved in a variety of ways. Comparison of the ITR sequences and sequence requirements of transcriptionally activating elements can provide insight into how to encode the elements within the ITRs. For example, transcriptional activity can be added to an ITR through the introduction of specific alterations in the ITR sequence that duplicate functional elements of the transcriptionally activating element. Several techniques exist in the art for efficiently adding, deleting, and / or modifying specific nucleotide sequences at specific sites (see, e.g., Deng and Nickoloff (1992) Anal. Biochem. 200:81-88). Another method for creating transcriptionally activating ITRs involves the introduction of restriction sites at desired positions in the ITRs. Furthermore, multiple transcriptionally activating elements can be incorporated into transcriptionally activating ITRs using methods known in the art.
[0152] By way of illustration, transcriptionally activating ITRs can be generated by the inclusion of one or more transcriptionally active elements, such as a TATA box, a GC box, a CCAAT box, an Sp1 site, an Inr region, a CRE (cAMP regulatory element) site, an ATF-1 / CRE site, an APBβ box, an APBα box, a CArG box, a CCAC box, or any other element involved in transcription as known in the art.
[0153] B. Therapeutic Proteins Certain aspects of the present disclosure are directed to nucleic acid molecules comprising a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein. In some embodiments, the gene cassette encodes one therapeutic protein. In some embodiments, the gene cassette encodes more than one therapeutic protein. In some embodiments, the gene cassette encodes two or more copies of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more variants of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more different therapeutic proteins.
[0154] Certain embodiments of the present disclosure are directed to nucleic acid molecules comprising a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a coagulation factor. In some embodiments, the coagulation factor is selected from the group consisting of FI, FII, FIII, FIV, FV, FVI, FVII, FVIII, FIX, FX, FXI, FXII, FXIII, VWF, prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), any zymogen thereof, any active form thereof, and any combination thereof. In one embodiment, the coagulation factor comprises FVIII or a variant or fragment thereof. In another embodiment, the clotting factor comprises FIX or a variant or fragment thereof. In another embodiment, the clotting factor comprises FVII or a variant or fragment thereof. In another embodiment, the clotting factor comprises VWF or a variant or fragment thereof.
[0155] 1. Clotting factors In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a factor VIII polypeptide. As used herein, "factor VIII," abbreviated as "FVIII" throughout this application, unless otherwise specified, refers to a FVIII polypeptide that is functional in its normal role in coagulation. Thus, the term FVIII includes functional variant polypeptides. "FVIII protein" is used interchangeably with FVIII polypeptide (or protein) or FVIII. Examples of FVIII function include, but are not limited to, the ability to activate coagulation, act as a cofactor for factor IX, or inhibit Ca. 2+ FVIII has the ability to form a tenase complex with factor IX in the presence of phospholipids and subsequently convert factor X to the activated form Xa. The FVIII protein can be human, porcine, canine, rat, or murine FVIII protein. Furthermore, comparisons between human-derived FVIII and FVIII from other species have identified conserved residues likely required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US Pat. No. 6,251,632). Full-length polypeptide and polynucleotide sequences are known, as are many functional fragments, mutants, and modified versions. Various FVIII amino acid and nucleotide sequences are disclosed, for example, in U.S. Patent Application Publication Nos. 2015 / 0158929, 2014 / 0308280, and 2014 / 0370035 and International Patent Application Publication No. 2015 / 106052. FVIII polypeptides include, for example, full-length FVIII, full-length FVIII minus the N-terminal Met, mature FVIII (minus the signal sequence), mature FVIII with an additional N-terminal Met, and / or FVIII with a full or partial deletion of the B domain. FVIII variants, whether partial or full deletions, include: Contains a deletion of the B domain.
[0156] a. Polynucleotide sequences encoding FVIII and FVIII proteins In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a factor VIII polypeptide. As used herein, "factor VIII," abbreviated as "FVIII" throughout this application, unless otherwise specified, refers to a FVIII polypeptide that is functional in its normal role in coagulation. Thus, the term FVIII includes functional variant polypeptides. "FVIII protein" is used interchangeably with FVIII polypeptide (or protein) or FVIII. Examples of FVIII function include, but are not limited to, the ability to activate coagulation, act as a cofactor for factor IX, or inhibit Ca. 2+ FVIII has the ability to form a tenase complex with factor IX in the presence of phospholipids and subsequently convert factor X to the activated form Xa. The FVIII protein can be human, porcine, canine, rat, or murine FVIII protein. Furthermore, comparisons between human-derived FVIII and FVIII from other species have identified conserved residues likely required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US Pat. No. 6,251,632). Full-length polypeptide and polynucleotide sequences are known, as are many functional fragments, mutants, and modified versions. Various FVIII amino acid and nucleotide sequences are disclosed, for example, in U.S. Patent Application Publication Nos. 2015 / 0158929, 2014 / 0308280, and 2014 / 0370035 and International Patent Application Publication No. 2015 / 106052. FVIII polypeptides include, for example, full-length FVIII, full-length FVIII minus the N-terminal Met, mature FVIII (minus the signal sequence), mature FVIII with an additional N-terminal Met, and / or FVIII with a total or partial deletion of the B domain. FVIII variants include deletions of the B domain, whether partial or total.
[0157] As used herein, the FVIII portion of the chimeric protein has FVIII activity, which can be measured by any method known in the art. Several tests are available to assess the function of the coagulation system: activated partial thromboplastin time (aPTT) test, chromogenic assays, ROTEM assay, prothrombin time (PT) test (also used to determine INR), fibrinogen test (often by the Clauss method), platelet count, platelet function test (often by PFA-100), TCT, bleeding time, mixing test (whether abnormalities are corrected when the patient's plasma is mixed with normal plasma), coagulation factor assays, antiphospholipid antibodies, D-dimer, genetic tests (e.g., factor V Leiden, prothrombin mutation G20210A), dilute Russell's viper venom time (dRVVT), various platelet function tests, thromboelastography (TEG or Sonoclot), thromboelastometry (TEM®, e.g., ROTEM®), or euglobulin lysis time (ELT).
[0158] The aPTT test is a performance indicator that measures the effectiveness of both the "intrinsic" coagulation pathway (also called the contact activation pathway) and the conventional coagulation pathway. This test is commonly used to measure the clotting activity of commercially available recombinant clotting factors, such as FVIII. This test is used in conjunction with the prothrombin time (PT), which measures the extrinsic pathway.
[0159] ROTEM analysis provides information on the entire dynamics of hemostasis: clotting time, clot formation, clot stability, and lysis. The various parameters of thromboelastometry depend on the activity of the plasma coagulation system, platelet function, fibrinolysis, or many factors that affect their interactions. This assay can provide a complete picture of secondary hemostasis.
[0160] The mechanism of the chromogenic assay is based on the principle of the blood coagulation cascade, in which activated FVIII, in the presence of activated factor IX, phospholipids, and calcium ions, accelerates the conversion of factor X to factor Xa. Factor Xa activity is assessed by the hydrolysis of the factor Xa-specific p-nitroanilide (pNA) substrate. The initial rate of p-nitroaniline release, measured at 405 nM, is directly proportional to the factor Xa activity, and hence the FVIII activity, in the sample.
[0161] Chromogenic assays are recommended by the FVIII and Factor IX Subcommittee of the Scientific and Standardization Committee (SSC) of the International Society on Thrombosis and Hemostatsis (ISTH). Since 1994, chromogenic assays have also been the reference method in the European Pharmacopoeia for the assignment of potency of FVIII concentrates. Thus, in one embodiment, the chimeric polypeptide comprising FVIII has FVIII activity comparable to that of a chimeric polypeptide comprising mature FVIII or BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).
[0162] In another embodiment, the chimeric proteins comprising FVIII of this disclosure have a factor Xa production rate comparable to that of chimeric proteins comprising mature FVIII or BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).
[0163] To activate factor X to form factor Xa, activated factor IX (factor IXa) must be activated with Ca 2+In the presence of FVIII, membrane phospholipids, and FVIII cofactors, FVIII hydrolyzes one arginine-isoleucine bond of factor X to form factor Xa. Therefore, the interaction of FVIII with factor IX is crucial in the aggregation pathway. In certain embodiments, chimeric polypeptides containing FVIII can interact with factor IXa at a rate comparable to that of chimeric polypeptides containing the mature FVIII sequence or BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).
[0164] Furthermore, FVIII binds to von Willebrand factor but is inactive in the circulating blood. When FVIII is not bound to VWF, it is rapidly degraded and released from VWF by the action of thrombin. In some embodiments, chimeric polypeptides comprising FVIII bind to von Willebrand factor at a level comparable to chimeric polypeptides comprising the sequence of mature FVIII or BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).
[0165] FVIII can be inactivated by activated protein C in the presence of calcium and phospholipids. Activated protein C cleaves the FVIII heavy chain after arginine 336 in the A1 domain, destroying the substrate interaction site for factor X, and after arginine 562 in the A2 domain, promoting dissociation of the A2 domain and destroying the interaction site with factor IXa. This cleavage also bisects the A2 domain (43 kDa), generating the A2-N domain (18 kDa) and the A2-C domain (25 kDa). Thus, activated protein C can catalyze multiple cleavages in the heavy chain. In one embodiment, a chimeric polypeptide containing FVIII is inactivated by activated protein C at levels comparable to those of a chimeric polypeptide containing the sequence of mature FVIII or BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®).
[0166] In other embodiments, the chimeric protein comprising FVIII has in vivo FVIII activity comparable to that of a chimeric polypeptide comprising mature FVIII sequence or BDD FVIII (e.g., ADVATE®, REFACTO®, or ELOCTATE®). In certain embodiments, the chimeric polypeptide comprising FVIII has the sequence or BDD FVIII of mature FVIII in the HemA mouse tail vein transection model. It is possible to protect HemA mice at levels comparable to chimeric polypeptides containing FVIII (eg, ADVATE®, REFACTO®, or ELOCTATE®).
[0167] The "B domain" of FVIII, as used herein, is identical to B domains known in the art, defined by internal amino acid sequence identity and the site of proteolytic cleavage by thrombin, e.g., residues Ser741 to Arg1648 of mature human FVIII. Other human FVIII domains are defined by the following amino acid residues relative to mature human FVIII: A1, residues Ala1 to Arg372 of mature FVIII; A2, residues Ser373 to Arg740 of mature FVIII; A3, residues Ser1690 to Ile2032 of mature FVIII; C1, residues Arg2033 to Asn2172 of mature FVIII; and C2, residues Ser2173 to Tyr2332 of mature FVIII. Unless otherwise specified, sequence residue numbers used herein without reference to any SEQ ID NO: correspond to the FVIII sequence without the signal peptide sequence (19 amino acids). The A3-C1-C2 sequence, also known as the FVIII heavy chain, includes residues Ser1690 to Tyr2332. The remaining sequence, residues Glu1649 to Arg1689, is commonly referred to as the FVIII light chain activation peptide. The locations of all of the domain boundaries, including the B domain, for porcine, mouse, and canine FVIII are also known in the art. In one embodiment, the B domain of FVIII is deleted ("B-domain deleted FVIII" or "BDD FVIII"). An example of a BDD FVIII is REFACTO® (recombinant BDD FVIII). In one specific embodiment, the B-domain deleted FVIII variant contains a deletion of amino acid residues 746 to 1648 of mature FVIII.
[0168] "B domain deleted FVIII" is a compound described in U.S. Patent Nos. 6,316,226, 6,346,513, 7,041,635, 5,789,203, 6,060,447, 5,595,886, 6,228,620, 5,972,885, 6,048,720, 5,595,886, and 5,595,886. Nos. 43,502, 5,610,278, 5,171,844, 5,112,950, 4,868,112, and 6,458,563, and International Patent Application Publication No. 2015106052 (PCT / US2015 / 010738). In some embodiments, the B domain-deleted FVIII sequence used in the disclosed methods comprises any one of the deletions disclosed in U.S. Patent No. 6,316,226 (also U.S. Patent No. 6,346,513) at column 4, line 4 to column 5, line 28 and in Examples 1-5. In another embodiment, the B-domain deleted FVIII is S743 / Q1638 B-domain deleted FVIII (SQ BDD FVIII) (e.g., a factor VIII having a deletion of amino acids 744 to 1637, e.g., a factor VIII having amino acids 1-743 and amino acids 1638-2332 of mature FVIII). In some embodiments, the B-domain deleted FVIII used in the methods of the disclosure has a deletion as disclosed in column 2, lines 26-51 and Examples 5-8 of U.S. Pat. No. 5,789,203 (also U.S. Pat. Nos. 6,060,447, 5,595,886, and 6,228,620). In some embodiments, the B-domain deleted factor VIII is described in U.S. Pat. No. 5,972,885, column 1, line 25 to column 2, line 40; U.S. Pat. No. 6,048,720, column 6, lines 1-22 and Example 1; U.S. Pat. No. 5,543,502, column 2, lines 17-46; U.S. Pat. No. 5,171,844, column 4, line 22 to column 5, line 36; column 2, lines 55-68, Figure 2, and Example 1 of U.S. Patent No. 5,112,950; column 2, line 2 to column 19, line 21 and Table 2 of U.S. Patent No. 4,868,112; column 2, line 1 to column 3, line 19, column 3, line 40 to column 4, line 67, column 7, line 43 to column 8, line 26, and column 11, line 5 to column 13, line 39 of U.S. Patent No. 7,041,635; or column 4, lines 25-53 of U.S. Patent No. 6,458,563. In some embodiments, as disclosed in WO 91 / 09122, the B domain-deleted FVIII has most of the B domain deleted but still contains the amino-terminal sequence of the B domain that is essential for in vivo protein processing of the primary translation product into two polypeptide chains. In some embodiments, a B-domain deleted FVIII is constructed by deleting amino acids 747-1638, i.e., substantially complete deletion of the B-domain. Hoeben RC et al. J. Biol. Chem. 265(13):7318-7323 (1990). B-domain deleted factor VIII may contain deletion of amino acids 771-1666 or amino acids 868-1562 of FVIII. Meulien P. et al., Protein Eng. 2(4):301-6 (1988). Additional B domain deletions that are part of the present disclosure include amino acids 982 to 1562 or 760 to 1639 (Toole et al., Proc. Natl. Acad. Sci. USA 83:5939-5942 (1986)), 797 to 1562 (Eaton et al., Biochemistry 25:8343-8347 (1986)), 741 to 1646 (Kaufman, WO 87 / 04187), 747 to 1560 (Sarver et al., DNA 6:553-564 (1987)), 741 to 1648 (Pasek, PCT Application No. 88 / 00831), or 816 to 1598 or 741 to 1648 (Lagner, Behring Inst. Mitt. (1988) No. 82:16-25; EP 295597). In one particular embodiment, the B-domain deleted FVIII comprises a deletion of amino acid residues 746 to 1648 of mature FVIII. In another embodiment, the B-domain deleted FVIII comprises a deletion of amino acid residues 745 to 1648 of mature FVIII. In some embodiments, the BDD FVIII comprises a single-chain FVIII containing a deletion of amino acids 765 to 1652, corresponding to full-length mature FVIII (also known as rVIII-SingleChain and AFSTYLA®). See U.S. Patent No. 7,041,635.
[0169] In other embodiments, BDD FVIII comprises a FVIII polypeptide containing a fragment of the B domain that retains one or more N-linked glycosylation sites, e.g., residues 757, 784, 828, 900, 963, or optionally 943, corresponding to the amino acid sequence of the full-length FVIII sequence. Examples of B domain fragments include 226 or 163 amino acids of the B domain (i.e., the first 226 or 163 amino acids of the B domain are retained) as disclosed in Miao, HZ et al., Blood 103(a):3412-3419 (2004), Kasuda, A et al., J. Thromb. Haemost. 6:1352-1359 (2008), and Pipe, SW et al., J. Thromb. Haemost. 9:2235-2242 (2011). In yet another embodiment, the BDD FVIII further comprises a point mutation at residue 309 (Phe to Ser) to improve expression of the BDD FVIII protein. See Miao, HZ et al., Blood 103(a):3412-3419 (2004). In yet another embodiment, the BDD FVIII comprises a FVIII polypeptide that contains a portion of the B domain but does not contain one or more furin cleavage sites (e.g., Arg1313 and Arg1648). Pipe, SW et al., J. See Thromb. Haemost. 9:2235-2242 (2011). In some embodiments, the BDD FVIII comprises a single-chain FVIII containing a deletion at amino acids 765 to 1652, corresponding to full-length mature FVIII (also known as rVIII-SingleChain and AFSTYLA®). See, e.g., US Pat. No. 7,041,635. Each of the above deletions can be made in any FVIII sequence.
[0170] Numerous functional FVIII variants are known, as discussed above and below. Furthermore, hundreds of non-functional mutations in FVIII have been identified in hemophilia patients, and it has been determined that the effects of these mutations on FVIII function are more likely to be due to their location within the three-dimensional structure of FVIII than to the nature of the substitution (Cutler et al., Hum. Mutat. 19:274-8 (2002)), which is incorporated herein by reference in its entirety. Furthermore, comparison of human-derived FVIII with FVIII from other species has identified conserved residues that may be required for function (Cameron et al., Thromb. Haemost. 79:317-22 (1998); US 6,251,632), which is incorporated herein by reference in its entirety.
[0171] In some embodiments, the FVIII polypeptide comprises a FVIII variant or a fragment thereof, wherein the FVIII variant or fragment thereof has FVIII activity. In some embodiments, the gene cassette encodes a full-length FVIII polypeptide. In other embodiments, the gene cassette encodes a B-domain deleted (BDD) FVIII polypeptide, wherein all or part of the B domain of FVIII is deleted. In a specific embodiment, the gene cassette encodes a polypeptide comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 106, 107, 109, 110, 111, or 112. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 17 or a fragment thereof. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 106, or a fragment thereof. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 107. In some embodiments, the gene cassette encodes a polypeptide having the amino acid sequence of SEQ ID NO: 109, or a fragment thereof. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 16. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 109.
[0172] In some embodiments, the gene cassette of the present disclosure encodes a FVIII polypeptide comprising a signal peptide or a fragment thereof. In other embodiments, the gene cassette encodes a FVIII polypeptide lacking the signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.
[0173] In some embodiments, the gene cassette comprises a nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence is codon-optimized. In certain embodiments, the gene cassette comprises a nucleotide sequence disclosed in International Patent Application No. PCT / US2017 / 015879, which is incorporated by reference in its entirety. In some embodiments, the gene cassette comprises a nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence is codon-optimized. In certain embodiments, the gene cassette comprises a nucleotide sequence selected from SEQ ID NOs: 1-14 and at least 85 In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 71. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19.
[0174] i. Codon-optimized nucleotide sequences encoding FVIII polypeptides In some embodiments, the nucleic acid molecule of the present disclosure comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the first ITR and the second ITR are derived from an AAV genome, and the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide. In some embodiments, the codon-optimized nucleotide sequence encodes a full-length FVIII polypeptide. In other embodiments, the codon-optimized nucleotide sequence encodes a B-domain deleted (BDD) FVIII polypeptide, in which all or part of the B domain of FVIII is deleted. In a specific embodiment, the codon-optimized nucleotide sequence encodes a polypeptide or fragment thereof comprising an amino acid sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 17. In one embodiment, the codon-optimized nucleotide sequence encodes a polypeptide having the amino acid sequence of SEQ ID NO: 17, or a fragment thereof.
[0175] In some embodiments, the codon-optimized nucleotide sequence encodes a FVIII polypeptide that includes a signal peptide or a fragment thereof. In other embodiments, the codon-optimized sequence encodes a FVIII polypeptide that lacks a signal peptide. In some embodiments, the signal peptide comprises amino acids 1-19 of SEQ ID NO: 17.
[0176] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence including a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide, wherein the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO:3 or (ii) nucleotides 58-1791 of SEQ ID NO:4, and the N-terminal and C-terminal portions together have FVIII polypeptide activity. In a specific embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO:3. In another embodiment, the first nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 1791 of SEQ ID NO: 4. In other embodiments, the first nucleotide sequence comprises nucleotides 58 to 1791 of SEQ ID NO: 3 or nucleotides 58 to 1791 of SEQ ID NO: 4.
[0177] In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide, wherein the first nucleic acid sequence is selected from the group consisting of (i) nucleotides 1 to 1791 of SEQ ID NO: 3 or (ii) nucleotides 1 to 1791 of SEQ ID NO: 3. ) has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1-1791 of SEQ ID NO:4, and the N-terminal and C-terminal portions together have FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1-1791 of SEQ ID NO:3 or nucleotides 1-1791 of SEQ ID NO:4. In another embodiment, the second nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO:3 or nucleotides 1792-4374 of SEQ ID NO:4. In a specific embodiment, the second nucleotide sequence comprises nucleotides 1792-4374 of SEQ ID NO:3 or nucleotides 1792-4374 of SEQ ID NO:4. In yet another embodiment, the second nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:3 or 1792-4374 of SEQ ID NO:4 (i.e., nucleotides 1792-4374 of SEQ ID NO:3 or 1792-4374 of SEQ ID NO:4 that do not include the nucleotides encoding the B domain or a fragment of the B domain). In a specific embodiment, the second nucleotide sequence comprises nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:3 or 1792 to 2277 and 2320 to 4374 of SEQ ID NO:4 (i.e., nucleotides 1792 to 4374 of SEQ ID NO:3 or 1792 to 4374 of SEQ ID NO:4 that do not include the nucleotides encoding the B domain or a fragment of the B domain).
[0178] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence including a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide, wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO:5 or (ii) nucleotides 1792-4374 of SEQ ID NO:6, and the N-terminal and C-terminal portions together have FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO:5. In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-4374 of SEQ ID NO: 6. In a particular embodiment, the second nucleic acid sequence comprises nucleotides 1792-4374 of SEQ ID NO: 5 or 1792-4374 of SEQ ID NO: 6. In some embodiments, the first nucleic acid sequence linked to the above-listed second nucleic acid sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 5 or nucleotides 58-1791 of SEQ ID NO: 6. In other embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1-1791 of SEQ ID NO:5 or nucleotides 1-1791 of SEQ ID NO:6.
[0179] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a FVIII polypeptide. and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide, wherein the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1792-4374 of SEQ ID NO:5 without the nucleotides encoding the B domain or a fragment of the B domain) or (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:6 without the nucleotides encoding the B domain or a fragment of the B domain), and the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In certain embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1792-4374 of SEQ ID NO:5 without the nucleotides encoding the B domain or a fragment of the B domain). In other embodiments, the second nucleic acid sequence has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:6 without the nucleotides encoding the B domain or a fragment of the B domain). In one specific embodiment, the second nucleic acid sequence comprises nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 or 1792-2277 and 2320-4374 of SEQ ID NO:6 (i.e., nucleotides 1792-4374 of SEQ ID NO:5 and 1792-4374 of SEQ ID NO:6 that do not include nucleotides encoding the B domain or a fragment of the B domain).In some embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-1791 of SEQ ID NO: 5 or nucleotides 58-1791 of SEQ ID NO: 6. In other embodiments, the first nucleic acid sequence linked to the second nucleic acid sequence listed above has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 1-1791 of SEQ ID NO: 5 or nucleotides 1-1791 of SEQ ID NO: 6.
[0180] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide, wherein the first nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO:1, (ii) nucleotides 58-1791 of SEQ ID NO:2, (iii) nucleotides 58-1791 of SEQ ID NO:70, or (iv) nucleotides 58-1791 of SEQ ID NO:71, and the N-terminal and C-terminal portions together have FVIII polypeptide activity. In other embodiments, the first nucleotide sequence comprises nucleotides 58-1791 of SEQ ID NO:1, nucleotides 58-1791 of SEQ ID NO:2, (iii) nucleotides 58-1791 of SEQ ID NO:70, or (iv) nucleotides 58-1791 of SEQ ID NO:71.
[0181] In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide, wherein the first nucleic acid sequence is selected from the group consisting of: (i) nucleotides 1 to 1791 of SEQ ID NO: 1; (ii) nucleotides 1 to 1791 of SEQ ID NO: 1; and (iv) nucleotides 1-1791 of SEQ ID NO: 71, and wherein the N-terminal and C-terminal portions together have FVIII polypeptide activity. In one embodiment, the first nucleotide sequence comprises nucleotides 1-1791 of SEQ ID NO: 1, nucleotides 1-1791 of SEQ ID NO: 2, (iii) nucleotides 1-1791 of SEQ ID NO: 70, or (iv) nucleotides 1-1791 of SEQ ID NO: 71. In another embodiment, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO: 1, (ii) 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In a particular embodiment, the second nucleotide sequence linked to the first nucleotide sequence comprises (i) nucleotides 1792-4374 of SEQ ID NO: 1, (ii) 1792-4374 of SEQ ID NO: 2, (iii) nucleotides 1792-4374 of SEQ ID NO: 70, or (iv) nucleotides 1792-4374 of SEQ ID NO: 71. In other embodiments, the second nucleotide sequence linked to the first nucleotide sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:71.In one embodiment, the second nucleotide sequence comprises (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:71.
[0182] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide, wherein the second nucleic acid sequence has at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to (i) nucleotides 1792-4374 of SEQ ID NO:1, (ii) nucleotides 1792-4374 of SEQ ID NO:2, (iii) nucleotides 1792-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-4374 of SEQ ID NO:71, and the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In a particular embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792 to 4374 of SEQ ID NO:1, (ii) nucleotides 1792 to 4374 of SEQ ID NO:2, (iii) nucleotides 1792 to 4374 of SEQ ID NO:70, or (iv) nucleotides 1792 to 4374 of SEQ ID NO:71. In some embodiments, the codon-optimized sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide, wherein the second nucleic acid sequence is selected from (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:71 (i.e., a SEQ ID NO:1 that does not include nucleotides encoding the B domain or a fragment of the B domain). and a C-terminal portion of the FVIII polypeptide having at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with (nucleotides 1792 to 4374 of SEQ ID NO:1, nucleotides 1792 to 4374 of SEQ ID NO:2, nucleotides 1792 to 4374 of SEQ ID NO:70, or nucleotides 1792 to 4374 of SEQ ID NO:71), and the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity. In one embodiment, the second nucleic acid sequence comprises (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:1, (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:2, (iii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:70, or (iv) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:71 (i.e., nucleotides 1792-4374 of SEQ ID NO:1, nucleotides 1792-4374 of SEQ ID NO:2, nucleotides 1792-4374 of SEQ ID NO:70, or nucleotides 1792-4374 of SEQ ID NO:71, which do not include the nucleotides encoding the B domain or a fragment of the B domain).
[0183] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 4374 of SEQ ID NO: 1. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58 to 2277 and 2320 to 4374 of SEQ ID NO: 1 (i.e., nucleotides 58 to 4374 of SEQ ID NO: 1 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 58-4374 of SEQ ID NO: 1 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 58-4374 of SEQ ID NO: 1. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 1 (i.e., nucleotides 1-4374 of SEQ ID NO: 1 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 1-4374 of SEQ ID NO: 1.
[0184] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-4374 of SEQ ID NO: 2. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 2. In other embodiments, the nucleic acid sequence has at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 2. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:2 (i.e., nucleotides 58-4374 of SEQ ID NO:2 without the nucleotides encoding the B domain or B domain fragment) or nucleotides 58 through 4374 of SEQ ID NO:2. 2 (i.e., nucleotides 1 to 4374 of SEQ ID NO:2 not including the nucleotides encoding the B domain or B domain fragment) or nucleotides 1 to 4374 of SEQ ID NO:2.
[0185] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-4374 of SEQ ID NO: 70. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 58-4374 of SEQ ID NO: 70 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 70. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 58-4374 of SEQ ID NO: 70 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 58-4374 of SEQ ID NO: 70. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 70 (i.e., nucleotides 1-4374 of SEQ ID NO: 70 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 1-4374 of SEQ ID NO: 70.
[0186] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-4374 of SEQ ID NO: 71. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 71. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 58-4374 of SEQ ID NO: 71. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 71 (i.e., nucleotides 1-4374 of SEQ ID NO: 71 without the nucleotides encoding the B domain or B domain fragment), or SEQ ID NO: 71. 1, including nucleotides 1 to 4374 of nucleotide 1.
[0187] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-4374 of SEQ ID NO: 3. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 58-4374 of SEQ ID NO: 3 without the nucleotides encoding the B domain or B domain fragment). In certain embodiments, the nucleic acid sequence has at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 3. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 58-4374 of SEQ ID NO: 3, not including the nucleotides encoding the B domain or B domain fragment), or nucleotides 58-4374 of SEQ ID NO: 3. In yet other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 3 (i.e., nucleotides 1-4374 of SEQ ID NO: 3, not including the nucleotides encoding the B domain or B domain fragment), or nucleotides 1-4374 of SEQ ID NO: 3.
[0188] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-4374 of SEQ ID NO: 4. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 58-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment). In other embodiments, the nucleic acid sequence has at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 4. In other embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 58-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 58-4374 of SEQ ID NO: 4. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 4 (i.e., nucleotides 1-4374 of SEQ ID NO: 4 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 1-4374 of SEQ ID NO: 4.
[0189] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence is at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 100% identical to nucleotides 58 to 4374 of SEQ ID NO:5. %, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO:5. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 58-4374 of SEQ ID NO:5 not including the nucleotides encoding the B domain or B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:5. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 58-4374 of SEQ ID NO:5 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 58-4374 of SEQ ID NO:5. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO:5 (i.e., nucleotides 1-4374 of SEQ ID NO:5 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 1-4374 of SEQ ID NO:5.
[0190] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-4374 of SEQ ID NO: 6. In other embodiments, the nucleotide sequence comprises a nucleic acid sequence having at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 58-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment). In certain embodiments, the nucleic acid sequence has at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 6. In some embodiments, the nucleotide sequence comprises nucleotides 58-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 58-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 58-4374 of SEQ ID NO: 6. In yet other embodiments, the nucleotide sequence comprises nucleotides 1-2277 and 2320-4374 of SEQ ID NO: 6 (i.e., nucleotides 1-4374 of SEQ ID NO: 6 without the nucleotides encoding the B domain or B domain fragment), or nucleotides 1-4374 of SEQ ID NO: 6.
[0191] In some embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide comprises a nucleic acid sequence encoding a signal peptide. In certain embodiments, the signal peptide is a FVIII signal peptide. In some embodiments, the nucleic acid sequence encoding the signal peptide is codon-optimized. In a specific embodiment, the nucleic acid sequence encoding the signal peptide is selected from the group consisting of: (i) nucleotides 1 to 57 of SEQ ID NO:1; (ii) nucleotides 1 to 57 of SEQ ID NO:2; (iii) nucleotides 1 to 57 of SEQ ID NO:3; (iv) nucleotides 1 to 57 of SEQ ID NO:4; (v) nucleotides 1 to 57 of SEQ ID NO:5; (vi) nucleotides 1 to 57 of SEQ ID NO:6; (vii) nucleotides 1 to 57 of SEQ ID NO:70; and (viii) nucleotides 1 to 57 of SEQ ID NO:71. or (ix) has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to nucleotides 1 to 57 of SEQ ID NO:68.
[0192] SEQ ID NOS: 1-6, 70, and 71 are optimized versions of SEQ ID NOS: 16, the starting or "parent" or "wild-type" FVIII nucleotide sequence. SEQ ID NOS: 16 encodes B-domain deleted human FVIII. While SEQ ID NOS: 1-6, 70, and 71 are derived from a specific B-domain deleted form of FVIII (SEQ ID NOS: 16), it should be understood that the present disclosure also includes optimized versions of nucleic acids encoding other versions of FVIII. For example, other versions of FVIII can include full-length FVIII, other B-domain deletions of FVIII (as described herein), or other fragments of FVIII that retain FVIII activity.
[0193] In one embodiment, the gene cassette comprises a FVIII construct comprising a polynucleotide sequence (6526 nucleotides) listed in Tables 2A-2C. In one embodiment, the gene cassette comprises a FVIII construct comprising a polynucleotide sequence (6526 nucleotides) listed in Table 2B.
[0194] In certain embodiments, the isolated nucleic acid molecule comprises a nucleotide sequence having at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to the nucleotide sequence of SEQ ID NO: 179 or 182, and retains the ability to express a functional FVIII protein.
[0195] [Table 3] [Table 4] [Table 5]
[0196] [Table 6] [Table 7] [Table 8] [Table 9]
[0197] [Table 10] [Table 11] [Table 12] [Table 13]
[0198] A. Codon Adaptation Index In one embodiment, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the human codon adaptation index of the codon-optimized nucleotide sequence is increased relative to SEQ ID NO: 16. For example, the codon-optimized nucleotide sequence may have a human codon adaptation index of at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 1.00 (1.01%), at least about 1.01 (1.02%), at least about 1.02 (1.03%), at least about 1.03 (1.04%), at least about 1.04 (1.05%), at least about 1.05 (1.06%), at least about 1.06 (1.07%), at least about 1.07 (1.08%), at least about 1.08 (1.09%), at least about 1.09 (2.10%), at least about 1.10 (2.11%), at least about 1.11 (2.12%), at least about 1.12 (2.13%), at least about 1.13 (2.14%), at least about 1.14 (2.15%), at least about 1.15 (2.16 In some embodiments, the human codon adaptation index may be at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%), at least about 0.91 (91%), at least about 0.92 (92%), at least about 0.93 (93%), at least about 0.94 (94%), at least about 0.95 (95%), at least about 0.96 (96%), at least about 0.97 (97%), at least about 0.98 (98%), or at least about 0.99 (99%). In some embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index that is at least about 0.88 (88%). In other embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index that is at least about 0.91 (91%). In other embodiments, the codon-optimized nucleotide sequence has a human codon adaptation index that is at least about 0.97 (97%).
[0199] In one particular embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the first nucleic acid sequence is at least about identical to (i) nucleotides 58 to 1791 of SEQ ID NO:3; (ii) nucleotides 1 to 1791 of SEQ ID NO:3; (iii) nucleotides 58 to 1791 of SEQ ID NO:4; or (iv) nucleotides 1 to 1791 of SEQ ID NO:4. having 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO:16. In some embodiments, the nucleotide sequence has a human codon adaptation index that is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), at least about 0.88 (88%), at least about 0.89 (89%), at least about 0.90 (90%), or at least about 0.91 (91%). In a particular embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.88 (88%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.91 (91%).
[0200] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence comprising a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the second nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:5 or (ii) nucleotides 1792-2277 and 2320-4374 of SEQ ID NO:6; the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO:16. In some embodiments, the nucleotide sequence has a human codon adaptation index that is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), or at least about 0.88 (88%). In one particular embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.83 (83%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.88 ( It has a human codon adaptation index of 88%).
[0201] In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a nucleotide sequence encoding a polypeptide having FVIII activity, wherein the nucleotide sequence has at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to nucleotides 58-2277 and 2320-4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 that do not contain nucleotides encoding the B domain or B domain fragment); and the human codon adaptation index of the nucleotide sequence is increased relative to SEQ ID NO: 16. In some embodiments, the nucleotide sequence has a human codon adaptation index that is at least about 0.75 (75%), at least about 0.76 (76%), at least about 0.77 (77%), at least about 0.78 (78%), at least about 0.79 (79%), at least about 0.80 (80%), at least about 0.81 (81%), at least about 0.82 (82%), at least about 0.83 (83%), at least about 0.84 (84%), at least about 0.85 (85%), at least about 0.86 (86%), at least about 0.87 (87%), or at least about 0.88 (88%). In a particular embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.75 (75%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.83 (83%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.88 (88%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.91 (91%). In another embodiment, the nucleotide sequence has a human codon adaptation index that is at least about 0.97 (97%).
[0202] In some embodiments, a codon-optimized nucleotide sequence encoding a FVIII polypeptide of the present disclosure has an increased frequency of optimal codon (FOP) relative to SEQ ID NO: 16. In certain embodiments, the FOP of a codon-optimized nucleotide sequence encoding a FVIII polypeptide is at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 64, at least about 65, at least about 70, at least about 75, at least about 79, at least about 80, at least about 85, or at least about 90.
[0203] In other embodiments, a codon-optimized nucleotide sequence encoding a FVIII polypeptide of the present disclosure has an increased relative synonymous codon usage (RCSU) relative to SEQ ID NO: 16. In some embodiments, the RCSU of the isolated nucleic acid molecule is greater than 1.5. In other embodiments, the RCSU of the isolated nucleic acid molecule is greater than 2.0. In certain embodiments, the RCSU of the isolated nucleic acid molecule is at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2.0, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, or at least about 2.7.
[0204] In still other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide of the present disclosure has a reduced effective number of codons relative to SEQ ID NO: 16. In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has fewer than about 50, fewer than about 45, fewer than about 40, or fewer than about 35 codons. , fewer than about 30, or fewer than about 25. In a particular embodiment, the isolated nucleic acid molecule has an effective number of codons of about 40, about 35, about 30, about 25, or about 20.
[0205] Optimization of BG / C content In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, or at least about 60%.
[0206] In one particular embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the first nucleic acid sequence is at least about 80%, at least about 85%, at least about 90%, or at least about 95% identical to (i) nucleotides 58 to 1791 of SEQ ID NO:3; (ii) nucleotides 1 to 1791 of SEQ ID NO:3; (iii) nucleotides 58 to 1791 of SEQ ID NO:4; or (iv) nucleotides 1 to 1791 of SEQ ID NO:4. the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides in SEQ ID NO: 16. In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, or at least about 58%. In one particular embodiment, the nucleotide sequence encoding the polypeptide having FVIII activity has a G / C content that is at least about 58%.
[0207] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the second nucleic acid sequence is selected from the group consisting of (i) nucleotides 1792 to 4374 of SEQ ID NO:5; (ii) nucleotides 1792 to 4374 of SEQ ID NO:6; (iii) nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:5 (i.e., nucleotides 1792 to 4374 of SEQ ID NO:5 not including the nucleotides encoding the B domain or B domain fragment), or (iv) nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:6. nucleotides 1792-4374 of SEQ ID NO: 6, excluding nucleotides encoding the B domain or B domain fragment); the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with nucleotides 1792-4374 of SEQ ID NO: 6, excluding nucleotides encoding the B domain or B domain fragment; The codon-optimized nucleotide sequence encoding a FVIII polypeptide contains a higher percentage of G / C nucleotides compared to the percentage of C nucleotides. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 45%, at least about 46%, at least about 47%, at least about 48%, at least about 49%, at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, or at least about 57%. In a specific embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 52%. In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 55%. In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide has a G / C content that is at least about 57%.
[0208] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58 to 4374 or (ii) nucleotides 58 to 2277 and 2320 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58 to 4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 that do not contain nucleotides encoding the B domain or B domain fragment); and the nucleotide sequence contains a higher percentage of G / C nucleotides compared to the percentage of G / C nucleotides of SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content that is at least about 45%. In one particular embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content that is at least about 52%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content that is at least about 55%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content that is at least about 57%. In another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content that is at least about 58%. In yet another embodiment, the codon-optimized nucleotide sequence encoding the FVIII polypeptide has a G / C content that is at least about 60%.
[0209] "G / C content" (or guanine-cytosine content), or "percentage of G / C nucleotides," refers to the percentage of nitrogenous bases in a DNA molecule that are either guanine or cytosine. G / C content is calculated using the following formula:
number
[0210] The G / C content of human genes is highly heterogeneous, with some genes having a low G / C content of around 20% and others having a high G / C content of around 95%. Generally, G / C-rich genes are more highly expressed. In fact, it has been demonstrated that an increase in the G / C content of a gene can lead to increased gene expression, which is mostly due to increased transcription and a much more stable mRNA level. Kudla et al., PLoS Bio l., 4(6):e180 (2006).
[0211] C. Matrix-binding domain-like sequence In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 6, up to 5, up to 4, up to 3, or up to 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains no MARS / ARS sequences.
[0212] In one particular embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the first nucleic acid sequence has at least about 80%, at least about 80%, or at least about 80% identity with (i) nucleotides 58 to 1791 of SEQ ID NO:3; (ii) nucleotides 1 to 1791 of SEQ ID NO:3; (iii) nucleotides 58 to 1791 of SEQ ID NO:4; or (iv) nucleotides 1 to 1791 of SEQ ID NO:4. 5%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences than SEQ ID NO: 16. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains up to 6, up to 5, up to 4, up to 3, or up to 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to one MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a MARS / ARS sequence.
[0213] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the second nucleic acid sequence is selected from the group consisting of: (i) nucleotides 1792 to 4374 of SEQ ID NO:5; (ii) nucleotides 1792 to 4374 of SEQ ID NO:6; (iii) nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:5 (i.e., nucleotides 1792 to 4374 of SEQ ID NO:5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:6. 74 (i.e., nucleotides 1792-4374 of SEQ ID NO: 6, not including the nucleotides encoding the B domain or B domain fragment); the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 6, up to 5, up to 4, up to 3, or up to 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 6, up to 5, up to 4, up to 3, or up to 2 MARS / ARS sequences. The codon-optimized nucleotide sequence encoding the peptide contains at most one MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a MARS / ARS sequence.
[0214] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence is (i) nucleotides 58 to 4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71, or (ii) nucleotides 58 to 2277 and 2320 to 4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 (i.e., SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71, excluding the nucleotides encoding the B domain or B domain fragment). The codon-optimized nucleotide sequence encoding a FVIII polypeptide includes a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 16 (nucleotides 58-4374 of SEQ ID NO: 6, 70, or 71); the codon-optimized nucleotide sequence contains fewer MARS / ARS sequences relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 6, up to 5, up to 4, up to 3, or up to 2 MARS / ARS sequences. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 1 MARS / ARS sequence. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a MARS / ARS sequence.
[0215] AT-rich elements in the human FVIII nucleotide sequence have been identified that share sequence similarity with the autonomously replicating sequence (ARS) and nuclear-matrix attachment region (MAR) of Saccharomyces cerevisiae (Fallux et al., Mol. Cell. Biol. 16:4264-4272 (1996)). One of these elements has been demonstrated to bind to a nuclear factor in vitro and repress expression of a chloramphenicol acetyltransferase (CAT) reporter gene. Ibid. It is hypothesized that these sequences may be responsible for transcriptional repression of the human FVIII gene. Thus, in one embodiment, all MAR / ARS sequences are eliminated in the codon-optimized nucleotide sequence encoding the FVIII polypeptide of the present disclosure. In the parent FVIII sequence (SEQ ID NO: 16), four MAR / ARS ATATTT sequences (SEQ ID NO: 21) and three MAR / ARS AAATAT sequences (SEQ ID NO: 22) are present. All of these sites were mutated to disrupt the MAR / ARS sequences of the optimized FVIII sequence (SEQ ID NOs: 1-6). The location of each of these elements and the corresponding nucleotide sequence in the optimized sequence are shown in Table 3 below.
[0216] [Table 14] [Table 15]
[0217] D. Destabilizing Sequences In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 9, up to 8, up to 7, up to 6, or up to 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 4, up to 3, up to 2, or up to 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains no destabilizing elements.
[0218] In one particular embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the first nucleic acid sequence has at least about 80%, at least about 80%, or at least about 80% identity with (i) nucleotides 58 to 1791 of SEQ ID NO:3; (ii) nucleotides 1 to 1791 of SEQ ID NO:3; (iii) nucleotides 58 to 1791 of SEQ ID NO:4; or (iv) nucleotides 1 to 1791 of SEQ ID NO:4. 5%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence contains fewer destabilizing elements than SEQ ID NO: 16. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains up to 9, up to 8, up to 7, up to 6, or up to 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 4, up to 3, up to 2, or up to 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain a destabilizing element.
[0219] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the second nucleic acid sequence is selected from the group consisting of: (i) nucleotides 1792 to 4374 of SEQ ID NO:5; (ii) nucleotides 1792 to 4374 of SEQ ID NO:6; (iii) nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:5 (i.e., nucleotides 1792 to 4374 of SEQ ID NO:5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:6 (i.e., nucleotides 1792 to 4374 of SEQ ID NO:5 without the nucleotides encoding the B domain or B domain fragment). i.e., at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 6 (nucleotides 1792-4374), not including nucleotides encoding the B domain or B domain fragment; the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the nucleotide sequence encoding a polypeptide having FVIII activity contains up to 9, up to 8, up to 7, up to 6, or up to 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 4, up to 3, or up to 5 destabilizing elements. , up to 2, or up to 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain a destabilizing element.
[0220] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence is (i) nucleotides 58 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, or (ii) nucleotides 58 to 2277 and 2320 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., a sequence that does not include nucleotides encoding a B domain or a B domain fragment). The codon-optimized nucleotide sequence comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 (nucleotides 58 to 4374); the codon-optimized nucleotide sequence contains fewer destabilizing elements relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 9, up to 8, up to 7, up to 6, or up to 5 destabilizing elements. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 4, up to 3, up to 2, or up to 1 destabilizing element. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains no destabilizing elements.
[0221] There are ten destabilizing elements in the parent FVIII sequence (SEQ ID NO: 16): six ATTTA sequences (SEQ ID NO: 23) and four TAAAT sequences (SEQ ID NO: 24). In one embodiment, the sequences at these sites were mutated to disrupt the destabilizing elements in optimized FVIII SEQ ID NOs: 1-6, 70, and 71. The location of each of these elements and the corresponding nucleotide sequence in the optimized sequence are shown in Table 3.
[0222] E. Potential promoter binding sites In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 9, up to 8, up to 7, up to 6, or up to 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 4, up to 3, up to 2, or up to 1 potential promoter binding site. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains no potential promoter binding sites.
[0223] In one particular embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97% similarity to (i) nucleotides 58 to 1791 of SEQ ID NO:3; (ii) nucleotides 1 to 1791 of SEQ ID NO:3; (iii) nucleotides 58 to 1791 of SEQ ID NO:4; or (iv) nucleotides 1 to 1791 of SEQ ID NO:4. , at least about 98%, or at least about 99% sequence identity; the N-terminal and C-terminal portions together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 9, up to 8, up to 7, up to 6, or up to 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 4, up to 3, up to 2, or up to 1 potential promoter binding site. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide does not contain any potential promoter binding sites.
[0224] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the second nucleic acid sequence is selected from the group consisting of (i) nucleotides 1792 to 4374 of SEQ ID NO:5; (ii) nucleotides 1792 to 4374 of SEQ ID NO:6; (iii) nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:5 (i.e., nucleotides 1792 to 4374 of SEQ ID NO:5 without the nucleotides encoding the B domain or B domain fragment); or (iv) nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:6 (i.e., nucleotides 1792 to 4374 of SEQ ID NO:5 without the nucleotides encoding the B domain or B domain fragment). the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 9, up to 8, up to 7, up to 6, or up to 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 4, up to 3, up to 2, or up to 1 potential promoter binding site. In still other embodiments, the codon-optimized nucleotide sequence encoding the FVIII polypeptide does not contain any potential promoter binding sites.
[0225] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, the nucleotide sequence being (i) nucleotides 58 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, or (ii) nucleotides 58 to 2277 and 2320 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, which do not include nucleotides encoding the B domain or B domain fragment). 16, 3, 4, 5, 6, 70, or 71 of nucleotides 58-4374; the codon-optimized nucleotide sequence contains fewer potential promoter binding sites relative to SEQ ID NO: 16. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 9, up to 8, up to 7, up to 6, or up to 5 potential promoter binding sites. In other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 4, up to 3, up to 2, or up to 1 potential promoter binding site. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 1, up to 2, up to 3, up to 4, or up to 1 potential promoter binding site. In yet other embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide contains up to 1, up to 2, up to 3, up to 4, or up to 1 potential promoter binding site. The selected nucleotide sequence does not contain any potential promoter binding sites.
[0226] TATA boxes are regulatory sequences often found in eukaryotic promoter regions. They serve as binding sites for the general transcription factor TATA-binding protein (TBP). TATA boxes usually contain the sequence TATAA (SEQ ID NO: 28) or a close variant thereof. However, a TATA box within a coding sequence can inhibit translation of the full-length protein. The wild-type BDD FVIII sequence (SEQ ID NO: 16) contains 10 potential promoter binding sites: five TATAA sequences (SEQ ID NO: 28) and five TTATA sequences (SEQ ID NO: 29). In some embodiments, at least one, at least two, at least three, or at least four promoter binding sites are eliminated in the FVIII gene of the present disclosure. In some embodiments, at least five promoter binding sites are eliminated in the FVIII gene of the present disclosure. In other embodiments, at least six, at least seven, or at least eight promoter binding sites are eliminated in the FVIII gene of the present disclosure. In one embodiment, at least nine promoter binding sites are eliminated in the FVIII gene of the present disclosure. In one particular embodiment, all promoter binding sites are eliminated in the FVIII gene of the present disclosure. The location of each potential promoter binding site and the corresponding nucleotide sequence in the optimized sequence are shown in Table 3.
[0227] F. Other Cis-Acting Negative Regulatory Regions In addition to the MAR / ARS sequences, destabilizing elements, and potential promoter sites described above, several additional potential inhibitory sequences can be identified in the wild-type BDD FVIII sequence (SEQ ID NO: 16). Two AU-rich sequence elements (AREs) can be identified (ATTTTATT (SEQ ID NO: 30); and ATTTTTAA (SEQ ID NO: 31)), along with a polyA site (AAAAAAA; SEQ ID NO: 26), a polyT site (TTTTTT; SEQ ID NO: 25), and a splice site (GGTGAT; SEQ ID NO: 27) in the non-optimized BDD FVIII sequence. One or more of these elements can be removed from the optimized FVIII sequence. The location of each of these sites and the corresponding nucleotide sequence in the optimized sequence are shown in Table 3.
[0228] In certain embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the first nucleic acid sequence is at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99% identical to (i) nucleotides 58 to 1791 of SEQ ID NO:3; (ii) nucleotides 1 to 1791 of SEQ ID NO:3; (iii) nucleotides 58 to 1791 of SEQ ID NO:4; or (iv) nucleotides 1 to 1791 of SEQ ID NO:4. the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, such as a splice site, a poly-T sequence, a poly-A sequence, an ARE sequence, or any combination thereof.
[0229] In another embodiment, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the second nucleic acid sequence comprises (i) nucleotides 1792 to 4374 of SEQ ID NO:5; (ii) nucleotides 1792 to 4374 of SEQ ID NO:6; (iii) nucleotides 1792 to 4374 of SEQ ID NO:5 (i.e., nucleotides 1792 to 4374 of SEQ ID NO:5 without any nucleotides encoding the B domain or a B domain fragment); or (iv) nucleotides 1792 to 2277 and 2320 to 4374 of SEQ ID NO:6 (i.e., nucleotides 1792 to 4374 of SEQ ID NO:6 without any nucleotides encoding the B domain or a B domain fragment), which are at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 1109%, at least about 1111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at having about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, such as a splice site, a poly-T sequence, a poly-A sequence, an ARE sequence, or any combination thereof.
[0230] In other embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, the nucleotide sequence being (i) nucleotides 58 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, or (ii) nucleotides 58 to 2277 and 2320 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides 58 to 2277 and 2320 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, or 71 that do not include nucleotides encoding a B domain or a B domain fragment). codon-optimized nucleotide sequences include nucleic acid sequences having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to a codon-optimized nucleotide sequence (SEQ ID NO: 58-4374); the codon-optimized nucleotide sequence does not contain one or more cis-acting negative regulatory elements, e.g., a splice site, a poly-T sequence, a poly-A sequence, an ARE sequence, or any combination thereof.
[0231] In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the first nucleic acid sequence is at least about 80%, at least about 100%, identical to (i) nucleotides 58 to 1791 of SEQ ID NO:3; (ii) nucleotides 1 to 1791 of SEQ ID NO:3; (iii) nucleotides 58 to 1791 of SEQ ID NO:4; or (iv) nucleotides 1 to 1791 of SEQ ID NO:4. having 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence does not contain the splice site GGTGAT (SEQ ID NO: 27). In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of the FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of the FVIII polypeptide; the first nucleic acid sequence has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to (i) nucleotides 58-1791 of SEQ ID NO:3; (ii) nucleotides 1-1791 of SEQ ID NO:3; (iii) nucleotides 58-1791 of SEQ ID NO:4; or (iv) nucleotides 1-1791 of SEQ ID NO:4; In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the first nucleic acid sequence has at least about 80%, at least about 100%, a sequence identical to (i) nucleotides 58 to 1791 of SEQ ID NO:3; (ii) nucleotides 1 to 1791 of SEQ ID NO:3; (iii) nucleotides 58 to 1791 of SEQ ID NO:4; or (iv) nucleotides 1 to 1791 of SEQ ID NO:4. the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence does not contain a polyA sequence (SEQ ID NO: 26). In some embodiments, the codon-optimized nucleotide sequence encoding a FVIII polypeptide comprises a first nucleic acid sequence encoding an N-terminal portion of a FVIII polypeptide and a second nucleic acid sequence encoding a C-terminal portion of a FVIII polypeptide; the first nucleic acid sequence is at least about 80%, at least about 85%, identical to (i) nucleotides 58 to 1791 of SEQ ID NO:3; (ii) nucleotides 1 to 1791 of SEQ ID NO:3; (iii) nucleotides 58 to 1791 of SEQ ID NO:4; or (iv) nucleotides 1 to 1791 of SEQ ID NO:4. %, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; the N-terminal portion and the C-terminal portion together have FVIII polypeptide activity; and the codon-optimized nucleotide sequence does not contain an ARE element (SEQ ID NO: 30 or SEQ ID NO: 31).
[0232] In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence is selected from (i) nucleotides 58 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, or (ii) nucleotides 58 to 2277 and 2320 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides encoding a B domain or a B domain fragment). The codon-optimized nucleotide sequence includes a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71; the codon-optimized nucleotide sequence does not contain the splice site GGTGAT (SEQ ID NO: 27). In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence is selected from (i) nucleotides 58 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, or (ii) nucleotides 58 to 2277 and 2320 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., nucleotides encoding a B domain or a B domain fragment). In some embodiments, the gene cassette comprises a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71 (nucleotides 58-4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71, not including the sequence (SEQ ID NO: 25); the codon-optimized nucleotide sequence does not contain a poly-T sequence (SEQ ID NO: 25). comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, the codon-optimized nucleotide sequence being selected from (i) nucleotides 58 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, or (ii) nucleotides 58 to 2277 and 2320 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, which do not include nucleotides encoding the B domain or B domain fragment). 1, 2, 3, 4, 5, 6, 70, or 71 of nucleotides 58-4374) has at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity; the codon-optimized nucleotide sequence does not contain a polyA sequence (SEQ ID NO:26). In some embodiments, the gene cassette comprises a codon-optimized nucleotide sequence encoding a FVIII polypeptide, wherein the codon-optimized nucleotide sequence is (i) nucleotides 58 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71, or (ii) nucleotides 58 to 2277 and 2320 to 4374 of an amino acid sequence selected from SEQ ID NOs: 1, 2, 3, 4, 5, 6, 70, and 71 (i.e., a sequence that does not include nucleotides encoding a B domain or a B domain fragment). a nucleic acid sequence having at least about 80%, at least about 85%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to a codon-optimized nucleotide sequence (nucleotides 58 to 4374 of SEQ ID NO: 1, 2, 3, 4, 5, 6, 70, or 71); wherein the codon-optimized nucleotide sequence does not contain an ARE element (SEQ ID NO: 30 or SEQ ID NO: 31).
[0233] In other embodiments, the optimized FVIII sequences of the present disclosure do not include one or more antiviral motifs, stem-loop structures, and repeat sequences.
[0234] In yet other embodiments, the nucleotides surrounding the transcription start site include the Kozak consensus sequence (GCCGCCACC ATG C (SEQ ID NO: 32), in which the underlined nucleotide is the start codon.) In other embodiments, restriction sites are added or removed to facilitate the cloning process.
[0235] b. FIX and polynucleotide sequences encoding the FIX protein In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a FIX polypeptide. In some embodiments, the FIX polypeptide comprises FIX or a variant or fragment thereof, wherein the FIX or variant or fragment thereof has FIX activity.
[0236] Human FIX is a serine protease that is a key component of the intrinsic pathway of the blood coagulation cascade. "Factor IX" or "FIX," as used herein, refers to coagulation factor proteins and species and sequence variants thereof, including, but not limited to, the 461 single-chain amino acid sequence of the human FIX precursor polypeptide ("prepro"), the 415 single-chain amino acid sequence of mature human FIX (SEQ ID NO: 125), and the R338L FIX (Padua) variant (SEQ ID NO: 126). FIX includes any form of FIX molecule that possesses the typical characteristics of blood coagulation FIX. As used herein, "Factor IX" and "FIX" refer to the G1A domain (a region containing a gamma-carboxyglutamic acid residue), EGF1 and EGF2 (regions containing sequences homologous to human epidermal growth factor), the activation peptide ("AP" formed by residues R136-R180 of mature FIX), and the C-terminal protease domain ("Pro"), or any combination of these domains known in the art. The term "FIX" is intended to encompass polypeptides containing synonyms of the native protein, or may be truncated fragments or sequence variants that retain at least some of the biological activity of the native protein. FIX or sequence variants have been cloned as described in U.S. Patent Nos. 4,770,999 and 7,700,734, and cDNA encoding human FIX has been isolated, characterized, and cloned into an expression vector (see, e.g., Choo et al., Nature 299:178-180 (1982); Fair et al., Blood 64:194-204 (1984); and Kurachi et al., Proc. Natl. Acad. Sci., USA 79:6461-6464 (1982)). One particular variant of FIX characterized by Simioni et al., 2009, the R338L FIX (Padua) variant (SEQ ID NO: 2), contains a gain-of-function mutation that correlates with an approximately 8-fold increase in activity of the Padua variant relative to native FIX (Table 4). FIX variants can also include any FIX polypeptide with one or more conservative amino acid substitutions that do not affect the FIX activity of the FIX polypeptide. In some embodiments, a FIX variant comprises rFIX albumin fused by a cleavable linker, e.g., IDELVION®. See U.S. Pat. No. 7,939,632, incorporated herein by reference in its entirety.
[0237] [Table 16] [Table 17] [Table 18]
[0238] The FIX polypeptide is 55 kDa and is synthesized as a prepropolypeptide chain (SEQ ID NO: 125) from three regions: a 28-amino acid signal peptide (amino acids 1 to 28 of SEQ ID NO: 127), an 18-amino acid propeptide (amino acids 29 to 46) required for gamma-carboxylation of glutamic acid residues, and the 415-amino acid mature factor IX (SEQ ID NO: 125 or 126). The propeptide is an 18-amino acid sequence N-terminal to the gamma-carboxyglutamate domain. The propeptide binds to vitamin K-dependent gamma-carboxylase and is then cleaved from the FIX precursor polypeptide by an endogenous protease, most likely PACE (paired basic amino acid cleaving enzyme), also known as furin or PCSK3. Without gamma-carboxylation, the Gla domain cannot bind calcium to assume the correct conformation required to anchor the protein to negatively charged phospholipid surfaces, rendering factor IX nonfunctional. Even when carboxylated, the Gla domain also depends on cleavage of the propeptide for proper function because the retained propeptide interferes with the conformational changes in the Gla domain necessary for optimal binding to calcium and phospholipids. In humans, the resulting mature factor IX is secreted into the bloodstream by hepatocytes as an inactive zymogen, a single-chain protein of 415 amino acid residues that contains approximately 17% carbohydrate by weight (Schmidt, AE et al. (2003) Trends Cardiovasc Med 13:39).
[0239] Mature FIX is composed of several domains, arranged from N to C termini: the GLA domain, EGF1 domain, EGF2 domain, activation peptide (AP) domain, and protease (or catalytic) domain. A short linker connects the EGF2 domain, including the AP domain. FIX contains two activation peptides, formed by R145-A146 and R180-V181, respectively. After activation, single-chain FIX becomes a two-chain molecule, with the two chains linked by a disulfide bond. Coagulation factors can be engineered by replacing their activation peptides, resulting in altered activation specificity. In mammals, mature FIX must be activated by activated factor XI to yield factor IXa. During activation of FIX to FIXa, the protease domain provides the catalytic activity of FIX. Activated factor VIII (FVIIIa) is the specific cofactor for full expression of FIXa activity.
[0240] In certain embodiments, the FIX polypeptide comprises the Thr148 allele of plasma-derived FIX and has structural and functional characteristics similar to endogenous FIX.
[0241] Many functional FIX variants are known in the art. International Patent Application Publication No. WO 02 / 040544 discloses, on page 4, lines 9-30 and on page 15, lines 6-31, mutants that exhibit increased resistance to inhibition by heparin. International Patent Application Publication No. WO 03 / 020764 discloses, in Tables 2 and 3 (pages 14-24) and on page 12, lines 1-27, FIX mutants with reduced T-cell immunogenicity. International Patent Application Publication No. WO 2007 / 149406 discloses, on page 4, line 1 to page 19, line 11, FIX mutants with increased protein stability, increased in vivo and in vitro half-life, and resistance to proteases. WO 2007 / 149406 also discloses chimeric and other variant FIX molecules on page 19, line 12 to page 20, line 9. International Patent Application Publication No. WO 08 / 118507 discloses FIX mutants on page 5, line 14 to page 6, line 5 that exhibit increased clotting activity. International Patent Application Publication No. WO 09 / 051717 discloses FIX mutants with an increased number of N-linked and / or O-linked glycosylation sites that result in increased and / or restored half-life on page 9, line 11 to page 20, line 2. International Patent Application Publication No. WO 09 / 137254 discloses factor IX mutants with an increased number of glycosylation sites on page 2, paragraphs
[0006] to
[0011] , and pages 16,
[0044] to
[0057] . International Patent Application Publication No. WO 09 / 130198 discloses functional mutant FIX molecules with an increased number of glycosylation sites, resulting in an increased half-life, on page 4, line 26 to page 12, line 6. International Patent Application Publication No. WO 09 / 140015 discloses functional FIX mutants with an increased number of Cys residues that can be used for conjugation of polymers (e.g., PEG) on page 11, paragraphs
[0043] to
[0053] . The FIX polypeptide described in International Patent Application No. PCT / US2011 / 043569, filed July 11, 2011, and published January 12, 2012, as WO2012 / 006624, is incorporated herein by reference in its entirety. In some embodiments, the FIX polypeptide comprises a FIX polypeptide fused to albumin, e.g., FIX-albumin. In certain embodiments, the FIX polypeptide is IDELVION® or rIX-FP.
[0242] Additionally, hundreds of non-functional mutations in FIX have been identified in hemophilia patients, many of which are disclosed in Table 6 on pages 11-14 of International Patent Application Publication No. WO 09 / 137254. Such non-functional mutations are not included in the present invention, but provide additional guidance as to which mutations are more or less likely to result in functional FIX polypeptides.
[0243] In one embodiment, the FIX polypeptide (or the Factor IX portion of the fusion polypeptide) comprises an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 1 or 2 (amino acids 1 to 415 of SEQ ID NO: 125 or 126), or alternatively, the propeptide sequence, or the propeptide and signal sequence (full-length FIX). In another embodiment, the FIX polypeptide comprises an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 2.
[0244] FIX clotting activity is expressed as International Units (IU). One unit of FIX activity roughly corresponds to the amount of FIX in one milliliter of normal human plasma. Several assays for measuring FIX activity are available, including one-stage clotting assays (activated partial thromboplastin time; aPTT), thrombin generation time (TGA), and rotational thromboelastometry (ROTEM®). The present invention contemplates FIX sequences, sequences homologous to sequence fragments, and non-naturally occurring sequence variants derived from natural sources, e.g., humans, non-human primates, and mammals (including domestic animals), that retain at least some of the biological activity or function of FIX and / or are useful for preventing, treating, mediating, or ameliorating coagulation factor-related diseases, deficiencies, disorders, or conditions (e.g., bleeding episodes related to trauma, surgery, or coagulation factor deficiencies). Sequences homologous to human FIX can be found by standard homology search techniques, e.g., NCBI BLAST.
[0245] In certain embodiments, the FIX sequence is codon-optimized. Examples of codon-optimized FIX sequences include, but are not limited to, SEQ ID NOS: 1 and 54-58 of International Patent Application Publication No. WO2016 / 004113, which is incorporated herein by reference in its entirety.
[0246] c. Polynucleotide sequences encoding FVII and FVII proteins In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a Factor VII polypeptide. In some embodiments, the FVII polypeptide comprises FVII or a variant or fragment thereof having FVII activity.
[0247] "Factor VII" ("FVII," or "F7"; also known as Factor 7, clotting factor VII, serum factor VII, serum prothrombin converting accelerator, SPCA, proconvertin, and eptacog alfa) is a serine protease that is part of the coagulation cascade. In one embodiment, the coagulation factor in the nucleic acids described herein is FVII. Recombinant activated factor VII ("FVII") has become widely used to treat major bleeding events, such as those that occur in patients with hemophilia A or B, clotting factor XI, FVII deficiency, defective platelet function, thrombocytopenia, or von Willebrand's disease.
[0248] Recombinant activated FVII (rFVIIa; NOVOSEVEN®) is used to treat bleeding episodes in (i) hemophilia patients with neutralizing antibodies to FVIII or FIX (inhibitors), (ii) patients with FVII deficiency, or (iii) patients with hemophilia A or B with inhibitors via surgical procedures. However, NOVOSEVEN® exhibits inadequate efficacy. Due to its low affinity for activated platelets in the absence of tissue factor, its short half-life, and its insufficient enzymatic activity, repeated administration of FVIIa at high concentrations is often required to control bleeding. Therefore, there is an unmet medical need for better treatment and prophylaxis options for hemophilia patients with FVIII and FIX inhibitors and / or FVII deficiency.
[0249] In one embodiment, the gene cassette encodes the mature form of FVII or a variant thereof. FVII contains a Gla domain, two EGF domains (EGF-1 and EGF-2), and a serine protease domain (or peptidase S1 domain) that is highly conserved among all members of the peptidase S1 family of serine proteases, such as chymotrypsin. FVII occurs as a single-chain zymogen (i.e., activatable FVII) and a fully activated two-chain form.
[0250] C. growth factors In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises a growth factor. The growth factor can be selected from any growth factor known in the art. In some embodiments, the growth factor is a hormone. In other embodiments, the growth factor is a cytokine. In some embodiments, the growth factor is a chemokine.
[0251] In some embodiments, the growth factor is adrenomedullin (AM). In some embodiments, the growth factor is angiopoietin (Ang). In some embodiments, the growth factor is an autocrine cell motility stimulating factor. In some embodiments, the growth factor is a bone morphogenetic protein (BMP). In some embodiments, the BMP is selected from BMP2, BMP4, BMP5, and BMP7. In some embodiments, the growth factor is a member of the ciliary neurotrophic factor family. In some embodiments, the ciliary neurotrophic factor family In some embodiments, the growth factor is a colony-stimulating factor. In some embodiments, the colony-stimulating factor is selected from macrophage colony-stimulating factor (m-CSF), granulocyte colony-stimulating factor (G-CSF), and granulocyte-macrophage colony-stimulating factor (GM-CSF). In some embodiments, the growth factor is an epidermal growth factor (EGF). In some embodiments, the growth factor is an ephrin. In some embodiments, the ephrin is selected from ephrinA1, ephrinA2, ephrinA3, ephrinA4, ephrinA5, ephrinB1, ephrinB2, and ephrinB3. In some embodiments, the growth factor is erythropoietin (EPO). In some embodiments, the growth factor is a fibroblast growth factor (FGF). In some embodiments, the FGF is selected from FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, and FGF23. In some embodiments, the growth factor is fetal bovine somatotrophin (FBS). In some embodiments, the growth factor is a member of the GDNF family. In some embodiments, the member of the GDNF family is selected from glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, and artemin. In some embodiments, the growth factor is growth differentiation factor-9 (GDF9). In some embodiments, the growth factor is hepatocyte growth factor (HGF). In some embodiments, the growth factor is hepatoma-derived growth factor (HDGF). In some embodiments, the growth factor is insulin. In some embodiments, the growth factor is an insulin-like growth factor. In some embodiments, the insulin-like growth factor is insulin-like growth factor-1 (IGF-1) or IGF-2. In some embodiments, the growth factor is an interleukin (IL). In some embodiments, the IL is selected from IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, and IL-7.In some embodiments, the growth factor is keratinocyte growth factor (KGF). In some embodiments, the growth factor is migration stimulating factor (MSF). In some embodiments, the growth factor is macrophage stimulating protein (MSP or hepatocyte growth factor-like protein (HGFLP)). In some embodiments, the growth factor is myostatin (GDF-8). In some embodiments, the growth factor is neuregulin. In some embodiments, the neuregulin is selected from neuregulin 1 (NRG1), NRG2, NRG3, and NRG4. In some embodiments, the growth factor is a neurotrophin. In some embodiments, the growth factor is brain-derived neurotrophic factor (BDNF). In some embodiments, the growth factor is nerve growth factor (NGF). In some embodiments, the NGF is neurotrophin-3 (NT-3) or NT-4. In some embodiments, the growth factor is placental growth factor (PGF). In some embodiments, the growth factor is platelet-derived growth factor (PDGF). In some embodiments, the growth factor is renalase (RNLS). In some embodiments, the growth factor is T cell growth factor (TCGF). In some embodiments, the growth factor is thrombopoietin (TPO). In some embodiments, the growth factor is a transforming growth factor. In some embodiments, the transforming growth factor is transforming growth factor-alpha (TGF-α) or TGF-β. In some embodiments, the growth factor is tumor necrosis factor-alpha (TNF-α). In some embodiments, the growth factor is vascular endothelial growth factor (VEGF).
[0252] D. MicroRNA (miRNA) MicroRNAs (miRNAs) are small non-coding RNA molecules (approximately 18-22 nucleotides) that negatively regulate gene expression by inhibiting translation or inducing the degradation of messenger RNA (mRNA). Since their discovery, miRNAs have been implicated in various cellular processes, including apoptosis, differentiation, and cell proliferation, and have been shown to play important roles in carcinogenesis. The ability of miRNAs to regulate gene expression has led to the development of gene therapy. In this method, miRNAs are expressed as useful tools in vivo.
[0253] Certain aspects of the present disclosure are directed to a plasmid-like nucleic acid molecule comprising a first ITR, a second ITR, and a gene cassette encoding an miRNA, wherein the first ITR and / or the second ITR are non-adeno-associated virus ITRs (e.g., the first ITR and / or the second ITR are derived from a non-AAV). The miRNA may be any miRNA known in the art. In some embodiments, the miRNA downregulates the expression of a target gene. In certain embodiments, the target gene is selected from SOD1, HTT, RHO, or any combination thereof.
[0254] In some embodiments, the gene cassette encodes one miRNA. In some embodiments, the gene cassette encodes more than one miRNA. In some embodiments, the gene cassette encodes two or more different miRNAs. In some embodiments, the gene cassette encodes two or more copies of the same miRNA. In some embodiments, the gene cassette encodes two or more variants of the same therapeutic protein. In certain embodiments, the gene cassette encodes one or more miRNAs and one or more therapeutic proteins.
[0255] In some embodiments, the miRNA is a naturally occurring miRNA. In some embodiments, the miRNA is an engineered miRNA. In some embodiments, the miRNA is an artificial miRNA. In certain embodiments, the miRNA comprises the miHTT-engineered miRNA disclosed in Evers et al., Molecular Therapy 26(9):1-15 (published electronically ahead of print in June 2018). In certain embodiments, the miRNA comprises the artificial miRNA miR SOD1 disclosed in Dirren et al., Annals of Clinical and Translational Neurology 2(2):167-84 (February 2015). In certain embodiments, the miRNA comprises miR-708, which targets RHO (see Behrman et al., JCB 192(6):919-27 (2011)).
[0256] In some embodiments, the miRNA upregulates the expression of a gene by downregulating the expression of an inhibitor of the gene. In some embodiments, the inhibitor is a natural, e.g., wild-type, inhibitor. In some embodiments, the inhibitor is due to a mutated, heterologous, and / or misexpressed gene.
[0257] E. Dissimilar parts In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, wherein the therapeutic protein comprises at least one heterologous moiety. In some embodiments, the heterologous moiety is fused to the N-terminus or C-terminus of the therapeutic protein. In other embodiments, the heterologous moiety is inserted between two amino acids within the therapeutic protein.
[0258] In some embodiments, a therapeutic protein comprises a FVIII polypeptide and a heterologous moiety inserted between two amino acids within the FVIII polypeptide. In some embodiments, the heterologous moiety is inserted into the FVIII polypeptide at one or more insertion sites selected from Table 5. In some embodiments, the heterologous amino acid sequence is inserted into the FVIII polypeptide at any site disclosed in International Patent Application Publication Nos. WO2013 / 123457A1, WO2015 / 106052A1, or U.S. Patent Application Publication No. 2015 / 0158929, which are incorporated by reference in their entireties. In one specific embodiment, the therapeutic protein comprises FVIII and a heterologous moiety, and the heterologous moiety is inserted within FVIII immediately downstream of amino acid 745 relative to mature FVIII. In one specific embodiment, the therapeutic protein comprises FVIII and an XTEN, and the XTEN is inserted within FVIII immediately downstream of amino acid 745 relative to mature FVIII. In one specific embodiment, the FVIII comprises a deletion of amino acids 746-1646, corresponding to human mature FVIII (SEQ ID NO: 15), and the heterologous moiety is inserted immediately downstream of amino acid 745, corresponding to human mature FVIII (SEQ ID NO: 15).
[0259] [Table 19]
[0260] In some embodiments, a therapeutic protein comprises a FIX polypeptide and a heterologous moiety inserted between two amino acids within the FIX polypeptide. In some embodiments, the heterologous moiety is inserted within the FIX polypeptide at one or more insertion sites selected from Table 5. In some embodiments, the heterologous amino acid sequence is inserted within a coagulation factor polypeptide encoded by a nucleic acid molecule of the present disclosure at any site disclosed in International Patent Application No. PCT / US2017 / 015879, which is incorporated herein by reference in its entirety. In a specific embodiment, a therapeutic protein comprises a FIX polypeptide and a heterologous moiety, wherein the heterologous moiety is inserted within the FIX polypeptide immediately downstream of amino acid 166 relative to mature FIX. In a specific embodiment, a therapeutic protein comprises a FIX polypeptide and an XTEN, wherein the XTEN is inserted between amino acid 166 relative to mature FVIII. It is inserted into the FIX immediately downstream.
[0261] [Table 20]
[0262] In other embodiments, a Therapeutic protein of the present disclosure further comprises 2, 3, 4, 5, 6, 7, or 8 heterologous nucleotide sequences. In some embodiments, all heterologous moieties are identical. In some embodiments, at least one heterologous moiety is different from the other heterologous moieties. In some embodiments, the present disclosure may comprise more than 2, 3, 4, 5, 6, or 7 heterologous moieties in tandem.
[0263] In some embodiments, the heterologous moiety increases the half-life of the therapeutic protein (is a "half-life extender").
[0264] In some embodiments, the heterologous moiety is a peptide or polypeptide having either non-structural or structural features associated with increased in vivo half-life when incorporated into a protein of the present disclosure. Non-limiting examples include albumin, an albumin fragment, an Fc fragment of an immunoglobulin, the C-terminal peptide (CTP) of the beta subunit of human chorionic gonadotropin, a HAP sequence, an XTEN sequence, transferrin or a fragment thereof, a PAS polypeptide, a polyglycine linker, a polyserine linker, an albumin-binding moiety, or any fragment, derivative, variant, or combination of these polypeptides. In a specific embodiment, the heterologous amino acid sequence is an immunoglobulin constant region or a portion thereof, transferrin, albumin, or a PAS sequence. In some aspects, the heterologous moiety comprises von Willebrand factor or a fragment thereof. In other related embodiments, the heterologous moiety may comprise an attachment site (e.g., a cysteine amino acid) for a non-polypeptide moiety, such as polyethylene glycol (PEG), hydroxyethyl starch (HES), polysialic acid, or any derivative, variant, or combination of these elements. In some embodiments, the heterologous moiety comprises a cysteine amino acid that serves as an attachment site for a non-polypeptide moiety, such as polyethylene glycol (PEG), hydroxyethyl starch (HES), polysialic acid, or any derivative, variant, or combination of these elements.
[0265] In a specific embodiment, the first heterologous moiety is a half-life extender known in the art, and the second heterologous moiety is a half-life extender known in the art. In certain embodiments, the first heterologous moiety (e.g., a first Fc moiety) and the second heterologous moiety (e.g., a second Fc moiety) associate with each other to form a dimer. In one embodiment, the second heterologous moiety is a second Fc moiety, and the second Fc moiety is linked to or associated with the first heterologous moiety, e.g., the first Fc moiety. For example, the second heterologous moiety (e.g., the second Fc moiety) can be linked to the first heterologous moiety (e.g., the first Fc moiety) by a linker or associated with the first heterologous moiety by a covalent or non-covalent bond.
[0266] In some embodiments, the heterologous moiety is a polypeptide comprising, consisting essentially of, or consisting of at least about 10, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, at least about 2000, at least about 2500, at least about 3000, or at least about 4000 amino acids. In other embodiments, the heterologous moiety is a polypeptide comprising, consisting essentially of, or consisting of about 100 to about 200 amino acids, about 200 to about 300 amino acids, about 300 to about 400 amino acids, about 400 to about 500 amino acids, about 500 to about 600 amino acids, about 600 to about 700 amino acids, about 700 to about 800 amino acids, about 800 to about 900 amino acids, or about 900 to about 1000 amino acids.
[0267] In certain embodiments, the heterologous moiety improves one or more pharmacokinetic properties of a therapeutic protein without significantly affecting its biological activity or function.
[0268] In certain embodiments, the heterologous moiety increases the in vivo and / or in vitro half-life of a Therapeutic protein of the present disclosure. In other embodiments, the heterologous moiety facilitates visualization or localization of a Therapeutic protein of the present disclosure or a fragment thereof (e.g., a fragment comprising a heterologous moiety following proteolytic cleavage of the FVIII protein). The visualization and / or localization of a Therapeutic protein of the present disclosure or a fragment thereof is in vivo, in vitro, ex vivo, or a combination thereof.
[0269] In other embodiments, the heterologous moiety increases the stability of a Therapeutic protein or fragment thereof (e.g., a fragment comprising a heterologous moiety following proteolytic cleavage of a Therapeutic protein, e.g., a coagulation factor). As used herein, the term "stability" refers to an art-recognized measure of the maintenance of one or more physical properties of a Therapeutic protein in response to environmental conditions (e.g., increased or decreased temperature). In certain aspects, the physical property is the maintenance of the covalent structure of the Therapeutic protein (e.g., absence of proteolysis, undesired oxidation, or deamidation). In other aspects, the physical property is also the presence of the Therapeutic protein in a correctly folded state (e.g., absence of soluble or insoluble aggregation or precipitation). In one aspect, the stability of a Therapeutic protein is measured by assaying a biophysical property of the Therapeutic protein, such as temperature stability, pH unfolding profile, stable removal of glycosylation, solubility, biochemical function (e.g., ability to bind to a protein, receptor, or ligand), and / or a combination thereof. In another aspect, biochemical function is demonstrated by the binding affinity of an interaction. In one aspect, the measure of protein stability is thermal stability, i.e., resistance to heat stress. Stability can be measured using methods known in the art, such as HPLC (High Performance Liquid Chromatography), SEC (Size Exclusion Chromatography), DLS (Dynamic Light Scattering), etc. Methods for measuring thermal stability include, but are not limited to: These include differential scanning calorimetry (DSC), differential scanning fluorimetry (DSF), circular dichroism (CD), and heat stress assays.
[0270] In certain embodiments, a therapeutic protein encoded by a nucleic acid molecule of the present disclosure comprises at least one half-life extender, i.e., a heterologous moiety that increases the in vivo half-life of the therapeutic protein relative to the in vivo half-life of a corresponding therapeutic protein lacking the heterologous moiety. The in vivo half-life of a therapeutic protein can be determined by any method known to those of skill in the art, such as an activity assay (e.g., a chromosomal assay or a one-stage clotting aPTT assay where the therapeutic protein comprises a FVIII polypeptide), ELISA, ROTEM®, etc.
[0271] In some embodiments, the presence of one or more half-life extenders increases the half-life of a therapeutic protein compared to the half-life of a corresponding protein lacking such one or more half-life extenders, such that the half-life of a therapeutic protein comprising a half-life extender is at least about 1.5-fold, at least about 2-fold, at least about 2.5-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 11-fold, or at least about 12-fold longer than the in vivo half-life of a corresponding therapeutic protein lacking such half-life extenders.
[0272] In one embodiment, the half-life of a therapeutic protein comprising a half-life extender is about 1.5 to about 20 times, about 1.5 to about 15 times, or about 1.5 to about 10 times longer than the in vivo half-life of the corresponding protein lacking such half-life extender. In another embodiment, the half-life of a therapeutic protein comprising a half-life extender is about 1.5 to about 20 times, about 1.5 to about 15 times, or about 1.5 to about 10 times longer than the in vivo half-life of the corresponding protein lacking such half-life extender. Compared to the in vivo half-life, it is extended by about 2-fold to about 10-fold, about 2-fold to about 9-fold, about 2-fold to about 8-fold, about 2-fold to about 7-fold, about 2-fold to about 6-fold, about 2-fold to about 5-fold, about 2-fold to about 4-fold, about 2-fold to about 3-fold, about 2.5-fold to about 10-fold, about 2.5-fold to about 9-fold, about 2.5-fold to about 8-fold, about 2.5-fold to about 7-fold, about 2.5-fold to about 6-fold, about 2.5-fold to about 5-fold, about 2.5-fold to about 4-fold, about 2.5-fold to about 3-fold, about 3-fold to about 10-fold, about 3-fold to about 9-fold, about 3-fold to about 8-fold, about 3-fold to about 7-fold, about 3-fold to about 6-fold, about 3-fold to about 5-fold, about 3-fold to about 4-fold, about 4-fold to about 6-fold, about 5-fold to about 7-fold, or about 6-fold to about 8-fold.
[0273] In other embodiments, the half-life of a therapeutic protein comprising a half-life extender is at least about 17 hours, at least about 18 hours, at least about 19 hours, at least about 20 hours, at least about 21 hours, at least about 22 hours, at least about 23 hours, at least about 24 hours, at least about 25 hours, at least about 26 hours, at least about 27 hours, at least about 28 hours, at least about 29 hours, at least about 30 hours, at least about 31 hours, at least about 32 hours, at least about 33 hours, at least about 34 hours, at least about 35 hours, at least about 36 hours, at least about 48 hours, at least about 60 hours, at least about 72 hours, at least about 84 hours, at least about 96 hours, or at least about 108 hours.
[0274] In still other embodiments, the half-life of the therapeutic protein comprising the half-life extender is from about 15 hours to about 2 weeks, from about 16 hours to about 1 week, from about 17 hours to about 1 week, from about 18 hours to about 1 week, from about 19 hours to about 1 week, from about 20 hours to about 1 week, from about 21 hours to about 1 week, from about 22 hours to about 1 week, from about 23 hours to about 1 week, from about 24 hours to about 1 week, from about 36 hours to about 1 week, from about 48 hours to about 1 week, from about 60 hours to about 1 week, from about 24 hours to about 6 days, from about 24 hours to about 5 days, from about 24 hours to about 4 days, from about 24 hours to about 3 days, or from about 24 hours to about 2 days.
[0275] In some embodiments, the average half-life per subject of a therapeutic protein comprising a half-life extender is about 15 hours, about 16 hours, about 17 hours, about 18 hours, about 19 hours, about 20 hours, about 21 hours, about 22 hours, about 23 hours, about 24 hours (1 day), about 25 hours, about 26 hours, about 27 hours, about 28 hours, about 29 hours, about 30 hours, about 31 hours, about 32 hours, about 33 hours, The incubation period is about 34 hours, about 35 hours, about 36 hours, about 40 hours, about 44 hours, about 48 hours (2 days), about 54 hours, about 60 hours, about 72 hours (3 days), about 84 hours, about 96 hours (4 days), about 108 hours, about 120 hours (5 days), about 6 days, about 7 days (1 week), about 8 days, about 9 days, about 10 days, about 11 days, about 12 days, about 13 days, or about 14 days.
[0276] One or more half-life extenders can be fused to the C-terminus or N-terminus of the therapeutic protein or inserted within the therapeutic protein.
[0277] 1. Immunoglobulin constant region or a part thereof In another aspect, the heterologous moiety comprises one or more immunoglobulin constant regions or portions thereof (e.g., Fc regions). In one embodiment, the isolated nucleic acid molecule of the present disclosure further comprises a heterologous nucleic acid sequence encoding an immunoglobulin constant region or portion thereof. In some embodiments, the immunoglobulin constant region or portion thereof is an Fc region.
[0278] The immunoglobulin constant region is composed of domains designated CH (constant heavy chain) domains (CH1, CH2, etc.). Depending on the isotype (i.e., IgG, IgM, IgA IgD, or IgE), the constant region is composed of three or four CH domains. Some isotype (e.g., IgG) constant regions also contain a hinge region. See Janeway et al. 2001, Immunobiology, Garland Publishing, NY, NY.
[0279] The immunoglobulin constant region or portion thereof of the present disclosure can be obtained from several different sources. In one embodiment, the immunoglobulin constant region or portion thereof is derived from a human immunoglobulin. However, it is understood that the immunoglobulin constant region or portion thereof can also be derived from the immunoglobulin of another mammalian species, including, for example, rodents (e.g., mice, rats, rabbits, guinea pigs) or non-human primates (e.g., chimpanzees, macaques). Furthermore, the immunoglobulin constant region or portion thereof can be derived from any immunoglobulin class, including IgM, IgG, IgD, IgA, and IgE, and any immunoglobulin isotype, including IgG1, IgG2, IgG3, and IgG4. In one embodiment, the human isotype IgG1 is used.
[0280] Various immunoglobulin constant region gene sequences (e.g., human constant region gene sequences) are available in the form of publicly accessible deposits. Constant region domain sequences can be selected that have specific effector functions (or lack specific effector functions) or with specific modifications that reduce immunogenicity. Many sequences of antibodies and antibody-encoding genes have been published, and suitable Ig constant region sequences (e.g., hinge, CH2, and / or CH3 sequences, or portions thereof) can be derived from these sequences using art-recognized techniques. The resulting genetic material, using any of the aforementioned methods, can then be modified or synthesized to obtain the polypeptides of the present disclosure. It will be further recognized that the scope of this disclosure encompasses alleles, variants, and mutations of constant region DNA sequences.
[0281] The sequence of an immunoglobulin constant region or a portion thereof can be cloned, for example, using the polymerase chain reaction and primers selected to amplify the domain of interest. To clone the immunoglobulin constant region sequence or a portion thereof from an antibody, mRNA can be isolated from hybridoma, spleen, or lymphocytes, reverse transcribed into DNA, and the antibody gene can be amplified by PCR. PCR amplification methods are described in detail in U.S. Patent Nos. 4,683,195; 4,683,202; 4,800,159; and 4,965,188, for example, in "PCR Protocols: A Guide to Methods and Applications," edited by Innis et al., Academic Press, San Diego, CA (1990); Ho et al., 1989. Gene 77:51; Horton et al., 1993. Methods Enzymol. 217:270. PCR can be initiated with consensus constant region primers or more specific primers based on published heavy and light chain DNA and amino acid sequences. PCR can also be used to isolate DNA clones encoding the light and heavy chains of an antibody. In this case, the library can be screened with consensus primers or larger homologous probes, such as mouse constant region probes. Numerous primer sets suitable for amplifying antibody genes are known in the art (e.g., 5' primers based on the N-terminal sequence of purified antibodies (Benhar and Pastan, 1994, Protein Engineering 7:1509); rapid amplification of cDNA ends (Ruberti, F. et al., 1994, J. Immunol. Methods 173:33); antibody leader sequences (Larrick et al., 1989, Biochem. Biophys. Res. Commun. 160:1250). Cloning of antibody sequences is further described in U.S. Patent No. 5,658,570, filed January 25, 1995, by Newman et al., which is incorporated herein by reference).
[0282] As used herein, an immunoglobulin constant region can include all domains and hinge regions or portions thereof. In one embodiment, an immunoglobulin constant region or portion thereof includes a CH2 domain, a CH3 domain, and a hinge region, i.e., an Fc region or an FcRn binding partner.
[0283] As used herein, the term "Fc region" is defined as the portion of a polypeptide corresponding to the Fc region of a native Ig, i.e., as formed by the dimeric association of the Fc domains of each of its two heavy chains. A native Fc region forms a homodimer with another Fc region. In contrast, the term "genetically fused Fc region" or "single-chain Fc region" (scFc region), as used herein, refers to a synthetic dimeric Fc region composed of Fc domains genetically linked (i.e., encoded by a single contiguous gene sequence) within a single polypeptide chain. See International Patent Application Publication No. WO 2012 / 006635, which is incorporated herein by reference in its entirety.
[0284] In one embodiment, "Fc region" refers to that portion of a single Ig heavy chain beginning with the hinge region just upstream of the papain cleavage site (i.e., residue 216 of IgG, with the first residue of the heavy chain constant region being 114) and ending at the C-terminus of the antibody. Thus, a complete Fc region includes at least the hinge, CH2, and CH3 domains.
[0285] The immunoglobulin constant region or a portion thereof may be an FcRn binding partner. FcRn is active in adult epithelial tissues and is expressed in the lumen of the intestine, the lung airways, the nasal cavity surface, the vaginal surface, the colon, and the rectal surface (U.S. Patent No. 6,485,726). An FcRn binding partner is a portion of an immunoglobulin that binds to FcRn.
[0286] FcRn receptors have been isolated from several mammalian species, including humans. The sequences of human FcRn, monkey FcRn, rat FcRn, and mouse FcRn are known (S Tory et al. 1994, J. Exp. Med. 180:2377. The FcRn receptor binds IgG (but not other immunoglobulin classes such as IgA, IgM, IgD, and IgE) at a relatively low pH, actively transports the IgG transcellularly in the lumen toward the serosal membrane, and then releases the IgG at the relatively high pH found in interstitial fluid. It is expressed in adult epithelial tissues, including lung and intestinal epithelium (Israel et al. 1997, Immunology 92:69), renal proximal tubular epithelium (Kobayashi et al. 2002, Am. J. Physiol. Renal Physiol. 282:F358), and nasal epithelium, vaginal surface, and biliary surface (U.S. Patent Nos. 6,485,726, 6,030,613, 6,086,875; WO03 / 077834; US2003-0235536).
[0287] FcRn binding partners useful in the present disclosure include molecules that are specifically bound by the FcRn receptor, including whole IgG, Fc fragments of IgG, and other fragments containing the complete binding region of the FcRn receptor. The region of the Fc portion of IgG that binds to the FcRn receptor has been described based on X-ray crystallography (Burmeister et al., 1994, Nature 372:379). The main contact region of the Fc with FcRn is near the junction of the CH2 and CH3 domains. All Fc-FcRn contacts are within a single Ig heavy chain. FcRn binding partners include whole IgG, Fc fragments of IgG, and other fragments of IgG containing the complete binding region of the FcRn. Major contact sites include amino acid residues 248, 250-257, 272, 285, 288, 290-291, 308-311, and 314 in the CH2 domain, and amino acid residues 385-387, 428, and 433-436 in the CH3 domain. All references to the amino acid numbering of immunoglobulins or immunoglobulin fragments or regions are based on Kabat et al. 1991, Sequences of Proteins of Immunological Interest, US Department of Public Health, Bethesda, Md.
[0288] Fc regions or FcRn-binding partners bound to FcRn can be efficiently transported across epithelial barriers by FcRn, providing a non-invasive means for systemic administration of desired therapeutic molecules. Furthermore, fusion proteins containing Fc regions or FcRn-binding partners are phagocytosed by cells expressing FcRn. However, these fusion proteins are not subject to degradation but are recycled and re-enter the blood circulation, thereby increasing the in vivo half-life of these proteins. In certain embodiments, the portion of the immunoglobulin constant region is an Fc region or FcRn-binding partner that typically associates with another Fc region or another FcRn-binding partner via disulfide bonds and other nonspecific interactions to form dimers and higher-order multimers.
[0289] Two FcRn receptors can bind to a single Fc molecule. Crystallographic data suggest that each FcRn molecule binds to a single polypeptide of an Fc homodimer. In one embodiment, an FcRn binding partner, e.g., an Fc fragment of IgG, is linked to a biologically active molecule, providing a means for oral, buccal, sublingual, rectal, vaginal, nasal aerosol, or pulmonary delivery, or topical ocular delivery of the biologically active molecule. In another embodiment, the coagulation factor protein can be administered invasively, e.g., subcutaneously or intravenously.
[0290] An FcRn binding partner region is a molecule or moiety that specifically binds to the FcRn receptor, thereby allowing active transport of the Fc region by the FcRn receptor. Specific binding refers to two molecules forming a complex that is relatively stable under physiological conditions. Specific binding is characterized by high affinity and low to moderate capacity, as distinguished from nonspecific binding, which is usually characterized by low affinity and moderate to high capacity. Typically, , the affinity constant K A is 10 6 M -1 More than or equal to 10 8 M -1 Binding is considered specific when the binding affinity exceeds 0.05. If necessary, non-specific binding can be reduced by changing the binding conditions without substantially affecting the specific binding affinity. Those skilled in the art can optimize appropriate binding conditions, such as the concentration of the molecule, the ionic strength of the solution, temperature, binding time, and the concentration of the blocking agent (e.g., serum albumin, milk casein), using routine techniques.
[0291] In certain embodiments, a Therapeutic protein encoded by a nucleic acid molecule of the present disclosure comprises one or more truncated Fc regions that, despite being truncated, are sufficient to confer Fc receptor (FcR) binding properties to the Fc region. For example, the portion of the Fc region that binds to FcRn (i.e., the FcRn-binding portion) comprises amino acids approximately 282-438 of IgG1, according to EU numbering (major contact sites are amino acids 248, 250-257, 272, 285, 288, 290-291, 308-311, and 314 of the CH2 domain and amino acid residues 385-387, 428, and 433-436 of the CH3 domain). Thus, an Fc region of the present disclosure may comprise or consist of an FcRn-binding portion. The FcRn-binding portion may be derived from a heavy chain of any isotype, including IgG1, IgG2, IgG3, and IgG4. In one embodiment, an FcRn-binding portion derived from an antibody of human isotype IgG1 is used. In another embodiment, an FcRn-binding portion derived from an antibody of human isotype IgG4 is used.
[0292] The Fc region can be obtained from several different sources. In one embodiment, the Fc region of the polypeptide is derived from a human immunoglobulin. However, the Fc portion may be derived from the immunoglobulin of another mammalian species, including, for example, rodents (e.g., mice, rats, rabbits, guinea pigs) or non-human primates (e.g., chimpanzees, macaques). Furthermore, the Fc domain or portion thereof may be derived from any immunoglobulin class, including IgM, IgG, IgD, IgA, and IgE, and any immunoglobulin isotype, including IgG1, IgG2, IgG3, and IgG4. In another embodiment, the human isotype IgG1 is used.
[0293] In certain embodiments, the Fc variant provides an alteration in at least one effector function conferred by the Fc portion comprising the wild-type Fc domain (e.g., an improved or decreased ability of the Fc region to bind to an Fc receptor (e.g., FcγRI, FcγRII, or FcγRIII) or a complement protein (e.g., C1q), or to induce antibody-dependent cellular cytotoxicity (ADCC), phagocytosis, or complement-dependent cytotoxicity (CDCC)). In other embodiments, the Fc variant provides an engineered cysteine residue.
[0294] The Fc regions of the present disclosure can employ art-recognized Fc variants known to alter (e.g., enhance or decrease) effector function and / or FcR or FcRn binding. Specifically, the Fc regions of the present disclosure can employ Fc variants described in, for example, International PCT Application Publication Nos. WO88 / 07089A1, WO96 / 14339A1, WO98 / 05787A1, WO98 / 23289A1, WO99 / 51642A1, WO99 / 58572A1, WO00 / 09560A2, WO00 / 32767A1, and WO00 / 42 072A2, WO02 / 44215A2, WO02 / 060919A2, WO03 / 074569A2, WO04 / 016750A2, WO04 / 029207A2, WO04 / 029207A2 04 / 035752A2, WO04 / 063351A2, WO04 / 074455A2, WO04 / 099249A2, WO05 / 040217A2, WO04 / 044859, Nos. WO05 / 070963A1, WO05 / 077981A2, WO05 / 092925A2, WO05 / 123780A2, WO06 / 019447A1, WO06 / 047350A2, and WO06 / 085967A2; U.S. Patent Application Publication No. US2007 / 0231 329, US2007 / 0231329, US2007 / 0237765, US2007 / 0237766, US2007 / 0237767, US2007 / 0243188, US20070248603, US20070286859, US20080057 No. 056; or U.S. Patent Nos. 5,648,260, 5,739,277, 5,834,250, 5,869,046, 6,096,871, 6,121,022, 6,194,551, 6,242,195, 6,277,375, 6,528,624, Nos. 6,538,124, 6,737,056, 6,821,505, 6,998,253, 7,083,784, 7,404,956, and 7,317,091 may contain modifications (e.g., substitutions) at one or more of the amino acid positions disclosed therein. In one embodiment, a specific modification (e.g., a specific substitution of one or more amino acids disclosed in the art) is made at one or more of the disclosed amino acid positions. In another embodiment, a different modification (e.g., a different substitution of one or more amino acid positions disclosed in the art) is made at one or more of the disclosed amino acid positions.
[0295] The Fc region of IgG or FcRn binding partners can be modified using well-recognized procedures, such as site-directed mutagenesis, to obtain modified IgG or Fc fragments or portions thereof that will be bound by FcRn. Such modifications include modifications at sites distant from the FcRn contact site and modifications within the contact site that retain or even enhance binding to FcRn. For example, the Fc region of human IgG1 can be modified without significant loss of Fc binding affinity to FcRn. The following single amino acid residues in Fc (Fcγ1) can be substituted: P238A, S239A, K246A, K248A, D249A, M252A, T256A, E258A, T260A, D265A, S267A, H268A, E269A, D270A, E272A, L274A, N276A, Y278A, D280A, V282A, E283A, H285A, N28 6A, T289A, K290A, R292A, E293A, E294A, Q295A, Y296F, N297A, S298A, Y300F, R301A, V303A, V305A, T307A , L309A, Q311A, D312A, N315A, K317A, E318A, K320A, K322A, S324A, K326A, A327Q, P329A, A330Q, P331A, E 333A, K334A, T335A, S337A, K338A, K340A, Q342A, R344A, E345A, Q347A, R355A, E356A, M358A, T359A, K3 60A, N361A, Q362A, Y373A, S375A, D376A, A378Q, E380A, E382A, S383A, N384A, Q386A, E388A, N389A, N390 A, Y391F, K392A, L398A, S400A, D401A, D413A, K414A, R416A, Q418A, Q419A, N421A, V422A, S424A, E430A, N434A, T437A, Q438A, K439A, S440A, S444A, and K447A (e.g., P238A represents a substitution of the wild-type proline at position 238 with alanine). By way of example, in specific embodiments, an N297A mutation is incorporated to remove a highly conserved N-glycosylation site.In addition to alanine, other amino acids can be substituted for the wild-type amino acids at the positions identified above. Mutations can be individually incorporated into an Fc, resulting in over 100 Fc regions that differ from the native Fc. Furthermore, combinations of two, three, or more of these individual mutations can be incorporated together, resulting in hundreds of additional Fc regions.
[0296] Certain of the above mutations can confer new functions to the Fc region or FcRn-binding partner. For example, in one embodiment, N297A is incorporated to remove a highly conserved N-glycosylation site. The effect of this mutation is to reduce immunogenicity, thereby increasing the circulating half-life of the Fc region and rendering the Fc region unable to bind to FcγRI, FcγRIIA, FcγRIIB, and FcγRIIIA without compromising affinity for FcRn (Routledge et al., 1995, Transplantation 60:847; Friend et al., 1999, Transplantation 68:1632; Shields et al., 1995, J. Biol. Chem. 276:6591). As a further example of new functions resulting from the above mutations, in some cases, affinity for FcRn can be increased compared to wild-type affinity. This increased affinity can reflect an increased "on" rate, a decreased "off" rate, or both an increased "on" rate and a decreased "off" rate. Examples of mutations that are thought to increase affinity for FcRn include, but are not limited to, T256A, T307A, E380A, and N434A (Shields et al. 2001, J. Biol. Chem. 276:6591).
[0297] Furthermore, at least three human Fc gamma receptors are believed to recognize a binding site on IgG within the downstream hinge region, generally amino acids 234-237. Therefore, another example of new function and potentially reduced immunogenicity could result from mutation of this region, for example, by substituting amino acids 233-236 "ELLG" (SEQ ID NO: 45) of human IgG1 with the corresponding sequence "PVA" (one amino acid deletion) of IgG2. It has been shown that when such mutations are introduced, FcγRI, FcγRII, and FcγRIII, which mediate various effector functions, no longer bind to IgG1. Ward and Ghetie 1995, Therapeutic Immunology 2:77; Armour et al. 1999, Eur. J. Immunol. 29:2613.
[0298] In another embodiment, the immunoglobulin constant region or portion thereof comprises an amino acid sequence in the hinge region or portion thereof that forms one or more disulfide bonds with a second immunoglobulin constant region or portion thereof. The second immunoglobulin constant region or portion thereof can be linked to a second polypeptide to unite the therapeutic protein and the second polypeptide. In some embodiments, the second polypeptide is an enhancer moiety. As used herein, the term "enhancer moiety" refers to a molecule, fragment thereof, or polypeptide component capable of enhancing the activity of a therapeutic protein. The enhancer moiety is a cofactor, such as soluble tissue factor (sTF), where the therapeutic protein is a clotting factor or a procoagulant peptide. Thus, upon activation of the clotting factor, the enhancer moiety is available to enhance the activity of the clotting factor.
[0299] In certain embodiments, the therapeutic proteins encoded by the nucleic acid molecules of the present disclosure comprise amino acid substitutions to an immunoglobulin constant region or portion thereof (e.g., an Fc variant) that alter the antigen-dependent effector function of the Ig constant region, particularly the circulating half-life of the protein.
[0300] 2.scFc region In another aspect, the heterologous moiety comprises an scFc (single-chain Fc) region. In one embodiment, the isolated nucleic acid molecule of the present disclosure further comprises a heterologous nucleic acid sequence encoding an ScFc region. An scFc region is a combination of at least two immunoglobulin constant regions or subunits capable of folding (e.g., intramolecularly or intermolecularly) into the same linear polypeptide chain to form a functional scFc region linked by an Fc peptide linker. or portions thereof (e.g., Fc moieties or Fc domains (e.g., 2, 3, 4, 5, 6, or more Fc moieties or domains)). For example, in one embodiment, a polypeptide of the disclosure is capable of binding via its ScFc region to at least one Fc receptor (e.g., FcRn, an FcγR receptor (e.g., FcγRIII), or a complement protein (e.g., C1q)) for purposes of improving half-life or eliciting immune effector function (e.g., antibody-dependent cellular cytotoxicity (ADCC), phagocytosis, or complement-dependent cytotoxicity (CDCC)) and / or improving manufacturability.
[0301] 3.CTP In another embodiment, the heterologous moiety comprises one C-terminal peptide (CTP) of the beta subunit of human chorionic gonadotropin, or a fragment, variant, or derivative thereof. One or more CTP peptides inserted into a recombinant protein are known to increase the in vivo half-life of that protein. See, e.g., U.S. Patent No. 5,712,122, incorporated herein by reference in its entirety.
[0302] Exemplary CTP peptides include DPRFQDSSSSKAPPPSLPSPSRLPGPSDTPIL (SEQ ID NO: 33) or SSSSKAPPPSLPSPSRLPGPSDTPILPQ (SEQ ID NO: 34). See, e.g., U.S. Patent Application Publication No. 2009 / 0087411A1, incorporated by reference.
[0303] 4. XTEN sequence In some embodiments, the heterologous moiety comprises one or more XTEN sequences, fragments, variants, or derivatives thereof. As used herein, "XTEN sequence" refers to an extended polypeptide having a non-naturally occurring, substantially non-repetitive sequence composed primarily of small hydrophilic amino acids and exhibiting little or no secondary or tertiary structure under physiological conditions. As a heterologous moiety, XTEN can function as a half-life extending moiety. Additionally, XTEN can provide desirable properties, including, but not limited to, enhanced pharmacokinetic parameters and solubility characteristics.
[0304] Incorporation of heterologous moieties, including XTEN sequences, into proteins of the present disclosure can confer one or more of the following advantageous properties to the protein: conformational flexibility, increased aqueous solubility, high protease resistance, low immunogenicity, low binding to mammalian receptors, or increased hydrodynamic (or Stokes) radius.
[0305] In certain embodiments, the XTEN sequence can provide improved pharmacokinetic properties, such as a longer in vivo half-life or increased area under the curve (AUC), such that the proteins of the disclosure remain in vivo and have procoagulant activity for an increased period of time compared to the same protein but without the XTEN heterologous moiety.
[0306] In some embodiments, XTEN sequences useful with the present disclosure are peptides or polypeptides having more than about 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1200, 1400, 1600, 1800, or 2000 amino acid residues. In certain embodiments, XTEN is a peptide or polypeptide having from greater than about 20 to about 3000 amino acid residues, from greater than 30 to about 2500 residues, from greater than 40 to about 2000 residues, from greater than 50 to about 1500 residues, from greater than 60 to about 1000 residues, from greater than 70 to about 900 residues, from greater than 80 to about 800 residues, from greater than 90 to about 700 residues, from greater than 100 to about 600 residues, from greater than 110 to about 500 residues, or from greater than 120 to about 400 residues. In a particular embod...
Claims
1. A nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette; wherein the first ITR and / or the second ITR are non-adeno-associated virus (non-AAV) ITRs, wherein the gene cassette is disposed between the first ITR and the second ITR, and wherein the gene cassette encodes a therapeutic protein, a miRNA, or both a therapeutic protein and a miRNA.
2. The nucleic acid molecule of claim 1 , wherein the therapeutic protein comprises a clotting factor.
3. 3. The nucleic acid molecule of claim 1 or 2, wherein the non-AAV is selected from the group consisting of members of the Parvoviridae family of viruses.
4. The nucleic acid molecule according to any one of claims 1 to 3, wherein the first ITR and the second ITR are non-AAV ITRs.
5. Members of the Parvoviridae family of viruses include Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Aveparvovirus, Copiparvovirus, and Protoparvovirus.
5. The nucleic acid molecule of claim 3, wherein the nucleic acid molecule is selected from the group consisting of: Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstildensovirus, Muscovy duck parvovirus (MDPV) strain, porcine parvovirus (U44978), minute virus of mice (U34256), canine parvovirus (M19296), and mink enteritis virus (D00765).
6. The nucleic acid molecule of any one of claims 3 to 5, wherein the member of the Parvoviridae family of viruses is the erythrovirus parvovirus B19 (human virus).
7. 6. The nucleic acid molecule of any one of claims 3 to 5, wherein the viral Parvoviridae is the Dependovirus goose parvovirus (GPV) strain.
8. The nucleic acid molecule of any one of claims 1 to 7, further comprising a tissue-specific promoter.
9. The nucleic acid molecule of claim 8 , wherein the promoter promotes expression of the therapeutic protein in hepatocytes, endothelial cells, muscle cells, sinusoidal cells, or any combination thereof.
10. The promoters include mouse thyretin promoter (mTTR), endogenous human factor VIII promoter (F8), human alpha-1-antitrypsin promoter (hAAT), human albumin minimal promoter, mouse albumin promoter, tristetraprolin (TTP) promoter, CASI promoter, CAG promoter, cytomegalovirus (CAG) promoter, and cytomegalovirus (CAG) promoter.
10. The nucleic acid molecule of claim 8 or 9, wherein the promoter is selected from the group consisting of: a cholera virus (CMV) promoter, alpha 1-antitrypsin (AAT), muscle creatine kinase (MCK), myosin heavy chain alpha (αMHC), myoglobin (MB), desmin (DES), SPc5-12, 2R5Sc5-12, dMCK, tMCK, and phosphoglycerate kinase (PGK) promoter.
11. The nucleotide sequence is (a) intron sequences; (b) post-transcriptional regulatory elements; (c) 3′UTR poly(A) tail sequence; (d) an enhancer sequence, or (e) Any combination of (a) to (d) The nucleic acid molecule according to any one of claims 1 to 10, further comprising:
12. (a) the intron sequence is located 5' to the nucleic acid sequence encoding the clotting factor; (b) the post-transcriptional regulatory element comprises a mutated woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, or any combination thereof; (c) the 3′UTR poly(A) tail sequence is selected from the group consisting of bGH poly(A), actin poly(A), hemoglobin poly(A), and any combination thereof; or (d) any combination of (a) to (c); The nucleic acid molecule of claim 11.
13. (a) the intron sequence comprises SEQ ID NO: 115; (b) whether the microRNA binding site contains a binding site for miR142-3p; (c) the 3′UTR poly(A) tail sequence comprises a bGH poly(A); or (d) any combination of (a) to (c); The nucleic acid molecule of claim 12.
14. (a) a first ITR that is an ITR of a non-AAV family member of the Parvoviridae family; (b) a tissue-specific promoter sequence comprising the TTP promoter; (c) an intron that is a synthetic intron; (d) a nucleotide sequence encoding a coagulation factor; (e) a post-transcriptional regulatory element containing WPRE; (f) a 3'UTR poly(A) tail sequence containing bGHpA; and (g) a second ITR that is an ITR of a non-AAV family member of the Parvoviridae family The nucleic acid molecule according to any one of claims 1 to 13, comprising, in this order:
15. 15. The nucleic acid molecule of any one of claims 2 to 14, wherein the coagulation factor comprises factor I (FI), factor II (FII), factor V (FV), factor VII (FVII), factor VIII (FVIII), factor IX (FIX), factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FXIII), von Willebrand factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2-antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor-1 (PAI-1), plasminogen activator inhibitor-2 (PAI2), or any combination thereof.
16. The nucleic acid molecule of claim 15, wherein the FVIII comprises an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to an amino acid sequence set forth in a sequence selected from SEQ ID NOs: 71, 106, 107, and 109.
17. 17. The nucleic acid molecule of any one of claims 1 to 16, wherein the coagulation factor comprises a heterologous moiety selected from the group consisting of albumin or a fragment thereof, an immunoglobulin Fc region, a C-terminal peptide (CTP) of the beta subunit of human chorionic gonadotropin, a PAS sequence, a HAP sequence, transferrin or a fragment thereof, an albumin binding moiety, a derivative thereof, or any combination thereof.
18. The nucleic acid molecule of any one of claims 15 to 17, wherein the FVIII further comprises an FcRn binding partner.
19. The nucleic acid molecule of any one of claims 1 to 18, formulated in a delivery agent comprising lipid nanoparticles.
20. A pharmaceutical composition comprising the nucleic acid molecule of any one of claims 1 to 19 and a pharmaceutically acceptable carrier.
21. A method for expressing a coagulation factor in a subject in need thereof, comprising administering to said subject a nucleic acid molecule according to any one of claims 1 to 19 or a pharmaceutical composition according to claim 20.