Therapeutic adeno-associated virus comprising liver-specific promoters for treating Pompe disease and lysosomal disorders
Patent Information
- Application Number
- AU2020388634
- Authority / Receiving Office
- AU · AU
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-05-12
- Filing Date
- 2020-11-19
- Publication Date
- 2026-08-27
- Estimated Expiration
- 2040-11-19
AI Technical Summary
Current methods for treating Pompe disease and lysosomal storage disorders face challenges in delivering lysosomal enzymes, such as alpha-glucosidase (GAA), to lysosomes efficiently, leading to frequent infusions and side effects like hypoglycemia due to non-specific cellular uptake and antibody formation.
Development of adeno-associated virus (AAV) vectors equipped with liver-specific promoters and targeting sequences, like IGF2 peptides, to specifically express and secrete GAA protein from the liver, ensuring targeted delivery to lysosomes in muscle and liver cells, reducing the need for frequent infusions and minimizing side effects.
The AAV vectors enable effective and sustained expression of GAA in affected tissues, improving treatment efficacy for Pompe disease by enhancing enzyme delivery and reducing systemic side effects, potentially leading to long-term non-invasive therapy.
Smart Images

Figure 00000001_0000 
Figure 00000175_0000 
Figure 00000176_0000
Abstract
Description
THERAPEUTIC ADENO-ASSOCIATED VIRUS COMPRISING LIVER-SPECIFIC PROMOTERS FOR TREATING POMPE DISEASE AND LYSOSOMAL DISORDERS SEQUENCE LISTING
[0001] This invention claims benefit under 35 U.S.C. §119(e) of U.S. Provisional Application 62 / 937,556 filed on November 19, 2019, and U.S. Provisional Application 62 / 937,583 filed on November 19, 2019, and U.S. Provisional Application 63 / 023,570 filed on May 12, 2020, the contents of each are incorporated herein in their entirety by reference. SEQUENCE LISTING
[0002] The instant application contains a Sequence Listing which has been submitted electronically in ASCII format, and is hereby incorporated by reference in its entirety. Said ASCII copy, created on November 17, 2020, is named 046192-096600WOPT SL.txt and is 840,179 bytes in size. FIELD OF THE INVENTION
[0003] The present invention relates to adeno-associated virus (AAV) particles, virions and vectors for targeted translocation of lysosomal enzymes, such as, e.g., an alpha-glucosidase (GAA) polypeptide, and method of use for the treatment of lysosomal storage diseases and disorders, such as, e.g., Pompe disease. BACKGROUND
[0004] More than forty lysosomal storage diseases (LSDs) are caused, directly or indirectly, by the absence of one or more lysosomal enzymes in the lysosome Enzyme replacement therapy for LSDs is being actively pursued. Therapy generally requires that LSD proteins be taken up and delivered to the lysosomes of a variety of cell types in an M6P-dependent fashion. One possible approach involves purifying an LSD protein and modifying it to incorporate a carbohydrate moiety with M6P. This modified material may be taken up by the cells more efficiently than unmodified LSD proteins due to interaction with M6P receptors on the cell surface.
[0005] As an alternative or adjunct to enzyme therapy, the feasibility of gene therapy approaches to treat GSD-II have been investigated (Amalfitano, A., et al., (1999) Proc. Natl. Acad. Sci. USA 96:8861-8866, Ding, E., et al. (2002) Mol. Ther. 5:436-446, Fraites, T. J., et al., (2002) Mol. Ther. 5:571-578, Tsujino, S., et al. (1998) Hum. Gene Ther. 9:1609-1616).
[0006] However, viral or AAV delivery of genes, in particular lysosomal proteins and enzymes for treatment of lysosomal storage diseases has challenges. Normally, mammalian lysosomal enzymes are synthesized in the cytosol and traverse the ER where they are glycosylated with N-linked, high mannose type carbohydrate. In the Golgi, the high mannose carbohydrate is modified on lysosomal proteins by the addition of mannose-6-phosphate (M6P) which targets these proteins to the lysosome. The M6P-modified proteins are delivered to the lysosome via interaction with either of two M6P receptors. However, recombinantly produced proteins used in enzyme replacement therapy often lack the addition of the M6P which is required for targeting them to the lysosomes, therefore, often requiring high doses of recombinantly produced enzymes to be administered to a patient and / or frequent infusions.
[0007] Acid alpha-glucosidase (GAA) is a lysosomal enzyme that hydrolyzes the alpha 1-4 linkage in maltose and other linear oligosaccharides, including the outer branches of glycogen, thereby breaking down excess glycogen in the lysosome (Hirschhom et al. (2001) in The Metabolic and Molecular Basis of Inherited Disease, Scriver, et al., eds. (2001), McGraw-Hill: New York, p. 3389- 3420). Like other mammalian lysosomal enzymes, GAA is synthesized in the cytosol and traverses the ER where it is glycosylated with N-linked, high mannose type carbohydrate. In the Golgi, the high mannose carbohydrate is modified on lysosomal proteins by the addition of mannose-6-phosphate (M6P) which targets these proteins to the lysosome. The M6P-modified proteins are delivered to the lysosome via interaction with either of two M6P receptors. The most favorable form of modification is when two M6Ps are added to a high mannose carbohydrate.
[0008] Insufficient GAA activity in the lysosome results in Pompe disease, a disease also known as acid maltase deficiency (AMD), glycogen storage disease type II (GSDII), glycogenosis type II, or GAA deficiency. The diminished enzymatic activity occurs due to a variety of missense and nonsense mutations in the gene encoding GAA. Consequently, glycogen accumulates in the lysosomes of all cells in patients with Pompe disease. In particular, glycogen accumulation is most pronounced in lysosomes of cardiac and skeletal muscle, liver, and other tissues. Accumulated glycogen ultimately impairs muscle function. In the most severe form of Pompe disease, death occurs before two years of age due to cardio-respiratory failure.
[0009] There is a need for an effective treatment of Pompe disease. Enzyme replacement therapeutics for Pompe require a recombinant GAA protein to be administered and taken up by muscle and liver cells in the subject where it is subsequently transported to the lysosomes in those cells in a M6P-dependent fashion. However, while enzyme therapy has demonstrated reasonable efficacy for severe infantile GSD II, the benefit of GAA enzyme therapy is limited by the need for frequent infusions as well as the subject developing inhibitor or neutralizing antibodies against recombinant hGAA protein (Amalfitano, A., et al. (2001) Genet. In Med. 3:132-138).
[0010] Gene therapy has the potential to not only cure genetic disorders, but to also facilitate the long-term non-invasive treatment of acquired and degenerative disease using a virus. One gene therapy vector is adeno-associated virus (AAV). AAV itself is a non-pathogenic-dependent parvovirus that needs helper viruses for efficient replication. AAV has been utilized as a virus vector for gene therapy because of its safety and simplicity. AAV has a broad host and cell type tropism capable of transducing both dividing and non-dividing cells.
[0011] However, AAV delivery of the GAA polypeptide has some challenges with respect to achieving sufficient expression in the liver and / or delivery to lysosomes with patients reporting to experience glycaemia. In particular, in human subjects, the administration of rAAV vectors encoding GAA polypeptide have resulted in a number of patients experiencing hypoglycemia or becoming hyperglycemic due to non- specific update in cells (see, e.g., Byme ef al, A study on the safety and efficacy of Reveglucosidease alfa in patients with late-onset Pompe disease; Orphanet J. of Rare diseases; 2017; 12: 144).
[0012] Accordingly, there is a need in the art for improved methods of producing lysosomal polypeptides such as GAA in vitro and in vivo, for example, to treat lysosomal polypeptide deficiencies, including modifications of GAA. Moreover, there is a need for improved secretion from the liver as well as improved targeting of GAA to the lysosomes to help reduce any side effects from overexpression of the GAA polypeptide, and reducing the risk of hypoglycemia. Further, there is a need for methods that result in systemic delivery of GAA and other lysosomal polypeptides to affected tissues and organs. In particular, there remains a need for more efficient methods for administering GAA protein to subjects and targeting GAA protein to patient lysosomes, while reducing any potential side effects. SUMMARY OF THE INVENTION
[0013] The technology described herein relates generally to gene therapy constructs, methods and composition, for the treatment lysosomal storage diseases and disorders, such as, for example but not limited to, Pompe Disease. More particularly, the technology relates to adeno-associated (AAV) virions configured for delivering a lysosomal enzyme, e.g., a GAA polypeptide to a subject, and more particularly for delivering a lysosomal enzyme, e.g., a GAA polypeptide to the liver of a subject where it is targeted to the lysosomes and secreted from the liver cells.
[0014] In particular, described herein are targeted viral vectors, e.g., using rAAV vectors as an exemplary example, that comprise a nucleotide sequence containing inverted terminal repeats (ITRs), a promoter, a heterologous gene, a poly-A tail and potentially other regulator elements for use to treat a lysosomal storage disease, such as those listed in Table 5A or Table 6A herein, wherein the heterologous gene is a lysosomal enzyme, such as, e.g., GAA, and wherein the vector, ¢.g., TAAV can be administered to a patient in a therapeutically effective dose that is delivered to the appropriate tissue and / or organ for expression of the heterologous lysosomal enzyme gene and treatment of the disease, e.g., Pompe disease.
[0015] Aspects of the present invention teach certain benefits in construction and use which give rise to the exemplary advantages described below.
[0016] Accordingly, in particular embodiments described herein are rAAV vectors that comprises a nucleotide sequence containing inverted terminal repeats (ITRs) and located between the ITRs, a liver specific promoter (LSP), a heterologous nucleic acid sequence that encodes the acid alpha-glucosidase (GAA) protein, a poly-A tail and potentially other regulator elements for use to treat Pompe Disease, and wherein the rAAV expressing GAA protein can be administered to a patient in a therapeutically effective dose that is delivered to the appropriate tissue and / or organ for expression of the heterologous gene encoding the GAA protein for the treatment of a subject with Pompe disease.
[0017] More specifically, the AAV virion or genome comprises a LSP selected from any promoter listed in Table 4 herein, or a functional variant or functional fragment thereof, or any LSP selected from SEQ ID NO: 86, 91-96, or 146-150 or a functional variant or functional fragment thereof, that enables the lysosomal protein, e.g., GAA protein to be preferentially expressed in the liver. In some embodiments, the liver-specific promoter, while preferentially expresses the hGAA protein in the liver, can also express the hGAA to some extent in another tissue of interest, e.g., the muscle, or CNS, or muscle and CNS tissues. In some embodiments, the expressed lysosomal enzyme, e.g., GAA protein can be configured as GAA-fusion protein with a targeting sequence, such as a IGF2 targeting peptide as disclosed herein that targets the GAA protein to lysosomes, and / or fused with a signal peptide (SP), the GAA protein is expressed by the rAAV genome in the liver, where it is secreted and taken up by lysosomes of mammalian cells, in particular muscle cells.
[0018] In some embodiments of the compositions and methods described herein, the rAAV vector disclosed herein comprises, in its genome: 5° and 3° AAV inverted terminal repeats (ITR) sequences, and located between the 5° and 3° ITRs, a liver specific promoter (LSP) operatively linked to a heterologous nucleic acid sequence encoding an alpha-glucosidase (GAA) polypeptide, wherein the liver-specific promoter (LSP) comprises a nucleic acid sequence selected from any promoter listed from SEQ ID NOS: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), SEQ ID NOS: 93 (SP0239), SEQ ID NO: 94 (SP0265, also referred to SP131_A1), SEQ ID NO: 95 (SP0240) or SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240-UTR), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-Al- UTR) or a functional fragment or variant thereof, or any LSP selected from SEQ ID NO: 270-341 or 342-430, or a functional fragment or variant thereof. In some embodiments, the GAA polypeptide is not fused to either a IGF2 targeting sequence, or signal sequence. In some embodiments, the GAA polypeptide is fused to a signal sequence as disclosed herein, and / or a IGF2 targeting sequence as disclosed herein.
[0019] In some embodiments of the compositions and methods described herein, the rAAV vector disclosed herein comprises, in its genome: 5° and 3° AAV inverted terminal repeats (ITR) sequences, and located between the 5° and 3’ ITRs, a liver specific promoter (LSP) operatively linked to a heterologous nucleic acid sequence encoding a fusion polypeptide comprising (i) a secretory signal peptide, and / or an IGF?2 targeting peptide; and (ii) an alpha-glucosidase (GAA) polypeptide, wherein the liver-specific promoter (LSP) is selected from any promoter listed in Table 4 herein, or a functional variant or functional fragment thereof, or any LSP selected from SEQ ID NO: 86, 91-96, or 146-150 or a functional variant or functional fragment thereof.
[0020] In some embodiments, the rAAV vector disclosed herein comprises, in its genome: 5° and 3’ AAV inverted terminal repeats (ITR) sequences, and located between the 5° and 3° ITRs, a heterologous nucleic acid sequence encoding a fusion polypeptide comprising (i) a secretory signal peptide (also referred to as a leader peptide), and (ii) an alpha-glucosidase (GAA) polypeptide, wherein the heterologous nucleic acid is operatively linked to a liver-specific promoter (LSP) selected from any promoter listed in Table 4 herein, or a functional variant or functional fragment thereof, or any LSP selected from SEQ ID NO: 86, 91-96, or 146-150, or a functional variant or functional fragment thereof or any LSP selected from Table 4 herein, or a functional variant or functional fragment thereof. Exemplary leader sequences include, but are not limited to the innate GAA leader sequence, AAT sequence, IL2(1-3), IL2 leader sequence (IL2 wt), a modified IL2 leader sequence (IL2 mut), fibronectin (FN1) signal sequence, or IgG leader sequence or functional variants thereof, as disclosed herein. In some embodiments, the AAV vector comprises a Kozak sequence located between the LSP and the leader sequence.
[0021] In some embodiments, the AAV vector disclosed herein comprises, in its genome: 5° and 3” AAV inverted terminal repeats (ITR) sequences, and located between the 5° and 3° ITRs, a heterologous nucleic acid sequence encoding a fusion polypeptide comprising (i) an IGF2 targeting peptide, and (ii) an alpha-glucosidase (GAA) polypeptide, wherein the heterologous nucleic acid is operatively linked to a liver-specific promoter (LSP) selected from any promoter listed in Table 4 herein, or a functional variant or functional fragment thereof, or any LSP selected from SEQ ID NO: 86, 91-96, or 146-150 or a functional variant or functional fragment thereof.
[0022] In a further embodiments, the rAAV vector disclosed herein comprises, in its genome: 5° and 3° AAV inverted terminal repeats (ITR) sequences, and located between the 5° and 3° ITRs, a heterologous nucleic acid sequence encoding an alpha-glucosidase (GAA) polypeptide (i.e., where the GAA polypeptide not fused to a heterologous signal peptide (or a leader sequence), or not fused to an IGF2 targeting sequence as described herein), wherein the heterologous nucleic acid is operatively linked to a liver-specific promoter (LSP) selected from any promoter listed in Table 4 herein, or a functional variant or functional fragment thereof, or any LSP selected from SEQ ID NO: 86, 91-96, or 146-150 or a functional variant or functional fragment thereof.
[0023] In some embodiments, the rAAV vector comprises a liver specific capsid, e.g., a liver specific capsid selected from XL32 and XL32.1, as disclosed in W02019 / 241324, which is incorporated herein in its entirety by reference. In some embodiments, the rAAV vector is a AAVXL32 or AAVXL32.1 as disclosed in W02019 / 241324, which is incorporated herein in its entirety by reference, or a AAV8 vector, or a haploid AAV vector comprising at least one AAVS capsid protein (e.g., at least one of VP1, VP2, or VP3 is from the AAVS serotype), and in some embodiments, the AAV vector is a haploid AAV vector comprising at least two AAV8 capsid proteins). In some embodiments, the AAV vector comprises a capsid disclosed in W02019241324A1, or International Patent application PCT / US2019 / 036676, which are incorporated herein in their entirety by reference. In some embodiments, the AAV vector comprises a capsid which is encoded by a nucleic acid AAV capsid coding sequence that is at least 90% identical to a nucleotide sequence of any one of SEQ ID NOs: 1-3 as disclosed in W02019241324A1; or (b) a nucleotide sequence encoding any one of SEQ ID NOS:4-6 as disclosed in W02019241324A1. In some embodiments, an AAV capsid comprises an amino acid sequence at least 90% identical to any one of SEQ ID NOS:4-6 as disclosed in W02019241324A 1, along with AAV particles comprising an AAV vector genome and the AAV capsid of the invention. In some embodiments, the rAAV vector comprises capsid proteins such that the AAV vector transduces liver cells, and in some embodiments the rAAV vector comprises the rAAV vector comprises capsid proteins such that the AAV vector transduces muscle and liver cells.
[0024] An exemplary LSP encompassed for use in the methods and compositions is SP0412 (SEQ ID NO: 91) or a functional variant thereof. In altemative embodiments, a LSP can be selected from any of SEQ ID NOS: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), SEQ ID NOs: 93 (SP0239), SEQ ID NO: 94 (SP0265, also referred to SP131_A1), SEQ ID NO: 95 (SP0240) or SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240-UTRY), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-A1-UTR), or functional fragments or variants thereof.
[0025] In some embodiments of the compositions and methods described herein, the secretory signal peptide is selected from any of: AAT signal peptide, a fibronectin signal peptide (FN1), a GAA signal peptide, innate GAA leader sequence, AAT sequence, IL2(1-3), IL2 leader sequence (IL2 wt), a modified IL2 leader sequence (IL2 mut), or IgG leader sequence or functional variants thereof having secretory signal activity.
[0026] In some embodiments of the compositions and methods described herein, the alpha- glucosidase (GAA) polypeptide is linked to the IGF2 targeting peptide at the N-terminal end of a GAA polypeptide. In some embodiments, the IGF2 targeting peptide is linked to the N-terminal at amino acid 70 of human acid alpha-glucosidase (GAA) polypeptide (SEQ ID NO: 10) (i.., linked to the N-terminal of residues 70-952 of human acid alpha-glucosidase (GAA) polypeptide), or a GAA polypeptide at least 85% sequence identity to amino acids 70-952 of SEQ ID NO: 10. In alternative embodiments, the IGF2 targeting peptide is linked to the N-terminal at amino acid 40 of human acid alpha-glucosidase (GAA) polypeptide (SEQ ID NO: 10) (i.e., linked to the N-terminal of residues 40- 952 of human acid alpha-glucosidase (GAA) polypeptide) ), or a GAA polypeptide at least 85% sequence identity to amino acids 40-952 of SEQ ID NO: 10. In some embodiments of the compositions and methods described herein, the GAA polypeptide is encoded by the wild-type GAA nucleic acid sequence (e.g., SEQ ID NO: 11 or SEQ ID NO: 72), or can be a codon optimized GAA nucleic acid sequence, e.g., for any one of increasing expression in vivo, reducing CpG islands and / or reducing innate immune response in a subject. Exemplary codon optimized GAA nucleic acid sequences include, but are not limited to SEQ ID NO; 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76 and SEQ ID NO: 182.
[0027] In some embodiments of the methods and compositions disclosed herein, the recombinant AAV vector comprises a liver-specific promoter (LSP), for example but not limited to, a liver specific promoter is selected from any in Table 4 herein or functional variants thereof, or functional variants thereof. Exemplary LSP encompassed for use in the methods and compositions include SP0412 and functional variants thereof. In alternative embodiments, a LSP can comprise a nucleic acid sequence selected from any of SP0422, SP0131A1, SP0239, SP0240 or SP0246, or a functional variant thereof as disclosed herein. For Example, the liver specific promoter can comprise a nucleic acid sequence selected from any of SEQ ID NOS: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), or a functional variant or functional fragment thereof. In alternative embodiments, the liver specific promoter can comprise a nucleic acid sequence selected from any of SEQ ID NOs: 93 (SP0239), SEQ ID NO: 94 (SP0265 also called SP131_A1), SEQ ID NO: 95 (SP0240) or SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240-UTR), SEQ ID NO: 149 (SP0246-UTR), or SEQ ID NO: 150 (SP0131-A1-UTR). In some embodiments of the compositions and methods disclosed herein, a liver-specific promoter, includes a liver-specific cis-regulatory element (CRE), a synthetic liver-specific cis-regulatory module (CRM) or a synthetic liver-specific promoter comprising a promoter sequence selected from any of SEQ ID NOs: 270-341 (minimal LSP, which can include a CRM) or SEQ ID NO: 342-430(exemplary synthetic LSP), or a functional fragment or functional variant thereof, as previously disclosed in Tables 4A or 4B of provisional application 62,937,556, which is encompassed in its entirety by reference herein. These liver-specific promoter elements can include minimal liver-specific promoters (see, e.g., SEQ ID NO: 86, 270-341 or liver-specific proximal promoters (see, e.g., SEQ ID Nos: 91- 96, 146-150 and 342-430). For Example, SEQ ID NOs: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), or a functional variant or functional fragment thereof.
[0028] In some embodiments of the methods and compositions disclosed herein, the recombinant AAV vector comprises a liver-specific promoter (LSP), for example but not limited to, a liver specific promoter is selected from any of SEQ ID NOs: 86, 91-96, 146-150, 370-430 or a functional variant or functional fragment thereof.
[0029] For example, a functional variant or a functional fragment of a liver-specific promoter disclosed in Table 4 herein, or any LSP selected from SEQ ID NO: 86, 91-96, or 146-150, or 370-430 or a functional variant or functional fragment thereof has at least about 75% sequence identity to, or at least about 80% sequence identity to, at least about 90% sequence identity to, at least about 95% sequence identity to, at least about 98% sequence identity to the original unmodified reference sequence, and also at least 35% of the promoter activity, or at least about 45% of the promoter activity, or at least about 50% of the promoter activity, or at least about 60% of the promoter activity, or at least about 75% of the promoter activity, or at least about 80% of the promoter activity, or at least about 85% of the promoter activity, or at least about 90% of the promoter activity, or at least about 95% of the promoter activity of the corresponding unmodified promoter sequence.
[0030] For example, a functional variant or a functional fragment of SEQ ID NO: 92 (SP0422) or SEQ ID NO: 91 (SP0412) has at least about 75% sequence identity to SEQ ID NO: 92 or SEQ ID NO: 91, or at least about 80% sequence identity to SEQ ID NO: 92 or SEQ ID NO: 91, at least about 90% sequence identity to SEQ ID NO: 92 or SEQ ID NO: 91, at least about 95% sequence identity to SEQ ID NO: 92 or SEQ ID NO: 91, at least about 98% sequence identity to SEQ ID NO: 92 or SEQ ID NO: 91, or the original unmodified sequence, and also at least 35% of the promoter activity, or at least about 45% of the promoter activity, or at least about 50% of the promoter activity, or at least about 60% of the promoter activity, or at least about 75% of the promoter activity, or at least about 80% of the promoter activity, or at least about 85% of the promoter activity, or at least about 90% of the promoter activity, or at least about 95% of the promoter activity of the corresponding unmodified promoter sequence of SEQ ID NO: 92 or SEQ ID NO: 91, respectively.
[0031] A functional fragment is a portion of the promoter that has at least 35%, or at least about 45%, or at least about 50%, or at least about 75%, or at least about 80%, or at least about 85%, or at least about 90% of the untrunkated promoter. In some embodiments, a functional fragment comprises a contiguous portion of the unmodified promoter sequence. While TTR (SEQ ID NO: 431) is disclosed in the Examples herein as an exemplary LSP, one of ordinary skill in the art can replace the TTR promoter (SEQ ID NO: 431) with any one or more of the liver-specific promoter listed in Table 4 herein, for example, a nucleic acid sequence comprising at least SEQ ID NO: 92 (SP0422) or SEQ ID NO: 91 (SP0412) or a functional variant or fragment of SEQ ID NO: 92 (SP0422) or SEQ ID NO: 91 (SP0412), or a nucleic acid sequence comprising any of SEQ ID NOs: 93 (SP0239), SEQ ID NO: 94 (SP131_A1), SEQ ID NO: 95 (SP0240), SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265- UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240-UTR), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-A1-UTR), or any sequence selected from SEQ ID NO: 270-341 or 342-430, or a functional variant or functional fragment thereof. In some embodiments, the LSP, while preferentially expresses the hGAA protein in the liver, can also express the hGAA to some extent in another tissue of interest, e.g., the muscle, or CNS, or muscle and CNS tissues.
[0032] In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises a heterologous nucleic acid sequence that encodes a wild-type GAA polypeptide (wtGAA) or a modified GAA polypeptide, as disclosed herein, where one or more amino acids of the GAA polypeptide is modified, .g., HI99R, R223H, H201L modifications. In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises a heterologous nucleic acid sequence encoding the GAA polypeptide that is the human GAA gene or a human codon optimized GAA gene (coGAA) or a modified GAA nucleic acid sequence that is codon optimized that encodes a modified GAA polypeptide comprising one or more of the modifications selected from: HI99R, R223H, H201L. In all aspects of the methods and compositions as disclosed herein, a nucleic acid sequence encoding the GAA polypeptide is codon optimized for any one or more of: enhanced expression ix vivo, to reduce CpG islands, or to reduce the innate immune response. In all aspects of the methods and compositions as disclosed herein, a nucleic acid sequence encoding the GAA polypeptide is codon optimized to reduce CpG islands and to reduce the innate immune response. In some embodiments, the nucleic acid sequence encoding the wild type GAA polypeptide comprises modifications as disclosed in SEQ ID NO: 182, and described herein.
[0033] Another aspect of the technology herein relates to a pharmaceutical composition comprising any of the recombinant AAV vector compositions disclosed herein, and a pharmaceutically acceptable carrier.
[0034] Another aspect of the technology herein relates to a composition comprising a nucleic acid sequence comprising in the following order: a 5° ITR, a liver specific promoter (LSP) operatively linked to a nucleic acid sequence comprising a nucleic acid encoding a modified GAA polypeptide comprising one or more of the modifications selected from: HI99R, R223H, H201L, and a 3° ITR. In one aspect the nucleic acid sequence optionally further comprises a nucleic acid sequence encoding a leader sequence (or signal sequence) located between the LSP and the nucleic acid encoding the GAA polypeptide, where the leader sequence is selected from any of: the innate GAA leader sequence, AAT sequence, IL2(1-3), IL2 leader sequence (IL2 wt), a modified IL2 leader sequence (IL2 mut), fibronectin (FN1), or IgG leader sequence or functional variants thereof, as disclosed herein. In some embodiments, the nucleic acid sequence optionally further comprises a kozak sequence located between the LSP and the leader sequence. In some embodiments, the nucleic acid sequence optionally further comprises am IGF2 targeting peptide located between the leader sequence and the nucleic acid encoding the GAA polypeptide. In some embodiments, the nucleic acid sequence optionally further comprises a 3° UTR located 3” of the nucleic acid encoding the GAA polypeptide and the polyA sequence. In some embodiments, the nucleic acid sequence optionally further comprises an intron sequence 3° of the LSP and 5” of the nucleic acid encoding the GAA polypeptide, preferably between the LSP and the kozak sequence. Exemplary constructs for the rAAV vector or rAAV genome are shown in FIGS. 5A-5G.
[0035] Another aspect of the technology herein relates to a composition comprising a nucleic acid sequence comprising a 5° ITR, a liver specific promoter (LSP) operatively linked to a nucleic acid sequence encoding a modified GAA polypeptide comprising one or more of the modifications selected from: HI99R, R223H, H201L, a polyA sequence and a 3’ ITR sequence, where the poly A sequence can be a full length or truncated polyA signal sequence. Another aspect of the technology herein relates to a composition comprising a nucleic acid sequence comprising a 5° ITR, a liver specific promoter (LSP) operatively linked to a nucleic acid sequence encoding a modified GAA polypeptide comprising one or more of the modifications selected from: HI99R, R223H, H201L, a full-length polyA sequence, a terminal repeat sequence and a 3° ITR sequence, where the nucleic acid lacks a AAV P5 promoter sequence.
[0036] Another aspect of the technology herein relates to a composition comprising a nucleic acid sequence comprising: a liver specific promoter (LSP) operatively linked to a nucleic acid sequence comprising, in the following order: (a) a nucleic acid encoding a secretory signal peptide, (b) a nucleic acid encoding a IGF2 targeting peptide, and (c) a nucleic acid encoding a GAA polypeptide.
[0037] Another aspect of the technology herein relates to a composition comprising a nucleic acid sequence for a recombinant adenovirus associated (rAAV) vector genome, the nucleic acid sequence comprising: (a) a 5° and a 3° AAV inverted terminal repeats (ITR) nucleic acid sequences, and (b) located between the 5° and 3° ITR sequence, a heterologous nucleic acid sequence encoding a polypeptide comprising a secretory signal peptide and an alpha-glucosidase (GAA) polypeptide, wherein the heterologous nucleic acid is operatively linked to a liver-specific promoter as described above. An exemplary liver-specific promoter is SP0412 or SP0422 or a functional variant thereof. In some embodiments, a liver-specific promoter for use in the methods and compositions as disclosed herein includes a liver-specific cis-regulatory element (CRE), a synthetic liver-specific cis-regulatory module (CRM) or a synthetic liver-specific promoter as disclosed in Table 4 herein.
[0038] In some embodiments of the methods and compositions disclosed herein, the nucleic acid sequence comprises a heterologous nucleic acid sequence encoding a GAA polypeptide, where the nucleic acid sequence is a human GAA gene or a human codon optimized GAA gene (coGAA) or a modified GAA nucleic acid sequence. In some embodiments of the methods and compositions disclosed herein, the nucleic acid sequence comprises a heterologous nucleic acid sequence that is a codon optimized (coGAA) GAA gene, for any one or more of enhanced expression in vivo, to reduce CpG islands or to reduce the innate immune response. In some embodiments of the methods and compositions disclosed herein, the nucleic acid sequence comprises a heterologous nucleic acid sequence that is a codon optimized (coGAA) GAA gene to reduce CpG islands and to reduce the innate immune response.
[0039] In some embodiments of the methods and compositions disclosed herein, the nucleic acid sequence comprises a heterologous nucleic acid sequence encoding a GAA polypeptide selected from any of SEQ ID NO: 11 (full length hGAA), SEQ ID NO: 55 (Dwight cDNA), SEQ ID NO: 56 (hGAA A1-66) or SEQ ID NO: 182 (modGAA, H199R, R223H), or a nucleic acid sequence encoding a GAA polypeptide having the amino acid sequence of SEQ ID NO: 170 (modGAA; H199R, R223H), SEQ ID NO: 171 (modGAA; H199R, R223H, H201L), or a nucleic acid sequence encoding a GAA polypeptide that is at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to any of SEQ ID NOs: 11, 55, 56 or 182.
[0040] In some embodiments of the methods and compositions disclosed herein, the nucleic acid sequence comprises a heterologous nucleic acid sequence encoding the GAA polypeptide, where the nucleic acid encoding the GAA polypeptide is selected from any of SEQ ID NO: 74 (codon optimized 1), SEQ ID NO: 75 (codon optimized 2), and SEQ ID NO: 76 (codon optimized 3), or SEQ ID NO: 182 (modGAA, HI199R, R223H), or a nucleic acid sequence at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to any of SEQ ID NOs: 74, 75, 76 or 182.
[0041] Another aspect of the technology herein relates to use of the rAAV and nucleic acid compositions disclosed herein in a method to treat a disease. In particular, one aspect of the technology herein relates to use of the rAAV vector compositions and nucleic acid compositions disclosed herein, in a method to treat a subject with a glycogen storage disease type II (GSD II, Pompe Disease, Acid Maltase Deficiency) or having a deficiency in alpha-glucosidase (GAA) polypeptide, the method comprising administering any of the recombinant AAV vector, or the rtAAV genome or the nucleic acid sequence disclosed herein to the subject. In some embodiments of the methods disclosed herein, the expressed GAA polypeptide is secreted from the subject’s liver and there is uptake of the secreted GAA by skeletal muscle tissue, cardiac muscle tissue, diaphragm muscle tissue or a combination thereof, wherein uptake of the secreted GAA results in a reduction in lysosomal glycogen stores in the tissue(s). In some embodiments in the disclosed methods, the recombinant AAV vector, or the rAAV genome or the nucleic acid sequence is administered to the subject by any suitable administration method, for example, but not limited to, an administration method selected from any of: intramuscular, sub-cutaneous, intraspinal, intracisternal, intrathecal, intravenous administration. In some embodiments, the pharmaceutical composition disclosed herein can be used in the methods disclosed herein.
[0042] Another aspect of the technology herein relates to a cell comprising any one or more of a rAAV composition, a rAAV genome composition, or a nucleic acid composition as disclosed herein. In some embodiments, the cell is a human cell, or a non-human cell mammalian cell, or an insect cell.
[0043] Another aspect of the technology herein relates to host animal comprising any one or more of arAAV composition, a rAAV genome composition, or a nucleic acid composition as disclosed herein, In some embodiments, the host animal is a mammal, a non-human mammal or a human.
[0044] Another aspect of the technology herein relates to host animal comprising at least one cell that comprises any one or more of a rAAV composition, a rAAV genome composition, or a nucleic acid composition as disclosed herein. In some embodiments, the host animal comprising such a modified cell is a mammal, a non-human mammal or a human.
[0045] In some embodiments, disclosed herein is a pharmaceutical formulation comprising an rAAV vectors, nucleic acid encoding a rAAV genome as disclosed herein, and a pharmaceutically acceptable carrier.
[0046] Aspects of the present invention teach certain benefits in construction and use which give rise to the exemplary advantages described below. Other features and advantages of aspects of the present invention will become apparent from the following more detailed description, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of aspects of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] This application file contains at least one drawing executed in color. Copies of this patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee. The accompanying drawings illustrate aspects of the present invention. In such drawings:
[0048] FIG. 1 is a graph illustrating a y-axis of vector genomes per diploid genome and an x-axis of different AAV serotypes AAV3b, AAV3ST, AAVS, and AAVY, as measured in whole blood, in accordance with at least one embodiment.
[0049] FIG. 2 is a graph illustrating a y-axis of vector genomes per diploid genome and an x-axis of different AAV serotypes AAV3b, AAV3ST, AAVS, and AAVY, as measured in left, median and right liver lobes, in accordance with at least one embodiment.
[0050] FIGS. 3A-3B are exemplary plasmids for production of rAAV vectors useful in the methods and compositions as disclosed herein. FIG. 3A is an illustration of a plasmid map of pAAV- LSPhGAA plasmid for production of a rAAV vector in a producer cell line, e.g., a pro-10 cell line, in accordance with at least one embodiment, where the plasmid comprises a 5° ITR, LSP, hGAA nucleic acid sequence, 3° UTR, polyA sequence, and 3° ITR, where the ITRs are from AAV2. FIG. 3B shows a more detailed map of the illustration of the plasmid map of FIG. 3A.
[0051] FIGS. 4A-4G are illustrations of exemplary nucleic acid constructs for a rAAV genome as disclosed herein that have a targeting peptide, using hGAA as the exemplary lysosomal protein being expressed. FIG. 4A shows a nucleic acid construct for a rAAV genome, comprising a 5° ITR, a Liver specific promoter (LSP), operatively linked to a heterologous nucleic acid encoding a secretory signal peptide (SS), a targeting peptide (TP) and a human GAA (hGAA) polypeptide, and a 3° ITR. FIG 4B shows an exemplary nucleic acid construct for a rAAV genome as disclosed herein, comprising the same elements as FIG 4A, and additionally comprising at least one polyA signal 3° of the hGAA polypeptide and 5° of the 3°-ITR. FIG. 4C shows an exemplary nucleic acid construct for a rAAV genome as disclosed herein, comprising the same elements as FIG 4B, except comprising with an intron sequence 3’ of the promoter. FIG. 4D shows an exemplary nucleic acid construct for a rAAV genome as disclosed herein, comprising the same elements as FIG 4C, except comprising a collagen stability (CS) sequence and / or a 3° UTR sequence located 3” of the hGAA polypeptide nucleic acid sequence and before the poly A sequence. FIG. 4E shows an exemplary nucleic acid construct for a rAAV genome as disclosed herein, comprising the same elements as FIG 4D, except also comprising a nucleic acid encoding a spacer of at least 1 amino acid that is located between the nucleic acid encoding the hGAA polypeptide and the nucleic acid encoding the targeting peptide (TP), e.g., IGF2 targeting peptide. FIG. 4F shows an exemplary nucleic acid construct for a rAAV genome as disclosed herein, comprising the same elements as FIG 4E, wherein the promoter is a liver promoter, the intron sequence is selected from a MVM or HBB2 intron sequence, the secretory signal peptide is selected from any of FN1 signal peptide (e.g., hFN1, ratFN1), a AAT signal peptide or a hGAA signal peptide; the targeting peptide is a IGF2 targeting peptide as disclosed herein, and the at least polyA sequence is selected from hGHpA or a synPA poly A sequence. FIG. 4G shows an exemplary nucleic acid construct for a rAAV genome as disclosed herein, comprising the same elements as FIG 4F, except where the IGF2 targeting peptide is a nucleic acid sequence selected from SEQ ID NO: 2 (IGF2 A2-7), SEQ ID NO: 3 (IGF2 A1-7), or SEQ ID NO: 4 (IGF2 V43M).
[0052] FIG. 5A-5G shows an exemplary nucleic acid constructs for a rAAV genome. FIG. 5A isa schematic of exemplary rAAV genome comprising a 5° ITR, a liver specific promoter, operatively linked to a nucleic acid encoding a hGAA polypeptide, a polyA sequence (e.g., any one or more of hGHpA, synPA, RBG or SV40 polyA sequences) and a 3’ ITR. FIG. 5B is a schematic of an exemplary rAAV genome comprising a 5° ITR, a liver specific promoter, operatively linked to a nucleic acid encoding a signal secretory peptide (e.g., selected from any of FN1, AAT or cognate GAA signal peptide, IL2, mutIL2, IgG), a nucleic acid encoding a human GAA polypeptide and a polyA sequence and a 3° ITR. FIG. 5C is a schematic of an exemplary rAAV genome comprising a 5° ITR, a liver specific promoter, operatively linked to an intron sequence (e.g., MVM, SV40 or HBB?2 intron sequence), a nucleic acid encoding a signal secretory peptide (e.g., selected from any of FN1, AAT or cognate GAA signal peptide, IL2, mutIL2, IgG), a nucleic acid encoding a human GAA polypeptide and a polyA sequence and a 3’ ITR. FIG. 5D is a schematic of a similar construct to FIG 5C, which includes a collagen stability (CS) sequence or 3° UTR located between the 3” of the nucleic acid encoding GAA and the at least one polyA sequence (e.g., h\GHpA and / or synPA polyA sequence). In some embodiments, the construct comprses both a CS sequence and a 3 UTR sequence as disclosed herein. In some embodiments, the CS sequence can be replaced by a 3° UTR sequence as disclosed herein. In FIGS 5A-5D, exemplary liver specific promoter can be selected from any of those disclosed in Table 4 herein, and include, but are not limited to SEQ ID NOs 86, 91-96, or 146-150, or a sequence with at least 85% sequence identity to SEQ ID NOs: 86, 91-96, or 146-150. FIG. SE is a schematic of one embodiment of a AAV vector useful in the methods and compositions as disclosed herein for treating Pompe Disease, comprising, flanked between a 5° ITR and a 3° ITR sequence, the nucleic acid comprising in a 5° to 3° direction: a LSP promoter, a kozak sequence, a signal sequence (referred to as leader sequence in FIG 5E), a nucleic acid encoding hGAA and a poly A sequence. In some embodiments, the leader sequence can be selected from any of: innate GAA leader sequence, IL2 leader sequence (IL2 wt), a modified IL2 leader sequence (IL2 mut) or IgG leader sequence or functional variants thereof; and the hGAA sequence can be selected from a consensus hGAA nucleic acid sequence or a hGAA nucleic acid with at least the H201L mutation, or other modifications as disclosed herein (e.g., HI99R, R223H). FIG. 5F is a schematic of another embodiment of a AAV vector useful in the methods and compositions as disclosed herein for treating Pompe Disease, comprising in a 5° to 3’ direction: a liver specific promoter, an intron sequence, a kozak sequence, a signal sequence (also referred to as a leader sequence), an IGF2 targeting peptide sequence (referred to in FIG. 5F as a “GILT"), a nucleic acid encoding hGAA, optionally a 3° UTR sequence, and a poly A sequence, showing that different embodiments, e.g., the promoter can be selected from any LSP as disclosed herein, e.g., LSP that have different levels of expression, such as a High-expression level LSP (LSP-H), a medium expression level LSP (LSP-M) or low-expressing LSP (LSP-L), the intron sequence can be selected from HBB2, MVM, SV40 and other intron sequences, the leader sequence can be selected from any of: innate GAA leader sequence, AAT sequence (referred to as AIAT in FIG. 5F), IL2(1-3), IL2 leader sequence (IL2 wt), a modified IL2 leader sequence (IL2 mut), fibronectin (FN1, referred to as FBN in FIG. 5F), or IgG leader sequence or functional variants thereof; an IGF2 targeting peptide sequence selected from any of the IGF2 targeting peptides described herein, e.g., WT IGF2 (SEQ ID NO: 1), A2-7, V43M (SEQ ID NO: 9), A2-7V43M, or functional variants thereof, and a hGAA nucleic acid sequence that is codon optimized as disclosed herein, e.g., C1-10, which can optionally also comprise at least the H201L mutation, and / or other modifications as disclosed herein (e.g., HI99R, R223H), and a polyA sequence, selected from, e.g., RBG or SV40 polyA. The LSP designated LSP-H, M-LSP and LSP-L represent liver specific promoters that predominantly and preferentially express hGAA in the liver, but can express hGAA in one or more other tissues, for example, in the muscle. Such LSPs allows for expression in the liver for systemic secretion and uptake by the muscle cells, as well as some expression in the muscle tissues. FIG 5G shows schematic of different embodiments of a AAV vector construct useful in the methods and compositions as disclosed herein for treating Pompe Disease, where construct 1 (top panel) shows arAAV vector construct comprising in a 5° to 3’ direction, a 5° ITR, AAV P5 promoter, liver-specific promoter (LSP), hGAA nucleic acid sequence, truncated polyA sequence (t-pA), and 3° ITR; and Construct 2 (bottom panel) shows an exemplary rAAV vector construct where the P5 AAV promoter fragment is removed, the construct comprising in a 5° to 3” direction, a 5° ITR, liver-specific promoter (LSP), hGAA nucleic acid sequence, full length polyA sequence (fl-pA), a terminator sequence in the antisense orientation and 3° ITR (in the sense orientation).
[0053] FIG. 6 shows an illustration of the Gibson cloning technique to generate rAAV genomes as disclosed herein. In particular, a triple ligation is performed to ligate 3 blocks of nucleic acid sequence together, which can then be cloned into a vector with the promoter, e.g., liver specific promoter, and 5’ and 3” ITRs to generate the rAAV genome. The Gibson cloning methodology was used to generate the following rAAV genomes: SEQ ID NO: 57 (AAT-V43M-wtGAA (deltal-69aa)); SEQ ID NO: 58 (ratFN1-IGF2V43M-wtGAA (deltal-69aa)); SEQ ID NO: 59 (hFN1-IGF2V43M-wtGAA (deltal- 69aa)); SEQ ID NO: 60 (AAT-IGF2A2-7-wtGAA (delta 1-69); SEQ ID NO: 61 (FNIrat- IGFA2-7- WiGAA (delta 1-69)); SEQ ID NO: 62 (hFN1- IGFA2-7-wiGAA (delta 1-69).
[0054] FIG. 7 shows the generation of an exemplary rAAV genome of SEQ ID NO: 57 comprising AAT-V43M-wtGAA (deltal-69aa)) using Gibson cloning of nucleic acid sequence blocks (1, 2 and 3). One of ordinary skill in the art can readily replace the TTR liver promoter with any of the liver specific promoters disclosed in Table 4 herein, including but not limited to a promoter selected from any of SEQ ID NO: 86, 91-96, 146-150, or a functional variant or functional fragment thereof. Also shown in the AAT-V43M-wtGAA (deltal-69aa)) vector is the location a 3 amino acid (3aa) spacer nucleic acid sequence (showing the exemplary 3aa sequence "G-A-P" as SEQ ID NO: 31) which is located 3° of the nucleic acid sequence encoding the IGF2(V43M) targeting peptide and 5°of the nucleic acid encoding wtGAA(A1-69) enzyme, and a stuffer nucleic acid sequence (referred to in FIG 8. as a “spacer” sequence) which is located 3° of the polyA sequence and 5° of the 3’ITR sequence.
[0055] FIG. 8 shows the generation of a rAAV genome of SEQ ID NO: 62 comprising hFN1- IGFA2-7-wtGAA (delta 1-69), using Gibson cloning of nucleic acid sequence blocks (8, 2 and 3). One of ordinary skill in the art can readily replace the TTR liver promoter with any of the liver specific promoters disclosed in Table 4 herein, including but not limited to a promoter selected from any of SEQ ID NO: 86, 91-96, 146-150, or a functional variant or functional fragment thereof. Also shown in the hFN1- IGFA2-7-wtGAA (delta 1-69) vector is the location a 3 amino acid (3aa) spacer nucleic acid sequence (showing the exemplary 3aa sequence "G-A-P" as SEQ ID NO: 31) which is located 3° of the nucleic acid sequence encoding the IGFA2-7targeting peptide and 5°of the nucleic acid encoding wtGAA(A1-69) enzyme, and a stuffer nucleic acid sequence (referred to in FIG 13. as a “spacer” sequence) which is located 3” of the polyA sequence and 5° of the 3’TTR sequence.
[0056] FIGS. 9A-9F shows schematics of exemplary constructs of rAAV genomes expressing wild-type GAA. FIG. 9A shows a schematic of exemplary rAAV genome construct of Candidate 1_AAT_hIGF2-V43M_wtGAA_del1-69_Stuffer.V02 (SEQ ID NO: 79). FIG. 9B shows a schematic of exemplary rAAV genome construct of Candidate 2_FIBrat_hIGF2-V43M_wtGAA_dell- 69_Stuffer.V02 (SEQ ID NO: 80). FIG. 9C shows a schematic of exemplary rAAV genome construct of Candidate 3_FIBhum_hIGF2-V43M_wtGAA_del1-69_Stuffer.V02 (SEQ ID NO: 81) FIG. 9D shows a schematic of exemplary rAAV genome construct of Candidate 4 AAT GILT wtGAA_dell- 69__Stuffer.V02 (SEQ ID NO: 82). FIG. 9E shows a schematic of exemplary rAAV genome construct of Candidate 5_FIBrat_GILT wtGAA_del1-69_Stuffer.V02 (SEQ ID NO: 83). FIG. 9F shows a schematic of exemplary rAAV genome construct of Candidate 6_FIBhum_GILT wtGAA_dell-69_Stuffer.V02 (SEQ ID NO: 84). One of ordinary skill in the art can readily replace the TTR liver promoter shown in FIGS 9A-9F for any LPS, ¢.g., any liver specific promoters disclosed in Table 4 herein, including but not limited to a promoter selected from any of SEQ ID NO: 86, 91-96, or 146-150. Moreover, TTR promoter can be replaced with a LSP that can express the hGAA polypeptide preferentially in the liver and also in at least one other tissue of interest, e.g., the muscle, or CNS, and in some embodiments, the TTR promoter can be replaced with a LSP that can express the h\GAA polypeptide preferentially in the liver and the muscle and CNS tissues. In some embodiments, the expressed lysosomal enzyme, e.g., GAA protein can be configured as GAA-fusion protein with a targeting sequence, such as a IGF2 targeting peptide as disclosed herein that targets the GAA protein to lysosomes, and / or fused with a signal peptide (SP), the GAA protein is expressed by the rAAV genome in the liver, where it is secreted and taken up by lysosomes of mammalian cells, in particular muscle cells.
[0057] As these are exemplary constructs for illustration purposes only, one can also readily substitute the wtGAA sequence with a codon optimized sequence as disclosed herein, or a GAA sequence that has been modified to reduce CpG islands and / or to reduce innate immunity as disclosed herein (see FIG. 11B).
[0058] FIG. 10 shows the mean in vivo luciferase expression in mice driven by exemplary liver- specific promoters SP0244 and SP0239. The expression level is shown as the mean bioluminescence intensity total flux (in photons per second). Error bars are standard error of the mean. When animals are injected with saline only (n=10), no luciferase bioluminescence is detected. When animals are injected with a construct comprising luciferase operably linked to the LP1 promoter (n=9), luciferase bioluminescence is detected. To test the activity of exemplary liver-specific promoters, animals are injected with an equivalent construct comprising luciferase operably inked to the SP0244 promoter (n=8) and the SP0239 promoter (n=10). Promoters SP0244 and SP0239 showed higher luciferase expression in vivo than the control LP1.
[0059] FIGS. 11A-11D shows exemplary modifications to the nucleic acid sequence encoding the GAA polypeptide, and the nucleic acid construct to optimize for GAA protein expression by AAV in vivo. FIG. 11A shows a schematic of the wild type GAA (wtGAA) nucleotide sequence operatively linked to a liver specific promoter as disclosed herein, e.g., a LSP of Table 4, with alternative reading frames shown by the arrows, and three CpG islands. FIG. 11B shows a schematic similar to FIG. 11A that shows in more detail modifications to the nucleic acid sequence encoding GAA to remove the CpG islands. FIG. 11C shows modifications to the wtGAA nucleic acid sequence of SEQ ID NO: 182 that has been modified to (i) reduce the alternative reading frames, (ii) the number of CpG islands and (iii) to include modifications for an optimal Kozak sequence. FIG. 11D is another schematic to show modifications in the nucleic acid sequence encoding the GAA polypeptide to reduce the alternative reading frames, the number of CpG islands and to modifications for an optimal Kozak sequence.
[0060] FIG. 12 shows schematics of exemplary rAAV constructs comprising LSP for expressing GAA under liver specific promoters. The LSP can be selected from any of the liver specific promoters disclosed in Table 4 herein, with or without a stuffer sequence.
[0061] FIGS. 13A-13B show GAA expression from construct comprising liver specific promoters SP0412 and SP0422 in Huh 7 cells and HEPG2 cells. FIG. 13A shows a western blot of GAA expression from construct comprising the liver specific promoter SP0412 (SEQ ID NO: 91) and SP0422 (SEQ ID NO: 92) in Huh 7 cells. FIG. 13A shows that expression of hGAA using promoters 412 (SEQ ID NO: 91) and 422 (SEQ ID NO: 92) leads to significantly higher expression of hGAA in Huh?7 cells as compared to the expression using the LP1 promoter (SEQ ID NO: 432) which is referred to as “LSP SS”. FIG. 13B shows a western blot of GAA expression from construct comprising the liver specific promoter SP0412 (SEQ ID NO: 91) and SP0422 (SEQ ID NO: 92) in HEPG?2 cells. GAA polypeptide was expressed from rAAV generated using the following plasmids: LSP NEW (SEQ ID NO: 160), 412 NEW (SEQ ID NO: 159), TTR NEW (SEQ ID NO: 155), LSP ss (AAV with LP-1), 412 TTR, 422 Stuffer (SEQ ID NO: 158), 422 TTR, 412 Stuffer (SEQ ID NO: 156). FIG. 13B shows that expression of hGAA using promoters 412 (SEQ ID NO: 91) and 422 (SEQ ID NO: 92) leads to significantly higher expression of hGAA in HepG2 cells as compared to the expression using the LP1 promoter (SEQ ID NO: 432) which is referred to as “LSP SS”.
[0062] The above described figures illustrate aspects of the invention in at least one of its exemplary embodiments, which are further defined in detail in the following description. Features, elements, and aspects of the invention that are referenced by the same numerals in different figures represent the same, equivalent, or similar features, elements, or aspects, in accordance with one or ‘more embodiments. DETAILED DESCRIPTION
[0063] The disclosure described herein generally relates to recombinant AAV (rAAV) vectors and constructs for rAAV genomes for gene therapy for delivering a lysosomal protein, such as a GAA polypeptide to a subject. In particular, the technology described herein relates in general to a rAAV vector, or a rAAV genome for producing a lysosomal protein, e.g., GAA polypeptide that is expressed in the liver and effectively targeted to the lysosomes of mammalian cells, for example, human cardiac and skeletal muscle cells. For example, the technology relates to a rAAV vector for transducing liver cells, where the transduced liver cells secrete the GAA polypeptide, and the secreted GAA polypeptide is targeted to lysosomes in skeletal muscle tissue, cardiac muscle tissue, diaphragm muscle tissue or a combination thereof
[0064] Accordingly, one aspect of the technology described herein provides a rAAV vector comprising a rAAV genome that can be used to produce a lysosomal protein, e.g., GAA or modified GAA, that is more effectively secreted from cells, e.g., liver cells, and then targeted to the lysosomes of mammalian cells, for example, human cardiac and skeletal muscle cells.
[0065] In particular, in some embodiments, the lysosomal protein, e.g., GAA polypeptide is expressed by itself. In some embodiments, the lysosomal protein is expressed as a fusion protein comprising at least a signal peptide that promotes secretion of the lysosomal protein, e.g., GAA polypeptide from the liver. In some embodiments, the GAA polypeptide, or modified GAA, is expressed as a fusion protein comprising at least a signal peptide that promotes secretion of the GAA polypeptide from the liver, and also a targeting sequence, that allows effective targeting to lysosomes in mammalian cells, e.g., muscle cells, for example, human cardiac and skeletal muscle cells. In some embodiments, the targeting peptide is a IGF2 targeting peptide a described herein.
[0066] One aspect of the technology described herein relates to a rAAV vector that comprises a nucleotide sequence containing inverted terminal repeats (ITRs), a liver specific promoter, a heterologous gene, a poly-A tail and potentially other regulator elements for use to treat a disease, such as Pompe Disease, and further, for the treatment of Pompe Disease, wherein the heterologous gene is a GAA and wherein the rAAV GAA can be administered to a patient in a therapeutically effective dose that is delivered to the appropriate tissue and / or organ for expression of the heterologous gene and treatment of the disease.
[0067] One aspect of the technology described herein relates to a rAAV vector that comprises in its genome the following in a 5” to 3” direction: 5’- and 3°-AAV inverted terminal repeats (ITR) sequences, and located between the 5° and 3° ITRs, a heterologous nucleic acid sequence encoding an alpha-glucosidase (GAA) polypeptide, wherein the heterologous nucleic acid is operatively linked to a liver specific promoter, for example, a liver specific promoter disclosed in Table 4 herein, ora functional variant thereof. Another aspect of the technology described herein relates to a rAAV vector that comprises in its genome the following in a 5” to 3” direction: 5°- and 3°-AAYV inverted terminal repeats (ITR) sequences, and located between the 5° and 3’ ITRs, a heterologous nucleic acid sequence encoding a secretory signal peptide (SS), a nucleic acid sequence encoding an alpha- glucosidase (GAA) polypeptide, wherein the heterologous nucleic acid is operatively linked to a liver specific promoter, for example, a liver specific promoter disclosed in Table 4 herein, ora functional variant thereof.
[0068] One aspect of the technology described herein relates to a rAAV vector that comprises in its genome the following in a 5° to 3° direction: 5’- and 3°-AAV inverted terminal repeats (ITR) sequences, and located between the 5” and 3° ITRs, a heterologous nucleic acid sequence encoding a fusion polypeptide comprising (i) a secretory signal peptide (SS), (ii) an IGF2 targeting peptide; and (iii) an alpha-glucosidase (GAA) polypeptide, wherein the heterologous nucleic acid is operatively linked to a liver specific promoter, for example, a liver specific promoter disclosed in Table 4 herein, or a functional variant thereof,
[0069] In all aspects of all embodiments of the technology described herein, the liver specific promoter expresses the lysosomal protein, e.g., h\GAA polypeptide preferentially in the liver. In all aspects of all embodiments of the technology described herein, the liver specific promoter expresses the lysosomal protein e.g., hGAA polypeptide preferentially in the liver and at least one other tissue of interest, e.g., the muscle, or CNS, and in some embodiments, the LSP can be replaced with a LSP that can express the hGAA polypeptide preferentially in the liver and the muscle and CNS tissues. In all aspects of all embodiments of the technology described herein, in some embodiments where the AAV vector comprises at least one capsid protein targeting the muscle, the liver specific promoter can be replaced with another promoter, e.g., a muscle promoter.
[0070] In some embodiments of the methods and compositions as disclosed herein, the secretory signal peptide is selected from any of: AAT signal peptide, a fibronectin signal peptide (FN1), a GAA signal peptide, or an active fragment thereof having secretory signal activity.
[0071] In some embodiments, the a rAAV vector described herein is from any serotype. In some embodiments, the rAAV vector is a AAV3Db serotype, including, but not limited to, an AAV3b265D virion, an AAV3b265D549A virion, an AAV3b549A virion, an AAV3bQ263Y virion, or an AAV3bSASTG virion (i.e., a virion comprising a AAV3b capsid comprising Q263A / T265 mutations). In some embodiments, the rAAV vector comprises a liver specific capsid, e.g., a liver specific capsid selected from XL32 and XL32.1, as disclosed in W02019 / 241324, which is incorporated herein in its entirety by reference. In some embodiments, the rAAV vector is a AAVXL32 or AAVXL32.1 as disclosed in W02019 / 241324, which is incorporated herein in its entirety by reference. In some embodiments, the rAAV vector is a rAAV8 vector, or a haploid rAAV vector comprising at least one capsid protein from AAVS (i.e., any one or more of VP1, VP2 or VP3 is from AAVS or a chimeric protein thereof). In some embodiments, the AAV vector comprises a capsid disclosed in W02019241324A 1, or International Patent application PCT / US2019 / 036676, which are incorporated herein in their entirety by reference. In some embodiments, the AAV vector comprises a capsid which is encoded by a nucleic acid AAV capsid coding sequence that is at least 90% identical to a nucleotide sequence of any one of SEQ ID NOs: 1-3 as disclosed in ‘WO02019241324A1; or (b) a nucleotide sequence encoding any one of SEQ ID NOS:4-6 as disclosed in W02019241324A1. In some embodiments, an AAV capsid comprises an amino acid sequence at least 90% identical to any one of SEQ ID NOS:4-6 as disclosed in W02019241324A1, along with AAV particles comprising an AAV vector genome and the AAV capsid of the invention. In some embodiments, the rAAV vector comprises capsid proteins such that the AAV vector transduces liver cells, and in some embodiments the rAAV vector comprises the rAAV vector comprises capsid proteins such that the AAV vector transduces muscle and liver cells. In such embodiments, where the TAAV comprises capsid proteins that enable transduction of muscle cells, the LSP can be replaced with another promoter, ¢.g., a muscle promoter, or promoter that expresses a protein in liver cells and muscle cells. I. Definitions
[0072] The following terms are used in the description herein and the appended claims:
[0073] The terms “a,” “an,” “the” and similar references used in the context of describing the present invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Further, ordinal indicators — such as “first,” “second,” “third,” etc. — for identified elements are used to distinguish between the elements, and do not indicate or imply a required or limited number of such elements, and do not indicate a particular position or order of such elements unless otherwise specifically stated. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein is intended merely to better illuminate the present invention and does not pose a limitation on the scope of the invention otherwise claimed. No language in the present specification should be construed as indicating any non-claimed element essential to the practice of the invention.
[0074] Furthermore, the term "about," as used herein when referring to a measurable value such as an amount of the length of a polynucleotide or polypeptide sequence, dose, time, temperature, and the like, is meant to encompass variations oft 20%, + 10%, + 5%, + 1%, = 0.5%, or even 0.1% of the specified amount.
[0075] Also as used herein, "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").
[0076] As used herein, the transitional phrase "consisting essentially of’ means that the scope of a claim is to be interpreted to encompass the specified materials or steps recited in the claim, "and those that do not materially affect the basic and novel characteristic(s)" of the claimed invention. See, In re Herz, 537 F.2d 549, 551-52, 190 USPQ 461,463 (CCPA 1976) (emphasis in the original); see also MPEP § 2111.03. Thus, the term "consisting essentially of when used in a claim of this invention is not intended to be interpreted to be equivalent to "comprising.” Unless the context indicates otherwise, it is specifically intended that the various features of the invention described herein can be used in any combination.
[0077] Moreover, the present invention also contemplates that in some embodiments of the invention, any feature or combination of features set forth herein can be excluded or omitted.
[0078] To illustrate further, if, for example, the specification indicates that a particular amino acid can be selected from A, G, I, Land / or V, this language also indicates that the amino acid can be selected from any subset of these amino acid(s) for example A, G,IorL; A, G,Ior V; A or G; only L; etc. as if each such subcombination is expressly set forth herein. Moreover, such language also indicates that one or more of the specified amino acids can be disclaimed (e.g., by negative proviso). For example, in particular embodiments the amino acid is not A, G or I is not A; is not G or V; etc. as if each such possible disclaimer is expressly set forth herein,
[0079] The term “parvovirus” as used herein encompasses the family Parvoviridae, including autonomously replicating parvoviruses and dependoviruses. The autonomous parvoviruses include members of the genera Parvovirus, Erythrovirus, Densovirus, Iteravirus, and Contravirus. Exemplary autonomous parvoviruses include, but are not limited to, minute virus of mouse, bovine parvovirus, canine parvovirus, chicken parvovirus, feline panleukopenia virus, feline parvovirus, goose parvovirus, HI parvovirus, Muscovy duck parvovirus, B19 virus, and any other autonomous parvovirus now known or later discovered. Other autonomous parvoviruses are known to those skilled in the art. See, e.g, BERNARD N. FIELDS et al., VIROLOGY, volume 2, chapter 69 (4th ed., Lippincott-Raven Publishers).
[0080] As used herein, the term "adeno-associated virus" (AAV), includes but is not limited to, AAV type 1, AAV type 2, AAV type 3 (including types 3A and 3B), AAV type 4, AAV type 5, AAV type 6, AAV type 7, AAV type 8, AAV type 9, AAV type 10, AAV type 11, avian AAV, bovine AAV, canine AAV, equine AAV, ovine AAV, and any other AAV now known or later discovered. See, e.g., BERNARD N. FIELDS et al, VIROLOGY, volume 2, chapter 69 (4th ed., Lippincott- Raven Publishers). A number of relatively new AAV serotypes and clades have been identified (see, e.g, Gao etal. (2004) J. Virology 78:6381-6388; Moris et al., (2004) Virology 33-:375- 383); and also Table 1 as disclosed in U.S. Provisional Application 62,937,556, filed on November 19, 2019 and Table 1 in International Applications W02020 / 102645, and W02020 / 102667, each of which is incorporated herein in their entirety.
[0081] The genomic sequences of various serotypes of AAV and the autonomous parvoviruses, as well as the sequences of the native inverted terminal repeats (ITRs), Rep proteins, and capsid subunits are known in the art. Such sequences may be found in the literature or in public databases such as GenBank. See, e.g., GenBank Accession Numbers NC_002077, NC_001401, NC_001729, NC_001863, NC_001829, NC_001862, NC_000883, NC_001701, NC_001510, NC_006152, NC_006261, AF063497, U89790, AF043303, AF028705, AF028704, J02275, J01901, J02275, X01457, AF288061, AH009962, AY 028226, AY028223, NC_001358, NC_001540, AF513851, AF513852, AY530579; the disclosures of which are incorporated by reference herein for teaching parvovirus and AAV nucleic acid and amino acid sequences. See also, e.g., Srivistava et al., (1983) J Virology 45:555; Chiarini et al., (1998) J. Virology 71:6823; Chiarini et al., (1999) J. Virology 73:1309; Bantel-Schaal et al., (1999) J. Virology 73:939; Xiao et al., (1999) J. Virology 73:3994; Muramatsu et al., (1996) Virology 221:208; Shade et al., (1986) J. Viral. 58:921; Gao et al., (2002) Proc. Nat. Acad. Sci. USA 99:11854; Morris et al., (2004) Virology 33-:375- 383; international patent publications WO 00 / 28061, WO 99 / 61601, WO 98 / 11244; and U.S. Patent No. 6,156,303; the disclosures of which are incorporated by reference herein for teaching parvovirus and AAV nucleic acid and amino acid sequences. See also Table 1 and Table 5 disclosed in 62,937,556, filed on November 19, 2019 or Table 1 as disclosed in Intemational Applications W02020 / 102645, and ‘W02020 / 102667, each of which is incorporated herein in their entirety. The capsid structures of autonomous parvoviruses and AAV are described in more detail in BERNARD N. FIELDS et al., VIROLOGY, volume 2, chapters 69 & 70 (4th ed., Lippincott-Raven Publishers). See also, description of the crystal structure of AAV2 (Xie et al., (2002) Proc. Nat. Acad. Sci. 99:10405-10), AAV4 (Padron et al., (2005) J. Viral. 79: 5047-58), AAVS5 (Walters et al., (2004) J. Viral. 78: 3361- 71) and CPV (Xie et al., (1996) J. Mal. Biol. 6:497-520 and Tsao et al., (1991) Science 251: 1456-64).
[0082] The term "tropism" as used herein refers to preferential entry of the virus into certain cells or tissues, optionally followed by expression (e.g., transcription and, optionally, translation) of a sequence(s) carried by the viral genome in the cell, e.g., for a recombinant virus, expression of a heterologous nucleic acid(s) of interest.
[0083] As used here, "systemic tropism" and "systemic transduction" (and equivalent terms) indicate that the virus capsid or virus vector of the invention exhibits tropism for and / or transduces tissues throughout the body (e.g., brain, lung, skeletal muscle, heart, liver, kidney and / or pancreas). In embodiments of the invention, systemic transduction of the central nervous system (e.g., brain, neuronal cells, etc.) is observed. In other embodiments, systemic transduction of cardiac muscle tissues is achieved.
[0084] As used herein, "selective tropism" or "specific tropism" means delivery of virus vectors to and / or specific transduction of certain target cells and / or certain tissues.
[0085] Unless indicated otherwise, “efficient transduction” or “efficient tropism,” or similar terms, can be determined by reference to a suitable control (e.g., at least about 50%, 60%, 70%, 80%, 85%, 90%, 95%, 100%, 125%, 150%, 175%, 200%, 250%, 300%, 350%, 400%, 500% or more of the transduction or tropism, respectively, of the control). In particular embodiments, the virus vector efficiently transduces or has efficient tropism for liver cells and muscle cells. Suitable controls will depend on a variety of factors including the desired tropism and / or transduction profile.
[0086] Similarly, it can be determined if a virus “does not efficiently transduce” or “does not have efficient tropism” for a target tissue, or similar terms, by reference to a suitable control. In particular embodiments, the virus vector does not efficiently transduce (i.¢., has does not have efficient tropism) for kidney, gonads and / or germ cells. In particular embodiments, transduction (e.g., undesirable transduction) of tissue(s) (e.g., kidney) is 20% or less, 10% or less, 5% or less, 1% or less, 0.1% or less of the level of transduction of the desired target tissue(s) (e.g., liver, skeletal muscle, diaphragm muscle, cardiac muscle and / or cells of the central nervous system).
[0087] In some embodiments of this invention, an AAV particle comprising a capsid of this invention can demonstrate multiple phenotypes of efficient transduction of 30 certain tissues / cells and very low levels of transduction (e.g., reduced transduction) for certain tissues / cells, the transduction of which is not desirable.
[0088] As used herein, the term "polypeptide" encompasses both peptides and proteins, unless indicated otherwise.
[0089] A "polynucleotide" is a sequence of nucleotide bases, and may be RNA, DNA or DNA- RNA hybrid sequences (including both naturally occurring and non-naturally occurring nucleotides), but in representative embodiments are either single or double stranded DNA sequences.
[0090] The terms “heteralogous nucleotide sequence” and “heterologous nucleic acid molecule” are used interchangeably herein and refer to a nucleic acid sequence that is not naturally occurring in the virus. Generally, the heterologous nucleic acid molecule or heterologous nucleotide sequence comprises an open reading frame that encodes a polypeptide and / or nontranslated RNA of interest {e.g., for delivery to a cell and / or subject).
[0091] A “chimeric nucleic acid” comprises two or more nucleic acid sequences covalently linked together to encode a fusion polypeptide. The nucleic acids may be DNA, RNA, or a hybrid thereof.
[0092] The term “fusion polypeptide” comprises two or more polypeptides covalently linked together, typically by peptide bonding.
[0093] As used herein, an "isolated" polynucleotide (e.g., an "isolated DNA" or an "isolated RNA") means a polynucleotide at least partially separated from at least some of the other components of the naturally occurring organism or virus, for example; the cell or viral structural components or other polypeptides or nucleic acids commonly found associated with the polynucleotide. In representative embodiments an "isolated" nucleotide is enriched by at least about 10-fold, 100'-fold, 1000-fold, 10,000-fold or more as compared with the starting material.
[0094] Likewise, an "isolated" polypeptide means a polypeptide that is at least partially separated from at least some of the other components of the naturally occurring organism or virus, for example, the cell or viral structural components or other polypeptides or nucleic acids commonly found associated with the polypeptide. In representative embodiments an "isolated" polypeptide is enriched by at least about 10-fold, 100-fold, 1000-fold, 10,000-fold or more as compared with the starting material
[0095] An "isolated cell" refers to a cell that is separated from other components with which it is normally associated in its natural state. For example, an isolated cell can be a cell in culture medium and / or a cell in a pharmaceutically acceptable carrier of this invention. Thus, an isolated cell can be delivered to and / or introduced into a subject. In some embodiments, an isolated cell can be a cell that is removed from a subject and manipulated as described herein ex vivo and then retumed to the subject.
[0096] A population of virions can be generated by any of the methods described herein. In one embodiment, the population is at least 101 virions. In one embodiment, the population is at least 102 virions, at least 103, virions, at least 104 virions, at least 105 virions, at least 106 virions, at least 107 virions, at least 108 virions, at least 109 virions, at least 1010 virions, at least 1011 virions, at least 1012 virions, at least 1013 virions, at least 1014 virions, at least 1015 virions, at least 1016 virions, or at least 1017 virions. A population of virions can be heterogeneous or can be homogeneous (e.g., substantially homogeneous or completely homogeneous).
[0097] A “substantially homogeneous population” as the term is used herein, refers to a population of virions that are mostly identical, with few to no contaminant virions (those that are not identical) therein. A substantially homogeneous population is at least 90% of identical virions (e.g., the desired virion), and can be at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% of identical virions.
[0098] A population of virions that is completely homogeneous contains only identical virions.
[0099] As used herein, by "isolate" or "purify" (or grammatical equivalents) a virus vector or virus particle or population of virus particles, it is meant that the virus vector or virus particle or population of virus particles is at least partially separated from at least some of the other components in the starting material. In representative embodiments an "isolated" or "purified" virus vector or virus particle or population of virus particles is enriched by at least about 10-fold, 100-fold, 1000-fold, 10,000-fold or more as compared with the starting material.
[00100] Unless indicated otherwise, "efficient transduction" or "efficient tropism," or similar terms, can be determined by reference to a suitable control (e.g., at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 125%, 150%, 175%, 200%, 250%, 300%, 350%, 400%, 500% or more of the transduction or tropism, respectively, of the control). In particular embodiments, the virus vector efficiently transduces or has efficient tropism for neuronal cells and cardiomyocytes. Suitable controls will depend on a variety of factors including the desired tropism and / or transduction profile.
[00101] A "therapeutic polypeptide" is a polypeptide that can alleviate, reduce, prevent, delay and / or stabilize symptoms that result from an absence or defect in a protein in a cell or subject and / or is a polypeptide that otherwise confers a benefit to a subject, e.g., enzyme replacement to reduce or eliminate symptoms of a disease, or improvement in transplant survivability or induction of an immune response.
[00102] The terms "heterologous nucleotide sequence" and "heterologous nucleic acid molecule" are used interchangeably herein and refer to a nucleic acid sequence that is not naturally occurring in the virus. Generally, the heterologous nucleic acid molecule or heterologous nucleotide sequence comprises an open reading frame that encodes a polypeptide and / or nontranslated RNA of interest (e.g., for delivery to a cell and / or subject), for example the GAA polypeptide.
[00103] As used herein, the terms "virus vector," "vector" or "gene delivery vector" refer to a virus (e.g.. AAV) particle that functions as a nucleic acid delivery vehicle, and which comprises the vector genome (e.g., viral DNA [VDNA]) packaged within a virion. Alternatively, in some contexts, the term "vector" may be used to refer to the vector genome / vDNA alone.
[00104] An "rAAYV vector genome" or "rAAV genome" is an AAV genome (i.e., vDNA) that comprises one or more heterologous nucleic acid sequences. AAV vectors generally require only the inverted terminal repeat(s) (TR(s)) in cis to generate virus. All other viral sequences are dispensable and may be supplied in trans (Muzyczka, (1992) Curr. Topics Microbial. Immunol. 158:97). Typically, the rAAV vector genome will only retain the one or more TR sequence so as to maximize the size of the transgene that can be efficiently packaged by the vector. The structural and non- structural protein coding sequences may be provided in trans (e.g., from a vector, such as a plasmid, or by stably integrating the sequences into a packaging cell). In embodiments of the invention the rAAV vector genome comprises at least one ITR sequence (e.g., AAV TR sequence), optionally two ITRs (e.g., two AAV TRs), which typically will be at the 5' and 3' ends of the vector genome and flank the heterologous nucleic acid, but need not be contiguous thereto. The TRs can be the same or different from each other.
[00105] The term "terminal repeat" or "TR" includes any viral terminal repeat or synthetic sequence that forms a hairpin structure and functions as an inverted terminal repeat (i.e., an ITR that mediates the desired functions such as replication, virus packaging, integration and / or provirus rescue, and the like). The TR can be an AAV TR or a non-AAV TR. For example, a non-AAV TR sequence such as those of other parvoviruses (e.g., canine parvovirus (CPV), mouse parvovirus (MVM), human parvovirus B-19) or any other suitable virus sequence (e.g., the SV40 hairpin that serves as the origin of SV40 replication) can be used as a TR, which can further be modified by truncation, substitution, deletion, insertion and / or addition. Further, the TR can be partially or completely synthetic, such as the "double-D sequence" as described in United States Patent No. 5,478,745 to Samulski et al.
[00106] An "AAV terminal repeat" or "AAV TR," including an “AAV inverted terminal repeat” or “AAV ITR” may be from any AAV, including but not limited to serotypes 1,2, 3,4, 5,6, 7, 8, 9, 10, 11 or 12 or any other AAV now known or later discovered. An AAV terminal repeat need not have the native terminal repeat sequence (e.g., a native AAV TR or AAV ITR sequence may be altered by insertion, deletion, truncation and / or missense mutations), as long as the terminal repeat mediates the desired functions, e.g., replication, virus packaging, integration, and / or provirus rescue, and the like.
[00107] AAV proteins VP1, VP2 and VP3 are capsid proteins that interact together to form an AAV capsid of an icosahedral symmetry. VP1.5 is an AAV capsid protein described in US Publication No. 2014 / 0037585.
[00108] The virus vectors of the invention can further be "targeted" virus vectors (e.g.. having a directed tropism) and / or a "hybrid" parvovirus (i.e., in which the viral TRs and viral capsid are from different parvoviruses) as described in intemational patent publication WO 00 / 28004 and Chao et al., (2000) Molecular Therapy 2:619.
[00109] The virus vectors of the invention can further be duplexed parvovirus particles as described in international patent publication WO 01 / 92551 (the disclosure of which is incorporated herein by reference in its entirety). Thus, in some embodiments, double stranded (duplex) genomes can be packaged into the virus capsids of the invention,
[00110] Further, the viral capsid or genomic elements can contain other modifications, including insertions, deletions and / or substitutions.
[00111] A "chimeric' capsid protein as used herein means an AAV capsid protein (e.g., any one or more of VP1, VP2 or VP3) that has been modified by substitutions in one or more (e.g, 2, 3, 4, 5, 6, 7,8, 9, 10, etc.) amino acid residues in the amino acid sequence of the capsid protein relative to wild type, as well as insertions and / or deletions of one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.) amino acid residues in the amino acid sequence relative to wild type. In some embodiments, complete or partial domains, functional regions, epitopes, etc., from one AAV serotype can replace the corresponding wild type domain, functional region, epitope, etc. of a different AAV serotype, in any combination, to produce a chimeric capsid protein of this invention. Production of a chimeric capsid protein can be carried out according to protocols well known in the art and a significant number of chimeric capsid proteins are described in the literature as well as herein that can be included in the capsid of this invention,
[00112] As used herein, the term “haploid AAV” shall mean that AAV as described in International Application W02018 / 170310, or US Application US2018 / 037149, which are incorporated herein in their entirety by reference. In some embodiments, a population of virions is a haploid AAV population where a virion particle can be constructed wherein at least one viral protein from the group consisting of AAV capsid proteins, VP1, VP2 and VP3, is different from at least one of the other viral proteins, required to form the virion particle capable of encapsulating an AAV genome. For each viral protein present (VP1, VP2, and / or VP3), that protein is the same type (e.g., all AAV2 VP1). In one instance, at least one of the viral proteins is a chimeric viral protein and at least one of the other two viral proteins is not a chimeric. In one embodiment VP1 and VP2 are chimeric and only VP3 is non- chimeric. For example, only the viral particle composed of VP1 / VP2 from the chimeric AAV2 / 8 (the N-terminus of AAV? and the C-terminus of AAV8) paired with only VP3 from AAV2; or only the chimeric VP1 / VP2 28m-2P3 (the N-terminal from AAV8 and the C-terminal from AAV2 without mutation of VP3 start codon) paired with only VP3 from AAV2. In another embodiment only VP3 is chimeric and VP1 and VP2 are non-chimeric. In another embodiment at least one of the viral proteins is from a completely different serotype. For example, only the chimeric VP1 / VP2 28m-2P3 paired with VP3 from only AAV3. In another example, no chimeric is present,
[00113] The term a "hybrid" AAV vector or parvovirus refers to a rAAV vector where the viral TRs or ITRs and viral capsid are from different parvoviruses. Hybrid vectors are described in international patent publication WO 00 / 28004 and Chao et al., (2000) Molecular Therapy 2:619. For example, a hybrid AAV vector typically comprises the adenovirus 5' and 3' cis ITR sequences sufficient for adenovirus replication and packaging (i.c., the adenovirus terminal repeats and PAC sequence).
[00114] The term “polyploid AAV” refers to a AAV vector which is composed of capsids from two or more AAV serotypes, €.g., and can take advantages from individual serotypes for higher transduction but not in certain embodiments eliminate the tropism from the parents.
[00115] The term “GAA” or “GAA polypeptide,” as used herein, encompasses mature (“76 or “67 kDa) and precursor (e.g., "110 kDa) GAA as well as modified (e.g., truncated or mutated by insertion(s), deletion(s) and / or substitution(s)) GAA proteins or fragments thereof that retain biological function (i.e., have at least one biological activity of the native GAA protein, e.g., can hydrolyze glycogen, as defined above) and GAA variants (e.g., GAA II as described by Kunita et al., (1997) Biochemica et Biophysica Acta 1362:269; GAA polymorphisms and SNPs are described by Hirschhorn, R. and Reuser, A. J. (2001) in The Metabolic and Molecular Basis for Inherited Disease (Scriver, C. R., Beaudet. A. L., Sly, W. S. & Valle, D. Eds.), pp. 3389-3419, McGraw-Hill, New York, see pages 3403-3403; each incorporated herein by reference in its entirety). Any GAA coding sequence known in the art may be used, for example, see the coding sequences of FIGS. 8 and 9; GenBank Accession number NM_00152 and Hoefsloot et al., (1988) EMBO J. 7:1697 and Van Hove etal., (1996) Proc. Natl. Acad. Sci. USA 93:65 (human), GenBank Accession number NM_008064 (mouse), and Kunita et al., (1997) Biochemica et Biophysics Acta 1362:269 (quail); the disclosures of which are incorporated herein by reference for their teachings of GAA coding and noncoding sequences.
[00116] The terms “cation-independent mannose-6-phosphate receptor (CI-MPR),” “M6P / IGF-1I receptor,” “CI-MPR / IGF-II receptor,” “IGF-II receptor” or “IGF2 Receptor,” or abbreviations thereof, are used interchangeably herein, referring to the cellular receptor which binds both M6P and IGF-II.
[00117] The term “targeting peptide” is also referred to as a “targeting sequence” as used herein is intended to refer to a peptide that targets a particular subcellular compartment, for example, a mammalian lysosome. A targeting peptide encompassed for use herein is a lysosome targeting peptide that is mannose-6-phosphate-independent. An exemplary targeting sequence is an IGF2 targeting peptide as disclosed herein.
[00118] The term “IGF2 sequence” is used in conjunction with “IGF2 targeting sequence” or “IGF2 leader sequence” and “IGF2 targeting peptide” are used interchangeably herein and refer to a sequence of the IGF2 polypeptide that binds to the CI-MBR on the surface of the cell. In particular, the IGF2 sequence is a peptide that comprises a part of the IGF2 uptake sequence of SEQ ID NO: 5, or comprises a modification in amino acid of SEQ ID NO:5. An IGF2 targeting peptide refers to a peptide sequence that binds to a receptor domain consisting essentially of repeats 11-12, repeat 11 or amino acids 1508-1566 of the human cation-independent mannose-6-phosphate receptor (CI-MPR or CA-MB6P receptor).
[00119] The term “leader sequence” is used interchangeably herein with the term “secretory signal sequence” or “signal sequence” or “signal peptide” or variations thereof, and intended to refer to amino acid sequences that function to enhance (as defined above) secretion of an operably linked polypeptide, (e.g., a GAA peptide or IGF2-GAA fusion protein) from the cell as compared with the level of secretion seen with the native polypeptide. As defined above, by “enhanced” secretion, it is meant that the relative proportion of lysosomal polypeptide synthesized by the cell that is secreted from the cell is increased, it is not necessary that the absolute amount of secreted protein is also increased. In particular embodiments of the invention, essentially all (i.e., at least 95%, 97%, 98%, 99% or more) of the GAA-polypeptide is secreted. It is not necessary, however, that essentially all or even most of the GAA polypeptide is secreted, as long as the level of secretion is enhanced as compared with the native GAA polypeptide. Exemplary leader sequences include, but are not limited to the innate GAA leader sequence (also referred to cognate GAA leader sequence), AAT sequence, IL2(1-3), IL2 leader sequence (IL2 wt), a modified IL2 leader sequence (IL2 mut), fibronectin (FN1, also referred to as FBN), or IgG leader sequence or functional variants thereof, as disclosed herein.
[00120] As used herein, the term "amino acid" encompasses any naturally occurring amino acid, modified forms thereof, and synthetic amino acids. Naturally occurring, levorotatory (L-) amino acids are disclosed in Table 2 of US Publication 2018 / 0371496, which is incorporated herein in its entirety. Alternatively, the amino acid can be a modified amino acid residue (nonlimiting examples are shown in Table 4 of US Publication of US Publication 2018 / 0371496) and / or can be an amino acid that is modified by post-translation modification (e.g., acetylation, amidation, formylation, hydroxylation, methylation, phosphorylation or sulfatation). Further, the non-naturally occurring amino acid can be an “unnatural” amino acid as described by Wang et al., Annu Rev Biophys Biomol Struct. 35:225-49 (2006). These unnatural amino acids can advantageously be used to chemically link molecules of interest to the AAV capsid protein.
[00121] To illustrate further, if, for example, the specification indicates that a particular amino acid can be selected from A, G, I, L and / or V, this language also indicates that the amino acid can be selected from any subset of these amino acid(s) for example A, G,IorL; A, G, Lor V; A or G; only L; etc. as if each such subcombination is expressly set forth herein. Moreover, such language also indicates that one or more of the specified amino acids can be disclaimed (e.g., by negative proviso). For example, in particular embodiments the amino acid is not A, G or I is not A; is not G or V; etc. as if each such possible disclaimer is expressly set forth herein,
[00122] The term “cis-regulatory element” or “CRE”, is a term well-known to the skilled person, and means a nucleic acid sequence such as an enhancer, promoter, insulator, or silencer, that can regulate or modulate the transcription of a neighboring gene (i.e. in cis). CREs are found in the vicinity of the genes that they regulate. CREs typically regulate gene transcription by binding to transcription factors (TFs), i.e. they include TF binding site (TFBS). A single TF may bind to many CREs, and hence control the expression of many genes (pleiotropy). CREs are usually, but not always, located upstream of the transcription start site (TSS) of the gene that they regulate. “Enhancers” are CREs that enhance (i.e. upregulate) the transcription of genes that they are operably associated with, and can be found upstream, downstream, and even within the introns of the gene that they regulate. Multiple enhancers can act in a coordinated fashion to regulate transcription of one gene. “Silencers” in this context relates to CREs that bind TFs called repressors, which act to prevent or downregulate transcription of a gene. The term "silencer" can also refer to a region in the 3' untranslated region of messenger RNA, that bind proteins which suppress translation of that mRNA molecule, but this usage is distinct from its use in describing a CRE. Generally, the CREs of the present invention are liver-specific enhancers (often referred to as liver-specific CREs, or liver-specific CRE enhancers, or suchlike). In the present context, it is preferred that the CRE is located 1500 nucleotides or less from the transcription start site (TSS), more preferably 1000 nucleotides or less from the TSS, more preferably 500 nucleotides or less from the TSS, and suitably 250, 200, 150, or 100 nucleotides or less from the TSS. CREs of the present invention are preferably comparatively short in length, preferably 100 nucleotides or less in length, for example they may be 90, 80, 70, 60 nucleotides or less in length.
[00123] The term “cis-regulatory module” or “CRM” means a functional module made up of two or more CRESs; in the present invention the CRESs are typically liver-specific enhancers. Thus, in the present application a CRM typically comprises a plurality of liver-specific enhancer CREs. Typically, the multiple CREs within the CRM act together (e.g. additively or synergistically) to enhance the transcription of a gene that the CRM is operably associated with. There is conservable scope to shuffle (i.e. reorder), invert (i.e. reverse orientation), and alter spacing in CREs within a CRM. Accordingly, functional variants of CRMs of the present invention include variants of the referenced CRMs wherein CREs within them have been shuffled and / or inverted, and / or the spacing between CREs has been altered.
[00124] As used herein, the phrase "promoter" refers to a region of DNA that generally is located upstream of a nucleic acid sequence to be transcribed that is needed for transcription to occur, i.e. which initiates transcription. Promoters permit the proper activation or repression of transcription of a coding sequence under their control. A promoter typically contains specific sequences that are recognized and bound by plurality of TFs. TFs bind to the promoter sequences and result in the recruitment of RNA polymerase, an enzyme that synthesizes RNA from the coding region of the gene. A great many promoters are known in the art.
[00125] The term “synthetic promoter” as used herein relates to a promoter that does not occur in nature. In the present context it typically comprises a synthetic CRE and / or CRM of the present invention operably linked to a minimal (or core) promoter or liver-specific proximal promoter. The CREs and / or CRMs of the present invention serve to enhance liver-specific transcription of a gene operably linked to the promoter. Parts of the synthetic promoter may be naturally occurring (e.g. the minimal promoter or one or more CREs in the promoter), but the synthetic promoter as a complete entity is not naturally occurring.
[00126] As used herein, “minimal promoter" (also known as the “core promoter”) refers to a short DNA segment which is inactive or largely inactive by itself, but can mediate transcription when combined with other transcription regulatory elements. Minimum promoter sequence can be derived from various different sources, including prokaryotic and eukaryotic genes. Examples of minimal promoters are discussed above, and include the dopamine beta-hydroxylase gene minimum promoter, cytomegalovirus (CMV) immediate early gene minimum promoter (CMV-MP), and the herpes thymidine kinase minimal promoter (MinTK). A minimal promoter typically comprises the transcription start site (TSS) and elements directly upstream, a binding site for RNA polymerase II, and general transcription factor binding sites (often a TATA box).
[00127] As used herein, “proximal promoter” relates to the minimal promoter plus the proximal sequence upstream of the gene that tends to contain primary regulatory elements. It often extends approximately 250 base pairs upstream of the TSS, and includes specific TFBS. In the present case, the proximal promoter is suitably a naturally occurring liver-specific proximal promoter that can be combined with one or more CREs or CRMs of the present invention. However, the proximal promoter can be synthetic.
[00128] A “functional variant” of a cis-regulatory element (CRE), cis-regulatory module (CRM), promoter or other nucleic acid sequence in the context of the present invention is a variant of a reference sequence that retains the ability to function in the same way as the reference sequence, e.g. as a liver-specific cis-regulatory enhancer element, liver-specific cis-regulatory module or liver- specific promoter. Altemative terms for such functional variants include “biological equivalents” or “equivalents”.
[00129] It will be appreciated that the ability of a given cis-regulatory element to function as a liver- specific enhancer is determined principally by the ability of the sequence to bind the same liver- specific transcription factors (TFs) that bind to the reference sequence. Accordingly, in most cases, a functional variant of a cis-regulatory element will contain TFBS for the same TFs as the reference cis- regulatory element. It is preferred, but not essential, that the transcription factor binding site (TFBS) of a functional variant are in the same relative positions (i.e. order) as the reference cis-regulatory element. It is also preferred, but not essential, that the TFBS of a functional variant are in the same orientation as the reference sequence (it will be noted that TFBS can in some cases be present in reverse orientation, e.g. as the reverse complement vis-a-vis the sequence in the reference sequence). It is also preferred, but not essential, that the TFBS of a functional variant are on the same strand as the reference sequence. Thus, in preferred embodiments, the functional variant comprises TFBS for the same TFs, in the same order, in the same orientation and on the same strand as the reference sequence. It will also be appreciated that the sequences lying between TFBS (referred to in some cases as spacer sequences, or suchlike) are of less consequence to the function of the cis-regulatory element. Such sequences can typically be varied considerably, and their lengths can be altered. However, in preferred embodiments the spacing (i.e. the distance between adjacent TFBS) is substantially the same (e.g. it does not vary by more than 20, preferably by not more than 10%, more preferably it is the same) in a functional variant as it is in the reference sequence. It will be apparent that in some cases a functional variant of a cis-regulatory enhancer element can be present in the reverse orientation, e.g. it can be the reverse complement of a cis-regulatory enhancer element as described above, or a variant thereof.
[00130] Levels of sequence identity between a functional variant and the reference sequence can also be an indicator or retained functionality. High levels of sequence identity in the TFBS of the cis- regulatory element is of generally higher importance than sequence identity in the spacer sequences (where there is little or no requirement for any conservation of sequence). However, it will be appreciated that even within the TFBS, a considerable degree of sequence variation can be accommodated, given that the sequence of a functional TFBS does not need to exactly match the consensus sequence.
[00131] The ability of one or more TFs to bind to a TFBS in a given functional variant can determined by any relevant means known in the art, including, but not limited to, electromobility shift assays (EMSA), binding assays, chromatin immunoprecipitation (ChIP), and ChIP-sequencing (ChIP-seq). In a preferred embodiment the ability of one or more TFs to bind a given functional variant is determined by EMSA. Methods of performing EMSA are well-known in the art. Suitable approaches are described in Sambrook et al. cited above. Many relevant articles describing this procedure are available, e.g. Hellman and Fried, Nat Protoc. 2007; 2(8): 1849-1861.
[00132] The terms “liver-specific” or “liver-specific expression” when in reference to a promoter refers to the ability of a cis-regulatory element, cis-regulatory module or promoter to enhance or drive expression of a gene in the liver (or in liver-derived cells) in a preferential or predominant manner as compared to other tissues (e.g. spleen, muscle, heart, lung, and brain). Expression of the gene can be in the form of mRNA or protein. In some embodiments, liver-specific expression is such that there is negligible expression in other (i.e. non-liver) tissues or cells, i.e. expression is highly liver-specific. In some embodiments, while a liver-specific promoter drives expression preferentially in the liver, it can also drive expression of the gene in another tissue of interest at a lower level, e.g., muscle.
[00133] The ability of a cis-regulatory element to function as a liver-specific cis-regulatory enhancer element can be readily assessed by the skilled person. The skilled person can thus easily determine whether any variant of the specific cis-regulatory elements recited above remains functional (i.e. itis a functional variant as defined above). For example, any given cis-regulatory element to be assessed can be operably linked to a minimal promoter (e.g. positioned upstream of CMV-MP) and the ability of the cis-regulatory element to drive liver-specific expression of a gene (typically a reporter gene) is measured. Alternatively, a variant of a cis-regulatory enhancer element can be substituted into a synthetic liver-specific promoter in place of a reference cis-regulatory enhancer element, and the effects on liver-specific expression driven by said modified promoter can be determined and compared to the unmodified form. Similarly, the ability of a cis-regulatory module or promoter to drive liver-specific expression can be readily assessed by the skilled person (e.g. as described in the examples below). Expression levels of a gene driven by a variant of a reference promoter can be compared to the expression levels driven by the reference sequence. In some embodiments, where liver-specific expression levels driven by a variant promoter are at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100% of the expression levels driven by the reference promoter, it can be said that the variant remains functional. Suitable nucleic acid constructs and reporter assays to assess liver-specific expression enhancement can easily be constructed, and the examples set out below give suitable methodologies.
[00134] Liver-specificity can be identified wherein the expression of a gene (e.g. a therapeutic or reporter gene) occurs preferentially or predominantly in liver-derived cells. Preferential or predominant expression can be defined, for example, where the level of expression is significantly greater in liver-derived cells than in other types of cells (i.e. non-liver-derived cells). For example, expression in liver-derived cells is suitably at least 5-fold higher than non-liver cells, preferably at least 10-fold higher than non-liver cells, and it may be 50-fold higher or more in some cases. For convenience, liver-specific expression can suitably be demonstrated via a comparison of expression levels in a hepatic cell line (e.g. liver-derived cell line such as Huh7 and / or HepG2 cells) or liver primary cells, compared with expression levels in a kidney-derived cell line (e.g. HEK-293), a cervical tissue-derived cell line (e.g. HeLa) and / or a lung-derived cell line (e.g. A549).
[00135] The synthetic liver-specific promoters of the present invention are preferably suitable for promoting expression in the liver of a subject, e.g. driving liver-specific expression of a transgene, preferably a therapeutic transgene. In some embodiments, the liver-specific promoters of the invention are suitable for promoting liver-specific transgene expression at a level at least 1.5-fold greater than the LP1 promoter of SEQ ID NO: 432, preferably 2-fold greater than the LP1 promoter, more preferably 3-fold greater than the LP1 promoter, and yet more preferably 5-fold greater than the LPI promoter (SEQ ID NO: 432). Such expression is suitably determined in liver-derived cells, e.g. in Huh7, and / or HepG?2 cells or primary liver cells (suitably primary human hepatocytes). In some embodiments, the synthetic liver-specific promoters of the present invention are suitable for promoting gene expression at a level of at least 1.5-fold less than an LP1 promoter (SEQ ID NO: 432) in non-liver-derived cells (e.g. HEK-293, HeLa, and / or A549 cells).
[00136] Preferred synthetic liver-specific promoters of the present invention are suitable for promoting liver-specific transgene expression and have an activity in liver cells which is at least 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 125%, 150%, 175%, 200%, 250%, 300%, 350% or 400% of the activity of the TBG promoter (SEQ ID NO: 435).
[00137] The synthetic liver-specific promoters of the present invention are preferably suitable for promoting liver-specific expression at a level at least 1.5-fold greater than a CMV-IE promoter of SEQ ID NO: 433 in liver-derived cells, preferably at least 2-fold greater than a CMV promoter in liver-derived cells (e.g. HEK-293, HeLa, and / or A549 cells). The synthetic liver specific promoters disclosed herein can be LSP-H, LSP-M and LSP-L promoters, referring to high, medium and low expression in the liver, and in some embodiments, the LSP-H, LSP-M and LSP-L can preferentially or predominantly express a protein in the liver, but can also express the protein on one or more other tissues, for example, in the muscle and / or brain. Such LSP-H, LSP-M and LSP-L promoters disclosed herein can preferentially express at least 90%, or at least 80%, or at least 70% or at least 60%, or at least 50% of a protein in the liver, and also express at least 10%, or at least 20%, or at least 30%, or at least 40% or at least 50% in another tissue, for example, in muscle tissue. In some embodiments, a LSP-H, LSP-M and LSP-L promoter useful in the method and compositions as disclosed herein, for example for the treatment of Pompe or a lysosomal disease drives or enhances gene expression in a preferential or predominant manner in the liver, but can also express at least some of the protein in muscle tissue.
[00138] The terms "identity" and "identical" and the like refer to the sequence similarity between two polymeric molecules, e.g., between two nucleic acid molecules, such as between two DNA molecules. Sequence alignments and determination of sequence identity can be done, e.g., using the Basic Local Alignment Search Tool (BLAST) originally described by Altschul et al. 1990 (J Mol Biol 215: 403- 10), such as the "Blast 2 sequences" algorithm described by Tatusova and Madden 1999 (FEMS Microbiol Lett 174; 247-250).
[00139] The term “synthetic” as used herein means a nucleic acid molecule that does not occur in nature. Synthetic nucleic acid expression constructs of the present invention are produced artificially, typically by recombinant technologies. Such synthetic nucleic acids may contain naturally occurring sequences (e.g. promoter, enhancer, intron, and other such regulatory sequences), but these are present in a non-naturally occurring context. For example, a synthetic gene (or portion of a gene) typically contains one or more nucleic acid sequences that are not contiguous in nature (chimeric sequences), and / or may encompass substitutions, insertions, and deletions and combinations thereof.
[00140] A “spacer sequence” or “spacer” as used herein is a nucleic acid sequence that separates two functional nucleic acid sequences. It can have essentially any sequence, provided it does not prevent the functional nucleic acid sequence (e.g. cis-regulatory element) from functioning as desired (e.g. this could happen if it includes a silencer sequence, prevents binding of the desired transcription factor, or suchlike). Typically, it is non-functional, as in it is present only to space adjacent functional nucleic acid sequences from one another.
[00141] The term "pharmaceutically acceptable” as used herein is consistent with the art and means compatible with the other ingredients of the pharmaceutical composition and not deleterious to the recipient thereof.
[00142] By the terms "treat," "treating" or "treatment of (and grammatical variations thereof) it is meant that the severity of the subject's condition is reduced, at least partially improved or stabilized and / or that some alleviation, mitigation, decrease or stabilization in at least one clinical symptom is achieved and / or there is a delay in the progression of the disease or disorder.
[00143] The terms "prevent," "preventing" and "prevention" (and grammatical variations thereof) refer to prevention and / or delay of the onset of a disease, disorder and / or a clinical symptom(s) in a subject and / or a reduction in the severity of the onset of the disease, disorder and / or clinical symptomy(s) relative to what would occur in the absence of the methods of the invention. The prevention can be complete, e.g., the total absence of the disease, disorder and / or clinical symptom(s). The prevention can also be partial, such that the occurrence of the disease, disorder and / or clinical symptomy(s) in the subject and / or the severity of onset is substantially less than what would occur in the absence of the present invention.
[00144] A "treatment effective” amount as used herein is an amount that is sufficient to provide some improvement or benefit to the subject. Altematively stated, a "treatment effective” amount is an amount that will provide some alleviation, mitigation, decrease or stabilization in at least one clinical symptom in the subject. Those skilled in the art will appreciate that the therapeutic effects need not be complete or curative, as long as some benefit is provided to the subject.
[00145] A "prevention effective" amount as used herein is an amount that is sufficient to prevent and / or delay the onset of a disease, disorder and / or clinical symptoms in a subject and / or to reduce and / or delay the severity of the onset of a disease, disorder and / or clinical symptoms in a subject relative to what would occur in the absence of the methods of the invention. Those skilled in the art will appreciate that the level of prevention need not be complete, as long as some preventative benefit is provided to the subject.
[00146] The phrase a “therapeutically effective amount” and like phrases mean a dose or plasma concentration in a subject that provides the desired specific pharmacological effect, e.g. to express a therapeutic gene in the liver, and secretion into the plasma. It is emphasized that a therapeutically effective amount may not always be effective in treating the conditions described herein, even though such dosage is deemed to be a therapeutically effective amount by those of skill in the art. The therapeutically effective amount may vary based on the route of administration and dosage form, the age and weight of the subject, and / or the disease or condition being treated.
[00147] The terms “individual,” “subject,” and “patient” are used interchangeably, and refer to any individual subject with a disease or condition in need of treatment. For the purposes of the present disclosure, the subject may be a primate, preferably a human, or another mammal, such as a dog, cat, horse, pig, goat, or bovine, and the like.
[00148] Additional patents incorporated for reference herein that are related to, disclose or describe an AAV or an aspect of an AAV, including the DNA vector that includes the gene of interest to be expressed are: U.S. Patent Nos. 6,491,907; 7,229,823; 7,790,154; 7,201898; 7,071,172; 7,892,809; 7,867,484; 8.889.641; 9,169,494: 9,169,492; 9,441,206; 9,409,953; and, 9,447,433; 9,592,247; and, 9,737,618. II. rAAV genome elements
[00149] As disclosed herein, one aspect of the technology relates to a rAAV vector comprising a capsid, and within its capsid, a nucleotide sequence referred to as the “rAAV vector genome”. The rAAYV vector genome (also referred to as “rAAV genome) includes multiple elements, including, but not limited to two inverted terminal repeats (ITRs, e.g., the 5°-ITR and the 3°-ITR), and located between the ITRs are additional elements, including a promoter, a heterologous gene and a poly-A tail
[00150] In some embodiments, the rAAV genome disclosed herein comprises a 5° ITR and 3° ITR sequence, and located between the 5°ITR and the 3° ITR, a promoter, .g., a liver specific promoter sequence as disclosed herein, which operatively linked to a heterologous nucleic acid encoding a nucleic acid encoding an alpha-glucosidase (GAA) polypeptide, where the heterologous nucleic acid sequence can optionally further comprise one or more of the following elements: an intron sequence, a nucleic acid encoding a secretory signal peptide, a nucleic acid encoding an IGF2 targeting peptide, and a poly A sequence.
[00151] In some embodiments, the rAAV genome disclosed herein comprises a 5° ITR and 3° ITR sequence, and located between the 5°ITR and the 3° ITR, a promoter operatively linked to a heterologous nucleic acid encoding a secretory peptide and nucleic acid encoding an alpha- glucosidase (GAA) polypeptide (i.e., the heterologous nucleic acid encodes a GAA fusion polypeptide comprising a signal peptide-GAA polypeptide), where the rAAV genome optionally further comprises one or more of: an intron sequence, a collagen stability (CS) sequence, a polyA tail and a nucleic acid encoding a spacer of at least 1 amino acid. In some embodiments, the rAAV genome disclosed herein comprises a 5° ITR and 3° ITR sequence, and located between the 5°ITR and the 3° ITR, a liver specific promoter as disclosed herein operatively linked to a heterologous nucleic acid encoding a secretory peptide (e.g., FN1, AAT or GAA signal peptides) and nucleic acid encoding an alpha- glucosidase (GAA) polypeptide, where the rAAV genome optionally further comprises one or more of: an intron sequence (e.g., MVM or HBB2 intron sequence), a collagen stability (CS) sequence, a polyA tail and a nucleic acid encoding a spacer of at least 1 amino acid.
[00152] In some embodiments, the rAAV genome disclosed herein comprises a 5° ITR and 3° ITR sequence, and located between the 5°ITR and the 3° ITR, a promoter operatively linked to a heterologous nucleic acid encoding a secretory peptide, a targeting peptide and a GAA polypeptide (i.e., the heterologous nucleic acid encodes a GAA fusion polypeptide comprising a signal peptide- targeting sequence-GAA polypeptide), where targeting peptide is a IGF2 targeting peptide as described herein, and where the rAAV genome can optionally further comprise one or more of: an intron sequence, a collagen stability (CS) sequence, a polyA tail and a nucleic acid encoding a spacer of at least 1 amino acid.
[00153] Each of the elements in the rAAV genome are discussed herein. A. Alpha-glucosidase (GAA) polypeptide
[00154] Alpha-glucosidase (GAA) polypeptide is a member of family 31 of glycoside hydrolyases. Human GAA is synthesized as a 110 kDal precursor (Wisselaar et al. (1993) J. Biol. Chem. 268(3): 2223-31). The mature form of the enzyme is a mixture of monomers of 70 and 76 kDal (Wissclaar et al. (1993) J. Biol. Chem. 268(3): 2223-31). The precursor enzyme has seven potential glycosylation sites and four of these are retained in the mature enzyme (Wisselaar et al. (1993) J. Biol. Chem. 268(3): 2223-31). The proteolytic cleavage events which produce the mature enzyme occur in late endosomes or in the lysosome (Wisselaar et al. (1993) J. Biol. Chem. 268(3): 2223-31).
[00155] The rAAV vector genome can encode a GAA polypeptide can include, for example, amino acid residues 40-952 or 70-952 of human GAA, or a smaller portion, such as amino acid residues 40- 790 or 70-790.
[00156] In one embodiment, the GAA polypeptide can be fused to an IGF2 targeting sequence. In some embodiments, a IGF2 targeting sequence is fused to amino acid 40, or amino acid 70, or to an amino acid within one or two positions of amino acid 40 or 70 of human GAA polypeptide. In some embodiments, the IGF2 targeting peptide as disclosed herein is a ligand for an extracellular receptor, for example, the IGF2 targeting peptide binds to human cation-independent mannose-6-phosphate receptor (CI-MPR) or the IGF2 receptor.
[00157] The C-terminal 160 amino acids are absent from the mature 70 and 76 kDal GAA polypeptide species. However, certain Pompe alleles resulting in the complete loss of GAA activity map to this region, for example Val949Asp (Becker et al. (1998) J. Hum. Genet. 62:991). The phenotype of this mutant indicates that the C-terminal portion of the protein, although not part of the 70 or 76 kDal species, plays an important role in the function of the protein. It has also been reported that the C- terminal portion of the protein, although cleaved from the rest of the protein during processing, remains associated with the major species (Moreland et al. (Nov. 1, 2004) J. Biol. Chem., Manuscript 404008200). Accordingly, the C-terminal residues could play a direct role in the catalytic activity of the protein, and / or may be involved in promoting proper folding of the N-terminal portions of the protein.
[00158] The native GAA gene encodes a precursor polypeptide which possesses a signal sequence and an adjacent putative trans-membrane domain, a trefoil domain (PFAM PF00088) which is a cysteine- rich domain of about 45 amino acids containing 3 disulfide linkages (Thim (1989) FEBS Lett. 250:85), the domain defined by the mature 70 / 76 kDal polypeptide, and the C-terminal domain. It has been reported that both the trefoil domain and the C-terminal domain are required for the production of functional GAA, and that it is possible that the C-terminal domain interacts with the trefoil domain during protein folding perhaps facilitating appropriate disulfide bond formation in the trefoil domain.
[00159] The GAA polypeptide is described in US patents 5,962,313 and 6,537,785, which are incorporated herein in their entireties by reference. One of ordinary skill in the art can appreciate particular positions of GAA to which a secretory signal peptide (SS) or alternatively, the targeting peptide (e.g., IGF2 targeting peptide) can be fused. Accordingly, in one aspect the invention relates to a GAA fusion protein, where the SP or IGF2 targeting peptide is fused to amino acid 40, 68, 69, 70, 71, 72,779, 787, 789, 790, 791, 792, 793, or 796 of human GAA of SEQ ID NO: 10, or a modified GAA protein of SEQ ID NO: 170-174, or a portion thereof.
[00160] In some embodiments of the methods and compositions as disclosed herein, the human GAA protein expressed by the AAV comprises amino acids of SEQ ID NO: 10, or fragments or variants thereof, for example a human GAA protein beginning at residue 40, 68, 69, 70, 71, 72, 779, 787, 789, 790, 791, 792, 793, or 796 of SEQ ID NO: 10. In some embodiments of the methods and compositions as disclosed herein, the human GAA protein expressed by the AAV comprises amino acids of SEQ ID NO: 10, or a protein at least 60%, or 70%, or 80%, 85% or 90% or 95%, or 98%, or 99% identical to SEQ ID NO: 10. In some embodiments of the methods and compositions as disclosed herein, the human GAA protein expressed by the AAV comprises amino acids is a human GAA protein beginning at residue 40, 68, 69, 70, 71, 72, 779, 787, 789, 790, 791, 792, 793, or 796 of SEQ ID NO: 10, or a protein at least 60%, or 70%, or 80%, 85% or 90% or 95%, or 98%, or 99% identical thereto. In some embodiments, the human GAA protein expressed by the AAV comprises amino acids of beginning at residue 40, 68, 69, 70, 71, 72, 779, 787, 789, 790, 791, 792, 793, or 796 of any of SEQ ID NO: 170 (modGAA; H199R, R223H) or SEQ ID NO: 171 (modGAA; HI99R, R223H, H201L) or a protein at least 60%, or 70%, or 80%, 85% or 90% or 95%, or 98%, or 99% identical thereto.
[00161] In some embodiments, one of ordinary skill in the art can appreciate particular positions of GAA to which a secretory signal peptide (SS) or altematively, the targeting peptide (e.g., IGF2 targeting peptide) can be fused. For example, International Patent application W02018046774A1, which is incorporated herein in its entirety, discloses truncated GAA polypeptides to which the secretory signal peptide (SS) or alternatively, the targeting peptide (e.g., IGF2 targeting peptide) can be attached. The signal peptide or IGF2 targeting peptide can be attached to any truncated GAA polypeptide or truncated modified GAA polypeptide, starting amino acids of GAA truncated proteins as disclosed in U.S. Provisional Application 62,937,556, filed on November 19, 2019 and International Application WO 2020 / 102667, filed Nov 15, 2019. which is incorporated herein in its entirety by reference.
[00162] In some embodiments, the GAA-fusion polypeptides encoded by the rAAV genome as described herein can include, for example, amino acid residues 40-952 or residues 70-952 of human GAA, or a smaller portion, such as amino acid residues 40-790 or 70-790. In one embodiment, a secretory signal peptide (SS) or targeting peptide, e.g., IGF2 targeting peptide is fused to amino acid 40, or to amino acid 70, or to an amino acid within one or two positions of amino acid 40 or 70.
[00163] In some embodiments, the fusion protein comprising the secretory signal peptide (SS) and GAA polypeptide and optionally an IGF2 targeting peptide (i.., a SS-GAA fusion polypeptide, or a SS-IGF2-GAA fusion protein) comprises amino acid residues 40-952 or residues 70-952 of human acid alpha-glucosidase (GAA) (SEQ ID NO: 10). In some embodiments, the N-terminal of the GAA polypeptide is attached to the C-terminus of the SS and in some embodiments, the N-terminal of the GAA polypeptide is attached to the C-terminus of the IGF2 targeting peptide, and the N-terminus of the IGF2 targeting peptide is attached to the C-terminus of the secretory signal peptide. 0 Modified GAA (modGAA)
[00164] In some embodiments, the GAA protein comprises a H201L variant, as disclosed in US2014 / 0186326, and Moreland et al., Gene, 2012; 491 (25-30), which are both incorporated herein in their entirety by reference. In particular, the histidine (His) at amino acid position 201 is changed to a leucine (L) residue to enables rapid processing of the 76kD GAA pre-protein into the mature 70kD GAA protein.
[00165] In particular, in some embodiments, a fusion protein as disclosed herein comprises a GAA polypeptide of SEQ ID NO: 10, with a modification of amino acids that results in increased hydrophobicity at or near the N-terminal 70-kDa processing site. In some instances, the GAA peptide is modified at one or more amino acids corresponding to positions 190-209 of SEQ ID NO: 10. In further embodiments, the polypeptide is modified at one or more amino acids corresponding to positions 195-209 of SEQ ID NO: 10. In further embodiments, the modification is at one or more amino acids corresponding to amino acid positions 200-204 of SEQ ID NO: 10. In certain embodiments, the modification is at the amino acid corresponding to position 201 of SEQ ID NO: 10. In further embodiments, the modification is substitution of one or more amino acids with a more hydrophobic amino acid. In other embodiments, the modification is insertion of one or more hydrophobic amino acids. In even further embodiments, the hydrophobic amino acid is chosen from leucine and tyrosine, or a conservative amino acid of leucine or tyrosine.
[00166] In certain embodiments, GAA is modified to increase its hydrophobicity at or near the N- terminal 70-kDa processing site by substituting at least one amino acid with a more hydrophobic amino acid. In some embodiments, the substitution may be made within 5 amino acids upstream or downstream of the N-terminal 70-kDa processing site. In certain examples, the amino acid substitution may be made at an amino acid corresponding to position 195 to 209 of SEQ ID NO: 10. In other instances, the amino acid substitution may be made at an amino acid corresponding to position 200 to 204 of SEQ ID NO: 10. In further embodiments, the modified human GAA contains a hydrophobic amino acid at the position corresponding to amino acid position 201 of SEQ ID NO: 10. In some embodiments, GAA is modified by inserting one or more hydrophobic amino acids at or near the N-terminal 70-kDa processing site. Additional modifications include deletion of one or more amino acids at or near the N-terminal 70-kDa processing site.
[00167] In certain embodiments, a modified human GAA is provided containing a hydrophobic amino acid (natural or synthetic) at more than one position at the N-terminal 70-kDa processing site, or within 5 amino acids of the N-terminal 70-kDa processing site. In one embodiment, one of the modified amino acids is at the position corresponding to amino acid 201 of SEQ ID NO: 10.
[00168] In various embodiments the hydrophobic amino acid is chosen from valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, tyrosine, cysteine or alanine. In further embodiments, the hydrophobic amino acid is leucine or tyrosine. In some embodiments, the modified human GAA contains a synthetic or non-natural amino acid that exhibits hydrophobic properties. Generally, the substituted amino acid is more hydrophobic than the wild-type amino acid, and thus increases the hydrophobicity at or near the N-terminal 70kDa processing site.
[00169] In one exemplary embodiment, the modified GAA has a leucine at the position corresponding to amino acid 201 of SEQ ID NO: 10. In another embodiment, the modified GAA has a tyrosine at the position corresponding to amino acid 201 of SEQ ID NO: 10.
[00170] In some embodiments, the modified human GAA protein comprises a polypeptide with a His (H) to Arginine (R) (H199R) modification at amino acid position 199 of SEQ ID NO: 10 (GAA(HI199R), or a modification of an arginine (R) to a histidine (H) (R223H) at amino acid position 223 of SEQ ID NO: 10 (GAA(R223H). In some embodiments, the modified human GAA protein comprises a polypeptide with a His (H) to Arginine (R) (H199R) modification at amino acid position 199 of SEQ ID NO: 10, and a modification of an arginine (R) to a histidine (H) (R223H) at amino acid position 223 of SEQ ID NO: 10 (GAA(H199R-R223H). In some embodiments, the modified human GAA protein comprises SEQ ID NO: 170 or a variant of at least 80%, 90%, 95%, or 99% homology to at least 500, 550, 600, 650, 700, 750, 800, 850, or 900 amino acids of SEQ ID NO: 170, having at least one modification of HI99R or R223H, or both. In some embodiments, the cognate leader sequence of GAA (i.e., SEQ ID NO: 175 or amino acids 1-27 of SEQ ID NO: 170) is replaced with an IGF2 targeting peptide as disclosed herein, or a leader sequence of SEQ ID NO: 176, or an IL2 wild type leader peptide (SEQ ID NO: 178), modified IL2 leader peptide (SEQ ID NO: 180) or leader peptides at least 90% sequence identity to SEQ ID Nos 176, 178 or 180.
[00171] In some embodiments, the modified human GAA protein comprises SEQ ID NO: 171 ora variant of at least 80%, 90%, 95%, or 99% homology to at least 500, 550, 600, 650, 700, 750, 800, 850, or 900 amino acids of SEQ ID NO: 171, comprising at least the modification of H210L. In some embodiments, the cognate leader sequence of GAA (i.e., SEQ ID NO: 175 or amino acids 1-27 of SEQ ID NO: 171) is replaced with an IGF2 targeting peptide as disclosed herein, or a leader sequence of SEQ ID NO: 176, or an IL2 wild type leader peptide (SEQ ID NO: 178), modified IL2 leader peptide (SEQ ID NO: 180) or leader peptides at least 90% sequence identity to SEQ ID Nos 176, 178 or 180
[00172] In some embodiments, the modified human GAA protein comprises a polypeptide with at least one modification selected from: HI99R, R223H, or H201L of SEQ ID NO: 10, or a variant of at least 80%, 90%, 95%, or 99% homology to at least 500, 550, 600, 650, 700, 750, 800, 850, or 900 amino acids of SEQ ID NO: 10 having at least one of these modification. In some embodiments, the modified human GAA protein comprises a polypeptide comprises at least two modifications selected from: HI99R, R223H, or H201L of SEQ ID NO: 10, or a variant of at least 80%, 90%, 95%, or 99% homology to at least 500, 550, 600, 650, 700, 750, 800, 850, or 900 amino acids of SEQ ID NO: 10 having at least two of these modification. In some embodiments, the modified human GAA protein comprises a polypeptide with three modifications H199R, R223H, and H201L of SEQ ID NO: 10 (GAA- HI99R-H201L- R223H), or a variant of at least 80%, 90%, 95%, or 99% homology to at least 500, 550, 600, 650, 700, 750, 800, 850, or 900 amino acids of SEQ ID NO: 10 having these three modifications.
[00173] In certain embodiments, modified human GAAs are provided having at least 80%, 90%, 95%, or 99% homology to at least 500, 550, 600, 650, 700, 750, 800, 850, or 900 amino acids of SEQ ID NO: 10, and wherein the modified human GAA has at least one amino acid in the N-terminal 70-kDa processing site substituted with a more hydrophobic amino acid.
[00174] In some embodiments, at least 50% of the modified human GAA is processed to a 70-kDa form in the lysosome within 20, 30, or 40 hours. In still further embodiments, substantially all of the modified human GAA is processed to a 70-kDa form in the lysosome within 55, 65, or 75 hours.
[00175] In certain embodiments, a modified human GAA of the invention can be identified by its more rapid proteolytic processing to a mature 70-kDa form, or a corresponding variant thereof. In other embodiments, a modified human GAA as described herein can be identified by the production of an 82-kDa intermediate polypeptide that is not produced during proteolytic processing of native human GAA. In further embodiments, a modified human GAA can be identified by the absence of a 76-kDa intermediate polypeptide that is produced during proteolytic processing of unmodified human GAA.
[00176] In certain embodiments, the polypeptide has at least 80% identity to at least 500 amino acids of SEQ ID NO: 10 or SEQ ID NO: 170-171. In some instances, the polypeptide has at least 90% identity to at least 500 amino acids of SEQ ID NO: 10 or SEQ ID NO: 170-171. In other instances, the polypeptide has at least 95% identity to at least 500 amino acids of SEQ ID NO: 10 or SEQ ID NO: 170-171.
[00177] In certain embodiments, the GAA polypeptide with a modification at amino acid 201 to a hydrophobic residue, e.g., for example a H201L modification, exhibits more rapid lysosomal protease processing when compared to an unmodified human acid alpha-glucosidase protein. In some embodiments, at least 50% of the GAA pre-polypeptide is proteolytically processed to a 70-kDa mature GAA form within 20 hours of expression. In other embodiments, substantially all the GAA pre-polypeptide is proteolytically processed to a 70-kDa mature GAA form within 55 hours of expression.
[00178] In some embodiments, the cognate GAA leader peptide of amino acids 1-27 of SEQ ID NO: 10 (i.e., MGVRHPPCSHRLLAVCALVSLATAALL, SEQ ID NO: 175) is replaced with a different signal peptide (leader peptide). For example, the cognate leader peptide of GAA (SEQ ID NO: 175) can be replaced with any of: (i) an IgG1 leader peptide (referred to herein as a “201 leader peptide” or “2011p” having an amino acid sequence of: MEFGLSWVFLVALLKGVQCE (SEQ ID NO: 176) encoded by nucleic acid sequence SEQ ID NO: 177, (ii) wtIL2 Ip: MYRMQLLSCIALSLALVTNS (SEQ ID NO: 178) encoded by nucleic acid sequence SEQ ID NO: 179, or (iii) mutIL2 Ip: MYRMQLLLLIALSLALVTNS (SEQ ID NO: 180) encoded by nucleic acid sequence SEQ ID NO: 181. In some embodiments, the cognate GAA leader peptide (SEQ ID NO: 175) remains present, and an additional signal peptide is added, e.g., any one or more of signal peptides AAT, FN1, an IgG1 leader peptide (referred to herein as a “201 leader peptide” or “201lp” having an amino acid sequence of: MEFGLSWVFLVALLKGVQCE (SEQ ID NO: 176) encoded by nucleic acid sequence SEQ ID NO: 177, (ii) wtIL2 Ip: MYRMQLLSCIALSLALVTNS (SEQ ID NO: 178) encoded by nucleic acid sequence SEQ ID NO: 179, or (iii) mutIL2 Ip: MYRMQLLLLJALSLALVTNS (SEQ ID NO: 180) encoded by nucleic acid sequence SEQ ID NO: 181
[00179] In some embodiments, GAA is modified to add or remove glycosylation sites such as N- linked glycosylation sites, O-linked glycosylation sites or both. In certain embodiments, the addition or removal of glycosylation sites are achieved by N-terminal deletions, C-terminal deletions, internal deletions, random point mutagenesis, or, site directed mutagenesis. In some embodiments, the exemplary GAA modification involve addition of one or more Asparagine (Asn) residue / s or, one or more mutation to yield Asparagine (Asn) residue / s or, deletion of one or more Asparagine (Asn) residue / s. In certain embodiments, all or some of the N-linked, and / or, O-linked glycosylation sites present in GAA are mutated. In some embodiments, GAA modifications will yield information pertaining to the biological activity, physical structure and / or substrate binding potential of GAA. (ii) Nucleic acid encoding GAA
[00180] In some embodiments, the rAAV genome comprises a heterologous nucleic acid sequence encoding the entire GAA polypeptide (e.g., the N-terminal / catalytic and the C-terminal domain), that is not fused to a heterologous signal sequence or a targeting peptide.
[00181] In some embodiments, the rAAV genome comprises a heterologous nucleic acid sequence encoding a secretory signal peptide or IGF2 targeting peptide fused in frame to the 3' terminus of a GAA nucleic acid sequence that encodes the entire GAA polypeptide (e.g., the N-terminal / catalytic and the C-terminal domain). For example, heterologous nucleic acid sequence encoding a secretory signal peptide, or IGF2 targeting peptide is fused in frame to the 3' terminus of a GAA nucleic acid sequence that encodes the 70kDa and 76 kDa GAA polypeptides, such both polypeptides are expressed from the rAAV genome when the rAAV vector transduces a mammalian cell. In some embodiments, expression of the GAA nucleic acid can be driven by two promoters in the rAAV genome or by one promoter driving expression of a bicistronic construct.
[00182] In some embodiments of the methods and compositions as disclosed herein, the rAAV vector comprises a nucleic acid sequence encoding a GAA protein is a wild type GAA nucleic acid sequence, ¢.g., SEQ ID NO: 11 or SEQ ID NO: 72 or SEQ ID NO: 182. In some embodiments of the methods and compositions as disclosed herein, the rAAV vector comprises a nucleic acid sequence encoding a GAA protein which is a codon optimized GAA nucleic acid sequence, for any one or more of (i) enhanced expression in vivo, (ii) to reduce CpG islands, (iii).to reduce the innate immune response. Exemplary codon optimized GAA nucleic sequences encompassed for use in the methods and rAAV compositions as disclosed herein can be selected from any of: SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76 or SEQ ID NO: 182, or a nucleic acid sequence having at least 60%, or 70%, or 80%, 85% or 90% or 95%, or 98%, or 99% sequence identity to SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76 or SEQ ID NO: 182.
[00183] In addition, in some embodiments, the GAA nucleic acid sequences encompassed for use in the methods and rAAV compositions as disclosed herein are further modified with at least one or more of the following modifications: (i) removal of at least one, or two or in some embodiments, all alternative reading frames, (ii) removal of one or more CpGs islands, (iii) modification of the Kozak sequence, (iv) modification of a translational terminator sequence, and (v) removal of a spacer between promoter and Kozak sequence.
[00184] For example, in some embodiments, the rAAV composition comprises a hGAA nucleotide sequence of SEQ ID NO: 182, or a nucleic acid sequence having at least 60%, or 70%, or 80%, 85% or 90% or 95%, or 98%, or 99% sequence identity to SEQ ID NO: 182, where SEQ ID NO: 182 comprises the following elements shown in Table 1A, as compared to the wild type nucleic acid sequence for GAA;
[00185] Table 1A: elements of modGAA nucleic acid sequence. modification Base pair (nucleotide numbers based on SEQ ID NO: 440 cognate leader peptide 954-1034 (81bp 1152-1154 1899-1901 1909-1910 1965-1967 GTT to GTG to remove CpG 2022-2024 T2024G, 2154-2156 CAC to GAT to remove CpG 2163-2165 C2165T 2391-2393 2259-2561 GTC to GTG to remove CpG 3399-3401 C3401G 3534-3536 Replacement of non optimal 3810-3821 (12bp) stop
[00186] In some in some embodiments, the rAAV composition comprises a hGAA nucleotide sequence of SEQ ID NO: 182, or a nucleic acid sequence having at least 60%, or 70%, or 80%, 85% or 90% or 95%, or 98%, or 99% sequence identity to SEQ ID NO: 182, where the hGAA nucleotide sequence has been modified to with a series of point mutations that eliminate 3 potentially pro- inflammatory CpG motifs and a number of alternative reading frames (ARFs), where SEQ ID NO: 182 comprises the following point mutations as shown in Table 1B, as compared to the wildtype nucleic acid sequence for GAA, where numbering in table 1B assumes “A” in the GAA start codon ATG is the first nucleotide.
[00187] Table 1B: NT# | Purpose Original NT | New NT | NT# | Purpose Original New NT NT IE C | 1212 | Destrov CpG C T T C 1440 | Remove ORF A G T C 1608 | Remove ORF T C T C 2448 | Destrov CoG C G T G 2583 | Remove ORF T C C Destroy CpG Cc Remove ORF Cc Remove ORF C Destroy CpG G 2583 | Remove ORF AO [mT 50 [Rem ORr | _ T T AE A 1203 | Remove ORF “A
[00188] In some embodiments, the nucleic acid sequence encoding the cognate leader peptide in SEQ ID NO: 182 (e.g., nucleotides 1-81 of SEQ ID NO: 182) can be replaced by nucleic acid sequences encoding any of 2011p, wtIL2 Ip or mutIL2 Ip. Accordingly, in some embodiments, the nucleic residues 1-81 of SEQ ID NO: 182 (encoding the cognate leader peptide of GAA) can be replaced by nucleic acid sequences of SEQ ID NO: 177 (2011p), SEQ ID NO: 179 (wtIL2 Ip) or SEQ ID NO: 181 (mutIL2 Ip), or a nucleic acid sequence having at least 60%, or 70%, or 80%, 85% or 90% or 95%, or 98%, or 99% sequence identity to SEQ ID NOS: 177, 179 or 181.
[00189] In some embodiments, the rAAV vector or rAAV genome comprises a heterologous nucleic acid sequence encoding a GAA polypeptide comprising SEQ ID NO: 170 (GAA polypeptide with a cognate GAA signal sequence and H199R, R223H modifications), or SEQ ID NO: 171 (GAA polypeptide with a cognate GAA signal sequence and H199R, H201L, R223H modifications). The GAA polypeptide of SEQ ID NO: 170 is encoded by the nucleic acid sequence of SEQ ID NO: 182. Accordingly, in some embodiments, the rAAV vector comprises a nucleic acid of SEQ ID NO: 182 encoding a modified GAA polypeptide comprising HI99R, R223H modifications. The GAA polypeptide of SEQ ID NO: 171 is encoded by the nucleic acid sequence of SEQ ID NO: 182 where basepairs (bp) 667-669 of SEQ ID NO: 182 are changed from CAC to any of: UUA, UUG, CUU, CUC CUA, CUG (resulting in a Histadine (H) to Leucine (L) amino acid change); or where bp 668 of SEQ ID NO: 182 is changed from A to U. Accordingly, in some embodiments, the rAAV vector comprises a nucleic acid of SEQ ID NO: 182, where bp 667-669 of SEQ ID NO: 182 are changed from CAC to any of: UUA, UUG, CUU, CUC CUA, CUG (which changes the amino acid from Histidine (H) to leucine (L)); or where bp 668 of SEQ ID NO: 182 is changed from A to U, which encodes a modified GAA polypeptide comprising HI99R, H201L and R223H modifications.
[00190] In some embodiments, the rAAV vector or rAAV genome comprises a heterologous nucleic acid sequence encoding a GAA polypeptide selected from any of: SEQ ID NO: 172 (GAA polypeptide where cognate signal peptide is replaced with a IgG signal sequence and HI99R, R223H modifications), or SEQ ID NO: 173 (GAA polypeptide where cognate signal peptide is replaced with a wtIL2 signal sequence and HI99R, R223H modifications), SEQ ID NO: 174 (GAA polypeptide where cognate signal peptide is replaced with a mutIL3 signal sequence and HI99R, R223H modifications).
[00191] In some embodiments, the rAAV vector or AAV genome comprises a heterologous nucleic acid sequence comprising SEQ ID NO: 182 where bp 1-81 of SEQ ID NO: 182 is replaced with the nucleic acid of SEQ ID NO: 177 (IgG signal sequence), which encodes a GAA polypeptide of SEQ ID NO: 172 (IgG leader-GAA with H199R, R223H modifications). In some embodiments, the rAAV vector comprises a heterologous nucleic acid sequence comprising SEQ ID NO: 182, where bp 668 of SEQ ID NO: 182 is changed from A to U and where bp 1-81 of SEQ ID NO: 182 is replaced with the nucleic acid of SEQ ID NO: 177 (IgG signal peptide), which encodes a GAA polypeptide of SEQ ID NO: 172 (IgG leader-GAA with H199R, H201L and R223H modifications).
[00192] In some embodiments, the rAAV vector or rAAV genome comprises a heterologous nucleic acid sequence comprising SEQ ID NO: 182 where bp 1-81 of SEQ ID NO: 182 is replaced with the nucleic acid of SEQ ID NO: 179 (wt IL2 signal peptide), which encodes a GAA polypeptide of SEQ ID NO: 173 (wt IL2 signal peptide-GAA with HI99R, R223H modifications). In some embodiments, the rAAV vector comprises a heterologous nucleic acid sequence comprising SEQ ID NO: 182, where bp 668 of SEQ ID NO: 182 is changed from A to U and where bp 1-81 of SEQ ID NO: 182 is replaced with the nucleic acid of SEQ ID NO: 179 (wt IL2 signal peptide), which encodes a GAA polypeptide of SEQ ID NO: 173 (wt IL2 signal peptide-GAA with HI99R, H201L and R223H modifications).
[00193] In some embodiments, the rAAV vector comprises a heterologous nucleic acid sequence comprising SEQ ID NO: 182 where bp 1-81 of SEQ ID NO: 182 is replaced with the nucleic acid of SEQ ID NO: 181 (mutIL2 signal peptide), which encodes a GAA polypeptide of SEQ ID NO: 174 (mutIL2 signal peptide-GAA with H199R, R223H modifications). In some embodiments, the rAAV vector comprises a heterologous nucleic acid sequence comprising SEQ ID NO: 182, where bp 668 of SEQ ID NO: 182 is changed from A to U and where bp 1-81 of SEQ ID NO: 182 is replaced with the nucleic acid of SEQ ID NO: 181 (mut IL2 signal peptide), which encodes a GAA polypeptide of SEQ ID NO: 174 (mut IL2 signal peptide-GAA with HI99R, H201L, and R223H modifications).
[00194] The C-terminal domain of GAA functions in trans in conjunction with the 70 / 76 kDal species to generate active GAA. The boundary between the catalytic domain and the C-terminal domain appears to be at about amino acid residue 791, based on its presence in a short region of less than 18 amino acids that is absent from most members of the family 31 hydrolyases and which contains 4 consecutive proline residues in GAA. It has been reported that the C-terminal domain associated with the mature species begins at amino acid residue 792 (Moreland et al. (Nov. 1, 2004) J. Biol. Chem., Manuscript 404008200). Accordingly, in some embodiments, the GAA nucleic acid sequence that encodes the entire GAA polypeptide, with the exception of the C-terminal domain. Thus, in such an embodiment, the rAAV vector can be used to transduce a mammalian cell that expresses the C- terminal domain of GAA as a separate polypeptide. B. Secretory Signal peptide
[00195] Native GAA signal peptide is not cleaved in the ER thereby causing native GAA polypeptide to be membrane bound in the ER (Tsuji et al. (1987) Biochem. Int. 15(5):945-952). Disruption of the membrane association of GAA can be accomplished by replacing the endogenous GAA signal peptide (and optionally adjacent sequences) with an alternate signal peptide for GAA.
[00196] Accordingly, in representative embodiments, the rAAV vector and rAAV genome as disclosed herein further comprises a heterologous nucleic acid encoding a GAA polypeptide to be transferred to a target cell, attached to a heterologous nucleic acid sequence that encodes a secretory signal peptide in the place of the endogenous GAA signal peptide. The heterologous nucleic acid is operatively associated with the segment encoding the secretory signal peptide, such that upon transcription and translation a fusion polypeptide is produced containing the secretory signal sequence operably associated with (e.g., directing the secretion of) the GAA polypeptide.
[00197] In some embodiments, the AAV vector encodes a GAA polypeptide that comprises the endogenous GAA signal peptide (e.g., amino acids 1-27 of SEQ ID NO: 10 (also referred to as “innate GAA” or “cognate GAA” signal peptide). In some embodiments, the AAV vector encodes a GAA polypeptide that comprises the endogenous GAA signal peptide (e.g., amino acids 1-27 of SEQ ID NO: 10 (also referred to as “innate GAA” or “cognate GAA” signal peptide) and an additional heterologous (non native) signal sequence. In some embodiments, the GAA polypeptide that lacks the endogenous signal peptide of amino acids 1-27 of GAA is fused to a secretory signal. In some embodiments of the compositions and methods described herein, the secretory signal serves a general purpose of assisting the secretion of the GAA polypeptide, or a fusion polypeptide, e.g., the IGF2 targeting peptide-GAA fusion polypeptide from the liver cells into the blood, where it can travel and be targeted to the lysosomes of mammalian cells, for example, human cardiac and skeletal muscle cells, as described herein. In some embodiments, a heterologous secretory signal is selected from any of: a AAT signal peptide, a fibronectin signal peptide (FN1), a GAA signal peptide, or an active fragment of AAT, FN1 or GAA signal peptide having secretory signal activity.
[00198] In some embodiments, the secretory signal peptide is heterologous to (i.¢., foreign or exogenous to) the polypeptide of interest. For example, a heterologous secretory signal peptide is a fibronectin secretory signal peptide, the polypeptide of interest is not fibronectin. In some embodiments, the secretory signal peptide is selected from any of: AAT signal peptide, a fibronectin signal peptide (FN1), or an active fragment of AAT, FN1 or GAA signal peptide having secretory signal activity. In alternative embodiments, the secretory signal peptide is not heterologous to GAA, i.e., the signal peptide is the GAA signal peptide (i.c., residues 1-27 of the native GAA polypeptide).
[00199] In some embodiments, the cognate GAA signal peptide of amino acids 1-27 of SEQ ID NO: 10 (i.e, MGVRHPPCSHRLLAVCALVSLATAALL, SEQ ID NO: 175) is replaced with a different or heterologous leader peptide. For example, the cognate leader peptide of GAA (SEQ ID NO: 175) can be replaced with any of the heterologous signal peptides selected from: (i) an IgG1 leader peptide (referred to herein as a “201 leader peptide” or “2011p” having an amino acid sequence of: MEFGLSWVFLVALLKGVQCE (SEQ ID NO: 176) encoded by nucleic acid sequence SEQ ID NO: 177, (ii) wtIL2 Ip: MYRMQLLSCIALSLALVTNS (SEQ ID NO: 178) encoded by nucleic acid sequence SEQ ID NO: 179, or (iii) mutIL2 Ip: MYRMQLLLLIALSLALVTNS (SEQ ID NO: 180) encoded by nucleic acid sequence SEQ ID NO: 181, or a leader peptide having at least 90% sequence identity to any of SEQ ID NOs 176, 178 or 180.
[00200] In general, the secretory signal peptide will be at the amino-terminus (N-terminus) of the fusion polypeptide (i.e., the nucleic acid segment encoding the secretory signal peptide is 5’ to the heterologous nucleic acid encoding the GAA peptide or GAA-fusion peptide in the rAAV vector or rAAV genome as disclosed herein). Alternatively, the secretory signal may be at the carboxyl- terminus or embedded within the GAA polypeptide or GAA fusion polypeptide (¢.g., IGF2-GAA fusion polypeptide), as long as the secretory signal is operatively associated therewith and directs secretion of the GAA polypeptide or GAA fusion polypeptide of interest (either with or without cleavage of the signal peptide from the GAA polypeptide) from the cell.
[00201] The secretory signal is operatively associated with the GAA polypeptide or GAA fusion polypeptide is targeted to the secretory pathway. Alternatively stated, the secretory signal is operatively associated with the GAA polypeptide such that the GAA-polypeptide or GAA fusion polypeptide is secreted from the cell at a higher level (i.c., a greater quantity) than in the absence of the secretory signal peptide. In general, typically at least about 20%, 30%, 40%, 50%, 70%, 80%, 85%, 90%, 95% or more of the GAA-polypeptide or IGF2-GAA fusion polypeptide (alone and / or fused with the signal peptide) is secreted from the cell when a signal peptide is attached as compared to in the absence of the attachment of a secretory signal peptide. In other embodiments, essentially all of the detectable polypeptide (alone and / or in the form of the fusion polypeptide) is secreted from the cell.
[00202] By the phrase “secreted from the cell”, the polypeptide may be secreted into any compartment (e.g., fluid or space) outside of the cell including but not limited to: the interstitial space, blood, lymph, cerebrospinal fluid, kidney tubules, airway passages (¢.g., alveoli, bronchioles, bronchia, nasal passages, etc.), the gastrointestinal tract (e.g., esophagus, stomach, small intestine, colon, etc.), vitreous fluid in the eye, and the cochlear endolymph, and the like.
[00203] In one embodiment, the rAAV genome comprises a heterologous nucleic acid that encodes a secretory signal peptide (SP) fused to the GAA-fusion polypeptide, where the GAA-fusion polypeptide comprises a targeting peptide (e.g., IGF2 targeting peptide) fused to a GAA polypeptide. As used herein GAA also refers to the modified GAA described above. Accordingly, the signal peptide disclosed herein increases the efficacy of secretion of the GAA polypeptide or IGF2-GAA fusion polypeptide from the cell transduced with the rAAV vector or comprising the rAAV genome as described herein
[00204] Accordingly, in some embodiments, the rAAV genome disclosed herein comprises a 5° ITR and 3” ITR sequence, and located between the 5°ITR and the 3° ITR, a promoter operatively linked to a heterologous nucleic acid encoding a secretory peptide and nucleic acid encoding an alpha- glucosidase (GAA) polypeptide (i.e., the heterologous nucleic acid encodes a GAA fusion polypeptide comprising a signal peptide-GAA polypeptide).
[00205] In alternative embodiments, the rAAV genome disclosed herein comprises a 5° ITR and 3° ITR sequence, and located between the 5°ITR and the 3° ITR, a promoter operatively linked to a heterologous nucleic acid encoding a secretory peptide and nucleic acid encoding an alpha- glucosidase (GAA) fusion polypeptide, where the fusion protein comprises IGF2 targeting peptide and a GAA polypeptide (i.., the heterologous nucleic acid encodes a GAA fusion polypeptide comprising a signal peptide-IGF2-GAA polypeptide).
[00206] Generally, secretory signal peptides are cleaved within the endoplasmic reticulum and, in some embodiments, the secretory signal peptide is cleaved from the GAA polypeptide prior to secretion. It is not necessary, however, that the secretory signal peptide is cleaved as long as secretion of the GAA polypeptide or IGF2-GAA fusion polypeptide from the cell is enhanced and the GAA polypeptide is functional. Thus, in some embodiments, the secretory signal peptide is partially or entirely retained.
[00207] In some embodiments, the rAAV genome, or an isolated nucleic acid as disclosed herein comprises a nucleic acid encoding a chimeric polypeptide comprising a GAA polypeptide operably linked to a secretory signal peptide, and the chimeric polypeptide is expressed and produced from a cell transduced with the rAAV vector and the GAA polypeptide is secreted from the cell. The GAA polypeptide or GAA fusion polypeptide (e.g., IGF2-GAA fusion polypeptide) can be secreted after cleavage of all or part of the secretory signal peptide. Alternatively, the GAA polypeptide or GAA fusion polypeptide (e.g., IGF2-GAA fusion polypeptide) can retain the secretory signal peptide (i.e., the secretory signal is not cleaved). Thus, in this context, the “GAA polypeptide or GAA fusion polypeptide” can be a chimeric polypeptide comprising the secretory peptide.
[00208] The secretory signal sequences of the invention are not limited to any particular length as long as they direct the polypeptide of interest to the secretory pathway. In representative embodiments, the signal peptide is at least about 6, 8, 10 12, 15, 20, 25, 30 or 35 amino acids in length up to a length of about 40, 50, 60, 75, or 100 amino acids or longer.
[00209] Secretory signal peptide encoded by the rAAV genome and in the rAAV vector as disclosed herein can comprise, consist essentially of or consist of a naturally occurring secretory signal sequence or a modification thereof. Numerous secreted proteins and sequences that direct secretion from the cell are known in the art, are disclosed in US Patent 9,873,868, which is incorporated herein in its entirety by reference. Exemplary secreted proteins (and their secretory signals) include but are not limited to: erythropoietin, coagulation Factor IX, cystatin, lactotransferrin, plasma protease C1 inhibitor, apolipoproteins (e.g., APO A, C, E), MCP-1, o-2-HS-glycoprotein, o-1-microgolubilin, complement (e.g., C1Q, C3), vitronectin, lymphotoxin-a, azurocidin, VIP, metalloproteinase inhibitor 2, glypican-1, pancreatic hormone, clusterin, hepatocyte growth factor, insulin, a-1-antichymotrypsin, growth hormone, type IV collagenase, guanylin, properdin, proenkephalin A, inhibin (e.g., A chain), prealbumin, angiocenin, lutropin (e.g., B chain), insulin-like growth factor binding protein 1 and 2, proactivator polypeptide, fibrinogen (e.g., p chain), gastric triacylglycerol lipase, midkine, neutrophil defensins 1, 2, and 3, o-1-antitrypsin, matrix gla-protein, a-tryptase, bile-salt-activated lipase, chymotrypsinogen B, elastin, IG lambda chain V region, platelet factor 4 variant, chromogranin A, ‘WNT-1 proto-oncogene protein, oncostatin M, -neoendorphin-dynorphin, von Willebrand factor, plasma serine protease inhibitor, serum amyloid A protein, nidogen, fibronectin, rennin, osteonectin, histatin 3, phospholipase A2, cartilage matrix Protein, GM-CSF, matrilysin, neuroendocrine protein 7B2, placental protein 11, gelsolin, M-CSF, transcobalamin I, lactase-phlorizin hydrolase, elastase 2B, pepsinogen A, MIP 1-B, prolactin, trypsinogen II, gastrin-releasing peptide II, atrial natriuretic factor, secreted alkaline phosphatase, pancreatic a-amylase, secretogranin I, B-casein, serotransferrin, tissue factor pathway inhibitor, follitropin B-chain, coagulation factor XII, growth hormone-releasing factor, prostate seminal plasma protein, interleukins (e.g., 2, 3, 4, 5, 9, 11), inhibin (e.g., alpha chain), angiotensinogen, thyroglobulin, IG heavy or light chains, plasminogen activator inhibitor-1, lysozyme C, plasminogen activator, antileukoproteinase 1, statherin, fibulin-1, isoform B, uromodulin, thyroxine-binding globulin, axonin-1, endometrial a-2 globulin, interferon (e.g., alpha, beta, gamma), B-2-microglobulin, procholecystokinin, progastricsin, prostatic acid phosphatase, bone sialoprotein II, colipase, Alzheimer's amyloid A4 protein, PDGF (e.g., A or B chain), coagulation factor V, triacylglycerol lipase, haptoglobuin-2, corticosteroid-binding globulin, triacylglycerol lipase, prorelaxin H2, follistatin 1 and 2, platelet glycoprotein IX, GCSF, VEGF, heparin cofactor II, antithrombin-III, leukemia inhibitory factor, interstitial collagenase, pleiotrophin, small inducible cytokine Al, melanin-concentrating hormone, angiotensin-converting enzyme, pancreatic trypsin inhibitor, coagulation factor VIII, a-fetoprotein, a-lactalbumin, senogelin II, kappa casein, glucagon, thyrotropin beta chain, transcobalamin II, thrombospondin 1, parathyroid hormone, vasopressin copeptin, tissue factor, motilin, MPIF-1, kininogen, neuroendocrine convertase 2, stem cell factor procollagen al chain, plasma kallikrein keratinocyte growth factor, as well as any other secreted hormone, growth factor, cytokine, enzyme, coagulation factor, milk protein, immunoglobulin chain, and the like.
[00210] In some embodiments, other secretory signal peptides encoded by the rAAV genome and in the rAAV vector as disclosed herein can be selected from, but are not limited to, the secretory signal sequences from prepro-cathepsin L (e.g., GenBank Accession Nos. KHRTL, NP_037288; NP_034114, AAB81616, AAA39984, P07154, CAA68691; the disclosures of which are incorporated by reference in their entireties herein) and prepro-alpha 2 type collagen (e.g., GenBank Accession Nos. CAA98969, CAA26320, CGHU2S, NP_000080, BAA25383, P08123; the disclosures of which are incorporated by reference in their entireties herein) as well as allelic variations, modifications and functional fragments thereof (as discussed above with respect to the fibronectin secretory signal sequence). Exemplary secretory signal sequences include for preprocathepsin L (Rattus norvegicus, MTPLLLLAVLCLGTALA [SEQ ID NO: 27]; Accession No. CAA68691) and for prepro-alpha 2 type collagen (Homo sapiens, MLSFVDTRTLLLLAVTLCLATC [SEQ ID NO: 28]; Accession No. CAA98969). Also encompassed are longer amino acid sequences comprising the full-length secretory signal sequence from preprocathepsin L and prepro-alpha 2 type collagen or functional fragments thereof (as discussed above with respect to the fibronectin secretory signal sequence).
[00211] In some embodiments, the secretory signal peptide is derived in part or in whole from a secreted polypeptide that is produced by liver cells. In some embodiments, a secretory signal peptide can further be in whole or in part synthetic or artificial. Synthetic or artificial secretory signal peptides are known in the art, see e.g., Barash et al., “Human secretory signal peptide description by hidden Markov model and generation of a strong artificial signal peptide for secreted protein expression,” Biochem. Biophys. Res. Comm. 294:835-42 (2002); the disclosure of which is incorporated herein in its entirety. In particular embodiments, the secretory signal peptide comprises, consists essentially of, or consists of the artificial secretory signal: MWWRLWWLLLLLLLLWPMVWA (SEQ ID NO: 29) or variations thereof having 1, 2, 3, 4, or 5 amino acid substitutions (optionally, conservative amino acid substitutions, conservative amino acid substitutions are known in the art).
[00212] Exemplary signal peptides for use in the methods and compositions as disclosed herein can be selected from any signal peptide disclosed in Table 2, or functional variants thereof. Exemplary signal peptides are Fibronectin (FN1), or AAT. In some embodiments of the methods and compositions disclosed herein, the rAAV vector composition comprises the nucleic acid encoding a secretory signal peptide, e.g., encoding a secretory signal peptide selected from an AAT signal peptide (e.g., SEQ ID NO: 17), a fibronectin signal peptide (FN1) (e.g., SEQ ID NO: 18-21), a GAA signal peptide, an hIGF2 signal peptide (e.g., SEQ ID NO: 22) or an active fragment thereof having secretory signal activity, e.g., a nucleic acid encoding an amino acid sequence that has at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to SEQ ID NOs: 17-22.
[00213] In some embodiments of the methods and compositions as disclosed herein, the nucleic acid encoding the secretory signal is selected from any of SEQ ID NO: 17, 81-21, 22-26, or a nucleic acid sequence at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to any of SEQ ID NOs: 17 or 22-26.
[00214] In some embodiments, one can readily substitute a FN1 or AAT signal peptide with any signal peptide, including signal peptides for over liver expressed proteins, or signal peptides disclosed in 62,937,556, filed on November 19, 2019, or PCT / US19 / 61633 filed on November 15, 2019.
[00215] Fibronectin secretory signal peptide:
[00216] In some embodiments, the secretory signal peptide is a fibronectin secretory signal peptide, which term includes modifications of naturally occurring sequences (as described in more detail below).
[00217] In some embodiments, the secretory signal peptide is a fibronectin signal peptide, e.g., a signal sequence of human fibronectin or a signal sequence from rat fibronectin. Fibronectin (FN1) signal sequences and modified FNI signal peptides encompassed for use in the rAAV genome and rAAV vectors described herein are disclosed in US patent 7,071,172, which is incorporated herein in its entirety by reference, and in Table 3 of provisional application 62 / 937,556, filed on November 19, 2019. Examples of exemplary fibronectin secretory signal sequences include, but are not limited to those listed in Table 1 of US patent 7,071,172, which is incorporated herein in its entirety by reference.
[00218] Table 2: Exemplary Fibronectin (FN1) secretory signal peptides Species Secretory Signal sequence Nucleic acid sequence H. Sapiens MLRGPGPGLLLLAVQCLGTAV | ATG CTT AGG GGT CCG GGG CCC GGG CTG PSTGA (SEQ ID NO: 20) CTG CTG CTG GCC GTC CAG TGC CTG GGG ACA GCG GTG CCC TCC ACG GGA GCC (SEQID NO: 25) [rR MLRGPGPGRLLLLAVLCLGTSV | 5'- ATGCTCAGGGGTCCGGGACCCGGGCGGCT X. laevis Norvegicus | RCTETGKSKR (SEQ ID NO: 18) | GCTGCTGCTAGCAGTCCTGTGCCTGGGGAC ATCGGTGCGCTGCACCGAAACCGGGAAGA GCAAGAGG-3 (SEQ ID NO: 23) (nucleotides 208-303 R MLRGPGPGRLLLLAVLCLGTSV | 5'-ATG CTC AGG GGT CCG GGA CCC GGG Norvegicus | RCTETGKSKR 1 LALQIV CGG CTG CTG CTG CTA GCA GTC CTG TGC | (SEQ ID NO: 19) CTG GGG ACA TCG GTG CGC TGC ACC GAA ACC GGG AAG AGC AAG AGG T CAG GCT CAG CAA ATC GTG-3'. (SEQID NO: 24) (1 denotes the cleavage site) X laevis ATG CGC CGG GGG GCC CTG ACC GGG CTG MRRGALTGLLLVLCLSVVLRA | CTC CTG GTC CTG TGC CTG AGT GTT GTG APSATSKKRR (SEQ ID NO: 21) CTA CGT GCA GCC CCC TCT GCA ACA AGC | AAG AAG CGC AGG (SEQ ID NO: 26)
[00219] An exemplary nucleotide sequence encoding the fibronectin secretory signal sequence of Rattus norvegicus is found at GenBank accession number X15906 (the disclosure of which is incorporated herein by reference). As yet another illustrative sequence, the nucleotide sequence encoding the secretory signal peptide of human fibronectin 1, transcript variant 1 (Accession No. NM_002026, nucleotides 268-345; the disclosure of Accession No. NM_002026 is incorporated herein by reference in its entirety). Another exemplary secretory signal sequence is encoded by the nucleotide sequence encoding the secretory signal peptide of the Xenopus laevis fibronectin protein (Accession No. M77820, nucleotides 98-190; the disclosure of Accession No. M77820 incorporated herein by reference in its entirety).
[00220] In another embodiment, the fibronectin signal sequence (FN1, nucleotides 208-303, 5-ATG CTC AGG GGT CCG GGA CCC GGG CGG CTG CTG CTG CTA GCA GTC CTG TGC CTG GGG ACA TCG GTG CGC TGC ACC GAA ACC GGG AAG AGC AAG AGG-3', SEQ ID NO: 23) was derived from the rat fibronectin mRNA sequence (Genbank accession #X15906) and codes for the following peptide signal sequence: Met Leu Arg Gly Pro Gly Pro Gly Arg Leu Leu Leu Leu Ala Val Leu Cys Leu Gly Thr Ser Val Arg Cys Thr Glu Thr Gly Lys Ser Lys Arg (SEQ ID NO: 18). In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises a heterologous nucleic acid sequence that encodes a secretory signal peptide which is a fibronectin signal peptide (FN1) or an active fragment thereof having secretory signal activity (e.g., a FNI1 signal peptide has the sequence of any of SEQ ID NO: 18-21, or an amino acid sequence at having at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to any of SEQ ID NOs: 18-21), and the heterologous nucleic acid sequence encodes a IGF2 targeting peptide selected from any of: SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8 or SEQ ID NO: 9, or a IGF2 peptide having at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to SEQ ID NOs: 5-9. In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises a heterologous nucleic acid sequence that encodes a secretory signal peptide is AAT signal peptide or an active fragment thereof having secretory signal activity, (e.g., a AAT signal peptide has the sequence of SEQ ID NO: 17, or an amino acid sequence at having at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to SEQ ID NO: 17), and the heterologous nucleic acid sequence encodes a IGF2 targeting peptide selected from any of: SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8 or SEQ ID NO: 9, or a IGF2 peptide having at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to SEQ ID NOs: 5-9.
[00221] Those skilled in the art will appreciate that the secretory signal sequence may encode one, two, three, four, five or all six or more of the amino acids at the C-terminal side of the peptidase cleavage site (identified by an T) (see e.g., SEQ ID NO: 19 and 24 in Table 2). Those skilled in the art will appreciate that additional amino acids (e.g., 1, 2, 3, 4, 5, 6 or more amino acids) on the carboxy-terminal side of the cleavage site may be included in the secretory signal sequence.
[00222] In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises, or consist of, located between the 5° ITR and the 3° ITR, a heterologous nucleic acid sequence that encodes a secretory signal peptide and nucleic acid encoding a hGAA polypeptide, where the nucleic acid sequence that encodes the signal sequence is selected from any of: an AAT signal peptide (e.g., SEQ ID NO: 17), a fibronectin signal peptide (FN1) (e.g., SEQ ID NO: 18-21), a cognate GAA signal peptide (SEQ ID NO: 175), an hIGF2 signal peptide (e.g., SEQ ID NO: 22), a IgGl leader peptide (SEQ ID NO: 177), wtIL2 leader peptide (SEQ ID NO: 179), mutant IL2 leader peptide (SEQ ID NO: 181) or an active fragment thereof having secretory signal activity, e.g., a nucleic acid encoding an amino acid sequence that has at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to SEQ ID NOs: 17-22, 175, 177, 179 or 181, and where the nucleic acid encoding the signal peptide is located 5” of a nucleic acid encoding a hGAA polypeptide as disclosed herein, and where the nucleic acid encoding the signal sequence and the hGAA polypeptide are operatively linked to any LSP disclosed herein in Table 4, or a functional variant thereof.
[00223] ). In embodiments of the invention, the functional fragment has at least about 50%, 70%, 80%, 90% or more secretory signal activity as compared with the sequences specifically disclosed herein or even has a greater level of secretory signal activity.
[00224] Peptidase cleavage sites
[00225] In some embodiments, one or more exogenous peptidase cleavage site may be inserted into the secretory signal peptide - GAA fusion polypeptide, e.g., between the secretory signal peptide and the GAA polypeptide. In particular embodiments, an autoprotease (¢.g., the foot and mouth disease virus 2A autoprotease) is inserted between the secretory signal peptide and the GAA polypeptide or IGF2-GAA fusion polypeptide. In other embodiments, a protease recognition site that can be controlled by addition of exogenous protease is employed (e.g., Lys—Arg recognition site for trypsin, the Lys—Arg recognition site of the Aspergillus KEX2-like protease, the recognition site for a metalloprotease, the recognition site for a serine protease, and the like). Modification of the GAA polypeptide to delete or inactivate native protease sites is encompassed herein and disclosed in U.S. Provisional Application 62,937,556, filed on November 19, 2019 and International Application PCT / US19 / 61653, filed Nov 15, 2019. C. IGF2 Targeting Peptide Sequence
[00226] In one embodiment, the rAAV genome comprises a heterologous nucleic acid that encodes a targeting peptide (TP) fused to the GAA polypeptide. In some embodiments, the targeting peptide is a ligand for an extracellular receptor, wherein the targeting peptide binds an extracellular domain of a receptor on the surface of a target cell and, upon intemalization of the receptor, permits localization of the polypeptide in a human lysosome. In one embodiment, the targeting peptide includes a urokinase- type plasminogen receptor moiety capable of binding the cation-independent mannose-6-phosphate receptor. In some embodiments, the targeting peptide incorporates one or more amino acid sequences of a IGF2 targeting peptide.
[00227] In some embodiments, the IGF2 targeting peptide as disclosed herein comprises at least part of a ligand for an extracellular receptor, for example, the IGF2 targeting peptide binds to human cation-independent mannose-6-phosphate receptor (CI-MPR) or the IGF2 receptor.
[00228] IGF2 is also known by alias; chromosome 11 open reading frame 43, insulin-like growth factor 2, IGF-II, FLJ44734; IGF2, somatomedin A and preptin. The mRNA of wild-type human IGF2 sequence is corresponds to: GCTTACCGCCCCAGTGAGACCCTGTGCGGCGGGGAGCTGGTGGACACCCTCCAGTTCGTC TGTGGGGACCGCGGCTTCTACTTCAGCAGGCCCGCAAGCCGTGTGAGCCGTCGCAGCCGT GGCATCGTTGAGGAGTGCTGTTTCCGCAGCTGTGACCTGGCCCTCCTGGAGACGTACTGT GCTACCCCCGCCAAGTCCGAG (SEQ ID NO: 1). The full length IGF2 protein (including the IGF?2 targeting sequence) is encoded by the nucleic acid sequence of NM_000612.6, and encodes the full length IGF2 protein NP_000603.1.
[00229] The mature human IGF2 targeting peptide is shown below: AYRPSETLCGGELVDTLQFVCGDRGFYFSRPASRVSRRSRGIVEEC CFRSCDLALLETYCATPAKSE(SEQIDNO: 5)
[00230] The coding sequence of human IGF2 is also disclosed in US patent 8,492,388 (see e.g., FIG. 2) which is incorporated herein in its entirety by reference. IGF2 protein is synthesized as a pre- pro-protein with a 24 amino acid signal peptide at the amino terminus and a 89 amino acid carboxy terminal region both of which are removed post-translationally, reviewed in O'Dell et al. (1998) Int. J. Biochem Cell Biol. 30(7): 767-71. The mature protein is 67 amino acids. A Leishmania codon optimized version of the mature IGF2 is disclosed in US patent 8,492,388 (see, e.g, FIG. 3 of 8,492,388) (Langford et al. (1992) Exp. Parasitol. 74(3):360-1). Additional cassettes containing a deletion of amino acids 1-7 or 2-7 of the mature polypeptide (A1-7), alteration of residue 27 from tyrosine to leucine (Y27L) or both mutations (A1-7,Y27L or A2-7,Y27L) were made to produce IGF- 2 cassettes with specificity for only the desired receptor as described below. Accordingly, in some embodiments, the IGF2 targeting sequence can be selected from any of: wildtype, Y27L, Al-7, A2-7 and Y27L-Al1-7, Y27L-A2-7, V43M, Y27L-V43M, Y27L-A1-7-V43M, Y27L-A2-7-V43M IGF2 variants are encompassed for use herein.
[00231] Exemplary IGF2 targeting peptide for use in the methods and compositions herein are disclosed in U.S. Provisional Application 62,937,556, filed on November 19, 2019 and International Application PCT / US19 / 61653, filed Nov 15, 2019, and International application PCT / US19 / 61701, filed November 15, 2019, each of which are incorporated herein in their entirety by reference.
[00232] In some embodiments, an IGF2 targeting peptide for use in the methods and compositions herein can have a modification of any one or more of: E6R, F268, Y27L, V43L, F48T, R495, S501, AS54R, L55R, K65R, as disclosed in US application 2019 / 0343968, which is incorporated herein in its entirety. In some embodiments, the IGF2 targeting peptide has a modification of V43M in addition to one or more modifications selected from: E6R, F26S, Y27L, V43L, F48T, R495, S501, A54R, L55R and K65R. In some embodiments, the IGF2 targeting peptide has a A1-7 or A2-7 modification in addition to one or more modifications selected from: E6R, F268, Y27L, V43L, F48T, R495, S501, A54R, L55R and K65R. In some embodiments, the IGF2 targeting peptide has a A1-7 or A2-7 modification, a V43M modification, and one or more modifications selected from: E6R, F26S, Y27L, V43L, F48T, R495, S501, A54R, L55R and K65R.
[00233] In particular embodiments, the IGF2 targeting peptide comprises a modification at valine 43, where valine is modified to a met (V43M), such that translation initiation starts at amino acid 43. A IGF?2 targeting peptide with a modification of V43M encompassed for use herein as a targeting peptide or IGF2 targeting peptide binds the cation-independent mannose-6-phosphate receptor. In alterative embodiments, the IGF2 targeting peptide is delta 1-42 of IGF2 with V43 changed to an Met (i.e., IGF2-A1-42 (SEQ ID NO: 8) or IGF2-V43M (SEQ ID NO:9).
[00234] In some embodiments, the rAAV genome comprises a nucleic acid encoding an IGF2-GAA fusion protein, where the nucleic acid encoding the mature IGF2 targeting peptide (SEQ ID NO: 5) or a IGF2 targeting peptide variant (e.g., SEQ ID NO: 6 (IGF2-A2-7); SEQ ID NO: 7 (IGF2-A1-7); SEQ ID NO: 8 (IGF2--A1-42), SEQ ID NO: 9 (IGF2-V43M)) or sequences having at least 85%, or 90% or 95% sequence identity to SEQ ID NO: 5-9, is fused to the 5’ end of nucleic acid encoding the GAA protein, fusion proteins (e.g., IGF2-GAA fusion polypeptides) are created that can be taken up by a variety of cell types and transported to the lysosome. Alternatively, a nucleic acid encoding a precursor IGF2 polypeptide can be fused to the 3’ end of a GAA gene; the precursor includes a carboxy-terminal portion that is cleaved in mammalian cells to yield the mature IGF2 polypeptide, but the IGF2 targeting peptide is preferably omitted (or moved to the 5' end of the GAA gene). This method has numerous advantages over methods involving glycosylation including simplicity and cost effectiveness, because once the protein is isolated, no further modifications need be made.
[00235] In some embodiments, the IGF2 targeting peptide encompassed for use herein is described US patents 7,785,856 and 9,873,868 which are each incorporated herein in their entirety by reference. (i) Deletion mutants of IGF2:
[00236] In some embodiments, the IGF2 targeting peptide is a modified or truncated IGF2 targeting peptide (also referred to as a deletion mutant of IGF2), as disclosed in International Application PCT / US19 / 61701, filed Nov 15, 2019, which is incorporated herein in its entirety by reference. For example, in some embodiments, the IGF2 targeting peptide comprises a V43M modification and also any deletion of one or more amino acids from amino acid 1-42. For example, in some embodiments of the methods and compositions as disclosed herein, the IGF2 targeting peptide comprises V43M and further comprises one or more deletions selected from any of: Al-3, Al-4, Al-5, Al-6, A1-8, A1-9, AL-10, AL-11, Al-12, AI-13, Al-14, AI-15, Al-16, AL-17, Al-18, AI-19, A1-20, Al-21, A1-22, A1-23, Al-24, A125, A1-26, A1-27, A1-28, A1-29, A1-30, Al-31, A1-32, A1-33, A1-34, A1-35, A1-36, AL-37, A1-38, A1-39, A1-40, Al-41 or Al-42 of SEQ ID NO: 5 and wherein residue 43 of SEQ ID NO: Sis a methionine (V43M). In some embodiments of the methods and compositions as disclosed herein, the IGF?2 targeting peptide comprises V43M and further comprises a Al-7 deletion (IGF2-A1-7,V43M).
[00237] In some embodiments of the methods and compositions as disclosed herein, the lysosomal IGF2 targeting peptide further comprises one or more modifications selected from any of: A2-3, A2-4, A2-5, A2-6, A2-8, A2-9, A2-10, A2-11, A212, A2-13, A2-14, A2-15, A2-16, A2-17, A2-18, A2-19, A2- 20, A2-21, A222, A2-23, A2-24, A2-25, A2-26, A2-2T, A2-28, A2-29, A2-30, A2-31, A2-32, A2-33, A2-34, A2-35, A2-36, A2-37, A2-38, A2-39, A2-40, A2-41 or A2-42 of SEQ ID NO: 5 and wherein residue 43 of SEQ ID NO: 5 is a methionine (V43M). In some embodiments of the methods and compositions as disclosed herein, the IGF2 targeting peptide comprises V43M and further comprises a A2-7 deletion (IGF2-A2-7,V43M).
[00238] In some embodiments, a IGF2 targeting peptide for fusion to a GAA-polypeptide can comprise amino acids 8-28 and 41-61 of IGF2. In some embodiments, these stretches of amino acids can be joined directly or separated by a linker. Alternatively, amino acids 8-28 and 41-61 can be provided on separate polypeptide chains. In some embodiments, amino acids 8-28 of IGF2, or a conservative substitution variant thereof, could be fused to GAA polypeptide to express a IGF2-GAA fusion protein from the rAVV vector, and a separate rAAV vector could express IGF2 amino acids 41-61, or a conservative substitution variant thereof.
[00239] In order to facilitate proper presentation and folding of the IGF2 targeting peptide, longer portions of IGF2 proteins can be used. For example, an IGF2 targeting peptide including amino acid residues 1-67, 1-87, or the entire precursor form can be used.
[00240] In some embodiments, the IGF2 targeting peptide is a nucleic acid sequence that encodes an IGF?2 targeting peptide of any of the following: residue 1 followed by residues 8-67 of wild-type mature human insulin-like growth factor II (IGF2) of SEQ ID NO: 5 (i.e., SEQ ID NO: 6; i.c., IGF2- delta 2-7); residues 8-67 of wild-type mature human insulin-like growth factor II (IGF2) of SEQ ID NO: 5 (i.e., SEQ ID NO: 7; IGF2-delta 1-7) or residues 43-67 of wild-type mature human insulin-like growth factor II (IGF2) of SEQ ID NO: 5 (i.e., IGF2-V43M (SEQ ID NO: 9) or IGF-delta 1-42 (SEQ ID NO: 8).
[00241] In some embodiments of the methods and compositions as disclosed herein, the IGF2 targeting peptide is a nucleic acid sequence selected from any nucleic acid sequence comprising any of: SEQ ID NO: 2 (i.e., IGF2-delta 2-7); SEQ ID NO: 3 (i.e., IGF2-delta 1-7) or SEQ ID NO: 4 (i.e., IGF2-V43M) or a sequence at least sequence at least 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto.
[00242] In some embodiments of the methods and compositions as disclosed herein, the IGF2(V43M) sequence is a nucleic acid sequence encoding a IGF2(V43M) sequence of any of SEQ ID NO: 65 (IGF2A2-7V43M) or an amino acid sequence having at least 85%, or 90%, or 95% or 96%, or 97%, or 98% or 99% or 100% identity to SEQ ID NO: 65, or SEQ ID NO: 66 (IGFAL- 7V43M) or an amino acid sequence having at least 85%, or 90%, or 95% or 96%, or 97%, or 98% or 99% or 100% identity to SEQ ID NO: 66. Table 3: Exemplary nucleic acid sequences encoding IGF2 targeting peptide: [1GF2 targeting | Sequence - peptide IGF2-delta 2-7 | GCTICTGTGCGGCGGGGAGCTGGTGGACACCCTCCAGTICGTCTGTGGGGA (IGFA2-7) CCGCGGCTTCTACTTCAGCAGGCCCGCAAGCCGTGTGAGCCGTCGCAGCC GTGGCATCGTTGAGGAGTGCTGTTTCCGCAGCTGTGACCTGGCCCTCCTGG AGACGTACTGTGCTACCCCCGCCAAGTCCGAG) (SEQ ID NO: 2) 1GF2-delta 1-7 | CTGTGCGGCGGGGAGCTGGTGGACACCCTCCAGTICGTCTGTGGGGACCG (IGFAI-T) CGGCTTCTACTTCAGCAGGCCCGCAAGCCGTGTGAGCCGTCGCAGCCGTG GCATCGTTGAGGAGTGCTGTTTCCGCAGCTGTGACCTGGCCCTCCTGGAG ACGTACTGTGCTACCCCCGCCAAGTCCGAG (SEQ ID NO: 3) [TGF2VaM GCTTACCGCCCCAGTGAGACCCTGTGCGGCGGGGAGCTGGTGGACACCCT CCAGTTCGTCTGTGGGGACCGCGGCTTCTACTTCAGCAGGCCCGCAAGCC GTGTGAGCCGTCGCAGCCGTGGCATCATGGAGGAGTGCTGTTTCCGCAGC TGTGACCTGGCCCTCCTGGAGACGTACTGTGCTACCCCCGCCAAGTCCGA | G (SEQ ID NO: 4)
[00243] In some embodiments, in order to facilitate proper presentation and folding of the IGF2 targeting peptide, longer portions of IGF2 proteins can be used. For example, an IGF2 targeting peptide including amino acid residues 1-67, 1-87, or the entire precursor form can be used.
[00244] In some embodiments of the methods and compositions disclosed herein, the recombinant AAV comprises a heterologous nucleic acid sequence encoding a signal peptide-GAA (SP-GAA) fusion polypeptide further comprises a IGF2 targeting peptide located between the secretory signal peptide (SP) and the an alpha-glucosidase (GAA) polypeptide.
[00245] In some embodiments of the methods and compositions disclosed herein, the recombinant AAV vector comprises a heterologous nucleic acid sequence that encodes a IGF2 targeting peptide which binds human cation-independent mannose-6-phosphate receptor (CI-MPR) or the IGF2 receptor, for example, the heterologous nucleic acid sequence encodes a IGF2 targeting peptide having the amino acid sequence of SEQ ID NO: 5 or comprises at least one amino modification in SEQ ID NO: 5 that binds to the IGF2 receptor. In some embodiments, the recombinant AAV vector comprises a heterologous nucleic acid sequence that encodes a IGF2 targeting peptide that has at least one amino modification in SEQ ID NO: 5 is a V43M amino acid modification (SEQ ID NO: 8 or SEQ ID NO: 9) or A2-7 (SEQ ID NO: 6) or Al-7 (SEQ ID NO: 7), or is a IGF2 peptide having at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to SEQ ID NOs: 5-9.
[00246] In some embodiments of the methods and compositions disclosed herein, the nucleic acid encoding a IGF2 targeting peptide is selected from any of SEQ ID NO: 2 (IGF2-A2-7), SEQ ID NO: 3 (IGF2-A1-7), or SEQ ID NO: 4 (IGF2 V43M), or a nucleic acid sequence at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to any of SEQ ID NOs: 2, 3 or 4.
[00247] In some embodiments of the compositions and methods described herein, the IGF2 targeting peptide is a nucleic acid sequence that encodes any of: residue 1 followed by residues 8-67 of wild-type mature human insulin-like growth factor II (IGF2) of SEQ ID NO: 5 (i.e., IGF2-delta 2-7 or IGF2A2-7; which corresponds to SEQ ID NO: 6); residues 8-67 of wild-type mature human insulin-like growth factor II (IGF2) of SEQ ID NO: 5 (i.e., IGF2-delta 1-7 or IGF2A1-7, which corresponds to SEQ ID NO: 7;) or residues 43-67 of wild-type mature human insulin-like growth factor IT (IGF2) of SEQ ID NO: 5 (i.e., IGF2 delta 1-42 or IGF2A1-42, which corresponds to SEQ ID NO: 8). In some embodiments of the compositions and methods described herein, the IGF2 targeting peptide is a nucleic acid sequence that has a modification of amino acid residue 43, for example residue 43 is modified to a start codon, for example IGF2-V43M (corresponding to SEQ ID NO: 9).
[00248] In some embodiments of the compositions and methods described herein, the IGF2 targeting peptide is a nucleic acid sequence comprising any of: SEQ ID NO: 2 (i.e., IGF2-delta 2-7); SEQ ID NO: 3 (i.e., IGF2-delta 1-7) or SEQ ID NO: 4 (i.e., IGF2-V43M).
[00249] In some embodiments of the compositions and methods described herein, the fusion protein comprising the GAA polypeptide and a IGF2 targeting peptide comprises amino acid residues 40-952 or residues 70-952 of human acid alpha-glucosidase (GAA) polypeptide (SEQ ID NO: 10) that is attached to an IGF2 targeting peptide that comprises residue 1 followed by residues 8-67 of wild-type mature human insulin-like growth factor II (IGF2) (SEQ ID NO: 5), (that is - residues 2-7 of mature human IGF2 (SEQ ID NO:5) are not present), wherein the IGF2 targeting peptide is linked to amino acid residue 70 of human GAA (SEQ ID NO: 10).
[00250] In some embodiments of the compositions and methods described herein, the fusion protein comprising the GAA polypeptide and a IGF2 targeting peptide comprises amino acid residues 40-952 or residues 70-952 of human acid alpha-glucosidase (GAA) polypeptide (SEQ ID NO: 10) that is attached to an IGF2 targeting peptide that comprises residues 8-67 of wild-type mature human insulin-like growth factor II (IGF2) (SEQ ID NO: 5), (that is - residues 1-7 of mature human IGF2 (ie, Y RP SET; SEQ ID NO: 63) are not present), wherein the IGF2 targeting peptide is linked to amino acid residue 70 of human GAA (SEQ ID NO: 10).
[00251] In some embodiments of the compositions and methods described herein, the fusion protein comprising the GAA polypeptide and a IGF2 targeting peptide comprises amino acid residues 40-952 or residues 70-952 of human acid alpha-glucosidase (GAA) (SEQ ID NO: 10) that is attached to a modified IGF2 targeting peptide that comprises residues 43-67 of wild-type mature human insulin- like growth factor II (IGF2) (SEQ ID NO: 5), (where residues 1-42 of mature human IGF2 (SEQ ID NO: 5) are not present), and where the IGF2 targeting peptide is linked to amino acid residue 70 of human GAA (SEQ ID NO: 10).
[00252] In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises a heterologous nucleic acid sequence that encodes an IGF2 peptide, where the IGF2 peptide sequence is SEQ ID NO: 8 or SEQ ID NO: 9, or a IGF2 peptide having at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to SEQ ID NO: 8 or 9. (ii) Modified IGF2 targeting peptides and IGF2 homologues
[00253] In some embodiments, the nucleic acid encoding IGF2 can be modified to diminish their affinity for IGFBPs, and / or decreasing affinity for binding to IGF-I receptor, thereby increasing targeting to the lysosomes and increasing the bioavailability of the fused GAA-polypeptide.
[00254] IGF?2 targeting peptide preferably specifically targets and binds to the M6P receptor. Particularly useful are IGF2 targeting peptides which have mutations in the IGF2 polypeptide that result in a protein that binds the CI-MPR / M6P receptor with high affinity while no longer binding the other two receptors with appreciable affinity.
[00255] IGF2(V43M) targeting peptide is preferably targeted specifically to the M6P receptor. Particularly useful are IGF2(V43M) targeting peptides which have mutations in the IGF2 polypeptide that result in a protein that binds the CI-MPR / M6P receptor with high affinity while no longer binding the other two receptors with appreciable affinity.
[00256] The IGF2(V43M) targeting peptide can also be modified to minimize binding to serum IGF-binding proteins (IGFBPs) (Baxter (2000) Am. J. Physiol Endocrinol Metab. 278(6):967-76) and to IGF-I receptor, in order to avoid sequestration of IGF2 constructs. A number of studies have localized residues in IGF-1 and IGF2 necessary for binding to IGF-binding proteins. Constructs with mutations at these residues can be screened for retention of high affinity binding to the M6P / IGF2 receptor and for reduced affinity for IGF-binding proteins. For example, replacing Phe 26 of IGF2 with Ser is reported to reduce affinity of IGF2 for IGFBP-1 and -6 with no effect on binding to the ME6P / IGF2 receptor (Bach et al. (1993) J. Biol. Chem. 268(13):9246-54). Other substitutions, such as Ser for Phe 19 and Lys for Glu 9, can also be advantageous. The analogous mutations, separately or in combination, in a region of IGF-I that is highly conserved with IGF2 result in large decreases in IGF- BP binding (Magee et al. (1999) Biochemistry 38(48): 15863-70).
[00257] The IGF2 targeting peptide can also be modified to minimize binding to serum IGF-binding proteins (IGFBPs) and to IGF-I receptor, in order to avoid sequestration of IGF2 constructs.
[00258] In some embodiments, a IGF2 targeting peptide is modified to be furin resistant, i.e., resistant to degradation by furin protease, which recognizes Arg-X-X-Arg cleavage sites. Such IGF2 targeting peptides are disclosed in US application 22012 / 0213762 which is incorporated herein in its entirety by reference. In some embodiments, a furin resistant IGF2 targeting peptide for use in a rAAV genome as described herein contains a mutation within a region corresponding to amino acids 30-40 (e.g., 31-40, 32-40, 33-40, 34-40, 30-39, 31-39, 32-39, 34-37, 32-39, 33-39, 34-39, 35-39, 36- 39, 37-40, 34-40) of SEQ ID NO: 5 (wt IGF?2 targeting peptide) can be substituted with any other amino acid or deleted. For example, substitutions at position 34 may affect furin recognition of the first cleavage site. Insertion of one or more additional amino acids within each recognition site may abolish one or both furin cleavage sites. Deletion of one or more of the residues in the degenerate positions may also abolish both furin cleavage sites.
[00259] In some embodiments, a furin-resistant IGF2 targeting peptide contains amino acid substitutions at positions corresponding to Arg37 (R37) or Arg40 (R40) of SEQ ID NO:5. In some embodiments, a furin-resistant IGF2 targeting peptide contains a Lys (K) or Ala (A) substitution at positions Arg37 or Arg40 of SEQ ID NO: 5. Other substitutions are possible, including combinations of Lys and / or Ala mutations at both positions 37 and 40, or substitutions of amino acids other than Lys (K) or Ala (A). In some embodiments, the IGF2 targeting peptide encompassed for use in the rAVV genome as disclosed herein is IGFA2-7-K37, or IGFA2-7-K40 or IGFA1-7-K37 or IGFA1-7- K40, indicating that the IGF2 targeting peptides has a deletion of aa 2-7 or 1-7 and a modification of a Arg (R) residue at position 37 to a lysine (i.e., R37K modification) or R40K respectively. In some embodiments, the IGF2 targeting peptide encompassed for use in the rAVV genome as disclosed herein is IGFA2-7-K37-K40, or IGFA1-7-R37K-R40K indicating that the IGF2 targeting peptides has a deletion of residues 2-7 or residues 1-7 and a modification of a R residue at position 37 and position 40 to lysinines (R37K and R40K). In some embodiments, the IGF2 targeting peptide encompassed for use in the rAVV genome as disclosed herein is selected from any of: IGFA2-7-R37A, or IGFA2- 7-R40A or IGFA1-7-R37A or IGFA1-7-R40A, IGFA2-7-R37A-R40A, or IGFA1-7-R37A-R40A. Exemplary constructs for the IGF2 targeting peptide encompassed for use in the rAVV genome as disclosed herein are disclosed in US application 2012 / 0213762, which is incorporated herein in its entirety by reference.
[00260] In some embodiments, the furin-resistant IGF2 targeting peptide suitable for the invention may contain additional mutations. For example, up to 30% or more of the residues of SEQ ID NO: 5 may be changed (e.g., up to 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30% or more residues may be changed). Thus, a furin-resistant IGF2 mutein suitable for the invention may have an amino acid sequence at least 70%, including at least 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99%, identical to SEQ ID NO: 5.
[00261] Moreover, use of a IGF2 targeting peptide as disclosed herein is also referred to in the art as Glycosylation Independent Lysosomal Targeting (GILT) because the IGF2 targeting peptide replaces MB6P as the moiety targeting the lysosomes. Details of the GILT technology are described in U.S. Application Publication Nos. 2003 / 0082176, 2004 / 0006008, 2004 / 0005309, 2003 / 0072761, 2005 / 0281805, 2005 / 0244400, and international publications WO 03 / 032913, WO 03 / 032727, WO 02 / 087510, WO 03 / 102583, WO 2005 / 078077, the disclosures of all of which are hereby incorporated by reference.
[00262] Other modifications to the amino acid sequence of the IGF2 targeting peptide for use in the methods and compositions as disclosed herein are disclosed in US provisional application 62,937,556, filed on November 19, 2019 and in PCT application PCT / US19 / 61653, filed Nov 15, 20109, both of which are incorporated herein in their entirety by reference.
[00263] IGF2 binds to the IGF2 / M6P and IGF-I receptors with relatively high affinity and binds with lower affinity to the insulin receptor. Substitution of IGF2 residues 48-50 (Phe Arg Ser) with the corresponding residues from insulin, (Thr Ser Ile), or substitution of residues 54-55 (Ala Leu) with the corresponding residues from IGF-I (Arg Arg) result in diminished binding to the IGF2 / M6P receptor but retention of binding to the IGF-I and insulin receptors (Sakano et al. (1991) J. Biol. Chem. 266(31):20626-35).
[00264] IGF2 binds to repeat 11 of the cation-independent M6P receptor. Indeed, a minireceptor in which only repeat 11 is fused to the transmembrane and cytoplasmic domains of the cation- independent M6P receptor is capable of binding IGF2 (with an affinity approximately one tenth the affinity of the full length receptor) and mediating internalization of IGF2 and its delivery to lysosomes (Grimme et al. (2000) J. Biol. Chem. 275(43):33697-33703). The structure of domain 11 of the M6P receptor is known (Protein Data Base entries 1GP0 and 1GP3; Brown et al. (2002) EMBO J. 21(5):1054-1062). The putative IGF2 binding site is a hydrophobic pocket believed to interact with hydrophobic amino acids of IGF2; candidate amino acids of IGF2 include leucine 8, phenylalanine 48, alanine 54, and leucine 55. Although repeat 11 is sufficient for IGF2 binding, constructs including larger portions of the cation-independent M6P receptor (e.g. repeats 10-13, or 1-15) generally bind IGF2 with greater affinity and with increased pH dependence (see, for example, Linnell et al. (2001) J. Biol. Chem. 276(26):23986-23991).
[00265] Substitution of IGF2 residues Tyr 27 with Leu, or Ser 26 with Phe diminishes the affinity of IGF2 for the IGF-I receptor by 94-, 56-, and 4-fold respectively (Torres et al. (1995) J. Mol. Biol. 248(2):385-401). Deletion of residues 1-7 of human IGF2 resulted in a 30-fold decrease in affinity for the human IGF-I receptor and a concomitant 12-fold increase in affinity for the rat IGF2 receptor (Hashimoto et al. (1995) J. Biol. Chem. 270(30):18013-8). Truncation of the C-terminus of IGF2 (residues 62-67) also appear to lower the affinity of IGF2 for the IGF-I receptor by 5 fold (Roth et al. (1991) Biochem. Biophys. Res. Commun. 181(2):907-14).
[00266] Substitution of IGF2 residue phenylalanine 26 with serine reduces binding to IGFBPs 1-5 by 5-75 fold (Bach et al. (1993) J. Biol. Chem. 268(13):9246-54). Replacement of IGF2 residues 48- 50 with threonine-serine-isoleucine reduces binding by more than 100 fold to most of the IGFBPs (Bach et al. (1993) J. Biol. Chem. 268(13):9246-54); these residues are, however, also important for binding to the cation-independent mannose-6-phosphate receptor. The Y27L substitution that disrupts binding to the IGF-I receptor interferes with formation of the ternary complex with IGFBP3 and acid labile subunit (Hashimoto et al. (1997) J. Biol. Chem. 272(44):27936-42); this ternary complex accounts for most of the IGF2 in the circulation (Yu et al. (1999) J. Clin. Lab Anal. 13(4):166-72). Deletion of the first six residues of IGF2 also interferes with IGFBP binding (Luthi et al. (1992) Eur. J. Biochem. 205(2):483-90).
[00267] Studies on IGF-I interaction with IGFBPs revealed additionally that substitution of serine for phenylalanine 16 did not affect secondary structure but decreased IGFBP binding by between 40 and 300 fold (Magee et al. (1999) Biochemistry 38(48):15863-70). Changing glutamate 9 to lysine also resulted in a significant decrease in IGFBP binding. Furthermore, the double mutant lysine 9 / serine 16 exhibited the lowest affinity for IGFBPs. The conservation of sequence between this region of IGF-I and IGF2 suggests that a similar effect will be observed when the analogous mutations are made in IGF2 (glutamate 12 lysine / phenylalanine 19 serine).
[00268] In some embodiments, the IGF2(V43M) sequence comprises at least amino acids 48-55; at least amino acids 8-28 and 41-61; or at least amino acids 8-87, or a sequence variant thereof (e.g. R68A) or truncated form thereof (e.g. C-terminally truncated from position 62) that binds the cation- independent mannose-6-phosphate receptor.
[00269] In another embodiment of the invention, the rAAV genome encoding the targeting peptide (e.g., IGF2 targeting peptide) is inserted into the native GAA coding sequence at the junction of the mature 70 / 76 kDal polypeptide and the C-terminal domain, for example at position 791. This creates a single chimeric polypeptide. In some embodiments, a protease cleavage site may be inserted just downstream of the targeting peptide (e.g., IGF2 targeting peptide).
[00270] In one embodiment, a targeting peptide, e.g., IGF2 targeting peptide as defined herein, is fused directly to the N- or C-terminus of the GAA polypeptide. In another embodiment, a IGF2 targeting peptide is fused to the N- or C-terminus of the GAA polypeptide by a spacer. In one specific embodiment, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer of 10-25 amino acids. In another embodiment, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer including glycine residues.
[00271] In some embodiments, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer of at least 1, 2, or 3 amino acids. In some embodiments, the spacer comprises amino acids GAP or Gly-Ala-Pro (SEQ ID NO: 31), or an amino acid sequence at least 50% identical thereto. In some embodiments, the spacer is GGG or GA or AP, or GP or variants thereof. In some embodiments, the spacer is encoded by nucleic acids GGC GCG CCG (SEQ ID NO: 30).
[00272] In some embodiments, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer including a helical structure. In another specific embodiment, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer at least 50% identical to the sequence GGGTVGDDDDK (SEQ ID NO: 35). In some embodiments of the methods and compositions as disclosed herein, the spacer is SEQ ID NO: 31 (encoded by nucleic acids of SEQ ID NO: 30). In some embodiments of the methods and compositions as disclosed herein, the spacer is selected from any of: SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34 or SEQ ID NO: 35, or a sequence at least sequence at least 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto. (iii) Alternative targeting peptides that bind to the Cation-independent M6P Receptor (CI-MPR).
[00273] In some embodiments, the targeting peptide is a lysosomal targeting peptide or protein, or other moiety other than the IGF2 targeting peptide disclosed herein that binds to the cation independent M6P / IGF2 receptor (CI-MPR) in a mannose-6-phosphate-independent manner. The CI- MPR also contains binding sites for at least three distinct ligands that can be used as targeting peptides. As disclosed herein, IGF2 ligand binds to CI-MPR with a dissociation constant of about 14 nM at or about pH 7.4, primarily through interactions with repeat 11. The CI-MPR is capable of binding high molecular weight O-glycosylated IGF2 forms. Accordingly, in some embodiments, the IGF2 targeting peptide can be post-transcriptionally modified to comprises O-glycosylation.
[00274] In an altemative embodiment, the targeting peptide that binds to CI-MPR is retinoic acid. Retinoic acid binds to the receptor with a dissociation constant of 2.5 nM. Affinity photolabeling of the cation-independent M6P receptor with retinoic acid does not interfere with IGF2 or M6P binding to the receptor, indicating that retinoic acid binds to a distinct site on the receptor. Binding of retinoic acid to the receptor alters the intracellular distribution of the receptor with a greater accumulation of the receptor in cytoplasmic vesicles and also enhances uptake of M6P modified B-glucuronidase. Retinoic acid has a photoactivatable moiety that can be used to link it to a therapeutic agent without interfering with its ability to bind to the cation-independent M6P receptor.
[00275] The urokinase-type plasminogen receptor (uPAR) also binds CI-MPR with a dissociation constant of 9 uM. uPAR is a GPI-anchored receptor on the surface of most cell types where it functions as an adhesion molecule and in the proteolytic activation of plasminogen and TGF-p. Binding of uPAR to the CI-M6P receptor targets it to the lysosome, thereby modulating its activity. Thus, fusing the extracellular domain of uPAR, or a portion thereof competent to bind the cation- independent M6P receptor, to a therapeutic agent permits targeting of the agent to a lysosome. D. Spacer and fusion junction of the GAA polypeptide
[00276] Where GAA is expressed as a fusion protein with a secretory signal peptide (e.g., SS-GAA fusion polypeptide) or with a targeting peptide (i.e., SS-IGF2-GAA polypeptide double fusion polypeptide), the signal peptide or IGF2 targeting peptide can be fused directly to the GAA polypeptide or can be separated from the GAA polypeptide by a linker. An amino acid linker (also referred to herein as a “spacer”) incorporates one or more amino acids other than that appearing at that position in the natural protein. Spacers can be generally designed to be flexible or to interpose a structure, such as an a-helix, between the two protein moieties.
[00277] Accordingly, in some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises a heterologous nucleic acid sequence encoding an IGF2-GAA fusion polypeptide, wherein the IGF2-GAA fusion protein further comprises a spacer comprising a nucleotide sequence of at least 1 amino acid in length, which is located N-terminal to the GAA polypeptide, and C-terminal to the IGF2 targeting peptide. In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises a heterologous nucleic acid sequence that comprises a nucleic acid encoding a spacer of at least 1 amino acids located between the nucleic acid encoding the IGF2 targeting peptide and the nucleic acid encoding the GAA polypeptide.
[00278] In one embodiment, the IGF2 targeting peptide is fused directly to the N- or C-terminus of the GAA polypeptide. In another embodiment, a IGF2 targeting peptide is fused to the N- or C- terminus of the GAA polypeptide by a spacer. In one specific embodiment, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer of 10-25 amino acids. In another specific embodiment, a IGF?2 targeting peptide is fused to the GAA polypeptide by a spacer including glycine residues. In another specific embodiment, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer including a helical structure. In another specific embodiment, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer at least 50% identical to the sequence GGGTVGDDDDK (SEQ ID NO: 35).
[00279] In some embodiments, a spacer or linker can be relatively short, e.g., atleast 1,2,3,4 or 5 amino acids, or such as the sequence Gly-Ala-Pro (SEQ ID NO: 31) or Gly-Gly-Gly-Gly-Gly-Pro (SEQ ID NO: 32), or can be longer, such as, for example, 5-10 amino acids in length or 10-25 amino acids in length. For example, flexible repeating linkers of 3-4 copies of the sequence (GGGGS (SEQ ID NO:33)) and a-helical repeating linkers of 2-5 copies of the sequence (EAAAK (SEQ ID NO:34)) have been described (Arai et al. (2004) Proteins: Structure, Function and Bioinformatics 57:829-838).
[00280] The use of another linker, GGGTVGDDDDK (SEQ ID NO: 35), in the context of an IGF2 fusion protein has also been reported (DiFalco et al. (1997) Biochem. J. 326:407-413) and is encompassed for use. Linkers incorporating an a-helical portion of a human serum protein can be used to minimize immunogenicity of the linker region.
[00281] In some embodiments, the spacer is encoded by nucleic acids GGC GCG CCG (SEQ ID NO: 30) which encodes the amino acid spacer comprising amino acids GAP or Gly-Ala-Pro (SEQ ID NO: 31).
[00282] The site of a fusion junction in the GAA polypeptide to fuse with either the signal peptide (to generate a SS-GAA fusion protein) or with the targeting peptide (e.g., to generate a SP-IGF2-GAA double fusion polypeptide) should be selected with care to promote proper folding and activity of each polypeptide in the fusion protein and to prevent premature separation of a signal peptide from a GAA polypeptide.
[00283] In some embodiments, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer including a helical structure. In another specific embodiment, a IGF2 targeting peptide is fused to the GAA polypeptide by a spacer at least 50% identical to the sequence GGGTVGDDDDK (SEQ ID NO: 35). In some embodiments of the methods and compositions as disclosed herein, the spacer is SEQ ID NO: 31 (encoded by nucleic acids of SEQ ID NO: 30). In some embodiments of the methods and compositions as disclosed herein, the spacer is selected from any of: SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34 or SEQ ID NO: 35.
[00284] Four exemplary strategies for creating a IGF2-GAA fusion protein can be generated, which are disclosed in provisional application 62,937,556, filed on November 19, 2019, and in PCT / US19 / 61653, filed Nov 15, 2019, which are incorporated herein in their entirety by reference.
[00285] In some embodiments, a targeting peptide (e.g., a IGF2 targeting peptide) can be fused, directly or by a spacer, to amino acid 40 or amino acid 70 of GAA, a position permitting expression of the protein, catalytic activity of the GAA protein, and proper targeting by the IGF2 targeting peptide as described herein in the Examples. Alternatively, a targeting peptide (e.g., a IGF2 targeting peptide) can be fused at or near the cleavage site separating the C-terminal domain of GAA from the mature polypeptide. This permits synthesis of a GAA protein with an internal targeting peptide (e.g., a IGF2 targeting peptide), which optionally can be cleaved to liberate the mature polypeptide or the C- terminal domain from the targeting domain, depending on placement of cleavage sites. Alternatively, the mature polypeptide can be synthesized as a fusion protein at about position 791 without incorporating C-terminal sequences in the open reading frame of the expression construct.
[00286] In order to facilitate folding of the IGF2 targeting peptide, GAA amino acid residues adjacent to the fusion junction can be modified. For example, since it is possible that GAA cysteine residues may interfere with proper folding of the targeting peptide (e.g., a IGF2 targeting peptide), the terminal GAA cysteine 952 can be deleted or substituted with serine to accommodate a C-terminal targeting peptide (e.g., a IGF2 targeting peptide). The targeting peptide (e.g., a IGF2 targeting peptide) can also be fused immediately preceding the final Cys952. The penultimate cys938 can be changed to proline in conjunction with a mutation of the final Cys952 to serine. E. CS sequence
[00287] In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises a heterologous nucleic acid sequence that further comprises at collagen stability (CS) sequence located 3° of the nucleic acid encoding the GAA polypeptide and 5° of the 3° ITR sequence. In some embodiments, the rAAV genome disclosed herein comprises a heterologous nucleic acid sequence that can optionally comprise a Collagen stability sequence (CS or CSS), which is positioned 3° of the GAA gene and 5’ of a polyA signal. In some embodiments, the CS sequence can be replaced by a 3° UTR sequence as disclosed herein.
[00288] Exemplary collagen stability sequences include CCCAGCCCACTTTTCCCCAA (SEQ ID NO: 65) or a sequence at least 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto. An exemplary collagen stability sequence can have an amino acid sequence of P S PL F P (SEQ ID NO: 66) or an amino acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto. CS sequences are disclosed in Holick and Liebhaber, Proc. Nat. Acad. Sci. 94: 2410- 2414, 1997 (See, e.g. Figure 3, p. 5205), which is incorporated herein its entirety by reference. F. Promoters
[00289] In some embodiments, to achieve appropriate levels of GAA expression, the AAV genotype comprises a liver specific promoter (LSP). A LSP enables expression of the operatively linked gene in the liver, and can in some embodiments, be and inducible LSP. In an embodiment, a LSP is located upstream 5° and is operatively linked to the heterologous nucleic acid sequence encoding the GAA protein. Exemplary liver-specific promoters are disclosed herein, and include for example, the LSP comprising SEQ ID NO: 86, 91-96 or 146-150, or functional variant or functional fragment thereof, or any LSP listed in Table 4 herein, or a functional fragment or functional variants thereof. In some embodiments of the compositions and methods disclosed herein, a liver-specific promoter includes a liver-specific cis-regulatory element (CRE), a synthetic liver-specific cis- regulatory module (CRM) or a synthetic liver-specific promoter is selected from any of SEQ ID NO: 270-341 (minimal LSP with CRM) or SEQ ID NO: 342-430 (synthetic liver specific proximal promoters) disclosed in Table 4 herein). In some embodiments, an rAAV vector genome can include one or more constitutive promoters, such as viral promoters or promoters from mammalian genes that are generally active in promoting transcription. (i) Synthetic Liver-specific promoters
[00290] In some embodiments of the methods and compositions as disclosed herein, the promoter is a liver specific promoter, and can be selected from promoters including, but not limited to, those listed in Table 4 disclosed herein or functional variants thereof, and or any selected from Tables 4A and 4B of U.S. provisional application 62,937,556, filed on November 19, 2019, or functional variants thereof
[00291] While transthyretin promoter (TTR) (SEQ ID NO: 431) and SP0412 (SEQ ID NO: 91) and SP0422 (SEQ ID NO: 92) are used as an exemplary liver specific promoters (see Examples 1, 12 and 13) in the specification and Examples, one of ordinary skill in the art can readily replace TTR with any liver specific promoter as disclosed herein in Table 4 or functional variants thereof, and or any selected from Tables 4A and 4B of U.S. provisional application 62,937,556, filed on November 19, 2019, or functional variants thereof. A liver-specific promoter can comprise a liver-specific cis- regulatory element (CRE), a synthetic liver-specific cis-regulatory module (CRM) or a synthetic liver- specific promoter as disclosed herein, in Tables 4A and 4B of U.S. provisional application 62,937,556, filed on November 19, 2019, or functional variants thereof,
[00292] Table 4 shows exemplary liver-specific promoters. The relatively small size of liver- specific promoters disclosed herein is advantageous because it takes up the minimal amount of the payload of the vector. This is particularly important when a LSP is used in a vector with limited capacity, such as an AAV-based vector.
[00293] Table 4: Exemplary LSP identified by SEQ ID NOs for use in the methods and compositions as disclosed herein Table 4 : Exemplary LSP Rein Name of LSP SEQ | Seqip Name of LSP 1D | Name of LSP Sealy | Name of LSP SEQ ID Name of LSP NO: 271 | CRM SP0109 331 | CRM SP0397 385 | CRM_SP0397 273 | CRM_SPO111 332 | CRM_SP0398 386 274 | CRM SPO112 333 | CRM SP0399 387 | CRM_SP0112 | CRM_SP0399 388 275 | CRM_SP0113 334 | CRM_SP0403 270 | CRM_SP0107 330 CRM_SP0396 384 SD vi 27m | CRM_SP0109 331 CRM_SP0397 sw (TUR 273 | crRM spot 332 | crM_spo30s 386 3 Al 274 | CRM_SPO112 333 | CRM_SP0399 387 au i 275 | CRM_SPO113 334 CRM_SP0403 983 Sw vi 276 | CRM_SPO115 335 | CRM_SP0404 59 | SPO256 276 | CRM_SP0115 335 | CRM_SP0404 389 390 | SP0257 277 | CRM_SP0116 336 | CRM_SP0405 278 | CRM_SP0121 337 | CRM_SP0406 391 | SP0258 338 Jo 279 | CRM_SP0124 | CRM_SP0407 392 | SP0259 CRM SP0409 SP0264 SP0265(LVR SPI CRM SP0411 31 Al) CRM SP0412 SP0266(LVR_SP1 31 VD) CRM SP0413 SP0267(LVR_SP1 31 V2) SP0107 SP0268 (LVR 132 Al) SP0109 SP0269 (LVR _132 VI) SPOL11 SP0270 (LVR 132 V2) SPOL12 SP0271 (LVR 133 Al) SPO113 SP0272 (LVR _133 VI) (LVR _133 V2) SPO116 SP0368 SP0121 SP0373 SP0124 SP0378 SP0273 CRM_SP0158 SPO115 LVR 1 CRM SP0163 SPO116 SP0368 CRM SP0236 SP0121 SP0373 CRM_SP0239 SPO124 SP0378 SP0127 | CRM_SP0240 (LVR SP127) SP0379 SPOI27A1 | CRM_SP0241 (LVR_SP127_Al | SP0380 SPO127V1 | CRM_SP0242 (LVR_SP127_V1 | SP0381 SP0127V2 SP0272 | CRM_SP0243 (LVR_SP127 V2 | ig (LVR 133 VI) SP0273 | CRM_SP0244 SP0128 (LVR 1 (LVR 133 V2) SP0131 | CRM_SP0246 (LVR SP131) SP0368 SP0132 | CRM_SP0247 (LVR SP132) SP0373 SP0378 LVR SP133 SP0155 SP0379 SP0O158 SP0380 SP0163 SP0381 SP0236 SP0384 SP0239 SP0388 SP0240 SP0396 SP0241 SP0397 SP0133 CRM_SP0248 Es CRM_SP0249 SP0155 CRM_SP0250 SP0158 CRM_SP0251 SP0163 CRM_SP0252 SP0236 CRM_SP0253 SP0239 CRM_SP0254 SP0240 CRM_SP0255 SP0241 310 | CRM SP0258 366 | SP0244 423 31 | CRM SP0259 96 | SP0246 424 312 | CRM SP0264 367 | SP0247 425 313 | ev ote 1 368 | SP0248 426 308 CRM _SP0256 364 | SP0242 421 SP0398 309 CRM _SP0257 365 | SP0243 422 SP0399 310 CRM SP0258 366 | SP0244 423 SP0403 3m CRM_SP0259 96 SP0246 424 SP0404 312 CRM _SP0264 367 | SP0247 425 SP0405 CRM_SP0265 313 (CRM LVR 131 Al) | 368 | SP0248 426 SP0406 364 | SP0242 365 | SP0243 366 | SP0244 96 SP0246 367 | SP0247 | 368 | SP0248 CRM_SP0266 t 314 | CRM LVR 131 vi) | 36° | spo249 427 | soso CRM_SP0267 y < 315 | (CRM LVR 131 vo | 37 | spozs0 428 | sposos CRM_SP0268 3 n 316 jen, ona an | 371 | spoasi a2 | spoatl CRM_SP0269 317 | crm LVR 132 vy | 372 | spo2s2 91 | spoa12 CRM_SP0270 v 318 | (CRM LVR 132 vy | 373 | SP0253 92 | SP0422 CRM_SP0271 319 | RM LVR 133 An | 374 | SP0254 146 | SP0265-UTR CRM_SP0272 . 320 | ceMLvR 133 vy | 373 | spozss 147 | spozzo-utr 323 CRM SP0373 378 | SP0258 150 324 | CRM SP0378 379 | SP0259 430 CRM_SP0273 321 CRM LVR 133 V2 376 | SP0256 148 SP0240-UTR 322 CRM _SP0368 377 | SP0257 149 SP0246-UTR 323 CRM SP0373 378 | SP0258 150 SP0131-A1-UTR 324 CRM _SP0378 379 | SP0259 430 SP0413 325 CRM SP0379 380 | SP0264 431 TTR promoter 308 CRM _SP0256 364 | SP0242 421 SPO: 309 CRM _SP0257 365 | SP0243 422 SPO: 310 CRM SP0258 366 | SP0244 423 SP04 3m CRM_SP0259 96 SP0246 424 SP0A 312 CRM _SP0264 367 | SP0247 425 SP04 CRM_SP0265 313 CRM LVR 131 Al 368 | SP0248 426 SP0A CRM_SP0266 314 CRM LVR 131 VI 369 SP0249 427 | SP0A CRM_SP0267 < 315 CRM LVR 131 V2 370 | SP0250 428 | SP04 CRM_SP0268 x 316 CRM LVR 132 Al 371 | SP0251 429 | SP04 CRM_SP0269 317 CRM LVR 132 VI 372 SP0252 91 | SP0A CRM_SP0270 318 CRM LVR 132 V2. 373 SP0253 92 | SP0A CRM_SP0271 319 CRM LVR 133 Al 374 | SP0254 146 | SP02 CRM_SP0272 x 320 CRM LVR 133 VI 375 | SP0255 147 | SP02 CRM_SP0273 321 CRM LVR 133 V2 376 | SP0256 148 SPO2 322 CRM _SP0368 377 | SP0257 149 SP02 323 CRM SP0373 378 | SP0258 150 SPO] 324 CRM _SP0378 379 | SP0259 430 SP04 325 CRM SP0379 380 | SP0264 431 TTR SP02635 326 | CRM_SP0380 94 Je sores 432 | LPI 327 | CRM_SP0381 381 | ” RALYE. Se | 433 | CMV-IE SP0267(LVR_SP 328 2 | crM_spo3s4 382 | Sp | 434 | cBa SP0268 329 | crM_spo3ss 383 | TER AD 435 | TBG promoter 1 SP0266(LVR_SP 131 V1 ) SP0267(LVR_SP 131 V2 3 SP0268 LVR 132 Al (ii) Functional variants of Liver-Specific promoters
[00294] In some embodiments, the synthetic liver-specific promoter useful in the methods and compositions as disclosed herein is a bi-specific, or tri-specific promoter as defined herein. As an illustrative example, a liver bi-specific promoter is active in the liver and one other tissue, for example, the muscle. Additionally, another illustrative example of a liver bi-specific promoter is active in the liver and one other tissue, e.g., the brain. As an illustrative example of a liver tri-specific promoter is active in the liver and two other tissues, for example, the muscle and brain. Additionally, another illustrative example of a liver tri-specific promoter is active in the liver and two other tissues, such as, e.g., the kidney and muscle.
[00295] In some embodiments, a synthetic liver specific promoter that is at least 50%, 60%, 70%, 80%, 90% or 95% identical to any of SEQ ID NO: 86, 91-96, 146-150, 270-430 comprises a source regulatory nucleic acid sequence which is preferentially active in liver, and is also active to a lesser extent (e.g., <50%, or about 49-40%, or about 39-30%, or about 29-20% or about 19-10% or <10% of total expression) in a second type of cell or tissue, e.g., muscle or CNS.
[00296] In some embodiments, the promoter is a synthetic liver-specific promoter comprising a combination of the cis-regulatory elements (CREs) CRE0051 (SEQ ID NO: 97) and CRE0042 (SEQ ID NO: 104), or functional variants thereof. Functional variants thereof may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. Typically, the CREs are operably linked to a promoter element. In some preferred embodiments, the liver-specific promoter comprises said CREs, or functional variants thereof, in the order CRE0051 (SEQ ID NO: 97), CRE0042 (SEQ ID NO: 104), and then the promoter element (order is given in an upstream to downstream direction, as is conventional in the art).
[00297] The promoter element can be any suitable proximal promoter or minimal promoter. In some embodiments, the promoter element is a minimal promoter. Where the promoter is a proximal promoter, it is generally preferred that the proximal promoter is liver-specific.
[00298] In some preferred embodiments, the promoter element is CRE0059 (SEQ ID NO: 110), ora functional variant thereof. CRE0059 is a proximal promoter, as is discussed further below.
[00299] Thus, in one embodiment the promoter comprises the following regulatory elements: CRE0051 (SEQ ID NO: 97), CRE0042 (SEQ ID NO: 104) and CRE0059 (SEQ ID NO: 110), or functional variants thereof.
[00300] Functional variants of CRE0051 (SEQ ID NO: 97) are regulatory elements with sequences which vary from CRE0051, but which substantially retain activity as liver-specific CREs. It will be appreciated by the skilled person that it is possible to vary the sequence of a CRE while retaining its ability to bind to the requisite transcription factors (TFs) and enhance expression. A functional variant can comprise substitutions, deletions and / or insertions compared to a reference CRE, provided they do not render the CRE substantially non-functional.
[00301] In some embodiments, a functional variant of CRE0051 can be viewed as a CRE which, when substituted in place of CRE0051 in a promoter, substantially retains its activity. For example, a liver-promoter which comprises a functional variant of CRE0051 substituted in place of CRE0051 preferably retains 80% of its activity, more preferably 90% of its activity, more preferably 95% of its activity, and yet more preferably 100% of its activity. For example, considering promoter SP0412 (SEQ ID NO: 91) as an example, CRE0051 in SP0412 can be replaced with a functional variant of CREO0051, and the promoter substantially retains its activity. Retention of activity can be assessed by comparing expression of a suitable reporter under the control of the reference promoter with an otherwise identical promoter comprising the substituted CRE under equivalent conditions.
[00302] In some embodiments the functional variant of CRE0051 comprises transcription factor binding sites (TFBS) for the same liver-specific TFs as CRE0051. The liver-specific TFBS present in CREO0051, listed in the order in which they are present, are: HNF1 (SEQ ID NO: 98), HNF4 (SEQ ID NO: 99), HNF3 (SEQ ID NO: 100), HNF1’ (SEQ ID NO: 101) and HNF3’ (SEQ ID NO: 102), see Table 5. The functional variant of CRE0051 thus preferably comprises all of these TFBS. Preferably, they are present in the same order that they are present in CRE0051, i.e. in the order HNF1 (SEQ ID NO: 98), HNF4 (SEQ ID NO: 99), HNF3 (SEQ ID NO: 100), HNF1’ (SEQ ID NO: 101) and HNF3’ (SEQ ID NO: 102),. When the cis-regulatory element is associated with a promoter and gene, this order is preferably considered in an upstream to downstream direction (i.e. in the direction from distal from the transcription start site (TSS) to proximal to the TSS). Spacer sequences may be provided between adjacent TFBS. In some embodiments the TFBS may suitably overlap, provided they remain functional, i.e. overlapping sequences are both able to bind their respective TFs to the extent required to regulate expression.
[00303] In some embodiments the functional variant of CRE0051 (SEQ ID NO: 97) comprises the following TFBS sequences: GTTAATTTTTAAA (HNF1) (SEQ ID NO: 98), GTGGCCCTTGG (HNF4) (SEQ ID NO: 99), TGTTTGC (HNF3) (SEQ ID NO: 100), TGGTTAATAATCTCA (HNF1’) (SEQ ID NO: 101) then ACAAACA (HNF3) (SEQ ID NO: 102), sequences complementary thereto, or functional variants of these TFBS sequences that maintain the ability to bind to their respective TF. These may be present in the same order as CRE0051, i.e. the order in which they are set out above. It is well-known in the art that there is sequence variability associated with TFBS, and that for a given TFBS there is typically a consensus sequence, from which some degree of deviation is typically present. Further information about the variation that occurs in a TFBS can be illustrated using a positional weight matrix (PWM), which represents the frequency with which a given nucleotide is typically found at a given location in the consensus sequence. Details of TF consensus sequences and associated PWM:s can be found in, for example, the Jaspar or Transfac databases (http: / / jaspar.genereg.net / and http: / / gene-regulation.com / pub / databases.html). This information allows the skilled person to modify the sequence in any given TFBS of a CRE in a manner which retains, and in some cases even increases, CRE functionality.
[00304] In some embodiments, the functional variant of CRE0051 comprises the sequence:
[00305] GTTAATTTITTAAA-Na-GTGGCCCTTGG-Nb-TGTTTGC-Nc-TGGTTAATAATCTCA- Nd-ACAAACA (SEQ ID NO: 103), or a sequence that is at least 70%, 80%, 90%, 95% or 99% identical thereto, wherein Na, Nb, Nc, and Nd represent optional spacer sequences. When present, Na optionally has a length of from 10 to 26 nucleotides, preferably from 14 to 22 nucleotides, and more preferably 18 nucleotides. When present, Nb optionally has a length of from 8 to 22 nucleotides, preferably from 12 to 20 nucleotides, more preferably 16 nucleotides. When present, Nc optionally has a length of from 1 to 10 nucleotides, preferably 1 to 5 nucleotides, and more preferably 2 nucleotides. When present, Nd suitably has a length of from 1 to 13 nucleotides, preferably from 2 to 9 nucleotides in length, and more preferably 5 nucleotides in length.
[00306] In some embodiments, the CRE consists of SEQ ID No: 98-102 or a functional variant thereof.
[00307] It will be noted that the CRE or functional variant thereof can be provided on either strand of a double stranded polynucleotide and can be provided in either orientation. As such, complementary and reverse complementary sequences of SEQ ID NO: 97-102 or a functional variant thereof fall within the scope of the invention. Single stranded nucleic acids comprising the sequence according to SEQ ID NO: 97 or 103 or a functional variant thereof also fall within the scope of the invention.
[00308] In some embodiments, the CRE comprising or consisting of CRE0051 (SEQ ID NO: 97), or a functional variant thereof, has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, or 100 or fewer nucleotides.
[00309] In some embodiments, the CRE comprising or consisting of CRE0042 (SEQ ID NO: 104) or a functional variant thereof, has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, or 100 or fewer nucleotides. Functional variants thereof may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[00310] Functional variants of CRE0042 (SEQ ID NO: 104) are regulatory elements with sequences which vary from CRE0042, but which substantially retain their activity as liver-specific CREs. It will be appreciated by the skilled person that it is possible to vary the sequence of a CRE while retaining its ability to bind to the requisite transcription factors (TFs) and enhance expression. A functional variant can comprise substitutions, deletions and / or insertions compared to a reference CRE, provided they do not render the CRE substantially non-functional,
[00311] In some embodiments, a functional variant of CRE0042 (SEQ ID NO: 104) can be viewed as a CRE which, when substituted in place of CRE0042 in a promoter, substantially retains its activity. For example, a promoter which comprises a functional variant of CRE0042 substituted in place of CRE0042 preferably retains 80% of its activity, more preferably 90% of its activity, more preferably 95% of its activity, and yet more preferably 100% of its activity (compared to the reference promoter comprising CRE0042 (SEQ ID NO: 104)). For example, considering promoter SP0412 as an example, CRE0042 (SEQ ID NO: 104) in SP412 (SEQ ID NO: 91) can be replaced with a functional variant of CRE0042, and the promoter substantially retains its activity. Retention of activity can be assessed by comparing expression of a suitable reporter under the control of the reference promoter with an otherwise identical promoter comprising the substituted CRE under equivalent conditions.
[00312] In some embodiments it is preferred that the functional variant of CRE0042 (SEQ ID NO: 104) comprises TFBS for the same liver-specific TFs as CRE0042. The liver-specific TFBS present in CRE0042, listed in the order in which they are present, are: HNF-3 (SEQ ID NO: 106), C / EBP (SEQ ID NO: 107), HNF-4 (SEQ ID NO: 108) and C / EBP’ (SEQ ID NO: 109). The functional variant of CRE0042 thus preferably comprises all of these TFBS. Preferably, they are present in the same order that they are present in CRE0042, i.e. in the order HNF-3, C / EBP, HNF-4 and then C / EBP. When the cis-regulatory element is associated with a promoter and gene, this order is preferably considered in an upstream to downstream direction (i.e. in the direction from distal from the transcription start site (TSS) to proximal to the TSS). Spacer sequences may be provided between adjacent TFBS. In some embodiments the TFBS may suitably overlap, provided they remain functional, i.e. overlapping sequences are both able to bind their respective TFs.
[00313] In some embodiments the functional variant of CRE042 (SEQ ID NO: 104) comprises the following TFBS sequences: GTTCAAACATG (HNF-3) (SEQ ID NO: 106), CTAATACTCTG (C / EBP) (SEQ ID NO: 107), TGCAAGGGTCAT (HINF-4) (SEQ ID NO: 108), and TTACTCAACA (C / EBP) (SEQ ID NO: 109) and sequences complementary thereto, or functional variants of these TFBS sequences that maintain the ability to bind to their respective TF. These may be present in the same order as CRE0042, i.e. the order in which they are set out above. As discussed above, it is well- known in the art that there is sequence variability associated with TFBS, and that for a given TFBS there is typically a consensus sequence, from which some degree of deviation is typically present.
[00314] In some embodiments of the invention, the functional variant of CRE0042 comprises the sequence:
[00315] GTTCAAACATG-Na-CTAATACTCTG-Nb-TGCAAGGGTCAT-Nc-TTACTCAACA (SEQ ID NO: 105) or a sequence that is at least 70%, 80%, 90%, 95% or 99% identical thereto, wherein Na, Nb and Nc represent optional spacer sequences. When present, Na optionally has a length of from 1 to 10 nucleotides, preferably from 1 to 5 nucleotides, and more preferably 2 nucleotides. When present, Nb optionally has a length of from 1 to 10 nucleotides, preferably from 2 to 6 nucleotides, and more preferably 4 nucleotides. When present, Nc optionally has a length of from 8 to 23 nucleotides, preferably from 10 to 20 nucleotides, and more preferably 15 nucleotides.
[00316] In some embodiments of the invention the cis-regulatory enhancer element consists of CRE0042 (SEQ ID NO: 104) or a functional variant thereof.
[00317] It will be noted that the CRE or functional variant thereof can be provided on either strand of a double stranded polynucleotide and can be provided in either orientation. As such, complementary and reverse complementary sequences of SEQ ID NO: 104 or 105 or a functional variant thereof fall within the scope of the invention. Single stranded nucleic acids comprising the sequence according to SEQ ID NO: 104 or 105 or a functional variant thereof also fall within the scope of the invention. In some embodiments, the CRE comprising or consisting of CRE0042 (SEQ ID NO: 104), ora functional variant thereof, has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, 100 or fewer nucleotides, or 80 or fewer nucleotides.
[00318] In some embodiments, the CRE comprising or consisting of CRE0059 (SEQ ID NO: 110) or a functional variant thereof, has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, or 100 or fewer nucleotides. Functional variants thereof may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto
[00319] As discussed above, functional variants of CRE0059 (SEQ ID NO: 110) substantially retain the ability of CRE00059 to act as a liver-specific promoter element. For example, when a functional variant of CRE0059 is substituted into liver-specific promoter SP0412, the modified promoter retains at least 80% of its activity, more preferably at least 90% of its activity, more preferably at least 95% of its activity, and yet more preferably 100% of the activity of SP0412 (SEQ ID NO: 91). Suitably the functional variant of CRE0059 comprises a sequence which has at least 70%, 80%, 90%, 95% or 99% identity to SEQ ID NO: 110.
[00320] CREO059 is a proximal promoter and comprises a TFBS for a liver-specific TF, namely HNF1, upstream of the TSS. The functional variant of CRE0059 thus preferably comprises a TFBS for HNF1 upstream of the TSS.
[00321] In some embodiments, a functional variant of CRE0059 comprises a sequence which is at least 70% identical to SEQ ID NO: 110 (preferably at least 80%, 90%, 95% or 99% identical to SEQ ID NO: 110), which contains a TFBS for HNF1 (SEQ ID NO: 111), and which contains a TSS sequence (referred to as pl@SERPINAL or pl@AFP) which is at least 80%, 90%, 95% or completely identical to SEQ ID NO: 112 downstream of said TFBS for HNF1.
[00322] In some embodiments, a functional variant of CRE0059 comprises a sequence which has at least 70%, 80%, 90%, 95% or 99% identity to SEQ ID NO: 110, and which further comprises a TFBS comprising SEQ ID NO: 111 for HNF1 at or near position 24-36; and which comprises the TSS sequence which is at least 80%, 90%, 95% or completely identical to SEQ ID NO: 112 at or near position 73-93, positions being numbered with reference to SEQ ID NO: 110. At or near in the present context suitably means within 10, 5, 4, 3, 2, or 1 nucleotide of the recited position with reference to SEQ ID NO: 110. Suitable TFBS sequences are SEQ ID NOS 111 and SEQ ID NO: 112, but alternative TFBS sequences can be used.
[00323] In some embodiments, a promoter element comprising or consisting of CRE0059 (SEQ ID NO: 110) or a functional variant thereof has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, 110 or fewer nucleotides, or 95 or fewer nucleotides.
[00324] In some embodiments the liver-specific promoter useful in the methods and compositions as disclosed herein comprises or consists of SEQ ID NO: 91, or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. The promoter having a sequence according to SEQ ID NO: 91 is referred to as SP0412. The SP0412 promoter is particularly preferred in some embodiments. This promoter has been found to be powerful and is also very short, which is advantageous in some circumstances. a. SP0265 (also known as SP131A1) and variants thereof
[00325] In some embodiments, the promoter is a synthetic liver-specific promoter comprising a combination of the CREs CRE0051 (SEQ ID NO: 97), CRE0058 (SEQ ID NO: 113), CRE0065 (SEQ ID NO: 117), and CRE0066 (SEQ ID NO: 122), or functional variants thereof. Typically, the CREs are operably linked to a promoter element. In some preferred embodiments, the liver-specific promoter comprises said CREs, or functional variants thereof, in the order CRE0051, CRE0058, CRE0065, CRE0066, and then the promoter element (in an upstream to downstream direction).
[00326] The promoter element can be any suitable proximal or minimal promoter. In some preferred embodiments, the promoter element is a minimal promoter. Where the promoter is a proximal promoter, it is generally preferred that the proximal promoter is liver-specific.
[00327] In some preferred embodiments, the promoter element is CRE0052 (also referred to as G6PC) (SEQ ID NO: 126). CRE0052 is a minimal promoter (also referred to as a core promoter).
[00328] In some embodiments, the liver-specific promoter comprises the following regulatory elements (or functional variants thereof): CRE0051, CRE0058, CRE0065, CRE0066 then CRE0052 (SEQ ID NO: 126). The sequence of CRE0051 (SEQ ID NO: 97) and variants thereof are set out above.
[00329] In some embodiments, the CRE comprising or consisting of CRE0058 (SEQ ID NO: 113), or a functional variant thereof, has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, 100 or fewer nucleotides, or 80 or fewer nucleotides. Functional variants thereof may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[00330] Functional variants of CRE0058 (SEQ ID NO: 113) are regulatory elements with sequences which vary from CRE0058, but which substantially retain their activity as liver-specific CREs. It will be appreciated by the skilled person that it is possible to vary the sequence of a CRE while retaining its ability to bind to the requisite TFs and enhance expression. A functional variant can comprise substitutions, deletions and / or insertions compared to a reference CRE, provided they do not render the CRE non-functional .
[00331] In some embodiments, a functional variant of CRE0058 (SEQ ID NO: 113) can be viewed as a CRE which, when substituted in place of CRE0058 in a promoter, substantially retains its activity. For example, a promoter which comprises a functional variant of CRE0058 substituted in place of CRE0058 preferably retains 80% of its activity, more preferably 90% of its activity, more preferably 95% of its activity, and yet more preferably 100% of its activity (compared to the reference promoter comprising CRE0058 (SEQ ID NO: 113)). For example, considering promoter SP0265 (SEQ ID NO: 94) as an example, CRE0058 in SP0265 can be replaced with a functional variant of CREO0058, and the promoter substantially retains its activity. Retention of activity can be assessed by comparing expression of a suitable reporter under the control of the reference promoter with an otherwise identical promoter comprising the substituted CRE under equivalent conditions.
[00332] In some embodiments it is preferred that the functional variant of CRE0058 (SEQ ID NO: 113) comprises transcription factor binding sites (TFBS) for the same liver-specific transcription factors (TF) as CRE0058. The liver-specific TFBS present in CRE0058, listed in the order in which they are present, are: HNF4 (SEQ ID NO: 115) and ¢ / EBP (SEQ ID NO: 116). The functional variant of CRE0058 thus preferably comprises all of these TFBS. Preferably, they are present in the same order that they are present in CRE0058, i.e. in the order HNF4 then ¢ / EBP. When the CRE is associated with a promoter and gene, this order is preferably considered in an upstream to downstream direction (i.e. in the direction from distal from the transcription start site (TSS) to proximal to the TSS). Spacer sequences may be provided between adjacent TFBS. In some embodiments the TFBS may suitably overlap, provided they remain functional, i.e. overlapping sequences are both able to bind their respective TFs.
[00333] In some embodiments, the functional variant of CRE0058 (SEQ ID NO: 113) comprises the following TFBS sequences: CGCCCTTTGGACC (HNF4) (SEQ ID NO: 115) and GACCTTTTGCAATCCTGG (c / EBP) (SEQ ID NO: 116), sequences complementary thereto, or functional variants of these TFBS sequences that maintain the ability to bind to their respective TF. These may be present in the same order as CRE0038, i.e. the order in which they are set out above. As discussed above, it is well-known in the art that there is sequence variability associated with TFBS, and that for a given TFBS there is typically a consensus sequence, from which some degree of deviation is typically present.
[00334] In some embodiments, the functional variant of CRE0058 comprises the sequence: GCGCCCTTTGGACCTTTTGCAATCCTGG (SEQ ID NO: 114), or a sequence that is at least 70%, 80%, 90%, 95% or 99% identical thereto. In some embodiments, the CRE consists of SEQ ID NO: 113 or 114 or a functional variant thereof.
[00335] It will be noted that the CRE or functional variant thereof can be provided on either strand of a double stranded polynucleotide and can be provided in either orientation. As such, complementary and reverse complementary sequences of SEQ ID NO: 113 or 114 or a functional variant thereof fall within the scope of the invention. Single stranded nucleic acids comprising the sequence according to SEQ ID NO: 113 or 114, or a functional variant thereof, also fall within the scope of the invention.
[00336] In some embodiments, the CRE comprising or consisting of CRE0058, or a functional variant thereof, has a length of 120 or fewer nucleotides, 80 or fewer nucleotides, 60 or fewer nucleotides, or 40 or fewer nucleotides.
[00337] In some embodiments, the CRE comprising or consisting of CRE0065 (SEQ ID NO: 117), or a functional variant thereof, has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, 100 or fewer nucleotides, or 80 or fewer nucleotides. Functional variants thereof may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 0999 identical thereto.
[00338] Functional variants of CRE0065 (SEQ ID NO: 117) are regulatory elements with sequences which vary from CRE0065, but which substantially retain their activity as liver-specific CREs. It will be appreciated by the skilled person that it is possible to vary the sequence of a CRE while retaining its ability to bind to the requisite TFs and enhance expression. A functional variant can comprise substitutions, deletions and / or insertions compared to a reference CRE, provided they do not render the CRE non-functional .
[00339] In some embodiments, a functional variant of CRE0065 can be viewed as a CRE which, when substituted in place of CRE0065 in a promoter, substantially retains its activity. For example, a promoter which comprises a functional variant of CRE0065 substituted in place of CRE0065 preferably retains 80% of its activity, more preferably 90% of its activity, more preferably 95% of its activity, and yet more preferably 100% of its activity (compared to the reference promoter comprising CRE0065). For example, considering promoter SP0265 (SEQ ID NO: 94) as an example, CRE0065 in SP0265 can be replaced with a functional variant of CRE0065, and the promoter substantially retains its activity. Retention of activity can be assessed by comparing expression of a suitable reporter under the control of the reference promoter with an otherwise identical promoter comprising the substituted CRE under equivalent conditions.
[00340] In some embodiments it is preferred that the functional variant of CRE0065 comprises TFBS for the same liver-specific TFs as CRE0065. The liver-specific TFBS present in CRE0065, listed in the order in which they are present, are: RXR Alpha (SEQ ID NO: 119), HNF3 (SEQ ID NO: 120) and HNF3 (SEQ ID NO: 121). The functional variant of CRE0065 thus preferably comprises all of these TFBS. Preferably, they are present in the same order that they are present in CRE0065, i.e. in the order RXR Alpha, HNF3 then HNF3. When the cis-regulatory element is associated with a promoter and gene, this order is preferably considered in an upstream to downstream direction (i.e. in the direction from distal from the transcription start site (TSS) to proximal to the TSS). Spacer sequences may be provided between adjacent TFBS. In some embodiments the TFBS may suitably overlap, provided they remain functional, i.e. overlapping sequences are both able to bind their respective TFs.
[00341] In some embodiments, the functional variant of CRE0065 comprises the following TFBS sequences: ACTGAACCCTTGACCCCTGCCCT (RXR Alpha) (SEQ ID NO: 119), CTGTTTGCCC (HNF3) (SEQ ID NO: 120), and CTATTTGCCC (HNF3) (SEQ ID NO: 121), sequences complementary thereto, or functional variants of these TFBS sequences that maintain the ability to bind to their respective TF. These may be present in the same order as CRE0065, i.e. the order in which they are set out above. As discussed above, it is well-known in the art that there is sequence variability associated with TFBS, and that for a given TFBS there is typically a consensus sequence, from which some degree of deviation is typically present.
[00342] In some embodiments, the functional variant of CRE0065 comprises the sequence:
[00343] ACTGAACCCTTGACCCCT-Na-CTGTTTGCCC-Nb-TATTTGCCC (SEQ ID NO: 118), or a sequence that is at least 70%, 80%, 90%, 95% or 99% identical thereto, wherein Na and Nb represent optional spacer sequences. When present, Na optionally has a length of from 14 to 30 nucleotides, preferably from 18 to 26 nucleotides, and more preferably 22 nucleotides. When present, Nb optionally has a length of from 1 to 10 nucleotides, preferably from 2 to 6 nucleotides, and more preferably 4 nucleotides. In some embodiments, the CRE consists of SEQ ID NO: 117 or 118, ora functional variant thereof.
[00344] It will be noted that the CRE or functional variant thereof can be provided on either strand of a double stranded polynucleotide and can be provided in either orientation. As such, complementary and reverse complementary sequences of SEQ ID NO: 117 or 118 or a functional variant thereof fall within the scope of the invention. Single stranded nucleic acids comprising the sequence according to SEQ ID NO: 117 or 118 or a functional variant thereof also fall within the scope of the invention.
[00345] In some preferred embodiments, the CRE comprising or consisting of CRE0065, or a functional variant thereof, has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, 90 or fewer nucleotides, or 72 or fewer nucleotides.
[00346] In some embodiments, the CRE comprising or consisting of CRE0066 (SEQ ID NO: 122), or a functional variant thereof, has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, 100 or fewer nucleotides, or 80 or fewer nucleotides Functional variants thereof may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 099, identical thereto.
[00347] Functional variants of CRE0066 (SEQ ID NO: 122) are regulatory elements with sequences which vary from CRE0066, but which substantially retain their activity as liver-specific CREs. It will be appreciated by the skilled person that it is possible to vary the sequence of a CRE while retaining its ability to bind to the requisite TFs and enhance expression. A functional variant can comprise substitutions, deletions and / or insertions compared to a reference CRE, provided they do not render the CRE non-functional
[00348] In some embodiments, a functional variant of CRE0066 can be viewed as a CRE which, when substituted in place of CRE0066 in a promoter, substantially retains its activity. For example, a promoter which comprises a functional variant of CRE0066 substituted in place of CRE0066 preferably retains 80% of its activity, more preferably 90% of its activity, more preferably 95% of its activity, and yet more preferably 100% of its activity (compared to the reference promoter comprising CRE0066 (SEQ ID NO: 122). For example, considering promoter SP0265 (SEQ ID NO: 94) as an example, CRE0066 in SP0265 can be replaced with a functional variant of CRE0066, and the promoter substantially retains its activity. Retention of activity can be assessed by comparing expression of a suitable reporter under the control of the reference promoter with an otherwise identical promoter comprising the substituted CRE under equivalent conditions.
[00349] In some embodiments, it is preferred that the functional variant of CRE0066 comprises transcription factor binding sites (TFBS) for the same liver-specific transcription factors (TF) as CRE0066. The liver-specific TFBS present in CRE0066, listed in the order in which they are present, are: HNF4G (SEQ ID NO: 124) and FOS::JUN (SEQ ID NO: 125). The functional variant of CREO0066 thus preferably comprises all of these TFBS. Preferably, they are present in the same order that they are present in CRE0066, i.e. in the order HNF4G then FOS:: JUN. When the cis-regulatory element is associated with a promoter and gene, this order is preferably considered in an upstream to downstream direction (i.e. in the direction from distal from the transcription start site (TSS) to proximal to the TSS). Spacer sequences may be provided between adjacent TFBS. In some embodiments the TFBS may suitably overlap, provided they remain functional, i.e. overlapping sequences are both able to bind their respective TFs.
[00350] In some embodiments, the functional variant of CRE0066 (SEQ ID NO: 122) comprises the following TFBS sequences: GCAGGGCAAAGTGCA (HNF4G) (SEQ ID NO: 124) and GATGACTCAG (FOS: JUN) (SEQ ID NO: 125), sequences complementary thereto, or functional variants of these TFBS sequences that maintain the ability to bind to their respective TF. These may be present in the same order as CRE0066, i.e. the order in which they are set out above. As discussed above, it is well-known in the art that there is sequence variability associated with TFBS, and that for a given TFBS there is typically a consensus sequence, from which some degree of deviation is typically present.
[00351] In some embodiments, the functional variant of CRE0066 (SEQ ID NO: 122) comprises the sequence: GCAGGGCAAAGTGCA-Na-GATGACTCAG (SEQ ID NO: 123) or a sequence that is at least 70%, 80%, 90%, 95% or 99% identical thereto, wherein Na represents an optional spacer sequence. When present, Na optionally has a length of from 10 to 28 nucleotides, preferably from 14 to 24 nucleotides, and more preferably 19 nucleotides. In some embodiments, the CRE consists of CRE0066 or a functional variant thereof.
[00352] It will be noted that the CRE or functional variant thereof can be provided on either strand of a double stranded polynucleotide and can be provided in either orientation. As such, complementary and reverse complementary sequences of SEQ ID NO: 122 or 123 or a functional variant thereof fall within the scope of the invention. Single stranded nucleic acids comprising the sequence according to SEQ ID NO: 122 or 123, or a functional variant thereof, also fall within the scope of the invention.
[00353] In some preferred embodiments, the CRE comprising or consisting of CRE0066 or a functional variant thereof has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, 100 or fewer nucleotides, or 87 or fewer nucleotides.
[00354] In some embodiments, the promoter comprises the promoter element CRE0052 (also referred to as G6PC) (SEQ ID NO: 126) or a functional variant or functional fragment thereof. Functional variants thereof may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[00355] Functional variants of CRE0052 (SEQ ID NO: 126) substantially retain the ability of CRE0052 to act as a liver-specific promoter element. For example, when a functional variant of CREO0052 is substituted into liver-specific promoter SP0265, the modified promoter retains at least 80% of its activity, more preferably at least 90% of its activity, more preferably at least 95% of its activity, and yet more preferably 100% of the activity of SP0265.
[00356] In one embodiment the liver-specific promoter comprises SEQ ID NO: 94, or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 94. The promoter having a sequence according to SEQ ID NO: 94 is referred to as SP0265 (also known as SP131A1 or LVR 131_Al). A promoter comprising or consisting of SEQ ID NO: 94 is particularly preferred in some embodiments.
[00357] In some embodiments, the liver-specific promoter is SEQ ID NO: 94 and comprises the following components: CRE0051 (SEQ ID NO: 97); CRE0058 (SEQ ID NO: 113); CRE0065 (SEQ ID NO: 117), CRE0066 (SEQ ID NO: 122), CRE0052 (SEQ ID NO: 126) ; or functional variants of SEQ ID NO: 97, 113, 117, 122 or 126 which may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. b.SP0239 and variants thereof
[00358] In some embodiments, the promoter is a synthetic liver-specific promoter comprising the following CREs: CRE0018 (SEQ ID NO: 151), CRE0051 (SEQ ID NO: 97), CRE0058 (SEQ ID NO: 113), CRE0065 (SEQ ID NO: 117) and CRE0066 (SEQ ID NO: 122), or functional variants thereof. Typically, the CREs are operably linked to a promoter element. In some preferred embodiments, the liver-specific promoter comprises said CREs, or functional variants thereof, in the order CRE0018, CRE0051, CRE0058, CRE0065, CRE0066, and then the promoter element (in an upstream to downstream direction).
[00359] The promoter element can be any suitable proximal or minimal promoter. In some preferred embodiments the promoter element is CRE0052 (also referred to as G6PC). CRE0052 is a minimal promoter (also referred to as a core promoter).
[00360] In some embodiments the liver-specific promoter comprises the following elements (or functional variants thereof): CRE0018, CRE0051, CRE0058, CRE0065, CRE0066 and then CRE0052.
[00361] The sequences of CRE0051, CRE0058, CRE0065, and CRE0066 and the promoter element CREO0052, and functional variants thereof, are set out above.
[00362] CREO0018 has the sequence of SEQ ID NO: 151 or a functional variant or functional fragment thereof. Functional variants thereof may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[00363] Functional variants of CRE0018 (SEQ ID NO: 151) are regulatory elements with sequences which vary from CRE0018, but which substantially retain their activity as liver-specific CREs. It will be appreciated by the skilled person that it is possible to vary the sequence of a CRE while retaining its ability to bind to the requisite TFs and enhance expression. A functional variant can comprise substitutions, deletions and / or insertions compared to a reference CRE, provided they do not render the CRE substantially non-functional.
[00364] In some embodiments, a functional variant of CRE0018 can be viewed as a CRE which, when substituted in place of CRE0018 in a promoter, substantially retains its activity. For example, a promoter which comprises a functional variant of CRE0018 substituted in place of CRE0018 preferably retains 80% of its activity, more preferably 90% of its activity, more preferably 95% of its activity, and yet more preferably 100% of its activity (compared to the reference promoter comprising CRE0018). For example, considering promoter SP0239 as an example, CRE0018 in SP0239 in can be replaced with a functional variant of CRE0018, and the promoter substantially retains its activity. Retention of activity can be assessed by comparing expression of a suitable reporter under the control of the reference promoter with an otherwise identical promoter comprising the substituted CRE under equivalent conditions,
[00365] In some embodiments, the functional variant of CRE0018 comprises TFBS for the same liver-specific TFs as CRE0018. The liver-specific TFBS present in CRE0018, listed in the order in which they are present, are: IRF (SEQ ID NO: 129), NF1 (SEQ ID NO: 130), HNF3 (SEQ ID NO: 131), HBLF (SEQ ID NO: 132), RXRa (SEQ ID NO: 133), EF-C (SEQ ID NO: 134), NF1 (SEQ ID NO: 135), and ¢ / EBP (SEQ ID NO: 136). The functional variant of CRE0018 thus preferably comprises all of these TFBS. Preferably, they are present in the same order that they are present in CREO0018, i.e. in the order IRF, NF1, HNF3, HBLF, RXRa, EF-C, NF], and then ¢ / EBP. When the CRE is associated with a promoter and gene, this order is preferably considered in an upstream to downstream direction (i.e. in the direction from distal from the transcription start site (TSS) to proximal to the TSS). Spacer sequences may be provided between adjacent TFBS. In some embodiments the TFBS may suitably overlap, provided they remain functional, i.e. overlapping sequences are both able to bind their respective TFs.
[00366] In some embodiments the functional variant of CRE0018 comprises the following TFBS sequences: CTTTCACTTTC (IRF) (SEQ ID NO: 129), TCGCCAA (NF1) (SEQ ID NO: 130), TGTGTAAACA (HNF3) (SEQ ID NO: 131), TGTAAACAATA (HBLF) (SEQ ID NO: 132), CTGAACCTTTACCC (RXRa) (SEQ ID NO: 133), GTTGCCCGGCAAC (EF-C) (SEQ ID NO: 134), CAGGTCTGTGCCAAG (NF1) (SEQ ID NO: 135), TGCCAAGTGTTTG (c / EBP) (SEQ ID NO: 136), sequences complementary thereto, or functional variants of these TFBS sequences that maintain the ability to bind to their respective TF of SEQ ID NO: 129-136. These may be present in the same order as CRE0018, i.e. the order in which they are set out above. As discussed above, it is well-known in the art that there is sequence variability associated with TFBS, and that for a given TFBS there is typically a consensus sequence, from which some degree of deviation is typically present.
[00367] In some embodiments of the invention, the functional variant of CRE0018 comprises the sequence: CTTTCACTTTCTCGCCAA-Na-TGTGTAAACAATA-Nb-CTGAACCTTTACCC-Nc- GTTGCCCGGCAAC-Nd-CAGGTCTGTGCCAAGTGTTTG (SEQ ID NO: 128), or a sequence that is at least 70%, 80%, 90%, 95% or 99% identical thereto, wherein Na, Nb, Nc, and Nd represent optional spacer sequences. When present, Na optionally has a length of from 10 to 20 nucleotides, preferably from 13 to 17 nucleotides, and more preferably 15 nucleotides. When present, Nb optionally has a length of from I to 10 nucleotides, preferably from 1 to 5 nucleotides, more preferably 1 nucleotide. When present, Nc optionally has a length of from 1 to 10 nucleotides, preferably 1 to 5 nucleotides, and more preferably 1 nucleotide. When present, Nd suitably has a length of from 1 to 10 nucleotides, preferably from 2 to 8 nucleotides in length, and more preferably 3 nucleotides in length.
[00368] In some embodiments of the invention the CRE consists of SEQ ID NO: 127 or 128 ora functional variant thereof.
[00369] It will be noted that the CRE or functional variant thereof can be provided on either strand of a double stranded polynucleotide and can be provided in either orientation. As such, complementary and reverse complementary sequences of SEQ ID NO: 128 or 129 or a functional variant thereof fall within the scope of the invention. Single stranded nucleic acids comprising the sequence according to SEQ ID NOS: 128 or 129 or a functional variant thereof also fall within the scope of the invention.
[00370] In some embodiments, the CRE comprising or consisting of CRE0018 (SEQ ID NO: 151), or a functional variant thereof, has a length of 200 or fewer nucleotides, 150 or fewer nucleotides, 125 or fewer nucleotides, or 103 or fewer nucleotides.
[00371] In one embodiment the liver-specific promoter comprises or consist of: SEQ ID NO: 93, or a functional variant thereof. The promoter having a sequence according to SEQ ID NO: 93 is referred to as SP0239. Functional variants of SP0239 can have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[00372] Accordingly, in some embodiments, the liver-specific promoter is SP0239 (SEQ ID NO: 93) and comprises the following components: CRE0018 (SEQ ID NO: 151), CRE0051 (SEQ ID NO: 97), CRE0058 (SEQ ID NO: 113), CRE0065 (SEQ ID NO: 117) and CRE0066 (SEQ ID NO: 122, and CRE0052 (SEQ ID NO: 126); or functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. ¢. SP0240 and variants thereof
[00373] In some embodiments, the promoter is a synthetic liver-specific promoter comprising CRE0018 operably linked to a promoter element. In some preferred embodiments, the liver-specific promoter comprises CRE0018, or immediately upstream of the promoter element.
[00374] The promoter element can be any suitable proximal or minimal promoter. In some preferred embodiments the promoter element is CRE0006 (SEQ ID NO: 137). CRE0006 is a liver- specific proximal promoter.
[00375] In some embodiments the liver-specific promoter comprises the following elements (or functional variants thereof): CRE0018 and then CRE0006.
[00376] The sequence of CRE0018 and variants thereof are set out above.
[00377] CREO006 is a proximal promoter and comprises TFBS for liver-specific TFs upstream of the TSS. The liver-specific TFBS present in CRE0006, listed in order, are HNF4 (SEQ ID NO: 138), RXRa (SEQ ID NO: 139), HNF4 (SEQ ID NO: 140), ¢ / EBP (SEQ ID NO: 141), and HNF3 (SEQ ID NO: 142), and optionally pl@VTN (SEQ ID NO: 143). The functional variant of CRE0006 thus preferably comprises these TFBS. Preferably, they are present in the same order that they are present in CREQ006, i.e. in the order HNF4, ¢ / EBP, HNF3, and HNF3. In some embodiments the TFBS overlap, provided they remain functional, i.e. overlapping sequences are both able to bind their respective TFs.
[00378] pl@VTN (SEQ ID NO: 143), represents the transcription start site (TSS) in CRE0006, as determined by Cap Analysis of Gene Expression (CAGE).
[00379] In some embodiments, a functional variant of CRE0006 comprises a sequence which is at least 70% identical to SEQ ID NO: 137 (preferably at least 80%, 90%, 95% or 99% identical to SEQ ID NO: 25), which contains TFBS for HNF4, RXRa, HNF4, ¢ / EBP, and HNF3, and preferably which contains a TSS sequence which is at least 80%, 90%, 95% or completely identical to TFBS for HNF4, RXRa, HNF4, ¢ / EBP, and HNF3 downstream of said TFBS.
[00380] In some embodiments, a functional variant of CRE0006 comprises a sequence which has at least 70%, 80%, 90%, 95% or 99% identity to SEQ ID NO: 137, and which further comprises the following TFBS: HNF4 (SEQ ID NO: 138) at or near position 25-37; RXRa (SEQ ID NO: 139) at or near position 73-83; HNF4 (SEQ ID NO: 140) at or near position 74-86; ¢ / EBP (SEQ ID NO: 141) at or near position 123-136; and HNF3 (SEQ ID NO: 142) at or near position 129-137; and which comprises a TSS sequence which is at least 80%, 90%, 95% or completely identical to SEQ ID NO: 143 at or near position 166-196, positions being numbered with reference to SEQ ID NO: 137. Ator near in the present context suitably means within 10, 5, 4, 3, 2, or 1 nucleotide of the recited position with reference to SEQ ID NO: 137. Suitable TFBS sequences are SEQ ID Nos: 138-142, but altemative TFBS sequences can be used.
[00381] In one embodiment the liver-specific promoter comprises or consist of SEQ ID NO: 95, ora functional variant thereof. The promoter having a sequence according to SEQ ID NO: 95 is referred to as SP0240. Functional variants of SP0240 can have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. d. SP0246 and variants thereof
[00382] In some embodiments, the promoter is a synthetic liver-specific promoter comprising the following CREs: CRE0051, CRE0058, and CRE0065, or functional variants thereof. Typically, the CRE:s are operably linked to a promoter element. In some preferred embodiments, the liver-specific promoter comprises said CREs, or functional variants thereof, in the order CRE0051, CRE0058, and CREO0065, and then the promoter element (in an upstream to downstream direction).
[00383] The promoter element can be any suitable proximal or minimal promoter. In some preferred embodiments the promoter element is CRE0052 (also referred to as G6PC). CRE0052 is a minimal promoter (also referred to as a core promoter).
[00384] In some embodiments the liver-specific promoter comprises the following elements (or functional variants thereof): CRE0051, CRE0058, CRE0065, and then CRE0052. The sequences of CRE0051, CRE0058, CRE0065 and the promoter element CRE0052, and functional variants thereof, are set out above.
[00385] In one embodiment the liver-specific promoter comprises or consist of SEQ ID NO: 96, or a functional variant thereof. The promoter having a sequence according to SEQ ID NO: 96 is referred to as SP0246. Functional variants of SP0246 can have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 96. e. SP0131 and variants thereof
[00386] In some embodiments, the promoter is a synthetic liver-specific promoter comprising the following CREs: CRE0058, CRE0065 and CRE0066, or functional variants thereof. Typically, the CRE: are operably linked to a promoter element. In some preferred embodiments, the liver-specific promoter comprises said CREs, or functional variants thereof, in the order CRE0058, CRE0065, CRE0066 and then the promoter element (in an upstream to downstream direction).
[00387] The promoter element can be any suitable proximal or minimal promoter. In some preferred embodiments the promoter element is CRE0052 (also referred to as G6PC). CRE0052 is a minimal promoter (also referred to as a core promoter).
[00388] The sequences of CRE0058, CRE0065, and CRE0066 and the promoter element CRE0052, and functional variants thereof, are set out above.
[00389] SEQ ID NO: 141, or a functional variant thereof. The promoter having a sequence according to SEQ ID NO: 141 is referred to as SPO131. Functional variants of SP0131 can have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. f. Composite Promoters:
[00390] In some embodiments, the liver-specific promoters as set out above are operably linked to one or more additional regulatory sequences. An additional regulatory sequence can, for example, enhance expression compared to the liver-specific promoter which is not operably linked the additional regulatory sequence. Generally, it is preferred that the additional regulatory sequence does not substantively reduce the specificity of the liver-specific promoter.
[00391] For example, the liver-specific promoter can be operably linked to a sequence encoding a UTR (e.g. a 5° and / or 3° UTR), an intron, or such.
[00392] In some embodiments, the liver-specific promoter is operably linked to sequence encoding aUTR, e.g. a5’ UTR. A 5' UTR can contain various elements that can regulate gene expression. The 5° UTR in a natural gene begins at the transcription start site and ends one nucleotide before the start codon of the coding region. It should be noted that 5' UTRs as referred to herein may be an entire naturally occurring 5° UTR or it may be a portion of a naturally occurring 5° UTR. The 5’UTR can also be partially or entirely synthetic. In eukaryotes, 5' UTRs have a median length of approximately 150 nt, but in some cases they can be considerably longer. Regulatory sequences that can be found in 5' UTRs include, but are not limited to: Binding sites for proteins, that may affect the mRNA's stability or translation; Riboswitches; Sequences that promote or inhibit translation initiation; and Introns within 5' UTRs have been linked to regulation of gene expression and mRNA export.
[00393] In some embodiments, a liver-specific promoter as set out above is operably linked to a sequence encoding a 5° UTR derived from the CMV major immediate gene (CMV-IE gene). For example, the 5° UTR from the CMV-IE gene suitably comprises the CMV-IE gene exon 1 and the CMV-IE gene exon 1, or portions thereof. In some cases, the promoter element may be modified in view of the linkage to the 5 ‘UTR, for example sequences downstream of the transcription start site (TSS) in the promoter element can be removed (e.g. replaced with the 5° UTR).
[00394] The CMV-IE 5°UTR is described in Simari, et al., Molecular Medicine 4: 700-706, 1998 “Requirements for Enhanced Transgene Expression by Untranslated Sequences from the Human Cytomegalovirus Immediate-Early Gene”, which is incorporated herein by reference. Variants of the CMV-IE 5° UTR sequences discussed in Simari, et al. are also set out in W02002 / 031137, incorporated by reference, and the regulatory sequences disclosed therein can also be used. Other UTRs that can be used in combination with a promoter are known in the art, e.g. in Leppek, K., Das, R. & Barna, M. “Functional 5' UTR mRNA structures in eukaryotic translation regulation and how to find them”. Nat Rev Mol Cell Biol 19, 158-174 (2018), incorporated by reference.
[00395] In some embodiments the sequence encoding the 5° UTR comprises SEQ ID NO: 145, or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. SEQ ID NO: 145 encodes a CMV-IE 5° UTR.
[00396] In some embodiments the 5° UTR comprises a nucleic acid motif that functions as the protein translation initiation site, e.g. sequences that define a Kozak sequence in the mRNA produced. For example, in some embodiments, the sequence encoding the 5° UTR comprises the sequence motif GCCACC (SEQ ID NO: 153) at or near its 3° end. Other Kozak sequences or other protein translation initiation sites can be used, as is known in the art (e.g. Marilyn Kozak, “Point Mutations Define a Sequence Flanking the AUG Initiator Codon That Modulates Translation by Eukaryotic Ribosomes” Cell, Vol. 44, 283-292, January 31, 1986; Marilyn Kozak “At Least Six Nucleotides Preceding the AUG Initiator Codon Enhance Translation in Mammalian Cells” J. Mol. Rid. (1987) 196, 947-950; Marilyn Kozak “An analysis of 5'-noncoding sequences from 699 vertebrate messenger RNAs” Nucleic Acids Research. Vol. 15 (20) 1987, all of which are incorporated herein by reference). The protein translation initiation site (e.g. Kozak sequence) is preferably positioned immediately adjacent to the start codon.
[00397] In some embodiments the sequence encoding the 5° UTR comprises SEQ ID NO: 438, or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. This 5° UTR comprises six nucleotides of SEQ ID NO: 153 which define a Kozak sequence at the 3” end of the CMV-IE 5° UTR.
[00398] In some embodiments, the SP0412 promoter, or variants thereof, as discussed above is linked to a sequence encoding a 5° UTR to provide a composite promoter / 5” UTR regulatory construct. Herein, such composite promoter / 5” UTR constructs may be referred to simply as “composite promoters”, or in some cases simply “promoters” for brevity.
[00399] In some embodiments, the composite promoter comprises or consists of SEQ ID NO: 92, or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 92.
[00400] This composite promoter comprises SP0412 operably linked to a sequence encoding the 5° UTR from the CMV-IE gene (SEQ ID NO: 145) and the GCCACC (SEQ ID NO: 153) Kozak sequence discussed above. This (composite) promoter is referred to as SP0422 (SEQ ID NO: 92). SP0422 is a preferred liver specific promoter in some embodiments. As discussed above, the 5° UTR suitably comprises a nucleic acid motif that functions as the protein translation initiation site, e.g. sequences that define a Kozak sequence. In the sequence above, the 5° UTR comprises the sequence motif GCCACC (SEQ ID NO: 153) at its 3° end, but this sequence motif can be omitted or alternative sequences can be used.
[00401] In some embodiments, the SP0265 promoter, or variants thereof, as discussed above is linked to a sequence encoding a 5° UTR to provide a composite promoter (SP0236-5UTR).
[00402] In some embodiments, the composite promoter comprises or consists of SEQ ID NO: 146, or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 146.
[00403] This composite promoter comprises SP0265(SEQ ID NO: 94) operably linked to the 5° UTR from the CMV-IE gene (SEQ ID NO: 145) and the GCCACC (SEQ ID NO: 153) Kozak sequence. This (composite) promoter is referred to as SP0420. In this promoter, a short sequence downstream of the TSS in the CRE0052 promoter element have been replaced with sequences from the 5° UTR from the CMV-IE. Thus, this promoter actually comprises a minor variant of SP0265 with a modification to CRE0052 whereby some sequence has been removed. SP0420 is preferred in some embodiments. As discussed above, the 5° UTR suitably comprises a nucleic acid motif that functions as the protein translation initiation site, e.g. sequences that define a Kozak sequence. In the sequence above, the 5° UTR comprises the sequence motif GCCACC (SEQ ID NO: 153) at its 3’ end, but this sequence motif can be omitted or alternative sequences can be used.
[00404] In some embodiments, the SP0239 promoter, or variants thereof, as discussed above is linked to a sequence encoding a 5° UTR to provide a composite promoter (SP0239-UTR).
[00405] In some embodiments, the composite promoter comprises or consists of SEQ ID NO: 147, or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 147.
[00406] This composite promoter / 5” UTR construct comprises SP0239 operably linked to the 5° UTR from the CMV-IE gene and the GCCACC (SEQ ID NO: 153) Kozak sequence. This (composite) promoter is referred to as SP0421. Again, in this promoter, a short sequence downstream of the TSS in the CRE0052 promoter element have been replaced with sequences from the 5° UTR from the CMV-IE. Thus, this promoter actually comprises a minor variant of SP0239 with a modification to CRE0052 whereby some sequence has been removed. SP0421 is preferred in some embodiments. As discussed above, the 5° UTR suitably comprises a nucleic acid motif that functions as the protein translation initiation site, e.g. sequences that define a Kozak sequence. In the sequence above, the 5° UTR comprises the sequence motif GCCACC (SEQ ID NO: 153) at its 3” end, but this sequence motif can be omitted or alternative sequences can be used.
[00407] In some embodiments, the SP0240 promoter, or variants thereof, as discussed above is linked to a sequence encoding a 5° UTR to provide a composite promoter.
[00408] In some embodiments, the composite promoter comprises or consists of SEQ ID NO: 148, or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[00409] This composite promoter / 5’ UTR construct comprises SP0240 operably linked to the 5° UTR from the CMV-IE gene and the GCCACC (SEQ ID NO: 153) Kozak sequence. This (composite) promoter is referred to as SP0240-UTR. Again, in this promoter, a short sequence downstream of the TSS in the CRE0006 promoter element have been replaced with sequences from the 5° UTR from the CMV-IE. Thus, this promoter actually comprises a minor variant of SP0240 with a modification to CRE0006 whereby some sequence has been removed. SP0240-UTR is preferred in some embodiments. As discussed above, the 5° UTR suitably comprises a nucleic acid motif that functions as the protein translation initiation site, e.g. sequences that define a Kozak sequence. In the sequence above, the 5° UTR comprises the sequence motif GCCACC (SEQ ID NO: 153) at its 3° end, but this sequence motif can be omitted or alternative sequences can be used.
[00410] In some embodiments, the SP0246 promoter, or variants thereof, as discussed above is linked to a sequence encoding a 5° UTR to provide a composite promoter.
[00411] In some embodiments, the composite promoter comprises or consists of SEQ ID NO: 149 (SP0246-UTR), or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[00412] This composite promoter / 5’ UTR construct comprises SP0246 operably linked to the 5° UTR from the CMV-IE gene and the GCCACC (SEQ ID NO: 153) Kozak sequence. This (composite) promoter is referred to as SP0246-UTR. Again, in this promoter, a short sequence downstream of the TSS in the CRE0052 promoter element have been replaced with sequences from the 5° UTR from the CMV-IE. Thus, this promoter actually comprises a minor variant of SP0246 with a modification to CRE0052 whereby some sequence has been removed. SP0246-UTR is preferred in some embodiments. As discussed above, the 5° UTR suitably comprises a nucleic acid motif that functions as the protein translation initiation site, e.g. sequences that define a Kozak sequence. In the sequence above, the 5° UTR comprises the sequence motif GCCACC (SEQ ID NO: 153) at its 3° end, but this sequence motif can be omitted or alternative sequences can be used.
[00413] In some embodiments, the SPO131_A1 promoter, or variants thereof, as discussed above is linked to a sequence encoding a 5° UTR to provide a composite promoter.
[00414] In some embodiments, the composite promoter comprises or consists of SEQ ID NO: 150 (SP0131 A1-UTR), or a functional variant thereof. In some embodiments, functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 999%, identical thereto.
[00415] This composite promoter / 5’ UTR construct comprises SP0131 operably linked to the 5° UTR from the CMV-IE gene and the GCCACC (SEQ ID NO: 153) Kozak sequence. This (composite) promoter is referred to as SP0131-UTR. Again, in this promoter, a short sequence downstream of the TSS in the CRE0052 promoter element have been replaced with sequences from the 5° UTR from the CMV-IE. Thus, this promoter actually comprises a minor variant of SP0131 with a modification to CRE0052 whereby some sequence has been removed. SP0131-UTR is preferred in some embodiments. As discussed above, the 5° UTR suitably comprises a nucleic acid motif that functions as the protein translation initiation site, e.g. sequences that define a Kozak sequence. In the sequence above, the 5° UTR comprises the sequence motif GCCACC (SEQ ID NO: 153) at its 3” end, but this sequence motif can be omitted or alternative sequences can be used.
[00416] In some embodiments, the liver-specific promoter is SP0412 (SEQ ID NO: 91) and comprises the following components: CRE0051 (SEQ ID NO: 97), CRE0067 (SEQ ID NO: 152), CRE0059 (SEQ ID NO: 110) and a Kozak sequence (SEQ ID NO: 153); or functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 099%, identical thereto.
[00417] In some embodiments, the liver-specific promoter is SP0422 (SEQ ID NO: 9) and comprises the following components: CRE0051 (SEQ ID NO: 97), CRE0067 (SEQ ID NO: 152), CRE0059 (SEQ ID NO: 110), CMV-IE 5°UTR (SEQ ID NO: 153) and a Kozak sequence (SEQ ID NO: 153, or functional variants may have a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. (iii) Functional variants of the synthetic liver-specific promoters
[00418] In some embodiments, a functional variant of a liver-specific promoter can be viewed as a promoter element which, when substituted in place of a reference promoter element in a promoter, substantially retains its activity. For example, a functional variant of liver-specific promoter which comprises a functional variant of a given promoter in Table 4 herein, or any promoter listed from SEQ ID NOS: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), SEQ ID NOS: 93 (SP0239), SEQ ID NO: 94 (SP0265, also referred to SP131_A1), SEQ ID NO: 95 (SP0240) or SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240-UTR), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-A1-UTR), or a functional fragment or variant thereof, or any LSP selected from SEQ ID NO: 270-341 or 342- 430, or a functional fragment or variant thereof where a functional variant or functional fragments preferably retains at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 70% or at least 80% of its activity, more preferably at least 90% of its activity, more preferably at least 95% of the activity of the unchanged promoter, and yet more preferably 100% of the activity (as compared to the unchanged promoter sequence comprising the unmodified promoter element). Suitable assays for assessing liver-specific promoter activity are disclosed herein, e.g. in Examples 12 and 13.
[00419] In some embodiments, a functional variant or a functional fragment of a liver-specific promoter disclosed in Table 4 herein, or any promoter listed from SEQ ID NOS: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), SEQ ID NOS: 93 (SP0239), SEQ ID NO: 94 (SP0263, also referred to SP131_A1), SEQ ID NO: 95 (SP0240) or SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240- UTR), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-A1-UTR), or any LSP selected from SEQ ID NO: 270-341 or 342-430, has at least about 75% sequence identity to, or at least about 80% sequence identity to, at least about 90% sequence identity to, at least about 95% sequence identity to, at least about 98% sequence identity to the original unmodified sequence, and also at least 35% of the promoter activity, or at least about 45% of the promoter activity, or at least about 50% of the promoter activity, or at least about 60% of the promoter activity, or at least about 75% of the promoter activity, or at least about 80% of the promoter activity, or at least about 85% of the promoter activity, or at least about 90% of the promoter activity, or at least about 95% of the promoter activity of the corresponding unmodified promoter sequence.
[00420] For example, a functional variant or a functional fragment of SEQ ID NO: 258 (SP0412) or SEQ ID NO: 92 (SP0422) has at least about 75% sequence identity to SEQ ID NO: 91(SP0412) or SEQ ID NO: 92(SP0422), or at least about 80% sequence identity to SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), at least about 90% sequence identity to SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), at least about 95% sequence identity to SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), at least about 98% sequence identity to SEQ ID NO: SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92(SP0422), or the original unmodified sequence, and also at least 35% of the promoter activity, or at least about 45% of the promoter activity, or at least about 50% of the promoter activity, or at least about 60% of the promoter activity, or at least about 75% of the promoter activity, or at least about 80% of the promoter activity, or at least about 85% of the promoter activity, or at least about 90% of the promoter activity, or at least about 95% of the promoter activity of the corresponding unmodified promoter sequence of SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), respectively.
[00421] In some embodiments, a functional variant of a liver-specific promoter disclosed herein retains a significant level of sequence identity to the unmodified promoter sequence. Suitably functional variants comprise a sequence that is at least 60% identical to the unmodified promoter sequence, more preferably at least 70%, 80%, 90%, 95% or 99% identical to the unmodified liver- specific promoter sequence.
[00422] In some embodiments, a functional fragment of a liver-specific promoter disclosed herein retains a significant level of sequence identity to the unmodified promoter sequence. Suitable functional fragments comprise a sequence that is at least 60% identical to the unmodified promoter sequence, more preferably at least 70%, 80%, 90%, 95% or 99% identical to the unmodified liver- specific promoter sequence.
[00423] In some embodiments, a functional variant of a promoter element can be viewed as a promoter element which, when substituted in place of a reference promoter element in a promoter, substantially retains its activity. For example, a liver-specific promoter which comprises a functional variant of a given promoter element preferably retains at least 80% of its activity, more preferably at least 90% of its activity, more preferably at least 95% of its activity, and yet more preferably 100% of its activity (compared to the reference promoter comprising the unmodified promoter element). Suitable assays for assessing liver-specific promoter activity are disclosed herein, e.g. in Examples 12 and 13.
[00424] It should be noted that the sequences of a liver-specific promoter as disclosed herein in Table 4, or any LSP selected from SEQ ID NO: 270-341 or 342-430 can be altered without causing a substantial loss of activity. Thus, functional variants of a liver-specific promoter are discussed below can be prepared by modifying the sequence of a liver-specific promoter disclosed in Table 4 herein, or any or any LSP selected from SEQ ID NO: 270-341 or 342-430, provided that modifications which are significantly detrimental to activity of the liver-specific promoter are avoided. In view of the information provided in the present disclosure, modification of a liver-specific promoter disclosed herein in Table 4, or any LSP selected from SEQ ID NO: 270-341 or 342-430 to provide functional variants is straightforward. Moreover, the present disclosure provides methodologies for simply assessing the functionality of any given liver-specific promoter variant. Functional variants for each liver-specific promoter are discussed below.
[00425] In some embodiments of the invention the synthetic liver-specific promoter comprises a sequence from the group consisting of: any promoter listed from SEQ ID NOS: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), SEQ ID NOS: 93 (SP0239), SEQ ID NO: 94 (SP0265, also referred to SP131_A1), SEQ ID NO: 95 (SP0240) or SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240- UTR), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-A1-UTR), or any LSP selected from SEQ ID NO: 270-341 or 342-430, or a functional variant of any thereof. Suitably the functional variant of any of said liver-specific promoter comprises a sequence that is at least 70% identical to the reference synthetic liver-specific promoter, more preferably at least 80%, 90%, 95% or 99% identical to the reference synthetic liver-specific promoter.
[00426] In some embodiments, the functional variant thereof may suitably comprise a sequence that is at least 60%, 70%, 80%, 90%, 95% or 99% identical to any one of the sequences listed in Table 4. Additionally or altematively, a functional variant of any one of the sequences listed in Table 4, suitably comprises a sequence which hybridizes under stringent conditions to the reference sequence. Functional variants of any one of the sequences listed in Table 4 include variants in which one or more of the sequence provided therein has been replaced with a functional variant thereof as defined above, and / or where the order of the sequences provided therein has been altered.
[00427] In some embodiments, a functional variant of any one of the liver-specific promoter sequences listed in Table 4 can be viewed as a liver-specific promoter, when at least one or more nucleotides are substituted and it substantially retains its activity. For example, a liver-specific promoter which comprises a functional variant of any one of the liver-specific promoter sequences listed in Table 4 preferably retains 80% of its activity, more preferably 90% of its activity, more preferably 95% of its activity, and yet more preferably 100% of its activity (compared to the reference promoter sequence). For example, if a LSP comprises a nucleic acid sequence comprising SP0412 (SEQ ID NO: 91) as an example, a portion of nucleotides in SP0412 (e.g., SEQ ID NO:91) in can be replaced with a functional variant of thereof, and the liver specific SP0412 promoter substantially retains its activity. Retention of activity can be assessed by comparing expression of a suitable reporter under the control of the reference promoter with an otherwise identical promoter comprising the substituted nucleic acids under equivalent conditions. Suitable assays for assessing liver-specific promoter activity are disclosed herein, e.g. in examples 12 and 13.
[00428] In some embodiments of the compositions and methods disclosed herein, a synthetic liver- specific promoter disclosed herein in Table 4, or any promoter listed from SEQ ID NOS: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), SEQ ID NOS: 93 (SP0239), SEQ ID NO: 94 (SP0263, also referred to SP131_Al), SEQ ID NO: 95 (SP0240) or SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240-UTR), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-A1-UTR), or any LSP selected from SEQ ID NO: 270-341 or 342-430is a functional variant thereof that has length of 700, 600, 500, 450, 400, 350, 300, 250 or 200 or 150 or fewer nucleotides.
[00429] In some embodiments of the compositions and methods disclosed herein, the synthetic liver- specific promoter disclosed herein in Table 4, any promoter listed from SEQ ID NOS: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), SEQ ID NOS: 93 (SP0239), SEQ ID NO: 94 (SP0265, also referred to SP131_A1), SEQ ID NO: 95 (SP0240) or SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240-UTR), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-A1-UTRY), or any LSP selected from SEQ ID NO: 270-341 or 342-430, comprises a synthetic liver-specific cis- regulatory element (CRE) or cis-regulatory module (CRM) operably linked to a minimal promoter. Examples of suitable minimal promoters for use in the present invention include, but are not limited to, the CMV-minimal promoter, MinTk minimal promoter, and the LVR_CRE0052_G6PC minimal promoter (SEQ ID NO: 126). In particular embodiments, the minimal promoter is the CMV-IE promoter comprising the sequence of SEQ ID NO: 145, a sequence that is at least 60%, 70%, 80%, 90%, 95% or 99% identical to SEQ ID NO: 145 Exemplary promoters comprising a CMV-IE for use in the methods and compositions disclosed herein can be selected from, but are not limited to, SEQ ID NO: 92 (SP0422), SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240-UTR), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-A1-UTR).
[00430] In some embodiments of the compositions and methods disclosed herein, a synthetic liver- specific promoter disclosed herein in Table 4, e.g., any promoter listed from SEQ ID NOS: 86 (CRM 0412), SEQ ID NO: 91 (SP0412) or SEQ ID NO: 92 (SP0422), SEQ ID NOS: 93 (SP0239), SEQ ID NO: 94 (SP0263, also referred to SP131_A1l), SEQ ID NO: 95 (SP0240) or SEQ ID NO: 96 (SP0246), or SEQ ID NO: 146 (SP0265-UTR), SEQ ID NO: 147 (SP0239-UTR), SEQ ID NO: 148 (SP0240-UTR), SEQ ID NO: 149 (SP0246-UTR) or SEQ ID NO: 150 (SP0131-A1-UTR), or any LSP selected from SEQ ID NO: 270-341 or 342-430 is able to increase expression of gene in the liver of a subject or in a liver cell by at least 20%, at least 40%, at least 60%, at least 80%, at least 100%, at least 200%, at least 300%, at least 500%, at least 1000% or more relative to the LP-1 promoter (SEQ ID NO: 432).
[00431] In some embodiments of the compositions and methods disclosed herein, the synthetic liver- specific promoter disclosed herein in Table 4,or any LSP promoter selected from SEQ ID NOS: 86, 91-96, 146-150, or any LSP selected from SEQ ID NO: 270-341 or 342-430, a synthetic liver-specific promoter is able to promote liver-specific transgene expression and has an activity in liver cells which is at least 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 125%, 150%, 175%, 200%, 250%, 300%, 350% or 400% of the activity of the TTR promoter (SEQ ID NO: 431).
[00432] In some embodiments of the compositions and methods disclosed herein, the synthetic liver- specific promoter disclosed herein in Table 4, or any LSP promoter selected from SEQ ID NOS: 86, 91-96, 146-150, or any LSP selected from SEQ ID NO: 270-341 or 342-430, a synthetic liver-specific promoter is able to promote liver-specific transgene expression and has an activity in liver cells which is at least 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 125%, 150%, 175%, 200%, 250%, 300%, 350% or 400% of the activity of the TBG promoter (SEQ ID NO: 435). G. Intron sequence
[00433] In some embodiments, the rAAV genotype comprises an intron sequence located 3” of the promoter sequence and 5” of the secretory signal peptide. Intron sequences serve to increase one or more of: mRNA stability, mRNA transport out of nucleus and / or expression and / or regulation of the expressed GAA fusion polypeptide (e.g., SS-GAA fusion polypeptide or SS-IGF2-GAA polypeptide). In alternative embodiments, a rAAV genotype does not comprise an intron sequence.
[00434] In some embodiments, the intron sequence is a MVM intron sequence, for example, but not limited to and intron sequence of SEQ ID NO: 13 or nucleic acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% nucleotide sequence identity thereto.
[00435] In some embodiments, the intron sequence is a HBB2 intron sequence, for example, but not limited to and intron sequence of SEQ ID NO: 14 or nucleic acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% nucleotide sequence identity thereto.
[00436] In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises a heterologous nucleic acid sequence that further comprises an intron sequence located 5° of the sequence encoding the secretory signal peptide, and 3° of the promoter. In some embodiments, the intron sequence comprises a MVM sequence or a HBB2 sequence, wherein the MVM sequence comprises the nucleic acid sequence of SEQ ID NO: 13, or a nucleic acid sequence at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to SEQ ID NO: 13, and the HBB2 sequence comprises the nucleic acid sequence of SEQ ID NO: 14, or a nucleic acid sequence at least about 75%, or 80%, or 85%, or 90%, or 95%, or 98%, or 99% sequence identity to SEQ ID NO: 14.
[00437] In some embodiments, the rAAV genotype comprises an intron sequence selected in the group consisting of a human beta globin b2 (or HBB2) intron, a FIX intron, a chicken beta-globin intron, and a SV40 intron. In some embodiments, the intron is optionally a modified intron such as a modified HBB2 intron (see, ¢.g., SEQ ID NO: 17 in of W02018046774A1): a modified FIX intron (see., e.g., SEQ ID NO: 19 in W02018046774A1), or a modified chicken beta-globin intron (e.g., see SEQ ID NO: 21 in W02018046774A1), or modified HBB2 or FIX introns disclosed in ‘WO02015 / 162302, which are incorporated herein in their entirety by reference. H. Poly-A
[00438] In some embodiments, an rAAV vector genome includes at least one poly-A tail that is located 3° and downstream from the heterologous nucleic acid gene encoding the in one embodiment, a GAA fusion polypeptide (e.g., SS-GAA fusion polypeptide or SS-IGF2-GAA polypeptide). In some embodiments, the polyA signal is 3” of a stability sequence or CS sequence as defined herein. Any polyA sequence can be used, including but not limited to hGH poly A, synpA polyA and the like. In some embodiments, the polyA is a synthetic polyA sequence. In some embodiments, the rAAV vector genome comprises two poly-A tails, e.g., a hGH poly A sequence and another polyA sequence, where a spacer nucleic acid sequence is located between the two poly A sequences. In some embodiments, the rAAV genome comprises 3” of the nucleic acid encoding the GAA fusion polypeptide (e.g., SS- GAA fusion polypeptide or SS-IGF2-GAA polypeptide), or altematively, 3° of the CS sequence the following elements; a first polyA sequence, a spacer nucleic acid sequence (of between 100-400bp, or about 250bp), a second poly A sequence, a spacer nucleic acid sequence, and the 3° ITR. In some embodiments, the first and second poly A sequence is a hGH poly A sequence, and in some embodiments, the first and second poly A sequences are a synthetic poly A sequence. In some embodiments, the first poly A sequence is a hGH poly A sequence and the second poly A sequence is a synthetic sequence, or vice versa — that is, in alternative embodiments, the first poly A sequence is a synthetic poly A sequence and the second poly A sequence is a hGH polyA sequence. An exemplary poly A sequence is, for example, SEQ ID NO: 15 (hGH poly A sequence), or a poly A nucleic acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% nucleotide sequence identity to SEQ ID NO: 15. In some embodiments, the hGHpoly sequence encompassed for use is described in Anderson et al. J. Biol. Chem 264(14); 8222-8229, 1989 (See, ¢.g. p. 8223, 2nd column, first paragraph) which is incorporated herein in its entirety by reference.
[00439] In some embodiments, a poly-A tail can be engineered to stabilize the RNA transcript that is transcribed from an rAAV vector genome, including a transcript for a heterologous gene, which in one embodiment is a GAA, and in alternative embodiments, the poly-A tail can be engineered to include elements that are destabilizing.
[00440] In some embodiments of the methods and compositions disclosed herein, a recombinant AAV vector comprises at least one polyA sequence located 3° of the nucleic acid encoding the GAA gene and 5” of the 3’ ITR sequence. In some embodiments, the poly A is a full length poly A (fl- polyA) sequence. In some embodiments, the polyA is a truncated polyA sequence, see. e.g., FIG. 5G.
[00441] In an embodiment, a poly-A tail can be engineered to become a destabilizing element by altering the length of the poly-A tail. In an embodiment, the poly-A tail can be lengthened or shortened. In some embodiments, the 3’ untranslated region comprises GAA 3° UTR (SEQ ID NO: 85) ora 3’ UTR (SEQ ID NO: 77).
[00442] In another embodiment, a destabilizing element is a microRNA (miRNA) that has the ability to silence (repress translation and promote degradation) the RNA transcripts the miRNA binds to that encode a heterologous gene. In an embodiment, addition or deletion of seed regions within the poly-A tail can increase or decrease expression of a protein, such as the GAA protein or modified GAA polypeptide.
[00443] In another embodiment, seed regions can also be engineered into the 3° untranslated regions located between the heterologous gene and the poly-A tail. In a further embodiment, the destabilizing agent can be an siRNA. The coding region of the siRNA can be included in an rAAV vector genome and is generally located downstream, 3” of the poly-A tail.
[00444] In all aspects of the methods and compositions as disclosed herein, the rAAV genome may also comprise a Stuffer DNA nucleic sequence. An exemplary stuffer DNA sequence is SEQ ID NO: 71, or a nucleic acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% nucleotide sequence identity thereto. As shown in FIGS 7-8 and FIGS 9A-9E, the stuffer sequence is located 3 of the poly A tail, for example, and is located 5° of the ‘3 ITR sequence. In some embodiments, the stuffer DNA sequence comprises a synthetic polyadenylation signal in the reverse orientation.
[00445] In some embodiments, a stuffer nucleic acid sequence (also referred to as a “spacer” nucleic acid fragment, see FIGS 7-8) can be located between the poly A sequence and the 3° ITR (ie.,a stuffer nucleic acid sequence is located 3” of the polyA sequence and 5° of the 3’ ITR) (see, e.g., FIG. 7-8). Such a stuffer nucleic acid sequence can be about 30bp, 50pb, 75bp, 100bp, 150bp, 200bp, 250bp, 300bp or longer than 300bp. In some embodiments of the methods and compositions as disclosed herein, a stuffer nucleic acid fragment is between 20-50bp, 50-100bp, 100-200bp, 200- 300bp, 300-500bp, or any integer between 20-500bp. Exemplary stuffer (or spacer) nucleic acid sequence comprise SEQ ID NO: 16, SEQ ID NO: 71 or SEQ ID NO: 78, or a nucleic acid sequence at least about 70, 71, 72, 73,74, 75,76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95. 96, 97, 98, 99%, identical to SEQ ID NO: 16 or SEQ ID NO: 71 or SEQ ID NO: 78. I. AAV ITRs
[00446] The rAAV genome as disclosed here comprises AAV ITRs that have desirable characteristics and can be designed to modulate the activities of, and cellular responses to vectors that incorporate the ITRs. In another embodiment, the AAV ITRs are synthetic AAV ITRs that has desirable characteristics and can be designed to manipulate the activities of and cellular responses to vectors comprising one or two synthetic ITRs, including, as set forth in U.S. Patent No. 9,447433, which is incorporated herein by reference.
[00447] In another embodiment, an ITR exhibits modified transcription activity relative to a naturally occurring ITR, e.g., ITR2 from AAV2. It is known that the ITR2 sequence inherently has promoter activity. It also inherently has termination activity, similar to a poly(A) sequence. The minimal functional ITR of the present invention exhibits transcription activity as shown in the examples, although at a diminished level relative to ITR2. Thus, in some embodiments, the ITR is functional for transcription. In other embodiments, the ITR is defective for transcription. In certain embodiments, the ITR can act as a transcription insulator, e.g., preventing transcription of a transgenic cassette present in the vector when the vector is integrated into a host chromosome.
[00448] One aspect of the invention relates to an rAAV vector genome comprising at least one synthetic AAV ITR, wherein the nucleotide sequence of one or more transcription factor binding sites in the ITR is deleted and / or substituted, relative to the sequence of a naturally occurring AAV ITR such as ITR2. In some embodiments, it is the minimal functional ITR in which one or more transcription factor binding sites are deleted and / or substituted. In some embodiments at least 1 transcription factor binding site is deleted and / or substituted, e.g., at least 5 or more or 10 or more transcription factor binding sites, e.g., at least 1,2, 3,4, 5, 6,7, 8,9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 transcription factor binding sites.
[00449] Another embodiment, a rAAV vector, including an rAAV vector genome as described herein comprises a polynucleotide comprising at least one synthetic AAV ITR, wherein one or more CpG islands (a cytosine base followed immediately by a guanine base (a CpG) in which the cytosines in such arrangement tend to be methylated) that typically occur at, or near the transcription start site in an ITR are deleted and / or substituted. In an embodiment, ...
Claims
1. A recombinant adeno-associated virus (AAV) vector comprising in its genome:a. 5’ and 3’ AAV inverted terminal repeats (ITR) sequences, andb. located between the 5’ and 3’ ITRs, a heterologous nucleic acid comprising an alphaglucosidase (GAA) coding sequence that encodes a GAA polypeptide or a contiguous portion thereof comprising at least amino acids 70-790 of SEQ ID NO: 10 or SEQ ID NO: 170, operatively linked to a liver-specific promoter selected from any one of: CRM_SP0412 (SEQ ID NO: 86) or SP0412 (SEQ ID NO: 91) or SP0422 (SEQ ID NO: 92); or a functional fragment of SP0412 that comprises a contiguous portion of the unmodified promoter that has at least 75% of the untruncated promoter, truncated from the 3’ end, and retains liver-specific promoter activity at least of CRM_SP0412; or a functional fragment of SP0422 that comprises a contiguous portion of the unmodified promoter that has at least 35% of the untruncated promoter, truncated from the 3’ end, and retains liver-specific promoter activity at least of CRM_SP0412.
2. The recombinant AAV vector of claim 1, wherein the heterologous nucleic acid further encodes a secretory signal polypeptide that is fused to the GAA polypeptide, or further encodes a targeting peptide that is fused to the GAA polypeptide, or further encodes a secretory signal polypeptide and a targeting peptide that are fused to the GAA polypeptide.
3. The recombinant AAV vector of claim 2, wherein the AAV genome comprises, in the 5’ to 3’ direction:a. the 5’ ITR,b. the liver-specific promoter,c. an intron sequence,d. a nucleic acid encoding a secretory signal polypeptide,e. a GAA coding sequence encoding a GAA polypeptide or a contiguous portion thereof comprising amino acids 70-790 of SEQ ID NO: 10 or SEQ ID NO: 170,f. a poly A sequence, andg. the 3’ ITR.
4. The recombinant AAV vector of claim 2 or 3, wherein the secretory signal polypeptide is: an AAT signal peptide, a fibronectin signal peptide (FN1), a GAA leader sequence, a IL-2 wt leader sequence, a modified IL-2 leader sequence, an IL2(1-3) leader sequence, an IgG leader sequence, an AAT leader sequence, or an active fragment thereof having secretory signal activity.
5. The recombinant AAV vector of claim 2, wherein the targeting peptide is an IGF2 targeting peptide which binds human cation-independent mannose-6-phosphate receptor (CI-MPR) or the IGF2 receptor.2020388634 04 Aug 20266. The recombinant AAV vector of claim 5, wherein the IGF2 targeting peptide comprises SEQ ID NO: 5 or comprises at least one amino modification in SEQ ID NO: 5 that binds to the IGF2 receptor.
7. The recombinant AAV vector of claim 6, wherein the at least one amino modification in SEQ ID NO: 5 is a V43M amino acid modification (SEQ ID NO: 8 or SEQ ID NO: 9) or A2-7 (SEQ ID NO: 6) or A1-7 (SEQ ID NO: 7).
8. The recombinant AAV vector of any of claims 1-7, wherein the GAA coding sequence is a human GAA gene or a human codon optimized GAA gene (coGAA) or a modified GAA nucleic acid sequence.
9. The recombinant AAV vector of any of claims 1-8, wherein the GAA coding sequence is modified from SEQ ID NO: 11 or SEQ ID NO: 55 for any one or more of: (i) codon optimized for enhanced expression in vivo, (ii) reduce CpG islands, (iii) modification of STOP sequences, (iv) reduction of alternative reading frames, and (v) to reduce the innate immune response.
10. The recombinant AAV vector of any of claims 5-9, wherein the encoded fusion polypeptide further comprises a spacer comprising at least 1 amino acid located amino-terminal to the GAA polypeptide, and C-terminal to the IGF2 targeting peptide.
11. The recombinant AAV vector of claim 10, wherein the spacer is located between the IGF2 targeting peptide and the GAA polypeptide.
12. The recombinant AAV vector of any of claims 1-11, further comprising at least one polyA sequence located 3’ of the GAA coding sequence and 5’ of the 3’ ITR.
13. The recombinant AAV vector of any of claims 1-12, wherein the heterologous nucleic acid further comprises a nucleic acid encoding a collagen stability (CS) sequence, or a 3’ UTR sequence, or a CS and 3’ UTR sequence located 3’ of the GAA coding sequence and 5’ of the 3’ ITR sequence.
14. The recombinant AAV vector of claim 12 or 13, wherein the heterologous nucleic acid further comprises a nucleic acid encoding a collagen stability (CS) sequence, or a 3’ UTR sequence, or a CS and 3’ UTR sequence, located between the GAA coding sequence and the poly A sequence.
15. The recombinant AAV vector of any of claims 1-14, wherein the heterologous nucleic acid further comprises an intron sequence located 5’ of the GAA coding sequence , and 3’ of the promoter.
16. The recombinant AAV vector of claim 15, wherein the intron sequence comprises a MVM sequence or a HBB2 sequence or a SV40 sequence.
17. The recombinant AAV vector of any of claims 1-16, wherein the 5’ and / or 3’ ITR comprises an insertion, deletion or substitution.2020388634 04 Aug 202618. The recombinant AAV vector of claim 17, wherein one or more CpG islands in the ITR are removed.
Citation Information
Patent Citations
Acid-alpha glucosidase variants and uses thereof
EP3293260A1
Improved constructs for expressing lysosomal polypeptides
WO2004064750A2
Transcription regulatory elements and uses thereof
WO2019153009A1