Neurotensin mutants and tagged proteins containing the neurotensin or sortilin propeptide

JP2024526286A5Pending Publication Date: 2025-07-09AMICUS THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024500052
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-01
Filing Date
2022-07-01
Publication Date
2025-07-09

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The disclosure provides amino acid sequences of neurotensin variants and recombinant proteins comprising neurotensin, sortilin propeptide, or variants thereof. The disclosure also provides methods of producing the recombinant proteins and methods for determining cellular uptake of the recombinant proteins. Gene therapy compositions, pharmaceutical compositions, methods of treatment, and uses of the gene therapy compositions and recombinant proteins are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to treating genetic disorders. [Background technology]

[0002] Genetic disorders such as lysosomal storage diseases arise from inherited or de novo mutations occurring within gene coding regions of the genome. Mutations within this region can result in defective proteins, which are critical for cellular metabolic activity. Many mutant proteins are unstable in the ER (Ishii et al., Biochem. Biophys. Res. Comm. 1996;220:812-815), so that the proteins are delayed in the normal trafficking pathway (ER→Golgi apparatus→endosomes→lysosomes) and are prematurely degraded. The resulting protein deficiency can cause the accumulation of toxic compounds and subsequent disruption of normal cellular functions.

[0003] Protein replacement therapy is one of the approved treatments for treating genetic disorders.This therapy typically involves intravenous injection of the purified form of corresponding wild-type protein.However, one of the main problems associated with protein replacement therapy is the acquisition and maintenance of therapeutically effective amounts of protein in vivo due to low intracellular bioavailability and rapid degradation and / or clearance of injected protein.The current approach to overcome this problem is to carry out many costly high-dose injections.

[0004] Gene therapy is another possible approach to treat genetic disorders. Gene therapy involves supplementing or complementing defective genes with nucleic acid sequences that code for functional proteins. Gene therapy can use recombinant vectors to deliver nucleic acid sequences that code for functional proteins or genetically modified human cells that express functional proteins. However, delivery of transgenes or transgene products to relevant target cells and organelles, either by direct transduction or cross-correction, remains a challenge in gene therapy. Summary of the Invention [Problem to be solved by the invention]

[0005] Thus, due to the challenges associated with reaching the relevant cells and organelles for both gene therapy and protein replacement therapy, new therapeutic approaches that provide more precise targeting for treating genetic disorders are needed. [Means for solving the problem]

[0006] Neurotensin is a neuropeptide encoded by the neurotensin gene. In some embodiments, neurotensin functions as one or more of a hormone, a growth factor, an apoptosis signaling peptide, and a lysosomal trafficking promoting peptide. One of the natural receptors of neurotensin is sortilin (also called neurotensin receptor 3, NTSR3), which is mainly present in the Golgi membrane, nuclear membrane, lysosomal membrane, cell plasma membrane, endoplasmic reticulum membrane, and endosomal membrane, and co-localizes with CI-MPR / IGF2R in lysosomes. Sortilin also has high levels of expression in certain tissues, such as kidney, heart, brain, and reproductive tissue. Thus, various aspects of the present invention provide precise targeting of therapeutic proteins to cells or organelles with high expression of sortilin using neurotensin or its variants.

[0007] Various aspects of the invention relate to variants of wild-type neurotensin peptide amino acid sequences. In one or more embodiments, the neurotensin peptide amino acid sequence comprises at least 70% sequence identity to SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13, such that the amino acid sequence is not SEQ ID NO:1. In one or more embodiments, the neurotensin peptide amino acid sequence comprises at least 90% sequence identity to SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13, such that the amino acid sequence is not SEQ ID NO:1. In one or more embodiments, the neurotensin peptide amino acid sequence comprises SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13.

[0008] One or more embodiments of the invention relate to a polynucleotide comprising a nucleotide sequence encoding a neurotensin peptide having at least 70% sequence identity to SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13, such that the resulting amino acid sequence is not SEQ ID NO:1. In one or more embodiments, the nucleotide sequence encodes a neurotensin peptide having at least 90% sequence identity to SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13, such that the resulting amino acid sequence is not SEQ ID NO:1. In one or more embodiments, the nucleotide sequence encodes a neurotensin peptide having SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13.

[0009] Various aspects of the invention relate to neurotensin polynucleotides that comprise at least 40% nucleotide sequence identity to SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26, such that the resulting polypeptide is not SEQ ID NO: 1. In one or more embodiments, the polynucleotide has at least 70% nucleotide sequence identity to SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26, such that the polynucleotide is not SEQ ID NO:14. In one or more embodiments, the polynucleotide has at least 97% nucleotide sequence identity to SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26, such that the polynucleotide is not SEQ ID NO:14. One or more embodiments include a neurotensin polynucleotide having SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26.

[0010] Various aspects of the invention refer to tagged proteins comprising a therapeutic protein and one or more tags selected from a neurotensin peptide and / or a sortilin propeptide. In one or more embodiments, the neurotensin peptide of the tagged protein comprises a polypeptide that is at least 70% identical to SEQ ID NO:1. In one or more embodiments, the neurotensin peptide of the tagged protein comprises a polypeptide that is at least 90% identical to SEQ ID NO:1. In one or more embodiments, the neurotensin peptide of the tagged protein comprises SEQ ID NO:1. In one or more embodiments, the neurotensin peptide of the tagged protein comprises a polypeptide that is at least 70% identical to SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13. In one or more embodiments, the neurotensin peptide of the tagged protein comprises a polypeptide that is at least 90% identical to SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13. In one or more embodiments, the neurotensin peptide of the tagged protein comprises SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13. In one or more embodiments, the sortilin propeptide of the tagged protein comprises a polypeptide that is at least 70% identical to SEQ ID NO:48. In one or more embodiments, the sortilin propeptide of the tagged protein comprises a polypeptide that is at least 90% identical to SEQ ID NO:48. In one or more embodiments, the sortilin propeptide of the tagged protein comprises SEQ ID NO:48. In other embodiments, the sortilin propeptide is a variant that is at least 40%, 70% or 97% identical to SEQ ID NO:48 but is not SEQ ID NO:48.

[0011] In one or more embodiments, the therapeutic protein comprises a soluble protein. In one or more embodiments, the therapeutic protein comprises a lysosomal protein. In one or more embodiments, the therapeutic protein of the tagged protein is palmitoyl protein thioesterase 1 (PPT1) (CLN1), tripeptidyl peptidase 1 (TPP1) (CLN2), cathepsin D (CTSD) (CLN10), progranulin (PGRN) (CLN11) and cathepsin F (CTSF) (CLN13), alpha-galactosidase A, β-galactosidase, β-hexosaminidase, galactosylceramidase, arylsulfatase, β-glucocerebrosidase, glucocerebrosidase, lysosomal acid lipase, lysosomal enzyme acid sphingomyelinase, formylglycine generating enzyme, iduronidase, acetyl-CoA:alpha-glucosaminide N-acetyltransferase ... Enzymatically active enzymes include, for example, glycosaminoglycan alpha-L-iduronohydrolase, heparan N-sulfatase, N-acetyl-α-D-glucosaminidase (NAGLU), iduronate 2-sulfatase, galactosamine-6-sulfatase, N-acetylgalactosamine-6-sulfatase, glycosaminoglycan N-acetylgalactosamine 4-sulfatase, β-glucuronidase, hyaluronidase, alpha-N-acetylneuraminidase (sialidase), ganglioside sialidase, phosphotransferase, alpha-glucosidase, alpha-D-mannosidase, beta-D-mannosidase, aspartylglucosaminidase, alpha-L-fucosidase, or enzymatically active fragments thereof. In one or more embodiments, the therapeutic protein of the tagged protein comprises a Batten-related protein selected from PPT1, TPP1, CTSD, PGRN, or CTSF, In one or more embodiments, the therapeutic protein of the tagged protein comprises lysosomal alpha-glucosidase (GAA).

[0012] In one or more embodiments, the tagged protein comprises one or more of an affinity tag, a linker peptide, and a secretory signal peptide.

[0013] Various aspects of the invention relate to polynucleotides comprising nucleotide sequences encoding tagged proteins. In one or more embodiments, the polynucleotide sequence comprises a therapeutic protein polynucleotide sequence and one or more tag sequences comprising a neurotensin polynucleotide sequence and / or a sortilin propeptide polynucleotide sequence. In one or more embodiments, the neurotensin polynucleotide sequence comprises at least 40% sequence identity to SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26. In one or more embodiments, the neurotensin polynucleotide sequence comprises at least 70% sequence identity to SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26. In one or more embodiments, the neurotensin polynucleotide sequence comprises at least 97% sequence identity to SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, or SEQ ID NO: 26. In one or more embodiments, the neurotensin polynucleotide sequence comprises SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, or SEQ ID NO: 26. In one or more embodiments, the sortilin propeptide polynucleotide sequence comprises at least 40%, 70%, 97% or 100% sequence identity to SEQ ID NO: 49.

[0014] Various aspects of the invention relate to methods of producing tagged proteins, including expressing the tagged protein and purifying the tagged protein. In one or more embodiments, the tagged protein is expressed in Expi293F cells, PC12 cells, COS7 cells, HAP1 cells, Chinese hamster ovary (CHO) cells, HeLa cells, human embryonic kidney (HEK) cells, mouse primary myoblasts, NIH 3T3 cells, Escherichia coli cells, baculovirus expression systems (e.g., Sf9 cells, Sf21 cells), yeast cells (e.g., Saccharomyces cerevisiae), or mutants thereof.

[0015] Various aspects of the invention relate to methods for determining cellular uptake of tagged proteins. In one or more embodiments, the methods include culturing cells, incubating the cells with tagged proteins, and identifying the tagged proteins that are transported into the cells. In one or more embodiments, the methods may include determining the functionality of the tagged proteins that are transported into the cells. In one or more embodiments, the methods may include quantifying the tagged proteins that are transported into the cells. In one or more embodiments, the methods are performed in Expi293F cells, PC12 cells, COS7 cells, HAP1 cells, Chinese Hamster Ovary (CHO) cells, HeLa cells, human embryonic kidney (HEK) cells, mouse primary myoblasts, NIH 3T3 cells, Escherichia coli cells, Sf9 cells, Sf21 cells, Saccharomyces cerevisiae, or mutants thereof.

[0016] Various aspects of the invention relate to methods for assessing sortilin binding efficiency of a tagged protein. In one or more embodiments, the method comprises incubating immobilized sortilin with the tagged protein and quantifying the percentage of sortilin-bound tagged protein. In one or more embodiments, the sortilin comprises a domain or full-length protein. In one or more embodiments, the sortilin domain comprises amino acid residues 78-755 of the full-length sortilin amino acid sequence. In one or more embodiments, the method comprises measuring functionality of the sortilin-bound tagged protein.

[0017] Various aspects of the invention relate to pharmaceutical formulations comprising neurotensin peptides, tagged proteins, and / or polynucleotides encoding neurotensin peptides or tagged proteins. The pharmaceutical formulations may also include a pharma- ceutical acceptable carrier and / or one or more pharma-ceutical acceptable excipients.

[0018] Various aspects of the invention relate to methods of treating a disease or disorder comprising administering a pharmaceutical formulation to a patient in need of such treatment. In one or more embodiments, the disease or disorder is Fabry disease and the therapeutic protein comprises alpha-galactosidase A. In one or more embodiments, the disease or disorder is Pompe disease and the therapeutic protein comprises acid alpha-glucosidase. In one or more embodiments, the disease or disorder is Batten disease and the therapeutic protein comprises PPT1 (CLN1), TPP1 (CLN2), CTSD (CLN10), PGRN (CLN11), or CTSF (CLN13). In one or more embodiments, the pharmaceutical formulation is administered intrathecally, intravenously, intracisternally, intraventricularly, intraocularly, intravitreally, retinatally, subretinatally, intramuscularly, subcutaneously, intracerebrally, surgically, or intraparenchymally.

[0019] Various aspects of the invention relate to gene therapy compositions comprising a gene therapy delivery system and a polynucleotide encoding a tagged protein. In one or more embodiments, the gene therapy delivery system comprises one or more of a vector, a liposome, a lipid-nucleic acid nanoparticle, an exosome, and a gene editing system. In one or more embodiments, the gene therapy editing system comprises one or more of clustered regularly interspaced short palindromic repeats (CRISPR) associated protein 9 (CRISPR-Cas-9), a transcription activator-like effector nuclease (TALEN), or a ZNF (zinc finger protein). In one or more embodiments, the gene therapy delivery system comprises a viral vector. In one or more embodiments, the viral vector comprises one or more of an adenoviral vector, an adeno-associated viral vector, a lentiviral vector, a retroviral vector, a poxviral vector, or a herpes simplex viral vector. In one or more embodiments, the viral vector comprises a viral polynucleotide operably linked to a polynucleotide encoding a tagged protein. In one or more embodiments, the viral vector comprises at least one inverted terminal repeat (ITR). In one or more embodiments, the viral vector comprises one or more of an SV40 intron, a polyadenylation signal (e.g., bovine growth hormone polyadenylation signal (bGHpolyA)), or a stabilizing element (e.g., marmot hepatitis virus (WHP) posttranscriptional regulatory element (WPRE)). In one or more embodiments, the gene therapy delivery system comprises a promoter (e.g., CBA). In one or more embodiments, the gene therapy delivery system comprises a polynucleotide encoding a secretory signal peptide (e.g., human BiP or a variant thereof). In one or more embodiments, the gene therapy composition comprises a pharma- ceutically acceptable carrier.

[0020] Various aspects of the invention relate to methods of treating a disease or disorder comprising administering a gene therapy composition to a patient in need of such treatment. In one or more embodiments, the disease or disorder is Fabry disease and the therapeutic protein comprises alpha-galactosidase A. In one or more embodiments, the disease or disorder is Pompe disease and the therapeutic protein comprises alpha-glucosidase. In one or more embodiments, the disease or disorder is Batten disease and the therapeutic protein comprises PPT1 (CLN1), TPP1 (CLN2), CTSD (CLN10), PGRN (CLN11), or CTSF (CLN13). In one or more embodiments, the gene therapy composition is administered intrathecally, intravenously, intracisternally, intraventricularly, intraocularly, intravitreally, retinatally, subretinatally, intramuscularly, subcutaneously, intracerebrally, surgically, or intraparenchymally. [Brief description of the drawings]

[0021] [Figure 1A-1B] Binding efficiency of GAA-tagged proteins to Sortilin. Two variants of purified GAA-tagged protein, GAA-tagged protein with a C-terminal NT tag (GAA-NT) (Figure 1A) and GAA-tagged protein with an N-terminal NT tag (NT-GAA) (Figure 1B), were separately co-immunoprecipitated with soluble Sortilin domain using a commercial co-immunoprecipitation kit (Thermo Fisher, 26149). The images show that both GAA-NT and NT-GAA have high affinity for the soluble Sortilin domain. [Diagram 2]α-D-glucopyranoside activity of sortilin-bound GAA-tagged proteins. Extracellular soluble sortilin domain with a C-terminal TwinStrep tag to be immobilized on a 96-well microplate. The wells were then incubated with various concentrations of NT-GAA (GAA-tagged protein with an N-terminal NT tag), GAA-NT (GAA-tagged protein with a C-terminal NT tag) or GAA (GAA without the NT tag). The α-D-glucopyranoside activity of sortilin-bound tagged proteins was measured using fluorescent 4-methylumbelliferyl-α-D-glucopyranoside substrate. Figure 2 shows the α-D-glucopyranoside activity of sortilin-bound NT-GAA and GAA-NT. Wells incubated with GAA without the NT tag had no α-D-glucopyranoside activity, indicating that the protein did not bind to sortilin. [Figure 3A-3B] α-D-glucopyranoside activity of GAA-tagged proteins transported into cells. The uptake of GAA-tagged proteins was tested in HAP1 KO cells by determining the α-D-glucopyranoside activity of GAA-tagged proteins transported into cells. Uptake assays for NT-GAA (GAA-tagged protein with an N-terminal NT tag), GAA-NT (GAA-tagged protein with a C-terminal NT tag) or GAA (GAA without an NT tag) were performed in the presence and absence of 10 mM mannose-6-phosphate (M6P). After the uptake assay, cells were lysed and α-D-glucopyranoside activity was measured using the fluorescent substrate 4-methylumbelliferyl-α-D-glucopyranoside at pH 4.8. Figure 3A shows the α-D-glucopyranoside activity for NT-GAA, GAA-NT and GAA. Similarly, Western blot analysis in FIG. 3B shows NT-GAA, GAA-NT, and GAA transported into cells in the presence of M6P. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0022] Before describing certain exemplary embodiments of the invention, it is to be understood that the invention is not limited to the details of construction or process steps set forth in the following description, as the invention is capable of other embodiments and of being practiced or carried out in various ways.

[0023] Various embodiments of the present invention relate to methods for delivering therapeutic molecules to cells, hi one or more embodiments, the methods are used to deliver therapeutic proteins for genetic disorders.

[0024] definition The term "Batten-related protein" refers to proteins involved in Batten disease, such as neuronal ceroid lipofuscinosis proteins, including CLN1, CLN2, CLN3, CLN4, CLN5, CLN6, CLN7, CLN8, CLN10, CLN11, CLN12, CLN13, and CLN14. Non-limiting exemplary neuronal ceroid lipofuscinosis proteins and corresponding Batten diseases are shown in Table 1.

[0025] [Table 1]

[0026] In one or more embodiments, the neuronal ceroid lipofuscinosis protein comprises a soluble protein. Examples of such soluble proteins include neuronal ceroid lipofuscinosis protein 1 (palmitoyl protein thioesterase, such as palmitoyl protein thioesterase 1 (PPT1)), neuronal ceroid lipofuscinosis protein 2 (tripeptidyl peptidase 1 (TPP1)), neuronal ceroid lipofuscinosis protein 10 (cathepsin D (CTSD)), neuronal ceroid lipofuscinosis protein 11 (progranulin (PGRN)), and neuronal ceroid lipofuscinosis protein 13 (cathepsin F (CTSF)).

[0027] The term "tagged protein" refers to a protein that comprises a tag and a therapeutic protein. In one or more embodiments, the tag comprises a neurotensin peptide and / or a sortilin propeptide. In one or more embodiments, the tagged protein may comprise at least one additional protein, peptide, or polypeptide linked thereto. The additional protein, peptide, or polypeptide may comprise one or more of a secretion signal peptide, a linker peptide, and an affinity tag.

[0028] The term "gene therapy delivery system" refers to any system that can be used to deliver an exogenous gene of interest to a target cell such that the gene of interest is expressed or overexpressed in the target cell. In one or more embodiments, the target cell is an in vivo patient cell. In one or more embodiments, the target cell is an ex vivo cell, and the cell is then administered to a patient.

[0029] The term "host cell" means any cell of any organism that is selected, modified, transformed, propagated, or used or manipulated in any way for the cellular production of a substance, such as the cellular expression of a gene, DNA or RNA sequence, protein or enzyme.

[0030] The term "minigene" refers to the combination of a transgene, a promoter / enhancer and 5' and 3' AAV ITRs, referred to herein for ease of reference as a "minigene." Given the teachings of this invention, the design of such a minigene can be generated using conventional techniques.

[0031] The wild-type neurotensin peptide (NT) consists of 13 amino acids and is in GeneBank Accession No. P30990. In one or more embodiments, the wild-type neurotensin peptide may have one or more deletions, additions, and / or substitutions. Non-limiting exemplary neurotensin peptides and their encoding polynucleotide sequences are shown in Table 2 below.

[0032] [Table 2]

[0033] Non-limiting exemplary mutant sequences include SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13. In one or more embodiments, the NT may have at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, or at least 90% sequence identity to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13. One skilled in the art can easily generate a polynucleotide sequence encoding the amino acid sequence. The polynucleotide sequence may also be codon-optimized for expression in a target cell using commercially available products. In one or more embodiments, a polynucleotide sequence encoding NT may have at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 97% sequence identity to SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26, such that the resulting polypeptide is not SEQ ID NO: 1. In one or more embodiments, the nucleotide sequence of a neurotensin peptide may include SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26.

[0034] The term "operably linked" refers to the functional relationship of a polynucleotide / gene with nucleotide regulatory and effector sequences, such as promoters, enhancers, transcription and translation termination sites, and other signal sequences. For example, operably linking a nucleic acid to a promoter refers to a physical and functional relationship between a polynucleotide and a promoter such that transcription of DNA is initiated from the promoter by an RNA polymerase that specifically recognizes and binds to the promoter. The promoter directs the transcription of RNA from the polynucleotide.

[0035] The term "patient" refers to any mammalian subject, particularly humans, for whom diagnosis, treatment or therapy is desired. A mammalian subject for purposes of treatment refers to any animal classified as a mammal, including humans, farm and laboratory animals, zoo animals, sport animals or pet animals, such as dogs, horses, cats, cows, sheep, goats, pigs, mice, rats, rabbits, guinea pigs, monkeys, etc. A mammalian subject may be a fetus, a newborn, a child, a juvenile, or an adult with a disability. In one or more embodiments, the mammalian subject is a human.

[0036] The term "pharmaceutical acceptable" refers to molecular entities and compositions that are physiologically acceptable and generally do not produce adverse reactions when administered to humans. Preferably, as used herein, the term "pharmaceutical acceptable" means approved by a federal or state regulatory agency or listed in the United States Pharmacopoeia or other generally recognized pharmacopoeias for use in animals and more specifically in humans. The term "carrier" refers to a diluent, adjuvant, excipient, or vehicle with which a compound is administered. Such pharmaceutical carriers can be sterile liquids, such as water and oils. Water or aqueous saline solutions and aqueous dextrose and glycerol solutions are preferably used as carriers, particularly for injectable solutions. Suitable pharmaceutical carriers are described in "Remington's Pharmaceutical Sciences" by E. W. Martin, 18th Edition, or other editions.

[0037] The term "promoter" refers to a DNA sequence to which the enzyme RNA polymerase binds and initiates the transcription of a DNA sequence into RNA. Promoters can be synthetic or natural, capable of conferring, activating or enhancing expression of a nucleic acid sequence in a cell. Promoters can contain one or more specific transcriptional regulatory sequences to further enhance expression and / or alter its spatial and / or temporal expression. Promoters can also contain distal enhancer or repressor elements, which can be located as far away as several thousand base pairs from the start site of transcription. Promoters can be from sources including viral, bacterial, fungal, plant, insect and animal. Promoters can regulate the expression of gene components constitutively or differentially, with respect to the cell, tissue or organ in which expression occurs, with respect to the developmental stage in which expression occurs, or in response to external stimuli such as physiological stress, pathogens, metal ions or inducers. Representative examples of promoters include, but are not limited to, CBA, P546, mouse mammary tumor virus (MMTV), human immunodeficiency virus (HIV) long terminal repeat (LTR) promoter, MoMuLV promoter, avian leukemia virus promoter, Epstein-Barr virus immediate early promoter, Rous sarcoma virus promoter, actin promoter, myosin promoter, elongation factor-1a promoter, hemoglobin promoter, and creatine kinase promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, SV40 early promoter or SV40 late promoter and CMV IE promoter.

[0038] The term "protein replacement therapy" refers to the introduction of exogenous purified proteins into individuals with deficiencies in such proteins. The administered proteins may be obtained from natural sources or by recombinant expression. The term also refers to the introduction of purified proteins in individuals who otherwise require or would benefit from the administration of purified proteins. In at least one embodiment, such individuals suffer from a protein deficiency. The term "protein deficiency" as used herein refers to one or more of a functional deficiency and a dietary deficiency. In one or more embodiments, the protein deficiency is a functional deficiency. The introduced protein may be a purified recombinant protein produced in vitro in a host cell, or a protein purified from an isolated tissue or body fluid, such as, for example, placenta or animal milk, or from a plant.

[0039] The term "serotype" is a characteristic of an AAV that has a capsid that is serologically distinct from other AAV serotypes. Serotype differences are determined based on the lack of cross-reactivity between antibodies and the AAV as compared to other AAVs.

[0040] Cross-reactivity is generally measured in a neutralizing antibody assay. For this assay, polyclonal sera are generated against a specific AAV in rabbits or other suitable animal models using adeno-associated virus. In this assay, the sera generated against a specific AAV are then tested for their ability to neutralize either the same (homologous) or heterologous AAV. The dilution that reaches 50% neutralization is considered the neutralizing antibody titer. If the quotient of the heterologous titer divided by the homologous titer for two AAVs is reciprocally lower than 16, these two vectors are considered to be of the same serotype. Conversely, if the ratio of the heterologous titer to the homologous titer is reciprocally 16 or higher, the two AAVs are considered to be of different serotypes.

[0041] The term "sortilin" refers to a membrane protein that is a single-pass transmembrane protein with a Vps10p extracellular domain. Sortilin is also called neurotensin receptor 3 (NTSR3). Residues 78-755 form the extracellular domain. Sortilin has high levels of expression in certain tissues, among which kidney, heart, brain, and reproductive tissues have the highest expression. At the cellular level, sortilin is primarily present in the Golgi membrane, nuclear membrane, lysosomal membrane, cell plasma membrane, endoplasmic reticulum membrane, and endosomal membrane. Sortilin is known to co-localize with CI-MPR / IGF2R in lysosomes and mediates mannose-6-phosphate-independent uptake of lysosomal proteins, such as alpha-galactosidase (GLA), acid sphingomyelinase (ASM), etc. In some embodiments, neurotensin is engineered to interact with sortilin. In some embodiments, based on the known crystal structure of sortilin in complex with NT (PBD:3f6k, 4po7) in silico modeling, 12 NT mutants SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13 were predicted to improve sortilin binding affinity to NT.

[0042] Non-limiting exemplary therapeutic proteins include lysosomal proteins such as alpha-glucosidase (GAA). Non-limiting exemplary therapeutic proteins include PPT1 (CLN1), TPP1 (CLN2), CTSD (CLN10), PGRN (CLN11), CTSF (CLN13), alpha-galactosidase, β-galactosidase, β-hexosaminidase, galactosylceramidase, arylsulfatase, β-glucocerebrosidase, glucocerebrosidase, lysosomal acid lipase, lysosomal enzyme acid sphingomyelinase, formylglycine generating enzyme, iduronidase, acetyl-CoA:alpha-glucosaminide N-acetyltransferase, glycosaminoglycan alpha-L-iduronohydrolase, heparan N-sulfatase, and the like. The enzymes and / or enzymes used in the present invention include, but are not limited to, N-acetylglucosamine (Aglu), N-acetylglucosamine (Glu ...

[0043] The term "treating" refers to administering an agent or performing a procedure for the purpose of obtaining a therapeutic effect, including inhibiting, attenuating, alleviating, preventing, or altering, in a statistically or clinically significant manner, at least one aspect or marker of a disorder. The term "treating" does not state or imply a cure for the underlying condition, but rather, prevents a disorder or a symptom of a disorder from occurring in a patient who may have a predisposition to, but has not yet been diagnosed as having, a disorder (e.g., including a disorder that may be associated with or caused by a primary disorder; (b) inhibiting a disorder, i.e., arresting its occurrence; (c) palliating a disorder, i.e., causing regression of a disorder; and (d) improving at least one symptom of a disorder. Treating may refer to any indication of success in treating or ameliorating or preventing a disorder, including, but not limited to, any indication of success in treating or ameliorating or preventing a disorder. The term "treatment" includes any objective or subjective parameter, such as, but not limited to, abatement; remission; relief; attenuation of symptoms or making the disorder more tolerable to the patient; slowing the rate of degeneration or decline; or making the end point of degeneration less debilitating. The treatment or amelioration of symptoms is based on one or more objective or subjective parameters; including the results of a physician's examination. Thus, the term "treating" includes administration of a compound or agent of the invention to prevent or slow the onset of, alleviate or arrest or inhibit the onset of symptoms or conditions associated with the disorder.

[0044] The term "vector" refers to a gene therapy delivery vehicle or carrier that delivers a therapeutic gene to a cell. The vector is any vector suitable for use in gene therapy, for example any vector suitable for therapeutic delivery of a nucleic acid polymer (encoding a polypeptide or a variant thereof) to a target cell of a patient. In some embodiments, the gene therapy vector delivers a nucleic acid encoding a tagged protein to a cell where the tagged protein is expressed and secreted from the cell. The vector can be of any type, for example it can be a plasmid vector or a minicircle DNA. Generally, the vector is a viral vector. The viral vector can be derived, for example, from an adeno-associated virus (AAV), a retrovirus, a lentivirus, a herpes simplex virus, or an adenovirus. Viral vectors may also contain additional natural or synthetic elements to increase expression and / or stabilize the vector, such as promoters (e.g., the hybrid CBA promoter (CBh) and the human synapsin 1 promoter (hSyn1)), polyadenylation signals (e.g., the bovine growth hormone polyadenylation signal (bGH polyA)), stabilizing elements (e.g., the marmot hepatitis virus (WHP) post-transcriptional regulatory element (WPRE)), and / or SV40 introns. In one or more embodiments, the vector can contain polynucleotide sequences flanked by regions that promote homologous recombination at a desired site in the genome, thus resulting in expression of a desired protein (see Koller and Smithies, 1989, Proc. Natl. Acad. Sci. USA, 86:8932-8935; Zijlstra et al., 1989, Nature 342:435-438; U.S. Pat. No. 6,244,113 to Zarling et al.; and U.S. Pati et al., U.S. Pat. No. 6,200,812 to Pati et al.).

[0045] AAV vectors AAV vectors for transgene delivery In various embodiments, the gene therapy compositions and / or methods utilize AAV vectors. Alternatively, other viral vectors or gene therapy delivery systems as described herein may be used.

[0046] AAV vectors from one of many families can be utilized to deliver transgenes in vivo for transduction, with many examples provided in detail in U.S. Patent No. 7,198,951, which is incorporated herein in its entirety. Using the genomes of these AAV vectors and the manufacturing processes described herein and known in the art, recombinant AAV (rAAV) can be generated as vectors for delivery of one or more of the transgenes provided herein.

[0047] I. phylogenetic group Phylogenetic groups are groups of AAVs that are phylogenetically related to each other as determined by alignment of AAV vp1 amino acid sequences using a neighbor-joining algorithm with a bootstrap value of at least 75% (out of at least 1000 replicates) and a Poisson-corrected distance measure not exceeding 0.05.

[0048] Neighbor-joining algorithms have been widely described in the literature. See, for example, M. Nei and S. Kumar, Molecular Evolution and Phylogenetics (Oxford University Press, New York (2000)). Computer programs are available that can be used to implement this algorithm. For example, the MEGA v2.1 program implements a modified Nei-Gojobori method. Using these techniques and computer programs and the sequence of the AAV vp1 capsid protein, one skilled in the art can easily determine whether a selected AAV is contained in one of the phylogenetic groups identified herein, in another phylogenetic group, or outside these phylogenetic groups.

[0049] While the phylogenetic groups defined herein are primarily based on naturally occurring AAV vp1 capsids, the phylogenetic groups are not limited to naturally occurring AAVs. The phylogenetic groups can encompass non-naturally occurring AAVs, including but not limited to recombinant, modified or engineered, chimeric, hybrid, synthetic, artificial AAVs, etc., that are phylogenetically related as determined based on alignment of the AAV vp1 amino acid sequence using a neighbor-joining algorithm in at least 75% (out of at least 1000 replicates) and a Poisson-corrected distance measure not exceeding 0.05.

[0050] The phylogenetic groups described herein include phylogenetic group A (represented by AAV1 and AAV6), phylogenetic group B (represented by AAV2) and phylogenetic group C (represented by AAV2-AAV3 hybrids), phylogenetic group D (represented by AAV7), phylogenetic group E (represented by AAV8) and phylogenetic group F (represented by human AAV9).

[0051] Phylogenetic group B (AAV2) and phylogenetic group C (AAV2-AAV3 hybrid) are the most common found in humans (22 isolates from 12 individuals for AAV2 and 17 isolates from 8 individuals for phylogenetic group C).

[0052] Phylogenetic group A (represented by AAV1 and AAV6) AAV vectors can include those of phylogenetic group A, which includes AAV1 and AAV6. See, e.g., WO 00 / 28061, May 18, 2000; Rutledge et al, J Virol, 72(1):309-319 (January 1998). In addition, this phylogenetic group includes the AAVs described in U.S. Patent No. 7,198,951.

[0053] Clade B (represented by the AAV2 clade) In other embodiments, the AAV vector comprises those of phylogenetic group B, which includes AAV2 and those described in U.S. Patent No. 7,198,951. In one or more embodiments, one or more members of this phylogenetic group have a capsid that has at least 85% amino acid identity, at least 90% identity, at least 95% identity, or at least 97% identity over the entire length of vp1, vp2, or vp3 of the AAV2 capsid.

[0054] Clade C (represented by the AAV2-AAV3 hybrid clade) In another embodiment, the AAV vector is characterized in that it contains an AAV that is a hybrid of the previously published AAV2 and AAV3, including those of phylogenetic group C, and that is described in U.S. Patent No. 7,198,951. In one embodiment, one or more members of this phylogenetic group have a capsid that has at least 85% identity, at least 90% identity, at least 95% identity, or at least 97% identity amino acid identity over the entire length of vp1, vp2, or vp3 of the hu.4 and / or hu.2 capsid.

[0055] Phylogenetic group D (AAV7 phylogenetic group) In another embodiment, the AAV vector comprises one of phylogenetic group D, which includes AAV7, as described in U.S. Patent No. 7,198,951. In one or more embodiments, one or more members of this phylogenetic group have a capsid with at least 85% amino acid identity, at least 90% identity, at least 95% identity, or at least 97% identity over the entire length of vp1, vp2, or vp3 of the AAV7 capsid.

[0056] Clade E (represented by the AAV8 clade) In one or more aspects, the AAV vector comprises those of phylogenetic group E, which includes AAV8 and those described in U.S. Patent Publication No. 7,198,951. In one or more embodiments, one or more members of this phylogenetic group have a capsid that has at least 85% identity, at least 90% identity, at least 95% identity, or at least 97% identity amino acid identity over the entire length of vp1, vp2, or vp3 of the AAV8 capsid. In another embodiment, the invention provides novel AAV vectors of phylogenetic group E, as described in U.S. Patent Application Publication No. 2003 / 0138772A1 (July 24, 2003).

[0057] Clade F (represented by the AAV9 clade) In another embodiment of the invention, the AAV vector comprises phylogenetic group F, which includes AAV9 and those described in U.S. Patent No. 7,198,951. In one or more embodiments, one or more members of this phylogenetic group have a capsid with at least 85% identity, at least 90% identity, at least 95% identity, or at least 97% identity amino acid identity over the entire length of vp1, vp2 or vp3 of the AAV9 capsid.

[0058] The AAV families are useful for a variety of purposes, including providing a ready-made collection of related AAVs for the generation of viral vectors and for the generation of targeting molecules. These families can also be used as tools for a variety of purposes that will be readily apparent to those of skill in the art.

[0059] Transgene A transgene is a nucleic acid sequence that is heterologous to the vector sequences that flank the transgene, which encodes a polypeptide, protein, or other product of interest. The nucleic acid coding sequence may be operably linked to regulatory components to enable or regulate transgene transcription, translation, and / or expression in a host cell.

[0060] Transgenes can be used to correct or improve gene defects, which may include defects in which normal genes are expressed below normal levels or in which a fully functional gene product is not expressed. Transgenes can be used to suppress gene expression in certain cells where gene expression is toxic or harmful (e.g., using miRs to suppress gene expression in dorsal root ganglia). Alternatively, transgenes can provide cells with products that are not natively expressed in the cell type or in the host. A preferred type of transgene sequence encodes a therapeutic protein or polypeptide to be expressed in the host cell. The invention further includes the use of multiple transgenes. In certain situations, different transgenes can be used to code for each subunit of a protein or to code for different peptides or proteins. This is desirable when the size of the DNA encoding the protein subunits is large. In order for the cells to produce a multisubunit protein, the cells are transfected with a recombinant virus containing each of the different subunits. Alternatively, the different subunits of a protein can be encoded by the same transgene. In this case, a single transgene contains DNA encoding each of the subunits, and the DNA for each subunit is separated by an internal ribozyme entry site (IRES). This is desirable when the size of the DNA encoding each of the subunits, e.g., the total size of the DNA encoding the subunits, is small, and the IRES is less than 5 kilobases. As an alternative to an IRES, the DNA can be separated by a sequence encoding a 2A peptide, which self-cleaves in a post-translational event. See, e.g., ML Donnelly, et al, J. Gen. Virol., 78(Pt 1):13-21 (January 1997); Furler, S., et al, Gene Ther., 8(11):864-873 (June 2001); Klump H., et al., Gene Ther., 8(10):811-817 (May 2001). This 2A peptide is significantly smaller than an IRES, making it highly suitable for use when space is a limiting factor.More frequently, when the transgene is large, it consists of multiple subunits, or two transgenes are delivered simultaneously. In some embodiments, rAAV carrying the desired transgenes or subunits can be administered simultaneously to allow them to concatenate in vivo to form a single vector genome. In such embodiments, the first AAV can carry an expression cassette expressing a single transgene, and the second AAV can carry an expression cassette expressing a different transgene for co-expression in the host cell. However, the selected transgene can code for any biologically active product or other product, such as a product desired for testing.

[0061] Therapeutic transgenes In some embodiments, gene products such as enzymes may be useful in enzyme replacement therapy, which is useful in a variety of pathologies resulting from deficiencies in the activity of enzymes. In some embodiments, a suitable gene comprises a polynucleotide sequence encoding β-glucuronidase.

[0062] In some embodiments, gene products include non-naturally occurring polypeptides, for example, chimeric or hybrid polypeptides having a non-naturally occurring amino acid sequence containing insertions, deletions, or amino acid substitutions.

[0063] NT and mutants The amino acid sequence of wild type NT is shown in SEQ ID NO:1. Various embodiments of the invention provide novel NT mutants. The amino acid sequence of the NT mutants is shown in SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13. In one or more embodiments, the NT amino acid sequence comprises at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% sequence identity to SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13. In one or more embodiments, the NT amino acid sequence may comprise one of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13.

[0064] In one or more embodiments, the polynucleotide comprises a nucleotide sequence such that the resulting polypeptide has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% sequence identity to SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13. In one or more embodiments, the polynucleotide comprises a nucleotide sequence that encodes SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13.

[0065] Similarly, in one or more embodiments, a neurotensin polynucleotide encoding NT comprises at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 97% nucleotide sequence identity to SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26, such that the resulting polypeptide is not SEQ ID NO: 1. In one or more embodiments, a polynucleotide encoding NT comprises SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26.

[0066] Sortilin has two possible interaction sites for NT, one with the amino terminus of NT and one with the carboxy terminus. Thus, a tagged protein including NT and a therapeutic protein can be cloned into a suitable plasmid. In some embodiments, when expressed, the therapeutic protein is at either the amino or carboxy terminus of NT. In one or more embodiments, NT can include an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% similar to SEQ ID NO:1. In one or more embodiments, NT can include amino acids with at least 5, at least 4, at least 3, at least 2, or at least 1 mutation, addition, or deletion to SEQ ID NO:1. In one or more embodiments, the NT amino acid sequence comprises at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence similarity to SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13. In one or more embodiments, the polynucleotide comprises a nucleotide sequence such that the resulting polypeptide has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence similarity to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, or SEQ ID NO:13.In one or more embodiments, the polynucleotide sequence encoding NT comprises at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or 100% sequence similarity to SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26.

[0067] Sortilin Propeptide Sortilin is synthesized using a sortilin propeptide. The sortilin propeptide becomes cleaved during maturation of sortilin. The sortilin propeptide has high binding affinity to mature sortilin. Thus, in one or more embodiments, a tagged protein comprising a sortilin propeptide and a therapeutic protein can be cloned into an appropriate plasmid. In some embodiments, when expressed, the therapeutic protein is at the amino and / or carboxy terminus of the sortilin propeptide.

[0068] In one or more embodiments, the sortilin propeptide comprises an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% similar to SEQ ID NO: 48. In one or more embodiments, the NT can comprise an amino acid sequence that has at least 5, at least 4, at least 3, at least 2, or at least 1 mutation, addition, or deletion to SEQ ID NO: 48. In one or more embodiments, the polynucleotide comprises a nucleotide sequence such that the resulting polypeptide has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence similarity to SEQ ID NO: 48. In one or more embodiments, the polynucleotide sequence encoding the sortilin propeptide comprises at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or 100% sequence similarity to SEQ ID NO:49.

[0069] In some embodiments, the binding affinity of the sortilin propeptide to sortilin is greater than or equal to the binding affinity of the NT to sortilin.

[0070] Tagged proteins One aspect of the invention relates to a tagged protein comprising a therapeutic protein and a tag. In some embodiments, the tag may be at the amino and / or carboxy terminus. In one or more embodiments, the tag comprises a neurotensin peptide and / or a sortilin propeptide. In some embodiments, the tag comprises a sortilin propeptide on the amino terminus and a neurotensin on the carboxy terminus. In some embodiments, the tag comprises a neurotensin on the amino terminus and a sortilin propeptide on the carboxy terminus. In one or more embodiments, the tagged protein comprises a neurotensin and / or a sortilin propeptide. In some embodiments, the tagged protein comprises a neurotensin on the amino terminus and a sortilin propeptide on the carboxy terminus. In some embodiments, the tagged protein comprises a sortilin propeptide on the amino terminus and a neurotensin on the carboxy terminus.

[0071] In one or more embodiments, the therapeutic protein comprises a lysosomal protein, hi one or more embodiments, the therapeutic protein comprises alpha-galactosidase, beta-galactosidase, beta-hexosaminidase, galactosylceramidase, arylsulfatase, beta-glucocerebrosidase, glucocerebrosidase, lysosomal acid lipase, lysosomal enzyme acid sphingomyelinase, formylglycine generating enzyme, iduronidase, acetyl-CoA:alpha-glucosaminide N-acetyltransferase, glycosaminoglycan alpha-L-iduronohydrolase, heparan N-sulfatase, N-acetyl-α-D-glucosaminidase (NAGLU), iduro In one or more embodiments, the Batten-related proteins include PPT1 (CLN1), TPP1 (CLN2), CTSD (CLN10), PGRN (CLN11), or CTSF (CLN13).

[0072] In one or more embodiments, the therapeutic protein comprises a non-naturally occurring polypeptide, for example, a chimeric or hybrid polypeptide having a non-naturally occurring amino acid sequence that contains insertions, deletions, or amino acid substitutions.

[0073] In some embodiments, the tagged protein comprises one or more additional amino acid sequences encoding one or more functional domains or sequences, including one or more of the other functional domains and sequences provided herein. Exemplary functional domains or sequences include, but are not limited to, affinity tags, linker peptides, protease cleavage sites, secretion signal peptides, cell targeting domains, reporters, enzymes, or combinations thereof. In one or more embodiments, the tagged protein comprises an affinity tag, where the affinity tag is linked to the therapeutic protein such that the end of the NT remains open. In some embodiments, the affinity tag can be one or more of mCherry, TwinStrep, polyhistidine, HA, FLAG, GST, and GFP.

[0074] In some embodiments, proteolytic cleavage sites can be engineered into tagged proteins to facilitate release of proteins of interest from vesicle-targeted proteins and / or other peptide functional domains that contain affinity tags in conjunction with synthesis or purification of the tagged proteins. Exemplary protease cleavage sites include, but are not limited to, cleavage sites sensitive to thrombin, furin, factor Xa, metalloproteases, enterokinase, and cathepsin.

[0075] The cell targeting domain can comprise an amino acid sequence that confers cell type-specific or cell differentiation-specific targeting.

[0076] The functional domains in the tagged proteins of the present invention may be separated from each other by a spacer or linker to promote independent folding of each peptide moiety relative to each other and ensure that the individual peptide moieties in the tagged protein do not interfere with each other. The spacer may comprise any amino acid or mixtures thereof. Preferably, the spacer selected improves the flexibility of the protein and promotes the adoption of extended conformations. Preferred peptide spacers include the amino acids proline, lysine, glycine, alanine, serine and combinations thereof. In one embodiment, the linker is a glycine-rich linker.

[0077] The linker or spacer sequence comprises 4-15 amino acids in length. Suitable amino acids for incorporation in the linker are alanine, arginine, serine, or glycine. In some embodiments, the linker sequence comprises a lysosomal cleavage sequence. In some embodiments, the lysosomal cleavage sequence comprises an amino acid according to SEQ ID NO: 27. In one or more embodiments, the tagged protein comprises a linker peptide between the affinity tag and the therapeutic protein. In one or more embodiments, the tagged protein comprises a linker peptide between the tag (e.g., a neurotensin peptide or a sortilin propeptide) and the therapeutic protein. In some embodiments, the linker peptide comprises repeated glycine residues, repeated glycine-serine residues, or a combination thereof. In some embodiments, the linker peptide comprises 5-20 amino acids, 5-15 amino acids, 5-10 amino acids, or 8-12 amino acids. In some embodiments, the linker peptide comprises about 5, 6, 7, 8, 9, 10, 11, 12, or 13 amino acids. Linkers suitable for the gene therapy and enzyme replacement therapy constructs herein include, but are not limited to, those shown in Table 3 below.

[0078] [Table 3]

[0079] In some embodiments, a linker for a gene therapy and / or enzyme replacement therapy construct comprises an amino acid sequence according to SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34 or SEQ ID NO:35.

[0080] In one or more embodiments, the tagged protein may include a secretory signal peptide. The signal peptide may be derived from a lysosomal protein, including but not limited to PPT1, TPP1, GAA, GLA, NAGLU, SGSH, or any other lysosomal enzyme that contains a signal peptide. Exemplary signal peptides are listed in Table 4 below.

[0081] [Table 4]

[0082] In some embodiments, the signal peptide sequence comprises an amino acid sequence according to SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, or SEQ ID NO:47.

[0083] Those skilled in the art can easily generate polynucleotide sequences that code amino acid sequences.Polynucleotide sequences can also be codon-optimized for expression in target cells using commercially available products.The present invention includes polynucleotide sequences that code tagged proteins.Polynucleotide sequences that code tagged proteins can be operably linked to expression control sequences, such as promoters that direct the expression of genes, such as secretory signal peptide genes.

[0084] Expression of tagged proteins In one or more embodiments, the method of producing a tagged protein comprises transfecting a host cell, allowing the host cell to transiently express the tagged protein, and purifying the tagged protein. In one or more embodiments, the tagged protein is expressed in the host cell. In one or more embodiments, the host cell comprises any one of Expi293F cells, PC12 cells, HAP1 cells, Chinese hamster ovary (CHO) cells, HeLa cells, human embryonic kidney (HEK) cells, mouse primary myoblasts, NIH 3T3 cells, Escherichia coli cells, baculovirus expression systems (e.g., Sf9 cells, Sf21 cells), yeast (e.g., Saccharomyces cerevisiae), or mutants thereof. In one or more embodiments, cell-free synthesis is used to express the tagged protein.

[0085] In one or more embodiments, gene transfer is performed into the host cell using a suitable vector, including viruses (such as adenoviruses, parvoviruses (e.g., adeno-associated viruses (AAV)), vaccinia viruses, herpes viruses, baculoviruses, and retroviruses or lentiviruses), bacteriophages, cosmids, plasmids, bacterial artificial chromosomes (BACs), fungal vectors, naked DNA, DNA-lipid complexes, naked RNA, RNA-lipid complexes, or other recombinant vehicles used in the art that have been described for expression in a variety of eukaryotic and prokaryotic hosts and can be used for gene therapy as well as for simple protein expression.

[0086] Following expression and secretion, the recombinant protein can be recovered and purified from the surrounding cell culture medium using standard techniques, or alternatively, the recombinant protein can be isolated and purified directly from the cells rather than from the medium.

[0087] Determination of binding efficiency of tagged proteins to sortilin In one or more embodiments, binding of the tagged protein to Sortilin may be confirmed by pull-down assay. In one or more embodiments, the pull-down assay is performed using a commercially available co-immunoprecipitation kit. In one or more embodiments, the pull-down assay is performed using a Sortilin domain or a full-length Sortilin protein. In one or more embodiments, the Sortilin domain comprises amino acid residues 78-755 of the full-length Sortilin amino acid sequence. In one or more embodiments, the method of the pull-down assay comprises immobilizing Sortilin, incubating the tagged protein with immobilized Sortilin, washing to remove unbound protein, and quantifying the percentage of Sortilin-bound tagged protein.

[0088] Measuring the activity of sortilin-binding tagged proteins In one or more embodiments, the functionality of the sortilin-binding tagged protein is measured. In one or more embodiments, the activity of the sortilin-binding tagged protein may be measured by a microplate binding assay, the steps for the assay including immobilizing sortilin on a plate, incubating the tagged protein with immobilized sortilin, and measuring the activity of the sortilin-binding tagged protein.

[0089] Determination of sortilin-mediated cellular uptake of tagged proteins Sortilin-mediated cellular uptake of tagged proteins can be determined by incubating cells with the tagged proteins and identifying the tagged proteins within the cells. In one or more embodiments, the method includes quantifying the transported tagged proteins. In one or more embodiments, immunofluorescence can be used to identify the tagged proteins within the cells. In one or more embodiments, immunofluorescence can be used to quantify the tagged proteins within the cells.

[0090] Determination of activity of tagged proteins following sortilin-mediated intracellular transport The activity of the tagged protein within a cell can be determined by incubating cells with the tagged protein and measuring the activity of the tagged protein within the cell. In one or more embodiments, intracellular protein activity can be assayed by measuring a fluorescent signal.

[0091] Delivery of tagged proteins to patients Another aspect of the invention relates to a pharmaceutical formulation comprising the tagged protein and a pharma- ceutically acceptable carrier.

[0092] Another aspect of the invention relates to a method of treating a disease or disorder comprising administering the pharmaceutical formulation to a patient in need of such treatment, hi one or more embodiments, the disease or disorder comprises a lysosomal storage disease and the therapeutic protein comprises a lysosomal enzyme. In some embodiments, the lysosomal storage disease is selected from the group consisting of aspartylglucosaminuria, Batten disease, cystinosis, Fabry disease, Gaucher disease type I, Gaucher disease type II, Gaucher disease type III, Pompe disease, Tay-Sachs disease, Sandhoff disease, metachromatic leukodystrophy, mucolipidosis type I, mucolipidosis type II, mucolipidosis type III, mucolipidosis type IV, Hurler disease, Hunter disease, Sanfilippo disease type A, Sanfilippo disease type B, Sanfilippo disease type C, Sanfilippo disease type D, Morquio disease type A, Morquio disease type B, Maroteaux-Lamy disease, Sly disease, Niemann-Pick disease type A, Niemann-Pick disease type B, Niemann-Pick disease type C1, Niemann-Pick disease type C2, Schindler disease type I, and Schindler disease type II. In some embodiments, the lysosomal storage disease is selected from the group consisting of activator deficiency, GM2-gangliosidosis; GM2-gangliosidosis, AB variant; alpha-mannosidosis (type 2, moderate; type 3, neonatal, severe); beta-mannosidosis; aspartylglucosaminuria; lysosomal acid lipase deficiency; cystinosis (late-onset juvenile or adolescent nephropathic; pediatric nephropathic); Shanarin-Dorfman syndrome; triglyceride storage disease with myopathy; NLSDM; Danon disease; Fabry disease; Fabry disease type II, late-onset; Farber disease; Farber lipogranulomatosis; fucosidosis; galactosidosis. neuraminidase and beta-galactosidase deficiency (combined deficiency);Gaucher disease;Gaucher disease type II;Gaucher disease type III;Gaucher disease type IIIC;Gaucher disease, atypical, due to saposin C deficiency;GM1-gangliosidosis (late infantile / juvenile GM1-gangliosidosis; adult / chronic GM1-gangliosidosis);Globoid cell leukodystrophy, Krabbe disease (late infantile-onset; juvenile-onset; adult-onset);Krabbe disease, atypical, due to saposin A deficiency;Metachromatic leukodystrophy (juvenile; adult);Partial cerebroside sulfate deficiency;Pseudoarylsulfatase A deficiency;Metachromatic leukodystrophy due to saposin B deficiency;Mucopolysaccharidosis disorders: MPS I, Hurler syndrome; MPS I, Hurler-Schey syndrome; MPS I, Scheie syndrome; MPS II, Hunter syndrome; MPS II, Hunter syndrome; Sanfilippo syndrome type A / MPS IIIA; Sanfilippo syndrome type B / MPS IIIB; Sanfilippo syndrome type C / MPS IIIC; Sanfilippo syndrome type D / MPS IIID; Morquio syndrome type A / MPS IVA; Morquio syndrome type B / MPS IVB; MPS IX hyaluronidase deficiency; MPS VI Maroteaux-Lamy syndrome; MPS VII Sly syndrome; Mucolipidosis I, Sialidosis type II; I-cell disease, Leroy disease, Mucolipidosis II; Pseudo-Hurler polydystrophy / Mucolipidosis type III; Mucolipidosis IIIC / ML III GAMMA;mucolipidosis type IV;multiple sulfatase deficiency;Niemann-Pick disease (type B; type C1 / chronic neuronopathic; type C2; ​​type D / Nova Scotia);neuronal ceroid lipofuscinosis:CLN6 disease-atypical late infantile, late variant, early juvenile;Batten-Spielmeyer-Voigt / juvenile NCL / CLN3 disease;Finland variant late infantile CLN5;Jansky-Beerschowski disease / late infantile CLN2 / TPP1 disease;Koufus disease / adult-onset NCL / CLN4 disease (type B);northern epilepsy / variant late infantile CLN8;Santavoli-Hartier disease / infantile CLN1 / PPT disease;Pompe disease (glycogen storage disease type II);late-onset Pompe disease ;pyknodysostosis;Sandhoff disease / GM2 gangliosidosis;Sandhoff disease / GM2 gangliosidosis;Schindler disease (type III / intermediate, variable);Kanzaki disease;Salla disease;Infantile free sialic acid storage disease (ISSD);Spinal muscular atrophy with progressive myoclonic epilepsy (SMAPME);Tay-Sachs disease / GM2 gangliosidosis;Early-onset Tay-Sachs disease;Late-onset Tay-Sachs disease;Christianson syndrome;Row oculocerebrorenal syndrome;Charcot-Marie-Tooth disease type 4J, CMT4J;Eunice-Baron syndrome;Bilateral temporo-occipital polymicrogyria (BTOP);X-linked hypercalciuric nephrolithiasis, Dent disease type 1;and Dent disease type 2. In one or more embodiments, the disease or disorder comprises Fabry disease and the therapeutic protein comprises alpha-galactosidase A. In one or more embodiments, the disease or disorder comprises Pompe disease and the therapeutic protein comprises alpha-glucosidase. In some embodiments, the therapeutic protein is associated with a lysosomal storage disease, the therapeutic protein being selected from the group consisting of GM2-activator protein; α-mannosidase; MAN2B1; lysosomal β-mannosidase; glycosylasparaginase; lysosomal acid lipase; cystinosin; CTNS; PNPLA2; lysosomal associated membrane protein-2; α-galactosidase A; GLA; acid ceramidase; α-L-fucosidase; protection protein / cathepsin A; acid β-glucosidase; GBA; PSAP; β-galactosidase-1; GLB1; galactosylceramide β-galactosidase; GALC; PSAP; arylsulfatase A; ARSA; α-L -iduronidase; iduronate 2-sulfatase; heparan N-sulfatase; N-α-acetylglucosaminidase; heparan acetyl-CoA:α-glucosaminide acetyltransferase; N-acetylglucosamine 6-sulfatase; galactosamine-6-sulfate sulfatase; β-galactosidase; hyaluronidase; arylsulfatase B; β-glucuronidase; neuraminidase; NEU1; gamma subunit of N-acetylglucosamine-1-phosphotransferase; mucolipin-1; sulfatase modifying factor-1; acid sphingomyelinase; SMPD1; NPC1; and NPC2. In one or more embodiments, the disease or disorder comprises Batten disease and the therapeutic protein comprises PPT1 (CLN1), TPP1 (CLN2), CTSD (CLN10), PGRN (CLN11), or CTSF (CLN13);

[0093] In one or more embodiments, the pharmaceutical formulation is administered intrathecally, intravenously, intracisternally, intraventricularly, intraocularly, intravitreally, retinatically, subretinatally, intramuscularly, subcutaneously, intracerebrally, surgically, or intraparenchymally.

[0094] Gene Therapy Another aspect of the present invention is applied in gene therapy. In one or more embodiments, the gene therapy composition comprises a gene therapy delivery system and a polynucleotide encoding a modified protein. In one or more embodiments, the gene therapy delivery system comprises one or more of a vector, a liposome, a lipid-nucleic acid nanoparticle, an exosome, and a gene editing system. In one or more embodiments, the gene therapy delivery system comprises one or more of a clustered regularly interspaced short palindromic repeats (CRISPR) associated protein 9 (CRISPR-Cas-9), a transcription activator-like effector nuclease (TALEN), or a ZNF (zinc finger protein). In one or more embodiments, the gene therapy delivery system comprises a promoter. In one or more embodiments, the gene therapy delivery system comprises a polynucleotide encoding a secretory signal peptide.

[0095] In one or more embodiments, the gene therapy delivery system comprises a viral vector. In one or more embodiments, the viral vector comprises one or more of an adenoviral vector, an adeno-associated viral vector, a lentiviral vector, a retroviral vector, a poxviral vector, or a herpes simplex viral vector. The viral vector may also comprise additional elements to increase expression and / or stabilize the vector, such as a promoter (e.g., a hybrid CBA promoter (CBh) and a human synapsin 1 promoter (hSyn1)), a polyadenylation signal (e.g., a bovine growth hormone polyadenylation signal (bGH polyA)), a stabilizing element (e.g., a marmot hepatitis virus (WHP) posttranscriptional regulatory element (WPRE)), and / or an SV40 intron. In one or more embodiments, the viral vector comprises a viral polynucleotide operably linked to a polynucleotide encoding a tagged protein. In one or more embodiments, the viral vector comprises at least one inverted terminal repeat (ITR).

[0096] In one or more embodiments, the gene therapy composition includes a pharma- ceutically acceptable carrier, hi one or more embodiments, a pharma- ceutically acceptable carrier includes a diluent, adjuvant, excipient, or vehicle with which the compound is administered.

[0097] Various aspects of the invention relate to methods of treating a disease or disorder comprising administering a gene therapy composition to a patient in need thereof. In one or more embodiments, the gene therapy composition is administered to a patient having a lysosomal storage disease, and the polynucleotide sequence comprises a nucleotide sequence encoding a lysosomal enzyme. In one or more embodiments, the gene therapy composition is administered to a patient having Fabry disease, and the polynucleotide sequence comprises a nucleotide sequence encoding alpha-galactosidase A. In one or more embodiments, the gene therapy composition is administered to a patient having Pompe disease, and the polynucleotide sequence comprises a nucleotide sequence encoding alpha-glucosidase. In one or more embodiments, the gene therapy composition is administered to a patient having Batten disease, and the polynucleotide sequence comprises a nucleotide sequence encoding PPT1 (CLN1), TPP1 (CLN2), CTSD (CLN10), PGRN (CLN11), or CTSF (CLN13). In one or more embodiments, the gene therapy composition is administered intrathecally, intravenously, intracisternally, intraventricularly, intraocularly, intravitreally, retinatally, subretinatally, intramuscularly, subcutaneously, intracerebrally, surgically, or intraparenchymally. EXAMPLES

[0098] Example 1 Molecular cloning and expression of [A]GAA-tagged proteins Two variants of GAA-tagged proteins were cloned into Amicus proprietary plasmids such that NT (SEQ ID NO:1) was inserted at either the 5' or 3' end of either the therapeutic protein, GAA (RefSeq:NM_000152.5), or the affinity tag, mCherry (RefSeq:MN781140.1), with a flexible 9 amino acid linker peptide between the affinity tag and the protein of interest. At the 5' end of each transcript there is a secretion signal. The tagged proteins were transiently expressed in the Expi293F cell line. The tagged proteins were then purified.

[0099] [B] Determination of the binding efficiency of GAA-tagged proteins to sortilin. Pull-down assays for soluble sortilin (R&D systems, 3154-ST-050) were performed by covalently coupling the soluble sortilin domain to agarose beads by amine coupling chemistry using a commercial co-immunoprecipitation kit (Thermo Fisher, 26149). Both NT-GAA and GAA-NT were independently incubated with the bound beads for 2 hours at room temperature with agitation. After incubation, the bound proteins were eluted in a split at low pH. The fractions and the unbound protein eluate were electrophoresed on a standard SDS-Page reducing gel, and then the proteins in the gel were transferred onto a nitrocellulose membrane for anti-mCherry staining (Invitrogen, M11217). After incubation with the primary antibody, the membrane was washed and then stained with a secondary antibody (Invitrogen, A21096) and then imaged. FIG. 1 shows that both GAA-NT (FIG. 1A) and NT-GAA (FIG. 1B) have high affinity for sortilin.

[0100] [C] Measurement of enzymatic activity of sortilin-bound GAA-tagged proteins. For the microplate binding assay, the extracellular soluble domain of sortilin, residues 78-755, was expressed in the Expi293F cell line (using the Expi293 Expression System; Thermo Fisher, A14635) with a C-terminal TwinStrep affinity tag. The product was harvested after 5 days. The conditioned medium was incubated with StrepTactin XT-coated microplates (IBA, 2-4101-001) for 1 h to immobilize sortilin to the plate. After several washes, the sortilin-coated wells were incubated with various doses of NT-GAA, GAA-NT or GAA diluted in PBS for 30 min. Unbound proteins were then washed away. The α-D-glucopyranoside activity of bound GAA-tagged proteins was measured using the fluorescent substrate, 4-methylumbelliferyl-α-D-glucopyranoside, at pH 4.8. Figure 2 shows the α-D-glucopyranoside activity of sortilin-bound NT-GAA and GAA-NT. The inflection points in both binding curves indicate the bivalency of sortilin towards neurotensin. GAA without the NT tag had no α-D-glucopyranoside activity, indicating that the protein did not bind to sortilin.

[0101] Determination of sortilin-mediated uptake of [D]GAA-tagged proteins COS7 cell line was used for immunofluorescence. Cells were seeded at approximately 40% confluency. The next day, cells were incubated with 200 nM AT-GAA (Amicus enzyme replacement therapy), NT-GAA, or GAA-NT in either the presence or absence of 10 mM mannose 6-phosphate in uptake medium for 3 hours. After uptake, cells were washed with PBS, pH 7.4, and fixed with 4% formaldehyde in PBS for 10 minutes. Cells were then permeabilized with 0.1% saponin in PBS for 5 minutes. Wells were incubated with anti-sortilin (Invitrogen, MA5-31437) and anti-GAA (from Amicus) antibodies for 1 hour. After washing with PBS, wells were incubated with fluorescent secondary antibodies (Thermo, A32723 and A11012) for 1 hour and then washed before imaging. DAPI was added for DNA staining in the final wash before imaging.

[0102] Intracellular enzymatic activity of [E]NT-tagged GAA HAP1 GAA KO cell line was used for cellular uptake assay. Cells were seeded in tissue culture plates at approximately 40% confluency. The next day, cells were incubated with various concentrations of AT-GAA (Amicus enzyme replacement therapy), NT-GAA, or GAA-NT for 18 hours either in the presence or absence of 10 mM mannose 6-phosphate in uptake medium. After 18 hours, cells were washed several times with PBS, pH 7.4, and lysed with 0.25% Triton X-100 in dH2O. The α-D-glucopyranoside activity of GAA-tagged proteins transported into cells was measured by providing the fluorescent substrate, 4-methylumbelliferyl-α-D-glucopyranoside, at pH 4.8 and measuring the 4-methylumbelliferone cleavage product. Cell lysates were electrophoresed on a standard SDS-page under reducing conditions followed by Western blotting (primary Ab: Abcam, ab137068; secondary Ab: Invitrogen, SA535571). Figure 3A shows that mannose-6-phosphate did not affect α-D-glucopyranoside activity in cells incubated with NT-GAA and GAA-NT. Similarly, Figure 3B shows that mannose-6-phosphate did not affect the uptake of NT-GAA and GAA-NT. In conclusion, the uptake of GAA-tagged proteins is independent of CI-MPR and is likely mediated by sortilin. Also, GAA-tagged proteins transported into cells retain α-D-glucopyranoside activity.

Claims

**Claim 1**: A neurotensin peptide comprising an amino acid sequence having at least 70% sequence identity, at least 90% sequence identity, or 100% sequence identity to SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, or SEQ ID NO: 13, such that the amino acid sequence is not SEQ ID NO:

1. **Claim 2** A polynucleotide comprising a nucleotide sequence encoding the neurotensin peptide according to Claim 1. **Claim 3**: A neurotensin polynucleotide comprising a polynucleotide having at least 40% nucleotide sequence identity, at least 70% nucleotide sequence identity, at least 97% nucleotide sequence identity, or 100% nucleotide sequence identity to SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, or SEQ ID NO: 26, such that the resulting polypeptide is not SEQ ID NO:

1. **Claim 4**: A sortilin propeptide comprising an amino acid sequence having at least 70% sequence identity, at least 90% sequence identity, or 100% sequence identity to SEQ ID NO: 48, such that the amino acid sequence is not SEQ ID NO:

48. **Claim 5** A sortilin propeptide polynucleotide comprising a polynucleotide having at least 40% nucleotide sequence identity, at least 70% nucleotide sequence identity, at least 97% nucleotide sequence identity, or 100% nucleotide sequence identity to SEQ ID NO:

49. **Claim 6** A tagged protein comprising a tag and a therapeutic protein, wherein the tag is a tagged protein comprising a neurotensin peptide and / or a sortilin propeptide. **Claim 7** The tagged protein according to Claim 6, wherein the neurotensin peptide comprises a polypeptide that is at least 70% identical, at least 90% identical, or 100% identical to SEQ ID NO:

1. **Claim 8** The tagged protein according to claim 6, wherein the neurotensin peptide comprises a polypeptide that is at least 70% identical, at least 90% identical, or 100% identical to SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, or SEQ ID NO:

13.

9. The tagged protein according to claim 6, wherein the polynucleotide encoding the neurotensin peptide comprises a polynucleotide having at least 40% sequence identity, at least 70% sequence identity, at least 97% sequence identity, or 100% sequence identity to SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, or SEQ ID NO:

26.

10. The tagged protein according to claim 6, wherein the sortilin propeptide comprises a polypeptide that is at least 70% identical, at least 90% identical, or 100% identical to SEQ ID NO:

48.

11. The tagged protein according to claim 6, wherein the polynucleotide encoding the sortilin propeptide comprises a polynucleotide having at least 70% sequence identity, at least 97% sequence identity, or 100% sequence identity to SEQ ID NO:

50.

12. The tagged protein according to claim 6, wherein the therapeutic protein comprises a lysosomal protein.

13. The tagged protein according to claim 6, wherein the therapeutic protein comprises alpha-galactosidase, beta-galactosidase, beta-hexosaminidase, galactosylceramidase, arylsulfatase, beta-glucocerebrosidase, glucocerebrosidase, lysosomal acid lipase, lysosomal enzyme acid sphingomyelinase, formylglycine-generating enzyme, iduronidase, acetyl-CoA:alpha-glucosaminide N-acetyltransferase, glycosaminoglycan alpha-L-iduronohydrolase, heparan N-sulfatase, N-acetyl-alpha-D-glucosaminidase (NAGLU), iduronic acid-2-sulfatase, galactosamine-6-sulfate sulfatase, N-acetylgalactosamine-6-sulfatase, glycosaminoglycan N-acetylgalactosamine 4-sulfatase, beta-glucuronidase, hyaluronidase, alpha-N-acetylneuraminidase (sialidase), ganglioside sialidase, phosphotransferase, alpha-glucosidase, alpha-D-mannosidase, beta-D-mannosidase, aspartylglucosaminidase, alpha-L-fucosidase, and other Batten-related proteins or enzymatically active fragments thereof.

14. The tagged protein according to claim 13, wherein the Batten-related protein comprises palmitoyl protein thioesterase 1 (PPT1) (CLN1), tripeptidyl peptidase 1 (TPP1) (CLN2), cathepsin D (CTSD) (CLN10), progranulin (PGRN) (CLN11), and cathepsin F (CTSF) (CLN13).

15. The tagged protein according to claim 6, wherein the therapeutic protein comprises alpha-glucosidase (GAA).

16. The tagged protein according to any one of claims 6 to 15, wherein the tagged protein comprises one or more of an affinity tag, a linker peptide, and a secretion signal peptide.

17. A polynucleotide comprising a nucleotide sequence encoding the tagged protein according to claim 16.

18. A method for producing the tagged protein according to claim 16, the method comprising expressing the tagged protein and purifying the tagged protein.

19. The method according to claim 18, wherein the tagged protein is expressed in Expi293F cells, PC12 cells, COS7 cells, HAP1 cells, Chinese hamster ovary (CHO) cells, HeLa cells, human embryonic kidney (HEK) cells, primary mouse myoblasts, NIH 3T3 cells, Escherichia coli cells, Sf9 cells, Sf21 cells, Saccharomyces cerevisiae or variants thereof.

20. A method for determining the cellular uptake of a tagged protein according to any one of claims 6 to 15, comprising culturing cells, incubating the cells with the tagged protein, and identifying the tagged protein transported into the cells.

21. The method according to claim 20, comprising determining the functionality of the tagged protein transported into the cells.

22. The method according to claim 20, comprising quantifying the tagged protein transported into the cells.

23. The method according to claim 20, wherein the cells are Expi293F cells, PC12 cells, COS7 cells, HAP1 cells, Chinese hamster ovary (CHO) cells, HeLa cells, human embryonic kidney (HEK) cells, primary mouse myoblasts, NIH 3T3 cells, Escherichia coli cells, Sf9 cells, Sf21 cells, yeast cells or variants thereof.

24. A pharmaceutical formulation comprising a tagged protein according to any one of claims 6 to 15, and a pharmaceutically acceptable carrier.

25. A medicament for treating a disease or disorder, comprising the pharmaceutical formulation according to claim 24.

26. The medicament according to claim 25, wherein the disease or disorder is Fabry disease and the therapeutic protein comprises alpha-galactosidase A.

27. The medicament according to claim 25, wherein the disease or disorder is Pompe disease and the therapeutic protein comprises alpha-glucosidase A.

28. The disease or disorder is Batten disease, and the therapeutic protein comprises palmitoyl-protein thioesterase 1 (PPT1) (CLN1), tripeptidyl peptidase 1 (TPP1) (CLN2), cathepsin D (CTSD) (CLN10), progranulin (PGRN) (CLN11) and cathepsin F (CTSF) (CLN13), the medicament according to claim 25.

29. The pharmaceutical formulation is administered into the subarachnoid space, intravenously, intracisternally, intraventricularly, intraocularly, intravitreally, to the retina, subretinally, intramuscularly, subcutaneously, intracerebrally, surgically or parenchymally, the medicament according to claim 25.

30. A gene therapy delivery system; and A polynucleotide encoding the tagged protein according to any one of claims 6 to 15 A gene therapy composition comprising the same.

31. The gene therapy delivery system comprises one or more of a vector, liposome, lipid-nucleic acid nanoparticle, exosome, and gene editing system, the gene therapy composition according to claim 30.

32. The gene therapy delivery system comprises one or more of clustered regularly interspaced short palindromic repeats (CRISPR)-associated protein 9 (CRISPR-Cas-9), transcription activator-like effector nuclease (TALEN) or ZNF (zinc finger protein), the gene therapy composition according to claim 30.

33. The gene therapy delivery system comprises a viral vector, the gene therapy composition according to claim 30.

34. The viral vector comprises one or more of an adenoviral vector, an adeno-associated viral vector, a lentiviral vector, a retroviral vector, a poxviral vector or a herpes simplex viral vector, the gene therapy composition according to claim 33.

35. The viral vector comprises a viral polynucleotide operably linked to the polynucleotide encoding the tagged protein, the gene therapy composition according to claim 33.

36. The viral vector comprises at least one terminal inverted repeat (ITR), the gene therapy composition according to claim 33.

37. The viral vector comprises one or more of an SV40 intron, a polyadenylation signal or a stabilizing element, the gene therapy composition according to claim 33.

38. The gene therapy composition according to claim 30, wherein the gene therapy delivery system comprises a promoter.

39. The gene therapy composition according to claim 30, wherein the gene therapy delivery system comprises a polynucleotide encoding a secretory signal peptide.

40. The gene therapy composition according to claim 30, comprising a pharmaceutically acceptable carrier.

41. A medicament for treating a disease or disorder, comprising a formulation of the gene therapy composition according to claim 40.

42. The medicament according to claim 41, wherein the disease or disorder is Fabry disease and the sequence of the polynucleotide comprises a nucleotide sequence encoding alpha-galactosidase A.

43. The medicament according to claim 41, wherein the disease or disorder is Pompe disease and the sequence of the polynucleotide comprises a nucleotide sequence encoding alpha-glucosidase.

44. The medicament according to claim 41, wherein the disease or disorder is Batten disease and the sequence of the polynucleotide comprises nucleotide sequences encoding palmitoyl protein thioesterase 1 (PPT1) (CLN1), tripeptidyl peptidase 1 (TPP1) (CLN2), cathepsin D (CTSD) (CLN10), progranulin (PGRN) (CLN11) and cathepsin F (CTSF) (CLN13).

45. The medicament according to claim 41, wherein the formulation of the gene therapy composition is administered into the subarachnoid space, intravenously, into the cisterna magna, intraventricularly, intravitreally, into the vitreous body, to the retina, subretinally, intramuscularly, subcutaneously, intracerebrally, surgically or intrathecally.