ANTI-TfR: GAA AND ANTI-CD63: GAA INSERTION FOR THE TREATMENT OF PONSITY DISEASE
By fusing lysosomal α-glucosidase and a delivery domain into a nucleic acid construct, the problem of insufficient ERT delivery across the blood-brain barrier and skeletal muscle in the treatment of Pompe disease was solved, enabling effective treatment for neonatal patients.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- REGENERON PHARMACEUTICALS INC
- Filing Date
- 2024-07-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing enzyme replacement therapy (ERT) has difficulty effectively crossing the blood-brain barrier when treating Pompe disease (PD), and skeletal muscle and central nervous system (CNS) are undertreated, with additional barriers, especially in neonate and adolescent patients.
Nucleic acid constructs containing multi-domain therapeutic proteins are provided. By fusing lysosomal α-glucosidase with delivery domains (such as CD63- or TfR-binding delivery domains) and using specific polyadenylation signals and splice acceptors, the nucleic acid constructs are integrated and expressed in the target genome to improve the delivery and accumulation of therapeutic proteins in cells.
It enhances the delivery of therapeutic proteins in skeletal muscle and CNS, reduces glycogen accumulation, and effectively treats or prevents symptoms of Pompe disease, especially in newborn patients.
Smart Images

Figure CN121909288A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims the benefit of U.S. Application No. 63 / 516,395, filed July 28, 2023, which is incorporated herein by reference in its entirety for all purposes.
[0002] References to sequence lists Submit as an XML file The sequence list written to file 616967SEQLIST.xml is 1,213,989 bytes long, created on July 24, 2024, and incorporated by reference. Background Technology
[0003] Pompe disease (PD), or type II glycogen storage disease, is a monogenic lysosomal disorder caused by a deficiency of the enzyme lysosomal acid alpha-glucosidase (GAA). GAA deficiency leads to the accumulation of its substrate, glycogen, in the lysosomes of tissues, including skeletal and cardiac muscle. This abnormal accumulation of glycogen in muscle fibers results in progressive damage to muscle tissue, with symptoms including cardiac hypertrophy, mild to severe muscle weakness, and ultimately death from heart or respiratory failure. Infantile PD (IOPD) is associated with <1% of normal GAA activity. It is severe and affects visceral organs, muscles, and the central nervous system (CNS). Late-onset PD (LOPD) is associated with 2%–40% GAA activity. It is less severe and primarily involves respiratory and skeletal muscles.
[0004] The only approved treatment for PD is enzyme replacement therapy (ERT). Recombinant human (rh) GAA is delivered to patients via intravenous infusion every other week. While ERT has been very successful in treating the cardiac manifestations of PD, skeletal muscle and the CNS remain rarely treated with it. The primary mechanism by which rhGAA reaches lysosomes is through uptake by the cation-independent mannose-6-phosphate (CIMPR) receptor, which binds to M6P on rhGAA. However, CIMPR expression in skeletal muscle is very low, and rhGAA is weakly mannose-6-phosphorylated. Furthermore, CIMPR may be misdirected into autophagosomes in affected cells instead of lysosomes, and significant amounts of the drug are also taken up by the liver, an organ without primary pathology in PD. ERT does not cross the blood-brain barrier. Additionally, the unique circumstances of neonatal and adolescent patients present additional barriers to treatment of PD that may require early in life. Summary of the Invention
[0005] This provides the ability to insert the coding sequences of multi-domain therapeutic proteins (e.g., GAA fusion proteins) into target genomic loci (such as endogenous...). ALBNucleic acid constructs and compositions that encode the multidomain therapeutic protein (e.g., GAA fusion protein) at a locus and / or express the coding sequence of such a multidomain therapeutic protein. These nucleic acid constructs and compositions can be used in methods of: integrating or inserting the nucleic acid of a multidomain therapeutic protein (e.g., GAA fusion protein) into a target genomic locus in a cell or cell population or subject; expressing a multidomain therapeutic protein (e.g., GAA fusion protein) in a cell or cell population or subject; reducing glycogen accumulation in a cell or cell population or subject; treating Pompe disease or GAA deficiency in a subject; and preventing or reducing the onset of signs or symptoms of Pompe disease in subjects (such as subjects with reduced GAA activity or expression) and subjects diagnosed with Pompe disease (including neonatal subjects). In some embodiments, these cells, cell populations, or subjects are neonatal cells, neonatal cell populations, or neonatal subjects.
[0006] In one aspect, a composition comprising a nucleic acid construct containing a coding sequence for a multi-domain therapeutic protein, the multi-domain therapeutic protein containing a delivery domain fused to a lysosomal α-glucosidase polypeptide, wherein the lysosomal α-glucosidase coding sequence is CpG depleted relative to a wild-type lysosomal α-glucosidase coding sequence, optionally wherein the delivery domain is a CD63-binding delivery domain or a TfR-binding delivery domain. In some such compositions, the nucleic acid construct contains a polyadenylation signal or sequence downstream of the coding sequence of the multi-domain therapeutic protein. In some such compositions, the polyadenylation signal comprises a bovine growth hormone (BGH) polyadenylation signal, a simian virus 40 (SV40) polyadenylation signal, or a combination of a bovine growth hormone polyadenylation signal and an SV40 polyadenylation signal. In some such compositions, the SV40 polyadenylation signal is a unidirectional SV40 late polyadenylation signal, wherein each instance of the sequence AATAAA in the back chain is mutated in the unidirectional SV40 late polyadenylation signal, optionally wherein the SV40 polyadenylation signal is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in SEQ ID NO: 752, and optionally wherein the SV40 polyadenylation signal comprises the sequence shown in SEQ ID NO: 752. In some such compositions, the polyadenylation signal comprises a BGH polyadenylation signal, optionally wherein the BGH polyadenylation signal is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in SEQ ID NO: 751, and optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO: 751. In some such compositions, the polyadenylation signal comprises a BGH polyadenylation signal and an SV40 polyadenylation signal, optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO: 751, and optionally wherein the SV40 polyadenylation signal comprises the sequence shown in SEQ ID NO: 752. Optionally, the polyadenylation signal comprising the BGH and SV40 polyadenylation signals is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in SEQ ID NO: 795, and optionally, the polyadenylation signal comprising the BGH and SV40 polyadenylation signals comprises the sequence shown in SEQ ID NO: 795. In some such compositions, the nucleic acid construct is a single-phase nucleic acid construct.
[0007] In some of these compositions, the coding sequence of the delivery domain is modified to remove one or more hidden splice sites, the coding sequence of the lysosomal α-glucosidase polypeptide is modified to remove one or more hidden splice sites, or the coding sequence of the multi-domain therapeutic protein is modified to remove one or more hidden splice sites. In some of these compositions, the coding sequence of the delivery domain is CpG depleted, or the coding sequence of the multi-domain therapeutic protein is CpG depleted. In some of these compositions, the coding sequence of the delivery domain is codon-optimized and CpG depleted, the coding sequence of the lysosomal α-glucosidase polypeptide is codon-optimized and CpG depleted, or the coding sequence of the multi-domain therapeutic protein is codon-optimized and CpG depleted.
[0008] In some of these compositions, the nucleic acid construct includes a splice acceptor located upstream of the coding sequence of a multidomain therapeutic protein. In some of these compositions, the nucleic acid construct does not include a homologous arm. In some of these compositions, the nucleic acid construct from 5' to 3' includes: a splice acceptor, the coding sequence of a multidomain therapeutic protein, and a polyadenylation signal or sequence, wherein the nucleic acid construct does not include a promoter driving the expression of the multidomain therapeutic protein, and wherein the nucleic acid construct does not include a homologous arm. In some of these compositions, the nucleic acid construct includes a homologous arm. In some of these compositions, the nucleic acid construct does not include a promoter driving the expression of a multidomain therapeutic protein. In some of these compositions, the coding sequence of the multidomain therapeutic protein is operatively linked to a promoter, optionally wherein the promoter is a liver-specific promoter.
[0009] In some such compositions, the C-terminus of the delivery domain is fused to the N-terminus of the lysosomal α-glucosidase polypeptide. In some such compositions, the delivery domain is fused to the lysosomal α-glucosidase polypeptide via a peptide linker. In some such compositions, the lysosomal α-glucosidase polypeptide lacks the lysosomal α-glucosidase signal peptide and propeptide. In some such compositions, the lysosomal α-glucosidase polypeptide comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 727. In some such compositions, the lysosomal α-glucosidase coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 750, optionally wherein the nucleotide at position 1095 is G, the nucleotide at position 1098 is C, and the nucleotide at position 2343 is G. In some such compositions, the lysosomal α-glucosidase coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any of SEQ ID NO: 750, and encodes a lysosomal α-glucosidase protein comprising SEQ ID NO: 727, optionally wherein the nucleotide at position 1095 is G, the nucleotide at position 1098 is C, and the nucleotide at position 2343 is G. In some such compositions, the lysosomal α-glucosidase coding sequence comprises, substantially comprises, or comprises of the sequence shown in any of SEQ ID NO: 750. In some such compositions, the lysosomal α-glucosidase coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 749, optionally wherein the nucleotide at position 2343 is G. In some such compositions, the lysosomal α-glucosidase coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any of SEQ ID NO: 749, and encodes the lysosomal α-glucosidase protein comprising SEQ ID NO: 727, optionally wherein the nucleotide at position 2343 is G. In some such compositions, the lysosomal α-glucosidase coding sequence comprises, substantially comprises, or comprises of the sequence shown in any of SEQ ID NO: 749.
[0010] In some such compositions, the delivery domain is a TfR-binding delivery domain. In some such compositions, the TfR-binding delivery domain comprises an anti-TfR antigen-binding protein, optionally wherein the antigen-binding protein is in the form of about 41 nM K. D Or, with a stronger affinity, bind to the human transferrin receptor, optionally wherein the antigen-binding protein binds at approximately 3 nM K. D It may bind to the human transferrin receptor with a stronger affinity, or optionally, the antigen-binding protein may bind to it with a K+ of approximately 0.45 nM to 3 nM. D Binding to human transferrin receptor. In some such compositions, the anti-TfR antigen-binding protein comprises: (i) an HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 171, 181, 191, 201, 211, 221, 231, 241, 251, 261, 271, 281, 291, 301, 311, 321, 331, 341, 351, 361, 371, 381, 391, 401, 411, 421, 431, 441, 451, 461, 471, or 481; and / or (ii) an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence ... The amino acid sequences (or variants thereof) shown as 176, 186, 196, 206, 216, 226, 236, 246, 256, 266, 276, 286, 296, 306, 316, 326, 336, 346, 356, 366, 376, 386, 396, 406, 416, 426, 436, 446, 456, 466, 476, or 486.
[0011] In some such compositions, the anti-TfR antigen-binding protein comprises: (1) an HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 171 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 176 (or a variant thereof); (2) an HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 181 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 186 (or a variant thereof); (3) an HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 191 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 191 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 191 (or a variant thereof). (3) The amino acid sequence shown in SEQ ID NO: 201 (or a variant thereof); (4) An HCVR containing HCDR1, HCDR2, and HCDR3, wherein the HCVR contains the amino acid sequence shown in SEQ ID NO: 201 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, wherein the LCVR contains the amino acid sequence shown in SEQ ID NO: 206 (or a variant thereof); (5) An HCVR containing HCDR1, HCDR2, and HCDR3, wherein the HCVR contains the amino acid sequence shown in SEQ ID NO: 211 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, wherein the LCVR contains the amino acid sequence shown in SEQ ID NO: 216 (or a variant thereof); (6) An HCVR containing HCDR1, HCDR2, and HCDR3, wherein the HCVR contains the amino acid sequence shown in SEQ ID NO: 221 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, wherein the LCVR contains the amino acid sequence shown in SEQ ID NO: 221 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, wherein the LCVR contains the amino acid sequence shown in SEQ ID NO: 216 (or a variant thereof); (7) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 231 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 236 (or a variant thereof); (8) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 241 (or a variant thereof);(9) An HCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 246 (or a variant thereof); and an LCVR comprising HCDR1, HCDR2, and HCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 251 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 256 (or a variant thereof); (10) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 261 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 266 (or a variant thereof); (11) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 266 (or a variant thereof); (12) An HCVR containing HCDR1, HCDR2, and HCDR3, which contains the amino acid sequence shown in SEQ ID NO: 271 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, which contains the amino acid sequence shown in SEQ ID NO: 281 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, which contains the amino acid sequence shown in SEQ ID NO: 286 (or a variant thereof); (13) An HCVR containing HCDR1, HCDR2, and HCDR3, which contains the amino acid sequence shown in SEQ ID NO: 291 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, which contains the amino acid sequence shown in SEQ ID NO: 296 (or a variant thereof); (14) An HCVR containing HCDR1, HCDR2, and HCDR3, which contains the amino acid sequence shown in SEQ ID NO: 296 (or a variant thereof); The amino acid sequence shown in SEQ ID NO: 301 (or a variant thereof); and an LCVR containing LCDR1, LCDR2 and LCDR3, the LCVR containing the amino acid sequence shown in SEQ ID NO: 306 (or a variant thereof); (15) an HCVR containing HCDR1, HCDR2 and HCDR3, the HCVR containing the amino acid sequence shown in SEQ ID NO: 311 (or a variant thereof); and an LCVR containing LCDR1, LCDR2 and LCDR3, the LCVR containing the amino acid sequence shown in SEQ ID NO: 316 (or a variant thereof);(16) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 321 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 326 (or a variant thereof); (17) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 331 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 336 (or a variant thereof); (18) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 341 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 341 (or a variant thereof); (19) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 351 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 356 (or a variant thereof); (20) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 361 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 366 (or a variant thereof); (21) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 371 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 371 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 371 (or a variant thereof). (22) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 381 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 386 (or a variant thereof); (23) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof).(24) An HCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof); and an LCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 401 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 406 (or a variant thereof); (25) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 411 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 416 (or a variant thereof); and (26) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 416 (or a variant thereof). (27) An HCVR containing HCDR1, HCDR2, and HCDR3, which contains the amino acid sequence shown in SEQ ID NO: 426 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, which contains the amino acid sequence shown in SEQ ID NO: 431 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, which contains the amino acid sequence shown in SEQ ID NO: 436 (or a variant thereof); (28) An HCVR containing HCDR1, HCDR2, and HCDR3, which contains the amino acid sequence shown in SEQ ID NO: 441 (or a variant thereof); and an LCVR containing LCDR1, LCDR2, and LCDR3, which contains the amino acid sequence shown in SEQ ID NO: 446 (or a variant thereof); (29) An HCVR containing HCDR1, HCDR2, and HCDR3, which contains the amino acid sequence shown in SEQ ID NO: 446 (or a variant thereof); The amino acid sequence shown in SEQ ID NO: 451 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 456 (or a variant thereof); (30) an HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 461 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 466 (or a variant thereof);(31) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 471 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 476 (or a variant thereof); or (32) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 481 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 486 (or a variant thereof).
[0012] In some such compositions, the anti-TfR antigen-binding protein comprises: (1) an HCVR comprising HCDR1, HCDR2, and HCDR3, the HCVR comprising the amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, the LCVR comprising the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof); or (2) an HCVR comprising HCDR1, HCDR2, and HCDR3, the HCVR comprising the amino acid sequence shown in SEQ ID NO: 411 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, the LCVR comprising the amino acid sequence shown in SEQ ID NO: 416 (or a variant thereof). In some such compositions, the anti-TfR antigen-binding protein comprises: an HCVR comprising HCDR1, HCDR2, and HCDR3, the HCVR comprising the amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, the LCVR comprising the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof).
[0013] In some such compositions, the anti-TfR antigen-binding protein comprises: (a) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 172, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 173, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 174; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 177, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 178, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 179; (b) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 182, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 183, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 184; and LC ... (c) An HCVR comprising: an LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 187, an LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 188, and an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 189; and an LCVR comprising: an LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 197, an LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 193, and an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 194; and an LCVR comprising: an LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 197, an LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 198, and an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 199; and an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 197; and an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 199; and an LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 197, an LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 198, an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 199; and an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 197; HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 203 and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 204; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 207, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 208 and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 209;(e) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 212, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 213, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 214; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 217, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 218, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 219; (f) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 222, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 223, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 224; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 227, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 212, HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 214, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 219 ...22, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 223, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 224; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID (g) HCVR comprising: LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 228 and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 229; (g) HCVR comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 232, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 233 and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 234; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 237, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 238 and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 239; (h) HCVR comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 242, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 243 and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 239; HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 244; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 247, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 248, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 249;(i) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 252, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 253, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 254; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 257, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 258, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 259; (j) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 262, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 263, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 264; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 267, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 259, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 262, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 264; (k) An HCVR comprising: an LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 268 and an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 269; (k) An HCVR comprising: an HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 272, an HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 273 and an HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 274; and an LCVR comprising: an LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 277, an LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 278 and an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 279; (l) An HCVR comprising: an HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 282, an HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 283 and an LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 274; HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 284; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 287, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 288, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 289;(m) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 292, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 293, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 294; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 297, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 298, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 299; (n) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 302, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 303, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 304; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 307, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 293, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 294; LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 308 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 309; (o) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 312, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 313 and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 314; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 317, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 318 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 319; (p) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 322, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 323 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 322 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 323 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 314; HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 324; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 327, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 328, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 329;(q) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 332, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 333, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 334; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 337, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 338, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 339; (r) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 342, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 343, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 344; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 347, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 339, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 342, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 344; LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 348 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 349; (s) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 352, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 353 and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 354; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 357, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 358 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 359; (t) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 362, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 363 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 354. HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 364; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 367, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 368, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 369;(u) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 372, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 373, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 374; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 377, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 378, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 379; (v) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 382, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 383, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 384; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 387, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 379, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 382, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 384; (w) An LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 388 and an LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 389; and an HCVR comprising: an HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 392, an HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 393 and an HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 394; and an LCVR comprising: an LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 397, an LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 398 and an LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 399; and an HCVR comprising: an LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 402, an HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 403 and an LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 394; HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 404; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 407, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 408, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 409;(y) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 412, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 413, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 414; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 417, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 418, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 419; (z) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 422, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 423, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 424; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 427, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 412, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 414, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 419; LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 428 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 429; (aa) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 432, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 433 and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 434; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 437, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 438 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 439; (ab) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 442, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 443 and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 439; HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 444; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 447, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 448, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 449;(ac) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 452, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 453, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 454; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 457, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 458, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 459; (ad) HCVR, comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 462, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 463, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 464; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 467, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 452, HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 454, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 459 ...62, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 464; and LCVR, comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 467, HCDR2 containing the amino acid sequence (or a variant thereof) LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 468 and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 469; (ae) HCVR comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 472, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 473 and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 474; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 477, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 478 and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 479; and / or (af) HCVR comprising: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 482, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 483 and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 474; HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 484; and LCVR comprising: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 487, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 488, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 489.
[0014] In some such compositions, the anti-TfR antigen-binding protein comprises: (a) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 392, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 393, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 394; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 397, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 398, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 399; or (b) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 412, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 413, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 414; and LCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 392, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 393, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 39 ... LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 417, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 418, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 419. In some such compositions, the anti-TfR antigen-binding protein comprises: HCVR, which comprises: HCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 392, HCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 393, and HCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 394; and LCVR, which comprises: LCDR1 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 397, LCDR2 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 398, and LCDR3 containing the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 399.
[0015] In some such compositions, the anti-TfR antigen-binding protein comprises: (i) HCVR, which comprises the amino acid sequence shown in SEQ ID NO: 171 (or a variant thereof); and LCVR, which comprises the amino acid sequence shown in SEQ ID NO: 176 (or a variant thereof); (ii) HCVR, which comprises the amino acid sequence shown in SEQ ID NO: 181 (or a variant thereof); and LCVR, which comprises the amino acid sequence shown in SEQ ID NO: 186 (or a variant thereof); (iii) HCVR, which comprises the amino acid sequence shown in SEQ ID NO: 191 (or a variant thereof); and LCVR, which comprises the amino acid sequence shown in SEQ ID NO: 196 (or a variant thereof); (iv) HCVR, which comprises the amino acid sequence shown in SEQ ID NO: 201 (or a variant thereof); and LCVR, which comprises the amino acid sequence shown in SEQ ID NO: 206 (or a variant thereof); (v) HCVR, which comprises the amino acid sequence shown in SEQ ID NO: 211 (or a variant thereof); and LCVR, which comprises the amino acid sequence shown in SEQ ID NO: 176 (or a variant thereof). (vi) HCVR containing the amino acid sequence shown in SEQ ID NO: 221 (or a variant thereof); and LCVR containing the amino acid sequence shown in SEQ ID NO: 226 (or a variant thereof); (vii) HCVR containing the amino acid sequence shown in SEQ ID NO: 231 (or a variant thereof); and LCVR containing the amino acid sequence shown in SEQ ID NO: 236 (or a variant thereof); (viii) HCVR containing the amino acid sequence shown in SEQ ID NO: 241 (or a variant thereof); and LCVR containing the amino acid sequence shown in SEQ ID NO: 246 (or a variant thereof); (ix) HCVR containing the amino acid sequence shown in SEQ ID NO: 251 (or a variant thereof); and LCVR containing the amino acid sequence shown in SEQ ID NO: 256 (or a variant thereof); (x) HCVR containing the amino acid sequence shown in SEQ ID NO: 261 (or a variant thereof); and LCVR containing the amino acid sequence shown in SEQ ID NO: 266 (or a variant thereof); (xi) HCVR containing the amino acid sequence shown in SEQ ID NO: 266 (or a variant thereof); The amino acid sequence (or a variant thereof) shown in SEQ ID NO: 271; and LCVR, which contains the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 276; (xii) HCVR, which contains the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 281; and LCVR, which contains the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 286.(xiii) HCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 291; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 296; (xiv) HCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 301; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 306; (xv) HCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 311; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 316; (xvi) HCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 321; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 326; (xvii) HCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 331; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 336; (xviii ...21; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 331; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 336; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 331; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 331; and LCVR comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 331; and The amino acid sequence shown in SEQ ID NO: 341 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 346 (or a variant thereof); (xix) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 351 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 356 (or a variant thereof); (xx) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 361 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 366 (or a variant thereof); (xxi) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 371 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 376 (or a variant thereof); (xxii) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 381 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 386 (or a variant thereof); (xxiii) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 386 (or a variant thereof); The amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof); (xxiv) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 401 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 406 (or a variant thereof).(xxv) HCVR, comprising the amino acid sequence shown in SEQ ID NO: 411 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 416 (or a variant thereof); (xxvi) HCVR, comprising the amino acid sequence shown in SEQ ID NO: 421 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 426 (or a variant thereof); (xxvii) HCVR, comprising the amino acid sequence shown in SEQ ID NO: 431 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 436 (or a variant thereof); (xxviii) HCVR, comprising the amino acid sequence shown in SEQ ID NO: 441 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 446 (or a variant thereof); (xxix) HCVR, comprising the amino acid sequence shown in SEQ ID NO: 451 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 456 (or a variant thereof); (xxx) HCVR, comprising the amino acid sequence shown in SEQ ID NO: The amino acid sequence shown in SEQ ID NO: 461 (or a variant thereof); and LCVR containing the amino acid sequence shown in SEQ ID NO: 466 (or a variant thereof); (xxxi) HCVR containing the amino acid sequence shown in SEQ ID NO: 471 (or a variant thereof); and LCVR containing the amino acid sequence shown in SEQ ID NO: 476 (or a variant thereof); and / or (xxxii) HCVR containing the amino acid sequence shown in SEQ ID NO: 481 (or a variant thereof); and LCVR containing the amino acid sequence shown in SEQ ID NO: 486 (or a variant thereof).
[0016] In some such compositions, the anti-TfR antigen-binding protein comprises: (i) HCVR, which comprises the amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof); and LCVR, which comprises the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof); or (ii) HCVR, which comprises the amino acid sequence shown in SEQ ID NO: 411 (or a variant thereof); and LCVR, which comprises the amino acid sequence shown in SEQ ID NO: 416 (or a variant thereof). In some such compositions, the anti-TfR antigen-binding protein comprises: HCVR, which comprises the amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof); and LCVR, which comprises the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof).
[0017] In some of these compositions, the TfR-binding delivery domain is an antigen-binding protein that binds one or more hTfR epitopes selected from the following: (a) an epitope containing the sequence LLNE (SEQ ID NO: 796) and / or an epitope containing the sequence TYKEL (SEQ ID NO: 706); (b) an epitope containing the sequence DSTDFTGT (SEQ ID NO: 797) and / or an epitope containing the sequence VKHPVTGQF (SEQ ID NO: 798) and / or an epitope containing the sequence IERIPEL (SEQ ID NO: 799); (c) an epitope containing the sequence LNENSYVPREAGSQKDEN (SEQ ID NO: 800); (d) an epitope containing the sequence FEDL (SEQ ID NO: 718); (e) an epitope containing the sequence IVDKNGRL (SEQ ID NO: 801); (f) an epitope containing the sequence IVDKNGRLVY (SEQ ID NO: 706). Epitopes containing: (g) the sequence DQTKF (SEQ ID NO: 803); (h) the sequence LVENPGGY (SEQ ID NO: 804) and / or the sequence PIVNAELSF (SEQ ID NO: 805) and / or the sequence PYLGTTMDT (SEQ ID NO: 806); (i) the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or the sequence TYKEL (SEQ ID NO: 706); (j) the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 802); (h) the sequence DQTKF (SEQ ID NO: 803); (h) the sequence LVENPGGY (SEQ ID NO: 804) and / or the sequence PIVNAELSF (SEQ ID NO: 805) and / or the sequence PYLGTTMDT (SEQ ID NO: 806); (i) the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or the sequence TYKEL (SEQ ID NO: 706); (j) the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 802); (g) the sequence DQTKF (SEQ ID NO: 803); (h) the sequence LVENPGGY (SEQ ID NO: 804) and / or the sequence PIVNAELSF (SEQ ID NO: 805) and / or the sequence PYLGTTMDT (SEQ ID NO: 806); (j) the Epitopes of (708) and / or epitopes containing the sequence LIERIPELNKVARAAAE (SEQ ID NO: 709); (k) epitopes containing the sequence LNENSYVPREAGSQKDENL (SEQ ID NO: 710); (l) epitopes containing the sequence GTKKDFEDL (SEQ ID NO: 711); (m) epitopes containing the sequence SVIIVDKNGRLVYLVENPGGYVAYSK (SEQ ID NO: 712);(n) Epitopes containing the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes containing the sequence DQTKFPIVNAEL (SEQ ID NO: 714) and / or epitopes containing the sequence TYKELIERIPELNK (SEQ ID NO: 715); (o) Epitopes containing the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes containing the sequence TYKELIERIPELNK (SEQ ID NO: 715); (p) Epitopes containing the sequence SVIIVDKNGRLVYLVENPGGYVAY (SEQ ID NO: 716); (q) Epitopes containing the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes containing the sequence FGNMEGDCPSDWKTDSTCRM (SEQ ID NO: 716). Epitopes of (717); (r) epitopes containing the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes containing the sequence LVENPGYVAYSKAATVTGKL (SEQ ID NO: 719) and / or epitopes containing the sequence IYMDQTKFPIVNAELSF (SEQ ID NO: 720) and / or epitopes containing the sequence ISRAAAEKL (SEQ ID NO: 721) and / or epitopes containing the sequence VTSESKNVKLTVSNVLKE (SEQ ID NO: 722) and / or epitopes containing the sequence FCEDTDYPYLGTTMDT (SEQ ID NO: 723); (s) epitopes contained within or overlapping with the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes contained within the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 717). (705) Epitopes contained in or overlapping with the sequence and / or included in or overlapping with the sequence TYKEL (SEQ ID NO: 706); (t) Epitopes contained in or overlapping with the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or included in or overlapping with the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 708) and / or included in or overlapping with the sequence LIERIPELNKVARAAAE (SEQ ID NO: 709); (u) Epitopes contained in or overlapping with the sequence LNENSYVPREAGSQKDENL (SEQ ID NO: 710);(v) Epitopes contained in or overlapping with the sequence GTKKDFEDL (SEQ ID NO: 711); (w) Epitopes contained in or overlapping with the sequence SVIIVDKNGRLVYLVENPGGYVAYSK (SEQ ID NO: 712); (x) Epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or Epitopes contained in or overlapping with the sequence DQTKFPIVNAEL (SEQ ID NO: 714) and / or Epitopes contained in or overlapping with the sequence TYKELIERIPELNK (SEQ ID NO: 715); (y) Epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or Epitopes contained in or overlapping with the sequence TYKELIERIPELNK (SEQ ID NO: 715). Epitopes contained in or overlapping with the sequence 715); (z) epitopes contained in or overlapping with the sequence SVIIVDKNGRLVYLVENPGGYVAY (SEQ ID NO: 716); (aa) epitopes contained in or overlapping with the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes contained in or overlapping with the sequence FGNMEGDCPSDWKTDSTCRM (SEQ ID NO: 717); and (ab) epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes contained in or overlapping with the sequence LVENPGYVAYSKAATVTGKL (SEQ ID NO: 719) and / or epitopes contained in or overlapping with the sequence IYMDQTKFPIVNAELSF (SEQ ID NO: 716). Epitopes contained in or overlapping with the sequence ISRAAAEKL (SEQ ID NO: 721) and / or contained in or overlapping with the sequence VTSESKNVKLTVSNVLKE (SEQ ID NO: 722) and / or contained in or overlapping with the sequence FCEDTDYPYLGTTMDT (SEQ ID NO: 723).
[0018] In some such compositions, the TfR-binding delivery domain comprises an antibody or an antigen-binding fragment thereof that binds one or more hTfR epitopes selected from the following: (a) an epitope consisting of the sequence LLNE (SEQ ID NO: 796) and / or an epitope consisting of the sequence TYKEL (SEQ ID NO: 706); (b) an epitope consisting of the sequence DSTDFTGT (SEQ ID NO: 797) and / or an epitope consisting of the sequence VKHPVTGQF (SEQ ID NO: 798) and / or an epitope consisting of the sequence IERIPEL (SEQ ID NO: 799); (c) an epitope consisting of the sequence LNENSYVPREAGSQKDEN (SEQ ID NO: 800); (d) an epitope consisting of the sequence FEDL (SEQ ID NO: 718); (e) an epitope consisting of the sequence IVDKNGRL (SEQ ID NO: 801); (f) an epitope consisting of the sequence IVDKNGRLVY (SEQ ID NO: 706). (g) Epitopes composed of the sequence DQTKF (SEQ ID NO: 803); (h) Epitopes composed of the sequence LVENPGGY (SEQ ID NO: 804) and / or the sequence PIVNAELSF (SEQ ID NO: 805) and / or the sequence PYLGTTMDT (SEQ ID NO: 806); (i) Epitopes composed of the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or the sequence TYKEL (SEQ ID NO: 706); (j) Epitopes composed of the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 702). Epitopes consisting of (708) and / or the sequence LIERIPELNKVARAAAE (SEQ ID NO: 709); (k) the sequence LNENSYVPREAGSQKDENL (SEQ ID NO: 710); (l) the sequence GTKKDFEDL (SEQ ID NO: 711); and (m) the sequence SVIIVDKNGRLVYLVENPGGYVAYSK (SEQ ID NO: 712).(n) Epitopes consisting of the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or the sequence DQTKFPIVNAEL (SEQ ID NO: 714) and / or the sequence TYKELIERIPELNK (SEQ ID NO: 715); (o) Epitopes consisting of the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or the sequence TYKELIERIPELNK (SEQ ID NO: 715); (p) Epitopes consisting of the sequence SVIIVDKNGRLVYLVENPGGYVAY (SEQ ID NO: 716); (q) Epitopes consisting of the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or the sequence FGNMEGDCPSDWKTDSTCRM (SEQ ID NO: 705). Epitopes consisting of (717); and (r) epitopes consisting of the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes consisting of the sequence LVENPGYVAYSKAATVTGKL (SEQ ID NO: 719) and / or epitopes consisting of the sequence IYMDQTKFPIVNAELSF (SEQ ID NO: 720) and / or epitopes consisting of the sequence ISRAAAEKL (SEQ ID NO: 721) and / or epitopes consisting of the sequence VTSESKNVKLTVSNVLKE (SEQ ID NO: 722) and / or epitopes consisting of the sequence FCEDTDYPYLGTTMDT (SEQ ID NO: 723).
[0019] In some such compositions, the TfR-binding delivery domain comprises an anti-TfR antibody, an antibody fragment, or a single-chain variable fragment (scFv). In some such compositions, the TfR-binding delivery domain is a single-chain variable fragment (scFv), optionally wherein the multi-domain therapeutic protein comprises domains arranged in the following orientation: N'-heavy chain variable region-light chain variable region-lysosomal α-glucosidase polypeptide-C' or N'-light chain variable region-heavy chain variable region-lysosomal α-glucosidase polypeptide-C', optionally wherein the scFv and the lysosomal α-glucosidase polypeptide are linked by a peptide linker, and optionally wherein the peptide linker is -(GGGGS). m - (SEQ ID NO: 537); where m is 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, optionally wherein the scFv variable region is linked by a peptide linker, and optionally wherein the peptide linker is -(GGGGS). m- (SEQ ID NO:537); where m is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some such compositions, the multi-domain therapeutic protein includes a heavy chain variable region (V... H ) and light chain variable region (V L ), and lysosomal α-glucosidase polypeptide, of which V H V L The polypeptide arrangement of lysosomal α-glucosidase is as follows: (i) V L -V H - Lysosomal α-glucosidase polypeptide; (ii) V H -V L - Lysosomal α-glucosidase polypeptide; (iii) V L -[(GGGGS)3(SEQ ID NO: 616)]-V H -[(GGGGS)2(SEQ ID NO: 617)]-lysosomal α-glucosidase polypeptide; or (iv) V H -[(GGGGS)3(SEQ ID NO: 616)]-V L -[(GGGGS)2(SEQID NO: 617)]-lysosomal α-glucosidase polypeptide.
[0020] In some such compositions, the scFv comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 508. In some such compositions, the scFv encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 532 and encodes the scFv comprising SEQ ID NO: 508. In some such compositions, the scFv encoding sequence comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 532. In some such compositions, the multi-domain therapeutic protein comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 746. In some such compositions, the multi-domain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 745, optionally wherein the nucleotide at position 1857 is G, the nucleotide at position 1860 is C, and the nucleotide at position 3105 is G. In some such compositions, the multi-domain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 745, and encodes a multi-domain therapeutic protein comprising SEQ ID NO: 746, optionally wherein the nucleotide at position 1857 is G, the nucleotide at position 1860 is C, and the nucleotide at position 3105 is G. In some such compositions, the multidomain therapeutic protein coding sequence comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 745.
[0021] In some such compositions, the nucleic acid construct from 5' to 3' comprises: a splice acceptor, a coding sequence for a multi-domain therapeutic protein, and a polyadenylation signal or sequence, wherein the coding sequence for the multi-domain therapeutic protein comprises SEQ ID NO: 745, optionally wherein the nucleic acid construct comprises the sequence shown in SEQ ID NO: 780, optionally wherein the nucleic acid construct comprises the sequence shown in SEQ ID NO: 764, wherein the polyadenylation signal comprises a BGH polyadenylation signal and a unidirectional SV40 late polyadenylation signal, optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO: 751, and the unidirectional SV40 late polyadenylation signal comprises the sequence shown in SEQ ID NO: 752, optionally the polyadenylation signal comprising the BGH polyadenylation signal and the unidirectional SV40 late polyadenylation signal comprises the sequence shown in SEQ ID NO: 795, wherein the nucleic acid construct does not contain a promoter driving the expression of the multi-domain therapeutic protein, and wherein the nucleic acid construct does not contain a homologous arm. In some such compositions, the nucleic acid construct from 5' to 3' comprises: a splice acceptor, a coding sequence for a multi-domain therapeutic protein, and a polyadenylation signal or sequence, wherein the coding sequence for the multi-domain therapeutic protein comprises SEQ ID NO: 745, optionally wherein the nucleic acid construct comprises the sequence shown in SEQ ID NO: 781, optionally wherein the nucleic acid construct comprises the sequence shown in SEQ ID NO: 765, wherein the polyadenylation signal comprises a BGH polyadenylation signal, optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO: 751, wherein the nucleic acid construct does not contain a promoter driving the expression of the multi-domain therapeutic protein, and wherein the nucleic acid construct does not contain a homologous arm.
[0022] In some such compositions, the delivery domain is a CD63-binding delivery domain. In some such compositions, the CD63-binding delivery domain comprises an anti-CD63 antigen-binding protein. In some such compositions, the CD63-binding delivery domain comprises an anti-CD63 antibody, an antibody fragment, or a single-chain variable fragment (scFv). In some such compositions, the CD63-binding delivery domain is a single-chain variable fragment (scFv). In some such compositions, the scFv comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 730. In some such compositions, the scFv encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 759, optionally wherein the nucleotide at position 3 is A, the nucleotide at position 132 is A, and the nucleotide at position 273 is T. In some such compositions, the scFv coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 759, and encodes the scFv comprising SEQ ID NO: 730, optionally wherein the nucleotide at position 3 is A, the nucleotide at position 132 is A, and the nucleotide at position 273 is T. In some such compositions, the scFv coding sequence comprises, substantially comprises, or comprises the sequence shown in SEQ ID NO: 759. In some such compositions, the scFv coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 760, optionally wherein the nucleotide at position 273 is T. In some such compositions, the scFv coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 760, and encodes the scFv comprising SEQ ID NO: 730, optionally wherein the nucleotide at position 273 is T. In some such compositions, the scFv coding sequence comprises, substantially comprises, or comprises the sequence shown in SEQ ID NO: 760. In some such compositions, the scFv coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 732.In some such compositions, the scFv encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 732 and encodes the scFv comprising SEQ ID NO: 730. In some such compositions, the scFv encoding sequence comprises, substantially comprises, or comprises of the sequence shown in SEQ ID NO: 732.
[0023] In some such compositions, the multidomain therapeutic protein comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 733. In some such compositions, the multidomain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to that in SEQ ID NO: 756, optionally wherein the nucleotide at position 3 is A, the nucleotide at position 132 is A, the nucleotide at position 273 is T, the nucleotide at position 723 is G, the nucleotide at position 1830 is G, the nucleotide at position 1833 is C, and the nucleotide at position 3078 is G. In some such compositions, the multi-domain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 756, and encodes a multi-domain therapeutic protein comprising SEQ ID NO: 733, optionally wherein the nucleotide at position 3 is A, the nucleotide at position 132 is A, the nucleotide at position 273 is T, the nucleotide at position 723 is G, the nucleotide at position 1830 is G, the nucleotide at position 1833 is C, and the nucleotide at position 3078 is G. In some such compositions, the multi-domain therapeutic protein encoding sequence comprises, substantially comprises, or comprises of the sequence shown in SEQ ID NO: 756. In some such compositions, the multidomain therapeutic protein coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 757, optionally wherein the nucleotide at position 273 is T, the nucleotide at position 723 is G, the nucleotide at position 1830 is G, the nucleotide at position 1833 is C, and the nucleotide at position 3078 is G. In some such compositions, the multi-domain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 757, and encodes a multi-domain therapeutic protein comprising SEQ ID NO: 733, optionally wherein the nucleotide at position 273 is T, the nucleotide at position 723 is G, the nucleotide at position 1830 is G, the nucleotide at position 1833 is C, and the nucleotide at position 3078 is G. In some such compositions, the multi-domain therapeutic protein encoding sequence comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 757.In some such compositions, the multidomain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 758, optionally wherein the nucleotide at position 3078 is G. In some such compositions, the multidomain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 758, and encodes a multidomain therapeutic protein comprising SEQ ID NO: 733, optionally wherein the nucleotide at position 3078 is G. In some such compositions, the multidomain therapeutic protein encoding sequence comprises, substantially comprises, or comprises the sequence shown in SEQ ID NO: 758.
[0024] In some such compositions, the nucleic acid construct from 5' to 3' comprises: a splice acceptor, a coding sequence for a multi-domain therapeutic protein, and a polyadenylation signal or sequence, wherein the coding sequence for the multi-domain therapeutic protein comprises SEQ ID NO: 756, optionally wherein the nucleic acid construct comprises the sequence shown in SEQ ID NO: 793, optionally wherein the nucleic acid construct comprises the sequence shown in SEQ ID NO: 777, wherein the polyadenylation signal comprises a BGH polyadenylation signal and a unidirectional SV40 late polyadenylation signal, optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO: 751, and the unidirectional SV40 late polyadenylation signal comprises the sequence shown in SEQ ID NO: 752, optionally the polyadenylation signal comprising the BGH polyadenylation signal and the unidirectional SV40 late polyadenylation signal comprises the sequence shown in SEQ ID NO: 795, wherein the nucleic acid construct does not contain a promoter driving the expression of the multi-domain therapeutic protein, and wherein the nucleic acid construct does not contain a homologous arm. In some such compositions, the nucleic acid construct from 5' to 3' comprises: a splice acceptor, a coding sequence for a multi-domain therapeutic protein, and a polyadenylation signal or sequence, wherein the coding sequence for the multi-domain therapeutic protein comprises SEQ ID NO: 756, optionally wherein the nucleic acid construct comprises the sequence shown in SEQ ID NO: 794, optionally wherein the nucleic acid construct comprises the sequence shown in SEQ ID NO: 778, wherein the polyadenylation signal comprises a BGH polyadenylation signal, optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO: 751, wherein the nucleic acid construct does not contain a promoter driving the expression of the multi-domain therapeutic protein, and wherein the nucleic acid construct does not contain a homologous arm.
[0025] In some of these compositions, the nucleic acid construct is contained within a nucleic acid carrier or lipid nanoparticles. In some of these compositions, the nucleic acid construct is contained within a nucleic acid carrier, optionally wherein the nucleic acid carrier is a viral vector. In some of these compositions, the nucleic acid carrier is an adeno-associated virus (AAV) vector, optionally wherein the nucleic acid construct has an inverted terminal repeat (ITR) sequence attached to each end, optionally wherein at least one end of the ITR contains, substantially constitutes, or is composed of SEQ ID NO: 160, and optionally wherein each end of the ITR contains, substantially constitutes, or is composed of SEQ ID NO: 160. In some of these compositions, the AAV vector is a single-stranded AAV (ssAAV) vector. In some of these compositions, the AAV vector is a recombinant AAV8 (rAAV8) vector, optionally wherein the AAV vector is a single-stranded rAAV8 vector.
[0026] In some of these compositions, the composition is combined with a nuclease preparation that targets a nuclease target site in a target genomic locus. In some of these compositions, the target genomic locus is an albumin gene, optionally wherein the albumin gene is a human albumin gene. In some of these compositions, the nuclease target site is located in intron 1 of the albumin gene. In some of these compositions, the nuclease preparation comprises: (a) a zinc finger nuclease (ZFN); (b) a transcription activator-like effector nuclease (TALEN); or (c) (i) a Cas protein or a nucleic acid encoding a Cas protein; and (ii) a guide RNA or one or more DNA encoding a guide RNA, wherein the guide RNA contains a DNA targeting segment that targets a guide RNA target sequence, and wherein the guide RNA binds to the Cas protein and targets the Cas protein to the guide RNA target sequence. In some of these compositions, the nuclease preparation comprises: (a) a Cas protein or a nucleic acid encoding a Cas protein; and (b) a guide RNA or one or more DNA encoding a guide RNA, wherein the guide RNA contains a DNA targeting segment that targets a guide RNA target sequence, and wherein the guide RNA binds to the Cas protein and targets the Cas protein to the guide RNA target sequence.
[0027] In some of these compositions, the guide RNA target sequence is located in intron 1 of the albumin gene. In some of these compositions, the DNA targeting segment comprises any one of SEQ ID NO: 30-61, optionally comprising any one of SEQ ID NO: 36, 30, 33, and 41, or wherein the DNA targeting segment consists of any one of SEQ ID NO: 30-61, optionally comprising any one of SEQ ID NO: 36, 30, 33, and 41. In some of these compositions, the guide RNA comprises any one of SEQ ID NO: 62-125, optionally comprising any one of SEQ ID NO: 68, 100, 62, 94, 65, 97, 73, and 105. In some of these compositions, the DNA targeting segment comprises or consists of SEQ ID NO: 36. In some of these compositions, the guide RNA comprises SEQ ID NO: 68 or 100. In some such compositions, the composition comprises guide RNA in the form of RNA. In some such compositions, the guide RNA comprises at least one modification. In some such compositions, the at least one modification comprises: (i) a phosphate thioester bond between the first four nucleotides at the 5' end of the guide RNA; (ii) a phosphate thioester bond between the last four nucleotides at the 3' end of the guide RNA; (iii) a 2'-O-methyl modified nucleotide at the first three nucleotides at the 5' end of the guide RNA; and (iv) a 2'-O-methyl modified nucleotide at the last three nucleotides at the 3' end of the guide RNA. In some of these compositions, the composition comprises a guide RNA in the form of SEQ ID NO: 100, and the guide RNA comprises: (i) a phosphate thioester bond between the first four nucleotides at the 5' end of the guide RNA; (ii) a phosphate thioester bond between the last four nucleotides at the 3' end of the guide RNA; (iii) a 2'-O-methyl modified nucleotide at the first three nucleotides at the 5' end of the guide RNA; and (iv) a 2'-O-methyl modified nucleotide at the last three nucleotides at the 3' end of the guide RNA.
[0028] In some such compositions, the Cas protein is the Cas9 protein, optionally wherein the Cas protein is derived from Streptococcus pyogenes (Streptococcus pyogenes). Streptococcus pyogenesCas9 protein. In some such compositions, the Cas protein comprises the sequence shown in SEQ ID NO:11. In some such compositions, the composition comprises a nucleic acid encoding the Cas protein, wherein the nucleic acid comprises mRNA encoding the Cas protein. In some such compositions, the mRNA encoding the Cas protein comprises at least one modification. In some such compositions, the mRNA encoding the Cas protein is completely substituted with N1-methyl-pseuuridine. In some such compositions, the mRNA encoding the Cas protein comprises the sequence shown in SEQ ID NO:1 or 2. In some such compositions, the composition comprises a nucleic acid encoding the Cas protein, wherein the nucleic acid comprises mRNA encoding the Cas protein, the mRNA encoding the Cas protein comprises the sequence shown in SEQ ID NO:1 or 2, and the mRNA encoding the Cas protein is completely substituted with N1-methyl-pseuuridine, comprises a 5' cap, and comprises a poly(A) tail.
[0029] In some of these compositions, the composition comprises a guide RNA in the form of RNA, and the guide RNA comprises SEQ ID NO: 68 or 100, and the composition comprises an administration of a nucleic acid encoding a Cas protein, wherein the nucleic acid comprises mRNA encoding a Cas protein, and the mRNA encoding a Cas protein comprises the sequence shown in SEQ ID NO: 1 or 2. In some of these compositions, the composition comprises a guide RNA in the form of RNA, the guide RNA comprising SEQ ID NO: 100, and the guide RNA comprising: (i) a phosphate thioester bond between the first four nucleotides at the 5' end of the guide RNA; (ii) a phosphate thioester bond between the last four nucleotides at the 3' end of the guide RNA; (iii) a 2'-O-methyl modified nucleotide at the first three nucleotides at the 5' end of the guide RNA; and (iv) a 2'-O-methyl modified nucleotide at the last three nucleotides at the 3' end of the guide RNA, and wherein the composition comprises a nucleic acid encoding a Cas protein, wherein the nucleic acid comprises mRNA encoding a Cas protein, the mRNA encoding a Cas protein comprising the sequence shown in SEQ ID NO: 1 or 2, and the mRNA encoding a Cas protein being completely substituted with N1-methyl-pseuuridine, comprising a 5' cap, and comprising a poly(A) tail.
[0030] In some such compositions, the Cas protein or nucleic acid encoding the Cas protein and guide RNA, or one or more DNA encoding guide RNA, are associated with lipid nanoparticles. In some such compositions, the lipid nanoparticles comprise cationic lipids, neutral lipids, accessory lipids, and octanoic lipids. In some such compositions, the cationic lipid is lipid A ((9Z,12Z)-3-((4,4-bis(octyloxy)butyryl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyloctadec-9,12-dienoate), and / or wherein the neutral lipid is distearylphosphatidylcholine or 1,2-distearyl-sn-glycerol-3-phosphocholine (DSPC), and / or wherein the accessory lipid is cholesterol, and / or wherein the octanoic lipid is 1,2-dimyristoyl-racemic-glycerol-3-methoxypolyethylene glycol-2000. In some of these compositions, the cationic lipid is lipid A, the neutral lipid is DSPC, the accessory lipid is cholesterol, and the occult lipid is PEG2k-DMG. In some of these compositions, the lipid nanoparticles comprise four lipids in the following molar ratios: approximately 50 mol% lipid A, approximately 9 mol% DSPC, approximately 38 mol% cholesterol, and approximately 3 mol% PEG2k-DMG.
[0031] In another aspect, cells comprising any of the above compositions are provided. In some such cells, the coding sequence of a nucleic acid construct or a multi-domain therapeutic protein is integrated into a target genomic locus, wherein the multi-domain therapeutic protein is expressed by that target genomic locus, or the coding sequence of a nucleic acid construct or a multi-domain therapeutic protein is integrated into intron 1 of an endogenous albumin locus, wherein the multi-domain therapeutic protein is expressed by that endogenous albumin locus. In some such cells, the percentage of unintended transcripts from the target genomic locus containing the integrated coding sequence of the nucleic acid construct or a multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%. In some such cells, the cells are liver cells or hepatocytes. In some such cells, the cells are human cells.
[0032] On another front, a method is provided for inserting a nucleic acid encoding a multidomain therapeutic protein comprising a delivery domain fused with lysosomal α-glucosidase into a target genomic locus in a cell or cell population. This method includes administering any of the above-described compositions to the cell or cell population, wherein a nuclease preparation cleaves a nuclease target site in the target genomic locus, and the nucleic acid construct or the nucleic acid encoding the multidomain therapeutic protein is inserted into the target genomic locus. In some such methods, the percentage of unintended transcripts from the target genomic locus containing the inserted nucleic acid construct or the nucleic acid encoding the multidomain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%. On another front, a method is provided for expressing a multidomain therapeutic protein comprising a delivery domain fused with lysosomal α-glucosidase in a cell or cell population, comprising administering any of the above-described compositions to the cell or cell population, wherein the coding sequence of the multidomain therapeutic protein is operatively linked to a promoter in the nucleic acid construct and expressed in the cell or cell population. On the other hand, a method is provided for expressing a multi-domain therapeutic protein comprising a delivery domain fused to a lysosomal α-glucosidase from a target genomic locus in cells or a cell population. This method includes administering any of the above-described compositions to cells or a cell population, wherein a nuclease preparation cleaves a nuclease target site in the target genomic locus, a coding sequence of a nucleic acid construct or multi-domain therapeutic protein is inserted into the target genomic locus to produce a modified target genomic locus, and the multi-domain therapeutic protein comprising a delivery domain fused to a lysosomal α-glucosidase is expressed from the modified target genomic locus. In some of these methods, the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the inserted nucleic acid construct or multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
[0033] In some of these methods, the cells are liver cells or hepatocytes, or the cell population is liver cells or a population of hepatocytes. In some of these methods, the cells are human cells, or the cell population is a population of human cells. In some of these methods, the cells are neonatal cells, or the cell population is a population of neonatal cells. In some of these methods, the neonatal cells or neonatal cell population are derived from human neonatal subjects within 24 weeks of birth, optionally from human neonatal subjects within 12 weeks of birth, optionally from human neonatal subjects within 8 weeks of birth, and optionally from human neonatal subjects within 4 weeks of birth. In some of these methods, the cells are not neonatal cells, or the cell population is not a neonatal cell population. In some of these methods, the cells are in vitro or ex vivo, or the cell population is in vitro or ex vivo. In some of these methods, the cells are in vivo within the subject, or the cell population is in vivo within the subject.
[0034] On another front, a method is provided for inserting a nucleic acid encoding a multidomain therapeutic protein comprising a delivery domain fused with a lysosomal α-glucosidase into a target genomic locus in a subject's cells. This method includes administering any of the above-described compositions to the subject, wherein a nuclease preparation cleaves a nuclease target site in the target genomic locus, and the nucleic acid construct or the nucleic acid encoding the multidomain therapeutic protein is inserted into the target genomic locus. In some such methods, the percentage of unintended transcripts from the target genomic locus containing the coding sequence of the inserted nucleic acid construct or the multidomain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%. On another front, a method is provided for expressing a multidomain therapeutic protein comprising a delivery domain fused with a lysosomal α-glucosidase protein in a subject's cells. This method includes administering any of the above-described compositions to the subject, wherein the coding sequence of the multidomain therapeutic protein is operatively linked to a promoter in the nucleic acid construct and expressed in the cells. On the other hand, a method is provided for expressing a multi-domain therapeutic protein comprising a delivery domain fused with lysosomal α-glucosidase from a target genomic locus in the cells of a subject. This method includes administering any of the above-described compositions to a subject, wherein a nuclease preparation cleaves a nuclease target site in the target genomic locus, a coding sequence of a nucleic acid construct or multi-domain therapeutic protein is inserted into the target genomic locus to produce a modified target genomic locus, and the multi-domain therapeutic protein comprising a delivery domain fused with lysosomal α-glucosidase is expressed from the modified target genomic locus. In some such methods, the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the inserted nucleic acid construct or multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
[0035] In some of these methods, the expressed multidomain therapeutic protein is delivered to and internalized in the skeletal muscle and cardiac tissues of the subject, or the expressed multidomain therapeutic protein is delivered to and internalized in the skeletal muscle, cardiac, and central nervous system tissues of the subject. In some of these methods, the cells are liver cells or hepatocytes. In some of these methods, the cells are human cells. In some of these methods, the cells are neonatal cells. In some of these methods, the neonatal subject is a human subject within 24 weeks of birth, optionally within 12 weeks of birth, optionally within 8 weeks of birth, and optionally within 4 weeks of birth. In some of these methods, the cells are not neonatal cells.
[0036] In another aspect, a method for treating lysosomal α-glucosidase deficiency in a subject in need is provided, the method comprising administering any of the above-described compositions to the subject, wherein the coding sequence of a multi-domain therapeutic protein is operatively linked to a promoter in a nucleic acid construct and expressed in the subject. In another aspect, a method for treating lysosomal α-glucosidase deficiency in a subject in need is provided, the method comprising administering any of the above-described compositions to the subject, wherein a nuclease preparation cleaves a nuclease target site in a target genomic locus, the coding sequence of a nucleic acid construct or a multi-domain therapeutic protein is inserted into the target genomic locus to produce a modified target genomic locus, and a multi-domain therapeutic protein comprising a delivery domain fused to lysosomal α-glucosidase is expressed from the modified target genomic locus. In some of these methods, the percentage of unintended transcripts from the target genomic locus containing the coding sequence of the inserted nucleic acid construct or multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%. In another aspect, a method for reducing glycogen accumulation in the tissues of a subject in need is provided, the method comprising administering any of the above-described compositions to the subject, wherein the coding sequence of a multi-domain therapeutic protein is operatively linked to a promoter in a nucleic acid construct and expressed in the subject to reduce glycogen accumulation in the tissues. In another aspect, a method for reducing glycogen accumulation in the tissues of a subject in need is provided, the method comprising administering any of the above-described compositions to the subject, wherein a nuclease preparation cleaves a nuclease target site in a target genomic locus, the coding sequence of a nucleic acid construct or a multi-domain therapeutic protein is inserted into the target genomic locus to produce a modified target genomic locus, and a multi-domain therapeutic protein comprising a delivery domain fused to lysosomal α-glucosidase is expressed from the modified target genomic locus to reduce glycogen accumulation in the tissues. In some of these methods, the percentage of unintended transcripts from target genomic loci containing the coding sequence of an inserted nucleic acid construct or a multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%. In some of these methods, the subject has Pompe disease. In another aspect, a method for treating Pompe disease in a subject in need is provided, comprising administering to the subject any of the above-described compositions, wherein the coding sequence of the multi-domain therapeutic protein is operatively linked to a promoter in the nucleic acid construct and expressed in the subject, thereby treating Pompe disease.On the other hand, a method for treating Pompe disease in a subject in need is provided, comprising administering any of the above-described compositions to the subject, wherein a nuclease preparation cleaves a nuclease target site in a target genomic locus, a coding sequence of a nucleic acid construct or a multi-domain therapeutic protein is inserted into the target genomic locus to produce a modified target genomic locus, and a multi-domain therapeutic protein comprising a delivery domain fused with a lysosomal α-glucosidase is expressed from the modified target genomic locus, thereby treating Pompe disease. In some of these methods, the percentage of unintended transcripts from the target genomic locus containing the coding sequence of the inserted nucleic acid construct or multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
[0037] In some of these methods, Pompe disease is infantile Pompe disease. In some of these methods, Pompe disease is late-onset Pompe disease. In some of these methods, the subject is a human subject. In some of these methods, the subject is a neonatal subject, optionally wherein the neonatal subject is a human subject within 24 weeks, 12 weeks, 8 weeks, or 4 weeks of birth. In some of these methods, the subject is not a neonatal subject. In some of these methods, the method results in the production of therapeutically effective levels of circulating multidomain therapeutic proteins or lysosomal α-glucosidase in the subject. In some such methods, the method reduces glycogen accumulation in the skeletal muscle, cardiac tissue, or central nervous system tissue of a subject, optionally reducing glycogen accumulation in the skeletal muscle, cardiac tissue, and central nervous system tissue of a subject, optionally resulting in a reduction in glycogen levels in the skeletal muscle, cardiac, and central nervous system tissues of a subject comparable to age-matched wild-type levels; or wherein the method reduces glycogen accumulation in the skeletal muscle or cardiac tissue of a subject, optionally reducing glycogen accumulation in the skeletal muscle and cardiac tissue of a subject, optionally resulting in a reduction in glycogen levels in the skeletal muscle and cardiac tissues of a subject comparable to age-matched wild-type levels. In some such methods, the method improves the subject's muscle strength or prevents muscle strength loss in a subject compared to a control subject. In some such methods, the method results in a subject having muscle strength comparable to age-matched wild-type levels.
[0038] In another aspect, a method is provided for preventing or reducing the onset of signs or symptoms of Pompe disease in a subject in need, the method comprising administering to the subject any of the above-described compositions, wherein the coding sequence of a multi-domain therapeutic protein is operatively linked to a promoter in a nucleic acid construct and expressed in the subject, thereby preventing or reducing the onset of signs or symptoms of Pompe disease in the subject. In another aspect, a method is provided for preventing or reducing the onset of signs or symptoms of Pompe disease in a subject in need, the method comprising administering to the subject any of the above-described compositions, wherein a nuclease preparation cleaves a nuclease target site, the coding sequence of a nucleic acid construct or a multi-domain therapeutic protein is inserted into a target genomic locus to produce a modified target genomic locus, and a multi-domain therapeutic protein comprising a delivery domain fused with lysosomal α-glucosidase is expressed from the modified target genomic locus, thereby preventing or reducing the onset of signs or symptoms of Pompe disease in the subject. In some of these methods, the percentage of unintended transcripts from target genomic loci containing the coding sequence of an inserted nucleic acid construct or a multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
[0039] In some of these methods, Pompe disease is infantile-onset Pompe disease. In some of these methods, Pompe disease is late-onset Pompe disease. In some of these methods, the method results in the production of therapeutically effective levels of circulating multi-domain therapeutic proteins or lysosomal α-glucosidase in the subject. In some of these methods, the method prevents or reduces glycogen accumulation in the subject's skeletal muscle, cardiac, or central nervous system tissues. In some of these methods, the method prevents or reduces glycogen accumulation in the subject's skeletal muscle, cardiac, and central nervous system tissues, or wherein the method prevents or reduces glycogen accumulation in the subject's skeletal muscle and cardiac tissues. In some of these methods, the subject is a human subject. In some of these methods, the subject is a neonatal subject. In some of these methods, the neonatal subject is a human subject within 24 weeks of birth, optionally within 12 weeks of birth, optionally within 8 weeks of birth, and optionally within 4 weeks of birth. In some of these methods, the subject is not a neonatal subject.
[0040] In some of these methods, compared to methods involving administration of a free expression vector encoding a multi-domain therapeutic protein to a control subject, this method increases the expression of the multi-domain therapeutic protein in the subject. In some of these methods, compared to methods involving administration of a free expression vector encoding a multi-domain therapeutic protein to a control subject, this method increases the serum level of the multi-domain therapeutic protein in the subject. In some of these methods, the serum level of the multi-domain therapeutic protein in the subject is at least about 1 μg / mL, at least about 2 μg / mL, at least about 3 μg / mL, at least about 4 μg / mL, at least about 5 μg / mL, at least about 6 μg / mL, at least about 7 μg / mL, at least about 8 μg / mL, at least about 9 μg / mL, or at least about 10 μg / mL. In some of these methods, the serum level of the multi-domain therapeutic protein in the subject is at least about 2 μg / mL or at least about 5 μg / mL. In some such methods, the method results in serum levels of the multidomain therapeutic protein in the subject between about 2 μg / mL and about 30 μg / mL, or between about 2 μg / mL and about 20 μg / mL. In some such methods, the method results in serum levels of the multidomain therapeutic protein in the subject between about 5 μg / mL and about 30 μg / mL, or between about 5 μg / mL and about 20 μg / mL. In some such methods, the method achieves lysosomal α-glucosidase activity levels of at least about 40%, at least about 45%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or 100% of the normal value. In some of these methods, (I) the subject has infancy-onset Pompe disease, and the method achieves at least about 1% or more of the normal value of lysosomal α-glucosidase expression or activity; or (II) the subject has late-onset Pompe disease, and the method achieves at least about 40% or more of the normal value of lysosomal α-glucosidase expression or activity. In some of these methods, the expression or activity of the multidomain therapeutic protein is at least 50% of the peak expression level measured against the subject at 24 weeks post-administration. In some of these methods, the expression or activity of the multidomain therapeutic protein is at least 50% of the peak expression level measured against the subject at one year post-administration. In some of these methods, the expression or activity of the multidomain therapeutic protein is at least 60% of the peak expression level measured against the subject at 24 weeks post-administration. In some of these methods, the expression or activity of the multidomain therapeutic protein is at least 50% of the peak expression level of the multidomain therapeutic protein measured against the subject two years after administration.In some of these methods, the expression or activity of the multidomain therapeutic protein is at least 60% of the peak expression level measured against the subject at 2 years post-administration. In some of these methods, the expression or activity of the multidomain therapeutic protein is at least 60% of the peak expression level measured against the subject at 24 weeks post-administration.
[0041] In some of these methods, the method further includes assessing the subject's pre-existing AAV immunity prior to administration of the nucleic acid construct. In some of these methods, the pre-existing AAV immunity is pre-existing AAV8 immunity. In some of these methods, assessing pre-existing AAV immunity includes evaluating immunogenicity using a total antibody immunoassay or a neutralizing antibody assay.
[0042] In some of these methods, the nucleic acid construct is administered simultaneously with a nuclease preparation or one or more nucleic acids encoding that nuclease preparation. In some of these methods, the nucleic acid construct is not administered simultaneously with a nuclease preparation or one or more nucleic acids encoding that nuclease preparation. In some of these methods, the nucleic acid construct is administered before the nuclease preparation or one or more nucleic acids encoding that nuclease preparation. In some of these methods, the nucleic acid construct is administered after the nuclease preparation or one or more nucleic acids encoding that nuclease preparation. Attached Figure Description
[0043] Figure 1 V was displayed k -3xG4S(SEQ ID NO: 616)-V H Amino acid sequences of various forms of anti-human transferrin receptor scFv molecules.
[0044] Figures 2A-2C The anti-human TFRC scFv antibody clone was shown to deliver GAA to Tfrc. hum Mouse brains. Anti-human TfR:GAA molecules 69261, 69329, 12839, 12841, 12843, and 12845 were tested. Figure 2A ), 69348, 12795, 12799, 12801, 12850 and 12798 ( Figure 2B ), and 12802, 69340, 12847, 12848, 69307 and 69323 ( Figure 2C Each lane = 1 mouse. Delivered via HDD.
[0045] Figure 3This image shows a subset of anti-hTFRC antibodies (12798, 12850, 69323, 12841, 12843, 12845, 12847, 12848, 12799, 69307, and 12839) delivering mature GAA to the brain parenchyma in scfv:GAA form (via HDD delivery). Lane E corresponds to the endothelium, while lane P corresponds to the parenchyma. The affinity ratio of mfTfR to human TfR is shown below the image (mf refers to cynomolgus monkey). Macaca fascicularis )).
[0046] Figure 4 The study demonstrated that anti-hTFRC antibodies (12799, 12843, 12847, and 12839) delivered mature GAA in scfv:GAA form to the brain parenchyma (AAV8 free liver reservoir gene therapy). Lane E corresponds to the endothelium, while lane P corresponds to the parenchyma.
[0047] Figure 5 The study demonstrated that the free AAV8 liver reservoir anti-hTFRC scfv:GAA antibody delivered GAA protein to Gaa. - / - / Tfrc hum The mouse's CNS (cerebellum, cerebrum, spinal cord), heart, and muscles (quadriceps).
[0048] Figure 6 The free AAV8 liver reservoir showed that anti-hTFRC scfv:GAA antibodies (12839, 12843, and 12847) corrected for Gaa - / - / Tfrc hum Glycogen storage in the central nervous system (CNS) (cerebellum, cerebrum, spinal cord), heart, and muscles (quadriceps) of mice.
[0049] Figures 7A-7D The free AAV8 liver reservoir showed that anti-hTFRC scfv:GAA antibodies (12847, 12843, and 12799) corrected for Gaa - / - / Tfrc hum Mouse brain (thalamus) Figure 7A ), cerebral cortex ( Figure 7B ), CA1 area of the equina ( Figure 7C )) and muscles (quadriceps ( Figure 7D Glycogen storage in ))
[0050] Figure 8 The insertion of anti-hTFRC 12847scfv:GAA demonstrates that it delivers mature GAA protein into the CNS and muscle of a Pompe disease model mouse.
[0051] Figure 9The insertion of anti-hTFRC 12847scfv:GAA showed that it corrected glycogen storage in the CNS and muscle of Pompe disease model mice. Univariate ANOVA (* p < 0.01; ** p < 0.001; *** p < 0.0001).
[0052] Figure 10 This study shows serum GAA activity following Cas9-mediated AAV-delivered insertion of anti-TfR1:GAA or anti-CD63:GAA into the cynomolgus macaque albumin locus. The mediator alone served as a negative control. One unit of GAA activity was defined as the amount of enzyme producing 1.0 µmol 4-MU per minute at pH 4.5 and 37 °C. Error bars are shown in SEM images. For the mediator, N=1; for all other conditions, N=2–4.
[0053] Figure 11 The study demonstrates how albumin insertion of anti-hTFRC 12847scfv:GAA delivers mature GAA protein into the CNS and muscle of cynomolgus monkeys. For the bar chart, mature GAA was quantified by Western blot analysis of tissue lysates; the error bar is SD.
[0054] Figure 12 Machuposa virus (MDV) is shown on two TfR molecules that are co-superimposed on a symmetric unit. Mammarenavirus machupoense GP1 protein (PDB 3KAS), human ferritin (PDB 6GSR), Plasmodium vivax ( Plasmodium vivax ) Sal-1 Interactions of PvRBP2b protein (PDB 6D04), human HFE protein (PDB 1DE4), and human transferrin (PDB 1SUV) molecules. For Machuposa virus GP1 protein and human ferritin, only one copy of the symmetric unit is shown for clarity to reduce the complexity of the diagram.
[0055] Figure 13 Hydrogen-deuterium exchange mass spectrometry (HDX) was used to depict the protection of the antibody, which was attributed to five regions in the TfR (PDB 1SUV) as tested in HDX-MS experiments.
[0056] Figure 14 The TfR regions protected by REGN17513 are shown. REGN17513 represents an antibody that induces HDX protection in the TfR apical domain. These protective regions overlap with the binding sites of Machuposa GP1 protein, human ferritin, and Plasmodium vivax PvRBP2b protein.
[0057] Figure 15 The TfR region protected by REGN17510 is shown, where REGN17510 represents a TfR apex structural domain that is not associated with... Figure 15 Other TfR-binding antibodies with shared HDX protection are shown.
[0058] Figure 16 The TfR region protected by REGN17515 is shown, where REGN17515 represents the presence of human ferritin and Plasmodium vivax in the TfR apical domain. Sal-1 An antibody protected by HDX with a common binding site for the PvRBP2b protein.
[0059] Figure 17 The TfR region protected by REGN17514 is shown. REGN17514 represents HDX protection within the TfR-like protease domain and its interaction with Plasmodium vivax. Sal-1 Antibodies that share a common binding site with the PvRBP2b protein.
[0060] Figure 18 The TfR region protected by REGN17508 is shown, where REGN17508 represents an antibody with HDX protection in the TfR protease domain. This region was not... Figure 18 Other TfR interacting molecules are shown.
[0061] Figure 19A and Figure 19B The study showed the GAA enzyme activity in culture medium after various anti-TfR:GAA insert templates (CpG depleted and native) were inserted into the albumin locus of primary human hepatocytes following delivery via rAAV2.
[0062] Figure 20A Western blot analysis showed anti-human TfR antibody clones (0 CpG and native), and GAA delivery to 3-month-old infants via intravenous administration of LNP-g666 (3 mg / kg) and various recombinant AAV8 anti-TfR:GAA or AAV8 anti-CD63:GAA insert templates. Gaa - / - / Tfrc hum mice or Gaa - / - / CD63 hum The mouse's brain. Each lane = 1 mouse.
[0063] Figure 20B The results showed that intravenous administration of LNP-g666 (3 mg / kg) and various recombinant AAV8 anti-TfR:GAA or AAV8 anti-CD63:GAA insert templates were effective. Gaa - / - / Tfrc hum mice or Gaa - / - / CD63 hum In mice, albumin insertion of anti-hTfR:GAA corrected glycogen storage in the brain, quadriceps femoris, diaphragm, and heart. Glycogen levels were measured 3 weeks after administration. Wt Untreated mice served as a positive control, and Gaa - / - Untreated mice served as a negative control.
[0064] Figure 21 The results showed that administering LNP-g666 (1 mg / kg) and recombinant AAV8 anti-CD63:GAA insert template (1.2e13 vg / kg) (“insert”) or free AAV encoding anti-CD63:GAA (4e12 vg / kg) (“free”) to adult male and female mice with Pompe disease model (n=12; GAA) demonstrated the effects of administering LNP-g666 (1 mg / kg) and recombinant AAV8 anti-CD63:GAA insert template (1.2e13 vg / kg) (“insert”) or free AAV encoding anti-CD63:GAA (4e12 vg / kg) (“free”) - / - CD63 hu / hu The serum anti-CD63:GAA level during the subsequent 10-month process.
[0065] Figure 22 The results showed that in adult male and female mice with Pompe disease (n=12, GAA) - / - CD63 hu / hu Ten months after administration of LNP-g666 and recombinant AAV8 anti-CD63:GAA inserted template, or ten months after administration of free AAV encoding anti-CD63:GAA, Pompe disease model mice (GAA) - / - CD63 hu / hu The levels of glycogen in the heart, quadriceps muscle, diaphragm, and spinal cord of wild-type GAA mice were measured. + / + CD63 hu / hu Mice (n=4) and untreated Pompe disease model mice (n=4) were used as controls. The horizontal dashed line represents the lower limit of detection for the assay.
[0066] Figures 23A to 23B The application of LNP-g666 and recombinant AAV8 anti-CD63:GAA inserted templates (n=10; males and females; "insertions") or free AAV encoding anti-CD63:GAA (n=6; males and females; "free") to neonatal (P1) Pompe disease model mice (GAA) is shown. - / - CD63 hu / hu The level of anti-CD63:GAA in serum over a subsequent 15-month period. The dashed line represents the lower limit of detection for this assay. Figure 23A The error bar is ±SD, and Figure 23B The error bar is ±SEM.
[0067] Figure 24AThe results show that, 3 months after administration of LNP-g666 and recombinant AAV8 anti-CD63:GAA inserted template (n=5; males and females, "I") to newborn (P1) mice or 3 months after administration of free AAV encoding anti-CD63:GAA (n=3; males and females, "E"), the results in Pompe disease model mice (GAA) - / - CD63 hu / hu Glycogen levels in the heart, quadriceps femoris, gastrocnemius, and diaphragm were measured. Untreated Pompe disease model mice were used as controls.
[0068] Figure 24B The results show that, 15 months after administration of LNP-g666 and recombinant AAV8 anti-CD63:GAA inserted template (n=10; males and females, "I") to newborn (P1) mice or 15 months after administration of free AAV encoding anti-CD63:GAA (n=6; males and females, "E"), the results in Pompe disease model mice (GAA) - / - CD63 hu / hu Glycogen levels in the heart, quadriceps, gastrocnemius, diaphragm, brain, and spinal cord were measured. Untreated Pompe disease model mice (“U”) and wild-type mice (“W”) were used as controls.
[0069] Figure 25 The results show that 15 months after administration of LNP-g666 and recombinant AAV8 anti-CD63:GAA inserted template (n=10; males and females, "P1 inserted AAV+LNP") to newborn (P1) mice or 15 months after administration of free AAV encoding anti-CD63:GAA (n=6; males and females, "P1 free AAV"), the results in Pompe disease model mice (GAA) - / - CD63 hu / hu The gripping force of wild-type GAA mice (GAA + / + CD63 hu / hu Wild-type mice and untreated Pompe disease model mice (untreated KO) were used as controls.
[0070] Figure 26 The IFNα response, as measured by IFNα ELISA, is shown in a primary human plasma cell-like DC-based assay. Various rAAV6 CpG-depleted anti-CD63:GAA templates were tested compared to first-generation (non-CpG-depleted) anti-CD63:GAA templates. rAAV6-GFP was used as a positive control, and CpG-depleted (0 CpG) templates were also tested. F9 The template was used as a negative control.
[0071] Figure 27The study showed the GAA enzyme activity in the culture medium after various anti-CD63:GAA and anti-TfR:GAA insert templates were inserted into the albumin locus of primary human hepatocytes following delivery via rAAV2.
[0072] Figure 28 The study showed the GAA enzyme activity in the culture medium after various anti-CD63:GAA insert templates were inserted into the albumin locus of primary human hepatocytes following delivery via rAAV6.
[0073] Figures 29A to 29B The study showed that after administration of LNP-g666 and various recombinant AAV8 anti-CD63:GAA insert templates, GAA... - / - Serum GAA expression in mice. Untreated KO and untreated WT mice were used as controls.
[0074] Figure 30A This study demonstrates serum GAA activity in cynomolgus monkeys administered recombinant AAV8 containing a CpG-depleted anti-CD63:GAA template and LNP-g9860, using a fluorescent substrate assay. Three different doses of AAV8 (0.3e13vg / kg, 1.5e13vg / kg, and 5.6e13vg / kg) and a dose of LNP at 3 mg / kg were used. N=1 in the mediator control group and N=3 in the administration group.
[0075] Figure 30B The expression of mature GAA in tissue lysates of cynomolgus monkeys administered with recombinant AAV8 containing a CpG-depleted anti-CD63:GAA template and LNP-g9860 was demonstrated. Three different doses of AAV8 (0.3e13vg / kg, 1.5e13vg / kg, and 5.6e13vg / kg) and a dose of LNP at 3 mg / kg were used. N=1 in the vector control group and N=3 in the administration group. Tissues were collected at sacrifice (day 89), and the presence of 76 kDa lysosomal GAA was detected by Western blotting.
[0076] Figure 31 A schematic diagram of LNP-g9860 and recombinant AAV8 (rAAV8) capsid is shown; the former contains a target human albumin (rAAV8). ALB The lipid nanoparticles of Cas9 mRNA and sgRNA 9860 containing intron 1, the latter packaged with an anti-CD63:GAA insert template.
[0077] Figure 32 A schematic diagram of GAA targeting lysosomes via fusion with anti-CD63 scFv is shown.
[0078] Figure 33This demonstrates CRISPR / Cas9-mediated... ALB A schematic diagram of the insertion of the anti-CD63:GAA template at the gene locus. It depicts a human... ALB The loci, where Cas9 cleavage sites are indicated by scissors, depict the cleavage acceptor sites for the CD63:GAA transgene laterally inserted into the template. In endogenous... ALB Following promoter-driven insertion and transcription, in ALB Splicing occurred between exon 1 and the inserted anti-CD63:GAA DNA template, as shown by the dashed line, resulting in a heterozygous strain. ALB -anti CD63: GAA mRNA. The ALB signal peptide promotes the secretion of anti-CD63:GAA and is removed during protein maturation to produce anti-CD63:GAA in plasma.
[0079] Figure 34 A schematic diagram of LNP-g9860 and recombinant AAV8 (rAAV8) capsid is shown; the former contains a target human albumin (rAAV8). ALB The lipid nanoparticles containing Cas9 mRNA and sgRNA 9860 in intron 1, the latter packaged with an anti-TfR:GAA insert template.
[0080] Figure 35 A schematic diagram showing multiple pathways for targeting GAA via fusion with anti-TfR scFv is presented.
[0081] Figure 36 This demonstrates CRISPR / Cas9-mediated... ALB A schematic diagram illustrating the insertion of an anti-TfR:GAA template at the gene locus. It depicts the human... ALB Loci, where Cas9 cleavage sites are indicated by scissors. The cleavage acceptor sites for the TfR:GAA transgene lateralized in the inserted template are depicted. In endogenous... ALB Following promoter-driven insertion and transcription, in ALB Splicing occurs between exon 1 and the inserted anti-TfR:GAADNA template, as shown by the dashed line, resulting in a heterozygous product. ALB -anti TfR:GAA mRNA. The ALB signal peptide promotes the secretion of anti-TfR:GAA and is removed during protein maturation to produce anti-TfR:GAA in plasma.
[0082] Figure 37 Showing the identification in cynomolgus monkeys ALB -anti CD63:GAA Transcript. During gene insertion mediated by construct VVT1254-LNP-g9860, the antibodies supplied by construct VVT1254 will be... CD63:GAA DNA template insertion ALB In intron 1 of the gene. Anti- CD63:GAA The polyadenylated sequence following transgenesis was labeled pA. RNA sequencing analysis was performed to identify the polyadenylated sequence produced in PHH co-incubated with the construct VVT1254-LNP-g9860. ALB -anti CD63:GAA Splicing patterns in fusion transcripts. Expected. ALB -anti CD63:GAA Fusion transcripts have only one splicing event, i.e., from ALB Exon 1 to insertion resistance CD63:GAA The splice acceptor site encoded within the DNA template. Hidden splice donor or acceptor sites leading to unintended transcripts are identified by arrows at the indicated locations. Splicing is performed from nucleotide position 3078 to… ALB Unexpected transcripts formed from exon 2 are shown as an example.
[0083] Figure 38 GAA activity in the supernatant of PXB human hepatocytes treated with an AAV encoding an anti-CD63:GAA gene insertion template of LNP-g9860+ was shown. This AAV had various modifications to the cryptic splice site and polyA sequence.
[0084] Figure 39 The activity of GAA in the supernatant of PXB human hepatocytes treated with an AAV encoding an anti-TfR:GAA gene insertion template of LNP-g9860+ was shown. This AAV had various modifications to the cryptic splice site and polyA sequence.
[0085] Figure 40 The study demonstrated GAA activity in the supernatant of PXB human hepatocytes treated with an AAV encoding an anti-CD63:GAA gene insertion template of LNP-g9860+, which had various modifications to the cryptic splice site and polyA sequence, compared to the original anti-CD63:GAA gene insertion template.
[0086] Figure 41 The study demonstrated GAA activity in the supernatant of PXB human hepatocytes treated with an AAV encoding an anti-TfR:GAA gene insertion template of LNP-g9860+, which had various modifications to the cryptic splice site and polyA sequence, compared to the original anti-TfR:GAA gene insertion template.
[0087] Figure 42 Showing Tfrc hum / hum ;Gaa - / - Experimental setup for template validation of anti-TfR:GAA in mice.
[0088] Figure 43 This demonstrates the quantification of transgenic anti-TfR:GAA DNA in liver nucleotide preparations and the quantification of anti-TfR:GAA mRNA expression in the liver using a standard protocol via Taqman.
[0089] Figure 44 Western blot analysis revealed that anti-human TfR antibody clones with mutations to remove cryptic splicing sites and different polyA sequences delivered GAA to 4-month-old infants. Gaa - / - / Tfrc hum The brain (cerebellum, cerebrum), muscles (quadriceps), liver, and serum of mice were analyzed. LNP-g666 (3 mg / kg) and various recombinant AAV8 anti-TfR:GAA insert templates were administered intravenously to mice. Each lane = 1 mouse.
[0090] Figures 45A-45B The results showed the effects of intravenous administration of LNP-g666 (3 mg / kg) and various recombinant AAV8 anti-TfR:GAA insert templates. Gaa - / - / Tfrc hum In mice, albumin insertion of anti-human TfR antibody clones with mutations to remove cryptic splicing sites and different polyA sequences corrected brain ( Figure 45A ) and muscles ( Figure 45B Glycogen stores in the body. Glycogen levels were measured 3 weeks after application. Wt Untreated mice served as a positive control, and Gaa - / - Untreated mice served as a negative control.
[0091] Figure 46 The experimental setup for anti-TfR:GAA template validation in albumin-humanized mice is shown.
[0092] Figures 47A-47C The protein blot was displayed. Figure 47A ) and the quantitative results of bands in the protein blot ( Figures 47B-47C The study showed the expression of anti-human TfR:GAA in the serum and liver of 3-month-old humanized albumin mice that were intravenously administered LNP-g9860 (3 mg / kg) and various recombinant AAV8 anti-TfR:GAA insert templates (3e12vg / kg). Figure 47A Each lane in the pool corresponds to one mouse.
[0093] Figure 48This demonstrates the quantification of transgenic anti-TfR:GAA DNA in liver nucleotide preparations and the quantification of anti-TfR:GAA mRNA expression in the liver using a standard protocol via Taqman.
[0094] definition The terms “protein,” “polypeptide,” and “peptide,” used interchangeably herein, encompass amino acids of any length in polymeric form, including coding and non-coding amino acids, as well as amino acids that are chemically or biochemically modified or derived. These terms also include modified polymers, such as polypeptides having a modified peptide backbone. The term “domain” refers to any portion of a protein or polypeptide that has a specific function or structure.
[0095] Proteins are considered to have an "N-terminus" and a "C-terminus". The term "N-terminus" refers to the starting point of a protein or polypeptide that terminates at an amino acid with a free amine group (-NH2). The term "C-terminus" refers to the end of the amino acid chain (protein or polypeptide) that terminates at a free carboxyl group (-COOH).
[0096] The terms “nucleic acid” and “polynucleotide”, used interchangeably herein, encompass nucleotides in polymeric form of any length, including ribonucleotides, deoxyribonucleotides, or their analogues or modified forms. These terms include single-stranded, double-stranded, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers containing purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.
[0097] Nucleic acids are considered to have both a "5' end" and a "3' end" because mononucleotides react in a certain way to form oligonucleotides, such that the 5' phosphate of one mononucleotide's pentose ring is attached to the 3' oxygen of its neighbor in one direction via a phosphodiester bond. If the 5' phosphate of an oligonucleotide is not attached to the 3' oxygen of a mononucleotide's pentose ring, then the end of that oligonucleotide is called the "5' end." If the 3' oxygen of an oligonucleotide is not attached to the 5' phosphate of another mononucleotide's pentose ring, then the end of that oligonucleotide is called the "3' end." Even if a nucleic acid sequence is inside a larger oligonucleotide, that nucleic acid sequence can be considered to have both a 5' end and a 3' end. In linear or circular DNA molecules, discrete elements are referred to as being "downstream" or "upstream" of 3' elements, or 5'.
[0098] The term "genome-integrated" refers to nucleic acids that have been introduced into cells so that their nucleotide sequences are integrated into the cell's genome. Nucleic acids can be stably incorporated into the cell's genome using any method.
[0099] The term "viral vector" refers to a recombinant nucleic acid containing at least one viral element and containing elements sufficient or permissible for packaging into a viral vector particle. Vectors and / or particles can be used for the purpose of transferring DNA, RNA, or other nucleic acids into cells in vitro, ex vivo, or in vivo. Various forms of viral vectors are known.
[0100] The term "isolated" in relation to cells, tissues (e.g., liver samples), proteins, and nucleic acids includes cells, tissues (e.g., liver samples), proteins, and nucleic acids that are relatively purified relative to other bacteria, viruses, cells, or other components that may normally be present in situ, up to and including substantially pure preparations of cells, tissues (e.g., liver samples), proteins, and nucleic acids. The term "isolated" also includes cells, tissues (e.g., liver samples), proteins, and nucleic acids that do not have naturally occurring counterparts, have been chemically synthesized, and are therefore substantially uncontaminated by other cells, tissues (e.g., liver samples), proteins, and nucleic acids, or have been isolated or purified from most of their naturally occurring accompanying components (e.g., cellular components) (e.g., other cellular proteins, polynucleotides, or cellular components).
[0101] The term "wildtype" includes entities that have the structure and / or activity found in normal states or conditions (compared to mutated, diseased, altered, etc.). Wildtype genes and polypeptides often exist in many different forms (e.g., alleles).
[0102] The term "endogenous sequence" refers to nucleic acid sequences that naturally exist within cells or animals. For example, in humans... ALB Sequence refers to what naturally exists in humans ALB Natural at the gene locus ALB sequence.
[0103] "Exogenous" molecules or sequences include molecules or sequences that are not normally present in cells in that form. Normal presence includes the presence of molecules in relation to a specific developmental stage of the cell and environmental conditions. For example, exogenous molecules or sequences may include mutant forms of corresponding endogenous sequences within the cell (such as humanized forms of endogenous sequences), or sequences that correspond to endogenous sequences within the cell but are in a different form (i.e., not within chromosomes). In contrast, endogenous molecules or sequences include molecules or sequences that are normally present in that form in a specific cell at a specific developmental stage under specific environmental conditions.
[0104] When used in the context of nucleic acids or proteins, the term "heterologous" indicates that a nucleic acid or protein contains at least two segments that are not naturally found together in the same molecule. For example, when used with respect to nucleic acid segments or protein segments, the term "heterologous" indicates that a nucleic acid or protein contains two or more subsequences that are not found in nature to be in the same relationship (e.g., conjugated together). As an example, a "heterologous" region of a nucleic acid vector is a nucleic acid segment within or attached to another nucleic acid molecule that is not found in nature to associate with other molecules. For example, a heterologous region of a nucleic acid vector may contain a coding sequence flanked by a sequence that is not found in nature to associate with a coding sequence. Similarly, a "heterologous" region of a protein is an amino acid segment within or attached to another peptide molecule (e.g., a fusion protein or a tagged protein) that is not found in nature to associate with other peptide molecules. Similarly, nucleic acids or proteins may contain heterologous tags or heterologous secretion or localization sequences.
[0105] “Codon optimization” (i.e., “codon-optimized” sequences) takes advantage of the degeneracy of codons, as demonstrated by the diversity of codon combinations for a given amino acid tribase pair, and typically involves modifying a nucleic acid sequence to enhance expression in a specific host cell by replacing at least one codon in the natural sequence with a codon that is more frequently or most frequently used in the host cell’s gene, while maintaining the natural amino acid sequence. For example, a nucleic acid encoding a target polypeptide can be modified to replace a codon that has a higher frequency of use in a given prokaryotic or eukaryotic cell, including bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, or any other host cell, compared to the naturally occurring nucleic acid sequence. Codon usage tables are readily available, for example, at “codon usage databases.” These tables can be modified in various ways. See Nakamura et al. (2000). Nucleic Acids Res. 28(1):292, which is incorporated herein by reference in its entirety for all purposes. Computer algorithms for codon optimization of specific sequences expressed in a particular host are also available (see, for example, Gene Forge).
[0106] The term "locus" refers to a specific location on a chromosome within an organism's genome, representing a gene (or significant sequence), DNA sequence, or polypeptide coding sequence. For example, " ALB "Locus" can refer to ALB Gene, ALB A specific location in a DNA sequence, an albumin-coding sequence, or on a chromosome in an organism's genome. ALB The ALB location has been identified as the site where this type of sequence resides. ALB "Locus" may include ALBRegulatory elements of a gene include, for example, enhancers, promoters, 5' and / or 3' untranslated regions (UTRs), or combinations thereof.
[0107] The term "gene" refers to a DNA sequence in a chromosome that, if naturally present, may contain at least one coding region and at least one non-coding region. The DNA sequence of a chromosome encoding a product (e.g., but not limited to RNA products and / or polypeptide products) may include a coding region interrupted by non-coding introns and a sequence located adjacent to the coding region at both the 5' and 3' ends such that the gene corresponds to a full-length mRNA sequence (containing 5' and 3' untranslated sequences). Additionally, other non-coding sequences, including regulatory sequences (e.g., but not limited to promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, isolation sequences, and matrix attachment regions, may be present in a gene. These sequences may be located near the coding region of the gene (e.g., but not limited to within 10 kb) or at distal sites, and these sequences may influence the level or rate of transcription and translation of the gene.
[0108] The term "allele" refers to a variant form of a gene. Some genes have multiple different forms, located at the same location or locus on a chromosome. Diploid organisms have two alleles at each locus. Each pair of alleles represents the genotype at a specific locus. If there are two identical alleles at a particular locus, the genotype is described as homozygous, and if the two alleles are different, the genotype is described as heterozygous.
[0109] A promoter is a regulatory region of DNA that typically contains a TATA box that guides RNA polymerase II to initiate RNA synthesis at the appropriate transcription start site of a specific polynucleotide sequence. Promoters may additionally contain other regions that influence the rate of transcription initiation. The promoter sequences disclosed herein regulate the transcription of operatively linked polynucleotides. Promoters may be active in one or more of the cell types disclosed herein (e.g., mouse cells, rat cells, pluripotent cells, single-cell embryos, differentiated cells, or combinations thereof). Promoters may be, for example, constitutively active promoters, conditional promoters, inducible promoters, time-restricted promoters (e.g., developmentally regulatory promoters), or spatially restricted promoters (e.g., cell-specific or tissue-specific promoters). Examples of promoters can be found, for example, in WO 2013 / 176772, which is incorporated herein by reference in its entirety for all purposes.
[0110] "Operable linkage" or "operably coupled" involves juxtaposing two or more components (e.g., a promoter and another sequence element) such that both components function normally and allow at least one component to mediate the function imposed on at least one of the other components. For example, if a promoter controls the transcriptional level of a coding sequence in response to the presence or absence of one or more transcriptional regulatory factors, then a promoter can be operably coupled to a coding sequence. Operable linkages may include sequences that are adjacent to each other or act in a trans-regulatory manner (e.g., a regulatory sequence that acts at a distance to control the transcription of a coding sequence).
[0111] The methods and compositions provided herein employ a variety of different components. Some components throughout the specification may have active variants and fragments. The term "functional" refers to the innate ability of a protein or nucleic acid (or a fragment or variant thereof) to exhibit biological activity or function. The biological function of a functional fragment or variant may be the same as, or may actually be altered (e.g., regarding its specificity or selectivity or efficacy) compared to the original molecule, but retains the essential biological function of the molecule.
[0112] The term "variant" refers to a nucleotide sequence that differs from the most common sequence in the population (e.g., by one nucleotide) or a protein sequence that differs from the most common sequence in the population (e.g., by one amino acid).
[0113] When referring to proteins, the term "fragment" means a protein that is shorter than the full-length protein or has fewer amino acids. When referring to nucleic acids, the term "fragment" means a nucleic acid that is shorter than the full-length nucleic acid or has fewer nucleotides. When referring to protein fragments, a fragment can be, for example, an N-terminal fragment (i.e., a portion of the C-terminus of the protein removed), a C-terminal fragment (i.e., a portion of the N-terminus of the protein removed), or an internal fragment (i.e., a portion of both the N-terminus and C-terminus of the protein removed). When referring to nucleic acid fragments, a fragment can be, for example, a 5' fragment (i.e., a portion of the 3' end of the nucleic acid removed), a 3' fragment (i.e., a portion of the 5' end of the nucleic acid removed), or an internal fragment (i.e., a portion of both the 5' end and 3' end of the nucleic acid removed).
[0114] In the context of two polynucleotide or polypeptide sequences, "sequence identity" or "identity" refers to the same residues in two sequences when compared for maximum correspondence within a specified comparison window. When using a sequence identity percentage relative to proteins, dissimilar residue positions are often distinguished by conserved amino acid substitutions, where an amino acid residue is substituted by another amino acid residue with similar chemical properties (e.g., charge or hydrophobicity) and therefore does not alter the functional properties of the molecule. When there are differences in the conserved substitutions of sequences, the sequence identity percentage can be adjusted upwards to correct for the conservatism of the substitution. Thus, sequences that are different despite having similar conserved substitutions are considered to have "sequence similarity" or "identity." The means of making such adjustments are well known. Typically, this involves counting conserved substitutions as partial mismatches rather than complete mismatches, thereby increasing the sequence identity percentage. Thus, for example, when the score for identical amino acids is 1 and the score for non-conserved substitutions is zero, the score for conserved substitutions is between zero and 1. The score for conserved substitutions is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).
[0115] The "Sequence Identity Percentage" is a value determined by comparing two optimally aligned sequences within a comparison window (the maximum number of perfectly matched residues). The portion of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., vacancies) compared to the reference sequence (excluding additions or deletions) to achieve optimal alignment. This percentage is calculated by determining the number of positions in both sequences where the same nucleic acid base or amino acid residue occurs to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the sequence identity percentage. Unless otherwise specified (e.g., the shorter sequence contains linked heterologous sequences), the comparison window is the full length of the shorter of the two compared sequences.
[0116] Unless otherwise stated, sequence identity / similarity values include values obtained using GAP version 10 with the following parameters: nucleotide sequence identity % and similarity using a GAP weight of 50 and a length weight of 3 and an nwsgapdna.cmp scoring matrix; amino acid sequence identity % and similarity using a GAP weight of 8 and a length weight of 2 and a BLOSUM62 scoring matrix; or any equivalent procedure thereof. "Equivalent procedure" includes any sequence comparison procedure that generates alignments with the same nucleotide or amino acid residue match and the same percentage of sequence identity for any two sequences in question when compared to corresponding alignments generated by GAP version 10.
[0117] The term "conservative amino acid substitution" refers to replacing a normally present amino acid in a sequence with a different amino acid having similar size, charge, or polarity. Examples of conservative substitution include replacing one nonpolar (hydrophobic) residue with another nonpolar residue, such as isoleucine, valine, or leucine. Similarly, examples of conservative substitution include replacing one polar (hydrophilic) residue with another polar residue, such as between arginine and lysine, glutamine and asparagine, or glycine and serine. Additionally, replacing one basic residue with a basic residue, such as lysine, arginine, or histidine, or replacing one acidic residue with another acidic residue, such as aspartic acid or glutamic acid, are further examples of conservative substitution. Examples of nonconservative substitution include replacing a polar (hydrophilic) residue, such as cysteine, glutamine, glutamic acid, or lysine, with a nonpolar amino acid residue, such as isoleucine, valine, leucine, alanine, or methionine, and / or replacing a nonpolar residue with a polar residue. Typical amino acid classifications are summarized below.
[0118] Table 1. Amino acid classification
[0119] "Homologous" sequences (e.g., nucleic acid sequences) include sequences that are identical or substantially similar to a known reference sequence, such that they are, for example, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known reference sequence. Homologous sequences can include, for example, orthologous and paralogous sequences. For example, homologous genes typically originate from a common ancestral DNA sequence through speciation events (orthologous genes) or gene duplication events (paralogous genes). "Orthologous" genes include genes that evolved from a common ancestral gene in different species through speciation. Orthologous genes typically retain the same function during evolution. "Paralogous" genes include genes associated with duplication within the genome. Paralogous genes may evolve new functions during evolution.
[0120] The term "in vitro" includes artificial environments and processes or reactions that occur within artificial environments (e.g., test tubes or isolated cells or cell lines). The term "in vivo" includes natural environments (e.g., cells, organisms, or bodies) and processes or reactions that occur within natural environments. The term "ex vivo" includes cells that have been removed from an organism and processes or reactions that occur within such cells.
[0121] As used herein, the term "antibody" refers to an immunoglobulin molecule comprising four polypeptide chains interconnected by disulfide bonds: two heavy (H) chains and two light (L) chains. Each heavy chain contains a heavy chain variable region (abbreviated herein as HCVR or VH) and a heavy chain constant region. The heavy chain constant region contains three domains, CH1, CH2, and CH3. Each light chain contains a light chain variable region (abbreviated herein as LCVR, VL, or VK) and a light chain constant region. The light chain constant region contains one domain, CL. The VH and VL regions can be further subdivided into hypervariable regions called complementarity-determining regions (CDRs), which are interspersed with more conserved regions called framework regions (FRs). Each VH and VL consists of three CDRs and four FRs, arranged in the following order from the amino terminus to the carboxyl terminus: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4 (heavy chain CDRs can be abbreviated as HCDR1, HCDR2, and HCDR3; light chain CDRs can be abbreviated as LCDR1, LCDR2, and LCDR3). The term "high affinity" antibody refers to an antibody with at least 10... -9 M, at least 10 -10 M, at least 10 -11 M, or at least 10 -12 Those antibodies that bind to M with affinity, such as those that do so via surface plasmon resonance (e.g., BIACORE). ™ This can be measured by solution affinity ELISA or other methods. The term "antibody" can cover any type of antibody, such as monoclonal or polyclonal antibodies. Furthermore, antibodies can be of any origin, such as mammalian or non-mammal origin. In one embodiment, the antibody can be mammalian or avian. In another embodiment, the antibody can be of human origin and can also be a human monoclonal antibody.
[0122] The phrase "bispecific antibody" refers to an antibody capable of selectively binding to two or more epitopes. Bispecific antibodies typically comprise two distinct heavy chains, each specifically binding to different epitopes on two different molecules (e.g., antigens) or on the same molecule (e.g., the same antigen). If a bispecific antibody is capable of selectively binding to two different epitopes (a first epitope and a second epitope), the affinity of the first heavy chain for the first epitope will typically be at least one to two, three, or four orders of magnitude lower than the affinity of the first heavy chain for the second epitope, and vice versa. The epitopes recognized by the bispecific antibody can be located on the same or different targets (e.g., on the same or different proteins). Bispecific antibodies can be prepared, for example, by combining heavy chains that recognize different epitopes of the same antigen. For example, a nucleic acid sequence encoding a variable sequence of a heavy chain recognizing different epitopes of the same antigen can be fused with a nucleic acid sequence encoding a constant region of a different heavy chain, and such sequences can be expressed in cells expressing immunoglobulin light chains. A typical bispecific antibody has two heavy chains and one immunoglobulin light chain. Each heavy chain has three heavy chain CDRs, followed by a (N-terminal to C-terminal) CH1 domain, a hinge, a CH2 domain, and a CH3 domain. The immunoglobulin light chain does not confer antigen-binding specificity but can associate with each heavy chain, or can associate with each heavy chain and bind one or more of the epitopes bound by the heavy chain antigen-binding region, or can associate with each heavy chain and enable one or two of the heavy chains to bind one or two epitopes.
[0123] The phrase "heavy chain" or "immunoglobulin heavy chain" contains a constant region sequence of an immunoglobulin heavy chain from any organism and, unless otherwise stated, contains a heavy chain variable domain. Unless otherwise stated, the heavy chain variable domain contains three heavy chain CDRs and four FR regions. Fragments of the heavy chain contain CDRs, CDRs, and FRs, as well as combinations thereof. A typical heavy chain has a CH1 domain, a hinge, a CH2 domain, and a CH3 domain following the variable domain (from the N-terminus to the C-terminus). Functional fragments of the heavy chain contain fragments capable of specifically recognizing antigens (e.g., antigens with KD in the micromolar, nanomolar, or picomolar range), capable of being expressed and secreted by cells, and including at least one CDR.
[0124] The phrase "light chain" encompasses the constant region sequence of an immunoglobulin light chain from any organism and, unless otherwise specified, includes human κ and λ light chains. Unless otherwise specified, the light chain variable (VL) domain typically comprises three light chain CDRs and four frame (FR) regions. Generally, a full-length light chain contains a VL domain and a light chain constant domain from the amino terminus to the carboxyl terminus, the VL domain comprising FR1-CDR1-FR2-CDR2-FR3-CDR3-FR4. Light chains that can be used herein include, for example, those that do not selectively bind to a first or second antigen selectively bound by an antigen-binding protein. Suitable light chains include those that can be identified by screening the most commonly used light chains in an existing antibody library (wet library or computer library), wherein the light chain substantially does not interfere with the affinity and / or selectivity of the antigen-binding domain of the antigen-binding protein. Suitable light chains include those that can bind one or both epitopes bound by the antigen-binding region of an antigen-binding protein.
[0125] The phrase “variable domain” includes an amino acid sequence (modified as needed) of the immunoglobulin light or heavy chain, which, from the N-terminus to the C-terminus (unless otherwise specified), contains the following amino acid regions in sequence: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. “Variable domain” includes an amino acid sequence capable of folding into a canonical domain (VH or VL) with a double β-sheet structure, wherein these β-sheets are linked by disulfide bonds between residues of the first and second β-sheets.
[0126] The phrase "complementarity-determining region" or the term "CDR" contains an amino acid sequence encoded by the nucleic acid sequence of an organism's immunoglobulin gene. This amino acid sequence typically (i.e., in wild-type animals) occurs between two framework regions within the variable region of the light or heavy chain of an immunoglobulin molecule (e.g., an antibody or T-cell receptor). The CDR can be encoded by, for example, germline sequences, rearranged sequences, or unrearranged sequences, and is encoded, for example, by naive or mature B cells or T cells. In some cases (e.g., for CDR3), the CDR can be encoded by two or more sequences (e.g., germline sequences) that are discontinuous (e.g., in an unrearranged nucleic acid sequence) but are continuous in the B-cell nucleic acid sequence, for example, due to splicing or linking sequences (e.g., VDJ recombination to form the heavy chain CDR3).
[0127] The term “antibody fragment” refers to one or more fragments of an antibody that retain the ability to specifically bind to an antigen. Examples of binding fragments encompassed within the term “antibody fragment” include: (i) Fab fragments, a monovalent fragment consisting of VL, VH, CL, and CH1 domains; (ii) F(ab')2 fragments, a bivalent fragment comprising two Fab fragments linked by disulfide bonds at hinge regions; (iii) Fd fragments consisting of VH and CH1 domains; (iv) Fv fragments consisting of the VL and VH domains of a single arm of an antibody; (v) dAb fragments (Ward et al. (1989) Nature 241:544-546) consisting of a VH domain; (vi) isolated CDRs; and (vii) scFv fragments consisting of the two domains VL and VH of an Fv fragment, which are joined by a synthetic linker to form a single protein chain, wherein the VL and VH regions pair to form a monovalent molecule. Other forms of single-chain antibodies, such as biantibodies, are also covered within the term "antibody" (see, for example, Holliger et al. (1993)). Proc.Natl.Acad.Sci. USA90:6444-6448; Poljak et al. (1994) Structure 2:1121-1123).
[0128] The phrase "Fc-containing protein" includes antibodies, bispecific antibodies, immunoadhesins, and other binding proteins that contain at least the functional portions of the CH2 and CH3 regions of immunoglobulins. "Functional portion" refers to the CH2 and CH3 regions that can bind to Fc receptors (e.g., FcyR; or FcRn, i.e., the neonatal Fc receptor) and / or participate in complement activation. If the CH2 and CH3 regions contain deletions, substitutions, and / or insertions or other modifications that prevent them from binding to any Fc receptors and from activating complement, then the CH2 and CH3 regions are inactive.
[0129] Fc-containing proteins may contain modifications in their immunoglobulin domains, including modifications that affect one or more effector functions of the bound protein (e.g., modifications affecting FcyR binding, FcRn binding, thereby affecting half-life and / or CDC activity). Such modifications include, but are not limited to, the following modifications and combinations thereof, referring to the EU numbers for the immunoglobulin constant region: 238, 239, 248, 249, 250, 252, 254, 255, 256, 258, 265, 267, 268, 269, 270, 272, 276, 278, 280, 283, 285, 286, 289, 290, 292, 293, 294, 295, 296, 297, 298, 301, 303, 305, 307, 308, 309, 311, 312, 315. 318, 320, 322, 324, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 337, 338, 339, 340, 342, 344, 356, 358, 359, 360, 361, 362, 373, 375, 376, 378, 380, 382, 383, 384, 386, 388, 389, 398, 414, 416, 419, 428, 430, 433, 434, 435, 437, 438, and 439.
[0130] For example, but not as a limitation, the binding protein is an Fc-containing protein that exhibits a prolonged serum half-life (compared to the same Fc-containing protein without the listed modifications) and has the following modifications: modifications at positions 250 (e.g., E or Q), 250 and 428 (e.g., L or F), 252 (e.g., L / Y / F / W or T), 254 (e.g., S or T), and 256 (e.g., S / R / Q / E / D or T); or modifications at positions 428 and / or 433 (e.g., L / R / SI / P / Q or K) and / or 434 (e.g., H / F or Y); or modifications at positions 250 and / or 428; or modifications at positions 307 or 308 (e.g., 308F, V308F) and 434. In another example, modifications may include: 428L (e.g., M428L) and 434S (e.g., N434S) modifications; 428L, 2591 (e.g., V259I) and 308F (e.g., V308F) modifications; 433K (e.g., H433K) and 434 (e.g., 434Y) modifications; 252, 254 and 256 (e.g., 252Y, 254T and 256E) modifications; 250Q and 428L modifications (e.g., T250Q and M428L); 307 and / or 308 modifications (e.g., 308F or 308P).
[0131] As used herein, the term "antigen-binding protein" refers to a polypeptide or protein (one or more polypeptides that combine to form a single functional unit) that specifically recognizes epitopes on antigens, such as the cell-specific antigens and / or target antigens provided herein. Antigen-binding proteins can be multispecific. The term "multispecific" with respect to antigen-binding proteins means that the protein recognizes different epitopes on the same antigen or on different antigens. The multispecific antigen-binding proteins provided herein can be a single multifunctional polypeptide, or they can be a polymeric complex of two or more polypeptides covalently or non-covalently associated with each other. The term "antigen-binding protein" includes antibodies or fragments thereof provided herein that can be linked to or co-expressed with another functional molecule, such as another peptide or protein. For example, antibodies or fragments thereof can be functionally linked to one or more other molecular entities, such as proteins or fragments thereof (e.g., through chemical coupling, gene fusion, non-covalent association, or other means), to produce bispecific or multispecific antigen-binding molecules with a second binding specificity.
[0132] As used herein, the term "epitope" refers to an antigenic moiety recognized by a multispecific antigen-binding polypeptide. A single antigen (such as an antigenic polypeptide) may have more than one epitope. Epitopes can be defined as structural or functional. Functional epitopes are typically a subset of structural epitopes and are defined as those residues that directly promote the affinity between the antigen-binding polypeptide and the antigen. Epitopes can also be conformational, i.e., composed of nonlinear amino acids. In some embodiments, epitopes may comprise determinants of chemically active surface groups (such as amino acids, sugar side chains, phosphoryl groups, or sulfonyl groups) as molecules, and in some embodiments, may have specific three-dimensional structural characteristics and / or specific charge characteristics. Epitopes formed from consecutive amino acids are generally retained upon exposure to denaturing solvents, while epitopes formed from tertiary folds are generally lost upon treatment with denaturing solvents.
[0133] The term "domain" refers to any portion of a protein or polypeptide that has a specific function or structure. Preferably, the domains provided herein bind to cell-specific antigens or target antigens. As used herein, cell-specific antigen-binding domains or target antigen-binding domains include any naturally occurring, enzymatically obtainable, synthetic, or genetically engineered polypeptides or glycoproteins that specifically bind to antigens.
[0134] The interchangeable terms "half-body" or "half-antibody" refer to half of an antibody that essentially contains one heavy chain and one light chain. Antibody heavy chains can form dimers, so the heavy chain of one half can associate with a different molecule (e.g., another half) or another heavy chain containing an Fc polypeptide. Two slightly different Fc domains can undergo "heterodimerization," as occurs in the formation of bispecific antibodies or other heterodimers, heterotrimers, heterotetramers, etc. See Vincent and Murini (2012). Biotechnol.J. 7(12):1444-1450; and Shimamoto et al. (2012) MAbs 4(5):586-91. In one embodiment, the hemisomal variable domain specifically recognizes internalization effectors, and the hemisomal Fc domain dimers with an Fc fusion protein including a substitute enzyme (e.g., a peptide antibody).
[0135] The term "single-chain variable fragment" or "scFv" refers to a single-chain fusion polypeptide containing both the variable region (VH) of the immunoglobulin heavy chain and the variable region (VL) of the immunoglobulin light chain. In some embodiments, the VH and VL are linked by a linker sequence of 10 to 25 amino acids. scFv polypeptides may also contain other amino acid sequences, such as CL or CH1 regions. scFv molecules can be prepared by phage display or by direct subcloning of the heavy and light chains from hybridoma or B cells. See Ahmad et al. (2012). Clin.Dev.Immunol. 2012:980250, this document is incorporated herein by reference in its entirety for all purposes.
[0136] As used herein, in the context of humans, the term "newborn" encompasses human subjects who are at most or less than 1 year old (52 weeks), preferably at most or less than 24 weeks, more preferably at most or less than 12 weeks, more preferably at most or less than 8 weeks, and even more preferably at most or less than 4 weeks old. In some embodiments, the newborn human subject is at most 4 weeks old. In some embodiments, the newborn human subject is at most 8 weeks old. In another embodiment, the newborn human subject is within 3 weeks of birth. In another embodiment, the newborn human subject is within 2 weeks of birth. In another embodiment, the newborn human subject is within 1 week of birth. In another embodiment, the newborn human subject is within 7 days of birth. In another embodiment, the newborn human subject is within 6 days of birth. In another embodiment, the newborn human subject is within 5 days of birth. In another embodiment, the newborn human subject is within 4 days of birth. In another embodiment, the newborn human subject is within 3 days of birth. In another embodiment, the newborn human subject is within 2 days of birth. In another embodiment, the newborn human subject is within 1 day of birth. The time windows disclosed above apply to human subjects and are also intended to cover corresponding developmental time windows in other animals. As used herein, “newborn cells” refer to cells from newborn subjects, and a newborn cell population refers to a population of cells from newborn subjects.
[0137] As used herein, a “control” in a control sample or control subject is a comparison used for a measurement (e.g., a diagnostic measurement of a sign or symptom of a disease). In some embodiments, a control may be a subject sample from an earlier time point (e.g., prior to a treatment intervention) of the same subject. In some embodiments, a control may be a measurement from a normal subject (i.e., a subject who does not have the disease of the treated subject) to provide a normal control, such as the concentration or activity of an enzyme in a subject sample. In some embodiments, a normal control may be a population control, i.e., the mean of subjects in a general population. In some embodiments, a control may be an untreated subject with the same disease. In some embodiments, a control may be a subject treated with a different therapy (e.g., a standard of care). In some embodiments, a control may be a subject or a group of subjects from a natural history study of subjects with the disease of the compared subject. In some embodiments, a control is matched to the tested subject for certain factors, such as age and sex. In some embodiments, a control may be a control level from a specific laboratory (e.g., a clinical laboratory). The selection of an appropriate control is within the competence of those skilled in the art.
[0138] A composition or method that “comprising” or “contains” one or more of the listed elements may include other elements not specifically listed. For example, a composition that “comprising” or “contains” a protein may contain a protein alone or in combination with other ingredients. The transitional phrase “consistently made of” means that the scope of the claims should be interpreted to cover the specified elements listed in the claims as well as those elements that do not substantially affect the essential and novel characteristics of the claimed invention. Therefore, when used in the claims of this invention, the term “consistently made of” should not be interpreted as equivalent to “comprising”.
[0139] "Optional" or "optionally" means that the event or situation described below may or may not occur, and the description includes instances in which the event or situation occurs as well as instances in which the event or situation does not occur.
[0140] The specification of a numerical range includes all integers within that range or defining that range, as well as all subranges defined by the integers within that range. For example, 5-10 nucleotides is understood to mean 5, 6, 7, 8, 9, or 10 nucleotides, while 5%-10% is understood to include 5% and all possible values up to 10%.
[0141] A sequence of at least 17 nucleotides in a 20-nucleotide sequence is understood to include 17, 18, 19, or 20 nucleotides in the provided sequence, thus providing an upper limit, even if no upper limit is explicitly provided as would be clearly understood. Similarly, a sequence of at most 3 nucleotides is understood to cover 0, 1, 2, or 3 nucleotides, thus providing a lower limit, even if no lower limit is explicitly provided. When “at least,” “at most,” or other similar language modifies a number, it is understood to modify each number in the series.
[0142] As used in this article, “no more than” or “less than” is understood to be the value adjacent to the phrase and a logically lower value or integer to zero that is reasonable from the context. For example, the double-stranded region of “no more than 2 nucleotide base pairs” has 2, 1, or 0 nucleotide base pairs. When “no more than” or “less than” precedes a series of numbers or ranges, it is to be understood that each number in the series or range is modified.
[0143] As used herein, "detection of analyte" is understood to mean performing a determination in which the analyte (if present) can be detected, wherein the analyte is present in an amount higher than the detection level of the assay.
[0144] As used herein, “loss of function” is understood to mean the absence of activity for any reason, such as the absence of enzyme activity. In some embodiments, the absence of activity may be due to the absence of a functional protein, for example, the protein is not transcribed or translated, the protein is translated but unstable, or it cannot be properly transported intracellularly or systemically. In some embodiments, the absence of activity may be due to the presence of mutations, such as point mutations, truncations, or aberrant splicing, that result in the presence of a protein without function. Loss of function can be partial or complete. In some embodiments, it is known that different degrees of loss of function can lead to various conditions, disease severity, or age of onset. As used herein, loss of function is preferably not a temporary loss of function, such as loss of function due to stress or other responses that result in a temporary loss of function of the protein. Therapeutic interventions to correct protein loss of function may include compensating for the loss of function with the lacking protein, or compensating for the loss of function with a protein that compensates for the loss of function but has a different sequence or structure than the protein with the loss of function. It is to be understood that the loss of function of a protein can be compensated for by providing or altering the activity of another protein in the same biological pathway. In some implementations, compensating for loss-of-function proteins includes one or more of truncated, mutated, or non-natural sequences to guide protein transport within cells or throughout the body, thereby overcoming the loss of protein function. Therapeutic interventions may or may not correct loss of protein function in all cell types or tissues. Therapeutic interventions may include expressing proteins to compensate for loss of function at sites distant from where the lacking protein is normally expressed, such as at sites where the deficiency leads to cellular or organ dysfunction. Therapeutic interventions may include expressing proteins in the liver to compensate for loss of function at sites distant from the liver. In humans and other species, many genetic mutations are associated with specific loss-of-function mutations.
[0145] As used herein, “enzyme deficiency” is understood as insufficient enzyme activity due to loss of protein function. Enzyme deficiency can be partial or complete and may result in variations in the timing of onset or the severity of signs or symptoms, depending on the extent and location of functional loss. As used herein, enzyme deficiency is preferably not a transient deficiency caused by stress or other factors. In humans and other species, many genetic mutations are associated with enzyme deficiency. In some embodiments, enzyme deficiency leads to congenital metabolic defects. In some embodiments, enzyme deficiency leads to lysosomal storage diseases. In some embodiments, enzyme deficiency leads to galactosemia. In some embodiments, enzyme deficiency leads to hemorrhagic diseases.
[0146] As used herein, it is to be understood that when the maximum value is expressed as 100% (e.g., 100% inhibition or 100% encapsulation), that value is limited by the detection method. For example, 100% inhibition is understood as inhibition to a level below the detection level of the assay, and 100% encapsulation is understood as the material intended for encapsulation not being detected outside the vesicle.
[0147] Unless the context otherwise requires, the term "about" covers values ±5% of the specified value. In some embodiments, the term "about" should be understood to cover variations or errors permissible in the art, such as within 2 standard deviations from the mean, or the sensitivity of the method used to make the measurement, or a percentage of values permissible in the art (e.g., in age descriptions). When "about" precedes the first value in a series, it can be understood to modify each value in that series.
[0148] The term “and / or” means and covers any and all possible combinations of one or more of the associated listed items, as well as the absence of such combinations when interpreted in the alternative (“or”) sense.
[0149] The term "or" refers to any one member of a particular list, and also includes any combination of members of that list.
[0150] Unless the context clearly indicates otherwise, the singular forms of the articles “a” and “the” include plural referents. For example, the terms “protein” or “at least one protein” can include multiple proteins, or mixtures thereof.
[0151] Statistically significant means p≤0.05.
[0152] In the event of a conflict between the sequence in this application and the specified login number or position within the login number, the sequence in this application shall prevail. Detailed Implementation
[0153] I. Overview Compositions and methods are also provided for inserting nucleic acids encoding multi-domain therapeutic proteins (e.g., GAA fusion proteins) into target genomic loci in cells, cell populations, or subjects (e.g., neonatal cells, neonatal cell populations, or neonatal subjects), or for expressing nucleic acids encoding multi-domain therapeutic proteins (e.g., GAA fusion proteins) from target genomic loci in cells, cell populations, or subjects (e.g., neonatal cells, neonatal cell populations, or neonatal subjects). Compositions and methods are also provided for treating GAA deficiency, reducing glycogen accumulation in tissues, treating Pompe disease, or preventing or reducing the onset of signs or symptoms of Pompe disease in subjects (e.g., neonatal subjects). Cells or cell populations comprising nucleic acid constructs (e.g., neonatal cells or neonatal cell populations) containing the coding sequence of a multi-domain therapeutic protein (e.g., GAA fusion protein) inserted into a target genomic locus are also provided.
[0154] This article also provides nucleic acid constructs and compositions (e.g., free expression vectors) for expressing multi-domain therapeutic proteins (e.g., GAA fusion proteins). This article also provides methods for inserting the coding sequences of multi-domain therapeutic proteins (e.g., GAA fusion proteins) into target genomic loci (such as endogenous...). ALB Nucleic acid constructs and compositions that encode the multidomain therapeutic protein (e.g., GAA fusion protein) at a locus and / or express the coding sequence of such multidomain therapeutic protein. These nucleic acid constructs and compositions can be used in methods of: integrating or inserting the nucleic acid of a multidomain therapeutic protein (e.g., GAA fusion protein) into a target genomic locus in a cell or cell population or subject; expressing a multidomain therapeutic protein (e.g., GAA fusion protein) in a cell or cell population or subject; reducing glycogen accumulation in a cell or cell population or subject; treating Pompe disease or GAA deficiency in a subject; and preventing or reducing the onset of signs or symptoms of Pompe disease in subjects (including neonatal cells and subjects).
[0155] Compositions or combinations or kits comprising nucleic acid constructs containing coding sequences for multi-domain therapeutic proteins are also provided, and are combined with a nuclease preparation or one or more nucleic acids encoding a nuclease preparation, wherein the nuclease preparation targets a nuclease target site in a target genomic locus. As used herein, the term “combined with” means that additional components may be administered before, simultaneously with, or after the application of the nucleic acid construct. The different components of the combination may be formulated as a single composition, for example for simultaneous delivery, or individually formulated as two or more compositions (e.g., kits comprising each component, for example, wherein the additional preparation is in a separate formulation).
[0156] More specifically, in some embodiments, this document describes therapeutic products based on CRISPR / Cas9 gene editing technology and optionally contained in a lipid nanoparticle (LNP) delivery system, which bind to a DNA gene insertion template of a multi-domain therapeutic protein (e.g., a GAA fusion protein) optionally contained in recombinant adeno-associated virus serotype 8 (rAAV8). The CRISPR / Cas9 component has been engineered to target gene loci (e.g., safe harbor loci, such as those in hepatocytes). ALB The gene locus targets and cleaves double-stranded DNA, allowing a multi-domain therapeutic protein (e.g., GAA fusion protein) DNA template to be inserted into the genome at the target genomic locus. The transgene insertion provides a functional multi-domain therapeutic protein (e.g., GAA fusion protein) gene encoding a missing or absent genomic component in patients with Pompe disease. GAA .
[0157] Compared to the native GAA coding sequence, some of the multi-domain therapeutic protein (e.g., GAA fusion protein) coding sequences in the constructs disclosed herein are optimized for expression. For example, the coding sequences in the constructs disclosed herein may contain one or more modifications, such as codon optimization (e.g., optimization of human codons), deletion of CpG dinucleotides, mutation of cryptic splicing sites, or any combination thereof. Other multi-domain therapeutic protein coding sequences in the constructs disclosed herein include the native GAA coding sequence.
[0158] In some implementations, to minimize missplicing in multidomain therapeutic protein nucleic acid constructs, we employed a two-pronged approach: (1) functionally identifying cryptic splice donors via RNA sequencing (RNA-Seq) (rather than predicting based on shared sequences) and introducing synonymous mutations to disrupt key “GU” nucleotide pairs that form the core of the splice donor sequence; and (2) adding additional elements with the net effect of extending the time from the transcription of polyA to the arrival of RNA polymerase at the next splice acceptor site. We introduced tandem polyA signals (e.g., bovine growth hormone (BGH) and SV40), MAZ elements that cause polymerase pause, or additional filler sequences to extend the time between the transcription of polyA by RNA polymerase and its transcription of the next splice acceptor. SV40 polyA is bidirectional, but “late” oriented polyadenylation is more effective than “early” oriented polyadenylation. In some implementations, to tandemly link the SV40 “late” polyA with the BGH polyA, we mutate the transcription terminator sequence that exists in the reverse direction of the “early” SV40, thus making this type of SV40 polyA unidirectional rather than bidirectional. Therefore, if our DNA insertion template is inserted into the genome in a nonfunctional “reverse” orientation, transcription should proceed directly through the entire locus (e.g., the albumin locus), and the nonfunctional insertion should be spliced out along with the first intron because there is no transcription terminator sequence in the “reverse” orientation.
[0159] II. Multi-domain therapeutic proteins and those used for insertion into cells encoding and / Or express multi-domain therapeutic proteins White nucleic acid construct composition It provides multi-domain therapeutic proteins containing a TfR-binding delivery domain fused to a lysosomal α-glucosidase (GAA) peptide or a CD63-binding delivery domain, and allows insertion of multi-domain therapeutic protein coding sequences into target genomic loci (such as endogenous). ALB Nucleic acid constructs and compositions that encode multidomain therapeutic proteins at loci and / or express multidomain therapeutic protein sequences. These multidomain therapeutic proteins and nucleic acid constructs and compositions can be applied to cells, cell populations, or subjects and can be used for methods of integrating multidomain therapeutic protein nucleic acids into target genomic loci, methods of expressing multidomain therapeutic proteins in cells or cell populations or subjects, methods of reducing glycogen accumulation in cells or cell populations or tissues of subjects, methods of treating Pompe disease or GAA deficiency in subjects, and methods of preventing or reducing the onset of signs or symptoms of Pompe disease or GAA deficiency in subjects.
[0160] This article provides a multi-domain therapeutic protein comprising a TfR-binding delivery domain fused to a lysosomal α-glucosidase (GAA) polypeptide or a CD63-binding delivery domain. The multi-domain therapeutic protein and compositions can be used for methods of introducing the multi-domain therapeutic protein into cells or cell populations or subjects, methods of treating Pompe disease or GAA deficiency in subjects, and methods of preventing or reducing the onset of signs or symptoms of Pompe disease or GAA deficiency in subjects.
[0161] This article provides nucleic acid constructs and compositions that allow the insertion of multi-domain therapeutic protein-coding sequences into target genomic loci (such as endogenous albumin). ALB This document describes the expression of multi-domain therapeutic protein-coding sequences at a target genomic locus. Nucleic acid constructs and compositions (e.g., free expression vectors) for expressing multi-domain therapeutic proteins are also provided. These constructs and compositions can be used in methods of introducing a nucleic acid construct containing a multi-domain therapeutic protein-coding sequence into cells or cell populations or subjects; integrating a multi-domain therapeutic protein nucleic acid into a target genomic locus; expressing a multi-domain therapeutic protein in cells; treating Pompe disease or GAA deficiency in subjects; and preventing or reducing the onset of signs or symptoms of Pompe disease or GAA deficiency in subjects. Nuclease preparations (e.g., targeted endogenous nucleases) are also provided. ALB (Locus) or nucleic acid encoding a nuclease preparation to facilitate the integration of nucleic acid constructs into target genomic loci (such as endogenous loci). ALB (in the locus).
[0162] A. Multi-domain therapeutic proteins and nucleic acid constructs encoding multi-domain therapeutic proteins The compositions and methods described herein include the use of multidomain therapeutic proteins comprising a lysosomal α-glucosidase (GAA) polypeptide (GAA or its bioactive portion thereof, to provide GAA enzyme alternative activity) linked or fused to a TfR-binding delivery domain or a CD63-binding delivery domain. The compositions and methods described herein also include the use of nucleic acid constructs comprising a coding sequence of a multidomain therapeutic protein. The compositions and methods described herein may also include the use of nucleic acid constructs comprising a coding sequence of a multidomain therapeutic protein or an inverse complementary sequence of a multidomain therapeutic protein coding sequence. Such nucleic acid constructs can be used to express multidomain therapeutic proteins in cells. Such nucleic acid constructs can be used for insertion into target genomic loci or into cleavage sites generated by nuclease preparations or CRISPR / Cas systems as disclosed elsewhere herein. The term cleavage site includes a DNA sequence in which a nick or double-strand break is formed by a nuclease preparation (e.g., a Cas9 protein complexed with guide RNA). In some implementations, double-strand breaks are generated by a Cas9 protein complexed with a guide RNA, such as a SpyCas9 protein complexed with a SpyCas9 guide RNA.
[0163] The length of the nucleic acid constructs disclosed herein can vary. The constructs can be, for example, from about 1 kb to about 5 kb, such as from about 1 kb to about 4.5 kb or from about 1 kb to about 4 kb. Exemplary nucleic acid constructs are between about 1 kb and about 5 kb or between about 1 kb and about 4 kb. Alternatively, the length of the nucleic acid constructs can be between about 1 kb and about 1.5 kb, about 1.5 kb and about 2 kb, about 2 kb and about 2.5 kb, about 2.5 kb and about 3 kb, about 3 kb and about 3.5 kb, about 3.5 kb and about 4 kb, about 4 kb and about 4.5 kb or about 4.5 kb and about 5 kb. Alternatively, the length of the nucleic acid constructs can be, for example, no more than 5 kb, no more than 4.5 kb, no more than 4 kb, no more than 3.5 kb, no more than 3 kb, or no more than 2.5 kb.
[0164] The construct may contain deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), and may be single-stranded, double-stranded, or partially single-stranded and partially double-stranded, and may be introduced into the host cell in linear or circular (e.g., small loop) form. See, for example, US 2010 / 0047805, US 2011 / 0281361, and US 2011 / 0207221, each of which is incorporated herein by reference in its entirety for all purposes. If introduced in a linear form, the ends of the construct may be protected by known methods (e.g., from exonuclease degradation). For example, one or more dideoxynucleotide residues may be added to the 3′ end of a linear molecule and / or self-complementary oligonucleotides may be ligated to one or both ends. See, for example, Chang et al. (1987). Proc.Natl.Acad.Sci.USA 84:4959-4963 and Nehls et al. (1996) Science 272:886-889, each of which is incorporated herein by reference in its entirety for all purposes. Other methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino groups and the use of modified internucleotide bonds, such as phosphate thioesters, aminophosphate esters, and O-methylribose or deoxyribose residues. The construct can be introduced into cells as part of a vector molecule with additional sequences, such as origin of replication, promoters, and genes encoding antibiotic resistance. Viral elements may be omitted from the construct. Furthermore, the construct can be introduced as a naked nucleic acid, as a nucleic acid complexed with a formulation (such as liposomes or poloxamer), or delivered via a virus (e.g., adenovirus, adeno-associated virus (AAV), herpesvirus, retrovirus, or lentivirus).
[0165] The constructs disclosed herein may be modified at one or both ends to incorporate one or more suitable structural features and / or confer one or more functional benefits as needed. For example, structural modifications may vary depending on the method used to deliver the constructs disclosed herein to host cells (e.g., delivery using a viral vector or packaging into lipid nanoparticles for delivery). Such modifications include, for example, terminal structures such as inverted terminal repeats (ITRs), hairpins, loops, and other structures such as ring-like structures. For example, the constructs disclosed herein may contain one, two, or three ITRs, or may contain no more than two ITRs. Various methods of structural modification are known.
[0166] Some constructs can be inserted such that their expression is driven by an endogenous promoter at the insertion site (e.g., when the construct is integrated into the host cell). ALB Endogenous in the locus ALB (Promoter). Such constructs may not contain a promoter that drives the expression of multi-domain therapeutic proteins. For example, the expression of multi-domain therapeutic proteins may be driven by a host cell promoter (e.g., when the transgene is integrated into the host cell). ALB Endogenous in the locus ALB(Promoter). In this case, the construct may lack control elements driving its expression (e.g., promoters and / or enhancers) (e.g., promoterless construct). In other cases, the construct may contain promoters and / or enhancers, such as constitutive promoters or inducible or tissue-specific (e.g., liver or platelet-specific) promoters that drive the expression of multi-domain therapeutic proteins in their free form or after integration. For example, the construct may be a construct for expression but not for insertion (e.g., a free construct). In some embodiments, the construct is not used for insertion. Non-limiting examples of constitutive promoters include the cytomegalovirus immediate early promoter (CMV), simian virus (SV40) promoter, adenovirus major late (MLP) promoter, Rous sarcoma virus (RSV) promoter, mouse mammary tumor virus (MMTV) promoter, phosphoglycerate kinase (PGK) promoter, elongation factor-α (EF1a) promoter, ubiquitin promoter, actin promoter, microtubule promoter, immunoglobulin promoter, functional fragments thereof, or combinations of any of the foregoing substances. For example, the promoter may be a CMV promoter or a truncated CMV promoter. In another example, the promoter may be an EF1a promoter. Non-limiting exemplary inducible promoters include those that can be induced by heat shock, light, chemical agents, peptides, metals, steroids, antibiotics, or alcohols. Inducible promoters may be promoters with low basal (non-inducible) expression levels, such as Tet-On. ®Promoters (Clontech). Although not essential for expression, constructs may contain transcriptional or translational regulatory sequences, such as promoters, enhancers, isolators, internal ribosome entry sites, additional sequences encoding peptides, and / or polyadenylation signals. Constructs may contain a sequence encoding a multi-domain therapeutic protein, which is downstream of and operatively linked to a signal sequence encoding a signal peptide. In some examples, nucleic acid constructs function in the homology-independent insertion of nucleic acids encoding multi-domain therapeutic proteins. Such nucleic acid constructs can function in, for example, non-dividing cells (e.g., cells where non-homologous end joining (NHEJ) rather than homologous recombination (HR) is the primary mechanism for repairing double-strand DNA breaks) or dividing cells (e.g., actively dividing cells). Such constructs can be, for example, homology-independent donor constructs. In preferred embodiments, the promoters and other regulatory sequences are adapted for human use, for example, recognized by regulatory factors in human cells (e.g., human liver cells), and are acceptable to regulatory agencies for human use. Examples of liver-specific promoters include TTR promoters, such as human or mouse TTR promoters. In one example, the construct may contain a TTR promoter, such as a mouse TTR promoter or a human TTR promoter (e.g., the coding sequence of a multi-domain therapeutic protein is operatively linked to the TTR promoter). In one example, the construct may contain a SERPINA1 enhancer, such as a mouse SERPINA1 enhancer or a human SERPINA1 enhancer (e.g., the coding sequence of a multi-domain therapeutic protein is operatively linked to the SERPINA1 enhancer). In one example, the construct may contain both a TTR promoter and a SERPINA1 enhancer, such as a human SERPINA1 enhancer and a mouse TTR promoter (e.g., the coding sequence of a multi-domain therapeutic protein is operatively linked to both the SERPINA1 enhancer and the TTR promoter).
[0167] The constructs disclosed herein can be modified to include or exclude any suitable structural features required for any particular purpose and / or to impart one or more desired functions. For example, some constructs disclosed herein do not contain homologous arms. Some constructs disclosed herein are capable of insertion into the cleavage site of a target genomic locus or target DNA sequence of a nuclease preparation via non-homologous end joining (e.g., capable of inserting into safe harbor genes, such as...). ALB (In the locus). For example, such constructs can be inserted into blunt-ended double-strand breaks after being cleaved with a nuclease preparation as disclosed herein (e.g., a CRISPR / Cas system, such as the SpyCas9 CRISPR / Cas system). In specific examples, the constructs can be delivered via AAV and may be able to be inserted via non-homologous end conjugation (e.g., the nucleic acid construct does not contain homologous arms).
[0168] In certain examples, the construct can be inserted via homology-independent targeted integration. For instance, a multi-domain therapeutic protein-coding sequence in the construct may have a target site of a nuclease agent lateralized on each side (e.g., the same target site in the target DNA sequence used for targeting the inserter (e.g., in a safe harbor gene), and the same nuclease agent used to cleave the target DNA sequence for targeted insertion). The nuclease agent can then cleave the target site lateralized with the multi-domain therapeutic protein. In a specific example, the construct is delivered via AAV-mediated delivery, and cleavage of the target site lateralized with the multi-domain therapeutic protein-coding sequence removes the inverted terminal repeat (ITR) sequence of the AAV. In some cases, if the multi-domain therapeutic protein-coding sequence is inserted into the cleavage site or target DNA sequence in the correct orientation, the target DNA sequence used for targeted insertion (e.g., the target DNA sequence in a safe harbor locus, such as a gRNA target sequence containing a lateralized protospacer motif) is no longer present, but if the multi-domain therapeutic protein-coding sequence is inserted into the cleavage site or target DNA sequence in the opposite orientation, it is reformed. This helps ensure that multi-domain therapeutic protein coding sequences are inserted with the correct expression orientation.
[0169] The constructs disclosed herein may contain polyadenylated sequences or polyadenylated tail sequences (e.g., downstream or 3' of a multi-domain therapeutic protein-coding sequence). Methods for designing suitable polyadenylated tail sequences are well known. Polyadenylated tail sequences may encode, for example, a "poly-A" segment downstream of a multi-domain therapeutic protein-coding sequence. The poly-A tail may contain, for example, at least 20, 30, 40, 50, 60, 70, 80, 90, or 100 adenine nucleotides, and optionally up to 300 adenine nucleotides. In specific examples, the poly-A tail contains 95, 96, 97, 98, 99, or 100 adenine nucleotides. Methods for designing suitable polyadenylated tail sequences and / or polyadenylated signal sequences are well known. For example, the polyadenylated signal sequence AAUAAA is commonly used in mammalian systems, but variants such as UAUAAA or AU / GUAAA have been identified. See, for example, Proudfoot (2011). Genes & Dev.25(17):1770-82, which is incorporated herein by reference in its entirety for all purposes. The term polyadenylation signal sequence refers to any sequence that directs transcription termination and the addition of a poly-A tail to the mRNA transcript. In eukaryotes, transcription terminators are recognized by protein factors, and polyadenylation occurs after termination, which is the process of adding a poly(A) tail to the mRNA transcript in the presence of poly(A) polymerase. Mammalian poly(A) signals typically consist of a core sequence of about 45 nucleotides in length, which may be flanked by various auxiliary sequences to enhance cleavage and polyadenylation efficiency. The core sequence consists of a highly conserved upstream element (AATAAA or AAUAAA) in the mRNA, referred to as the poly A recognition motif or poly A recognition sequence, which is recognized by the cleavage and polyadenylation specific factor (CPSF); and an undefined downstream region (rich in U or G and U) which is bound by the cleavage stimulating factor (CstF). Examples of usable transcription terminators include, for example, human growth hormone (HGH) polyadenylation signals, simian virus 40 (SV40) late polyadenylation signals, rabbit β-globin polyadenylation signals, bovine growth hormone (BGH) polyadenylation signals, phosphoglycerate kinase (PGK) polyadenylation signals, AOX1 transcription termination sequences, CYC1 transcription termination sequences, or any transcription termination sequence known to be suitable for regulating gene expression in eukaryotic cells. In one example, the polyadenylation signal is the simian virus 40 (SV40) late polyadenylation signal. For example, the polyadenylation signal may contain, consist essentially of, or consist of SEQ ID NO: 615, 169, or 161. For example, the polyadenylation signal may contain, consist essentially of, or consist of SEQ ID NO: 169 or 161. For example, the polyadenylation signal may contain, consist essentially of, or consist of SEQ ID NO: 169. For example, a polyadenylation signal may include, consist substantially of, or consist of SEQ ID NO: 615. In another example, the polyadenylation signal is a bovine growth hormone (BGH) polyadenylation signal or a CpG-depleted BGH polyadenylation signal. For example, a polyadenylation signal may include, consist substantially of, or consist of SEQ ID NO: 162.
[0170] In one example, the polyadenylation signal may include a BGH polyadenylation signal. For example, the BGH polyadenylation signal may include, substantially consist of, or consist of SEQ ID NO: 751. In another example, the polyadenylation signal may include an SV40 polyadenylation signal. For example, the SV40 polyadenylation signal may be a unidirectional SV40 late polyadenylation signal. For example, the transcription terminator sequence present in the “early” reverse direction of SV40 may be mutated (e.g., by mutating the reverse strand AAUAAA sequence to AAUCAA). SV40 polyA is bidirectional, but “late” oriented polyadenylation is more efficient than “early” oriented polyadenylation. For example, a unidirectional SV40 late polyadenylation signal may include, substantially consist of, or consist of SEQ ID NO: 752. In another example, a synthetic polyadenylation signal may be used. For example, a synthetic polyadenylation signal may include, substantially consist of, or consist of SEQ ID NO: 753. In another example, two or more polyadenylation signals may be used in combination. For example, a polyadenylation signal may comprise a combination of a BGH polyadenylation signal and an SV40 polyadenylation signal (e.g., a late SV40 polyadenylation signal, such as a unidirectional late SV40 polyadenylation signal). For example, a polyadenylation signal may comprise a combination of a BGH polyadenylation signal and a unidirectional late SV40 polyadenylation signal. For example, a BGH polyadenylation signal may comprise, substantially constitute, or consist of SEQ ID NO: 751, while a unidirectional late SV40 polyadenylation signal may comprise, substantially constitute, or consist of SEQ ID NO: 752. In a specific example, a BGH polyadenylation signal may be located upstream (5') of an SV40 polyadenylation signal (e.g., a unidirectional late SV40 polyadenylation signal). For example, a combined polyadenylation signal may comprise the sequence shown in SEQ ID NO: 795. In another example, the polyadenylation signal may comprise a combination of a BGH polyadenylation signal and a synthetic polyadenylation signal. For example, the BGH polyadenylation signal may comprise, substantially constitute, or consist of SEQ ID NO: 751, while the synthetic polyadenylation signal may comprise, substantially constitute, or consist of SEQ ID NO: 753. In some embodiments, the nucleic acid construct is a unidirectional construct.
[0171] In some embodiments, the filler sequence can be used to extend the time between RNA polymerase transcription of polyA and its transcription of the next splice acceptor. For example, the filler sequence can be used between two different polyadenylation signals (e.g., between a BGH polyadenylation signal and a synthetic polyadenylation signal). For example, the filler sequence may contain, consist substantially of, or consist of SEQ ID NO: 754.
[0172] In some embodiments, the MAZ element that causes polymerase pausing is used in combination with a polyadenylation signal (e.g., a BGH polyadenylation signal or an SV40 polyadenylation signal). For example, one or more (e.g., at least 1, at least 2, at least 3, at least 4, or about 1 to about 4, about 2 to about 4, about 3 to about 4, or 1, 2, 3, or 4) MAZ elements may be used in combination with a polyadenylation signal. For example, the MAZ element may comprise, substantially consist of, or consist of SEQ ID NO: 755.
[0173] In some embodiments, a unidirectional SV40 late polyadenylation signal is used. SV40 polyA is bidirectional, but polyadenylation in the "late" orientation is more effective than polyadenylation in the "early" orientation. The unidirectional SV40 late polyadenylation signal described herein is localized in the "late" orientation, while polyadenylation signals present in the "early" orientation are mutated or inactivated. In some embodiments, each instance of the sequence AATAAA in the reverse strand is mutated in the unidirectional SV40 late polyadenylation signal. For example, the two conserved AATAAA poly(A) signals present in the SV40 "early" poly(A) are mutated to AATCATA. In some embodiments, the unidirectional SV40 late polyadenylation signal is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in SEQ ID NO: 752. In some embodiments, the unidirectional SV40 late polyadenylation signal comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 752.
[0174] The unidirectional SV40 late polyadenylation signal can be used in combination (e.g., in tandem) with one or more additional polyadenylation signals. Examples of transcription terminators that can be used include, for example, human growth hormone (HGH) polyadenylation signals, simian virus 40 (SV40) late polyadenylation signals, rabbit β-globin polyadenylation signals, bovine growth hormone (BGH) polyadenylation signals, phosphoglycerate kinase (PGK) polyadenylation signals, AOX1 transcription termination sequences, CYC1 transcription termination sequences, or any transcription termination sequence known to be suitable for regulating gene expression in eukaryotic cells. For example, the unidirectional SV40 late polyadenylation signal can be used in combination (e.g., in tandem) with the bovine growth hormone (BGH) polyadenylation signal, optionally wherein the BGH polyadenylation signal is upstream (5') of the unidirectional SV40 late polyadenylation signal. In some embodiments, the BGH polyadenylation signal is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in SEQ ID NO: 751. In some embodiments, the BGH polyadenylation signal comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 751. In some embodiments, the combination of the BGH polyadenylation signal and the unidirectional SV40 late polyadenylation signal is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in SEQ ID NO: 795. In some embodiments, the combination of the BGH polyadenylation signal and the unidirectional SV40 late polyadenylation signal comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 795.
[0175] In some embodiments, the filler sequence can be used to extend the time between RNA polymerase transcription of polyA and its transcription of the next splice acceptor. For example, the filler sequence can be used between two different polyadenylation signals (e.g., between a BGH polyadenylation signal and a synthetic polyadenylation signal). For example, the filler sequence may contain, consist substantially of, or consist of SEQ ID NO: 754.
[0176] In some embodiments, the MAZ element that causes polymerase pausing is used in combination with a polyadenylation signal (e.g., a BGH polyadenylation signal or an SV40 polyadenylation signal). For example, one or more (e.g., at least 1, at least 2, at least 3, at least 4, or about 1 to about 4, about 2 to about 4, about 3 to about 4, or 1, 2, 3, or 4) MAZ elements may be used in combination with a polyadenylation signal. For example, the MAZ element may comprise, substantially consist of, or consist of SEQ ID NO: 755.
[0177] The constructs disclosed herein may also include a splice acceptor site (e.g., operatively linked to a multi-domain therapeutic protein-coding sequence, such as upstream or 5' of a multi-domain therapeutic protein-coding sequence). The splice acceptor site may, for example, contain or consist of NAGs. In a specific example, the splice acceptor is ALB Scissor acceptor (e.g., used to cut) ALB Exon 1 and exon 2 were spliced together. ALB Scissor acceptor (i.e., ALB Exon 2 scissor acceptor). For example, this scissor acceptor can be derived from humans. ALB Gene. In another example, the splice acceptor can be derived from mice. Alb Genes (e.g., used to modify mice) Alb Exon 1 and exon 2 were spliced together. ALB Scission receptor (i.e., mouse) Alb Exon 2 splice acceptor). In another example, the splice acceptor is a splice acceptor from a gene encoding the target polypeptide (e.g., the GAA splice acceptor). For example, such a splice acceptor could be derived from human... GAA Gene. Alternatively, this splice acceptor can be derived from mice. GAA Genes. Other suitable scission acceptor sites (including artificial scission acceptors) that can be used in eukaryotes are well known. See, for example, Shapiro et al. (1987). Nucleic Acids Res .15:7155-7174, and Burset et al. (2001) Nucleic Acids Res .29:255-259, each of these references is incorporated herein by reference in its entirety for all purposes. In a specific example, the scissor acceptor is the mouse. Alb Exon 2 splicing acceptor. In specific examples, the splicing acceptor may contain, consist substantially of, or consist of SEQ ID NO:163.
[0178] In some examples, the nucleic acid constructs disclosed herein may be bidirectional constructs, which are described in more detail below. In some examples, the nucleic acid constructs disclosed herein may be unidirectional constructs, which are described in more detail below. Similarly, in some examples, the nucleic acid constructs disclosed herein may be contained in vectors (e.g., viral vectors, such as AAV or rAAV8) and / or lipid nanoparticles, as described in more detail elsewhere herein.
[0179] (1) Multi-domain therapeutic proteins Multidomain therapeutic proteins as described herein include lysosomal α-glucosidase peptides (GAA or its bioactive portion, to provide GAA enzyme alternative activity) linked or fused to a TfR-binding delivery domain or a CD63-binding delivery domain. The TfR-binding domain, the CD63-binding delivery domain, and the GAA peptide are described in more detail below. Examples of multidomain therapeutic proteins can be found in WO 2013 / 138400, WO 2017 / 007796, WO 2017 / 190079, WO2017 / 100467, WO 2018 / 226861, WO 2019 / 157224, and WO 2019 / 222663, each of which is incorporated herein by reference in its entirety for all purposes. For example, the multidomain therapeutic proteins described herein may include a TfR-binding delivery domain linked or fused to a GAA peptide. The TfR-binding domain provides binding to the internalizing factor TfR. Multidomain therapeutic proteins produced by the liver target muscle and the CNS by targeting TfR, which is expressed in muscle and on brain endothelial cells. Endocytotic transport of TfR in these cells enables crossing the blood-brain barrier. For example, the multidomain therapeutic proteins described herein may include a CD63-binding delivery domain linked to or fused with a GAA peptide. The CD63-binding domain provides binding to the internalizing factor CD63. Multidomain therapeutic proteins target muscle by targeting CD63, a rapidly internalizing protein highly expressed in muscle. In some multidomain therapeutic proteins, the delivery domain is covalently linked to the GAA. Covalent linkage can be any type of covalent bond (i.e., any bond involving electron sharing). In some cases, the covalent bond is a peptide bond between two amino acids, causing the GAA and the delivery domain to form a continuous polypeptide chain, either whole or partially, as in fusion proteins. In some cases, the GAA portion and the delivery domain portion are directly linked. In other cases, linkers (such as peptide linkers) are used to tether the two parts. Any suitable linker can be used. See Chen et al., “Fusion protein linkers: property, design and functionality,” 65(10) Adv DrugDeliv Rev. 1357-69 (2013). In some cases, cleavable linkers are used. For example, a cathepsin-cleavable linker can be inserted between the delivery domain and the GAA to facilitate the removal of the delivery domain from the lysosome. In another example, the linker may contain an amino acid sequence, such as one, two, three, four, five, six, seven, eight, eight, or ten repeats of Gly4Ser (SEQ ID NO:537), which is about 10 amino acids long.In one example, the linker comprises three such repeating sequences (SEQ ID NO: 616), substantially consisting of or consisting of them. For example, the coding sequence of the linker may comprise any one of SEQ ID NO: 618-622 and 747, substantially consisting of or consisting of them. In another example, the linker comprises two such repeating sequences (SEQ ID NO: 617), substantially consisting of or consisting of them. For example, the coding sequence of the linker may comprise any one of SEQ ID NO: 623-629, substantially consisting of or consisting of them. In yet another example, the linker comprises one such repeating sequence (SEQ ID NO: 537), substantially consisting of or consisting of it. For example, the coding sequence of the linker may comprise SEQ ID NO: 630 or 748, substantially consisting of or consisting of it. In another example, a rigid linker, such as the 2XH4 linker, may be used. In one example, the connector contains, is substantially composed of, or is composed of AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA (SEQ ID NO: 842). For example, the encoded sequence of the connector may contain SEQ ID NO: 841, is substantially composed of, or is composed of.
[0180] In a specific multidomain therapeutic protein, the GAA (e.g., N-terminus) is covalently linked to the C-terminus of the heavy chain or light chain of an anti-TfR or anti-CD63 antibody (i.e., the multidomain therapeutic protein is in the form of anti-TfR:GAA or anti-CD63:GAA from the N-terminus to the C-terminus). In another specific multidomain therapeutic protein, the GAA is covalently linked to the N-terminus of the heavy chain or light chain of an anti-TfR or anti-CD63 antibody (i.e., the multidomain therapeutic protein is in the form of GAA:anti-TfR or GAA:anti-CD63 from the N-terminus to the C-terminus). In yet another specific embodiment, the GAA (e.g., N-terminus) is linked to the C-terminus of the anti-TfR or anti-CD63 scFv domain (i.e., the multidomain therapeutic protein is in the form of anti-TfR-scFv:GAA or anti-CD63-scFv:GAA from the N-terminus to the C-terminus, such as anti-TfR-scFv(V L V H ):GAA or anti-CD63-scFv(V L V HIn another specific embodiment, the GAA (e.g., the N-terminus) is attached to the C-terminus of the anti-TfR or anti-CD63 Fab heavy chain (i.e., the multidomain therapeutic protein is in the form of anti-TfR-Fab (light chain / heavy chain):GAA or anti-CD63-Fab (light chain / heavy chain):GAA from the N-terminus to the C-terminus). In another specific embodiment, the GAA (e.g., the N-terminus) is attached to the C-terminus of the anti-TfR or anti-CD63 Fab light chain (i.e., the multidomain therapeutic protein is in the form of anti-TfR-Fab (heavy chain / light chain):GAA or anti-CD63-Fab (heavy chain / light chain):GAA from the N-terminus to the C-terminus).
[0181] (a) Lysosomal α-glucosidase (GAA) Lysosomal α-glucosidase (GAA; also known as acid α-glucosidase, acid α-glucosidase proproteinogen, acid maltase, α-glucosidase, α-1,4-glucosidase, amylase, glucoamylase, LYAG) is produced by... GAA Encoded. This enzyme is active in lysosomes, where it breaks down glycogen into glucose.
[0182] people GAA The gene (NCBI GeneID 2548) encodes a protein of 952 amino acids. In lysosomes, human GAA is subsequently processed by proteases into peptides of 76 kDa, 19.4 kDa, and 3.9 kDa that maintain association. Further cleavage between R(200) and A(204) does not efficiently convert the 76 kDa peptide into the mature 70 kDa form with an additional 10.4 kDa peptide. GAA maturation increases its affinity for glycogen by 7-10 times. The signal peptide is encoded by amino acids 1-27, the propeptide by amino acids 28-69, the lysosomal α-glucosidase after removing the signal peptide and propeptide is encoded by amino acids 70-952, the 76 kDa lysosomal α-glucosidase by amino acids 123-952, and the 70 kDa lysosomal α-glucosidase by amino acids 204-952.
[0183] The GAA expressed by the compositions and methods disclosed herein can be any wild-type or variant GAA. In one example, the GAA is a human GAA protein. Human GAA is designated as UniProt reference number P10253. An exemplary amino acid sequence of human GAA is designated as NCBI accession number NP_000143.2 and is shown in SEQ ID NO: 724. GAA The mRNA (cDNA) sequence is designated as NCBI accession number NM_000152.5 and is shown in SEQ ID NO: 725. Exemplary human GAAThe coding sequence is designated as CCDS ID CCDS32760.1 and is shown in SEQ ID NO: 726. An exemplary mature human GAA amino acid sequence starting at amino acid 70 (i.e., the human GAA sequence after removing the signal peptide and propeptide) (i.e., GAA 70-952) is shown in SEQ ID NO: 727. An exemplary coding sequence of GAA 70-952 is shown in SEQ ID NO: 728.
[0184] In some examples, the GAA (e.g., human GAA) is a wild-type GAA (e.g., wild-type human GAA) sequence or a fragment thereof. For example, the GAA may be a fragment containing the mature GAA amino acid sequence (i.e., the GAA sequence after removing the signal peptide and propeptide), a fragment containing a GAA in the form of 77 kDa, or a fragment containing a GAA in the form of 70 kDa. In a specific example, the GAA may contain SEQ ID NO: 727, or may be at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to SEQ ID NO: 727. In another specific example, the GAA may consist substantially of SEQ ID NO: 727. In yet another specific example, the GAA may consist of SEQ ID NO: 727.
[0185] The GAA coding sequence in the constructs disclosed herein may contain one or more modifications, such as codon optimization (e.g., optimization of human codons), deletion of CpG dinucleotides, mutation of cryptic splicing sites, addition of one or more glycosylation sites, or any combination thereof. CpG dinucleotides in the constructs can limit the therapeutic efficacy of the construct. First, unmethylated CpG dinucleotides can interact with the host toll-like receptor-9 (TLR-9) to stimulate an innate pro-inflammatory immune response. Second, once CpG dinucleotides are methylated, they can lead to inhibition of transgene expression coordinated by methyl-CpG binding proteins. Cryptic splicing sites are sequences in premessenger RNA that are not normally used as splicing sites but can be activated, for example, by inactivating typical splicing sites or by forming mutations in splicing sites that were not previously present. Accurate splicing site selection is crucial for successful gene expression, and removal of cryptic splicing sites can facilitate the use of normal or intended splicing sites.
[0186] In one example, the GAA coding sequence in the construct disclosed herein has been mutated or one or more cryptic splicing sites have been removed. In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”. In some embodiments, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”. In some embodiments, the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”. In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”, and the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”. In another example, the GAA coding sequence in the construct disclosed herein has been mutated or all identified cryptic splicing sites have been removed. In another example, the GAA coding sequence in the construct disclosed herein has had one or more CpG dinucleotides removed (i.e., it is CpG depleted). In another example, the GAA coding sequence in the construct disclosed herein has all CpG dinucleotides removed (i.e., it is completely CpG depleted). In another example, the GAA coding sequence in the construct disclosed herein is codon-optimized (e.g., codon-optimized for expression in humans or mammals). In a specific example, the GAA coding sequence in the construct disclosed herein has one or more CpG dinucleotides removed (i.e., it is CpG depleted) and has been mutated or one or more cryptic splicing sites removed. In another specific example, the GAA coding sequence in the construct disclosed herein has all CpG dinucleotides removed and has been mutated or one or more identified cryptic splicing sites removed. In another specific example, the GAA coding sequence in the construct disclosed herein has had one or more CpG dinucleotides removed (i.e., is CpG depleted) and is codon-optimized (e.g., codon-optimized for expression in humans or mammals). In yet another specific example, the GAA coding sequence in the construct disclosed herein has all CpG dinucleotides removed (i.e., is completely CpG depleted) and is codon-optimized (e.g., codon-optimized for expression in humans or mammals).
[0187] Various codon-optimized GAA coding sequences are provided. The GAA coding sequence can be, for example, CpG-depleted (e.g., fully CpG-depleted) and / or codon-optimized (e.g., CpG-depleted (e.g., fully CpG-depleted) and codon-optimized). In one example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to any of SEQ ID NO: 750, 749, and 649. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to any of SEQ ID NO: 750, 749, and 649. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to any of SEQ ID NO: 750, 749, and 649. In another example, the GAA coding sequence comprises the sequence shown in any of SEQ ID NO: 750, 749, and 649. In another example, the GAA coding sequence consists substantially of the sequence shown in any of SEQ ID NO: 750, 749, and 649. In another example, the GAA coding sequence consists of the sequence shown in any of SEQ ID NO: 750, 749, and 649. Optionally, the GAA coding sequence encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to (and, for example, retains the activity of natural GAA) of SEQ ID NO: 727. Optionally, the GAA-coding sequence encodes a GAA protein (or a GAA protein containing that sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA-coding sequence in the above examples encodes a GAA protein (or a GAA protein containing that sequence) that is at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA-coding sequence in the above examples encodes a GAA protein containing the sequence shown in SEQ ID NO: 727. Optionally, the GAA-coding sequence in the above examples encodes a GAA protein that is substantially composed of the sequence shown in SEQ ID NO: 727.Optionally, the GAA coding sequence in the above examples encodes a GAA protein consisting of the sequence shown in SEQ ID NO: 727. In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”. In some embodiments, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”. In some embodiments, the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”. In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”, and the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”.
[0188] In one example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 750. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 750, and encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 750, and encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 750. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 750, and encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 750, and encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 750. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 750, and encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727.In another example, the GAA coding sequence is (or comprises) at least 99%, at least 99.5%, or 100% identical to the sequence in SEQ ID NO: 750, and encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. In another example, the GAA coding sequence comprises the sequence shown in SEQ ID NO: 750. In another example, the GAA coding sequence consists substantially of the sequence shown in SEQ ID NO: 750. In another example, the GAA coding sequence consists of the sequence shown in SEQ ID NO: 750. The GAA coding sequence can be, for example, CpG depleted (e.g., fully CpG depleted) and / or codon-optimized. For example, the GAA coding sequence can be CpG depleted (e.g., fully CpG depleted) and codon-optimized. Optionally, the GAA coding sequence encodes a GAA protein (or a GAA protein containing the sequence) that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA coding sequence encodes a GAA protein (or a GAA protein containing the sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA coding sequence in the above examples encodes a GAA protein (or a GAA protein containing the sequence) that is at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA coding sequence in the above examples encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. Optionally, the GAA coding sequence in the above examples encodes a GAA protein consisting substantially of the sequence shown in SEQ ID NO: 727. Optionally, the GAA coding sequence in the above examples encodes a GAA protein consisting of the sequence shown in SEQ ID NO: 727. In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”. In some embodiments, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”. In some embodiments, the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”.In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”, and the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”.
[0189] In one example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 749. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727, and encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 749, and encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 749. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 749, and encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 749, and encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 749. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 749, and encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727.In another example, the GAA coding sequence is (or comprises) at least 99%, at least 99.5%, or 100% identical to the sequence in SEQ ID NO: 749, and encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. In another example, the GAA coding sequence comprises the sequence shown in SEQ ID NO: 749. In another example, the GAA coding sequence consists substantially of the sequence shown in SEQ ID NO: 749. In another example, the GAA coding sequence consists of the sequence shown in SEQ ID NO: 749. The GAA coding sequence can be, for example, CpG depleted (e.g., fully CpG depleted) and / or codon-optimized. For example, the GAA coding sequence can be CpG depleted (e.g., fully CpG depleted) and codon-optimized. Optionally, the GAA coding sequence encodes a GAA protein (or a GAA protein containing the sequence) that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA coding sequence encodes a GAA protein (or a GAA protein containing the sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA coding sequence in the above examples encodes a GAA protein (or a GAA protein containing the sequence) that is at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA coding sequence in the above examples encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. Optionally, the GAA coding sequence in the above examples encodes a GAA protein consisting substantially of the sequence shown in SEQ ID NO: 727. Optionally, the GAA coding sequence in the above examples encodes a GAA protein consisting of the sequence shown in SEQ ID NO: 727. In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”. In some embodiments, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”. In some embodiments, the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”.In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”, and the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”.
[0190] In one example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 649. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727, and encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 649, and encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 649. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 649, and encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 649, and encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 649. In another example, the GAA coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 649, and encodes a GAA protein (or a GAA protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 727.In another example, the GAA coding sequence is (or comprises) at least 99%, at least 99.5%, or 100% identical to the sequence in SEQ ID NO: 649, and encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. In another example, the GAA coding sequence comprises the sequence shown in SEQ ID NO: 649. In another example, the GAA coding sequence consists substantially of the sequence shown in SEQ ID NO: 649. In another example, the GAA coding sequence consists of the sequence shown in SEQ ID NO: 649. The GAA coding sequence can be, for example, CpG depleted (e.g., fully CpG depleted) and / or codon-optimized. For example, the GAA coding sequence can be CpG depleted (e.g., fully CpG depleted) and codon-optimized. Optionally, the GAA coding sequence encodes a GAA protein (or a GAA protein containing the sequence) that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA coding sequence encodes a GAA protein (or a GAA protein containing the sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA coding sequence in the above examples encodes a GAA protein (or a GAA protein containing the sequence) that is at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 727 (and, for example, retains the activity of natural GAA). Optionally, the GAA coding sequence in the above examples encodes a GAA protein comprising the sequence shown in SEQ ID NO: 727. Optionally, the GAA coding sequence in the above examples encodes a GAA protein consisting substantially of the sequence shown in SEQ ID NO: 727. Optionally, the GAA coding sequence in the above examples encodes a GAA protein consisting of the sequence shown in SEQ ID NO: 727. In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”. In some embodiments, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”. In some embodiments, the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”.In some embodiments, the nucleotide at position 1095 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”, the nucleotide at position 1098 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “C”, and the nucleotide at position 2343 (or the corresponding position when the GAA coding sequence is aligned with SEQ ID NO: 750) is “G”.
[0191] Various other codon-optimized GAA coding sequences are provided. GAA coding sequences can be, for example, CpG-depleted (e.g., fully CpG-depleted) and / or codon-optimized (e.g., CpG-depleted (e.g., fully CpG-depleted) and codon-optimized).
[0192] When this article discloses specific GAA When referring to a nucleic acid construct sequence of a multi-domain therapeutic protein, it means encompassing the disclosed sequence or its inverse complementary sequence. For example, if the sequence disclosed herein... GAA If a multi-domain therapeutic protein nucleic acid construct consists of the hypothetical sequence 5'-CTGGACCGA-3', it also implies the inverse complementary sequence (5'-TCGGTCCAG-3') covering that sequence. Similarly, when construct elements are disclosed herein in a specific 5' to 3' order, it also implies the inverse complementary sequence covering the order of those elements. One reason for this is that in many embodiments disclosed herein, GAA Multidomain therapeutic protein nucleic acid constructs are part of single-stranded recombinant AAV vectors. Single-stranded AAV genomes are packaged as sense (positive strand) or antisense (negative strand) genomes, and single-stranded AAV genomes of both + and - polarity are packaged into mature rAAV virions at equal frequencies. See, for example, LING et al. (2015). J. Mol.Genet.Med. 9(3):175, Zhou et al. (2008) Mol.Ther. 16(3):494-499, and Samulski et al. (1987) J. Virol. References 61:3096-3101, each of which is incorporated herein by reference in its entirety for all purposes.
[0193] (b) Delivery domain combined with CD63 The multi-domain therapeutic protein disclosed herein may include a CD63-binding delivery domain fused to a GAA peptide. The CD63-binding domain provides binding to the internalizing factor CD63 (UniProt Ref. P08962-1). CD63 (also known as CD63 antigen, granulolysin, lysosome-associated membrane protein 3, LAMP-3, lysosome-integrated membrane protein 1, Limp1, melanoma-associated antigen ME491, OMA81H, ocular melanoma-associated antigen, four-transmembrane protein-30, or Tspan-30) is a member of the four-transmembrane protein superfamily of cell surface proteins that cross the cell membrane four times. It is composed of… CD63 Genes (also known as genes) MLA1 or TSPAN30 CD63 is encoded by [the protein name is missing here]. It is expressed in almost all tissues and is believed to be involved in the formation and stabilization of signal transduction complexes. CD63 is located in the cell membrane, lysosomal membrane, and late endosome membrane. CD63 is known to associate with integrins and may be involved in the epithelial-mesenchymal transition.
[0194] In some multidomain therapeutic proteins, the CD63-binding delivery domain is an antibody, antibody fragment, or other antigen-binding protein. Examples of antigen-binding proteins include, for instance, receptor-fusion molecules, capture molecules, receptor-Fc fusion molecules, antibodies, Fab fragments, F(ab')2 fragments, Fd fragments, Fv fragments, single-chain Fv (scFv) molecules, dAb fragments, separated complementarity-determining regions (CDRs), CDR3 peptides, restricted FR3-CDR3-FR4 peptides, domain-specific antibodies, single-domain antibodies, domain-deficient antibodies, chimeric antibodies, CDR-transplanted antibodies, biantibodies, triantibodies, tetraantibodies, microantibodies, nanobodies, monovalent nanobodies, divalent nanobodies, small modular immunopharmaceuticals (SMIPs), camelid antibodies (VHH heavy chain homodimer antibodies), and shark variable IgNAR domains. Examples of delivery domains incorporating CD63 can be found in WO 2013 / 138400, WO 2017 / 007796, WO 2017 / 190079, WO 2017 / 100467, WO 2018 / 226861, WO 2019 / 157224 and WO 2019 / 222663, each of which is incorporated herein by reference in its entirety for all purposes.
[0195] In certain multi-domain therapeutic proteins, the CD63-binding delivery domain is the anti-CD63 scFv. In a specific example, the anti-CD63 scFv may comprise SEQ ID NO: 730, or may be at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to SEQ ID NO: 730. In another specific example, the anti-CD63 scFv may consist substantially of SEQ ID NO: 730. In yet another specific example, the anti-CD63 scFv may consist of SEQ ID NO: 730.
[0196] The CD63-binding delivery domain in the constructs disclosed herein may contain one or more modifications, such as codon optimization (e.g., optimization of human codons), deletion of CpG dinucleotides, mutation of cryptic splicing sites, addition of one or more glycosylation sites, or any combination thereof. CpG dinucleotides in the constructs can limit the therapeutic efficacy of the construct. First, unmethylated CpG dinucleotides can interact with the host toll-like receptor-9 (TLR-9) to stimulate an innate pro-inflammatory immune response. Second, once CpG dinucleotides are methylated, they can lead to inhibition of transgene expression coordinated by methyl-CpG-binding proteins. Cryptic splicing sites are sequences in premessenger RNA that are not normally used as splicing sites but can be activated, for example, by inactivating typical splicing sites or by forming mutations in previously absent splicing sites. Accurate splicing site selection is crucial for successful gene expression, and removal of cryptic splicing sites can facilitate the use of normal or intended splicing sites.
[0197] In one example, the CD63-binding delivery domain coding sequence in the constructed organism disclosed herein has been mutated or has one or more hidden splicing sites removed. In another example, the CD63-binding delivery domain coding sequence in the constructed organism disclosed herein has been mutated or has all identified hidden splicing sites removed. In another example, the CD63-binding delivery domain coding sequence in the constructed organism disclosed herein has one or more CpG dinucleotides removed (i.e., it is CpG depleted). In another example, the CD63-binding delivery domain coding sequence in the constructed organism disclosed herein has all CpG dinucleotides removed (i.e., it is completely CpG depleted). In another example, the CD63-binding delivery domain coding sequence in the constructed organism disclosed herein is codon-optimized (e.g., codon-optimized for expression in humans or mammals). In a specific example, the CD63-binding delivery domain coding sequence in the constructed organism disclosed herein has one or more CpG dinucleotides removed (i.e., it is CpG depleted) and has been mutated or has one or more hidden splicing sites removed. In another specific example, the CD63-binding delivery domain coding sequence in the construct disclosed herein has all CpG dinucleotides removed and has been mutated or removed one or more of the identified cryptic splicing sites. In another specific example, the CD63-binding delivery domain coding sequence in the construct disclosed herein has one or more CpG dinucleotides removed (i.e., is CpG depleted) and is codon-optimized (e.g., codon-optimized for expression in humans or mammals). In another specific example, the CD63-binding delivery domain coding sequence in the construct disclosed herein has all CpG dinucleotides removed (i.e., is completely CpG depleted) and is codon-optimized (e.g., codon-optimized for expression in humans or mammals).
[0198] Various anti-CD63 scFv coding sequences are provided. In one example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to any one of SEQ ID NO: 759, 760, and 732. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to any one of SEQ ID NO: 759, 760, and 732. In yet another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to any one of SEQ ID NO: 759, 760, and 732. In another example, the anti-CD63 scFv coding sequence comprises the sequence shown in any of SEQ ID NO: 759, 760, and 732. In another example, the anti-CD63 scFv coding sequence consists substantially of the sequence shown in any of SEQ ID NO: 759, 760, and 732. In another example, the anti-CD63 scFv coding sequence consists of the sequence shown in any of SEQ ID NO: 759, 760, and 732. In one example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence comprises the sequence shown in SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence consists essentially of the sequence shown in SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence consists of the sequence shown in SEQ ID NO: 759.Optionally, the anti-CD63 scFv coding sequence encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein containing the sequence) that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to (and, for example, retains CD63 binding activity) of SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein containing the sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to (and, for example, retains CD63 binding activity) of SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above examples encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein containing that sequence) that is at least 99%, at least 99.5%, or 100% identical to (and for example retains CD63 binding activity) of SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above examples encodes an anti-CD63 scFv protein comprising substantially the sequence shown in SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above examples encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In some embodiments, the nucleotide at position 3 (or the corresponding position when the anti-CD63 scFv coding sequence is compared with SEQ ID NO: 759) is “A”. In some embodiments, the nucleotide at position 132 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A". In some embodiments, the nucleotide at position 273 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "T". In some embodiments, the nucleotide at position 3 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A", the nucleotide at position 132 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A", and the nucleotide at position 273 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "T".
[0199] In one example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 759, and encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence is (or comprises) at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to the sequence shown in SEQ ID NO: 759, and encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence is (or comprises) at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to the sequence shown in SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 759, and encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 759, and encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence is (or contains) a sequence that is at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 759.In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 759, and encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 759, and encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence comprises the sequence shown in SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence consists essentially of the sequence shown in SEQ ID NO: 759. In another example, the anti-CD63 scFv coding sequence consists of the sequence shown in SEQ ID NO: 759. The anti-CD63 scFv coding sequence may be, for example, CpG depleted (e.g., completely CpG depleted) and / or codon-optimized. For example, the anti-CD63 scFv coding sequence may be CpG depleted (e.g., completely CpG depleted) and codon-optimized. Optionally, the anti-CD63 scFv coding sequence encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein containing the sequence) that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical (and, for example, retains CD63 binding activity) to the sequence in SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to (and for example, retains CD63 binding activity) of SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above example encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to (and for example, retains CD63 binding activity) of SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above example encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730.Optionally, the anti-CD63 scFv coding sequence in the above examples encodes an anti-CD63 scFv protein consisting substantially of the sequence shown in SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above examples encodes an anti-CD63 scFv protein consisting of the sequence shown in SEQ ID NO: 730. In some embodiments, the nucleotide at position 3 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A". In some embodiments, the nucleotide at position 132 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A". In some embodiments, the nucleotide at position 273 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "T". In some implementations, the nucleotide at position 3 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A", the nucleotide at position 132 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A", and the nucleotide at position 273 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "T".
[0200] In one example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 760. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 760, and encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence is (or comprises) at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to the sequence shown in SEQ ID NO: 760, and encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence is (or comprises) at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to the sequence shown in SEQ ID NO: 760. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 760, and encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 760, and encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence is (or contains) a sequence that is at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 760.In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 760, and encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 760. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 760, and encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence comprises the sequence shown in SEQ ID NO: 760. In another example, the anti-CD63 scFv coding sequence consists essentially of the sequence shown in SEQ ID NO: 760. In another example, the anti-CD63 scFv coding sequence consists of the sequence shown in SEQ ID NO: 760. The anti-CD63 scFv coding sequence may be, for example, CpG depleted (e.g., completely CpG depleted) and / or codon-optimized. For example, the anti-CD63 scFv coding sequence may be CpG depleted (e.g., completely CpG depleted) and codon-optimized. Optionally, the anti-CD63 scFv coding sequence encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein containing the sequence) that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical (and, for example, retains CD63 binding activity) to the sequence in SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to (and for example, retains CD63 binding activity) of SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above example encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to (and for example, retains CD63 binding activity) of SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above example encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730.Optionally, the anti-CD63 scFv coding sequence in the above examples encodes an anti-CD63 scFv protein consisting substantially of the sequence shown in SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above examples encodes an anti-CD63 scFv protein consisting of the sequence shown in SEQ ID NO: 730. In some embodiments, the nucleotide at position 3 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A". In some embodiments, the nucleotide at position 132 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A". In some embodiments, the nucleotide at position 273 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "T". In some implementations, the nucleotide at position 3 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A", the nucleotide at position 132 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A", and the nucleotide at position 273 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "T".
[0201] In one example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 732. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 732, and encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence is (or comprises) at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to the sequence shown in SEQ ID NO: 732, and encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence is (or comprises) at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to the sequence shown in SEQ ID NO: 732. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 732, and encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 732. In yet another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 732, and encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence is (or contains) a sequence that is at least 99%, at least 99.5%, or 100% identical to SEQ ID NO: 732.In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 732, and encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 732. In another example, the anti-CD63 scFv coding sequence is (or comprises) a sequence that is at least 99%, at least 99.5%, or 100% identical to that of SEQ ID NO: 732, and encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730. In another example, the anti-CD63 scFv coding sequence comprises the sequence shown in SEQ ID NO: 732. In another example, the anti-CD63 scFv coding sequence consists essentially of the sequence shown in SEQ ID NO: 732. In another example, the anti-CD63 scFv coding sequence consists of the sequence shown in SEQ ID NO: 732. The anti-CD63 scFv coding sequence may be, for example, CpG depleted (e.g., completely CpG depleted) and / or codon-optimized. For example, the anti-CD63 scFv coding sequence may be CpG depleted (e.g., completely CpG depleted) and codon-optimized. Optionally, the anti-CD63 scFv coding sequence encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein containing the sequence) that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical (and, for example, retains CD63 binding activity) to SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to (and for example, retains CD63 binding activity) of SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above example encodes an anti-CD63 scFv protein (or an anti-CD63 scFv protein comprising that sequence) that is at least 99%, at least 99.5%, or 100% identical to (and for example, retains CD63 binding activity) of SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above example encodes an anti-CD63 scFv protein comprising the sequence shown in SEQ ID NO: 730.Optionally, the anti-CD63 scFv coding sequence in the above examples encodes an anti-CD63 scFv protein consisting substantially of the sequence shown in SEQ ID NO: 730. Optionally, the anti-CD63 scFv coding sequence in the above examples encodes an anti-CD63 scFv protein consisting of the sequence shown in SEQ ID NO: 730. In some embodiments, the nucleotide at position 3 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A". In some embodiments, the nucleotide at position 132 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A". In some embodiments, the nucleotide at position 273 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "T". In some implementations, the nucleotide at position 3 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A", the nucleotide at position 132 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "A", and the nucleotide at position 273 (or the corresponding position when the anti-CD63 scFv coding sequence is aligned with SEQ ID NO: 759) is "T".
[0202] When a specific anti-CD63 scFv or multidomain therapeutic protein nucleic acid construct sequence is disclosed herein, it means that the disclosed sequence or its reverse complement is covered. For example, if the anti-CD63 scFv or multidomain therapeutic protein nucleic acid construct disclosed herein consists of the hypothetical sequence 5'-CTGGACCGA-3', it also means that the reverse complement of that sequence (5'-TCGGTCCAG-3') is covered. Similarly, when construct elements are disclosed herein in a specific 5' to 3' order, it also means that the reverse complement of the order of those elements is covered. One reason for this is that in many embodiments disclosed herein, the anti-CD63 scFv or multidomain therapeutic protein nucleic acid construct is part of a single-stranded recombinant AAV vector. The single-stranded AAV genome is packaged as sense (positive strand) or antisense (negative strand) genomes, and the + and - polar single-stranded AAV genomes are packaged into mature rAAV virions at equal frequencies. See, for example, LING et al. (2015). J. Mol.Genet.Med. 9(3):175, Zhou et al. (2008) Mol.Ther. 16(3):494-499, and Samulski et al. (1987) J. Virol.References 61:3096-3101, each of which is incorporated herein by reference in its entirety for all purposes.
[0203] (c) Delivery domain combined with TfR The multi-domain therapeutic proteins disclosed herein may include a TfR-binding delivery domain fused to a GAA peptide. The TfR-binding domain provides binding to the internalizing factor transferrin receptor protein 1 (TfR; UniProt Ref. P02786). TfR (also known as TR, TfR1, and Trfr) is composed of… TFRC Genetically encoded. TfR is expressed in muscle and on brain endothelial cells. Endocytotic transport of TfR in these cells enables crossing the blood-brain barrier. In some embodiments, multi-domain therapeutic proteins comprising a TfR-binding delivery domain fused to a GAA peptide (e.g., scFv) do not alter transferrin uptake. In some embodiments, multi-domain therapeutic proteins comprising a TfR-binding delivery domain fused to a GAA peptide (e.g., scFv) do not alter iron homeostasis. In some embodiments, multi-domain therapeutic proteins comprising a TfR-binding delivery domain fused to a GAA peptide (e.g., scFv) do not alter transferrin uptake or iron homeostasis.
[0204] Transferrin receptor 1 (TfR) is a membrane receptor involved in controlling iron supply to cells via the binding of transferrin (the main iron carrier protein). Transferrin receptor 1 is composed of… TFRCGene expression. Transferrin receptor 1 may be referred to herein as TFRC. This receptor plays a crucial role in the control of cell proliferation because iron is essential for maintaining the activity of ribonucleotide reductase, and is the only enzyme that catalyzes the conversion of ribonucleotides to deoxyribonucleotides. Preferably, the TfR is human TfR (hTfR). See, for example, accession numbers NP_001121620.1; BAD92491.1 and NP_001300894.1; and e!Ensembl entry: ENSG00000072274. Human transferrin receptor 1 is expressed in several tissues, including but not limited to: cerebral cortex, cerebellum, hippocampus, caudate nucleus, parathyroid glands, adrenal glands, bronchi, lungs, oral mucosa, esophagus, stomach, duodenum, small intestine, colon, rectum, liver, gallbladder, pancreas, kidneys, bladder, testes, epididymis, prostate, vagina, ovaries, fallopian tubes; endometrium, cervix, placenta, breast, myocardium, smooth muscle, soft tissue, skin, appendix, lymph nodes, tonsils, and bone marrow. The associated transferrin receptor is transferrin receptor 2 (TfR2). Human transferrin receptor 2 shares approximately 45% sequence identity with human transferrin receptor 1. Trinder and Baker, Transferrin receptor 2: a new molecule iniron metabolism. Int J Biochem Cell Biol. March 2003; 35(3):292-6. Unless otherwise stated, the transferrin receptor as used herein generally refers to transferrin receptor 1 (e.g., human transferrin receptor 1).
[0205] Human transferrin (Tf) is a single-chain 80 kDa member of the anion-binding protein superfamily. Transferrin is a 698-amino acid precursor, broken down into a 19-amino acid signal sequence followed by a 679-amino acid mature segment typically containing 19 intrachain disulfide bonds. N-terminal and C-terminal side-binding regions (or domains) bind ferric iron via interactions with specific anions (e.g., bicarbonate) and four amino acids (His, Asp, and two Tyr). Iron-deferrotransferrin initially binds an iron atom at the C-terminus, followed by subsequent iron binding through the N-terminus to form all-ferrotransferrin (diferrotransferrin, all-ferrotransferrin). Through its C-terminal iron-binding domain, all-ferrotransferrin interacts with TfR on the cell surface, where it is internalized into acidified endosomes. Iron dissociates from the Tf molecules within these endosomes and is transported as ferrous iron into the cytosol. In addition to TfR, transferrin has also been reported to bind to cupulin, IGFBP3, microbial iron-binding protein, and liver-specific TfR2.
[0206] The blood-brain barrier (BBB) is located within the microvascular system of the brain and regulates the pathways through which molecules enter the brain from the blood. Burkhart et al., Accessing targeted nanoparticles to the brain: the vascularroute. Curr Med Chem. 2014; 21(36):4092-9. Transcellular channels through brain capillary endothelial cells can occur via: 1) leukocyte entry into cells; 2) carrier-mediated influx, such as glucose via glucose transporter 1 (GLUT-1), amino acids via, for example, L-amino acid transporter 1 (LAT-1), and small peptides via, for example, organic anion transporter peptide B (OATP-B); 3) paracellular channels for small hydrophobic molecules; 4) adsorption-mediated endocytosis, such as albumin and cationized molecules; 5) passive diffusion of lipid-soluble nonpolar solutes (including CO2 and O2); and 6) receptor-mediated endocytosis, such as insulin via insulin receptor endocytosis and TfR via Tf endocytosis. Johnsen et al., Targeting the transferrin receptor for brain drug delivery, ProgNeurobiol. 2019 Oct; 181:101665.
[0207] For example, an anti-TfR:GAA fusion protein was provided, exhibiting high affinity for the transferrin receptor and excellent blood-brain barrier crossing. Surprisingly, the fusion exhibiting high binding affinity for TfR crossed the blood-brain barrier more efficiently than low-affinity binders. We found that the high-affinity antibody, in the form of anti-hTFRscfv:GAA, conferred optimal delivery to the CNS and muscle. This contrasts with previous findings using monovalent and bivalent anti-TFR antibodies, in which low-affinity antibodies crossed the BBB more efficiently. The fusion presented herein possesses the ability to efficiently deliver GAA to the brain, thus enabling effective treatment of diseases such as GAA deficiency (e.g., Pompe disease).
[0208] This document provides antigen-binding proteins that specifically bind to transferrin receptors, preferably human transferrin receptor 1 (anti-hTfR), such as antibodies and their antigen-binding fragments, such as Fab and scFv. For example, in an embodiment, anti-hTfR is in the form of a fusion protein. The fusion protein comprises an anti-hTfR antigen-binding protein fused to a GAA peptide. Anti-hTfR effectively crosses the blood-brain barrier (BBB) and thereby delivers the fused GAA to the brain.
[0209] Antigen-binding proteins that specifically bind to transferrin receptors (e.g., human transferrin receptor (e.g., REGN2431) or monkey transferrin receptor (e.g., REGN2054)) and their fusions (e.g., tags, such as His6 and / or myc) at approximately 25°C, for example, in a surface plasmon resonance assay at approximately 20 nM K. D Or bind with a higher affinity. This antigen-binding protein may be referred to as "anti-TfR". In some embodiments, the antigen-binding protein binds at approximately 0.41 nM K. D Or it may bind with a stronger affinity to the human transferrin receptor. In some embodiments, the antigen-binding protein binds at approximately 3 nM K. D Or it may bind with a stronger affinity to the human transferrin receptor. In some embodiments, the antigen-binding protein binds at a K+ level of approximately 0.45 nM to 3 nM. D It binds to the human transferrin receptor. In some embodiments, Fab with HCVR and LCVR binds at approximately 0.65 nM K. D Or it may bind with a stronger affinity to the human transferrin receptor. In some embodiments, the fusion protein disclosed herein binds at approximately 1 × 10⁻⁶. -7 M of K D Or it may bind to the human transferrin receptor with a stronger affinity.
[0210] In the embodiments, the anti-hTfR scFv:GAA fusion protein comprises a scFv containing an arrangement of variable regions as follows: LCVR-HCVR or HCVR-LCVR, wherein HCVR and LCVR are optionally linked by a linker, and the scFv is optionally linked by a linker to a GAA peptide (e.g., LCVR-(Gly4Ser)3(SEQ ID NO: 616)-HCVR-(Gly4Ser)2(SEQ ID NO: 617))-GAA; or LCVR-(Gly4Ser)3(SEQ ID NO: 616)-HCVR-(Gly4Ser)2(SEQ ID NO: 617))-GAA (Gly4Ser = SEQ ID NO: 537)). In one example, the scFv contains an arrangement of variable regions as follows: LCVR-HCVR. In another example, the scFv contains an arrangement of variable regions as follows: HCVR-LCVR. In one example, the linker between HCVR and LCVR comprises three such repeating sequences (SEQ ID NO: 616), substantially consisting of or consisting of them. For example, the coding sequence of the linker may comprise any one of SEQ ID NO: 618-622 and 747, substantially consisting of or consisting of them. In another example, the linker between HCVR and LCVR comprises two such repeating sequences (SEQ ID NO: 617), substantially consisting of or consisting of them. For example, the coding sequence of the linker may comprise any one of SEQ ID NO: 623-629, substantially consisting of or consisting of them. In yet another example, the linker between HCVR and LCVR comprises one such repeating sequence (SEQ ID NO: 537), substantially consisting of or consisting of it. For example, the coding sequence of the linker may comprise SEQ ID NO: 630 or 748, substantially consisting of or consisting of them. In one example, the linker between scFv and GAA contains three such repeating sequences (SEQ ID NO: 616), substantially consisting of or consisting of them. For example, the coding sequence of the linker may contain any of SEQ ID NO: 618-622 and 747, substantially consisting of or consisting of them. In another example, the linker between scFv and GAA contains two such repeating sequences (SEQ ID NO: 617), substantially consisting of or consisting of them. For example, the coding sequence of the linker may contain any of SEQ ID NO: 623-629, substantially consisting of or consisting of them. In yet another example, the linker between scFv and GAA contains one such repeating sequence (SEQ ID NO: 537), substantially consisting of or consisting of it.For example, the encoded sequence of the connector may contain, consist substantially of, or consist of SEQ ID NO: 630 or 748. In another example, a rigid connector, such as the 2XH4 connector, may be used. In one example, the connector contains AEAAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA (SEQ ID NO: 842), consists substantially of, or consists of. For example, the encoded sequence of the connector may contain SEQ ID NO: 841, consists substantially of, or consists of.
[0211] The anti-hTfR:GAA optionally includes a signal peptide linked to an antigen-binding protein that specifically binds to a transferrin receptor (TfR), preferably (optionally via a linker) a human transferrin receptor (hTfR) fused to the GAA. In embodiments, the signal peptide is an mROR signal sequence (e.g., mROR signal sequence -LCVR-(Gly4Ser)3 (SEQ ID NO: 616) -HCVR-(Gly4Ser)2 (SEQ ID NO: 617)) -GAA; or LCVR-(Gly4Ser)3 (SEQ ID NO: 616) -HCVR-(Gly4Ser)2 (SEQ ID NO: 617)) -GAA (Gly4Ser = SEQ ID NO: 537)). The terms "fused" or "tethered" in relation to fusion peptides refer to peptides that are directly or indirectly bound (e.g., via a linker or other peptide).
[0212] In the implementation scheme, amino acids are assigned to each frame or CDR domain in the immunoglobulin according to the definitions in the following literature: Sequences of Proteins of Immunological Interest, Kabat et al.; National Institutes of Health, Bethesda, Md.; 5th edition; NIH Publ. No. 91-3242 (1991); Kabat (1978) Adv. Prot. Chem. 32:1-75; Kabat et al., (1977) J. Biol. Chem. 252:6609-6616; Chothia et al., (1987) J Mol. Biol. 196:901-917; or Chothia et al., (1989) Nature 342: 878-883. Therefore, it includes antibody and antigen-binding fragments containing V... H CDR and V L The CDR, the V H and VL It contains amino acid sequences as shown herein (see, for example, sequences in Table 2, or variants thereof), wherein the CDR is as defined according to Kabat and / or Chothia.
[0213] In some multidomain therapeutic proteins, the delivery domain that binds to the TfR is an antibody, antibody fragment, or other antigen-binding protein. Examples of antigen-binding proteins include, for instance, receptor-fusion molecules, capture molecules, receptor-Fc fusion molecules, antibodies, Fab fragments, F(ab')2 fragments, Fd fragments, Fv fragments, single-chain Fv (scFv) molecules, dAb fragments, separated complementarity-determining regions (CDRs), CDR3 peptides, restricted FR3-CDR3-FR4 peptides, domain-specific antibodies, single-domain antibodies, domain-deficient antibodies, chimeric antibodies, CDR-transplanted antibodies, biantibodies, triantibodies, tetraantibodies, microantibodies, nanobodies, monovalent nanobodies, divalent nanobodies, small modular immunopharmaceuticals (SMIPs), camelid antibodies (VHH heavy chain homodimer antibodies), and shark variable IgNAR domains.
[0214] This document provides antibodies that specifically bind to human transferrin receptor 1. As used herein, the term "antibody" refers to an immunoglobulin molecule comprising four polypeptide chains interconnected by disulfide bonds: two heavy chains (HC) and two light chains (LC). In embodiments, each antibody heavy chain (HC) contains a heavy chain variable region ("HCVR" or "V"). H (e.g., containing SEQ ID NO: 171, 680, 181, 681, 191, 682, 201, 211, 221, 685, 231, 687, 241, 689, 251, 261, 691, 271, 281, 692, 291, 301, 311, 694, 321, 331, 696, 341, 351, 697, 361, 699, 371, 700, 381, 391, 401, 411, 421, 701, 431, 441, 451, 461, 471, 702 and / or 481, or variants thereof) and a heavy chain constant region (e.g., human IgG, human IgG1 or human IgG4); and each antibody light chain (LC) contains a light chain variable region (“LCVR” or “V”). L(e.g., SEQ ID NO: 176, 186, 196, 206, 683, 216, 684, 226, 686, 236, 688, 246, 690, 256, 266, 276, 286, 693, 296, 306, 316, 695, 326, 336, 346, 356, 698, 366, 376, 386, 396, 406, 416, 426, 436, 446, 456, 466, 476, 632, 486 and / or 703, or variants thereof) and a light chain constant region (e.g., human κ or human λ). In embodiments, each antibody heavy chain (HC) contains a heavy chain variable region (“HCVR” or “V”). H (e.g., containing SEQ ID NO: 391 or 411, or a variant thereof) and a heavy chain constant region (e.g., human IgG, human IgG1, or human IgG4); and each antibody light chain (LC) contains a light chain variable region (“LCVR” or “V”). L (e.g., SEQ ID NO: 396 or 416, or variants thereof) and a light chain constant region (e.g., human κ or human λ). In embodiments, each antibody heavy chain (HC) contains a heavy chain variable region (“HCVR” or “V”). H (e.g., containing SEQ ID NO: 391, or a variant thereof) and a heavy chain constant region (e.g., human IgG, human IgG1, or human IgG4); and each antibody light chain (LC) contains a light chain variable region (“LCVR” or “V”). L (e.g., SEQ ID NO: 396, or a variant thereof) and a light chain constant region (e.g., human κ or human λ). In embodiments, each antibody heavy chain (HC) contains a heavy chain variable region (“HCVR” or “V”). H (e.g., containing SEQ ID NO: 411, or a variant thereof) and a heavy chain constant region (e.g., human IgG, human IgG1, or human IgG4); and each antibody light chain (LC) contains a light chain variable region (“LCVR” or “V”). L (e.g., SEQ ID NO: 416, or a variant thereof) and light chain constant regions (e.g., human κ or human λ). V H District and V L The region can be further subdivided into highly variable regions known as Complementarity Determining Regions (CDRs), interspersed with more conservative regions known as Framing Regions (FRs). Each V H and V L It contains three CDRs and four FRs. The anti-TfR antibody disclosed in this paper can also be fused with GAA.
[0215] The anti-TfR antigen-binding proteins described herein can be antigen-binding fragments of antibodies capable of tethering to GAAs. As used herein, the term "antigen-binding portion" or "antigen-binding fragment" of an antibody refers to an immunoglobulin molecule that binds an antigen but does not contain the complete antibody (preferably, the complete antibody is IgG) of its entire sequence. Non-limiting examples of antigen-binding fragments include: (i) Fab fragments; (ii) F(ab')2 fragments; (iii) Fd fragments; (iv) Fv fragments; (v) single-chain Fv (scFv) molecules; and (vi) dAb fragments; consisting of amino acid residues mimicking the hypervariable region of an antibody (e.g., a separated complementarity-determining region (CDR), such as a CDR3 peptide) or a bound FR3-CDR3-FR4 peptide. Other engineered molecules, such as domain-specific antibodies, single-domain antibodies, domain-deficient antibodies, chimeric antibodies, CDR-transplanted antibodies, biantibodies, triantibodies, tetraantibodies, microantibodies, and small modular immunopharmaceuticals (SMIPs), are also covered under the term "antigen-binding fragment" as used herein.
[0216] Anti-TfR antigen-binding proteins can be scFv, which can tether to the GAA. scFv (single-stranded variable fragment) has a variable region (V) with a restructured domain. H ) and the variable region (V) of the light structural domain L (In any order), these variable regions are preferably joined together by flexible linkers (e.g., peptide linkers). The length of the flexible linker used to connect the two V regions can be important for the proper folding of the polypeptide chain. It has been previously estimated that the peptide linker must span 3.5 nm (35 Å) between the carboxyl terminus of the variable domain and the amino terminus of the other domain without affecting the domain's ability to fold and form an intact antigen-binding site (Huston et al., Protein engineering of single-chain Fv analogs and fusion proteins. Methods in Enzymology. 1991; 203:46–88). In an embodiment, the linker contains an amino acid sequence of such length that the variable domains are spaced approximately 3.5 nm apart.
[0217] In some embodiments, the anti-TfR antigen-binding proteins described herein comprise monovalent or “single-arm” antibodies. As used herein, a monovalent or “single-arm” antibody refers to an immunoglobulin containing a single variable domain. For example, a single-arm antibody may contain a single variable domain within a Fab, wherein the Fab is linked to at least one Fc fragment. In some embodiments, the single-arm antibody comprises: (i) a heavy chain containing a heavy chain constant region and a heavy chain variable region, (ii) a light chain containing a light chain constant region and a light chain variable region, and (iii) a polypeptide containing an Fc fragment or a truncated heavy chain. In some embodiments, the Fc fragment or truncated heavy chain contained in a separate polypeptide is a “pseudo-Fc,” which refers to an Fc fragment not linked to an antigen-binding domain. The single-arm antibodies of this disclosure may comprise any of the HCVR / LCVR pairs or CDR amino acid sequences shown herein in Table 2. Standard methodologies can be used to construct single-arm antibodies comprising a full-length heavy chain, a full-length light chain, and an additional Fc domain polypeptide (see, for example, WO2010151792, the full text of which is incorporated herein by reference), wherein the heavy chain constant region differs from the Fc domain polypeptide by at least two amino acids (e.g., H95R and Y96F according to the IMGT exon numbering system; or H435R and Y436F according to the EU numbering system). Such modifications can be used to purify monovalent antibodies (see WO2010151792).
[0218] In this embodiment, the antigen-binding fragment of the antibody will contain at least one variable domain. The variable domain can have any size or amino acid composition and typically contains at least one CDR, which is adjacent to or within one or more frame sequences. L V of domain association H In the antigen-binding fragment of the domain, V H and V L Domains can be positioned relative to each other in any suitable arrangement. For example, variable regions can be dimers and contain V. H -V H V H -V L or V L -V L Dimer. Alternatively, the antigen-binding fragment of the antibody may contain monomer V. H or V L Structural domain.
[0219] In some embodiments, the antigen-binding fragment of the antibody may contain at least one variable domain covalently linked to at least one constant domain. Non-limiting exemplary configurations of the variable and constant domains found within the antigen-binding fragment of the antibody described herein include: (i) V H -CH1、(ii) V H-CH2、(iii) V H -CH3、(iv) V H -CH1-CH2、(v) V H -CH1-CH2-CH3、(vi) V H -CH2-CH3、(vii) V H -CL、(viii) V L -CH1、(ix) V L -CH2、(x)V L -CH3、(xi) V L -CH1-CH2, (xii) VL-CH1-CH2-CH3, (xiii) V L -CH2-CH3 and (xiv) V L -CL. In any configuration of the variable and constant domains (including any of the exemplary configurations listed above), the variable and constant domains may be directly connected to each other or connected via complete or partial hinge regions or linker regions. The hinge region may consist of at least two (e.g., 5, 10, 15, 20, 40, 60, or more) amino acids, forming a flexible or semi-flexible connection between adjacent variable and / or constant domains in a single polypeptide molecule. Furthermore, the antigen-binding fragments of the antibodies described herein may comprise homodimers or heterodimers (or other multimers) having any of the variable and constant domain configurations listed above, non-covalently associated with each other and / or with one or more monomers V H or V L Domains (e.g., via disulfide bonds) are non-covalently associated. This disclosure includes antigen-binding fragments of antigen-binding proteins, such as antibodies shown herein.
[0220] Antigen-binding proteins (e.g., antibodies and antigen-binding fragments) can be monospecific or multispecific (e.g., bispecific). Multispecific antigen-binding proteins are discussed further herein. This disclosure includes both monospecific and multispecific (e.g., bispecific) antigen-binding fragments that contain one or more variable domains from antigen-binding proteins specifically described herein.
[0221] The term "specific binding" or "specific binding" refers to proteins that bind to antigens (such as human TfR protein, mouse TfR protein, or monkey TfR protein) with a binding affinity of K. D At least about 10 -9The binding affinity of an antigen-binding protein (e.g., an antibody or its antigen-binding fragment) to M (e.g., 0.01 nM, 0.1 nM, 0.2 nM, 0.3 nM, 0.4 nM, 0.5 nM, 0.6 nM, 0.7 nM, 0.8 nM, 0.9 nM, or 1.0 nM), such binding affinity is determined by real-time label-free biolayer interference, for example at 25 °C or 37 °C (e.g., Octet). ® HTX biosensors), or through surface plasmon resonance (e.g., BIACORE). ™ This can be measured by solution affinity ELISA, or by other means. This disclosure includes antigen-binding proteins that specifically bind to TfR proteins. "Anti-TfR" refers to antigen-binding proteins (or other molecules) that specifically bind to TfR, such as antibodies or their antigen-binding fragments.
[0222] "Isolated" antigen-binding proteins (e.g., antibodies or antigen-binding fragments thereof), polypeptides, polynucleotides, and carriers are at least partially free of other biomolecules from the cells or cell cultures in which they are produced. Such biomolecules include nucleic acids, proteins, other antibodies or antigen-binding fragments, lipids, carbohydrates, or other materials such as cell debris and growth media. Isolated antigen-binding proteins may also be at least partially free of expression system components, such as biomolecules from host cells or their growth media. Generally, the term "isolated" is not intended to refer to the complete absence of such biomolecules (e.g., the possible retention of small or insignificant amounts of impurities), or the absence of water, buffers, or salts, or components of pharmaceutical formulations containing antigen-binding proteins (e.g., antibodies or antigen-binding fragments).
[0223] This disclosure includes antigen-binding proteins, such as antibodies or antigen-binding fragments, that bind to the same epitopes as the antigen-binding proteins described herein. In some embodiments, an antigen-binding protein is provided that specifically binds to the transferrin receptor or its antigen fragments or variants thereof, binding one or more hTfR epitopes selected from the following: (a) an epitope containing the sequence LLNE (SEQ ID NO: 796) and / or an epitope containing the sequence TYKEL (SEQ ID NO: 706); (b) an epitope containing the sequence DSTDFTGT (SEQ ID NO: 797) and / or an epitope containing the sequence VKHPVTGQF (SEQ ID NO: 798) and / or an epitope containing the sequence IERIPEL (SEQ ID NO: 799); (c) an epitope containing the sequence LNENSYVPREAGSQKDEN (SEQ ID NO: 800); (d) an epitope containing the sequence FEDL (SEQ ID NO: 718); (e) an epitope containing the sequence IVDKNGRL (SEQ ID NO: 801); (f) an epitope containing the sequence IVDKNGRLVY (SEQ ID NO: 706). Epitopes containing: (g) the sequence DQTKF (SEQ ID NO: 803); (h) the sequence LVENPGGY (SEQ ID NO: 804) and / or the sequence PIVNAELSF (SEQ ID NO: 805) and / or the sequence PYLGTTMDT (SEQ ID NO: 806); (i) the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or the sequence TYKEL (SEQ ID NO: 706); (j) the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 802); (h) the sequence DQTKF (SEQ ID NO: 803); (h) the sequence LVENPGGY (SEQ ID NO: 804) and / or the sequence PIVNAELSF (SEQ ID NO: 805) and / or the sequence PYLGTTMDT (SEQ ID NO: 806); (i) the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or the sequence TYKEL (SEQ ID NO: 706); (j) the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 702); (g) the sequence DQTKF (SEQ ID NO: 803); (h) the sequence LVENPGGY (SEQ ID NO: 804) and / or the sequence PIVNAELSF (SEQ ID NO: 805) and / or the sequence PYLGTTMDT (SEQ ID NO: 806); (f) the Epitopes of (708) and / or epitopes containing the sequence LIERIPELNKVARAAAE (SEQ ID NO: 709); (k) epitopes containing the sequence LNENSYVPREAGSQKDENL (SEQ ID NO: 710); (l) epitopes containing the sequence GTKKDFEDL (SEQ ID NO: 711); (m) epitopes containing the sequence SVIIVDKNGRLVYLVENPGGYVAYSK (SEQ ID NO: 712);(n) Epitopes containing the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes containing the sequence DQTKFPIVNAEL (SEQ ID NO: 714) and / or epitopes containing the sequence TYKELIERIPELNK (SEQ ID NO: 715); (o) Epitopes containing the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes containing the sequence TYKELIERIPELNK (SEQ ID NO: 715); (p) Epitopes containing the sequence SVIIVDKNGRLVYLVENPGGYVAY (SEQ ID NO: 716); (q) Epitopes containing the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes containing the sequence FGNMEGDCPSDWKTDSTCRM (SEQ ID NO: 716). (r) Epitopes containing the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes containing the sequence LVENPGYVAYSKAATVTGKL (SEQ ID NO: 719) and / or epitopes containing the sequence IYMDQTKFPIVNAELSF (SEQ ID NO: 720) and / or epitopes containing the sequence ISRAAAEKL (SEQ ID NO: 721) and / or epitopes containing the sequence VTSESKNVKLTVSNVLKE (SEQ ID NO: 722) and / or epitopes containing the sequence FCEDTDYPYLGTTMDT (SEQ ID NO: 723); (s) Epitopes contained within or overlapping with the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes contained within the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 717). (705) Epitopes contained in or overlapping with the sequence and / or included in or overlapping with the sequence TYKEL (SEQ ID NO: 706); (t) Epitopes contained in or overlapping with the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or included in or overlapping with the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 708) and / or included in or overlapping with the sequence LIERIPELNKVARAAAE (SEQ ID NO: 709); (u) Epitopes contained in or overlapping with the sequence LNENSYVPREAGSQKDENL (SEQ ID NO: 710);(v) Epitopes contained in or overlapping with the sequence GTKKDFEDL (SEQ ID NO: 711); (w) Epitopes contained in or overlapping with the sequence SVIIVDKNGRLVYLVENPGGYVAYSK (SEQ ID NO: 712); (x) Epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or Epitopes contained in or overlapping with the sequence DQTKFPIVNAEL (SEQ ID NO: 714) and / or Epitopes contained in or overlapping with the sequence TYKELIERIPELNK (SEQ ID NO: 715); (y) Epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or Epitopes contained in or overlapping with the sequence TYKELIERIPELNK (SEQ ID NO: 715). Epitopes contained in or overlapping with the sequence 715); (z) epitopes contained in or overlapping with the sequence SVIIVDKNGRLVYLVENPGGYVAY (SEQ ID NO: 716); (aa) epitopes contained in or overlapping with the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes contained in or overlapping with the sequence FGNMEGDCPSDWKTDSTCRM (SEQ ID NO: 717); and (bb) epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes contained in or overlapping with the sequence LVENPGYVAYSKAATVTGKL (SEQ ID NO: 719) and / or epitopes contained in or overlapping with the sequence IYMDQTKFPIVNAELSF (SEQ ID NO: 716). Epitopes contained in or overlapping with the sequence ISRAAAEKL (SEQ ID NO: 721) and / or contained in or overlapping with the sequence VTSESKNVKLTVSNVLKE (SEQ ID NO: 722) and / or contained in or overlapping with the sequence FCETDYPYLGTTMDT (SEQ ID NO: 723). In some embodiments, an antigen-binding protein is provided, wherein the antigen-binding protein comprises an antibody or an antigen-binding fragment thereof that binds one or more hTfR epitopes selected from the following: (a) an epitope consisting of the sequence LLNE (SEQ ID NO: 796) and / or an epitope consisting of the sequence TYKEL (SEQ ID NO: 706);(b) Epitopes consisting of the sequence DSTDFTGT (SEQ ID NO: 797) and / or the sequence VKHPVTGQF (SEQ ID NO: 798) and / or the sequence IERIPEL (SEQ ID NO: 799); (c) Epitopes consisting of the sequence LNENSYVPREAGSQKDEN (SEQ ID NO: 800); (d) Epitopes consisting of the sequence FEDL (SEQ ID NO: 718); (e) Epitopes consisting of the sequence IVDKNGRL (SEQ ID NO: 801); (f) Epitopes consisting of the sequence IVDKNGRLVY (SEQ ID NO: 802); (g) Epitopes consisting of the sequence DQTKF (SEQ ID NO: 803); (h) Epitopes consisting of the sequence LVENPGGY (SEQ ID NO: 804) and / or the sequence PIVNAELSF (SEQ ID NO: 799). Epitopes consisting of (SEQ ID NO: 805) and / or epitopes consisting of the sequence PYLGTTMDT (SEQ ID NO: 806); (i) epitopes consisting of the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes consisting of the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes consisting of the sequence TYKEL (SEQ ID NO: 706); (j) epitopes consisting of the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or epitopes consisting of the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 708) and / or epitopes consisting of the sequence LIERIPELNKVARAAAE (SEQ ID NO: 709); (k) epitopes consisting of the sequence LNENSYVPREAGSQKDENL (SEQ ID NO: 805) and / or epitopes consisting of the sequence PYLGTTMDT (SEQ ID NO: 806); Epitopes consisting of (710) the sequence GTKKDFEDL (SEQ ID NO: 711); (m) the sequence SVIIVDKNGRLVYLVENPGGYVAYSK (SEQ ID NO: 712); (n) the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or the sequence DQTKFPIVNAEL (SEQ ID NO: 714) and / or the sequence TYKELIERIPELNK (SEQ ID NO: 715);(o) an epitope consisting of the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or an epitope consisting of the sequence TYKELIERIPELNK (SEQ ID NO: 715); (p) an epitope consisting of the sequence SVIIVDKNGRLVYLVENPGGYVAY (SEQ ID NO: 716); (q) an epitope consisting of the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or an epitope consisting of the sequence FGNMEGDCPSDWKTDSTCRM (SEQ ID NO: 717); and (r) an epitope consisting of the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or an epitope consisting of the sequence LVENPGYVAYSKAATVTGKL (SEQ ID NO: 719) and / or an epitope consisting of the sequence IYMDQTKFPIVNAELSF (SEQ ID NO: 716). Epitopes consisting of the sequence ISRAAAEKL (SEQ ID NO: 720) and / or the sequence VTSESKNVKLTVSNVLKE (SEQ ID NO: 722) and / or the sequence FCEDTDYPYLGTTMDT (SEQ ID NO: 723).
[0224] An antigen is a molecule to which a peptide (e.g., TfR or a fragment thereof (antigen fragment)) binds, such as an antibody or an antigen-binding fragment thereof. A specific region on an antigen that an antibody recognizes and binds to is called an epitope. Antigen-binding proteins (e.g., antibodies) that specifically bind to such antigens, as described herein, are part of this disclosure.
[0225] The term "epitope" refers to an antigenic determinant (e.g., on a TfR) that interacts with a specific antigen-binding site of an antigen-binding protein, such as a variable region of an antibody, called a complementary site. A single antigen may have more than one epitope. Thus, different antibodies can bind to different regions on an antigen and may have different biological effects. The term "epitope" may also refer to a site on an antigen to which B cells and / or T cells respond, and / or a region on an antigen to which an antibody binds. Epitopes can be defined as structural or functional. Functional epitopes are typically a subset of structural epitopes and have those residues that directly contribute to the affinity of the interaction. Epitopes can be linear or conformational, i.e., composed of non-linear amino acids. In some embodiments, an epitope may comprise a determinant of chemically active surface groups (such as amino acids, sugar side chains, phosphoryl groups, or sulfonyl groups) as molecules, and in some embodiments, may have specific three-dimensional structural properties and / or specific charge properties. Epitopes bound by the antigen-binding proteins described herein may be contained in fragments of the TfR (e.g., its extracellular domain). The antigen-binding proteins (e.g., antibodies) that bind to such epitopes as described herein are part of this disclosure.
[0226] Methods for identifying epitopes of antigen-binding proteins (e.g., antibodies, fragments, or peptides) include alanine scanning mutation analysis, peptide blotting (Reineke (2004) Methods Mol. Biol. 248: 443-63), peptide cleavage analysis, crystallographic studies, and NMR analysis. Alternatively, methods such as epitope excision, epitope extraction, and antigen chemical modification can be employed (Tomer (2000) Prot. Sci. 9: 487-496). Another method for ...
Claims
1. A composition comprising a nucleic acid construct containing a coding sequence for a multi-domain therapeutic protein, the multi-domain therapeutic protein containing a delivery domain fused to a lysosomal α-glucosidase polypeptide, wherein the lysosomal α-glucosidase coding sequence is CpG depleted relative to a wild-type lysosomal α-glucosidase coding sequence, optionally wherein the delivery domain is a TfR-binding delivery domain or a CD63-binding delivery domain.
2. The composition of claim 1, wherein the nucleic acid construct comprises a polyadenylation signal or sequence located downstream of the coding sequence of the multidomain therapeutic protein.
3. The composition according to claim 2, wherein the polyadenylation signal comprises bovine growth hormone (BGH) polyadenylation signal, simian virus 40 (SV40) polyadenylation signal, or a combination of the bovine growth hormone polyadenylation signal and the SV40 polyadenylation signal.
4. The composition of claim 3, wherein the SV40 polyadenylation signal is a unidirectional SV40 late polyadenylation signal, wherein each instance of the sequence AATAAA in the reverse strand is mutated in the unidirectional SV40 late polyadenylation signal, optionally wherein the SV40 polyadenylation signal is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in SEQ ID NO: 752, and optionally wherein the SV40 polyadenylation signal comprises the sequence shown in SEQ ID NO:
752.
5. The composition according to claim 3 or 4, wherein the polyadenylation signal comprises the BGH polyadenylation signal, optionally wherein the BGH polyadenylation signal is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in SEQ ID NO: 751, optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO:
751.
6. The composition according to any one of claims 3-5, wherein the polyadenylation signal comprises the BGH polyadenylation signal and the SV40 polyadenylation signal, optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO: 751, and optionally wherein the SV40 polyadenylation signal comprises the sequence shown in SEQ ID NO: 752, optionally wherein the polyadenylation signal comprising the BGH polyadenylation signal and the SV40 polyadenylation signal is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in SEQ ID NO: 795, and optionally wherein the polyadenylation signal comprising the BGH polyadenylation signal and the SV40 polyadenylation signal comprises the sequence shown in SEQ ID NO:
795.
7. The composition according to any one of claims 1-6, wherein the nucleic acid construct is a single-direction nucleic acid construct.
8. The composition according to any one of claims 1-7, wherein the coding sequence of the delivery domain is modified to remove one or more hidden splice sites, the coding sequence of the lysosomal α-glucosidase polypeptide is modified to remove one or more hidden splice sites, or the coding sequence of the multi-domain therapeutic protein is modified to remove one or more hidden splice sites.
9. The composition according to any one of claims 1-8, wherein the coding sequence of the delivery domain is CpG depleted, or the coding sequence of the multi-domain therapeutic protein is CpG depleted.
10. The composition according to any one of claims 1-9, wherein the coding sequence of the delivery domain is codon-optimized and CpG-depleted, the coding sequence of the lysosomal α-glucosidase polypeptide is codon-optimized and CpG-depleted, or the coding sequence of the multi-domain therapeutic protein is codon-optimized and CpG-depleted.
11. The composition according to any one of claims 1-10, wherein the nucleic acid construct comprises a splice acceptor located upstream of the coding sequence of the multi-domain therapeutic protein.
12. The composition according to any one of claims 1-11, wherein the nucleic acid construct does not contain homologous arms.
13. The composition according to any one of claims 1-12, wherein the nucleic acid construct from 5' to 3' comprises: a splice acceptor, the coding sequence of the multi-domain therapeutic protein, and a polyadenylation signal or sequence. The nucleic acid construct described herein does not contain a promoter that drives the expression of the multi-domain therapeutic protein, and The nucleic acid constructs described herein do not contain homologous arms.
14. The composition according to any one of claims 1-11, wherein the nucleic acid construct comprises a homologous arm.
15. The composition according to any one of claims 1-14, wherein the nucleic acid construct does not contain a promoter that drives the expression of the multi-domain therapeutic protein.
16. The composition according to any one of claims 1-14, wherein the coding sequence of the multi-domain therapeutic protein is operatively linked to a promoter, optionally wherein the promoter is a liver-specific promoter.
17. The composition according to any one of claims 1-16, wherein the C-terminus of the delivery domain is fused to the N-terminus of the lysosomal α-glucosidase polypeptide.
18. The composition according to any one of claims 1-17, wherein the delivery domain is fused to the lysosomal α-glucosidase polypeptide via a peptide linker.
19. The composition according to any one of claims 1-18, wherein the lysosomal α-glucosidase polypeptide lacks the lysosomal α-glucosidase signal peptide and the propeptide.
20. The composition according to any one of claims 1-19, wherein the lysosomal α-glucosidase polypeptide comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO:
727.
21. The composition according to any one of claims 1-20, wherein the lysosomal α-glucosidase encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 750, optionally wherein the nucleotide at position 1095 is G, the nucleotide at position 1098 is C, and the nucleotide at position 2343 is G.
22. The composition according to any one of claims 1-21, wherein the lysosomal α-glucosidase encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NO: 750, and encodes a lysosomal α-glucosidase protein comprising SEQ ID NO: 727, optionally wherein the nucleotide at position 1095 is G, the nucleotide at position 1098 is C, and the nucleotide at position 2343 is G.
23. The composition according to any one of claims 1-22, wherein the lysosomal α-glucosidase encoding sequence comprises, is substantially composed of, or is composed of any one of the sequences shown in SEQ ID NO:
750.
24. The composition according to any one of claims 1-20, wherein the lysosomal α-glucosidase encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 749, optionally wherein the nucleotide at position 2343 is G.
25. The composition according to any one of claims 1-20 and 24, wherein the lysosomal α-glucosidase encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NO: 749, and encodes a lysosomal α-glucosidase protein comprising SEQ ID NO: 727, optionally wherein the nucleotide at position 2343 is G.
26. The composition according to any one of claims 1-20, 24 and 25, wherein the lysosomal α-glucosidase encoding sequence comprises, is substantially composed of or is composed of the sequence shown in any one of SEQ ID NO:
749.
27. The composition according to any one of claims 1-26, wherein the delivery domain is the TfR-binding delivery domain.
28. The composition of claim 27, wherein the TfR-binding delivery domain comprises an anti-TfR antigen-binding protein, optionally wherein the antigen-binding protein is in a K+ of about 41 nM. D Or with a stronger affinity, it binds to the human transferrin receptor, optionally wherein the antigen-binding protein binds at approximately 3 nM K. D Or, with a stronger affinity, bind to the human transferrin receptor, or optionally, wherein the antigen-binding protein binds at a K+ of about 0.45 nM to 3 nM. D It binds to the human transferrin receptor.
29. The composition of claim 28, wherein the anti-TfR antigen-binding protein comprises: (i) an HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 391, 171, 181, 191, 201, 211, 221, 231, 241, 251, 261, 271, 281, 291, 301, 311, 321, 331, 341, 351, 361, 371, 381, 401, 411, 421, 431, 441, 451, 461, 471, or 481; and / or (ii) An LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 396, 176, 186, 196, 206, 216, 226, 236, 246, 256, 266, 276, 286, 296, 306, 316, 326, 336, 346, 356, 366, 376, 386, 406, 416, 426, 436, 446, 456, 466, 476, or 486.
30. The composition according to claim 28 or 29, wherein the anti-TfR antigen-binding protein comprises: (1) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof). (2) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 171 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 176 (or a variant thereof). (3) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 181 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 186 (or a variant thereof). (4) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 191 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 196 (or a variant thereof). (5) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 201 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 206 (or a variant thereof). (6) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 211 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 216 (or a variant thereof). (7) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 221 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 226 (or a variant thereof). (8) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 231 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 236 (or a variant thereof). (9) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 241 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 246 (or a variant thereof). (10) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 251 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 256 (or a variant thereof). (11) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 261 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 266 (or a variant thereof). (12) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 271 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 276 (or a variant thereof). (13) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 281 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 286 (or a variant thereof). (14) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 291 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 296 (or a variant thereof). (15) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 301 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 306 (or a variant thereof). (16) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 311 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 316 (or a variant thereof). (17) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 321 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 326 (or a variant thereof). (18) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 331 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 336 (or a variant thereof). (19) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 341 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 346 (or a variant thereof). (20) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 351 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 356 (or a variant thereof). (21) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 361 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 366 (or a variant thereof). (22) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 371 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 376 (or a variant thereof). (23) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 381 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 386 (or a variant thereof). (24) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 401 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 406 (or a variant thereof). (25) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 411 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 416 (or a variant thereof). (26) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 421 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 426 (or a variant thereof). (27) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 431 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 436 (or a variant thereof). (28) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 441 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 446 (or a variant thereof). (29) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 451 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 456 (or a variant thereof). (30) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 461 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 466 (or a variant thereof). (31) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 471 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 476 (or a variant thereof); or (32) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 481 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 486 (or a variant thereof).
31. The composition according to any one of claims 28-30, wherein the anti-TfR antigen-binding protein comprises: (1) An HCVR comprising HCDR1, HCDR2, and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2, and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof); or (2) An HCVR comprising HCDR1, HCDR2 and HCDR3, wherein the HCVR comprises the amino acid sequence shown in SEQ ID NO: 411 (or a variant thereof); and an LCVR comprising LCDR1, LCDR2 and LCDR3, wherein the LCVR comprises the amino acid sequence shown in SEQ ID NO: 416 (or a variant thereof).
32. The composition according to any one of claims 28-31, wherein the anti-TfR antigen-binding protein comprises: (a) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 392, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 393, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 394; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 397, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 398, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 399; (b) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 172, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 173, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 174; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 177, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 178, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 179; (c) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 182, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 183, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 184; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 187, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 188, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 189; (d) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 192, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 193, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 194; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 197, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 198, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 199; (e) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 202, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 203, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 204; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 207, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 208, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 209; (f) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 212, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 213, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 214; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 217, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 218, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 219; (g) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 222, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 223, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 224; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 227, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 228, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 229; (h) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 232, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 233, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 234; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 237, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 238, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 239; (i) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 242, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 243, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 244; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 247, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 248, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 249; (j) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 252, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 253, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 254; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 257, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 258, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 259; (k) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 262, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 263, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 264; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 267, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 268, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 269; (l) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 272, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 273, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 274; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 277, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 278, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 279; (m)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 282, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 283, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 284; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 287, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 288, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 289; (n)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 292, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 293, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 294; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 297, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 298, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 299; (o)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 302, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 303, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 304; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 307, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 308, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 309; (p)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 312, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 313, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 314; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 317, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 318, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 319; (q)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 322, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 323, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 324; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 327, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 328, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 329; (r)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 332, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 333, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 334; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 337, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 338, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 339; (s)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 342, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 343, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 344; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 347, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 348, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 349; (t)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 352, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 353, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 354; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 357, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 358, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 359; (u)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 362, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 363, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 364; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 367, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 368, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 369; (v) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 372, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 373, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 374; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 377, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 378, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 379; (w)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 382, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 383, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 384; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 387, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 388, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 389; (x)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 402, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 403, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 404; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 407, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 408, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 409; (y)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 412, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 413, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 414; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 417, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 418, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 419; (z)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 422, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 423, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 424; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 427, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 428, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 429; (aa)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 432, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 433, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 434; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 437, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 438, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 439; (ab)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 442, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 443, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 444; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 447, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 448, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 449; (ac)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 452, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 453, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 454; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 457, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 458, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 459; (ad)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 462, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 463, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 464; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 467, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 468, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 469; (ae)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 472, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 473, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 474; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 477, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 478, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 479; and / or (af)HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 482, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 483, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 484; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 487, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 488, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO:
489.
33. The composition according to any one of claims 28-32, wherein the anti-TfR antigen-binding protein comprises: (a) HCVR, comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 392, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 393, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 394; and LCVR, comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 397, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 398, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 399; or (b) HCVR comprising: HCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 412, HCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 413, and HCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 414; and LCVR comprising: LCDR1 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 417, LCDR2 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO: 418, and LCDR3 comprising the amino acid sequence (or a variant thereof) shown in SEQ ID NO:
419.
34. The composition according to any one of claims 28-33, wherein the anti-TfR antigen-binding protein comprises: (i) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof). (ii) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 171 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 176 (or a variant thereof). (iii) HCVR, comprising the amino acid sequence shown in SEQ ID NO: 181 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 186 (or a variant thereof). (iv) HCVR, comprising the amino acid sequence shown in SEQ ID NO: 191 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 196 (or a variant thereof). (v)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 201 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 206 (or a variant thereof). (vi) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 211 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 216 (or a variant thereof). (vii)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 221 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 226 (or a variant thereof); (viii) HCVR, comprising the amino acid sequence shown in SEQ ID NO: 231 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 236 (or a variant thereof); (ix)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 241 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 246 (or a variant thereof). (x)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 251 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 256 (or a variant thereof); (xi)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 261 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 266 (or a variant thereof). (xii)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 271 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 276 (or a variant thereof); (xiii) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 281 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 286 (or a variant thereof); (xiv)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 291 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 296 (or a variant thereof). (xv)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 301 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 306 (or a variant thereof); (xvi)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 311 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 316 (or a variant thereof). (xvii)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 321 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 326 (or a variant thereof); (xviii)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 331 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 336 (or a variant thereof); (xix)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 341 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 346 (or a variant thereof). (xx)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 351 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 356 (or a variant thereof); (xxi)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 361 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 366 (or a variant thereof). (xxii)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 371 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 376 (or a variant thereof); (xxiii)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 381 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 386 (or a variant thereof); (xxiv)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 401 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 406 (or a variant thereof); (xxv)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 411 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 416 (or a variant thereof). (xxvi)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 421 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 426 (or a variant thereof); (xxvii)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 431 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 436 (or a variant thereof); (xxviii)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 441 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 446 (or a variant thereof); (xxix)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 451 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 456 (or a variant thereof); (xxx)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 461 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 466 (or a variant thereof). (xxxi)HCVR, comprising the amino acid sequence shown in SEQ ID NO: 471 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 476 (or a variant thereof); and / or (xxxii)HCVR, which contains the amino acid sequence shown in SEQ ID NO: 481 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 486 (or a variant thereof).
35. The composition according to any one of claims 28-34, wherein the anti-TfR antigen-binding protein comprises: (i) HCVR, comprising the amino acid sequence shown in SEQ ID NO: 391 (or a variant thereof); and LCVR, comprising the amino acid sequence shown in SEQ ID NO: 396 (or a variant thereof); or (ii) HCVR, which contains the amino acid sequence shown in SEQ ID NO: 411 (or a variant thereof); and LCVR, which contains the amino acid sequence shown in SEQ ID NO: 416 (or a variant thereof).
36. The composition according to any one of claims 27-35, wherein the TfR-binding delivery domain is an antigen-binding protein that binds one or more hTfR epitopes selected from: (a) Epitopes containing the sequence LLNE (SEQ ID NO: 796) and / or epitopes containing the sequence TYKEL (SEQ ID NO: 706); (b) Epitopes containing the sequence DSTDFTGT (SEQ ID NO: 797) and / or epitopes containing the sequence VKHPVTGQF (SEQ ID NO: 798) and / or epitopes containing the sequence IERIPEL (SEQ ID NO: 799); (c) Epitopes containing the sequence LNENSYVPREAGSQKDEN (SEQ ID NO: 800); (d) Epitopes containing the sequence FEDL (SEQ ID NO: 718); (e) Epitopes containing the sequence IVDKNGRL (SEQ ID NO: 801); (f) Epitopes containing the sequence IVDKNGRLVY (SEQ ID NO: 802); (g) Epitopes containing the sequence DQTKF (SEQ ID NO: 803); (h) Epitopes containing the sequence LVENPGGY (SEQ ID NO: 804) and / or epitopes containing the sequence PIVNAELSF (SEQ ID NO: 805) and / or epitopes containing the sequence PYLGTTMDT (SEQ ID NO: 806); (i) Epitopes containing the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes containing the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes containing the sequence TYKEL (SEQ ID NO: 706); (j) Epitopes containing the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or epitopes containing the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 708) and / or epitopes containing the sequence LIERIPELNKVARAAAE (SEQ ID NO: 709); (k) contains epitopes of the sequence LNENSYVPREAGSQKDENL (SEQ ID NO: 710); (l) Epitopes containing the sequence GTKKDFEDL (SEQ ID NO: 711); (m) contains an epitope of the sequence SVIIVDKNGRLVYLVENPGGYVAYSK (SEQ ID NO: 712); (n) Epitopes containing the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes containing the sequence DQTKFPIVNAEL (SEQ ID NO: 714) and / or epitopes containing the sequence TYKELIERIPELNK (SEQ ID NO: 715); (o) Epitopes containing the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes containing the sequence TYKELIERIPELNK (SEQ ID NO: 715); (p) contains epitopes of the sequence SVIIVDKNGRLVYLVENPGGYVAY (SEQ ID NO: 716); (q) Epitopes containing the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes containing the sequence FGNMEGDCPSDWKTDSTCRM (SEQ ID NO: 717); (r) Epitopes containing the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes containing the sequence LVENPGYVAYSKAATVTGKL (SEQ ID NO: 719) and / or epitopes containing the sequence IYMDQTKFPIVNAELSF (SEQ ID NO: 720) and / or epitopes containing the sequence ISRAAAEKL (SEQ ID NO: 721) and / or epitopes containing the sequence VTSESKNVKLTVSNVLKE (SEQ ID NO: 722) and / or epitopes containing the sequence FCEDTDYPYLGTTMDT (SEQ ID NO: 723); (s) Epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes contained in or overlapping with the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes contained in or overlapping with the sequence TYKEL (SEQ ID NO: 706); (t) Epitopes contained in or overlapping with the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or epitopes contained in or overlapping with the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 708) and / or epitopes contained in or overlapping with the sequence LIERIPELNKVARAAAE (SEQ ID NO: 709); (u) Epitopes contained in or overlapping with the sequence LNENSYVPREAGSQKDENL (SEQ ID NO: 710); (v) Epitopes contained in or overlapping with the sequence GTKKDFEDL (SEQ ID NO: 711); (w) Epitopes contained in or overlapping with the sequence SVIIVDKNGRLVYLVENPGGYVAYSK (SEQ ID NO: 712); (x) Epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes contained in or overlapping with the sequence DQTKFPIVNAEL (SEQ ID NO: 714) and / or epitopes contained in or overlapping with the sequence TYKELIERIPELNK (SEQ ID NO: 715); (y) Epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes contained in or overlapping with the sequence TYKELIERIPELNK (SEQ ID NO: 715); (z) Epitopes contained in or overlapping with the sequence SVIIVDKNGRLVYLVENPGGYVAY (SEQ ID NO: 716); (aa) Epitopes contained in or overlapping with the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes contained in or overlapping with the sequence FGNMEGDCPSDWKTDSTCRM (SEQ ID NO: 717); and (bb) Epitopes contained in or overlapping with the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes contained in or overlapping with the sequence LVENPGYVAYSKAATVTGKL (SEQ ID NO: 719) and / or epitopes contained in or overlapping with the sequence IYMDQTKFPIVNAELSF (SEQ ID NO: 720) and / or epitopes contained in or overlapping with the sequence ISRAAAEKL (SEQ ID NO: 721) and / or epitopes contained in or overlapping with the sequence VTSESKNVKLTVSNVLKE (SEQ ID NO: 722) and / or epitopes contained in or overlapping with the sequence FCEDTDYPYLGTTMDT (SEQ ID NO: 723) 37. The composition of claim 36, wherein the TfR-binding delivery domain comprises an antibody or an antigen-binding fragment thereof that binds one or more hTfR epitopes selected from: (a) Epitopes consisting of the sequence LLNE (SEQ ID NO: 796) and / or epitopes consisting of the sequence TYKEL (SEQ ID NO: 706); (b) Epitopes consisting of the sequence DSTDFTGT (SEQ ID NO: 797) and / or the sequence VKHPVTGQF (SEQ ID NO: 798) and / or the sequence IERIPEL (SEQ ID NO: 799); (c) An epitope consisting of the sequence LNENSYVPREAGSQKDEN (SEQ ID NO: 800); (d) Epitopes consisting of the sequence FEDL (SEQ ID NO: 718); (e) Epitopes consisting of the sequence IVDKNGRL (SEQ ID NO: 801); (f) Epitopes consisting of the sequence IVDKNGRLVY (SEQ ID NO: 802); (g) Epitopes consisting of the sequence DQTKF (SEQ ID NO: 803); (h) an epitope consisting of the sequence LVENPGGY (SEQ ID NO: 804) and / or an epitope consisting of the sequence PIVNAELSF (SEQ ID NO: 805) and / or an epitope consisting of the sequence PYLGTTMDT (SEQ ID NO: 806); (i) Epitopes consisting of the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes consisting of the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or epitopes consisting of the sequence TYKEL (SEQ ID NO: 706); (j) Epitopes consisting of the sequence KRKLSEKLDSTDFTGTIKL (SEQ ID NO: 707) and / or epitopes consisting of the sequence YTLIEKTMQNVKHPVTGQFL (SEQ ID NO: 708) and / or epitopes consisting of the sequence LIERIPELNKVARAAAE (SEQ ID NO: 709); (k) An epitope consisting of the sequence LNENSYVPREAGSQKDENL (SEQ ID NO: 710); (l) An epitope consisting of the sequence GTKKDFEDL (SEQ ID NO: 711); (m) An epitope consisting of the sequence SVIIVDKNGRLVYLVENPGGYVAYSK (SEQ ID NO: 712); (n) Epitopes consisting of the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes consisting of the sequence DQTKFPIVNAEL (SEQ ID NO: 714) and / or epitopes consisting of the sequence TYKELIERIPELNK (SEQ ID NO: 715); (o) Epitopes consisting of the sequence LLNENSYVPREAGSQKDEN (SEQ ID NO: 713) and / or epitopes consisting of the sequence TYKELIERIPELNK (SEQ ID NO: 715); (p) An epitope consisting of the sequence SVIIVDKNGRLVYLVENPGGYVAY (SEQ ID NO: 716); (q) an epitope consisting of the sequence IYMDQTKFPIVNAEL (SEQ ID NO: 705) and / or an epitope consisting of the sequence FGNMEGDCPSDWKTDSTCRM (SEQ ID NO: 717); and (r) Epitopes consisting of the sequence LLNENSYVPREAGSQKDENLAL (SEQ ID NO: 704) and / or epitopes consisting of the sequence LVENPGYVAYSKAATVTGKL (SEQ ID NO: 719) and / or epitopes consisting of the sequence IYMDQTKFPIVNAELSF (SEQ ID NO: 720) and / or epitopes consisting of the sequence ISRAAAEKL (SEQ ID NO: 721) and / or epitopes consisting of the sequence VTSESKNVKLTVSNVLKE (SEQ ID NO: 722) and / or epitopes consisting of the sequence FCEDTDYPYLGTTMDT (SEQ ID NO: 723).
38. The composition according to any one of claims 27-37, wherein the TfR-binding delivery domain comprises an anti-TfR antibody, an antibody fragment, or a single-chain variable fragment (scFv).
39. The composition of claim 38, wherein the TfR-binding delivery domain is the single-chain variable fragment (scFv), optionally wherein the multi-domain therapeutic protein comprises domains arranged in the following orientation: N'-heavy chain variable region-light chain variable region-lysosomal α-glucosidase polypeptide-C' or N'-light chain variable region-heavy chain variable region-lysosomal α-glucosidase polypeptide-C'. Optionally, the scFv and the lysosomal α-glucosidase polypeptide are linked by a peptide linker, and optionally, the peptide linker is -(GGGGS). m - (SEQ ID NO: 537); where m is 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, Optionally, the scFv variable region is linked by a peptide linker, and optionally, the peptide linker is -(GGGGS). m - (SEQ ID NO: 537); where m is 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.
40. The composition of claim 39, wherein the multi-domain therapeutic protein comprises a heavy chain variable region (V... H ) and light chain variable region (V L ), and lysosomal α-glucosidase polypeptide, wherein the V H V L The polypeptides of lysosomal α-glucosidase are arranged as follows: (i)V L -V H - Lysosomal α-glucosidase polypeptide; (ii)V H -V L - Lysosomal α-glucosidase polypeptide; (iii)V L -[(GGGGS)3(SEQ ID NO: 616)]-V H -[(GGGGS)2(SEQ ID NO: 617)]-lysosomal α-glucosidase polypeptide; or (iv)V H -[(GGGGS)3(SEQ ID NO: 616)]-V L -[(GGGGS)2(SEQ ID NO: 617)]-lysosomal α-glucosidase polypeptide.
41. The composition according to claim 39 or 40, wherein the scFv comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO:
508.
42. The composition according to any one of claims 39-41, wherein the scFv encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 532, and encodes an scFv comprising SEQ ID NO:
508.
43. The composition according to any one of claims 39-42, wherein the scFv encoding sequence comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO:
532.
44. The composition according to any one of claims 27-43, wherein the multidomain therapeutic protein comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO:
746.
45. The composition according to any one of claims 27-44, wherein the multi-domain therapeutic protein coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 745, optionally wherein the nucleotide at position 1857 is G, the nucleotide at position 1860 is C, and the nucleotide at position 3105 is G.
46. The composition according to any one of claims 27-45, wherein the multi-domain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 745, and encodes a multi-domain therapeutic protein comprising SEQ ID NO: 746, optionally wherein the nucleotide at position 1857 is G, the nucleotide at position 1860 is C, and the nucleotide at position 3105 is G.
47. The composition according to any one of claims 27-46, wherein the multi-domain therapeutic protein coding sequence comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO:
745.
48. The composition according to any one of claims 27-47, wherein the nucleic acid construct from 5' to 3' comprises: a splice acceptor, the coding sequence of the multi-domain therapeutic protein, and a polyadenylation signal or sequence. The coding sequence of the multi-domain therapeutic protein comprises SEQ ID NO: 745, and optionally the nucleic acid construct comprises the sequence shown in SEQ ID NO: 780, or optionally the nucleic acid construct comprises the sequence shown in SEQ ID NO:
764. The polyadenylation signal comprises a BGH polyadenylation signal and a unidirectional SV40 late polyadenylation signal, optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO: 751 and the unidirectional SV40 late polyadenylation signal comprises the sequence shown in SEQ ID NO: 752, optionally the polyadenylation signal comprising the BGH polyadenylation signal and the unidirectional SV40 late polyadenylation signal comprises the sequence shown in SEQ ID NO:
795. The nucleic acid construct described herein does not contain a promoter that drives the expression of the multi-domain therapeutic protein, and The nucleic acid constructs described herein do not contain homologous arms.
49. The composition according to any one of claims 27-47, wherein the nucleic acid construct from 5' to 3' comprises: a splice acceptor, the coding sequence of the multi-domain therapeutic protein, and a polyadenylation signal or sequence. The coding sequence of the multi-domain therapeutic protein comprises SEQ ID NO: 745, and optionally the nucleic acid construct comprises the sequence shown in SEQ ID NO: 781, or optionally the nucleic acid construct comprises the sequence shown in SEQ ID NO:
765. The polyadenylation signal includes a BGH polyadenylation signal, optionally wherein the BGH polyadenylation signal includes the sequence shown in SEQ ID NO:
751. The nucleic acid construct described herein does not contain a promoter that drives the expression of the multi-domain therapeutic protein, and The nucleic acid constructs described herein do not contain homologous arms.
50. The composition according to any one of claims 1-26, wherein the delivery domain is a CD63-bound delivery domain.
51. The composition of claim 50, wherein the CD63-binding delivery domain comprises an anti-CD63 antigen-binding protein.
52. The composition of claim 50 or 51, wherein the CD63-binding delivery domain comprises an anti-CD63 antibody, an antibody fragment, or a single-chain variable fragment (scFv).
53. The composition of claim 52, wherein the delivery domain binding CD63 is the single-stranded variable fragment (scFv).
54. The composition of claim 53, wherein the scFv comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO:
730.
55. The composition according to any one of claims 53 or 54, wherein the scFv encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 759, optionally wherein the nucleotide at position 3 is A, the nucleotide at position 132 is A, and the nucleotide at position 273 is T.
56. The composition according to any one of claims 53-55, wherein the scFv encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 759, and encodes the scFv comprising SEQ ID NO: 730, optionally wherein the nucleotide at position 3 is A, the nucleotide at position 132 is A, and the nucleotide at position 273 is T.
57. The composition according to any one of claims 53-56, wherein the scFv encoding sequence comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO:
759.
58. The composition according to any one of claims 53 or 54, wherein the scFv coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:760, optionally wherein the nucleotide at position 273 is T.
59. The composition according to any one of claims 53, 54 and 58, wherein the scFv encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 760, and encodes an scFv comprising SEQ ID NO: 730, optionally wherein the nucleotide at position 273 is T.
60. The composition according to any one of claims 53, 54, 58 and 59, wherein the scFv encoding sequence comprises, is substantially composed of or is composed of the sequence shown in SEQ ID NO:
760.
61. The composition according to any one of claims 53 or 54, wherein the scFv encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:
732.
62. The composition according to any one of claims 53, 54 and 61, wherein the scFv encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 732, and encodes an scFv comprising SEQ ID NO:
730.
63. The composition according to any one of claims 53, 54, 61 and 62, wherein the scFv encoding sequence comprises, is substantially composed of or is composed of the sequence shown in SEQ ID NO:
732.
64. The composition according to any one of claims 50-63, wherein the multidomain therapeutic protein comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO:
733.
65. The composition according to any one of claims 50-57 and 64, wherein the multi-domain therapeutic protein-coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 756, optionally wherein the nucleotide at position 3 is A, the nucleotide at position 132 is A, the nucleotide at position 273 is T, the nucleotide at position 723 is G, the nucleotide at position 1830 is G, the nucleotide at position 1833 is C, and the nucleotide at position 3078 is G.
66. The composition according to any one of claims 50-57, 64, and 65, wherein the multi-domain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 756, and encodes a multi-domain therapeutic protein comprising SEQ ID NO: 733, optionally wherein the nucleotide at position 3 is A, the nucleotide at position 132 is A, the nucleotide at position 273 is T, the nucleotide at position 723 is G, the nucleotide at position 1830 is G, the nucleotide at position 1833 is C, and the nucleotide at position 3078 is G.
67. The composition according to any one of claims 50-57 and 64-66, wherein the multi-domain therapeutic protein coding sequence comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO:
756.
68. The composition according to any one of claims 50-54, 58-60, and 64, wherein the multi-domain therapeutic protein coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 757, optionally wherein the nucleotide at position 273 is T, the nucleotide at position 723 is G, the nucleotide at position 1830 is G, the nucleotide at position 1833 is C, and the nucleotide at position 3078 is G.
69. The composition according to any one of claims 50-54, 58-60, 64 and 68, wherein the multi-domain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 757, and encodes a multi-domain therapeutic protein comprising SEQ ID NO: 733, optionally wherein the nucleotide at position 273 is T, the nucleotide at position 723 is G, the nucleotide at position 1830 is G, the nucleotide at position 1833 is C, and the nucleotide at position 3078 is G.
70. The composition according to any one of claims 50-54, 58-60, 64, 68 and 69, wherein the multi-domain therapeutic protein coding sequence comprises, is substantially composed of or is composed of the sequence shown in SEQ ID NO:
757.
71. The composition according to any one of claims 50-54 and 64, wherein the multi-domain therapeutic protein coding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 758, optionally wherein the nucleotide at position 3078 is G.
72. The composition according to any one of claims 50-54, 64 and 71, wherein the multi-domain therapeutic protein encoding sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 758, and encodes a multi-domain therapeutic protein comprising SEQ ID NO: 733, optionally wherein the nucleotide at position 3078 is G.
73. The composition according to any one of claims 50-54, 64, 71 and 72, wherein the multi-domain therapeutic protein coding sequence comprises, is substantially composed of or is composed of the sequence shown in SEQ ID NO:
758.
74. The composition according to any one of claims 50-57 and 64-67, wherein the nucleic acid construct from 5' to 3' comprises: a splice acceptor, the coding sequence of the multi-domain therapeutic protein, and a polyadenylation signal or sequence. The coding sequence of the multi-domain therapeutic protein comprises SEQ ID NO: 756, and optionally the nucleic acid construct comprises the sequence shown in SEQ ID NO: 793, or optionally the nucleic acid construct comprises the sequence shown in SEQ ID NO:
777. The polyadenylation signal comprises a BGH polyadenylation signal and a unidirectional SV40 late polyadenylation signal, optionally wherein the BGH polyadenylation signal comprises the sequence shown in SEQ ID NO: 751 and the unidirectional SV40 late polyadenylation signal comprises the sequence shown in SEQ ID NO: 752, optionally the polyadenylation signal comprising the BGH polyadenylation signal and the unidirectional SV40 late polyadenylation signal comprises the sequence shown in SEQ ID NO:
795. The nucleic acid construct described herein does not contain a promoter that drives the expression of the multi-domain therapeutic protein, and The nucleic acid constructs described herein do not contain homologous arms.
75. The composition according to any one of claims 50-57 and 64-67, wherein the nucleic acid construct from 5' to 3' comprises: a splice acceptor, the coding sequence of the multi-domain therapeutic protein, and a polyadenylation signal or sequence. The coding sequence of the multi-domain therapeutic protein comprises SEQ ID NO: 756, and optionally the nucleic acid construct comprises the sequence shown in SEQ ID NO: 794, or optionally the nucleic acid construct comprises the sequence shown in SEQ ID NO:
778. The polyadenylation signal includes a BGH polyadenylation signal, optionally wherein the BGH polyadenylation signal includes the sequence shown in SEQ ID NO:
751. The nucleic acid construct described herein does not contain a promoter that drives the expression of the multi-domain therapeutic protein, and The nucleic acid constructs described herein do not contain homologous arms.
76. The composition according to any one of claims 1-75, wherein the nucleic acid construct is contained in a nucleic acid carrier or lipid nanoparticles.
77. The composition of claim 76, wherein the nucleic acid construct is contained in the nucleic acid vector, optionally wherein the nucleic acid vector is a viral vector.
78. The composition according to claim 76 or 77, wherein the nucleic acid vector is an adeno-associated virus (AAV) vector. Optionally, the nucleic acid construct is side-attached with an inverted terminal repeat (ITR) sequence at each end, and optionally, at least one of the ITRs contains, is substantially composed of, or is composed of SEQ ID NO: 160, and optionally, the ITR at each end contains, is substantially composed of, or is composed of SEQ ID NO:
160.
79. The composition of claim 78, wherein the AAV vector is a single-chain AAV (ssAAV) vector.
80. The composition according to claim 78 or 79, wherein the AAV vector is a recombinant AAV8 (rAAV8) vector, optionally wherein the AAV vector is a single-stranded rAAV8 vector.
81. The composition according to any one of claims 1-80, wherein the composition is combined with a nuclease preparation targeting a nuclease target site in a target genomic locus.
82. The composition of claim 81, wherein the target genomic locus is an albumin gene, optionally wherein the albumin gene is a human albumin gene.
83. The composition according to claim 82, wherein the nuclease target site is located in intron 1 of the albumin gene.
84. The composition according to any one of claims 81-83, wherein the nuclease preparation comprises: (a) Zinc finger nucleases (ZFN); (b) Transcription activator-like effector nucleases (TALENs); or (c)(i) the Cas protein or the nucleic acid encoding the Cas protein; and (ii) a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises a DNA targeting segment that targets the guide RNA target sequence, and wherein the guide RNA binds to the Cas protein and targets the Cas protein to the guide RNA target sequence.
85. The composition according to any one of claims 81-83, wherein the nuclease preparation comprises: (a) The Cas protein or the nucleic acid encoding the Cas protein; as well as (b) A guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises a DNA targeting segment that targets the guide RNA target sequence, and wherein the guide RNA binds to the Cas protein and targets the Cas protein to the guide RNA target sequence.
86. The composition of claim 85, wherein the guide RNA target sequence is located in intron 1 of the albumin gene.
87. The composition according to claim 85 or 86, wherein the DNA targeting segment comprises any one of SEQ ID NO: 36, 30-35, and 37-61, optionally wherein the DNA targeting segment comprises any one of SEQ ID NO: 36, 30, 33, and 41, or The DNA targeting segment is composed of any one of SEQ ID NO: 36, 30-35 and 37-61, and optionally the DNA targeting segment is composed of any one of SEQ ID NO: 36, 30, 33 and 41.
88. The composition according to any one of claims 85 to 87, wherein the guide RNA comprises any one of SEQ ID NO: 68, 100, 62-67, 69-99 and 101-125, and optionally wherein the guide RNA comprises any one of SEQ ID NO: 68, 100, 62, 94, 65, 97, 73 and 105.
89. The composition according to any one of claims 85-88, wherein the DNA targeting segment comprises or is composed of SEQ ID NO:
36.
90. The composition according to any one of claims 85-89, wherein the guide RNA comprises SEQ ID NO: 68 or 100.
91. The composition according to any one of claims 85-90, wherein the composition comprises the guide RNA in the form of RNA.
92. The composition according to any one of claims 85-91, wherein the guide RNA comprises at least one modification.
93. The composition according to claim 92, wherein the at least one modification comprises: (i) the phosphate thioester bond between the first four nucleotides at the 5' end of the guide RNA; (ii) the phosphate thioester bond between the last four nucleotides at the 3' end of the guide RNA; (iii) the 2'-O-methyl modified nucleotides at the first three nucleotides at the 5' end of the guide RNA; and (iv) the 2'-O-methyl modified nucleotides at the last three nucleotides at the 3' end of the guide RNA.
94. The composition according to any one of claims 85-93, wherein the composition comprises the guide RNA in RNA form, the guide RNA comprising SEQ ID NO: 100, and the guide RNA comprising: (i) a phosphate thioester bond between the first four nucleotides at the 5' end of the guide RNA; (ii) a phosphate thioester bond between the last four nucleotides at the 3' end of the guide RNA; (iii) a 2'-O-methyl modified nucleotide at the first three nucleotides at the 5' end of the guide RNA; and (iv) a 2'-O-methyl modified nucleotide at the last three nucleotides at the 3' end of the guide RNA.
95. The composition according to any one of claims 85-94, wherein the Cas protein is a Cas9 protein, optionally wherein the Cas protein is derived from Streptococcus pyogenes Cas9 protein.
96. The composition according to any one of claims 85-95, wherein the Cas protein comprises the sequence shown in SEQ ID NO:
11.
97. The composition according to any one of claims 85-96, wherein the composition comprises the nucleic acid encoding the Cas protein, wherein the nucleic acid comprises mRNA encoding the Cas protein.
98. The composition of claim 97, wherein the mRNA encoding the Cas protein comprises at least one modification.
99. The composition of claim 98, wherein the mRNA encoding the Cas protein is completely replaced by N1-methyl-pseuuridine.
100. The composition according to any one of claims 97-99, wherein the mRNA encoding the Cas protein comprises the sequence shown in SEQ ID NO: 1 or 2.
101. The composition according to any one of claims 85-100, wherein the composition comprises the nucleic acid encoding the Cas protein, wherein the nucleic acid comprises mRNA encoding the Cas protein, the mRNA encoding the Cas protein comprises the sequence shown in SEQ ID NO: 1 or 2, and the mRNA encoding the Cas protein is completely substituted with N1-methyl-pseuuridine, comprises a 5' cap, and comprises a poly(A) tail.
102. The composition according to any one of claims 85-101, wherein the composition comprises the guide RNA in RNA form, and the guide RNA comprises SEQ ID NO: 68 or 100, and The composition comprises administering the nucleic acid encoding the Cas protein, wherein the nucleic acid comprises mRNA encoding the Cas protein, and the mRNA encoding the Cas protein comprises the sequence shown in SEQ ID NO: 1 or 2.
103. The composition according to any one of claims 85-102, wherein the composition comprises the guide RNA in RNA form, the guide RNA comprising SEQ ID NO: 100, and the guide RNA comprising: (i) a phosphate thioester bond between the first four nucleotides at the 5' end of the guide RNA; (ii) a phosphate thioester bond between the last four nucleotides at the 3' end of the guide RNA; (iii) a 2'-O-methyl-modified nucleotide at the first three nucleotides at the 5' end of the guide RNA; and (iv) a 2'-O-methyl-modified nucleotide at the last three nucleotides at the 3' end of the guide RNA, and The composition comprises the nucleic acid encoding the Cas protein, wherein the nucleic acid comprises mRNA encoding the Cas protein, the mRNA encoding the Cas protein comprises the sequence shown in SEQ ID NO: 1 or 2, and the mRNA encoding the Cas protein is completely substituted with N1-methyl-pseuuridine, comprises a 5' cap, and comprises a poly(A) tail.
104. The composition according to any one of claims 85-103, wherein the Cas protein or the nucleic acid encoding the Cas protein and the guide RNA or one or more DNAs encoding the guide RNA are associated with lipid nanoparticles.
105. The composition of claim 104, wherein the lipid nanoparticles comprise cationic lipids, neutral lipids, accessory lipids, and stealth lipids.
106. The composition according to claim 105, wherein the cationic lipid is lipid A ((9Z,12Z)-3-((4,4-bis(octyloxy)butyryl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyloctadec-9,12-dienoate), and / or The neutral lipids mentioned above are distearylphosphatidylcholine or 1,2-distearyl-sn-glycerol-3-phosphate choline (DSPC), and / or The accessory lipid mentioned above is cholesterol, and / or The elusive lipid is 1,2-dimyristic-racemic-glycerol-3-methoxypolyethylene glycol-2000.
107. The composition of claim 106, wherein the cationic lipid is lipid A, the neutral lipid is DSPC, the accessory lipid is cholesterol, and the occult lipid is PEG2k-DMG.
108. The composition according to any one of claims 105-107, wherein the lipid nanoparticles comprise four lipids in the following molar ratios: about 50 mol% lipid A, about 9 mol% DSPC, about 38 mol% cholesterol, and about 3 mol% PEG2k-DMG.
109. A cell comprising the composition according to any one of claims 1-108.
110. The cell of claim 109, wherein the coding sequence of the nucleic acid construct or the multi-domain therapeutic protein is integrated into a target genomic locus, and wherein the multi-domain therapeutic protein is expressed by the target genomic locus, or wherein the coding sequence of the nucleic acid construct or the multi-domain therapeutic protein is integrated into intron 1 of an endogenous albumin locus, and wherein the multi-domain therapeutic protein is expressed by the endogenous albumin locus.
111. The cell of claim 110, wherein the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the integrated nucleic acid construct or the multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
112. The cell according to any one of claims 109-111, wherein the cell is a liver cell or hepatocyte.
113. The cell according to any one of claims 109-112, wherein the cell is a human cell.
114. A method for inserting a nucleic acid encoding a multi-domain therapeutic protein into a target genomic locus in a cell or cell population, said multi-domain therapeutic protein comprising a delivery domain fused to a lysosomal α-glucosidase, said method comprising administering to said cell or cell population a composition according to any one of claims 81-108. The nuclease preparation cleaves the nuclease target site in the target genomic locus, and the nucleic acid construct or the nucleic acid encoding the multi-domain therapeutic protein is inserted into the target genomic locus.
115. The method of claim 114, wherein the percentage of unintended transcripts from the target genomic locus of the nucleic acid containing the inserted nucleic acid construct or encoding the multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
116. A method for expressing a multi-domain therapeutic protein in a cell or cell population, said multi-domain therapeutic protein comprising a delivery domain fused to a lysosomal α-glucosidase, said method comprising administering to said cell or cell population a composition according to any one of claims 1-80. The coding sequence of the multi-domain therapeutic protein is operatively linked to a promoter in the nucleic acid construct and expressed in the cell or cell population.
117. A method for expressing a multi-domain therapeutic protein from a target genomic locus in a cell or cell population, said multi-domain therapeutic protein comprising a delivery domain fused to a lysosomal α-glucosidase, said method comprising administering to said cell or cell population a composition according to any one of claims 81-108. The nuclease preparation cleaves the nuclease target site in the target genome locus, the coding sequence of the nucleic acid construct or the multi-domain therapeutic protein is inserted into the target genome locus to produce a modified target genome locus, and the multi-domain therapeutic protein containing the delivery domain fused with the lysosomal α-glucosidase is expressed by the modified target genome locus.
118. The method of claim 117, wherein the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the inserted nucleic acid construct or the multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
119. The method according to any one of claims 114-118, wherein the cell is a liver cell or hepatocyte, or the cell population is a liver cell or hepatocyte population.
120. The method according to any one of claims 114-119, wherein the cell is a human cell, or the cell population is a human cell population.
121. The method according to any one of claims 114-120, wherein the cell is a neonatal cell, or the cell population is a neonatal cell population.
122. The method of claim 121, wherein the neonatal cells or the neonatal cell population are derived from a human neonatal subject within 24 weeks of birth, optionally wherein the neonatal cells or the neonatal cell population are derived from a human neonatal subject within 12 weeks of birth, optionally wherein the neonatal cells or the neonatal cell population are derived from a human neonatal subject within 8 weeks of birth, and optionally wherein the neonatal cells or the neonatal cell population are derived from a human neonatal subject within 4 weeks of birth.
123. The method according to any one of claims 114-120, wherein the cells are not neonatal cells, or the cell population is not a neonatal cell population.
124. The method according to any one of claims 114-123, wherein the cells are in vitro or ex vivo, or the cell population is in vitro or ex vivo.
125. The method according to any one of claims 114-123, wherein the cells are in the body of the subject, or the cell population is in the body of the subject.
126. A method of inserting a nucleic acid encoding a multi-domain therapeutic protein into a target genomic locus in a subject's cells, said multi-domain therapeutic protein comprising a delivery domain fused to a lysosomal α-glucosidase, the method comprising administering to the subject the composition according to any one of claims 81-108. The nuclease preparation cleaves the nuclease target site in the target genomic locus, and the nucleic acid construct or the nucleic acid encoding the multi-domain therapeutic protein is inserted into the target genomic locus.
127. The method of claim 126, wherein the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the inserted nucleic acid construct or the multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
128. A method for expressing a multi-domain therapeutic protein in the cells of a subject, said multi-domain therapeutic protein comprising a delivery domain fused to a lysosomal α-glucosidase protein, said method comprising administering to the subject the composition according to any one of claims 1-80. The coding sequence of the multi-domain therapeutic protein is operatively linked to a promoter in the nucleic acid construct and expressed in the cell.
129. A method for expressing a multi-domain therapeutic protein from a target genomic locus in a subject's cells, said multi-domain therapeutic protein comprising a delivery domain fused to a lysosomal α-glucosidase protein, said method comprising administering to the subject the composition according to any one of claims 81-108. The nuclease preparation cleaves the nuclease target site in the target genome locus, the coding sequence of the nucleic acid construct or the multi-domain therapeutic protein is inserted into the target genome locus to produce a modified target genome locus, and the multi-domain therapeutic protein containing the delivery domain fused with the lysosomal α-glucosidase is expressed by the modified target genome locus.
130. The method of claim 129, wherein the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the inserted nucleic acid construct or the multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
131. The method according to any one of claims 128-130, wherein the expressed multi-domain therapeutic protein is delivered to and internalized in the skeletal muscle and cardiac tissues of the subject, or wherein the expressed multi-domain therapeutic protein is delivered to and internalized in the skeletal muscle, cardiac, and central nervous system tissues of the subject.
132. The method according to any one of claims 126-131, wherein the cell is a liver cell or hepatocyte.
133. The method according to any one of claims 126-132, wherein the cell is a human cell.
134. The method according to any one of claims 126-133, wherein the cell is a neonatal cell.
135. The method of claim 134, wherein the newborn subject is a human subject within 24 weeks of birth, optionally wherein the newborn subject is a human subject within 12 weeks of birth, optionally wherein the newborn subject is a human subject within 8 weeks of birth, and optionally wherein the newborn subject is a human subject within 4 weeks of birth.
136. The method according to any one of claims 126-133, wherein the cell is not a neonatal cell.
137. A method for treating lysosomal α-glucosidase deficiency in a subject of need, the method comprising administering to the subject the composition according to any one of claims 1-80, The coding sequence of the multi-domain therapeutic protein is operatively linked to a promoter in the nucleic acid construct and expressed in the subject.
138. A method for treating lysosomal α-glucosidase deficiency in a subject of need, the method comprising administering to the subject the composition according to any one of claims 81-108, The nuclease preparation cleaves the nuclease target site in the target genome locus, the coding sequence of the nucleic acid construct or the multi-domain therapeutic protein is inserted into the target genome locus to produce a modified target genome locus, and the multi-domain therapeutic protein containing the delivery domain fused with the lysosomal α-glucosidase is expressed by the modified target genome locus.
139. The method of claim 138, wherein the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the inserted nucleic acid construct or the multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
140. A method for reducing glycogen accumulation in the tissues of a subject in need, the method comprising administering to the subject the composition according to any one of claims 1-80. The coding sequence of the multi-domain therapeutic protein is operatively linked to a promoter in the nucleic acid construct and expressed in the subject to reduce glycogen accumulation in the tissue.
141. A method for reducing glycogen accumulation in the tissues of a subject in need, the method comprising administering to the subject the composition according to any one of claims 81-108, The nuclease preparation cleaves the nuclease target site in the target genomic locus, the coding sequence of the nucleic acid construct or the multi-domain therapeutic protein is inserted into the target genomic locus to produce a modified target genomic locus, and the multi-domain therapeutic protein containing the delivery domain fused with the lysosomal α-glucosidase is expressed by the modified target genomic locus and reduces glycogen accumulation in the tissue.
142. The method of claim 141, wherein the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the inserted nucleic acid construct or the multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
143. The method according to any one of claims 125-142, wherein the subject suffers from Pompe disease.
144. A method of treating Pompe disease in a subject of need, the method comprising administering to the subject the composition according to any one of claims 1-80, The coding sequence of the multi-domain therapeutic protein is operatively linked to a promoter in the nucleic acid construct and expressed in the subject, thereby treating Pompe disease.
145. A method of treating Pompe disease in a subject of need, the method comprising administering to the subject the composition according to any one of claims 81-108, The nuclease preparation cleaves the nuclease target site in the target genomic locus, the coding sequence of the nucleic acid construct or the multi-domain therapeutic protein is inserted into the target genomic locus to produce a modified target genomic locus, and the multi-domain therapeutic protein containing the delivery domain fused with the lysosomal α-glucosidase is expressed by the modified target genomic locus, thereby treating Pompe disease.
146. The method of claim 145, wherein the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the inserted nucleic acid construct or the multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
147. The method according to any one of claims 143-146, wherein the Pompe disease is an infancy-onset Pompe disease.
148. The method according to any one of claims 143-146, wherein the Pompe disease is late-onset Pompe disease.
149. The method according to any one of claims 125-148, wherein the subject is a human subject.
150. The method according to any one of claims 125-149, wherein the subject is a neonatal subject, optionally wherein the neonatal subject is a human subject within 24 weeks, 12 weeks, 8 weeks or 4 weeks after birth.
151. The method according to any one of claims 125-149, wherein the subject is not a neonatal subject.
152. The method according to any one of claims 125-151, wherein the method results in the production of a therapeutically effective level of a circulating multidomain therapeutic protein or lysosomal α-glucosidase in the subject.
153. The method according to any one of claims 125-152, wherein the method reduces glycogen accumulation in the skeletal muscle, cardiac tissue, or central nervous system tissue of the subject, optionally wherein the method reduces glycogen accumulation in the skeletal muscle, cardiac tissue, and central nervous system tissue of the subject, optionally wherein the method reduces glycogen levels in the skeletal muscle, cardiac, and central nervous system tissues of the subject to levels comparable to wild-type levels at the same age, or The method wherein said method reduces glycogen accumulation in the skeletal muscle or cardiac tissue of the subject, optionally said method reduces glycogen accumulation in the skeletal muscle and cardiac tissue of the subject, optionally said method reduces glycogen levels in the skeletal muscle and cardiac tissue of the subject to levels comparable to wild-type levels at the same age.
154. The method according to any one of claims 125-153, wherein the method improves the muscle strength of the subject or prevents the subject from losing muscle strength compared to a control subject.
155. The method of claim 154, wherein the method results in the subject having muscle strength comparable to that of a wild-type subject of the same age.
156. A method for preventing or reducing the onset of signs or symptoms of Pompe disease in a subject in need, the method comprising administering to the subject the composition according to any one of claims 1-80. The coding sequence of the multi-domain therapeutic protein is operatively linked to a promoter in the nucleic acid construct and expressed in the subject, thereby preventing or reducing the onset of the signs or symptoms of Pompe disease in the subject.
157. A method for preventing or reducing the onset of signs or symptoms of Pompe disease in a subject in need, the method comprising administering to the subject the composition according to any one of claims 81-108. The nuclease preparation cleaves the nuclease target site, the coding sequence of the nucleic acid construct or the multi-domain therapeutic protein is inserted into the target genomic locus to produce a modified target genomic locus, and the multi-domain therapeutic protein containing the delivery domain fused with the lysosomal α-glucosidase is expressed by the modified target genomic locus, thereby preventing or reducing the onset of the signs or symptoms of Pompe disease in the subject.
158. The method of claim 157, wherein the percentage of unintended transcripts from the target genomic locus comprising the coding sequence of the inserted nucleic acid construct or the multi-domain therapeutic protein is less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%.
159. The method according to any one of claims 156-158, wherein the Pompe disease is an infancy-onset Pompe disease.
160. The method according to any one of claims 156-158, wherein the Pompe disease is late-onset Pompe disease.
161. The method according to any one of claims 156-160, wherein the method results in the production of a therapeutically effective level of a circulating multidomain therapeutic protein or lysosomal α-glucosidase in the subject.
162. The method according to any one of claims 156-161, wherein the method prevents or reduces glycogen accumulation in the skeletal muscle, heart or central nervous system tissues of the subject.
163. The method according to any one of claims 156-162, wherein the method prevents or reduces glycogen accumulation in the skeletal muscle, heart, and central nervous system tissues of the subject, or The method described therein prevents or reduces glycogen accumulation in the skeletal muscle and cardiac tissues of the subject.
164. The method according to any one of claims 156-163, wherein the subject is a human subject.
165. The method according to any one of claims 156-164, wherein the subject is a neonatal subject.
166. The method of claim 165, wherein the newborn subject is a human subject within 24 weeks of birth, optionally wherein the newborn subject is a human subject within 12 weeks of birth, optionally wherein the newborn subject is a human subject within 8 weeks of birth, and optionally wherein the newborn subject is a human subject within 4 weeks of birth.
167. The method according to any one of claims 156-164, wherein the subject is not a neonatal subject.
168. The method according to any one of claims 125-167, wherein, compared to a method comprising administering to a control subject a free expression vector encoding the multi-domain therapeutic protein, the method results in an increase in the expression of the multi-domain therapeutic protein in the subject.
169. The method according to any one of claims 125-168, wherein, compared to a method comprising administering to a control subject a free expression vector encoding the multi-domain therapeutic protein, the method results in an increase in serum levels of the multi-domain therapeutic protein in the subject.
170. The method according to any one of claims 125-169, wherein the method results in the serum level of the multi-domain therapeutic protein in the subject being at least about 1 μg / mL, at least about 2 μg / mL, at least about 3 μg / mL, at least about 4 μg / mL, at least about 5 μg / mL, at least about 6 μg / mL, at least about 7 μg / mL, at least about 8 μg / mL, at least about 9 μg / mL, or at least about 10 μg / mL.
171. The method according to any one of claims 125-170, wherein the method results in the serum level of the multi-domain therapeutic protein in the subject being at least about 2 μg / mL or at least about 5 μg / mL.
172. The method according to any one of claims 125-171, wherein the method results in a serum level of the multi-domain therapeutic protein in the subject being between about 2 μg / mL and about 30 μg / mL, or between about 2 μg / mL and about 20 μg / mL.
173. The method according to any one of claims 125-172, wherein the method results in a serum level of the multi-domain therapeutic protein in the subject being between about 5 μg / mL and about 30 μg / mL, or between about 5 μg / mL and about 20 μg / mL.
174. The method according to any one of claims 125-173, wherein the method achieves a lysosomal α-glucosidase activity level of at least about 40%, at least about 45%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or 100% of the normal value.
175. The method according to any one of claims 125-174, wherein: (I) The subject has Pompe disease with onset in infancy, and the method achieves at least about 1% or more of the normal value of lysosomal α-glucosidase expression or activity level; or (II) The subject has late-onset Pompe disease and the method achieves at least about 40% of the normal value or more than about 40% of the normal value of lysosomal α-glucosidase expression or activity level.
176. The method of any one of claims 125 to 175, wherein the expression or activity of the multi-domain therapeutic protein at 24 weeks after administration is at least 50% of the expression or activity of the multi-domain therapeutic protein at the peak expression level measured against the subject.
177. The method of any one of claims 125 to 176, wherein the expression or activity of the multi-domain therapeutic protein at one year after administration is at least 50% of the expression or activity of the multi-domain therapeutic protein at the peak expression level measured against the subject.
178. The method of any one of claims 125 to 177, wherein the expression or activity of the multidomain therapeutic protein at 24 weeks after administration is at least 60% of the expression or activity of the multidomain therapeutic protein at the peak expression level measured against the subject.
179. The method of any one of claims 125 to 178, wherein the expression or activity of the multidomain therapeutic protein at two years after administration is at least 50% of the expression or activity of the multidomain therapeutic protein at the peak expression level measured against the subject.
180. The method of any one of claims 125 to 179, wherein the expression or activity of the multi-domain therapeutic protein at 2 years after administration is at least 60% of the expression or activity of the multi-domain therapeutic protein at the peak expression level measured against the subject.
181. The method of any one of claims 125 to 180, wherein the expression or activity of the multi-domain therapeutic protein at 24 weeks after administration is at least 60% of the expression or activity of the multi-domain therapeutic protein at the peak expression level measured against the subject.
182. The method according to any one of claims 125-181, wherein the method further comprises assessing the subject's pre-existing AAV immunity prior to administering the nucleic acid construct to the subject.
183. The method of claim 182, wherein the pre-existing AAV immunity is pre-existing AAV8 immunity.
184. The method of claim 182 or 183, wherein assessing pre-existing AAV immunity comprises assessing immunogenicity using a total antibody immunoassay or a neutralizing antibody assay.
185. The method according to any one of claims 114-184, wherein the nucleic acid construct is administered simultaneously with the nuclease preparation or one or more nucleic acids encoding the nuclease preparation.
186. The method according to any one of claims 114-184, wherein the nucleic acid construct is not administered simultaneously with the nuclease preparation or one or more nucleic acids encoding the nuclease preparation.
187. The method of claim 186, wherein the nucleic acid construct is administered prior to the nuclease preparation or one or more nucleic acids encoding the nuclease preparation.
188. The method of claim 186, wherein the nucleic acid construct is administered after the nuclease preparation or one or more nucleic acids encoding the nuclease preparation.
Citation Information
Patent Citations
Semiconductor device
US11211123B2
Load-sustainer.
US1202530A
Methods and compositions for using zinc finger endonucleases to enhance homologous recombination
US20030232410A1
Use of chimeric nucleases to stimulate gene targeting
US20050026157A1
Targeted chromosomal mutagenasis using zinc finger nucleases
US20050208489A1