ATXN2 RNA interference agents
ATXN2 RNAi agents with a TfR binding domain are developed to cross the BBB and reduce ATXN2 expression, addressing the need for effective treatment of ATXN2-associated neurological diseases.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ELI LILLY & CO
- Filing Date
- 2025-11-21
- Publication Date
- 2026-05-28
AI Technical Summary
There is a need for therapeutic agents that can inhibit or adjust the expression of ATXN2 for treating ATXN2-associated neurological diseases and deliver ATXN2 RNAi agents across the blood-brain barrier (BBB) into the central nervous system (CNS) effectively.
Development of ATXN2 RNA interference (RNAi) agents comprising a double-stranded RNA (dsRNA) with a monovalent human transferrin receptor (TfR) binding domain, which can cross the BBB and reduce ATXN2 expression by administering intravenously or subcutaneously.
The ATXN2 RNAi agents effectively cross the BBB and reduce ATXN2 expression, providing a potential treatment for ATXN2-associated neurological diseases.
Smart Images

Figure IMGF000049_0001 
Figure IMGF000052_0001 
Figure IMGF000052_0002
Abstract
Description
ATXN2 RNA INTERFERENCE AGENTSSEQUENCE LISTING
[0001] The present application is being filed along with a Sequence Listing in ST.26 XML format. The Sequence Listing is provided as a file titled “31271 _WO” created September 8, 2025, and is 183,063 bytes in size. The Sequence Listing information in the ST.26 XML format is incorporated herein by reference in its entirety.BACKGROUND
[0002] Ataxin-2 is encoded by the gene ATXN2 (also known as ATX2, SCA2, TNRC13). The ATXN2 gene includes CAG trinucleotide repeats, which result in a polyglutamine (polyQ) stretch in the N-terminal region of the Ataxin-2 protein. More than 90% of normal individuals have an ATXN2 allele with 22 polyQ repeats. CAG expansions of more than 22 repeats in ATXN2 are associated with certain neurodegenerative diseases.
[0003] Spinocerebellar ataxia 2 (SCA2) is caused by CAG expansion of 33 repeats or more in ATXN-2. The most common SCA2-associated ATXN2 alleles have 37-39 CAG repeats; the longer CAG repeat expansions are associated with earlier onset of SCA2. SCA2 is an autosomal dominant neurodegenerative disease characterized by progressive degeneration of neurons in the cerebellum, brain stem, and / or spinal cord. Patients with SCA2 show progressive incoordination of gait and often poor coordination of hands, speech and eye movements, likely due to cerebellum degeneration with variable involvement of the brainstem and spinal cord. Moderate CAG expansion (23 or more repeats but below the threshold for SCA2) in the ATXN2 gene is also associated with amyotrophic lateral sclerosis (ALS).
[0004] RNA interference (RNAi) is a highly conserved regulatory mechanism in which RNA molecules are involved in sequence-specific suppression of gene expression by doublestranded RNA molecules (dsRNA) (Fire et al., Nature 391:806-811, 1998).
[0005] The blood brain barrier (BBB) is a selective semipermeable border of capillary endothelial cells that prevents solutes, including pathogens, from passing into the central nervous system (CNS). The BBB allows the passage of some small molecules by passive diffusion and the cells of BBB actively transport metabolic products crucial to neural function such as glucose and amino acids across the barrier using specific transport proteins. The BBB has neuroprotective function by tightly controlling access to the brain; but it also impedes access of therapeutic agents to CNS. Antibodies directed to transferrin receptor(“TfR’’) have been used for modulating BBB transport (see e.g., W02024 / 036096).However, there are no approved TfR shuttles or conjugates for the treatment of CNS diseases in the United States and no disease modifying treatment available for ATXN2 associated diseases.
[0006] There remains a need for therapeutic agents that can inhibit or adjust the expression of ATXN2 for treating ATXN2-associated neurological diseases, e.g., by utilizing RNAi. There is also need for conjugates that can deliver ATXN2 RNAi agent across the BBB into the CNS and achieve good distribution.SUMMARY OF INVENTION
[0007] Provided herein are ATXN2 RNAi agents, including ATXN2 RNAi agents capable of crossing BBB, and compositions comprising such ATXN2 RNAi agent. Such ATXN2 RNAi agent and composition can be administered intravenously or subcutaneously. Also provided herein are methods of using such ATXN2 RNAi agents or compositions comprising a ATXN2 RNAi agent for reducing ATXN2 expression, and / or treating a ATXN2-associated neurological disease in a subject.
[0008] In one aspect, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the antisense strand is complementary7to ATXN2 mRNA; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain ("‘human TfR binding protein’7); and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3.
[0009] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the antisense strand is complementary to ATXN2 mRNA; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain, wherein the human TfR binding domain comprises a heavy chain variable region (VH) and a light chain variable region (VL), wherein the VH comprises heavy chain complementarity determining regions HCDR1, HCDR2, and HCDR3, and the VL comprises light chain complementarity determining regions LCDR1, LCDR2. and LCDR3. wherein HCDR1 comprises SEQ ID NO: 1, HCDR2 comprises SEQ ID NO: 2, HCDR3 comprises SEQ ID NO: 3, LCDR1 comprises SEQ ID NO: 4, LCDR2 comprises SEQ ID NO: 5, and LCDR3 comprises SEQ ID NO: 6; and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In someembodiments, VH comprises SEQ ID NO: 7 and VL comprises SEQ ID NO: 8. In some embodiments, VH comprises a sequence having at least 95% sequence identity to SEQ ID NO: 7 and VL comprises a sequence having at least 95% sequence identity to SEQ ID NO: 8. Exemplary sequences of human TfR binding domains and proteins are provided in Table la and lb.
[0010] In some embodiments, L is a SMCC linker, OD linker, or MSPT linker (see Table 5). In some embodiments, L is a MSPT linker in Table 5.
[0011] Exemplary unmodified sense strand and antisense strand sequences of dsRNA targeting human ATXN2 mRNA are provided in Table 3a. In some embodiments, the sense strand and the antisense strand comprise a pair of nucleic acid sequences selected from the group consisting of:(a) the sense strand comprises SEQ ID NO: 35, and the antisense strand comprises SEQ ID NO: 36,(b) the sense strand comprises SEQ ID NO: 37, and the antisense strand comprises SEQ ID NO: 38,(c) the sense strand comprises SEQ ID NO: 39, and the antisense strand comprises SEQ ID NO: 40,(d) the sense strand comprises SEQ ID NO: 41, and the antisense strand comprises SEQ ID NO: 36,(e) the sense strand comprises SEQ ID NO: 42, and the antisense strand comprises SEQ ID NO: 38, and(f) the sense strand comprises SEQ ID NO: 43, and the antisense strand comprises SEQ ID NO: 40,wherein optionally one or more nucleotides of the sense strand and the antisense strand are independently modified nucleotides, and wherein optionally one or more intemucleotide linkages of the sense strand and the antisense strand are modified intemucleotide linkages.
[0012] The dsRNA can include modifications. The modifications can be made to one or more nucleotides of the sense and / or antisense strand or to the intemucleotide linkages. In some embodiments, one or more nucleotides of the sense strand and / or the antisense strand are independently modified nucleotides, which means the sense strand and the antisense strand can have different modified nucleotides. In some embodiments, each nucleotide of the sense strand is a modified nucleotide. In some embodiments, each nucleotide of the antisense strand is a modified nucleotide. In some embodiments, the modified nucleotide is a 2'-fluoro modified nucleotide, 2'-O-methyl modified nucleotide, 2’ deoxy nucleotide (DNA), or 2'-O-alkyl (e.g., 2 -0-C 16 alkyl) modified nucleotide. In some embodiments, each nucleotide of the sense strand and the antisense strand is independently a modified nucleotide, e.g., a 2'-fluoro modified nucleotide, 2'-O-methyl modified nucleotide, 2’ deoxy nucleotide (DNA), or 2'-O-alkyl (e.g., 2 -O-C 16 alkyl) modified nucleotide.
[0013] In some embodiments, the sense strand has four 2'-fluoro modified nucleotides, e.g., at positions 7, 9, 10, 11 from the 5’ end of the sense strand. In some embodiments, at least one nucleotide of the sense strand is an unmodified RNA nucleotide. In some embodiments, at least one nucleotide of the sense strand is 2’ deoxy nucleotide (DNA). In some embodiments, the other nucleotides of the sense strand are 2'-O-methyl modified nucleotides. In some embodiments, the antisense strand has four 2'-fluoro modified nucleotides, e.g.. at positions 2, 6, 14, 16 from the 5' end of the antisense strand. In some embodiments, the other nucleotides of the antisense strand are 2'-O-methyl modified nucleotides.
[0014] In some embodiments, the sense strand has three 2'-fluoro modified nucleotides, e.g., at positions 9, 10, 11 from the 5’ end of the sense strand. In some embodiments, at least one nucleotide of the sense strand is an unmodified RNA nucleotide. In some embodiments, at least one nucleotide of the sense strand is 2’ deoxy nucleotide (DNA). In some embodiments, the other nucleotides of the sense strand are 2'-O-methyl modified nucleotides. In some embodiments, the antisense strand has five 2'-fluoro modified nucleotides, e.g.. at positions 2, 5, 7. 14. 16 from the 5’ end of the antisense strand. In some embodiments, the antisense strand has five 2'-fluoro modified nucleotides, e.g., at positions 2, 5, 8, 14, 16 from the 5’ end of the antisense strand. In some embodiments, the antisense strand has five 2'-fluoro modified nucleotides, e.g., at positions 2, 3, 7, 14, 16 from the 5’ end of the antisense strand. In some embodiments, the antisense strand has three 2'-fluoro modified nucleotides, e.g., at positions 2, 14, 16 from the 5’ end of the antisense strand. In some embodiments, the other nucleotides of the antisense strand are 2'-O-methyl modified nucleotides.
[0015] In some embodiments, the 5' end of the antisense strand has a phosphate analog, e.g., 5’-vinylphosphonate (5 -VP).
[0016] In some embodiments, the sense strand or the antisense strand comprises an abasic moiety or inverted abasic moiety.
[0017] In some embodiments, the sense strand and the antisense strand have one or more modified intemucleotide linkages. In some embodiments, the modified intemucleotide linkage is phosphorothioate linkage. In some embodiments, the sense strand has four or fivephosphorothioate linkages. In some embodiments, the antisense strand has four or five phosphorothioate linkages. In some embodiments, the sense strand and the antisense strand each has four or five phosphorothioate linkages. In some embodiments, the sense strand has four phosphorothioate linkages and the antisense strand has five phosphorothioate linkages.
[0018] Exemplary7modified sense strand and antisense strand sequences of dsRNA targeting human ATXN2 mRNA are provided in Table 3b. In some embodiments, the sense strand and the antisense strand comprise a pair of nucleic acid sequences selected from the group consisting of:(a) the sense strand comprises SEQ ID NO: 44, and the antisense strand comprises SEQ ID NO: 45,(b) the sense strand comprises SEQ ID NO: 46, and the antisense strand comprises SEQ ID NO: 47,(c) the sense strand comprises SEQ ID NO: 48, and the antisense strand comprises SEQ ID NO: 49.(d) the sense strand comprises SEQ ID NO: 50, and the antisense strand comprises SEQ ID NO: 45,(e) the sense strand comprises SEQ ID NO: 51, and the antisense strand comprises SEQ ID NO: 47, and(I) the sense strand comprises SEQ ID NO: 52, and the antisense strand comprises SEQ ID NO: 49.
[0019] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the antisense strand is complementary to ATXN2 mRNA; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain, and P is selected from TBP1, TBP2, TBP3, TBP4, or TBP5 in Table lb; and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, L is a linker in Table 5.
[0020] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the dsRNA is any dsRNA in Table 3a or 3b; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain; and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, L is a linker in Table 5.
[0021] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the dsRNA is any dsRNA in Table 3a or 3b; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain, and P is selected from TBP1, TBP2, TBP3, TBP4, or TBP5 in Table lb; and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, L is a linker in Table 5.
[0022] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the antisense strand is complementary to ATXN2 mRNA; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain; wherein the human TfR binding domain comprises two heavy chains HC1 and HC2 and one light chain LC1, wherein HC1 comprises SEQ ID NO: 14, LC1 comprises SEQ ID NO: 10, HC2 comprises SEQ ID NO: 15, and wherein n is 1 or 2. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, L is a linker in Table 5.
[0023] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the antisense strand is complementary to ATXN2 mRNA; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain; wherein the human TfR binding domain comprises two heavy chains HC1 and HC2 and one light chain LC1, wherein HC1 comprises SEQ ID NO: 16, LC1 comprises SEQ ID NO: 10, HC2 comprises SEQ ID NO: 17, and wherein n is 1. In some embodiments, L is a linker in Table 5.
[0024] In another aspect, provided herein are methods of treating a ATXN2-associated neurological disease in a patient in need thereof, and such the method comprises administering to the patient an effective amount of the ATXN2 RNAi agent or a pharmaceutical composition described herein. The ATXN2 RNAi agent or a pharmaceutical composition comprising ATXN2 RNAi agent can be administered to the patient intravenously or subcutaneously.
[0025] In another aspect, provided herein are ATXN2 RNAi agents or pharmaceutical compositions comprising a ATXN2 RNAi agent for use in a therapy. Also provided herein are ATXN2 RNAi agents or pharmaceutical compositions comprising a ATXN2 RNAi agent for use in the treatment of a ATXN2-associated neurological disease. Also provided hereinare uses of the ATXN2 RNAi agent in the manufacture of a medicament for treating a ATXN2-associated neurological disease.DETAILED DESCRIPTION
[0026] Provided herein are ATXN2 RNAi agents and compositions comprising a ATXN2 RNAi agent. Also provided herein are methods of using the ATXN2 RNAi agents or compositions comprising a ATXN2 RNAi agent for reducing ATXN2 expression and / or treating a ATXN2-associated neurological disease in a subject.
[0027] In one aspect, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, wherein the antisense strand is complementary to ATXN2 mRNA; wherein L is a linker, or absent: wherein P is a protein comprising one monovalent human TfR binding domain ('‘human TfR binding protein’’); and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. For clarity, when L is absent, R is directly linked to P through a direct bond.
[0028] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (1): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the antisense strand is complementary to ATXN2 mRNA; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain; wherein the human TfR binding domain comprises a heavy chain variable region (VH) and a light chain variable region (VL), wherein the VH comprises heavy chain complementarity determining regions HCDR1, HCDR2, and HCDR3, and the VL comprises light chain complementarity determining regions LCDR1, LCDR2, and LCDR3, wherein HCDR1 comprises SEQ ID NO: 1, HCDR2 comprises SEQ ID NO: 2, HCDR3 comprises SEQ ID NO: 3, LCDR1 comprises SEQ ID NO: 4, LCDR2 comprises SEQ ID NO: 5, and LCDR3 comprises SEQ ID NO: 6; and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, L is a linker in Table 5. For clarity, when L is absent, R is directly linked to P through a direct bond.
[0029] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the antisense strand is complementary to ATXN2 mRNA; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain, and P is selected from TBP1, TBP2, TBP3, TBP4, or TBP5 in Table lb;and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, L is a linker in Table 5. For clarity, when L is absent, R is directly linked to P through a direct bond.
[0030] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the dsRNA is any dsRNA in Table 3a or 3b; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain; and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, L is a linker in Table 5. For clarity, when L is absent, R is directly linked to P through a direct bond.
[0031] In some embodiments, provided herein are ATXN2 RNAi agents comprising Formula (I): (R-L)n-P, wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and the dsRNA is any dsRNA in Table 3a or 3b; wherein L is a linker, or absent; wherein P is a protein comprising one monovalent human TfR binding domain, and P is selected from TBP1, TBP2, TBP3, TBP4, or TBP5 in Table lb; and wherein n is an integer of 1 to 3. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, L is a linker in Table 5. For clarity, when L is absent, R is directly linked to P through a direct bond.
[0032] In another aspect, provided herein are ATXN2 RNAi agents comprising a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, wherein the sense stand and antisense strand sequences are selected from Table 3a or 3b. In some embodiments, ATXN2 RNAi agents comprising any dsRNA in Table 3a or 3b.Human TfR binding proteins
[0033] The ATXN2 RNAi agents described herein comprise a protein comprising one monovalent human TfR binding domain (‘'human TfR binding protein”). Human TfR binding protein of the ATXN2 RNAi agents can bind TfR on BBB and transport the dsRNA into the CNS.
[0034] Exemplary sequences of human TfR binding domains and proteins are provided in Table la and lb. In some embodiments, the monovalent human TfR binding domain comprises a heavy chain variable region (VH) and a light chain variable region (VL), and the VH comprises heavy chain complementarity determining regions HCDR1, HCDR2, and HCDR3, and the VL comprises light chain complementarity determining regionsLCDR1, LCDR2, and LCDR3. In some embodiments, HCDR1 comprises SEQ ID NO: 1, HCDR2 comprises SEQ ID NO: 2, HCDR3 comprises SEQ ID NO: 3, LCDR1 comprises SEQ ID NO: 4, LCDR2 comprises SEQ ID NO: 5, and LCDR3 comprises SEQ ID NO: 6. In some embodiments, VH comprises SEQ ID NO: 7, and VL comprises SEQ ID NO: 8. In some embodiments, VH comprises a sequence having at least 95% sequence identity to SEQ ID NO: 7, and VL comprises a sequence having at least 95% sequence identity to SEQ ID NO: 8.Table la. Exemplary sequences of human TfR binding domains and proteins Region Sequence SEQ ID NO HCDR1 SYSMN 1(KABAT)HCDR2 SISSSSSYIYYADSVKG 2(KABAT)HCDR3 RHGYSNSDAFDN 3(KABAT)LCDR1 RASQGISHYLV 4(KABAT)LCDR2 AASSLQS 5(KABAT)LCDR3 LQHNSYPWT 6(KABAT)VH EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMN 7WVRQAPGKGLEWVSSISSSSSYIYYADSVKGRFTIS RDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSN SDAFDNWGQGTLVTVSS VL DIQMTQSPSAMSASVGDRVTITCRASQGISHYLVW 8FQQKPGKVPKRLIYAASSLQSGVPSRFSGSGSGTEF TLTISSLQPEDFATYYCLQHNSYPWTFGQGTKVEI KFab HC EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMN 9WVRQAPGKGLEWVSSISSSSSYIYYADSVKGRFTIS RDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSN SDAFDNWGQGTLVTVSSASTKGPSVFPLAPSSKST SGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHT FPAVLQSSGLYSLSSVVTVPSSSLGTQTYICNVNHK PSNTKVDKRVEPKCFab / Fab- DIQMTQSPSAMSASVGDRVTITCRASQGISHYLVW 10 VHH / OAH LC FQQKPGKVPKRLIYAASSLQSGVPSRFSGSGSGTEF TLTISSLQPEDFATYYCLQHNSYPWTFGQGTKVEI KRTVAAPSVFIFPPSDEQLKSGTASVVCLLNNFYPR EAKVQWKVDNALQSGNSQESVTEQDSKDSTYSLS STLTLSKADYEKHKVYACEVTHQGLSSPVTKSFNR GECFab-VHH HC EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMN 11 WVRQAPGKGLEWVSSISSSSSYIYYADSVKGRFTIS RDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSN SDAFDNWGQGTLVTVSSASTKGPCVFPLAPSSKST SGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHT FPAVLQSSGLYSLSSVVTVPSSSLGTQTYICNVNHK PSNTKVDKRVEPKCDKTHTGGGGQGGGGQGGGG QGGGGQGGGGQEVQLLESGGGLVQPGGSLRLSCA ASGRYIDETAVAWFRQAPGKGREFVAGIGGGVDI TYYADSVKGRFTISRDNSKNTLYLQMNSLRPEDTA VYYCGARPGRPLITSKVADLYPYWGQGTLVTVSS PPh!gG4 PAA HC EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMN 13WVRQAPGKGLEWVSSISSSSSYIYYADSVKGRFTIS RDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSN SDAFDNWGQGTLVTVSSASTKGPXVFPLAPCSRST SESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTF PAVLQSSGLYSLSSVVTVPSSSLGTKTYTCNVDHK PSNTKVDKRVESKYGPPCPPCPAPEAAGGPSVFLF PPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNW YVDGVEVHNAKTKPREEQFNSTYRVVSVLTVLHQ DWLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREP QVYTLPPSQEEMTKNQVSLTCLVKGFYPSDIAVE WESNGQPENNYKTTPPVLDSDGSFFLYSRLTVDKS RWQEGNVFSCSVMHEALHNHYTQKSLSLSLG, wherein X is S or C.0AH1 (one arm EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMN 14 heteromab) HC1 WVRQAPGKGLEWVSSISSSSSYIYYADSVKGRFTIS (A378C) RDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSN SDAFDNWGQGTLVTVSSASTKGPSVFPLAPCSRST SESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTF PAVLQSSGLYSLSSVVTVPSSSLGTKTYTCNVDHK PSNTKVDKRVESKYGPPCPPCPAPEAAGGPSVFLF PPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNW YVDGVEVHNAKTKPREEQFNSTYRVVSVLTVLHQ DWLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREP QVSTLPPSQEEMTKNQVSLMCLVYGFYPSDIXVE WESNGQPENNYKTTPPVLDSDGSFFLYSVLTVDKS RWQEGNVFSCSVMHEALHNHYTQKSLSLSLG. wherein X is A or C.OAH1 HC2 ESKYGPPCPPCPAPEAAGGPSVFLFPPKPKDTLMIS 15 (A378C) RTPEVTCVVVDVSQEDPEVQFNWYVDGVEVHNA KTKPREEQFNSTYRVVSVLTVLHQDWLNGKEYKC KVSNKGLPSSIEKTISKAKGQPREPQVYTLPPSQGD MTKNQVQLTCLVKGFYPSDIXVEWESNGQPENNY KTTPPVLDSDGSFFLASRLTVDKSRWQEGNVFSCS VMHEALHNHYTQKSLSLSLG, wherein X is A or C. OAH2 HC1 EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMN 16 (S124C) WVRQAPGKGLEWVSSISSSSSYIYYADSVKGRFTIS RDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSN SDAFDNWGQGTLVTVSSASTKGPCVFPLAPCSRSTSESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYSLSSVVTVPSSSLGTKTYTCNVDHK PSNTKVDKRVESKYGPPCPPCPAPEAAGGPSVFLF PPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNW YVDGVEVHNAKTKPREEQFNSTYRVVSVLTVLHQ DWLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREP QVSTLPPSQEEMTKNQVSLMCLVYGFYPSDIAVE WESNGQPENNYKTTPPVLDSDGSFFLYSVLTVDKS RWQEGNVFSCSVMHEALHNHYTQKSLSLSLG0AH2 HC2 ESKYGPPCPPCPAPEAAGGPSVFLFPPKPKDTLMIS 17 (S124C) RTPEVTCVVVDVSQEDPEVQFNWYVDGVEVHNA KTKPREEQFNSTYRVVSVLTVLHQDWLNGKEYKC KVSNKGLPSSIEKTISKAKGQPREPQVYTLPPSQGD MTKNQVQLTCLVKGFYPSDIAVEWESNGQPENNY KTTPPVLDSDGSFFLASRLTVDKSRWQEGNVFSCS VMHEALHNHYTQKSLSLSLGNull Arm HC QVQLVQSGAEVKKPGSSVKVSCKASGYTFSSYAIE 18WVRQAPGQGLEWMGGILPGSGTINYNEKFKGRVT ITADKSTSTAYMELSSLRSEDTAVYYCARMSSNSD QGFDLWGQGTLVTVSSASTKGPXVFPLAPCSRSTS ESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFP AVLQSSGLYSLSSVVTVPSSSLGTKTYTCNVDHKP SNTKVDKRVESKYGPPCPPCPAPEAAGGPSVFLFP PKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWY VDGVEVHNAKTKPREEQFNSTYRVVSVLTVLHQD WLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREPQ VYTLPPSQEEMTKNQVSLTCLVKGFYPSDIAVEWE SNGQPENNYKTTPPVLDSDGSFLLYSKLTVDKSR WQEGNVFSCSVMHEALHNHYTQKSLSLSLG, wherein X is S or C.Null Arm LC DIQMTQSPSSLSASVGDRVTITCKASQGISRFLSWF 19QQKPGKAPKSLIYAVSSLVDGVPSRFSGSGSGTDF TLTISSLQPEDFATYYCVQYNSYPYGFGGGTKVEI KRTVAAPSVFIFPPSDEQLKSGTASVVCLLNNFYPR EAKVQWKVDNALQSGNSQESVTEQDSKDSTYSLS STLTLSKADYEKHKVYACEVTHQGLSSPVTKSFNRGECTable lb. Exemplary sequences of human TfR binding proteinsHuman TfR HC1 LC1 HC2 LC2 bindingprotein (TBP)TBP1 SEQ ID NO: 9 SEQ ID NO: 10 N / A N / A (Fab)TBP2 SEQ ID NO: 11 SEQ ID NO: 10 N / A N / A(Fab-VHH)TBP3 SEQ ID NO: 13 SEQ ID NO: 10 SEQ ID NO: 18 SEQ ID NO: 19 (HeterodimericAb)TBP4 SEQ ID NO: 14 SEQ ID NO: 10 SEQ ID NO: 15 N / A*(One ArmHeteromab 1,A378C)TBP5 SEQ ID NO: 16 SEQ ID NO: 10 SEQ ID NO: 17 N / A*(One ArmHeteromab 2,S124C)
[0035] In some embodiments, the monovalent human TfR binding domain is an antibody fragment, e.g., Fab, scFv, Fv, or scFab (single chain Fab). In some embodiments, the monovalent human TfR binding domain is Fab. In some embodiments, the human TfR binding domain further comprises a heavy chain constant region and / or a light chain constant region.
[0036] In some embodiments, the human TfR binding protein further comprises a half-life extender, e.g.. an immunoglobulin Fc region or a VHH that binds human serum albumin (HSA).
[0037] In some embodiments, the human TfR binding protein further comprises an immunoglobulin Fc region, e.g., a modified human IgG4 Fc region, or a modified human IgGl Fc region. In some embodiments, the human TfR binding protein further comprises a modified human IgG4 Fc region comprising proline at residue 228, and alanine at residues 234 and 235 (all residues are numbered according to the EU Index numbering, also called hIgG4PAA Fc region). In some embodiments, the human TfR binding protein further comprises a modified human IgGl Fc region comprising alanine at residues 234, 235, and 329, serine at position 265, aspartic acid at position 436 (all residues are numbered according to the EU Index numbering, also called hlgGl effector null or hlgGlEN Fc region).
[0038] In some embodiments, the human TfR binding protein further comprise a VHH that binds human HSA. In some embodiments, the VHH also binds mouse, rat, and / or cynomolgus monkey albumin. An exemplary VHH that binds human HSA is shown in Table 1c. In some embodiments, such a VHH comprises CDR1 comprising SEQ ID NO: 20, CDR2 comprising SEQ ID NO: 21, and CDR3 comprising SEQ ID NO: 22. In some embodiments, such a VHH comprises SEQ ID NO: 23. In some embodiments, the VHH is linked to the TfR binding domain through a peptide linker, e.g., (GGGGQ)4 (SEQ ID NO: 12). In some embodiments, the VHH is linked to the C-terminus of the TfR binding domain.Table 1c. Exemplary sequences of VHH that binds human serum albumin (HSA) Region Sequence SEQ ID NO CDR1 ETAVA 20(KABAT)CDR2 GIGGGVDITYYADSVKG 21 (KABAT)CDR3 RPGRPLITSKVADLYPY 22 (KABAT)VHH full EVQLLESGGGLVQPGGSLRLSCAASGRYIDETAV 23length AWFRQAPGKGREFVAGIGGGVDITYYADSVKGR FTISRDNSKNTLYLQMNSLRPEDTAVYYCGARPG RPLITSKVADLYPYWGQGTLVTVSSPPOptional GGGGQGGGGQGGGGQGGGGQ 12linker
[0039] In some embodiments, the human TIR binding protein is heterodimeric antibody that comprises a first arm comprising one monovalent human TfR binding domain and a second arm that is a null arm, e.g., an arm that does not bind any known human target (e.g., an isotype arm). Heterodimeric antibodies such as heteromab, orthomab or duobody have been described in WO2014150973, WO2016118742, WO2018118616, and WO2011131746. In some embodiments, the first arm comprises any monovalent human TfR binding domain described herein. In some embodiments, the second arm is a null arm that does not bind any known human target (e.g., an isotype arm) comprises the sequences in Table la. In some embodiments, the second arm comprises a heavy chain (HC) and a light chain (LC), wherein the HC comprises SEQ ID NO: 18, and the LC comprises SEQ ID NO: 19.
[0040] In some embodiments, the human TfR binding protein comprises heterodimeric mutations. In some embodiments, the human TfR binding protein comprises a modified Fc region comprising a first Fc CH3 domain comprising serine at residue 349, methionine at residue 366. tyrosine at residue 370, and valine at residue 409. and a second Fc CH3 domain comprising glycine at residue 356, aspartic acid at residue 357, glutamine at residue 364 and alanine at residue 407 (all residues are numbered according to the EU Index numbering). In some embodiments, the human TfR binding protein comprises a modified Fc region comprising a first Fc CH3 domain comprising leucine at residue 405, and a second Fc CH3 domain comprising arginine at residue 409 (all residues are numbered according to the EU Index numbering).
[0041] In some embodiments, the human TfR binding protein comprises one or more native cysteine residues, which can be used for conjugation. For example, in some embodiments, the human TfR binding protein comprises a native cysteine at position 220 of the light chain and / or a native cysteine at position 226 of the heavy chain, which can be used for conjugation (all residues according to the EU Index numbering).
[0042] In some embodiments, the human TfR binding protein comprises engineered cysteine residues for conjugation. The approach of including engineered cysteines as a means for conjugation has been described in WO 2018 / 232088. In some embodiments, the human TfR binding protein comprises a heavy chain comprising one or more cysteines at the following residues: 124, 157, 162, 262, 373, 375, 378, 397, 415 (all residues according to the EU Index numbering). In some embodiments, the human TfR binding protein comprises a light chain (e.g., a kappa light chain) comprising one or more cysteines at the following residues: 156, 171, 191, 193, 202, 208 (all residues according to the EU Index numbering). In some embodiments, the human TfR binding protein comprises a heavy chain constant region comprising cysteine at residue 124 (according to the EU Index numbering). In some embodiments, the human TfR binding protein comprises a light chain constant region comprising cysteine at residue 156 (according to the EU Index numbering). In some embodiments, the human TfR binding protein comprises an immunoglobulin Fc region comprising cysteine at residue 378 (according to the EU Index numbering).
[0043] In some embodiments, the human TfR binding protein is any one of the human TfR binding proteins in Table lb, e.g., TBP1, TBP2, TBP3, TBP4, TBP5.
[0044] In some embodiments, the human TfR binding protein has a Fab format, e g., TBP1. In some embodiments, the human TfR binding protein comprises one HC and one LC, and wherein the HC comprises SEQ ID NO: 9 and the LC comprises SEQ ID NO: 10.
[0045] In some embodiments, the human TfR binding protein has a Fab-VHH format, e.g., TBP2. In some embodiments, the human TfR binding proteins comprises one HC and one LC, wherein the HC comprises SEQ ID NO: 11 and the LC comprises SEQ ID NO: 10.
[0046] In some embodiments, the human TfR binding protein has a heterodimeric antibody format, e.g., TBP3. In some embodiments, the human TfR binding protein comprises two heavy chains HC1 and HC2 and two light chains LC1 and LC2, wherein HC1 comprises SEQ ID NO: 13, LC1 comprises SEQ ID NO: 10, HC2 comprises SEQ ID NO: 18, and LC2 comprises SEQ ID NO: 19.
[0047] In some embodiments, the human TfR binding protein has a one arm heteromab format, e.g., TBP4 or TBP5. In some embodiments, the human TfR binding protein comprises two heavy chains HC1 and HC2 and one light chain LC1, wherein HC1 comprises SEQ ID NO: 14, LC1 comprises SEQ ID NO: 10, HC2 comprises SEQ ID NO: 15. In some embodiments, provided herein are human TfR binding proteins comprise two heavy chains HC1 and HC2 and one light chain LC1, wherein HC1 comprises SEQ ID NO: 16. LC1 comprises SEQ ID NO: 10, HC2 comprises SEQ ID NO: 17.
[0048] The human TfR binding proteins described herein can be recombinantly produced in a host cell, for example, using an expression vector. For example, an expression vector may include a sequence that encodes one or more signal peptides that facilitate secretion of the polypeptide(s) from a host cell. Expression vectors containing a polynucleotide of interest (e.g., a polynucleotide encoding a heavy chain or light chain of the TfR binding proteins) may be transferred into a host cell by well-known methods.Additionally, expression vectors may contain one or more selection markers, e.g., tetracycline, neomycin, and dihydrofolate reductase, to aide in detection of host cells transformed with the desired polynucleotide sequences.
[0049] A host cell includes cells stably or transiently transfected, transformed, transduced or infected with one or more expression vectors expressing all or a portion of the TfR binding proteins described herein. According to some embodiments, a host cell may be stably or transiently transfected, transformed, transduced or infected with an expression vector expressing HC polypeptides and an expression vector expressing LC polypeptides of the TfR binding proteins described herein. In some embodiments, a host cell may be stably or transiently transfected, transformed, transduced or infected with an expression vector expressing HC and LC polypeptides of the TfR binding proteins described herein. The TfR binding proteins may be produced in mammalian cells such as CHO, NS0, HEK293 or COS cells according to techniques well known in the art.
[0050] Medium, into which the TfR binding proteins has been secreted, may be purified by conventional techniques, such as mixed-mode methods of ion-exchange and hydrophobic interaction chromatography. For example, the medium may be applied to and eluted from a Protein A or G column using conventional methods; mixed-mode methods of ion-exchange and hydrophobic interaction chromatography may also be used. Soluble aggregate and multimers may be effectively removed by common techniques, including size exclusion, hydrophobic interaction, ion exchange, or hydroxyapatite chromatography.Various methods of protein purification may be employ ed, and such methods are known in the art and described, for example, in Deutscher, Methods in Enzymology’ 182: 83-89 (1990) and Scopes, Protein Purification: Principles and Practice, 3rd Edition, Springer, NY (1994).Mouse TfR binding proteins
[0051] Some conjugates used in the Examples below comprise a protein comprising one monovalent mouse TfR binding domain (‘‘mouse TfR binding proteins” or mTBP). Exemplary sequences of mouse TfR binding proteins are provided in Tables 2A and 2B. Suchconjugates comprising a mouse TfR binding protein can serve as surrogate molecules in mouse models.Table 2A. Exemplary sequences of mouse TfR binding proteinsRegion Sequence SEQ ID NO HCDR1 GSYWIC 24 (KABAT)HCDR2 CIYSTSGGRTYYASWVKG 25 (KABAT)HCDR3 GDDSISDAYFDL 26 (KABAT)LCDR1 QSSQSVYNNNRLA 27 (KABAT)LCDR2 DASTLAS 28 (KABAT)LCDR3 QGTYFSSGWSWA 29 (KABAT)VH QSLEESGGDLVKPEGSLTLTCTASGFSFSGSYWI 30CWVRQAPGKGLEWIGCIYSTSGGRTYYASWVK GRFTISKTSSTTVTLQMTSLTAADTATYFCARG DDSISDAYFDLWGPGTLVTVSS VL ALDMTQTASPVSAAVGGTVTINCQSSQSVYNN 31NRLAWYQQKPGQPPKLLIYDASTLASGVPSRFK GSGSGTQFTLTISGVQSDDSATYYCQGTYFSSG WSWAFGGGTEVVVK HC1 QSLEESGGDLVKPEGSLTLTCTASGFSFSGSYWI 32CWVRQAPGKGLEWIGCIYSTSGGRTYYASWVK GRFTISKTSSTTVTLQMTSLTAADTATYFCARG DDSISDAYFDLWGPGTLVTVSSASTKGPCVFPL APCSRSTSESTAALGCLVKDYFPEPVTVSWNSG ALTSGVHTFPAVLQSSGLYSLSSVVTVPSSSLGT KTYTCNVDHKPSNTKVDKRVESKYGPPCPPCPA PEAAGGPSVFLFPPKPKDTLMISRTPEVTCVVVD VSQEDPEVQFNWYVDGVEVHNAKTKPREEQFN STYRVVSVLTVLHQDWLNGKEYKCKVSNKGLP SSIEKTISKAKGQPREPQVYTLPPSQEEMTKNQV SLTCLVKGFYPSDIAVEWESNGQPENNYKTTPP VLDSDGSFFLYSRLTVDKSRWQEGNVFSCSVM HEALHNHYTQKSLSLSLG LC1 ALDMTQTASPVSAAVGGTVTINCQSSQSVYNN 33NRLAWYQQKPGQPPKLLIYDASTLASGVPSRFK GSGSGTQFTLTISGVQSDDSATYYCQGTYFSSG WSWAFGGGTEVVVKRTVAAPSVFIFPPSDEQLK SGTASVVCLLNNFYPREAKVQWKVDNALQSGN SQESVTEQDSKDSTYSLSSTLTLSKADYEKHKV YACEVTHQGLSSPVTKSFNRGEC HC2 ESKYGPPCPPCPAPEAAGGPSVFLFPPKPKDTLM 34ISRTPEVTCVVVDVSQEDPEVQFNWYVDGVEV HNAKTKPREEQFNSTYRVVSVLTVLHQDWLNG KEYKCKVSNKGLPSSIEKTISKAKGQPREPQVYTLPPSQEEMTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFLLYSKLTVDKSRWQEGNVFSCSVMHEALHNHYTQKSLSLSLGTable 2B. Exemplary sequences of mouse TfR binding proteinsMouse TfR binding HC1 LC1 HC2protein (mTBP)mTBPl SEQ ID NO: 32 SEQ ID NO: 33 SEQ ID NO: 34dsRNA
[0052] The ATXN2 RNAi agents described herein comprise a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, and wherein the antisense strand is complementary to ATXN2 mRNA. After the antisense strand of the dsRNA is incorporated into the RNA-induced silencing complex (RISC), the RISC can bind and degrade target ATXN2 mRNA.
[0053] In some embodiments, the sense strand and the antisense strand of the dsRNA are each 15-30 nucleotides in length, e.g., 20-25 nucleotides in length. In some embodiments, the dsRNA has a sense strand of 21 nucleotides and an antisense strand of 23 nucleotides. In some embodiments, the sense strand and antisense strand of the dsRNA may have overhangs at either the 5‘ end or the 3’ end (i.e., 5’ overhang or 3‘ overhang). For example, the sense strand and the antisense strand may have 5’ or 3’ overhangs of 1 to 5 nucleotides or 1 to 3 nucleotides. In some embodiments, the antisense strand comprises a 3’ overhang of two nucleotides.
[0054] Exemplary unmodified sense strand and antisense strand sequences of dsRNA targeting human ATXN2 mRNA are provided in Table 3 a.
[0055] In some embodiments, the sense strand comprises SEQ ID NO: 35, and the antisense strand comprises SEQ ID NO: 36. In some embodiments, the sense strand comprises SEQ ID NO: 37, and the antisense strand comprises SEQ ID NO: 38. In some embodiments, the sense strand comprises SEQ ID NO: 39, and the antisense strand comprises SEQ ID NO: 40. In some embodiments, the sense strand comprises SEQ ID NO: 41, and the antisense strand comprises SEQ ID NO: 36. In some embodiments, the sense strand comprises SEQ ID NO: 42, and the antisense strand comprises SEQ ID NO: 38. In some embodiments, the sense strand comprises SEQ ID NO: 43, and the antisense strand comprises SEQ ID NO: 40.Table 3a. Unmodified sequences of dsRNA targeting human ATXN2 mRNAdsRNA Sense Strand (5' to 3') SEQ Antisense Strand (5' to 3') SE Start No. ID Q position of NO ID antisense NO strand target region of human ATXN2 transcript NM_002973.4 (SEQ ID NO: 53)1 GAAUCUAUGGAUCA 35 UAGUAGUUGAUCCAUA 36 2240 ACUACUA GAUUCAG2 AAGAAUGAUUUUA 37 UUGUAACCUAAAAUCA 38 2204 GGUUACAA UUCUUAA3 AGGAUGGUUCAUAU 39 UGUAAGUAUAUGAACC 40 602 ACUUACA AUCCUCA4 UAAUCUAUGGAUCA 41 UAGUAGUUGAUCCAUA 36 2240 ACUACUA GAUUCAG5 UAGAAUGAUUUUA 42 UUGUAACCUAAAAUCA 38 2204 GGUUACAA UUCUUAA6 UGGAUGGUUCAUAU 43 UGUAAGUAUAUGAACC 40 602 ACUUACA AUCCUCATable 3b. Modified sequences of dsRNA targeting human ATXN2 mRNAdsR Sense Strand (5' to 3') SE Antisense Strand (5' to 3') SEQ Start position NA Q ID of antisense No. ID NO strand target NO region of human ATXN2 transcript NM_002973.4 (SEQ ID NO: 53)7 mG*mA*mAmUmCmUmAmU 44 mU*fA*mGmUfAmGmUf 45 2240 fGf Gf AmUm Cm Am AmCmU m UmGmAmUmCmCfAmUfAmC*mU*mA AmGm AmUmUmC *mA*mG8 mA*mA*mGmAmAmUmGmA 46 mU*fU*mGmUfAmAmCf 47 2204 fUfUfUmUmAmGmGmUmUm CmUmAmAmAmAfUmCfAmC*mA*mA AmUmUmCmUmU*mA*mA9 mA*mG*mGmAmUmGmGmU 48 mU*fG*mUmAfAmGmUf 49 602 fUfCfAmUmAmUmAmCmUm AmUm AmUm Gm Af AmCfUmA*mC*mA CmAmUmCmCmU*mC *mA10 [NH2mU] *mA*m AmUmCmU 50 mU*fA*mGmUfAmGmUf 45 2240 mAmUfGfGfAmUmCmAmAm UmGmAmUmCmCfAmUf CmUmAmC*mU*mA Am Gm AmUmUmC *mA*mG11 [NH2mU] *mA*mGmAmAmU 51 mU*fU*mGmUfAmAmCf 47 2204 mGmAfUfUfUmUmAmGmGm CmUmAmAmAmAfUmCf UmUmAmC*mA*mA AmUmUmCmUmU*mA*mA12 [NH2mU]*mG*mGmAmUmG 52 mU*fG*mUmAfAmGmUf 49 602 mGmUfUfCfAmUmAmUmAm AmUm AmUm Gm AfAmCf CmUmUmA*mC*mA CmAmUmCmCmU*mC*mAAbbreviations - ”m" indicates 2’-0Me; “f ’ indicates 2’-fluoro; indicates phosphorothioate linkage; unless otherwise noted, the 5’ position of the AS can include 5’-phosphate or 5’-vinylphosphonate (VP).
[0056] The dsRNA can include modifications. The modifications can be made to one or more nucleotides of the sense and / or antisense strand or to the intemucleotide linkages, which are the bonds between two nucleotides in the sense or antisense strand. For example, some 2’ -modifications of ribose or deoxyribose can increase RNA or DNA stability and halflife. Such 2’ -modifications can be 2’-fluoro, 2’-O-methyl (i.e., 2’-methoxy), or 2'-O-alkyl (e.g., 2’-O-Ci6 alkyl).
[0057] In some embodiments, one or more nucleotides of the sense strand and / or the antisense strand are independently modified nucleotides, which means the sense strand and the antisense strand can have different modified nucleotides. In some embodiments, each nucleotide of the sense strand is a modified nucleotide. In some embodiments, at least one nucleotide of the sense strand is an unmodified RNA nucleotide. In some embodiments, each nucleotide of the antisense strand is a modified nucleotide. In some embodiments, the modified nucleotide is a 2'-fluoro modified nucleotide, 2'-O-methyl modified nucleotide, 2’ deoxy nucleotide (DNA), or 2'-O-alkyl (e.g., 2’-O-Ci6 alkyl) modified nucleotide. In some embodiments, each nucleotide of the sense strand and the antisense strand is independently a modified nucleotide, e g., a 2'-fluoro modified nucleotide, 2'-O-methyl modified nucleotide,deoxy nucleotide (DNA), or 2'-O-alkyl (e.g., 2'-O-C16alkyl) modified nucleotide. In some embodiments, at least one nucleotide of the sense strand is 2’ deoxy nucleotide (DNA).
[0058] In some embodiments, the sense strand has four 2'-fluoro modified nucleotides, e.g., at positions 7, 9, 10, 11 from the 5’ end of the sense strand. In some embodiments, at least one nucleotide of the sense strand is an unmodified RNA nucleotide. In some embodiments, at least one nucleotide of the sense strand is 2‘ deoxy nucleotide (DNA). In some embodiments, the other nucleotides of the sense strand are 2'-O-methyl modified nucleotides.
[0059] In some embodiments, the antisense strand has four 2'-fluoro modified nucleotides, e.g., at positions 2, 6, 14, 16 from the 5’ end of the antisense strand. In some embodiments, the other nucleotides of the antisense strand are 2'-O-methyl modified nucleotides.
[0060] In some embodiments, the sense strand has three 2'-fluoro modified nucleotides, e.g., at positions 9, 10, 11 from the 5‘ end of the sense strand. In some embodiments, at least one nucleotide of the sense strand is an unmodified RNA nucleotide. In some embodiments, at least one nucleotide of the sense strand is 2’ deoxy nucleotide (DNA). In some embodiments, the other nucleotides of the sense strand are 2'-O-methyl modified nucleotides.
[0061] In some embodiments, the antisense strand has five 2'-fluoro modified nucleotides, e.g.. at positions 2, 5, 7. 14. 16 from the 5’ end of the antisense strand. In some embodiments, the antisense strand has five 2'-fluoro modified nucleotides, e.g., at positions 2, 5, 8, 14, 16 from the 5’ end of the antisense strand. In some embodiments, the antisense strand has five 2'-fluoro modified nucleotides, e.g., at positions 2, 3, 7, 14, 16 from the 5’ end of the antisense strand. In some embodiments, the antisense strand has three 2'-fluoro modified nucleotides, e.g., at positions 2, 14, 16 from the 5’ end of the antisense strand. In some embodiments, the other nucleotides of the antisense strand are 2'-O-methyl modified nucleotides.
[0062] In some embodiments, the 5 ' end of the antisense strand has a phosphate analog, e.g., 5’-vinylphosphonate (5 -VP).
[0063] In some embodiments, the sense strand or the antisense strand comprises an abasic moiety or inverted abasic moiety, e.g., a moiety shown in Table 4.Table 4. Abasic or inverted abasic (iAb) moieties“5”’ and '‘3’” indicate the 5’ to 3’ direction of the sequences.
[0064] In some embodiments, the sense strand and the antisense strand have one or more modified intemucleotide linkages. In some embodiments, the modified intemucleotide linkage is phosphorothioate linkage. In some embodiments, the sense strand has four or five phosphorothioate linkages. In some embodiments, the antisense strand has four or five phosphorothioate linkages. In some embodiments, the sense strand and the antisense strand each has four or five phosphorothioate linkages. In some embodiments, the sense strand has four phosphorothioate linkages and the antisense strand has five phosphorothioate linkages.
[0065] Exemplary modified sense strand and antisense strand sequences of dsRNA targeting human ATXN2 mRNA are provided in Table 3b.
[0066] In some embodiments, the sense strand comprises SEQ ID NO: 44, and the antisense strand comprises SEQ ID NO: 45. In some embodiments, the sense strand comprises SEQ ID NO: 46, and the antisense strand comprises SEQ ID NO: 47. In some embodiments, the sense strand comprises SEQ ID NO: 48, and the antisense strand comprises SEQ ID NO: 49. In some embodiments, the sense strand comprises SEQ ID NO: 50, and the antisense strand comprises SEQ ID NO: 45. In some embodiments, the sense strand comprises SEQ ID NO: 51, and the antisense strand comprises SEQ ID NO: 47. In some embodiments, the sense strand comprises SEQ ID NO: 52, and the antisense strand comprises SEQ ID NO: 49.
[0067] In some embodiments, the sense strand consists of SEQ ID NO: 44, and the antisense strand consists of SEQ ID NO: 45. In some embodiments, the sense strand consists of SEQ ID NO: 46, and the antisense strand consists of SEQ ID NO: 47. In some embodiments, the sense strand consists of SEQ ID NO: 48, and the antisense strand consistsof SEQ ID NO: 49. In some embodiments, the sense strand comprises SEQ ID NO: 50, and the antisense strand comprises SEQ ID NO: 45. In some embodiments, the sense strand comprises SEQ ID NO: 51, and the antisense strand comprises SEQ ID NO: 47. In some embodiments, the sense strand comprises SEQ ID NO: 52, and the antisense strand comprises SEQ ID NO: 49.
[0068] In some embodiments, the dsRNA comprises a sense strand that comprises a sequence that has 1, 2, or 3 differences from a sense strand sequence in Table 3a or 3b. In some embodiments, the dsRNA comprises an antisense strand that comprises a sequence that has 1, 2, or 3 differences from an antisense strand sequence in Table 3a or 3b.
[0069] The sense strand and antisense strand of dsRNA can be synthesized using any nucleic acid polymerization methods known in the art, for example, solid-phase synthesis by employing phosphoramidite chemistry methodology (e.g.. Current Protocols in Nucleic Acid Chemistry, Beaucage, S. L. et al. (Edrs.), John Wiley & Sons, Inc., New York, NY, USA), H-phosphonate, phosphortriester chemistry', or enzymatic synthesis. Automated commercial synthesizers can be used, for example. MerMade™ 12 from LGC Biosearch Technologies, or other synthesizers from BioAutomation or Applied Biosystems. Phosphorothioate linkages can be introduced using a sulfurizing reagent such as phenylacetyl disulfide or DDTT (((dimethylaminomethylidene) amino)-3H-l,2,4-dithiazaoline-3-thione). It is well known to use similar techniques and commercially available modified amidites and controlled-pore glass (CPG) products to synthesize modified oligonucleotides or conjugated oligonucleotides.
[0070] Purification methods can be used to exclude the unwanted impurities from the final oligonucleotide product. Commonly used purification techniques for single stranded oligonucleotides include reverse-phase ion pair high performance liquid chromatography (RP-IP-HPLC), capillary gel electrophoresis (CGE), anion exchange HPLC (AX-HPLC), and size exclusion chromatography (SEC). After purification, oligonucleotides can be analyzed by mass spectrometry and quantified by spectrophotometry at a wavelength of 260 nm. The sense strand and antisense strand can then be annealed to form a dsRNA.
[0071] The RNAi agent described herein can be made by a variety of procedures known to one of ordinary skill in the art, some of which are illustrated in the preparations and examples below, e.g., in Examples 1-3. One of ordinary skill in the art recognizes that the specific synthetic steps for each of the routes described may be combined in different ways, or in conjunction with steps from different schemes, to prepare the RNAi agent. The product of each step can be recovered by conventional methods well known in the art, including extraction, evaporation, precipitation, chromatography, filtration, trituration, andcrystallization. The reagents and starting materials are readily available to one of ordinary skill in the art.
[0072] In some embodiments, the TfR binding protein with native or engineered cysteines described herein can be first treated with a reducing agent, e.g., DTT, and then reoxidized with an oxidizing agent, e.g., DHAA. The resulting oxidized TfR binding protein is then incubated with a linker functionalized dsRNA, e.g., linker-dsRNA, to produce the conjugated RNAi agent.Linker
[0073] In some embodiments, the ATXN2 RNAi agents described herein comprises a linker that links the human TfR binding protein to the dsRNA. Exemplary linker structures are shown in Table 5. In some embodiments, the linker comprises a SMCC linker, OD linker, or MSPT tinker. In some embodiments, the tinker is a SMCC tinker, OD tinker, or MSPT tinker. In some embodiments, the linker is a SMCC linker. In some embodiments, the linker is a MSPT tinker.
[0074] In some embodiments, the linker is conjugated to an engineered cysteine in the human TfR binding protein. In some embodiments, the linker is conjugated to the sense strand of the dsRNA, e.g., the 5’ end or 3’ end of the sense strand.Table 5. Exemplary linker structuresLinker StructureSMCC linker 1*O1 / H )-TBP sense strand P^^. 'HO0T oO SMCC linker 2O2sense strand J / >TBPT0OHydrolyzed ring open form of SMCC linker 1 *O3? H OO / -TBP sense strand P^ > / HOHydrolyzed ring open form of SMCC linker 20TBPsense strand / I HOMal-Tet-TCO linker 1*HN •NJU 10\ J 0.1 nTBPnX3 % M sense st.rand.-O' Fpt.0 z^Z— n U k i z iNxn / H0O / X ) O > ( N0Mal-Tet-TCO linker2ZTx / O\HN -NJU 10\ 7 k' OU™ sense strand 9« k ^NV01 I VOx J0oGDM linker 1*O 0HOoII IIssense strand ' / \zz\^^ / \z / x'TBPO *xH H / \GDM linker20 0sense strand / k ^x k- N ^x X xS, TBPH / \3’-OD linker*Osense strand-0'N'N 3’-MSPT linker*O10HOXTBP\ J sense strand-0xv N v ikNlN^N 5‘-OD linkerO11r^\f'^'x / ^'O / Xxx'(^'xxX^x'sense strandTBP-^ YN'N5’-MSPT linkerO12TBP r??Y^<^x^xO'x^ / <^x^xsense strandN^N*Note - X is O or SPharmaceutical Composition
[0075] In another aspect, provided herein are pharmaceutical compositions comprising any of the ATXN2 RNAi agents described herein and a pharmaceutically acceptable carrier. Such pharmaceutical compositions can also comprise one or more pharmaceutically acceptable excipient, diluent, or carrier. Pharmaceutical compositions can be prepared by methods well known in the art (e.g., Remington: The Science and Practice of Pharmacy, 23rd edition (2020), A. Loyd et al., Academic Press).Method of Treatment and Therapeutic Use
[0076] In another aspect, provided herein are methods of treating an ATXN2-associated neurological disease in a patient in need thereof, and such method comprises administering to the patient an effective amount of the ATXN2 RNAi agent or a pharmaceutical composition described herein. Abnormal C AG trinucleotide expansion in ATXN2 gene is associated with spinocerebellar ataxia type 2 (SC A2) and amyotrophic lateralsclerosis (ALS). In another aspect. Ataxin-2 is known to be required for the toxicity mediated by abnormal aggregation of TAR DNA binding protein (TDP-43), which is believed to be the driver in 95% of sporadic and familial amyotrophic lateral sclerosis cases. Exemplary ATXN2-associated neurological diseases, includes, but are not limited to, spinocerebellar ataxia type 2 (SCA2), amyotrophic lateral sclerosis (ALS), primary lateral sclerosis (PLS), Parkinson's disease, Alzheimer's disease, frontotemporal lobar degeneration (FTLD), progressive muscular atrophy (PMA), multiple system proteinopathy, Perry disease, and TDP-43 proteinopathy, e.g., neurological disease associated with abnormal TDP-43 aggregation (de Boer et al, 2021, J Neurol. Neurosurg. Psychiatry'. 2020 Nov 11;92(1 ): 86-95).
[0077] In another aspect, provided herein are methods of reducing ATXN2 expression in a patient in need thereof, and such method comprises administering to the patient an effective amount of an ATXN2 RNAi agent or a pharmaceutical composition described herein.
[0078] The ATXN2 RNAi agent or a pharmaceutical composition comprising ATXN2 RNAi agent can be administered to the patient intrathecally, intravenously or subcutaneously.
[0079] ATXN2 RNAi agent dosage regimens may be adjusted to provide the optimum desired response (e.g., a therapeutic response). For example, a single bolus may be administered, several divided doses may be administered over time, or the dose may be proportionally reduced or increased as indicated by the exigencies of the therapeutic situation.
[0080] Dosage values may vary with the type and severity of the condition to be alleviated. It is further understood that for any particular subject, specific dosage regimens should be adjusted over time according to the individual need and the professional judgment of the person administering or supervising the administration of the compositions.
[0081] In another aspect, provided herein are ATXN2 RNAi agents or pharmaceutical compositions comprising a ATXN2 RNAi agent for use in reducing ATXN2 expression. Also provided herein are ATXN2 RNAi agents or the pharmaceutical composition comprising a ATXN2 RNAi agent for use in a therapy. Also provided herein are ATXN2 RNAi agents or pharmaceutical compositions comprising a ATXN2 RNAi agent for use in the treatment of a ATXN2-associated neurological disease. Also provided herein are uses of ATXN2 RNAi agents in the manufacture of a medicament for the treatment of a ATXN2-associated neurological disease.Definitions
[0082] As used herein, the terms 'a." “an,” “the,” and similar terms used in the context of the present disclosure (especially in the context of the claims) are to be construed to cover both the singular and plural unless otherwise indicated herein or clearly contradicted by the context.
[0083] As used herein, the term "‘alkyl” means saturated linear or branched-chain monovalent hydrocarbon radical, containing the indicated number of carbon atoms. For example, “C1-C20 alkyl” means a radical having 1-20 carbon atoms in a linear or branched arrangement.
[0084] The term “antibody,” as used herein, refers to a molecule that binds an antigen. Embodiments of an antibody include a monoclonal antibody, polyclonal antibody, human antibody, humanized antibody, chimeric antibody, heterodimeric antibody, bispecific or multispecific antibody, or conjugated antibody. The antibodies can be of any class (e.g., IgG, IgE, IgM, IgD, IgA), and any subclass (e.g., IgGl, IgG2, IgG3, IgG4).
[0085] An immunoglobulin G (IgG) type antibody comprised of four polypeptide chains: two heavy chains (HC) and two light chains (LC) that are cross-linked via inter-chain disulfide bonds. The amino-terminal portion of each of the four polypeptide chains includes a variable region of about 100-125 or more amino acids primarily responsible for antigen recognition. The carboxyl-terminal portion of each of the four polypeptide chains contains a constant region primarily responsible for effector function. Each heavy chain is comprised of a heavy chain variable region (VH) and a heavy chain constant region. Each light chain is comprised of a light chain variable region (VL) and a light chain constant region. The IgG isotype may be further divided into subclasses (e.g., IgGl, IgG2, IgG3, and IgG4).
[0086] The VH and VL regions can be further subdivided into regions of hypervariability, termed complementarity determining regions (CDRs), interspersed with regions that are more conserved, termed framework regions (FR). The CDRs are exposed on the surface of the protein and are important regions of the antibody for antigen binding specificity. Each VH and VL is composed of three CDRs and four FRs, arranged from amino-terminus to carboxyl-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. Herein, the three CDRs of the heavy chain are referred to as “HCDR1, HCDR2, and HCDR3” and the three CDRs of the light chain are referred to as “LCDR1, LCDR2 and LCDR3”. The CDRs contain most of the residues that form specific interactions with the antigen. Assignment of amino acid residues to the CDRs may be done according to the well-known schemes, including those described in Kabat (Kabat et al., “Sequences of Proteins ofImmunological Interest,'’ National Institutes of Health, Bethesda, Md. (1991)), Chothia (Chothia et al., “Canonical structures for the hypervariable regions of immunoglobulins”, Journal of Molecular Biology, 196, 901-917 (1987); Al-Lazikani et al., “Standard conformations for the canonical structures of immunoglobulins”, Journal of Molecular Biology', 273, 927-948 (1997)), North (North et al., “A New Clustering of Antibody CDR Loop Conformations”, Journal of Molecular Biology, 406, 228-256 (2011)), or IMGT (the international ImMunoGeneTics database available on at www.imgt.org; see Lefranc et al.. Nucleic Acids Res. 1999; 27:209-212).
[0087] Embodiments of the present disclosure also include antibody fragments or antigen-binding fragments that, as used herein, comprise at least a portion of an antibody retaining the ability to specifically interact with an antigen or an epitope of the antigen, such as Fab, Fab’, F(ab’)2, Fv fragments, scFv antibody fragments, scFab, disulfide-linked Fvs (sdFv), a Fd fragment.
[0088] The term “antigen binding domain”, as used herein, refers to a portion of an antibody or antibody fragment that binds an antigen or an epitope of the antigen. For example, “TfR binding domain” refers to a portion of an antibody or antibody fragment that binds TfR or an epitope of TfR.
[0089] The term “heterodimeric antibody”, as used herein, refers to an antibody that comprises two distinct antigen-binding domains.
[0090] As used herein, “antisense strand” means a single-stranded oligonucleotide that is complementary to a region of a target sequence. Likewise, and as used herein, “sense strand” means a single-stranded oligonucleotide that is complementary to a region of an antisense strand.
[0091] The terms “bind” and “binds” as used herein are intended to mean, unless indicated otherwise, the ability of a protein or molecule to form a chemical bond or attractive interaction with another protein or molecule, which results in proxi mi ty of the two proteins or molecules as determined by common methods known in the art.
[0092] As used herein, “complementary” means a structural relationship between two nucleotides (e.g., on two opposing nucleic acids or on opposing regions of a single nucleic acid strand, e.g., a hairpin) that permits the two nucleotides to form base pairs with one another. For example, a purine nucleotide of one nucleic acid that is complementary to a pyrimidine nucleotide of an opposing nucleic acid may base pair together by forming hydrogen bonds with one another. Complementary nucleotides can base pair in the Watson-Crick manner or in any other manner that allows for the formation of stable duplexes.Likewise, two nucleic acids may have regions of multiple nucleotides that are complementary with each other to form regions of complementarity, as described herein.
[0093] As used herein, “duplex,’’ in reference to nucleic acids or oligonucleotides, means a structure formed through complementary base pairing of two antiparallel sequences of nucleotides (i.e., in opposite directions), whether formed by two separate nucleic acid strands or by a single, folded strand (e.g., via a hairpin).
[0094] An “effective amount” refers to an amount necessary (for periods of time and for the means of administration) to achieve the desired therapeutic result. An effective amount of a protein or conjugate may vary according to factors such as the disease state, age, sex, and weight of the individual, and the ability of the protein or conjugate to elicit a desired response in the individual. An effective amount is also one in which any toxic or detrimental effects of the protein or conjugate are outweighed by the therapeutically beneficial effects.
[0095] The term “Fc region” as used herein refers to a polypeptide comprising the CH2 and CH3 domains of a constant region of an immunoglobulin, e.g., IgGl, IgG2, IgG3, or IgG4. Optionally, the Fc region may include a portion of the hinge region or the entire hinge region of an immunoglobulin, e.g., IgGl, IgG2, IgG3, or IgG4. In some embodiments, the Fc region is a human IgG Fc region, e.g., a human IgGl Fc region, human IgG2 Fc region, human IgG3 Fc region or human IgG4 Fc region. In some embodiments, the Fc region is a modified IgG Fc region with reduced or eliminated effector functions compared to the corresponding wild type IgG Fc region. The numbering of the residues in the Fc region is based on the EU index as described in Kabat (Kabat et al, Sequences of Proteins of Immunological Interest, 5th edition, Bethesda, MD: U. S. Dept, of Health and Human Services, Public Health Service, National Institutes of Health, 1991). The boundaries of the Fc region of an immunoglobulin heavy chain might vary, and the human IgG heavy chain Fc region is usually defined as the stretch from the N-terminus of the CH2 domain (e.g., the amino acid residue at position 231 according to the EU index numbering) to the C-terminus of the CH3 domain (or the C-terminus of the immunoglobulin).
[0096] The term “knockdown” or “expression knockdown” refers to reduced mRNA or protein expression of a gene after treatment of a reagent.
[0097] As used herein, “modified intemucleotide linkage” means an intemucleotide linkage having one or more chemical modifications when compared with a reference intemucleotide linkage having a phosphodi ester bond. A modified intemucleotide linkage can be a non-naturally occurring linkage. In some embodiments, the modified intemucleotide linkage is phosphorothioate linkage.
[0098] As used herein, “modified nucleotide” refers to a nucleotide having one or more chemical modifications when compared with a corresponding reference nucleotide selected from: adenine ribonucleotide, guanine ribonucleotide, cytosine ribonucleotide, uracil ribonucleotide, adenine deoxyribonucleotide, guanine deoxyribonucleotide, cytosine deoxyribonucleotide, and thymidine deoxyribonucleotide. A modified nucleotide can have, for example, one or more chemical modification in its sugar, nucleobase, and / or phosphate group. Additionally, or alternatively, a modified nucleotide can have one or more chemical moieties conjugated to a corresponding reference nucleotide. In some embodiments, the modified nucleotide is a 2'-fluoro modified nucleotide, 2'-O-methyl modified nucleotide, 2’ deoxy nucleotide (DNA), or 2'-O-alkyl (e.g., 2’-O-Ci6 alkyl) modified nucleotide. In some embodiments, the modified nucleotide has a phosphate analog, e.g., 5’-vinylphosphonate. In some embodiments, the modified nucleotide has an abasic moiety or inverted abasic moiety, e.g., a moiety shown in Table 4.
[0099] As used herein, “nucleotide” means an organic compound having a nucleoside (a nucleobase, e.g.. adenine, cytosine, guanine, thymine, or uracil, and a pentose sugar, e.g., ribose or 2'-deoxyribose) linked to a phosphate group. A “nucleotide” can serve as a monomeric unit of nucleic acid polymers such as deoxyribonucleic acid (DNA) and ribonucleic acid (RNA).[000100] As used herein, a “null arm” means an antibody arm that does not bind any known human target.[000101] As used herein, “oligonucleotide” means a polymer of linked nucleotides, each of which can be modified or unmodified. An oligonucleotide is ty pically less than about 100 nucleotides in length.[000102] As used herein, “overhang” means the unpaired nucleotide or nucleotides that protrude from the duplex structure of a double stranded oligonucleotide. An overhang may include one or more unpaired nucleotides extending from a duplex region at the 5’ terminus or 3’ terminus of a double stranded oligonucleotide. The overhang can be a 3’ or 5’ overhang on the antisense strand or sense strand of a double stranded oligonucleotide.[000103] The term “patient”, as used herein, refers to a human patient.[000104] As used herein, “phosphate analog” means a chemical moiety that mimics the electrostatic and / or steric properties of a phosphate group. In some embodiments, a phosphate analog is positioned at the 5' end of an oligonucleotide in place of a 5 ’-phosphate, which is sometimes susceptible to enzymatic removal. A 5' phosphate analog can include a phosphatase-resistant linkage. Examples of phosphate analogs include 5’ methylenephosphonate (5 ’-MP) and 5’-(E)-vinylphosphonate (5 ’-VP). In some embodiments, the phosphate analog is 5 ’-VP.[000105] As used herein, “ATXN2” (also known as ATX2, SCA2, TNRC13) refers to a human ATXN2 mRNA transcript. The nucleotide sequence of human ATXN2 mRNA isoform 1 can be found atNM_002973.4:i AGAGCTCGCC TCCCTCCGCC TCAGACTGTT TTGGTAGCAA CGGCAACGGC GGCGGCGCGT 61 TTCGGCCCGG CTCCCGGCGG CTCCTTGGTC TCGGCGGGCC TCCCCGCCCC TTCGTCGTCC 121 TCCTTCTCCC CCTCGCCAGC CCGGGCGCCC CTCCGGCCGC GCCAACCCGC GCCTCCCCGC 181 TCGGCGCCCG CGCGTCCCCG CCGCGTTCCG GCGTCTCCTT GGCGCGCCCG GCTCCCGGCT 241 GTCCCCGCCC GGCGTGCGAG CCGGTGTATG GGCCCCTCAC CATGTCGCTG AAGCCCCAGC 301 AGCAGCAGCA GCAGCAGCAG CAGCAGCAGC AGCAGCAACA GCAGCAGCAG CAGCAGCAGC 361 AGCAGCCGCC GCCCGCGGCT GCCAATGTCC GCAAGCCCGG CGGCAGCGGC CTTCTAGCGT 421 CGCCCGCCGC CGCGCCTTCG CCGTCCTCGT CCTCGGTCTC CTCGTCCTCG GCCACGGCTC 481 CCTCCTCGGT GGTCGCGGCG ACCTCCGGCG GCGGGAGGCC CGGCCTGGGC AGAGGTCGAA 541 ACAGTAACAA AGGACTGCCT CAGTCTACGA TTTCTTTTGA TGGAATCTAT GCAAATATGA 601 GGATGGTTCA TATACTTACA TCAGTTGTTG GCTCCAAATG TGAAGTACAA GTGAAAAATG 661 GAGGTATATA TGAAGGAGTT TTTAAAACTT ACAGTCCGAA GTGTGATTTG GTACTTGATG 721 CCGCACATGA GAAAAGTACA GAATCCAGTT CGGGGCCGAA ACGTGAAGAA ATAATGGAGA 781 GTATTTTGTT CAAATGTTCA GACTTTGTTG TGGTACAGTT TAAAGATATG GACTCCAGTT 841 ATGCAAAAAG AGATGCTTTT ACTGACTCTG CTATCAGTGC TAAAGTGAAT GGCGAACACA 901 AAGAGAAGGA CCTGGAGCCC TGGGATGCAG GTGAACTCAC AGCCAATGAG GAACTTGAGG 961 CTTTGGAAAA TGACGTATCT AATGGATGGG ATCCCAATGA TATGTTTCGA TATAATGAAG 1021 AAAATTATGG TGTAGTGTCT ACGTATGATA GCAGTTTATC TTCGTATACA GTGCCCTTAG 1081 AAAGAGATAA CTCAGAAGAA TTTTTAAAAC GGGAAGCAAG GGCAAACCAG TTAGCAGAAG 1141 AAATTGAGTC AAGTGCCCAG TACAAAGCTC GAGTGGCCCT GGAAAATGAT GATAGGAGTG 1201 AGGAAGAAAA AT C CAGCA GTTCAGAGAA ATTCCAGTGA ACGTGAGGGG CACAGCATAA 1261 ACACTAGGGA AAATAAATAT ATTCCTCCTG GACAAAGAAA TAGAGAAGTC ATATCCTGGG 1321 GAAGTGGGAG ACAGAATTCA CCGCGTATGG GCCAGCCTGG ATCGGGCTCC ATGCCATCAA 1381 GATCCACTTC TCACACTTCA GATTTCAACC CGAATTCTGG TTCAGACCAA AGAGTAGTTA 1441 ATGGAGGTGT TCCCTGGCCA TCGCCTTGCC CATCTCCTTC CTCTCGCCCA CCTTCTCGCT 1501 ACCAGTCAGG TCCCAACTCT CTTCCACCTC GGGCAGCCAC CCCTACACGG CCGCCCTCCA 1561 GGCCCCCCTC GCGGCCATCC AGACCCCCGT CTCACCCCTC TGCTCATGGT TCTCCAGCTC 1621 CTGTCTCTAC TATGCCTAAA CGCATGTCTT CAGAAGGGCC TCCAAGGATG TCCCCAAAGG 1681 CCCAGCGACA TCCTCGAAAT CACAGAGTTT CTGCTGGGAG GGGTTCCATA TCCAGTGGCC 1741 TAGAATTTGT ATCCCACAAC CCACCCAGTG AAGCAGCTAC TCCTCCAGTA GCAAGGACCA 1801 GTCCCTCGGG GGGAACGTGG TCATCAGTGG TCAGTGGGGT TCCAAGATTA TCCCCTAAAA 1861 CTCATAGACC CAGGTCTCCC AGACAGAACA GTATTGGAAA TACCCCCAGT GGGCCAGTTC 1921 TTGCTTCTCC CCAAGCTGGT ATTATTCCAA CTGAAGCTGT TGCCATGCCT ATTCCAGCTG 1981 CATCTCCTAC GCCTGCTAGT CCTGCATCGA ACAGAGCTGT TACCCCTTCT AGTGAGGCTA 2041 AAGATTCCAG GCTTCAAGAT CAGAGGCAGA ACTCTCCTGC AGGGAATAAA GAAAATATTA 2101 AACCCAATGA AACATCACCT AGCTTCTCAA AAGCTGAAAA CAAAGGTATA TCACCAGTTG 2161 TTTCTGAACA TAGAAAACAG ATTGATGATT TAAAGAAATT TAAGAATGAT TTTAGGTTAC 2221 AGCCAAGTTC TACTTCTGAA TCTATGGATC AACTACTAAA CAAAAATAGA GAGGGAGAAA 2281 AATCAAGAGA TTTGATCAAA GACAAAATTG AACCAAGTGC TAAGGATTCT TTCATTGAAA 2341 ATAGCAGCAG CAACTGTACC AGTGGCAGCA GCAAGCCGAA TAGCCCCAGC ATTTCCCCTT 2401 CAATACTTAG TAACACGGAG CACAAGAGGG GACCTGAGGT CACTTCCCAA GGGGTTCAGA 2461 CTTCCAGCCC AGCATGTAAA CAAGAGAAAG ACGATAAGGA AGAGAAGAAA GACGCAGCTG 2521 AGCAAGTTAG GAAATCAACA TTGAATCCCA ATGCAAAGGA GTTCAACCCA CGTTCCTTCT 2581 CTCAGCCAAA GCCTTCTACT ACCCCAACTT CACCTCGGCC TCAAGCACAA CCTAGCCCAT 2641 CTATGGTGGG TCATCAACAG CCAACTCCAG TTTATACTCA GCCTGTTTGT TTTGCACCAA 2701 ATATGATGTA TCCAGTCCCA GTGAGCCCAG GCGTGCAACC TTTATACCCA ATACCTATGA 2761 CGCCCATGCC AGTGAATCAA GCCAAGACAT ATAGAGCAGT ACCAAATATG CCCCAACAGC 2821 GGCAAGACCA GCATCATCAG AGTGCCATGA TGCACCCAGC GTCAGCAGCG GGCCCACCGA 2881 TTGCAGCCAC CCCACCAGCT TACTCCACGC AATATGTTGC CTACAGTCCT CAGCAGTTCC 2941 CAAATCAGCC CCTTGTTCAG CATGTGCCAC ATTATCAGTC TCAGCATCCT CATGTCTATA 3001 GTCCTGTAAT ACAGGGTAAT GCTAGAATGA TGGCACCACC AACACACGCC CAGCCTGGTT 3061 TAGTATCTTC TTCAGCAACT CAGTACGGGG CTCATGAGCA GACGCATGCG ATGTATGCAT3121 GTCCCAAATT ACCATACAAC AAGGAGACAA GCCCTTCTTT CTACTTTGCC ATTTCCACGG 3181 GCTCCCTTGC TCAGCAGTAT GCGCACCCTA ACGCTACCCT GCACCCACAT ACTCCACACC 3241 CTCAGCCTTC AGCTACCCCC ACTGGACAGC AGCAAAGCCA ACATGGTGGA AGTCATCCTG 3301 CACCCAGTCC TGTTCAGCAC CATCAGCACC AGGCCGCCCA GGCTCTCCAT CTGGCCAGTC 3361 CACAGCAGCA GTCAGCCATT TACCACGCGG GGCTTGCGCC AACTCCACCC TCCATGACAC 3421 CTGCCTCCAA CACGCAGTCG CCACAGAATA GTTTCCCAGC AGCACAACAG ACTGTCTTTA 3481 CGATCCATCC TTCTCACGTT CAGCCGGCGT ATACCAACCC ACCCCACATG GCCCACGTAC 3541 CTCAGGCTCA TGTACAGTCA GGAATGGTTC CTTCTCATCC AACTGCCCAT GCGCCAATGA 3601 TGCTAATGAC GACACAGCCA CCCGGCGGTC CCCAGGCCGC CCTCGCTCAA AGTGCACTAC 3661 AGCCCATTCC AGTCTCGACA ACAGCGCATT TCCCCTATAT GACGCACCCT TCAGTACAAG 3721 CCCACCACCA ACAGCAGTTG TAAGGCTGCC CTGGAGGAAC CGAAAGGCCA AATTCCCTCC 3781 TCCCTTCTAC TGCTTCTACC AACTGGAAGC ACAGAAAACT AGAATTTCAT TTATTTTGTT 3841 TTTAAAATAT ATATGTTGAT TTCTTGTAAC ATCCAATAGG AATGCTAACA GTTCACTTGC 3901 AGTGGAAGAT ACTTGGACCG AGTAGAGGCA TTTAGGAACT TGGGGGCTAT TCCATAATTC 3961 CATATGCTGT TTCAGAGTCC CGCAGGTACC CCAGCTCTGC TTGCCGAAAC TGGAAGTTAT 4021 TTATTTTTTA ATAACCCTTG AAAGTCATGA ACACATCAGC TAGCAAAAGA AGTAACAAGA 4081 GTGATTCTTG CTGCTATTAC TGCTAAAAAA AAAAAAAAAA AAAAATCAAG ACTTGGAACG 4141 CCCTTTTACT AAACTTGACA AAGTTTCAGT AAATTCTTAC CGTCAAACTG ACGGATTATT 4201 ATTTATAAAT CAAGTTTGAT GAGGTGATCA CTGTCTACAG TGGTTCAACT TTTAAGTTAA 4261 GGGAAAAACT TTTACTTTGT AGATAAT TA AAATAAAAAC TTAAAAAAAA TTTAAAAAAT 4321 AAAAAAAGTT TTAAAAACTG A (SEQ ID NO: 53).[000106] The corresponding amino acid sequence of human ATXN2 protein isoform 1 can be found at NP_002964.4:1 MSLKPQQQQQ QQQQQQQQQQ QQQQQQQQPP PAAANVRKPG GSGLLASPAA APSPSSSSVS 61 SSSATAPSSV VAATSGGGRP GLGRGRNSNK GLPQSTISFD GIYANMRMVH ILTSWGSKC 121 EVQVKNGGIY EGVFKTYSPK CDLVLDAAHE KSTESSSGPK REEIMESILF KCSDFVWQF 181 KDMDSSYAKR DAFTDSAISA KVNGEHKEKD LEPWDAGELT ANEELEALEN DVSNGWDPND 241 MFRYNEENYG WSTYDSSLS SYTVPLERDN SEEFLKREAR ANQLAEEIES SAQYKARVAL 301 ENDDRSEEEK YTAVQRNSSE REGHSINTRE NKYIPPGQRN REVISWGSGR QNSPRMGQPG 361 SGSMPSRSTS HTSDFNPNSG SDQRVVNGGV PWPSPCPSPS SRPPSRYQSG PNSLPPRAAT 421 PTRPPSRPPS RPSRPPSHPS AHGSPAPVST MPKRMSSEGP PRMSPKAQRH PRNHRVSAGR 481 GSISSGLEFV SHNPPSEAAT PPVARTSPSG GTWSSWSGV PRLSPKTHRP RSPRQNSIGN 541 TPSGPVLASP QAGI IPTEAV AMPIPAASPT PASPASNRAV TPSSEAKDSR LQDQRQNSPA 601 GNKENIKPNE TSPSFSKAEN KGISPWSEH RKQIDDLKKF KNDFRLQPSS TSESMDQLLN 661 KNREGEKSRD LIKDKIEPSA KDSFIENSSS NCTSGSSKPN SPSISPSILS NTEHKRGPEV 721 TSQGVQTSSP ACKQEKDDKE EKKDAAEQVR KSTLNPNAKE FNPRSFSQPK PSTTPTSPRP 781 QAQPSPSMVG HQQPTPVYTQ PVCFAPNMMY PVPVSPGVQP LYPIPMTPMP VNQAKTYRAV 841 PNMPQQRQDQ HHQSAMMHPA SAAGPPIAAT PPAYSTQYVA YSPQQFPNQP LVQHVPHYQS 901 QHPHVYSPVI QGNARMMAPP THAQPGLVSS SATQYGAHEQ THAMYACPKL PYNKETSPSF 961 YFAISTGSLA QQYAHPNATL HPHTPHPQPS ATPTGQQQSQ HGGSHPAPSP VQHHQHQAAQ 1021 ALHLASPQQQ SAIYHAGLAP TPPSMTPASN TQSPQNSFPA AQQTVFTIHP SHVQPAYTNP 1081 PHMAHVPQAH VQSGMVPSHP TAHAPMMLMT TQPPGGPQAA LAQSALQPIP VSTTAHFPYM 1141 THPSVQAHHQ QQL (SEQ ID NO: 54).[000107] The human ATXN2 isoform 2 mRNA sequence can be found at NM_001310121.1 (SEQ ID NO: 55); and the corresponding protein sequence can be found at NP_001297050.1 (SEQ ID NO: 56).1 CCCGAGAAAG CAACCCAGCG CGCCGCCCGC TCCTCACGTG TCCCTCCCGG CCCCGGGGCC 61 ACCTCACGTT CTGCTTCCGT CTGACCCCTC CGACTTCCGA GGTCGAAACA GTAACAAAGG 121 ACTGCCTCAG TCTACGATTT CTTTTGATGG AATCTATGCA AATATGAGGA TGGTTCATAT 181 ACTTACATCA GTTGTTGGCT CCAAATGTGA AGTACAAGTG AAAAATGGAGGTATATATGA241 AGGAGTTTTT AAAACTTACA GTCCGAAGTG TGATTTGGTA CTTGATGCCG CACATGAGAA301 AAGTACAGAA TCCAGTTCGG GGCCGAAACG TGAAGAAATA ATGGAGAGTA TTTTGTTCAA361 ATGTTCAGAC TTTGTTGTGG TACAGTTTAA AGATATGGAC TCCAGTTATG CAAAAAGAGA421 TGCTTTTACT GACTCTGCTA TCAGTGCTAA AGTGAATGGC GAACACAAAG AGAAGGACCT481 GGAGCCCTGG GATGCAGGTG AACTCACAGC CAATGAGGAA CTTGAGGCTT TGGAAAATGA541 CGTATCTAAT GGATGGGATC CCAATGATAT GTTTCGATAT AATGAAGAAA ATTATGGTGT601 AGTGTCTACG TATGATAGCA GTTTATCTTC GTATACAGTG CCCTTAGAAA GAGATAACTC 661 AGAAGAATTT TTAAAACGGG AAGCAAGGGC AAACCAGTTA GCAGAAGAAA TTGAGTCAAG721 TGCCCAGTAC AAAGCTCGAG TGGCCCTGGA AAATGATGAT AGGAGTGAGG AAGAAAAATA781 CACAGCAGTT CAGAGAAATT CCAGTGAACG TGAGGGGCAC AGCATAAACA CTAGGGAAAA841 TAAATATATT CCTCCTGGAC AAAGAAATAG AGAAGTCATA TCCTGGGGAA GTGGGAGACA901 GAATTCACCG CGTATGGGCC AGCCTGGATC GGGCTCCATG CCATCAAGAT CCACTTCTCA961 CACTTCAGAT TTCAACCCGA ATTCTGGTTC AGACCAAAGA GTAGTTAATG GAGGTGTTCC1021 CTGGCCATCG CCTTGCCCAT CTCCTTCCTC TCGCCCACCT TCTCGCTACC AGTCAGGTCC 1081 CAACTCTCTT CCACCTCGGG CAGCCACCCC TACACGGCCG CCCTCCAGGC CCCCCTCGCG1141 GCCATCCAGA CCCCCGTCTC ACCCCTCTGC TCATGGTTCT CCAGCTCCTG TCTCTACTAT 1201 GCCTAAACGC ATGTCTTCAG AAGGGCCTCC AAGGATGTCC CCAAAGGCCC AGCGACATCC1261 TCGAAATCAC AGAGTTTCTG CTGGGAGGGG TTCCATATCC AGTGGCCTAG AATTTGTATC1321 CCACAACCCA CCCAGTGAAG CAGCTACTCC TCCAGTAGCA AGGACCAGTC CCTCGGGGGG1381 AACGTGGTCA TCAGTGGTCA GTGGGGTTCC AAGATTATCC CCTAAAACTC ATAGACCCAG1441 GTCTCCCAGA CAGAACAGTA TTGGAAATAC CCCCAGTGGG CCAGTTCTTG CTTCTCCCCA1501 AGCTGGTATT ATTCCAACTG AAGCTGTTGC CATGCCTATT CCAGCTGCAT CTCCTACGCC1561 TGCTAGTCCT GCATCGAACA GAGCTGTTAC CCCTTCTAGT GAGGCTAAAG ATTCCAGGCT1621 TCAAGATCAG AGGCAGAACT CTCCTGCAGG GAATAAAGAA AATATTAAAC CCAATGAAAC1681 ATCACCTAGC TTCTCAAAAG CTGAAAACAA AGGTATATCA CCAGTTGTTT CTGAACATAG1741 AAAACAGATT GATGATTTAA AGAAATTTAA GAATGATTTT AGGTTACAGC CAAGTTCTAC1801 TTCTGAATCT ATGGATCAAC TACTAAACAA AAATAGAGAG GGAGAAAAAT CAAGAGATTT1861 GATCAAAGAC AAAATTGAAC CAAGTGCTAA GGATTCTTTC ATTGAAAATA GCAGCAGCAA1921 CTGTACCAGT GGCAGCAGCA AGCCGAATAG CCCCAGCATT TCCCCTTCAA TACTTAGTAA1981 CACGGAGCAC AAGAGGGGAC CTGAGGTCAC TTCCCAAGGG GTTCAGACTT CCAGCCCAGC2041 ATGTAAACAA GAGAAAGACG ATAAGGAAGA GAAGAAAGAC GCAGCTGAGC AAGTTAGGAA2101 ATCAACATTG AATCCCAATG CAAAGGAGTT CAACCCACGT TCCTTCTCTC AGCCAAAGCC2161 TTCTACTACC CCAACTTCAC CTCGGCCTCA AGCACAACCT AGCCCATCTA TGGTGGGTCA2221 TCAACAGCCA ACTCCAGTTT ATACTCAGCC TGTTTGTTTT GCACCAAATA TGATGTATCC 2281 AGTCCCAGTG AGCCCAGGCG TGCAACCTTT ATACCCAATA CCTATGACGC CCATGCCAGT2341 GAATCAAGCC AAGACATATA GAGCAGTACC AAATATGCCC CAACAGCGGC AAGACCAGCA2401 TCATCAGAGT GCCATGATGC ACCCAGCGTC AGCAGCGGGC CCACCGATTG CAGCCACCCC2461 ACCAGCTTAC TCCACGCAAT ATGTTGCCTA CAGTCCTCAG CAGTTCCCAA ATCAGCCCCT2521 TGTTCAGCAT GTGCCACATT ATCAGTCTCA GCATCCTCAT GTCTATAGTC CTGTAATACA 2581 GGGTAATGCT AGAATGATGG CACCACCAAC ACACGCCCAG CCTGGTTTAG TATCTTCTTC2641 AGCAACTCAG TACGGGGCTC ATGAGCAGAC GCATGCGATG TATGCATGTC CCAAATTACC2701 ATACAACAAG GAGACAAGCC CTTCTTTCTA CTTTGCCATT TCCACGGGCT CCCTTGCTCA 2761 GCAGTATGCG CACCCTAACG CTACCCTGCA CCCACATACT CCACACCCTCAGCCTTCAGC2821 TACCCCCACT GGACAGCAGC AAAGCCAACA TGGTGGAAGT CATCCTGCAC CCAGTCCTGT2881 TCAGCACCAT CAGCACCAGG CCGCCCAGGC TCTCCATCTG GCCAGTCCAC AGCAGCAGTC2941 AGCCATTTAC CACGCGGGGC TTGCGCCAAC TCCACCCTCC ATGACACCTG CCTCCAACAC3001 GCAGTCGCCA CAGAATAGTT TCCCAGCAGC ACAACAGACT GTCTTTACGA TCCATCCTTC3061 TCACGTTCAG CCGGCGTATA CCAACCCACC CCACATGGCC CACGTACCTC AGTGCGCCAG3121 TGAGGCTCTG GCAAGGTGTG GGCTAGAGAT GCGACTCAGT TGGATCTATC TCTCAGAAGG3181 CTACCTTGCT CATGTACAGT CAGGAATGGT TCCTTCTCAT CCAACTGCCC ATGCGCCAAT 3241 GATGCTAATG ACGACACAGC CACCCGGCGG TCCCCAGGCC GCCCTCGCTC AAAGTGCACT3301 ACAGCCCATT CCAGTCTCGA CAACAGCGCA TTTCCCCTAT ATGACGCACC CTTCAGTACA3361 AGCCCACCAC CAACAGCAGT TGTAAGGCTG CCCTGGAGGA ACCGAAAGGC CAAATTCCCT3421 CCTCCCTTCT ACTGCTTCTA CCAACTGGAA GCACAGAAAA CTAGAATTTC ATTTATTTTG 3481 TTTTTAAAAT ATATATGTTG ATTTCTTGTA ACATCCAATA GGAATGCTAA CAGTTCACTT 3541 GCAGTGGAAG ATACTTGGAC CGAGTAGAGG CATTTAGGAA CTTGGGGGCT ATTCCATAAT3601 TCCATATGCT GTTTCAGAGT CCCGCAGGTA CCCCAGCTCT GCTTGCCGAA ACTGGAAGTT3661 ATTTATTTTT TAATAACCCT TGAAAGTCAT GAACACATCA GCTAGCAAAA GAAGTAACAA3721 GAGTGATTCT TGCTGCTATT ACTGCTAAAA AAAAAAAAAA AAAAAAATCA AGACTTGGAA3781 CGCCCTTTTA CTAAACTTGA CAAAGTTTCA GTAAATTCTT ACCGTCAAAC TGACGGATTA3841 TTATTTATAA ATCAAGTTTG ATGAGGTGAT CACTGTCTAC AGTGGTTCAA CTTTTAAGTT 3901 AAGGGAAAAA CTTTTACTTT GTAGATAATA TAAAATAAAA ACTTAAAAAA AATTTAAAAA3961 ATAAAAAAAG TTTTAAAAAC TGAAAAAAAA AAA (SEQ ID NO: 55)1 MRMVHILTSV VGSKCEVQVK NGGIYEGVFK TYSPKCDLVL DAAHEKSTES SSGPKREEIM 61 ESILFKCSDF WVQFKDMDS SYAKRDAFTD SAISAKVNGE HKEKDLEPWD AGELTANEEL 121 EALENDVSNG WDPNDMFRYN EENYGWSTY DSSLSSYTVP LERDNSEEFL KREARANQLA 181 EEIESSAQYK ARVALENDDR SEEEKYTAVQ RNSSEREGHS INTRENKYIP PGQRNREVIS 241 WGSGRQNSPR MGQPGSGSMP SRSTSHTSDF NPNSGSDQRV VNGGVPWPSP CPSPSSRPPS 301 RYQSGPNSLP PRAATPTRPP SRPPSRPSRP PSHPSAHGSP APVSTMPKRM SSEGPPRMSP361 KAQRHPRNHR VS GRGSISS GLEFVSHNPP SEAATPPVAR TSPSGGTWSS VVSGVPRLSP 421 KTHRPRSPRQ NSIGNTPSGP VLASPQAGI I PTEAVAMPIP AASPTPASPA SNRAVTPSSE 481 AKDSRLQDQR QNSPAGNKEN IKPNETSPSF SKAENKGISP WSEHRKQID DLKKFKNDFR 541 LQPSSTSESM DQLLNKNREG EKSRDLIKDK IEPSAKDSFI ENSSSNCTSG SSKPNSPSIS 601 PSILSNTEHK RGPEVTSQGV QTSSPACKQE KDDKEEKKDA AEQVRKSTLN PNAKEFNPRS 661 FSQPKPSTTP TSPRPQAQPS PSMVGHQQPT PVYTQPVCFA PNMMYPVPVS PGVQPLYPIP 721 MTPMPVNQAK TYRAVPNMPQ QRQDQHHQSA MMHPASAAGP PIAATPPAYS TQYVAYSPQQ 781 FPNQPLVQHV PHYQSQHPHV YSPVIQGNAR MMAPPTHAQP GLVSSSATQY GAHEQTHAMY 841 ACPKLPYNKE TSPSFYFAIS TGSLAQQYAH PNATLHPHTP HPQPSATPTG QQQSQHGGSH 901 PAPSPVQHHQ HQAAQALHLA SPQQQSAIYH AGLAPTPPSM TPASNTQSPQ NSFPAAQQTV 961 FTIHPSHVQP AYTNPPHMAH VPQCASEALA RCGLEMRLSW IYLSEGYLAH VQSGMVPSHP 1021 TAHAPMMLMT TQPPGGPQAA LAQSALQPIP VSTTAHFPYM THPSVQAHHQ QQL (SEQ ID NO: 56)[000108] The human ATXN2 isoform 3 mRNA sequence can be found at NM_001310123.1 (SEQ ID NO: 57); and the corresponding protein sequence can be found at NP_001297052.1 (SEQ ID NO: 58).1 CCCGAGAAAG CAACCCAGCG CGCCGCCCGC TCCTCACGTG TCCCTCCCGG CCCCGGGGCC 61 ACCTCACGTT CTGCTTCCGT CTGACCCCTC CGACTTCCGA TTTCTTTTGA TGGAATCTAT 121 GCAAATATGA GGATGGTTCA TATACTTACA TCAGTTGTTT GTGATTTGGT ACTTGATGCC 181 GCACATGAGA AAAGTACAGA ATCCAGTTCG GGGCCGAAAC GTGAAGAAAT AATGGAGAGT241 ATTTTGTTCA AATGTTCAGA CTTTGTTGTG GTACAGTTTA AAGATATGGA CTCCAGTTAT 301 GCAAAAAGAG ATGCTTTTAC TGACTCTGCT ATCAGTGCTA AAGTGAATGG CGAACACAAA361 GAGAAGGACC TGGAGCCCTG GGATGCAGGT GAACTCACAG CCAATGAGGA ACTTGAGGCT421 TTGGAAAATG ACGTATCTAA TGGATGGGAT CCCAATGATA TGTTTCGATA TAATGAAGAA481 AATTATGGTG TAGTGTCTAC GTATGATAGC AGTTTATCTT CGTATACAGT GCCCTTAGAA 541 AGAGATAACT CAGAAGAATT TTTAAAACGG GAAGCAAGGG CAAACCAGTT AGCAGAAGAA601 ATTGAGTCAA GTGCCCAGTA CAAAGCTCGA GTGGCCCTGG AAAATGATGA TAGGAGTGAG661 GAAGAAAAAT ACACAGCAGT TCAGAGAAAT TCCAGTGAAC GTGAGGGGCA CAGCATAAAC721 ACTAGGGAAA ATAAATATAT TCCTCCTGGA CAAAGAAATA GAGAAGTCAT ATCCTGGGGA781 AGTGGGAGAC AGAATTCACC GCGTATGGGC CAGCCTGGAT CGGGCTCCAT GCCATCAAGA841 TCCACTTCTC ACACTTCAGA TTTCAACCCG AATTCTGGTT CAGACCAAAG AGTAGTTAAT 901 GGAGGTGTTC CCTGGCCATC GCCTTGCCCA TCTCCTTCCT CTCGCCCACC TTCTCGCTAC 961 CAGTCAGGTC CCAACTCTCT TCCACCTCGG GCAGCCACCC CTACACGGCCGCCCTCCAGG1021 CCCCCCTCGC GGCCATCCAG ACCCCCGTCT CACCCCTCTG CTCATGGTTC TCCAGCTCCT 1081 GTCTCTACTA TGCCTAAACG CATGTCTTCA GAAGGGCCTC CAAGGATGTC CCCAAAGGCC1141 CAGCGACATC CTCGAAATCA CAGAGTTTCT GCTGGGAGGG GTTCCATATC CAGTGGCCTA1201 GAATTTGTAT CCCACAACCC ACCCAGTGAA GCAGCTACTC CTCCAGTAGC AAGGACCAGT1261 CCCTCGGGGG GAACGTGGTC ATCAGTGGTC AGTGGGGTTC CAAGATTATC CCCTAAAACT1321 CATAGACCCA GGTCTCCCAG ACAGAACAGT ATTGGAAATA CCCCCAGTGG GCCAGTTCTT1381 GCTTCTCCCC AAGCTGGTAT TATTCCAACT GAAGCTGTTG CCATGCCTAT TCCAGCTGCA 1441 TCTCCTACGC CTGCTAGTCC TGCATCGAAC AGAGCTGTTA CCCCTTCTAG TGAGGCTAAA1501 GATTCCAGGC TTCAAGATCA GAGGCAGAAC TCTCCTGCAG GGAATAAAGA AAATATTAAA1561 CCCAATGAAA CATCACCTAG CTTCTCAAAA GCTGAAAACA AAGGTATATC ACCAGTTGTT1621 TCTGAACATA GAAAACAGAT TGATGATTTA AAGAAATTTA AGAATGATTT TAGGTTACAG1681 CCAAGTTCTA CTTCTGAATC TATGGATCAA CTACTAAACA AAAATAGAGA GGGAGAAAAA1741 TCAAGAGATT TGATCAAAGA CAAAATTGAA CCAAGTGCTA AGGATTCTTT CATTGAAAAT1801 AGCAGCAGCA ACTGTACCAG TGGCAGCAGC AAGCCGAATA GCCCCAGCAT TTCCCCTTCA1861 ATACTTAGTA ACACGGAGCA CAAGAGGGGA CCTGAGGTCA CTTCCCAAGG GGTTCAGACT1921 TCCAGCCCAG CATGTAAACA AGAGAAAGAC GATAAGGAAG AGAAGAAAGA CGCAGCTGAG1981 CAAGTTAGGA AATCAACATT GAATCCCAAT GCAAAGGAGT TCAACCCACG TTCCTTCTCT2041 CAGCCAAAGC CTTCTACTAC CCCAACTTCA CCTCGGCCTC AAGCACAACC TAGCCCATCT2101 ATGGTGGGTC ATCAACAGCC AACTCCAGTT TATACTCAGC CTGTTTGTTT TGCACCAAAT2161 ATGATGTATC CAGTCCCAGT GAGCCCAGGC GTGCAACCTT TATACCCAAT ACCTATGACG2221 CCCATGCCAG TGAATCAAGC CAAGACATAT AGAGCAGTAC CAAATATGCCCCAACAGCGG2281 CAAGACCAGC ATCATCAGAG TGCCATGATG CACCCAGCGT CAGCAGCGGG CCCACCGATT2341 GCAGCCACCC CACCAGCTTA CTCCACGCAA TATGTTGCCT ACAGTCCTCA GCAGTTCCCA2401 AATCAGCCCC TTGTTCAGCA TGTGCCACAT TATCAGTCTC AGCATCCTCA TGTATAGТCC 2461 CCTGTAATAC AGGGTAATGC TAGAATGATG GCACCACCAA CACACGCCCA GCCTGGTTTA2521 GTATCTTCTT CAGCAACTCA GTACGGGGCT CATGAGCAGA CGCATGCGAT GTATGTTTCC2581 ACGGGCTCCC TTGCTCAGCA GTATGCGCAC CCTAACGCTA CCCTGCACCC ACATACTCCA2641 CACCCTCAGC CTTCAGCTAC CCCCACTGGA CAGCAGCAAA GCCAACATGG TGGAAGTCAT2701 CCTGCACCCA GTCCTGTTCA GCACCATCAG CACCAGGCCG CCCAGGCTCT CCATCTGGCC2761 AGTCCACAGC AGCAGTCAGC CATTTACCAC GCGGGGCTTG CGCCAACTCC ACCCTCCATG2821 ACACCTGCCT CCAACACGCA GTCGCCACAG AATAGTTTCC CAGCAGCACA ACAGACTGTC2881 TTTACGATCC ATCCTTCTCA CGTTCAGCCG GCGTATACCA ACCCACCCCA CATGGCCCAC2941 GTACCTCAGG CTCATGTACA GTCAGGAATG GTTCCTTCTC ATCCAACTGC CCATGCGCCA3001 ATGATGCTAA TGACGACACA GCCACCCGGC GGTCCCCAGG CCGCCCTCGC TCAAAGTGCA3061 CTACAGCCCA TTCCAGTCTC GACAACAGCG CATTTCCCCT ATATGACGCA CCCTTCAGTA3121 CAAGCCCACC ACCAACAGCA GTTGTAAGGC TGCCCTGGAG GAACCGAAAG GCCAAATTCC3181 CTCCTCCCTT CTACTGCTTC TACCAACTGG AAGCACAGAA AACTAGAATT TCATTTATTT 3241 TGTTTTTAAA ATATATATGT TGATTTCTTG TAACATCCAA TAGGAATGCT AACAGTTCAC 3301 TTGCAGTGGA AGATACTTGG ACCGAGTAGA GGCATTTAGG AACTTGGGGG CTATTCCATA3361 ATTCCATATG CTGTTTCAGA GTCCCGCAGG TACCCCAGCT CTGCTTGCCG AAACTGGAAG3421 TTATTTATTT TTTAATAACC CTTGAAAGTC ATGAACACAT CAGCTAGCAA AAGAAGTAAC3481 AAGAGTGATT CTTGCTGCTA TTACTGCTAA AAAAAAAAAA AAAAAAAAATCAAGACTTGG3541 AACGCCCTTT TACTAAACTT GACAAAGTTT CAGTAAATTC TTACCGTCAA ACTGACGGAT3601 TATTATTTAT AAATCAAGTT TGATGAGGTG ATCACTGTCT ACAGTGGTTC AACTTTTAAG 3661 TTAAGGGAAA AACTTTTACT TTGTAGATAA TATAAAATAA AAACTTAAAA AAAATTTAAA3721 AAATAAAAAA AGTTTTAAAA ACTGAAAAAA AAAAA (SEQ ID NO: 57)1 MRMVHILTSV VCDLVLDAAH EKSTESSSGP KREEIMESIL FKCSDFVVVQ FKDMDSSYAK 61 RDAFTDSAIS AKVNGEHKEK DLEPWDAGEL TANEELEALE NDVSNGWDPN DMFRYNEENY 121 GVVSTYDSSL SSYTVPLERD NSEEFLKREA RANQLAEEIE SSAQYKARVA LENDDRSEEE 181 KYTAVQRNSS EREGHSINTR ENKYIPPGQR NREVISWGSG RQNSPRMGQP GSGSMPSRST 241 SHTSDFNPNS GSDQRVVNGG VPWPSPCPSP SSRPPSRYQS GPNSLPPRAA TPTRPPSRPP 301 SRPSRPPSHP SAHGSPAPVS TMPKRMSSEG PPRMSPKAQR HPRNHRVSAG RGSISSGLEF 361 VSHNPPSEAA TPPVARTSPS GGTWSSVVSG VPRLSPKTHR PRSPRQNSIG NTPSGPVLAS 421 PQAGIIPTEA VAMPIPAASP TPASPASNRA VTPSSEAKDS RLQDQRQNSP AGNKENIKPN 481 ETSPSFSKAE NKGISPVVSE HRKQIDDLKK FKNDFRLQPS STSESMDQLL NKNREGEKSR 541 DLIKDKIEPS AKDSFIENSS SNCTSGSSKP NSPSISPSIL SNTEHKRGPE VTSQGVQTSS 601 PACKQEKDDK EEKKDAAEQV RKSTLNPNAK EFNPRSFSQP KPSTTPTSPR PQAQPSPSMV 661 GHQQPTPVYT QPVCFAPNMM YPVPVSPGVQ PLYPIPMTPM PVNQAKTYRA VPNMPQQRQD 721 QHHQSAMMHP ASAAGPPIAA TPPAYSTQYV AYSPQQFPNQ PLVQHVPHYQ SQHPHVYSPV 781 IQGNARMMAP PTHAQPGLVS SSATQYGAHE QTHAMYVSTG SLAQQYAHPN ATLHPHTPHP 841 QPSATPTGQQ QSQHGGSHPA PSPVQHHQHQ AAQALHLASP QQQSAIYHAG LAPTPPSMTP 901 ASNTQSPQNS FPAAQQTVFT IHPSHVQPAY TNPPHMAHVP QAHVQSGMVP SHPTAHAPMM 961 LMTTQPPGGP QAALAQSALQ PIPVSTTAHF PYMTHPSVQA HHQQQL (SEQ ID NO: 58) [000109] The human ATXN2 isoform 4 mRNA sequence can be found at NM 001372574.1 (SEQ ID NO: 59); the corresponding protein sequence can be found at NP_001359503.1 (SEQ ID NO: 60).1 AGAGCTCGCC TCCCTCCGCC TCAGACTGTT TTGGTAGCAA CGGCAACGGC GGCGGCGCGT 61 TTCGGCCCGG CTCCCGGCGG CTCCTTGGTC TCGGCGGGCC TCCCCGCCCC TTCGTCGTCC 121 TCCTTCTCCC CCTCGCCAGC CCGGGCGCCC CTCCGGCCGC GCCAACCCGC GCCTCCCCGC 181 TCGGCGCCCG CGCGTCCCCG CCGCGTTCCG GCGTCTCCTT GGCGCGCCCG GCTCCCGGCT 241 GTCCCCGCCC GGCGTGCGAG CCGGTGTATG GGCCCCTCAC CATGTCGCTG AAGCCCCAGC301 AGCAGCAGCA GCAGCAGCAG CAGCAGCAGC AGCAGCAACA GCAGCAGCAG CAGCAGCAGC361 AGCAGCCGCC GCCCGCGGCT GCCAATGTCC GCAAGCCCGG CGGCAGCGGC CTTCTAGCGT421 CGCCCGCCGC CGCGCCTTCG CCGTCCTCGT CCTCGGTCTC CTCGTCCTCG GCCACGGCTC 481 CCTCCTCGGT GGTCGCGGCG ACCTCCGGCG GCGGGAGGCC CGGCCTGGGC AGAGGTCGAA541 ACAGTAACAA AGGACTGCCT CAGTCTACGA TTTCTTTTGA TGGAATCTAT GCAAATATGA601 GGATGGTTCA TATACTTACA TCAGTTGTTG GCTCCAAATG TGAAGTACAAGTGAAAAATG661 GAGGTATATA TGAAGGAGTT TTTAAAACTT ACAGTCCGAA GTGTGATTTG GTACTTGATG721 CCGCACATGA GAAAAGTACA GAATCCAGTT CGGGGCCGAA ACGTGAAGAA ATAATGGAGA781 GTATTTTGTT CAAATGTTCA GACTTTGTTG TGGTACAGTT TAAAGATATG GACTCCAGTT 841 ATGCAAAAAG AGATGCTTTT ACTGACTCTG CTATCAGTGC TAAAGTGAAT GGCGAACACA901 AAGAGAAGGA CCTGGAGCCC TGGGATGCAG GTGAACTCAC AGCCAATGAG GAACTTGAGG961 CTTTGGAAAA TGACGTATCT AATGGATGGG ATCCCAATGA TATGTTTCGA TATAATGAAG1021 AAAATTATGG TGTAGTGTCT ACGTATGATA GCAGTTTATC TTCGTATACA GTGCCCTTAG1081 AAAGAGATAA CTCAGAAGAA TTTTTAAAAC GGGAAGCAAG GGCAAACCAG TTAGCAGAAG1141 AAATTGAGTC AAGTGCCCAG TACAAAGCTC GAGTGGCCCT GGAAAATGAT GATAGGAGTG1201 AGGAAGAAAA ATACACAGCA GTTCAGAGAA ATTCCAGTGA ACGTGAGGGG CACAGCATAA1261 ACACTAGGGA AAATAAATAT ATTCCTCCTG GACAAAGAAA TAGAGAAGTC ATATCCTGGG1321 GAAGTGGGAG ACAGAATTCA CCGCGTATGG GCCAGCCTGG ATCGGGCTCC ATGCCATCAA1381 GATCCACTTC TCACACTTCA GATTTCAACC CGAATTCTGG TTCAGACCAA AGAGTAGTTA1441 ATGGAGGTGT TCCCTGGCCA TCGCCTTGCC CATCTCCTTC CTCTCGCCCA CCTTCTCGCT 1501 ACCAGTCAGG TCCCAACTCT CTTCCACCTC GGGCAGCCAC CCCTACACGG CCGCCCTCCA1561 GGCCCCCCTC GCGGCCATCC AGACCCCCGT CTCACCCCTC TGCTCATGGT TCTCCAGCTC 1621 CTGTCTCTAC TATGCCTAAA CGCATGTCTT CAGAAGGGCC TCCAAGGATG TCCCCAAAGG1681 CCCAGCGACA TCCTCGAAAT CACAGAGTTT CTGCTGGGAG GGGTTCCATA TCCAGTGGCC1741 TAGAATTTGT ATCCCACAAC CCACCCAGTG AAGCAGCTAC TCCTCCAGTA GCAAGGACCA1801 GTCCCTCGGG GGGAACGTGG TCATCAGTGG TCAGTGGGGT TCCAAGATTA TCCCCTAAAA1861 CTCATAGACC CAGGTCTCCC AGACAGAACA GTATTGGAAA TACCCCCAGT GGGCCAGTTC1921 TTGCTTCTCC CCAAGCTGGT ATTATTCCAA CTGAAGCTGT TGCCATGCCT ATTCCAGCTG 1981 CATCTCCTAC GCCTGCTAGT CCTGCATCGA ACAGAGCTGT TACCCCTTCT AGTGAGGCTA2041 AAGATTCCAG GCTTCAAGAT CAGAGGCAGA ACTCTCCTGC AGGGAATAAA GAAAATATTA2101 AACCCAATGA AACATCACCT AGCTTCTCAA AAGCTGAAAA CAAAGGTATA TCACCAGTTG2161 TTTCTGAACA TAGAAAACAG ATTGATGATT TAAAGAAATT TAAGAATGAT TTTAGGTTAC2221 AGCCAAGTTC TACTTCTGAA TCTATGGATC AACTACTAAA CAAAAATAGA GAGGGAGAAA2281 AATCAAGAGA TTTGATCAAA GACAAAATTG AACCAAGTGC TAAGGATTCT TTCATTGAAA2341 ATAGCAGCAG CAACTGTACC AGTGGCAGCA GCAAGCCGAA TAGCCCCAGC ATTTCCCCTT2401 CAATACTTAG TAACACGGAG CACAAGAGGG GACCTGAGGT CACTTCCCAA GGGGTTCAGA2461 CTTCCAGCCC AGCATGTAAA CAAGAGAAAG ACGATAAGGA AGAGAAGAAA GACGCAGCTG2521 AGCAAGTTAG GAAATCAACA TTGAATCCCA ATGCAAAGGA GTTCAACCCACGTTCCTTCT2581 CTCAGCCAAA GCCTTCTACT ACCCCAACTT CACCTCGGCC TCAAGCACAA CCTAGCCCAT2641 CTATGGTGGG TCATCAACAG CCAACTCCAG TTTATACTCA GCCTGTTTGT TTTGCACCAA 2701 ATATGATGTA TCCAGTCCCA GTGAGCCCAG GCGTGCAACC TTTATACCCA ATACCTATGA2761 CGCCCATGCC AGTGAATCAA GCCAAGACAT ATAGAGCAGG TAAAGTACCA AAATATGCCCC2821 AACAGCGGCA AGACCAGCAT CATCAGAGTG CCATGATGCA CCCAGCGTCA GCAGCGGGCC2881 CACCGATTGC AGCCACCCCA CCAGCTTACT CCACGCAATA TGTTGCCTAC AGTCCTCAGC2941 AGTTCCCAAA TCAGCCCCTT GTTCAGCATG TGCCACATTA TCAGTCTCAG CATCCTCATG 3001 TCTATAGTCC TGTAATACAG GGTAATGCTA GAATGATGGC ACCACCAACA CACGCCCAGC3061 CTGGTTTAGT ATCTTCTTCA GCAACTCAGT ACGGGGCTCA TGAGCAGACG CATGCGATGT3121 ATGCATGTCC CAAATTACCA TACAACAAGG AGACAAGCCC TTCTTTCTAC TTTGCCATTT3181 CCACGGGCTC CCTTGCTCAG CAGTATGCGC ACCCTAACGC TACCCTGCAC CCACATACTC3241 CACACCCTCA GCCTTCAGCT ACCCCCACTG GACAGCAGCA AAGCCAACAT GGTGGAAGTC3301 ATCCTGCACC CAGTCCTGTT CAGCACCATC AGCACCAGGC CGCCCAGGCT CTCCATCTGG3361 CCAGTCCACA GCAGCAGTCA GCCATTTACC ACGCGGGGCT TGCGCCAACT CCACCCTCCA3421 TGACACCTGC CTCCAACACG CAGTCGCCAC AGAATAGTTT CCCAGCAGCA CAACAGACTG3481 TCTTTACGAT CCATCCTTCT CACGTTCAGC CGGCGTATAC CAACCCACCC CACATGGCCC 3541 ACGTACCTCA GGCTCATGTA CAGTCAGGAA TGGTTCCTTC TCATCCAACT GCCCATGCGC3601 CAATGATGCT AATGACGACA CAGCCACCCG GCGGTCCCCA GGCCGCCCTC GCTCAAAGTG3661 CACTACAGCC CATTCCAGTC TCGACAACAG CGCATTTCCC CTATATGACG CACCCTTCAG3721 TACAAGCCCA CCACCAACAG CAGTTGTAAG GCTGCCCTGG AGGAACCGAA AGGCCAAATT3781 CCCTCCTCCC TTCTACTGCT TCTACCAACT GGAAGCACAG AAAACTAGAA TTTCATTTAT 3841 TTTGTTTTTA AAATATATAT GTTGATTTCT TGTAACATCC AATAGGAATG CTAACAGTTC 3901 ACTTGCAGTG GAAGATACTT GGACCGAGTA GAGGCATTTA GGAACTTGGG GGCTATTCCA3961 TAATTCCATA TGCTGTTTCA GAGTCCCGCA GGTACCCCAG CTCTGCTTGC CGAAACTGGA4021 AGTTATTTAT TTTTTAATAA CCCTTGAAAG TCATGAACAC ATCAGCTAGC AAAAGAAGTA4081 ACAAGAGTGA TTCTTGCTGC TATTACTGCT AAAAAAAAAA AAAAAAAAAA ATCAAGACTT4141 GGAACGCCCT TTTACTAAAC TTGACAAAGT TTCAGTAAAT TCTTACCGTC AAACTGACGG4201 ATTATTATTT ATAAATCAAG TTTGATGAGG TGATCACTGT CTACAGTGGT TCAACTTTTA 4261 AGTTAAGGGA AAAACTTTTA CTTTGTAGAT AATATAAAAT AAAAACTTAA AAAAAATTTA4321 AAAAATAAAA AAAGTTTTAA AAACTGA (SEQ ID NO: 59)1 MSLKPQQQQQ QQQQQQQQQQ QQQQQQQQPP PAAANVRKPG GSGLLASPAA APSPSSSSVS 61 SSSATAPSSV VAATSGGGRP GLGRGRNSNK GLPQSTISFD GIYANMRMVH ILTSVVGSKC 121 EVQVKNGGIY EGVFKTYSPK CDLVLDAAHE KSTESSSGPK REEIMESILF KCSDFVVVQF 181 KDMDSSYAKR DAFTDSAISA KVNGEHKEKD EEPWDAGEET ANFEEFAEFN DVSNGWDPND 241 MFRYNEENYG VVSTYDSSLS SYTVPLERDN SEEFLKREAR ANQLAEEIES SAQYKARVAL 301 ENDDRSEEEK YTAVQRNSSE REGHSINTRE NKYIPPGQRN REVISWGSGR QNSPRMGQPG361 SGSMPSRSTS HTSDFNPNSG SDQRVVNGGV PWPSPCPSPS SRPPSRYQSG PNSLPPRAAT 421 PTRPPSRPPS RPSRPPSHPS AHGSPAPVST MPKRMSSEGP PRMSPKAQRH PRNHRVSAGR 481 GSISSGLEFV SHNPPSEAAT PPVARTSPSG GTWSSVVSGV PRLSPKTHRP RSPRQNSIGN 541 TPSGPVLASP QAGIIPTEAV AMPIPAASPT PASPASNRAV TPSSEAKDSR LQDQRQNSPA 601 GNKENIKPNE TSPSFSKAEN KGISPVVSEH RKQIDDLKKF KNDFRLQPSS TSESMDQLLN 661 KNREGEKSRD LIKDKIEPSA KDSFIENSSS NCTSGSSKPN SPSISPSILS NTEHKRGPEV 721 TSQGVQTSSP ACKQEKDDKE EKKDAAEQVR KSTLNPNAKE FNPRSFSQPK PSTTPTSPRP 781 QAQPSPSMVG HQQPTPVYTQ PVCFAPNMMY PVPVSPGVQP LYPIPMTPMP VNQAKTYRAG 841 KVPNMPQQRQ DQHHQSAMMH PASAAGPPIA ATPPAYSTQY VAYSPQQFPN QPLVQHVPHY 901 QSQHPHVYSP VIQGNARMMA PPTHAQPGLV SSSATQYGAH EQTHAMYACP KLPYNKETSP 961 SFYFAISTGS LAQQYAHPNA TLHPHTPHPQ PSATPTGQQQ SQHGGSHPAP SPVQHHQHQA 1021 AQALHLASPQ QQSAIYHAGL APTPPSMTPA SNTQSPQNSF PAAQQTVFTI HPSHVQPAYT 1081 NPPHMAHVPQ AHVQSGMVPS HPTAHAPMML MTTQPPGGPQ AALAQSALQP IPVSTTAHFP 1141 YMTHPSVQAH HQQQL (SEQ ID NO: 60)[000110] Another human ATXN2 isoform (that includes UTRs) mRNA sequence (SEQ ID NO: 61) and the corresponding protein sequence (SEQ ID NO: 62) can be found at ENST00000608853.5.ATGTCGCT AAGCCCCAGCAGCAGCAGCAGCAGCAGCAGC AGCAGCAGCAGCAGCAACAGCAGC AGC GCAGCAGCAGCAGCAGCCGCCGCCCGCGGCTGCCAATGTCCGCAAGCCCGGCGGCAGCGGC CTTCTAGCGTCGCCCGCCGCCGCGCCTTCGCCGTCCTCGTCCTCGGTCTCCTCGTCCTCGGCCACGG CTCCCTCCTCGGTGGTCGCGGCGACCTCCGGCGGCGGGAGGCCCGGCCTGGGCAGAGGTCGAAAC AGTAACAAAGGACTGCCTCAGTCTACGATTTCTTTTGATGGAATCTATGCAAATATGAGGATGGTT CAT TACTTACATCAGTTGTTGGCTCCAAATGTGAAGTACAAGTGAAAAATGGAGGTATATATGAA GGAGTTTTTAAAACTTACAGTCCGAAGTGTGATTTG TACTTGATGCCGCACATGAGAAAAGTACA GA ATCC AGTTCGGGGCN G AACG TGAAG AAAT AATGG AG AGT ATITTGTTCA ATGTTC AG ACTTT I TG I GG 1 ACAG I T I AAAGA't' IGGACTGCAG Pi 'A. JCAAAAAGAGATGC T T 14 C H'GAC I C KY'! ATCAGIXKGAAAGTGAA1X1GCGAACACAAAGAGAAGGACCTGGAGCCCTGGGATGCAGGTGAAC TCACAR3CCAATGAGGAACTTGAGGCTTTGGAAAATGACGTATCTAATGGATGGGATCCCAATGAT ATGTTTCGATATAATGAAGAAAATTATGGTGTAGTGTCTACGTATGATAGCAGTTTATCTTCGTAT ACAGTGCCCTTAGAAAGAGATAACTCAGAAGAATTTTTAAAACGGGAAGCAAGGGCAAACCAGTT AGCAGAAGAAATTGAGTCAAGTGGCCAGTACAAAGCTCGAGTGGCCCTGGAAAATGATGATAGGA GTGAGGAAGAAAAATACACAGCAGTTCAGAGAAATTCCAGTGAACGTGAGGGGCACAGCATAAA CAC J AGG GA AAAI AAA TATA TTCC TCCTGGACAAAGAAATAGAGAAGTCATA'i'CCTGGGGAAGTG GGAGACAG ATTCACCGCG’IATGGGCCAGCCTGGATCGGGCTCCATGCCAFCAAGATCCACTTCFC ACACTTCAGATTTCAACCCGAATTCTGGTTCAGACCAAAGAGTAGTTAATGGAGGTGTTCCCTGGC CATCGCCTTGCCCATCTCCTTCCTCTCGCCCACCTTCTCGCTACCAGTCAGGTCCCAACTCTCTTCC ACCTCGGGCAGCCACCCCTACACGGCCGCCCTCCAGGCCCCCCTCGCGGCCATCCAGACCCCCGTC TCACCCCTCTGCTCATGGTTCTCCAGCTCCTGTCTCTACTATGCCTAAACGCATGTCTTCAGAAGGG CCTCCAAGGATGTCCCCAAAGGCCCAGCGACATCCTCGAAATCACAGAGTTTCTGCTGGGAGGGG T TCCATATCCAG't'GGCC'l AGAATTTGlATCGCACAACCGACGCAGTGAAGCAGC'l ACTCC TCCAGT AGCAAGGACCAGTCCCGCGGGGGGAACGTGGTCATCAGTGGIGAxGTGGGGTTCCAAGATTATCCC CTAAA / xCTCATAGACCCAiGGTC / rCCCAGACAGAACAxGTATTGGAAATACCCCCAGTGGGCCAGTT CTTGCTTCTCCCCAAGCTGGTATTATTCCAACTGAAGCTGTTGCCATGCCTATTCCAGCTGCATCTCCTACGCCTGCTAGTCCTGCATCGAACAGAGCTGTTACCCCTTCTAGTGAGGCTAAAGATTCCAGGC TTCAAGATCAGAGGCAGAACTCTCCTGCAGGGAATAAAGAAAATATTAAACCCAATGAAACATCA CCTAGCTTCTCAAAAGCTGAAAACAAAGGTATATCACCAGTTGTTTCTGAACATAGA / XAACAGATT G ATG ATTT A AAG AA ATTT A AG AATG ATTTT AGGTT AC AG CC A AGTTCT ACTTCTG A ATCT ATGG AT CAACTACTAAACAAAAATAGAGAGGGAGAAAAATCAAGAGATTTGATCAAAGACAAAATTGAAC C AAG TGCT A A( 1G ATTCTT TC ATTG A A AAT AGC AGC AGC AACTG T ACC AG TGGC AGC AGC AAGCC G A AT AGCC CC AGC ATT TCCCCTT C A AT A CT T AG T A AC ACGG AGC AC AAG AG GGG ACCT G AG GTC A C TTCCCAAGGGGTTCAGACTTCCAGCCCAGCATGTAAACAAGAGAAAGACGATAAGGAAGAGAAG AAAGACGCAGCTGAGCAAGTTAGGAAATCAACATTGAATCCCAATGCAAAGGAGTTCAACCCACG TTCCTTCTCTCAGCCAAAGCCTTCTACTACCCCAACTTCACCTCGGCCTCAAGCACAACCTAGCCCA TCTATGGTGGGTCATCAACAGCCAACTCCAGTTTATACTCAGCCTGTTTGTTTTGCACCAAATATGA TGTATCCAGTCCCAGTGAGCCCAGGCGTGCAACCTTTATACCCAATACCTATGACGCCCATGCCAG TGAATCAAGCCAAGACATATAGAGCAGTACCAAATATGCCCCAACAGCGGCAAGACCAGCATCAT CAGAGTGCCATGATGCACCCAGCGTCAGCAGCGGGCCCACCGATTGCAGCCACCCCACCAGCTTA CTCCACGCAATATGTTGCCTACAGTCCTCAGCAGTTCCCAAATCAGCCCCTTGTTCAGCATGTGCC ACATTATCAGTCTCAGCATCCTCATGTCTATAGTCCTGTAATACAGGGTAATGCTAGAATGATGGC ACCACCAACACACGCCCAGCCTGGTTTAGTATCTTCTTCAGCAACTCAGTACGGGGCTCATGAGCA GACGCATGCGATGTATGCATGTCCCAAATTACCATACAACAAGGAGACAAGCCCTTCTTTCTACTT TGCCATTTCCACGGGCTCCCTTGCTCAGCAGTATGCGCACCCTAACGCTACCCTGCACCCACATAC TCCACACCCTCAGCCTTCAGCTACCCCCACTGGACAGCAGCAAAGCCAACATGGTGGAAGTCATCC TGCACCCAGTCCTGTTCAGCACCATCAGCACCAGGCCGCCCAGGCTCTCCATCTGGCCAGTCCACA GCAGCAGTCAGCCATTTACCACGCGGGGCTTGCGCCAACTCCACCCTCCATGACACCTGCCTCCAA CACGCAGTCGCCACAGAATAGTTTCCCAGCAGCACAACAGACTGTCTTTACGATCCATCCTTCTCA CGTTCAGCCGGCGTATACCAACCCACCCCACATGGCCCACGTACCTCAGGCTCATGTACAGTCAGG AATGGTTCCTTCTCATCCAACTGCCCATGCGCCAATGATGCTAATGACGACACAGCCACCCGGCGG TCCCCAGGCCGCCCTCGCTCAAAGTGCACTACAGCCCATTCCAGTCTCGACAACAGCGCATTTCCC CTATATGACGCACCCTTCAGTACAAGCCCACCACCAACAGCAGTTGTAA (SEQID NO: 61 ) MSLKPQQQQQQQQQQQQQQQQQQQQQQPPPAAANVRKPGGSGLLASPAAAPSPSSSSVSSSSATAPSSVVAATSGGGRPGLGRGRNSNKGLPQSTISFDGIYANMRMVHILTSVVGSKCEVQVKNGGIYEGVFKT YSPKCDLVLDAAHEKSTESSSGPKREEIMESILFKCSDFVVVQFKDMDSSYAKRDAFTDSAISAKVNGE HKEKDEEPWDAGEETANFEEFAEFNDVSNGWDPNDMFRYNEENYGVVSTYDSSLSSYTVPLERDNSE EFLKREARANQLAEEIESSAQYKARVALENDDRSEEEKYTAVQRNSSEREGHSINTRENKYIPPGQRNR EVISWGSGRQNSPRMGQPGSGSMPSRSTSHTSDFNPNSGSDQRVVNGGVPWPSPCPSPSSRPPSRYQSG PNSLPPRAAATPTRPPSRPPSRPSRPPSHPSAHGSPAPVSTMPKRMSSEGPPRMSPKAQRHPRNHRVSAGR GSISSGLEFVSHNPPSEAATPPVARTSPSGGTWSSVVSGVPRLSPKTHRPRSPRQNSIGNTPSGPVLASPQ AGIIPTEAVAMPIPAASPTPASPASNRAVTPSSEAKDSRI. QDQRQNSPAGNKENIKPNETSPSFSKAENKG! SPVVSEHRKQIDDLKKFKNDFRLQPSSTSESMDQLLNKNREGEKSRDLIKDKIEPSAKDSFIENSSSSNCT SGSSKPNSPSJSPSILSNTEHKRGPEVTSQGVQTSSPACKQEKDDKEEKKDAAEQVRKSTI-NPNAKEFNP RSFSQPKPSTTPTSPRPQAQPSPSMVGHQQPTPVYTQPVCFAPNMMYPVPVSPGVQPLYPIPMTPMPVN QAKTYRAVPNMPQQRQDQHHQSAMMHPASAAGPPIAATPPA. YSTQYVAYSPQQFPNQPLVQHVPHYQSQHPHVYSPVIQGNARMMAPPTHAQPGLVSSSATQYGAHEQTHAMYACPKLPYNKETSPSFYFAISTGSLAQQYAHPNATLHPHTPHPQPSATPTGQQQSQHGGSHPAPSPVQHHQHQAAQALHLASPQQQSAIYHAGLAPTPPSMTPASNTQSPQNSFPAAQQTVFTIHPSHVQPAYTNPPHMAHVPQAHVQSGMVPSHPTAHAPMMLMTTQPPGGPQAALAQSALQPIPVSTTAHFPYMTHPSVQAHHQQQL (SEQ ID NO: 62).[000111] As used herein. “ATXN2-associated neurological disease’" means a neurological disease associated with abnormal ATXN2 expression, activity, or function, CAG repeat or polyglutamine expansion in ATXN2 gene, or abnormal TDP-43 aggregation.[000112] The term “% sequence identity” or “percentage sequence identity” with respect to a reference nucleic acid sequence is defined as the percentage of nucleotides, nucleosides, or nucleobases in a candidate sequence that are identical with the nucleotides, nucleosides, or nucleobases in the reference nucleic acid sequence, after optimally aligning the sequences and introducing gaps or overhangs, if necessary, to achieve the maximum percent sequence identity. Alignment for purposes of determining percent nucleic acid sequence identity can be achieved in various ways that are within the skill in the art, for instance, using publicly available computer software programs, for example, those described in Current Protocols in Molecular Biology' (Ausubel et al., eds., 1987, Supp. 30, section 7.7.18, Table 7.7.1), and including BLAST, BLAST-2, ALIGN, Megalign (DNASTAR), Clustal W2.0 or Clustal X2.0 software. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared. Percentage of “sequence identity ” can be determined by comparing two optimally aligned sequences over a comparison window, where the fragment of the nucleic acid sequence in the comparison window may comprise additions or deletions (e.g., gaps or overhangs) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage can be calculated by determining the number of positions at which the identical nucleotide, nucleoside, or nucleobase occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison, and multiplying the result by 100 to yield the percentage of sequence identity. The output is the percent identity of the subject sequence with respect to the query sequence.[000113] The term “polypeptide” or “protein”, as used herein, refers to a polymer of amino acid residues. The term applies to polymers comprising naturally occurring amino acids and polymers comprising one or more non-naturally occurring amino acids.[000114] As used herein, “RNAi,” “RNAi agent,” “iRNA,” “iRNA agent,” or “RNA interference agent” means an agent that mediates sequence-specific degradation of a target mRNA by RNA interference, e.g., via RNA-induced silencing complex (RISC) pathway. In some embodiments, the RNAi agent has a sense strand and an antisense strand, and the sense strand and the antisense strand form a duplex (e.g., a double stranded RNA).[000115] As used herein, “strand” refers to a single, contiguous sequence of nucleotides linked together through intemucleotide linkages (e.g., phosphodiester linkages or phosphorothioate linkages). A strand can have two free ends (e.g., a 5’ end and a 3’ end).[000116] As used herein, “treatment” or “treating” refers to all processes wherein there may be a slowing, controlling, delaying, or stopping of the progression of the disorders or disease disclosed herein, or ameliorating disorder or disease symptoms, but does not necessarily indicate a total elimination of all disorder or disease symptoms. Treatment includes administration of a protein or nucleic acid or vector or composition for treatment of a disease or condition in a patient, particularly in a human.[000117] The following examples are offered to illustrate, but not to limit, the claimed inventions.EXAMPLESExample 1: Generation and Characterization of TfR binding proteinsGeneration of human TfR binding proteins[000118] Antibody against human TfR was generated by immunizing AlivaMab® transgenic mice with the extracellular domains of human Transferrin Receptor 1 protein with a His tag (hTfR-ECD-6His, SEQ ID NO: 64, see Table 6) and mouse Transferrin Receptor protein with a His tag (mTfR-ECD-6His, SEQ ID NO: 63). Antigen positive B-cells were sorted from pooled spleens. Binding of individual antibodies cloned from those B-cells to hi s-tagged hTfR-ECD was verified.[000119] Additional antibody against human TfR was generated by immunizing AlivaMab® transgenic mice with the apical domain of human Transferrin Receptor 1 protein with a His tag (hTfR-ApD-6His, SEQ ID NO: 65, see Table 6). Antigen positive B-cells were sorted from pooled spleens. Binding of individual antibodies cloned from those B-cells to his-tagged hTfR-ECD was verified.Table 6. Sequences of the immunogens used to generate human or mouse TfR antibodies.Immunogen Sequence SEQ ID NO mTIR-ECD-6His HHHHHHCKRVEQKEECVKLAETEETDKSETMETEDV 63PTSSRLYWADLKTLLSEKLNSIEFADTIKQLSQNTYTP REAGSQKDESLAYYIENQFHEFKFSKVWRDEHYVKI QVKSSIGQNMVTIVQSNGNLDPVESPEGYVAFSKPTE VSGKLVHANFGTKKDFEELSYSVNGSLVIVRAGEITF AEKVANAQSFNAIGVLIYMDKNKFPVVEADLALFGH AHLGTGDPYTPGFPSFNHTQFPPSQSSGLPNIPVQTISR AAAEKLFGKMEGSCPARWNIDSSCKLELSQNQNVKL IVKNVLKERRILNIFGVIKGYEEPDRYVVVGAQRDAL GAGVAAKSSVGTGLLLKLAQVFSDMISKDGFRPSRSII FASWTAGDFGAVGATEWLEGYLSSLHLKAFTYINLD KVVLGTSNFKVSASPLLYTLMGKIMQDVKHPVDGKS LYRDSNWISKVEKLSFDNAAYPFLAYSGIPAVSFCFCE DADYPYLGTRLDTYEALTQKVPQLNQMVRTAAEVA GQLIIKLTHDVELNLDYEMYNSKLLSFMKDLNQFKT DIRDMGLSLQWLYSARGDYFRATSRLTTDFHNAEKT NRFVMREINDRIMKVEYHFLSPYVSPRESPFRHIFWGS GSHTLSALVENLKLRQKNITAFNETLFRNQLALATWT IQGVANALSGDIWNIDNEFhTfR-ECD-6His HHHHHHCKGVEPKTECERLAGTESPVREEPGEDFPA 64ARRLYWDDLKRKLSEKLDSTDFTGTIKLLNENSYVPR EAGSQKDENLALYVENQFREFKLSKVWRDQHFVKIQ VKDSAQNSVIIVDKNGRLVYLVENPGGYVAYSKAAT VTGKLVHANFGTKKDFEDLYTPVNGSIVIVRAGKITF AEKVANAESLNAIGVLIYMDQTKFPIVNAELSFFGHA HLGTGDPYTPGFPSFNHTQFPPSRSSGLPNIPVQTISRA AAEKLFGNMEGDCPSDWKTDSTCRMVTSESKNVKL TVSNVLKEIKILNIFGVIKGFVEPDHYVVVGAQRDAW GPGAAKSGVGTALLLKLAQMFSDMVLKDGFQPSRSII FASWSAGDFGSVGATEWLEGYLSSLHLKAFTYINLD KAVLGTSNFKVSASPLLYTLIEKTMQNVKHPVTGQFL YQDSNWASKVEKLTLDNAAFPFLAYSGIPAVSFCFCE DTDYPYLGTTMDTYKELIERIPELNKVARAAAEVAG QFVIKLTHDVELNLDYERYNSQLLSFVRDLNQYRADI KEMGLSLQWLYSARGDFFRATSRLTTDFGNAEKTDR FVMKKLNDRVMRVEYHFLSPYVSPKESPFRHVFWGS GSHTLPALLENLKLRKQNNGAFNETLFRNQLALATW TIQGAANALSGDVWDIDNEFhTfR-ApD-6His HHHHHHHHGKPIPNPLLGLDSTGGGGSDSAQNSVIIV 65DKNGRLVYLVENPGGYVAYSKAATVTGKLVHANFG TKKDFEDLYTPVNGSIVIVRAGKITFAEKVANAESLN AIGVLIYMDQTKFPIVNAELSFFGHAHLGGGGGGLPN IPVQT1SRAAAEKLFGNMEGDCPSDWKTDSTCRMVTSESKNVKLTVS[000120] Affinity variants of the generated human TfR antibodies were made by systematically introducing mutations into individual CDR of each antibody and the resulting variants were subjected to multiple rounds of selection with decreasing concentrations of antigen and / or increasing periods of dissociation to isolate clones with improved affinities.The sequences of individual variants were used to construct a combinatorial library which was subjected to an additional round of selection with increased stringency to identify additive or synergistic mutational pairings between the individual CDR regions. Individual combinatorial clones are sequenced. The heavy chain and light chain CDRs and VH / VL sequences of the human TfR binding domains and proteins are provided in Table la.[000121] Human TfR binding proteins were generated by recombinant DNA technology. Such human TfR binding proteins can be expressed in a mammalian cell line such as HEK293 or CHO, either transiently or stably transfected with an expression system using an optimal predetermined HC: LC vector ratio or a single vector system encoding both HC and LC. Clarified media, into which the protein has been secreted, can be purified using the commonly used techniques.Binding affinity[000122] Binding affinity and binding stoichiometry of the exemplified human TfR binding proteins to human and cynomolgus TfR was characterized using a surface plasmon resonance assay on a Biacore 8K instrument primed with HBS-EP+ (lOmM Hepes pH 7.4 + 150mM NaCl + 3mM EDTA + 0.05% (w / v) surfactant P20) running buffer and analysis temperature set at 37 °C. Target human and cynomologus TfR ECD’s were immobilized on a CM4 chip (Cytiva P / N 29104989) using standard NHS-EDC amine coupling. The TfR binding proteins were prepared at a final concentration of 0.3, 0.1, 0.033, 0.01, 0.0033, 0.001, 0.00033, 0.0001 pM respectively by dilution of stock solution into running buffer.[000123] Binding analysis was performed in a multi-cycle kinetics manner. Each analysis cycle consists of (1) injection of the lowest to highest concentration proteins over all Fc at 50 pL / min for 140 seconds followed by return to buffer flow for 400 seconds to monitor dissociation phase; (2) regeneration of chip surfaces with injection of 3M magnesium chloride, for 30 seconds at 100 pL / min over all cells; and (3) equilibration of chip surfaces with a 50 pL (30-sec) injection of HBS-EP+. Data were processed using standard doublereferencing and fit to a 2-state binding model using Biacore 8K Evaluation software, to determine the association rate (kon, M⁻¹s⁻¹ units), dissociation rate (koff, s⁻¹ units), and Rmax(RU units). The equilibrium dissociation constant (KD) is calculated from the relationship KD = koff / kon, and is in molar units. Results are provided in Table 7.Table 7. Binding Affinity of Exemplified human TfR binding proteins to human or cynomolgus TfR at 37 °CHuman TfR Standard error Standard error of binding Human TfR KD of the mean, Cyno TfR KD the mean, Cyno proteins (Biacore, nM) Human TfR KD (Biacore, nM) TfR KD(TBP) at 37 °C (Biacore, nM) at 37 °C (Biacore, nM)n=3 n=3TBP3 32.087 11.795 66.565 11.695TBP4 153.642 7.949 300.180 2.565TBP5 0.522 0.284 502.210 8.129Example 2: Synthesis and characterization of dsRNA targeting ATXN2[000124] Single strands (sense and antisense) of the dsRNA duplexes were typically synthesized on solid support via a K& A H-8 (K& A Labs GmbH) or a similar automated oligonucleotide synthesizer. The sense strands were synthesized using an appropriate CPG such as 3'-Cholesterol-TEG CNA CPG 500 (LGC Biosearch Technologies), 3 -TEG-Tocopherol (LGC Biosearch Technologies) or phthalamido amino C6 Icaa CPG 500 A (Chemgenes) whereas the antisense strands used standard support (LGC Biosearch Technologies). The sequences of the sense and antisense strands were show n in Table 3a or 3b.[000125] Standard reagents were used in the oligo synthesis (Table 8), where 0. IM xanthane hydride in pyridine was used as the sulfurization reagent. All monomers (Table 9a) except for OMe U (0.1 M in 20% DMF / ACN) were made at 0. IM in ACN and contained a molecular sieves trap bag.[000126] The oligonucleotides were cleaved and deprotected (C / D) using AMA (1: 1 mixture of concentrated ammonia and 40% wt methylamine in water) at room temperature for 2 hours. C / D was determined complete by IP-RP LCMS when the resulting mass data confirmed the identity of sequence. Dependent on scale, the CPG w as filtered via Acrodisc® 32mm syringe filter with 0.2 pm Supor® membrane. The CPG was back washed / rinsed with RNAse free water then filtered through the same filtering device and combined with the first filtrate. This was repeated twice. The material was then divided evenly into 50 mL falcon tubes to remove organics via Genevac™.[000127] The crude oligonucleotides w ere purified via AKTA™ Pure purification system using anion-exchange (AEX). A Sepax Source 15Q column, 15 urn, 10x250 mm was used at room temperature with mobile phase A (MPA): 20mM sodium phosphate buffer, 20% ACN, pH 7.0 and mobile phase B (MPB): 20mM sodium phosphate buffer, 1.5M NaBr, 20% ACN, pH 7.0. In all cases, fractions which contained a mass purity greater than 85% without impurities >5% where combined.[000128] The purified oligonucleotides were desalted using 15 mL 3K MWCO centrifugal spin tubes at 3500xg for ~30 min. The oligonucleotides were rinsed with RNAse free water until the eluent conductivity reached < 100 usemi / cm. After desalting was complete, approximately 1 -2 mL of sample was recovered and transferred to a 5 mL Eppendorf tube. The final desalted oligonucleotides were analyzed for concentration (nano drop at A260), characterized by IP-RP LCMS for mass purity and UV-purity.[000129] For the preparation of duplexes, 1 equivalent of sense strand and 1.03-1.05 equivalents of antisense strand were combined and monitored for their annealing via UPLC (ensuring no greater than 5% of excess antisense strand was present). Further integrity of the duplex was confirmed by LCMS using IP-RP. For in vivo analysis, the appropriate amount of duplex was either lyophilized or diluted with IX PBS for rodent studies and a CSF for nonhuman primate studies.[000130] For in-vitro testing, cholesterol or tocopherol-conjugated oligonucleotides were annealed at this stage to give cholesterol or tocopherol conjugated dsRNA by mixing equimolar aliquots of sense and antisense strands at room temperature for 30 minutes. The final desalted oligonucleotides were analyzed for concentration (nano drop at A260), characterized by IP-RP LC / MS for mass purity and UPLC for UV-purity.Table 8 - Oligonucleotide Synthesis ReagentsReagentsActivator Solution (0.5M ETT in ACN)Cap A (Acetic Anhydride, 2,6-lutidine in THF, 1:1:8)Cap B (1 -Methylimidazole in THF, 16:84)Oxidation Solution (0.02M Iodine in THF / Pyridine / Water,70:20:10)Deblock Solution, 3% TCA in DCM (w / v)Acetonitrile (Anhydrosolv, Water max. 10 ppm)Xanthane Hydride (0.1 M in Pyridine)Diethylamine (20 % in Acetonitrile)Table 9a- PhosphoramiditesPhosphoramidite Abbreviation Supplier Catalog # CASDMT-2'-F-A(Bz)- fA Chemgenes ANP-9151 136834-22-5 CEPhosphoramiditeDMT-2'-F-C(Ac)- fC Chemgenes ANP-9152 159414-99-0 CEPhosphoramiditeDMT-2'-F-G(Ac)- fG Chemgenes ANP-9158 159414-99-0 CEPhosphoramiditeDMT-2'-F-U-CE fU Chemgenes ANP-9154 146954-75-8 PhosphoramiditeDMT-2'-O-Me- mA Chemgenes ANP-5751 110782-31-5 A(Bz)-CEPhosphoramiditeDMT-2'-O-Me- mC Chemgenes ANP-6756 199593-09-4 C(Ac)-CEPhosphoramiditeDMT-2'-O-Me- mG Chemgenes ANP-5763 150780-67-9 G(Ac)-CEPhosphoramiditeDMT-2'-O-Me-U- mU Chemgenes ANP-5754 110764-79-9 CEPhosphoramidite5'bis(POM) vinyl POM-VPmU Hongene PR5-032 BVPMUP23B2A1 phosphate-2'-Ome- U3'CEphosphoramiditeTable 9b. Cholesterol or Tocopherol Conjugate StructuresStructure1(Cholesterol 41 Jb conjugate) wm p T o p.. oA-0- zv - A'.g JL VX2 I,■ ■■ (Tocopherol 0. G. o.-v. n: I • ser.sa strand • Xy vconjugate);Example 3: Generation of ATXN2 RNAi agents[000131] Certain abbreviations are defined as follows: “ACN” refers to acetonitrile; “aAEX” refers to analytical anion exchange; “AS” refers to antisense strand; “DAR” refers to drug / siRNA to antibody / protein ratio; “DCM” refers to dichloromethane; “DHAA” refers to dehydroascorbic acid; “dsRNA” refers to double stranded ribonucleic acid; “DTT” refers to dithiothreitol; “h” refers to hours; “HPLC” refers to high-performance liquid chromatography; “LC / MS” refers to liquid chromatography mass spectrometry; “LTQ / MS” refers to linear ion trap mass spectrometer; “min” refers to minutes; “MSPT” refers to 4-(5-methylsulfonyl-lH-tetrazole-lyl)phenol; “MW” refers to molecular weight; “MWCO” refers to molecular weight cut-off; “NHS” refers to N-hydroxysuccinimide; “OD” refers to 4-(5-(methylsulfonyl)-l,3,4-oxadiazol-2-yl)phenol; “PBS” phosphate-buffered saline; “PEG” refers to polyethylene glycol; “RNAi” refers to RNA interference; “rpm” refers to revolutions per minute; “SEC” refers to size exclusion chromatography; “siRNA” refers to small interfering RNA; “SMCC” refers to succinimidyl-4-(N-maleimidomethyl)cyclohexane-1 -carboxylate; “SS” refers to sense strand; “TCO” refers to trans-cyclo-octene; “TfR” refers to transferrin receptor; “TEIF” refers to tetrahydrofuran; “TRIS” refers to tris(hydroxymethyl)aminomethane; “UPLC” refers to ultra performance liquid chromatography; and “UV” refers to ultraviolet.Scheme 1Step C[000132] Scheme 1, step A depicts the methylation of the thiol on compound (1) using iodomethane and a suitable base such as DIEA in a solvent such as THF to give compound (2). Step B shows an alkylation of compound (2) with tert-butyl 2-(2-(2-bromoethoxy)ethoxy)acetate using a base such as potassium carbonate in a solvent such as acetone to give compound (3). Step C shows the oxidation of compound (3) with hydrogen peroxide and ammonium molybdate (VI) tetrahydrate in a solvent such as EtOH followed by an acidic deprotection using an acid such as TFA in a solvent such as DCM to give compound (4). Note that in the case of the I H-tetrazole. the deprotection took place during the oxidation step. Step D depicts a coupling of compound (4) and l-hydroxypyrrolidine-2,5-dione using EDCI in a solvent system such as DCM and THF to give compound (5).Scheme 2o6 7 [000133] Scheme 2, step A shows the coupling of compound (6) and isoindoline-1, 3-dione using DIAD and tributyl phosphine in a solvent such as THF to give compound (7).Step B depicts the phosphorylation of compound (7) with 2-cyanoethyl-N, N-diisopropylchlorophosphoramidite using a base such as DIEA in a solvent such as DCM to give compound (8).Preparation 14-(5-(Methylthio)-lH-tetrazol-l-yl)phenol[000134] A solution of 4-(5-mercapto-lH-tetrazol-l-yl)phenol (4.00 g, 20.6 mmol) in THF (50 mL) was cooled to 0 °C. DIEA (4.31 g, 33.3 mmol) was added then stirred for 10 minutes before adding iodomethane (1.54 mL, 24.7 mmol) dropwise over a period of 1 minute. The mixture was stirred at 0 °C for 20 minutes, and then stirred at ambient temperature for 12 hours. After this time, the mixture was diluted with EtOAc (100 mL) and washed with saturated aqueous NEUCl (2 x 50 mL). The organic layer was separated, dried over sodium sulfate, and concentrated in vacuo to give the title compound (4.2 g, 93%). ES / MS m z: 209 (M+H).Preparation 2 / e / - Butyl 2-(2-(2-(4-(5-(methylthio)-lH-tetrazol-l-yl)phenoxy)ethoxy)ethoxy)acetateN'N[000135] In a pressure vessel, potassium carbonate (3.15 g, 22.8 mmol) was added to / c 7-butyl 2-(2-(2-bromoethoxy)ethoxy)acetate (4.33 g, 14.8 mmol) and 4-(5-(methylthio)-lH-tetrazol-l-yl)phenol (2.5 g, 11.4 mmol) in acetone (60 mL). The pressure vessel was sealed and heated at 80 °C for 8 hours with vigorous stirring. After this time, the mixture was cooled to ambient temperature then filtered while washing through with acetone / EtOAc / DCM (30 mL each). The filtrate was concentrated in vacuo and purified via silica gel column chromatography eluting with 0-100% EtOAc / DCM to give the title compound as a white solid (3.98 g, 85%). ES / MS m / z: 411 (M+H).Preparation 32-(2-(2-(4-(5-(Methylsulfonyl)-lH-tetrazol-l-yl)phenoxy)ethoxy)ethoxy)acetic acidO[000136] tert- Butyl 2-(2-(2-(4-(5-(methylthio)-lH-tetrazol-l-yl)phenoxy)ethoxy)ethoxy)acetate (3.98 g, 9.21 mmol) was dissolved in EtOH (100 mL) and cooled to 5-10 °C. Then, 30% hydrogen peroxide (19 mL, 184 mmol) was added, followed by ammonium molybdate (VI) tetrahydrate (1.14 g, 0.921 mmol). The mixture was allowed to warm to ambient temperature and then stirred for 4 hours, after which it was diluted with DCM (150 mL) and washed with saturated aqueous sodium chloride solution. The organic phase was separated, dried over sodium sulfate, and concentrated in vacuo. The resulting residue was purified via silica gel column chromatography eluting with 0-100% EtOAc / DCM to give the title compound as a white solid (3.00 g, 80%). ES / MS m / z: 385 (M-H).Preparation 42,5-Dioxopyrrolidin-l-yl 2-(2-(2-(4-(5-(methylsulfonyl)-lH-tetrazol-l- yl)phenoxy)ethoxy)ethoxy)acetate[000137] EDCI (1.60 g, 10.3 mmol) was added to a solution of 2-(2-(2-(4-(5-(methylsulfonyl)-lH-tetrazol-l-yl)phenoxy)ethoxy)ethoxy)acetic acid (2.80 g, 7.25 mmol) and l-hydroxypyrrolidine-2, 5-dione (1.33 g, 11.6 mmol) in DCM (50 mL) and THF (70 mL). Another 20 mL of DCM was added to bring the mixture into a solution followed by stirring at ambient temperature for 12 hours. After this time, concentrated in vacuo and purified via silica gel column chromatography eluting with 0-100% EtOAc / DCM to give the title compound (2.61g, 65%). ES / MS m / z 484 (M+H).Preparation 52-(((2R,3R,4R,5R)-5-(2,4-Dioxo-3,4-dihydropyrimidin-l(2H)-yl)-3-hydroxy-4- methoxytetrahydrofuran-2-yl)methyl)isoindoline-l, 3-dioneOI[000138] A solution of l-((2R,3R,4R,5R)-4-hydroxy-5-(hydroxymethyl)-3-methoxytetrahydrofuran-2-yl)pyrimidine-2,4(lH,3H)-dione (30 g, 120 mmol), isoindoline-1,3-dione (21 g, 140 mmol), DIAD (27 mL, 140 mmol), tributyl phosphine (36 mL, 150 mmol), and THF (300 mL) was stirred at ambient temperature for 12 h. The crude reaction was filtered, concentrated in vacuo, and purified via silica gel flash chromatography eluting with 0-100% EtOAc / hexanes to give the title compound as a white solid (6.0 g, 13%).Preparation 62-Cyanoethyl ((2R,3R,4R,5R)-5-(2.4-dioxo-3,4-dihydropyrimidin-l(2H)-yl)-2-((l,3- dioxoisoindolin-2-yl)methyl)-4-methoxytetrahydrofuran-3-yl) diisopropylphosphoramidite[000139] A solution of 2-(((2R,3R,4R,5R)-5-(2,4-dioxo-3,4-dihydropyrimidin-l(2H)-yl)-3-hydroxy-4-methoxytetrahydrofuran-2-yl)methyl)isoindoline-l, 3-dione (3.00 g, 7.74 mmol), 2-cyanoethyl-N, N-diisopropylchlorophosphoramidite (2.47 mL, 11.6 mmol), DIEA (4.05 mL, 23.2 mmol), and DCM (40 mL) was stirred at ambient temperature. After 1 hour, additional 2-cyanoethyl-N, N-diisopropylchlorophosphoramidite (0.82 mL, 3.8 mmol) was added. After 1 hour, the crude reaction was poured into a slurry of silica gel (15 g) in 30 mL of 1% TEA / DCM, concentrated in vacuo to a dry powder, and purified via silica gel flash chromatography eluting with 40-100% EtOAc / hexanes (0.5% TEA) to give the title compound as a white foam (3.70 g, 81%).1H NMR (d6-DMSO) d 11.4 (br s, 1 H), 7.96-7.78 (m, 5H), 5.83 (dd, 1H), 5.71 (dd, 1H), 4.46-3.47 (m, 9H), 3.39 (s, 1.5H), 3.35 (s, 1.5H), 2.82-2.73 (m, 2H), 1.16-0.97 (m, 12H).31P NMR (d6-DMSO) d 149.7, 149.4.[000140]Preparation 73’ Tetrazole linker-functionalized sense strand0N~N [000141] A 50 mL Falcon tube was charged with ATXN2-C6Am sense strand (15.81 mg, 5.750 mL, 2.20 pmol) and 0.863 mL 20x borate buffer. A solution of 2,5-dioxopyrrolidin-1-yl 2-(2-(2-(4-(5-(methylsulfonyl)-1H-tetrazol-1-yl)phenoxy)ethoxy)ethoxy)acetate (10.6 mg, 22.0 pmol) in ACN (5.750 mL) was then added. The mixture was vortexed for one minute, then shook at lOOOrpm at 40 °C for 30 minutes. The solution was then diluted to 60 mL using RNAse free water to bring ACN concentration to < 10%. The solution was filtered using 20 mL 3K MWCO centrifugal spin tubes at 3900 rpm for ~30 minutes. The oligonucleotide was rinsed with RNAse free water three times. The retentate was quantitatively transferred from the spin tube to a 15 mL Falcon tube by adding 1 mL of RNAse free water, vortexing, and transferring to the Falcon tube. This was repeated until complete transfer of oligo by measuring concentration of compound on filter via nanodrop. The final oligonucleotide was analyzed for concentration (nano drop at A260), characterized by IP-RP, LCMS for mass purity, and UPLC for UV-purity. The optical density measurement of the product solution (average of 2 measurements) was 82 OD / mL, 0.369 mM, 5 mL, 13.94 mg. ES / MS (m / z): 7551.31 (M+H).Preparation 85' Tetrazole linker-functionalized sense strand[000142] A 50 mL Falcon tube was charged with ATXN2-5’Am sense strand (16.63 mg, 1.150 mL, 2.39 pmol) and 0.173 mL of 20x borate buffer. A solution of 2,5-dioxopyrrolidin-1-yl 2-(2-(2-(4-(5-(methylsulfonyl)-1H-tetrazol-1-yl)phenoxy)ethoxy)ethoxy)acetate (11.6 mg, 23.9 pmol) in ACN (1.150 mL) was then added. The mixture was vortexed for one minute, then shook at lOOOrpm at 40 °C for 30 minutes. The solution was then diluted to 20 mL using RNAse free water to bring the ACN concentration to < 10%. The solution was filtered using 20 mL 3K MWCO centrifugal spin tubes at 3900 rpm for ~30 minutes. The oligonucleotide was rinsed with RNAse free water three times. The retentate was quantitatively transferred from the spin tube to a 15 mL Falcon tube by adding 1 mL of RNAse free water, vortexing, and transferring to the Falcon tube. This was repeated until complete transfer of oligo by measuring concentration of compound on filter via nanodrop. The final oligonucleotide was analyzed for concentration (nano drop at A260), characterized by IP-RP, LCMS for mass purity, and UPLC for UV-purity. The optical density measurement of the product solution (average of 2 measurements) was 91 OD / mL.0.413 mM, 4.5 mL, 13.58 mg. ES / MS (m / z): 7316.06 (M+H).Linker -ATXN 2 Duplex[000143] The nanodrop concentrations of aqueous solutions of each strand (average of 3x) were measured as SS = 268.0pM and AS = 1138pM. The sense strand and antisense strand are annealed to form a dsRNA. 10 mL of SS and 2.41 mL of AS are mixed and shook for 30 min at 20 °C. The amount of residual SS strand was measured until completion. Removed endotoxins by filtering through a 0.45 pM filter. The resulting 12.41 mL of solution measured (Nanodrop™ Lite. 3x average) 76.2 OD / mL equating to 199.6pM and a total of 37.9 mg. LTQ / MS m / z 7457,7845; UV purity 95.1%.Conjugation of dsRNA to TfR binding proteins[000144] Site-specific native or engineered cysteine amino acid residues in the TfR binding proteins were used to conjugate dsRNA. Cysteines can be engineered into theprimary amino acid sequence of the TfR binding proteins. The approach of introducing cysteines as a means for conjugation has been described in WO 2018 / 232088, which is incorporated specifically in relation to conjugation via cysteine residues. For engineered cysteine conjugation, the TfR binding proteins were first reduced with 40 molar equivalents reducing agent dithiothreitol (DTT) at 37 °C for two hours, followed by desalting to remove reducing agent via dialysis or desalting columns. This is followed by re-oxidation of the TfR binding protein to reform the structural disulfides with 10 molar equivalent dehydroascorbic acid (DHAA) incubation at ambient temperature for two hours. A follow up desalting was performed to remove oxidizing agent.[000145] Conjugation of dsRNA onto TfR binding proteins were done using the following methods.Conjugation Scheme 1[000146] The first conjugation method utilized the 3’SS tetrazole (MSPT) -functionalized dsRNA for conjugating onto the engineered cysteine of the TfR binding proteins. For this method, TfR binding protein was prepared similarly as above to make the engineered thiol available for conjugation by undergoing a reduction and oxidation process of the TfR binding proteins. This is followed by incubating the MSPT-dsRNA with the TfR binding proteins at 1.2 to 2 molar equivalents for overnight conjugation at ambient temperature.TfR binding protein conjugation with 3’ MSPT linkerConjugation Scheme 2[000147] The second conjugation method utilized the 5'SS tetrazole (MSPT) - functionalized dsRNA for conjugating onto the engineered cysteine of the TfR binding proteins. For this method, TfR binding protein was prepared similarly as above to make the engineered thiol available for conjugation by undergoing a reduction and oxidation process of the TfR binding proteins. This is followed by incubating the MSPT-dsRNA with the TfR binding proteins at 1.2 to 2 molar equivalents for overnight conjugation at ambient temperature.TfR binding protein conjugation with 5’ MSPT linkerSH[000148] Synthesis of other linkers in Table 5 and conjugation of linker-functionalized dsRNA to the engineered cysteine of the TfR binding proteins have been described in W02024 / 036096.[000149] Conjugation was monitored using analytical anion exchange chromatography. A ProPac™ SAX-10 HPLC Column, 10 pm particle, 4 mm diameter, 250 mm length was utilized with the following method. Flow rate of 1 mL / min, Buffer A: 20mM TRIS pH 7.0, Buffer B: 20 mM TRIS pH 7.0 + IM NaCl, at 30 °C.[000150] Drug / siRNA to antibody / protein ratio (DAR) was calculated based on peak area % from the analytical anion exchange (aAEX) chromatogram.[000151] Post conjugation of dsRNA to the TfR binding protein, excess dsRNA and unconjugated protein was removed by further purification. Either preparative size exclusion chromatography (SEC) or preparative anion exchange chromatography was utilized forpurification of the final conjugate. Preparative SEC was performed using Cytiva Superdex® 200 in IX PBS pH 7.2 under an isocratic condition. Alternatively, anion exchange, e g., ThermoFisher POROS™ XQ, was used with starting buffer of 20mM TRIS pH 7.0 and eluting with 20 column volume gradient with a buffer containing 20mM TRIS pH 7.0 and IM NaCl. These resulted in purified TfR binding protein-dsRNA conjugate devoid of excess dsRNA and minimal unconjugated protein. The resulting conjugate profile was analyzed by analytical anion exchange for final DAR quantitation (Table 10).Table 10. siRNA / drug to TBP / antibody ratio (DAR)Average DAR % of DAR0 % of DARI % of DAR2TBP5- 5’MSPT- 1.01 1.4 96.2 2.4dsRNA No. 10TBP5- 5’MSPT- 1.00 1.4 96.7 1.9dsRNA No. 11TBP5- 5’MSPT- 1.00 1.5 96.6 1.9dsRNA No. 12TBP5- 3’MSPT- 1.00 1.3 97.0 1.7dsRNA No. 7TBP5- 3’MSPT- 1.01 1.2 96.4 2.4dsRNA No. 8TBP5- 3’MSPT- 1.01 1.5 96.2 2.3dsRNA No. 9Example 4: In vitro characterization of ATXN2 RNAi agentsSelected ATXN2 RNAi agents were tested in SH-SY5Y cells.SH-SY5Y Cell Culture and RNAi Treatment and Analysis:[000152] SH-SY5Y cells (ATCC CRL-2266) were derived from the SK-N-SH neuroblastoma cell line (Ross, R. A., et al., 1983. J Natl Cancer Inst 71, 741-747). The base medium was composed of a 1:1 mixture of ATCC-formulated Eagle's Minimum Essential Medium, (Cat No. 30-2003), and F12 Medium. The complete growth medium was supplemented with 10% fetal bovine serum. IX amino acids, IX sodium bicarbonate, and IXpenicillin-streptomycin (Gibco) and cells incubated at 37 °C in a humidified atmosphere of 5% CO2. On Day One. SH-SY5Y cells were plated in 96 well fibronectin coated tissue culture plates and allowed to attach overnight. On Day Two, complete media was removed and replaced with RNAi agent in serum free media. Cells were incubated with RNAi agent for 72 hours, followed by media change with RNAi reagent for another 72 hours for a total of 144 hours of drug incubation before analysis of gene expression. Analysis of changes in gene expression in RNAi treated SH-SY5Y cells was measured using Cells-to-Ct Kits following the manufacturer’s protocol (ThermoFisher A35377). Predesigned gene expression assays (supplied as 20X mixtures) were selected from Applied Bio-systems (Foster City, CA, USA). The efficiencies of these assays (ThermoFisher Hs00268077_m1 ATXN2 and Hs99999903_m1 ACTB) were characterized with a dilution series of cDNA. RT-QPCR was performed in MicroAmp Optical 384-well reaction plates using QuantStudio 7 Flex system. The delta-delta CT method of normalizing to the housekeeping gene GAPDH was used to determine relative amounts of gene expression. GraphPad Prism v9.0 was used to determine IC50 with a four-parameter logistic fit.The percentage knockdown and absolute IC50 of the exemplary ATXN2 RNAi agents are shown in Table 11.Table 11: In vitro activity of ATXN2 RNAi agents in SH-SY5Y cellsATXN2 RNAi Agent No. SHSY5Y cell Absolute%KD @0.1 jiMIC50 (nM)TBP5-5'MSPT-dsRNA No. 1049.27 2.987TBP5-5’MSPT-dsRNA No. 1150.77 1.388TBP5-5’MSPT-dsRNA No. 1225.02 0.547Example 5. In vivo characterization of mouse TBP-dsRNA conjugates in human TfR transgenic mice[000153] To determine the efficacy of the TfR binding proteins-dsRNA conjugates against ATXN2, they were tested in human TfR transgenic mice in a proof-of-concept study. TBP5-5’MSPT-dsRNANo. 10, TBP5-5’MSPT-dsRNA No. 11, and TBP5-5'MSPT-dsRNANo. 12 were dosed intravenously at 10 mg / kg siRNA concentration in human TfR transgenic mice and compared to the PBS dosed control group (n=4 per group). 28 days following initial dosing, mice were euthanized under carbon monoxide and sacrificed, then brain, spinal cord, samples were collected and processed for assessment of gene expression changes using RT-qPCR. ATXN2 primer from Thermo Fisher (Mm00485946_m1) was used as the target primer and ACTB primer from Thermo Fisher (Mm02619580_g1) was used as the housekeeping gene primer. ATXN2 protein expression changes were also assessed from tissue samples.[000154] Table 12 shows ATXN2 mRNA knockdown data.Table 12. ATXN2 mRNA knockdown by ATXN2 RNAi agents%ATXN2 %ATXN2%ATXN2ATXN2 mRNA mRNAmRNARNAi remaining remainingremainingAgent (Hemi(Spinal (Cerebellum)brain) Cord )TBP5- 5’MSPT- dsRNA18 22 19No. 10TBP5- 5’MSPT- dsRNA25 36 28No. 11TBP5- 5’MSPT- dsRNA35 41 32No. 12[000155] The level of ATXN2 protein in the protein lysate was measured using an inhouse developed MSD assay. Briefly, the MSD GOLD 96-well Small Spot Streptavidin SECTOR Plate (Meso Scale Diagnostics, Rockville, Maryland) was simultaneously blocked with bovine serum albumin and coated with the capture antibody, the biotinylate-rabbitmonoclonal anti-ATXN2 antibody with shaking at room temperature for 1 hour. After washing, the wells on each plate were incubated with the protein lysate or the recombinant human ATXN2 protein in the presence of MSD Blocker A (Meso Scale Diagnostics), with shaking at room temperature for 2 hours. The plates were washed again, then incubated with the detection antibody, the biotinylated-mouse monoclonal anti-ATXN2 antibody (BD Bioscience, Clone 22, 611378) in the presence of MSD Blocker A with shaking at room temperature for 1 hour. After the incubation, the plates were washed, then added with 2X MSD Read Buffer T (Meso Scale Diagnostics). The electrochemiluminescence signal was then measured on an MSD SQ120MM plate reader within 5 minutes of addition of the MSD Read Buffer.[000156] Table 13 shows ATXN2 protein knockdown data.Table 13. ATXN2 protein knockdown by ATXN2 RNAi agents%ATXN2 %ATXN2%ATXN2protein proteinATXN2 RNAi proteinremaining remainingAgent remaining(Hemi(Spinal (Cerebellum)brain) Cord)TBP5- 5’MSPT- dsRNANo. 10 22 24 19TBP5- 5’MSPT- dsRNANo. 11 29 36 35TBP5- 5'MSPT- dsRNA No. 12 34 35 29[000157] To further evaluate the efficacy of the TfR binding proteins-dsRNA conjugates against ATXN2, TBP5-5’MSPT-dsRNANo. 10, TBP5-5 MSPT-dsRNANo. 11, and TBP5-5'MSPT-dsRNA No. 12 were dosed intravenously at 10, 2 and 0.5 mg / kg siRNA concentration in transgenic mice and compared to the PBS dosed control group (n= 3 or 4 per group). 28 and 70 days following initial dosing. Gene expression changes was evaluated as mentioned above using RT-qPCR. ATXN2 protein expression changes were also assessed from tissue samples with the aforementioned method.[000158] Table 14 shows ATXN2 mRNA knockdown data from the dose-response and durability study in human TfR transgenic mice.Table 14. ATXN2 mRNA knockdown by ATXN2 RNAi agents in dose-response and durability study% ATXN2 mRNA remaining (Hemi-brain)Dose 10 mg / kg 2 mg / kg 0.5 mg / kgDuration 28d 70d 28d 70d 28d 70dTBP5-5’MSPT- dsRNANo. 10 20 26 30 45 42 51 TBP5-5 MSPT- dsRNANo. 11 25 38 38 57 39 57% ATXN2 mRNA remaining (Cerebellum) Dose 10 mg / kg 2 mg / kg 0.5 mg / kg Duration 28d 70d 28d 70d 28d 70d TBP5-5’MSPT- dsRNANo. 10 26 34 33 53 50 80 TBP5-5’MSPT- dsRNA No. 11 40 61 47 52 56 98% ATXN2 mRNA remaining (Lumbar Spinal cord) Dose 10 mg / kg 2 mg / kg 0.5 mg / kg Duration 28d 70d 28d 70d 28d 70d TBP5-5 MSPT- dsRNANo. 10 16 23 22 33 35 53 TBP5-5 MSPT- dsRNANo. 11 23 37 29 52 47 64[000159] Table 15 shows ATXN2 protein knockdown data from the dose-response and durability study in human TfR transgenic mice.Table 15. ATXN2 protein knockdown by ATXN2 RNAi agents in dose-response and durability study% ATXN2 protein remaining i Hemi-brain) Dose 10 mg / kg 2 mg / kg 0.5 mg / kg Duration 28d 70d 28d 70d 28d 70d TBP5-5’MSPT- dsRNANo. 10 20 31 28 50 50 55 TBP5-5’MSPT- dsRNANo. 11 28 42 31 56 52 69% ATXN2 protein remaining (Cerebellum) Dose 10 mg / kg 2 mg / kg 0.5 mg / kgDuration 28d 70d 28d 70d 28d 70dTBP5-5’MSPT- dsRNANo. 10 ND* 28 ND* 45 ND* 59 TBP5-5 MSPT- dsRNA No. 11 ND* 53 ND* 59 ND* 65% ATXN2 protein remaining (Lumbar Spinal cord) Dose 10 mg / kg 2 mg / kg 0.5 mg / kg Duration 28d 70d 28d 70d 28d 70d TBP5-5’MSPT- dsRNANo. 10 12 18 27 39 43 55 TBP5-5’MSPT- dsRNA No. 11 24 42 34 51 48 61*ND: Not determined[000160] To evaluate the effect of site of conjugation on the sense strand, conjugates with TfR binding proteins conjugated to either 5’ end of the sense strand or 3’end of sense strand were assessed for knockdown efficacy in human TfR transgenic mice. TBP5-5’MSPT-dsRNANo. 10, TBP5-5'MSPT-dsRNANo. 11, TBP5-3’MSPT-dsRNA No. 10 and TBP5-3’MSPT-dsRNA No. 11 were dosed intravenously at 0.5 mg / kg siRNA concentration in TfR transgenic mice and compared to the PBS dosed control group (n= 4 or 6 per group). 28 days following initial dosing, gene expression changes were evaluated as mentioned above using RT-qPCR.[000161] Table 16 shows ATXN2 mRNA knockdown data by TBP5 conjugated to either 5’ end of the sense strand (5'MSPT) or 3’end of sense strand (3’MSPT) of dsRNA No.10 or dsRNA No. 11. The data shows the 3’ MSPT conjugates are as efficacious as the 5’ MSPT conjugates.Table 16. ATXN2 mRNA knockdown by ATXN2 RNAi agents with different TBP conjugate sites%ATXN2 mRNA remaining ATXN2 RNAi Agent(Hemi-brain)TBP5-5’MSPT-dsRNA No. 1037TBP5-5'MSPT-dsRNA No. 1147TBP5-3’MSPT-dsRNA No. 10 68TBP5-3’MSPT-dsRNA No. 1179SEQUENCE LISTINGSEQ SequenceID NO1 SYSMN2 SISSSSSYIYYADSVKG3 RHGYSNSDAFDN4 RASQGISHYLV5 AASSLQS6 LQHNSYPWT7 EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMNWVRQAPGKGLEWVSSISSSSSYIYYA DSVKGRFTISRDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSNSDAFDNWGQGTLVTVSS8 DIQMTQSPSAMSASVGDRVTITCRASQGISHYLVWFQQKPGKVPKRLIYAASSLQSGVPSRF SGSGSGTEFTLTISSLQPEDFATYYCLQHNSYPWTFGQGTKVEIK9 EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMNWVRQAPGKGLEWVSSISSSSSYIYYA DSVKGRFTISRDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSNSDAFDNWGQGTLVTVS SASTKGPSVFPLAPSSKSTSGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSG LYSLSSVVTVPSSSLGTQTYICNVNHKPSNTKVDKRVEPKC10 DIQMTQSPSAMSASVGDRVTITCRASQGISHYLVWFQQKPGKVPKRLIYAASSLQSGVPSRF SGSGSGTEFTLTISSLQPEDFATYYCLQHNSYPWTFGQGTKVEIKRTVAAPSVFIFPPSDEQL KSGTASVVCLLNNFYPREAKVQWKVDNALQSGNSQESVTEQDSKDSTYSLSSTLTLSKAD YEKHKVYACEVTHQGLSSPVTKSFNRGEC11 EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMNWVRQAPGKGLEWVSSISSSSSYIYYA DSVKGRFTISRDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSNSDAFDNWGQGTLVTVS SASTKGPCVFPLAPSSKSTSGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSG LYSLSSVVTVPSSSLGTQTYICNVNHKPSNTKVDKRVEPKCDKTHTGGGGQGGGGQGGGG QGGGGQGGGGQEVQLLESGGGLVQPGGSLRLSCAASGRYIDETAVAWFRQAPGKGREFV AGIGGGVDITYYADSVKGRFTISRDNSKNTLYLQMNSLRPEDTAVYYCGARPGRPLITSKV ADLYPYWGQGTLVTVSSPP12 GGGGQGGGGQGGGGQGGGGQ13 EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMNWVRQAPGKGLEWVSSISSSSSYIYYA DSVKGRFTISRDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSNSDAFDNWGQGTLVTVS SASTKGPXVFPLAPCSRSTSESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSG LYSLSSVVTVPSSSLGTKTYTCNVDHKPSNTKVDKRVESKYGPPCPPCPAPEAAGGPSVFLF PPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTYRVV SVLTVLHQDWLNGKEYKCKVSNKGLPSSIEKT1SKAKGQPREPQVYTLPPSQEEMTKNQVS LTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSRLTVDKSRWQEGNVFSC SVMHEALHNHYTQKSLSLSLG, wherein X is S or C.14 EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMNWVRQAPGKGLEWVSSISSSSSYIYYA DSVKGRFTISRDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSNSDAFDNWGQGTLVTVS SASTKGPSVFPLAPCSRSTSESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSG LYSLSSVVTVPSSSLGTKTYTCNVDHKPSNTKVDKRVESKYGPPCPPCPAPEAAGGPSVFLF PPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTYRVV SVLTVLHQDWLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREPQVSTLPPSQEEMTKNQVS LMCLVYGFYPSDIXVEWESNGQPENNYKTTPPVLDSDGSFFLYSVLTVDKSRWQEGNVFSC SVMHEALHNHYTQKSLSLSLG, wherein X is A or C.15 ESKYGPPCPPCPAPEAAGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVD GVEVHNAKTKPREEQFNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKGLPSSIEKTISKAK GQPREPQVYTLPPSQGDMTKNQVQLTCLVKGFYPSDIXVEWESNGQPENNYKTTPPVLDSDGSFFLASRLTVDKSRWQEGNVFSCSVMHEALHNHYTQKSLSLSLG. wherein X is A or C.EVQLVESGGGLVKPGGSLRLSCVASGFTFSSYSMNWVRQAPGKGLEWVSSISSSSSYIYYA DSVKGRFTISRDNAKNSLYLQMNSLRAEDTAVYYCARRHGYSNSDAFDNWGQGTLVTVS SASTKGPCVFPLAPCSRSTSESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSG LYSLSSVVTVPSSSLGTKTYTCNVDHKPSNTKVDKRVESKYGPPCPPCPAPEAAGGPSVFLF PPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTYRVV SVLTVLHQDWLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREPQVSTLPPSQEEMTKNQVS LMCLVYGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSVLTVDKSRWQEGNVFSC SVMHEALHNHYTQKSLSLSLG ESKYGPPCPPCPAPEAAGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVD GVEVHNAKTKPREEQFNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKGLPSSIEKTISKAK GQPREPQVYTLPPSQGDMTKNQVQLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSD GSFFLASRLTVDKSRWQEGNVFSCSVMHEALHNHYTQKSLSLSLG QVQLVQSGAEVKKPGSSVKVSCKASGYTFSSYAIEWVRQAPGQGLEWMGGILPGSGTINY NEKFKGRVTITADKSTSTAYMELSSLRSEDTAVYYCARMSSNSDQGFDLWGQGTLVTVSS ASTKGPXVFPLAPCSRSTSESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGL YSLSSVVTVPSSSLGTKTYTCNVDHKPSNTKVDKRVESKYGPPCPPCPAPEAAGGPSVFLFP PKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTYRVVS VLTVLHQDWLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREPQVYTLPPSQEEMTKNQVSL TCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFLLYSKLTVDKSRWQEGNVFSCS VMHEALHNHYTQKSLSLSLG, wherein X is S or C.D1QMTQSPSSLSASVGDRVTITCKASQGISRFLSWFQQKPGKAPKSL1YAVSSLVDGVPSRFS GSGSGTDFTLTISSLQPEDFATYYCVQYNSYPYGFGGGTKVEIKRTVAAPSVFIFPPSDEQLK SGTASVVCLLNNFYPREAKVQWKVDNALQSGNSQESVTEQDSKDSTYSLSSTLTLSKADYE KHKVYACEVTHQGLSSPVTKSFNRGEC ETAVA GIGGGVDITYYADSVKG RPGRPLITSKVADLYPY EVQLLESGGGLVQPGGSLRLSCAASGRYIDETAVAWFRQAPGKGREFVAGIGGGVDITYYA DSVKGRFTISRDNSKNTLYLQMNSLRPEDTAVYYCGARPGRPLITSKVADLYPYWGQGTLV TVSSPP GSYWIC CIYSTSGGRTYYASWVKG GDDSISDAYFDL QSSQSVYNNNRLA DASTLAS QGTYFSSGWSWA QSLEESGGDLVKPEGSLTLTCTASGFSFSGSYWICWVRQAPGKGLEWIGCIYSTSGGRTYYA SWVKGRFT1SKTSSTTVTLQMTSLTAADTATYFCARGDDS1SDAYFDLWGPGTLVTVSS ALDMTQTASPVSAAVGGTVTINCQSSQSVYNNNRLAWYQQKPGQPPKLLIYDASTLASGV PSRFKGSGSGTQFTLTISGVQSDDSATYYCQGTYFSSGWSWAFGGGTEVVVK QSLEESGGDLVKPEGSLTLTCTASGFSFSGSYWICWVRQAPGKGLEWIGCIYSTSGGRTYYA SWVKGRFTISKTSSTTVTLQMTSLTAADTATYFCARGDDSISDAYFDLWGPGTLVTVSSAST KGPCVFPLAPCSRSTSESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYSL SSVVTVPSSSLGTKTYTCNVDHKPSNTKVDKRVESKYGPPCPPCPAPEAAGGPSVFLFPPKP KDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTYRVVSVLT VLHQDWLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREPQVYTLPPSQEEMTKNQVSLTCL VKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSRLTVDKSRWQEGNVFSCSVM HEALHNHYTQKSLSLSLG ALDMTQTASPVSAAVGGTVTINCQSSQSVYNNNRLAWYQQKPGQPPKLLIYDASTLASGV PSRFKGSGSGTQFTLTISGVQSDDSATYYCQGTYFSSGWSWAFGGGTEVVVKRTVAAPSVF IFPPSDEQLKSGTASVVCLLNNFYPREAKVQWKVDNALQSGNSQESVTEQDSKDSTYSLSSTLTLSKADYEKHKVYACEVTHQGLSSPVTKSFNRGECESKYGPPCPPCPAPEAAGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVD GVEVHNAKTKPREEQFNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKGLPSSIEKTISKAK GQPREPQVYTLPPSQEEMTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSD GSFLLYSKLTVDKSRWQEGNVFSCSVMHEALHNHYTQKSLSLSLG GAAUCUAUGGAUCAACUACUA UAGUAGUUGAUCCAUAGAUUCAG AAGAAUGAUUUUAGGUUACAA UUGUAACCUAAAAUCAUUCUUAA AGGAUGGUUCAUAUACUUACA UGUAAGUAUAUGAACCAUCCUCA UAAUCUAUGGAUCAACUACUA UAGAAUGAUUUUAGGUUACAA UGGAUGGUUCAUAUACUUACAmG*mA*mAmUmCmUmAmUfGfGfAmUmCmAmAmCmUmAmC*mU*mA mU*fA*mGmUfAmGmUfUmGmAmUmCmCfAmUfAmGmAmUmUmC*mA*mG mA*mA*mGmAmAmUmGmAfUfUfUmUmAmGmGmUmUmAmC*mA*mA mU*fU*mGmUf AmAmCfCmUm Am Am AmAfUmCf AmUmUmCmUmU*m A*m A mA*mG*mGmAmUmGmGmUfUfGfAmUmAmUmAmCmUmUmA*mC*mA mU*fG*mUmAfAmGmUfAmUmAmUmGmAfAmCfCmAmUmCmCmU*mC*mA [NH2mU]*mA*mAmUmCmUmAmUfGfGfAmUmCmAmAmCmUmAmC*mU*mA[NH2mU]*mA*mGmAmAmUmGmAfUfUfUmUmAmGmGmUmUmAmC*mA*mA[NH2mU]*mG*mGmAmUmGmGmUfUfCfAmUmAmUmAmCmUmUmA*mC*mA1 AGAGCTCGCC TCCCTCCGCC TCAGACTGTT TTGGTAGCAA CGGCAACGGC GGCGGCGCGT 61 TTCGGCCCGG CTCCCGGCGG CTCCTTGGTC TCGGCGGGCC TCCCCGCCCC TTCGTCGTCC 121 TCCTTCTCCC CCTCGCCAGC CCGGGCGCCC CTCCGGCCGC GCCAACCCGC GCCTCCCCGC 181 TCGGCGCCCG CGCGTCCCCG CCGCGTTCCG GCGTCTCCTT GGCGCGCCCG GCTCCCGGCT 241 GTCCCCGCCC GGCGTGCGAG CCGGTGTATG GGCCCCTCAC CATGTCGCTG AAGCCCCAGC 301 AGCAGCAGCA GCAGCAGCAG CAGCAGCAGC AGCAGCAACA GCAGCAGCAG CAGCAGCAGC 361 AGCAGCCGCC GCCCGCGGCT GCCAATGTCC GCAAGCCCGG CGGCAGCGGC CTTCTAGCGT 421 CGCCCGCCGC CGCGCCTTCG CCGTCCTCGT CCTCGGTCTC CTCGTCCTCG GCCACGGCTC 481 CCTCCTCGGT GGTCGCGGCG ACCTCCGGCG GCGGGAGGCC CGGCCTGGGC AGAGGTCGAA 541 ACAGTAACAA AGGACTGCCT CAGTCTACGA TTTCTTTTGA TGGAATCTAT GCAAATATGA 601 GGATGGTTCA TATACTTACA TCAGTTGTTG GCTCCAAATG TGAAGTACAA GTGAAAAATG 661 GAGGTATATA TGAAGGAGTT TTTAAAACTT ACAGTCCGAA GTGTGATTTG GTACTTGATG 721 CCGCACATGA GAAAAGTACA GAATCCAGTT CGGGGCCGAA ACGTGAAGAA ATAATGGAGA 781 GTATTTTGTT CAAATGTTCA GACTTTGTTG TGGTACAGTT TAAAGATATG GACTCCAGTT 841 ATGCAAAAAG AGATGCTTTT ACTGACTCTG CTATCAGTGC TAAAGTGAAT GGCGAACACA 901 AAGAGAAGGA CCTGGAGCCC TGGGATGCAG GTGAACTCAC AGCCAATGAG GAACTTGAGG 961 CTTTGGAAAA TGACGTATCT AATGGATGGG ATCCCAATGA TATGTTTCGA TATAATGAAG 1021 AAAATTATGG TGTAGTGTCT ACGTATGATA GCAGTTTATC TTCGTATACA GTGCCCTTAG 1081 AAAGAGATAA CTCAGAAGAA TTTTTAAAAC GGGAAGCAAG GGCAAACCAG TTAGCAGAAG 1141 AAATTGAGTC AAGTGCCCAG TACAAAGCTC GAGTGGCCCT GGAAAATGAT GATAGGAGTG 1201 AGGAAGAAAA ATACACAGCA GTTCAGAGAA ATTCCAGTGA ACGTGAGGGG CACAGCATAA 1261 ACACTAGGGA AAATAAATAT ATTCCTCCTG GACAAAGAAA TAGAGAAGTC ATATCCTGGG 1321 GAAGTGGGAG ACAGAATTCA CCGCGTATGG GCCAGCCTGG ATCGGGCTCC ATGCCATCAA 1381 GATCCACTTC TCACACTTCA GATTTCAACC CGAATTCTGG TTCAGACCAA AGAGTAGTTA 1441 ATGGAGGTGT TCCCTGGCCA TCGCCTTGCC CATCTCCTTC CTCTCGCCCA CCTTCTCGCT 1501 ACCAGTCAGG TCCCAACTCT CTTCCACCTC GGGCAGCCAC CCCTACACGG CCGCCCTCCA 1561 GGCCCCCCTC GCGGCCATCC AGACCCCCGT CTCACCCCTC TGCTCATGGT TCTCCAGCTC 1621 CTGTCTCTAC TATGCCTAAA CGCATGTCTT CAGAAGGGCC TCCAAGGATG TCCCCAAAGG 1681 CCCAGCGACA TCCTCGAAAT CACAGAGTTT CTGCTGGGAG GGGTTCCATA TCCAGTGGCC 1741 TAGAATTTGT ATCCCACAAC CCACCCAGTG AAGCAGCTAC TCCTCCAGTA GCAAGGACCA 1801 GTCCCTCGGG GGGAACGTGG TCATCAGTGG TCAGTGGGGT TCCAAGATTA TCCCCTAAAA1861 CTCATAGACC CAGGTCTCCC AGACAGAACA GTATTGGAAA TACCCCCAGT GGGCCAGTTC1921 TTGCTTCTCC CCAAGCTGGT ATTATTCCAA CTGAAGCTGT TGCCATGCCT ATTCCAGCTG 1981 CATCTCCTAC GCCTGCTAGT CCTGCATCGA ACAGAGCTGT TACCCCTTCT AGTGAGGCTA 2041 AAGATTCCAG GCTTCAAGAT CAGAGGCAGA ACTCTCCTGC AGGGAATAAA GAAAATATTA 2101 AACCCAATGA AACATCACCT AGCTTCTCAA AAGCTGAAAA CAAAGGTATA TCACCAGTTG 2161 TTTCTGAACA TAGAAAACAG ATTGATGATT TAAAGAAATT TAAGAATGAT TTTAGGTTAC 2221 AGCCAAGTTC TACTTCTGAA TCTATGGATC AACTACTAAA CAAAAATAGA GAGGGAGAAA 2281 AATCAAGAGA TTTGATCAAA GACAAAATTG AACCAAGTGC TAAGGATTCT TTCATTGAAA 2341 ATAGCAGCAG CAACTGTACC AGTGGCAGCA GCAAGCCGAA TAGCCCCAGC ATTTCCCCTT 2401 CAATACTTAG TAACACGGAG CACAAGAGGG GACCTGAGGT CACTTCCCAA GGGGTTCAGA 2461 CTTCCAGCCC AGCATGTAAA CAAGAGAAAG ACGATAAGGA AGAGAAGAAA GACGCAGCTG 2521 AGCAAGTTAG GAAATCAACA TTGAATCCCA ATGCAAAGGA GTTCAACCCA CGTTCCTTCT 2581 CTCAGCCAAA GCCTTCTACT ACCCCAACTT CACCTCGGCC TCAAGCACAA CCTAGCCCAT 2641 CTATGGTGGG TCATCAACAG CCAACTCCAG TTTATACTCA GCCTGTTTGT TTTGCACCAA 2701 ATATGATGTA TCCAGTCCCA GTGAGCCCAG GCGTGCAACC TTTATACCCA ATACCTATGA 2761 CGCCCATGCC AGTGAATCAA GCCAAGACAT ATAGAGCAGT ACCAAATATG CCCCAACAGC 2821 GGCAAGACCA GCATCATCAG AGTGCCATGA TGCACCCAGC GTCAGCAGCG GGCCCACCGA 2881 TTGCAGCCAC CCCACCAGCT TACTCCACGC AATATGTTGC CTACAGTCCT CAGCAGTTCC 2941 CAAATCAGCC CCTTGTTCAG CATGTGCCAC ATTATCAGTC TCAGCATCCT CATGTCTATA 3001 GTCCTGTAAT ACAGGGTAAT GCTAGAATGA TGGCACCACC AACACACGCC CAGCCTGGTT 3061 TAGTATCTTC TTCAGCAACT CAGTACGGGG CTCATGAGCA GACGCATGCG ATGTATGCAT 3121 GTCCCAAATT ACCATACAAC AAGGAGACAA GCCCTTCTTT CTACTTTGCC ATTTCCACGG 3181 GCTCCCTTGC TCAGCAGTAT GCGCACCCTA ACGCTACCCT GCACCCACAT ACTCCACACC 3241 CTCAGCCTTC AGCTACCCCC ACTGGACAGC AGCAAAGCCA ACATGGTGGA AGTCATCCTG 3301 CACCCAGTCC TGTTCAGCAC CATCAGCACC AGGCCGCCCA GGCTCTCCAT CTGGCCAGTC 3361 CACAGCAGCA GTCAGCCATT TACCACGCGG GGCTTGCGCC AACTCCACCC TCCATGACAC 3421 CTGCCTCCAA CACGCAGTCG CCACAGAATA GTTTCCCAGC AGCACAACAG ACTGTCTTTA 3481 CGATCCATCC TTCTCACGTT CAGCCGGCGT ATACCAACCC ACCCCACATG GCCCACGTAC 3541 CTCAGGCTCA TGTACAGTCA GGAATGGTTC CTTCTCATCC AACTGCCCAT GCGCCAATGA 3601 TGCTAATGAC GACACAGCCA CCCGGCGGTC CCCAGGCCGC CCTCGCTCAA AGTGCACTAC 3661 AGCCCATTCC AGTCTCGACA ACAGCGCATT TCCCCTATAT GACGCACCCT TCAGTACAAG 3721 CCCACCACCA ACAGCAGTTG TAAGGCTGCC CTGGAGGAAC CGAAAGGCCA AATTCCCTCC 3781 TCCCTTCTAC TGCTTCTACC AACTGGAAGC ACAGAAAACT AGAATTTCAT TTATTTTGTT 3841 TTTAAAATAT ATATGTTGAT TTCTTGTAAC ATCCAATAGG AATGCTAACA GTTCACTTGC 3901 AGTGGAAGAT ACTTGGACCG AGTAGAGGCA TTTAGGAACT TGGGGGCTAT TCCATAATTC 3961 CATATGCTGT TTCAGAGTCC CGCAGGTACC CCAGCTCTGC TTGCCGAAAC TGGAAGTTAT 4021 TTATTTTTTA ATAACCCTTG AAAGTCATGA ACACATCAGC TAGCAAAAGA AGTAACAAGA 4081 GTGATTCTTG CTGCTATTAC TGCTAAAAAA AAAAAAAAAA AAAAATCAAG ACTTGGAACG 4141 CCCTTTTACT AAACTTGACA AAGTTTCAGT AAATTCTTAC CGTCAAACTG ACGGATTATT 4201 ATTTATAAAT CAAGTTTGAT GAGGTGATCA CTGTCTACAG TGGTTCAACT TTTAAGTTAA 4261 GGGAAAAACT TTTACTTTGT AGATAATATA AAATAAAAAC TTAMAAAAA TTTAAAAAAT 4321 AAAAAAAGTT TTAAAAACTG A1 MSLKPQQQQQ QQQQQQQQQQ QQQQQQQQPP PAAANVRKPG GSGLLASPAA APSPSSSSVS 61 SSSATAPSSV VAATSGGGRP GLGRGRNSNK GLPQSTISFD GIYANMRMVH ILTSVVGSKC 121 EVQVKNGGIY EGVFKTYSPK CDLVLDAAHE KSTESSSGPK REEIMESILF KCSDFVWQF 181 KDMDSSYAKR DAFTDSAISA KVNGEHKEKD LEPWDAGELT ANEELEALEN DVSNGWDPND 241 MFRYNEENYG WSTYDSSLS SYTVPLERDN SEEFLKREAR ANQLAEEIES SAQYKARVAL 301 ENDDRSEEEK YTAVQRNSSE REGHSINTRE NKYIPPGQRN REVISWGSGR QNSPRMGQPG 361 SGSMPSRSTS HTSDFNPNSG SDQRWNGGV PWPSPCPSPS SRPPSRYQSG PNSLPPRAAT 421 PTRPPSRPPS RPSRPPSHPS AHGSPAPVST MPKRMSSEGP PRMSPKAQRH PRNHRVSAGR 481 GSISSGLEFV SHNPPSEAAT PPVARTSPSG GTWSSWSGV PRLSPKTHRP RSPRQNSIGN 541 TPSGPVLASP QAGIIPTEAV AMPIPAASPT PASPASNRAV TPSSEAKDSR LQDQRQNSPA 601 GNKENIKPNE TSPSFSKAEN KGISPVVSEH RKQIDDLKKF KNDFRLQPSS TSESMDQLLN 661 KNREGEKSRD LIKDKIEPSA KDSFIENSSS NCTSGSSKPN SPSISPSILS NTEHKRGPEV 721 TSQGVQTSSP ACKQEKDDKE EKKDAAEQVR KSTLNPNAKE FNPRSFSQPK PSTTPTSPRP 781 QAQPSPSMVG HQQPTPVYTQ PVCFAPNMMY PVPVSPGVQP LYPIPMTPMP VNQAKTYRAV 841 PNMPQQRQDQ HHQSAMMHPA SAAGPPIAAT PPAYSTQYVA YSPQQFPNQP LVQHVPHYQS 901 QHPHVYSPVI QGNARMMAPP THAQPGLVSS SATQYGAHEQ THAMYACPKL PYNKETSPSF 961 YFAISTGSLA QQYAHPNATL HPHTPHPQPS ATPTGQQQSQ HGGSHPAPSP VQHHQHQAAQ 1021 ALHLASPQQQ SAIYHAGLAP TPPSMTPASN TQSPQNSFPA AQQTVFTIHP SHVQPAYTNP 1081 PHMAHVPQAH VQSGMVPSHP TAHAPMMLMT TQPPGGPQAA LAQSALQPIP VSTTAHFPYM1141 THPSVQAHHQ QQL1 CCCGAGAAAG CAACCCAGCG CGCCGCCCGC TCCTCACGTG TCCCTCCCGG CCCCGGGGCC 61 ACCTCACGTT CTGCTTCCGT CTGACCCCTC CGACTTCCGA GGTCGAAACA GTAACAAAGG 121 ACTGCCTCAG TCTACGATTT CTTTTGATGG AATCTATGCA AATATGAGGA TGGTTCATAT 181 ACTTACATCA GTTGTTGGCT CCAAATGTGA AGTACAAGTG AAAAATGGAG GTATATATGA 241 AGGAGTTTTT AAAACTTACA GTCCGAAGTG TGATTTGGTA CTTGATGCCG CACATGAGAA 301 AAGTACAGAA TCCAGTTCGG GGCCGAAACG TGAAGAAATA ATGGAGAGTA TTTTGTTCAA 361 ATGTTCAGAC TTTGTTGTGG TACAGTTTAA AGATATGGAC TCCAGTTATG CAAAAAGAGA 421 TGCTTTTACT GACTCTGCTA TCAGTGCTAA AGTGAATGGC GAACACAAAG AGAAGGACCT 481 GGAGCCCTGG GATGCAGGTG AACTCACAGC CAATGAGGAA CTTGAGGCTT TGGAAAATGA 541 CGTATCTAAT GGATGGGATC CCAATGATAT GTTTCGATAT AATGAAGAAA ATTATGGTGT 601 AGTGTCTACG TATGATAGCA GTTTATCTTC GTATACAGTG CCCTTAGAAA GAGATAACTC 661 AGAAGAATTT TTAAAACGGG AAGCAAGGGC AAACCAGTTA GCAGAAGAAA TTGAGTCAAG 721 TGCCCAGTAC AAAGCTCGAG TGGCCCTGGA AAATGATGAT AGGAGTGAGG AAGAAAAATA 781 CACAGCAGTT CAGAGAAATT CCAGTGAACG TGAGGGGCAC AGCATAAACA CTAGGGAAAA 841 TAAATATATT CCTCCTGGAC AAAGAAATAG AGAAGTCATA TCCTGGGGAA GTGGGAGACA 901 GAATTCACCG CGTATGGGCC AGCCTGGATC GGGCTCCATG CCATCAAGAT CCACTTCTCA 961 CACTTCAGAT TTCAACCCGA ATTCTGGTTC AGACCAAAGA GTAGTTAATG GAGGTGTTCC 1021 CTGGCCATCG CCTTGCCCAT CTCCTTCCTC TCGCCCACCT TCTCGCTACC AGTCAGGTCC 1081 CAACTCTCTT CCACCTCGGG CAGCCACCCC TACACGGCCG CCCTCCAGGC CCCCCTCGCG 1141 GCCATCCAGA CCCCCGTCTC ACCCCTCTGC TCATGGTTCT CCAGCTCCTG TCTCTACTAT 1201 GCCTAAACGC ATGTCTTCAG AAGGGCCTCC AAGGATGTCC CCAAAGGCCC AGCGACATCC 1261 TCGAAATCAC AGAGTTTCTG CTGGGAGGGG TTCCATATCC AGTGGCCTAG AATTTGTATC 1321 CCACAACCCA CCCAGTGAAG CAGCTACTCC TCCAGTAGCA AGGACCAGTC CCTCGGGGGG 1381 AACGTGGTCA TCAGTGGTCA GTGGGGTTCC AAGATTATCC CCTAAAACTC ATAGACCCAG 1441 GTCTCCCAGA CAGAACAGTA TTGGAAATAC CCCCAGTGGG CCAGTTCTTG CTTCTCCCCA 1501 AGCTGGTATT ATTCCAACTG AAGCTGTTGC CATGCCTATT CCAGCTGCAT CTCCTACGCC 1561 TGCTAGTCCT GCATCGAACA GAGCTGTTAC CCCTTCTAGT GAGGCTAAAG ATTCCAGGCT 1621 TCAAGATCAG AGGCAGAACT CTCCTGCAGG GAATAAAGAA AATATTAAAC CCAATGAAAC 1681 ATCACCTAGC TTCTCAAAAG CTGAAAACAA AGGTATATCA CCAGTTGTTT CTGAACATAG 1741 AAAACAGATT GATGATTTAA AGAAATTTAA GAATGATTTT AGGTTACAGC CAAGTTCTAC 1801 TTCTGAATCT ATGGATCAAC TACTAAACAA AAATAGAGAG GGAGAAAAAT CAAGAGATTT 1861 GATCAAAGAC AAAATTGAAC CAAGTGCTAA GGATTCTTTC ATTGAAAATA GCAGCAGCAA 1921 CTGTACCAGT GGCAGCAGCA AGCCGAATAG CCCCAGCATT TCCCCTTCAA TACTTAGTAA 1981 CACGGAGCAC AAGAGGGGAC CTGAGGTCAC TTCCCAAGGG GTTCAGACTT CCAGCCCAGC 2041 ATGTAAACAA GAGAAAGACG ATAAGGAAGA GAAGAAAGAC GCAGCTGAGC AAGTTAGGAA 2101 ATCAACATTG AATCCCAATG CAAAGGAGTT CAACCCACGT TCCTTCTCTC AGCCAAAGCC 2161 TTCTACTACC CCAACTTCAC CTCGGCCTCA AGCACAACCT AGCCCATCTA TGGTGGGTCA 2221 TCAACAGCCA ACTCCAGTTT ATACTCAGCC TGTTTGTTTT GCACCAAATA TGATGTATCC 2281 AGTCCCAGTG AGCCCAGGCG TGCAACCTTT ATACCCAATA CCTATGACGC CCATGCCAGT 2341 GAATCAAGCC AAGACATATA GAGCAGTACC AAATATGCCC CAACAGCGGC AAGACCAGCA 2401 TCATCAGAGT GCCATGATGC ACCCAGCGTC AGCAGCGGGC CCACCGATTG CAGCCACCCC2461 ACCAGCTTAC TCCACGCAAT ATGTTGCCTA CAGTCCTCAG CAGTTCCCAA ATCAGCCCCT 2521 TGTTCAGCAT GTGCCACATT ATCAGTCTCA GCATCCTCAT GTCTATAGTC CTGTAATACA 2581 GGGTAATGCT AGAATGATGG CACCACCAAC ACACGCCCAG CCTGGTTTAG TATCTTCTTC 2641 AGCAACTCAG TACGGGGCTC ATGAGCAGAC GCATGCGATG TATGCATGTC CCAAATTACC 2701 ATACAACAAG GAGACAAGCC CTTCTTTCTA CTTTGCCATT TCCACGGGCT CCCTTGCTCA 2761 GCAGTATGCG CACCCTAACG CTACCCTGCA CCCACATACT CCACACCCTC AGCCTTCAGC 2821 TACCCCCACT GGACAGCAGC AAAGCCAACA TGGTGGAAGT CATCCTGCAC CCAGTCCTGT 2881 TCAGCACCAT CAGCACCAGG CCGCCCAGGC TCTCCATCTG GCCAGTCCAC AGCAGCAGTC 2941 AGCCATTTAC CACGCGGGGC TTGCGCCAAC TCCACCCTCC ATGACACCTG CCTCCAACAC 3001 GCAGTCGCCA CAGAATAGTT TCCCAGCAGC ACAACAGACT GTCTTTACGA TCCATCCTTC 3061 TCACGTTCAG CCGGCGTATA CCAACCCACC CCACATGGCC CACGTACCTC AGTGCGCCAG 33121 TGAGGCTCTG GCAAGGTGTG GGCTAGAGAT GCGACTCAGT TGGATCTATC TCTCAGAAGG 3181 CTACCTTGCT CATGTACAGT CAGGAATGGT TCCTTCTCAT CCAACTGCCC ATGCGCCAAT 3241 GATGCTAATG ACGACACAGC CACCCGGCGG TCCCCAGGCC GCCCTCGCTC AAAGTGCACT 3301 ACAGCCCATT CCAGTCTCGA CAACAGCGCA TTTCCCCTAT ATGACGCACC CTTCAGTACA 3361 AGCCCACCAC CAACAGCAGT TGTAAGGCTG CCCTGGAGGA ACCGAAAGGC CAAATTCCCT 3421 CCTCCCTTCT ACTGCTTCTA CCAACTGGAA GCACAGAAAA CTAGAATTTC ATTTATTTTG 3481 TTTTTAAAAT ATATATGTTG ATTTCTTGTA ACATCCAATA GGAATGCTAA CAGTTCACTT 3541 GCAGTGGAAG ATACTTGGAC CGAGTAGAGG CATTTAGGAA CTTGGGGGCT ATTCCATAAT 3601 TCCATATGCT GTTTCAGAGT CCCGCAGGTA CCCCAGCTCT GCTTGCCGAA ACTGGAAGTT 3661 ATTTATTTTT TAATAACCCT TGAAAGTCAT GAACACATCA GCTAGCAAAA GAAGTAACAA 3721 GAGTGATTCT TGCTGCTATT ACTGCTAAAA AAAAAAAAAA AAAAAAATCA AGACTTGGAA 3781 CGCCCTTTTA CTAAACTTGA CAAAGTTTCA GTAAATTCTT ACCGTCAAAC TGACGGATTA 3841 TTATTTATAA ATCAAGTTTG ATGAGGTGAT CACTGTCTAC AGTGGTTCAA CTTTTAAGTT 3901 AAGGGAAAAA CTTTTACTTT GTAGATAATA TAAAATAAAA ACTTAAAAAA AATTTAAAAA 3961 ATAAAAAAAG TTTTAAAAAC TGAAAAAAAA AAA1 MRMVHILTSV VGSKCEVQVK NGGIYEGVFK TYSPKCDLVL DAAHEKSTES SSGPKREEIM 61 ESILFKCSDF VWQFKDMDS SYAKRDAFTD SAISAKVNGE HKEKDLEPWD AGELTANEEL 121 EALENDVSNG WDPNDMFRYN EENYGVVSTY DSSLSSYTVP LERDNSEEFL KREARANQLA 181 EEIESSAQYK ARVALENDDR SEEEKYTAVQ RNSSEREGHS INTRENKYIP PGQRNREVIS 241 WGSGRQNSPR MGQPGSGSMP SRSTSHTSDF NPNSGSDQRV VNGGVPWPSP CPSPSSRPPS 301 RYQSGPNSLP PRAATPTRPP SRPPSRPSRP PSHPSAHGSP APVSTMPKRM SSEGPPRMSP 361 KAQRHPRNHR VSAGRGSISS GLEFVSHNPP SEAATPPVAR TSPSGGTWSS WSGVPRLSP 421 KTHRPRSPRQ NSIGNTPSGP VLASPQAGII PTEAVAMPIP AASPTPASPA SNRAVTPSSE 481 AKDSRLQDQR QNSPAGNKEN IKPNETSPSF SKAENKGISP WSEHRKQID DLKKFKNDFR 541 LQPSSTSESM DQLLNKNREG EKSRDLIKDK IEPSAKDSFI ENSSSNCTSG SSKPNSPSIS 601 PSILSNTEHK RGPEVTSQGV QTSSPACKQE KDDKEEKKDA AEQVRKSTLN PNAKEFNPRS 661 FSQPKPSTTP TSPRPQAQPS PSMVGHQQPT PVYTQPVCFA PNMMYPVPVS PGVQPLYPIP 721 MTPMPVNQAK TYRAVPNMPQ QRQDQHHQSA MMHPASAAGP PIAATPPAYS TQYVAYSPQQ 781 FPNQPLVQHV PHYQSQHPHV YSPVIQGNAR MMAPPTHAQP GLVSSSATQY GAHEQTHAMY 841 ACPKLPYNKE TSPSFYFAIS TGSLAQQYAH PNATLHPHTP HPQPSATPTG QQQSQHGGSH 901 PAPSPVQHHQ HQAAQALHLA SPQQQSAIYH AGLAPTPPSM TPASNTQSPQ NSFPAAQQTV 961 FTIHPSHVQP AYTNPPHMAH VPQCASEALA RCGLEMRLSW IYLSEGYLAH VQSGMVPSHP 1021 TAHAPMMLMT TQPPGGPQAA LAQSALQPIP VSTTAHFPYM THPSVQAHHQ QQL 1 CCCGAGAAAG CAACCCAGCG CGCCGCCCGC TCCTCACGTG TCCCTCCCGG CCCCGGGGCC 61 ACCTCACGTT CTGCTTCCGT CTGACCCCTC CGACTTCCGA TTTCTTTTGA TGGAATCTAT 121 GCAAATATGA GGATGGTTCA TATACTTACA TCAGTTGTTT GTGATTTGGT ACTTGATGCC181 GCA CATGAGA AAAGTACAGA ATCCAGTTCG GGGCCGAAAC GTGAAGAAAT AATGGAGAGT 241 ATTTTGTTCA AATGTTCAGA CTTTGTTGTG GTACAGTTTA AAGATATGGA CTCCAGTTAT 301 GCAAAAAGAG ATGCTTTTAC TGACTCTGCT ATCAGTGCTA AAGTGAATGG CGAACACAAA 361 GAGAAGGACC TGGAGCCCTG GGATGCAGGT GAACTCACAG CCAATGAGGA ACTTGAGGCT 421 TTGGAAAATG ACGTATCTAA TGGATGGGAT CCCAATGATA TGTTTCGATA TAATGAAGAA 481 AATTATGGTG TAGTGTCTAC GTATGATAGC AGTTTATCTT CGTATACAGT GCCCTTAGAA 541 AGAGATAACT CAGAAGAATT TTTAAAACGG GAAGCAAGGG CAAACCAGTT AGCAGAAGAA 601 ATTGAGTCAA GTGCCCAGTA CAAAGCTCGA GTGGCCCTGG AAAATGATGA TAGGAGTGAG 661 GAAGAAAAAT ACACAGCAGT TCAGAGAAAT TCCAGTGAAC GTGAGGGGCA CAGCATAAAC 721 ACTAGGGAAA ATAAATATAT TCCTCCTGGA CAAAGAAATA GAGAAGTCAT ATCCTGGGGA 781 AGTGGGAGAC AGAATTCACC GCGTATGGGC CAGCCTGGAT CGGGCTCCAT GCCATCAAGA 841 TCCACTTCTC ACACTTCAGA TTTCAACCCG AATTCTGGTT CAGACCAAAG AGTAGTTAAT 901 GGAGGTGTTC CCTGGCCATC GCCTTGCCCA TCTCCTTCCT CTCGCCCACC TTCTCGCTAC 961 CAGTCAGGTC CCAACTCTCT TCCACCTCGG GCAGCCACCC CTACACGGCC GCCCTCCAGG 1021 CCCCCCTCGC GGCCATCCAG ACCCCCGTCT CACCCCTCTG CTCATGGTTC TCCAGCTCCT 1081 GTCTCTACTA TGCCTAAACG CATGTCTTCA GAAGGGCCTC CAAGGATGTC CCCAAAGGCC 1141 CAGCGACATC CTCGAAATCA CAGAGTTTCT GCTGGGAGGG GTTCCATATC CAGTGGCCTA 1201 GAATTTGTAT CCCACAACCC ACCCAGTGAA GCAGCTACTC CTCCAGTAGC AAGGACCAGT 1261 CCCTCGGGGG GAACGTGGTC ATCAGTGGTC AGTGGGGTTC CAAGATTATC CCCTAAAACT 1321 CATAGACCCA GGTCTCCCAG ACAGAACAGT ATTGGAAATA CCCCCAGTGG GCCAGTTCTT 1381 GCTTCTCCCC AAGCTGGTAT TATTCCAACT GAAGCTGTTG CCATGCCTAT TCCAGCTGCA 1441 TCTCCTACGC CTGCTAGTCC TGCATCGAAC AGAGCTGTTA CCCCTTCTAG TGAGGCTAAA 1501 GATTCCAGGC TTCAAGATCA GAGGCAGAAC TCTCCTGCAG GGAATAAAGA AAATATTAAA 1561 CCCAATGAAA CATCACCTAG CTTCTCAAAA GCTGAAAACA AAGGTATATC ACCAGTTGTT 1621 TCTGAACATA GAAAACAGAT TGATGATTTA AAGAAATTTA AGAATGATTT TAGGTTACAG 1681 CCAAGTTCTA CTTCTGAATC TATGGATCAA CTACTAAACA AAAATAGAGA GGGAGAAAAA 1741 TCAAGAGATT TGATCAAAGA CAAAATTGAA CCAAGTGCTA AGGATTCTTT CATTGAAAAT 1801 AGCAGCAGCA ACTGTACCAG TGGCAGCAGC AAGCCGAATA GCCCCAGCAT TTCCCCTTCA 1861 ATACTTAGTA ACACGGAGCA CAAGAGGGGA CCTGAGGTCA CTTCCCAAGG GGTTCAGACT 1921 TCCAGCCCAG CATGTAAACA AGAGAAAGAC GATAAGGAAG AGAAGAAAGA CGCAGCTGAG 1981 CAAGTTAGGA AATCAACATT GAATCCCAAT GCAAAGGAGT TCAACCCACG TTCCTTCTCT 2041 CAGCCAAAGC CTTCTACTAC CCCAACTTCA CCTCGGCCTC AAGCACAACC TAGCCCATCT 2101 ATGGTGGGTC ATCAACAGCC AACTCCAGTT TATACTCAGC CTGTTTGTTT TGCACCAAAT 2161 ATGATGTATC CAGTCCCAGT GAGCCCAGGC GTGCAACCTT TATACCCAAT ACCTATGACG 2221 CCCATGCCAG TGAATCAAGC CAAGACATAT AGAGCAGTAC CAAATATGCC CCAACAGCGG 2281 CAAGACCAGC ATCATCAGAG TGCCATGATG CACCCAGCGT CAGCAGCGGG CCCACCGATT 2341 GCAGCCACCC CACCAGCTTA CTCCACGCAA TATGTTGCCT ACAGTCCTCA GCAGTTCCCA 2401 AATCAGCCCC TTGTTCAGCA TGTGCCACAT TATCAGTCTC AGCATCCTCA TGTCTATAGT 2461 CCTGTAATAC AGGGTAATGC TAGAATGATG GCACCACCAA CACACGCCCA GCCTGGTTTA 2521 GTATCTTCTT CAGCAACTCA GTACGGGGCT CATGAGCAGA CGCATGCGAT GTATGTTTCC 2581 ACGGGCTCCC TTGCTCAGCA GTATGCGCAC CCTAACGCTA CCCTGCACCC ACATACTCCA2641 CACCCTCAGC CTTCAGCTAC CCCCACTGGA CAGCAGCAAA GCCAACATGG TGGAAGTCAT 2701 CCTGCACCCA GTCCTGTTCA GCACCATCAG CACCAGGCCG CCCAGGCTCT CCATCTGGCC 2761 AGTCCACAGC AGCAGTCAGC CATTTACCAC GCGGGGCTTG CGCCAACTCC ACCCTCCATG 2821 ACACCTGCCT CCAACACGCA GTCGCCACAG AATAGTTTCC CAGCAGCACA ACAGACTGTC 2881 TTTACGATCC ATCCTTCTCA CGTTCAGCCG GCGTATACCA ACCCACCCCA CATGGCCCAC 2941 GTACCTCAGG CTCATGTACA GTCAGGAATG GTTCCTTCTC ATCCAACTGC CCATGCGCCA 3001 ATGATGCTAA TGACGACACA GCCACCCGGC GGTCCCCAGG CCGCCCTCGC TCAAAGTGCA 3061 CTACAGCCCA TTCCAGTCTC GACAACAGCG CATTTCCCCT ATATGACGCA CCCTTCAGTA 3121 CAAGCCCACC ACCAACAGCA GTTGTAAGGC TGCCCTGGAG GAACCGAAAG GCCAAATTCC 3181 CTCCTCCCTT CTACTGCTTC TACCAACTGG AAGCACAGAA AACTAGAATT TCATTTATTT 3241 TGTTTTTAAA ATATATATGT TGATTTCTTG TAACATCCAA TAGGAATGCT AACAGTTCAC 3301 TTGCAGTGGA AGATACTTGG ACCGAGTAGA GGCATTTAGG AACTTGGGGG CTATTCCATA 3361 ATTCCATATG CTGTTTCAGA GTCCCGCAGG TACCCCAGCT CTGCTTGCCG AAACTGGAAG 3421 TTATTTATTT TTTAATAACC CTTGAAAGTC ATGAACACAT CAGCTAGCAA AAGAAGTAAC 3481 AAGAGTGATT CTTGCTGCTA TTACTGCTAA AAAAAAAAAA AAAAAAAAAT CAAGACTTGG 3541 AACGCCCTTT TACTAAACTT GACAAAGTTT CAGTAAATTC TTACCGTCAA ACTGACGGAT 3601 TATTATTTAT AAATCAAGTT TGATGAGGTG ATCACTGTCT ACAGTGGTTC AACTTTTAAG 3661 TTAAGGGAAA AACTTTTACT TTGTAGATAA TATAAAATAA AAACTTAAAA AAAATTTAAA 3721 AAATAAAAAA AGTTTTAAAA ACTGAAAAAA AAAAA1 MRMVHILTSV VCDLVLDAAH EKSTESSSGP KREEIMESIL FKCSDFVVVQ FKDMDSSYAK 61 RDAFTDSAIS AKVNGEHKEK DLEPWDAGEL TANEELEALE NDVSNGWDPN DMFRYNEENY 121 GVVSTYDSSL SSYTVPLERD NSEEFLKREA RANQLAEEIE SSAQYKARVA LENDDRSEEE 181 KYTAVQRNSS EREGHSINTR ENKYIPPGQR NREVISWGSG RQNSPRMGQP GSGSMPSRST 241 SHTSDFNPNS GSDQRVVNGG VPWPSPCPSP SSRPPSRYQS GPNSLPPRAA TPTRPPSRPP 301 SRPSRPPSHP SAHGSPAPVS TMPKRMSSEG PPRMSPKAQR HPRNHRVSAG RGSISSGLEF 361 VSHNPPSEAA TPPVARTSPS GGTWSSVVSG VPRLSPKTHR PRSPRQNSIG NTPSGPVLAS 421 PQAGIIPTEA VAMPIPAASP TPASPASNRA VTPSSEAKDS RLQDQRQNSP AGNKENIKPN 481 ETSPSFSKAE NKGISPVVSE HRKQIDDLKK FKNDFRLQPS STSESMDQLL NKNREGEKSR 541 DLIKDKIEPS AKDSFIENSS SNCTSGSSKP NSPSISPSIL SNTEHKRGPE VTSQGVQTSS 601 PACKQEKDDK EEKKDAAEQV RKSTLNPNAK EFNPRSFSQP KPSTTPTSPR PQAQPSPSMV 661 GHQQPTPVYT QPVCFAPNMM YPVPVSPGVQ PLYPIPMTPM PVNQAKTYRA VPNMPQQRQD 721 QHHQSAMMHP ASAAGPPIAA TPPAYSTQYV AYSPQQFPNQ PLVQHVPHYQ SQHPHVYSPV 781 IQGNARMMAP PTHAQPGLVS SSATQYGAHE QTHAMYVSTG SI. AQQYAf IPN ATLHPHTPHP 841 QPSATPTGQQ QSQHGGSHPA PSPVQHHQHQ AAQALHLASP QQQSAIYHAG LAPTPPSMTP 901 ASNTQSPQNS FPAAQQTVFT IHPSHVQPAY TNPPHMAHVP QAHVQSGMVP SHPTAHAPMM 961 LMTTQPPGGP QAALAQSALQ PIPVSTTAHF PYMTHPSVQA HHQQQL1 AGAGCTCGCC TCCCTCCGCC TCAGACTGTT TTGGTAGCAA CGGCAACGGC GGCGGCGCGT 61 TTCGGCCCGG CTCCCGGCGG CTCCTTGGTC TCGGCGGGCC TCCCCGCCCC TTCGTCGTCC 121 TCCTTCTCCC CCTCGCCAGC CCGGGCGCCC CTCCGGCCGC GCCAACCCGC GCCTCCCCGC 181 TCGGCGCCCG CGCGTCCCCG CCGCGTTCCG GCGTCTCCTT GGCGCGCCCG GCTCCCGGCT 241 GTCCCCGCCC GGCGTGCGAG CCGGTGTATG GGCCCCTCAC CATGTCGCTG AAGCCCCAGC 301 AGCAGCAGCA GCAGCAGCAG CAGCAGCAGC AGCAGCAACA GCAGCAGCAGCAGCAGCAGC361 AGCAGCCGCC GCCCGCGGCT GCCAATGTCC GCAAGCCCGG CGGCAGCGGC CTTCTAGCGT421 CGCCCGCCGC CGCGCCTTCG CCGTCCTCGT CCTCGGTCTC CTCGTCCTCG GCCACGGCTC 481 CCTCCTCGGT GGTCGCGGCG ACCTCCGGCG GCGGGAGGCC CGGCCTGGGC AGAGGTCGAA541 ACAGTAACAA AGGACTGCCT CAGTCTACGA TTTCTTTTGA TGGAATCTAT GCAAATATGA 601 GGATGGTTCA TATACTTACA TCAGTTGTTG GCTCCAAATG TGAAGTACAA GTGAAAAATG 661 GAGGTATATA TGAAGGAGTT TTTAAAACTT ACAGTCCGAA GTGTGATTTG GTACTTGATG 721 CCGCACATGA GAAAAGTACA GAATCCAGTT CGGGGCCGAA ACGTGAAGAA ATAATGGAGA781 GTATTTTGTT CAAATGTTCA GACTTTGTTG TGGTACAGTT TAAAGATATG GACTCCAGTT 841 ATGCAAAAAG AGATGCTTTT ACTGACTCTG CTATCAGTGC TAAAGTGAAT GGCGAACACA 901 AAGAGAAGGA CCTGGAGCCC TGGGATGCAG GTGAACTCAC AGCCAATGAG GAACTTGAGG961 CTTTGGAAAA TGACGTATCT AATGGATGGG ATCCCAATGA TATGTTTCGA TATAATGAAG 1021 AAAATTATGG TGTAGTGTCT ACGTATGATA GCAGTTTATC TTCGTATACA GTGCCCTTAG 1081 AAAGAGATAA CTCAGAAGAA TTTTTAAAAC GGGAAGCAAG GGCAAACCAG TTAGCAGAAG1141 AAATTGAGTC AAGTGCCCAG TACAAAGCTC GAGTGGCCCT GGAAAATGAT GATAGGAGTG1201 AGGAAGAAAA ATACACAGCA GTTCAGAGAA ATTCCAGTGA ACGTGAGGGG CACAGCATAA1261 ACACTAGGGA AAATAAATAT ATTCCTCCTG GACAAAGAAA TAGAGAAGTC ATATCCTGGG1321 GAAGTGGGAG ACAGAATTCA CCGCGTATGG GCCAGCCTGG ATCGGGCTCC ATGCCATCAA1381 GATCCACTTC TCACACTTCA GATTTCAACC CGAATTCTGG TTCAGACCAA AGAGTAGTTA 1441 ATGGAGGTGT TCCCTGGCCA TCGCCTTGCC CATCTCCTTC CTCTCGCCCA CCTTCTCGCT 1501 ACCAGTCAGG TCCCAACTCT CTTCCACCTC GGGCAGCCAC CCCTACACGG CCGCCCTCCA1561 GGCCCCCCTC GCGGCCATCC AGACCCCCGT CTCACCCCTC TGCTCATGGT TCTCCAGCTC 1621 CTGTCTCTAC TATGCCTAAA CGCATGTCTT CAGAAGGGCC TCCAAGGATG TCCCCAAAGG1681 CCCAGCGACA TCCTCGAAAT CACAGAGTTT CTGCTGGGAG GGGTTCCATA TCCAGTGGCC1741 TAGAATTTGT ATCCCACAAC CCACCCAGTG AAGCAGCTAC TCCTCCAGTA GCAAGGACCA1801 GTCCCTCGGG GGGAACGTGG TCATCAGTGG TCAGTGGGGT TCCAAGATTA TCCCCTAAAA1861 CTCATAGACC CAGGTCTCCC AGACAGAACA GTATTGGAAA TACCCCCAGT GGGCCAGTTC1921 TTGCTTCTCC CCAAGCTGGT ATTATTCCAA CTGAAGCTGT TGCCATGCCT ATTCCAGCTG 1981 CATCTCCTAC GCCTGCTAGT CCTGCATCGA ACAGAGCTGT TACCCCTTCT AGTGAGGCTA 2041 AAGATTCCAG GCTTCAAGAT CAGAGGCAGA ACTCTCCTGC AGGGAATAAA GAAAATATTA2101 AACCCAATGA AACATCACCT AGCTTCTCAA AAGCTGAAAA CAAAGGTATA TCACCAGTTG2161 TTTCTGAACA TAGAAAACAG ATTGATGATT TAAAGAAATT TAAGAATGAT TTTAGGTTAC2221 AGCCAAGTTC TACTTCTGAA TCTATGGATC AACTACTAAA CAAAAATAGA GAGGGAGAAA2281 AATCAAGAGA TTTGATCAAA GACAAAATTG AACCAAGTGC TAAGGATTCT TTCATTGAAA2341 ATAGCAGCAG CAACTGTACC AGTGGCAGCA GCAAGCCGAA TAGCCCCAGC ATTTCCCCTT2401 CAATACTTAG TAACACGGAG CACAAGAGGG GACCTGAGGT CACTTCCCAA GGGGTTCAGA2461 CTTCCAGCCC AGCATGTAAA CAAGAGAAAG ACGATAAGGA AG AG A AG AAAGACGCAGCTG2521 AGCAAGTTAG GAAATCAACA TTGAATCCCA ATGCAAAGGA GTTCAACCCA CGTTCCTTCT2581 CTCAGCCAAA GCCTTCTACT ACCCCAACTT CACCTCGGCC TCAAGCACAA CCTAGCCCAT 2641 CTATGGTGGG TCATCAACAG CCAACTCCAG TTTATACTCA GCCTGTTTGT TTTGCACCAA 2701 ATATGATGTA TCCAGTCCCA GTGAGCCCAG GCGTGCAACC TTTATACCCA ATACCTATGA 2761 CGCCCATGCC AGTGAATCAA GCCAAGACAT ATAGAGCAGG TAAAGTACCA AATATGCCCC2821 AACAGCGGCA AGACCAGCAT CATCAGAGTG CCATGATGCA CCCAGCGTCA GCAGCGGGCC2881 CACCGATTGC AGCCACCCCA CCAGCTTACT CCACGCAATA TGTTGCCTAC AGTCCTCAGC 2941 AGTTCCCAAA TCAGCCCCTT GTTCAGCATG TGCCACATTA TCAGTCTCAG CATCCTCATG 3001 TCTATAGTCC TGTAATACAG GGTAATGCTA GAATGATGGC ACCACCAACA CACGCCCAGC3061 CTGGTTTAGT ATCTTCTTCA GCAACTCAGT ACGGGGCTCA TGAGCAGACG CATGCGATGT 3121 ATGCATGTCC CAAATTACCA TACAACAAGG AGACAAGCCC TTCTTTCTAC TTTGCCATTT 3181 CCACGGGCTC CCTTGCTCAG CAGTATGCGC ACCCTAACGC TACCCTGCAC CCACATACTC 3241 CACACCCTCA GCCTTCAGCT ACCCCCACTG GACAGCAGCA AAGCCAACAT GGTGGAAGTC3301 ATCCTGCACC CAGTCCTGTT CAGCACCATC AGCACCAGGC CGCCCAGGCT CTCCATCTGG 3361 CCAGTCCACA GCAGCAGTCA GCCATTTACC ACGCGGGGCT TGCGCCAACT CCACCCTCCA3421 TGACACCTGC CTCCAACACG CAGTCGCCAC AGAATAGTTT CCCAGCAGCA CAACAGACTG3481 TCTTTACGAT CCATCCTTCT CACGTTCAGC CGGCGTATAC CAACCCACCC CACATGGCCC 3541 ACGTACCTCA GGCTCATGTA CAGTCAGGAA TGGTTCCTTC TCATCCAACT GCCCATGCGC 3601 CAATGATGCT AATGACGACA CAGCCACCCG GCGGTCCCCA GGCCGCCCTC GCTCAAAGTG3661 CACTACAGCC CATTCCAGTC TCGACAACAG CGCATTTCCC CTATATGACG CACCCTTCAG 3721 TACAAGCCCA CCACCAACAG CAGTTGTAAG GCTGCCCTGG AGGAACCGAA AGGCCAAATT3781 CCCTCCTCCC TTCTACTGCT TCTACCAACT GGAAGCACAG AAAACTAGAA TTTCATTTAT 3841 TTTGTTTTTA AAATATATAT GTTGATTTCT TGTAACATCC AATAGGAATG CTAACAGTTC 3901 ACTTGCAGTG GAAGATACTT GGACCGAGTA GAGGCATTTA GGAACTTGGG GGCTATTCCA3961 TAATTCCATA TGCTGTTTCA GAGTCCCGCA GGTACCCCAG CTCTGCTTGC CGAAACTGGA 4021 AGTTATTTAT TTTTTAATAA CCCTTGAAAG TCATGAACAC ATCAGCTAGC AAAAGAAGTA 4081 ACAAGAGTGA TTCTTGCTGC TATTACTGCT AAAAAAAAAA AAAAAAAAAA ATCAAGACTT4141 GGAACGCCCT TTTACTAAAC TTGACAAAGT TTCAGTAAAT TCTTACCGTC AAACTGACGG 4201 ATTATTATTT ATAAATCAAG TTTGATGAGG TGATCACTGT CTACAGTGGT TCAACTTTTA 4261 AGTTAAGGGA AAAACTTTTA CTTTGTAGAT AATATAAAAT AAAAACTTAA AAAAAATTTA4321 AAAAATAAAA AAAGTTTTAA AAACTGA1 MSLKPQQQQQ QQQQQQQQQQ QQQQQQQQPP PAAANVRKPG GSGLLASPAA APSPSSSSVS 61 SSSATAPSSV VAATSGGGRP GLGRGRNSNK GLPQSTISFD GIYANMRMVH ILTSVVGSKC 121 EVQVKNGGIY EGVFKTYSPK CDLVLDAAHE KSTESSSGPK REEIMESILF KCSDFVVVQF 181 KDMDSSYAKR DAFTDSAISA KVNGEHKEKD LEPWDAGELT ANEELEALEN DVSNGWDPND 241 MFRYNEENYG VVSTYDSSLS SYTVPLERDN SEEFLKREAR ANQLAEEIES SAQYKARVAL 301 ENDDRSEEEK YTAVQRNSSE REGHSINTRE NKYIPPGQRN REVISWGSGR QNSPRMGQPG 361 SGSMPSRSTS HTSDFNPNSG SDQRVVNGGV PWPSPCPSPS SRPPSRYQSG PNSEPPRAAT 421 PTRPPSRPPS RPSRPPSHPS AHGSPAPVST MPKRMSSEGP PRMSPKAQRH PRNHRVSAGR 481 GSISSGLEFV SHNPPSEAAT PPVARTSPSG GTWSSVVSGV PRLSPKTHRP RSPRQNSIGN 541 TPSGPVLASP QAG11P1EAV AMP1PAASPT PASPASNRAV TPSSEAKDSR LQE)QRQNSPA 601 GNKENIKPNE TSPSFSKAEN KGISPVVSEH RKQIDDLKKF KNDFRLQPSS TSESMDQLLN 661 KNREGEKSRD LIKDKIEPSA KDSFIENSSS NCTSGSSKPN SPSISPSILS NTEHKRGPEV 721 TSQGVQTSSP ACKQEKDDKE EKKDAAEQVR KSTLNPNAKE FNPRSFSQPK PSTTPTSPRP 781 QAQPSPSMVG HQQPTPVYTQ PVCFAPNMMY PVPVSPGVQP LYPIPMTPMP VNQAKTYRAG 841 KVPNMPQQRQ DQHHQSAMMH PASAAGPPIA ATPPAYSTQY VAYSPQQFPN QPLVQHVPHY 901 QSQHPHVYSP VIQGNARMMA PPTHAQPGLV SSSATQYGAH EQTHAMYACP KLPYNKETSP961 SFYFAISTGS LAQQYAHPNA TLHPHTPHPQ PSATPTGQQQ SQHGGSHPAP SPVQHHQHQA1021 AQALHLASPQ QQSAIYHAGL APTPPSMTPA SNTQSPQNSF PAAQQTVFTI HPSHVQPAYT 1081 NPPHMAHVPQ AHVQSGMVPS HPTAHAPMML MTTQPPGGPQ AALAQSALQP IPVSTTAHFP 1141 YMTHPSVQAH HQQQL ATGTCGCTGAAGCCCCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAACAGCAG CAGCAGCAGCAGCAGCAGCAGCCGCCGCCCGCGGCTGCCAATGTCCGCAAGCCCGGCGGCAGC GGCCTTCTAGCGTCGCCCGCCGCCGCGCCTTCGCCGTCCTCGTCCTCGGTCTCCTCGTCCTCGGCC ACGGC rCCCTCCICGGl GGTCGCGGCG ACCTCCGGCGGCGGG AGGCCCGGCC I GGGCAG AGGTC GAAACAGTAACAAAGGACTGCCTCAGTCTACGATTTCTTTTGATGGAATCTATGCAAATATGAGG ATGGTTCATATACTTACATCAGTTGTTGGCTCCAAATGTGAAGTACAAGTGAAAAATGGAGGTAT AT ATG AAGG AGTTTTT AAAACTTAC AG TCCGAAG TGTG Al IT GG TACTTG ATGCCGC AC ATGAG A AAAGTACAGAATCCAGTTCGGGGCCGAzAACGTGAAGzVsATAATGGAGAGTATTTTGTTCAAATG TTCAGACTTTGTTGTGGTACAGTTTAAAGATATGGACTCCAGTTATGCAAAAAGAGATGCTTTTA CTGACTCTGCTATCAGTGCTAAAGTGAATGGCGAACACAAAGAGAAGGACCTGGAGCCCTGGGA TGCAGGTGAACTCACAGCCAATGAGGAACTTGAGGCTTTGGAAAATGACGTATCTAATGGATGG GATCCCAATGATATGTTTCGATATAATGAAGAAAATTzATGGTGTAGTGTCTzACGlATGATAGCAG TTTATCTTCGTATACAGTGCCCTTAGAAAGAGATAACTCAGAAGAATTTTTAAAACGGGAAGCA AGGGCAAACCAGTTAGCAGAAGAAATTGAGTCAAGTGCCCAGTACAAAGCTCGAGTGGCCCTGG AAA ATG ATG AT AGO AG TGAGGAAG A AAAATACACAGC AGTTC AG AG A AATTCC AG TG AACGT G AGGGGCACAGCATAAACACTAGGGAAAATAAATATATTCCTCCTGGACA / ViGAAATAGAGAAG TCATATCCTGGGGAAGTGGGAGACAGAATTCACCGCGTATGGGCCAGCCTGGATCGGGCTCCAT GCCATCAAGATCCACTTCTCACACTTCAGATTTCAACCCGAATTCTGGTTCAGACCAAAGAGTAG TTAATGGAGGTGTTCCCTGGCCATCGCCTTGCCCATCTCCTTCCTCTCGCCCACCTTCTCGCTACC AGTCAGGTCCCAACTCTCTTCCACCTCGGGCAGCCACCCCTACACGGCCGCCCTCCAGGCCCCCC TCGCGGCCATCCAGACCCCCGTCTCACCCCTCTGCTCATGGTTCTCCAGCTCCTGTCTCTACTATG CCTAAACGCATGTCTTCAGAAGGGCCTCCAAGGATGTCCCCAAAGGCCCAGCGACATCCTCGAA ATCACAGAGTTTCTGCTGGGAGGGGTTCCATATCCAGTGGCCTAGAATTTGTATCCCACAACCCA CCCAGTGAAGCAGCTACTCCTCCAGTAGCAAGGACCAGTCCCTCGGGGGGAACGTGGTCATCAG TGGTCAGTGGGGTTCCAAGATTATCCCCTAAAACTCATAGACCCAGGTCTCCCAGACAGAACAG TA'T'Tl’KlAAATACXXlCC^AGTClGCKXlAG’IlTlI'l'GCl'T'lTll'CCXXZAAGClT'GTll'Ar’rAl'l'CXlAAClT'GAACl CTGTTG CC ATGCCTAT TCC AG CTG CATCTCCTACG CCTGCTAGTCCTGC ATCG AAC AGAGCTGTTA CCCCTTCTAGTGAGGCTAAAGATTCCAGGCTTCAAGATCAGAGGCAGAACTCTCCTGCAGGGAA TAAAGAAAATATTAAACCCAATGAAACATCACCTAGCTTCTCAAAAGCTGAAAACAAAGGTATA TCACCAGTTGTTTCTGAACATAGAAAACAGATTGATGATTTAAAGAAATTTAAGAATGATTTTAG GTTACAGCCAAGTTCTACTTCTGAATCTATGGATCAACTACTAAACAAAAATAGAGAGGGAGAA AAATCAAGAGATTTGATCAAAGACAAAATTGAACCAAGTGCTAAGGATTCTTTCATTGAAAATA GCAGCAGCAA CIG TACCAGTGGCAGCAGCAAGCCGAATAGCCCCAGCATTTCCCC1TCAA TACT T AGTAACACGGAGC ACAAGAGGGG ACCTGAGGTC ACTTCCC AAGGGGTTCAGACTTCC AG CCC A GCATGTAAACAAGAGAAAGACGATAAGGAAGAGAAGAAAGACGCAGCTGAGCAAGTTAGGAA ATCAACATTGAATCCCAATGCAAAGGAGTTCAACCCACGTTCCTTCTCTCAGCCAAAGCCTTCTA CTACCCCAACTTCACCTCGGCCTCAAGCACAACCTAGCCCATCTATGGTGGGTCATCAACAGCCA ACTCCAGTTTATACTCAGCCTGTTTGTTTTGCACCAAATATGATGTATCCAGTCCCAGTGAGCCCA GGCGTGCAACCTTTATACCCAATACCTATGACGCCCATGCCAGTGAATCAAGCCAAGACATATA GAGCAGTACCAAATATGCCCCAACAGCGGCAAGACCAGCATCATCAGAGTGCCATGATGCACCC AGCGTCAGCAGCGGGCCCACCGATTGCAGCCACCCCACCAGCTTACTCCACGCAATATGTTGCCT ACAGTCCTCAGCAGTTCCCAAATCAGCCCCTTGTTCAGCATGTGCCACATTATCAGTCTCAGCAT CCTCATGTCTATAGTCCTGTAATACAGGGTAATGCTAGAATGATGGCACCACCAACACACGCCCA GCCTGGTTTAGTATCTTCTTCAGCAACTCAGTACGGGGCTCATGAGCAGACGCATGCGATGTATG CATGTCCCAAATTACCATACAACyXAGGAGAC / XAGCCCT'rCTTTCTACTTTGCCATTTCCACGGGC TCCCTTGCTCAGCAGTATGCGCACCCTAACGCTACCCTGCACCCACATACTCCACACCCTCAGCC TTCAGCTACCCCCACTGGACAGCAGCAAAGCCAACATGGTGGAAGTCATCCTGCACCCAGTCCT GT T'CAGC ACC ATC AG C ACCAGGCCGCCC AG GCTCTCC ATCTGGCCAGTCC AC AGC AG C AGTC AG CCATTTACCACGCGGGGCTTGCGCCAACTCCACCCTCCATGACACCTGCCTCCAACACGCAGTCG CCACAGAATAGTTTCCCAGCAGCACAACAGACTGTCTTTACGATCCATCC1TCTCACGTTCAGCC GGCGTATACCAACCCACCCCACATGGCCCACGTACCTCAGGCTCATGTACAGTCAGGAATGGTTC CTTCTCATCCAACTGCCCATGCGCCAATGATGCTAATGACGACACAGCCACCCGGCGGTCCCCAG GCCGCCCTCGCTCAAAGTGCACTACAGCCCATTCCAGTCTCGAC / AACAGCGCATTTCCCCTATAT GACGCACCCTTCAGTACAAGCCCACCACCAACAGCAGTTGTAA MSLKPQQQQQQQQQQQQQQQQQQQQQQQPPPAAANVRKPGGSGLLASPAAAPSPSSSSVSSSSATA PSSVVAATSGGGRPGLGRGRNSNKGLPQSTISFIX3IYANMRMVHILTSVVGSKCEVQVKNGGIYEGVFKTYSPKCDLVLDAAHEKSTESSSGPKREEIMESILFKCSDFWVQFKDMDSSYAKRbAFTDSAISAKVN!GEHKEKDL. EPWDAGELTANEELEAI, ENDVSN(}WI)PN'DMFRYN'EENYGVVSTYDSSLSSYTVPLE RDNSEEFLKREARANQLAEEIESSAQYKARVALENDDRSEEEKn'AVQRNSSEREGHSIN'rRENKYlP PGQRNREMSWGSGRQNSPRMGQPGSGSMPSRSTSHTSDFNPNSGSDQRVVNGGVPVTSPCPSPSSRP PSRYQSGPNSLPPRAATPTRPPSRPPSRPSRPPSHPSAHGSPAPVSTMPKRMSSEGPPRMSPKAQRHPR NHRVSAGRGSISSGLEFVSI-INPPSEAATPPVARTSPSGGTWSSVVSGVPRLSPKTHRPRSPRQNSIGNT PSGPVLASPQAGIIPTEAVAMP[PAASPTPASPASNRAVTPSSEAKDSRLQDQRQNSPAGNKENIKPNE TSPSFSKAENkGlSPVVSEHRKQroDLKKFKNDFRLQPSSTSESMDQLLNKNREGEKSRDLIKDKlEPS AKDSHENSSSNCTSGSSKPNSPSISPSILSNTEHKRGPEVTSQGVQTSSPACKQEKDDKEEKKDAAEQ VRKSTLNPNAKEFNPRSFSQPKPSTTPTSPRPQAQPSPSMVGHQQPTPVYTQPVCFAPNMMYPVPVSP GVQPLYPIPWPWVNQAKTYRAVPNMPQQRQDQHFIQSAJ'vtMHPASAAGPPIAATPPAYSTQYVAY SPQQFPNQPLVQHVPHYQSQHPHVYSPVIQGNARMMAPPTHAQPGLVSSSATQYGAHEQTHAMYAC PKLPYNKETSPSFYFAISTGSLAQQYAHPNATLHPHTPHPQPSATPTGQQQSQHGGSHPAPSPVQHHQ HQAAQALHLASPQQQSAIYHAGLAPTPPSMTPASNTQSPQNSFPAAQQTVFTIHPSHVQPAYTNPPHM AIIVPQ^IVQSGMVPSlIPTMIAPAfl' / fLMTTQPPGGPQykALAQSALQPIPVST'IAIIFPYMriIPSVQAHH QQQL HHHHHHCKRVEQKEEC VKL AETEETDK SETMETED VPT SSRL YWADLKTLL SEKLN SIEF ADTIKQL SQNTYTPREAGSQKDESLAYYIENQFHEFKFSKVWRDEHYVKIQVKSSIGQNMVTIVQSNGNLDPVE SPEGYVAFSKPTEVSGKLVHANFGTKKDFEELSYSVNGSLVIVRAGEITFAEKVANAQSFNAIGVLIY MDKNKFPVVEADLALFGHAHLGTGDPYTPGFPSFNHTQFPPSQSSGLPNIPVQTISRAAAEKLFGKME GSCPARWNIDSSCKLELSQNQNVKLIVKNVLKERRILNIFGVIKGYEEPDRYVVVGAQRDALGAGVA AKSS VGTGLLLKLAQ VF SDMI SKDGFRPSRSIIF AS WT AGDFG A VG ATE WLEG YL S SLHLK AFT YINL DKVVLGTSNFKVSASPLLYTLMGKIMQDVKIIPVDGKSLYRDSNWISKVEKLSFDNAAYPFLAYSGIP AVSFCFCEDADYPYLGTRLDTYEALTQKVPQLNQMVRTAAEVAGQLIIKLTHDVELNLDYEMYNSK LLSFMKDLNQFKTDIRDMGLSLQWLYSARGDYFRATSRLTTDFHNAEKTNRFVMREINDRIMKVEY HFLSPYVSPRESPFRHIFWGSGSHTLSALVENLKLRQKNITAFNETLFRNQLALATWTIQGVANALSG DIWNIDNEF HHHHHHCKGVEPKTECERLAGTESPVREEPGEDFPAARRLYWDDLKRKLSEKLDSTDFTGTIKLLNE NSYVPREAGSQKDENLALYVENQFREFKLSKVWRDQHFVKIQVKDSAQNSVIIVDKNGRLVYLVEN PGGYVAYSKAATVTGKLVHANFGTKKDFEDLYTPVNGSIVIVRAGKITFAEKVANAESLNAIGVLIY MDQTKFPI VNAEL SFFGHAHLGTGDP YTPGFPSFNHTQFPPSRS SGLPNIP VQTI SRAAAEKLFGNMEG DCPSDWKTDSTCRMVTSESKNVKLTVSNVLKEIKILNIFGVIKGFVEPDHYVVVGAQRDAWGPGAA KSGVGTALLLKLAQMFSDMVLKDGFQPSRSIIFASWSAGDFGSVGATEWLEGYLSSLHLKAFTYINL DKAVEGTSNFKVSASPEEYTETEKTMQNVKHPVTGQFLYQDSNWASKVEKETLDNAAFPFT, AYSGIP AVSFCFCEDTDYPYLGTTMDTYKELIERIPELNKVARAAAEVAGQFVIKLTHDVELNLDYERYNSQL LSFVRDLNQYRADIKEMGLSLQWLYSARGDFFRATSRLTTDFGNAEKTDRFVMKKLNDRVMRVEY HFLSPYVSPKESPFRHVFWGSGSHTLPALLENLKLRKQNNGAFNETLFRNQEALATW11QGAANAES GDVWDIDNEF HHHHHHHHGKPIPNPLLGLDSTGGGGSDSAQNSVIIVDKNGRLVYLVENPGGYVAYSKAATVTGKL VHANFGTKKDFEDLYTPVNGSIVIVRAGKITFAEKVANAESLNAIGVLIYMDQTKFPIVNAELSFFGHAHLGGGGGGLPNIPVQTISRAAAEKLFGNMEGDCPSDWKTDSTCRMVTSESKNVKLTVS
Claims
CLAIMS1. An ATXN2 RNAi agent comprising Formula (I): (R-L)n-P,wherein R is a double stranded RNA (dsRNA) comprising a sense stand and an antisense strand, wherein the antisense strand is complementary to ATXN2 mRNA; wherein L is a linker, or absent; andwherein P is a protein comprising one monovalent human TfR binding domain, wherein the human TfR binding domain comprises a heavy chain variable region (VH) and a light chain variable region (VL). wherein the VH comprises heavy chain complementarity determining regions HCDR1, HCDR2, and HCDR3, and the VL comprises light chain complementarity determining regions LCDR1, LCDR2, and LCDR3, wherein HCDR1 comprises SEQ ID NO: 1, HCDR2 comprises SEQ ID NO:
2. HCDR3 comprises SEQ ID NO: 3, LCDR1 comprises SEQ ID NO: 4, LCDR2 comprises SEQ ID NO: 5, and LCDR3 comprises SEQ ID NO: 6; andwherein n is an integer of 1 to 3.
2. The ATXN2 RNAi agent of claim 1, wherein n is 1.
3. The ATXN2 RNAi agent of claim 1, wherein n is 2.
4. The ATXN2 RNAi agent of any one of claims 1-3, wherein the sense strand and the antisense strand comprise a pair of nucleic acid sequences selected from the group consisting of:(a) the sense strand comprises SEQ ID NO: 35, and the antisense strand comprises SEQ ID NO: 36,(b) the sense strand comprises SEQ ID NO: 37, and the antisense strand comprises SEQ ID NO: 38,(c) the sense strand comprises SEQ ID NO: 39, and the antisense strand comprises SEQ ID NO: 40,(d) the sense strand comprises SEQ ID NO: 41, and the antisense strand comprises SEQ ID NO: 36,(e) the sense strand comprises SEQ ID NO: 42, and the antisense strand comprises SEQ ID NO: 38, and(f) the sense strand comprises SEQ ID NO: 43, and the antisense strand comprises SEQ ID NO: 40,wherein optionally one or more nucleotides of the sense strand and the antisense strand are independently modified nucleotides, and wherein optionally one or more intemucleotide linkages of the sense strand and the antisense strand are modified intemucleotide linkages.
5. The ATXN2 RNAi agent of any one of claims 1-4, wherein VH comprises SEQ ID NO: 7 and VL comprises SEQ ID NO: 8.
6. The ATXN2 RNAi agent of any one of claims 1-5, wherein the human TfR binding domain is a Fab, scFv, Fv, or scFab.
7. The ATXN2 RNAi agent of any one of claims 1-6, wherein the human TfR binding domain further comprises a heavy chain constant region comprising cysteine at residue 124 (according to the EU Index numbering).
8. The ATXN2 RNAi agent of any one of claims 1-7, wherein P further comprises a half-life extender.
9. The ATXN2 RNAi agent of claim 8. wherein the half-life extender is an immunoglobulin Fc region or a VHH that binds human serum albumin (HSA).
10. The ATXN2 RNAi agent of claim 8 or 9, wherein the half-life extender is an immunoglobulin Fc region.
11. The ATXN2 RNAi agent of claim 10, wherein the immunoglobulin Fc region is a modified human IgG4 Fc region.
12. The ATXN2 RNAi agent of claim 11, wherein the modified human IgG4 Fc region comprises proline at residue 228, and alanine at residues 234 and 235 (all residues are numbered according to the EU Index numbering).
13. The ATXN2 RNAi agent of any one of claims 10-12, wherein P comprises an immunoglobulin Fc region comprising cysteine at residue 378 (according to the EU Index numbering).
14. The ATXN2 RNAi agent of any one of claims 10-13, wherein the immunoglobulin Fc region comprises:(a) a first Fc CH3 domain comprising a serine at position 349, a methionine at position 366, a tyrosine at position 370, and a valine at position 409; and a second Fc CH3 domain comprising a glycine at position 356, an aspartic acid at position 357, a glutamine at position 364, and an alanine at position 407 (all residues are numbered according to the EU Index numbering); or(b) a first Fc CH3 domain comprising leucine at residue 405, and a second Fc CH3 domain comprising arginine at residue 409 (all residues are numbered according to the EU Index numbering).
15. The ATXN2 RNAi agent of any one of claims 1-7, wherein P comprises one heavy chain (HC) and one light chain (LC), wherein HC comprises SEQ ID NO: 9 and LC comprises SEQ ID NO: 10.
16. The ATXN2 RNAi agent of any one of claims 1-14, wherein P comprises two heavy chains HC1 and HC2 and one light chain LC1, wherein HC1 comprises SEQ ID NO:
14. LC1 comprises SEQ ID NO: 10, HC2 comprises SEQ ID NO: 15.
17. The ATXN2 RNAi agent of any one of claims 1-14, wherein P comprises two heavy chains HC1 and HC2 and one light chain LC1, wherein HC1 comprises SEQ ID NO:
16. LC1 comprises SEQ ID NO: 10, HC2 comprises SEQ ID NO: 17.
18. The ATXN2 RNAi agent of claim 8 or 9, wherein the half-life extender is a VHH that binds HSA.
19. The ATXN2 RNAi agent of claim 18, wherein the VHH comprises CDR1 comprising SEQ ID NO: 20, CDR2 comprising SEQ ID NO: 21, and CDR3 comprising SEQ ID NO: 22.
20. The ATXN2 RNAi agent of claim 18 or 19, wherein the VHH comprises SEQ ID NO:
21. The ATXN2 RNAi agent of any one of claims 18-20, wherein P comprises one heavy chain (HC) and one light chain (LC), and wherein the HC comprises SEQ ID NO: 11 and the LC comprises SEQ ID NO: 10.
22. The ATXN2 RNAi agent of any one of claims 1-14, wherein P is a heterodimeric antibody that comprises a first arm comprising one monovalent human TfR binding domain and a second arm that is a null arm.
23. The ATXN2 RNAi agent of claim 22, wherein the second arm comprises one heavy chain (HC) and one light chain (LC), and wherein the HC comprises SEQ ID NO: 18 and the LC comprises SEQ ID NO: 19.
24. The ATXN2 RNAi agent of claim 22 or 23, wherein P comprises two heavy chains HC1 and HC2 and two light chains LC1 and LC2, wherein HC1 comprises SEQ ID NO: 13, LC1 comprises SEQ ID NO: 10, HC2 comprises SEQ ID NO: 18, and LC2 comprises SEQ ID NO: 19.
25. The ATXN2 RNAi agent of any one of claims 1-24, wherein L is a SMCC linker, OD linker, or MSPT linker.
26. The ATXN2 RNAi agent of any one of claims 1-25, wherein L is a MSPT linker.
27. The ATXN2 RNAi agent of any one of claims 1-26, wherein P is linked to the 3' end of the sense strand of dsRNA via the linker.
28. The ATXN2 RNAi agent of any one of claims 1-26. wherein P is linked to the 5‘ end of the sense strand of dsRNA via the linker.
29. The ATXN2 RNAi agent of any one of claims 1-28, wherein one or more nucleotides of the sense strand are modified nucleotides.
30. The ATXN2 RNAi agent of claim 29, wherein each nucleotide of the sense strand is a modified nucleotide.
31. The ATXN2 RNAi agent of any one of claims 1-30, wherein one or more nucleotides of the antisense strand are modified nucleotides.
32. The ATXN2 RNAi agent of claim 31, wherein each nucleotide of the antisense strand is a modified nucleotide.
33. The ATXN2 RNAi agent of any one of claims 29-32, wherein the modified nucleotide is a 2'-fluoro modified nucleotide, 2'-O-methyl modified nucleotide, 2’ deoxy nucleotide (DNA), or 2'-O-Ci6 alkyl modified nucleotide.
34. The ATXN2 RNAi agent of any one of claims 29-33, wherein the sense strand has four 2'-fluoro modified nucleotides at positions 7, 9, 10, and 11 from the 5’ end of the sense strand.
35. The ATXN2 RNAi agent of claim 34, wherein nucleotides at positions other than positions 7, 9, 10, and 11 of the sense strand are 2'-O-methyl modified nucleotides.
36. The ATXN2 RNAi agent of any one of claims 29-35, wherein the antisense strand has four 2'-fluoro modified nucleotides at positions 2, 6, 14, and 16 from the 5‘ end of the antisense strand.
37. The ATXN2 RNAi agent of claim 36, wherein nucleotides at positions other than positions 2, 6, 14 and 16 of the antisense strand are 2'-O-methyl modified nucleotides.
38. The ATXN2 RNAi agent of any one of claims 29-33, wherein the sense strand has three 2'-fluoro modified nucleotides at positions 9, 10, and 11 from the 5’ end of the sense strand.
39. The ATXN2 RNAi agent of claim 38, wherein nucleotides at positions other than positions 9, 10, and 11 of the sense strand are 2'-O-methyl modified nucleotides.
40. The ATXN2 RNAi agent of any one of claims 29-35, 38, 39, wherein the antisense strand has five 2'-fluoro modified nucleotides at positions 2, 5, 7, 14, and 16 from the 5’ end of the antisense strand.
41. The ATXN2 RNAi agent of claim 40, wherein nucleotides at positions other than positions 2, 5, 7, 14, and 16 of the antisense strand are 2'-O-methyl modified nucleotides.
42. The ATXN2 RNAi agent of any one of claims 29-35, 38, 39, wherein the antisense strand has five 2'-fluoro modified nucleotides at positions 2, 5, 8, 14, and 16 from the 5’ end of the antisense strand.
43. The ATXN2 RNAi agent of claim 42, wherein nucleotides at positions other than positions 2, 5, 8, 14, and 16 of the antisense strand are 2'-O-methyl modified nucleotides.
44. The ATXN2 RNAi agent of any one of claims 29-35, 38, 39, wherein the antisense strand has five 2'-fluoro modified nucleotides at positions 2, 3, 7, 14, and 16 from the 5’ end of the antisense strand.
45. The ATXN2 RNAi agent of claim 44, wherein nucleotides at positions other than positions 2, 3, 7, 14, and 16 of the antisense strand are 2'-O-methyl modified nucleotides.
46. The ATXN2 RNAi agent of any one of claims 29-35, 38, 39, wherein the antisense strand has five 2'-fluoro modified nucleotides at positions 2, 14, and 16 from the 5‘ end of the antisense strand.
47. The ATXN2 RNAi agent of claim 46, wherein nucleotides at positions other than positions 2, 14. and 16 of the antisense strand are 2'-O-methyl modified nucleotides.
48. The ATXN2 RNAi agent of any one of claims 1-47, wherein the sense strand and the antisense strand have one or more modified intemucleotide linkages.
49. The ATXN2 RNAi agent of claim 48, wherein the modified intemucleotide linkage is phosphorothioate linkage.
50. The ATXN2 RNAi agent of claim 48 or 49, wherein the sense strand has four or five phosphorothioate linkages.
51. The ATXN2 RNAi agent of any one of claims48-50, wherein the antisense strand has four or five phosphorothioate linkages.
52. The ATXN2 RNAi agent of any one of claims 1-51, wherein the antisense strand has a phosphate analog at the 5’ end.
53. The ATXN2 RNAi agent of claim 52, wherein the phosphate analog is 5‘-vinylphosphonate.
54. The ATXN2 RNAi agent of any one of claims 1-53. wherein the sense strand or antisense strand comprises an abasic moiety or inverted abasic moiety.
55. The ATXN2 RNAi agent of any one of claims 1-54. wherein the sense strand and the antisense strand comprise a pair of nucleic acid sequences selected from the group consisting of:(a) the sense strand comprises SEQ ID NO: 44, and the antisense strand comprises SEQ ID NO: 45,(b) the sense strand comprises SEQ ID NO: 46, and the antisense strand comprises SEQ ID NO: 47,(c) the sense strand comprises SEQ ID NO: 48, and the antisense strand comprises SEQ ID NO: 49(d) the sense strand comprises SEQ ID NO: 50, and the antisense strand comprises SEQ ID NO: 45,(e) the sense strand comprises SEQ ID NO: 51, and the antisense strand comprises SEQ ID NO: 47, and(I) the sense strand comprises SEQ ID NO: 52, and the antisense strand comprises SEQ ID NO: 49.
56. The ATXN2 RNAi agent of any one of claims 1-55. wherein the sense strand and the antisense strand consist of a pair of nucleic acid sequences selected from the group consisting of:(a) the sense strand consists of SEQ ID NO: 44, and the antisense strand consists of SEQ ID NO: 45,(b) the sense strand consists of SEQ ID NO:
46. and the antisense strand consists of SEQ ID NO: 47,(c) the sense strand consists of SEQ ID NO: 48, and the antisense strand consists of SEQ ID NO: 49(d) the sense strand comprises SEQ ID NO: 50, and the antisense strand comprises SEQ ID NO: 45,(e) the sense strand comprises SEQ ID NO: 51, and the antisense strand comprises SEQ ID NO: 47, and(I) the sense strand comprises SEQ ID NO: 52, and the antisense strand comprises SEQ ID NO: 49.
57. A pharmaceutical composition comprising the ATXN2 RNAi agent of any one of claims 1-56 and a pharmaceutically acceptable carrier.
58. A method of treating a ATXN2-associated neurological disease in a patient in need thereof, the method comprising administering to the patient an effective amount of the ATXN2 RNAi agent of any one of claims 1-56, or the pharmaceutical composition of claim 57.
59. The method of claim 58, wherein the ATXN2-associated neurological disease is spinocerebellar ataxia type 2 (SCA2), amyotrophic lateral sclerosis (ALS), primary lateral sclerosis (PLS), Parkinson’s disease, Alzheimer’s disease, frontotemporal lobar degeneration (FTLD), progressive muscular atrophy (PMA), multiple system proteinopathy, Perry disease, or TDP-43 proteinopathy.
60. The method of any one of claims 58-61, wherein the ATXN2 RNAi agent is administered to the patient intrathecally. intravenously or subcutaneously.
61. The ATXN2 RNAi agent of any one of claims 1-56, or the pharmaceutical composition of claim 57, for use in a therapy.
62. The ATXN2 RNAi agent of any one of claims 1-56, or the pharmaceutical composition of claim 57, for use in the treatment of a ATXN2-associated neurological disease.
63. The ATXN2 RNAi agent or pharmaceutical composition for use of claim 62. wherein the ATXN2-associated neurological disease is spinocerebellar ataxia type 2(SCA2), amyotrophic lateral sclerosis (ALS), primary lateral sclerosis (PLS), Parkinson’s disease, Alzheimer’s disease, frontotemporal lobar degeneration (FTLD), progressivemuscular atrophy (PMA), multiple system proteinopathy, Perry disease, or TDP-43 proteinopathy.
64. Use of the ATXN2 RNAi agent of any one of claims 1-56 in the manufacture of a medicament for treating a ATXN2-associated neurological disease.
65. The use of claim 64, wherein the ATXN2-associated neurological disease is spinocerebellar ataxia type 2 (SCA2), amyotrophic lateral sclerosis (ALS). primary lateral sclerosis (PLS), Parkinson’s disease, Alzheimer’s disease, frontotemporal lobar degeneration (FTLD), progressive muscular atrophy (PMA), multiple system proteinopathy, Perry' disease, or TDP-43 proteinopathy.