Aggregation-resistant variants of TDP-43
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-08-04
- Publication Date
- 2026-08-14
AI Technical Summary
The mechanisms by which RNA-binding proteins such as TDP-43 cause neurodegeneration are not fully understood, and mutations in its prion-like domain (PLD) are associated with ALS, but the correlation between the physiological function of TDP-43 and these diseases remains unclear.
TDP-43 variants with mutated prion-like domains containing more aromatic amino acids and/or more evenly spaced aromatic amino acids are developed, which are less prone to aggregation and maintain splicing regulation functions, potentially treating or preventing TDP-43 proteinopathies like ALS.
The TDP-43 variants reduce aggregation and maintain normal splicing functions, offering therapeutic potential for ALS by expressing in cells or animals, potentially improving survival and splicing regulation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Application No. 63 / 370,527, filed August 05, 2022, which is incorporated herein by reference in its entirety for all purposes.
[0002] (Reference to sequence listing submitted as a text file via EFS WEB) The sequence listing set forth in file 599567SEQLIST.xml is 122 kilobytes, was created on August 4, 2023, and is incorporated herein by reference. [Background technology]
[0003] TAR DNA-binding protein 43 (TDP-43) is a ubiquitous protein encoded by the highly conserved TARDBP gene. TDP-43 is a primarily nuclear RNA-binding protein similar to members of the heterogeneous nuclear ribonucleoprotein (hnRNP) family, which are developmentally regulated and essential for early embryonic development. TDP-43 binds to thousands of RNAs with a strong preference for UG-rich sequences. TDP-43 autoregulates pre-mRNA synthesis by promoting selective spiking events at the 3' end of mRNA. Although the full 3D structure of TDP-43 has not yet been elucidated, several structural features of the TDP-43 protein have been identified. These include a nuclear localization signal (NLS), two RNA recognition motifs (RRM1 and RRM2), a putative nuclear export signal (NES), and a large domain in the carboxyl-terminal half of the protein that has been described as either less complex, less ordered, or a prion-like domain (PLD). TDP-43 has also been shown to be a disease signature protein associated with several neurodegenerative diseases, including amyotrophic lateral sclerosis (ALS), with 97% of ALS cases showing postmortem pathology of cytoplasmic TDP-43 aggregates. These aggregates are ubiquitinated, hyperphosphorylated, and excised. Additionally, mutations in TDP-43 have been associated with ALS. Of these rare TDP-43 mutations, the majority are found in its prion-like domain. The correlation between the physiological function of TDP-43 and these diseases remains unknown. Therefore, it is essential to understand the normal physiological role of TDP-43.
[0004] The mechanisms by which RNA-binding proteins such as TDP-43 cause neurodegeneration are not fully understood. TDP-43 is ubiquitously expressed and is thought to be involved in multiple levels of RNA metabolism, including transcription, splicing, transport, and translation. Despite the important role that TDP-43 plays in maintaining cellular lifespan, the structural and functional domains in which these functions are maintained remain poorly defined. Summary of the Invention
[0005] Provided herein are TDP-43 variants, nucleic acids encoding such TDP-43 variants, cells and animals containing such variants or nucleic acids, methods of making such cells and animals, and methods of using such TDP-43 variants, e.g., methods of treating TDP-43 proteinopathies such as amyotrophic lateral sclerosis (ALS), in which the prion-like domain (PLD) of the TDP-43 variants has been mutated to have more aromatic amino acids and / or more evenly spaced aromatic amino acids than the PLD from wild-type TDP-43.
[0006] In one aspect, a TAR DNA-binding protein 43 (TDP-43) variant is provided, wherein the prion-like domain (PLD) of the TDP-43 variant is mutated to have more aromatic amino acids and / or more evenly spaced aromatic amino acids than the PLD from wild-type TDP-43. In some such TDP-43 variants, the PLD is mutated to have more aromatic amino acids. In some such TDP-43 variants, the PLD is mutated to have more evenly spaced aromatic amino acids than the PLD from wild-type TDP-43. In some such TDP-43 variants, the PLD is mutated to have more aromatic amino acids and more evenly spaced aromatic amino acids than the PLD from wild-type TDP-43.
[0007] In some such TDP-43 variants, the PLD in the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 22. In some such TDP-43 variants, the PLD in the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 63. In some such TDP-43 variants, the PLD in the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 22. In some such TDP-43 variants, the PLD in the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 63. In some such TDP-43 variants, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 23. In some such TDP-43 variants, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 64. In some such TDP-43 variants, the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 23. In some such TDP-43 variants, the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 64.
[0008] In some such TDP-43 variants, a portion of the PLD from wild-type TDP-43 is replaced with at least a portion of the PLD from a different RNA-binding protein in the variant TDP-43. In some such TDP-43 variants, the replaced portion of the PLD from wild-type TDP-43 is at least about 10, at least about 15, at least about 20, at least about 25, or at least about 28 amino acids. In some such TDP-43 variants, the replaced portion of the PLD from wild-type TDP-43 is about 10 to about 50, about 20 to about 40, or about 25 to about 35 amino acids. In some such TDP-43 variants, the replaced portion of the PLD is about 28 amino acids. In some such TDP-43 variants, the replaced portion of the PLD comprises, consists essentially of, or consists of SEQ ID NO:6, or the replaced portion of the PLD comprises, consists essentially of, or consists of SEQ ID NO:47. In some such TDP-43 variants, the different RNA-binding protein is hnRNPA2B1 (e.g., human hnRNPA2B1 or mouse hnRNPA2B1). In some such TDP-43 variants, a portion of the PLD from hnRNPA2B1 is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO:24. In some such TDP-43 variants, a portion of the PLD from hnRNPA2B1 is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO:65. In some such TDP-43 variants, the portion of the PLD from hnRNPA2B1 comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:24.In some such TDP-43 variants, a portion of the PLD from hnRNPA2B1 comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 65. In some such TDP-43 variants, the PLD in the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 25. In some such TDP-43 variants, the PLD in the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 66. In some such TDP-43 variants, the PLD in the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 25. In some such TDP-43 variants, the PLD in the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 66. In some such TDP-43 variants, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 26. In some such TDP-43 variants, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 67. In some such TDP-43 variants, the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 26. In some such TDP-43 variants, the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:67.
[0009] In some such TDP-43 variants, the PLD from wild-type TDP-43 is replaced with a PLD from a different RNA-binding protein in the variant TDP-43. In some such TDP-43 variants, the different RNA-binding protein is hnRNPA2B1 (e.g., human hnRNPA2B1 or mouse hnRNPA2B1). In some such TDP-43 variants, the PLD in the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 27 or 13. In some such TDP-43 variants, the PLD in the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 68 or 58. In some such TDP-43 variants, the PLD in the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 27 or 13. In some such TDP-43 variants, the PLD in the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 68 or 58. In some such TDP-43 variants, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 28. In some such TDP-43 variants, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 69. In some such TDP-43 variants, the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:28.In some such TDP-43 variants, the TDP-43 variant comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:69.
[0010] In some such TDP-43 variants, the TDP-43 variant is less prone to aggregation than wild-type TDP-43. In some such TDP-43 variants, the TDP-43 variant retains the function of wild-type TDP-43 in splicing regulation. In some such TDP-43 variants, the TDP-43 variant is predominantly nuclear and / or retains the subcellular distribution of wild-type TDP-43. In some such TDP-43 variants, the TDP-43 variant retains the function of wild-type TDP-43 during embryonic development. In some such TDP-43 variants, the TDP-43 variant is a human TDP-43 variant.
[0011] Some such TDP-43 variants are for use in treating a TDP-43 proteinopathy in a subject, optionally wherein the TDP-43 proteinopathy is amyotrophic lateral sclerosis (ALS). Some such TDP-43 variants are for use in preventing a TDP-43 proteinopathy in a subject, optionally wherein the TDP-43 proteinopathy is amyotrophic lateral sclerosis (ALS). Some such TDP-43 variants are for the manufacture of a medicament for the treatment of a TDP-43 proteinopathy, optionally wherein the TDP-43 proteinopathy is amyotrophic lateral sclerosis (ALS). In another aspect, there is provided the use of any of the above TDP-43 variants for the manufacture of a medicament for the prevention of a TDP-43 proteinopathy, optionally wherein the TDP-43 proteinopathy is amyotrophic lateral sclerosis (ALS).
[0012] In another aspect, a nucleic acid encoding any of the above TDP-43 variants is provided. In some such nucleic acids, the nucleic acid is messenger RNA. In some such nucleic acids, the nucleic acid comprises DNA, optionally, the DNA comprises complementary DNA (cDNA). In some such nucleic acids, the nucleic acid is in an expression construct comprising a promoter operably linked to the nucleic acid encoding the TDP-43 variant, optionally, the promoter is a neuron-specific promoter or a constitutive promoter. In some such nucleic acids, the promoter is a constitutive promoter, a tissue-specific promoter, or an inducible promoter. In some such nucleic acids, the promoter is a neuron-specific promoter, optionally, the promoter is a synapsin-1 promoter, optionally, the promoter is a human synapsin-1 promoter.
[0013] In some such nucleic acids, the nucleic acid is in a vector. In some such nucleic acids, the vector is a viral vector. In some such nucleic acids, the viral vector is a lentiviral vector or an adeno-associated virus (AAV) vector. In some such nucleic acids, the vector is an AAV vector, optionally, the AAV vector is an AAV-PHP.eB vector.
[0014] In some such nucleic acids, the nucleic acids are codon-optimized for expression in human cells.
[0015] In another aspect, a cell is provided that contains any of the above TDP-43 variants or any of the above nucleic acids encoding a TDP-43 variant. In some such cells, the TDP-43 variant is expressed. In some such cells, the cell is a mammalian cell. In some such cells, the mammalian cell is a human cell, a rodent cell, a mouse cell, or a rat cell. In some such cells, the cell is a human cell. In some such cells, the cell is a neuron, a glial cell, or a muscle cell. In some such cells, the cell is in vivo in a subject. In some such cells, the cell is a neuron in the brain of a subject.
[0016] In some such cells, endogenous TDP-43 is not expressed in the cell. In some such cells, the endogenous TARDBP genomic locus contains a mutation that prevents expression of endogenous TDP-43 in the cell. In some such cells, the cell further contains an agent that reduces or eliminates expression of endogenous TDP-43 in the cell. In some such cells, the agent contains an antisense oligonucleotide or RNAi agent that targets endogenous TARDBP messenger RNA, or a nucleic acid encoding the antisense oligonucleotide or RNAi agent. In some such cells, the agent contains a nuclease agent that targets the endogenous TARDBP genomic locus, or one or more nucleic acids encoding a nuclease agent. In some such cells, the nuclease agent is a Zinc Finger Nuclease (ZFN), a Transcription Activator-Like Effector Nuclease (TALEN), or a Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-associated (Cas) protein and guide RNA. In some such cells, the nuclease agent is a Cas protein and guide RNA, and optionally, the Cas protein is a Cas9 protein.
[0017] In some such cells, the cell has a genetically modified endogenous TARDBP genomic locus, and the nucleic acid is integrated at the endogenous TARDBP genomic locus. In some such cells, the cell is heterozygous for the integrated nucleic acid. In some such cells, the cell is homozygous for the integrated nucleic acid. In some such cells, the nucleic acid is operably linked to an endogenous TARDBP promoter. In some such cells, the TDP-43 variant is expressed from the endogenous TARDBP genomic locus. In some such cells, the TDP-43 variant replaces expression of endogenous TDP-43.
[0018] In some such cells, the cells have reduced TDP-43 aggregation compared to control cells that do not contain the TDP-43 variant or nucleic acid.
[0019] In another aspect, a non-human animal is provided that comprises any of the above TDP-43 variants or any of the above nucleic acids encoding a TDP-43 variant. In some such non-human animals, the TDP-43 variant is expressed. In some such non-human animals, the non-human animal is a mammal. In some such non-human animals, the non-human animal is a rodent, mouse, or rat. In some such non-human animals, the non-human animal is a mouse. In some such non-human animals, the TDP-43 variant or nucleic acid is in a neuron, glial cell, or muscle cell of the non-human animal. In some such non-human animals, the TDP-43 variant or nucleic acid is in a neuron. In some such non-human animals, the neuron is in the brain of the non-human animal.
[0020] In some such non-human animals, endogenous TDP-43 is not expressed in the non-human animal. In some such non-human animals, the endogenous TARDBP genomic locus contains a mutation that prevents expression of endogenous TDP-43 in the non-human animal. In some such non-human animals, the non-human animal contains an agent that reduces or eliminates expression of endogenous TDP-43 in the non-human animal. In some such non-human animals, the agent contains an antisense oligonucleotide or RNAi agent that targets endogenous TARDBP messenger RNA, or a nucleic acid encoding the antisense oligonucleotide or RNAi agent. In some such non-human animals, the agent contains a nuclease agent that targets the endogenous TARDBP genomic locus, or one or more nucleic acids encoding a nuclease agent. In some such non-human animals, the nuclease agent is a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), or a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein and guide RNA. In some such non-human animals, the nuclease agent is a Cas protein and guide RNA, and optionally, the Cas protein is a Cas9 protein.
[0021] In some such non-human animals, the non-human animal has a genetically modified endogenous TARDBP genomic locus, and the nucleic acid is integrated at the endogenous TARDBP genomic locus. In some such non-human animals, the non-human animal is heterozygous for the integrated nucleic acid. In some such non-human animals, the non-human animal is homozygous for the integrated nucleic acid. In some such non-human animals, the nucleic acid is operably linked to an endogenous TARDBP promoter. In some such non-human animals, the TDP-43 variant is expressed from the endogenous TARDBP genomic locus. In some such non-human animals, the TDP-43 variant replaces expression of endogenous TDP-43.
[0022] In some such non-human animals, the non-human animal has reduced TDP-43 aggregation compared to a control non-human animal that does not contain the TDP-43 variant or nucleic acid.
[0023] In another aspect, there is provided a method of producing any of the above non-human animals, comprising administering a TDP-43 variant or nucleic acid to the non-human animal. In another aspect, methods are provided for making any of the above-described non-human animals, comprising: (I) (a) modifying the genome of a pluripotent non-human animal cell to comprise a genetically modified endogenous TARDBP genomic locus; (b) identifying or selecting a genetically modified pluripotent non-human animal cell that comprises a genetically modified endogenous TARDBP genomic locus; (c) introducing the genetically modified pluripotent non-human animal cell into a non-human animal host embryo; and (d) gestating the non-human animal host embryo in a surrogate mother; or (II) (a) modifying the genome of a one-cell stage embryo of the non-human animal to comprise a genetically modified endogenous TARDBP genomic locus; (b) selecting a genetically modified one-cell stage embryo of the non-human animal that comprises a genetically modified endogenous TARDBP genomic locus; and (c) gestating the genetically modified one-cell stage embryo of the non-human animal in a surrogate mother.
[0024] In another aspect, a method is provided comprising administering to a cell (i) any of the TDP-43 variants described above, or (ii) any of the nucleic acids encoding a TDP-43 variant described above. In some such methods, the TDP-43 variant is expressed. In some such methods, the cell is a mammalian cell. In some such methods, the mammalian cell is a human cell, a rodent cell, a mouse cell, or a rat cell. In some such methods, the cell is a human cell. In some such methods, the cell is a neuron, a glial cell, or a muscle cell. In some such methods, the cell is in vivo in a subject. In some such methods, the cell is a neuron in the brain of a subject. In some such methods, the TDP-43 variant or nucleic acid is administered to the subject via intraventricular, intracranial, or intrathecal injection.
[0025] In some such methods, the method includes administering to a cell an agent that reduces or eliminates expression of endogenous TDP-43 in the cell. In some such methods, the agent includes an antisense oligonucleotide or RNAi agent that targets endogenous TARDBP messenger RNA, or a nucleic acid encoding the antisense oligonucleotide or RNAi agent. In some such methods, the agent includes a nuclease agent that targets the endogenous TARDBP genomic locus, or one or more nucleic acids encoding the nuclease agent. In some such methods, the nuclease agent is a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), or a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein and a guide RNA. In some such methods, the nuclease agent is a Cas protein and a guide RNA, and optionally, the Cas protein is a Cas9 protein.
[0026] In some such methods, the method includes administering a nucleic acid encoding a TDP-43 variant, wherein the nuclease agent cleaves the endogenous TARDBP genomic locus, and the nucleic acid encoding the TDP-43 variant is inserted at or recombines with the cleaved endogenous TARDBP genomic locus, and the TDP-43 variant is expressed from the endogenous TARDBP genomic locus. In some such methods, the TDP-43 variant replaces expression of endogenous TDP-43.
[0027] In some such methods, endogenous TDP-43 in the cell is prone to aggregation, and the method reduces TDP-43 aggregation in the cell. In some such methods, there is aberrant splicing regulation by endogenous TDP-43 in the cell, and the TDP-43 variant rescues the aberrant TDP-43 splicing regulation in the cell. In some such methods, there is aberrant subcellular distribution of endogenous TDP-43 in the cell, and the TDP-43 variant rescues the aberrant subcellular distribution of endogenous TDP-43 in the cell.
[0028] In another aspect, methods for treating a TDP-43 proteinopathy in a subject are provided. Some such methods comprise administering to one or more cells in the subject (i) any of the TDP-43 variants described above, or (ii) any of the nucleic acids encoding a TDP-43 variant described above. In some such methods, the TDP-43 variant is expressed in one or more cells in the subject. In another aspect, methods for preventing a TDP-43 proteinopathy in a subject are provided. Some such methods comprise administering to one or more cells in the subject (i) any of the TDP-43 variants described above, or (ii) any of the nucleic acids encoding a TDP-43 variant described above. In some such methods, the TDP-43 variant is expressed in one or more cells in the subject.
[0029] In some such methods, the TDP-43 proteinopathy is amyotrophic lateral sclerosis (ALS). In some such methods, the subject is a mammal. In some such methods, the subject is a human. In some such methods, the one or more cells comprise neurons in the brain of the subject. In some such methods, the TDP-43 variant or nucleic acid is administered to the subject via intraventricular, intracranial, or intrathecal injection.
[0030] In some such methods, the method further comprises administering to the one or more cells an agent that reduces or eliminates expression of endogenous TDP-43 in the one or more cells. In some such methods, the agent comprises an antisense oligonucleotide or RNAi agent that targets endogenous TARDBP messenger RNA, or a nucleic acid encoding the antisense oligonucleotide or RNAi agent. In some such methods, the agent comprises a nuclease agent that targets the endogenous TARDBP genomic locus, or one or more nucleic acids encoding the nuclease agent. In some such methods, the nuclease agent is a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), or a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein and a guide RNA. In some such methods, the nuclease agent is a Cas protein and a guide RNA, and optionally, the Cas protein is a Cas9 protein.
[0031] In some such methods, the method includes administering a nucleic acid encoding a TDP-43 variant, wherein the nuclease agent cleaves an endogenous TARDBP genomic locus in one or more cells, the nucleic acid encoding the TDP-43 variant is inserted at or recombines with the cleaved endogenous TARDBP genomic locus, and the TDP-43 variant is expressed from the endogenous TARDBP genomic locus, replacing expression of endogenous TDP-43.
[0032] In some such methods, endogenous TDP-43 in one or more cells is prone to aggregation, and the method reduces TDP-43 aggregation in one or more cells. In some such methods, there is aberrant splicing regulation by endogenous TDP-43 in one or more cells, and the TDP-43 variant rescues the aberrant TDP-43 splicing regulation in one or more cells. In some such methods, there is aberrant subcellular distribution of endogenous TDP-43 in one or more cells, and the TDP-43 variant rescues the aberrant subcellular distribution of endogenous TDP-43 in one or more cells. [Brief explanation of the drawings]
[0033] [Figure 1] The amino acid composition of the prion-like domains (PLDs) from TDP-43, hnRNPA1, and hnRNPA2B1 is shown. Aromatic residues are highlighted in orange. Biochemical studies have shown that even spacing of aromatic residues throughout the PLD allows efficient liquid-liquid phase separation (LLPS) but prevents irreversible association that leads to aggregation. The PLD of TDP-43 contains fewer aromatic amino acids and is less evenly spaced than other RNA-binding proteins, suggesting that reconfiguring the spacing of aromatic residues throughout the PLD may prevent aggregation. [Figure 2A] The structure of TDP-43 is shown. [Figure 2B] To determine whether the PLD of another RNA-binding protein could replace the PLD of TDP-43, we show that we replaced the PLD of TDP-43 with the PLD from hnRNPA2B1, an RNA-binding protein that is less prone to aggregation. [Figure 3]The results show that wild-type and PLDswap TDP-43 proteins were detected at the expected size of approximately 43 kDa. Using an antibody recognizing the N-terminus of TDP-43, mutant TDP-43 polypeptides lacking functional PLD were redistributed from the nucleus to the cytoplasm in ES cell-derived motor neurons. Similar to wild-type TDP-43, TDP-43 PLDswap is predominantly nuclear. As expected, an antibody recognizing the C-terminus of the wild-type TDP-43 protein failed to detect TDP-43 PLDswap protein (data not shown). [Figure 4] We show that embryos lacking functional TDP-43 protein (TDP-43- / -) were not viable and did not survive beyond the E3.5 stage. Similarly, embryos expressing only TDP-43 protein lacking a functional NLS (ΔNLS / -) or only TDP-43 protein lacking a functional PLD (TDP-43ΔPLD / -) were not viable, but these embryos survived longer, suggesting that they may compensate for some of the embryonic functions of TDP-43. Embryos expressing only the PLDswap protein were able to survive to birth, indicating that this form can fully complement wild-type TDP-43 during embryogenesis. [Figure 5]RT-PCR analysis of the indicated TDP-43-dependent splicing events is shown. The specific splicing events being monitored are indicated in the mRNA schematic on the right, with primer locations indicated by black arrows. Cryptic exons are indicated as red boxes in the schematic. In ES cell-derived MNs in which ΔNLS and ΔNES are the only forms of TDP-43 protein, there is a clear loss of function in splicing regulation. This loss of function is less severe in mutants in which ΔPLD is the only form of TDP-43. TDP-43 function is nearly restored in ES cell-derived MNs in which the only form of TDP-43 is the PLDswap chimeric protein. Adnp2 and Dnajc5 assays monitor aberrant inclusion of cryptic exons, whereas Poldip3 and Tsn monitor alternative exon skipping and Sortilin1 monitor alternative exon inclusion. [Figure 6A] Survival of PLDswap neonates is shown. Mice with PLDswap protein as the only form of TDP-43 after Cre-mediated removal of the WT allele either ubiquitously (CAG-Cre) or neuronally restricted (SYN-Cre) (Figure 6A) survive at similar levels compared to WT heterozygous control mice (Figure 6B). Survival plots of TDP-43ΔEx3 / ΔEx3 after Cre-mediated removal of the conditional WT allele across genotypes (CAG-Cre, n = 18; SYN-Cre, n = 12), TDP-43ΔEx3 / WT (CAG-Cre, n = 7; SYN-Cre, n = 9), and TDP-43ΔEx3 / PLDswap (CAG-Cre, n = 5; SYN-Cre, n = 7). Median survival times were as follows: for CAG-Cre: TDP-43ΔEx3 / ΔEx3: 4 weeks, TDP-43ΔEx3 / WT: 17.36 weeks, TDP-43ΔEx3 / PLDswap: 29.43 weeks; for SYN-Cre: TDP-43ΔEx3 / ΔEx3: 3.93 weeks, TDP-43ΔEx3 / WT: 52 weeks, TDP-43ΔEx3 / PLDswap: 37 weeks. [Figure 6B]Survival of PLDswap neonates is shown. Mice with PLDswap protein as the only form of TDP-43 after Cre-mediated removal of the WT allele either ubiquitously (CAG-Cre) or neuronally restricted (SYN-Cre) (Figure 6A) survive at similar levels compared to WT heterozygous control mice (Figure 6B). Survival plots of TDP-43ΔEx3 / ΔEx3 after Cre-mediated removal of the conditional WT allele across genotypes (CAG-Cre, n = 18; SYN-Cre, n = 12), TDP-43ΔEx3 / WT (CAG-Cre, n = 7; SYN-Cre, n = 9), and TDP-43ΔEx3 / PLDswap (CAG-Cre, n = 5; SYN-Cre, n = 7). Median survival times were as follows: for CAG-Cre: TDP-43ΔEx3 / ΔEx3: 4 weeks, TDP-43ΔEx3 / WT: 17.36 weeks, TDP-43ΔEx3 / PLDswap: 29.43 weeks; for SYN-Cre: TDP-43ΔEx3 / ΔEx3: 3.93 weeks, TDP-43ΔEx3 / WT: 52 weeks, TDP-43ΔEx3 / PLDswap: 37 weeks. [Figure 7] Figure 1 shows neonatal RT-PCR analysis of TDP-43-dependent splicing in PLDswap. RT-PCR analysis of the indicated TDP-43-dependent splicing events is shown. The specific splicing events being monitored are indicated in the mRNA schematic on the right, with primer positions indicated by black arrows. Adnp2 monitors aberrant inclusion of cryptic exons, whereas Tsn monitors alternative exon skipping. The cryptic exon in Adnp2 is shown as a box between exons 2 and 3 in the schematic. Loss of TDP-43 (ΔEx3 / ΔEx3) causes missplicing of both Adnp2 and Tsn, and both tend to be rescued when PLDswap is the only form of TDP-43.
[0034] definition The terms "protein," "polypeptide," and "peptide," used interchangeably herein, include polymeric forms of amino acids of any length, including coded and non-coded amino acids, and amino acids that are chemically or biochemically modified or derivatized. These terms also include modified polymers, such as polypeptides having modified peptide backbones. The term "domain" refers to any portion of a protein or polypeptide having a specific function or structure.
[0035] The terms "nucleic acid" and "polynucleotide," used interchangeably herein, include polymeric forms of nucleotides of any length, containing ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof. These include single-, double-, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers that contain purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.
[0036] The terms "expression vector" or "expression construct" or "expression cassette" refer to a recombinant nucleic acid comprising a desired coding sequence operably linked to appropriate nucleic acid sequences necessary for the expression of the operably linked coding sequence in a particular host cell or organism. Nucleic acid sequences necessary for expression in prokaryotes usually include a promoter, an operator (optional), a ribosome binding site, and other sequences. Eukaryotic cells are well known to generally utilize promoters, enhancers, termination and polyadenylation signals, although some elements can be deleted and others added without sacrificing the necessary expression.
[0037] The term "viral vector" refers to a recombinant nucleic acid that contains at least one element of viral origin and contains elements sufficient for or that allow packaging into a viral vector particle. The vector and / or particle can be used to transfer DNA, RNA, or other nucleic acids into cells either ex vivo or in vivo. Many forms of viral vectors are known.
[0038] The term "isolated," with respect to proteins, nucleic acids, and cells, includes proteins, nucleic acids, and cells that are relatively purified, usually with respect to other cellular or biological components that may be present in situ, and includes substantially pure preparations of proteins, nucleic acids, or cells. The term "isolated" may also include proteins and nucleic acids that have no naturally occurring counterpart, or proteins or nucleic acids that are chemically synthesized and thus are substantially uncontaminated by other proteins or nucleic acids. The term "isolated" may also include proteins, nucleic acids, or cells that have been separated or purified from most other cellular or biological components with which they are naturally associated (e.g., other cellular proteins, nucleic acids, or cellular or extracellular components).
[0039] The term "wild-type" includes entities having a structure and / or activity as found in a normal state or context (as opposed to a diseased, altered, etc. mutant). Wild-type genes and polypeptides often exist in multiple alternative forms (e.g., alleles).
[0040] The term "endogenous sequence" refers to a nucleic acid sequence that occurs naturally within a cell or animal. For example, an endogenous TARDBP sequence in an animal refers to the native TARDBP sequence that occurs naturally at the TARDBP locus in the animal.
[0041] "Exogenous" molecules or sequences include molecules or sequences that are not normally present in a cell in that form, or molecules or sequences that are introduced into a cell from an external source. Normal presence includes presence in relation to a particular developmental stage and environmental conditions of the cell. An exogenous molecule or sequence can include, for example, a mutated version of a corresponding endogenous sequence in the cell, such as a humanized version of an endogenous sequence, or can include a sequence that corresponds to an endogenous sequence within the cell but in a different form (i.e., not within a chromosome). In contrast, an endogenous molecule or sequence includes a molecule or sequence that is normally present in that form, in a particular cell, at a particular developmental stage, and under particular environmental conditions.
[0042] The term "heterologous" when used in the context of a nucleic acid or a protein indicates that the nucleic acid or protein comprises at least two segments that do not naturally occur together in the same molecule. For example, when used with respect to a segment of a nucleic acid or a segment of a protein, the term "heterologous" indicates that the nucleic acid or protein comprises two or more subsequences that are not found in the same relationship to each other (e.g., linked together) in nature. As an example, a "heterologous" region of a nucleic acid vector is a segment of nucleic acid within or attached to another nucleic acid molecule that is not found in association with that other molecule in nature. For example, a heterologous region of a nucleic acid vector can include a coding sequence adjacent to a heterologous promoter that is not found in association with the coding sequence in nature. Similarly, a "heterologous" region of a protein is a segment of amino acids within or attached to another peptide molecule that is not found in association with that other peptide molecule in nature (e.g., a fusion protein or a tagged protein). Similarly, a nucleic acid or protein can include a heterologous tag or a heterologous secretion or localization sequence.
[0043] "Codon optimization" refers to the process of modifying a nucleic acid sequence for enhanced expression in a particular host cell by taking advantage of codon degeneracy, as indicated by the diversity of three-base pair codon combinations that specify an amino acid, generally by replacing at least one codon of the native sequence with a codon more frequently or most frequently used in the host cell's genes while maintaining the native amino acid sequence. For example, a nucleic acid encoding a TAR DNA-binding protein 43 (TDP-43) protein can be modified to use alternative codons that are more frequently used in a given prokaryotic or eukaryotic cell, including bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, or any other host cell, compared to the naturally occurring nucleic acid sequence. Codon usage tables are readily available, for example, in "codon usage databases." These tables can be applied in a variety of ways. See Nakamura et al. (2000) Nucleic Acids Research 28:292, incorporated herein by reference in its entirety for all purposes. Computer algorithms are also available for codon optimization of a particular sequence for expression in a particular host (see, eg, Gene Forge).
[0044] A "promoter" is a regulatory region of DNA that typically contains a TATA box that can direct RNA polymerase II to begin RNA synthesis at the appropriate transcription start site for a particular polynucleotide sequence. Promoters may additionally contain other regions that affect the rate of transcription initiation. The promoter sequences disclosed herein regulate transcription of an operably linked polynucleotide. The promoters may be active in one or more cell types disclosed herein (e.g., eukaryotic cells, non-human mammalian cells, human cells, rodent cells, pluripotent cells, one-cell stage embryos, differentiated cells, or combinations thereof). The promoters may be, for example, constitutively active, conditional, inducible, temporally restricted (e.g., developmentally regulated), or spatially restricted (e.g., cell-specific or tissue-specific) promoters. Examples of promoters can be found, for example, in International Publication No. WO 2013 / 176772, which is incorporated herein by reference in its entirety for all purposes.
[0045] Constitutive promoters are promoters that are active in all tissues or in specific tissues at all stages of development. Examples of constitutive promoters include human cytomegalovirus immediate early (hCMV), mouse cytomegalovirus immediate early (mCMV), human elongation factor 1 alpha (hEF1α), mouse elongation factor 1 alpha (mEF1α), mouse phosphoglycerate kinase (PGK), chicken beta actin hybrid (CAG or CBh), SV40 early, and beta2 tubulin promoters.
[0046] Examples of inducible promoters include chemically regulated promoters and physically regulated promoters. Chemically regulated promoters include alcohol-regulated promoters (e.g., alcohol dehydrogenase (alcA) gene promoter), tetracycline-regulated promoters (e.g., tetracycline-responsive promoter, tetracycline operator sequence (tetO), tet-On promoter, or tet-Off promoter), steroid-regulated promoters (e.g., rat glucocorticoid receptor, estrogen receptor promoter, or ecdysone receptor promoter), or metal-regulated promoters (e.g., metalloprotein promoters). Physically regulated promoters include temperature-regulated promoters (e.g., heat shock promoters) and light-regulated promoters (e.g., light-inducible promoters or light-repressible promoters).
[0047] The tissue-specific promoter can be, for example, a neuron-specific promoter or a glia-specific promoter or a muscle-specific promoter.
[0048] Developmentally-regulated promoters include, for example, promoters that are active only during embryonic stages of development or only in adult cells.
[0049] "Operable linkage" or "operably linked" includes the juxtaposition of two or more components (e.g., a promoter and another sequence element) that allows for the normal function of both components and the potential for at least one of the components to mediate the function of at least one of the other components. For example, a promoter can be operably linked to a coding sequence if it controls the level of transcription of the coding sequence depending on the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include proximity of such sequences to each other, or acting in trans (e.g., regulatory sequences can act at a distance to control transcription of the coding sequence).
[0050] The term "in vitro" includes an artificial environment and processes or reactions that occur within an artificial environment (e.g., a test tube or an isolated cell or cell line). The term "in vivo" includes a natural environment (e.g., a cell or organism or body) and processes or reactions that occur within a natural environment. The term "ex vivo" includes cells removed from an individual's body and processes or reactions that occur within such cells.
[0051] A composition or method "comprising" or "including" one or more recited elements may include other elements not specifically recited. For example, a composition "comprises" or "includes" a protein may contain the protein alone or in combination with other components. The transitional phrase "consisting essentially of" means that the scope of the claim shall be construed to include the specific elements recited in the claim as well as elements that do not materially affect the basic and novel characteristics of the claimed invention. Thus, the term "consisting essentially of," when used in the claims of the present invention, is not intended to be construed as equivalent to "comprising."
[0052] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes examples where the event or circumstance occurs and examples where it does not occur.
[0053] Designation of a range of values includes all integers within or defining that range, and all subranges defined by integers within that range, e.g., 5-10 nucleotides is understood to mean 5, 6, 7, 8, 9, or 10 nucleotides, while 5-10% is understood to include all possible values between 5% and 10%.
[0054] At least 17 nucleotides of a 20 nucleotide sequence is understood to include 17, 18, 19, or 20 nucleotides of the provided sequence, thereby providing an upper limit even if one is not specifically provided, as is clearly understood. Similarly, up to 3 nucleotides is understood to encompass 0, 1, 2, or 3 nucleotides, providing a lower limit even if one is not specifically provided. When "at least," "up to," or other similar language modifies a number, it can be understood to modify each number in the series.
[0055] As used herein, "less than" or "less than" is understood as the value adjacent to the phrase and any logically lower value or integer logically from the context, up to 0. For example, a duplex region of "two or fewer nucleotide base pairs" has 2, 1, or 0 nucleotide base pairs. When "less than" or "less than" appears before a series of numbers or ranges, it is understood that each of the numbers in the series or range is modified.
[0056] As used herein, when a maximum value is expressed as 100% (e.g., 100% inhibition), it is understood that the value is limited by the detection method. For example, 100% inhibition is understood as inhibition relative to a level below the detection level of the assay.
[0057] Unless otherwise clear from the context, the term "about" includes values that are ±5% of the stated value. In certain embodiments, the term "about" is understood to include variations or errors accepted within the art, such as two standard deviations from the mean, or the sensitivity of the method used to make the measurement, or percentages of values accepted in the art, such as those with age. When "about" is present before the first value in a series, it can be understood to modify each value in the series.
[0058] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").
[0059] The term "or" refers to any one member of a particular list and also includes any combination of members of that list.
[0060] The singular articles "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "protein" or "at least one protein" can include a plurality of proteins, including mixtures thereof.
[0061] Statistically significant means p≦0.05.
[0062] In the event of a discrepancy between the sequence in this application and the indicated accession number or position in the accession number, the sequence in this application will control. DETAILED DESCRIPTION OF THE INVENTION
[0063] I. Overview Provided herein are TDP-43 variants, nucleic acids encoding such TDP-43 variants, and methods of using such TDP-43 variants, for example, methods for treating TDP-43 proteinopathies such as amyotrophic lateral sclerosis (ALS), in which the prion-like domain (PLD) of the TDP-43 variant is mutated to have more aromatic amino acids and / or more evenly spaced aromatic amino acids than the PLD from wild-type TDP-43. ALS is a devastating neurodegenerative disease affecting motor neurons, resulting in paralysis and eventual death. A nearly universal pathological finding in postmortem examination of ALS patient tissue is the accumulation of TDP-43 (interactive response DNA-binding protein 43 kDa) in cytoplasmic inclusions. TDP-43 is a primarily nuclear RNA-binding protein similar in structure to members of the heterogeneous nuclear ribonucleoprotein (hnRNP) family, which is required for the viability of all mammalian cells and for normal development and life in animals. The redistribution of TDP-43 from the nucleus to the cytoplasm and its accumulation in insoluble aggregates are two key diagnostic features of ALS disease. Although the biological function of TDP-43 may not yet be fully understood, there is evidence that this protein is involved in regulating pre-messenger RNA (pre-mRNA) splicing by preventing the use of cryptic exons in large introns and by affecting the alternative splicing of some pre-mRNAs. TDP-43 has also been proposed to have functions in the cytoplasm, possibly in shuttling RNA between the nucleus and cytoplasm, and in the transport of mRNA within neuronal axons.
[0064] Several structural features of the TDP-43 protein have been identified, including a nuclear localization signal (NLS), two RNA recognition motifs (RRM1 and RRM2), a putative nuclear export signal (NES), and a large domain in the carboxyl-terminal half of the protein that has been described as a less complex, less ordered, or prion-like domain (PLD). Of the mutations in TDP-43 associated with familial cases of ALS, the majority are found in the PLD.
[0065] Studies of TDP-43 domain mutants have identified two mutants (ΔPLD and ΔNLS) that are viable as embryonic stem cell-derived motor neurons and mislocalize the TDP-43 protein from the nucleus to the cytoplasm, inducing cytoplasmic aggregation. Mice expressing only these mutants (ΔPLD / - and ΔNLS / -) are embryonic lethal. Mice heterozygous for ΔNLS (ΔNLS / +) and ΔPLD (ΔPLD / +) are viable and exhibit mislocalized, aggregated, insoluble, and phosphorylated TDP-43.
[0066] In ΔPLD / + mice, we found that the mutant (approximately 30 kDa) protein could be distinguished from the native (approximately 43 kDa) WT protein. The TDP-43ΔPLD protein was found almost exclusively in the cytoplasmic fraction, while some wild-type TDP-43 was also found in the cytoplasmic fraction. The wild-type protein, but not the ΔPLD protein, was phosphorylated and found in the detergent-insoluble fraction. These results suggest that the presence of the mutant TDP-43 protein affects the behavior of the wild-type protein. In other experiments, ΔNLS / + mice were found to have progressive degeneration of the motor system. Mice heterozygous for ΔPLD exhibited degeneration, but unlike ΔNLS, this degeneration did not progress by 4 to 6 months. Adult motor neurons were able to survive with TDP-43 without PLD, and fewer aggregates were observed in cells with ΔPLD as the only form of TDP-43. This suggests that the TDP-43 PLD mediates aggregate formation, suggesting that removal or re-engineering of the TDP-43 PLD may provide a novel therapeutic approach. Biochemical studies have shown that even spacing of aromatic residues throughout the PLD allows efficient liquid-liquid phase separation but prevents irreversible association that leads to aggregation. The TDP-43 PLD contains fewer aromatic amino acids and is less evenly spaced than other RNA-binding proteins, suggesting that reconfiguring the spacing of aromatic residues throughout the PLD may prevent aggregation.
[0067] To test this, we designed and generated TDP-43 variants in which the prion-like domain (PLD) of the TDP-43 variant was mutated to have more aromatic amino acids and / or more evenly spaced aromatic amino acids than the PLD from wild-type TDP-43. As an example, we generated a PLDswap allele, replacing the TDP-43 PLD with the PLD from hnRNPA2B1. The hnRNPA2B1 PLD is composed of more evenly spaced aromatic residues than the TDP-43 PLD. Embryonic stem (ES) cells and ES cell-derived motor neurons carrying the PLDswap allele as the only form of TDP-43 are viable and exhibit normal TDP-43 subcellular distribution. Replacement of the TDP-43 PLD with the hNRNPA2B1 PLD also rescues the splicing defect associated with loss of functional TDP-43. This suggests that a potential therapeutic modality could be to remove wild-type TDP-43 and replace it with an engineered, aggregation-resistant form. Simply replacing the wild-type protein could also rescue the phenotypes associated with loss of functional TDP-43.
[0068] II. TDP-43 variants Variants of TAR DNA-binding protein 43 (TDP-43) are provided herein. Human TDP-43 has been assigned the UniProt reference number Q13148. The human gene encoding TDP-43 (TARDBP or TDP43) has been assigned the NCBI GeneID 23435 and is found on chromosome 1 at location 1p36.22 (construct: GRCh38.p14(GCF_000001405.40), location: NC_000001.11(11012654..11030528)). At least two isoforms of human TDP-43 are known. The first isoform is 414 amino acids long and has been assigned the UniProt reference number Q13148-1 and the NCBI reference number NP_031401.1 (SEQ ID NO: 1). An exemplary coding sequence has been assigned the reference number CCDS122.1 (SEQ ID NO: 2), and an exemplary mRNA (cDNA) sequence has been assigned the reference number NM_007375.4 (SEQ ID NO: 3). The second isoform is 298 amino acids and has been assigned the UniProt reference number Q13148-2 (SEQ ID NO: 4).
[0069] Mouse TDP-43 has been assigned the UniProt reference number Q921F2. The mouse gene encoding TDP-43 (TARDBP or TDP43) has been assigned NCBI GeneID 230908 and is found at position 4E2 on chromosome 4 (construct: GRCm39 (GCF_000001635.27), position: NC_000070.7 (148696839..148711672, complement)). The canonical isoform is 414 amino acids long and has been assigned the UniProt reference number Q921F2-1 and the NCBI reference number NP_663531.1 (SEQ ID NO: 43). An exemplary coding sequence has been assigned the reference number CCDS38971.1 (SEQ ID NO: 44), and an exemplary mRNA (cDNA) sequence has been assigned the reference number NM_145556.4 (SEQ ID NO: 45).
[0070] TDP-43 is a highly conserved, ubiquitously expressed RNA / DNA-binding protein belonging to the heterogeneous nuclear ribonucleoprotein (hnRNP) family. TDP-43 is crucial for multiple cellular functions, including regulation of RNA metabolism, mRNA transport, microRNA maturation, and stress granule formation. Consistent with its nuclear and cytoplasmic functions, TDP-43 can shuttle between the nucleus and cytoplasm, but under normal physiological conditions, its localization is primarily nuclear. In relation to brain function, TDP-43 appears to be important for the normal development of central neurons during early embryogenesis.
[0071] Dysfunction of TDP-43-related pathways is increasingly recognized as a key pathogenic mechanism in neurodegenerative diseases. Hyperphosphorylated and ubiquitinated TDP-43 cytoplasmic inclusions have been identified as a pathological hallmark of amyotrophic lateral sclerosis (ALS) and frontotemporal lobar disease (FTLD). Pathogenic missense mutations in the TARDBP gene, which encodes the TDP-43 protein, were subsequently identified as the causative genetic mutation in both ALS and FTLD, although only in a small percentage of familial cases. The majority of patients with ALS and FTLD do not have mutations in the TARDBP gene but demonstrate a wide range of abnormalities involving TDP-43. TDP-43 deposition has been associated with an increasing number of neurodegenerative diseases and has been identified as a major pathogenic factor, resulting in these disorders, termed TDP-43 proteinopathies.
[0072] The protein structure of TDP-43 consists of an N-terminal region, a nuclear localization signal (NLS), two RNA recognition motifs: RRM1 and RRM2, a nuclear export signal (NES), and a C-terminal region encompassing a prion-like domain (PLD). The PLD is a subset of low-complexity regions rich in uncharged polar amino acids and glycine, which bear similarity to the yeast prion protein, as defined using a hidden Markov algorithm. PLDs are often found in RNA-binding proteins that drive protein aggregation in neurodegenerative disorders such as amyotrophic lateral sclerosis. Although primarily localized in the nucleus, TDP-43 shuttles between the nucleus and cytoplasm, a process mediated by active and passive transport, where it exerts its physiological functions. In addition, TDP-43 localizes to mitochondria, where it associates with the mitochondrial genome and is important in the respiratory chain pathway. TDP-43's best-known function is regulating the splicing of cryptic and alternative exons. Cryptic exons are intronic sequences that are incorrectly recognized as exons by the splicing machinery. GU-rich TDP-43 binding sites are often found near cryptic exons in large introns. TDP-43 suppresses the recognition of cryptic exons in large introns and promotes normal splicing. Loss of TDP-43 leads to aberrant splicing of cryptic exons, resulting in the loss of normal RNA and protein. In addition, TDP-43 regulates alternative splicing of specific target RNAs. This regulation, including both exon inclusion and exon exclusion, can affect transcripts involved in ALS pathogenesis.
[0073] The N-terminal domain is important for the formation of functional homodimers, which is crucial for proper TDP-43 physiological function. Located within the N-terminal region is the NLS domain (amino acids 82–98), which mediates the import of TDP-43 into the nucleus, where it exerts its physiological functions. The RNA-binding motifs (RRM1 (amino acids 106–176) and RRM2 (amino acids 191–262)) are essential for the TDP-43 protein to bind to RNA / DNA molecules, regulate mRNA transcription, translation, splicing, and stability, and mediate RNA export. Additionally, TDP-43 forms ribonucleoprotein (RNP) granules, which are important for transporting mRNA molecules and promoting the biogenesis of non-coding RNAs, such as microRNAs (miRNAs). Separately, the prion-like domain (PLD; amino acids 274-414; SEQ ID NO: 5 in human TDP-43; SEQ ID NO: 46 in mouse TDP-43) has been implicated in the pathogenesis of TDP-43 because this region regulates protein solubility and mediates pathological aggregation. TDP-43 is also important for the formation of stress granules, protecting neurons from cellular insults such as oxidative stress.
[0074] Biochemical studies have shown that even spacing of aromatic residues throughout the PLD allows for efficient liquid-liquid phase separation but prevents irreversible association leading to aggregation. The PLD of TDP-43 contains fewer aromatic amino acids and is less evenly spaced than other RNA-binding proteins, suggesting that reconfiguring the spacing of aromatic residues throughout the PLD may prevent aggregation. Provided herein are TDP-43 variants engineered to be less prone to aggregation than wild-type TDP-43. Provided herein are TDP-43 variants (e.g., human TDP-43 variants) in which the PLD of the TDP-43 variant has been mutated to have more aromatic amino acids and / or more evenly spaced aromatic amino acids than the PLD from wild-type TDP-43. The three aromatic amino acids are tyrosine, phenylalanine, and tryptophan. In some cases, the TDP-43 variants described herein are mutated to have more aromatic amino acids than the PLD from wild-type TDP-43. In some cases, the TDP-43 variants described herein are mutated to have aromatic amino acids that are more evenly spaced than the PLD from wild-type TDP-43. In some cases, the TDP-43 variants described herein are mutated to have more aromatic amino acids and more evenly spaced than the PLD from wild-type TDP-43.
[0075] In one example, the PLD of the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO:22. In one example, the PLD of the TDP-43 variant comprises the sequence set forth in SEQ ID NO:22. In one example, the PLD of the TDP-43 variant consists essentially of the sequence set forth in SEQ ID NO:22. In one example, the PLD of the TDP-43 variant consists of the sequence set forth in SEQ ID NO:22. In one example, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO:23. In one example, the TDP-43 variant comprises the sequence set forth in SEQ ID NO:23. In one example, the TDP-43 variant consists essentially of the sequence set forth in SEQ ID NO:23. In one example, the TDP-43 variant consists of the sequence set forth in SEQ ID NO:23.
[0076] In one example, the PLD of the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO:63. In one example, the PLD of the TDP-43 variant comprises the sequence set forth in SEQ ID NO:63. In one example, the PLD of the TDP-43 variant consists essentially of the sequence set forth in SEQ ID NO:63. In one example, the PLD of the TDP-43 variant consists of the sequence set forth in SEQ ID NO:63. In one example, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO:64. In one example, the TDP-43 variant comprises the sequence set forth in SEQ ID NO:64. In one example, the TDP-43 variant consists essentially of the sequence set forth in SEQ ID NO:64. In one example, the TDP-43 variant consists of the sequence set forth in SEQ ID NO:64.
[0077] In one example, a portion of PLD from wild-type TDP-43 is replaced with at least a portion of PLD from a different RNA-binding protein in the variant TDP-43. In one example, the replaced portion of PLD from wild-type TDP-43 and / or the replaced portion of PLD from the different RNA-binding protein is at least about 10, at least about 15, at least about 20, at least about 25, or at least about 28 amino acids. In one example, the replaced portion of PLD from wild-type TDP-43 and / or the replaced portion of PLD from the different RNA-binding protein is about 10 to about 50, about 20 to about 40, about 25 to about 35, about 10 to about 45, about 10 to about 40, about 10 to about 35, about 10 to about 30, about 15 to about 50, about 20 to about 50, about 25 to about 50, or about 28 to about 50 amino acids. In another example, the replaced portion of PLD from wild-type TDP-43 and / or the portion of PLD from a different RNA-binding protein is about 10 to about 140, about 20 to about 140, about 30 to about 140, about 40 to about 140, about 50 to about 140, about 60 to about 140, about 70 to about 140, about 80 to about 140, about 90 to about 140, about 100 to about 140, about 110 to about 140, about 120 to about 140, about 130 to about 140, about 28 to about 130, about 28 to about 120, about 28 to about 110, about 28 to about 100, about 28 to about 90, about 28 to about 80, about 28 to about 70, about 28 to about 60, about 28 to about 50, or about 28 to about 40 amino acids. In another example, the replaced portion of PLD from wild-type TDP-43 and / or the replaced portion of PLD from the different RNA-binding protein is about 10 to about 50, about 20 to about 40, or about 25 to about 35 amino acids. In another example, the replaced portion of PLD from wild-type TDP-43 and / or the replaced portion of PLD from the different RNA-binding protein is about 28 amino acids.
[0078] For example, a highly conserved 28 amino acid stretch in the PLD of TDP-43, which has been shown to be important for flocculation in yeast, can be replaced with a sequence from a PLD from a different RNA-binding protein. For example, the portion of the PLD from wild-type TDP-43 that is replaced can include the sequence set forth in SEQ ID NO: 6. In another example, the portion of the PLD from wild-type TDP-43 that is replaced essentially consists of the sequence set forth in SEQ ID NO: 6. In another example, the portion of the PLD from wild-type TDP-43 that is replaced consists of the sequence set forth in SEQ ID NO: 6. For example, the portion of the PLD from wild-type TDP-43 that is replaced can include the sequence set forth in SEQ ID NO: 47. In another example, the portion of the PLD from wild-type TDP-43 that is replaced essentially consists of the sequence set forth in SEQ ID NO: 47. In another example, the portion of the PLD from wild-type TDP-43 that is replaced consists of the sequence set forth in SEQ ID NO: 47.
[0079] In another example, the portions of the PLDs from the different RNA binding proteins are at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 24. In another example, the portions of the PLDs from the different RNA binding proteins comprise the sequence set forth in SEQ ID NO: 24. In another example, the portions of the PLDs from the different RNA binding proteins consist essentially of the sequence set forth in SEQ ID NO: 24. In another example, the portions of the PLDs from the different RNA binding proteins consist of the sequence set forth in SEQ ID NO: 24. In one example, the PLDs of the TDP-43 variants are at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 25. In one example, the PLDs of the TDP-43 variants comprise the sequence set forth in SEQ ID NO: 25. In one example, the PLDs of the TDP-43 variants consist essentially of the sequence set forth in SEQ ID NO: 25. In one example, the PLD of the TDP-43 variant consists of the sequence set forth in SEQ ID NO:25. In one example, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO:26. In one example, the TDP-43 variant comprises the sequence set forth in SEQ ID NO:26. In one example, the TDP-43 variant consists essentially of the sequence set forth in SEQ ID NO:26. In one example, the TDP-43 variant consists of the sequence set forth in SEQ ID NO:26.
[0080] In another example, the portion of the PLD from the different RNA binding protein is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 65. In another example, the portion of the PLD from the different RNA binding protein comprises the sequence set forth in SEQ ID NO: 65. In another example, the portion of the PLD from the different RNA binding protein consists essentially of the sequence set forth in SEQ ID NO: 65. In another example, the portion of the PLD from the different RNA binding protein consists of the sequence set forth in SEQ ID NO: 65. In one example, the PLD of the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 66. In one example, the PLD of the TDP-43 variant comprises the sequence set forth in SEQ ID NO: 66. In one example, the PLD of the TDP-43 variant consists essentially of the sequence set forth in SEQ ID NO: 66. In one example, the PLD of the TDP-43 variant consists of the sequence set forth in SEQ ID NO: 66. In one example, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 67. In one example, the TDP-43 variant comprises the sequence set forth in SEQ ID NO: 67. In one example, the TDP-43 variant consists essentially of the sequence set forth in SEQ ID NO: 67. In one example, the TDP-43 variant consists of the sequence set forth in SEQ ID NO: 67.
[0081] In one example, the PLD from wild-type TDP-43 is replaced with a PLD from a different RNA-binding protein in the variant TDP-43. In one example, the PLD from the different RNA-binding protein is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 13 or 27. In one example, the PLD from the different RNA-binding protein comprises the sequence set forth in SEQ ID NO: 13 or 27. In another example, the PLD from the different RNA-binding protein consists essentially of the sequence set forth in SEQ ID NO: 13 or 27. In another example, the PLD from the different RNA-binding protein consists of the sequence set forth in SEQ ID NO: 13 or 27. In one example, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 28. In one example, the TDP-43 variant comprises the sequence set forth in SEQ ID NO: 28. In one example, the TDP-43 variant consists essentially of the sequence set forth in SEQ ID NO: 28. In one example, the TDP-43 variant consists of the sequence set forth in SEQ ID NO: 28.
[0082] In one example, the PLD from wild-type TDP-43 is replaced with a PLD from a different RNA-binding protein in the variant TDP-43. In one example, the PLD from the different RNA-binding protein is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 58 or 68. In one example, the PLD from the different RNA-binding protein comprises the sequence set forth in SEQ ID NO: 58 or 68. In another example, the PLD from the different RNA-binding protein consists essentially of the sequence set forth in SEQ ID NO: 58 or 68. In another example, the PLD from the different RNA-binding protein consists of the sequence set forth in SEQ ID NO: 58 or 68. In one example, the TDP-43 variant is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 69. In one example, the TDP-43 variant comprises the sequence set forth in SEQ ID NO: 69. In one example, the TDP-43 variant consists essentially of the sequence set forth in SEQ ID NO: 69. In one example, the TDP-43 variant consists of the sequence set forth in SEQ ID NO: 69.
[0083] In one example, the PLD from wild-type TDP-43 is replaced with a PLD from a different RNA-binding protein in the variant TDP-43. In one example, the PLD from the different RNA-binding protein is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 21. In one example, the PLD from the different RNA-binding protein comprises the sequence set forth in SEQ ID NO: 21. In another example, the PLD from the different RNA-binding protein consists essentially of the sequence set forth in SEQ ID NO: 21. In another example, the PLD from the different RNA-binding protein consists of the sequence set forth in SEQ ID NO: 21.
[0084] In one example, the PLD from wild-type TDP-43 is replaced with a PLD from a different RNA-binding protein in the variant TDP-43. In one example, the PLD from the different RNA-binding protein is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the sequence set forth in SEQ ID NO: 73. In one example, the PLD from the different RNA-binding protein comprises the sequence set forth in SEQ ID NO: 73. In another example, the PLD from the different RNA-binding protein consists essentially of the sequence set forth in SEQ ID NO: 73. In another example, the PLD from the different RNA-binding protein consists of the sequence set forth in SEQ ID NO: 73.
[0085] The different RNA-binding protein may be less prone to aggregation than wild-type TDP-43. Similarly, the different RNA-binding protein may have more aromatic residues in the PLD and / or more evenly spaced aromatic residues in the PLD than wild-type TDP-43.
[0086] One example of such an RNA-binding protein is heterogeneous nuclear ribonucleoprotein A2 / B1 (hnRNP A2 / B1 or hnRNPA2B1). hnRNPA2B1 is a heterogeneous nuclear ribonucleoprotein (hnRNP) that associates with nascent pre-mRNAs and packages them into hnRNP particles. hnRNP particle positioning on nascent hnRNAs is nonrandom but sequence-dependent, and serves to condense and stabilize transcripts, minimizing tangles and knotting. Packaging plays a role in various processes, including transcription, pre-mRNA processing, RNA export, subcellular location, mRNA translation, and mature mRNA stability. Human hnRNPA2B1 has been assigned the UniProt reference number P22626. The human gene encoding hnRNPA2B1 (HNRNPA2B1 or HNRPA2B1) has been assigned NCBI GeneID 3181 and is found on chromosome 7 at location 7p15.2 (construct: GRCh38.p14 (GCF_000001405.40), location: NC_000007.14 (26189927..26200746, complement)). At least two isoforms of human hnRNPA2B1 are known. The first isoform has been assigned UniProt reference number P22626-1 and NCBI reference number NP_112533.1 (SEQ ID NO: 7). An exemplary coding sequence has been assigned reference number CCDS43557.1 (SEQ ID NO: 8), and an exemplary mRNA (cDNA) sequence has been assigned reference number NM_031243.3 (SEQ ID NO: 9). The second isoform has been assigned UniProt reference number P22626-2 and NCBI reference number NP_002128.1 (SEQ ID NO: 10). An exemplary coding sequence has been assigned reference number CCDS5397.1 (SEQ ID NO: 11), and an exemplary mRNA (cDNA) sequence has been assigned reference number NM_002137.3 (SEQ ID NO: 12). The PLD of human hnRNPA2B1 is set forth in SEQ ID NO: 13 or 27.
[0087] Mouse hnRNPA2B1 has been assigned the UniProt reference number O88569. The mouse gene encoding hnRNPA2B1 (Hnrnpa2b1 or Hnrpa2b1) has been assigned NCBI GeneID 53379 and is found at position 6, 6B3 on chromosome 6 (construct: GRCm39 (GCF_000001635.27), position: NC_000072.7 (51437414..51448054, complement)). At least three isoforms of mouse hnRNPA2B1 are known. The first isoform has been assigned the UniProt reference number O88569-1, and the NCBI reference numbers NP_001361674.1 and XP_006506436.2 (SEQ ID NO: 48). An exemplary coding sequence has been assigned reference number CCDS90046.1 (SEQ ID NO: 49), and exemplary mRNA (cDNA) sequences have been assigned reference numbers XM_006506373.3 (SEQ ID NO: 50) and NM_001374745.1 (SEQ ID NO: 51). A second isoform has been assigned UniProt reference number O88569-2 and NCBI reference number NP_058086.2 (SEQ ID NO: 52). An exemplary coding sequence has been assigned reference number CCDS51774.1 (SEQ ID NO: 53), and exemplary mRNA (cDNA) sequences have been assigned reference number NM_016806.3 (SEQ ID NO: 54). A third isoform has been assigned UniProt reference number O88569-3 and NCBI reference number NP_872591.1 (SEQ ID NO: 55). An exemplary coding sequence has been assigned the reference number CCDS51773.1 (SEQ ID NO: 56), and an exemplary mRNA (cDNA) sequence has been assigned the reference number NM_182650.4 (SEQ ID NO: 57). The PLD of mouse hnRNPA2B1 is set forth in SEQ ID NO: 58 or 68.
[0088] Another example of such an RNA-binding protein is heterogeneous nuclear ribonucleoprotein A1 (hnRNPA1). hnRNPA1 is involved in packaging pre-mRNA into hnRNP particles, transport of poly(A) mRNA from the nucleus to the cytoplasm, and regulating splice site selection. Human hnRNPA1 has been assigned the UniProt reference number P09651. The human gene encoding hnRNPA1 (HNRNPA1 or HNRPA1) has been assigned NCBI GeneID 3178 and is found on chromosome 12 at location 12q13.13 (construct: GRCh38.p14 (GCF_000001405.40), location: NC_000012.12 (54280726..54287087)). At least three isoforms of human hnRNPA1 are known. The first isoform has been assigned UniProt reference number P09651-1 and NCBI reference number NP_112420.1 (SEQ ID NO: 14). An exemplary coding sequence has been assigned reference number CCDS44909.1 (SEQ ID NO: 15), and an exemplary mRNA (cDNA) sequence has been assigned reference number NM_031157.4 (SEQ ID NO: 16). The second isoform has been assigned UniProt reference number P09651-2 and NCBI reference number NP_002127.1 (SEQ ID NO: 17). An exemplary coding sequence has been assigned reference number CCDS41793.1 (SEQ ID NO: 18), and an exemplary mRNA (cDNA) sequence has been assigned reference number NM_002136.4 (SEQ ID NO: 19). The PLD (second isoform) of human hnRNPA2B1 is set forth in SEQ ID NO: 21. The third isoform has been assigned the UniProt reference number P09651-3 (SEQ ID NO: 20).
[0089] Mouse hnRNPA1 has been assigned the UniProt reference number P49312. The human gene encoding hnRNPA1 (Hnrnpa1 or Hnrpa1) has been assigned the NCBI GeneID 15382 and is found at position 15 F3; 15 58.58 cM on chromosome 15 (construct: GRCm39 (GCF_000001635.27), position: NC_000081.7 (103148370..103155125)). At least two isoforms of mouse hnRNPA1 are known. The first isoform has been assigned the UniProt reference number P49312-1 and the NCBI reference number NP_034577.1 (SEQ ID NO: 59). An exemplary coding sequence has been assigned the reference number CCDS37233.1 (SEQ ID NO: 60), and an exemplary mRNA (cDNA) sequence has been assigned the reference number NM_010447.5 (SEQ ID NO: 61). A second isoform has been assigned the UniProt reference number P49312-2 (SEQ ID NO: 62). The PLD of mouse hnRNPA2B1 is set forth in SEQ ID NO: 73.
[0090] In some cases, the TDP-43 variant is less prone to aggregation than wild-type TDP-43. In some cases, the TDP-43 variant retains the function of wild-type TDP-43 in splicing regulation (e.g., retains most of the functionality of wild-type TDP-43 in regulating the splicing of Adnp2, Dnajc5, Tsn, and / or Sortilin1). In some cases, the TDP-43 variant is primarily nuclear. In some cases, the TDP-43 variant retains the subcellular distribution of wild-type TDP-43. In some cases, the TDP-43 variant retains the function of wild-type TDP-43 during embryonic development.
[0091] III. Nucleic acids encoding TDP-43 variants Nucleic acids or nucleic acid constructs encoding TDP-43 variants are provided herein. The nucleic acids or nucleic acid constructs may be isolated nucleic acid constructs.
[0092] In some cases, nucleic acids encoding TDP-43 variants can be codon-optimized (e.g., codon-optimized for expression in humans or mice). For example, the nucleic acids can be modified to alternative codons more frequently used in human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest.
[0093] The nucleic acid encoding the TDP-43 variant may be DNA or RNA. The nucleic acid may optionally be messenger RNA (mRNA) encoding the TDP-43 variant. The nucleic acid may optionally be complementary DNA (cDNA) encoding the TDP-43 variant. Examples of coding sequences for some of the TDP-43 variants described herein are set forth, for example, in SEQ ID NOS: 70-72. For example, such a nucleic acid may comprise only the coding sequence without any intervening introns. In other cases, the nucleic acid may comprise one or more introns separating exons in the TDP-43 variant coding sequence. For example, the nucleic acid may comprise a genomic sequence comprising both exons and introns.
[0094] In some cases, the nucleic acid is in an expression construct comprising a nucleic acid encoding a TDP-43 variant operably linked to a promoter. The promoter can be any suitable promoter for in vivo expression in an animal or in vitro expression in an isolated cell. The promoter can be a constitutively active promoter (e.g., a CAG promoter or a U6 promoter), a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific promoter or a tissue-specific promoter). Such promoters are well known and are discussed elsewhere herein. In a specific example, the promoter is active in neurons. In a specific example, the promoter is active in glial cells. In a specific example, the promoter is active in muscle cells. In some cases, the promoter is a heterologous promoter (i.e., a promoter to which the TDP-43 nucleic acid is not naturally operably linked). In other cases, the promoter may be an endogenous promoter (i.e., a TDP-43 variant nucleic acid operably linked to a TDP-43 promoter). The heterologous promoter may be any type of promoter disclosed elsewhere herein. For example, the promoter may be a constitutive promoter, such as the EF1α promoter. Alternatively, the promoter may be a tissue-specific promoter or an inducible promoter. For example, the promoter may be a neuron-specific promoter. One example of a suitable neuron-specific promoter is the synapsin-1 promoter (e.g., the human synapsin-1 promoter). For example, the promoter may be a glia-specific promoter. For example, the promoter may be a muscle-specific promoter.
[0095] The nucleic acids and expression constructs disclosed herein can also include post-transcriptional regulatory elements, such as the post-transcriptional regulatory elements of the woodchuck hepatitis virus.
[0096] The nucleic acid and expression construct may further comprise one or more polyadenylation signal sequences. For example, the nucleic acid construct may comprise a polyadenylation signal sequence located 3' of the TDP-43 variant coding sequence. Any suitable polyadenylation signal sequence may be used. The term polyadenylation signal sequence refers to any sequence that directs the termination of transcription and the addition of a poly(A) tail to an mRNA transcript. In eukaryotes, transcription terminators are recognized by protein factors, and termination is followed by polyadenylation, the process of adding a poly(A) tail to an mRNA transcript in the presence of poly(A) polymerase. Mammalian poly(A) signals typically consist of a core sequence approximately 45 nucleotides long, which may be flanked by various auxiliary sequences that help increase the efficiency of cleavage and polyadenylation. The core sequence, called the poly(A) recognition motif or poly(A) recognition sequence, consists of a highly conserved upstream element (AATAAA or AAUAAA) in mRNA that is recognized by cleavage and polyadenylation-specificity factor (CPSF) and a poorly defined downstream region (rich in U or G and U) that is bound by cleavage stimulation factor (CstF). Examples of transcription terminators that can be used include, for example, the human growth hormone (HGH) polyadenylation signal, the simian virus 40 (SV40) late polyadenylation signal, the rabbit beta globin polyadenylation signal, the bovine growth hormone (BGH) polyadenylation signal, the phosphoglycerate kinase (PGK) polyadenylation signal, the AOX1 transcription termination sequence, the CYC1 transcription termination sequence, or any transcription termination sequence known to be suitable for regulating gene expression in eukaryotic cells.
[0097] The nucleic acid and expression construct may also include a polyadenylation signal sequence upstream of the TDP-43 variant coding sequence. The polyadenylation signal sequence upstream of the TDP-43 variant coding sequence may be flanked by recombinase recognition sites recognized by a site-specific recombinase. In some constructs, the recombinase recognition sites also flank a selection cassette containing, for example, a coding sequence for a drug resistance protein. In some constructs, the recombinase recognition sites do not flank the selection cassette. The polyadenylation signal sequence prevents transcription and expression of the protein or RNA encoded by the coding sequence. However, upon exposure to a site-specific recombinase, the polyadenylation signal sequence may be excised, allowing the protein or RNA to be expressed.
[0098] Such a configuration can allow for tissue- or developmental stage-specific expression if the polyadenylation signal sequence is excised in a tissue- or developmental stage-specific manner. Excision of the polyadenylation signal sequence in a tissue- or developmental stage-specific manner can be achieved if the animal containing the nucleic acid or expression construct further contains a coding sequence for a site-specific recombinase operably linked to a tissue- or developmental stage-specific promoter. The polyadenylation signal sequence is then excised only in those tissues or those developmental stages, allowing for tissue- or developmental stage-specific expression. In one example, the TDP-43 variant encoded by the nucleic acid or expression construct can be expressed in a neuron-specific manner. In one example, the TDP-43 variant encoded by the nucleic acid or expression construct can be expressed in a glia- or developmental stage-specific manner. In one example, the TDP-43 variant encoded by the nucleic acid or expression construct can be expressed in a muscle-specific manner.
[0099] Site-specific recombinases include enzymes that can promote recombination between recombinase recognition sites where the two recombination sites are physically separated within a single nucleic acid or on separate nucleic acids. Examples of recombinases include Cre, Flp, and Dre recombinases. One example of a Cre recombinase gene is Crei, in which the two exons encoding Cre recombinase are separated by an intron, preventing expression in prokaryotic cells. Such recombinases may further contain a nuclear localization signal to promote nuclear localization (e.g., NLS-Crei). Recombinase recognition sites include nucleotide sequences that are recognized by site-specific recombinases and can serve as substrates for recombination events. Examples of recombinase recognition sites include FRT, FRT11, FRT71, attp, att, rox, and lox sites such as loxP, lox511, lox2272, lox66, lox71, loxM2, and lox5171.
[0100] The nucleic acids disclosed herein may comprise deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), which may be single-stranded or double-stranded, and which may be in linear or circular form. Nucleic acid constructs may be naked nucleic acids or may be delivered by vectors, such as AAV vectors, as described elsewhere herein. When in linear form, the ends of the nucleic acid may be protected (e.g., from exonuclease degradation) by well-known methods. For example, one or more dideoxynucleotide residues may be added to the 3' end of the linear molecule, and / or self-complementary oligonucleotides may be attached to one or both ends. See, for example, Chang et al. (1987) Proc. Natl. Acad. Sci. USA 84:4959-4963 and Nehls et al. (1996) Science 272:886-889, each of which is incorporated herein by reference in its entirety for all purposes. Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino group(s) and the use of modified internucleotide linkages such as phosphorothioates, phosphoramidates, and O-methylribose or deoxyribose residues. A nucleic acid or expression construct may optionally include one or more of the following terminal structures: a hairpin, a loop, an inverted terminal repeat (ITR), or a toroid. For example, a nucleic acid or expression construct may include an ITR.
[0101] The nucleic acid or expression construct may include modifications or sequences that provide additional desirable characteristics (e.g., altered or controlled stability, tracking or detection using fluorescent labels, binding sites for proteins or protein complexes, etc.). For example, modifications may be made to one or more nucleosides within the mRNA. Examples of chemical modifications to mRNA nucleobases include pseudouridine, 1-methyl-pseudouridine, and 5-methyl-cytidine. The mRNA may also be capped. The mRNA may also be polyadenylated (to include a poly(A) tail). As an example, a capped polyadenylated mRNA containing N1-methyl-pseudouridine can be used (e.g., completely replaced with N1-methyl-pseudouridine). The nucleic acid construct may include one or more fluorescent labels, purification tags, epitope tags, or combinations thereof. For example, the nucleic acid construct may include one or more fluorescent labels (e.g., fluorescent proteins or other fluorophores or dyes), such as at least one, at least two, at least three, at least four, or at least five fluorescent labels. Exemplary fluorescent labels include fluorophores such as fluorescein (e.g., 6-carboxyfluorescein (6-FAM)), Texas Red, HEX, Cy3, Cy5, Cy5.5, Pacific Blue, 5-(and-6)-carboxytetramethylrhodamine (TAMRA), and Cy7. A wide range of fluorescent dyes for labeling oligonucleotides are commercially available (e.g., from Integrated DNA Technologies). The label or tag can be at the 5' end, 3' end, or internally located on the nucleic acid construct. For example, the nucleic acid construct can be conjugated to the 5' end with the IR700 fluorophore (5'IRDYE® 700) from Integrated DNA Technologies.
[0102] Nucleic acids and expression constructs can also include conditional alleles. Conditional alleles can be multifunctional alleles, as described in U.S. Patent Application Publication No. 2011 / 0104799, which is incorporated by reference in its entirety for all purposes. For example, a conditional allele can include: (a) an actuating sequence in the sense orientation relative to transcription of a target gene; (b) a drug selection cassette (DSC) in the sense or antisense orientation; (c) a nucleotide sequence of interest (NSI) in the antisense orientation; and (d) a conditional by inversion (COIN) module in the reverse orientation, which utilizes an exon-splitting intron and a reversible gene trap-like module. See, e.g., U.S. Patent Application Publication No. 2011 / 0104799. The conditional allele can further comprise recombinable units that recombine upon exposure to a first recombinase to form a conditional allele that (i) lacks an actuating sequence and a DSC, and (ii) comprises an NSI in a sense orientation and a COIN in an antisense orientation. See, e.g., U.S. Patent Application Publication No. 2011 / 0104799.
[0103] The nucleic acids and expression constructs may also include a polynucleotide encoding a selection marker. Alternatively, the nucleic acids and expression constructs may lack a polynucleotide encoding a selection marker. The selection marker may be included in a selection cassette. Optionally, the selection cassette may be a self-deletion cassette. See, for example, U.S. Pat. No. 8,697,851 and U.S. Patent Application Publication No. 2013 / 0312129, each of which is incorporated by reference in its entirety for all purposes. By way of example, the self-deletion cassette may include a Crei gene (comprising two exons encoding Cre recombinase separated by an intron) operably linked to the mouse Prm1 promoter and a neomycin resistance gene operably linked to a human ubiquitin promoter. By using the Prm1 promoter, the self-deletion cassette can be deleted specifically in the male germ cells of the F0 animal. Exemplary selection markers include neomycin phosphotransferase (neomycin phosphotransferase) and the neomycin resistant gene. r ), hygromycin B phosphotransferase (hyg r ), puromycin-N-acetyltransferase (puro r ), blasticidin S deaminase (bsr r Selectable markers include guanine phosphoribosyltransferase (GPT), xanthine / guanine phosphoribosyltransferase (GPT), or herpes simplex virus thymidine kinase (HSV-K), or a combination thereof. The polynucleotide encoding the selectable marker may be operably linked to a promoter active in the targeted cells. Examples of promoters are described elsewhere herein.
[0104] The nucleic acid or expression construct may also contain a reporter gene. Exemplary reporter genes include those encoding luciferase, β-galactosidase, green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, Cerulean, T-Sapphire, and alkaline phosphatase. Such reporter genes can be operably linked to a promoter active in the targeted cells. Examples of promoters are described elsewhere herein.
[0105] IV. Vector Also provided herein are vectors containing nucleic acids, nucleic acid constructs, or expression constructs encoding TDP-43 variants. Vectors may contain additional sequences, such as, for example, origins of replication, promoters, and genes encoding antibiotic resistance.
[0106] Some vectors may be circular. Alternatively, vectors may be linear. Vectors may be packaged for delivery via lipid nanoparticles, liposomes, non-lipid nanoparticles, or viral capsids. Non-limiting exemplary vectors include plasmids, phagemids, cosmids, artificial chromosomes, minichromosomes, transposons, viral vectors, and expression vectors.
[0107] The nucleic acid or expression construct can be in a vector, such as a viral vector. The viral vector can be, for example, an adeno-associated virus (AAV) vector or a lentivirus (LV) vector (i.e., a recombinant AAV vector or a recombinant LV vector). Other exemplary viruses / viral vectors include retroviruses, adenoviruses, vaccinia viruses, poxviruses, and herpes simplex viruses. Viruses can infect dividing cells, non-dividing cells, or both dividing and non-dividing cells. Viruses can integrate into the host genome, or alternatively, do not integrate into the host genome. Such viruses can also be engineered to reduce immunity. Viruses can be replication-competent or replication-deficient (e.g., defective in one or more genes required for additional rounds of virion replication and / or packaging). Viruses can cause transient expression, long-term expression (e.g., at least 1 week, 2 weeks, 1 month, 2 months, or 3 months), or persistent expression. Exemplary viral titers (e.g., AAV titers) include titers of about 10 12 , about 10 13 , about 10 14 , about 10 15 , and about 10 16 Other exemplary viral titers (e.g., AAV titers) include about 10 vector genomes / mL. 12 , about 10 13 , about 10 14 , about 10 15 , and about 10 16 Vector genome (vg) / kg body weight is included.
[0108] In one example, the nucleic acid or expression construct is in an AAV vector. AAV may be of any suitable serotype and may be single-stranded (ssAAV) or self-complementary (scAAV). The ssDNA AAV genome consists of two open reading frames, Rep and Cap, flanked by two inverted terminal repeats that allow for the synthesis of complementary DNA strands. When constructing an AAV transfer plasmid, the transgene is placed between the two ITRs, and Rep and Cap can be supplied in trans. In addition to Rep and Cap, AAV may require a helper plasmid containing genes from adenovirus. These genes (E4, E2a, and VA) mediate AAV replication. For example, the transfer plasmid, Rep / Cap, and helper plasmid can be transfected into HEK293 cells containing the adenovirus gene E1+ to generate infectious AAV particles. Alternatively, the Rep, Cap, and adenoviral helper genes may be combined into a single plasmid. Similar packaging cells and methods can be used for other viruses, such as retroviruses.
[0109] Several AAV serotypes have been identified. These serotypes differ in the type of cells they infect (i.e., their tropism), allowing for preferential transduction of specific cell types. Serotypes for CNS tissues include AAV1, AAV2, AAV4, AAV5, AAV8, and AAV9. The selectivity of AAV serotypes for gene delivery in neurons is described, for example, in Hammond et al. (2017) PLoS One 12(12):e0188830, the entire contents of which are incorporated herein by reference for all purposes. In a specific example, the AAV-PHP.eB vector is used. The AAV-PHP.eB vector exhibits a high ability to cross the blood-brain barrier, increasing its CNS transduction efficiency. In another specific example, the AAV9 vector is used. Serotypes for use in skeletal muscle transduction include, for example, AAV1, AAV2, AAV6, and AAV9. See, e.g., Riaz et al. (2015) Skeletal Muscle Vol. 5, Article 37, incorporated herein by reference in its entirety for all purposes.
[0110] Tropism can be further refined through pseudotyping, which is a mixture of capsids and genomes from different viral serotypes. For example, AAV2 / 5 refers to a virus containing a serotype 2 genome packaged in a serotype 5 capsid. The use of pseudotyped viruses not only improves transduction efficiency but can also alter tropism. Hybrid capsids derived from different serotypes can also be used to alter viral tropism. For example, AAV-DJ contains hybrid capsids from eight serotypes and exhibits high infectivity across a wide range of cell types in vivo. AAV-DJ8 is another example that exhibits the properties of AAV-DJ but with enhanced brain uptake. AAV serotypes can also be modified by mutations. Examples of mutational modifications in AAV2 include Y444F, Y500F, Y730F, and S662V. Examples of mutational modifications in AAV3 include Y705F, Y731F, and T492V. Examples of mutational modifications of AAV6 include S663V and T492V. Other pseudotyped / modified AAV variants include AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, AAV8.2, and AAV / SASTG.
[0111] To accelerate transgene expression, self-complementary AAV (scAAV) variants can be used. Because AAV relies on the cell's DNA replication machinery to synthesize the complementary strand of its single-stranded DNA genome, transgene expression can be delayed. To address this delay, scAAV can be used, which contain complementary sequences that can spontaneously anneal upon infection, eliminating the need for host cell DNA synthesis. However, single-stranded AAV (ssAAV) vectors can also be used.
[0112] To increase packaging capacity, a long transgene can be split between two AAV transfer plasmids, one containing a 3' splice donor and the other a 5' splice acceptor. Upon co-infection of cells, these viruses can form concatemers that can be spliced together to express the full-length transgene. This allows for expression of longer transgenes, but at a reduced efficiency. A similar method for increasing capacity utilizes homologous recombination. For example, the transgene can be split between two transfer plasmids, but with substantial sequence overlap, such that co-expression induces homologous recombination and expression of the full-length transgene.
[0113] V. Lipid Nanoparticles Also provided herein are lipid nanoparticles comprising a TDP-43 variant or nucleic acid, nucleic acid construct, expression construct, or vector encoding a TDP-43 variant.
[0114] Lipid formulations can protect biomolecules from degradation while improving cellular uptake. Lipid nanoparticles are particles containing multiple lipid molecules physically associated with each other through intermolecular forces. These include microspheres (including unilamellar and multilamellar vesicles, e.g., liposomes), the dispersed phase in an emulsion, micelles, or the internal phase in a suspension. Such lipid nanoparticles can be used to encapsulate one or more nucleic acids or proteins for delivery. Formulations containing cationic lipids are useful for delivering polyanions such as nucleic acids. Other lipids that can be included are neutral lipids (i.e., uncharged or zwitterionic lipids), anionic lipids, helper lipids that enhance transfection, and stealth lipids that increase the length of time the nanoparticles can persist in vivo. Examples of suitable cationic lipids, neutral lipids, anionic lipids, helper lipids, and stealth lipids can be found in International Publication No. WO 2016 / 010840 A1, which is incorporated herein by reference in its entirety for all purposes. Exemplary lipid nanoparticles may include a cationic lipid and one or more other components. In one example, the other components may include a helper lipid such as cholesterol. In another example, the other components may include a helper lipid such as cholesterol and a neutral lipid such as DSPC. In another example, the other components may include a helper lipid such as cholesterol, an optional neutral lipid such as DSPC, and a stealth lipid such as S010, S024, S027, S031, or S033.
[0115] LNPs may contain one or more or all of the following: (i) lipids for encapsulation and endosomal escape, (ii) neutral lipids for stabilization, (iii) helper lipids for stabilization, and (iv) stealth lipids. See, e.g., Finn et al. (2018) Cell Rep. 22(9):2227-2235 and WO 2017 / 173054(A1), each of which is incorporated herein by reference in its entirety for all purposes. For a specific example of using LNPs for brain delivery, see Nabhan et al. (2016) Sci. Rep. 6:20019, which is incorporated herein by reference in its entirety for all purposes.
[0116] VI. Composition Also provided herein are compositions comprising a TDP-43 variant or nucleic acid, nucleic acid construct, expression construct, vector, or lipid nanoparticle disclosed herein. Such compositions can be, for example, for use in administering a TDP-43 variant to a cell or a subject, or for use in expressing a TDP-43 variant in a cell or a subject. Such compositions can be, for example, for use in inhibiting or reducing TDP-43 aggregation in a cell or a subject. Such compositions can be, for example, for use in rescuing aberrant TDP-43 splicing regulation in a cell or a subject. Such compositions can be, for example, for use in rescuing aberrant subcellular distribution of endogenous TDP-43 in a cell or a subject. Such compositions can be, for example, for use in treating a TDP-43 proteinopathy in a subject. Such compositions can be, for example, for use in preventing a TDP-43 proteinopathy in a subject.
[0117] VII. Cells or Animals Also provided are cells or subjects (e.g., animals or non-human animals) comprising a TDP-43 variant or nucleic acid, nucleic acid construct, expression construct, vector, or lipid nanoparticle disclosed herein. The cells or subjects can express a TDP-43 variant.
[0118] The cell or subject may be, for example, a mammal, a non-human mammal, or a human. A mammal may be, for example, a non-human mammal, a human, a rodent, a rat, a mouse, or a hamster. Other non-human mammals include, for example, non-human primates, monkeys, apes, cats, dogs, rabbits, horses, bulls, deer, bison, and livestock (e.g., bovine species such as cows and steers; ovine species such as sheep and goats; and porcine species such as pigs and wild boars). The term "non-human" excludes humans.
[0119] The cell can be an isolated cell (e.g., in vitro) or can be in vivo within a subject (e.g., an animal or mammal). The cell can also be in any type of undifferentiated or differentiated state. In one example, the cell is a neuron. In one example, the cell is a glial cell. In one example, the cell is a muscle cell.
[0120] The cells provided herein can be normal, healthy cells, or diseased cells containing TDP-43 aggregates or abnormal TDP-43 function (e.g., abnormal TDP-43 splicing regulation or abnormal TDP-43 subcellular localization). The cells can, for example, be susceptible to TDP-43 aggregation, or they can have pre-existing TDP-43 aggregation.
[0121] In one example, the cell is a human cell, a rodent cell, a mouse cell, or a rat cell, e.g., a human neuron, a rodent neuron, a mouse neuron, or a rat neuron, or a human glial cell, a rodent glial cell, a mouse glial cell, or a rat glial cell, or a human muscle cell, a rodent muscle cell, a mouse muscle cell, or a rat muscle cell. In a specific example, the cell is a human neuron. In a specific example, the cell is a human glial cell. In a specific example, the cell is a human muscle cell. In a specific example, the cell is in vivo in a subject (e.g., a neuron in the brain of a subject, or a glial cell in the brain of a subject, or a muscle cell of a subject).
[0122] In some such cells or animals, endogenous TDP-43 is not expressed in the cell or animal. For example, the endogenous TARDBP genomic locus may contain a mutation that prevents expression of endogenous TDP-43 in the cell or animal. Similarly, the cell or animal may contain an agent that reduces or eliminates expression of endogenous TDP-43 in the cell. Such agents are described in more detail below. For example, the agent may include an antisense oligonucleotide or RNAi agent that targets endogenous TARDBP messenger RNA, or a nucleic acid encoding an antisense oligonucleotide or RNAi agent. The agent may include a nuclease agent that targets the endogenous TARDBP genomic locus, or one or more nucleic acids encoding a nuclease agent. For example, the nuclease agent may be a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), or a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein and guide RNA. In a particular example, the nuclease agent is a Cas protein and a guide RNA (e.g., a Cas9 protein and a guide RNA). Examples of other suitable agents are described in more detail elsewhere herein.
[0123] Some cells or animals have a genetically modified endogenous TARDBP genomic locus, and a nucleic acid encoding a TDP-43 variant is integrated at the endogenous TARDBP genomic locus. The cells or animals can be heterozygous or homozygous for the integrated nucleic acid. The integrated nucleic acid can be operably linked to an endogenous TARDBP promoter or an exogenous promoter. In one example, the TDP-43 variant is expressed from the endogenous TARDBP genomic locus and replaces expression of endogenous TDP-43.
[0124] Such cells or animals can have any of the phenotypes described herein associated with the TDP-43 variant, for example, the cells or animals can have reduced TDP-43 aggregation compared to control cells that do not contain the TDP-43 variant or nucleic acid.
[0125] Such cells or animals can be produced by any suitable method. For example, the method of producing such cells or animals can include administering a TDP-43 variant or nucleic acid to the cells or animals. Suitable administration methods are described in more detail elsewhere herein.
[0126] As disclosed elsewhere herein, various methods are provided for generating non-human animal genomes, non-human animal cells, or non-human animals containing a genetically modified endogenous TARDBP genomic locus. Any convenient method or protocol for producing genetically modified organisms is suitable for producing such genetically modified non-human animals. See, for example, Cho et al. (2009) Current Protocols in Cell Biology 42:19.11:19.11.1-19.11.22 and Gama Sosa et al. (2010) Brain Struct. Funct. 214(2-3):91-109, each of which is incorporated by reference in its entirety for all purposes. Such genetically modified non-human animals can be generated, for example, through gene knock-in at a targeted TARDBP locus.
[0127] For example, a method for producing a non-human animal comprising a genetically modified endogenous TARDBP genomic locus includes (1) modifying the genome of a pluripotent cell to comprise the genetically modified endogenous TARDBP genomic locus, (2) identifying or selecting a genetically modified pluripotent cell comprising the genetically modified endogenous TARDBP genomic locus, (3) introducing the genetically modified pluripotent cell into a non-human animal host embryo, and (4) gestation of the host embryo in a surrogate mother. Optionally, the host embryo comprising the modified pluripotent cell (e.g., non-human ES cell) can be implanted in a surrogate mother and incubated to the blastocyst stage before gestation to produce an F0 non-human animal. The surrogate mother can then produce an F0 generation non-human animal comprising the genetically modified endogenous TARDBP genomic locus.
[0128] The method can further include identifying a cell or animal that has the modified target genomic locus. A variety of methods can be used to identify cells and animals that have the targeted genetic modification.
[0129] The step of modifying the genome can, for example, utilize an exogenous donor nucleic acid (e.g., a targeting vector) to modify the TARDBP locus to contain a genetically modified endogenous TARDBP genomic locus as disclosed herein. As an example, the targeting vector can be for generating a genetically modified endogenous TARDBP genomic locus, where the targeting vector includes a 5' homology arm targeting a 5' target sequence of the endogenous TARDBP locus and a 3' homology arm targeting a 3' target sequence of the endogenous TARDBP locus. The exogenous donor nucleic acid can also include a nucleic acid insert comprising a segment of DNA to be integrated into the locus. Integration of the nucleic acid insert at the TARDBP locus can result in the addition of a nucleic acid sequence of interest at the TARDBP locus, a deletion of a nucleic acid sequence of interest at the TARDBP locus, or a replacement (i.e., deletion and insertion) of a nucleic acid sequence of interest at the TARDBP locus. The homology arms can flank an insert nucleic acid that includes a nucleic acid encoding a TDP-43 variant to generate a genetically modified endogenous TARDBP genomic locus.
[0130] The exogenous donor nucleic acid can be for insertion via non-homologous end joining or homologous recombination. The exogenous donor nucleic acid can comprise deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), which can be single-stranded or double-stranded, and can be in linear or circular form. For example, the repair template can be a single-stranded oligodeoxynucleotide (ssODN).
[0131] The exogenous donor nucleic acid can also contain heterologous sequences that are not present in the untargeted endogenous TARDBP locus. For example, the exogenous donor nucleic acid can contain a selection cassette, such as a selection cassette flanked by recombinase recognition sites.
[0132] Some exogenous donor nucleic acids include homology arms. When the exogenous donor nucleic acid also includes a nucleic acid insert, the homology arms can flank the nucleic acid insert. For ease of reference, the homology arms are referred to herein as 5' and 3' (i.e., upstream and downstream) homology arms. This term refers to the relative position of the homology arms to the nucleic acid insert within the exogenous donor nucleic acid. The 5' and 3' homology arms correspond to regions within the TARDBP locus, and are referred to herein as the "5' target sequence" and "3' target sequence," respectively.
[0133] A homology arm and a target sequence "correspond" or "correspond" to one another if the two regions share a sufficient level of sequence identity with each other to act as substrates for a homologous recombination reaction. The term "homology" includes DNA sequences that are identical to or share sequence identity with the corresponding sequence. The sequence identity between a given target sequence and the corresponding homology arm found in the exogenous donor nucleic acid can be any degree of sequence identity that allows homologous recombination to occur. For example, the amount of sequence identity shared by the homology arms of the exogenous donor nucleic acid (or fragment thereof) and the target sequence (or fragment thereof) can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, allowing the sequences to undergo homologous recombination. Furthermore, the corresponding regions of homology between the homology arms and the corresponding target sequence can be of any length sufficient to promote homologous recombination. In some targeting vectors, the intended mutation of the endogenous TARDBP locus is contained in the insert nucleic acid flanked by the homology arms.
[0134] In cells other than one-cell stage embryos, the exogenous donor nucleic acid can be a "large targeting vector" or "LTVEC," which includes targeting vectors containing homology arms corresponding to and derived from larger nucleic acid sequences than those typically used in other approaches aimed at achieving homologous recombination in cells. LTVEC also includes targeting vectors containing nucleic acid inserts with larger nucleic acid sequences than those typically used in other approaches aimed at achieving homologous recombination in cells. For example, LTVEC enables the modification of large loci that cannot be accommodated by conventional plasmid-based targeting vectors due to size limitations. For example, the targeted locus can be a cellular locus (i.e., the 5' and 3' homology arms can correspond) that cannot be targeted using conventional methods or that can be targeted only imprecisely or with significantly lower efficiency in the absence of a nick or double-strand break induced by a nuclease agent (e.g., a Cas protein). LTVECs can be of any length, typically at least 10 kb in length. The combined length of the 5' and 3' homology arms of the LTVEC is typically at least 10 kb.
[0135] The screening step can include, for example, a quantitative assay to evaluate the allelic modification (MOA) of the parent chromosome. For example, the quantitative assay can be performed via quantitative PCR, such as real-time PCR (qPCR). The real-time PCR can utilize a first primer set that recognizes the target locus and a second primer set that recognizes a non-target reference locus. The primer set can include a fluorescent probe that recognizes the amplified sequence.
[0136] Other examples of suitable quantitative assays include fluorescence-mediated in situ hybridization (FISH), comparative genomic hybridization, isothermal DNA amplification, quantitative hybridization to immobilized probes, INVADER® probes, TAQMAN® molecular beacon probes, or ECLIPSE™ probe technology (see, e.g., U.S. Patent Application Publication No. 2005 / 0144655, which is incorporated by reference in its entirety for all purposes).
[0137] An example of a suitable pluripotent cell is an embryonic stem (ES) cell (e.g., a mouse ES cell or a rat ES cell). Modified pluripotent cells can be generated by recombination, for example, by (a) introducing into a cell one or more exogenous donor nucleic acids (e.g., targeting vectors) containing an insert nucleic acid, e.g., an insert nucleic acid flanked by 5' and 3' homology arms corresponding to 5' and 3' target sites, where the insert nucleic acid comprises a nucleic acid encoding a TDP-43 variant such that a genetically modified endogenous TARDBP genomic locus is generated, and (b) identifying at least one cell containing the insert nucleic acid integrated into its genome at the endogenous TARDBP locus (i.e., identifying at least one cell containing a genetically modified endogenous TARDBP genomic locus).
[0138] Alternatively, modified pluripotent cells can be generated by (a) introducing into a cell (i) a nuclease agent or a nucleic acid encoding the nuclease agent, where the nuclease agent induces a nick or double-stranded break at a target site within the endogenous TARDBP locus, and (ii) one or more exogenous donor nucleic acids (e.g., targeting vectors) containing an insert nucleic acid, e.g., an insert nucleic acid flanked by 5' and 3' homology arms corresponding to the 5' and 3' target sites located sufficiently close to the nuclease target site, where the insert nucleic acid comprises a nucleic acid encoding a TDP-43 variant such that a genetically modified endogenous TARDBP genomic locus is generated, and (c) identifying at least one cell containing the insert nucleic acid integrated into its genome at the endogenous TARDBP locus (i.e., identifying at least one cell containing a genetically modified endogenous TARDBP genomic locus). Any nuclease agent that induces a nick or double-stranded break at the desired recognition site can be used. Examples of suitable nucleases include transcription activator-like effector nucleases (TALENs), zinc finger nucleases (ZFNs), meganucleases, and clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) systems (e.g., CRISPR / Cas9 systems) or components of such systems (e.g., CRISPR / Cas9). See, for example, U.S. Patent Application Publication No. 2013 / 0309670 and U.S. Patent Application Publication No. 2015 / 0159175, each of which is incorporated herein by reference in its entirety for all purposes.
[0139] The donor cells can be introduced into the host embryo at any stage, such as the blastocyst stage or pre-morula stage (i.e., 4-cell or 8-cell stage). Offspring are generated that can transmit the genetic modification through the germline. See, e.g., U.S. Patent No. 7,294,754, which is incorporated herein by reference in its entirety for all purposes.
[0140] Alternatively, the methods of producing a non-human animal described elsewhere herein may include (1) modifying the genome of a one-cell stage embryo to contain a genetically modified endogenous TARDBP genomic locus, (2) selecting the genetically modified embryo, and (3) gestation of the genetically modified embryo in a surrogate mother, producing offspring capable of transmitting the genetic modification through the germline.
[0141] Nuclear transfer techniques can also be used to generate non-human mammals. Briefly, methods for nuclear transfer can include the steps of: (1) enucleating an oocyte or providing an enucleated oocyte; (2) isolating or providing a donor cell or nucleus combined with the enucleated oocyte; (3) inserting the cell or nucleus into the enucleated oocyte to form a reconstituted cell; (4) implanting the reconstituted cell in the uterus of an animal to form an embryo; and (5) allowing the embryo to develop. In such methods, oocytes are typically recovered from deceased animals, but can also be isolated from the oviducts and / or ovaries of living animals. Oocytes can be matured in a variety of well-known media before enucleation. Enucleation of oocytes can be performed by many well-known methods. Insertion of a donor cell or nucleus into an enucleated oocyte to form a reconstituted cell can be performed by microinjecting the donor cell beneath the zona pellucida prior to fusion. Fusion can be induced by applying a DC electric pulse across the contact / fusion surface (electrofusion), exposing the cells to fusion-promoting chemicals such as polyethylene glycol, or by inactivated viruses such as Sendai virus. Reconstituted cells can be activated by electrical and / or non-electrical means before, during, and / or after fusion of the nuclear donor and recipient oocytes. Activation methods include electrical pulses, chemically induced shock, penetration with sperm, increasing the level of divalent cations in the oocyte, and reducing the phosphorylation of cellular proteins in the oocyte (with kinase inhibitors). Activated reconstituted cells, or embryos, can be cultured in known media and then implanted into the uterus of an animal. See, for example, U.S. Patent Application Publication No. 2008 / 0092249, WO 1999 / 005266, U.S. Patent Application Publication No. 2004 / 0177390, WO 2008 / 017234, and U.S. Patent No. 7,612,250, each of which is incorporated by reference herein in its entirety for all purposes.
[0142] Various methods provided herein enable the generation of genetically modified non-human F0 animals, wherein the cells of the genetically modified F0 animals contain a genetically modified endogenous TARDBP genomic locus. It is recognized that the number of cells within the F0 animal that contain the genetically modified endogenous TARDBP genomic locus will vary depending on the method used to generate the F0 animal. For example, introduction of donor ES cells into pre-morula stage embryos (e.g., 8-cell stage mouse embryos) from a corresponding organism via the VELOCIMOUSE® method allows a greater percentage of the F0 animal's cell population to contain cells that contain the nucleotide sequence of interest, including the targeted genetic modification. For example, at least 50%, 60%, 65%, 70%, 75%, 85%, 86%, 87%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the cellular contribution of the non-human F0 animal can comprise a cell population with a targeted modification.
[0143] The cells of the genetically modified F0 animal can be heterozygous for the genetically modified endogenous TARDBP genomic locus or can be homozygous for the genetically modified endogenous TARDBP genomic locus.
[0144] VIII. Method Provided herein are methods for administering a TDP-43 variant described herein to a cell or a subject, or administering a nucleic acid encoding a TDP-43 variant to a cell or a subject, such that the TDP-43 variant is expressed. Methods for reducing or inhibiting TDP-43 aggregation in a cell or a subject, rescuing aberrant TDP-43 splicing regulation in a cell or a subject, or rescuing aberrant intracellular distribution of TDP-43 in a cell or a subject are also provided. Methods for treating a TDP-43 proteinopathy in a subject or preventing a TDP-43 proteinopathy in a subject are also provided. Such methods may include administering a TDP-43 variant described herein (e.g., a therapeutically effective amount) to a cell or a subject, or administering a nucleic acid encoding a TDP-43 variant (e.g., a therapeutically effective amount) to a cell or a subject, such that the TDP-43 variant is expressed. A therapeutically effective amount is the amount that produces the desired effect for which it is administered. The exact amount will depend on the purpose of the treatment and can be ascertained by one of skill in the art using known techniques. See, e.g., Lloyd (1999) The Art, Science and Technology of Pharmaceutical Compounding.
[0145] Some such methods include administering a TDP-43 variant to a cell or a subject. Some such methods include administering a nucleic acid encoding the TDP-43 variant to a cell or a subject. The nucleic acid may be a nucleic acid construct, as described in more detail elsewhere herein. In some cases, the nucleic acid encoding the TDP-43 variant may be codon-optimized (e.g., codon-optimized for expression in humans or mice). For example, the nucleic acid may be modified to alternative codons more frequently used in human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest.
[0146] The nucleic acid encoding a TDP-43 variant may be DNA or RNA. The nucleic acid may optionally be messenger RNA (mRNA) encoding a TDP-43 variant. The nucleic acid may optionally be complementary DNA (cDNA) encoding a TDP-43 variant. For example, such a nucleic acid may include only a coding sequence without any intervening introns. Examples of coding sequences for some of the TDP-43 variants described herein are set forth, for example, in SEQ ID NOS: 70-72. In other cases, the nucleic acid may include one or more introns separating exons in the TDP-43 variant coding sequence. For example, the nucleic acid may include a TDP-43 variant genomic sequence including both exons and introns.
[0147] In some methods, the nucleic acid is in an expression construct comprising a nucleic acid encoding a TDP-43 variant operably linked to a promoter. The promoter can be any suitable promoter for in vivo expression in an animal or in vitro expression in an isolated cell. The promoter can be a constitutively active promoter (e.g., a CAG promoter or a U6 promoter), a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific promoter or a tissue-specific promoter). Such promoters are well known and are discussed elsewhere herein. In a specific example, the promoter is active in neurons. In a specific example, the promoter is active in glial cells. In a specific example, the promoter is active in muscle cells. In some cases, the promoter is a heterologous promoter (i.e., a promoter to which the TDP-43 nucleic acid is not naturally operably linked). In other cases, the promoter may be an endogenous promoter (i.e., a TDP-43 variant nucleic acid operably linked to a TARDBP promoter). The heterologous promoter may be any type of promoter disclosed elsewhere herein. For example, the promoter may be a constitutive promoter, such as the EF1α promoter. Alternatively, the promoter may be a tissue-specific promoter or an inducible promoter. For example, the promoter may be a neuron-specific promoter. An example of a suitable neuron-specific promoter that is highly specific with low levels of expression is the synapsin-1 promoter (e.g., the human synapsin-1 promoter). For example, the promoter may be a glia-specific promoter. For example, the promoter may be a muscle-specific promoter.
[0148] The nucleic acids and expression constructs disclosed herein can also include post-transcriptional regulatory elements, such as the woodchuck hepatitis virus post-transcriptional regulatory elements. For example, the promoter can be a glia-specific promoter.
[0149] The nucleic acid and expression construct may further comprise one or more polyadenylation signal sequences. For example, the nucleic acid construct may comprise a polyadenylation signal sequence located 3' of the TDP-43 variant coding sequence. Any suitable polyadenylation signal sequence may be used. The term polyadenylation signal sequence refers to any sequence that directs the termination of transcription and the addition of a poly(A) tail to an mRNA transcript. In eukaryotes, transcription terminators are recognized by protein factors, and termination is followed by polyadenylation, the process of adding a poly(A) tail to an mRNA transcript in the presence of poly(A) polymerase. Mammalian poly(A) signals typically consist of a core sequence approximately 45 nucleotides long, which may be flanked by various auxiliary sequences that help increase the efficiency of cleavage and polyadenylation. The core sequence, called the polyA recognition motif or polyA recognition sequence, consists of a highly conserved upstream element (AATAAA or AAUAAA) in mRNA that is recognized by the cleavage and polyadenylation specificity factor (CPSF), and a poorly defined downstream region (rich in U or G and U) that is bound by the cleavage stimulation factor (CstF). Examples of transcription terminators that can be used include, for example, the human growth hormone (HGH) polyadenylation signal, the simian virus 40 (SV40) late polyadenylation signal, the rabbit beta globin polyadenylation signal, the bovine growth hormone (BGH) polyadenylation signal, the phosphoglycerate kinase (PGK) polyadenylation signal, the AOX1 transcription termination sequence, the CYC1 transcription termination sequence, or any transcription termination sequence known to be suitable for regulating gene expression in eukaryotic cells.
[0150] The nucleic acid and expression construct may also optionally include a polyadenylation signal sequence upstream of the TDP-43 variant coding sequence. The polyadenylation signal sequence upstream of the TDP-43 variant coding sequence may be adjacent to a recombinase recognition site recognized by a site-specific recombinase. In some constructs, the recombinase recognition site also flanks a selection cassette containing, for example, a coding sequence for a drug resistance protein. In some constructs, the recombinase recognition site does not flank the selection cassette. The polyadenylation signal sequence prevents transcription and expression of the protein or RNA encoded by the coding sequence. However, upon exposure to a site-specific recombinase, the polyadenylation signal sequence may be excised, allowing the protein or RNA to be expressed.
[0151] Such a configuration can allow for tissue- or developmental stage-specific expression if the polyadenylation signal sequence is excised in a tissue- or developmental stage-specific manner. Excision of the polyadenylation signal sequence in a tissue- or developmental stage-specific manner can be achieved if the animal containing the nucleic acid or expression construct further contains a coding sequence for a site-specific recombinase operably linked to a tissue- or developmental stage-specific promoter. The polyadenylation signal sequence is then excised only in those tissues or those developmental stages, allowing for tissue- or developmental stage-specific expression. In one example, the TDP-43 variant encoded by the nucleic acid or expression construct can be expressed in a neuron-specific manner. In one example, the TDP-43 variant encoded by the nucleic acid or expression construct can be expressed in a glia- or developmental stage-specific manner. In one example, the TDP-43 variant encoded by the nucleic acid or expression construct can be expressed in a muscle-specific manner.
[0152] Site-specific recombinases include enzymes that can promote recombination between recombinase recognition sites where the two recombination sites are physically separated within a single nucleic acid or on separate nucleic acids. Examples of recombinases include Cre, Flp, and Dre recombinases. One example of a Cre recombinase gene is Crei, in which the two exons encoding Cre recombinase are separated by an intron, preventing expression in prokaryotic cells. Such recombinases may further contain a nuclear localization signal to promote nuclear localization (e.g., NLS-Crei). Recombinase recognition sites include nucleotide sequences that are recognized by site-specific recombinases and can serve as substrates for recombination events. Examples of recombinase recognition sites include FRT, FRT11, FRT71, attp, att, rox, and lox sites such as loxP, lox511, lox2272, lox66, lox71, loxM2, and lox5171.
[0153] The nucleic acids disclosed herein may comprise deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), which may be single-stranded or double-stranded, and which may be in linear or circular form. Nucleic acid constructs may be naked nucleic acids, as described elsewhere herein, or may be delivered by vectors such as AAV vectors. When in linear form, the ends of the nucleic acid may be protected (e.g., from exonuclease degradation) by well-known methods. For example, one or more dideoxynucleotide residues may be added to the 3' end of the linear molecule, and / or self-complementary oligonucleotides may be attached to one or both ends. See, e.g., Chang et al. (1987) Proc. Natl. Acad. Sci. USA 84:4959-4963 and Nehls et al. (1996) Science 272:886-889, each of which is incorporated herein by reference in its entirety for all purposes. Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino group(s) and the use of modified internucleotide linkages such as phosphorothioates, phosphoramidates, and O-methylribose or deoxyribose residues. The nucleic acid or expression construct may optionally include one or more of the following terminal structures: a hairpin, a loop, an inverted terminal repeat (ITR), or a toroid. For example, the nucleic acid or expression construct may include an ITR.
[0154] The nucleic acid or expression construct may include modifications or sequences that provide additional desirable characteristics (e.g., altered or controlled stability, tracking or detection using fluorescent labels, binding sites for proteins or protein complexes, etc.). For example, modifications may be made to one or more nucleosides within the mRNA. Examples of chemical modifications to mRNA nucleobases include pseudouridine, 1-methyl-pseudouridine, and 5-methyl-cytidine. The mRNA may also be capped. The mRNA may also be polyadenylated (to include a poly(A) tail). As an example, a capped polyadenylated mRNA containing N1-methyl-pseudouridine can be used (e.g., completely replaced with N1-methyl-pseudouridine). The nucleic acid construct may include one or more fluorescent labels, purification tags, epitope tags, or combinations thereof. For example, the nucleic acid construct may include one or more fluorescent labels (e.g., fluorescent proteins or other fluorophores or dyes), such as at least one, at least two, at least three, at least four, or at least five fluorescent labels. Exemplary fluorescent labels include fluorophores such as fluorescein (e.g., 6-carboxyfluorescein (6-FAM)), Texas Red, HEX, Cy3, Cy5, Cy5.5, Pacific Blue, 5-(and-6)-carboxytetramethylrhodamine (TAMRA), and Cy7. A wide range of fluorescent dyes for labeling oligonucleotides are commercially available (e.g., from Integrated DNA Technologies). The label or tag can be at the 5' end, 3' end, or internally located on the nucleic acid construct. For example, the nucleic acid construct can be conjugated to the 5' end with the IR700 fluorophore (5'IRDYE® 700) from Integrated DNA Technologies.
[0155] The nucleic acids and expression constructs may also include conditional alleles. The conditional allele may be a multifunctional allele, as described in U.S. Patent Application Publication No. 2011 / 0104799, which is incorporated herein by reference in its entirety for all purposes. For example, the conditional allele may include: (a) an actuating sequence in the sense orientation relative to transcription of the target gene; (b) a drug selection cassette (DSC) in the sense or antisense orientation; (c) a nucleotide sequence of interest (NSI) in the antisense orientation; and (d) a conditional inversion module (COIN, which utilizes an exon-splitting intron and a reversible gene trap-like module) in the reverse orientation. See, for example, U.S. Patent Application Publication No. 2011 / 0104799. The conditional allele may further include recombinable units that recombine upon exposure to a first recombinase to form a conditional allele that (i) lacks the actuating sequence and DSC, and (ii) includes the NSI in the sense orientation and the COIN in the antisense orientation. See, for example, U.S. Patent Application Publication No. 2011 / 0104799.
[0156] The nucleic acids and expression constructs may also include a polynucleotide encoding a selection marker. Alternatively, the nucleic acids and expression constructs may lack a polynucleotide encoding a selection marker. The selection marker may be included in a selection cassette. Optionally, the selection cassette may be a self-deletion cassette. See, for example, U.S. Pat. No. 8,697,851 and U.S. Patent Application Publication No. 2013 / 0312129, each of which is incorporated by reference in its entirety for all purposes. By way of example, the self-deletion cassette may include a Crei gene (comprising two exons encoding Cre recombinase separated by an intron) operably linked to the mouse Prm1 promoter and a neomycin resistance gene operably linked to a human ubiquitin promoter. By using the Prm1 promoter, the self-deletion cassette can be deleted specifically in the male germ cells of the F0 animal. Exemplary selection markers include neomycin phosphotransferase (neomycin phosphotransferase) and the neomycin resistant gene.r ), hygromycin B phosphotransferase (hyg r ), puromycin-N-acetyltransferase (puro r ), blasticidin S deaminase (bsr r ), xanthine / guanine phosphoribosyltransferase (gpt), or herpes simplex virus thymidine kinase (HSV-k), or a combination thereof. The polynucleotide encoding the selectable marker may be operably linked to a promoter active in the targeted cells. Examples of promoters are described elsewhere herein.
[0157] The nucleic acid or expression construct may also contain a reporter gene. Exemplary reporter genes include those encoding luciferase, β-galactosidase, green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, Cerulean, T-Sapphire, and alkaline phosphatase. Such reporter genes can be operably linked to a promoter active in the targeted cells. Examples of promoters are described elsewhere herein.
[0158] The nucleic acid or expression construct can be in a vector, such as a viral vector, which can include additional sequences such as, for example, an origin of replication, a promoter, and genes encoding antibiotic resistance.
[0159] Some vectors may be circular. Alternatively, vectors may be linear. Vectors may be packaged for delivery via lipid nanoparticles, liposomes, non-lipid nanoparticles, or viral capsids. Non-limiting exemplary vectors include plasmids, phagemids, cosmids, artificial chromosomes, minichromosomes, transposons, viral vectors, and expression vectors.
[0160] The nucleic acid or expression construct may be in a vector, such as a viral vector. The viral vector may be, for example, an adeno-associated viral (AAV) vector or a lentiviral (LV) vector (i.e., a recombinant AAV vector or a recombinant LV vector). Other exemplary viruses / viral vectors include retroviruses, adenoviruses, vaccinia viruses, poxviruses, and herpes simplex viruses. Viruses can infect dividing cells, non-dividing cells, or both dividing and non-dividing cells. Viruses may integrate into the host genome, or alternatively, not integrate into the host genome. Such viruses may also be engineered to reduce immunity. Viruses may be replication-competent or replication-deficient (e.g., defective in one or more genes required for additional rounds of virion replication and / or packaging). Viruses may cause transient expression, long-term expression (e.g., at least 1 week, 2 weeks, 1 month, 2 months, or 3 months), or persistent expression. Exemplary viral titers (e.g., AAV titers) include titers of about 10 12 , about 10 13 , about 10 14 , about 10 15 , and about 10 16 Other exemplary viral titers (e.g., AAV titers) include about 10 vector genomes / mL. 12 , about 10 13 , about 10 14 , about 10 15 , and about 10 16 Vector genomes (vg) / kg body weight are included.
[0161] In one example, the nucleic acid or expression construct is in an AAV vector. AAV may be of any suitable serotype and may be single-stranded AAV (ssAAV) or self-complementary AAV (scAAV). The ssDNA AAV genome consists of two open reading frames, Rep and Cap, flanked by two inverted terminal repeats that allow for the synthesis of complementary DNA strands. When constructing an AAV transfer plasmid, the transgene is placed between the two ITRs, and Rep and Cap can be supplied in trans. In addition to Rep and Cap, AAV may require a helper plasmid containing genes from adenovirus. These genes (E4, E2a, and VA) mediate AAV replication. For example, the transfer plasmid, Rep / Cap, and helper plasmid can be transfected into HEK293 cells containing the adenovirus gene E1+ to produce infectious AAV particles. Alternatively, Rep, Cap, and adenovirus helper genes can be combined into a single plasmid. Similar packaging cells and methods can be used for other viruses, such as retroviruses.
[0162] Several AAV serotypes have been identified. These serotypes differ in the type of cells they infect (i.e., their tropism), allowing for preferential transduction of specific cell types. Serotypes for CNS tissues include AAV1, AAV2, AAV4, AAV5, AAV8, and AAV9. The selectivity of AAV serotypes for gene delivery in neurons is described, for example, in Hammond et al. (2017) PLoS One 12(12):e0188830, the entire contents of which are incorporated herein by reference for all purposes. In a specific example, the AAV-PHP.eB vector is used. The AAV-PHP.eB vector exhibits a high ability to cross the blood-brain barrier, increasing its CNS transduction efficiency. In another specific example, the AAV9 vector is used. Serotypes for use in skeletal muscle transduction include, for example, AAV1, AAV2, AAV6, and AAV9. See, e.g., Riaz et al. (2015) Skeletal Muscle Vol. 5, Article 37, incorporated herein by reference in its entirety for all purposes.
[0163] Tropism can be further refined through pseudotyping, which is a mixture of capsids and genomes from different viral serotypes. For example, AAV2 / 5 refers to a virus containing a serotype 2 genome packaged in a serotype 5 capsid. The use of pseudotyped viruses not only improves transduction efficiency but can also alter tropism. Hybrid capsids derived from different serotypes can also be used to alter viral tropism. For example, AAV-DJ contains hybrid capsids from eight serotypes and exhibits high infectivity across a wide range of cell types in vivo. AAV-DJ8 is another example that exhibits the properties of AAV-DJ but with enhanced brain uptake. AAV serotypes can also be modified by mutations. Examples of mutational modifications in AAV2 include Y444F, Y500F, Y730F, and S662V. Examples of mutational modifications in AAV3 include Y705F, Y731F, and T492V. Examples of mutational modifications of AAV6 include S663V and T492V. Other pseudotyped / modified AAV variants include AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, AAV8.2, and AAV / SASTG.
[0164] To accelerate transgene expression, self-complementary AAV (scAAV) variants can be used. Because AAV relies on the cell's DNA replication machinery to synthesize the complementary strand of its single-stranded DNA genome, transgene expression can be delayed. To address this delay, scAAV can be used, which contain complementary sequences that can spontaneously anneal upon infection, eliminating the need for host cell DNA synthesis. However, single-stranded AAV (ssAAV) vectors can also be used.
[0165] To increase packaging capacity, a long transgene can be split between two AAV transfer plasmids, one containing a 3' splice donor and the other a 5' splice acceptor. Upon co-infection of cells, these viruses can form concatemers that can be spliced together to express the full-length transgene. This allows for expression of longer transgenes, but at a reduced efficiency. A similar method for increasing capacity utilizes homologous recombination. For example, the transgene can be split between two transfer plasmids, but with substantial sequence overlap, such that co-expression induces homologous recombination and expression of the full-length transgene.
[0166] In some methods, the TDP-43 variant or a nucleic acid encoding the TDP-43 variant is associated with a lipid nanoparticle. Lipid formulations can protect biomolecules from degradation while improving cellular uptake. Lipid nanoparticles are particles containing multiple lipid molecules physically associated with each other through intermolecular forces. These include microspheres (including unilamellar and multilamellar vesicles, e.g., liposomes), the dispersed phase in an emulsion, micelles, or the internal phase in a suspension. Such lipid nanoparticles can be used to encapsulate one or more nucleic acids or proteins for delivery. Formulations containing cationic lipids are useful for delivering polyanions such as nucleic acids. Other lipids that can be included are neutral lipids (i.e., uncharged or zwitterionic lipids), anionic lipids, helper lipids that enhance transfection, and stealth lipids that increase the length of time the nanoparticles can persist in vivo. Examples of suitable cationic lipids, neutral lipids, anionic lipids, helper lipids, and stealth lipids can be found in WO 2016 / 010840(A1), which is incorporated herein by reference in its entirety for all purposes. Exemplary lipid nanoparticles can include a cationic lipid and one or more other components. In one example, the other components can include a helper lipid such as cholesterol. In another example, the other components can include a helper lipid such as cholesterol and a neutral lipid such as DSPC. In another example, the other components can include a helper lipid such as cholesterol, an optional neutral lipid such as DSPC, and a stealth lipid such as S010, S024, S027, S031, or S033.
[0167] LNPs may contain one or more or all of the following: (i) lipids for encapsulation and endosomal escape, (ii) neutral lipids for stabilization, (iii) helper lipids for stabilization, and (iv) stealth lipids. See, e.g., Finn et al. (2018) Cell Rep. 22(9):2227-2235 and WO 2017 / 173054(A1), each of which is incorporated herein by reference in its entirety for all purposes. For a specific example of using LNPs for brain delivery, see Nabhan et al. (2016) Sci. Rep. 6:20019, which is incorporated herein by reference in its entirety for all purposes.
[0168] The TDP-43 variant or a nucleic acid encoding a TDP-43 variant can be administered to a cell or a subject by any suitable means. Various methods and compositions are provided herein to enable the introduction of molecules (e.g., nucleic acids or proteins) into a cell or a subject.
[0169] The methods provided herein do not depend on a particular method for introducing a nucleic acid or protein into a cell, but only on the ability of the nucleic acid or protein to access the interior of the cell. Methods for introducing nucleic acids and proteins into various cell types are well known in the art and include, for example, stable transfection methods, transient transfection methods, and viral-mediated methods.
[0170] Transfection protocols and protocols for introducing molecules (e.g., nucleic acids or proteins) into cells can vary. Non-limiting transfection methods include chemical transfection methods using liposomes, nanoparticles, and calcium phosphate (Graham et al. (1973) Virology 52(2):456-67; Bacchetti et al. (1977) Proc. Natl. Acad. Sci. USA 74(4):1590-4; and Kriegler, M. (1991). Transfer and Expression: A Laboratory Manual. New York: W.H. Freeman and Company. pp.96-97, each of which is incorporated herein by reference in its entirety for all purposes). Chemical-based transfection methods include dendrimers or cationic polymers such as DEAE-dextran or polyethyleneimine. Non-chemical methods include electroporation, sonoporation, and phototransfection. Particle-based transfection includes the use of a gene gun or magnet-assisted transfection (Bertram (2006) Current Pharmaceutical Biotechnology 7, 277-28, which is incorporated herein by reference in its entirety in this application). Viral methods can also be used for transfection.
[0171] Introduction of molecules (e.g., nucleic acids or proteins) into cells can also be mediated by electroporation, intracytoplasmic injection, viral infection, adenovirus, adeno-associated virus, lentivirus, retrovirus, transfection, lipid-mediated transfection, or nucleofection. Nucleofection is an improved electroporation technique that allows nucleic acid substrates to be delivered not only to the cytoplasm but also through the nuclear membrane to the nucleus. In addition, the use of nucleofection in the methods disclosed herein typically requires far fewer cells than conventional electroporation (e.g., only about 2 million cells compared to 7 million cells for conventional electroporation). In one example, nucleofection is performed using the LONZA® NUCLEOFECTOR™ system.
[0172] Introduction of molecules (e.g., nucleic acids or proteins) into cells can also be achieved by microinjection. Microinjection of mRNA is preferably into the cytoplasm (e.g., to deliver mRNA directly to the translation machinery), while microinjection of proteins or DNA encoding proteins is preferably into the nucleus. Alternatively, microinjection can be performed by injection into both the nucleus and the cytoplasm, but the needle can be first introduced into the nucleus and the first amount can be injected, and the second amount can be injected into the cytoplasm while the needle is removed from the cell. Methods for performing microinjection are well known. See, e.g., Nagy et al. (Nagy A, Gertsenstein M, Wintersten K, Behringer R., 2003, Manipulating the Mouse Embryo. Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press), Meyer et al. (2010) Proc. Natl. Acad. Sci. USA 107:15022-15026, and Meyer et al. (2012) Proc. Natl. Acad. Sci. USA 109:9354-9359, each of which is incorporated by reference in its entirety for all purposes.
[0173] Other methods for introducing molecules (e.g., nucleic acids or proteins) into cells may include, for example, vector delivery, particle-mediated delivery, exosome-mediated delivery, lipid nanoparticle-mediated delivery, cell-penetrating peptide-mediated delivery, or implantable device-mediated delivery. Methods for administering nucleic acids or proteins to a subject to modify cells in vivo are disclosed elsewhere herein. As specific examples, molecules (e.g., nucleic acids or proteins) can be introduced into cells or non-human animals in carriers such as poly(lactic acid) (PLA) microspheres, poly(D,L-lactic-coglycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid cochleates, or lipid microtubules. Some specific examples of delivery to non-human animals include hydrodynamic delivery, viral-mediated delivery (e.g., lentivirus-mediated delivery or adeno-associated virus (AAV)-mediated delivery), and lipid nanoparticle-mediated delivery.
[0174] In one example, the TDP-43 variant or a nucleic acid encoding the TDP-43 variant can be administered via viral transduction, such as lentiviral transduction or adeno-associated viral transduction. In another example, the TDP-43 variant or a nucleic acid encoding the TDP-43 variant can be administered via lipid nanoparticle (LNP)-mediated delivery.
[0175] In vivo administration can be via any suitable route such that the TDP-43 variant or nucleic acid encoding the TDP-43 variant reaches the intended target cells (e.g., neurons in the brain of a subject, and / or glial cells in the brain of a subject, and / or muscle cells of a subject) or target tissue (e.g., brain or muscle). Examples of administration routes include parenteral, intravenous, oral, subcutaneous, intraarterial, intracranial, intrathecal, intraperitoneal, topical, intranasal, or intramuscular. Systemic administration modes include, for example, oral and parenteral routes. Examples of parenteral routes include intravenous, intraarterial, intraosseous, intramuscular, intradermal, subcutaneous, intranasal, and intraperitoneal routes. A specific example is intravenous infusion. Intranasal infusion and intravitreal injection are other specific examples. Local administration modes include, for example, intrathecal, intracerebroventricular, intraparenchymal (e.g., localized intraparenchymal delivery to the striatum (e.g., to the caudate nucleus or putamen), cerebral cortex, precentral gyrus, hippocampus (e.g., dentate gyrus or CA3 region), temporal cortex, amygdala, frontal cortex, thalamus, cerebellum, medulla, thalamus, optic tectum, tegmentum, or substantia nigra), intraocular, intraorbital, subconjunctival, intravitreal, subretinal, and transscleral routes. Significantly smaller amounts of components (compared to systemic approaches) may be effective when administered locally (e.g., intraparenchymal or intravitreal) compared to when administered systemically (e.g., intravenously). Local administration modes may also reduce or eliminate the incidence of potentially toxic side effects that can occur when therapeutically effective amounts of components are administered systemically. For example, a TDP-43 variant or a nucleic acid encoding a TDP-43 variant may be administered directly to the subject's brain or to neurons or glial cells within the subject's brain. In a specific example, administration to a subject is by intrathecal injection or intracranial injection (e.g., stereotactic surgery for injection into the hippocampus and other brain regions, or intraventricular injection). In a specific example, administration to a subject is by intraventricular injection. In another specific example, administration to a subject is by intracranial injection. In another specific example, administration to a subject is by intrathecal injection. In another example, the TDP-43 variant or a nucleic acid encoding a TDP-43 variant can be administered directly to a muscle or to a muscle cell of a subject (e.g., intramuscular injection).
[0176] The frequency and number of doses may depend, among other factors, on the half-life and route of administration of the administered composition. The introduction of the nucleic acid or protein into the cell or non-human animal may be performed one or more times over a period of time. For example, the introduction may be performed at least two times over a period of time, at least three times over a period of time, at least four times over a period of time, at least five times over a period of time, at least six times over a period of time, at least seven times over a period of time, at least eight times over a period of time, at least nine times over a period of time, at least ten times over a period of time, at least eleven times over a period of time, at least 12 times over a period of time, at least 13 times over a period of time, at least 14 times over a period of time, at least 15 times over a period of time, at least 16 times over a period of time, at least 17 times over a period of time, at least 18 times over a period of time, at least 19 times over a period of time, or at least 20 times over a period of time.
[0177] The cell or subject in the method can be, for example, a mammal, a non-human mammal, or a human. The mammal can be, for example, a non-human mammal, a human, a rodent, a rat, a mouse, or a hamster. Other non-human mammals include, for example, non-human primates, monkeys, apes, cats, dogs, rabbits, horses, bulls, deer, bison, and livestock (e.g., bovine species such as cows and steers; ovine species such as sheep and goats; and porcine species such as pigs and wild boars). The term "non-human" excludes humans. In a specific example, the cell or subject is a human.
[0178] The cell can be an isolated cell (e.g., in vitro) or can be in vivo within a subject (e.g., an animal or mammal). The cell can also be in any type of undifferentiated or differentiated state. In one example, the cell is a neuron. In one example, the cell is a glial cell. In one example, the cell is a muscle cell.
[0179] The cells provided herein can be normal, healthy cells, or diseased cells containing TDP-43 aggregates or abnormal TDP-43 function (e.g., abnormal TDP-43 splicing regulation or abnormal TDP-43 subcellular localization). The cells can, for example, be susceptible to TDP-43 aggregation, or they can have pre-existing TDP-43 aggregation.
[0180] In one example, the cell is a human cell, a rodent cell, a mouse cell, or a rat cell, e.g., a human neuron, a rodent neuron, a mouse neuron, or a rat neuron, or a human glial cell, a rodent glial cell, a mouse glial cell, or a rat glial cell, or a human muscle cell, a rodent muscle cell, a mouse muscle cell, or a rat muscle cell. In a specific example, the cell is a human neuron. In a specific example, the cell is a human glial cell. In a specific example, the cell is a human muscle cell. In a specific example, the cell is in vivo in a subject (e.g., a neuron or glial cell in the brain of a subject, or a muscle cell of a subject). For example, such a method can be a method of inhibiting or reducing TDP-43 aggregation, reducing TDP-43 phosphorylation, reducing or rescuing aberrant regulation of splicing by TDP-43, or reducing or rescuing aberrant TDP-43 subcellular localization (e.g., reducing or rescuing aberrant nuclear depletion of TDP-43) in a subject's cell (e.g., a neuron or glial cell in the subject's brain, or a subject's muscle cell).
[0181] The methods described herein may further include administering to a cell or a subject an agent that reduces or eliminates the expression of endogenous TDP-43 in the cell or subject. Any suitable agent can be used to reduce or inhibit the expression of endogenous TDP-43. Examples of agents that can reduce the expression of endogenous TDP-43 include nuclease agents (e.g., ZFN, TALEN, or CRISPR / Cas), DNA binding proteins fused to transcriptional repressors (e.g., transcriptional repressors such as catalytically inactive / dead Cas (dCas) fused to a KRAB domain (dCas-KRAB)), or antisense oligonucleotides, siRNAs, shRNAs, or antisense RNAs.
[0182] Nuclease agents can be used to reduce the expression of endogenous TDP-43. For example, such nuclease agents can be designed to target and cleave a region of the TARDBP gene that disrupts its expression. As a specific example, the nuclease agent can be designed to cleave a region of TARDBP near the start codon. For example, the target sequence can be within approximately 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon, and cleavage by the nuclease agent can destroy the start codon. Alternatively, a nuclease agent designed to cleave a region near the start codon and stop codon can be used to delete the coding sequence between the two nuclease target sequences. DNA-binding proteins fused to a transcriptional repressor domain can also be used to reduce the expression of endogenous TDP-43. For example, a DNA binding protein fused to a transcriptional repressor domain (e.g., a catalytically inactive Cas fused to a KRAB transcriptional repressor domain) can be designed to target a region of TARDBP near the start codon, e.g., within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon).
[0183] Cleavage by a nuclease agent can cause double-strand breaks that can be repaired by non-homologous end joining (NHEJ). NHEJ involves the repair of double-strand breaks in nucleic acids by direct ligation of the cut ends to each other or to an exogenous sequence without the need for a homologous template. Ligation of non-adjacent sequences by NHEJ can often result in deletions, insertions, or translocations near the site of the double-strand break. These insertions and deletions (indels) can disrupt expression of target genes, for example, through frameshift mutations or disruption of start codons.
[0184] Any nuclease agent that induces a nick or double-stranded break at the desired recognition site can be used in the methods and compositions disclosed herein. Naturally occurring or native nuclease agents can be used, as long as the nuclease agent induces a nick or double-stranded break at the desired recognition site. Alternatively, modified or engineered nuclease agents can be used. An "engineered nuclease agent" includes a nuclease that has been engineered (modified or derived) from its native form to specifically recognize and induce a nick or double-stranded break at the desired recognition site. Thus, an engineered nuclease agent can be derived from a native, naturally occurring nuclease agent or can be artificially created or synthesized. An engineered nuclease can, for example, induce a nick or double-stranded break at a recognition site that is not a sequence that would be recognized by a native (unengineered or unmodified) nuclease agent. The modification of the nuclease agent can be as little as one amino acid in a protein cleaving agent or one nucleotide in a nucleic acid cleaving agent. Creating a nick or double-strand break in the recognition site or other DNA can be referred to herein as "cutting" or "cleaving" the recognition site or other DNA.
[0185] Active variants and fragments of the exemplary recognition sites are also provided. Such active variants can include at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a given target sequence, and the active variants retain biological activity and can therefore be recognized and cleaved by nuclease agents in a sequence-specific manner. Assays for measuring double-stranded cleavage of recognition sites by nuclease agents are known in the art (e.g., TaqMan® qPCR assays, Frendewey et al. (2010) Methods in Enzymology 476:295-307, the entire contents of which are incorporated herein by reference for all purposes).
[0186] One type of nuclease agent is transcription activator-like effector nucleases (TALENs). TAL effector nucleases are a class of sequence-specific nucleases that can be used to create double-strand breaks at specific target sequences within the genomes of prokaryotes or eukaryotes. TAL effector nucleases are created by fusing a natural or engineered transcription activator-like (TAL) effector, or a functional portion thereof, to the catalytic domain of an endonuclease, such as FokI. The unique modular TAL effector DNA-binding domain allows for the design of proteins with specific DNA recognition specificity. Therefore, the DNA-binding domain of a TAL effector nuclease can be engineered to recognize specific DNA target sites and can therefore be used to create double-strand breaks at desired target sequences. See WO 2010 / 079430, Morbitzer et al. (2010) Proc. Natl. Acad. Sci. USA 107(50):21617-21622, Scholze & Boch (2010) Virulence 1:428-432, Christian et al. Genetics (2010) 186:757-761, Li et al. (2010) Nucleic Acids Res. (2011) 39(1):359-372, and Miller et al. (2011) Nature Biotechnology 29:143-148, each of which is incorporated by reference in its entirety for all purposes.
[0187] Examples of suitable TAL nucleases and methods for preparing suitable TAL nucleases are disclosed, for example, in U.S. Patent Application Publication Nos. 2011 / 0239315(A1), 2011 / 0269234(A1), 2011 / 0145940(A1), 2003 / 0232410(A1), 2005 / 0208489(A1), 2005 / 0026157(A1), 2005 / 0064474(A1), 2006 / 0188987(A1), and 2006 / 0063231(A1), each of which is incorporated by reference in its entirety for all purposes. In various embodiments, for example, TAL effector nucleases are engineered to cleave within or near a target nucleic acid sequence at a genetic or genomic locus of interest, where the target nucleic acid sequence is at or near a sequence modified by a targeting vector. TAL nucleases suitable for use in the various methods and compositions provided herein include those specifically designed to bind to or near a target nucleic acid sequence modified by a targeting vector described herein.
[0188] In some TALENs, each TALEN monomer contains 33-35 TAL repeats that recognize a single base pair via two hypervariable residues. In some TALENs, the nuclease agent is a chimeric protein containing a TAL repeat-based DNA-binding domain operably linked to an independent nuclease, such as a FokI endonuclease. For example, the nuclease agent can include a first TAL repeat-based DNA-binding domain and a second TAL repeat-based DNA-binding domain, each of which is operably linked to a FokI nuclease, wherein the first and second TAL repeat-based DNA-binding domains recognize two consecutive target DNA sequences on each strand of the target DNA sequence, separated by spacer sequences of various lengths (12-20 bp), and the FokI nuclease subunits dimerize to generate an active nuclease that makes a double-stranded break in the target sequence.
[0189] Nuclease agents used in the various methods and compositions disclosed herein can further include zinc finger nucleases (ZFNs). In some ZFNs, each monomer of the ZFN contains three or more zinc finger-based DNA-binding domains, and each zinc finger-based DNA-binding domain binds to a 3-bp subsite. In other ZFNs, the ZFN is a chimeric protein containing a zinc finger-based DNA-binding domain operably linked to a separate nuclease, such as a FokI endonuclease. For example, the nuclease agent can include a first ZFN and a second ZFN, each operably linked to a FokI nuclease subunit, which recognize two adjacent target DNA sequences on each strand of the target DNA sequence, separated by an approximately 5-7 bp spacer, and the FokI nuclease subunits dimerize to generate an active nuclease that makes a double-stranded break. See, e.g., U.S. Patent Application Publication No. 20060246567, U.S. Patent Application Publication No. 20080182332, U.S. Patent Application Publication No. 20020081614, U.S. Patent Application Publication No. 20030021776, WO 2002 / 057308(A2), U.S. Patent Application Publication No. 20130123484, U.S. Patent Application Publication No. 20100291048, WO 2011 / 017293(A2), and Gaj et al. (2013) Trends in Biotechnology, 31(7):397-405, each of which is incorporated by reference in its entirety for all purposes.
[0190] The methods and compositions disclosed herein can utilize clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) systems or components of such systems to modify genomes or alter the expression of genes in cells. CRISPR / Cas systems include transcripts and other elements that are involved in the expression of or direct the activity of Cas genes. CRISPR / Cas systems can be, for example, Type I, Type II, Type III, or Type V systems (e.g., subtype VA or subtype VB). The methods and compositions disclosed herein can employ CRISPR / Cas systems by utilizing CRISPR complexes (including guide RNAs (gRNAs) complexed with Cas proteins) for site-specific binding or cleavage of nucleic acids. See, for example, WO 2013 / 176772, WO 2014 / 065596, WO 2014 / 089290, WO 2014 / 093622, WO 2014 / 099750, WO 2013 / 142578, and WO 2014 / 131833, each of which is incorporated by reference in its entirety for all purposes.
[0191] The CRISPR / Cas systems used in the compositions and methods disclosed herein may not be naturally occurring. A "non-naturally occurring" system includes any that exhibits human intervention, such as one or more components of the system being altered or mutated from their naturally occurring state, being at least substantially free of at least one other component with which they are naturally associated in nature, or being associated with at least one other component with which they are not naturally associated. For example, some CRISPR / Cas systems use non-naturally occurring CRISPR complexes that include a non-naturally occurring gRNA and Cas protein together, use non-naturally occurring Cas proteins, or use non-naturally occurring gRNAs.
[0192] Active variants and fragments of nuclease agents (i.e., engineered nuclease agents) are also provided. Such active variants may have at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to a native nuclease agent, and the active variants retain the ability to cleave at a desired recognition site and thus retain nick or double-strand break-inducing activity. For example, any of the nuclease agents described herein may be modified from a native endonuclease sequence and designed to recognize and induce a nick or double-strand break at a recognition site not recognized by the native nuclease agent. Thus, some engineered nucleases have the specificity to induce a nick or double-strand break at a recognition site that is different from the corresponding native nuclease agent recognition site. Assays for nick or double-strand break-inducing activity are known and generally measure the overall activity and specificity of an endonuclease on a DNA substrate containing a recognition site.
[0193] The nuclease agent may be introduced into cells by any known means. A polypeptide encoding the nuclease agent may be directly introduced into cells. Alternatively, one or more nucleic acids encoding the nuclease agent may be introduced into cells. Once the nucleic acid encoding the nuclease agent is introduced into cells, the nuclease agent may be expressed transiently, conditionally, or constitutively in the cells. Thus, the nucleic acid encoding the nuclease agent may be contained in an expression cassette and operably linked to a conditional promoter, an inducible promoter, a constitutive promoter, or a tissue-specific promoter. Such promoters of interest are discussed in more detail elsewhere herein. Alternatively, the nuclease agent is introduced into cells as an mRNA encoding the nuclease agent.
[0194] The nucleic acid encoding the nuclease agent can be stably integrated into the genome of the cell and operably linked to a promoter active within the cell. Alternatively, the nucleic acid encoding the nuclease agent can be within a target vector (e.g., the target vector containing the insert polynucleotide, or a vector or plasmid separate from the target vector containing the insert polynucleotide).
[0195] When a nuclease agent is provided to a cell by introducing a nucleic acid encoding the nuclease agent, the nucleic acid encoding the nuclease agent can be modified to replace codons that are more frequently used in the target cell compared to the naturally occurring polynucleotide sequence encoding the nuclease agent. For example, the polynucleotide encoding the nuclease agent can be modified to replace codons that are more frequently used in a given prokaryotic or eukaryotic cell, including bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other target host cell, compared to the naturally occurring polynucleotide sequence.
[0196] Antisense oligonucleotides, antisense RNA, small interfering RNA (siRNA), or short hairpin RNA (shRNA) can also be used to reduce endogenous TDP-43 expression. Such antisense RNA, siRNA, or shRNA can be designed to target any region of the TARDBP mRNA.
[0197] The term "antisense RNA" refers to a single-stranded RNA that is complementary to a messenger RNA strand transcribed within a cell. The term "small interfering RNA (siRNA)" refers to a typical double-stranded RNA molecule that induces the RNA interference (RNAi) pathway. These molecules vary in length (generally between 18 and 30 base pairs) and contain varying degrees of complementarity of the antisense strand to the target mRNA. Some, but not all, siRNAs have unpaired overhanging bases at the 5' or 3' end of the sense and / or antisense strands. The term "siRNA" includes duplexes of two separate strands as well as single strands that can form hairpin structures containing the double-stranded region. The double-stranded structure can be, for example, less than 20, 25, 30, 35, 40, 45, or 50 nucleotides in length. For example, the double-stranded structure can be about 21 to 23 nucleotides in length, about 19 to 25 nucleotides in length, or about 19 to 23 nucleotides in length. The term "short hairpin RNA (shRNA)" refers to a single strand of RNA bases that self-hybridizes in a hairpin structure and can induce the RNA interference (RNAi) pathway upon processing. These molecules vary in length (typically about 50-90 nucleotides, sometimes up to more than 250 nucleotides in length (e.g., in the case of microRNA-compatible shRNAs)). shRNA molecules are processed intracellularly to form siRNAs, which can knock down gene expression. shRNAs can be incorporated into vectors. The term "shRNA" also refers to DNA molecules into which short hairpin RNA molecules may be transcribed.
[0198] Antisense oligonucleotides and RNAi agents can also be used to reduce endogenous TDP-43 expression. Such antisense oligonucleotides or RNAi agents can be designed to target any region of the TARDBP mRNA.
[0199] An "RNAi agent" is a composition comprising small double-stranded RNA or RNA-like (e.g., chemically modified RNA) oligonucleotide molecules that can facilitate the degradation or inhibition of translation of a target RNA, such as messenger RNA (mRNA), in a sequence-specific manner. The oligonucleotides of an RNAi agent are polymers of linked nucleosides, each of which can be independently modified or unmodified. RNAi agents function via the RNA interference mechanism (i.e., induce RNA interference through interaction with the RNA interference pathway machinery (RNA-induced silencing complex, or RISC) in mammalian cells). Although RNAi agents, as that term is used herein, are believed to act primarily via the RNA interference mechanism, the disclosed RNAi agents are not constrained or limited to any particular pathway or mechanism of action. RNAi agents disclosed herein include sense and antisense strands, and include, but are not limited to, short interfering RNA (siRNA), double-stranded RNA (dsRNA), microRNA (miRNA), short hairpin RNA (shRNA), and Dicer substrates. The antisense strand of an RNAi agent described herein is at least partially complementary to the sequence (i.e., the sequence or order of nucleobases or nucleotides described by a series of letters using standard nomenclature) of a target RNA.
[0200] Single-stranded antisense oligonucleotides (ASOs) and RNA interference (RNAi) share the fundamental principle that oligonucleotides bind to target RNA via Watson-Crick base pairing. Without wishing to be bound by theory, during RNAi, a small RNA duplex (RNAi agent) binds to an RNA-induced silencing complex (RISC), where one strand (the passenger strand) is lost and the remaining strand (the guide strand) binds to complementary RNA in cooperation with the RISC. Argonaute 2 (Ago2), a catalytic component of RISC, then cleaves the target RNA. The guide strand is always associated with either the complementary sense strand or a protein (RISC). In contrast, ASOs must survive and function as single strands. ASOs bind to target RNA and either block other factors, such as ribosomes or splicing factors, from binding to the RNA or recruit proteins such as nucleases. Various modifications and target regions are selected for ASOs based on the desired mechanism of action. Gapmers are ASO oligonucleotides containing 2-5 chemically modified nucleotides (e.g., LNA or 2'-MOE) at each end flanking a central 8-10 base gap in the DNA. After binding to the target RNA, the DNA-RNA hybrid serves as a substrate for RNase H.
[0201] ASOs are DNA oligos, typically 15–25 bases long, designed in an antisense orientation to an RNA of interest. Examples of ASOs targeting TARDBPs are provided, for example, in U.S. Patent Application Publication No. 2020-0165610, incorporated herein by reference in its entirety for all purposes. Hybridization of ASOs to target RNAs can mediate RNase H cleavage of the RNA, potentially preventing protein translation of mRNA. To enhance nuclease resistance, phosphorothioate (PS) modifications can be added to oligos. Phosphorothioate linkages also facilitate binding to serum proteins, increasing ASO bioavailability and facilitating productive cellular uptake. In phosphorothioates, sulfur atoms replace non-bridging oxygen atoms in the oligophosphate backbone. ASOs can be chimeras containing both DNA and modified RNA bases. The use of modified RNAs such as 2'-O-methoxy-ethyl (2'-MOE) RNA, 2'-O-methyl (2'OMe) RNA, or Affinity Plus Locked Nucleic Acid bases in chimeric antisense designs has been shown to improve nuclease stability and the affinity (T) of antisense oligos for target RNAs. mHowever, these modifications do not activate RNase H cleavage (i.e., ASOs composed entirely of sugar-modified RNA-like nucleotides (e.g., 2'-MOE) do not support RNase H cleavage of complementary RNA). Therefore, one antisense strategy is the "gapmer" design, which incorporates 2'-O-modified RNA or affinity-locked nucleobases into a chimeric antisense oligo that retains an RNase H-activating domain. Standard gapmers retain a central region of PS-modified DNA bases sufficient to induce RNase H cleavage. These bases are flanked on both sides by blocks of 2'-O-alkyl modifications that enhance binding affinity to the target. For example, gapmers can contain a central section of deoxynucleotides that allows induction of RNase H cleavage, flanked by blocks of 2'-O-alkyl-modified ribonucleotides that protect the central section from nuclease degradation. Once delivered to cells, ASOs enter the nucleus and bind to complementary endogenous RNA targets. Hybridization of the ASO gapmer to the target RNA forms a DNA:RNA heteroduplex in the central region, which becomes a substrate for cleavage by the enzyme RNase H1.
[0202] In one example, the agent that reduces or eliminates endogenous TDP-43 expression includes an antisense oligonucleotide or RNAi agent that targets endogenous TARDBP messenger RNA, or a nucleic acid encoding the antisense oligonucleotide or RNAi agent. In another example, the agent includes a nuclease agent that targets the endogenous TARDBP genomic locus, or one or more nucleic acids encoding the nuclease agent. The nuclease agent can be, for example, a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), or a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein and guide RNA. In a specific example, the nuclease agent is a Cas protein and guide RNA, and optionally, the Cas protein is a Cas9 protein. For example, the methods described herein can further include administering a nucleic acid encoding TDP-43, wherein the nuclease agent cleaves the endogenous TARDBP genomic locus, and the nucleic acid encoding the TDP-43 variant is inserted or recombined with the cleaved endogenous TARDBP genomic locus, such that the TDP-43 variant is expressed from the endogenous TARDBP genomic locus, replacing expression of the variant endogenous TDP-43.
[0203] Such methods may further include screening cells or subjects to confirm the presence of a TDP-43 variant or a nucleic acid encoding a TDP-43 variant. Screening cells or subjects for a TDP-43 variant or a nucleic acid encoding a TDP-43 variant may be performed by any known means. Such methods may further include screening cells or subjects to confirm expression of a nucleic acid encoding a TDP-43 variant. Screening cells or subjects for expression of a nucleic acid encoding a TDP-43 variant may be performed by any known means. For example, methods for measuring protein expression and methods for measuring expression of mRNA encoded by a coding sequence are well known.
[0204] One example of an assay that can be used is the BASESCOPE™ RNA In Situ Hybridization (ISH) assay, a method that can quantify cell-specific edited transcripts, including single-nucleotide changes, in the context of intact, fixed tissue. The BASESCOPE™ RNA ISH assay can complement next-generation sequencing (NGS) and qPCR in characterizing gene editing. While NGS / qPCR can provide quantitative averages of wild-type and edited sequences, it does not provide information about the heterogeneity or percentage of edited cells within a tissue. The BASESCOPE™ ISH assay can provide a landscape view of the entire tissue and quantification of wild-type versus edited transcripts at single-cell resolution, allowing for quantification of the actual number of cells in a target tissue that contain edited mRNA transcripts. The BASESCOPE™ assay uses paired oligo ("ZZ") probes to amplify the signal and achieve single-molecule RNA detection without nonspecific background. However, the BASESCOPE™ probe design and signal amplification system allows for the detection of single molecule RNA using 1ZZ probes, which can differentially detect single nucleotide edits and mutations in intact, fixed tissues.
[0205] As another example, a reporter gene can be used for screening. For example, a nucleic acid encoding a TDP-43 variant can encode the TDP-43 variant fused to a reporter gene such as a fluorescent protein. Exemplary reporter genes include those encoding luciferase, β-galactosidase, green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, Cerulean, T-Sapphire, and alkaline phosphatase.
[0206] As another example, selectable markers can be used to screen for cells harboring TDP-43 variants or nucleic acids encoding TDP-43 variants. Exemplary selectable markers include neomycin phosphotransferase (neomycin phosphotransferase). r ), hygromycin B phosphotransferase (hyg r ), puromycin-N-acetyltransferase (puro r ), blasticidin S deaminase (bsr r ), xanthine / guanine phosphoribosyltransferase (gpt), or herpes simplex virus thymidine kinase (HSV-k).
[0207] The method may further include assessing one or more signs or symptoms of a TDP-43 proteinopathy by any suitable means. Examples of such signs and symptoms are discussed in more detail elsewhere herein and include, for example, TDP-43 hyperphosphorylation, TDP-43 aggregation, aberrant splicing regulation by TDP-43, and aberrant TDP-43 subcellular distribution (e.g., aberrant nuclear depletion). This may be performed, for example, about 1 day, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 5 weeks, about 6 weeks, or more after introduction of the TDP-43 variant or a nucleic acid encoding the TDP-43 variant. For example, the assessment may be performed about 2 weeks to about 6 weeks, or about 3 weeks to about 5 weeks after introduction of the TDP-43 variant or a nucleic acid encoding the TDP-43 variant.
[0208] The methods described herein may, for example, reduce the amount of new TDP-43 aggregate formation (new TDP-43 aggregation) in a cell or a subject, and / or reduce the amount of existing TDP-43 aggregate formation (existing TDP-43 aggregation) in a cell or a subject. For example, the methods described herein may prevent new TDP-43 aggregate formation and / or reverse existing TDP-43 aggregate formation.
[0209] The methods described herein can, for example, reduce the amount of aberrant regulation of splicing by TDP-43 in a cell or subject. For example, the methods described herein can rescue the aberrant regulation of splicing by TDP-43.
[0210] The methods described herein can, for example, reduce the abnormal subcellular localization of TDP-43 in a cell or subject (e.g., reduce the abnormal nuclear depletion of TDP-43). For example, the methods described herein can rescue the abnormal subcellular localization of TDP-43 (e.g., rescue the abnormal nuclear depletion of TDP-43).
[0211] The methods described herein can, for example, reduce the level of phospho-TDP-43 in a cell or a subject.
[0212] In some of the methods described herein, the method is for treating or preventing a TDP-43 proteinopathy in a subject. Therapeutic or pharmaceutical compositions, including those disclosed herein, can be administered with suitable carriers, excipients, and other agents incorporated into the formulation to provide improved entry, delivery, tolerance, and the like. Many suitable formulations can be found in Remington's Pharmaceutical Sciences, Mack Publishing Company, Easton, PA, a formulary known to all pharmacists. See also Powell et al., "Compendium of excipients for parenteral formulations," PDA (1998) J. Pharm. Sci. Technol. 52:238-311. The compositions disclosed herein can be administered to alleviate, prevent, or reduce the severity of one or more of the signs or symptoms of a TDP-43 proteinopathy, as described in more detail elsewhere herein.
[0213] In some methods (e.g., methods for treatment), the subject has one or more signs or symptoms of a TDP-43 proteinopathy. For example, the subject may have pre-existing TDP-43 aggregate formation in one or more cells. TDP-43 proteinopathies are classified based on the degree of altered TDP-43 inclusions and include a growing number of neurodegenerative diseases, including amyotrophic lateral sclerosis (ALS), frontotemporal lobar degeneration with ubiquitin-immunoreactive tau-negative inclusions (FTLD-U), and FTLD with motor neuron disease (FTLD-MND). In addition, TDP-43 inclusions have also been identified in several other neurodegenerative disorders, including Alzheimer's disease, corticobasal degeneration, Lewy body-associated disease, and Pick's disease. TDP-43 proteinopathies are typically characterized by abnormal phosphorylation, ubiquitination, cleavage, and / or nuclear depletion of TDP-43 in neurons and glial cells. Various major neurodegenerative diseases exhibit similar TDP-43 pathological manifestations in neurons and glia, including the accumulation of detergent-resistant, ubiquitinated, or hyperphosphorylated TDP-43 inclusions in the cytoplasm, usually accompanied by depletion of TDP-43 from the nucleus. These TDP-43-associated pathological features are typically referred to as TDP-43 proteinopathies. See, e.g., Gao et al. (2018) J.Neurochem.doi:10.1111 / jnc.14327, incorporated herein by reference in its entirety for all purposes.
[0214] TDP-43 proteinopathies encompass a wide range of neurodegenerative diseases and phenotypes that may be inherited in a Mendelian pattern or may be apparently sporadic. TDP-43 has been found to aggregate in several diseases, including Alzheimer's disease (brain), LATE (brain), and inclusion body myositis (muscle). Numerous genes and diseases have been associated with TDP-43 proteinopathies (Table 2). See, e.g., de Boer et al. (2021) J. Neurol. Neurosurg. Psychiatry 92:86-95, incorporated herein by reference in its entirety for all purposes.
[0215] [Table 2] * Multisystem proteinopathy - a familial disorder in which patients exhibit ALS, FTLD, inclusion body myositis, Paget's disease of bone, or a combination of these phenotypes. ALS, amyotrophic lateral sclerosis; bi, behavioral impairment; CARTS, cerebral age-related TDP-43 with sclerosis; ci, cognitive impairment; CTE, chronic traumatic encephalopathy; FOSMN, facial onset sensory and motor neuron disorder; FTLD, frontotemporal lobar degeneration; HS, hippocampal sclerosis; LATE, limbic-predominant age-related TDP-43 encephalopathy; na, not applicable; PPA, primary progressive aphasia; sIBM, sporadic inclusion body myositis; TDP-43, TAR DNA-binding protein 43.
[0216] TDP-43 aggregates are evident in approximately 97% of all amyotrophic lateral sclerosis (ALS) cases. These TDP-43 inclusions are evident in both demented and non-demented patients with ALS, and their density increases with disease progression, particularly with the onset of cognitive impairment. Three predominant cell-type-specific patterns of TDP-43 pathology have been identified in ALS, including (1) glia (22% of cases), (2) mixed neurons and glia (59% of cases), and (3) neurons (7% of cases). The extent of TDP-43 pathology distinguishes ALS-FTLD from ALS without FTLD, and the presence of TDP-43 pathology outside the motor cortex, as assessed by the Edinburgh Cognitive and Behavioral ALS Screen (ECAS), has been associated with cognitive impairment in ALS. TDP-43 pathology in the orbitofrontal, dorsolateral prefrontal, medial prefrontal cortex, and ventral anterior cingulate cortex was associated with executive dysfunction. Language dysfunction was associated with TDP-43 pathology in the inferior frontal gyrus, transverse temporal area, middle and inferior temporal gyrus, and angular gyrus. Verbal fluency dysfunction was associated with TDP-43 pathology in the prefrontal cortex, inferior frontal gyrus, ventral anterior cingulate, and transverse temporal area. However, behavioral abnormalities were associated with TDP-43 pathology in the orbitofrontal and prefrontal cortices and ventral anterior cingulate cortex.
[0217] The methods described herein can alleviate one or more signs and symptoms of TDP-43 proteinopathy in a cell or subject. Some examples of signs and symptoms of TDP-43 proteinopathy at the cellular level include abnormal phosphorylation, ubiquitination, cleavage, and / or nuclear depletion of TDP-43 in neurons, glial cells, or muscle cells, or the accumulation of detergent-resistant, ubiquitinated, or hyperphosphorylated TDP-43 inclusions in the cytoplasm, usually accompanied by depletion of TDP-43 from the nucleus. Other signs and symptoms at the organismal level can include neurodegeneration, a term that refers to the progressive loss of neuronal structure and function.
[0218] All patent applications, websites, other publications, accession numbers, etc. cited above or below are incorporated by reference in their entirety for all purposes to the same extent as if each individual item was specifically and individually indicated to be incorporated by reference. Where different versions of a sequence are associated with an accession number at different times, the version associated with the accession number as of the effective filing date of this application is meant. The effective filing date means the filing date of the actual application associated with the accession number or of the priority application, if applicable. Similarly, where different versions of a publication, website, etc. are published at different times, the version published most recently before the effective filing date of this application is meant, unless otherwise specified. Unless specifically indicated otherwise, any feature, step, element, embodiment, or aspect of the invention may be used in combination with any other. While the invention has been described in some detail through diagrams and examples for purposes of clarity and understanding, it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims.
[0219] A brief description of arrays The nucleotide and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases and three-letter code for amino acids. The nucleotide sequences follow the standard convention of beginning at the 5'-end of the sequence and proceeding forward (i.e., left to right on each line) to the 3'-end. Only one strand of each nucleotide sequence is shown, but the complementary strand is understood to be included by any reference to the shown strand. Where a nucleotide sequence encoding an amino acid sequence is provided, it is understood that codon-degenerate variants thereof that encode the same amino acid sequence are also provided. The amino acid sequences follow the standard convention of beginning at the amino-terminus of the sequence and proceeding forward (i.e., left to right on each line) to the carboxy-terminus.
[0220] [Table 3-1]
[0221] [Table 3-2] [Example]
[0222] Example 1. Replacement of the prion-like domain of TDP-43 with the PLD of hnRNPA2B1 Several structural features of the TDP-43 protein have been identified, including a nuclear localization signal (NLS), two RNA recognition motifs (RRM1 and RRM2), a putative nuclear export signal (NES), and a large domain in the carboxyl-terminal half of the protein that has been described as a less complex, less ordered, or prion-like domain (PLD). Among the mutations in TDP-43 associated with familial cases of ALS, the majority are found in the PLD. The TDP-43 PLD is largely responsible for mediating TDP-43's tendency to form aggregates. Biochemical studies have shown that the even spacing of aromatic residues throughout the PLD allows efficient liquid-liquid phase separation but prevents irreversible association that leads to aggregation. The TDP-43 PLD contains fewer aromatic amino acids and is less evenly spaced than other RNA-binding proteins (Figure 1), suggesting that reconfiguring the spacing of aromatic residues throughout the PLD may prevent aggregation. To test this hypothesis, we created a PLDswap allele (amino acid sequence set forth in SEQ ID NO: 69, coding sequence set forth in SEQ ID NO: 70) that replaces the PLD of mouse TDP-43 with the PLD of mouse hnRNPA2B1 (Figure 2B). The PLD of hnRNPA2B1 is composed of evenly spaced aromatic residues compared to the PLD of TDP-43 (Figure 1). Western blot analysis of subcellular fractions obtained from motor neurons derived from ES cells carrying the PLDswap allele as the only form of TDP-43 showed normal TDP-43 subcellular distribution (Figure 3). To determine whether this chimeric protein could replace TDP-43, we attempted to generate mice in which this was the only form of TDP-43. We previously confirmed that genetic ablation of TDP-43 results in embryonic lethality between e3.5 and e5.5, and mice with mutant NLS (ΔNLS) or PLD deletion (ΔPLD) do not survive past embryonic days 12.5 and 15.5, respectively.Rather surprisingly, mice carrying the PLDswap allele as their only form of TDP-43 died at birth, suggesting that this form of TDP-43 may replace TDP-43 function during embryogenesis (Fig. 4).
[0223] Finally, we wanted to determine whether this chimeric protein could retain TDP-43 function. The best-characterized function of TDP-43 is in regulating RNA splicing. Specifically, TDP-43 binds to intronic sequences to suppress the aberrant inclusion of cryptic exons in mRNA transcripts and can also control alternative splicing events. To assay for TDP-43 function, we performed semiquantitative RT-PCR for specific splicing events in Adnp2, Dnajc5, Poldip3, Tsn, and Sortilin1, which were determined to be TDP-43 dependent. We previously showed that these events were disrupted by mutations that rendered the NLS nonfunctional (ΔNLS), deleted the nuclear export signal (ΔNES), and deleted the prion-like domain (ΔPLD). Strikingly, in ES cell-derived motor neurons, the replacement of the TDP-43 PLD by the hnRNPA2B1 TDP-43 splicing functional PLD is largely retained (Fig. 5 ).
[0224] Collectively, these data demonstrate that the prion-like domain of TDP-43 is amenable to manipulation and suggest that removing wild-type TDP-43 and replacing it with an engineered, aggregation-resistant form may prove therapeutically valuable. Simply replacing the wild-type protein may also rescue phenotypes associated with loss of functional TDP-43.
[0225] method Cell culture. The ability of the chimeric TDP-43-hnRNP-A2B1 PLDswap, as the only form of protein expressed by the cells, to support the viability of embryonic stem (ES) cells and their derived motor neurons (ESMN) was tested by differentiation in culture. ES cells were cultured in embryonic stem cell medium (ESM; DMEM + 15% fetal bovine serum + penicillin / streptomycin + glutamine + non-essential amino acids + nucleosides + β-mercaptoethanol + sodium pyruvate + LIF) for 2 days, during which the medium was changed daily. One hour before trypsinization, the ES medium was replaced with 7 mL of ADFNK medium (advanced DMEM / F12 + neurobasal medium + 10% knockout serum + penicillin / streptomycin + glutamine + β-mercaptoethanol). The ADFNK medium was aspirated, and the ESCs were trypsinized with 0.05% trypsin-EDTA. The pelleted cells were resuspended in 12 mL of ADFNK and grown in suspension for 2 days. The cells were cultured for an additional 4 days in ADFNK supplemented with retinoic acid (RA), smooth muscle agonist, and purmorphamine to obtain limb-like motor neurons (ESMN). Dissociated motor neurons were seeded and matured in embryonic stem cell-derived motor neuron medium (ESMN; neurobasal medium + 2% horse serum + B27 + glutamine + penicillin / streptomycin + β-mercaptoethanol + 10 ng / mL GDNF, BDNF, and CNTF). The conditional knockout allele was activated at the ES cell and ESMN stages using cre recombinase delivered via electroporation.
[0226] Subcellular fractionation of TDP-43. The subcellular localization of the chimeric TDP-43-PLDswap protein was analyzed using an antibody recognizing the N-terminus of the TDP-43 polypeptide (α-TDP-43 N-term) and an antibody recognizing the C-terminal prion-like domain of the TDP-43 polypeptide (α-TDP-43 C-term) (Proteintech, Rosemont, IL). Soluble cytoplasmic protein extracts were prepared by incubating ES cell-derived MNs on ice for 10 min in ice-cold lysis buffer (10 mM KCl, 10 mM Tris-HCl, pH 7.4, 1 mM MgCl2, 1 mM DTT, 0.01% NP-40) supplemented with protease and phosphatase inhibitors (Roche). The cells were then passed five times through a 27-gauge syringe. After centrifugation at 4000 rpm for 5 min at 4°C, the protein supernatant containing the soluble cytoplasmic extract was collected. Insoluble nuclear protein extracts were prepared by resuspending the pellet in an equal volume of RBS-100 buffer (10 mM Tris-HCl pH 7.4, 2.5 mM MgCl, 100 mM NaCl, 0.1% NP-40) supplemented with proteases and phosphatases. An equal volume of 2x SDS sample buffer was added to each fraction, and the samples were heated to 90°C. Equal volumes of each fraction were then loaded onto a 10% SDS gel and electroporated at 225 V for 50 min. Western blotting for TDP-43 was then performed using either the α-TDP-43-N-term antibody or the α-TDP-43-C-term antibody (the latter does not recognize the PLDswap mutant).
[0227] Generation of mice expressing chimeric TDP-43-hnRNPA2B1 "PLDswap" proteins. Although TDP-43 deletion results in embryonic lethality, embryonic stem cells expressing only the mutant ΔNLS TDP-43 gene or the mutant ΔPLD TDP-43 gene from the endogenous TARDBP locus are viable and can differentiate into motor neurons in vitro. These data raise the possibility that embryonic stem cells expressing mutant TDP-43 polypeptides lacking functional structural domains from the endogenous TARDBP locus may be viable and useful for creating animal models of TDP-43 proteinopathies. For example, such embryonic stem cells can be used to generate non-human animals, e.g., mice, expressing mutant TDP-43 proteins lacking functional structural domains to investigate the role of TDP-43 structural domains in normal and pathological biological processes.
[0228] To generate embryos or animals expressing the chimeric TDP-43-hnRNPA2B1 protein, we used the VELOCIMOUSE® method (Dechiara (2009) Methods Mol. Biol. 530:311-324 and Poueymirou et al. (2007) Nat. Biotechnol. 25:91-99, each of which is incorporated herein by reference in its entirety for all purposes). Targeted ES cells containing (i) the TARDBP gene containing the chimeric TDP-43-hnRNPA2B1 "PLDswap" allele at the endogenous TARDBP locus and (ii) a null TDP-43 allele generated by Cre-mediated deletion of the floxed exon 3(-) were injected into uncompressed 8-cell Swiss Webster embryos. The viability of the embryos after fertilization was examined, and their ability to generate live-born F0 generation mice was assessed.
[0229] RT-PCR of splicing events. Total RNA was isolated from ES cell-derived motor neurons using Trizol reagent, followed by DNase treatment. For cDNA, 1 μg of total RNA was used as a template for cDNA synthesis using the SuperScript IV First-Strand Synthesis System (ThermoFisher catalog number 18091050). Reactions were performed in a 20 μL volume, and then the final volume was brought to 100 μL after cDNA synthesis was completed. PCR reactions were performed using Q5 2X MasterMix (NEB) with 2 μL of cDNA template, 1 mM of each forward primer, and 1 mM of each reverse primer in a total reaction volume of 25 μL. For each transcript, PCR was first optimized to determine the number of cycles that would allow amplification within the linear range while remaining nonsaturating. Reactions were run on a 1.8% agarose gel in 1x TAE, and bands were visualized using SybrSafe. The primers used and the corresponding cycle numbers are listed below.
[0230] [Table 4]
[0231] Example 2. Function of TDP-43 PLDswap in neonatal mice The requirement for TDP-43 function during embryogenesis and fetal development is essential and may encompass functions of TDP-43 unrelated to postnatal neuronal survival in vivo. Therefore, to further characterize the functionality of the PLDswap form of TDP-43 while circumventing any prenatal requirement for TDP-43, we performed a series of experiments to test whether the PLDswap form of TDP-43 can complement wild-type TDP-43 in postnatal mice. To do this, we utilized mice carrying an exon 3 floxed conditional knockout (cKO) allele (exon 3 floxed, "flEx3") that undergoes Cre-mediated recombination to generate a ΔEx3 knockout allele in the presence of Cre. We then transformed this conditional allele into a PLDswap mutant (TDP-43 flEx3 / PLDswap ), allowing for the conditional ablation of WT TDP-43, leaving PLDswap as the only form of protein in Cre-expressing cells. As negative and positive controls, respectively, we also used mice homozygous for the flEx3 conditional allele (TDP-43 flEx3 / flEx3 ) and mice with one WT allele paired with a conditional allele (TDP-43 flEx3 / WT ) was used. To induce conditional knockout of the floxed WT allele, PHP.eB.AAV viruses expressing a Cre-2A-mCherry cassette under the control of either the ubiquitous CAG promoter (PHP.eB.AAV-CAG-Cre-2A-mCherry) or the neuron-specific human synapsin promoter (PHP.eB.AAV-hSYN-Cre-2A-mCherry) were injected intracerebroventricularly (icv) into P0 neonatal mice of all three genotypes. See Figure 6A. Given the essential requirement for TDP-43, we used survival as a readout of TDP-43 function in these animals. As expected due to the essentiality of TDP-43, the cKO allele (TDP-43 ΔEx3 / ΔEx3Mice homozygous for TDP-43 did not survive beyond 5 weeks after Cre delivery, either in the ubiquitous (CAG) or neuron-specific (SYN) contexts. In contrast, TDP-43 flEx3 / WT and TDP-43 flEx3 / PLDswap Both mice survived significantly longer and to a remarkably similar extent, suggesting that the PLDswap form of TDP-43 can function long-term as a replacement for wild-type TDP-43 protein in neurons (see Figure 6B).
[0232] One caveat of this experimental system is that high expression of Cre recombinase in neurons can be toxic, especially in the context of highly expressed CAG-Cre, as seen in our positive control TDP-43. flEx3 / WT This likely accounts for the shorter than expected survival of the mice, and therefore, this system may underestimate the extent to which the PLDswap form of TDP-43 can compensate for TDP-43 function in vivo.
[0233] To support the survival data, we also directly tested the function of the PLDswap protein in TDP-43-dependent splicing events. To do this, we collected spinal cord tissue from mice with or without Cre-mediated removal of the conditional wild-type TDP-43 allele and performed semiquantitative RT-PCR for specific TDP-43-dependent splicing events. We tested two mRNAs known to undergo TDP-43 splicing: Adnp2 mRNA and Tsn mRNA. All animals without CAG-Cre expression showed normal transcript processing, whereas those without the conditional allele (and TDP-43) showed normal transcript processing. ΔEx3 / ΔEx3 Homozygous deletion of Adnp2 and Tsn resulted in mis-splicing of both transcripts (i.e., cryptic exon inclusion in Adnp2 and exon 5 skipping in Tsn). In comparison, Cre-treated TDP-43 ΔEx3 / WT Both transcripts were spliced normally in mice. Cre-treated TDP-43 flEx3 / PLDswapIn mice, there was a trend toward rescue of splicing defects, particularly in the cryptic exon inclusion in Adnp2 (see Figure 7). This analysis of two genes that depend on TDP-43 for normal splicing provides molecular support for the notion that chimeric PLDswap forms of TDP-43 have the ability to function somewhat like normal TDP-43.
[0234] Example 3. Replacement of a portion of the prion-like domain of TDP-43 with a portion of the PLD of hnRNPA2B1 As described in Example 1, the PLD of TDP-43 contains fewer aromatic amino acids and is less evenly spaced than other RNA-binding proteins (Figure 1), suggesting that reorganizing the spacing of aromatic residues throughout the PLD may prevent aggregation. To test this hypothesis, we generated the PLD28aa allele (amino acid sequence set forth in SEQ ID NO: 67, coding sequence set forth in SEQ ID NO: 71), which replaces a highly conserved 28-amino acid stretch within the PLD of mouse TDP-43, which has been shown to be important for aggregation in yeast, with a sequence from the PLD of mouse hnRNPA2B1. The PLD of hnRNPA2B1 is composed of more evenly spaced aromatic residues than the PLD of TDP-43 (Figure 1). The PLD28aa allele was confirmed to be viable as the only form of TDP-43 in mouse embryonic stem cells (data not shown). As in Example 1, Western blot analysis of subcellular fractions obtained from ES cell-derived motor neurons bearing the PLD28aa allele as the only form of TDP-43 was performed to analyze the intracellular distribution of TDP-43. As in Example 1, to determine whether this chimeric protein can replace TDP-43, mice were generated in which this was the only form of TDP-43. Embryonic development was evaluated in the mice. To determine whether this chimeric protein can retain the function of TDP-43, experiments were also performed in Example 1, specifically by evaluating TDP-43 splicing function in ES cell-derived motor neurons.
[0235] Example 4. Introduction of uniformly spaced aromatic residues ("stickers") into the prion-like domain of TDP-43 to promote liquid-liquid phase separation (LLPS) but inhibit aggregation As described in Example 1, the PLD of TDP-43 contains fewer aromatic amino acids and is less evenly spaced than other RNA-binding proteins (Figure 1), suggesting that reconfiguring the spacing of aromatic residues throughout the PLD may prevent aggregation. To test this hypothesis, we introduced evenly spaced aromatic residues ("stickers") into the prion-like domain of TDP-43 to create a PLD aromatic allele (amino acid sequence set forth in SEQ ID NO: 64, coding sequence set forth in SEQ ID NO: 72) that promotes liquid-liquid phase separation (LLPS) but inhibits aggregation. The PLD aromatic allele was confirmed to be viable as the sole form of TDP-43 in mouse embryonic stem cells (data not shown). Western blot analysis of subcellular fractions obtained from motor neurons derived from ES cells harboring the PLD aromatic allele as the sole form of TDP-43, as in Example 1, was performed to analyze TDP-43 subcellular distribution. As in Example 1, to determine whether this chimeric protein can replace TDP-43, mice in which this is the only form of TDP-43 are generated. Embryonic development is evaluated in the mice. To determine whether this chimeric protein can retain the function of TDP-43, experiments are also performed in Example 1, specifically by evaluating TDP-43 splicing function in ES cell-derived motor neurons.
Claims
1. A TAR DNA-binding protein 43 (TDP-43) variant, wherein the prion-like domain (PLD) of the TDP-43 variant is mutated to have more aromatic amino acids and / or more evenly spaced aromatic amino acids than the PLD from wild-type TDP-43.
2. The TDP-43 variant according to claim 1, wherein the PLD of wild-type TDP-43 is replaced with a PLD from hnRNPA2B1 in the TDP-43 variant.
3. The PLD in the TDP-43 variant is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence described in SEQ ID NO: 27 or 13, or The TDP-43 variant according to claim 1 or 2, wherein the PLD in the TDP-43 variant is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence described in sequence number 68 or 58.
4. The PLD in the TDP-43 variant includes the sequence described in Sequence ID No. 27 or 13, or The TDP-43 variant according to claim 1 or 2, wherein the PLD in the TDP-43 variant includes the sequence described in sequence number 68 or 58.
5. The TDP-43 variant is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence described in Sequence ID No. 28, or The TDP-43 variant according to claim 1 or 2, wherein the TDP-43 variant is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence described in Sequence ID No.
69.
6. The TDP-43 variant includes the sequence described in Sequence ID No. 28, or The TDP-43 variant according to claim 1 or 2, wherein the TDP-43 variant includes the sequence described in sequence number 69.
7. The aforementioned TDP-43 variant is (I) Less prone to aggregation than the wild-type TDP-43, and / or (II) Maintaining the function of wild-type TDP-43 in splicing regulation, and / or (III) Primarily nucleate and / or retaining the intracellular distribution of wild-type TDP-43, and / or (IV) Maintains the function of wild-type TDP-43 during embryonic development. The TDP-43 variant according to claim 1 or 2.
8. The TDP-43 variant according to claim 1 or 2, wherein the TDP-43 variant is a human TDP-43 variant.
9. A composition for use in the treatment or prevention of TDP-43 proteinosis in a subject, comprising the TDP-43 variant described in claim 1 or 2.
10. The composition according to claim 9, wherein the TDP-43 protein disorder is amyotrophic lateral sclerosis (ALS).
11. A nucleic acid encoding the TDP-43 variant described in claim 1.
12. The nucleic acid according to claim 11, wherein the nucleic acid comprises DNA, and the nucleic acid is located within an expression construct which includes a promoter operably linked to the nucleic acid encoding the TDP-43 variant.
13. The nucleic acid according to claim 12, wherein the promoter is a neuron-specific promoter or a constitutive promoter.
14. The nucleic acid according to claim 11, wherein the nucleic acid is present in a viral vector.
15. The nucleic acid according to claim 14, wherein the viral vector is a lentiviral vector or an adeno-associated virus (AAV) vector.
16. It is a cell, (i) The TDP-43 variant according to claim 1 or 2, or (ii) A cell comprising the nucleic acid described in claim 11, wherein the TDP-43 variant is expressed.
17. Non-human animals, (i) The TDP-43 variant according to claim 1 or 2, or (ii) comprising the nucleic acid according to claim 11, wherein the TDP-43 variant is expressed, A non-human animal, wherein the aforementioned non-human animal is a mouse or a rat.
18. A method for producing a non-human animal according to claim 17, comprising administering the TDP-43 variant or the nucleic acid to the non-human animal.
19. A composition characterized in that the composition is administered to cells, (i) The TDP-43 variant according to claim 1 or 2, or (ii) A composition comprising the nucleic acid described in claim 11, wherein the TDP-43 variant is expressed.