Expression vectors and compositions for treating neurological disorders
An expression vector with an artificial transcription factor and regulatory elements induces neuronal differentiation of non-neuronal cells in vivo, overcoming the in vitro-to-in vivo translation challenge.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- シャンハイ ジェネマジック バイオサイエンシズ カンパニーリミティド
- Filing Date
- 2024-03-30
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods struggle to induce neuronal differentiation of non-neuronal cells into functional neurons in vivo due to the complex environment of the body, despite successful induction in vitro.
An expression vector containing a nucleic acid sequence encoding an artificial transcription factor with an RZFD core domain, regulated by a glial cell-specific promoter and enhanced by various elements, capable of binding to RE1 to promote neuronal differentiation.
The vector effectively differentiates non-neuronal cells into functional neurons in vivo, addressing the challenge of in vitro-to-in vivo translation.
Smart Images

Figure 2026511737000025 
Figure 2026511737000026 
Figure 2026511737000027
Abstract
Description
Technical Field
[0001] This application relates to molecular biology and therapeutic drugs for nervous system diseases, and particularly to an expression vector and a composition containing a nucleic acid sequence encoding an artificial transcription factor capable of binding to RE1.
[0002] Cross-reference
[0003] This application claims the benefit of Patent Application No. PCT / CN2023 / 085228 filed on March 30, 2023 and Patent Application No. PCT / CN2024 / 072004 filed on January 12, 2024, the contents of which are incorporated herein by reference in their entirety.
[0004] The sequence listing file submitted simultaneously The entire content of the following XML file is incorporated herein by reference: Sequence listing in computer-readable format (CRF) (file name: TFH00974PCT-Sequence listing, date: 20240329, size: 131 KB).
Background Art
[0005] The repressor element 1 / neuron-restrictive silencer element (RE1 / NRSE) is a negative regulatory element that binds to the RE1 silencing transcription factor (REST) and can affect the development and maturation of neurons.
[0006] In previous studies, scientists have succeeded in inducing many special types of neurons in in vitro culture systems by combining the expression of multiple genes and transcription factors while simultaneously adding different culture media and small molecules. However, these in vitro experiments are difficult to apply to the body. Due to the complex environment of the body, various factors screened in vitro cannot function as they did in vitro. Therefore, many factors that induce neuronal differentiation in vitro cannot induce the differentiation of glial cells into neurons in the body. How to regenerate special types of neurons in the body is an extremely difficult challenge. [Overview of the Initiative]
[0007] The present invention provides an expression vector and further provides a composition, the expression vector and composition provided by the present invention can promote the differentiation of non-neuronal cells into functional neurons.
[0008] In a first embodiment of the present invention, the expression vector provided by the present invention comprises (a) a nucleic acid sequence encoding an artificial transcription factor capable of binding to RE1, including an RZFD core domain, and (b) a regulatory element that modulates the nucleic acid expression of the artificial transcription factor, wherein the amino acid sequence of the RZFD core domain includes any 5 to 8 sequences from SEQ ID NOs. 3 to 10, or is any 5 to 8 sequences from SEQ ID NOs. 3 to 10, and does not include the sequences shown in SEQ ID NOs. 12 and 13. Preferably, the amino acid sequence of the RZFD core domain includes any 6 to 8 sequences from SEQ ID NOs. 3 to 10, or is any 6 to 8 sequences from SEQ ID NOs. 3 to 10.
[0009] In some embodiments, the amino acid sequence of the RZFD core domain includes, or is a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 76, preferably the amino acid sequence of the RZFD core domain is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, or at least 92% identity with SEQ ID NO: 74 or 75. , includes or is a sequence having at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 2.
[0010] In some embodiments, the amino acid sequence of the RZFD core domain includes, or is, a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 14.
[0011] In some embodiments, the coding nucleic acid sequence of the RZFD core domain includes, or is, a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with sequence number 15 or 64.
[0012] In some embodiments, the regulatory element includes a promoter that modulates the nucleic acid expression of an artificial transcription factor, and preferably, the promoter includes a glial cell-specific promoter.
[0013] Furthermore, the above-mentioned glial cell-specific promoter is an astrocyte-specific promoter, a Müller cell-specific promoter, or a cochlear glial cell-specific promoter. In addition, the above-mentioned glial cell-specific promoter includes the GFAP promoter, ALDH1L1 promoter, EAAT1 / GLAST promoter, glutamine synthetase promoter, S100β promoter, EAAT2 / GLT-1 promoter, Plp1 promoter, or Rlbp1 promoter.
[0014] Furthermore, the above-mentioned cochlear glial cell-specific promoters include the GFAP promoter, the ALDH1L1 promoter, the EAAT1 / GLAST promoter, or the Plp1 promoter.
[0015] Furthermore, the above-mentioned astrocyte-specific promoter or Müller cell-specific promoter includes the GFAP promoter, ALDH1L1 promoter, EAAT1 / GLAST promoter, glutamine synthetase promoter, S100β promoter, EAAT2 / GLT-1 promoter, Plp1 promoter, or Rlbp1 promoter.
[0016] In some embodiments, the regulatory element further comprises an enhancer that increases the transcription of the artificial transcription factor, preferably the enhancer of the expression vector comprising a CMV enhancer, an SV40 intron enhancer, an MVM intron enhancer, a β-globulin intron enhancer, or a synthetic intron enhancer, or any combination thereof (e.g., a combination of a CMV enhancer and an MVM enhancer). In some embodiments, the nucleic acid sequence of the CMV enhancer comprises or is a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 16. In some embodiments, the nucleic acid sequence of the SV40 intron type enhancer includes, or is a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 17. In some embodiments, the nucleic acid sequence of the MVM intron type enhancer includes, or is a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 18. In some embodiments, the nucleic acid sequence of the β-globulin intron type enhancer includes, or is, a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 19.In some embodiments, the nucleic acid sequence of the synthetic intron-type enhancer includes, or is, a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence shown in any one of SEQ ID NOs.
[0017] In some specific embodiments, the above-mentioned adjustment element may include one enhancer or may include multiple enhancers, and the multiple enhancers included may be the same or different.
[0018] In some embodiments, the regulatory element further comprises a post-transcriptional enhancing element (WPRE) of woodchuck hepatitis virus. For example, the regulatory element comprises a promoter and a WPRE that regulate the nucleic acid expression of an artificial transcription factor, or the regulatory element comprises a promoter, a WPRE, and an enhancer that increases the transcription of the artificial transcription factor.
[0019] In some embodiments, the artificial transcription factor further comprises at least one nuclear localization sequence (NLS). Preferably, the NLS is selected from the following: SV40 NLS, nucleoplasmin NLS, c-myc NLS, hRNPA1M9 NLS, nuclear importin α NLS, p53 NLS, influenza virus NS1 NLS, hepatitis D antigen NLS, or bpNLS.
[0020] In some embodiments, the NLS is located at the N-terminus or C-terminus of the RZFD core domain.
[0021] In some embodiments, the NLS is located at the N-terminus and C-terminus of the RZFD core domain.
[0022] In some embodiments, the amino acid sequence of the NLS includes, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 24. In some embodiments, the coding nucleic acid sequence of the NLS includes, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence shown in SEQ ID NO: 25.
[0023] In some embodiments, the artificial transcription factor further comprises a gene activation domain, preferably located at the C-terminus and / or N-terminus of the RZFD core domain.
[0024] In some embodiments, the gene activation domain includes the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination thereof.
[0025] In some embodiments, the gene activation domain includes the VP64 activation domain, the P65 activation domain, the RTA activation domain, or the HSF1 activation domain, or any combination thereof.
[0026] In some embodiments, the activation domain is located at the C-terminus or N-terminus of the RZFD core domain. Furthermore, the activation domain of VP64 is located at both the C-terminus and N-terminus of the RZFD core domain.
[0027] In some embodiments, the amino acid sequence of the activating domain of VP64 includes, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 26. In some embodiments, the coding nucleic acid sequence of the activating domain of VP64 includes, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 27. In some embodiments, the amino acid sequence of the activation domain of P65 includes, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 69. In some embodiments, the amino acid sequence of the activation domain of RTA includes, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 11. In some embodiments, the amino acid sequence of the activation domain of HSF1 includes, or is, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 70.
[0028] In some embodiments, the gene activation domain includes the activation domain of VP64 and the activation domain of P65. Preferably, the gene activation domain further includes the activation domain of HSF1 or the activation domain of RTA.
[0029] In some embodiments, the activation domain is located at the N-terminus or C-terminus of the RZFD core domain. Further, the activation domain is located at the N-terminus and C-terminus of the RZFD core domain.
[0030] In some embodiments, the gene activation domain includes the activation domains of P65 and HSF1. In some embodiments, the amino acid sequence of the activation domain includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 28, or is one of those sequences. In some embodiments, the activation domain is encoded by a nucleotide sequence having at least 70%, at least 8%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 29.
[0031] In some embodiments, the gene activation domain includes the activation domain of VP64, the activation domain of P65, and the activation domain of HSF1. Preferably, the sequence of the activation domain includes, or is a sequence of, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 73.
[0032] In some embodiments, the gene activation domain includes the VP64 activation domain, the P65 activation domain, and the RTA activation domain. Preferably, the sequence of the activation domain includes, or is a sequence of, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 72.
[0033] In some embodiments, the sequence of the gene activation domain includes, or is, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NOs. 28, 69, 70, or 71.
[0034] In some embodiments, the artificial transcription factor further comprises a nuclear localization sequence (NLS) and a gene activation domain, preferably the NLS and / or gene activation domain being located at the N-terminus and / or C-terminus of the RZFD core domain, and more preferably the method of linking the N-terminus to the C-terminus of the artificial transcription factor is RZFD core domain-NLS-gene activation domain, RZFD core domain-gene activation domain-NLS, NLS-gene activation domain-RZFD core domain, gene activation domain-NLS-RZFD core domain, NLS-RZFD core domain-gene activation domain, gene activation domain-RZFD core domain-NLS, gene activation domain-NLS-RZFD core domain-NLS-gene activation domain, or NLS-gene activation domain-RZFD core domain-gene activation domain-NLS.
[0035] In some embodiments, the gene activation domain includes the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination thereof, preferably the gene activation domain includes the activation domain of VP64, the activation domain of P65, the activation domain of RTA, or the activation domain of HSF1, or any combination thereof.
[0036] In some embodiments, the activation domain of VP64 contains, or is a sequence of amino acids having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 26, and the activation domain of P65 contains, or is a sequence of amino acids having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 69. The sequence is such that the activation domain of the above RTA contains, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 11, and the activation domain of the above HSF1 contains, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 70.
[0037] In some embodiments, the amino acid sequence of the above NLS includes, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 24.
[0038] In some embodiments, the RZFD core domain and the NLS are linked via a linker, or the RZFD core domain and the gene activation domain are linked via a linker, or the NLS and the gene activation domain are linked via a linker.
[0039] In some embodiments, the artificial transcription factor comprises, or is a sequence of amino acids having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100%, identity with any one of the sequences shown in SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92.
[0040] In some embodiments, the coding nucleic acid sequence of the artificial transcription factor includes, or is a sequence of nucleotides having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence shown in any one of SEQ ID NOs: 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93.
[0041] In some embodiments, the expression vector contains, or is a sequence of nucleotides having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence shown in any one of SEQ ID NOs: 36, 37, or 62.
[0042] In a second aspect of the present invention, a composition is provided comprising (a) an artificial transcription factor capable of binding to RE1, comprising an RZFD core domain, a first NLS, and a first gene activation domain, or (b) a nucleic acid encoding the artificial transcription factor described in (a), wherein the RZFD core domain comprises any 5 to 8 sequences from SEQ ID NOs. 3 to 10, or any 5 to 8 sequences from SEQ ID NOs. 3 to 10, and does not include the sequences shown in SEQ ID NOs. 12 and 13. Preferably, the amino acid sequence of the RZFD core domain comprises any 6 to 8 sequences from SEQ ID NOs. 3 to 10, or any 6 to 8 sequences from SEQ ID NOs. 3 to 10.
[0043] In some embodiments, the amino acid sequence of the RZFD core domain includes, or is a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 76, preferably the amino acid sequence of the RZFD core domain is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, or at least 92% identity with SEQ ID NO: 74 or 75. , includes or is a sequence having at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 2.
[0044] In some embodiments, the amino acid sequence of the RZFD core domain includes, or is, a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 14.
[0045] In some embodiments, the coding nucleic acid sequence of the RZFD core domain includes, or is, a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with sequence number 15 or 64.
[0046] In some embodiments, the method for ligating the RZFD core domain, the first NLS, and the first gene activation domain from the N-terminus to the C-terminus is as follows:
[0047] 1) RZFD core domain -- first NLS -- first gene activation domain, 2) First gene activation domain -- first NLS -- RZFD core domain, 3) RZFD core domain -- first gene activation domain -- first NLS, 4) First NLS -- First gene activation domain -- RZFD core domain, 5) First NLS--RZFD core domain--First gene activation domain, 6) The first gene activation domain -- RZFD core domain -- is selected from the first NLS.
[0048] In some embodiments, the RZFD core domain, the first NLS, and the first gene activation domain may be linked via a linker or without a linker.
[0049] In some embodiments, the artificial transcription factor further comprises a second NLS. In specific embodiments, the first NLS and the second NLS may be the same or different.
[0050] In some embodiments, the first NLS is located at the N-terminus of the RZFD core domain, and the second NLS is located at the C-terminus of the RZFD core domain.
[0051] In some embodiments, the first gene activation domain is located at the N-terminus of the first NLS.
[0052] In some embodiments, the artificial transcription factor further comprises a second gene activation domain. In specific embodiments, the first NLS and the second NLS may be the same or different.
[0053] In some embodiments, the second gene activation domain and the first activation domain are located at opposite ends of the RZFD core domain, respectively.
[0054] In some embodiments, the second gene activation domain is located at the C-terminus of the second NLS.
[0055] In some embodiments, the first NLS comprises, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 24. In some embodiments, the coding nucleic acid sequence of the first NLS comprises, or is a sequence thereof, a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 25.
[0056] In some embodiments, the second NLS comprises, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 24. In some embodiments, the coding nucleic acid sequence of the second NLS comprises, or is a sequence thereof, a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 25.
[0057] In some embodiments, the first gene activation domain includes the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination thereof, preferably the first gene activation domain includes the activation domain of VP64, the activation domain of P65, the activation domain of RTA, or the activation domain of HSF1, or any combination thereof. In some embodiments, the amino acid sequence of the activating domain of VP64 has, or is, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 26, and the activating domain of P65 contains, or is, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 69. Furthermore, the amino acid sequence of the activation domain of the above-mentioned RTA has, or is, an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identical to SEQ ID NO: 11, and the amino acid sequence of the activation domain of the above-mentioned HSF1 has, or is, an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identical to SEQ ID NO: 70.In some embodiments, the coding nucleic acid sequence of the activating domain of VP64 includes, or is a sequence of, a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 27.
[0058] In some embodiments, the first gene activation domain comprises an activating domain for VP64 and an activating domain for P65, and preferably, the first gene activation domain further comprises an activating domain for HSF1 or an activating domain for RTA.
[0059] In some embodiments, the first gene activation domain comprises the VP64 activation domain, the P65 activation domain, and the HSF1 activation domain, and preferably the amino acid sequence of the first activation domain comprises, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 73.
[0060] In some embodiments, the first gene activation domain comprises the VP64 activation domain, the P65 activation domain, and the RTA activation domain, and preferably the amino acid sequence of the first activation domain comprises, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 72.
[0061] In some embodiments, the amino acid sequence of the first gene activation domain includes, or is, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NOs. 28, 69, 70, or 71.
[0062] In some embodiments, the coding nucleic acid sequence of the first gene activation domain comprises or is a sequence of nucleotides having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 29.
[0063] In some embodiments, the second gene activation domain comprises the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination thereof, preferably the second gene activation domain comprises the activation domain of VP64, more preferably the second gene activation domain comprises the activation domain of VP64 and the activation domain of P65, and even more preferably the second gene activation domain further comprises the activation domain of HSF1 or the activation domain of RTA. In some embodiments, the amino acid sequence of the VP64 activation domain has, or is, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 26. In some embodiments, the coding nucleic acid sequence of the activating domain of VP64 includes, or is a sequence of, a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 27.In some embodiments, the amino acid sequence of the activation domain of P65 contains, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 69, and in some embodiments, the amino acid sequence of the activation domain of RTA has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, or at least The amino acid sequence has, or is, an amino acid sequence having 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence number 70. In some embodiments, the amino acid sequence of the activation domain of HSF1 has, or is, an amino acid sequence having, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence number 70.
[0064] In some embodiments, the second gene activation domain includes the VP64 activation domain and the P65 activation domain, and preferably the second gene activation domain further includes the HSF1 activation domain or the RTA activation domain.
[0065] In some embodiments, the second gene activation domain comprises the VP64 activation domain, the P65 activation domain, and the HSF1 activation domain, and preferably the sequence of the second activation domain comprises, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 73.
[0066] In some embodiments, the second gene activation domain comprises the VP64 activation domain, the P65 activation domain, and the RTA activation domain, and preferably the sequence of the second activation domain comprises, or is a sequence thereof, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 72.
[0067] In some embodiments, the sequence of the second gene activation domain includes, or is a sequence of, an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NOs. 28, 69, 70, or 71.
[0068] In some embodiments, the artificial transcription factor of the composition contains, or is a sequence of amino acids having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100%, identity with the sequence shown in any one of SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92. In some embodiments, the coding nucleic acid sequence of the artificial transcription factor includes, or is a sequence of nucleotides having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NOs: 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93.
[0069] In some embodiments, the composition further comprises a regulatory element that modulates the nucleic acid expression of the artificial transcription factor, preferably the regulatory element comprising a promoter that modulates the nucleic acid expression of the artificial transcription factor, and more preferably the promoter comprising a glial cell-specific promoter. In some embodiments, the glial cell-specific promoter comprises an astrocyte-specific promoter, a Müller cell-specific promoter, or a cochlear glial cell-specific promoter, and preferably the glial cell-specific promoter comprises a GFAP promoter, an ALDH1L1 promoter, an EAAT1 / GLAST promoter, a glutamine synthetase promoter, an S100β promoter, an EAAT2 / GLT-1 promoter, a Plp1 promoter, or an Rlbp1 promoter.
[0070] In some embodiments, the regulatory element further comprises an enhancer that increases the transcription of the artificial transcription factor, preferably the enhancer comprising a CMV enhancer, an SV40 intron-type enhancer, an MVM intron-type enhancer, a β-globulin intron-type enhancer, or a synthetic intron-type enhancer, or any combination thereof. In some specific embodiments, the nucleic acid sequence of the CMV enhancer comprises or is a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 16. The nucleic acid sequence of the above SV40 intron type enhancer contains, or is a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 17. The nucleic acid sequence of the above MVM intron type enhancer contains, or is a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 18. The nucleic acid sequence of the above-mentioned β-globulin intron type enhancer contains, or is a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NO: 19.The nucleic acid sequence of the above-mentioned synthetic intron-type enhancer contains, or is, a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence shown in any one of SEQ ID NOs.
[0071] In some specific embodiments, the above-mentioned adjustment element may include one enhancer or may include multiple enhancers, and the multiple enhancers included may be the same or different.
[0072] In some embodiments, the regulatory element further comprises a post-transcriptional enhancing element (WPRE) of woodchuck hepatitis virus. For example, the regulatory element comprises a promoter and a WPRE that regulate the nucleic acid expression of an artificial transcription factor, or the regulatory element comprises a promoter, a WPRE, and an enhancer that increases the transcription of the artificial transcription factor.
[0073] In some embodiments, the composition comprises, or is a sequence thereof, a nucleic acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with SEQ ID NOs.
[0074] A third aspect of the present invention provides a method for differentiating non-neuronal cells of an individual into functional neurons, wherein an expression vector described in any one aspect of the first aspect of the present invention or a composition described in any one aspect of the second aspect of the present invention is brought into contact with the non-neuronal cells. In some specific embodiments, the contact is in vitro. In some specific embodiments, the contact is in vivo.
[0075] In some embodiments, the functional neurons include dopamine neurons, retinal ganglion cells, photoreceptor cells, cochlear spiral ganglion cells, GABA neurons, 5-HT neurons, glutamatergic neurons, ChAT neurons, NE neurons, motor neurons, spinal neurons, spinal motor neurons, spinal sensory neurons, bipolar cells, amacrine neurons, pyramidal cells, interneurons, medium spiny neurons (MSNs), Purkinje cells, granule cells, olfactory sensory neurons, juxtaglomerular cells, or any combination thereof. In some embodiments, the dopamine neurons express one or more markers from among NeuN, TH, FoxA2, Nurr1, Pitx3, Vmat2, or DAT. In some embodiments, the optic ganglion neurons express one or more markers from among RBPMS, Pax6, Brn3a, Brn3b, Brn3c, and Map2. In some embodiments, the photoreceptor cells express one or more markers from among rhodopsin, mCAR, m-opsin, and S-opsin. In some embodiments, the cochlear spiral ganglion cells express one or more markers from among NeuN, Prox1, Tuj-1, and Map2.
[0076] In some embodiments, the functional neuron includes an axon. In some embodiments, the non-neuronal cells include glial cells, fibroblasts, stem cells, neural progenitor cells, or neural stem cells.
[0077] In some embodiments, the glial cells include astrocytes, oligodendrocytes, ependymal cells, Schwann cells, NG2 cells, satellite cells, Müller cells, and cochlear nerve cells.
[0078] In some embodiments, the success rate of the conversion of non-neuronal cells into functional neurons is at least 1%.
[0079] In some embodiments, glial cells are located in the brain, spinal cord, eyes, or ears. In some embodiments, glial cells are located in the striatum, substantia nigra, ventral tegmental area of the midbrain, medulla oblongata, hypothalamus, dorsal midbrain, or cerebral cortex of the brain.
[0080] In some embodiments, glial cells include astrocytes, and functional neurons include dopaminergic neurons. In some embodiments, the method provided herein includes a method for transdifferentiating astrocytes into dopaminergic neurons in an individual. In some embodiments, astrocytes are located in the striatum and / or substantia nigra. In some embodiments, the method includes administering an expression vector or composition provided herein, or a product containing an expression vector or composition provided herein, to the striatum and / or substantia nigra of an individual.
[0081] In some embodiments, glial cells include Müller cells, and functional neuronal cells include retinal ganglion cells (RGCs) and / or photoreceptor cells. In some embodiments, the method provided herein relates to a method for transdifferentiating Müller cells into retinal ganglion cells (RGCs) and / or photoreceptor cells in an individual. In some embodiments, Müller cells are located subretinally or in the vitreous cavity. In some embodiments, the method includes administering an expression vector or composition provided herein, or a product containing an expression vector or composition provided herein, into the subretinal or vitreous cavity of an individual.
[0082] In some embodiments, glial cells include cochlear glial cells, and functional neuronal cells include cochlear spiral ganglion cells. In some embodiments, the methods provided herein relate to methods for transdifferentiating cochlear glial cells into cochlear spiral ganglion cells in an individual. In some embodiments, cochlear glial cells are located in the inner ear. In some embodiments, the methods involve administering an expression vector or composition provided herein, or a product containing an expression vector or composition provided herein, into the inner ear of an individual.
[0083] A fourth aspect of the present invention further provides a method for treating a disease in an individual requiring treatment, comprising administering a therapeutically effective amount of an expression vector described in any one embodiment of the first aspect of the present invention, or administering a therapeutically effective amount of a composition described in any one embodiment of the second aspect of the present invention.
[0084] In some embodiments, the diseases include Parkinson's disease, Alzheimer's disease, stroke, schizophrenia, Huntington's disease, depression, motor neuron disease, ALS, spinal muscular atrophy, epilepsy, ataxia, visual impairment due to RGC cell death, glaucoma, age-related RGC damage, optic nerve damage, focal ischemia or hemorrhage of the retina, hereditary optic nerve disease, degeneration or death of photoreceptor cells due to trauma or degeneration, macular degeneration, retinitis pigmentosa, diabetes-related blindness, night blindness, color blindness, hereditary blindness, spiral ganglion cell death, or any combination thereof.
[0085] In some embodiments, the disease includes Parkinson's disease. In some embodiments, the expression vectors and compositions described herein are administered to the brain of the individual in need. In some embodiments, the expression vectors and compositions described herein are administered to the striatum or substantia nigra of the individual in need.
[0086] In some embodiments, the disease includes optic nerve disease. In some embodiments, the expression vectors and compositions described herein are administered to the eye of the individual in need. In some embodiments, the expression vectors and compositions described herein are administered to the vitreous cavity or subretina of the individual in need. In some embodiments, the optic nerve disease includes, but is not limited to, visual impairment due to RGC cell death, glaucoma, or age-related RGC damage.
[0087] In this application, when it is stated that a predetermined domain is located at the N or C terminus of the RZFD core domain, it means that the predetermined domain is located on one side of the N terminus of the RZFD core domain, that is, it may be in close proximity to the N terminus of the RZFD core domain, or it may be located on one side of the N terminus of the RZFD core domain separated by another domain, or it may be in close proximity to the C terminus of the RZFD core domain, or it may be located on one side of the C terminus of the RZFD core domain separated by another domain.
[0088] For example, the fact that the NLS is located at the N-terminus and / or C-terminus of the RZFD core domain is understood to mean that the NLS is either closely adjacent to the N-terminus of the RZFD core domain, or located on one side of the N-terminus of the RZFD core domain separated by several amino acid sequences, or the NLS is either closely adjacent to the C-terminus of the RZFD core domain, or located on one side of the C-terminus of the RZFD core domain separated by several amino acid sequences.
[0089] The present application also relates to the use of the expression vector or composition for the conversion of non-neuronal cells into functional neurons, or to the use of the expression vector or composition for the production of a drug for the conversion of non-neuronal cells into functional neurons.
[0090] The above-described methods for differentiating non-neuronal cells into functional neurons apply to the use of non-neuronal cells for differentiating them into functional neurons, or to the use of drugs for differentiating non-neuronal cells into functional neurons.
[0091] The present application also relates to the use of the expression vector or composition for treating a disease, or for manufacturing a therapeutic agent for a disease. The diseases include Parkinson's disease, Alzheimer's disease, stroke, schizophrenia, Huntington's disease, depression, motor neuron disease, ALS, spinal muscular atrophy, epilepsy, ataxia, visual impairment due to RGC cell death, glaucoma, age-related RGC damage, optic nerve damage, focal ischemia or hemorrhage of the retina, hereditary optic nerve disease, degeneration or death of photoreceptor cells due to trauma or degeneration, macular degeneration, retinitis pigmentosa, diabetes-related blindness, night blindness, color blindness, hereditary blindness, spiral ganglion cell death, or any combination thereof.
[0092] The present invention also relates to the artificial transcription factor comprising an RZFD core domain, wherein the amino acid sequence of the RZFD core domain comprises any 5 to 8 sequences from SEQ ID NOs. 3 to 10 and does not include the sequences shown in SEQ ID NOs. 12 and 13, and preferably, the amino acid sequence of the RZFD core domain comprises any 6 to 8 sequences from SEQ ID NOs. 3 to 10. Any of the artificial transcription factors described in the other embodiments above are applicable to the artificial transcription factor described in this embodiment.
[0093] Through the detailed description, additional aspects and advantages of the present invention will be apparent to those skilled in the art, and the title of this method discloses and describes only exemplary embodiments. As should be understood, the title of this method includes other different embodiments, and some details further include various obvious substitutions. All of these are within the scope of this disclosure. Accordingly, the drawings and description should be considered exemplary and not limiting.
[0094] The present invention further relates to the following: 1. An expression vector comprising (a) a mutant encoding the RE1 silencing transcription factor REST, which includes the DNA-binding domain of REST; (b) a promoter; and (c) an enhancer. 2. The expression vector according to item 1, wherein the enhancer comprises a CMV enhancer, an SV40 intron-type enhancer, an MVM intron-type enhancer, a β-globulin intron-type enhancer, or a synthetic intron-type enhancer. 3. The expression vector according to item 2, wherein the CMV enhancer comprises a nucleic acid sequence having at least 70% similarity to SEQ ID NO: 16. 4. The expression vector according to item 2, wherein the SV40 intron-type enhancer comprises a nucleic acid sequence having at least 70% similarity to SEQ ID NO: 17. 5. The expression vector according to item 2, wherein the MVM intron-type enhancer comprises a nucleotide sequence having at least 70% homology to SEQ ID NO: 18. 6. The expression vector according to item 2, wherein the β-globulin intron-type enhancer comprises a nucleotide sequence having at least 70% homology with SEQ ID NO: 19. 7. The expression vector according to item 2, wherein the synthetic intron-type enhancer comprises a nucleotide sequence having at least 70% homology with any one of SEQ ID NOs: 20-22. 8. The expression vector according to any one of items 1-7, wherein the promoter comprises a glial cell-specific promoter. 9. The expression vector according to item 8, wherein the glial cell-specific promoter comprises an astrocyte-specific promoter or a Müller cell-specific promoter. 10. The expression vector according to item 8, wherein the glial cell-specific promoter comprises a GFAP promoter, an ALDH1L1 promoter, an EAAT1 / GLAST promoter, a glutamine synthetase promoter, an S100β promoter, an EAAT2 / GLT-1 promoter, or an Rlbp1 promoter. 11. In some embodiments, the expression vector according to item 8, wherein the glial cell-specific promoter comprises a cochlear glial cell-specific promoter. 12. The cochlear glial cell-specific promoter is the expression vector described in item 11, comprising the GFAP promoter, ALDH1L1 promoter, EAAT1 / GLAST promoter, or Plp1 promoter.13. An expression vector according to any one of items 1 to 12, further comprising a post-transcriptional enhancing element (WPRE) of woodchuck hepatitis virus. 14. An expression vector according to any one of items 1 to 13, wherein the DNA-binding domain of REST comprises an RE1-binding domain. 15. An expression vector according to item 14, wherein the RE1-binding domain comprises a sequence having at least 70% homology to SEQ ID NO: 2 or 14. 16. An expression vector according to item 15, wherein the RE1-binding domain is encoded by a sequence having at least 70% homology to SEQ ID NO: 15. 17. An expression vector according to any one of items 1 to 16, further comprising a nuclear localization sequence (NLS). 18. An expression vector according to item 17, wherein the NLS is located at the N-terminus or C-terminus of the DNA-binding domain of REST. 19. An expression vector according to item 18, wherein the NLS is located at both the N-terminus and C-terminus of the DNA-binding domain of REST. 20. An expression vector according to item 17, wherein the NLS has an amino acid sequence having at least 70% identity with SEQ ID NO: 24. 21. An expression vector according to item 20, wherein the NLS is encoded by a nucleotide sequence having at least 70% identity with the sequence shown in SEQ ID NO: 25. 22. An expression vector according to any one of items 1 to 21, wherein the REST variant further comprises a gene activation domain. 23. An expression vector according to item 22, wherein the gene activation domain comprises VP64, P65-HSF1, VP16, RTA, Suntag, P300, CBP, or any combination thereof. 24. An expression vector according to item 23, wherein the gene activation domain comprises VP64. 25. An expression vector according to item 24, wherein VP64 is located at the C-terminus or N-terminus of the DNA-binding domain of REST. 26. An expression vector according to item 25, wherein VP64 is located at both the C-terminus and N-terminus of the DNA-binding domain of REST. 27. VP64 is an expression vector according to any one of items 23-26, having an amino acid sequence that is at least 70% identical to sequence number 26. 28. VP64 is an expression vector according to item 27, encoded by a nucleotide sequence that is at least 70% identical to sequence number 27.29. The gene activation domain described above is P65-HSF1, and the expression vector described in item 23. 30. The gene activation domain described above is located at the N-terminus or C-terminus of the DNA binding domain of REST, and the expression vector described in item 29. 31. The gene activation domain described above is P65-HSF1, and the expression vector described in item 30, and the P65-HSF1 is located at the N-terminus and C-terminus of the DNA binding domain of REST. 32. The gene activation domain described above is P65-HSF1, and the expression vector described in any one of items 29 to 31 has an amino acid sequence that is at least 70% identical to SEQ ID NO: 28. 33. The gene activation domain described above is P65-HSF1, and the expression vector described in item 32 is encoded by a nucleotide sequence that is at least 70% identical to SEQ ID NO: 29. 34. The gene activation domain described above is P65-HSF1, and the expression vector described in any one of items 1 to 33 further comprises NLS and a gene activation domain. 35. The gene activation domain described above is VP64, and the expression vector described in item 34. 36. The expression vector described in item 35, wherein the above NLS and VP64 are located at the N-terminus and C-terminus of the DNA-binding domain of REST. 37. The expression vector described in item 36, wherein the above REST variant has an amino acid sequence having at least 70% identity with SEQ ID NO: 32. 38. The expression vector described in item 37, wherein the above REST variant is encoded by a nucleotide sequence having at least 70% identity with SEQ ID NO: 33. 39. The expression vector described in item 34, wherein the above gene activation domain contains P65-HSF1. 40. The expression vector described in item 39, wherein the above NLS and P65-HSF1 are located at the N-terminus and C-terminus of the DNA-binding domain of REST. 41. The expression vector described in any one of items 1 to 40, wherein the above REST variant lacks the inhibitor domains of REST at the N-terminus and C-terminus. 42. The expression vector according to item 40, wherein the REST variant has an amino acid sequence that is at least 70% identical to SEQ ID NOs. 34, 67, 84, 86, 90, or 92. 43. The expression vector according to item 40, wherein the REST variant is encoded by a nucleotide sequence that is at least 70% identical to SEQ ID NOs. 35, 68, 85, 87, 91, or 93.44. A composition comprising (a) a REST variant comprising a DNA-binding domain of REST, a first NLS, a second NLS, and a first gene-activating domain, or (b) a nucleic acid encoding a mutant of REST. 45. The composition according to item 44, wherein the DNA-binding domain comprises an RE1-binding domain. 46. The composition according to item 44, wherein the first NLS is located at the N-terminus of the DNA-binding domain of REST. 47. The composition according to item 46, wherein the second NLS is located at the C-terminus of the DNA-binding domain of REST. 48. The composition according to item 46, wherein the first gene-activating domain is located at the N-terminus of the first NLS. 49. The composition according to any one of items 44 to 47, wherein the REST variant further comprises a second gene-activating domain. 50. The composition according to item 49, wherein the second gene-activating domain is located at the C-terminus of the second NLS. 51. The composition according to any one of items 44 to 50, wherein the DNA-binding domain of REST comprises an amino acid sequence having at least 70% identity with SEQ ID NO: 2 or 14. 52. The composition according to item 51, wherein the DNA-binding domain is encoded by a nucleotide sequence having at least 70% identity with SEQ ID NO: 15. 53. The composition according to any one of items 44 to 50, wherein the first NLS comprises an amino acid sequence having at least 70% identity with SEQ ID NO: 24. 54. The composition according to item 53, wherein the first NLS is encoded by a sequence containing nucleotides having at least 70% identity with SEQ ID NO: 25. 55. The composition according to any one of items 44 to 54, wherein the second NLS comprises an amino acid sequence having at least 70% identity with SEQ ID NO: 24. 56. The composition according to item 55, wherein the second NLS is encoded by a sequence containing nucleotides having at least 70% identity with SEQ ID NO: 25. 57. A composition according to any one of items 44 to 56, wherein the first gene activation domain comprises VP64. 58. A composition according to item 57, wherein VP64 has an amino acid sequence having at least 70% identity with SEQ ID NO: 26. 59. A composition according to item 58, wherein VP64 is encoded by a nucleotide sequence having at least 70% identity with SEQ ID NO: 27.60. The composition according to any one of items 44 to 56, wherein the first gene activation domain is P65-HSF1. 61. The expression vector according to item 60, wherein the P65-HSF1 has an amino acid sequence having at least 70% identity with SEQ ID NO: 28. 62. The expression vector according to item 61, wherein the P65-HSF1 is encoded by a nucleotide sequence having at least 70% identity with SEQ ID NO: 29. 63. The composition according to item 49, wherein the second gene activation domain is VP64. 64. The composition according to item 63, wherein the VP64 has an amino acid sequence having at least 70% identity with SEQ ID NO: 26. 65. The composition according to item 64, wherein the VP64 is encoded by a nucleotide sequence having at least 70% identity with SEQ ID NO: 27. 66. The composition according to item 49, wherein the second gene activation domain is P65-HSF1. 67. The composition according to item 66, wherein P65-HSF1 has an amino acid sequence having at least 70% identity with SEQ ID NO: 28. 68. The expression vector composition according to item 67, wherein P65-HSF1 is encoded by a nucleotide sequence having at least 70% identity with SEQ ID NO: 29. 69. The composition according to any one of items 44 to 68, wherein the REST variant has an amino acid sequence having at least 70% identity with SEQ ID NO: 32, 34, 67, 84, 86, 90, or 92. 70. The composition according to item 69, wherein the REST variant is encoded by a nucleotide sequence having at least 70% identity with SEQ ID NO: 33, 35, 68, 85, 87, 91, or 93. 71. A method for converting non-neuronal cells into functional neurons in a single organism, comprising administering to the organism an expression vector according to any one of items 1 to 43 or a composition according to any one of items 44 to 66. 72. A method for in vitro converting non-neuronal cells into functional neurons, comprising administering to cells an expression vector described in any one of items 1 to 43 or a composition described in any one of items 44 to 66.73. The method according to item 71 or 72, wherein the functional neurons include dopamine neurons, retinal ganglion cells, photoreceptor cells, cochlear spiral ganglion cells, GABA neurons, 5-HT neurons, glutamatergic neurons, ChAT neurons, NE neurons, motor neurons, spinal neurons, spinal motor neurons, spinal sensory neurons, bipolar cells, amacrine neurons, pyramidal cells, interneurons, medium spiny neurons (MSNs), Purkinje cells, granule cells, olfactory sensory neurons, juxtaglomerular cells, or any combination thereof. 74. The method according to item 73, wherein the dopamine neurons express one or more markers selected from NeuN, TH, FoxA2, Nurr1, Pitx3, Vmat2, or DAT. 75. The method according to item 73, wherein the optic ganglion neurons express one or more markers selected from RBPMS, Pax6, Brn3a, Brn3b, Brn3c, and Map2. 76. The photoreceptor cells described above express one or more markers selected from rhodopsin, mCAR, m-opsin, and S-opsin, according to the method of item 73. 77. The cochlear spiral ganglion cells described above express one or more markers selected from NeuN, Prox1, Tuj-1, and Map2, according to the method of item 73. 78. The functional neurons described above include axons, according to the method of any one of items 71 to 77. 79. The non-neuronal cells described above include glial cells, fibroblasts, stem cells, neural progenitor cells, or neural stem cells, according to the method of any one of items 71 to 77. 80. The glial cells described above include astrocytes, oligodendrocytes, ependymal cells, Schwann cells, NG2 cells, satellite cells, Müller cells, and cochlear nerve glial cells, according to the method of item 79. 81. The method according to any one of items 71-80, wherein the efficacy rate of the conversion of non-neuronal cells to functional neurons is at least 1%. 82. A method for treating a disease in an individual in need, comprising administering a therapeutically effective amount of an expression vector according to any one of items 1-43 or a composition according to any one of items 44-70.83. The above diseases include Parkinson's disease, Alzheimer's disease, stroke, schizophrenia, Huntington's disease, depression, motor neuron disease, ALS, spinal muscular atrophy, Pick's disease, sleep disorders, epilepsy, ataxia, visual impairment due to RGV cell death, glaucoma, age-related RGC damage, optic nerve damage, and focal ischemia or effusion of the retina. The method described in item 82, including blood, hereditary optic nerve disease, degeneration or death of photoreceptor cells due to trauma or degeneration, macular degeneration, retinitis pigmentosa, diabetes-related blindness, night blindness, color blindness, hereditary blindness, spiral ganglion cell death, or any combination thereof.
[0095] Built-in reference information All publications, patents, and patent applications referenced herein are incorporated herein by reference to the same extent as each individual publication, patent, or patent application is explicitly and individually cited and incorporated herein by reference. Where there is any conflict between the publications, patents, or patent applications incorporated by reference and the disclosures contained in the specification, the specification may substitute and / or prevail over any such conflicting material. [Brief explanation of the drawing]
[0096] [Figure 1]This shows the in vivo differentiation-transition effects of different versions of artificial transcription factors. (Figure 1A) Schematic diagram of the REST protein functional domain and schematic diagrams of the designs of different versions of artificial transcription factors, of which NLS, VP16, and VP64 are sequences that are increased. (Figure 1B) Schematic diagram of the design of each AAV expression vector, using the GFAP promoter to achieve specific expression in astrocytes, of which mCherry is a red fluorescent protein coding sequence, and dRT-V1, dRT-V2, dRT-V3, and dRT-V4 are different versions of artificial transcription factors. (Figure 1C) In vivo differentiation-transition results of different versions of artificial transcription factors, of which design-1 is the control group, and designs-2 to design-5 are representative diagrams of the in vivo differentiation-transition results of each artificial transcription factor, red fluorescence is the fluorescence expressed by mCherry, NeuN is a neuron-specific protein marker stain, TH is a dopamine neuron-specific protein marker stain, and DAPI is the cell nucleus. The scale is 50 μm. (Figure 1D) This shows the percentage of astrocytes that differentiate into neurons in each group. Design-1 is the control group. Less than 2% of mCherry+NeuN double-positive cells were present in Design-1, approximately 5% in Design-2 and Design-3, and a clear improvement in mCherry+NeuN double-positive cells was observed in Design-4 and Design-5. As can be seen from the statistical results for dopamine neuron-specific protein marker TH staining, there were no TH-positive cells in Design-1, Design-2, and Design-3, while TH-positive cells were present in Design-4 and Design-5, and the number of TH-positive cells in Design-5 was approximately twice that of Design-4. Difference analysis was performed using a t-test. Error lines were obtained using SEM. * indicates P<0.05, ** indicates P<0.01, and *** indicates P<0.001. [Figure 2]This is a schematic diagram of a combination of artificial transcription factors. RZFD is a DNA-binding sequence, NLS is a nuclear localization sequence, TAD is a transcriptional activation domain, and EPD is an epigenetic regulatory domain. RZFD may be used alone, or it may be arbitrarily combined with one or more of NLS, TAD, or EPD. Of these, NLS, TAD, EPD, or combinations thereof may be located at the N-terminus or C-terminus of RZFD. [Figure 3] These are exemplary vector designs. (Figure 3A) Different vector designs are shown, using the astrocyte-specific promoter GFAP to promote the expression of artificial transcription factors (i.e., genes of interest, GOI in the figure) (plasmid design 1), or adding a WPRE regulatory sequence to the artificial transcription factor (plasmid design 2), or adding an enhancer before the promoter (plasmid design 3), or increasing the enhancer and WPRE (plasmid design 4), or increasing a different enhancer (plasmid design 5). (Figure 3B) Designs for each artificial transcription factor (ATF) used in the examples are shown, and each ATF is named RZFD-V1 to RZFD-V12. [Figure 4]This is a comparison of the effects of dopamine neurons generated by in vivo transdifferentiation using different vector designs. (Figure 4A) Schematic diagrams of each vector design used in this experiment (Figure 4B) are shown. Design-V1 promotes RZFD-V5 using the glial cell-specific promoter GFAP, Design-V2 adds two enhancers to Design-V1 (including a CMV enhancer and an MVM intronic enhancer), Design-V3 expresses RZFD-V6 compared to Design-V2, Design-V2 expresses RZFD-V5, Design-V3 adds one VP64 at each end of RZFD-V5, and the GOI in Design-V4 is RZFD-V2, compared to RZFD-V6, RZFD-V2 uses only one VP64 and one NLS located at each end of RZFD. (Figure 4B) This figure shows representative diagrams of the differentiation-transition effects of various designs in Figure 4A. The injection given to the control group was AAV-pGFAP-mCherry, and designs V1 to V4 were AAV-pGFAP-dRT-V5, AAV-CMV enhancer-pGFAP-MVM-dRT-V5, AAV-CMV enhancer-pGFAP-MVM-dRT-V6, and AAV-CMV enhancer-pGFAP-MVM-dRT-V2. TH is a dopamine neuron-specific protein marker, DAPI is a cell nuclear stain, and white arrows indicate TH-positive neurons. The scale is 50 μm. (Figure 4C) This figure shows the statistical results of the number of TH-positive cells in each group, with design V1 as the baseline and each group being a multiple of the number of TH-positive cells. The control group had no TH-positive cells. The number of TH-positive cells in Design-V2 (increased enhancer) was approximately twice that of Design-V1. The number of TH-positive cells in Design-V3 was approximately four times that of Design-V1 and approximately twice that of Design-V2. The number of TH-positive cells in Design-V4 was close to that of Design-V3, and in all cases approximately twice that of Design-V2 and approximately four times that of Design-V1. Difference analysis was performed using a t-test. Error lines were generated using SEM. * indicates P<0.05, ** indicates P<0.01, and *** indicates P<0.001. [Figure 5]This is a comparison of the differentiation and conversion effects of different RZFD cleavage versions. (Figure 5A) shows schematic diagrams of the designs of each vector used in this experiment (Figure 5B). The ATF region of design-V5 is NLS-RZFD(159-412) containing eight zinc finger domains, the ATF region of design-V6 is NLS-RZFD(209-412) containing seven zinc finger domains (ZF2-8), the ATF region of design-V7 is NLS-RZFD(159-388) containing seven zinc finger domains (ZF1-7), and the ATF region of design-V8 is NLS-RZFD(209-388) containing six zinc finger domains (ZF2-7). The design of the rest of the vector is consistent. (Figure 5B) This figure shows representative diagrams of the differentiation conversion effects of various ATF designs in Figure 5A. TH is a dopamine neuron-specific protein marker, DAPI is cell nuclear staining, and white arrows point to TH-positive neurons. Designs V5 to V8 can all differentiate astrocytes into dopamine neurons. The scale is 50 μm. (Figure 5C) This figure shows the statistical results of the number of TH-positive cells in each group, with design V2 as the baseline, and the number of TH-positive cells in each other group being a multiple of the number of TH-positive cells in design V2. Designs V5 to V8 are close to the number of TH-positive cells generated by design V2, and there is no statistically significant difference. Difference analysis is performed using a t-test. The error lines are SEM. * indicates P<0.05, ** indicates P<0.01, and *** indicates P<0.001. [Figure 6]This is a comparison of the differentiation-transformational effects of different activation domains. (Figure 6A) Schematic diagrams of the designs of each vector used in this experiment (Figure 6B) are shown. Using design-V2 as the baseline, the ATF of design-V9 is RZFD-V9 and the activation domain is VP64-p65-HSF1, the ATF of design-V10 is RZFD-V10 and the activation domain is VP64-p65-Rta, and the ATF of design-V11 is RZFD-V11 and the activation domain is VP64. (Figure 6B) Representative diagrams of the differentiation-transformational effects of various ATF designs in Figure 6A are shown. TH is a dopamine neuron-specific protein marker, DAPI is cell nuclear staining, and white arrows point to TH-positive neurons. Designs-V2 to V10 can all generate dopamine neurons through differentiation-transformational effects. The scale is 50 μm. (Figure 6C) The statistical results for the number of TH-positive cells in each group are shown, with design-V2 as the baseline, and the numbers of TH-positive cells in each other group are shown as multiples of the number of TH-positive cells in design-V2. Design-V9 is close to the number of TH-positive cells generated in design-V10, at 1.42 times and 1.47 times that of design-V2, respectively, while the number of TH-positive cells generated in design-V10 is even higher, at 1.99 times that of design-V2. Difference analysis is performed using a t-test. Error lines are from SEM. * indicates P<0.05, ** indicates P<0.01, and *** indicates P<0.001. [Figure 7]This is a comparison of the differentiation-transition effects of RZFD-V6 and Ptbp1 expression knockdown. (Figure 7A) Schematic diagrams of the designs of each vector used in this experiment (Figure 7B) are shown. In the Ptbp1-KD vector design, U6 promotes gRNA expression and GFAP promotes CasRx expression, and the design of GFAP and enhancer is consistent with the design of the RZFD-V6 expression vector (design-V3 in this study). (Figure 7B) Representative figures and statistical results of the differentiation-transition effects of Ptbp1 knockdown or RZFD-V6 overexpression are shown. TH is a dopamine neuron-specific protein marker, and white arrows point to TH-positive neurons. The number of TH-positive neurons in the RZFD-V6 group is approximately twice that of the Ptbp1-KD group. The scale is 50 μm. Difference analysis is performed using a t-test. Error lines are SEM. * indicates P<0.05, ** indicates P<0.01, and *** indicates P<0.001. [Figure 8] This study investigates the differentiation effect of RZFD-V6 using 6-OHDA modeling in Dat-Cre::Ai9 strain tracking mice. (Figure 8A) A schematic diagram of the vector design is shown. Vector 1 is AAV used for labeling, promoting EGFP expression in astrocytes using the astrocyte-specific promoter GFAP, and Vector 2 promotes the expression of the artificial transcription factor (ATF:RZFD-V6) using the astrocyte-specific promoter GFAP. (Figure 7B) Dat-Cre::Ai9 mice were used to label mature dopamine neurons, and after AAV injection into 6-OHDA modeled Dat-Cre::Ai9 mice, astrocytes differentiated into neurons or dopaminergic neurons. A representative cell staining diagram is shown, where red fluorescence is tdTomato fluorescence expressed by mature dopamine neurons Dat-Cre::Ai9, white signal is NeuN staining, a neuron-specific protein marker, and white arrows point to tdTomato-positive mature dopamine neurons. The scale is 50 μm. [Figure 9]This study explores the therapeutic effect of the artificial transcription factor RZFD-V6 on Parkinson's disease (PD) mice. AAV containing RZFD-V6 was injected into a 6-OHDA modeled Parkinson's mouse model, and behavioral analysis of the therapeutic effect was performed 1.5 to 3 months later. In the control group, there was no difference before and after treatment, while the group injected with the artificial transcription factor RZFD-V6 showed a clear improvement in behavior after treatment. [Modes for carrying out the invention]
[0097] The various embodiments disclosed and described below are merely illustrative. Those skilled in the art can make various modifications and substitutions without departing from the technical concept of this application. It should be understood that various substitutions of the specific embodiments described herein should also be included.
[0098] This application provides an expression vector and composition capable of modulating inhibitor element 1 / neuron restriction silencer element (RE1 / NRSE). This application further provides methods and uses of the expression vector and composition for the conversion of non-neuronal cells into functional neurons, and uses and methods of the expression vector and composition for the treatment of diseases, particularly neurological diseases.
[0099] Repressor element 1 / neuron-restrictive silencer element (RE1 / NRSE) is a specific DNA sequence, also called RE1 in textbooks, approximately 21 bp long (extending from 20 to 23 bp), and primarily binds to REST (RE1 silencing transcription factor, also called neuron-restrictive silencer factor (NRSF)), regulating the expression of genes related to neuronal development and maturation. RE1 is a negative regulator involved in neurogenesis, first found at the 5' end of the promoters of NaV1.2 and SCG10, and can regulate the expression of genes related to neuronal development and maturation. In non-neuronal cells, the RE1 site binds to a silencing complex consisting of histone deacetylase and methylase, inhibiting the expression of neuron-related genes. However, because there are more than 1800 RE1 elements in mice and humans, regulation based on conventional techniques is difficult. For example, while CRISPR-mediated gene regulation and epigenetic modification technologies offer very high precision and can accurately regulate the expression of specific genes, it is difficult to regulate the expression of RE1-regulated genes in this manner.
[0100] Neurological disorders refer to diseases associated with neuronal death or reduction, including Parkinson's disease, Alzheimer's disease, stroke, schizophrenia, Huntington's disease, depression, motor neuron disease, ALS, spinal muscular atrophy, epilepsy, ataxia, visual impairment due to RGC cell death, glaucoma, age-related RGC damage, optic nerve damage, focal ischemia or hemorrhage of the retina, hereditary optic nerve disease, degeneration or death of photoreceptor cells due to trauma or degeneration, macular degeneration, retinitis pigmentosa, diabetes-related blindness, night blindness, color blindness, hereditary blindness, spiral ganglion cell death, and other conditions.
[0101] Motor neuron diseases (MNDs) refer to diseases in which the motor neurons in the human body are damaged and gradually progress, including amyotrophic lateral sclerosis (ALS, also known as Lou Gehrig's disease), progressive bulbar palsy, primary lateral sclerosis, and progressive muscular atrophy.
[0102] The expression vectors and compositions provided herein can modulate the binding of REST and RE1 / NRSE and enable the transdifferentiation of glial cells into neurons. In some embodiments, they can induce the transdifferentiation of astrocytes into dopamine neurons. In some embodiments, they can induce the transdifferentiation of Müller cells into photoreceptor cells or optic ganglion cells.
[0103] In some embodiments, the expression vectors and compositions provided herein can convert non-neuronal cells into functional neuronal cells in vivo in an individual. In some embodiments, the expression vectors and compositions can be administered to a site of interest in the body (e.g., a site affected by disease) to block the binding of REST to the RE1 / NRSE element, thereby enabling the conversion of non-neuronal cells into functional neurons.
[0104] In some embodiments, the expression vector provided herein comprises (a) a nucleic acid sequence encoding an artificial transcription factor capable of binding to RE1, including an RZFD core domain, and (b) a regulatory element that modulates the nucleic acid expression of the artificial transcription factor, wherein the RZFD core domain includes any 5 to 8 sequences from SEQ ID NOs. 3 to 10, but does not include the sequences shown in SEQ ID NOs. 12 and 13. In some embodiments, the amino acid sequence of the RZFD core domain includes any 6 to 8 sequences from SEQ ID NOs. 3 to 10. In some examples, the amino acid sequence of the RZFD core domain may include any of the sequences shown in SEQ ID NOs. 2, 14, 74, 76, and 75. In some embodiments, the amino acid sequence of the RZFD core domain may include a sequence that has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence shown in any of SEQ ID NOs: 2, 14, 74, 76, or 75.
[0105] In some embodiments, the composition provided herein comprises (a) an artificial transcription factor comprising an RZFD core domain, a first NLS, and a first gene activation domain, or (b) a nucleic acid encoding an artificial transcription factor, wherein the RZFD core domain comprises any 5 to 8 sequences from SEQ ID NOs. 3 to 10, but does not include the sequences shown in SEQ ID NOs. 12 and 13. In some embodiments, the amino acid sequence of the RZFD core domain comprises any 6 to 8 sequences from SEQ ID NOs. 3 to 10. In some examples, the amino acid sequence of the RZFD core domain may include any of the sequences shown in SEQ ID NOs. 2, 14, 74, 76, and 75. In some embodiments, the amino acid sequence of the RZFD core domain may include a sequence that has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence shown in any of SEQ ID NOs: 2, 14, 74, 76, or 75.
[0106] The composition or expression vector provided herein further includes optimization of the artificial transcription factor, such as the addition of NLS or gene activation domains, thereby enabling differentiation from non-neuronal cells to functional neurons.
[0107] It should be understood that the numbering of specific positions or residues in each sequence of this application is determined by the specific protein and numbering scheme used. Those skilled in the art can identify any homologous protein and corresponding residues in the corresponding coding nucleic acid by methods well known in the art, such as sequence alignment and homologous residue measurement.
[0108] definition The terms “at least,” “greater than,” or “greater than or equal to” apply to all numbers in a sequence of numbers when the first number is followed by two or more numbers. For example, 1, 2, or 3 or more is equivalent to 1 or more, 2 or more, or 3 or more.
[0109] The terms “not exceeding,” “less than,” and “less than or equal to” apply to all numbers in a sequence of numbers when the first number is followed by two or more numbers. For example, 3, 2, or 1 or less is equivalent to 3 or less, 2 or less, or 1 or less.
[0110] Where used herein, the singular form should also include the plural form unless otherwise explicitly indicated in the context. Furthermore, where the terms “include,” “incorporate,” “have,” “are,” “attach,” or any other terms having the same meaning are used in the detailed specification and / or claims, these terms are intended to have an inclusive meaning in a manner similar to that of the term “include.”
[0111] The terms “determination,” “measurement,” “evaluation,” “assessment,” and “analysis” are often used interchangeably herein and refer to forms of measurement. These terms include determining whether an element is present or not (e.g., detection). These terms may include quantitative, qualitative, or quantitative and qualitative measurements. Assessments may be relative or absolute. In this specification, “detection” may include detecting the quantity of something that is present, and further include detecting whether it is present or absent.
[0112] The terms “approximately” or “nearly” refer to the acceptable range of error for a particular value as determined by those skilled in the art, which is determined in part by how the value is measured, for example, by the limitations of the measuring system. For example, “approximately” can mean that in practice a given value, it is within one or more standard deviations. Where a particular value is described in an application or claim, unless otherwise specified, the term “approximately” should be assumed to refer to the acceptable range of error for that value.
[0113] As used herein, the phrases “at least one,” “one or more,” and “and / or” are open forms of expression in operations of integration or separation. For example, in “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, and C,” “one or more of A, B, or C,” and “A, B, and / or C,” each represents A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.
[0114] As used herein, the term “nucleic acid” refers to a polymer comprising at least two nucleotides (e.g., deoxyribonucleotides or ribonucleotides) formed in single-stranded or double-stranded form, including DNA and RNA. A “nucleotide” comprises one deoxyribose (DNA) or ribose (RNA), one base, and one phosphate group. Nucleotides are linked by phosphate groups. A “base” comprises purines and pyrimidines, and further comprises natural adenine, thymine, guanine, cytosine, uracil, inosine, and natural analogs, as well as synthetic derivatives of purines and pyrimidines, and includes, but is not limited to, novel substitutions of reactive groups, such as amines, alcohols, thiols, carboxylates, and alkyl halides. Nucleic acids include nucleic acids comprising known nucleotide analogs or modified skeletal residues or bonds, which are synthetic, naturally occurring, and non-natural, and which have similar binding properties to the reference nucleic acid. Examples of such analogues and / or modified residues include, but are not limited to, phosphorothioates, phosphoramidates, methylphosphonates, chiral methylphosphonates, 2'-O-methylribonucleotides, and peptide nucleic acids (PNAs).
[0115] As used herein, the terms “protein,” “polypeptide,” and “peptide” may be used interchangeably and refer to a polymer of amino acid residues linked by peptide bonds, which may consist of two or more polypeptide chains. The terms “polypeptide,” “protein,” and “peptide” refer to a polymer in which at least two amino acid monomers are linked by amide bonds. The amino acids may be L-optical isomers or D-optical isomers. More specifically, the terms “polypeptide,” “protein,” and “peptide” refer to a molecule consisting of two or more amino acids in a specific order, for example, the order determined by the sequence of nucleotides in a protein-coding gene or RNA. In some cases, a protein may be a fragment of a protein, for example, a protein domain, subdomain, or motif. In some cases, a protein may be a variant (or mutant) of a protein having one or more insertions, deletions, and / or substitutions of amino acid residues in the amino acid sequence of a naturally occurring (or at least known) protein. A polypeptide may be a single linear polymer chain linked by peptide bonds between the carboxyl and amino groups of adjacent amino acid residues. Polypeptides can be modified, for example, by adding carbohydrates or by phosphorylation. Proteins can contain one or more polypeptides.
[0116] Proteins or their variants may be naturally occurring or recombinant. Methods for detecting and / or measuring polypeptides in biomaterials are well known in the art and include, but are not limited to, Western blotting, flow cytometry, ELISA, RIA, and various proteomics techniques. An exemplary method for measuring or detecting polypeptides is immunoassay, e.g., ELISA. This type of protein quantification can be based on an antibody capable of capturing a specific antigen and a second antibody capable of detecting the captured antigen.
[0117] As used herein, the term "fragment" or equivalent term may refer to a portion of a protein whose length is less than the total length of the protein and which optionally maintains the function of the protein. Furthermore, when aligning such a portion of a protein with the protein, the sequence of the portion of the protein may have, for example, at least 80% identity with the sequence of the portion of the protein.
[0118] The ranges provided herein are understood to be abbreviations for all values within that range. For example, the range 1 to 50 is understood to include any digit, combination of digits or subranges from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50, as well as all intermediate decimal values between the above integers such as 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8 and 1.9. The range further includes subranges and "nested subranges" extending from any one endpoint of the range. For example, a nested subrange of 1-50 may include 1-10, 1-20, 1-30, and 1-40 in one direction, or 50-40, 50-30, 50-20, and 50-10 in another direction.
[0119] Sequence identity can be measured using sequence analysis software (e.g., BLAST, BESTFIT, GAP, or PILEUP / PRITTYBOX programs). Such software matches similar or identical sequences by evaluating the identity of various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions at the following groups: glycine, alanine, valine, isoleucine, leucine, aspartic acid, glutamic acid, asparagine, glutamine, serine, threonine, lysine, arginine, and phenylalanine, tyrosine.
[0120] The term "nuclear localization sequence," or NLS, also known as a nuclear localization signal, is typically a single short amino acid sequence that, after fusing with a protein, can guide that protein into the cell nucleus. Non-exclusive examples of NLS include NLS sequences derived from: the SV40 viral large T antigen (e.g., the sequence shown in SEQ ID NO: 38), the nucleoplasmin NLS (e.g., the sequence shown in SEQ ID NO: 39), the c-myc NLS (e.g., the sequence shown in SEQ ID NO: 40 or 41), the hRNPA1 M9 NLS (e.g., the sequence shown in SEQ ID NO: 42), the sequence derived from the IBB domain of input protein-α (e.g., the sequence shown in SEQ ID NO: 43), the myoma T protein sequence (e.g., the sequence shown in SEQ ID NO: 44 or 45), the human p53 sequence (e.g., the sequence shown in SEQ ID NO: 46), and mouse c-abl Sequences of IV (e.g., the sequence shown in SEQ ID NO: 47), influenza virus NS1 (e.g., the sequence shown in SEQ ID NO: 48 or 49), hepatitis virus δ antigen (e.g., the sequence shown in SEQ ID NO: 50), mouse Mx1 protein (e.g., the sequence shown in SEQ ID NO: 51), human poly(ADP-ribose) polymerase (e.g., the sequence shown in SEQ ID NO: 52), and steroid hormone receptor (human) glucocorticoid (e.g., the sequence shown in SEQ ID NO: 53).
[0121] In some embodiments, the RZFD core domain may be fused with one or more nuclear localization sequences (NLSs), for example, about one, two, three, four, five, six, seven, eight, nine, ten, or more NLSs. In some embodiments, the RZFD core domain contains about one, two, three, four, five, six, seven, eight, nine, ten, or more NLSs at or near its N-terminus and / or C-terminus.
[0122] In one embodiment, if the NLS is located within approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain from the N-terminus or C-terminus of the RZFD core domain, then the NLS is considered to be close to the N-terminus or C-terminus of the RZFD core domain.
[0123] The term "gene activation domain" refers to a domain capable of mediating transcription and activation, which can facilitate the binding of RNA polymerase to a promoter, thereby regulating the expression of a gene (e.g., the RZFD core domain in this specification). Non-exclusive examples of gene activation domains include gene activation domain sequences derived from: VP64, P65, HSF1, VP16, MyoD1, RTA, SET7 / 9, Suntag, P300, CBP, or histone acetyltransferase activation domains, or combinations thereof, e.g., combinations of VP64 and P65, VP64, P65 and HSF1, MyoD1 and RTA, SET7 / 9, Suntag and P300, etc.
[0124] In some embodiments, the RZFD core domain may be fused with one or more gene activation domains, for example, about one, two, three, four, five, six, seven, eight, nine, ten, or more gene activation domains. In some embodiments, the RZFD core domain contains about one, two, three, four, five, six, seven, eight, nine, ten, or more gene activation domains at or near the N-terminus and / or C-terminus.
[0125] In one embodiment, if the gene activation domain is located within approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain from the N-terminus or C-terminus of the RZFD core domain, the gene activation domain is considered to be close to the N-terminus or C-terminus of the RZFD core domain.
[0126] In some embodiments, the NLS may be fused with one or more gene activation domains, for example, about one, two, three, four, five, six, seven, eight, nine, ten, or more gene activation domains. In some embodiments, the NLS contains about one, two, three, four, five, six, seven, eight, nine, ten, or more gene activation domains at or near its N-terminus and / or C-terminus.
[0127] In one embodiment, the gene activation domain is considered to be close to the N-terminus or C-terminus of the NLS if it is located within approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain from the N-terminus or C-terminus of the NLS.
[0128] The term "artificial transcription factor" refers to a protein that can bind to RE1 and can promote neuron production or regulate the expression of neuron-related genes. In some embodiments, the artificial transcription factor can bind to RE1 / NRSE competitively with endogenous REST. The artificial transcription factor contains an RZFD core domain that can bind to RE1.
[0129] In this specification, the term “RZFD core domain” is also referred to as “core domain,” and includes any 5 to 8 sequences from SEQ ID NOs: 3 to 10, but does not include the sequences shown in SEQ ID NOs: 12 and 13. In some embodiments, the amino acid sequence of the RZFD core domain includes any 6 to 8 sequences from SEQ ID NOs: 3 to 10. In some examples, the amino acid sequence of the RZFD core domain may include the sequence shown in any of SEQ ID NOs: 2, 14, 74, 76, and 75. In some examples, the amino acid sequence of the RZFD core domain may include a sequence that has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% identity with the sequence shown in any of SEQ ID NOs: 2, 14, 74, 76, and 75.
[0130] In this specification, "homology," "identity," or "similarity" of an amino acid sequence or nucleotide sequence have the same meaning and refer to the proportion of identical bases or amino acids in the entire sequence between one sequence and another sequence being compared.
[0131] The term "enhancer" refers to a short region on DNA that can bind to a protein, and after binding to the protein, can enhance the transcription of the target gene. Enhancers may be located upstream or downstream of the target gene. The sequence of an enhancer may or may not be adjacent to the target gene.
[0132] The term "WPRE," which refers to the post-transcriptional regulatory element of woodchuck hepatitis virus, can increase the expression efficiency of exogenous gene fragments. For example, by incorporating WPRE into the 3'UTR region of rAAV (e.g., the coding nucleic acid of an artificial transcription factor) containing an exogenous gene, mRNA expression levels and translation efficiency can be increased, and the expression of recombination can be further enhanced.
[0133] The term "regulatory element" refers to a nucleic acid sequence that can regulate the transcription of a target gene, enhance the stability of the transcript, or improve the protein expression level of the target gene. Exemplary regulatory elements may include promoters, enhancers, WPREs, etc.
[0134] Artificial transcription factors and RZFD core domains RE1 binds to REST (RE1 silencing transcription factor) and can regulate the expression of genes related to neuronal development and maturation. REST is an endogenous protein that binds to RE1. The sequence of the human wild-type REST protein is shown in SEQ ID NO: 1.
[0135] Domain prediction and protein structure modeling revealed that positions 159–412 of the amino acid sequence of human wild-type REST protein contain eight zinc finger domains (RZFDs). Positions 1–83 of the novel N-terminal amino acid sequence of REST contain an N-terminal inhibitory region that can bind to proteins, such as Sin3a and Sin3b. Positions 1008–1097 of the C-terminal amino acid sequence of REST contain an inhibitory domain and one zinc finger domain that can bind to proteins, such as RCOR1.
[0136] The artificial transcription factor of this application comprises an RZFD core domain, which can bind to RE1 and promote neuronal production or regulate the expression of neuron-related genes. In some embodiments, the artificial transcription factor can bind to RE1 competitively with endogenous REST.
[0137] In some embodiments, the artificial transcription factor may include the sequences shown in SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92. In some embodiments, the artificial transcription factor may also be a mutant sequence of the sequences shown in SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92, and possible mutations in the mutant include, but are not limited to, insertions, deletions, point mutations, InDel mutations, substitutions, or any combination thereof. In some embodiments, the artificial transcription factor may be a sequence that has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the sequence shown in SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92.
[0138] In some examples, the nucleic acid sequence of the mutant may contain at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, or more mutations.
[0139] The nucleic acid sequence of the artificial transcription factor may include a codon-optimized form of REST.
[0140] In some specific embodiments, the coding sequence of the codon-optimized artificial transcription factor may include the nucleotide sequence of SEQ ID NO: 64. Selectively, the coding sequence of the codon-optimized artificial transcription factor may include a nucleotide sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 94%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with SEQ ID NO: 64.
[0141] Compared to the coding sequences of non-codon-optimized artificial transcription factors, codon optimization can improve expression levels. The coding sequences of codon-optimized artificial transcription factors can improve expression by at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 100%, or more, compared to the coding sequences of non-codon-optimized artificial transcription factors.
[0142] Wild-type REST contains an N-terminal repressor domain, an intermediate DNA-binding domain that binds to DNA, and a C-terminal transcriptional repressor domain.
[0143] In some embodiments, the artificial transcription factor includes the RZFD core domain but does not include the sequences shown in SEQ ID NOs: 12 and 13. For example, the artificial transcription factor may include positions 159-412 of the REST amino acid sequence but does not include the N-terminal and / or C-terminal repression domains of REST. SEQ ID NOs: 12 shows the N-terminal amino acid sequence of REST, and SEQ ID NOs: 13 shows the C-terminal amino acid sequence of REST.
[0144] The core domain may contain any 5 to 8 sequences from SEQ ID NOs. 3 to 10. In some embodiments, the amino acid sequence of the RZFD core domain contains any 6 to 8 sequences from SEQ ID NOs. 3 to 10. In some examples, the amino acid sequence of the RZFD core domain may include the sequence shown in any of SEQ ID NOs. 2, 14, 74, 76, or 75. In some specific embodiments, the core domain may contain an amino acid sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with the sequence shown in any of SEQ ID NOs. 3, 14, 74, 76, or 75.
[0145] In some embodiments, the artificial transcription factor may include an RZFD core domain, an NLS, and / or an activation domain, or the artificial transcription factor may include an RZFD core domain, an NLS, and / or an epigenetic regulatory domain, or the artificial transcription factor may include an RZFD core domain, an NLS, an activation domain, and an epigenetic regulatory domain. In some specific embodiments, the artificial transcription factor may employ the design scheme shown in Figure 1, where RZFD represents the RZFD core domain described in this application, NLS represents an optional nuclear localization signal, TAD represents an optional activation domain, and ERD represents an optional epigenetic regulatory domain.
[0146] In some embodiments, the RZFD core domains, NLS, TAD, or ERD may be linked together using appropriate linkers.
[0147] The artificial transcription factor may contain the amino acid sequence of any one of SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92. The RZFD variant may contain an amino acid sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92. Artificial transcription factors may be encoded by nucleic acid sequences containing any one of sequence numbers 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93. Artificial transcription factors may also be encoded by nucleic acid sequences containing sequences having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of sequence numbers 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93.
[0148] Nuclear localization signals The artificial transcription factor may further contain a nuclear localization signal (NLS). A nuclear localization signal is an amino acid sequence that can guide a protein into the cell nucleus via nuclear transport. The NLS may be located at the N-terminus of a protein (e.g., the RZFD core domain). Alternatively, the NLS may be located at the C-terminus of a protein (e.g., the RZFD core domain). The NLS may be located at the N-terminus of the RZFD core domain. Alternatively, the NLS may be located at the C-terminus of the RZFD core domain. In some embodiments, the NLS is located at both the N-terminus and the C-terminus of the RZFD core domain.
[0149] The artificial transcription factor may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS sequences. The NLS may contain the amino acid sequence of SEQ ID NO: 24 or any one of 38-53. Selectively, the NLS may contain an amino acid sequence that is at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 94%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to one of SEQ ID NO: 24 or any one of 38-53. The NLS may be encoded by a nucleic acid sequence containing the sequence of SEQ ID NO: 25. Alternatively, the NLS may be encoded by a nucleic acid sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 94%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with the sequence of sequence number 25.
[0150] An artificial transcription factor having an NLS may contain the amino acid sequence shown in any of SEQ ID NOs: 30, 32, 34, 57, 59, 67, or 61. Alternatively, an artificial transcription factor having an NLS may contain an amino acid sequence that has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 94%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the sequences from SEQ ID NOs: 30, 32, 34, 57, 59, 67, or 61. The coding nucleic acid sequence of an artificial transcription factor having an NLS may contain the nucleic acid sequence shown in any one of SEQ ID NOs: 31, 33, 35, 56, 58, 60, 68, or 64. Alternatively, the coding nucleic acid sequence of an artificial transcription factor having an NLS may include a nucleic acid sequence that has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 94%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of sequence numbers 31, 33, 35, 56, 58, 60, 68, or 64.
[0151] The artificial transcription factor may further contain a nuclear export signal (NES). The nuclear export signal is an amino acid sequence that labels the export of a protein from the cell nucleus by nuclear transport. The NES may be fused to the N-terminus of a protein (e.g., the core domain). Alternatively, the NES may be fused to the C-terminus of a protein (core domain). The protein (e.g., the artificial transcription factor) may contain NES sequences 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. The NES may contain one amino acid sequence from sequence number 54 or 55. Alternatively, the NES may contain an amino acid sequence that has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 94%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with one of sequence number 54 or 55.
[0152] Gene activation domain In some embodiments, the artificial transcription factor further comprises a gene activation domain (also abbreviated as activation domain in this application). The activation domain may comprise an epigenetic modification protein or a gene activation regulator, and selectively comprises the activation domains of VP64, P65, HSF1, VP16, MyoD1, RTA, SET7 / 9, Suntag, P300, CBP, or histone acetyltransferase, or any combination thereof. In some embodiments, the gene activation domain comprises the activation domain of VP64, the activation domain of P65, the activation domain of RTA, or the activation domain of HSF1, or any combination thereof.
[0153] In some embodiments, the activation domain may include multiple activation domains, for example, the VP64 activation domain and the P65 activation domain. In some embodiments, the activation domain further includes the VP64 activation domain, the P65 activation domain and the HSF1 activation domain. In some embodiments, the activation domain further includes the VP64 activation domain, the P65 activation domain and the RTA activation domain.
[0154] The activation domain may be located at the N-terminus of the RZFD core domain. Alternatively, the activation domain may be located at the C-terminus of the RZFD core domain. Alternatively, the activation domain may be located at both the C-terminus and N-terminus of the RZFD core domain simultaneously.
[0155] In some embodiments, the NLS may be fused with one or more gene activation domains, for example, about one, two, three, four, five, six, seven, eight, nine, ten, or more gene activation domains. In some embodiments, the N-terminus and / or C-terminus of the NLS, or its vicinity, contains about one, two, three, four, five, six, seven, eight, nine, ten, or more gene activation domains.
[0156] In one embodiment, the gene activation domain is considered to be close to the N-terminus or C-terminus of the NLS if it is located within approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain from the N-terminus or C-terminus of the NLS.
[0157] The activating domain of VP64 may contain the amino acid sequence of SEQ ID NO: 26. Alternatively, the activating domain of VP64 may contain an amino acid sequence having at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 92%, at least approximately 94%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity with the sequence of SEQ ID NO: 26. The activating domain of VP64 may be encoded by a nucleic acid sequence containing the sequence of SEQ ID NO: 27. The activating domain of VP64 may be encoded by a nucleic acid sequence containing a sequence having at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 92%, at least approximately 94%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity with the sequence of SEQ ID NO: 27.
[0158] The activation domain of P65 contains an amino acid sequence that has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 69.
[0159] The activation domain of RTA has an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 11.
[0160] The activation domain of HSF1 has an amino acid sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 70.
[0161] Some fusion forms of the activation domains may include a fusion form of the activation domains of P65 and HSF1 (hereinafter abbreviated as P65-HSF1 activation domain or P65-HSF1), where P65-HSF1 may include the amino acid sequence of SEQ ID NO: 28. Alternatively, the P65-HSF1 activation domain may include an amino acid sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 94%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with the sequence of SEQ ID NO: 28. The P65-HSF1 activation domain may be encoded by a nucleic acid sequence including the sequence of SEQ ID NO: 29. The P65-HSF1 activation domain may be encoded by a nucleic acid sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 94%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with the sequence of Sequence ID No. 29.
[0162] Some fusion forms of the activation domain may include fusion forms of the VP64 activation domain, the P65 activation domain, and the RTA activation domain. In specific examples, the activation domain of such fusion form may include an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 72.
[0163] Some fusion forms of the activation domain may include fusion forms of the VP64 activation domain, the P65 activation domain, and the HSF1 activation domain. In specific examples, the activation domain of such fusion form may include an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 73.
[0164] Some fusion forms of the activation domain may include a fusion form of the P65 activation domain and the RTA activation domain. In specific examples, the activation domain of such fusion form may include an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 71.
[0165] Artificial transcription factor design direction In some embodiments, the artificial transcription factor comprises an RZFD core domain, an NLS, and a gene activation domain.
[0166] The method of linking the RZFD core domain, NLS, and gene activation domain from the N-terminus to the C-terminus may be chosen at will.
[0167] In some specific embodiments, the artificial transcription factor may include an NLS located at the N-terminus of the RZFD core domain and a gene activation domain. Alternatively, the artificial transcription factor may include an NLS located at the C-terminus of the RZFD core domain and a gene activation domain. Alternatively, the artificial transcription factor may include an NLS located at the N-terminus of the RZFD core domain and a gene activation domain located at the C-terminus of the RZFD core domain. Alternatively, the artificial transcription factor may include an NLS located at the C-terminus of the RZFD core domain and a gene activation domain located at the N-terminus of the RZFD core domain.
[0168] In some specific embodiments, the NLS may be located at the N-terminus of the gene activation domain, or at the C-terminus of the gene activation domain. The NLS may be adjacent to the gene activation domain, or it may not be adjacent to the gene activation domain. The artificial transcription factor may contain one, two, three, four, five or more NLS domains and one, two, three, four, five or more gene activation domains. In some embodiments, the artificial transcription factor may contain at least two NLS domains and at least one gene activation domain. In some embodiments, the artificial transcription factor may contain at least two NLS domains and at least two gene activation domains.
[0169] Linker In some embodiments of the present invention, if any two nuclear localization signals, core domains, gene activation domains, enhancers, or WPREs are fused to each other, they may be linked together using a linker, also known as a connexon.
[0170] The term "linker" refers to a short peptide or its coding sequence that links different polypeptide sequences or amino acid sequences. In some specific embodiments, the short peptide may consist of 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids, or it may contain more amino acids, for example, 11 to 15. Typically, such short sequences have no special physiological activity other than linking protein or nucleic acid fragments or maintaining some minimum distance or other spatial relationship between protein or nucleic acid fragments. However, in some embodiments, the linker may be selected to influence specific properties of the linker and / or fusion molecule, such as linker folding, net charge, or hydrophobicity.
[0171] Linkers suitable for the method of the present invention are well known to those skilled in the art and include, but are not limited to, linear or branched carbon linkers, heterocyclic carbon linkers, or peptide linkers. The linker may further be covalent (carbon-carbon bond or carbon-heteroatom bond).
[0172] Expression vector The nucleic acid encoding the artificial transcription factor may be located in the expression vector. The expression vector may contain a nucleic acid sequence encoding the artificial transcription factor, of which the artificial transcription factor includes a core domain.
[0173] Expression vectors, also known as expression constructs, may be bacterial plasmids or viral vectors designed for gene expression in cells. In some embodiments, viral vectors may include recombinant adeno-associated virus vectors (rAAV), adeno-associated virus (AAV) vectors, adenovirus vectors, lentiviral vectors, retrovirus vectors, poxvirus vectors, herpesviruses, SV40 virus vectors, or any combination thereof.
[0174] In several specific examples, the adeno-associated virus (AAV) vector can be selected from AAV1, AAV2, AAV5, AAV6, AAV8, AAV9, AAV10, AAVrh, AAV-DJ, AAV-PHP.B, AAV-PHP.S, and AAV-PHP.eB, or their mutants.
[0175] The expression vector may further contain regulatory elements. These regulatory elements are nucleic acid regions that regulate the transcription (e.g., expression or translation rate) of the relevant gene. The regulatory elements may be protein transcription start sites, promoters, enhancers, inhibitors, or post-transcriptional regulators.
[0176] The expression vector may further include a protein translation initiation site (e.g., a Kozak sequence). The Kozak sequence may include the sequence of SEQ ID NO: 23. The Kozak sequence may include a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with SEQ ID NO: 23. In SEQ ID NO: 23, R is a purine (A or G).
[0177] Adjustment element The regulatory element is a nucleic acid region that regulates the transcription (e.g., expression or translation rate) of the relevant gene (e.g., artificial transcription factor).
[0178] In some embodiments, the regulatory element includes a promoter that modulates the nucleic acid expression of the relevant gene, and preferably, the promoter includes a glial cell-specific promoter.
[0179] In some embodiments, the regulatory element includes an enhancer that enhances gene expression. Preferably, the enhancer includes a CMV enhancer, an SV40 intron-type enhancer, an MVM intron-type enhancer, a β-globulin intron-type enhancer, or a synthetic intron-type enhancer, or any combination thereof.
[0180] In some embodiments, the regulatory element includes a post-transcriptional regulatory element (WPRE) of woodchuck hepatitis virus.
[0181] In some embodiments, the regulatory element includes an enhancer and a post-transcriptional regulatory element (WPRE) for woodchuck hepatitis virus.
[0182] promoter Regulatory elements include promoters that regulate the nucleic acid expression of related genes.
[0183] A promoter is a cis-regulatory element, a nucleic acid region capable of initiating the transcription of a downstream, relevant nucleic acid sequence. Promoters may be tissue-specific (e.g., specific to a particular cell type). Non-limiting examples of promoters include stem cell promoters, osteocyte promoters, hematopoietic cell promoters, and neuron promoters, muscle cell promoters, or spermatophore promoters.
[0184] A promoter is a gene that helps drive the expression of a specific gene. Non-exclusive examples of promoters include photoreceptor promoters, cholinergic neuron promoters, adrenergic neuron promoters, peptide-glucan neuron promoters, glial cell promoters (e.g., astrocyte-specific promoters or Müllerian glial cell promoters), Schwann cell promoters, Purkinje cell promoters, and medium spiny neuron promoters.
[0185] In some specific embodiments, the glial cell-specific promoter may include an astrocyte-specific promoter. Non-limiting examples of the glial cell-specific promoter may include the GFAP promoter, ALDH1L1 promoter, EAAT1 / GLAST promoter, glutamine synthetase promoter, S100β promoter, EAAT2 / GLT-1 promoter, or Rlbp1 promoter. The glial cell-specific promoter may also include a Müller glial (MG) cell-specific promoter. The glial cell-specific promoter may also include a cochlear glial cell-specific promoter. Non-limiting examples of the cochlear glial cell-specific promoter may include the GFAP promoter, ALDH1L1 promoter, EAAT1 / GLAST promoter, or Plp1 promoter.
[0186] In some embodiments, the glial cell-specific promoter may include the GFAP promoter. In some embodiments, the cochlear glial cell-specific promoter may include the GFAP promoter. The GFAP promoter may include the nucleotide sequence of SEQ ID NO: 65 or 66. The GFAP promoter may include a nucleotide sequence having at least about 70%, 80%, 85%, 90%, 95%, 99%, or more sequence identity with SEQ ID NO: 65 or 66.
[0187] Enhancer The expression vector may further contain enhancers. Enhancers are cis-regulatory elements, which are nucleic acid regions that can bind to a protein to increase the transcriptionability of the relevant nucleic acid sequence. Non-limiting examples of enhancers include CMV enhancers, SV40 intron enhancers, MVM intron enhancers, β-globulin intron enhancers, or synthetic intron enhancers, or any combination thereof.
[0188] The present invention may include one enhancer, or multiple enhancers, for example, 2, 3, 4, 5, 6, 7, 8, 9, or 10 enhancers.
[0189] The enhancer may be a CMV enhancer. For example, the CMV enhancer may contain the sequence of sequence number 16. The CMV enhancer may contain a sequence that has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with sequence number 16.
[0190] The enhancer may be an SV40 intron enhancer. For example, the SV40 intron enhancer may contain the sequence of sequence number 17. The SV40 intron enhancer may contain a sequence that has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with sequence number 17.
[0191] The enhancer may be an MVM intron enhancer. For example, the MVM intron enhancer may contain the sequence of sequence number 18. The MVM intron enhancer may contain a sequence that has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with sequence number 18.
[0192] The enhancer may be a β-globin intron enhancer. For example, the β-globin intron enhancer may contain the sequence of SEQ ID NO: 19. The β-globin intron enhancer may contain a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with SEQ ID NO: 19.
[0193] The enhancer may be a synthetic intron enhancer. For example, the synthetic intron enhancer may contain the sequence of any one of sequence numbers 20 to 22. The synthetic intron enhancer may contain a sequence that has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of sequence numbers 20 to 22.
[0194] Differentiation This specification describes a method for differentiating non-neuronal cells into functional neurons by regulating the binding of endogenous REST to RE1 / NRSE by administering an expression vector or composition described herein.
[0195] In some embodiments, non-neuronal cells may be terminally differentiated cells, progenitor cells, stem cells, etc.
[0196] In some embodiments, the RE1 sequence can be targeted using the expression vector or composition described herein, thereby blocking the binding of endogenous REST to RE1 by competitively binding to endogenous REST.
[0197] In some embodiments, the expression vector or composition may be an adeno-associated virus (AAV) vector or lentiviral vector having an artificial transcription factor. Selective AAV vectors include, but are not limited to, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10 or their mutants.
[0198] In non-neuronal cells (e.g., glial cells), expression of the RZFD core domain can reduce the inhibition of neuron-related gene expression by the REST silencing complex. In some embodiments, activating domains such as VP64, P65, HSF1, and Rta are fused to the RZFD core domain to further promote the expression of neuron-related genes and accelerate the differentiation of glial cells into neurons. In some embodiments, NLS such as bpNLS is fused to the RZFD core domain to further promote the expression of neuron-related genes and accelerate the differentiation of glial cells into neurons. In some embodiments, fusion of NLS, activating domains, and the RZFD core domain further promotes the expression of neuron-related genes and accelerates the differentiation of glial cells into neurons.
[0199] In some embodiments, the NLS, activation domain, and RZFD core domain may employ the linking method shown in Figure 1, and linkers may be used to link the RZFD core domain, NLS, TAD, or ERD.
[0200] In some embodiments, the expression vectors or compositions described herein can convert glial cells, particularly astrocytes, into dopaminergic neurons. In some specific embodiments, the design of the artificial transcription factors of the expression vectors or compositions described herein, as shown in Figure 2B, allows the expression vectors or compositions containing RZFD-V1, RZFD-V2, RZFD-V3, RZFD-V4, RZFD-V5, RZFD-V6, RZFD-V7, or RZFD-V8 as artificial transcription factors to convert glial cells into dopaminergic neurons. In specific embodiments, the expression vector or composition may be an AAV vector. In some specific embodiments, dopaminergic neurons can be detected by immunofluorescence staining. In some specific embodiments, astrocytes can be converted into dopaminergic neurons by injecting an AAV vector containing artificial transcription factors into the brain (e.g., the striatum or substantia nigra).
[0201] In some embodiments, the expression vectors or compositions described herein can convert Müller cells into retinal ganglion cells or photoreceptor cells. For example, Müller cells can be converted into retinal ganglion cells or photoreceptor cells by injecting an AAV vector AAV containing an artificial transcription factor subretinally or into the vitreous cavity. Retinal ganglion cells are the only cells in the visual pathway that transmit visual signals to the brain, and their deletion or death can cause permanent blindness.
[0202] Non-neuronal cells may include glial cells, fibroblasts, stem cells, neural progenitor cells, or neural stem cells. In some embodiments, non-neuronal cells are glial cells. In some embodiments, glial cells are selected from astrocytes, oligodendrocytes, ependymal cells, Schwann cells, NG2 cells, satellite cells, Müllerian glial cells, ostratus glial cells, and any combination thereof.
[0203] In some embodiments, glial cells are located in the brain, spinal cord, eyes, or ears. In some embodiments, glial cells are located in the striatum, substantia nigra, ventral tegmental area of the midbrain, medulla oblongata, hypothalamus, dorsal midbrain, or cerebral cortex of the brain.
[0204] In some embodiments, glial cells include astrocytes, and functional neurons include dopaminergic neurons. In some embodiments, the methods provided herein relate to methods for transdifferentiating astrocytes into dopaminergic neurons in an individual. In some embodiments, astrocytes are located in the striatum and / or substantia nigra. In some embodiments, the methods include administering the active material provided herein to the striatum and / or substantia nigra of an individual.
[0205] In some embodiments, glial cells include Müller cells, and functional neuronal cells include retinal ganglion cells (RGCs) and / or photoreceptor cells. In some embodiments, the methods provided herein relate to methods for transdifferentiating Müller cells into retinal ganglion cells (RGCs) and / or photoreceptor cells in an individual. In some embodiments, Müller cells are located subretinally or in the vitreous cavity. In some embodiments, the methods include administering the active material provided herein subretinally or in the vitreous cavity of an individual.
[0206] In some embodiments, glial cells include cochlear glial cells, and functional neuronal cells include cochlear spiral ganglion cells. In some embodiments, the methods provided herein relate to methods for transdifferentiating cochlear glial cells into cochlear spiral ganglion cells in an individual. In some embodiments, cochlear glial cells are located in the inner ear. In some embodiments, the methods include administering the activators provided herein into the inner ear of an individual.
[0207] Functional neurons may refer to neuronal cells having a specific function, such as dopaminergic neurons, retinal ganglion cells, photoreceptor cells, and other neurons having a specific function. In some embodiments, functional neurons have at least one morphological feature of a neuron (e.g., synapses, axons). In some embodiments, functional neurons include axons. In some embodiments, functional neurons express at least one marker of a mature neuron, such as a NeuN gene expression product. In some embodiments, functional neurons have electrophysiological properties.
[0208] In some embodiments, functional neurons express the NeuN gene. The NeuN gene is a known specific marker for mature neurons. Detection of NeuN gene expression products (e.g., NeuN protein) in non-neuronal cells indicates that the non-neuronal cells have been converted into functional neurons.
[0209] Functional neurons may have different functions. In some embodiments, functional neurons include dopaminergic neurons, ganglion cells, retinal ganglion cells, photoreceptor cells and cochlear spiral ganglion cells, GABAergic neurons, 5-HT neurons, glutamatergic neurons, ChAT neurons, NE neurons, motor neurons, spinal neurons, spinal motor neurons, pyramidal neurons, interneurons, medium spiny neurons (MSNs), Purkinje cells, granule cells, olfactory neurons, periglomerular cells, or any combination thereof.
[0210] In some embodiments, functional neurons include dopaminergic neurons. Dopaminergic neurons are neurons that contain and release dopamine (DA) as a neurotransmitter. Dopaminergic neurons are the primary source of dopamine in the central nervous system. Dopamine is a catecholamine neurotransmitter that can influence neuronal functions such as mood and reward and exerts important biological effects in the central nervous system. In the brain, dopaminergic neurons are mainly concentrated in the substantia nigra pars compacta (SNc), ventral tegmental area (VTA), hypothalamus, and periventricular region of the midbrain. Dopaminergic neurons may be associated with a variety of diseases in the human body, most typically Parkinson's disease. The progressive loss of dopaminergic neurons causes many of the motor symptoms associated with Parkinson's disease.
[0211] In some embodiments, dopaminergic neurons may express one or more markers selected from NeuN, tyrosine hydroxylase (TH), FoxA2, Nurrl, Pitx3, Vmat2, and DAT. The markers may refer to gene expression products such as mRNA or proteins. Detection of the expression of one or more markers in functional neuronal cells indicates that the functional neuron is a dopaminergic neuron. In some embodiments, dopaminergic neurons express TH. TH is an enzyme responsible for catalyzing the conversion of the amino acid L-tyrosine to dihydroxyphenylalanine (DOPA), and is also an enzyme involved in the synthesis and metabolism of dopamine in dopaminergic neurons. In some embodiments, dopaminergic neurons express NeuN, TH, and DAT.
[0212] In some embodiments, functional neurons include retinal ganglion cells. Retinal ganglion cells are neurons located near the inner surface of the retina (ganglion cell layer) and receive visual information from photoreceptors via two types of interneurons (bipolar cells and amacrine cells). Their dendrites primarily establish synaptic connections with bipolar cells, and their axons extend to the optic head, forming the optic nerve, which extends to the brain.
[0213] In some embodiments, retinal ganglion cells (RGCs) express one or more markers selected from RBPMS, Pax6, Brn3a, Brn3b, Brn3c, and Map2. RBPMS is a specific marker for RGCs. When RBPMS expression is detected in functional neurons, it indicates that the functional neurons are RGCs. In some embodiments, retinal ganglion cells (RGCs) express NeuN and RBPMS.
[0214] In some embodiments, functional neurons include photoreceptor cells. Photoreceptor cells are specialized neuroepithelial cells found in the retina that have the function of sensing light and transmitting light. They are processed by bipolar cells and ganglion cells, which convert light signals into electrical signals that can be transmitted to the brain. Photoreceptor cells include rod cells and cone cells.
[0215] In some embodiments, photoreceptor cells express one or more markers selected from rhodopsin, mCAR, m-opsin, and S-opsin. Rhodopsin, mCAR, m-opsin, and S-opsin are all specific markers for photoreceptor cells. Detection of rhodopsin, mCAR, m-opsin, and / or S-opsin expression in functional neuron cells indicates that the functional neuron is a photoreceptor cell. In some embodiments, photoreceptor cells express NeuN, Rhodopsin, and / or mCAR.
[0216] In some embodiments, functional neurons include cochlear spiral ganglion cells. Cochlear spiral ganglion cells are bipolar ganglion cells and are primary neurons in the auditory transmission pathway. Their peripheral processes connect with hair cells, and their intermediate processes are involved in the formation of the auditory nerve. Spiral ganglion cells play a crucial role in the transmission and coding of acoustic signals.
[0217] In some embodiments, cochlear spiral ganglion cells express one or more markers selected from NeuN, Prox1, Tuj-1, and Map2. Detection of Prox1 and Map2 expression in functional neurons indicates that the functional neurons are cochlear spiral ganglion cells. In some embodiments, cochlear spiral ganglion cells express NeuN, Prox1, Tuj-1, and Map2.
[0218] In some embodiments, transdifferentiation is achieved by administering the fusion protein and / or expression vector described herein to glial cells. The administration of glial cells may occur locally at one or more of the following locations in the individual: 1) glial cells in the striatum, 2) glial cells in the substantia nigra of the brain, 3) glial cells in the retina, 4) glial cells in the inner ear, 5) glial cells in the spinal cord, 6) glial cells in the prefrontal cortex, 7) glial cells in the motor cortex, 8) glial cells in the hypothalamus, and 9) glial cells in the ventral tegmental area (VTA).
[0219] In some embodiments, the efficiency of differentiation conversion from non-neuronal cells to functional neurons is at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 100%, or more. In some embodiments, the efficiency of differentiation conversion from glial cells to functional neurons is at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 100%, or more.
[0220] The differentiation efficiency may be detected and calculated by methods known to those skilled in the art. For example, fluorescent proteins such as GFAP-mCherry, GFAP-tdTomato, and GFAP-EGFP may be used to label the initially differentiated cells (e.g., glial cells). For example, an Ai9 transgenic mouse with fluorescently labeled glial cells may be used. Since the differentiated cells also have fluorescence, the differentiation efficiency may be calculated by calculating the percentage of the number of differentiated cells to the number of initially labeled cells. Alternatively, the differentiation efficiency may be calculated as the percentage of the number of cells produced by differentiation to the number of this type of cell at the administration site, for example, in the substantia nigra, as the percentage of newly produced dopaminergic neurons to the dopaminergic neurons in the substantia nigra.
[0221] active material In this specification, the active material refers to a substance comprising an expression vector or composition provided herein. In some specific embodiments, the active material may further comprise a pharmaceutically acceptable carrier.
[0222] In some specific embodiments, the expression vector described herein comprises (a) a nucleic acid sequence encoding an artificial transcription factor capable of binding to RE1, including an RZFD core domain, and (b) a regulatory element that modulates the expression of the artificial transcription factor nucleic acid, wherein the RZFD core domain includes any 5 to 8 sequences from SEQ ID NOs. 3 to 10, but does not include the sequences shown in SEQ ID NOs. 12 and 13. In some embodiments, the amino acid sequence of the RZFD core domain includes any 6 to 8 sequences from SEQ ID NOs. 3 to 10. In some examples, the amino acid sequence of the RZFD core domain may include any of the sequences shown in SEQ ID NOs. 2, 14, 74, 76, and 75. In some specific embodiments, the RZFD core domain includes sequences that have at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with sequence numbers 2, 14, 74, 76, and 75. In some specific embodiments, the RZFD core domain includes sequences that have at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with sequence number 14. In some specific embodiments, the coding nucleic acid sequence of the RZFD core domain includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with sequence number 15.
[0223] In some specific embodiments, the artificial transcription factor of the expression vector comprises an RZFD core domain, an NLS, and / or a gene activation domain. In some specific embodiments, the NLS may comprise one or more domains. In some specific embodiments, the gene activation domain may comprise one or more domains.
[0224] In some specific embodiments, the regulatory element includes a promoter that modulates the expression of artificial transcription factor nucleic acid. In some specific embodiments, the regulatory element of the expression vector includes a promoter that modulates the expression of artificial transcription factor nucleic acid and an enhancer that increases the transcription of the artificial transcription factor, or the regulatory element of the expression vector includes a promoter that modulates the expression of artificial transcription factor nucleic acid and a post-transcriptional enhancing element (WPRE) of woodchuck hepatitis virus, or the regulatory element of the expression vector includes a promoter that modulates the expression of artificial transcription factor nucleic acid, an enhancer that increases the transcription of the artificial transcription factor, and a post-transcriptional enhancing element (WPRE) of woodchuck hepatitis virus.
[0225] In some specific embodiments, the composition described herein comprises (a) an artificial transcription factor capable of binding to RE1, comprising an RZFD core domain, a first NLS, and a first gene activation domain, or (b) a nucleic acid encoding the artificial transcription factor described in (a), wherein the RZFD core domain comprises any 5 to 8 sequences from SEQ ID NOs. 3 to 10, but does not include the sequences shown in SEQ ID NOs. 12 and 13. In some embodiments, the amino acid sequence of the RZFD core domain comprises any 6 to 8 sequences from SEQ ID NOs. 3 to 10. In some examples, the amino acid sequence of the RZFD core domain may include any of the sequences shown in SEQ ID NOs. 2, 14, 74, 76, and 75. In some specific embodiments, the RZFD core domain includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 2, 14, 74, 76, and 75. In some specific embodiments, the coding nucleic acid sequence of the RZFD core domain includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 15.
[0226] In some specific embodiments, the artificial transcription factor of the composition further comprises a second NLS, in some specific embodiments, the artificial transcription factor of the composition further comprises a second gene activation domain, and in some specific embodiments, the artificial transcription factor of the composition further comprises a second NLS and a second gene activation domain.
[0227] In some specific embodiments, the composition further comprises a regulatory element that modulates the expression of the artificial transcription factor nucleic acid. In some specific embodiments, the regulatory element comprises a promoter that modulates the expression of the artificial transcription factor nucleic acid. In some specific embodiments, the regulatory element comprises a promoter that modulates the expression of the artificial transcription factor nucleic acid and an enhancer that increases the transcription of the artificial transcription factor, or the regulatory element of the expression vector comprises a promoter that modulates the expression of the artificial transcription factor nucleic acid and a post-transcriptional enhancing element (WPRE) of woodchuck hepatitis virus, or the regulatory element of the expression vector comprises a promoter that modulates the expression of the artificial transcription factor nucleic acid, an enhancer that increases the transcription of the artificial transcription factor, and a post-transcriptional enhancing element (WPRE) of woodchuck hepatitis virus.
[0228] Treatment of disease In some embodiments, administration of the expression vector or composition described herein may be used to prevent and / or treat diseases associated with loss of neuronal function or death of neurons in individuals requiring such treatment. In some embodiments, neurons obtained by the transdifferentiation described herein may be used to prevent and / or treat diseases associated with loss of neuronal function or death of neurons in individuals requiring such treatment.
[0229] Diseases or conditions associated with loss of neuronal function or death of neurons may include diseases associated with loss of function or death of dopaminergic neurons, or diseases associated with visual impairment due to loss or death of optic ganglia or photoreceptor cells. In some embodiments, diseases associated with loss of neuronal function or death of neurons are selected from Parkinson's disease, Alzheimer's disease, stroke, schizophrenia, Huntington's disease, depression, motor neuropathy, amyotrophic lateral sclerosis, spinal muscular atrophy, Pick's disease, sleep disorders, epilepsy, ataxia, glaucoma, age-related RGC lesions, optic nerve injury, retinal ischemia or hemorrhage, Leber hereditary optic nerve lesions, degeneration or death of photoreceptor cells due to injury or degeneration, macular degeneration, retinitis pigmentosa, diabetes-related blindness, night blindness, color blindness, hereditary blindness, congenital amaurosis, spiral nerve hearing loss or hearing loss due to ganglion cell death.
[0230] In some embodiments, the present application provides a method for preventing and / or treating diseases associated with loss of neuronal function or death in an individual in need thereof, comprising administering a therapeutically effective amount of the expression vector or composition described herein into the subretinal or vitreous cavity of the individual. The expression vector provided herein can convert Müller cells in the subretinal or vitreous cavity into retinal ganglion cells (RGCs) and / or photoreceptor cells, of which diseases associated with loss of neuronal function or death are selected from RGC cell death due to optic nerve damage, glaucoma, age-related RGC lesions, optic nerve injury, retinal ischemia or hemorrhage, Leber hereditary optic nerve lesions, degeneration or death of photoreceptor cells due to injury or degeneration, macular degeneration, retinitis pigmentosa, diabetes-related blindness, night blindness, color blindness, hereditary blindness and amaurosis.
[0231] The compositions described herein (also referred to herein as pharmaceutical compositions) may include the expression vectors described herein. The pharmaceutical compositions may further include carriers, excipients, adhesives, fillers, suspensions, flavorings, sweeteners, disintegrants, dispersants, surfactants, lubricants, colorants, diluents, solubilizers, plasticizers, stabilizers, penetration enhancers, wetting agents, defoamers, antioxidants, preservatives, or combinations thereof. The pharmaceutical compositions may further include vectors for delivering polynucleotides, of which the vectors include viral vectors, liposomes, nanoparticles, exosomes, or viral particles.
[0232] The pharmaceutical composition may be administered topically, parenterally, intravenously, intradermally, intramuscularly, or intraperitoneally. In some embodiments, the pharmaceutical composition is suitable for topical administration to glial cells in one or more of the following sites: i) glial cells in the striatum, ii) glial cells in the substantia nigra, iii) glial cells in the retina, iv) glial cells in the inner ear, v) glial cells in the spinal cord, vi) glial cells in the prefrontal cortex, vii) glial cells in the motor cortex, viiii) glial cells in the hypothalamus, and ix) glial cells in the ventral tegmental area (VTA). In some embodiments, the pharmaceutical composition is suitable for intracranial or intraocular administration.
[0233] In some embodiments, the pharmaceutical composition further comprises i) one or more dopaminergic neuron-related factors, or ii) one or more retinal ganglion cell-related factors. The one or more dopaminergic neuron-related factors may be selected from FoxA2, Lmx1a, Lmx1b, Nurr1, Pbx1a, Pitx3, Gata2, Gata3, FGF8, BMP, En1, En2, PET1, Pax family proteins (Pax3, Pax6, etc.), SHH, Wnt family proteins, TGF-β family proteins, or any combination thereof. One or more retinal ganglion cell-related factors may be selected from β-catenin, Oct4, Sox2, Klf4, Crx, CamKII, Brn3a, Brn3b, Brn3C, Math5, Otx2, Ngn2, Ngn1, AscL1, miRNA9, miRNA-124, Nr2e3, Nrl, or any combination thereof.
[0234] In some embodiments, the present application provides a method for preventing and / or treating a disease associated with loss of neuronal function or death in an individual in need thereof, comprising administering a therapeutically effective amount of the active material of the present composition to the inner ear of the individual in order to differentiate cochlear glial cells of the inner ear into cochlear spiral ganglion cells, wherein the disease associated with loss of neuronal function or death thereof is selected from spiral ganglion cell death or hearing loss.
[0235] In some embodiments, the present application provides a method for preventing and / or treating a disease associated with loss of neuronal function or death in an individual in need thereof, comprising administering a therapeutically effective amount of the active substance provided herein to the striatum and / or substantia nigra of the individual in order to convert astrocytes of the striatum and / or substantia nigra into dopaminergic neurons, wherein the disease associated with loss of neuronal function or death thereof is selected from Parkinson's disease, depression and Alzheimer's disease.
[0236] The methods for preventing and / or treating neuronal diseases described herein may include replacing non-neuronal cells, converting non-neuronal cells into functional neurons in vitro, and administering functional neurons to individuals in need. Alternatively, the methods for preventing and / or treating neuronal diseases described herein may include converting non-neuronal cells into functional neurons in vivo. Alternatively, the methods for preventing and / or treating neuronal diseases described herein may include administering a therapeutically effective amount of nucleic acid, fusion protein, and / or expression vector described herein, thereby reducing the binding of REST to the RE1 / NRSE element or reducing the amount or activity of REST, thereby converting non-neuronal cells into functional neurons at disease-affected sites.
[0237] The subject (also called the individual) may be, for example, a non-human primate such as a chimpanzee, as well as a mammal such as other apes and monkey species, a farm animal such as a cow, horse, sheep, goat, or pig, a domestic animal such as a rabbit, dog, or cat, or an experimental animal including a rodent such as a rat, mouse, or guinea pig. The mammal may be a human. The subject may be diagnosed or suspected of being at high risk of developing a particular disease. In certain circumstances, the subject may not necessarily be diagnosed or suspected of being at high risk of developing the disease. In some embodiments, the subject has one or more symptoms indicating a disease or condition related to a neurological disorder or medical condition.
[0238] Several experimental means and methods in this study can be implemented using the usual procedures of those skilled in the art, for example.
[0239] Plasmid construction: The AAV backbone vector was enzymatically cleaved with restriction endonucleases and recovered by agarose gel electrophoresis. PCR was performed using cell cDNA as a template, and the PCR fragments were recovered after agarose gel electrophoresis. The backbone vector and fragments were ligated using the ClonExpress MultiS One Step Cloning Kit (Vazyme, C113-02) from Nuowizan Biotechnology Co., Ltd. After ligation, the cells were transformed into DH5α Escherichia coli and plated. The following day, monoclones were collected and identified, positive clones were sequenced, and expansion culture and plasmid extraction were performed on clones with completely correct sequencing.
[0240] AAV injection into mouse brains: Stereotactic injection was performed using the Reward stereotactic injection system (C57BL / 6 or Dat-Cre:Ai9 mice, age over 2 months). The titer was approximately 5 × 10⁻⁶. 12 The dosage was vg / ml (1-3 μL per injection), and AAV was injected into the striatum (AP +0.8 mm, ML ±1.6 mm, and DV -2.8 mm) or substantia nigra (AP -3.0 mm, ML ±1.25 mm, and DV -4.5 mm).
[0241] Mouse tissue immunofluorescence staining: Samples were collected after re-injection of AAV, sectioned, and immunofluorescence staining was performed. After perfusing mice with physiological saline and 4% PFA, the brain was excised, fixed overnight with 4% paraformaldehyde (PFA), and then dehydrated with 30% sucrose for at least 12 hours until the tissue sank to the bottom of the solution. After embedding by OCT, frozen sections were prepared, with section thicknesses of 30 μm or 40 μm. Before immunofluorescence staining, brain sections were washed three times with 0.1 M phosphate buffer (PBS) for 5-10 minutes each. Primary antibody was incubated overnight at 4°C, followed by three-four washes with PBS for 10-15 minutes each. Next, secondary antibody diluted with antibody diluent was added and incubated at room temperature for 2-3 hours. After incubation, the sections were washed three-four more times with PBS for 10-15 minutes each. Finally, the sections were mounted with a fluorescent fading-preventing mounting medium (Life Technology) and stored.
[0242] 6-OHDA PD mouse model: For example, adult C57BL / 6 mice (7-10 weeks old) were used and 25 mg / kg of desipramine hydrochloride (D3900, Sigma-Aldrich) was intraperitoneally injected 0.5 hours before anesthesia. After anesthesia, 3 μg of 6-OHDA (H116, Sigma-Aldrich) or saline was injected into the right medial forebrain bundle of the mice: anteroposterior (A / P) = -1.2 mm, medial-lateral (M / L) = -1.1 mm, dorsal-ventral (D / V) = -5 mm. One hour postoperatively, 1 mL of 4% glucose-saline solution was subcutaneously injected into the mice.
[0243] Examples The following examples are merely for illustrative purposes of illustrating the technical aspects of the present invention and are not intended to limit the scope of the invention.
[0244] Example 1: Comparison of the in vivo differentiation-transition effects of different versions of artificial transcription factors (ATFs). Due to the significant differences between in vivo and in vitro environments, many in vitro studies are ineffective in vivo. To investigate whether different artificial transcription factors can enable the differentiation of glial cells into neurons in vivo, this study constructed different versions of artificial transcription factors (Figure 1A). Of these, dRT-V1, dRT-V2, and dRT-V3 each have different transcriptional activation domains added to their terminals to enhance their transcriptional activity, and dRT-V3 does not increase VP64. dRT-V1 uses amino acids from positions 85 to 1008 of REST (amino acid sequence as shown in SEQ ID NO: 88, nucleotide sequence as shown in SEQ ID NO: 94), and its C-terminus is fused with the VP64 activation domain. dRT-V2 uses amino acids from positions 73 to 508 of REST (amino acid sequence as shown in SEQ ID NO: 89, nucleotide sequence as shown in SEQ ID NO: 96), and its C-terminus is fused with the VP64 activation domain. In this study, the astrocyte-specific promoter GFAP was used to drive the expression of the target gene (GOI), which is an artificial transcription factor (ATF). The control group in the study used the AAV-GFAP-mCherry fluorescent protein (design-1) (Figure 1B). AAV from each group was injected into the striatum of C57 mice in a 6-OHDA model. Samples were collected and analyzed 1.5 to 2 months after AAV injection. The results showed that the red fluorescently labeled cells in the AAV-GFAP-mCherry control group still retained typical astrocyte morphology, with relatively small cell bodies and a large number of coarse cellular protrusions and branching (Figure 1C). In the groups injected with Design-2 (AAV-GFAP-dRT-V1) and Design-3 (AAV-GFAP-dRT-V2), the red fluorescently labeled cells mostly retained the typical astrocyte morphology, with relatively small cell bodies. Some of these cells exhibited a rounded shape and elongated cell processes. However, staining with the neuron-specific protein marker NeuN revealed that these cells were NeuN-negative, with only about 5% of the red fluorescently labeled cells expressing NeuN. These results indicate that AAV-GFAP-dRT-V1 and AAV-GFAP-dRT-V2 have a very weak ability to convert astrocytes into neurons in vivo.Furthermore, in the groups injected with Design-4 (AAV-GFAP-dRT-V3) and Design-5 (AAV-GFAP-dRT-V4), a large number of positive cells co-labeled with mCherry and NeuN appeared, indicating that Design-4 (AAV-GFAP-dRT-V3) and Design-5 (AAV-GFAP-dRT-V4) can efficiently differentiate astrocytes into neurons.
[0245] To investigate whether dRT-V1, dRT-V2, dRT-V3, and dRT-V4 can differentiate glial cells into dopamine neurons, we further stained sections from each group with dopamine neuron-specific protein markers. It was found that neither AAV-GFAP-dRT-V1 nor AAV-GFAP-dRT-V2, which were AAV-GFAP-mCherry control groups, possessed TH-positive cells. Furthermore, the groups injected with Design-4 (AAV-GFAP-dRT-V3) and Design-5 (AAV-GFAP-dRT-V4) possessed TH cells, and the number of TH-positive cells in the Design-5 (AAV-GFAP-dRT-V4) group was approximately twice that of the Design-4 (AAV-GFAP-dRT-V3) group. These results suggest that (1) in vitro study results are likely to be ineffective in the complex in vivo environment, (2) an increase in NLS may be a key factor in improving neuronal differentiation, and (3) an increase in VP64 activation domains can significantly improve the efficiency of glial cell differentiation into neurons, as well as significantly improve the efficiency of dopamine neuron production.
[0246] Example 2: Design of an artificial transcription factor and its expression vector based on RZFD The results of Example 1 demonstrate that the RZFD domain can be designed as an artificial transcription factor, enabling efficient differentiation of astrocytes into neurons or dopamine neurons in vivo. To further improve and optimize the design of RZFD-based artificial transcription factors (ATFs), RZFD is used in conjunction with one or more of the following: a nuclear localization sequence (NLS), a transcriptional activation domain (TAD), and / or an epigenetic regulatory domain (EPD). The ATF must contain a DNA-binding domain, which in this study is RZFD, and may be used alone or in combination with one or more of the nuclear localization sequence (NLS), transcriptional activation domain (TAD), and / or epigenetic regulatory domain (EPD), and the NLS, TAD, and / or EPD may be located at the N-terminus or the C-terminus of RZFD (Figure 2).
[0247] Due to the complexity of the in vivo environment, the in vivo differentiation-transition effect depends on various factors, and not only does the design of the artificial transcription factor affect the final effect, but the design of its expression vector also affects the final effect. As shown in Figure 3A, in this study, we used the astrocyte-specific promoter GFAP to drive the expression of the target gene (the target gene (GOI) is i.e., the artificial transcription factor (ATF)). Several vectors were designed so that GFAP initiates the expression of the target gene (e.g., Vector Design 1), several vectors were designed to improve mRNA stability by adding a WPRE regulatory sequence after the GOI (e.g., Vector Design 2), several vectors were designed to further include an enhancer before the promoter (e.g., Vector Design 3), and several vectors were designed to add both the enhancer and the WPRE regulatory sequence simultaneously (e.g., Vector Designs 4 and 5).
[0248] As shown in Figure 3B, the specific ATF design proposals used in this study are as follows: The RZFD core domain has the amino acid sequence shown in SEQ ID NO: 14. RZFD-V1 has an NLS added to the N-terminus of the RZFD core domain (containing eight zinc finger structures), with the nucleic acid sequence shown in SEQ ID NO: 56 and the amino acid sequence shown in SEQ ID NO: 57. RZFD-V2 has the activation domain of the transcriptional regulatory element VP64 added to the C-terminus of the core domain and an NLS added to the N-terminus, with the nucleic acid sequence shown in SEQ ID NO: 58 and the amino acid sequence shown in SEQ ID NO: 59. RZFD-V3 has the activation domain of P65-HSF1 added to the C-terminus of the core domain and an NLS added to the N-terminus, with the nucleic acid sequence shown in SEQ ID NO: 60 and the amino acid sequence shown in SEQ ID NO: 61. RZFD-V4 has bpNLS added to the C-terminus of the core domain. RZFD-V5 has bpNLS added to the N-terminus and C-terminus of the core domain, the nucleic acid sequence is shown in SEQ ID NO: 31, and the amino acid sequence is shown in SEQ ID NO: 30. RZFD-V6 has the VP64 activation domain further added to the N-terminus and C-terminus of RZFD-V5, the nucleic acid sequence is shown in SEQ ID NO: 33, and the amino acid sequence is shown in SEQ ID NO: 32. RZFD-V7 has the P65 and HSF1 activation domains further added to the C-terminus of RZFD-V6, the N-terminus remains the same as V6, the nucleic acid sequence is shown in SEQ ID NO: 35, and the amino acid sequence is shown in SEQ ID NO: 34. RZFD-V8 has the P65 and Rta activation domains further added to the C-terminus of RZFD-V6, the N-terminus remains the same as V6, and the amino acid sequence is shown in SEQ ID NO: 67. RZFD-V9 is RZFD-V5 with an additional VP64-P65-HSF1 activation domain added to the C-terminus, and its amino acid sequence is shown in SEQ ID NO: 73. RZFD-V10 is RZFD-V5 with an additional VP64-P65-Rta activation domain added to the C-terminus, and its amino acid sequence is shown in SEQ ID NO: 72. RZFD-V11 is RZFD-V5 with an additional VP64 activation domain added to the C-terminus, and its nucleic acid sequence is shown in SEQ ID NO: 91, and its amino acid sequence is shown in SEQ ID NO: 90. RZFD-V12 is RZFD-V6 with the linking order of NLS and VP64 reversed, and its nucleic acid sequence is shown in SEQ ID NO: 93, and its amino acid sequence is shown in SEQ ID NO: 92.
[0249] Of these, VP64, P65, HSF1, and Vta are all activation domains. In specific examples, bpNLS is used for the NLS, which has the amino acid sequence shown in SEQ ID NO: 24, and the nucleic acid sequence is shown in SEQ ID NO: 25. NLS is a nuclear localization signal and can improve the efficiency of entry into the cell nucleus. The activation domain of VP64 has the amino acid sequence shown in SEQ ID NO: 26, and the nucleic acid sequence is shown in SEQ ID NO: 27. The fusion form of the activation domains of P65 and HSF1 has the amino acid sequence shown in SEQ ID NO: 28, and the nucleic acid sequence is shown in SEQ ID NO: 29. The activation domain of Vta has the amino acid sequence shown in SEQ ID NO: 11.
[0250] Example 3: Increasing the enhancer and / or the transcriptional activation domain can increase the production of dopamine neurons.
[0251] To investigate whether increasing enhancer sequences in an expression vector can increase the efficiency of in vivo differentiation and thus the efficiency of obtaining dopamine neurons, this study designed designs V1 and V2, as shown in Figure 4A, based on RZFD-V5. As shown in Figure 4A, the RZFD core domain sequence originated from positions 155-419 of REST, driving the expression of the target gene using the astrocyte-specific promoter GFAP, with ATF (RZFD-V5) expressed in the middle, and WPRE and PloyA sequences downstream. Design V2 was based on Design V1, but with an increased CMV enhancer upstream of the GFAP promoter and an increased MVM enhancer downstream of the GFAP promoter. Design-V3 is based on Design-V2, but with the ATF (RZFD-V5) region replaced by ATF (RZFD-V6), and Design-V4 is based on Design-V3, but with the ATF (RZFD-V6) region replaced by ATF (RZFD-V2) (Figure 4A).
[0252] The vectors designed above were packaged into AAVs, and the control group AAV and experimental group AAVs were injected into the striatum of C57 mice used to construct the 6-OHDA model. Samples were collected 1.5 to 3 months after AAV injection and analyzed to detect the differentiation of astrocytes into dopaminergic neurons. Staining with the dopamine neuron-specific marker TH was performed, and as shown in Figure 4B, no TH-positive neurons were produced in the control group, but a large number of TH-positive dopamine neurons were produced in the striatum of mice injected with design-V1, design-V2, design-V3, or design-V4 (Figure 4B). Cell counting was performed on the TH-positive cells in the four groups (design-V1, design-V2, design-V3, or design-V4), and the statistical results are shown in Figure 4C. The results show that Design-V1 drove ATF(RZFD-V5) expression using GFAP, producing a minimum number of dopamine neurons, and that Design-V2, with increased enhancer sequences, produced approximately twice as many dopamine neurons as Design-V1. Design-V3 produced approximately twice as many dopamine neurons as Design-V2, and approximately 4.2 times as many as Design-V1 (Figure 4C). These results indicate that not only can increasing the enhancer sequence increase the number of TH-positive dopamine neurons, but increasing the activation domain can further increase the number of dopamine neurons. Statistical analysis of the number of dopamine neurons produced in Design-V4 revealed that the number of dopamine neurons produced in Design-V4 was close to that of Design-V3 (Figure 4C). This suggests that having one NLS and one VP64 located at opposite ends of the RZFD core domain is sufficient to achieve an effect similar to that of using two NLS-VP64s in combination.
[0253] Example 4: Partial deletion of the zinc finger domain in RZFD does not affect its in vivo differentiation effect.
[0254] The results above demonstrate that RZDF-V5 (NLS-RZFD(155-419)-NLS) is sufficient to convert astrocytes into dopamine neurons in vivo. To further investigate whether the RZFD(155-419) core domain can be further shortened, this study designed different RZFD cleavage sequences (Figure 5A), with the ATF region sequences being cleavage sequences at positions 159-412 of REST (amino acid sequence shown in SEQ ID NO: 2), 209-412 (amino acid sequence shown in SEQ ID NO: 74), 159-388 (amino acid sequence shown in SEQ ID NO: 75), and 209-388 (amino acid sequence shown in SEQ ID NO: 76).
[0255] As shown in Figure 5A, different AAV expression vectors were constructed. Of these, the RZFD core domain used by the artificial transcription factor of design-V2 is located at positions 155-419 of REST, with bpNLS located at the N-terminus and C-terminus of the RZFD core domain, respectively. The RZFD core domain used by the artificial transcription factor of design-V5 is located at positions 159-412 of REST (zinc finger domains 1-8), with bpNLS located at the N-terminus of the RZFD core domain. The RZFD core domain used by the artificial transcription factor of design-V6 is located at positions 209-412 of REST (zinc finger domains 1-8). The RZFD core domains are 1-7, and bpNLS is located at the N-terminus of the RZFD core domain. The RZFD core domain adopted by the artificial transcription factor of design-V7 is at positions 159-388 of REST (zinc finger domains 2-8), and bpNLS is located at the N-terminus of the RZFD core domain. The RZFD core domain adopted by the artificial transcription factor of design-V8 is at positions 209-388 of REST (zinc finger domains 2-7), and bpNLS is located at the N-terminus of the RZFD core domain. In plasmid construction, the expression of the artificial transcription factor was driven using the astrocyte-specific promoter GFAP, the CMV enhancer was increased upstream of the GFAP promoter, the MVM enhancer was increased downstream of the GFAP promoter, and WPRE was located downstream of the artificial transcription factor.
[0256] The above vectors were packaged into AAVs, and each group of AAVs was injected into the striatum of C57 mice used to construct the 6-OHDA model. Samples were collected 1.5 to 2 months after AAV injection and analyzed to detect the differentiation of astrocytes into dopaminergic neurons. Staining with the dopamine neuron-specific marker TH was performed, and the staining results are shown in Figure 5C. Compared to design-V2, design-V5 adopted a smaller core sequence (159-412) in design-V8, and the difference from design-1 was very small. The number of TH-positive neurons produced by this differentiation was also close to design-1, with no significant difference. Compared to design-V5, design-V6 contained only 7 of the 8 zinc finger domains (ZFD1-7), but the number of TH-positive neurons produced by this differentiation was also close to design-1, with no significant difference, indicating that the 7 zinc finger domains (ZFD2-8) were sufficient for function. The study shows that, compared to design-V5, design-V7 contains only 7 of the 8 zinc finger domains (ZFD2-8), but the number of TH-positive neurons produced by its transdifferentiation was close to that of design-1, with no significant difference, indicating that the 7 zinc finger domains (ZFD2-8) are sufficient for function. Similarly, compared to design-V5, design-V8 contains only 6 of the 8 zinc finger domains (ZFD2-7), but the number of TH-positive neurons produced by its transdifferentiation was close to that of design-1, with no significant difference, indicating that the 6 zinc finger domains (ZFD2-7) are sufficient for function. Statistical analysis of TH-positive cells in each group, as shown in Figure 5C, reveals that the in vivo transdifferentiation effects of these designs are equivalent, with similar numbers of dopamine neurons produced and similar transdifferentiation efficiencies.
[0257] Example 5: All different transcriptional activation domains function. To further investigate the effects of using different activation domains in conjunction with the RZFD core domain, this study designed Design-V9, Design-V10, and Design-V11, each different version, based on Design-V2. In this embodiment, the plasmid designs for Design-V2, Design-V9, Design-V10, and Design-V11 are shown in Figure 6A. The RZFD core domain sequence originates from positions 155-419 of REST, and the expression of the target gene is driven using the astrocyte-specific promoter GFAP. CMV is increased as an enhancer upstream of the GFAP promoter, MVM is increased as an enhancer downstream of the GFAP promoter, and WPRE is increased downstream of the target gene. Of these, the ATF of design-V9 is RZFD-V9, the activating domain is VP64-p65-HSF1, shown in SEQ ID NO: 84, and the nucleic acid sequence is shown in SEQ ID NO: 85. The ATF of design-V10 is RZFD-V10, the activating domain is VP64-p65-Rta, the amino acid sequence of RZFD-V10 is shown in SEQ ID NO: 86, and the nucleic acid sequence is shown in SEQ ID NO: 87. The ATF of design-V11 is RZFD-V11, the activating domain is VP64, the amino acid sequence of RZFD-V11 is shown in SEQ ID NO: 90, and the nucleic acid sequence is shown in SEQ ID NO: 91.
[0258] The expression vectors described above were packaged into AAVs, and the control group AAVs and experimental group AAVs were injected into the striatum of C57 mice used to construct the 6-OHDA model. Samples were collected 1.5 to 3 months after AAV injection and analyzed to detect the differentiation of astrocytes into dopaminergic neurons. Staining with the dopamine neuron-specific marker TH was performed, and as shown in Figure 6B, the striatums of mice injected with designs-V2, V9, V10, and V11 produced a large number of dopamine neurons (TH positive), while the control group had almost no dopamine neurons. Cell counting was performed on TH-positive cells in the four groups of designs-V2, V9, V10, and V11, and the statistical figures are shown in Figure 6C. Using Design-V2 (without the activation domain) as a baseline, the number of dopamine neurons in Design-V9 and Design-V10 groups was 1.42 times and 1.47 times that of Design-V2, respectively, while the number of dopamine neurons produced by Design-V11 group was approximately 1.99 times that of Design-V2 group. These results indicate that each transcriptional activator in each group can significantly improve the efficiency of differentiation of astrocytes into dopamine neurons.
[0259] Example 6: Comparative status of differentiation into dopaminergic neurons Conventional techniques have disclosed that inhibiting PTBP1 can induce differentiation of astrocytes into dopaminergic neurons. To compare differentiation efficiency, PTBP1 was knocked down using CRISPR / CasRx, and the differentiation efficiency was compared with that of the RZFD-V6 AAV in this example. The expression of the target gene was driven using the astrocyte-specific promoter GFAP (Figure 7A). Nucleic acids (the nucleic acid sequence is shown in SEQ ID NO: 83) containing gRNAs targeting CasRx and PTBP1 were packaged into an AAV vector. The packaging of the AAV virus followed standard experimental procedures in this field, and this group was designated the PTBP1-KD group. The AAV from the PTBP1-KD group and the AAV from RZFD-V6 were injected into the striatum of mice in a 6-OHDA model. Samples were collected and analyzed 1.5 to 2 months after AAV injection to detect the differentiation status of astrocytes into dopaminergic neurons. As shown in Figure 7B, the RZFD-V6 group produced more dopaminergic neurons compared to the PTBP1-KD group, approximately twice as many. This suggests that overexpression of RZFD-V6 can efficiently differentiate astrocytes into dopaminergic neurons.
[0260] Example 7: Analysis of differentiation of astrocytes into dopaminergic neurons in Dat-Cre::Ai9 tracking mice
[0261] The Dat-Cre::Ai9 mouse is a strain-following mouse that can be used to label mature dopamine neurons. In this mouse, only dopamine neurons that maturely express Dat can be labeled with red fluorescent protein (details of the mouse can be found in the technical documentation in this field). To label astrocytes in Dat-Cre::Ai9 mice, this study constructed the AAV expression vector shown in Figure 8A, with plasmid 1 for the control group and plasmid 2 for the experimental group. RZFD-V6 was selected as the artificial transcription factor (target gene, GOI), the amino acid sequence of RZFD-V6 is shown in SEQ ID NO: 32, and the nucleic acid sequence is shown in SEQ ID NO: 64. The control group (plasmid 1) used the astrocyte-specific promoter GFAP to drive EGFP expression in astrocytes for labeling, while the experimental group (plasmid 2) used the astrocyte-specific promoter GFAP to drive RZFD-V6 expression, increasing the CMV enhancer upstream of the GFAP promoter and increasing the MVM enhancer downstream of the GFAP promoter. The nucleic acid sequence of the CMV enhancer is shown in SEQ ID NO: 16, and the nucleic acid sequence of the MVM enhancer is shown in SEQ ID NO: 18.
[0262] The above vectors were packaged into AAV (AAV8) and injected into the striatum of Dat-Cre::Ai9 mice used to construct a 6-OHDA model. Samples were collected and analyzed 1.5 to 3 months after injection. The results are shown in Figure 8B. Green fluorescence indicates cells labeled with AAV-GFAP-EGFP, red fluorescence indicates tdTomato fluorescence expressed by mature dopamine neurons, white fluorescence is NeuN immunofluorescence staining (a neuron-specific protein marker), and blue indicates cell nuclei stained with DAPI. In the control group injected with AAV-GFAP-EGFP, a large amount of green fluorescent protein was expressed in astrocytes. These glial cells still retained a very typical astrocyte morphology, were not co-labeled with NeuN, and showed no tdTomato positive signal. In the group injected with GFAP-RZFD-V6, cells expressing green fluorescent protein exhibited typical neuronal morphology, and the green fluorescent protein signal was co-labeled with NeuN (e.g., cells indicated by yellow arrows in the GFAP-RZFD-V6 group in Figure 4), indicating that these cells had already been reprogrammed from astrocytes into neurons. In the brains of DAT-Cre::Ai9 mice, only mature dopamine neurons were labeled with the tdTomato red fluorescent signal, while in the striatum of mice injected with GFAP-RZFD-V6, many cells expressed tdTomato red fluorescence, and these cells were co-labeled with EGFP (e.g., cells indicated by white arrows in Figure 4), indicating that RZFD-V6 can reprogram astrocytes into mature dopamine neurons.
[0263] Example 8: Behavioral studies of PD model mice treated with RZFD-V6 To investigate the therapeutic effect of RZFD-V6 differentiation into dopamine neurons in the treatment of Parkinson's disease, this study constructed a mouse model using 6-OHDA in C57 mice of nearly matched age, and obtained a Parkinson's disease model. Baseline-level behavioral tests (cylinder tests) were performed on the obtained mouse model, and mice that met the group inclusion criteria were selected and randomly divided into two groups, each consisting of 12 mice. One group was a blank control group, and the other was a treatment group. The plasmid for the blank control group was designated as the control group (GFAP-mCherry) and drove mCherry expression with GFAP. In this example, the vector for the treatment group was as shown in the RZFD-V6 expression vector design in Figure 7A, with RZFD-V6 as the target gene. RZFD-V6 expression was driven using the astrocyte-specific promoter GFAP, the CMV enhancer was increased upstream of the GFAP promoter, the MVM enhancer was increased downstream of the GFAP promoter (abbreviated as eGFAP), and WPRE was increased downstream of the target gene. The striatal regions of the two groups of mice were injected with the GFAP-mCherry control and the experimental group AAV virus (AAV-eGFAP-RZFD-V6), respectively, and behavioral experiments (cylinder test) were performed 1.5 months after injection. As shown in Figure 9, the control group showed no behavioral differences before and after administration, while the treatment group exhibited significant behavioral improvement after administration compared to before administration. These results suggest that dopamine neurons produced by the reprogramming of astrocytes can exert dopamine neuronal function and significantly improve the symptoms of Parkinson's disease, potentially leading to the development of drugs to treat Parkinson's disease.
[0264] Example 9: When RZFD-V7 and RZFD-V8 reprogram astrocytes into dopamine neurons.
[0265] To further investigate whether all ligations of different activation domains to the RZFD core domain were functional, astrocytes were reprogrammed into neurons or dopamine neurons, and RZFD-V7 and RZFD-V8 were constructed as artificial transcription factors, as shown in Figure 2B. The amino acid sequence of RZFD-V7 is shown in SEQ ID NO: 34, and the nucleic acid sequence is shown in SEQ ID NO: 35. The amino acid sequence of RZFD-V8 is shown in SEQ ID NO: 67, and the nucleic acid sequence is shown in SEQ ID NO: 68. The expression vectors were packaged as AAVs, and GFAP-EGFP or GFAP-mCherry were used as controls.
[0266] Control group AAV and experimental group AAV were injected into the striatum of Parkinson's disease model mice constructed using the 6-OHDA model. Samples were collected and analyzed 1.5 to 3 months after AAV injection to detect the differentiation of astrocytes into neurons or dopaminergic neurons. Staining with the neuron-specific marker NeuN antibody revealed that the red fluorescent signal in the control group was not co-labeled with NeuN, and these cells still retained typical astrocyte morphology. In the AAV groups injected with RZFD-V7 and RZFD-V8, cells expressing the red fluorescent protein showed typical neuronal morphology and were co-labeled with NeuN. This indicates that overexpression of RZFD-V7 and RZFD-V8 in astrocytes can reprogram astrocytes into neurons. Simultaneously, to further investigate whether RZFD-V7 and RZFD-V8 can reprogram astrocytes into dopamine neurons, we stained them with the dopamine neuron-specific marker TH. The results showed that mice injected with RZFD-V7 and RZFD-V8 produced a large number of dopamine neurons in their striatum (TH positive), while the control group had almost no dopamine neurons. In summary, the results indicate that the linkage of different activation domains to the RZFD core domain can both enable the reprogramming of astrocytes into neurons or dopamine neurons.
[0267] Example 10: Differentiating Müller glial cells into retinal ganglion cells. Neurons are also present in retinal tissue, but their structure and cell types differ significantly from those in the striatum. The retina contains a large number of Müllerian glial cells and a small number of astrocytes, and there is a clear distinction between these two cell types. To further investigate whether Müllerian glial cells can be differentiated into retinal ganglion cells, this embodiment employs the specific promoter GFAP to drive the target gene RZFD-V6 to be overexpressed in retinal Müllerian glial cells.
[0268] First, Ai9 mice were modeled using NMDA, and endogenous retinal ganglion cells were removed. Two to three weeks after NMDA injection, the control group (AAV(GFAP-Cre)) and the experimental group (AAV(GFAP-Cre + GFAP-RZFD-V6)) were injected, respectively. One and a half to two months after injection, samples were collected and analyzed, and the number of red axons in the optic nerve was imaged and statistically analyzed. In the group injected with the control group AAV(GFAP-Cre), there were almost no red fluorescent signals (tdTomato-positive axons) in the optic nerve, but in the experimental group injected with GFAP-Cre + GFAP-RZFD-V6, many tdTomato-positive axons were observed. Staining with the RGC-specific marker Brn3a and RBPMS revealed clear tdTomato-positive RGC cells in the retina of the experimental group. These results indicate that the proposed design can differentiate Müller glial cells into RGCs.
[0269] Preferred embodiments of the present application are described herein, but these embodiments are provided merely as examples to those skilled in the art. Without departing from the technical application of the present disclosure, those skilled in the art can refer to the disclosure of embodiments and make various modifications and substitutions. When implementing the technical application of the present disclosure, it should be understood that various alternatives to the embodiments of the present disclosure described herein may be adopted.
[0270] The following are the sequence number and sequence information of the sequence relating to this application. [Table 1-1] [Table 1-2] Table 1-3 Table 1-4 Table 1-5 Table 1-6 Table 1-7 Table 1-8 Table 1-9 Table 1-10 Table 1-11 Table 1-12 Table 1-13 Table 1-14 Table 1-15 Table 1-16 Table 1-17 Table 1-18 Table 1-19 Table 1-20 Table 1-21 Table 1-22 Table 1-23 Table 1-24
Claims
1. An expression vector, (a) A nucleic acid sequence encoding an artificial transcription factor that can bind to RE1 and includes the RZFD core domain, (b) A regulatory element that modulates the nucleic acid expression of the artificial transcription factor, Of these, the amino acid sequence of the RZFD core domain includes any 5 to 8 sequences from sequence numbers 3 to 10, and does not include the sequences shown in sequence numbers 12 and 13. Preferably, the amino acid sequence of the RZFD core domain includes any 6 to 8 sequences from SEQ ID NOs: 3 to 10. Expression vector.
2. The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 76, or includes a sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:
76. The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 74 or 75, or includes a sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 74 or 75, or The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 2, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
2. The expression vector according to claim 1.
3. The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 14, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
14. The expression vector according to claim 2.
4. The coding nucleic acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 15 or 64, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 15 or 64. The expression vector according to claim 2.
5. The regulatory element includes a promoter that regulates the nucleic acid expression of an artificial transcription factor, and preferably the promoter includes a glial cell-specific promoter. An expression vector according to any one of claims 1 to 4.
6. The aforementioned glial cell-specific promoter is an astrocyte-specific promoter, a Müller cell-specific promoter, or a cochlear cell-specific promoter. The expression vector according to claim 5.
7. The aforementioned glial cell-specific promoters include the GFAP promoter, ALDH1L1 promoter, EAAT1 / GLAST promoter, glutamine synthetase promoter, S100β promoter, EAAT2 / GLT-1 promoter, Plp1 promoter, or Rlbp1 promoter. The expression vector according to claim 6.
8. The regulatory element further includes an enhancer that increases the transcription of the artificial transcription factor, Preferably, the enhancer includes a CMV enhancer, an SV40 intron-type enhancer, an MVM intron-type enhancer, a β-globulin intron-type enhancer, or a synthetic intron-type enhancer, or any combination thereof. The expression vector according to claim 5.
9. The nucleic acid sequence of the enhancer includes the sequence shown in any one of SEQ ID NOs: 16 to 22, or includes a sequence that has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the sequence shown in any one of SEQ ID NOs: 16 to 22. The expression vector according to claim 8.
10. The regulatory element further comprises a post-transcriptional enhancing element (WPRE) of woodchuck hepatitis virus. An expression vector according to any one of claims 5 to 9.
11. The artificial transcription factor further contains at least one nuclear localization sequence (NLS), Preferably, the NLS is selected from SV40 NLS, nucleoplasmin NLS, c-myc NLS, hRNPA1M9 NLS, nuclear importin α NLS, p53 NLS, influenza virus NS1 NLS, hepatitis D antigen NLS, or bpNLS. An expression vector according to any one of claims 1 to 10.
12. NLS is located at the N-terminus and / or C-terminus of the RZFD core domain. The expression vector according to claim 11.
13. The amino acid sequence of NLS includes the sequence shown in SEQ ID NO: 24, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% identity with SEQ ID NO:
24. The expression vector according to claim 11.
14. The coding nucleic acid sequence of the NLS includes the sequence shown in SEQ ID NO: 25, or includes a sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% identical to the sequence shown in SEQ ID NO:
25. The expression vector according to claim 13.
15. The artificial transcription factor further comprises a gene activation domain, preferably the gene activation domain being located at the C-terminus and / or N-terminus of the RZFD core domain. An expression vector according to any one of claims 1 to 10.
16. The gene activation domain includes the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination of these activation domains. Preferably, the gene activation domain includes the VP64 activation domain, the P65 activation domain, the RTA activation domain, or the HSF1 activation domain, or any combination thereof. The expression vector according to claim 15.
17. The activation domain is located at the C-terminus and / or N-terminus of the RZFD core domain. The expression vector according to claim 16.
18. The amino acid sequence of the activating domain of VP64 includes the sequence described in SEQ ID NO: 26, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
26. The amino acid sequence of the activation domain of P65 comprises the sequence shown in SEQ ID NO: 69, or comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
69. The amino acid sequence of the activation domain of the RTA includes the sequence shown in SEQ ID NO: 11, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
11. The amino acid sequence of the activation domain of HSF1 includes the sequence shown in SEQ ID NO: 70, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
70. The expression vector according to claim 16 or 17.
19. The coding nucleic acid sequence of the VP64 activation domain includes the sequence shown in SEQ ID NO: 27, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
27. The expression vector according to claim 18.
20. The gene activation domain comprises the VP64 activation domain and the P65 activation domain, and preferably further comprises the HSF1 activation domain or the RTA activation domain. The expression vector according to claim 16.
21. The gene activation domain is located at the N-terminus and / or C-terminus of the RZFD core domain. The expression vector according to claim 20.
22. The gene activation domain includes the VP64 activation domain, the P65 activation domain, and the HSF1 activation domain. Preferably, the amino acid sequence of the activation domain includes the sequence shown in SEQ ID NO: 73, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
73. The expression vector according to claim 20 or 21.
23. The gene activation domain includes the VP64 activation domain, the P65 activation domain, and the RTA activation domain. Preferably, the amino acid sequence of the activation domain includes the sequence shown in SEQ ID NO: 72, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
72. An expression vector according to either claim 20 or 21.
24. The amino acid sequence of the gene activation domain includes the sequence shown in SEQ ID NOs. 28, 69, 70, or 71, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NOs. 28, 69, 70, or 71. The expression vector according to claim 16.
25. The artificial transcription factor further comprises a nuclear localization sequence (NLS) and a gene activation domain, preferably the NLS and / or gene activation domain are located at the N-terminus and / or C-terminus of the RZFD core domain, and more preferably the method of linking the N-terminus to the C-terminus of the artificial transcription factor is RZFD core domain - NLS gene activation domain, RZFD core domain - gene activation domain - NLS, NLS-gene activation domain-RZFD core domain, Gene activation domain - NLS-RZFD core domain, NLS-RZFD core domain - gene activation domain, Gene activation domain - RZFD core domain - NLS, Gene activation domain-NLS-RZFD core domain-NLS-gene activation domain, or NLS-gene activation domain-RZFD core domain-gene activation domain-NLS, An expression vector according to any one of claims 1 to 10.
26. The gene activation domain includes the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination thereof. Preferably, the gene activation domain includes the VP64 activation domain, the P65 activation domain, the RTA activation domain, or the HSF1 activation domain, or any combination thereof. The expression vector according to claim 25.
27. The amino acid sequence of the activation domain of VP64 includes the sequence shown in SEQ ID NO: 26, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
26. The amino acid sequence of the activation domain of P65 comprises the sequence shown in SEQ ID NO: 69, or comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
69. The amino acid sequence of the activation domain of the RTA includes the sequence shown in SEQ ID NO: 11, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
11. The amino acid sequence of the activation domain of HSF1 includes the sequence shown in SEQ ID NO: 70, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
70. The expression vector according to claim 26.
28. The amino acid sequence of the NLS includes the sequence shown in SEQ ID NO: 24, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% identity with SEQ ID NO:
24. The expression vector according to claim 25.
29. The RZFD core domain and the NLS are linked via a linker, or The RZFD core domain and the gene activation domain are linked via a linker, or The NLS and the gene activation domain are linked via a linker. An expression vector according to any one of claims 25 to 28.
30. The artificial transcription factor includes a sequence represented by any one of SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with a sequence represented by any one of SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92. An expression vector according to any one of claims 1 to 29.
31. The coding nucleic acid sequence of the artificial transcription factor includes the sequence shown in any one of SEQ ID NOs: 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93, or includes a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the sequence shown in any one of SEQ ID NOs: 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93. An expression vector according to any one of claims 1 to 30.
32. The expression vector contains the sequence shown in any one of SEQ ID NOs. 36, 37, or 62, or contains a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the sequence shown in any one of SEQ ID NOs. 36, 37, or 62. The expression vector according to claim 1.
33. A composition, (a) an artificial transcription factor capable of binding to RE1, comprising an RZFD core domain, a first NLS, and a first gene activation domain, or (b) comprising a nucleic acid encoding the artificial transcription factor described in (a), Of these, the RZFD core domain contains any 5 to 8 sequences from sequence numbers 3 to 10, and does not contain the sequences shown in sequence numbers 12 and 13. Preferably, the amino acid sequence of the RZFD core domain includes any 6 to 8 sequences from SEQ ID NOs: 3 to 10. composition.
34. The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 76, or includes a sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:
76. The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 74 or 75, or includes a sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 74 or 75, or The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 2, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
2. The composition according to claim 33.
35. The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 14, or includes a sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:
14. The coding nucleic acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 15 or 64, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 15 or 64. The composition according to claim 33.
36. The method for ligating the RZFD core domain, the first NLS, and the first gene activation domain from the N-terminus to the C-terminus is as follows: 1) RZFD core domain - first NLS - first gene activation domain, 2) First gene activation domain - first NLS - RZFD core domain, 3) RZFD core domain - first gene activation domain - first NLS, 4) First NLS - First gene activation domain - RZFD core domain, 5) First NLS-RZFD core domain - First gene activation domain, 6) Selected from the first gene activation domain - RZFD core domain - first NLS, The composition according to claim 33.
37. The artificial transcription factor further comprises a second NLS, The composition according to claim 33.
38. The first NLS is located at the N-terminus of the RZFD core domain, and the second NLS is located at the C-terminus of the RZFD core domain. The composition according to claim 37.
39. The artificial transcription factor further comprises a second gene activation domain. The composition according to any one of claims 33 to 38.
40. The second gene activation domain and the first activation domain are located at opposite ends of the RZFD core domain, The composition according to claim 39.
41. The first NLS and / or the second NLS comprises the sequence shown in SEQ ID NO: 24, or comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% identity with SEQ ID NO:
24. The composition according to any one of claims 33 to 40.
42. The coding nucleic acid sequences of the first NLS and / or the second NLS include the sequence shown in SEQ ID NO: 25, or include a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% identity with SEQ ID NO:
25. The composition according to claim 41.
43. The first gene activation domain includes the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination thereof. Preferably, the first gene activation domain includes the VP64 activation domain, the P65 activation domain, the RTA activation domain, or the HSF1 activation domain, or any combination thereof. The composition according to any one of claims 33 to 42.
44. The amino acid sequence of the activating domain of VP64 includes the sequence shown in SEQ ID NO: 26, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
26. The amino acid sequence of the activation domain of P65 comprises the sequence shown in SEQ ID NO: 69, or comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
69. The amino acid sequence of the activation domain of the RTA includes the sequence shown in SEQ ID NO: 11, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
11. The amino acid sequence of the activation domain of HSF1 includes the sequence shown in SEQ ID NO: 70, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
70. The composition according to claim 43.
45. The coding nucleic acid sequence of the VP64 activation domain includes the sequence shown in SEQ ID NO: 27, or includes a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
27. The composition according to claim 44.
46. The first gene activation domain includes the VP64 activation domain and the P65 activation domain. Preferably, the first gene activation domain further comprises an HSF1 activation domain or an RTA activation domain. The composition according to claim 43.
47. The first gene activation domain includes the VP64 activation domain, the P65 activation domain, and the HSF1 activation domain. Preferably, the amino acid sequence of the first activation domain includes the sequence shown in SEQ ID NO: 73, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
73. The composition according to claim 46.
48. The first gene activation domain includes the VP64 activation domain, the P65 activation domain, and the RTA activation domain. Preferably, the amino acid sequence of the first activation domain includes the sequence shown in SEQ ID NO: 72, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
72. The composition according to claim 46.
49. The amino acid sequence of the first gene activation domain includes the sequence shown in SEQ ID NOs. 28, 69, 70, or 71, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NOs. 28, 69, 70, or 71. The composition according to claim 43.
50. The second gene activation domain includes the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination of these activation domains. Preferably, the second gene activation domain includes the VP64 activation domain, More preferably, the second gene activation domain comprises the VP64 activation domain and the P65 activation domain, and even more preferably, the second gene activation domain further comprises the HSF1 activation domain or the RTA activation domain. The composition according to claim 39.
51. The amino acid sequence of the activating domain of VP64 includes the sequence shown in SEQ ID NO: 26, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
26. The amino acid sequence of the activation domain of P65 comprises the sequence shown in SEQ ID NO: 69, or comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
69. The amino acid sequence of the activation domain of the RTA includes the sequence shown in SEQ ID NO: 11, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
11. The amino acid sequence of the activation domain of HSF1 includes the sequence shown in SEQ ID NO: 70, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
70. The composition according to claim 50.
52. The coding nucleic acid sequence of the VP64 activation domain includes the sequence shown in SEQ ID NO: 27, or includes a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
27. The composition according to claim 51.
53. The second gene activation domain comprises the VP64 activation domain and the P65 activation domain, and preferably further comprises the HSF1 activation domain or the RTA activation domain. The composition according to claim 39.
54. The aforementioned second gene activation domain includes the VP64 activation domain, the P65 activation domain, and the HSF1 activation domain. Preferably, the amino acid sequence of the second activation domain includes the sequence shown in SEQ ID NO: 73, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
73. The composition according to claim 53.
55. The aforementioned second gene activation domain includes the VP64 activation domain, the P65 activation domain, and the RTA activation domain. Preferably, the amino acid sequence of the second activation domain includes the sequence shown in SEQ ID NO: 72, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
72. The composition according to claim 53.
56. The amino acid sequence of the second gene activation domain includes the sequence shown in SEQ ID NOs. 28, 69, 70, or 71, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NOs. 28, 69, 70, or 71. The composition according to claim 39.
57. The artificial transcription factor includes the sequence shown in SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the sequence shown in any one of SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92. The composition according to any one of claims 33 to 56.
58. The coding nucleic acid sequence of the artificial transcription factor includes the sequence shown in SEQ ID NOs: 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93, or includes a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NOs: 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93. The composition according to claim 57.
59. The composition further comprises a regulatory element that modulates the nucleic acid expression of the artificial transcription factor, preferably the regulatory element comprising a promoter that modulates the nucleic acid expression of the artificial transcription factor, and more preferably the promoter comprising a glial cell-specific promoter. The composition according to any one of claims 33 to 58.
60. The aforementioned glial cell-specific promoter includes an astrocyte-specific promoter, a Müller cell-specific promoter, or a cochlear glial cell-specific promoter. Preferably, the glial cell-specific promoter includes the GFAP promoter, ALDH1L1 promoter, EAAT1 / GLAST promoter, glutamine synthetase promoter, S100β promoter, EAAT2 / GLT-1 promoter, Plp1 promoter, or Rlbp1 promoter. The composition according to claim 59.
61. The regulatory element further includes an enhancer that increases the transcription of the artificial transcription factor, Preferably, the enhancer includes a CMV enhancer, an SV40 intron-type enhancer, an MVM intron-type enhancer, a β-globulin intron-type enhancer, or a synthetic intron-type enhancer, or any combination thereof. The composition according to claim 59.
62. The regulatory element further comprises a post-transcriptional enhancing element (WPRE) of woodchuck hepatitis virus. The composition according to claim 59.
63. A nucleic acid sequence comprising the sequence shown in SEQ ID NO: 36, 37, or 62, or comprising a nucleic acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 36, 37, or 62, The composition according to claim 59.
64. A method for differentiating non-neuronal cells of an individual into functional neurons, comprising contacting the non-neuronal cells with an expression vector according to any one of claims 1 to 32, or a composition according to any one of claims 33 to 63. Preferably, the contact is an in vitro contact or an in vivo contact. method.
65. The functional neurons include dopamine neurons, retinal ganglion cells, photoreceptor cells, cochlear spiral ganglion cells, GABA neurons, 5-HT neurons, glutamatergic neurons, ChAT neurons, NE neurons, motor neurons, spinal neurons, spinal motor neurons, spinal sensory neurons, bipolar cells, amacrine neurons, pyramidal cells, interneurons, medium spiny neurons (MSNs), Purkinje cells, granule cells, olfactory sensory neurons, juxtaglomerular cells, or any combination thereof. The method according to claim 64.
66. The dopamine neurons express one or more of the following markers: NeuN, TH, FoxA2, Nurr1, Pitx3, Vmat2, or DAT. The optic ganglion neurons express one or more of the following markers: RBPMS, Pax6, Brn3a, Brn3b, Brn3c, and Map2. The aforementioned photoreceptor cells express one or more markers from the following: rhodopsin, mCAR, m-opsin, and S-opsin. The aforementioned cochlear spiral ganglion cells express one or more of the following markers: NeuN, Prox1, Tuj-1, and Map2. The method according to claim 65.
67. The functional neuron includes an axon, The method according to any one of claims 64 to 66.
68. The aforementioned non-neuronal cells include glial cells, fibroblasts, stem cells, neural progenitor cells, or neural stem cells. Preferably, the glial cells include astrocytes, oligodendrocytes, ependymal cells, Schwann cells, NG2 cells, satellite cells, Müller cells, and cochlear nerve cells. The method according to any one of claims 64 to 66.
69. The success rate of converting non-neuronal cells into functional neurons is at least 1%. The method according to any one of claims 64 to 68.
70. A method for treating a disease in an individual in need, comprising administering a therapeutically effective amount of an expression vector according to any one of claims 1 to 32, or a therapeutically effective amount of a composition according to any one of claims 33 to 63. Preferably, the diseases include Parkinson's disease, Alzheimer's disease, stroke, schizophrenia, Huntington's disease, depression, motor neuron disease, ALS, spinal muscular atrophy, epilepsy, ataxia, visual impairment due to RGC cell death, glaucoma, age-related RGC damage, optic nerve damage, focal ischemia or hemorrhage of the retina, hereditary optic nerve disease, degeneration or death of photoreceptor cells due to trauma or degeneration, macular degeneration, retinitis pigmentosa, and diabetes-related blindness. Treatment method.
71. An artificial transcription factor comprising an RZFD core domain, wherein the amino acid sequence of the RZFD core domain includes any 5 to 8 sequences from SEQ ID NOs. 3 to 10, and does not include the sequences shown in SEQ ID NOs. 12 and 13, preferably the amino acid sequence of the RZFD core domain includes any 6 to 8 sequences from SEQ ID NOs. 3 to 10. Artificial transcription factors.
72. The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 76, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
76. Preferably, the amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 74 or 75, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 74 or 75. Preferably, the amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 2, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
2. The artificial transcription factor according to claim 71.
73. The amino acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 14, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
14. The artificial transcription factor according to claim 72.
74. The coding nucleic acid sequence of the RZFD core domain includes the sequence shown in SEQ ID NO: 15 or 64, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 15 or 64. The artificial transcription factor according to claim 72.
75. The artificial transcription factor further contains at least one nuclear localization sequence (NLS), Preferably, the NLS is selected from SV40 NLS, nucleoplasmin NLS, c-myc NLS, hRNPA1M9 NLS, nuclear importin α NLS, p53 NLS, influenza virus NS1 NLS, hepatitis D antigen NLS, or bpNLS. An artificial transcription factor according to any one of claims 71 to 74.
76. NLS is located at the N-terminus and / or C-terminus of the RZFD core domain. Preferably, the amino acid sequence of the NLS includes the sequence shown in SEQ ID NO: 24, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% identity with SEQ ID NO:
24. Preferably, the coding nucleic acid sequence of the NLS includes the sequence shown in SEQ ID NO: 25, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% identity with the sequence shown in SEQ ID NO:
25. The artificial transcription factor according to claim 75.
77. Artificial transcription factors further contain gene activation domains, Preferably, the gene activation domain is located at the C-terminus and / or N-terminus of the RZFD core domain. An artificial transcription factor according to any one of claims 71 to 76.
78. The gene activation domain includes the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination of these activation domains. Preferably, the gene activation domain includes the VP64 activation domain, the P65 activation domain, the RTA activation domain, or the HSF1 activation domain, or any combination thereof. The artificial transcription factor according to claim 77.
79. The activation domain is located at the C-terminus and / or N-terminus of the RZFD core domain. The artificial transcription factor according to claim 78.
80. The amino acid sequence of the activating domain of VP64 includes the sequence shown in SEQ ID NO: 26, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
26. The amino acid sequence of the activation domain of P65 comprises the sequence shown in SEQ ID NO: 69, or comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
69. The amino acid sequence of the activation domain of the RTA includes the sequence shown in SEQ ID NO: 11, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
11. The amino acid sequence of the activation domain of HSF1 includes the sequence shown in SEQ ID NO: 70, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
70. The artificial transcription factor according to claim 78.
81. The coding nucleic acid sequence of the VP64 activation domain includes the sequence shown in SEQ ID NO: 27, or includes a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
27. The artificial transcription factor according to claim 80.
82. The gene activation domain comprises the VP64 activation domain and the P65 activation domain, and preferably further comprises the HSF1 activation domain or the RTA activation domain. The artificial transcription factor according to claim 78.
83. The gene activation domain includes the VP64 activation domain, the P65 activation domain, and the HSF1 activation domain. Preferably, the amino acid sequence of the activation domain includes the sequence shown in SEQ ID NO: 73, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
73. An artificial transcription factor according to any one of claims 77 to 82.
84. The gene activation domain includes the VP64 activation domain, the P65 activation domain, and the RTA activation domain. Preferably, the amino acid sequence of the activation domain includes the sequence shown in SEQ ID NO: 72, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
72. An artificial transcription factor according to any one of claims 77 to 82.
85. The amino acid sequence of the gene activation domain includes the sequence shown in any one of SEQ ID NOs. 28, 69, 70, or 71, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NOs. 28, 69, 70, or 71. The artificial transcription factor according to claim 78.
86. Artificial transcription factors further include a nuclear localization sequence (NLS) and a gene activation domain. Preferably, the NLS and / or gene activation domain is located at the N-terminus and / or C-terminus of the RZFD core domain. More preferably, the method for ligating the N-terminus to the C-terminus of the artificial transcription factor is: RZFD core domain - NLS gene activation domain, RZFD core domain - gene activation domain - NLS, NLS-gene activation domain-RZFD core domain, Gene activation domain - NLS-RZFD core domain, NLS-RZFD core domain - gene activation domain, Gene activation domain - RZFD core domain - NLS, Gene activation domain-NLS-RZFD core domain-NLS-gene activation domain, or NLS-gene activation domain-RZFD core domain-gene activation domain-NLS, An artificial transcription factor according to any one of claims 71 to 74.
87. The gene activation domain includes the activation domains of VP64, P65, HSF1, VP16, RTA, Suntag, P300, or CBP, or any combination of any activation domains, or any combination of these activation domains. Preferably, the gene activation domain includes the VP64 activation domain, the P65 activation domain, the RTA activation domain, or the HSF1 activation domain, or any combination thereof. Preferably, the amino acid sequence of the activation domain of VP64 includes the sequence shown in SEQ ID NO: 26, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
26. The amino acid sequence of the activation domain of P65 includes the sequence shown in SEQ ID NO: 69, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 69, or The amino acid sequence of the activation domain of the RTA includes the sequence shown in SEQ ID NO: 11, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 11, or The amino acid sequence of the activation domain of HSF1 includes the sequence shown in SEQ ID NO: 70, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:
70. The artificial transcription factor according to claim 86.
88. The amino acid sequence of the NLS includes the sequence shown in SEQ ID NO: 24, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% identity with SEQ ID NO:
24. The artificial transcription factor according to claim 86.
89. The RZFD core domain and the NLS are linked via a linker, or the RZFD core domain and the gene activation domain are linked via a linker, or the NLS and the gene activation domain are linked via a linker. An artificial transcription factor according to any one of claims 86 to 88.
90. The artificial transcription factor includes a sequence represented by any one of SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92, or includes an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with a sequence represented by any one of SEQ ID NOs: 2, 14, 30, 32, 34, 57, 59, 61, 63, 67, 77, 78, 79, 80, 81, 82, 84, 86, 90, or 92. An artificial transcription factor according to any one of claims 71 to 89.
91. The coding nucleic acid sequence of the artificial transcription factor includes the sequence shown in any one of SEQ ID NOs: 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93, or includes a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the sequence shown in any one of SEQ ID NOs: 15, 31, 33, 35, 56, 58, 60, 64, 68, 85, 87, 91, or 93. An artificial transcription factor according to any one of claims 71 to 90.