Method for editing target RNA
Artificial DYW proteins with DYW domains enable efficient C-to-U and U-to-C RNA editing, addressing the lack of such methods in existing technologies and offering applications in gene therapy and gene mutagenesis.
Patent Information
- Application Number
- JP2025073407
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-03-31
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2041-03-31
AI Technical Summary
Existing methods lack a systematic approach to design molecules that can convert cytidine (C) to uridine (U) or vice versa in RNA sequences using the DYW domain and artificial RNA-binding proteins, particularly in plant groups that exhibit U-to-C RNA editing.
The development of artificial DYW proteins containing specific DYW domains, such as DYW:PG, DYW:WW, and DYW:KP, which are fused with PPR or PLS arrays to achieve C-to-U and U-to-C RNA editing activities, utilizing the PPR-code for targeted RNA binding.
These DYW proteins enable efficient conversion of C to U or U to C in RNA sequences, with editing efficiencies ranging from 3% to 70%, applicable in various fields including medicine and agriculture for gene therapy and gene mutagenesis.
Smart Images

Figure 2025107239000014 
Figure 2025107239000015 
Figure 2025107239000016
Abstract
Description
Technical Field
[0001] The present invention relates to an RNA editing technique using a protein capable of binding to a target RNA. The present invention is useful in a wide range of fields such as medicine (drug discovery support, treatment), agriculture (agricultural, forestry, livestock, and fishery product production, breeding), and chemistry (biological substance production).
Background Art
[0002] In plant mitochondria and chloroplasts, RNA editing in which specific bases in the genome are substituted at the RNA level frequently occurs, and it is known that this phenomenon is mediated by the pentatricopeptide repeat (PPR) protein, an RNA-binding protein.
[0003] PPR proteins are classified into two families, P and PLS, according to the structure of the PPR motifs that make up this protein (Non-Patent Document 1). While P-class PPR proteins are composed of simple repeats of the standard 35-amino acid PPR motif (P), PLS proteins contain, in addition to P, two motifs called L and S that are similar to it. The PPR array (the arrangement of PPR motifs) of PLS proteins is composed of three PPR motifs, P1 (about 35 amino acids), L1 (about 35 amino acids), and S1 (about 31 amino acids), as a P-L-S repeating unit, and on the C-terminal side of P1L1S1, a P-L-S with a slightly different sequence, that is, P2 (35 amino acids), L2 (36 amino acids), and S2 (32 amino acids) motifs follow. In addition to being composed of P-L-S repeating units, SS (31 amino acids) may be repeated. Further, on the C-terminal side of this last P2L2S2 motif, two PPR-like motifs called E1 and E2, and a DYW domain having a 136-amino acid cytidine deaminase domain-like sequence may follow (Non-Patent Document 2).
[0004] The interaction between PPR proteins and RNA is defined by the PPR code that specifies the bound RNA bases by combinations of several amino acids in each PPR motif, and these amino acids bind to the corresponding nucleotides via hydrogen bonds (Patent Document 1, Patent Document 2, Non-Patent Document 3, Non-Patent Document 4, Non-Patent Document 5, Non-Patent Document 6, Non-Patent Document 7, Non-Patent Document 8).
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Non-Patent Documents
[0006]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Non-Patent Document 12
Summary of the Invention
Problems to be Solved by the Invention
[0007] In the organelles of land plants, C-to-U RNA editing, in which RNA bases are substituted from cytidine to uridine, generally occurs. However, in mosses, some leptosporangiate ferns, and ferns, U-to-C RNA editing, in which uridine is substituted with cytidine, is also observed (Non-Patent Document 9). From two different bioinformatics studies, unique DYW domains with sequences different from the standard DYW domain were discovered in plants with U-to-C RNA editing (Non-Patent Document 10, Non-Patent Document 11).
[0008] Two DYW:PG proteins derived from Physcomitrella patens have been reported to exhibit C-to-U RNA editing activity in Escherichia coli (Non-Patent Document 12). In plant groups that branched early in the evolution of land plants, the DYW domain is broadly classified into two groups. The first is called the DYW:PG / WW group, which includes the DYW:PG type, which is a standard DYW domain, and the DYW:WW type, in which one tryptophan (W) is added to the PG box. The second group is called DYW:KP. This DYW domain has a sequence in the PG box and the three amino acid sequences at the C-terminus that are different from those of the PG / WW group, and is present only in plants with U-to-C RNA editing. The technique of changing one base in an arbitrary RNA sequence to a specific base is useful in gene therapy and gene mutagenesis techniques in industrial applications. To date, various RNA-binding molecules have been developed, but a method for designing a molecule that converts cytidine (C) of an arbitrary target RNA to uridine (U) or vice versa, uridine (U) to cytidine (C), using the DYW domain and an artificial RNA-binding protein has not been established.
Means for Solving the Problems
[0009] In this study, we show that the portion containing the C-terminal DYW domain of the PPR-DYW protein can be used as a modular editing domain for RNA editing. When each of the three DYW domains designed in this study was fused with a PPR-P or PLS array, the DYW:PG and DYW:WW domains showed C-to-U, and the DYW:KP showed U-to-C RNA editing activity.
[0010] The present invention provides the following. [1] A method for editing a target RNA, which comprises applying an artificial DYW protein containing a DYW domain consisting of any one of the following polypeptides a, b, c, and bc to the target RNA. a. x a1 PGx a2 SWIEx a3 -xa16 HP … Hx aa E … Cx a17 x a18 A polypeptide having CH … DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity b. x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx bb E … Cx b17 x b18 A polypeptide having CH … DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity c. KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH … x c6 x c7 x c8 Having, having at least 40% sequence identity with the sequence of SEQ ID NO: 3, and having C-to-U / U-to-C editing activity bc. x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx cc E … Cx c4 x c5 CH … Dx bc1 x bc2 Having, having at least 40% sequence identity with the 90 sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity (In the sequence, x represents any amino acid, and … represents any polypeptide fragment.) [2] A DYW domain consisting of any one of the following polypeptides a, b, c, and bc a. x a1 PGx a2 SWIEx a3 -x a16 HP … Hx aa E … Cxa17 x a18 A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity b. x b1 PGx b2 SWWTDx b3 -x b16 HP…Hx bb E…Cx b17 x b18 A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity c. KPAx c1 Ax c2 IEx c3 …Hx cc E…Cx c4 x c5 CH…x c6 x c7 x c8 A polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity bc. x b1 PGx b2 SWWTDx b3 -x b16 HP…Hx cc E…Cx c4 x c5 CH…Dx bc1 x bc2 A polypeptide having at least 40% sequence identity with the 90 sequence of SEQ ID NO: 2 and having C-to-U / U-to-C editing activity [3] A DYW protein comprising at least one PPR motif, an RNA binding domain capable of specifically binding to a target RNA, and the DYW domain described in 2 [4] The DYW protein according to 2, wherein the RNA binding domain is of the PLS type [5] A method for editing a target RNA, comprising applying a DYW domain consisting of the polypeptide of c or bc below to the target RNA to convert the target U to C c. KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH … x c6 x c7 x c8 having and having at least 40% sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity, a polypeptide bc. x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx cc E … Cx c4 x c5 CH … Dx bc1 x bc2 having and having at least 40% sequence identity with the 90 sequence of SEQ ID NO: 2 and having C-to-U / U-to-C editing activity, a polypeptide [6] The method according to 5, wherein the DYW domain comprises at least one PPR motif and is fused to an RNA binding domain capable of specifically binding to a target RNA according to the rules of the PPR-code. [7] A composition comprising the DYW domain according to 2 for editing RNA in a eukaryotic cell. [8] A nucleic acid encoding the DYW domain according to 2, or the DYW protein according to 3 or 4. [9] A vector comprising the nucleic acid according to 8.
[10] A cell (excluding human individuals) comprising the vector according to 9.
[0011] [1] A method for editing a target RNA, applying an artificial DYW protein comprising a DYW domain consisting of any one of the following polypeptides a to c to the target RNA. a. x a1 PGx a2 SWIEx a3 -x a16 HP … HSE … Cx a17 x a18A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity b. x b1 PGx b2 SWWTDx b3 -x b16 HP…HSE…Cx b17 x b18 A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity c. KPAx c1 Ax c2 IEx c3 …HAE…Cx c4 x c5 CH…Dx c6 x c7 A polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3, and having C-to-U / U-to-C editing activity (In the sequence, x represents any amino acid, and … represents any polypeptide fragment.) [2] A DYW domain consisting of any one of the following polypeptides a to c a. x a1 PGx a2 SWIEx a3 -x a16 HP…HSE…Cx a17 x a18 A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity b. x b1 PGx b2 SWWTDx b3 -x b16 HP…HSE…Cx b17 x b18 A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity c. KPAx c1 Ax c2 IEx c3 …HAE…Cxc4 x c5 CH … Dx c6 x c7 having and having at least 40% sequence identity with the sequence of SEQ ID NO: 3, and having C-to-U / U-to-C editing activity [3] A DYW protein comprising at least one PPR motif, an RNA binding domain capable of specifically binding to a target RNA, and the DYW domain described in 2. [4] The DYW protein according to 2, wherein the RNA binding domain is of the PLS type. [5] A method for editing a target RNA, comprising the step of applying a DYW domain consisting of the polypeptide of c below to the target RNA to convert the editing target U to C. c. KPAx c1 Ax c2 IEx c3 … HAE … Cx c4 x c5 CH … Dx c6 x c7 having and having at least 40% sequence identity with the sequence of SEQ ID NO: 3, and having C-to-U / U-to-C editing activity [6] The method according to 5, wherein the DYW domain comprises at least one PPR motif and is fused to an RNA binding domain capable of specifically binding to a target RNA according to the rules of the PPR-code. [7] A composition for editing RNA in a eukaryotic cell, comprising the DYW domain described in 2. [8] A nucleic acid encoding the DYW domain described in 2, or the DYW protein described in 3 or 4. [9] A vector comprising the nucleic acid described in 8.
[10] A cell (excluding human individuals) comprising the vector described in 9. [Advantages of the Invention]
[0012] According to the present invention, the editing target C contained in the target RNA can be converted to U, or the editing target U can be converted to C. [Brief Description of the Drawings]
[0013]
Figure 1
Figure 2
Figure 3
Figure 4-1
Figure 4-2
Figure 4-3
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0014] The present invention relates to a method for editing a target RNA by applying a DYW protein containing a DYW domain consisting of any one of the following polypeptides a, b, c, and bc to the target RNA. a. x a1 PGx a2 SWIEx a3 -x a16 HP … Hx aa E … Cx a17 x a18 Having CH … DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity b. x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx bb E … Cx b17 x b18 Having CH … DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity c. KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH … x c6 x c7 x c8 Having, having at least 40% sequence identity with the sequence of SEQ ID NO: 3, and having C-to-U / U-to-C editing activity bc. x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx cc E … Cx c4 x c5 CH … Dx bc1 x bc2having and having at least 40% sequence identity with the sequence of 90 of SEQ ID NO: 2 and having C-to-U / U-to-C editing activity
[0015] With respect to the present invention, the C-to-U / U-to-C editing activity refers to the activity capable of converting an editing target C contained in a target RNA to U or an editing target U to C when an editing assay is performed by ligating a target polypeptide to the C-terminal side of an RNA binding domain capable of specifically binding to the target RNA. The conversion may be such that at least about 3%, preferably about 5% of the editing target bases are replaced with the target bases under appropriate conditions.
[0016] [DYW domain] x a1 PGx a2 SWIEx a3 -x a16 HP … Hx aa E … Cx a17 x a18 CH … DYW, x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx bb E … Cx b17 x b18 CH … DYW, and KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH … x c6 x c7 x c8 each represents an amino acid sequence. In the sequences, x each independently represents any amino acid, and … each independently represents a polypeptide fragment consisting of an amino acid sequence of any length. With respect to the present invention, the DYW domain can be represented by any one of these three amino acid sequences. In particular, x a1 PGx a2 SWIEx a3 -x a16 HP … Hx aa E … Cx a17 x a18The DYW domain consisting of CH … DYW is DYW:PG, x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx bb E … Cx b17 x b18 The DYW domain consisting of CH … DYW is DYW:WW, and KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH … x c6 x c7 x c8 The DYW domain consisting of... may be represented as DYW:KP.
[0017] The DYW domain has a region containing a PG box consisting of approximately 15 amino acids at the N-terminus, a zinc-binding domain in the center (HxEx n CxxCH, where x n is any number n of consecutive arbitrary amino acids.), and three regions of DYW at the C-terminus. The zinc-binding domain can be further divided into an HxE region and a CxxCH region. These regions of each DYW domain can be represented as shown in the following table.
[0018]
Table 1
[0019] (DYW:PG) DYW:PG is x a1 PGx a2 SWIEx a3 -x a16 HP … Hx aa E … Cx a17 x a18 A polypeptide consisting of CH … DYW. Preferably, x a1 PGx a2 SWIEx a3 -x a16 HP … Hx aa E … Cx a17 xa18 A polypeptide having CH…DYW, having sequence identity with the sequence of SEQ ID NO:1 (detailed in the section of [terms]), and having C-to-U / U-to-C editing activity. DYW:PG has the activity of converting the editing target C to U (C-to-U editing activity). SEQ ID NO:1 shows the sequence of DYW:PG consisting of 136 amino acids in total length used in the experiments shown in the Examples section of this specification. This sequence is disclosed for the first time in this application and is novel.
[0020] The full length of DYW:PG is not particularly limited as long as it can exhibit C-to-U editing activity. For example, it is 110 to 160 amino acids in length, preferably 124 to 148 amino acids in length, more preferably 128 to 144 amino acids in length, and even more preferably 132 to 140 amino acids in length.
[0021] In the region (x a1 PGx a2 SWIEx a3 -x a16 HP) of DYW:PG: x a1 is not particularly limited as long as DYW:PG can exhibit C-to-U editing activity, but is preferably E (glutamic acid) or an amino acid with similar properties, and more preferably G. x a2 is not particularly limited as long as DYW:PG can exhibit C-to-U editing activity, but is preferably C (cysteine) or an amino acid with similar properties, and more preferably C. x a3 -x a16 Each amino acid of is not particularly limited as long as DYW:PG can exhibit C-to-U editing activity, but is preferably the same as or an amino acid with similar properties to the corresponding amino acids at positions 9 to 22 of the sequence of SEQ ID NO:1, and more preferably the same amino acids as the corresponding amino acids at positions 9 to 22 of the sequence of SEQ ID NO:1.
[0022] In one preferred embodiment, the HxE region of DYW:PG is HSE regardless of the situation of other regions.
[0023] In the CxxCH region of DYW:PG, that is, Cx a17 x a18 CH: x a17 is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG, but is preferably G (glycine) or an amino acid with similar properties, more preferably G. x a18 is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG, but is preferably D (aspartic acid) or an amino acid with similar properties, more preferably D.
[0024] In DYW:PG, the region containing the PG box and the part connecting to Hx aa E, the part connecting the Hx aa E region and the CxxCH region, and the part connecting the CxxCH region and DYW are referred to as the first connecting part, the second connecting part, and the third connecting part in order (the same applies to other DYW domains).
[0025] The total length of the first connecting part of DYW:PG is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG. For example, it is 39 - 47 amino acids in length, preferably 40 - 46 amino acids in length, more preferably 41 - 45 amino acids in length, and even more preferably 42 - 44 amino acids in length. The amino acid sequence of the first connecting part is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG, but is preferably the same as the part at positions 25 - 67 of the sequence of SEQ ID NO:1, or a sequence in which 1 - 22 amino acids are substituted, deleted, or added in that partial sequence, or a sequence having sequence identity with that partial sequence, and more preferably, the same sequence as that partial sequence.
[0026] One of the preferred embodiments of the first linking portion of DYW:PG is a polypeptide represented by the following formula, which is 43 amino acids long, regardless of the sequence of other parts of the DYW domain.
[0027] N a25 -N a26 -N a27 - … -N a65 -N a66 -N a67
[0028] The above polypeptide is preferably the same as the portion at positions 25 to 67 of the sequence of SEQ ID NO: 1, or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit C-to-U editing activity as DYW:PG. At this time, the amino acid substitution is preferably such that the amino acids with large bits values (for example, N a29 、N a30 、N a32 、N a33 、N a35 、N a36 、N a40 、N a44 、N a45 、N a47 、N a48 、N a52 、N a53 、N a54 、N a55 、N a58 、N a61 、N a65 、N a67 ) are the same as in Figure 4, and it is preferred that the other amino acids are substituted.
[0029] The total length of the second linker of DYW:PG is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG. For example, it is 21 to 29 amino acids in length, preferably 22 to 28 amino acids in length, more preferably 23 to 27 amino acids in length, and even more preferably 24 to 26 amino acids in length. The amino acid sequence of the second linker is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG. Preferably, it is the same as the portion at positions 71 to 95 of the sequence of SEQ ID NO:1, or a sequence in which 1 to 13 amino acids are substituted, deleted, or added in that partial sequence, or a sequence having sequence identity with that partial sequence. More preferably, it is the same sequence as that partial sequence.
[0030] One of the preferred embodiments of the second linker of DYW:PG is a polypeptide represented by the following formula with a length of 25 amino acids, regardless of the sequence of other parts of the DYW domain.
[0031] N a71 -N a72 -N a73 - … -N a93 -N a94 -N a95
[0032] The above polypeptide is preferably the same as the portion at positions 71 to 95 of the sequence of SEQ ID NO:1, or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit C-to-U editing activity as DYW:PG. At this time, the amino acid substitution is an amino acid with a large bits value (for example, N a71 、N a72 、N a73 、N a76 、N a77 、N a78 、N a79 、N a81 、N a82 、N a86 、N a88 、N a89 、N a91 、N a92 、N a93 、N a94) is the same as FIG. 4, and it is preferable that other amino acids are substituted.
[0033] The total length of the third linker of DYW:PG is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG. For example, it is 29 to 37 amino acids in length, preferably 30 to 36 amino acids in length, more preferably 31 to 35 amino acids in length, and even more preferably 32 to 34 amino acids in length. The amino acid sequence of the third linker is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG. Preferably, it is the same as the portion at positions 101 to 133 of the sequence of SEQ ID NO: 1, or a sequence in which 1 to 17 amino acids are substituted, deleted, or added in that partial sequence, or a sequence having sequence identity with that partial sequence. More preferably, it is the same sequence as that partial sequence.
[0034] One of the preferred embodiments of the third linker of DYW:PG is a polypeptide represented by the following formula with a length of 33 amino acids, regardless of the sequence of other parts of the DYW domain.
[0035] N a101 -N a102 -N a103 - … -N a131 -N a132 -N a133
[0036] The above polypeptide is preferably the same as the portion at positions 101 to 133 of the sequence of SEQ ID NO: 1, or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit C-to-U editing activity as DYW:PG. At this time, the amino acid substitution is an amino acid with a large bits value (for example, N a102 、N a104 、N a107 、N a112 、N a114 、N a117 、N a118 、N a121 、N a122 、N a123 、N a124 、N a125 、Na128 , N a130 , N a131 , N a132 ) is the same as that in FIG. 4, and it is preferable that other amino acids are substituted.
[0037] By introducing a mutation into the PG domain consisting of the sequence of SEQ ID NO: 1, the RNA editing activity can be improved. A preferred example of such a domain with the mutation introduced is PG11 (a polypeptide consisting of the amino acid sequence of SEQ ID NO: 50) shown in the Examples section of this specification.
[0038] (DYW:WW) DYW:WW is x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx bb E … Cx b17 x b18 A polypeptide consisting of CH … DYW. Preferably, x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx bb E … Cx b17 x b18 A polypeptide having CH … DYW, having sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity. DYW:WW has the activity of converting the editing target C to U (C-to-U editing activity). SEQ ID NO: 2 shows the sequence of DYW:WW consisting of 137 amino acids in total length used in the experiments shown in the Examples section of this specification. This sequence is disclosed for the first time in this application and is novel.
[0039] The full length of DYW:WW is not particularly limited as long as it can exhibit C-to-U editing activity. For example, it is 110 to 160 amino acids in length, preferably 125 to 149 amino acids in length, more preferably 129 to 145 amino acids in length, and even more preferably 133 to 141 amino acids in length.
[0040] Region containing the PG box of DYW:WW, i.e., x b1 PGx b2 SWWTDx b3 -x b16 In HP, the part consisting of WTD may be WSD.
[0041] In the region containing the PG box of DYW:WW: x b1 is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW, but is preferably K (lysine) or an amino acid with similar properties, more preferably K. x b2 is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW, but is preferably Q (glutamine) or an amino acid with similar properties, more preferably Q. x b3 -x b16 Each amino acid of is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW, but is preferably the same as or an amino acid with similar properties to the corresponding amino acids at positions 10 - 23 of the sequence of SEQ ID NO: 2, more preferably the same amino acids as the corresponding amino acids at positions 10 - 23 of the sequence of SEQ ID NO: 2.
[0042] In one preferred embodiment, the HxE region of DYW:WW is HSE regardless of the sequences of other parts.
[0043] The CxxCH region of DYW:WW, i.e., Cx b17 x b18 In CH: x b17 is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW, but is preferably D (aspartic acid) or an amino acid with similar properties, more preferably D. x b18 is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW, but is preferably D or an amino acid with similar properties, more preferably D.
[0044] The total length of the first linker of DYW:WW is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW. For example, it is 39 to 47 amino acids in length, preferably 40 to 46 amino acids in length, more preferably 41 to 45 amino acids in length, and even more preferably 42 to 44 amino acids in length. The amino acid sequence of the first linker is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW. Preferably, it is the same as the portion at positions 25 to 67 of the sequence of SEQ ID NO: 2, or a sequence in which 1 to 22 amino acids are substituted, deleted, or added in that partial sequence, or a sequence having sequence identity with that partial sequence. More preferably, it is the same sequence as that partial sequence.
[0045] One preferred embodiment of the first linker of DYW:WW is a polypeptide represented by the following formula with a length of 43 amino acids, regardless of the sequence of other parts of the DYW domain.
[0046] N b26 -N b27 -N b28 - … -N b66 -N b67 -N b68
[0047] The above polypeptide is preferably the same as the portion at positions 26 to 68 of the sequence of SEQ ID NO: 2, or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit C-to-U editing activity as DYW:PG. At this time, the amino acid substitution is an amino acid with a large bits value at the corresponding position in FIG. 4 (for example, N b26 、N b30 、N b33 、N b34 、N b37 、N b41 、N b45 、N b46 、N b48 、N b49 、N b51 、N b52 、N b53 、N b55 、N b56 、N b57 、Nb59 , N b61 , N b62 , N b63 , N b64 , N b66 , N b67 , N b68 ) is the same as that in FIG. 4, and it is preferable that other amino acids are substituted.
[0048] The total length of the second linker of DYW:WW is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW. For example, it is 21 to 29 amino acids in length, preferably 22 to 28 amino acids in length, more preferably 23 to 27 amino acids in length, and even more preferably 24 to 26 amino acids in length. The amino acid sequence of the second linker is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW. Preferably, it is the same as the portion at positions 71 to 95 of the sequence of SEQ ID NO: 2, or a sequence in which 1 to 13 amino acids are substituted, deleted, or added in that partial sequence, or a sequence having sequence identity with that partial sequence. More preferably, it is the same sequence as that partial sequence.
[0049] One of the preferred embodiments of the second linker of DYW:WW is a polypeptide represented by the following formula with a length of 25 amino acids, regardless of the sequence of other parts of the DYW domain.
[0050] N b72 -N b73 -N b74 - … -N b94 -N b95 -N b96
[0051] The above polypeptide is preferably the same as the portion at positions 72 to 96 of the sequence of SEQ ID NO: 2, or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit C-to-U editing activity as DYW:WW. At this time, the amino acid substitution is an amino acid with a large bits value at the corresponding position in FIG. 4 (for example, N b72 , N b73 , N b74 , N b75 , Nb77 , N b78 , N b79 , N b81 , N b82 , N b84 , N b88 , N b89 , N b90 , N b91 , N b92 , N b93 , N b94 , N b95 , N b96 ) is the same as Figure 4, and it is preferably performed such that other amino acids are substituted.
[0052] The total length of the third linker of DYW:WW is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW. For example, it is 29 to 37 amino acids in length, preferably 30 to 36 amino acids in length, more preferably 31 to 35 amino acids in length, and even more preferably 32 to 34 amino acids in length. The amino acid sequence of the third linker is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW. Preferably, it is the same as the portion at positions 101 to 133 of the sequence of SEQ ID NO: 2, or a sequence in which 1 to 17 amino acids are substituted, deleted, or added in that partial sequence, or a sequence having sequence identity with that partial sequence. More preferably, it is the same sequence as that partial sequence.
[0053] One preferred embodiment of the third linker of DYW:WW is a polypeptide represented by the following formula having a length of 33 amino acids, regardless of the sequence of other parts of the DYW domain.
[0054] N b102 -N b103 -N b104 - … -N b132 -N b133 -N b134
[0055] The above polypeptide is preferably the same as the portion at positions 102 to 134 of the sequence of SEQ ID NO: 2, or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit C-to-U editing activity as DYW:WW. At this time, the amino acid substitution is an amino acid with a large bits value (for example, N b104 , N b105 , N b107 , N b108 , N b109 , N b110 , N b111 , N b113 , N b115 , N b116 , N b117 , N b118 , N b119 , N b121 , N b122 , N b123 , N b124 , N b126 , N b129 , N b131 , N b132 , N b133 N b134 ) is the same as in Figure 4, and it is preferable that the substitution of other amino acids is carried out.
[0056] By introducing a mutation into the WW domain consisting of the sequence of SEQ ID NO: 2, the RNA editing activity can be improved. Preferred examples of such a mutated domain are WW2 to 11 and WW13 shown in the Examples section of this specification, and particularly, the one with high editing activity is WW11 (a polypeptide consisting of the amino acid sequence of SEQ ID NO: 63).
[0057] (DYW:KP) DYW:KP is a polypeptide consisting of KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH … x c6 x c7 x c8 . Preferably, KPAx c1 Ax c2 IEx c3… Hx cc E … Cx c4 x c5 CH … x c6 x c7 x c8 It is a polypeptide having the sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity. DYW:KP has the activity of converting the editing target U to C (U-to-C editing activity). SEQ ID NO: 3 shows the sequence of DYW:KP consisting of 133 amino acids in total length used in the experiments shown in the Examples section of this specification. This sequence is disclosed for the first time in this application and is novel.
[0058] The full length of DYW:KP is not particularly limited as long as it can exhibit U-to-C editing activity. For example, it is 110 to 160 amino acids in length, preferably 121 to 145 amino acids in length, more preferably 125 to 141 amino acids in length, and even more preferably 129 to 137 amino acids in length.
[0059] The region containing the PG box of DYW:KP, that is, KPAx c1 Ax c2 IEx c3 wherein: x c1 is not particularly limited as long as DYW:KP can exhibit U-to-C editing activity, but is preferably S (serine) or an amino acid with similar properties, and more preferably S. x c2 is not particularly limited as long as DYW:KP can exhibit U-to-C editing activity, but is preferably L (leucine) or an amino acid with similar properties, and more preferably L. x c3 is not particularly limited as long as DYW:KP can exhibit U-to-C editing activity, but is preferably V (valine) or an amino acid with similar properties, and more preferably V.
[0060] In one of the preferred embodiments, the HxE region of DYW:KP is HAE regardless of the sequences of other parts.
[0061] The CxxCH region of DYW:KP, i.e., Cx c4 x c5 in CH: x c4 is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably N (asparagine) or an amino acid with similar properties, more preferably N. x b5 is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably D (aspartic acid) or an amino acid with similar properties, more preferably D.
[0062] The part corresponding to DYW of DYW:KP, i.e., x c6 x c7 x c8 in: x c6 is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably D (aspartic acid) or an amino acid with similar properties, more preferably D. x c7 is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably M (methionine) or an amino acid with similar properties, more preferably M. x b8 is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably F (phenylalanine) or an amino acid with similar properties, more preferably F. In one preferred embodiment, x c6 x c7 x c8 is Dx c7 x c8 regardless of the sequences of other parts. In another preferred embodiment, x c6 x c7 x c8 is GRP regardless of the sequences of other parts.
[0063] The total length of the first linker of DYW:KP is not particularly limited as long as DYW:KP can exhibit U-to-C editing activity. For example, it is 51 to 59 amino acids in length, preferably 52 to 58 amino acids in length, more preferably 53 to 57 amino acids in length, and even more preferably 54 to 56 amino acids in length. The amino acid sequence of the first linker is not particularly limited as long as DYW:KP can exhibit U-to-C editing activity. Preferably, it is the same as the portion at positions 10 to 64 of the sequence of SEQ ID NO: 3, or a sequence in which 1 to 28 amino acids are substituted, deleted, or added in that partial sequence, or a sequence having sequence identity with that partial sequence. More preferably, it is the same sequence as that partial sequence.
[0064] One preferred embodiment of the first linker of DYW:KP is a polypeptide represented by the following formula with a length of 55 amino acids, regardless of the sequence of other parts of the DYW domain.
[0065] N c10 -N c11 -N c12 - … -N c62 -N c63 -N c64
[0066] The above polypeptide is preferably the same as the portion at positions 10 to 64 of the sequence of SEQ ID NO: 3, or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit U-to-C editing activity as DYW::KP. At this time, for amino acid substitution, amino acids with large bits values (for example, N c10 、N c13 、N c14 、N c15 、N c16 、N c17 、N c18 、N c19 、N c25 、N c26 、N c29 、N c30 、N c33 、N c34 、N c36 、N c38 、N c41 、Nc42 , N c44 , N c45 , N c47 , N 49 , N c55 , N c58 , N c59 , N c62 , N c63 , N c64 ) is the same as that in Figure 4, and it is preferably performed such that other amino acids are substituted.
[0067] The full length of the second linker of DYW:KP is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP. For example, it is 21 to 29 amino acids in length, preferably 22 to 28 amino acids in length, more preferably 23 to 27 amino acids in length, and even more preferably 24 to 26 amino acids in length. The amino acid sequence of the second linker is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP. Preferably, it is the same as the portion at positions 68 to 92 of the sequence of SEQ ID NO: 3, or a sequence in which 1 to 13 amino acids are substituted, deleted, or added in that partial sequence, or a sequence having sequence identity with that partial sequence. More preferably, it is the same sequence as that partial sequence.
[0068] One of the preferred embodiments of the second linker of DYW:KP is a polypeptide represented by the following formula with a length of 25 amino acids, regardless of the sequence of other parts of the DYW domain.
[0069] N c68 -N c69 -N c70 - … -N c90 -N c91 -N c92
[0070] The above polypeptide is preferably the same as the portion at positions 68 to 92 of the sequence of SEQ ID NO: 3, or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit U-to-C editing activity as DYW:KP. At this time, the amino acid substitution is an amino acid with a large bits value at the corresponding position in Figure 4 (for example, N c68 , Nc70 , N c71 , N c72 , N c73 , N c74 , N c75 , N c76 , N c77 , N c78 , N c79 , N c80 , N c81 , N c83 , N c84 , N c85 , N c86 , N c87 , N c88 , N c89 , N c90 , N c91 , N c92 ) is the same as that in FIG. 4, and it is preferably carried out such that other amino acids are substituted.
[0071] The full length of the third linker of DYW:KP is not particularly limited as long as DYW:KP can exhibit U-to-C editing activity. For example, it is 29 to 37 amino acids in length, preferably 30 to 36 amino acids in length, more preferably 31 to 35 amino acids in length, and even more preferably 32 to 34 amino acids in length. The amino acid sequence of the third linker is not particularly limited as long as DYW:KP can exhibit U-to-C editing activity. Preferably, it is the same as the portion at positions 98 to 130 of the sequence of SEQ ID NO: 3, or a sequence in which 1 to 17 amino acids are substituted, deleted, or added in that partial sequence, or a sequence having sequence identity with that partial sequence. More preferably, it is the same sequence as that partial sequence.
[0072] One preferred embodiment of the third linker of DYW:KP is a polypeptide represented by the following formula with a length of 33 amino acids, regardless of the sequence of other parts of the DYW domain.
[0073] N c98 -N c99 -N c100 - … -N c128 -N c129 -N c130
[0074] The above polypeptide is preferably the same as the portion at positions 98 to 130 of the sequence of SEQ ID NO: 3, or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit U-to-C editing activity as DYW:KP. At this time, the amino acid substitution is such that the amino acid with a large bits value (for example, N c98 、N c100 、N c101 、N c103 、N c104 、N c105 、N c107 、N c109 、N c110 、N c111 、N c112 、N c114 、N c115 、N c117 、N c118 、N c119 、N c120 、N c121 、N c122 、N c123 、N c124 、N c125 、N c127 、N c129 、N c130 ) is the same as in FIG. 4, and it is preferable that the other amino acids are substituted.
[0075] By introducing a mutation into the KP domain consisting of the sequence of SEQ ID NO: 3, the editing activity can be improved. A preferred example of such a mutated domain is KP2 to 23 (SEQ ID NOs: 68 to 89) shown in the Examples section of this specification. KP22 (SEQ ID NO: 88) has the highest editing activity from U to C and the lowest editing activity from C to U, and the RNA editing activity is improved as compared with the KP domain consisting of the sequence of SEQ ID NO: 3.
[0076] (Chimeric DYW) The DYW domain can be divided into several regions based on the conservation of the amino acid sequence. By Chimeric DYW in which these regions are exchanged, the C-to-U editing activity or U-to-C editing activity can be improved. One of the preferred embodiments is the region containing the PG box of DYW:WW, that is, xb1 PGx b2 SWWTDx b3 -x b16 HP and DYW are fused to other regions of DYW:KP (... Hx cc E... Cx c4 x c5 CH...) and are fused at this time. The x b1 , x b2 , x b3 -x b16 is as described above for DYW:WW. Also, the first connecting part, the second connecting part, and the third connecting part are as described above for DYW:KP. The DYW part is Dx bc1 x bc2 It may be. The full length is not particularly limited as long as it can exhibit U-to-C editing activity. For example, it is, for example, 110 to 160 amino acids in length, preferably 125 to 149 amino acids in length, more preferably 129 to 145 amino acids in length, and even more preferably 133 to 141 amino acids in length.
[0077] One of the preferred Chimeric domains is x b1 PGx b2 SWWTDx b3 -x b16 HP... Hx cc E... Cx c4 x c5 CH... Dx bc1 x bc2 It has, has at least 40% sequence identity with the sequence of 90 of SEQ ID NO: 2, and consists of a polypeptide having C-to-U / U-to-C editing activity.
[0078] One of the particularly preferred Chimeric domains is a polypeptide having sequence identity with the sequence of SEQ ID NO: 90 and having U-to-C editing activity. SEQ ID NO: 90 shows the sequence of the Chimeric domain used in the experiment shown in the Examples section of this specification. This domain shows higher U to C editing activity than DYW:KP consisting of the sequence of SEQ ID NO: 3 and has almost no C to U editing activity.
[0079] (Comparison with known sequences) As described above, the DYW domain consisting of the sequences of SEQ ID NOs: 1, 2, and 3 is novel. The results of an investigation using the PPR database (https: / / ppr.plantenergy.uwa.edu.au / onekp / ; Non-Patent Document 11 cited above) are shown below.
[0080] [Table 2-1]
[0081] Accordingly, the present invention also provides a polypeptide of any one of d to i below: d. A polypeptide having a sequence identity of more than 78% with the sequence of SEQ ID NO: 1, preferably 80% or more, more preferably 85% or more, still more preferably 90%, still more preferably 95%, still more preferably 97%, and having C-to-U editing activity; e. A polypeptide having a sequence identity of more than 84% with the sequence of SEQ ID NO: 2, preferably 85% or more, more preferably 90% or more, still more preferably 95%, still more preferably 97%, and having C-to-U editing activity; f. A polypeptide having a sequence identity of more than 86% with the sequence of SEQ ID NO: 3, preferably 87% or more, more preferably 90% or more, still more preferably 95%, still more preferably 97%, and having U-to-C editing activity. g. A polypeptide having a sequence in which 1 to 29, preferably 1 to 25, more preferably 1 to 21, still more preferably 1 to 17, still more preferably 1 to 13, still more preferably 1 to 9, still more preferably 1 to 5 amino acids in the sequence of SEQ ID NO: 1 are substituted, deleted, or added, and having C-to-U editing activity; h. A polypeptide having a sequence in which 1 to 21, preferably 1 to 18, more preferably 1 to 15, still more preferably 1 to 12, still more preferably 1 to 9, still more preferably 1 to 6 amino acids are substituted, deleted, or added in the sequence of SEQ ID NO: 2, and having C-to-U editing activity; i. A polypeptide having a sequence in which 1 to 18, preferably 1 to 16, more preferably 1 to 14, still more preferably 1 to 12, still more preferably 1 to 10, still more preferably 1 to 8 amino acids, and more preferably 1 to 6 amino acids are substituted, deleted, or added in the sequence of SEQ ID NO: 3, and having U-to-C editing activity.
[0082] Also, the alignment of these sequences is shown below. The alignment was created using AliView (Larsson, A. (2014). AliView: a fast and lightweight alignment viewer and editor for large data sets. Bioinformatics 30(22): 3276 - 3278. http: / / dx.doi.org / 10.1093 / bioinformatics / btu531). Note that identical amino acids are represented by dots.
[0083] [Table 2 - 2]
[0084] [RNA binding domain] In the present invention, when converting an editing target using the DYW domain, in order to target the RNA containing the editing target, an RNA binding protein is used as the RNA binding domain.
[0085] A preferred example of the RNA binding protein used as the RNA binding domain is a PPR protein composed of PPR motifs.
[0086] (PPR motif) Unless otherwise specified, the PPR motif refers to a polypeptide composed of 30 to 38 amino acids having an amino acid sequence with an E-value of a predetermined value or less (preferably E-03) obtained in PF01535 in Pfam and PS51375 in Prosite when analyzing the amino acid sequence with a protein domain search program on the web. The position numbers of the amino acids constituting the PPR motif defined in the present invention are almost synonymous with PF01535, while corresponding to the number obtained by subtracting 2 from the position of the amino acid in PS51375 (e.g., position 1 in the present invention → position 3 in PS51375). However, when referring to the amino acid at position "ii" (-2), it means the second amino acid from the end (C-terminal side) of the amino acids constituting the PPR motif, or two amino acids on the N-terminal side with respect to the first amino acid of the next PPR motif, that is, the -2 amino acid. When the next PPR motif is not clearly identified, the amino acid two positions before the first amino acid of the next helix structure is defined as "ii". For Pfam, refer to http: / / pfam.sanger.ac.uk / , and for Prosite, refer to http: / / www.expasy.org / prosite / .
[0087] The conserved amino acid sequence of the PPR motif has low conservation at the amino acid level, but two α-helices are well conserved in the secondary structure. A typical PPR motif is composed of 35 amino acids, but its length is variable from 30 to 38 amino acids.
[0088] More specifically, the PPR motif consists of a polypeptide 30 to 38 amino acids in length represented by Formula 1.
[0089]
Chemical formula
[0090] In the formula: Helix A is a part capable of forming an α-helix structure 12 amino acids in length and is represented by Formula 2,
[0091]
Chemical formula
[0092] In Formula 2, A1 to A 12 each independently represents an amino acid; X is a portion that is absent or consists of 1 to 9 amino acids in length; Helix B is a portion consisting of 11 to 13 amino acids in length and capable of forming an α-helix structure; L is a portion represented by Formula 3 and consisting of 2 to 7 amino acids in length;
[0093] [Chemical formula]
[0094] In Formula 3, each amino acid is numbered from the C-terminal side as "i" (-1), "ii" (-2), and so on, provided that iii ~L vii may be absent.
[0095] (PPR-code) In the PPR motif, the combination of the three amino acids at positions 1, 4, and ii is important for specific binding to a base, and these combinations can determine which base is bound. The relationship between the combination of the three amino acids at positions 1, 4, and ii and the base to which they can bind is known as the PPR-code (Patent Document 2 cited above) and is as follows.
[0096] (1) When the combination of the three amino acids of A1, A4, and L ii is valine, asparagine, and aspartic acid in that order, the PPR motif has selective RNA base binding ability to strongly bind to U, then to C, and then to A or G. (2) A1, A4, and L iiWhen the combination of three amino acids is valine, threonine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, then to G, then to C, but not to U. (3) A1, A4, and L ii When the combination of three amino acids is valine, asparagine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to C, then to A or U, but not to G. (4) A1, A4, and L ii When the combination of three amino acids is glutamic acid, glycine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G, but not to A, U, or C. (5) A1, A4, and L ii When the combination of three amino acids is isoleucine, asparagine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to C, then to U, then to A, but not to G. (6) A1, A4, and L ii When the combination of three amino acids is valine, threonine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G, then to U, but not to A or C. (7) A1, A4, and L ii When the combination of three amino acids is lysine, threonine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G, then to A, but not to U or C. (8) A1, A4, and L ii When the combination of three amino acids is phenylalanine, serine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, then to C, then to G and U. (9) A1, A4, and L iiWhen the combination of three amino acids is valine, asparagine, and serine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to C and then to U, but does not bind to A and G. (10) A1, A4, and L ii When the combination of three amino acids is phenylalanine, threonine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, but does not bind to G, U, and C. (11) A1, A4, and L ii When the combination of three amino acids is isoleucine, asparagine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to U and then to A, but does not bind to G and C. (12) A1, A4, and L ii When the combination of three amino acids is threonine, threonine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, but does not bind to G, U, and C. (13) A1, A4, and L ii When the combination of three amino acids is isoleucine, methionine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to U and then to C, but does not bind to A and G. (14) A1, A4, and L ii When the combination of three amino acids is phenylalanine, proline, and aspartic acid in that order for PPR, its motif has selective RNA base-binding ability such that it binds strongly to U and then to C, but does not bind to A and G. (15) A1, A4, and L ii When the combination of three amino acids is tyrosine, proline, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to U, but does not bind to A, G, and C. (16) A1, A4, and L iiIn the case where the combination of three amino acids is leucine, threonine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G but not to A, U, and C.
[0097] (P array) PPR proteins are classified into two families, P and PLS, according to the structure of the constituent PPR motifs. The P-type PPR protein is composed of a simple repeat (P array) of a standard 35-amino acid PPR motif (P). The DYW domain of the present invention can be used by being linked to the P-type PPR protein.
[0098] (PLS array) The arrangement of the PPR motifs of the PLS-type PPR protein is composed of repeating units of three PPR motifs, P1, L1, and S1, and on the C-terminal side of the repeat, P2, L2, and S2 motifs follow. Further, on the C-terminal side of this last P2L2S2 motif, two PPR-like motifs called E1 and E2, and a DYW domain may follow.
[0099] The full length of P1 is not particularly limited as long as it can bind to the target base, but for example, it is 33 to 37 amino acids in length, preferably 34 to 36 amino acids in length, and more preferably 35 amino acids in length. The full length of L1 is not particularly limited as long as it can bind to the target base, but for example, it is 33 to 37 amino acids in length, preferably 34 to 36 amino acids in length, and more preferably 35 amino acids in length. The full length of S1 is not particularly limited as long as it can bind to the target base, but for example, it is 30 to 33 amino acids in length, preferably 30 to 32 amino acids in length, and more preferably 31 amino acids in length.
[0100] The full length of P2 is not particularly limited as long as it can bind to the target base. For example, it is 33 to 37 amino acids in length, preferably 34 to 36 amino acids in length, and more preferably 35 amino acids in length. The full length of L2 is not particularly limited as long as it can bind to the target base. For example, it is 34 to 38 amino acids in length, preferably 35 to 37 amino acids in length, and more preferably 36 amino acids in length. The full length of S2 is not particularly limited as long as it can bind to the target base. For example, it is 30 to 34 amino acids in length, preferably 31 to 33 amino acids in length, and more preferably 32 amino acids in length. SEQ ID NOs: 17, 22, and 27 show the sequences of P2 used in the examples section of this specification. SEQ ID NOs: 18, 23, and 28 show the sequences of L2 used in the examples section of this specification. SEQ ID NOs: 19, 24, and 29 show the sequences of S2 used in the examples section of this specification.
[0101] The full length of E1 is not particularly limited as long as it can bind to the target base. For example, it is 32 to 36 amino acids in length, preferably 33 to 35 amino acids in length, and more preferably 34 amino acids in length. SEQ ID NOs: 20, 25, and 30 show the sequences of E1 used in the examples section of this specification.
[0102] The full length of E2 is not particularly limited as long as it can bind to the target base. For example, it is 30 to 34 amino acids in length, preferably 31 to 33 amino acids in length, and more preferably 33 amino acids in length. SEQ ID NOs: 21, 26, and 31 show the sequences of E2 used in the examples section of this specification.
[0103] In the PLS-type PPR protein, the repetitive part of P1L1S1 and the part up to P2 can be designed according to the above-mentioned PPR-code rules according to the sequence of the target RNA.
[0104] The S2 motif has a correlation with the nucleotide corresponding to the amino acid at position ii (N at position 31 of SEQ ID NO: 19) (Non-Patent Document 8 cited above). Also, note that the C or U four positions to the right of the target base of the S2 motif is the editing target base by the DYW domain, and it can be incorporated into the PLS-type PPR protein.
[0105] In the E1 motif, the correlation with nucleotides is only observed in the amino acid at the fourth position (G at the fourth position of SEQ ID NO: 20) (Ruwe et al. (2019) New Phytol. 222 218-229.).
[0106] The fourth (V at the fourth position of SEQ ID NO: 21) and the last (K at the 33rd position of SEQ ID NO: 21) amino acids in the E2 motif are highly conserved and are not involved in specific PPR-RNA recognition (Non-Patent Document 2 cited above).
[0107] The number of repeats of P1L1S1 is not particularly limited as long as it can bind to the target base sequence. For example, it is 1 to 5, preferably 2 to 4, and more preferably 3. In principle, even 1 unit (3 pieces) can be used. MEF8 (L1-S1-P2-L2-S2-E-DYW) composed of 5 PPR motifs is known to be involved in about 60 editing sites.
[0108] In natural PPR proteins, P1L1S1 located at the beginning and the end shows a clear difference in amino acid residues at specific positions and is different from the internal P1L1S1. From the viewpoint of designing an artificial PLS array as close as possible to the natural one, it is advisable to divide it into P1L1S1 located at the first (N-terminal side), internal P1L1S1, and the last (C-terminal side) P1L1S1 located immediately before P2L2S2, and design three types of P1L1S1 according to the position of the PPR motif. In addition to those composed of the repeating unit of P-L-S in nature, there are cases where SS (31 amino acids) is repeated, and such cases can also be used in the present invention.
[0109] [DYW protein] The present invention provides a DYW protein for editing a target RNA, which includes an RNA binding domain containing at least one PPR motif and capable of specifically binding to the target RNA according to the rules of the PPR-code, and a DYW domain which is any one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.
[0110] Such DYW proteins can be artificial. Artificial means not a natural product but a product synthesized artificially. Being artificial applies, for example, when having a DYW domain with a sequence not found in nature, when having a PPR binding domain with a sequence not found in nature, when having a combination of an RNA binding domain and a DYW domain not found in nature, or when a protein for RNA editing in animal cells has additional parts not present in natural DYW proteins derived from plants, such as a nuclear localization signal or a human mitochondrial localization signal. Examples of nuclear localization signal sequences include PKKKRKV (SEQ ID NO:32) derived from SV40 large T antigen, and KRPAATKKAGQAKKKK (SEQ ID NO:33) which is the NLS of nucleoplasmin.
[0111] The number of PPR motifs can be appropriately determined according to the sequence of the target RNA. The number of PPR motifs may be at least one, and may also be two or more. It is known that two PPR motifs can bind to RNA (Nucleic Acids Research, 2012, Vol. 40, No. 6, 2712 - 2723).
[0112] One preferred embodiment of the DYW protein is as follows. An artificial DYW protein for editing a target RNA, which includes an RNA binding domain containing at least one, preferably 2 - 25, more preferably 5 - 20, and even more preferably 10 - 18 PPR motifs and capable of specifically binding to the target RNA according to the rules of the PPR-code, and a DYW domain which is any one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.
[0113] One of the preferred embodiments of other DYW proteins is as follows. An RNA-binding domain (preferably an RNA-binding domain of a PLS-type PPR protein) that contains at least 1, preferably 2 to 25, more preferably 5 to 20, and even more preferably 10 to 18 PPR motifs and can bind specifically to a target RNA of an animal according to the rules of the PPR-code, and a DYW domain that is any one of DYW:PG, DYW:WW, or DYW:KP described above, a DYW protein for editing a target RNA.
[0114] The DYW domain, the PPR protein that is an RNA-binding domain, and the DYW protein of the present invention can be prepared in a relatively large amount by methods well known to those skilled in the art. Such methods may include determining the nucleic acid sequence encoding it from the amino acid sequence of the target domain or protein, cloning it, and creating a transformant that produces the target domain or protein.
[0115] [Use of DYW Protein] (Nucleic acid, vector, cell encoding DYW protein, etc.) The present invention also provides the above-described PPR motif, DYW protein, or nucleic acid encoding a DYW protein, a vector containing the nucleic acid (for example, a vector for amplification, an expression vector). The vector also includes a viral vector. The vector for amplification can use Escherichia coli or yeast as a host. In the present specification, the expression vector means, for example, a vector containing DNA having a promoter sequence, DNA encoding a desired protein, and DNA having a terminator sequence from upstream, but as long as it exhibits the desired function, it is not necessarily arranged in this order. In the present invention, various vectors commonly used by those skilled in the art can be recombinantly used.
[0116] Specifically, the present invention provides a nucleotide sequence encoding a DYW protein, which comprises at least one PPR motif, an RNA binding domain (preferably an RNA binding domain which is a PLS type PPR protein) capable of specifically binding to a target RNA (preferably a target RNA of an animal) according to the rules of PPR-code, and a DYW domain which is any one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.
[0117] Specifically, the present invention also provides a vector for editing a target RNA, which comprises a nucleotide sequence encoding a DYW protein, which comprises at least one PPR motif, an RNA binding domain (preferably an RNA binding domain which is a PLS type PPR protein) capable of specifically binding to a target RNA (preferably a target RNA of an animal) according to the rules of PPR-code, and a DYW domain which is any one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.
[0118] The DYW protein of the present invention can function in cells of eukaryotes (such as animals, plants, microorganisms (yeast, etc.), protists). The DYW protein of the present invention can particularly function in animal cells (in vitro or in vivo). Examples of animal cells into which the DYW protein of the present invention or a vector expressing the DYW protein can be introduced include cells derived from humans, monkeys, pigs, cows, horses, dogs, cats, mice, and rats. Examples of cultured cells into which the DYW protein of the present invention or a vector expressing the DYW protein can be introduced include, but are not limited to, Chinese hamster ovary (CHO) cells, COS-1 cells, COS-7 cells, VERO (ATCC CCL-81) cells, BHK cells, dog kidney-derived MDCK cells, hamster AV-12-664 cells, HeLa cells, WI38 cells, HEK293 cells, HEK293T cells, and PER.C6 cells.
[0119] (Use) The DYW protein of the present invention can convert the editing target C contained in the target RNA to U, or the editing target U to C. RNA-binding PPR proteins are involved in all steps of RNA processing, cleavage, RNA editing, translation, splicing, and RNA stabilization found in organelles.
[0120] In addition, the DYW protein of the present invention can perform mitochondrial RNA single-base editing. Mitochondria have their own genome, and the constituent proteins of important complexes involved in respiration and ATP production are encoded. It is known that various diseases are caused by their mutations. Treatment of various diseases can be expected by mutation repair using the present invention.
[0121] On the other hand, the CRISPR-Cas system has been developed as a C-to-U RNA editing tool in the cytoplasm by fusing the Cas protein to a modified ADAR domain (Abudayyeh et al., 2019). However, the CRISPR-Cas system is composed of a protein and a guide RNA, and efficient mitochondrial transport of the guide RNA is difficult. On the other hand, a PPR protein can perform RNA editing with a single molecule, and generally, a protein can be delivered to mitochondria by fusing a mitochondrial localization signal sequence to the N-terminal side. Therefore, in order to confirm whether this technology can be used for mitochondrial RNA editing, PPRs targeting MT-ND2 and MT-ND5 were designed, and genes fused with the PG or WW domain were created (Figure 10a).
[0122] These proteins consisting of a mitochondrial targeting sequence (MTS) and a PPR-P sequence target the third positions of codons 178 and 301 of MT-ND2 and MT-ND5 so as not to adversely affect HEK293T cells (Figure 10a). After introduction of the plasmid into HEK293T cells, editing was confirmed within the mRNAs of MT-ND2 and MT-ND5 (Figures 10b and c). Editing activities of up to 70% were detected for these four proteins against the target, and no off-target mutations were detected for the same mRNA molecules (Figures 10b and c).
[0123] Therefore, the RNA base editing method provided by the present invention can be expected to be used in various fields as follows.
[0124] (1) Medicine · Recognize and edit specific RNAs associated with specific diseases. By using the present invention, genetic diseases caused by single-base mutations can be treated. Many mutation directions in genetic diseases are mutations from C to U. Therefore, the method of the present invention that can convert U to C in particular may be useful.
[0125] · Create cells with controlled RNA suppression and expression. Such cells include stem cells (e.g., iPS cells) that monitor the differentiated / undifferentiated state, model cells for evaluating cosmetics, and cells that can turn on / off the expression of functional RNAs for the purpose of elucidating the mechanism of drug discovery and pharmacological tests.
[0126] (2) Agriculture, forestry, and fisheries · Improve yields and quality in crops, forest products, fishery products, etc. · Breed organisms with improved disease resistance, improved environmental tolerance, and improved or new functionality.
[0127] For example, regarding first-generation hybrid (F1) crops, artificial creation of F1 crops by editing mitochondrial RNA with DYW protein may be able to improve yield and quality. RNA editing with DYW protein enables variety improvement and breeding (genetically improving organisms) more accurately and rapidly than prior art. Also, since RNA editing with DYW protein does not transform traits with foreign genes like genetic recombination, it can be said to be close to traditional breeding methods such as mutant selection and backcrossing. Therefore, it can surely and rapidly respond to global food and environmental problems.
[0128] (3) Chemistry · In the production of useful substances using microorganisms, cultured cells, plants, and animals (e.g., insects), protein expression levels are controlled by RNA manipulation. Thereby, the productivity of useful substances can be improved. Examples of useful substances include proteinaceous substances such as antibodies, vaccines, and enzymes, as well as relatively low-molecular-weight compounds such as pharmaceutical intermediates, fragrances, and pigments.
[0129] · By modifying the metabolic pathways of algae and microorganisms, the production efficiency of biofuels is improved.
[0130] [Terms] The numerical range x~y includes the values x and y at both ends, unless otherwise specified.
[0131] Regarding the amino acid sequences of proteins and polypeptides, amino acid residues may simply be referred to as amino acids.
[0132] Regarding a base sequence (sometimes referred to as a nucleotide sequence) or an amino acid sequence, "identity" means, unless otherwise specified, the percentage of the number of matching bases or amino acids shared between two sequences when the two sequences are aligned in an optimal manner. That is, identity = (the number of matching positions / the total number of positions) × 100, and it can be calculated using commercially available algorithms. Such algorithms are incorporated into the NBLAST and XBLAST programs described in Altschul et al., J. Mol. Biol. 215 (1990) 403 - 410. More specifically, searches and analyses regarding the identity of a base sequence or an amino acid sequence can be performed by algorithms or programs well-known to those skilled in the art (for example, BLASTN, BLASTP, BLASTX, ClustalW). When using a program, parameters can be appropriately set by those skilled in the art, or the default parameters of each program can also be used. Specific methods of these analysis methods are also well-known to those skilled in the art.
[0133] Regarding a base sequence or an amino acid sequence, unless otherwise specified, higher sequence identity is preferred. Specifically, it is preferably 40% or more, more preferably 45% or more, still more preferably 50% or more, still more preferably 55% or more, still more preferably 60% or more, still more preferably 65% or more. Also, it is preferably 70% or more, more preferably 80% or more, still more preferably 85% or more, still more preferably 90% or more, still more preferably 95% or more, and still more preferably 97.5% or more.
[0134] Regarding a polypeptide or protein, unless otherwise specified, the number of amino acids to be substituted, deleted, or added when referring to a "substituted, deleted, or added sequence" is not particularly limited as long as the motif or protein consisting of the amino acid sequence has the desired function, regardless of the motif or protein. It is about 1 to 9 or 1 to 4, or there can be a larger number of substitutions or the like if the substitution is to an amino acid with similar properties. Means for preparing polynucleotides or proteins related to such amino acid sequences are well known to those skilled in the art.
[0135] Amino acids with similar properties refer to amino acids with similar physical properties such as hydrophobicity, charge, pKa, solubility, etc., and for example, include the following. Hydrophobic (nonpolar) amino acids: alanine, valine, glycine, isoleucine, leucine, phenylalanine, proline, tryptophan, tyrosine Non-hydrophobic amino acids: arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine, cysteine, histidine, methionine; Hydrophilic amino acids: arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine; Acidic amino acids: aspartic acid, glutamic acid; Basic amino acids: lysine, arginine, histidine; Neutral amino acids: alanine, asparagine, cysteine, glutamine, glycine, isoleucine, leucine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine; Sulfur-containing amino acids: methionine, cysteine; Aromatic ring-containing amino acids: tyrosine, tryptophan, phenylalanine.
Examples
[0136] [1. Design of the DYW domain and measurement of RNA editing activity in E. coli] [Results] (Design of the DYW domain) Since most of the currently available plant genomic information is that of angiosperms, the development of artificial DYW proteins is limited to the DYW:PG type. On the other hand, although the transcriptome data is a partial sequence, it also includes information on early land plant species with both C-to-U and U-to-C RNA editing. Therefore, we designed a complete artificial DYW domain using a database of PPR proteins (Non-Patent Document 11 cited above) constructed based on the transcriptome dataset created by the 1000 Plants (1KP) International Consortium.
[0137] In designing the artificial DYW domain, in addition to the cytidine deaminase-like DYW domain, PPR (P2 and L2) motifs and PPR-like (S2, E1, and E2) motifs were added to the design. This is because the importance of these motifs in RNA editing activity is not yet known at present. We constructed a phylogenetic tree of the DYW domains of mosses, ferns, and fern allies and classified them into three groups: DYW:PG, DYW:WW, and DYW:KP. Then, a new phylogenetic tree was created for each of these groups of DYW domains, and a group of proteins available for the design of the artificial DYW domain was selected (Figure 1). Since the DYW:PG group had a large difference in sequences, it was difficult to select a group of proteins for use in the design (Figure 1a). Since the sequences of fern proteins were rich in diversity and there were many of them, we decided to focus on this protein (PG1). For the design of the DYW:WW domain, a group of proteins with short branches found only in mosses was selected (Figure 1b, WW1). Since this group of branches is short, it is presumed that the gene mutations between proteins are small. And for the design of the DYW:KP domain, we focused on a group of DYW domains specific to fern allies. The sequences of the proteins included in this group have large amino acid mutations but the lengths are conserved (Figure 1c, KP1).
[0138] (Design of RNA Binding Domain) We fused each DYW domain with an artificial P or PLS array designed based on the PPR motif sequences identified from plant genomic information (Non-Patent Document 2 cited above) (Fig. 2a). The PPR motif of the P array was constructed based on the consensus sequence obtained from the alignment of 35-amino acid P motifs and some amino acids were replaced to enhance RNA recognition. For the design of the PLS array, motifs of standard lengths of 35 amino acids (P1 and L1) and 31 amino acids (S1) were selected. In natural PPR proteins, the P1L1S1 located at the first and last positions show distinct differences in amino acid residues at specific positions and are different from the internal P1L1S1. To design an artificial PLS array as close as possible to the natural one, it was divided into three types of P1L1S1 according to the position of the PPR motif: the P1L1S1 located at the first (N-terminal side), the internal P1L1S1, and the last (C-terminal side) P1L1S1 located immediately before P2L2S2. These proteins will be named based on their structures in the future. For example, the one with the DYW:WW domain fused to the P array will be designated as P-DYW:WW.
[0139] (Artificial DYW proteins can specifically edit target sequences) In Arabidopsis thaliana, the PPR protein CLB19 recognizes RNA editing sites on chloroplast rpoA and clpP RNAs. We decided to design a PPR protein targeting the rpoA editing site. The fourth and ii-th amino acids in the P1, L1, S1, and P2 motifs followed the PPR code. On the other hand, since the PPR code for the C-terminal PPR-like motifs (L2, S2, E1, E2) is unknown, the fourth and ii-th amino acids of CLB19 were used for the L2, S2, E1, and E2 motifs (Fig. 2b).
[0140] The gene region encoding the recombinant artificial DYW protein was cloned into an expression vector, and a target sequence was added downstream of its stop codon. Based on previous studies on two types of PPR proteins (PPR56 and PPR65) in *Marchantia polymorpha* (Non-Patent Document 12 cited above), a method for testing the RNA editing activity of the designed PPR proteins in *Escherichia coli* was developed. While no editing activity was observed for the designed PLS-DYW:PG1 and PLS-DYW:WW1 against DNA, cytidine was substituted with uridine at an editing efficiency of over 90% against RNA (Figs. 3a, b, d). When a P array was used instead of the PLS domain, the editing efficiency decreased by 10 - 40%. On the other hand, U-to-C RNA editing activity was observed for both P-DYW:KP and PLS-DYW:KP (Figs. 3c, e). However, the editing activity was lower compared to DYW:PG and DYW:WW.
[0141] [Method] [Phylogenetic tree] The P2-L2-S2-E1-E2-DYW region containing the DYW domain (minimum 132 amino acids) was extracted from the PPR database (https: / / ppr.plantenergy.uwa.edu.au / onekp / ; Non-Patent Document 11 cited above). The alignment of the obtained sequences was created using MAFFT L-INS-i (v7.407 automatic mode) (K. Katoh, D. M. Standley (2013). Mol Biol Evol. 30(4): 772 - 780.), and then trimmed using trimAl (v.1.4.rev15) (Salvador Capella-Gutierrez, et al. (2009). Bioinformatics. 25(15): 1972 - 1973.). The parameters for trimming were set to gt 0.2 cons 20. Sequences with mutations in the active site (HxEx n CxxCH) were removed from the alignment (Fig. 4).
[0142] A phylogenetic tree was constructed from the remaining sequences using FastTree (v2.1.10) (Price, M. N., et al. (2010). PLoS One 5:e9490.), and DYW:PG, DYW:WW, and DYW:KP were identified. The parameters at this time were wag and cat 8.
[0143] (Cloning of Trx-PPR-DYW proteins and target sequences) The consensus sequences of each DYW domain (including P2, L2, S2, E1, and E2) were designed using EMBOSS:cons (v.6.6.0.0). The protein expression vector was modified from pET21b+PA to remove the original Esp3I and BpiI restriction enzyme sites and add two Esp3I sites as cloning sites. The gene was divided into four sections (Trx, PPR array, DYW domain, and RNA editing site) and constructed using a two-step Golden Gate method. First, three parts, 1) the thioredoxin-6×His-TEV gene region (including the BpiI restriction enzyme site at the 3'), 2) the P2-L2-S2-E1-E2-DYW gene region (including the BpiI restriction enzyme site at the 5'), and 3) the coding sequence region of the RNA editing site, were cloned into the Esp3I site of the modified pET21b vector. Then, the full-length PLS domain or P domain was cloned into the BpiI site to produce the PLS:DYW (SEQ ID NOs: 35-37) or P:DYW (SEQ ID NOs: 38-40) proteins.
[0144] (Measurement of RNA editing activity in Escherichia coli) To analyze the RNA editing activity of the recombinant protein in Escherichia coli, we modified the protocol developed by Oldenkott et al. (supra Non-Patent Document 12). The plasmid DNA prepared above was introduced into Escherichia coli Rosetta 2 strain and cultured overnight at 37°C in 1 mL of LB medium (containing 50 μg / mL carbenicillin and 17 μg / mL chloramphenicol). 5 mL of LB medium containing the appropriate antibiotics was prepared in a deep-well 24-well plate, and 100 μL of the preculture was seeded here. This culture solution was used to measure the absorbance (OD 600Until it reached 0.4 - 0.6, it was grown at 37°C and 200 rpm, and then the plate was cooled at 4°C for 10 - 15 minutes. Next, 0.4 mM of ZnSO4 and 0.4 mM of IPTG were added, and it was further cultured at 16°C and 180 rpm for 18 hours. 750 μL of the culture solution was collected and centrifuged, and then the cell pellet was frozen in liquid nitrogen and stored at -80°C.
[0145] The frozen cell mass was resuspended in 200 μL of 1 - thioglycerol / homogenization solution, sonicated at 40 W for 10 seconds to separate the cells, and then 200 μL of lysis buffer was added. RNA was extracted using the Maxwell® RSC simplyRNA Tissue Kit (Promega). The RNA was treated with DNaseI (Takara Bio), and cDNA was synthesized using SuperScript® III Reverse Transcriptase (Invitrogen) with 1 μg of the treated RNA and 1.25 μM of random primers (6 mer). The region containing the editing site was amplified using NEBNext High - Fidelity 2x PCR Master Mix (New England Biolabs), 1 μL of cDNA, and primers for thioredoxin and the T7 terminator sequence. The PCR product was purified using NucleoSpin® Gel and PCR Cleanup (Takara Bio), and sequence analysis was performed using a forward primer specific to the DYW domain sequence to determine the bases at the RNA editing site. For measuring the RNA editing efficiency, the ratio of the peak heights of the waveforms of C and U at the editing site was used, and the C - to - U RNA editing efficiency was calculated as U / (C + U)×100, and the U - to - C editing efficiency was calculated as C / (C + U)×100. The independent experiment was repeated 3 times.
[0146] [List of cited sequences]
[0147]
Table 3 - 1
[0148]
Table 3 - 2
[0149] [Table 3-3]
[0150] [2. Examples in animal cultured cells] [Results] A gene in which PLS-type PPR was fused with each DYW domain (PG1 (SEQ ID NO: 1), WW1 (SEQ ID NO: 2), KP1 (SEQ ID NO: 3)), and a plasmid containing the target sequence were transfected into HEK293T cells, and RNA was recovered after culturing. The conversion efficiency from cytidine (C) to uridine (U), or from uridine (U) to cytidine (C) at the target site was analyzed by Sanger sequencing (Figure 5). When the PG1 or WW1 domain was fused, there was more than 90% C to U activity, while no U to C activity was detected (Figure 5a, c). When the KP1 domain was fused, 25% U to C activity was detected, but about 10% C to U activity was also detected (Figure 5b, c). From these results, it was found that the editing enzyme also functions in animal cultured cells.
[0151] [Methods] (Preparation of PPR expression plasmid for animal cultured cell test) From the plasmid used in Figure 3, the gene sequence of 6xHis-PPR-DYW protein (the same protein as used in the experiment in E. coli (SEQ ID NOs: 35 to 37)) and the region containing the editing site (SEQ ID NO: 34) were amplified by PCR and cloned into an animal cell expression vector by the Golden Gate Assembly method. The vector is under the control of a promoter containing a CMV promoter and a human β-globin chimeric intron, and PPR is expressed, and a poly A signal is added by the SV40 polyadenylation signal.
[0152] (HEK293T cell culture) HEK293T cells were cultured at 37°C and 5% CO2 using Dulbecco's Modified Eagle Medium (DMEM) supplemented with high glucose, glutamine, phenol - RED, sodium pyruvate (FUJIFILM Wako Pure Chemical Corporation), 10% fetal bovine serum (Capricom), and 1% penicillin - streptomycin (FUJIFILM Wako Pure Chemical Corporation). The cells were passaged every 2 - 3 days when they reached 80 - 90% confluence.
[0153] (Transfection) For the RNA editing assay, approximately 8.0 x 10 4 HEK293T cells were seeded into each well of a 24 - well flat - bottom cell culture plate and cultured at 37°C and 5% CO2 for 24 hours. To each well, 500 ng of plasmid was added with 18.5 μl of Opti - MEM® I Reduced Serum Medium (ThermoFisher) and 1.5 μl of FuGENE® HD Transfection Reagent (Promega) to a final volume of 25 μl. The mixture was incubated at room temperature for 10 minutes before adding to the cells. The cells were harvested 24 hours after transfection.
[0154] (RNA Extraction, Reverse Transcription, Sequencing) The procedure was the same as that in the above - mentioned E. coli assay. In the following examples, unless otherwise specified, experiments were conducted in the same manner as the E. coli assay and this example.
[0155] [3. Effect on RNA Editing Activity by Domain Swap] The DYW domain can be divided into several regions based on the conservation of the amino acid sequence, but the relationship with these RNA editing activities is not known. Here, a part of the KP1 and WW1 domains was swapped, and the effect on RNA editing activity was examined in HEK239T cells (Figure 6). In the chimKP1a (SEQ ID NO: 90) in which the PG box and DYW of the WW1 domain were fused to the central part containing the Active site of the KP1 domain, a U to C activity nearly 50% higher than that of the KP1 domain was observed, and it was found that the C to U activity almost disappeared. By domain swapping, we succeeded in improving the U to C editing performance of the KP domain.
[0156] [4. Improvement of RNA editing activity of the KP domain by mutagenesis] Various mutations were introduced into the KP domain with the aim of improving the RNA editing activity of the KP domain. KP2 - KP23 (SEQ ID NO: 68 - 89) were designed, and the C to U or U to C RNA editing activities were examined in E. coli (Figure 7a) and HEK293T cells (Figure 7b, c). KP22 (SEQ ID NO: 88) had the highest editing activity for U to C and the lowest editing activity for C to U, and we succeeded in improving the RNA editing activity compared to KP1.
[0157] [5. Improvement of RNA editing activity of the PG domain by mutagenesis] Various mutations were introduced into the PG domain with the aim of improving the RNA editing activity of the PG domain. PG2 - PG13 (SEQ ID NO: 41 - 53) were designed, and the C to U RNA editing activity was examined in E. coli (Figure 8a) and HEK293T cells (Figure 8b, c). PG11 (SEQ ID NO: 50) had the highest editing activity for C to U, and we succeeded in improving the RNA editing activity.
[0158] [6. Improvement of RNA editing activity of the WW domain by mutagenesis] Various mutations were introduced into the WW domain with the aim of improving the RNA editing activity of the WW domain. WW2 to WW14 were designed, and their C to U RNA editing activities were examined in E. coli (Figure 9a) and HEK293T cells (Figures 9b and c). WW11 (SEQ ID NO: 63) had the highest editing activity for C to U and successfully improved the RNA editing activity compared to WW1.
[0159] [7. Human mitochondrial RNA editing using PPR proteins] [Results] Mitochondria have their own genomes, and the constituent proteins of important complexes involved in respiration and ATP production are encoded. It is known that various diseases are caused by these mutations, and methods for repairing mutations are required.
[0160] The CRISPR-Cas system has been developed as a C-to-U RNA editing tool in the cytoplasm by fusing the Cas protein with a modified ADAR domain (Abudayyeh et al. 2019 Science Vol.365, Issue 6451, pp.382-386). However, the CRISPR-Cas system is composed of a protein and a guide RNA, and efficient mitochondrial transport of the guide RNA is difficult. On the other hand, PPR proteins can perform RNA editing with a single molecule, and generally, proteins can be delivered to mitochondria by fusing a mitochondrial localization signal sequence to the N-terminal side. Therefore, in order to confirm whether this technology can be used for mitochondrial RNA editing, PPRs targeting MT-ND2 and MT-ND5 were designed, and genes fused with the PG1 or WW1 domain were created (Figure 10a).
[0161] These proteins consisting of a mitochondrial targeting sequence (MTS) and a P-DYW sequence target the third positions of codons 178 and 301 of MT-ND2 and MT-ND5 so as not to have an adverse effect on HEK293T cells (Figure 10a). After introduction into HEK293T cells with a plasmid, editing was confirmed in MT-ND2 and MT-ND5 mRNAs (Figure 10b, c). Editing activities of up to 70% were detected for these four proteins against the target, and no off-target mutations were detected for the same mRNA molecules (Figure 10b, c).
[0162] [Method] (Cloning for mitochondrial editing) The mitochondrial targeting sequence from the LOC100282174 protein of Zea mays (Chin et al. 2018), ten PPR-P and PPR-like motifs, and the DYW domain portion (the DYW domain uses PG1 and WW1) were cloned into an expression plasmid under the control of the CMV promoter by Golden Gate Assembly (SEQ ID NOs: 91-94).
[0163] [List of cited sequences]
[0164] [Table 4-1]
[0165] [Table 4-2]
[0166] [Table 4-3]
[0167] [Table 4-4]
Claims
1. A nucleic acid encoding a DYW protein, wherein the DYW protein has an RNA binding domain containing at least one PPR motif and capable of specifically binding to a target RNA, and a DYW domain consisting of any one of the following polypeptides The nucleic acid is a DYW protein containing. ・A polypeptide consisting of the sequence of SEQ ID NO: 1; a polypeptide having at least 90% sequence identity with the sequence of SEQ ID NO: 1 and having C-to-U editing activity; a polypeptide having a sequence in which 1 to 13 amino acids in the sequence of SEQ ID NO: 1 are substituted, deleted or added and having C-to-U editing activity; or a polypeptide consisting of the sequence of SEQ ID NO: 41 ・A polypeptide consisting of the sequence of SEQ ID NO: 2; a polypeptide having at least 90% sequence identity with the sequence of SEQ ID NO: 2 and having C-to-U editing activity; a polypeptide having a sequence in which 1 to 12 amino acids in the sequence of SEQ ID NO: 2 are substituted, deleted or added and having C-to-U editing activity; or a polypeptide consisting of the sequence of SEQ ID NO: 54
2. The nucleic acid according to claim 1, wherein the DYW domain consists of any one of the following polypeptides. ・A polypeptide consisting of the sequence of SEQ ID NO: 1 ・A polypeptide consisting of the sequence of SEQ ID NO: 2
3. The nucleic acid according to claim 1 or 2, wherein the RNA binding domain is of the PLS type.
4. A vector comprising the nucleic acid according to any one of claims 1 to 3.
5. A cell (excluding human individuals) comprising the vector according to claim 4.
6. A composition for editing RNA in eukaryotic cells, comprising the nucleic acid according to any one of claims 1 to 3, or the vector according to claim 4.
Citation Information
Patent Citations
Design method for RNA-binding protein using PPR motif, and use thereof
WO2013058404A1
DNA binding protein using PPR motif, and use thereof
WO2014175284A1