Methods for editing target RNA

By employing DYW domains from PPR-P or PLS arrays, the method addresses the challenge of converting cytidine to uridine or vice versa in RNA, achieving efficient RNA editing for applications in gene therapy and mutation introduction.

JP7897387B2Active Publication Date: 2026-07-29EDITFORCE INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
EDITFORCE INC
Filing Date
2025-06-06
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Current methods lack the ability to design molecules that can convert cytidine (C) to uridine (U) or vice versa in any target RNA using a DYW domain and an artificial RNA-binding protein, particularly in plant groups that exhibit U-to-C RNA editing.

Method used

The use of a modular DYW domain from PPR-P or PLS arrays, specifically designed DYW:PG, DYW:WW, and DYW:KP domains, to achieve C-to-U or U-to-C RNA editing by fusing them with PPR-P or PLS arrays, enabling sequence-specific binding to target RNA.

Benefits of technology

This approach allows for the conversion of editing target C to U or U to C in target RNA with at least 3-5% efficiency, facilitating applications in gene therapy and gene mutation introduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007897387000014
    Figure 0007897387000014
  • Figure 0007897387000015
    Figure 0007897387000015
  • Figure 0007897387000016
    Figure 0007897387000016
Patent Text Reader

Abstract

To provide a method for converting an editing target C contained in a target RNA to U or an editing target U contained in a target RNA to C.SOLUTION: A method for editing a target RNA is provided, comprising applying, to the target RNA, an artificial DYW protein containing a DYW domain consisting of any one of the following polypeptides a, b, c, and bc: a. a polypeptide having xa1 PGxa2SWIExa3-xa16HP...HxaaE...Cxa17xa18CH...DYW, having a sequence identity of at least 40% to the sequence of SEQ ID NO: 1 and having a C-to-U / U-to-C editing activity; b. a polypeptide having xb1 PGxb2SWWTDxb3-xb16HP...HxbbE...Cxb17xb18CH...DYW, having a sequence identity of at least 40% to the sequence of SEQ ID NO: 2 and having a C-to-U / U-to-C editing activity; c. a polypeptide having KPAxc1Axc2IExc3...HxccE...Cxc4xc5CH...xc6xc7xc8, having a sequence identity of at least 40% to the sequence of SEQ ID NO: 3 and having a C-to-U / U-to-C editing activity; bc. a polypeptide having xb1 PGxb2SWWTDxb3-xb16HP...HxccE...Cxc4xc5CH...Dxbc1xbc2, having a sequence identity of at least 40% to a 90-th sequence of SEQ ID NO: 2 and having a C-to-U / U-to-C editing activity (in the sequences, x represents an arbitrary amino acid, and ... represents an arbitrary polypeptide fragment).SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to RNA editing technology using a protein capable of binding to target RNA. This invention is useful in a wide range of fields, including medicine (drug discovery support, treatment), agriculture (agricultural, fishery, and livestock product production, breeding), and chemistry (biological substance production). [Background technology]

[0002] In plant mitochondria and chloroplasts, RNA editing, in which specific bases in the genome are replaced at the RNA level, occurs frequently, and it is known that pentatricopeptide repeat (PPR) proteins, which are RNA-binding proteins, are involved in this phenomenon.

[0003] PPR proteins are classified into two families, P and PLS, based on the structure of the PPR motifs that make up the protein (Non-Patent Literature 1). While P-class PPR proteins consist of simple repeats of a standard 35-amino acid PPR motif (P), PLS proteins contain two similar motifs called L and S in addition to P. The PPR array (arrangement of PPR motifs) of a PLS protein consists of three PPR motifs, P1 (approximately 35 amino acids), L1 (approximately 35 amino acids), and S1 (approximately 31 amino acids), which form the PLS repeating unit. At the C-terminus of this P1L1S1, there are PLS motifs with slightly different sequences, namely P2 (35 amino acids), L2 (36 amino acids), and S2 (32 amino acids). In addition to those composed of PLS ​​repeating units, there are also cases where SS (31 amino acids) is repeated. Furthermore, this final P2L2S2 motif may be followed by two PPR-like motifs called E1 and E2 at its C-terminus, and a DYW domain with a 136-amino acid cytidine deaminase domain-like sequence (Non-Patent Literature 2).

[0004] The interaction between PPR proteins and RNA is defined by a PPR code that specifies the binding RNA base through a combination of several amino acids within each PPR motif, and these amino acids bind to the corresponding nucleotide via hydrogen bonds (Patent Documents 1, 2, 3, 4, 5, 6, 7, and 8). [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] International release WO2013 / 058404 [Patent Document 2] International release WO2014 / 175284 [Non-patent literature]

[0006] [Non-Patent Document 1] Lurin, C., Andres, C., Aubourg, S., Bellaoui, M., Bitton, F., Bruyere, C., Caboche, M., Debast, C., Gualberto, J., Hoffmann, B., et al. (2004). Genome-wide analysis of Arabidopsis pentatricopeptide repeat proteins reveals their essential role in organelle biogenesis. Plant Cell 16:2089-2103. [Non-Patent Document 2] Cheng , S. , Gutmann , B. , Zhong , X. , Ye , Y. , Fisher , MF , Bai , F. , Castleden , I. , Song , Y. , Song , B. , Huang , J. , et al. (2016). Redefining the structural motifs that determine RNA binding and RNA editing by pentatricopeptide repeat proteins in land plants. Plant J. 85:532–547.

Outdoor Tools3

Outdoor Tools 4

Direct Environment 5

Outdoor Configuration6

Direct Environment 7

Outdoor Track 8

Outdoor Tools9

[0007] In terrestrial plant organelles, C-to-U RNA editing, in which the RNA base is substituted from cytidine to uridine, commonly occurs. However, in hornworts, and some phyllum and ferns, U-to-C RNA editing, in which uridine is substituted to cytidine, is also observed (Non-Patent Literature 9). Two different bioinformatics studies have revealed unique DYW domains with sequences different from standard DYW domains in plants exhibiting U-to-C RNA editing (Non-Patent Literature 10, Non-Patent Literature 11).

[0008] Two types of DYW:PG proteins derived from Physcomitrella patens have been reported to exhibit C-to-U RNA editing activity in E. coli (Non-Patent Literature 12). In plant groups that diverged early in the evolution of land plants, DYW domains can be broadly classified into two groups. The first is called the DYW:PG / WW group, which includes the standard DYW domain, the DYW:PG type, and the DYW:WW type, in which one tryptophan (W) molecule is added to the PG box. The second group is called DYW:KP, and this DYW domain has a different sequence of three amino acids in the PG box and the C-terminus from the PG / WW group, and is also found only in plants that possess U-to-C RNA editing. The technology to change a single base in any RNA sequence to a specific base is useful in gene therapy and gene mutation introduction technology for industrial use. Although various RNA-binding molecules have been developed to date, a method for designing molecules that convert cytidine (C) to uridine (U) or vice versa of any target RNA using a DYW domain and an artificial RNA-binding protein has not yet been established. [Means for solving the problem]

[0009] In this study, we demonstrate that the portion of the PPR-DYW protein containing the C-terminal DYW domain can be used as a modular editing domain for RNA editing. When the three DYW domains designed in this study were fused with PPR-P or PLS arrays, the DYW:PG and DYW:WW domains showed C-to-U RNA editing activity, while DYW:KP showed U-to-C RNA editing activity.

[0010] The present invention provides the following: [1] A method for editing target RNA, comprising applying an artificial DYW protein to the target RNA, the DYW domain comprising one polypeptide from a, b, c, and bc below. a. x a1 PGx a2 SWISE a3 -xa16 HP … Hx aa E … Cx a17 x a18 A polypeptide having CH … DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity b. x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx bb E … Cx b17 x b18 A polypeptide having CH … DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity c. KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH … x c6 x c7 x c8 Having, having at least 40% sequence identity with the sequence of SEQ ID NO: 3, and having C-to-U / U-to-C editing activity bc. x b1 PGx b2 SWWTDx b3 -x b16 HP … Hx cc E … Cx c4 x c5 CH … Dx bc1 x bc2 Having, having at least 40% sequence identity with the 90 sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity (In the sequence, x represents any amino acid, and … represents any polypeptide fragment.) [2] A DYW domain consisting of any one of the following polypeptides a, b, c, and bc a. x a1 PGx a2 SWIEx a3 -x a16 HP … Hx aa E … Cxa17 x a18 A polypeptide having CH…DYW, possessing at least 40% sequence identity with the sequence of SEQ ID NO: 1, and exhibiting C-to-U / U-to-C editing activity. b. x b1 PGx b2 SWWTDx b3 -x b16 HP… Hx bb E … Cx b17 x b18 A polypeptide having CH … DYW, possessing at least 40% sequence identity with the sequence of SEQ ID NO: 2, and exhibiting C-to-U / U-to-C editing activity. c. KPAx c1 Ax c2 IEx c3 ... Hx cc E … Cx c4 x c5 CH… x c6 x c7 x c8 A polypeptide having a sequence identity of at least 40% with sequence number 3, and possessing C-to-U / U-to-C editing activity. bc. x b1 PGx b2 SWWTDx b3 -x b16 HP… Hx cc E … Cx c4 x c5 CH … Dx bc1 x bc2 A polypeptide having a sequence identity of at least 40% with sequence 90 of sequence number 2, and possessing C-to-U / U-to-C editing activity. [3] A DYW protein comprising an RNA-binding domain containing at least one PPR motif and capable of sequence-specifically binding to a target RNA, and the DYW domain described in 2. [4] The DYW protein described in 2, wherein the RNA-binding domain is of the PLS type. [5] A method for editing a target RNA, A method comprising the step of applying a DYW domain consisting of the following polypeptide c or bc to a target RNA to convert the editing target U to C. c. KPAx c1 Ax c2 IEx c3 ... Hx cc E … Cx c4 x c5 CH… x c6 x c7 x c8 A polypeptide having a sequence identity of at least 40% with sequence number 3, and possessing C-to-U / U-to-C editing activity. bc. x b1 PGx b2 SWWTDx b3 -x b16 HP… Hx cc E … Cx c4 x c5 CH … Dx bc1 x bc2 A polypeptide having a sequence identity of at least 40% with sequence 90 of sequence number 2, and possessing C-to-U / U-to-C editing activity. [6] The method of 5, wherein the DYW domain is fused to an RNA-binding domain that contains at least one PPR motif and is sequence-specifically bound to a target RNA according to the rules of the PPR-code. [7] A composition comprising the DYW domain described in 2 for editing RNA in eukaryotic cells. [8] A nucleic acid encoding the DYW domain described in 2, or the DYW protein described in 3 or 4. [9] A vector containing the nucleic acid described in 8.

[10] Cells (excluding human individuals) containing the vector described in 9.

[0011] [1] A method for editing target RNA, comprising applying an artificial DYW protein containing a DYW domain consisting of one polypeptide from a to c below to the target RNA. a. x a1 PGx a2 SWISE a3 -x a16 HP … HSE … Cx a17 x a18A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity b. x b1 PGx b2 SWWTDx b3 -x b16 HP…HSE…Cx b17 x b18 A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity c. KPAx c1 Ax c2 IEx c3 …HAE…Cx c4 x c5 CH…Dx c6 x c7 A polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3, and having C-to-U / U-to-C editing activity (In the sequence, x represents any amino acid, and … represents any polypeptide fragment.) [2] A DYW domain consisting of any one of the following polypeptides a to c a. x a1 PGx a2 SWIEx a3 -x a16 HP…HSE…Cx a17 x a18 A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity b.x b1 PGx b2 SWWTDx b3 -x[[ID=Z57]] b16 HP…HSE…Cx b17 x b18 A polypeptide having CH…DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity c. KPAx c1 Ax c2 IEx c3 …HAE…Cxc4 x c5 CH … Dx c6 x c7 A polypeptide having a sequence identity of at least 40% with sequence number 3, and possessing C-to-U / U-to-C editing activity. [3] A DYW protein comprising an RNA-binding domain containing at least one PPR motif and capable of sequence-specifically binding to a target RNA, and the DYW domain described in 2. [4] The DYW protein described in 2, wherein the RNA-binding domain is of the PLS type. [5] A method for editing a target RNA, A method comprising the step of applying a DYW domain consisting of the polypeptide c below to a target RNA to convert the editing target U to C. c. KPAx c1 Ax c2 IEx c3 … HAE … Cx c4 x c5 CH … Dx c6 x c7 A polypeptide having a sequence identity of at least 40% with sequence number 3, and possessing C-to-U / U-to-C editing activity. [6] The method of 5, wherein the DYW domain is fused to an RNA-binding domain that contains at least one PPR motif and is sequence-specifically bound to a target RNA according to the rules of the PPR-code. [7] A composition comprising the DYW domain described in 2 for editing RNA in eukaryotic cells. [8] A nucleic acid encoding the DYW domain described in 2, or the DYW protein described in 3 or 4. [9] A vector containing the nucleic acid described in 8.

[10] Cells (excluding human individuals) containing the vector described in 9. [Effects of the Invention]

[0012] The present invention makes it possible to convert editing target C in target RNA to U, or editing target U to C. [Brief explanation of the drawing]

[0013] [Figure 1] Approximate maximum likelihood phylogenetic trees of the C-terminal domains in DYW proteins. Phylogenetic trees for (a) DYW:PG type, (b) DYW:WW type, and (c) DYW:KP type domains were constructed using FastTree. The DYW:PG domain was designed based on DYW:PG protein derived from phylogenetic algae. The phylogenetic groups (clades) of proteins selected to design the DYW:WW and DYW:KP domains are shown by black lines. The phylogenetic trees were visualized using iTOL (Letunic, I. and Bork, P. (2016) Nucleic Acids Res. 44 W242-245), with red representing proteins from hornworts, green representing proteins from phylogenetic algae, and blue representing proteins from ferns. [Figure 2] The structure of the PPR protein designed in this study. (a) The PPR protein contains thioredoxin, a His-tag and a TEV site at the N-terminus, followed by a PPR array consisting of a P or PLS motif, and finally a DYW domain (DYW:PG, DYW:WW, or DYW:KP). The target sequence of the PPR protein (including the RNA editing site) was inserted downstream of the stop codon. (b) The 4th and 2nd amino acids involved in RNA recognition within the P, P1, L1, S1, and P2 motifs. The 4th and 2nd amino acids of L2, S2, E1, and E2 are identical to those of CLB19 (Chateigner-Boutin, AL, et al. (2008). Plant J. 56 590-602.). [Figure 3] RNA editing assay using E. coli. (a) Target base (C or T) on the DNA of each PPR-DYW. (b) C-to-U RNA editing activity of PPR-DYW:PG and WW and (c) U-to-C RNA editing activity of PPR-DYW:KP are shown. One of three independent experimental results is shown as an example. (d) Average value of C-to-U and (e) U-to-C RNA editing efficiency measured three times (except for P-DYW:KP, which was measured twice). Error bars indicate the standard deviation. [Figure 4-1]The frequency of amino acid occurrences in each selected DYW domain strain. Visualized using sequence logos created with WebLogo. The same applies to Figures 4-2 and 4-3. (a) DYW:PG [Figure 4-2] (b) DYW:WW [Figure 4-3] (c)DYW:KP [Figure 5] RNA editing activity (C to U or U to C) in HEK293T cells. Sequencing results of target sites when a PG1 or WW1 domain (a) or a KP1 domain (b) is fused to a PLS type. RNA editing activity (c) in three independent experiments. The bar graph shows the average, and each point represents the RNA editing activity in the three experiments. The left bar graph shows C to U editing activity, and the right bar graph shows U to C editing activity. [Figure 6] Improving KP domain performance through domain swapping. The RNA editing activity of Chimeric KP1a was investigated by swapping the preceding and succeeding domains with WW1 domains, while retaining the central portion containing the KP1 active site (HxExnCxxCH). Schematic diagram of domain swapping (a), actual RNA editing activity (b, c). [Figure 7] Measurement of RNA editing activity of KP domain mutants. RNA editing activity of KP1, KP2, KP3, and KP4 in E. coli (a) and RNA editing activity with HEK293T (b), and RNA editing activity of KP5-KP23 with HEK293T (c). [Figure 8] Measurement of RNA editing activity of PG domain mutants. RNA editing activity of PG1 and PG2 in E. coli (a) and RNA editing activity with HEK293T (b). RNA editing activity of PG3-PG13 with HEK293T. [Figure 9] Measurement of RNA editing activity of WW domain mutants. RNA editing activity in E. coli WW1 and WW2 (a) and RNA editing activity in HEK293T (b). RNA editing activity in HEK293T of WW3-WW14. [Figure 10] Human mitochondrial RNA editing. Target sequence information (a). RNA editing activity (b, c). [Modes for carrying out the invention]

[0014] The present invention relates to a method for editing target RNA, comprising applying a DYW protein containing a DYW domain consisting of one of the polypeptides a, b, c, and bc described below to the target RNA. a. x a1 PGx a2 SWISE a3 -x a16 HP… Hx aa E … Cx a17 x a18 A polypeptide having CH…DYW, possessing at least 40% sequence identity with the sequence of SEQ ID NO: 1, and exhibiting C-to-U / U-to-C editing activity. b. x b1 PGx b2 SWWTDx b3 -x b16 HP… Hx bb E … Cx b17 x b18 A polypeptide having CH … DYW, possessing at least 40% sequence identity with the sequence of SEQ ID NO: 2, and exhibiting C-to-U / U-to-C editing activity. c. KPAx c1 Ax c2 IEx c3 ... Hx cc E … Cx c4 x c5 CH… x c6 x c7 x c8 A polypeptide having a sequence identity of at least 40% with sequence number 3, and possessing C-to-U / U-to-C editing activity. bc. x b1 PGx b2 SWWTDx b3 -x b16 HP… Hx cc E … Cx c4 x c5 CH … Dx bc1 x bc2A polypeptide having a sequence identity of at least 40% with sequence 90 of sequence number 2, and possessing C-to-U / U-to-C editing activity.

[0015] In relation to the present invention, C-to-U / U-to-C editing activity refers to the activity that, when a target polypeptide is ligated to the C-terminal side of an RNA-binding domain that can sequence-specifically bind to the target RNA and an editing assay is performed, can convert editing target C to U, or editing target U to C, contained in the target RNA. The conversion is sufficient if, under appropriate conditions, at least about 3%, preferably about 5%, of the editing target bases are replaced with the desired bases.

[0016] [DYW domain] x a1 PGx a2 SWISE a3 -x a16 HP… Hx aa E … Cx a17 x a18 CH… DYW, x b1 PGx b2 SWWTDx b3 -x b16 HP… Hx bb E … Cx b17 x b18 CH… DYW, and KPAx c1 Ax c2 IEx c3 ... Hx cc E … Cx c4 x c5 CH… x c6 x c7 x c8 Each of these represents an amino acid sequence. In the sequence, x independently represents any amino acid, and ... independently represents a polypeptide fragment consisting of any amino acid sequence of any length. With respect to the present invention, the DYW domain can be represented by any one of these three amino acid sequences. In particular, x a1 PGx a2 SWISE a3 -x a16 HP… Hx aa E … Cx a17 x a18CH... The DYW domain consisting of DYW is DYW:PG, x b1 PGx b2 SWWTDx b3 -x b16 HP… Hx bb E … Cx b17 x b18 CH... The DYW domain consisting of DYW is referred to as DYW:WW, and KPAx c1 Ax c2 IEx c3 ... Hx cc E … Cx c4 x c5 CH… x c6 x c7 x c8 The DYW domain consisting of these elements is sometimes represented as DYW:KP.

[0017] The DYW domain consists of a region containing a PG box of approximately 15 amino acids at the N-terminus, and a central zinc-binding domain (HxEx n CxxCH, x n A DYW domain is a sequence of any number of amino acids n, and has three regions: the DYW domain, the HxE domain, and the C-terminal DYW domain. The zinc-binding domain can be further divided into the HxE domain and the CxxCH domain. These regions of each DYW domain can be represented as shown in the table below.

[0018] [Table 1]

[0019] (DYW:PG) DYW:PG is x a1 PGx a2 SWISE a3 -x a16 HP… Hx aa E … Cx a17 x a18 CH is a polypeptide consisting of DYW. Preferably, x a1 PGx a2 SWISE a3 -x a16 HP… Hx aa E … Cx a17 xa18 CH…DYW is a polypeptide having sequence identity (detailed in the [Terminology] section) with the sequence of SEQ ID NO: 1, and possessing C-to-U / U-to-C editing activity. DYW:PG has the activity to convert the editing target C to U (C-to-U editing activity). SEQ ID NO: 1 shows the sequence of DYW:PG, consisting of 136 amino acids in length, used in the experiments described in the Examples section of this specification. This sequence is novel and is disclosed for the first time in this application.

[0020] The total length of DYW:PG is not particularly limited as long as it exhibits C-to-U editing activity, but is for example 110 to 160 amino acids long, preferably 124 to 148 amino acids long, more preferably 128 to 144 amino acids long, and even more preferably 132 to 140 amino acids long.

[0021] Area containing the PG box of DYW:PG (x a1 PGx a2 SWISE a3 -x a16 On HP: x a1 The amino acid is not particularly limited as long as it exhibits C-to-U editing activity as DYW:PG, but is preferably E (glutamic acid) or an amino acid with similar properties, and more preferably G. x a2 The amino acid is not particularly limited as long as it exhibits C-to-U editing activity as DYW:PG, but is preferably C (cysteine) or an amino acid with similar properties, and more preferably C. x a3 -x a16 Each of the amino acids is not particularly limited as long as it exhibits C-to-U editing activity as DYW:PG, but is preferably the same as or similar in properties to the corresponding amino acids at positions 9 to 22 in the sequence of SEQ ID NO: 1, and more preferably is the same as the corresponding amino acids at positions 9 to 22 in the sequence of SEQ ID NO: 1.

[0022] In one preferred embodiment, the HxE region of DYW:PG is HSE regardless of the other regions.

[0023] DYW:PG's CxxCH region, i.e., Cx a17 x a18 In CH: x a17 The amino acid is not particularly limited as long as it exhibits C-to-U editing activity as DYW:PG, but is preferably G (glycine) or an amino acid with similar properties, and more preferably G. x a18 The amino acid is not particularly limited as long as it exhibits C-to-U editing activity as DYW:PG, but is preferably D (aspartic acid) or an amino acid with similar properties, and more preferably D.

[0024] In DYW:PG, the area including the PG box and Hx aa Joining E...part, Hx aa The portion that joins the E region and the CxxCH region, and the portion that connects the CxxCH region and the DYW, are referred to as the first junction, the second junction, and the third junction, respectively (the same applies to other DYW domains).

[0025] The total length of the first linkage portion of DYW:PG is not particularly limited as long as it exhibits C-to-U editing activity as DYW:PG, but is, for example, 39 to 47 amino acids long, preferably 40 to 46 amino acids long, more preferably 41 to 45 amino acids long, and even more preferably 42 to 44 amino acids long. The amino acid sequence of the first linkage portion is not particularly limited as long as it exhibits C-to-U editing activity as DYW:PG, but is preferably the same as the portion at positions 25 to 67 of sequence number 1, or a sequence in which 1 to 22 amino acids are substituted, deleted, or added in that subsequence, or a sequence having sequence identity with that subsequence, and more preferably the same as that subsequence.

[0026] One preferred embodiment of the first ligation site of DYW:PG is a polypeptide represented by the following formula, which is 43 amino acids long, regardless of the sequence of the other parts of the DYW domain.

[0027] N a25 -N a26 -N a27 - … -N a65 -N a66 -N a67

[0028] The polypeptide described above is preferably a sequence that is the same as the portion of the sequence at positions 25-67 of Sequence ID No. 1, or a sequence in which multiple amino acids are substituted in that subsequence, and which can exhibit C-to-U editing activity as DYW:PG. In this case, the amino acid substitution is made by substituting an amino acid with a large bits value at the corresponding position in Figure 4 (for example, N a29 , N a30 , N a32 , N a33 , N a35 , N a36 , N a40 , N a44 , N a45 , N a47 , N a48 , N a52 , N a53 , N a54 , N a55 , N a58 , N a61 , N a65 , N a67 ) is the same as in Figure 4, and it is preferable that the substitution is carried out so that the other amino acids are replaced.

[0029] The total length of the second linkage of DYW:PG is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG, but is, for example, 21 to 29 amino acids long, preferably 22 to 28 amino acids long, more preferably 23 to 27 amino acids long, and even more preferably 24 to 26 amino acids long. The amino acid sequence of the second linkage is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG, but is preferably the same as the portion of the sequence at positions 71 to 95 of sequence number 1, or a sequence in which 1 to 13 amino acids are substituted, deleted, or added in that subsequence, or a sequence having sequence identity with that subsequence, and more preferably the same as that subsequence.

[0030] One preferred embodiment of the second ligation site of DYW:PG is a polypeptide represented by the following formula, which is 25 amino acids long, regardless of the sequence of the other parts of the DYW domain.

[0031] N a71 -N a72 -N a73 - … -N a93 -N a94 -N a95

[0032] The polypeptide described above is preferably a sequence that is the same as the portion of the sequence at positions 71-95 of Sequence ID No. 1, or a sequence in which multiple amino acids are substituted in that subsequence, and which can exhibit C-to-U editing activity as DYW:PG. In this case, the amino acid substitution is made by substituting an amino acid with a large bits value at the corresponding position in Figure 4 (for example, N a71 , N a72 , N a73 , N a76 , N a77 , N a78 , N a79 , N a81 , N a82 , N a86 , N a88 , N a89 , N a91 , N a92 , N a93 , N a94) is the same as in Figure 4, and it is preferable that the substitution is carried out so that the other amino acids are replaced.

[0033] The total length of the third linkage of DYW:PG is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG, but is, for example, 29 to 37 amino acids long, preferably 30 to 36 amino acids long, more preferably 31 to 35 amino acids long, and even more preferably 32 to 34 amino acids long. The amino acid sequence of the third linkage is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:PG, but is preferably the same as the portion of the sequence of Sequence ID No. 1 at positions 101 to 133, or a sequence in which 1 to 17 amino acids are substituted, deleted, or added in that subsequence, or a sequence having sequence identity with that subsequence, and more preferably the same as that subsequence.

[0034] One preferred embodiment of the third ligation site of DYW:PG is a polypeptide represented by the following formula, which is 33 amino acids long, regardless of the sequence of the other parts of the DYW domain.

[0035] N a101 -N a102 -N a103 - … -N a131 -N a132 -N a133

[0036] The polypeptide described above is preferably a sequence that is the same as the portion of the sequence of sequence number 101-133, or a sequence in which multiple amino acids are substituted in that subsequence, and which can exhibit C-to-U editing activity as DYW:PG. In this case, the amino acid substitution is made by an amino acid with a large bits value at the corresponding position in Figure 4 (for example, N a102 , N a104 , N a107 , N a112 , N a114 , N a117 , N a118 , N a121 , N a122 , N a123 , N a124 , N a125 , Na128 , N a130 , N a131 , N a132 ) is the same as in Figure 4, and it is preferable that the substitution is carried out so that the other amino acids are replaced.

[0037] RNA editing activity can be improved by introducing a mutation into the PG domain consisting of the sequence of SEQ ID NO: 1. A preferred example of such a mutated domain is PG11 (a polypeptide consisting of the amino acid sequence of SEQ ID NO: 50) shown in the Examples section of this specification.

[0038] (DYW:WW) DYW:WW is x b1 PGx b2 SWWTDx b3 -x b16 HP… Hx bb E … Cx b17 x b18 CH is a polypeptide consisting of DYW. Preferably, x b1 PGx b2 SWWTDx b3 -x b16 HP… Hx bb E … Cx b17 x b18 CH…DYW is a polypeptide having sequence identity with the sequence of SEQ ID NO:2 and possessing C-to-U / U-to-C editing activity. DYW:WW has the activity to convert the editing target C to U (C-to-U editing activity). SEQ ID NO:2 shows the sequence of DYW:WW, consisting of 137 amino acids in length, used in the experiments described in the Examples section of this specification. This sequence is novel and is disclosed for the first time in this application.

[0039] The total length of DYW:WW is not particularly limited as long as it exhibits C-to-U editing activity, but is for example 110 to 160 amino acids long, preferably 125 to 149 amino acids long, more preferably 129 to 145 amino acids long, and even more preferably 133 to 141 amino acids long.

[0040] The region containing the PG box of DYW:WW, i.e., x b1 PGx b2 SWWTDx b3 -x b16 In HP, the portion consisting of WTD may also be WSD.

[0041] In the area containing the PG box of DYW:WW: x b1 The amino acid is not particularly limited as long as it exhibits C-to-U editing activity as DYW:WW, but is preferably K (lysine) or an amino acid with similar properties, and more preferably K. x b2 The amino acid is not particularly limited as long as it exhibits C-to-U editing activity as DYW:WW, but is preferably Q (glutamine) or an amino acid with similar properties, and more preferably Q. x b3 -x b16 Each of the amino acids is not particularly limited as long as it exhibits C-to-U editing activity as DYW:WW, but is preferably the same as or similar in properties to the corresponding amino acids at positions 10-23 of the sequence of Sequence ID No. 2, and more preferably is the same as the corresponding amino acids at positions 10-23 of the sequence of Sequence ID No. 2.

[0042] In one preferred embodiment, the HxE region of DYW:WW is HSE regardless of the arrangement of the other parts.

[0043] The CxxCH region of DYW:WW, i.e., Cx b17 x b18 In CH: x b17 The amino acid is not particularly limited as long as it exhibits C-to-U editing activity as DYW:WW, but is preferably D (aspartic acid) or an amino acid with similar properties, and more preferably D. x b18 The amino acid is not particularly limited as long as it exhibits C-to-U editing activity as DYW:WW, but is preferably D or an amino acid with similar properties, and more preferably D.

[0044] The total length of the first linkage of DYW:WW is not particularly limited as long as it exhibits C-to-U editing activity as DYW:WW, but is, for example, 39 to 47 amino acids long, preferably 40 to 46 amino acids long, more preferably 41 to 45 amino acids long, and even more preferably 42 to 44 amino acids long. The amino acid sequence of the first linkage is not particularly limited as long as it exhibits C-to-U editing activity as DYW:WW, but is preferably the same as the portion of the sequence of Sequence ID No. 2 at positions 25 to 67, or a sequence in which 1 to 22 amino acids are substituted, deleted, or added in that subsequence, or a sequence having sequence identity with that subsequence, and more preferably the same as that subsequence.

[0045] One preferred embodiment of the first ligation site of DYW:WW is a polypeptide represented by the following formula, which is 43 amino acids long, regardless of the sequence of the other parts of the DYW domain.

[0046] N b26 -N b27 -N b28 - … -N b66 -N b67 -N b68

[0047] The polypeptide described above is preferably a sequence that is the same as the portion of the sequence from positions 26 to 68 of sequence number 2, or a sequence in which multiple amino acids are substituted in that subsequence, and which can exhibit C-to-U editing activity as DYW:PG. In this case, the amino acid substitution is made by substituting an amino acid with a large bits value at the corresponding position in Figure 4 (for example, N b26 , N b30 , N b33 , N b34 , N b37 , N b41 , N b45 , N b46 , N b48 , N b49 , N b51 , N b52 , N b53 , N b55 , N b56 , N b57 , Nb59 , N b61 , N b62 , N b63 , N b64 , N b66 , N b67 , N b68 ) is the same as in Figure 4, and it is preferable that the substitution is carried out so that the other amino acids are replaced.

[0048] The total length of the second linkage of DYW:WW is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW, but is, for example, 21 to 29 amino acids long, preferably 22 to 28 amino acids long, more preferably 23 to 27 amino acids long, and even more preferably 24 to 26 amino acids long. The amino acid sequence of the second linkage is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW, but is preferably the same as the portion of the sequence at positions 71 to 95 of sequence number 2, or a sequence in which 1 to 13 amino acids are substituted, deleted, or added in that subsequence, or a sequence having sequence identity with that subsequence, and more preferably the same as that subsequence.

[0049] One preferred embodiment of the second ligation site of DYW:WW is a polypeptide represented by the following formula, which is 25 amino acids long, regardless of the sequence of the other parts of the DYW domain.

[0050] N b72 -N b73 -N b74 - … -N b94 -N b95 -N b96

[0051] The polypeptide described above is preferably a sequence that is the same as the portion of the sequence of Sequence ID No. 2 from positions 72 to 96, or a sequence in which multiple amino acids are substituted in that subsequence, and which can exhibit C-to-U editing activity as DYW:WW. In this case, the amino acid substitution is made by an amino acid with a large bits value at the corresponding position in Figure 4 (for example, N b72 , N b73 , N b74 , N b75 , Nb77 , N b78 , N b79 , N b81 , N b82 , N b84 , N b88 , N b89 , N b90 , N b91 , N b92 , N b93 , N b94 , N b95 , N b96 ) is the same as in Figure 4, and it is preferable that the substitution is carried out so that the other amino acids are replaced.

[0052] The total length of the third linkage of DYW:WW is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW, but is, for example, 29 to 37 amino acids long, preferably 30 to 36 amino acids long, more preferably 31 to 35 amino acids long, and even more preferably 32 to 34 amino acids long. The amino acid sequence of the third linkage is not particularly limited as long as it can exhibit C-to-U editing activity as DYW:WW, but is preferably the same as the portion of the sequence of Sequence ID No. 2 at positions 101 to 133, or a sequence in which 1 to 17 amino acids are substituted, deleted, or added in that subsequence, or a sequence having sequence identity with that subsequence, and more preferably the same as that subsequence.

[0053] One preferred embodiment of the third ligation site of DYW:WW is a polypeptide represented by the following formula, which is 33 amino acids long, regardless of the sequence of the other parts of the DYW domain.

[0054] N b102 -N b103 -N b104 - … -N b132 -N b133 -N b134

[0055] The above polypeptide is preferably the same as the portion at positions 102 to 134 of the sequence of SEQ ID NO: 2 or a sequence in which a plurality of amino acids are substituted in that partial sequence and can exhibit C-to-U editing activity as DYW:WW. At this time, the amino acid substitution is such that the amino acid with a large bits value (for example, N b104 、N b105 、N b107 、N b108 、N b109 、N b110 、N b111 、N b113 、N b115 、N b116 、N b117 、N b118 、N b119 、N b121 、N b122 、N b123 、N b124 、N b126 、N b129 、N b131 、N b132 、N b133 N b134 ) is the same as that in FIG. 4, and it is preferable that the other amino acids are substituted. <00008�4><00008�5><00008�6>By introducing a mutation into the WW domain consisting of the sequence of SEQ ID NO: 2, the RNA editing activity can be improved. Preferred examples of such a domain with a mutation introduced are WW2 to 11 and WW13 shown in the Examples section of this specification, and particularly, the one with high editing activity is WW11 (a polypeptide consisting of the amino acid sequence of SEQ ID NO: 63). <00008�7><00008�8><00008�9>(DYW:KP) DYW:KP is a polypeptide consisting of KPAx[[ID=5۷]] c1 [[ID=5۸]]Ax[[ID=5۹]] c2 [[ID=6۰]]IEx[[ID=6۱]]<000041۷>[[ID=6۲]]… Hx[[ID=6۳]]<000041۸>[[ID=6۴]]E … Cx[[ID=6۵]]<000041۹>[[ID=6۶]]x[[ID=6۷]]<000042۰>[[ID=6۸]]CH … x[[ID=6۹]]<000042۱>[[ID=۷۰]]x[[ID=۷۱]]<000042۲>[[ID=۷۲]]x[[ID=۷۳]]<000042۳>[[ID=۷۴]]からなるポリペプチドである。好ましくは、KPAx[[ID=۷۵]]<000042۴> [[ID=۷۶]]Ax[[ID=۷۷]]<000042۵>[[ID=۷۸]]IEx [[ID=۷۹]]<000042۶>… Hx cc E … Cx c4 x c5 CH … x c6 x c7 x c8 It is a polypeptide that has, has sequence identity with the sequence of SEQ ID NO: 3, and has C-to-U / U-to-C editing activity. DYW:KP has the activity of converting the editing target U to C (U-to-C editing activity). SEQ ID NO: 3 shows the sequence of DYW:KP consisting of 133 amino acids in total length used in the experiments shown in the Examples section of this specification. This sequence is disclosed for the first time in this application and is novel.

[0058] The full length of DYW:KP is not particularly limited as long as it can exhibit U-to-C editing activity. For example, it is 110 to 160 amino acids in length, preferably 121 to 145 amino acids in length, more preferably 125 to 141 amino acids in length, and even more preferably 129 to 137 amino acids in length.

[0059] The region containing the PG box of DYW:KP, that is, KPAx c1 Ax c2 IEx c3 In: x c1 is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably S (serine) or an amino acid with similar properties, and more preferably S. x c2 is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably L (leucine) or an amino acid with similar properties, and more preferably L. x c3 is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably V (valine) or an amino acid with similar properties, and more preferably V.

[0060] In one preferred embodiment, the HxE region of DYW:KP is HAE regardless of the sequences of other parts.

[0061] The CxxCH region of DYW:KP, i.e., Cx c4 x c5 In CH: x c4 The amino acid is not particularly limited as long as it exhibits U-to-C editing activity as DYW:KP, but is preferably N (asparagine) or an amino acid with similar properties, and more preferably N. x b5 The amino acid is not particularly limited as long as it exhibits U-to-C editing activity as DYW:KP, but is preferably D (aspartic acid) or an amino acid with similar properties, and more preferably D.

[0062] DYW: The part corresponding to DYW in KP, i.e., x c6 x c7 x c8 In: x c6 The amino acid is not particularly limited as long as it exhibits U-to-C editing activity as DYW:KP, but is preferably D (aspartic acid) or an amino acid with similar properties, and more preferably D. x c7 The amino acid is not particularly limited as long as it exhibits U-to-C editing activity as DYW:KP, but is preferably M (methionine) or an amino acid with similar properties, and more preferably M. x b8 The amino acid is not particularly limited as long as it exhibits U-to-C editing activity as DYW:KP, but is preferably F (phenylalanine) or an amino acid with similar properties, and more preferably F. In one preferred embodiment, x c6 x c7 x c8 Regardless of the arrangement of the other parts, Dx c7 x c8 That is the case. In another preferred embodiment, x c6 x c7 x c8 It is a GRP regardless of the arrangement of the other parts.

[0063] The total length of the first linkage of DYW:KP is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is, for example, 51 to 59 amino acids long, preferably 52 to 58 amino acids long, more preferably 53 to 57 amino acids long, and even more preferably 54 to 56 amino acids long. The amino acid sequence of the first linkage is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably the same as the portion of the sequence of Sequence ID No. 3 at positions 10 to 64, or a sequence in which 1 to 28 amino acids are substituted, deleted, or added in that subsequence, or a sequence having sequence identity with that subsequence, and more preferably the same as that subsequence.

[0064] One preferred embodiment of the first ligation site of DYW:KP is a polypeptide represented by the following formula, which is 55 amino acids long, regardless of the sequence of the other parts of the DYW domain.

[0065] N c10 -N c11 -N c12 - … -N c62 -N c63 -N c64

[0066] The polypeptide described above is preferably a sequence that is the same as the portion of the sequence of sequence number 3 from positions 10 to 64, or a sequence in which multiple amino acids are substituted in that subsequence, and which can exhibit U-to-C editing activity as DYW: :KP. In this case, the amino acid substitution is made by an amino acid with a large bits value at the corresponding position in Figure 4 (for example, N c10 , N c13 , N c14 , N c15 , N c16 , N c17 , N c18 , N c19 , N c25 , N c26 , N c29 , N c30 , N c33 , N c34 , N c36 , N c38 , N c41 , Nc42 , N c44 , N c45 , N c47 , N 49 , N c55 , N c58 , N c59 , N c62 , N c63 , N c64 ) is the same as in Figure 4, and it is preferable that the substitution is carried out so that the other amino acids are replaced.

[0067] The total length of the second linkage of DYW:KP is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is, for example, 21 to 29 amino acids long, preferably 22 to 28 amino acids long, more preferably 23 to 27 amino acids long, and even more preferably 24 to 26 amino acids long. The amino acid sequence of the second linkage is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably the same as the portion of the sequence of Sequence ID No. 3 at positions 68 to 92, or a sequence in which 1 to 13 amino acids are substituted, deleted, or added in that subsequence, or a sequence having sequence identity with that subsequence, and more preferably the same as that subsequence.

[0068] One preferred embodiment of the second ligation site of DYW:KP is a polypeptide represented by the following formula, which is 25 amino acids long, regardless of the sequence of the other parts of the DYW domain.

[0069] N c68 -N c69 -N c70 - … -N c90 -N c91 -N c92

[0070] The polypeptide described above is preferably a sequence that is the same as the portion of the sequence 68-92 of Sequence ID No. 3, or a sequence in which multiple amino acids are substituted in that subsequence, and which can exhibit U-to-C editing activity as DYW:KP. In this case, the amino acid substitution is made by substituting an amino acid with a large bits value at the corresponding position in Figure 4 (for example, N c68 , Nc70 , N c71 , N c72 , N c73 , N c74 , N c75 , N c76 , N c77 , N c78 , N c79 , N c80 , N c81 , N c83 , N c84 , N c85 , N c86 , N c87 , N c88 , N c89 , N c90 , N c91 , N c92 ) is the same as in Figure 4, and it is preferable that the substitution is carried out so that the other amino acids are replaced.

[0071] The total length of the third linkage of DYW:KP is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is for example 29 to 37 amino acids long, preferably 30 to 36 amino acids long, more preferably 31 to 35 amino acids long, and even more preferably 32 to 34 amino acids long. The amino acid sequence of the third linkage is not particularly limited as long as it can exhibit U-to-C editing activity as DYW:KP, but is preferably the same as the portion of the sequence of Sequence ID No. 3 at positions 98 to 130, or a sequence in which 1 to 17 amino acids are substituted, deleted, or added in that subsequence, or a sequence having sequence identity with that subsequence, and more preferably the same as that subsequence.

[0072] One preferred embodiment of the third ligation site of DYW:KP is a polypeptide represented by the following formula, which is 33 amino acids long, regardless of the sequence of the other parts of the DYW domain.

[0073] N c98 -N c99 -N c100 - … -N c128 -N c129 -N c130

[0074] The polypeptide described above is preferably the same as the portion of the sequence 98-130 of Sequence ID No. 3, or a sequence in which multiple amino acids are substituted in that subsequence, and which can exhibit U-to-C editing activity as DYW:KP. In this case, the amino acid substitution is made by substituting an amino acid with a large bits value at the corresponding position in Figure 4 (for example, N c98 , N c100 , N c101 , N c103 , N c104 , N c105 , N c107 , N c109 , N c110 , N c111 , N c112 , N c114 , N c115 , N c117 , N c118 , N c119 , N c120 , N c121 , N c122 , N c123 , N c124 , N c125 , N c127 , N c129 , N c130 ) is the same as in Figure 4, and it is preferable that the substitution is carried out so that the other amino acids are replaced.

[0075] Editing activity can be improved by introducing mutations into the KP domain consisting of the sequence of Sequence ID No. 3. Preferred examples of such mutated domains are KP2-23 (Sequence IDs No. 68-89) shown in the Examples section of this specification. KP22 (Sequence ID No. 88) exhibits the highest U to C editing activity and the lowest C to U editing activity, showing improved RNA editing activity compared to the KP domain consisting of the sequence of Sequence ID No. 3.

[0076] (Chimeric DYW) The DYW domain can be divided into several regions based on the conservation of its amino acid sequence. By exchanging these regions, Chimeric DYW can improve C-to-U editing activity or U-to-C editing activity. One preferred embodiment is the region containing the PG box of DYW:WW, i.e., xb1 PGx b2 SWWTDx b3 -x b16 HP and DYW are fused to other regions of DYW:KP (... Hx cc E... Cx c4 x c5 CH...) is fused. At this time, x b1 , x b2 , x b3 -x b16 is as described above for DYW:WW. Also, the first connecting part, the second connecting part, and the third connecting part are as described above for DYW:KP. The DYW part is Dx bc1 x bc2 may be. The full length is not particularly limited as long as it can exhibit U-to-C editing activity. For example, it is, for example, 110 to 160 amino acids long, preferably 125 to 149 amino acids long, more preferably 129 to 145 amino acids long, and even more preferably 133 to 141 amino acids long.

[0077] One of the preferred Chimeric domains is x b1 PGx b2 SWWTDx b3 -x b16 HP... Hx cc E... Cx c4 x c5 CH... Dx bc1 x bc2 has, has at least 40% sequence identity with the sequence of 90 of SEQ ID NO:2, and consists of a polypeptide having C-to-U / U-to-C editing activity.

[0078] One of the particularly preferred Chimeric domains is a polypeptide having sequence identity with the sequence of SEQ ID NO:90 and having U-to-C editing activity. SEQ ID NO:90 shows the sequence of the Chimeric domain used in the experiments shown in the Examples section of this specification. This domain shows higher U to C editing activity than DYW:KP consisting of the sequence of SEQ ID NO:3, and furthermore has almost no C to U editing activity.

[0079] (Comparison with known sequences) As mentioned above, the DYW domain consisting of sequences 1, 2, and 3 is novel. The results of an investigation using the PPR database (https: / / ppr.plantenergy.uwa.edu.au / onekp / ; see Non-Patent Document 11 above) are shown below.

[0080] [Table 2-1]

[0081] Therefore, the present invention also provides one polypeptide of any one of the following d to i: d. A polypeptide having more than 78% sequence identity with sequence number 1, preferably 80% or more, more preferably 85% or more, even more preferably 90%, even more preferably 95%, and even more preferably 97%, and having C-to-U editing activity; e. A polypeptide having more than 84% sequence identity with sequence number 2, preferably 85% or more, more preferably 90% or more, even more preferably 95%, and even more preferably 97% sequence identity, and having C-to-U editing activity; f. A polypeptide having more than 86% sequence identity with the sequence of Sequence ID No. 3, preferably 87% or more, more preferably 90% or more, even more preferably 95% sequence identity, and even more preferably 97% sequence identity, and having U-to-C editing activity. g. A polypeptide having a sequence in which 1 to 29, preferably 1 to 25, more preferably 1 to 21, even more preferably 1 to 17, even more preferably 1 to 13, even more preferably 1 to 9, and even more preferably 1 to 5 amino acids are substituted, deleted, or added to the sequence of SEQ ID NO: 1, and having C-to-U editing activity; h. Polypeptides having a sequence in which 1 to 21, preferably 1 to 18, more preferably 1 to 15, even more preferably 1 to 12, even more preferably 1 to 9, and even more preferably 1 to 6 amino acids are substituted, deleted, or added in the sequence of Sequence ID No. 2, and which have C-to-U editing activity; i. A polypeptide having a sequence in which 1 to 18 amino acids, preferably 1 to 16, more preferably 1 to 14, even more preferably 1 to 12, even more preferably 1 to 10, even more preferably 1 to 8, and even more preferably 1 to 6 amino acids are substituted, deleted, or added in the sequence of SEQ ID NO: 3, and having U-to-C editing activity.

[0082] The alignment of these sequences is shown below. The alignment was created using AliView (Larsson, A. (2014). AliView: a fast and lightweight alignment viewer and editor for large data sets. Bioinformatics30(22): 3276-3278. http: / / dx.doi.org / 10.1093 / bioinformatics / btu531). Identical amino acids are represented by dots.

[0083] [Table 2-2]

[0084] [RNA-binding domain] In this invention, when converting an editing target using a DYW domain, an RNA-binding protein is used as the RNA-binding domain to target the RNA containing the editing target.

[0085] A preferred example of an RNA-binding protein used as an RNA-binding domain is a PPR protein composed of a PPR motif.

[0086] (PPR motif) Unless otherwise specified, a PPR motif refers to a polypeptide consisting of 30 to 38 amino acids whose E value obtained at PF01535 in Pfam and PS51375 in Prosite, when the amino acid sequence is analyzed using a web-based protein domain search program, is less than or equal to a predetermined value (preferably E-03). The positional numbers of the amino acids constituting the PPR motif as defined in this invention are almost synonymous with PF01535, while corresponding to the number obtained by subtracting 2 from the amino acid position in PS51375 (e.g., No. 1 in this invention → No. 3 in PS51375). However, when referring to the amino acid at position "ii" (-2), it refers to the second amino acid from the end (C-terminus) of the amino acids constituting the PPR motif, or the amino acid two positions N-terminus relative to the first amino acid of the next PPR motif, i.e., the -2 amino acid. If the next PPR motif is not clearly identified, the amino acid two positions prior to the first amino acid of the next helix structure is referred to as "ii". For information on Pfam, please refer to http: / / pfam.sanger.ac.uk / , and for information on Prosite, please refer to http: / / www.expasy.org / prosite / .

[0087] The conserved amino acid sequence of the PPR motif has low conservation at the amino acid level, but the two α-helices are well conserved in the secondary structure. A typical PPR motif consists of 35 amino acids, but its length is variable, ranging from 30 to 38 amino acids.

[0088] More specifically, the PPR motif consists of a polypeptide with a length of 30 to 38 amino acids, represented by Formula 1.

[0089] [ka]

[0090] During the ceremony: Helix A is a 12-amino acid length region capable of forming an α-helix structure, represented by Equation 2.

[0091] [ka]

[0092] In formula 2, A1~A 12 Each of these independently represents an amino acid; X is either absent or a portion consisting of 1 to 9 amino acids in length; Helix B is a region consisting of 11-13 amino acids that can form an α-helix structure; L is the portion represented by formula 3, which is 2 to 7 amino acids long;

[0093] [ka]

[0094] In Equation 3, each amino acid is numbered from the C-terminus as "i" (-1), "ii" (-2), and so on. However, L iii ~L vii It may not exist.

[0095] (PPR-code) The PPR motif relies on a combination of three amino acids, positions 1, 4, and ii, for specific binding to a base. This combination determines which base will bind. The relationship between the combination of the three amino acids (1, 4, and ii) and the binding base is known as the PPR-code (see Patent Document 2 above), and is as follows.

[0096] (1) A1, A4, and L ii When the combination of these three amino acids is valine, asparagine, and aspartic acid, the PPR motif has selective RNA base binding ability, strongly binding to U, then to C, and then to A or G. (2) A1, A4, and L iiIn the case of the three amino acid combinations, valine, threonine, and asparagine, respectively, the PPR motif has selective RNA base binding ability, binding strongly to A, then to G, then to C, but not to U. (3) A1, A4, and L ii The combination of these three amino acids, in order, valine, asparagine, and asparagine, respectively, has a selective RNA base-binding ability in which its PPR motif strongly binds to C, then to A or U, but does not bind to G. (4) A1, A4, and L ii When the combination of these three amino acids is glutamic acid, glycine, and aspartic acid, the PPR motif has selective RNA base binding ability, strongly binding to G but not to A, U, and C. (5) A1, A4, and L ii The combination of these three amino acids, in order, isoleucine, asparagine, and in the case of asparagine, its PPR motif has selective RNA base binding ability, binding strongly to C, then to U, then to A, but not to G. (6) A1, A4, and L ii When the combination of these three amino acids is valine, threonine, and aspartic acid, the PPR motif has selective RNA base binding ability, strongly binding to G, then to U, but not to A and C. (7) A1, A4, and L ii When the combination of these three amino acids is lysine, threonine, and aspartic acid, the PPR motif has selective RNA base binding ability, strongly binding to G, then to A, but not to U and C. (8) A1, A4, and L ii When the combination of these three amino acids is phenylalanine, serine, and asparagine, the PPR motif has selective RNA base binding ability, binding strongly to A, then to C, and then to G and U. (9) A1, A4, and L iiWhen the combination of these three amino acids is valine, asparagine, and serine, the PPR motif has selective RNA base binding ability, strongly binding to C, then to U, but not to A and G. (10) A1, A4, and L ii When the combination of these three amino acids is phenylalanine, threonine, and asparagine, the PPR motif has selective RNA base binding ability, strongly binding to A but not to G, U, and C. (11) A1, A4, and L ii When the combination of these three amino acids is, in order, isoleucine, asparagine, and aspartic acid, its PPR motif exhibits selective RNA base binding ability, strongly binding to U, then to A, but not to G and C. (12) A1, A4, and L ii When the combination of these three amino acids is threonine, threonine, and asparagine, the PPR motif exhibits selective RNA base binding ability, strongly binding to A but not to G, U, and C. (13) A1, A4, and L ii When the combination of these three amino acids is, in order, isoleucine, methionine, and aspartic acid, its PPR motif exhibits selective RNA base binding ability, strongly binding to U, then to C, but not to A and G. (14) A1, A4, and L ii The combination of these three amino acids, in the case of phenylalanine, proline, and aspartic acid, is called PPR, and its motif has selective RNA base binding ability, strongly binding to U, then to C, but not to A and G. (15) A1, A4, and L ii When the combination of these three amino acids is tyrosine, proline, and aspartic acid, the PPR motif has selective RNA base binding ability, strongly binding to U but not to A, G, and C. (16) A1, A4, and L iiWhen the combination of these three amino acids is leucine, threonine, and aspartic acid, the PPR motif has selective RNA base binding ability, strongly binding to G but not to A, U, and C.

[0097] (P array) PPR proteins are classified into two families, P and PLS, based on the structure of the PPR motifs they comprise. P-type PPR proteins consist of a simple repeat (P array) of a standard 35-amino acid PPR motif (P). The DYW domain of the present invention can be used by ligating it to a P-type PPR protein.

[0098] (PLS array) The PPR motifs in PLS-type PPR proteins are arranged as repeating units of three PPR motifs: P1, L1, and S1. The P2, L2, and S2 motifs follow at the C-terminus of this repeat. Furthermore, two PPR-like motifs called E1 and E2, and a DYW domain may follow at the C-terminus of this final P2L2S2 motif.

[0099] The total length of P1 is not particularly limited as long as it can bind to the target base, but is for example 33 to 37 amino acids long, preferably 34 to 36 amino acids long, and more preferably 35 amino acids long. The total length of L1 is not particularly limited as long as it can bind to the target base, but is for example 33 to 37 amino acids long, preferably 34 to 36 amino acids long, and more preferably 35 amino acids long. The total length of S1 is not particularly limited as long as it can bind to the target base, but is for example 30 to 33 amino acids long, preferably 30 to 32 amino acids long, and more preferably 31 amino acids long.

[0100] The total length of P2 is not particularly limited as long as it can bind to the target base, but is for example 33 to 37 amino acids long, preferably 34 to 36 amino acids long, and more preferably 35 amino acids long. The total length of L2 is not particularly limited as long as it can bind to the target base, but is for example 34 to 38 amino acids long, preferably 35 to 37 amino acids long, and more preferably 36 amino acids long. The total length of S2 is not particularly limited as long as it can bind to the target base, but is for example 30 to 34 amino acids long, preferably 31 to 33 amino acids long, and more preferably 32 amino acids long. Sequence IDs 17, 22, and 27 show the sequences of P2 used in the Examples section of this specification. Sequence IDs 18, 23, and 28 show the sequences of L2 used in the Examples section of this specification. Sequence IDs 19, 24, and 29 show the sequences of S2 used in the Examples section of this specification.

[0101] The total length of E1 is not particularly limited as long as it can bind to the target base, but is, for example, 32 to 36 amino acids long, preferably 33 to 35 amino acids long, and more preferably 34 amino acids long. Sequence IDs 20, 25, and 30 show the sequences of E1 used in the Examples section of this specification.

[0102] The total length of E2 is not particularly limited as long as it can bind to the target base, but is, for example, 30 to 34 amino acids long, preferably 31 to 33 amino acids long, and more preferably 33 amino acids long. Sequence IDs 21, 26, and 31 show the sequences of E2 used in the Examples section of this specification.

[0103] In PLS-type PPR proteins, the P1L1S1 repeat portion and the portion up to P2 can be designed according to the PPR-code rules described above, depending on the sequence of the target RNA.

[0104] The S2 motif correlates with the nucleotide corresponding to amino acid ii (the 31st N in SEQ ID NO: 19) (see Non-Patent Document 8 cited above). Furthermore, it can be incorporated into PLS-type PPR proteins, keeping in mind that the C or U four positions to the right of the target base of the S2 motif is the editing target base for the DYW domain.

[0105] In the E1 motif, the correlation with the nucleotide is only observed at the fourth amino acid (the fourth G in sequence number 20) (Ruwe et al. (2019) New Phytol. 222 218-229).

[0106] The fourth amino acid (the fourth V in SEQ ID NO: 21) and the last amino acid (the 33rd K in SEQ ID NO: 21) in the E2 motif are highly conserved and are not involved in the recognition of specific PPR-RNAs (see Non-Patent Document 2 above).

[0107] The number of P1L1S1 repeats is not particularly limited as long as it can bind to the target base sequence, but is for example 1 to 5, preferably 2 to 4, and more preferably 3. In principle, even one unit (3 repeats) can be used. MEF8 (L1-S1-P2-L2-S2-E-DYW), which consists of 5 PPR motifs, is known to be involved in approximately 60 editing sites.

[0108] In natural PPR proteins, the first and last P1L1S1 units show clear differences in the amino acid residues at specific positions, distinguishing them from the internal P1L1S1 units. From the perspective of designing an artificial PLS array that is as close as possible to naturally occurring ones, it is advisable to design three types of P1L1S1 units corresponding to the positions of the PPR motif: the first (N-terminal) P1L1S1 unit, the internal P1L1S1 unit, and the last (C-terminal) P1L1S1 unit located immediately before P2L2S2. In addition to those composed of repeating PLS units, natural PPR proteins also sometimes contain repeating SS units (31 amino acids), and these can also be used in this invention.

[0109] [DYW Protein] The present invention provides a DYW protein for editing target RNA, comprising an RNA-binding domain that includes at least one PPR motif and is sequence-specifically bound to target RNA according to the rules of the PPR-code, and a DYW domain that is one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0110] Such DYW proteins can be artificial. Artificial means that they are not natural products but are artificially synthesized. Artificiality can be defined, for example, by having a DYW domain with a sequence not found in nature, a PPR-binding domain with a sequence not found in nature, a combination of RNA-binding domain and DYW domain not found in nature, or by having additional parts not found in natural plant-derived DYW proteins, such as nuclear localization signals or human mitochondrial localization signals, which are used for RNA editing in animal cells. Examples of nuclear localization signal sequences include PKKKRKV (SEQ ID NO:32) derived from the SV40 large T antigen, and KRPAATKKAGQAKKKK (SEQ ID NO:33), the NLS of nucleoplasmin.

[0111] The number of PPR motifs can be adjusted as appropriate depending on the sequence of the target RNA. At least one PPR motif is required, but two or more may be present. It is known that two PPR motifs can bind to RNA (Nucleic Acids Research, 2012, Vol. 40, No. 6, 2712-2723).

[0112] One preferred embodiment of the DYW protein is as follows: An artificial DYW protein for editing target RNA, comprising at least one, preferably two to 25, more preferably five to 20, and even more preferably ten to 18 PPR motifs, an RNA-binding domain capable of sequence-specifically binding to target RNA according to the rules of the PPR-code, and a DYW domain which is one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0113] Another preferred embodiment of the DYW protein is as follows: A DYW protein for editing target RNA, comprising at least one, preferably two to 25, more preferably five to 20, and even more preferably ten to 18 PPR motifs, an RNA-binding domain (preferably an RNA-binding domain that is a PLS-type PPR protein) capable of sequence-specifically binding to an animal target RNA according to the rules of the PPR-code, and a DYW domain that is one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0114] The DYW domain, the RNA-binding domain PPR protein, and the DYW protein of the present invention can be prepared in relatively large quantities by methods well known to those skilled in the art. Such methods may include determining the nucleic acid sequence encoding the domain or protein from the amino acid sequence of the domain or protein of interest, cloning it, and creating a transformant that produces the domain or protein of interest.

[0115] [Utilization of DYW protein] (Nucleic acids, vectors, and cells encoding DYW proteins, etc.) The present invention also provides the above-mentioned PPR motif, DYW protein, or nucleic acid encoding the DYW protein, and vectors containing nucleic acid (e.g., amplification vectors, expression vectors). Vectors also include viral vectors. Amplification vectors can use E. coli or yeast as hosts. In this specification, an expression vector means a vector containing, for example, DNA having a promoter sequence, DNA encoding the desired protein, and DNA having a terminator sequence from upstream, but the sequences do not necessarily have to be in this order as long as the desired function is performed. In the present invention, various vectors that are commonly used by those skilled in the art can be rearranged and used.

[0116] Specifically, the present invention provides a nucleotide sequence encoding a DYW protein, comprising at least one PPR motif, an RNA-binding domain (preferably an RNA-binding domain which is a PLS-type PPR protein) which is sequence-specifically bound to a target RNA (preferably an animal target RNA) according to the rules of the PPR-code, and a DYW domain which is one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0117] More specifically, the present invention provides a vector for editing target RNA, comprising a nucleotide sequence encoding a DYW protein, which includes at least one PPR motif and an RNA-binding domain (preferably an RNA-binding domain that is a PLS-type PPR protein) capable of sequence-specifically binding to target RNA (preferably an animal target RNA) according to the rules of the PPR-code, and a DYW domain that is one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0118] The DYW protein of the present invention can function in eukaryotic cells (e.g., animals, plants, microorganisms (yeast, etc.), protists). In particular, the DYW protein of the present invention can function in animal cells (in vitro or in vivo). Examples of animal cells into which the DYW protein of the present invention, or a vector expressing the DYW protein, can be introduced include cells derived from humans, monkeys, pigs, cattle, horses, dogs, cats, mice, and rats. Examples of cultured cells into which the DYW protein of the present invention, or a vector expressing the DYW protein, can be introduced include, but are not limited to, Chinese hamster ovary (CHO) cells, COS-1 cells, COS-7 cells, VERO (ATCC CCL-81) cells, BHK cells, canine kidney-derived MDCK cells, hamster AV-12-664 cells, HeLa cells, WI38 cells, HEK293 cells, HEK293T cells, and PER.C6 cells.

[0119] (Application) The DYW protein of the present invention can convert editing target C to U, or editing target U to C, within the target RNA. RNA-binding PPR proteins are involved in all RNA processing steps found in organelles: cleavage, RNA editing, translation, splicing, and RNA stabilization.

[0120] Furthermore, the DYW protein of the present invention enables single-base editing of mitochondrial RNA. Mitochondria have their own unique genomes and encode constituent proteins of important complexes involved in respiration and ATP production. It is known that mutations in these proteins can cause various diseases. Mutation repair using the present invention is expected to provide treatment for a variety of diseases.

[0121] On the other hand, the CRISPR-Cas system has been developed as a cytoplasmic C-to-U RNA editing tool by fusing the Cas protein to a modified ADAR domain (Abudayyeh et al., 2019). However, the CRISPR-Cas system consists of a protein and guide RNA, making efficient mitochondrial transport of the guide RNA difficult. In contrast, PPR proteins can perform RNA editing with a single molecule, and generally, proteins can be delivered to mitochondria by fusing a mitochondrial localization signal sequence to their N-terminus. Therefore, to confirm whether this technology can be used for mitochondrial RNA editing, we designed PPRs targeting MT-ND2 and MT-ND5 and created genes fused with PG or WW domains (Figure 10a).

[0122] These proteins, consisting of mitochondrial target sequences (MTS) and PPR-P sequences, target the third position of codons 178 and 301 of MT-ND2 and MT-ND5 to avoid adverse effects on HEK293T cells (Figure 10a). After plasmid introduction into HEK293T cells, editing was confirmed within the mRNA of MT-ND2 and MT-ND5 (Figure 10b,c). These four proteins showed up to 70% editing activity against their targets, and no off-target mutations were detected in the same mRNA molecules (Figure 10bc).

[0123] Therefore, the RNA base editing method provided by the present invention can be expected to have the following uses in various fields.

[0124] (1) Medical • Recognizes and edits specific RNAs associated with specific diseases. The present invention can be used to treat genetic diseases caused by single nucleotide mutations. Many mutations in genetic diseases are oriented from C to U. Therefore, the method of the present invention, which can convert U to C, may be particularly useful.

[0125] • Create cells with controlled RNA repression and expression. Such cells include stem cells (e.g., iPS cells) with monitored differentiated and undifferentiated states, model cells for evaluating cosmetics, and cells in which the expression of functional RNA can be switched ON / OFF for the purpose of elucidating drug discovery mechanisms and conducting pharmacological tests.

[0126] (2) Agriculture, forestry and fisheries • To improve yield and quality in agricultural products, forestry products, fishery products, etc. • To breed organisms with improved disease resistance, improved environmental tolerance, or enhanced or new functionalities.

[0127] For example, with regard to first-generation hybrid (F1) crops, it may be possible to artificially create F1 crops by editing mitochondrial RNA with DYW proteins, thereby improving yield and quality. RNA editing with DYW proteins allows for more accurate and rapid improvement of biological varieties and breeding (genetic improvement of organisms) than conventional techniques. Furthermore, since RNA editing with DYW proteins does not involve altering traits with foreign genes like genetic modification, it is closer to traditional breeding methods such as mutant selection and backcrossing. Therefore, it can reliably and quickly address global food and environmental problems.

[0128] (3) Chemistry In the production of useful substances using microorganisms, cultured cells, plants, and animals (e.g., insects), protein expression levels can be controlled by manipulating RNA. This can improve the productivity of useful substances. Examples of useful substances include proteinaceous substances such as antibodies, vaccines, and enzymes, as well as relatively low-molecular-weight compounds such as pharmaceutical intermediates, fragrances, and dyes.

[0129] • Improve the efficiency of biofuel production by modifying the metabolic pathways of algae and microorganisms.

[0130] [term] The numerical range x~y includes the values ​​x and y at both ends unless otherwise specified.

[0131] In relation to the amino acid sequences of proteins and polypeptides, amino acid residues are sometimes simply referred to as amino acids.

[0132] With respect to base sequences (sometimes called nucleotide sequences) or amino acid sequences, "identity," unless otherwise specified, refers to the percentage of matching bases or amino acids shared between two sequences when the two sequences are aligned in the most optimal manner. That is, identity can be calculated as (number of matching positions / total number of positions) × 100, and can be calculated using commercially available algorithms. Such algorithms are incorporated into the NBLAST and XBLAST programs described in Altschul et al., J.Mol.Biol. 215(1990) 403-410. More specifically, the search and analysis of identity between base sequences or amino acid sequences can be performed using algorithms or programs well known to those skilled in the art (e.g., BLASTN, BLASTP, BLASTX, ClustalW). When using a program, the parameters can be appropriately set by those skilled in the art, or the default parameters of each program may be used. The specific methods of these analysis methods are also well known to those skilled in the art.

[0133] With respect to the base sequence or amino acid sequence, a high degree of sequence identity is preferred unless otherwise specified. Specifically, it is preferable to have 40% or more, more preferably 45% or more, even more preferably 50% or more, even more preferably 55% or more, even more preferably 60% or more, and even more preferably 65% ​​or more. Furthermore, it is preferable to have 70% or more, more preferably 80% or more, even more preferably 85% or more, even more preferably 90% or more, even more preferably 95% or more, and even more preferably 97.5% or more.

[0134] With respect to a polypeptide or protein, the number of amino acids substituted, deleted, or added in a "substituted, deleted, or added sequence" is not particularly limited in any motif or protein, as long as the motif or protein consisting of that amino acid sequence has the desired function, unless otherwise specified. However, it is usually around 1 to 9 or 1 to 4 amino acids, or even more if the substitutions are with similar amino acids. Means for preparing polynucleotides or proteins relating to such amino acid sequences are well known to those skilled in the art.

[0135] Similar amino acids refer to amino acids with similar physical properties such as hydroxyl, charge, pKa, and solubility. Examples include the following: Hydrophobic (nonpolar) amino acids: alanine, valine, glycine, isoleucine, leucine, phenylalanine, proline, tryptophan, tyrosine Nonhydrophobic amino acids: arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine, cysteine, histidine, methionine; Hydrophilic amino acids; arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine; Acidic amino acids: aspartic acid, glutamic acid; Basic amino acids: lysine, arginine, histidine; Neutral amino acids: alanine, asparagine, cysteine, glutamine, glycine, isoleucine, leucine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine; Sulfur-containing amino acids: methionine, cysteine; Aromatic ring amino acids: tyrosine, tryptophan, phenylalanine. [Examples]

[0136] [1. Design of the DYW domain and measurement of RNA editing activity in E. coli] [result] (Design of the DYW domain) Since most of the currently available plant genome information is from angiosperms, the development of artificial DYW proteins is limited to the DYW:PG type. On the other hand, transcriptome data, although partial sequences, also includes information on early land plant species that possess both C-to-U and U-to-C RNA editing. Therefore, we designed a complete artificial DYW domain using a database of PPR proteins constructed based on transcriptome datasets created by the 1000 Plants (1KP) international consortium (see Non-Patent Literature 11 above).

[0137] In designing the artificial DYW domain, we incorporated PPR (P2 and L2) motifs and PPR-like (S2, E1 and E2) motifs in addition to the cytidine deaminase-like DYW domain. This is because the importance of these motifs to RNA editing activity is not currently understood. We constructed phylogenetic trees of DYW domains from hornworts, chlorophyllans, and ferns and classified them into three groups: DYW:PG, DYW:WW, and DYW:KP. Next, we constructed new phylogenetic trees for the DYW domains within each of these groups and selected protein lineages that could be used for the design of artificial DYW domains (Figure 1). The DYW:PG group showed large differences between sequences, making it difficult to select protein lineages for design (Figure 1a). Because the sequences of chlorophyll proteins were highly diverse and numerous, we decided to focus on this protein (PG1). For the design of the DYW:WW domain, we selected a lineage of short-branched proteins found only in hornworts (Figure 1b, WW1). The short branching suggests that gene variation between proteins in this lineage is small. For the design of the DYW:KP domain, we focused on a lineage of DYW domains specific to ferns. The protein sequences in this lineage show significant amino acid variation, but their length is conserved (Figure 1c, KP1).

[0138] (Design of RNA-binding domain) We fused each DYW domain with an artificial P or PLS array designed based on PPR motif sequences identified from plant genomic information (see Non-Patent Literature 2 above) (Figure 2a). The PPR motifs in the P array were constructed based on a common sequence obtained from the alignment of 35-amino acid P motifs, with some amino acids replaced to enhance RNA recognition. For the design of the PLS array, we selected a standard-length 35-amino acid motif (P1 and L1) and a 31-amino acid motif (S1). In natural PPR proteins, the first and last P1L1S1 show clear differences in amino acid residues at specific positions and are different from the internal P1L1S1. To design an artificial PLS array as close as possible to the naturally occurring one, we divided it into the first (N-terminal) P1L1S1, the internal P1L1S1, and the last (C-terminal) P1L1S1 located immediately before P2L2S2, designing three types of P1L1S1 according to the position of the PPR motif. In the future, these proteins will be named based on their structure. For example, a protein with a DYW:WW domain fused to a P array will be named P-DYW:WW.

[0139] (Artificial DYW proteins can be specifically edited for target sequences.) In Arabidopsis thaliana, the PPR protein CLB19 recognizes RNA editing sites on chloroplast rpoA and clpP RNA. We decided to design a PPR protein that targets the rpoA editing site. The fourth and second amino acids in the P1, L1, S1, and P2 motifs followed the PPR code. However, since the PPR code for the C-terminal PPR-like motifs (L2, S2, E1, E2) is unknown, we used the fourth and second amino acids of CLB19 for the L2, S2, E1, and E2 motifs (Figure 2b).

[0140] The gene region encoding the recombinant artificial DYW protein was cloned into an expression vector, and a target sequence was added downstream of its stop codon. Based on previous research on two types of PPR proteins (PPR56 and PPR65) from the moss Physcomitrella patens (see Non-Patent Literature 12 above), a method was developed to test the RNA editing activity of the designed PPR protein in E. coli. The designed PLS-DYW:PG1 and PLS-DYW:WW1 showed no editing activity on DNA, but substituted cytidine with uridine on RNA with an editing efficiency of over 90% (Figures 3a, b, d). When a P array was used instead of the PLS domain, the editing efficiency decreased by 10-40%. On the other hand, U-to-C RNA editing activity was observed in both P-DYW:KP and PLS-DYW:KP (Figures 3c, e). However, their editing activity was lower compared to DYW:PG and DYW:WW.

[0141] [method] (phylogenetic tree) The P2-L2-S2-E1-E2-DYW region, containing the DYW domain (shortest 132 amino acids), was extracted from the PPR database (https: / / ppr.plantenergy.uwa.edu.au / onekp / ; see Non-Patent Document 11 above). Alignment of the obtained sequence was created using MAFFT L-INS-i (v7.407 automatic mode) (K. Katoh, DM Standley (2013). Mol Biol Evol. 30(4): 772-780.), and then trimmed using trimAl (v.1.4.rev15) (Salvador Capella-Gutierrez, et al. (2009). Bioinformatics. 25(15): 1972-1973.). The trimming parameters were set to gt 0.2 and cons 20. Active site (HxEx n Sequences with mutations in CxxCH were excluded from the alignment (Figure 4).

[0142] From the remaining sequences, a phylogenetic tree was constructed using FastTree (v2.1.10) (Price, MN, et al. (2010). PLoS One 5:e9490.), and DYW:PG, DYW:WW, and DYW:KP were identified. The parameters used were wag and cat 8.

[0143] (Cloning of the Trx-PPR-DYW protein and target sequence) The common sequences of each DYW domain (including P2, L2, S2, E1, and E2) were designed using EMBOSS:cons (v.6.6.0.0). The protein expression vector was modified from pET21b+PA, with the original Esp3I and BpiI restriction enzyme sites removed and two Esp3I sites added as cloning sites. The gene was divided into four sections (Trx, PPR array, DYW domain, and RNA editing site) and constructed using a two-step Golden Gate method. First, three parts were cloned into the Esp3I site of the modified pET21b vector: 1) the thioredoxin-6×His-TEV gene region (containing the BpiI restriction enzyme site at 3'), 2) the P2-L2-S2-E1-E2-DYW gene region (containing the BpiI restriction enzyme site at 5'), and 3) the coding sequence region of the RNA editing site. Next, a full-length PLS domain or P domain was cloned into the BpiI region to produce PLS:DYW (SEQ ID NOs: 35-37) or P:DYW (SEQ ID NOs: 38-40) proteins.

[0144] (Measurement of RNA editing activity in E. coli) To analyze the RNA editing activity of recombinant proteins in E. coli, we modified the protocol developed by Oldenkott et al. (Non-Patent Document 12, cited above). Plasmid DNA prepared above was introduced into two strains of E. coli Rosetta and cultured overnight at 37°C in 1 mL of LB medium (containing carbenicillin 50 μg / mL and chloramphenicol 17 μg / mL). 5 mL of LB medium containing appropriate antibiotics was prepared in a deep-bottom 24-well plate, and 100 μL of the pre-culture solution was seeded therein. This culture solution was then analyzed by measuring its absorbance (OD 600The cells were grown at 37°C and 200 rpm until the saturation reached 0.4-0.6, after which the plates were cooled to 4°C for 10-15 minutes. Next, 0.4 mM ZnSO4 and 0.4 mM IPTG were added, and the cells were cultured for a further 18 hours at 16°C and 180 rpm. 750 μL of the culture solution was collected and centrifuged, and the cell pellet was frozen in liquid nitrogen and stored at -80°C.

[0145] Frozen cell aggregates were resuspended in 200 μL of 1-thioglycerol / homogenized solution, sonicated at 40 W for 10 seconds to separate the cells, and then 200 μL of lysis buffer was added. RNA was extracted using Maxwell® RSC simplyRNA Tissue Kit (Promega). The RNA was treated with DNase I (Takara Bio), and cDNA was synthesized using 1 μg of the treated RNA and 1.25 μM random primers (6 mer) with SuperScript® III Reverse Transcriptase (Invitrogen). The region containing the editing site was amplified using NEBNext High-Fidelity 2x PCR Master Mix (New England Biolabs), 1 μL of cDNA, and primers for thioredoxin and the T7 terminator sequence. The PCR product was purified with NucleoSpin® Gel and PCR Cleanup (Takara Bio), and sequencing analysis was performed using forward primers specific to the DYW domain sequence to determine the bases of the RNA editing site. RNA editing efficiency was measured using the ratio of the waveform peak heights of C and U at the editing site. C-to-U RNA editing efficiency was calculated as U / (C+U)×100, and U-to-C editing efficiency as C / (C+U)×100. Three independent experiments were repeated.

[0146] [List of cited sequences]

[0147] [Table 3-1]

[0148] [Table 3-2]

[0149] [Table 3-3]

[0150] [2. Examples using animal cultured cells] [result] Genes fused with PLS-type PPR and each DYW domain (PG1 (SEQ ID NO: 1), WW1 (SEQ ID NO: 2), KP1 (SEQ ID NO: 3)) and plasmids containing the target sequence were transfected into HEK293T cells, and RNA was recovered after culture. The conversion efficiency from cytidine (C) to uridine (U) or from uridine (U) to cytidine (C) at the target site was analyzed by Sanger sequencing (Figure 5). When the PG1 or WW1 domain was fused, more than 90% C to U activity was observed, while U to C activity was not detected (Figure 5a, c). When the KP1 domain was fused, 25% U to C activity was detected, but approximately 10% C to U activity was also detected (Figure 5b, c). These results indicate that editing enzymes function even in animal cultured cells.

[0151] [method] (Preparation of PPR expression plasmids for animal cell culture testing) From the plasmid used in Figure 3, the gene sequence of the 6xHis-PPR-DYW protein (the same protein used in the E. coli experiment (sequence numbers 35-37)) and the region containing the editing site (sequence number 34) were amplified by PCR and cloned into an animal cell expression vector using the Golden Gate Assembly method. In the vector, PPR is expressed under the control of a promoter containing the CMV promoter and a human β-globin chimerichiontron, and the poly A signal is conferred by the SV40 polyadenylation signal.

[0152] (HEK293T cell culture) HEK293T cells were cultured at 37°C and 5% CO2 in Dulbecco's Modified Eagle Medium (DMEM) medium containing high glucose, glutamine, phenol-RED, and sodium pyruvate (Fujifilm Wako Pure Chemical Corporation), supplemented with 10% fetal bovine serum (Capricorn) and 1% penicillin-streptomycin (Fujifilm Wako Pure Chemical Corporation). Cells were passaged every 2-3 days once they reached 80-90% confluence.

[0153] (Transfection) In the RNA editing assay, HEK293T cells were placed in each well of a 24-well flat-bottom cell culture plate, approximately 8.0 x 10⁶ cells. 4 Cells were placed in individual wells and cultured at 37°C in 5% CO2 for 24 hours. 500 ng of plasmid was added to each well, along with 18.5 μl of Opti-MEM® I Reduced Serum Medium (ThermoFisher) and 1.5 μl of FuGENE® HD Transfection Reagent (Promega), to a final volume of 25 μl. The mixture was incubated at room temperature for 10 minutes before adding it to the cells. Cells were harvested 24 hours after transfection.

[0154] (RNA extraction, reverse transcription, sequencing) The assay was carried out in the same manner as described above for E. coli. In the following examples, unless otherwise specified, the experiments were performed in the same manner as the assay for E. coli and in these examples.

[0155] [3. Effects of domain swapping on RNA editing activity] The DYW domain can be divided into several regions based on the conservation of its amino acid sequence, but the relationship between these regions and their RNA editing activity is unknown. Here, we swapped parts of the KP1 and WW1 domains and investigated the effect on RNA editing activity in HEK239T cells (Figure 6). In a fused WW1 domain PG box and DYW to the central region containing the active site of the KP1 domain (chimKP1a SEQ ID NO: 90), nearly 50% higher U to C activity was observed compared to the KP1 domain, and C to U activity was almost completely eliminated. Domain swapping successfully improved the U to C editing performance of the KP domain.

[0156] [4. Improvement of RNA editing activity of the KP domain by introducing mutations] We aimed to improve the RNA editing activity of the KP domain by introducing various mutations into the KP domain. We designed KP2 to KP23 (sequence numbers: 68 to 89) and examined their C to U or U to C RNA editing activity in E. coli (Figure 7a) and HEK293 T cells (Figure 7b, c). KP22 (sequence number: 88) showed the highest U to C editing activity and the lowest C to U editing activity, successfully improving RNA editing activity compared to KP1.

[0157] [5. Improvement of RNA editing activity of the PG domain by introducing mutations] We aimed to improve the RNA editing activity of the PG domain by introducing various mutations into it. PG2 to PG13 (SEQ ID NOs: 41 to 53) were designed, and their C to U RNA editing activity was examined in E. coli (Figure 8a) and HEK293 T cells (Figure 8b, c). PG11 (SEQ ID NO: 50) showed the highest C to U editing activity, demonstrating successful improvement in RNA editing activity.

[0158] [6. Improvement of RNA editing activity of the WW domain by introducing mutations] We aimed to improve the RNA editing activity of the WW domain by introducing various mutations into it. WW2-WW14 were designed, and their C to U RNA editing activity was examined in E. coli (Figure 9a) and HEK293T cells (Figure 9b, c). WW11 (SEQ ID NO: 63) showed the highest C to U editing activity, successfully improving RNA editing activity compared to WW1.

[0159] [7. Human mitochondrial RNA editing using PPR proteins] [result] Mitochondria have their own unique genomes, encoding essential proteins for complexes involved in respiration and ATP production. Mutations in these proteins are known to cause various diseases, and methods for repairing these mutations are needed.

[0160] The CRISPR-Cas system has been developed as a C-to-U RNA editing tool in the cytoplasm by fusing a modified ADAR domain to the Cas protein (Abudayyeh et al. 2019 Science Vol.365, Issue 6451, pp.382-386). However, the CRISPR-Cas system consists of a protein and guide RNA, making efficient mitochondrial transport of the guide RNA difficult. On the other hand, PPR proteins can perform RNA editing with a single molecule, and generally, proteins can be delivered to mitochondria by fusing a mitochondrial localization signal sequence to the N-terminus. Therefore, to confirm whether this technology can be used for mitochondrial RNA editing, we designed PPRs targeting MT-ND2 and MT-ND5 and created genes fused with PG1 or WW1 domains (Figure 10a).

[0161] These proteins, consisting of mitochondrial target sequences (MTS) and P-DYW sequences, target the third position of codons 178 and 301 of MT-ND2 and MT-ND5, respectively, to avoid adverse effects on HEK293T cells (Figure 10a). After plasmid introduction into HEK293T cells, editing was confirmed within MT-ND2 and MT-ND5 mRNA (Figure 10b,c). These four proteins showed up to 70% editing activity against their targets, and no off-target mutations were detected in the same mRNA molecules (Figure 10bc).

[0162] [method] (Cloning for mitochondrial editing) The mitochondrial target sequence (Chin et al. 2018), 10 PPR-P and PPR-like motifs, and the DYW domain (using PG1 and WW1 for the DYW domain) from the LOC100282174 protein of Zea mays were cloned into expression plasmids under the control of the CMV promoter using Golden Gate Assembly (SEQ ID NOs: 91-94).

[0163] [List of cited sequences]

[0164] [Table 4-1]

[0165] [Table 4-2]

[0166] [Table 4-3]

[0167] [Table 4-4]

Claims

1. A composition for treating disease-related RNA, comprising nucleic acid encoding a DYW protein, or a vector containing such nucleic acid, The DYW protein, An RNA-binding domain comprising at least one PPR motif and capable of sequence-specifically binding to a target RNA, and A DYW domain consisting of one of the following polypeptides A composition containing DYW protein. • A polypeptide consisting of the sequence of SEQ ID NO: 1; a polypeptide having at least 90% sequence identity with the sequence of SEQ ID NO: 1 and possessing C-to-U editing activity; a polypeptide having a sequence in which 1 to 13 amino acids are substituted, deleted, or added to the sequence of SEQ ID NO: 1 and possessing C-to-U editing activity; or a polypeptide consisting of the sequence of SEQ ID NO: 41 • A polypeptide consisting of the sequence of SEQ ID NO: 2; a polypeptide having at least 90% sequence identity with the sequence of SEQ ID NO: 2 and possessing C-to-U editing activity; a polypeptide having a sequence in which 1 to 12 amino acids are substituted, deleted, or added to the sequence of SEQ ID NO: 2 and possessing C-to-U editing activity; or a polypeptide consisting of the sequence of SEQ ID NO: 54 Polypeptides consisting of the sequence of SEQ ID NO: 3; polypeptides having at least 90% sequence identity with the sequence of SEQ ID NO: 3 and possessing U-to-C editing activity; polypeptides having sequences in which 1 to 12 amino acids are substituted, deleted, or added to the sequence of SEQ ID NO: 3 and possessing U-to-C editing activity; or polypeptides consisting of the sequences of SEQ ID NO: 68, 69, 70, 72, 74, 76, 77, 78, 80, 81, 82, 84, 86, 87, 88, or 89 • A polypeptide consisting of the sequence of SEQ ID NO: 90; a polypeptide having at least 90% sequence identity with the sequence of SEQ ID NO: 90 and possessing U-to-C editing activity.

2. The composition according to claim 1, wherein the DYW domain consists of any one polypeptide described below. • Sequence ID: A polypeptide consisting of sequence ID 1. • Polypeptide consisting of sequence number 2 • Sequence ID: A polypeptide consisting of sequence 3 • Sequence ID: A polypeptide consisting of sequence number 90.

3. The composition according to claim 1 or 2, wherein the RNA-binding domain is of the PLS type.