Method for editing target RNA

The fusion of C-terminal DYW domains with PPR or PLS arrays in artificial proteins allows for efficient C-to-U/U-to-C RNA editing, addressing the lack of methods for converting cytidine to uridine in target RNAs, particularly in hornworts and ferns.

JP2025120313AActive Publication Date: 2025-08-15EDITFORCE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025094806
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-31
Filing Date
2025-06-06
Publication Date
2025-08-15
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

Current methods lack the ability to design molecules that can convert cytidine (C) to uridine (U) or vice versa in any target RNA using the DYW domain and artificial RNA-binding proteins, particularly for U-to-C RNA editing observed in hornworts, microphyllous plants, and some ferns.

Method used

The use of a C-terminal DYW domain from PPR-DYW proteins, fused with PPR or PLS arrays, to create artificial DYW proteins with C-to-U/U-to-C editing activity, specifically designed polypeptides such as DYW:PG, DYW:WW, and DYW:KP, which exhibit RNA editing activities.

Benefits of technology

Enables the conversion of editing target C to U or U to C in target RNAs, demonstrating RNA editing efficiency under appropriate conditions, with at least 3-5% base replacement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025120313000014
    Figure 2025120313000014
  • Figure 2025120313000015
    Figure 2025120313000015
  • Figure 2025120313000016
    Figure 2025120313000016
Patent Text Reader

Abstract

To provide a method for converting an editing target C contained in a target RNA to U or an editing target U contained in a target RNA to C.SOLUTION: A method for editing a target RNA is provided, comprising applying, to the target RNA, an artificial DYW protein containing a DYW domain consisting of any one of the following polypeptides a, b, c, and bc: a. a polypeptide having xa1 PGxa2SWIExa3-xa16HP...HxaaE...Cxa17xa18CH...DYW, having a sequence identity of at least 40% to the sequence of SEQ ID NO: 1 and having a C-to-U / U-to-C editing activity; b. a polypeptide having xb1 PGxb2SWWTDxb3-xb16HP...HxbbE...Cxb17xb18CH...DYW, having a sequence identity of at least 40% to the sequence of SEQ ID NO: 2 and having a C-to-U / U-to-C editing activity; c. a polypeptide having KPAxc1Axc2IExc3...HxccE...Cxc4xc5CH...xc6xc7xc8, having a sequence identity of at least 40% to the sequence of SEQ ID NO: 3 and having a C-to-U / U-to-C editing activity; bc. a polypeptide having xb1 PGxb2SWWTDxb3-xb16HP...HxccE...Cxc4xc5CH...Dxbc1xbc2, having a sequence identity of at least 40% to a 90-th sequence of SEQ ID NO: 2 and having a C-to-U / U-to-C editing activity (in the sequences, x represents an arbitrary amino acid, and ... represents an arbitrary polypeptide fragment).SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an RNA editing technology that uses proteins capable of binding to target RNA. The present invention is useful in a wide range of fields, including medicine (drug discovery support, treatment), agriculture (agricultural and livestock production, breeding), and chemistry (biological substance production). [Background technology]

[0002] In plant mitochondria and chloroplasts, RNA editing, in which specific bases in the genome are replaced at the RNA level, frequently occurs, and this phenomenon is known to be mediated by pentatricopeptide repeat (PPR) proteins, which are RNA-binding proteins.

[0003] PPR proteins are classified into two families, P and PLS, based on the structure of the PPR motifs that make up the protein (Non-Patent Document 1). While P class PPR proteins consist of simple repeats of the standard 35-amino acid PPR motif (P), PLS proteins contain two similar motifs, L and S, in addition to P. The PPR array (arrangement of PPR motifs) of PLS proteins consists of three PPR motifs, P1 (approximately 35 amino acids), L1 (approximately 35 amino acids), and S1 (approximately 31 amino acids), as the PLS repeat unit. At the C-terminus of P1L1S1, PLS motifs with slightly different sequences, namely P2 (35 amino acids), L2 (36 amino acids), and S2 (32 amino acids), follow. In addition to the PLS repeat unit, SS (31 amino acids) repeats may also occur. Furthermore, this final P2L2S2 motif may be followed at the C-terminal end by two PPR-like motifs called E1 and E2, and a DYW domain having a 136-amino acid cytidine deaminase domain-like sequence (Non-Patent Document 2).

[0004] The interaction between PPR proteins and RNA is regulated by the PPR code, which specifies the binding RNA base by the combination of amino acids at several positions in each PPR motif, and these amino acids bind to the corresponding nucleotide through hydrogen bonds (Patent Document 1, Patent Document 2, Non-Patent Document 3, Non-Patent Document 4, Non-Patent Document 5, Non-Patent Document 6, Non-Patent Document 7, Non-Patent Document 8). [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication WO2013 / 058404 [Patent Document 2] International Publication WO2014 / 175284 [Non-patent literature]

[0006] [Non-Patent Document 1] Lurin, C., Andres, C., Aubourg, S., Bellaoui, M., Bitton, F., Bruyere, C., Caboche, M., Debast, C., Gualberto, J., Hoffmann, B., et al. (2004). Genome-wide analysis of Arabidopsis pentatricopeptide repeat proteins reveals their essential role in organelle biogenesis. Plant Cell 16:2089-2103. [Non-patent document 2] Cheng , S. , Gutmann , B. , Zhong , X. , Ye , Y. , Fisher , MF , Bai , F. , Castleden , I. , Song , Y. , Song , B. , Huang , J. , et al. (2016). Redefining the structural motifs that determine RNA binding and RNA editing by pentatricopeptide repeat proteins in land plants. Plant J. 85:532–547.

Outdoor Tools3

Outdoor Tools 4

Direct Environment 5

Outdoor Configuration6

Direct Environment 7

Outdoor Track 8

Outdoor Tools9

[0007] C-to-U RNA editing, in which RNA bases are replaced by uridine from cytidine, commonly occurs in organelles of land plants. However, U-to-C RNA editing, in which uridine is replaced by cytidine, is also observed in hornworts, microphyllous plants, and some ferns (Non-Patent Document 9). Two different bioinformatics studies have discovered unique DYW domains in plants with U-to-C RNA editing, whose sequences differ from those of standard DYW domains (Non-Patent Document 10, Non-Patent Document 11).

[0008] Two DYW:PG proteins from the moss Physcomitrella patens have been reported to exhibit C-to-U RNA editing activity in Escherichia coli (Non-Patent Document 12). In the plant group that diverged early in the evolution of land plants, DYW domains are broadly classified into two groups. The first group is called the DYW:PG / WW group, which includes the standard DYW domain DYW:PG type and the DYW:WW type, which has an additional tryptophan (W) in the PG box. The second group is called the DYW:KP group. This DYW domain has a different sequence from the PG / WW group in the PG box and the three amino acids at the C-terminus. Furthermore, it is found only in plants with U-to-C RNA editing. A technique for changing a single base in an arbitrary RNA sequence to a specific base is useful in gene therapy and gene mutagenesis for industrial applications. Although various RNA-binding molecules have been developed to date, no method has been established for designing molecules that can convert cytidine (C) to uridine (U) in any target RNA, or vice versa, by using the DYW domain and an artificial RNA-binding protein. [Means for solving the problem]

[0009] In this study, we demonstrate that the C-terminal DYW domain of the PPR-DYW protein can be used as a modular editing domain for RNA editing. When the three DYW domains designed in this study were fused to PPR-P or PLS arrays, the DYW:PG and DYW:WW domains exhibited C-to-U RNA editing activity, while the DYW:KP domain exhibited U-to-C RNA editing activity.

[0010] The present invention provides the following: [1] A method for editing a target RNA, comprising applying an artificial DYW protein to the target RNA, the artificial DYW protein comprising a DYW domain consisting of any one of the following polypeptides: a, b, c, and bc: a.x a1 PGx a2 SWIEx a3 -xa16 HP...Hx aa E … Cx a17 x a18 A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity. b.x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx bb E … Cx b17 x b18 A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity. c. KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH...x c6 x c7 x c8 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity. bc.x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx cc E … Cx c4 x c5 CH... Dx bc1 x bc2 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity. (In the sequences, x represents any amino acid, and ... represents any polypeptide fragment.) [2] A DYW domain consisting of any one of the following polypeptides: a, b, c, and bc: a.x a1 PGx a2 SWIEx a3 -x a16 HP...Hx aa E … Cxa17 x a18 A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity. b.x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx bb E … Cx b17 x b18 A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity. c. KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH...x c6 x c7 x c8 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity. bc.x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx cc E … Cx c4 x c5 CH... Dx bc1 x bc2 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity. [3] A DYW protein comprising an RNA-binding domain containing at least one PPR motif and capable of sequence-specifically binding to a target RNA, and the DYW domain described in 2. [4] The DYW protein according to 2, wherein the RNA-binding domain is a PLS type. [5] A method for editing a target RNA, comprising: A method comprising the step of applying a DYW domain consisting of the polypeptide c or bc below to a target RNA to convert the editing target U to C. c. KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH...x c6 x c7 x c8 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity. bc.x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx cc E … Cx c4 x c5 CH... Dx bc1 x bc2 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity. [6] The method according to 5, wherein the DYW domain is fused to an RNA-binding domain that contains at least one PPR motif and is capable of sequence-specifically binding to a target RNA according to the rules of the PPR-code. [7] A composition for editing RNA in eukaryotic cells, comprising the DYW domain described in 2. [8] A nucleic acid encoding the DYW domain described in 2 or the DYW protein described in 3 or 4. [9] A vector comprising the nucleic acid described in 8.

[10] A cell (excluding a human individual) containing the vector described in 9.

[0011] [1] A method for editing a target RNA, comprising applying an artificial DYW protein containing a DYW domain consisting of any one of the following polypeptides a to c to the target RNA: a.x a1 PGx a2 SWIEx a3 -x a16 HP … HSE … Cx a17 x a18A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity. b.x b1 PGx b2 SWWTDx b3 -x b16 HP … HSE … Cx b17 x b18 A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity. c. KPAx c1 Ax c2 IEx c3 … HAE … Cx c4 x c5 CH... Dx c6 x c7 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity. (In the sequences, x represents any amino acid, and ... represents any polypeptide fragment.) [2] A DYW domain consisting of any one of the following polypeptides a to c: a.x a1 PGx a2 SWIEx a3 -x a16 HP … HSE … Cx a17 x a18 A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity. b.x b1 PGx b2 SWWTDx b3 -x b16 HP … HSE … Cx b17 x b18 A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity. c. KPAx c1 Ax c2 IEx c3 … HAE … Cxc4 x c5 CH... Dx c6 x c7 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity. [3] A DYW protein comprising an RNA-binding domain containing at least one PPR motif and capable of sequence-specifically binding to a target RNA, and the DYW domain described in 2. [4] The DYW protein according to 2, wherein the RNA-binding domain is a PLS type. [5] A method for editing a target RNA, comprising: A method comprising the step of applying a DYW domain consisting of the polypeptide of c below to a target RNA to convert the editing target U to C. c. KPAx c1 Ax c2 IEx c3 … HAE … Cx c4 x c5 CH... Dx c6 x c7 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity. [6] The method according to 5, wherein the DYW domain is fused to an RNA-binding domain that contains at least one PPR motif and is capable of sequence-specifically binding to a target RNA according to the rules of the PPR-code. [7] A composition for editing RNA in eukaryotic cells, comprising the DYW domain described in 2. [8] A nucleic acid encoding the DYW domain described in 2 or the DYW protein described in 3 or 4. [9] A vector comprising the nucleic acid described in 8.

[10] A cell (excluding a human individual) containing the vector described in 9. [Effects of the Invention]

[0012] According to the present invention, an editing target C contained in a target RNA can be converted to U, or an editing target U can be converted to C. [Brief explanation of the drawings]

[0013] [Figure 1] Approximate maximum likelihood phylogenetic tree of the C-terminal domains of DYW proteins. Phylogenetic trees of (a) DYW:PG, (b) DYW:WW, and (c) DYW:KP domains were constructed using FastTree. The DYW:PG domain was designed based on the DYW:PG protein from microphyllous plants. The clades of proteins selected for designing the DYW:WW and DYW:KP domains are indicated by black lines. The phylogenetic tree was visualized using iTOL (Letunic, I. and Bork, P. (2016) Nucleic Acids Res. 44 W242-245). Red indicates proteins from hornworts, green indicates proteins from microphyllous plants, and blue indicates proteins from ferns. [Figure 2] Structure of the PPR protein designed in this study. (a) The PPR protein contains a thioredoxin, His-tag, and TEV site at the N-terminus, followed by a PPR array consisting of a P or PLS motif, and finally a DYW domain (DYW:PG, DYW:WW, or DYW:KP). The target sequence of the PPR protein (including the RNA editing site) was inserted downstream of the stop codon. (b) The fourth and second amino acids within the P, P1, L1, S1, and P2 motifs are involved in RNA recognition. The fourth and second amino acids of L2, S2, E1, and E2 are identical to those of CLB19 (Chateigner-Boutin, AL, et al. (2008). Plant J. 56 590-602.). [Figure 3] RNA editing assay using E. coli. (a) Target base (C or T) on DNA for each PPR-DYW. (b) C-to-U RNA editing activity of PPR-DYW:PG and WW, and (c) U-to-C RNA editing activity of PPR-DYW:KP. One example is shown from three independent experiments. (d) C-to-U and (e) U-to-C RNA editing efficiencies were measured three times (except for P-DYW:KP, which was measured twice). Error bars indicate standard deviation. [Figure 4-1]The frequency of amino acid occurrence in each selected DYW domain family. Visualized using sequence logos created with WebLogo. Same as Figure 4-2 and Figure 4-3. (a) DYW:PG [Figure 4-2] (b)DYW:WW [Figure 4-3] (c)DYW:KP [Figure 5] RNA editing activity (C to U or U to C) in HEK293T cells. Sequencing results of target sites when the PLS type is fused with the PG1 or WW1 domain (a), or the KP1 domain (b). RNA editing activity in three independent experiments (c). Bar graphs show averages, and each point represents the RNA editing activity in three experiments. The left bar shows C to U editing activity, and the right bar shows U to C editing activity. [Figure 6] Improving KP domain performance through domain swapping. We examined the RNA editing activity of Chimeric KP1a, in which the central part containing the active site (HxExnCxxCH) of KP1 was left intact and the surrounding domains were swapped with the WW1 domain. Schematic diagram of domain swapping (a), actual RNA editing activity (b, c). [Figure 7] Measurement of RNA editing activity of KP domain mutants. RNA editing activity of KP1, KP2, KP3, and KP4 in E. coli (a) and HEK293T (b), and RNA editing activity of KP5 to KP23 in HEK293T (c). [Figure 8] Measurement of RNA editing activity of PG domain mutants. RNA editing activity of PG1 and PG2 in E. coli (a) and HEK293T (b). RNA editing activity of PG3 to PG13 in HEK293T. [Figure 9] Measurement of RNA editing activity of WW domain mutants. RNA editing activity of WW1 and WW2 in E. coli (a) and in HEK293T (b). RNA editing activity of WW3 to WW14 in HEK293T. [Figure 10] Human mitochondrial RNA editing. Target sequence information (a). RNA editing activity (b, c). DETAILED DESCRIPTION OF THE INVENTION

[0014] The present invention relates to a method for editing a target RNA, which comprises applying a DYW protein containing a DYW domain consisting of any one of the following polypeptides a, b, c, and bc to the target RNA: a.x a1 PGx a2 SWIEx a3 -x a16 HP...Hx aa E … Cx a17 x a18 A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 1, and having C-to-U / U-to-C editing activity. b.x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx bb E … Cx b17 x b18 A polypeptide having the sequence CH...DYW, having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity. c. KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH...x c6 x c7 x c8 and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity. bc.x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx cc E … Cx c4 x c5 CH... Dx bc1 x bc2and a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity.

[0015] In the present invention, C-to-U / U-to-C editing activity refers to the activity of converting an editing target C contained in a target RNA to U or an editing target U to C when an editing assay is performed in which a subject polypeptide is linked to the C-terminus of an RNA-binding domain capable of sequence-specifically binding to a target RNA. Conversion is sufficient if at least about 3%, preferably about 5%, of the editing target bases are replaced with the desired base under appropriate conditions.

[0016] [DYW domain] x a1 PGx a2 SWIEx a3 -x a16 HP...Hx aa E … Cx a17 x a18 CH…DYW, x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx bb E … Cx b17 x b18 CH...DYW, and KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH...x c6 x c7 x c8 Each of the x's represents an amino acid sequence. In the sequences, each x independently represents an amino acid, and each ... independently represents a polypeptide fragment of any length consisting of any amino acid sequence. In the present invention, the DYW domain can be represented by any one of these three amino acid sequences. In particular, x a1 PGx a2 SWIEx a3 -x a16 HP...Hx aa E … Cx a17 x a18The DYW domain consisting of CH … DYW is DYW:PG, x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx bb E … Cx b17 x b18 The DYW domain consisting of CH … DYW is called DYW:WW, and KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH...x c6 x c7 x c8 The DYW domain consisting of these is sometimes expressed as DYW:KP.

[0017] The DYW domain consists of an N-terminal region containing a PG box consisting of approximately 15 amino acids, a central zinc-binding domain (HxEx n CxxCH,x n is a sequence of any number n of amino acids. The zinc-binding domain has three regions: the HxE region, the CxxCH region, and the C-terminal DYW region. The zinc-binding domain can be further divided into the HxE region and the CxxCH region. These regions of each DYW domain can be represented as shown in the table below.

[0018] [Table 1]

[0019] (DYW:PG) DYW:PG is x a1 PGx a2 SWIEx a3 -x a16 HP...Hx aa E … Cx a17 x a18 A polypeptide consisting of CH...DYW. Preferably, x a1 PGx a2 SWIEx a3 -x a16 HP...Hx aa E … Cx a17 xa18 CH...DYW, and is a polypeptide that has sequence identity to the sequence of SEQ ID NO:1 (described in detail in the "Terminology" section) and has C-to-U / U-to-C editing activity. DYW:PG has the activity of converting the editing target C to U (C-to-U editing activity). SEQ ID NO:1 shows the sequence of DYW:PG, which is 136 amino acids in length and was used in the experiments shown in the Examples section of this specification. This sequence is disclosed for the first time in the present application and is therefore novel.

[0020] The total length of DYW:PG is not particularly limited as long as it can exhibit C-to-U editing activity, but is, for example, 110 to 160 amino acids in length, preferably 124 to 148 amino acids in length, more preferably 128 to 144 amino acids in length, and even more preferably 132 to 140 amino acids in length.

[0021] DYW: The area containing the PG box of PG (x a1 PGx a2 SWIEx a3 -x a16 HP): x a1 is not particularly limited as long as it can exert C-to-U editing activity as DYW:PG, but is preferably E (glutamic acid) or an amino acid similar in properties thereto, and more preferably G. x a2 is not particularly limited as long as it can exert C-to-U editing activity as DYW:PG, but is preferably C (cysteine) or an amino acid similar in properties to C, and more preferably C. x a3 -x a16 Each amino acid is not particularly limited as long as it can exert C-to-U editing activity as DYW:PG, but is preferably the same as the corresponding amino acids at positions 9 to 22 of the sequence of SEQ ID NO: 1 or an amino acid with properties similar to the corresponding amino acids, and more preferably the same as the corresponding amino acids at positions 9 to 22 of the sequence of SEQ ID NO: 1.

[0022] In one preferred embodiment, the HxE region of DYW:PG is an HSE whatever the other regions may be.

[0023] The CxxCH region of DYW:PG, i.e., Cx a17 x a18 In CH: x a17 is not particularly limited as long as it can exert C-to-U editing activity as DYW:PG, but is preferably G (glycine) or an amino acid similar in properties thereto, and more preferably G. x a18 is not particularly limited as long as it can exert C-to-U editing activity as DYW:PG, but is preferably D (aspartic acid) or an amino acid similar in properties to it, and more preferably D.

[0024] In DYW:PG, the region containing the PG box and Hx aa E bonded to the... moiety, Hx aa The ... portion connecting the E domain and the CxxCH domain, and the ... portion connecting the CxxCH domain and the DYW domain are referred to as the first linking portion, the second linking portion, and the third linking portion, respectively (the same applies to other DYW domains).

[0025] The total length of the first linking portion of DYW:PG is not particularly limited as long as it can exert C-to-U editing activity as DYW:PG, but is, for example, 39 to 47 amino acids in length, preferably 40 to 46 amino acids in length, more preferably 41 to 45 amino acids in length, and even more preferably 42 to 44 amino acids in length. The amino acid sequence of the first linking portion is not particularly limited as long as it can exert C-to-U editing activity as DYW:PG, but is preferably the same as the portion of positions 25 to 67 of the sequence of SEQ ID NO: 1, or a partial sequence of this sequence in which 1 to 22 amino acids have been substituted, deleted, or added, or a sequence that shares sequence identity with this partial sequence, and more preferably the same sequence as this partial sequence.

[0026] One preferred embodiment of the first junction of DYW:PG is a 43 amino acid long polypeptide represented by the following formula, regardless of the sequence of other parts of the DYW domain:

[0027] N a25 -N a26 -N a27 - … -N a65 -N a66 -N a67

[0028] The polypeptide preferably has a sequence identical to positions 25 to 67 of SEQ ID NO: 1 or a partial sequence thereof with multiple amino acid substitutions, and is capable of exhibiting C-to-U editing activity as DYW:PG. In this case, the amino acid substitutions are performed by substituting amino acids with a larger bits value (e.g., N) at the corresponding positions in Figure 4. a29 , N a30 , N a32 , N a33 , N a35 , N a36 , N a40 , N a44 , N a45 , N a47 , N a48 , N a52 , N a53 , N a54 , N a55 , N a58 , N a61 , N a65 , N a67 ) are the same as in FIG. 4, and it is preferable that the substitution is carried out so that other amino acids are substituted.

[0029] The total length of the second linking portion of DYW:PG is not particularly limited as long as it can exert C-to-U editing activity as DYW:PG, but is, for example, 21 to 29 amino acids in length, preferably 22 to 28 amino acids in length, more preferably 23 to 27 amino acids in length, and even more preferably 24 to 26 amino acids in length. The amino acid sequence of the second linking portion is not particularly limited as long as it can exert C-to-U editing activity as DYW:PG, but is preferably the same as positions 71 to 95 of SEQ ID NO: 1, or a partial sequence thereof in which 1 to 13 amino acids have been substituted, deleted, or added, or a sequence having sequence identity to the partial sequence, and more preferably the same sequence as the partial sequence.

[0030] A preferred embodiment of the second junction of DYW:PG is a 25 amino acid long polypeptide represented by the following formula, regardless of the sequence of the other parts of the DYW domain:

[0031] N a71 -N a72 -N a73 - … -N a93 -N a94 -N a95

[0032] The polypeptide preferably has a sequence identical to the portion of positions 71 to 95 of SEQ ID NO: 1 or a partial sequence thereof with multiple amino acid substitutions, and is capable of exhibiting C-to-U editing activity as DYW:PG. In this case, the amino acid substitutions are performed by substituting amino acids with a larger bits value (e.g., N) at the corresponding positions in Figure 4. a71 , N a72 , N a73 , N a76 , N a77 , N a78 , N a79 , N a81 , N a82 , N a86 , N a88 , N a89 , N a91 , N a92 , N a93 , N a94) are the same as in FIG. 4, and it is preferable that the substitution is carried out so that other amino acids are substituted.

[0033] The total length of the third linking portion of DYW:PG is not particularly limited as long as it can exert the C-to-U editing activity as DYW:PG, but is, for example, 29 to 37 amino acids in length, preferably 30 to 36 amino acids in length, more preferably 31 to 35 amino acids in length, and even more preferably 32 to 34 amino acids in length. The amino acid sequence of the third linking portion is not particularly limited as long as it can exert the C-to-U editing activity as DYW:PG, but is preferably the same as the portion of positions 101 to 133 of the sequence of SEQ ID NO: 1, or a partial sequence of this sequence in which 1 to 17 amino acids have been substituted, deleted, or added, or a sequence that shares sequence identity with this partial sequence, and more preferably the same sequence as this partial sequence.

[0034] A preferred embodiment of the third junction of DYW:PG is a 33 amino acid long polypeptide represented by the following formula, regardless of the sequence of other parts of the DYW domain:

[0035] N a101 -N a102 -N a103 - … -N a131 -N a132 -N a133

[0036] The polypeptide preferably has a sequence identical to positions 101 to 133 of SEQ ID NO: 1 or a partial sequence thereof with multiple amino acid substitutions, and is capable of exhibiting C-to-U editing activity as DYW:PG. In this case, the amino acid substitutions are performed by substituting amino acids with a larger bits value (e.g., N) at the corresponding positions in Figure 4. a102 , N a104 , N a107 , N a112 , N a114 , N a117 , N a118 , N a121 , N a122 , N a123 , N a124 , N a125 , Na128 , N a130 , N a131 , N a132 ) are the same as in FIG. 4, and it is preferable that the substitution is carried out so that other amino acids are substituted.

[0037] RNA editing activity can be improved by introducing a mutation into the PG domain consisting of the sequence of SEQ ID NO: 1. A preferred example of such a mutated domain is PG11 (a polypeptide consisting of the amino acid sequence of SEQ ID NO: 50) shown in the Examples section of this specification.

[0038] (DYW:WW) DYW:WW is x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx bb E … Cx b17 x b18 A polypeptide consisting of CH...DYW. Preferably, x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx bb E … Cx b17 x b18 It is a polypeptide having the sequence CH...DYW, which has sequence identity to the sequence of SEQ ID NO: 2, and has C-to-U / U-to-C editing activity. DYW:WW has the activity of converting the editing target C to U (C-to-U editing activity). SEQ ID NO: 2 shows the sequence of DYW:WW, which is 137 amino acids in length and was used in the experiments shown in the Examples section of this specification. This sequence is disclosed for the first time in the present application and is therefore novel.

[0039] The total length of DYW:WW is not particularly limited as long as it can exhibit C-to-U editing activity, but is, for example, 110 to 160 amino acids in length, preferably 125 to 149 amino acids in length, more preferably 129 to 145 amino acids in length, and even more preferably 133 to 141 amino acids in length.

[0040] The region containing the PG box of DYW:WW, i.e., x b1 PGx b2 SWWTDx b3 -x b16 In the HP, the portion consisting of WTD may be WSD.

[0041] In the area containing the PG box of DYW:WW: x b1 is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is preferably K (lysine) or an amino acid similar in properties to K, and more preferably K. x b2 is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is preferably Q (glutamine) or an amino acid similar in properties thereto, and more preferably Q. x b3 -x b16 Each amino acid is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is preferably the same as the corresponding amino acids at positions 10 to 23 of the sequence of SEQ ID NO:2 or an amino acid with similar properties to the corresponding amino acids, and more preferably the same as the corresponding amino acids at positions 10 to 23 of the sequence of SEQ ID NO:2.

[0042] In one preferred embodiment, the HxE region of DYW:WW is an HSE, whatever the sequence of the other portions.

[0043] DYW:WW CxxCH region, i.e., Cx b17 x b18 In CH: x b17 is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is preferably D (aspartic acid) or an amino acid similar in properties thereto, and more preferably D. x b18 is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is preferably D or an amino acid similar thereto, and more preferably D.

[0044] The total length of the first linking portion of DYW:WW is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is, for example, 39 to 47 amino acids in length, preferably 40 to 46 amino acids in length, more preferably 41 to 45 amino acids in length, and even more preferably 42 to 44 amino acids in length. The amino acid sequence of the first linking portion is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is preferably the same as positions 25 to 67 of the sequence of SEQ ID NO:2, or a partial sequence thereof in which 1 to 22 amino acids have been substituted, deleted, or added, or a sequence having sequence identity to the partial sequence, and more preferably the same sequence as the partial sequence.

[0045] One preferred embodiment of the first junction of DYW:WW is a 43 amino acid long polypeptide represented by the following formula, regardless of the sequence of the other parts of the DYW domain:

[0046] N b26 -N b27 -N b28 - … -N b66 -N b67 -N b68

[0047] The polypeptide preferably has a sequence identical to the portion of positions 26 to 68 of the sequence of SEQ ID NO: 2 or a partial sequence thereof with multiple amino acid substitutions, and is capable of exhibiting C-to-U editing activity as DYW:PG. In this case, the amino acid substitutions are performed by substituting amino acids with a larger bits value (e.g., N) at the corresponding positions in Figure 4. b26 , N b30 , N b33 , N b34 , N b37 , N b41 , N b45 , N b46 , N b48 , N b49 , N b51 , N b52 , N b53 , N b55 , N b56 , N b57 , Nb59 , N b61 , N b62 , N b63 , N b64 , N b66 , N b67 , N b68 ) are the same as in FIG. 4, and it is preferable that the substitution is carried out so that other amino acids are substituted.

[0048] The total length of the second linking portion of DYW:WW is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is, for example, 21 to 29 amino acids in length, preferably 22 to 28 amino acids in length, more preferably 23 to 27 amino acids in length, and even more preferably 24 to 26 amino acids in length. The amino acid sequence of the second linking portion is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is preferably the same as positions 71 to 95 of the sequence of SEQ ID NO:2, or a partial sequence thereof in which 1 to 13 amino acids have been substituted, deleted, or added, or a sequence having sequence identity to the partial sequence, and more preferably the same sequence as the partial sequence.

[0049] A preferred embodiment of the second junction of DYW:WW is a 25 amino acid long polypeptide represented by the following formula, regardless of the sequence of the other parts of the DYW domain:

[0050] N b72 -N b73 -N b74 - … -N b94 -N b95 -N b96

[0051] The polypeptide preferably has a sequence identical to the portion of positions 72 to 96 of SEQ ID NO: 2 or a partial sequence thereof with multiple amino acid substitutions, and is capable of exhibiting C-to-U editing activity as DYW:WW. In this case, the amino acid substitutions are performed by substituting amino acids with a larger bits value (e.g., N) at the corresponding positions in Figure 4. b72 , N b73 , N b74 , N b75 , Nb77 , N b78 , N b79 , N b81 , N b82 , N b84 , N b88 , N b89 , N b90 , N b91 , N b92 , N b93 , N b94 , N b95 , N b96 ) are the same as in FIG. 4, and it is preferable that the substitution is carried out so that other amino acids are substituted.

[0052] The total length of the third linkage of DYW:WW is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is, for example, 29 to 37 amino acids in length, preferably 30 to 36 amino acids in length, more preferably 31 to 35 amino acids in length, and even more preferably 32 to 34 amino acids in length. The amino acid sequence of the third linkage is not particularly limited as long as it can exert C-to-U editing activity as DYW:WW, but is preferably the same as the portion of positions 101 to 133 of the sequence of SEQ ID NO:2, or a partial sequence of this sequence in which 1 to 17 amino acids have been substituted, deleted, or added, or a sequence that shares sequence identity with this partial sequence, and more preferably the same sequence as this partial sequence.

[0053] A preferred embodiment of the third junction of DYW:WW is a 33 amino acid long polypeptide represented by the following formula, regardless of the sequence of the other parts of the DYW domain:

[0054] N b102 -N b103 -N b104 - … -N b132 -N b133 -N b134

[0055] The polypeptide preferably has a sequence identical to the portion of positions 102 to 134 of SEQ ID NO: 2 or a partial sequence thereof with multiple amino acid substitutions, and is capable of exhibiting C-to-U editing activity as DYW:WW. In this case, the amino acid substitutions are performed by substituting amino acids with a larger bits value (e.g., N) at the corresponding positions in Figure 4. b104 , N b105 , N b107 , N b108 , N b109 , N b110 , N b111 , N b113 , N b115 , N b116 , N b117 , N b118 , N b119 , N b121 , N b122 , N b123 , N b124 , N b126 , N b129 , N b131 , N b132 , N b133 N b134 ) are the same as in FIG. 4, and it is preferable that the substitution is carried out so that other amino acids are substituted.

[0056] RNA editing activity can be improved by introducing a mutation into the WW domain consisting of the sequence of SEQ ID NO: 2. Preferred examples of domains into which such mutations have been introduced are WW2 to WW11 and WW13 shown in the Examples section of this specification, and WW11 (a polypeptide consisting of the amino acid sequence of SEQ ID NO: 63) has particularly high editing activity.

[0057] (DYW:KP) DYW:KP is KPAx c1 Ax c2 IEx c3 … Hx cc E … Cx c4 x c5 CH...x c6 x c7 x c8 Preferably, KPAx is a polypeptide consisting of c1 Ax c2 IEx c3… Hx cc E … Cx c4 x c5 CH...x c6 x c7 x c8 DYW:KP is a polypeptide having sequence identity to the sequence of SEQ ID NO: 3 and having C-to-U / U-to-C editing activity. DYW:KP has the activity of converting the editing target U to C (U-to-C editing activity). SEQ ID NO: 3 shows the sequence of DYW:KP, which is 133 amino acids in length and was used in the experiments shown in the Examples section of this specification. This sequence is disclosed for the first time in the present application and is therefore novel.

[0058] The total length of DYW:KP is not particularly limited as long as it can exhibit U-to-C editing activity, but is, for example, 110 to 160 amino acids in length, preferably 121 to 145 amino acids in length, more preferably 125 to 141 amino acids in length, and even more preferably 129 to 137 amino acids in length.

[0059] The region containing the PG box of DYW:KP, i.e., KPAx c1 Ax c2 IEx c3 In: x c1 is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably S (serine) or an amino acid similar in properties, and more preferably S. x c2 is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably L (leucine) or an amino acid similar in properties to it, and more preferably L. x c3 is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably V (valine) or an amino acid similar in properties thereto, and more preferably V.

[0060] In one preferred embodiment, the HxE region of DYW:KP is HAE, whatever the sequence of the other portions.

[0061] DYW:KP CxxCH region, i.e., Cx c4 x c5 In CH: x c4 is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably N (asparagine) or an amino acid similar in properties to asparagine, and more preferably N. x b5 is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably D (aspartic acid) or an amino acid similar in properties thereto, and more preferably D.

[0062] DYW: The part of KP that corresponds to DYW, i.e., x c6 x c7 x c8 In: x c6 is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably D (aspartic acid) or an amino acid similar in properties thereto, and more preferably D. x c7 is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably M (methionine) or an amino acid similar in properties thereto, and more preferably M. x b8 is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably F (phenylalanine) or an amino acid similar in properties thereto, and more preferably F. In one preferred embodiment, x c6 x c7 x c8 Dx c7 x c8 is. In another preferred embodiment, x c6 x c7 x c8 is GRP, whatever the sequence of the other parts.

[0063] The total length of the first linking portion of DYW:KP is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is, for example, 51 to 59 amino acids in length, preferably 52 to 58 amino acids in length, more preferably 53 to 57 amino acids in length, and even more preferably 54 to 56 amino acids in length. The amino acid sequence of the first linking portion is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably the same as the portion of positions 10 to 64 of the sequence of SEQ ID NO: 3, or a partial sequence of this sequence in which 1 to 28 amino acids have been substituted, deleted, or added, or a sequence that shares sequence identity with this partial sequence, and more preferably the same sequence as this partial sequence.

[0064] One preferred embodiment of the first junction of DYW:KP is a 55 amino acid long polypeptide represented by the following formula, regardless of the sequence of other parts of the DYW domain:

[0065] N c10 -N c11 -N c12 - … -N c62 -N c63 -N c64

[0066] The polypeptide preferably has a sequence identical to positions 10 to 64 of SEQ ID NO: 3 or a partial sequence thereof with multiple amino acid substitutions, and is capable of exhibiting U-to-C editing activity as DYW::KP. In this case, the amino acid substitutions are performed by substituting amino acids with a larger bits value (e.g., N) at the corresponding positions in Figure 4. c10 , N c13 , N c14 , N c15 , N c16 , N c17 , N c18 , N c19 , N c25 , N c26 , N c29 , N c30 , N c33 , N c34 , N c36 , N c38 , N c41 , Nc42 , N c44 , N c45 , N c47 , N 49 , N c55 , N c58 , N c59 , N c62 , N c63 , N c64 ) are the same as in FIG. 4, and it is preferable that the substitution is carried out so that other amino acids are substituted.

[0067] The total length of the second linking portion of DYW:KP is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is, for example, 21 to 29 amino acids in length, preferably 22 to 28 amino acids in length, more preferably 23 to 27 amino acids in length, and even more preferably 24 to 26 amino acids in length. The amino acid sequence of the second linking portion is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably the same as the portion of positions 68 to 92 of the sequence of SEQ ID NO: 3, or a partial sequence of this sequence in which 1 to 13 amino acids have been substituted, deleted, or added, or a sequence that shares sequence identity with this partial sequence, and more preferably the same sequence as this partial sequence.

[0068] A preferred embodiment of the second junction of DYW:KP is a 25 amino acid long polypeptide represented by the following formula, regardless of the sequence of other parts of the DYW domain:

[0069] N c68 -N c69 -N c70 - … -N c90 -N c91 -N c92

[0070] The polypeptide preferably has a sequence identical to the portion of positions 68 to 92 of SEQ ID NO: 3 or a partial sequence thereof with multiple amino acid substitutions, and is capable of exhibiting U-to-C editing activity as DYW:KP. In this case, the amino acid substitutions are performed by substituting amino acids with a larger bits value (e.g., N) at the corresponding positions in Figure 4. c68 , Nc70 , N c71 , N c72 , N c73 , N c74 , N c75 , N c76 , N c77 , N c78 , N c79 , N c80 , N c81 , N c83 , N c84 , N c85 , N c86 , N c87 , N c88 , N c89 , N c90 , N c91 , N c92 ) are the same as in FIG. 4, and it is preferable that the substitution is carried out so that other amino acids are substituted.

[0071] The total length of the third linking part of DYW:KP is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is, for example, 29 to 37 amino acids in length, preferably 30 to 36 amino acids in length, more preferably 31 to 35 amino acids in length, and even more preferably 32 to 34 amino acids in length. The amino acid sequence of the third linking part is not particularly limited as long as it can exert U-to-C editing activity as DYW:KP, but is preferably the same as the part at positions 98 to 130 of the sequence of SEQ ID NO: 3, or a partial sequence thereof in which 1 to 17 amino acids have been substituted, deleted, or added, or a sequence having sequence identity to the partial sequence, and more preferably the same sequence as the partial sequence.

[0072] A preferred embodiment of the third junction of DYW:KP is a 33 amino acid long polypeptide represented by the following formula, regardless of the sequence of other parts of the DYW domain:

[0073] N c98 -N c99 -N c100 - … -N c128 -N c129 -N c130

[0074] The polypeptide preferably has a sequence identical to the portion of positions 98 to 130 of SEQ ID NO: 3 or a partial sequence thereof with multiple amino acid substitutions, and is capable of exhibiting U-to-C editing activity as DYW:KP. In this case, the amino acid substitutions are performed by substituting amino acids with a larger bits value (e.g., N) at the corresponding positions in Figure 4. c98 , N c100 , N c101 , N c103 , N c104 , N c105 , N c107 , N c109 , N c110 , N c111 , N c112 , N c114 , N c115 , N c117 , N c118 , N c119 , N c120 , N c121 , N c122 , N c123 , N c124 , N c125 , N c127 , N c129 , N c130 ) are the same as in FIG. 4, and it is preferable that the substitution is carried out so that other amino acids are substituted.

[0075] Editing activity can be improved by introducing a mutation into the KP domain consisting of the sequence of SEQ ID NO: 3. Preferred examples of domains into which such mutations have been introduced are KP2 to 23 (SEQ ID NOs: 68 to 89) shown in the Examples section of this specification. KP22 (SEQ ID NO: 88) has the highest U to C editing activity and the lowest C to U editing activity, and thus has improved RNA editing activity compared to the KP domain consisting of the sequence of SEQ ID NO: 3.

[0076] (Chimeric DYW) The DYW domain can be divided into several regions based on the conservation of the amino acid sequence. Chimeric DYWs that have been exchanged with each of these regions can improve C-to-U editing activity or U-to-C editing activity. In one preferred embodiment, the region containing the PG box of DYW:WW, i.e., xb1 PGx b2 SWWTDx b3 -x b16 HP and DYW, other areas of DYW:KP (... Hx cc E … Cx c4 x c5 CH ...) x b1 , x b2 , x b3 -x b16 The first linking portion, the second linking portion, and the third linking portion are as described above for DYW:WW. The DYW portion is Dx bc1 x bc2 The total length is not particularly limited as long as it can exhibit U-to-C editing activity, but is, for example, 110 to 160 amino acids in length, preferably 125 to 149 amino acids in length, more preferably 129 to 145 amino acids in length, and even more preferably 133 to 141 amino acids in length.

[0077] One preferred chimeric domain is x b1 PGx b2 SWWTDx b3 -x b16 HP...Hx cc E … Cx c4 x c5 CH... Dx bc1 x bc2 and consists of a polypeptide having at least 40% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U / U-to-C editing activity.

[0078] One particularly preferred Chimeric domain is a polypeptide that has sequence identity with the sequence of SEQ ID NO: 90 and has U-to-C editing activity. SEQ ID NO: 90 shows the sequence of the Chimeric domain used in the experiments described in the Examples section of the present specification. This domain has been observed to have higher U to C editing activity than DYW:KP consisting of the sequence of SEQ ID NO: 3, and further has almost no C to U editing activity.

[0079] (Comparison with known sequences) As mentioned above, the DYW domain consisting of the sequences of SEQ ID NOs: 1, 2, and 3 is novel. Below, the results of an investigation using the PPR database (https: / / ppr.plantenergy.uwa.edu.au / onekp / ; Non-Patent Document 11 cited above) are shown.

[0080] [Table 2-1]

[0081] Therefore, the present invention also provides any one of the following polypeptides d to i: d. a polypeptide having greater than 78% sequence identity to the sequence of SEQ ID NO: 1, preferably greater than 80%, more preferably greater than 85%, even more preferably greater than 90%, even more preferably greater than 95%, and even more preferably greater than 97%, and having C-to-U editing activity; e. a polypeptide having greater than 84% sequence identity, preferably greater than 85%, more preferably greater than 90%, even more preferably greater than 95%, and even more preferably greater than 97% sequence identity with the sequence of SEQ ID NO: 2, and having C-to-U editing activity; f. A polypeptide having greater than 86% sequence identity with the sequence of SEQ ID NO: 3, preferably 87% or more, more preferably 90% or more, even more preferably 95% sequence identity, and even more preferably 97% sequence identity, and having U-to-C editing activity. g. A polypeptide having a sequence in which 1 to 29, preferably 1 to 25, more preferably 1 to 21, even more preferably 1 to 17, even more preferably 1 to 13, even more preferably 1 to 9, and even more preferably 1 to 5 amino acids have been substituted, deleted, or added in the sequence of SEQ ID NO: 1, and which has C-to-U editing activity; h. A polypeptide having a sequence in which 1 to 21, preferably 1 to 18, more preferably 1 to 15, even more preferably 1 to 12, even more preferably 1 to 9, and still more preferably 1 to 6 amino acids in the sequence of SEQ ID NO: 2 have been substituted, deleted, or added, and which has C-to-U editing activity; i. A polypeptide having a sequence in which 1 to 18, preferably 1 to 16, more preferably 1 to 14, even more preferably 1 to 12, even more preferably 1 to 10, even more preferably 1 to 8, and even more preferably 1 to 6 amino acids have been substituted, deleted, or added in the sequence of SEQ ID NO: 3, and having U-to-C editing activity.

[0082] The alignment of these sequences is shown below. The alignment was created using AliView (Larsson, A. (2014). AliView: a fast and lightweight alignment viewer and editor for large data sets. Bioinformatics30(22): 3276-3278. http: / / dx.doi.org / 10.1093 / bioinformatics / btu531). Identical amino acids are represented by dots.

[0083] [Table 2-2]

[0084] [RNA binding domain] In the present invention, when converting an editing target using a DYW domain, an RNA binding protein is used as an RNA binding domain to target the RNA containing the editing target.

[0085] A preferred example of an RNA-binding protein used as an RNA-binding domain is a PPR protein composed of a PPR motif.

[0086] (PPR motif) Unless otherwise specified, a PPR motif refers to a polypeptide consisting of 30 to 38 amino acids whose E value obtained by analyzing amino acid sequences using online protein domain search programs (PF01535 in Pfam and PS51375 in Prosite) is a predetermined value or less (preferably E-03). The position numbers of the amino acids constituting the PPR motif as defined herein are nearly synonymous with PF01535, but correspond to the position of the amino acid in PS51375 minus 2 (e.g., position 1 in the present invention corresponds to position 3 in PS51375). However, when referring to the amino acid at position "ii" (-2), it refers to the second amino acid from the end (C-terminal) of the amino acid constituting the PPR motif, or the amino acid two amino acids N-terminal to amino acid 1 of the next PPR motif, i.e., the -2 amino acid. If the next PPR motif is not clearly identified, the amino acid two amino acids before the first amino acid in the next helix structure is designated as "ii." For more information on Pfam, see http: / / pfam.sanger.ac.uk / . For more information on Prosite, see http: / / www.expasy.org / prosite / .

[0087] The conserved amino acid sequence of the PPR motif is not highly conserved at the amino acid level, but the two α-helices in the secondary structure are well conserved. A typical PPR motif consists of 35 amino acids, but its length varies from 30 to 38 amino acids.

[0088] More specifically, the PPR motif is composed of a polypeptide of 30 to 38 amino acids in length represented by formula 1.

[0089] [ka]

[0090] During the ceremony: Helix A is a 12 amino acid long portion capable of forming an α-helical structure and is represented by Formula 2:

[0091] [ka]

[0092] In formula 2, A1~A 12 each independently represents an amino acid; X is absent or a moiety consisting of 1 to 9 amino acids in length; Helix B is a portion consisting of 11 to 13 amino acids that can form an α-helical structure; L is a moiety of formula 3 that is 2 to 7 amino acids in length;

[0093] [ka]

[0094] In Formula 3, each amino acid is numbered from the C-terminus, as "i" (-1), "ii" (-2), etc. However, L iii ~L vii may not exist.

[0095] (PPR-code) In the PPR motif, the combination of the three amino acids 1, 4, and ii is important for specific binding to a base, and this combination determines which base will bind. The relationship between the combination of the three amino acids 1, 4, and ii and the bases that can bind is known as the PPR-code (see Patent Document 2, cited above), and is as follows:

[0096] (1) A1, A4, and L ii When the combination of these three amino acids is, in order, valine, asparagine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to U, then C, and then A or G. (2) A1, A4, and L iiWhen the three amino acids in the PPR motif are, in order, valine, threonine, and asparagine, the PPR motif binds strongly to A, then G, then C, but not to U, showing selective RNA base binding ability. (3) A1, A4, and L ii When the combination of these three amino acids is, in order, valine, asparagine, asparagine, the PPR motif has selective RNA base-binding ability, binding strongly to C, then to A or U, but not to G. (4) A1, A4, and L ii When the combination of these three amino acids is, in order, glutamic acid, glycine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to G but not to A, U, or C. (5) A1, A4, and L ii When the three amino acid combinations are isoleucine, asparagine, and asparagine, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to C, then U, then A, but not to G. (6) A1, A4, and L ii When the three amino acid combinations are valine, threonine, and aspartic acid, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to G, then to U, but not to A or C. (7) A1, A4, and L ii When the three amino acid combinations are lysine, threonine, and aspartic acid, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to G, then to A, but not to U or C. (8) A1, A4, and L ii When the combination of these three amino acids is, in order, phenylalanine, serine, and asparagine, the PPR motif has selective RNA base-binding ability, binding strongly to A, then C, and then G and U. (9) A1, A4, and L iiWhen the three amino acid combinations are valine, asparagine, and serine, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to C, then to U, but not to A or G. (10) A1, A4, and L ii When the combination of these three amino acids is, in order, phenylalanine, threonine, and asparagine, the PPR motif has selective RNA base-binding ability, binding strongly to A but not to G, U, or C. (11) A1, A4, and L ii When the combination of these three amino acids is, in order, isoleucine, asparagine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to U, then to A, but not to G or C. (12) A1, A4, and L ii When the combination of these three amino acids is threonine, threonine, and asparagine, respectively, the PPR motif has selective RNA base binding ability, binding strongly to A but not to G, U, or C. (13) A1, A4, and L ii When the three amino acids in the PPR motif are, in order, isoleucine, methionine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to U, then C, but not A or G. (14) A1, A4, and L ii When the three amino acid combinations are phenylalanine, proline, and aspartic acid, respectively, the motif is called PPR. It binds strongly to U, then C, but not A or G, and has selective RNA base binding ability. (15) A1, A4, and L ii When the combination of these three amino acids is tyrosine, proline, and aspartic acid, respectively, the PPR motif has selective RNA base binding ability, binding strongly to U but not to A, G, or C. (16) A1, A4, and L iiWhen the combination of these three amino acids is leucine, threonine, and aspartic acid, respectively, the PPR motif has selective RNA base binding ability, binding strongly to G but not to A, U, or C.

[0097] (P array) PPR proteins are classified into two families, P and PLS, depending on the structure of the PPR motif they comprise. P-type PPR proteins are composed of simple repeats (P arrays) of a standard 35-amino acid PPR motif (P). The DYW domain of the present invention can be used by linking it to a P-type PPR protein.

[0098] (PLS array) The PPR motif arrangement in PLS-type PPR proteins consists of a repeating unit of three PPR motifs, P1, L1, and S1, followed by P2, L2, and S2 motifs at the C-terminus of the repeat. This final P2L2S2 motif may be further followed by two PPR-like motifs, E1 and E2, and a DYW domain at the C-terminus.

[0099] The total length of P1 is not particularly limited as long as it can bind to the target base, but is, for example, 33 to 37 amino acids, preferably 34 to 36 amino acids, and more preferably 35 amino acids. The total length of L1 is not particularly limited as long as it can bind to the target base, but is, for example, 33 to 37 amino acids, preferably 34 to 36 amino acids, and more preferably 35 amino acids. The total length of S1 is not particularly limited as long as it can bind to the target base, but is, for example, 30 to 33 amino acids, preferably 30 to 32 amino acids, and more preferably 31 amino acids.

[0100] The total length of P2 is not particularly limited as long as it can bind to the target base, but is, for example, 33 to 37 amino acids, preferably 34 to 36 amino acids, and more preferably 35 amino acids. The total length of L2 is not particularly limited as long as it can bind to the target base, but is, for example, 34 to 38 amino acids, preferably 35 to 37 amino acids, and more preferably 36 amino acids. The total length of S2 is not particularly limited as long as it can bind to the target base, but is, for example, 30 to 34 amino acids, preferably 31 to 33 amino acids, and more preferably 32 amino acids. SEQ ID NOs: 17, 22, and 27 show the sequences of P2 used in the Examples section of this specification. SEQ ID NOs: 18, 23, and 28 show the sequences of L2 used in the Examples section of this specification. SEQ ID NOs: 19, 24, and 29 show the sequences of S2 used in the Examples section of this specification.

[0101] The total length of E1 is not particularly limited as long as it can bind to the target base, but is, for example, 32 to 36 amino acids in length, preferably 33 to 35 amino acids in length, and more preferably 34 amino acids in length. SEQ ID NOs: 20, 25, and 30 show the sequences of E1 used in the Examples section of this specification.

[0102] The total length of E2 is not particularly limited as long as it can bind to the target base, but is, for example, 30 to 34 amino acids in length, preferably 31 to 33 amino acids in length, and more preferably 33 amino acids in length. SEQ ID NOs: 21, 26, and 31 show the sequences of E2 used in the Examples section of this specification.

[0103] In a PLS-type PPR protein, the P1L1S1 repeat portion and the portion up to P2 can be designed according to the sequence of the target RNA, following the PPR-code rules described above.

[0104] The S2 motif correlates with the nucleotide corresponding to amino acid ii (N at position 31 in SEQ ID NO: 19) (Non-Patent Document 8, cited above). Furthermore, the C or U four positions to the right of the target base in the S2 motif is the target base for editing by the DYW domain, and this should be kept in mind when incorporating it into a PLS-type PPR protein.

[0105] In the E1 motif, correlation with the nucleotide is only found in the fourth amino acid (the fourth G in SEQ ID NO: 20) (Ruwe et al. (2019) New Phytol. 222 218-229.).

[0106] The fourth (V at position 4 in SEQ ID NO: 21) and last (K at position 33 in SEQ ID NO: 21) amino acids in the E2 motif are highly conserved and are not involved in specific PPR-RNA recognition (Non-Patent Document 2 cited above).

[0107] The number of P1L1S1 repeats is not particularly limited as long as it can bind to the target nucleotide sequence, but is, for example, 1 to 5, preferably 2 to 4, and more preferably 3. In principle, even a single unit (3 repeats) can be used. MEF8 (L1-S1-P2-L2-S2-E-DYW), which consists of five PPR motifs, is known to be involved in editing at approximately 60 sites.

[0108] In natural PPR proteins, the first and last P1L1S1 have clear differences in amino acid residues at specific positions and are different from internal P1L1S1. From the perspective of designing an artificial PLS array that is as close as possible to a naturally occurring one, it is advisable to design three types of P1L1S1 according to the position of the PPR motif: the first (N-terminal) P1L1S1, the internal P1L1S1, and the last (C-terminal) P1L1S1 located just before P2L2S2. Note that in nature, in addition to those composed of PLS repeating units, there are also cases where SS (31 amino acids) are repeated, and these can also be used in the present invention.

[0109] [DYW protein] The present invention provides a DYW protein for editing a target RNA, which comprises an RNA-binding domain that contains at least one PPR motif and is capable of sequence-specifically binding to the target RNA according to the rules of the PPR-code, and a DYW domain that is any one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0110] Such DYW proteins may be artificial. "Artificial" refers to proteins that are not naturally occurring but artificially synthesized. Examples of artificial proteins include those that have a DYW domain with a sequence not found in nature, those that have a PPR-binding domain with a sequence not found in nature, those that have an RNA-binding domain and a DYW domain with a combination not found in nature, and those that have an additional portion, such as a nuclear localization signal or a human mitochondrial localization signal, that is not present in naturally occurring DYW proteins derived from plants. Examples of nuclear localization signal sequences include PKKKRKV (SEQ ID NO: 32) derived from SV40 large T antigen and KRPAATKKAGQAKKKK (SEQ ID NO: 33), an NLS of nucleoplasmin.

[0111] The number of PPR motifs can be appropriately determined depending on the sequence of the target RNA. The number of PPR motifs may be at least one, but may be two or more. It is known that two PPR motifs are sufficient to bind to RNA (Nucleic Acids Research, 2012, Vol. 40, No. 6, 2712-2723).

[0112] A preferred embodiment of the DYW protein is as follows: An artificial DYW protein for editing a target RNA, comprising at least one, preferably 2 to 25, more preferably 5 to 20, and even more preferably 10 to 18 PPR motifs, an RNA-binding domain capable of sequence-specifically binding to a target RNA according to the PPR-code rules, and a DYW domain which is any one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0113] Another preferred embodiment of the DYW protein is as follows: A DYW protein for editing target RNA, which comprises at least one, preferably 2 to 25, more preferably 5 to 20, and even more preferably 10 to 18 PPR motifs, an RNA-binding domain (preferably an RNA-binding domain that is a PLS-type PPR protein) that can bind sequence-specifically to an animal's target RNA according to the PPR-code rules, and a DYW domain that is any one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0114] The DYW domain, PPR protein as an RNA-binding domain, and DYW protein of the present invention can be prepared in relatively large quantities by methods well known to those skilled in the art. Such methods may include determining the nucleic acid sequence encoding the domain or protein of interest from its amino acid sequence, cloning it, and preparing a transformant that produces the domain or protein of interest.

[0115] [Use of DYW protein] (nucleic acids, vectors, and cells encoding DYW proteins, etc.) The present invention also provides the above-mentioned PPR motif, DYW protein, or nucleic acid encoding the DYW protein, and vectors containing the nucleic acid (e.g., vectors for amplification, expression vectors). Vectors also include viral vectors. Vectors for amplification can use Escherichia coli or yeast as hosts. As used herein, an expression vector refers to a vector containing, from upstream, DNA having a promoter sequence, DNA encoding a desired protein, and DNA having a terminator sequence, although these do not necessarily have to be arranged in this order as long as the desired function is exerted. In the present invention, various vectors that are commonly used by those skilled in the art can be recombined and used.

[0116] Specifically, the present invention provides a nucleotide sequence encoding a DYW protein, which includes an RNA-binding domain (preferably an RNA-binding domain that is a PLS-type PPR protein) that contains at least one PPR motif and is capable of sequence-specifically binding to a target RNA (preferably an animal target RNA) according to the rules of the PPR code, and a DYW domain that is any one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0117] More specifically, the present invention provides a vector for editing a target RNA, which comprises an RNA-binding domain (preferably an RNA-binding domain that is a PLS-type PPR protein) that includes at least one PPR motif and is capable of sequence-specifically binding to a target RNA (preferably an animal target RNA) according to the rules of the PPR code, and a nucleotide sequence encoding a DYW protein that includes a DYW domain that is any one of the aforementioned DYW:PG, DYW:WW, or DYW:KP.

[0118] The DYW proteins of the present invention can function in eukaryotic cells (e.g., animal, plant, microorganism (yeast, etc.), protist) and particularly in animal cells (in vitro or in vivo). Examples of animal cells into which the DYW proteins of the present invention or vectors expressing the DYW proteins can be introduced include cells derived from humans, monkeys, pigs, cows, horses, dogs, cats, mice, and rats. Examples of cultured cells into which the DYW proteins of the present invention or vectors expressing the DYW proteins can be introduced include, but are not limited to, Chinese hamster ovary (CHO) cells, COS-1 cells, COS-7 cells, VERO (ATCC CCL-81) cells, BHK cells, canine kidney-derived MDCK cells, hamster AV-12-664 cells, HeLa cells, WI38 cells, HEK293 cells, HEK293T cells, and PER.C6 cells.

[0119] (Application) The DYW protein of the present invention can convert the editing target C contained in the target RNA to U, or convert the editing target U to C. RNA-binding PPR proteins are involved in all RNA processing steps found in organelles, including cleavage, RNA editing, translation, splicing, and RNA stabilization.

[0120] Furthermore, the DYW protein of the present invention can edit mitochondrial RNA1 bases. Mitochondria have their own genomes, encoding proteins that make up important complexes involved in respiration and ATP production. Mutations in these proteins are known to cause various diseases. Mutation repair using the present invention is expected to be useful in treating a variety of diseases.

[0121] On the other hand, the CRISPR-Cas system has been developed as a cytoplasmic C-to-U RNA editing tool by fusing Cas proteins with engineered ADAR domains (Abudayyeh et al., 2019). However, the CRISPR-Cas system consists of a protein and a guide RNA, making efficient mitochondrial transport of the guide RNA difficult. On the other hand, PPR proteins can perform RNA editing with a single molecule, and proteins can generally be delivered to mitochondria by fusing a mitochondrial localization signal sequence to their N-terminus. Therefore, to confirm whether this technology can be used for mitochondrial RNA editing, we designed PPRs targeting MT-ND2 and MT-ND5 and created genes fused with the PG or WW domain (Figure 10a).

[0122] These proteins, consisting of a mitochondrial targeting sequence (MTS) and a PPR-P sequence, target the third position of codons 178 and 301 of MT-ND2 and MT-ND5 to avoid adverse effects on HEK293T cells (Fig. 10a). After plasmid introduction into HEK293T cells, editing was confirmed within the MT-ND2 and MT-ND5 mRNAs (Fig. 10b, c). These four proteins demonstrated up to 70% editing activity against their targets, and no off-target mutations were detected in the same mRNA molecules (Fig. 10bc).

[0123] Therefore, the RNA base editing method provided by the present invention is expected to be used in various fields as follows.

[0124] (1) Medical care The method recognizes and edits specific RNAs associated with specific diseases. Genetic diseases caused by single-base mutations can be treated by using the present invention. The direction of most mutations in genetic diseases is from C to U. Therefore, the method of the present invention, which can convert U to C in particular, can be useful.

[0125] - Creating cells with controlled RNA suppression and expression. Such cells include stem cells (e.g., iPS cells) whose differentiated and undifferentiated states are monitored, model cells for evaluating cosmetics, and cells in which functional RNA expression can be turned on and off for the purpose of elucidating drug discovery mechanisms and pharmacological testing.

[0126] (2) Agriculture, Forestry and Fisheries -Improve yields and quality of agricultural, forestry and marine products. -Breeding organisms with improved disease resistance, improved environmental tolerance, and improved or new functionality.

[0127] For example, with regard to first generation hybrid (F1) crops, it may be possible to artificially create F1 crops by editing mitochondrial RNA with DYW proteins, potentially improving yield and quality. RNA editing with DYW proteins makes it possible to selectively breed and genetically improve organisms more accurately and quickly than conventional techniques. Furthermore, because RNA editing with DYW proteins does not transform traits using foreign genes as in genetic modification, it is closer to traditional breeding methods such as mutant selection and backcrossing. This could potentially address global food and environmental issues reliably and quickly.

[0128] (3)Chemistry In the production of useful substances using microorganisms, cultured cells, plants, and animals (e.g., insects), RNA manipulation controls the amount of protein expression, thereby improving the productivity of useful substances. Examples of useful substances include proteinaceous substances such as antibodies, vaccines, and enzymes, as well as relatively low-molecular-weight compounds such as pharmaceutical intermediates, fragrances, and pigments.

[0129] -Improve the efficiency of biofuel production by modifying the metabolic pathways of algae and microorganisms.

[0130] [term] Numerical ranges from x to y include the endpoints x and y unless otherwise specified.

[0131] With respect to the amino acid sequence of a protein or polypeptide, an amino acid residue may be simply referred to as an amino acid.

[0132] Unless otherwise specified, "identity" with respect to a base sequence (sometimes referred to as a nucleotide sequence) or an amino acid sequence refers to the percentage of matching bases or amino acids shared between the two sequences when the two sequences are optimally aligned. That is, identity can be calculated as follows: identity = (number of matching positions / total number of positions) × 100, and can be calculated using commercially available algorithms. Such algorithms are incorporated into the NBLAST and XBLAST programs described in Altschul et al., J. Mol. Biol. 215 (1990) 403-410. More specifically, searches and analyses for the identity of base sequences or amino acid sequences can be performed using algorithms or programs well known to those skilled in the art (e.g., BLASTN, BLASTP, BLASTX, ClustalW). When using a program, parameters can be appropriately set by those skilled in the art, or the default parameters of each program can be used. Specific techniques for these analysis methods are also well known to those skilled in the art.

[0133] With respect to nucleotide sequences or amino acid sequences, unless otherwise specified, the higher the sequence identity, the more preferable. Specifically, it is preferably 40% or more, more preferably 45% or more, even more preferably 50% or more, even more preferably 55% or more, even more preferably 60% or more, and even more preferably 65% or more. Also, it is preferably 70% or more, more preferably 80% or more, even more preferably 85% or more, even more preferably 90% or more, even more preferably 95% or more, and even more preferably 97.5% or more.

[0134] With respect to a polypeptide or protein, the number of amino acids to be substituted, deleted, or added when referring to a "substituted, deleted, or added sequence" is not particularly limited, unless otherwise specified, in any motif or protein, as long as the motif or protein consisting of that amino acid sequence has the desired function, but may be about 1 to 9 or 1 to 4, or may be an even larger number of substitutions, etc., as long as the substitutions are with amino acids with similar properties. Means for preparing polynucleotides or proteins with such amino acid sequences are well known to those skilled in the art.

[0135] Amino acids with similar properties refer to amino acids with similar physical properties such as hydropathy, charge, pKa, and solubility, and include, for example, the following: Hydrophobic (non-polar) amino acids: alanine, valine, glycine, isoleucine, leucine, phenylalanine, proline, tryptophan, tyrosine Non-hydrophobic amino acids; arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine, cysteine, histidine, methionine; Hydrophilic amino acids; arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine; Acidic amino acids: aspartic acid, glutamic acid; Basic amino acids: lysine, arginine, histidine; Neutral amino acids: alanine, asparagine, cysteine, glutamine, glycine, isoleucine, leucine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine; Sulfur-containing amino acids: methionine, cysteine; Aromatic ring-containing amino acids: tyrosine, tryptophan, phenylalanine. [Example]

[0136] [1. Design of DYW domain and measurement of RNA editing activity in E. coli] [result] (DYW Domain Design) Because most of the currently available plant genome information is from angiosperms, the development of artificial DYW proteins is limited to the DYW:PG type. Meanwhile, transcriptome data, although partial, also contain information on early land plant species that possess both C-to-U and U-to-C RNA editing. Therefore, we designed a complete artificial DYW domain using a database of PPR proteins constructed based on the transcriptome dataset compiled by the 1000 Plants (1KP) International Consortium (see above, non-patent document 11).

[0137] To design artificial DYW domains, we added PPR (P2 and L2) motifs and PPR-like (S2, E1, and E2) motifs to the cytidine deaminase-like DYW domain, because the importance of these motifs for RNA editing activity is currently unknown. We constructed a phylogenetic tree of DYW domains from hornworts, microphyllous plants, and ferns, and classified them into three groups: DYW:PG, DYW:WW, and DYW:KP. Next, we constructed a new phylogenetic tree for the DYW domains from each group and selected phylogenetic groups of proteins that could be used for designing artificial DYW domains (Fig. 1). The DYW:PG group exhibited significant sequence variation, making it difficult to select a phylogenetic group for use in the design (Fig. 1a). Because of the high sequence diversity and abundance of microphyllous proteins, we focused on this group (PG1). For the design of the DYW:WW domain, we selected a phylogenetic group of short-branched proteins found only in hornworts (Fig. 1b, WW1). The short branching of this phylogenetic group suggests that there is little genetic variation between proteins. For the design of the DYW:KP domain, we focused on a phylogenetic group of DYW domains specific to ferns. The protein sequences in this phylogenetic group show large amino acid variations but are conserved in length (Fig. 1c, KP1).

[0138] (Design of RNA-binding domain) We fused each DYW domain to an artificial P or PLS array designed based on the PPR motif sequence identified from plant genome information (see Non-Patent Document 2, supra) (Figure 2a). The PPR motif in the P array was constructed based on a consensus sequence obtained from an alignment of 35-amino acid P motifs, with several amino acid substitutions to enhance RNA recognition. Standard-length 35-amino acid motifs (P1 and L1) and a 31-amino acid motif (S1) were selected for the design of the PLS array. In natural PPR proteins, the first and last P1L1S1 motifs differ from internal P1L1S1 motifs by distinct amino acid residues at specific positions. To design an artificial PLS array as close as possible to the naturally occurring one, we divided the P1L1S1 motifs into the first (N-terminal) P1L1S1, the internal P1L1S1, and the last (C-terminal) P1L1S1 located immediately before P2L2S2, and designed three types of P1L1S1 motifs according to the position of the PPR motif. These proteins will be named based on their structure, for example, a DYW:WW domain fused to a P array will be called P-DYW:WW.

[0139] (The artificial DYW protein can specifically edit target sequences.) In Arabidopsis, the PPR protein CLB19 recognizes RNA editing sites on the chloroplast rpoA and clpP RNAs. We designed a PPR protein that targets the rpoA editing site. The fourth and second amino acids in the P1, L1, S1, and P2 motifs follow the PPR code. However, because the PPR code for the C-terminal PPR-like motifs (L2, S2, E1, and E2) is unknown, we used the fourth and second amino acids of CLB19 for the L2, S2, E1, and E2 motifs (Figure 2b).

[0140] The gene region encoding the recombinant artificial DYW protein was cloned into an expression vector, and a target sequence was added downstream of the stop codon. Based on previous research on two PPR proteins (PPR56 and PPR65) from the moss Physcomitrella patens (see Non-Patent Document 12), we developed a method to test the RNA editing activity of the designed PPR proteins in E. coli. The designed PLS-DYW:PG1 and PLS-DYW:WW1 showed no DNA editing activity, whereas RNA editing efficiency exceeded 90% when cytidine was replaced with uridine (Figure 3a, b, d). When the P array was used instead of the PLS domain, editing efficiency decreased by 10–40%. On the other hand, both P-DYW:KP and PLS-DYW:KP showed U-to-C RNA editing activity (Figure 3c, e). However, their editing activity was reduced compared to DYW:PG and DYW:WW.

[0141] [method] (phylogenetic tree) The P2-L2-S2-E1-E2-DYW region containing the DYW domain (minimum 132 amino acids) was extracted from the PPR database (https: / / ppr.plantenergy.uwa.edu.au / onekp / ; see Non-Patent Document 11). The resulting sequence alignment was generated using MAFFT L-INS-i (v7.407 automatic mode) (K. Katoh, D.M. Standley (2013). Mol. Biol. Evol. 30(4): 772-780.), followed by trimming using trimAl (v.1.4.rev15) (Salvador Capella-Gutierrez, et al. (2009). Bioinformatics. 25(15): 1972-1973.). The trimming parameters were gt 0.2 cons 20. The active site (HxEx n Sequences with mutations in CxxCH) were excluded from the alignment (Fig. 4).

[0142] A phylogenetic tree was constructed from the remaining sequences using FastTree (v2.1.10) (Price, MN, et al. (2010). PLoS One 5:e9490.) to identify DYW:PG, DYW:WW, and DYW:KP. The parameters used were wag and cat 8.

[0143] (Cloning of Trx-PPR-DYW protein and target sequence) Consensus sequences for each DYW domain (including P2, L2, S2, E1, and E2) were designed using EMBOSS:cons (v. 6.6.0.0). The protein expression vector was modified from pET21b+PA by removing the original Esp3I and BpiI restriction enzyme sites and adding two Esp3I sites for cloning. The gene was divided into four sections (Trx, PPR array, DYW domain, and RNA editing site) and constructed using the two-step Golden Gate method. First, three parts were cloned into the Esp3I site of the modified pET21b vector: 1) the thioredoxin-6×His-TEV gene region (containing a BpiI restriction enzyme site at the 3' end), 2) the P2-L2-S2-E1-E2-DYW gene region (containing a BpiI restriction enzyme site at the 5' end), and 3) the coding sequence for the RNA editing site. Next, the full-length PLS domain or P domain was cloned into the BpiI site to prepare PLS:DYW (SEQ ID NOs: 35 to 37) or P:DYW (SEQ ID NOs: 38 to 40) proteins.

[0144] (Assessment of RNA editing activity in E. coli) To analyze the RNA-editing activity of the recombinant protein in E. coli, we modified the protocol developed by Oldenkott et al. (Non-Patent Document 12). The plasmid DNA prepared above was introduced into E. coli Rosetta 2 strain and cultured overnight at 37°C in 1 mL of LB medium (containing 50 μg / mL of carbenicillin and 17 μg / mL of chloramphenicol). 5 mL of LB medium containing the appropriate antibiotic was prepared in a deep-bottom 24-well plate, and 100 μL of the preculture was inoculated into the well. The culture was analyzed by optical density (OD 600The cells were grown at 37°C and 200 rpm until the β-kappa-methyl-1-methyl-2-methyl-1 ...

[0145] Frozen cell pellets were resuspended in 200 μL of 1-thioglycerol / homogenization solution and sonicated at 40 W for 10 seconds to dissociate the cells, followed by the addition of 200 μL of lysis buffer. RNA was extracted using the Maxwell® RSC simplyRNA Tissue Kit (Promega). The RNA was treated with DNase I (Takara Bio), and cDNA was synthesized using 1 μg of treated RNA and 1.25 μM random primers (6-mer) with SuperScript® III Reverse Transcriptase (Invitrogen). The region containing the editing site was amplified using NEBNext High-Fidelity 2x PCR Master Mix (New England BioLabs), 1 μL of cDNA, and primers targeting thioredoxin and T7 terminator sequences. The PCR product was purified using NucleoSpin® Gel and PCR Cleanup (Takara Bio), and sequence analysis was performed using a forward primer specific for the DYW domain sequence to identify the base sequence at the RNA editing site. The RNA editing efficiency was measured using the ratio of the peak heights of the C and U waveforms at the editing site. The C-to-U RNA editing efficiency was calculated as U / (C+U) × 100, and the U-to-C editing efficiency was calculated as C / (C+U) × 100. Experiments were performed three times.

[0146] [List of cited sequences]

[0147] [Table 3-1]

[0148] [Table 3-2]

[0149] [Table 3-3]

[0150] [2. Example of cultured animal cells] [result] HEK293T cells were transfected with a gene fused to a PLS-type PPR and each DYW domain (PG1 (sequence number: 1), WW1 (sequence number: 2), KP1 (sequence number: 3)) and a plasmid containing the target sequence, and RNA was collected after culture. The conversion efficiency of cytidine (C) to uridine (U) or uridine (U) to cytidine (C) at the target site was analyzed by Sanger sequencing (Figure 5). When the PG1 or WW1 domain was fused, C to U activity was over 90%, while U to C activity was undetectable (Figure 5a, c). When the KP1 domain was fused, U to C activity was detected at 25%, but C to U activity was also detected at approximately 10% (Figure 5b, c). These results demonstrate that the editing enzyme functions in cultured animal cells.

[0151] [method] (Preparation of PPR expression plasmid for animal cell culture tests) From the plasmid used in Figure 3, the gene sequence of the 6xHis-PPR-DYW protein (the same protein used in the E. coli experiments (SEQ ID NOs: 35-37)) and the region containing the editing site (SEQ ID NO: 34) were amplified by PCR and cloned into an animal cell expression vector using the Golden Gate Assembly method. The vector expresses PPR under the control of a promoter containing a CMV promoter and a human β-globin chimeric intron, and a poly(A) signal is provided by the SV40 polyadenylation signal.

[0152] (HEK293T cell culture) HEK293T cells were cultured in Dulbecco's Modified Eagle Medium (DMEM) containing high glucose, glutamine, phenol-RED, and sodium pyruvate (Fujifilm Wako Pure Chemical Corporation) supplemented with 10% fetal bovine serum (Capricom) and 1% penicillin-streptomycin (Fujifilm Wako Pure Chemical Corporation) at 37°C and 5% CO2. Cells were passaged every 2–3 days when they reached 80–90% confluence.

[0153] (Transfection) For the RNA editing assay, approximately 8.0 x 10 HEK293T cells were plated in each well of a 24-well flat-bottom cell culture plate. 4 The cells were cultured at 37°C and 5% CO2 for 24 hours. 500 ng of plasmid was added to each well in 18.5 μl of Opti-MEM® I Reduced Serum Medium (ThermoFisher) and 1.5 μl of FuGENE® HD Transfection Reagent (Promega) for a final volume of 25 μl. The mixture was incubated at room temperature for 10 minutes before adding to the cells. 24 hours after transfection, the cells were harvested.

[0154] (RNA extraction, reverse transcription, sequencing) The assay was carried out in the same manner as in the assay using E. coli described above. In the following examples, unless otherwise specified, experiments were carried out in the same manner as in the assay using E. coli and in this example.

[0155] [3. Effect of domain swapping on RNA editing activity] The DYW domain can be divided into several regions based on the amino acid sequence conservation, but the relationship between these regions and RNA editing activity is unknown. Here, we swapped parts of the KP1 and WW1 domains and examined the effect on RNA editing activity in HEK239T cells (Figure 6). When the PG box and DYW of the WW1 domain were fused to the central region containing the active site of the KP1 domain (chimKP1a, SEQ ID NO: 90), we observed U-to-C activity nearly 50% higher than that of the KP1 domain, and almost no C-to-U activity. Domain swapping successfully improved the U-to-C editing performance of the KP domain.

[0156] [4. Improving the RNA editing activity of the KP domain by introducing mutations] We aimed to improve the RNA editing activity of the KP domain by introducing various mutations into it. KP2 to KP23 (SEQ ID NOs: 68 to 89) were designed, and their C to U or U to C RNA editing activity was examined in E. coli (Figure 7a) and HEK293T cells (Figures 7b, c). KP22 (SEQ ID NO: 88) had the highest U to C editing activity and the lowest C to U editing activity, demonstrating successful improvement in RNA editing activity compared to KP1.

[0157] [5. Improving RNA editing activity of the PG domain by mutation introduction] We aimed to improve the RNA editing activity of the PG domain by introducing various mutations into it. PG2 to PG13 (SEQ ID NOs: 41 to 53) were designed, and their C to U RNA editing activity was examined in E. coli (Figure 8a) and HEK293T cells (Figures 8b and 8c). PG11 (SEQ ID NO: 50) showed the highest C to U editing activity, successfully improving RNA editing activity.

[0158] 6. Improving RNA editing activity of WW domains by mutation introduction We aimed to improve the RNA editing activity of the WW domain by introducing various mutations into the WW domain. WW2 to WW14 were designed, and their C to U RNA editing activity was examined in E. coli (Figure 9a) and HEK293T cells (Figures 9b and 9c). WW11 (SEQ ID NO: 63) had the highest C to U editing activity, successfully improving RNA editing activity compared to WW1.

[0159] [7. Human mitochondrial RNA editing using PPR proteins] [result] Mitochondria have their own genome, encoding the proteins that make up the important complexes involved in respiration and ATP production. Mutations in these proteins are known to cause various diseases, and methods for repairing these mutations are needed.

[0160] The CRISPR-Cas system has been developed as a cytoplasmic C-to-U RNA editing tool by fusing a modified ADAR domain to a Cas protein (Abudayyeh et al. 2019 Science Vol. 365, Issue 6451, pp. 382-386). However, the CRISPR-Cas system consists of a protein and a guide RNA, making efficient mitochondrial transport of the guide RNA difficult. On the other hand, PPR proteins can perform RNA editing with a single molecule, and proteins can generally be delivered to mitochondria by fusing a mitochondrial localization signal sequence to their N-terminus. Therefore, to confirm whether this technology can be used for mitochondrial RNA editing, we designed PPRs targeting MT-ND2 and MT-ND5 and created genes fused with the PG1 or WW1 domain (Figure 10a).

[0161] These proteins, consisting of a mitochondrial targeting sequence (MTS) and a P-DYW sequence, target the third position of codons 178 and 301 of MT-ND2 and MT-ND5 to avoid adverse effects on HEK293T cells (Fig. 10a). After plasmid introduction into HEK293T cells, editing was confirmed within MT-ND2 and MT-ND5 mRNA (Fig. 10b, c). These four proteins demonstrated up to 70% editing activity against their targets, and no off-target mutations were detected in the same mRNA molecules (Fig. 10bc).

[0162] [method] (Cloning for mitochondrial editing) The mitochondrial targeting sequence from the Zea mays LOC100282174 protein (Chin et al. 2018), 10 PPR-P and PPR-like motifs, and the DYW domain portion (PG1 and WW1 are used for the DYW domain) were cloned into an expression plasmid under the control of the CMV promoter by Golden Gate Assembly (SEQ ID NOs: 91–94).

[0163] [List of cited sequences]

[0164] [Table 4-1]

[0165] [Table 4-2]

[0166] [Table 4-3]

[0167] [Table 4-4]

Claims

1. A composition for treating disease-associated RNA, comprising a nucleic acid encoding a DYW protein or a vector comprising said nucleic acid, DYW protein, an RNA-binding domain comprising at least one PPR motif and capable of sequence-specifically binding to a target RNA; and A DYW domain consisting of any one of the following polypeptides: The composition is a DYW protein. a polypeptide consisting of the sequence of SEQ ID NO: 1; a polypeptide having at least 90% sequence identity with the sequence of SEQ ID NO: 1 and having C-to-U editing activity; a polypeptide having a sequence in which 1 to 13 amino acids are substituted, deleted, or added in the sequence of SEQ ID NO: 1 and having C-to-U editing activity; or a polypeptide consisting of the sequence of SEQ ID NO: 41 a polypeptide consisting of the sequence of SEQ ID NO: 2; a polypeptide having at least 90% sequence identity with the sequence of SEQ ID NO: 2 and having C-to-U editing activity; a polypeptide having a sequence in which 1 to 12 amino acids are substituted, deleted, or added in the sequence of SEQ ID NO: 2 and having C-to-U editing activity; or a polypeptide consisting of the sequence of SEQ ID NO: 54 a polypeptide consisting of the sequence of SEQ ID NO: 3; a polypeptide having at least 90% sequence identity with the sequence of SEQ ID NO: 3 and having U-to-C editing activity; a polypeptide having a sequence in which 1 to 12 amino acids are substituted, deleted, or added in the sequence of SEQ ID NO: 3 and having U-to-C editing activity; or a polypeptide consisting of the sequence of SEQ ID NO: 68, 69, 70, 72, 74, 76, 77, 78, 80, 81, 82, 84, 86, 87, 88, or 89 - A polypeptide consisting of the sequence of SEQ ID NO: 90; a polypeptide having at least 90% sequence identity with the sequence of SEQ ID NO: 90 and having U-to-C editing activity.

2. The composition of claim 1, wherein the DYW domain consists of any one of the following polypeptides: A polypeptide consisting of the sequence of SEQ ID NO: 1 A polypeptide consisting of the sequence of SEQ ID NO: 2 A polypeptide consisting of the sequence of SEQ ID NO: 3 A polypeptide consisting of the sequence of SEQ ID NO: 90

3. The composition of claim 1 or 2, wherein the RNA-binding domain is of the PLS type.

Citation Information

Patent Citations

  • Design method for RNA-binding protein using PPR motif, and use thereof

    WO2013058404A1

  • DNA binding protein using PPR motif, and use thereof

    WO2014175284A1