Design method of RNA-binding protein using PPR motif and its use

By identifying key amino acids in PPR motifs, the method allows for the design of proteins that can selectively bind to RNA sequences, addressing the lack of specificity in existing RNA-binding proteins and enabling targeted RNA modification in cells.

JP7789426B2Active Publication Date: 2025-12-22KYUSHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024194918
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2011-10-21
Filing Date
2024-11-07
Publication Date
2025-12-22
Estimated Expiration
2032-10-22

AI Technical Summary

Technical Problem

The correlation between the amino acid composition and function of PPR proteins, which act as RNA adaptors, remains largely unknown, limiting the ability to construct proteins that can bind to RNA sequences of any length and specificity.

Method used

The method involves identifying the specific amino acids (A1, A4, and Lii) in the PPR motif, which determine RNA base binding, allowing for the design of proteins that can selectively bind to specific RNA bases or sequences by manipulating the motif's structure or combination.

Benefits of technology

This approach enables the creation of proteins that can bind to target RNAs of any sequence or length, facilitating the prediction and identification of target RNAs and proteins, and enabling functional modification of genetic material in cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007789426000010
    Figure 0007789426000010
  • Figure 0007789426000011
    Figure 0007789426000011
  • Figure 0007789426000012
    Figure 0007789426000012
Patent Text Reader

Abstract

To provide a composition comprising a nucleic acid encoding a PPR protein, the composition being used for treating a disease mediated by any modification selected from the group consisting of translation, splicing, and stability of an RNA having a target sequence to which the PPR protein can bind.SOLUTION: The protein to be used is a protein containing one or more (preferably 2 to 14) PPR motifs consisting of a polypeptide of 30 to 38 amino acids in length represented by formula 1 (where: Helix A is a portion of 12 amino acids in length that can form an α-helical structure, represented by formula 2, and A1 to A12 each independently represent an amino acid in the formula 2; X is absent or is a portion consisting of 1 to 9 amino acids in length; Helix B is a portion of 11 to 13 amino acids in length that can form an α-helical structure; L is a portion of 2 to 7 amino acids in length that is represented by formula 3; in the formula 3, each amino acid is numbered from a C-terminus side as "i" (-1), "ii" (-2), except that Liii to Lvii may not exist). The combination of three amino acids A1, A4, and Lii, or the combination of two amino acids A4 and Lii, is determined according to a target RNA base or a base sequence.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to proteins capable of selectively or specifically binding to intended RNA bases or RNA sequences. The present invention utilizes pentatricopeptide repeat (PPR) motifs. The present invention can be used to identify and design RNA-binding proteins, identify target RNAs of PPR proteins, and regulate RNA function. The present invention is useful in fields such as medicine and agriculture. [Background technology]

[0002] In recent years, various analytical techniques have been developed to utilize nucleic acid-binding protein factors that bind to specific sequences. By utilizing this sequence-specific binding, it is becoming possible to analyze the intracellular location of target nucleic acids (DNA or RNA), remove target DNA sequences, or control (activate or inactivate) the expression of downstream protein-coding genes.

[0003] Research and development has been conducted using zinc finger proteins (Non-Patent Document 1) and TAL effectors (Non-Patent Document 2, Patent Document 1) as protein engineering materials for protein factors that act on DNA. However, the development of protein factors that specifically act on RNA remains very limited. This is because the general rules regarding the affinity of the amino acid sequence constituting a protein with RNA and the relationship with the RNA sequence to which it binds are largely unknown, or no rules have been found. An exception has been made regarding the pumilio protein, which is composed of multiple repeats of the 38-amino acid puf motif. It has been shown that one puf motif binds to a single base in RNA (Non-Patent Document 3). Attempts have been made to create proteins with novel RNA-binding properties using pumilio proteins, as well as techniques for modifying RNA-binding properties (Non-Patent Document 4). However, puf motifs are highly conserved and exist in limited numbers. Therefore, they have only been used to create protein factors that act on limited RNA sequences.

[0004] On the other hand, genome sequence information has identified a large family of proteins, numbering as many as 500 in plants alone: ​​PPR proteins (proteins with pentatricopeptide repeat (PPR) motifs) (Non-Patent Document 5). PPR proteins are nuclear-encoded, but they act in a gene-specific manner exclusively on RNA-level regulation, cleavage, translation, splicing, RNA editing, and RNA stability in organelles (chloroplasts and mitochondria). PPR proteins typically have a structure consisting of approximately 10 consecutive PPR motifs, a 35-amino acid motif with low conservation. It is believed that the combination of PPR motifs is responsible for sequence-selective binding to RNA. Most PPR proteins consist of only approximately 10 repeats of the PPR motif, and in many cases, the domain required for catalytic activity is not found. Therefore, these PPR proteins are thought to function as RNA adapters (Non-Patent Document 6).

[0005] The present inventors have proposed a method for modifying RNA-binding proteins using this PPR motif (Patent Document 2). [Prior art documents] [Patent documents]

[0006] [Patent Document 1] WO2011 / 072246 [Patent Document 2] WO2011 / 111829 [Non-patent literature]

[0007] [Non-Patent Document 1] Maeder , ML , Thibodeau-Beganny , S , Osiak , A , Wright , DA , Anthony , RM , Eichtinger , M , Jiang , T , Foley , JE , Winfrey , RJ , Townsend , JA , et al. (2008). Rapid "open-source" engineering of customized zinc-finger nucleases for highly efficient gene modification. Mol. Cell 31, 294–301.

Outdoor Tool2

Outdoor Tools3

Outdoor Tools 4

Direct Environment 5

[0008] The properties of PPR proteins as RNA adaptors are expected to be determined by the properties of each PPR motif that constitutes the PPR protein, as well as the combination of multiple PPR motifs. However, the correlation between the amino acid composition and function remains largely unknown. If the amino acids that function when PPR motifs exert their RNA-binding properties are identified and the relationship between the structure of the PPR motif and the target base is clarified, it may be possible to construct proteins that can bind to RNA of any sequence and length by artificially manipulating the structure of the PPR motif or the combination of multiple motifs. [Means for solving the problem]

[0009] To achieve the above-mentioned object, the present inventors have studied genetically analyzed PPR proteins, particularly PPR proteins involved in RNA editing (modification of genetic information at the RNA level, particularly the conversion of cytosine (hereinafter, C) to uracil (hereinafter, U)), and their target RNA sequences. Using computational science techniques, they have revealed that three amino acids (amino acids 1, 4, and "ii" (-2)) in the PPR motif contain information that controls binding to specific RNA bases. More specifically, the present inventors have found that the selectivity (also referred to as specificity) of the RNA base binding of the PPR motif is determined by three amino acids: the first and fourth amino acids in the first helix of the two α-helices that make up the motif, and the second amino acid from the end (C-terminal side) in the portion of the second helix that can form a loop structure ("ii"; -2). Based on this finding, the present invention has been completed.

[0010] The present invention provides the following: [1] A method for designing a protein capable of binding to an RNA base selectively or specifically to an RNA base sequence, comprising: The protein contains one or more (preferably 2 to 14) PPR motifs consisting of a polypeptide having a length of 30 to 38 amino acids and represented by formula 1.

[0011] [ka] (In the formula: Helix A is a 12 amino acid long portion capable of forming an α-helical structure and is represented by Formula 2:

[0012] [ka] In formula 2, A1~A 12 each independently represents an amino acid; X is absent or a moiety consisting of 1 to 9 amino acids in length; Helix B is a portion consisting of 11 to 13 amino acids that can form an α-helical structure; L is a moiety of formula 3 that is 2 to 7 amino acids in length;

[0013] [ka] In Formula 3, each amino acid is numbered from the C-terminus side as "i" (-1), "ii" (-2), etc. However, L iii ~L vii may not exist.) A1, A4, and L ii A combination of three amino acids, or A4, L ii The combination of the two amino acids is determined according to the target RNA base or base sequence. [2] A1, A4, and L ii The method according to [1], wherein the combination of three amino acids is determined based on an RNA base or a target base sequence, and the combination of amino acids is determined based on any of the following: (3-1) A1, A4, and L ii When the three amino acids are, in order, valine, asparagine, and aspartic acid, the PPR motif can selectively bind to U (uracil); (3-2) A1, A4, and L ii When the three amino acids are, in order, valine, threonine, and asparagine, the PPR motif can selectively bind to A (adenine); (3-3) A1, A4, and L ii When the three amino acids are, in order, valine, asparagine, asparagine, the PPR motif can selectively bind to C (cytosine); (3-4) A1, A4, and L ii When the three amino acids are, in order, glutamic acid, glycine, and aspartic acid, the PPR motif can selectively bind to G (guanine); (3-5) A1, A4, and L ii When the three amino acids are, in order, isoleucine, asparagine, asparagine, the PPR motif can selectively bind to C or U; (3-6) A1, A4, and L ii When the three amino acids are, in order, valine, threonine, and aspartic acid, the PPR motif can selectively bind to G; (3-7) A1, A4, and L ii When the three amino acids are, in order, lysine, threonine, and aspartic acid, the PPR motif can selectively bind to G; (3-8) A1, A4, and L ii When the three amino acids are, in order, phenylalanine, serine, and asparagine, the PPR motif can selectively bind to A; (3-9) A1, A4, and L ii When the three amino acids are, in order, valine, asparagine, and serine, the PPR motif can selectively bind to C; (3-10) A1, A4, and L ii When the three amino acids are, in order, phenylalanine, threonine, and asparagine, the PPR motif can selectively bind to A; (3-11) A1, A4, and L ii When the three amino acids are, in order, isoleucine, asparagine, and aspartic acid, the PPR motif can selectively bind to U or A; (3-12) A1, A4, and L ii When the three amino acids are, in order, threonine, threonine, and asparagine, the PPR motif can selectively bind to A; (3-13) A1, A4, and L ii When the three amino acids are, in order, isoleucine, methionine, and aspartic acid, the PPR motif can selectively bind to U or C; (3-14) A1, A4, and L ii When the three amino acids are, in order, phenylalanine, proline, and aspartic acid, the PPR motif can selectively bind to U; (3-15) A1, A4, and L iiWhen the three amino acids are, in order, tyrosine, proline, and aspartic acid, the PPR motif can selectively bind to U; (3-16) A1, A4, and L ii When the three amino acids are leucine, threonine, and aspartic acid, respectively, the PPR motif can selectively bind to G. [3] A4 and L ii The method according to [1], wherein the combination of two amino acids is determined based on an RNA base or a target base sequence, and the combination of amino acids is determined based on any of the following: (2-1) A4, L ii are, in order, asparagine, aspartic acid, the motif can selectively bind to U; (2-2) A4, L ii is, in order, asparagine, asparagine, the motif can selectively bind to C; (2-3) A4, L ii is, in order, threonine, asparagine, the motif can selectively bind to A; (2-4) A4, L ii is, in order, threonine, aspartic acid, the motif can selectively bind to G; (2-5) A4, L ii is, in order, serine, asparagine, the motif can selectively bind to A; (2-6) A4, L ii are, in order, glycine, aspartic acid, the motif can selectively bind to G; (2-7) A4, L ii are, in order, asparagine and serine, the motif can selectively bind to C; (2-8) A4, L ii is, in order, proline, aspartic acid, the motif can selectively bind to U; (2-9) A4, L ii are, in order, glycine, asparagine, the motif can selectively bind to A; (2-10) A4, L ii are, in order, methionine, aspartic acid, the motif can selectively bind to U; (2-11) A4, L ii is, in order, leucine, aspartic acid, the motif can selectively bind to C; (2-12) A4, L ii When α, β ... [4] A method for identifying a base or base sequence that is a target of an RNA-binding protein and that contains one or more (preferably 2 to 14) PPR motifs defined in [1], comprising: The identification is based on any one of (3-1) to (3-16) described in 2 or any one of (2-1) to (2-12) described in 3, and PPR motifs A1, A4, and L ii A combination of three amino acids, or A4, L ii The method is carried out by identifying the presence or absence of a base according to the combination of two amino acids. [5] A method for identifying a PPR protein that is capable of binding to a target RNA base or a target RNA having a specific base sequence and contains one or more (preferably 2 to 14) PPR motifs defined in [1], comprising the steps of: The identification is based on any one of (3-1) to (3-16) described in 2 or any one of (2-1) to (2-12) described in 3, and the PPR motifs A1, A4, and L corresponding to the target RNA base or a specific base constituting the target RNA are identified. ii The method is carried out by identifying the presence or absence of a combination of three amino acids. [6] A method for controlling RNA function using a protein designed by the method described in [1]. [7] A complex comprising a region consisting of a protein designed by the method described in [1] and a functional region linked together. [8] A method for modifying the genetic material of a cell, comprising the steps of: providing cells containing RNA having a target sequence; and [7] A method for modifying RNA having a target sequence by introducing the complex described in [7] into a cell, whereby the protein domain of the complex binds to RNA having a target sequence, and thus the functional domain modifies the RNA having the target sequence. [9] detecting amino acid polymorphisms in the PPR protein gene, which acts as a fertility restorer for cytoplasmic male sterility, found among various varieties; identifying the association between polymorphisms in the gene and fertility; a step of identifying the base sequence of the PPR protein gene obtained from the test sample and determining the fertility of the test sample; A method for determining the fertility of a PPR protein gene, comprising:

[10] The method according to [9], wherein the PPR protein is a protein containing one or more (preferably 2 to 16) PPR motifs consisting of a 30 to 38 amino acid long polypeptide represented by formula 1 defined in [1].

[11] The method according to [9] or

[10] , wherein the amino acid polymorphism is identified as a polymorphism for each PPR motif.

[12] The polymorphism of the PPR motif is A1, A4, and L of the motif of Formula 1. ii A combination of three amino acids, or A4, L ii The method according to any one of [9] to

[11] , wherein the amino acid sequence is specified by a combination of two amino acids.

[13] The method of claim 12, wherein the polymorphism of the PPR motif is identified by a polymorphism of the fourth amino acid (A4) of the motif of formula 1.

[14] The method described in claim 13, wherein the fourth amino acid of all PPR motifs on the PPR protein gene is identical to the fourth amino acid of all corresponding PPR motifs in Enko B, indicating fertility.

[15] A method according to any one of claims 9 to 14, wherein the PPR protein gene is an orf687-like gene (i.e., a family gene located at a locus homologous to the "687 gene" encoding Enko B, a gene having 90% or more amino acid sequence identity with Enko B, or a gene having 90% or more nucleotide sequence identity with the "ORF687 gene" encoding Enko B).

[16] The method according to any one of claims 9 to 15, wherein the proteins encoded by the orf687-like genes of various varieties are any of SEQ ID NOs: 576 to 578, 585 to 591. [Effects of the Invention]

[0014] The present invention provides a PPR motif capable of binding to a target RNA base and a protein containing the same. By arranging multiple PPR motifs, proteins capable of binding to target RNAs of any sequence or length can be provided.

[0015] The present invention makes it possible to predict and identify the target RNA of any PPR protein, and conversely, to predict and identify PPR proteins that bind to any RNA. Predicting the target RNA sequence clarifies its genetic identity and expands the possibilities for its use. For example, when fertility is considered as the function of a PPR protein in the present invention, the functionality of homologous genes with various amino acid polymorphisms can be tested based on differences in their target RNA sequences for an industrially useful PPR protein gene that acts as a restorer of cytoplasmic male sterility.

[0016] Furthermore, a functional group can be attached to the PPR motif or PPR protein provided by the present invention to prepare a complex.

[0017] Furthermore, the present invention can be used to deliver the above-mentioned complex into a living body and make it function, or to create transformants using nucleic acid sequences (DNA, RNA) encoding the proteins obtained by the present invention, or to specifically modify, control, and impart functions to living organisms (cells, tissues, individuals) in various situations. [Brief explanation of the drawings]

[0018] [Figure 1]Figure 1 shows the conserved sequences and amino acid numbers of the PPR motif. (A) The amino acids constituting the PPR motif as defined in this invention and their amino acid numbers are shown. (B) The positions of the three amino acids (numbers 1, 4, and "ii" (-2)) that control binding base selectivity are shown in the predicted structure. (C) The positions of the amino acids in the predicted structure. Using the entire amino acid sequence of Arabidopsis thaliana CRR4 (SEQ ID NO: 6) and CRR21 (SEQ ID NO: 3) as query sequences, the predicted structures were analyzed using the PHYRE (http: / / www.sbg.bio.ic.ac.uk / phyre / ) program. Using O-GlucNAc transferase (1w3b) as a template, the predicted structures were predicted with high scores (4.3e-17 and 4.7e-16, respectively) for CRR4 and CRR21. The fifth PPR motif of CRR4 (left) and the eighth PPR motif of CRR21 (right) are shown. Numbers 1, 4, and "ii" (-2) are shown as magenta sticks (dark gray in black and white). [Figure 2] Figure 2 shows the RNA-editing PPR proteins that have been analyzed so far and their target RNA-editing sites. [Figure 3-1] FIG. 3 shows the PPR motif sequences and amino acid numbers of Arabidopsis RNA-editing PPR proteins. [Figure 3-2] Figure 3-2 shows a continuation of Figure 3-1. [Figure 3-3] Figure 3-3 shows a continuation of Figure 3-2. [Figure 3-4] Figure 3-4 shows a continuation of Figure 3-3. [Figure 4]Figure 4 shows the amino acids in the PPR motif involved in RNA recognition. (A) Identification of amino acids with base-specific binding ability in the PPR motif. The PPR motif of an RNA-editing PPR protein was aligned with the upstream sequence of the RNA editing site in various configurations. The alignment was performed by sequentially arranging the motif and base one-to-one. Alignment P1 corresponds to the last PPR motif, which corresponds to the base immediately before the edited C. Alignments P2–P6 were obtained by shifting the base sequence one base to the right. Squares represent PPR motifs, and diamonds represent additional motifs (E, E+, DYW) at the C-terminus. If an amino acid at a specific position within the motif (e.g., an amino acid in a green motif (dark gray in monochrome display)) is responsible for RNA base recognition, low randomness can be expected between the corresponding base in a specific alignment (bottom right). Otherwise, high randomness is expected (top right). (B) The RNA base-specific binding ability of amino acids 1, 4, and “ii” (-2). The low randomness of the amino acid and base in each alignment is shown as a P value. (C) The RNA base-binding ability of amino acids 1, 4, and "ii"(-2) in various nucleic acid classifications. Same as (B). However, nucleic acids are classified as purines or pyrimidines (RY; A&G or U&C) or hydrogen-bonding groups (WS; A&U, G&C). (D) Further analysis of the base-binding ability of the RNA-recognizing amino acids in the PPR motif shown in (C) above revealed that in addition to the primary function of amino acid 4 in distinguishing between purines and pyrimidines (RY), amino acid ii"(-2) also distinguishes between amino (A and C) and keto (G and U) forms (MK) of the base (Figure 4D). (E) Examples of RNA recognition codes (PPR codes) for several PPR motifs. The white background indicates the type of amino acid 1, 4, and "ii"(-2). The frequency of occurrence of each code is shown as a number, and the frequency of occurrence of the corresponding nucleic acid is shown as a nucleotide frequency. [Figure 5]Figure 5 shows an example of the identification of amino acids in PPR motifs involved in RNA recognition. Using the dataset of RNA bases corresponding to the PPR motifs in each alignment, we searched for amino acids involved in RNA recognition. For example, using the data of RNA bases corresponding to the PPR motifs in alignment P4, we analyzed the ability of amino acids 4 and 5 to bind to RNA. For each alignment, we first sorted the data by amino acid type and calculated the number of RNA bases included (top left). Next, a theoretical value was created based on the median occurrence frequency of all RNAs included in the dataset (top right). A chi-square test was performed using these two data sets to calculate P values. The top panel shows the analysis results for amino acid 4 in alignment P4, which yielded a significant P value, while the bottom panel shows amino acid 5 in alignment 4, which did not yield a significant P value. [Figure 6] Figure 6 shows the search for amino acids responsible for RNA base specification ability. (A) The low randomness between the type of amino acid and the frequency of occurrence of the base was calculated for amino acids at all positions in the alignment of P1 to P6. Amino acids that showed a significant P value (P<0.01) are shown in magenta (dark gray in black and white). The cyan line (horizontal line in the graph; dark gray in black and white) indicates a P value of 0.01. (B) Summary of low randomness in each alignment. The product of the P values ​​for the amino acids at each position shown in (A) is shown as the overall low randomness value for that alignment. [Figure 7] Figure 7 shows the ability of two amino acids to specify the bases of the RNA. The ability of two amino acids in different combinations (amino acids 1 & 4, 1 & "ii", and 4 & "ii") to specify the bases of the RNA was analyzed using low randomness between the amino acids and the corresponding bases, as in Figure 4. [Figure 8] FIG. 8 shows the RNA recognition code of the PPR motif extracted from Arabidopsis thaliana. [Figure 9]Figure 9 shows the sequences of the Physcomitrella patens RNA-editing PPR proteins and the RNA-editing site where each protein functions. The amino acid sequences of positions 1, 4, and "ii" (-2) within each PPR motif are shown along with the protein motif structure. Magenta and cyan (both dark gray in black-and-white display) letters indicate amino acid combinations homologous to the triPPR or diPPR codes extracted from Arabidopsis thaliana, respectively. Additional motifs (E, E+, DYW) at the C-terminus are also shown. The sequence of the RNA-editing site where each protein functions (the upstream sequence containing the C to be edited) is indicated by the alignment P4 shown in Figure 4. [Figure 10] Figure 10 shows a flowchart for calculating the compatibility score between PPR proteins and RNA sequences at the RNA editing site. Obtain a protein PPR model from the Uniprot or PROSITE database and assign each amino acid number according to Figure 1. Extract the first, fourth, and "ii" amino acids. The moss PPR protein, PpPPR71, is shown as an example. Next, matched amino acid combinations are converted into a triPPR code matrix. Motifs that could not be converted into the triPPR code are then converted into a diPPR code matrix. In parallel, the 30-nt RNA editing site (the final edited C) is converted into a formula matrix. The ccmFCeU122SF sequence, which is acted upon by the PpPPR71 protein, is shown as an example. Next, the product of corresponding squares in the protein coding matrix and the RNA formula matrix is ​​calculated, and the sum is used to calculate the compatibility score. The last row of the protein coding matrix must be matched to the row corresponding to the four bases before the edited C. This calculation is performed separately for the protein coding matrices created from the triPPR code and the diPPR code. Using a standard distribution curve constructed from the match values ​​with multiple RNA sequences, provisional P values ​​for each RNA sequence are calculated for the triPPR and diPPR codes. The final match value (P value) is calculated as the product of the provisional P values ​​for the triPPR and diPPR codes. [Figure 11]Figure 11 shows the predicted target RNA sequences of PPR proteins using the PPR code. (A) Amino acids 1, 4, and "ii" (-2) were extracted from the moss PPR protein and converted to triPPR or diPPR code as shown in Figure 10. The match scores with the RNA editing sites were calculated and shown as P values. Thirteen RNA editing sites in moss were used as the RNA editing sites, and 34 RNA editing sites in Arabidopsis chloroplasts were used as the reference sequences. Only the match scores with the 13 moss RNA editing sites are shown in the figure. Diamonds indicate the match scores of each protein with each editing site. Correct editing sites are shown in magenta (solid gray in black-and-white display). (B) The P values ​​shown in (A) are shown in a table. [Figure 12] Figure 12 shows the verification of the accuracy of RNA editing site prediction using Arabidopsis RNA editing proteins. The prediction accuracy was verified using the Arabidopsis PPR proteins used for code extraction. (A) Predicted RNA editing sites for 13 known PR proteins for all 34 chloroplast RNA editing sites. Each diamond indicates the match score between the protein and RNA editing site sequence. Correct RNA editing sites are shown in magenta (solid gray in black and white display). (B) Predicted RNA editing sites for 11 known PR proteins for all 488 mitochondrial RNA editing sites. [Figure 13]Figure 13 shows the predicted and experimentally verified target RNA editing sites of the Arabidopsis PPR protein AHG11. (A) Motif structure of AHG11. It has a typical RNA-editing PPR protein structure, consisting of 12 PPR motifs and additional C-terminal motifs (E, E+, DYW). In the Ahg11 mutant, a point mutation at the position indicated by an asterisk (Trp 295) results in a new translation termination codon within the coding region. (B) Target RNA editing sites predicted using all RNA editing sites in Arabidopsis chloroplasts and mitochondria. The top 10 editing sites with the highest P values ​​are shown. The presence or absence of RNA editing in wild-type and mutant strains was experimentally verified and is shown as the editing status. Sites where editing was detected in both wild-type and mutant strains are indicated as E, and sites where RNA editing was not detected in the mutant strain only are indicated as Un. (C) Prediction results are shown graphically. (D) Experimental verification of the target RNA editing sites of AHG11. The results of sequence analysis of the region containing mitochondrial nad4 are shown. RNA was extracted from the wild-type strain and the ahg11 mutant strain, and cDNA was prepared by reverse transcription and sequenced. Two RNA editing sites (nsd4362 and nsd376) are present in this region. The edited site is indicated by a black arrow, and the unedited site is indicated by a white arrow. [Figure 14] Figure 14 shows the prediction of target sites from the chloroplast genome sequence. Using six PPR proteins, target sites were predicted from the entire Arabidopsis chloroplast genome sequence (154,478 bp). For the prediction, codes extracted from Arabidopsis (At code) or codes extracted from Arabidopsis and moss (At + Pp code) were used. [Figure 15] FIG. 15 shows the RNA recognition codes of PPR motifs extracted from Arabidopsis thaliana and Physcomitrella patens. [Figure 16-1] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 16-2] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 16-3] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 16-4] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 16-5] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 16-6] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 16-7] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 16-8] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 16-9] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 16-10] FIG. 16 shows amino acid sequences or nucleotide sequences related to the present invention. [Figure 17] FIG. 17 shows the binding analysis of Enko B protein with RNA containing the cytoplasmic male sterility (CMS) gene. [Figure 18] FIG. 18 shows the binding of ORF687-like proteins to RNA. [Figure 19] FIG. 19 shows predicted binding sequences of fertility restorers that act on Ogura-type cytoplasm. [Figure 20] FIG. 20 shows the secondary structure and structural changes of the candidate binding RNA region of the ORF687-like protein. [Figure 21-1] FIG. 21 shows an alignment of ORF687-like proteins. [Figure 21-2] FIG. 21 shows an alignment of ORF687-like proteins. [Figure 22] FIG. 22 shows a list of amino acid sequences with base designations for ORF687-like proteins contained in various radish varieties. DETAILED DESCRIPTION OF THE INVENTION

[0019] [PPR motifs and PPR proteins] Unless otherwise specified, the term "PPR motif" as used herein refers to a polypeptide consisting of 30 to 38 amino acids whose E value obtained by analyzing amino acid sequences using online protein domain search programs (PF01535 in Pfam and PS51375 in Prosite) is a predetermined value or less (preferably E-03). The position numbers of the amino acids constituting the PPR motif as defined herein are nearly synonymous with those of PF01535, but correspond to the position of the amino acid in PS51375 minus two (e.g., position 1 in the present invention corresponds to position 3 in PS51375). However, when referring to the amino acid at position "ii" (-2), it refers to the second amino acid from the end (C-terminal) of the amino acid constituting the PPR motif, or the amino acid two amino acids N-terminal to amino acid 1 of the next PPR motif, i.e., the -2 amino acid (Figure 1). If the next PPR motif is not clearly identified, the amino acid two amino acids before the first amino acid in the next helix structure is designated as "ii." For more information on Pfam, see http: / / pfam.sanger.ac.uk / . For more information on Prosite, see http: / / www.expasy.org / prosite / .

[0020] The conserved amino acid sequence of the PPR motif is not highly conserved at the amino acid level, but the two α-helices in the secondary structure are well conserved. A typical PPR motif consists of 35 amino acids, but its length varies from 30 to 38 amino acids.

[0021] More specifically, the PPR motif in the present invention consists of a polypeptide of 30 to 38 amino acids in length, represented by formula 1.

[0022] [ka] During the ceremony: Helix A is a 12 amino acid long portion capable of forming an α-helical structure and is represented by Formula 2:

[0023] [ka] In formula 2, A1~A 12 each independently represents an amino acid; X is absent or a moiety consisting of 1 to 9 amino acids in length; Helix B is a portion consisting of 11 to 13 amino acids that can form an α-helical structure; L is a moiety of formula 3 that is 2 to 7 amino acids in length;

[0024] [ka] In Formula 3, each amino acid is numbered from the C-terminus, as "i" (-1), "ii" (-2), etc. However, L iii ~L vii may not exist.

[0025] In the present invention, the term "PPR protein" refers to a PPR protein having one or more, preferably two or more, of the above-mentioned PPR motifs, unless otherwise specified. In this specification, the term "protein" refers to any substance consisting of a polypeptide (a chain of multiple amino acids peptide-bonded together), unless otherwise specified, and includes those consisting of relatively low-molecular-weight polypeptides. In the present invention, the term "amino acid" may refer to a normal amino acid molecule, as well as to an amino acid residue constituting a peptide chain. It will be clear to those skilled in the art from the context which is being referred to.

[0026] PPR proteins are abundant in plants, with 500 proteins and approximately 5,000 motifs found in Arabidopsis. Many land plants, including rice, poplar, and Selaginella, also contain PPR motifs and PPR proteins with diverse amino acid sequences. Some PPR proteins are known to be important genes for F1 seed production, as fertility restorers that act in pollen formation (male gamete formation), resulting in hybrid vigor. Similar to fertility restorers, some PPR proteins have been shown to function in speciation. Most PPR proteins are also known to act on RNA in mitochondria or chloroplasts.

[0027] In animals, abnormalities in the PPR protein identified as LRPPRC are known to cause Leigh syndrome French Canadian (LSFC; subacute necrotizing encephalomyelopathy).

[0028] In the present invention, the term "selective" in relation to the binding of a PPR motif to an RNA base means that the binding activity for one of the RNA bases is higher than the binding activity for the other bases, unless otherwise specified. This selectivity can be confirmed by a person skilled in the art through experimental design, or it can be calculated as disclosed in the Examples of the present specification.

[0029] In the present invention, unless otherwise specified, the term "RNA base" refers to the base of a ribonucleotide that constitutes RNA, specifically any of adenine (A), guanine (G), cytosine (C), or uracil (U). Note that PPR proteins may have selectivity for bases in RNA, but do not bind to nucleic acid monomers.

[0030] Prior to the present invention, a method for searching sequences for conserved amino acids as PPR motifs had been established, but no rules regarding selective binding to RNA bases had been discovered.

[0031] The present invention provides the following findings.

[0032] (I) Information on the positions of amino acids important for selective binding. Specifically, the combination of three amino acids (A1, A4, L) in the PPR motif, 1, 4, and “ii”(-1), ii ), or a combination of two amino acids numbered 4, "ii"(-1) (A4, L ii ) are important for selective binding to RNA bases, and the combination of these determines which RNA base will bind.

[0033] The present invention relates to A1, A4, and L iiA combination of three amino acids, A4, and L ii This is based on knowledge of the combination of two amino acids.

[0034] (II) A 1 、A 4 , and L ii Information on the correspondence between the three amino acid combinations and the NA bases. Specifically, it is as follows: (3-1) A1, A4, and L ii When the combination of these three amino acids is, in order, valine, asparagine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to U, then C, and then A or G. (3-2) A1, A4, and L ii When the three amino acids in the PPR motif are, in order, valine, threonine, and asparagine, the PPR motif binds strongly to A, then G, then C, but not to U, showing selective RNA base binding ability. (3-3) A1, A4, and L ii When the combination of these three amino acids is, in order, valine, asparagine, asparagine, the PPR motif has selective RNA base-binding ability, binding strongly to C, then to A or U, but not to G. (3-4) A1, A4, and L ii When the combination of these three amino acids is, in order, glutamic acid, glycine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to G but not to A, U, or C. (3-5) A1, A4, and L ii When the three amino acid combinations are isoleucine, asparagine, and asparagine, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to C, then U, then A, but not to G. (3-6) A1, A4, and L iiWhen the three amino acids in the PPR motif are, in order, valine, threonine, and aspartic acid, the PPR motif binds strongly to G, then to U, but does not bind to A or C, showing selective RNA base binding ability. (3-7) A1, A4, and L ii When the three amino acid combinations are lysine, threonine, and aspartic acid, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to G, then to A, but not to U or C. (3-8) A1, A4, and L ii When the combination of these three amino acids is, in order, phenylalanine, serine, and asparagine, the PPR motif has selective RNA base-binding ability, binding strongly to A, then C, and then G and U. (3-9) A1, A4, and L ii When the three amino acid combinations are valine, asparagine, and serine, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to C, then to U, but not to A or G. (3-10) A1, A4, and L ii When the combination of these three amino acids is, in order, phenylalanine, threonine, and asparagine, the PPR motif has selective RNA base-binding ability, binding strongly to A but not to G, U, or C. (3-11) A1, A4, and L ii When the combination of these three amino acids is, in order, isoleucine, asparagine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to U, then to A, but not to G or C. (3-12) A1, A4, and L ii When the combination of these three amino acids is threonine, threonine, and asparagine, respectively, the PPR motif has selective RNA base binding ability, binding strongly to A but not to G, U, or C. (3-13) A1, A4, and Lii When the three amino acids in the PPR motif are, in order, isoleucine, methionine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to U, then C, but not A or G. (3-14) A1, A4, and L ii When the three amino acid combinations are phenylalanine, proline, and aspartic acid, respectively, the motif is called PPR. It binds strongly to U, then C, but not A or G, and has selective RNA base binding ability. (3-15) A1, A4, and L ii When the combination of these three amino acids is tyrosine, proline, and aspartic acid, respectively, the PPR motif has selective RNA base binding ability, binding strongly to U but not to A, G, or C. (3-16) A1, A4, and L ii When the combination of these three amino acids is leucine, threonine, and aspartic acid, respectively, the PPR motif has selective RNA base binding ability, binding strongly to G but not to A, U, or C.

[0035] (II) A 4 , and L ii Information on the correspondence between the combination of two amino acids and the NA base. Specifically, it is as follows: (2-1) A4, L ii However, in the case of asparagine and aspartic acid, the PPR motif has selective RNA base binding ability, binding strongly to U, then C, and then A and G. (2-2) A4, L ii However, in the case of asparagine, the PPR motif has selective RNA base binding ability, binding strongly to C, then U, and then A and G. (2-3) A4, L iiHowever, in the case of threonine and asparagine, the PPR motif binds strongly to A, and weakly to G, U, and C, respectively, and has selective RNA base binding ability. (2-4) A4, L ii However, in the case of threonine and aspartic acid, the PPR motif binds strongly to G, and then weakly to A, U, and C, showing selective RNA base binding ability. (2-5) A4, L ii However, in the case of serine and asparagine, the PPR motif binds strongly to A, followed by G, U, and C, showing selective RNA base binding ability. (2-6) A4, L ii However, in the case of glycine and aspartic acid, respectively, the PPR motif has selective RNA base binding ability, binding strongly to G, then U, then A, but not C. (2-7) A4, L ii However, in the case of asparagine and serine, the PPR motif binds strongly to C, then U, and then A and G, showing selective RNA base binding ability. (2-8) A4, L ii However, in the case of proline and aspartic acid, the PPR motif has selective RNA base binding ability, binding strongly to U, then G, C, and C, but not A. (2-9) A4, L ii However, in the case of glycine and asparagine, the PPR motif binds strongly to A, then to G, but does not bind to C or U, showing selective RNA base binding ability. (2-10) A4, L ii However, in the case of methionine and aspartic acid, the PPR motif has selective RNA base binding ability, binding strongly to U, followed by weaker binding to A, G, and C, respectively. (2-11) A4, L iiHowever, in the case of leucine and aspartic acid, the PPR motif binds strongly to C, then to U, but does not bind to A or G, showing selective RNA base binding ability. (2-12) A4, L ii However, in the case of valine and threonine, the PPR motif binds strongly to U, then to A, but does not bind to G or C, showing selective RNA base binding ability.

[0036] In the Examples herein, the above findings were obtained by further computational analysis of the binding between a protein and its potential RNA target sequence, which had been partially analyzed genetically or molecularly. More specifically, the binding or selective binding between a protein and RNA was analyzed using a P value (probability) as an index. In the present invention, a P value of 0.05 or less (5% or less chance of chance), which is a common significance level, preferably a P value of 0.01 or less (1% or less chance of chance), and more preferably a more significant P value, is evaluated as indicating a sufficiently high probability of binding between the protein and RNA. Such judgment based on P values ​​is well understood by those skilled in the art.

[0037] The binding affinity of a specific combination of amino acids at specific positions to an RNA base can be confirmed experimentally. Experiments for this purpose include preparation of a protein containing a PPR motif or multiple PPR motifs, preparation of substrate RNA, and binding tests (e.g., gel shift assay). Each of these experiments is well known to those skilled in the art, and for more specific procedures and conditions, see, for example, Patent Document 2.

[0038] [Utilization of PPR motifs and PPR proteins] Identification and Design: A PPR motif can recognize a specific base in RNA. According to the present invention, by selecting appropriate amino acids at specific positions, PPR motifs selective for A, U, G, and C can be selected or designed, and proteins containing an appropriate series of such PPR motifs can recognize the corresponding specific sequence. Therefore, according to the present invention, it is possible to predict and identify natural PPR proteins that selectively bind to RNAs with specific base sequences, and conversely, it is possible to predict and identify RNAs that are targets for PPR protein binding. Target prediction and identification is useful in clarifying genetic entities and expanding the availability of targets.

[0039] Furthermore, the present invention makes it possible to design a PPR motif capable of selectively binding to a desired RNA base, and a protein having multiple PPR motifs capable of sequence-specifically binding to a desired RNA. When designing, the sequence information of a natural PPR motif can be used as a reference for the portion of the PPR motif other than the amino acids at key positions. Alternatively, the design may be performed by using a natural PPR motif as a whole and substituting only the amino acids at the relevant positions. The number of repeats of the PPR motif can be appropriately determined depending on the target sequence, and can be, for example, 2 or more, e.g., 2 to 20.

[0040] In designing, consideration may be given to a combination of amino acids 1, 4, and "ii" or a combination of amino acids 4 and "ii." For example, consideration of amino acids 8 and 12 described in the aforementioned Patent Document 2 may be important in some cases for exhibiting RNA binding activity. According to the studies of the present inventors, the combination of A8 of a certain PPR motif and A9 of the same PPR motif 12 A8 may be a basic amino acid, preferably lysine, or an acidic amino acid, preferably aspartic acid; A 12 can be a basic amino acid, a neutral amino acid, or a hydrophobic amino acid.

[0041] Designed motifs or proteins can be prepared by methods well known to those skilled in the art. Specifically, the present invention provides PPR motifs that selectively bind to specific RNA bases, focusing on a combination of amino acids 1, 4, and "ii," or a combination of amino acids 4 and "ii," and PPR proteins that specifically bind to RNAs having specific sequences. In particular, when considering the effect on fertility as the function of the PPR protein, it has been found that amino acid 4 (A4) and amino acid ii are effective in either the above-mentioned combination of three amino acids or the above-mentioned combination of two amino acids. Such motifs and proteins can be prepared in relatively large quantities by methods well known to those skilled in the art. Such methods may include determining the nucleic acid sequence encoding the desired motif or protein from its amino acid sequence, cloning it, and generating a transformant that produces the desired motif or protein.

[0042] Preparation of the complex and its use: The PPR motifs or PPR proteins provided by the present invention can be linked to functional regions to form complexes. A functional region refers to a portion that has a specific biological function in a living organism or cell, such as an enzymatic function, catalytic function, inhibitory function, or enhancer function, or a portion that functions as a label. Such regions include proteins, peptides, nucleic acids, physiologically active substances, and drugs. An example of a protein functional region is ribonuclease (RNase). Examples of RNases include RNase A (e.g., bovine pancreatic ribonuclease A: PDB 2AAS) and RNase H. Such complexes do not exist in nature and are novel.

[0043] Furthermore, the complex obtained by the present invention may be able to deliver and function a functional region in an RNA sequence-specific manner in vivo or within cells. This may enable RNA sequence-specific modification or destruction in vivo or within cells, similar to zinc finger proteins (Non-Patent Document 1, supra) and TAL effectors (Non-Patent Document 2, supra; Patent Document 1, supra), and may also enable the conferring of new functions. Furthermore, it may enable the delivery of drugs in an RNA sequence-specific manner. Therefore, the present invention also provides a method for delivering RNA sequence-specific functional substances.

[0044] Some PPR proteins are known to be important as fertility restorers that function in pollen formation (male gamete formation) and in obtaining F1 seeds for hybrid vigor. The present invention is expected to lead to the identification of as-yet-unidentified fertility restorers and the development of technologies for utilizing these factors in an advanced manner. For example, in the case revealed in the examples of the present application, by detecting an amino acid polymorphism in a specific PPR motif in a PPR protein gene that functions as a fertility restorer for cytoplasmic male sterility, it is possible to determine whether the PPR protein gene of a test sample is a genotype related to fertility or a genotype related to sterility, based on the correlation between the polymorphism in the gene and fertility. In this case, the PPR protein genes to be detected for polymorphism include, for example, in the case of radish, family genes located at a locus homologous to the "ORF687 gene" encoding the ORF687 protein of Kosena (named Enko B), genes with 90% or more amino acid sequence identity to Enko B, and genes with 90% or more nucleotide sequence identity to the "ORF687 gene" encoding Enko B. Here, the family genes located at a locus homologous to the "ORF687 gene" encoding the ORF687 protein of Kosena (named Enko B) include the genes shown in Figures 21 and 22 (Kosena B, Comet B, Enko A, Comet A, Icicle CA, rrORF690-1, rrORF690-2, PC PPR-A, PC PPR motifs include, but are not limited to, all PPR motifs (e.g., PPR-BL). Genes with 90% or more amino acid sequence identity to Enko B, including the "ORF687 gene" encoding Enko B, and genes with 90% or more nucleotide sequence identity can be obtained by searching gene databases, and the species of origin is not limited to Japanese radish. The PPR motif may be a PPR motif consisting of a 30-38 amino acid polypeptide represented by Formula 1 above. A PPR protein may be characterized by containing one or more (preferably 2-16) such PPR motifs. Polymorphisms in this PPR motif may refer to polymorphisms based on the combination of amino acids 1, 4, and "ii," or the combination of amino acids 4 and "ii," which have been shown in the present invention to be responsible for RNA binding in each PPR motif. As shown by the P values ​​calculated in Figures 4B and 4D, among the amino acid combinations responsible for RNA binding in such PPR motifs, amino acid 4 plays the most important role, followed by amino acid ii. Furthermore, by comparing the PPR proteins of Enko B, we demonstrated that the identity of amino acid 4 of all PPR motifs in the proteins encoded by the tested genes to Enko B, or the identity of amino acid ii of all corresponding PPR motifs to Enko B, is important for the function of the fertility restorer. Furthermore, similar to fertility restoration, several PPR proteins have been shown to function in speciation. Identification and modification of the target RNAs of these PPR proteins raises the possibility of previously impossible interspecies hybridization. Furthermore, because most PPR proteins act on RNAs in mitochondria and chloroplasts, the novel PPR proteins provided by the present invention will likely contribute to the modification and improvement of functions related to photosynthesis, respiration, and the synthesis of useful metabolites.

[0045] Meanwhile, in animals, abnormalities in the PPR protein identified as LRPPRC are known to cause Leigh syndrome French Canadian (LSFC; subacute necrotizing encephalomyelopathy). The present invention may contribute to the treatment (prevention, treatment, and inhibition of progression) of LSFC.

[0046] Furthermore, PPR proteins are involved in all RNA processing steps found in organelles, including cleavage, RNA editing, translation, splicing, and RNA stabilization. The present invention is expected to enable the expression of desired RNAs to be altered by modifying the binding base selectivity of PPR motifs.

[0047] The PPR proteins used in the present invention function exclusively to specify editing sites in RNA editing (conversion of genetic information on RNA; in most cases, C → U) (see References 2 and 3 below). This type of PPR protein has an additional motif on the C-terminus that is suggested to interact with an RNA-mutagenase. PPR proteins with such a structure are expected to be useful for introducing nucleotide polymorphisms or treating diseases or conditions caused by nucleotide polymorphisms.

[0048] Some PPR proteins also have RNA cleavage enzymes attached to their C-terminus. By modifying the RNA base binding selectivity of the N-terminal PPR motif, RNA sequence-specific RNA cleavage enzymes can be constructed. Furthermore, complexes with labeled moieties such as GFP can be used to visualize desired RNAs in vivo.

[0049] On the other hand, some existing PPR proteins act on DNA. One is reported to be a transcriptional activator of mitochondrial genes, and the other is a transcriptional activator localized in the nucleus. Therefore, based on the findings of this invention, it may be possible to design protein factors that bind to desired DNA sequences. [Example]

[0050] Example 1: Collection of PPR proteins involved in RNA editing and their target sequences With reference to the information shown in Figure 2, the PPR proteins involved in RNA editing in Arabidopsis that have been analyzed so far (SEQ ID NOs: 2 to 24) were listed in the Arabidopsis Genome Information Database (MATDB: http: / / mips.gsf.de / proj / thal / db / index.html), and the sequences surrounding the target RNA editing sites (SEQ ID NOs: 48, 50, 53, 55, 57, 59, 60, 61, 62, 63, 64, 65, 68, 69, 70, 71, 73, 74, 76, 78, 80, 122, 206, 228, 232, 252, 284, 316, 338, 339, 358, 430, 433, 455, 552, 563) were listed in the RNA Editing Database (http: / / biologia.unical.it / py The RNA sequences were collected from the nucleotide sequence database (http: / / www.ncbi.nlm.nih.gov / script / overview.html). The RNA sequences were collected from the 31 bases upstream of the edited C (cytosine) residue. All collected proteins and their corresponding RNA editing sites are shown in Figure 2.

[0051] The PPR motif structures in the proteins were assigned amino acid numbers as defined in the present invention, along with information from the Uniprot database (http: / / www.uniprot.org / ). The PPR motifs contained in the 24 Arabidopsis PPR proteins (SEQ ID NOS: 2 to 25) used in the experiment and their amino acid numbers are shown in Figure 3.

[0052] Example 2: Identification of amino acids that confer binding base selectivity Previous studies have revealed that PPR proteins involved in RNA editing contain motifs with specific conserved amino acid sequences (E, E+, and DYW motifs, although DYW is often absent) at their C-terminus. It has been suggested that the dozen or so amino acids in the E+ motif are required for C (cytosine) to U (uracil) conversion rather than for selective binding to RNA (Reference 3). Furthermore, a previous non-patent paper suggested that the information required for recognizing the C to be edited is contained within 20 bases upstream and 5 bases downstream of the C. Thus, it is predicted that multiple PPR motifs in PPR proteins recognize "somewhere" in the upstream sequence of the C to be edited, with the E+ motif located near the C to be edited. Furthermore, it is possible that specific amino acids in the PPR motif recognize RNA residues in the upstream sequence to which they bind (Figure 4A).

[0053] This possibility was examined using the 24 RNA-editing PPR proteins from Arabidopsis thaliana described in Example 1 and their target RNA sequences. First, the last PPR motif in the protein was positioned at the first base of the C to be edited, and all PPR motifs were aligned with the RNA residues in a one-to-one correspondence and linear sequence (Figure 4A, alignment P1). Next, the RNA sequence was shifted one base to the right to obtain alignments P2 to P6. Information on the RNA residues corresponding to each PPR motif was collected from the data set of alignments P1 to P6.

[0054] For PPR proteins that edit RNA at one site, each RNA residue (A, U, G, or C) was assigned a score of 1. For PPR proteins that edit RNA at two or three sites, each RNA residue was assigned a score of 0.5 or 0.3, respectively. Next, the amino acid sequence was sorted by amino acid type for each amino acid number in the PPR motif. Generally, the amino acid sequence and the RNA residue sequence are predicted to be random (high-randomness or high-entropy) (e.g., Figure 4A, upper right). However, if amino acids at specific sites have the ability to select RNA bases, the correct alignment (P1–P6 above) predicts that the corresponding RNA bases will converge to one or a limited number of types (low randomness or low entropy; e.g., Figure 4A, lower right).

[0055] For the alignment data set of P1 to P6 created above, we calculated the low randomness for all amino acid numbers in the PPR motif. Low randomness was calculated by a chi-square test against the theoretical value (average of all base occurrence frequencies) (e.g., Figure 5).

[0056] As a result, a significance value of P<0.01 (probability of less than 1%) was calculated for amino acids 1, 4, and "ii" (-2) in the P4 alignment (Figure 4B). This indicates that the final PPR motif in the RNA-editing PPR protein is located four bases before the edited C, and that three amino acids (1, 4, and "ii") are responsible for RNA base selection. Furthermore, no significant P values ​​were calculated for the P3 and P5 alignments, indicating no interference from the flanking PPR motifs; that is, each PPR motif recognizes a single RNA residue, independent of motif structure. No significant P values ​​were obtained for the other amino acids in alignment P4 or any of the other alignments (Figure 6). Furthermore, when similar calculations were performed using RNA bases classified as purines (A and G) or pyrimidines (C and U) (RY), only amino acid 4 obtained a highly significant P value (P<0.01) (Figure 4C). This indicates that amino acid 4 primarily distinguishes between purine and pyrimidine bases in the RNA to which it binds. Further analysis of the base-specific binding ability of the RNA-recognizing amino acids in the PPR motif shown in Figure 4C was performed. The results showed that, in addition to primarily distinguishing between purine and pyrimidine bases (RY) to which it binds, amino acid "ii" (-2) also functions to distinguish between amino (A and C) and keto (G and U) forms of the base (MK) (Figure 4D).

[0057] Combinations of three amino acids (number 1, number 4, and number "ii") that occur three or more times were designated as triPPR codes, and their P values ​​were calculated to assess their ability to specify RNA bases. Some of the identified triPPR codes are shown in Figure 4E.

[0058] Because the amino acids at these three positions are highly diverse, we calculated the RNA base-specific binding ability of two amino acids (1&4, 1&"ii", or 4&"ii") and found a significant P value for the combination of amino acids 4&"ii" (Figure 7). Therefore, the combination of amino acids 4&"ii" that was used three or more times was designated as the diPPR code, which is one of the RNA recognition codes for the PPR motif. The identified triPPR code and diPPR code are shown in Figure 8.

[0059] Example 3: Verification of the identified RNA recognition code The RNA recognition code of the PPR motif identified using the RNA-editing PPR protein of Arabidopsis thaliana was verified. For the verification, the RNA-editing PPR protein of Physcomitrella patens was used. It has already been revealed that RNA editing occurs at a total of 13 sites (11 sites in mitochondria and 2 sites in chloroplasts; SEQ ID NOs: 32 to 44) in Physcomitrella patens (hereinafter referred to as moss). Furthermore, it was found that six PPR proteins (PpPPR 56, 71, 77, 78, 79, and 91) have been shown to function in RNA editing at nine sites. The proteins and their corresponding RNA editing sites are shown in Figure 9.

[0060] The verification was performed as shown in Figure 10. First, the amino acid sequence information of the moss PPR protein was obtained from a non-patent literature (SEQ ID NOS: 26-31; Figures 2 and 9). Three amino acids (1, 4, and "ii") were extracted from each PPR motif according to the PPR motif model defined in Figure 1. If the combination of the three extracted amino acids matched the triPPR code identified from Arabidopsis thaliana, it was replaced with the binding base scoring matrix represented by that code. Next, among the PPR motifs that could not be converted using the triPPR code, if they matched the diPPR code, the motif was replaced with the binding base scoring matrix of the diPPR code. In parallel, sequences surrounding the RNA editing site (31-mer sequences with the edited C at the 3' end) were obtained from a non-patent literature (SEQ ID NOS: 32-44; Figures 2, 9, and 16) and replaced with the numerical matrix of the RNA sequence shown in Figure 10. Next, following the alignment P4 (corresponding to the four bases before the C where the final PPR motif is edited), the protein binding base score matrix and the RNA sequence number matrix were multiplied by each square, and the sum of the obtained values ​​was calculated as the matching score between the protein and the RNA sequence. This calculation was performed using the triPPR code, diPPR code, and their respective PPR binding base score matrices (PPR scoring matrix). It was held at.

[0061] This calculation was performed for all 13 RNA editing sites in moss for one protein. Similar calculations were also performed for 34 RNA sequences in the RNA editing sites of Arabidopsis chloroplasts (Figure 16, SEQ ID NOs: 45 to 78) as reference sequences for the sequences surrounding the RNA editing sites.

[0062] Next, a normal distribution curve was drawn from the protein match values ​​for each RNA sequence, and a provisional P value for the match value for each RNA sequence was calculated for each triPPR code and diPPR code.

[0063] The final P value (matching score between protein and RNA sequence) was calculated as the product of the provisional P values ​​of the triPPR code and the diPPR code.

[0064] Figure 11 shows the matching scores for each moss PPR protein and the 13 moss RNA editing sites. As a result of the analysis, six of the seven proteins were computationally identified as having the correct RNA editing site. This analysis indicates that three amino acids (1, 4, and "ii") contain all the information required for the PPR motif to specify the bases that bind to the RNA. In other words, this indicates that by referencing the amino acid information for the three or two combinations of amino acids shown in Figure 8 (triPPR, diPPR code), it is possible to search for PPR proteins that bind to a specific RNA sequence. This also indicates that by using or linking PPR motifs containing this amino acid information, it is possible to synthesize artificial proteins that bind to a specific RNA sequence.

[0065] Example 4: Identification of target molecules of uncharacterized RNA-editing PPR proteins Next, we performed an analysis using Arabidopsis thaliana, which contains more RNA editing sites than moss (34 sites in the chloroplast genome (SEQ ID NOs: 45-78) and 488 sites in the mitochondrial genome (SEQ ID NOs: 79-566) (see Figure 6). To verify the accuracy of our predictions, we predicted the RNA variant sites of the 24 PPR proteins used in the code extraction. As a result, for chloroplast-localized PPR proteins, 10 out of 13 proteins predicted at least one correct RNA editing site with the highest P value. For mitochondrial-localized PPR proteins, 8 out of 11 proteins predicted correct RNA editing sites within the top 20 (Figure 12). Based on this verification of prediction accuracy, we predicted the target RNA editing sites of PPR proteins with unknown functions. The AHG11 mutant is a mutant that exhibits abnormalities in the abscisic acid pathway, and the protein encoded by its gene (ahg11, at2g44880) has a typical RNA-editing PPR protein-like motif structure (Figure 13; SEQ ID NO: 1). We predicted RNA editing sites and experimentally verified the RNA editing of 405 mitochondrial and 30 chloroplast sites, including the top 20. As a result, the mitochondrial nad4 predicted with the seventh highest P value was Only 376 RNA editing sequences were found to be abnormal in the mutants (Fig. 13).

[0066] Next, we obtained the complete organelle genome sequence, i.e., approximately 3 × 10 5 We attempted to identify target RNA sequences from this RNA sequence dataset. For this analysis, we used the probability matrix of the PPR code shown in Figure 8. In addition, background frequencies were applied to motifs with amino acid combinations that did not match the diPPR or triPPR codes. The probability matrix of the protein we created was used in FIMO analysis of the MEME suite (http: / / meme.nbcr.net / meme4) together with the complete chloroplast sequence of Arabidopsis thaliana (AP000423). 6 1 / fimo-intro.html).

[0067] As a result, we were able to accurately predict the target RNA sequences of CRR4 and CRR21. Furthermore, we improved the PPR codes by extracting them from moss PPR proteins (Figure 15), and the prediction accuracy improved significantly for several proteins.

[0068] These results demonstrate that by using the identified PPR code, it is possible to identify a single correct target sequence from hundreds of thousands of RNA sequence patterns. Conversely, by searching for PPR motifs that have amino acids in accordance with the code at the relevant positions (1, 4, "ii"), it is possible to identify proteins that bind to a desired useful RNA sequence. Alternatively, it demonstrates that by linking PPR motifs, it is possible to create artificial RNA-binding proteins with high sequence selectivity. Those skilled in the art will also understand that by introducing mutations, it is possible to obtain the desired RNA-binding selectivity by combining amino acids at the relevant positions in accordance with the PPR code.

[0069] In Figure 15, the RNA base selectivity of triPPR code and diPPR code was evaluated using P values. PPR codes that showed significant P values ​​(P < 0.05) can be inferred to have high RNA base selectivity.

[0070] Example 5: Prediction of target RNA sequences of radish Rf Next, based on the findings obtained in the present invention, the function of the PPR protein, which acts as a fertility restorer for cytoplasmic male sterility, was assessed (Examples 5 to 9).

[0071] Cytoplasmic male sterility (CMS) is a trait in which male gametes do not function normally due to mutations in the cytoplasmic genome, particularly the mitochondrial genome. This trait can often be counteracted by a fertility restorer (Rf) gene present in the nucleus, resulting in normal male gametes. It is used in hybrid breeding and is an important agricultural trait. In this CMS-Rf system, the Rf gene is often known to encode a PPR protein.

[0072] The Ogura (also known as Kosena) cytoplasm used in the hybrid breeding of radish and rapeseed is derived from the expression of the orf125 gene in the mitochondrial genome, and the presence of the nuclear-encoded orf687 gene overcomes sterility, making the plant fertile. The orf687 gene product is a PPR protein, which is thought to act on the RNA containing orf125, thereby inactivating its expression and thereby overcoming sterility.

[0073] However, previous breeding analyses have revealed that orf687-like genes in various radish lines contain amino acid polymorphisms, and that these amino acid polymorphisms affect the functionality of the genes as fertility restorers. However, no method has been established to infer their functionality from the amino acid sequences of the genes.

[0074] Therefore, we first identified the PPR motif from the amino acids of the ORF687 protein (named Enko B) of the radish cultivar Enko, which is known to function as a dominant Rf. We then extracted the amino acids (1, 4, ii) responsible for base specification and converted them into a PPR code. After that, we performed target RNA sequence prediction for transcripts containing mitochondrial orf125 (Figure 19).

[0075] In parallel, we biochemically analyzed the characteristics of three ORF687-like proteins: the ORF687 protein (named Enko B) from the radish cultivar Sonobeni, which is known to function as a dominant Rf; an ORF687-like protein (named enko A) also found in Sonobeni and which is very similar to ORF687 but functions as a recessive gene; and a gene (named kosena B; a recessive gene) homologous to Sonobeni ORF687 present in the genome of a different radish cultivar, Kosena.

[0076] (5-1) Preparation of genomic DNA from radish Radish plants were cultured for three weeks in Murashige and Sook medium (containing 2% sucrose and 0.5% gellanthus). Green leaves (0.5 g) from the cultured plants were extracted with phenol / chloroform, and then ethanol was added to insolubilize the DNA. The recovered DNA was dissolved in 100 μl of TE solution (10 mM Tris-HCl (pH 8.0), 1 mM EDTA), and 10 units of RNase A (DNase-free, Takara Bio) was added. The mixture was incubated at 37°C for 30 minutes. The reaction mixture was then extracted again with phenol / chloroform, and the DNA was recovered by ethanol precipitation. 10 μg of DNA was obtained.

[0077] (5-2) Cloning of the gene encoding ORF687-like protein Radish genomic DNA was used as a template, and Enko B was used as an oligonucleotide primer (Enko BF primer and Enko BR primers; set forth in SEQ ID NOs: 567 and 568, respectively), Kosena B is an oligonucleotide primer (kosena BF primer and kosena BR primers; set forth in SEQ ID NOs: 569 and 570, respectively), Enko A is an oligonucleotide primer (Enko AF Primer and Enko AR primers (sequence numbers 571 and 572, respectively) were used to amplify each fragment by PCR using 50 μl of the reaction solution at 95°C for 30 seconds, 60°C for 30 seconds, and 72°C for 30 seconds for 25 cycles, with KOD-FX (TOYOBO) as the DNA elongation enzyme.

[0078] The resulting DNA fragments were cloned into the pBAD / Thio-TOPO vector (Invitrogen) according to the protocol provided with the product. The DNA sequences were determined and confirmed to be homologous to the target DNA sequences (Enko B, SEQ ID NO: 573; Kosena B, SEQ ID NO: 574; Enko A, SEQ ID NO: 575).

[0079] (5-3) Preparation of recombinant ORF687-like protein The resulting plasmid was transformed into Escherichia coli TOP10 (Invitrogen). The E. coli was cultured at 37°C in 300 ml of LB medium (1 L Erlenmeyer flask containing 300 ml of medium) containing 100 μg / ml ampicillin. When the turbidity of the culture medium reached an absorbance of 0.5 at 600 nm, the inducer L-arabinose was added to a final concentration of 0.2%, and the culture was continued for an additional 4 hours.

[0080] After harvesting by centrifugation, the cells were suspended in 200 ml of buffer A (50 mM Tris-HCl pH 8.0, 500 mM KCl, 2 mM imidazole, 10 mM MgCl2, 0.5% Triton X100, 10% glycerol) containing 1 mg / ml lysozyme, and disrupted by sonication and freeze-thawing. After centrifugation at 15,000 × g for 20 minutes, the supernatant was collected as a crude extract.

[0081] This crude extract was applied to a column packed with nickel column resin (ProBond A, Invitrogen) equilibrated with buffer A.

[0082] Column chromatography was performed using a two-step gradient, washing thoroughly with Buffer A containing 20 mM imidazole, followed by elution of the target protein with Buffer A containing 200 mM imidazole. The resulting protein is a fusion protein with the amino acid sequence set forth in SEQ ID NO: (Enko B, SEQ ID NO: 576; Kosena B, SEQ ID NO: 577; Enko A, SEQ ID NO: 578), a thioredoxin amino acid sequence at the N-terminus for increased solubility, and a histidine tag sequence at the C-terminus. A 100 μl aliquot of the purified fraction was dialyzed against 500 mL of Buffer E (20 mM Tris-HCl pH 7.9, 60 mM KCl, 12.5 mM MgCl2, 0.1 mM EDTA, 17% glycerol, 2 mM DTT) and used as a purified preparation.

[0083] (5-4) Preparation of substrate RNA As substrate RNAs, three types of RNAs, RNAa, RNAb, and RNAc, containing the mitochondrial DNA sequence of Ogura-type cytoplasmic radish were used.

[0084] RNAa was amplified using oligonucleotide primers AF and AR (SEQ ID NOs: 579 and 580, respectively), RNAb using oligonucleotide primers BF and BR (SEQ ID NOs: 581 and 582, respectively), and RNAc using oligonucleotide primers CF and CR (SEQ ID NOs: 583 and 584). A 50-μl reaction mixture containing 10 ng of Ogura-type cytoplasmic radish DNA as template was subjected to PCR amplification for 25 cycles at 95°C for 30 seconds, 60°C for 30 seconds, and 72°C for 30 seconds using KOD FX (TOYOBO) as the DNA elongation enzyme. Each forward primer (-F) contained a T7 promoter sequence for in vitro synthesis of substrate RNA.

[0085] The resulting DNA fragment was purified by running it on an agarose gel and then cutting it out from the gel. The purified DNA fragment was used as a template in a 500-kDa PCR reaction with NTP mix (10 nmol GTP, CTP, ATP, 0.5 nmol UTP), 4 μl [ 32 P] α-UTP (GE Healthcare, 3000 Ci / mmol) and T7 RNA polymerase (Takara Bio) were added to 20 μl of the reaction solution, which was then reacted at 37°C for 60 minutes to synthesize substrate RNA.

[0086] The substrate RNA was extracted with phenol / chloroform and precipitated with ethanol. The entire volume was then subjected to electrophoresis on a denaturing 6% polyacrylamide gel containing 6 M urea, and exposed to X-ray film for 60 seconds. 32 P-labeled RNA was detected.

[0087] next, 32 P-labeled RNA was excised from the gel and immersed in 200 μl of gel elution buffer (0.3 M sodium acetate, 2.5 mM EDTA, 0.01% SDS) at 4°C for 12 hours to elute the RNA. The radioactivity of 1 μl of the RNA was measured, and the total amount of synthesized RNA was calculated. After ethanol precipitation, the RNA was dissolved in ultrapure water to a concentration of 2500 cpm / μl (1 fmol / μl). This preparation method typically yielded approximately 100 μl of RNA at 2500 cpm / μl.

[0088] (5-5) Protein-RNA binding experiment Recombinant proteins of Enko B (Rf), Kosena B (rf), and Enko A (rf; an ORF687-like protein present in Kosena cultivar) were produced and their RNA-binding activities were examined.

[0089] The RNA binding activity of the recombinant proteins (Enko B (SEQ ID NO: 576), Kosena B (SEQ ID NO: 577), and Enko A (SEQ ID NO: 578)) was analyzed by gel shift analysis. 375 pM (7.5 fmol / 20 μL) of the substrate RNA (BD120) and 0–2500 nM of the recombinant protein were mixed in 20 μL of reaction solution (10 mM Tris-HCl pH 7.9, 30 mM KCl, 6 mM MgCl2, 2 mM DTT, 8% glycerol, 0.0067% Triton X-100) and incubated at 25°C for 15 minutes. Four μL of 80% glycerol was then added to the reaction mixture, and 10 μL of the mixture was loaded onto a 10% native polyacrylamide gel containing 1x TBE (89 mM Tris-HCl, 89 mM boric acid, 2 mM EDTA). The gel was then dried after electrophoresis.

[0090] The radioactivity of the RNA in the gel was measured using a bioimaging analyzer BAS2000 (Fujifilm).

[0091] Example 6: RNA binding experiments using recombinant proteins Figure 17 shows the binding analysis of Enko B protein with RNA containing the cytoplasmic male sterility (CMS) gene. Figure 17A shows a schematic diagram of the vicinity of mitochondrial orf125, along with a schematic diagram of the regions of RNA a, RNA bc, RNA b, and RNA c used in the binding experiment. Figure 17B shows the RNA binding of Enko B protein. Enko B protein (1.4 nmol) and 32 A gel shift competition experiment was performed by reacting P-labeled RNA bc (0.1 ng) with unlabeled RNA a, RNA bc, RNA b, and RNA c (used as competitive inhibitors at 5 and 10 w / w concentrations relative to RNA bc) in a 20 μL reaction mixture. Complex (▽) on the left of the figure represents a complex of protein and RNA, while Free (▼) represents RNA alone.

[0092] As shown in these figures, the binding of proteins and RNA occurs as follows: 32 This appears as a difference in the mobility of P-labeled RNA. 32The molecular weight of the P-labeled RNA-protein complex is 32 This is because the molecular weight of the P-labeled RNA is larger than that of the unlabeled RNA, resulting in slower electrophoretic mobility. In this experiment, recombinant EnkoB protein was prepared and its binding to mitochondrial RNA containing orf125 was examined by competitive gel shift analysis. RI-labeled RNAb and protein were mixed, and then unlabeled RNA was added. The greater the decrease in signal intensity of the band indicated by "Complex," the greater the binding of the competitor RNA and protein. This indicates that this is the RNA region to which EnkoB binds with high affinity. The results demonstrated that EnkoB binds strongly to the region of RNAb.

[0093] The candidate sequence for No. 208 showed the most significant P value in the predicted binding sequence shown in Figure 19, and it is located exactly at the 3' end of the tRNA methionine. However, previous analyses have shown that there is no difference in the amount of tRNA or the shape of the RNA containing orf125 (presence or absence of cleavage) between sterile and restored lines, and in vitro binding experiments (Figure 17B) have shown that the RNAa sequence containing No. 208 does not bind to Enko B, so we have concluded that this region is not related to the fertility or sterility of Ogura-type cytoplasm.

[0094] Therefore, we focused on the regions 316, 352, and 373 contained in RNAb. RNAb consists of 125 base pairs. We attempted to narrow down the binding region to 20 base pairs by scanning mutations, but were unable to narrow it down to a single site (data not shown). Therefore, we considered the possibility that there are multiple Enko B binding sites in RNAb.

[0095] Example 7: RNA binding activity of Rf-like proteins Figure 18 shows the binding of ORF687-like proteins to RNA. Figure 18A shows the results of gel shift analysis of the binding of Enko B (Rf), Kosena B (rf), and Enko A (rf) to RNAb, which is an important factor in determining the RNA-binding properties of ORF687-like proteins. Figure 18B is a graph of the results of (A), from which the dissociation constant (KD), which represents the RNA-binding ability of each protein, was calculated. Figure 18C shows the match scores of Enko B (Rf), Kosena B (rf), and Enko A (rf) to potential binding regions, calculated in a manner similar to that used in Figure 19.

[0096] As a result, under non-competitive conditions, all three proteins (Enko B, Kosena B, and Enko A) bound to RNA with high affinity to RNA. When the RNA-binding activity of Kosena B was analyzed under competitive conditions, no clear difference was observed between it and Enko B (Fig. 18A and 18B).

[0097] Kosena B often exhibits slightly lower RNA-binding activity than Enko B (approximately 2-fold lower KD). However, the difference in RNA-binding activity is often detected as more than 10-fold, so this difference cannot be considered significant.

[0098] Even in predictions based on the PPR code, there was no clear difference in the match scores for the corresponding region between the proteins (Figure 18C). Based on this, we decided to investigate the possibility that the difference between EnkoB and kosenaB was not simply a difference in RNA-binding affinity, but rather a difference in the action after binding.

[0099] Furthermore, Figure 19 shows the predicted binding sequence of the fertility restorer that acts on the Ogura cytoplasm. Figure 19A shows the predicted binding sequence of the Enko B protein using the PPR code, and the lower diagram in Figure 19A shows the structure of the RNA containing the CMS gene orf125. For the regions RNAa to RNAc in Figure 19A, see Figure 17. In Figure 19A, we focused on regions 208, 230, 316, 352, and 373, which showed significantly high P values ​​(Figure 19A).

[0100] Next, Figure 19B shows the logo of the target RNA sequence predicted from the ORF687 protein sequence (the sequence of the regions (Nos. 208, 316, 352, and 373) that showed significant P values), candidate binding RNA sequences, and the logo of the target RNA sequence predicted from the ORF687-like protein sequence (Kosena B), a radish cultivar with recessive rf. The predicted binding bases of Kosena B, a recessive rf, are also shown.

[0101] EnkoB and KosenaB were found to have different bases designated by amino acid polymorphisms in the second and third PPR motifs (Rf is UA, rf is GC). This difference was predicted to be directly linked to the functional differences between Rf and rf.

[0102] Example 8: RNA structure prediction and analysis Computer prediction and in vitro RNA binding experiments suggested that Rf might bind to regions of RNAb, particularly positions 316, 352, and 373. In vitro analysis suggested that there might be multiple binding sites within RNAb. Therefore, we performed secondary structure prediction of the RNAb sequence and focused on the relevant regions.

[0103] The results are shown in Figure 20. Figure 20 shows the secondary structure and structural changes of the candidate binding RNA region of ORF687-like protein. Figure 20A shows the secondary structure of the region containing No. 306 and the predicted binding site of ORF687-like protein. Each PPR motif is indicated by a box with the corresponding base. The second and third PPR motifs, which show significant differences between EnkoB (Rf) and Kosena B (rf), are highlighted. Figure 20B shows the secondary structure of the region containing Nos. 352 and 373 and the predicted binding site of ORF687-like protein. Figure 20C shows the structural changes of RNAb by Enko B. RNAb and Enko B protein were mixed, and then a double-strand-selective RNase (RNase V1) was added.

[0104] The results revealed that region 316 corresponds to the stem-loop structure immediately downstream of the initiation codon of orf125 (Fig. 20A). Furthermore, PPR motifs 2 and 3, which are polymorphic between Enko B and Kosena B, are located in the double-stranded region at the base of the stem-loop. In particular, the corresponding base of PPR motif 3 is A in Enko B, whereas it is C in Kosena B (Fig. 19B). Based on these findings, a working hypothesis was proposed that EnkoB binds to this region, promotes the formation of a stem-loop structure, and thereby inhibits translation of orf125.

[0105] A double-stranded structure was also predicted for regions 352 and 373, suggesting that Rf protein binds to both sides (Figure 20B). However, in this case, Rf binding would be expected to disrupt the structure (promote single-strand formation). Furthermore, differences in the bases and structure corresponding to PPR motifs 2 and 3, which are the differences between Rf and rf, were not considered, and no specific molecular mechanism could be predicted.

[0106] Therefore, we mixed the internally labeled RNA with protein and added RNase V1 to perform limited degradation of the labeled RNA. RNase V1 is an RNase that selectively cleaves only the double-stranded region of RNA. The results showed that the substrate RNA was rapidly degraded in the presence of protein, i.e., the formation of double-stranded RNA was promoted in the presence of Rf (Enko B) (Fig. 20C). Thus, we suggest that the translational inhibition of orf125 mRNA by Rf, resulting in double-stranded RNA formation, is the main cause of fertility restoration in Ogura cytoplasmic male sterility.

[0107] Example 9: Functional assessment of the fertility restoration ability of ORF687-like genes ORF687-like genes have been isolated from various radish cultivars, and their Rf functionality has been inferred through crossbreeding experiments. However, the amino acid sequences of each gene are very similar, and it is not possible to determine their functionality as Rf based on overall amino acid conservation.

[0108] In this example, first, sequence analysis of ORF687-like proteins was performed. Specifically, sequence analysis was performed as a PPR protein using the protein sequences shown in SEQ ID NOs: 576 to 578 and 585 to 591. All sequences were used as query sequences to obtain sequence alignments using CLUSTALW (http: / / www.genome.jp / tools / clustalw / ). Web-based domain analysis software, Pfam (http: / / pfam.sanger.ac.uk / ), InterProScan (http: / / www.ebi.ac.uk / Tools / InterProScan / ), Prosite(http: / / www.expasy.org / prosite / ), We used this to create an alignment of ORF687-like proteins and analyzed the PPR motif structure of each protein. The results are shown in Figure 21. All ORF687-like genes are composed of 16 PPR motifs (Figure 21).

[0109] According to the obtained PPR motif model and the amino acid numbering shown in Non-Patent Document 5, amino acids 1, 2, and "ii" (-2) were extracted and used to determine the function of the fertility restoration ability of the ORF-like protein.

[0110] Therefore, we evaluated the functions of nine Rf-like genes using PPR codes. As with EnkoB, we extracted the amino acids (1, 4, and ii) responsible for base selection and converted them into PPR codes. The amino acid sequences were then used to determine their functionality as RNA binding windows (Figure 22). EnkoB and KosenaB share 99.4% overall homology, but two amino acid polymorphisms exist within the RNA binding window, which are thought to be deeply involved in the dominant / recessive nature of ORF687-like genes for fertility restoration (Non-patent Paper 4). On the other hand, in the Comet cultivar, CometB, which resides at a locus homologous to EnkoB, shares 98.0% homology with EnkoB and has an identical RNA binding window. This supported the finding that CometB is a dominant gene obtained in previous crossbreeding experiments. Furthermore, EnkoA, a duplicate gene located near EnkoB, is also suggested to be recessive in terms of RNA recognition. These data suggest that the dominant / recessive nature of ORF687-like genes for fertility restoration depends on the amino acids (1, 4, ii) that control base specification ability, especially the same 4th amino acid (A4) or the same iith amino acid, in all the corresponding PPR motifs in ORF687-like genes, when compared with the RNA binding window of Enko B. It is especially important to have the same 4th amino acid (A4). From this point of view, it is possible to compare the dominant / recessive nature of ORF687-like genes for fertility restoration in various radish lines whose fertility information is unknown, with the genes rrORF690-1, rrORF690-2, and icicle, which are located at loci homologous to Enko B. pprCA, P.C. PPR-A, PC PPR-BL has a different RNA binding window from the dominant gene Enko B, and these genes are also considered to be recessive rf genes.

[0111] These results suggest that the PPR code described in this invention can speed up the functional determination of industrially useful PPR proteins, such as those that function as fertility restorers. This technology makes it possible to determine the fertility restorer ability of candidate Rf genes from their sequences when applying new lines to hybrid breeding methods using the CMS-Rf system. The inventors have performed functional determinations of ORF687-like genes in 21 new radish varieties and have successfully determined the dominant / recessive fertility restorer ability of 19 ORF-like genes (data not published). This technology is applicable not only to radishes with Ogura-type cytoplasm, but also to various cytoplasms and plant species that use PPR proteins as Rf.

[0112] [Papers cited in the examples] Reference 1: Small, ID, and Peeters, N. (2000). The PPR motif - a TPR-related motif prevalent in plant organellar proteins. Trends Biochem. Sci. 25, 46-47. Reference 2: Lurin, C., Andres, C., Aubourg, S., Bellaoui, M., Bitton, F., Bruyere, C., Caboche, M., Debast, C., Gualberto, J., Hoffmann, B., et al. (2004). Genome-wide analysis of Arabidopsis pentatricopeptide proteins repeat reveals their essential role in organelle biogenesis. Plant Cell 16, 2089-2103. Reference 3: Okuda, K., Myouga, F., Motohashi, R., Shinozaki, K., and Shikanai, T. (2007). Conserved domain structure of pentatricopeptide repeat proteins involved in chloroplast RNA editing. Proc Natl Acad Sci USA 104, 8178-8183. Reference 4: Koizuka N, Imai R, Fujimoto H, Hayakawa T, Kimura Y, et al. (2003) Genetic characterization of a pentatricopeptide repeat protein gene, orf687, that restores fertility in the cytoplasmic male-sterile Kosena radish. Plant J 34: 407-415. Reference 5: Nakamura T, Yagi Y, Kobayashi K (2012) Mechanistic insight into pentatricopeptide repeat proteins as sequence-specific RNA-binding proteins for organellar RNAs in plants. Plant & Cell Physiology 53: 1171-1179

Claims

1. A composition comprising a nucleic acid encoding a PPR protein, a composition for use in treating a disease mediated by any modification selected from the group consisting of translation, splicing, and stability of an RNA having a target sequence to which the PPR protein can bind; The PPR protein contains 2 to 20 PPR motifs, which are polypeptides having a length of 30 to 38 amino acids and are represented by the following formula 1: 【Chemistry 1】 (In formula 1: Helix A is a 12 amino acid long portion capable of forming an α-helix structure, Formula 2: 【Chemistry 2】 (In formula 2, A 1 ~A 12 each independently represents an amino acid), In Formula 1, X is absent or a moiety consisting of 1 to 9 amino acids in length; In Formula 1, Helix B is a portion consisting of 11 to 13 amino acids and capable of forming an α-helix structure, In formula 1, L is a 2-7 amino acid long molecule of formula 3: 【Transformation 3】 (In Formula 3, each amino acid is numbered from the C-terminal side as "i" (-1), "ii" (-2), etc.) However, L iii ~L vii may not exist.) A PPR motif in the PPR protein that selectively binds to one RNA base in the RNA having the target sequence 1 , A 4 , and L ii A composition characterized in that the combination of the three amino acids is any of the following: (3-1) A, which constitutes a PPR motif that selectively binds to U (uracil) 1 , A 4 , and L ii The combination of these three amino acids is, in order, valine, asparagine, and aspartic acid; (3-2) A, which constitutes the PPR motif that selectively binds to A (adenine) 1 , A 4 , and L ii The combination of these three amino acids is, in order, valine, threonine, and asparagine; (3-3) A, which constitutes a PPR motif that selectively binds to C (cytosine) 1 , A 4 , and L ii a combination of three amino acids in the order of valine, asparagine, and asparagine; (3-4) A, which constitutes the PPR motif that selectively binds to G (guanine) 1 , A 4 , and L ii The combination of these three amino acids is, in order, glutamic acid, glycine, and aspartic acid; (3-5) A, which constitutes a PPR motif that selectively binds to C or U 1 , A 4 , and L ii The combination of these three amino acids is, in order, isoleucine, asparagine, asparagine; (3-6) A, which constitutes the PPR motif that selectively binds to G 1 , A 4 , and L ii The combination of these three amino acids is, in order, valine, threonine, and aspartic acid; (3-7) A, which constitutes a PPR motif that selectively binds to G 1 , A 4 , and L ii The combination of these three amino acids is, in order, lysine, threonine, and aspartic acid; (3-8) A, which constitutes a PPR motif that selectively binds to A 1 , A 4 , and L ii The combination of these three amino acids is, in order, phenylalanine, serine, and asparagine; (3-9) A, which constitutes a PPR motif that selectively binds to C 1 , A 4 , and L ii The combination of the three amino acids is, in order, valine, asparagine, and serine; (3-10) A, which constitutes a PPR motif that selectively binds to A 1 , A 4 , and L ii The combination of these three amino acids is, in order, phenylalanine, threonine, and asparagine; (3-11) A, which constitutes a PPR motif that selectively binds to U or A 1 , A 4 , and L ii The combination of these three amino acids is, in order, isoleucine, asparagine, and aspartic acid; (3-12) A, which constitutes a PPR motif that selectively binds to A 1 , A 4 , and L ii The combination of these three amino acids is, in order, threonine, threonine, asparagine; (3-13) A, which constitutes a PPR motif that selectively binds to U or C 1 , A 4 , and L ii The combination of these three amino acids is, in order, isoleucine, methionine, and aspartic acid; (3-14) A, which constitutes a PPR motif that selectively binds to U 1 , A 4 , and L ii The combination of these three amino acids is, in order, phenylalanine, proline, and aspartic acid; (3-15) A, which constitutes a PPR motif that selectively binds to U 1 , A 4 , and L ii The combination of these three amino acids is, in order, tyrosine, proline, and aspartic acid; (3-16) A, which constitutes a PPR motif that selectively binds to G 1 , A 4 , and L ii The combination of these three amino acids is, in order, leucine, threonine, and aspartic acid.

2. A composition comprising a nucleic acid encoding a PPR protein, a composition for use in treating a disease mediated by any modification selected from the group consisting of translation, splicing, and stability of an RNA having a target sequence to which the PPR protein can bind; the PPR protein comprises 2 to 20 PPR motifs, which are polypeptides having a length of 30 to 38 amino acids and represented by formula 1 as defined in claim 1; A PPR motif in the PPR protein that selectively binds to one RNA base in the RNA having the target sequence 4 , and L ii A composition characterized in that the combination of two amino acids is any of the following: (2-1) A, which constitutes the PPR motif that selectively binds to U 4 , and L ii The combination of these two amino acids is, in order, asparagine and aspartic acid; (2-2) A, which constitutes a PPR motif that selectively binds to C 4 , and L ii a combination of two amino acids in the order of asparagine, asparagine; (2-3) A, which constitutes a PPR motif that selectively binds to A 4 , and L ii The combination of these two amino acids is, in order, threonine and asparagine; (2-4) A, which constitutes the PPR motif that selectively binds to G 4 , and L ii The combination of these two amino acids is, in order, threonine and aspartic acid; (2-5) A, which constitutes a PPR motif that selectively binds to A 4 , and L ii The combination of these two amino acids is, in order, serine and asparagine; (2-6) A, which constitutes the PPR motif that selectively binds to G 4 , and L ii The combination of these two amino acids is, in order, glycine and aspartic acid; (2-7) A, which constitutes a PPR motif that selectively binds to C 4 , and L ii The combination of these two amino acids is, in order, asparagine and serine; (2-8) A, which constitutes a PPR motif that selectively binds to U 4 , and L ii The combination of these two amino acids is, in order, proline and aspartic acid; (2-9) A, which constitutes a PPR motif that selectively binds to A 4 , and L ii The combination of these two amino acids is, in order, glycine and asparagine; (2-10) A, which constitutes a PPR motif that selectively binds to U 4 , and L ii The combination of these two amino acids is, in order, methionine and aspartic acid; (2-11) A, which constitutes a PPR motif that selectively binds to C 4 , and L ii The combination of these two amino acids is, in order, leucine and aspartic acid; (2-12) A, which constitutes a PPR motif that selectively binds to U 4 , and L ii The combination of these two amino acids is, in order, valine and threonine.

3. The composition according to claim 1 or 2, wherein the PPR protein comprises 2 to 16 PPR motifs represented by formula 1.

4. A composition comprising a nucleic acid encoding a PPR protein, a composition for use in treating a disease mediated by any modification selected from the group consisting of translation, splicing, and stability of an RNA having a target sequence to which the PPR protein can bind; the PPR protein comprises 2 to 20 PPR motifs, which are polypeptides having a length of 30 to 38 amino acids and represented by formula 1 as defined in claim 1; A 1 , A 4 , and L ii The composition is characterized in that the combination of the three amino acids is any one of (3-1) to (3-16) defined in claim 1, depending on one RNA base in the RNA having the target sequence.

5. A composition comprising a nucleic acid encoding a PPR protein, a composition for use in treating a disease mediated by any modification selected from the group consisting of translation, splicing, and stability of an RNA having a target sequence to which the PPR protein can bind; the PPR protein comprises 2 to 20 PPR motifs, which are polypeptides having a length of 30 to 38 amino acids and represented by formula 1 as defined in claim 1; A 4 , and L ii The composition is characterized in that the combination of the two amino acids is any one of (2-1) to (2-12) defined in claim 2, depending on one RNA base in the RNA having the target sequence.

6. The composition according to claim 4 or 5, wherein the PPR protein comprises 2 to 16 PPR motifs represented by formula I.

Citation Information

Patent Citations

  • TAL effector-mediated DNA modification

    WO2011072246A2

  • Method for modifying RNA binding protein using PPR motif

    WO2011111829A1