Efficient PPR protein production method and use thereof
Novel PPR motifs and proteins with enhanced binding capabilities address the limitations of existing PPR proteins by allowing specific binding to longer RNA sequences, improving RNA manipulation and splicing efficiency.
Patent Information
- Application Number
- JP2025076735
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-05-29
- Filing Date
- 2025-05-02
- Publication Date
- 2025-07-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing PPR proteins are limited in their ability to specifically bind to long nucleotide sequences and perform desired operations in cells, as they typically consist of 7 to 14 motifs, which is insufficient for targeting specific sequences in complex genomes like the human genome with 6 billion bases.
Development of novel PPR motifs with specific amino acid substitutions and modifications, such as SEQ ID NO: 9, 10, 11, and 12, allowing for PPR proteins to bind to longer RNA sequences by incorporating 15 or more motifs, enhancing binding performance and enabling operations like RNA splicing and detection.
The novel PPR motifs and proteins demonstrate improved binding specificity and affinity to target RNA sequences, facilitating efficient RNA manipulation and splicing, even in complex genetic environments.
Smart Images

Figure 2025107351000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a nucleic acid manipulation technique using a protein capable of binding to a target nucleic acid. The present invention is useful in a wide range of fields such as medicine (drug discovery support, treatment), agriculture (agricultural, forestry, fishery and livestock product production, breeding), and chemistry (biological substance production).
Background Art
[0002] The PPR protein is a protein containing a repeat of PPR motifs, each about 35 amino acids long, and one PPR motif can specifically bind to one base. The combination of the first, fourth, and ii-th (two before the next motif) amino acids in the PPR motif determines which of adenine, cytosine, guanine, uracil (or thymine) it binds to (Patent Documents 1 and 2).
[0003] Since one PPR motif recognizes and binds to one base, for example, when designing a PPR protein that binds sequence-specifically to an 18-base-long nucleic acid, 18 PPR motifs need to be linked. So far, the production of artificial PPR proteins with 7 to 14 linked PPR motifs has been reported (Non-Patent Documents 1 to 6).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] In order to specifically bind to a target RNA molecule in a cell and further perform desired operations, a PPR protein with high performance is required.
[0007] In addition, in order to specifically bind to a target RNA molecule in a cell and further perform desired operations, a PPR protein that links more motifs than the conventional 7 to 14 and binds to a long base sequence is required. For example, the human genome has 6 billion bases and is composed of four bases (A, C, G, T or U). Therefore, in order to specify a single base sequence from the sequence order, at least a 17-base sequence is required (4 16 to the power of 40 billion, 4 17 to the power of 16 billion).
Means for Solving the Problems
[0008] The present invention provides the following as novel PPR motifs and the like. [1] Any one of the following PPR motifs: (A-1) The PPR motif consisting of the sequence of SEQ ID NO: 9, or in the sequence of SEQ ID NO: 9, a substitution of the amino acid at position 10 with tyrosine, a substitution of the amino acid at position 15 with lysine, a substitution of the amino acid at position 16 with leucine, a substitution of the amino acid at position 17 with glutamic acid, a substitution of the amino acid at position 18 with aspartic acid, and a substitution of the amino acid at position 28 with glutamic acid; a PPR motif consisting of an amino acid sequence having any substitution selected from the group consisting of these substitutions; a PPR motif consisting of the sequence of SEQ ID NO: 401, or in the sequence of SEQ ID NO: 401, a substitution of the amino acid at position 10 with tyrosine, a substitution of the amino acid at position 16 with leucine, a substitution of the amino acid at position 17 with glutamic acid, a substitution of the amino acid at position 18 with aspartic acid, and a substitution of the amino acid at position 28 with glutamic acid; a PPR motif consisting of an amino acid sequence having any substitution selected from the group consisting of these substitutions; (A-2) A PPR motif consisting of a sequence in which 1 to 20 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 9 or 401 are substituted, deleted, or added, and which is adenine-binding; (A-3) A PPR motif having at least 42% sequence identity with the sequence of SEQ ID NO: 9 or 401, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are identical, and which is adenine-binding; (C-1) The PPR motif consisting of the sequence of SEQ ID NO: 10, or in the sequence of SEQ ID NO: 10, a substitution of the amino acid at position 2 with serine, a substitution of the amino acid at position 5 with isoleucine, a substitution of the amino acid at position 7 with leucine, a substitution of the amino acid at position 8 with lysine, a substitution of the amino acid at position 10 with phenylalanine or tyrosine, a substitution of the amino acid at position 15 with arginine, a substitution of the amino acid at position 22 with valine, a substitution of the amino acid at position 24 with arginine, a substitution of the amino acid at position 27 with leucine, and a substitution of the amino acid at position 29 with arginine; a PPR motif consisting of an amino acid sequence having any substitution selected from the group consisting of these substitutions; (C-2) A PPR motif consisting of a sequence in which 1 to 25 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 10 are substituted, deleted, or added, and which is cytosine-binding; (C-3) A PPR motif having at least 25% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is cytosine-binding; (G-1) A PPR motif consisting of the sequence of SEQ ID NO: 11, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with phenylalanine, substitution of the amino acid at position 15 with aspartic acid, substitution of the amino acid at position 27 with valine, substitution of the amino acid at position 28 with serine, and substitution of the amino acid at position 35 with isoleucine is made in the sequence of SEQ ID NO: 11; (G-2) A PPR motif consisting of a sequence in which 1 to 21 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 11 are substituted, deleted, or added, and which is guanine-binding; (G-3) A PPR motif having at least 40% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is guanine-binding; (U-1) A PPR motif consisting of the sequence of SEQ ID NO: 12, or in the sequence of SEQ ID NO: 12, any substitution selected from the group consisting of substitution of the amino acid at position 10 with phenylalanine, substitution of the amino acid at position 13 with serine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 20 with leucine, substitution of the amino acid at position 21 with lysine, substitution of the amino acid at position 23 with phenylalanine, substitution of the amino acid at position 24 with aspartic acid, substitution of the amino acid at position 27 with lysine, substitution of the amino acid at position 28 with lysine, substitution of the amino acid at position 29 with arginine, and substitution of the amino acid at position 31 with leucine, and a PPR motif consisting of the resulting amino acid sequence (U-2) A PPR motif consisting of a sequence in which 1 to 22 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 12 are substituted, deleted, or added, and which is uracil-binding (U-3) A PPR motif having at least 37% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are the same, and which is uracil-binding [2] Use of the PPR motif according to 1 for the production of a PPR protein in which the target RNA is 15 bases or longer [3] Use for the production of a PPR protein of the PPR motif according to 1, which is for enhancing the binding performance of the PPR protein to the target RNA [4] A protein containing n PPR motifs capable of binding to a target RNA consisting of n base sequences, wherein the PPR motif for adenine in the base sequence is the PPR motif of (A-1), (A-2), or (A-3) defined in 1; wherein the PPR motif for cytosine in the base sequence is the PPR motif of (C-1), (C-2), or (c-3) defined in 1; wherein the PPR motif for guanine in the base sequence is the PPR motif of (G-1), (G-2), or (G-3) defined in 1; A PPR protein in which the PPR motif for uracil in the base sequence is the PPR motif of (U-1), (U-2), or (U-3) defined in 1. [5] The protein according to 4, wherein n is 15 or more. [6] The protein according to 4 or 5, wherein the first PPR motif from the N-terminus is any one of the following: (1st A-1) A PPR motif in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 402 are substituted so as to satisfy any one of the combinations defined below; (1st A-2)(1st A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of A-1) are substituted, deleted, or added, and which is adenine-binding; (1st A-3)(1st A PPR motif having at least 80% sequence identity with the sequence of A-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and which is adenine-binding; (1st C-1) A PPR motif consisting of the sequence of SEQ ID NO: 403; (1st C-2)(1st A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of C-1) are substituted, deleted, or added, and which is cytosine-binding; (1st C-3)(1st A PPR motif having at least 80% sequence identity with the sequence of C-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and which is cytosine-binding; (1st G-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 404 are substituted so as to satisfy any one of the combinations defined below; (1st G-2)(1st A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of G-1 are substituted, deleted, or added, and which is guanine-binding; (1st G-3)(1st A PPR motif having at least 80% sequence identity with the sequence of G-1, provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and which is guanine-binding; (1st U-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 405 are substituted so as to satisfy any one of the combinations defined below; (1st U-2)(1st A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of U-1 are substituted, deleted, or added, and which is uracil-binding; (1st U-3)(1st A PPR motif having at least 80% sequence identity with the sequence of U-1, provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and which is uracil-binding. · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamic acid · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamine · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is lysine · A combination in which the amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine [7] A method for controlling RNA splicing, characterized by using the protein according to any one of items 4 to 6. [8] A method for detecting RNA, characterized by using the protein according to any one of items 4 to 6. [9] A fusion protein comprising at least one selected from the group consisting of a fluorescent protein, a nuclear localization signal peptide, and a tag protein, and the protein according to any one of items 4 to 6.
[10] The PPR motif according to 1, or a nucleic acid encoding the protein according to any one of items 4 to 6.
[11] A vector comprising the nucleic acid according to 10.
[12] A cell (excluding human individuals) comprising the vector according to 11.
[13] A method for manipulating RNA (excluding implementation in human individuals) using the PPR motif according to 1, the protein according to any one of items 4 to 6, or the vector according to 11.
[14] A method for producing an organism comprising the manipulation method according to 13.
[0009] [1] Any one of the following PPR motifs: (A-1) A PPR motif consisting of the sequence of SEQ ID NO: 9, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid is made in the sequence of SEQ ID NO: 9; (A-2) A PPR motif consisting of a sequence in which 1 to 20 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 9 are substituted, deleted, or added, and which is adenine-binding; (A-3) A PPR motif having at least 42% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are the same, and which is adenine-binding; (C-1) A PPR motif consisting of the sequence of SEQ ID NO: 10, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the serine at position 2 of the sequence of SEQ ID NO: 10 with serine, substitution of the isoleucine at position 5 with isoleucine, substitution of the leucine at position 7 with leucine, substitution of the lysine at position 8 with lysine, substitution of the phenylalanine or tyrosine at position 10 with phenylalanine or tyrosine, substitution of the arginine at position 15 with arginine, substitution of the valine at position 22 with valine, substitution of the arginine at position 24 with arginine, substitution of the leucine at position 27 with leucine, and substitution of the arginine at position 29 with arginine is performed; (C-2) A PPR motif consisting of a sequence in which 1 to 25 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 10 are substituted, deleted, or added, and which is cytosine-binding; (C-3) A PPR motif having at least 25% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is cytosine-binding; (G-1) A PPR motif consisting of the sequence of SEQ ID NO: 11, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the phenylalanine at position 10 of the sequence of SEQ ID NO: 11 with phenylalanine, substitution of the aspartic acid at position 15 with aspartic acid, substitution of the valine at position 27 with valine, substitution of the serine at position 28 with serine, and substitution of the isoleucine at position 35 with isoleucine is performed; (G-2) A PPR motif consisting of a sequence in which 1 to 21 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 11 are substituted, deleted, or added, and which is guanine-binding; (G-3) A PPR motif having at least 40% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is guanine-binding; (U-1) A PPR motif consisting of the sequence of SEQ ID NO: 12, or in the sequence of SEQ ID NO: 12, any substitution selected from the group consisting of substitution of the amino acid at position 10 with phenylalanine, substitution of the amino acid at position 13 with serine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 20 with leucine, substitution of the amino acid at position 21 with lysine, substitution of the amino acid at position 23 with phenylalanine, substitution of the amino acid at position 24 with aspartic acid, substitution of the amino acid at position 27 with lysine, substitution of the amino acid at position 28 with lysine, substitution of the amino acid at position 29 with arginine, and substitution of the amino acid at position 31 with leucine, a PPR motif consisting of an amino acid sequence with such a substitution (U-2) A PPR motif consisting of a sequence in which 1 to 22 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 12 are substituted, deleted, or added, and which is uracil-binding; (U-3) A PPR motif having at least 37% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are the same, and which is uracil-binding. [2] Use of the PPR motif according to 1 for the production of a PPR protein in which the target RNA is 15 bases or longer. [3] Use for the production of a PPR protein of the PPR motif according to 1, for enhancing the binding performance of the PPR protein to the target RNA. [4] A protein containing n PPR motifs capable of binding to a target RNA consisting of n base sequences, wherein the PPR motif for adenine in the base sequence is the PPR motif of (A-1), (A-2), or (A-3) defined in 1; wherein the PPR motif for cytosine in the base sequence is the PPR motif of (C-1), (C-2), or (c-3) defined in 1; wherein the PPR motif for guanine in the base sequence is the PPR motif of (G-1), (G-2), or (G-3) defined in 1; A PPR protein in which the PPR motif for uracil in the nucleotide sequence is the PPR motif of (U-1), (U-2), or (U-3) defined in 1. [5] The protein according to 4, wherein n is 15 or more. [6] A method for controlling RNA splicing, comprising using the protein according to 4 or 5. [7] A method for detecting RNA, comprising using the protein according to 4 or 5. [8] A fusion protein of at least one selected from the group consisting of a fluorescent protein, a nuclear localization signal peptide, and a tag protein, and the protein according to 4 or 5. [9] A nucleic acid encoding the PPR motif according to 1, or the protein according to 4 or 5.
[10] A vector containing the nucleic acid according to 9.
[11] A cell (excluding human individuals) containing the vector according to 10.
[12] A method for manipulating RNA (excluding implementation in human individuals), using the PPR motif according to 1, the protein according to 4 or 5, or the vector according to 10.
[13] A method for producing an organism, comprising the manipulation method according to 12.
[14] A method for producing a gene encoding a protein containing n PPR motifs capable of binding to a target nucleic acid consisting of n nucleotide sequences, comprising the following steps: Selecting m PPR parts necessary for producing a target gene from a library of at least 20 × m types of PPR parts, wherein each of at least 20 types of polynucleotides encoding each of the PPR motifs having adenine, cytosine, guanine, or uracil or thymine binding properties, and each of 16 types of linkers of two PPR motifs is inserted into each of the intermediate vectors Dest-a... designed to be ligatable in at least m types of order; Subjecting the selected m types of PPR parts to a Golden Gate reaction together with vector parts to obtain a vector into which a linker of m polynucleotides is inserted (where n is not less than m and not more than m × 2).
[15] The production method according to 14 for producing a gene encoding a protein containing 15 or more PPR motifs, where m is 10.
[16] A method for detecting or quantifying a protein containing n PPR motifs that can bind to a target nucleic acid consisting of n base sequences, comprising the following steps: A step of providing a solution containing a candidate protein to the immobilized target nucleic acid and detecting or quantifying the protein bound to the target nucleic acid.
[17] The method according to 16, wherein the candidate protein is fused with a labeled protein.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Mode for Carrying Out the Invention
[0011] [PPR Motif, PPR Protein] (Definition) When referring to the PPR motif in the present invention, unless otherwise specified, when analyzing the amino acid sequence with a protein domain search program on the web, it refers to a polypeptide composed of 30 to 38 amino acids having an amino acid sequence with an E-value obtained from PF01535 in Pfam and PS51375 in Prosite being less than or equal to a predetermined value (preferably E-03). The position numbers of the amino acids constituting the PPR motif defined in the present invention are almost synonymous with PF01535, while corresponding to the number obtained by subtracting 2 from the position of the amino acid in PS51375 (e.g., position 1 in the present invention → position 3 in PS51375). However, when referring to the amino acid at the "ii" (-2) position, it refers to the second amino acid from the end (C-terminal side) of the amino acids constituting the PPR motif, or two amino acids on the N-terminal side with respect to the first amino acid of the next PPR motif, that is, the -2 amino acid. When the next PPR motif cannot be clearly identified, the amino acid two positions before the first amino acid of the next helix structure is defined as "ii". For Pfam, refer to http: / / pfam.sanger.ac.uk / , and for Prosite, refer to http: / / www.expasy.org / prosite / .
[0012] The conserved amino acid sequence of the PPR motif has low conservation at the amino acid level, but the two α-helices are well conserved in the secondary structure. A typical PPR motif is composed of 35 amino acids, but its length is variable from 30 to 38 amino acids.
[0013] More specifically, the PPR motif referred to in the present invention consists of a polypeptide having a length of 30 to 38 amino acids represented by Formula 1.
[0014]
Chemical formula
[0015]
Chemical formula
[0016]
Chemical formula
[0017] When referring to a PPR protein in the present invention, unless otherwise specified, it refers to a PPR protein having one or more, preferably two or more of the above-described PPR motifs. When referring to a protein in this specification, unless otherwise specified, it refers to all substances composed of polypeptides (chains in which a plurality of amino acids are peptide-bonded), and includes those composed of relatively low-molecular-weight polypeptides. When referring to an amino acid in the present invention, it may refer to a normal amino acid molecule, or may refer to an amino acid residue constituting a peptide chain. Which one is being referred to is clear to those skilled in the art from the context.
[0018] In the present invention, regarding the binding property to a base in the target nucleic acid of a PPR motif, when referring to specificity / specific, unless otherwise specified, it means that the binding activity to any one of the four types of bases is higher than the binding activity to other bases.
[0019] When referring to a nucleic acid in the present invention, it refers to RNA or DNA. Note that a PPR protein may have specificity for a base in RNA or DNA, but does not bind to a nucleic acid monomer.
[0020] The PPR motif is such that the combination of three amino acids at positions 1, 4, and ii is important for specific binding to bases, and these combinations can determine which base is bound (Patent Documents 1 and 2 cited above).
[0021] Specifically, regarding the RNA-binding PPR motif, the relationship between the combination of three amino acids at positions 1, 4, and ii and the bases that can bind is as follows (see Patent Document 1 cited above). (3-1) For the combination of three amino acids A1, A4, and L ii in the order of valine, asparagine, and aspartic acid, the PPR motif has selective RNA base-binding ability to strongly bind to U, then to C, and then to A or G. (3-2) For the combination of three amino acids A1, A4, and L ii in the order of valine, threonine, and asparagine, the PPR motif has selective RNA base-binding ability to strongly bind to A, then to G, and then to C, but not to U. (3-3) For the combination of three amino acids A1, A4, and L ii in the order of valine, asparagine, and asparagine, the PPR motif has selective RNA base-binding ability to strongly bind to C, and then to A or U, but not to G. (3-4) For the combination of three amino acids A1, A4, and L ii in the order of glutamic acid, glycine, and aspartic acid, the PPR motif has selective RNA base-binding ability to strongly bind to G, but not to A, U, or C. (3-5) For the combination of three amino acids A1, A4, and L ii in the order of isoleucine, asparagine, and asparagine, the PPR motif has selective RNA base-binding ability to strongly bind to C, then to U, and then to A, but not to G. (3-6) For the combination of three amino acids A1, A4, and L iiWhen the combination of three amino acids is valine, threonine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G and then to U, but does not bind to A and C. (3-7) A1, A4, and L ii When the combination of three amino acids is lysine, threonine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G and then to A, but does not bind to U and C. (3-8) A1, A4, and L ii When the combination of three amino acids is phenylalanine, serine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, then to C, and then to G and U. (3-9) A1, A4, and L ii When the combination of three amino acids is valine, asparagine, and serine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to C and then to U, but does not bind to A and G. (3-10) A1, A4, and L ii When the combination of three amino acids is phenylalanine, threonine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, but does not bind to G, U, and C. (3-11) A1, A4, and L ii When the combination of three amino acids is isoleucine, asparagine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to U and then to A, but does not bind to G and C. (3-12) A1, A4, and L ii When the combination of three amino acids is threonine, threonine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, but does not bind to G, U, and C. (3-13) A1, A4, and Lii When the combination of three amino acids is isoleucine, methionine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to U and then to C, but does not bind to A and G. (3-14) A1, A4, and L ii When the combination of three amino acids is phenylalanine, proline, and aspartic acid in that order for the PPR, its motif has selective RNA base-binding ability such that it binds strongly to U and then to C, but does not bind to A and G. (3-15) A1, A4, and L ii When the combination of three amino acids is tyrosine, proline, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to U, but does not bind to A, G, and C. (3-16) A1, A4, and L ii When the combination of three amino acids is leucine, threonine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G, but does not bind to A, U, and C.
[0022] Specifically, regarding the DNA-binding PPR motif, the relationship between the combination of three amino acids at positions 1, 4, and ii and the bases to which they can bind is as follows (see Patent Document 2 cited above). (2-1) A1, A4, and L ii When the combination of three amino acids is any amino acid, glycine, and aspartic acid in that order, the PPR motif selectively binds to G; (2-2) A1, A4, and L ii When the combination of three amino acids is glutamic acid, glycine, and aspartic acid in that order, the PPR motif selectively binds to G; (2-3) A1, A4, and L ii When the combination of three amino acids is any amino acid, glycine, and asparagine in that order, the PPR motif selectively binds to A; (2-4) Combinations of three amino acids, A1, A4, and L ii When the combination of the three amino acids is, in order, glutamic acid, glycine, and asparagine, the PPR motif selectively binds to A; (2-5) Combinations of three amino acids, A1, A4, and L ii When the combination of the three amino acids is, in order, any amino acid, glycine, and serine, the PPR motif selectively binds to A and then binds to C; (2-6) Combinations of three amino acids, A1, A4, and L ii When the combination of the three amino acids is, in order, any amino acid, isoleucine, and any amino acid, the PPR motif selectively binds to T and C; (2-7) Combinations of three amino acids, A1, A4, and L ii When the combination of the three amino acids is, in order, any amino acid, isoleucine, and asparagine, the PPR motif selectively binds to T and then binds to C; (2-8) Combinations of three amino acids, A1, A4, and L ii When the combination of the three amino acids is, in order, any amino acid, leucine, and any amino acid, the PPR motif selectively binds to T and C; (2-9) Combinations of three amino acids, A1, A4, and L ii When the combination of the three amino acids is, in order, any amino acid, leucine, and aspartic acid, the PPR motif selectively binds to C; (2-10) Combinations of three amino acids, A1, A4, and L ii When the combination of the three amino acids is, in order, any amino acid, leucine, and lysine, the PPR motif selectively binds to T; (2-11) Combinations of three amino acids, A1, A4, and L ii When the combination of the three amino acids is, in order, any amino acid, methionine, and any amino acid, the PPR motif selectively binds to T; (2-12) Combinations of three amino acids, A1, A4, and L ii When the combination of the three amino acids is, in order, any amino acid, methionine, and aspartic acid, the PPR motif selectively binds to T; (2-13) The combination of three amino acids, A1, A4, and L ii When the combination of the three amino acids is isoleucine, methionine, and aspartic acid in that order, the PPR motif selectively binds to T and then binds to C; (2-14) The combination of three amino acids, A1, A4, and L ii When the combination of the three amino acids is any amino acid, asparagine, and any amino acid in that order, the PPR motif selectively binds to C and T; (2-15) The combination of three amino acids, A1, A4, and L ii When the combination of the three amino acids is any amino acid, asparagine, and aspartic acid in that order, the PPR motif selectively binds to T; (2-16) The combination of three amino acids, A1, A4, and L ii When the combination of the three amino acids is phenylalanine, asparagine, and aspartic acid in that order, the PPR motif selectively binds to T; (2-17) The combination of three amino acids, A1, A4, and L ii When the combination of the three amino acids is glycine, asparagine, and aspartic acid in that order, the PPR motif selectively binds to T; (2-18) The combination of three amino acids, A1, A4, and L ii When the combination of the three amino acids is isoleucine, asparagine, and aspartic acid in that order, the PPR motif selectively binds to T; (2-19) The combination of three amino acids, A1, A4, and L ii When the combination of the three amino acids is threonine, asparagine, and aspartic acid in that order, the PPR motif selectively binds to T; (2-20) The combination of three amino acids, A1, A4, and L ii When the combination of the three amino acids is valine, asparagine, and aspartic acid in that order, the PPR motif selectively binds to T and then binds to C; (2-21) The combination of three amino acids, A1, A4, and L iiWhen the combination of three amino acids is tyrosine, asparagine, and aspartic acid in sequence, the PPR motif selectively binds to T and then binds to C; (2-22) A1, A4, and L ii When the combination of three amino acids is any amino acid, asparagine, and asparagine in sequence, the PPR motif selectively binds to C; (2-23) A1, A4, and L ii When the combination of three amino acids is isoleucine, asparagine, and asparagine in sequence, the PPR motif selectively binds to C; (2-24) A1, A4, and L ii When the combination of three amino acids is serine, asparagine, and asparagine in sequence, the PPR motif selectively binds to C; (2-25) A1, A4, and L ii When the combination of three amino acids is valine, asparagine, and asparagine in sequence, the PPR motif selectively binds to C; (2-26) A1, A4, and L ii When the combination of three amino acids is any amino acid, asparagine, and serine in sequence, the PPR motif selectively binds to C; (2-27) A1, A4, and L ii When the combination of three amino acids is valine, asparagine, and serine in sequence, the PPR motif selectively binds to C; (2-28) A1, A4, and L ii When the combination of three amino acids is any amino acid, asparagine, and threonine in sequence, the PPR motif selectively binds to C; (2-29) A1, A4, and L ii When the combination of three amino acids is valine, asparagine, and threonine in sequence, the PPR motif selectively binds to C; (2-30) A1, A4, and L iiWhen the combination of three amino acids is, in order, any amino acid, asparagine, and tryptophan, the PPR motif selectively binds to C and then binds to T; (2-31) A1, A4, and L ii When the combination of three amino acids is, in order, isoleucine, asparagine, and tryptophan, the PPR motif selectively binds to T and then binds to C; (2-32) A1, A4, and L ii When the combination of three amino acids is, in order, any amino acid, proline, and any amino acid, the PPR motif selectively binds to T; (2-33) A1, A4, and L ii When the combination of three amino acids is, in order, any amino acid, proline, and aspartic acid, the PPR motif selectively binds to T; (2-34) A1, A4, and L ii When the combination of three amino acids is, in order, phenylalanine, proline, and aspartic acid, the PPR motif selectively binds to T; (2-35) A1, A4, and L ii When the combination of three amino acids is, in order, tyrosine, proline, and aspartic acid, the PPR motif selectively binds to T; (2-36) A1, A4, and L ii When the combination of three amino acids is, in order, any amino acid, serine, and any amino acid, the PPR motif selectively binds to A and G; (2-37) A1, A4, and L ii When the combination of three amino acids is, in order, any amino acid, serine, and asparagine, the PPR motif selectively binds to A; (2-38) A1, A4, and L ii When the combination of three amino acids is, in order, phenylalanine, serine, and asparagine, the PPR motif selectively binds to A; (2-39) A1, A4, and L iiWhen the combination of three amino acids is valine, serine, and asparagine in that order, the PPR motif selectively binds to A; (2-40) A1, A4, and L ii When the combination of three amino acids is any amino acid, threonine, and any amino acid in that order, the PPR motif selectively binds to A and G; (2-41) A1, A4, and L ii When the combination of three amino acids is any amino acid, threonine, and aspartic acid in that order, the PPR motif selectively binds to G; (2-42) A1, A4, and L ii When the combination of three amino acids is valine, threonine, and aspartic acid in that order, the PPR motif selectively binds to G; (2-43) A1, A4, and L ii When the combination of three amino acids is any amino acid, threonine, and asparagine in that order, the PPR motif selectively binds to A; (2-44) A1, A4, and L ii When the combination of three amino acids is phenylalanine, threonine, and asparagine in that order, the PPR motif selectively binds to A; (2-45) A1, A4, and L ii When the combination of three amino acids is isoleucine, threonine, and asparagine in that order, the PPR motif selectively binds to A; (2-46) A1, A4, and L ii When the combination of three amino acids is valine, threonine, and asparagine in that order, the PPR motif selectively binds to A; (2-47) A1, A4, and L ii When the combination of three amino acids is any amino acid, valine, and any amino acid in that order, the PPR motif binds to A, C, and T but not to G; (2-48) A1, A4, and L iiWhen the combination of three amino acids is isoleucine, valine, and aspartic acid in that order, the PPR motif selectively binds to C and then binds to A; (2-49) A1, A4, and L ii When the combination of three amino acids is any amino acid, valine, and glycine in that order, the PPR motif selectively binds to C; (2-50) A1, A4, and L ii When the combination of three amino acids is any amino acid, valine, and threonine in that order, the PPR motif selectively binds to T.
[0023] (Novel PPR motif) The present invention provides a novel PPR motif. The novel PPR motifs provided by the present invention that are adenine-binding are as follows: (A-1), (A-2), and (A-3): (A-1) A PPR motif consisting of the sequence of SEQ ID NO: 9, or a PPR motif consisting of an amino acid sequence obtained by performing any substitution selected from the group consisting of substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid in the sequence of SEQ ID NO: 9; (A-2) A PPR motif consisting of a sequence in which 1 to 20 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 9 are substituted, deleted, or added, and which is adenine-binding; (A-3) A PPR motif having at least 42% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are identical and which is adenine-binding.
[0024] (The substitution in (A-1) may be 1, may be 2 or more, or may be all of the above.)
[0025] In (A-2), 1 to 20 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, which are amino acids that can be substituted, etc. in the sequence of SEQ ID NO: 9, preferably, they are amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, and are 1 to 11 amino acids other than the amino acids at positions 5, 8, 13, 21, 22, 23, 25, 29, and 35, more preferably, they are amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, and are amino acids other than the amino acids at positions 5, 8, 13, 21, 22, 23, 25, 29, and 35, and are 1 to 7 amino acids other than the amino acids at positions 20, 24, 31, and 32, even more preferably, any one of the amino acids at positions 10, 15, 16, 17, 18, and 28.
[0026] (A-3) has at least 42% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are identical, preferably, it has at least 71% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, and the amino acids at positions 5, 8, 13, 21, 22, 23, 25, 29, and 35 are identical, more preferably, it has at least 80% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, the amino acids at positions 5, 8, 13, 21, 22, 23, 25, 29, and 35, and the amino acids at positions 20, 24, 31, and 32 are identical, More preferably, it has at least 82% sequence identity with the sequence of SEQ ID NO: 9, provided that the non-identical amino acids are any of the amino acids at positions 10, 15, 16, 17, 18, and 28.
[0027] The novel PPR motif provided by the present invention, which is cytosine-binding, is as follows: (C-1), (C-2), and (C-3): (C-1) A PPR motif consisting of the sequence of SEQ ID NO: 10, or a PPR motif consisting of an amino acid sequence obtained by performing any substitution selected from the group consisting of substitution of the amino acid at position 2 with serine, substitution of the amino acid at position 5 with isoleucine, substitution of the amino acid at position 7 with leucine, substitution of the amino acid at position 8 with lysine, substitution of the amino acid at position 10 with phenylalanine or tyrosine, substitution of the amino acid at position 15 with arginine, substitution of the amino acid at position 22 with valine, substitution of the amino acid at position 24 with arginine, substitution of the amino acid at position 27 with leucine, and substitution of the amino acid at position 29 with arginine in the sequence of SEQ ID NO: 10; (C-2) A PPR motif consisting of a sequence in which 1 to 25 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 10 are substituted, deleted, or added, and which is cytosine-binding; (C-3) A PPR motif having at least 25% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are identical and which is cytosine-binding.
[0028] (C-1) The substitution may be 1, may be 2 or more, or may be all of the above.
[0029] (C-2) In the sequence of SEQ ID NO: 10, the 1 to 25 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34, which are amino acids that can be substituted, deleted, or added, Preferably, it is an amino acid other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34, and is 1 to 14 amino acids other than the amino acids at positions 6, 9, 11, 12, 17, 20, 21, 23, 25, 28, and 35, More preferably, it is an amino acid other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34, and is an amino acid other than the amino acids at positions 6, 9, 11, 12, 17, 20, 21, 23, 25, 28, and 35, and is 1 to 10 amino acids other than the amino acids at positions 13, 16, 31, and 32, Even more preferably, it is any one of the amino acids at positions 2, 5, 7, 8, 10, 15, 22, 24, 27, and 29.
[0030] (C-3) has at least 25% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are identical, Preferably, it has at least 60% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34, and the amino acids at positions 6, 9, 11, 12, 17, 20, 21, 23, 25, 28, and 35 are identical, More preferably, it has at least 71% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34, the amino acids at positions 6, 9, 11, 12, 17, 20, 21, 23, 25, 28, and 35, and the amino acids at positions 13, 16, 31, and 32 are identical, Even more preferably, it has at least 71% sequence identity with the sequence of SEQ ID NO: 10, provided that the non-identical amino acids are any one of the amino acids at positions 2, 5, 7, 8, 10, 15, 22, 24, 27, and 29.
[0031] The novel PPR motif provided by the present invention, which is guanine-binding, is the following (G-1), (G-2), and (G-3): (G-1) A PPR motif consisting of the sequence of SEQ ID NO: 11, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the phenylalanine at position 10, substitution of the aspartic acid at position 15, substitution of the valine at position 27, substitution of the serine at position 28, and substitution of the isoleucine at position 35 is made in the sequence of SEQ ID NO: 11; (G-2) A PPR motif consisting of a sequence in which 1 to 21 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 11 are substituted, deleted, or added, and which is guanine-binding; (G-3) A PPR motif having at least 40% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are identical, and which is guanine-binding.
[0032] (G-1) The substitution in (G-1) may be 1, may be 2 or more, or may be all of the above.
[0033] (G-2) In (G-2), the 1 to 21 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34, which are amino acids that can be substituted, etc. in the sequence of SEQ ID NO: 11, are preferably 1 to 12 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 and other than the amino acids at positions 5, 11, 12, 17, 20, 21, 22, 23, and 25, more preferably 1 to 5 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34, other than the amino acids at positions 5, 11, 12, 17, 20, 21, 22, 23, and 25, and other than the amino acids at positions 8, 13, 16, 24, 29, 31, and 32, More preferably, it is any one of the amino acids at positions 10, 15, 27, 28, and 35.
[0034] (G-3) has at least 40% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are the same. Preferably, it has at least 65% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34, and the amino acids at positions 5, 11, 12, 17, 20, 21, 22, 23, and 25 are the same. More preferably, it has at least 85% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34, the amino acids at positions 5, 11, 12, 17, 20, 21, 22, 23, and 25, and the amino acids at positions 8, 13, 16, 24, 29, 31, and 32 are the same. Even more preferably, it has at least 85% sequence identity with the sequence of SEQ ID NO: 11, provided that the non-identical amino acids are any one of the amino acids at positions 10, 15, 27, 28, and 35.
[0035] The novel PPR motif provided by the present invention, which is uracil-binding, is the following (U-1), (U-2), and (U-3): (U-1) A PPR motif consisting of the sequence of SEQ ID NO: 12, or an amino acid sequence obtained by performing any substitution selected from the group consisting of substitution of the phenylalanine at position 10, substitution of the serine at position 13, substitution of the lysine at position 15, substitution of the glutamate at position 17, substitution of the leucine at position 20, substitution of the lysine at position 21, substitution of the phenylalanine at position 23, substitution of the aspartic acid at position 24, substitution of the lysine at position 27, substitution of the lysine at position 28, substitution of the arginine at position 29, and substitution of the leucine at position 31 in the sequence of SEQ ID NO: 12; (U-2) A PPR motif consisting of a sequence in which 1 to 22 amino acids other than those at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 12 are substituted, deleted, or added, and which is uracil-binding; (U-3) A PPR motif having at least 37% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are identical and which is uracil-binding.
[0036] (U-1) The substitution may be 1, may be 2 or more, or may be all of the above.
[0037] (U-2) In (U-2), the 1 to 22 amino acids other than those at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34, which are amino acids that can be substituted, etc. in the sequence of SEQ ID NO: 12, are preferably 1 to 14 amino acids other than those at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 and other than the amino acids at positions 5, 7, 9, 16, 18, 22, 25, and 35; More preferably, it is one to twelve amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34, and other than the amino acids at positions 5, 7, 9, 16, 18, 22, 25, and 35, and other than the amino acids at positions 8 and 32, Even more preferably, it is any one of the amino acids at positions 10, 13, 15, 17, 20, 21, 23, 24, 27, 28, 29, and 31.
[0038] (U-3) has at least 37% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are the same, Preferably, it has at least 60% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34, and the amino acids at positions 5, 7, 9, 16, 18, 22, 25, and 35 are the same, More preferably, it has at least 65% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34, the amino acids at positions 5, 7, 9, 16, 18, 22, 25, and 35, and the amino acids at positions 8 and 32 are the same, Even more preferably, it has at least 65% sequence identity with the sequence of SEQ ID NO: 12, provided that the non-identical amino acids are any one of the amino acids at positions 10, 13, 15, 17, 20, 21, 23, 24, 27, 28, 29, and 31.
[0039] The PPR motif v2 created by the present inventors A (SEQ ID NO: 9), v2 C (SEQ ID NO: 10), v2 G (SEQ ID NO: 11), v2 U (SEQ ID NO:12) is disclosed for the first time in this application and does not exist in nature. For each of their homologs (among the embodiments shown as (A-1), (A-2), (A-3), (C-1), (C-2), (C-3), (G-1), (G-2), (G-3), (U-1), (U-2), (U-3) and the preferred embodiments thereof, the embodiments consisting of sequences other than SEQ ID NOs: 9-12), (regardless of whether each homolog is disclosed for the first time in this application or whether it exists in nature), combinations of at least two or more, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 of any of the homologs are considered not to exist in nature. Regarding the present invention, when referring to "any", the number selected is arbitrary.
[0040] (Explanation of the sequence of the novel PPR motif) In FIGS. 1 to 4, among the PPR motif sequences of Arabidopsis thaliana, combinations of amino acids at positions 1, 4, and ii are collected for the PPR motif that recognizes adenine, which is VTN, for the PPR motif that recognizes cytosine, which is VSN, for the PPR motif that recognizes guanine, which is VTD, and for the PPR motif that recognizes uracil, which is VND. The types and numbers of amino acids appearing at each position are summarized. The sequence v2 of the new PPR motif A (SEQ ID NO:9), v2 C (SEQ ID NO:10), v2 G (SEQ ID NO:11), v2 In U (SEQ ID NO:12), the amino acids at each position have a high frequency of occurrence. In FIG. 6A, together with these novel sequences, in the sequence of the dPPR motif, v1 with the combination of amino acids at positions 1, 4, and ii made the same as v2 A (SEQ ID NO:13), v1 C (SEQ ID NO:14), v1 G (SEQ ID NO:15), v1 U (SEQ ID NO:16) is also shown.
[0041] In addition, FIG. 6A shows the amino acid sequence of the v3.1 motif. V3.1 is obtained by introducing a D15K mutation into the adenine recognition motif in v2 (SEQ ID NO: 401), and other points are the same as those in v2. By using v3.1 in the PPR protein, it may be possible to obtain a protein with improved binding ability compared to v2.
[0042] In addition, Tables 1 to 4 below summarize how far the amino acid occurrence frequencies in each of the sequences of SEQ ID NOs: 9 - 12 deviate from those in the case of random (for example, if 100 PPR motifs are collected, and the occurrence frequency of an amino acid at a certain position is randomly occurring, then each of the 20 types of amino acids will appear 5 times). At a certain position, if the amino acid occurrence frequency is far from that in the case of random and the occurrence frequency is high, it is considered to be evolutionarily convergent, and the amino acid at that position is considered to have a high association with the function. For amino acids with a high association with the function, even if they are replaced with other amino acids with a high occurrence frequency and far from those in the case of random, the function as a PPR motif can be maintained.
[0043] [Table 1]
[0044] [Table 2]
[0045] [Table 3]
[0046] [Table 4]
[0047] (Novel PPR protein) The present invention provides a novel PPR protein containing a novel PPR motif. The novel PPR protein provided by the present invention is as follows. A protein containing n PPR motifs that can bind to a target RNA consisting of n base sequences, wherein the PPR motif for adenine in the base sequence is the PPR motif of (A-1), (A-2), or (A-3) described above; the PPR motif for cytosine in the base sequence is the PPR motif of (C-1), (C-2), or (c-3) described above; the PPR motif for guanine in the base sequence is the PPR motif of (G-1), (G-2), or (G-3) described above; and the PPR motif for uracil in the base sequence is the PPR motif of (U-1), (U-2), or (U-3) described above.
[0048] Preferred examples of the PPR motif contained in the PPR protein are as described above for (A-1), (A-2), (A-3), (C-1), (C-2), (C-3), (G-1), (G-2), (G-3), (U-1), (U-2), or (U-3) regarding the PPR motif, and the description applies as it is.
[0049] In the PPR protein of the present invention, n (representing an integer of 1 or more) is not particularly limited, but can be 10 or more, preferably 12 or more, more preferably 15 or more, and even more preferably 18 or more. By increasing the number of motifs, a PPR protein having high binding strength to many targets can be produced.
[0050] Conventionally, as shown in the following table, while the production of artificial PPR proteins consisting of 7 to 14 motifs has been reported, the construction of genes that contain many PPR motifs and inevitably have many repeats in the nucleotide sequence has been considered difficult. Also, generally, when producing a gene containing a repeat sequence, there may be cases where production is difficult, such as the repeat part being rearranged during the cloning process (Trinh, T. et al. An Escherichia coli strain for the stable propagation of retroviral clones and direct repeat sequences. Focus, 16, 78 - 80(1994)). In the table, the Kd values represent the lowest values indicated in each literature.
[0051]
Table 5
[0052] When constructing a gene for a PPR protein having 15 or more PPR motifs, by utilizing the degeneracy of codons and making the nucleotide sequences encoding the amino acids (excluding 1, 4, ii involved in binding in each motif, and for the case of using the GoldenGate method described later, the 29 amino acids at positions 5 to 33 excluding the vicinity of both ends to be made common) different between each motif as appropriate, a gene with reduced repeats in the nucleotide sequence can be constructed. The degree of difference can be appropriately designed by those skilled in the art. For example, 4.5% or more (4 or more positions in 87 bases), 15% or more, or 30% or more (26 or more positions in 87 bases) of the bases can be made different.
[0053] For example, regarding the nucleotide sequences encoding the existing v1 to v4 motifs (SEQ ID NO:13 - 16), examples of the nucleotide sequences encoding motifs utilizing the degeneracy of codons can include sequences as shown in the following table.
[0054]
Table 6
[0055] Note that "preparation" can be paraphrased as "production" or "manufacture". Also, regarding genes and the like, when preparing by combining parts, it is sometimes referred to as "construction", but "construction" can also be paraphrased as "production" or "manufacture".
[0056] (PPR motif, nucleic acid encoding PPR protein) The present invention provides a novel PPR motif and a nucleic acid encoding a novel PPR protein containing the same. The nucleotide sequence encoding the novel PPR motif has several variations due to codon degeneracy.
[0057] Amino acid sequence v2 of the novel PPR motif of the present invention A (SEQ ID NO:9), v2 C (SEQ ID NO:10), v2 G (SEQ ID NO:11), v2 Preferred examples of the nucleotide sequences encoding U (SEQ ID NO:12) are shown in the following table.
[0058]
Table 7-1
[0059] In the dPPR motif, the amino acid sequence of the PPR motif with the amino acid combinations at positions 1, 4, and ii being the same as v2, v1 A (SEQ ID NO:13), v1 C (SEQ ID NO:14), v1 G (SEQ ID NO:15), v1 The nucleotide sequences encoding U (SEQ ID NO:16) are shown in the following table.
[0060]
Table 7-2
[0061] The nucleotide sequence encoding the PPR protein can be constituted by any combination of the above sequences. Amino acid sequence v2 A (SEQ ID NO:9), v2 C (SEQ ID NO:10), v2 G (SEQ ID NO:11), v2 The nucleotide sequence encoding U (SEQ ID NO:12) and v1 A (SEQ ID NO:13), v1 C (SEQ ID NO:14), v1 G (SEQ ID NO:15), v1 The nucleotide sequence encoding U (SEQ ID NO:16) may be appropriately combined to constitute the amino acids encoding the protein.
[0062] Amino acid sequence v3.1 of the novel PPR motif of the present invention A (SEQ ID NO:401), 1st A (SEQ ID NO:402), 1st C (SEQ ID NO:403), 1st G (SEQ ID NO:404), 1st Preferred examples of the nucleotide sequence encoding U (SEQ ID NO:405) are shown in the following table.
[0063]
Table 8
[0064] The nucleotide sequence encoding the PPR protein can be constituted by any combination of the above sequences. As the nucleotide sequence encoding the first PPR motif from the N-terminus, the above v3.2 is used. Any one selected from X is used, and as the nucleotide sequence encoding the subsequent PPR motif, the nucleotide sequence encoding the PPR motif for adenine is the above v3.1 Select A, and appropriately combine those selected from the above v2 series as the base sequences encoding PPR motifs for cytosine, guanine, and uracil.
[0065] (Improvement in Aggregation) The present inventors found that from the amino acid information of existing naturally occurring PPR motifs, the amino acid at the 6th position of the PPR motif is often hydrophobic (especially leucine), and the amino acid at the 9th position is often a non-hydrophilic amino acid (especially glycine). From the structures of PPR proteins for which crystal structures have already been obtained (Non-Patent Document 6: Coquille et al., 2014 Nat. Commun.; PDB ID: 4PJQ, 4WN4, 4WSL, 4PJR; Non-Patent Document 7: Shen et al., 2015 Nat. Commun., PDB ID: 5I9D, 5I9F, 5I9G, 5I9H), since those at the 6th and 9th positions of the first motif (N-terminal side) are exposed to the outside, it was imagined that aggregation was shown due to the exposed hydrophobic amino acids (Figure 6A). On the other hand, in the second motif and later, since the amino acids at the 6th and 9th positions are buried in the protein and form a hydrophobic core, it was considered that putting hydrophilic residues at the 6th and 9th positions of all motifs might collapse the protein structure. Therefore, it was decided to reduce the aggregation of PPR by making only the amino acids at the 6th position, preferably the 6th and 9th positions, of the first motif hydrophilic amino acids (asparagine, aspartic acid, glutamine, glutamic acid, lysine, arginine, serine, threonine).
[0066] Specifically, do as follows. In a protein capable of binding to a target nucleic acid having a specific base sequence, in the first PPR motif (M1) from the N-terminus: (1) Make the A6 amino acid a hydrophilic amino acid, preferably make the A6 amino acid asparagine or aspartic acid. (2) Further, make the A9 amino acid a hydrophilic amino acid or glycine, preferably glutamine, glutamic acid, lysine, or glycine. (3) Alternatively, the A6 amino acid and the A9 amino acid are any of the following combinations. · A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamic acid · A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamine · A combination in which the A6 amino acid is asparagine and the A9 amino acid is lysine · A combination in which the A6 amino acid is aspartic acid and the A9 amino acid is glycine
[0067] Among such PPR motifs, the particularly preferred ones are as follows. (1st A-1) In the sequence of SEQ ID NO: 402, a PPR motif in which the amino acids at positions 6 and 9 are replaced to satisfy any one of the combinations defined below; (1st A-2)(1st A-1) consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 are replaced, deleted, or added, and having adenine-binding property; (1st A-3)(1st A-1) having at least 80% sequence identity with the sequence of A-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and having adenine-binding property; (1st C-1) A PPR motif consisting of the sequence of SEQ ID NO: 403; (1st C-2)(1st C-1) consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 are replaced, deleted, or added, and having cytosine-binding property; (1st C-3)(1st A PPR motif having at least 80% sequence identity with the sequence of (C-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are identical and it is cytosine-binding; (1st G-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 404 are substituted so as to satisfy any one of the combinations defined below; (1st G-2)(1st G-1) A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of G-1 are substituted, deleted, or added, and it is guanine-binding; (1st G-3)(1st G-1) A PPR motif having at least 80% sequence identity with the sequence of G-1, provided that the amino acids at positions 1, 4, 6, 9, and 34 are identical and it is guanine-binding; (1st U-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 405 are substituted so as to satisfy any one of the combinations defined below; (1st U-2)(1st U-1) A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of U-1 are substituted, deleted, or added, and it is uracil-binding; (1st U-3)(1st U-1) A PPR motif having at least 80% sequence identity with the sequence of U-1, provided that the amino acids at positions 1, 4, 6, 9, and 34 are identical and it is uracil-binding. · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamic acid · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamic acid · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is lysine ·The combination where the amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine
[0068] In FIG. 6A, the amino acid sequence of the v3.2 motif is shown together with the v3.1 motif. V3.2, the first motif is 1st A (SEQ ID NO: 402), 1st C (SEQ ID NO: 403), 1st G (SEQ ID NO: 404), 1st U (SEQ ID NO: 405) is selected, and from the second motif onwards is v2 C, v2 G, v2 U, v3.1 A is selected. By using any of v3.2 as the first PPR motif from the N-terminus in the PPR protein, aggregation in cells can be improved.
[0069] (Others) When the term "identity" is used with respect to a base sequence (sometimes also referred to as a nucleotide sequence) or an amino acid sequence in the present invention, unless otherwise specified, it means the percentage of the number of matching bases or amino acids shared between two sequences when the two sequences are aligned in an optimal manner. That is, identity = (number of matching positions / total number of positions) × 100 and can be calculated using commercially available algorithms. Such algorithms are incorporated into the NBLAST and XBLAST programs described in Altschul et al., J. Mol. Biol. 215 (1990) 403 - 410. More specifically, searches and analyses regarding the identity of base sequences or amino acid sequences can be performed using algorithms or programs well-known to those skilled in the art (for example, BLASTN, BLASTP, BLASTX, ClustalW). Parameters when using programs can be appropriately set by those skilled in the art, or the default parameters of each program may be used. Specific methods of these analysis methods are also well-known to those skilled in the art.
[0070] In this specification, regarding a nucleotide sequence or an amino acid sequence, when identity is expressed as a percentage, unless otherwise specified, in any case, a higher identity percentage value is preferred. Specifically, it is preferably 70% or more, more preferably 80% or more, even more preferably 85% or more, even more preferably 90% or more, even more preferably 95% or more, and even more preferably 97.5% or more.
[0071] Also, regarding the PPR motif or protein in the present invention, the number of amino acids to be substituted, deleted, or added when referred to as "substituted, deleted, or added sequence" is not particularly limited in any motif or protein, unless otherwise specified, as long as the motif or protein composed of the amino acid sequence has the desired function. However, it is about 1 to 9 or about 1 to 4, or if it is a substitution with amino acids having similar properties, there can be even more substitutions and the like. Means for preparing polynucleotides or proteins related to such amino acid sequences are well known to those skilled in the art.
[0072] Amino acids with similar properties refer to amino acids having similar physical properties such as hydrophobicity, charge, pKa, solubility, etc. For example, it refers to the following. Hydrophobic (non-polar) amino acids: alanine, valine, glycine, isoleucine, leucine, phenylalanine, proline, tryptophan, tyrosine Non-hydrophobic amino acids: arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine, cysteine, histidine, methionine; Hydrophilic amino acids: arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine; Acidic amino acids: aspartic acid, glutamic acid; Basic amino acids: lysine, arginine, histidine; Neutral amino acids: alanine, asparagine, cysteine, glutamine, glycine, isoleucine, leucine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine; Sulfur-containing amino acids: methionine, cysteine; Aromatic ring-containing amino acids: tyrosine, tryptophan, phenylalanine. Those skilled in the art can prepare the PPR motif of the present invention, the protein containing the same, or the nucleic acid encoding them using the prior art.
[0073] [Performance of the novel PPR motif] (Binding ability) The PPR protein prepared using the novel PPR motif (SEQ ID NOs: 9 - 12) of the present invention is not only suitable for preparing PPR proteins for relatively long target RNAs, but may also have higher RNA binding performance than the PPR protein prepared using the existing PPR (SEQ ID NOs: 13 - 16) motif for the same target RNA.
[0074] That is, by using the novel PPR motif of the present invention in the PPR protein, the binding ability to the target RNA can be increased compared to the case of using the existing PPR motif. By increasing the binding ability, the efficiency of RNA manipulation in cells by the PPR protein can be improved. For example, the splicing efficiency in cells can be improved by using a PPR protein with a high binding ability to the target (see Example 5).
[0075] The degree to which the binding ability is increased seems to also depend on the sequence and length of the target. For example, the binding ability can be 1.1 times or more, specifically 1.3 times or more, 2.0 times or more, 3.0 times or more, 3.6 times or more.
[0076] The binding affinity for the target sequence can be evaluated by methods using EMSA (Electrophoretic Mobility Shift Assay) or Biacore. EMSA is a method that utilizes the property that the mobility of nucleic acid molecules changes when a sample in which a protein and a nucleic acid are bound is electrophoresed, as compared to the case where they are not bound. Molecular interaction analysis instruments typified by Biacore enable detailed protein-nucleic acid binding analysis because they can perform kinetic analysis of the reaction.
[0077] The binding affinity for the target sequence can also be evaluated by RPB-ELISA described later. In RPB-ELISA, the value obtained by subtracting the background signal (the luminescence signal value when the target PPR protein is added without adding the target RNA) from the luminescence amount of the sample to which the target PPR protein and its target RNA are added can be defined as the binding affinity between the target PPR and its target RNA.
[0078] (Specificity) The PPR protein produced using the novel PPR motif of the present invention may have a higher ability in terms of specificity for the target sequence than the PPR protein produced using the existing PPR motif for the same target RNA.
[0079] That is, by using the novel PPR motif of the present invention in a PPR protein, the specificity for the target RNA can be increased compared to the case of using the existing PPR motif. If a PPR protein has high specificity for the target RNA, when it is used to manipulate the target RNA in cells, it is possible to avoid unintended effects resulting from binding to unintended RNAs.
[0080] The affinity for the target sequence can be evaluated by a conventional method by those skilled in the art. In RPB-ELISA, an appropriate non-target RNA is designed for the target PPR protein, and the binding strength (luminescence signal value) in this case is similarly determined, whereby the binding signal value for the target sequence / the binding signal value for the non-target sequence (S / N) can be determined as an index of the specificity (affinity) for the target RNA.
[0081] (Kd value) The PPR protein prepared using the novel PPR motif of the present invention can have a high affinity (equilibrium dissociation constant, Kd value) for the target RNA.
[0082] The Kd value for the target sequence can be calculated by an existing method such as EMSA. When referring to the Kd value in the present invention, unless otherwise specified, it refers to the value measured by EMSA under the conditions described in the Examples section below.
[0083] The Kd value of the PPR protein prepared using the novel PPR motif of the present invention seems to also depend on the target sequence and length, but when the length of the target sequence is 18 bases or more, it can be -7 10 -8 M or less, or can be on the order of 10 -9 M. According to the studies of the present inventors, under the conditions of the Examples when the length of the target sequence is 18 bases, the minimum value (high affinity) of the Kd value is 1.95 x 10 -9 and is lower than any of the Kd values of the designed PPR proteins reported so far (see Table 1). It should be noted that the Kd value has been found to correlate with the signal value obtained in the binding experiment by RPB-ELISA. When the luminescence value in RPB-ELISA (under the conditions described in the Examples section) is 1-2 x 10 7 , the Kd value is 10 -6 -10 -7 M, and when the luminescence value in RPB-ELISA is 2-4 x 10 7 , the Kd value is 10 -7 -10 -8When the luminescence value in M, RPB-ELISA is greater than 4 x 10 7 it can be estimated that the Kd value is ~10 -8 or less.
[0084] (Construction efficiency of PPR protein) By using the novel PPR motif of the present invention, a desired PPR protein can be efficiently constructed. The construction efficiency can be calculated by determining the ratio of the PPR proteins with high Kd values that can be constructed using existing methods. Instead of the Kd value, the luminescence signal value by RPB-ELISA may be determined as described above and calculated in the same manner.
[0085] Specifically, by using the novel PPR motif of the present invention, when the length of the target sequence is 18 bases, the Kd value is 10 -6 M or less (RPB-ELISA value is 1 x 10 7 or more) PPR proteins can be obtained with an efficiency of 50% or more, specifically 60% or more, more specifically 70% or more, and even more specifically 80% or more. Also, according to the present invention, PPR proteins with a Kd value of 10 -7 M or less (RPB-ELISA value is 2 x 10 7 or more) and a target sequence length of 18 bases can be obtained with an efficiency of 50% or more, specifically 55% or more, more specifically 65% or more, and even more specifically 75% or more. Also, according to the present invention, PPR proteins with a Kd value of 10 -8 M or less (RPB-ELISA value is 4 x 10 7 or more) and a target sequence length of 18 bases can be obtained with an efficiency of 20% or more, specifically 25% or more, more specifically 30% or more, and even more specifically 35% or more.
[0086] The construction efficiency can be calculated based on the binding signal value to the target sequence / binding signal value to the non-target sequence (S / N) using the RPB-ELISA method.
[0087] Specifically, by using the novel PPR motif of the present invention, PPR proteins with a target sequence 18 bases in length and an S / N higher than 10 can be obtained with an efficiency of 50% or more, 55% or more when more specifically identified, 65% or more when further specifically identified, and 75% or more when even further specifically identified. Also, according to the present invention, PPR proteins with a target sequence 18 bases in length and an S / N higher than 100 can be obtained with an efficiency of 15% or more, 20% or more when more specifically identified, 25% or more when further specifically identified, and 30% or more when even further specifically identified.
[0088] [Seamless cloning of PPR protein genes using a parts library] The present invention also provides a method for producing a gene encoding a protein containing n PPR motifs that can bind to a target nucleic acid consisting of n base sequences, the method including the following steps: Selecting m PPR parts necessary for producing a target gene from a library of at least 20×m types of PPR parts, wherein at least 20 types of polynucleotides, including 4 types encoding each of the PPR motifs that are adenine, cytosine, guanine, or uracil / thymine binding, and 16 types encoding each of the dimers of the PPR motifs, are inserted into respective intermediate vectors Dest-a... designed to be ligatable in at least m types of sequences; Subjecting the selected m types of PPR parts to a Golden Gate reaction together with vector parts to obtain a vector into which a ligate of m polynucleotides is inserted. n is an integer that is m or more and 2×m or less. n can be, for example, 10 to 20.
[0089] The method of the present invention utilizes the Golden Gate reaction. In the Golden Gate reaction, a plurality of DNA fragments are inserted into a vector using a Type IIS restriction enzyme and T4 DNA Ligase. Since the Type IIS restriction enzyme cleaves outside the recognition sequence, the sticky ends can be freely set. In addition, it is highly efficient because the 4-base overhangs are used for ligation. Furthermore, no recognition sequences remain in the annealed and ligated constructs. Therefore, the polynucleotide encoding the PPR motif can be seamlessly ligated (Figure 5). A particularly preferred example of the Type IIS restriction enzyme is BsaI.
[0090] The method of the present invention can efficiently generate a gene even when there are many repeat sequences by using a parts library appropriately designed in consideration of the characteristics of the PPR protein and the Golden Gate reaction. Therefore, this method is beneficial when generating a gene for a protein containing 15 or more PPR motifs that can bind to a target nucleic acid with a length of 15 bases or more with an increasing number of repeat sequences. If m is set to 10, a library of 200 types of PPR parts is used, and 10 PPR parts required for generating the target gene are selected from the library, a gene encoding a protein containing 10 to 20 PPR motifs can be freely generated. In the following, although the case of generating a gene encoding an RNA-binding PPR protein with a target sequence of 10 to 20 bases in length may be described as an example, this method can also be applied to prepare PPR proteins for other lengths of target sequences and can also be applied to prepare DNA-binding PPR proteins.
[0091] In the method of the present invention, a parts library containing one or two arrays encoding the PPR motif is prepared (STEP1 and STEP2 in FIG. 5) and used. The parts library can be prepared, for example, by inserting the PPR motif array into 10 types of intermediate vectors Dest-a, b, c, d, e, f, g, h, i, j. The intermediate vectors are designed such that Dest-a to Dest-j are seamlessly concatenated in order by the Golden Gate reaction. The PPR motif arrays to be inserted can be at least 20 types, including 4 types (A, C, G, U) encoding each of them and 16 types (AA, AC, AG, AU, CA, CC, CG, CU, GA, GC, GG, GU, UA, UC, UG, UU) encoding each of the two concatenates of the PPR motif. In this case, the parts library contains at least 200 types of parts.
[0092] Next, the necessary parts are selected according to the target base sequence. Specifically, for example, one part is selected from each of the parts libraries of Dest-a, b, c, d, e, f, g, h, i, j, and the Golden Gate reaction is performed together with the vector parts (STEP3 in FIG. 5). If 10 arrays with 1 motif in each of all the intermediate vectors are selected, 10 arrays will be concatenated, and if those with 2 motifs are used, 20 arrays will be concatenated. When concatenating 11 to 19, one with 1 motif can be selected from any Dest-x library.
[0093] The vector parts used in STEP3 can be selected from 3 types of CAP-x vectors (see Non-Patent Document 1 above, where consideration needs to be given to the second amino acid of the PPR motif arranged on the most C-terminal side, and the second amino acid of the guanine-binding PPR motif and the uracil-binding PPR motif is the same). When the base sequence recognized by the motif located on the most C-terminal side is adenine, CAP-A can be used, when it is cytosine, CAP-C can be used, and when it is guanine or uracil, CAP-GU can be used respectively.
[0094] The obtained plasmid can be transformed into Escherichia coli and amplified and extracted.
[0095] [Method for detecting or analyzing PPR protein] The present invention provides a method for detecting or quantifying a protein containing n PPR motifs that can bind to a target nucleic acid consisting of n base sequences, including the following steps: A step of providing a solution containing a candidate protein to the immobilized target nucleic acid and detecting or quantifying the protein bound to the target nucleic acid. This detection and analysis method of the present invention is useful as a method for evaluating the binding performance of high-throughput PPR proteins.
[0096] Since the detection and analysis method of the present invention applies ELISA (Enzyme-Linked Immuno Sorbent Assay) (FIG. 7A), it may be referred to as the RPB-ELISA (RNA-protein binding ELISA) method. Although the method of the present invention is described in this specification as a method for evaluating RNA-binding PPR proteins, similarly, it can also be applied to evaluate the binding performance of DNA-binding PPR proteins to target DNA.
[0097] The step of providing a solution containing a candidate protein to the immobilized target nucleic acid can be specifically carried out by flowing a solution containing the target binding protein over the target nucleic acid molecule immobilized on the plate. For the immobilization of the target nucleic acid molecule, various existing immobilization methods can be used. For example, it can be achieved by providing a nucleic acid probe containing a biotinylated target nucleic acid molecule to a well plate coated with streptavidin.
[0098] On the other hand, the candidate protein to be measured can be fused with a labeled protein, such as an enzyme like luciferase or a fluorescent protein. Fusion with the labeled protein makes detection and quantification easier.
[0099] The RPB-ELISA method has the advantage of not requiring a special device such as Biacore. Also, with the RPB-ELISA method, the throughput is high, and the binding between protein and nucleic acid can be evaluated in a short period. Furthermore, under the conditions of the examples, the RPB-ELISA method can be sufficiently detected at a protein concentration of 6.25 nM or higher, and can also be detected similarly in an E. coli lysate. Therefore, it has the advantage that it is not necessary to purify the target nucleic acid-binding protein.
[0100] [Use of PPR protein] (Complex, fusion protein) The PPR motif or PPR protein provided by the present invention can be linked with a functional region to form a complex. Also, it can be linked with a proteinaceous functional region to form a fusion protein. The functional region refers to a part that has a specific biological function, such as an enzyme function, a catalytic function, an inhibitory function, a promoting function, etc., or a part that has a function as a label, in vivo or intracellularly. Such a region consists of, for example, a protein, a peptide, a nucleic acid, a physiologically active substance, or a drug. Hereinafter, the present invention may be described by taking a fusion protein as an example with respect to the complex, but those skilled in the art can also understand the case of a complex other than the fusion protein according to this description.
[0101] In one of the preferred embodiments, the functional region is ribonuclease (RNase). Examples of RNase are RNase A (for example, bovine pancreatic ribonuclease A: PDB 2AAS), RNase H.
[0102] In one preferred embodiment, the functional region is a fluorescent protein. Examples of fluorescent proteins are mCherry, EGFP, GFP, Sirius, EBFP, ECFP, mTurquoise, TagCFP, AmCyan, mTFP1, MidoriishiCyan, CFP, TurboGFP, AcGFP, TagGFP, Azami-Green, ZsGreen, EmGFP, HyPer, TagYFP, EYFP, Venus, YFP, PhiYFP, PhiYFP-m, TurboYFP, ZsYellow, mBanana, KusabiraOrange, mOrange, TurboRFP, DsRed-Express, DsRed2, TagRFP, DsRed-Monomer, AsRed2, mStrawberry, TurboFP602, mRFP1, JRed, KillerRed, HcRed, KeimaRed, mRasberry, mPlum, PS-CFP, Dendra2, Kaede, EosFP, KikumeGR. From the viewpoint of improving aggregation and / or efficient localization to the nucleus as a fusion protein, a preferred example is mClover3.
[0103] In one preferred embodiment, when the target is mRNA, the functional region is a functional domain that improves the protein expression level from the target mRNA (WO2017 / 209122). Examples of the functional domain that improves the protein expression level from mRNA may be all or a functional part of the functional domain of a protein known to directly or indirectly promote the translation of mRNA. More specifically, it may be a domain that induces ribosomes to mRNA, a domain related to the initiation or promotion of mRNA translation, a domain related to the transport of mRNA out of the nucleus, a domain related to the binding to the endoplasmic reticulum membrane, a domain containing an endoplasmic reticulum retention signal sequence, or a domain containing an endoplasmic reticulum signal sequence. Even more specifically, the above domain that induces ribosomes to mRNA may be a domain containing all or a functional part of a polypeptide selected from the group consisting of DENR (Density-regulated protein), MCT-1 (Malignant T-cell amplified sequence 1), TPT1 (Translationally-controlled tumor protein), and Lerepo4 (Zinc finger CCCH-domain). Also, the above domain related to the initiation or promotion of mRNA translation may be a domain containing all or a functional part of a polypeptide selected from the group consisting of eIF4E and eIF4G. Also, the above domain related to the transport of mRNA out of the nucleus may be a domain containing all or a functional part of SLBP (Stem-loop binding protein). Also, the above domain related to the binding to the endoplasmic reticulum membrane may be a domain containing all or a functional part of a polypeptide selected from the group consisting of SEC61B, TRAP-alpha (Translocon associated protein alpha), SR-alpha, Dia1 (Cytochrome b5 reductase 3), and p180. Also, the above endoplasmic reticulum retention signal sequence may be a signal sequence containing the KDEL (KEEL) sequence.Further, the endoplasmic reticulum signal sequence may be a signal sequence including MGWSCIILFLVATATGAHS.
[0104] In the present invention, the functional region may be fused to the N-terminal side of the PPR protein, may be fused to the C-terminal side, or may be fused to both the N-terminal side and the C-terminal side. Further, the complex or fusion protein may contain a plurality of functional regions (for example, 2 to 5). Furthermore, in the complex or fusion protein of the present invention, the functional region and the PPR protein may be indirectly fused via a linker or the like.
[0105] (Nucleic acid, vector, cell encoding a PPR protein or the like) The present invention also provides a nucleic acid encoding the above-described PPR motif, PPR protein or fusion protein, and a vector containing the nucleic acid (for example, a vector for amplification, an expression vector). The vector for amplification can use Escherichia coli or yeast as a host. In the present specification, the expression vector means, for example, a vector containing DNA having a promoter sequence, DNA encoding a desired protein, and DNA having a terminator sequence from upstream, but it is not necessarily arranged in this order as long as it exhibits a desired function. In the present invention, various vectors that can be normally used by those skilled in the art can be recombinantly used.
[0106] The PPR protein or fusion protein of the present invention can function in cells of eukaryotes (e.g., animals, plants, microorganisms (such as yeast), protists). The fusion protein of the present invention can particularly function in animal cells (in vitro or in vivo). Examples of animal cells into which the PPR protein, fusion protein, or vector expressing the same of the present invention can be introduced include cells derived from humans, monkeys, pigs, cows, horses, dogs, cats, mice, and rats. Examples of cultured cells into which the PPR protein, fusion protein, or vector expressing the same of the present invention can be introduced include, but are not limited to, Chinese hamster ovary (CHO) cells, COS-1 cells, COS-7 cells, VERO (ATCC CCL-81) cells, BHK cells, canine kidney-derived MDCK cells, hamster AV-12-664 cells, HeLa cells, WI38 cells, 293 cells, 293T cells, and PER.C6 cells.
[0107] (Use) The PPR protein or fusion protein of the present invention may be capable of delivering and functioning a functional region specifically to a nucleic acid sequence in vivo or in cells. A complex linked with a labeling moiety such as GFP can be used to visualize a desired RNA in vivo.
[0108] In addition, the PPR protein or fusion protein of the present invention can modify and disrupt a nucleic acid sequence specifically in cells or in vivo, and may be able to confer a new function. In particular, RNA-binding PPR proteins are involved in all steps of RNA processing, cleavage, RNA editing, translation, splicing, and RNA stabilization found in organelles. Therefore, the methods related to the modification of the PPR protein provided by the present invention, and the PPR motif and PPR protein provided by the present invention can be expected to be used as follows in various fields.
[0109] (1) Medicine ·Produce a PPR protein that recognizes and binds to a specific RNA associated with a specific disease. Also, analyze the target sequence and the accompanying protein for a specific RNA. The results of these analyses can be used to search for compounds for the treatment of diseases.
[0110] For example, in animals, it is known that an abnormality in a PPR protein identified as LRPPRC causes Leigh syndrome French Canadian (LSFC; Leigh syndrome, subacute necrotizing encephalomyelopathy). The present invention can contribute to the treatment (prevention, treatment, suppression of progression) of LSFC. Many existing PPR proteins function to specify the editing site of RNA manipulation (conversion of genetic information on RNA; in many cases, C→U). This type of PPR protein has an additional motif that is suggested to interact with an RNA editing enzyme at the C-terminal side. It is expected that a PPR protein having such a structure can introduce a base polymorphism or treat a disease or condition caused by a base polymorphism.
[0111] ·Produce cells that control the suppression / expression of RNA. Such cells include stem cells (e.g., iPS cells) that monitor the differentiated / undifferentiated state, model cells for the evaluation of cosmetics, and cells that can turn on / off the expression of functional RNA for the purpose of elucidating the mechanism of drug discovery and pharmacological tests.
[0112] ·Produce a PPR protein that specifically binds to a specific RNA associated with a specific disease. Introduce such a PPR protein into cells using a plasmid, viral vector, mRNA, or purified protein. By binding of the PPR protein to its target RNA in the cell, the RNA function that is the cause of the disease can be changed (improved). Means of changing the function include, for example, changes in the RNA structure due to binding, knockdown by degradation, changes in the splicing reaction by splicing, and base substitution.
[0113] (2) Agriculture, Forestry, and Fisheries ·Improve the yield and quality of agricultural crops, forest products, fishery products, etc. · Breed organisms with improved disease resistance, improved environmental tolerance, and improved or new functionality.
[0114] For example, regarding first-generation hybrid (F1) crops, artificially creating F1 crops by using the PPR protein to stabilize mitochondrial RNA and control translation may improve yield and quality. RNA manipulation and genome editing using the PPR protein enable plant variety improvement and breeding (genetically improving organisms) more accurately and quickly than the prior art. Also, RNA manipulation and genome editing using the PPR protein do not transform traits with foreign genes like genetic recombination, but rather handle the RNA and genome originally possessed by animals and plants, and thus can be said to be similar to traditional breeding methods such as mutant selection and backcrossing. Therefore, it can surely and quickly respond to global food and environmental problems.
[0115] (3) Chemistry · In the production of useful substances using microorganisms, cultured cells, plants, and animals (e.g., insects), control the protein expression level by manipulating DNA and RNA. This can improve the productivity of useful substances. Examples of useful substances include proteinaceous substances such as antibodies, vaccines, and enzymes, as well as relatively low-molecular-weight compounds such as pharmaceutical intermediates, fragrances, and pigments.
[0116] · Improve the production efficiency of biofuels by modifying the metabolic pathways of algae and microorganisms.
Example
[0117] [Example 1: Establishment of a method for producing the PPR gene] (Motif design) First, the PPR motif was designed. The PPR motif sequences used in previously reported artificial PPR proteins are consensus sequences of PPR motif sequences existing in nature extracted by various methods. Among them, the PPR protein prepared using the motif sequence of dPPR (Non-Patent Documents 2, 3, and 6 cited above) has a low Kd value (high affinity). This PPR motif sequence is hereinafter referred to as the v1 PPR motif.
[0118] As another PPR motif sequence, a consensus sequence was generated using only the PPR motifs containing the representative first, fourth, and ii-th amino acid combinations that recognize each base. Specifically, the representative amino acid combinations that recognize each base are: for the combination that recognizes adenine, the first is valine, the fourth is threonine, and the ii-th is asparagine; for the combination that recognizes cytosine, the first is valine, the fourth is asparagine, and the ii-th is serine; for the combination that recognizes guanine, the first is valine, the fourth is threonine, and the ii-th is aspartic acid; for the combination that recognizes uracil, the first is valine, the fourth is asparagine, and the ii-th is aspartic acid. Therefore, the consensus amino acid sequence was extracted from the PPR motif sequences containing the first, fourth, and ii-th combinations, and this sequence was used as the PPR motif sequence that specifically recognizes adenine, cytosine, guanine, and uracil, respectively (Figs. 1 to 4, SEQ ID NOs: 9-12). This is hereinafter referred to as the v2 PPR motif. In the v1 PPR motif, the same first, fourth, and ii-th amino acid combinations were also used (SEQ ID NOs: 13-16).
[0119] (Seamless cloning using 1-motif and 2-motif libraries) A cloning method was constructed to seamlessly link these designed PPR motif sequences (Figure 5). Cloning is carried out in three steps. In STEP1, the design and production of each motif sequence; in STEP2, the production of a plasmid library in which one or two motifs are cloned; and in STEP3, the target PPR gene is completed by linking the required number of motifs.
[0120] First, a plasmid in which one PPR motif sequence (from the 4th to the ii-th) was cloned was prepared (STEP1). The STEP1 plasmid contains PPR motif sequences that recognize A, C, G, and U respectively. In the subsequent STEP2, a DNA fragment containing the PPR motif sequence in the STEP1 plasmid is cloned into an intermediate vector (Dest-x, the sequence is as follows). The STEP1 plasmid into which only one motif can be inserted was named P1a-vx-X, and the plasmid into which two motifs can be inserted was named P2a-vx-X on the N side and P2b-vx-X on the C side (vx is v1 or v2, and X is A, C, G, U). For cloning into Dest-x, the BsaI restriction enzyme site (the BsaI restriction enzyme recognizes and cleaves the GGTCTCnXXXX sequence (SEQ ID NO:17), and the XXXX part becomes a 4-base overhang (hereinafter referred to as the tag sequence).) and the base sequences for seamless ligation were designed as follows.
[0121] On the 5' side and 3' side of each motif sequence In the case of P1a, ggtctca atac (SEQ ID NO:18), gtgg tgagacc(SEQ ID NO:19), In the case of P2a, ggtctca atac (SEQ ID NO:18 above), gtggtca cata tgagacc(SEQ ID NO:20), In the case of P2b, ggtctca cata c(SEQ ID NO:21), gtgg tgagacc(SEQ ID NO:19 above) The base sequences with the respective arrays added were prepared by gene synthesis technology and cloned into pUC57-amp.
[0122] There are 10 types of Dest-x, namely Dest-a, b, c, d, e, f, g, h, i, j, and the base sequences were designed so as to be seamlessly linked in order from Dest-a to Dest-j. Dest-a is gaagacataaactccgtggtcacATACagagaccaaggtctcaGTGGtcacatacatgtcttc(SEQ ID NO:1), Dest-b is gaagacatATACagagaccaaggtctcaGTGGtgacataatgtcttc(SEQ ID NO:22), Dest-c is gaagacatcATACagagaccaaggtctcaGTGGttacatatgtcttc(SEQ ID NO:23), Dest-d is gaagacatacATACagagaccaaggtctcaGTGGttacaatgtcttc(SEQ ID NO:24), Dest-e is gaagacattacATACagagaccaaggtctcaGTGGtgacatgtcttc(SEQ ID NO:25), Dest-f is gaagacattgacATACagagaccaaggtctcaGTGGttaatgtcttc(SEQ ID NO:26), Dest-g is gaagacatgttacATACagagaccaaggtctcaGTGGtcatgtcttc(SEQ ID NO:27), Dest-h is gaagacatggtcacATACagagaccaaggtctcaGTGGtatgtcttc(SEQ ID NO:28), Dest-i is gaagacattggttacATACagagaccaaggtctcaGTGGatgtcttc(SEQ ID NO:29), Dest-j is gaagacatgtggtgacATACagagaccaaggtctcaGTGGtcttc (SEQ ID NO:30) It was prepared by gene synthesis technology and cloned into pUC57-kan.
[0123] For all Dest-x, those with PPR motifs corresponding to A, C, G, and U inserted, and those with two PPR motifs inserted that recognize each base combination of AA, AC, AG, AU, CA, CC, CG, CU, GA, GC, GG, GU, UA, UC, UG, and UU were prepared, and the STEP1 plasmid libraries of V1 and V2 were prepared (200 types each). To achieve the above combinations, only 40 ng of the P1a plasmid, or 40 ng of the P2a plasmid, 40 ng of the P2b plasmid, 0.2 μL of 10x ligase buffer (NEB, B0202S), 0.1 μL of BsaI (NEB, R0535S), and 0.1 μL of Quick ligase (NEB, M2200S) were added, and the volume was adjusted to 1.9 μL with sterilized water. A reaction of repeating 37°C for 5 minutes and 16°C for 5 minutes alternately 5 times was carried out using a thermal cycler (Biorad, 1861096J1). Furthermore, 0.1 μL of 10x Cut smart buffer (NEB, B7204) and 0.1 μL of BsaI (NEB, R0535S) were added, and the reaction was carried out at 37°C for 60 minutes and 80°C for 10 minutes. 2.5 μL of the reaction solution was transformed into XL1-blue and selected on LB medium containing 30 μg / ml kanamycin. It was confirmed by sequencing that the target sequence was inserted.
[0124] In STEP3, Dest-a to Dest-j are selected along the target base sequence and cloned into the CAP-x vector (Non-Patent Document 1 cited above). If all intermediate vectors contain 1 motif, 10 motifs will be ligated; if they contain 2 motifs, 20 motifs will be ligated. When ligating 11 to 19 motifs, it can be prepared by using the one with 1 motif in Dest-x at a preferred position. For example, when preparing an 18-motif PPR sequence, plasmids with 1 motif in Dest-a and Dest-b and 2 motifs in the others are used.
[0125] The intermediate vectors used in the cloning of STEP3 need to be selected from three types of vectors. When the base sequence recognized by the motif located on the most C-terminal side is adenine, CAP-A is used; when it is cytosine, CAP-C is used; when it is guanine or uracil, CAP-GU is used respectively. By being cloned into the intermediate vector for STEP3, it is designed such that the amino acid sequence of MGNSV (SEQ ID NO:31) is added to the N-side of the PPR repeat and the amino acid sequence of ELTYNTLISGLGKAGRARDPPV (SEQ ID NO:32) is added to the C-side.
[0126] 20 ng each of 10 types of intermediate plasmids, 1 μL of 10x ligase buffer (NEB, B0202S), 0.5 μL of BpiL (Thermo, ER1012), and 0.5 μL of Quick ligase (NEB, M2200S) were added, and finally adjusted to 10 μl with sterilized water. The reaction was carried out for 15 cycles at 37°C for 5 minutes and 16°C for 7 minutes. Furthermore, 0.4 μL of BpiL was added and the reaction was carried out at 37°C for 30 minutes and at 75°C for 6 minutes. Subsequently, 0.3 μL of 1 mM ATP and 0.15 μL of Plasmid safe nuclease (Epicentre, E3110K) were added and the reaction was carried out at 37°C for 15 minutes. 3.5 μl of the reaction solution was transformed into Escherichia coli (Competent cell of XL-1 Blue strain, Nippon Gene) and cultured in LB medium containing 100 μg / mL spectinomycin at 37°C for 16 hours for selection. A part of the grown colonies was pCR8 Fw: 5'-TTGATGCCTGGCAGTTCCCT -3' (SEQ ID NO:33) and pCR8 Rv: 5'-CGAACCGAACAGGCTTATGT -3' (SEQ ID NO:34) primers were used to amplify the inserted gene region. 5 μL of 2 x Go-taq (Promega, M7123), 1.5 μL of 10 μM pCR8 Fw, 1.5 μL of 10 μM pCR8 Rv, 2 μL of sterilized water were added to a 0.2 mL tube. After reacting at 98°C for 2 minutes in a thermal cycler, DNA amplification reaction was performed with 15 cycles of 5 seconds at 98°C, 10 seconds at 55°C, and 2.5 minutes at 72°C. A part of the reaction solution was electrophoresed using MultiNA (SHIMADZU, MCE202) to confirm the size of the inserted DNA fragment. Three types of 18-motif PPR proteins (PPR1, PPR2, PPR3) (SEQ ID NOs: 35-37, 40-42) were each prepared with 3 clones using v1 motif or v2 motif (v1 PPR1, v1 PPR2, v1 PPR3, v2 PPR1, v2 PPR2, v2 PPR3). The results are shown in Fig. 6B. In v1, bands of the correct size were obtained for all clones except the second clone of PPR2. In v2, bands of the correct size were obtained for all clones. Furthermore, their sequences were confirmed by sequencing. These results indicated that the PPR protein gene can be efficiently constructed by cloning with this method.
[0127] [Example 2: Construction of a high-throughput RNA-binding protein binding performance evaluation system] Generally, the evaluation of the binding between a nucleic acid-binding protein and a nucleic acid molecule is performed by methods using EMSA or Biacore. Electrophoretic Mobility Shift Assay (EMSA) is a method that utilizes the property that the mobility of a nucleic acid molecule changes when a sample in which a protein and a nucleic acid are bound is electrophoresed, as compared with the case where the nucleic acid molecule is not bound. This method has the drawbacks of requiring purified protein, being cumbersome in operation, and not being able to analyze many samples at once. Molecular interaction analysis instruments typified by Biacore enable detailed protein-nucleic acid binding analysis because kinetic analysis of the reaction is possible, but this also requires purified protein and a special apparatus. Therefore, a method with high throughput that can evaluate the binding between a protein and a nucleic acid in a short period of time was considered.
[0128] Enzyme-Linked Immuno Sorbent Assay (ELISA) is generally used for analyzing the binding between an antibody (protein) and a protein. In this method, a primary antibody is immobilized on a well plate, a solution containing the protein to be detected is added thereto, and after washing, a secondary antibody capable of color development or luminescence detection is reacted to quantify the remaining amount of the protein to be analyzed. Applying this, a system was devised in which a nucleic acid molecule is immobilized on a plate, a solution containing a target nucleic acid-binding protein is flowed thereon, and the amount of the bound protein is quantified (Figure 7A). The method for immobilizing the nucleic acid molecule to be examined is performed by adding a nucleic acid probe modified with biotin at its end to a well plate coated with streptavidin. The nucleic acid-binding protein to be measured can be made easier to detect by fusing it with luciferase or a fluorescent protein. In addition, purification of the protein to be measured is not essential, and it is also possible to use a crude extract of cells (animal cultured cells, yeast, Escherichia coli, etc.) in which the nucleic acid-binding protein to be measured is expressed, and the time for purification can be shortened (Figure 7B). When measuring the binding between RNA and an RNA-binding protein by this method, it is described below as RPB-ELISA (RNA-protein binding ELISA).
[0129] To establish an experimental system, recombinant MS2 protein and its binding RNA probe were prepared. The gene of a protein with a luciferase protein fused to the N-terminal side of the MS2 protein and a 6x histidine tag fused to the C-terminal side was prepared by gene synthesis and cloned into the pET21b vector (NL MS2 HIS, SEQ ID NO:357). The RNA probe was synthesized by biotinylating the 5'-ends of a target sequence (RNA 4, SEQ ID NO:64) containing the MS2 binding sequence and a non-target sequence (RNA 51, SEQ ID NO:247) not containing it (Greiner). The MS2 protein expression plasmid was transformed into Escherichia coli of the Rosetta(DE3) strain and cultured overnight at 37 °C in 2 mL of LB medium containing 100 μg / mL ampicillin. Then, 2 mL of the culture solution was added to 300 mL of LB medium containing 100 μg / mL ampicillin, and OD 600It was cultured at 37°C until it reached 0.5 to 0.8. After lowering the cultured medium to 15°C, IPTG was added to a final concentration of 0.1 mM and further cultured for 12 hours. The culture solution was centrifuged at 5000 x g, 4°C for 10 minutes to collect the cells, 5 mL of lysis buffer (20 mM Tris-HCl (pH 8.0), 150 mM NaCl, 0.5% NP-40, 1 mM DTT, 1 mM EDTA) was added, and after stirring with a vortex mixer, the cells were disrupted by sonication. It was centrifuged at 15,000 rpm, 4°C for 10 minutes, and the supernatant was collected. Half of the supernatant was stored at -80°C until used as an E. coli lysate, and the other half was subjected to affinity purification using a histidine tag and Ni-NTA. First, 200 μl of Ni-NTA agarose beads (Qiagen, Cat no. 30230) was spin-down to collect the beads. 100 μL of washing buffer was added thereto, and the beads were equilibrated by stirring with a rotator at 4°C for 1 hour. The entire amount of the equilibrated beads was mixed with the protein solution and reacted at 4°C for 1 hour. Then, after centrifuging at 2,000 rpm for 2 minutes to collect the beads, factors that non-specifically bind to the beads were removed with 10 ml of washing buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40, 10 mM imidazole). Elution was performed with 60 μL of elution buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40, 500 mM imidazole). The purification degree was confirmed by SDS-PAGE. It was dialyzed overnight at 4°C with 20 mM Tris-HCl, pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM DTT, 1 mM EDTA).
[0130] The luminescence of luciferase in the Escherichia coli lysate and the purified MS2 protein solution was measured. 40 μL of the luciferase substrate (Promega, E151A) diluted 2500-fold with the luminescence buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT), 40 μL of the Escherichia coli lysate, or 40 μL of the purified MS2 protein solution was added to a 96-well white plate and reacted for 5 minutes, after which the luminescence was measured with a plate reader (PerkinElmer, 5103-35). From the obtained luminescence, it was diluted with the lysis buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT, 0.1% BSA) to become 0.01 x 10 8 、0.02 x 10 8 、0.09 x 10 8 、0.38 x 10 8 、1.50 x 10 8 、6.00 x 10 8 LU / μL.
[0131] Add 2.5 pmol of biotinylated RNA probe to a 96-well streptavidin-coated white plate (Thermo Fisher, 15502), react at room temperature for 30 minutes, and wash with lysis buffer. Wells containing lysis buffer instead of biotinylated RNA were also prepared (-Probe) for background measurement. Then, add blocking buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT, 1% BSA) and block the plate surface at room temperature for 30 minutes. Then, add 100 μL of the E. coli lysate or purified protein solution diluted above and perform a binding reaction at room temperature for 30 minutes. Then, wash 5 times with 200 μL of wash buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT). Add 40 μL of luciferase substrate (Promega, E151A) diluted 2,500-fold with wash buffer to the wells, react for 5 minutes, and then measure the luminescence with a plate reader (PerkinElmer, 5103-35).
[0132] The value obtained by subtracting the background (the luminescence signal value when the PPR protein was added without adding RNA) from the luminescence of the sample to which the solution containing each RNA and MS2 protein was added was defined as the binding affinity between MS2 protein and RNA.
[0133] The results are shown in Fig. 7C. Specific binding between MS2 protein and target RNA (Target seq.) was detected in both the purified protein solution and the E. coli lysate. Since 6.0 x 10 8 LU / μL corresponds to 100 nM purified MS2 protein, it was found that it can be sufficiently detected at a protein concentration of 6.25 nM (0.38 x 10 8 LU / μL) or higher. Furthermore, since it was also detectable with the E. coli lysate, it was found that it was not necessary to purify the protein.
[0134] [Example 3: RNA binding performance comparison experiment of 18-motif PPR proteins prepared using existing PPR motif sequences or novel PPR motif sequences] To evaluate the RNA binding performance of PPR proteins prepared using v1 or v2 PPR motif, recombinant proteins were prepared in E. coli and their binding performance was evaluated using RPB-ELISA. For comparison, five types of target sequences (T 1, T 2, T 3, T 4, T 5, SEQ ID NOs:46-50) were set, PPR proteins binding to each of them were designed, and genes encoding each of them were prepared (v1 PPR1, v1 PPR2, v1 PPR3, v1 PPR4, v1 PPR5, v2 PPR1, v2 PPR2, v2 PPR3, v2 PPR4 v2 PPR5, SEQ ID NOs:35-39, 40-45). A luciferase protein gene was added to the N-terminal side of the prepared PPR gene, and a histag sequence was added to the C-terminal side, followed by cloning into the pET21 vector (NL v1 PPR1, NL v1 PPR2, NL v1 PPR3, NL v1 PPR4, NL v1 PPR5, NL v2 PPR1, NL v2 PPR2, NL v2 PPR3, NL v2 PPR4, NL v2 PPR5, SEQ ID NOs: 51 - 60). The PPR expression plasmid was transformed into Rosetta(DE3) strain. This Escherichia coli was cultured in 2 mL of LB medium containing 100 μg / mL ampicillin at 37°C for 12 hours, and when OD 600 reached 0.5 to 0.8, the culture solution was transferred to an incubator at 15°C and allowed to stand for 30 minutes. Then, 100 μL (final concentration 0.1 mM IPTG) was added and cultured at 15°C for 16 hours. The Escherichia coli pellet was collected by centrifugation at 5,000 x g, 4°C for 10 minutes, and 1.5 mL of lysis buffer (20 mM Tris-HCl, pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM MgCl2, 2 mg / mL lysozyme, 1 mM PMSF, 2 μL of 10 mg / mL DNase) was added and frozen at -80°C for 20 minutes. Cell freeze-thaw disruption was performed while permeating at 25°C for 30 minutes. Subsequently, centrifugation was carried out at 3,700 rpm, 4°C for 15 minutes, and the supernatant (Escherichia coli lysate) containing soluble PPR protein was collected.
[0135] A 30-base sequence containing 18 bases of the target sequence was designed, and an RNA probe (RNA 1, RNA 2, RNA 3, RNA 4, RNA 5, SEQ ID NOs: 61 - 65) was synthesized (Grainer). The 5'-end biotinylated RNA probe was added to a streptavidin-coated plate (Thermo fisher), reacted at room temperature for 30 minutes, and washed with lysis buffer (20 mM Tris-HCl, pH 7.6, 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT, 0.1% BSA). For background measurement, wells were also prepared with 100 μL of lysis buffer, 1 μL of 100 mM DTT, and 1 μL of 40 unit / μL RNase inhibitor (Takara, 2313A) without adding RNA. Then, 200 μL of blocking buffer (20 mM Tris-HCl, pH 7.6, 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT, 1% BSA) was added, and the plate surface was blocked at room temperature for 30 minutes. Then, 100 μL of Escherichia coli lysate containing luciferase-fused PPR protein with a luminescence of 1.5 x 10 8 LU / μL was added, and a binding reaction was carried out at room temperature for 30 minutes. Then, it was washed 5 times with 200 μL of washing buffer (20 mM Tris-HCl, pH 7.6, 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT). 40 μL of luciferase substrate (Promega, E151A) diluted 2,500-fold with washing buffer was added to the wells and reacted for 5 minutes, and then the luminescence was measured with a plate reader (PerkinElmer, 5103-35). The value obtained by subtracting the background signal (the luminescence signal value when adding PPR protein without adding RNA) from the luminescence of the sample with each RNA and PPR protein added was taken as the binding affinity between PPR and RNA.
[0136] The results are shown in Fig. 8. When prepared with the motif sequence v2, an increase in the binding affinity to the target sequence (1.3 - 3.6 times) was observed in all cases compared to those prepared with the motif sequence v1. Also, two types of RNA probes having non-target sequences (off target 1, off target 2) (SEQ ID NOs: 66, 69) were prepared, and as a result of examining the binding thereto, the target binding signal / non-target binding signal (S / N) of v2 was higher than that of v1 in all cases (upper left in Fig. 8). From this, it was found that v2 has higher affinity and higher specificity for the target than v1.
[0137] [Example 4: Detailed analysis of the RNA binding performance of the PPR protein prepared using the v2 motif] (Specificity evaluation) Using the v2 motif, PPR proteins against 23 types of target sequences (T 1 - 3, T 5 - T 24, SEQ ID NOs: 46 - 48, 51 - 69) were prepared (NL v2 PPR1 - 3, NL v2 PPR5 - 24, SEQ ID NOs: 56 - 58, 70 - 88), and all binding combinations were analyzed using RPB - ELISA. The experimental method was the same as in Example 3.
[0138] The results are shown in Fig. 9 (upper). It was found that 21 types of PPR proteins except for v2 17 and v2 24 have the strongest binding affinity to the target. From these results, it was shown that the PPR protein prepared by using the v2 motif can be stably prepared and has high binding specificity.
[0139] Using the V3.1 motif, PPR proteins for the same 23 types of target sequences were similarly prepared (base sequences: SEQ ID NOs: 411 - 433, amino acid sequences: SEQ ID NOs: 434 - 456), and all binding combinations were analyzed using RPB - ELISA. The experimental method was the same as in Example 3.
[0140] The results are shown in Figure 9 (below) and the following table. Those with improved binding ability compared to V2 were obtained.
[0141]
Table 9
[0142] In addition, the RNA - binding performance of the PPR proteins shown in Figure 9 is shown in the following table in numerical values (log2 values).
[0143]
Table 10 - 1
[0144]
Table 10 - 2
[0145] (Affinity evaluation) Furthermore, to calculate the affinity (Kd value) of each PPR protein for the target RNA, EMSA was performed. Among the above 23 types, a streptavidin - binding peptide sequence was added to the N - terminal side and a 6x histag sequence was added to the C - terminal side of the gene sequences encoding 10 types of PPR proteins to construct an E. coli expression plasmid (SBP v2 PPR1 HIS, SBP v2 PPR 2 HIS, SBP v2 PPR 3 HIS, SBP v2 PPR6 HIS, SBP v2 PPR9 HIS, SBP v2 PPR12 HIS, SBP v2 PPR15 HIS, SBP v2 PPR16 HIS, SBP v2 PPR20 HIS, SBP v2 PPR24 HIS, SEQ ID NOs: 89 - 97). Transformed into Escherichia coli strain Rosetta(DE3), cultured overnight at 37°C in 2 mL of LB medium containing 100 μg / mL ampicillin. Then, 2 mL of the culture solution was transferred to 300 mL of LB medium containing 100 μg / mL ampicillin, and cultured at 37°C until the OD 600 reached 0.5 to 0.8. After cooling the cultured medium to 15°C, 0.1 mM IPTG was added and further cultured for 12 hours. The culture solution was centrifuged at 5000 x g, 4°C for 10 minutes to collect the cells, 5 mL of lysis buffer (20 mM Tris-HCl, pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM DTT, and 1 mM EDTA) was added, stirred with a Voltex, and then sonicated to disrupt the cells. Centrifuged at 15000 rpm, 4°C for 10 minutes, and the supernatant was collected.
[0146] Subsequently, the target protein was purified by affinity chromatography using the SBP tag. 100 μL of Streptavidine Sepharose High Performance (GE Healthcare, 17511301) was taken, and after the beads were collected by spin-down, they were equilibrated with a washing buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40). The equilibrated beads were gently mixed with the cell extract collected earlier and allowed to permeate at 4°C for 10 minutes. Subsequently, the entire volume of this bead solution was applied to the column, and the beads were washed with 10 mL of the washing buffer. Elution was performed with an elution buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 2 mM biotin).
[0147] Subsequently, affinity purification using the histidine tag was performed. First, 200 μL of Ni-NTA agarose (Qiagen, 30230) was collected, and after centrifugation, the beads were collected. 100 μL of the washing buffer was added, and the beads were equilibrated by allowing permeation at 4°C for 1 hour. The entire volume of the equilibrated beads was mixed with the protein solution eluted from the SBP beads and reacted at 4°C for 1 hour. Then, after centrifuging at 2,000 rpm for 2 minutes to collect the beads, factors that non-specifically bind to the beads were removed with 10 mL of the washing buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40, and 10 mM imidazole). Elution was performed with 60 μl of the elution buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40, and 500 mM imidazole). The degree of purification was confirmed by SDS-PAGE. Dialysis was performed overnight at 4°C with (20 mM Tris-HCl, pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM DTT, and 1 mM EDTA).
[0148] The Pierce 660nm Protein Assay Kit (Thermo fisher, 22662) was used to estimate the total amount of protein after dialysis. Next, to determine the amount of the target protein, first, the sample after dialysis was subjected to SDS-PAGE on a 10% polyacrylamide gel and stained with CBB. The gel image after staining was captured using a ChemiDoc Touch MP imaging system (Biorad). The total band intensity and that of the target band were derived from this gel image. The amount of the target protein was calculated by multiplying the ratio of the intensity of the target band to all band intensities by the total protein amount. Using this value, the molar concentration of the purified protein in the dialysis sample was calculated. Based on the molar concentration determined above, diluted protein solutions of 400 nM, 200 nM, 100 nM, 50 nM, 20 nM, 10 nM, 5 nM, 2 nM, and 1 nM were prepared. At that time, they were diluted using a binding buffer (20 mM Tris-HCl, pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM DTT, 1 mM EDTA). The RNA probe (RNA 1, RNA 2, RNA 3, RNA 6, RNA 9, RNA 12, RNA 15, RNA 16, RNA 20, RNA 24) biotinylated on the 5'-side was adjusted to a final concentration of 20 nM with the binding buffer. The RNA sample heat-treated at 75°C for 1 minute and then rapidly cooled was used in the following experiment.
[0149] The 20 nM RNA probe solution was mixed with each of the protein solutions at the concentrations adjusted above, and a binding reaction was carried out at 25°C for 20 minutes.
[0150] After the reaction, 2 μL of 80% glycerol was added and well suspended, and then 10 μL of it was applied to an ATTO 7.5% gel and electrophoresis was carried out at a C.V. of 150 V for 30 minutes.
[0151] The gel after electrophoresis was transferred to a Hybond N + membrane (GE, RPN203B). Subsequently, RNA was UV cross-linked to the membrane using an ASTEC Dual UV Transiluminator UVA-15 (Astec, 49909-06). The membrane was blocked with a blocking buffer (6.7 mM NaH2PO4·2H2O, 6.7 mM Na2HPO4·2H2O, 125 mM NaCl, 5% SDS). At this time, 0.5 μL of Stereptavidine-HRP (Abcam, ab7403) was pre-added to the blocking buffer, and the antigen-antibody reaction was carried out while infiltrating for 15 minutes. The blocking buffer was discarded, 20 mL of a washing buffer (0.67 mM NaH2PO4·2H2O, 0.67 mM Na2HPO4·2H2O, 12.5 mM NaCl, 0.5% SDS) was added, and the membrane was washed. Subsequently, after repeating this washing operation 5 times, 20 ml of an equilibration buffer (100 mM Tris-HCl, pH9.5, 100 mM NaCl, 10 mM MgCl2) was added, and the membrane was infiltrated for 5 minutes. Thereafter, Immunobilon Western Chemiluminisecnet HRP Substrate (Millipore, Cat No. WBKLS0100) was added to the membrane, and the biotinylated RNA was detected by chemiluminescence to detect the band. The gel image was captured using a ChemiDoc Touch MP imaging system (Biorad). The band intensity of the band shifted by binding of the unbound RNA probe to the protein was calculated. The equilibrium dissociation constant (Kd value) was calculated by fitting to the Hill equation from the molar concentration of the protein and the ratio of the corresponding shifted band.
[0152] The results are shown in Fig. 10. The prepared PPR protein was found to have a Kd value of 10 -9 - 10 -7 M. The minimum value (high affinity) was 1.95 x 10 -9Yes, it has the lowest Kd value among the reported designed PPR proteins so far (see Table 1). Also, these Kd values were found to correlate with the signal values obtained from the binding experiments in RPB-ELISA (R 2 = -0.85). From these results, when the luminescence value in RPB-ELISA is 1 - 2 x 10 7 , the Kd value is 10 -6 - 10 -7 M; when the luminescence value in RPB-ELISA is 2 - 4 x 10 7 , the Kd value is 10 -7 - 10 -8 M; when the luminescence value in RPB-ELISA is greater than 4 x 10 7 , it can be estimated that the Kd value is ~10 -8 or less.
[0153] (Evaluation of construction success rate) Furthermore, PPR proteins against 72 types of target sequences (T1 - T3, T6 - T76, SEQ ID NOs: 46 - 48, 51 - 69, 117 - 168) were prepared using the v2 motif (NL v2 PPR1 - 3, NL v2 PPR6 - 76, SEQ ID NOs: 56 - 58, 70 - 88, 169 - 220), and the construction success rate was calculated using RPB-ELISA. Biotinylated RNA probes containing the target sequences (RNA 1 - 3, RNA 6 - 76, SEQ ID NOs: 61 - 63, 98 - 116, 221 - 272) and biotinylated RNA probes containing non-target sequences (T 51, SEQ ID NO: 143) (RNA51, SEQ ID NO: 247) were prepared (Greinar). The experimental method was the same as in Example 3. The results are shown in Figure 11.
[0154] Among the 72 PPR proteins, those with a Kd value of 10 -6 M or less (RPB-ELISA value of 1 x 10 7are presumed to be 63 (88%), with a Kd value of 10 -7 M or less (RPB-ELISA value of 2 x 10 7 are presumed to be 57 (79%), with a Kd value of 10 -8 M or less (RPB-ELISA value of 4 x 10 7 are presumed to be 43 (40%). Also, the value obtained by dividing the target binding signal by the non-target binding signal was defined as the specificity evaluation value (S / N). Among them, 54 (75%) had an S / N higher than 10, and 23 (32%) had an S / N higher than 100. From these results, it was shown that sequence-specific RNA-binding proteins can be efficiently produced by using the v2 motif to produce PPR proteins.
[0155] (Evaluation of target binding activity related to the number of PPR motifs) Analysis was performed on the number of PPR motifs and target binding activity. Thirteen 18-base target sequences were set, and 3 bases and 6 bases on the 5' side of each sequence were removed, respectively, to obtain 15-base (T 1a, T 49a, T 3a, T 14a, T 40a, T 12a, T 13a, T 2a, T 38a, T 37a, T 39a, T 56a, T 68a, SEQ ID NOs, 273, 275, 277, 279, 281, 283, 285, 287, 289, 291, 293, 395, 297), 12-base (T 1b, T 49b, T 3b, T 14b, T 40b, T 12b, T 13b, T 2b, T 38b, T 37b, T 39b, T 56b, T 68b, SEQ ID NOs: 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298) target sequences were set. Corresponding PPR proteins (the 15-motif was named PPRxa, and the 12-motif was named PPRxb) were prepared (NL v2 PPR1, 1a, 1b; NL v2 PPR49, 49a, 49b; NL v2 PPR3, 3a, 3b; NL v2 PPR14, 14a, 14b; NL v2 PPR40, 40a, 40b; NL v2 PPR12, 12a, 12b; NL v2 PPR13, 13a, 13b; NL v2 PPR2, 2a, 2b; NL v2 PPR38, 38a, 38b; NL v2 PPR37, 37a, 37b; NL v2 PPR39, 39a, 39b; NL v2 PPR56, 56a, 56b; NL v2 PPR68, 68a, 68b; SEQ ID NOs:56, 299 - 324). To perform analysis by RPB-ELISA, biotinylated RNA probes containing the target sequences (T 1, T 49, T 3, T 14, T 40, T 12, T 13, T 2, T 38, T 37, T 39, T 56, T 68) and a non-target sequence (T 51, SEQ ID NO:143), a biotinylated RNA probe (RNA 51, SEQ ID NO:247) was prepared, and the binding activities of the target (on target) and non-target (off target) with their respective PPR proteins were analyzed by RPB-ELISA.
[0156] The results for each target sequence are shown in Fig. 12A. Fig. 12B shows the average values of the 18-motif, 15-motif, and 12-motif values plotted as box-and-whisker plots. It was found that the higher the number of motifs, the higher the binding strength, and when comparing the 18-motif and 15-motif, proteins with higher binding strength could be stably produced in the case of the 18-motif.
[0157] [Example 5: Artificial splicing control by PPR protein] To demonstrate that the PPR protein binds to the target RNA molecule intracellularly and enables the desired RNA manipulation, an experiment using a splicing reporter was conducted (Figure 13A). The splicing reporter (RG6) has a gene structure consisting of exon 1, intron 1, exon 2, intron 2, exon 3, etc. (Orengo et al., 2006 NAR). In intron 1, exon 2, and intron 2, the intron 4, intron 5 of chicken cTNT, and an artificially created alternative exon sequence are inserted. This reporter has two splicing forms, and the quantitative ratio of the mRNA with exon 2 skipped to the mRNA without exon 2 skipped is approximately 1:1. Also, in exon 3, the RFP and GFP genes are encoded, but due to the change in the reading frame depending on the presence or absence of exon 2, RFP is expressed in the mRNA with exon 2 skipped, and GFP is expressed in the mRNA without exon 2 skipped. It is known that the amount of the splicing form of this reporter is controlled by splicing factors that bind to the regions of intron 1, exon 2, and intron 2 (Orengo et al., 2006 NAR). Therefore, 18-base sequences were selected from the regions of intron 1, exon 2, and intron 2, and an experiment was conducted to determine whether the splicing form of the RG-6 reporter could be changed by the PPR proteins that bind to them.
[0158] Seven target sequences (T77 - T83, SEQ ID NOs: 325 - 330) were selected from the RG6 reporter. The PPR protein genes were designed with both the v1 motif and the v2 motif (v1 PPRsp1 - 6, v2 PPRsp1 - 6, SEQ ID NOs: 331 - 342). The protein in which the nuclear localization signal was fused to the N-terminal side and the FLAG epitope tag sequence was fused to the C-terminal side of the PPR protein gene was cloned into pcDNA3.1 so that it would be expressed (NLS v1PPRsp1 - 6, NLS v2PPRsp1-6, SEQ ID NOs: 343-354). pcDNA3.1 has a CMV promoter and an SV40 polyA signal (terminator), and the PPR protein gene was inserted between them.
[0159] HEK293T cells were seeded at 1 x 10 6 cells / well in a 10 cm dish containing 9 mL of DMEM and 1 mL of FBS. After culturing for 2 days in an environment of 37°C and 5% CO2, the cells were harvested. The harvested cells were seeded at 4 x 10 4 cells / well into a PLL-coated 96-well plate and cultured for 1 day in an environment of 37°C and 5% CO2. 100 ng of PPR expression plasmid DNA, 100 ng of RG-6, 0.6 μL of Fugene®-HD (Promega, E2311), and 200 μL of Opti-MEM were mixed and added to the full volume of each well, and then cultured for 2 days in an environment of 37°C and 5% CO2. As a control, a sample without adding PPR expression plasmid DNA was also prepared. After culturing, GFP fluorescence and RFP fluorescence images of each well were acquired using a fluorescence microscope DMi8 (Leica). For the imaging conditions, first, using the sample transfected with only the RG-6 plasmid, the exposure time and gain were determined such that the intensities of GFP and RFP were comparable, and then fluorescence images of each sample were acquired under the same conditions.
[0160] After image acquisition, total RNA was extracted using the Maxwell (registered trademark) RSC simplyRNA Cells Kit. 500 ng of the extracted total RNA, 0.5 μL of 100 μM dT20 primer, and 0.5 μL of 10 mM dNTPs were added to a 0.2 mL tube, incubated at 65 °C for 5 minutes, and then immediately cooled on ice. To this, 2 μL of 5x RT-buffer (Invitrogen, 18080-051), 0.5 μL of 0.1 M DTT, 0.5 μL of 40 U / μL RNaseOUT (Invitrogen, 18080-051), and 0.5 μL of 200 unit / μL SupperScript III (Invitrogen, 18080-051) were added, reacted in a thermal cycler at 50 °C for 50 minutes, then reacted at 85 °C for 5 minutes, and cooled to 16 °C. The reverse-transcribed sample was diluted 10-fold with sterile water. 2 μL of this, 10 μL of 5x GXL buffer (TAKARA, R050A), 4 μL of 2.5 mM dNTPs, 1.5 μL of 10 μM RT-Fw primer (5'-CAAAGTGGAGGACCCAGTACC-3') (SEQ ID NO:355), 1.5 μL of 10 μM RT-Rv primer (5'-GCGCATGAACTCCTTGATGAC-3') (SEQ ID NO:356), 1 μL of GXL (TAKARA, R050A), and 31.5 μL of sterile water were added to a 0.2 mL tube, reacted in a thermal cycler at 98 °C for 2 minutes, then 98 °C for 10 seconds, 58 °C for 15 seconds, and 68 °C for 5 seconds were repeated 35 times, and then cooled to 12 °C. The reaction solution was diluted 10-fold and electrophoresed using MultiNA (SHIMADZU, MCE202). The band around 114 bp was taken as the band of exon-skipped RNA, the band around 142 bp was taken as the band of non-skipped RNA, and the band intensity in each sample was calculated. The value obtained by dividing the 114 bp band intensity by the sum of the 114 bp band intensity and the 142 bp band intensity was defined as the splicing ratio.
[0161] The results are shown in FIGS. 13B and C. When only the RG6 reporter was introduced, the splicing ratio was 0.48. When PPRsp4 was introduced, it was about the same, but when other PPRs were introduced, it was found that there was a large change. Comparing v1 and v2, except for PPRsp4, v2 changed more significantly. These splicing ratios were also consistent with the RFP and GFP expression ratios in FIG. 13B. From these results, it was demonstrated that exon skipping can be changed by using PPR proteins, and it was also found that splicing can be changed more efficiently by using the v2 motif.
[0162] [Example 6: Control of Aggregation of PPR Protein] The PPR protein using the V2 motif (base sequence: SEQ ID NO: 457, amino acid sequence: SEQ ID NO: 458) and the PPR protein using the v3.2 motif (base sequence: SEQ ID NO: 459, amino acid sequence: SEQ ID NO: 460) were each prepared and purified in an E. coli expression system and separated by gel filtration chromatography.
[0163] (Expression and Purification of Protein) Using the pE-SUMOpro Kan plasmid containing the DNA sequence encoding the target PPR protein, transform the Escherichia coli Rosetta strain. After culturing at 37°C and when the OD600 reaches 0.6, lower the temperature to 20°C and add IPTG to a final concentration of 0.5 mM to express the target PPR protein as a SUMO fusion protein in Escherichia coli. After culturing overnight, collect the bacterial cells by centrifugation and resuspend them in Lysis Buffer (50 mM Tris-HCl pH 8.0, 500 mM NaCl). Disrupt Escherichia coli by sonication, and after centrifugation at 17,000 g for 30 min, apply the supernatant fraction to a Ni-Agarose column. After washing the column with Lysis Buffer containing 20 mM imidazole, elute the SUMO fusion target PPR protein with Lysis Buffer containing 400 mM imidazole. After elution, cleave the SUMO protein from the target PPR protein using Ulp1, and at the same time, replace the protein solution with ion exchange Buffer (50 mM Tris-HCl pH 8.0, 200 mM NaCl) by dialysis. Then, perform cation exchange chromatography using an SP column. After applying the column, elute the protein by gradually increasing the NaCl concentration from 200 mM to 1 M. Final purification of the fraction containing the target PPR protein was performed by gel filtration chromatography using a Superdex200 column. Apply the target PPR protein eluted from ion exchange to a gel filtration column equilibrated with Gel Filtration Buffer (25 mM HEPES pH 7.5, 200 mM NaCl, 0.5 mM tris(2-carboxyethyl)phosphine (TCEP)). Finally, concentrate the fraction containing the target PPR protein, freeze it in liquid nitrogen, and store it at -80°C until use in the next analysis.
[0164] (Gel filtration chromatography) The purified recombinant PPR protein was adjusted to a concentration of 1 mg / ml. Gel filtration chromatography was performed using Superdex 200 increase 10 / 300 GL (GE Helthcare). The adjusted protein was applied to a gel filtration column equilibrated with 25 mM HEPES pH 7.5, 200 mM NaCl, 0.5 mM tris(2-carboxyethyl)phosphine (TCEP), and the properties of the protein were analyzed by measuring the absorbance at 280 nm of the solution eluted from the gel filtration column.
[0165] (Results) The results are shown in Figure 14. The smaller the elution fraction (Elution vol.), the larger the molecular size. At V2, elution occurred in the elution fraction of 8 to 10 mL, while at v3.2, a peak was observed in the elution fraction of 12 to 14 mL. From this, it was suggested that at v2, the protein size was large and it might be aggregated, and it was found that the aggregation was improved at v3.2.
Claims
【Claim 1】 A protein, or a fusion protein of the protein and at least one selected from the group consisting of a fluorescent protein, a nuclear localization signal peptide, and a tag protein A composition for treating a disease, comprising the protein has the following formula 1 【Chemical 1】 (wherein: Helix A is a 12-amino acid-long part capable of forming an α-helix structure, represented by formula 2, 【Chemical 2】 In Formula 2, A 1 ~ A 12 each independently represents an amino acid; X is absent or a part consisting of 1 to 9 amino acids; Helix B is a part capable of forming an α-helix structure, consisting of 11 to 13 amino acids; L is a part represented by formula 3, consisting of 2 to 7 amino acids; [Chemical 3] In formula 3, each amino acid is numbered from the C-terminal side as "i" (-1), "ii" (-2), However, L iii ~L vii may not exist.) a protein containing a polypeptide which is a PPR motif represented by and capable of binding to a target nucleic acid having a specific base sequence A composition, wherein at least one polypeptide is any of the following. (A-1) A polypeptide consisting of the sequence of SEQ ID NO: 9, or a polypeptide consisting of an amino acid sequence obtained by performing any substitution selected from the group consisting of substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid in the sequence of SEQ ID NO: 9, or a polypeptide consisting of the sequence of SEQ ID NO: 401, or a polypeptide consisting of an amino acid sequence obtained by performing any substitution selected from the group consisting of substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid in the sequence of SEQ ID NO: 401; (A-2') A polypeptide which is a PPR motif represented by the following formula 1, which consists of a sequence obtained by substituting, deleting, or adding 1 to 20 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 9 or 401 and is adenine-binding (provided that the polypeptides in the following group A are excluded.); A polypeptide which is a PPR motif represented by the following formula 1, having at least 42% sequence identity with the sequence of SEQ ID NO: 9 or 401, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are the same and it is adenine-binding (however, excluding the polypeptides in group A below); A polypeptide consisting of the sequence of SEQ ID NO: 10, or a polypeptide consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 2 with serine, substitution of the amino acid at position 5 with isoleucine, substitution of the amino acid at position 7 with leucine, substitution of the amino acid at position 8 with lysine, substitution of the amino acid at position 10 with phenylalanine or tyrosine, substitution of the amino acid at position 15 with arginine, substitution of the amino acid at position 22 with valine, substitution of the amino acid at position 24 with arginine, substitution of the amino acid at position 27 with leucine, and substitution of the amino acid at position 29 with arginine is made in the sequence of SEQ ID NO: 10; A polypeptide which is a PPR motif represented by the following formula 1, consisting of a sequence in which 1 to 25 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 10 are substituted, deleted, or added, and it is cytosine-binding (however, excluding the polypeptides in group C below); A polypeptide which is a PPR motif represented by the following formula 1, having at least 25% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are the same and it is cytosine-binding (however, excluding the polypeptides in group C below); A polypeptide consisting of the sequence of SEQ ID NO: 11, or a polypeptide consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with phenylalanine, substitution of the amino acid at position 15 with aspartic acid, substitution of the amino acid at position 27 with valine, substitution of the amino acid at position 28 with serine, and substitution of the amino acid at position 35 with isoleucine is made in the sequence of SEQ ID NO: 11; A polypeptide which is a PPR motif represented by the following formula 1, consisting of a sequence in which 1 to 21 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 11 are substituted, deleted, or added, and which is guanine-binding (however, excluding the polypeptides in the following Group G); A polypeptide which is a PPR motif represented by the following formula 1, having at least 40% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is guanine-binding (however, excluding the polypeptides in the following Group G); A polypeptide consisting of the sequence of SEQ ID NO: 12, or a polypeptide consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with phenylalanine, substitution of the amino acid at position 13 with serine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 20 with leucine, substitution of the amino acid at position 21 with lysine, substitution of the amino acid at position 23 with phenylalanine, substitution of the amino acid at position 24 with aspartic acid, substitution of the amino acid at position 27 with lysine, substitution of the amino acid at position 28 with lysine, substitution of the amino acid at position 29 with arginine, and substitution of the amino acid at position 31 with leucine is made in the sequence of SEQ ID NO: 12; A polypeptide which is a PPR motif represented by the following formula 1, consisting of a sequence in which 1 to 22 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 12 are substituted, deleted, or added, and which is uracil-binding (however, excluding the polypeptides in the following Group U); A polypeptide which is a PPR motif represented by the following formula 1, having at least 37% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are the same, and which is uracil-binding (however, excluding the polypeptides in the following Group U). Group A: polypeptides consisting of the sequences at positions 545 - 579 of SEQ ID NO: 462, polypeptides consisting of the sequences at positions 409 - 443 of SEQ ID NO: 464, polypeptides consisting of the sequences at positions 431 - 465 of SEQ ID NO: 465, polypeptides consisting of the sequences at positions 437 - 471 of SEQ ID NO: 466 Group C: polypeptides consisting of the sequences at positions 318 - 352 of SEQ ID NO: 461 Group G: polypeptides consisting of the sequences at positions 869 - 903 of SEQ ID NO: 463, polypeptides consisting of the sequences at positions 374 - 408 of SEQ ID NO: 464, polypeptides consisting of the sequences at positions 396 - 430 of SEQ ID NO: 465, polypeptides consisting of the sequences at positions 402 - 436 of SEQ ID NO: 466 Group U: polypeptides consisting of the sequences at positions 353 - 387 of SEQ ID NO: 461, polypeptides consisting of the sequences at positions 475 - 509 of SEQ ID NO: 462, polypeptides consisting of the sequences at positions 380 - 414 of SEQ ID NO: 463
Citation Information
Patent Citations
Method for modifying RNA binding protein using PPR motif
WO2011111829A1
Design method for RNA-binding protein using PPR motif, and use thereof
WO2013058404A1
DNA binding protein using PPR motif, and use thereof
WO2014175284A1
DNA-binding protein using PPR motif and use of said DNA-binding protein
WO2018030488A1
Proteins and their use for nucleotide binding
WO2019232588A1