Method for producing efficient PPR protein and its use
Novel PPR motifs and proteins are designed to address the challenge of binding to long target RNA sequences, achieving enhanced binding affinity and specificity, thereby improving RNA manipulation efficiency.
Patent Information
- Application Number
- JP2023126196
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-29
- Filing Date
- 2023-08-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-05-29
AI Technical Summary
Current PPR proteins struggle to specifically bind to long target RNA sequences, which is essential for efficient RNA manipulation and operations in cells.
Development of novel PPR motifs and proteins with enhanced binding performance, specifically designed to target RNA sequences of 15 bases or longer, by optimizing the amino acid sequences and combinations at specific positions within the PPR motifs.
The novel PPR motifs and proteins demonstrate improved binding affinity and specificity to target RNA sequences, enhancing RNA manipulation efficiency and reducing unintended binding to non-target RNAs.
Smart Images

Figure 0007683865000019 
Figure 0007683865000020 
Figure 0007683865000021
Abstract
Description
Technical Field
[0001] The present invention relates to a nucleic acid manipulation technique using a protein capable of binding to a target nucleic acid. The present invention is useful in a wide range of fields such as medicine (drug discovery support, treatment), agriculture (agricultural, forestry, livestock and fishery product production, breeding), and chemistry (biological substance production).
Background Art
[0002] The PPR protein is a protein containing a repeat of PPR motifs, each about 35 amino acids long, and one PPR motif can specifically bind to one base. The combination of the first, fourth, and ii-th (two before the next motif) amino acids in the PPR motif determines which of adenine, cytosine, guanine, and uracil (or thymine) it binds to (Patent Documents 1 and 2).
[0003] Since one PPR motif recognizes and binds to one base, for example, when designing a PPR protein that binds specifically to an 18-base-long nucleic acid sequence, 18 PPR motifs need to be linked. So far, the production of artificial PPR proteins with 7 to 14 linked PPR motifs has been reported (Non-Patent Documents 1 to 6).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Non-Patent Documents
[0005]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Non-Patent Document 6
Summary of the Invention
Problems to be Solved by the Invention
[0006] In order to specifically bind to a target RNA molecule in a cell and further perform desired operations, a PPR protein with high performance is required.
[0007] Also, in order to specifically bind to a target RNA molecule in a cell and further perform desired operations, a PPR protein that links more motifs than the conventional 7 - 14 and binds to a long base sequence is required. For example, the human genome has 6 billion bases and is composed of 4 bases (A, C, G, T or U). Therefore, in order to specify a single base sequence from the sequence arrangement, at least a 17 - base sequence is required (4 16 to the power of 4 billion, 4 17 to the power of 16 billion).
Means for Solving the Problems
[0008] The present invention provides the following as novel PPR motifs and the like. [1] Any one of the following PPR motifs: (A-1) A PPR motif consisting of the sequence of SEQ ID NO: 9, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid is made in the sequence of SEQ ID NO: 9, or a PPR motif consisting of the sequence of SEQ ID NO: 401, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid is made in the sequence of SEQ ID NO: 401; (A-2) A PPR motif consisting of a sequence in which 1 to 20 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are substituted, deleted, or added in the sequence of SEQ ID NO: 9 or 401, and which is adenine-binding; (A-3) A PPR motif having at least 42% sequence identity with the sequence of SEQ ID NO: 9 or 401, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are identical, and which is adenine-binding; (C-1) A PPR motif consisting of the sequence of SEQ ID NO: 10, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 2 with serine, substitution of the amino acid at position 5 with isoleucine, substitution of the amino acid at position 7 with leucine, substitution of the amino acid at position 8 with lysine, substitution of the amino acid at position 10 with phenylalanine or tyrosine, substitution of the amino acid at position 15 with arginine, substitution of the amino acid at position 22 with valine, substitution of the amino acid at position 24 with arginine, substitution of the amino acid at position 27 with leucine, and substitution of the amino acid at position 29 with arginine is made in the sequence of SEQ ID NO: 10; (C-2) A PPR motif consisting of a sequence in which 1 to 25 amino acids other than those at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 10 are substituted, deleted, or added, and which is cytosine-binding; (C-3) A PPR motif having at least 25% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is cytosine-binding; (G-1) A PPR motif consisting of the sequence of SEQ ID NO: 11, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with phenylalanine, substitution of the amino acid at position 15 with aspartic acid, substitution of the amino acid at position 27 with valine, substitution of the amino acid at position 28 with serine, and substitution of the amino acid at position 35 with isoleucine has been made in the sequence of SEQ ID NO: 11; (G-2) A PPR motif consisting of a sequence in which 1 to 21 amino acids other than those at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 11 are substituted, deleted, or added, and which is guanine-binding; (G-3) A PPR motif having at least 40% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is guanine-binding; (U-1) A PPR motif consisting of the sequence of SEQ ID NO: 12, or an amino acid sequence obtained by performing any substitution selected from the group consisting of substitution of the phenylalanine at position 10, substitution of the serine at position 13, substitution of the lysine at position 15, substitution of the glutamate at position 17, substitution of the leucine at position 20, substitution of the lysine at position 21, substitution of the phenylalanine at position 23, substitution of the aspartic acid at position 24, substitution of the lysine at position 27, substitution of the lysine at position 28, substitution of the arginine at position 29, and substitution of the leucine at position 31 in the sequence of SEQ ID NO: 12 (U-2) A PPR motif consisting of a sequence in which 1 to 22 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 12 are substituted, deleted, or added, and which is uracil-binding (U-3) A PPR motif having at least 37% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are identical and which is uracil-binding [2] Use of the PPR motif according to 1 for the production of a PPR protein in which the target RNA is 15 bases or longer [3] Use for the production of a PPR protein of the PPR motif according to 1, which is used to enhance the binding performance of the PPR protein to the target RNA [4] A protein containing n PPR motifs capable of binding to a target RNA consisting of n base sequences, wherein the PPR motif for adenine in the base sequence is the PPR motif of (A-1), (A-2), or (A-3) defined in 1; wherein the PPR motif for cytosine in the base sequence is the PPR motif of (C-1), (C-2), or (c-3) defined in 1; wherein the PPR motif for guanine in the base sequence is the PPR motif of (G-1), (G-2), or (G-3) defined in 1; A PPR protein in which the PPR motif for uracil in the base sequence is the PPR motif of (U-1), (U-2), or (U-3) defined in 1. [5] The protein according to 4, wherein n is 15 or more. [6] The protein according to 4 or 5, wherein the first PPR motif from the N-terminus is any one of the following: (1st (A-1) A PPR motif in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 402 are substituted so as to satisfy any one of the combinations defined below; (1st (A-2) (1st It consists of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of (A-1) are substituted, deleted, or added, and is an adenine-binding PPR motif; (1st (A-3) (1st It has at least 80% sequence identity with the sequence of (A-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and is an adenine-binding PPR motif; (1st (C-1) A PPR motif consisting of the sequence of SEQ ID NO: 403; (1st (C-2) (1st It consists of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of (C-1) are substituted, deleted, or added, and is a cytosine-binding PPR motif; (1st (C-3) (1st It has at least 80% sequence identity with the sequence of (C-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and is a cytosine-binding PPR motif; (1st (G-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 404 are substituted so as to satisfy any one of the combinations defined below; (1st G-2)(1st A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of G-1 are substituted, deleted, or added, and which is guanine-binding; (1st G-3)(1st A PPR motif having at least 80% sequence identity with the sequence of G-1, provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and which is guanine-binding; (1st U-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 405 are substituted so as to satisfy any one of the combinations defined below; (1st U-2)(1st A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of U-1 are substituted, deleted, or added, and which is uracil-binding; (1st U-3)(1st A PPR motif having at least 80% sequence identity with the sequence of U-1, provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and which is uracil-binding. · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamic acid · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamine · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is lysine · A combination in which the amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine [7] A method for controlling RNA splicing, comprising using the protein according to any one of items 4 to 6. [8] A method for detecting RNA, comprising using the protein according to any one of items 4 to 6. [9] A fusion protein comprising at least one selected from the group consisting of a fluorescent protein, a nuclear localization signal peptide, and a tag protein, and the protein according to any one of items 4 to 6.
[10] The nucleic acid encoding the PPR motif according to 1, or the protein according to any one of items 4 to 6.
[11] A vector containing the nucleic acid according to 10.
[12] A cell (excluding human individuals) containing the vector according to 11.
[13] A method for manipulating RNA (excluding implementation in human individuals) using the PPR motif according to 1, the protein according to any one of items 4 to 6, or the vector according to 11.
[14] A method for producing an organism comprising the manipulation method according to 13.
[0009] [1] Any one of the following PPR motifs: (A-1) A PPR motif consisting of the sequence of SEQ ID NO: 9, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid is made in the sequence of SEQ ID NO: 9; (A-2) A PPR motif consisting of a sequence in which 1 to 20 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 9 are substituted, deleted, or added, and which is adenine-binding; (A-3) A PPR motif having at least 42% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are the same, and which is adenine-binding; (C-1) A PPR motif consisting of the sequence of SEQ ID NO: 10, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of: substitution of the serine at position 2 in the sequence of SEQ ID NO: 10 with isoleucine; substitution of the isoleucine at position 5 with leucine; substitution of the leucine at position 7 with lysine; substitution of the lysine at position 8 with arginine; substitution of the phenylalanine or tyrosine at position 10 with arginine; substitution of the arginine at position 15 with valine; substitution of the valine at position 22 with arginine; substitution of the leucine at position 27 with arginine; and substitution of the arginine at position 29 has been performed; (C-2) A PPR motif consisting of a sequence in which 1 to 25 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 10 have been substituted, deleted, or added, and which is cytosine-binding; (C-3) A PPR motif having at least 25% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is cytosine-binding; (G-1) A PPR motif consisting of the sequence of SEQ ID NO: 11, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of: substitution of the phenylalanine at position 10 in the sequence of SEQ ID NO: 11 with aspartic acid; substitution of the aspartic acid at position 15 with valine; substitution of the valine at position 27 with serine; substitution of the serine at position 28 with isoleucine; and substitution of the isoleucine at position 35 has been performed; (G-2) A PPR motif consisting of a sequence in which 1 to 21 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 11 have been substituted, deleted, or added, and which is guanine-binding; (G-3) A PPR motif having at least 40% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is guanine-binding; (U-1) A PPR motif consisting of the sequence of SEQ ID NO: 12, or in the sequence of SEQ ID NO: 12, a substitution of the phenylalanine at position 10, a substitution of the serine at position 13, a substitution of the lysine at position 15, a substitution of the glutamate at position 17, a substitution of the leucine at position 20, a substitution of the lysine at position 21, a substitution of the phenylalanine at position 23, a substitution of the aspartic acid at position 24, a substitution of the lysine at position 27, a substitution of the lysine at position 28, a substitution of the arginine at position 29, and a substitution of the leucine at position 31, and a PPR motif consisting of an amino acid sequence with any substitution selected from the group consisting of these substitutions. (U-2) A PPR motif consisting of a sequence in which 1 to 22 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 12 are substituted, deleted, or added, and which is uracil-binding. (U-3) A PPR motif having at least 37% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are the same, and which is uracil-binding. [2] Use of the PPR motif according to 1 for the production of a PPR protein in which the target RNA is 15 bases or longer. [3] Use for the production of a PPR protein of the PPR motif according to 1, which is for enhancing the binding performance of the PPR protein to the target RNA. [4] A protein containing n PPR motifs capable of binding to a target RNA consisting of n base sequences, wherein the PPR motif for adenine in the base sequence is the PPR motif of (A-1), (A-2), or (A-3) defined in 1; wherein the PPR motif for cytosine in the base sequence is the PPR motif of (C-1), (C-2), or (c-3) defined in 1; wherein the PPR motif for guanine in the base sequence is the PPR motif of (G-1), (G-2), or (G-3) defined in 1; A PPR protein in which the PPR motif for uracil in the base sequence is the PPR motif of (U-1), (U-2), or (U-3) defined in 1. [5] The protein according to 4, wherein n is 15 or more. [6] A method for controlling RNA splicing, comprising using the protein according to 4 or 5. [7] A method for detecting RNA, comprising using the protein according to 4 or 5. [8] A fusion protein of at least one selected from the group consisting of a fluorescent protein, a nuclear localization signal peptide, and a tag protein, and the protein according to 4 or 5. [9] A nucleic acid encoding the PPR motif according to 1, or the protein according to 4 or 5.
[10] A vector containing the nucleic acid according to 9.
[11] A cell (excluding human individuals) containing the vector according to 10.
[12] A method for manipulating RNA (excluding implementation in human individuals), using the PPR motif according to 1, the protein according to 4 or 5, or the vector according to 10.
[13] A method for producing an organism, comprising the manipulation method according to 12.
[14] A method for preparing a gene encoding a protein containing n PPR motifs capable of binding to a target nucleic acid consisting of n base sequences, comprising the following steps: Selecting m PPR parts necessary for preparing a target gene from a library of at least 20×m types of PPR parts, wherein each of at least 20 types of polynucleotides encoding each of the PPR motifs having adenine, cytosine, guanine, or uracil or thymine-binding property and each of the 16 types of linkers of two PPR motifs is inserted into each of the intermediate vectors Dest-a... designed to be ligatable in at least m types of order; Subjecting the selected m types of PPR parts to a Golden Gate reaction together with vector parts to obtain a vector into which a linker of m polynucleotides is inserted (where n is not less than m and not more than m×2).
[15] The production method according to 14 for producing a gene encoding a protein containing 15 or more PPR motifs, where m is 10.
[16] A method for detecting or quantifying a protein containing n PPR motifs capable of binding to a target nucleic acid consisting of n base sequences, comprising the following steps: A step of subjecting a solution containing a candidate protein to a solid-phase target nucleic acid and detecting or quantifying the protein bound to the target nucleic acid.
[17] The method according to 16, wherein the candidate protein is fused with a labeled protein.
Brief Description of Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
[0011] [PPR Motif, PPR Protein] (Definition) When referring to the PPR motif in the present invention, unless otherwise specified, when analyzing the amino acid sequence with a protein domain search program on the web, it refers to a polypeptide composed of 30 to 38 amino acids having an E value of a predetermined value or less (preferably E-03) obtained by PF01535 in Pfam and PS51375 in Prosite. The position numbers of the amino acids constituting the PPR motif defined in the present invention are almost synonymous with PF01535, while they correspond to the numbers obtained by subtracting 2 from the positions of the amino acids in PS51375 (e.g., position 1 in the present invention → position 3 in PS51375). However, when referring to the amino acid at the "ii" (-2) position, it refers to the second amino acid from the end (C-terminal side) of the amino acids constituting the PPR motif, or two amino acids on the N-terminal side with respect to the first amino acid of the next PPR motif, that is, the -2nd amino acid. When the next PPR motif is not clearly identified, the amino acid two positions before the first amino acid of the next helix structure is defined as "ii". For Pfam, refer to http: / / pfam.sanger.ac.uk / , and for Prosite, refer to http: / / www.expasy.org / prosite / .
[0012] The conserved amino acid sequence of the PPR motif has low conservation at the amino acid level, but the two α-helices are well conserved in the secondary structure. A typical PPR motif is composed of 35 amino acids, but its length is variable from 30 to 38 amino acids.
[0013] More specifically, the PPR motif referred to in the present invention consists of a polypeptide having a length of 30 to 38 amino acids represented by Formula 1.
[0014]
Chemical formula
[0015]
Chemical formula
[0016]
Chemical formula
[0017] When referring to a PPR protein in the present invention, unless otherwise specified, it refers to a PPR protein having one or more, preferably two or more, of the above-described PPR motifs. When referring to a protein in this specification, unless otherwise specified, it refers to all substances composed of polypeptides (chains in which multiple amino acids are peptide-bonded), including those composed of relatively low-molecular-weight polypeptides. When referring to an amino acid in the present invention, it may usually refer to a normal amino acid molecule, or may refer to an amino acid residue constituting a peptide chain. Which one is being referred to will be clear to those skilled in the art from the context.
[0018] Regarding the binding property to a base in the target nucleic acid of the PPR motif in the present invention, when referring to specificity / specific, unless otherwise specified, it means that the binding activity to any one of the four types of bases is higher than the binding activity to other bases.
[0019] When referring to a nucleic acid in the present invention, it refers to RNA or DNA. Note that a PPR protein may have specificity for a base in RNA or DNA, but does not bind to a nucleic acid monomer.
[0020] The PPR motif is such that the combination of three amino acids at positions 1, 4, and ii is important for specific binding to bases, and these combinations can determine which base is to be bound (see Patent Documents 1 and 2 cited above).
[0021] Specifically, regarding the RNA-binding PPR motif, the relationship between the combination of three amino acids at positions 1, 4, and ii and the bases capable of binding is as follows (see Patent Document 1 cited above). (3-1) A 1 、A 4 、and L ii When the combination of the three amino acids is valine, asparagine, and aspartic acid in that order, the PPR motif has a selective RNA base-binding ability to strongly bind to U, then to C, and then to A or G. (3-2) A 1 、A 4 、and L ii When the combination of the three amino acids is valine, threonine, and asparagine in that order, the PPR motif has a selective RNA base-binding ability to strongly bind to A, then to G, and then to C, but not to U. (3-3) A 1 、A 4 、and L ii When the combination of the three amino acids is valine, asparagine, and asparagine in that order, the PPR motif has a selective RNA base-binding ability to strongly bind to C, and then to A or U, but not to G. (3-4) A 1 、A 4 、and L ii When the combination of the three amino acids is glutamic acid, glycine, and aspartic acid in that order, the PPR motif has a selective RNA base-binding ability to strongly bind to G, but not to A, U, and C. (3-5) A 1 、A 4 、and L iiWhen the combination of three amino acids is isoleucine, asparagine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to C, then to U, and then to A, but does not bind to G. (3-6) A 1 、A 4 、and L ii When the combination of three amino acids is valine, threonine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G and then to U, but does not bind to A and C. (3-7) A 1 、A 4 、and L ii When the combination of three amino acids is lysine, threonine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G and then to A, but does not bind to U and C. (3-8) A 1 、A 4 、and L ii When the combination of three amino acids is phenylalanine, serine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, then to C, and then to G and U. (3-9) A 1 、A 4 、and L ii When the combination of three amino acids is valine, asparagine, and serine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to C and then to U, but does not bind to A and G. (3-10) A 1 、A 4 、and L ii When the combination of three amino acids is phenylalanine, threonine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, but does not bind to G, U, and C. (3-11) A 1 、A 4 、and Lii When the combination of three amino acids is isoleucine, asparagine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to U and then to A, but does not bind to G and C. (3-12) A 1 , A 4 , and L ii When the combination of three amino acids is threonine, threonine, and asparagine in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to A, but does not bind to G, U, and C. (3-13) A 1 , A 4 , and L ii When the combination of three amino acids is isoleucine, methionine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to U and then to C, but does not bind to A and G. (3-14) A 1 , A 4 , and L ii When the combination of three amino acids is phenylalanine, proline, and aspartic acid in that order for PPR, its motif has selective RNA base-binding ability such that it binds strongly to U and then to C, but does not bind to A and G. (3-15) A 1 , A 4 , and L ii When the combination of three amino acids is tyrosine, proline, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to U, but does not bind to A, G, and C. (3-16) A 1 , A 4 , and L ii When the combination of three amino acids is leucine, threonine, and aspartic acid in that order, the PPR motif has selective RNA base-binding ability such that it binds strongly to G, but does not bind to A, U, and C.
[0022] Specifically, regarding the DNA-binding PPR motif, the relationship between the combination of three amino acids at positions 1, 4, and ii and the bases that can bind thereto is as follows (see Patent Document 2 cited above). (2-1) A 1 、A 4 、and L ii When the combination of the three amino acids is, in order, any amino acid, glycine, and aspartic acid, the PPR motif selectively binds to G; (2-2) A 1 、A 4 、and L ii When the combination of the three amino acids is, in order, glutamic acid, glycine, and aspartic acid, the PPR motif selectively binds to G; (2-3) A 1 、A 4 、and L ii When the combination of the three amino acids is, in order, any amino acid, glycine, and asparagine, the PPR motif selectively binds to A; (2-4) A 1 、A 4 、and L ii When the combination of the three amino acids is, in order, glutamic acid, glycine, and asparagine, the PPR motif selectively binds to A; (2-5) A 1 、A 4 、and L ii When the combination of the three amino acids is, in order, any amino acid, glycine, and serine, the PPR motif selectively binds to A and then binds to C; (2-6) A 1 、A 4 、and L ii When the combination of the three amino acids is, in order, any amino acid, isoleucine, and any amino acid, the PPR motif selectively binds to T and C; (2-7) A 1 、A 4 、and L iiWhen the combination of three amino acids is, in order, any amino acid, isoleucine, and asparagine, the PPR motif selectively binds to T and then binds to C; (2-8) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, leucine, and any amino acid, the PPR motif selectively binds to T and C; (2-9) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, leucine, and aspartic acid, the PPR motif selectively binds to C; (2-10) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, leucine, and lysine, the PPR motif selectively binds to T; (2-11) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, methionine, and any amino acid, the PPR motif selectively binds to T; (2-12) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, methionine, and aspartic acid, the PPR motif selectively binds to T; (2-13) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, isoleucine, methionine, and aspartic acid, the PPR motif selectively binds to T and then binds to C; (2-14) A 1 、A 4 、and L iiWhen the combination of three amino acids is, in order, any amino acid, asparagine, and any amino acid, the PPR motif selectively binds to C and T; (2-15) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-16) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, phenylalanine, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-17) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, glycine, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-18) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, isoleucine, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-19) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, threonine, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-20) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, valine, asparagine, and aspartic acid, the PPR motif selectively binds to T and then binds to C; (2-21) A 1 、A 4 、and L iiWhen the combination of three amino acids is tyrosine, asparagine, and aspartic acid in sequence, the PPR motif selectively binds to T and then binds to C; (2-22) A 1 、A 4 、and L ii When the combination of three amino acids is any amino acid, asparagine, and asparagine in sequence, the PPR motif selectively binds to C; (2-23) A 1 、A 4 、and L ii When the combination of three amino acids is isoleucine, asparagine, and asparagine in sequence, the PPR motif selectively binds to C; (2-24) A 1 、A 4 、and L ii When the combination of three amino acids is serine, asparagine, and asparagine in sequence, the PPR motif selectively binds to C; (2-25) A 1 、A 4 、and L ii When the combination of three amino acids is valine, asparagine, and asparagine in sequence, the PPR motif selectively binds to C; (2-26) A 1 、A 4 、and L ii When the combination of three amino acids is any amino acid, asparagine, and serine in sequence, the PPR motif selectively binds to C; (2-27) A 1 、A 4 、and L ii When the combination of three amino acids is valine, asparagine, and serine in sequence, the PPR motif selectively binds to C; (2-28) A 1 、A 4 、and L ii When the combination of three amino acids is any amino acid, asparagine, and threonine in sequence, the PPR motif selectively binds to C; (2-29) A 1 、A 4 、and L ii When the combination of three amino acids of, in order, valine, asparagine, and threonine, its PPR motif selectively binds to C; (2-30) A 1 、A 4 、and L ii When the combination of three amino acids of, in order, any amino acid, asparagine, and tryptophan, its PPR motif selectively binds to C and then binds to T; (2-31) A 1 、A 4 、and L ii When the combination of three amino acids of, in order, isoleucine, asparagine, and tryptophan, its PPR motif selectively binds to T and then binds to C; (2-32) A 1 、A 4 、and L ii When the combination of three amino acids of, in order, any amino acid, proline, and any amino acid, its PPR motif selectively binds to T; (2-33) A 1 、A 4 、and L ii When the combination of three amino acids of, in order, any amino acid, proline, and aspartic acid, its PPR motif selectively binds to T; (2-34) A 1 、A 4 、and L ii When the combination of three amino acids of, in order, phenylalanine, proline, and aspartic acid, its PPR motif selectively binds to T; (2-35) A 1 、A 4 、and L ii When the combination of three amino acids of, in order, tyrosine, proline, and aspartic acid, its PPR motif selectively binds to T; (2-36) A 1 、A 4 、and Lii When the combination of three amino acids is, in order, any amino acid, serine, and any amino acid, the PPR motif selectively binds to A and G; (2-37) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, serine, and asparagine, the PPR motif selectively binds to A; (2-38) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, phenylalanine, serine, and asparagine, the PPR motif selectively binds to A; (2-39) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, valine, serine, and asparagine, the PPR motif selectively binds to A; (2-40) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, threonine, and any amino acid, the PPR motif selectively binds to A and G; (2-41) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, threonine, and aspartic acid, the PPR motif selectively binds to G; (2-42) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, valine, threonine, and aspartic acid, the PPR motif selectively binds to G; (2-43) A 1 、A 4 、and L iiWhen the combination of three amino acids is, in order, any amino acid, threonine, asparagine, the PPR motif selectively binds to A; (2-44) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, phenylalanine, threonine, asparagine, the PPR motif selectively binds to A; (2-45) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, isoleucine, threonine, asparagine, the PPR motif selectively binds to A; (2-46) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, valine, threonine, asparagine, the PPR motif selectively binds to A; (2-47) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, valine, any amino acid, the PPR motif binds to A, C, and T but not to G; (2-48) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, isoleucine, valine, aspartic acid, the PPR motif selectively binds to C and then to A; (2-49) A 1 、A 4 、and L ii When the combination of three amino acids is, in order, any amino acid, valine, glycine, the PPR motif selectively binds to C; (2-50) A 1 、A 4 、and L iiWhen the combination of three amino acids is, in order, any amino acid, valine, and threonine, the PPR motif selectively binds to T.
[0023] (Novel PPR motif) The present invention provides a novel PPR motif. The novel PPR motifs provided by the present invention that are adenine-binding are as follows: (A-1), (A-2), and (A-3): (A-1) A PPR motif consisting of the sequence of SEQ ID NO: 9, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid is made in the sequence of SEQ ID NO: 9; (A-2) A PPR motif consisting of a sequence in which 1 to 20 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 9 are substituted, deleted, or added, and which is adenine-binding; (A-3) A PPR motif having at least 42% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are the same and which is adenine-binding.
[0024] (A-1) The substitutions may be 1, may be 2 or more, or may be all of the above.
[0025] (In (A-2), the 1 to 20 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, which are amino acids that can be substituted, etc. in the sequence of SEQ ID NO: 9, are Preferably, it is 1 to 11 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, and other than the amino acids at positions 5, 8, 13, 21, 22, 23, 25, 29, 35, More preferably, it is 1 to 7 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, and other than the amino acids at positions 5, 8, 13, 21, 22, 23, 25, 29, 35, and other than the amino acids at positions 20, 24, 31, and 32, Even more preferably, it is any one of the amino acids at positions 10, 15, 16, 17, 18, and 28.
[0026] (A-3) has at least 42% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are the same, Preferably, it has at least 71% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, and the amino acids at positions 5, 8, 13, 21, 22, 23, 25, 29, 35 are the same, More preferably, it has at least 80% sequence identity with the sequence of SEQ ID NO: 9, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34, the amino acids at positions 5, 8, 13, 21, 22, 23, 25, 29, 35, and the amino acids at positions 20, 24, 31, and 32 are the same, Even more preferably, it has at least 82% sequence identity with the sequence of SEQ ID NO: 9, provided that the non-identical amino acids are any one of the amino acids at positions 10, 15, 16, 17, 18, and 28.
[0027] The novel PPR motif provided by the present invention, which is cytosine-binding, is the following (C-1), (C-2), and (C-3): (C-1) A PPR motif consisting of the sequence of SEQ ID NO: 10, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the serine at position 2 of the sequence of SEQ ID NO: 10 with isoleucine, substitution of the amino acid at position 5 with isoleucine, substitution of the amino acid at position 7 with leucine, substitution of the amino acid at position 8 with lysine, substitution of the amino acid at position 10 with phenylalanine or tyrosine, substitution of the amino acid at position 15 with arginine, substitution of the amino acid at position 22 with valine, substitution of the amino acid at position 24 with arginine, substitution of the amino acid at position 27 with leucine, and substitution of the amino acid at position 29 with arginine is made; (C-2) A PPR motif consisting of a sequence in which 1 to 25 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 10 are substituted, deleted, or added, and which is cytosine-binding; (C-3) A PPR motif having at least 25% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are the same and which is cytosine-binding.
[0028] The substitution in (C-1) may be 1, may be 2 or more, or may be all of the above.
[0029] In (C-2), the 1 to 25 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34, which are amino acids that can be substituted, etc. in the sequence of SEQ ID NO: 10, are preferably 1 to 14 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 and also other than the amino acids at positions 6, 9, 11, 12, 17, 20, 21, 23, 25, 28, and 35; More preferably, there are 1 to 10 amino acids other than those at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34, and other than those at positions 6, 9, 11, 12, 17, 20, 21, 23, 25, 28, and 35, and other than those at positions 13, 16, 31, and 32, Even more preferably, it is any one of the amino acids at positions 2, 5, 7, 8, 10, 15, 22, 24, 27, and 29.
[0030] (C-3) has at least 25% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are identical, Preferably, it has at least 60% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34, and the amino acids at positions 6, 9, 11, 12, 17, 20, 21, 23, 25, 28, and 35 are identical, More preferably, it has at least 71% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34, the amino acids at positions 6, 9, 11, 12, 17, 20, 21, 23, 25, 28, and 35, and the amino acids at positions 13, 16, 31, and 32 are identical, Even more preferably, it has at least 71% sequence identity with the sequence of SEQ ID NO: 10, provided that the non-identical amino acids are any one of the amino acids at positions 2, 5, 7, 8, 10, 15, 22, 24, 27, and 29.
[0031] The novel PPR motif provided by the present invention, which is guanine-binding, is as follows: (G-1), (G-2), and (G-3): (G-1) A PPR motif consisting of the sequence of SEQ ID NO: 11, or a PPR motif consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the phenylalanine at position 10, substitution of the aspartic acid at position 15, substitution of the valine at position 27, substitution of the serine at position 28, and substitution of the isoleucine at position 35 is made in the sequence of SEQ ID NO: 11; (G-2) A PPR motif consisting of a sequence in which 1 to 21 amino acids other than those at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 11 are substituted, deleted, or added, and which is guanine-binding; (G-3) A PPR motif having at least 40% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are the same, and which is guanine-binding.
[0032] (G-1) The substitution in (G-1) may be 1, may be 2 or more, or may be all of the above.
[0033] (G-2) In (G-2), the 1 to 21 amino acids other than those at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34, which are amino acids that can be substituted, etc. in the sequence of SEQ ID NO: 11, are preferably 1 to 12 amino acids other than those at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 and other than those at positions 5, 11, 12, 17, 20, 21, 22, 23, and 25, more preferably 1 to 5 amino acids other than those at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34, other than those at positions 5, 11, 12, 17, 20, 21, 22, 23, and 25, and other than those at positions 8, 13, 16, 24, 29, 31, and 32, More preferably, it is any one of the amino acids at positions 10, 15, 27, 28, and 35.
[0034] (G-3) has at least 40% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are identical. Preferably, it has at least 65% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34, and the amino acids at positions 5, 11, 12, 17, 20, 21, 22, 23, and 25 are identical. More preferably, it has at least 85% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34, the amino acids at positions 5, 11, 12, 17, 20, 21, 22, 23, and 25, and the amino acids at positions 8, 13, 16, 24, 29, 31, and 32 are identical. Even more preferably, it has at least 85% sequence identity with the sequence of SEQ ID NO: 11, provided that the non-identical amino acids are any one of the amino acids at positions 10, 15, 27, 28, and 35.
[0035] The novel PPR motif provided by the present invention, which is uracil-binding, is the following (U-1), (U-2), and (U-3): (U-1) A PPR motif consisting of the sequence of SEQ ID NO: 12, or in the sequence of SEQ ID NO: 12, a substitution of the amino acid at position 10 with phenylalanine, a substitution of the amino acid at position 13 with serine, a substitution of the amino acid at position 15 with lysine, a substitution of the amino acid at position 17 with glutamic acid, a substitution of the amino acid at position 20 with leucine, a substitution of the amino acid at position 21 with lysine, a substitution of the amino acid at position 23 with phenylalanine, a substitution of the amino acid at position 24 with aspartic acid, a substitution of the amino acid at position 27 with lysine, a substitution of the amino acid at position 28 with lysine, a substitution of the amino acid at position 29 with arginine, and a substitution of the amino acid at position 31 with leucine; a PPR motif consisting of an amino acid sequence having any substitution selected from the group consisting of these substitutions. (U-2) A PPR motif consisting of a sequence in which 1 to 22 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 12 are substituted, deleted, or added, and which is uracil-binding. (U-3) A PPR motif having at least 37% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are identical, and which is uracil-binding.
[0036] The substitution in (U-1) may be 1, may be 2 or more, or may be all of the above.
[0037] In (U-2), the 1 to 22 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34, which are amino acids that can be substituted, etc. in the sequence of SEQ ID NO: 12, are preferably 1 to 14 amino acids other than the amino acids at positions 5, 7, 9, 16, 18, 22, 25, and 35 and other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34. More preferably, they are amino acids other than those at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34, and amino acids other than those at positions 5, 7, 9, 16, 18, 22, 25, and 35, and 1 to 12 of amino acids other than those at positions 8 and 32, Even more preferably, they are any of the amino acids at positions 10, 13, 15, 17, 20, 21, 23, 24, 27, 28, 29, and 31.
[0038] (U-3) has at least 37% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are identical, Preferably, it has at least 60% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34, and the amino acids at positions 5, 7, 9, 16, 18, 22, 25, and 35 are identical, More preferably, it has at least 65% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34, the amino acids at positions 5, 7, 9, 16, 18, 22, 25, and 35, and the amino acids at positions 8 and 32 are identical, Even more preferably, it has at least 65% sequence identity with the sequence of SEQ ID NO: 12, provided that the non-identical amino acids are any of the amino acids at positions 10, 13, 15, 17, 20, 21, 23, 24, 27, 28, 29, and 31.
[0039] In addition, the PPR motif v2 created by the present inventors A (SEQ ID NO: 9), v2 C (SEQ ID NO: 10), v2 G (SEQ ID NO: 11), v2 U (SEQ ID NO:12) is disclosed for the first time in the present application and does not exist in nature. For each of their homologs (among the embodiments shown as (A-1), (A-2), (A-3), (C-1), (C-2), (C-3), (G-1), (G-2), (G-3), (U-1), (U-2), (U-3) above and the embodiments shown as their preferred cases, the embodiments consisting of sequences other than SEQ ID NOs: 9-12), (regardless of whether each homolog is disclosed for the first time in the present application or not, and regardless of whether it exists in nature or not), combinations of two or more of at least any one of the homologs, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 are considered not to exist in nature. Regarding the present invention, when "any one of" is mentioned, the number selected is arbitrary.
[0040] (Explanation of the sequence of the novel PPR motif) In FIGS. 1 to 4, among the PPR motif sequences of Arabidopsis thaliana, the combinations of amino acids at positions 1, 4, and ii are VTN for the PPR motif that recognizes adenine, VSN for the PPR motif that recognizes cytosine, VTD for the PPR motif that recognizes guanine, and VND for the PPR motif that recognizes uracil, and the types and numbers of amino acids appearing at each position are summarized. The sequence v2 of the new PPR motif A (SEQ ID NO:9), v2 C (SEQ ID NO:10), v2 G (SEQ ID NO:11), v2 In U (SEQ ID NO:12), the amino acids at each position have a high frequency of appearance. In FIG. 6A, together with these novel sequences, in the sequence of the dPPR motif, v1 with the combinations of amino acids at positions 1, 4, and ii made the same as v2 A (SEQ ID NO:13), v1 C (SEQ ID NO:14), v1 G (SEQ ID NO:15), v1 U (SEQ ID NO:16) is also shown.
[0041] Figure 6A also shows the amino acid sequence of the v3.1 motif. V3.1 is obtained by introducing a D15K mutation into the adenine recognition motif in v2 (SEQ ID NO: 401), and other points are the same as those in v2. By using v3.1 in the PPR protein, a protein with improved binding affinity compared to v2 may be obtained.
[0042] In addition, Tables 1 to 4 below summarize how far the amino acid occurrence frequencies in each of the sequences of SEQ ID NOs: 9 - 12 deviate from those in the random case (for example, if 100 PPR motifs are collected, and the occurrence frequency of an amino acid at a certain position is randomly distributed, then each of the 20 types of amino acids would occur 5 times). At a certain position, if the amino acid occurrence frequency is far from that in the random case and the occurrence frequency is high, it is considered to be evolutionarily convergent, and the amino acid at that position is thought to be highly related to the function. For an amino acid highly related to the function, even if it is replaced with another amino acid with a high occurrence frequency that is also far from the random case, the function as a PPR motif can be maintained.
[0043] [Table 1]
[0044] [Table 2]
[0045] [Table 3]
[0046] [Table 4]
[0047] (Novel PPR Protein) The present invention provides a novel PPR protein containing a novel PPR motif. The novel PPR protein provided by the present invention is as follows. A protein containing n PPR motifs that can bind to a target RNA consisting of n base sequences, wherein the PPR motif for adenine in the base sequence is the PPR motif of (A-1), (A-2), or (A-3) described above; the PPR motif for cytosine in the base sequence is the PPR motif of (C-1), (C-2), or (c-3) described above; the PPR motif for guanine in the base sequence is the PPR motif of (G-1), (G-2), or (G-3) described above; and the PPR motif for uracil in the base sequence is the PPR motif of (U-1), (U-2), or (U-3) described above, a PPR protein.
[0048] Preferred examples of the PPR motif contained in the PPR protein are as follows: the descriptions of (A-1), (A-2), (A-3), (C-1), (C-2), (C-3), (G-1), (G-2), (G-3), (U-1), (U-2), or (U-3) regarding the PPR motif apply as they are.
[0049] In the PPR protein of the present invention, n (representing an integer of 1 or more) is not particularly limited, but can be 10 or more, preferably 12 or more, more preferably 15 or more, and even more preferably 18 or more. By increasing the number of motifs, a PPR protein having a high binding strength to many targets can be produced.
[0050] Conventionally, as shown in the following table, while the production of artificial PPR proteins consisting of 7 to 14 motifs has been reported, the construction of genes that contain many PPR motifs and inevitably have many repeats in the nucleotide sequence has been considered difficult. Also, generally, when producing a gene containing a repeat sequence, there may be cases where production is difficult, such as the repeat portion being rearranged during the cloning process (Trinh, T. et al. An Escherichia coli strain for the stable propagation of retroviral clones and direct repeat sequences. Focus, 16, 78 - 80(1994)). In the table, the Kd value represents the lowest value shown in each document.
[0051]
Table 5
[0052] When constructing a gene for a PPR protein having 15 or more PPR motifs, by utilizing the degeneracy of codons and making the nucleotide sequences encoding amino acids (excluding 1, 4, and ii involved in binding in each motif, and for the case of using the GoldenGate method described later, the 29 amino acids at positions 5 to 33 excluding the vicinity of both ends to be made common) different between each motif as appropriate, a gene with reduced repeats in the nucleotide sequence can be constructed. The degree of difference can be appropriately designed by those skilled in the art. For example, 4.5% or more (4 or more positions out of 87 bases), 15% or more, or 30% or more (26 or more positions out of 87 bases) of the bases can be made different.
[0053] For example, regarding the nucleotide sequences encoding the existing v1 - v4 motifs (SEQ ID NO:13 - 16), examples of nucleotide sequences encoding motifs utilizing the degeneracy of codons can be sequences as shown in the following table.
[0054]
Table 6
[0055] Note that "production" can be rephrased as "manufacture" or "fabrication". When parts are combined to produce genes or the like, it may be referred to as "construction", which can also be rephrased as "manufacture" or "fabrication".
[0056] (PPR motif, nucleic acid encoding PPR protein) The present invention provides a novel PPR motif and a nucleic acid encoding a novel PPR protein containing the same. The nucleotide sequence encoding the novel PPR motif has several variations due to codon degeneracy.
[0057] Amino acid sequence v2 of the novel PPR motif of the present invention A (SEQ ID NO:9), v2 C (SEQ ID NO:10), v2 G (SEQ ID NO:11), v2 Preferred examples of the nucleotide sequences encoding U (SEQ ID NO:12) are shown in the following table.
[0058] [Table 7-1]
[0059] In the dPPR motif, the amino acid sequences of the PPR motifs with the same combinations of amino acids at positions 1, 4, and ii as v2, v1 A (SEQ ID NO:13), v1 C (SEQ ID NO:14), v1 G (SEQ ID NO:15), v1 The nucleotide sequences encoding U (SEQ ID NO:16) are shown in the following table.
[0060] [Table 7-2]
[0061] The nucleotide sequence encoding the PPR protein can be composed of any combination of the above sequences. Amino acid sequence v2 A (SEQ ID NO:9), v2 C (SEQ ID NO:10), v2 G (SEQ ID NO:11), v2 The nucleotide sequence encoding U (SEQ ID NO:12), and v1 A (SEQ ID NO:13), v1 C (SEQ ID NO:14), v1 G (SEQ ID NO:15), v1 The nucleotide sequence encoding U (SEQ ID NO:16) may be appropriately combined to constitute the amino acids encoding the protein.
[0062] The amino acid sequence v3.1 of the novel PPR motif of the present invention A (SEQ ID NO:401), 1st A (SEQ ID NO:402), 1st C (SEQ ID NO:403), 1st G (SEQ ID NO:404), 1st Preferred examples of the nucleotide sequence encoding U (SEQ ID NO:405) are shown in the following table.
[0063]
Table 8
[0064] The nucleotide sequence encoding the PPR protein can be composed of any combination of the above sequences. As the nucleotide sequence encoding the first PPR motif from the N-terminus, the above v3.2 is used Any one selected from X is used, and as the nucleotide sequence encoding the subsequent PPR motif, as the nucleotide sequence encoding the PPR motif for adenine, the above v3.1 Select A and appropriately combine those selected from the above v2 series as the base sequences encoding PPR motifs for cytosine, guanine, and uracil.
[0065] (Improvement in Aggregation) The present inventors found that, from the amino acid information of existing naturally occurring PPR motifs, the amino acid at the 6th position of the PPR motif is often hydrophobic (especially leucine), and the amino acid at the 9th position is often a non-hydrophilic amino acid (especially glycine). From the structures of PPR proteins for which crystal structures have already been obtained (Non-Patent Document 6: Coquille et al., 2014 Nat. Commun.; PDB ID: 4PJQ, 4WN4, 4WSL, 4PJR; Non-Patent Document 7: Shen et al., 2015 Nat. Commun., PDB ID: 5I9D, 5I9F, 5I9G, 5I9H), since those at the 6th and 9th positions of the first motif (N-terminal side) are exposed to the outside, it was imagined that the exposed hydrophobic amino acids cause aggregation (Figure 6A). On the other hand, in the second and subsequent motifs, since the amino acids at the 6th and 9th positions are buried in the protein and form a hydrophobic core, it was considered that putting hydrophilic residues at the 6th and 9th positions of all motifs might collapse the protein structure. Therefore, it was decided to reduce the aggregation of PPR by making the amino acids at the 6th position, preferably the 6th and 9th positions, of only the first motif hydrophilic amino acids (asparagine, aspartic acid, glutamine, glutamic acid, lysine, arginine, serine, threonine).
[0066] Specifically, it is done as follows. In a protein capable of binding to a target nucleic acid having a specific base sequence, in the first PPR motif (M 1 ) counted from the N-terminus: (1) The A 6 amino acid is made into a hydrophilic amino acid, preferably the A 6 amino acid is made into asparagine or aspartic acid. (2) Further, the A 9The amino acid is a hydrophilic amino acid or glycine, preferably glutamine, glutamic acid, lysine, or glycine. (3) Or, A 6 Amino acid and A 9 The amino acid is any of the following combinations. · A 6 The amino acid is asparagine and A 9 The combination where the amino acid is glutamic acid · A 6 The amino acid is asparagine and A 9 The combination where the amino acid is glutamine · A 6 The amino acid is asparagine and A 9 The combination where the amino acid is lysine · A 6 The amino acid is aspartic acid and A 9 The combination where the amino acid is glycine
[0067] Among such PPR motifs, particularly preferred are the following. (1st A-1) In the sequence of SEQ ID NO: 402, the PPR motif in which the amino acids at positions 6 and 9 are substituted to satisfy any one of the combinations defined below; (1st (1st A-2) Consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of (1st A-1) are substituted, deleted, or added, and having adenine-binding property; (1st (1st A-3) Having at least 80% sequence identity with the sequence of (1st A-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and having adenine-binding property; (1st C-1) The PPR motif consisting of the sequence of SEQ ID NO: 403; (1st (1st A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of (C-1) are substituted, deleted, or added, and which is cytosine-binding; (1st C-3)(1st A PPR motif having at least 80% sequence identity with the sequence of (C-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and which is cytosine-binding; (1st G-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 404 are substituted so as to satisfy any one of the combinations defined below; (1st G-2)(1st A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of (G-1) are substituted, deleted, or added, and which is guanine-binding; (1st G-3)(1st A PPR motif having at least 80% sequence identity with the sequence of (G-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and which is guanine-binding; (1st U-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 405 are substituted so as to satisfy any one of the combinations defined below; (1st U-2)(1st A PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in the sequence of (U-1) are substituted, deleted, or added, and which is uracil-binding; (1st U-3)(1st A PPR motif having at least 80% sequence identity with the sequence of (U-1), provided that the amino acids at positions 1, 4, 6, 9, and 34 are the same, and which is uracil-binding. · A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamic acid · A combination where the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamine · A combination where the amino acid at position 6 is asparagine and the amino acid at position 9 is lysine · A combination where the amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine
[0068] In Figure 6A, the amino acid sequence of the v3.2 motif is shown together with the v3.1 motif. For V3.2, for the first motif, select 1st A (SEQ ID NO:402), 1st C (SEQ ID NO:403), 1st G (SEQ ID NO:404), 1st U (SEQ ID NO:405), and for the motifs after the second motif, select v2 C, v2 G, v2 U, v3.1 A. By using any of v3.2 as the first PPR motif from the N-terminus in the PPR protein, aggregation in the cell can be improved.
[0069] (Others) When the term "identity" is used in the present invention with respect to a nucleotide sequence (sometimes also referred to as a base sequence) or an amino acid sequence, unless otherwise specified, it means the percentage of the number of identical bases or amino acids shared between two sequences when the two sequences are aligned in an optimal manner. That is, identity = (the number of matching positions / the total number of positions) × 100, and it can be calculated using commercially available algorithms. Such algorithms are incorporated into the NBLAST and XBLAST programs described in Altschul et al., J. Mol. Biol. 215 (1990) 403-410. More specifically, searches and analyses regarding the identity of nucleotide sequences or amino acid sequences can be performed using algorithms or programs well-known to those skilled in the art (for example, BLASTN, BLASTP, BLASTX, ClustalW). Parameters when using a program can be appropriately set by those skilled in the art, or the default parameters of each program can also be used. Specific methods of these analysis methods are also well-known to those skilled in the art.
[0070] In this specification, when representing the identity with respect to a nucleotide sequence or an amino acid sequence in %, unless otherwise specified, in any case, a higher identity % value is preferred. Specifically, it is preferably 70% or more, more preferably 80% or more, still more preferably 85% or more, still more preferably 90% or more, still more preferably 95% or more, and still more preferably 97.5% or more.
[0071] Also, when the term "substituted, deleted, or added sequence" is used in the present invention with respect to a PPR motif or a protein, unless otherwise specified, the number of amino acids to be substituted, etc. in any motif or protein is not particularly limited as long as the motif or protein composed of the amino acid sequence has the desired function, but it is about 1 to 9 or 1 to 4, or if it is a substitution with amino acids having similar properties, there can be even more substitutions, etc. Means for preparing polynucleotides or proteins related to such amino acid sequences are well-known to those skilled in the art.
[0072] Amino acids with similar properties refer to amino acids with similar physical properties such as hydrophobicity, charge, pKa, solubility, etc. For example, it refers to the following. Hydrophobic (non-polar) amino acids; alanine, valine, glycine, isoleucine, leucine, phenylalanine, proline, tryptophan, tyrosine Non-hydrophobic amino acids; arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine, cysteine, histidine, methionine; Hydrophilic amino acids; arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine; Acidic amino acids: aspartic acid, glutamic acid; Basic amino acids: lysine, arginine, histidine; Neutral amino acids: alanine, asparagine, cysteine, glutamine, glycine, isoleucine, leucine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine; Sulfur-containing amino acids: methionine, cysteine; Aromatic ring-containing amino acids: tyrosine, tryptophan, phenylalanine. The PPR motif of the present invention, the protein containing it, or the nucleic acid encoding them can be prepared by those skilled in the art using the prior art.
[0073] [Performance of the novel PPR motif] [Binding ability] The PPR protein prepared using the novel PPR motif (SEQ ID NOs: 9-12) of the present invention is not only suitable for preparing PPR proteins for relatively long target RNAs, but may also have higher RNA binding performance than the PPR proteins prepared using the existing PPR (SEQ ID NOs: 13-16) motif for the same target RNA.
[0074] That is, by using the novel PPR motif of the present invention in the PPR protein, the binding ability to the target RNA can be enhanced as compared with the case of using the existing PPR motif. By increasing the binding ability, the efficiency of RNA manipulation in cells by the PPR protein can be improved. For example, the splicing efficiency in cells can be improved by using a PPR protein having a high binding ability to the target (see Example 5).
[0075] The degree to which the binding ability is enhanced is thought to also depend on the sequence and length of the target. For example, the binding ability can be made 1.1 times or more, specifically 1.3 times or more, 2.0 times or more, 3.0 times or more, 3.6 times or more.
[0076] The binding ability to the target sequence can be evaluated by methods using EMSA (Electrophoretic Mobility Shift Assay) or Biacore. EMSA is a method that utilizes the property that the mobility of nucleic acid molecules changes when a sample in which a protein and a nucleic acid are bound is electrophoresed, as compared with the case where they are not bound. Since molecular interaction analysis instruments typified by Biacore can perform kinetic analysis, detailed protein-nucleic acid binding analysis is possible.
[0077] The binding ability to the target sequence can also be evaluated by RPB-ELISA described later. In RPB-ELISA, the value obtained by subtracting the background signal (the luminescence signal value when the target PPR protein is added without adding the target RNA) from the luminescence amount of a sample to which the target PPR protein and its target RNA are added can be defined as the binding ability between the target PPR and its target RNA.
[0078] (Specificity) The PPR protein prepared using the novel PPR motif of the present invention may have a higher ability in terms of specificity for the target sequence than the PPR protein prepared using the existing PPR motif for the same target RNA.
[0079] That is, by using the novel PPR motif of the present invention in the PPR protein, the specificity for the target RNA can be enhanced as compared with the case of using the existing PPR motif. If a PPR protein has high specificity for the target RNA, when it is used to manipulate the target RNA in cells, it is possible to avoid unintended effects resulting from binding to unintended RNAs.
[0080] The affinity for the target sequence can be evaluated by a conventional method by those skilled in the art. Also, in RPB-ELISA, an appropriate non-target RNA is designed for the target PPR protein, and by similarly determining the binding force (luminescence signal value) in this case, the binding signal value for the target sequence / the binding signal value for the non-target sequence (S / N) can be determined as an index of the specificity (affinity) for the target RNA.
[0081] (Kd value) The PPR protein produced using the novel PPR motif of the present invention may have a high affinity (equilibrium dissociation constant, Kd value) for the target RNA.
[0082] The Kd value for the target sequence can be calculated by an existing method such as EMSA. When referring to the Kd value regarding the present invention, unless otherwise specified, it refers to the value measured by EMSA under the conditions described in the Examples section below.
[0083] The Kd value of the PPR protein produced using the novel PPR motif of the present invention seems to also depend on the target sequence and length, but when the length of the target sequence is 18 bases or longer, it can be -7 M or less, and can be -8 M or less, or on the order of -9 M. According to the studies of the present inventors, under the conditions of the Examples when the length of the target sequence is 18 bases, the minimum value (high affinity) of the Kd value is 1.95 x 10 -9Yes, it is lower than any of the Kd values of the reported design PPR proteins (see Table 1). It should be noted that the Kd value is known to correlate with the signal value obtained in the binding experiment by RPB-ELISA. When the luminescence value in RPB-ELISA (under the conditions described in the Examples section) is 1 - 2 x 10 7 in the case of, the Kd value is 10 -6 ~10 -7 M. When the luminescence value in RPB-ELISA is 2 - 4 x 10 7 in the case of, the Kd value is 10 -7 ~ 10 -8 M. When the luminescence value in RPB-ELISA is greater than 4 x 10 7 it can be estimated that the Kd value is ~10 -8 or less.
[0084] (Construction efficiency of PPR protein) By using the novel PPR motif of the present invention, a desired PPR protein can be efficiently constructed. The construction efficiency can be calculated by determining the ratio of the PPR proteins having a high Kd value that can be constructed using existing methods. Instead of the Kd value, the luminescence signal value by RPB-ELISA may be determined as described above and calculated in the same manner.
[0085] Specifically, by using the novel PPR motif of the present invention, when the length of the target sequence is 18 bases, PPR proteins with a Kd value of 10 -6 M or less (RPB-ELISA value of 1 x 10 7 or more) can be obtained with an efficiency of 50% or more, specifically 60% or more, more specifically 70% or more, and even more specifically 80% or more. Also, according to the present invention, PPR proteins with a target sequence of 18 bases in length and a Kd value of 10 -7 M or less (RPB-ELISA value of 2 x 10 7 or more) can be obtained with an efficiency of 50% or more, specifically 55% or more, more specifically 65% or more, and even more specifically 75% or more. Also, according to the present invention, PPR proteins with a Kd value of 10 -8 M or less (RPB-ELISA value of 4 x 10 7PPR proteins with a target sequence of 18 bases in length can be obtained with an efficiency of 20% or more, specifically 25% or more, more specifically 30% or more, and even more specifically 35% or more.
[0086] The construction efficiency can be calculated based on the binding signal value for the target sequence / the binding signal value for the non-target sequence (S / N) using the RPB-ELISA method.
[0087] Specifically, by using the novel PPR motif of the present invention, PPR proteins with a target sequence of 18 bases in length and an S / N higher than 10 can be obtained with an efficiency of 50% or more, specifically 55% or more, more specifically 65% or more, and even more specifically 75% or more. Further, according to the present invention, PPR proteins with a target sequence of 18 bases in length and an S / N higher than 100 can be obtained with an efficiency of 15% or more, specifically 20% or more, more specifically 25% or more, and even more specifically 30% or more.
[0088] [Seamless Cloning of PPR Protein Genes Using a Parts Library] The present invention also provides a method for producing a gene encoding a protein containing n PPR motifs capable of binding to a target nucleic acid consisting of n base sequences, comprising the following steps: Selecting m PPR parts necessary for producing the target gene from a library of at least 20×m types of PPR parts, each of which is inserted into each of at least 20 types of intermediate vectors Dest-a... designed such that at least 20 types of polynucleotides, including 4 types encoding each of the PPR motifs that are adenine, cytosine, guanine, or uracil or thymine-binding, and 16 types encoding each of the 2-linkers of the PPR motif, can be ligated in at least m types of sequences; Subjecting the selected m types of PPR parts to a Golden Gate reaction together with vector parts to obtain a vector into which a linker of m polynucleotides is inserted. n is an integer of m or more and 2×m or less. n can be, for example, 10 to 20.
[0089] The method of the present invention utilizes the Golden Gate reaction. In the Golden Gate reaction, multiple DNA fragments are inserted into a vector using a Type IIS restriction enzyme and T4 DNA Ligase. Since the Type IIS restriction enzyme cleaves outside the recognition sequence, the sticky ends can be freely set. In addition, it is highly efficient because the 4-base overhangs are used for ligation. Furthermore, no recognition sequence remains in the annealed and ligated construct. Therefore, the polynucleotide encoding the PPR motif can be seamlessly ligated (Figure 5). A particularly preferred example of the Type IIS restriction enzyme is BsaI.
[0090] The method of the present invention can efficiently generate genes even in the case of a large number of repeat sequences by using a parts library appropriately designed in consideration of the characteristics of the PPR protein and the Golden Gate reaction. Therefore, this method is beneficial when generating a gene for a protein containing 15 or more PPR motifs that can bind to a target nucleic acid with a length of 15 bases or more with an increasing number of repeat sequences. If m is set to 10, a library of 200 types of PPR parts is used, and 10 PPR parts necessary for generating the target gene are selected from the library, a gene encoding a protein containing 10 to 20 PPR motifs can be freely generated. In the following, although the case of generating a gene encoding an RNA-binding PPR protein with a target sequence of 10 to 20 bases in length may be described as an example, this method can also be applied to prepare PPR proteins for other lengths of target sequences and can also be applied to prepare DNA-binding PPR proteins.
[0091] In the method of the present invention, a parts library containing one or two arrays encoding the PPR motif is prepared (STEP1 and STEP2 in FIG. 5) and used. The parts library can be prepared, for example, by inserting the PPR motif array into 10 types of intermediate vectors Dest-a, b, c, d, e, f, g, h, i, j. The intermediate vectors are designed so that Dest-a to Dest-j are seamlessly concatenated in order by the Golden Gate reaction. The PPR motif arrays to be inserted can be at least 20 types, including 4 types (A, C, G, U) each encoding one, and 16 types (AA, AC, AG, AU, CA, CC, CG, CU, GA, GC, GG, GU, UA, UC, UG, UU) each encoding two linkers of the PPR motif. In this case, the parts library contains at least 200 types of parts.
[0092] Next, the necessary parts are selected according to the target base sequence. Specifically, for example, one part is selected from each of the parts libraries of Dest-a, b, c, d, e, f, g, h, i, j, and the Golden Gate reaction is performed together with the vector parts (STEP3 in FIG. 5). If 10 arrays with 1 motif in all intermediate vectors are selected, 10 sequences will be concatenated, and if those with 2 motifs are used, 20 sequences will be concatenated. When concatenating 11 to 19, one with 1 motif can be selected from any Dest-x library.
[0093] The vector parts used in STEP3 can be selected from 3 types of CAP-x vectors (see Non-Patent Document 1 above, taking into account the second amino acid of the PPR motif arranged on the most C-terminal side, and the second amino acid of the guanine-binding PPR motif and the uracil-binding PPR motif is the same). When the base sequence recognized by the motif located on the most C-terminal side is adenine, CAP-A can be used, when it is cytosine, CAP-C can be used, and when it is guanine or uracil, CAP-GU can be used respectively.
[0094] The obtained plasmid can be transformed into Escherichia coli and amplified and extracted.
[0095] [Method for detecting or analyzing PPR protein] The present invention provides a method for detecting or quantifying a protein containing n PPR motifs capable of binding to a target nucleic acid consisting of n base sequences, which comprises the following steps: A step of providing a solution containing a candidate protein to the immobilized target nucleic acid and detecting or quantifying the protein bound to the target nucleic acid. This detection and analysis method of the present invention is useful as a method for evaluating the binding performance of high-throughput PPR proteins.
[0096] Since the detection and analysis method of the present invention applies ELISA (Enzyme-Linked Immuno Sorbent Assay) (Fig. 7A), it may be referred to as the RPB-ELISA (RNA-protein binding ELISA) method. Although the method of the present invention is described in this specification as a method for evaluating RNA-binding PPR proteins, it can also be similarly applied to evaluate the binding performance of DNA-binding PPR proteins to target DNA.
[0097] The step of providing a solution containing a candidate protein to the immobilized target nucleic acid can be specifically carried out by flowing a solution containing the target binding protein through the target nucleic acid molecule immobilized on the plate. For the immobilization of the target nucleic acid molecule, various existing immobilization methods can be used. For example, it can be achieved by providing a nucleic acid probe containing a biotinylated target nucleic acid molecule to a well plate coated with streptavidin.
[0098] On the other hand, the candidate protein to be measured can be fused with a labeled protein, such as an enzyme like luciferase or a fluorescent protein. Fusion with the labeled protein makes detection and quantification easier.
[0099] The RPB-ELISA method has the advantage of not requiring a special device such as Biacore. Also, with the RPB-ELISA method, the throughput is high and the binding between protein and nucleic acid can be evaluated in a short period. Furthermore, under the conditions of the examples, the RPB-ELISA method can be sufficiently detected at a protein concentration of 6.25 nM or higher, and can also be detected in an E. coli lysate, so it has the advantage that it is not necessary to purify the target nucleic acid-binding protein.
[0100] [Use of PPR protein] (Complex, fusion protein) The PPR motif or PPR protein provided by the present invention can be linked with a functional region to form a complex. Also, it can be linked with a proteinaceous functional region to form a fusion protein. The functional region refers to a part having a specific biological function, such as an enzyme function, a catalytic function, an inhibitory function, a promoting function, etc. in vivo or in a cell, or a part having a function as a label. Such regions are composed of, for example, proteins, peptides, nucleic acids, bioactive substances, and drugs. Hereinafter, the present invention may be described by taking a fusion protein as an example with respect to a complex, but those skilled in the art can understand the case of a complex other than the fusion protein according to the description.
[0101] In one of the preferred embodiments, the functional region is ribonuclease (RNase). Examples of RNase are RNase A (for example, bovine pancreatic ribonuclease A: PDB 2AAS), RNase H.
[0102] In one preferred embodiment, the functional region is a fluorescent protein. Examples of fluorescent proteins are mCherry, EGFP, GFP, Sirius, EBFP, ECFP, mTurquoise, TagCFP, AmCyan, mTFP1, MidoriishiCyan, CFP, TurboGFP, AcGFP, TagGFP, Azami-Green, ZsGreen, EmGFP, HyPer, TagYFP, EYFP, Venus, YFP, PhiYFP, PhiYFP-m, TurboYFP, ZsYellow, mBanana, KusabiraOrange, mOrange, TurboRFP, DsRed-Express, DsRed2, TagRFP, DsRed-Monomer, AsRed2, mStrawberry, TurboFP602, mRFP1, JRed, KillerRed, HcRed, KeimaRed, mRasberry, mPlum, PS-CFP, Dendra2, Kaede, EosFP, KikumeGR. From the viewpoint of improving aggregation and / or efficient localization to the nucleus as a fusion protein, a preferred example is mClover3.
[0103] In one preferred embodiment, when the target is mRNA, the functional region is a functional domain that improves the protein expression level from the target mRNA (WO2017 / 209122). Examples of functional domains that improve the protein expression level from mRNA may be all or a functional part of the functional domains of proteins known to directly or indirectly promote mRNA translation. More specifically, it may be a domain that induces ribosomes to mRNA, a domain related to the initiation or promotion of mRNA translation, a domain related to the export of mRNA from the nucleus, a domain related to binding to the endoplasmic reticulum membrane, a domain containing an endoplasmic reticulum retention signal sequence, or a domain containing an endoplasmic reticulum signal sequence. Even more specifically, the above domain that induces ribosomes to mRNA may be a domain containing all or a functional part of a polypeptide selected from the group consisting of DENR (Density-regulated protein), MCT-1 (Malignant T-cell amplified sequence 1), TPT1 (Translationally-controlled tumor protein), and Lerepo4 (Zinc finger CCCH-domain). Also, the above domain related to the initiation or promotion of mRNA translation may be a domain containing all or a functional part of a polypeptide selected from the group consisting of eIF4E and eIF4G. Also, the above domain related to the export of mRNA from the nucleus may be a domain containing all or a functional part of SLBP (Stem-loop binding protein). Also, the above domain related to binding to the endoplasmic reticulum membrane may be a domain containing all or a functional part of a polypeptide selected from the group consisting of SEC61B, TRAP-alpha (Translocon associated protein alpha), SR-alpha, Dia1 (Cytochrome b5 reductase 3), and p180. Also, the above endoplasmic reticulum retention signal sequence may be a signal sequence containing the KDEL (KEEL) sequence.Further, the endoplasmic reticulum signal sequence may be a signal sequence including MGWSCIILFLVATATGAHS.
[0104] In the present invention, the functional region may be fused to the N-terminal side of the PPR protein, may be fused to the C-terminal side, or may be fused to both the N-terminal side and the C-terminal side. Further, the complex or fusion protein may include a plurality of functional regions (for example, 2 to 5). Furthermore, in the complex or fusion protein of the present invention, the functional region and the PPR protein may be indirectly fused via a linker or the like.
[0105] (Nucleic acid, vector, cell encoding PPR protein, etc.) The present invention also provides a nucleic acid encoding the above-described PPR motif, PPR protein, or fusion protein, and a vector containing the nucleic acid (for example, a vector for amplification, an expression vector). The vector for amplification can use Escherichia coli or yeast as a host. In the present specification, the expression vector means, for example, a vector containing DNA having a promoter sequence, DNA encoding a desired protein, and DNA having a terminator sequence from upstream, but it is not necessarily arranged in this order as long as it exhibits the desired function. In the present invention, various vectors that can be usually used by those skilled in the art can be recombinantly used.
[0106] The PPR protein or fusion protein of the present invention can function in cells of eukaryotes (e.g., animals, plants, microorganisms (yeasts, etc.), protists). The fusion protein of the present invention can particularly function in animal cells (in vitro or in vivo). Examples of animal cells into which the PPR protein, fusion protein, or vector expressing the same of the present invention can be introduced include cells derived from humans, monkeys, pigs, cows, horses, dogs, cats, mice, and rats. Examples of cultured cells into which the PPR protein, fusion protein, or vector expressing the same of the present invention can be introduced include, but are not limited to, Chinese hamster ovary (CHO) cells, COS-1 cells, COS-7 cells, VERO (ATCC CCL-81) cells, BHK cells, canine kidney-derived MDCK cells, hamster AV-12-664 cells, HeLa cells, WI38 cells, 293 cells, 293T cells, and PER.C6 cells.
[0107] (Use) The PPR protein or fusion protein of the present invention may be capable of delivering and functioning a functional region specifically to a nucleic acid sequence in vivo or in cells. A complex linked with a labeling moiety such as GFP can be used to visualize a desired RNA in vivo.
[0108] Furthermore, the PPR protein or fusion protein of the present invention can specifically modify and disrupt a nucleic acid sequence in cells or in vivo, and may be able to confer a new function. In particular, RNA-binding PPR proteins are involved in all steps of RNA processing, cleavage, RNA editing, translation, splicing, and RNA stabilization found in organelles. Therefore, the method related to the modification of the PPR protein provided by the present invention, and the PPR motif and PPR protein provided by the present invention can be expected to be used as follows in various fields.
[0109] (1) Medicine ·Produce a PPR protein that recognizes and binds to a specific RNA associated with a specific disease. Also, analyze the target sequence and the accompanying proteins for a specific RNA. The results of these analyses can be used to search for compounds for the treatment of diseases.
[0110] For example, in animals, it is known that an abnormality in a PPR protein identified as LRPPRC causes Leigh syndrom French Canadian (LSFC; Leigh syndrome, subacute necrotizing encephalomyelopathy). The present invention can contribute to the treatment (prevention, treatment, suppression of progression) of LSFC. Many existing PPR proteins function to specify the editing site of RNA manipulation (conversion of genetic information on RNA; often C→U). This type of PPR protein has an additional motif that is suggested to interact with an RNA editing enzyme at the C-terminal side. It is expected that a PPR protein having such a structure can introduce a base polymorphism or treat a disease or condition caused by a base polymorphism.
[0111] ·Produce cells that control RNA suppression / expression. Such cells include stem cells (e.g., iPS cells) that monitor the differentiated / undifferentiated state, model cells for the evaluation of cosmetics, and cells that can turn on / off the expression of functional RNA for the purpose of elucidating the mechanism of drug discovery and pharmacological tests.
[0112] ·Produce a PPR protein that specifically binds to a specific RNA associated with a specific disease. Introduce such a PPR protein into cells using a plasmid, viral vector, mRNA, or purified protein, and by binding of the PPR protein to its target RNA in the cell, the RNA function that is the cause of the disease can be changed (improved). Means of changing the function include, for example, changes in the RNA structure by binding, knockdown by degradation, changes in the splicing reaction by splicing, and base substitution.
[0113] (2) Agriculture, Forestry and Fisheries ·Improve the yield and quality of agricultural crops, forest products, fishery products, etc. · Breed organisms with improved disease resistance, environmental tolerance, or enhanced or new functionality.
[0114] For example, regarding first-generation hybrid (F1) crops, it may be possible to artificially create F1 crops by using the stabilization and translational control of mitochondrial RNA by PPR proteins, which could improve yield and quality. RNA manipulation and genome editing using PPR proteins enable the breeding and improvement of organisms (genetically modifying organisms) more accurately and quickly than conventional techniques. Also, RNA manipulation and genome editing using PPR proteins do not transform traits with foreign genes like genetic recombination but rather deal with the RNA and genomes originally possessed by animals and plants, making it similar to traditional breeding methods such as mutant selection and backcrossing. Therefore, it can surely and quickly respond to global-scale food and environmental problems.
[0115] (3) Chemistry · In the production of useful substances using microorganisms, cultured cells, plants, and animals (e.g., insects), control the protein expression level by manipulating DNA and RNA. This can improve the productivity of useful substances. Examples of useful substances include proteinaceous substances such as antibodies, vaccines, and enzymes, as well as relatively low-molecular-weight compounds such as pharmaceutical intermediates, fragrances, and pigments.
[0116] · Improve the production efficiency of biofuels by modifying the metabolic pathways of algae and microorganisms.
Example
[0117] [Example 1: Establishment of a method for producing the PPR gene] (Motif design) First, the PPR motif was designed. In the artificial PPR proteins reported so far, the consensus sequences of the PPR motif sequences existing in nature extracted by various methods are used. Among them, the PPR protein prepared using the motif sequence of dPPR (Non-Patent Documents 2, 3, and 6 cited above) has a low Kd value (high affinity). This PPR motif sequence is hereinafter referred to as the v1 PPR motif.
[0118] As another PPR motif sequence, a consensus sequence was generated using only the PPR motifs containing the representative first, fourth, and second amino acid combinations that recognize each base. Specifically, the representative amino acid combinations that recognize each base are: for the combination that recognizes adenine, the first is valine, the fourth is threonine, and the second is asparagine; for the combination that recognizes cytosine, the first is valine, the fourth is asparagine, and the second is serine; for the combination that recognizes guanine, the first is valine, the fourth is threonine, and the second is aspartic acid; for the combination that recognizes uracil, the first is valine, the fourth is asparagine, and the second is aspartic acid. Therefore, the consensus amino acid sequence was extracted from the PPR motif sequences containing the first, fourth, and second combinations, and this sequence was used as the PPR motif sequence that specifically recognizes adenine, cytosine, guanine, and uracil, respectively (Figs. 1 to 4, SEQ ID NOs: 9 - 12). This is hereinafter referred to as the v2 PPR motif. In the v1 PPR motif, the same first, fourth, and second amino acid combinations were also used (SEQ ID NOs: 13 - 16).
[0119] (Seamless cloning using 1 - motif and 2 - motif libraries) A cloning method for seamlessly linking these designed PPR motif arrays was constructed (Figure 5). Cloning is performed in three steps. In STEP1, the design and production of each motif array are carried out. In STEP2, a plasmid library in which one or two motifs are cloned is produced. In STEP3, the target PPR gene is completed by linking the required number of motifs.
[0120] First, a plasmid in which one PPR motif array (from the 4th to the ii-th) was cloned was produced (STEP1). The STEP1 plasmid contains PPR motif arrays that recognize A, C, G, and U respectively. In the subsequent STEP2, a DNA fragment containing the PPR motif array in the STEP1 plasmid is cloned into an intermediate vector (Dest-x, the sequence is as follows). The STEP1 plasmid into which only one motif can be inserted was named P1a-vx-X, the plasmid into which two motifs can be inserted was named P2a-vx-X on the N side and P2b-vx-X on the C side (vx is v1 or v2, X is A, C, G, U). For cloning into Dest-x, the BsaI restriction enzyme site (the BsaI restriction enzyme recognizes and cleaves the GGTCTCnXXXX sequence (SEQ ID NO:17), and the XXXX part becomes a 4-base overhang (hereinafter referred to as the tag sequence). And the base sequences for seamless ligation were designed as follows.
[0121] On the 5' side and 3' side of each motif array In the case of P1a, ggtctca atac (SEQ ID NO:18), gtgg tgagacc(SEQ ID NO:19), In the case of P2a, ggtctca atac (SEQ ID NO:18 above), gtggtca cata tgagacc(SEQ ID NO:20), In the case of P2b, ggtctca cata c(SEQ ID NO:21), gtgg tgagacc(SEQ ID NO:19 above) The base sequences with the respective arrays added were prepared by gene synthesis technology and cloned into pUC57-amp.
[0122] There are 10 types of Dest-x, namely Dest-a, b, c, d, e, f, g, h, i, and j. The base sequences were designed so that they are seamlessly linked in order from Dest-a to Dest-j. Dest-a is gaagacataaactccgtggtcacATACagagaccaaggtctcaGTGGtcacatacatgtcttc (SEQ ID NO:1), Dest-b is gaagacatATACagagaccaaggtctcaGTGGtgacataatgtcttc (SEQ ID NO:22), Dest-c is gaagacatcATACagagaccaaggtctcaGTGGttacatatgtcttc (SEQ ID NO:23), Dest-d is gaagacatacATACagagaccaaggtctcaGTGGttacaatgtcttc (SEQ ID NO:24), Dest-e is gaagacattacATACagagaccaaggtctcaGTGGtgacatgtcttc (SEQ ID NO:25), Dest-f is gaagacattgacATACagagaccaaggtctcaGTGGttaatgtcttc (SEQ ID NO:26), Dest-g is gaagacatgttacATACagagaccaaggtctcaGTGGtcatgtcttc (SEQ ID NO:27), Dest-h is gaagacatggtcacATACagagaccaaggtctcaGTGGtatgtcttc (SEQ ID NO:28), Dest-i is gaagacattggttacATACagagaccaaggtctcaGTGGatgtcttc (SEQ ID NO:29), Dest-j was prepared by gene synthesis technology and cloned into pUC57-kan. The sequence is gaagacatgtggtgacATACagagaccaaggtctcaGTGGtcttc (SEQ ID NO:30). It was cloned into pUC57-kan by gene synthesis technology.
[0123] For all Dest-x, constructs with PPR motifs corresponding to A, C, G, and U inserted, and constructs with two PPR motifs inserted to recognize each of the base combinations AA, AC, AG, AU, CA, CC, CG, CU, GA, GC, GG, GU, UA, UC, UG, and UU were prepared, and V1 and V2 STEP1 plasmid libraries (200 types each) were prepared. To achieve the above combinations, only 40 ng of the P1a plasmid, or 40 ng of the P2a plasmid, 40 ng of the P2b plasmid, 0.2 μL of 10x ligase buffer (NEB, B0202S), 0.1 μL of BsaI (NEB, R0535S), and 0.1 μL of Quick ligase (NEB, M2200S) were added, and the volume was adjusted to 1.9 μL with sterile water. A reaction of alternately repeating 37°C for 5 minutes and 16°C for 5 minutes 5 times was performed using a thermal cycler (Biorad, 1861096J1). Furthermore, 0.1 μL of 10x Cut smart buffer (NEB, B7204) and 0.1 μL of BsaI (NEB, R0535S) were added, and the reaction was carried out at 37°C for 60 minutes and 80°C for 10 minutes. 2.5 μL of the reaction solution was transformed into XL1-blue and selected on LB medium containing 30 μg / ml kanamycin. The insertion of the target sequence was confirmed by sequencing.
[0124] In STEP3, Dest-a to Dest-j are selected along the target base sequence and cloned into the CAP-x vector (Non-Patent Document 1 cited above). If all intermediate vectors contain 1 motif, 10 motifs will be ligated; if they contain 2 motifs, 20 motifs will be ligated. When ligating 11 to 19 motifs, it can be prepared by using the one with 1 motif in Dest-x at a desired position. For example, when preparing an 18-motif PPR sequence, plasmids with 1 motif in Dest-a and Dest-b and 2 motifs in the others are used.
[0125] The intermediate vectors used in the cloning of STEP3 need to be selected from three types of vectors. When the base sequence recognized by the motif located on the most C-terminal side is adenine, CAP-A is used; when it is cytosine, CAP-C is used; when it is guanine or uracil, CAP-GU is used respectively. By being cloned into the intermediate vector for STEP3, it is designed such that the amino acid sequence of MGNSV (SEQ ID NO:31) is added to the N-side of the PPR repeat and the amino acid sequence of ELTYNTLISGLGKAGRARDPPV (SEQ ID NO:32) is added to the C-side.
[0126] 20 ng each of 10 types of intermediate plasmids, 1 μL of 10x ligase buffer (NEB, B0202S), 0.5 μL of BpiL (Thermo, ER1012), and 0.5 μL of Quick ligase (NEB, M2200S) were added, and finally adjusted to 10 μl with sterilized water. The reaction was carried out for 15 cycles at 37°C for 5 minutes and 16°C for 7 minutes. Further, 0.4 μL of BpiL was added and reacted at 37°C for 30 minutes and at 75°C for 6 minutes. Subsequently, 0.3 μL of 1 mM ATP and 0.15 μL of Plasmid safe nuclease (Epicentre, E3110K) were added and reacted at 37°C for 15 minutes. 3.5 μl of the reaction solution was transformed into Escherichia coli (Competent cell of XL-1 Blue strain, Nippon Gene), and cultured in LB medium containing 100 μg / mL spectinomycin at 37°C for 16 hours for selection. A part of the grown colonies was pCR8 Forward: 5′-TTGATGCCTGGCAGTTCCCT -3′ (SEQ ID NO:33) and pCR8 Reverse: 5′-CGAACCGAACAGGCTTATGT -3′ (SEQ ID NO:34) primers were used to amplify the inserted gene region. 5 μL of 2 x Go-taq (Promega, M7123), 1.5 μL of 10 μM pCR8 Forward, 1.5 μL of 10 μM pCR8 Reverse, 2 μL of sterilized water were added to a 0.2 mL tube. After reacting at 98°C for 2 minutes in a thermal cycler, DNA amplification reaction was carried out with 15 cycles of 98°C for 5 seconds, 55°C for 10 seconds, and 72°C for 2.5 minutes. A part of the reaction solution was electrophoresed using MultiNA (SHIMADZU, MCE202) to confirm the size of the inserted DNA fragment. Three types of 18-motif PPR proteins (PPR1, PPR2, PPR3) (SEQ ID NOs: 35-37, 40-42) were each prepared with 3 clones using the v1 motif or v2 motif (v1 PPR1, v1 PPR2, v1 PPR3, v2 PPR1, v2 PPR2, v2 PPR3). The results are shown in Fig. 6B. In v1, correct-sized bands were obtained for all clones except the second clone of PPR2. In v2, correct-sized bands were obtained for all clones. Furthermore, their sequences were confirmed by sequencing. These results indicated that the PPR protein gene can be efficiently constructed by cloning with this method.
[0127] [Example 2: Construction of a high-throughput RNA-binding protein binding performance evaluation system] Generally, the evaluation of the binding between a nucleic acid-binding protein and a nucleic acid molecule is performed by methods using EMSA or Biacore. Electrophoretic Mobility Shift Assay (EMSA) is a method that utilizes the property that the mobility of a nucleic acid molecule changes when a sample in which a protein and a nucleic acid are bound is electrophoresed, as compared with the case where the nucleic acid molecule is not bound. This method has drawbacks such as the need for purified protein, complicated operations, and the inability to analyze many samples at once. Molecular interaction analysis instruments typified by Biacore enable detailed protein-nucleic acid binding analysis because kinetic analysis of the reaction is possible, but this also requires purified protein and a special apparatus. Therefore, a method with high throughput that can evaluate the binding between a protein and a nucleic acid in a short period was considered.
[0128] Enzyme-Linked Immuno Sorbent Assay (ELISA) is generally used when analyzing the binding between an antibody (protein) and a protein. In this method, a primary antibody is immobilized on a well plate, a solution containing the protein to be detected is added thereto, and after washing, a secondary antibody capable of colorimetric or luminescence detection is reacted to quantify the remaining amount of the protein to be analyzed. Applying this, a system was devised in which a nucleic acid molecule is immobilized on a plate, a solution containing the target nucleic acid-binding protein is flowed through, and the amount of the bound protein is quantified (Figure 7A). The method for immobilizing the nucleic acid molecule to be examined is performed by adding a nucleic acid probe modified with biotin at its end to a well plate coated with streptavidin. The nucleic acid-binding protein to be measured can be made easier to detect by fusing it with luciferase or a fluorescent protein. In addition, purification of the protein to be measured is not essential, and it is also possible to use a crude extract of cells (animal cultured cells, yeast, Escherichia coli, etc.) in which the nucleic acid-binding protein to be measured is expressed, and the time for purification can be shortened (Figure 7B). When measuring the binding between RNA and an RNA-binding protein by this method, it is described below as RPB-ELISA (RNA-protein binding ELISA).
[0129] To establish an experimental system, recombinant MS2 protein and its binding RNA probe were prepared. The gene of a protein with a luciferase protein fused to the N-terminal side of the MS2 protein and a 6x histidine tag fused to the C-terminal side was prepared by gene synthesis and cloned into the pET21b vector (NL MS2 HIS, SEQ ID NO:357). The RNA probe was synthesized by biotinylating the 5'-ends of a target sequence (RNA 4, SEQ ID NO:64) containing the MS2 binding sequence and a non-target sequence (RNA 51, SEQ ID NO:247) not containing it (Greiner). The MS2 protein expression plasmid was transformed into Escherichia coli strain Rosetta(DE3) and cultured overnight at 37 °C in 2 mL of LB medium containing 100 μg / mL ampicillin. Then, 2 mL of the culture solution was added to 300 mL of LB medium containing 100 μg / mL ampicillin, and OD 600It was cultured at 37°C until it reached 0.5 to 0.8. After the cultured medium was cooled to 15°C, IPTG was added to a final concentration of 0.1 mM and further cultured for 12 hours. The culture solution was centrifuged at 5000 x g, 4°C for 10 minutes to recover the cells, 5 mL of lysis buffer (20 mM Tris-HCl (pH 8.0), 150 mM NaCl, 0.5% NP-40, 1 mM DTT, 1 mM EDTA) was added, and after stirring with a vortex mixer, the cells were disrupted by sonication. It was centrifuged at 15,000 rpm, 4°C for 10 minutes, and the supernatant was recovered. Half of the supernatant was stored at -80°C until used as an E. coli lysate, and the remaining was subjected to affinity purification using a histidine tag and Ni-NTA. First, 200 μl of Ni-NTA agarose beads (Qiagen, Cat no. 30230) was spin-down to recover the beads. 100 μL of washing buffer was added thereto, and the beads were equilibrated by stirring with a rotator at 4°C for 1 hour. The total amount of the equilibrated beads was mixed with the protein solution and reacted at 4°C for 1 hour. Then, after centrifuging at 2,000 rpm for 2 minutes to recover the beads, factors that non-specifically bind to the beads were removed with 10 ml of washing buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40, 10 mM imidazole). Elution was performed with 60 μL of elution buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40, 500 mM imidazole). The purification degree was confirmed by SDS-PAGE. It was dialyzed overnight at 4°C with 20 mM Tris-HCl, pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM DTT, 1 mM EDTA).
[0130] The luminescence amounts of luciferase of the E. coli lysate and the purified MS2 protein solution were measured. Luminescence buffer (20 mM Tris -HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl 2, 0.5% NP-40, 1 mM DTT), 40 μL of the luciferase substrate (Promega, E151A) diluted 2,500-fold and 40 μL of the E. coli lysate or 40 μL of the purified MS2 protein solution were added to a 96-well white plate and reacted for 5 minutes, after which the luminescence was measured with a plate reader (PerkinElmer, 5103-35). From the obtained luminescence, it was diluted with a lysis buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl 8 , 0.02 x 10 8 , 0.09 x 10 8 , 0.38 x 10 8 , 1.50 x 10 8 , 6.00 x 10 8 LU / μL so that it became. 2 , 0.5% NP-40, 1 mM DTT, 0.1% BSA).
[0131] 2.5 pmol of the biotinylated RNA probe was added to a 96-well streptavidin-coated white plate (Thermo fisher, 15502), reacted at room temperature for 30 minutes, and washed with the lysis buffer. Wells were also prepared (-Probe) with the lysis buffer added instead of the biotinylated RNA for background measurement. Then, a blocking buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl 2 , 0.5% NP-40, 1 mM DTT, 1% BSA) was added, and the plate surface was blocked at room temperature for 30 minutes. Then, 100 μL of the E. coli lysate or the purified protein solution diluted above was added, and a binding reaction was carried out at room temperature for 30 minutes. Then, 200 μL of the washing buffer (20 mM Tris -HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl 2, 0.5% NP-40, 1 mM DTT) and washed five times. 40 μL of luciferase substrate (Promega, E151A) diluted 2,500-fold with the wash buffer was added to the wells and allowed to react for 5 minutes, after which the luminescence was measured using a plate reader (PerkinElmer, 5103-35).
[0132] From the luminescence of the samples to which the solutions containing each RNA and MS2 protein were added, the value obtained by subtracting the background (the luminescence signal value when the PPR protein was added without adding RNA) was taken as the binding affinity between the MS2 protein and RNA.
[0133] The results are shown in Fig. 7C. Specific binding between the MS2 protein and the target RNA (Target seq.) was detected in both the purified protein solution and the E. coli lysate. 6.0 x 10 8 LU / μL corresponds to 100 nM purified MS2 protein, so it was found that it could be sufficiently detected at a protein concentration of 6.25 nM (0.38 x 10 8 LU / μL) or higher. Furthermore, since it was also detectable in the E. coli lysate in the same way, it was found that there was no need to purify the protein.
[0134] [Example 3: Comparative experiment on RNA binding performance of 18-motif PPR proteins prepared using existing PPR motif sequences or novel PPR motif sequences] To evaluate the RNA binding performance of PPR proteins prepared using the v1 or v2 PPR motif, recombinant proteins were prepared in E. coli and their binding performance was evaluated using RPB-ELISA. For comparison, five types of target sequences (T 1, T 2, T 3, T 4, T 5, SEQ ID NOs: 46-50) were set, PPR proteins that bind to each were designed, and genes encoding each were prepared (v1 PPR1, v1 PPR2, v1 PPR3, v1 PPR4, v1 PPR5, v2 PPR1, v2 PPR2, v2 PPR3, v2 PPR4 v2 PPR5, SEQ ID NOs: 35 - 39, 40 - 45). A luciferase protein gene was added to the N-terminal side of the prepared PPR gene, and a histag sequence was added to the C-terminal side, followed by cloning into the pET21 vector (NL v1 PPR1, NL v1 PPR2, NL v1 PPR3, NL v1 PPR4, NL v1 PPR5, NL v2 PPR1, NL v2 PPR2, NL v2 PPR3, NL v2 PPR4, NL v2 PPR5, SEQ ID NOs: 51 - 60). The PPR expression plasmid was transformed into Rosetta(DE3) strain. This Escherichia coli was cultured in 2 mL of LB medium containing 100 μg / mL ampicillin at 37°C for 12 hours, and when the OD 600 reached 0.5 to 0.8, the culture solution was transferred to an incubator at 15°C and allowed to stand for 30 minutes. Then, 100 μL (final concentration 0.1 mM IPTG) was added, and the culture was carried out at 15°C for 16 hours. The Escherichia coli pellet was recovered by centrifugation at 5,000 x g, 4°C for 10 minutes, and 1.5 mL of lysis buffer (20 mM Tris-HCl, pH8.0, 150 mM NaCl, 0.5% NP-40, 1 mM MgCl 2, 2 mg / mL lysozyme, 1 mM PMSF, and 2 μL of 10 mg / mL DNase were added, and the mixture was frozen at -80°C for 20 minutes. Cell lysis was performed while infiltrating at 25°C for 30 minutes. Subsequently, centrifugation was carried out at 3,700 rpm, 4°C for 15 minutes to collect the supernatant (E. coli lysate) containing soluble PPR protein.
[0135] A 30-base sequence containing 18 bases of the target sequence was designed, and the 5'-end of the RNA probe (RNA 1, RNA 2, RNA 3, RNA 4, RNA 5, SEQ ID NOs: 61 - 65) was synthesized (Grainer). The 5'-end biotinylated RNA probe was added to a streptavidin-coated plate (Thermo fisher) and reacted at room temperature for 30 minutes, then washed with lysis buffer (20 mM Tris-HCl, pH 7.6, 150 mM NaCl, 5 mM MgCl 2 , 0.5% NP-40, 1 mM DTT, 0.1% BSA). For background measurement, wells were also prepared by adding 100 μL of lysis buffer, 1 μL of 100 mM DTT, and 1 μL of 40 unit / μL RNase inhibitor (Takara, 2313A) without adding RNA. Then, 200 μL of blocking buffer (20 mM Tris-HCl, pH 7.6, 150 mM NaCl, 5 mM MgCl 2 , 0.5% NP-40, 1 mM DTT, 1% BSA) was added, and the plate surface was blocked at room temperature for 30 minutes. Then, 100 μL of E. coli lysate containing luciferase-fused PPR protein with a luminescence of 1.5 x 10 8 LU / μL was added, and the binding reaction was carried out at room temperature for 30 minutes. Subsequently, 200 μL of washing buffer (20 mM Tris -HCl, pH 7.6, 150 mM NaCl, 5 mM MgCl 2, 0.5% NP-40, 1 mM DTT) was washed five times. 40 μL of a 2,500-fold diluted luciferase substrate (Promega, E151A) in the wash buffer was added to the wells and allowed to react for 5 minutes, after which the luminescence was measured using a plate reader (PerkinElmer, 5103-35). The binding affinity between PPR and RNA was defined as the value obtained by subtracting the background signal (the luminescence signal value when PPR protein was added without adding RNA) from the luminescence of the samples to which each RNA and PPR protein were added.
[0136] The results are shown in Fig. 8. When prepared with motif sequence v2, an increase in the binding affinity to the target sequence (1.3 - 3.6-fold) was observed in all cases compared to those prepared with motif sequence v1. Also, two types of RNA probes with non-target sequences (off target 1, off target 2) (SEQ ID NOs: 66, 69) were prepared, and as a result of examining the binding to them, the target binding signal / non-target binding signal (S / N) of v2 was higher than that of v1 in all cases (upper left of Fig. 8). From this, it was found that v2 has higher affinity and higher specificity for the target than v1.
[0137] [Example 4: Detailed analysis of RNA binding performance of PPR protein prepared using v2 motif] (Specificity evaluation) Using the v2 motif, PPR proteins against 23 types of target sequences (T 1 - 3, T 5 - T 24, SEQ ID NOs: 46 - 48, 51 - 69) were prepared (NL v2 PPR1 - 3, NL v2 PPR5 - 24, SEQ ID NOs: 56 - 58, 70 - 88), and all binding combinations were analyzed using RPB-ELISA. The experimental method was the same as in Example 3.
[0138] The results are shown in Fig. 9 (upper). v2 17, v2 Among 21 types of PPR proteins excluding 24, it was found that they had the strongest binding ability to the target. From these results, it was shown that PPR proteins prepared by utilizing the motif of v2 could be stably prepared and had high binding specificity.
[0139] Using the V3.1 motif, PPR proteins against the same 23 types of target sequences were similarly prepared (base sequences are SEQ ID NOs: 411 - 433, amino acid sequences are SEQ ID NOs: 434 - 456), and all binding combinations were analyzed using RPB - ELISA. The experimental method was the same as in Example 3.
[0140] The results are shown in Figure 9 (below) and the table below. Those with improved binding ability compared to V2 were obtained.
[0141]
Table 9
[0142] In addition, the RNA - binding performance of the PPR proteins shown in Figure 9 is shown in the table below in numerical values (log2 values).
[0143]
Table 10 - 1
[0144]
Table 10 - 2
[0145] (Affinity evaluation) Furthermore, to calculate the affinity (Kd value) of each PPR protein for the target RNA, EMSA was performed. Among the above 23 types, a streptavidin - binding peptide sequence was added to the N - terminal side and a 6x histag sequence was added to the C - terminal side of the gene sequences encoding 10 types of PPR proteins to construct an Escherichia coli expression plasmid (SBP v2 PPR1 HIS, SBP v2 PPR 2 HIS, SBP v2 PPR 3 HIS, SBP v2 PPR6 HIS, SBP v2 PPR9 HIS, SBP v2 PPR12 HIS, SBP v2 PPR15 HIS, SBP v2 PPR16 HIS, SBP v2 PPR20 HIS, SBP v2 PPR24 HIS, SEQ ID NOs: 89 - 97). It was transformed into E. coli strain Rosetta(DE3) and cultured overnight at 37°C in 2 mL of LB medium containing 100 μg / mL ampicillin. Then, 2 mL of the culture solution was transferred to 300 mL of LB medium containing 100 μg / mL ampicillin, and cultured at 37°C until the OD 600 reached 0.5 to 0.8. After the cultured medium was cooled to 15°C, 0.1 mM IPTG was added and further cultured for 12 hours. The culture solution was centrifuged at 5000 x g, 4°C for 10 minutes to collect the cells, 5 mL of lysis buffer (20 mM Tris - HCl, pH 8.0, 150 mM NaCl, 0.5% NP - 40, 1 mM DTT, and 1 mM EDTA) was added, stirred with a Voltex, and then sonicated to disrupt the cells. Centrifugation was performed at 15000 rpm, 4°C for 10 minutes, and the supernatant was collected.
[0146] Subsequently, the target protein was purified by affinity chromatography using the SBP tag. 100 μL of Streptavidine Sepharose High Performance (GE Healthcare, 17511301) was taken, and after collecting the beads by spin-down, it was equilibrated with a washing buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40). The equilibrated beads were gently mixed with the cell extract collected earlier and allowed to permeate at 4°C for 10 minutes. Subsequently, the entire volume of this bead solution was applied to the column, and the beads were washed with 10 mL of the washing buffer. Elution was performed with an elution buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 2 mM biotin).
[0147] Subsequently, affinity purification using the histidine tag was performed. First, 200 μL of Ni-NTA agarose (Qiagen, 30230) was collected, and after centrifugation, the beads were collected. 100 μL of the washing buffer was added, and the beads were equilibrated by allowing permeation at 4°C for 1 hour. The entire volume of the equilibrated beads was mixed with the protein solution eluted from the SBP beads and reacted at 4°C for 1 hour. Then, after centrifuging at 2,000 rpm for 2 minutes to collect the beads, factors that non-specifically bind to the beads were removed with 10 mL of the washing buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40, and 10 mM imidazole). Elution was performed with 60 μl of the elution buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.5% NP-40, and 500 mM imidazole). The degree of purification was confirmed by SDS-PAGE. Dialysis was performed overnight at 4°C with (20 mM Tris-HCl, pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM DTT, and 1 mM EDTA).
[0148] The Pierce 660nm Protein Assay Kit (Thermo fisher, 22662) was used to estimate the total amount of protein after dialysis. Next, to determine the amount of the target protein, first, the sample after dialysis was subjected to SDS-PAGE on a 10% polyacrylamide gel and stained with CBB. The gel image after staining was captured using a ChemiDoc Touch MP imaging system (Biorad). The total band intensity and that of the target band were derived from this gel image. The amount of the target protein was calculated by multiplying the ratio of the intensity of the target band to all band intensities by the total protein amount. Using this value, the molar concentration of the purified protein in the dialysis sample was calculated. Based on the molar concentration determined above, diluted protein solutions of 400 nM, 200 nM, 100 nM, 50 nM, 20 nM, 10 nM, 5 nM, 2 nM, and 1 nM were prepared. At that time, dilution was performed using a binding buffer (20 mM Tris-HCl, pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM DTT, 1 mM EDTA). An RNA probe (RNA 1. RNA 2. RNA 3. RNA 6. RNA 9. RNA 12. RNA 15. RNA 16. RNA 20. RNA 24) biotinylated on the 5'-side was adjusted to a final concentration of 20 nM with the binding buffer. The RNA sample that had been heat-treated at 75°C for 1 minute and then rapidly cooled was used in the following experiment.
[0149] The 20 nM RNA probe solution was mixed with the protein solutions at each concentration adjusted above, and a binding reaction was carried out at 25°C for 20 minutes.
[0150] After the reaction, 2 μL of 80% glycerol was added and well suspended, and then 10 μL of this was applied to an ATTO 7.5% gel and electrophoresis was carried out at a constant voltage of 150 V for 30 minutes.
[0151] The gel after electrophoresis was transferred to a Hybond N + membrane (GE, RPN203B). Subsequently, RNA was UV cross-linked to the membrane using an ASTEC Dual UV Transiluminator UVA-15 (Astec, 49909-06). The membrane was blocked with a blocking buffer (6.7 mM NaH 2 PO 4 ·2H 2 O, 6.7 mM Na 2 HPO 4 ·2H 2 O, 125 mM NaCl, 5% SDS). At this time, 0.5 μL of Stereptavidine-HRP (Abcam, ab7403) was pre-added to the blocking buffer, and the antigen-antibody reaction was carried out while infiltrating for 15 minutes. The blocking buffer was discarded, and 20 mL of a washing buffer (0.67 mM NaH 2 PO 4 ·2H 2 O, 0.67 mM Na 2 HPO 4 ·2H 2 O, 12.5 mM NaCl, 0.5% SDS) was added, and the membrane was washed. Subsequently, after repeating this washing operation 5 times, 20 ml of an equilibration buffer (100 mM Tris-HCl, pH 9.5, 100 mM NaCl, 10 mM MgCl 2 ) was added, and the membrane was infiltrated for 5 minutes. Then, Immunobilon Western Chemiluminisecnet HRP Substrate (Millipore, Cat No. WBKLS0100) was added to the membrane, and the biotinylated RNA was detected by chemiluminescence to detect the band. The gel image was captured with a ChemiDoc Touch MP imaging system (Biorad). The band intensity of the band of the unbound RNA probe and the band shifted by binding to the protein was calculated. The equilibrium dissociation constant (Kd value) was calculated by fitting to the Hill equation from the molar concentration of the protein and the ratio of the corresponding shifted band.
[0152] The results are shown in Fig. 10. The Kd value of the prepared PPR protein for the target was found to be 10 -9 - 10 -7 M. The minimum value (high affinity) was 1.95 x 10 -9 and it had the lowest Kd value among the reported designed PPR proteins so far (see Table 1). Also, these Kd values were found to correlate with the signal values obtained from the binding experiments in RPB-ELISA (R 2 = -0.85). From these results, when the luminescence value in RPB-ELISA is 1 - 2 x 10 7 , the Kd value is 10 -6 ~ 10 -7 M; when the luminescence value in RPB-ELISA is 2 - 4 x 10 7 , the Kd value is 10 -7 ~ 10 -8 M; when the luminescence value in RPB-ELISA is greater than 4 x 10 7 , it can be estimated that the Kd value is ~10 -8 or less.
[0153] (Evaluation of construction success rate) Furthermore, PPR proteins against 72 types of target sequences (T1 - T3, T6 - T76, SEQ ID NOs: 46 - 48, 51 - 69, 117 - 168) were prepared using the v2 motif (NL v2 PPR1 - 3, NL v2 PPR6 - 76, SEQ ID NOs: 56 - 58, 70 - 88, 169 - 220), and the construction success rate was calculated using RPB-ELISA. Biotinylated RNA probes containing the target sequences (RNA 1 - 3, RNA 6 - 76, SEQ ID NOs: 61 - 63, 98 - 116, 221 - 272) and non-target sequences (T A biotinylated RNA probe (RNA51, SEQ ID NO: 247) containing 51, SEQ ID NO: 143) was prepared (Greinar). The experimental method was the same as in Example 3. The results are shown in Fig. 11.
[0154] Among the 72 PPR proteins, those with a Kd value of 10 -6 M or less (RPB-ELISA value of 1 x 10 7 or more) were estimated to be 63 (88%), and those with a Kd value of 10 -7 M or less (RPB-ELISA value of 2 x 10 7 or more) were estimated to be 57 (79%), and those with a Kd value of 10 -8 M or less (RPB-ELISA value of 4 x 10 7 or more) were 43 (40%). Also, the value obtained by dividing the target binding signal by the non-target binding signal was defined as the specificity evaluation value (S / N). Among them, 54 (75%) had an S / N higher than 10, and 23 (32%) had an S / N higher than 100. From these results, it was shown that sequence-specific RNA-binding proteins can be efficiently produced by using the v2 motif to produce PPR proteins.
[0155] (Evaluation of target binding activity related to the number of PPR motifs) Analysis was performed on the number of PPR motifs and target binding activity. Thirteen types of 18-base target sequences were set, and 3 bases and 6 bases on the 5' side of each sequence were removed respectively to obtain 15-base (T 1a, T 49a, T 3a, T 14a, T 40a, T 12a, T 13a, T 2a, T 38a, T 37a, T 39a, T 56a, T 68a, SEQ ID NOs, 273, 275, 277, 279, 281, 283, 285, 287, 289, 291, 293, 395, 297), 12 bases (T 1b, T 49b, T 3b, T 14b, T 40b, T 12b, T 13b, T 2b, T 38b, T 37b, T 39b, T 56b, T 68b, SEQ ID NOs: 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298) target sequences were set. Corresponding PPR proteins (15 motifs were named PPRxa, 12 motifs were named PPRxb) were prepared (NL v2 PPR1, 1a, 1b; NL v2 PPR49, 49a, 49b; NL v2 PPR3, 3a, 3b; NL v2 PPR14, 14a, 14b; NL v2 PPR40, 40a, 40b; NL v2 PPR12, 12a, 12b; NL v2 PPR13, 13a, 13b; NL v2 PPR2, 2a, 2b; NL v2 PPR38, 38a, 38b; NL v2 PPR37, 37a, 37b; NL v2 PPR39, 39a, 39b; NL v2 PPR56, 56a, 56b; NL v2 PPR68, 68a, 68b; SEQ ID NOs: 56, 299 - 324). To perform analysis by RPB - ELISA, biotinylated RNA probes (T 1, T 49, T 3, T 14, T 40, T 12, T 13, T 2, T 38, T 37, T 39, T 56, T 68) and biotinylated RNA probes (RNA 51, SEQ ID NO: 143) containing non - target sequences were prepared, and the binding activities of on - target, off - target, and their respective PPR proteins were analyzed by RPB - ELISA. 51, SEQ ID NO: 247) were prepared, and the binding activities of on - target, off - target, and their respective PPR proteins were analyzed by RPB - ELISA.
[0156] The results for each target sequence are shown in Fig. 12A. The average values of the values for each of the 18 - motif, 15 - motif, and 12 - motif were plotted as box - and - whisker plots in Fig. 12B. It was found that the higher the number of motifs, the higher the binding strength, and when comparing the 18 - motif and 15 - motif, proteins with higher binding strength could be stably produced for the 18 - motif.
[0157] [Example 5: Artificial Splicing Control by PPR Protein] To demonstrate that the PPR protein binds to the target RNA molecule intracellularly and enables the desired RNA manipulation, an experiment using a splicing reporter was conducted (Figure 13A). The splicing reporter (RG6) has a gene structure consisting of exon 1, intron 1, exon 2, intron 2, exon 3, etc. (Orengo et al., 2006 NAR). In intron 1, exon 2, and intron 2, intron 4 and intron 5 of chicken cTNT and an artificially created alternative exon sequence are inserted. This reporter has two splicing forms, and the quantitative ratio of the mRNA with exon 2 skipped to the mRNA without skipping is approximately 1:1. Also, in exon 3, the RFP and GFP genes are encoded, but depending on the presence or absence of exon 2, the reading frame changes, so that RFP is expressed in the mRNA with exon 2 skipped, and GFP is expressed in the mRNA without skipping. It is known that the amount of the splicing form of this reporter is controlled by splicing factors that bind to the regions of intron 1, exon 2, and intron 2 (Orengo et al., 2006 NAR). Therefore, an experiment was conducted to determine whether the splicing form of the RG-6 reporter could be changed by PPR proteins that bind to 18-base sequences selected from the regions of intron 1, exon 2, and intron 2.
[0158] Seven types of target sequences (T77 - T83, SEQ ID NOs: 325 - 330) were selected from the RG6 reporter. The PPR protein genes were designed with both the v1 motif and the v2 motif (v1 PPRsp1 - 6, v2 PPRsp1 - 6, SEQ ID NOs: 331 - 342). The PPR protein genes were cloned into pcDNA3.1 so that a protein with a nuclear localization signal fused to the N-terminal side and a FLAG epitope tag sequence fused to the C-terminal side would be expressed (NLS v1PPRsp1 - 6, NLS v2PPRsp1-6, SEQ ID NOs: 343-354). pcDNA3.1 has a CMV promoter and an SV40 polyA signal (terminator), and the PPR protein gene was inserted between them.
[0159] HEK293T cells were seeded at 1 x 10 6 cells / well in a 10-cm dish containing 9 mL of DMEM and 1 mL of FBS. After culturing for 2 days at 37°C in a 5% CO 2 environment, the cells were harvested. The harvested cells were seeded at 4 x 10 4 cells / well into a PLL-coated 96-well plate and cultured for 1 day at 37°C in a 5% CO 2 environment. 100 ng of PPR expression plasmid DNA, 100 ng of RG-6, 0.6 μL of Fugene®-HD (Promega, E2311), and 200 μL of Opti-MEM were mixed and added to each well in full volume, and cultured for 2 days at 37°C in a 5% CO 2 environment. As a control, samples without adding PPR expression plasmid DNA were also prepared. After culturing, GFP fluorescence and RFP fluorescence images of each well were obtained using a fluorescence microscope DMi8 (Leica). The imaging conditions were first determined using a sample transfected with only the RG-6 plasmid to set the exposure time and gain such that the intensities of GFP and RFP were comparable, and then fluorescence images of each sample were obtained under the same conditions.
[0160] After image acquisition, total RNA was extracted using the Maxwell (registered trademark) RSC simplyRNA Cells Kit. 500 ng of the extracted total RNA, 0.5 μL of 100 μM dT20 primer, and 0.5 μL of 10 mM dNTPs were added to a 0.2 mL tube, incubated at 65°C for 5 minutes, and then immediately cooled on ice. To this, 2 μL of 5x RT-buffer (Invitrogen, 18080-051), 0.5 μL of 0.1 M DTT, 0.5 μL of 40 U / μL RNaseOUT (Invitrogen, 18080-051), and 0.5 μL of 200 unit / μL SupperScript III (Invitrogen, 18080-051) were added, reacted at 50°C for 50 minutes in a thermal cycler, then reacted at 85°C for 5 minutes, and cooled to 16°C. The reverse-transcribed sample was diluted 10-fold with sterile water. 2 μL of this, 10 μL of 5x GXL buffer (TAKARA, R050A), 4 μL of 2.5 mM dNTPs, 1.5 μL of 10 μM RT-Fw primer (5'-CAAAGTGGAGGACCCAGTACC-3') (SEQ ID NO:355), 1.5 μL of 10 μM RT-Rv primer (5'-GCGCATGAACTCCTTGATGAC-3') (SEQ ID NO:356), 1 μL of GXL (TAKARA, R050A), and 31.5 μL of sterile water were added to a 0.2 mL tube, reacted at 98°C for 2 minutes in a thermal cycler, then 98°C for 10 seconds, 58°C for 15 seconds, and 68°C for 5 seconds were repeated 35 times, and then cooled to 12°C. The reaction solution was diluted 10-fold and electrophoresed using MultiNA (SHIMADZU, MCE202). The band around 114 bp was taken as the band of exon-skipped RNA, and the band around 142 bp was taken as the band of non-skipped RNA, and the band intensity in each sample was calculated. The value obtained by dividing the 114 bp band intensity by the sum of the 114 bp band intensity and the 142 bp band intensity was defined as the splicing ratio.
[0161] The results are shown in FIGS. 13B and C. When only the RG6 reporter was introduced, the splicing ratio was 0.48. When PPRsp4 was introduced, it was about the same, but when other PPRs were introduced, it was found to change greatly. Comparing v1 and v2, except for PPRsp4, v2 changed more significantly. These splicing ratios also agreed with the RFP / GFP expression ratios in FIG. 13B. From these results, it was demonstrated that exon skipping can be changed by using PPR proteins, and it was also found that splicing can be changed more efficiently by using the v2 motif.
[0162] [Example 6: Control of Aggregation of PPR Protein] The PPR protein using the V2 motif (base sequence: SEQ ID NO: 457, amino acid sequence: SEQ ID NO: 458) and the PPR protein using the v3.2 motif (base sequence: SEQ ID NO: 459, amino acid sequence: SEQ ID NO: 460) were each prepared and purified in an E. coli expression system and separated by gel filtration chromatography.
[0163] (Expression and Purification of Protein) Using the pE-SUMOpro Kan plasmid containing the DNA sequence encoding the target PPR protein, transform the E. coli Rosetta strain. After culturing at 37°C, when the OD600 reaches 0.6, lower the temperature to 20°C and add IPTG to a final concentration of 0.5 mM to express the target PPR protein as a SUMO fusion protein in E. coli. After culturing overnight, collect the bacterial cells by centrifugation and resuspend them in Lysis Buffer (50 mM Tris-HCl pH 8.0, 500 mM NaCl). Disrupt E. coli by sonication. After centrifugation at 17,000 g for 30 min, apply the supernatant fraction to a Ni-Agarose column. Wash the column with Lysis Buffer containing 20 mM imidazole, and then elute the SUMO fusion target PPR protein with Lysis Buffer containing 400 mM imidazole. After elution, simultaneously cleave the SUMO protein from the target PPR protein with Ulp1 and replace the protein solution with ion exchange Buffer (50 mM Tris-Hcl pH 8.0, 200 mM NaCl) by dialysis. Then, perform cation exchange chromatography using an SP column. After applying the column, elute the protein by gradually increasing the NaCl concentration from 200 mM to 1 M. Final purification of the fraction containing the target PPR protein was performed by gel filtration chromatography using a Superdex200 column. Apply the target PPR protein eluted from ion exchange to a gel filtration column equilibrated with Gel Filtration Buffer (25 mM HEPES pH 7.5, 200 mM NaCl, 0.5 mM tris(2-carboxyethyl)phosphine (TCEP)). Finally, concentrate the fraction containing the target PPR protein, freeze it in liquid nitrogen, and store it at -80°C until use in the next analysis.
[0164] (Gel filtration chromatography) The purified recombinant PPR protein was adjusted to a concentration of 1 mg / ml. Gel filtration chromatography was performed using Superdex 200 increase 10 / 300 GL (GE Helthcare). The adjusted protein was applied to a gel filtration column equilibrated with 25 mM HEPES pH 7.5, 200 mM NaCl, 0.5 mM tris(2-carboxyethyl)phosphine (TCEP), and the properties of the protein were analyzed by measuring the absorbance at 280 nm of the solution eluting from the gel filtration column.
[0165] (Results) The results are shown in Figure 14. The smaller the elution fraction (Elution vol.), the larger the molecular size. At V2, elution occurred in the elution fraction of 8 to 10 mL, while at v3.2, a peak was observed in the elution fraction of 12 to 14 mL. From this, it was suggested that at v2, the protein size was large and there was a possibility of aggregation, and it was found that the aggregation was improved at v3.2.
Claims
1. A protein, or a fusion protein comprising the protein and a nuclear localization signal peptide A composition for controlling splicing, comprising: the protein is a protein comprising a polypeptide consisting of n PPR motifs capable of binding to a target RNA consisting of n base sequences, the PPR motif for adenine in the base sequence is a polypeptide of the following (A-1), (A-2), or (A-3); the PPR motif for cytosine in the base sequence is a polypeptide of the following (C-1), (C-2), or (c-3); the PPR motif for guanine in the base sequence is a polypeptide of the following (G-1), (G-2), or (G-3); the PPR motif for uracil in the base sequence is a polypeptide of the following (U-1), (U-2), or (U-3); at this time, an n is 15 or more, a PPR protein. (A-1) A polypeptide consisting of the sequence of SEQ ID NO: 9, or in the sequence of SEQ ID NO: 9, substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid A polypeptide consisting of an amino acid sequence in which any substitution selected from the group consisting of substitutions is made, or a polypeptide consisting of the sequence of SEQ ID NO: 401, or in the sequence of SEQ ID NO: 401, substitution of the amino acid at position 10 with tyrosine, substitution of the amino acid at position 16 with leucine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 18 with aspartic acid, and substitution of the amino acid at position 28 with glutamic acid A polypeptide consisting of an amino acid sequence in which any substitution selected from the group consisting of substitutions is made; (A-2) In the sequence of SEQ ID NO: 9 or 401, 1 to 4 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are substituted, deleted, or added, and a PPR motif represented by the following formula 1, which is adenine-binding, is a polypeptide; A polypeptide which is a PPR motif represented by the following formula 1, having at least 90% sequence identity with the sequence of SEQ ID NO: 9 or 401, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 11, 12, 14, 19, 26, 30, 33, and 34 are identical and it has adenine-binding property; A polypeptide consisting of the sequence of SEQ ID NO: 10, or a polypeptide consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the serine at position 2 with isoleucine, substitution of the isoleucine at position 5 with leucine, substitution of the leucine at position 7 with leucine, substitution of the lysine at position 8 with lysine, substitution of the phenylalanine or tyrosine at position 10 with phenylalanine or tyrosine, substitution of the arginine at position 15 with arginine, substitution of the valine at position 22 with valine, substitution of the arginine at position 24 with arginine, substitution of the leucine at position 27 with leucine, and substitution of the arginine at position 29 with arginine is made in the sequence of SEQ ID NO: 10; A polypeptide which is a PPR motif represented by the following formula 1, consisting of a sequence in which 1 to 4 amino acids other than the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 10 are substituted, deleted, or added, and having cytosine-binding property; A polypeptide which is a PPR motif represented by the following formula 1, having at least 90% sequence identity with the sequence of SEQ ID NO: 10, provided that the amino acids at positions 1, 3, 4, 14, 18, 19, 26, 30, 33, and 34 are identical and it has cytosine-binding property; A polypeptide consisting of the sequence of SEQ ID NO: 11, or a polypeptide consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the phenylalanine at position 10 with phenylalanine, substitution of the aspartic acid at position 15 with aspartic acid, substitution of the valine at position 27 with valine, substitution of the serine at position 28 with serine, and substitution of the isoleucine at position 35 with isoleucine is made in the sequence of SEQ ID NO: 11; A polypeptide which is a PPR motif represented by the following formula 1, consisting of a sequence in which 1 to 4 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 11 are substituted, deleted, or added, and having guanine-binding property; A polypeptide which is a PPR motif represented by the following formula 1, having at least 90% sequence identity with the sequence of SEQ ID NO: 11, provided that the amino acids at positions 1, 2, 3, 4, 6, 7, 9, 14, 18, 19, 26, 30, 33, and 34 are identical and is guanine-binding; A polypeptide consisting of the sequence of SEQ ID NO: 12, or a polypeptide consisting of an amino acid sequence in which any substitution selected from the group consisting of substitution of the amino acid at position 10 with phenylalanine, substitution of the amino acid at position 13 with serine, substitution of the amino acid at position 15 with lysine, substitution of the amino acid at position 17 with glutamic acid, substitution of the amino acid at position 20 with leucine, substitution of the amino acid at position 21 with lysine, substitution of the amino acid at position 23 with phenylalanine, substitution of the amino acid at position 24 with aspartic acid, substitution of the amino acid at position 27 with lysine, substitution of the amino acid at position 28 with lysine, substitution of the amino acid at position 29 with arginine, and substitution of the amino acid at position 31 with leucine is made in the sequence of SEQ ID NO: 12; A polypeptide which is a PPR motif represented by the following formula 1, consisting of a sequence in which 1 to 4 amino acids other than the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 in the sequence of SEQ ID NO: 12 are substituted, deleted, or added, and is uracil-binding; A polypeptide which is a PPR motif represented by the following formula 1, having at least 90% sequence identity with the sequence of SEQ ID NO: 12, provided that the amino acids at positions 1, 2, 3, 4, 6, 11, 12, 14, 19, 26, 30, 33, and 34 are identical and is uracil-binding. 【Chemical 1】 (In the formula: Helix A is a part capable of forming an α-helix structure 12 amino acids in length, represented by formula 2, 【Chemical 2】 In Formula 2, A 1 ~A 12 each independently represents an amino acid; X is a part that does not exist or consists of 1 to 9 amino acids in length; Helix B is a part capable of forming an α-helix structure consisting of 11 to 13 amino acids in length; L is a part represented by formula 3, 2 to 7 amino acids in length; [Chemical 3] In formula 3, each amino acid is numbered from the C-terminal side as "i" (-1), "ii" (-2), However, L iii ~L vii may not exist.)
2. A nucleic acid encoding the protein or fusion protein defined in claim 1, or A vector containing the nucleic acid A composition for controlling splicing, comprising.
Citation Information
Patent Citations
Method for modifying RNA binding protein using PPR motif
WO2011111829A1
Design method for RNA-binding protein using PPR motif, and use thereof
WO2013058404A1
DNA binding protein using PPR motif, and use thereof
WO2014175284A1
DNA-binding protein using PPR motif and use of said DNA-binding protein
WO2018030488A1
Proteins and their use for nucleotide binding
WO2019232588A1