PPR protein with low aggregation and uses thereof
Mutating the sixth and ninth amino acids in PPR proteins to hydrophilic residues addresses aggregation issues, ensuring stable and effective binding to target nucleic acids.
Patent Information
- Application Number
- JP2025158040
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-05-29
- Filing Date
- 2025-09-24
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2040-05-29
AI Technical Summary
Some PPR proteins produced by linking multiple motifs exhibit aggregation, particularly when expressed in cultured animal cells.
Mutate the sixth and ninth amino acids in the first motif of the PPR protein to hydrophilic amino acids to reduce aggregation.
The mutation significantly reduces protein aggregation while maintaining the ability to specifically bind to target nucleic acids, enhancing the stability and functionality of PPR proteins.
Smart Images

Figure 2025181985000013 
Figure 2025181985000014 
Figure 2025181985000015
Abstract
Description
[Technical Field]
[0001] The present invention relates to a nucleic acid manipulation technology that uses proteins capable of binding to targeted nucleic acids. The present invention is useful in a wide range of fields, including medicine (drug discovery support, disease treatment), agriculture (agricultural production, breeding), and chemistry (biological substance production). [Background technology]
[0002] PPR proteins are proteins that contain repeats of PPR motifs, each consisting of approximately 35 amino acids in length, and each PPR motif can specifically bind to one base. The combination of the first, fourth, and second amino acids (two amino acids before the next motif) in the PPR motif determines which of adenine, cytosine, guanine, or uracil (or thymine) it will bind to (Patent Documents 1 and 2).
[0003] Among naturally occurring RNA-binding PPR motifs, the most frequently occurring combinations corresponding to each base are: adenine, valine at position 1, threonine at position 4, and asparagine at position 2; cytosine, valine at position 1, asparagine at position 4, and serine at position 2; guanine, valine at position 1, threonine at position 4, and aspartic acid at position 2; and uracil, valine at position 1, asparagine at position 4, and aspartic acid at position 2 (Non-Patent Documents 1 to 5). Utilizing these amino acid combinations, it is possible to design PPR proteins that can specifically bind to any sequence. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] International Publication WO2013 / 058404 [Patent Document 2] International Publication WO2014 / 175284 [Patent Document 3] Patent application with serial number 195056K filed on the same day as this application [Non-patent literature]
[0005] [Non-Patent Document 1] Coquille, S. et al. An artificial PPR scaffold for programmable RNA recognition. Nature Communications 5, Article number: 5729(2014) [Non-patent document 2] Shen, C. et al. Specific RNA Recognition by Designer Pentatricopeptide Repeat Protein. Molecular Plant 8, 667-670(2015) [Non-patent document 3] Shen, C. et al. Structural basis for specific single-stranded RNA recognition by designer pentatricopeptide repeat proteins. Nature Communications volume 7, Article number: 11285 (2016) [Non-patent document 4] Miranda, RG et al. RNA-binding specificity landscapes of designer pentatricopeptide repeat proteins elucidate principles of PPR-RNA interactions. Nucleic Acids Research, 46(5), 2613-2623(2018) [Non-Patent Document 5] Yan, J. et al. Delineation of pentatricopeptide repeat codes for target RNA prediction. Nucleic Acids Research, gkz075(2019) Summary of the Invention [Problem to be solved by the invention]
[0006] The present inventors have investigated the use of the above amino acid combinations to produce PPR proteins with high performance and many, for example, 15 or more, PPR motifs linked together (Patent Document 3). However, the present inventors' investigations have revealed that some PPR proteins produced by this method exhibit aggregation. In particular, aggregation has sometimes been observed when PPR proteins are expressed in cultured animal cells. [Means for solving the problem]
[0007] Therefore, we investigated ways to solve this problem by mutating amino acids within the PPR motif, and found that the aggregation tendency of PPR can be improved by changing the sixth amino acid, or preferably the sixth and ninth amino acids, in the first motif (N-terminal side) of the PPR protein to hydrophilic amino acids, thereby completing the present invention.
[0008] The present invention provides the following: [1] One of the following PPR motifs: (C-1) a PPR motif consisting of any one of SEQ ID NOs: 4 to 7; (C-2) a cytosine-binding PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, and 34 in any one of SEQ ID NOs: 4 to 7 have been substituted, deleted, or added; (C-3) a PPR motif having at least 80% sequence identity with any one of SEQ ID NOs: 4 to 7, with the proviso that the amino acids at positions 1, 4, 6, and 34 are identical, and which is cytosine-binding; (A-1) A PPR motif consisting of the sequence of SEQ ID NO: 8, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (A-2) A PPR motif consisting of the sequence of (A-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which has adenine-binding properties; (A-3) a PPR motif having at least 80% sequence identity with the sequence of (A-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and having adenine-binding properties; (G-1) A PPR motif consisting of the sequence of SEQ ID NO: 9, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (G-2) A PPR motif consisting of the sequence of (G-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which has guanine-binding properties; (G-3) a PPR motif having at least 80% sequence identity with the sequence of (G-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and which is guanine-binding; (U-1) a PPR motif consisting of the sequence of SEQ ID NO: 10, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (U-2) a PPR motif consisting of the sequence of (U-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which is uracil-binding; (U-3) A PPR motif having at least 80% sequence identity with the sequence of (U-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and which is uracil-binding. [2] One of the following PPR motifs: (C-1) a PPR motif consisting of any one of SEQ ID NOs: 4 to 7; (C-2) a cytosine-binding PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in any one of SEQ ID NOs: 4 to 7 have been substituted, deleted, or added; (C-3) a PPR motif having at least 80% sequence identity with any one of SEQ ID NOs: 4 to 7, provided that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and having cytosine-binding properties; (A-1) A PPR motif in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 8 are substituted so as to satisfy any one of the combinations defined below; (A-2) A PPR motif consisting of the sequence of (A-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which has adenine-binding properties; (A-3) a PPR motif having at least 80% sequence identity with the sequence of (A-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and having adenine-binding properties; (G-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 9 have been substituted so as to satisfy any one of the combinations defined below; (G-2) A PPR motif consisting of the sequence of (G-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which has guanine-binding properties; (G-3) a PPR motif having at least 80% sequence identity with the sequence of (G-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and which is guanine-binding; (U-1) a PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 10 have been substituted so as to satisfy any one of the combinations defined below; (U-2) a PPR motif consisting of the sequence of (U-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which is uracil-binding; (U-3) A PPR motif having at least 80% sequence identity with the sequence of (U-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and which is uracil-binding. A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamic acid A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamine The amino acid at position 6 is asparagine and the amino acid at position 9 is lysine. The amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine [3] The PPR motif according to 1 or 2, which is one of the following: (C-4) a PPR motif consisting of the sequence of SEQ ID NO: 4; (A-4) a PPR motif consisting of the sequence of SEQ ID NO: 58; (G-4) a PPR motif consisting of the sequence of SEQ ID NO: 59; (U-4) PPR motif consisting of the sequence of SEQ ID NO: 60. [4] Use of the PPR motif according to any one of 1 to 3 as the first PPR motif from the N-terminus in a PPR protein. [5] The use according to 4 for reducing the aggregation of PPR proteins. [6] A protein capable of binding to a target nucleic acid having a specific base sequence, which contains 1 to 30 PPR motifs represented by the following formula 1, wherein the A6 amino acid of the first PPR motif (M1) from the N-terminus is a hydrophilic amino acid. [ka] (In the formula: Helix A is a 12 amino acid long portion capable of forming an α-helical structure and is represented by Formula 2: [ka] In formula 2, A1~A 12 each independently represents an amino acid; X is absent or a moiety consisting of 1 to 9 amino acids in length; Helix B is a portion consisting of 11 to 13 amino acids that can form an α-helical structure; L is a moiety of formula 3 that is 2 to 7 amino acids in length; [ka] In Formula 3, each amino acid is numbered from the C-terminus, as "i" (-1), "ii" (-2), etc. However, L iii ~L vii may not exist.) [7] The protein according to 6, wherein the A9 amino acid of M1 is a hydrophilic amino acid or glycine. [8] The protein according to 6 or 7, wherein the A6 amino acid of M1 is asparagine or aspartic acid. [9] The protein according to any one of items 6 to 8, wherein the A9 amino acid of M1 is glutamine, glutamic acid, lysine, or glycine.
[10] The protein according to any one of items 6 to 9, wherein the A6 amino acid of M1 and the A9 amino acid of M1 are any of the following combinations: A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamic acid A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamine A combination in which the A6 amino acid is asparagine and the A9 amino acid is lysine The A6 amino acid is aspartic acid and the A9 amino acid is glycine
[11] A fusion protein of at least one selected from the group consisting of a fluorescent protein, a nuclear localization signal peptide, and a tag protein, and a PPR protein containing the PPR motif described in any one of items 1 to 3 as the first PPR motif from the N-terminus, or a protein described in any one of items 6 to 10.
[12] A method for modifying a PPR protein capable of binding to a target nucleic acid having a specific base sequence, which contains a PPR motif as defined in 6, by making the A6 amino acid of the first PPR motif (M1) from the N-terminus more hydrophilic.
[13] A method for detecting nucleic acids, characterized by using a PPR protein containing the PPR motif described in any one of 1 to 3 as the first PPR motif from the N-terminus, a protein described in any one of 6 to 10, or a fusion protein described in 11.
[14] A nucleic acid encoding the PPR motif described in any one of 1 to 3, a PPR protein comprising the PPR motif described in any one of 1 to 3 as the first PPR motif from the N-terminus, or a protein described in any one of 6 to 10.
[15] A vector comprising the nucleic acid according to 14.
[16] A cell (excluding a human individual) containing the vector described in 15.
[17] A method for manipulating nucleic acids (excluding implementation in human individuals), using a PPR motif according to any one of 1 to 3, a PPR protein comprising a PPR motif according to any one of 1 to 3 as the first PPR motif from the N-terminus, a protein according to any one of 6 to 10, or a vector according to 15.
[18] A method for producing an organism, comprising the method of operation described in 17.
[0009] The present invention also provides the following: [1] One of the following PPR motifs: (C-1) a PPR motif consisting of any one of SEQ ID NOs: 4 to 7; (C-2) a cytosine-binding PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, and 34 in any one of SEQ ID NOs: 4 to 7 have been substituted, deleted, or added; (C-3) a PPR motif having at least 80% sequence identity with any one of SEQ ID NOs: 4 to 7, with the proviso that the amino acids at positions 1, 4, 6, and 34 are identical, and which is cytosine-binding; (A-1) A PPR motif consisting of the sequence of SEQ ID NO: 8, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (A-2) a PPR motif consisting of the sequence of (A-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which has adenine-binding properties; (A-3) a PPR motif having at least 80% sequence identity with the sequence of (A-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and having adenine-binding properties; (G-1) A PPR motif consisting of the sequence of SEQ ID NO: 9, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (G-2) A PPR motif consisting of the sequence of (G-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which has guanine-binding properties; (G-3) a PPR motif having at least 80% sequence identity with the sequence of (G-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and which is guanine-binding; (U-1) a PPR motif consisting of the sequence of SEQ ID NO: 10, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (U-2) a PPR motif consisting of the sequence of (U-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which is uracil-binding; (U-3) A PPR motif having at least 80% sequence identity with the sequence of (U-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and which is uracil-binding. [2] One of the following PPR motifs: (C-1) a PPR motif consisting of any one of SEQ ID NOs: 4 to 7; (C-2) a cytosine-binding PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in any one of SEQ ID NOs: 4 to 7 have been substituted, deleted, or added; (C-3) a PPR motif having at least 80% sequence identity with any one of SEQ ID NOs: 4 to 7, provided that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and having cytosine-binding properties; (A-1) A PPR motif in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 8 are substituted so as to satisfy any one of the combinations defined below; (A-2) A PPR motif consisting of the sequence of (A-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which has adenine-binding properties; (A-3) a PPR motif having at least 80% sequence identity with the sequence of (A-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and having adenine-binding properties; (G-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 9 have been substituted so as to satisfy any one of the combinations defined below; (G-2) A PPR motif consisting of the sequence of (G-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which has guanine-binding properties; (G-3) a PPR motif having at least 80% sequence identity with the sequence of (G-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and which is guanine-binding; (U-1) a PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 10 have been substituted so as to satisfy any one of the combinations defined below; (U-2) a PPR motif consisting of the sequence of (U-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which is uracil-binding; (U-3) A PPR motif having at least 80% sequence identity with the sequence of (U-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and which is uracil-binding. A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamic acid A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamine A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is lysine A combination in which the amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine [3] Use of the PPR motif according to 1 or 2 as the first PPR motif from the N-terminus in a PPR protein. [4] The use according to 3 for reducing the aggregation of PPR proteins. [5] A protein capable of binding to a target nucleic acid having a specific base sequence, which contains 1 to 30 PPR motifs represented by the following formula 1, wherein the A6 amino acid of the first PPR motif (M1) from the N-terminus is a hydrophilic amino acid. [ka] (In the formula: Helix A is a 12 amino acid long portion capable of forming an α-helical structure and is represented by Formula 2: [ka] In formula 2, A1~A 12 each independently represents an amino acid; X is absent or a moiety consisting of 1 to 9 amino acids in length; Helix B is a portion consisting of 11 to 13 amino acids that can form an α-helical structure; L is a moiety of formula 3 that is 2 to 7 amino acids in length; [ka] In Formula 3, each amino acid is numbered from the C-terminus, as "i" (-1), "ii" (-2), etc. However, L iii ~L vii may not exist.) [6] The protein according to 5, wherein the A9 amino acid of M1 is a hydrophilic amino acid or glycine. [7] The protein according to 5 or 6, wherein the A6 amino acid of M1 is asparagine or aspartic acid. [8] The protein according to any one of items 5 to 7, wherein the A9 amino acid of M1 is glutamine, glutamic acid, lysine, or glycine. [9] The protein according to any one of items 5 to 8, wherein the A6 amino acid of M1 and the A9 amino acid of M1 are any of the following combinations: A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamic acid A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamine A combination in which the A6 amino acid is asparagine and the A9 amino acid is lysine The A6 amino acid is aspartic acid and the A9 amino acid is glycine
[10] A fusion protein of at least one selected from the group consisting of a fluorescent protein, a nuclear localization signal peptide, and a tag protein with a PPR protein containing the PPR motif described in 1 or 2 as the first PPR motif from the N-terminus, or a protein described in any one of 5 to 9.
[11] A method for modifying a PPR protein capable of binding to a target nucleic acid having a specific base sequence, which contains the PPR motif defined in 3, by making the A6 amino acid of the first PPR motif (M1) from the N-terminus more hydrophilic.
[12] A method for detecting nucleic acids, characterized by using a PPR protein containing the PPR motif described in 1 or 2 as the first PPR motif from the N-terminus, a protein described in any one of 5 to 9, or a fusion protein described in 10.
[13] A nucleic acid encoding the PPR motif according to 1 or 2, a PPR protein comprising the PPR motif according to 1 or 2 as the first PPR motif from the N-terminus, or a protein according to any one of 5 to 9.
[14] A vector comprising the nucleic acid described in 13.
[15] A cell (excluding a human individual) containing the vector described in 14.
[16] A method for manipulating nucleic acids (excluding implementation in human individuals) using a PPR motif according to 1 or 2, a PPR protein containing the PPR motif according to 1 or 2 as the first PPR motif from the N-terminus, or a protein according to any one of 5 to 9, or a vector according to 14.
[17] A method for producing an organism, comprising the method of operation described in 16. [Brief explanation of the drawings]
[0010] [Figure 1] PPR motif design method. A: The sixth and ninth amino acids of the first motif are exposed. B: For the sixth and ninth amino acids of the first motif, which recognize cytosine, the typical combination was leucine and glycine (C6L9G), and the variants selected were leucine and glutamic acid (C6L9E), asparagine and glutamine (C6N9Q), asparagine and glutamic acid (C6N9E), asparagine and lysine (C6N9K), and aspartic acid and glycine (C6D9G). [Figure 2] Aggregation and nuclear localization of each PPR protein. Expression of PPRs fused to GFP and nuclear localization signal sequences was confirmed in cells using fluorescence microscopy. When fused to EGFP, PPRcag1 (6L9G) and PPRcag2 (6L9E) did not localize to the nucleus but were strongly aggregated around the nucleus. On the other hand, PPRcag3 (6N9Q), PPRcag4 (6N9E), PPRcag5 (6N9K), and PPRcag6 (6D9G) showed low aggregation but did not localize to the nucleus. When fused to mClover3, PPRcag1 (6L9G) and PPRcag2 (6L9E) localized to the nucleus but were aggregated within the nucleus. PPRcag3 (6N9Q), PPRcag4(6N9E), PPRcag5(6N9K), and PPRcag6(6D9G) localized to the nucleus and did not exhibit aggregation. [Figure 3]Binding experiments between PPR proteins and RNA. All PPRs, including those with mutations at positions 6 and 9, were found to specifically bind to their target, CAGx6. Compared to PPRcag1, PPRcag2 had similar binding avidity to the target sequence, while PPRcag3 was approximately 80%, PPRcag4 was approximately 60%, PPRcag5 was approximately 120%, and PPRcag6 was approximately 130%. [Figure 4] The effect of the first PPR motif from the N-terminus on aggregation. Each PPR protein was produced in an E. coli expression system, purified, and separated by gel filtration chromatography. The smaller the elution volume, the larger the molecular size. V2 eluted in the 8-10 mL elution fraction, while v3.2 showed a peak in the 12-14 mL elution fraction. This suggests that v2 may have aggregated due to its larger protein size, and that aggregation was improved in v3.2. DETAILED DESCRIPTION OF THE INVENTION
[0011] [PPR motif, PPR protein] (definition) Unless otherwise specified, the term "PPR motif" as used herein refers to a polypeptide consisting of 30 to 38 amino acids, whose amino acid sequence, when analyzed using an online protein domain search program, yields an E value of PF01535 in Pfam and PS51375 in Prosite that is equal to or less than a predetermined value (preferably E-03). The position numbers of the amino acids constituting the PPR motif as defined herein are essentially the same as those for PF01535, but correspond to the position of the amino acid in PS51375 minus two (e.g., position 1 in the present invention is replaced by position 3 in PS51375). However, when referring to the amino acid at position "ii" (-2), it refers to the second amino acid from the end (C-terminal) of the amino acid constituting the PPR motif, or the amino acid two amino acids N-terminal to amino acid 1 of the next PPR motif, i.e., the -2 amino acid. If the next PPR motif is not clearly identified, the amino acid two amino acids before the first amino acid in the next helix structure is designated as "ii." For more information on Pfam, see http: / / pfam.sanger.ac.uk / . For more information on Prosite, see http: / / www.expasy.org / prosite / .
[0012] The conserved amino acid sequence of the PPR motif is not highly conserved at the amino acid level, but the two α-helices in the secondary structure are well conserved. A typical PPR motif consists of 35 amino acids, but its length varies from 30 to 38 amino acids.
[0013] More specifically, the PPR motif in the present invention consists of a polypeptide of 30 to 38 amino acids in length, represented by formula 1.
[0014] [ka] During the ceremony: Helix A is a 12 amino acid long portion capable of forming an α-helical structure and is represented by Formula 2:
[0015] [ka] In formula 2, A1~A 12 each independently represents an amino acid; X is absent or a moiety consisting of 1 to 9 amino acids in length; Helix B is a portion consisting of 11 to 13 amino acids that can form an α-helical structure; L is a moiety of formula 3 that is 2 to 7 amino acids in length;
[0016] [ka] In Formula 3, each amino acid is numbered from the C-terminus, as "i" (-1), "ii" (-2), etc. However, L iii ~L vii may not exist.
[0017] In the present invention, the term "PPR protein" refers to a PPR protein having one or more, preferably two or more, of the above-mentioned PPR motifs, unless otherwise specified. In this specification, the term "protein" refers to any substance consisting of a polypeptide (a chain of multiple amino acids peptide-bonded together), unless otherwise specified, and includes those consisting of relatively low-molecular-weight polypeptides. In the present invention, the term "amino acid" may refer to a normal amino acid molecule, as well as to amino acid residues that constitute a peptide chain. It will be clear to those skilled in the art from the context which is being referred to.
[0018] In the present invention, when referring to the binding ability of a PPR motif to a base in a target nucleic acid, the term "specific" means that the binding activity to any one of the four types of bases is higher than the binding activity to the other bases, unless otherwise specified.
[0019] In the present invention, the term nucleic acid refers to RNA or DNA. Note that PPR proteins may have specificity for bases in RNA or DNA, but do not bind to nucleic acid monomers.
[0020] In the PPR motif, the combination of three amino acids, 1, 4, and ii, is important for specific binding to a base, and this combination determines which base will bind (Patent Documents 1 and 2 cited above).
[0021] Specifically, for the RNA-binding PPR motif, the relationship between the combination of three amino acids 1, 4, and ii and the bases that can bind is as follows (see Patent Document 1 cited above). (3-1) A1, A4, and L ii When the combination of these three amino acids is, in order, valine, asparagine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to U, then C, and then A or G. (3-2) A1, A4, and L ii When the three amino acids in the PPR motif are, in order, valine, threonine, and asparagine, the PPR motif binds strongly to A, then G, then C, but not to U, showing selective RNA base binding ability. (3-3) A1, A4, and L ii When the combination of these three amino acids is, in order, valine, asparagine, asparagine, the PPR motif has selective RNA base-binding ability, binding strongly to C, then to A or U, but not to G. (3-4) A1, A4, and L ii When the combination of these three amino acids is, in order, glutamic acid, glycine, and aspartic acid, the PPR motif has selective RNA base binding ability, binding strongly to G but not to A, U, or C. (3-5) A1, A4, and L ii When the three amino acid combinations are isoleucine, asparagine, and asparagine, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to C, then U, then A, but not to G. (3-6) A1, A4, and L iiWhen the three amino acids in the PPR motif are, in order, valine, threonine, and aspartic acid, the PPR motif binds strongly to G, then to U, but does not bind to A or C, showing selective RNA base binding ability. (3-7) A1, A4, and L ii When the three amino acid combinations are lysine, threonine, and aspartic acid, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to G, then to A, but not to U or C. (3-8) A1, A4, and L ii When the combination of these three amino acids is, in order, phenylalanine, serine, and asparagine, the PPR motif has selective RNA base-binding ability, binding strongly to A, then C, and then G and U. (3-9) A1, A4, and L ii When the three amino acid combinations are valine, asparagine, and serine, respectively, the PPR motif has selective RNA base-binding ability, binding strongly to C, then to U, but not to A or G. (3-10) A1, A4, and L ii When the combination of these three amino acids is, in order, phenylalanine, threonine, and asparagine, the PPR motif has selective RNA base-binding ability, binding strongly to A but not to G, U, or C. (3-11) A1, A4, and L ii When the combination of these three amino acids is, in order, isoleucine, asparagine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to U, then to A, but not to G or C. (3-12) A1, A4, and L ii When the combination of these three amino acids is threonine, threonine, and asparagine, respectively, the PPR motif has selective RNA base binding ability, binding strongly to A but not to G, U, or C. (3-13) A1, A4, and Lii When the three amino acids in the PPR motif are, in order, isoleucine, methionine, and aspartic acid, the PPR motif has selective RNA base-binding ability, binding strongly to U, then C, but not A or G. (3-14) A1, A4, and L ii When the three amino acid combinations are phenylalanine, proline, and aspartic acid, respectively, the motif is called PPR. It binds strongly to U, then C, but not A or G, and has selective RNA base binding ability. (3-15) A1, A4, and L ii When the combination of these three amino acids is tyrosine, proline, and aspartic acid, respectively, the PPR motif has selective RNA base binding ability, binding strongly to U but not to A, G, or C. (3-16) A1, A4, and L ii When the combination of these three amino acids is leucine, threonine, and aspartic acid, respectively, the PPR motif has selective RNA base binding ability, binding strongly to G but not to A, U, or C.
[0022] Specifically, with regard to the DNA-binding PPR motif, the relationship between the combination of three amino acids 1, 4, and ii and the bases that can bind is as follows (see Patent Document 2 cited above). (2-1) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, glycine, and aspartic acid, the PPR motif selectively binds to G; (2-2) A1, A4, and L ii When the three amino acid combinations are, in order, glutamic acid, glycine, and aspartic acid, the PPR motif selectively binds to G; (2-3) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, glycine, and asparagine, the PPR motif selectively binds to A; (2-4) A1, A4, and L ii When the combination of the three amino acids is, in order, glutamic acid, glycine, and asparagine, the PPR motif selectively binds to A; (2-5) A1, A4, and L ii When the combination of three amino acids is, in order, any amino acid, glycine, and serine, the PPR motif selectively binds to A and then to C; (2-6) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, isoleucine, any amino acid, the PPR motif selectively binds to T and C; (2-7) A1, A4, and L ii When the three amino acid combinations are, in order, any amino acid, isoleucine, and asparagine, the PPR motif selectively binds to T and then to C; (2-8) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, leucine, any amino acid, the PPR motif selectively binds to T and C; (2-9) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, leucine, and aspartic acid, the PPR motif selectively binds to C; (2-10) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, leucine, and lysine, the PPR motif selectively binds to T; (2-11) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, methionine, any amino acid, the PPR motif selectively binds to T; (2-12) A1, A4, and L ii When the three amino acid combinations are, in order, any amino acid, methionine, and aspartic acid, the PPR motif selectively binds to T; (2-13) A1, A4, and L ii When the three amino acid combinations are, in order, isoleucine, methionine, and aspartic acid, the PPR motif selectively binds to T and then C; (2-14) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, asparagine, any amino acid, the PPR motif selectively binds to C and T; (2-15) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-16) A1, A4, and L ii When the three amino acid combinations are, in order, phenylalanine, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-17) A1, A4, and L ii When the combination of the three amino acids is, in order, glycine, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-18) A1, A4, and L ii When the three amino acid combinations are, in order, isoleucine, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-19) A1, A4, and L ii When the three amino acid combinations are threonine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-20) A1, A4, and L ii When the three amino acid combinations are, in order, valine, asparagine, and aspartic acid, the PPR motif selectively binds to T and then to C; (2-21) A1, A4, and L iiWhen the three amino acid combinations are tyrosine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T and then C; (2-22) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, asparagine, asparagine, the PPR motif selectively binds to C; (2-23) A1, A4, and L ii When the combination of the three amino acids is, in order, isoleucine, asparagine, asparagine, the PPR motif selectively binds to C; (2-24) A1, A4, and L ii When the combination of these three amino acids is serine, asparagine, asparagine, in that order, the PPR motif selectively binds to C; (2-25) A1, A4, and L ii When the combination of these three amino acids is, in order, valine, asparagine, asparagine, the PPR motif selectively binds to C; (2-26) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, asparagine, and serine, the PPR motif selectively binds to C; (2-27) A1, A4, and L ii When the combination of the three amino acids is, in order, valine, asparagine, and serine, the PPR motif selectively binds to C; (2-28) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, asparagine, and threonine, the PPR motif selectively binds to C; (2-29) A1, A4, and L ii When the three amino acid combinations are, in order, valine, asparagine, and threonine, the PPR motif selectively binds to C; (2-30) A1, A4, and L iiWhen the three amino acid combinations are, in order, any amino acid, asparagine, and tryptophan, the PPR motif selectively binds to C and then to T; (2-31) A1, A4, and L ii When the three amino acid combinations are, in order, isoleucine, asparagine, and tryptophan, the PPR motif selectively binds to T and then C; (2-32) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, proline, any amino acid, the PPR motif selectively binds to T; (2-33) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, proline, and aspartic acid, the PPR motif selectively binds to T; (2-34) A1, A4, and L ii When the three amino acid combinations are, in order, phenylalanine, proline, and aspartic acid, the PPR motif selectively binds to T; (2-35) A1, A4, and L ii When the three amino acid combinations are tyrosine, proline, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-36) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, serine, any amino acid, the PPR motif selectively binds to A and G; (2-37) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, serine, and asparagine, the PPR motif selectively binds to A; (2-38) A1, A4, and L ii When the combination of the three amino acids is, in order, phenylalanine, serine, and asparagine, the PPR motif selectively binds to A; (2-39) A1, A4, and L iiWhen the combination of the three amino acids is, in order, valine, serine, and asparagine, the PPR motif selectively binds to A; (2-40) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, threonine, any amino acid, the PPR motif selectively binds to A and G; (2-41) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, threonine, and aspartic acid, the PPR motif selectively binds to G; (2-42) A1, A4, and L ii When the three amino acid combinations are, in order, valine, threonine, and aspartic acid, the PPR motif selectively binds to G; (2-43) A1, A4, and L ii When the combination of three amino acids is, in order, any amino acid, threonine, and asparagine, the PPR motif selectively binds to A; (2-44) A1, A4, and L ii When the combination of the three amino acids is, in order, phenylalanine, threonine, and asparagine, the PPR motif selectively binds to A; (2-45) A1, A4, and L ii When the combination of the three amino acids is, in order, isoleucine, threonine, and asparagine, the PPR motif selectively binds to A; (2-46) A1, A4, and L ii When the combination of the three amino acids is, in order, valine, threonine, and asparagine, the PPR motif selectively binds to A; (2-47) A1, A4, and L ii When the three amino acid combinations are, in order, any amino acid, valine, any amino acid, the PPR motif binds to A, C, and T, but not to G; (2-48) A1, A4, and L iiWhen the three amino acid combinations are isoleucine, valine, and aspartic acid, respectively, the PPR motif selectively binds to C and then to A; (2-49) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, valine, and glycine, the PPR motif selectively binds to C; (2-50) A1, A4, and L ii When the combination of these three amino acids is, in order, any amino acid, valine, and threonine, the PPR motif selectively binds to T; and the protein has selective DNA base binding ability.
[0023] (particularly preferred combinations of three amino acids) In RNA-binding PPR motifs, there are representative combinations of amino acids at positions 1, 4, and ii that can recognize and specifically bind to each base. Specifically, the combination that recognizes adenine is valine at position 1, threonine at position 4, and asparagine at position ii; the combination that recognizes cytosine is valine at position 1, asparagine at position 4, and serine at position ii; the combination that recognizes guanine is valine at position 1, threonine at position 4, and aspartic acid at position ii; and the combination that recognizes uracil is valine at position 1, asparagine at position 4, and aspartic acid at position ii (Non-Patent Documents 1 to 5 cited above). In one preferred embodiment of the present invention, these combinations are used.
[0024] (Improved cohesion) Based on the amino acid information of existing naturally occurring PPR motifs, the present inventors have found that the amino acid at position 6 of a PPR motif is often hydrophobic (especially leucine) and the amino acid at position 9 is often non-hydrophilic (especially glycine). Based on the crystal structures of PPR proteins (Non-Patent Document 6: Coquille et al., 2014 Nat. Commun.; PDB IDs: 4PJQ, 4WN4, 4WSL, 4PJR; Non-Patent Document 7: Shen et al., 2015 Nat. Commun., PDB IDs: 5I9D, 5I9F, 5I9G, 5I9H), the inventors hypothesized that the exposed hydrophobic amino acids at positions 6 and 9 of the first motif (N-terminal side) would be responsible for the aggregation tendency (Figure 1A). On the other hand, since the 6th and 9th amino acids in the second and subsequent motifs are buried within the protein and form a hydrophobic core, it was thought that placing hydrophilic residues at the 6th and 9th positions in all motifs could cause the protein structure to collapse. Therefore, we decided to reduce the aggregation tendency of PPR by replacing the 6th amino acid, preferably the 6th and 9th amino acids, in the first motif only with hydrophilic amino acids (asparagine, aspartic acid, glutamine, glutamic acid, lysine, arginine, serine, and threonine).
[0025] Specifically, do the following: In a protein capable of binding to a target nucleic acid having a specific base sequence, the first PPR motif (M1) from the N-terminus is: (1) The A6 amino acid is a hydrophilic amino acid, preferably asparagine or aspartic acid. (2) Furthermore, the A9 amino acid is a hydrophilic amino acid or glycine, preferably glutamine, glutamic acid, lysine, or glycine. (3) Alternatively, the A6 amino acid and the A9 amino acid are any of the following combinations: A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamic acid A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamine A combination in which the A6 amino acid is asparagine and the A9 amino acid is lysine The A6 amino acid is aspartic acid and the A9 amino acid is glycine
[0026] (New PPR motif) The present invention provides the novel PPR motif discovered above, which has improved aggregation properties, and a novel PPR protein containing the same.
[0027] The novel PPR motifs provided by the present invention are as follows: (C-1) a PPR motif consisting of any one of SEQ ID NOs: 4 to 7; (C-2) a cytosine-binding PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, and 34 in any one of SEQ ID NOs: 4 to 7 have been substituted, deleted, or added; (C-3) a PPR motif having at least 80% sequence identity with any one of SEQ ID NOs: 4 to 7, with the proviso that the amino acids at positions 1, 4, 6, and 34 are identical, and which is cytosine-binding; (A-1) A PPR motif consisting of the sequence of SEQ ID NO: 8, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (A-2) A PPR motif consisting of the sequence of (A-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which has adenine-binding properties; (A-3) a PPR motif having at least 80% sequence identity with the sequence of (A-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and having adenine-binding properties; (G-1) A PPR motif consisting of the sequence of SEQ ID NO: 9, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (G-2) A PPR motif consisting of the sequence of (G-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which has guanine-binding properties; (G-3) a PPR motif having at least 80% sequence identity with the sequence of (G-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and which is guanine-binding; (U-1) a PPR motif consisting of the sequence of SEQ ID NO: 10, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (U-2) a PPR motif consisting of the sequence of (U-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which is uracil-binding; (U-3) A PPR motif having at least 80% sequence identity with the sequence of (U-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and which is uracil-binding.
[0028] Among such PPR motifs, the following are particularly preferred: (C-1) a PPR motif consisting of any one of SEQ ID NOs: 4 to 7; (C-2) a cytosine-binding PPR motif consisting of a sequence in which 1 to 9 amino acids other than the amino acids at positions 1, 4, 6, 9, and 34 in any one of SEQ ID NOs: 4 to 7 have been substituted, deleted, or added; (C-3) a PPR motif having at least 80% sequence identity with any one of SEQ ID NOs: 4 to 7, provided that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and having cytosine-binding properties; (A-1) A PPR motif in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 8 are substituted so as to satisfy any one of the combinations defined below; (A-2) A PPR motif consisting of the sequence of (A-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which has adenine-binding properties; (A-3) a PPR motif having at least 80% sequence identity with the sequence of (A-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and having adenine-binding properties; (G-1) A PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 9 have been substituted so as to satisfy any one of the combinations defined below; (G-2) A PPR motif consisting of the sequence of (G-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which has guanine-binding properties; (G-3) a PPR motif having at least 80% sequence identity with the sequence of (G-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and which is guanine-binding; (U-1) a PPR motif consisting of a sequence in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 10 have been substituted so as to satisfy any one of the combinations defined below; (U-2) a PPR motif consisting of the sequence of (U-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which is uracil-binding; (U-3) A PPR motif having at least 80% sequence identity with the sequence of (U-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and which is uracil-binding. A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamic acid A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamine The amino acid at position 6 is asparagine and the amino acid at position 9 is lysine. The amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine
[0029] Specific sequences of SEQ ID NOs: 4 to 10 are shown in FIG. 1 and the sequence listing.
[0030] Among such PPR motifs, the following are more preferred: (C-4) a PPR motif consisting of the sequence of SEQ ID NO: 4; (A-4) a PPR motif consisting of the sequence of SEQ ID NO: 58; (G-4) a PPR motif consisting of the sequence of SEQ ID NO: 59; (U-4) PPR motif consisting of the sequence of SEQ ID NO: 60.
[0031] The sequences of SEQ ID NOs: 58 to 60 are shown below and in the sequence listing. Sequence of SEQ ID NO: 58 VTYTTNIDQLCKAGKVDEALELFKEMRSKGVKPNV Sequence of SEQ ID NO: 59 VTYTTNIDQLCKAGKVDEALELFDEMKERGIKPDV Sequence of SEQ ID NO: 60 VTYNTNIDQLCKAGRLDEAEELLEEMEEKGIKPDV
[0032] (PPR protein with improved aggregation properties) The present invention also provides the PPR protein discovered above, which has improved aggregation properties.
[0033] In a preferred embodiment, the A9 amino acid of M1 is a non-hydrophobic amino acid or glycine, regardless of the other amino acids in M1 or the amino acid sequences of motifs other than M1. The non-hydrophobic amino acid is a hydrophilic amino acid, cysteine, or histidine; preferably, it is a hydrophilic amino acid, i.e., arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, or threonine; more preferably, it is glutamine, glutamic acid, or lysine.
[0034] In a preferred embodiment, the A9 amino acid of M1 is glutamine, glutamic acid, lysine, or glycine, regardless of any other amino acid in M1 or any amino acid sequence in the motif other than M1.
[0035] In a preferred embodiment, the A6 amino acid of M1 is a non-hydrophobic amino acid, regardless of the other amino acids in M1 or the amino acid sequences of motifs other than M1. The non-hydrophobic amino acid is, for example, a hydrophilic amino acid, cysteine, or histidine; preferably, a hydrophilic amino acid, such as arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, or threonine; more preferably, asparagine or aspartic acid.
[0036] In a particularly preferred embodiment, the A6 and A9 amino acids of M1 are any of the following combinations, regardless of the other amino acids in M1 or the amino acid sequence of the motif other than M1: A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamic acid A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamine A combination in which the A6 amino acid is asparagine and the A9 amino acid is lysine The A6 amino acid is aspartic acid and the A9 amino acid is glycine
[0037] In a preferred embodiment of the RNA-binding protein, the A6 and A9 amino acids of M1 satisfy the above-mentioned conditions, and at least one, preferably more than half, and more preferably all of the PPR motifs contained therein satisfy any of the following: If the base to be bound is cytosine, A1 is valine, A4 is asparagine, and A ii is serine If the base to be bound is adenine, A1 is valine, A4 is threonine, and A ii is asparagine If the base to be bound is guanine, A1 is valine, A4 is threonine, and A ii is aspartic acid When the base to be bound is uracil or thymine, A1 is valine, A4 is asparagine, and A ii is aspartic acid
[0038] In one preferred embodiment of the RNA-binding protein, M1 is the novel PPR motif described above.
[0039] In a particularly preferred embodiment, M1 is a PPR motif consisting of any one of the following polypeptides: When the base to be bound is cytosine, a polypeptide consisting of any one of SEQ ID NOs: 4-7 If the base to be linked is adenine, a polypeptide in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 8 are substituted so as to satisfy one of the combinations defined in the following paragraphs. If the base to be linked is guanine, a polypeptide in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO:9 have been substituted so as to satisfy one of the combinations defined in the following paragraphs. If the base to be linked is uracil, a polypeptide in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 10 are substituted so as to satisfy one of the combinations defined in the following paragraphs. At least one of the PPR motifs other than M1 is a PPR motif consisting of any one of the following polypeptides: If the base to be bound is cytosine, a polypeptide having the sequence of SEQ ID NO:2 If the base to be bound is adenine, a polypeptide having the sequence of SEQ ID NO:8 If the base to be bound is guanine, a polypeptide having the sequence of SEQ ID NO:9 If the base to be bound is uracil, the polypeptide has the sequence of SEQ ID NO: 10.
[0040] The combination referred to in the above paragraph is any of the following: A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamic acid A combination in which the A6 amino acid is asparagine and the A9 amino acid is glutamine A combination in which the A6 amino acid is asparagine and the A9 amino acid is lysine The A6 amino acid is aspartic acid and the A9 amino acid is glycine
[0041] In a particularly preferred embodiment, M1 is a PPR motif consisting of any one of the following polypeptides: If the base to be bound is cytosine, a polypeptide having the sequence of SEQ ID NO:4 If the base to be bound is adenine, a polypeptide having the sequence of SEQ ID NO: 58 If the base to be bound is guanine, a polypeptide having the sequence of SEQ ID NO: 59 If the base to be bound is uracil, a polypeptide having the sequence of SEQ ID NO: 60 At least one of the PPR motifs other than M1 is a PPR motif consisting of any one of the following polypeptides: If the base to be bound is cytosine, a polypeptide having the sequence of SEQ ID NO:2 If the base to be bound is adenine, a polypeptide consisting of the sequence of SEQ ID NO:8 in which the amino acid at position 15 is substituted with lysine. If the base to be bound is guanine, a polypeptide having the sequence of SEQ ID NO:9 If the base to be bound is uracil, the polypeptide has the sequence of SEQ ID NO: 10.
[0042] (Utilizing the highly efficient PPR motif framework) In one preferred embodiment of the present invention, in the PPR motifs for cytosine, adenine, guanine, and uracil (or thymine), amino acids other than positions 1, 4, 6, 9, and ii can be selected as specific amino acids. Specifically, among Arabidopsis thaliana PPR motif sequences, those in which the combination of amino acids at positions 1, 4, and ii is VTN for a PPR motif that recognizes adenine, VSN for a PPR motif that recognizes cytosine, VTD for a PPR motif that recognizes guanine, and VND for a PPR motif that recognizes uracil are collected, and the types and numbers of amino acids appearing at each position are compiled. By selecting amino acids that appear frequently at each position, the performance of the PPR motif can be improved.
[0043] From the viewpoint of selecting amino acids other than 1, 4, 6, 9, and ii as frequently occurring amino acids as described above, the amino acid sequences of the PPR motifs listed below can be used as a reference in order to obtain RNA-binding PPR proteins. The PPR motif corresponding to cytosine includes a PPR motif consisting of any one of the sequences of SEQ ID NOs: 4-7; Examples of the PPR motif corresponding to adenine include a PPR motif consisting of any one of the sequences of SEQ ID NO:8; The PPR motif corresponding to guanine includes a PPR motif consisting of any one of the sequences of SEQ ID NO:9; The PPR motif corresponding to uracil is a PPR motif consisting of any one of the sequences of SEQ ID NO:10.
[0044] (Explanation of terms, etc.) In the present invention, the term "identity" in reference to a base sequence (sometimes referred to as a nucleotide sequence) or an amino acid sequence refers to the percentage of matching bases or amino acids shared between the two sequences when the two sequences are optimally aligned, unless otherwise specified. That is, identity can be calculated as follows: identity = (number of matching positions / total number of positions) × 100, and can be calculated using commercially available algorithms. Such algorithms are incorporated into the NBLAST and XBLAST programs described in Altschul et al., J. Mol. Biol. 215 (1990) 403-410. More specifically, searches and analyses for the identity of base sequences or amino acid sequences can be performed using algorithms or programs well known to those skilled in the art (e.g., BLASTN, BLASTP, BLASTX, ClustalW). When using a program, parameters can be appropriately set by those skilled in the art, or the default parameters of each program can be used. Specific techniques for these analysis methods are also well known to those skilled in the art.
[0045] In this specification, when a base sequence or an amino acid sequence is said to be identical (or highly identical), unless otherwise specified, it means that the sequence has at least 70% identity, preferably 80% or more, more preferably 85% or more, even more preferably 90% or more, even more preferably 95% or more, even more preferably 97.5% or more, and even more preferably 99% or more identity.
[0046] Furthermore, in the present invention, when a "substituted, deleted, or added sequence" is used in relation to a PPR motif or protein, the number of amino acids to be substituted, etc. is not particularly limited, unless otherwise specified, in any motif or protein, as long as the motif or protein consisting of that amino acid sequence has the desired function, but may be about 1 to 9 or 1 to 4 amino acids, or may include an even greater number of substitutions, etc., as long as the substitutions are with amino acids with similar properties. Means for preparing polynucleotides or proteins with such amino acid sequences are well known to those skilled in the art.
[0047] Amino acids with similar properties refer to amino acids with similar physical properties such as hydropathy, charge, pKa, and solubility, and include, for example, the following: Hydrophobic amino acids: alanine, valine, glycine, isoleucine, leucine, phenylalanine, proline, tryptophan, tyrosine Non-hydrophobic amino acids; arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine, cysteine, histidine; Hydrophilic amino acids; arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine; Acidic amino acids: aspartic acid, glutamic acid; Basic amino acids: lysine, arginine, histidine; Neutral amino acids: alanine, asparagine, cysteine, glutamine, glycine, isoleucine, leucine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine; Sulfur-containing amino acids: methionine, cysteine; Aromatic ring-containing amino acids: tyrosine, tryptophan, phenylalanine.
[0048] With respect to genes, nucleic acids, polynucleotides, proteins, motifs, etc., "production" can be replaced with "production" or "manufacturing." Furthermore, with respect to genes, etc., when parts are combined to create them, the term "construction" is sometimes used, and "construction" can also be replaced with "production" or "manufacturing."
[0049] The PPR motifs of the present invention, proteins containing them, or nucleic acids encoding them can be produced by those skilled in the art using conventional techniques and the description in the Examples section of this specification.
[0050] [Characteristics and uses of PPR proteins] (Improvement of PPR protein aggregation) PPR proteins prepared using the novel PPR motif of the present invention have reduced intracellular aggregation tendency. Those skilled in the art can evaluate the aggregation tendency of PPR proteins by expressing the PPR protein in cells and confirming the presence or absence of aggregation. Confirmation is easier if the PPR protein is expressed in fusion with a fluorescent protein. Studies by the present inventors have shown that appropriate modification of amino acids in the first motif of PPR proteins improves the intracellular aggregation tendency of PPR proteins and enhances their nuclear translocation.
[0051] (binding force) PPR proteins prepared using the novel PPR motif of the present invention not only have reduced intracellular aggregation but also have RNA-binding activity equivalent to or higher than that of PPR proteins prepared using existing PPR motifs for the same target RNA. "Equivalent" means 55% or more, preferably about 75%.
[0052] The binding strength to the target sequence can be evaluated using methods such as EMSA (Electrophoretic Mobility Shift Assay) and Biacore. EMSA is a method that utilizes the property that the mobility of nucleic acid molecules changes when a sample containing bound protein and nucleic acid is electrophoresed compared to when the nucleic acid is not bound. Molecular interaction analysis devices such as Biacore are capable of reaction kinetic analysis, making it possible to perform detailed analysis of protein-nucleic acid binding.
[0053] The binding strength to the target sequence can also be analyzed by applying a solution containing a candidate protein to the immobilized target nucleic acid and detecting or quantifying the protein bound to the target nucleic acid. This method, which is an adaptation of ELISA (Enzyme-Linked Immunosorbent Assay), is sometimes referred to as RPB-ELISA (RNA-protein binding ELISA). The step of applying a solution containing a candidate protein to the immobilized target nucleic acid can be specifically performed by pouring a solution containing the target binding protein over the target nucleic acid molecule immobilized on a plate. Various existing immobilization methods can be used to immobilize the target nucleic acid molecule. For example, this can be achieved by applying a nucleic acid probe containing a biotin-modified target nucleic acid molecule to a streptavidin-coated well plate. Detailed experimental conditions can be found in the experimental method described in detail in the Examples section of this invention. In RPB-ELISA, the binding strength between the target PPR protein and its target RNA can be determined by subtracting the background signal (the luminescence signal value when the target PPR protein is added without the target RNA) from the luminescence intensity of a sample to which the target PPR protein and its target RNA have been added.
[0054] [Utilization of PPR proteins] (complex, fusion protein) The PPR motif or PPR protein provided by the present invention can be linked to a functional region to form a complex. Furthermore, a proteinaceous functional region can be linked to form a fusion protein. A functional region refers to a portion that has a specific biological function in vivo or in a cell, such as an enzymatic function, catalytic function, inhibitory function, or enhancer function, or a portion that functions as a label. Such a portion can be composed of, for example, a protein, peptide, nucleic acid, physiologically active substance, or drug. Hereinafter, the present invention will be described in relation to a complex, using a fusion protein as an example. However, those skilled in the art will be able to understand the case of complexes other than fusion proteins in accordance with this description.
[0055] In a preferred embodiment, the functional region is a ribonuclease (RNase). Examples of RNases include RNase A (e.g., bovine pancreatic ribonuclease A: PDB 2AAS) and RNase H.
[0056] In one preferred embodiment, the functional region is a fluorescent protein. Examples of fluorescent proteins include mCherry, EGFP, GFP, Sirius, EBFP, ECFP, mTurquoise, TagCFP, AmCyan, mTFP1, MidoriishiCyan, CFP, TurboGFP, AcGFP, TagGFP, Azami-Green, ZsGreen, EmGFP, HyPer, TagYFP, EYFP, Venus, YFP, PhiYFP, PhiYFP-m, TurboYFP, ZsYellow, mBanana, KusabiraOrange, mOrange, TurboRFP, DsRed-Express, DsRed2, TagRFP, DsRed-Monomer, AsRed2, mStrawberry, TurboFP602, mRFP1, JRed, KillerRed, HcRed, KeimaRed, mRasberry, mPlum, PS-CFP, Dendra2, Kaede, EosFP, and KikumeGR. As a fusion protein, a preferred example is mClover3 from the viewpoint of improved aggregation properties and / or efficient nuclear translocation.
[0057] In one preferred embodiment, when the target is mRNA, the functional region is a functional domain that improves the amount of protein expressed from the target mRNA (WO 2017 / 209122). Examples of functional domains that improve the amount of protein expressed from mRNA include, for example, the entirety or a functional portion of a functional domain of a protein known to directly or indirectly promote mRNA translation. More specifically, the functional domain may be a domain that guides ribosomes to mRNA, a domain involved in mRNA translation initiation or translation promotion, a domain involved in mRNA nuclear export, a domain involved in binding to the endoplasmic reticulum membrane, a domain containing an ER retention signal sequence, or a domain containing an endoplasmic reticulum signal sequence. More specifically, the domain that guides ribosomes to mRNA may be a domain containing the entirety or a functional portion of a polypeptide selected from the group consisting of DENR (density-regulated protein), MCT-1 (malignant T-cell amplified sequence 1), TPT1 (translationally-controlled tumor protein), and Lerepo4 (zinc finger CCCH domain). Furthermore, the domain involved in the initiation or promotion of mRNA translation may be a domain containing all or a functional portion of a polypeptide selected from the group consisting of eIF4E and eIF4G. The domain involved in the export of mRNA from the nucleus may be a domain containing all or a functional portion of SLBP (Stem-Loop Binding Protein). The domain involved in binding to the endoplasmic reticulum membrane may be a domain containing all or a functional portion of a polypeptide selected from the group consisting of SEC61B, TRAP-alpha (Translocon Associated Protein alpha), SR-alpha, Dia1 (Cytochrome b5 Reductase 3), and p180.The ER retention signal sequence may be a signal sequence containing the KDEL (SEQ ID NO: 55) or KEEL (SEQ ID NO: 56) sequence, or a signal sequence containing MGWSCIILFLVATATGAHS (SEQ ID NO: 57).
[0058] In the present invention, the functional region may be fused to the N-terminus, the C-terminus, or both the N-terminus and the C-terminus of the PPR protein. The complex or fusion protein may contain multiple functional regions (e.g., 2 to 5). Furthermore, the complex or fusion protein of the present invention may have a functional region indirectly fused to the PPR protein via a linker or the like.
[0059] (Nucleic acids, vectors, and cells encoding PPR proteins, etc.) The present invention also provides nucleic acids encoding the above-mentioned PPR motifs, PPR proteins, or fusion proteins, and vectors containing the nucleic acids (e.g., vectors for amplification, expression vectors). E. coli or yeast can be used as hosts for amplification vectors. As used herein, an expression vector refers to a vector containing, from upstream, DNA having a promoter sequence, DNA encoding a desired protein, and DNA having a terminator sequence; however, these sequences do not necessarily have to be arranged in this order as long as the desired function is exhibited. Various vectors commonly used by those skilled in the art can be recombined and used in the present invention.
[0060] The PPR proteins or fusion proteins of the present invention can function in eukaryotic cells (e.g., animals, plants, microorganisms (yeast, etc.), and protists). The fusion proteins of the present invention can function particularly in animal cells (in vitro or in vivo). Animal cells into which the PPR proteins or fusion proteins of the present invention, or vectors expressing them, can be introduced include, for example, cells derived from humans, monkeys, pigs, cows, horses, dogs, cats, mice, and rats. Furthermore, cultured cells into which the PPR proteins or fusion proteins of the present invention, or vectors expressing them, can be introduced include, but are not limited to, Chinese hamster ovary (CHO) cells, COS-1 cells, COS-7 cells, VERO (ATCC CCL-81) cells, BHK cells, canine kidney-derived MDCK cells, hamster AV-12-664 cells, HeLa cells, WI38 cells, 293 cells, 293T cells, and PER.C6 cells.
[0061] (Application) The PPR protein or fusion protein of the present invention may be able to deliver and function a functional region in a nucleic acid sequence-specific manner in vivo or within a cell. A complex linked to a labeling moiety such as GFP can be used to visualize a desired RNA in vivo.
[0062] Furthermore, the PPR proteins or fusion proteins of the present invention can modify or destroy nucleic acid sequences specifically within cells or living organisms, and may also be able to confer new functions. In particular, RNA-binding PPR proteins are involved in all RNA processing steps found in organelles, including cleavage, RNA editing, translation, splicing, and RNA stabilization. Therefore, the methods for modifying PPR proteins provided by the present invention, as well as the PPR motifs and PPR proteins provided by the present invention, are expected to be used in a variety of fields, including the following:
[0063] (1) Medical care We create PPR proteins that recognize and bind to specific RNAs associated with specific diseases. We also analyze the target sequences of specific RNAs and the associated proteins. These analysis results can be used to search for compounds for treating diseases.
[0064] For example, it is known that in animals, abnormalities in a PPR protein identified as LRPPRC cause Leigh syndrome French Canadian (LSFC; subacute necrotizing encephalomyelopathy). The present invention may contribute to the treatment (prevention, cure, and inhibition of progression) of LSFC. Many existing PPR proteins function to specify editing sites for RNA manipulation (conversion of genetic information on RNA; often C → U). This type of PPR protein has an additional motif on the C-terminus that is suggested to interact with RNA editing enzymes. PPR proteins with such structures are expected to be able to introduce nucleotide polymorphisms or treat diseases or conditions caused by nucleotide polymorphisms.
[0065] - Creating cells with controlled RNA suppression and expression. Such cells include stem cells (e.g., iPS cells) whose differentiated and undifferentiated states are monitored, model cells for evaluating cosmetics, and cells in which functional RNA expression can be turned on and off for the purpose of elucidating drug discovery mechanisms and pharmacological testing.
[0066] We create PPR proteins that specifically bind to specific RNAs associated with specific diseases. By introducing these PPR proteins into cells using plasmids, viral vectors, mRNA, or purified proteins, the PPR proteins can bind to their target RNAs within the cells, altering (improving) the RNA function that causes the disease. Methods for altering function include, for example, altering RNA structure through binding, knockdown through degradation, altering the splicing reaction through splicing, and base substitution.
[0067] (2) Agriculture, Forestry and Fisheries -Improve yields and quality of agricultural, forestry and marine products. -Breeding organisms with improved disease resistance, improved environmental tolerance, and improved or new functionality.
[0068] For example, with regard to first-generation hybrid (F1) crops, it may be possible to artificially produce F1 crops by using PPR proteins to stabilize mitochondrial RNA and regulate translation, potentially improving yield and quality. RNA manipulation and genome editing using PPR proteins enable the genetic improvement and breeding of organisms more accurately and quickly than conventional techniques. Furthermore, unlike genetic modification, which involves the transformation of traits using foreign genes, PPR protein-based RNA manipulation and genome editing are techniques that manipulate the RNA and genomes inherent in plants and animals, making them closer to traditional breeding methods such as mutant selection and backcrossing. This could potentially address global food and environmental issues reliably and quickly.
[0069] (3) Chemistry In the production of useful substances using microorganisms, cultured cells, plants, and animals (e.g., insects), the amount of protein expression can be controlled by manipulating DNA and RNA. This can improve the productivity of useful substances. Examples of useful substances include proteinaceous substances such as antibodies, vaccines, and enzymes, as well as relatively low-molecular-weight compounds such as pharmaceutical intermediates, fragrances, and pigments.
[0070] -Improve the efficiency of biofuel production by modifying the metabolic pathways of algae and microorganisms. [Example]
[0071] [Example 1: Intracellular analysis of fluorescent protein-fused PPR protein] (motif design) The target sequence was CAGCAGCAGCAGCAGCAG (SEQ ID NO: 1), which consists of six repeats of the CAG sequence. The bases recognized by the PPR motif are determined by the amino acid sequence at positions 1, 4, and ii. The PPR motif recognizing cytosine has valine at position 1, asparagine at position 4, and serine at position ii. The PPR motif recognizing adenine has valine at position 1, threonine at position 4, and asparagine at position ii. The PPR motif recognizing guanine has valine at position 1, threonine at position 4, and aspartic acid at position ii. The PPR motif recognizing uracil has valine at position 1, asparagine at position 4, and aspartic acid at position ii.
[0072] Furthermore, the typical combination of the 6th and 9th amino acids in the first motif (mutated motif in Figure 1A) that recognizes cytosine is leucine and glycine (C 6L9G, PPRcag 1, supra, SEQ ID NO:2), and variants containing leucine and glutamic acid (C 6L9E, PPRcag 2, SEQ ID NO:3), asparagine and glutamine (C 6N9Q, PPRcag 3, SEQ ID NO:4), asparagine and glutamic acid (C 6N9E, PPRcag 4, SEQ ID NO:5), asparagine and lysine (C 6N9K, PPRcag 5, SEQ ID NO:6), aspartic acid and glycine (C 6D9G, PPRcag SEQ ID NO: 6, SEQ ID NO: 7) were selected (Figure 1B). These PPR motif sequences were arranged so that they would bind to the CAGCAGCAGCAGCAGCAG sequence (SEQ ID NO: 1, supra), and PPR genes were constructed (SEQ ID NOs: 11-16). In order to efficiently and accurately link the 18 DNAs encoding each PPR motif, amino acids other than the first, fourth, sixth, ninth, and second amino acids in the PPR motifs for cytosine, adenine, and guanine, respectively, were selected to be amino acids that occur frequently, as described above (SEQ ID NOs: 8-9; see Patent Document 1, supra).
[0073] (Plasmid construction) Plasmids containing PPR genes were constructed using the Golden Gate method. More specifically, 10 intermediate vectors, Dest-a, b, c, d, e, f, g, h, i, and j, were prepared and designed to be seamlessly linked in order. 20 motifs, including one and two motifs (PPR motifs corresponding to A, C, G, and U, and two PPR motifs recognizing the base combinations AA, AC, AG, AU, CA, CC, CG, CU, GA, GC, GG, GU, UA, UC, UG, and UU), were inserted into each of the 10 vectors to create 200 parts.
[0074] Dest-a is gaagacataaactccgtggtcacATACagagaccaaggtctcaGTGGtcacatacatgtcttc(SEQ ID NO:43), Dest-b is gaagacatATACagagaccaaggtctcaGTGGtgacataatgtcttc(SEQ ID NO:44), Dest-c is gaagacatcATACagagaccaaggtctcaGTGGttacatatgtcttc(SEQ ID NO:45), Dest-d is gaagacatacATACagagaccaaggtctcaGTGGttacaatgtcttc(SEQ ID NO:46), Dest-e is gaagacattacATACagagaccaaggtctcaGTGGtgacatgtcttc(SEQ ID NO:47), Dest-f is gaagacattgacATACagagaccaaggtctcaGTGGttaatgtcttc(SEQ ID NO:48), Dest-g is gaagacatgttacATACagagaccaaggtctcaGTGGtcatgtcttc(SEQ ID NO:49), Dest-h is gaagacatggtcacATACagagaccaaggtctcaGTGGtatgtcttc(SEQ ID NO:50), Dest-i is gaagacattggttacATACagagaccaaggtctcaGTGGatgtcttc(SEQ ID NO:51), Dest-j is gaagacatgtggtgacATACagagaccaaggtctcaGTGGtcttc(SEQ ID NO:52) was prepared by gene synthesis and cloning into pUC57-kan.
[0075] Dest-a through Dest-j were selected along the target nucleotide sequence and cloned into a vector using the Golden Gate reaction. The vector used here was designed so that the amino acid sequence of MGNSV (SEQ ID NO: 53) was added to the N-terminus of the 18-unit PPR sequence and the amino acid sequence of ELTYNTLISGLGKAGRARDPPV (SEQ ID NO: 54) was added to the C-terminus. The cloned gene was confirmed to be of the correct size, and the sequence of the cloned gene was confirmed by sequencing.
[0076] (Detection of expression in cells) The pcDNA3.1 expression plasmid for cultured animal cells contains a CMV promoter and an SV40 polyA signal sequence, allowing insertion of a gene of interest. To detect PPR protein expression in cells, we expressed PPR proteins fused with fluorescent proteins and analyzed their intracellular aggregation and nuclear localization using fluorescent images. The protein genes fused with EGFP, a nuclear localization signal sequence, PPR protein, and a FLAG epitope tag were inserted into pcDNA3.1 in the following order (SEQ ID NOs: 17-22). Additionally, the gene fused with mClover3, PPR protein, a nuclear localization signal sequence, and a FLAG epitope tag was inserted into pcDNA3.1 in the following order (SEQ ID NOs: 23-28). A PPR-free plasmid was also prepared as a control (SEQ ID NOs: 35-36).
[0077] HEK293T cells were cultured at 1 x 10 in a 10 cm dish in 9 mL DMEM and 1 mL FBS. 6 After culturing for 2 days at 37°C in a 5% CO2 environment, the cells were harvested. The harvested cells were seeded at 4 x 10 per well in a PLL-coated 96-well plate. 4 Cells were seeded at 1000 cells / well and cultured at 37°C for 1 day in a 5% CO2 environment. 200 ng of plasmid DNA, 0.6 μL Fugene®-HD (Promega, E2311), and 200 μL Opti-MEM were mixed and added to the wells. Cultures were then cultured at 37°C for 1 day in a 5% CO2 environment. After culture, the medium was removed and the wells were washed once with 50 μL PBS. Then, 1 μL Hoechst (1 mg / mL, Dojindo Laboratories, 346-07951) and 50 μL PBS were added. The wells were incubated at 37°C for 10 minutes in a 5% CO2 environment, followed by a 50 μL wash with PBS. After washing, 50 μL PBS was added, and GFP and Hoechst fluorescence images of each well were captured using a DMi8 (Leica) fluorescence microscope.
[0078] The results are shown in Figure 2. As a result of confirming the intracellular expression of PPR fused with EGFP and a nuclear localization signal sequence, PPRcag 1 (6L9G) and PPRcag In the case of 2(6L9E), it was confirmed that it did not localize to the nucleus but rather aggregated strongly around the nucleus. 3 (6N9Q), PPRcag 4(6N9E), PPRcag 5(6N9K), PPRcag In the case of 6(6D9G), although aggregation was low, it did not localize to the nucleus. 1 (6L9G) and PPRcag In the case of 2(6L9E), although it was localized to the nucleus, it was observed that it aggregated within the nucleus. 3 (6N9Q), PPRcag 4(6N9E), PPRcag 5(6N9K), PPRcag 6(6D9G) localized to the nucleus and did not aggregate. Therefore, the 6N9E, 6N9Q, 6N9K, and 6D9G mutations were effective in improving aggregation, and mClover3 was better than EGFP for efficient nuclear localization.
[0079] [Example 2: RNA binding analysis of CAG-binding PPR proteins] PPRcag 1. PPRcag 2. PPRcag 3. PPRcag 4. PPRcag 5. PPRcag To confirm the binding of 6 to the target RNA, we prepared a recombinant protein and performed a binding experiment.
[0080] We designed the gene encoding each PPR protein by fusing luciferase to the N-terminus and a 6 x histidine tag sequence to the C-terminus, and cloned it into an E. coli expression plasmid (SEQ ID NOs: 29-34). As a control, we also constructed the Nluc-Hisx6 protein gene (SEQ ID NO: 37), which does not contain PPR protein.
[0081] The completed plasmid was transformed into Escherichia coli Rosetta (DE3) strain. This Escherichia coli was cultured in 2 mL of LB medium containing 100 μg / mL ampicillin at 37°C for 12 hours, and the OD 600 When the pH reached 0.5 to 0.8, the culture was transferred to a 15°C incubator and left to stand for 30 minutes. Then, 100 μL (final concentration: 0.1 mM IPTG) was added and cultured at 15°C for 16 hours. The E. coli pellet was collected by centrifugation at 5,000 xg and 4°C for 10 minutes. 1.5 mL of lysis buffer (20 mM Tris-HCl, pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM MgCl2, 2 mg / ml lysozyme, 1 mM PMSF, 2 μL of DNase) was added and frozen at -80°C for 20 minutes. The cells were then freeze-disrupted at 25°C for 30 minutes with shaking. The supernatant (E. coli lysate) containing soluble PPR proteins was then collected by centrifugation at 3700 rpm and 4°C for 15 minutes.
[0082] The binding experiment between PPR protein and RNA was carried out by the same method as in the binding experiment between PPR protein and biotinylated RNA on a streptavidin plate.
[0083] RNA probes were synthesized (Grainer) using 30-nt RNAs (SEQ ID NOs: 38-42, respectively) containing the target CAGx6 sequence, non-target CGGx6, CUGx6, CCGx6, and D1b (UGGUGUAUCUUGUCUUUA) sequences (SEQ ID NO: 42, positions 8-25), each modified with biotin at the 5' end. 2.5 pmol of biotinylated RNA probe was added to a streptavidin-coated plate (Cat No. 15502, Thermo Fisher Scientific) and incubated for 30 minutes at room temperature. The plate was then washed with probe wash buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT, 0.1% BSA). For background measurements, wells containing lysis buffer without biotinylated RNA were also prepared (-Probe). Then, blocking buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT, 1% BSA) was added, and the plate surface was blocked at room temperature for 30 minutes. 8 100 μL of E. coli lysate containing luciferase-fused PPR proteins with luminescence output of LU / μL was added, and the binding reaction was carried out at room temperature for 30 minutes. The wells were then washed five times with 200 μL of wash buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT). 40 μL of luciferase substrate (Promega, E151A) diluted 2500-fold in wash buffer was added to the wells, and the reaction was allowed to proceed for 5 minutes. Luminescence was then measured using a plate reader (PerkinElmer, Cat No. 5103-35).
[0084] The results are shown in Figure 3. All PPRs were found to bind specifically to the target CAGx6. The binding strength to the target sequence was PPRcag compared to 1 2 is the same, PPRcag 3 is about 80%, PPRcag 4 is about 60%, PPRcag 5 is about 120%, PPRcag 6 was about 130%. It was found that except for 4, there was almost no change in binding ability due to mutation.
[0085] [Example 3: Control of PPR protein aggregation] A PPR protein using the V2 motif (nucleotide sequence: SEQ ID NO:61, amino acid sequence: SEQ ID NO:62) and a PPR protein using the v3.2 motif (nucleotide sequence: SEQ ID NO:63, amino acid sequence: SEQ ID NO:64) were produced in an E. coli expression system, purified, and separated by gel filtration chromatography. The v2 motif refers to the PPR motif having the sequence of SEQ ID NO:2 or SEQ ID NOs:8-10. The v3.2 motif refers to the PPR motif having the sequence of SEQ ID NO:4 or SEQ ID NOs:58-60 when the first amino acid from the N-terminus is the PPR motif. Otherwise, for adenine, the aspartic acid at position 15 in SEQ ID NO:8 is substituted with lysine, and for bases other than adenine, the PPR motif has a sequence selected from SEQ ID NOs:2, 9, and 10.
[0086] (Protein expression and purification) E. coli Rosetta strain was transformed with the pE-SUMOpro Kan plasmid containing the DNA sequence encoding the target PPR. After incubation at 37°C, the temperature was lowered to 20°C when the OD600 reached 0.6. IPTG was added to a final concentration of 0.5 mM to express the target PPR protein as a SUMO fusion protein in the E. coli. After overnight incubation, the cells were harvested by centrifugation and resuspended in lysis buffer (50 mM Tris-HCl pH 8.0, 500 mM NaCl). The E. coli cells were disrupted by sonication and centrifuged at 17,000 xg for 30 min. The supernatant was applied to a Ni-Agarose column. After washing with lysis buffer containing 20 mM imidazole, the SUMO-fused target PPR protein was eluted with lysis buffer containing 400 mM imidazole. After elution, the Ulp1-mediated SUMO protein was cleaved from the target PPR protein, and the protein solution was replaced with ion exchange buffer (50 mM Tris-HCl pH 8.0, 200 mM NaCl) by dialysis. Cation exchange chromatography was then performed using an SP column. After loading onto the column, the protein was eluted by gradually increasing the NaCl concentration from 200 mM to 1 M. The fraction containing the target PPR protein was then purified by gel filtration chromatography using a Superdex 200 column. The target PPR protein eluted from the ion exchange column was loaded onto a gel filtration column equilibrated with gel filtration buffer (25 mM HEPES pH 7.5, 200 mM NaCl, 0.5 mM tris(2-carboxyethyl)phosphine (TCEP)). Finally, the fraction containing the target PPR protein was concentrated, frozen in liquid nitrogen, and stored at -80°C until further analysis.
[0087] (Gel filtration chromatography) The purified recombinant PPR protein was adjusted to a concentration of 1 mg / ml. Gel filtration chromatography was performed using Superdex 200 increase 10 / 300 GL (GE Healthcare). The prepared protein was applied to a gel filtration column equilibrated with 25 mM HEPES pH 7.5, 200 mM NaCl, and 0.5 mM tris(2-carboxyethyl)phosphine (TCEP). The absorbance of the solution eluted from the gel filtration column was measured at 280 nm to analyze the protein properties.
[0088] (result) The results are shown in Figure 4. The smaller the elution volume, the larger the molecular size. In V2, the protein eluted in the 8 to 10 mL elution fraction, while in v3.2, a peak was observed in the 12 to 14 mL elution fraction. This suggests that the larger protein size in v2 may have caused aggregation, and that this aggregation was improved in v3.2.
Claims
1. A composition for the treatment of a disease, comprising a vector containing a nucleic acid encoding a protein, The protein is represented by the following formula 1 【Chemistry 1】 (In the formula: Helix A is a 12 amino acid long portion capable of forming an α-helical structure and is represented by Formula 2: 【Chemistry 2】 In formula 2, A 1 ~A 12 each independently represents an amino acid; X is absent or a moiety consisting of 1 to 9 amino acids in length; Helix B is a portion consisting of 11 to 13 amino acids that can form an α-helical structure; L is a moiety of formula 3 that is 2 to 7 amino acids in length; 【Transformation 3】 In Formula 3, each amino acid is numbered from the C-terminus side as "i"(-1), "ii"(-2), etc. However, L iii ~L vii may not exist.) a protein that contains 1 to 30 PPR motifs represented by the formula: The first PPR motif from the N-terminus (M 1 )but: A 6 a polypeptide in which the amino acid is a hydrophilic amino acid; or Any one of the polypeptides listed in Group A and Group B below That is, protein. Group A: (C-1) a polypeptide consisting of any one of SEQ ID NOs: 4 to 7; (C-2) A polypeptide consisting of any one of SEQ ID NOs: 4 to 7, in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which is cytosine-binding; (C-3) A polypeptide having at least 70% sequence identity with any one of SEQ ID NOs: 4 to 7, with the proviso that the amino acids at positions 1, 4, 6, and 34 are identical and are cytosine-binding; (A-1) a polypeptide consisting of the sequence of SEQ ID NO: 8, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (A-2) A polypeptide consisting of the sequence of (A-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which has adenine-binding properties; (A-3) A polypeptide having at least 70% sequence identity with the sequence of (A-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and which is adenine-binding; (G-1) a polypeptide consisting of the sequence of SEQ ID NO: 9, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (G-2) A polypeptide consisting of the sequence of (G-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which is guanine-binding; (G-3) A polypeptide having at least 70% sequence identity with the sequence of (G-1), except that the amino acids at positions 1, 4, 6, and 34 are identical, and which is guanine-binding; (U-1) a polypeptide consisting of the sequence of SEQ ID NO: 10, in which the amino acid at position 6 is substituted with asparagine or aspartic acid; (U-2) A polypeptide consisting of the sequence of (U-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, and 34 have been substituted, deleted, or added, and which is uracil-binding; (U-3) A polypeptide having at least 70% sequence identity with the sequence of (U-1), except that the amino acids at positions 1, 4, 6, and 34 are identical and are uracil-linked. Group B: (C-1) a polypeptide consisting of any one of SEQ ID NOs: 4 to 7; (C-2) A polypeptide consisting of any one of SEQ ID NOs: 4 to 7, in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which is cytosine-binding; (C-3) A polypeptide having at least 70% sequence identity with any one of SEQ ID NOs: 4 to 7, with the proviso that the amino acids at positions 1, 4, 6, 9, and 34 are identical and are cytosine-binding; (A-1) A polypeptide in which the amino acids at positions 6 and 9 in the sequence of SEQ ID NO: 8 have been substituted so as to satisfy any one of the combinations defined below; (A-2) A polypeptide consisting of the sequence of (A-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which has adenine-binding properties; (A-3) A polypeptide having at least 70% sequence identity with the sequence of (A-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and which is adenine-binding; (G-1) A polypeptide consisting of the sequence of SEQ ID NO: 9 in which the amino acids at positions 6 and 9 are substituted so as to satisfy any one of the combinations defined below; (G-2) A polypeptide consisting of the sequence of (G-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which is guanine-binding; (G-3) A polypeptide having at least 70% sequence identity with the sequence of (G-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical, and which is guanine-binding; (U-1) A polypeptide consisting of the sequence of SEQ ID NO: 10 in which the amino acids at positions 6 and 9 have been substituted so as to satisfy any one of the combinations defined below; (U-2) A polypeptide consisting of the sequence of (U-1) in which 1 to 9 amino acids other than those at positions 1, 4, 6, 9, and 34 have been substituted, deleted, or added, and which is uracil-binding; (U-3) A polypeptide having at least 70% sequence identity with the sequence of (U-1), except that the amino acids at positions 1, 4, 6, 9, and 34 are identical and are uracil-linked. A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamic acid A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is glutamine A combination in which the amino acid at position 6 is asparagine and the amino acid at position 9 is lysine The amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine.
2. M 1 A 9 The composition of claim 1 , wherein the amino acid is a hydrophilic amino acid or glycine.
3. M 1 A 6 3. The composition of claim 1, wherein the amino acid is asparagine or aspartic acid.
4. M 1 A 9 The composition according to any one of claims 1 to 3, wherein the amino acid is glutamine, glutamic acid, lysine, or glycine.
5. M 1 A 6 Amino acids and M 1 A 9 The composition according to any one of claims 1 to 4, wherein the amino acids are any combination of the following: ・A 6 The amino acid is asparagine, and A 9 Combination of amino acids containing glutamic acid ・A 6 The amino acid is asparagine, and A 9 Combination of amino acids containing glutamine ・A 6 The amino acid is asparagine, and A 9 Combination of amino acid lysine ・A 6 The amino acid aspartic acid and A 9 Combination of amino acid glycine
6. M 1 The composition of any one of claims 1 to 5, wherein is any one of the following: (C-4) a polypeptide consisting of the sequence of SEQ ID NO: 4; (A-4) a polypeptide consisting of the sequence of SEQ ID NO: 58; (G-4) a polypeptide consisting of the sequence of SEQ ID NO: 59; (U-4) A polypeptide consisting of the sequence of SEQ ID NO:
60.
7. A composition for the manipulation of nucleic acids, comprising a vector containing a nucleic acid encoding a protein as defined in any one of claims 1 to 6.
8. The composition according to any one of claims 1 to 7, wherein the nucleic acid encoding the protein is a nucleic acid encoding a fusion protein comprising said protein.
9. The composition according to claim 8, wherein the fusion protein comprising the protein is a fusion protein in which the protein is linked to a functional region capable of acting on RNA.
10. The composition of claim 9 , wherein the functional region has an enzymatic function, a catalytic function, an inhibitory function, or an activating function.
11. 10. The composition of claim 9, wherein the functional domain is a ribonuclease.
Citation Information
Patent Citations
JP195056B
Design method for RNA-binding protein using PPR motif, and use thereof
WO2013058404A1
DNA binding protein using PPR motif, and use thereof
WO2014175284A1