Low-aggregation ppr proteins and uses thereof
By performing hydrophilic substitutions on the 6th and 9th amino acids of the PPR motif, the aggregation problem of PPR protein in animal cell expression was solved, and its intracellular localization and nucleic acid binding ability were improved.
Patent Information
- Application Number
- CN202080040065.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-29
- Filing Date
- 2020-05-29
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2040-05-29
AI Technical Summary
When expressing PPR protein using animal cultured cells, there is an aggregation problem that affects its functional performance.
The aggregation property of PPR proteins can be improved by replacing the 6th and 9th amino acids in the PPR motif with hydrophilic amino acids. Specifically, replacing the 6th amino acid with asparagine, aspartic acid, or other hydrophilic amino acids can reduce aggregation property.
It effectively reduced the aggregation of PPR proteins, improved their intracellular localization and functional performance, and enhanced their binding ability to nucleic acids with specific base sequences.
Smart Images

Figure HDA0003380792870000011 
Figure HDA0003380792870000021 
Figure HDA0003380792870000031
Abstract
Description
Technical Field
[0001] This invention relates to nucleic acid manipulation techniques using proteins capable of binding to desired nucleic acids. This invention has applications in a wide range of fields, including medicine (innovative drug development, disease treatment), agriculture (agricultural product production, breeding), and chemistry (production of biological substances). Background Technology
[0002] PPR proteins are proteins containing repeating motifs consisting of PPR motifs of approximately 35 amino acids in length. Each PPR motif can specifically bind to one base. The binding to adenine, cytosine, guanine, uracil (or thymine) is determined by the combination of amino acids at positions 1, 4, and ii (two bases before the next motif) within the PPR motif (Patent Documents 1 and 2).
[0003] The most common combinations of bases in naturally occurring RNA-binding PPR motifs are as follows: for adenine, valine at position 1, threonine at position 4, and asparagine at position ii; for cytosine, valine at position 1, asparagine at position 4, and serine at position ii; for guanine, valine at position 1, threonine at position 4, and aspartic acid at position ii; and for uracil, valine at position 1, asparagine at position 4, and aspartic acid at position ii (Non-Patent Literature 1-5). By utilizing these amino acid combinations, PPR proteins that can specifically bind to any sequence can be designed.
[0004] Existing technical documents
[0005] Patent documents
[0006] Patent Document 1: International Publication WO2013 / 058404
[0007] Patent Document 2: International Publication WO2014 / 175284
[0008] Patent Document 3: Japanese Patent Application 2019-100551
[0009] Non-patent literature
[0010] Non-patent literature 1: Coquille, S. et al. An artificial PPR scaffold for programmable RNA recognition. Nature Communications 5, Article number: 5729 (2014)
[0011] Non-patent literature 2: Shen, C. et al. Specific RNA Recognition by Designer Pentatricopeptide Repeat Protein. Molecular Plant 8, 667-670 (2015)
[0012] Non-patent document 3: Shen, C. et al. Structural basis for specific single-stranded RNA recognition by designer pentatricopeptide repeat proteins. Nature Communications volume 7, Article number: 11285 (2016)
[0013] Non-patent literature 4: Miranda, RGet al. RNA-binding specificity landscapes of designer pentatricopeptide repeat proteins elucidate principles of PPR-RNAinteractions. Nucleic Acids Research, 46(5), 2613-2623(2018)
[0014] Non-patent literature 5: Yan, J. et al. Delineation of pentatricopeptide repeat codes for target RNAprediction. Nucleic Acids Research, gkz075 (2019). Summary of the Invention
[0015] The problem that the invention aims to solve
[0016] The inventors conducted research on the production of high-performance PPR proteins by utilizing the above-mentioned amino acid combinations and linking multiple (e.g., more than 15) PPR motifs together (Patent Document 3). On the other hand, according to the inventors' research, some PPR proteins produced using this method exhibit aggregation properties. Aggregation is sometimes observed, particularly when expressing PPR proteins using animal cultured cells.
[0017] Methods for solving problems
[0018] Therefore, studies were conducted to address this issue through amino acid variations within the PPR motif. It was discovered that by making the 6th, preferably the 6th and 9th, amino acids at the N-terminal position of the first motif of the PPR protein hydrophilic, the cohesiveness of the PPR could be improved, thus completing this invention.
[0019] The present invention provides the following solution.
[0020] [1] A PPR motif, which is any of the following PPR motifs:
[0021] (C-1) A PPR motif consisting of any sequence from sequence numbers 4 to 7;
[0022] (C-2) is a cytosine-binding PPR motif consisting of 1 to 9 amino acids other than amino acids at positions 1, 4, 6 and 34 in any of the sequences 4 to 7, which have been replaced, deleted or added.
[0023] (C-3) has at least 80% sequence identity with any of the sequences in sequence numbers 4 to 7, wherein the amino acids at positions 1, 4, 6 and 34 are the same and are cytosine-binding PPR motifs;
[0024] (A-1) The PPR motif consisting of sequence number 8 has an amino acid at position 6 replaced with an asparagine or aspartic acid.
[0025] (A-2) is a sequence consisting of 1 to 9 amino acids other than positions 1, 4, 6 and 34 in the sequence of (A-1) that have been replaced, deleted or added, and is an adenine-binding PPR motif.
[0026] (A-3) and (A-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6 and 34 being identical, and both being adenine-binding PPR motifs;
[0027] (G-1) The PPR motif consisting of sequence number 9 has an amino acid at position 6 replaced by asparagine or aspartic acid.
[0028] (G-2) is a guanine-binding PPR motif consisting of 1 to 9 amino acids other than positions 1, 4, 6 and 34 in the sequence of (G-1) that have been replaced, deleted or added.
[0029] (G-3) and (G-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6 and 34 being identical and being guanine-binding PPR motifs;
[0030] (U-1) The PPR motif consisting of sequence number 10 has an amino acid at position 6 replaced by asparagine or aspartic acid.
[0031] (U-2) is a uracil-binding PPR motif consisting of 1 to 9 amino acids other than amino acids 1, 4, 6 and 34 in the sequence of (U-1) that have been replaced, deleted or added.
[0032] The sequences of (U-3) and (U-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6 and 34 being identical, and both being uracil-binding PPR motifs.
[0033] [2] A PPR motif, which is any of the following PPR motifs:
[0034] (C-1) A PPR motif consisting of any sequence from sequence numbers 4 to 7;
[0035] (C-2) is a cytosine-binding PPR motif consisting of 1 to 9 amino acids other than amino acids 1, 4, 6, 9 and 34 in any of the sequences in sequence numbers 4 to 7, which have been replaced, deleted or added.
[0036] (C-3) has at least 80% sequence identity with any of the sequences in sequence numbers 4 to 7, wherein the amino acids at positions 1, 4, 6, 9 and 34 are the same and are cytosine-binding PPR motifs;
[0037] (A-1) A PPR motif in sequence number 8 in which amino acids at positions 6 and 9 are substituted in any combination that satisfies the following definition;
[0038] (A-2) is a PPR motif that consists of 1 to 9 amino acids other than amino acids 1, 4, 6, 9 and 34 in the sequence of (A-1) that have been replaced, deleted or added, and is adenine-binding.
[0039] (A-3) and (A-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6, 9 and 34 being identical and being adenine-binding PPR motifs;
[0040] (G-1) A PPR motif consisting of a sequence in which amino acids at positions 6 and 9 of sequence number 9 are substituted in any combination that satisfies the following definition.
[0041] (G-2) is a guanine-binding PPR motif consisting of 1 to 9 amino acids other than positions 1, 4, 6, 9 and 34 in the sequence of (G-1) through substitution, deletion or addition.
[0042] (G-3) and (G-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6, 9 and 34 being identical and being guanine-binding PPR motifs;
[0043] (U-1) A PPR motif consisting of a sequence in which amino acids at positions 6 and 9 of sequence number 10 are substituted in any combination that satisfies the following definition.
[0044] (U-2) is a uracil-binding PPR motif consisting of 1 to 9 amino acids other than positions 1, 4, 6, 9 and 34 in the sequence of (U-1) through substitution, deletion or addition.
[0045] The sequences of (U-3) and (U-1) have at least 80% sequence identity, with the amino acids at positions 1, 4, 6, 9 and 34 being identical, and both being uracil-binding PPR motifs.
[0046] • The amino acid at position 6 is asparagine, and the amino acid at position 9 is glutamic acid.
[0047] • The amino acid at position 6 is asparagine, and the amino acid at position 9 is glutamine.
[0048] • The amino acid at position 6 is asparagine, and the amino acid at position 9 is lysine.
[0049] • The amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine.
[0050] [3] The PPR motif as described in 1 or 2, wherein the PPR motif is any of the following:
[0051] (C-4) The PPR motif consisting of the sequence number 4;
[0052] (A-4) The PPR motif consisting of the sequence number 58;
[0053] (G-4) The PPR motif consisting of sequence number 59;
[0054] (U-4) is a PPR motif consisting of the sequence number 60.
[0055] [4] The PPR motif described in any one of 1 to 3 is used as the first PPR motif from the N-terminus in the PPR protein.
[0056] [5] As described in 4, it is used to reduce the cohesiveness of PPR proteins.
[0057] [6] A protein comprising 1 to 30 PPR motifs represented by Formula 1 below, capable of binding to a target nucleic acid having a specific base sequence, wherein the A6 amino acid of the first PPR motif (M1) from the N-terminus is a hydrophilic amino acid.
[0058] [Chemistry 1]
[0059] (Helix A)-X-(Helix B)-L(Equation 1)
[0060] (In the formula:
[0061] Helix A is a portion of 12 amino acids in length capable of forming an α-helix structure, represented by Equation 2.
[0062] [Chemistry 2]
[0063] A1-A2-A3-A4-A5-A6-A7-A8-A9-A 10 -A 11 -A 12 (Equation 2)
[0064] In Equation 2, A1~A 12 Each amino acid can be represented independently;
[0065] X either does not exist or is composed of 1 to 9 amino acids in length;
[0066] Helix B is a portion composed of 11 to 13 amino acids in length that can form an α-helix structure;
[0067] L is the portion represented by Equation 3, which has a length of 2 to 7 amino acids;
[0068] [Chemistry 3]
[0069] L vii -L vi -L v -L iv -L iii -L ii -L i (Equation 3)
[0070] In Formula 3, each amino acid is numbered "i" (-1) and "ii" (-2) starting from the C-terminus.
[0071] Where L iii ~L vii (There are cases where it does not exist.)
[0072] [7] The protein as described in 6, wherein the A9 amino acid of M1 is a hydrophilic amino acid or glycine.
[0073] [8] The protein as described in 6 or 7, wherein the A6 amino acid of M1 is asparagine or aspartic acid.
[0074] [9] The protein as described in any one of 6 to 8, wherein the A9 amino acid of M1 is glutamine, glutamic acid, lysine or glycine.
[0075]
[10] The protein as described in any one of 6 to 9, wherein the A6 amino acid of M1 and the A9 amino acid of M1 are any combination of the following.
[0076] • A6 amino acid is asparagine, and A is a combination of amino acids that are glutamic acid.
[0077] • A6 amino acid is asparagine, and A. amino acid is glutamine.
[0078] • The combination of amino acid A6 being asparagine and amino acid A9 being lysine
[0079] • The combination of A6 amino acid being aspartic acid and A9 amino acid being glycine.
[0080]
[11] A fusion protein, which is a fusion protein of at least one of the group consisting of fluorescent protein, nuclear transport signal peptide and tag protein, and a PPR protein containing the PPR motif of any one of 1 to 3 as the first PPR motif from the N-terminus, or a protein of any one of 6 to 10.
[0081]
[12] A modification method comprising a PPR protein containing the PPR motifs defined in 6, capable of binding to a target nucleic acid having a specific base sequence, wherein the A6 amino acid of the first PPR motif (M1) from the N-terminus is made more hydrophilic.
[0082]
[13] A method for detecting nucleic acid, characterized in that a PPR protein containing any one of 1 to 3 as the first PPR motif from the N-terminus, a protein containing any one of 6 to 10, or a fusion protein as described in 11 is used.
[0083]
[14] A nucleic acid encoding any one of the PPR motifs in 1 to 3, a PPR protein comprising any one of the PPR motifs in 1 to 3 as the first PPR motif from the N-terminus, or any one of 6 to 10.
[0084]
[15] A vector comprising the nucleic acid described in 14.
[0085]
[16] A cell (excluding individual human cells) comprising the carrier described in 15.
[0086]
[17] A method of manipulating a nucleic acid (excluding implementation in a human individual) using any one of the PPR motifs in 1 to 3, a PPR protein comprising any one of the PPR motifs in 1 to 3 as the first PPR motif from the N-terminus, or any one of the proteins in 6 to 10, or the vector described in 15.
[0087]
[18] A method for producing a biological substance, comprising the operating method described in 17.
[0088] The present invention also provides the following solutions.
[0089] [1] A PPR motif, which is any of the following PPR motifs:
[0090] (C-1) A PPR motif consisting of any sequence from sequence numbers 4 to 7;
[0091] (C-2) is a cytosine-binding PPR motif consisting of 1 to 9 amino acids other than amino acids at positions 1, 4, 6 and 34 in any of the sequences 4 to 7, which have been replaced, deleted or added.
[0092] (C-3) has at least 80% sequence identity with any of the sequences in sequence numbers 4 to 7, wherein the amino acids at positions 1, 4, 6 and 34 are the same and are cytosine-binding PPR motifs;
[0093] (A-1) The PPR motif consisting of sequence number 8 has an amino acid at position 6 replaced with an asparagine or aspartic acid.
[0094] (A-2) is a sequence consisting of 1 to 9 amino acids other than positions 1, 4, 6 and 34 in the sequence of (A-1) that have been replaced, deleted or added, and is an adenine-binding PPR motif.
[0095] (A-3) and (A-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6 and 34 being identical, and both being adenine-binding PPR motifs;
[0096] (G-1) The PPR motif consisting of sequence number 9 has an amino acid at position 6 replaced by asparagine or aspartic acid.
[0097] (G-2) is a guanine-binding PPR motif consisting of 1 to 9 amino acids other than positions 1, 4, 6 and 34 in the sequence of (G-1) that have been replaced, deleted or added.
[0098] (G-3) and (G-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6 and 34 being identical and being guanine-binding PPR motifs;
[0099] (U-1) The PPR motif consisting of sequence number 10 has an amino acid at position 6 replaced by asparagine or aspartic acid.
[0100] (U-2) is a uracil-binding PPR motif consisting of 1 to 9 amino acids other than amino acids 1, 4, 6 and 34 in the sequence of (U-1) that have been replaced, deleted or added.
[0101] The sequences of (U-3) and (U-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6 and 34 being identical, and both being uracil-binding PPR motifs.
[0102] [2] A PPR motif, which is any of the following PPR motifs:
[0103] (C-1) A PPR motif consisting of any sequence from sequence numbers 4 to 7;
[0104] (C-2) is a cytosine-binding PPR motif consisting of 1 to 9 amino acids other than amino acids 1, 4, 6, 9 and 34 in any of the sequences in sequence numbers 4 to 7, which have been replaced, deleted or added.
[0105] (C-3) has at least 80% sequence identity with any of the sequences in sequence numbers 4 to 7, wherein the amino acids at positions 1, 4, 6, 9 and 34 are the same and are cytosine-binding PPR motifs;
[0106] (A-1) A PPR motif in sequence number 8 in which amino acids at positions 6 and 9 are substituted in any combination that satisfies the following definition;
[0107] (A-2) is a PPR motif that consists of 1 to 9 amino acids other than amino acids 1, 4, 6, 9 and 34 in the sequence of (A-1) that have been replaced, deleted or added, and is adenine-binding.
[0108] (A-3) and (A-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6, 9 and 34 being identical and being adenine-binding PPR motifs;
[0109] (G-1) A PPR motif consisting of a sequence in which amino acids at positions 6 and 9 of sequence number 9 are substituted in any combination that satisfies the following definition.
[0110] (G-2) is a guanine-binding PPR motif consisting of 1 to 9 amino acids other than positions 1, 4, 6, 9 and 34 in the sequence of (G-1) through substitution, deletion or addition.
[0111] (G-3) and (G-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6, 9 and 34 being identical and being guanine-binding PPR motifs;
[0112] (U-1) A PPR motif consisting of a sequence in which amino acids at positions 6 and 9 of sequence number 10 are substituted in any combination that satisfies the following definition.
[0113] (U-2) is a uracil-binding PPR motif consisting of 1 to 9 amino acids other than positions 1, 4, 6, 9 and 34 in the sequence of (U-1) through substitution, deletion or addition.
[0114] The sequences of (U-3) and (U-1) have at least 80% sequence identity, with the amino acids at positions 1, 4, 6, 9 and 34 being identical, and both being uracil-binding PPR motifs.
[0115] • The amino acid at position 6 is asparagine, and the amino acid at position 9 is glutamic acid.
[0116] • The amino acid at position 6 is asparagine, and the amino acid at position 9 is glutamine.
[0117] • The amino acid at position 6 is asparagine, and the amino acid at position 9 is lysine.
[0118] • The amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine.
[0119] [3] The PPR motif described in 1 or 2 is used as the first PPR motif from the N-terminus in the PPR protein.
[0120] [4] As described in 3, it is used to reduce the aggregation of PPR proteins.
[0121] [5] A protein comprising 1 to 30 PPR motifs represented by Formula 1 below, capable of binding to a target nucleic acid having a specific base sequence, wherein the A6 amino acid of the first PPR motif (M1) from the N-terminus is a hydrophilic amino acid.
[0122] [Chemistry 4]
[0123] (Helix A)-X-(Helix B)-L (Equation 1)
[0124] (In the formula:
[0125] Helix A is a portion of 12 amino acids in length capable of forming an α-helix structure, represented by Equation 2.
[0126] [Chemistry 5]
[0127] A1-A2-A3-A4-A5-A6-A7-A8-A9-A 10 -A 11 -A 12 (Equation 2)
[0128] In Equation 2, A1~A 12 Each amino acid can be represented independently;
[0129] X either does not exist or is composed of 1 to 9 amino acids in length;
[0130] Helix B is a portion composed of 11 to 13 amino acids in length that can form an α-helix structure;
[0131] L is the portion represented by Equation 3, which has a length of 2 to 7 amino acids;
[0132] [Chemistry 6]
[0133] Lv li -L vi -L v -L iv -L iii -L ii -L i (Equation 3)
[0134] In Formula 3, each amino acid is numbered "i" (-1) and "ii" (-2) from the C-terminus side.
[0135] Where L iii ~L vii (There are cases where it does not exist.)
[0136] [6] The protein as described in 5, wherein the A9 amino acid of M1 is a hydrophilic amino acid or glycine.
[0137] [7] The protein as described in 5 or 6, wherein the A6 amino acid of M1 is asparagine or aspartic acid.
[0138] [8] The protein as described in any one of 5 to 7, wherein the A9 amino acid of M1 is glutamine, glutamic acid, lysine or glycine.
[0139] [9] The protein as described in any one of 5 to 8, wherein the A6 amino acid of M1 and the A9 amino acid of M1 are any combination of the following.
[0140] • The combination of amino acid A6 being asparagine and amino acid A9 being glutamic acid
[0141] • The combination of amino acid A6 being asparagine and amino acid A9 being glutamine
[0142] • The combination of amino acid A6 being asparagine and amino acid A9 being lysine
[0143] • The combination of A6 amino acid being aspartic acid and A9 amino acid being glycine.
[0144]
[10] A fusion protein, which is a fusion protein of at least one of the group consisting of fluorescent proteins, nuclear transport signal peptides and tag proteins, and a PPR protein containing the PPR motif described in 1 or 2 as the first PPR motif from the N-terminus, or a protein of any one of 5 to 9.
[0145]
[11] A modification method comprising a PPR protein containing the PPR motifs defined in 3, capable of binding to a target nucleic acid having a specific base sequence, wherein the A6 amino acid of the first PPR motif (M1) from the N-terminus is made more hydrophilic.
[0146]
[12] A method for detecting nucleic acid, characterized in that a PPR protein containing the PPR motif described in 1 or 2 as the first PPR motif from the N-terminus, a protein described in any one of 5 to 9, or a fusion protein described in 10 is used.
[0147]
[13] A nucleic acid encoding the PPR motif described in 1 or 2, a PPR protein comprising the PPR motif described in 1 or 2 as the first PPR motif from the N-terminus, or a protein as described in any one of 5 to 9.
[0148]
[14] A vector comprising the nucleic acid described in 13.
[0149]
[15] A cell (excluding human individual cells) comprising the carrier described in 14.
[0150]
[16] A method of manipulating a nucleic acid (excluding implementation in a human individual) using the PPR motif described in 1 or 2, a PPR protein comprising the PPR motif described in 1 or 2 as the first PPR motif from the N-terminus, or a protein as described in any one of 5 to 9, or a vector as described in 14.
[0151]
[17] A method for producing a biological substance, comprising the operating method described in 16. Attached Figure Description
[0152] Figure 1This is the PPR motif design method. A: The amino acids at positions 6 and 9 of the first motif are exposed to the outside. B: Regarding the amino acids at positions 6 and 9 of the first cytosine motif, leucine and glycine (C_6L9G) are selected as representative combinations, and leucine and glutamic acid (C_6L9E), asparagine and glutamine (C_6N9Q), asparagine and glutamic acid (C_6N9E), asparagine and lysine (C_6N9K), and aspartic acid and glycine (C_6D9G) are selected as variants.
[0153] Figure 2 This study investigated the aggregation and nuclear transport of various PPR proteins. Fluorescence microscopy confirmed the expression of PPRs fused with GFP and nuclear transport signaling sequences within the cell. With EGFP fusion, PPRcag_1 (6L9G) and PPRcag_2 (6L9E) were found to be significantly aggregated around the nucleus, not localized within it. Conversely, PPRcag_3 (6N9Q), PPRcag_4 (6N9E), PPRcag_5 (6N9K), and PPRcag_6 (6D9G) exhibited low aggregation but were not localized within the nucleus. With mClover3 fusion, PPRcag_1 (6L9G) and PPRcag_2 (6L9E) were found to be localized within the nucleus but aggregated there. PPRcag_3(6N9Q), PPRcag_4(6N9E), PPRcag_5(6N9K), and PPRcag_6(6D9G) are located in the core, and no cohesion was observed.
[0154] Figure 3 This is a binding experiment of PPR protein to RNA. It shows that all PPR proteins, including those with amino acid variations at positions 6 and 9, specifically bind to CAGx6 as their target. Regarding binding affinity to the target sequence, compared to PPRcag_1, PPRcag_2 is at the same level, PPRcag_3 is approximately 80%, PPRcag_4 is approximately 60%, PPRcag_5 is approximately 120%, and PPRcag_6 is approximately 130%.
[0155] Figure 4 This study investigated the effect of the first PPR motif (from the N-terminus) on aggregation. PPR proteins were prepared and purified using an *E. coli* expression system, and separated by gel filtration chromatography. Smaller elution volumes indicated larger molecular sizes. V2 eluted in 8–10 mL elution volumes, while V3.2 showed peaks in 12–14 mL elution volumes. This suggests that V2, with its larger protein size, implied a higher likelihood of aggregation, which was improved in V3.2. Detailed Implementation
[0156] [PPR motif, PPR protein]
[0157] (definition)
[0158] In this invention, when referring to a PPR motif, unless otherwise specified, it refers to a polypeptide consisting of 30 to 38 amino acids having the following amino acid sequence, wherein the amino acid sequence, when analyzed using a protein domain search program on the internet, has an E value below a specified value (preferably E-03) obtained using PF01535 in Pfam and PS51375 in Prosite. The position number of the amino acid constituting the PPR motif as defined in this invention is substantially synonymous with PF01535, and on the other hand, it is equivalent to the number obtained by subtracting 2 from the position of the amino acid in PS51375 (for example, position 1 in this invention → position 3 in PS51375). When referring to the amino acid at position "ii" (-2), it is the second amino acid from the last position (C-terminal side) of the amino acid constituting the PPR motif; or the -2nd amino acid, counting two positions from the N-terminus relative to position 1 of the next PPR motif. In the absence of a clearly identified next PPR motif, the amino acid two positions preceding the first amino acid of the next helical structure is designated as "ii". For information about Pfam, please refer to http: / / pfam.sanger.ac.uk / . For information about Prosite, please refer to http: / / www.expasy.org / prosite / .
[0159] The conserved amino acid sequence of the PPR motif is low at the amino acid level, but the two α-helices in the secondary structure are highly conserved. A representative PPR motif consists of 35 amino acids, but its length varies in the range of 30 to 38 amino acids.
[0160] More specifically, the PPR motif mentioned in this invention is composed of a polypeptide of 30 to 38 amino acids in length as represented by Formula 1.
[0161] [Chemistry 7]
[0162] (Helix A)-X-(Helix B)-L (Equation 1)
[0163] In the formula:
[0164] Helix A is a portion of 12 amino acids in length capable of forming an α-helix structure, represented by Equation 2.
[0165] [Chemistry 8]
[0166] A1-A2-A3-A4-A5-A6-A7-A8-A g -A 10-A 11 -A 12 (Equation 2)
[0167] In Equation 2, A1~A 12 Each amino acid can be represented independently;
[0168] X is absent or is a part consisting of 1 to 9 amino acids in length;
[0169] Helix B is a portion composed of 11 to 13 amino acids in length that can form an α-helix structure;
[0170] L is the portion represented by Equation 3, which consists of 2 to 7 amino acids in length;
[0171] [Chemistry 9]
[0172] L vii -L vi -Lv-L lv -L iii -L ii -L i (Equation 3)
[0173] In Formula 3, each amino acid is numbered "i" (-1) and "ii" (-2) starting from the C-terminus.
[0174] Among them, L iii ~L vii There are cases where it does not exist.
[0175] When PPR proteins are mentioned in this invention, unless specifically stated otherwise, it refers to PPR proteins having one or more, preferably two or more, of the aforementioned PPR motifs. When proteins are mentioned in this specification, unless specifically stated otherwise, it refers to all substances composed of polypeptides (chains of two or more amino acids linked by peptide bonds), including substances composed of lower molecular weight polypeptides. When amino acids are mentioned in this invention, sometimes it refers to ordinary amino acid molecules, and sometimes it refers to amino acid residues that constitute a peptide chain. Those skilled in the art will understand which term is being referred to based on the context.
[0176] In this invention, when referring to the binding activity of the PPR motif to the bases in the target nucleic acid, unless otherwise specified, it means that the binding activity to any one of the four bases is higher than the binding activity to the other bases.
[0177] In this invention, when nucleic acids are mentioned, it refers to RNA or DNA. It should be noted that PPR proteins can be specific to bases in RNA or DNA, but they do not bind to nucleic acid monomers.
[0178] The combination of the three amino acids at positions 1, 4, and ii of the PPR motif is important for the specific binding to a base, and their combination determines which base is being bound (see Patent Documents 1 and 2 above).
[0179] Specifically, for the RNA-binding PPR motif, the relationship between the combination of the three amino acids at positions 1, 4, and ii and the bases that can be bound is as follows (see Patent Document 1 above).
[0180] (3-1) A1, A4 and L ii When the combination of these three amino acids is valine, asparagine, and aspartic acid in sequence, this PPR motif has a selective RNA base binding ability, which is to bind strongly to U, then to C, and then to A or G.
[0181] (3-2)A1, A4 and L ii When the combination of these three amino acids is valine, threonine, and asparagine in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to A, then to G, then to C, but not to U.
[0182] (3-3)A1, A4 and L ii When the combination of these three amino acids is valine, asparagine, and asparagine in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to C, then binds to A or U, but does not bind to G.
[0183] (3-4)A1, A4 and L ii When the combination of these three amino acids is glutamic acid, glycine, and aspartic acid in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to G but does not bind to A, U, or C.
[0184] (3-5)A1, A4 and L ii When the combination of these three amino acids is isoleucine, asparagine, and asparagine in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to C, then to U, then to A, but not to G.
[0185] (3-6)A1, A4 and L ii When the combination of these three amino acids is valine, threonine, and aspartic acid in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to G, then to U, but not to A and C.
[0186] (3-7)A1, A4 and L iiWhen the combination of these three amino acids is lysine, threonine, and aspartic acid in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to G, then binds to A, but does not bind to U or C.
[0187] (3-8)A1, A4 and L ii When the combination of these three amino acids is phenylalanine, serine, and asparagine in sequence, this PPR motif has a selective RNA base binding ability, which is strong with A, followed by C, and then G and U.
[0188] (3-9)A1, A4 and L ii When the combination of these three amino acids is valine, asparagine, and serine in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to C, then to U, but not to A and G.
[0189] (3-10)A1, A4 and L ii When the combination of these three amino acids is phenylalanine, threonine, and asparagine, the PPR motif has a selective RNA base binding ability that strongly binds to A but not to G, U, or C.
[0190] (3-11) A1, A4 and L ii When the combination of these three amino acids is isoleucine, asparagine, and aspartic acid, the PPR motif exhibits selective RNA base binding ability, with strong binding to U, followed by binding to A, but not binding to G or C.
[0191] (3-12)A1, A4 and L ii When the combination of these three amino acids is threonine, threonine, and asparagine, the PPR motif has a selective RNA base binding ability that strongly binds to A but not to G, U, and C.
[0192] (3-13)A1, A4 and L ii When the combination of these three amino acids is isoleucine, methionine, and aspartic acid, the PPR motif exhibits selective RNA base binding ability, with strong binding to U, followed by binding to C, but not binding to A and G.
[0193] (3-14) A1, A4 and L ii When the combination of these three amino acids is phenylalanine, proline, and aspartic acid in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to U, then binds to C, but does not bind to A and G.
[0194] (3-15)A1, A4 and Lii When the combination of these three amino acids is tyrosine, proline, and aspartic acid in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to U but not to A, G, and C.
[0195] (3-16)A1, A4 and L ii When the combination of these three amino acids is leucine, threonine, and aspartic acid in sequence, this PPR motif has a selective RNA base binding ability that strongly binds to G but not to A, U, and C.
[0196] For the PPR motif for DNA binding, the relationship between the combination of the three amino acids at positions 1, 4, and ii and the bases that can be bound is as follows (see Patent Document 2 above).
[0197] (2-1) A1, A4 and L ii When the combination of these three amino acids is any amino acid, glycine, and aspartic acid in sequence, the PPR motif selectively binds to G.
[0198] (2-2)A1, A4 and L ii When the combination of these three amino acids is glutamic acid, glycine, and aspartic acid in sequence, the PPR motif selectively binds to G.
[0199] (2-3) A1, A4 and L ii When the combination of these three amino acids is any amino acid, glycine, and asparagine in sequence, the PPR motif selectively binds to A.
[0200] (2-4)A1, A4 and L ii When the combination of these three amino acids is glutamic acid, glycine, and asparagine in sequence, the PPR motif selectively binds to A.
[0201] (2-5)A1, A4 and L ii When the combination of these three amino acids is any amino acid, glycine, and serine in sequence, the PPR motif selectively binds to A and then to C.
[0202] (2-6)A1, A4 and L ii When the combination of these three amino acids is any amino acid, isoleucine, and any amino acid in sequence, the PPR motif selectively binds to T and C.
[0203] (2-7)A1, A4 and L ii When the combination of these three amino acids is any amino acid, isoleucine, and asparagine in sequence, the PPR motif selectively binds to T and then to C.
[0204] (2-8)A1, A4 and L ii When the combination of these three amino acids is any amino acid, leucine, and any amino acid in sequence, the PPR motif selectively binds to T and C.
[0205] (2-9)A1, A4 and L ii When the combination of these three amino acids is any amino acid, leucine, and aspartic acid in sequence, the PPR motif selectively binds to C.
[0206] (2-10)A1, A4 and L ii When the combination of these three amino acids is any amino acid, leucine, and lysine in sequence, the PPR motif selectively binds to T.
[0207] (2-11) A1, A4 and L ii When the combination of these three amino acids is any amino acid, methionine, and any amino acid in sequence, the PPR motif selectively binds to T.
[0208] (2-12)A1, A4 and L ii When the combination of these three amino acids is any amino acid, methionine, and aspartic acid in sequence, the PPR motif selectively binds to T.
[0209] (2-13) A1, A4 and L ii When the combination of these three amino acids is isoleucine, methionine, and aspartic acid, the PPR motif selectively binds to T and then to C.
[0210] (2-14) A1, A4 and L ii When the combination of these three amino acids is any amino acid, asparagine, and any amino acid in sequence, the PPR motif selectively binds to C and T.
[0211] (2-15) A1, A4 and L ii When the combination of these three amino acids is any amino acid, asparagine, and aspartic acid in sequence, the PPR motif selectively binds to T.
[0212] (2-16) A1, A4 and L ii When the combination of these three amino acids is phenylalanine, asparagine, and aspartic acid in sequence, the PPR motif selectively binds to T.
[0213] (2-17)A1, A4 and L ii When the combination of these three amino acids is glycine, asparagine, and aspartic acid in sequence, the PPR motif selectively binds to T.
[0214] (2-18) A1, A4 and L iiWhen the combination of these three amino acids is isoleucine, asparagine, and aspartic acid in sequence, the PPR motif selectively binds to T.
[0215] (2-19) A1, A4 and L ii When the combination of these three amino acids is threonine, asparagine, and aspartic acid in sequence, the PPR motif selectively binds to T.
[0216] (2-20)A1, A4 and L ii When the combination of these three amino acids is valine, asparagine, and aspartic acid in sequence, the PPR motif selectively binds to T and then to C.
[0217] (2-21)A1, A4 and L ii When the combination of these three amino acids is tyrosine, asparagine, and aspartic acid in sequence, the PPR motif selectively binds to T and then to C.
[0218] (2-22)A1, A4 and L ii When the combination of these three amino acids is any amino acid, asparagine, and asparagine in sequence, the PPR motif selectively binds to C.
[0219] (2-23)A1, A4 and L ii When the combination of these three amino acids is isoleucine, asparagine, and asparagine in sequence, the PPR motif selectively binds to C.
[0220] (2-24) A1, A4 and L ii When the combination of these three amino acids is serine, asparagine, and asparagine in sequence, the PPR motif selectively binds to C.
[0221] (2-25) A1, A4 and L ii When the combination of these three amino acids is valine, asparagine, and asparagine in sequence, the PPR motif selectively binds to C.
[0222] (2-26) A1, A4 and L ii When the combination of these three amino acids is any amino acid, asparagine, and serine in sequence, the PPR motif selectively binds to C.
[0223] (2-27)A1, A4 and L ii When the combination of these three amino acids is valine, asparagine, and serine in sequence, the PPR motif selectively binds to C.
[0224] (2-28) A1, A4 and L ii When the combination of these three amino acids is any amino acid, asparagine, and threonine in sequence, the PPR motif selectively binds to C.
[0225] (2-29) A1, A4 and L ii When the combination of these three amino acids is valine, asparagine, and threonine in sequence, the PPR motif selectively binds to C.
[0226] (2-30)A1, A4 and L ii When the combination of these three amino acids is any amino acid, asparagine, and tryptophan in sequence, the PPR motif selectively binds to C and then to T.
[0227] (2-31)A1, A4 and L ii When the combination of these three amino acids is isoleucine, asparagine, and tryptophan in sequence, the PPR motif selectively binds to T and then to C.
[0228] (2-32)A1, A4 and L ii When the combination of these three amino acids is any amino acid, proline, and any amino acid in sequence, the PPR motif selectively binds to T.
[0229] (2-33)A1, A4 and L ii When the combination of these three amino acids is any amino acid, proline, and aspartic acid in sequence, the PPR motif selectively binds to T.
[0230] (2-34)A1, A4 and L ii When the combination of these three amino acids is phenylalanine, proline, and aspartic acid in sequence, the PPR motif selectively binds to T.
[0231] (2-35)A1, A4 and L ii When the combination of these three amino acids is tyrosine, proline, and aspartic acid in sequence, the PPR motif selectively binds to T.
[0232] (2-36)A1, A4 and L ii When the combination of these three amino acids is any amino acid, serine, and any amino acid in sequence, the PPR motif selectively binds to A and G.
[0233] (2-37)A1, A4 and L ii When the combination of these three amino acids is any amino acid, serine, and asparagine in sequence, the PPR motif selectively binds to A.
[0234] (2-38)A1, A4 and L ii When the combination of these three amino acids is phenylalanine, serine, and asparagine in sequence, the PPR motif selectively binds to A.
[0235] (2-39)A1, A4 and Lii When the combination of these three amino acids is valine, serine, and asparagine in sequence, the PPR motif selectively binds to A.
[0236] (2-40)A1, A4 and L ii When the combination of these three amino acids is any amino acid, threonine, and any amino acid in sequence, the PPR motif selectively binds to A and G.
[0237] (2-41)A1, A4 and L ii When the combination of these three amino acids is any amino acid, threonine, and aspartic acid in sequence, the PPR motif selectively binds to G.
[0238] (2-42)A1, A4 and L ii When the combination of these three amino acids is valine, threonine, and aspartic acid in sequence, the PPR motif selectively binds to G.
[0239] (2-43)A1, A4 and L ii When the combination of these three amino acids is any amino acid, threonine, and asparagine in sequence, the PPR motif selectively binds to A.
[0240] (2-44)A1, A4 and L ii When the combination of these three amino acids is phenylalanine, threonine, and asparagine in sequence, the PPR motif selectively binds to A.
[0241] (2-45)A1, A4 and L ii When the combination of these three amino acids is isoleucine, threonine, and asparagine in sequence, the PPR motif selectively binds to A.
[0242] (2-46)A1, A4 and L ii When the combination of these three amino acids is valine, threonine, and asparagine in sequence, the PPR motif selectively binds to A.
[0243] (2-47)A1, A4 and L ii When the combination of these three amino acids is any amino acid, valine, and any amino acid in sequence, the PPR motif binds to A, C, and T, but not to G.
[0244] (2-48)A1, A4 and L ii When the combination of these three amino acids is isoleucine, valine, and aspartic acid, the PPR motif selectively binds to C and then to A.
[0245] (2-49)A1, A4 and L iiWhen the combination of these three amino acids is any amino acid, valine, and glycine in sequence, the PPR motif selectively binds to C.
[0246] (2-50)A1, A4 and L ii When the combination of these three amino acids is any amino acid, valine, and threonine in sequence, the PPR motif selectively binds to T; proteins determined by the above relationship have selective DNA base binding ability.
[0247] (A particularly preferred combination of 3 amino acids)
[0248] In the RNA-binding PPR motif, there are representative combinations of amino acids at positions 1, 4, and ii that recognize and specifically bind to each base. Specifically, in the combination recognizing adenine, position 1 is valine, position 4 is threonine, and position ii is asparagine; in the combination recognizing cytosine, position 1 is valine, position 4 is asparagine, and position ii is serine; in the combination recognizing guanine, position 1 is valine, position 4 is threonine, and position ii is aspartic acid; and in the combination recognizing uracil, position 1 is valine, position 4 is asparagine, and position ii is aspartic acid (as described in Non-Patent Literature 1-5 above). These combinations are used in a preferred embodiment of the present invention.
[0249] (Improvement in cohesion)
[0250] Based on the amino acid information of naturally occurring PPR motifs, the inventors discovered that it is very common for the amino acid at position 6 of the PPR motif to be hydrophobic (especially leucine) and the amino acid at position 9 to be non-hydrophilic (especially glycine). According to the structures of PPR proteins with known crystal structures (Non-Patent Literature 6: Coquille et al., 2014 Nat. Commun.; PDB ID: 4PJQ, 4WN4, 4WSL, 4PJR; Non-Patent Literature 7: Shenet et al., 2015 Nat. Commun., PDB ID: 5I9D, 5I9F, 5I9G, 5I9H), positions 6 and 9 of the first motif (N-terminal side) are exposed on the outside. Therefore, it is conceivable that these exposed hydrophobic amino acids would contribute to aggregation (…). Figure 1 A). On the other hand, after the second motif, the amino acids at positions 6 and 9 are embedded within the protein, forming a hydrophobic core. Therefore, it is believed that adding hydrophilic residues at positions 6 and 9 of all motifs might disrupt the protein structure. Thus, it was decided to reduce the cohesiveness of the PPR by using only hydrophilic amino acids (asparagine, aspartic acid, glutamine, glutamic acid, lysine, arginine, serine, threonine) at positions 6, preferably 6 and 9 of the first motif.
[0251] Specifically, as described below.
[0252] In proteins capable of binding to target nucleic acids with a specific base sequence, in the first PPR motif (M1) from the N-terminus:
[0253] (1) Make A6 amino acid a hydrophilic amino acid, preferably make A6 amino acid asparagine or aspartic acid.
[0254] (2) The A9 amino acid is then made to be a hydrophilic amino acid or glycine, preferably glutamine, glutamic acid, lysine or glycine.
[0255] (3) Or make A6 amino acid and A9 amino acid any of the following combinations.
[0256] • The combination of amino acid A6 being asparagine and amino acid A9 being glutamic acid
[0257] • The combination of amino acid A6 being asparagine and amino acid A9 being glutamine
[0258] • The combination of amino acid A6 being asparagine and amino acid A9 being lysine
[0259] • The combination of A6 amino acid being aspartic acid and A9 amino acid being glycine.
[0260] (New PPR motif)
[0261] This invention provides a novel PPR motif that improves cohesion, discovered through the above-described scheme, and a novel PPR protein containing the motif.
[0262] The novel PPR motif provided by this invention is as follows.
[0263] (C-1) A PPR motif consisting of any sequence from sequence numbers 4 to 7;
[0264] (C-2) is a cytosine-binding PPR motif consisting of 1 to 9 amino acids other than amino acids at positions 1, 4, 6 and 34 in any of the sequences 4 to 7, which have been replaced, deleted or added.
[0265] (C-3) has at least 80% sequence identity with any of the sequences in sequence numbers 4 to 7, wherein the amino acids at positions 1, 4, 6 and 34 are the same and are cytosine-binding PPR motifs;
[0266] (A-1) The PPR motif consisting of sequence number 8 has an amino acid at position 6 replaced with an asparagine or aspartic acid.
[0267] (A-2) is a sequence consisting of 1 to 9 amino acids other than positions 1, 4, 6 and 34 in the sequence of (A-1) that have been replaced, deleted or added, and is an adenine-binding PPR motif.
[0268] (A-3) and (A-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6 and 34 being identical, and both being adenine-binding PPR motifs;
[0269] (G-1) The PPR motif consisting of sequence number 9 has an amino acid at position 6 replaced by asparagine or aspartic acid.
[0270] (G-2) is a guanine-binding PPR motif consisting of 1 to 9 amino acids other than positions 1, 4, 6 and 34 in the sequence of (G-1) that have been replaced, deleted or added.
[0271] (G-3) and (G-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6 and 34 being identical and being guanine-binding PPR motifs;
[0272] (U-1) The PPR motif consisting of sequence number 10 has an amino acid at position 6 replaced by asparagine or aspartic acid.
[0273] (U-2) is a uracil-binding PPR motif consisting of 1 to 9 amino acids other than amino acids 1, 4, 6 and 34 in the sequence of (U-1) that have been replaced, deleted or added.
[0274] The sequences of (U-3) and (U-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6 and 34 being identical, and both being uracil-binding PPR motifs.
[0275] Among such PPR motifs, the following PPR motifs are particularly preferred.
[0276] (C-1) A PPR motif consisting of any sequence from sequence numbers 4 to 7;
[0277] (C-2) is a cytosine-binding PPR motif consisting of 1 to 9 amino acids other than amino acids 1, 4, 6, 9 and 34 in any of the sequences in sequence numbers 4 to 7, which have been replaced, deleted or added.
[0278] (C-3) has at least 80% sequence identity with any of the sequences in sequence numbers 4 to 7, wherein the amino acids at positions 1, 4, 6, 9 and 34 are the same and are cytosine-binding PPR motifs;
[0279] (A-1) A PPR motif in sequence number 8 in which amino acids at positions 6 and 9 are substituted in any combination that satisfies the following definition;
[0280] (A-2) is a PPR motif that consists of 1 to 9 amino acids other than amino acids 1, 4, 6, 9 and 34 in the sequence of (A-1) that have been replaced, deleted or added, and is adenine-binding.
[0281] (A-3) and (A-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6, 9 and 34 being identical and being adenine-binding PPR motifs;
[0282] (G-1) A PPR motif consisting of a sequence in which amino acids at positions 6 and 9 of sequence number 9 are substituted in any combination that satisfies the following definition.
[0283] (G-2) is a guanine-binding PPR motif consisting of 1 to 9 amino acids other than positions 1, 4, 6, 9 and 34 in the sequence of (G-1) through substitution, deletion or addition.
[0284] (G-3) and (G-1) have at least 80% sequence identity, with amino acids at positions 1, 4, 6, 9 and 34 being identical and being guanine-binding PPR motifs;
[0285] (U-1) A PPR motif consisting of a sequence in which amino acids at positions 6 and 9 of sequence number 10 are substituted in any combination that satisfies the following definition.
[0286] (U-2) is a uracil-binding PPR motif consisting of 1 to 9 amino acids other than positions 1, 4, 6, 9 and 34 in the sequence of (U-1) through substitution, deletion or addition.
[0287] The sequences of (U-3) and (U-1) have at least 80% sequence identity, with the amino acids at positions 1, 4, 6, 9 and 34 being identical, and both being uracil-binding PPR motifs.
[0288] • The amino acid at position 6 is asparagine, and the amino acid at position 9 is glutamic acid.
[0289] • The amino acid at position 6 is asparagine, and the amino acid at position 9 is glutamine.
[0290] • The amino acid at position 6 is asparagine, and the amino acid at position 9 is lysine.
[0291] • The amino acid at position 6 is aspartic acid and the amino acid at position 9 is glycine.
[0292] The specific sequences of serial numbers 4 to 10 are as follows: Figure 1 And as shown in the sequence list.
[0293] Among such PPR motifs, the following PPR motifs are even more preferred.
[0294] (C-4) The PPR motif consisting of the sequence number 4;
[0295] (A-4) The PPR motif consisting of the sequence number 58;
[0296] (G-4) The PPR motif consisting of sequence number 59;
[0297] (U-4) is a PPR motif consisting of the sequence number 60.
[0298] The sequences of serial numbers 58 to 60 are shown in the following sequences and sequence list.
[0299] Serial number 58, sequence VTYTTNIDQLCKAGKVDEALELFKEMRSKGVKPNV
[0300] Serial number 59, sequence VTYTTNIDQLCKAGKVDEALELFDEMKERGIKPDV
[0301] Serial number 60, sequence VTYNTNIDQLCKAGRLDEAEELLEEMEEKGIKPDV
[0302] (PPR protein with improved cohesiveness)
[0303] The present invention also provides a PPR protein with improved cohesiveness discovered through the above-described scheme.
[0304] In a preferred embodiment, regarding the A9 amino acid of M1, regardless of the other amino acids in M1 and regardless of the amino acid sequence of the motif other than M1, the A9 amino acid of M1 is always a non-hydrophobic amino acid or glycine. The non-hydrophobic amino acid is a hydrophilic amino acid, or cysteine or histidine; preferably a hydrophilic amino acid, i.e., arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, or threonine; more preferably glutamine, glutamic acid, or lysine.
[0305] In a preferred embodiment, the A9 amino acid of M1 is glutamine, glutamic acid, lysine, or glycine, regardless of the other amino acids in M1 and the amino acid sequence of motifs other than M1.
[0306] In a preferred embodiment, the A6 amino acid of M1 is a non-hydrophobic amino acid, regardless of the other amino acids in M1 and the amino acid sequence of the motif other than M1. The non-hydrophobic amino acid is, for example, a hydrophilic amino acid, or cysteine or histidine; preferably, a hydrophilic amino acid, i.e., arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, or threonine; more preferably, asparagine or aspartic acid.
[0307] In a particularly preferred embodiment, regarding the A6 and A9 amino acids of M1, regardless of the other amino acids of M1 and regardless of the amino acid sequence of the motif other than M1, the A6 and A9 amino acids of M1 are any of the following combinations:
[0308] • The combination of amino acid A6 being asparagine and amino acid A9 being glutamic acid
[0309] • The combination of amino acid A6 being asparagine and amino acid A9 being glutamine
[0310] • The combination of amino acid A6 being asparagine and amino acid A9 being lysine
[0311] • The combination of A6 amino acid being aspartic acid and A9 amino acid being glycine.
[0312] In a preferred embodiment of an RNA-binding protein, the A6 and A9 amino acids of M1 satisfy the above conditions, and at least one, preferably more than half, and more preferably all of the PPR motifs contained therein satisfy any of the following conditions:
[0313] • When the bound base is cytosine, A1 is valine, A4 is asparagine, and A... ii For serine
[0314] • When the bound base is adenine, A1 is valine, A4 is threonine, and A ii Asparagine
[0315] • When the bound base is guanine, A1 is valine, A4 is threonine, and A... ii Aspartic acid
[0316] • When the bound base is uracil or thymine, A1 is valine, A4 is asparagine, and A... ii Aspartic acid
[0317] In a preferred manner for RNA-binding proteins, M1 is the aforementioned novel PPR motif.
[0318] In a particularly preferred embodiment, M1 is a PPR motif composed of any of the following polypeptides:
[0319] • When the bound base is cytosine, the polypeptide consisting of any sequence in SEQ ID NOs: 4-7
[0320] • In the case where the bound base is adenine, the polypeptide in which amino acids at positions 6 and 9 of SEQ ID NO: 8 are substituted in any combination satisfying the combination defined in the following paragraphs.
[0321] • In the case where the bound base is guanine, a polypeptide in which amino acids at positions 6 and 9 of SEQ ID NO: 9 are substituted in any combination satisfying the combination defined in the following paragraph.
[0322] • In the case where the bound base is uracil, a polypeptide in which amino acids at positions 6 and 9 of SEQ ID NO: 10 are substituted in any combination satisfying the combinations defined in the following paragraphs.
[0323] At least one of the PPR motifs other than M1 is a PPR motif composed of any of the following polypeptides:
[0324] • When the bound base is cytosine, the polypeptide consisting of the sequence of SEQ ID NO: 2
[0325] • When the bound base is adenine, the polypeptide consisting of the sequence of SEQ ID NO: 8
[0326] • When the bound base is guanine, the polypeptide consisting of the sequence of SEQ ID NO: 9
[0327] • When the bound base is uracil, the polypeptide consisting of the sequence of SEQ ID NO: 10 is...
[0328] The combination mentioned in the above paragraph is any one of the following combinations:
[0329] • The combination of amino acid A6 being asparagine and amino acid A9 being glutamic acid
[0330] • The combination of amino acid A6 being asparagine and amino acid A9 being glutamine
[0331] • The combination of amino acid A6 being asparagine and amino acid A9 being lysine
[0332] • The combination of A6 amino acid being aspartic acid and A9 amino acid being glycine.
[0333] In a particularly preferred embodiment, M1 is a PPR motif composed of any of the following polypeptides:
[0334] • When the bound base is cytosine, the polypeptide consisting of the sequence of SEQ ID NOs: 4
[0335] • When the bound base is adenine, the polypeptide consisting of the sequence of SEQ ID NO: 58
[0336] • When the bound base is guanine, the polypeptide consisting of the sequence of SEQ ID NO: 59
[0337] • When the bound base is uracil, the polypeptide consisting of the sequence of SEQ ID NO: 60
[0338] At least one of the PPR motifs other than M1 is a PPR motif composed of any of the following polypeptides:
[0339] • When the bound base is cytosine, the polypeptide consisting of the sequence of SEQ ID NO: 2
[0340] • When the bound base is adenine, the polypeptide is formed by replacing amino acid position 15 of SEQ ID NO: 8 with lysine.
[0341] • When the bound base is guanine, the polypeptide consisting of the sequence of SEQ ID NO: 9
[0342] • When the bound base is uracil, the polypeptide consisting of the sequence of SEQ ID NO: 10 is...
[0343] (Application of high-performance PPR motif backbones)
[0344] In a preferred embodiment of the invention, in the PPR motif targeting each base of cytosine, adenine, guanine, and uracil (or thymine), the amino acids other than positions 1, 4, 6, 9, and ii can be specific amino acids. Specifically, when collecting sequences from Arabidopsis thaliana PPR motifs where the amino acid combinations at positions 1, 4, and ii form VTN (to become a PPR motif for recognizing adenine), VSN (to become a PPR motif for recognizing cytosine), VTD (to become a PPR motif for recognizing guanine), and VND (to become a PPR motif for recognizing uracil), and summarizing the types and numbers of amino acids appearing at each position, the performance of the PPR motif can be improved by selecting amino acids that appear frequently at each position.
[0345] To obtain an RNA-binding PPR protein, the amino acid sequence of the PPR motif below can be referenced from the perspective of making the amino acids other than those at positions 1, 4, 6, 9, and ii appear at a high frequency as described above.
[0346] As a PPR motif corresponding to cytosine, a PPR motif consisting of any of the sequences in SEQ ID NOs: 4-7;
[0347] As a PPR motif corresponding to adenine, a PPR motif consisting of any sequence in SEQ ID NO: 8;
[0348] As a PPR motif corresponding to guanine, a PPR motif consisting of any sequence in SEQ ID NO: 9;
[0349] As a PPR motif corresponding to uracil, a PPR motif is formed by any of the sequences in SEQ ID NO: 10.
[0350] (Explanation of terminology, etc.)
[0351] In this invention, when referring to "identity" in relation to base sequences (sometimes also called nucleotide sequences) or amino acid sequences, unless specifically stated otherwise, it refers to the percentage of identical bases or amino acids shared between two sequences when they are aligned in the optimal manner. That is, it can be calculated using the formula: Identity = (Number of identical positions / Total number of positions) × 100, and can be calculated using commercially available algorithms. Furthermore, such algorithms are incorporated into the NBLAST and XBLAST programs described in Altschul et al., J. Mol. Biol. 215 (1990) 403-410. More specifically, the relevant retrieval / analysis of the identity of base sequences or amino acid sequences can be performed by those skilled in the art using well-known algorithms or programs (e.g., BLASTN, BLASTP, BLASTX, ClustalW). The parameters used when using the program can be appropriately set by those skilled in the art, or the default parameters of each program can be used. The specific methods of these analyses are also well-known to those skilled in the art.
[0352] In this specification, when referring to the identity (or high identity) of a base sequence or amino acid sequence, unless otherwise specified, it means that the identity is at least 70%, preferably 80% or more, more preferably 85% or more, further preferably 90% or more, even more preferably 95% or more, even more preferably 97.5% or more, even more preferably 99% or more.
[0353] Furthermore, in this invention, when referring to "replacement, deletion, or addition of sequences" in relation to PPR motifs or proteins, the number of amino acids involved in the substitutions, etc., is not particularly limited, unless specifically stated otherwise. In any motif or protein, as long as the motif or protein composed of that amino acid sequence has the desired function, it is not particularly limited to about 1 to 9 or 1 to 4, or, if the substitution is with an amino acid of similar properties, it may have a greater number of substitutions, etc. The means for preparing polynucleotides or proteins associated with such amino acid sequences are well known to those skilled in the art.
[0354] Amino acids with similar properties refer to amino acids with similar physical properties such as hydrophobicity / hydrophilicity, charge, pKa, and solubility, for example, the following amino acids.
[0355] Hydrophobic amino acids: alanine, valine, glycine, isoleucine, leucine, phenylalanine, proline, tryptophan, tyrosine
[0356] Non-hydrophobic amino acids: arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine, cysteine, histidine;
[0357] Hydrophilic amino acids: arginine, asparagine, aspartic acid, glutamic acid, glutamine, lysine, serine, threonine;
[0358] Acidic amino acids: aspartic acid, glutamic acid;
[0359] Basic amino acids: lysine, arginine, histidine;
[0360] Neutral amino acids: alanine, asparagine, cysteine, glutamine, glycine, isoleucine, leucine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine;
[0361] Sulfur-containing amino acids: methionine, cysteine;
[0362] Aromatic ring amino acids: tyrosine, tryptophan, phenylalanine.
[0363] Regarding genes, nucleic acids, polynucleotides, proteins, and motifs, "making" can be replaced with "producing" or "manufacturing." Additionally, regarding genes, when creating them by combining parts, it is sometimes called "constructing," and "constructing" can also be replaced with "producing" or "manufacturing."
[0364] The PPR motif of the present invention, the protein containing the PPR motif, or the nucleic acid encoding the thereof can be manufactured by those skilled in the art using existing technology and the descriptions in the embodiments of this specification.
[0365] [Characteristics and uses of PPR protein]
[0366] (Improved PPR protein aggregation)
[0367] The PPR protein prepared using the novel PPR motif of this invention exhibits reduced intracellular aggregation. The aggregation of the PPR protein can be evaluated by those skilled in the art by expressing the PPR protein in cells and confirming the presence or absence of aggregation. Confirmation is even easier if the PPR protein is expressed by fusing it with a fluorescent protein. According to the inventors' research, by appropriately altering the amino acids in the first motif of the PPR protein, the intracellular aggregation of the PPR protein is improved, and its transport to the nucleus is enhanced.
[0368] (binding force)
[0369] PPR proteins prepared using the novel PPR motif of this invention not only exhibit reduced intracellular aggregation but also possess equal or higher RNA binding performance than PPR proteins prepared using existing PPR motifs targeting the same RNA. "Equal" means 55% or more, preferably around 75%.
[0370] The binding affinity to target sequences can be evaluated using EMSA (Electrotrophoretic Mobility Shift Assay) or Biacore methods. EMSA utilizes the property that the migration rate of nucleic acid molecules changes during electrophoresis of samples containing bound proteins and nucleic acids compared to unbound samples. Intermolecular interaction analysis equipment, such as Biacore, can perform reaction rate-based analysis, thus enabling detailed protein-nucleic acid binding analysis.
[0371] The binding affinity to the target sequence can also be analyzed by supplying a solution containing the candidate protein to the immobilized target nucleic acid and detecting or quantifying the protein bound to the target nucleic acid. This method utilizes ELISA (Enzyme-Linked Immunosorbent Assay) and is therefore sometimes called RPB-ELISA (RNA-protein binding ELISA). Specifically, the step of supplying a solution containing the candidate protein to the immobilized target nucleic acid can be performed by passing a solution containing the target binding protein through the target nucleic acid molecules immobilized on a plate. Immobilization of the target nucleic acid molecules can be achieved using various existing immobilization methods, for example, by supplying a biotin-modified target nucleic acid molecule nucleic acid probe to a streptavidin-coated well plate. Detailed experimental conditions can be found in the experimental methods described in the embodiments of this invention. In RPB-ELISA, the binding affinity between the target PPR protein and its target RNA can be obtained by subtracting the background signal (the luminescence signal value when the target PPR protein is added but not the target RNA) from the luminescence intensity of the sample containing the target PPR protein and its target RNA.
[0372] [Applications of PPR protein]
[0373] (Complex, Fusion Protein)
[0374] The PPR motifs or PPR proteins provided by this invention can be linked to functional regions to form complexes. They can also be linked to protein-like functional regions to form fusion proteins. A functional region refers to a portion within an organism or cell that has a specific biological function, such as enzyme function, catalytic function, inhibitory function, or hyperactive function, or a portion that serves as a marker. Such regions are, for example, composed of proteins, peptides, nucleic acids, physiologically active substances, or pharmaceuticals. It should be noted that the following description uses fusion proteins as an example to illustrate the complexes of this invention, but those skilled in the art can understand the situation of complexes other than fusion proteins based on this description.
[0375] In a preferred embodiment, the functional region is a ribonuclease (RNase). Examples of RNases include RNase A (e.g., bovine pancreatic ribonuclease A: PDB 2AAS) and RNase H.
[0376] In a preferred embodiment, the functional region is a fluorescent protein. Examples of fluorescent proteins are mCherry, EGFP, GFP, Sirius, EBFP, ECFP, mTurquoise, TagCFP, AmCyan, mTFP1, MidoriishiCyan, CFP, TurboGFP , AcGFP, TagGFP, Azami-Green, ZsGreen, EmGFP, HyPer, TagYFP, EYFP, Venus, YFP, PhiYFP, PhiYFP-m, TurboYFP, ZsYellow , mBanana, KusabiraOrange, mOrange, TurboRFP, DsRed-Express, DsRed2, TagRFP, DsRed-Monomer, AsRed2, mStrawberry, TurboFP602, mRFP1, JRed, KillerRed, HcRed, KeimaRed, mRasberry, mPlum, PS-CFP, Dendra2, Kaede, EosFP, KikumeGR. From the perspective of improving the cohesiveness and / or efficient transport to the nucleus of fusion proteins, mClover3 is a preferred example.
[0377] In a preferred embodiment, when the target is mRNA, the functional region is a functional domain that enhances the expression level of protein from the target mRNA (WO2017 / 209122). Examples of functional domains that enhance the expression level of protein from mRNA include, for example, all or a functional portion of the functional domains of proteins known to directly or indirectly promote the translation of mRNA. More specifically, these could be domains that induce ribosomes to mRNA, domains associated with the initiation or promotion of mRNA translation, domains associated with the extranuclear transport of mRNA, domains associated with binding to the endoplasmic reticulum membrane, domains containing an endoplasmic reticulum retention signal sequence, or domains containing an endoplasmic reticulum signal sequence. More specifically, the aforementioned domains that induce ribosomes to mRNA can be domains containing all or a functional portion of peptides selected from the group consisting of DENR (Density-regulated protein), MCT-1 (Malignant T-cell amplified sequence 1), TPT1 (Translationally-controlled tumor protein), and Lerepo4 (Zinc finger CCCH domain). Additionally, the aforementioned domains related to mRNA translation initiation or promotion can be domains containing all or a functional portion of peptides selected from the group consisting of eIF4E and eIF4G. Furthermore, the aforementioned domains related to mRNA transnuclear transport can be domains containing all or a functional portion of SLBP (Stem-loop binding protein). Furthermore, the domains associated with binding to the endoplasmic reticulum membrane may be domains comprising all or a functional portion of a polypeptide selected from the group consisting of SEC61B, TRAP-alpha (translocon associated protein alpha), SR-alpha, Dia1 (cytochrome b5 reductase 3), and p180. Additionally, the ER retention signal sequence may be a signal sequence containing the KDEL (SEQ ID NO: 55) or KEEL (SEQ ID NO: 56) sequence. Furthermore, the ER signal sequence may be a signal sequence containing MGWSCIILFLVATATGAHS (SEQ ID NO: 57).
[0378] In this invention, the functional regions can be fused to the N-terminal side of the PPR protein, the C-terminal side, or both the N-terminal and C-terminal sides. Furthermore, the complex or fusion protein can contain more than two functional regions (e.g., two to five). Additionally, in the complex or fusion protein of this invention, the functional regions can be indirectly fused to the PPR protein via linkers or the like.
[0379] (Nucleic acids, vectors, and cells encoding PPR proteins, etc.)
[0380] This invention also provides nucleic acids encoding the aforementioned PPR motif, PPR protein, or fusion protein, and vectors containing the nucleic acids (e.g., vectors for amplification, expression vectors). Vectors for amplification can use *E. coli* or yeast as a host. In this specification, an expression vector refers to, for example, a vector comprising, from upstream, DNA with a promoter sequence, DNA encoding the desired protein, and DNA with a terminator sequence; however, this order is not necessarily required as long as the desired function is achieved. Various vectors commonly used by those skilled in the art can be recombined in this invention.
[0381] The PPR protein or fusion protein of the present invention can function in eukaryotic cells (e.g., animals, plants, microorganisms (yeast, etc.), protozoa). The fusion protein of the present invention particularly functions in animal cells (in vitro or in vivo). Examples of animal cells from humans, monkeys, pigs, cattle, horses, dogs, cats, mice, and rats that can be used as vectors for introducing or expressing the PPR protein or fusion protein of the present invention are: * **Cultural cells from Chinese hamster ovary (CHO) cells, COS-1 cells, COS-7 cells, VERO (ATCC CCL-81) cells, BHK cells, canine kidney-derived MDCK cells, hamster AV-12-664 cells, HeLa cells, WI38 cells, 293 cells, 293T cells, and PER.C6 cells, but are not limited to these.
[0382] (use)
[0383] The PPR protein or fusion protein of the present invention has the potential to deliver functional regions into organisms or cells in a nucleic acid sequence-specific manner to enable them to function. Complexes linked to marker motifs such as GFP can be used to visualize desired RNA in vivo.
[0384] Furthermore, the PPR proteins or fusion proteins of the present invention have the potential to be specifically modified / disrupted within cells or organisms, and to confer new functions. In particular, RNA-binding PPR proteins are involved in all steps of RNA processing observed in organelles, including cleavage, RNA editing, translation, splicing, and RNA stabilization. Therefore, the methods for modifying PPR proteins provided by the present invention, as well as the PPR motifs and PPR proteins provided by the present invention, are expected to have applications in various fields as described below.
[0385] (1) Medical
[0386] • Develop PPR proteins that recognize and bind to specific RNAs associated with specific diseases. Additionally, analyze the target sequence and accompanying proteins for each specific RNA. These analytical results can be used in the exploration of compounds for disease treatment.
[0387] For example, in animals, abnormalities in PPR proteins identified as LRPPRC are known to cause Leigh syndrom (LSFC); Leigh syndrome, subacute necrotizing encephalomyelopathy. This invention can aid in the management (prevention, treatment, and inhibition of progression) of LSFC. Most existing PPR proteins function as editing sites for designated RNA manipulations (transformation of genetic information on RNA; mostly C→U). This type of PPR protein possesses an additional motif at its C-terminus that suggests interaction with RNA editing enzymes. With PPR proteins having such a structure, it is expected that base polymorphisms can be introduced, or that diseases or conditions arising from base polymorphisms can be addressed.
[0388] • Create cells that regulate RNA inhibition / expression. Such cells include stem cells (e.g., iPS cells) used to monitor differentiated / undifferentiated states; model cells for cosmetic evaluation; and cells capable of switching functional RNA expression on / off for the purpose of elucidating drug development mechanisms or conducting pharmacological experiments.
[0389] • Create PPR proteins that specifically bind to specific RNAs associated with specific diseases. Introduce such PPR proteins into cells using plasmids, viral vectors, mRNA, or purified proteins. Within the cell, the PPR protein binds to its target RNA, thereby altering (improving) the function of the RNA that is the cause of the disease. Methods of altering function include, for example, changes in RNA structure caused by binding, gene knockdown due to degradation, changes in splicing responses due to splicing, and base substitution.
[0390] (2) Agriculture, forestry and water industry
[0391] • Improve yield and quality in agricultural, forestry, and aquatic products.
[0392] • Cultivate organisms with improved disease resistance, improved environmental tolerance, and enhanced or new functionalities.
[0393] For example, regarding first-generation hybrid (F1) crops, F1 crops can be artificially created by stabilizing or regulating the translation of mitochondrial RNA based on PPR proteins, potentially improving yield and quality. PPR protein-based RNA manipulation and genome editing can more accurately and rapidly improve biological varieties and breeds (genetically modify organisms) than existing technologies. Furthermore, PPR protein-based RNA manipulation and genome editing do not utilize foreign genes for trait transformation like gene recombination; instead, they manipulate the RNA or genome inherent in plants and animals. From this perspective, it can be said to be similar to previous breeding methods such as mutant screening and backcrossing. Therefore, it can provide a more effective and rapid response to global food and environmental issues.
[0394] (3) Chemistry
[0395] • In the production of useful substances using microorganisms, cultured cells, plants, or animals (such as insects), protein expression levels are regulated through manipulation of DNA and RNA. This can increase the productivity of useful substances. Examples of useful substances include not only protein-based substances such as antibodies, vaccines, and enzymes, but also lower molecular weight compounds such as pharmaceutical intermediates, fragrances, and pigments.
[0396] • Improve biofuel production efficiency by altering the metabolic pathways of algae and microorganisms.
[0397] [Example]
[0398] [Example 1: Intracellular analysis of fluorescent protein fusion with PPR protein]
[0399] (Sequence Design)
[0400] The target sequence is CAGCAGCAGCAGCAGCAG (SEQ ID NO: 1), formed by repeating the CAG sequence six times. The PPR motif determines the recognized base by the amino acid sequence at positions 1, 4, and ii. In the PPR motif recognizing cytosine, valine is placed at position 1, asparagine at position 4, and serine at position ii; in the PPR motif recognizing adenine, valine is placed at position 1, threonine at position 4, and asparagine at position ii; in the PPR motif recognizing guanine, valine is placed at position 1, threonine at position 4, and aspartic acid at position ii. Additionally, in the PPR motif recognizing uracil, valine is placed at position 1, asparagine at position 4, and aspartic acid at position ii.
[0401] Furthermore, regarding the first motif for recognizing cytosine ( Figure 1 The amino acids at positions 6 and 9 of the mutated motif (A) were selected as representative combinations, specifically leucine and glycine (C_6L9G, PPRcag_1, SEQ ID NO: 2 above); as variants, leucine and glutamic acid (C_6L9E, PPRcag_2, SEQ ID NO: 3), asparagine and glutamine (C_6N9Q, PPRcag_3, SEQ ID NO: 4), asparagine and glutamic acid (C6N9E, PPRcag_4, SEQ ID NO: 5), asparagine and lysine (C_6N9K, PPRcag_5, SEQ ID NO: 6), and aspartic acid and glycine (C_6D9G, PPRcag_6, SEQ ID NO: 7) were selected. Figure 1 B). Using these PPR motif sequences, PPR genes (SEQ ID NOs: 11-16) were created by arranging them in a manner that binds to the CAGCAGCAGCAGCAGCAG sequence (SEQ ID NO: 1 above). It should be noted that, in order to efficiently and accurately link the 18 DNA molecules encoding each PPR motif, in the PPR motifs targeting cytosine, adenine, and guanine bases, the amino acids other than positions 1, 4, 6, 9, and ii were amino acids that appeared at high frequencies as described above (SEQ ID NOs: 8-9, see Patent Document 1 above).
[0402] (Plasmid preparation)
[0403] Plasmids containing the PPR gene were constructed using the Golden Gate method. More specifically, 10 intermediate vectors (Dest-a, b, c, d, e, f, g, h, i, j) were designed in a sequential, seamless ligation manner. Twenty motifs, including one motif and two motifs (PPR motifs corresponding to A, C, G, and U, and two PPR motifs recognizing base combinations of AA, AC, AG, AU, CA, CC, CG, CU, GA, GC, GG, GU, UA, UC, UG, and UU), were inserted into each of the 10 vectors, thereby producing 200 elements.
[0404] Dest-a was prepared by using gene synthesis technology to create gaagacataaactccgtggtcacATACagagaccaaggtctcaGTGGtcacatacatgtcttc (SEQ ID NO: 43) and cloning it into pUC57-kan;
[0405] Dest-b was prepared by using gene synthesis technology to create gaagacatATACagagaccaaggtctcaGTGGtgacataatgtcttc (SEQ ID NO: 44) and cloning it into pUC57-kan;
[0406] Dest-c was prepared by using gene synthesis technology to create gaagacatcATACagagaccaaggtctcaGTGGttacatatgtcttc (SEQ ID NO: 45) and cloning it into pUC57-kan;
[0407] Dest-d was prepared by using gene synthesis technology to create gaagacatacATACagagaccaaggtctcaGTGGttacaatgtcttc (SEQ ID NO: 46) and cloning it into pUC57-kan;
[0408] Dest-e was prepared by using gene synthesis technology to create gaagacattacATACagagaccaaggtctcaGTGGtgacatgtcttc (SEQ ID NO: 47) and cloning it into pUC57-kan;
[0409] Dest-f was prepared by using gene synthesis technology to create gaagacattgacATACagagaccaaggtctcaGTGGttaatgtcttc (SEQ ID NO: 48) and cloning it into pUC57-kan;
[0410] Dest-g was prepared by using gene synthesis technology to create gaagacatgttacATACagagaccaaggtctcaGTGGtcatgtcttc (SEQ ID NO: 49) and cloning it into pUC57-kan;
[0411] Dest-h was prepared by using gene synthesis technology to create gaagacatggtcacATACagagaccaaggtctcaGTGGtatgtcttc (SEQ ID NO: 50) and cloning it into pUC57-kan;
[0412] Dest-i was prepared by using gene synthesis technology to create gaagacattggttacATACagagaccaaggtctcaGTGGatgtcttc (SEQ ID NO: 51) and cloning it into pUC57-kan;
[0413] Dest-j was prepared by using gene synthesis technology to create gaagacatgtggtgacATACagagaccaaggtctcaGTGGtcttc (SEQ ID NO: 52) and clone it into pUC57-kan.
[0414] Dest-a to Dest-j were selected based on the target base sequence and cloned into a vector via a Golden Gate reaction. The vector used here was designed with the amino acid sequence MGNSV (SEQ ID NO: 53) appended to the N-terminus of the 18 linked PPR sequences and the amino acid sequence ELTYNTLISGLGKAGRARDPPV (SEQ ID NO: 54) appended to the C-terminus. The correct gene size was confirmed to be cloned, and the sequence of the cloned gene was confirmed by sequencing.
[0415] (Detection of expression within cells)
[0416] The expression plasmid pcDNA3.1 in animal cultured cells contains a CMV promoter and an SV40 polyA signal sequence, between which a gene to be expressed can be inserted. To detect PPR protein expression in cells, it was decided to express a PPR protein fused with a fluorescent protein, and analyze intracellular aggregation and nuclear transport based on its fluorescence image. A protein gene fused sequentially from the N-terminus to EGFP, the nuclear transport signal sequence, the PPR protein, and the FLAG epitope tag was inserted into pcDNA3.1 (SEQ ID NOs: 17-22). Additionally, a protein gene fused sequentially from the N-terminus to mClover3, the PPR protein, the nuclear transport signal sequence, and the FLAG epitope tag was also inserted into pcDNA3.1 (SEQ ID NOs: 23-28). As a control, a plasmid without PPR was also prepared (SEQ ID NOs: 35-36).
[0417] HEK293T cells were used at a rate of 1×10 6 Cells were seeded in 10cm culture dishes containing 9mL DMEM and 1mL FBS. After culturing at 37℃ and 5% CO2 for 2 days, the cells were harvested. The harvested cells were then seeded at a rate of 4 × 10⁶ cells per well. 4 Cells were seeded in 96-well plates coated with PLL and cultured at 37°C and 5% CO2 for 1 day. 200 ng plasmid DNA, 0.6 μL Fugene (registered trademark)-HD (Promega, E2311), and 200 μL Opti-MEM were mixed and added to the wells. The cells were then cultured at 37°C and 5% CO2 for 1 day. After culture, the culture medium was removed, and the cells were washed once with 50 μL PBS. Then, 1 μL Hoechst (1 mg / mL, Tongren Chemical, 346-07951) and 50 μL PBS were added, and the cells were incubated at 37°C and 5% CO2 for 10 minutes. Afterward, the cells were washed once with 50 μL PBS. Following washing, 50 μL PBS was added, and GFP and Hoechst fluorescence images of each well were acquired using a DMi8 (Leica) fluorescence microscope.
[0418] The results are shown in Figure 2The results of confirming the intracellular expression of PPRs fused with EGFP and nuclear transport signaling sequences showed that PPRcag_1(6L9G) and PPRcag_2(6L9E) were not localized in the nucleus but were significantly aggregated around the nucleus. On the other hand, PPRcag_3(6N9Q), PPRcag_4(6N9E), PPRcag_5(6N9K), and PPRcag_6(6D9G), while exhibiting low aggregation, were not localized in the nucleus. With the fusion of mClover3, PPRcag_1(6L9G) and PPRcag_2(6L9E) were localized in the nucleus but aggregated therein. PPRcag_3(6N9Q), PPRcag_4(6N9E), PPRcag_5(6N9K), and PPRcag_6(6D9G) were localized in the nucleus, and no aggregation was observed. Therefore, it can be seen that in order to improve cohesion, the 6N9E, 6N9Q, 6N9K, and 6D9G variants are preferred. In addition, in order to effectively locate in the nucleus, mClover3 is better than EGFP.
[0419] [Example 2: RNA binding analysis of CAG to PPR protein]
[0420] To confirm the binding of PPRcag_1, PPRcag_2, PPRcag_3, PPRcag_4, PPRcag_5, and PPRcag_6 to the target RNA, recombinant proteins were prepared and binding experiments were conducted.
[0421] Protein genes with luciferase fused to the N-terminus and a 6×histidine tag sequence fused to the C-terminus of each PPR protein were designed and cloned into E. coli expression plasmids (SEQ ID NOs: 29-34). Additionally, as a control, an Nluc-Hisx6 protein gene (SEQ ID NO: 37) without PPR proteins was also created.
[0422] The completed plasmid was transformed into *E. coli* Rosetta(DE3) strain. The *E. coli* strain was cultured in 2 mL of LB medium containing 100 μg / mL ampicillin at 37°C for 12 hours. The OD... 600When the PPR value reached 0.5, the culture medium was transferred to a 15°C incubator and incubated for 30 minutes. Then, 100 μL (final concentration 0.1 mM) of IPTG was added, and the cells were incubated at 15°C for 16 hours. The *E. coli* clumps were recovered by centrifugation at 5,000 × g, 4°C, for 10 minutes. 1.5 mL of lysis buffer (20 mM Tris-HCl pH 8.0, 150 mM NaCl, 0.5% NP-40, 1 mM MgCl2, 2 mg / mL lysozyme, 1 mM PMSF, 2 μL DNase) was added, and the cells were frozen at -80°C for 20 minutes. Cell lysis was then performed by infiltration at 25°C for 30 minutes. Finally, the cells were centrifuged at 3700 rpm, 4°C, for 15 minutes, and the supernatant containing soluble PPR protein (*E. coli* lysis buffer) was recovered.
[0423] The binding assay of PPR protein to RNA was performed using the method of binding PPR protein to biotinylated RNA on streptavidin plates.
[0424] An RNA probe, Grainer, was synthesized by biotinyl-modifying the 5' end of a 30-base RNA containing the target CAGx6 sequence and non-target CGGx6, CUGx6, CCGx6, and D1b (UGGUGUAUCUUGUCUUUA) sequences (SEQ ID NO: 42, positions 8-25, respectively). 2.5 pmol of the biotinylated RNA probe was added to a streptavidin-coated plate (Cat No. 15502, Thermo Fisher) and reacted at room temperature for 30 min. The plate was then washed with probe washing buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT, 0.1% BSA). For background determination, wells containing lysis buffer but without biotinylated RNA (-probe) were also prepared. Then, blocking buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT, 1% BSA) was added, and the plate surface was blocked at room temperature for 30 minutes. 100 μL of a solution containing 1.5 × 10⁻⁶ ppm was then added. 8The luciferase fusion protein with PPR protein, measured at LU / μL luminescence, was used as an E. coli lysate for binding reaction at room temperature for 30 minutes. The mixture was then washed five times with 200 μL of washing buffer (20 mM Tris-HCl (pH 7.6), 150 mM NaCl, 5 mM MgCl2, 0.5% NP-40, 1 mM DTT). 40 μL of luciferase substrate (Promega, E151A), diluted 2500-fold with washing buffer, was added to the wells. After reacting for 5 minutes, the luminescence intensity was measured using a microplate reader (PerkinElmer, Cat No. 5103-35).
[0425] The results are shown in Figure 3 It can be seen that all PPRs specifically bind to CAGx6 as the target. Regarding the binding affinity to the target sequence, compared with PPRcag_1, PPRcag_2 is at the same level, PPRcag_3 is about 80%, PPRcag_4 is about 60%, PPRcag_5 is about 120%, and PPRcag_6 is about 130%. These results indicate that, except for PPRcag_4, almost no changes in binding performance due to variation were observed.
[0426] [Example 3: Regulation of PPR protein aggregation]
[0427] PPR proteins using the V2 motif (base sequence SEQ ID NO: 61, amino acid sequence SEQ ID NO: 62) and the v3.2 motif (base sequence SEQ ID NO: 63, amino acid sequence SEQ ID NO: 64) were prepared and purified using an *E. coli* expression system, and separated by gel filtration chromatography. It should be noted that the v2 motif refers to a PPR motif having the sequence SEQ ID NO: 2 and SEQ ID NOs: 8–10. Regarding the v3.2 motif, in the case of the first motif from the N-terminus, it refers to a PPR motif having any of the sequences SEQ ID NO: 4 and SEQ ID NOs: 58–60. Otherwise, for adenine, it refers to a PPR motif having a sequence in which the aspartic acid at position 15 of SEQ ID NO: 8 is replaced with a lysine; for bases other than adenine, it refers to a PPR motif having a sequence selected from SEQ ID NOs: 2, 9, and 10.
[0428] (Protein expression / purification)
[0429] The pE-SUMOpro Kan plasmid, containing the DNA sequence encoding the target PPR, was used to transform *E. coli* Rosetta strain. The cells were cultured at 37°C, and then the temperature was lowered to 20°C when the OD600 reached 0.6. IPTG was added to achieve a final concentration of 0.5 mM, allowing the target PPR protein to be expressed in *E. coli* as a SUMO fusion protein. After one night of culture, the bacterial cells were collected by centrifugation and resuspended in lysis buffer (50 mM Tris-HCl pH 8.0, 500 mM NaCl). *E. coli* was lysed by sonication, and after centrifugation at 17000 g for 30 minutes, the supernatant fraction was fed onto a Ni-agarose column. The column was washed with lysis buffer containing 20 mM imidazole, followed by elution with lysis buffer containing 400 mM imidazole to remove the SUMO fusion target PPR protein. After elution, SUMO protein was cleaved from the target PPR protein using Ulp1, while the protein solution was replaced with ion exchange buffer (50 mM Tris-HCl pH 8.0, 200 mM NaCl) by dialysis. Cation exchange chromatography using an SP column was then performed. After feeding the protein onto the column, the NaCl concentration was slowly increased from 200 mM to 1 M for elution. The fraction containing the target PPR protein was then finally purified using gel filtration chromatography on a Superdex 200 column. The ion-exchange eluted target PPR protein was fed onto a gel filtration column equilibrated with gel filtration buffer (25 mM HEPES pH 7.5, 200 mM NaCl, 0.5 mM tris(2-carboxyethyl)phosphine (TCEP)). Finally, the fraction containing the target PPR protein was concentrated, frozen in liquid nitrogen, and stored at -80°C until the next analysis.
[0430] (Gel filtration chromatography)
[0431] The purified recombinant PPR protein was adjusted to a concentration of 1 mg / mL. Gel filtration chromatography was performed using a Superdex 200 increase 10 / 300GL (GE Helthcare). The adjusted protein was fed onto a gel filtration column equilibrated with 25 mM HEPES pH 7.5, 200 mM NaCl, and 0.5 mM tris(2-carboxyethyl)phosphine (TCEP). The absorbance of the eluted solution at 280 nm was measured to analyze the properties of the protein.
[0432] (result)
[0433] The results are shown in Figure 4The smaller the elution volume, the larger the molecular size. v2 eluted in 8 to 10 mL elution volumes, while peaks were observed in 12 to 14 mL elution volumes for v3.2. This suggests that v2, due to the increased protein size, indicated a higher likelihood of aggregation, which was improved in v3.2. sequence list <110> EditForce Inc. (a Japanese company specializing in gene editing) National University Corporation Kyushu University <120> Low-cohesive PPR Proteins and Their Use <130> F20689K <150> JP2019-100553 <151> 2019-05-29 <160> 64 <170> PatentIn version 3.5 <210> 1 <211> 18 <212> DNA <213> Artificial sequence <220> <223> Target sequence <400> 1 cagcagcagc agcagcag 18 <210> 2 <211> 35 <212> PRT <213> Artificial sequence <220> <223> PPR motif <400> 2 Val Thr Tyr Asn Thr Leu Ile Asp Gly Leu Cys Lys Ser Gly Lys Ile 1 5 10 15 Glu Glu Ala Leu Lys Leu Phe Lys Glu Met Glu Glu Lys Gly Ile Thr 20 25 30 Pro Ser Val 35 <210> 3 <211> 35 <212> PRT <213> Artificial sequence <220> <223> PPR motif <400> 3 Val Thr Tyr Asn Thr Leu Ile Asp Glu Leu Cys Lys Ser Gly Lys Ile 1 5 10 15 Glu Glu Ala Leu Lys Leu Phe Lys Glu Met Glu Glu Lys Gly Ile Thr 20 25 30 Pro Ser Val 35 <210> 4 <211> 35 <212> PRT <213> Artificial sequence <220> <223> PPR motif, 1st_U (v3.2_U) <400> 4 Val Thr Tyr Asn Thr Asn Ile Asp Gln Leu Cys Lys Ser Gly Lys Ile 1 5 10 15 Glu Glu Ala Leu Lys Leu Phe Lys Glu Met Glu Glu Lys Gly Ile Thr 20 25 30 Pro Ser Val 35 <210> 5 <211> 35 <212> PRT <213> Artificial sequence <220> <223> PPR motif <400> 5 Val Thr Tyr Asn Thr Asn Ile Asp Glu Leu Cys Lys Ser Gly Lys Ile 1 5 10 15 Glu Glu Ala Leu Lys Leu Phe Lys Glu Met Glu Glu Lys Gly Ile Thr 20 25 30 Pro Ser Val 35 <210> 6 <211> 35 <212> PRT <213> Artificial sequence <220> <223> PPR motif <400> 6 Val Thr Tyr Asn Thr Asn Ile Asp Lys Leu Cys Lys Ser Gly Lys Ile 1 5 10 15 Glu Glu Ala Leu Lys Leu Phe Lys Glu Met Glu Glu Lys Gly Ile Thr 20 25 30 Pro Ser Val 35 <210> 7 <211> 35 <212> PRT <213> Artificial sequence <220> <223> PPR motif <400> 7 Val Thr Tyr Asn Thr Asp Ile Asp Gly Leu Cys Lys Ser Gly Lys Ile 1 5 10 15 Glu Glu Ala Leu Lys Leu Phe Lys Glu Met Glu Glu Lys Gly Ile Thr 20 25 30 Pro Ser Val 35 <210> 8 <211> 35 <212> PRT <213> Artificial sequence <220> <223> PPR motif <400> 8 Val Thr Tyr Thr Thr Leu Ile Asp Gly Leu Cys Lys Ala Gly Asp Val 1 5 10 15 Asp Glu Ala Leu Glu Leu Phe Lys Glu Met Arg Ser Lys Gly Val Lys 20 25 30 Pro Asn Val 35 <210> 9 <211> 35 <212> PRT <213> Artificial sequence <220> <223> PPR motif <400> 9 Val Thr Tyr Thr Thr Leu Ile Asp Gly Leu Cys Lys Ala Gly Lys Val 1 5 10 15 Asp Glu Ala Leu Glu Leu Phe Asp Glu Met Lys Glu Arg Gly Ile Lys 20 25 30 Pro Asp Val 35 <210> 10 <211> 35 <212> PRT <213> Artificial sequence <220> <223> PPR motif <400> 10 Val Thr Tyr Asn Thr Leu Ile Asp Gly Leu Cys Lys Ala Gly Arg Leu 1 5 10 15 Asp Glu Ala Glu Glu Leu Leu Glu Glu Met Glu Glu Lys Gly Ile Lys 20 25 30 Pro Asp Val 35 <210> 11 <211> 2303 <212> DNA <213> Artificial sequence <220> <223> PPR gene <220> <221> misc_feature <222> (1894) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1935) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1963) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2004) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2032) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2073) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2101)..(2101) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2142)..(2142) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2170)..(2170) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2211)..(2211) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2239)..(2239) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2280)..(2280) <223> n is a, c, g, t, or u <400> 11 gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggtcac acaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtgacataca ccaccctgat cgacggcctg 240 tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg 420 gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 480 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc 540 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 600 atgaaggaga ggggcatcaa gcccgacgtg gttacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaaggggc 720. atcaccccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 780 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 840 gtgacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg 900 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1140 gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg 1320 aagctgttca aggagatgga ggagaagggc atcacccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gttacataca ccaccctgat cgacggcctg 1500 tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc 1560 atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag 1620 atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg 1680 gtgacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 1740 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac ataccacc 1800 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 1860 atgaaggaga ggggcatcaa gcccgacgag vtyntdgcks gkakkmkgts vvtyttdgck 1920 agdvdakmrs kgvknvvtyt tdgckagkvd admkrgkdvv tyntdgcksg kakkmkgtsv 1980 vtyttdgcka gdvdakmrsk gvknvvtytt dgckagkvda dmkrgkdvvt yntdgcksgk 2040 akkmkgtsvv tyttdgckag dvdakmrskg vknvvtyttd gckagkvdad mkrgkdvvty 2100 ntdgcksgka kkmkgtsvvt yttdgckagd vdakmrskgv knvvtyttdg ckagkvdadm 2160 krgkdvvtyn tdgcksgkak kmkgtsvvty ttdgckagdv dakmrskgvk nvvtyttdgc 2220 kagkvdadmk rgkdvvtynt dgcksgkakk mkgtsvvtyt tdgckagdvd akmrskgvkn 2280 vvtyttdgck agkvdadmkr gkd 2303 <210> 12 <211> 2302 <212> DNA <213> Artificial sequence <220> <223> PPR gene <220> <221> misc_feature <222> (1894) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1934) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1962) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2003) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2031) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2072)..(2072) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2100)..(2100) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2141)..(2141) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2169)..(2169) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2210)..(2210) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2238)..(2238) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2279)..(2279) <223> n is a, c, g, t, or u <400> 12 gtcacataca acaccctgat cgacgaactg tgcaagagcg gcaagatcga ggaggccctg 60 aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggtcac atacaccaca 120 ctgatcgacg gactgtgtaa agccggcgac gtggacgaag ccctcgagct gttcaaagag 180 atgcggagca agggcgtgaa gcccaacgtg gtgacataca ccaccctgat cgacggcctg 240 tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg 420 gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 480 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc 540 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 600 atgaaggaga ggggcatcaa gcccgacgtg gttacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaaggggc 720. atcaccccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 780 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 840 gtgacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg 900 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1140 gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg 1320 aagctgttca aggagatgga ggagaagggc atcacccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gttacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag 1620. atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg 1680 gtgacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 1740 gagctgttca aagagatcg gagcaagggc gtgaagccca acgtggtcac atacaccacc 1800 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 1860 atgaaggaga ggggcatcaa gcccgacgag vtyntdcksg kakkmkgtsv vtyttdgcka 1920 gdvdakmrsk gvknvvtytt dgckagkvda dmkrgkdvvt yntdgcksgk akkmkgtsvv 1980 tyttdgckag dvdakmrskg vknvvtyttd gckagkvdad mkrgkdvvty ntdgcksgka 2040 kkmkgtsvvt yttdgckagd vdakmrskgv knvvtyttdg ckagkvdadm krgkdvvtyn 2100 tdgcksgkak kmkgtsvvty ttdgckagdv dakmrskgvk nvvtyttdgc kagkvdadmk 2160 rgkdvvtynt dgcksgkakk mkgtsvvtyt tdgckagdvd acmrskgvkn vvtyttdgck 2220 agkvdadmkr gkdvvtyntd gcksgkakkm kgtsvvtytt dgckagdvda kmrskgvknv 2280 vtyttdgcka gkvdadmkrg kd 2302 <210> 13 <211> 2303 <212> DNA <213> Artificial sequence <220> <223> PPR gene <220> <221> misc_feature <222> (1894) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1896) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1935) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1963) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2004) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2032) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2073) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2101)..(2101) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2142)..(2142) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2170)..(2170) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2211)..(2211) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2239)..(2239) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2280)..(2280) <223> n is a, c, g, t or u <400> 13 gtcacataca acaccaacat cgaccagctg tgcaagagcg gcaagatcga ggaggccctg 60 aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggtcac atacaccaca 120 ctgatcgacg gactgtgtaa agccggcgac gtggacgaag ccctcgagct gttcaaagag 180 atgcggagca agggcgtgaa gcccaacgtg gtgacataca ccaccctgat cgacggcctg 240 tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc 300 atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg 420 gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 480 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc 540 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 600 atgaaggaga ggggcatcaa gcccgacgtg gttacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaaggggc 720. atcaccccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 780 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 840 gtgacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg 900 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1140 gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg 1320 aagctgttca aggagatgga ggagaagggc atcacccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gttacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag 1620. atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtgacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 1740 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc 1800 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 1860 atgaaggaga ggggcatcaa gcccgacgag vtyntndcks gkakkmkgts vvtyttdgck 1920 agdvdakmrs kgvknvvtyt tdgckagkvd admkrgkdvv tyntdgcksg kakkmkgtsv 1980 vtyttdgcka gdvdakmrsk gvknvvtytt dgckagkvda dmkrgkdvvt yntdgcksgk 2040 akkmkgtsvv tyttdgckag dvdakmrskg vknvvtyttd gckagkvdad mkrgkdvvty 2100 ntdgcksgka kkmkgtsvvt yttdgckagd vdakmrskgv knvvtyttdg ckagkvdadm 2160 krgkdvvtyn tdgcksgkak kmkgtsvvty ttdgckagdv dakmrskgvk nvvtyttdgc 2220 kagkvdadmk rgkdvvtynt dgcksgkakk mkgtsvvtyt tdgckagdvd akmrskgvkn 2280 vvtyttdgck agkvdadmkr gkd 2303 <210> 14 <211> 2303 <212> DNA <213> Artificial sequence <220> <223> PPR gene <220> <221> misc_feature <222> (1894) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1896) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1935) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1963) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2004) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2032) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2073) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2101)..(2101) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2142)..(2142) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2170)..(2170) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2211)..(2211) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2239)..(2239) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2280)..(2280) <223> n is a, c, g, t or u <400> 14 gtcacataca acaccaacat cgacgaactg tgcaagagcg gcaagatcga ggaggccctg 60 aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggtcac atacaccaca 120 ctgatcgacg gactgtgtaa agccggcgac gtggacgaag ccctcgagct gttcaaagag 180 atgcggagca agggcgtgaa gcccaacgtg gtgacataca ccaccctgat cgacggcctg 240 tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc 300 atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag 360 atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg 420 gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 480 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc 540 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 600 atgaaggaga ggggcatcaa gcccgacgtg gttacataca acaccctgat cgacggcctg 660 tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc 720 atcaccccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 780 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 840 gtgacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg 900 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc 960 ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag 1020 atggaggaga agggcatcac ccccagcgtg gttacataca ccacactgat cgacggactg 1080 tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1140 gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg 1320 aagctgttca aggagatgga ggagaagggc atcacccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gttacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag 1620. atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtgacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 1740 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 1860. atgaaggaga ggggcatcaa gcccgacgag vtyntndcks gkakkmkgts vvtyttdgck 1920 agdvdakmrs kgvknvvtyt tdgckagkvd admkrgkdvv tyntdgcksg kakkmkgtsv 1980 vtyttdgcka gdvdakmrsk gvknvvtytt dgckagkvda dmkrgkdvvt yntdgcksgk 2040 akkmkgtsvv tyttdgckag dvdakmrskg vknvvtyttd gckagkvdad mkrgkdvvty 2100 ntdgcksgka kkmkgtsvvt yttdgckagd vdakmrskgv knvvtyttdg ckagkvdadm 2160 krgkdvvtyn tdgcksgkak kmkgtsvvty ttdgckagdv dakmrskgvk nvvtyttdgc 2220 kagkvdadmk rgkdvvtynt dgcksgkakk mkgtsvvtyt tdgckagdvd akmrskgvkn 2280 vvtyttdgck agkvdadmkr gkd 2303 <210> 15 <211> 2304 <212> DNA <213> Artificial sequence <220> <223> PPR gene <220> <221> misc_feature <222> (1894)..(1894) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (1896)..(1896) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (1936) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (1964) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2005) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2033) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2074) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2102)..(2102) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2143)..(2143) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2171)..(2171) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2212)..(2212) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2240)..(2240) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2281)..(2281) <223> n is a, c, g, t or u <400> 15 gtcacataca acaccaacat cgacaaactg tgcaagagcg gcaagatcga ggaggccctg 60 aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggtcac atacaccaca 120 ctgatcgacg gactgtgtaa agccggcgac gtggacgaag ccctcgagct gttcaaagag 180 atgcggagca agggcgtgaa gcccaacgtg gtgacataca ccaccctgat cgacggcctg 240 tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc 300 atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag 360 atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg 420 gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 480 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc 540 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 600 atgaaggaga ggggcatcaa gcccgacgtg gttacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaaggggc 720. atcaccccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 780 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 840 gtgacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg 900 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1140 gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg 1320 aagctgttca aggagatgga ggagaagggc atcacccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gttacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag 1620. atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtgacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 1740 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 1860. atgaaggaga ggggcatcaa gcccgacgag vtyntndkck sgkakkmkgt svvtyttdgc 1920. kagdvdakmr skgvknvvty ttdgckagkv dadmkrgkdv vtyntdgcks gkakkmkgts vvtyttdgck agdvdakmrs kgvknvvtyt tdgckagkvd admkrgkdvv tyntdgcksg 2040 kakkmkgtsv vtyttdgcka gdvdakmrsk gvknvvtytt dgckagkvda dmkrgkdvvt 2100 yntdgcksgk akkmkgtsvv tyttdgckag dvdakmrskg vknvvtyttd gckagkvdad 2160 mkrgkdvvty ntdgcksgka kkmkgtsvvt yttdgckagd vdakmrskgv knvvtyttdg 2220 ckagkvdadm krgkdvvtyn tdgcksgkak kmkgtsvvty ttdgckagdv dakmrskgvk 2280 nvvtyttdgc kagkvdadmk rgkd 2304 <210> 16 <211> 2304 <212> DNA <213> Artificial sequence <220> <223> PPR gene <220> <221> misc_feature <222> (1894)..(1894) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (1936)..(1936) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (1964)..(1964) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2005) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2033) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2074) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2102)..(2102) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2143)..(2143) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2171)..(2171) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2212)..(2212) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2240)..(2240) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2281)..(2281) <223> n is a, c, g, t, or u <400> 16 gtcacataca acaccgatat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggtcac acaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtgacataca ccaccctgat cgacggcctg 240 tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg 420 gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 480 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc 540 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 600 atgaaggaga ggggcatcaa gcccgacgtg gttacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaaggggc 720. atcaccccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 780 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 840 gtgacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg 900 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1140 gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg 1320 aagctgttca aggagatgga ggagaagggc atcacccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gttacataca ccaccctgat cgacggcctg 1500 tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc 1560 atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag 1620 atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg 1680 gtgacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 1740 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac ataccacc 1800 ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag 1860 atgaaggaga ggggcatcaa gcccgacgag vtyntddgck sgkakkkmkgt svvtyttdgc 1920 kagdvdakmr skgvknvvty ttdgckagkv dadmkrgkdv vtyntdgcks gkakkkmkgts 1980 vvtyttdgck agdvdakmrs kgvknvvtyt tdgckagkvd admkrgkdvv tyntdgcksg 2040 kakkmkgtsv vtyttdgcka gdvdakmrsk gvknvvtytt dgckagkvda dmkrgkdvvt 2100 yntdgcksgk akkmkgtsvv tyttdgckag dvdakmrskg vknvvtyttd gckagkvdad 2160 mkrgkdvvty ntdgcksgka kkmkgtsvvt yttdgckagd vdakmrskgv knvvtyttdg 2220 ckagkvdadm krgkdvvtyn tdgcksgkak kmkgtsvvty ttdgckagdv dakmrskgvk 2280 nvvtyttdgc kagkvdadmk rgkd 2304 <210> 17 <211> 3481 <212> DNA <213> Artificial sequence <220> <223> PPR vector <220> <221> misc_feature <222> (2858)..(2858) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2914)..(2914) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2926)..(2926) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2934)..(2934) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2939)..(2939) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2941)..(2941) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2944)..(2944) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2952)..(2952) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2956)..(2956) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2960)..(2960) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2969)..(2969) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2976)..(2976) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2986)..(2986) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3023)..(3023) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3029)..(3029) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3070)..(3070) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3098)..(3098) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3139)..(3139) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3167)..(3167) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3208)..(3208) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3236)..(3236) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3277)..(3277) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3305)..(3305) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3346)..(3346) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (3374)..(3374) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (3415)..(3415) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (3441)..(3441) <223> n is a, c, g, t or u <400> 17 atggccggag tgagcaaggg cgaggagctg ttcaccgggg tggtgcccat cctggtcgag 60 ctggacggcg acgtaaacgg ccacaagttc agcgtgtccg gcgagggcga gggcgatgcc 120 acctacggca agctgaccct gaagttcatc tgcaccaccg gcaagctgcc cgtgccctgg 180 cccaccctcg tgaccaccct gacctacggc gtgcagtgct tcagccgcta ccccgaccac 240 atgaagcagc acgacttctt caagtccgcc atgcccgaag gctacgtcca ggagcgcacc 300 atcttcttca aggacgacgg caactacaag acccgcgccg aggtgaagtt cgagggcgac 360 accctggtga accgcatcga gctgaagggc atcgacttca aggaggacgg caacatcctg gggcacaagc tggagtacaa ctacaacagc cacaacgtct atatcatggc cgacaagcag aagaacggca tcaaggtgaa cttcaagatc cgccacaaca tcgaggacgg cagcgtgcag 540 ctcgccgacc actaccagca gaacaccccc atcggcgacg gccccgtgct gctgcccgac aaccactacc tgagcaccca gtccgccctg agcaaagacc ccaacgagaa gcgcgatcac atggtcctgc tggagttcgt gaccgccgcc gggatcactc tcggcatgga cgagctgtac 720 aagccaaaga aaaagagaa ggttagccat ggctccggcg gcagcgggggg agggctccat 780 atgggaaact ccgtggtcac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 960 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 1260 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 1380 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggttac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1620 gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac acaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggttac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag 2340 atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggtgac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 2520 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 2580 gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 2640 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgagctgac ctacaacacc 2700 ctgatcagcg gcctgggcaa ggccggcaga gccagagacc cccccgtgct cagtagcggg 2760 gactataagg accacgacgg agactacaag gatcatgata ttgattacaa agacgatgac 2820 gataagatgg ccggacgcta gmagvskgtg vvvdgdvngh ksvsgggdat ygktkcttgk 2880 vwtvtttygv csrydhmkhd ksamgyvrtk ddgnyktrav kgdtvnrkgd kdgnghkyny 2940 nshnvymadk kngkvnkrhn dgsvadhynt gdgvdnhyst saskdnkrdh mvvtaagtgm 3000 dykkkkrkvs hgsggsgggh mgnsvvtynt dgcksgkakk mkgtsvvtyt tdgckagdvd 3060 akmrskgvkn vvtyttdgck agkvdadmkr gkdvvtyntd gcksgkakkm kgtsvvtytt 3120 dgckagdvda kmrskgvknv vtyttdgcka gkvdadmkrg kdvvtyntdg cksgkakkmk 3180 gtsvvtyttd gckagdvdak mrskgvknvv tyttdgckag kvdadmkrgk dvvtyntdgc 3240 ksgkakkmkg tsvvtyttdg ckagdvdakm rskgvknvvt yttdgckagk vdadmkrgkd 3300 vvtyntdgck sgkakkmkgt svvtyttdgc kagdvdakmr skgvknvvty ttdgckagkv 3360 dadmkrgkdv vtyntdgcks gkakkmkgts vvtyttdgck agdvdakmrs kgvknvvtyt 3420 tdgckagkvd admkrgkdty ntsggkagra rdvssgdykd hdgdykdhdd ykddddkmag 3480 r 3481 <210> 18 <211> 3480 <212> DNA <213> Artificial sequence <220> <223> PPR vector <220> <221> misc_feature <222> (2858)..(2858) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2914)..(2914) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2926)..(2926) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2934)..(2934) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2939)..(2939) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2941)..(2941) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2944)..(2944) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2952)..(2952) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2956)..(2956) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2960)..(2960) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2969)..(2969) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2976)..(2976) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2986)..(2986) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3023)..(3023) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3029)..(3029) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3069)..(3069) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3097)..(3097) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3138)..(3138) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3166)..(3166) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3207)..(3207) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3235)..(3235) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3276)..(3276) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3304)..(3304) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (3345)..(3345) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (3373)..(3373) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (3414)..(3414) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (3440)..(3440) <223> n is a, c, g, t or u <400> 18 atggccggag tgagcaaggg cgaggagctg ttcaccgggg tggtgcccat cctggtcgag 60 ctggacggcg acgtaaacgg ccacaagttc agcgtgtccg gcgagggcga gggcgatgcc 120 acctacggca agctgaccct gaagttcatc tgcaccaccg gcaagctgcc cgtgccctgg 180 cccaccctcg tgaccaccct gacctacggc gtgcagtgct tcagccgcta ccccgaccac 240 atgaagcagc acgacttctt caagtccgcc atgcccgaag gctacgtcca ggagcgcacc 300 atcttcttca aggacgacgg caactacaag acccgcgccg aggtgaagtt cgagggcgac 360 accctggtga accgcatcga gctgaagggc atcgacttca aggaggacgg caacatcctg gggcacaagc tggagtacaa ctacaacagc cacaacgtct atatcatggc cgacaagcag aagaacggca tcaaggtgaa cttcaagatc cgccacaaca tcgaggacgg cagcgtgcag 540 ctcgccgacc actaccagca gaacaccccc atcggcgacg gccccgtgct gctgcccgac aaccactacc tgagcaccca gtccgccctg agcaaagacc ccaacgagaa gcgcgatcac atggtcctgc tggagttcgt gaccgccgcc gggatcactc tcggcatgga cgagctgtac 720 aagccaaaga aaaagagaa ggttagccat ggctccggcg gcagcgggggg agggctccat 780 atgggaaact ccgtggtcac attackacacc ctgatcgacg aactgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 960 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 1260 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 1380 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggttac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1620 gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac acaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggttac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag 2340 atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggtgac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 2520 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 2580 gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 2640 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgagctgac ctacaacacc 2700 ctgatcagcg gcctgggcaa ggccggcaga gccagagacc cccccgtgct cagtagcggg 2760 gactataagg accacgacgg agactacaag gatcatgata ttgattacaa agacgatgac 2820 gataagatgg ccggacgcta gmagvskgtg vvvdgdvngh ksvsgggdat ygktkcttgk 2880 vwtvtttygv csrydhmkhd ksamgyvrtk ddgnyktrav kgdtvnrkgd kdgnghkyny 2940 nshnvymadk kngkvnkrhn dgsvadhynt gdgvdnhyst saskdnkrdh mvvtaagtgm 3000 dykkkkrkvs hgsggsgggh mgnsvvtynt dcksgkakkm kgtsvvtytt dgckagdvda 3060 kmrskgvknv vtyttdgcka gkvdadmkrg kdvvtyntdg cksgkakkmk gtsvvtyttd 3120 gckagdvdak mrskgvknvv tyttdgckag kvdadmkrgk dvvtyntdgc ksgkakkmkg 3180 tsvvtyttdg ckagdvdakm rskgvknvvt yttdgckagk vdadmkrgkd vvtyntdgck 3240 sgkakkmkgt svvtyttdgc kagdvdakmr skgvknvvty ttdgckagkv dadmkrgkdv 3300 vtyntdgcks gkakkmkgts vvtyttdgck agdvdakmrs kgvknvvtyt tdgckagkvd 3360 admkrgkdvv tyntdgcksg kakkmkgtsv vtyttdgcka gdvdakmrsk gvknvvtytt 3420 dgckagkvda dmkrgkdtyn tsggkagrar dvssgdykdh dgdykdhddy kddddkmagr 3480 <210> 19 <211> 3481 <212> DNA <213> Artificial sequence <220> <223> PPR vector <220> <221> misc_feature <222> (2858)..(2858) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2914)..(2914) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2926)..(2926) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2934)..(2934) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2939)..(2939) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2941)..(2941) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2944)..(2944) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2952)..(2952) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2956)..(2956) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2960)..(2960) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2969)..(2969) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2976)..(2976) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2986)..(2986) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3023)..(3023) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3029)..(3029) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3031)..(3031) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3070)..(3070) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3098)..(3098) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3139)..(3139) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3167)..(3167) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3208)..(3208) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3236)..(3236) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3277)..(3277) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3305)..(3305) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3346)..(3346) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3374)..(3374) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3415)..(3415) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3441)..(3441) <223> n is a, c, g, t, or u <400> 19 atggccggag tgagcaaggg cgaggagctg ttcaccgggg tggtgcccat cctggtcgag 60 ctggacggcg acgtaaacgg ccacaagttc agcgtgtccg gcgagggcga gggcgatgcc 120 acctacggca agctgaccct gaagttcatc tgcaccaccg gcaagctgcc cgtgccctgg 180 cccaccctcg tgaccaccct gacctacggc gtgcagtgct tcagccgcta ccccgaccac 240 atgaagcagc acgacttctt caagtccgcc atgcccgag gctacgtcca ggagcgcacc 300 atcttcttca aggacgacgg caactacaag acccgcgccg aggtgaagtt cgagggcgac 360 accctggtga accgcatcga gctgaagggc atcgacttca aggaggacgg caacatcctg gggcacaagc tggagtacaa ctacaacagc cacaacgtct atatcatggc cgacaagcag aagaacggca tcaaggtgaa cttcaagatc cgccacaaca tcgaggacgg cagcgtgcag 540 ctcgccgacc actaccagca gaacaccccc atcggcgacg gccccgtgct gctgcccgac aaccactacc tgagcaccca gtccgccctg agcaaagacc ccaacgagaa gcgcgatcac atggtcctgc tggagttcgt gaccgccgcc gggatcactc tcggcatgga cgagctgtac 720 aagccaaaga aaaagagaa ggttagccat ggctccggcg gcagcgggggg agggctccat 780 atgggaaact ccgtggtcac attackacacc aacatcgacc agctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 960 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtgac atacaccacc 1020 ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag 1080 atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg 1140 tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc 1200 atcaccccca gcgtggttac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 1260 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 1320 gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 1380 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggttac atacaacacc 1440 ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag 1500 atggaggaga agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg 1560 tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1620 gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac acaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggttac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag 2340 atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg 2400 tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc 2460 atcaccccca gcgtggtgac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 2520 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 2580 gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 2640 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgagctgac ctacaacacc 2700 ctgatcagcg gcctgggcaa ggccggcaga gccagagacc cccccgtgct cagtagcggg 2760 gactataagg accacgacgg agactacaag gatcatgata ttgattacaa agacgatgac 2820 gataagatgg ccggacgcta gmagvskgtg vvvdgdvngh ksvsgggdat ygktkcttgk 2880 vwtvtttygv csrydhmkhd ksamgyvrtk ddgnyktrav kgdtvnrkgd kdgnghkyny 2940 nshnvymadk kngkvnkrhn dgsvadhynt gdgvdnhyst saskdnkrdh mvvtaagtgm 3000 dykkkkrkvs hgsggsgggh mgnsvvtynt ndcksgkakk mkgtsvvtyt tdgckagdvd 3060 akmrskgvkn vvtyttdgck agkvdadmkr gkdvvtyntd gcksgkakkm kgtsvvtytt 3120 dgckagdvda kmrskgvknv vtyttdgcka gkvdadmkrg kdvvtyntdg cksgkakkmk 3180 gtsvvtyttd gckagdvdak mrskgvknvv tyttdgckag kvdadmkrgk dvvtyntdgc 3240 ksgkakkmkg tsvvtyttdg ckagdvdakm rskgvknvvt yttdgckagk vdadmkrgkd 3300 vvtyntdgck sgkakkmkgt svvtyttdgc kagdvdakmr skgvknvvty ttdgckagkv 3360 dadmkrgkdv vtyntdgcks gkakkmkgts vvtyttdgck agdvdakmrs kgvknvvtyt 3420 tdgckagkvd admkrgkdty ntsggkagra rdvssgdykd hdgdykdhdd ykddddkmag 3480 r 3481 <210> 20 <211> 3481 <212> DNA <213> Artificial sequence <220> <223> PPR vector <220> <221> misc_feature <222> (2858)..(2858) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2914)..(2914) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (2926)..(2926) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2934)..(2934) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2939)..(2939) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2941)..(2941) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2944)..(2944) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2952)..(2952) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2956)..(2956) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2960)..(2960) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2969)..(2969) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2976)..(2976) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2986)..(2986) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3023)..(3023) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3029)..(3029) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3031)..(3031) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3070)..(3070) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3098)..(3098) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3139)..(3139) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3167)..(3167) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3208)..(3208) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3236)..(3236) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3277)..(3277) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3305)..(3305) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3346)..(3346) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3374)..(3374) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3415)..(3415) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3441)..(3441) <223> n is a, c, g, t, or u <400> 20 atggccggag tgagcaaggg cgaggagctg ttcaccgggg tggtgcccat cctggtcgag 60 ctggacggcg acgtaaacgg ccacaagttc agcgtgtccg gcgagggcga gggcgatgcc 120 180. acctacggca agctgaccct gaagttcatc tgcaccaccg gcaagctgcc cgtgccctgg cccaccctcg tgaccaccct gacctacggc gtgcagtgct tcagccgcta ccccgaccac 240 atgaagcagc acgacttctt caagtccgcc atgcccgag gctacgtcca ggagcgcacc 300 atcttcttca aggacgacgg caactacaag acccgcgccg aggtgaagtt cgagggcgac 360 accctggtga accgcatcga gctgaagggc atcgacttca aggaggacgg caacatcctg gggcacaagc tggagtacaa ctacaacagc cacaacgtct atatcatggc cgacaagcag aagaacggca tcaaggtgaa cttcaagatc cgccacaaca tcgaggacgg cagcgtgcag 540 ctcgccgacc actaccagca gaacaccccc atcggcgacg gccccgtgct gctgcccgac aaccactacc tgagcaccca gtccgccctg agcaaagacc ccaacgagaa gcgcgatcac atggtcctgc tggagttcgt gaccgccgcc gggatcactc tcggcatgga cgagctgtac 720 aagccaaaga aaaagagaa ggttagccat ggctccggcg gcagcgggggg agggctccat 780 atgggaaact ccgtggtcac attackacacc aacatcgacg aactgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 960 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 1260 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 1380 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggttac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1620 gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac acaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggttac atacaccacc 2280 ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag 2340 atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg 2400 tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc 2460 atcaccccca gcgtggtgac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 2520 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 2580 gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 2640 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgagctgac ctacaacacc 2700 ctgatcagcg gcctgggcaa ggccggcaga gccagagacc cccccgtgct cagtagcggg 2760 gactataagg accacgacgg agactacaag gatcatgata ttgattacaa agacgatgac 2820 gataagatgg ccggacgcta gmagvskgtg vvvdgdvngh ksvsgggdat ygktkcttgk 2880 vwtvtttygv csrydhmkhd ksamgyvrtk ddgnyktrav kgdtvnrkgd kdgnghkyny 2940 nshnvymadk kngkvnkrhn dgsvadhynt gdgvdnhyst saskdnkrdh mvvtaagtgm 3000 dykkkkrkvs hgsggsgggh mgnsvvtynt ndcksgkakk mkgtsvvtyt tdgckagdvd 3060 akmrskgvkn vvtyttdgck agkvdadmkr gkdvvtyntd gcksgkakkm kgtsvvtytt 3120 dgckagdvda kmrskgvknv vtyttdgcka gkvdadmkrg kdvvtyntdg cksgkakkmk 3180 gtsvvtyttd gckagdvdak mrskgvknvv tyttdgckag kvdadmkrgk dvvtyntdgc 3240 ksgkakkmkg tsvvtyttdg ckagdvdakm rskgvknvvt yttdgckagk vdadmkrgkd 3300 vvtyntdgck sgkakkmkgt svvtyttdgc kagdvdakmr skgvknvvty ttdgckagkv 3360 dadmkrgkdv vtyntdgcks gkakkmkgts vvtyttdgck agdvdakmrs kgvknvvtyt 3420 tdgckagkvd admkrgkdty ntsggkagra rdvssgdykd hdgdykdhdd ykddddkmag 3480 r 3481 <210> 21 <211> 3482 <212> DNA <213> Artificial sequence <220> <223> PPR vector <220> <221> misc_feature <222> (2858)..(2858) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2914)..(2914) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2926)..(2926) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2934)..(2934) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2939)..(2939) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2941)..(2941) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2944)..(2944) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2952)..(2952) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2956)..(2956) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2960)..(2960) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2969)..(2969) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2976)..(2976) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2986)..(2986) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3023)..(3023) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3029)..(3029) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3031)..(3031) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3071)..(3071) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3099)..(3099) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3140)..(3140) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3168)..(3168) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3209)..(3209) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3237)..(3237) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3278)..(3278) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3306)..(3306) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3347)..(3347) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3375)..(3375) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3416)..(3416) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3442)..(3442) <223> n is a, c, g, t or u <400> 21 atggccggag tgagcaaggg cgaggagctg ttcaccgggg tggtgcccat cctggtcgag 60 ctggacggcg acgtaaacgg ccacaagttc agcgtgtccg gcgagggcga gggcgatgcc 120 acctacggca agctgaccct gaagttcatc tgcaccaccg gcaagctgcc cgtgccctgg 180 cccaccctcg tgaccaccct gacctacggc gtgcagtgct tcagccgcta ccccgaccac 240 atgaagcagc acgacttctt caagtccgcc atgcccgaag gctacgtcca ggagcgcacc 300 atcttcttca aggacgacgg caactacaag acccgcgccg aggtgaagtt cgagggcgac 360 accctggtga accgcatcga gctgaagggc atcgacttca aggaggacgg caacatcctg 420 gggcacaagc tggagtacaa ctacaacagc cacaacgtct atatcatggc cgacaagcag 480 aagaacggca tcaaggtgaa cttcaagatc cgccacaaca tcgaggacgg cagcgtgcag 540 ctcgccgacc actaccagca gaacaccccc atcggcgacg gccccgtgct gctgcccgac 600 aaccactacc tgagcaccca gtccgccctg agcaaagacc ccaacgagaa gcgcgatcac atggtcctgc tggagttcgt gaccgccgcc gggatcactc tcggcatgga cgagctgtac 720 aagccaaaga aaaagagaa ggttagccat ggctccggcg gcagcgggggg agggctccat 780 atgggaaact ccgtggtcac attackacacc aacatcgaca aactgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 960 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 1260 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 1380 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggttac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1620 gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac acaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggttac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag 2340 atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcacccca gcgtggtgac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 2520 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 2640 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgagctgac ctacaacacc ctgatcagcg gcctgggcaa ggccggcaga gccagagacc cccccgtgct cagtagcggg 2760 gactataagg accacgacgg agactacaag gatcatgata ttgattacaa agacgatgac 2820 gataagatgg ccggacgcta gmagvskgtg vvvdgdvngh ksvsgggdat ygktkcttgk 2880 vwtvtttygv csrydhmkhd ksamgyvrtk ddgnyktrav kgdtvnrkgd kdgnghkyny 2940 nshnvymadk kngkvnkrhn dgsvadhynt gdgvdnhyst saskdnkrdh mvvtaagtgm 3000 dykkkkrkvs hgsggsgggh mgnsvvtynt ndkcksgkak kmkgtsvvty ttdgckagdv 3060 dakmrskgvk nvvtyttdgc kagkvdadmk rgkdvvtynt dgcksgkakk mkgtsvvtyt 3120 tdgckagdvd akmrskgvkn vvtyttdgck agkvdadmkr gkdvvtyntd gcksgkakkm 3180 kgtsvvtytt dgckagdvda kmrskgvknv vtyttdgcka gkvdadmkrg kdvvtyntdg 3240 cksgkakkmk gtsvvtyttd gckagdvdak mrskgvknvv tyttdgckag kvdadmkrgk 3300 dvvtyntdgc ksgkakkmkg tsvvtyttdg ckagdvdakm rskgvknvvt yttdgckagk 3360 vdadmkrgkd vvtyntdgck sgkakkmkgt svvtyttdgc kagdvdakmr skgvknvvty 3420 ttdgckagkv dadmkrgkdt yntsggkagr ardvssgdyk dhdgdykdhd dykddddkma 3480 gr 3482 <210> twenty two <211> 3482 <212> DNA <213> Artificial sequence <220> <223> PPR vector <220> <221> misc_feature <222> (2858)..(2858) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2914)..(2914) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2926)..(2926) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2934)..(2934) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2939)..(2939) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2941)..(2941) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2944)..(2944) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2952)..(2952) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2956)..(2956) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2960)..(2960) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2969)..(2969) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2976)..(2976) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2986)..(2986) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3023)..(3023) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3029)..(3029) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3071)..(3071) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3099)..(3099) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3140)..(3140) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3168)..(3168) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3209)..(3209) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3237)..(3237) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3278)..(3278) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3306)..(3306) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3347)..(3347) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3375)..(3375) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3416)..(3416) <223> n is a, c, g, t or u <220> <221> misc_feature <222> (3442)..(3442) <223> n is a, c, g, t or u <400> 22 atggccggag tgagcaaggg cgaggagctg ttcaccgggg tggtgcccat cctggtcgag 60 ctggacggcg acgtaaacgg ccacaagttc agcgtgtccg gcgagggcga gggcgatgcc 120 acctacggca agctgaccct gaagttcatc tgcaccaccg gcaagctgcc cgtgccctgg 180 cccaccctcg tgaccaccct gacctacggc gtgcagtgct tcagccgcta ccccgaccac 240 atgaagcagc acgacttctt caagtccgcc atgcccgaag gctacgtcca ggagcgcacc 300 atcttcttca aggacgacgg caactacaag acccgcgccg aggtgaagtt cgagggcgac 360 accctggtga accgcatcga gctgaagggc atcgacttca aggaggacgg caacatcctg 420 gggcacaagc tggagtacaa ctacaacagc cacaacgtct atatcatggc cgacaagcag 480 aagaacggca tcaaggtgaa cttcaagatc cgccacaaca tcgaggacgg cagcgtgcag 540 ctcgccgacc actaccagca gaacaccccc atcggcgacg gccccgtgct gctgcccgac aaccactacc tgagcaccca gtccgccctg agcaaagacc ccaacgagaa gcgcgatcac atggtcctgc tggagttcgt gaccgccgcc gggatcactc tcggcatgga cgagctgtac 720 aagccaaaga aaaagagaa ggttagccat ggctccggcg gcagcgggggg agggctccat 780 atgggaaact ccgtggtcac attackacacc gatatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc 960 gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 1260 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 1380 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggttac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1620 gtgaagccca acgtggtgac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggttac acaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggag agggcatcac ccccagcgtg gtcacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggttac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaaa gtggacgagg ccctggagct gttcgacgag 2340 atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc atcacccca gcgtggtgac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 2520 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtcacataca ccaccctgat cgacggcctg tgcaaggccg gcaaggtgga tgaggccctg 2640 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgagctgac ctacaacacc ctgatcagcg gcctgggcaa ggccggcaga gccagagacc cccccgtgct cagtagcggg 2760 gactataagg accacgacgg agactacaag gatcatgata ttgattacaa agacgatgac 2820 gataagatgg ccggacgcta gmagvskgtg vvvdgdvngh ksvsgggdat ygktkcttgk 2880 vwtvtttygv csrydhmkhd ksamgyvrtk ddgnyktrav kgdtvnrkgd kdgnghkyny 2940 nshnvymadk kngkvnkrhn dgsvadhynt gdgvdnhyst saskdnkrdh mvvtaagtgm 3000 dykkkkrkvs hgsggsgggh mgnsvvtynt ddgcksgkak kmkgtsvvty ttdgckagdv 3060 dakmrskgvk nvvtyttdgc kagkvdadmk rgkdvvtynt dgcksgkakk mkgtsvvtyt 3120 tdgckagdvd akmrskgvkn vvtyttdgck agkvdadmkr gkdvvtyntd gcksgkakkm 3180 kgtsvvtytt dgckagdvda kmrskgvknv vtyttdgcka gkvdadmkrg kdvvtyntdg 3240 cksgkakkmk gtsvvtyttd gckagdvdak mrskgvknvv tyttdgckag kvdadmkrgk 3300 dvvtyntdgc ksgkakkmkg tsvvtyttdg ckagdvdakm rskgvknvvt yttdgckagk 3360 vdadmkrgkd vvtyntdgck sgkakkmkgt svvtyttdgc kagdvdakmr skgvknvvty 3420 ttdgckagkv dadmkrgkdt yntsggkagr ardvssgdyk dhdgdykdhd dykddddkma 3480 gr 3482 <210> twenty three <211> 3551 <212> DNA <213> Artificial sequence <220> <223> PPR vector <220> <221> misc_feature <222> (2912)..(2912) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2925)..(2925) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2982)..(2982) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2990)..(2990) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2995) (2996) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3007) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3011)..(3011) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3015) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3025)..(3025) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3032)..(3032) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3042)..(3042) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3066)..(3066) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3072)..(3072) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3113)..(3113) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3141)..(3141) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3182)..(3182) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3210)..(3210) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3251)..(3251) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3279)..(3279) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3320)..(3320) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3348)..(3348) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3389)..(3389) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3417)..(3417) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3458)..(3458) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3484)..(3484) <223> n is a, c, g, t or u <400> 23 atggccggag tgtccaaagg cgaggagctg tttaccggcg tcgtgcctat tctggtggag 60 ctggacggcg acgtgaacgg ccacaagttc tccgtgaggg gcgagggcga aggcgatgcc 120 acaaacggca agctgaccct caagttcatc tgcaccactg gtaaactgcc cgttccttgg 180 cccacactgg tgaccacctt cggctacggc gtggcttgtt tctctcgtta ccccgaccat 240 atgaagcagc acgacttctt caagtccgcc atgcccgagg gatacgtgca agaaaggacc 300 atctccttca aggacgatgg cacctacaag accagagccg aggtgaagtt cgagggcgac 360 acactggtga atcgtatcga actgaagggc atcgacttca aagaggacgg caacattctg 420 ggccacaagc tggagtacaa cttcaacagc cactacgtgt acatcaccgc cgataagcag 480 aagaactgca tcaaggccaa cttcaagatt cgtcacaacg tggaggatgg ctccgtgcag 540 ctggccgatc actaccagca gaacacaccc atcggcgatg gacccgtttt actgcccgac 600 aaccactatt tagccacca gagcaagctg tccaaggacc ccaacgaga gcgtgatcat atggtgctgc tcgagtttgt gaccgccgcc ggcatcaccc atggaatgga cgagctgtac 720 aagagccggc tccatatggg aaactccgtg gtcacataca acaccctgat cgacggcctg 780 tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaagggc 840 atcaccccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtgacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac attackacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 1260 gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gttacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg 1440 aagctgttca aggagatgga ggagaagggc atcaccccca gcgtggtcac acaccaca ctgatcgacg gactgtgtaa agccggcgac gtggacgag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg gtgacataca ccaccctgat cgacggcctg 1620. tgcaaggccg gcaaagtgga cgaggccctg gagctgttcg acgagatgaa ggagaggggc 1680. atcaagcccg acgtggtcac atacaacacc ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag atggaggaga agggcatcac ccccagcgtg gttacataca ccacactgat cgacggactg tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag gtggatgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgtg gtcacataca acaccctgat cgacggcctg tgcaagagcg gcaagatcga ggaggccctg aagctgttca aggagatgga ggagaaggggc 2100 atcaccccca gcgtggtcac atacaccaca ctgatcgacg gactgtgtaa agccggcgac 2160 gtggacgaag ccctcgagct gttcaaagag atgcggagca agggcgtgaa gcccaacgtg 2220 gttacataca ccaccctgat cgacggcctg tgcaaggccg gcaaagtgga cgaggccctg 2280 gagctgttcg acgagatgaa ggagaggggc atcaagcccg acgtggtcac atacaacacc 2340 ctgatcgacg gcctgtgcaa gagcggcaag atcgaggagg ccctgaagct gttcaaggag 2400 atggagaga agggcatcac ccccagcgtg gtgacataca ccacactgat cgacggactg 2460 tgtaaagccg gcgacgtgga cgaagccctc gagctgttca aagagatgcg gagcaagggc 2520 gtgaagccca acgtggtcac atacaccacc ctgatcgacg gcctgtgcaa ggccggcaag 2580 gtggatgagg ccctggagct gttcgacgag atgaaggaga ggggcatcaa gcccgacgag 2640 ctgacctaca acaccctgat cagcggcctg ggcaaggccg gcagagccag agacccccc 2700 gtgctcagta gccccaagaa gaacgcaaa gtcgaggatc caaaagaa aaggaaggtt 2760 gaacccca agaaaaagag gaaggtgggt tccgactata aggaccacga cggagactac 2820 aaggatcatg atattgatta caaagacgat gacgataaga tggccccaaa gaaagcgg 2880 aaggtcggac gctagmagvs kgtgvvvdgd vnghksvrgg gdatngktkc ttgkvwtvtt 2940 gygvacsryd hmkhdksamg yvrtskddgt yktravkgdt vnrkgdkdgn ghkynnshyv 3000 ytadkkncka nkrhnvdgsv adhyntgdgv dnhyshsksk dnkrdhmvvt aagthgmdyk 3060 srhmgnsvvt yntdgcksgk akkmkgtsvv tyttdgckag dvdakmrskg vknvvtyttd 3120 gckagkvdad mkrgkdvvty ntdgcksgka kkmkgtsvvt yttdgckagd vdakmrskgv 3180 knvvtyttdg ckagkvdadm krgkdvvtyn tdgcksgkak kmkgtsvvty ttdgckagdv 3240 dakmrskgvk nvvtyttdgc kagkvdadmk rgkdvvtynt dgcksgkakk mkgtsvvtyt 3300 tdgckagdvd akmrskgvkn vvtyttdgck agkvdadmkr gkdvvtyntd gcksgkakkm 3360 kgtsvvtytt dgckagdvda kmrskgvknv vtyttdgcka gkvdadmkrg kdvvtyntdg 3420 cksgkakkmk gtsvvtyttd gckagdvdak mrskgvknvv tyttdgckag kvdadmkrgk 3480 dtyntsggka grardvsskk krkvdkkkrk vdkkkrkvgs dykdhdgdyk dhddykdddd 3540 kmakkkrkvg r 3551 <210> twenty four <211> 3550 <212> DNA <213> Artificial sequence <220> <223> PPR vector <220> <221> misc_feature <222> (2912)..(2912) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2925)..(2925) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2982)..(2982) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2990)..(2990) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (2995) (2996) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3007) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3011)..(3011) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3015) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3025)..(3025) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3032)..(3032) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3042)..(3042) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3066)..(3066) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3072)..(3072) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3112)..(3112) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3140)..(3140) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3181)..(3181) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3209)..(3209) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3250)..(3250) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3278)..(3278) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3319)..(3319) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3347)..(3347) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3388)..(3388) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3416)..(3416) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3457)..(3457) <223> n is a, c, g, t, or u <220> <221> misc_feature <222> (3483)..(3483) <223> n is a, c, g, t or u <400> 24 atggccggag tgtccaaagg cgaggagctg tttaccggcg tcgtgcctat tctggtggag 60 ctggacggcg acgtgaacgg ccacaagttc tccgtgaggg gcgagggcga aggcgatgcc
END
Claims
1. A polypeptide that is any of the following PPR motifs: The PPR motif is composed of any sequence from sequence number 4 to 7.
2. The polypeptide of claim 1, wherein, The polypeptide is a PPR motif composed of the sequence number 4.
3. The use of the polypeptide of claim 1 or 2 as a polypeptide that is the first PPR motif from the N-terminus in a PPR protein.
4. The application as described in claim 3, which is used to reduce the aggregation of PPR proteins.
5. A PPR protein, which is a protein capable of binding to a target nucleic acid having a sequence having sequence number 1, characterized in that, The protein is encoded by any of the sequences 11 to 16, and the polypeptide that serves as the first PPR motif from the N-terminus is the polypeptide of claim 1 or 2.
6. A fusion protein, which is a fusion protein of at least one selected from the group consisting of fluorescent proteins, nuclear transport signal peptides and tag proteins with the PPR protein of claim 5.
7. A modification method for improving the aggregation properties of a PPR protein composed of a sequence of sequence number 62, which is capable of binding to a target nucleic acid having a sequence having sequence number 1, wherein, The polypeptide that is the first PPR motif from the N-terminus is the polypeptide as described in claim 1 or 2.
8. Use of the PPR protein of claim 5 or the fusion protein of claim 6 in the preparation of a kit for detecting a target nucleic acid having sequence number 1.
9. A nucleic acid encoding the polypeptide of claim 1 or 2, the PPR protein of claim 5, or the fusion protein of claim 6.
10. A vector comprising the nucleic acid of claim 9.
11. A cell that does not include individual human cells, comprising the carrier of claim 10.
Citation Information
Patent Citations
FUSION PROTEIN FOR IMPROVING PROTEIN EXPRESSION FROM TARGET mRNA
WO2017209122A1