DNA-binding proteins using PPR motifs and their uses
PPR motifs are used to create custom DNA-binding proteins with specific amino acid combinations for precise DNA binding and cleavage, addressing the limitations of existing enzymes and enabling targeted genome editing and functional regulation.
Patent Information
- Application Number
- JP2023083041
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-04-22
- Filing Date
- 2023-05-19
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2034-04-22
AI Technical Summary
Existing DNA cleaving enzymes like ZFNs and TALENs are limited by high production costs, inefficiency, and inaccuracies in target gene recognition, necessitating the development of new artificial DNA cleaving enzymes with improved specificity and flexibility.
Utilizing PPR motifs to design custom DNA-binding proteins by applying the RNA recognition code of PPR motifs, enabling selective binding to desired DNA sequences through specific amino acid combinations in PPR proteins, and creating complexes with functional regions for DNA cleavage or transcriptional regulation.
Enables precise and flexible DNA binding and cleavage, allowing for targeted genome editing and functional regulation, with applications in living organisms and transformants.
Smart Images

Figure 0007754430000012 
Figure 0007754430000013 
Figure 0007754430000014
Abstract
Description
[Technical Field]
[0001] The present invention relates to proteins capable of selectively or specifically binding to intended DNA bases or DNA sequences. The present invention utilizes a pentatricopeptide repeat (PPR) motif. The present invention can be used to identify and design DNA-binding proteins, identify target DNAs of proteins with PPR motifs, and regulate DNA function. The present invention is useful in fields such as medicine and agriculture. The present invention also relates to a new DNA-cleaving enzyme that utilizes a protein containing a PPR motif and a complex with a protein that defines a functional region. [Background technology]
[0002] In recent years, various analyses have revealed that nucleic acid-binding proteins bind to specific sequences, and this sequence-specific binding has been established and is now being used. By utilizing this sequence-specific binding, it is becoming possible to analyze the intracellular location of target nucleic acids (DNA or RNA), remove target DNA sequences, or control (activate or inactivate) the expression of downstream protein-coding genes.
[0003] Research and development is being conducted on zinc finger proteins (Non-Patent Document 1, Non-Patent Document 2), TAL effectors (TALE, Non-Patent Document 3, Patent Document 1), and CRISPR (Non-Patent Document 4, Non-Patent Document 5) as protein engineering materials that act on DNA, but the types of such protein factors are still very limited.
[0004] For example, zinc finger nucleases (ZFNs), known as artificial DNA cleaving enzymes, are chimeric proteins in which a single DNA cleavage domain from a bacterial DNA cleaving enzyme (e.g., FokI) is linked to a portion that recognizes base sequences in units of 3-4 bases, consisting of 3-6 linked zinc fingers that specifically recognize and bind to 3-4 base DNA (Non-Patent Document 2). Such chimeric proteins are based on the fact that zinc finger domains are protein domains known to bind to DNA, and that many transcriptional regulators possess this domain and regulate gene expression by binding to specific DNA sequences. Using two of these ZFNs, each with three zinc fingers, theoretically allows for cleavage at one site per approximately 70 billion bases.
[0005] However, this method using ZFNs has not been widely adopted due to the high cost of production. Furthermore, the efficiency of selecting functional ZFNs has been suggested to be problematic. Furthermore, a zinc finger domain consisting of n zinc fingers tends to recognize the sequence (GNN)n, which limits the flexibility of target gene sequences.
[0006] Meanwhile, TALENs (TAL effectors), proteins consisting of a combinatorial sequence of modular parts capable of recognizing individual bases, have been developed by combining the DNA cleavage domain of a bacterial DNA cleavage enzyme (e.g., FokI) with TAL effectors (TALEs), and are being investigated as artificial enzymes to replace ZFNs (Non-Patent Document 3). TALENs are enzymes that combine the DNA-binding domain of a transcription factor found in the plant pathogenic bacterium Xanthomonas with the DNA cleavage domain of the DNA restriction enzyme FokI. TALENs are known to bind to adjacent DNA sequences, form dimers, and cleave double-stranded DNA. The DNA-binding domain of TALEs, discovered in bacteria that infect plants, recognizes a single base with a combination of two amino acids in a 34-amino acid TALE motif. Therefore, the binding affinity to target DNA can be selected by selecting the repeating structure of the TALE module. TALENs, which utilize DNA-binding domains with these characteristics, have the same advantage as ZFNs in that they can introduce mutations into target genes. However, their major advantages over ZFNs are that they offer significantly greater flexibility in target gene (base sequence) and can encode the binding bases.
[0007] However, because the complete three-dimensional structure of TALENs has not yet been elucidated, it is currently not possible to identify the DNA cleavage site of TALENs. Therefore, compared to ZFNs, TALENs have the drawback of being inaccurate and inconsistent in their cleavage site, leading to the possibility of cleaving similar sequences as well. This has led to the problem that the target base sequence cannot be accurately cleaved by DNA cleaving enzymes. For these reasons, there is a need for the development and provision of new artificial DNA cleaving enzymes that do not have the above-mentioned drawbacks.
[0008] Genome sequence information has identified a large family of proteins, numbering as many as 500 in plants alone: PPR proteins (proteins with pentatricopeptide repeat (PPR) motifs) (Non-Patent Document 6). PPR proteins are nuclear-encoded but are known to act in a gene-specific manner, primarily in organelles (chloroplasts and mitochondria) at the RNA level, including regulation, cleavage, translation, splicing, RNA editing, and RNA stability. PPR proteins typically have a structure consisting of approximately 10 consecutive PPR motifs, a 35-amino acid motif with low conservation. It is believed that the combination of PPR motifs is responsible for sequence-selective binding to RNA. Most PPR proteins consist of only approximately 10 repeats of the PPR motif, and in many cases, the domain required for catalytic activity is missing. Therefore, PPR proteins are thought to function as RNA adaptors (Non-Patent Document 7).
[0009] Generally, the binding between proteins and DNA and the binding between proteins and RNA are based on different molecular mechanisms, and DNA-binding proteins generally do not bind to RNA, and conversely, RNA-binding proteins generally do not bind to DNA. For example, in the case of Pumilio proteins, which are known as RNA-binding factors and can encode recognition RNAs, binding to DNA has not been reported (Non-Patent Documents 8 and 9).
[0010] However, in the course of examining the properties of various types of PPR proteins, it became clear that some types of PPR proteins are suggested to function as DNA-binding factors.
[0011] Wheat p63 is a PPR protein with nine PPR motifs, and gel shift assays suggest that it binds to DNA in a sequence-specific manner (Non-Patent Document 10).
[0012] The Arabidopsis GUN1 protein has 11 PPR motifs, and pull-down assays suggest that it binds to DNA (Non-Patent Document 11).
[0013] Arabidopsis pTac2 (a protein with 15 PPR motifs, Non-Patent Document 12) and Arabidopsis DG1 (a protein with 10 PPR motifs, Non-Patent Document 13) have been shown by run-on assay to be directly involved in transcription, which uses DNA as a template to produce RNA, and are thought to bind to DNA.
[0014] Arabidopsis GRP23 (a protein with 11 PPR motifs, non-patent document 14) gene knockout strains exhibit embryonic lethal phenotypes, but the protein has been shown to physically interact with the major subunit of eukaryotic RNA transcription polymerase 2, a DNA-dependent RNA transcriptase, suggesting that GRP23 also plays a role in DNA binding.
[0015] However, the binding of these PPR proteins to DNA has only been indirectly suggested, and there is insufficient evidence that they actually bind in a sequence-specific manner. Furthermore, even if these proteins bind to DNA in a sequence-specific manner, it is generally believed that the binding between proteins and DNA and the binding between proteins and RNA are based on different molecular mechanisms, so it is not even possible to predict what specific sequence rules are involved in the binding. [Prior art documents] [Patent documents]
[0016] [Patent Document 1] WO2011 / 072246 [Patent Document 2] WO2011 / 111829 [Non-patent literature]
[0017] [Non-Patent Document 1] Maeder, ML, et al. (2008). Rapid "open-source" engineeringof customized zinc-finger nucleases for highly efficient gene modification. Mol.Cell 31, 294-301. [Non-patent document 2] Urnov, FD, et al., (2010) Genome editing with engineered zinc finger nucleases, Nature Review Genetics, 11, 636-646 [Non-patent document 3] Miller, JC, et al. (2011). A TALE nuclease architecture for efficient genome editing. Nature biotech. 29, 143-148. [Non-patent document 4] Mali P, et al. (2013) RNA-guided human genome engineering via Cas9. Science. 339, 823-826. [Non-Patent Document 5] Cong L, et al. (2013) Multiplex genome engineering using CRISPR / Cas systems. Science. 339, 819-823 [Non-patent document 6] Small, ID, and Peeters, N. (2000). The PPR motif - a TPR-related motif prevalent in plant organellar proteins. Trends Biochem. Sci. 25, 46-47. [Non-Patent Document 7] Woodson, JD, and Chory, J. (2008). Coordination of gene expression between organellar and nuclear genomes. Nature Rev. Genet. 9, 383-395. [Non-patent document 8] Wang, X., et al. (2002). Modular recognition of RNA by a human pumilio-homology domain. Cell 110, 501-512. [Non-Patent Document 9] Cheong, CG, and Hall, TM (2006). Engineering RNA sequence specificity of Pumilio repeats. Proc. Natl. Acad. Sci. USA 103, 13635-13639. [Non-Patent Document 10] Ikeda TM and Gray MW (1999) Characterization of a DNA-binding protein implicated in transcription in wheat mitochondria. Mol Cell Biol19(12):8113-8122 [Non-Patent Document 11] Koussevitzky S, et al. (2007) Signals from chloroplasts converge to regulate nuclear gene expression. Science 316: 715-719. [Non-Patent Document 12] Pfalz J, et al. (2006) pTAC2, -6, and -12 are components of the transcriptionally active plastid chromosome that are required for plastid gene expression. Plant Cell 18: 176-197. [Non-Patent Document 13] Chi W, et al. (2008) The pentratricopeptide repeat proteinDELAYED GREENING1 is involved in the regulation of early chloroplast development and chloroplast gene expression in Arabidopsis. Plant Physiol. 147: 573-584. [Non-Patent Document 14] Ding YH, et al. (2006) Arabidopsis GLUTAMINE-RICH PROTEIN23 is essential for early embryogenesis and encodes a novel nuclear PPR motif protein that interacts with RNA polymerase II subunit III. Plant Cell 18: 815-830. Summary of the Invention [Problem to be solved by the invention]
[0018] The present inventors have predicted that the properties of PPR proteins (proteins with PPR motifs) as RNA adaptors are determined by the properties of each PPR motif that constitutes the PPR protein and the combination of multiple PPR motifs, and have proposed a method for modifying RNA-binding proteins using these PPR motifs (Patent Document 2). They have demonstrated that PPR motifs and RNA bind in a one-to-one correspondence, that consecutive PPR motifs recognize consecutive RNA bases in the RNA sequence, and that RNA recognition is determined by a combination of three specific amino acids out of the 35 amino acids that constitute the PPR motif. They have filed patent applications for a method for designing custom RNA-binding proteins using the RNA recognition code of the PPR motif and its use (PCT / JP2012 / 077274; Yagi, Y., et al. (2013) PLoS One, 8, e57286; and Barkan, A., et al. (2012) PLoS Genet., 8, e1002910).
[0019] It has generally been thought that protein-DNA binding and protein-RNA binding are based on different molecular mechanisms. In contrast, we predicted that the RNA recognition rules of PPR motifs could also be applied to DNA recognition. We analyzed PPR proteins that function in DNA binding and aimed to identify PPR proteins with such characteristics. Furthermore, we used the PPR proteins thus obtained that can specifically bind to DNA to prepare custom DNA-binding proteins that bind to desired sequences. We also aimed to provide new artificial enzymes by using these proteins together with proteins that define functional regions, and to provide new artificial DNA-cleaving enzymes by using these proteins together with DNA-cleaving active regions as functional regions. [Means for solving the problem]
[0020] In the case of PPR proteins, various domain search programs (Pfam, Prosite, Interpro, etc.) have revealed that there is no particular distinction between the PPR motifs found in general RNA-binding PPR proteins and the PPR motifs found in the several types of DNA-binding PPR proteins mentioned above. Therefore, it is thought that PPR proteins may contain amino acids (amino acid groups) that determine their binding to DNA or RNA, in addition to the amino acids necessary for nucleic acid recognition.
[0021] In PCT / JP2012 / 077274, the present inventors revealed that RNA-binding PPR motifs bind to RNA in a one-to-one correspondence, and that consecutive PPR motifs recognize consecutive RNA bases in the RNA sequence. They also revealed that base-selective binding to RNA is determined by the combination of three specific amino acids among the 35 amino acids that make up the PPR motif (i.e., the first and fourth amino acids (AA-1 and AA-4) of the first helix (Helix A) of the two α-helical structures that make up the motif, and the second amino acid from the C-terminus (AA "ii" (-2)). They then filed a patent application for a method for designing custom RNA-binding proteins that utilize the RNA recognition code of the PPR motif and its use.
[0022] Thus, among the PPR proteins, those suggested to bind to DNA, such as wheat p63 (Non-Patent Document 11; the amino acid sequence of the homologous protein in Arabidopsis is SEQ ID NO: 1), Arabidopsis GUN1 protein (Non-Patent Document 12; the amino acid sequence is SEQ ID NO: 2), Arabidopsis pTac2 (Non-Patent Document 13; the amino acid sequence is SEQ ID NO: 3), DG1 (Non-Patent Document 14; the amino acid sequence is SEQ ID NO: 4), and Arabidopsis GRP23 (Non-Patent Document 15; the amino acid sequence is SEQ ID NO: 5), were compared in terms of amino acid occurrence frequency with the RNA-binding motif for three amino acids (AA 1, AA 4, and AA "ii"(-2)) that are thought to be important for targeting RNA and that serve as the nucleic acid recognition code. This revealed that the trends in amino acid occurrence frequency between the PPR motifs of these PPR proteins suggested to bind to DNA and the RNA-binding motif were nearly identical.
[0023] This suggests that the nucleic acid recognition code of RNA-binding PPR motifs can also be applied to DNA-binding PPR motifs. Thymine (T), also known as 5-methyluracil, is a uracil (U) derivative with a methylated carbon at the 5th position of uracil (U). Based on the properties of the bases that make up such nucleic acids, it was suggested that the combination of amino acids that recognizes uracil (U) in RNA-binding PPR motifs could also be used to recognize thymine (T) in DNA.
[0024] Based on these findings, we used the DNA-binding PPR proteins p63 (amino acid sequence of SEQ ID NO: 1), Arabidopsis GUN1 protein (amino acid sequence of SEQ ID NO: 2), Arabidopsis pTac2 (amino acid sequence of SEQ ID NO: 3), DG1 (amino acid sequence of SEQ ID NO: 4), and Arabidopsis GRP23 (amino acid sequence of SEQ ID NO: 5) as templates and applied the findings obtained from the investigation of RNA-binding PPR motifs to these PPR proteins to demonstrate that custom DNA-binding proteins that bind to any DNA sequence can be created by arranging amino acids at three positions (AA 1, AA 4, and AA "ii"(-2)).
[0025] That is, the present inventors have accomplished the present invention by providing a protein capable of binding to DNA bases selectively or in a DNA base sequence-specific manner, which contains a plurality of PPR motifs, preferably 2 to 30, more preferably 5 to 25, and most preferably 9 to 15, each of which is represented by the amino acid sequence of SEQ ID NO: 1, the amino acid sequence of SEQ ID NO: 2, the amino acid sequence of SEQ ID NO: 3, the amino acid sequence of SEQ ID NO: 4, or the amino acid sequence of SEQ ID NO: 5, in which three amino acids (AA at position 1, AA at position 4, and AA at position "ii"(-2)) are replaced with specific amino acids as described below.
[0026] The present invention provides the following: [1] A PPR motif having the structure of Formula 1 below:
[0027] [ka]
[0028] (In formula 1: Helix A is a portion capable of forming an α-helical structure; X is absent or a moiety consisting of 1 to 9 amino acids in length; Helix B is a portion capable of forming an α-helical structure; and L is a portion consisting of 2 to 7 amino acids in length. A protein comprising one or more of: A PPR motif (M n )but, the first amino acid of Helix A is amino acid 1 (AA-1), and the fourth amino acid is amino acid 4 (AA-4), and PPR motif (M n ) and the next PPR motif (M n+1 ) is present (when there is no amino acid insertion between the PPR motifs), n ) the second amino acid from the end (C-terminal side) of the amino acids that make up the nucleotide sequence; PPR motif (M n ) and the next PPR motif (M n+1 ), if a non-PPR motif of 1 to 20 amino acids is found between the two, the next PPR motif (M n+1 ) two amino acids upstream of the first amino acid, i.e., the second amino acid; or PPR motif (M n ) at the C-terminal side of the following PPR motif (M n+1 ) is not found, or the next PPR motif (M n+1 ), if 21 or more amino acids constituting a non-PPR motif are found between the PPR motif (M n ) -2nd amino acid from the end (C-terminal side) of the amino acids that make up When "ii"(-2) is the amino acid ("ii"(-2) AA), A protein that can bind to DNA base-selectively or DNA base sequence-specifically, which is a PPR motif that has three amino acids, AA 1, AA 4, and AA "ii"(-2), in a specific amino acid combination that corresponds to the target DNA base or target DNA base sequence. [2] The protein according to [1], wherein the combination of the three amino acids AA 1, AA 4, and AA "ii"(-2) corresponds to a target DNA base or a target DNA base sequence, and the combination of amino acids is any of the following: (1-1) When AA 4 is glycine (G), AA 1 can be any amino acid, and AA ii (-2) is aspartic acid (D), asparagine (N), or serine (S); (1-2) When AA at position 4 is isoleucine (I), AA at position 1 and AA at position “ii” (-2) can both be any amino acid; (1-3) When AA at position 4 is leucine (L), AA at position 1 and AA at position “ii” (-2) can both be any amino acid; (1-4) When AA 4 is methionine (M), AA 1 and AA “ii” (−2) can both be any amino acid; (1-5) When AA 4 is asparagine (N), AA 1 and AA “ii” (−2) can both be any amino acid; (1-6) When AA 4 is proline (P), AA 1 and AA “ii” (−2) can both be any amino acid; (1-7) When AA 4 is serine (S), AA 1 and AA “ii” (-2) can both be any amino acid; (1-8) When AA 4 is threonine (T), AA 1 and AA “ii” (−2) can both be any amino acid; (1-9) When AA at position 4 is valine (V), AA at position 1 and AA at position “ii” (−2) can both be any amino acid; Protein, determined based on. [3] The protein according to [1], wherein the combination of the three amino acids AA 1, AA 4, and AA "ii"(-2) corresponds to a target DNA base or a target DNA base sequence, and the combination of amino acids is any of the following: (2-1) When the three amino acids AA 1, AA 4, and AA “ii”(-2) are, in order, any amino acid, glycine, and aspartic acid, the PPR motif selectively binds to G; (2-2) When the three amino acids of AA 1, AA 4, and AA “ii”(-2) are glutamic acid, glycine, and aspartic acid, respectively, the PPR motif selectively binds to G; (2-3) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, glycine, and asparagine, the PPR motif selectively binds to A; (2-4) When the three amino acids of AA 1, AA 4, and AA “ii”(-2) are glutamic acid, glycine, and asparagine, respectively, the PPR motif selectively binds to A; (2-5) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, glycine, and serine, the PPR motif selectively binds to A and then to C; (2-6) When the three amino acids at AA 1, AA 4, and AA “ii”(-2) are, in order, any amino acid, isoleucine, any amino acid, the PPR motif selectively binds to T and C; (2-7) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, isoleucine, and asparagine, the PPR motif selectively binds to T and then to C; (2-8) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, leucine, any amino acid, the PPR motif selectively binds to T and C; (2-9) When the three amino acids AA 1, AA 4, and AA “ii”(-2) are, in order, any amino acid, leucine, and aspartic acid, the PPR motif selectively binds to C; (2-10) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, leucine, and lysine, the PPR motif selectively binds to T; (2-11) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, methionine, any amino acid, the PPR motif selectively binds to T; (2-12) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, methionine, and aspartic acid, the PPR motif selectively binds to T; (2-13) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are isoleucine, methionine, and aspartic acid, respectively, the PPR motif selectively binds to T and then C; (2-14) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, any amino acid, the PPR motif selectively binds to C and T; (2-15) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-16) When the three amino acids at AA 1, AA 4, and AA “ii”(−2) are phenylalanine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-17) When the three amino acids at AA 1, AA 4, and AA “ii”(−2) are glycine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-18) When the three amino acids at AA 1, AA 4, and “ii”(-2) are isoleucine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-19) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are threonine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-20) When the three amino acids at AA 1, AA 4, and AA “ii”(−2) are, in order, valine, asparagine, and aspartic acid, the PPR motif selectively binds to T and then to C; (2-21) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are tyrosine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T and then to C; (2-22) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and asparagine, the PPR motif selectively binds to C; (2-23) When the three amino acids at AA 1, AA 4, and “ii”(-2) are isoleucine, asparagine, and asparagine, respectively, the PPR motif selectively binds to C; (2-24) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are serine, asparagine, and asparagine, respectively, the PPR motif selectively binds to C; (2-25) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are valine, asparagine, and asparagine, respectively, the PPR motif selectively binds to C; (2-26) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and serine, the PPR motif selectively binds to C; (2-27) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are valine, asparagine, and serine, respectively, the PPR motif selectively binds to C; (2-28) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and threonine, the PPR motif selectively binds to C; (2-29) When the three amino acids of AA 1, AA 4, and “ii”(-2) are valine, asparagine, and threonine, respectively, the PPR motif selectively binds to C; (2-30) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and tryptophan, the PPR motif selectively binds to C and then to T; (2-31) When the three amino acids AA 1, AA 4, and “ii”(-2) are isoleucine, asparagine, and tryptophan, respectively, the PPR motif selectively binds to T and then C; (2-32) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, proline, any amino acid, the PPR motif selectively binds to T; (2-33) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, proline, and aspartic acid, the PPR motif selectively binds to T; (2-34) When the three amino acids at AA 1, AA 4, and “ii”(-2) are phenylalanine, proline, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-35) When the three amino acids of AA 1, AA 4, and “ii”(−2) are tyrosine, proline, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-36) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, serine, any amino acid, the PPR motif selectively binds to A and G; (2-37) When the three amino acids AA 1, AA 4, and “ii”(−2) AA are, in order, any amino acid, serine, and asparagine, the PPR motif selectively binds to A; (2-38) When the three amino acids of AA 1, AA 4, and “ii”(-2) are phenylalanine, serine, and asparagine, respectively, the PPR motif selectively binds to A; (2-39) When the three amino acids of AA 1, AA 4, and “ii”(-2) are valine, serine, and asparagine, respectively, the PPR motif selectively binds to A; (2-40) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, threonine, any amino acid, the PPR motif selectively binds to A and G; (2-41) When the three amino acids AA 1, AA 4, and “ii”(-2) AA are, in order, any amino acid, threonine, and aspartic acid, the PPR motif selectively binds to G; (2-42) When the three amino acids at AA 1, AA 4, and “ii”(-2) are valine, threonine, and aspartic acid, respectively, the PPR motif selectively binds to G; (2-43) When the three amino acids AA 1, AA 4, and “ii”(-2) AA are, in order, any amino acid, threonine, and asparagine, the PPR motif selectively binds to A; (2-44) When the three amino acids of AA 1, AA 4, and “ii”(-2) are phenylalanine, threonine, and asparagine, respectively, the PPR motif selectively binds to A; (2-45) When the three amino acids of AA 1, AA 4, and “ii”(-2) are isoleucine, threonine, and asparagine, respectively, the PPR motif selectively binds to A; (2-46) When the three amino acids of AA 1, AA 4, and “ii”(-2) are valine, threonine, and asparagine, respectively, the PPR motif selectively binds to A; (2-47) When the three amino acids at AA 1, AA 4, and “ii”(−2) are, in order, any amino acid, valine, any amino acid, the PPR motif binds to A, C, and T, but not to G; (2-48) When the three amino acids at AA 1, AA 4, and “ii”(-2) are isoleucine, valine, and aspartic acid, respectively, the PPR motif selectively binds to C and then to A; (2-49) When the three amino acids AA 1, AA 4, and “ii”(-2) AA are, in order, any amino acid, valine, and glycine, the PPR motif selectively binds to C; (2-50) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, valine, and threonine, the PPR motif selectively binds to T; The protein is determined based on the following: [4] The PPR motif (M n The protein according to any one of [1] to [3], comprising 2 to 30 of the above-mentioned amino acid sequences. [5] The PPR motif (M n The protein according to any one of [1] to [3], comprising 5 to 25 of the above-mentioned amino acid sequences. [6] The PPR motif (M n The protein according to any one of [1] to [3], comprising 9 to 15 of the above-mentioned amino acid sequences. [7] The PPR protein according to [6], which has a sequence selected from the amino acid sequence of SEQ ID NO: 1 having 9 PPR motifs, the amino acid sequence of SEQ ID NO: 2 having 11 PPR motifs, the amino acid sequence of SEQ ID NO: 3 having 15 PPR motifs, the amino acid sequence of SEQ ID NO: 4 having 10 PPR motifs, and the amino acid sequence of SEQ ID NO: 5 having 11 PPR motifs. [8] The PPR motif (M n ) or a DNA base sequence that is a target of a DNA-binding protein, the method comprising: The method is performed by determining the presence or absence of DNA bases corresponding to the combination of three amino acids, AA 1, AA 4, and AA "ii"(-2), of the PPR motif, based on either (1-1) to (1-9) described in [1] or (2-1) to (2-50) described in [3]. [9] A PPR motif (M) as defined in [1], which can bind to a target DNA base or a target DNA having a specific base sequence. n ) comprising one or more (preferably 2 to 30) PPR proteins, the method comprising: The method is performed by determining the presence or absence of a combination of three amino acids, AA 1, AA 4, and AA "ii"(-2), of the PPR motif, depending on the target DNA base or specific bases constituting the target DNA, based on either (1-1) to (1-9) described in [1] or (2-1) to (2-50) described in [1].
[10] A method for controlling DNA function using the protein described in [1].
[11] A complex comprising a region of the protein described in [1] and a functional region linked together.
[12] The complex according to
[11] , wherein the complex comprises a functional region fused to the C-terminus of the protein according to [1].
[13] The complex according to
[11] or
[12] , wherein the functional region is a DNA cleavage enzyme or its nuclease domain, or a transcriptional regulatory domain, and the complex functions as a target sequence-specific DNA cleavage enzyme or a transcriptional regulatory factor.
[14] The complex according to
[13] , wherein the DNA cleaving enzyme is the nuclease domain of FokI (SEQ ID NO: 6).
[15] A method for modifying the genetic material of a cell, comprising the steps of: providing cells containing DNA having a target sequence; and
[11] The method, wherein the complex described in
[11] is introduced into a cell, so that a region consisting of the protein of the complex binds to DNA having a target sequence, and thus the functional region modifies the DNA having the target sequence.
[16] A method for identifying, recognizing, or targeting DNA bases or DNA having a specific base sequence using a PPR protein containing one or more PPR motifs.
[17] The method according to
[16] , wherein the protein contains one or more PPR motifs, three of which are a specific combination of amino acids.
[18] The method according to
[16] or
[17] , wherein the protein contains one or more PPR motifs (Mn) defined in 1. [Effects of the Invention]
[0029] The present invention provides a PPR motif capable of binding to a target DNA base and a protein containing the same. By arranging multiple PPR motifs, proteins capable of binding to target DNA of any sequence or length can be provided.
[0030] The present invention makes it possible to predict and identify the target DNA of any PPR protein, and conversely, to predict and identify PPR proteins that bind to any DNA. Predicting the target DNA sequence clarifies its genetic identity and expands its potential applications. Furthermore, the present invention makes it possible to test the functionality of homologous genes with various amino acid polymorphisms for industrially useful PPR protein genes based on differences in their target DNA sequences.
[0031] Furthermore, the present invention can provide a new DNA cleaving enzyme that utilizes a PPR motif; that is, a functional region protein can be linked to the PPR motif or PPR protein provided by the present invention to prepare a complex containing a protein that has binding activity for a specific nucleic acid sequence and has specific functionality.
[0032] The functional region that can be used in the present invention refers to a region that can impart any of a variety of functions, such as DNA cleavage, transcription, replication, repair, synthesis, and modification. By adjusting the sequence of the PPR motif, which is a feature of the present invention, and determining the base sequence of the target DNA, almost any DNA sequence can be used as a target, and genome editing can be achieved using the functions of the functional region, such as DNA cleavage, transcription, replication, repair, synthesis, and modification.
[0033] For example, when the functional region has a DNA cleavage function, a complex is provided in which the PPR protein portion prepared in the present invention is linked to a DNA cleavage region. Such a complex can function as an artificial DNA cleaving enzyme that recognizes the target DNA base sequence with the PPR protein portion and then cleaves the DNA with the DNA cleavage region. When the functional region has a transcriptional control function, a complex is provided in which the PPR protein portion prepared in the present invention is linked to a DNA transcriptional control region. Such a complex can function as an artificial transcriptional control factor that recognizes the target DNA base sequence with the PPR protein portion and then promotes transcription of the target DNA.
[0034] Furthermore, the present invention can be used to deliver the above-mentioned complex into a living body and make it function, or to create transformants using nucleic acid sequences (DNA, RNA) encoding the proteins obtained by the present invention, or to specifically modify, control, and impart functions to living organisms (cells, tissues, individuals) in various situations. [Brief explanation of the drawings]
[0035] [Figure 1] Figure 1 shows the conserved sequences and amino acid numbers of the PPR motif. (A) The amino acids that make up the PPR motif defined in this invention and their amino acid numbers are shown. (B) The positions of three amino acids (AA 1, AA 4, and AA "ii"(-2)) that control binding base selectivity on the predicted structure are shown. (C) Two structural examples of the PPR motif and the positions of the amino acids on the predicted structure for each are shown. Here, AA 1, AA 4, and AA "ii"(-2) are shown as magenta sticks (dark gray in black-and-white display) on the protein 3D structure diagram. [Figure 2]Figure 2 summarizes the structural outlines of the DNA-binding PPR proteins that function in DNA metabolism: Arabidopsis p63 (amino acid sequence of SEQ ID NO: 1), Arabidopsis GUN1 protein (amino acid sequence of SEQ ID NO: 2), Arabidopsis pTac2 (amino acid sequence of SEQ ID NO: 3), DG1 (amino acid sequence of SEQ ID NO: 4), and Arabidopsis GRP23 (amino acid sequence of SEQ ID NO: 5), as well as an outline of the assay systems demonstrating their DNA binding. [Figure 3] FIG. 3 shows a summary of the amino acid occurrence frequencies of three amino acids (AA 1, AA 4, and AA "ii"(-2)) that are responsible for the nucleic acid recognition code in the PPR motifs of PPR proteins (SEQ ID NOs: 1 to 5) that are suggested to have DNA binding properties and known RNA-binding motifs. [Figure 4-1] Figure 4-1 shows the positions of the PPR motifs contained within (A) Arabidopsis thaliana p63 (amino acid sequence of SEQ ID NO: 1) and (B) Arabidopsis thaliana GUN1 protein (amino acid sequence of SEQ ID NO: 2), as well as the positions of three amino acids (AA 1, AA 4, and AA "ii"(-2)) that are responsible for the nucleic acid recognition code within the PPR motif. [Figure 4-2] Figure 4-2 shows the positions of the PPR motifs contained within Arabidopsis thaliana pTac2 (amino acid sequence of SEQ ID NO: 3) and (D) DG1 (amino acid sequence of SEQ ID NO: 4), as well as the positions of the three amino acids (AA 1, AA 4, and AA "ii"(-2)) that are responsible for the nucleic acid recognition code within the PPR motif. [Figure 4-3]Figure 4-3 shows (E) the position of the PPR motif contained within Arabidopsis thaliana GRP23 (amino acid sequence of SEQ ID NO: 5) and the positions of three amino acids (AA 1, AA 4, and AA "ii"(-2)) that are responsible for the nucleic acid recognition code within the PPR motif. [Figure 5] Evaluation of sequence-specific DNA binding ability of PPR molecules. We created artificial transcription factors by fusing the transcription activation domain VP64 to three types of PPR molecules (potentially DNA-binding), and examined whether they could activate luciferase reporters with their respective target sequences in human cultured cells. [Figure 6] We compared the luciferase activity of pTac2-VP64 and GUN1-VP64 when they were co-transfected with the negative control pminCMV-luc2 and with reporter vectors containing four or eight target sequences. The results showed that the activity of both vectors increased with the number of target sequences, demonstrating that these PPR-VP64 molecules bind specifically to their respective target sequences and function as site-specific transcriptional activators. DETAILED DESCRIPTION OF THE INVENTION
[0036] [PPR motifs and PPR proteins] Unless otherwise specified, the term "PPR motif" as used herein refers to a polypeptide consisting of 30 to 38 amino acids, having an amino acid sequence whose E value obtained using PF01535 in Pfam (http: / / pfam.sanger.ac.uk / ) or PS51375 in Prosite (http: / / www.expasy.org / prosite / ) is a predetermined value or less (preferably E-03) when the amino acid sequence is analyzed using an online protein domain search program (e.g., Pfam, Prosite, Uniprot, etc.). The Uniprot database (http: / / www.uniprot.org) also defines PPR motifs in various proteins.
[0037] The PPR motif of the present invention has low conservation of the amino acid sequence of the PPR motif, but the secondary structure of helix, loop, helix, loop as shown in the diagram below is well conserved.
[0038] [ka]
[0039] The position numbers of the amino acids constituting the PPR motif defined in the present invention are in accordance with the paper by the present inventors (Kobayashi K, et al., Nucleic Acids Res., 40, 2712-2723 (2012)). That is, the position numbers of the amino acids constituting the PPR motif defined in the present invention are almost synonymous with the amino acid numbering of PF01535 in Pfam, but also correspond to the number obtained by subtracting 2 from the amino acid numbering of PS51375 in Prosite (e.g., from number 1 in the present invention to number 3 in PS51375), and also correspond to the number obtained by subtracting 2 from the amino acid numbering of the PPR motif defined in Uniprot.
[0040] Specifically, in the present invention, amino acid 1 is the first amino acid at the beginning of Helix A shown in Formula 1. Amino acid 4 is the fourth amino acid counting from amino acid 1. However, when referring to amino acid "ii" (-2), PPR motif (M n ) and the next PPR motif (M n+1 ) exists (when there is no amino acid insertion between PPR motifs, for example, motifs Nos. 1, 2, 3, 4, 6, and 7 in Figure 4-1 (A)). n ) refers to the second amino acid from the end (C-terminal side) of the amino acids that make up the nucleotide sequence; PPR motif (M n ) and the next PPR motif (M n+1) (for example, motifs 5 and 8 in Figure 4-1 (A), and motifs 1, 2, 7, and 8 in Figure 4-3 (D)). n+1 ) is the amino acid "ii" (-2) that is two amino acids upstream of amino acid 1, i.e., the -2 amino acid (see Figure 1); and PPR motif (M n ) at the C-terminal side of the following PPR motif (M n+1 ) cannot be found (for example, in Figure 4-1, Motif No. 9 in (A) and Motif No. 11 in (B)), or the next PPR motif (M n+1 ), if 21 or more amino acids constituting a non-PPR motif are found between the PPR motif (M n The amino acid at the -2 position from the end (C-terminal side) of the amino acids constituting the nucleotide sequence is designated as amino acid "ii" (-2).
[0041] In the present invention, the term "PPR protein" refers to a PPR protein having multiple PPR motifs described above, unless otherwise specified. In this specification, the term "protein" refers to any substance consisting of a polypeptide (a chain of multiple amino acids peptide-bonded together), unless otherwise specified, and includes those consisting of relatively low molecular weight polypeptides. In the present invention, the term "amino acid" may refer to a normal amino acid molecule, as well as to amino acid residues that make up a peptide chain. It will be clear to those skilled in the art from the context which is being referred to.
[0042] PPR proteins are abundant in plants, with 500 proteins and approximately 5,000 motifs found in Arabidopsis. Many land plants, including rice, poplar, and Selaginella, also contain PPR motifs and PPR proteins with diverse amino acid sequences. Some PPR proteins are known to act as fertility restorers, acting in pollen formation (male gamete formation), and are important genes for F1 seed production for hybrid vigor. Similar to fertility restorers, some PPR proteins have been shown to play a role in speciation. Most PPR proteins are also known to act on RNA in mitochondria or chloroplasts.
[0043] In animals, abnormalities in the PPR protein identified as LRPPRC are known to cause Leigh syndrome French Canadian (LSFC; subacute necrotizing encephalomyelopathy).
[0044] In the present invention, the term "selective" in relation to the binding of a PPR motif to DNA bases means that the binding activity for one DNA base is higher than the binding activity for the other bases, unless otherwise specified. This selectivity can be confirmed by a person skilled in the art through experimental design, or it can be calculated as disclosed in the Examples of the present specification.
[0045] In the present invention, unless otherwise specified, the term "DNA base" refers to the base of the deoxyribonucleotide that constitutes DNA, specifically adenine (A), guanine (G), cytosine (C), or thymine (T). Note that PPR proteins may have selectivity for bases in DNA, but do not bind to nucleic acid monomers.
[0046] Prior to the present invention, a method for searching sequences for conserved amino acids as PPR motifs had been established, but no rules regarding selective binding to DNA bases had been discovered.
[0047] [Discovery provided by the present invention] The present invention provides the following findings.
[0048] (I) Information on the positions of amino acids important for selective binding. Specifically, the first amino acid of Helix A of the PPR motif is designated as amino acid 1, and the fourth amino acid is designated as amino acid 4, and PPR motif (M n ) and the next PPR motif (M n+1 ) is present (when there is no amino acid insertion between the PPR motifs), n ) the second amino acid from the end (C-terminal side) of the amino acids that make up the nucleotide sequence; PPR motif (M n ) and the next PPR motif (M n+1 ), if a non-PPR motif of 1 to 20 amino acids is found between the two, the next PPR motif (M n+1 ) two amino acids upstream of the first amino acid, i.e., the second amino acid; or PPR motif (M n ) at the C-terminal side of the following PPR motif (M n+1 ) is not found, or the next PPR motif (M n+1 ), if 21 or more amino acids constituting a non-PPR motif are found between the PPR motif (M n ) -2nd amino acid from the end (C-terminal side) of the amino acids that make up is the amino acid at position "ii"(-2), the three amino acid combinations of the first and fourth amino acids of helix A, AA1, AA4, and AA at position "ii"(-2) defined above (AA1, AA4, and AA at position "ii"(-2)), are important for selective binding to DNA bases, and these combinations determine which DNA bases will bind.
[0049] The present invention is based on the findings of the present inventors regarding the combination of three amino acids, AA 1, AA 4, and AA "ii"(-2). Specifically, When AA (1-1)4 is glycine (G), AA (1) may be any amino acid, and AA (ii)(-2) is aspartic acid (D), asparagine (N), or serine (S). For example, the combination of AA (1) and AA (ii)(-2) is: Any amino acid in combination with aspartic acid (D) (*GD), Preferably, a combination of glutamic acid (E) and aspartic acid (D) (EGD), Any amino acid in combination with asparagine (N) (*GN), Preferably, a combination of glutamic acid (E) and asparagine (N) (EGN), or Any amino acid in combination with serine (S) (*GS); When (1-2) AA 4 is isoleucine (I), AA 1 and AA "ii"(-2) may be any amino acid. For example, the combination of AA 1 and AA "ii"(-2) is: any amino acid in combination with asparagine (N); (1-3) When AA at position 4 is leucine (L), AA at position 1 and AA at position "ii"(-2) may be any amino acid. For example, the combination of AA at position 1 and AA at position "ii"(-2) is: Any amino acid in combination with aspartic acid (D) (*LD), or Any amino acid in combination with lysine (K) (*LK); (1-4) When AA 4 is methionine (M), AA 1 and AA "ii"(-2) may be any amino acid. For example, the combination of AA 1 and AA "ii"(-2) is: Any amino acid in combination with aspartic acid (D) (*MD), or A combination of isoleucine (I) and aspartic acid (D) (IMD); When (1-5) AA 4 is asparagine (N), AA 1 and AA "ii"(-2) may be any amino acid. For example, the combination of AA 1 and AA "ii"(-2) is: Any amino acid in combination with aspartic acid (D) (*ND), Phenylalanine (F), glycine (G), isoleucine (I), threonine (T), valine (V), or tyrosine (Y) in combination with aspartic acid (D) (FND, GND, IND, TND, VND, or YND), Any amino acid in combination with asparagine (N) (*NN), A combination of isoleucine (I), serine (S), or valine (V) with asparagine (N) (INN, SNN, or VNN), Any amino acid in combination with serine (S) (*NS), · A combination of valine (V) and serine (S) (VNS), Any amino acid in combination with threonine (T) (*NT), · A combination of valine (V) and threonine (T) (VNT), Any amino acid in combination with tryptophan (W) (*NW), or a combination of isoleucine (I) and tryptophan (W) (INW); When (1-6) AA 4 is proline (P), AA 1 and AA "ii"(-2) may be any amino acid. For example, the combination of AA 1 and AA "ii"(-2) is: Any amino acid in combination with aspartic acid (D) (*PD), Phenylanine (F) in combination with aspartic acid (D) (FPD), or A combination of tyrosine (Y) and aspartic acid (D) (YPD); When (1-7) AA 4 is serine (S), AA 1 and AA "ii"(-2) may be any amino acid. For example, the combination of AA 1 and AA "ii"(-2) is: Any amino acid in combination with asparagine (N) (*SN), Phenylalanine (F) in combination with asparagine (N) (FSN), or A combination of valine (V) and asparagine (N) (VSN); When (1-8) AA 4 is threonine (T), AA 1 and AA "ii"(-2) may be any amino acid. For example, the combination of AA 1 and AA "ii"(-2) is: Any amino acid in combination with aspartic acid (D) (*TD), Valine (V) and aspartic acid (D) in combination (VTD), Any amino acid in combination with asparagine (N) (*TN), Phenylalanine (F) and asparagine (N) in combination (FTN), a combination of isoleucine (I) and asparagine (N) (ITN), or Valine (V) and asparagine (N) in combination (VTN); When (1-9) AA 4 is valine (V), AA 1 and AA "ii"(-2) may be any amino acid. For example, the combination of AA 1 and AA "ii"(-2) is: Isoleucine (I) in combination with aspartic acid (D) (IVD), Any amino acid in combination with glycine (G) (*VG), or Any amino acid may be combined with threonine (T) (*VT).
[0050] (II) Information on the correspondence between the combination of the three amino acids AA 1, AA 4, and AA "ii"(-2) and the DNA base. Specifically, it is as follows: (2-1) When the three amino acids AA 1, AA 4, and AA “ii”(-2) are, in order, any amino acid, glycine, and aspartic acid, the PPR motif selectively binds to G; (2-2) When the three amino acids of AA 1, AA 4, and AA “ii”(-2) are glutamic acid, glycine, and aspartic acid, respectively, the PPR motif selectively binds to G; (2-3) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, glycine, and asparagine, the PPR motif selectively binds to A; (2-4) When the three amino acids of AA 1, AA 4, and AA “ii”(-2) are glutamic acid, glycine, and asparagine, respectively, the PPR motif selectively binds to A; (2-5) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, glycine, and serine, the PPR motif selectively binds to A and then to C; (2-6) When the three amino acids at AA 1, AA 4, and AA “ii”(-2) are, in order, any amino acid, isoleucine, any amino acid, the PPR motif selectively binds to T and C; (2-7) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, isoleucine, and asparagine, the PPR motif selectively binds to T and then to C; (2-8) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, leucine, any amino acid, the PPR motif selectively binds to T and C; (2-9) When the three amino acids AA 1, AA 4, and AA “ii”(-2) are, in order, any amino acid, leucine, and aspartic acid, the PPR motif selectively binds to C; (2-10) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, leucine, and lysine, the PPR motif selectively binds to T; (2-11) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, methionine, any amino acid, the PPR motif selectively binds to T; (2-12) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, methionine, and aspartic acid, the PPR motif selectively binds to T; (2-13) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are isoleucine, methionine, and aspartic acid, respectively, the PPR motif selectively binds to T and then C; (2-14) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, any amino acid, the PPR motif selectively binds to C and T; (2-15) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-16) When the three amino acids at AA 1, AA 4, and AA “ii”(−2) are phenylalanine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-17) When the three amino acids at AA 1, AA 4, and AA “ii”(−2) are glycine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-18) When the three amino acids at AA 1, AA 4, and “ii”(-2) are isoleucine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-19) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are threonine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-20) When the three amino acids at AA 1, AA 4, and AA “ii”(−2) are, in order, valine, asparagine, and aspartic acid, the PPR motif selectively binds to T and then to C; (2-21) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are tyrosine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T and then to C; (2-22) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and asparagine, the PPR motif selectively binds to C; (2-23) When the three amino acids at AA 1, AA 4, and “ii”(-2) are isoleucine, asparagine, and asparagine, respectively, the PPR motif selectively binds to C; (2-24) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are serine, asparagine, and asparagine, respectively, the PPR motif selectively binds to C; (2-25) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are valine, asparagine, and asparagine, respectively, the PPR motif selectively binds to C; (2-26) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and serine, the PPR motif selectively binds to C; (2-27) When the three amino acids of AA 1, AA 4, and AA “ii”(−2) are valine, asparagine, and serine, respectively, the PPR motif selectively binds to C; (2-28) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and threonine, the PPR motif selectively binds to C; (2-29) When the three amino acids of AA 1, AA 4, and “ii”(-2) are valine, asparagine, and threonine, respectively, the PPR motif selectively binds to C; (2-30) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, asparagine, and tryptophan, the PPR motif selectively binds to C and then to T; (2-31) When the three amino acids AA 1, AA 4, and “ii”(-2) are isoleucine, asparagine, and tryptophan, respectively, the PPR motif selectively binds to T and then C; (2-32) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, proline, any amino acid, the PPR motif selectively binds to T; (2-33) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, proline, and aspartic acid, the PPR motif selectively binds to T; (2-34) When the three amino acids at AA 1, AA 4, and “ii”(-2) are phenylalanine, proline, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-35) When the three amino acids of AA 1, AA 4, and “ii”(−2) are tyrosine, proline, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-36) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, serine, any amino acid, the PPR motif selectively binds to A and G; (2-37) When the three amino acids AA 1, AA 4, and “ii”(−2) AA are, in order, any amino acid, serine, and asparagine, the PPR motif selectively binds to A; (2-38) When the three amino acids of AA 1, AA 4, and “ii”(-2) are phenylalanine, serine, and asparagine, respectively, the PPR motif selectively binds to A; (2-39) When the three amino acids of AA 1, AA 4, and “ii”(-2) are valine, serine, and asparagine, respectively, the PPR motif selectively binds to A; (2-40) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, threonine, any amino acid, the PPR motif selectively binds to A and G; (2-41) When the three amino acids AA 1, AA 4, and “ii”(-2) AA are, in order, any amino acid, threonine, and aspartic acid, the PPR motif selectively binds to G; (2-42) When the three amino acids at AA 1, AA 4, and “ii”(-2) are valine, threonine, and aspartic acid, respectively, the PPR motif selectively binds to G; (2-43) When the three amino acids AA 1, AA 4, and “ii”(-2) AA are, in order, any amino acid, threonine, and asparagine, the PPR motif selectively binds to A; (2-44) When the three amino acids of AA 1, AA 4, and “ii”(-2) are phenylalanine, threonine, and asparagine, respectively, the PPR motif selectively binds to A; (2-45) When the three amino acids of AA 1, AA 4, and “ii”(-2) are isoleucine, threonine, and asparagine, respectively, the PPR motif selectively binds to A; (2-46) When the three amino acids of AA 1, AA 4, and “ii”(-2) are valine, threonine, and asparagine, respectively, the PPR motif selectively binds to A; (2-47) When the three amino acids at AA 1, AA 4, and “ii”(−2) are, in order, any amino acid, valine, any amino acid, the PPR motif binds to A, C, and T, but not to G; (2-48) When the three amino acids at AA 1, AA 4, and “ii”(-2) are isoleucine, valine, and aspartic acid, respectively, the PPR motif selectively binds to C and then to A; (2-49) When the three amino acids AA 1, AA 4, and “ii”(-2) AA are, in order, any amino acid, valine, and glycine, the PPR motif selectively binds to C; (2-50) When the three amino acids AA 1, AA 4, and AA “ii”(−2) are, in order, any amino acid, valine, and threonine, the PPR motif selectively binds to T; the protein has selective DNA base-binding ability.
[0051] The binding of a specific combination of amino acids at a specific position to a DNA base can be confirmed by experiment. Experiments for this purpose include the preparation of a protein containing a PPR motif or multiple PPR motifs, the preparation of substrate DNA, and a binding test (e.g., gel shift assay). Each experiment is well known to those skilled in the art, and more specific procedures are described below. Regarding the conditions, for example, Patent Document 2 can be referred to.
[0052] [Utilization of PPR motifs and PPR proteins] Identification and Design: One PPR motif recognizes a specific base in DNA, and multiple consecutive PPR motifs can recognize consecutive bases in a DNA sequence. According to the present invention, by appropriately selecting amino acids at specific positions, it is possible to select or design PPR motifs that are selective for A, T, G, and C, respectively, and proteins containing an appropriate sequence of such PPR motifs can recognize the corresponding specific sequence. Therefore, according to the present invention, it is possible to predict natural PPR proteins that selectively bind to DNA having a specific base sequence. - Identify and, conversely, predict the DNA targets of PPR protein binding -Can identify and predict targets Identification is useful in that it helps clarify genetic entities and may broaden the availability of targets.
[0053] Furthermore, the present invention makes it possible to design PPR motifs capable of selectively binding to desired DNA bases, and proteins having multiple PPR motifs capable of sequence-specifically binding to desired DNA. When designing, the portions of the PPR motif other than the amino acids at key positions can be determined by reference to the sequence information of natural PPR motifs in DNA-binding PPR proteins, such as those set forth in SEQ ID NOs: 1 to 5. Alternatively, the design may be performed by using the natural PPR motif as a whole and substituting only the amino acids at the relevant positions. The number of repeats of the PPR motif can be appropriately determined depending on the target sequence, and can be, for example, 2 or more, 2 to 30, more preferably 5 to 25, and most preferably 9 to 15.
[0054] In designing a PPR motif, consideration may be given to combinations other than those of the amino acids at AA 1, AA 4, and AA "ii"(-2). For example, consideration of the amino acids at 8 and 12, as described in the aforementioned Patent Document 2, may be important for exhibiting DNA-binding activity. According to the studies of the present inventors, the amino acid at 8 of a certain PPR motif and the amino acid at 12 of the same PPR motif may cooperate in DNA binding. The amino acid at 8 can be a basic amino acid, preferably lysine, or an acidic amino acid, preferably aspartic acid, and the amino acid at 12 can be a basic amino acid, a neutral amino acid, or a hydrophobic amino acid.
[0055] Designed motifs or proteins can be prepared by methods well known to those skilled in the art. That is, the present invention provides PPR motifs that selectively bind to specific DNA bases, focusing on the combination of amino acids AA 1, AA 4, and AA "ii"(-2), and PPR proteins that specifically bind to DNA having a specific sequence. Such motifs and proteins can be prepared in relatively large quantities by methods well known to those skilled in the art. Such methods may include determining the nucleic acid sequence encoding the desired motif or protein from its amino acid sequence, cloning it, and preparing a transformant that produces the desired motif or protein.
[0056] Preparation of the complex and its use: The PPR motifs or PPR proteins provided by the present invention can be linked to functional regions to form complexes. A functional region generally refers to a portion that has a specific biological function in vivo or in a cell, such as an enzymatic function, catalytic function, inhibitory function, or activating function, or a portion that functions as a label. Such a portion can be, for example, a protein, peptide, nucleic acid, physiologically active substance, or drug.
[0057] In the present invention, by linking a functional region to a PPR protein, the target DNA sequence binding function exerted by the PPR protein and the function exerted by the functional region can be exerted in combination. For example, by using a protein with DNA cleavage function (e.g., a restriction enzyme such as FokI) or its nuclease domain as the functional region, the complex can function as an artificial DNA cleaving enzyme.
[0058] To produce such a complex, techniques generally available in the art can be used, including methods of synthesizing the complex as a single protein molecule and methods of separately synthesizing multiple protein components and then combining those components to form the complex.
[0059] For example, in the case of synthesizing a complex as a single protein molecule, a protein complex can be designed in which a cleavage enzyme is fused to the C-terminus of a PPR protein via an amino acid linker, an expression vector construct for expressing the protein complex is constructed, and the complex of interest can be expressed from the construct. Examples of such a preparation method include the method described in Japanese Patent Application No. 2011-242250.
[0060] The PPR protein and the functional region protein may be linked by any linking means known in the art, such as linking via an amino acid linker, linking via specific affinity such as avidin-biotin, or linking via other chemical linkers.
[0061] The functional region that can be used in the present invention refers to a region that can impart any of a variety of functions, such as DNA cleavage, transcription, replication, repair, synthesis, and modification. By adjusting the sequence of the PPR motif, which is a feature of the present invention, and determining the base sequence of the target DNA, almost any DNA sequence can be used as a target, and genome editing can be achieved using the functions of the functional region, such as DNA cleavage, transcription, replication, repair, synthesis, and modification.
[0062] For example, when the function of the functional region is DNA cleavage, a complex is provided in which the PPR protein portion prepared in the present invention is linked to the DNA cleavage region. Such a complex can function as an artificial DNA cleaving enzyme, recognizing the target DNA base sequence with the PPR protein portion and then cleaving the DNA with the DNA cleavage region.
[0063] An example of a functional region having a cleavage function that can be used in the present invention is a deoxyribonuclease (DNase) that functions as an endodeoxyribonuclease. Examples of such DNases include endodeoxyribonucleases such as DNase A (e.g., bovine pancreatic ribonuclease A: PDB 2AAS), DNase H, and DNase I, as well as various bacterial restriction enzymes (e.g., FokI (SEQ ID NO: 6)) and their nuclease domains. Such a complex comprising a PPR protein and a functional region does not exist in nature and is therefore novel.
[0064] When the function of the functional region is transcriptional regulation, the present invention provides a complex in which the PPR protein portion prepared in the present invention is linked to a transcriptional regulatory region of DNA. Such a complex can function as an artificial transcriptional regulator that regulates the transcription of the target DNA after recognizing the base sequence of the target DNA with the PPR protein portion.
[0065] The functional region having transcriptional regulatory function that can be used in the present invention may be a domain that activates transcription or a domain that represses transcription. Examples of transcriptional regulatory domains include VP16, VP64, TA2, STAT-6, and p65. Such complexes containing a PPR protein and a transcriptional regulatory domain do not exist in nature and are novel.
[0066] Furthermore, the complex obtained by the present invention may be able to deliver and function a functional region in a DNA sequence-specific manner in vivo or within cells. This makes it possible to perform DNA sequence-specific modification or destruction in vivo or within cells, similar to protein complexes using zinc finger proteins (Non-Patent Documents 1 and 2, supra) or TAL effectors (Non-Patent Document 3, supra; Patent Document 1, supra), thereby imparting new functions such as DNA cleavage and genome editing utilizing its functions. Specifically, a PPR protein in which multiple PPR motifs that can bind to specific bases are linked can recognize a specific DNA sequence. Then, by using the functional region linked to the PPR protein, genome editing of the recognized DNA region can be achieved by utilizing the function of the functional region.
[0067] Furthermore, by binding a drug to a PPR protein that specifically binds to a DNA sequence, it may be possible to deliver the drug to the area surrounding that DNA sequence. Therefore, the present invention also provides a method for delivering a functional substance specifically to a DNA sequence.
[0068] The PPR proteins used in the present invention function to specify editing sites in DNA editing, and it has been revealed that PPR motifs with specific amino acids at residue positions AA 1, AA 4, and AA "ii"(-2) recognize specific bases in DNA and exhibit DNA binding activity. Based on these characteristics, it is expected that PPR proteins of this type with specific amino acids at residue positions AA 1, AA 4, and AA "ii"(-2) will recognize specific bases in DNA, thereby introducing nucleotide polymorphisms or treating diseases or conditions caused by nucleotide polymorphisms. In addition, by combining these proteins with other functional regions as described above, they will be able to cleave DNA and contribute to the modification and improvement of functions for genome editing.
[0069] In addition, an exogenous DNA cleavage enzyme can be fused to the C-terminus of the PPR protein. Alternatively, by improving the DNA base binding selectivity of the N-terminal PPR motif, a DNA sequence-specific DNA cleavage enzyme can be constructed. Furthermore, complexes with a labeling moiety such as GFP can also be used to visualize the desired DNA in vivo. [Example]
[0070] Example 1: Collection of PPR proteins involved in DNA editing and their target sequences With reference to the information disclosed in prior art documents (Non-Patent Documents 11 to 15), the structures and functions of p63 protein (SEQ ID NO: 1), GUN1 protein (SEQ ID NO: 2), pTac2 protein (SEQ ID NO: 3), DG1 protein (SEQ ID NO: 4), and GRP23 protein (SEQ ID NO: 5) were analyzed.
[0071] The PPR motif structures in these proteins were assigned amino acid numbers as defined in this invention, along with information from the Uniprot database (http: / / www.uniprot.org / ). The PPR motifs contained in the five Arabidopsis PPR proteins (SEQ ID NOs: 1 to 5) used in the experiment and their amino acid numbers are shown in Figure 3.
[0072] Specifically, for the aforementioned p63 protein (SEQ ID NO: 1), GUN1 protein (SEQ ID NO: 2), pTac2 protein (SEQ ID NO: 3), DG1 protein (SEQ ID NO: 4), and GRP23 protein (SEQ ID NO: 5), the amino acid occurrence frequencies of three amino acids (AA 1, AA 4, and AA "ii"(-2)) that are responsible for the nucleic acid recognition code in the PPR motif, which are thought to be important when targeting RNA, were compared with those of the RNA-binding motif.
[0073] The Arabidopsis p63 protein (SEQ ID NO: 1) has nine PPR motifs, and the positions of AA 1, AA 4, and AA "ii" (-2) in its amino acid sequence are summarized in the table below and in Figure 3.
[0074] [Table 1]
[0075] The Arabidopsis GUN1 protein (SEQ ID NO: 2) has 11 PPR motifs, and the positions of AA 1, AA 4, and AA "ii" (-2) in its amino acid sequence are summarized in the table below and in Figure 3.
[0076] [Table 2] The Arabidopsis pTac2 protein (SEQ ID NO: 3) has 15 PPR motifs, and the positions of AA 1, AA 4, and AA "ii" (-2) in its amino acid sequence are summarized in the table below and in Figure 3.
[0077] [Table 3]
[0078] The Arabidopsis DG1 protein (SEQ ID NO: 4) has 10 PPR motifs, and the positions of AA 1, AA 4, and AA "ii" (-2) in its amino acid sequence are summarized in the table below and in Figure 3.
[0079] [Table 4] The Arabidopsis GRP23 protein (SEQ ID NO: 5) has 11 PPR motifs, and the positions of AA 1, AA 4, and AA "ii" (-2) in its amino acid sequence are summarized in the table below and in Figure 3.
[0080] [Table 5]
[0081] The amino acid frequencies at these positions were confirmed for each protein and compared with the amino acid frequencies at the same positions in the RNA-binding motif. The results are shown in Figure 2. It was revealed that the amino acid frequency trends between the PPR motifs of these PPR proteins, which are suggested to have DNA binding properties, and the RNA-binding motif are nearly identical. In other words, it was revealed that PPR proteins that function in DNA binding bind to nucleic acids using the same sequence rules as PPR proteins that function in RNA binding, and that the RNA recognition code described in the inventors' pending patent application (PCT / JP2012 / 077274) can be used as a DNA recognition code for PPR proteins that function in DNA binding.
[0082] The DNA-binding PPR motifs that selectively bind to each base were evaluated with reference to the RNA recognition code in a non-patent document (Yagi, Y. et al., Plos One, 2013, 8, e57286). Specifically, the nucleotide occurrence frequency shown in Table 6 was calculated using a chi-square test based on the expected nucleotide frequency calculated from the background frequency. The test was performed for each base (NT), purine, and nucleotide. Pyrimidine (AG or CT; PY), hydrogen-bonding group (AT or GC; HB), or amino The test was performed on the keto form (AC or GT). The significance value was set at P<0.06 (5.E-02; 5% significance level), and if a significant value was obtained in any test, the combination of amino acid 1, amino acid 4, and amino acid "ii" (-2) was selected.
[0083] [Table 6-1] [Table 6-2]
[0084] Table 6 lists amino acid combinations that showed significant base selectivity. These results indicate that PPR motifs containing amino acid 1, amino acid 4, and amino acid "ii"(-2) (represented in the table as (NSRs; 1, 4, ii)) that yielded significant P values confer base-selective binding ability, and that the larger the "positive" value after background subtraction, the higher the base selectivity for that base. Among amino acids 1, 4, and "ii"(-2), amino acid 4 is most strongly involved in base selectivity, amino acid "ii"(-2) is second most involved, and amino acid 1 is the least involved of the three amino acids in base selectivity.
[0085] Example 2: Evaluation of sequence-specific DNA binding ability of PPR molecules In this example, we created artificial transcription factors by fusing the transcription activation domain VP64 to three types of DNA-binding (potentially DNA-binding) PPR molecules: p63, pTac2, and GUN1. We then investigated whether each PPR molecule has sequence-specific DNA binding ability by examining whether it can activate luciferase reporters containing the respective target sequences in human cultured cells (Figure 5).
[0086] (Experimental Method) 1. Construction of PPR-VP64 expression vector Of the coding sequences of p63, pTac2, and GUN1, only the portions corresponding to the PPR motif were artificially synthesized. DNA synthesis was performed using Biomatik's artificial gene synthesis service. The pCS2P vector, which contains a CMV promoter, was used as the backbone vector, and the synthesized PPR sequence was inserted into it. Furthermore, a Flag tag and nuclear localization signal were inserted at the N-terminus of the PPR sequence, and a VP64 sequence was inserted at the C-terminus. The sequences of the constructed p63-VP64, pTac2-VP64, and GUN1-VP64 are shown in SEQ ID NOs: 7 to 9 in the sequence listing.
[0087] 2. Preparation of reporter vectors carrying PPR target sequences A reporter vector (pminCMV-luc2, SEQ ID NO: 10) was constructed by ligating the firefly luciferase gene downstream of the minimal CMV promoter and placing a multicloning site upstream of the promoter. The predicted target sequences of each PPR were inserted into the multicloning site of this vector. The target sequences of each PPR (TCTATCACT for p63, AACTTTCGTCACTCA for pTac2, and AATTTGTCGAT for GUN1; SEQ ID NOs: 11–13 in the Sequence Listing) were determined by predicting the motif-DNA recognition code for DNA-binding PPRs from the motif-RNA recognition code for RNA-binding PPRs. For each PPR, vectors with four and eight target sequences inserted were constructed and used in the following assays. The nucleotide sequences of each vector are shown in SEQ ID NOs: 14–19 in the Sequence Listing.
[0088] 3. Transfection into HEK293T cells The PPR-VP64 expression vector prepared in 1, the firefly luciferase expression vector prepared in 2, and the Promega pRL-CMV vector (a Renilla luciferase expression vector) were transfected using Life Technologies' Lipofectamine LTX. 25 μl of DMEM medium was added to each well of a 96-well plate, followed by a mixture of 400 ng of PPR-VP64 expression vector, 100 ng of firefly luciferase expression vector, and 20 ng of pRL-CMV vector. A mixture of 25 μl of DMEM medium and 0.7 μl of Lipofectamine LTX was then added to each well. The plate was then left to stand at room temperature for 30 minutes, after which 6 × 10 cells suspended in 100 μl of DMEM medium containing 15% fetal bovine serum were transfected. 4 A certain number of HEK293T cells were added and cultured in a CO2 incubator at 37°C for 24 hours.
[0089] 4. Luciferase assay Luciferase assays were performed using the Promega Dual-Glo Luciferase Assay System according to the kit's instructions, and luciferase activity was measured using a Berthold TriStar LB 941 plate reader.
[0090] (Results / Discussion) We compared the luciferase activity of pTac2-VP64 and GUN1-VP64 when they were co-transfected with the negative control pminCMV-luc2 and with reporter vectors containing four or eight target sequences (Table 6, below). Activity comparisons were based on a normalized score (Fluc / Rluc) calculated by dividing the Fluc (firefly luciferase) measurement by the reference Rluc (renilla luciferase) measurement. The results showed that activity increased with increasing target sequence, demonstrating that these PPR-VP64 molecules specifically bind to their respective target sequences and function as site-specific transcription activators.
[0091]
Table 7
Claims
1. A nucleic acid encoding a fusion protein in which a binding region that binds to DNA bases or DNA having a specific base sequence is fused with a functional region, the functional region is a DNA cleavage enzyme or its nuclease domain, or a transcriptional regulatory domain; The binding region is A PPR motif having the structure of Formula 1 below: 【Chemical 1】 (In formula 1: Helix A is a portion capable of forming an α-helical structure; X is absent or a moiety consisting of 1 to 9 amino acids in length; Helix B is a portion capable of forming an α-helical structure; and L is a portion consisting of 2 to 7 amino acids in length.) Contains 9 or more of the following; One PPR motif (M n )but, the first amino acid of Helix A is amino acid 1 (AA-1) and the fourth amino acid is amino acid 4 (AA-4), and ・PPR motif (M n ) and the next PPR motif (M n+1 ) is present (when there is no amino acid insertion between the PPR motifs), n The second amino acid from the end (C-terminal side) of the amino acids constituting the nucleotide sequence; ・PPR motif (M n ) and the next PPR motif (M n+1 ), if a non-PPR motif of 1 to 20 amino acids is found between the two, the next PPR motif (M n+1 ) two amino acids upstream of the first amino acid, i.e., the -2nd amino acid; or ・PPR motif (M n ) at the C-terminal side of the following PPR motif (M n+1 ) is not found, or the next PPR motif (M n+1 ) and 21 or more amino acids that constitute a non-PPR motif are found between the PPR motif (M n ) the second amino acid from the end (C-terminal side) of the amino acids that make up the is the amino acid "ii" (-2) (AA "ii" (-2)), The combination of the three amino acids AA 1, AA 4, and AA "ii" (-2) is one of the following: (2-1) When the three amino acids AA 1, AA 4, and AA "ii" (-2) are, in order, any amino acid, glycine, and aspartic acid, the PPR motif selectively binds to G; (2-2) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are glutamic acid, glycine, and aspartic acid, respectively, the PPR motif selectively binds to G; (2-3) When the three amino acids AA 1, AA 4, and AA "ii" (-2) are, in order, any amino acid, glycine, and asparagine, the PPR motif selectively binds to A; (2-4) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are glutamic acid, glycine, and asparagine, respectively, the PPR motif selectively binds to A; (2-5) When the three amino acids AA 1, AA 4, and AA "ii" (-2) are, in order, any amino acid, glycine, and serine, the PPR motif selectively binds to A and then to C; (2-6) When the three amino acids at AA 1, AA 4, and AA "ii" (-2) are, in order, any amino acid, isoleucine, and any amino acid, the PPR motif selectively binds to T and C; (2-7) When the three amino acids AA 1, AA 4, and AA "ii" (-2) are, in order, any amino acid, isoleucine, and asparagine, the PPR motif selectively binds to T and then to C; (2-8) When the three amino acids at AA 1, AA 4, and AA "ii" (-2) are, in order, any amino acid, leucine, any amino acid, the PPR motif selectively binds to T and C; (2-9) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, leucine, and aspartic acid, the PPR motif selectively binds to C; (2-10) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, leucine, and lysine, the PPR motif selectively binds to T; (2-11) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, methionine, and any amino acid, the PPR motif selectively binds to T; (2-12) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, methionine, and aspartic acid, the PPR motif selectively binds to T; (2-13) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are isoleucine, methionine, and aspartic acid, respectively, the PPR motif selectively binds to T and then C; (2-14) When the three amino acids at AA 1, AA 4, and AA "ii" (-2) are, in order, any amino acid, asparagine, any amino acid, the PPR motif selectively binds to C and T; (2-15) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, asparagine, and aspartic acid, the PPR motif selectively binds to T; (2-16) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are phenylalanine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-17) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are glycine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-18) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are isoleucine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-19) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are threonine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-20) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are, in order, valine, asparagine, and aspartic acid, the PPR motif selectively binds to T and then to C; (2-21) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are tyrosine, asparagine, and aspartic acid, respectively, the PPR motif selectively binds to T and then to C; (2-22) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, asparagine, and asparagine, the PPR motif selectively binds to C; (2-23) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are isoleucine, asparagine, and asparagine, respectively, the PPR motif selectively binds to C; (2-24) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are serine, asparagine, and asparagine, respectively, the PPR motif selectively binds to C; (2-25) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are valine, asparagine, and asparagine, respectively, the PPR motif selectively binds to C; (2-26) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, asparagine, and serine, the PPR motif selectively binds to C; (2-27) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are valine, asparagine, and serine, respectively, the PPR motif selectively binds to C; (2-28) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, asparagine, and threonine, the PPR motif selectively binds to C; (2-29) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are valine, asparagine, and threonine, respectively, the PPR motif selectively binds to C; (2-30) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, asparagine, and tryptophan, the PPR motif selectively binds to C and then to T; (2-31) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are isoleucine, asparagine, and tryptophan, respectively, the PPR motif selectively binds to T and then to C; (2-32) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, proline, any amino acid, the PPR motif selectively binds to T; (2-33) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, proline, and aspartic acid, the PPR motif selectively binds to T; (2-34) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are phenylalanine, proline, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-35) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are tyrosine, proline, and aspartic acid, respectively, the PPR motif selectively binds to T; (2-36) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, serine, any amino acid, the PPR motif selectively binds to A and G; (2-37) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, serine, and asparagine, the PPR motif selectively binds to A; (2-38) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are phenylalanine, serine, and asparagine, respectively, the PPR motif selectively binds to A; (2-39) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are valine, serine, and asparagine, respectively, the PPR motif selectively binds to A; (2-40) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, threonine, any amino acid, the PPR motif selectively binds to A and G; (2-41) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, threonine, and aspartic acid, the PPR motif selectively binds to G; (2-42) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are valine, threonine, and aspartic acid, respectively, the PPR motif selectively binds to G; (2-43) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, threonine, and asparagine, the PPR motif selectively binds to A; (2-44) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are phenylalanine, threonine, and asparagine, respectively, the PPR motif selectively binds to A; (2-45) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are isoleucine, threonine, and asparagine, respectively, the PPR motif selectively binds to A; (2-46) When the three amino acids of AA 1, AA 4, and AA “ii” (−2) are valine, threonine, and asparagine, respectively, the PPR motif selectively binds to A; (2-47) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, valine, any amino acid, the PPR motif binds to A, C, and T, but not to G; (2-48) When the three amino acids at AA 1, AA 4, and AA “ii” (−2) are isoleucine, valine, and aspartic acid, respectively, the PPR motif selectively binds to C and then to A; (2-49) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, valine, and glycine, the PPR motif selectively binds to C; (2-50) When the three amino acids AA 1, AA 4, and AA “ii” (−2) are, in order, any amino acid, valine, and threonine, the PPR motif selectively binds to T; The nucleic acid is defined based on the above formula (1), wherein the fusion protein functions as a target sequence-specific DNA cleavage enzyme or a transcriptional regulator.
2. The nucleic acid of claim 1, wherein the functional region is the nuclease domain of FokI (SEQ ID NO: 6).
3. A vector containing the nucleic acid according to claim 1 or 2.
4. A cell into which the vector according to claim 3 has been introduced.
5. The cell of claim 4, which is Escherichia coli.
Citation Information
Patent Citations
TAL effector-mediated DNA modification
WO2011072246A2
Method for modifying RNA binding protein using PPR motif
WO2011111829A1