Methods for editing plant genomes
The fusion of TALEN enzymes with cytidine deaminase proteins, equipped with localization signals, addresses the inefficiencies in plant genome editing by achieving stable and high-efficiency modification of specific bases in nuclear, plastid, and mitochondrial genomes, improving genetic modification efficiency and regulatory compliance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Filing Date
- 2022-01-21
- Publication Date
- 2026-04-28
AI Technical Summary
Current technologies for editing plant genomes, particularly the plastid and mitochondrial genomes, are inefficient and lack methods for accurately modifying specific single bases, posing challenges in improving plant traits and being subject to international regulations.
A method involving the use of TALEN enzymes fused with cytidine deaminase proteins, equipped with specific localization signals, to target and modify C:G pairs to T:A pairs in plant nuclear, plastid, and mitochondrial genomes, achieving stable and high-efficiency editing.
This approach enables the modification of almost all copies of target bases in plant genomes, enhancing genetic modification efficiency and potentially exempting it from certain international regulations.
Smart Images

Figure 0007852925000003 
Figure 0007852925000004 
Figure 0007852925000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to methods for editing or modifying plant genomes, specifically nuclear genomes, mitochondrial genomes, and plastid genomes. [Background technology]
[0002] Nuclear genome editing or modification is considered an effective method for improving the varieties of higher plants. Furthermore, genomes in plastids such as mitochondria and chloroplasts also contain genes that play important roles, and genome editing of these intracellular organelles is also considered effective in improving plant varieties.
[0003] The plastid genome of higher plants is about 150 kb and contains approximately 120 genes, which are involved in photosynthesis, antibiotic resistance, herbicide resistance, and other processes. Among the plastid genes, for example, key genes of the photosystem. psbA It is a key enzyme in the dark reaction CO2 fixation. rbcL These are important genes that govern plant functions, and improving these genes is expected to contribute to optimizing the use of light energy by plants, increasing food production, increasing bioethanol production and biomass production, and improving the utilization of CO2 as a resource. Gene transfer into the plastid genome has been performed for approximately 30 years. Gene transfer into the plastid genome has advantages that differ from gene transfer into the nuclear genome. For example, because the plastid genome is maternally inherited, it is possible to prevent the spread of recombinant genes through pollen. Also, since gene silencing, which occurs during nuclear genetic recombination, does not occur, the expression of the desired gene product is relatively easy.
[0004] However, introducing foreign genes into the plastid genome is not so easy. Gene introduction requires specialized equipment (e.g., particle guns) and culture techniques. Furthermore, the number of plant species that can be genetically introduced is limited, and even in model plants such as Arabidopsis thaliana and rice, introducing foreign genes into the chloroplast genome is difficult (Non-Patent Documents 1 and 2). Although there have been some successful examples of gene introduction into the plastid genome (e.g., Patent Document 1), it remains a difficult technique. Furthermore, there are currently no practical technologies for genome editing that modify only a specific single base in the plastid genome. Recombinant plants created through the aforementioned gene introduction are internationally regulated under the Cartagena Protocol. In contrast, modifying only a specific single base in the plastid genome that is naturally present in plants may be exempt from the Cartagena Protocol, although the treatment varies from country to country. Therefore, the development of technologies that modify only a specific single base in the plastid genome, rather than introducing genes into the plastid genome, is eagerly awaited.
[0005] The plant mitochondrial genome encodes not only genes involved in the electron transport chain, ATP synthesis, and mitochondrial gene translation, but also many open reading frames (ORFs) with unknown functions. The limited utilization and characterization of plant mitochondrial genomes are thought to be partly due to the limited tools available for modification and the difficulty in identifying single nucleotide polymorphisms (SNPs) in the genome that affect crop traits after modification. To date, stable gene introduction into the mitochondrial genome has been achieved using the particle gun method in two single-celled organisms, the green alga Chlamydomonas (Non-Patent Literature 3) and yeast (Non-Patent Literature 4 and 5), but there have been no successful examples of stable transformation (gene introduction) of the mitochondrial genome of higher plants.
[0006] Recently, Mok et al. Burkholderia cenocepaciaThe cytidine deaminase (CD) gene of the DddA protein was split into two, and each was fused with a DNA-binding domain of a uracil glycosylase inhibitor (UGI) and a transcription activator-like effector (TALE). These resulting proteins were transiently expressed in mammalian cells (Non-Patent Literature 6). As a result, we succeeded in replacing target C:G pairs with T:A pairs in the mitochondrial genome. This conversion from C:G pairs to T:A pairs occurred in up to 50% of the mitochondrial genome within the cell.
[0007] Furthermore, Kang et al. applied Mok et al.'s technology to substitute target base pairs in the mitochondrial genome of lettuce and rapeseed callus (conversion from C:G to T:A), and transiently expressed a fusion protein of UGI and TALE in lettuce and rapeseed callus, reporting that the mitochondrial genome editing frequency was up to approximately 25% (Non-Patent Literature 7).
[0008] As described above, although single-nucleotide editing technology for plant genomes is improving year by year, its editing efficiency is currently low, and further technological improvements are needed. [Prior art documents] [Patent Documents]
[0009] [Patent Document 1] Japanese Patent Publication No. 2009-225721 [Non-patent literature]
[0010] [Non-Patent Document 1] Yu et al., Plant physiology 175, 186-193 2017. [Non-Patent Document 2] Ruf et al., Nature Plants 5, 282-289 2019. [Non-Patent Document 3] Remacle et al., Proc. Natl. Acad. Sci. 103, 4771 - 4776 2006. [Non - Patent Document 4] Fox et al., Proc. Natl. Acad. Sci. 85, 7288 - 7292 1988. [Non - Patent Document 5] Johnston et al., Science 240, 1538 - 1541 1988. [Non - Patent Document 6] Mok et al., Nature 583, 631 - 637 2020. [Non - Patent Document 7] Kang et al., Nat. Plants 7, 899 - 905 2021. [Non - Patent Document 8] Gualberto et al., Biochimie 100, 107 - 120 2014. [Non - Patent Document 9] Smith et al., Proc Natl Acad Sci USA 100, 892 - 897 2003 [Summary of the Invention] [Problems to be Solved by the Invention]
[0011] In view of the above circumstances, an object of the present invention is to provide a method for editing or modifying a plant genome, that is, a nuclear genome, plastid (e.g., chloroplast) genome, and mitochondrial genome of a plant, particularly a method for accurately and highly efficiently editing or modifying a target single base. [Means for Solving the Problems]
[0012] The present inventors intensively studied whether the technique reported by Mok et al. (Non - Patent Document 6) can be used for editing the nuclear genome, plastid genome, and mitochondrial genome of plants. First, the inventors designed DNA-binding sequences called TALE repeats, which are used in the genome editing enzyme TALEN (transcription activator-like effector nuclease) that recognize 7 to 21 bp intervals before and after a 10-20 bp interval containing a single base that is the target of editing. They then designed protein sequences (TALECD) by fusing these repeats with a pair of halves of DddA cytidine deaminase. Next, we constructed expression vectors for each of these proteins (vectors that stably introduce the DNA encoding each of the three peptide-modified proteins into the nuclear genome) by adding a nuclear localization signal (nTALECD), a chloroplast localization signal (ptpTALECD), or a mitochondrial localization signal (mtpTALECD) to these two proteins, and transformed the nucleus of plant stem cells with these vectors (the DNA encoding each TALECD is incorporated into the plant nuclear genome DNA, allowing for the stable (not transient) expression of each TALECD). We confirmed that nTALECD, ptpTALECD, or mtpTALECD expressed from these three expression vectors translocates to the nucleus, chloroplast, or mitochondria, respectively, and performs target single-nucleotide editing (conversion from a C:G pair to a T:A pair). We have found that using the plant genome editing method according to the present invention described above, targeted C:G pairs contained in the plant genome (nuclear genome, plastid genome, and mitochondrial genome) are homoplasmically modified. In other words, for example, taking the plastid genome as an example, it is possible to modify almost all of the target C:G pairs in the plastid genome, of which there are more than 1000 copies in the cells of the plant individual, into T:A pairs.
[0013] Incidentally, both plastids and mitochondria are organelles that developed as a result of endosymbiosis with free-living bacteria, and they each possess their own unique genomic DNA. However, compared to mitochondria, which have been endosymbiotic for a longer period, the plastid genome has a sequence and structure that is closer to that of bacteria. Furthermore, unlike the mitochondrial genome, the plastid genome has transcription, translation, and DNA replication / repair systems that clearly exhibit a bacterial type. Moreover, plant mitochondria partially reuse and duplicate some of the enzymes in the DNA replication / repair system used in plastids, and have their own unique hybrid system that differs from both the plastid genome and the mammalian mitochondrial genome. In other words, these three types of organelle genomes have three distinct patterns. In fact, among the molecules identified as repair factors for plastid genomic DNA and mammalian mitochondrial genomic DNA, there are many completely different repair molecules. Therefore, the genomic DNA repair and changes that occur when the mitochondrial and plastid genomic DNA are modified also differ (see Non-Patent Literature 8 and Non-Patent Literature 9, etc.). As described above, mammalian mitochondria and plant plastids and mitochondria are completely different intracellular organelles. Therefore, editing technologies applicable to the mammalian mitochondrial genome are not necessarily applicable to plant mitochondrial genome editing and plastid genome editing. Therefore, the result that "the targeted C:G pair is homoplasmically modified" is a remarkable effect that could not have been predicted from the result disclosed in Non-Patent Literature 6, which stated that "only about 42% of the targeted C:G pairs in mammalian cells were modified." Furthermore, regarding the editing techniques for plant mitochondrial genomes and plastid genomes disclosed in Non-Patent Literature 7, the single-nucleotide modification rates were approximately 25% and 38%, respectively. Based on these results, the plant genome editing method according to the present invention can be said to be an extremely efficient plant genome editing method compared to the method disclosed in Non-Patent Literature 7.
[0014] In other words, the present invention is as follows (1) to (6). (1) A method for editing plant genomic DNA, comprising modifying a target base on the genomic DNA with another base. The modification may be carried out by cytidine deaminase. (2) The method for editing plant genomic DNA may also be such that the cytidine deaminase is either of the proteins described in (a) or (b) below; (a) A protein consisting of the amino acid sequence represented by Sequence ID No. 35, (b) A protein having an amino acid sequence that has 90% or more sequence identity with the amino acid sequence represented by Sequence ID No. 35, and which has cytidine deaminase activity. (3) The method for editing plant genome DNA may involve fusing the N-terminal portion and the remaining portion of the cytidine deaminase with separate TALEs (transcription activator-like effectors). (4) The method for editing plant genomic DNA may include introducing the encoding DNA of a fusion of a part or all of the cytidine deaminase and TALE, to which a nuclear localization signal peptide, a plastid localization signal peptide, or a mitochondrial localization signal peptide has been added, into the nuclear genome of a plant cell (integrating it into the nuclear genomic DNA), and expressing the fusion with the added signal peptide in the plant cell, thereby modifying target bases in the nuclear genomic DNA, plastid genomic DNA, or mitochondrial genomic DNA of the plant to other bases. (5) A plant genome comprising plant genome DNA edited by the plant genome DNA editing method described above, a plant cell having the plant genome, a seed or plant containing the plant cell. (6) A method for producing a plant whose plant genome has been edited, comprising editing the plant genome using the plant genome DNA editing method described in any of (1) to (4) above. In this specification, the symbol "~" indicates a numerical range that includes the values to its left and right. [Effects of the Invention]
[0015] According to the present invention, it is possible to modify a single base in a plant genome, specifically in the nuclear genome, plastid genome, or mitochondrial genome of a plant. Furthermore, according to the present invention, it is possible to modify the target base in almost all copies of the nuclear genome, plastid genome, or mitochondrial genome within a plant individual. [Brief explanation of the drawing]
[0016] [Figure 1] Mechanism of action and expression vector of ptpTALECD, which targets the genes of plastids. Figure a schematically shows the target regions in the pTALECD and 16S rRNA genes. The 16S rRNA sequences in the figure are SEQ ID NOs. 39 and 40 from top to bottom. Figure b shows the T-DNA region of the ptpTALECD tandem expression vector. "1333C" is a protein consisting of amino acids from position 45 to 138 at the C-terminus of the DddAtox amino acid sequence represented by SEQ ID NOs. 35, and "1333N" is a protein consisting of amino acids from position 1 to 44 at the N-terminus of the DddAtox amino acid sequence represented by SEQ ID NOs. 35. "1397C" is a protein consisting of the C-terminal amino acid sequence of DddAtox, represented by SEQ ID NO: 35, from amino acid position 95 to 138, while "1397N" is a protein consisting of the N-terminal amino acid sequence of DddAtox, represented by SEQ ID NO: 35, from amino acid position 1 to 94.
[0017] [Figure 2] Schematic diagram of the ptpTALECD expression vector construction process. Figure a shows the assembly steps for constructing the pTALECD ORF. While the Platinum TALEN Kit was primarily used, step 2 of the entry vector was prepared using the procedure shown in Figure 8. Figure b shows the ptpTALECD expression vector construction process. The ptpTALECD expression vector was constructed using LR Clonase™ II Plus enzyme (Thermo Fisher Scientific).
[0018] [Figure 3] Substitution of the FokI coding sequence with the coding sequence of one half (referred to herein as "CD half") of a split cytidine deaminase (i.e., DddAtox). The FokI coding sequence and the coding sequences of the CD half (sequence numbers 7-10) inserted into the Step 2 entry vector used by Arimura et al., The Plant Journal 2020 104, 1459-1471 were amplified by PCR. The purified PCR amplification products were mixed with 5 x In-Fusion HD Cloning Enzyme Premix (TaKaRa) and incubated at 50°C for 15 minutes.
[0019] [Figure 4] Editing results of cytidine within the target region. ac shows the number of plant individuals with cytidine base substitutions, editing efficiency, and expected amino acid substitutions. The sequences shown in a are, from top to bottom, SEQ ID NOs. 41 and 42; in b, SEQ ID NOs. 43 and 44; and in c, SEQ ID NOs. 45 and 46. df shows representative analysis results for Sanger sequencing of ptpTALECD target sequences in T1 individuals 23 days after dormancy-awakening cold and humid treatment (hereinafter referred to as "23DAS"). The sequences shown in d are, from top to bottom, SEQ ID NOs. 47, 47, 48, 49, and 50; in e, SEQ ID NOs. 51, 52, 51, and 52; and in f, SEQ ID NOs. 53, 53, and 54. g shows the number of plant individuals grouped by target base substitution mutation type for T1 individuals in 11DAS and 23DAS. h / c (heteroplasmically or chimerically): heteroplasmic or chimeric substitution; homo: homoplasmic substitution; Cp: target cytosine where preferential substitution is expected; Cp*: cytosine where biological effects are expected.
[0020] [Figure 5]The following shows the results of analysis of leaves that have undergone chimeric base editing. Image a shows a leaf with partially different coloration of the 16S rRNA 1397NC (1397N-1397C) series 3 of 23DAS. Image b shows the results of genotype analysis of the ptpTALECD target region. The sequences shown in b are, from top to bottom, sequence number 55, sequence number 56, and sequence number 57.
[0021] [Figure 6] Analysis results for the T2 generation. The genotypes and phenotypes of six T2 individuals of the 16S rRNA1397CN lineage 2 are shown. The upper figure a shows the PCR amplification results of GFP and the target sequence 16S rRNA for three GFP-positive and three GFP-negative (i.e., individuals that inherited the T-DNA vector in the nucleus (positive) and individuals that did not inherit it (negative)) seeds, and the lower figure shows the genotype analysis results and phenotypes of the G5 single nucleotide substitution (SNP). b shows the representative phenotype of the T2 generation of the 16S rRNA1397CN lineage 2. The bars represent 1 mm. c and d show the phenotypes of the T2 generation of the 16S rRNA1397CN lineage 2 and 16S rRNA1397CN lineage 15 in the presence of Spm (spectinomycin). C shows images of T2 generation and wild-type seeds (0DAS) and seedlings (8DAS) of the two lineages on 1 / 2 MS medium containing 50 mg / L Spm (spectinomycin). D summarizes the relationship between the presence or absence of GFP fluorescence in seeds and the color of 8DAS individuals. W / G: Individuals with white or red cotyledons and green true leaves; ng: Did not germinate.
[0022] [Figure 7] Analysis results regarding the genotype and phenotype of T2 individuals. 'a' summarizes the genotype and phenotype of T2 individuals obtained by self-pollination of 16S rRNA1397CN lineages 2, 8, and 1397NC lineage 3. 'b' is an image of a representative phenotype of the T2 individuals shown in 'a'. The bar represents 0.5 mm.
[0023] [Figure 8]Construction of the 2nd entry vector and destination vector. Figure a shows the construction process for the 2nd entry vector. The 2nd entry vector (used in Arimura et al., The Plant Journal 104, 1459-1471 2020) and the RECA1 plastid-transfer peptide coding sequence were amplified by PCR. The purified PCR amplification product was mixed with 5 x In-Fusion HD Cloning Enzyme Premix (TaKaRa) and incubated at 50°C for 15 minutes. Figure b shows the construction process for the destination vector. The destination vector (used in Arimura et al., The Plant Journal 104, 1459-1471 2020) was amplified by PCR. The purified PCR amplification product was mixed with 5 x In-Fusion HD Cloning Enzyme Premix (TaKaRa) and incubated at 50°C for 15 minutes. The assembled destination vector was cleaved with KpnI, and the purified product was mixed with the OLE1GFP coding sequence amplified from 5 x In-Fusion HD Cloning Enzyme Premix (TaKaRa) and pFAST02 (INPLANTAINNOVATIONS INC). The mixture was incubated at 50°C for 15 minutes to construct the ptpTALECD expression vector.
[0024] [Figure 9] Genotypes of cotyledons in 13DAS of Spmr (spectinomycin-resistant) and Spms-like (spectinomycin-sensitive) individuals. The presence or absence of seed GFP fluorescence, the presence or absence of G5 SNPs, and the phenotypes of Spmr individuals (16S rRNA1397CN series 15 T2) and Spms-like individuals (16S rRNA1397CN series 2 T2) in 13DAS are shown in Figure 6c. W / G: White or red cotyledons and green true leaves
[0025] [Figure 10]Introduction of homoplasmic mutations to target bases in apt1. Figure a schematically shows a pair of pTALECD proteins, target bases, and target regions. See Figure 1 for the explanation of the CD splitting locations. The N-terminal half and C-terminal half of CD were fused to TALE, respectively. UGI (uracil glycosylase inhibitor): Uracil glycosylase inhibitor. The sequences shown in a are SEQ ID NOs. 58 and 59 from top to bottom. Figure b shows the number of plant individuals with cytidine base substitutions, editing efficiency, and expected amino acid substitutions in T1 individuals 11 days (11DAS) after dormancy-breaking cold and humid treatment. Cp: C at the T position of the 3' side chain, Cp*: Special target of opt87, No.: Total number of T1 individuals, h / c: Heteroplasmic and / or chimeric substitution, homo: Homoplasmic substitution. The sequences shown in b are SEQ ID NOs. 60 and 61 from top to bottom. c shows four representative examples of Sanger sequencing of PCR-amplified target sequences. The sequences shown in c are, from top to bottom, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, and SEQ ID NO: 65. d shows the number of plant individuals grouped by target base substitution mutation type for T1 individuals of 11DAS and 23DAS. The mutation stability rate (%) was calculated by dividing the number of bases that changed due to the mutation by the total number of bases that were substituted. An "unstable" mutation means that the type of mutation differs between individuals of 11DAS and 23DAS.
[0026] [Figure 11]Analysis results for T2 individuals. 'a' shows the genotypes of eight T2 generation individuals from atp1 1397NC 4. Seed-specific GFP expression derived from T-DNA was confirmed by fluorescence. The positive signal for mtpTALECD amplification indicates that the mtpTALECD gene introduced into the nuclear genome was inherited. atp1 is a positive control for PCR amplification of mtpTALECD. Sanger data for two bases in the target window (G4 and C10: locations where the parent plant has mutations) are shown in the figure below. NTC: no template control. 'b' shows the genotypes of the T2 generation of four series from 20DAS, Col-0 and otp87. Five nuclear mtpTALECD gene-free T2 generations (T2 no. 9-13, listed in Figures 16 and 17) from the 4T1 lineage (atp1 1333CN 3, 1333NC 7, 1397CN 24, and 1397NC 4) inherited mitochondrial homoplasmic mutations, grew to a similar extent to Col-0, and grew better than otp87. Bars represent 1 cm. c shows the results of on-target and off-target SNP analysis in the mitochondrial genomes of eight representative T2 individuals (two offspring from each of the four T1 lineages). None of these individuals contained the mtpTALECD gene. The X and Y axes show the location and frequency of mutated SNPs, respectively (different from the reference genome (BK010421.1) by ≥5%). Allele frequencies were calculated using AFmu-AFWT. AFmu is the allele frequency of the SNP for each mutation, and AFWT is the mean value of the same SNP in three wild-type individuals.
[0027] [Figure 12]Mitochondrial atp1 RNA repair in otp87 mutants using mtpTALECD. The left figure shows representative plant individuals in 13DAS for Col-0, the otp87 mutant, and otp87 modified with mtpTALECD. The right figure shows the DNA and RNA sequences near 393Leu in atp1. In the top figure, the C in the 393Leu codon is normally converted to T by RNA editing by OTP87. In the otp87 mutant (middle figure), this conversion does not occur, resulting in a substitution from Leu to Ser, which hinders plant growth. To restore the mutant's growth to normal, the C in atp1 was replaced with T using mtpTALECD (bottom figure). In this case, RNA editing by OTP87 was not necessary. This substitution restored the growth of the otp87 mutant to the same level as that of the wild type. Other experimental results are shown in Figures 21a and 21b. Bars represent 1 cm. The sequences shown in the figure, from top to bottom, are sequence number 66, sequence number 67, sequence number 66, sequence number 66, sequence number 67, and sequence number 67.
[0028] [Figure 13]The effect of mutations in the OTP87 predicted binding site within the atp1 sequence on OTP87 RNA editing. 'a' shows RNA sequence logos indicating the probability of occurrence of the bases to which each PPR motif of OTP87 binds, based on two key amino acids at positions 5 and 35 of each PPR motif of OTP87. The actual RNA sequence corresponding to the predicted binding site is located upstream of the OTP87 RNA editing site in atp1 (A shows this sequence (SEQ ID NO: 68)). PPR motifs are numbered from the C-terminal amino acid. The C-terminal S2 domain and N-terminal S domain correspond to the 4th base (-4A) and the 25th base upstream from the editing site (-25G), respectively. Target bases for mtpTALECD (see explanation in 'b') are enclosed in squares. 'b' shows the RNA sequence and RNA editing site of the expected OTP87 binding site in apt1 (see the top sequence). The -20G, -13G, and -6G in the sequence were each substituted with A by three pairs of mtpTALECD. The allele obtained by editing, the plant number of each allele, and the RNA editing from 1178C to U are also shown. TALE binding sequences are underlined. h / c (heteroplasmically or chimerically): heteroplasmic or chimeric substitution, homo: homoplasmic substitution. The sequences shown in b are, from top to bottom, SEQ ID NOs. 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, and 81. c shows representative examples of RNA (complementary DNA) sequences near the RNA editing site of the obtained allele. The bottom example in c shows data from the case where C was converted to T(U) at the highest level among the five (little) edited individuals (i.e., RNA editing hardly occurs in these individuals). Images of all the analyzed plant individuals and their genotypes are shown in Figures 22b and 22c, and Figure 23.
[0029] [Figure 14] A schematic diagram of the mtpTALECD tandem expression vector is shown. The primers used in Figure 11a are shown.
[0030] [Figure 15] Sanger sequencing results of amplicons amplified with primers that bind to both nuclear mitochondrial (NUMT) DNA sequences and mitochondrial DNA sequences (1). Representative examples of Sanger sequencing results of PCR amplification products amplified using primers that bind to both nuclear mitochondrial DNA sequences and mitochondrial DNA sequences (left) and primers that specifically bind to mitochondrial DNA (right) are shown. Data shown in the same position on both sides are results from the same plant individual. h / c (heteroplasmically or chimerically): heteroplasmic or chimeric substitution, homo: homoplasmic substitution. (That is, in these individuals, mitochondrial DNA is edited homoplasmically, and at the same time, homologous sequences exist in the nucleus but these sequences are not edited.) The sequences shown in the figure are, from top left, SEQ ID NOs. 82, 83, 84, 85, and from top right, SEQ ID NOs. 86, 87, 88, and 89.
[0031] [Figure 16] Sanger sequencing results of amplicons amplified with primers that bind to both nuclear mitochondrial (NUMT) DNA sequences and mitochondrial DNA sequences (2). Genotype lists for 11DAS and 23DAS are shown. *; DNA extracted from cotyledons. **; These base substitutions replace amino acids from G to N (when bases G3 and G4 are substituted with A), S (when only G3 is substituted with A), or D (when only G4 is substituted with A). ne; Not analyzed.
[0032] [Figure 17]Sanger sequencing results of amplicons amplified with primers that bind to both nuclear mitochondrial (NUMT) DNA sequences and mitochondrial DNA sequences (3). Genotype lists for 11DAS and 23DAS are shown. **; These base substitutions replace amino acids from G to N (when bases G3 and G4 are substituted with A), S (when only G3 is substituted with A), or D (when only G4 is substituted with A).
[0033] [Figure 18] Genotype of T2 individuals. The results of DNA sequencing of the target region of T2 individuals are shown. Mitochondrial genome-specific primers (NUMT primers do not amplify) were used for PCR. The far right column shows the Sanger sequencing results of the target region of a representative individual (number 9) from each of the 13 individuals in each lineage. Some bases that mutated homoplasmic and / or heteroplasmic in the T1 generation changed to a uniform genotype in the T2 generation. For example, in 1397CN 24, G4 was h / c in the 11DAS of the T1 generation, but reverted to the wild type in the T2 generation. The sequences shown in the far right column are, from top to bottom, sequence numbers 90, 91, 92, and 93. *The genotype of T1 is the same for both the 11DAS and 23DAS genotypes. **The genotype of individuals (numbers 9 to 13 in each lineage) is the genotype of 20DAS. h / c (heteroplasmically or chimerically): heteroplasmic or chimeric substitution; homo: homoplasmic substitution.
[0034] [Figure 19]Comparison of mitochondrial genome coverage analysis patterns of NGS short reads obtained from T2 individuals treated with mitoTALEN and mtpTALECD. Coverage data for T2 individuals treated with mitoTALEN was obtained from a previously reported study (Arimura et al., Plant J. 104 1459-1471 2020). The sequence information is the same as that shown in Figure 2c. The narrow gap common to all plant individuals, including Col-0, is an artifact resulting from the removal of reads homologous to sequences in the plastid genome. The white and black circles in the figure indicate the target sites of mtpTALECD and mitoTALEN, respectively.
[0035] [Figure 20] Amplicon sequences of the atp1-like NUMT sequence from T2 individuals. Individuals numbered 9-12 from each of the four series were selected as representative examples. The C corresponding to 1178C in atp1 is indicated by an arrow. The sequencing results showed that no significant substitutions occurred in sequences homologous to the target region. All sequences shown in the figure are sequence number 94.
[0036] [Figure 21] Growth status and genotype of T1 opt87 individuals transformed with apt1 1397CN. Figure a shows an image of the plant individual in 13DAS. Bars represent 1 cm. Figure b shows the genotype of the T1 individual shown in Figure a.
[0037] [Figure 22] Phenotypes and genotypes of all T1 individuals analyzed from those in which the predicted binding sequence of OTP87 was edited (1). a shows the predicted binding RNA sequence of OTP87 in apt1 and its RNA editing site. It shows amino acid sequence substitutions and RNA editing induced by C:G to T:A conversion using mtpTALECD. b shows the appearance of all plant individuals analyzed in 12DAS. c shows the genotype of the T1 individuals shown in b. Data is shown only for individuals in which mutations were confirmed out of 15 individuals.
[0038] [Figure 23] Phenotypes and genotypes (2) of all T1 individuals analyzed after editing the predicted binding sequence of OTP87. Representative examples of Sanger sequencing of mutant alleles and the presence or absence of 1178CRNA editing are shown. The sequences shown in the figure from top to bottom are SEQ ID NOs. 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, and 110.
[0039] [Figure 24] Editing of the CYO1 gene by nTALECD. a shows representative examples of the cyo1 mutant and wild-type phenotypes at true leaf emergence (11DAS). b-d show representative examples of phenotypes in 7DAS of the T1 generation with nTALECD introduced. e shows the cotyledon phenotype (7DAS) of the T1 generation with nTALECD introduced. f shows the number of individuals for each cotyledon phenotype in the T1 individual population and WT individual population of CYO1 ex1 (Example 1) and ex2 (Example 2). DAS: Days after stratification.
[0040] [Figure 25] Introduction of site-specific base substitutions into the target sequence in CYO1. The number of individuals with mutations in each base of the CYO1 ex1 / ex2 target sequence, as determined by PCR Sanger sequencing at the 21DAS timeframe, is shown. h / c indicates a heterozygous or chimeric individual of wild-type and mutant. In both ex1 and ex2, these mutations result in the formation of stop codons (ex1: CGA to TGA, ex2: TGG to TGA or TAG or TAA).
[0041] [Figure 26] Introduction of site-specific base substitutions into target sequences in PKT3 or MSH1. The number of individuals with mutations per base sequence, as determined by PCR Sanger sequencing at 21DAS time points, is shown for the target sequences of PKT3 and MSH1. h / c indicates a heterozygous or chimeric individual with both wild-type and mutant forms.
[0042] [Figure 27] Examination of the presence or absence of off-target editing near the target sequence. This section presents information on off-target mutations in the regions near 200 bp (a) and 1 kbp (b) of the target sequence, as examined by PCR Sanger sequencing at 35 DAS time points, along with the ratio of individuals in which mutations were detected compared to the individuals examined. [Modes for carrying out the invention]
[0043] The following describes embodiments for carrying out the present invention. The first embodiment is a method for editing plant genomic DNA, which includes modifying a target base on the genomic DNA with another base. In this embodiment, "plant genome" refers to the genome contained in the nucleus of a plant (nuclear genome), the genome contained in the plastids (plastid genome), or the genome contained in the mitochondria (mitochondrial genome). In this embodiment, "plastids" are organelles present in the cells of plants and algae, which perform assimilation such as photosynthesis, storage of sugars and fats, and synthesis of various compounds. Examples of plastids include chloroplasts, leucoplasts, and chromoplasts.
[0044] Modification of target bases may be carried out using base-modifying enzymes such as deaminase introduced into the nucleus, plastids, or mitochondria, although this is not particularly limited. Examples of such enzymes include cytidine deaminase, which modifies cytosine (C) in DNA to uridine (U). Particularly preferred are enzymes that modify C to U in double-stranded DNA, such as DddA of Burkholderia senosepacia. Burkholderia cenocepacia The cytidine deaminase domain of DddA (hereinafter referred to as DddA) tox For example: Sequence ID 35), or DddA tox It is essentially the same protein as DddA. toxA protein that is substantially identical to the above is, without any particular limitations, a protein that contains an amino acid sequence having 70% or more, preferably 80% or more, more preferably 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, and most preferably 99% or more amino acid identity with the amino acid sequence represented by Sequence ID No. 35, and that has cytidine deaminase activity (activity to change C to U in double-stranded DNA).
[0045] To specifically modify target bases in plant nuclear genomic DNA, plastid genomic DNA, or mitochondrial genomic DNA, it is necessary to enable a modifying enzyme, such as a deaminase (e.g., cytidine deaminase), to recognize the target bases. One method for this is to link a modifying enzyme to a transcription activator-like effector (TALE) that binds to genomic DNA near the target base (e.g., 0 to 1000 bases, preferably 5 to 100 bases, more preferably 5 to 50 bases) in the nuclear genomic DNA, plastid genomic DNA, or mitochondrial genomic DNA, and then introduce the modified enzyme-TALE fusion protein into the plant nucleus, plastid, or mitochondria. More specifically, for example, DNA encoding the modified enzyme-TALE fusion protein may be introduced into the nuclear genomic DNA (integrated into the nuclear genomic DNA), and the modified enzyme-TALE fusion protein expressed in the cytoplasm may be transported (introduced) into the nucleus, plastid, or mitochondria. In this case, it is desirable to introduce DNA encoding a fusion protein into the nuclear genomic DNA, which is a modified enzyme-TALE fusion protein to which various signal peptides (nuclear localization signal peptide, plastid localization signal peptide, or mitochondrial localization signal peptide) described later have been attached (bound).
[0046] One method for transporting modified enzyme-TALE fusion proteins into the nucleus is to fuse the modified enzyme-TALE fusion protein with a nuclear localization signal / sequence (NLS) peptide and express it. While not limited to these examples, nuclear localization signal peptides usable in embodiments of the present invention include, for example, the NLS peptide of SV40 large T antigen (PKKKRKV, SEQ ID NO: 111), the NLS peptide of nucleoplasmin (AVKRPAATKKAGQAKKKKLD, SEQ ID NO: 112), the NLS peptide of EGL-13 (MSRRRKANPTKLSENAKKLAKEVEN, SEQ ID NO: 113), the NLS peptide of c-Myc (PAAKRVKLD, SEQ ID NO: 114), and the NLS peptide of TUS protein (KLKIKRPVK, SEQ ID NO: 115). Other nuclear localization signal peptides are also available; for example, refer to the nuclear localization signal database NLSdb (https: / / rostlab.org / services / nlsdb / browse / signals).
[0047] One method for transporting the modified enzyme-TALE fusion protein into plastids is to fuse the modified enzyme-TALE fusion protein with a plastid-transfer signal peptide (a peptide that does not have clear higher-order structure or sequence homology, but for example, is rich in basic amino acids and several hydrophobic amino acids and has few acidic amino acids, and exhibits the function of being specifically selectively transported to chloroplasts or plastids when added to the N-terminus of the protein amino acid sequence) and express it. In the embodiments of the present invention, the plastid-transfer signal peptide that can be used is preferably a signal peptide possessed by a protein localized in a plant plastid. Preferred signal peptides, though not limited to them, include, for example, signal peptides derived from proteins such as RECA1, RBCS, CAB, NEP, SIG1-5, and GUN2-5, as well as signal peptides derived from nuclear-coded chloroplast ribosomal proteins such as RPL12 and RPS9, signal peptides derived from nuclear-coded chloroplast tRNA aminoacyltransferases, signal peptides derived from nuclear-coded chloroplast heat shock proteins, signal peptides derived from proteins such as FtsZ, FtsH, MinC, MinD, and MinE, signal peptides derived from the nuclear-coded chloroplast photosynthesis-related enzyme complex, signal peptides derived from nuclear-coded plastid lipid metabolism enzymes, and signal peptides derived from nuclear-coded thylakoid constituent proteins. For plastid-transfer signal peptides, see, for example, von HEIJNE et al., Eur. J. Biochem. 180, 535-545 1989.
[0048] One method for transporting the modified enzyme-TALE fusion protein into mitochondria is to fuse the modified enzyme-TALE fusion protein with a mitochondrial localization signal peptide (a peptide that does not have clear higher-order structure or sequence homology, but for example, one that exhibits the characteristic of alternating basic amino acids and multiple hydrophobic amino acids) and express it. In the embodiments of the present invention, the plastid localization signal peptide that can be used is preferably a signal peptide possessed by a protein localized in plant mitochondria. Preferred signal peptides are not limited to those listed below, but include, for example, the signal peptide derived from the ATPase δ' subunit of Arabidopsis thaliana (MFKQASRLLS RSVAAASSKS VTTRAFSTEL PSTLDS, SEQ ID NO: 116), the signal peptide derived from the ALDH2a gene product of rice (MAARRAASSL LSRGLIARPS AASSTGDSAI LGAGSARGFL PGSLHRFSAA PAAAATAAAT EEPIQPPVDV KYTKLLINGN FVDAASGKTF ATVDP, SEQ ID NO: 117), and the signal peptide derived from cytochrome c oxidase Vb-3 of pea (MWRRLFTSPH LKTLSSSSLS RPRSAVAGIR CVDLSRHVAT QSAASVKKRV EDVV, SEQ ID NO: 118), as well as the signal peptide derived from the ATPase β subunit of Arabidopsis thaliana and the signal peptide derived from chaperonin CPN-60 (Logan et al., Journal of Experimental Botany 50 Examples include the signal peptide of rice ALDH (865-871 2000), the signal peptide of rice F1F0-ATPase inhibitor protein (Nakazono et al., Plant Physiology 124 587-598 2000), and the signal peptide of rice F1F0-ATPase inhibitor protein (Nakazono et al., Plant 210 188-194 2000).
[0049] Alternatively, methods for directly introducing plasmid DNA, mRNA, and modified enzyme-TALE fusion proteins into cells (such as viral methods, particle gun methods, PEG methods, and cell membrane-permeable peptide methods) can also be used.
[0050] To modify target bases in plant genomic DNA with high probability, two modified enzyme-TALE fusion proteins (e.g., TALE left and TALE right shown in Figure 1, illustrating modification of the plastid genome) may be simultaneously expressed in a single Ti plasmid, and a tandem expression Ti plasmid may be used to localize it to the nucleus, plastid, or mitochondria by adding a nuclear localization signal peptide, plastid localization signal peptide, or mitochondrial localization signal peptide (see, for example, Non-Patent Literature 6). Additionally, DddA is used as a target base modifying enzyme. tox In cases where using the full-length protein directly would cause adverse effects on cells due to toxicity, the full-length protein may be cleaved at an appropriate position to produce partial proteins, which are then fused to the aforementioned TALE left and TALE right proteins, respectively, and each fused protein is transferred into the chromosome. The two partial proteins, separated at an appropriate position, reassemble when they bind near the target base and can exert the desired activity (see Examples). DddA is used as the modified enzyme. tox When using this, for example, DddA represented by sequence number 35 tox In the amino acid sequence, the split may occur between any of the amino acids in the sequence from the 40th to the 100th amino acid position, for example, between the 44th and 45th amino acids, or between the 94th and 95th amino acids. Furthermore, the modified enzyme-TALE fusion protein may be fused with other proteins that have a function to enhance the action of the fusion protein. Examples of such proteins include uracil glycosylase inhibitors (UGIs). UGIs inhibit the activity of uracil glycosylase, which removes U. Therefore, when cytidine deaminase is used as the modified enzyme, UGIs play a role in preventing the removal of U, which has been modified from C, and maintaining the modification by the cytidine deaminase-TALE fusion protein.
[0051] In the first embodiment, for example, DddA is the cytidine deaminase (CD) mentioned above. tox When used as a modifying enzyme, the target base C in nuclear genomic DNA, plastid genomic DNA, and mitochondrial genomic DNA can be homoplasmically modified to T (a state in which the mutation is identical throughout the cell, tissue, or organism). Therefore, the present invention provides a very effective means for improving plant organisms. The second embodiment is a nuclear genome in which target bases in the nuclear genomic DNA of a plant have been modified by the plant genomic DNA editing method according to the first embodiment, a plastid genome in which target bases in the plastid genomic DNA of a plant have been modified, or a mitochondrial genome in which target bases in the mitochondrial DNA of a plant have been modified, a nucleus having the nuclear genome, a plastid having the plastid genome, or a mitochondria having the mitochondrial genome, a plant cell having the nuclear genome, the plastid genome, or the mitochondrial genome, the cytoplasm of the plant cell, or a seed or plant (adult plant) containing the plant cell. In this embodiment, the plant (adult plant) includes not only the generation that differentiated into an adult plant (T0, or T1 depending on the plant) from transformed cells in which the target base in the nuclear genomic DNA, the target base in the plastid genomic DNA, or the target base in the mitochondrial genomic DNA was modified, but also the offspring generation obtained from T0 / T1. Furthermore, the seeds in the second embodiment include not only the seeds obtained from the T0 / T1 generation, but also the seeds obtained from the offspring generation.
[0052] The third embodiment is a method for producing a plant with an edited plant genome, which includes editing the plant genome using the plant genome DNA editing method according to the first embodiment. In other words, the third embodiment is a method for producing a plant with an edited nuclear genome, which includes editing the nuclear genome using the plant genome DNA editing method according to the first embodiment. A method for producing a plant with an edited plastid genome, comprising editing the plastid genome using the plant genome DNA editing method according to the first embodiment, or This is a method for producing a plant in which the mitochondrial genome has been edited, which includes editing the mitochondrial genome using the plant genomic DNA editing method according to the first embodiment.
[0053] The plants according to the first, second, and third embodiments are not particularly limited and may be any seed plants. Examples include grasses such as rice, wheat, maize, barley, rye, and sorghum, or plants of the Brassicaceae family such as *Capsella bursa-pastoris*, *Arabidopsis thaliana* (e.g., *Arabidopsis thaliana*), *Horsetail* (e.g., *Horsetail*), *Tricholoma rhodopolium*, and *Brassica napus* (e.g., *Tatsai*, *Brassica rapa*, *Rapeseed*, *Mizuna*, *Kale*, *Cabbage*, *Cauliflower*, *Cabbage*, *Brassica napus*, *Brussels sprouts*, *Broccoli*, *Chingensai*, *Rapeseed*, *Chinese cabbage*, *Komatsuna*, and *Turnip*). Plants belonging to the genera *Dracaena*, *Capsella*, *Cardamine*, *Cardamine*, *Dracaena*, *Dracaena*, *Dracaena* (including arugula), *Raphanus*, *Raphanus*, *Ionopsidium*, *Dracaena*, *Dracaena*, *Dracaena*, *Marcolmia*, *Mallotus*, *Nasturtium*, *Mallotus*, *Raphanus* (including radish and wild radish), *Rorippa indica*, *Rorippa indica*, *Arabis*, *Dracaena*, *Wasabi* (including wasabi), etc. can be used. Furthermore, examples include plants of the Solanaceae family such as tomatoes, potatoes, bell peppers, shishito peppers, and petunias; plants of the Asteraceae family such as sunflowers and dandelions; plants of the Convolvulaceae family such as morning glories and sweet potatoes; plants of the Araceae family such as konjac, taro, taro, and hoopoe; plants of the Fabaceae family such as soybeans, adzuki beans, and green beans; plants of the Cucurbitaceae family such as pumpkins, cucumbers, and melons; and plants of the Amaryllidaceae family such as onions, leeks, and garlic.
[0054] All disclosures of referenced documents in this specification are incorporated by reference as a whole. Furthermore, throughout this specification, where the singular words “a,” “an,” and “the” are included, they are considered plural as well as singular unless the context clearly indicates otherwise. The present invention will be further explained below with reference to examples, but these examples are merely illustrative of embodiments of the present invention and do not limit the scope of the present invention. [Examples]
[0055] I. Editing of plastid genomes I-1. Materials and Methods I-1-1. Plant materials and cultivation conditions Wild-type Arabidopsis thaliana Col-0 (Col-0) and genetically modified strains were cultivated at 22°C under long-day conditions (16 hours of light, 8 hours of darkness). Col-0 seeds were sown on 1 / 2 MS medium (pH=5.7) containing Murashige-Skoog medium salts (Wako, Japan) (2.3 g / L), MES (500 mg / L), and sucrose (10 g / L), as well as on 1 / 2 MS medium containing Plant Preservative Mixture (Plant Cell Technology, USA) (1 mL / L), Gamborg's Vitamin Solution (Sigma-Aldrich, USA) (1 mL / L), and agar (8 g / L). Seedlings were transplanted to Jiffy-7 (Jiffy Products International BV, Netherlands) 1-2 weeks after sowing and then used for Agrobacterium transfection. Some slow-growing T1 plants were transplanted to plant boxes containing 1 / 2 MS medium 23 days after stratification (DAS) (23DAS).
[0056] I-1-2. Design of TALE conjugate sequences The TALE target sequence was designed to bind to both sides of the cytidine deaminase target region using the Old TALEN Targeter (https: / / tale-nt.cac.cornell.edu / node / add / talen-old). The first base recognized should be as close as possible to the 3' side adjacent to T. The minimum length of the TALE target sequence was set to 15 bp for TALE to bind sequence-specifically. The binding sequence of TALE is shown below. 16S rRNA TALE left associative sequence: 5'-TAACCCAACACCTTACGGCACG-3' (SEQ ID NO: 1) TALE right associative sequence: 5'-CGGACACAGGTGGTGCAT-3' (SEQ ID NO: 2) rpoC1 TALE left associative linkage sequence: 5'-TGTTGATGTTTATACCGA-3' (SEQ ID NO: 3) TALE right associative sequence: 5'-TCGGAATGAATCACAAAAT-3' (SEQ ID NO: 4) psbA TALE left associative linkage sequence: 5'-TTTCGCGTCTCTCTAA-3' (SEQ ID NO: 5) TALE right associative sequence: 5'-TTAAATAAACCAAGGATTT-3' (SEQ ID NO: 6)
[0057] I-1-3. Construction of TALECD expression vector For each target, a pair of left and right ptpTALECDs (Figure 2) were incorporated into a Ti plasmid and constructed using the Platinum Gate assembling kit and multisite Gateway (Thermo Fisher), following the previously reported method for constructing mitoTALENs (Kazama et al., Nature plants 5, 722-730 2019). The DNA-binding domain of ptpTALECDs was assembled using the Platinum Gate TALEN system (Sakuma et al., Scientific reports 3, 1-8 2013.) (Figure 2a). The FokI coding sequence of mitoTALENs used in the reported assembly-step2 was replaced in advance with the coding sequences of CD half and UGI using the In-Fusion HD cloning Kit (TaKaRa, Japan, Figure 3). The coding sequences of CD half and UGI were designed to encode the same amino acid sequences as those disclosed in Non-Patent Document 3, and were synthesized by commissioning Eurofins Genomics (https: / / www.eurofinsgenomics.jp / jp / orderpages / gsy / gene-synthesis-multiple / ) using codons optimized for Arabidopsis thaliana. The assembled 1 st entry vector, 3 rd entry vectors and 2 nd The ORFs of the entry vectors were subjected to a multi-LR reaction using LR Clonase TM II Plus enzyme (Thermo Fisher Scientific) (Figure 2b) and were incorporated into the Ti plasmid (Arimura et al., The Plant Journal 104, 1459-1471 2020.). 2 nd The entry vectors have the terminator of the Arabidopsis thaliana heat shock protein (Nagaya et al., Plant and cell physiology 51, 328-332 2010.), the Arabidopsis thaliana RPS5A promoter and the N-terminal peptide (51 amino acids) of the plastid transit peptide (PTP) of Arabidopsis thaliana RECA1 (Figure 8a). This Ti plasmid has the CaMV 35S promoter of the Gateway destination Ti plasmid pK7WG2 (Karimi et al., Trends in plant science 7, 193-195 2002) replaced with the Arabidopsis thaliana RPS5A The sequence was reconstructed by substituting with the promoter (Tsutsui et al., Plant and Cell Physiology 58, 46-56 2017.) and inserting the PTP coding sequence and proOleosin::Ole1-GFP derived from pFAST02 (http: / / www.inplanta.jp / pfast.html, INPLANTA INNOVATIONS INC., Japan) (Figure 8b).
[0058] Below are the CD half-UGI sequences and RecA1 The PTP sequence is shown. G1333C+UGI array: (Sequence ID 7) "G1333C" is represented by sequence number 35, DddA toxThis protein consists of the amino acid sequence from position 45 to 138 at the C-terminus of the amino acid sequence. Furthermore, UGI (Uracil Glycosylase Inhibitor) consists of the amino acid sequence represented by SEQ ID NO: 36, and is linked to "G1333C" by a linker peptide (SEQ ID NO: 37) (hereinafter, the same applies to the UGI amino acid sequence and linker peptide).
[0059] G1333N+UGI array: GGATCTGGTAGCTATGCGTTAGGACCCTATCAGATTTCAGCTCCTCAATTGCCTGCCTATAATGGGCAAACTGTTGGCACCTTTTACTACGTCAATGATGCTGGAGGGTTAGAATCCAAGGTGTTCTCAAGTGGTGGTTCTGGAGGTAGTACGAATCTTTCGGACATCATAGAGAAGGAAACTGGAAAACAGCTCGTTATCCA AGAGAGCATTCTCATGTTGCCAGAAGAAGTTGAAGAGGTTATAGGCAACAAACCGGAATCTGACATTCTGGTACATACCGCTTATGATGAGTCAACAGATGAGAACGTCATGCTTTTGACATCTGATGCACCAGAATACAAACCTTGGGCACTTGTGATTCAGGATTCCAATGGTGAGAACAAGATCAAGATGCTA (SEQ ID NO: 8) "G1333N" is represented by sequence number 35, DddA tox It is a protein consisting of the amino acid sequence from the N-terminus, specifically from the 1st to the 44th amino acid.
[0060] G1397C+UGI array: GGTTCTGCGATTCCAGTTAAGAGAGGAGCTACAGGAGAAACGAAAGTCTTTACTGGGAATTCCAATTCTCCCAAATCACCGACTAAAGGCGGATGTAGTGGTGGTAGTACCAATCTTTCCGACATTATCGAGAAGGAAACAGGTAAACAACTCGTAATCCAAGAAAGCATACTGATGCTTCC TGAAGAGGTTGAAGAGGTCATAGGGAACAAACCTGAAAGCGACATTTTGGTTCATACTGCCTATGATGAGTCTACAGATGAGAACGTGATGTTGCTAACCTCAGATGCACCTGAATACAAGCCATGGGCTTTAGTGATTCAGGATTCGAATGGAGAGAACAAGATCAAGATGCTC (SEQ ID NO: 9) "G1397C" is represented by sequence number 35, DddA tox It is a protein consisting of the amino acid sequence from position 95 to 138 at the C-terminus of the amino acid sequence.
[0061] G1397N+UGI: (Sequence ID 10) "G1397N" is represented by sequence number 35, DddA tox It is a protein consisting of the amino acid sequence from the N-terminus, specifically from the 1st to the 94th amino acid.
[0062] RecA1 PTP code sequence: ATGGATTCACAGCTAGTCTTGTCTCTGAAGCTGAATCCAAGCTTCACTCCTCTTTCTCCTCTCTTCCCTTTCACTCCATGTTCTTCTTTTTCGCCGTCGCTCCGGTTTTCTTCTTGCTACTCCCGCCGCCTCTATTCTCCGGTTACCGTCTACGCCGCGAAG (SEQ ID NO: 11) "PTP" stands for Arabidopsis thaliana. RECA1It is a plastid-transfer peptide (its amino acid sequence is shown in SEQ ID NO: 38).
[0063] The primer sequences used for vector construction are shown in Table 1 below. [Table 1]
[0064] I-1-4. Plant transformation and screening of transformants Col-0 is conditioned to retain one of the above transformation vectors by the floral dipping method (Clough et al., The Plant Journal 16, 735-743 1998). Agrobacterium tumefaciens The plants were transformed using strain C58C1. First, transgenic T1 seeds were selected using fluorescence from GFP as an indicator. GFP-positive seeds were seeded on 1 / 2 MS medium containing 125 mg / L claforan. Then, GFP-negative seeds were seeded on 1 / 2 MS medium containing 50 mg / L kanamycin and 125 mg / L claforan.
[0065] I-1-5. Sanger sequencing and next-generation sequencing (NGS) Total DNA was extracted from the second true leaves of selected seedlings using the Maxwell® RSC Plant DNA Kit (Promega, USA). For genotyping of the transgenic strains, the plastid DNA sequence region near the cytidine deaminase target sequence was amplified using the primer set shown below, corresponding to the target gene. To detect target base substitutions, the purified PCR product was sequenced using the Sanger sequencing method. 16S rRNA Forward primer: 5'-GGTTCCAAACTCAACGGTGG-3' (SEQ ID NO: 27) Reverse primer: 5'-TAGGGGCAGAGGGAATTTCC-3' (SEQ ID NO: 28) psbA Forward primer: 5'-GGTATTATTTTAGTGGCCCA-3' (SEQ ID NO: 29) Reverse primer: 5'-GCCTGTGATAATAGGAAAGC-3' (SEQ ID NO: 30) rpoC Forward primer: 5'-AGACGGTTTTCAGTGCTAGT-3' (SEQ ID NO: 31) Reverse primer: 5'- TTTGGGGAGGGGTTTTTTAC-3' (SEQ ID NO: 32)
[0066] Using all DNA sequence data, single nucleotide polymorphisms (SNPs) in the plastid and mitochondrial genomes were identified. First, Macrogen Japan was commissioned to prepare a PE library using the Nextera XT DNA library Prep Kit (Illumina), and sequencing was performed using the Illumina NovaSeq 6000 platform. Sequence reads with 150 bp paired ends were analyzed using Geneious prime (Biomatters Ltd). Sequence reads were attached to the Arabidopsis thaliana chloroplast genome sequence, and sequences detected as SNPs with the reference chloroplast genome sequence in more than 50% of the reads are shown in Table 2. [Table 2]
[0067] I-1-6. Genotyping of T2 individuals T2 seeds obtained from T1 individuals corresponding to each target gene were sown on 1 / 2 M medium. The cotyledons of 7DAS or 13DAS seedlings were... 16S rRNA Genotyping was performed in the same manner as for T1 individuals. GFP PCR was performed using the primers shown below. Forward primer: 5'- GGTGATATCCCGCGGATGGTGAGCAAGGGCGAGGA-3' (SEQ ID NO: 33) Reverse primer: 5'- ACGTAACATGCCGGGCTTGTACAGCTCGTCCATGC-3' (SEQ ID NO: 34)
[0068] I-1-7. Screening for Spectinomycin-Resistant Individuals For 11DAS and 23DAS, 16S rRNA T2 seeds derived from T1 individuals in which C5 was replaced with homoprosmic were seeded in 1 / 2 MS medium containing 0, 10, or 50 mg / L spectinomycin. The phenotype of germinated cotyledons was observed in 8DAS.
[0069] I-1-8. Image Processing Plant images were taken with an iPhone® Xs (Apple Inc., US) and a LEICA MC 170 HD (Leica, Germany). Gel images were taken with ChemiDoc. TM The images were taken using MP Imaging System (BIORAD, USA). The images were also processed using Adobe Photoshop 2021 (Adobe, USA).
[0070] I-2. Results I-2-1. TALECD expression vector DddA represented by sequence number 35 tox In the amino acid sequence, between the 44th and 45th amino acids, or between the 94th and 95th amino acids, DddA tox The molecule was split, and either the N-terminus or C-terminus was ligated to the C-terminus of the platinum TALE DNA binding domain (Sakuma et al., Scientific reports 3, 1-8 2013) (pTALECD, Figure 1a). The N-terminus of pTALECD was ligated to the C-terminus of Arabidopsis thaliana. RECA1A protein plastid-targeting signal peptide (PTP) (Figure 1b) was ligated to the protein. Furthermore, a uracil glycosylase inhibitor (UGI) (Non-Patent Literature 3) was ligated to inhibit the hydrolysis of uracil (U) produced by cytidine deaminase (Figure 1b). DddA tox The nucleotide sequences of (CD) and UGI were optimized for codon usage frequency in Arabidopsis thaliana. The PTP-pTALECD-UGI (ptpTALECD) pair (a pair including the N-terminus and C-terminus of CD) RPS5A Expression was performed using a single-plant transformation vector under the promoter (Arimura et al., The Plant Journal 104, 1459-1471 2020) (Figure 1b). By modifying the method disclosed in a previous report (Kazama et al., Nature plants 5, 722-730 2019), an assembly system was established to easily construct tandem ptpTALECD expression vectors for each target sequence on a Ti plasmid (Figures 2a and b). In this example, the vector used in the previously disclosed method... FokI This was replaced with CD-UGI (Figure 3). The constructed vector was introduced into the nucleus of Arabidopsis thaliana using the floral dip method, and three regions of the plastid genome, namely, 16S rRNA Gene region (Figure 4a), rpoC1 Region (Figure 4b) and psbA In the region (Figure 4c), we attempted to replace C / G with T / A. In this manner, twelve different ptpTALECD expression vectors (expression vectors targeting three regions using combinations of four CD halves (see Figure 1a)) were constructed.
[0071] Each expression vector was introduced into Arabidopsis thaliana, and the target region of T1 was sequenced in 23DAS using the Sanger sequencing method. Only constructs from which T1 was obtained are shown in Figures 4a, b, and c. In multiple T1s, substitution of C / G pairs with T / A was confirmed in all target sequences of the three regions (Figures 4a-f). In addition to strains with heteroplasmic substitution or chimeric substitution (h / c; Figures 4a-f), surprisingly, a large number of strains with homoplasmic substitution of target bases (homo) were observed. Not all C / G pairs in the target region were substituted, and the substituted C / G pairs showed a bias in all three regions (Figures 4a-c). In the three regions, the homoplasmic substituted base was the C of (5')TC(3'), which Mok et al. (Non-Patent Literature 3) has indicated is more prone to mutation (Figure 4a-c). 16S rRNA The C in the (5')AC(3') of the gene was also homoplasmically substituted (Figure 4a).
[0072] To investigate the stability of mutations during individual growth, the base sequences of total DNA extracted from neonatal leaves of T1 in 11DAS and 23DAS (or cotyledons of slow-growing individuals in 11DAS) were examined. In 11DAS and 23DAS, among individuals in which base mutations occurred within the target region, some individuals retained the mutant base in a heteroplasmic or chimeric (h / c) state at both time points (30.0% of the total, 15 / 50, Figure 4g). In other individuals, the mutation status differed at both time points (for example, 4.0% of the total, 2 / 50, homo (homoplasmic mutation) became h / c; 14.0% of the total, 7 / 50, h / c became wild-type; 8.0% of the total, 4 / 50, h / c became homo; 2.0% of the total, 1 / 50, wild-type became h / c) (Figure 4g). The remaining majority of individuals retained the mutant base homoplasmically at both time points (42.0%, 21 / 50, Figure 4g). Interestingly, T1 individuals ( 16S rRNA In the cotyledons of 1397VC3), there are wild-type-like green areas and lighter-colored areas, and in each region 16S rRNAThe mutation rates in Cp* (cytosine, which is expected to cause biological effects) differed (Figures 5a and 5b). Surprisingly, most of the homoplasmic substituted bases in 11DAS remained homoplasmic in 23DAS (91.3%, 21 / 23). This result suggests that the target bases of T1 transformed with the ptpTALECD expression vector are frequently homoplasmically substituted, and that these mutations are stably maintained throughout the growth process.
[0073] Next, we investigated the off-target effects (substitution of non-target bases) of ptpTALECD in the maternally inherited plastid genome and mitochondrial genome (Table 2 above). The total genome sequences of 14 T1 individuals were determined (Novaseq, Illumina). In 13 individuals, most of the target C bases were homoplasmically substituted with T. 16S rRNA 1397C-1397N (1397CN) Series 2, Series 7, Series 8, Series 12, Series 16, 1397N-1397C (1397NC) Series 1, Series 2, Series 3: psbA 1397C-1397N(1397CN) Series 6, 1397N-1397C(1397NC) Series 1, Series 5: rpoC1 1397C-1397N (1397CN) Series 16) The remaining 1 target ( rpoC1 The 1397C-1397N (1397CN) series 3 (see Figures 4a-c) was heteroplasmically or chimerically substituted. Table 2 shows plastid SNPs in at least one T1 individual where more than 50% of the reads differed from the reference genome. Duplicate mutations in repetitive sequences of the plastid genome were counted as one mutation. Most of the target bases in 13 individuals were confirmed to be homoplasmically substituted. In the other individual, the base was confirmed to be heteroplasmically or chimerically substituted (Table 2). The main off-target point mutations (substitution frequency >50%) were: 16S rRNA Six off-target point mutations were found in the 1397C-1397N (1397CN) series 1, but no off-target point mutations were detected in other series (Table 2). 16S rRNAThe 1397CN series 1 died on 23DAS without developing true leaves. Regarding the mitochondrial genome, 16S rRNA No significant off-target mutations were detected in the mitochondrial genomes of any of the 14 individuals, including the 1397CN lineage 1. These results indicate that ptpTALECD rarely introduces off-target point mutations into organelle genomes and specifically and homoplasmically replaces C / G in target regions with T / A.
[0074] 16S rRNA Transform with the target ptpTALECD vector, and the first Cp*(G5) and / or C 10 In T1 individuals, one individual was replaced with homoplasmic ( 16S rRNA With the exception of the 1397C-1397N series 1), all were fertile. To investigate whether the C-to-T substitution mutation is inherited by offspring, these three strains ( 16S rRNA Genotyping was performed on T2 individuals of the 1397C-1397N series 2, series 8, and 1397N-1397C series 3) (Figures 6a and 7a). Based on the results of seed-specific GFP (green fluorescent protein) derived from Ole1 pro::Ole1-GFP13 on T-DNA (Figure 1b) and GFP PCR (Figure 6a), T2 individuals were classified into T-DNA transgene-free individuals (null isolates) and transgene-transfected individuals. All T2 individuals stably retained homoplasmic mutations (Figures 6a and 7a). Interestingly, the cotyledons of some T2 individuals were white, red, or variegated (Figures 6b and 7b), differing from the phenotype of their parent individuals. All such individuals were GFP-positive (Figures 6a and 7a), and many of them (8 out of 9 individuals) were 16S rRNA In the ~400 bp region examined within the sequence, other mutations were found (Figure 7a). Used for ptpTALECD expression. RPS5ASince the promoter has been reported to be significantly expressed in egg cells, it is thought that the de novo mutation occurred in the early stages of development of the T2 individual, resulting in the formation of abnormal cotyledons. Unlike these T2 individuals, the T2 individual that was a null segregator retained the target mutation without exhibiting the additional phenotypes described above. These results indicate that plastid genomes with artificially introduced point mutations are stably inherited by offspring, and moreover, this is independent of the inheritance of nuclear T-DNA. Furthermore, these results also indicate that null segregators with target point mutations in the plastid genome can be successfully established.
[0075] 16S rRNA The G5 gene is found in E. coli ( E. coli ) 16S rRNA This corresponds to G, which is expected to cause biological effects in this E. coli. 16S rRNA The G substitution mutation is a form of spectinomycin resistance (Spm r It is known that this confers ) to T1 individuals in which G5 is homoplasmically replaced with A ( 16S rRNA T2 seeds collected from the 1397C-1397N series 2) were sown in a spectinomycin-containing medium. Regardless of the presence or absence of GFP fluorescence from the seeds, most of the seedlings that germinated from these seeds showed resistance to spectinomycin (Figure 6c). However, 16S rRNA Some T2 individuals from the 1397C-1397N series 2 are susceptible to spectinomycin (Spm s They exhibited a phenotype similar to that of (a white, underdeveloped plant with purple cotyledons, Figure 6c). All of these spectinomycin-sensitive underdeveloped individuals germinated from GFP-positive seeds (Figure 6c), and many of them (5 out of 5 individuals, Figure 9) showed multiple de novo mutations. 16S rRNA It was present in the gene. This result indicates that the de novo mutation was 16S rRNA This causes dysfunction of the phenotype, resulting in a spectinomycin-sensitive phenotype (spectinomycin is 16S rRNA This suggests that it will be a drug that inhibits ( 16S rRNASome offspring of the 1397C-1397N series15) also showed spectinomycin resistance. These offspring (18 individuals) germinated from GFP-positive seeds (Figure 6c). In 5 of these offspring, G5 was homoplasmically replaced with A, and in the remaining 13 individuals, many G5s were replaced with A (Figure 9). This result suggests that the inherited T-DNA caused a de novo mutation in G5. These results suggest that a homoplasmic mutation of G5 to A confers spectinomycin resistance to Arabidopsis thaliana. Furthermore, the result that GFP-negative T2 individuals exhibit the spectinomycin resistance or susceptibility phenotype predicted from the G5 SNP in T1 individuals suggests that null-isolated T2 individuals are more likely to inherit mutations from their parents. The results above demonstrate that ptpTALECD can introduce target-region-specific and homoplasmic mutations that convert C to T in the plastid genome of Arabidopsis thaliana, and that these mutations are stably inherited by offspring (presumably following a maternal inheritance pattern).
[0076] II. Editing of the Mitochondrial Genome II-1. Materials and Methods II-1-1. Plant material, growth conditions, transformation, and screening of transformants Arabidopsis thaliana Col-0, otp87 (Homozygous T-DNA insertion line, GK-073C06-011724) and transformants were cultivated under long-day conditions at 22°C (16 hours light, 8 hours dark). Col-0 seeds were sown on 1 / 2 MS-Agar plates (Non-Patent Literature 7). Seedlings 2-3 weeks old were transferred to Jiffy-7 (Jiffy Products International) and then subjected to Agrobacterium infection. Col-0 and otp87Mature plants were transformed using the floral dip method (Clough et al., The Plant Journal 16, 735-743 1998). The resulting T1 seeds were selected by their seed-specific GFP fluorescence (Non-patent Literature 7; Shimada et al., Plant J. 61, 519-528 2010). These T1 seeds were seedled in the above medium containing 125 mg / L of claphoran. The T1 plants were transplanted into Jiffy-7 using 23DAS. otp87 Seeds (GABI_073C06) were obtained from the ABRC Stock Center. Homozygosity of OTP87 T-DNA insertion in plants was confirmed by PCR (Hammani et al., J. Biol. Chem. 286, 21361-21371 2011).
[0077] II-1-2. Design of TALE conjugate sequences and vector construction The TALE binding sequence is shown in Figures 10a and 13b. The base recognized by TALE is adjacent to the 3′ side of thymine and its length was set to approximately 20 bp. The target window length (16 bp) and the position of the special target cytosine (C10) were set based on successful examples disclosed in previous reports (Nakazato et al., Nature Plants 7 906-913 2021). The binary vector expressing mtpTALECD was constructed using the Platinum Gate TALEN system (Sakuma et al., Scientific Reports 3 1-8 2013) and a multisite Gateway (Thermo Fisher), in much the same manner as in previous reports (Nakazato et al., Nature Plants 7 906-913 2021). However, for the destination vector and entry vector used in the multi-LR reaction, those with mitochondrial localization signals instead of chloroplast localization signals were used.
[0078] II-1-3. Genotyping of T1 and T2 plant individuals PCR for Sanger sequencing (Figures 10, 11, 15, 16, 17, and 20) was performed using KOD One PCR Master Mix (Toyobo) with crude DNA extracted from true leaves or cotyledons, following the standard protocol. Nucleic acid templates for Sanger sequencing PCR (Figures 12, 13, 21, and 23) were extracted using the Maxwell RSC Plant RNA Kit (Promega) without using the included DNase I. The DNA in the extracted nucleic acids was degraded with Deoxyribonuclease (RT Grade) for Heat Stop (Nippon Gene) to prepare RNA templates for RT-PCR. RT-PCR was performed using PrimeScript. TM The procedure was performed using the II High Fidelity One Step RT-PCR Kit (TaKaRa). A portion of the mtpTALECD reading frame was amplified with primers to identify transformants. Sequences around the target window of mitochondrial DNA and cDNA, and their homologous sequences in nuclear DNA, were amplified. The purified PCR products were read by Sanger sequencing, and the data was analyzed using Geneious Prime (v. 2021.2.2).
[0079] Total DNA for NGS was extracted from mature leaves using the DNeasy Plant Pro Kit (QIAGEN). Paired-end libraries of 11 samples using the VAHTS Universal Pro DNA Library Prep Kit for Illumina (Vazyme, China) and 5Gbase / sample sequencing using the Illumina NovaSeq 6000 platform were performed at GENEWIZ Japan. Whole genome sequence data for SNP calling was obtained for 3 wild-type plant samples and 8 T2 plant samples (2 samples each from 4 strains). As a preprocessing step for analysis, low-quality sequences and adapter sequences in the reads were trimmed using PEAT [v1.2.4 (Li et al., BMC Bioinformatics, (BioMed Central, 2015), pp. 1-11.)]. Paired-end reads from each strain were mapped to reference sequences (mitochondrial genome BK010421.1 and chloroplast genome AP000423.1) in single-end mode using BWA (v 0.7.12) (Durbin, Bioinformatics 25 1754-1760 2009). Inappropriate mapped reads with sequence identity less than 97% or alignment coverage less than 80% were filtered out. SNPs were called using the samtools mpileup command (-uf -d 50000 -L 2000) and the bcftools call command (-m -A -P 0.1 (Li et al., Bioinformatics 25 207-2079 2009)). Using allele frequencies (AF) calculated with bcftools, SNPs with a final value of (AF of T1 sample) - (average AF of 3 wild-type individuals) ≥ 0.05 were detected as off-target SNP candidates, and many artifact SNPs originating from chloroplast genome sequences similar to NUMTs and sequences within the mitochondrial genome were removed (Figure 11c).
[0080] II-1-4. Prediction of PPR binding sequences atp1To predict the binding site of OTP87 in , we used the PPR code (Takanaka et al., PLos one 8 e65343 2013; Yan et al., Nucleic acids research 4 3728-3738 2019). This code was used to calculate which nucleotides each PPR repeat might recognize, based on the combination of two key amino acid residues at positions 5 and 35 of each PPR repeat. The binding probabilities for each motif are plotted in the weblog shown in Figure 13a (http: / / weblogo.berkeley.edu / ).
[0081] II-1-5. Image Processing The plant photos were taken with a digital camera (OLYMPUS OM-D E-M5) and processed with Adobe Photoshop 2021.
[0082] II-2.Results II-2-1. atp1 Target single nucleotide substitution Mitochondria as a target for base editing ATPase subunit 1 ( atp1 The base pair corresponding to the RNA editing site of ) atp1 -1178C was selected. In wild-type plants, this C is converted to U on post-transcriptional RNA and translated. Therefore, when evaluating the efficiency of single-nucleotide substitutions and their heritability, the substitution from C:G to T:A is not expected to have adverse effects on plants. In order to substitute this target base, Burkholderia cenocepaciaFour vectors containing the cytidine deaminase (CD) domain at the C-terminus of the DddA protein (1,427 amino acid, Non-Patent Literature 6) were constructed. Similar to previously reported methods (Non-Patent Literature 6; Non-Patent Literature 7; Nakazato et al., Nat. Plants 7, 906-913 2021; Lee et al., Nat. Commun. 12, 1-6 2021), the coding sequence of the CD domain was split at the nucleotide immediately following the codon of Gly 1333 or Gly 1397. Each of the split CD halves (N-terminus and C-terminus) was fused to the 3′ side of the DNA-binding domain sequence (hereinafter referred to as pTALE) of a platinum TALEN (Sakuma et al., Sci. Rep. 3 1-8 2013) that recognizes up to 21 bases. To prevent the removal of uracil generated from cytosine, the pTALE-CD sequence was fused to the 5′ side of the UGI sequence (Non-patent Literature 6; Mol et al., Cell 82, 701-708 1995, pTALE-CD-UGI). The nucleotide sequences of CD and UGI are the same as those previously reported (Nakazato et al., Nat. Plants 7, 906-913 2021) and have been optimized for codon usage in Arabidopsis thaliana. Arabidopsis thaliana ATPase delta prime subunit The mitochondrial target signal sequence (Arimura et al., Plant J. 104 1459-1471 2020) was ligated to the 5′ side of pTALE-CD-UGI (mtpTALECD, Figure 14). Cassettes expressing each of the two mtpTALECDs were constructed in tandem in a single binary vector. Each mtpTALECD has been used in highly efficient genome editing of Arabidopsis thaliana. RPS5AFour binary vectors were constructed under the control of a promoter (Figure 14) (Arimura et al., Plant J. 104 1459-1471 2020; Nakazato et al., Nat. Plants 7, 906-913 2021; Tsutsui et al., Plant Cell Physiol. 58 46-56 2017). These were named 1333C-1333N (abbreviated as 1333CN, meaning that the C-terminal half of the CD domain, separated by Gly 1333, is fused to the left TALE domain, and the N-terminal half to the right), 1333N-1333C (1333NC), 1397C-1397N (1397CN), and 1397N-1397C (1397NC) (Figure 10a).
[0083] To replace the target C:G pair in the mitochondrial genome with a T:A pair, each vector was used to transform the nuclear genome of Arabidopsis thaliana using the floral dip method (Clough et al., Plant J. 16 735-743 1998). The total DNA of the leaves of T1 transformants was amplified by PCR, and the PCR product sequence was determined by Sanger sequencing. Of the 78 T1 transformants examined (the number of transformants obtained from all four vectors), 36 individuals showed C:G to T:A substitution within the target window (Figures 16 and 17). Plant nuclear genomes often contain large sequence fragments with high homology to mitochondrial DNA, called nuclear mitochondrial DNA or NUMT (Noutsos et al., Genome Res. 15 616-628 2005; Zhang et al., Int. J. Mol. Sci. 21 707 2020). During the sequencing process, a part of the NUMT on chromosome 2 of Arabidopsis thaliana Col-0 was discovered. atp1 It was found that a nuclear sequence almost identical to that of NUMT (At2g07698) was amplified (Noutsos et al., Genome Res. 15 616-628 2005). Therefore, a new primer was designed to specifically amplify mitochondrial DNA so as not to amplify the NUMT sequence, and this was used in subsequent analyses.
[0084] For T1 plants in which mutations were detected during the initial genotyping, genotyping was performed again using new primers. Many transformants appeared to have homoplasmic substitutions of bases within the target window (Figures 10B and C). In addition to the mutation in the 10th target C, the 3rd, 4th, and 7th G in the target window were substituted in some T1 plants. Most of the substituted Cs were located on the 3' side of T or A, as previously reported (Figure 10b). Base substitution activity and preference for the position of the substituted base within the target window differed among the four vectors, with the 10th C being the most frequently homoplasmically substituted C within the target window when the vector was 1397C-1397N (1397CN, Figure 10b). As a result, five mitochondrial mutants were obtained in which only the true target base (10th C) was substituted in the target window at both 11 and 23 days after stratification (DAS) following the end of cold-humidification treatment to promote germination.
[0085] To determine whether the type of introduced mutation changes during plant development, the sequences of PCR fragments using total DNA from 11 DAS and 23 DAS of different leaves as templates were determined by Sanger sequencing for each transformant, and the type of mutation was identified. A total of 76 mutant bases were detected on at least one of these days (Figure 10d). Of these, 14 bases were substituted heteroplasmically or chimerically (h / c; i.e., not homoplasmic) on both days, and 25 bases were substituted in different ways on both days (see Figure 10D for the number and percentage of substituted bases of each type). The remaining 37 bases, accounting for about half of the detected mutant bases, were substituted homoplasmically on both days [48.7% (37 / 76), Figure 10d]. These results indicate that mtpTALECD efficiently substitutes C:G pairs within the target window with T:A, and that there are transformants in which homoplasmic mutations can be stably detected in leaves at two time points even in the T1 generation.
[0086] II-2-2. Inheritance of introduced mutations to seed progeny To confirm whether the introduced mutations are inherited by seed progeny, 13 T2 progeny were genotyped for each of the four T1 plants in which a C:G pair within the target window was homoplasmically substituted. All T2 individuals examined inherited the parental homoplasmic mutation, regardless of whether they had the mtpTALECD gene in their nucleus (Figures 11a and 18). This indicates that the homoplasmic mutations in the mitochondrial genome introduced by mtpTALECD were stably inherited by seed progeny. For each of the four lines, offspring without the mtpTALECD gene grew similarly to wild-type plants, even in plants with two different mutations causing amino acid substitutions [G391D and S392N (Figure 11b)]. Some of the bases that mutated heteroplasmically or chimerically in the T1 generation were observed to have a uniform genotype in the T2 generation (Figure 18).
[0087] II-2-3. Off-target mutations in the mitochondrial genome To investigate the off-target effects of mtpTALECD in the mitochondrial genome, SNP frequencies were measured in T2 plants (Figure 18) that had already been confirmed to have inherited parental homoplasmic mutations occurring within the target window. The locations and frequencies of lineage-specific mutant SNPs different from the reference sequence (BK010421.1) are shown as dots in Fig. 2C. These data indicated that the frequency of off-target mutations outside the target window was less than 10% of the mitochondrial DNA copies in each plant.
[0088] In these eight individuals, the overall mitochondrial genome coverage pattern was very similar to that of wild-type plants (Figure 19). Furthermore, there were no signs of structural changes in the mitochondrial genome, such as deletions, sequence rearrangements, or the generation of new repetitive sequences, as seen in previous studies using mitoTALEN (Kazama et al., Nat. Plants 5, 722-730 2019; Arimura et al., Plant J. 104 1459-1471 2020).
[0089] Approximately 20% of the reads located at SNPs within the target window did not contain mutant bases (Figure 11c). However, the mitochondria of these eight plants atp1 In the PCR product sequences, homoplasmic substitutions from C:G pairs to T:A pairs are observed (Figure 18), and the nuclear genome atp1 The PCR product sequence of the sequence (At2g07698) did not show any substitutions in the sequence corresponding to the target window (Figure 20). These results suggest that wild-type C:G SNPs detected in whole-genome sequencing are located in the nucleus. atp1 This supported the idea that the sequence originated from a similar sequence. Furthermore, this sequence is basically free of base substitutions (Figure 20), and low-frequency off-target mutations in the sequence (1397CN 24-10 and 12, Figure 20) can be removed by mating. In any case, no major off-target mutations were detected in either the mitochondrial genome (Figure 11c) or in nuclear DNA sequences similar to the target window (Figure 20).
[0090] II-2-4. Using mtpTALECD ppr Complementary phenotypes of mutants RNA editing is a characteristic feature of the mitochondrial and chloroplast genomes of land plants, where specific cytoplasm (C) in a post-transcribed RNA molecule is converted to u. This is mediated by nucleoencoded mitochondrial target PPR proteins (Small et al., Plant J. 101 1040-1056 2020). To verify the usefulness of mtpTALECD in molecular analysis of the mitochondrial genome, we performed two experiments related to RNA editing. First, we investigated the effects of growth retardation. otp87 We investigated the mutants. In wild-type plants, the PPR protein OTP87 atp1 The transcript 1178C (target window C10, Figure 10a) and nad7 The 27C in the transcript is converted to U (Hammani et al., J. Biol. Chem. 286 21361-21371 2011). Only the former RNA editing causes an amino acid substitution (S393L), so this does not occur. otp87It has been proposed that this is the cause of the growth delay. Therefore, mtpTALECD atp1 We investigated whether replacing 1178C with T at the DNA level would improve RNA editing deficiencies and, consequently, growth delays. We used 1397CN (Figure 10b), one of the mtpTALECD expression vectors. otp87 The mutant was introduced into the nuclear genome. Of the 14 T1 plants examined, 7 grew similarly to wild-type plants (Figure 12, Figure 21a). These 7 plants had homoplasmic substitutions from 1178C (C10) to T (or U) at the DNA and RNA levels in the true leaves (Figure 12, Figure 21a). These results indicate that atp1 The inability to edit the transcript 1178C is otp87 This indicates that it is the cause of the delayed growth of the mutant.
[0091] II-2-5. According to OTP87 atp1 Recognition In the second experiment, it is predicted that OTP87 will bind. atp1 The sequences were investigated (Takenaka et al., PloS One 8 e65343 2013, Figure 13a, Figure 22a). The nucleotides to which the PLS-type PPR protein OTP87 is predicted to bind, and their probabilities, are shown as nucleotide logos at the top of Figure 13a. These were predicted by combinations of two important amino acid residues at positions 5 and 35 of each PPR motif [e.g., P, L, S (Takenaka et al., PloS One 8 e65343 2013; Yan et al., Nucleic Acids Res. 47 3728-3738 2019; Barkan et al., PLoS Genet. 8, e1002910 2012; Yagi et al., PloS One 8, e57286 2013)]. The actual upstream RNA editing site to which OTP87 is predicted to bind. atp1The sequence is shown at the bottom of Figure 13a. In this experiment, to confirm whether this sequence is necessary for RNA editing, and if so, which bases are involved, several C:G pairs in this sequence were replaced with T:A pairs. Three mtpTALECD expression vectors were constructed in which three Gs at 20, 13, and 6 bases upstream of 1178C were replaced with As, respectively (denoted as -20G, -13G, and -6G, Figures 13a and 22a). Fifteen T1 seeds from each line (Col-0 background) were sown, and the DNA and RNA sequences of the seedlings were analyzed to confirm the pattern of DNA mutations induced by mtpTALECD and its impact on RNA editing efficiency at 1178C. In this study, the substitution of -13G was unsuccessful, but mitochondrial genome mutants with the following four allele patterns were obtained in the OTP87 binding prediction sequence. (i) Replace -24C with T, (ii) Replace -20G with A, (iii) Replace -24C and -20G with T and A respectively, (iv) Replace both -7G and -6G with A (Figure 13b). atp1 The RNA editing efficiency, expressed as Sanger sequencing data of the RT-PCR products of the transcript, decreased only in allele pattern (iv) (Figures 13b and c, 22a and c, and 23). These results indicate that at least 1-2 bases of the predicted binding sequence of OTP87 actually affect the RNA editing efficiency, with -7G and / or -6G affecting editing at 1178C, and possibly atp1 These are necessary for recognizing and binding the transcript, and the substitution of -24C and -20G with U and A respectively does not affect their activity in this case (or at least does not affect it significantly).
[0092] III. Editing of the Nuclear Genome III-1. Materials and Methods III-1-1. Plant material, growth conditions, transformation, and screening of transformants Arabidopsis thaliana Col-0 and transformants were cultivated under long-day conditions at 22°C (16 hours light, 8 hours dark). Col-0 seeds were sown on 1 / 2 MS-Agar plates (Non-Patent Literature 7). Seedlings 2-3 weeks old were transferred to Jiffy-7 (Jiffy Products International) and subsequently subjected to Agrobacterium infection. Mature Col-0 plants were transformed using the floral dip method (Clough et al., The Plant Journal 16, 735-743 1998). The resulting T1 generation was analyzed.
[0093] III-1-2. Design of TALE conjugate sequences and vector construction Based on the ptpTALECD construct (Nakazato et al., Nature Plants 7 906-913 2021), we created nTALECD by replacing the chloroplast localization signal (PTP) with the SV40 nuclear localization signal (SV40NLS). The target is AtCYO1 , AtPKT3 , AtMSH1 Target sequences were designed with the aim of introducing stop codons or amino acid substitutions that are presumed to have a significant impact on gene function at two locations for each of the three gene loci. A total of six nTALECD expression vector constructs corresponding to each target sequence were created, and these were transformed into Col-0 through infection with Agrobacterium using the floral dip method.
[0094] III-1-3. Genotyping of T1 plant individuals PCR for Sanger sequencing was performed using KOD One PCR Master Mix (Toyobo) with crude DNA extracted from true leaves or cotyledons, following the standard protocol. The nucleic acid template for Sanger sequencing was extracted using the Maxwell RSC Plant RNA Kit (Promega) without using the included DNase I. The DNA in the extracted nucleic acid was degraded with Deoxyribonuclease (RT Grade) for Heat Stop (Nippon Gene) to prepare the RNA template for RT-PCR. RT-PCR was performed using PrimeScript. TM The procedure was performed using the II High Fidelity One Step RT-PCR Kit (TaKaRa). A portion of the mtpTALECD reading frame was amplified with primers to identify transformants. Sequences around the target window of mitochondrial DNA and cDNA, and their homologous sequences in nuclear DNA, were amplified. The purified PCR products were read by Sanger sequencing, and the data was analyzed using Geneious Prime (v. 2021.2.2).
[0095] III-1-4. Image Processing The plant photos were taken with a digital camera (OLYMPUS OM-D E-M5) and processed with Adobe Photoshop 2021.
[0096] III-2.Results III-2-1. CYO1 Target single nucleotide substitution 11DAS cyo1 Representative examples of the cotyledon phenotypes of 7DAS in mutants, wild-type (Figure 24a), and T1 transformants with nTALECD introduced (Figures 24b-d) are shown in Figure 25. cyo1 The mutant exhibits a phenotype in which only the cotyledons are albino. cyo1 Since loss-of-function mutations are recessive, it was suggested that many T1 individuals have loss-of-function mutations introduced either entirely (Figure 24c) or partially (Figure 24d) in a biallelic or homozygous manner.
[0097] CYO1 The base sequence within the target sequence was sequenced using the Sanger method. As a result, it was confirmed that base substitutions occurred with high efficiency (>40%) for specific C molecules in the base sequence, and that biallelic / homozygous mutants could be easily obtained in the T1 generation (Figure 25).
[0098] Episode 2-2. PKT31 and MSH1 Target single nucleotide substitution next, CYO1 As a different target sequence PKT31 and MSH1 By selecting this option, the base sequences within the target window of both alleles were sequenced using the Sanger method. As a result, it was confirmed that the bases C10-C11 or G4-G6 had been edited (Figure 26). Therefore, CYO1 It was revealed that single-nucleotide editing is possible stably in target sequences other than those mentioned above, and that biallelic / homozygous variants of the targeted single-nucleotide edit can be easily obtained in the T1 generation for all of these sequences.
[0099] III-2-3. Off-target editing near the target window When performing a single-nucleotide substitution using the method of the present invention, we investigated the extent to which editing other than the target base, i.e., off-target editing, occurs. As a result, although off-target base substitutions occurred (all TC→TT), they were infrequent, and no indels (insertions and / or deletions of base sequences) were observed around the target sequence (Figure 27). [Industrial applicability]
[0100] The method according to the present invention enables single-nucleotide editing of plant genomes (nuclear genome, plastid genome, and mitochondrial genome). Therefore, plants modified using the method according to the present invention are expected to contribute to improvements in areas such as increased food production and biofuel production.
Claims
1. A method for editing plant genome DNA, comprising altering a target base on the genome DNA to another base, The modification is carried out by cytidine deaminase, and the cytidine deaminase is one of the proteins described in (a) or (b) below. (a) A protein consisting of the amino acid sequence represented by Sequence ID No. 35, (b) A protein having an amino acid sequence that has 90% or more sequence identity with the amino acid sequence represented by Sequence ID No. 35, and which has cytidine deaminase activity. The N-terminal portion and the rest of the cytidine deaminase are each fused with separate TALEs (transcription activator-like effectors). The method comprising introducing the coding DNA of each fusion, in which a nuclear localization signal peptide, a plastid localization signal peptide, or a mitochondrial localization signal peptide is attached to each of the fusions of the N-terminal portion of cytidine deaminase and TALE, and the portion other than the N-terminal portion and TALE, into the nuclear genome of a plant cell, and expressing each of the fusions with the attached signal peptide in the plant cell.
2. A method for producing an edited plant genome, comprising editing a plant genome using the plant genome DNA editing method described in claim 1.
3. A method for creating plant cells in which the plant genome has been edited, The method comprising editing a genome using the plant genome DNA editing method described in claim 1.
4. A method for producing seeds or plants in which the plant genome has been edited, The method comprising editing a plant genome using the plant genome DNA editing method described in claim 1.
Citation Information
Patent Citations
JP1538154119A
Method for transferring gene into plant pigment
JP2009225721A
US100,892-8972003
Method for modifying genome sequence to introduce specific mutation to targeted DNA sequence by base-removal reaction, and molecular complex used therein
WO2016072399A1
Method for converting monocot plant genome sequence in which nucleic acid base in targeted DNA sequence is specifically converted, and molecular complex used therein
WO2017090761A1