Methods and compositions for engineering condensed tannins in maize
By engineering maize with mutant ANS gene edits and heterologous ANR/TT2 homologs, the production of catechin-based proanthocyanidins and condensed tannins is enhanced, addressing the lack of efficient PA biosynthesis in maize and improving feed nutritional value and animal health.
Patent Information
- Application Number
- PCT/US2025/021151
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2025-03-24
- Publication Date
- 2025-10-02
AI Technical Summary
The biosynthetic routes to proanthocyanidins (PAs) and their building blocks in monocotyledonous crops like maize remain largely unexplored, with low levels or unusual flavan-3-ol anthocyanin conjugates detected, and existing enzymes like ZmANRl functionally distinct, limiting efficient production of long chain oligomeric PAs.
Engineering maize plants with mutant alleles or genomic edits that downregulate anthocyanidin synthase (ANS) gene function, combined with heterologous ANR and TT2 homologs, to enhance production of catechin-based proanthocyanidins and condensed tannins, utilizing site-specific genome modifying enzymes like RNA-guided nucleases.
Increased production of unconventional PA monomers and their conjugates, improving nutritional quality of feeds and animal health by enhancing nitrogen retention and reducing greenhouse gas emissions.
Smart Images

Figure 00000081_0000 
Figure 00000081_0001 
Figure 00000081_0002
Abstract
Description
TITLE OF THE INVENTIONMETHODS AND COMPOSITIONS FOR ENGINEERING CONDENSED TANNINS IN MAIZECROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application Ser. No. 63 / 570,485, filed March 27, 2024, which is incorporated herein by reference in its entirety.INCORPORATION OF SEQUENCE LISTING
[0002] A sequence listing containing the file named “UTNT002WO_ST26.xml” which is 30.7 kilobytes (measured in MS-Windows®) and created on March 16, 2025, and comprises 19 sequences, is incorporated herein by reference in its entirety.FIELD OF THE INVENTION
[0003] This present disclosure relates to the field of agriculture and plant genetics. In particular, the present disclosure relates to methods and compositions for producing condensed tannins in maize.BACKGROUND OF THE INVENTION
[0004] Proanthocyanidins (PAs), or condensed tannins, are the second most abundant plant phenolic compounds after lignin. They are oligomers and polymers of flavan-3-ols and are produced in several tissues of vascular plants, providing protection from herbivores, fungal pathogens, and ultraviolet radiation. The antibacterial activities of PAs and their precursor flavan-3-ols and their beneficial effects in preventing cardiovascular disease make these compounds popular health supplements and targets for increasing food nutritional value. Addition of PAs to ruminant animal feed can improve nitrogen retention in the rumen and reduce pasture bloat caused by production of methane gas, improving animal health and productivity and reducing greenhouse gas emissions from livestock. Compared to the well- characterized flavan-3-ol and PA biosynthesis pathways in model species such as Arabidopsis thaliana and Medicago truncatula, the biosynthetic routes to PAs and their building blocks in monocotyledonous crops remain largely unexplored.
[0005] As the most important cereal crop cultivated worldwide, maize (Zea mays) provides food, animal feed, and biofuels. Through domestication, breeders have generated a large collection of maize cultivars with distinct seed pigmentation, determined by the accumulation of different types and levels of flavonoids, including anthocyanins, quercetin, maysin, and phlobaphenes. In maize seeds, anthocyanins mainly accumulate in aleurone tissues, where their biosynthesis has been extensively studied and the enzymes involved identified and characterized. The transcriptional regulation of anthocyanin biosynthesis depends on a ternary complex of transcription factors, namely MYB-bHLH-WD40. Two anthocyanin-related MYB transcription factors, Cl and Pl, are involved in tissue-specific anthocyanin deposition in maize. Besides anthocyanins, phlobaphenes are detected in pericarp tissues of some maize varieties with brick-red seeds. The biosynthesis of phlobaphenes is regulated by an R2R3-type MYB transcription factor, pericarp color 1 (Pl). The key enzyme for phlobaphene biosynthesis, a bifunctional dihydroflavonol 4-reductase / flavanone 4-reductase, encoded by the Al locus, is induced by Pl and diverts the flux of substrates from anthocyanin biosynthesis toward the phlobaphene pathway. Homologs of Pl and Al have been identified in Sorghum bicolor, demonstrating a conserved regulatory mechanism for the phlobaphene pathway. The biosynthesis of anthocyanins, phlobaphenes and PAs shares common precursors and intermediates. However, few reports exist on the investigation of PAs and their precursors in maize, and, where available, suggest only very low levels or the existence of unusual flavan- 3-ol anthocyanin conjugates in some lines.
[0006] A key enzyme in the PA biosynthesis pathway, anthocyanidin reductase (ANR), converts anthocyanidins to the flavan-3-ol building blocks of PAs (i.e., catechins and epicatechins). Loss-of-function of ANR leads to reduced levels of PAs and increased anthocyanins in Arabidopsis thaliana seeds. Maize ANR (ZmANRl), although able to convert cyanidin to catechin and epicatechin in vitro, does so at a much lower rate than the reaction catalyzed by ANRs from the dicots soybean (Glycine max), Medicago truncatula, or A. thaliana, and the major product of ZmANRl is (+)-epicatechin rather than the typical (-)- epicatechin produced by other ANRs. When ZmANRl is ectopically expressed in the A. thaliana ANR mutant ban, the epicatechin level is slightly increased in developing seeds, but procyanidin B2 dimer is not detected, demonstrating that ZmANRl from maize is functionally distinct from the ANRs of A. thaliana and other PA-rich dicot plants.
[0007] Like those in the anthocyanin pathways, many PA-related enzymes are transcriptionally regulated by the ternary MBW complex consisting of MYB (TT2 or MYB5-type), bHLH(TT8) and WD40 transcription factors. The maize bHLH family transcription factors Lc and Sn were able to induce the production of anthocyanins and PAs when ectopically expressed in alfalfa (Medicago sativa) and lotus Nelumbo nucifera). In addition, ectopic expression of pale aleurone color 1 (Pacl), a maize WD40 family transcription factor, in the A. thaliana ttgll mutant restored the levels of both anthocyanins and PAs. Therefore, some enzymes and transcription factors for PA biosynthesis may exist in maize, but their in vivo functions remain unclear in light of the general lack of long chain oligomeric PAs in this species.
[0008] The present disclosure describes the occurrence and diversity of PAs and their precursors in different maize varieties and demonstrates the consequences of expressing heterologous ANR and TT2 homologs on production of PAs or their biosynthetic intermediates in these lines. The present disclosure further provides in planta evidence for a PA pathway in maize that generates unconventional monomers and their conjugates but does not support efficient polymerization to the long chain forms that improve the nutritional quality of feeds. In addition, the present disclosure provides a significant advance in the art by providing compositions and methods for producing and engineering condensed tannins in maize utilizing plants that have a mutant allele or a genome modification that results in loss of function of anthocyanidin synthase.SUMMARY OF THE INVENTION
[0009] In one aspect, the present disclosure provides a maize plant, seed, cell, or plant part comprising: a) a mutant allele or a genomic edit that downregulates anthocyanidin synthase (ans gene function; and b) a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19.
[0010] In one embodiment, the maize plant, seed, cell or plant part comprises a recombinant polynucleotide molecule comprising a nucleotide sequence encoding the polypeptide. In certain embodiments, the maize plant, seed, cell, or plant part comprises increased catechin based proanthocyanidins or increased condensed tannins. In one embodiment, the maize plant, seed, cell, or plant part comprises at least one truncated allele of an endogenous ans gene. In another embodiment, the truncated ans gene encodes a truncated polypeptide comprising a fragment of a sequence having at least about 85% sequence identity to SEQ ID NO: 12. In still yet another embodiment, the mutant allele or the genomic edit is present in at least one allele of an endogenous ans gene. The mutant allele or the genomic edit, in one embodiment, is inan endogenous ans gene encoding a protein having at least about 70% sequence identity to SEQ ID NO: 12. The genomic edit, in another embodiment, is in a transcribable region of the ans gene. The genomic edit, in yet another embodiment, is in an intron region of the ans gene. The genomic edit, in still yet another embodiment, is in an exon region of the ans gene. In one embodiment, the mutant allele or the genomic edit comprises a deletion, an insertion, a substitution, an inversion, a duplication, or any combination thereof. In another embodiment, the genomic edit is in an intron region and an exon region of the ans gene. In yet another embodiment, the maize plant, seed, cell, or plant part is heterozygous for the mutant allele or the genomic edit. In still yet another embodiment, the maize plant, seed, cell, or plant part is homozygous for the mutant allele or the genomic edit. The mutant allele or the genomic edit, in one embodiment, reduces or disrupts the activity of ANS protein compared to the activity of ANS protein in an otherwise identical maize plant, seed, cell, or plant part that lacks the mutant allele or the genomic edit.
[0011] In another aspect, the present disclosure provides a method for producing a modified maize plant cell comprising: a) introducing a genomic edit into at least one target site of an endogenous ans gene of a maize plant cell that downregulates ans gene function, wherein the maize plant cell comprises a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19; and b) selecting at least one plant cell comprising the genomic edit. The selecting, in one embodiment comprises detecting the genomic edit. The selecting, in another embodiment, comprises measuring anthocyanidin synthase (ans) gene function. In yet another embodiment, the maize plant cell comprises a recombinant polynucleotide molecule comprising a nucleotide sequence encoding the polypeptide. In one embodiment, the method further comprises regenerating at least one modified maize plant or plant part from the at least one selected plant cell or a descendant thereof comprising the genomic edit. In particular embodiments, a) the modified maize plant cell comprises at least one truncated allele of an endogenous ans gene; b) the genomic edit is present in at least one allele of an endogenous ans gene; c) the genomic edit is in an endogenous ans gene encoding a protein having at least about 70% sequence identity to SEQ ID NO: 12; d) the target site is located in a transcribable region of the ans gene; e) the target site is located in an intron region of the ans gene; f) the target site is located in an exon region of the ans gene; g) the target site is located in an intron region and an exon region of the ans gene; h) the modified maize plant cell is heterozygous for the genomic edit; i) the modified maize plant cell is homozygous for the genomic edit; or j) thegenomic edit comprises a deletion, an insertion, a substitution, an inversion, a duplication, or any combination thereof. In another embodiment, the method further comprises selecting at least one modified maize plant or plant part comprising increased catechin based proanthocyanidins or increased condensed tannins. In yet another embodiment, introducing the genomic edit comprises use of at least one site-specific genome modifying enzyme in the plant cell. The site-specific genome modifying enzyme, in particular embodiments: a) is selected from the group consisting of an RNA-guided nuclease, a zinc-finger nuclease, a meganuclease, a TALE-nuclease, a recombinase, a transposase, and combinations of any thereof; b) is an RNA-guided nuclease; or c) creates at least one strand break at the target site. Non-limiting examples of RNA-guided nucleases include a Cas nuclease, a Cpf 1 nuclease, or a variant of either thereof.
[0012] In yet another aspect, the present disclosure provides a method for producing a modified maize plant cell comprising: a) introducing a recombinant polynucleotide molecule comprising a nucleotide sequence encoding a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19 into a maize plant cell, wherein the maize plant cell comprises a mutant allele or a genomic edit in an endogenous ans gene that downregulates ans gene function; and b) selecting at least one plant cell comprising the recombinant polynucleotide molecule. The selecting, in one embodiment, comprises detecting the recombinant polynucleotide molecule. The selecting, in another embodiment, comprises detecting the expression of the polypeptide. In one embodiment, the method further comprises regenerating at least one modified maize plant or plant part from the at least one selected plant cell or a descendant thereof comprising the recombinant polynucleotide molecule. In another embodiment, the modified maize plant cell comprises at least one truncated allele of an endogenous ans gene. In yet another embodiment, the truncated ans gene encodes a truncated polypeptide comprising a fragment of a sequence having at least about 85% sequence identity to SEQ ID NO: 12. In particular embodiments, a) the mutant allele or the genomic edit is present in at least one allele of an endogenous ans gene; b) the mutant allele or the genomic edit is in an endogenous ans gene encoding a protein having at least about 70% sequence identity to SEQ ID NO: 12; c) the genomic edit is located in a transcribable region of the ans gene; d) the genomic edit is located in an intron region of the ans gene; e) the genomic edit is located in an exon region of the ans gene; f) the genomic edit is located in an intron region and an exon region of the ans gene; g) the modified maize plant cell is heterozygous for the mutantallele or the genomic edit; h) the modified maize plant cell is homozygous for the mutant allele or the genomic edit; or i) the mutant allele genomic edit comprises a deletion, an insertion, a substitution, an inversion, a duplication, or any combination thereof. The method, in still yet another embodiment, further comprises selecting at least one modified maize plant or plant part comprising increased catechin based proanthocyanidins or increased condensed tannins.
[0013] In still yet another aspect, the present disclosure provides a method for producing a modified maize plant cell comprising: a) introducing a genomic edit into at least one target site of an endogenous ans gene of a maize plant cell that downregulates ans gene function; b) introducing a recombinant polynucleotide molecule comprising a nucleotide sequence encoding a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19 into the maize plant cell; and b) selecting at least one plant cell comprising the genomic edit and the recombinant polynucleotide molecule. In one embodiment, the method further comprises regenerating at least one modified maize plant or plant part from the at least one selected plant cell or a descendant thereof comprising the genomic edit and the recombinant polynucleotide molecule.
[0014] In one aspect, the present disclosure provides a method for producing a hybrid maize plant, the method comprising crossing a first maize plant comprising a) a mutant allele or a genomic edit that downregulates anthocyanidin synthase (ans) gene function; and b) a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19 with a second, non-isogenic maize plant. In one embodiment, the present disclosure provides a hybrid maize plant produced by the methods described herein. In another embodiment, a hybrid maize plant of the present disclosure comprises increased catechin based proanthocyanidins or increased condensed tannins. In yet another aspect, the present disclosure provides a method for producing a maize plant, the method comprising crossing a first maize plant comprising a) a mutant allele or a genomic edit that downregulates anthocyanidin synthase (ans gene function; and b) a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19 with itself or a second maize plant. In some embodiments, the first maize plant comprises increased catechin based proanthocyanidins or increased condensed tannins.
[0015] In still yet another aspect, the present disclosure provides a method for producing a maize progeny plant or seed comprising increased catechin based proanthocyanidins or increased condensed tannins, the method comprising: a) crossing a first maize plant with a second maize plant, wherein the first maize plant comprises a mutant allele or a genomic edit that downregulates anthocyanidin synthase ans) gene function, and the second maize plant comprises a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19; and b) selecting a maize progeny plant or seed comprising the mutant allele or the genomic edit and the polypeptide. In one embodiment, selecting the maize progeny plant or seed comprises detecting the mutant allele or the genomic edit. In another embodiment, selecting the maize progeny plant or seed comprises measuring anthocyanidin synthase ans) gene function. Selecting the maize progeny plant or seed, in yet another embodiment, comprises detecting the polypeptide. Selecting the maize progeny plant or seed, in still yet another embodiment, comprises detecting the increased catechin based proanthocyanidins or increased condensed tannins in a sample derived from the progeny plant or seed. In one embodiment, the second maize plant comprises a recombinant polynucleotide molecule comprising a nucleotide sequence encoding the polypeptide. Selecting the progeny plant or seed, in another embodiment, comprises detecting the recombinant polynucleotide molecule. In yet another embodiment, a method of the present disclosure may comprise a) selfing the maize progeny plant or crossing the maize progeny plant with the first maize plant or the second maize plant; and b) selecting a further maize progeny plant or seed comprising the mutant allele or the genomic edit and the polypeptide. The present disclosure further provides, in still yet another embodiment, a maize progeny plant or seed produced by a method of the present disclosure. In certain embodiments, the progeny maize plant or seed comprises the mutant allele or the genomic edit and the polypeptide. In still yet another embodiment, a method of the present disclosure may comprise selfing or crossing a maize plant or a maize progeny plant of the present disclosure, wherein the crossing comprises crossing with any other maize plant.BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present invention. The invention may be betterunderstood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[0017] Ectopic expression of ZmANRl fails to generate PA precursors in high anthocyanin producing tobacco. FIG. 1A Transcript levels of ZmANRl and ZmANR2 in developing seeds (14-days after pollination) of maize cultivars ST (Suntava) and BM (Black Mexican). Data are presented as mean ± S.D. (n = 3, independent biological replicates). Differences in transcript levels of ZmANRl and ZmANR2 between ST and BM maize are significant (P < 0.01) as determined by two-tailed Student’s Mest. ZmEFla was used as the reference gene. FIG. IB Confocal microscopy images of GFP-tagged GmANRl and ZmANRl showing their subcellular localization in A. thaliana protoplasts. Chloroplasts appear autofluorescent but are distinguishable from GFP-tagged GmANRl and ZmANRl . Images are representative of three independent replicates. FIG. 1C Images of tobacco flowers expressing PAP1, PAP1 with AtANR, PAP1 with GmANRl, and PAP1 with ZmANRl before (top) and after (bottom) DMACA staining. The distribution of PAs in tobacco flowers is indicated by DMACA staining. AtANR and GmANRl were cloned from Arabidopsis thaliana (Col-0) and Glycine max (cv Clark), as described (Jun et al. Sci. Adv. 7, eabg4682, 2021). FIG. ID Selected ion chromatogram of catechin and epicatechin (left, m / z = 289.0718 ± 10 ppm), 4 / ?-(S-cysteinyl)- catechin and 4 / ?-(S-cysteinyl)-epicatechin (middle, m / z = 408.0759 ± 10 ppm), procyanidin dimers Bl and B2 (right, m / z = 577.1360 ± 10 ppm) in PA extracts from tobacco flowers. FIG. IE Transcript levels of AtANR, GmANRl and ZmANRl in tobacco plants expressing PAP1, PAP1 with AtANR, PAP1 with GmANRl, and PAP1 with ZmANRl. Data are presented as mean ± S.D. (n = 3, independent biological replicates). Asterisks indicate significant difference relative to the untransformed control at P < 0.01 as determined by Student’s Mest. Transcripts not detected are labeled as n.d. NtEFla was used as the reference gene.
[0018] FIG. 2 Analysis of PA starter and extension units and anthocyanins in maize seeds. Left panels, selected ion chromatograms of catechin and epicatechin (m / z = 289.0718 ± 10 ppm). Middle panels, selected ion chromatograms of 4 / ?-(S-cysteinyl)-catechin and 4 ?-(S- cysteinyl)-epicatechin (m / z = 408.0759 ± 10 ppm). Right panels, selected ion chromatograms of cyanidin 3-O-glucoside (m / z = 447.0940 ± 10 ppm). SD, chemical standards; c, catechin; epi, epicatechin; c-cys, cysteinyl-catechin; epi-cys, cysteinyl-epicatechin; C3G, cyanidin 3-O- glucoside. Maize varieties are FBLL; AREQ, Arequipa; ST, Suntava; OSA, Osage; BM, Black Mexican.
[0019] Formation of PA precursors in different maize varieties expressing GmANRl . Selected ion chromatograms of catechin and epicatechin (m / z = 289.0718 ± 10 ppm), as well as 4 ?-(S- cysteinyl)-catechin and 4 / ?-(S-cysteinyl)-epicatechin (m / z = 408.0759 ± 10 ppm) in seeds of untransformed and transgenic ST (FIG. 3A) and OSA (FIG. 3C) maize. Peak areas of 4 ?-(S- cysteinyl)-epicatechin and ratio of 4 / ?-(S-cysteinyl)-epicatechin / 4 / ?-(S-cysteinyl)-catechin in ST, ST-GmANRl, OSA and OSA-GmANRl are shown in histograms. Data are presented as mean ± S.D. (n = 3, independent biological replicates). Asterisks denote significant difference relative to wild-type ST or OSA at P < 0.01 determined by two-tailed Student’s Z-test. FIG. 3B Phloroglucinolysis of PAs extracted from seeds of ST, ST-GmANRl and Hill-GmANRl using procyanidin dimer B2 as standard. Left, selected ion chromatograms of epicatechin monomers (m / z = 289.0718 ± 10 ppm); right, selected ion chromatograms of epicatechin-phloroglucinol (m / z = 413.0876 ± 10 ppm). FIG. 3D MS / MS spectra of 4 / ?-(S-cysteinyl)-epicatechin in the standard and PAs extracted from seeds expressing GmANRl. SD, chemical standards; ST, Suntava maize; OSA, Osage maize; Hill-GmANRl, Hill maize expressing GmANRl (line GmANRl-24); ST-GmANRl, seeds from genetic cross between ST and GmANRl-24; OSA- GmANRl, seeds from genetic cross between OSA and GmANRl-24. Arrows indicate peak representing cysteinyl-epicatechin.
[0020] Expression of SbTT2 or SbMYB5 reduces growth and enhances anthocyanin accumulation in yellow- seeded maize. FIG. 4A Top, images of seeds from untransformed Bl 04, transgenic lines expressing SbTT2 (SbTT2-21 and SbTT2-31), and transgenic lines expressing SbMYB5 (SbMYB5-12 and SbMYB5-51); middle, zoom-in images of seeds in the top row showing differences in seed size and color; bottom, images of maize coleoptile at 5 days after seed germination showing differences in color. FIG. 4B Seed weight of untransformed B104 and transgenic lines expressing SbTT2 or SbMYB5. Data are presented as mean ± S.D. (n = 6, independent biological replicates). Asterisks denote significant difference relative to untransformed B104 at P < 0.01 determined by two-tailed Student’s t- test. FIG. 4C Images of maize silk showing differences in color.
[0021] PA and PA precursor content and composition in seeds of Bl 04 maize expressing SbTT2 or SbMYB5. FIG. 5A Transcript levels of SbTT2, SbMYB5, ZmANRl and ZmANR2 in developing seeds (14-day after pollination) of untransformed Bl 04, SbTT2-OX and SbMYB5- OX analyzed by qRT-PCR. Data are presented as mean ± S.D. (n = 3, independent biological replicates). ZmEFla was used as reference gene. Asterisks indicate significant difference relative to the untransformed control at P < 0.01 as determined by two-tailed Student’s Z-test.FIG. 5B Contents of soluble and insoluble PAs in seeds of B 104, SbTT2-OX and SbMYB5- OX. Data are presented as mean ± S.D. (n = 3, independent biological replicates). Asterisks indicate significant difference relative to the untransformed control at P < 0.01 as determined by two-tailed Student’s Mest. FIG. 5C Selected ion chromatograms of catechin and epicatechin (m / z = 289.0718 ± 10 ppm), as well as 4 / ?-(S-cysteinyl)-catechin and 4 ?-(S- cysteinyl)-epicatechin (m / z = 408.0759 ± 10 ppm) in seeds of B 104, SbTT2-OX (line SbTT2- 31) and SbMYB5-OX (line SbMYB5-51). Peak areas of epicatechin (epi), 4 / ?-(S-cysteinyl)- catechin (c-cys), and 4 / ?-(S-cysteinyl)-epicatechin (epi-cys) in B 104 and SbTT2-OX are shown in the histograms. Data are presented as mean ± S.D (n = 3, independent biological replicates). Asterisks denote significant difference relative to untransformed B 104 at P < 0.01 determined by two-tailed Student’s Mest. Compounds not detected are labeled as n.d. (D) Phloroglucinolysis of PAs extracted from seeds of B 104, SbTT2-OX and SbMYB5-OX using procyanidin dimers B2 and B3 as references. Left, selected ion chromatograms of catechin and epicatechin monomers (m / z = 289.0718 ± 10 ppm); right, selected ion chromatograms for detection of catechin-phloroglucinol and epicatechin-phloroglucinol (m / z = 413.0876 ± 10 PPm).
[0022] Analysis of procyanidin dimers and trimers in maize seeds expressing GmANRl or SbTT2. Selected ion chromatograms of procyanidin dimers in standards and PAs extracted from seeds of untransformed and transgenic maize ST (FIG. 6A), OSA (FIG. 6B), and Bl 04 (FIG. 6C). FIG. 6D Selected ion chromatograms of procyanidin trimers in PAs extracted from seeds of Medicago (Mt), soybean (Gm), untransformed ST maize, and ST maize expressing GmANRl. SD, chemical standards; ST, Suntava maize; ST-GmANRl, transgenic maize seeds obtained from crosses between Hill-GmANRl and ST; OSA, Osage maize; OSA-GmANRl, transgenic maize seeds obtained from crosses between Hill-GmANRl and OSA; B1-B4, procyanidin dimers B1-B4 (m / z = 577.1360 ± 10 ppm); iso-B2, procyanidin B2 isomer with (+)-epicatechin as the starter unit; iso-B4, procyanidin B4 isomer with (+)-epicatechin as the starter unit; Cl, procyanidin Cl trimer (m / z = 865.1981 ± 10 ppm).
[0023] Analysis of anthocyanins and PA precursors in seeds of BZ1 and bzl mutant maize. FIG. 7A Selected ion chromatograms of catechin and epicatechin (left, m / z = 289.0718 ± 10 ppm), 4 / ?-(S-cysteinyl)-catechin and 4 / ?-(S-cysteinyl)-epicatechin (middle, m / z = 408.0759 ± 10 ppm), as well as cyanidin 3-O-glucoside (right, m / z = 447.0940 ± 10 ppm) in seeds of BZ1 and bzl maize. FIG. 7B Anthocyanin contents in BZ1 and bzl seeds. Data are presented asmean ± S.D. (n = 3, independent biological replicates). Asterisks indicate significant difference between BZ1 and bzl samples at P < 0.01 as determined by two-tailed Student’s Mest. FIG. 7C Chiral-HPLC analysis of the stereochemistry of epicatechin monomers in bzl seeds. Arrow indicates the (+)-epicatechin in PAs extracted from bzl seeds. FIG. 7D Selected ion chromatograms of procyanidin dimers (m / z = 577.1360 ± 10 ppm) in BZl and bzl maize seeds. SD, chemical standard; (-)-epi and (+)-epi, (-)-epicatechin and (+)-epicatechin; c-cys and epi- cys, 4 / ?-(S-cysteinyl)-catechin and 4 / ?-(S-cysteinyl)-epicatechin; C3G, cyanidin 3-0- glucoside; B3, procyanidin B3 dimer; iso-B2, procyanidin B2 isomer with (+)-epicatechin as the starter unit; iso-B4, procyanidin B4 isomer with (+)-epicatechin as the starter unit.
[0024] FIG. 8 Schematic diagram of PA, anthocyanin and phlobaphene biosynthesis pathways in maize seeds. CHS, chaicone synthase, c2 CHI, chaicone isomerase, chil; F3H, flavanone 3-hydroxylase, f3h F3’H, flavanone 3 '-hydroxylase, prl; DFR, dihydroflavanol 4-reductase, al; ANS, anthocyanidin synthase, a2; ANR, anthocyanidin reductase; UFGT, UDP-glucose: flavonoid glucosyltransferase, BZ1; GST, glutathione S-transferase, BZ2. Anthocyanins accumulate in aleurone cells, and phlobaphenes in the pericarp.
[0025] FIG. 9 shows LC-MS analysis of (epi)catechin (c and epi, left), (epi)catechin-cysteine (c-cys and epi-cys, middle), and cyanidin 3-O-glucoside (C3G, right) in wild-type maize (M142A) and in maize seeds with mutation in ANS / A2 (M142C), UFGT / Bzl (M142D) or GST / Bz2 (M142E).
[0026] FIG. 10 shows phloroglucinolysis of PAs extracted from maize seeds of M142A (WT), M142C ans!a2), M142D ufgt / bzl), and M142E (gst / bz ). Catechin and epicatechin (c and epi, left) are PA starter units, while catechin-phloroglucinol and epicatechin-phloroglucinol (c- phloro and epi-phloro, right) are PA extension units.
[0027] FIG. 11 shows LC-MS analysis of PA dimers in M142A (WT) and selected mutant seeds. M142C (ans / a2) has procyanidin B3 dimer.
[0028] FIG. 12 shows LC-MS analysis of (epi)catechin (left), (epi)catechin-cysteine (middle), and PA B3 dimer (right) in WT and antl9 (lar mutant) of barley seeds.
[0029] FIG. 13A and FIG. 13B show LC-MS analysis of PAs extracted from Ml 42V (ANS / A2, Rl-nj) and M541G (ans / a2, Rl-nj) maize seeds. FIG. 13A shows detection of (epi)catechin (left) and (epi)catechin-cysteine (right) from M142V and M541G seeds; FIG. 13B shows detection of (epi)catechin (left, PA starter units) and (epi)catechin-phloroglucinol(right, PA extension units) after phloroglucinolysis using PAs extracted from Ml 42 and M541G seeds.
[0030] FIG. 14 shows genotyping of maize plants with the A2 or a2 genotypes. Top, phenotype of Ml 42V (Rl-nj; A2) and M541G (Rl-nj; a2) seeds. Bottom, genotyping of Ml 42V and M541G plants using maize EFla as reference gene. The arrow indicates the anticipated PCR product amplifying the coding sequence of the A2 gene and this band is missing in M541G.
[0031] FIG. 15 shows a schematic of a simplified biosynthetic pathway of PA, anthocyanin, and phlobaphenes in colored maize seeds.
[0032] FIG. 16 shows the phenotype of maize seedlings before and after DMACA staining. Left, seedling from M142V 7-days after germination before and after DMACA staining; right, seedling from M541G 7-days after germination before and after DMACA staining. Arrows highlight maize coleoptile and root tissues.
[0033] FIG. 17 shows the phenotype of maize seedlings after DMACA staining. A purple color is present from line M541G on the right, indicating that PAs accumulate in maize vegetative tissues, but are not present in M142V (left).
[0034] FIG. 18 shows maize seeds 21-, 28-, 35-days after pollination, as well as mature dry seeds, cut in half and stained using DMACA reagent. A purple color is present from line M541G in the lower panels, particularly prevalent in dry seeds, indicating that PAs accumulate in this line in the embryo (due to the site of expression of the TT8 allele in this line). The PAs are not present in dry seeds of the M142V line.
[0035] FIG. 19 shows phloroglucinolysis of PAs extracted from Ml 42V (A2) and M541G (a2) leaf and root tissues using procyanidin B2 and B3 as standards.
[0036] FIG. 20 shows seeds from additional sets of maize lines with the mutant a2 gene. Left to right, seeds from 501G, 507 A, 507 AB, 507G and 511H obtained from MaizeGDB.
[0037] FIG. 21 shows phloroglucinolysis of PAs extracted from seeds of additional maize fines with the mutant a2 gene using procyanidin B2 and B3 as standards.
[0038] FIG. 22 shows seeds of the ans mutant line M541G (a2 pll bl), the anthocyanin-rich line maize M141 A (A2 Pll Bl), the F2 progeny of the cross, and the F3 progeny of the cross.
[0039] FIG. 23 shows DMACA-stained leaf and root tissues of dissected seedlings of maize line 541G, maize line 141A, and their F3 cross
[0040] FIG. 24 shows seeds of the ans mutant line M142C (a2 pll bl), the anthocyanin overexpressing line M142Y (A2 Pll Bl), the F2 progeny of the cross, and the F3 progeny of the cross.
[0041] FIG. 25 shows DMACA-stained leaf and root tissues of dissected seedlings of maize lines 142C, 141 Y and their F3 cross.
[0042] FIG. 26 shows maize seeds from 501G, 971-31, and the F2 seeds of the cross between 501G and 971-31.
[0043] FIG. 27 shows DMACA-stained leaf and root tissues of dissected seedlings of maize lines 971-31, 501G and their F2 cross.
[0044] FIG. 28 shows DMACA-stained seeds of maize lines 1971-31, 501G, and their F2 cross.
[0045] FIG. 29 shows sequences of the A2 gene in wild-type (Ml 42V) and mutant (M541G) maize near the DNA insertion site. Numbers on the top show distance from the start codon. SEQ ID NO: 14 represents the A2 gene in wild-type maize. SEQ ID NO: 16 represents inserted DNA sequences in the mutant allele. The highlighted sequence (TGA) represents the early stop codon caused by the inserted sequence.BRIEF DESCRIPTION OF THE SEQUENCES
[0046] SEQ ID NO: 1 - representative nucleotide sequence that encodes a ZmPLl transcription factor.
[0047] SEQ ID NO:2 - representative amino acid sequence of a ZmPLl transcription factor.
[0048] SEQ ID NO:3 - representative nucleotide sequence that encodes a SbTT2 transcription factor.
[0049] SEQ ID NO:4 - representative amino acid sequence of a SbTT2 transcription factor.
[0050] SEQ ID NO:5 - representative nucleotide sequence that encodes a ZmR2 transcription factor.
[0051] SEQ ID NO:6 - representative amino acid sequence of a ZmR2 transcription factor.
[0052] SEQ ID NO:7 representative nucleotide sequence that encodes a ZmRl transcription factor.
[0053] SEQ ID NO:8 - representative amino acid sequence of a ZmRl transcription factor.
[0054] SEQ ID NO:9 -representative nucleotide sequence that encodes a ZmCl transcription factor.
[0055] SEQ ID NO: 10 - representative amino acid sequence of a ZmCl transcription factor.
[0056] SEQ ID NO: 11 - representative nucleotide sequence comprising the ANSIA2 wild-type allele found in Ml 42 A
[0057] SEQ ID NO: 12 - representative amino acid sequence encoded by the ANSIA2 wild-type allele found in M142A.
[0058] SEQ ID NO: 13 - representative maize genomic DNA sequence comprising the ans gene.
[0059] SEQ ID NO: 14 - representative nucleotide sequence of nucleotides 238-282 (from the start codon) of the A2 gene in wild-type maize.
[0060] SEQ ID NO: 15 - representative amino acid sequence encoded by SEQ ID NO: 14.
[0061] SEQ ID NO: 16 - representative inserted DNA sequence in the a.2 mutant allele.
[0062] SEQ ID NO: 17 - representative amino acid sequence encoded by SEQ ID NO: 16.
[0063] SEQ ID NO: 18 - representative nucleotide sequence that encodes a ZmB 1 transcription factor.
[0064] SEQ ID NO: 19 - - representative amino acid sequence of a ZmBl transcription factor.DETAILED DESCRIPTION OF THE INVENTION
[0065] The present disclosure provides methods and compositions for producing condensed tannins in maize. The present disclosure provides a significant advance in the art by providing maize plants comprising condensed tannins. Proanthocyanidins (PAs), or condensed tannins, are the second most abundant plant phenolic compounds after lignin. They are oligomers and polymers of flavan-3-ols and are produced in several tissues of vascular plants, providing protection from herbivores, fungal pathogens, and ultraviolet radiation. The antibacterial activities of PAs and their precursor flavan-3-ols and their beneficial effects in preventing cardiovascular disease make these compounds popular health supplements and targets for increasing food nutritional value. Addition of PAs to ruminant animal feed can improve nitrogen retention in the rumen and reduce pasture bloat caused by production of methane gas, improving animal health and productivity and reducing greenhouse gas emissions from livestock.
[0066] There is little evidence for accumulation of PAs in maize (Zea mays), the most important cereal crop cultivated worldwide, although maize makes anthocyanins and possessesthe key enzyme of the PA pathway, anthocyanidin reductase (ANR). The present disclosure describes that the endogenous PA biosynthetic machinery in maize preferentially produces the unusual PA precursor (+)-epicatechin, as well as 4 / ?-(S-cysteinyl)-catechin, as potential PA starter and extension units. Furthermore, uncommon procyanidin dimers with (+)-epicatechin as starter unit are also found in maize. The present disclosure further describes that expression of soybean (Glycine max) anthocyanidin reductase 1 (ANRI) in maize seeds increases the levels of 4 / ?-(S-cysteinyl)-epicatechin and procyanidin dimers mainly using (-)-epicatechin as starter units, and that introduction of a Sorghum bicolor transcription factor (SbTT2) specifically regulating PA biosynthesis into a maize inbred deficient in anthocyanin biosynthesis activates both anthocyanin and PA biosynthesis pathways, demonstrating conservation of the PA regulatory machinery across species. In addition, the present disclosure provides a significant advance in the art by providing compositions and methods for engineering condensed tannins in maize.A. Condensed Tannins and Maize
[0067] The chemistry of proanthocyanidins has been studied for decades. Their name reflects the fact that, on acid hydrolysis, the extension units are converted to colored anthocyanidins, and this forms the basis of the classical assay for these compounds. The building blocks of most proanthocyanidins are (+)-catechin and (-)-epicatechin. (-)-Epicatechin has 2,3-cis stereochemistry and (+)-catechin has 2,3-trans-stereochemistry. These stereochemical differences are of major importance in proanthocyanidins biosynthesis since all chiral intermediates in the flavonoid pathway up to and including leucoanthocyanidin are of the 2,3- trans stereochemistry. Condensed tannins are also commonly termed proanthocyanidins due to the red anthocyanidins that are produced upon heating in acidic alcohol solutions. The most common anthocyanidins produced are cyanidin (from procyanidin) and delphinidin (from prodelphinidin). Condensed tannins may contain from about 2 to about 50 or more flavonoid units. Condensed tannin polymers have complex structures because of variations in the flavonoid units and the sites for interflavan bonds. Depending on their chemical structure and degree of polymerization, condensed tannins may or may not be soluble in aqueous organic solvents.
[0068] In Arabidopsis thaliana, condensed tannins accumulate predominantly in the endothelium layer of the seed coat. Many mutations that affect the seed coat color and condensed tannin accumulation have been characterized and the corresponding genes cloned. These genes include BAN, TTG1 TTG2, T , TT2, TT8 and TT12. The ban mutation leads toaccumulation of anthocyanins in the seed coat instead of condensed tannins. TTG1, TTG2, TT1, TT2 and TT8 are regulatory genes and encode a WD-repeat protein, a WRKY family protein, a WIP subfamily plant zinc finger protein, an R2R3 MYB domain protein, and a basic helix-loop-helix domain protein, respectively. Expression of the TT2 gene has been shown to induce the TT8 and BAN genes but did not lead to accumulation of condensed tannins in Arabidopsis (Nesi et al., 2001). All the above genes are predominantly expressed in the seed coat endothelium.
[0069] Maize is known for having great diversity of flavonoids across cultivars, associated with seed coloration. Although anthocyanin and phlobaphene biosynthesis pathways in maize have been well established, oligomeric PAs appear to be absent in this species. The biosynthetic pathways of anthocyanins, phlobaphenes, PAs, and other flavonoids partially overlap, sharing several common precursors and enzymes. The present disclosure demonstrates that among different maize cultivars with an active anthocyanin biosynthetic pathway, those producing less anthocyanins tend to accumulate more epicatechin and other PA precursors. The increased PA precursor accumulation in bzl maize deficient in the enzyme converting cyanidins into anthocyanins demonstrate that PA biosynthesis is closely related to the flux into anthocyanin biosynthesis. The biosynthesis of phlobaphenes diverts prior to the anthocyanin pathway and is localized to pericarp tissues. Despite the lack of an apparent homolog of TT2, a transcription factor specifically activating PA biosynthesis in PA-rich plant species, overexpression of SbTT2 in maize induced the production of PAs. Surprisingly though, expression of SbTT2 in a maize inbred with an inactive anthocyanin biosynthesis pathway also led to anthocyanin accumulation, resulting in purple-colored seeds in transgenic tines. Cl and R family transcription factors regulate anthocyanin biosynthesis in maize with dark-colored seeds. The PA precursor and anthocyanin accumulation in seeds of SbTT2-OX transgenic lines resembled that of maize cultivars with functional transcription factors for anthocyanin biosynthesis, demonstrating that SbTT2 and maize anthocyanin-related transcription factors share conserved activities in regulating PA and anthocyanin biosynthesis in maize.
[0070] ANR is a key enzyme that produces both PA starter and extension units, and ANR product specificity determines PA composition and polymerization. Unlike ANRs from PA- rich species such as soybean that can efficiently produce (-)-epicatechin via coupled reactions with LDOX (absent in maize) using (+)-catechin as substrate, ZmANRl produces (+)- epicatechin at low efficiency in such in vitro LDOX-coupled reactions and favors theproduction of (+)-epicatechin over other flavan-3-ol stereoisomers when using cyanidin as substrate (23). Biochemical analyses of PAs in maize show that (+)-epicatechin is the epicatechin stereoisomer detected in bzl seeds, and procyanidin dimers with (+)-epicatechin as starter units (i.e., iso-B2 and iso-B4 dimers) are the predominant procyanidin dimers in seeds of SbTT2-OX, where endogenous ZmANRl expression was upregulated. Seeds of ST- GmANRl and OSA-GmANRl, on the other hand, produce classical procyanidin B2 and B4 dimers, which have (-)-epicatechin as starter units and are commonly found in PA-rich plant species. In addition, expression of GmANRl in ST or OSA maize dramatically increased 4 / ?- (S-cysteinyl)-epicatechin level in seeds, consistent with the finding that GmANRl is directly involved in the formation of 4 / ?-(S-cysteinyl)-epicatechin extension units. Interestingly, while ZmANRl only weakly catalyzes in vitro conversion of flaven-3, 4-diol to 2,3-czs-leucocyandin and 4 / ?-(S-cysteinyl)-epicatechin, trace amounts of 4 / ?-(S-cysteinyl)-epicatechin were detected in seeds of some maize varieties and expression of SbTT2 greatly enhanced its accumulation, demonstrating that there may be other factors or pathways within maize cells contributing to the production of 4 / ?-(S-cysteinyl)-epicatechin and that this mechanism can be activated by SbTT2.
[0071] Although expression of GmANRl in colored maize increased the levels of epicatechin and 4 / ?-(S-cysteinyl)-epicatechin starter and extension units, the level of PA polymers remained low, demonstrating an inefficient PA polymerization mechanism. Given that the assembly of PA polymers is likely nonenzymatic, maize may lack an appropriate cellular mechanism to stabilize and protect reactive PA intermediates during delivery to appropriate cell compartments for polymerization. This explains the presence of unusual (epi)catechin- anthocyanin conjugates in some maize varieties. The procyanidin dimers iso-B2 and iso-B4 were also detected in some maize varieties and in transgenic maize overexpressing SbTT2, showing that the (+)-epicatechin produced by ZmANRl can be used as starter unit to generate procyanidin dimers.B. Nucleic Acids and Proteins
[0072] The present disclosure provides plants, seeds, cells, and plant parts comprising the nucleic acid or proteins of the present disclosure. As used herein, the term "wild type" refers to the endogenous version of a molecule that naturally occurs in an organism. In some aspects of the present disclosure, a wild-type version of a protein or polypeptide may be employed. In many aspects of the present disclosure, however, a recombinant protein or polypeptide is employed. The terms recombinant protein and recombinant polypeptide may be usedinterchangeably with the terms modified protein, modified polypeptide, and variant. In certain embodiments, a recombinant protein refers to a protein having an altered chemical structure or amino acid sequence compared to the wild-type protein. In some aspects, a recombinant protein may have at least one modified activity or function compared to the wild-type protein. As is known in the art, proteins may have multiple activities or functions. In some embodiments, the function of a recombinant protein may be altered with respect to one activity or function but retain the activity or function of the wild-type protein in other respects. In certain embodiments, the proteins of the present disclosure may include those which comprise a mutation compared to the wild- type protein. In one embodiment, the mutation may comprise an insertion, a deletion, a truncation, or at least one amino acid substitution.
[0073] The nucleotide and protein sequences for various genes have been previously disclosed, and may be found in computerized databases known in the art. Two such databases are the National Center for Biotechnology Information's GenBank and GenPept databases (on the World Wide Web at ncbi.nlm.nih.gov / ) and The Universal Protein Resource (UniProt; on the World Wide Web at uniprot.org). The coding regions for these genes may be amplified or expressed using techniques known in the art or those disclosed herein.
[0074] As is known in the art, amino acid residues may be changed in a polypeptide sequence to create an equivalent, or even improved, second-generation variant polypeptide. For example, certain amino adds may be substituted for other amino acids in a polypeptide sequence without appreciable loss of binding capacity or specificity to structures such as, for example, binding sites on substrate molecules. Since the binding capacity, the binding specificity, and the nature of a protein define the functional activity of a protein, certain amino acid substitutions can be made in a protein sequence, and / or in its corresponding DNA coding sequence, such that the resultant variant protein comprises similar or desirable properties of the original protein. Thus, in some embodiments, the polynucleotide and polypeptide sequences of the present disclosure may comprise various amino acid or nucleic acid substitutions, deletions, and / or insertions without appreciable loss of biological utility or activity. As used herein the term "functionally equivalent codon" refers to codons that encode the same amino acid, such as the six different codons known in the art which code for arginine. As used herein, the terms "neutral substitutions" and "neutral mutations" refer to a change in a polypeptide sequence, or the encoding nucleotide sequence, such that the sequence comprises or encodes a biologically equivalent amino acid compared to that found in the original sequence. In certain embodiments, any amino acid of any polypeptide described herein maybe substituted with any biologically equivalent amino acid. Biologically equivalent amino acids are known in the art.
[0075] Nucleic acid or amino acid sequence variants of the disclosure may comprise, in some embodiments, a substitution, an insertion, or a deletion. A polypeptide variant of the disclosure may affect 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more non-contiguous or contiguous amino acids of the polypeptide, as compared to the referenced polypeptide or to the wild-type polypeptide, including any range derivable therebetween. A variant polypeptide may comprise, for example, an amino acid sequence having at least 50%, 60%, 70%, 80%, or 90% sequence identity to a sequence comprising or encoded by any one of SEQ ID NOs: 1-19 or fragments thereof, including all ranges derivable therebetween. A polypeptide may include, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino add substitutions.
[0076] It also will be understood that amino acid and nucleic acid sequences may include additional residues, such as additional N- or C-terminal amino acids, or 5' or 3' sequences, respectively, and yet still be essentially identical to the sequences provided by the present disclosure. Such essentially identical sequences, in some embodiments, may maintain the biological activity described herein for the sequences of the present disclosure. The addition of terminal sequences particularly applies to nucleic acid sequences that may, for example, include various non-coding sequences flanking either of the 5' or 3' portions of the coding region.
[0077] Deletion variants typically lack one or more amino acid residues compared to the protein from which the variant was derived, the native protein, or the wild-type protein. In certain embodiments, individual amino acid residues may be deleted, or a number of contiguous amino acids may be deleted. In one embodiment, a stop codon may be introduced, for example by substitution or insertion, into an encoding nucleic acid sequence to generate a truncated protein variant. In certain embodiments of the present disclosure a truncated protein variant may comprise about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 260, about 270, about 280, about 290, about 300, about 310, about 320, about 330, about 340, about 350, about 360, about 370, about 380, or about 390 contiguous amino acids of SEQ ID NO: 12, including all ranges and variable derivable therebetween. Methodsfor producing such truncated proteins are known in the art and any such method may be utilized according to certain embodiments of the present disclosure. A truncated protein, in particular embodiments, may have reduced, disrupted, or altered activity compared to the wild type protein. In some embodiments, a truncated protein may result in reduced, disrupted, or altered activity of the protein encoded by the ans gene in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant. According to further embodiments, the present disclosure provides a truncated protein encoded by an ans gene that results in reduced, disrupted, or altered activity in at least one plant tissue by 5%-20%, 5%-25%, 5%-30%, 5%-40%, 5%-50%, 5%-60%, 5%-70%, 5%-75%, 5%-80%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%-90%, 50%-75%, 25%- 75%, 30%-80%, or 10%-75%, as compared to a control plant.
[0078] Insertional variants typically involve the addition of one or more amino acid residues at a non-terminal point of a polypeptide. Terminal additions may also be generated and can include fusion proteins. Non-limiting examples of such fusion proteins include multimers or concatemers of one or more polypeptides provided by the present disclosure.
[0079] Substitutional variants typically comprise the exchange of one amino acid for another at one or more sites within the polypeptide. Substitutional variants may be designed to modulate one or more properties of the polypeptide, with or without the loss of other functions or properties of the polypeptide. Amino acid substitutions may be conservative amino acid substitutions. As used herein, the term “conservative amino acid substitution” refers to an amino acid substitution wherein one amino acid is replaced with another amino acid having similar chemical properties. Conservative amino acid substitutions may involve, for example, the exchange of a member of one amino acid class with another member of the same class. Conservative substitutions are well known in the art and include, for example, the changes of: alanine to serine; arginine to lysine; asparagine to glutamine or histidine; aspartate to glutamate; cysteine to serine; glutamine to asparagine; glutamate to aspartate; glycine to proline; histidine to asparagine or glutamine; isoleucine to leucine or valine; leucine to valine or isoleucine; lysine to arginine; methionine to leucine or isoleucine; phenylalanine to tyrosine, leucine or methionine; serine to threonine; threonine to serine; tryptophan to tyrosine; tyrosine to tryptophan or phenylalanine; and valine to isoleucine or leucine. Conservative amino acid substitutions may, in some embodiments, encompass non-naturally occurring amino acid residues, which are typically incorporated by chemical peptide synthesis rather than bysynthesis in biological systems. These include peptidomimetics or other reversed or inverted forms of amino acid moieties.
[0080] Alternatively, substitutions may be non-conservative. As used herein, the term “nonconservative amino acid substitution” refers to an amino acid substitution that affects a function of the polypeptide. Non-conservative amino acid substitutions typically involve substituting an amino acid residue with one that is chemically dissimilar. Non-conservative amino acid substitutions may include, for example, the substitution of a polar or charged amino acid for a nonpolar or uncharged amino acid, and vice versa. Non-conservative substitutions may also include, for example, the substitution of a member of one of the amino acid classes for a member from another class.
[0081] One skilled in the art can determine suitable polypeptide variants, as set forth herein, using well-known techniques. For example, one skilled in the art may identify suitable areas of the polypeptide molecule that may be changed without affecting activity by targeting regions not believed to be critical for activity. The skilled artisan will also be able to identify amino acid residues and portions of the polypeptide molecules that are conserved among similar proteins or polypeptides. In some embodiments, regions of a polypeptide molecule that may be important for biological activity or for structure may be subject to conservative amino acid substitutions without significantly altering the biological activity or adversely affecting the protein structure.
[0082] When making conservative or non-conservative amino acid substitutions, the hydropathy index of amino acids may be considered. The hydropathy profile of a protein is calculated by assigning each amino acid a numerical value ("hydropathy index") and then repetitively averaging these values along the peptide chain. Each amino acid has been assigned a value based on its hydrophobicity and charge characteristics. These values are isoleucine (+4.5); valine (+4.2); leucine (+3.8); phenylalanine (+2.8); cysteine (+2.5); methionine (+1.9); alanine (+1.8); glycine (-0.4); threonine (-0.7); serine (-0.8); tryptophan (-0.9); tyrosine (-1.3); proline (1.6); histidine (-3.2); glutamate (-3.5); glutamine (-3.5); aspartate (-3.5); asparagine (- 3.5); lysine (-3.9); and arginine (-4.5). The importance of the hydropathy amino acid index in conferring interactive biologic function on a protein is generally understood in the art (Kyte et al., J. Mol. Biol. 157:105-131 (1982)). It is accepted that the relative hydropathic character of the amino acid contributes to the secondary structure of the resultant protein or polypeptide, which in turn defines the interaction of the protein or polypeptide with other molecules. It is also known that certain amino acids may be substituted for other amino acids having a similarhydropathy index or score, and still retain a similar biological activity. In making changes based upon the hydropathy index, in certain aspects, the substitution of amino acids whose hydropathy indices are within ±2 is included. In some aspects of the disclosure, those that are within ±1 are included, and in other aspects of the disclosure, those within ±0.5 are included.
[0083] As is known in the art, the substitution of like amino acids can be effectively made based on hydrophilicity. U.S. Patent 4,554,101, incorporated herein by reference, states that the greatest local average hydrophilicity of a protein, as governed by the hydrophilicity of its adjacent amino acids, correlates with a biological property of the protein. The following hydrophilicity values have been assigned to these amino acid residues: arginine (+3.0); lysine (+3.0); aspartate (+3.0+1); glutamate (+3.0+1); serine (+0.3); asparagine (+0.2); glutamine (+0.2); glycine (0); threonine (-0.4); proline (-0.5+1); alanine (-0.5); histidine (-0.5); cysteine (-1.0); methionine (-1.3); valine (-1.5); leucine (-1.8); isoleucine (-1.8); tyrosine (-2.3); phenylalanine (-2.5); and tryptophan (-3.4). In some embodiments, amino acid substitutions based upon similar hydrophilicity values may include the substitution of amino acids whose hydrophilicity values are within +2 of each other. In one embodiment, the substitution of amino acids whose hydrophilicity values are within +1 or within +0.5 are included.
[0084] In certain aspects, the present disclosure provides polynucleotide molecules that encode the polypeptide molecules describes herein. Non-limiting example of such polynucleotide molecules include isolated polynucleotide segments, recombinant vectors, and recombinant polynucleotide molecules. A nucleic acid molecule is the “complement” of another nucleic acid molecule if they exhibit complete complementarity. As used herein, two molecules exhibit “complete complementarity” if when aligned every nucleotide of the first molecule is complementary to every nucleotide of the second molecule. Two molecules are “minimally complementary” if they can hybridize to one another with sufficient stability to permit them to remain annealed to one another under at least conventional “low stringency” conditions. Similarly, the molecules are “complementary” if they can hybridize to one another with sufficient stability to permit them to remain annealed to one another under conventional “high stringency” conditions. Departures from complete complementarity are therefore permissible, as long as such departures do not completely preclude the capacity of the molecules to form a double-stranded structure. As used herein, with respective to a given sequence, a “complement”, a “complementary sequence” and a “reverse complement” are used interchangeably. All three terms refer to the inversely complementary sequence of a nucleotidesequence, i.e., to a sequence complementary to a given sequence in reverse order of the nucleotides.
[0085] Appropriate stringency conditions that promote DNA hybridization, for example, 6.0 x sodium chloride / sodium citrate (SSC) at about 45 °C, followed by a wash of 2.0 x SSC at 50°C, are known to those skilled in the art or can be found in Current Protocols in Molecular Biology, John Wiley & Sons, N.Y. (1989), 6.3.1-6.3.6. For example, the salt concentration in the wash step can be selected from a low stringency of about 2.0 x SSC at 50°C to a high stringency of about 0.2 x SSC at 50°C. In addition, the temperature in the wash step can be increased from low stringency conditions at room temperature, about 22°C, to high stringency conditions at about 65°C. Both temperature and salt may be varied, or either the temperature or the salt concentration may be held constant while the other variable is changed.
[0086] By convention, the DNA sequences of the present disclosure and fragments thereof are disclosed with reference to only one strand of the two complementary DNA sequence strands. By implication and intent, the complementary sequences of the sequences provided here (the sequences of the complementary strand), also referred to in the art as the reverse complementary sequences, are within the scope of the disclosure and are expressly intended to be within the scope of the subject matter claimed. Thus, as used herein reference to any one of SEQ ID NOs:l, 3, 5, 7, 9, 11, 13, 14, 16, or 18 and fragments thereof include and refer to the sequence of the complementary strand and fragments thereof.
[0087] A polynucleotide molecule of the present disclosure may, in some embodiments, comprise a contiguous nucleic acid sequence that encodes all or part of a polypeptide described herein. In certain embodiments, a polypeptide described herein may be encoded by variant nucleic acid sequences that encode the same or a substantially similar protein.
[0088] The polynucleotide molecules and fragments thereof provided by the present disclosure may, in some embodiments be combined with other polynucleotide molecules which comprise elements, such as promoters, polyadenylation signals, additional restriction enzyme sites, multiple cloning sites, other coding segments, and the like, such that the overall length of the polynucleotide molecule may vary considerably. The polynucleotide molecule provided by the present disclosure can be any length. In some cases, a nucleic acid sequence may encode a polypeptide sequence that comprises additional heterologous coding sequences.
[0089] In some aspects of the present disclosure, a polypeptide or polynucleotide of the present disclosure may comprise an amino acid sequence or be encoded by a nucleotide sequencecomprising a sequence having at least about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity to any one of SEQ ID NOs:l-19, including any range derivable therebetween. In certain embodiments, a polypeptide or polynucleotide of the present disclosure comprise or be encoded by a fragment of a sequence having at least about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity to any one of SEQ ID NOs:l-19, including any range derivable therebetween. A polypeptide or polynucleotide of the present disclosure may comprise, for example, at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32,33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57,58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82,83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150,160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 275,300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 625, 650, 675, 700, 725, 750, 775, 800, 825, 850, 875, 900, 925, 950, 1000, 1100, or 1200 nucleotides or amino acid residues of a sequence having at least about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-19, including any range derivable therebetween.C. Transcription Factors
[0090] The present disclosure provides modified plants, seeds, cells, and plant parts that express one or more transcription factors having an amino acid sequence with at least about 70% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19.
[0091] ZmCl in maize belongs to the MYB family of transcription factors. ZmCl can activate anthocyanin biosynthesis in the aleurone layer of maize seeds. To activate anthocyanin biosynthesis, in some embodiments, ZmCl also requires involvement of the bHLH family of transcription factors, ZmRl or ZmR2. ZmCl activates the anthocyanin biosynthesis pathway by binding to the promoters of genes encoding enzymes in the pathway to induce transcription.
[0092] ZmPLl is a homologous gene to ZmCl, also belonging to the MYB family of transcription factors. ZmPLl activates anthocyanin biosynthesis in maize vegetative tissues. The activation of the anthocyanin pathway in maize, in certain embodiment, may also requireinvolvement of the bHLH family of transcription factors, ZmRl or ZmR2. They activate the pathway by binding to the promoters of enzymes in the pathway to induce transcription.
[0093] SbTT2 is a TT2-type MYB transcription factor from sorghum identified by phylogenetic analysis of sorghum, rice, M. truncatula and A. thaliana as a clade distinct from Cl-type MYBs. PA-rich monocots, including sorghum and rice, possess Cl-type, TT2-type and MYB5-type MYBs, whereas maize only appears to contain the Cl- and MYB5-type. Furthermore, TT8 (bHLH) and TTG1 (WD40) transcription factors are required for expression of PA biosynthesis in A. thaliana and M. truncatula.
[0094] ZmR2 is a homologous gene to ZmRl, belonging to bHLH family of transcription factors. ZmR2, when combined with ZmCl, ZmPLl or SbTT2, can activate anthocyanin and proanthocyanidin biosynthesis pathways in maize. Notably, this leads to increased leucoanthocyanidin which is converted to PAs in maize in the absence of ANS.
[0095] ZmRl in maize belongs to the bHLH family of transcription factors. ZmRl, when combined with ZmCl, ZmPLl or SbTT2, can activate anthocyanin and proanthocyanidin biosynthesis pathways in maize.
[0096] ZmBl in maize belongs to the bHLH family of transcription factors. ZmBl, when combined with ZmCl, ZmPLl, or SbTT2, can activate anthocyanin and proanthocyanidin biosynthesis pathways in maize.D. Genome Editing
[0097] The present disclosure provides, in certain embodiments, maize plants, plant parts, plant cells, and seeds produced through genome modification. As used herein the terms “genome modification” or “genomic modification” refer to genomic change in the expression level and / or endogenous sequence of one or more genes of interest relative to a wild-type gene. In some embodiments, a genomic modification may include, but is not limited to, a mutant allele, introduction of a transgene, a site-specific integration, and a genome edit. Genome editing can be used to make one or more edit(s) or mutation(s) at a desired target site in the genome of a plant, such as to change expression and / or activity of one or more genes, or to integrate an insertion sequence or transgene at a desired location in a plant genome. Any site or locus within the genome of a plant may potentially be chosen for making a genomic edit (or gene edit) or site-directed integration of a transgene, construct, or transcribable DNA sequence. As used herein, a “target site” for genome editing or site-directed integration refers to the location of a polynucleotide sequence within a plant genome that is bound and cleavedby a site-specific nuclease to introduce a double-stranded break (DSB) or single-stranded nick into the nucleic acid backbone of the polynucleotide sequence and / or its complementary DNA strand within the plant genome. A target site may comprise, for example, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 29, or at least 30 consecutive nucleotides. A “target site” for an RNA-guided nuclease may comprise the sequence of either complementary strand of a double-stranded nucleic acid (DNA) molecule or chromosome at the target site. A site-specific nuclease may bind to a target site, such as via a non-coding guide RNA e.g., without being limiting, a CRISPR RNA (crRNA) or a single-guide RNA (sgRNA) as described further herein). A noncoding guide RNA provided herein may be complementary to a target site e.g., complementary to either strand of a double- stranded nucleic acid molecule or chromosome at the target site). It will be appreciated that perfect identity or complementarity may not be required for a non-coding guide RNA to bind or hybridize to a target site. For example, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 mismatches (or more) between a target site and a non-coding RNA may be tolerated. A “target site” also refers to the location of a polynucleotide sequence within a plant genome that is bound and cleaved by any other site-specific nuclease that may not be guided by a non-coding RNA molecule, such as a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, etc., to introduce a DSB or single- stranded nick into the polynucleotide sequence and / or its complementary DNA strand. As used herein, a “target region” or a “targeted region” refers to a polynucleotide sequence or region that is flanked by two or more target sites. Without being limiting, in some embodiments a target region may be subjected to a mutation, deletion, insertion, substitution, inversion, or duplication. As used herein, “flanked” when used to describe a target region of a polynucleotide sequence or molecule, refers to two or more target sites of the polynucleotide sequence or molecule surrounding the target region, with one target site on each side of the target region.
[0098] As used herein, a “targeted genome editing technique” refers to any method, protocol, or technique that allows the precise and / or targeted editing of a specific location in a genome of a plant (i.e., the editing is largely or completely non-random) using a site-specific nuclease, such as a meganuclease, a zinc-finger nuclease (ZFN), an RNA-guided endonuclease (e.g., the CRISPR / Cas9 system or the CRISPR / Cpfl system), a TALE (transcription activator-like effector)-endonuclease (TALEN), a recombinase, or a transposase. As used herein, “editing”or “genome editing” refers to generating a targeted mutation, deletion, insertion, substitution, inversion, or duplication of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1000, at least 2500, at least 5000, at least 10,000, or at least 25,000 nucleotides of an endogenous plant genome nucleic acid sequence. As used herein, “editing” or “genome editing” may also encompass the targeted insertion or site-directed integration of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 10,000, or at least 25,000 nucleotides into the endogenous genome of a plant. An “edit” or “genomic edit” in the singular refers to one such targeted mutation, deletion, insertion, substitution, inversion, or duplication, whereas “edits” or “genomic edits” refers to two or more targeted mutation(s), deletion(s), insertion(s), substitution(s), inversion(s), and / or duplication(s), with each “edit” being introduced via a targeted genome editing technique.
[0099] According to some embodiments, a site-specific nuclease may be co-delivered with a donor template molecule to serve as a template for making a desired edit, mutation or insertion into the genome at the desired target site through repair of the double strand break (DSB) or nick created by the site-specific nuclease. According to some embodiments, a site-specific nuclease may be co-delivered with a DNA molecule comprising a selectable or screenable marker gene.
[0100] A site-specific nuclease provided herein may be selected from the group consisting of a zinc-finger nuclease (ZFN), a TALE-endonuclease (TALEN), a meganuclease, an RNA- guided endonuclease (e.g., Cas9 and Cpfl), a recombinase, a transposase, or any combination thereof. See, e.g., Khandagale et al. (Plant Biotechnol Rep 10:327-343, 2016); and Gaj et al. (Trends Biotechnol. 31 (7):397-405, 2013). Zinc finger nucleases (ZFN) are synthetic proteins consisting of an engineered zinc finger DNA-binding domain fused to a cleavage domain (or a cleavage half-domain), which may be derived from a restriction endonuclease (e.g., Fold). The DNA binding domain may be canonical (C2H2) or non-canonical (e.g., C3H or C4). The DNA-binding domain can comprise one or more zinc fingers (e.g., 2, 3, 4, 5, 6, 7, 8, 9 or more zinc fingers) depending on the target site but may typically be composed of 3-4 (or more) zinc-fingers. Multiple zinc fingers in a DNA-binding domain may be separated by linkersequence(s). ZFNs can be designed to cleave almost any stretch of double-stranded DNA by modification of the zinc finger DNA-binding domain. ZFNs form dimers from monomers composed of a non-specific DNA cleavage domain (e.g., derived from the FokI nuclease) fused to a DNA-binding domain comprising a zinc finger array engineered to bind a target site DNA sequence. The amino acids at positions -1, +2, +3, and +6 relative to the start of the zinc finger a-helix, which contribute to site-specific binding to the target site, can be changed and customized to fit specific target sequences. The other amino acids may form a consensus backbone to generate ZFNs with different sequence specificities.
[0101] Methods and rules for designing ZFNs for targeting and binding to specific target sequences are known in the art. See, e.g., U.S. Patent App. Pub. Nos. 2005 / 0064474, 2009 / 0117617, and 2012 / 0142062. The Fold nuclease domain may require dimerization to cleave DNA and therefore two ZFNs with their C-terminal regions are needed to bind opposite DNA strands of the cleavage site (separated by 5-7 bp). The ZFN monomer can cut the target site if the two-ZF-binding sites are palindromic. A ZFN, as used herein, is broad and includes a monomeric ZFN that can cleave double stranded DNA without assistance from another ZFN. The term ZFN may also be used to refer to one or both members of a pair of ZFNs that are engineered to work together to cleave DNA at the same site. Because the DNA-binding specificities of zinc finger domains can be re-engineered using one of various methods, customized ZFNs can theoretically be constructed to target nearly any target sequence (e.g., at or near a gene in a plant genome). Publicly available methods for engineering zinc finger domains include Context-dependent Assembly (CoDA), Oligomerized Pool Engineering (OPEN), and Modular Assembly.
[0102] Transcription activator-like effectors (TALEs) can be engineered to bind practically any DNA sequence, such as at or near the genomic locus of a gene in a plant. TALE has a central DNA-binding domain composed of 13-28 repeat monomers of 33-34 amino acids. The amino acids of each monomer are highly conserved, except for hypervariable amino acid residues at positions 12 and 13. The two variable amino acids are called repeat- variable diresidues (RVDs). The amino acid pairs NI, NG, HD, and NN of RVDs preferentially recognize adenine, thymine, cytosine, and guanine / adenine, respectively, and modulation of RVDs can recognize consecutive DNA bases. This simple relationship between amino acid sequence and DNA recognition has allowed for the engineering of specific DNA binding domains by selecting a combination of repeat segments containing the appropriate RVDs.
[0103] TALENs are artificial restriction enzymes generated by fusing the TALE DNA binding domain to a nuclease domain. In some aspects, the nuclease is selected from a group consisting of PvuII, MutH, TevI, FokI, Alwl, Mlyl, Sbfl, Sdal, StsI, CleDORF, Clo051, and Pept071. When each member of a TALEN pair binds to the DNA sites flanking a target site, the Fold monomers dimerize and cause a double-stranded DNA break at the target site. The term TALEN, as used herein, is broad and includes a monomeric TALEN that can cleave double stranded DNA without assistance from another TALEN. The term TALEN also refers to one or both members of a pair of TALENs that work together to cleave DNA at the same site.
[0104] Besides the wild-type Fold cleavage domain, variants of the Fold cleavage domain with mutations have been designed to improve cleavage specificity and cleavage activity. The Fold domain functions as a dimer, requiring two constructs with unique DNA binding domains for sites in the target genome with proper orientation and spacing. Both the number of amino acid residues between the TALEN DNA binding domain and the Fold cleavage domain and the number of bases between the two individual TALEN binding sites are parameters for achieving high levels of activity. PvuII, MutH, and TevI cleavage domains are useful alternatives to Fold and Fold variants for use with TALEs. PvuII functions as a highly specific cleavage domain when coupled to a TALE (see Yank et al., PLoS One 8:e82539, 2013). MutH is capable of introducing strand-specific nicks in DNA (see Gabsalilow et al., Nucleic Acids Research. 41:e83, 2013). TevI introduces double-stranded breaks in DNA at targeted sites (see Beurdeley et al., Nature Communications 4:1762, 2013).
[0105] The relationship between amino acid sequence and DNA recognition of the TALE binding domain allows for designable proteins. Software programs such as DNAWorks can be used to design TALE constructs. Other methods of designing TALE constructs are known to those of skill in the art. See Doyle et al. (Nucleic Acids Research 40: W117-122, 2012); Cermak et al. (Nucleic Acids Research 39:e82, 2011); and tale-nt.cac.comell.edu / about. In another aspect, a TALEN provided herein is capable of generating a targeted DSB.
[0106] A site-specific nuclease may be a meganuclease. Meganucleases, which are commonly identified in microbes, such as the LAGLID ADG family of homing endonucleases, are unique enzymes with high activity and long recognition sequences (> 14 bp) resulting in site-specific digestion of target DNA. Engineered versions of naturally occurring meganucleases typically have extended DNA recognition sequences (for example, 14 to 40 bp). The engineering of meganucleases can be more challenging than ZFNs and TALENsbecause the DNA recognition and cleavage functions of meganucleases are intertwined in a single domain. Specialized methods of mutagenesis and high-throughput screening have been used to create novel meganuclease variants that recognize unique sequences and possess improved nuclease activity.
[0107] A site-specific nuclease may be an RNA-guided nuclease. In an aspect, the targeted genome editing described herein may comprise the use of an RNA-guided endonuclease. As used herein, an “RNA-guided nuclease” refers to an RNA-guided DNA endonuclease associated with the CRISPR system. According to some embodiments, an RNA-guided endonuclease may be selected from the group consisting of Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslO, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, Cpfl (also known as Casl2a, see e.g., Safari, F. et al., Cell Biosci 9:36, 2019), CasX, CasY, and homologs or modified versions of any thereof, as well as Argonaute proteins (non-limiting examples of Argonaute proteins include Thermus thermophilus Argonaute (TtAgo), Pyrococcus furiosus Argonaute (PfAgo), Natronobacterium gregoryi Argonaute (NgAgo), and homologs or modified versions of any thereof). According to some embodiments, an RNA-guided endonuclease is a Cas9 or Cpfl enzyme. According to some embodiments, an RNA-guided endonuclease is a Cpfl enzyme.
[0108] The CRISPR system, in its native context, provide bacteria and archaea with immunity to invading foreign nucleic acids and relies on an RNA-guided endonuclease to cleave the invading DNA or RNA into short sequence fragments and incorporating them into the bacterial CRISPR genomic locus. The incorporated short sequences, referred to as “protospacers”, and flanking direct repeats are transcribed and processed into CRISPR RNAs (crRNAs). These crRNAs hybridize with frans-activating crRNAs (tracrRNAs) to activate the RNA-guided Cas endonuclease to form a ribonucleoprotein (RNP) complex that is guided to a target site. A prerequisite for cleavage of the target site, however, is the presence of a conserved genomic protospacer-adjacent motif sequence recognized by the Cas endonuclease. A “protospacer adjacent motif’ (PAM) herein refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that is recognized (targeted) by a guide polynucleotide / Cas endonuclease system described herein. A PAM may be present in the genome immediately adjacent and upstream to the 5’ end of the genomic target site sequence complementary to the targeting sequence of the guide RNA - i.e., immediately downstream (3’) to the sense (+)strand of the genomic target site (relative to the targeting sequence of the guide RNA) as known in the art. See, e.g., Wu et al. {Quant Biol. 2(2):59-70, 2014). The Cas endonuclease may not successfully recognize a target DNA sequence if the target DNA sequence is not followed by a PAM sequence. The sequence and length of a PAM sequence herein can differ depending on the Cas endonuclease used. The PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides long.
[0109] CRISPR / Cas9, which is the CRISPR system from Streptococcus pyogenes, was adapted for use in eukaryotes and has been widely used for gene editing in plants. The CRISPR / Cas9 system requires both crRNA and tracrRNA to guide the Cas9 protein to recognize and cleave the target DNA double helix. Cas9 recognizes the genomic PAM sequence 5’-NGG-3’ (where N is any nucleotide) and, when located on the sense (+) strand adjacent to the target site, will create a blunt-end DSB at the target site, specifically the 5'-end of the PAM site. Cas9 has been observed to recognize other PAM sequences, such as 5’- NAG-3’and 5’-NGA-3,’ which may result in cleavage of non-specific DNA sequences. However, the corresponding sequence of the guide RNA {i.e., immediately downstream (3’) to the targeting sequence of the guide RNA) may generally not be complementary to the genomic PAM sequence.
[0110] Recently, the CRISPR / Cpfl system was discovered as an alternative to the CRISPR / Cas9 system for genome editing. While CRISPR / Cpfl functions in a manner similar to CRISPR / Cas9, it is an even simpler system than CRISPR / Cas9. CRISPR / Cpfl requires only one crRNA molecule and no tracrRNA to cleave DNA. Cpfl recognizes the genomic PAM sequence 5’-TTTV-3’ (where V is A, G, or C) or 5’-TTN-3’, depending on the Cpfl ortholog. See e.g., Alok et al. {Front. Plant Sci. 11:264, 2020). When Cpfl recognizes the genomic PAM located on the sense (+) strand adjacent to the target site, it will generate a staggered DSB with a 4 or 5-nt 5' overhang at the target site, specifically the 3 '-end of the PAM site.
[0111] The RNA-guided nuclease may be delivered as a protein with or without a guide RNA, or the guide RNA may be complexed with the RNA-guided nuclease enzyme and delivered as a ribonucleoprotein (RNP).
[0112] For RNA-guided endonucleases, a guide RNA molecule may be further provided to direct the endonuclease to a target site in the genome of the plant via base-pairing or hybridization to cause a DSB or nick at or near the target site. The guide RNA may betransformed or introduced into a plant cell or tissue as a gRNA molecule, or as a recombinant DNA molecule, construct or vector comprising a transcribable DNA sequence encoding the guide RNA operably linked to a promoter. As understood in the art, a guide RNA may comprise, for example, a CRISPR RNA (crRNA), a single-chain guide RNA (sgRNA), or any other RNA molecule that may guide or direct an endonuclease to a specific target site in the genome. A prototypical CRISPR associated protein, Cas9 from S. pyogenes, naturally binds two RNAs, a CRISPR RNA (crRNA) guide and a trans-acting CRISPR RNA (tracrRNA), to assemble a CRISPR ribonucleoprotein (crRNP). A “single-chain guide RNA” (or “sgRNA”) is an RNA molecule comprising a crRNA covalently linked to a tracrRNA by a linker sequence, which may be expressed as a single RNA transcript or molecule. The guide RNA comprises a guide or targeting sequence (also referred to herein as a “spacer sequence”) that is identical or complementary to a target site within the plant genome, such as at or near a gene. The guide RNA is typically a non-coding RNA molecule that does not encode a protein. The guide sequence of the guide RNA may be at least 10 nucleotides in length, such as 12-40 nucleotides, 12-30 nucleotides, 12-20 nucleotides, 12-35 nucleotides, 12-30 nucleotides, 15- 30 nucleotides, 17-30 nucleotides, or 17-25 nucleotides in length, or about 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more nucleotides in length. The guide sequence may be at least 95%, at least 96%, at least 97%, at least 99% or 100% identical or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or more consecutive nucleotides of a DNA sequence at the genomic target site.
[0113] In addition to the guide sequence, a guide RNA may further comprise one or more other structural or scaffold sequence(s), which may bind or interact with an RNA-guided endonuclease. Such scaffold or structural sequences may further interact with other RNA molecules (e.g., tracrRNA). Methods and techniques for designing targeting constructs and guide RNAs for genome editing and site-directed integration at a target site within the genome of a plant using an RNA-guided endonuclease are known in the art.
[0114] As mentioned above, a target gene for genome editing may be the Zea mays ans gene. For modification of the ans gene through genome editing, an RNA-guided endonuclease may be targeted to a transcribable DNA sequence (i.e., a transcribable region) of said gene. For example, in certain embodiments a transcribable DNA sequence targeted for genome editing may comprise an exon / intron boundary or may be in close proximity to an exon / intron boundary. If the resulting modification spans an exon / intron boundary, the modification maybe referred to as a modification in an exon region and an intron region. For genetic modification of the ans gene, a guide RNA may be used, which comprises a guide sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or more consecutive nucleotides of SEQ ID NO: 13 or a sequence complementary thereto (e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more consecutive nucleotides of SEQ ID NO: 13 or a sequence complementary thereto), although alternative splicing and different exon / intron boundaries may occur. As used herein, the term “consecutive” in reference to a polynucleotide or protein sequence means without deletions or gaps in the sequence.
[0115] As used herein, the term “antisense” refers to DNA or RNA sequences that are complementary to a specific DNA or RNA sequence. Antisense RNA molecules are singlestranded nucleic acids which can combine with a sense RNA strand or sequence or mRNA to form duplexes due to complementarity of the sequences. The term “antisense strand” refers to a nucleic acid strand that is complementary to the “sense” strand. The “sense strand” of a gene or locus is the strand of DNA or RNA that has the same sequence as an RNA molecule transcribed from the gene or locus (with the exception of uracil in RNA and thymine in DNA).
[0116] A protospacer-adjacent motif (PAM) may be present in the genome immediately adjacent and upstream to the 5’ end of the genomic target site sequence complementary to the targeting sequence of the guide RNA - i.e., immediately downstream (3’) to the sense (+) strand of the genomic target site (relative to the targeting sequence of the guide RNA) as known in the art. See, e.g., Wu et al. (Quant Biol. 2(2):59-70, 2014). However, the corresponding sequence of the guide RNA (i.e., immediately downstream (3’) to the targeting sequence of the guide RNA) may generally not be complementary to the genomic PAM sequence.
[0117] In some embodiments, a site-specific nuclease is a recombinase. Non-limiting examples of recombinases that may be used include a serine recombinase attached to a DNA recognition motif, a tyrosine recombinase attached to a DNA recognition motif, or any recombinase enzyme known in the art attached to a DNA recognition motif. In certain embodiments, the site-specific nuclease is a recombinase or transposase, which may be a DNA transposase or recombinase attached or fused to a DNA binding domain. Non-limiting examples of recombinases include a tyrosine recombinase selected from the group consistingof a Cre recombinase, a Gin recombinase, a Flp recombinase, and a Tnpl recombinase attached to a DNA recognition motif provided herein. In one aspect of the present disclosure, a Cre recombinase or a Gin recombinase provided herein is tethered to a zinc-finger DNA- binding domain, a TALE DNA-binding domain, or a Cas9 nuclease. In another aspect, a serine recombinase selected from the group consisting of a PhiC31 integrase, an R4 integrase, and a TP-901 integrase may be attached to a DNA recognition motif provided herein. In yet another aspect, a DNA transposase selected from the group consisting of a TALE-piggyBac and TALE-Mutator may be attached to a DNA binding domain provided herein.
[0118] Several site-specific nucleases, such as recombinases, zinc finger nucleases (ZFNs), meganucleases, and TALENs, are not RNA-guided and instead rely on their protein structure to determine their target site for causing the DSB or nick, or they are fused, tethered or attached to a DNA-binding protein domain or motif. The protein structure of the site- specific nuclease (or the fused / attached / tethered DNA binding domain) may target the site-specific nuclease to the target site. According to many of these embodiments, non-RNA-guided sitespecific nucleases, such as recombinases, zinc finger nucleases (ZFNs), meganucleases, and TALENs, may be designed, engineered and constructed according to known methods to target and bind to a target site at or near the genomic locus of an endogenous gene of a plant to create a DSB or nick at such a genomic locus. The DSB or nick created by the non-RNA-guided site-specific nuclease may lead to knockdown of gene expression, or a change in the activity of the protein encoded by the endogenous gene, via repair of the DSB or nick, which may result in a mutation or insertion of a sequence at the site of the DSB or nick through cellular repair mechanisms. Such cellular repair mechanism may be guided by a donor template molecule.
[0119] As used herein, a “donor molecule”, “donor template”, or “donor template molecule” (collectively a “donor template”), which may be a recombinant polynucleotide, DNA or RNA donor template or sequence, is defined as a nucleic acid molecule having a homologous nucleic acid template or sequence (e.g., homology sequence) and / or an insertion sequence for site-directed, targeted insertion or recombination into the genome of a plant cell via repair of a nick or DSB in the genome of a plant cell. A donor template may be a separate DNA molecule comprising one or more homologous sequence(s) and / or an insertion sequence for targeted integration, or a donor template may be a sequence portion (z.e., a donor template region) of a DNA molecule further comprising one or more other expression cassettes, genes / transgenes, and / or transcribable DNA sequences. For example, a “donor template” maybe used for site-directed integration of a transgene or construct, or as a template to introduce a mutation, such as an insertion, deletion, substitution, etc., into a target site within the genome of a plant. A targeted genome editing technique provided herein may comprise the use of one or more, two or more, three or more, four or more, or five or more donor molecules or templates. A donor template provided herein may comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten gene(s) or transgene(s) and / or transcribable DNA sequence(s). Alternatively, a donor template may comprise no genes, transgenes or transcribable DNA sequences.
[0120] Without being limiting, a gene / transgene or transcribable DNA sequence of a donor template may include, for example, an insecticidal resistance gene, an herbicide tolerance gene, a nitrogen use efficiency gene, a water use efficiency gene, a yield enhancing gene, a nutritional quality gene, a DNA binding gene, a selectable marker gene, an RNAi or suppression construct, a site- specific genome modification enzyme gene, a single guide RNA of a CRISPR / Cas9 system, a geminivirus-based expression cassette, or a plant viral expression vector system. According to other embodiments, an insertion sequence of a donor template may comprise a protein encoding sequence or a transcribable DNA sequence that encodes a non-coding RNA molecule, which may target an endogenous gene for suppression. A donor template may comprise a promoter operably linked to a coding sequence, gene, or transcribable DNA sequence, such as a constitutive promoter, a tissue-specific or tissuepreferred promoter, a developmental stage promoter, or an inducible promoter. A donor template may comprise a leader, enhancer, promoter, transcriptional start site, 5’-UTR, one or more exon(s), one or more intron(s), transcriptional termination site, region or sequence, 3’-UTR, and / or polyadenylation signal, which may each be operably linked to a coding sequence, gene (or transgene) or transcribable DNA sequence encoding a non-coding RNA, a guide RNA, an mRNA and / or protein. A donor template may be a single- stranded or doublestranded DNA or RNA molecule or plasmid.
[0121] An “insertion sequence” of a donor template is a sequence designed for targeted insertion into the genome of a plant cell, which may be of any suitable length. For example, the insertion sequence of a donor template may be between 2 and 50,000, between 2 and 10,000, between 2 and 5000, between 2 and 1000, between 2 and 500, between 2 and 250, between 2 and 100, between 2 and 50, between 2 and 30, between 15 and 50, between 15 and 100, between 15 and 500, between 15 and 1000, between 15 and 5000, between 18 and 30, between 18 and 26, between 20 and 26, between 20 and 50, between 20 and 100, between 20and 250, between 20 and 500, between 20 and 1000, between 20 and 5000, between 20 and 10,000, between 50 and 250, between 50 and 500, between 50 and 1000, between 50 and 5000, between 50 and 10,000, between 100 and 250, between 100 and 500, between 100 and 1000, between 100 and 5000, between 100 and 10,000, between 250 and 500, between 250 and 1000, between 250 and 5000, or between 250 and 10,000 nucleotides or base pairs in length. A donor template may also have at least one homology sequence or homology arm, such as two homology arms, to direct the integration of a mutation or insertion sequence into a target site within the genome of a plant via homologous recombination, wherein the homology sequence or homology arm(s) are identical or complementary, or have a percent identity or percent complementarity, to a sequence at or near the target site within the genome of the plant. When a donor template comprises homology arm(s) and an insertion sequence, the homology arm(s) will flank or surround the insertion sequence of the donor template. Each homology arm may be at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 99% or 100% identical or complementary to at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 500, at least 1000, at least 2500, or at least 5000 consecutive nucleotides of a target DNA sequence within the genome of a plant.
[0122] Any method known in the art for site-directed integration may be used with the present disclosure. In the presence of a donor template molecule with an insertion sequence, the DSB or nick can be repaired by homologous recombination between homology arm(s) of the donor template and the plant genome, or by non-homologous end joining (NHEJ), resulting in site-directed integration of the insertion sequence into the plant genome to create the targeted insertion event at the site of the DSB or nick. Thus, site-specific insertion or integration of a transgene, transcribable DNA sequence, construct, or sequence may be achieved if the transgene, transcribable DNA sequence, construct or sequence is located in the insertion sequence of the donor template.
[0123] The introduction of a DSB or nick may also be used to introduce targeted mutations in the genome of a plant, including genomic modifications that reduce or disrupt the activity of ANS, as compared to the activity of ANS in an otherwise identical maize plant, maize plant seed, maize plant part, or maize plant cell that lacks the modification. As used herein, a “mutation” refers to the permanent alteration of the nucleotide sequence of the genome of an organism, the extrachromosomal DNA, or other genetic elements, e.g., targeted mutationswithin a genomic region comprising SEQ ID NO: 13. According to this approach, mutations, such as deletions, insertions, substitutions, inversions, and / or duplications may be introduced at a target site via imperfect repair of the DSB or nick to produce a genetic modification within a gene. Such mutations may be generated by imperfect repair of the targeted locus even without the use of a donor template molecule. A modification of a gene may be achieved by inducing a DSB or nick at or near the endogenous locus of the gene that results in expression of a non-functional protein, interfering protein, or a protein having reduced, disrupted, or altered activity as compared to a protein expressed from the gene lacking said modification.
[0124] As used herein, the term “insertion” as it relates to a mutation, refers to the addition of one or more extra nucleotides into the DNA. Insertions in the coding region of a gene may alter splicing of the mRNA (splice site mutation) or cause a shift in the reading frame (frameshift), both of which can significantly alter the gene product.
[0125] As used herein, the term “deletion” as it relates to a mutation refers to the removal of one or more nucleotides from the DNA. Like insertion mutations, these mutations can alter the reading frame of the gene.
[0126] As used herein, the term “substitution” as it relates to a mutation refers to an exchange of a single nucleotide for another.
[0127] As used herein, the term “inversion” refers to reversing the orientation of a chromosomal segment. An inversion can be accompanied by a loss of nucleotides flanking either one or both sites of the inversion due to DNA repair mechanisms occurring at the cut and ligation sites during the formation of an inversion.
[0128] As used herein, the term “duplication” refers to the creation of multiple copies of chromosomal regions, increasing the dosage of the genes located within them.
[0129] As used herein, a “missense mutation” refers to a single nucleotide change that results in a codon that codes for a different amino acid. For example, the codon “CGU” encodes an arginine amino acid. If a missense mutation changes the G to a U, producing a “CUU” codon, the codon now encodes a leucine amino acid. Missense mutations can be caused by an insertion, deletion, substitution, duplication, or inversion. The frameshift, missense, or nonsense mutations described herein lead to loss of function or expression of a targeted gene, such as an ans gene. A “loss-of-function mutation” is a mutation in the coding sequence of a gene, which causes the function of the gene product, usually a protein, to be either reduced or completely absent, e.g., reduction of peptidase activity. A loss-of-function mutation can, forinstance, be caused by the truncation of the gene product. A phenotype associated with an allele with a loss of function mutation can be either recessive or dominant.
[0130] Similarly, such targeted mutations of a gene may be generated with a donor template molecule to direct a particular or desired mutation at or near the target site via repair of the DSB or nick. The donor template molecule may comprise a homologous sequence with or without an insertion sequence and comprising one or more mutations, such as one or more deletions, insertions, substitutions, inversions, and / or duplications, relative to the targeted genomic sequence at or near the site of the DSB or nick. For example, targeted mutations of a gene may be achieved by deleting, inserting, substituting, inverting, or duplicating at least a portion of the gene, such as by introducing a frame shift or premature stop codon into the coding sequence of the gene or introducing a modification into a transcribable DNA sequence. A deletion of a portion of a gene may also be introduced by generating DSBs or nicks at two target sites and causing a deletion of the intervening target region flanked by the target sites. A modification of a targeted gene may result in expression of a non-functional protein, interfering protein, or a protein having reduced, disrupted, or altered activity as compared to a protein expressed from the gene lacking said modification.
[0131] In an aspect, the present disclosure provides a modified maize plant, or plant seed, plant part or plant cell thereof, comprising a mutant allele of the ans gene. As used herein the term “mutant allele” refers to an allele that differs from the allele found in the standard or wild type organism. In some embodiments, a mutant allele comprises at least one genome modification involving at least 1, at least 2, at least 3, at least 4, at least 6, at least 8, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 115, at least 125, at least 150, at least 200, at least 300, at least 600, or at least 650 consecutive nucleotides of a transcribable region of the endogenous ans gene. A transcribable DNA sequence of the ans gene comprises the sequence of SEQ ID NO: 11, which is an approximately 1.5 kb polynucleotide sequence within the ans gene. The genome modification may be a deletion of a region comprising at least 1, at least 2, at least, 3, at least 4, at least 6, at least 8, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 115, at least 125, least 150, at least 200, at least 300, at least 600 or at least 650 consecutive nucleotides within the sequence of SEQ ID NO: 13. In an aspect, the genome modification may also comprise a deletion and nucleotide substitutions or nucleotide insertions of at least 1, at least 2, at least3, at least 4, at least 6, at least 8, at least 10, or at least 20 consecutive nucleotides around the deletion. Other targeted modifications may be made in the transcribable region to generate novel alleles in the ans gene. For example, one or more modification sites may be located upstream from the 3’ end of reference sequence SEQ ID NO: 13, or one or more modification sites may also be located downstream from the 5’ end of reference sequence SEQ ID NO: 13.
[0132] In another aspect, a modified maize plant, maize plant seed, maize plant part, or maize plant cell provided herein may, in certain embodiments, comprise at least one modification, wherein the modification comprises a deletion, an insertion, a substitution, an inversion, a duplication, or any combination thereof, in at least one allele of an endogenous ans gene.E. Constructs for Genome Editing
[0133] Recombinant DNA constructs and vectors are provided comprising a polynucleotide sequence encoding a site-specific nuclease, such as a zinc-finger nuclease (ZFN), a meganuclease, an RNA-guided endonuclease, a TALE-endonuclease (TALEN), a recombinase, or a transposase, wherein the coding sequence is operably linked to a plant expressible promoter. For RNA-guided endonucleases, recombinant DNA constructs and vectors are further provided comprising a polynucleotide sequence encoding a guide RNA, wherein the guide RNA comprises a guide sequence of sufficient length having a percent identity or complementarity to a target site within the genome of a plant, such as at or near a targeted ans gene. A polynucleotide sequence of a recombinant DNA construct and vector that encodes a site-specific nuclease or a guide RNA may be operably linked to a plant expressible promoter, such as an inducible promoter, a constitutive promoter, a tissue-specific promoter, etc.
[0134] In an aspect, vectors comprising polynucleotides encoding a site- specific nuclease, and optionally one or more, two or more, three or more, or four or more gRNAs are provided to a plant cell by transformation methods known in the art (e.g., without being limiting, particle bombardment, PEG-mediated protoplast transfection or Agrobacterium-mediated transformation). In an aspect, vectors comprising polynucleotides encoding a Cpfl nuclease, and optionally one or more, two or more, three or more, or four or more gRNAs are provided to a plant cell by transformation methods known in the art (e.g., without being limiting, particle bombardment, PEG-mediated protoplast transfection or Agrobacterium-mediated transformation). In another aspect, vectors comprising polynucleotides encoding a Cpfl and,optionally one or more, two or more, three or more, or four or more crRNAs are provided to a cell by transformation methods known in the art {e.g., without being limiting, viral transfection, particle bombardment, PEG-mediated protoplast transfection or Agrobacterium- mediated transformation).
[0135] As used herein, a “gene” refers to a nucleic acid sequence forming a genetic and functional unit and coding for one or more sequence-related RNA and / or polypeptide molecules. A gene generally contains a coding region operably linked to appropriate regulatory sequences that regulate the expression of a gene product e.g., a polypeptide or a functional RNA). A gene can have various sequence elements, including, but not limited to, a promoter, an untranslated region (UTR), exons, introns, and other upstream or downstream regulatory sequences.
[0136] As used herein, “locus” is a chromosomal locus or region where a polymorphic nucleic acid, trait determinant, gene, or marker is located. A “locus” can be shared by two homologous chromosomes to refer to their corresponding locus or region. As used herein, an “allele” refers to an alternative nucleic acid sequence of a gene or at a particular locus e.g., a nucleic acid sequence of a gene or locus that is different than other alleles for the same gene or locus). Such an allele can be considered (i) wild-type or (ii) mutant if one or more mutations or edits are present in the nucleic acid sequence of the mutant allele relative to the wild-type allele. A mutant or edited allele for a gene may have reduced, disrupted, altered, or eliminated activity, or a reduced or eliminated expression level for the gene relative to the wild-type allele. For example, a mutant or edited allele for the ans gene may have a deletion in the transcribable region of the endogenous ans gene that reduces, disrupts, or alters the activity of the protein encoded by the mutant allele as compared to the activity of the protein encoded by the wild-type allele in an otherwise identical maize plant. For diploid organisms such as maize, a first allele can occur on one chromosome, and a second allele can occur at the same locus on a second homologous chromosome. If one allele at a locus on one chromosome of a plant is a mutant or edited allele and the other corresponding allele on the homologous chromosome of the plant is wild type, then the plant is described as being heterozygous for the mutant or edited allele. However, if both alleles at a locus are mutant or edited alleles, then the plant is described as being homozygous or biallelic for the mutant or edited alleles. As used herein, the term “homozygous” refers to a genotype comprising two identical alleles at a given locus in a diploid genome. Given that maize is a diploid organism, CRISPR-mediated gene editing can result in biallelic (that is, different edits are made to thesame locus on corresponding homologous chromosomes) edits resulting in a genotype comprising two non-identical mutant alleles at a given locus in a diploid genome in Ro plants. When used in the context of edited alleles, plants comprising such genotypes may also be referred to as comprising a heteroallelic combination or biallelic edits.
[0137] As used herein, a “wild-type gene” or “wild-type allele” refers to a gene or allele having a sequence or genotype that is most common in a particular plant species, or another sequence or genotype having only natural variations, polymorphisms, or other silent mutations relative to the most common sequence or genotype that do not significantly impact the expression and activity of the gene or allele. Indeed, a “wild type” gene or allele contains no variation, polymorphism, or any other type of mutation that substantially affects the normal function, activity, expression, or phenotypic consequence of the gene or allele relative to the most common sequence or genotype.
[0138] In general, the term “variant” refers to molecules with some differences, generated synthetically or naturally, in their nucleotide or amino acid sequences as compared to a reference (native) polynucleotides or polypeptides, respectively. These differences include substitutions, insertions, deletions, inversions, duplications, or any desired combinations of such changes in a native polynucleotide or amino acid sequence.
[0139] As used herein, the term “expression” refers to the biosynthesis of a gene product, and typically the transcription and / or translation of a nucleotide sequence, such as an endogenous gene, a heterologous gene, a transgene or an RNA and / or protein coding sequence, in a cell, tissue, organ, or organism, such as a plant, plant part or plant cell, tissue or organ.
[0140] The term “recombinant” in reference to a polynucleotide (DNA or RNA) molecule, protein, construct, vector, etc., refers to a polynucleotide or protein molecule or sequence that is man-made and not normally found in nature, and / or is present in a context in which it is not normally found in nature, including a polynucleotide (DNA or RNA) molecule, protein, construct, etc., comprising a combination of two or more polynucleotide or protein sequences that would not naturally occur together in the same manner without human intervention, such as a polynucleotide molecule, protein, construct, etc., comprising at least two polynucleotide or protein sequences that are operably linked but heterologous with respect to each other. For example, the term “recombinant” can refer to any combination of two or more DNA or protein sequences in the same molecule (e.g., a plasmid, construct, vector, chromosome, protein, etc.)where such a combination is man-made and not normally found in nature. As used in this definition, the phrase “not normally found in nature” means not found in nature without human introduction. A recombinant polynucleotide or protein molecule, construct, etc., can comprise polynucleotide or protein sequence(s) that is / are (i) separated from other polynucleotide or protein sequence(s) that exist in proximity to each other in nature, and / or (ii) adjacent to (or contiguous with) other polynucleotide or protein sequence(s) that are not naturally in proximity with each other.
[0141] Such a recombinant polynucleotide molecule, protein, construct, etc., can also refer to a polynucleotide or protein molecule or sequence that has been genetically engineered and / or constructed outside of a cell. For example, a recombinant DNA molecule can comprise any engineered or man-made plasmid, vector, etc., and can include a linear or circular DNA molecule. Such plasmids, vectors, etc., can contain various maintenance elements including a prokaryotic origin of replication and selectable marker, as well as one or more transgenes or expression cassettes perhaps in addition to a plant selectable marker gene, etc. The term “operably linked” refers to a functional linkage between a promoter or other regulatory element and an associated transcribable DNA sequence or coding sequence of a gene (or transgene), such that the promoter, etc., operates or functions to initiate, assist, affect, cause, and / or promote the transcription and expression of the associated transcribable DNA sequence or coding sequence, at least in certain cell(s), tissue(s), developmental stage(s), and / or condition(s).
[0142] Reference in this application to an “isolated DNA molecule” or an “isolated polynucleotide”, or an equivalent term or phrase, is intended to mean that the DNA molecule or polynucleotide is one that is present alone or in combination with other compositions, but not within its natural environment. For example, nucleic acid elements such as a coding sequence, intron sequence, untranslated leader sequence, promoter sequence, transcriptional termination sequence, and the like, that are naturally found within the DNA of the genome of an organism are not considered to be “isolated” so long as the element is within the genome of the organism and at the location within the genome in which it is naturally found. However, each of these elements, and subparts of these elements, would be “isolated” within the scope of this disclosure so long as the element is not within the genome of the organism and at the location within the genome in which it is naturally found. Similarly, a nucleotide sequence encoding a protein or any naturally occurring variant of that protein would be an isolated nucleotide sequence so long as the nucleotide sequence was not within the DNA of theorganism in which the sequence encoding the protein is naturally found. A synthetic nucleotide sequence encoding the amino acid sequence of the naturally occurring protein would be considered to be isolated for the purposes of this disclosure. For the purposes of this disclosure, any transgenic nucleotide sequence, i.e., the nucleotide sequence of the DNA inserted into the genome of the cells of a plant or bacterium, or present in an extrachromosomal vector, would be considered to be an isolated nucleotide sequence whether it is present within the plasmid or similar structure used to transform the cells, within the genome of the plant or bacterium, or present in detectable amounts in tissues, progeny, biological samples or commodity products derived from the plant or bacterium.
[0143] As commonly understood in the art, the term “promoter” can generally refer to a DNA sequence that contains an RNA polymerase binding site, transcription start site, and / or TATA box and assists or promotes the transcription and expression of an associated transcribable polynucleotide sequence and / or gene (or transgene). A promoter can be synthetically produced, varied or derived from a known or naturally occurring promoter sequence or other promoter sequence. A promoter can also include a chimeric promoter comprising a combination of two or more heterologous sequences. A promoter of the present disclosure can thus include variants or fragments of promoter sequences that are similar in composition, but not identical to, other promoter sequence(s) known or provided herein. A promoter provided herein, or variant or fragment thereof, may comprise a “minimal promoter” which provides a basal level of transcription and is comprised of a TATA box or equivalent DNA sequence for recognition and binding of the RNA polymerase II complex for initiation of transcription. A promoter can be classified according to a variety of criteria relating to the pattern of expression of an associated coding or transcribable sequence or gene (including a transgene) operably linked to the promoter, such as constitutive, developmental, tissuespecific, inducible, etc. Promoters that drive expression in all or most tissues of the plant are referred to as “constitutive” promoters. Promoters that drive expression during certain periods or stages of development are referred to as “developmental” promoters. Promoters that drive enhanced expression in certain tissues of the plant relative to other plant tissues are referred to as “tissue-enhanced” or “tissue-preferred” promoters. Thus, a “tissue-preferred” promoter causes relatively higher or preferential expression in a specific tissue(s) of the plant, but with lower levels of expression in other tissue(s) of the plant. Promoters that express within a specific tissue(s) of the plant, with little or no expression in other plant tissues, are referred to as “tissue-specific” promoters. An “inducible” promoter is a promoter that initiatestranscription in response to an environmental stimulus such as cold, drought or light, or other stimuli, such as wounding or chemical application. A promoter can also be classified in terms of its origin, such as being heterologous, homologous, chimeric, synthetic, etc.
[0144] As used herein, a “plant-expressible promoter” refers to a promoter that can initiate, assist, affect, cause, and / or promote the transcription and expression of its associated transcribable DNA sequence, coding sequence or gene in a plant cell or tissue.
[0145] The term “heterologous” in reference to a promoter or other regulatory sequence in relation to an associated polynucleotide sequence {e.g., a transcribable DNA sequence or coding sequence or gene) is a promoter or regulatory sequence that is not operably linked to such associated polynucleotide sequence in nature without human introduction - e.g., the promoter or regulatory sequence has a different origin relative to the associated polynucleotide sequence and / or the promoter or regulatory sequence is not naturally occurring in a plant species to be transformed with the promoter or regulatory sequence.
[0146] As used herein, an “endogenous gene” or an “endogenous locus” refers to a gene or locus at its natural and original chromosomal location. As used herein, the “endogenous ans gene” refers to the ans genomic locus at its original chromosomal location.
[0147] As used herein, in the context of a protein-coding gene, an “exon” refers to a segment of a DNA or RNA molecule containing information coding for a protein or polypeptide sequence.
[0148] As used herein, an “intron” of a gene refers to a segment of a DNA or RNA molecule, which does not contain information coding for a protein or polypeptide, and which is first transcribed into an RNA sequence but then spliced out from a mature RNA molecule.
[0149] As used herein, an “untranslated region (UTR)” of a gene refers to a segment of an RNA molecule or sequence e.g., a mRNA molecule) expressed from a gene (or transgene) but excluding the exon and intron sequences of the RNA molecule. An “untranslated region (UTR)” also refers to a DNA segment or sequence encoding such a UTR segment of an RNA molecule. An untranslated region can be a 5 '-UTR or a 3 '-UTR depending on whether it is located at the 5' or 3' end of a DNA or RNA molecule or sequence relative to a coding region of the DNA or RNA molecule or sequence {i.e., upstream (5') or downstream (3') of the exon and intron sequences, respectively).
[0150] As used herein, a “transcribable region” or “transcribable DNA sequence” refers to a nucleic acid sequence expressed from a gene (or transgene).
[0151] As used herein, a “transcription termination sequence” refers to a nucleic acid sequence containing a signal that triggers the release of a newly synthesized transcript RNA molecule from an RNA polymerase complex and marks the end of transcription of a gene or locus.
[0152] The terms “percent identity,” “% identity” or “percent identical” as used herein in reference to two or more nucleotide or protein sequences is calculated by (i) comparing two optimally aligned sequences (nucleotide or protein) over a window of comparison, (ii) determining the number of positions at which the identical nucleic acid base (for nucleotide sequences) or amino acid residue (for proteins) occurs in both sequences to yield the number of matched positions, (iii) dividing the number of matched positions by the total number of positions in the window of comparison, and then (iv) multiplying this quotient by 100% to yield the percent identity. If the “percent identity” is being calculated in relation to a reference sequence without a particular comparison window being specified, then the percent identity is determined by dividing the number of matched positions over the region of alignment by the total length of the reference sequence. Accordingly, for purposes of the present application, when two sequences (query and subject) are optimally aligned (with allowance for gaps in their alignment), the “percent identity” for the query sequence is equal to the number of identical positions between the two sequences divided by the total number of positions in the query sequence over its length (or a comparison window), which is then multiplied by 100%. When percentage of sequence identity is used in reference to proteins it is recognized that residue positions which are not identical often differ by conservative amino acid substitutions, where amino acid residues are substituted for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity) and therefore do not change the functional properties of the molecule. When sequences differ in conservative substitutions, the percent sequence identity can be adjusted upwards to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have “sequence similarity” or “similarity.” Sequences having a percent identity to a base sequence may exhibit the activity of the base sequence.
[0153] Degeneracy of the genetic code provides the possibility to substitute at least one base of the protein encoding sequence of a gene with a different base without causing the amino acid sequence of the polypeptide produced from the gene to be changed. When optimally aligned, homolog proteins, or their corresponding nucleotide sequences, have typically at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, atleast about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or even at least about 99.5% identity over the full length of a protein or its corresponding nucleotide sequence identified as being associated with imparting an altered phenotype when expressed in plant cells. According to embodiments of the present invention, an ans gene encodes a protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to SEQ ID NO: 12.
[0154] Homologs are inferred from sequence similarity, by comparison of protein sequences, for example, manually or by use of a computer-based tool. For optimal alignment of sequences to calculate their percent identity, various pairwise or multiple sequence alignment algorithms and programs are known in the art, such as ClustalW or Basic Local Alignment Search Tool® (BLAST), etc., that can be used to compare the sequence identity or similarity between two or more nucleotide or protein sequences. BLAST can also be used, for example to search query protein sequences of a base organism against a database of protein sequences of various organisms, to find similar sequences. The generated summary Expectation value (E-value) can be used to measure the level of sequence similarity. Because a protein hit with the lowest E-value for a particular organism may not necessarily be an ortholog or be the only ortholog, a reciprocal query is used to filter hit sequences with significant E-values for ortholog identification. The reciprocal query entails search of the significant hits against a database of protein sequences of the base organism. A hit can be identified as an ortholog, when the reciprocal query's best hit is the query protein itself or a paralog of the query protein. With the reciprocal query process orthologs are further differentiated from paralogs among all the homologs, which allows for the inference of functional equivalence of genes.
[0155] The terms “percent complementarity” or “percent complementary”, as used herein in reference to two nucleotide sequences, is similar to the concept of percent identity but refers to the percentage of nucleotides of a query sequence that optimally base-pair or hybridize to nucleotides of a subject sequence when the query and subject sequences are linearly arranged and optimally base paired without secondary folding structures, such as loops, stems or hairpins. Such a percent complementarity may be between two DNA strands, two RNA strands, or a DNA strand and an RNA strand. The “percent complementarity” is calculated by (i) optimally base-pairing or hybridizing the two nucleotide sequences in a linear and fullyextended arrangement (i.e., without folding or secondary structures) over a window of comparison, (ii) determining the number of positions that base-pair between the two sequences over the window of comparison to yield the number of complementary positions, (iii) dividing the number of complementary positions by the total number of positions in the window of comparison, and (iv) multiplying this quotient by 100% to yield the percent complementarity of the two sequences. Optimal base pairing of two sequences may be determined based on the known pairings of nucleotide bases, such as G-C, A-T, and A-U, through hydrogen bonding. If the “percent complementarity” is being calculated in relation to a reference sequence without specifying a particular comparison window, then the percent identity is determined by dividing the number of complementary positions between the two linear sequences by the total length of the reference sequence. Thus, for purposes of the present disclosure, when two sequences (query and subject) are optimally base-paired (with allowance for mismatches or non-base-paired nucleotides but without folding or secondary structures), the “percent complementarity” for the query sequence is equal to the number of base-paired positions between the two sequences divided by the total number of positions in the query sequence over its length (or by the number of positions in the query sequence over a comparison window), which is then multiplied by 100%.
[0156] As used herein, a “fragment” of a polynucleotide or polypeptide refers to a sequence comprising at least about 10, at least about 50, at least about 75, at least about 95, at least about 100, at least about 125, at least about 150, at least about 175, at least about 200, at least about 225, at least about 250, at least about 275, at least about 300, at least about 500, at least about 600, at least about 700, at least about 750, at least about 800, at least about 900, or at least about 1000 contiguous nucleotides or amino acids, or longer, of a DNA molecule or protein as disclosed herein. Methods for producing such fragments from a starting polynucleotide or polypeptide molecule are well known in the art. Fragments of a DNA molecule or protein, in some embodiments, may exhibit the activity of the DNA molecule or protein from which they are derived.
[0157] A plant selectable marker transgene in a transformation vector or construct of the present disclosure may be used to assist in the selection of transformed cells or tissue due to the presence of a selection agent, such as an antibiotic or herbicide, wherein the plant selectable marker transgene provides tolerance or resistance to the selection agent. Thus, the selection agent may bias or favor the survival, development, growth, proliferation, etc., of transformed cells expressing the plant selectable marker gene, such as to increase theproportion of transformed cells or tissues in the Ro plant. Commonly used plant selectable marker genes include, for example, those conferring tolerance or resistance to antibiotics, such as kanamycin and paromomycin {nptll), hygromycin B {aph IV), streptomycin or spectinomycin {aadA) and gentamycin {aac3 and aacC4), or those conferring tolerance or resistance to herbicides such as glufosinate bar or pat), dicamba (DMO) and glyphosateproA or EPSPS). Plant screenable marker genes may also be used, which provide an ability to visually screen for transformants, such as luciferase or green fluorescent protein (GFP), or a gene expressing a beta glucuronidase or uidA gene (GUS) for which various chromogenic substrates are known. Plant transformation may also be carried out in the absence of selection during one or more steps or stages of culturing, developing or regenerating transformed explants, tissues, plants and / or plant parts.F. Transformation Methods
[0158] Methods and compositions are provided for transforming a plant cell, tissue or explant with a recombinant DNA molecule or construct encoding one or more molecules required for targeted genome editing {e.g., guide RNA(s) and / or site-directed nuclease(s)). Suitable methods for transformation of host plant cells include virtually any method by which DNA or RNA can be introduced into a cell (for example, where a recombinant DNA construct is stably integrated into a plant chromosome or where a recombinant DNA construct or an RNA is transiently provided to a plant cell) and are well known in the art. Two effective methods for cell transformation are bacterially mediated transformation, such as Agrobacterium-m&d aXed or R / zzzofezwm-mediated transformation, and microprojectile or particle bombardment-mediated transformation. Microprojectile bombardment methods are illustrated, for example, in U.S. Patent Nos. 5,550,318; 5,538,880; 6,160,208; and 6,399,861. Agrobacterium-m&d aXed transformation methods are described, for example in U.S. Patent No. 5,591,616. Other methods for plant transformation, such as microinjection, electroporation, vacuum infiltration, pressure, sonication, silicon carbide fiber agitation, PEG- mediated transformation, etc., are also known in the art.
[0159] Transformation of plant material is practiced in tissue culture on nutrient media, for example a mixture of nutrients that allow cells to grow in vitro. Recipient cell targets include, but are not limited to, meristem cells, shoot tips, hypocotyls, calli, immature or mature embryos, and gametic cells such as microspores and pollen. Callus can be initiated from tissue sources including, but not limited to, immature or mature embryos, hypocotyls, seedling apical meristems, microspores and the like. Cells containing a transgenic nucleus are growninto transgenic plants, also referred to as Ro plants. As used herein, “Ro plant” refers to an initial regenerated transformant. As used herein, “Ri seed” refers to seed produced from selfing Ro plants. As used herein, “Ri plant” refers to a plant grown from Ri seed. As used herein, “R2 seed” refers to seed produced from selfing Ri plants. As used herein, “R2 plant” refers to a plant grown from R2 seed. Following one to two generations of self-crossing of Ro plants, plants homozygous for edited alleles of the ans edited region may be produced. Furthermore, such modified plants may be crossed with a different WT male maize plant line to produce hybrid plants.
[0160] Any suitable method or technique for transformation of a plant cell known in the art may be used according to present methods. In transformation, DNA is typically introduced into only a small percentage of target plant cells in any one transformation experiment. Marker genes are used to provide an efficient system for identification of those cells that are stably transformed by receiving and integrating a recombinant DNA molecule into their genomes.
[0161] As used herein, the terms “regeneration” and “regenerating” refer to a process of growing or developing a plant from one or more plant cells through one or more culturing steps. Transformed or edited cells, tissues or explants containing a DNA sequence insertion or edit may be grown, developed or regenerated into transgenic plants in culture, plugs, or soil according to methods known in the art. Certain embodiments of the disclosure therefore relate to methods and constructs for regenerating a plant from a cell with modified genomic DNA resulting from genome editing. The regenerated plant can then be used to propagate additional plants.
[0162] According to an aspect of the present disclosure, regenerated plants or a progeny plant, plant part or seed thereof can be screened or selected based on a marker, trait, or phenotype produced by the edit or mutation, or by the site-directed integration of an insertion sequence, transgene, etc., in the developed or regenerated plant, or a progeny plant, plant part or seed thereof. If a given mutation, edit, trait or phenotype is recessive, one or more generations or crosses e.g., selfing) from the initial Ro plant may be necessary to produce a plant homozygous for the edit or mutation so the trait or phenotype can be observed. Progeny plants, such as plants grown from Ri seed or in subsequent generations, can be tested for zygosity using any known zygosity assay, such as by using a single nucleotide polymorphism (SNP) assay, DNA sequencing, thermal amplification, or polymerase chain reaction (PCR),and / or Southern blotting that allows for the distinction between heterozygote, homozygote and wild-type plants.
[0163] Methods and techniques are provided for screening for, and / or identifying, cells or plants, etc., for the presence of targeted edits or transgenes, and selecting cells or plants comprising targeted edits or transgenes, which may be based on one or more phenotypes or traits, or on the presence or absence of a molecular marker or polynucleotide or protein sequence in the cells or plants. As used herein, a “molecular technique” refers to any method known in the fields of molecular biology, biochemistry, genetics, plant biology, or biophysics that involves the use, manipulation, or analysis of a nucleic acid, a protein, or a lipid. Without being limiting, molecular techniques useful for detecting the presence of a modified sequence in a genome include phenotypic screening; molecular marker technologies such as SNP analysis by TaqMan® or Hlumina / Infinium technology; Southern blot; PCR (including amplicon sequencing which consists of the generation of one or more unique PCR products across the genomic region of interest for further sequencing analysis, e.g., using Next-Gen Sequencing techniques known in the art. Sequence data from each sample is then mapped to a reference sequence to identify consensus differences); enzyme-linked immunosorbent assay (ELISA); and sequencing e.g., Sanger, Illumina®, 454, Pac-Bio, Ion Torrent™). In one aspect, a method of detection provided herein comprises phenotypic screening. In another aspect, a method of detection provided herein comprises SNP analysis. In a further aspect, a method of detection provided herein comprises a Southern blot. In a further aspect, a method of detection provided herein comprises PCR. In a further aspect, a method of detection provided herein comprises amplicon sequencing. In an aspect, a method of detection provided herein comprises ELISA. In a further aspect, a method of detection provided herein comprises determining the sequence of a nucleic acid or a protein. Without being limiting, nucleic acids can be detected using hybridization. Hybridization between nucleic acids is discussed in detail in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).
[0164] Nucleic acids can be isolated using techniques routine in the art. For example, nucleic acids can be isolated using any method including, without limitation, recombinant nucleic acid technology, and / or PCR. General PCR techniques are described, for example in PCR Primer: A Laboratory Manual, Dieffenbach & Dveksler, Eds., Cold Spring Harbor Laboratory Press, 1995. Recombinant nucleic acid techniques include, for example, restriction enzyme digestion and ligation, which can be used to isolate a nucleic acid. Isolatednucleic acids also can be chemically synthesized, either as a single nucleic acid molecule or as a series of oligonucleotides.
[0165] Detection e.g., of an amplification product, of a hybridization complex, of a polypeptide) can be accomplished using detectable labels that may be attached or associated with a hybridization probe or antibody. The term “label” is intended to encompass the use of direct labels as well as indirect labels. Detectable labels include enzymes, prosthetic groups, fluorescent materials, luminescent materials, bioluminescent materials, and radioactive materials. The screening and selection of modified e.g., edited) plants or plant cells can be through any methodologies known to those skilled in the art of molecular biology. Examples of screening and selection methodologies include, but are not limited to, Southern analysis, PCR amplification for detection of a polynucleotide (including amplicon sequencing), Northern blots, RNase protection, primer-extension, RT-PCR amplification for detecting RNA transcripts, Sanger sequencing, Next Generation sequencing technologies e.g., Illumina®, PacBio®, Ion Torrent™, etc.) enzymatic assays for detecting enzyme or ribozyme activity of polypeptides and polynucleotides, and protein gel electrophoresis, Western blots, immunoprecipitation, and enzyme-linked immunoassays to detect polypeptides. Other techniques such as in situ hybridization, enzyme staining, and immunostaining also can be used to detect the presence or expression of polypeptides and / or polynucleotides. Methods for performing all of the referenced techniques are known in the art.
[0166] Polypeptides can be purified from natural sources e.g., a biological sample) by known methods such as DEAE ion exchange, gel filtration, and hydroxyapatite chromatography. A polypeptide also can be purified, for example, by expressing a nucleic acid in an expression vector. In addition, a purified polypeptide can be obtained by chemical synthesis. The extent of purity of a polypeptide can be measured using any appropriate method, e.g., column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis. Polypeptides can be detected using antibodies. Techniques for detecting polypeptides using antibodies include enzyme linked immunosorbent assays (ELISAs), Western blots, immunoprecipitations and immunofluorescence. An antibody provided herein can be a polyclonal antibody or a monoclonal antibody. An antibody having specific binding affinity for a polypeptide provided herein can be generated using methods well known in the art. An antibody provided herein can be attached to a solid support such as a microtiter plate using methods known in the art.G. Genetically Modified Plants
[0167] As used herein, “modified” in the context of a maize plant, maize plant seed, maize plant part, maize plant cell, and / or maize plant genome, refers to a maize plant, plant seed, plant part, plant cell, and / or plant genome comprising an engineered change in the expression level and / or endogenous sequence of one or more genes of interest relative to a wild-type or control maize plant, plant seed, plant part, plant cell, and / or plant genome. Indeed, the term “modified” may further refer to a maize plant, plant seed, plant part, plant cell, and / or plant genome having one or more deletions and / or one or more nucleotide substitutions or nucleotide insertions affecting an endogenous ans gene introduced through chemical mutagenesis, transposon insertion or excision, or any other known mutagenesis technique, or introduced through genome editing. In an aspect, a modified plant, plant seed, plant part, plant cell, and / or plant genome can comprise one or more transgenes. In particular embodiments, a modified plant, seed, plant part, plant cell, and / or plant genome may comprise a transgene encoding one or more polypeptides having a sequence with at least about 70% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO: 8, SEQ ID NO: 10, or SEQ ID NO: 19. For clarity, therefore, a modified maize plant, plant seed, plant part, plant cell, and / or plant genome includes but is not limited to a mutated, edited, and / or transgenic maize plant, plant seed, plant part, plant cell, and / or plant genome having a modified sequence of an ans gene relative to a wild-type or control plant, plant seed, plant part, plant cell, and / or plant genome. Furthermore, the modification may reduce, disrupt, or alter the activity of the protein encoded by the ans gene as compared to the activity of the protein encoded by the ans gene in an otherwise identical maize plant.
[0168] Modified maize plants, plant parts, seeds, etc., may have been subjected to mutagenesis, genome editing or site-directed integration, genetic transformation, or a combination thereof. Such “modified” maize plants, plant seeds, plant parts, and plant cells include plants, plant seeds, plant parts, and plant cells that are offspring or derived from “modified” maize plants, plant seeds, plant parts, and plant cells that retain the molecular change (e.g., change in expression level and / or activity) to the ans gene and / or retain a polynucleotide molecule having at least about 70% sequence identity to SEQ ID NOs: 1, 3, 5, 7, 9, 11, 14, 16, or 18 or fragments thereof. A modified seed provided herein may give rise to a modified plant provided herein. A modified plant, plant seed, plant part, plant cell, or plant genome provided herein may comprise a recombinant DNA construct or vector or genome edit as provided herein.
[0169] A “modified plant product” may be any product made from a modified plant, plant part, plant cell, or plant chromosome provided herein, or any portion or component thereof. For example, in some embodiments a modified plant product may be a commodity product produced from a modified plant or part thereof containing the recombinant DNA molecule as described herein, such as those having at least about 70% sequence identity to SEQ ID NOs: 1, 3, 5, 7, 9, 11, 14, 16, or 18. In some embodiments, commodity products contain a detectable amount of DNA comprising a DNA sequence selected from the group consisting of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 14, 16, or 18 or fragments or variants thereof. As used herein, a “commodity product” refers to any composition or product which is comprised of material derived from a modified plant, seed, plant cell, or plant part containing the DNA molecule as described herein. Commodity products include but are not limited to processed seeds, grains, plant parts, and meal, protein concentrate, protein isolate, grain, starch, flour, biomass, or seed oil. A commodity product containing a detectable amount of DNA corresponding to the recombinant DNA molecule as described herein is contemplated. Detection of one or more of this DNA in a sample may be used for determining the content or the source of the commodity product. Any standard method of detection for DNA molecules may be used, including methods of detection disclosed herein.
[0170] The plants of the present disclosure may be further crossed to themselves or other plants to produce plant seeds and progeny. In one embodiment, a first plant that comprises a mutant allele or a genomic edit that that downregulates anthocyanidin synthase ans) gene function may be crossed with a second plant comprising a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19 to produce plant seeds and progeny. In another embodiment, the first plant comprises the mutant allele that downregulates anthocyanidin synthase (ans) gene function. A non-limiting example of such an allele is the ansla.2 maize mutant allele. In still yet another embodiment, the polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19, is expressed from an endogenous gene that encodes the polypeptide. Non-limiting examples of such endogenous genes include the Pll(ZmPll ) and Bl(ZmBl ) genes. A modified plant may also be prepared by crossing a first plant comprising a DNA sequence or construct or an edit (e.g., a genomic deletion) with a second plant lacking the DNA sequence or construct or edit. For example, a DNA sequence or inversion may be introduced into a first plant line that is amenable totransformation or editing, which may then be crossed with a second plant line to introgress the DNA sequence or edit (e.g., deletion) into the second plant line. Progeny of these crosses can be further backcrossed into the desirable line multiple times, such as through 6 to 8 generations or back crosses, to produce a progeny plant with substantially the same genotype as the original parental line, but for the introduction of the DNA sequence or edit. A modified plant, plant cell, or seed provided herein may be a hybrid plant, plant cell, or seed. As used herein, a “hybrid” is created by crossing two plants from different varieties, lines, inbreds, or species, such that the progeny comprises genetic material from each parent. Skilled artisans recognize that higher order hybrids can be generated as well.
[0171] A modified maize plant, plant part, plant cell, or seed provided herein may be of an elite variety or an elite line. An “elite variety” or an “elite line” refers to a variety that has resulted from breeding and selection for superior agronomic performance.
[0172] As used herein, the term “control plant” (or likewise a “control” plant seed, plant part, plant cell, and / or plant genome) refers to a maize plant (or plant seed, plant part, plant cell, and / or plant genome) that is used for comparison to a modified plant (or modified plant seed, plant part, plant cell, and / or plant genome) and has the same or similar genetic background (e.g., same parental tines, hybrid cross, inbred line, testers, etc.) as the modified plant (or plant seed, plant part, plant cell, and / or plant genome), except for genome edit(s) e.g., a deletion) affecting an ans gene and / or transgenes comprising a nucleotide sequence having at least about 70% sequence identity to SEQ ID NO: 1, 3, 5, 7, and / or 9. For example, a control plant may be an inbred line that is the same as the inbred line used to make the modified maize plant, or a control plant may be the product of the same hybrid cross of inbred parental lines as the modified plant, except for the absence in the control plant of any transgenic events or genome edit(s) affecting an ans gene and / or any transgenes comprising a nucleotide sequence having at least about 70% sequence identity to SEQ ID NO: 1, 3, 5, 7, and / or 9. Similarly, an “unmodified control plant” refers to a plant that shares a substantially similar or essentially identical genetic background as a modified plant, but without the one or more engineered changes to the genome (e.g., mutation or edit) of the modified plant. For purposes of comparison to a modified plant, plant seed, plant part, plant cell, and / or plant genome, a “wild-type plant” (or likewise a “wild-type” plant seed, plant part, plant cell, and / or plant genome) refers to a non-transgenic and non-genome edited control plant, plant seed, plant part, plant cell, and / or plant genome. As used herein, a “control” plant, plant seed, plant part, plant cell, and / or plant genome may also be a plant, plant seed, plant part, plant cell,and / or plant genome having a similar (but not the same or identical) genetic background to a modified plant, plant seed, plant part, plant cell, and / or plant genome, if deemed sufficiently similar for comparison of the characteristics or traits to be analyzed.
[0173] As used herein, the term “activity” refers to the biological function of a gene or protein. A gene or a protein may provide one or more distinct functions. A reduction, disruption, or alteration in “activity” thus refers to a lowering, reduction, or elimination of one or more functions of a gene or a protein in a maize plant, plant cell, or plant tissue at one or more stage(s) of plant development, as compared to the activity of the gene or protein in a wild-type or control plant, cell, or tissue at the same stage(s) of plant development. Additionally, an increase in “activity” thus refers to an elevation of one or more functions of a gene or a protein in a maize plant, plant cell, or plant tissue at one or more stage(s) of plant development, as compared to the activity of the gene or protein in a wild-type or control plant, cell, or tissue at the same stage(s) of plant development.
[0174] According to some embodiments, a modified maize plant is provided having a genomic modification in an ans gene that results in downregulation of ans gene function, or reduced, disrupted, or altered activity of the protein encoded by the ans gene in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant. According to further embodiments, a modified maize plant is provided having a protein encoded by an ans gene that results in reduced, disrupted, or altered activity in at least one plant tissue by 5%-20%, 5%-25%, 5%-30%, 5%-40%, 5%-50%, 5%- 60%, 5%-70%, 5%-75%, 5%-80%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%-90%, 50%-75%, 25%- 75%, 30%-80%, or 10%-75%, as compared to a control plant.
[0175] According to some embodiments, a modified plant is provided having an ans, ZmPLl, SbTT2, ZmR2, and / or ZmRl mRNA level that is reduced or increased in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant. According to some embodiments, a modified plant is provided having an ans, ZmPLl, SbTT2, ZmR2, and / or ZmRl mRNA expression level that is reduced or increased in at least one plant tissue by 5%-20%, 5%-25%, 5%- 30%, 5%-40%, 5%-50%, 5%-60%, 5%-70%, 5%-75%, 5%-80%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%- 90%, 50%-75%, 25%-75%, 30%-80%, or 10%-75%, as compared to a control plant. According to some embodiments, a modified plant is provided having an ANS, ZmPLl,SbTT2, ZmR2, and / or ZmRl protein expression level that is reduced or increased in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant. According to some embodiments, a modified plant is provided having an ANS, ZmPLl, SbTT2, ZmR2, and / or ZmRl protein expression level that is reduced or increased in at least one plant tissue by 5%-20%, 5%- 25%, 5%-30%, 5%-40%, 5%-50%, 5%-60%, 5%-70%, 5%-75%, 5%-80%, 5%-90%, 5%-100%, 75%-100%, 50%- 100%, 50%-90%, 50%-75%, 25%-75%, 30%-80%, or 10%-75%, as compared to a control plant.
[0176] The present disclosure relates to a plant with improved economically important characteristics, including but not limited to increased production of proanthocyanidins or condensed tannins as compared to a control plant. Modified plants comprising or derived from plant cells that comprise a genome modification of this disclosure can be further enhanced with stacked traits, for example, a modified crop plant having an enhanced trait resulting from expression of DNA disclosed herein in combination with one or more additional genome modifications that provide a beneficial agronomic trait or further improve the enhanced trait.
[0177] As used herein, a “plant” includes a whole plant, explant, plant part, seedling, or plantlet at any stage of regeneration or development.
[0178] As used herein, a “plant part” can refer to any organ or intact tissue of a plant, such as a meristem, shoot organ / structure (e.g., leaf, stem or node), root, flower or floral organ / structure (e.g., bract, sepal, petal, stamen, carpel, anther and ovule), seed, embryo, endosperm, seed coat, fruit, the mature ovary, propagule, or other plant tissues e.g., vascular tissue, dermal tissue, ground tissue, and the like), or any portion thereof. Plant parts of the present disclosure can be viable, nonviable, regenerable, and / or non-regenerable. A “propagule” can include any plant part that can grow into an entire plant.
[0179] An “embryo” is a part of a plant seed, consisting of precursor tissues (e.g., meristematic tissue) that can develop into all or part of an adult plant. An “embryo” may further include a portion of a plant embryo.
[0180] A “meristem” or “meristematic tissue” comprises undifferentiated cells or meristematic cells, which are able to differentiate to produce one or more types of plant parts, tissues or structures, such as all or part of a shoot, stem, root, leaf, seed, etc.
[0181] The term "about" is used to indicate that a value includes the standard deviation of the mean for the device or method being employed to determine the value. The use of the term "or" in the claims is used to mean "and / or" unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive. When used in conjunction with the word "comprising" or other open language in the claims, the words "a" and "an" denote "one or more," unless specifically noted otherwise. The terms "comprise," "have," and "include" are open-ended linking verbs. Any forms or tenses of one or more of these verbs, such as "comprises," "comprising," "has," "having," "includes," and "including," are also open-ended. For example, any method that "comprises," "has," or "includes" one or more steps is not limited to possessing only those one or more steps and also covers other unlisted steps. Similarly, any system or method that "comprises," "has," or "includes" one or more components is not limited to possessing only those components and covers other unlisted components.
[0182] Other objects, features, and advantages of the present disclosure are apparent from detailed description provided herein. It should be understood, however, that the detailed description and any specific examples provided, while indicating specific embodiments of the disclosure, are given by way of illustration only, since various changes and modifications within the spirit and scope of the disclosure will become apparent to those skilled in the art from this detailed description. Any embodiment of the present disclosure may be used in combination with any other embodiment described herein.
[0183] All references herein are incorporated herein by reference in their entirety.EXAMPLES
[0184] The following examples are included to illustrate embodiments of the present disclosure. It should be appreciated by those of skill in the art that the techniques disclosed in the examples that follow represent techniques discovered by the inventor to function well in the practice of the invention. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments which are disclosed and still obtain a like or similar result without departing from the concept, spirit and scope of the invention. More specifically, it will be apparent that certain agents which are both chemically and physiologically related may be substituted for the agents described herein while the same or similar results would be achieved. All such similar substitutes andmodifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims.Example 1: Mining PA-related Genes in Monocotyledonous Plants.
[0185] To interrogate the potential for PA biosynthesis in maize, a search for homologs of genes known to be involved in PA biosynthesis in other plant species was conducted. Anthocyanidin Reductase (ANR) encodes an enzyme functionally conserved across species that is essential for PA biosynthesis. Using GmANRl to BLAST against the maize genome, two strong candidate ANR genes, namely ZmANRl (GRMZM2g097854) and ZmANR2 (GRMZM2gO97841) were identified. These two genes locate in close proximity to each other on chromosome 10. Transcriptional analysis of these two ANR homologs in colored maize seeds of varieties Suntava (ST, purple seeds) and Black Mexican (BM, blue / black seeds) indicated that ZmANRl is expressed at a much higher level than ZmANR2 in developing seeds (14 days-after-pollination, DAP) (FIG. 1A). Subcellular localization experiments using Arabidopsis protoplasts showed that ZmANRl localizes to the cytosol, similar to soybean GmANRl (FIG. IB). Ectopic expression of ZmANRl in the A. thaliana ban mutant led to increased levels of free epicatechin but no procyanidin B2 dimers [(-)-epicatechin-(-)- epicatechin] (23). To determine whether ZmANR differs from dicot ANRs in its ability to direct PA formation in planta, ZmANRl, A. thaliana ANR (AtANR), and GmANRl were expressed in tobacco (Nicotiana tabacum cv Xanthi) overexpressing the A. thaliana PAP1 (PRODUCTION OF ANTHOCYANIN PIGMENT 1) gene. PAP1 is a conserved MYB transcription factor regulating anthocyanin biosynthesis in plants. These PAP1-OX tobacco plants produce high levels of anthocyanins, resulting in red-colored petals, and overexpression of AtANR or GmANRl in the PAP 1 -OX background diverted flavonoid precursors from anthocyanins toward PA biosynthesis and consequently loss of the red floral pigmentation (FIG. 1C). However, PAPl-OX / ZmANRl lines displayed red-colored petals similar to PAP1-OX (FIG. 1C). Staining with p-dimethylaminocinnamaldehyde (DMACA) showed that PAs accumulated in petals of PAPl-OX / AtANR and PAPl-OX / GmANRl plants, as indicated by the purple color, but not in flowers of PAPl-OX / ZmANRl or PAP1-OX plants. Further analyses of PAs using liquid chromatography-mass spectrometry (LC-MS) showed that both epicatechin and procyanidin B2 dimers [(-)-epicatechin-(-)-epicatechin] accumulated in PAPl-OX / AtANR and PAPl-OX / GmANRl flowers, whereas only a small increase of epicatechin but not procyanidin B2 was observed in PAPl-OX / ZmANRl (FIG. ID). Furthermore, cysteinyl-epicatechin, an extension unit for PA polymerization, wassubstantially accumulated in PAPl-OX / AtANR and PAPl-OX / GmANRl lines but was only slightly elevated in the PAPl-OX / ZmANRl line with similar level of transgene expression (FIG. ID, FIG. IE).
[0186] Leucoanthocyanidin reductase (LAR) generates catechin from leucocyanidin and also participates in PA polymerization in some species by controlling the ratio of PA starter (epicatechin) and extension (cysteinyl-epicatechin) units. Searching the maize and sorghum genomes using MtLAR from M. truncatula and VvLARl from grapevine (Vitis vinifera) returned no apparent LAR homolog, with an annotated 2 '-hydroxyisoflavone reductase (Zm00001d040173) and an uncharacterized protein (Sb03g043200) showing the highest similarity in maize and sorghum, respectively. At least one OsLAR (LOC_Os03g 15360) was present in the rice genome. Further sequence analysis showed that Sb03g043200 and LOC_Os03g 15360 lack the ICCNSIA and THDIFI domains conserved in functional LAR homologs and share no obvious sequence similarity with LAR homologs in other plant species. Thus, sorghum and maize appear not to contain apparent LAR homologs.
[0187] TT8 (bHLH) and TTG1 (WD40) transcription factors are required for expression of PA biosynthesis in A. thaliana and M. truncatula. However, no apparent TT2-type MYB has yet been identified in maize. The maize MYB transcription factors with the highest sequence similarity to TT2 homologs in A. thaliana (AtTT2) and M. truncatula (MtMYB14) are Cl and Pl, both of which have been identified as anthocyanin regulators. Further sequence analysis showed that both rice and sorghum have both Cl- and TT2-type MYBs, i.e., OsCl (LOC_Os06gl0350), OsMYB3 / OsTT2 (LOC_Os03g29614), SbCl (Sbl0g006700), and SbTT2 (Sb01g032770). Phylogenetic analysis indicated that the TT2 homologs from sorghum, rice, M. truncatula and A. thaliana belong to a clade distinct from Cl-type MYBs. MYB5-type transcription factors also play important roles in plant development and PA biosynthesis in dicot species, and homologous genes can be identified in the maize, rice, and sorghum genomes. Taken together, PA-rich monocots, including sorghum and rice, possess Cl-type, TT2-type and MYB5-type MYBs, whereas maize only appears to contain the Cl- and MYB5-type.
[0188] Overall, these results indicate that maize possesses neither an effective ANR or LAR that facilitate PA biosynthesis nor a TT2-family transcription factor to regulate expression of their genes. Lu et al., Nature Communications, 14:4349, 2023, is incorporated herein by reference in its entirety.Example 2: Occurrence and Diversity of PA Precursors in Different Maize Varieties.
[0189] Given the lack of information regarding the presence or levels of flavan-3-ols and other PA-related metabolites in maize, a search for these metabolites was conducted in seeds of various maize varieties. In yellow maize seeds (FBLL), only trace amounts of epicatechin were detected (FIG. 2), whereas in brown (Osage, OSA), black (Black Mexican, BM), and purple (Suntava, ST; Arequipa, AREQ) seeds, larger amounts of catechin and epicatechin were detected, with more epicatechin than catechin (FIG. 2), consistent with the demonstration that ZmANRl preferentially produces epicatechin over catechin. The PA extension units cysteinyl-(epi)catechin were detected in brown (Osage, OSA), black (Black Mexican, BM), and purple (Suntava, ST) seeds, with the level of 4 / ?-(S-cysteinyl)-catechin generally higher than that of 4 / ?-(S-cysteinyl)-epicatechin, or, in the case of AREQ, exclusively 4 / ?-(S-cysteinyl)-catechin being present (FIG. 2). The levels of the anthocyanin cyanidin 3-O-glucoside were generally negatively correlated with free epicatechin levels in OSA, BM, ST and AREQ seeds (FIG. 2). In maize seeds naturally accumulating phlobaphenes, no (epi)catechin or cysteinyl-(epi)catechin were detected. In sorghum seeds included as a positive control, catechin and procyanidin B3 (catechin dimer) were observed.
[0190] To compare the PA profile of maize with that from a typical PA-rich crop, the soluble PA fraction extracted from seeds of soybean was analyzed, in which epicatechin was the primary flavan-3-ol monomer, and B2 [(-)-epicatechin-(-)-epicatechin] was the main procyanidin dimer. In contrast to maize, the PA extension unit 4 / ?-(S-cysteinyl)-epicatechin was detected in developing seeds but not in mature seeds.Example 3: Expression of GmANRl Leads to Accumulation of 4 / ?-(S-cysteinyl)- epicatechin in Maize Seeds with Active Flavonoid Biosynthesis.
[0191] ANR participates in the generation of both starter and extension units in PA-rich plants. To test whether the ANR from a PA-rich species can boost PA production in maize, soybean GmANRl was introduced into Hill maize that only produces trace amounts of epicatechin and no cysteinyl-epicatechin extension units. Hill was selected for transformation because of its high transformation efficiency and the short time needed for plant regeneration. Three independent GmANRl -OX lines with confirmed GmANRl expression were selected for further biochemical analyses. The contents of anthocyanins and soluble PAs were low and not significantly different between seeds of untransformed plants and the three GmANRl - OX lines. Epicatechin levels were similar between untransformed and GmANRl -OX seeds, and no 4 / ?-(S-cysteinyl)-catechin or 4 / ?-(S-cysteinyl)-epicatechin were detected. This showsthat the PA biosynthesis pathway may be substrate-limited since the general flavonoid pathway is inactive in white- or yellow-colored seeds such as Hill. To provide more substrate for PA biosynthesis, GmANRl-OX (line GmANRl-24) was crossed with Suntava (ST), an anthocyanin-rich purple maize, and the Fi plants were backcrossed for two more generations using ST as the recurrent parent to obtain BC2 seeds. The content of 4 / ?-(S-cysteinyl)- epicatechin in ST-GmANRl BC2 transgenic seeds was increased by up to 19-fold compared with that in untransformed ST seeds. In addition, the ratio of 4 / ?-(S-cysteinyl)-epicatechin to 4 / ?-(S-cysteinyl)-catechin was increased from 0.3 in untransformed seeds to 8 in ST- GmANRl seeds (FIG. 3A). These data are consistent with previous studies indicating that GmANRl is involved in the formation of 4 / ?-(S-cysteinyl)-epicatechin and that ZmANRl may not be able to efficiently catalyze the conversion of flaven-3 ,4-diol to 2,3-czs-leucocyandin and 4 / ?-(S-cysteinyl)-epicatechin. The lower expression level of endogenous ZmANRl compared to ectopically expressed GmANRl may partially contribute to different PA levels in different maize lines. Further phloroglucinolysis of PA polymers showed that ST- GmANRl had higher levels of released (epi)catechin-phloroglucinol (representing PA extension units from polymers) than either parental line (FIG. 3B).
[0192] To confirm that the increased 4 / ?-(S-cysteinyl)-epicatechin in ST-GmANRl is independent of the genetic background that introduces an active flavonoid pathway, GmANRl-24 (Hill) was crossed with another anthocyanin-rich maize, Osage (OSA), and the Fi generation was backcrossed with OSA as recurrent parent to obtain BCi seeds. Similar to the effect of GmANRl-OX in the ST background, the level of 4 / ?-(S-cysteinyl)-epicatechin and the ratio of 4 / ?-(S-cysteinyl)-epicatechin to 4 / ?-(S-cysteinyl)-catechin were increased by approximately 6-fold and 9-fold, respectively, in OSA-GmANRl BCi seeds compared with untransformed OSA seeds (FIG. 3C), indicating that the altered composition of soluble PA precursors was caused by GmANRl expression. These results highlight the endogenous ANR as a bottleneck for PA biosynthesis in maize due to its inability to facilitate formation of extension units.Example 4: Expressing SbTT2 or SbMYB5 in Yellow-Seed Maize Activates Both the Anthocyanin and PA Precursor Pathways.
[0193] Since maize does not have an apparent homolog of TT2, this shows that, in addition to the issue of ZmANR specificity, maize may lack a functional transcription factor to activate the PA biosynthesis pathway. Sorghum is a PA-rich monocot and a close-relative to maize, so it was investigated whether SbTT2 and SbMYB5 could be used to activate PAaccumulation in maize. Transactivation assays using Arabidopsis protoplasts indicated that both SbTT2 and SbMYB5, when combined with GmTT8 and GmWD40, were able to activate the GrnANRl promoter thereby driving firefly luciferase gene expression, showing that SbTT2 and SbMYB5 may function in a manner similar to that of GmTT2 and GmMYB5 in soybean and related MYBs in other species.
[0194] To first test the in vivo activity of the transcription factors, SbTT2, SbMYB5 and GUS (control) were individually expressed in soybean hairy roots. DMACA-stained PAs (purple color) were observed in vascular tissues of SbTT2 roots but were undetectable in SbMYB5 or GUS expressing roots. Quantitative reverse transcription (qRT)-PCR showed that transcript levels of GmANRl, GmLAR2 and GmANS were significantly higher in SbTT2 and SbMYB5 tines than in GUS lines, with GmANRl primarily induced by SbTT2 and GmLAR2 more effectively induced by SbMYB5. Phloroglucinolysis analysis showed increased levels of oligomer / polymer-derived epicatechin-phloroglucinol in SbTT2-expressing compared with the GUS control. These results show that, while both SbTT2 and SbMYB5 can induce the expression of PA-related genes, SbTT2 is the more effective in activating the PA biosynthesis / polymerization pathway, at least in soybean.
[0195] SbTT2 and SbMYB5 were then introduced individually into Bl 04 maize, an inbred deficient in Cl and B-peru activities, transcription factors regulating anthocyanin biosynthesis. Two independent lines of SbTT2-OX (SbTT2-21 and SbTT2-31) and two independent lines of SbMYB5-OX (SbMYB5-51 and SbMYB5-12) were selected for further analyses (FIG. 4A). Seeds from both SbTT2-OX and SbMYB5-OX lines were smaller than untransformed seeds (FIG. 4A) and weighed less (FIG. 4B). In contrast to the yellow seed color of the untransformed maize (B 104), seeds of the two SbTT2-OX tines were purple, and seeds from SbMYB5-OX tines were slightly darker than untransformed seeds (FIG. 4A). Anthocyanin-related pigmentation was also observed in the coleoptile tissue and later in silks (FIG. 4A [lower panels] and Fig. 4C). Qualitative analysis of anthocyanins using high performance LC (HPLC) and LC-MS indicated that the anthocyanins produced in SbTT2-OX and SbMYB5-OX are predominantly cyanidin-based and include cyanidin 3-O-glucoside. Quantification of total anthocyanins showed that expression of SbTT2 or SbMYB5 induced anthocyanin biosynthesis in maize seeds, and the level of anthocyanin was higher in SbTT2- OX than in SbMYB5-OX.
[0196] Transcript analysis showed that expression of ZmANRl and ZmANR2 was up- regulated in developing seeds of SbTT2-OX and SbMYB5-OX, with ZmANRl transcripts toa much higher level than ZmANR2 (FIG. 5A). Both soluble and insoluble PA levels were increased, by approximately 10-15-fold and 4-8-fold, in seeds of SbTT2-OX and SbMYB5- OX, respectively, compared to the untransformed maize (FIG. 5B), although the absolute levels were still low. Analysis of PA precursors using LC-MS indicated that, compared with untransformed Bl 04 seeds, where these precursors are barely detectable, the levels of epicatechin, 4 / ?-(S-cysteinyl)-catechin, and 4 / ?-(S-cysteinyl)-epicatechin were increased by over 100-fold in SbTT2-OX and to a lesser extent in SbMYB5-OX (FIG. 5C). In contrast to the ST-GmANRl and OSA-GmANRl lines, where 4 / ?-(S-cysteinyl)-epicatechin content is higher than 4 / ?-(S-cysteinyl)-catechin, SbTT2-OX and untransformed maize (BM, ST, and OSA) contain about twice as much 4 / ?-(S-cysteinyl)-catechin as 4 / ?-(S-cysteinyl)-epicatechin (FIG. 3A; FIG. 3C, FIG. 5C), which is likely due to the endogenous ZmANR not being able to efficiently produce 2,3-czs-leucocyanidin. Phloroglucinolysis assays showed that the amount of (epi)catechin-phloroglucinol produced was higher in SbTT2-OX lines than in untransformed plants (FIG. 5D), indicating increased accumulation of PA oligomers and polymers in SbTT2-OX seeds. This demonstrates that, due to the inefficiency of ZmANR excess substrates are re-directed toward anthocyanin biosynthesis, resulting in hyperaccumulation of anthocyanins.Example 5: Maize Produces Procyanidin Dimers with Unusual Stereochemistry.
[0197] In contrast to most ANRs, ZmANR 1 converts cyanidin to (+)-epicatechin in vitro, which can subsequently form procyanidin dimers with 2,3-czs-leucocyanidin or 4 ?-(S- cysteinyl)-epicatechin (23). Unlike (+)-epicatechin and (-)-epicatechin monomers that cannot be separated on LC-MS, the iso-B2 dimer ((-)-epicatechin-(+)-epicatechin] and iso-B4 dimer [(+)-catechin-(+)-epicatechin] have different retention times than B2 [(-)-epicatechin-(-)- epicatechin] and B4 [(+)-catechin-(-)-epicatechin]. To test whether maize seeds contain different stereoisomers of the common procyanidin dimers, standards for iso-B2 and iso-B4 dimers were chemically synthesized, and procyanidin dimers in polyphenol extracts from maize seeds were analyzed by LC-MS (see Methods). In ST maize seeds, iso-B2 and iso-B4 were detected as the major procyanidin dimers, both of which use (+)-epicatechin as starter unit and have not been detected in any other plant species (FIG. 6A). These dimers could also be detected in seeds of BM and to a lesser extent AREQ, which contains less free epicatechin than BM and ST. The level of iso-B4 tended to be higher than iso-B2 in these seeds, consistent with maize seeds preferentially containing more cysteinyl-catechin than cysteinyl-epicatechin (FIG. 2).
[0198] In seeds of ST-GmANRl and OSA-GmANRl, in addition to iso-B2 and iso-B4 dimers, procyanidin B2 and B4 dimers were also detected (FIG. 6A and FIG. 6B), which use (-)-epicatechin as starter unit and are commonly found in soybean and other PA-rich species, supporting the idea that ANR determines the C2-C3 stereochemistry of procyanidin dimers. While the yellow-colored seeds of B104 do not contain detectable procyanidin dimers, overexpression of SbTT2 in Bl 04 induced production of iso-B2 and iso-B4 in seeds (FIG. 6C), similar to untransformed BM and ST varieties. Furthermore, procyanidin trimer Cl (epicatechin-epicatechin-epicatechin) was also detected in seeds of ST-GmANRl using Cl trimers extracted from soybean and M. truncatula seeds as reference standards (FIG. 6D).Example 6: Crosstalk Between PA and Other Flavonoid Pathways in Maize.
[0199] Although ectopic expression of SbTT2 or SbMYB5 induced production of PAs in maize, their content remained low compared to that in sorghum and other PA-rich plants, showing that maize may lack a mechanism to efficiently divert flux toward PA biosynthesis. Therefore, it was speculated that upstream substrates may be routed to the synthesis of other flavonoid compounds.
[0200] Flavanol-anthocyanin conjugates (flavan-3-ol-anthocyanin) are dimeric flavonoids composed of a flavan-3-ol moiety and an anthocyanin moiety, with anthocyanin serving as the starter unit and flavan-3-ol as the extension unit. These compounds have been characterized in red wine, Guatemalan bean, blackcurrant and strawberry (39-44). Over the past decade, several studies have also confirmed the existence of flavanol-anthocyanin dimers (i.e., catechin-cyanidin-3,5-O-diglucoside) in some purple maize seeds (21, 45-47). Due to the lack of commercial standards, catechin-cyanidin-3,5-O-diglucoside (Ca-C35G, m / z = 897.2127 ± 10 ppm) extracted from seeds of AREQ maize was used as a reference, as identified in a Paulsmeyer et al. J. Agric. Food Chem. 65, 4341-4350, 2017. In this way, catechin-cyanidin- 3, 5-O-di glucoside was detected in purple (ST) but not in yellow (Bl 04) or blue (BM) maize seeds. Among the samples, AREQ had the highest levels of catechin-cyanidin-3,5-0- diglucoside and cysteinyl-catechin, the proposed extension unit for the conjugates (FIG. 2). In addition, overexpression of SbTT2 or SbMYB5 in maize with yellow seeds (Bl 04) induced production of catechin-cyanidin-3,5-O-diglucoside in the seeds. Interestingly, even though anthocyanins and PAs were present at higher levels in SbTT2-OX than in SbMYB5-OX, SbMYB5-OX produced more catechin-cyanidin-3,5-O-diglucoside than SbTT2-OX, likely due to the increased flux of flavonoid substrates.
[0201] To further explore the relationship between anthocyanin and PA biosynthesis, PAs were analyzed in a classic maize mutant line bzl, where anthocyanin biosynthesis and deposition are blocked due to the deficiency of a UDP-glucose: flavonoid 3-0- glucosyltransferase (UFGT), a key enzyme in anthocyanidin glycosylation. The bzl seeds had much higher levels of epicatechin but dramatically lower levels of cyanidin 3-O-glucoside than BZ1 seeds (FIG. 7A), likely because ANR and UFGT compete for common substrates and the disruption of UFGT results in substrates being redirected to PA biosynthesis. In fact, the accumulation of anthocyanins was almost abolished in bzl seeds (FIG. 7B). The relatively high levels of epicatechin in bzl seeds allowed its stereochemistry to be determined using chiral-HPLC, confirming the (+)-epicatechin stereoisomer (FIG. 7C). In contrast, (-)- epicatechin is the only epicatechin stereoisomer detected in soybean seeds, likely generated by the LAR-LDOX-ANR complex and consistent with the result that only procyanidin B2 [(-)- epicatechin-(-)-epicatechin] was present in soybean seed coats.
[0202] Levels of procyanidin dimers, iso-B2 and particularly iso-B4, were higher in bzl seeds compared with purple BZ1 seeds with the natural UFGT allele (FIG. 7D). The higher level of iso-B4 than iso-B2 in bzl seeds is consistent with bzl accumulating more cysteinyl-catechin than cysteinyl-epicatechin (FIG. 7A). Phloroglucinolysis confirmed the increased epicatechin levels, but the low relative level of released (epi)catechin-phloroglucinol in bzl relative to BZ1 seeds, showing that epicatechin extension units were not present in longer PA polymers.
[0203] The biosynthetic pathways of anthocyanins, phlobaphenes, PAs, and other flavonoids partially overlap, sharing several common precursors and enzymes (FIG. 8).Example 7: Engineering Condensed Tannins in the ansla.2 Maize Mutant
[0204] The ansla2 maize mutant (Maizegdb web ID: M142C, in the W23 background) has a higher level of PA extension units (catechin-cysteine) compared with other candidate genes such as ufgt / bzl and gst / bz2 (FIG. 9). Unlike maize seeds with wild-type ANS / A2 allele (Maizegdb web ID: M142A, also in W23 background) that contain mostly epicatechin and catechin-cysteine as PA starter and extension units, respectively, in PA-rich species, ans / a2 mutant seeds contain catechin and catechin-cysteine as the major forms of PA starter and extension units, respectively. This was further confirmed by phloroglucinolysis (FIG. 10). The analysis of procyanidin dimers in maize seeds demonstrated that procyanidin B3 (catechincatechin) is the predominant form of PA dimers in ans / a2 seeds (FIG. 11), which is exactly the same composition of PAs found in another major monocot crop, barley (FIG. 12).
[0205] While there is an LAR gene in barley, no apparent LAR homolog has been identified in maize or sorghum based on sequence analysis. An isoflavone reductase-like enzyme (same family as LAR), however, could function as an LAR in the ans background, where leucocyanidin is hyper-accumulated. The Arabidopsis ans / ldox mutant does not produce catechin, and introduction of the LcLAR from Lotus into the Arabidopsis ans mutant can induce the production of catechin. In addition, both Arabidopsis and maize ans mutants accumulate catechin-cysteine in seeds.
[0206] To confirm that the PA phenotype in the maize ans mutant is consistent in different maize varieties, another maize ans / a2 mutant, M541G (a2, Rl-nj) from maizegdb was analyzed. The corresponding control material is M142V (A2, Rl-nj). M142A and M142C have the normal Rl-r allele (R is bHLH transcription factor equivalent to TT8 in Arabidopsis), meaning anthocyanins in wild type plants (Ml 42 A) accumulate throughout the aleurone layer of the seeds, while some anthocyanins also accumulate in coleoptile and roots. In contrast Ml 42V and M541G have the Rl-nj allele which restricts anthocyanin accumulation to only the crown and embryo of the seeds and in the coleoptile and roots. The compositions of PAs extracted from Ml 42V and M541G are consistent with that from M142A and M142C (FIG. 13A and FIG. 13B). Furthermore, M142C and M541G both possess a truncated a.2 gene due to transposon insertion. Primers were designed to amplify the A2 gene, and the PCR results show a clear band representing the wild type A2 allele in M142V, while the band is missing in M541G (FIG. 14).
[0207] Another advantage of using maize ans / a2 mutant as the background for engineering PAs is that this strategy does not require ANR. ZmANR is responsible for generating an unusual stereoisomer of epicatechins, (+)-epicatechin, which may negatively influence the polymerization of PA (FIG. 15). Mutation of the ANS gene in maize causes the plant to synthesize catechin-based PAs, which can bypass ANR and avoid the undesired form of epicatechin.
[0208] To determine whether PAs could be detected in vegetative tissues of these maize plants, DMACA staining of maize seedlings was performed 7-days post germination, focusing on coleoptile and roots where anthocyanins normally accumulate. Before DMACA staining, a red coloration can be seen in coleoptiles and roots in Ml 42V seedlings but not in M541G. After DMACA staining and repeated washing with 70% ethanol, intense purple color could be seen in coleoptile and roots of M541G, but not in Ml 42V, indicating accumulation of PAs in M541G vegetative tissues but not in M142V (FIG. 16 and FIG. 17). Phloroglucinolysisanalysis, which detects the nature of the PA starter and extension units, indicated that both leaves and roots of M541G contain PAs with catechin extension units and epicatechin starter units (FIG. 18). Preliminary quantitative analysis using DMACA reagent estimated PA levels of around 25 mg epicatechin equivalents per g dry weight in leaves, and around 15 mg per g dry weight in roots. Importantly, these results show that PAs can accumulate in both seeds and vegetative tissues in suitably engineered maize.
[0209] To further verify the relationship between PA accumulation and ANS gene function, new sets of maize a2 mutants from MaizeGDB, including 501G, 507 A, 507 AB, 507G and 511H were obtained (FIG. 19). Phloroglucinolysis of PAs extracted from seed of lines 511H, 501G, and 507A indicates that all contain PAs with catechin as a both starter and extension unit (FIG. 20). Thus, the evidence conclusively points to the fact that loss of function of ANS in maize leads to formation of catechin- based PAs. This appears to be the result of the leucocyanidin substrate of ANS being rechanneled as a PA extension unit.
[0210] To engineer maize for producing high levels of PAs, in certain embodiments, the maize ANS gene is knocked out by, for example, a site-specific genome modifying enzyme in cultivars of interest. Any ans mutant is then crossed with an anthocyanin-rich plant to generate plants with increased catechin-based PAs. In further embodiments, a maize plant comprising a mutant alle of the ANS gene, one example of which is the ansla.2 maize mutant allele, may be crossed with an anthocyanin-rich plant to generate plants with increased catechin-based PAs. In one embodiment, the anthocyanin-rich plant either endogenously expresses or has been engineered to express anthocyanin and / or PA regulatory transcription factors as described herein.Example 8: Engineering Condensed Tannins in Maize by Crossing ans Mutant Lines with Lines Expressing High levels of Anthocyanins.
[0211] One approach to increasing PA production in maize fines possessing mutant ans alleles is to activate the anthocyanin biosynthesis pathway while disrupting the ANS gene to redirect substrates from anthocyanin biosynthesis to condensed tannin biosynthesis. Toward this end, two strategies have been demonstrated, one non-transgenic strategy involving natural anthocyanin-rich maize and another strategy using transgenic maize expressing Sorghum TT2 (SbTT2; SEQ ID NO:4).
[0212] For the non-transgenic approach, the ans mutant line M541G a2 pll bl) was first crossed with an anthocyanin-rich maize M141A (A2 Pll Bl) to obtain Fl seeds. ZmPll (Pll)and ZMB 1 (Bl) are two maize transcription factors (MYB and bHLH family genes, respectively) that cooperatively activate anthocyanin biosynthesis in maize, particularly in vegetative tissues. These transcription factors are active in M141A but not in M541G, making M141A an ideal material for genetic crossing in this approach. F2 seeds were obtained by selfpollinating Fi plants. F3 seeds were obtained after self-pollinating selected F2 plants with homozygous a.2 alleles (FIG. 22). Unlike the yellow color of M541G seeds or the purple color of M141A seeds, the F3 seeds are brown in color, demonstrating altered flavonoid accumulation (FIG. 22).
[0213] Seeds of M541G, M141A, and F3 (M541G x M141A), and dissected roots and leaves (including coleoptiles) of seedlings at 5-days after germination were collected and stained with DMACA reagent to estimate condensed tannin accumulation and distribution. In contrast to the barely detectable condensed tannins in both parent lines (M541G and M141A), intense purple color was observed after DMACA staining in F3 seedlings, especially in roots, indicating high levels of condensed tannin accumulation (FIG. 23). Because the expression of Pll and B 1 can be up-regulated by strong light conditions, it is anticipated that more condensed tannins will be produced in leaf tissues when plants are moved from growth chambers into greenhouses. The data indicate that combining the Pll and Bl transcription factors that can activate anthocyanin biosynthesis with a disrupted A2 gene represents an effective approach to increase condensed tannin accumulation in maize, including in vegetative tissues.
[0214] As an alternative non-transgenic approach to confirm the broad applicability of the strategy to other combinations of mutant ans and anthocyanin over-expressing fines, M142C (a2 pll bl) was crossed with M142Y (A2 Pll Bl), both of which are in the W23 background. M142C accumulates catechin-based condensed tannins only in seed aleurone tissues but not in vegetative tissues, while M142Y does not appear to accumulate condensed tannins in any tissue. F3 seeds were obtained after two rounds of self-pollination (FIG. 24). To estimate levels of condensed tannin accumulation, seeds of F3 (M142C x M142Y), M142C and M142Y were germinated, and the leaves and roots of seedlings were harvested at 5-days after germination. After staining the tissues with DMACA reagent, purple color was observed only in F3 (M142C x M142Y) seedlings, particularly in root tissues, while the purple color was absent in the parent lines, M142C and M142Y (FIG. 25). Interestingly, only one side of the roots of F3 seedlings showed purple color, and it is likely that this is the side of the roots that is facing the light. These data further demonstrate that introducing functional Pll and Blalleles in a.2 plants can effectively enhance the accumulation of condensed tannins in maize vegetative tissues.
[0215] In an alternative approach, a transgenic plant over expressing anthocyanins as a result of expression of a transgene encoding a suitable transcription factor can be crossed with an ans mutant. As an example, the sorghum gene that encodes a MYB transcription activator for anthocyanin biosynthesis (SbTT2, SEQ ID NO:4) (Lu et al., Nature Communications, 14:4349, 2023) was introduced into a maize a2 mutant (501G) by crossing the a2 mutant with transgenic maize expressing SbTT2 in the B 104 background (line 971-31). F2 seeds were obtained after self-pollination of Fi plants (501Gx 971) (FIG. 26). F2 seeds (501Gx 971-65) were germinated along with seeds from the parent lines (501G and 971-31), and the leaves (including coleoptiles) and roots of the seedlings at 5-days after germination stained with DMACA reagents. Purple color was observed in selected F2 seedlings, in both leaf and root tissues, while the purple color was barely detectable in seedlings of the parent lines (501G and 971- 31) (FIG. 27). Similar to vegetative tissues, F2 seeds contained significantly higher levels of condensed tannins compared to the parent lines 971-31 or 501G (FIG 28). These data demonstrate that introducing SbTT2, an anthocyanin activator, into a2 maize plants significantly enhances the accumulation of condensed tannins in both seed and vegetative tissues.
[0216] In yet another approach, increasing PA production in maize may be achieved by reducing or disrupting the function of the ANS protein in lines or cultivars of interest by editing the ans gene to downregulate gene function. Gene editing may be performed using any method or enzyme described herein. In certain embodiments, it may be desirable to edit the ans gene of an anthocyanin-rich plant that either endogenously expresses or has been engineered to express anthocyanin and / or PA regulatory transcription factors as described herein. Nonlimiting examples of such plants include plants that either endogenously express or have been engineered to express a protein having at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19.Example 9: Identification of a2 Mutations in Different Maize a2 Alleles.
[0217] To genotype maize mutants with different a2 alleles, reliable PCR programs were developed to confirm the a2 mutations, and the promoter and coding regions of several natural a2 alleles were sequenced. Sequencing results show that in most cases there is a transposon insertion in the promoter region and / or in the coding region. The transposon insertion in thecoding region likely causes frameshift during protein translation, resulting in nonfunctional protein (FIG. 29). The sequences of the A2 promoter and coding regions in wild- type A2 alleles are similar to those deposited in GenBank (GenBank Accession No. X55314.1).Example 10: Materials and Methods
[0218] Chemicals and seed stocks. Chemical standards of (+)-catechin, (-)-epicatechin, cyanidin chloride, and cyanidin 3-O-glucocide were purchased from Sigma-Aldrich. The standard of (+)-epicatechin was purchased from Nacalai USA. Standards of 4 / ?-(S-cysteinyl)- catechin and 4 / ?-(S-cysteinyl)-epicatechin were synthesized using procyanidin dimers B3 or B2 and L-cysteine as described in Liu et al., Nat. Plants 2, 16182, 2016. FBLL (PI 546481), Black Mexican (PI162573), and Arequipa (PI 571427) seeds were ordered from USDA. Suntava and Osage seeds were ordered from Burpee Gardens and Baker Creek Heirloom Seeds, respectively. Maize seeds rich in phlobaphene (Pl-rr) in the A619 background were provided by Drs. Nan Jiang and Erich Grotewold at Michigan State University. Tobacco {Nicotiana tabacum, cv Xanthi) seeds constitutively expressing AtPAPl were obtained from Dr. De-Yu Xie at North Carolina State University. BZ1 and bzl maize seeds both in the W23 Cl, R-r) genetic background were provided by Dr. Virginia Walbot at Stanford University.
[0219] Plant growth conditions and generation of transgenic plants. Plants of maize, tobacco, and soybean were grown in greenhouses with temperature and light conditions set at 25 °C / 22 °C day / night, and 14 h / 10 h of light / dark. The coding sequence of GmANRl was cloned from soybean (cv Clark) as previously described (Lu et al., Plant Biotech. J. 19, 1429- 1442, 2021). The coding sequences of SbTT2 and SbMYB5 were synthesized at Gene Universal. Sequences encoding GmANRl, SbTT2, and SbMYB5 were individually cloned in pMCG1005 vector between Asci and Xma restriction enzyme sites and driven by the maize Ubiquitin promoter. Agrobacterium-mediated maize transformation was carried out by the Plant Transformation Facility at Iowa State University. GmANRl was transformed into Hill maize, and Fi seeds were obtained by self-pollination of transgenic plants. SbTT2 and SbMYB5 were transformed into B 104 maize, and Fi seeds were obtained by pollinating transgenic plants with pollen collected from untransformed B 104 plants.
[0220] Transgenic soybean (cv Clark) hairy roots were generated using an Agrobacterium- mediated transformation method. Briefly, coding sequences of SbTT2, SbMYB5, and GUS were cloned into pB7WG2D gateway vector driven by p35S promoter, and the recombinant plasmids were subsequently transformed into Agrobacterium rhizogenes strain K599.Cotyledons from soybean (cv. Clark) were used for transformation and transgenic roots were maintained in B5 medium and sub-cultured every 3 weeks. For tobacco transformation, coding sequences of AtANR, GmANRl, and ZmANRl were cloned into pB7WG2D gateway vector, and the recombinant plasmids were transformed into Agrobacterium tumefaciens strain LB 4404.
[0221] Enzymatic assays. GmANRl and ZmANRl were individually cloned in pDEST17 Gateway destination vector and transformed in Rosetta™ (DE3) cells (Millipore Sigma). Protein extraction and purification were based on protocols described previously (Kovinich et al., J. Agric. Food Chem. 60, 574-84, 2012). Enzymatic assays were carried out in a total volume of 100 pL containing 0.1 mM cyanidin, 1 mM NADPH, 50 mM MES buffer pH 5.7 and protein (50 pg) at 420C for 1 h and stopped by adding 200 pL ethyl acetate.
[0222] Transactivation assays and protein subcellular localization. Transactivation assays using protoplasts isolated from Arabidopsis leaves were performed following the protocol described previously (Yoo et al., Nat. Protoc. 2, 1565-1572, 2007), and firefly luciferase activity was quantified as previously described (Lu etal., Planta 246, 323-35, 2017). For subcellular localization, coding sequences of GmANRl and ZmANRl were inserted into pMDC43 gateway vector. The recombinant plasmids with GFP-tagged GmANRl and ZmANRl were then introduced into Arabidopsis protoplasts to visualize the subcellular localization of recombinant proteins. Confocal images of protoplasts expressing the GFP- tagged proteins were acquired by a Zeiss LSM710 confocal laser-scanning microscope. GFP was excited by a 488-nm laser and emission signals were collected at 500-540 nm.
[0223] RNA extraction and qRT-PCR analysis. RNA samples were extracted from maize leaves, seeds, and soybean hairy roots using PureLink Plant RNA Reagent (Invitrogen), and then treated with DNA- / ree DNA removal kit (Invitrogen) according to the manufacturers’ manuals. Then, cDNA was synthesized using iScript Select cDNA synthesis kit (Bio-Rad). Gene expression (transcript) levels were analyzed by qRT-PCR using PowerUP SYBR Green Master Mix (Applied Biosystems) and a QuantStudio 6 Flex Real-Time PCR system.
[0224] Extraction and quantification of anthocyanins and PAs. Anthocyanins were extracted from maize seeds by adding 0.5 mL of 0.1% HC1 in methanol to about 10 mg of ground seed powders, sonicating for 1 h, and shaking overnight in a dark cold-room. Samples were then centrifuged at 13,000 rpm for 5 min, and supernatants were transferred to new tubes. Anthocyanins were further fractionated by adding equal volumes of water and 0.5 mL ofchloroform and centrifuging at 13,000 rpm for 5 min. Total anthocyanins in the upper aqueous phase were quantified based on the absorbance at 520 nm using a standard curve generated with cyanidin 3-O-glucoside. Hydrolysis of anthocyanins was performed by heating under acidic conditions (2 N HC1) at 100°C for 1 h.
[0225] Soluble PAs were extracted by adding 1 mL of extraction buffer (70 % acetone, 0.5 % acetic acid) to 100-200 mg of ground whole maize seeds and sonicating for 1 h at room temperature. Samples were centrifuged at 13,000 rpm for 5 min, and supernatants transferred to new tubes. PAs in the supernatants were further fractionated by washing with equal volumes of chloroform three times and washing with hexane once. The cleaned PA samples were dried under vacuum and dissolved in 50 % methanol. PAs were quantified by reaction with DMACA and measuring the absorbance at 640 nm. Insoluble PAs were quantified using the butanol- HC1 method as described previously (Liu et al., Nat. Plants 2, 16182, 2016).
[0226] The distribution of insoluble PAs in soybean hairy roots was determined by staining cross sections with DMACA staining buffer (0.2 % DMACA in methanol / HaO containing 3 N HC1) overnight and washing repeatedly with 70 % ethanol prior to imaging. Phloroglucinolysis was performed by adding 50 pL of fresh phloroglucinol solution (5% phloroglucinol, 1% ascorbic acid, 1 N HC1 dissolved in methanol) to dried PA extract and incubating at 37°C for 20 min. The reaction was stopped by adding 50 pL of sodium acetate (0.2 M) and immediately subjecting to LC-MS analysis.
[0227] Anthocyanins were analyzed using an Agilent 1290 Infinity II HPLC system equipped with an Eclipse Plus C18 column (4.6 x 250 mm, 5 pm). The samples were separated at a rate of 1 mL / min using solvent A (0.1 % formic acid) and solvent B (acetonitrile) with the following protocol: 0 to 5 min, 5 % B; 5 to 10 min, 5-10 % B; 10 to 25 min, 10-17 % B; 25 to 40 min, 17-50 % B; 40 to 50 min, 50-90 % B; 50 to 60 min, 90-100 % B. Data were collected at 520 nm.
[0228] Chiral HPLC analysis of epicatechin was performed using an Agilent HP1100 system equipped with a chiral column (Chiral Technologies #80325) with solvent A (0.5 % acetic acid in hexane) and solvent B (0.5 % acetic acid in ethanol) at 1 mL / min. The protocol used was as follows: 0 to 20 min, 20 % B; 20 to 23 min, 20-50 % B; 23-38 min, 50 % B; 38-40 min, 50- 20 % B. Signals were recorded at 280 nm.
[0229] Soluble PAs were analyzed with an Exion ultra high-performance liquid chromatography system coupled with a high resolution TripleTOF6600+ mass spectrometerfrom AB Sciex. Specifically, compounds were separated using a Cl 8 Acquity UPLC HSS T3 (100 x 2.1 mm, 1.8 pm) column from Waters. The temperatures of the column compartment and the autosampler were kept at 42 °C and 15 °C, respectively. The analytes were eluted using a gradient of 0.1 % formic acid in water (Solvent A) and 0.1 % formic acid in methanol (Solvent B) under a flow rate of 0.4 mL / min. The following gradient was applied: 0-1.0 min, 5.0 % B;I.0-2.0 min, 5.0-10.0 % B; 2.0-7.0 min, 10.0-28.2 % B; 7.0-11.0 min, 28.2-70.0 % B; 11.0-II.1 min, 70.0-95.0 % B; 11.1-13.0, 95 % B; 13.0-13.1 min, 95.0-5.0 % B; 13.1-15.0 min, 5.0 % B. In short, the mass spectrometer was set to scan metabolites from m / z 250-1000 amu in negative mode with an ion spray voltage of 4000 V and the accumulation time was 100 msec. MS / MS spectra were acquired over m / z 30-1000 amu with an accumulation time of 25 msec and parameters such as declustering potential, collision energy, and collision energy spread were set to 50 V, 25 V and 10 V, respectively. The temperature of the source was 500 °C. The analysis of the data was performed using Sciex OS software.
[0230] Extraction of polyphenols from maize seeds. To analyze procyanidin dimers and trimers in maize seeds, a modified polyphenol extraction and purification method was used. Briefly, maize seeds were ground into fine powders using liquid nitrogen and 1 mL of 80% methanol was added to 100-200 mg homogenized samples. After sonication for 1 h, samples were centrifuged for 5 min at 13,000 rpm. Supernatants were transferred to new tubes and dried under vacuum. Next, 100 pL of water was added to each sample, and 200 pL of ethyl acetate was then added, followed by vortexing and centrifugation at 13,000 rpm for 5 min. The upper ethyl acetate phases were transferred to new tubes and dried under vacuum. Finally, the samples were dissolved in 20 pL of 50 % methanol and stored at -20 °C until use.
[0231] Phylogenetic analysis. Multiple protein sequence alignments were performed using the ClustalW program, and the phylogenetic tree was generated by the MEGA program following the Neighbor-Joining method with 2000 bootstrap replicates (Tamura, et al., Mol. Biol. Evol. 30, 2725-2729, 2013).
[0232] Statistical analysis. Significant differences (P < 0.05, 0.01) were determined by Student’s Mest. All statistical analyses were performed using three to five biological replicates, and n values in different experiments are listed in figure legends.* * *
[0233] All of the methods disclosed and claimed herein can be made and executed without undue experimentation in light of the present disclosure. While the compositions and methodsof this invention have been described in terms of preferred embodiments or aspects, it will be apparent to those of skill in the art that variations may be applied to the methods and in the steps or in the sequence of steps of the method described herein without departing from the concept, spirit, and scope of the invention. More specifically, it will be apparent that certain agents which are both chemically and physiologically related may be substituted for the agents described herein while the same or similar results would be achieved. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims.
Claims
CLAIMS1. A maize plant, seed, cell, or plant part comprising: a) a mutant allele or at least one genomic edit that downregulates anthocyanidin synthase (ans) gene function; and b) a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19.
2. The maize plant, seed, cell, or plant part of claim 1 , wherein the maize plant, seed, cell or plant part comprises a recombinant polynucleotide molecule comprising a nucleotide sequence encoding the polypeptide.
3. The maize plant, seed, cell, or plant part claim 1, wherein the modified maize plant, seed, cell, or plant part comprises increased catechin based proanthocyanidins or increased condensed tannins.
4. The maize plant, seed, cell, or plant part of claim 1, wherein in the maize plant, seed, cell, or plant part comprises at least one truncated allele of an endogenous ans gene.
5. The maize plant, seed, cell, or plant part of claim 4, wherein the truncated ans gene encodes a truncated polypeptide comprising a fragment of a sequence having at least about 85% sequence identity to SEQ ID NO: 12.
6. The maize plant, seed, cell, or plant part of claim 1, wherein: a) the mutant allele or the genomic edit is present in at least one allele of an endogenous ans gene; b) the mutant allele or the genomic edit is in an endogenous ans gene encoding a protein having at least about 70% sequence identity to SEQ ID NO: 12; c) the genomic edit is in a transcribable region of the ans gene; d) the genomic edit is in an intron region of the ans gene; e) the genomic edit is in an exon region of the ans gene; f) the mutant allele or the genomic edit comprises a deletion, an insertion, a substitution, an inversion, a duplication, or any combination thereof;g) the genomic edit is in an intron region and an exon region of the ans gene; h) the maize plant, seed, cell, or plant part is heterozygous for the mutant allele or the genomic edit; or i) the maize plant, seed, cell, or plant part is homozygous for the mutant allele or the genomic edit.
7. The maize plant, seed, cell or plant part of claim 1, wherein the mutant allele or the genomic edit reduces or disrupts the activity of ANS protein compared to the activity of ANS protein in an otherwise identical maize plant, seed, cell, or plant part that lacks the mutant allele or the genomic edit.
8. A method for producing a modified maize plant cell comprising: a) introducing a genomic edit into at least one target site of an endogenous ans gene of a maize plant cell that downregulates ans gene function, wherein said maize plant cell comprises a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19; and b) selecting at least one plant cell comprising said genomic edit.
9. The method of claim 8, wherein said selecting comprises detecting said genomic edit or measuring anthocyanidin synthase (ans gene function.
10. The method of claim 8, wherein the maize plant cell comprises a recombinant polynucleotide molecule comprising a nucleotide sequence encoding said polypeptide.
11. The method of claim 8, the method further comprising regenerating at least one modified maize plant or plant part from the at least one selected plant cell or a descendant thereof comprising said genomic edit.
12. The method of claim 8, wherein: a) the modified maize plant cell comprises at least one truncated allele of an endogenous ans gene; b) the genomic edit is present in at least one allele of an endogenous ans gene; c) the genomic edit is in an endogenous ans gene encoding a protein having at least about 70% sequence identity to SEQ ID NO: 12;d) the target site is located in a transcribable region of the ans gene; e) the target site is located in an intron region of the ans gene; f) the target site is located in an exon region of the ans gene; g) the target site is located in an intron region and an exon region of the ans gene; h) the modified maize plant cell is heterozygous for the genomic edit; i) the modified maize plant cell is homozygous for the genomic edit; or j) the genomic edit comprises a deletion, an insertion, a substitution, an inversion, a duplication, or any combination thereof.
13. The method of claim 11, the method further comprising selecting at least one modified maize plant or plant part comprising increased catechin based proanthocyanidins or increased condensed tannins.
14. The method of claim 8, wherein introducing the genomic edit comprises use of at least one site-specific genome modifying enzyme in said plant cell.
15. The method of claim 14, wherein the site-specific genome modifying enzyme: a) is selected from the group consisting of an RNA-guided nuclease, a zinc-finger nuclease, a meganuclease, a TALE-nuclease, a recombinase, a transposase, and combinations of any thereof; b) is an RNA-guided nuclease; or c) creates at least one strand break at the target site.
16. The method of claim 15, wherein the RNA-guided nuclease is selected from the group consisting of a Cas nuclease, a Cpf 1 nuclease, or a variant of either thereof.
17. A method for producing a modified maize plant cell comprising: a) introducing a recombinant polynucleotide molecule comprising a nucleotide sequence encoding a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19 into a maize plant cell, wherein the maize plant cell comprises a mutant allele or a genomic edit in an endogenous ans gene that downregulates ans gene function; andb) selecting at least one plant cell comprising said recombinant polynucleotide molecule.
18. The method of claim 17, wherein selecting the plant cell comprises: a) detecting said recombinant polynucleotide molecule; or b) detecting expression of said polypeptide.
19. The method of claim 17, the method further comprising regenerating at least one modified maize plant or plant part from the at least one selected plant cell or a descendant thereof comprising said recombinant polynucleotide molecule.
20. The method of claim 17, wherein in the modified maize plant cell comprises at least one truncated allele of an endogenous ans gene.
21. The method of claim 20, wherein the truncated ans gene encodes a truncated polypeptide comprising a fragment of sequence having at least about 85% sequence identity to SEQ ID NO:12.
22. The method of claim 17, wherein: a) the mutant allele or the genomic edit is present in at least one allele of an endogenous ans gene; b) the mutant allele or the genomic edit is in an endogenous ans gene encoding a protein having at least about 70% sequence identity to SEQ ID NO: 12; c) the genomic edit is located in a transcribable region of the ans gene; d) the genomic edit is located in an intron region of the ans gene; e) the genomic edit is located in an exon region of the ans gene; f) the genomic edit is located in an intron region and an exon region of the ans gene; g) the modified maize plant cell is heterozygous for the mutant allele or the genomic edit; h) the modified maize plant cell is homozygous for the mutant allele or the genomic edit; or i) the mutant allele or the genomic edit comprises a deletion, an insertion, a substitution, an inversion, a duplication, or any combination thereof.
23. The method of claim 19, the method further comprising selecting at least one modified maize plant or plant part comprising increased catechin based proanthocyanidins or increased condensed tannins.
24. A method for producing a modified maize plant cell comprising: a) introducing a genomic edit into at least one target site of an endogenous ans gene of a maize plant cell that downregulates ans gene function; b) introducing a recombinant polynucleotide molecule comprising a nucleotide sequence encoding a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19 into the maize plant cell; and b) selecting at least one plant cell comprising said genomic edit and said recombinant polynucleotide molecule.
25. The method of claim 24, the method further comprising regenerating at least one modified maize plant or plant part from the at least one selected plant cell or a descendant thereof comprising said genomic edit and said recombinant polynucleotide molecule.
26. A method for producing a hybrid maize plant, the method comprising crossing the maize plant of claim 1 with a second, non-isogenic maize plant.
27. A method for producing a maize plant, the method comprising crossing the plant of claim 1 with itself or a second maize plant.
28. A method for producing a maize progeny plant or seed comprising increased catechin based proanthocyanidins or increased condensed tannins, the method comprising: a) crossing a first maize plant with a second maize plant, wherein the first maize plant comprises a mutant allele or a genomic edit that downregulates anthocyanidin synthase (ans gene function, and the second maize plant comprises a polypeptide having an amino acid sequence with at least about 85% sequence identity to SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, or SEQ ID NO: 19; and b) selecting a maize progeny plant or seed comprising said mutant allele or said genomic edit and said polypeptide.
29. The method of claim 28, wherein selecting the maize progeny plant or seed comprises:a) detecting the mutant allele or the genomic edit; b) measuring anthocyanidin synthase ans) gene function; c) detecting said polypeptide; or d) detecting the increased catechin based proanthocyanidins or increased condensed tannins in a sample derived from the progeny plant or seed.
30. The method of claim 28, wherein the second maize plant comprises a recombinant polynucleotide molecule comprising a nucleotide sequence encoding said polypeptide.
31. The method of claim 30, wherein selecting the maize progeny plant or seed comprises: a) detecting said mutant allele or said genomic edit; b) measuring anthocyanidin synthase ans) gene function; c) detecting said polypeptide; or d) detecting said recombinant polynucleotide molecule.
32. The method of claim 28, further comprising: a) selfing the maize progeny plant or crossing the maize progeny plant with the first maize plant or the second maize plant; and b) selecting a further maize progeny plant or seed comprising said mutant allele or said genomic edit and said polypeptide.
33. A maize progeny plant or seed produced by the method of claim 24, wherein said maize progeny plant or seed comprises said mutant allele or said genomic edit and said polypeptide.
Citation Information
Patent Citations
Sequence-determined DNA fragments and corresponding polypeptides encoded thereby
US20060048240A1
Manipulation of proanthocyanidin (PA) composition by affecting anthocyanidin synthase (ANS) and leucoanthocyanidin dioxygenase (LDOX)
US20190017060A1
Compositions and methods comprising plants with modified anthocyanin content
US20230270073A1
A plant, its use as a nutraceutical and the identification thereof
WO2005103258A1
Cited By
Molecular marker primer group for identifying content of tannin in corn kernels and application of molecular marker primer group
CN121737339A