uORF::reporter gene ligation to select sequence changes for gene editing of uORFs to regulate the expression of ascorbic acid genes
By editing uORFs in the GGP gene to regulate ascorbic acid production, the method addresses the limitations of existing technologies in enhancing stress tolerance and growth morphology in non-food crops, achieving desirable traits without adverse effects.
Patent Information
- Application Number
- JP2025519914
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-05
- Filing Date
- 2023-10-04
- Publication Date
- 2025-10-22
AI Technical Summary
Existing methods for increasing intracellular ascorbic acid levels in plants to enhance stress tolerance and growth morphology have not fully explored the potential of commercial non-food crops, and high levels of ascorbic acid can lead to negative effects such as reduced yield and abnormal growth.
Introduce mutations into the upstream open reading frames (uORFs) of the GGP gene using gene editing techniques to regulate ascorbic acid production, selecting mutations that result in a moderate increase in ascorbic acid without deleterious effects on growth, and use a reporter system to identify desirable mutations.
This approach allows for the creation of plants with enhanced stress tolerance and desirable morphological traits, such as increased abiotic stress tolerance and altered flowering time, without negative phenotypes like dwarfism or sterility, suitable for commercial applications.
Smart Images

Figure 2025535075000001 
Figure 2025535075000002 
Figure 2025535075000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to the creation of mutations in upstream open reading frames (uORFs) that control DNA and protein expression and their use in plants. [Background technology]
[0002] Previous research has focused on developing plants with elevated levels of ascorbic acid (vitamin C) to improve their nutritional value upon human or animal consumption. However, the potential for increasing intracellular ascorbic acid levels in crop plants to confer abiotic and / or biotic stress tolerance and / or modified growth morphology has not been fully explored. In particular, most previous studies have focused on genetically modifying ascorbic acid levels in food crops. Significant potential remains for the creation of commercial non-food crops (e.g., crops grown for energy generation, carbon capture, chemical processing, materials, or fiber applications) that are protected from stress by expressing elevated ascorbic acid levels.
[0003] Ascorbic acid is recognized to play an important role in protecting cells from the harmful effects of free radicals produced by cells under conditions of oxidative stress. Plants are exposed to a variety of stresses throughout their lives due to environmental factors, including temperature, water and nutrient availability, light intensity, and the presence of pathogens. All of these factors have the common effect of increasing the levels of oxidative stress and free radicals in plant cells, thereby causing damage to macromolecules such as nucleic acids. In crops, this damage often leads to impaired growth, abnormal growth, and poor yield quality. The invention detailed herein provides novel crop traits and methods for creating such traits that protect against the harmful effects of oxidative stress-induced cell damage by increasing the levels of ascorbic acid in plant cells through increasing the levels of the GDP-L-galactose phosphorylase (GGP) gene product. This can be achieved by generating transgenic plants expressing GGP from a heterologous promoter or by introducing mutations into a uORF previously known to negatively regulate the production of the GGP gene product. In the latter approach, the resulting allele typically contains a mutation absent within the start codon or a mutation that substantially deletes the uORF and results in a moderate increase in ascorbic acid levels compared to wild-type plants. Many researchers have reported the creation of transgenic or gene-edited plants with elevated ascorbic acid levels through modification of the GGP gene (see, for example, Koukounaras et al., Plant Physiology and Biochemistry 193, 15 December 2022, 124-138; Yang et al., 2023 Protoplasma 260(2):625-635). The production of extremely high levels of ascorbic acid can result in negative effects such as reduced yield, dwarfism, abnormal stem morphology, and abnormal organs. Therefore, alleles that result in a low-to-moderate increase in ascorbic acid are usually highly desirable. Herein, we present a novel reporter system-based approach to identify desirable mutations in the die uORF of the GGP gene that result in a desirable increase in GGP protein and / or ascorbic acid without deleterious effects on growth.Such or equivalent changes can be reproduced in the endogenous GGP gene of target crops by gene editing or TILLING-based approaches, although for ornamental or turfgrass species, dwarfism or growth retardation combined with stress tolerance may be highly desirable, leading to selection for stronger alleles.
[0004] It is important to note that because GGP acts in a dose-dependent manner, the GGP uORF-disrupted allele behaves semidominantly. Therefore, to optimize dosage and eliminate undesirable negative phenotypes such as sterility and dwarfism, genetic crosses are effective, in which homozygous lines carrying the GGP uORF mutant allele are crossed with heterozygous or wild-type lines to produce Fl heterozygotes. Indeed, Fl heterozygous seeds carrying such alleles would be desirable for commercial seed producers and / or cultivars.
[0005] The inventions disclosed herein are applicable to any gene encoding GGP, as GGP evolved as a key regulatory enzyme in ascorbic acid biosynthesis in plants. Indeed, this enzyme is conserved throughout the plant kingdom, with GGP genes recently identified in over 70 plant species by Tao et al. (AoB Plants, 2020, November 4;12(6):plaa055). Exemplary GGP proteins are disclosed herein as SEQ ID NOs: 157-242. However, the set of disclosed sequences is merely illustrative and not limiting. The range of GGPs to which these inventions are applicable is very broad and includes any that contain functional homologs of the disclosed GGP proteins.
[0006] Upstream open reading frames (uORFs) belong to a class of small, conserved open reading frames located upstream of the protein-coding major open reading frame (mORF) in the leader sequence of an mRNA (also known as the 5'-untranslated region (5'UTR)). uORFs function as cis-acting elements that regulate the activity of downstream polypeptide-encoding sequences. Therefore, genetic manipulations, such as gene editing techniques, that introduce mutations into uORF sequences offer new possibilities for activating the expression of downstream open reading frames encoding polypeptides of interest.
[0007] uORF regulatory elements are widely present in eukaryotic mRNAs. However, not all eukaryotic genes contain uORFs. In some cases, uORFs are thought to regulate the rate of translation initiation of downstream coding sequences (CDSs) by sequestering ribosomes. In other cases, uORFs encode short, evolutionarily conserved peptides that function as cis-acting inhibitory peptides for downstream mORFs. In many cases, the actual presence of uORFs is highly conserved across species. Therefore, once a uORF is identified at a locus of interest in one species, the homologous locus in another species will usually also contain a uORF and be subject to uORF suppression.
[0008] Genome-wide studies have revealed the broad regulatory functions of uORFs in different species under different biological conditions (Zhang et al. 2019. Trends Biochem. Sei. 44:782-794. doi: 10.1016 / j.tibs.2019.03.002). Specific uORFs may function as translational control elements, regulating the expression of associated downstream major open reading frames (mORFs). Translational regulation of mORFs by highly conserved uORFs in response to intracellular metabolite levels has been reported in plant studies (Hayden CA and Jorgensen RA 2007. BMC Biol. 5:32; Tran MK, et al. 2008. BMC Genomics 9:361). This allows for increased levels of a polypeptide of interest by, for example, knocking out the expression of a negatively acting uORF located upstream of the polypeptide-encoding sequence. Genetic modifications to uORF sequences can generate new desirable phenotypes with diverse applications depending on the specific cell or species. These applications include new crop traits (e.g., increased vigor, stress tolerance, delayed or early flowering, morphological variation, increased branching, reduced apical dominance, increased yield, and / or improved nutritional content), weed and other pest control, and activation of gene networks that switch on cell death, cell death or tumor suppressor genes in cancer cells, and / or production of desirable metabolites or peptides in fermentation systems.
[0009] The present disclosure relates to methods and compositions for the production of commercially valuable plants and crops, as well as methods for making and using them.
[0010] The uORFs provided and characterized in this disclosure may be modified to produce plants with altered traits that meet the needs of agriculture, food production, materials production, and environmental restoration and carbon sequestration. For example, GGP gene activity can be regulated by gene editing of polynucleotides encoding the uORF peptides identified by Liang et al. (U.S. Patent No. 9,648,813).
[0011] These traits may provide significant value by enabling plants to thrive in harsh environments where temperature, water or nutrient availability, or salinity limit or prevent growth of plants without the modified trait. These traits may also include desirable morphological changes, including altered flowering time, increased or decreased size, pest and disease resistance, light response, altered biochemical composition, and other desirable phenotypes. With increasing interest in producing crops under controlled indoor environments, traits such as delayed flowering and a more compact morphology are often desirable, especially in leafy vegetables.
[0012] Other aspects and embodiments of the disclosure are described below and may be derived from the disclosure as a whole. Summary of the Invention [Means for solving the problem]
[0013] The present disclosure is directed to methods for producing plants that enhance, enhance, or increase a desirable trait in the plant. The plant or cells thereof may be ornamental, turf, weed, or crop. The disclosure is directed to plants produced by the disclosed methods. Plants with enhanced desirable traits are modified by introducing a nucleic acid comprising a 5' UTR (untranslated region) derived from the upstream region of a gene of interest that directly or indirectly enhances the desirable trait. The 5' UTR comprises a mutation in an upstream open reading frame (uORF). The 5' UTR comprises any of the following: (1) a sequence having at least 70% identity to any of SEQ ID NOs: 41-100, 111-131, 144-145, 147, 149-150, or 153-155; (2) a sequence encoding a polypeptide having a sequence having at least 70% identity to any of SEQ ID NOs: 1-40, 108, 132-137, 146, 148, 151-152, or 156; The mutation disrupts the function of an upstream open reading frame (uORF) encoded by the 5'UTR, and the modification may be characterized by either a deletion, insertion, or substitution of at least one nucleotide in the 5'UTR.
[0014] Mutations can be introduced by means well known to those skilled in the art, including, but not limited to, radiation, photomutagenesis, or chemical mutagenesis. Mutations can also be created by targeted gene editing, in which the sequence changes generated by editing can be preselected or chosen through a process of conjugating a mutated 5' UTR with a polynucleotide encoding a reporter protein (e.g., luciferase (LUC), 3'-glucuronidase (GUS), or green fluorescent protein (GFP)). The resulting conjugates are then introduced into test plants, plant cells, or protoplasts. Conjugates in which the function of the uORF is completely eliminated typically exhibit substantially higher levels of reporter product (sometimes several-fold) compared to reporter conjugates in which the 5' UTR is fully retained. This system can be used to separate the function of different regions within a uORF contained within the 5' UTR. Researchers often select "weak" mutations within the uORF that derepress, but do not completely derepress, the uORF, thereby reducing the level of reporter product compared to complete elimination of uORF function. In such cases, the level of reporter gene product in the test plant or plant cell or protoplast is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the level of reporter product produced in a control plant or plant cell or protoplast containing a control 5'UTR:reporter polynucleotide in which the uORF has been partially or completely deleted from the 5' UTR linked to the reporter gene, thereby rendering the uORF inoperative.
[0015] Once a desirable mutation in a uORF is selected by reporter protein analysis, the practitioner uses gene editing to introduce the same or a corresponding mutation into the 5'UTR of a gene of interest at an endogenous locus in the genome of a plant species in which the desired trait is to be produced. This typically involves introducing a guide nucleic acid targeting the selected endogenous locus for editing into cells of the plant species, followed by regeneration of plants from tissue culture containing the novel allele resulting from the editing process. Individual plants with desirable traits may be selected from a population of plants containing nucleic acid changes in the uORF. "Desirable traits" referred to herein (including at least one of the present "desirable traits") include, but are not limited to, increased or decreased free radical damage, reduced free radical levels, increased stress tolerance, increased abiotic stress tolerance (e.g., salt tolerance, radiation tolerance, pollution tolerance, heat tolerance, cold tolerance, and / or drought tolerance), increased biotic stress tolerance, increased disease tolerance, increased disease resistance, increased nutrient utilization efficiency, increased nitrogen utilization efficiency, increased branching, and / or increased oxidative stress tolerance. The present disclosure also provides a method for producing a plant cell or plant selected for a current desirable trait compared to a control plant, the method comprising modifying the 5'-UTR of the GGP gene in the plant cell or plant, wherein the 5'-UTR comprises a sequence having at least 70% identity to any one of SEQ ID NOs: 41-100, 111-131, 144-145, 147, 149-150, or 153-155, or a sequence encoding a polypeptide having at least 70% identity to any one of SEQ ID NOs: 1-40, 108, 132-137, 146, 148, 151-152, or 156, wherein the modification disrupts function of an upstream open reading frame (uORF) encoded by the 5'-UTR, and the modification is at least one of a deletion, addition, or substitution of at least one nucleotide in the 5'-UTR.
[0016] The present specification also provides a method for enhancing a desired trait in a plant. The method includes providing a nucleic acid comprising a 5'-UTR (untranslated region) obtained from the upstream region of a gene that directly or indirectly enhances the desired trait. The 5'-UTR includes an upstream open reading frame (uORF). The 5'-UTR includes any of the following: (1) a sequence having at least 70% identity to any one of SEQ ID NOs: 41-100, 111-131, 144-145, 147, 149-150, or 153-155; and (2) A sequence encoding a polypeptide having a sequence identical to any one of SEQ ID NOs: 1-40, 108, 132-137, 146, 148, 151-152, or 156, wherein the modification disrupts the function of an upstream open reading frame (uORF) encoded by the 5'-UTR, and the modification is at least one of a deletion, addition, or substitution of at least one nucleotide in the 5'-UTR.
[0017] The nucleic acid is linked to a reporter gene (e.g., LUC, GUS, or GFP) to create a 5'UTR::reporter polynucleotide. Mutations are introduced into the nucleic acid (e.g., by direct synthesis) to create a 5'-UTR:reporter polynucleotide with a mutation in the uORF or putative uORF. The 5'UTR:reporter polynucleotide is introduced into a plant cell. The amount of reporter gene product in the plant cell is measured. Mutations that produce less product than the amount of product in control plant cells containing a control 5'UTR:reporter polynucleotide in which the uORF has been removed from the 5'UTR linked to the reporter gene are then selected or identified. The amount of product in the plant cell may be less than about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100% of the amount of reporter product produced in the control plant. The selected mutation in the uORF that causes the desired change in the level of the reporter product is then introduced into the genome of a cell of a crop plant, and the cell is regenerated into a plant that exhibits an enhancement of the desired trait. The resulting plant with the updated desired trait is used in breeding to introduce the trait into a commercial germplasm or to produce a population of plants with the same allele and trait through vegetative propagation. Typically, in plant breeding, a plant with the desired allele with the desired trait is backcrossed into a desired variety for up to eight generations to "fix" the allele into the genetic background of the desired variety. As used herein, "desirable trait" includes, but is not limited to, at least any of the following: reduced free radical damage, reduced free radical levels, increased abiotic stress tolerance, increased biotic stress tolerance, increased disease resistance, increased disease resistance, increased nutrient utilization efficiency, increased nitrogen utilization efficiency, and increased oxidative stress tolerance.
[0018] [Brief explanation of sequence listing] The Sequence Listing provides representative polynucleotide and polypeptide sequences of the present disclosure. The Sequence Listing is named GXTR-OOOlPCT.xml, was created on October 3, 2023, and has a size of 319,872 bytes. Where necessary, the entire contents of the Sequence Listing are incorporated herein by reference. DETAILED DESCRIPTION OF THE INVENTION
[0019] [Definition] "Expression" refers to the production of mRNA from a gene by transcription, or the production of a polypeptide from RNA by translation. "uORFs" refers to upstream open reading frames, often found within the leader sequence of an mRNA transcript, located upstream of the main ORF of the protein-coding region. (Note: mORFs are sometimes also called long ORFs or major ORFs; in this application, the terms mORF, main ORF, long ORF, and major ORF are used interchangeably.) uORFs are a group of small ORFs that function as repressors of downstream mORFs. uORFs may encode evolutionarily conserved functional peptides (e.g., cis-acting regulatory peptides) that function as repressors, for example, through translational repression. uORFs are generally defined by an upstream (i.e., 5'-direction) start codon (a three-base pair codon containing at least two of the following bases, in order: AUG) and an in-frame stop codon (UAA, UAG, or UGA) in the 5'-UTR, and do not overlap the main coding sequence.
[0020] "5'UTR" means "5' untranslated region," and as used herein refers to the leader sequence at the 5' end of an mRNA molecule, located upstream of the main ORF. Note that the term "untranslated" is inappropriate in some cases, since the uORF present in the 5'UTR may itself be translated into a short peptide.
[0021] A "polypeptide" is an amino acid sequence comprising a plurality of contiguous polymerized amino acid residues, e.g., at least about 15 contiguous polymerized amino acid residues, alternatively at least about 30 contiguous polymerized amino acid residues, or at least about 50 contiguous polymerized amino acid residues. Often, a polypeptide comprises a sequence of polymerized amino acid residues that is a transcription factor or a domain, portion, or fragment thereof. Additionally, a polypeptide may comprise 1) a localization domain, 2) an activation domain, 3) a repression domain, 4) an oligomerization domain, or 5) a DNA-binding domain, or the like. A polypeptide may optionally comprise modified amino acid residues, naturally occurring amino acid residues that are not encoded by a codon, or non-naturally occurring amino acid residues.
[0022] "Identity" or "similarity" refers to the sequence similarity between two polynucleotide sequences or two polypeptide sequences, with identity being a more strict comparison. "Percent identity" and "% identity" refer to the percentage of sequence identity found in a comparison of two or more polynucleotide sequences or two or more polypeptide sequences. "Sequence similarity" refers to the percentage of base pair sequence similarity (determined by an appropriate method) between two or more polynucleotide sequences. Two or more sequences may have a similarity anywhere from 0 to 100%, or any integer value therebetween. Identity or similarity is determined by comparing corresponding positions in each sequence aligned for comparison. If corresponding positions in the compared sequences are occupied by the same nucleotide base or amino acid, the molecules are identical at that position. The degree of identity or similarity between polynucleotide sequences is a function of the number of identical or matching nucleotides at positions shared by the polynucleotide sequences. The degree of identity between polypeptide sequences is a function of the number of identical amino acids at positions shared by the polypeptide sequences. A degree of homology or similarity of polypeptide sequences is a function of the number of amino acids at positions shared by the polypeptide sequences.
[0023] As further described and used herein, a "homolog" refers to a variant of a polypeptide or transcription factor from the same or a different species that has a substantial level of identity in its conserved regions and / or overall sequence, with a level of identity of at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or more preferably at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, or 79% compared to the first polypeptide or transcription factor. 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity, and the polypeptide or transcription factor has a similar or equivalent function in a cell or organism to the first polypeptide or transcription factor. Furthermore, homologous nucleotide sequences hybridize under more stringent conditions.
[0024] As used herein, the term "GGP protein" refers to a group of proteins produced in plants by genes encoding GDP-L-galactose phosphorylase, including, for example, the protein products encoded by the Arabidopsis locus AT4G26850 and the Arabidopsis VTC5 locus AT5G55120 (referred to as "Arabidopsis GGP proteins"), including the two protein sequences represented by SEQ ID NO: 239 and SEQ ID NO: 240.
[0025] As used herein, the term "homologue" refers to a polypeptide that, when subjected to BLAST analysis against a group of proteins encoded by the Arabidopsis proteome, shares a higher level of sequence identity with the Arabidopsis GGP protein than with any other protein in the Arabidopsis proteome, and has an HSP bit score of 50 or greater similarity to the Arabidopsis GGP protein.
[0026] "Orthologs" are evolutionarily related genes that have similar sequence and similar function. Orthologs are structurally related genes in different species that have been derived by a speciation event.
[0027] As used herein, "variant" refers to a polynucleotide or polypeptide sequence that differs from a specifically identified sequence by the deletion, substitution, or addition of one or more nucleotides or amino acid residues. Such variants include naturally occurring allelic variants or non-naturally occurring variants. Such variants may be from the same or other species and may include homologs, paralogs, and orthologs. In certain embodiments, the polypeptides of the invention and variants thereof have the same or similar biological activity as the polypeptides or polypeptides of the invention. The term "variant" with respect to polypeptides and polypeptides includes all forms of polypeptides and polypeptides as defined herein (Liang et al., U.S. Patent No. 9,648,813).
[0028] "Inverted repeats" refer to repetitive sequences in which the second half of the repeat is on the complementary strand of DNA. In such cases, read-through transcription produces transcripts that form "hairpin" structures through complementary base pairing, provided there is a 3-5 base pair spacer between the repeat regions.
[0029] "Non-food plants" refers to crops grown, for example, for turfgrass, ornamental plants, or for the purpose of carbon dioxide sequestration or the production of manufactured goods (e.g., fibers, biofuels, specialty chemicals, lubricants, building materials, pharmaceuticals, biopolymers, etc.).
[0030] [Polynucleotide variants] Variant polynucleotide sequences are preferably at least 50%, more preferably at least 51%, even more preferably at least 52%, more preferably at least 53%, more preferably at least 54%, more preferably at least 55%, more preferably at least 56%, more preferably at least 57%, more preferably at least 58%, more preferably at least 59%, more preferably at least 60%, more preferably at least 61%, more preferably at least 62% similar to the sequences of the present disclosure. More preferably, it has at least 63% identity, more preferably at least 64% identity, more preferably at least 65% identity, more preferably at least 66% identity, more preferably at least 67% identity, more preferably at least 68% identity, more preferably at least 69% identity, more preferably at least 70% identity, more preferably at least 71% identity, more preferably at least 72% identity, more preferably at least 73% identity, more preferably at least 74% identity, more preferably at least 75% identity, more preferably at least 76% identity, more preferably at least 77% identity, more preferably at least 78% identity, more preferably at least 79% identity, more preferably at least 80% identity, more preferably at least 81% identity, more preferably at least 82% identity, more preferably at least 83% identity, more preferably at least 84% identity, more preferably at least 85% identity, more preferably at least 86% identity, more preferably at least 87% identity, more preferably at least 88% identity, more preferably at least 89% identity, more preferably at least 90% identity, more preferably at least 91% identity, more preferably at least 92% identity, more preferably at least 93% identity, more preferably at least 94% identity, more preferably at least 95% identity, more preferably at least 96% identity, more preferably at least 97% identity, more preferably at least 98% identity, and most preferably at least 99% identity. Identity is found over a comparison window of at least 20 base pairs, preferably over a comparison window of at least 50 base pairs, more preferably over a comparison window of at least 100 base pairs, and most preferably over the entire length of a polynucleotide of the disclosure.
[0031] Polynucleotide sequence identity can be determined by the following method: Compare the target nucleotide sequence to a candidate polynucleotide sequence using BLASTN (BLAST program suite, version 2.2.5 [November 2002]) bl2seq (Tatiana A. Tatusova, Thomas L. Madden (1999), "Blast 2 sequences—a new tool for comparing protein and nucleotide sequences," FEMS Microbiol. Lett. 174:247-250). This tool is publicly available from NCBI (ftp colon slash slash file transfer protocol.ncbi.nih.gov / blast / ). Use the default parameters of bl2seq, but turn off low-complexity filtering.
[0032] Polynucleotide sequence identity may be determined using the following Unix command line parameters: bl2seq -i nucleotideseql -j nucleotideseq2 -FF -p blastn The parameter -FF disables filtering of low complexity regions. The parameter -p selects the appropriate algorithm for pairwise alignment. The bl2seq program reports the number and percentage of identical nucleotides in the "Identities=" line.
[0033] Polynucleotide sequence identity may be calculated over the entire length of overlap between the candidate and subject sequences using a global sequence alignment program (e.g., Needleman, SB and Wunsch, CD (1970) J. Mol. Biol. 48, 443-453). A complete implementation of the Needleman-Wunsch global alignment algorithm can be found in the needle program of the EMBOSS package (Rice, P. Longden, I. and Bleasby, A. EMBOSS: The European Molecular Biology Open Software Suite, Trends in Genetics June 2000, vol. 16, No. 6, pp. 276-277). The European Bioinformatics Institute server provides the ability to perform an online global alignment of two sequences using EMBOSS-needle (worldwide web.ebi.ac.uk / emboss / align / ).
[0034] Alternatively, one may use the GAP program, which calculates the optimal global alignment of two sequences without penalizing terminal gaps. GAP is described in: Huang, X. (1994) On Global Sequence Alignment. Computer Applications in the Biosciences 10. 227-235.
[0035] The recommended method for calculating percent sequence identity of polynucleotides is based on aligning the sequences to be compared using Clustal X (Jeanmougin et al., 1998, Trends Biochem. Sci. 23, 403-5).
[0036] Polynucleotide variants of the present disclosure include those that exhibit similarity to one or more of the specifically identified sequences such that they maintain functional equivalence of those sequences and would not reasonably be expected to occur by random chance. Such sequence similarity to polypeptides may be determined using the bl2seq program (version 2.2.5 [November 2002]) of the BLAST suite of programs available from NCBI (ftp colon slash slash file transfer protocol ftp.ncbi.nih.gov / blast / ).
[0037] Polynucleotide sequence similarity may be examined using the following Unix command line parameters: bl2seq -i nucleotideseql -j nucleotideseq2 -FF -p tblastx The parameter -FF disables filtering of low-complexity sections. The parameter -p selects the appropriate algorithm for paired alignments. The program finds regions of similarity between sequences and reports an "E-value" for each region. The E-value is the expectation that such a match would be observed by chance in a reference database (containing random sequences) of a fixed size. The size of this database is set to the default in the bl2seq program. If the E-value is small (much smaller than 1), the E-value approximately corresponds to the probability of that random match.
[0038] A variant polynucleotide sequence has an E value of less than 1x10-6, more preferably less than 1x10-9, more preferably less than 1x10-12, more preferably less than 1x10-15, more preferably less than 1x10-18, more preferably less than 1x10-21, more preferably less than 1x10-30, more preferably less than 1x10-40, more preferably less than 1x10-50, more preferably less than 1x10-60, more preferably less than 1x10-70, more preferably less than 1x10-80, more preferably less than 1x10-90, and most preferably less than 1x10-100, when compared to any of the specifically specified sequences.
[0039] Alternatively, a variant polynucleotide of the disclosure hybridizes under stringent conditions to a specified polynucleotide sequence or its complement.
[0040] "Hybridize under stringent conditions" and grammatical equivalents refer to the ability of a polynucleotide molecule to hybridize to a target polynucleotide molecule (e.g., a target polynucleotide molecule immobilized on a DNA or RNA blot (e.g., a Southern blot or Northern blot) under defined conditions of temperature and salt concentration. The ability to hybridize under stringent hybridization conditions can be determined by first conducting hybridization under less stringent conditions and then increasing the stringency to the desired level.
[0041] For polynucleotide molecules longer than about 100 bases, typical stringent hybridization conditions are 25-30°C (e.g., 10°C) below the melting temperature (Tm) of the native duplex. (See generally Sambrook et al., Eds., 1987, Molecular Cloning, A Laboratory Manual, 2nd Ed., Cold Spring Harbor Press; Ausubel et al., 1987, Current Protocols in Molecular Biology, Greene Publishing.) For polynucleotide molecules longer than about 100 bases, Tm can be calculated using the following formula: Tin = 81.5 + 0.41 × (G + C) - log(Na+). (Sambrook et al., Eds., 1987, Molecular Cloning, A Laboratory Manual, 2nd Ed., Cold Spring Harbor Press; Bolton and McCarthy, 1962, PNAS 84:1390). Typical stringent conditions for polynucleotides longer than 100 bases include a prewash in 6x Saline-Sodium Citrate (SSC)-0.2% SDS solution, overnight hybridization at 65°C in 6x SSC, 0.2% SDS, followed by two 30-minute washes at 65°C in 1x SSC, 0.1% SDS, and two additional 30-minute washes at 65°C in 0.2x SSC, 0.1% SDS.
[0042] For polynucleotide molecules less than 100 bases long, typical stringent hybridization conditions are 5-10°C below the Tm. On average, the Tm of a polynucleotide molecule less than 100 bp is reduced by approximately (500 / oligonucleotide length)°C.
[0043] For DNA mimics called peptide nucleic acids (PNAs) (Nielsen et al., Science. 1991 Dec. 6; 254(5037):1497-500), the Tm is higher than that of DNA-DNA or DNA-RNA hybrids and can be calculated using the formula described by Giesen et al. (Nucleic Acids Res. 1998 Nov. 1; 26(21):5004-6). Typical stringent hybridization conditions for DNA-PNA hybrids with a total length of less than 100 bases are temperatures 5-10°C below the Tm.
[0044] Variant polynucleotides of the present disclosure also include polynucleotides that encode polypeptides that differ from the sequences of the present disclosure but that, as a result of the degeneracy of the genetic code, have similar activity to the polypeptides encoded by the polynucleotides of the present disclosure. Sequence changes that do not change the amino acid sequence of the polypeptide are "silent mutations." Other codons that encode the same amino acids, with the exception of methionine (ATG) and tryptophan (TGG), can be altered using techniques recognized by those of skill in the art, for example, to optimize codon expression in a particular host organism.
[0045] Mutations in polynucleotide sequences that conservatively substitute one or more amino acids in the encoded polypeptide sequence without significantly altering its biological activity are also included in the present disclosure. Those skilled in the art are aware of methods for making phenotypically silent amino acid substitutions (e.g., Bowie et al., 1990, Science 247, 1306).
[0046] Variant polynucleotides due to silent mutations and conservative substitutions may be determined using the tblastx algorithm of the bl2seq program (version 2.2.5 [November 2002]) of the BLAST suite of programs available from NCBI (ftp: / / ncbi.nih.gov / blast / ), as described above.
[0047] The functionality of variant polynucleotides of the present disclosure may be assessed by cloning such sequences into bacteria and testing the activity of the encoded proteins, e.g., as described in the Examples. The functionality of variants can also be tested for their ability to alter GGP activity, ascorbic acid content, desirable properties, or the content of useful compounds, nutrients, or components in plants, as described in the Examples herein. (Liang et al., U.S. Patent No. 9,648,813)
[0048] [Polypeptide variants] The term "variant" with respect to a polypeptide includes naturally occurring, recombinantly, or chemically synthesized polypeptides. A variant 4444 polypeptide sequence is preferably at least 30%, preferably at least 50%, more preferably at least 51%, more preferably at least 52%, more preferably at least 53%, more preferably at least 54%, more preferably at least 55%, more preferably at least 56%, more preferably at least 57%, more preferably at least 58%, more preferably at least 59%, more preferably at least 60%, more preferably at least 61%, more preferably at least 62%, more preferably at least 63%, more preferably at least 64%, more preferably at least 65%, more preferably at least 66%, more preferably at least 67%, more preferably at least 68%, more preferably at least 69%, more preferably at least 70%, more preferably at least 71%, more preferably at least 72%, more preferably at least 73%, more preferably at least 74%, more preferably at least 75%, more preferably at least 76%, more preferably at least 77%, more preferably at least 78%, more preferably at least 79% identical to the sequences of the present disclosure. More preferably, it has at least 80%, more preferably at least 81%, more preferably at least 82%, more preferably at least 83%, more preferably at least 84%, more preferably at least 85%, more preferably at least 86%, more preferably at least 87%, more preferably at least 88%, more preferably at least 89%, more preferably at least 90%, more preferably at least 91%, more preferably at least 92%, more preferably at least 93%, more preferably at least 94%, more preferably at least 95%, more preferably at least 96%, more preferably at least 97%, more preferably at least 98%, and most preferably at least 99% identity.Identity is found in a comparison window that includes at least 20 amino acid residue positions, preferably in a comparison window that includes at least 50 amino acid residue positions, more preferably in a comparison window that includes at least 100 amino acid residue positions, and most preferably over the entire length of the polypeptides herein.
[0049] Polypeptide sequence identity can be determined as follows: A subject polypeptide sequence is compared to a candidate polypeptide sequence using BLASTP, included in version 2.2.5 (November 2002) of the BLAST suite of programs in bl2seq, publicly available from NCBI (ftp: / / ftp.ncbi.nih.gov / blast / ), using the default parameters of bl2seq, but turning off filtering of low-complexity regions.
[0050] Polypeptide sequence identity can also be calculated across the overlap between candidate and subject polynucleotide sequences using a global sequence alignment program. As described above, EMBOSS-needle (world wide web.ebi.ac.uk / emboss / align / ) and GAP (Huang, X. (1994) On Global Sequence Alignment. Computer Applications in the Biosciences 10, 227-235.) are suitable global sequence alignment programs for calculating polypeptide sequence identity.
[0051] A preferred method for calculating % sequence identity of polypeptides is based on aligning the sequences to be compared using Clustal X (Jeanmougin et al. 1998. Trends Biochem. Sci. 23, 403-5.).
[0052] Polypeptide variants of the present disclosure include those that exhibit similarity to one or more of the specifically identified sequences, may maintain functional equivalence of those sequences, and would not reasonably be expected to occur by random chance. Such sequence similarity for polypeptides can be determined using the bl2seq program (version 2.2.5 [November 2002]) from the BLAST suite of programs available from NCBI (ftp: / / ncbi.nih.gov / blast / ). Polypeptide sequence similarity can be determined using the following Unix command line parameters: bl2seq -i peptideseql -j peptideseq2 -FF - p blastp.
[0053] Preferably, the variant polypeptide sequence has an E value of less than 1x10-6, more preferably less than 1x10-9, more preferably less than 1x10-12, more preferably less than 1x10-15, more preferably less than 1x10-18, more preferably less than 1x10-21, more preferably less than 1x10-30, more preferably less than 1x10-40, more preferably less than 1x10-50, even more preferably less than 1x10-60, even more preferably less than 1x10-70, even more preferably less than 1x10-80, even more preferably less than 1x10-90, and most preferably less than 1x10-100, when compared to any of the specifically specified sequences.
[0054] The parameter -FF disables filtering of low complexity sections. The parameter -p selects the appropriate algorithm for paired sequences. The program finds regions of similarity between sequences and reports an "E-value" for each region. The E-value is the expected number of times such a match would be found by chance in a reference database of fixed size (containing random sequences). If the E-value is small, i.e., much smaller than 1, this approximates the probability of such a random match.
[0055] Variant polypeptides include polypeptides whose amino acid sequence differs from that of the polypeptide by one or more conservative amino acid substitutions, deletions, additions, or insertions that do not abolish the biological activity of the polypeptide. Conservative substitutions typically involve substitutions of amino acids with similar properties, such as substitutions within the following groups: valine, glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid; asparagine, glutamine; serine, threonine; lysine, arginine; phenylalanine, tyrosine.
[0056] Non-conservative substitutions involve exchanging a member of one of these classes for a member of another class.
[0057] Analysis of evolved biological sequences shows that not all sequence changes occur with equal probability, and that at least some reflect differences between conservative and non-conservative substitutions at the biological level. For example, some amino acid substitutions occur frequently, while others are very rare. Evolutionary changes and substitutions of amino acid residues can be modeled using scoring matrices, also known as substitution matrices. Such matrices are used in bioinformatics analyses to identify relationships between sequences, and a representative example is the BLOSUM62 matrix shown below (Table 1).
[0058] [Table 1]
[0059] The BLOSUM62 matrix shown is used to generate a score for each aligned amino acid pair found at the intersection of the corresponding column and row. For example, a substitution score of 2 is obtained for a glutamic acid residue (E) to an aspartic acid residue (D). The diagonal line displays the scores for unchanged amino acids. Most substitution changes have negative scores. This matrix contains only integers.
[0060] It is believed to be within the skill of one in the art to determine an appropriate scoring matrix to produce an optimal alignment for a given set of sequences. The BLOSUM62 matrix in Table 1 is used as the default matrix in BLAST searches, but is not intended to be limiting.
[0061] Other variants include peptides with modifications that affect the stability of the peptide. Such analogs may include, for example, those that contain one or more non-peptide bonds replacing the peptide bonds in the peptide sequence. Also included are analogs that contain residues other than naturally occurring L-amino acids, such as D-amino acids, non-naturally occurring synthetic amino acids (beta or gamma amino acids), and cyclic analogs.
[0062] The function of a polypeptide variant as a GGP may be assessed by the methods described in the Examples herein.
[0063] The function of a polypeptide variant as a GDP-D-mannose epimerase can be assessed by the method described in the Examples section of this specification (Liang et al., U.S. Patent No. 9,648,813).
[0064] [Constructs, vectors and their components] "Genetic construct" or "construct" refers to a polynucleotide molecule, usually double-stranded DNA, into which another polynucleotide molecule (an insert polynucleotide molecule) may be inserted. Insert polynucleotide molecules include, but are not limited to, cDNA molecules.
[0065] A genetic construct may contain the necessary elements to allow transcription of an inserted polynucleotide molecule and, optionally, translation of the transcript into a polypeptide. The inserted polynucleotide molecule may be derived from the host cell, from a different cell or organism, or may be a recombinant polynucleotide. When introduced into a host cell, the genetic construct may be integrated into the host's chromosomal DNA. The genetic construct may be linked to a vector.
[0066] The term "vector" refers to a polynucleotide molecule (usually double-stranded DNA) used to transport genetic constructs into a host cell. The vector may be in the form of a plasmid and / or may be replicable in at least one additional host system (e.g., E. coli).
[0067] The term "expression construct" refers to a genetic construct containing the necessary elements that allow for the transcription of an inserted polynucleotide molecule, and optionally, the translation of the transcript into a polypeptide. Expression constructs generally contain, in 5' to 3' direction, the following elements: (1) a promoter that functions in the host cell into which the construct will be transformed; (2) the polynucleotide to be expressed, and (3) A terminator that is functional in the host cell into which the construct will be transformed.
[0068] The term "coding region" or "open reading frame" (ORF) refers to the sense strand of a genomic DNA or cDNA sequence capable of producing a transcript and / or polypeptide under the control of appropriate regulatory sequences. A coding sequence is identified by the presence of a 5' translation start codon and a 3' translation stop codon. When inserted into a genetic construct, a "coding sequence" is expressible when it is operably linked to promoter and terminator sequences.
[0069] "Operably linked" means that a sequence of interest, such as a sequence to be expressed, is under the control of other sequences, including regulatory elements, including promoters, tissue-specific regulatory elements, temporal regulatory elements, enhancers, repressors, terminators, 5'-UTR sequences, 5'-UTR sequences containing uORFs, and uORFs.
[0070] In a preferred embodiment, the regulatory element comprises a polynucleotide sequence of the present disclosure.
[0071] Preferably, the sequences of the present disclosure comprise a 5'-UTR sequence. Preferably, the 5'-UTR sequence comprises a uORF.
[0072] The term "untranslated region" refers to the untranslated sequences located upstream of the translation start site and downstream of the translation stop site. These sequences are also called 5'-UTR and 3'-UTR, respectively. These regions contain elements necessary for transcription initiation and termination, as well as elements necessary for regulating translation efficiency.
[0073] The 5'-UTR sequence is the sequence located between the transcription start site and the translation start site of the major ORF.
[0074] A 5'-UTR sequence is an mRNA sequence encoded by genomic DNA, although as used herein the term 5'-UTR sequence is intended to include the genomic DNA sequence encoding the 5'-UTR sequence, the complement of that genomic sequence, and the 5'-UTR mRNA sequence.
[0075] Terminators are sequences that terminate transcription and are found at the 3' untranslated end of genes downstream of the translated sequence. Terminators are important determinants of mRNA stability and, in some cases, have been found to have spatial regulatory functions.
[0076] The terms "altering expression" and "altered expression" of a polynucleotide or polypeptide of the present disclosure are intended to encompass situations in which genomic DNA corresponding to a polynucleotide of the present disclosure is modified, resulting in altered expression or levels of the polynucleotide or polypeptide of the present disclosure. Polynucleotide or polypeptide levels can be measured by any of a range of techniques, including transcript profiling, RT-PCR, Northern blot, Western blot, microarray analysis, RNASeq, or fluorescence or luminescence measurements. Modification of genomic DNA may be performed by transformation or other methods known to those skilled in the art to introduce genetic mutations. "Altered expression" may relate to an increase or decrease in the amount of messenger RNA and / or polypeptide, and may also result in altered activity of the polypeptide due to changes in the sequence of the polynucleotide and polypeptide.
[0077] An "introduced targeted gene modification" or "targeted gene modification" refers to a change in the DNA sequence at a specific chromosomal location (also called a locus) within the genome of a plant, selected by a specialist, such as a plant breeder or molecular biologist, and introduced by a process of gene editing and / or selection using a specific complementary nucleic acid molecule sequence as a guide or probe.
[0078] A "native genomic locus" or "endogenous locus" refers to a gene or DNA sequence present at a specific chromosomal location within the genome of a wild-type plant of a particular species. A "native genomic locus" typically refers to the region from the start codon to the stop codon, including intervening introns, as well as upstream regulatory elements, including promoter regions and uORFs, which control the activity of mORFs. This region is transcribed to produce a major ORF that encodes a long polypeptide, typically greater than about 100 amino acids in length. uORFs are contained in the same mRNA transcript as the mORF they regulate. Therefore, uORFs and mORFs can be considered part of the same native genomic locus. Native genomic loci are often designated by their GenBank accession number, which indicates, for example, the DNA sequence present at that location and the encoded polypeptide. Furthermore, a locus may encode multiple protein variants due to alternative splicing of mRNA; these variants are represented by different "gene models," denoted by the accession number followed by a dot and a number. (Liang et al. US Pat. No. 9.648.813)
[0079] A "non-naturally occurring allele of a gene" refers to a sequence variant of a gene (where "gene" includes the protein-coding region and upstream regulatory elements such as uORFs) in a particular plant species that has arisen through human intervention, such as gene editing or selection, and / or a nucleotide sequence that does not naturally occur in the genome of a wild-type plant of that species or in the genome of a plant taken from a naturally occurring wild population of that species.
[0080] As used herein, the term "variant" refers to a polynucleotide or polypeptide that differs in sequence from the presently disclosed polynucleotides or polypeptides, respectively, and falls within the ranges defined above and below.
[0081] With respect to polynucleotide variants, differences between the presently disclosed polynucleotides and polynucleotide variants are limited so that the nucleotide sequences of the former and latter are overall closely similar and, in many regions, identical. Due to the degeneracy of the genetic code, differences between the former and latter nucleotide sequences may be silent (i.e., the amino acids encoded by the polynucleotides are the same, and the variant polynucleotide sequence encodes the same amino acid sequence as the presently disclosed polynucleotide). The variant nucleotide sequence may encode a different amino acid sequence, in which case such nucleotide differences result in amino acid substitutions, additions, deletions, insertions, truncations, or combinations relative to the analogous disclosed polynucleotide sequence. These variations result in polynucleotide variants encoding polypeptides that share at least one functional characteristic. Due to the degeneracy of the genetic code, many different variant polynucleotides can encode identical or substantially similar polypeptides in addition to the sequences set forth in the sequence listing.
[0082] Also included within the scope of this disclosure are variants of the nucleic acids set forth in the Sequence Listing, i.e., those having sequences that differ from or are complementary to any of the polynucleotide sequences set forth in the Sequence Listing and that encode functionally equivalent polypeptides (i.e., polypeptides having some degree of the same or similar biological activity), but which differ in sequence from the sequences set forth in the Sequence Listing due to the degeneracy of the genetic code. This definition includes polymorphisms, whether or not readily detectable using specific oligonucleotide probes for detecting polypeptides encoded by polynucleotides, as well as inappropriate or unexpected hybridization of polynucleotides and allelic variants at loci other than the normal chromosomal locus of the polypeptide encoded by the polynucleotide sequence.
[0083] The term "plant" includes the entire plant, stem vegetative organs / structures (e.g., leaves, stems, tubers), roots, rhizomes, flowers, and floral organs / structures (e.g., bracts, sepals, petals, stamens, pistils, anthers, ovules), seeds (including embryos, endosperms, and seed coats), fruits (mature ovaries), plant tissues (e.g., vascular tissue, ground tissue, etc.), and cells (e.g., stomatal cells, egg cells, etc.), and their progeny. Plant classifications usable in the methods of the present disclosure are generally equivalent to classifications of higher and lower plants amenable to transformation techniques, and include angiosperms (monocotyledons and dicotyledons), gymnosperms, femus, horsetails, syllophie, lycophytes, bryophytes, and multicellular algae. See, e.g., Daly et al. (2001) Plant Physiol. 127: 1328-1333; Ku et al. (2000) Proc. Natl. Acad. Sci. 97: 9121-9126; and Tudge, The Variety of Life, Oxford University Press, New York, NY (2000) pp. 547-606.
[0084] A "trait" refers to a physiological, morphological, biochemical, or physical characteristic of a cell or organism (including a plant, a particular plant material, or a plant cell). Some traits are visible to the human eye, such as seed size, plant size, or pigmentation; biochemical techniques, such as detecting seed protein, starch, or oil content; observations of metabolic or physiological processes, such as measuring carbon dioxide uptake; observations of gene expression levels, such as using Northern analysis, RT-PCR, microarray gene expression assays, RNA sequencing, or reporter gene expression systems; or agronomic observations, such as stress tolerance or pathogen resistance. However, any technique can be used to measure amounts, comparative levels, or differences in selected compounds or macromolecules in transformed plants.
[0085] "Trait modification" refers to a detectable difference in a trait in a plant that ectopically expresses a polynucleotide or polypeptide of the present disclosure compared to a plant that does not ectopically express the polynucleotide or polypeptide (e.g., a wild-type plant). In some cases, the trait modification can be quantitatively assessed. For example, the trait modification can include an increase or decrease of at least about 2%, a difference of at least about 5%, a difference of at least about 10%, a difference of at least about 20%, a difference of at least about 30%, at least about 50%, at least about 70%, at least about 85%, or about 100% in the observed trait (difference). Or, it can include an even greater difference compared to a wild-type plant. Natural variation in modified traits is known to exist. Thus, the observed trait modification involves a change in the normal distribution of the trait in the plant compared to the distribution observed in wild-type plants.
[0086] As used herein, "wild-type" refers to a cell, tissue, or plant in which one or more of the presently disclosed transcription factors have not been genetically modified to delete or overexpress them. Wild-type cells, tissues, or plants can be used as controls to compare expression levels and the extent and nature of trait modifications with cells, tissues, or plants in which transcription factor expression has been altered or ectopically expressed (e.g., knocked out or overexpressed). A wild-type allele refers to the DNA sequence at a particular endogenous locus within the genome of a wild-type plant.
[0087] "Crop" plants include cultivated plants or agricultural products, including grains, vegetables, fruit trees, ornamental plants, or plants (generally classified as groups) used for energy or material production. Crop plants include those grown in commercially useful numbers or quantities.
[0088] [Description of specific embodiments] The present disclosure is directed to a method for improving, increasing, or enhancing a desirable trait in a plant. The method utilizes a nucleic acid comprising a 5' UTR (untranslated region) with an upstream open reading frame (uORF) from the upstream region of a gene of interest. The nucleic acid is linked to a reporter gene to create a 5' UTR:reporter polynucleotide, which is used to identify useful mutations that can be introduced into a uORF present at an endogenous locus in the plant genome by gene editing or allelic selection to alter the activity of a downstream protein-coding ORF that the uORF regulates. Typically, the selected uORF mutation increases the level of protein produced from the downstream ORF, resulting in a plant that expresses a desirable trait, such as tolerance to abiotic or biotic stress. This approach begins by introducing a mutation into a nucleic acid within the uORF region or the sequence of the 5' UTR. The mutation can be introduced by a variety of known means, including direct DNA synthesis to generate the 5' UTR:reporter polynucleotide. It can also be generated by other means, including but not limited to, radiation, photomutagenesis, chemical mutagenesis, DNA repair enzymes, or gene editing. The 5'UTR::reporter polynucleotide is introduced into test cells or protoplasts to examine its activity. The cells may be derived from plants or algae, although in some cases, microbial cells such as yeast may be used. The amount of reporter gene product in the cells is measured, and this amount may be used to identify mutations in which the amount of reporter product in the plant cells is greater or less than the amount of product in control plant cells. The control plant cells may contain, for example, a control 5'UTR::reporter polynucleotide in which the uORF is inactivated by completely or partially deleting the uORF from the 5'UTR linked to the reporter gene. The mutation that increases the level of reporter product in the reporter system described above is then introduced into the genome of a crop plant of interest or cells of a target crop plant of interest (or a homologous uORF).Mutations are typically introduced by selecting individual plants or cells carrying the mutation from a population (by molecular genotyping, e.g., PCR followed by enzyme digestion and sequencing of the PCR products) or by gene editing. Cells or plants carrying the mutation are regenerated, if necessary, and whole plants are regenerated and tested to confirm that they retain the desired trait. Typically, multiple plants (lines or events) carrying the mutation are generated, and one or more are selected for experimental or commercial use based on the strength of the desired trait they express.
[0089] The present disclosure also relates to plants and constructs comprising the 5'UTR::reporter polynucleotide having a mutation in the 5'uORF.
[0090] In certain embodiments of the above methods, the gene of interest encodes a GGP protein, the level of which, when increased, has a trait including increased levels of ascorbic acid in the plant, thereby providing an increased level of protection against abiotic or biotic stress, resulting in a plant exhibiting the trait of increased stress tolerance, which is typically associated with increased growth or vigor, or altered stem architecture, increased stem number, decreased levels of anthocyanins, reduced tissue damage, or reduced wilting compared to a control plant.
[0091] In one aspect, the present disclosure provides an isolated polynucleotide comprising a sequence encoding an amino acid sequence selected from SEQ ID NOs: 1-20 and 132-134 (uORF peptides), or a variant or fragment thereof.
[0092] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to an amino acid sequence selected from SEQ ID NOs: 21-40 and 135-137 (conserved regions of uORF peptides).
[0093] In a further embodiment, the variant or fragment comprises an amino acid sequence selected from SEQ ID NOs: 21-40 and 135-137 (conserved regions of uORF peptides). In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to an amino acid sequence selected from SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions of uORF peptides in dicotyledonous plants). In a further embodiment, the variant or fragment comprises an amino acid sequence selected from SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions of uORF peptides in dicotyledonous plants).
[0094] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to the amino acid sequence of SEQ ID NO: 108 (consensus motif uORF peptide). In a further embodiment, the variant or fragment comprises the amino acid sequence SEQ ID NO: 108 (consensus motif uORF peptide).
[0095] In a further embodiment, the variant comprises a sequence (uORF peptide) having at least 70% identity to an amino acid sequence selected from SEQ ID NOs: 1-20 and 132-134. In a further embodiment, the variant comprises an amino acid sequence (uORF peptide) selected from SEQ ID NOs: 1-20 and 132-134.
[0096] In a further embodiment, the variant comprises a sequence having at least 70% identity to an amino acid sequence selected from SEQ ID NOs: 1-10, 13-17 and 132-134 (dicotyledonous plant uORF peptide). In a further embodiment, the variant comprises an amino acid sequence selected from SEQ ID NOs: 1-10, 13-17 and 132-134 (dicotyledonous plant uORF peptides).
[0097] In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 61-80 and 138-140 (conserved regions of uORF DNA sequences).
[0098] In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 61-80 and 138-140 (conserved regions of uORF DNA sequences). In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 61-70, 73-77 and 138-140 (a conserved region of dicotyledonous plant uORF DNA sequences). In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 61-70, 73-77 and 138-140 (conserved regions of dicotyledonous plant uORF DNA sequences).
[0099] In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 41-60 and 129-131 (uORF DNA sequence). In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 41-60 and 129-131 (uORF DNA sequence).
[0100] In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 41-50, 53-57 and 129-131 (dicotyledonous plant uORF DNA sequences). In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 41-50, 53-57 and 129-131 (dicotyledonous plant uORF DNA sequences).
[0101] In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 111-125 (5'UTR subsequence). In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 111-125 (5'UTR subsequence).
[0102] In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 81-100 and 126-128 (the entire 5'UTR sequence). In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 81-100 and 126-128 (the entire 5'UTR sequence).
[0103] In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 81-90, 93-97 and 125-128 (all 5'UTR sequences of dicotyledonous plants). In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 81-90, 93-97 and 125-128 (the entire 5'UTR sequence of a dicotyledonous plant).
[0104] In yet another aspect, the present disclosure provides an isolated polynucleotide comprising a sequence selected from SEQ ID NOs: 41-60 and 129-131 (uORF DNA sequence), or a variant or fragment thereof.
[0105] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 61-80 and 138-140 (conserved regions of uORF DNA sequences).
[0106] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 61-70, 73-77 and 138-140 (conserved regions of dicotyledonous plant uORF DNA sequences).
[0107] In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 41-60 and 129-131 (uORF DNA sequence). In one embodiment, the variant comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 41-60 and 129-131 (uORF DNA sequence).
[0108] In one embodiment, the variant comprises a sequence having at least 70% identity with a sequence selected from SEQ ID NOs: 41-50, 53-57 and 129-131 (dicotyledonous plant uORF DNA sequences).
[0109] In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 111-125 (5'UTR subsequence). In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 111-125 (5'UTR subsequence).
[0110] In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 81-100 and 126-128 (the entire 5'UTR sequence). In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 81-100 and 126-128 (the entire 5'UTR sequence).
[0111] In a further embodiment, the isolated polynucleotide comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 81-90, 93-97 and 126-128 (all 5'UTR sequences of dicotyledonous plants). In a further embodiment, the isolated polynucleotide comprises a sequence selected from SEQ ID NOs: 81-90, 93-97 and 125-128 (the entire 5'UTR sequence of a dicotyledonous plant).
[0112] In a further aspect, the present disclosure provides an isolated polynucleotide having a sequence selected from SEQ ID NOs: 81-100 and 126-128 (complete 5'UTR sequences of dicotyledonous plants), or a variant or fragment thereof. In one embodiment, the variant comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 81-100 and 126-128 (all 5'UTR sequences of dicotyledonous plants).
[0113] In a further aspect, the present disclosure provides an isolated polynucleotide having a sequence selected from SEQ ID NOs: 111-125 (5'UTR subsequences), or a variant or fragment thereof. In one embodiment, the variant comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 111-125 (5'UTR subsequence).
[0114] In a further embodiment, the variant or fragment comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 41-60 and 129-131 (uORF DNA sequences). In a further embodiment, the variant or fragment comprises a sequence selected from SEQ ID NOs: 41-60 and 129-131 (uORF DNA sequences).
[0115] In a further embodiment, the variant or fragment comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 41-50, 53-57 and 129-131 (dicotyledonous plant uORF DNA sequences). In a further embodiment, the variant or fragment comprises a sequence selected from SEQ ID NOs: 41-50, 53-57 and 129-131 (dicotyledonous plant uORF DNA sequences).
[0116] In a further embodiment, the variant or fragment comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 61-80 and 138-140 (conserved regions of uORF DNA sequences). In a further embodiment, the variant or fragment comprises a sequence selected from SEQ ID NOs: 61-80 and 138-140 (conserved regions of uORF DNA sequences).
[0117] In a further embodiment, the variant or fragment comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 61-70, 73-77 and 138-140 (conserved regions of uORF DNA sequences in dicotyledonous plants). In a further embodiment, the variant or fragment comprises a sequence selected from SEQ ID NOs: 61-70, 73-77 and 138-140 (conserved regions of uORF DNA sequences in dicotyledonous plants).
[0118] In a further embodiment, the variant encodes a sequence having at least 70% identity to at least one of SEQ ID NOs: 21-40 and 135-137 (conserved regions of the uORF peptides). In a further embodiment, the variant encodes a sequence selected from at least one of SEQ ID NOs: 21-40 and 135-137 (conserved regions of uORF peptides).
[0119] In a further embodiment, the variant encodes a sequence having at least 70% identity to at least one of SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions of uORF peptides in dicotyledonous plants). In a further embodiment, the variant encodes a sequence selected from at least one of SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions of uORF peptides in dicotyledonous plants).
[0120] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to the amino acid sequence of SEQ ID NO: 108 (the consensus motif of the uORF peptide). In a further embodiment, the variant or fragment comprises the amino acid sequence of SEQ ID NO: 108 (the uORF peptide consensus motif).
[0121] In a further embodiment, the variant encodes a sequence having at least 70% identity to at least one of SEQ ID NOs: 1-20 and 132-134 (uORF peptides). In a further embodiment, the variant or fragment encodes a sequence selected from at least one of SEQ ID NOs: 1-20 and 132-134 (uORF peptides).
[0122] In a further embodiment, the variant encodes a sequence having at least 70% identity to at least one of SEQ ID NOs: 1-10, 13-17 and 132-134 (dicotyledonous plant uORF peptides). In a further embodiment, the variant or fragment encodes a sequence selected from at least one of SEQ ID NOs: 1-10, 13-17 and 132-134 (dicotyledonous plant uORF peptides).
[0123] In one embodiment, the isolated polynucleotide is modified. In one embodiment, the modification is at least one of a deletion, addition, or substitution of at least one nucleotide in the sequence encoding the 5'-UTR. In one embodiment, the modification reduces, inhibits, or prevents translation of a uORF polypeptide having the sequence of SEQ ID NOs: 1-20 and 132-134 (uORF peptides) or a variant thereof. In a further embodiment, the modification reduces, inhibits, or eliminates the activity of a uORF polypeptide having the sequence of SEQ ID NOs: 1-20 and 132-134 (uORF peptides) or a variant thereof.
[0124] In one embodiment, the variant comprises a sequence having at least 70% identity to the sequence of any one of SEQ ID NOs: 1-20 and 132-134 (uORF peptides). In a further embodiment, the variant comprises a sequence having at least 70% identity to the sequence of any one of SEQ ID NOs: 1-10, 13-17 and 132-134 (dicotyledonous plant uORF peptides). In a further embodiment, the variant comprises a sequence having at least 70% identity to at least one of SEQ ID NOs: 21-40 and 135-137 (conserved regions of uORF peptides). In a further embodiment, the variant comprises at least one sequence of SEQ ID NOs: 21-40 and 135-137 (conserved regions of uORF peptides). In a further embodiment, the variant comprises a sequence having at least 70% identity to at least one of SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions of uORF peptides in dicotyledonous plants). In a further embodiment, the variant comprises at least one sequence of SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions of uORF peptides in dicotyledonous plants).
[0125] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to the amino acid sequence of SEQ ID NO: 108 (the consensus motif of the uORF peptide). In a further embodiment, the variant or fragment comprises the amino acid sequence of SEQ ID NO: 108 (the uORF peptide consensus motif).
[0126] In one embodiment, the polynucleotide, or a variant or fragment thereof, is operably linked to a nucleotide sequence of interest.
[0127] In a further embodiment, the nucleotide sequence of interest encodes a protein of interest. In one embodiment, the polynucleotide and the nucleotide sequence of interest are not normally associated in nature.
[0128] When the polynucleotide is modified as described above to inhibit the expression or activity of a uORF polypeptide, the operably linked sequence may be a GGP sequence. In this embodiment, the modification eliminates uORF-mediated repression through an ascorbic acid-mediated feedback loop. By expressing GGP under the control of the modified polynucleotide, it is possible to avoid negative regulation of expression by ascorbic acid via the uORF polypeptide while maintaining the spatial and / or temporal expression pattern of GGP similar to that of the endogenous GGP promoter and 5'-UTR. In this embodiment, the polynucleotide and the nucleotide sequence of interest may be those normally associated in nature, but the polynucleotide is in a modified form as described above.
[0129] [Polypeptide] In a further aspect, the disclosure provides an isolated polypeptide comprising a sequence selected from SEQ ID NOs: 1-20 and 132-134 (uORF peptides), or a variant or fragment thereof.
[0130] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 21-40 and 135-137 (conserved regions of uORF peptides). In a further embodiment, the variant or fragment comprises a sequence selected from SEQ ID NOs: 21-40 and 135-137 (conserved regions of uORF peptides).
[0131] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions of uORF peptides in dicotyledonous plants). In a further embodiment, the variant or fragment comprises a sequence selected from SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions of uORF peptides in dicotyledonous plants).
[0132] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to the amino acid sequence of SEQ ID NO: 108 (the consensus motif of the uORF peptide). In a further embodiment, the variant or fragment comprises the amino acid sequence of SEQ ID NO: 108 (the uORF peptide consensus motif).
[0133] In a further embodiment, the variant comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 1-20 and 132-134 (uORF peptides).
[0134] In a further embodiment, the variant comprises a sequence having at least 70% identity to a sequence selected from SEQ ID NOs: 1-10, 13-17 and 132-134 (dicotyledonous plant uORF peptides).
[0135] [construct] In a further embodiment, the present disclosure provides a construct comprising a polynucleotide of the present disclosure. In one embodiment, the polynucleotide is operably linked to a nucleotide sequence of interest. In a further embodiment, the polynucleotide and the nucleotide sequence of interest are not normally associated in nature. In a further embodiment, the nucleotide sequence of interest encodes a protein of interest.
[0136] [Polynucleotide activity] In one embodiment, the polynucleotides of the present disclosure are controllable by a compound. In this embodiment, the expression of any nucleotide sequence operably linked to a polynucleotide of the present disclosure is controlled by the compound. Preferably, the regulation occurs at the post-transcriptional level. Preferably, expression of the polypeptide encoded by the operably linked nucleotide sequence is controlled by the compound.
[0137] In one embodiment, expression of the operably linked nucleotide sequence is controlled by interaction of the compound with a uORF peptide expressed by the polynucleotide of the present disclosure. In one embodiment, the interaction is direct. In a further embodiment, the interaction is indirect. In a further embodiment, the indirect interaction is via another protein.
[0138] In one embodiment, the compound is ascorbic acid or a related metabolite. Preferably, the compound is ascorbic acid.
[0139] When modified as described above, the polynucleotide may no longer be regulated by the compound. An example of this embodiment is the expression of a GGP coding sequence under the control of a modified polynucleotide sequence. In this embodiment, the modification reduces or eliminates the inhibitory effect of the compound. If the compound is ascorbic acid, this means that GGP translational repression is lost, resulting in increased GGP production and ascorbic acid accumulation. In this embodiment, the uORF or its encoding sequence may be modified in terms of the promoter and 5'-UTR sequences. Examples of the GGP promoter and entire 5'-UTR sequences are provided in SEQ ID NOs: 101-107 or variants thereof. Using a uORF with a modified promoter and 5'-UTR sequence can avoid ascorbic acid-mediated repression of GGP translation while retaining spatial or temporal expression of a portion of the GGP sequence.
[0140] [cell] In a further embodiment, the present disclosure provides a cell comprising a polynucleotide of the present disclosure or a construct of the present disclosure. Preferably, the cell or its progenitor cell has been genetically modified to contain a polynucleotide of the present disclosure or a construct of the present disclosure. Preferably, the cell or a progenitor cell thereof has been transformed to contain a polynucleotide of the present disclosure or a construct of the present disclosure.
[0141] [Plant cells and plants] In a further embodiment, the present disclosure provides a plant cell or plant comprising a polynucleotide or construct of the present disclosure. Preferably, the plant cell or plant, or a progenitor plant cell or plant thereof, has been genetically modified to contain a polynucleotide of the present disclosure or a construct of the present disclosure. Preferably, the plant cell or plant, or a progenitor plant cell or plant thereof, has been transformed to contain a polynucleotide of the present disclosure or a construct of the present disclosure.
[0142] [Plant parts or propagules] In a further embodiment, the present disclosure provides a plant part or propagule comprising a polynucleotide or construct of the present disclosure. Preferably, the plant part or propagule, or a progenitor plant cell or plant thereof, has been genetically modified to include a polynucleotide of the present disclosure or a construct of the present disclosure. Preferably, the plant part or propagule, or a progenitor plant cell or plant thereof, has been transformed to contain a polynucleotide of the present disclosure or a construct of the present disclosure.
[0143] [Methods for expression control] In a further aspect, the present disclosure provides a method of controlling or regulating the expression of at least one nucleotide sequence in a cell, the method comprising introducing into the cell a polynucleotide or construct of the present disclosure. In a further aspect, there is provided a method of regulating the expression of at least one nucleotide sequence in a plant cell or plant, the method comprising the step of introducing into said plant cell or plant a polynucleotide or construct of the present disclosure. In a further aspect, the present disclosure provides a method of producing a cell or plant cell / plant with modified gene expression, the method comprising introducing into said cell or plant cell / plant a polynucleotide or construct of the present disclosure. In a further aspect, a method of modifying the phenotype of a plant is provided, the method comprising stably introducing a polynucleotide or construct of the present disclosure into the genome of said plant.
[0144] Those skilled in the art will understand that introducing a polynucleotide of the present disclosure into a cell, plant cell, or plant can control the expression of a nucleotide sequence operably linked to the polynucleotide, and the polynucleotide and the nucleotide sequence of interest can be located on a construct of the present disclosure and introduced simultaneously. In another embodiment, a polynucleotide of the present disclosure is inserted into a genome and controls the expression of nucleotide sequences (eg, protein-coding sequences) adjacent to the insertion site. In a preferred embodiment, the cell, plant cell or plant produces a compound whose expression is controlled via a polynucleotide of the present disclosure or a uORF polypeptide encoded thereby. Alternatively, the compound may be administered to a cell, plant cell or plant. In a further aspect, there is provided a plant cell or plant obtained by the method of the present disclosure.
[0145] [Method for producing plant cells or plants with increased ascorbic acid production] In one aspect, the present disclosure provides a method for increasing ascorbic acid production in a plant cell or plant, the method comprising modifying the 5'-UTR of the GGP gene in said plant cell or plant.
[0146] In one embodiment, the 5'-UTR is in terms of a polynucleotide sequence derived from SEQ ID NO: 101-107 (GGP genomic sequence including promoter and 5'-UTR) or a variant thereof.
[0147] Preferably, the variant has at least 70% identity with any of SEQ ID NOs: 101-107.
[0148] In a further embodiment, the 5'-UTR comprises a polynucleotide sequence selected from any of SEQ ID NOs: 81-100 and 126-128 (full length 5'-UTR sequences), or variants thereof.
[0149] Preferably, the variant has at least 70% identity with any of the sequences of SEQ ID NOs: 81-100 and 126-128. In a preferred embodiment, the modification is in the uORF sequence within the 5'-UTR. In a preferred embodiment, the uORF has the sequence of any of SEQ ID NOs: 41-60 and 129-131 (uORF DNA sequences), or a variant thereof.
[0150] In a preferred embodiment, the variant preferably has at least 70% identity to any of SEQ ID NOs: 41-60 and 129-131. In a preferred embodiment, the variant has at least 70% identity with any of SEQ ID NOs: 41-50, 53-57 and 129-131 (dicotyledonous plant uORF DNA sequences).
[0151] In a further embodiment, the variant comprises a sequence having at least 70% identity to any of SEQ ID NOs: 61-80 and 138-140 (conserved regions of uORF DNA sequences). In a further embodiment, the variant comprises any of the sequences of SEQ ID NOs: 61-80 and 138-140 (conserved regions of uORF DNA sequences).
[0152] In a further embodiment, the variant comprises a sequence having at least 70% identity to any of SEQ ID NOs: 61-70, 73-77 and 138-140 (conserved uORF DNA sequences in dicotyledonous plants). In a further embodiment, the variant comprises the sequence of any of SEQ ID NOs: 61-70, 73-77 and 138-140 (conserved uORF DNA sequences in dicotyledonous plants).
[0153] In a further embodiment, the uORF comprises the sequence of any of SEQ ID NOs: 1-20 and 132-134 (uORF polypeptide sequences), or a variant thereof. In a further embodiment, the variant comprises a sequence having at least 70% identity to any of SEQ ID NOs: 1-20 and 132-134 (uORF polypeptide sequences). In a further embodiment, the variant comprises the sequence of any of SEQ ID NOs: 1-20 and 132-134 (uORF polypeptide sequences).
[0154] In a further embodiment, the variant comprises the sequence of any of SEQ ID NOs: 1-10, 13-17 and 132-134 (dicotyledonous plant uORF polypeptide sequences).
[0155] In a further embodiment, the variant comprises a sequence having at least 70% identity to any of SEQ ID NOs: 21-40 and 135-137 (the uORF polypeptide conserved region). In a further embodiment, the variant comprises the sequence of any of SEQ ID NOs: 21-40 and 135-137 (uORF polypeptide conserved region).
[0156] In a further embodiment, the variant comprises a sequence having at least 70% identity to any of SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions in dicotyledonous plants). In a further embodiment, the variant comprises the sequence of any of SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved regions in dicotyledonous plants).
[0157] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to the amino acid sequence of SEQ ID NO: 108 (the consensus motif of the uORF peptide).
[0158] [Modification] In one embodiment, the alteration (also referred to as a mutation) is at least one of a deletion, addition, or substitution of at least one nucleotide in the sequence encoding the 5'-UTR.
[0159] In one embodiment, the modification reduces, inhibits, or prevents translation of a uORF polypeptide encoded by any of SEQ ID NOs: 1-20 and 132-134, or a variant thereof. In a further embodiment, the modification reduces, inhibits, or destroys the activity of a uORF polypeptide encoded by any of SEQ ID NOs: 1-20 and 132-134, or a variant thereof. In one embodiment, the variant comprises a sequence having at least 70% identity to any of SEQ ID NOs: 1-20 and 132-134 (uORF peptides).
[0160] In a further embodiment, the variant comprises a sequence having at least 70% identity to any of SEQ ID NOs: 1-10, 13-17 and 132-134 (dicotyledonous plant uORF peptides).
[0161] In a further embodiment, the variant comprises a sequence having at least 70% identity to at least one of SEQ ID NOs: 21-40 and 135-137 (uORF peptide conserved region). In a further embodiment, the variant comprises at least one sequence of SEQ ID NOs: 21-40 and 135-137 (uORF peptide conserved region).
[0162] In a further embodiment, the variant comprises a sequence having at least 70% identity to at least one of SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved uORF peptide regions in dicotyledonous plants). In a further embodiment, the variant comprises at least one sequence of SEQ ID NOs: 21-30, 33-37 and 135-137 (conserved uORF peptide regions in dicotyledonous plants).
[0163] In one embodiment, the variant or fragment comprises a sequence having at least 70% identity to the amino acid sequence of SEQ ID NO: 108 (the consensus motif of the uORF peptide). In a further embodiment, the variant or fragment comprises the amino acid sequence of SEQ ID NO: 108 (the uORF peptide consensus motif).
[0164] [How to select plants] In a further embodiment, there is provided a method for selecting a plant with increased ascorbic acid production, the method comprising determining in the plant the presence of a first polymorphism in a polynucleotide of the present disclosure, or a second polymorphism linked to the first polymorphism. In one embodiment, the presence of the first polymorphism or the second polymorphism linked to the first polymorphism indicates at least one of the following a) to d): In a further embodiment, the second polymorphism is in linkage disequilibrium (LD) with the first polymorphism. In a further embodiment, the method includes separating the selected plants from one or more non-selected plants. In a further aspect, the present disclosure provides a plant selected by the method of the present disclosure. In a further aspect, the present disclosure provides a group of plants selected by the method of the present disclosure, preferably comprising at least two plants, more preferably at least three, even more preferably at least four, even more preferably at least five, even more preferably at least six, even more preferably at least seven, even more preferably at least eight, even more preferably at least nine, even more preferably at least ten, even more preferably at least 11, even more preferably at least 12, even more preferably at least 13, even more preferably at least 14, even more preferably at least 15, even more preferably at least 16, even more preferably at least 17, even more preferably at least 18, even more preferably at least 19, and even more preferably at least 20 plants.
[0165] [Production method of ascorbic acid] In a further embodiment, the present disclosure provides a method for producing ascorbic acid comprising extracting ascorbic acid from a plant cell or plant of the present disclosure.
[0166] [antibody] In further embodiments, there are provided antibodies raised against the polypeptides of the present disclosure. In further embodiments, there are provided antibodies specific for the polypeptides of the present disclosure.
[0167] [Original and applied plant species] The polynucleotides, polypeptides, variants, and fragments of the present disclosure may be derived from any species. The polynucleotides, polypeptides, variants, and fragments may be naturally occurring or non-naturally occurring. The polynucleotides, variants, and fragments may be produced recombinantly or may be the product of gene shuffling.
[0168] In one embodiment, the polynucleotides, polypeptides, variants and fragments are derived from any plant species. The plants modified or transformed by the methods of the present disclosure may be derived from any plant species. The plant cells modified or transformed by the methods of the present disclosure may be derived from any plant species.
[0169] In further embodiments, the plant is derived from a gymnosperm or angiosperm, a dicotyledonous plant, or a fruit tree, vegetable crop, or tree, such as those exemplified below: In a further embodiment, the plant is derived from a gymnosperm species. In a further embodiment, the plant is derived from an angiosperm species. In a further embodiment, the plant is derived from a dicotyledonous plant species.
[0170] In a further embodiment, the plant is derived from a fruit or nut species selected from the following genera, but not limited to: Citrus, Prunus, Actinidia, Malus, Fragaria, Vaccinium. Particularly preferred fruit plant species include: Citrus × sinensis, Citrus limon, Citrus × paradisi, Prunus dulcis, Prunus domestica, Prunus persica, Prunus avium, Actinidia deliciosa, A. chinensis, A. eriantha, A. arguta, hybrids of four Actinidia species, Malus domestica and Malus sieboldii.
[0171] In a further embodiment, the plant is selected from the group consisting of: Actinidia eriantha, Cucumis sativus, Glycine max, Solanum lycopersicum, Vitis vinifera, Arabidopsis thaliana, Malus × domestica, Medicago truncatula, Populus trichocarpa, Actinidia arguta, Actinidia chinensis, Fragaria vulgaris, Solanum tuberosum, Zea mays.
[0172] In a further embodiment, the plant is derived from a vegetable species selected from the following genera, but not limited to: Brassica, Lycopersicon, Solanum. Particularly preferred vegetable plant species are Lycopersicon esculentum and Solanum tuberosum.
[0173] In a further embodiment, the plant is derived from a monocotyledonous plant species. In a further embodiment, the plant is derived from a crop species selected from, but not limited to, the following genera: Glycine, Zea, Triticum, Hordeum, Miscanthus, Saccharum, Beta, Oryza. Particularly preferred crop plant species include: Hordeum vulgare, Miscanthus × giganteus, Saccharum officinarum, Triticum aestivum, Beta vulgaris, Oryza sativa, Glycine max, Cenchrus purpureus (formerly Pennisetum purpureum), Zea mays.
[0174] In a further embodiment, the plant is derived from a plantation forestry tree species selected from the following genera, but not limited to: Eucalyptus, Salix, Betula, Populus, Pinus. Particularly preferred plantation tree species are Pinus taeda and Eucalyptus grandis. In a further embodiment, the plant is selected from the group consisting of: Actinidia eriantha, Cucumis sativus, Glycine max, Solanum lycopersicum, Vitis vinifera, Arabidopsis thaliana, Malus × domestica, Medicago truncatula, Populus trichocarpa, Actinidia arguta, Actinidia chinensis, Fragaria vulgaris, Solanum tuberosum, Zea mays.
[0175] [Detailed explanation] Citations herein to patent documents, external documents, and other sources are generally intended to provide a context for discussing features of the present disclosure. Unless otherwise expressly stated, citations to such external documents should not be construed as an admission that the documents or sources are prior art in any jurisdiction or are part of the general knowledge of one of ordinary skill in the art.
[0176] [Polynucleotides and their fragments] As used herein, the term "polynucleotide" refers to a single- or double-stranded deoxyribonucleic acid or ribonucleic acid polymer, preferably 15 nucleotides or more in length. This term includes, but is not limited to, coding and non-coding regions of genes, sense and antisense sequences, exons, introns, genomic DNA, cDNA, pre-mRNA, mRNA, rRNA, siRNA, miRNA, tRNA, ribozymes, recombinant polypeptides, isolated or purified naturally-occurring DNA / RNA sequences, synthetic RNA / DNA sequences, nucleic acid probes, primers, and fragments. As used herein, "polynucleotide" preferably includes both the designated sequence and its complementary sequence. A "fragment" of a polynucleotide refers to a subsequence of consecutive nucleotides, and includes, for example, a sequence having a length of 15 nucleotides or more. A fragment of the present disclosure comprises 15 consecutive nucleotides, more preferably 20 or more nucleotides, even more preferably 30 or more nucleotides, even more preferably 50 or more nucleotides, and most preferably 60 or more nucleotides.
[0177] A "primer" is a short polynucleotide, usually with a free hydroxyl group at its 3' end, that is used to hybridize to a template and initiate synthesis of a target complementary polynucleotide.
[0178] [Polypeptides and their fragments] "Polypeptide" refers to a chain of amino acids of any length, particularly preferably five or more amino acids, including full-length proteins, in which the amino acid residues are linked by covalent peptide bonds. Polypeptides of the present disclosure may be purified from natural sources or produced partially or completely by recombinant or synthetic techniques. The term may refer to a single polypeptide, an aggregate such as a dimer or multimer, a linked polypeptide, a polypeptide fragment, a variant, or a derivative. A "fragment" of a polypeptide refers to a subsequence of the polypeptide, preferably one that has a function necessary for biological activity and contributes to the tertiary structure of the polypeptide.
[0179] The term "isolated" as applied to polynucleotides or polypeptides of the present disclosure is used to mean a sequence that has been removed from its natural cellular environment. In one embodiment, the sequence is separated from flanking sequences with which it is naturally present. Isolated molecules may be obtained by any or a combination of biochemical, recombinant, or synthetic techniques.
[0180] The term "recombinant" refers to a polynucleotide that is separated from sequences that naturally flank it and / or that has been recombined with sequences that are not found in the natural state. A "recombinant polypeptide sequence" is produced by translation from a recombinant polynucleotide sequence.
[0181] The term "derived from," with respect to a polynucleotide or polypeptide of the present disclosure derived from a particular genus or species, means having a sequence identical to that of a polynucleotide or polypeptide naturally occurring in that genus or species. Such polynucleotides or polypeptides may be chemically synthesized or recombinantly produced.
[0182] [Methods for isolating or preparing polynucleotides] The polynucleotide molecules of the present disclosure can be isolated using various techniques well known to those skilled in the art. For example, the polypeptides can be isolated using the polymerase chain reaction (PCR) as described in Mullis et al. (eds., 1994, The Polymerase Chain Reaction, Birkhauser, incorporated herein by reference). The polypeptides of the present disclosure can be amplified based on the polynucleotide sequences of the present disclosure using primers defined herein.
[0183] Further, other methods for isolating the polynucleotides of the present disclosure include using all or part of a polypeptide having the sequence disclosed herein as a hybridization probe. Exemplary hybridization and washing conditions for screening a genomic or cDNA library using a technique in which a labeled polynucleotide probe is hybridized to a polynucleotide immobilized on a solid support such as a nitrocellulose filter or nylon membrane are as follows: Hybridization: 20 hours at 65°C in 5× SSC, 0.5% sodium lauryl sulfate (SDS), 1× Denhardt's. Washes: 3 times 20 minutes at 55°C in 1x SSC, 1% (w / v) SDS. Optionally, an additional 20-minute wash at 60°C in 0.5x SSC, 1% (w / v) SDS. Optionally, an additional 20-minute wash at 60°C in 0.1x SSC, 1% (w / v) SDS is also possible.
[0184] Polynucleotide fragments of the present disclosure can be produced by techniques well known to those of skill in the art, such as restriction endonuclease digestion, oligonucleotide synthesis, and PCR amplification.
[0185] Partial polynucleotide sequences can be used to identify corresponding full-length polynucleotide sequences. These methods include PCR-based methods, 5' RACE (Frohman MA, 1993, Methods Enzymol. 218:340-56), hybridization-based methods, and computer / database-based methods. As a further example, inverse PCR allows unknown sequences flanking the polynucleotide sequences disclosed herein to be obtained using primers based on known regions (Triglia et al., 1998, Nucleic Acids Res. 16:8186, incorporated herein by reference). This method uses multiple restriction enzymes to generate fragments appropriate to the known region of a gene, which are then circularized by intramolecular ligation and used as PCR templates. Divergent primers are designed from the known region. Physical construction of full-length clones can be achieved using standard molecular biology techniques (Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, 1987).
[0186] When generating transgenic plants from a particular plant species, it may be beneficial to transform the plant with a sequence derived from that plant species. This would help alleviate public concerns about interspecies gene introgression in the generation of transformed organisms. Furthermore, when silencing gene expression, it may be necessary to use a sequence that is identical to, or at least highly homologous to, the gene to be silenced in the plant. For these reasons, among others, it would be desirable to be able to identify and isolate orthologs of a particular gene across multiple plant species.
[0187] The methods described herein allow for the identification of variants (including orthologues).
[0188] [Variant Identification Method] [Physical method] Variant polynucleotides can be identified using PCR-based methods (Mullis et al., eds., 1994, The Polymerase Chain Reaction, Birkhauser). Typically, the sequences of primers useful for amplifying variants of the polynucleotide molecules of the present disclosure by PCR are designed based on sequences encoded by conserved regions of the corresponding amino acid sequences.
[0189] Alternatively, library screening techniques well known to those skilled in the art (Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, 1987) may be used. When identifying variants of a probe sequence, hybridization and / or washing stringency is generally lower than when searching for an exact match.
[0190] Variants of the polypeptides can also be obtained by physical techniques, such as screening expression libraries with antibodies raised against the polypeptides of the present disclosure or using said antibodies to identify naturally occurring polypeptides.
[0191] [Computer-based methods] Variant sequences (including polynucleotides and polypeptides) of the present disclosure can also be identified using computer-based methods well known to those skilled in the art. This includes searching sequence databases such as GenBank, EMBL, Swiss-Prot, and PIR using public sequence alignment algorithms and sequence similarity search tools (e.g., Nucleic Acids Res. 29: 1-10 and 11-16, 2001). In a similarity search, a target sequence is obtained and aligned for comparison with a subject sequence (query sequence). Sequence comparison algorithms use a scoring matrix to assign an overall score to each alignment.
[0192] One example of a set of programs useful for identifying variants in sequence databases is the BLAST suite (version 2.2.5 [November 2002]), which includes BLASTN, BLASTP, BLASTX, tBLASTN, and tBLASTX. These are available at ftp: / / ftp.ncbi.nih.gov / blast / and from the National Center for Biotechnology Information (NCBI; National Library of Medicine, Building 38A, Room 8N805, Bethesda, Md. 20894, USA). The NCBI server also provides functionality to use these programs to screen multiple public sequence databases. BLASTN compares a nucleotide query sequence with a nucleotide sequence database. BLASTP compares an amino acid query sequence with a protein database. BLASTX translates a nucleotide query sequence in all translation frames and compares it with a protein database. tBLASTN compares a protein query sequence to a translated nucleotide sequence. tBLASTX compares both a nucleotide query and a nucleotide database, translated in six frames. BLAST programs can be used with default parameters or modified as needed to improve search accuracy.
[0193] The use of the BLAST family of algorithms, including BLASTN, BLASTP, and BLASTX, is described in detail in Altschul et al. (Nucleic Acids Res. 25: 3389-3402, 1997). "Hits" generated by a query sequence to one or more database sequences by BLASTN, BLASTP, BLASTX, tBLASTN, tBLASTX, or similar algorithms align and identify similar portions of the sequences. Hits are ordered by the degree of similarity or length of sequence overlap. Typically, hits to a database sequence match only a small portion of the length of the query sequence. The BLASTN, BLASTP, BLASTX, tBLASTN, and tBLASTX algorithms output an "Expectation Value (E-value)" for an alignment. The E-value indicates the expected number of hits that will match by chance when searching a database of random, contiguous sequences of the same size. For example, an E-value of 0.1 assigned to a polynucleotide hit means that a match of that score would likely be observed 0.1 times by chance in a database the size of the database being screened. For sequences with an E-value of 0.01 or less across the aligned and matched portions, the probability of finding that match in the database by chance is interpreted as 1% or less using BLASTN, BLASTP, BLASTX, tBLASTN, and tBLASTX.
[0194] For multiple alignment of a group of similar sequences, there are several methods available, such as CLUSTALW (Thompson, JD, Higgins, DG and Gibson, TJ (1994) CLUSTALW: Improving die sensitivity of progressive multiple sequence alignment through sequence weighting, positions-specific gap penalties and weight matrix choice. Nucleic Acids Research, 22:4673-4680, world wide web -igbmc.u-strasbg.fr / BioInfo / ClustalW / Top.html) and T-COFFEE (Cedric Notredame, Desmond G. Higgins, Jaap Heringa. T-Coffee: A novel method for fast and accurate multiple sequence alignment. J. Mol. Biol. (2000) 302: 205-217)), and PILEUP (Feng and Doolittle, 1987, J. Mol. Evol. 25, 351) is used.
[0195] Pattern recognition software applications are available for finding motif or signature sequences. For example, MEME (Multiple Em for Motif Elicitation) detects motif and signature sequences in a group of sequences, and MAST (Motif Alignment and Search Tool) uses the motif to identify similar or identical motifs present in a query sequence. The results of MAST are presented as a series of alignments with appropriate statistics and a visual summary of the detected motifs. MEME and MAST were developed at the University of California, San Diego.
[0196] PROSITE (Bairoch and Bucher, 1994, Nucleic Acids Res. 22, 3583; Hofmann et al., 1999, Nucleic Acids Res. 27, 21) is a method for identifying the function of uncharacterized proteins translated from genomic or cDNA sequences. The PROSITE database (http: / / www.expasy.org / prosite) contains biologically significant patterns and profiles and is designed to be used with appropriate computational tools to assign new sequences to known protein families or to determine known domains present in the sequence (Falquet et al., 2002, Nucleic Acids Res. 30, 235). Prosearch is a tool that can search the SWISS-PROT and EMBL databases using a given sequence pattern or signature.
[0197] [Method for isolating polypeptides] Polypeptides of the present disclosure (including variant polypeptides) can be prepared by peptide synthesis techniques well known to those skilled in the art, such as direct peptide synthesis using solid-phase synthesis techniques (e.g., Stewart et al., 1969, Solid-Phase Peptide Synthesis, W.H. Freeman, San Francisco, CA) or automated synthesis (e.g., using an Applied Biosystems 431A peptide synthesizer (Foster City, CA)). Mutant forms of the polypeptides of the present disclosure may also be produced during the synthesis.
[0198] The polypeptides and variant polypeptides of the present disclosure can also be purified from naturally occurring sources using a variety of techniques well known to those skilled in the art (e.g., Deutscher, ed., 1990, Methods in Enzymology, Vol. 182, Guide to Protein Purification).
[0199] Alternatively, the polypeptides and variant polypeptides of the disclosure may be obtained by recombinant expression in a suitable host cell and isolation from the host cell, as described below.
[0200] [How to modify the sequence] Methods for modifying a protein sequence or a polynucleotide sequence encoding the same are well known to those skilled in the art. A protein sequence can be conveniently modified by modifying / altering the sequence encoding the protein and expressing the modified protein. Techniques such as site-directed mutagenesis may be applied to modify an existing polynucleotide sequence. Alternatively, a portion of the existing sequence may be excised using a restriction endonuclease. The modified polynucleotide sequence may be synthesized in a modified form.
[0201] [Construct and vector construction methods] A genetic construct according to the present disclosure comprises one or more polynucleotide sequences of the present disclosure and / or a polynucleotide encoding a polypeptide of the present disclosure and may be used to transform, for example, a bacterial, fungal, insect, mammalian, or plant organism. Genetic constructs according to the present disclosure are understood to include expression constructs as defined herein. Methods for making and using genetic constructs and vectors are widely known to those skilled in the art and are generally described, for example, in the following references: Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, 1987; Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing, 1987.
[0202] [Methods for producing host cells containing polynucleotides, constructs, or vectors] The present disclosure provides a host cell comprising a genetic construct or vector of the present disclosure. The host cell may be, for example, of bacterial, fungal, insect, mammalian, or plant origin. Host cells containing genetic constructs, e.g., expression constructs, of the invention are useful in methods for recombinantly producing polypeptides of the invention. Such methods are widely known to those skilled in the art (e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, 1987; Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing, 1987) and may involve culturing host cells in an appropriate medium under conditions suitable for or promoting expression of the polypeptide of the invention. The expressed recombinant polypeptide may be secreted into the medium, if desired, or may be separated from the medium, host cells, or culture medium by methods well known to those skilled in the art (e.g., Deutscher, ed., 1990, Methods in Enzymology, Vol. 182, Guide to Protein Purification).
[0203] [Methods for producing plant cells and plants containing constructs and vectors] The present invention further provides plant cells comprising the genetic constructs of the invention, as well as plant cells modified to alter expression of the polynucleotides or polypeptides of the invention. Plants comprising such cells also form an aspect of the invention. Methods for transforming plant cells, plants, and parts thereof with polypeptides are described in the following publications: Draper et al., 1988, "Plant Genetic Transformation and Gene Expression. A Laboratory Manual," Blackwell Sci. Pub., Oxford, p. 365; Potrykus and Spangenburg, 1995, "Gene Transfer to Plants," Springer-Verlag, Berlin; and Gelvin et al., 1993, "Plant Molecular Biol. Manual," Kluwer Acad. Pub. Dordrecht. A review of transgenic plants, including transformation techniques, is provided in Galun and Breiman, 1997, "Transgenic Plants," Imperial College Press, London.
[0204] [Methods for genetically manipulating plants] Numerous strategies for plant transformation exist (e.g., Birch, 1997, Ann Rev Plant Phys Plant Mol Biol, 48, 297; Hellens RP et al., 2000, Plant Mol Biol 42: 819-832; Hellens R. et al., 2005, Plant Meth 1: 13). For example, strategies can be designed to increase the expression of a polynucleotide / polypeptide in a plant cell, organ, or at a particular developmental stage where it is normally expressed, or to ectopically express it in a cell, tissue, organ, or developmental stage where it is not normally expressed. The polynucleotide / polypeptide to be expressed can be derived from the plant species to be transformed or from another plant species. Transformation strategies may be designed to reduce expression of a polynucleotide / polypeptide in plant cells, tissues, organs, or developmental stages where it is normally expressed, known as gene silencing strategies.
[0205] Gene constructs intended for gene expression in transgenic plants typically include a promoter that drives expression of one or more cloned polynucleotides, a terminator, and may also include a selectable marker sequence for detecting the presence of the gene construct in a transformed plant. Promoters that can be used in the constructs disclosed herein can function in cells, tissues, or organs of monocotyledonous or dicotyledonous plants and may include cell-specific, tissue-specific, and organ-specific promoters, cell cycle-specific promoters, time-specific promoters, inducible promoters, constitutive promoters active in most plant tissues, and recombinant promoters. The choice of promoter depends on the desired temporal and spatial polynucleotide expression. The promoter may be one normally associated with the foreign gene of interest or may be derived from genes from other plants, viruses, plant pathogenic bacteria, or fungi. Those skilled in the art will be able to select suitable promoters for modifying or modulating plant traits using genetic constructs containing the polynucleotide sequences disclosed herein without undue experimentation. Examples of constitutive plant promoters include the CaMV 35S promoter, the nopaline synthase promoter, the octopine synthase promoter, and the Ubi1 promoter from maize. Plant promoters active in specific tissues, promoters that respond to internal developmental signals, or external abiotic or biotic stresses are described in the scientific literature. Exemplary promoters are described, for example, in WO 02 / 00894 (incorporated herein by reference).
[0206] [Termination sequences commonly used in gene constructs for plant transformation] Exemplary termination sequences (terminators) commonly used in plant transformation include, for example, the cauliflower mosaic virus (CaMV) 35S termination sequence, the napalin synthase or octopine synthase termination sequences from Agrobacterium tumefaciens, the zein gene termination sequence from maize (Zea mays), the ADP-glucose pyrophosphorylase termination sequence from rice (Oryza sativa), and the PI-II termination sequence from potato (Solanum tuberosum).
[0207] [Selectable markers commonly used in plant transformation] Selectable markers commonly used in plant transformation include the neomycin phosphotransferase II gene (NPT II), which confers kanamycin resistance, the aadA gene, which confers spectinomycin and streptomycin resistance, the phosphinothricin acetyltransferase (bar gene), which confers resistance to ignite (AgrEvo) and basta (Hoechst), and the hygromycin phosphotransferase gene (hpt), which confers hygromycin resistance. [Use of gene constructs containing reporter genes] Coding sequences that express activities foreign to the host, typically enzyme activities and / or visible signals (e.g., luciferase (LUC), β-glucuronidase (GUS), or green fluorescent protein (GFP)) that can be used to analyze promoter activity in plants and plant tissues, are considered. Reporter genes are described in Herrera-Estrella et al., 1993, Nature 303, 209, and Schrott, 1995, "In: Gene Transfer to Plants" (Potrykus T., Spangenberg, ed., Springer Verlag, Berlin, pp. 325-336).
[0208] [Gene silencing strategies] Gene silencing strategies may focus on the target gene itself or on regulatory elements that affect expression of the encoded polypeptide. The term "regulatory element" is used herein in the broadest possible sense to include other genes that interact with the gene of interest. Genetic constructs designed to reduce or suppress expression of polynucleotides / polypeptides of the present disclosure may contain antisense copies of the polynucleotides of the present disclosure, in which the polynucleotide is positioned in an antisense orientation relative to the promoter and termination sequences. An "antisense" polynucleotide is one that is obtained by inverting a polynucleotide or a portion thereof so that the transcript produced is complementary to the mRNA transcript of the gene of interest.
[0209] [Gene silencing constructs containing inverted repeats] Gene constructs designed for gene silencing may contain inverted repeats. An "inverted repeat" is a repeated sequence in which the second half of the sequence is repeated to form a complementary strand. The resulting transcript may form a hairpin structure through complementary base pairing. Typically, a spacer of at least 3 to 5 base pairs is required between the repeat regions to enable the formation of a hairpin structure.
[0210] [Silencing strategy using small antisense RNAs] Another silencing approach is to use small antisense RNAs that target transcripts corresponding to miRNAs (Llave et al., 2002, Science 297, 2053).The use of such small antisense RNAs corresponding to the polynucleotides of the present invention is also expressly contemplated. As used herein, the term "gene construct" may also include small antisense RNAs and other polypeptides that cause gene silencing. Transformation with an expression construct as defined herein may result in gene silencing through a process known as "sense suppression" (e.g., Napoli et al., 1990, Plant Cell 2, 279; de Carvalho Niebel et al., 1995, Plant Cell 7, 347). In some cases, sense suppression may involve overexpression of the full-length or partial coding region of the gene, but may also involve expression of non-coding regions such as introns or 5' or 3' untranslated regions (UTRs). Chimeric partial-sense constructs may be used to coordinately silence multiple genes (Abbott et al., 2002, Plant Physiol. 128(3): 844-53; Jones et al., 1998, Planta 204: 499-505). The use of such sense suppression strategies to silence expression of polynucleotides of the present invention is also expressly contemplated. The polynucleotides inserted into gene constructs designed for gene silencing purposes may correspond to coding and / or non-coding regions, such as promoters and / or introns and / or 5' or 3' untranslated regions (UTRs) or the corresponding gene. Other gene silencing strategies may include dominant-negative approaches or the use of ribozyme constructs (McIntyre, 1996, Transgenic Res, 5, 257). Pre-transcriptional silencing may be caused by mutations in the gene itself or its regulatory elements, which may include point mutations, frameshifts, insertions, deletions, and substitutions.
[0211] The following are representative references that disclose genetic transformation protocols that can be used to genetically transform the following plant species: Rice (Alam et al., 1999, Plant Cell Rep. 18. 572); apple (Yao et al., 1995, Plant Cell Reports 14, 407-412); maize (U.S. Pat. Nos. 5,177,010 and 5,981.840); wheat (Ortiz et al., 1996, Plant Cell Rep. 15, 1996, 877); tomato (U.S. Pat. No. 5,159,135); potato (Kumar et al., 1996 Plant J. 9, 821); cassava (Li et al., 1996 Nat. Biotechnology' 14, 736); lettuce (Michelmore et al., 1987. Plant Cell Rep. 6, 439); tobacco (Horsch et al., 1985, Science 227, 1229); cotton (U.S. Pat. Nos. 5,846,797 and 5,004,863); grasses (U.S. Pat. Nos. 5,187,073 and 6,020,539); peppermint (Niu et al., 1998, Plant Cell Rep. 17, 165); citrus plants (Pena et al., 1995, Plant Sci. 104, 183); caraway (Krcns ct al., 1997, Plant Cell Rep, 17, 39); banana (U.S. Pat. No. 5,792,935); soybean (U.S. Pat. Nos. 5,416,011; 5,569,834; 5,824,877; 5,563,04455 and 5,968,830); pineapple (U.S. Pat. No. 5,952,543); poplar (U.S. Pat. No. 4.795,855); monocots in general (U.S. Pat. Nos. 5,591,616 and 6.037,522); brassica (U.S. Pat. Nos.5,188,958; 5,463,174 and 5,750,871); cereals (US Pat. No. 6.074,877); pear (Matsuda et al., 2005, Plant Cell Rep. 24(1):45-51); Primus (Ramesh et al., 2006 Plant Cell Rep. 25(8):821-8; Song and Sink 2005 Plant Cell Rep. 2006; 25(2): 117-23; Gonzalez Padilla et al., 2003 Plant Cell Rep. 22(l):38-45); strawberry (Oosumi et al., 2006 Planta. 223(6): 1219-30; Folta et al., 2006 Planta April 14; PMID: 16614818), rose (Li et al., 2003), Rubus (Graham et al., 1995 Methods Mol Biol. 1995; 44: 129-33), tomato (Dan et al., 2006, Plant Cell Reports V25:432-441), apple (Yao et al., 1995, Plant Cell Rep. 14, 407-412), and Actinidia eriantha (Wang et al. 2006. Plant Cell Rep. 25, 5: 425-31). The present disclosure also explicitly contemplates transformation of other plant species. Suitable methods and protocols are available in the scientific literature.
[0212] Furthermore, several additional methods known to those skilled in the art may be used to modify the expression of the nucleotides and / or polypeptides of the present disclosure. These methods include, but are not limited to, TILLING (Targeting Induced Local Lesions IN Genomes) (Till et al., 2003, Methods Mol Biol, 205), the so-called "Deletagene" technology (Li et al., 2001, Plant Journal 27(3), 235), and the use of artificial transcription factors, such as synthetic zinc finger transcription factors (e.g., Jouvenot et al., 2003, Gene Therapy 10, 513). In addition, antibodies or fragments thereof targeting specific polypeptides may be expressed in plants to modulate the activity of the polypeptides (Jobling et al., 2003, Nat. Biotechnol. 21(1), 35). Transposon tagging approaches may also be applied. Additionally, peptides that interact with the polypeptides of the present disclosure may be identified through techniques such as phage-display (Dyax Corporation). The interacting peptides may be expressed in plants or administered to plants to affect the activity of the polypeptides of the present disclosure. The use of any of the above methods to alter expression of the nucleotides and / or polypeptides of the present disclosure is specifically contemplated.
[0213] [Method for modifying endogenous DNA sequences in plants] Those skilled in the art know methods for modifying endogenous genomic DNA sequences in plants, which may involve the use of sequence-specific nucleases to induce targeted double-stranded DNA breaks in the gene of interest. Examples of such methods used in plants include zinc finger nucleases (Curtin et al., 2011, Plant Physiol. 156:466-473; Sander et al., 2011, Nat. Methods 8:67-69), transcription activator-like effector nucleases, so-called "TALENs" (Cermak et al., 2011, Nucleic Acids Res. 39:e82; Mahfouz et al., 2011, Proc. Natl. Acad. Sci. USA 108:2623-2628; Li et al., 2012, Nat. Biotechnol. 30:390-392), and LAGLIDADG homing endonucleases, also called "meganucleases" (Tzfira et al., 2012, Plant Biotechnol. J. 10:373-389).
[0214] A preferred method of practicing the present invention is the use of genome editing to achieve "targeted genetic modification," as referred to herein. The terms "genome editing," "genome edited," "genome modified," and "genetically modified" are used interchangeably to describe plants that have changes to specific DNA sequences within the plant genome, changes that include mutations of specific nucleotides, deletions of specific sequences, or insertions of specific sequences. As used herein, the term "technology for introducing targeted gene modifications" refers to any method, protocol, or technique that enables precise and targeted editing at a specific location (also called a "locus" or "native locus") in a plant genome using site-specific nucleases (e.g., meganucleases, zinc finger nucleases (ZFNs), RNA-guided endonucleases (e.g., CRISPR / Cas9 systems), TALE endonucleases (TALENs), recombinases, or transposases).
[0215] CRISPR stands for "clustered, regularly interspaced, short, palindromic repeats," and Cas stands for "CRISPR-associated protein" (for a review, see Khandagal and Nadal, Plant Biotechnol. Rep., 2016, 10, 327). Other gene-editing techniques that can be used include genetically engineered meganucleases, zinc finger nucleases (ZFNs), and TALENs. U.S. Patent Application 2016 / 0032297 provides detailed methods for these methods. Another gene-editing technique that can be used is the so-called ARCUS nuclease, which exploits the properties of the naturally occurring gene-editing enzyme, the homing endonuclease I-CreI, which evolved in nature to have a built-in safety switch that shuts off its activity after making a single, specific DNA modification.
[0216] Genome editing tools allow precise modification of genome structure at specific target locations. These tools can be used efficiently to create high-yield crops, modify plant composition, and generate plants resistant to biotic or abiotic stress. Because it can be difficult to achieve all desired modifications with a single genome editing tool, multiple genome editing tools have been developed to enable efficient genome editing. The main genome editing tools used in plant genome editing include homologous recombination (HR), zinc finger nucleases (ZFNs), TALENs, pentatricopeptide repeat proteins (PPRs), CRISPR / Cas9 systems, RNA interference (RNAi), cisgenesis, and intragenesis. Furthermore, site-specific sequence editing and oligonucleotide-directed mutagenesis have the potential to edit genomes at the single-base level. Recently, adenine base editors (ABEs) have been developed, which can mutate AT base pairs to GC base pairs. ABEs utilize a catalytically inactive Cas9 nickase in combination with deoxyadenine deaminase (TadA). A summary of these methods and their applicability is given in Mohanta et al., Genes (Basel), 2017 Dec; 8(12): 399.
[0217] These genome editing methods can precisely reduce or increase the expression of target genes in plant genomes by removing genes or fragments thereof, altering DNA sequences, inserting novel sequences, or modifying coding or regulatory sequences (Belhaj et al., 2013, Plant Methods. 9, 39; Khandagale and Nadal, 2016, Plant Biotechnol Rep. 10, 327). Preferably, methods use nuclease enzymes and the host plant's DNA repair mechanisms to achieve in vivo site-specific cleavage (double-strand breaks) at specific DNA sequences in genomic DNA. Many approaches exist for generating double-strand breaks in genomic DNA, including the use of CRISPR / Cas systems.
[0218] A comprehensive overview of the CRISPR / Cas system and its applications is available. www.addgene.org / guides / crispr / is described in. The CRISPR / Cas genome editing system offers the flexibility to target and modify specific sequences within the genome, enabling a variety of modifications, including activation, enhanced expression, or knockout of target loci. This method involves introducing a Cas enzyme and a guide RNA (gRNA) containing a short guide sequence (approximately 20 bases) complementary to the target DNA sequence into the plant genome. Depending on the type of Cas, a guide polynucleotide consisting of DNA, an RNA / DNA hybrid, or double-stranded DNA can be used. The guide portion of the guide polynucleotide guides the Cas enzyme to the desired cleavage site via a recognition sequence to which the Cas enzyme binds.
[0219] The target in the plant genome can be any 20-base sequence that is unique relative to other genomic sequences and is immediately followed by a protospacer adjacent motif (PAM) sequence. The PAM sequence serves as a binding signal for Cas9, and its exact sequence depends on the Cas protein used. A list of Cas proteins and their corresponding PAM sequences can be found at www.addgene.org / guides / crispr / #pam-table.
[0220] The simplest application of CRISPR / Cas is the generation of knockout or loss-of-function alleles at a target gene locus. The gRNA guides the Cas enzyme to a specific genomic locus, where it creates a double-strand break. The resulting DSB (double-strand break) is then repaired by one of the cell's existing repair pathways. This process typically results in the insertion or deletion of a small number of nucleotides (indels) at the break site. Often, these indels result in deletion, insertion, or frameshift mutations of amino acids within the open reading frame (ORF) of the target gene, resulting in a premature stop codon. The ideal outcome is a loss-of-function mutation within the target gene.
[0221] However, the robustness of the knockout phenotype in the mutant cells needs to be experimentally verified by detecting the presence or absence of the target ORF transcript using RT-PCR or hybridization-based methods. These features make the CRISPR / Cas system a suitable tool for knocking out uORFs.
[0222] CRISPR / Cas can also be used to make more sophisticated modifications to the native sequence of a target locus in the genome. This can involve inserting, substituting, or editing specific bases to introduce or create new domains in the polypeptide encoded by the locus of interest. One method for making these precise modifications is to utilize the HDR (homologous recombination repair) pathway, which is known to have high fidelity but low efficiency in cells. To achieve precise modifications using HDR, a DNA repair template containing the desired genomic modification must be introduced into the target cell along with a gRNA and Cas9 or Cas9 nickase. In addition to the sequence containing the desired modification, the repair template must also contain homologous sequences (left and right homology arms) upstream and downstream of the target site. The length of the homology arms depends on the size of the inserted sequence. Due to the high cleavage efficiency of Cas9 and the low efficiency of HDR, the majority of DSBs are repaired with edits that do not contain the desired modification. This necessitates an additional confirmation or screening step to select cells containing the desired modification from the edited cell population. This selection can be achieved by inserting an easily screenable marker sequence, or by PCR or hybridization-based techniques.
[0223] CRISPR-related gene editing systems can also be used to change specific bases without the need for double-strand breaks. This approach is known in the art as a "base editing" system. Base editing allows for the irreversible conversion of specific DNA bases to other bases at targeted genomic loci. For example, C can be converted to T, or A to G. Unlike other genome editing tools, base editing can be achieved without the need for double-strand breaks. Base editing is more efficient than traditional genome editing techniques at introducing point mutations at targeted loci. Because many genetic diseases result from point mutations, base editing has important applications in disease research. Using these systems, skilled engineers can create targeted gene modifications, including amino acid substitutions or the creation of start / stop codons.
[0224] To avoid relying on the inefficient homology-directed repair (HDR), researchers have developed two classes of base editing enzymes: cytosine base editors (CBEs) and adenine base editors (ABEs).
[0225] CBEs are created by combining Cas9 nickase or catalytically inactive "dead" Cas9 (dCas9) with an APOBEC-like cytidine deaminase. Similar to traditional CRISPR technology, the base editor is guided to a specific locus by a gRNA and can convert cytidine to uridine within a small editing window near the PAM site. The uridine is then converted to thymidine by base excision repair, resulting in a C to T change. Similarly, adenine base editors are designed to convert adenosine to inosine, which is treated as guanosine in the cell, resulting in an A to G change.
[0226] Adenine DNA deaminases do not occur in nature, but these enzymes were generated by directed evolution of tRNA adenine deaminase (TadA) from Escherichia coli. Similar to CBEs, the evolved TadA domain is bound to the Cas9 protein to form ABEs. Both base editing enzymes can be used in combination with multiple Cas9 variants, including high-fidelity Cas9s. Further advances include optimizing the expression of the binding protein, tailoring the editing window by modifying the linker region between the Cas variant and the deaminase, and improving product purity by adding a DNA glycosylase inhibitor (UGI) or the Gam protein from the bacteriophage Mu (Mu-GAM).
[0227] While many base editing enzymes are designed to function within a very narrow window near the PAM sequence, some base editing systems can induce large numbers of single-nucleotide mutations (somatic hypermutations) within a broader editing window, making them suitable for directed evolution applications. Examples of these systems include TAM (targeted AID-mediated mutagenesis) and CRISPR-X, in which Cas9 is coupled to activation-induced cytidine deaminase (AID).
[0228] Other CRISPR systems, particularly the Type VI CRISPR enzymes Cas13a / C2c2 and Cas13b, target RNA rather than DNA. The RNA-active adenosine deaminase ADAR2 (E488Q) is linked to catalytically inactive Cas13b to create a programmable RNA base editing enzyme (REPAIR) that converts adenosines in RNA to inosines. Inosine is functionally equivalent to guanosine, resulting in an A→G conversion in RNA. The catalytically inactive Cas13b orthologue dPspCas13b from the Prevotella genus does not appear to require flanking sequences in the RNA target, making it a highly flexible editing system. Editors based on another ADAR variant, ADAR2 (E488Q / T375G), have improved specificity, and an editor with the Δ984-1090 truncation retains RNA editing activity and is small enough to be packaged in an AAV vector.
[0229] As used herein, the term "Cas nuclease" refers to any nuclease that site-specifically recognizes a CRISPR sequence based on a gRNA or DNA sequence, including Cas9, Cpf1, and others described below. Many authors have identified CRISPR / Cas genome editing as a suitable tool for editing the genomes of complex organisms (Sander and Joung, 2013, Nat Biotech, 2014. 32, 347; Wright et al., 2016, Cell. 164, 29), including plants (Zhang et al., 2016, Journal of Genetics and Genomics, 43. 151; Puchta 2016, Plant J., 87. 5; Khandagale and Nadaf, 2016, Plant Biotechnol. Rep., 10, 327). U.S. Patent Application 2016 / 020822 provides a detailed description of materials and methods useful for genome editing using the CRISPR / Cas9 system in plants and describes many applications of the CRISPR / Cas9 system for genome editing of various gene targets in crop plants.
[0230] Furthermore, it is recognized that many variations of the CRISPR / Cas system can be used to apply the present invention, including the use of wild-type Cas9 from Streptococcus pyogenes (Type II Cas) (Barakate and Stephens, 2016, Frontiers in Plant Science, 7, 765; Bortesi and Fischer, 2015, Biotechnology Advances 5, 33. 41; Cong et al., 2013, Science, 339, 819; Rani et al., 2016, Biotechnology Letters, 1-16; Tsai et al., 2015, Nature Biotechnology, 33, 187). Other examples include Tru-gRNA / Cas9, which significantly reduces off-target mutations (Fu et al., 2014, Nature Biotechnology, 32, 279; Osakabe et al., 2016, Scientific Reports, 6, 26685; Smith et al., 2016, Genome Biology, 17, 1; Zhang et al., 2016, Scientific Reports, 6, 28566), and highly specific Cas9 (mutant S. pyogenes Cas9) with little or no off-target activity (Kleinstiver et al., 2016, Nature, 529, 490; Slaymaker et al., 2016, Science, 351. 84).Furthermore, there are Type I and Type III systems in which multiple Cas proteins are expressed to achieve editing (Li et al., 2016, Nucleic Acids Research, 44:e34; Luo et al., 2015, Nucleic Acids Research, 43, 674), Type V Cas systems using the Cpf1 enzyme (Kim et al., 2016, Nature Biotechnology, 34, 863; Toth et al., 2016, Biology Direct, 11, 46; Zetsche et al., 2015, Cell, 163, 759), DNA-guided editing using the NgAgo Argonaute enzyme from Natronobacterium gregoryi with guide DNA (Xu et al., 2016, Genome Biology, 17, 186), and dual-vector systems in which Cas9 and gRNA expression cassettes are carried on separate vectors (Cong et al., 2013, Science, 339, 819). An alternative to Cas9 is the use of a unique nuclease called Cpf1, which has the advantage over the Cas9 system of reducing off-target editing that can cause unwanted mutations in the host genome. Examples of genome editing in crops using the CRISPR / Cpf1 system include rice (Tang et al., 2017, Nature Plants 3, 1-5; Wu et al., 2017, Molecular Plant, March 16, 2017) and soybean (Kim et al., 2017, Nat Commun. 8, 14406).Other authors have described the use of Argonaute-related proteins as an alternative to the CRISPR system for gene editing (Hegge et al., Nature Rev. Microbiol. 2017. Epub 2017 / 07 / 25. pmid:28736447; Swarts et al., Nucleic Acids Res., 2015; 43(10):5120-9. Epub 2015 / 05 / 01. pmid:25925567; Swarts et al., Nature, 2014; 507(7491):258-61. Epub 2014 / 02 / 18. pmid:24531762; see also PCT Application No. PCT / US2019 / 025163 and / or Publication No. WO2019204266A1).
[0231] Detailed methodologies for gene editing to create new crop traits in plants, including selection of cells containing the desired edits, and methods for introducing CRISPR system components into initial target plant cells are described in Patent Application Publication WO2019195157. As specified therein, the "guide polynucleotide" in the CRISPR system also refers to a polynucleotide sequence that can form a complex with the Cas endonuclease and enables the Cas endonuclease to recognize and optionally cleave a DNA target site. The guide polynucleotide may be a single molecule (i.e., a single guide RNA (gRNA) that is a synthetic combination of a crRNA and a portion of the tracrRNA sequence) or two molecules (i.e., a crRNA and a tracrRNA found in the naturally occurring Cas9 system in bacteria). The guide polynucleotide sequence may be provided as an RNA sequence or may be transcribed from a DNA sequence into an RNA sequence. The guide polynucleotide sequence may also be provided as a combined RNA-DNA sequence (see, for example, Yin, H. et al., 2018, Nature Chemical Biology, 14, 311). As used herein, a "guide RNA" sequence refers to a sequence comprising a variable targeting domain called a "guide" that is complementary to the target genomic site, and a "guide RNA scaffold," which is an RNA sequence that interacts with Cas9 or Cpf1 endonuclease. A guide polynucleotide composed solely of ribonucleotides is also referred to as a "guide RNA." As used herein, a "guide target sequence" refers to a sequence of genomic DNA adjacent to the PAM site where the gRNA binds to cleave DNA. The "guide target sequence" is usually complementary to the "guide" portion of the gRNA, but depending on its position, some mismatches may still allow DNA cleavage by Cas. The present method also provides for the introduction of a single guide RNA (gRNA) into a plant. The gRNA comprises a nucleotide sequence complementary to the target chromosomal DNA.The gRNA may be, for example, an engineered single-stranded guide RNA containing a crRNA sequence (complementary to the target DNA sequence) and a common tracrRNA sequence, or a crRNA-tracrRNA hybrid. The gRNA may be introduced into a cell or organism as DNA with an appropriate promoter, in vitro transcribed RNA, or synthetic RNA. Basic guidelines for designing guide RNA for any target gene are well known to those skilled in the art and are described, for example, by Brazelton et al. (Brazelton, VA et al., 2015, GM Crops & Food. 6. 266-276) and Zhu (Zhu, LJ 2015, Frontiers in Biology, 10. 289-296).
[0232] Published patent applications WO2019195157 and WO2019204266A1 also provide examples of mutations that may enhance the activity of transcription factor polypeptides, including mutations to the coding sequence that result in amino acid changes in the encoded protein.
[0233] In a preferred embodiment of the invention, the guide polynucleotide / Cas endonuclease system can be used to insert promoters or promoter elements (e.g., enhancer elements) for any of the transcription factor sequences of the invention, and the promoter insertion (or promoter element deletion) may result in any one or more combinations of a permanently activated locus, increased promoter activity (enhanced promoter strength), increased promoter tissue specificity, decreased tissue specificity, new promoter activity, extended temporal window of gene expression, altered expression timing or developmental stage, DNA binding element mutations, and / or additional DNA binding elements.
[0234] The guide RNA / Cas endonuclease system can also be used to insert promoter elements to enhance expression of the transcription factor sequences of the present invention. Promoter elements, such as enhancer elements, are often introduced in multiple copies into promoters driving gene expression cassettes for trait gene testing or for creating transgenic plants expressing specific traits. Examples of enhancer elements include, but are not limited to, the CaMV 35S enhancer element (Benfey et al., EMBO J., 1989; 8: 2195-2202). In some plant (event) species, enhancer elements can cause desirable phenotypes, increased yield, or altered expression patterns of desired traits. It may be desirable to remove redundant enhancer element copies while retaining the desired trait gene cassette at its integration site in the genome. Guide RNA / Cas endonuclease can be used to remove unnecessary enhancer elements from the plant genome. Guide RNAs can be designed to contain variable target regions corresponding to 12-30 bp of target sequences flanked by NGG (PAM) within the enhancer. Cas endonucleases can perform cleavage to insert one or more enhancers.
[0235] To suppress the function of a target uORF and activate a downstream mORF, bases may be deleted from the uORF or an additional stop codon may be created. Other mutations or edits that can be used to suppress the function of a target uORF may include mutation of the initiation ATG codon, deletion of amino acids, insertion, frameshift mutation resulting in the formation of a premature stop codon, or deleterious mutations within the uORF. In some cases, substituting bases within the uORF may optimize a desired trait, such as increased yield, increased stress tolerance, or altered biochemical composition, without significant off-type effects such as organ abnormalities or dwarfism.
[0236] Delivery of gene-editing constructs into plant cells and plants. Sandhya et al. (2020, J. Genet. Eng. Biotechnol. 18: 25, published online July 7, 2020, doi:10.1186 / s43141-020-00036-8) present a method for introducing gene editing tools, such as CRISPR / Cas9 components, into plants to perform gene editing. Effective delivery of CRISPR / Cas9 components, including guide sequences, Cas9, and, if applicable, a DNA repair template containing the desired sequence mutation, into plant cells is crucial for enhancing editing efficiency. Practitioners may choose from a variety of delivery methods for introducing gene editing components into plant cells, including Agrobacterium-mediated transformation, biolistic (particle gun) methods, floral dip methods, and PEG-assisted protoplast transformation. Additional methods include nanoparticle and pollen magnetofection delivery systems (Kwak et al., 2019, Nature Nanotechnology, DOI 10.1038 / S41565-019-0375-4; Demirer et al., 2019, Nature Nanotechnology, DOI 10.1038 / S41565-019-0382-5). CRISPR constructs can be coated onto gold particles for biolistic delivery, transfected into protoplasts using PEG, or introduced using Agrobacterium strains harboring CRISPR vectors. Constructs can also be introduced via pollen tube pathway-based methods or floral dip (Castel et al., 2019, PLoS One 14:e0204778). The next step involves selecting plants or plant cells containing the targeted genetic modification created by the introduced CRISPR system. This may include redifferentiating the cells containing the modification into explants, plant tissues, or whole plants.In some cases, the explants containing genome editing are selected on selective medium, and then redifferentiated into whole plants.Finally, PCR and Sanger sequencing are generally used to confirm that the desired sequence editing has been successfully introduced into the selected plants.The selected plants are then evaluated to confirm that they actually express the targeted traits expected to result from the introduction of the genome edit.
[0237] (2020) also present a table showing applicable methods for specific crops. For example, the following plants have all been successfully gene edited using PEG-mediated delivery of CRISPR system components: apple (Apple), rapeseed (Brassica oleracea), turnip (Brassica rapa), watermelon (Citrullus lanatus), soybean (Glycine max), grapevine (Grapevine), rice (Oryza sativa), petunia (Petunia), liverwort (Physcomitrella patens), tomato (Solanum lycopersicum), wheat (Triticum aestivum), and corn (Zea mays). As further examples, the following plants have been successfully gene edited by delivery of CRISPR system components via particle bombardment: soybean (Glycine max), barley (Hordeum vulgare), rice (Oryza sativa), wheat (Triticum aestivum), and corn (Zea mays). As further examples, the following plants have also been successfully gene edited using particle bombardment-mediated delivery of CRISPR system components: Arabidopsis thaliana, banana, orange, cucumber, kiwi fruit, lotus japonicus, marshmallow, medicinal plant, Nicotiana benthamiana, tobacco, poplar, danshen, tomato, sorghum, wheat, and corn.
[0238] In certain embodiments of the disclosure, one or more base pairs within a uORF may be modified using one of the techniques described above (e.g., TALENs or zinc finger nucleases) to disable it from being translated.
[0239] In one embodiment, this may be achieved by changing the first base pair of the ACG start codon to TCG, which neutralizes the ascorbic acid feedback regulation of GGP translation, allowing for increased ascorbic acid concentrations in the plant.
[0240] Alternatively, codons for highly conserved amino acids in the uORF can be mutated to abolish the uORF's ability to repress GGP translation under high ascorbic acid concentrations, for example, by substituting a histidine residue in the conserved region with a leucine.
[0241] In a further embodiment, a premature base pair in the uORF may be mutated to introduce a stop codon, causing premature termination of the uORF and preventing ascorbic acid feedback regulation of GGP translation.
[0242] Thus, one skilled in the art will appreciate that there are numerous ways to disrupt uORFs, remove negative regulation by ascorbic acid, and increase ascorbic acid production, all of which are within the scope of the present disclosure.
[0243] [plant] The term "plant" is intended to include the whole plant, any part of the plant, propagules, and progeny of the plant. The term "propagule" means any part of the plant that can be used for propagation or multiplication, whether sexual or asexual, i.e., seeds, cuttings, etc. Plants of the present disclosure may be grown, self-crossed, or crossed with plants of different lines, and progeny from two or more generations also constitute an aspect of the present disclosure, so long as they retain the modification or transgene of the present invention. The present disclosure is also intended to be broadly organized with respect to the portions, elements, and features referred to or shown in the application specification, and is to be construed to include any and all combinations thereof, individually or collectively. And, where specific components are referred to herein and there are known equivalents in the art, such equivalents are deemed to be incorporated as if individually set forth herein. [Example]
[0244] The present specification, having been outlined above, will be more readily understood by reference to the following examples. The examples are included solely for the purpose of illustrating certain aspects and embodiments of the present disclosure and are not intended to limit the disclosure. The scope of the present invention is not intended to be limited to these examples alone. As will be appreciated by those skilled in the art, many variations can be made without departing from the scope of the present disclosure.
[0245] Example 1: Elucidation of the regulation of GGP expression [overview] Ascorbic acid (vitamin C) is an essential antioxidant and enzyme cofactor in both plants and animals. Ascorbic acid concentrations are tightly regulated in plants in response to abiotic and biotic stresses. Applicants have shown that ascorbic acid concentrations are controlled through post-transcriptional repression of GDP-L-galactose phosphorylase (GGP), the rate-limiting enzyme in the ascorbic acid biosynthetic pathway. This regulation requires translation of a cis-acting uORF (upstream open reading frame) that initiates translation from a non-canonical start codon and represses translation of the downstream GGP ORF under high ascorbic acid conditions. Removal of this uORF allows plants to produce high levels of ascorbic acid. This uORF is present in the GGP gene in both lower and higher plants, suggesting that it is an ancient mechanism for controlling ascorbic acid concentrations.
[0246] Ascorbic acid (vitamin C) is an essential biochemical found in most living organisms. It plays a central role in regulating cellular redox potential (Asensi-Fabado et al., 2010, Trends Plant Sci. 15, 582; Foyer et al., 2011, Plant Physiol. 155, 2) and also functions as an enzyme cofactor (Mandi et al., 2009, Br. J. Pharmacol. 157, 1097). Ascorbic acid concentration is regulated according to demand. For example, leaf ascorbic acid concentration increases under high light conditions, when ascorbic acid demand is greatest (Bartoli et al., 2006, J. Exp. Bot. 57, 1621; Gatzek et al., 2002, Plant J. 30, 541). However, the mechanism by which ascorbic acid biosynthesis is regulated remains unclear. Applicants have previously shown that the enzyme GDP-L-galactose phosphorylase (GGP) is central to ascorbic acid determination in plants (Bulley et al., 2012, Plant Biotechnol. J. 10, 390; Bulley et al., 2009, J. Exp. Bot. 60, 765), suggesting that the enzyme may play a regulatory role.
[0247] [Results and Discussion] To investigate whether the GGP gene is regulated by ascorbic acid concentration, Applicants linked the kiwifruit GGP promoter and its 5'-UTR (SEQ ID NO: 101) to a luciferase (LUC) reporter gene and transiently expressed the construct in Nicotiana benthamiana leaves (Hellens et al., 2005, Plant Methods 1, 13). To manipulate ascorbic acid, Applicants expressed the GGP coding sequence under a strong constitutive promoter. As a result, doubling the ascorbic acid concentration from approximately 2 mM (20 mg / 100 g fresh weight) to 4 mM reduced relative LUC activity by 50%, and increasing ascorbic acid to nearly 10 mM suppressed LUC activity by more than 90%. Similarly, the Arabidopsis-derived GGP (VTC2; At4g26850) promoter and 5′-UTR (SEQ ID NO: 102) also conferred an ascorbic acid-dependent repressive effect on the LUC reporter gene.
[0248] In contrast, when the promoter of a gene unrelated to ascorbic acid metabolism (T-8, an Arabidopsis bHLH transcription factor that controls polyphenol biosynthesis) was used, high ascorbic acid did not affect relative LUC activity. Furthermore, additional control experiments demonstrated that this regulation was specific to the GGP sequence, was independent of transgene expression levels, and was reflected as a change in LUC protein (Tables 2–4).
[0249] [Table 2]
[0250] [Table 3]
[0251] [Table 4]
[0252] In the experiments described herein, we manipulated leaf ascorbate levels through altered expression of the GGP coding sequence. To separate the effects of ascorbate from potential effects of the GGP protein (Müller-Moule, P., 2008, Plant Mol. Biol. 68, 31), we expressed both GGP and GME individually and together. We previously showed (Bulley et al., 2009, J. Exp. Bot. 60, 765) that expression of GGP alone in tobacco plants had a moderate effect and GME had little effect, whereas coexpression of the two genes resulted in a strong, synergistic stimulation of ascorbate levels. Thus, by altering the ratio of these two genes, we could manipulate ascorbate levels independently of GGP protein abundance. The relative response of LUC activity to ascorbate followed a smooth curve, despite the different GGP protein levels associated with different ascorbate concentrations. This indicates that the factor that reduces LUC activity is ascorbic acid or a metabolite related to it.
[0253] To verify whether the effect of ascorbic acid was mediated by the nontranscribed promoter or the 5'-UTR, we performed two experiments. First, we swapped the 5'-UTR regions between the GGP and TT8 promoters. These were transiently expressed in leaves, and relative LUC activity was measured. Increased ascorbic acid reduced LUC activity only in the presence of the GGP-derived 5'-UTR (see Table 5). Second, we deleted two regions (-387 to -432 bp and -514 to -597 bp) within the 5'-UTR that were particularly highly conserved at the DNA level between species, while retaining the remaining GGP promoter. We then verified this by reporter assay. All deletion mutations abolished the ability to down-regulate reporter gene expression by ascorbic acid. These experiments demonstrate that the 5'-UTR is both necessary and sufficient for ascorbic acid-induced down-regulation.
[0254] [Table 5]
[0255] To examine whether ascorbic acid regulation is at the transcriptional or post-transcriptional level, we measured the transcript levels of a reporter gene construct. Our data show that ascorbic acid has a minimal effect on LUC mRNA levels (see Table 6). This suggests that ascorbic acid acts directly or indirectly through the 5'-UTR to regulate GGP translation.
[0256] [Table 6]
[0257] To examine whether the 5'-UTR directly affects leaf ascorbic acid concentrations, we constructed constructs with and without the GGP 5'-UTR preceding the 35S promoter-driven GGP coding sequence. Both constructs increased leaf ascorbic acid in a transient expression system, but the construct lacking the 5'-UTR had approximately 30% higher ascorbic acid than the construct containing the 5'-UTR. Furthermore, to further increase ascorbic acid levels beyond those of GGP alone, we co-injected GME into leaves (Bulley et al., 2009, J. Exp. Bot. 60, 765). The construct lacking the 5'-UTR produced more than twice the amount of ascorbic acid compared with the construct containing the 5'-UTR. Therefore, under high-ascorbic acid conditions, the GGP 5'-UTR limited both GGP production and ascorbic acid synthesis. Removal of this control provides a means to create plants with high ascorbic acid concentrations.
[0258] Given that the effects of ascorbic acid are mediated through the 5'-UTR region of GGP, Applicants investigated the properties of the 5'-UTR. GGP is unique in that it possesses a long 5'-UTR (over 500 base pairs) and highly conserved sequence elements across many plant species. Alignment of the GGP 5'-UTR from multiple species, including algae and two bryophytes, revealed the presence of a highly conserved uORF that potentially encodes a peptide of 60-65 amino acids. Interestingly, translation of this peptide requires initiation at a non-canonical start codon, ACG. Although there are limited reports of non-canonical translation initiation (Ivanov et al., 2008, Proc. Natl. Acad. Sci. USA 105, 10079), this ORF contains a Kozak sequence required for efficient translation. To test whether this uORF is required for ascorbic acid-dependent GGP gene regulation, Applicants mutated the ACG start codon to TCG. LUC activity from the mutant construct remained high even under high ascorbic acid conditions. To further verify the necessity of this uORF, we replaced the highly conserved His residue (position 36, CGG) with Leu (CTG). This mutation abolished ascorbic acid-dependent regulation. Furthermore, mutating an internal ATG codon (encoding a short uORF consisting of 10 amino acids) more than doubled LUC activity, presumably due to removal of a competing initiation codon, but this did not affect the promoter's sensitivity to ascorbic acid concentrations.
[0259] Next, we used an ACG uORF mutant that was not responsive to ascorbic acid to verify whether the uORF functions in a cis or trans configuration. We separately expressed the ACG uORF in the mutant vector and tested whether it could restore ascorbic acid repression of LUC activity. The presence of the ACG uORF had no effect on any of the treatments and did not complement the mutant ACG uORF. This is consistent with the uORF functioning in a cis configuration with the CDS of GGP.
[0260] Here, we present evidence that ascorbic acid or its precursors interact directly or indirectly with peptides translated by a non-canonical uORF present in the 5'-UTR of GGP, inhibiting the translation of GGP, the rate-limiting enzyme in ascorbic acid biosynthesis. In eukaryotes, reports of protein expression regulation by metabolic pathway products are rare. Typically, gene expression is regulated by signal transduction cascades mediated by receptors acting on transcription factors or by post-translational modification of target proteins (Smeekens et al., 2010, Curr. Opin. Plant Biol. 13, 273). There are reports that 5'-UTR sequences are important for regulating protein expression (Hulzink et al., 2003, Plant Physiol. 132, 75), but reports of gene expression regulation by small molecules via 5'-UTRs are rare (Rahmani et al., 2009, Plant Physiol. 150, 1356), and regulation by uORFs with non-canonical start codons is extremely rare.
[0261] A simple model for the mechanism of action is that the ACG uORF is translated, but under high ascorbic acid conditions, ribosomes stall on the uORF. Under low ascorbic acid conditions, translation terminates at the stop codon and resumes immediately downstream at the initiation ATG of GGP. While the major initiation codon of GGP does not have a clear Kozak sequence, ACG1 does have a relatively strong Kozak sequence, which may allow ribosomes to queue on the GGP mRNA under high ascorbic acid conditions, enabling a rapid response to stress-induced ascorbic acid decline.
[0262] The function of this feedback loop may require the involvement of another factor. This is suggested by the fact that the GGP 5'-UTR from A. eriantha, a kiwifruit species with extremely high fruit ascorbic acid, exhibits ascorbic acid-regulated function in N. benthamiana. The high ascorbic acid content in A. eriantha may indicate a breakdown in the ascorbic acid feedback control of GGP translation in this species. However, the fact that the ascorbic acid control of GGP in A. eriantha functions in N. benthamiana suggests that a factor mediating the interaction between ascorbic acid and the ACG1 uORF exists and is functional in N. benthamiana. This factor is likely a protein.
[0263] There are two types of uORFs (Tran et al., 2008, BMC Genomics 9, 361): sequence-independent uORFs, in which translation of the uORF influences the reinitiation efficiency of downstream ORFs but the peptide sequence encoded by the uORF is not important (the 10-amino acid ATG short uORF in the 5'-UTR of GGP is thought to be one of these types), and sequence-dependent uORFs, in which the peptide synthesized by the uORF causes ribosome stalling during translation elongation and termination. The GGP uORF encodes a peptide that is highly conserved across a wide range of plant taxa, and a single amino acid mutation in the uORF abolishes ascorbic acid repression, suggesting that it is of the latter type. In plants, regulation by polyamines and sucrose (Rahmani et al., 2009, Plant Physiol. 150, 1356; Gong and Pua, 2005, Plant Physiol. 138, 276) are examples of sequence-dependent uORF regulation. Our novel example differs from these in that translation is initiated from a highly conserved non-canonical start codon.
[0264] In conclusion, we have shown that ascorbate concentration in leaves can be regulated via feedback regulation by ascorbate through a non-canonical uORF contained in the long 5'-UTR of GGP, a regulatory gene for ascorbate biosynthesis. We present evidence that this feedback acts post-transcriptionally by controlling the translational level of the GGP enzyme. We propose that this is the primary mechanism by which ascorbate concentration is regulated in the control of ascorbate biosynthesis in the L-galactose pathway.
[0265] Materials and Methods Plant materials and chemical assays A transient reporter gene system in Nicotiana benthamiana leaves was developed using luciferase (LUC) as a promoter-specific reporter and renilla (REN) as a transformation reporter, as previously described (Hellens et al., 2005, Plant Methods 1, 13). Leaf ascorbic acid concentrations were manipulated by co-injecting Agrobacterium tumefaciens with the pGreen vector (Hellens et al., 2000, Plant Mol. Biol. 42, 819), which contains the coding sequence for GGP from Actinidia chinensis under the 35S promoter. Alternatively, a knockout vector was constructed using the GGP sequence from N. benthamiana (constructed from seven ESTs in GenBank; Snowden et al., The Plant Cell 17, 746). Furthermore, we also used the previously reported CDS of GME from A. eriantha (GenBank accession number: FG424114) to synergistically enhance ascorbic acid with GGP (Bulley et al., 2009, J. Exp. Bot. 60, 765). Ascorbic acid was measured using an HPLC-based assay in extracts from the same leaves (Rassam et al., 2005, J. Agric. Food Chem. 53, 2322). Total ascorbic acid was measured by reducing the extracts prior to HPLC (ibid.). Measurement of the redox state of ascorbic acid revealed a significant decrease in redox potential with increasing ascorbic acid concentration. This suggests that the effects of ascorbic acid may be exerted via ascorbic acid itself or a decrease in redox potential with increasing ascorbic acid.
[0266] In different leaves of N. benthamiana, ascorbic acid affects LUC activity but has little effect on REN activity.
[0267] To verify whether the ascorbic acid effect was mediated by the promoter or the 5'-UTR, we constructed two vectors in which the 5'-UTR regions of the GGP and TT8 promoters were swapped. The resulting constructs were the TT8 core promoter (TT8P') followed by the GGP 5'-UTR (GGPUTR), and vice versa. These vectors were transiently transfected into leaves, and the relative LUC activity was measured.
[0268] To examine the effect of the 5'-UTR preceding the GGP coding sequence on ascorbic acid concentration, we used a different GGP gene (GenBank accession number: FG460629) from the GGP used in other experiments. This protein shares 96% identity with the standard GGP, and a version lacking the 5'-UTR increased ascorbic acid concentrations similarly to standard GGP. The 5'-UTR-containing version contained the entire 5'-UTR, while the UTR-less version had a deletion of 37 base pairs upstream of the ATG start codon using XhoI restriction enzyme. This region (the 3' end of the 5'-UTR) had low homology to the GGPs. Both versions were ligated into the pART277 vector (Gleave, A., 1992, Plant Mol. Biol. 20, 1203).
[0269] LUC protein levels were measured by Western blot analysis using 50 μg of soluble cellular protein from each construct transiently expressed in N. benthamiana leaves (extraction buffer: 40 mM phosphate buffer (pH 7.4), 150 mM NaCl) with anti-LUC antibody (Promega). The RuBisCO large subunit was visualized by SYPRO Tangerine protein staining as an indicator of loading.
[0270] To separate the effects of GGP protein from those of ascorbic acid, we initially attempted to inject ascorbic acid or its precursor directly into leaves using a syringe, but failed to obtain sustained changes in ascorbic acid concentrations in the leaves. We also introduced LUC / REN constructs using Agrobacterium, then imbibed the precursors into leaf segments or leaf disks for uptake. However, although these resulted in a significant increase in ascorbic acid in these leaves, the leaves deteriorated before LUC / REN levels reached measurable levels. Next, we attempted to reduce ascorbic acid alone without reducing GGP concentrations by knocking out two genes involved in ascorbic acid synthesis (galactose dehydrogenase and GDP-mannose epimerase), but failed to produce significant changes in ascorbic acid in the leaves. This suggests that the target enzymes may be stable or may have been present in excess during the experimental period (7 days).
[0271] The experiment was repeated at least twice with similar results. In some cases, high concentrations of ascorbic acid reduced both REN and LUC expression, but this did not alter the effect of ascorbic acid on the GGP promoter (i.e., the decrease in the slope of the LUC / REN ratio), whereas no such effect was observed on the TT8 promoter.
[0272] [Gene cloning and plasmids] The GGP promoter from Actinidia eriantha (SEQ ID NO: 101) was cloned by genome walking and has been deposited in GenBank under accession number JX486682. Genomic DNA from A. eriantha (2.0 pg) was digested with seven blunt-end-generating restriction enzymes: DraI, Ecl136II, EcoRV, HpaI, MscI, ScaI, SspI, and StuI. The digested product was purified and eluted using a PCR Clean and Concentrate spin column (Zymogen), and the eluate was reduced to 10 μL. A double-stranded adapter sequence (Clontech, containing nested PCR primer sites) was ligated to the fragment using T4 Rapid Ligase (Roche) overnight at 16°C. The ligated product was again column-purified and eluted in 30 μL. First-round PCR was carried out using 1 μL of each digest with primers 319998NRWLK1 and RPH-149 and Ex Taq polymerase (Takara) under the following two-step cycle conditions: * First high stringency step: 94°C, denaturation for 2 minutes, followed by 7 cycles of 94°C (25 seconds) and 72°C (3 minutes) * Second step: 32 cycles of 94°C (25 s), 67°C (3 min), followed by a final extension at 67°C for 3 min The first-round PCR product was electrophoresed on a 1% agarose gel, and 1 μL of the 1:50 diluted PCR product was used as a template for second-round PCR using primers 319998NRWLK2 and RPH-150. This was also a two-step PCR, and the conditions were as follows: * Initial denaturation: 94°C (2 min) → 5 cycles of 94°C (25 sec) and 72°C (3 min) * Then: 20 cycles of 94°C (25 sec), 67°C (3 min), followed by a final extension at 67°C (3 min) PCR products (ranging from 500 bp to 2 kb) were cloned into the pGem T Easy vector (Promega) according to the manufacturer's instructions. DNA sequencing of the clones confirmed overlap with known 5'-UTRs. Primers were designed to the ends of the sequences obtained in the first promoter walk to extend the A. eriantha promoter sequence to 2 kb. This 2 kb promoter sequence was PCR amplified from A. eriantha gDNA using primers Eriantha gDNA PCR 5' and 319998NRWLK2 and subcloned into the pGreen0800-5'_LUC vector (Hellens et al., 2000, Plant Mol. Biol. 42, 819) using EcoRV and NcoI restriction enzymes. The final construct is referred to as "GGP-promoter-pGreenII0800-5 LUC vector."
[0273] The GGP promoter (SEQ ID NO: 102) from Arabidopsis thaliana (At4g26850) and control promoters from several sources were also cloned by PCR, including TT8 (AT4G09820), EF1α (AT1G07940), Act2P (AT3G18780), and Act7P (AT5G09810).
[0274] Inactivation of the start codon in the uORF of the 5'-UTR and other deletions or mutations were performed by chemical synthesis of the mutation and control sequences (GenScript, Inc., www.genscript.com). For the inactivated versions, the ATG or ACG start codon of the uORF was replaced with TTG. Other mutations were performed by site-directed mutagenesis. The StuI restriction enzyme site, located 28 bp from the 5' end of the 5'-UTR, served as the 5' end of the synthetic fragment. CC was added to the 3' end to form an NcoI site (ccatgg), and the corresponding sequence was removed from GGP-promoter-pGreenII0800-5 LUC with StuI and NcoI. The synthetic fragments were individually cloned into vectors to construct two versions, one with the uORF and one without.
[0275] [RNA extraction and cDNA synthesis] Total RNA was extracted from 100 mg of leaf tissue using the RNasy Plant Mini kit (Qiagen), and its concentration was measured using a Nanodrop 1000 spectrophotometer (Thermo Fisher Scientific). Next, cDNA was synthesized using 1 μg of total RNA and random hexamers using the BluePrint Reagent Kit for Real Time PCR (Takara Bio) according to the manufacturer's instructions. After synthesis, the cDNA preparation was diluted 75-fold for quantitative real-time PCR.
[0276] [Quantitative PCR] Quantitative PCR was performed using a LightCycler® 480 Real-Time PCR System (Roche Diagnostics) in a total reaction volume of 5 μL, and the following primer pairs were used: * LUC1 / 2: 5'-TATCCGCTGGAAGATGGAAC-3'(SEQ ID NO: 109) * 5'-TCCACCTCGATATGTGCATC-3'(SEQ ID NO: 110) Primers were designed using Primer3 (Rozen and Skaletsky, 2000, Methods Mol Biol. 132, 365) with an annealing temperature of 60°C. The luciferase primer pair amplifies the 5' end region of the luciferase ORF. The reaction components were as follows: 2 μM of each primer, 1.25 μL of diluted cDNA, and LightCycler® 480 SYBR Green I Master Mix. A standard cycling protocol with a Tm of 60°C was used, and relative quantification normalized to the Renilla transcript was performed using the LightCycler® 480 software.
[0277] [Control test of the system used] Because the control gene promoter TT8 expresses approximately 10-fold more strongly than the GGP promoter, we suspected that the TT8-LUC construct might saturate mRNA transcription or LUC translation in tobacco cells, preventing the observed inhibitory effect. To test this hypothesis, we gradually increased the amount of Agrobacterium containing the TT8-LUC construct over a 200-fold range. The LUC / REN ratio and the slope between LUC and REN did not change with the amount of Agrobacterium injected, indicating no saturation of reporter gene expression (Table 2). The LUC values driven by the TT8 promoter overlapped with those driven by the GGP promoter. We also tested whether several alternative promoters (EF1α, Act2P, and Act7P) were inhibited by ascorbic acid, but none of these were negatively affected by ascorbic acid (Table 3). In this experiment, TT8 promoter activity was actually increased by ascorbic acid. As a third test, to confirm whether the effect of ascorbic acid on GGP promoter activity was specific to kiwifruit GGP, we performed a similar test with the Arabidopsis thaliana GGP promoter (At4g26850), and found that the LUC / REN ratio responses were essentially identical between promoters from different species.
[0278] As a final test, we investigated whether expressing the GGP gene (derived from kiwi) to increase ascorbic acid would affect the results. To do so, we introduced an additional control gene, a methyltransferase. In this experiment, the LUC / REN ratio of the added control gene slightly decreased (Table 4), but this did not change our conclusion that increasing ascorbic acid reduces the strength of the GGP promoter. Other control promoters had little effect (Liang et al., U.S. Patent No. 9,648,813).
[0279] <Example 2> Verification of the effect of ascorbic acid on other 5'UTR sequences in the LUC / REN reporter assay [method] In this example, a 35S promoter-driven LUC construct was prepared by deriving it from pGreen 0800 LUC (Hellens et al., 2000, Plant Mol. Biol. 42, 819). A second copy of the 35S promoter, sans the 5' UTR, was cloned into the multiple cloning site immediately preceding the LUC coding sequence. The 5' UTRs from apple, potato, and tomato (SEQ ID NOs: 126, 127, and 128, respectively) were then inserted between the 35S promoter and the start of the LUC coding sequence, replacing the original 5' UTR of 35S. These 5' UTR sequences are shown in uppercase in SEQ ID NOs: 126-128.
[0280] Tomato GGP 5’UTR atttgttcggtatactgtaaccccctgtttgcgattggccttgtagccccgttttacatcttccagagactccatttgtatcggttcacatacagtagcaaagcgccattatcttactctaccccattggcaaacccacagccacaattttccaatcctccattatcccttctacaattttctatataaatacccacatctctctgctctactcccttattatcaacaacaaccaccaaatttcttcttttttttcttcgatagtagcaatctatcaacaaaaacagagaccccatcacaagaatcttggaattttagtgttgggtttaagaggaaaaggggttattgtattttgcagttttgagggtaaagcccagtttaacaagttgtagacatcACGGCTATACACAAAGTAAACCGCCGACCACTTTTACATGTTCCAGCAGTACGTCGTAAGGGTTGTGTAACAGCTACTAACCCTGCGCCGCACGGTGGACGTGGCGCTTTGCCTTCTGAAGGTGGTAGTCCTTCCGACCTCCTCTTCCTTGCCGGCGGCGGTTCTTTCCTCTCCTTCTCCTACtagatatagttatacttactatagatctctagcttattacgtacagttgtatctagtattctattgattattcgaagaaaacacacaaaaagaagtaaagcc (SEQ ID NO: 126) Potato GGP 5’ UTR taagggggtgcttatataaagttggggagtctaccaatgagacgaactcattgaccaaatacgtctgcaggagaaagaccaccggagcaccaaacgccacccaacaaccacccattaaattcttccagaaaaaaacatcttcctcaaaattatcgatgaaggatcgttccttagtagttgl80ttcgttgatcctacaaattcaatcACGGCTCTTCTTGGATCTTTCGTTTGTATTCTCACAATTCATCATCACCGCAAAGTGTTGACCCTTAATCCAACTCTTCTGGTGGACGATAAGCACCGGACCCCTTCCCCTCACGGAGGTAGGGGTGCCTCACCCGCTGAAGGCGGTTGCCCCTCCGATCTCCTCTTCCTCGCCGGCGGCGGTCCAATTCTTCCTTTCTCTTTCTCCTTCTCCtaatttttcgtgtaagaattgtatttttgattatccatccaagaacaggaccgcc (SEQ ID NO: 127) Apple GGP 5' UTR ccacggtacaccctcagccacgaacaccccttcttctccccacacctataaatccaccccctcatctctcctccccacacccccactcacttcagttcgaaacaggcgatcctcgcctttctgggttgtttcctattttatctgagg gagaagaaaggaaggtgtttgatcaattttttggtatatttttaggggtaagacccaggttcgacgagttgtagacatcACGGCTATACACGGAGCTCCTCGGCCGCTCATTCATGTCCGGGCTGTCCGACGAAAGGGTTGTGTA ATTGAGAGCAACCCTTCGCCGCACGGCGGGCGTGGCGCTTTGCCTTCCGAAGGCGGTAGCCCCTCCGACCTGCTCTTCCTCGCTGGTGGCGGTTCTGCATCCTCTGTTTTTCTCTTCTGCTTATATtagcttttttagactttct tggttagattcttaggagattttagagatttttttcttctataaagcgcacgagtagatcgtattgttgttttcggggggttttgggtttggtggtgttgattttactgagaattaagaaaaaataaaaggaaaaaaaagaga gagagaaagaaggggaggagcatgcc (SEQ ID NO: 128)
[0281] These constructs were tested as described for the other GGP promoter constructs in Example 1 above.
[0282] [result] The LUC gene expression driven by each of the different GGP-derived 5' UTR insertion sequences was repressed by ascorbic acid (see Figures 16, 17, and 18, Table 7). In each case, both the LUC / REN ratio and the slope of the relationship between LUC and REN were reduced by ascorbic acid. In these experiments, the increase in ascorbic acid was less than that shown in Example 1 (see Table 7). This may be due to differences in growth conditions (e.g., reduced light and increased temperature). However, increasing ascorbic acid significantly reduced LUC values. Corresponding data for a typical non-responsive promoter-5' UTR (EF1 alpha) and the canonical GGP promoter-5' UTR are also shown. Furthermore, other variant 5' UTR sequences disclosed herein can, of course, be tested using similar methods (see Liang et al., U.S. Patent No. 9,648,813).
[0283] [Table 7]
[0284] <Example 3> Orthogonal Tandem by BLAST Analysis Using Polypeptides Encoded by Major ORFs Identification of uORFs in proteins Sequence similarities between mORFs (major open reading frames) in different species can be used to identify uORFs present in heterologous genes encoding homologs. As an example, AT1G01060.1 (LHY), encoding a putative MYB-related transcription factor involved in the circadian clock, was identified as a novel uORF-containing gene candidate. The Arabidopsis LHY protein sequence was used to identify an orthogonal LHY gene in Brassica oleracea. The AT1G01060.1 sequence was subjected to a sequence homology alignment search against the Brassica oleracea genome using BLAST (tblastn) at genomevolution.org / coge / CoGeBlast.pl, and similar analyses were performed for multiple species. The results are shown in Table 8. Similarly, in this application, the mORF encoding GGP from one species is used to identify GGP proteins in other species.
[0285] [Table 8]
[0286] Example 4: Improvement of rice traits by increasing the expression of GGP protein Rice plants exhibiting increased shoot number, improved nutrient utilization efficiency, increased drought tolerance, increased tolerance to other abiotic stresses, increased disease resistance, increased tolerance to oxidative stress, or tolerance to high accumulation of reactive oxygen species were produced using the following method. A gene encoding a rice-derived GGP protein or its homologue was ligated to a heterologous promoter (e.g., rice actin, tubulin, 35S, or drought-inducible promoters (RD29A, RD29B) or the promoter of the Arabidopsis thaliana gene AT5G43840, or a disease-inducible promoter (e.g., AT1G35230 or the disease-inducible promoter described in U.S. Patent No. 7,994,394)) and introduced into rice cells using an expression vector. Transformants (so-called events) were regenerated and then evaluated by oxidative stress, disease, heat, salt stress, or water deprivation tests (water deprivation until control plants wilted). Individuals exhibiting milder stress symptoms than the control were selected.
[0287] Example 5: Improvement of rice traits through increased expression of GGP protein by creating a non-natural allele containing a mutation in the uORF of the GGP gene In this example, rice plants exhibiting phenotypes such as increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved oxidative stress tolerance, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: Mutant variants with sequence changes within the uORF (or putative uORF) in the 5' UTR from the gene encoding the rice GGP protein are constructed and linked to a reporter gene encoding luciferase. A control construct is constructed in which the entire uORF in the 5' UTR from GGP is deleted. These constructs are then introduced into Nicotiana benthamiana cells or other eukaryotic cells, and reporter expression levels are compared. Mutations that exhibit luciferase activity levels lower than those of the uORF deletion mutants but higher than those of a construct containing the non-mutated 5' UTR of the GGP gene are selected for introduction into the rice genome. The selected mutation is then introduced into rice cells via guide RNA to generate an allele containing the edit or with comparable strength. Cells with the targeted mutation introduced into their genomes are selected and regenerated into plant lines. Each regenerated plant line, or its progeny (obtained by selfing, crossing, or cloning), is subjected to a salt stress assay, an oxidative stress assay, a heat tolerance assay, or a drought tolerance assay. In the drought tolerance assay, water supply is stopped and continued until visible stress symptoms, such as wilting, appear in control plants that do not have the gene editing mutation. Individual lines that show milder stress symptoms than the control are selected.
[0288] Example 6: Wheat trait improvement through increased expression of GGP protein In this example, wheat plants exhibiting the phenotype of increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved oxidative stress tolerance, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: A gene encoding a wheat GGP protein or a homolog thereof is ligated to a heterologous promoter (for example, but not limited to, rice actin, tubulin, 35S, drought-inducible promoter, RD29A, RD29B, or a promoter derived from the Arabidopsis thaliana gene AT5G43840, or a promoter derived from the Arabidopsis thaliana gene AT1G35230, or a disease-inducible promoter described in U.S. Pat. No. 7,994,394) and introduced into wheat cells using an expression vector. Multiple transformant lines (so-called events) carrying the transgene insertion are regenerated and subjected to oxidative stress assays, disease stress assays, heat stress assays, salt stress assays, or drought assays. In drought assays, water supply from the plants is stopped until visible stress symptoms (e.g., wilting) appear in a group of control plants that do not contain the transgene. Individual lines that show milder stress symptoms than the control are selected.
[0289] <Example 7> Wheat trait improvement through increased expression of GGP protein by creating a non-natural allele containing a mutation in the uORF of the GGP gene Wheat plants exhibiting phenotypes such as increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved oxidative stress tolerance, and / or improved tolerance to high levels of intracellular free radicals are generated as follows: Mutant variants of the 5' untranslated region (5' UTR) from the wheat GGP protein (SEQ ID NO: 1)-encoding gene are generated, each with a sequence change in the uORF (or predicted uORF). Each of these is linked to a reporter gene encoding luciferase. A control construct is generated in which the entire uORF is deleted in the 5' UTR from GGP. These constructs are introduced into Nicotiana benthamiana cells or other eukaryotic cells, and the reporter gene expression levels are compared. The mutation selected for introduction is one that exhibits lower luciferase activity than the uORF-deleted mutant and higher activity than a construct containing the unmutated 5' UTR of the GGP gene. The selected mutation is introduced into wheat cells using guide RNA, resulting in the selected gene edit or an allele of equivalent strength. Cells that have integrated the targeted mutation into their genome are selected and regenerated into plant lines. Each regenerated plant line or its progeny (obtained by selfing, crossing, or cloning) is subjected to one of the following assays: salt tolerance assay, oxidative stress assay, heat tolerance assay, or drought stress assay (dry-down assay), in which water supply is withheld until visible stress symptoms (e.g., wilting) appear in a group of control plants (plants without the gene-edited mutation). Individual lines that exhibit less severe stress symptoms than the control plants are selected in the assay.
[0290] Example 8: Improvement of corn traits by increasing GGP protein expression Maize plants exhibiting phenotypes such as increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved oxidative stress tolerance, and / or improved tolerance to high levels of intracellular free radicals are produced as follows: A gene encoding the maize GGP protein (SEQ ID NO: 19) or a homolog thereof is linked to an exogenous promoter (e.g., rice actin, tubulin, 35S promoter, drought-inducible promoters (RD29A, RD29B), the promoter of the Arabidopsis thaliana AT5G43840 gene, or the disease-inducible promoter derived from AT1G35230, or the disease-inducible promoter described in U.S. Patent No. 7,994,394, etc.), and introduced into maize cells using an expression vector. After regenerating multiple transformed lines (events) containing the exogenous gene, an oxidative stress assay, a disease stress assay, a heat stress assay, a salt stress assay, or a drought stress assay (water is withheld from control plants until they exhibit stress symptoms such as wilting) is performed, and individuals with milder stress symptoms than non-transformed controls are selected.
[0291] Example 9: Maize trait improvement through increased GGP protein expression by creating a non-natural allele containing a mutation in the uORF of the GGP gene Maize plants exhibiting phenotypes such as increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved oxidative stress tolerance, and / or improved tolerance to high levels of intracellular free radicals are produced as follows: Mutations are introduced into the uORF (or hypothetical uORF) present in the 5'-UTR of the gene encoding the maize GGP protein (SEQ ID NO: 99) to produce mutant variants, and constructs are constructed in which each mutant 5'-UTR is linked to a luciferase gene. As a control, a construct in which the uORF is completely deleted is also produced. Each of these constructs is introduced into Nicotiana benthamiana cells or other eukaryotic cells, and luciferase activity is measured to compare expression levels. Mutations that exhibit luciferase activity lower than the deletion construct but higher than the unmutated construct are selected and designated as candidate mutations for introduction into the maize genome. The selected mutations are then introduced into maize cells via gene editing via guide RNA, creating the mutation or an allele with equivalent functionality. Cells containing the edited mutations are selected and regenerated into individual plants. The regenerated plants, or their progeny through selfing, crossing, or cloning, are subjected to salt stress, oxidative stress, heat stress, or drought stress assays, and grown until the unedited control exhibits stress symptoms. Plants with milder stress symptoms than the control are selected.
[0292] Example 10: Improvement of soybean traits by increasing GGP protein expression Soybean plants exhibiting phenotypes such as increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced as follows: A gene encoding the soybean GGP protein (SEQ ID NO: 3) or a homolog thereof is ligated to an exogenous promoter (e.g., rice actin, tubulin, 35S promoter, drought-inducible promoters (RD29A, RD29B), the promoter of Arabidopsis thaliana AT5G43840, the disease-inducible promoter derived from AT1G35230, or the disease-inducible promoter described in U.S. Patent No. 7,994,394, etc.), and the construct is introduced into soybean cells via an expression vector. After the introduction, multiple transformation lines (events) are regenerated from the transformants, and each line is subjected to oxidative stress assays, disease stress assays, heat stress assays, salt stress assays, or drought stress assays. Comparisons are made at the point when untransformed control plants show clear stress symptoms. Individuals with mild stress symptoms are selected.
[0293] Example 11: Soybean trait improvement through increased GGP protein expression by creating a non-natural allele containing a mutation in the uORF of the GGP gene Soybean plants exhibiting phenotypes such as increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: Multiple variants are produced by introducing mutations into the uORF (or hypothetical uORF) present in the 5'-UTR region of the gene encoding the soybean GGP protein shown in SEQ ID NO: 83, and constructs are constructed by linking each mutant 5'-UTR to a reporter gene encoding luciferase. As a control, a construct is also produced in which the entire uORF in the 5'-UTR of GGP is deleted. These constructs are introduced into Nicotiana benthamiana cells or other eukaryotic cells, and reporter gene expression levels are compared. Mutations that exhibit luciferase activity lower than the deletion construct but higher than the unmutated construct are selected. These mutations are then introduced into soybean cells using genome editing with guide RNA to create an allele with the desired edit or equivalent strength. Cells carrying the targeted mutation in their genomes are selected and regenerated into individual plants. Each regenerated plant, or its selfed, crossed, or clonal progeny, is subjected to salt stress, oxidative stress, heat stress, or drought stress (water deprivation) assays, and evaluated at the point at which unedited control plants exhibit stress symptoms (such as wilting). Plants with milder stress symptoms than the control plants are selected.
[0294] Example 12: Improvement of alfalfa traits by increasing GGP protein expression Alfalfa plants that express increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved oxidative stress resistance, and / or tolerance to high intracellular free radical concentrations are produced as follows: A gene encoding an alfalfa GGP protein or a homolog thereof is linked to an exogenous promoter (e.g., rice-derived actin, tubulin, or 35S promoter, drought-inducible promoter RD29A or RD29B, or disease-inducible promoters AT5G43840 and AT1G35230 derived from Arabidopsis thaliana, or disease-inducible promoters described in U.S. Patent No. 7,994,394, etc.) to construct an expression construct, and then introduced into alfalfa cells. Regenerate multiple transformed lines (events) containing the exogenous gene and perform oxidative stress, disease stress, heat stress, salt stress, or drought stress (water deprivation) assays on each line. Compare these with untransformed control plants at the time when they exhibit stress symptoms. Select plants with mild stress symptoms.
[0295] Example 13: Improvement of alfalfa traits through increased expression of GGP protein by creating a non-natural allele containing a mutation in the uORF of the GGP gene Alfalfa plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: Mutations are introduced into the uORF present in the 5'-UTR region of the gene encoding the alfalfa GGP protein, creating variants, which are then linked to a reporter gene (luciferase) to construct constructs. Control constructs in which the entire uORF is deleted are also prepared. The constructs are introduced into Nicotiana benthamiana cells or other eukaryotic cells, and the luciferase expression levels are compared. Mutations that exhibit expression levels lower than those of the deleted type but higher than those of the unmutated type are selected and used as candidates for introduction into the alfalfa genome. The selected mutations are introduced into alfalfa cells using genome editing via guide RNA to create the targeted mutation or an allele with an equivalent effect. The edited cells are selected and regenerated into plants. The regenerated plants or their progeny, whether selfed, crossed, or cloned, are subjected to salt, oxidative, heat, or drought stress assays and evaluated at the time when unedited control plants show clear stress symptoms. Plants with milder stress symptoms than the control plants are selected.
[0296] Example 14: Improvement of sugarcane traits by increasing GGP protein expression Sugarcane plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: A gene encoding a sugarcane-derived GGP protein (e.g., SEQ ID NO: 153) or a homolog thereof is fused to an exogenous promoter (e.g., rice actin, tubulin, 35S, drought-inducible promoter RD29A or RD29B, or Arabidopsis-derived AT5G43840 or disease-inducible promoter AT1G35230, or a disease-inducible promoter disclosed in U.S. Patent No. 7,994,394) to construct an expression construct. This construct is then introduced into sugarcane cells, and multiple transformation events (transformants) are isolated and regenerated into whole plants. The regenerated individuals are evaluated by oxidative stress assays, disease stress assays, heat stress assays, salt stress assays, or drought stress assays (water deprivation), and individuals showing mild stress symptoms are selected under conditions in which untransformed control plants show stress symptoms such as wilting.
[0297] <Example 15> Improvement of sugarcane traits through increased expression of GGP protein by creating a non-natural allele containing a mutation in the uORF of the GGP gene Sugarcane plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: Mutations are introduced into the uORF (or hypothetical uORF) in the 5'-untranslated region (5'UTR) of the sugarcane gene encoding the GGP protein (e.g., SEQ ID NO: 153), creating multiple variants, and each of these mutant 5'UTRs is linked to a reporter gene encoding luciferase. A control construct containing a 5'UTR lacking all uORFs is also produced. These constructs are introduced into Nicotiana benthamiana cells or other eukaryotic cells, and luciferase expression levels are compared and evaluated. Mutations that show expression levels lower than the deletion type but higher than the unmutated type are selected and designated as mutations to be introduced into the sugarcane genome. The selected mutations are introduced into sugarcane cells using guide RNA, and the corresponding mutation or a non-natural allele with equivalent strength is constructed by genome editing. The edited cells are selected and regenerated into plants. The regenerated plants or their progeny derived from selfing, crossing, or cloning are subjected to salt stress, oxidative stress, heat stress, or drought stress assays, and individuals with a mild stress response are selected under conditions in which unedited control plants show stress symptoms such as wilting.
[0298] <Example 16> Character improvement of Miscanthus by increasing expression of GGP protein Miscanthus plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: A gene encoding the Miscanthus GGP protein (SEQ ID NO: 151) or a homolog thereof is fused to an exogenous promoter (e.g., rice actin, tubulin, 35S, RD29A, RD29B, or disease-inducible promoters such as AT5G43840 and AT1G35230 derived from Arabidopsis thaliana, or those disclosed in U.S. Patent No. 7,994,394) to produce an expression construct. This construct is then introduced into Miscanthus cells, and multiple transgenic lines (events) are regenerated to produce individual plants. The obtained plants are subjected to oxidative stress assays, disease stress assays, heat stress assays, salt stress assays, or drought stress assays (water deprivation), and individuals with high stress tolerance are selected under conditions in which non-transformed control individuals exhibit stress symptoms.
[0299] Example 17: Improvement of Miscanthus traits through increased expression of GGP protein by creating a non-natural allele with a mutation in the uORF of the GGP gene Miscanthus plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method. Specifically, multiple mutants are created by introducing sequence mutations into the uORF (or hypothesized uORF) present in the 5' untranslated region (5'UTR) of the gene encoding the Miscanthus GGP protein (SEQ ID NO: 150). These mutants are then linked to a reporter gene encoding luciferase to construct reporter constructs. For comparison, a control construct is also prepared, containing a 5'UTR in which the entire GGP uORF has been deleted. Each of these constructs is then introduced into eukaryotic cells, such as Nicotiana benthamiana cells, and the reporter gene expression levels are compared. Mutations that exhibit luciferase activity lower than that of deletion mutations but higher than that of the unmutated 5' UTR are selected and introduced into the Miscanthus genome. The selected mutation is introduced into Miscanthus cells via guide RNA, and a non-natural allele with equivalent effect is created by genome editing. The edited cells are selected and regenerated into plants. The regenerated plants, or their descendants (self-pollinated, cross-bred, clones, etc.), are subjected to salt stress assays, oxidative stress assays, heat stress assays, or drought stress assays (water deprivation), and individuals that exhibit milder symptoms than control plants are selected under conditions that show stress symptoms such as wilting.
[0300] <Example 18> Improvement of almond traits by increasing GGP protein expression Almond plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: A gene encoding an almond-derived GGP protein or its homolog is fused to an exogenous promoter (e.g., rice actin, tubulin, 35S promoter, RD29A, RD29B, or Arabidopsis-derived AT5G43840, or the disease-inducible promoter AT1G35230, or the disease-inducible promoter disclosed in U.S. Patent No. 7,994,394, etc.) to produce an expression construct. This is then introduced into almond cells to obtain transformed plants. The resulting transformants are subjected to oxidative stress assays, disease stress assays, heat stress assays, salt stress assays, or drought stress assays, and individuals with high stress tolerance are selected under conditions in which non-transformants show symptoms such as wilting.
[0301] Example 19: Improvement of almond traits through increased expression of GGP protein by introducing a non-natural allele with a mutation in the uORF of the GGP gene Almond plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are generated by the following method: Multiple mutant sequences are generated by introducing mutations into the uORF (or hypothetical uORF) in the 5' UTR of the almond GGP gene, and each is linked to a reporter gene encoding luciferase to create a reporter construct. A control construct containing a 5' UTR lacking the entire uORF is also generated. These constructs are then introduced into Nicotiana benthamiana cells or other tissues, and the luciferase expression levels are compared and evaluated. Mutations that exhibit lower activity than the deleted construct but higher activity than the unmutated construct are selected and designated as the mutations to be introduced. The selected mutations are then introduced into almond cells using guide RNA and then generated as unnatural alleles by corresponding genome editing. The edited cells are selected and regenerated into plants, or their progeny (self-pollination, cross-breeding, clones, etc.) are subjected to salt stress assays, oxidative stress assays, heat stress assays, and drought stress assays, and individuals with milder stress symptoms than the control are selected.
[0302] Example 20: Improvement of Eucalyptus traits by increasing GGP protein expression Eucalyptus plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: Specifically, a construct is constructed in which a gene encoding a Eucalyptus-derived GGP protein or a homolog thereof is fused to an exogenous promoter (e.g., rice actin, tubulin, 35S promoter, drought-inducible promoters RD29A and RD29B, or Arabidopsis thaliana-derived AT5G43840 promoter, or disease-inducible promoters such as AT1G35230, or the disease-inducible promoters described in U.S. Pat. No. 7,994,394), and this construct is then introduced into Eucalyptus cells using an expression vector. Various transformed lines (so-called events) are regenerated, and each line is subjected to oxidative stress assays, disease stress assays, heat stress assays, salt stress assays, or drought stress assays involving water deprivation. Lines that show fewer stress symptoms than untransformed control plants, such as wilting, are selected under these conditions.
[0303] Example 21: Transformation of Eucalyptus through increased expression of GGP protein by introducing a non-natural allele with a mutation in the uORF of the GGP gene Eucalyptus plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method. In this method, multiple mutant sequences are created by introducing mutations into a uORF (or hypothesized uORF) contained in the 5' untranslated region (5' UTR) of the Eucalyptus GGP gene, and reporter constructs are constructed by linking each of these mutant sequences to a reporter gene encoding luciferase. A control construct in which the entire uORF is deleted is also constructed. These constructs are introduced into Nicotiana benthamiana cells or other eukaryotic cells, and the luciferase expression levels are measured and compared. Mutations that exhibit expression levels lower than those of the deleted construct but higher than those of the unmutated construct are selected, and these mutations are introduced into the Eucalyptus genome. The selected mutations are then introduced into eucalyptus cells via guide RNA, resulting in the desired genome editing or a non-natural allele of equivalent strength. The cells are then selected and regenerated, and the resulting regenerated plants or their progeny (self-pollinated, cross-bred, cloned, etc.) are then subjected to salt stress assays, oxidative stress assays, heat stress assays, and drought stress assays. Individuals that exhibit reduced stress symptoms compared to controls are selected.
[0304] Example 22: Improvement of tomato traits by increasing GGP protein expression Tomato plants exhibiting phenotypes such as increased vitamin C concentration, increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: A gene encoding a tomato-derived GGP protein or its homolog is fused to an exogenous promoter (including rice actin, tubulin, 35S promoter, drought-inducible promoters RD29A and RD29B, or disease-inducible promoters such as Arabidopsis thaliana AT5G43840 or AT1G35230, and the promoter described in U.S. Patent No. 7,994,394) to prepare a construct, which is then introduced into tomato cells using an expression vector. Transformants are regenerated, and the regenerated lines are subjected to various stress assays, including oxidative stress, disease, heat, salt, and drought, to select individuals that exhibit milder symptoms than those observed in untransfected control plants, such as wilting.
[0305] Example 23: Transformation of tomato through increased expression of GGP protein by introducing a non-natural allele with a mutation in the uORF of the GGP gene Tomato plants exhibiting phenotypes such as increased vitamin C concentration, increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are generated by the following method. Specifically, multiple mutants with sequence mutations in the uORF (or hypothesized uORF) in the 5' untranslated region (5'UTR) of the GGP gene from tomato are generated, and reporter constructs are constructed by linking each of these to a luciferase gene. As a control, a construct containing a 5'UTR in which all uORFs are deleted is also generated. These constructs are introduced into Nicotiana benthamiana cells or other eukaryotic cells, and the reporter expression levels are compared. Mutants that exhibit luciferase activity higher than that of the unmutated plant but lower than that of the deleted plant are selected, and gene editing is performed using selected guide RNAs to introduce the mutations into the tomato genome. The cells containing the mutations are selected and regenerated, and the resulting regenerated plants or their progeny (through selfing, crossing, or cloning) are subjected to various assays of salt stress, oxidative stress, heat stress, and drought stress, and lines that show milder stress symptoms than control plants are selected.
[0306] Example 24: Improvement of potato traits by increasing GGP protein expression Potato plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: A gene encoding a potato-derived GGP protein or its homolog is fused to an exogenous promoter (rice actin, tubulin, 35S promoter, drought-inducible promoters RD29A and RD29B, or promoters derived from Arabidopsis thaliana AT5G43840 or disease-inducible promoters such as AT1G35230, or the disease-inducible promoter described in U.S. Patent No. 7,994,394) to prepare a construct, which is then introduced into potato cells using an expression vector. Each transformant line (so-called event) is regenerated and subjected to stress assays involving oxidative stress, disease, heat, salt, drought, and other conditions to select individuals exhibiting reduced stress symptoms compared to untransformed control plants.
[0307] Example 25: Potato trait improvement by introducing a non-natural allele with a mutation in the uORF of the GGP gene This paper provides a method for producing potato plants that exhibit increased biomass, increased shoot number, improved nutrient utilization efficiency, drought tolerance, disease tolerance, oxidative stress tolerance, and free radical tolerance. Multiple mutants with sequence mutations in the uORF (or hypothesized uORF) present in the 5'UTR of the potato GGP gene are created, and each is linked to a luciferase gene to create a reporter construct. A control construct in which the entire uORF is deleted is also created. These are then introduced into Nicotiana benthamiana or other eukaryotic cells, and luciferase activity is compared. Mutations that exhibit activity higher than the unmutated form but lower than the deleted form are selected, and potato cells are gene-edited using the corresponding guide RNA. Mutant-introduced cells are selected and regenerated, and various stress assays are performed using the regenerated plants or their progeny (self-pollination, cross-breeding, clones, etc.) to select individuals with milder symptoms than the control.
[0308] <Example 26> Improvement of avocado traits by increasing GGP protein expression Avocado plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are produced by the following method: A gene encoding the avocado GGP protein or its homolog is fused to an exogenous promoter (e.g., rice actin, tubulin, 35S promoter, RD29A, RD29B, AT5G43840 promoter, AT1G35230, or a disease-inducible promoter described in U.S. Patent No. 7,994,394) to create a construct, which is then introduced into avocado cells using an expression vector. Transformants are regenerated, and stress assays, such as those for oxidative stress, disease, heat, salt, and drought, are performed to select individuals exhibiting reduced stress symptoms compared to untransfected control plants.
[0309] Example 27: Improvement of avocado traits through increased expression of GGP protein by introducing a non-natural allele containing a mutation in the uORF of the GGP gene Avocado plants exhibiting phenotypes such as increased biomass, increased shoot number, improved nutrient utilization efficiency, improved drought tolerance, improved tolerance to other abiotic stresses, improved disease resistance, improved tolerance to oxidative stress, and / or improved tolerance to high levels of intracellular free radicals are generated by the following method: Mutants with sequence mutations in the uORF (or hypothetical uORF) present in the 5' untranslated region (5' UTR) of the avocado GGP protein-encoding gene are generated, and each is linked to a reporter gene encoding luciferase. A control construct is generated in which the entire uORF present in the 5' UTR of GGP is deleted. These constructs are introduced into Nicotiana benthamiana cells or other eukaryotic cells, and the reporter expression levels are compared. Mutations that exhibit lower luciferase activity than deletion mutations but higher than that of constructs containing the unmutated 5' UTR are selected, and avocado cells are gene-edited using guide RNA to introduce the desired gene edit or an allele of equivalent strength. Cells carrying the targeted mutation in their genome are selected and regenerated as plants. Each regenerated plant line or its progeny (obtained by selfing, crossing, or cloning) is subjected to salt stress assays, oxidative stress assays, heat stress assays, or drought assays in which water is not supplied, and the line is observed until stress symptoms (such as wilting) are observed in control plants without the gene-edited mutation. Lines showing milder stress symptoms than the control are selected.
[0310] Example 28: Use of a mutant uORF allele of the tomato GGP gene as a heterozygote to obtain desired traits without adverse effects on growth or organogenesis Tomato plants with elevated GGP protein expression levels can be obtained by carrying a non-natural allele containing a mutation in the uORF of the tomato GGP gene, which results in dwarfism and / or abnormal floral organs (e.g., abnormal anthers, fused floral organs, failure of pollen formation, or reduced pollen yield). The plant is used as the female parent and pollinated by crossing with a wild-type tomato or a heterozygote for the allele to obtain F1 seeds, which, when grown, appear wild-type or essentially wild-type, exhibiting few or no abnormalities, yet have higher ascorbic acid content in the fruit compared to wild-type tomatoes.
[0311] <Example 29> Method for selecting targeted mutations in the 5'UTR upstream of a gene that result in desirable traits in plants The method includes the steps of introducing a mutation into a uORF in the 5' untranslated region (5'UTR), constructing a polynucleotide conjugate in which the mutant 5'UTR is linked to a polynucleotide encoding a reporter protein, and introducing the polynucleotide conjugate into a test plant or plant cell. The 5'UTR contains any of the following: (a) a sequence having 70% or more sequence identity to any one of SEQ ID NOs: 41-100, 111-131, 144-145, 147, 149-150, and 153-155; or (b) a sequence encoding a polypeptide having 70% or more sequence identity to any one of SEQ ID NOs: 1-40, 108, 132-137, 146, 148, 151-152 and 156; The mutation disrupts the function of the upstream open reading frame (uORF) encoded by the 5'UTR, and the modification consists of at least one of the deletion, insertion, or substitution of at least one nucleotide in the 5'UTR. The mutation results in an amount of reporter gene product in the test plant or plant cell that is about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the product level in a control plant containing a control 5'UTR:reporter polynucleotide that partially or completely deletes the uORF and is linked to a reporter gene. The mutation that results in this level of reporter gene product, or an equivalent mutation, is selected as the target mutation to be introduced into the 5'UTR by gene editing at the endogenous locus in the plant genome. Example 30: Selected plants exhibiting desirable traits In this example, the selected plants include: a nucleic acid comprising a 5' UTR (untranslated region) located upstream of a gene of interest that directly or indirectly controls a desired trait; the 5' UTR comprising a mutation in an upstream open reading frame (uORF) created by targeted gene editing, wherein the sequence change resulting from the editing was selected by linking the 5' UTR containing the mutation to a polynucleotide encoding a reporter protein and introducing the polynucleotide into a test plant or plant cell; and the 5'UTR comprises at least one of the following: i) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, about 100%, or 100% sequence identity to any one of SEQ ID NOs: 41-100, 111-131, 144-145, 147, 149-150, or 153-155; or ii) a sequence encoding a polypeptide having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, about 100%, or 100% sequence identity to any one of SEQ ID NOs: 1-40, 108, 132-137, 146, 148, 151-152, or 156; wherein the mutation disrupts the function of an upstream open reading frame (uORF) encoded by the 5'UTR, and the modification is at least one of a deletion, addition, or substitution of at least one nucleotide in the 5'UTR; and the level of reporter gene product is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the level of reporter product produced in a control plant containing a control 5′UTR:reporter polynucleotide in which a non-functional 5′UTR from which the uORF is partially or completely deleted is linked to the reporter gene; and the plant is selected from a population of plants containing the nucleic acid as having a desirable trait; and the desirable trait is selected from the following: increased shoot number, reduced free radical damage, reduced free radical levels, increased stress tolerance, increased abiotic stress tolerance, increased biotic stress tolerance, increased disease tolerance, increased disease resistance, increased nutrient utilization efficiency, increased nitrogen utilization efficiency, and increased oxidative stress tolerance.
[0312] Example 31: Selected plants of Example 30, wherein the level of product in the test plant or plant cell is less than about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%, or 100% of the level of reporter product produced in the control plant cell.
[0313] Example 32 The selected plant of Example 30, wherein the desirable trait is selected from the group consisting of reduced free radical damage, reduced free radical levels, increased stress tolerance, increased abiotic stress tolerance, increased biotic stress tolerance, increased disease tolerance, increased disease resistance, increased nutrient utilization efficiency, increased nitrogen utilization efficiency, and increased oxidative stress tolerance.
[0314] Example 33. Methods for producing plant cells or plants selected for desirable traits. In this method, the desirable trait is selected from the group consisting of increased ascorbic acid levels, increased biomass, increased shoot number, reduced free radical damage, reduced free radical levels, increased stress tolerance, increased abiotic stress tolerance, increased biotic stress tolerance, increased disease tolerance, increased disease resistance, increased nutrient utilization efficiency, increased nitrogen utilization efficiency, and increased oxidative stress tolerance compared to a control plant. The method comprises the steps of: 1. A method for introducing a mutation into the 5′-UTR of the GGP gene, wherein the 5′-UTR comprises at least one of the following: (i) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, about 100%, or 100% sequence identity to any one of SEQ ID NOs: 41-100, 111-131, 144-145, 147, 149-150, or 153-155; or (ii) a sequence encoding a polypeptide having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, about 100%, or 100% sequence identity to any one of SEQ ID NOs: 1-40, 108, 132-137, 146, 148, 151-152, or 156; Here, the mutation disrupts the function of the upstream open reading frame (uORF) encoded by the 5'UTR, and the mutation is at least one of a deletion, addition, or substitution of at least one nucleotide in the 5'UTR.
[0315] Example 34 The method of Example 33, wherein the mutation is present in a uORF sequence in the 5'UTR.
[0316] Example 35 The method of Example 33, wherein the uORF is a sequence selected from SEQ ID NOs: 41 to 60 and 129 to 131, or a variant thereof having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or about 100%, or 100% sequence identity to any of SEQ ID NOs: 41 to 60 and 129 to 131.
[0317] Example 36: Methods for selecting plants with one or more desirable traits. In this example, the one or more desirable traits are selected from the group consisting of increased ascorbic acid levels, increased shoot number, increased biomass, reduced free radical damage, reduced free radical levels, increased stress tolerance, and increased oxidative stress tolerance compared to a control plant. The method comprises the steps of: selecting in the plant for the presence of a first polymorphism in a polynucleotide comprising a sequence encoding a polypeptide having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or about 100%, or 100% sequence identity to an amino acid sequence selected from SEQ ID NOs: 1-20 and 132-134, or selecting for a further polymorphism linked to said first polymorphism; wherein said first polymorphism inhibits expression of said polypeptide, thereby producing a selected plant having said one or more desirable traits.
[0318] <Example 37> The method of Example 36, wherein the selected plant is further selected for one or more of increased biomass, increased shoot number, reduced free radical damage, reduced free radical levels, increased stress tolerance, increased abiotic stress tolerance, increased biotic stress tolerance, increased disease tolerance, increased disease resistance, increased nutrient utilization efficiency, increased nitrogen utilization efficiency, or increased oxidative stress tolerance compared to the control plant.
[0319] Example 38 The method of Example 36, wherein the method comprises separating the selected plants from one or more non-selected plants.
[0320] [Table 9] JPEG2025535075000010.jpg187159JPEG2025535075000011.jpg211159JPEG2025535075000012.jpg240159 JPEG2025535075000013.jpg254163JPEG2025535075000014.jpg241159JPEG2025535075000015.jpg254165
Claims
1. a selected transformed crop plant comprising a polynucleotide sequence comprising a constitutive, tissue-specific, or inducible promoter linked to a DNA coding sequence, the DNA coding sequence encodes a polypeptide that is a homolog of an Arabidopsis GGP protein or has at least 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or about 100%, or 100% identity to any of SEQ ID NOs: 157-242; the transformed plants are selected for desirable traits; The selected transformed crop plants exhibit desirable traits selected from the group consisting of increased ascorbic acid levels, increased biomass, increased shoot number, reduced free radical damage, reduced free radical levels, increased stress tolerance, increased abiotic stress tolerance, increased biotic stress tolerance, increased disease resistance, increased disease resistance, increased nutrient utilization efficiency, increased nitrogen utilization efficiency, and increased oxidative stress tolerance.
2. 2. The selected transformed crop plant of claim 1, wherein said transformed plant is selected for said desirable trait from a group of plants containing said polynucleotide sequence.
3. 2. The selected transformed crop plant of claim 1, wherein the crop plant is a tomato or a member of the Solanaceae family.
4. 2. The selected transformed crop plant of claim 1, wherein the crop is a non-food crop.
5. 2. The selected transformed crop plant of claim 1, wherein said crop plant is Miscanthus or other perennial herbaceous plant.
6. 2. The selected transformed crop plant of claim 1, wherein the crop plant is Eucalyptus, almond, or other tree species.
7. 2. The selected transformed crop plant of claim 1, wherein the promoter sequence has at least 70% identity to a promoter sequence selected from 35S, rice actin, maize tubulin, RD29A, RD29B, ubiquitin promoter, drought-inducible promoter (including the drought-inducible promoters disclosed in U.S. Patent No. 8,895,305), or disease-inducible promoter.
8. 2. The selected transformed crop plant of claim 1, wherein the plant has been confirmed to exhibit desirable traits without other undesirable traits such as organ abnormalities and / or stunted growth.
9. A selected plant or plant part carrying a non-native allele of a gene of interest that controls a desired trait, the allele comprises a selected sequence alteration in a uORF within a 5'UTR; the alteration increases the level of the polypeptide compared to the level of the polypeptide produced by the gene of interest in a control plant; the polypeptide is a homolog of an Arabidopsis GGP protein or has at least 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or about 100% identity to any of SEQ ID NOs: 157-242; 10. A plant or plant part, wherein said plant exhibits a trait selected from the group consisting of the desirable traits of claim 1.
10. 10. The plant or plant part of claim 9, characterized in that it is a plant selected for a desirable trait from a group of plants having a sequence variation in the uORF.
11. 10. The plant or plant part of claim 9, wherein the plant is a tomato or a member of the Solanaceae family.
12. 10. The plant or plant part of claim 9, wherein the crop is a non-food crop.
13. 10. The plant or plant part of claim 9, wherein the crop plant is Miscanthus or other perennial herbaceous plant.
14. 10. The plant or plant part of claim 9, wherein the crop plant is Eucalyptus, almond, or other tree species.
15. 10. The plant or plant part of claim 9, wherein the plant has been confirmed to exhibit desirable traits without other undesirable traits such as organ abnormalities and / or stunted growth.
16. 1. A method for enhancing a desirable trait in a plant, comprising: a. Providing a nucleic acid comprising a 5'UTR (untranslated region) containing an upstream open reading frame (uORF) obtained from a region located upstream of a gene of interest that directly or indirectly enhances the desired trait; b. Attaching the nucleic acid to a reporter gene to create a 5'UTR::reporter polynucleotide; c) introducing a mutation into the nucleic acid to prepare a 5'UTR reporter polynucleotide having a mutation in the uORF; d. introducing the 5'UTR:reporter polynucleotide into a plant cell; e. measuring the level of a product from the reporter gene in the plant cell; f. comparing the reporter gene to a control plant cell containing a control 5'UTR:reporter polynucleotide in which the uORF has been removed from the 5'UTR to identify mutations that result in lower levels of product; and g. introducing the uORF mutation identified in the reporter analysis, or an equivalent mutation, into a uORF at an endogenous locus in the genome of a target plant or target plant cell by gene editing or selection, selecting from the plant or plant cell population, and confirming the presence of the enhanced desirable trait; The 5'UTR comprises at least one of the following (iii) to (v): (iii) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, about 100%, or 100% identity to any of SEQ ID NOs: 41-100, 111-131, 144-145, 147, 149-150, or 153-155; or (iv) a sequence encoding a polypeptide having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, about 100%, or 100% identity to any of SEQ ID NOs: 1-40, 108, 132-137, 146, 148, 151-152, or 156; (v) The method, wherein the 5'UTR has a mutation that is likely to disrupt the function of the uORF encoded by the 5'UTR, and the mutation is a deletion, insertion, or substitution of at least one nucleotide in the 5'UTR.
17. 17. The method of claim 16, wherein the level of reporter product in the plant cells is less than about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%, or 100% of the level of reporter product produced in the control plant.
18. 17. The method of claim 16, wherein the desirable trait is selected from the group consisting of increased ascorbic acid levels, increased biomass, increased shoot number, reduced free radical damage, reduced free radical levels, increased stress tolerance, increased abiotic stress tolerance, increased biotic stress tolerance, increased disease tolerance, increased disease resistance, increased nutrient utilization efficiency, increased nitrogen utilization efficiency, and increased oxidative stress tolerance.
19. 17. The method of claim 16, wherein the selected plant or plant cell is confirmed to exhibit the upregulated desirable trait without other undesirable traits, such as organ abnormalities.
20. 17. The method of claim 16, wherein the selected plant or plant cell is or is derived from a tomato.