BENEFICIAL PLANT TRAITS THROUGH uORF-BASED CONTROL OF CIRCADIAN CLOCK-ASSOCIATED GENES

By mutating uORFs in clock regulator genes using gene editing, the method achieves precise control of protein dosage in plants, enabling the development of crop varieties with improved traits like increased yield and stress tolerance, while maintaining native gene expression patterns.

WO2026015401A1PCT designated stage Publication Date: 2026-01-15GENXTRAITS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/036508
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-07
Filing Date
2025-07-03
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing methods for modifying plant traits through traditional transgenic approaches often disrupt the native transcriptional pattern of clock regulator genes, limiting the ability to fine-tune protein dosage and achieve desired phenotypes such as altered flowering time, yield, or growth optimization.

Method used

Introducing mutations into the upstream open reading frames (uORFs) of clock regulator genes using gene editing techniques, such as CRISPR/Cas, to alter the expression levels of circadian clock-associated proteins, allowing for precise control of protein dosage without disrupting the native expression pattern.

Benefits of technology

This approach enables the development of new crop varieties with enhanced traits like increased yield, biomass, or stress tolerance by fine-tuning protein levels, while preserving the native transcriptional pattern, thus overcoming limitations of traditional transgenic methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025036508_15012026_PF_FP_ABST
    Figure US2025036508_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to methods for disrupting the function of upstream open reading frames (uORFs) in the 5' untranslated regions (UTRs) of plant clock regulator genes by introducing or selecting mutations, for example, using a gene editing construct, such as CRISPR / Cas (including Cas9, Cas12, MAD7, or SHARC™), which contains a polynucleotide encoding a Cas enzyme and at least one guide RNA targeting a uORF sequence in a plant. These guide RNAs specifically target uORFs in 5'UTRs associated with main open reading frames (ORFs) encoding polypeptides that are homologous to, or share at least 30%–100% identity with sequences set forth specified SEQ ID Nos disclosed therein. The disclosure also covers polypeptides (uPEPs) encoded by uORFs within sequences set forth in additional specified SEQ ID NOs or their homologous nucleotide sequences.
Need to check novelty before this filing date? Find Prior Art

Description

BENEFICIAL PLANT TRAITS THROUGH uORF-BASED CONTROL OF CIRCADIAN CLOCK-ASSOCIATED GENES RELATED APPLICATION

[0001] This application claim priority to U.S. Provisional Patent Application No.: 63 / 668,271, filed on July 7, 2024. The entire content of the provisional application is herein incorporated by reference for all purposes.SEQUENCE LISTING

[0002] The instant application contains a Sequence Listing which has been submitted herewith and is hereby incorporated by reference in its entirety. Said .xml copy, created on July 03, 2025, is named 115704-1514268-000610WO, and is 378,131 bytes in size.FIELD OF THE INVENTION

[0003] This disclosure relates to compositions and methods for the development of plant varieties.BACKGROUND

[0004] The biology of most complex multicellular organisms is fine tuned to the twenty-four-hour day night cycles which prevail on planet earth. At the core of this biology, is a set of genes which encode regulatory proteins comprising the so-called “circadian clock” (also sometime simply called the “clock”). These genes have been the subject of intensive research by the plant science community for more than 30 years and much is now known about the interactions between different clock components and the processes they regulate. In particular, the clock controls a host of critical processes including key aspects of development processes such as the timing of flowering and the rate at which a crop matures as well as components of physiology such as the opening and closing of stomata and rhythms of daily leaf movements. In this application, we disclose our finding that the set of genes which encode proteins which are plant clock regulators are subject to so-called uORF regulation. uORFs are short open reading frames which sometimes reside upstream of the main start codon of a gene, and which disrupt the extent to which ribosomes translate the main ORF to produce its encoded protein. Thus, uORFs have likely evolved to fine tune the dosage of protein that is produced from key regulator genes. By introducing mutations into the uORFs of genes that encode clock regulators, we herein provide a novel approach for modifying the cellular level of clock proteins and thereby delivering desired traits such as changes in flowering time, increased yield or biomass, or optimization of a crop’s growth and maturation with respect to the prevailing light regime. For example, the latter traits can fine tune crop varieties to grow at different latitudes or for growth in modem indoor agricultural systems where artificial light can be provided continuously. A recent functional model for the plant clock has been published by Yin Hoon Chew et al., in silico Plants, Volume 4, Issue 2, 2022, diacOlO, doi.org / 10.1093 / insilicoplants / diac010

[0005] Many of the regulator proteins that make up the clock encode transcription factors or related interacting proteins. Herein, we disclose that a majority of the genes that encode key clock components possess uORF sequences in their upstream 5’ leader sequences. Similarly, we have observed that genes from several clock-associated pathways, such as the light signaling network, are also subject to uORF control. In particular, we note that the bZIP transcription factor HY5 and its homologs possess uORFs in their upstream regions. Generally, Upstream Open Reading Frames (uORFs), which reside in the 5' leader region of an mRNA, act as repressor elements that suppress translation of the downstream coding sequence (CDS; Calvo et al, 2009. Proc. Natl. Acad. Sci. USA 106: 7507-7512; Arribere & Gilbert, 2013. Genome Res 23: 977-987). Introducing mutations into the clock or light signaling gene uORF sequences thus provides a means of modifying the extent to which the transcripts from main ORFS of these genes are translated, and in turn, the dose of the clock proteins, that are produced in the cell. As such, introducing sequence changes into the uORFs of clock genes and / or light signaling network components offers a powerful approach to produce new crop varieties with enhanced traits. This offers a number of advantages over traditional transgenic approaches, since unlike transgenesis, a uORF mutation can preserve the native transcriptional pattern (or “expression” pattern) of a gene but markedly alter the dose of protein product that is made in that expression pattern.

[0006] uORFs offer substantial potential for fine tuning the level of cellular regulator proteins to deliver new traits. For example, mutations which delete or portions of the uORFs, or which substitute amino acids, can de-prepress translation of the downstream main ORF and produce a dominant gain of function phenotype. Such an example is afforded by the clock component, GIGANTEA, which when overexpressed in certain “long day” plants (where flowering is typically induced by a long daylength of 16 h of light or more), promotes the floral transition and accelerates flowering and / or maturation. Mutating the uORF of GIGANTEA or its homologs may thus provide a means to speed up or delay the onset of flowering in a target crop, depending on the photoperiodic responses exhibited by that crop. Additionally, clock components such as homologs of CIRCADIAN CLOCK ASSCOCIATED 1 (CCA1) or its paralog LATE ELONGATED HYPOCOTYL (LHY) can delay or prevent flowering when overexpressed. Mutation of the uORFs of LHY / CCA1 homologs may thus produce delayed flowering and / or maturation.

[0007] Conversely, in some instances, the uORF can be manipulated to increase the repression of an associated main ORF, by for example, introducing sequence changes that enhance the extent to which ribosomes will associate with the uORF. Examples of such mutations include modifying a uORF with a non-canonical uORF start codon so as to generate an ATG start codon, or introducing or strengthening a Kozak sequence. A particular instantiation of this approach is to strengthen the start codon of a uORF associated with a HY5 homolog to reduce or eliminate the production of HY5 protein. This will reduce the extent to which light regulated development is promoted, and could produce beneficial traits such as increased root growth and / or increased yield in crops such as soybean.

[0008] The present disclosure provides a means of creating mutations in uORFs DNA repressor elements for the purpose of generating new crop traits.SUMMARY

[0009] The present disclosure is directed to approaches that disrupt the function of uORFs in the 5’UTRs of clock regulator genes by creating or selecting mutations therein. One approach involves introducing gene editing construct or constructs into a plant cell, for example, a CRISPR / Cas construct, that comprises a polynucleotide sequence that encodes a Cas enzyme [e.g., Cas9, Casl2, MAD7, or SHARC™, a proprietary enzyme technology marketed by the company Pairwise (www.pairwise.com / the-fulcrum-platform)], and at least one guide RNA sequence with complementarity to a uORF sequence. The guide RNA sequence targets a target site in a nucleotide sequence that comprises a uORF in a plant. The present description provides for uORFs that reside in 5’UTR regions linked to main ORF sequences that encode polypeptides that are Homologs of or that are at least 30%, 40%, 50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to any of SEQ ID NO: 4, 1, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332, or for a polypeptide (uPEP) encoded by a uORF within any of SEQ ID NO: 5, 6, 2, 3, 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 21, 22, 24, 26, 27, 28, 29, 30, 32, 34, 35, 36, 37, 39, 41, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 55, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132,133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153,155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175,176, 177, 178, 179, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 195, 196, 197, 198,199, 200, 201, 202, 203, 204, 205, 206, 208, 209, 210, 211, 212, 214, 215, 216, 217, 218, 219, 220, 221,222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 240, 241, 242, 243,244, 245, 246, 247, 248, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 263, 264, 265, 266,267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 278, 279, 280, 281, 282, 283, 285, 286, 288, 290, 292,293, 295, 297, 298, 299, 300, 301, 303, -304, 330, 331, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, or 344, or within a homologous polynucleotide sequence of any of the foregoing.

[0010] In the present disclosure, the gene editing construct interacts with a target site that encodes a uORF and edits the uORF which results in the alteration of the expression level or activity of a circadian clock-associated polypeptide in the plant cell. The plant cell is then regenerated into a plant or plant part, which, optionally, is selected to have an elevated level or activity of a circadian clock-associatedpolypeptide and / or a Beneficial Trait. As a result, the plant exhibits at least one desirable attribute or phenotype which may include a trait chosen from the list: increased yield, increased biomass, reduced vegetative biomass, accelerated flowering, delayed flowering, accelerated maturation, delayed maturation, increased root growth, increased root biomass, increased root branching, increased drought tolerance, increased abiotic stress tolerance, improved nutrient use efficiency, reduced lodging, increased pod number, increased stem thickness, increased internode length and reduced internode length.Optionally, the resulting plant can be vegetatively propagated, selfed, grafted, and / or crossed to a second plant. The resulting progeny plants then exhibit a desirable phenotype.

[0011] The present disclosure is also directed to a plant, plant part, or plant cell, the genome of which harbors a non-naturally occurring allele of a gene that encodes a clock-associated protein homolog, or to a uORF encoded small polypeptide that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to a uPEP encoded by a uORF located upstream of a nucleotide sequence that encodes any of SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332, wherein the non- naturally occurring allele comprises a mutation in a uORF within a 5’UTR linked to the main ORF that encodes the clock-associated protein . Examples of potential uORF containing UTR regions are provided by, but not limited to SEQ ID NO: 2, 3, 5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 21, 22, 24, 26, 27, 28,29, 30, 32, 34, 35, 36, 37, 39, 41, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 55, 57, 58, 59, 60, 61, 62, 63,64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91,92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114,115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135,136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 155, 156, 157,158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178,179, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 195, 196, 197, 198, 199, 200, 201,202, 203, 204, 205, 206, 208, 209, 210, 211, 212, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224,225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 240, 241, 242, 243, 244, 245, 246,247, 248, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 263, 264, 265, 266, 267, 268, 269,270, 271, 272, 273, 274, 275, 276, 278, 279, 280, 281, 282, 283, 285, 286, 288, 290, 292, 293, 295, 297,298, 299, 300, 301, 303, 304, 330, 331, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, or 344.The allele in question may be produced by introducing or selecting mutation in the into a uORF that is at least at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or about 100% to a sequence contained within SEQ ID NO: 2, 3, 5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 21, 22, 24, 26, 27, 28, 29, 30, 32, 34, 35, 36, 37, 39, 41, 43, 44, 45,46, 47, 48, 49, 50, 51, 52, 53, 55, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123,124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144,145, 146, 147, 148, 149, 150, 151, 152, 153, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166,167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 181, 182, 183, 184, 185, 186, 187, 188,189, 190, 191, 192, 193, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 208, 209, 210, 211,212, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233,234, 235, 236, 237, 238, 240, 241, 242, 243, 244, 245, 246, 247, 248, 250, 251, 252, 253, 254, 255, 256,257, 258, 259, 260, 261, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 278, 279,280, 281, 282, 283, 285, 286, 288, 290, 292, 293, 295, 297, 298, 299, 300, 301, 303, 304, 330, 331, 333,334, 335, 336, 337, 338, 339, 340, 341, 342, 343, or 344.

[0012] The allele in question may be produced by a directed technique such as gene editing or by selecting a mutant plant carrying the allele from a mutagenized population (such as a TILLING population), with the selection being made based on either the presence of a Beneficial Trait, and / or an elevated level of the polypeptide, or by DNA sequencing to identify mutations in the UTR, as compared to a wild-type plant.

[0013] The plant part carrying the above non-naturally occurring allele may be a root stock that is grafted to a scion to produce a plant that exhibits the desired Beneficial Trait when grown to maturity.

[0014] The present disclosure also pertains to a method of producing new plant variety comprising introducing or selecting a mutation in a cell of the plant species. The mutation is in a uORF sequence within the 5 ’UTR of gene that encodes a polypeptide that is an clock associate protein or a polypeptide that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to a uPEP encoded by a uORF located upstream of a nucleotide sequence that encodes any of SEQ ID NO: 1,4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332, or a polypeptide encoded within any of SEQ ID NO: 2, 3,5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 21, 22, 24, 26, 27, 28, 29, 30, 32, 34, 35, 36, 37, 39, 41, 43,44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 55, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121,122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142,143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164,165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 181, 182, 183, 184, 185, 186,187, 188, 189, 190, 191, 192, 193, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 208, 209,210, 211, 212, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231,232, 233, 234, 235, 236, 237, 238, 240, 241, 242, 243, 244, 245, 246, 247, 248, 250, 251, 252, 253, 254,255, 256, 257, 258, 259, 260, 261, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276,278, 279, 280, 281, 282, 283, 285, 286, 288, 290, 292, 293, 295, 297, 298, 299, 300, 301, 303, 304, 330,331, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, or 344. The cell is then regenerated into a plant or plant part which is selected for an increased level of the polypeptide and the presence of a desirable phenotype, and then multiplying the selected plant or plant part through a process of vegetative propagation, grafting or tissue culture or crossing.BRIEF DESCRIPTION OF THE SEQUENCE LISTING AND DRAWINGS

[0015] The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences of the invention. Traits associated with the use of the sequences are included in the Examples.

[0016] Figures 1-5. These figures represent candidate open reading frames from the 5UTR regions of loci that encode soybean HY5 homologs, which are flanked by two stop codons (*). For each of Figs. 1- 5 the intervening amino acid sequence between the stop codons is identified on the x=0 axis. Potential canonical start codons (ATG), where present, are annotated with a black diamond below the nucleotide sequence and a continuous black lying representing the potential open reading frame(ORF) below the amino acid sequence. Below the amino acid sequence, intervening amino acids encoded by near-cognate codons were identified. Potential nucleotide sequences that could be modified by gene editing or allele selection to create canonical open-reading frames are shown. These are labelled with a black circle, the gene name (e.g. GLYMA_08G302500) and stop-stop fragment number (.1), then the target amino acid, the amino acid coordinate and the potential sequence change (e.g. T29M: denotes a modification where amino acid T at position 29 could be changed to M). The parentheses detail the nucleotide modification needed to convert near cognate to ATG codon (e.g. converting ACG to ATG would result in T to M in translation). Above the amino acid sequence, similar to near cognate Start codons, we've identified potential near cognate stop codons. These are marked with black triangles. Amino acids that could be converted to stop codons are detailed and the necessary nucleotide changes are detailed in parentheses.

[0017] Figure 1 shows GLYMA 08G302500.1 candidate open reading frames.

[0018] Figure 2 shows GLYMA 08G302500.2 candidate open reading frames.

[0019] Figure 3 shows GLYMA 08G302500.4 candidate open reading frames.

[0020] Figure 4 shows GLYMA 08G302500.2 candidate open reading frames.

[0021] Figure 5 shows GLYMA 18G117100.1 candidate open reading frames.

[0022] Figure 6. Vectors used in a protocol to assess the translational regulatory activity of a leader sequence in a reporter system. DNA from dual luciferase vectors (A) is introduced into protoplasts frommaize leaf cells by incubating with polyethene glycol (PEG) or a similar compound that facilitates DNA uptake by cells. These DNA migrate to the nucleus where transcription and translation are initiated from the vectors 35S promoter 2 to 3 days after incubation. The dual luciferase vector contains two promoters that drive the expression of two reported genes. The first of these is firefly luciferase, and the second of these is renilla luciferase. Using the duel luciferase detection kit, the enzymatic activity of the two chemiluminescent reported genes is determined. Transcriptional and / or translational efficiency is then calculated as a ratio of the level of firefly luciferase compared to the level of renilla. In our example (B), the translation efficiency of a uORF -containing (top) leader sequence, is compared to a modified leader sequence where the uORF sequence has been deleted (bottom). The system can be used to test the effect of any mutation within the uORF sequence being analyzed, ranging from a simple base substitution to a complete deletion.

[0023] Figure 7 shows PHY C expression results from analysis of a uORF containing region from PHYC in the dual luciferase assay, using vectors shown in Figure 6. The study was repeated twice; in each instance, many fold higher levels of reporter protein were produced when the uORF was absent from the leader sequence of a polynucleotide encoding the reporter protein, demonstrating that that the uORF has a repressive effect on translation of the reporter protein.

[0024] Figure 8 shows further expression results from analysis of uORF containing regions from the 5’UTR of a selection of genes encoding circadian clock associated proteins in the dual luciferase assay, using vectors shown in Figure 6. Data are plotted as regression slopes with Error Bars. Luciferase (LUC) and Renilla (REN) levels were determined for six uORF -containing genes encoding clock proteins: CCA1 (SEQ ID NO. 306), ELF3 (SEQ ID NO. 12), GI (SEQ ID NO. 4), PHY A (SEQ ID NO. 4), PHYC (SEQ ID NO. 38), and TOC1 (SEQ ID NO. 42). Where the leader (5' UTR) regions contained an intron (e.g., GI), both “intron” and “No-intron” versions were tested. The ratio of LUC to REN is plotted as the slope of a regression analysis, where LUC is under the transcriptional regulation of 35S and translation regulation of the leader sequence of interest. The REN gene is under the transcriptional control of the 35 S promoter, and the error bars represent the residual standard error (RSE) of the regression calculated from six replicates of the LUC and REN measurements for each construct. Leader sequences from the aforementioned clock protein encoding genes that included a candidate uORF are labelled “with uORF” and the corresponding version with a deleted candidate uORF is labelled as such. “uORFl del” or “uORF2 del” is used if the leader sequence contains more than one candidate uORF has been tested. Substantially higher levels of reporter protein were produced when the uORF was absent from the leader sequence of a polynucleotide encoding the reporter protein in the case of CCA 1 , GI (uORF 1), PHY 1 and PHY C, demonstrating that that the uORF has a repressive effect on translation of the reporter protein. By contrast the uORFs from ELF3 and TOC1 appeared to have a repressive effect on translation.

[0025] Figure 9 shows the phenotypic effects of an allele comprising a mutation targeted to a uORF containing region of the gene that encodes CCA1 (SEQ ID NO. 306). Wild-type control plants (Col-0)are shown on the left compared to plants carrying the allele of CCA1, which are shown on the right. The plants are approximately 2 months old and were grown under high-stress glasshouse conditions, where fluctuating heat stress conditions of greater than 24 degrees Centigrade regularly occurred throughout the growth cycle. The control plants showed symptoms of cellular stress including visible purple coloration in the leaf tissue, resulting from anthocyanin accumulation. The plants carrying the CCA1 allele showed much reduced symptoms, were visibly greener and less purple, suggesting that the tissue was protected against the stress by the presence of the allele.DETAILED DESCRIPTION

[0026] The present description is directed to compositions of plants with a desirable phenotype and methods for producing plants with a desirable phenotype.

[0027] The present description relates to polynucleotides and polypeptides, for example, for modifying phenotypes of plants. Throughout this disclosure, various information sources are referred to and / or are specifically incorporated. The information sources include scientific journal articles, patent documents, textbooks, and World Wide Web browser-inactive page addresses, for example. The contents and teachings of the information sources can be relied on and used to make and use embodiments of the invention.

[0028] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural reference unless the context clearly dictates otherwise. Thus, for example, a reference to “a plant” includes a plurality of such plants, and, for example, a reference to “a stress” is a reference to one or more stresses and equivalents thereof known to those skilled in the art, and so forth.DEFINITIONS

[0029] " Beneficial Trait” refers to a characteristic or phenotype exhibited by a plant carrying a mutation or a transgene in its genome, as compared to a control or wild-type plant at a given time in a given growth environment, wherein the characteristic or phenotypes is selected from the following list: increased yield, increased grain yield, increased biomass yield, increased fruit yield, increased grain quality, increased biomass quality, increased fruit quality, delayed flowering, accelerated flowering, delayed maturation, early maturation, increased abiotic stress tolerance, increased drought tolerance, increased salinity tolerance, increased heat tolerance, increased tolerance to low temperature, increased pathogen resistance, increase fungus resistance, increased virus resistance, increased bacterial resistance, increased nematode resistance, increased herbivore resistance, reduced internode length, increased internode length, increased root mass, increased branching, reduced branching, reduced lodging, increased tiller number, increased number of seed bearing stems, increased number of fruit bearing stems, increased harvest index, increased plant height, increased stem width, increased seedling vigor,improved organ growth, increased growth rate, increased pod number, increased number of grains per plant, increased grain weight per planted area, increased biomass per planted area, increased shade tolerance, or suitable for planting at higher planting density per planted area, increased carbon fixation, increased carbon sequestration, increased nutrient use efficiency, increased nutrient uptake, increased nutrient assimilation, increased photosynthesis, increased transpiration, reduced transpiration, reduced leaf temperature, altered circadian rhythm, altered leaf angle, increased protein content of plant tissue, increased lipid content of plant tissue, increased nutrient content of plant tissue, increased carotenoid level of plant tissue, increased flavonoid level of plant tissue.

[0030] “Identity" or "similarity" refers to sequence similarity between two polynucleotide sequences or between two polypeptide sequences, with identity being a more strict comparison. The phrases "percent identity" and "% identity" refer to the percentage of sequence similarity found in a comparison of two or more polynucleotide sequences or two or more polypeptide sequences. "Sequence similarity" refers to the percent similarity in base pair sequence (as determined by any suitable method) between two or more polynucleotide sequences. Two or more sequences can be anywhere from 0-100% similar, or any integer value therebetween. Identity or similarity can be determined by comparing a position in each sequence that may be aligned for purposes of comparison. When a position in the compared sequence is occupied by the same nucleotide base or amino acid, then the molecules are identical at that position. A degree of similarity or identity between polynucleotide sequences is a function of the number of identical or matching nucleotides at positions shared by the polynucleotide sequences. A degree of identity of polypeptide sequences is a function of the number of identical amino acids at positions shared by the polypeptide sequences. A degree of homology or similarity of polypeptide sequences is a function of the number of amino acids at positions shared by the polypeptide sequences.

[0031] The term "variant", as used herein, may refer to polynucleotides or polypeptides that differ from the presently disclosed polynucleotides or polypeptides, respectively, in sequence from each other, and as set forth below.

[0032] With regard to polynucleotide variants, differences between presently disclosed polynucleotides and polynucleotide variants are limited so that the nucleotide sequences of the former and the latter are closely similar overall and, in many regions, identical. Due to the degeneracy of the genetic code, differences between the former and latter nucleotide sequences o may be silent (i.e., the amino acids encoded by the polynucleotide are the same, and the variant polynucleotide sequence encodes the same amino acid sequence as the presently disclosed polynucleotide. Variant nucleotide sequences may encode different amino acid sequences, in which case such nucleotide differences will result in amino acid substitutions, additions, deletions, insertions, truncations or fusions with respect to the similar disclosed polynucleotide sequences. These variations result in polynucleotide variants encoding polypeptides that share at least one functional characteristic. The degeneracy of the geneticcode also dictates that many different variant polynucleotides can encode identical and / or substantially similar polypeptides in addition to those sequences illustrated in the Sequence Listing.

[0033] Also within the scope of the invention is a variant of a nucleic acid listed in the Sequence Listing, that is, one having a sequence that differs from the one of the polynucleotide sequences in the Sequence Listing, or a complementary sequence, that encodes a functionally equivalent polypeptide (i.e., a polypeptide having some degree of equivalent or similar biological activity) but differs in sequence from the sequence in the Sequence Listing, due to degeneracy in the genetic code. Included within this definition are polymorphisms that may or may not be readily detectable using a particular oligonucleotide probe of the polynucleotide encoding polypeptide, and improper or unexpected hybridization to allelic variants, with a locus other than the normal chromosomal locus for the polynucleotide sequence encoding polypeptide.

[0034] “Non-naturally occurring allele” means an allele of a gene that has been produced or selected through human intervention including through application of gene editing, or selection of plants or plant cells harboring a desired sequence change from among a larger population of mutated plants or plant cells.

[0035] The term "plant" includes whole plants, shoot vegetative organs / structures (e.g., leaves, stems and tubers), roots, flowers, and floral organs / structures (e.g., bracts, sepals, petals, stamens, carpels, anthers, and ovules), seed (including embryo, endosperm, and seed coat) and fruit (the mature ovary), plant tissue (e.g., vascular tissue, ground tissue, and the like) and cells (e.g., guard cells, egg cells, and the like), and progeny of same. The class of plants that can be used in the method of the invention is generally as broad as the class of higher and lower plants amenable to transformation techniques, including angiosperms (monocotyledonous and dicotyledonous plants), gymnosperms, fems, horsetails, psilophytes, lycophytes, bryophytes, and multicellular algae. (See for example, from Daly et al. (2001) Plant Physiol. 127: 1328-1333 Ku et al. (2000) Proc. Natl. Acad. Sci. 97: 9121-9126; and see also Tudge, in The Variety of Life, Oxford University Press, New York, NY (2000) pp. 547-606).

[0036] In the context of a plant, “dwarf” refers to an individual or variety which exhibits, as compared to wild-type plant of the same species, a phenotype that comprises a shorter or bushier stature, reduced apical dominance, increased outgrowth of secondary or higher order shoots, increased branching, reduced height, shorter internodes, and / or flowering branches which exhibit a reduced number of vegetative nodes prior to the production of floral nodes.

[0037] A “control plant” as used in the present invention refers to a plant cell, seed, plant component, plant tissue, plant organ or whole plant used to compare against a mutant, gene-edited, transgenic, or genetically modified plant for the purpose of identifying an enhanced phenotype in the mutant, gene- edited, transgenic, or genetically modified plant. A control plant may in some cases be a transgenic plant line that comprises an empty vector or marker gene, but does not contain a recombinant polynucleotideof the present invention that is expressed in the transgenic or genetically modified plant being evaluated. In general, a control plant is a plant of the same line or variety as the transgenic or genetically modified plant being tested. A suitable control plant would include a genetically unaltered or non-transgenic wildtype plant of the parental line used to generate a gene edited, mutant, or transgenic plant herein.

[0038] “Homology” refers to sequence similarity between a reference sequence and at least a fragment of a sequence of interest or a newly sequenced clone insert or its encoded amino acid sequence.

[0039] Homologous protein sequences are those with a common evolutionary origin. There are two types of protein homologs, depending on how they originated: paralogs, derived from a gene duplication event, and orthologs, originated from a speciation event. The ortholog conjecture postulates that orthologous sequences are functionally more similar than paralogous sequences in comparable divergence times (Mier, Perez-Pulido and Andrade-Navarro. BMC Bioinformatics 19, 431 (2018). doi .org / 10.1186 / s 12859-018-2457-y) .

[0040] “CIRCADIAN CLOCK ASSOCIATED PROTEIN Homolog” or “CLOCK ASSOCIATED PROTEIN Homolog” or “Homolog of a CIRCADIAN CLOCK ASSOCIATED PROTEIN” or “CIRCADIAN CLOCK ASSOCIATED POLYPEPTIDE Homolog” or similar, means a protein, which upon performing a BLAST analysis against the set of proteins encoded by the Arabidopsis proteome, returns a higher level of sequence identity to an Arabidopsis circadian clock-associated protein or a paralog of circadian clock-associated protein encoded by an Arabidopsis locus identified by one of the following genome identifier numbers: AT1G01060, AT2G46830, AT2G46790, AT5G02810, AT5G61380, AT1G22770, AT3G46640, AT2G40080, AT2G25930, AT2G43010, AT3G59060, AT4G16780, AT4G16780, AT1G09570, AT2G18790, AT5G35840, AT4G16250, AT5G17300, AT5G37260, AT5G57360 (ZTL), AT1G68050 (FKF1) AT5G11260 (HY5), AT3G17609 (HYH), AT2G32950 (COP1), or AT5G62430 (CDF1), than to any other protein in the Arabidopsis proteome, and which has a level of similarity to the Arabidopsis circadian clock-associated protein equal or higher than an HSP of bit score 50.

[0041] “uPEP” or “uPeptide” as used herein refers to the short peptide which is encoded by a uORF.

[0042] In general, the term “variant” refers to molecules with some differences, generated synthetically or naturally, in their base or amino acid sequences as compared to a reference (native) polynucleotide or polypeptide, respectively. These differences include substitutions, insertions, deletions, or any desired combinations of such changes in a native polynucleotide of amino acid sequence.

[0043] A “conserved domain”, with respect to presently disclosed polypeptides refers to a domain within a polypeptide family that exhibits a higher degree of sequence homology amongst family members than other regions of the polypeptide, such as at least at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 96%, at least 97%, at least 98%, at least 99% or about 100% amino acid residue sequence identity of a polypeptide of consecutive amino acid residues. A fragment or domain can be referred to as outside a conserved domain, outside a consensus sequence, or outside a consensus DNA-binding site that is known to exist or that exists for a particular protein class, family, or sub-family. In this case, the fragment or domain will not include the exact amino acids of a consensus sequence or consensus DNA-binding site of a polypeptide class, family or sub-family, or the exact amino acids of a particular protein consensus sequence or consensus DNA-binding site. Furthermore, a particular fragment, region, or domain of a polypeptide, or a polynucleotide encoding a polypeptide, can be “outside a conserved domain” if all the amino acids of the fragment, region, or domain fall outside of a defined conserved domain(s) for a polypeptide or protein. Sequences having lesser degrees of identity but comparable biological activity are considered to be equivalents.

[0044] As one of ordinary skill in the art recognizes, conserved domains may be identified as regions or domains of identity to a specific consensus sequence. Thus, by using alignment methods well known in the art, the conserved domains of the plant proteins for clock-associated protein homologs, orthologs, or paralogs can be identified.

[0045] A “trait” refers to a physiological, morphological, biochemical, or physical characteristic of a plant or particular plant material or cell. In some instances, this characteristic is visible to the human eye, such as seed or plant size, or can be measured by biochemical techniques, such as detecting the protein, starch, or oil content of seed or leaves, or by observation of a metabolic or physiological process, e.g. by measuring uptake of carbon dioxide, or by the observation of the expression level of a gene or genes, e.g., by employing Northern analysis, RT-PCR, microarray gene expression assays, or reporter gene expression systems, or by agricultural observations such as stress tolerance, yield, or pathogen tolerance. A trait can also include an increased level of a polypeptide, which may be detected by Western blotting, antibody-based techniques, or mass spectrometry. A variety of analytical techniques can be used to measure the amount of, comparative level of, or difference in any selected chemical compound or macromolecule in artificially modified plants, however.

[0046] ‘ ‘Trait modification” refers to producing a detectable difference in a characteristic in a mutant or gene edited plant or transgenic plant ectopically expressing or containing an elevated level of a polynucleotide or polypeptide of the present invention relative to a plant not doing so, such as a wild-type plant. In some cases, the trait modification can be evaluated quantitatively. For example, the trait modification can entail at least about a 2% increase or decrease in an observed trait (difference), at least a 5% difference, at least about a 10% difference, at least about a 20% difference, at least about a 30%, at least about a 50%, at least about a 70%, or at least about a 100%, or an even greater difference compared with a wild-type plant. It is known that there can be a natural variation in the modified trait. Therefore,the trait modification observed entails a change of the normal distribution of the trait in the plants compared with the distribution observed in wild-type plant.

[0047] “Wild type” or “wild-type”, as used herein, refers to a plant cell, seed, plant component, plant tissue, plant organ or whole plant that has not been genetically modified or treated in an experimental sense. Wild-type cells, seed, components, tissue, organs, or whole plants may be used as controls to compare levels of expression and the extent and nature of trait modification with cells, tissue, or plants of the same species in which a polypeptide's expression is altered, e.g., in that it has been knocked out, elevated overexpressed, or ectopically expressed.

[0048] ‘ ‘TILLING” is an acronym for a technique that involves “Targeting Induced Local Lesions inGenomes” whereby a plant breeder creates a mutant population through means of mutagens (such as EMS, X-rays, T-DNAs or transposon) and then individual plants carrying mutations in a target DNA sequence of interest are selected by sequencing that target region from a large number of plants and choosing the individuals that have changes in that target sequence compared to the same sequence from a wild-type plant.

[0049] “Yield” or “plant yield” refers to for example, increased plant growth, increased crop growth, increased pod yield, increased grain yield, increased weight or grain per planted area, increased number of fruit per planted area, increased weight of fruit per planted area, increased fruit size, increased grain size, increased biomass, and / or increased plant product production, and is usually dependent to some extent on temperature, plant size, organ size, planting density, light, water and nutrient availability, and on how the plant copes with various stresses, such as through temperature acclimation and water or nutrient use efficiency.

[0050] “Planting density” refers to the number of plants that can be grown per acre. For crop species, planting or population density varies from a crop to a crop, from one growing region to another, and from year to year. Using com as an example, the average prevailing density in 2000 was in the range of 20,000-25,000 plants per acre in Missouri, USA. A desirable higher population density (a measure of yield) would be at least 22,000 plants per acre, and a more desirable higher population density would be at least 28,000 plants per acre, more preferably at least 34,000 plants per acre, and most preferably at least 40,000 plants per acre. The average prevailing densities per acre of a few other examples of crop plants in the USA in the year 2000 were: wheat 1,000,000-1,500,000; rice 650,000-900,000; soybean 150,000-200,000, canola 260,000-350,000, sunflower 17,000-23,000 and cotton 28,000-55,000 plants per acre (Cheikh et al. (2003) U.S. Patent Application No. 20030101479). A desirable higher population density for each of these examples, as well as other valuable species of plants, would be at least 10% higher than the average prevailing density or yield. In particular, the creation of dwarf varieties can enable planting at high densities.

[0051] ‘ ‘Rootstock” refers to a lower portion of a plant comprising the roots and a portion of the stem which is fused through grafting to an aerial portion of the plant (the “Scion”), the latter generally having a different genotype. The grafting method allows for root beneficial and shoot beneficial phenotypes to be combined in the same plant by bringing together plant parts with different genetic compositions.

[0052] “5 ’untranslated region (UTR)” or “5’UTR” or “5’ UTR” or “5 UTR” means the leader sequence in the 5’ region of a gene, which is transcribed into the messenger RNA, and which is located upstream of the start codon of the main ORF of the gene. In some cases, the 5’ UTR contains one or more short open reading frames or “upstream open reading frames” “uORFs” which may or may not be translated into short peptides (a.k.a “uPEPs”). In some cases, therefore, the term “UTR” sequence is technically a misnomer, and the term “leader” would be a better descriptor of this upstream region of the gene.

[0053] The terms “about” and “approximately” shall generally mean an acceptable degree of error for the quantity measured given the nature or precision of the measurements. Typical, exemplary degrees of error are within 20 percent (%), preferably within 10%, and more preferably within 5% of a given value or range of values. Alternatively, and particularly in biological systems, the terms “about” and “approximately” may mean values that are within an order of magnitude, preferably within 5 -fold and more preferably within 2-fold of a given value. Numerical quantities given herein are approximate unless stated otherwise, meaning that the term “about” or “approximately” can be inferred when not expressly stated.Desirability of crop phenotypes and Beneficial Traits

[0054] Growers seek to improve crop productivity or the attractiveness of their harvested produce to end consumers by planting crop varieties with Beneficial Traits. In many cases it is advantageous to grow smaller or bushier plants. For example, plants with smaller stature may be more resistant to damage from wind and rain, have improved lodging resistance, or may be more resistant to heat, low humidity, or water deficit. Certain of the circadian clock-associated proteins disclosed herein, such as GIGANTEA, may be increased in dosage to produce plants of smaller stature. Such dwarf plants are also of significant interest to the ornamental horticulture industry, and particularly for home garden applications for which space availability may be limited. Conversely, growers of biomass crops typically seek larger plants which offer the capacity for maximizing the amount of carbon fixed and / or sequestered per planting area.

[0055] Growers of grain crops, such as wheat, rice, com, and sorghum, often seek dwarf varieties in order to improve Harvest Index. Harvest index is the ratio of the total grain production of a crop to its biomass and high values are sought so as to capture as much grain yield as possible per area of land. In particular, trait developers are seeking “short com” varieties with increased harvest index. Application of the inventions detailed herein can address this need.

[0056] Growers of fruit trees often have to wait for the plants to progress through a juvenile phase (which in trees may span for many years) before the plant enters the adult phase in which it yields harvestable fruit. Mutation of the uORFs associated with the genes encoding the circadian clock- associated proteins specified herein can reduce the duration of juvenile phase of fruit bearing plants, enabling growers to obtain an earlier harvest. Citrus or Prunus species and avocado are particular examples where such an application might be desired.

[0057] Dwarf plants may exhibit an altered phenotype, as compared to a wild-type plant or a plant of typical stature, characterized by at least one of the following attributes: a) altered auxin transport, b) slower auxin transport, c) reduced apical dominance, d) an altered xylem / phloem ratio, e) an increased number of phloem elements, f) smaller phloem elements, g) thicker bark, h) a bushier habit, i) reduced root mass, j) reduced vigor, k) less vegetative growth, 1) earlier termination of shoot growth, m) earlier competence to flower, n) precocity, o) earlier phase change, p) smaller canopy, q) reduced stem circumference, r) reduced branch diameter, s) fewer sylleptic branches, t) shorter sylleptic branches, u) more axillary flowers, v) an earlier terminating primary axis, w) earlier terminating secondary axes, and x) shorter internode length y) reduced scion mass, or z) smaller size (see patent publication US20180371481, 2018, Foster).

[0058] The plant may be a root stock that is grafted to a scion to produce a plant that exhibits a dwarf phenotype when grown to maturity.

[0059] In contrast to the above examples, growers of commercial forestry or bioenergy crop plantations often seek plants, which have and extended vegetative phase, or which do not enter a reproductive phase, thereby improving biomass yield. As an example, mutation of uORFs associated with genes encoding homologs of the clock-associated proteins CCA and / or LHY may deliver this benefit for plantation forestry.CIRCADIAN CLOCK COMPONENTS

[0060] The plant circadian clock senses photoperiod and establishes the daily rhythms for many developmental processes including seedling growth, organ movements and the onset of flowering and maturation (Kim et al., Mol Plant. 2012 May; 5(3): 152-161). A detailed review of the control of the floral transition by the clock and its effects in crops has recently been published by Yang et al. The Crop Journal Volume 12, Issue 1, February 2024, Pages 17-27). Clock models have described multiple interlocking feedback loops referred to as the morning, core, and evening loops. These loops are interlocked in a complex manner: (1) the core loop including TIMING OF CAB EXPRESSION 1 (TOC1), CIRCADIAN CLOCK ASSOCIATED 1 (CCA1), and LATE ELONGATED HYPOCOTYL (LHY) (Alabadi D, Oyama T, Yanovsky MJ, Harmon FG, Mas P, Kay SA. Reciprocal regulation between TOC1 and LHY / CCA1 within the Arabidopsis circadian clock. Science. 2001;293:880-883; Locke JC, et al. Extension of a genetic network model by iterative experimentation and mathematicalanalysis. Mol. Syst. Biol. 1, 0013. 2005a; Locke JC, Millar AJ, Turner MS. Modelling genetic networks with noisy and varied experimental data: the circadian clock in Arabidopsis thaliana. J. Theor. Biol. 2005b;234:383-393.); (2) the morning loop, inducing PSEUDO RESPONSE REGULATOR 9 (PRR9) and PSEUDO RESPONSE REGULATOR 7 (PRR7), which are linked to CCA1 / LHY (Locke JC, et al. Experimental validation of a predicted feedback loop in the multi-oscillator clock of Arabidopsis thaliana. Mol. Syst. Biol. 2006;2:59; Zeilinger MN, Farre EM, Taylor SR, Kay SA, Doyle FJ., Ill A novel computational model of the circadian clock in Arabidopsis that incorporates PRR7 and PRR9. Mol. Syst. Biol. 2006;2:58.); and (3) the evening loop, including GI and ZEITLUPE (ZTL), which are connected to TOC1 in the core loop (Pokhilko A, et al. Data assimilation constrains new connections and components in a complex, eukaryotic circadian clock model. Mol. Syst. Biol. 2010;6:416). In addition, EARLY FLOWERING 3 (ELF3) acts as a component of the circadian clock input pathway (McWatters HG, Bastow RM, Hall A, Millar AJ. The ELF3 zeitnehmer regulates light signaling to the circadian clock. Nature. 2000;408:716-720) and EARLY FLOWERING 4 (ELF4) has been suggested as another component in the core loop (Doyle MR, et al. The ELF4 gene controls circadian rhythms and flowering time in Arabidopsis thaliana. Nature. 2002;419:74-77; McWatters HG, et al. ELF4 is required for oscillatory properties of the circadian clock. Plant Physiol. 2007;144:391-401).GIGANTEA and its Homologs

[0061] GIGANTEA (GI), is of particular interest as a modulator of crop maturation because it is a well characterized transcriptional regulator of flowering time that is associated with the circadian clock (Fowler et al., 1999, EMBO J. 18(17): 4679-88). GIGANTEA (aka GI, AT1G22770) is an Arabidopsis protein that together with CONSTANS (CO) and FLOWERING LOCUS T (FT), promotes the transition to reproductive development in a circadian clock-controlled pathway. GI acts upstream of CO and FT and promotes increases in CO and FT mRNA abundance. Located in the nucleus, GI regulates several developmental processes, including photoperiod-mediated flowering, phytochrome B signaling, the circadian clock, carbohydrate metabolism, and cold stress responses. GI transcription is controlled by the circadian clock and it is post-transcriptionally regulated by light and dark. The protein is known to form a complex with FKF1 on the CO promoter to regulate CO expression, which in turn results in FT upregulation which triggers flowering. (Niwa, 2007, Plant Cell Physiol 48(7), 925-937).

[0062] GI also interacts with ELF3 , which recruits CONSTITUTIVELY PHOTOMORPHOGENIC 1 (COP1), leading to GI protein degradation during the night (Y u JW, et al. COP1 and ELF3 control circadian function and photoperiodic flowering by regulating GI stability. Mol. Cell. 2008;32:617-630). Additionally, GI can be found at CONSTANS (CO) and FLOWERING LOCUS T (FT) promoters where interactions between GI and FLAVIN-BINDING, KELCH REPEAT F-BOX 1 (FKF1) modulate CO mRNA expression through degradation of the CO repressor, CYCLING DOF FACTOR 1 (CDF1) (Sawa M, Nusinow DA, Kay SA, Imaizumi T. FKF 1 and GIGANTEA complex formation is required for day- length measurement in Arabidopsis. Science. 2007;318:261-265.). GI interacts with SHORTVEGETATIVE PHASE (SVP), TEMPRANILLO (TEM) 1, and TEM2 in vivo and controls expression at the FT promoter (Sawa M, Kay SA. GIGANTEA directly activates Flowering Locus T in Arabidopsis thaliana. Proc. Natl Acad. Sci. U S A. 2011;108: 11698-11703.). Furthermore, GI also functions in hypocotyl growth at the seedling stage (Huq E, Tepperman JM, Quail PH. GIGANTEA is a nuclear protein involved in phytochrome signaling in Arabidopsis. Proc. Natl Acad. Sci. U S A. 2000;97:9789- 9794.; Nozue K, et al. Rhythmic growth explained by coincidence between internal and external cues. Nature. 2007;448:358-361).

[0063] ELF4-deficient mutants show an early flowering phenotype with increased CO expression, while ELF4 overexpressors exhibit delayed flowering (Doyle MR, et al. The ELF4 gene controls circadian rhythms and flowering time in Arabidopsis thaliana. Nature. 2002;419:74-77.; McWatters HG, et al. ELF4 is required for oscillatory properties of the circadian clock. Plant Physiol. 2007;144:391- 401.). ELF4 is also involved in PHYTOCHROME B (PHYB)-dependent seedling growth and it was shown that that an ELF3 / ELF4 / LUX complex binds to PHYTOCHROME INTERACTING FACTOR (PIF) 4 and PIF5 promoters to control their expression (Nusinow DA, et al. The ELF4-ELF3-LUX complex links the circadian clock to diurnal control of hypocotyl growth. Nature. 2011;475:398-402). ELF4 is involved in many of the same physiological processes as GI but the ELF4 / GI interaction has been rarely investigated. Studies in pea (Pisum sativum) have suggested that DIE NEUTRALIS (DNE) and LATE BLOOMER 1 (LATE1), orthologs of ELF4 and GI, respectively, interact genetically to regulate flowering time (Liew LC, et al. DIE NEUTRALIS and LATE BLOOMER 1 contribute to regulation of the pea circadian clock. Plant Cell. 2009;21:3198-3211.)

[0064] Furthermore, overexpression of GI can accelerate or delay flowering depending on the species and its photoperiodic response habit (Mizoguchi, 2005, The Plant Cell, Volume 17, Issue 8, 2255-2270). A detailed explanation of the role of GI and the overall role of the clock and its component proteins in the control of flowering has been published by Yang et al. 2024 (The Crop Journal Volume 12, Issue 1, February 2024, Pages 17-27 and references therein); Yang et al. offer the following discussion, “In the LD plant pea, LATEl / PsGI, the pea homolog of GI promotes flowering by inducing the expression of FT-like genes whereas in SD soybean and maize, the homologs GmGI and ZmGH respectively act as flowering repressors by down-regulating the expression of FT-like genes. The molecular mechanism by which OsGI represses flowering provides an example of flowering regulation in SD plants. In rice, OsGI has been identified as a target of PHOTOPERIODIC SENSITIVITY 5 (SE5), a heme oxygenase involved in phytochrome-chromophore biosynthesis. Overexpression of OsGI results in delayed flowering under both SD and LD conditions. However, mutation of OsGI results in both early and delayed flowering depending on the photoperiod and temperature conditions. Consequently, OsGI is also recognized as a dual-function floral regulator in rice. Like Arabidopsis GI, OsGI physically interacts with OsFKFl and OsCDFl, promoting the expression of Heading date 1 (Hdl), a rice homolog of Arabidopsis CO. Although the GI-CO / Hdl-FT / Hd3a pathway is conserved, Hdl acts as a repressor of Hd3aexpression in rice under LD conditions. This divergence offers an explanation for the differing photoperiodic responses in rice and Arabidopsis. OsGI functions in flowering regulation by modulating a Ghd7-Ehdl-Hd3a pathway, a mechanism absent in Arabidopsis but conserved in SD monocots. OsGI antagonizes the function of OsPHYs in interacting with Ghd7, leading to the facilitation of Ghd7 protein degradation.” However, it has not previously been recognized that the protein levels of GI and its homologs can be modulated and fine-tuned by mutating their associated uORFs to control flowering and / or maturation. uORFs

[0065] An upstream open reading frame (uORF) is a member of a class of small, conserved ORFs located upstream of protein-coding major ORFs (mORFs) (also known as “main ORFs” or “long ORFs”) in the 5 '-untranslated regions (5’UTR) of mRNAs. uORFs are regulatory elements that are prevalent in eukaryotic mRNAs. uORFs act as cis acting elements that modify the activity of a downstream sequence that encodes a polypeptide. As such, they offer a novel opportunity to activate the expression of the longer downstream open reading frames encoding polypeptides of interest, through gene editing approaches that introduce mutations into the uORF sequences. Upregulation of the level of a target polypeptide can thereby be achieved by modulating the activity, or expression, for example, by knocking -out, of a negatively acting uORF upstream of the mORF sequence encoding the polypeptide. Not all eukaryotic genes contain uORFs but hitherto the barrier to the aforementioned approach has been in identifying the uORF sequences; existing algorithms often fail to accurately identify these elements due to their short length and also given that they are often initiated via non-AUG start codons (Hellens et al., 2016. Trend Plant Sci. 21:317-328). In some instances, uORFs are believed to modulate the translation initiation rate of downstream coding sequences (CDSs) by sequestering ribosomes. In other cases, uORFs encode evolutionarily conserved short peptides that may function directly or indirectly as cis-acting modulator peptides of the downstream mORF. In many cases the actual presence of a uORF is strongly conserved across species. Thus, once a uORF has been identified in a target locus from a given species, the homologous locus from another species will typically also contain a uORF and be subject to uORF repression. Identification of the uORFs associated with the main ORFs that encode the polypeptides described herein (SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329 and 332) therefore provides a roadmap for identifying the uORFs contained in the UTRs the genes encoding clock-associated protein homologs from other target species via homology searches that first identify loci that encode homologous polypeptides to the aforementioned list. The UTRs from those homologous loci can then be inspected and analyzed for the presence of uORFs within “stop-stop” regions using the methods detailed herein, including but not limited to dual luciferase reporter-based methods exemplified in Figure 6.

[0066] A small number of examples have been described in plants whereby mutation of a uORF derepresses a downstream mORF encoding a polypeptide of interest, following the identification of the uORF in a first species. An example concerns the GGP clade of proteins which regulate ascorbate levels. A desirable trait comprising increased ascorbate levels was identified through overexpression of the gene in Arabidopsis; that locus contains a uORF, and the equivalent homologous gene in a target crop can be activated to obtain the same desired trait by mutation of the uORF in the crop gene, which will typically reside at a similar position upstream of the mORF in the homologous locus of the target crop. A specific instance concerns editing the uORF of LsGGP2, which encodes a key enzyme in vitamin C biosynthesis in lettuce, which was targeted based on the homologous gene having been demonstrated as being subject to uORF control in Arabidopsis by Laing et al (Liang et al., Plant Cell. 2015 Mar; 27(3): 772-786).Editing the uORF of the lettuce homolog not only increased oxidation stress tolerance, but also increased ascorbate content by -150% (Zhang et al. 2018. Nature Biotechnol. 36:894-898.

[0067] Genome-wide studies have revealed the widespread regulatory functions of uORFs in different species in different biological contexts (Zhang et al. 2019. Trends Biochem. Sci. 44:782-794. doi: 10.1016 / j .tibs.2019.03.002). A given uORF may act as a translational control element for regulating expression of its associated downstream major open reading frame (mORF). The translational regulation of mORFs by highly conserved uORFs in response to cellular metabolite levels has been documented in plant studies (Hayden C.A. and Jorgensen R.A. 2007. BMC Biol. 5:32; Tran M.K., et al. 2008. BMC Genomics 9:361).

[0068] Various methods to identify uORFs in eukaryotes have been described. For example, to identify conserved peptide uORFs, Hayden and Jorgensen created "uORF -Finder", a Perl program that compares the mORF amino acid sequence of cDNAs from one collection with the mORF sequences of another species' collection to identify putative mORF homologs, and then compares uORFs in the 5' UTRs of the two paired sequences to identify uORFs with conserved amino acid sequences (Hayden and Jorgensen, 2007. BMC Biology 5:32). By comparing full-length cDNA sequences from Arabidopsis and rice, distinct homology groups of conserved peptide uORFs are so identified. Skarshewski et al. describe the use of “uPEPperoni”, an online tool for upstream open reading frame location and analysis of transcript conservation (Skarshewski, A., et al. 2014. BMC Bioinform. 15: 36. doi: 10.1186 / 1471-2105- 15-36).

[0069] Rather than making use of bioinformatics-based analysis, Ingolia et al. describe methods for ribosome profiling: identifying uORFs by evaluating ribosome occupancy of upstream open reading frames and other sequences. See, for example, US patent 9,677,068; Ingolia N.T., 2014. Cell Reports 8: 5, 1365-1379. See also Ingolia N.T. 2011. Cell 11; 147: 789-802 in which the authors describe how the majority of putative lincRNAs contain regions of high translation comparable to protein-coding genes. Specific start sites marked by harringtonine followed by ribosome footprints extended to the first inframe stop codon. The majority of novel near-cognate initiation sites detected drive the translation ofuORFs. This is consistent with the high level of translation that is observed on many 5' UTRs as opposed to 3' UTRs, which are almost devoid of ribosomes.Introduction of targeted genetic modifications through gene editing

[0070] One exemplary method of practicing the invention is to use genome editing to produce a“targeted genetic modification” to produce a “non-naturally occurring allele” as referenced herein. The terms “genome editing”, “genome edited”, “genome modified”, “genetically modified” are used interchangeably to describe plants with specific DNA sequence changes in their genomes wherein those DNA sequence changes include changes of specific nucleotides, the deletion of specific nucleotide sequences or the insertion of specific nucleotide sequences.

[0071] As used herein, a technique for introducing a “targeted genetic modification” refers to any method, protocol, or technique that allows the precise and / or targeted editing at a specific location (also referred to a “locus” or “native locus” in a genome of a plant (i.e., the editing is largely or completely non-random) using a site-specific nuclease, such as a meganuclease, a zinc-finger nuclease (ZFN), an RNA-guided endonuclease (e.g., the CRISPR / Cas system), a TALE -endonuclease (TALEN), a recombinase, or a transposase. CRISPR is an acronym for clustered, regularly interspaced, short, palindromic repeats and Cas an abbreviation for CRISPR-associated protein; for a review, see Khandagal & Nadal, Plant Biotechnol. Rep. (2016) 10: 327. Engineered meganucleases, zinc finger nucleases (ZFN), transcription activator-like effector nucleases (TALENs) can also be used. US Patent Application 2016 / 0032297 provides detailed methodology for these methods.

[0072] Genome editing tools can accurately change the architecture of a genome at specific target locations. These tools can be efficiently used for the generation of crop plants with high yields, desired alterations in composition, and resistance to biotic and abiotic stresses. It may be challenging to achieve all desired modifications using a particular genome editing tool. Thus, multiple genome editing tools have been developed to facilitate efficient genome editing. Some of the major suitable genome mechanisms, editing tools, or DNA modification enzymes used to edit plant genomes are: homologous recombination (HR), zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), pentatricopeptide repeat proteins (PPRs), the CRISPR / Cas9 system, the Retron Library Recombineering (RLR) Platform , Cas-CLOVER nucleases, Mini-Cas9 enzymes, RNA interference (RNAi), cisgenesis, and intragenesis. In addition, site-directed sequence editing and oligonucleotide- directed mutagenesis have the potential to edit the genome at the single-nucleotide level. Adenine base editors (ABEs) have been developed to mutate A-T base pairs to G-C base pairs. ABEs use deoxyadeninedeaminase (TadA) with catalytically impaired Cas9 nickase to mutate A-T base pairs to G- C base pairs. A summary of these methods an applicability is provided by Mohanta et al., Genes (Basel). 2017 Dec; 8(12): 399. Gene editing techniques are now available that use a variety of alternative CAS enzymes to CAS9

[0073] Such genome editing methods encompass a wide range of approaches to precisely remove genes, gene fragments, to alter the DNA sequence of coding sequences or control sequences, or to insert new DNA sequences into genes or protein coding regions to reduce or increase the expression of target genes in plant genomes (Belhaj, K. 2013, Plant Methods, 9, 39; Khandagale & Nadal (2016) Plant Biotechnol. Rep. 10: 327). Preferred methods involve the in vivo site-specific cleavage to achieve double stranded breaks in the genomic DNA of the plant genome at a specific DNA sequence using nuclease enzymes and the host plant DNA repair system. Multiple approaches are available for producing double stranded breaks in genomic DNA, and thus achieve genome editing, including the use of the CRISPR / Cas system.

[0074] An extensive overview of the CRISPR / Cas system and useful applications thereof can be found at www.addgene.org / guides / crispr /

[0075] The CRISPR / Cas genome editing system provides flexibility in targeting specific sequences for modification within the genome and enables the execution of a range of different edits including the activation or upregulation of target loci or the knock-out of target loci. The method relies on providing the Cas enzyme and a short guide RNA “gRNA” containing a short guide sequence (~20 bp), with sequence complementarity to the target DNA sequence in the plant genome, for example, a guide RNA with complementarity to a uORF. Depending on the type of Cas enzyme, alternatively a DNA, an RNA / DNA hybrid, or a double stranded DNA guide polynucleotide can be used. The guide portion of this guide polynucleotide directs the Cas enzyme to the desired cut site for cleavage with a recognition sequence for binding the Cas enzyme.

[0076] The target in the plant genome can be any of ~20 nucleotide DNA sequence, provided that the sequence is unique compared to the rest of the genome and also provided that the target sequence is present immediately adjacent to a Protospacer Adjacent Motif (PAM). The PAM sequence serves as a binding signal for the editing enzyme, but the exact sequence depends on which Cas protein is being used. A list of Cas proteins and PAM sequences can be found at www.addgene.org / guides / crispr / #pam- table

[0077] The simplest application of CRISPR / Cas is to produce knockout or loss of function alleles in a target locus. The gRNA targets the Cas enzyme to a specific locus in the genome, which then produces a double stranded break. The resulting DSB is then repaired by one of the general repair pathways present in the cell. This typically causes small nucleotide insertions or deletions (indels) at the DSB site. In most cases, small indels in the target DNA result in amino acid deletions, insertions, or frameshift mutations leading to premature stop codons within the open reading frame (ORF) of the targeted gene. The ideal result is a loss-of-function mutation within the targeted gene. However, the strength of the knockout phenotype for a given mutant cell must be validated experimentally, for example for testing for thepresence of transcript from the target ORF by RT-PCR or hybridization-based approaches. These features make the CRISPR / Cas system a suitable tool for knockout of uORFs.

[0078] CRISPR / Cas can also be used to produce more sophisticated changes to the native sequence at targeted loci in the genome. This can involve inserting sequences, replacing sequences, or editing specific bases so as to insert or create new domains within a polypeptide encoded at a desired locus. One way to introduce such changes is to make use of the high fidelity but low efficiency high fidelity homology directed (HDR) repair pathway within the cell. In order to make such precise modifications using HDR, a DNA repair template incorporating the desired genome modification that the practitioner desires to create at the target locus must be delivered into the cell type of interest with the gRNA(s) and Cas9 or Cas9 nickase. The repair template must contain the desired edit as well as additional homologous sequence immediately upstream and downstream of the target (termed left & right homology arms). The length of each homology arm is dependent on the size of the change being introduced, with larger insertions requiring longer homology arms. Since the efficiency of Cas9 cleavage is relatively high and the efficiency of HDR is relatively low, a large portion of the Cas9-induced DSBs will be repaired to produce edits not comprising the specific desired change. Thus, an additional confirmation / screening step is required to select one of more cells from the edited population that contain the desired change. These cells then can be regenerated into a population of cells, a tissue, organ or whole plant or plant population. Such selection can be achieved by incorporating a marker sequence into the edit, which is readily screened or by PCR or hybridization-based methods.

[0079] CRISPR-related gene editing systems can also be deployed to change specific bases without the need for double stranded breaks. Such approaches are referred to in the art as “base editing” systems. Using these systems, the skilled practitioner can create a targeted genetic modification comprising an amino acid substitution or the creation of start or stop codon.

[0080] To avoid relying on HDR, which has low efficiency, researchers have developed two classes of base editors: cytosine base editors (CBEs) and adenine base editors (ABEs). Cytosine base editors are created by fusing Cas9 nickase or catalytically inactive “dead” Cas9 (dCas9) to a cytidine deaminase like APOBEC. As with traditional CRISPR techniques, base editors are targeted to a specific locus by a gRNA, and they can convert cytidine to uridine within a small editing window near the PAM site. Uridine is subsequently converted to thymidine through base excision repair, creating a C to T change. Likewise, adenosine base editors have been engineered to convert adenosine to inosine, which is treated like guanosine by the cell, creating an A to G change.

[0081] Adenine DNA deaminases do not exist in nature, but these enzymes have been created by directed evolution of the Escherichia coli TadA, atRNA adenine deaminase. Like cytosine base editors, the evolved TadA domain is fused to a Cas9 protein to create the adenine base editor. Both types of base editors are available with multiple Cas9 variants including high fidelity Cas9’s. Further advancementshave been made by optimizing expression of the fusions, modifying the linker region between Cas variant and deaminase to adjust the editing window, or adding fusions that increase product purity such as the DNA glycosylase inhibitor (UGI) or the bacteriophage Mu- derived Gam protein (Mu- GAM).

[0082] While many base editors are designed to work in a very narrow window proximal to the PAM sequence, some base editing systems create a wide spectrum of single -nucleotide variants (somatic hypermutation) in a wider editing window, and are thus well suited to directed evolution applications. Examples of these base editing systems include targeted AID-mediated mutagenesis (TAM) and CRISPR-X, in which Cas9 is fused to activation-induced cytidine deaminase (AID).

[0083] Other CRISPR systems, specifically the Type VI CRISPR enzymes Casl3a / C2c2 and Casl3b, target RNA rather than DNA. Fusing a hyperactive adenosine deaminase that acts on RNA, ADAR2(E488Q), to catalytically dead Casl3b creates a programmable RNA base editor that converts adenosine to inosine in RNA (termed REPAIR). Since inosine is functionally equivalent to guanosine, the result is an A->G change in RNA. The catalytically inactive Casl3b ortholog from Prevotella sp., dPspCasl3b, does not appear to require a specific sequence adjacent to the RNA target, making this a very flexible editing system. Editors based on a second ADAR variant, ADAR2(E488Q / T375G), display improved specificity, and editors carrying the delta-984-1090 ADAR truncation retain RNA editing capabilities and are small enough to be packaged in AAV particles.

[0084] In the context of this disclosure, it is recognized that the term Cas nuclease includes any nuclease which site-specifically recognizes CRISPR sequences based on gRNA or DNA sequences and includes Cas9, Cpfl and others described below. Many authors have identified that CRISPR / Cas genome editing, is a preferred way to edit the genomes of complex organisms (Sander & Joung, 2013, Nat Biotech, 2014, 32, 347; Wright et al., 2016, Cell, 164, 29) including plants (Zhang et al., 2016, Journal of Genetics and Genomics, 43, 151; Puchta, EL, 2016, Plant J., 87, 5; Khandagale & Nadaf, 2016, PLANT BIOTECHNOL REP, 10, 327). US Patent Application 2016 / 020822 provides extensive description of the materials and methods useful for genome editing in plants using the CRISPR / Cas9 system and describes many of the uses of the CRISPR / Cas9 system for genome editing of a range of gene targets in crops.

[0085] It is further recognized that many variations of the CRISPR / Cas system can be used for applying the invention herein, including the use of wild-type Cas9 from Streptococcus pyogenes (Type II Cas) (Barakate & Stephens (2016) Frontiers Plant Sci. 7: 765; Bortesi & Fischer (2015) Biotechnol. Advances 5, 33, 41; Cong et al. (2013) Science 339: 819; Rani et al., (2016) Biotechnol. Letters, 1-16; Tsai et al., (2015) Nature Biotechnol. 33: 187). Other examples include Tru-gRNA / Cas9 in which off- target mutations are significantly decreased (Fu et al. (2014) Nature Biotechnol. 32: 279; Osakabe et al. (2016) Scientific Reports 6: 26685; Smith et al. (2016) Genome Biol. 17: 1; Zhang et al. (2016) Scientific Reports, 6: 28566), a high specificity Cas9 (mutated S. pyogenes Cas9) with little to no off-target activity (Kleinstiver et al.( 2016) Nature 529: 490; Slaymaker et al., (2016) Science 351: 84). Further variationscomprise the Type I and Type III systems in which multiple Cas proteins are expressed to achieve editing (Li et al.( 2016) Nucleic Acids Res. 44:e34; Luo et al., (2015) Nucleic Acids Res. 43, 674), the Type V Cas system using the Cpfl enzyme (Kim et al. (2016) Nature Biotechnol. 34, 863; Toth et al., (2016) Biol. Direct, 11: 46; Zetsche et al., (2015) Cell, 163: 759), DNA-guided editing using the NgAgo Argonaute enzyme from Natronobacterium gregoryi that employs guide DNA (Xu et al. (2016) Genome Biol. 17: 186), and the use of a two vector system in which Cas9 and gRNA expression cassettes are carried on separate vectors (Cong et al. (2013) Science 339: 819). A unique nuclease Cpfl, an alternative to Cas9 has advantages over the Cas9 system in reducing off-target edits which creates unwanted mutations in the host genome. Examples of crop genome editing using the CRISPR / Cpfl system include rice (Tang et. al., (2017) Nature Plants 3: 1- 5; Wu et. al. (2017) Molec. Plant 16: 2017) and soybean (Kim et., al., (2017) Nat. Comm. 8: 14406). Other authors have described the use of Argonaute related proteins as an alternative to CRISPR systems for gene editing (Hegge et al. Nature Reviews Microbiol. (2017) Epub 2017 / 07 / 25. pmid:28736447; Swarts et al. Nucleic acids research. 2015;43(10):5120-9. Epub 2015 / 05 / 01. pmid:25925567; Swarts et al. Nature (2014);507(7491):258-61. Epub 2014 / 02 / 18. pmid:24531762. See also PCT Application Number PCT / US2019 / 025163 and / or Publication Number WO2019204266A1.

[0086] Detailed methodologies for gene editing in plants to create new crop traits, including the selection of cells containing the desired edits, and methods for introducing the CRISPR system components into an initial target plant cell are set forth in published patent application WO2019195157. As specified therein, the “guide polynucleotide” in a CRISPR system also relates to a polynucleotide sequence that can form a complex with a Cas endonuclease and enables the Cas endonuclease to recognize and optionally cleave a DNA target site. The guide polynucleotide can be a single molecule (i.e., a single guide RNA (gRNA) that is a synthetic fusion between a crRNA and part of the tracrRNA sequence) or two molecules (i.e. the crRNA and tracrRNA as found in natural Cas9 systems in bacteria). The guide polynucleotide sequence can be provided as an RNA sequence or can be transcribed from a DNA sequence to produce an RNA sequence. The guide polynucleotide sequence can also be provided as a combination RNA-DNA sequence (see for example, Yin, H. et al., 2018, Nature Chemical Biology, 14, 311). As used herein “guide RNA” sequences comprise a variable targeting domain, called the “guide”, complementary to the target site in the genome (for example, a guide RNA with complementarity to a uORF), and an RNA sequence that interacts with the Cas9 or Cpfl endonuclease, called the “guide RNA scaffold”. A guide polynucleotide that solely comprises ribonucleic acids is also referred to as a “guide RNA”. As used herein the “guide target sequence” refers to the sequence of the genomic DNA adjacent to a PAM site, where the gRNA will bind to cleave the DNA. The “guide target sequence” is often complementary to the “guide” portion of the gRNA, however several mismatches, depending on their position, can be tolerated and still allow Cas mediated cleavage of the DNA. The method also provides introducing single guide RNAs (gRNAs) into plants. The single guide RNAs (gRNAs) include nucleotide sequences that are complementary to the target chromosomal DNA. The gRNAs can be, for example,engineered single chain guide RNAs that comprise a crRNA sequence (complementary to the target DNA sequence) and a common tracrRNA sequence, or as crRNA-tracrRNA hybrids. The gRNAs can be introduced into the cell or the organism as a DNA with an appropriate promoter, as an in vitro transcribed RNA, or as a synthesized RNA. Basic guidelines for designing the guide RNAs for any target gene of interest are well known in the art as described for example by Brazelton et al. (Brazelton, V.A. et al. (2015) GM Crops & Food, 6: 266-276) and Zhu (Zhu, L. J. (2015) Frontiers Biol. 10: 289-296).

[0087] Published patent applications WO2019195157 and WO2019204266A1 also provide example of the types of mutation that can lead to increased activity of transcription factor polypeptides. These include mutations to the coding sequence that give rise to amino acid changes in the encoded protein.

[0088] In certain preferred embodiments of the present invention, the guide polynucleotide / Cas endonuclease system can be used to allow for the insertion or deletion of a promoter or promoter element, such as an enhancer element, upstream of the coding sequence of any one the polypeptide sequences of the invention, wherein the promoter insertion (or promoter element deletion) results in any one of the following or any one combination of the following: an elevated level of the polypeptide, a permanently activated gene locus, an increased promoter activity (increased promoter strength), an increased promoter tissue specificity, a decreased promoter tissue specificity, a new promoter activity, an extended window of gene expression, a modification of the timing or developmental progress of gene expression, a mutation of DNA binding elements and / or an addition of DNA binding elements.

[0089] The guide RNA / Cas endonuclease system can be used to allow for the insertion of a promoter element to increase the expression of the polypeptide sequences of this disclosure. Promoter elements, such as enhancer elements, are often introduced in promoters driving gene expression cassettes in multiple copies for trait gene testing or to produce transgenic plants expressing specific traits. Enhancer elements can be, but are not limited to, a 35S enhancer element (Benfey et al, EMBO J, (1989) 8(8): 2195-2202). In some plants (events), the enhancer elements can cause a desirable phenotype, a yield increase, or a change in expression pattern of the trait of interest that is desired. It may be desired to remove the extra copies of the enhancer element while keeping the trait gene cassettes intact at their integrated genomic location. The guide RNA / Cas endonuclease can be used to remove the unwanted enhancing element from the plant genome. A guide RNA can be designed to contain a variable targeting region targeting a target site sequence of 12-30 bps adjacent to a NGG (PAM) in the enhancer. The Cas endonuclease can make cleavage to insert one or multiple enhancers.

[0090] To repress the function of a target polypeptide encoding locus, the promoter elements to be deleted can be, but are not limited to, promoter core elements, promoter enhancer elements or 35 S enhancer elements (CaMV35S enhancers (Benfey et al, EMBO J, August 1989; 8(8): 2195-2202)). The promoter or promoter fragment to be deleted can be endogenous, artificial, pre-existing, or transgenic to the cell that is being edited. Preferably the promoter element is endogenous to the cell that is beingedited. These approaches, removing promoters or promoter elements, can be applied to knock-out polypeptides of interest.

[0091] In a further embodiment, the targeted genetic modification of interest is an intron site of any one of the polypeptide sequences of the disclosure, wherein the modification consists of inserting an intron enhancing motif into the intron which results in modulation of the transcriptional activity of the gene containing said intron. In yet another embodiment, methods provide for modifying alternative splicing sites of any one of the polypeptide sequences of the disclosure resulting in enhanced production of the functional gene transcripts and polypeptides (proteins).

[0092] In additional embodiments, the modification of the polypeptide sequences of the disclosure includes editing the intron borders of alternatively spliced genes to alter the accumulation of splice variants. In other embodiments, the guide polynucleotide / Cas endonuclease system can be used to modify or replace a coding sequence of the polypeptide in the genome of a plant cell, wherein the modification or replacement results in any one of the following, or any one combination of the following: an increased protein activity, an increased protein functionality, a site specific mutation, a protein domain swap, a protein knock-out, a new protein functionality, insertion or creation of a transcriptional activation domain, insertion or creation of a transcriptional repression domain, or a modified protein functionality.Delivery of gene editing components into plant cells and plants:

[0093] Sandhya et al., J Genet Eng Biotechnol. 2020 Dec; 18: 25. Published online 2020 Jul 7. doi: 10.1186 / s43141-020-00036-8, present methods for delivering gene editing tools such CRISPR / Cas9 components into plants to execute the gene editing process. The effective delivery of CRISPR / Cas9 components, including the guide sequence, the Cas9, and where applicable a DNA-repair template containing the desired sequence edit, into plant cells is critical for editing to be efficient. The practitioner can select from a variety of delivery methods to introduce the gene editing components into plant cells. These include Agrobacterium-mediated, bombardment or biolistic method, floral-dip, and PEG-mediated protoplast transformation. Additional methods include nanoparticle and pollen magnetofection-mediated delivery systems (Kwak et ai., 2019, Nature Nanotechnology, DOI 10. 1038 / S41565-019-0375-4) (Demirer et al, 2019, Nature Nanotechnology, DOI 10.1038 / S41565-019-0382-5) can be used. CRISPR constructs can be coated onto gold particles for gene gun mediated introduction into plant cells, CRISPR constructs can be transfected into protoplasts using PEG, or introduced via an Agrobacterium strain harboring a CRISPR vector. Components may also be introduced via floral dip (Castel et al., PLoS One. 2019;14(l):e0204778) or a pollen-tube tube pathway-based method. In the next step, a plant cell containing targeted genetic modification produced by the introduced CRISPR system is selected and regenerated into an explant and then a plant tissue or whole plant. In some instances, this procedure involves selecting explants harboring the genome edit on selection plates and regenerating a whole plant. Finally, PCR and Sanger sequencing are generally used for confirmation that the desired sequence edithas been successfully introduced into the selected plant. The selected plant is then examined to confirm that it exhibits the target trait of interest that was initially sought by introducing the genome modification.

[0094] Sandhya et al., J Genet Eng Biotechnol. 2020 Dec; 18: 25. Published online 2020 Jul 7. doi: 10.1186 / s43141-020-00036-8, provide tables showing which methods can be successfully applied to particular crops. For example, the following plants can all be successfully gene edited using PEG mediated delivery of CRISPR system components: Apple, Brassica oleracea, Brassica rapa, Citrullus lanatus, Glycine max, Grapevine, Oryza sativa, Petunia, Physcomitrella patens, Solanum lycopersicum, Triticum aestivum, and Zea mays. By way of further example, the following plants can all be successfully gene edited using particle bombardment mediated delivery of CRISPR system components: Glycine max, Hordeum vulgare, Oryza sativa, Triticum aestivum and Zea mays. By way of a further example, the following plants can all be successfully gene edited using particle bombardment mediated delivery of CRISPR system components: Arabidopsis thaliana, Banana, Citrus sinensis, Cucumis sativum, Glycine max, Kiwi fruit, Lotus japonicus, Marchantia polymorpha, Medicago truncatula, Nicotiana benthamaina, Nicotiana tabacum, Oryza sativa, Populus, Salvia miltiorrhiza, Solanum lycopersicum, Sorghum bicolor, Triticum aestivum and Zea mays.

[0095] Upstream ORFs (uORFs) encode short mRNAs that are encoded within the 5' UTR of a gene encoding a regulatory protein of interest (the coding sequence of which is typically referred to as the long ORF, main ORF, major ORF, or mORF). The uORF is typically out-of-frame with the main coding sequence of the gene of interest, which in an embodiment of the invention herein is the GIGANTEA protein-encoding gene. A substantial proportion of eukaryotic mRNAs contain uORFs in the 5' leader sequence preceding the main functional protein-encoding ORF (Kochetov, (2008) BioEssays 30: 683- 691). uORFs often encode short peptides that modulate the activity of the regulatory proteins encoded by the genes of which they are upstream. Furthermore, uORFs sometimes initiate at a non-canonical codon (e.g., ACG rather than AUG) and the encoded peptide is often much less than 100 residues in length. These features make uORFs challenging to identify via automated bioinformatic searches, and experimentation is typically needed to confirm that a putative uORF functions as a negative regulator of a downstream coding sequence. For example, Laing et al. (2015) Plant Cell 27: 772-786; DOI: 10.1105 / tpc. 114. 133777, removed a uORF that encodes a 60- to 65-residue peptide in the upstream region of GGP in lettuce, and showed this was sufficient to deliver a trait comprising increase levels of ascorbate. In fact, the peptides encoded by uORFs are very short indeed in some instances; for example, in humans, a functional peptide of only 6 amino acids was identified as being encoded by a uORF. Several plant uORFs have been shown to modulate mORF translation in response to the levels of various key metabolites within the cell (e.g., polyamines, sucrose, phosphocholine and ascorbate. A proposed function of several of the peptides encoded by these uORFs is to slow or stall the ribosomes and as a consequence limit translation of the downstream main ORF which encodes to regulatory protein (seeHellens et al., 2016, Trends in Plant Science, Vol. 21, No. 4 pp317. dx.doi.org / 10.1016 / j.tplants.2015.11.005, and references therein).

[0096] Recently, it has been proposed that gene editing of uORFs may offer a general approach to activate crop genes to produce traits of interest in a highly targeted manner (Zhang and Voytas, May (2019) National Science Review 6:(3)391, doi.org / 10.1093 / nsr / nwy 123). In particular, this can avoid many of the drawbacks associated with traditional approaches, which often involve large insertions of foreign DNA fragments, such as sequences of strong promoters, enhancers or engineered artificial transcription activators, in the genome. Indeed, many of the problems associated with genetic modifications that comprise transgene integrations, including lack of consumer acceptance of GM products, may be eliminated in a next generation of crop traits produced through knock-out or mutation of uORFs in regulator genes by targeted gene editing. Knock-out as used herein, generally means removing or reducing the activity of a uORF by creating sequence changes within it. It can include deleting the entire uORF containing region, but not necessarily. Example sequence changes that can be introduced to the uORFs include deletions, insertions, inversions, duplications, translocations and substitutions provided, that they can result in desired Beneficial Traits. These beneficial mutated uORFs can be identified or selected using the methods described herein, for example, using a reporter constructs as disclosed in Example 13. In some instances, a practitioner may design a guide nucleic targeting the uORF containing region associated with a gene of interest and then produce a so-called “allelic series” comprising a set of plants each of which harbors a different gene edited mutation in the uORF containing region. The practitioner then selects the individual plant(s) which possess allele(s) that deliver(s) the desired level of the Beneficial Trait of interest and / or the desired level of protein from the gene of interest.

[0097] An increasing number of uORFs are being identified in the upstream regions of genes that encode transcriptional regulators and, in many cases, the uORF and / or its encoded short peptide appear to be controlled by a metabolic signal (van der Horst 2020, Plant Physiol. (2020) 182(1): 110-122, Published online 2019 Aug 26. doi: 10. 1104 / pp.19.00940 and references therein). These include these the SI -group bZIPs, including the (HG1) bZIP transcription factor, which controls amino acid and sugar metabolism and which in turn has its activity regulated by sucrose. SAC51 (HG15) is bHLH transcription factor which is involved in xylem differentiation and regulated in response to thermospermine. Another example is the HsfBl / TBFl (HG18) HSF transcription factor which is involved in heat tolerance and growth-to-defense transition which is regulated by galactinol.

[0098] A further example of transcription factor regulation by uORFs concerns a multiple uORF- containing genes identified in the transcription factors that are involved in circadian clock regulation, which are the subject of this application.Identifying Homologs (including Orthologs and Paralogs) for targeted genetic modification:

[0099] A homolog of a polypeptide is a related polypeptide that possesses an equivalent or similar function to a given polypeptide, and which can be subjected to a targeted genetic modification to deliver an equivalent or similar trait to a given polypeptide.

[0100] Homologous sequences as described herein can comprise orthologous or paralogous sequences. As used herein, a paralog is a homolog from the same species and an ortholog is a homolog from a different species. Several different methods are known by those of skill in the art for identifying and defining these functionally homologous sequences. General methods for identifying orthologs and paralogs, including phylogenetic methods, sequence similarity and hybridization methods, are described herein; an ortholog or paralog, including equivalogs, may be identified by one or more of the methods described below.

[0101] As described by Eisen, (1998) 10.1101 / gr.8.3.163 Genome Res. 8: 163-167), evolutionary information may be used to predict gene function. It is common for groups of genes that are homologous in sequence to have diverse, although usually related, functions. However, in many cases, the identification of homologs is not sufficient to make specific predictions because not all homologs have the same function. Thus, an initial analysis of functional relatedness based on sequence similarity alone may not provide one with a means to determine where similarity ends, and functional relatedness begins. Fortunately, it is well known in the art that protein function can be classified using phylogenetic analysis of gene trees combined with the corresponding species. Functional predictions can be greatly improved by focusing on how the genes became similar in sequence (i.e., by evolutionary processes) rather than on the sequence similarity itself (Eisen, 1998) supra. In fact, many specific examples exist in which gene function has been shown to correlate well with gene phylogeny (Eisen, 1998) supra. Thus, “[t]he first step in making functional predictions is the generation of a phylogenetic tree representing the evolutionary history of the gene of interest and its homologs. Such trees are distinct from clusters and other means of characterizing sequence similarity because they are inferred by techniques that help convert patterns of similarity into evolutionary relationships. After the gene tree is inferred, biologically determined functions of the various homologs are overlaid onto the tree. Finally, the structure of the tree and the relative phylogenetic positions of genes of different functions are used to trace the history of functional changes, which is then used to predict functions of [as yet] uncharacterized genes” (Eisen, 1998) supra.

[0102] Within a single plant species, gene duplication may cause two copies of a particular gene, giving rise to two or more genes with similar sequence and often similar function known as paralogs. A paralog is therefore a similar gene formed by duplication within the same species. Paralogs typically cluster together or in the same clade (a group of similar genes) when a gene family phylogeny is analyzed using programs such as CLUSTAL (Thompson et al. (1994) Nucleic Acids Res. 22: 4673-4680; Higgins et al. (1996) Methods Enzymol. 266: 383-402). Groups of similar genes can also be identified with pair- wise BLAST analysis (Feng and Doolittle (1987) J. Mol. Evol. 25: 351-360). For example, a clade ofvery similar MADS domain transcription factors from Arabidopsis all share a common function in flowering time (Ratcliffe et al. (2001) Plant Physiol. 126: 122-132), and a group of very similar AP2 domain transcription factors from Arabidopsis are involved in tolerance of plants to freezing (; Gilmour et al. (1998) Plant J.16:433-442). Analysis of groups of similar genes with similar function that fall within one clade can yield sub-sequences that are particular to the clade. These sub-sequences, known as consensus sequences, can not only be used to define the sequences within each clade, but define the functions of these genes; genes within a clade may contain paralogous sequences, or orthologous sequences that share the same function (see also, for example, Mount (2001), in Bioinformatics: Sequence and Genome Analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, page 543).

[0103] Transcription factor gene sequences (CCA1 and LHY are examples of transcription factors) are conserved across diverse eukaryotic species lines (Wang, Z., et al. (1997). Plant Cell (9), 491-507; Wang, Z, et al. (1998). Cell (93), 1207-1217. Goodrich et ai. (1993) Cell 75: 519-530; Lin et al. (1991) Nature 353: 569-571; Sadowski et al. (1988) Nature 335: 563-564). Plants are no exception to this observation; diverse plant species possess transcription factors that have similar sequences and functions. Speciation, the production of new species from a parental species, gives rise to two or more genes with similar sequence and similar function. These genes, termed orthologs, often have an identical function within their host plants and are often interchangeable between species without losing function. Because plants have common ancestors, many genes in any plant species will have a corresponding orthologous gene in another plant species. Once a phylogenic tree for a gene family of one species has been constructed using a program such as CLUSTAL (Thompson et al. (1994) Nucleic Acids Res. 22: 4673-4680); Higgins et al. (1996) Methods Enzymol. 266: 383-402) potential orthologous sequences can be placed into the phylogenetic tree and their relationship to genes from the species of interest can be determined.Orthologous sequences can also be identified by a reciprocal BLAST strategy. Once an orthologous sequence has been identified, the function of the ortholog can be deduced from the identified function of the reference sequence.

[0104] By using a phylogenetic analysis, one skilled in the art would recognize that the ability to predict similar functions conferred by closely related polypeptides is predictable. This predictability has been confirmed by our own many studies in which we have found that a wide variety of polypeptides have orthologous or closely-related homologous sequences that function as does the first, closely-related reference sequence.

[0105] The polypeptides sequences belong to distinct clades of polypeptides that include members from diverse species. In each case, most or all of the clade member sequences derived from both dicots and monocots have been shown to confer increased tolerance to one or more abiotic stresses when the sequences were overexpressed and hence will likely increase yield and or crop quality. These studies each demonstrate that evolutionarily conserved genes from diverse species are likely to function similarly(i.e., by regulating similar target sequences and controlling the same traits), and that polynucleotides from one species may be transformed into closely related or distantly related plant species to confer or improve traits.

[0106] At the nucleotide level, homologs of the sequences of the invention will typically share at least about 30% or 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%.55%, 56%, 57%, 58, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%.73%, 74%, 75%, 76%, 77%, 78%, 79%, 80% , 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%,91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or about 100% nucleotide sequence identity to one or more of the listed full-length sequences, or to a region of a listed sequence excluding or outside of the region(s) encoding a known consensus sequence or consensus DNA-binding site, or outside of the region(s) encoding one or all conserved domains. The degeneracy of the genetic code enables major variations in the nucleotide sequence of a polynucleotide while maintaining the amino acid sequence of the encoded protein.

[0107] At the polypeptide level, homologs of sequences of the invention will typically share preferably at least about 30%, or 35%, or 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%.51%, 52%, 53%, 54%, 55%, 56%, 57%, 58, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%.69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80% , 81%, 82%, 83%, 84%, 85%, 86%,87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or about 100% sequence identity to one or more of the listed full-length sequences. Within a conserved domain, which is often a DNA binding domain, homologs of sequences of the invention will typically share preferably at least about 60%, about 65%, about 70% or about 80% sequence identity, and more preferably about 85%, about 90%, about 95% or about 97% or more sequence identity to the conserved domains of the listed sequences.

[0108] Percent identity can be determined electronically, e.g., by using the MEGALIGN program (DNASTAR, Inc. Madison, Wis.). The MEGALIGN program can create alignments between two or more sequences according to different methods, for example, the clustal method (see, for example, Higgins and Sharp (1988) Gene 73: 237-244). The clustal algorithm groups sequences into clusters by examining the distances between all pairs. The clusters are aligned pairwise and then in groups. Other alignment algorithms or programs may be used, including ENTREZ, FASTA and BLAST, and which may be used to calculate percent similarity. These are available as a part of the GCG sequence analysis package (University of Wisconsin, Madison, WI), and can be used with or without default settings. ENTREZ is available through the National Center for Biotechnology Information. In one embodiment, the percent identity of two sequences can be determined by the GCG program with a gap weight of 1, e.g., each amino acid gap is weighted as if it were a single amino acid or nucleotide mismatch between the two sequences (see USPN 6,262,333).

[0109] Software for performing BLAST analyses is publicly available, e.g., through the National Center for Biotechnology Information (see www.ncbi.nlm.nih.gov). This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al. (1990) J. Mol. Biol. 215: 403-410; Altschul (1993) J. Mol. Evol. 36: 290-300. These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always > 0) and N (penalty score for mismatching residues; always < 0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a wordlength (W) of 11, an expectation (E) of 10, a cutoff of 100, M=5, n=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff (1989) Proc. Natl. Acad. Sci. USA 89: 10915). Unless otherwise indicated for comparisons of predicted polynucleotides, “sequence identity” refers to the % sequence identity generated from a protein sequence blast or a tblastx using the NCBI version of the algorithm at the default settings using gapped alignments with the filter “off” (see, for example, www.ncbi.nlm.nih.gov / ).

[0110] Other techniques for alignment are described by Doolittle, R. F. (1996) Methods in Enzymology: Computer Methods for Macromolecular Sequence Analysis, vol. 266, Academic Press, Orlando, FL, USA. Preferably, an alignment program that permits gaps in the sequence is utilized to align the sequences. The Smith-Waterman is one type of algorithm that permits gaps in sequence alignments (see Shpaer (1997) Methods Mol. Biol. 70: 173-187). Also, the GAP program using the Needleman and Wunsch alignment method can be utilized to align sequences. An alternative search strategy uses MPSRCH software, which runs on a MASPAR computer. MPSRCH uses a Smith- Waterman algorithm to score sequences on a massively parallel computer. This approach improves ability to pick up distantly related matches, and is especially tolerant of small gaps and nucleotide sequence errors. Nucleic acid-encoded amino acid sequences can be used to search both protein and DNA databases.[oni] The percentage similarity between two polypeptide sequences, e.g., sequence A and sequence B, is calculated by dividing the length of sequence A, minus the number of gap residues in sequence A,minus the number of gap residues in sequence B, into the sum of the residue matches between sequence A and sequence B, times one hundred. Gaps of low or of no similarity between the two amino acid sequences are not included in determining percentage similarity. Percent identity between polynucleotide sequences can also be counted or calculated by other methods known in the art, e.g., the Jotun Hein method (see, for example, Hein (1990) Methods Enzymol. 183: 626-645) Identity between sequences can also be determined by other methods known in the art, e.g., by varying hybridization conditions (see Published US Patent Application No. 20010010913). As used herein, percentage identity is ideally based on a comparison of the identity across two stretches of nucleotides or amino acids which are at least ten residues in length, and preferably at least 20 residues and most preferably at least 50 residues.

[0112] Thus, the invention provides methods for identifying a sequence similar or paralogous or orthologous or homologous to one or more polynucleotides as noted herein, or one or more target polypeptides encoded by the polynucleotides, or otherwise noted herein and may include linking or associating a given plant phenotype or gene function with a sequence. In the methods, a sequence database is provided (locally or across an internet or intranet) and a query is made against the sequence database using the relevant sequences herein and associated plant phenotypes or gene functions.

[0113] In addition, one or more polynucleotide sequences or one or more polypeptides encoded by the polynucleotide sequences may be used to search against a BLOCKS (Bairoch et al., 1997), PFAM, and other databases which contain previously identified and annotated motifs, sequences, and gene functions. Methods that search for primary sequence patterns with secondary structure gap penalties (Smith et al. (1992) Protein Engineering 5: 35-51) as well as algorithms such as Basic Local Alignment Search Tool (BLAST; BLAST; Altschul (1993) J. Mol. Evol. 36: 290-300; Altschul et al. (1990) J. Mol. Biol. 215: 403-410), BLOCKS ((Henikoff and Henikoff (1991) Nucleic Acids Res. 19: 6565-6572), Hidden Markov Models (HMM; Eddy (1996) Curr. Opin. Str. Biol. 6: 361-365; Sonnhammer et al. (1997) Proteins 28: 405-420), and the like, can be used to manipulate and analyze polynucleotide and polypeptide sequences encoded by polynucleotides. These databases, algorithms and other methods are well known in the art and are described in Ausubel et al. (1997; Short Protocols in Molecular Biology, John Wiley & Sons, New York, NY, unit 7.7) and in Meyers (1995; Molecular Biology and Biotechnology, Wiley VCH, New York, NY, p 856-853), and in Meyers (1995; Molecular Biology and Biotechnology, Wiley VCH, New York, NY, p 856-853).

[0114] Furthermore, methods using manual alignment of sequences similar or homologous to one or more polynucleotide sequences or one or more polypeptides encoded by the polynucleotide sequences may be used to identify regions of similarity and conserved domains characteristic of a particular transcription factor family, for example, sequences related to CCA / LHY or HY 5. Such manual methods are well-known of those of skill in the art and can include, for example, comparisons of tertiary structure between a polypeptide sequence encoded by a polynucleotide that comprises a known function and a polypeptide sequence encoded by a polynucleotide sequence that has a function not yet determined. Suchexamples of tertiary structure may comprise predicted alpha helices, beta-sheets, amphipathic helices, leucine zipper motifs, zinc finger motifs, proline-rich regions, cysteine repeat motifs, and the like.

[0115] Orthologs and paralogs of presently disclosed polypeptides may be directly synthesized or cloned using compositions provided by the present invention according to methods well known in the art. cDNAs can be cloned using mRNA from a plant cell or tissue that expresses one of the present sequences. Appropriate mRNA sources may be identified by interrogating Northern blots with probes designed from the present sequences, after which a library is prepared from the mRNA obtained from a positive cell or tissue. Polypeptide-encoding cDNA is then isolated using, for example, PCR, using primers designed from a presently disclosed gene sequence, or by probing with a partial or complete cDNA or with one or more sets of degenerate probes based on the disclosed sequences. The cDNA library may be used to transform plant cells. Expression of the cDNAs of interest is detected using, for example, microarrays, Northern blots, quantitative PCR, or any other technique for monitoring changes in expression. Genomic clones may be isolated using similar techniques to those.

[0116] This disclosure encompasses isolated nucleotide sequences that are phylogenetically and structurally similar to sequences listed in the Sequence Listing and can function in a plant by conferring a Beneficial Trait when gene edited or otherwise modified in a plant. One skilled in the art would predict that other similar, phylogenetically related sequences falling within the present clades of disclosed sequences would also perform similar functions when similarly upregulated or downregulated.Identifying Polynucleotides or Nucleic Acids by Hybridization

[0117] Polynucleotides homologous to those encoding the polypeptides sequences illustrated in the Sequence Listing and Table 1 can be identified, e.g., by hybridization to each other under stringent or under highly stringent conditions. Lor example, Selected example circadian clock-associated protein sequences were used for the queries including those provided as SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 305, 306, 307, 308, 309, 310, 311, 312, and 313. Single stranded polynucleotides hybridize when they associate based on a variety of well characterized physical-chemical forces, such as hydrogen bonding, solvent exclusion, base stacking, and the like. The stringency of a hybridization reflects the degree of sequence identity of the nucleic acids involved, such that the higher the stringency, the more similar are the two polynucleotide strands. Stringency is influenced by a variety of factors, including temperature, salt concentration and composition, organic and non-organic additives, solvents, etc. present in both the hybridization and wash solutions and incubations (and number thereof), as described in more detail in the references cited below (e.g., Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y; Berger and Kimmel (1987) Guide to Molecular Cloning Techniques, Methods in Enzymology, vol. 152 Academic Press, Inc., San Diego, CA; and Anderson and Young (1985) "Quantitative Lifter Hybridisation." In: Hames and Higgins, ed., Nucleic Acid Hybridisation, A Practical Approach. Oxford, IRL Press, 73-111.Table.1 Exemplary circadian clock-associated protein sequences and potential uORFs containing regions found within the 5’UTR of the gene encoding the clock proteincodons. The stop-stop fragment is the region containing the uORF, with the second stop being the stop codon of the uORF.

[0118] Encompassed by this disclosure are polynucleotide sequences that are capable of hybridizing to the polynucleotides that encode the claimed polypeptide sequences, including any of the polypeptides within the Sequence Listing, and fragments thereof under various conditions of stringency (see, for example, Wahl and Berger (1987) Methods Enzymol. 152: 399-407; and Kimmel (1987) Methods Enzymol. 152: 507-511). In addition to the nucleotide sequences that encode the polypeptides listed in the Sequence Listing, full length cDNA, orthologs, and paralogs of the polynucleotides encoding the present polypeptide sequences may be identified and isolated using well-known methods. The cDNA libraries, orthologs, and paralogs of the present nucleotide sequences may be screened using hybridization methods to determine their utility as hybridization target or amplification probes.

[0119] With regard to hybridization, conditions that are highly stringent, and means for achieving them, are well known in the art. See, for example, Sambrook et al., 1989 supra; Berger and Kimmel, 1987 supra, pages 467-469; and Anderson and Young, 1985 supra.

[0120] Stability of DNA duplexes is affected by such factors as base composition, length, and degree of base pair mismatch. Hybridization conditions may be adjusted to allow DNAs of different sequence relatedness to hybridize. The melting temperature (Tm) is defined as the temperature when 50% of the duplex molecules have dissociated into their constituent single strands. The melting temperature of a perfectly matched duplex, where the hybridization buffer contains formamide as a denaturing agent, may be estimated by the following equations:(I) DNA-DNA:Tm(° C)=81.5+16.6(log [Na+])+0.41(% G+C)- 0.62(% formamide)-500 / L(II) DNA-RNA:Tm(° C)=79.8+l 8.5 (log [Na+])+0.58(% G+C)+ O.I2(%G+C)2 - 0.5(% formamide) - 820 / L(III) RNA-RNA:Tm(° Q=79.8+ 18.5 (log [Na+])+0.58(% G+C)+ 0.12(%G+C)2 - 0.35(% formamide) - 820 / L where L is the length of the duplex formed, [Na+] is the molar concentration of the sodium ion in the hybridization or washing solution, and % G+C is the percentage of (guanine+cytosine) bases in the hybrid. For imperfectly matched hybrids, approximately 1° C is required to reduce the melting temperature for each 1% mismatch.

[0121] Hybridization experiments are generally conducted in a buffer of pH between 6.8 to 7.4, although the rate of hybridization is nearly independent of pH at ionic strengths likely to be used in the hybridization buffer (Anderson and Young, 1985, supra). In addition, one or more of the following may be used to reduce non-specific hybridization: sonicated salmon sperm DNA or another non- complementary DNA, bovine serum albumin, sodium pyrophosphate, sodium dodecylsulfate (SDS), polyvinyl-pyrrolidone, ficoll and Denhardt's solution. Dextran sulfate and polyethylene glycol 6000 act to exclude DNA from solution, thus raising the effective probe DNA concentration and the hybridization signal within a given unit of time. In some instances, conditions of even greater stringency may be desirable or required to reduce non-specific and / or background hybridization. These conditions may be created with the use of higher temperature, lower ionic strength, and higher concentration of a denaturing agent such as formamide.

[0122] Stringency conditions can be adjusted to screen for moderately similar fragments such as homologous sequences from distantly related organisms, or to highly similar fragments such as genes that duplicate functional enzymes from closely related organisms. The stringency can be adjusted either during the hybridization step or in the post-hybridization washes. Salt concentration, formamide concentration, hybridization temperature and probe lengths are variables that can be used to alter stringency (as described by the formula above). As a general guidelines high stringency is typically performed at Tm-5° C to Tm-20° C, moderate stringency at Tm-20° C to Tm-35° C and low stringency at Tm-35° C to Tm-50° C for duplex >150 base pairs. Hybridization may be performed at low to moderate stringency (25-50° C below Tm), followed by post-hybridization washes at increasing stringencies. Maximum rates of hybridization in solution are determined empirically to occur at Tm-25° C for DNA- DNA duplex and Tm-15° C for RNA-DNA duplex. Optionally, the degree of dissociation may be assessed after each wash step to determine the need for subsequent, higher stringency wash steps.

[0123] High stringency conditions may be used to select for nucleic acid sequences with high degrees of identity to the disclosed sequences. An example of stringent hybridization conditions obtained in a filter-based method such as a Southern or Northern blot for hybridization of complementary nucleic acids that have more than 100 complementary residues is about 5°C to 20°C lower than the thermal melting point (Tm) for the specific sequence at a defined ionic strength and pH. Conditions used for hybridizationmay include about 0.02 M to about 0. 15 M sodium chloride, about 0.5% to about 5% casein, about 0.02% SDS or about 0. 1% N-laurylsarcosine, about 0.001 M to about 0.03 M sodium citrate, at hybridization temperatures between about 50° C and about 70° C. More preferably, high stringency conditions are about 0.02 M sodium chloride, about 0.5% casein, about 0.02% SDS, about 0.001 M sodium citrate, at a temperature of about 50° C. Nucleic acid molecules that hybridize under stringent conditions will typically hybridize to a probe based on either the entire DNA molecule or selected portions, e.g., to a unique subsequence, of the DNA.

[0124] Stringent salt concentration will ordinarily be less than about 750 mM NaCl and 75 mM trisodium citrate. Increasingly stringent conditions may be obtained with less than about 500 mM NaCl and 50 mM trisodium citrate, to even greater stringency with less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of organic solvent, e.g., formamide, whereas high stringency hybridization may be obtained in the presence of at least about 35% formamide, and more preferably at least about 50% formamide. Stringent temperature conditions will ordinarily include temperatures of at least about 30° C, more preferably of at least about 37° C, and most preferably of at least about 42° C with formamide present. Varying additional parameters, such as hybridization time, the concentration of detergent, e.g., sodium dodecyl sulfate (SDS) and ionic strength, are well known to those skilled in the art. Various levels of stringency are accomplished by combining these various conditions as needed.

[0125] The washing steps that follow hybridization may also vary in stringency; the posthybridization wash steps primarily determine hybridization specificity, with the most critical factors being temperature and the ionic strength of the final wash solution. Wash stringency can be increased by decreasing salt concentration or by increasing temperature. Stringent salt concentration for the wash steps will preferably be less than about 30 mM NaCl and 3 mM trisodium citrate, and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate.

[0126] Thus, hybridization and wash conditions that may be used to bind and remove polynucleotides with less than the desired homology to the nucleic acid sequences or their complements that encode the present polypeptides include, for example:0.5X, LOX, 1.5X, or 2X SSC, 0.1% SDS at 50°, 55°, 60° or 65° C , or 6X SSC at 65° C;50% formamide, 4X SSC at 42° C; or0.5X SSC, 0.1% SDS at 65° C; with, for example, two wash steps of 10 - 30 minutes each. Useful variations on these conditions will be readily apparent to those skilled in the art. A formula for “SSC, 20X” may be found, for example, in Ausubel et al., Current Protocols in Molecular Biology, Ausubel et al. eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (supplemented through 2000).

[0127] A person of skill in the art would not expect substantial variation among polynucleotide species encompassed within the scope of the present invention because the highly stringent conditions set forth in the above formulae yield structurally similar polynucleotides.

[0128] If desired, one may employ wash steps of even greater stringency, including about 0.2x SSC, 0.1% SDS at 65° C and washing twice, each wash step being about 30 minutes, or about 0.1 x SSC, 0.1% SDS at 65° C and washing twice for 30 minutes. The temperature for the wash solutions will ordinarily be at least about 25° C, and for greater stringency at least about 42° C. Hybridization stringency may be increased further by using the same conditions as in the hybridization steps, with the wash temperature raised about 3° C to about 5° C, and stringency may be increased even further by using the same conditions except the wash temperature is raised about 6° C to about 9° C. For identification of less closely related homologs, wash steps may be performed at a lower temperature, e.g., 50o C.

[0129] An example of a low stringency wash step employs a solution and conditions of at least 25° C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS over 30 minutes. Greater stringency may be obtained at 42° C in 15 mM NaCl, with 1.5 mM trisodium citrate, and 0.1% SDS over 30 minutes. Even higher stringency wash conditions are obtained at 65° C -68° C in a solution of 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Wash procedures will generally employ at least two final wash steps. Additional variations on these conditions will be readily apparent to those skilled in the art (see, for example, US Patent Application No. 20010010913).

[0130] Stringency conditions can be selected such that an oligonucleotide that is perfectly complementary to the coding oligonucleotide hybridizes to the coding oligonucleotide with at least about a 5-1 Ox higher signal to noise ratio than the ratio for hybridization of the perfectly complementary oligonucleotide to a nucleic acid encoding a polypeptide known as of the filing date of the application. It may be desirable to select conditions for a particular assay such that a higher signal to noise ratio, that is, about 15x or more, is obtained. Accordingly, a subject nucleic acid will hybridize to a unique coding oligonucleotide with at least a 2x or greater signal to noise ratio as compared to hybridization of the coding oligonucleotide to a nucleic acid encoding known polypeptide. The particular signal will depend on the label used in the relevant assay, e.g., a fluorescent label, a colorimetric label, a radioactive label, or the like. Labeled hybridization or PCR probes for detecting related polynucleotide sequences may be produced by oligolabeling, nick translation, end-labeling, or PCR amplification using a labeled nucleotide.

[0131] Encompassed by the invention are polynucleotide sequences capable of hybridizing to the polynucleotide sequences that encode the claimed polypeptides, including any of the polypeptides within the Sequence Listing, and fragments thereof under various conditions of stringency (see, for example, Wahl and Berger, 1987, supra, pages 399-407; and Kimmel, 1987, supra). In addition to the nucleotide sequences in the Sequence Listing, full length cDNA, orthologs, and paralogs of the present nucleotidesequences may be identified and isolated using well-known methods. The cDNA libraries, orthologs, and paralogs of the present nucleotide sequences may be screened using hybridization methods to determine their utility as hybridization target or amplification probes.Identification of Homologs by the BLAST-out and BLAST-back method, followed by generation or selection of a uORF mutation

[0132] One skilled in the art can apply the following method to identify a homolog in target (crop) species in which the practitioner wishes to generate a trait. The practitioner first selects a query polynucleotide from a query plant species that delivers a trait of interest, and BLASTs the protein sequence encoded by the query polynucleotide against the proteome, or translated DNA sequences, from the target (crop) species of interest. A particular crop protein sequence identified from the BLAST-out can be considered a homolog if (i) it matches the original query protein with a HSP of bit score 50 or better and (ii) when the protein is compared back to the set of protein sequences encoded by the genome of the plant from which the query sequence was obtained, and the protein is more similar to the protein encoded by the query polynucleotide, or any paralog of the query polynucleotide, than it is to any other protein sequence encoded by the genome of the query plant species. The practitioner may then target a genetic modification to upregulate or downregulate the identified (crop) protein to produce the desired trait, for example, through the modification of a uORF that may be found upstream of the polynucleotide encoding the identified protein. A particular embodiment of the inventions detailed herein involves selecting or generating so called “weak” non naturally occurring alleles, which result in only a moderate increase in the level of a circadian clock-associated protein in the crop of interest. In particular, mutations which completely eliminate the activity of the uORF are less desirable, as these can produce very high levels of the circadian clock-associated protein (e.g., greater than 100% higher than the levels typically found in a wild-type plant) and may lead to developmental defects and / or a reduction in yield. Alleles might be favored which maintain the function of the uORF and result in an elevation in the level of the clock-associated polypeptide to a level that is only between 5 and 99% higher than the level found in a wild-type control plant. Such desirable mutations are typically those which fall outside of the start codon as well as mutations which do not produce a frame shift in the uORF.

[0133] All publications and patent applications mentioned in this disclosure are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.

[0134] In the specification, numerous specific details are set forth in order to provide a thorough understanding of the present embodiments. It will be apparent, however, to one having ordinary skill in the art that the specific detail need not be employed to practice the present embodiments. In other instances, well-known materials or methods have not been described in detail in order to avoid obscuring the present embodiments.

[0135] Throughout this specification, quantities are defined by ranges, and by lower and upper boundaries of ranges. Each lower boundary can be combined with each upper boundary to define a range. The lower and upper boundaries should each be taken as a separate element.

[0136] Additionally, any examples or illustrations given herein are not to be regarded in any way as restrictions on, limits to, or express definitions of any term or terms with which they are utilized. Instead, these examples or illustrations are to be regarded as being described with respect to one particular embodiment and as being illustrative only. Those of ordinary skill in the art will appreciate that any term or terms with which these examples or illustrations are utilized will encompass other embodiments which may or may not be given therewith or elsewhere in the specification and all such embodiments are intended to be included within the scope of that term or terms. Language designating such non-limiting examples and illustrations includes, but is not limited to: “for example,” “for instance,” “e.g.,” “In one instantiation,” and “in one embodiment.”

[0137] In this specification, groups of various parameters containing multiple members are described. Within a group of parameters, each member may be combined with any one or more of the other members to make additional sub-groups. For example, if the members of a group are a, b, c, d, and e, additional sub-groups specifically contemplated include any one, two, three, or four of the members, e.g., a and c; a, d, and e; b, c, d, and e; etc.Embodiments

[0138] This specification includes the following non-limiting, exemplary embodiments.A plant, plant part, or plant cell, that harbors in its genome a non-naturally occurring allele of a gene that contains a mutation in a 5 ’untranslated region (UTR) that lies upstream of a main ORF that encodes a homolog of a circadian clock-associated polypeptide.

[0139] Embodiment 1 A plant, plant part, or plant cell, the genome of which harbors a non-naturally occurring allele of a gene that contains a 5 ’untranslated region (UTR) that lies upstream of a main ORF polynucleotide sequence that encodes a Homolog of a circadian clock-associated polypeptide or a polypeptide that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to any of SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332, and wherein the 5 ’UTR contains a uORF, and wherein the non-naturally occurring allele comprises a mutation in the uORF.

[0140] Embodiment 2. The plant part, plant part, or plant cell, of Embodiment 1 wherein the non- naturally occurring allele of a gene comprises a mutation in uORF that encodes a uORF -encoded peptide (uPEP) that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%,84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to a polypeptide encoded within any of SEQ ID NO: 2, 3, 5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 21, 22, 24, 26, 27, 28, 29, 30, 32, 34, 35, 36, 37, 39, 41, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 55, 57, 58,59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86,87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110,111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131,132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152,153, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174,175, 176, 177, 178, 179, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 195, 196, 197,198, 199, 200, 201, 202, 203, 204, 205, 206, 208, 209, 210, 211, 212, 214, 215, 216, 217, 218, 219, 220,221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 240, 241, 242,243, 244, 245, 246, 247, 248, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 263, 264, 265,266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 278, 279, 280, 281, 282, 283, 285, 286, 288, 290,292, 293, 295, 297, 298, 299, 300, 301, 303, 304, 330, 331, 333, 334, 335, 336, 337, 338, 339, 340, 341,342, 343, or 344.

[0141] Embodiment 3. The plant part, plant part, or plant cell, of Embodiments 1 or 2 wherein the plant or plant part exhibits a Beneficial Trait.

[0142] Embodiment 4. The plant, plant part, or plant cell of Embodiments 1-3 wherein the plant is a grain crop, fruit crop, fruit tree or vine.

[0143] Embodiment 5. The plant of Embodiment 4, wherein the plant is selected from the group comprising; maize, wheat, rice, soybean, canola, cotton, sorghum, avocado, mango, mangosteen, breadfruit jackfruit, grape, apple, cherry, plum, pear, peach, nectarine, almond, pistachio, walnut, hazelnut, tomato, kiwi, and citrus.A plant, plant part, or plant cell, that harbors within its genome a non-naturally occurring allele of a gene which when compared to a wild-type allele of the gene, expresses an elevated level of a polypeptide that is a Homolog of a circadian clock-associated polypeptide .

[0144] Embodiment 6. A plant, plant part, or plant cell, the genome of which harbors a non-naturally occurring allele of a gene which when compared to a wild-type allele of the gene, expresses an elevated level of a polypeptide that is a Homolog of a circadian clock-associated polypeptide that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to any of SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332, and the 5’UTR of the gene contains a uORF that encodes a uPEP that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%,84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to a polypeptide encoded within any of SEQ ID NO: 2, 3, 5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 21,22, 24, 26, 27, 28, 29, 30, 32, 34, 35, 36, 37, 39, 41, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 55, 57, 58,59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86,87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110,111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131,132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152,153, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174,175, 176, 177, 178, 179, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 195, 196, 197,198, 199, 200, 201, 202, 203, 204, 205, 206, 208, 209, 210, 211, 212, 214, 215, 216, 217, 218, 219, 220,221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 240, 241, 242,243, 244, 245, 246, 247, 248, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 263, 264, 265,266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 278, 279, 280, 281, 282, 283, 285, 286, 288, 290,292, 293, 295, 297, 298, 299, 300, 301, 303, 304, 330, 331, 333, 334, 335, 336, 337, 338, 339, 340, 341,342, 343, or 344, and wherein the non-naturally occurring allele comprises a mutation in the uORF, and wherein the plant exhibits a Beneficial Trait.

[0145] Embodiment 7. The plant part, plant part, or plant cell of Embodiment 6 wherein the plant part is a root stock that is grafted to a scion to produce a plant that exhibits a dwarf phenotype when grown to maturity.

[0146] Embodiment 8. The plant, plant part, or plant cell of Embodiment 6 or 7 wherein the plant is a grain crop, fruit crop, fruit tree or vine.

[0147] Embodiment 9. The plant, plant part, or plant cell of Embodiment 8 wherein the plant is selected from the group comprising; maize, rice, wheat, soybean, canola, cotton, sorghum, avocado, mango, mangosteen, breadfruitjackfruit, grape, apple, cherry, plum, pear, peach, nectarine, almond, pistachio, walnut, hazelnut, tomato, kiwi, and citrus.

[0148] Embodiment 10. The non-naturally occurring allele of Embodiments 1 or 6 wherein the allele is produced through gene editing or use of a Cas enzyme.

[0149] Embodiment 11. The plant, plant part, or plant cell of Embodiment 1 or 6 wherein the allele is selected from a mutated population using the technique of TILLING.

[0150] Embodiment 12. The plant of Embodiment 1 or Embodiment 5 wherein the uORF is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or about 100% identical to a nucleotide sequence within a sequence from the group: SEQ ID NO: 2, 3, 5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 21, 22, 24, 26, 27, 28, 29, 30, 32, 34, 35, 36, 37, 39, 41, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 55, 57, 58, 59, 60, 61, 62, 63, 64, 65,66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136,137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 155, 156, 157, 158,159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179,181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 195, 196, 197, 198, 199, 200, 201, 202,203, 204, 205, 206, 208, 209, 210, 211, 212, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225,226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 240, 241, 242, 243, 244, 245, 246, 247,248, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 263, 264, 265, 266, 267, 268, 269, 270,271, 272, 273, 274, 275, 276, 278, 279, 280, 281, 282, 283, 285, 286, 288, 290, 292, 293, 295, 297, 298,299, 300, 301, 303, 304, 330, 331, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, or 344.A method of producing a population of plants with a Beneficial Trait

[0151] Embodiment 13. A method of producing a population of plants with a Beneficial Trait comprising crossing a plant of Embodiments 1 or 5 to a second plant and growing seed from the descendants of a resulting cross.A method of producing a new plant variety comprising introducing or selecting a mutation in a cell of the plant species wherein the mutation is in a uORF sequence within the 5 ’UTR of gene that encodes a polypeptide that is a Homolog of a circadian clock-associated polypeptide.

[0152] Embodiment 14. A method of producing a new plant variety comprising introducing or selecting a mutation in a cell of the plant species wherein the mutation is in a uORF sequence within the 5 ’UTR of gene that encodes a polypeptide that is a Homolog of a circadian clock-associated polypeptide or a polypeptide with at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity to any of SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332, and wherein the cell is regenerated into a plant or plant part which is selected for an increased level of the polypeptide and the presence of a Beneficial Trait, and then multiplying the selected plant or plant part through a process of vegetative propagation, grafting or tissue culture or crossing.

[0153] Embodiment 15. The method of Embodiment 14 wherein the uORF is at 30%, 40%, 50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or about 100% identical to a nucleotide sequence within a sequence from SEQ ID NO: 2, 3, 5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 21, 22, 24, 26, 27, 28, 29,30, 32, 34, 35, 36, 37, 39, 41, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 55, 57, 58, 59, 60, 61, 62, 63, 64,65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92,93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115,116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136,137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 155, 156, 157, 158,159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179,181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 195, 196, 197, 198, 199, 200, 201, 202,203, 204, 205, 206, 208, 209, 210, 211, 212, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225,226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 240, 241, 242, 243, 244, 245, 246, 247,248, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 263, 264, 265, 266, 267, 268, 269, 270,271, 272, 273, 274, 275, 276, 278, 279, 280, 281, 282, 283, 285, 286, 288, 290, 292, 293, 295, 297, 298,299, 300, 301, 303, 304, 330, 331, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, or 344.A gene-editing vector comprising a polynucleotide sequence encoding a gene editing or DNA repair enzyme, and at least one guide RNA sequence targeting a target site in the 5 ’UTR of gene that contains a uORF that encodes a uPEP

[0154] Embodiment 16. A gene-editing vector comprising a polynucleotide sequence encoding a gene editing or DNA repair enzyme, and at least one guide RNA sequence targeting a target site in the 5 ’UTR of gene that contains a uORF that encodes a uPEP that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to a peptide encoded within any of SEQ ID NO: 2, 3, 5, 6, 7, 8,10, 11, 13, 14, 15, 16, 17, 18, 19, 21, 22, 24, 26, 27, 28, 29, 30, 32, 34, 35, 36, 37, 39, 41, 43, 44, 45, 46,47, 48, 49, 50, 51, 52, 53, 55, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76,77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103,104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124,125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145,146, 147, 148, 149, 150, 151, 152, 153, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167,168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 181, 182, 183, 184, 185, 186, 187, 188, 189,190, 191, 192, 193, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 208, 209, 210, 211, 212,214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234,235, 236, 237, 238, 240, 241, 242, 243, 244, 245, 246, 247, 248, 250, 251, 252, 253, 254, 255, 256, 257,258, 259, 260, 261, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 278, 279, 280,281, 282, 283, 285, 286, 288, 290, 292, 293, 295, 297, 298, 299, 300, 301, 303, 304, 330, 331, 333, 334,335, 336, 337, 338, 339, 340, 341, 342, 343, or 344; wherein the construct alters the expression level or activity of a Homolog of a circadian clock-associated polypeptide or a polypeptide with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to any of SEQ ID NO: 1, 4,9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289,291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332.

[0155] Embodiment 17. The gene-editing vector of Embodiment 16, wherein the gene-editing construct is a CRISPR / Cas9 construct.

[0156] Embodiment 18. The gene-editing vector of Embodiment 16, wherein a plant expressing exhibits a Beneficial Trait or at least one attribute selected from the group consisting of a) altered auxin transport, b) slower auxin transport, c) reduced apical dominance, d) an altered xylem / phloem ratio, e) an increased number of phloem elements, f) smaller phloem elements, g) thicker bark, h) a bushier habit, i) reduced root mass, and j) reduced vigor, k) less vegetative growth, 1) earlier termination of shoot growth, m) earlier competence to flower, n) precocity, o) earlier phase change, p) smaller canopy, q) reduced stem circumference, r) reduced branch diameter, s) fewer sylleptic branches, t) shorter sylleptic branches, u) more axillary flowers, v) an earlier terminating primary axis, w) earlier terminating secondary axes, x) shorter internode length, y) reduced scion mass and z) early maturation.

[0157] Embodiment 19. The vector of Embodiment 15, wherein the guide RNA sequence is complementary to a uORF contained within a sequence comprising SEQ ID NO: 2, 3, 5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 21, 22, 24, 26, 27, 28, 29, 30, 32, 34, 35, 36, 37, 39, 41, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 55, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 208, 209, 210, 211, 212, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 240, 241, 242, 243, 244, 245, 246, 247, 248, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 278, 279, 280, 281, 282, 283, 285, 286, 288, 290, 292, 293, 295, 297, 298, 299, 300, 301, 303, 304, 330, 331, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, or 344.

[0158] Embodiment 20. A host cell comprising the vector of Embodiment 19.

[0159] Embodiment 21. The host cell of Embodiment 20, wherein the host cell is a bacterial cell.

[0160] Embodiment 22. The host cell of Embodiment 21 wherein the host is an Agrobacterium.

[0161] Embodiment 23. A plant tissue transformed with the host cell of Embodiment 21.

[0162] Embodiment 24. The plant tissue of Embodiment 22, wherein the tissue is a root tissue.A method of producing a plant with reduced expression level of a Homolog of a circadian clock- associated polypeptide.

[0163] Embodiment 25. A method of producing a plant with reduced expression level of a Homolog of a circadian clock-associated polypeptide that is at least 50%, 70%, 71%, 72%, 73%, 74%, 75%, 76%,77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%,97%, 98%, 99% or about 100% identical to SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309,310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332, the method comprising expressing the construct of Embodiment 15 in a plant cell and selecting a regenerated plant that shows a Beneficial Trait compared to a control plant.

[0164] Embodiment 26. The method of Embodiment 25, wherein the expression level of a Homolog of a circadian clock-associated polypeptide is reduced to a level of about 0%, or to about or at most 1%,2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%.21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%,49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%.67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%.85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, or 99%, or to 0% to 10%, or to or to 10% to 20%, or to2 0% to 30%, or to 30% to 40%, or to 40% to 50%, or to 50% to 60%, or to 60% to70%, or to 70% to 80%, or to 80% to 90%, or to 90% to 99%, of the expression level of the Homolog of a circadian clock-associated polypeptide in a control plant.A method of increasing the production of an active protein from a polynucleotide in a plant cell.

[0165] Embodiment 27. A method of increasing the production of an active protein from a polynucleotide in a plant cell, the method comprising introducing or selecting a sequence change within a uORF that resides upstream of the start codon of the main ORF that encodes the active protein and wherein the uORF sequence was first identified in the 5’UTR of a gene that encodes a protein that is a Homolog of a circadian clock-associated protein or a protein that has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% percent identity to a SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332.

[0166] Embodiment 28. A plant regenerated from the plant cell with increased production of an active protein produced by the method of Embodiment 27.

[0167] Embodiment 29. The method of Embodiment 27, wherein the active protein is luciferase, renilla, a luminescent protein, green fluorescent protein, or a homolog thereof, or a fluorescent protein.A plant, the genome of which harbors a non-naturally occurring allele of a gene which expresses an increased level of a polypeptide that is a Homolog of a circadian clock-associated polypeptide

[0168] Embodiment 30. A plant, the genome of harbors a non-naturally occurring allele of a gene which when compared to a wild-type allele of the gene, expresses an increased level of a polypeptide that is a Homolog of a circadian clock-associated polypeptide that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to any of SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, or 328; and wherein the 5’UTR of the gene contains a uORF, and wherein the non-naturally occurring allele comprises a mutation in the uORF.

[0169] Embodiment 31. The plant of Embodiment 30, wherein the plant has been selected for the presence of a Beneficial Trait.A plant, the genome of which harbors a non-naturally occurring allele of a gene which expresses a reduced level of a polypeptide that is a Homolog of a circadian clock-associated polypeptide

[0170] Embodiment 32. A plant, the genome of which harbors a non-naturally occurring allele of a gene which when compared to a wild-type allele of the gene, expresses a reduced level of a polypeptide that is a Homolog of a circadian clock-associated polypeptide that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to any of SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, or 332, and wherein the 5’UTR of the gene contains a uORF, and wherein the non- naturally occurring allele comprises a mutation in the uORF.

[0171] Embodiment 33. The plant of Embodiment 32, wherein the plant has been selected for the presence of a Beneficial Trait.

[0172] Embodiment 34. A plant, the genome of which harbors a non-naturally occurring allele of a gene which when compared to a wild-type allele of the gene, expresses a reduced level of a polypeptide that is a Homolog of the HY5 or HYH polypeptide or that is at least 50%, 60%, 65%, 70%, 71%, 72%,73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%,91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to any of SEQ ID NO: 54 and 311, andwherein the 5’UTR of the gene contains a uORF, and wherein the non-naturally occurring allele comprises a mutation in the uORF.

[0173] Embodiment 35. The plant of Embodiment 34, wherein the plant is a soybean plant.

[0174] Embodiment 36. The plant of Embodiments 34 or 35, wherein the soybean plant has been selected for the presence of a Beneficial Trait as compared to a control plant.

[0175] Embodiment 37. A method of increasing the production of a protein from a polynucleotide in a plant cell, the method comprising introducing or selecting a sequence change within a region containing an upstream open reading frame (uORF) that resides upstream of the start codon of the main ORF that encodes the active protein and wherein the uORF sequence was first identified in the 5’UTR of a gene that encodes a Homolog of a circadian clock-associated protein. The protein can be any protein the expression of which can be detected and / or quantified. In some embodiments, the protein is a reporter polypeptide, the expression of which produces a measurable signal — such as fluorescence, luminescence, or an enzymatic activity — that can be readily measured using standard laboratory equipment.

[0176] Embodiment 38. The method of Embodiment 37, wherein the Homolog has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% percent identity to a polypeptide selected from the group consisting of SEQ ID NO: 4, 1, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, and 332.

[0177] Embodiment 39. The method of Embodiment 37, wherein the protein is luciferase, renilla, a luminescent protein, green fluorescent protein, or a homolog thereof, or a fluorescent protein.

[0178] Embodiment 40. The method of Embodiment 37, wherein a plant is regenerated from the plant cell; and as compared to a control plant, the plant has increased production of the active protein.

[0179] Embodiment 41. A plant, or plant part, the genome of which harbors a non-naturally occurring allele of a gene which when compared to a wild-type allele of the gene, expresses an increased level of a polypeptide that is a Homolog of a circadian clock-associated polypeptide that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to a protein selected from the group consisting of SEQ ID NO: 4, 1, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, and 332; and wherein the 5’UTR of the gene contains a uORF; and wherein the non-naturally occurring allele comprises a mutation in the uORF; and wherein the plant or plant part exhibits a Beneficial Trait.

[0180] Embodiment 42. A guide RNA with complementarity to a uORF within the 5’UTR sequence of the gene encoding a homolog of a circadian clock-associated protein; the Homolog is introduced into a cell of the target plant species that expresses a CAS enzyme, or making use of a genome mechanism, editing tool or DNA modification enzyme; the cell is regenerated into a plant; wherein the plant is selected for a reduction in levels of the target circadian clock-associated protein; wherein the plant comprises a mutation within the uORF sequence; the selected plant is grown and propagated either sexually or asexually; and a progeny plant is selected that exhibits a Beneficial Trait.

[0181] Embodiment 43. The guide RNA of Embodiment 42, wherein the Homolog has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% percent identity to a polypeptide selected from the group consisting of SEQ ID NO: 4, 1, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, and 332.

[0182] Embodiment 44. The guide RNA of Embodiment 42, wherein the selected plant is a rootstock and is grafted to a scion from a second plant.

[0183] Embodiment 45. The guide RNA of Embodiment 42, wherein the CAS enzyme, genome mechanism, editing tool or DNA modification enzyme is selected from the group consisting of homologous recombination, zinc finger nucleases, transcription activator-like effector nucleases, pentatricopeptide repeat proteins, the CRISPR / Cas9 system, the CRISPR / Casl2a (Cpfl) system, adenine base editing system, the Retron Library Recombineering Platform , Cas-CLOVER nucleases, Mini-Cas9 enzymes, RNA interference (RNAi), cisgenesis, and intragenesis.

[0184] Embodiment 46. The plant of Embodiment 41 where the plant is from the mustard family.

[0185] Embodiment 47. The plant of Embodiment 41 where the plant is Brassica napus.

[0186] Embodiment 48. A method of decreasing the production of a protein from a polynucleotide in a plant cell, the method comprising introducing or selecting a sequence change within a region of the polynucleotide containing an upstream open reading frame (uORF) that resides upstream of the start codon of the main ORF that encodes the protein and wherein the uORF containing region was first identified in the 5’UTR of a gene that encodes a Homolog of a circadian clock-associated protein.

[0187] Embodiment 49. The method of Embodiment 48, wherein the Homolog has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% percent identity to a polypeptide selected from the group consisting of SEQ ID NO: 12, 42, 54, and 311.

[0188] Embodiment 50. The method of Embodiment 48, wherein the protein is luciferase, renilla, a luminescent protein, green fluorescent protein, or a homolog thereof, or a fluorescent protein.

[0189] Embodiment 51. The method of Embodiment 48, wherein a plant is regenerated from the plant cell; and as compared to a control plant, the plant has increased production of the protein.EXAMPLES

[0190] The invention, now being generally described, will be more readily understood by reference to the following examples, which are included merely for purposes of illustration of certain aspects and embodiments of the present invention and are not intended to limit the invention. It will be recognized by one of skill in the art that a genetic modification that is associated with a particular first trait may also be associated with at least one other, unrelated, and inherent second trait which was not predicted by the first trait.EXAMPLE 1. Identifying a selection of Homologs of circadian clock-associated proteins

[0191] Sequence similarity between major open reading frames of different species can be used to identify heterologous genes from which associated uORFs can then be identified. The sequences of a selection of circadian clock-associated proteins from Arabidopsis were used to identify the corresponding orthologs in various selected plant species. The query sequence in each case was used in a sequence homology alignment search of the genomes of plant species using BLAST at genomevolution.org / coge / CoGeBlast.pl. In each case the best hit proteins were blasted back against the Arabidopsis proteome. If the top hit from the Arabidopsis proteome was the original Arabidopsis query protein (or a paralog) with a match with a bit score of greater than 50, the identified crop protein was considered to be a homolog of a circadian clock-associated protein. Results are shown in Table 1.Selected example circadian clock-associated protein sequences were used for the queries including those provided as SEQ ID NO: 1, 4, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 305, 306, 307, 308, 309, 310, 311, 312, and 313.EXAMPLE 2.

[0192] Putative uORF -containing regions within the 5’ UTRs of circadian clock-associated protein encoding sequences in various plant species were identified through application of an algorithm to ribosome profiling data. The algorithm identifies the potential presence of a uORF based on the existence of ribosome enrichment in the interval from one stop codon to the next stop codon within the same open reading frame,

[0193] The later stop codon represents the end of a putative uORF. The sequence immediately upstream of the later stop codon represents a potential target for gene editing that disrupts the function of the putative uORF. Results are shown in Table 1, above.EXAMPLE 3: Selection of a tomato plant with a Beneficial Trait carrying a non-naturally occurring allele comprising a uORF mutation in a circadian clock associated protein encoding gene

[0194] A plant is selected from a tomato mutant population wherein the selected plant contains a mutation in a uORF within the 5 ’UTR of a sequence encoding a tomato circadian clock-associated protein homolog (SEQ ID NO: 56, 154, 180, 194, 207, 213, 239, 249, 317, 318, 319, 320, or 321), which is identified by sequencing the genomic region containing the sequence of the homolog from a number of plants from the mutant population and comparing the sequences to the genomic region from a wild-type tomato plant of the same variety. The selected plant is then grown and crossed, either to itself or to a second plant, and the progeny seed are then germinated and a plant is selected that exhibits a Beneficial Trait. Optionally, the selected plant may be grafted as a rootstock to a scion from a second plant.EXAMPLE 4: Selection of a rice plant with a Beneficial Trait carrying a non-naturally occurring allele comprising a uORF mutation in a circadian clock associated protein encoding gene

[0195] A plant is selected from a rice mutant population wherein the selected plant contains a mutation in a uORF within the 5 ’UTR of a sequence encoding a rice circadian clock-associated protein homolog (SEQ ID NO: 262, 277, 284, 287, 289, 291, 294, or 325), which is identified by sequencing the genomic region containing the sequence of the homolog from a number of plants from the mutant population and comparing the sequences to the genomic region from a wild-type rice plant of the same variety. The selected plant is then grown and crossed, either to itself or to a second plant, and the progeny seed are then germinated and a plant is selected which exhibits a Beneficial Trait.EXAMPLE 5: Generation of a tomato plant that exhibits a Beneficial Trait and harbors a non- naturally occurring allele comprising a uORF mutation in a gene encoding a circadian clock- associated protein

[0196] A guide RNA with complementarity to a uORF within the 5’UTR sequences (SEQ ID NO: 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131,132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152,153, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175,176, 177, 178, 179, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 196, 197, 198, 199, 200,201, 202, 203, 204, 205, 206, , 209, 210, 211, 212, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 241, 242, 243, 244, 245, 246, 247, 248,251, 252, 253, 254, 255, 256, 257, 258, 259, 260, or 261) ofa gene encoding atomato circadian clock- associated protein homolog (SEQ ID NO: 56, 154, 180, 194, 207, 213, 239, 249, 317, 318, 319, 320, or 321) is introduced into tomato cells that express a CAS enzyme (or other suitable genome mechanisms, editing tools, or DNA modification enzymes) and regenerated plants are obtained. A plant is then selected from the regenerated plants which contains a mutation within the uORF sequence as compared to a wild type plant. The selected plant is then grown and propagated, either sexually or asexually and a progeny plant is selected which exhibits a Beneficial Trait. Optionally, the selected plant may be grafted as a rootstock to a scion from a second plant.EXAMPLE 6: Generation of a rice plant that exhibits a Beneficial Trait and harbors a non- naturally occurring allele comprising a uORF mutation in a gene encoding a circadian clock- associated protein

[0197] A guide RNA with complementarity to a uORF within the 5’UTR sequences (SEQ ID NO: 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 279, 280, 281, 282, 283, 286, or 293) of a gene encoding a rice circadian clock-associated protein homolog (SEQ ID NO: 262, 277, 284, 287, 289, 291, 294, or 325) is introduced into cells that express a CAS enzyme (or other suitable genome mechanisms, editing tools, or DNA modification enzymes) and regenerated plants are obtained. A plant is then selected from the regenerated plants which contains a mutation within the uORF sequence as compared to a wild type plant. The selected plant is then grown and propagated, either sexually or asexually and a progeny plant is selected which exhibits a Beneficial Trait.EXAMPLE 7: Generation of a soybean plant that exhibits a Beneficial Trait and harbors a non- naturally occurring allele comprising a uORF mutation in a gene encoding a circadian clock- associated protein

[0198] A guide RNA with complementarity to a uORF within the 5’UTR sequences (SEQ ID NO: 298, 299, 300, 301, or 304) of a gene encoding a soybean circadian clock-associated protein homolog (SEQ ID NO: 296, 302, or 322) is introduced into cells that express a CAS enzyme (or other suitable genome mechanisms, editing tools, or DNA modification enzymes) and regenerated plants are obtained. A plant is then selected from the regenerated plants which contains a mutation within the uORF sequence as compared to a wild type plant. The selected plant is then grown and propagated, either sexually or asexually and a progeny plant is selected which exhibits a Beneficial Trait.EXAMPLE 8: Generation of a plant of a grain crop with a Beneficial Trait harboring a non- naturally occurring allele comprising a uORF mutation in a gene encoding a circadian clock- associated protein

[0199] A guide RNA with complementarity to a uORF within the 5’UTR sequence of the gene encoding a Homolog of a circadian clock associate protein from a grain bearing crop such as rice, wheat,sorghum, or com is introduced into cells of that crop that express a CAS enzyme (or other suitable genome mechanisms, editing tools, or DNA modification enzymes) and regenerated plants are obtained. A plant is then selected from the regenerated plants, which contains a mutation within the uORF sequence as compared to a wildtype plant. The selected plant is then grown and propagated, either sexually or asexually and a progeny plant is selected which exhibits a Beneficial Trait.EXAMPLE 9: Generation of a Miscanthus plant with a Beneficial Trait harboring a non-naturally occurring allele comprising a uORF mutation in a gene encoding a circadian clock-associated protein

[0200] A guide RNA with complementarity to a uORF within the 5’UTR sequence of the gene encoding a Miscanthus circadian clock-associated protein homolog is introduced into Miscanthus cells that express a CAS enzyme (or other suitable genome mechanisms, editing tools, or DNA modification enzymes) and regenerated plants are obtained. A plant is then selected from the regenerated plants which contains a mutation within the uORF sequence as compared to a wild type plant, wherein the mutation results in a reduction in levels of the circadian clock-associated protein. The selected plant is then grown and propagated, either sexually or asexually and a progeny plant is selected which exhibits a desirable Beneficial Trait.EXAMPLE 10: Generation of a tree with a Beneficial Trait that harbors a non-naturally occurring allele comprising a uORF mutation in gene encoding a circadian clock-associated protein

[0201] A guide RNA with complementarity to a uORF within the 5’UTR sequence of the gene encoding a tree homolog of a circadian clock-associated protein (e.g. SEQ ID NO: 314, 315, 316) is introduced into cells of the target tree species that express a CAS enzyme (or other suitable genome mechanisms, editing tools, or DNA modification enzymes) and regenerated plants are obtained. A plant is then selected from the regenerated plants which contains a mutation within the uORF sequence as compared to a wild type plant, wherein the mutation results in a reduction in levels of the target circadian clock-associated protein. The selected plant is then grown and propagated, either sexually or asexually and a progeny plant is selected which exhibits a Beneficial Trait.EXAMPLE 11: Generation of an improved crop variety by modulating HY5 Homolog protein levels through altered uORF control

[0202] The HY 5 transcription factor is a positive regulator of light regulated development. Based on work performed previously by Khanna et al. and Preuss et al., focused on overexpression of BBOX transcription factors, suppression of light regulated development can produce Beneficial Traits in plants. However, traits based on BBOX transcription factor interventions have not been commercialized, possibly because it was difficult to fine tune levels of activity of such proteins using transgenicapproaches. HY5 related genes possess candidate uORFs in their 5’UTR regions; the invention herein provides an alternative, more precise means of controlling light signaling, by targeting mutations to these uORFs in HY5 gene leader sequences. In particular, the approach can be applied to soybean to produce soybean plants with Beneficial Traits.

[0203] In one instantiation of this example, a guide RNA with complementarity to a uORF within the 5’UTR sequence of the gene encoding a Homolog of a HY5 protein from a seed-bearing crop such as rice, wheat, sorghum, millet, canola, or com is introduced into cells of that crop that express a CAS enzyme and regenerated plants are obtained. A plant is then selected from the regenerated plants, which contains a mutation within the uORF sequence as compared to a wildtype plant. The selected plant is then grown and propagated, either sexually or asexually and, optionally, a progeny plant is selected which exhibits a Beneficial Trait.

[0204] In a further instantiation of this example, a guide RNA with complementarity to a uORF within the 5’UTR sequence of the gene encoding a Homolog of a HY5 protein from soybean is introduced into cells of that crop that express a CAS enzyme (or other suitable genome mechanisms, editing tools, or DNA modification enzymes) and regenerated plants are obtained. A plant is then selected from the regenerated plants, which contains a mutation within the uORF sequence as compared to a wildtype plant. The selected plant is then grown and propagated, either sexually or asexually and, optionally, a progeny plant is selected which exhibits a Beneficial Trait. In a first instance of the example, the HY 5 homolog is a product of the locus identified GUYMA_18G117100 and the sequence change results in the conversion of a non-canonical start codon to start codon, which would increase the strength of uORF based repression of the HY5 protein translation. In a second instance, the HY5 homolog is a product of the locus identified by GUYMA_08G302500 and the sequence change results in the conversion of a non- canonical start codon to start codon, which would increase the strength of uORF based repression of the HY5 protein translation. By way of further example, potential “stop-stop” uORF containing regions from the UTR of these genes are shown in Figures Irrespectively, with candidate non-canonical start codons (ncStart) that could be converted to an ATG start codon indicated.EXAMPLE 12: Generation of a plant with a Beneficial Trait harboring a non-naturally occurring allele comprising a uORF mutation in a gene encoding a GIGANTEA Homolog

[0205] A guide RNA with complementarity to a uORF within the 5’UTR sequence of the gene encoding a Homolog of GIGANTEA is introduced into cells of a crop that express a CAS enzyme (or other suitable genome mechanisms, editing tools, or DNA modification enzymes) and regenerated plants are obtained. A plant is then selected from the regenerated plants, which contains a mutation within the uORF sequence as compared to a wildtype plant. The selected plant is then grown and propagated, either sexually or asexually and a progeny plant is selected which exhibits a Beneficial Trait.

[0206] In a particular instantiation of the example, the crop is Eucalyptus and the GI Homolog is at least 50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%.85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical toSEQ ID NO: 314, 315 or 316.

[0207] In a further instantiation of the example, the crop is tomato and the GI Homolog is at least50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%.86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to any of SEQ ID NO: 317 -321.

[0208] In a further instantiation of the example, the crop is soybean and the GI Homolog is at least 50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%. 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to any of GI soybean sequences represented by GenBank accession BAJ22595.1, KAH1139549.1,NP_001239995.2, KAH1036527.1, NP_001341719.1, ACJ65312.1, KAH1191303.1, KAG4919220.1, KAG4908008.1, BAN82581.1, BAN82586.1, KAH1230527.1, BAN82589.1, KAH1191305.1, KAH1191304.1, KAG5152664.1, XP_014624480.2, KAG5004875.1, ACJ65314.1, KAG4380382.1, KAG4395160.1, KAG5128058.1, KAH1206751.1, KAG4998117.1, KAH1139552.1, KAG4397849.1,ACA24490.1, XP_040861536.1, KAH1230528.1, KAH1191306.1, KAH1191302.1, KAG5062466.1, XP 040863520.1, AMZ01609.1, AMZ01659.1, KAG4951485.1, orto SEQ ID NO: 322.

[0209] In a further instantiation of the example, the crop is maize and the GI Homolog is at least 50%,60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%.87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to SEQ ID NO: 323.

[0210] In a further instantiation of the example, the crop is potato and the GI Homolog is at least 50%,60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%.87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to SEQ ID NO:324 or to any of the GI potato sequences represented by GenBank accession XP_006359040.1, KAH0649545.1, KAH0710770.1, XP_006359039.1, XP_006361616.1, KAH0707598.1, KAH0736444.1, or XP_006361617.1.

[0211] In a further instantiation of the example, the crop is rice and the GI Homolog is at least 50%,60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%.87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to SEQ ID NO:325 or to any of the sequences represented by GenBank accession NP_001396355.1, EEC70061.1, EEE54000.1, CAB56058.1, KAF2948772.1, AAX83420.1, AAX83424.1, AAX83418.1, orBAD68053.1.

[0212] In a further instantiation of the example, the crop is wheat and the GI Homolog is at least 50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to SEQ ID NO: 326 or to any of the sequences represented by GenBank accession XP_044347620. 1, XP_044347619. 1, XP 044339096.1, AAT79487.1, AAQ11738.1, XP_044355721.1, AAT79486.1, KAF7028336.1, KAF7035449.1, KAF7028337.1, or KAF7021482.1

[0213] In a further instantiation of the example, the crop is a tree of the genus Prunus and the GI Homolog is at least 50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to SEQ ID NO:327 or to any of the sequences represented by GenBank accessionXP 021818210.1, AJC01622.1, XP_007199688.2, XP_034196939.1, KAH0976342.1, XP_008237481.1, XP 021818214.1, XP 034196940.1, XP_020426395.1, XP_016650762.1, PQQ13557.1,XP 034196942.1, PQM43026.1, CAB4289400.1, CAB4319797.1, XP_034196943.1, XP_021818215.1, XP 034196945.1, or XP_034196944.1.

[0214] In a further instantiation of the example, the crop is a tree of the genus Citrus and the GIHomolog is at least 50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to SEQ ID NO: 328 or to any of the sequences represented by GenBank accession XP 052298603.1, KAH9757707.1, XP_024039988.1, KAH9691868.1, KAK9219975.1, KAH9691869.1, KAK9216391.1, KDO83806.1, ESR47759.1, KAH9757708.1, KAH9691870.1,KAH9757709.1, or KAH9691871.1.

[0215] In a further instantiation of the example, the crop is Brassica napus and the GI Homolog is at least 50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to SEQ ID NO: 329 or 332. By way of further example, the Beneficial Trait obtained in Brassica napus is early maturation and / or heat tolerance.EXAMPLE 13: Demonstration of the capacity of sequences within the 5’UTR of genes encoding example circadian clock components to regulate the level of a protein encoded by an operably linked main ORF

[0216] The identified uORF containing region from the 5 ’ UTR sequence of the gene encoding PHY C(SEQ ID NO. 38) was introduced into a pair of reporter constructs comprising the dual luciferase vectors shown in Figure 6. DNA from these vectors was then introduced into protoplasts from maize leaf cells and using a dual luciferase detection kit, the enzymatic activity of the two chemiluminescent reporter genes was determined. Translational efficiency was calculated as a ratio of the level of firefly luciferase compared to the level of renilla. In this example, translational efficiency is much higher in the absence ofthe uORF containing region, demonstrating that mutations, or sequence changes in, including deletions, in that region within a leader sequence associated with a downstream main ORF can increase translation of the main ORF. The results of a first study with PHYC are shown in Figure 7. Further studies were performed with identified uORF containing regions from the 5 ’ UTR sequences of the genes encoding CCA1, ELF3, GI, PHYA, PHYC and TOC1. In each case the identified uORF containing region was introduced into a pair of reporter constructs comprising the dual luciferase vectors shown in Figure 6, and the expression of the reporter in the presence or absence of the respective uORFs were compared. The results of these further studies are shown in Figure 8.EXAMPLE 14: Mutation of the 5’UTR uORF containing region of CCA1 delivers a Beneficial Trait

[0217] A gene editing construct targeting the 5’UTR of the gene encoding the CCA1 protein (SEQ ID NO. 306) was introduced into Arabidopsis. The resulting plants were grown under heat stress conditions in a glasshouse compared to wild-type control plants under the same conditions. The results are shown in Figure 9.

Claims

WHAT IS CLAIMED IS:

1. A method of increasing the production of a protein from a polynucleotide in a plant cell, the method comprising introducing or selecting a sequence change within a region of the polypeptide containing an upstream open reading frame (uORF) that resides upstream of the start codon of the main ORF that encodes the active protein and wherein the uORF containing region was first identified in the 5’UTR of a gene that encodes a Homolog of a circadian clock-associated protein.

2. The method of claim 1, wherein the Homolog has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% percent identity to a polypeptide selected from the group consisting of SEQ ID NO: 4, 1, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, and 332.

3. The method of claim 1, wherein the protein is luciferase, renilla, a luminescent protein, green fluorescent protein, or a homolog thereof, or a fluorescent protein.

4. The method of claim 1, wherein a plant is regenerated from the plant cell; and as compared to a control plant, the plant has increased production of the protein.

5. A plant, or plant part, the genome of which harbors a non-naturally occurring allele of a gene which when compared to a wild-type allele of the gene, expresses an increased level of a polypeptide that is a Homolog of a circadian clock-associated polypeptide that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% identical to a protein selected from the group consisting of SEQ ID NO: 4, 1, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, and 332; and wherein the 5’UTR of the gene contains a uORF; and wherein the non-naturally occurring allele comprises a mutation in the uORF; and wherein the plant or plant part exhibits a Beneficial Trait.

6. A guide RNA with complementarity to a uORF containing sequence within the 5’UTR of the gene encoding a Homolog of a circadian clock-associated protein; the Homolog is introduced into a cell of the target plant species that expresses a CAS enzyme, or making use of a genome mechanism, editing tool or DNA modification enzyme; the cell is regenerated into a plant; wherein the plant is selected for a reduction in levels of the target circadian clock-associated protein;wherein the plant comprises a mutation within the uORF sequence; the selected plant is grown and propagated either sexually or asexually; and a progeny plant is selected that exhibits a Beneficial Trait.

7. The guide RNA of claim 6, wherein the Homolog has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% percent identity to a polypeptide selected from the group consisting of SEQ ID NO: 4, 1, 9, 12, 20, 23, 25, 31, 33, 38, 40, 42, 54, 56, 154, 180, 194, 207, 213, 239, 249, 262, 277, 284, 287, 289, 291, 294, 296, 302, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, and 332.

8. The guide RNA of claim 6, wherein the selected plant is a rootstock and is grafted to a scion from a second plant.

9. The guide RNA of claim 6, wherein the CAS enzyme, genome mechanism, editing tool or DNA modification enzyme is selected from the group consisting of homologous recombination, zinc finger nucleases, transcription activator-like effector nucleases, pentatricopeptide repeat proteins, the CRISPR / Cas9 system, the CRISPR / Casl2a (Cpfl) system, adenine base editing system, the Retron Library Recombineering Platform , Cas-CLOVER nucleases, Mini-Cas9 enzymes, RNA interference (RNAi), cisgenesis, and intragenesis.

10. The plant of claim 5 where the plant is from the mustard family.

11. The plant of claim 5 where the plant is Brassica napus.

12. A method of decreasing the production of a protein from a polynucleotide in a plant cell, the method comprising introducing or selecting a sequence change within a region of the polynucleotide containing an upstream open reading frame (uORF) that resides upstream of the start codon of the main ORF that encodes the protein and wherein the uORF containing region was first identified in the 5’UTR of a gene that encodes a Homolog of a circadian clock-associated protein.

13. The method of claim 12, wherein the Homolog has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 96%, 97%, 98%, 99% or about 100% percent identity to a polypeptide selected from the group consisting of SEQ ID NO: 12, 42, 54, and 311.

14. The method of claim 12, wherein the protein is luciferase, renilla, a luminescent protein, green fluorescent protein, or a homolog thereof, or a fluorescent protein.

15. The method of claim 12, wherein a plant is regenerated from the plant cell; and as compared to a control plant, the plant has increased production of the protein.

Citation Information

Patent Citations

  • Plant control genes

    US6887708B1

  • Uorf::reporter gene fusions to select sequence changes to gene edit into uorfs to regulate ascorbate genes

    WO2024077110A2