Novel method for identifying and using upstream open reading frame
By analyzing the ribosome profile data to identify uORF and modifying the uORF sequence through gene editing technology, the problem of difficulty in identifying and utilizing uORF in the prior art is solved, and the accurate identification and modification of uORF is achieved, and the desired phenotype is generated.
Patent Information
- Application Number
- CN202380056417.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-05-27
- Filing Date
- 2023-05-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to accurately identify the upstream open reading frame (uORF) in the eukaryotic genome because these elements are short in length and are usually initiated by non-AUG initiation codons.
By analyzing the ribosome profile data, we identify locations in the genome where ribosomes are enriched and the downstream ribosome occupancy has plummeted to identify the presence of uORF, and modify the uORF sequence through gene editing techniques to produce the desired phenotype.
Accurate identification and modification of uORF is achieved, which can induce new desired phenotypes and be applied to improve crop traits, control weeds and cancer cells, etc.
Smart Images

Figure CN120077140A_ABST
Abstract
Description
Technical Field
[0001] This article deals with the identification and utilization of upstream open reading frames in eukaryotic genomes. Background Art
[0002] Upstream open reading frames (uORFs) are members of a class of short, conserved ORFs located upstream of the protein-encoding primary ORF (mORF) in the 5'-untranslated region (5'UTR) of an mRNA. uORFs act as cis-acting elements, regulating the activity of downstream sequences encoding polypeptides. Therefore, gene editing methods that introduce mutations into uORF sequences offer new opportunities for activating the expression of downstream open reading frames encoding polypeptides of interest. Thus, upregulation of target polypeptide levels can be achieved by regulating the expression of negatively acting uORFs located upstream of the polypeptide-encoding sequence (e.g., by knockout or mutation). Not all eukaryotic genes contain uORFs, and to date, the aforementioned methods have been hindered by the identification of uORF sequences; due to their short length and the fact that they typically initiate via non-AUG start codons, existing algorithms are often unable to accurately identify these elements (Hellens et al., 2016. Trend Plant Sci. 21:317-328). Here, we propose a novel algorithm that identifies the presence of uORFs by analyzing so-called ribosome profiling data. The target uORF sequence is then genetically modified to produce a new desired phenotype, which has a variety of applications depending on the specific cell or organism. Such applications include new crop traits (e.g., increased yield, vigor, stress tolerance, delayed or accelerated flowering, morphological changes, or improved nutrient content), control of weeds and other pests by activating gene networks that initiate cell death, activation of cell death or tumor suppressor genes in cancer cells, and / or production of desired metabolites or peptides in fermentation systems.
[0003] uORFs are ubiquitous regulatory elements in eukaryotic mRNAs. uORFs are located upstream of the protein-coding primary ORF (also called long ORF, main ORF, or mORF) in the 5'-untranslated region (5'UTR) of an mRNA. In some cases, uORFs are thought to regulate the translation initiation rate of downstream coding sequences (CDS) by sequestering ribosomes. In other cases, uORFs encode evolutionarily conserved short peptides (sometimes called "uPEPs") that act as cis-repressors for the downstream mORF or its protein product. In many cases, the actual presence of uORFs is highly conserved across species. Therefore, once a uORF is identified in a target locus in a given species, homologous loci from another species will often also contain the uORF and be repressed by it. Here, we applied a novel algorithm to identify a set of loci containing uORFs from the model plant Arabidopsis thaliana. These data now provide a roadmap for identifying loci containing uORFs in target crops based on homology searches of the polypeptides encoded by the mORFs located at these loci. Therefore, when a given desired trait is identified by overexpressing a gene in Arabidopsis, if the locus contains a uORF, the desired trait can be obtained by mutating the uORF in the crop gene to activate the equivalent homologous gene in the target crop, which is usually located in a similar position upstream of the mORF in the homologous locus of the target crop. A specific example is editing the uORF of LsGGP2, which encodes a key enzyme in vitamin C biosynthesis in lettuce. The targeting of this gene is based on Laing et al.'s demonstration that its homologous gene is controlled by a uORF in Arabidopsis (Liang et al., Plant Cell. 2015 Mar; 27(3): 772–786). Editing the uORF of the lettuce homolog not only improved antioxidant stress tolerance, but also increased ascorbic acid content by approximately 150% (Zhang et al. Nature Biotechnology Vol. 36, pp. 894–898 (2018)).
[0004] Whole-genome studies have revealed that uORFs have a wide range of regulatory functions in different species under different biological backgrounds (Zhang et al. 2019. Trends Biochem. Sci. 44: 782-794. doi: 10.1016 / j.tibs.2019.03.002). A given uORF can act as a translational control element to regulate the expression of its associated downstream major open reading frame (mORF). Plant studies have confirmed that highly conserved uORFs regulate the translation of mORFs in response to cellular metabolite levels (Hayden CA and Jorgensen RA 2007. BMC Biol. 5: 32; Tran MK et al. 2008. BMC Genomics 9: 361).
[0005] Various methods for identifying uORFs in eukaryotes have been described. For example, to identify conserved peptide uORFs, Hayden and Jorgensen created the Perl program "uORF-Finder," which compares the amino acid sequences of mORFs from a collection of cDNAs with those from another species to identify putative mORF homologs. The uORFs in the 5'UTRs of the two paired sequences are then compared to identify uORFs with conserved amino acid sequences (Hayden and Jorgensen, 2007. BMC Biology 5:32). By comparing full-length cDNA sequences from Arabidopsis thaliana and rice, different homologous groups of conserved peptide uORFs can be identified. Skarshewski et al. described the use of “uPEPperoni,” an online tool for upstream open reading frame mapping and transcript conservation analysis (Skarshewski, A., et al. 2014. BMC Bioinform. 15:36. doi: 10.1186 / 1471-2105-15-36).
[0006] Rather than using bioinformatics-based analysis, Ingolia et al. describe a method for ribosome profiling: uORFs are identified by assessing ribosome occupancy of upstream open reading frames and other sequences. See, e.g., US Patent 9,677,068; Ingolia NT, 2014. Cell Reports 8:5, 1365–1379. See also Ingolia NT, 2011. Cell 11; 147:789–802, where the authors describe how most putative lincRNAs contain highly translated regions comparable to protein-coding genes. Specific start sites are marked with harringtonine, followed by ribosome footprints extending to the first in-frame stop codon. Most of the novel near-cognate start sites detected drive translation of uORFs. This is consistent with the high levels of translation observed in many 5' UTRs, in contrast to the 3' UTRs, which are nearly ribosome-free.
[0007] In contrast to previously described methods, the novel approach of the present invention identifies uORFs based on the ability to clearly demarcate stop codons based on the sudden drop (i.e., sharp decrease) in ribosome occupancy at these positions, rather than identifying start codons, which are often non-canonical and difficult to define.
[0008] The present invention relates to methods and compositions for identifying and characterizing uORFs in eukaryotic organisms, particularly plants, and methods for modifying uORFs to produce desired traits. Methods for producing commercially valuable plants and crops, as well as methods for making and using them, are thereby determined.
[0009] The uORFs identified and characterized by this method can be modified to produce plants with improved traits, specifically traits that meet the needs of agriculture, food production and material production, as well as environmental remediation and carbon sequestration. These traits can provide significant value because they allow plants to thrive in harsh environments, for example, where temperature, water and nutrient availability or salinity may limit or prevent the growth of plants lacking the improved traits. These traits can also include desired morphological changes, including changes in flowering time, larger or smaller size, pest resistance, light response, changes in biochemical composition, and other desired phenotypes. In particular, with the growing interest in producing crops under controlled indoor conditions, traits such as delayed flowering or more compact structure are often desired, especially in leafy greens.
[0010] The present invention also relates to methods and compositions for eliminating unwanted vegetation (eg, weeds) in beds or fields of crops or ornamental plants, lawns, sports fields, or municipal environments.
[0011] Other aspects and embodiments of the invention are described below and can be derived from the overall teachings of the present disclosure. Summary of the Invention
[0012] This specification relates to novel methods for identifying regulatory regions within eukaryotic genomes, comprising one or more upstream open reading frames (uORFs) upstream of one or more downstream open reading frames encoding one or more polypeptides (including regulatory polypeptides or transcription factors). Once identified, the uORF sequences can be modified using gene editing techniques to induce new desired phenotypes in cells (i.e., target cells) or organisms.
[0013] In one embodiment, the present description relates to a method for identifying uORFs by applying an algorithm to ribosome profile data. Unlike traditional but generally unsuccessful methods for finding ORF sequences by looking for classic ATG start codons or even alternative start codons (with or without ribosome enrichment information), this algorithm and non-traditional methods are based on the presence of ribosome enrichment in the interval from one stop codon to the next stop codon within the same open reading frame to identify the presence of uORFs in the genome of an organism. The latter stop codon represents the end of the putative uORF. The sequence upstream of the latter stop codon represents a potential target for gene editing that destroys uORF function. Once uORF function is destroyed, the translation of the downstream main ORF will increase, and the polypeptide encoded by the main ORF will produce improved traits, i.e., desired phenotypes, in the target cells of the organism or the organism.
[0014] The method identifies putative uORFs by applying an algorithm that evaluates ribosome profiling data, the method comprising the following steps:
[0015] a) Identify existing or putative major open reading frames (ORFs) in the genome of an organism and obtain ribosome profiling data in the genome of the organism. ORFs can be identified by obtaining functional or putative functional gene sequences from original research (i.e., de novo research) or from existing public or private knowledge;
[0016] b) assessing ribosome occupancy of at least one region in the genome upstream of the ORF;
[0017] c) identifying a genomic location upstream of the ORF where ribosomes are enriched and downstream of which there is a sudden drop in ribosome occupancy;
[0018] d) identifying the stop codon of a putative or actual uORF in the genome by a sudden drop in ribosome occupancy at said position;
[0019] e) identifying the preceding or "first" stop codon upstream and in frame with the stop codon of the putative uORF; and
[0020] f) Thus, the algorithm identifies the presence of putative uORFs in the genome from ribosome enrichment data within the interval from the first stop codon to the stop codon of the putative uORF within the same open reading frame.
[0021] The present description also relates to cells, plant cells, plants or other organisms that have a targeted genetic modification introduced at a native genomic locus comprising a uORF mutation located in the 5'UTR of a gene encoding a polypeptide having cell regulatory activity. %, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity to the polypeptide provided in the sequence listing of the present application. The targeted genetic modification increases the expression level and / or activity of the encoded polypeptide having cellular regulatory activity.
[0022] This specification also relates to crops, turf, weeds or ornamental plants containing the target genetic modification introduced.The target genetic modification introduced comprises non-native alleles, and this non-native allele further comprises the mutation in the uORF in the 5'UTR of the gene encoding the polypeptide with cell regulation activity.This polypeptide has the amino acid sequence identical with the sequence provided in the sequence table provided in the application.The sequence table identifies the locus of the polypeptide of interest that coding is controlled by upstream uORF, and the identification position and sequence of uORF in the reference plant genome (Arabidopsis thaliana).The existence of the uORF upstream of the mORF encoding homologous polypeptides in the target crop can be retained conventionally.Compared with the reference or control plant of the same species lacking non-native alleles, the genetically modified plant shows improved traits. The amino acid sequence is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identical to the sequences provided in the sequence listing.
[0023] This specification also relates to a method for producing improved traits in crops, which comprises introducing a targeted genetic modification into the genome of the crop. The targeted genetic modification produces a non-native allele of a gene, which further comprises a mutation of a uORF in the 5'UTR of a gene encoding a polypeptide with cell regulatory activity. The amino acid sequence of the polypeptide has a percentage identity with the polypeptide provided in the sequence table submitted with this specification. Plants of the crop are then selected, and the selected plants contain the non-native allele and show improved traits when compared to reference or control plants of the same species that lack the non-native allele. The targeted genetic modification modulates the expression level and / or activity of the encoded polypeptide with transcriptional regulatory activity; and the percentage identity is at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or about 100%.
[0024] This specification also relates to a method for killing plant cells, which is related to a mutation of the uORF upstream of the main ORF, and the main ORF encoding can trigger the necrosis-induced polypeptide of plant cell death. In this method, each part of the plant is contacted with a suspension containing Agrobacterium (Agrobacterium) strain cells, and the Agrobacterium strain contains a nucleic acid construct. The nucleic acid construct can also be delivered to plants by other mechanisms, including coated nanoparticles (such as but not limited to DNA nanoparticles, carbon nanotubes, silicon carbide powder, magnetic transfection, peptide nanoparticles and clay nanosheets). Lv et al. 2020, Plant Journal, Vol. 104, 880-891 (doi.org / 10.1111 / tpj.14973) describe in detail the various methods of delivering nucleic acid constructs to plant cells. The nucleic acid construct includes a gene editing system, and the system expresses guide RNA in plant cells, and the guide RNA introduces mutations in the uORF upstream of the main ORF in the plant genome.
[0025] This specification also relates to a herbicidal composition for contact with a plant (such as a weed), wherein the plant's genome contains a major ORF encoding a necrosis-inducing polypeptide that triggers plant cell death. The herbicidal composition comprises a suspension containing cells of an Agrobacterium strain containing a nucleic acid construct comprising a gene editing system. The gene editing system expresses a guide RNA in cells of the target weed, and the guide RNA introduces a mutation in a uORF upstream of the major ORF in the plant or weed genome.
[0026] This specification also relates to a genetically modified cell comprising a non-naturally occurring polynucleotide produced by gene editing. The non-naturally occurring polynucleotide encodes a polypeptide that results in increased production of a target molecule or enzyme compared to a control microorganism that does not include the non-naturally occurring polynucleotide. The non-naturally occurring polynucleotide comprises a mutation in a uORF that is located in the same transcript as the primary ORF encoding the polypeptide.
[0027] This specification also relates to a method for controlling cancer cells or tumor cells, comprising contacting the cancer cells or tumor cells with a delivery vector containing a nucleic acid construct comprising a gene editing system that expresses a guide RNA in the cell, wherein the guide RNA introduces a mutation into a uORF upstream of a main ORF in the cell genome. The main ORF encodes a polypeptide that triggers cancer cell or tumor cell death or inhibits cell division.
[0028] The present disclosure also relates to a method for improving plant traits by exogenously applying short peptides encoded by uORFs (so-called "uPEPs"). Through such methods, uPEPs can be used as biostimulants to enhance the growth, yield, quality, harvestability, and / or performance of crops.
[0029] Sequence Listing and Brief Description of the Figures
[0030] The sequence listing provides exemplary polynucleotide and polypeptide sequences of the present disclosure. The sequence listing is named UOR-0002P_ST25, created on May 27, 2022, and is 6,751,561 bytes in size. The entire content of this sequence listing is incorporated herein by reference.
[0031] Figures 1 to 3 A graphic representation of the ribosome profile and analysis is shown. The black line represents the ribosome coverage over the length of the cDNA sequence. The long ORF (mORF) is a solid black line above the line, and the AUG is annotated as a black triangle at coverage = 0. All ribosome profile ratios (cross (X)) and ribosome profile differences (plus sign (+)) are shown above coverage = 0 (and scaled to fit the maximum ribosome profile count). All possible AUG codons are black circles at coverage = 0. Stop-stop intervals > 50 are shown below coverage = 0 as dashed, dotted, and dotted horizontal lines (representing three reading frames) (note that the long stop-stop corresponding to the long ORF here is a dashed line extending upstream of the ATG start codon). Stop-stop segments with statistically significant ratios (cross (X)) and / or differences (plus sign (+)) are selected for annotation at the 3' end.
[0032] Figure 4Describe a binary vector that can be delivered to plant cells by Agrobacterium or other methods (such as using nanoparticles). The gene constrained by the T-DNA border contains a gene editing system that will express in plant cells a guide RNA (encoded by a DNA sequence annotated as "guide") with the goal of knocking out or reducing the activity of a uORF. The uORF has a native effect of inhibiting the expression of a mORF, and the mORF encodes a polypeptide that causes a trait of interest. Transformation selection usually encodes resistance to herbicides or antibiotics, which enables the transformed edited cells to be selected and regenerated into whole plants. Through systemic expression, the polypeptide is upregulated in plant cells and the desired traits are obtained. In one embodiment of the invention, the plant in question is a weedy plant, and the T-DNA is not necessarily integrated into the host weed genome, but the gene editing system transiently expresses a guide RNA that knocks out or reduces the uORF activity of a polypeptide that can naturally inhibit cell necrosis. When the system is activated in the cells of the target weeds, necrosis is induced and the weeds are killed or controlled.
[0033] Figure 5 and Figure 6 All Arabidopsis genes (whose leader sequences are greater than 200 nucleotides and / or trailer sequences are longer than 200 nucleotides) are shown to start at AUG ( Figure 5 ) and terminate( Figure 6 ) Relative ribosome profile coverage (ribosome coverage / RNA-Seq coverage) at 100 nucleotides on either side of the codon. The histogram shows the relative expression (ribosome coverage / RNA-Seq coverage) of counts relative to Figure 5 AUG start codons in (1 = -70 to -42, 2 = -41 to -14, 3 = -13 to 14, 4 = -15 to 42 and 5 = 43 to 70) and Figure 6 The stop codons in the five regions (1 = -70 to -37, 2 = -36 to -4, 3 = -3 to 14, 4 = -15 to 42 and 5 = 43 to 70) were detected.
[0034] Figure 7Amaranthus hybridus subspecies (hybrid) is shown (hybrid contigs scaffolded to Amaranthus hypochondriacus): Modified genomic contigs of Amaranthus hybridus scaffolded to the pseudochromosome of Amaranthus hypochondriacus, which is the completed (v1.0, id57429) version. The gray columns represent the AUG codons in each of the three reading frames (note that the genes are in reverse order), and the black columns represent the stop codons (UAG, UAA, and UGA) in each of the three reading frames (note that the genes are in reverse order). Potential uORFs are defined by two adjacent stop-stop spacers in the same open reading frame. In this way, putative uORFs of genes can be identified and tested as candidates for gene editing.
[0035] Figure 8 A method for optimizing transgene activity by including a uORF in the transcript to inhibit translation of the encoded protein is described. If a transgenic event has already been generated in the target plant, the uORF can be inserted through gene editing. Alternatively, the uORF can be engineered into the transformation construct prior to the transformation process. In the latter case, the uORF is preferably introduced into the leader sequence of the transgene to be overexpressed by direct synthesis or ligation during assembly of the transformation construct. Figure 8 Footnote 1 indicated by the italicized number "1" shows the mRNA derived from the transgene; the introduced uORF represses the translation of the mORF. Figure 8 Footnote 2 indicated by the italicized number "2" shows the uORF that was intentionally introduced into this region (by gene editing) to inhibit the activity of the transgene by reducing the translation of the encoded mRNA into protein. DETAILED DESCRIPTION
[0036] definition
[0037] "uORF" is an upstream open reading frame / frame that is typically located in an mRNA transcript upstream of a major ORF encoding a protein (note that mORFs are sometimes also referred to as long ORFs or major ORFs, and in this application, the terms mORF, major ORF, long ORF, and major ORF are used interchangeably). uORFs are a type of small ORF that can act as repressors for their downstream mORFs. uORFs sometimes encode evolutionarily conserved functional peptides (such as cis-acting regulatory peptides) that can act as repressors, including, for example, through translational inhibition.
[0038] A "polypeptide" is an amino acid sequence comprising a plurality of consecutive polymerized amino acid residues, for example, at least about 15 consecutive polymerized amino acid residues, optionally at least about 30 consecutive polymerized amino acid residues, at least about 50 consecutive polymerized amino acid residues. In many cases, the polypeptide comprises a sequence of polymerized amino acid residues that is a transcription factor or a domain, portion, or fragment thereof. In addition, the polypeptide may comprise 1) a localization domain, 2) an activation domain, 3) an inhibition domain, 4) an oligomerization domain, or 5) a DNA binding domain, etc. The polypeptide optionally comprises modified amino acid residues, naturally occurring amino acid residues not encoded by codons, or non-naturally occurring amino acid residues.
[0039] "Identity" or "similarity" refers to the sequence similarity between two polynucleotide sequences or two polypeptide sequences, where identity is a more stringent comparison. The phrases "percent identity" and "% identity" refer to the percentage of sequence identity possessed in a comparison of two or more polynucleotide sequences or two or more polypeptide sequences. "Sequence similarity" refers to the percentage similarity of base pair sequences between two or more polynucleotide sequences (determined by any suitable method). The similarity of two or more sequences can be any value between 0-100%, or any integer value therebetween. Identity or similarity can be determined by comparing the positions in each sequence, which can be aligned for comparison purposes. When a position in the compared sequences is occupied by the same base or amino acid, the molecules are identical at that position. The degree of similarity or identity between polynucleotide sequences is a function of the number of identical or matching nucleotides at the shared positions of the polynucleotide sequences. The degree of identity of polypeptide sequences is a function of the number of identical amino acids at the shared positions of the polypeptide sequences. The degree of homology or similarity of polypeptide sequences is a function of the number of identical amino acids at the shared positions of the polypeptide sequences.
[0040] The terms "homolog" or "homologue" as further described and used herein refer to polypeptides or transcription factors derived from the same species or a different species that have a significant level of identity within their conserved domains and / or throughout their sequence, wherein the level of identity is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, or more compared to the first polypeptide or transcription factor. , 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identical to the first polypeptide or transcription factor, and the polypeptide or transcription factor has a similar or comparable function in the cell or organism as compared to the first polypeptide or transcription factor.
[0041] "Orthologs" refer to evolutionarily related genes that have similar sequences and similar functions. Orthologs are structurally related genes in different species that originated from a speciation event.
[0042] The term "introduced targeted genetic modification" or "targeted genetic modification" refers to an alteration of the DNA sequence of a plant at a specific chromosomal location (also called a locus) in the plant genome, which change is selected by a person skilled in the art (such as a plant breeder or molecular biologist) and introduced by a gene editing and / or selection process using a specific complementary nucleic acid molecule sequence as a guide or probe to achieve the process.
[0043] The term "native genomic locus" refers to a gene or DNA sequence present in the wild-type plant genome at a specific chromosomal position of a given species. A "native genomic locus" typically comprises a region spanning from the start to the stop codon, and any intermediate introns, which is transcribed to generate a major ORF encoding a long polypeptide, typically about 100 amino acids or more in length, and associated upstream regulatory elements, including a promoter region and any elements (such as uORFs) that control mORF activity. The uORF and the mORF regulated by the uORF are present in the same mRNA transcript; therefore, both uORF and mORF can be considered as part of the locus of the same overall native genome. The locus of the native genome is typically specified by reference to an accession number deposited in GenBank, for example, the accession number indicates the DNA sequence and the polypeptide encoded at that position. It should also be noted that the locus can encode a variety of protein variants produced by alternative splicing of mRNA, which are represented by different "gene models" represented by an accession number followed by a dot and a number.
[0044] The term "non-native allele of a gene" or "non-naturally occurring allele of a gene" refers to a sequence variant of a gene from a given plant species (wherein the term "gene" may include the protein coding region encoded by the main ORF, as well as upstream control elements (such as a promoter region) and elements (such as uORFs)), the nucleotide sequence of which is the result of human intervention (e.g., by gene editing or selection, such as by TILLING), and which is generally not naturally present in the genome of a wild-type plant of that species or in the genome of a plant of that species taken from a naturally occurring wild population. The term "TILLING" is an acronym for "targeted induced local lesionsingenome" and has been reviewed by Kurowska et al. in Appl Genet. 52(4):371–390, 2011.
[0045] As used herein, the term "variant" may refer to polynucleotides or polypeptides that differ from the presently disclosed polynucleotides or polypeptides in sequence, as shown below.
[0046] In some embodiments, the present invention relates to polynucleotide variants, wherein the difference between the polynucleotide disclosed at present and the polynucleotide variant is limited, so the nucleotide sequence of polynucleotide and polynucleotide variant is very similar generally, and is identical in many regions. Due to the degeneracy of genetic code, the difference between the nucleotide sequence of polynucleotide and polynucleotide variant may be meaningless (that is, the amino acid of polynucleotide encoding is identical, and the variant polynucleotide sequence encodes the amino acid sequence identical with the currently disclosed polynucleotide). Variant nucleotide sequence can encode different amino acid sequences, and in this case, this nucleotide difference will result in amino acid replacement, addition, disappearance, insertion, brachymemma or fusion relative to similar public polynucleotide sequence. These changes produce the polynucleotide variants of the polypeptide with at least one functional characteristic of coding. The degeneracy of genetic code has also determined that many different variant polynucleotides can also encode identical and / or substantially similar polypeptides except those sequences shown in the sequence table.
[0047] The scope of the present invention also includes variants of the nucleic acids listed in the sequence listing, i.e., nucleic acids having a sequence that is different from or complementary to one of the polynucleotide sequences in the sequence listing, which nucleic acids encode functionally equivalent polypeptides (i.e., polypeptides having a certain degree of equivalent or similar biological activity), but whose sequences differ from those in the sequence listing due to the degeneracy of the genetic code. This definition includes polymorphisms that may be easily or difficult to detect using specific oligonucleotide probes for the polynucleotide encoding the polypeptide, as well as inappropriate or unintended hybridization to allelic variants at loci other than the normal chromosomal loci of the polynucleotide sequence encoding the polypeptide.
[0048] The term "plant" includes whole plants, vegetative organs / structures (e.g., leaves, stems, and tubers), roots, flowers, and floral organs / structures (e.g., bracts, sepals, petals, stamens, carpels, anthers, and ovules), seeds (including embryos, endosperms, and seed coats), and fruits (mature ovaries), plant tissues (e.g., vascular tissues, ground tissues, etc.), and cells (e.g., guard cells, egg cells, etc.), and their progeny. The class of plants that can be used in the methods of the present invention is generally as broad as the class of higher and lower plants suitable for transformation techniques, including angiosperms (monocots and dicots), gymnosperms, ferns, horsetails, gymnophytes, lycophytes, mosses, and multicellular algae. See, e.g., Daly et al. (2001) Plant Physiol. 127: 1328-1333; Ku et al. (2000) Proc. Natl. Acad. Sci. 97: 9121-9126; see also Tudge, The Variety of Life, Oxford University Press, New York, NY (2000) pp. 547-606.
[0049] "Trait" is sometimes used interchangeably with the term "phenotype" and refers to a physiological, morphological, biochemical or physical characteristic of a cell or organism (including a plant or a specific plant material or plant cell). In some cases, this characteristic is visible to the human eye, such as seed or plant size or pigmentation, or can be measured by biochemical techniques, such as detecting the protein, starch or oil content of seeds or leaves, or by observing metabolic or physiological processes, for example, by measuring the uptake of carbon dioxide, or by observing the expression level of a gene, for example, by using Northern analysis, RT-PCR, microarray gene expression assays, RNA Seq or reporter gene expression systems, or by agricultural observations (such as stress tolerance, yield or pathogen tolerance). However, any technique can be used to measure the amount, relative level or difference of any selected compound or macromolecule in a transgenic plant.
[0050] "Property modification" refers to the detectable difference of the characteristics of a plant that ectopically expresses a polynucleotide or polypeptide of the present invention relative to a plant that does not express a polynucleotide or polypeptide of the present invention (such as a wild-type plant). In some cases, property modification can be quantitatively assessed. For example, compared to wild-type plants, property modification can cause the observed property to increase or decrease by at least about 2% (difference), at least 5% difference, at least about 10% difference, at least about 20% difference, at least about 30%, at least about 50%, at least about 70%, at least about 85%, or about 100% or even greater difference. It is known that there can be natural variation in the modified property. Therefore, compared to the distribution observed in wild-type plants, the observed property modification means that the normal distribution of property in the plant has changed.
[0051] As used herein, "wild type" or "wild-type" refers to cells, tissues, or plants that have not been genetically modified to mutate, knock out, ectopically express, or overexpress one or more of the presently disclosed target genes (such as genes encoding transcription factors). Wild-type cells, tissues, or plants can be used as controls to compare the extent and nature of expression levels and trait modifications with cells, tissues, or plants in which target gene expression is altered or ectopically expressed (e.g., the target gene has been knocked out or overexpressed).
[0052] "Yield" or "plant yield" refers to increased plant growth, increased crop growth, increased biomass and / or increased yield of plant products, and depends, in part, on temperature, plant size, organ size, planting density, light, water and nutrient availability, and how the plant responds to various stresses, such as through temperature acclimatization and water or nutrient use efficiency.
[0053] "Crops" include cultivated plants or agricultural products, which may be cereals, vegetables or fruit plants, usually considered as a group. Crops may be grown in quantities or amounts that have commercial value.
[0054] "Improved traits" that can be conferred on plants and provide environmental, commercial or ornamental advantages to crops may include, but are not limited to, traits selected from the following groups:
[0055] Increased yield, improved flavor, improved texture, altered circadian rhythm, accelerated flowering, accelerated senescence, delayed senescence, increased branching, decreased branching, increased apical dominance, decreased apical dominance, increased shade tolerance, increased root mass, increased number of root hairs, increased nutritional quality, increased fruit mass, improved fruit quality, increased germination rate, increased trichome length, decreased trichome length, decreased thorns, decreased spines, thornless, changed leaf shape, increased leaf number, decreased leaf number, changed leaf angle, changed leaf position, changed branch angle, improved peelability, decreased cell adhesion, increased cell adhesion, decreased peel thickness, decreased fruit peelreduced skin thickness, increased pericarp thickness, increased seed coat thickness, reduced seed coat thickness, seedlessness, reduced seed size, increased seed size, apomixis, increased embryogenesis, increased susceptibility to transgenic transformation, increased callus formation, increased embryo formation, increased root formation, increased cell division, reduced cell division, sterility, male sterility, pollen inactivity, stamen deficiency, carpel deficiency, increased number of carpels, increased number of petals, decreased number of petals, increased number of trichomes, decreased number of trichomes, increased stem width, decreased stem width, increased internode length, decreased internode length, increased floral organ size, altered floral organ shape, decreased floral organ size, decreased fruit abscission, decreased pod shattering, altered organ abscission Variation, mottled, increased hypocotyl length, decreased hypocotyl length, delayed flowering, cessation of flowering, sterility, increased tolerance to osmotic stress, increased photosynthesis, increased nitrogen use efficiency, increased phosphorus use efficiency, increased potassium use efficiency, increased nutrient use efficiency, increased nutrient uptake, increased metal ion uptake, increased heavy metal sequestration, increased tolerance to oxidative stress, increased pigment levels, increased salt tolerance, increased cold tolerance, increased frost tolerance, increased frost tolerance, increased dehydration stress tolerance, increased drought tolerance, improved recovery from drought, reduced wilting, increased number of plastids, increased chlorophyll content, increased thylakoid density, increased photosynthetic capacity, increased respiration, decreased respiration, increased photorespiration, decreased photorespiration , increased transpiration, decreased transpiration, increased stomatal conductance, decreased stomatal conductance, increased carbon fixation, increased carbon sequestration, increased photosynthetic rate, increased carotenoid levels, decreased carotenoid levels, increased electron transport, improved non-photochemical quenching, increased ion transport, decreased ion transport, altered carbon-nitrogen balance, increased sensitivity to hormones, decreased sensitivity to hormones, decreased sensitivity to ethylene, increased auxin levels, decreased auxin levels, increased auxin transport, increased auxin sensitivity, decreased auxin sensitivity, increased gibberellin levels, decreased gibberellin levels, increased gibberellin sensitivity, increased abscisic acid levels, decreased abscisic acid levels, increased abscisic acid sensitivity, decreased abscisic acid sensitivity, increased cytokinin levels , decreased cytokinin levels, increased cytokinin levels, decreased cytokinin levels, increased cytokinin sensitivity, decreased cytokinin sensitivity, increased jasmonic acid levels, decreased jasmonic acid levels, increased jasmonic acid sensitivity, decreased jasmonic acid sensitivity, increased salicylic acid levels, decreased salicylic acid levels, decreased salicylic acid sensitivity, increased salicylic acid sensitivity, increased strigolactone levels, decreased strigolactone levels, increased sensitivity to strigolactones, decreased sensitivity to strigolactones, decreased sensitivity to ethylene, increased sensitivity to ethylene, accelerated ripening, delayed ripening, reduced fruit spoilage, extended shelf life, improved heat stress, improved tolerance to low nitrogen conditions, seedling vigor (seedlingvigor), improved disease resistance, improved resistance to fungal pathogens, improved resistance to bacterial pathogens, improved resistance to viral pathogens, improved resistance to Botrytis, improved resistance to Erysiphe, improved resistance to Fusarium, improved resistance to Sclerotinia, enhanced resistance to rust, improved resistance to Phytophthora, and improved resistance to black leaf spot. sigatoka, increased resistance to Xanthomonas, increased resistance to necrotrophic fungi, increased resistance to biotrophic fungi, increased resistance to nematodes, increased resistance to insects, increased resistance to herbivores, increased resistance to mollusks, increased protein levels, increased oil levels, decreased lignin levels, increased CBD levels, increased anthocyanin levels, decreased anthocyanin levels, increased tissue nutrient levels, increased tissue vitamin levels, increased carbohydrate levels, decreased carbohydrate levels, increased starch levels, increased sugar levels, increased BRIX, increased protein levels , decreased protein levels, increased metabolite levels, increased photosynthetic pigment levels, increased lipid levels, decreased lipid levels, changes in fatty acid saturation, increased saturated fat levels, decreased saturated fat levels, increased tocopherol levels, decreased tocopherol levels, increased prenyllipid levels, increased tissue nutrient content, improved processability, increased caloric value, decreased chloride levels, increased alkaloid levels, decreased alkaloid levels, increased wax levels, decreased wax levels, increased wax ester levels, increased tannin levels, increased paclitaxel levels, increased lutein levels, increased bioplastic levels, increased biopolymer levels, decreased biopolymer levels, changes in starch composition, increased latex levels, increased rubber levels.
[0056] As used herein, "uPEP" refers to a small peptide encoded by a uORF.
[0057] Description of Specific Embodiments
[0058] The upstream open reading frame (uORF) is a short open reading frame that may encode a peptide and is located within the leader sequence of the messenger RNA. To avoid confusion, in this specification, the term "leader" sequence is generally used instead of the five major untranslated regions (5'UTR). This distinction is made because, by definition, uORF means translation, so the name "untranslated" may be misleading. Similarly, the three major untranslated regions (3'UTR) may also be able to translate, so in this specification, the "tail" sequence is sometimes used to refer to this region. This specification relates to a novel method for identifying regulatory regions in eukaryotic genomes, wherein the regulatory region comprises one or more uORFs that encode one or more polypeptides, including regulatory polypeptides or transcription factors. Once identified, the uORF sequence can be modified by gene editing technology to induce a new desired phenotype in the target cell or organism.
[0059] Challenges in annotating uORFs
[0060] The size of the upstream open reading frame is very small, which makes de novo annotation extremely challenging. This is because, even in a very small eukaryotic genome, the statistical probability of finding an open reading frame containing 100 amino acids or 300 nucleotides or less by chance is also very high. Therefore, it is difficult to distinguish between functional short open reading frames and open reading frames that exist by chance. For this reason, most gene prediction tools only consider open reading frames greater than 100 amino acids. The exception is that shorter amino acids are determined by experimental evidence, by homology with known short amino acid sequence genes or other related short sequences. Small peptides of less than 100 amino acids are still underrepresented in almost all genome annotations (Hellens 2016. Trend Plant Sci. 21: 317-328).
[0061] uORF annotation has become more complicated due to increasing evidence that these short open reading frames do not follow the normal convention of most annotated peptides, which is to begin with an AUG codon and a methionine amino acid. Many well-documented uORF sequences, including that in GDP-galactose pyrophosphorylase (GGP), have been shown to begin with non-AUG (also called near-cognate or non-canonical) start codons.
[0062] Taken together, these two characteristics of uORFs, namely small size and non-canonical start codons, make annotation prediction particularly difficult using computational methods alone.
[0063] Predicting and annotating uORFs using data and novel methods
[0064] Ribosome profiling uses next-generation sequencing to visualize regions on messenger RNA molecules where ribosomes reside. While a footprint does not reveal translation, it does indicate ribosome occupancy and, therefore, may implicate translation. This information is crucial for annotation of upstream open reading frames, as peptide sequences alone are rarely detected using accurate mass-based peptide detection methods. In fact, for most upstream open reading frames, ribosome profiling, along with mutational analysis, is the only evidence indicating a functional uORF.
[0065] Many methods use ribosome profiling data to guide uORF annotation; however, all methods to date rely on identifying potential uORF start and stop sites and then searching for ribosome enrichment along the candidate uORF. While most of these methods assume an ATG start, recent modifications have expanded the potential start sites to include all possible near-cognate start sites (where one nucleotide, A, U, or G, is replaced by a different nucleotide). Therefore, detecting uORFs by searching for their translation start sites is challenging because start codons other than AUG are often used. Furthermore, ribosomes accumulate in the leader sequence before translation initiation. In contrast, ribosome profiling data accurately maps translation stop sites. Three stop codons: UAA, UAG, and UGA appear to be commonly used in long open reading frames and shorter upstream open reading frames. Therefore, by using ribosome profiling data to predict stop codons, it can be inferred that the corresponding sequence interval between two in-frame stop codons contains the upstream open reading frame.
[0066] Our novel approach, the basis of the invention detailed herein, does not make any assumptions about the start position of uORFs. Instead, the open reading frame intervals from one stop codon to the next stop codon in the same open reading frame are determined, and it is assumed that if there is ribosome enrichment in this region, the start of the uORF is present downstream of the first stop codon (5' stop). The ability to identify these stop-stop intervals and to use stop-stop intervals to indicate the presence of uORFs is the novelty of this approach. By determining the boundaries of the region where a uORF is present in this way, the region can then be targeted by mutation and / or gene editing to modify or remove the uORF. The inventive method detailed herein has been put into practice by applying it to various datasets, including datasets derived from Arabidopsis thaliana.
[0067] Raw data preparation
[0068] The original data were downloaded and trimmed using the Trim Sequence (1.0.2) tool in the Galaxy environment according to the release method of each dataset.
[0069] The gene sequences of 5'UTR, 3'UTR and cDNA were downloaded from TAIR (the website used TAIR10 (updated 20101214) annotation files).
[0070] In Galaxy, ribosome profile short reads were mapped to Arabidopsis gene models using BWA (0.7.17.4). Unmapped reads were removed using the BAM filter. BED genome coverage (2.29.2 and -dz coverage output) was then used to generate a BEDgraph file for download. The file was imported into the R package in .csv format for further analysis.
[0071] Ribosome profiles at the start and end of long open reading frames
[0072] To determine ribosome profiling of long ORFs, we calculated ribosome coverage for all annotated long open reading frame cDNAs in the Arabidopsis genome. We generated ribosome profiles around the start and stop codons of long ORFs. We determined ribosome profiling coverage for all Arabidopsis genes using a 100-nucleotide window before and after the start and stop codons. We used the Liu dataset (Liu et al. 2013. Plant Cell 25:3699–3710) to generate ribosome profiling data and RNA-Seq data coverage. We determined the relative ribosome profiling value (ribosome profiling coverage divided by RNA-Seq coverage) for each nucleotide position in the sequence before and after the start and stop codons. Figure 5 and Figure 6 A graphical representation of these data is shown.
[0073] Given that ribosome profiles show clear differences in stop codons, these observations can be used to support the annotation of novel upstream open reading frames. Because ribosome profiles are less accurate in predicting start codons, potential regions where uORFs may reside were identified by identifying all stop-stop intervals within the mRNA sequence of a given gene.
[0074] Code1 identifies all stop-stop intervals where k<-50 # sets the ORF minimum to 50 bases, start_codons<-c("TGA","TAA","TAG") and stop_codons<-c("TGA","TAA","TAG"), s2 is the sequence of the gene of interest (GOI) under consideration.
[0075]
[0076] Code 2. Identify the longest stop-stop region using code modified from www.montefiore.ulg.ac.be / ~kbessonov / archived_data / GBIO009-1course2012 / presentations / HW1_2_review_slides_ORF_Finder.pdf (highlighted in bold).
[0077]
[0078] Code 3. Considering the above 15 nucleotides (offset < -15, window < -30, read_length < -29), calculate the NGS read length (read_length) + the ribosome profile count within the window at the end of the stop-stop fragment). Where GOI is the gene of interest.
[0079]
[0080]
[0081] The output of this analysis is summarized in Figure 2 wherein the gene of interest is AT4G26850.1 (VTC2 gene, GDP galactose pyrophosphorylase AT4G26850.1).
[0082] Genome-wide analysis of ribosome profiling by stop-terminated fragments
[0083] Repeat codes 1 to 3 were used for each of the 41,671 cDNAs in the TAIR annotation. Stop-stop fragments were filtered to obtain ratios that differed fivefold before and after termination, where the post-termination count was greater than one. For each of the three datasets used, boxplot analysis of the stop-stop coverage data was used to determine the quartile 1 (Q1) ribosome profile count. For gene models where the ribosome count after the stop codon was 0 and therefore a ratio could not be calculated, we used this value to determine the appropriate ribosome profile difference (called Delta).
[0084] The sequence listing includes the results of a genomic analysis of all the stop-stop regions annotated for each Arabidopsis gene using three datasets (Liu et al. 2013. Plant Cell 25:3699–3710; Hsu et al. 2016. PNAS 113:E7126-E7135; and Bazin et al. 2017. PNAS 114:E10018-E10027). Statistically significant stop-stop fragment counts are reported before and after the fragments, as well as in the leader, long ORF, and tail regions. For fragments in the leader region, these fragments are divided into the total number of short uORFs, the number of unique fragments therein (note that some leader sequences can have multiple short uORFs), and the mORF count (note that according to the definition provided here, all long open reading frames will extend to the upstream stop codon, thus entering the start AUG). For unique uORF counts (excluding multiple uORFs within the leader sequence), the number detected using ratio calculation, delta calculation (when the ribosome count after the stop codon is equal to zero) or both was summarized. Ultimately, each data list in the three data sets was compared, and only candidate uORFs that appeared in at least two of the three data sets were included in the final list of 1999 gene models with potential uORFs. In the sequence table, the identified Arabidopsis sequences containing uORFs are provided as odd-numbered sequences SEQ ID NO: 1-3997, while the corresponding predicted PEPs in each case are provided as subsequent even-numbered sequences, i.e. SEQ ID NO: 2, 4, 6 ... 3998. The polypeptide products of the major ORFs from the loci identified as containing upstream uORFs are provided in SEQ ID NO: 3999-5155. Ultimately, non-Arabidopsis DNA sequence examples containing uORFs are provided in SEQ ID NO: 5156-5227.
[0085] Descriptions of the sequences appearing in the sequence listing are shown in Table 1.
[0086] Table 1. Description of the sequences in the sequence listing
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138]
[0139]
[0140]
[0141]
[0142]
[0143]
[0144]
[0145]
[0146]
[0147]
[0148]
[0149]
[0150]
[0151]
[0152]
[0153]
[0154]
[0155]
[0156]
[0157]
[0158]
[0159]
[0160]
[0161]
[0162]
[0163]
[0164]
[0165]
[0166]
[0167]
[0168]
[0169]
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192]
[0193]
[0194]
[0195]
[0196]
[0197]
[0198]
[0199] Gene ontology analysis of candidate uORF genes
[0200] The gene lists were analyzed using the DAVID Bioinformatics Resource version 6.8 (david.ncifcrf.gov / ). Table 2 summarizes the output of this analysis. Among the functional categories analyzed, keywords captured 99.4% of the unique genes. Of these, the majority included one or more of the terms alternative splicing, transcription factor, transcription, nucleus, DNA binding, coiled-coil, and kinase. In the gene ontology analysis, the biological process categories enriched for transcription (DNA template), transcription regulation (DNA template), protein phosphorylation, and kinase were also noteworthy. Equally surprising was the detection of 18 genes involved in the response to ethylene and 20 genes involved in flower development. In the cellular function category, the largest proportion, 46.9%, was assigned to the nucleus. In the molecular function category, transcription factors, protein kinases, DNA binding, and ATP binding were all identified as enriched in the generated uORF list.
[0201] Table 2. Gene list output using DAVID bioinformatics resource analysis
[0202]
[0203] Numbering rules, transcription factors, and uPEP
[0204] The uORFs and uPEPs identified in this program are included in the sequence listing as SEQ ID NO: 2n-1 (SEQ ID NO: 1, 3, 5, 7, 9 and each other odd-numbered sequence to 3997) and SEQ ID NO: 2n (SEQ ID NO: 2, 4, 6, 8, 10 and each other even-numbered sequence to 3998), respectively, where n = 1-1999 (i.e., the uORFs identified are SEQ ID NO: 1, 3, 5, 7, 9 and each other odd-numbered sequence to 3997; the translations of the uORFs identified are SEQ ID NO: 2, 4, 6, 8, 10 and each other even-numbered sequence to 3998, respectively). These, along with a description of each transcription factor class, are included in the sequence listing; a summary of each sequence includes the AGI number, TAIR_ID, uORF 5' end, uORF 3' end, leader sequence, and each sequence is followed by the uORF or uPEP (the latter being the translation of the uORF).
[0205] We performed uORF and uORF translation (uPEP) analysis using a total of 130 transcription factor-encoding loci in Arabidopsis identified in our analysis. Our analysis identified representatives from nearly all transcription factor classes, but some families appeared particularly enriched for uORFs; these included the AP2 gene family, the homeodomain leucine zipper family, the sNF family, and one STAT transcription factor. A gene ontology list of genes involved in flowering is included in the sequence listing and contains many transcription factors, including AP2, ARF, homeodomain, and MYB transcription factors. Also of note, the flowering control protein FCA, which is thought to function as an RNA-binding protein, is included in our list of uORF-regulated loci.
[0206] Examples of individual genetic profiles
[0207] Below is a more detailed description of some example loci, where ribosome coverage, stop-stop fragments, and statistically significant ratios and increments (based on ribosome profiles before and after 3' termination) are shown. Figure 1-3 shown.
[0208] There are two clear ribosome clusters upstream of the long ORF of At2g23340.1 ( Figure 1 The uORFs indicated by open circles were selected because of their significantly different ratios ( Figure 1 ).
[0209] At4g36900.1: Figure 2 Shows a clear enrichment of ribosomes in the leader sequence (selected for significantly different ratios and increments) ( Figure 2 ).
[0210] AT4g16280.2 is a functional gene model for FCA. There are four gene models, but only .2 and .4 contain open reading frames, and alternative splicing may interfere with the candidate uORF, which is clearly identified as the enrichment of ribosomes upstream of the long ORF ( Figure 3 ).
[0211] HD-ZIP transcription factors are particularly enriched in uORFs. Figure 1 、 2 and 3 highlight some candidate genes where ribosome profiles indicate ribosome enrichment upstream of long open reading frames.
[0212] Bioinformatics analysis and identification of homologs
[0213] An important aspect of the present invention is that if the locus containing the mORF encoding the polypeptide with the desired function is identified as having an upstream uORF in the reference species, then the equivalent locus crop with the mORF encoding the polypeptide homolog in the target species also generally has a uORF. That is, the presence of uORF is generally conserved between homologous loci. Examples of this phenomenon have been confirmed in ascorbic acid biosynthesis genes between species. For example, see Zhang et al., 2018, Nature Biotechnology Vol. 36, pp. 894-898. Therefore, the implementer can apply the method herein to identify uORFs in the Arabidopsis locus, then identify the locus with the main ORF that can encode the homolog in the target crop, and then deploy gene editing to mutate the uORF sequence upstream of the mORF in the crop to eliminate the inhibition imposed by the uORF. Typically, this involves a mutant sequence between 1-1100bp upstream of the mORF start codon. Although uORFs can be edited twice or multiple times, a single edit is sufficient to achieve this purpose.
[0214] Homologs can be identified by various bioinformatics methods, as exemplified herein. The present invention can be an integrated system, computer, or computer-readable medium comprising an instruction set for determining the identity of one or more sequences in a database. In addition, the instruction set can be used to generate or identify sequences that meet any specified criteria. Further, the instruction set can be used to associate or link certain functional advantages (such as improved traits) with one or more identified sequences.
[0215] For example, the instruction set may include, for example, sequence comparison or other alignment programs, such as those available, for example, the Wisconsin Package version 10.0, such as BLAST, FASTA, PILEUP, FINDPATTERNS, etc. (GCG, Madison, WI). Public sequence databases such as GenBank, EMBL, Swiss-Prot, and PIR or private sequence databases may be searched.
[0216] Can carry out sequence alignment for comparison by the local homology algorithm of Smith and Waterman (1981) Adv.Appl.Math.2:482-489, the homology comparison algorithm of Needleman and Wunsch (1970) J.Mol.Biol.48:443-453, the similarity search method of Pearson and Lipman (1988) Proc.Natl.Acad.Sci.85:2444-2448 and the computerized realization of these algorithms.After alignment, usually carry out sequence alignment between two (or more) polynucleotides or polypeptides by identifying and comparing the local region of sequence similarity by comparing the sequence of two sequences on the comparison window.The comparison window can be the section of at least about 10 continuous positions, usually about 50 to about 200, more usually about 100 to about 150 continuous positions.Ausubel etc. provide the description of this method in its above text.
[0217] Can use a variety of methods to determine sequence relationships, including manual comparison and computer-assisted sequence comparison and analysis. Because computer-assisted methods can improve throughput, the latter method is the preferred method among the present invention. As mentioned above, there are a variety of computer programs for performing sequence comparisons available for use, or can be made by a technician.
[0218] An example of an algorithm suitable for determining percent sequence identity and sequence similarity is the BLAST algorithm, which is described in Altschul et al. (1990) J. Mol. Biol. 215: 403-410. Software for performing BLAST analysis is publicly available, for example, from the National Center for Biotechnology Information (http: / / ncbi.nlm.nih.gov / ) at the National Library of Medicine's National Center. This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that, when aligned with a sequence of the same length in the database sequence, either match or satisfy a positive threshold score, T. T is called the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood sequence hits are used as seeds to initiate the search to find longer HSPs containing them. The sequence hits are then extended in both directions along each sequence until the cumulative alignment score is improved. For nucleotide sequences, cumulative scores are calculated using the parameters M (reward score for a pair of matching residues; always > 0) and N (penalty score for mismatching residues; always < 0). For amino acid sequences, cumulative scores are calculated using a scoring matrix. Extension of sequence hits in each direction are halted when: the cumulative alignment score falls below its maximum achieved value by X; the cumulative score goes to zero or below due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) of 10, a cutoff of 100, M=5, N=-4, and a comparison of both chains. For amino acid sequences, the BLASTP program uses as defaults: word length 3, expectation (E) 10, BLOSUM62 scoring matrix (see Henikoff and Henikoff, (1992) Proc. Natl. Acad. Sci. 89: 10915-10919). Unless otherwise indicated, "sequence identity" herein refers to the percent sequence identity generated from tblastx using the NCBI version of the algorithm at default settings using gapped alignments and filter "off" (see, e.g., the NIH NLM NCBI website www.ncbi.nlm.nih.gov / , supra).
[0219] In addition to calculating percent sequence identity, the BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, for example, Karlin and Altschul (1993) Proc. Natl. Acad. Sci. 90:5873-5787). A similarity metric provided by the BLAST algorithm is the minimum probability sum (P(N)), which indicates the probability that a match will occur by chance between two nucleotide or amino acid sequences. For example, if the minimum probability sum when comparing a test nucleic acid to a reference nucleic acid is less than about 0.1, or less than about 0.01, or even less than about 0.001, then the nucleic acid is considered similar to the reference sequence (thus, homologous in this case). Another example of a useful sequence alignment algorithm is PILEUP. PILEUP creates a multiple sequence alignment from a group of related sequences using progressive, pairwise alignments. The program can align, for example, up to 300 sequences of a maximum length of 5,000 letters.
[0220] The integrated system or computer typically includes a user input interface that allows the user to selectively view one or more sequence records corresponding to one or more character strings, and an instruction set for aligning one or more character strings with each other or identifying one or more regions of sequence similarity using additional character strings. The system may include links to one or more character strings associated with a specific phenotype or gene function. Typically, the system includes a user-readable output element that displays the alignment generated by the alignment instruction set.
[0221] The methods of the present invention can be implemented in a localized or distributed computing environment. In a distributed environment, the methods can be implemented on a single computer containing multiple processors, or on multiple computers. These computers can be interconnected, but more preferably, the computers are nodes on a network. The network can be a general or private local area network or wide area network. In certain preferred embodiments, the computers can be components of an intranet or the Internet, or a "cloud" computing platform such as provided by Amazon Web Services.
[0222] Thus, the present invention provides methods for identifying sequences that are similar or homologous to one or more polynucleotides described herein, or one or more target polypeptides encoded by the polynucleotides, or other polynucleotides described herein, which can include linking or associating a given phenotype (such as the cellular biosynthetic ability of the target molecule) with a sequence. In the methods, a sequence database (local or across the Internet or intranet) is provided and the sequence database is queried using the relevant sequences herein and the relevant phenotype or function in the cellular biosynthesis of the target molecule.
[0223] Any sequence in this article can be input into the database before or after the query database.This both can expand the database, also can insert the control sequence into the database before the query step.Can detect the control sequence by query to guarantee the overall integrity of database and query.As previously mentioned, can use the interface execution query based on Web browser.For example, the database can be a centralized public database (such as GenBank) or a private database, and query can be carried out from a remote terminal or computer through the Internet or intranet.
[0224] Any sequence herein can be used to identify similar, homologous sequences in a genome (such as a target plant genome) using the above methods. Homologous polynucleotide sequences that have conserved or equivalent functions in delivering a desired trait typically have at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 200%, 201 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% sequence identity.
[0225] Generation of transgenic plant cells and plants
[0226] Can produce by multiple mature technology and comprise polynucleotide described in the present invention and / or express the genetically modified cell, vegetable cell, plant explant, plant tissue, plant organ or the plant of polypeptide described in the present invention.After making up transformation vector (most typical is expression cassette, comprises one or more polynucleotide of the present invention or its section), can use standard technique that polynucleotide is introduced cell to produce genetically modified cell or clone.Optionally, can regenerate genetically modified plant cell to produce explant, tissue or transgenic plant.
[0227] Currently, transformation and proliferation and / or regeneration of cells have become routine, and the choice of the most appropriate transformation technology will be determined by the practitioner. The choice of method will vary depending on the organism to be transformed; those skilled in the art will recognize the applicability of a particular method for a given organism. Suitable methods may include, but are not limited to: electroporation of protoplasts; liposome-mediated transformation; polyethylene glycol (PEG)-mediated transformation; Li-mediated transformation, using, for example, virus transformation; microinjection of cells; micro-projectile bombardment of cells; vacuum infiltration; or Agrobacterium-mediated transformation. In the present invention, transformation involves introducing a recombinant nucleotide sequence into a host cell in a manner that allows for stable or transient expression of the sequence, thereby resulting in the expression of the encoded polypeptide, which in turn leads to the production of the desired target molecule.
[0228] After transformation, preferably use the dominant selection marker that is included in the transformation vector to select genetically modified cells.Usually, such markers will give the transformed cells antibiotic or herbicide resistance, and the selection of transformants can be completed by exposing the cells to the antibiotic or herbicide of appropriate concentration.Alternatively, markers based on color (such as GFP or GUS) can be used to select transformed cells, or the target molecule that can be selected based on the expression of the polynucleotide in the expression cassette introduced by RT-PCR detection or detection of transgenic cells produced.
[0229] Methods for generating genetically modified plants (including transgenic plants) are reviewed in Keshavareddy et al. 2018, Int. J. Curr. Microbiol. App. Sci (2018) 7(7): 2656-2668. Detailed methods have also been published in U.S. Patents No. 7,345,217 (Zhang, March 18, 2008), No. 7,511,190 (Creelman et al., March 31, 2009), No. 7,196,245 (Jiang et al., March 27, 2007), and No. 7,663,025 (Heard et al., February 16, 2010).
[0230] Introduction of targeted genetic modifications by gene editing
[0231] A preferred method of practicing the present invention is to use genome editing to produce the "target genetic modifications" described herein. The terms "genome editing," "genome-edited," "genome modification," and "genetic modification" are used interchangeably to describe plants having specific DNA sequence changes in their genomes, wherein those DNA sequence changes include changes in specific nucleotides, deletions of specific nucleotide sequences, or insertions of specific nucleotide sequences.
[0232] As used herein, the technology of introducing "target genetic modification" refers to any method, protocol or technology that allows precise and / or targeted editing of a specific location (also referred to as a "locus" or "biological locus") in a plant genome using site-specific nucleases (such as meganucleases, zinc finger nucleases (ZFNs), RNA-guided endonucleases (e.g., CRISPR / Cas9 systems), TALE-endonucleases (TALENs), recombinases or transposases (i.e., the editing is largely or completely non-random). CRISPR is an abbreviation for clustered, regularly interspaced, short, palindromic repeats, and Cas is an abbreviation for CRISPR-associated protein; for review, see Khandagal and Nadal, Plant Biotechnol. Rep., 2016, 10, 327. Engineered meganucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs) can also be used. U.S. Patent Application 2016 / 0032297 provides details of these methods. Another gene editing method that can be applied is the use of so-called ARCUS nucleases, which exploit the properties of a naturally occurring gene editing enzyme, the homing endonuclease I-CreI, which has evolved in nature to make a single, highly specific DNA edit and then turn itself off using its built-in safety switch.
[0233] Genome editing tools can accurately change the genome structure at a specific target location. These tools can be effectively used to generate plants with high crop yields, desired composition changes, and resistance to biotic and abiotic stresses. It is challenging to achieve all the desired modifications using a specific genome editing tool. Therefore, a variety of genome editing tools have been developed to facilitate efficient genome editing. Some of the main genome editing tools used to edit plant genomes are: homologous recombination (HR), zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), pentapeptide repeat proteins (PPRs), CRISPR / Cas9 systems, RNA interference (RNAi), homologous gene modification (cisgenesis), and endogenous gene modification (intragenesis). In addition, site-directed sequence editing and oligonucleotide-directed mutagenesis can edit the genome at the single nucleotide level. Recently, adenine base editors (ABEs) have been developed to mutate AT base pairs to GC base pairs. ABE uses deoxyadenine deaminase (deoxyadeninedeaminase) (TadA) and a catalytically impaired Cas9 nickase to mutate AT base pairs to GC base pairs. Mohanta et al., Genes (Basel). 2017 Dec;8(12):399 summarizes these methods and their applicability.
[0234] Such genome editing methods encompass a variety of methods to precisely remove genes, gene fragments, to change the DNA sequence of coding sequences or control sequences, or to insert new DNA sequences into genes or protein coding regions to reduce or increase the expression of target genes in plant genomes (Belhaj, K. 2013, Plant Methods, 9, 39; Khandogale and Nadal, 2016, Plant Biotechnol Rep, 10, 327). Preferred methods involve in vivo site-specific cleavage to achieve double-strand breaks in plant genome genomic DNA at specific DNA sequences using nucleases and host plant DNA repair systems. There are a variety of methods that can be used to create double-strand breaks in genomic DNA to achieve genome editing, including the use of CRISPR / Cas systems.
[0235] A detailed overview of the CRISPR / Cas system and its useful applications can be found at www.addgene.org / guides / crispr /
[0236] The CRISPR / Cas genome editing system provides the flexibility to modify specific sequences within the genome and is able to perform a range of different edits, including activating or upregulating a target locus or knocking out a target locus. The method relies on providing a Cas enzyme and a short guide RNA "gRNA" containing a short guide sequence (~20bp) that has sequence complementarity with a target DNA sequence in the plant genome. Depending on the type of Cas enzyme, DNA, RNA / DNA hybrids, or double-stranded DNA guide polynucleotides can be used. The guide portion of the guide polynucleotide directs the Cas enzyme to the desired cleavage site and carries a recognition sequence for binding to the Cas enzyme.
[0237] The target in a plant genome can be any DNA sequence of approximately 20 nucleotides, provided that the sequence is unique compared to the rest of the genome and that the target is immediately adjacent to a protospacer adjacent motif (PAM). The PAM sequence serves as a binding signal for Cas9, but the exact sequence depends on the Cas protein used. A list of Cas proteins and PAM sequences can be found at www.addgene.org / guides / crispr / #pam-table
[0238] The simplest application of CRISPR / Cas is to produce a knockout or loss-of-function allele in a target locus. gRNA targets the Cas enzyme to a specific locus in the genome, which then produces a double-strand break. The resulting DSB is then repaired by one of the general repair pathways present in the cell. This typically results in a small nucleotide insertion or deletion (inDel) at the DSB site. In most cases, small inDels in the target DNA can lead to amino acid deletions, insertions, or frameshift mutations, resulting in premature stop codons within the open reading frame (ORF) of the target gene. The ideal result is a loss-of-function mutation within the target gene. However, the strength of the knockout phenotype of a given mutant cell must be experimentally verified, such as by testing the presence of transcripts in the target ORF using RT-PCR or hybridization-based methods. These traits make the CRISPR / Cas system a suitable tool for knocking out uORFs.
[0239] CRISPR / Cas can also be used to produce more complex changes to the native sequence of the target site in the genome. This may involve inserting sequences, replacing sequences, or editing specific bases to insert or create new domains in the polypeptide encoded by the desired locus. One method of introducing such changes is to utilize a high-fidelity but inefficient high-fidelity homology-directed (HDR) repair pathway within the cell. In order to use HDR for such precise modifications, a DNA repair template containing the desired genomic modification that the implementer wishes to create at the target site must be delivered to the cell type of interest using gRNA and Cas9 or Cas9 nickase. The repair template must contain the desired editing and additional homologous sequences (called left homology arms and right homology arms) immediately upstream and downstream of the target. The length of each homology arm depends on the size of the introduced changes, and the larger the insertion, the longer the homology arm. Since Cas9 cleavage efficiency is relatively high and HDR efficiency is relatively low, most of the DSBs induced by Cas9 will be repaired to produce edits that do not contain specific desired changes. Therefore, additional confirmation / screening steps are required to select one or more cells containing the desired changes from the editing population. These cells can then be regenerated into cell populations, tissues, organs, or whole plants or plant populations.Such selection can be achieved by incorporating marker sequences into the editing that can be easily screened, or by PCR or hybridization-based methods.
[0240] CRISPR-related gene editing systems can also be used to change specific bases without the need for double-strand breaks. Such methods are referred to as "base editing" systems in the art. Base editing can irreversibly convert a specific DNA base to another base at the target genomic locus, such as converting C to T, or converting A to G. Unlike other genome editing tools, base editing can be achieved without causing double-strand breaks. When point mutations are to be introduced at the target locus, base editing is more effective than traditional genome editing techniques. Since many genetic diseases originate from point mutations, base editing has important applications in disease research. Using these systems, skilled implementers can create target genetic modifications comprising amino acid substitutions or creating start or stop codons.
[0241] To avoid reliance on inefficient HDR, researchers have developed two classes of base editors: cytosine base editors (CBEs) and adenine base editors (ABEs). Cytosine base editors are created by fusing the Cas9 nickase or a catalytically inactive "dead" Cas9 (dCas9) to a cytidine deaminase such as APOBEC. Like traditional CRISPR technology, base editors are targeted to specific loci via gRNA, where they can convert cytidine to uridine within a small editing window near the PAM site. Uridine is then converted to thymidine through base excision repair, resulting in a C to T change. Similarly, adenosine base editors have been engineered to convert adenosine to inosine, which is processed by cells like guanosine, resulting in an A to G change.
[0242] Adenine DNA deaminases do not exist in nature, but these enzymes were generated through the directed evolution of Escherichia coli TadA (tRNA adenine deaminase). As with cytosine base editors, the evolved TadA domain was fused to the Cas9 protein to generate adenine base editors. There are multiple Cas9 variants for both types of base editors, including high-fidelity Cas9. Further progress has been made by optimizing fusion expression, modifying the linker region between the Cas variant and the deaminase to adjust the editing window, or adding fusions that improve product purity, such as DNA glycosylase inhibitors (UGIs) or phage Mu-derived Gam proteins (Mu-GAM).
[0243] While many base editors are designed to work in a very narrow window close to the PAM sequence, some base editing systems create a wide range of single-nucleotide variants (somatic hypermutations) within a wider editing window and are therefore well-suited for directed evolution applications. Examples of these base editing systems include targeted AID-mediated mutagenesis (TAM) and CRISPR-X, in which Cas9 is fused to activation-induced cytidine deaminase (AID).
[0244] Other CRISPR systems, particularly the type VI CRISPR enzymes Cas13a / C2c2 and Cas13b, target RNA rather than DNA. The overactive adenosine deaminase ADAR2 (E488Q) acting on RNA was fused to a catalytically inactive Cas13b to create a programmable RNA base editor that converts adenosine in RNA to inosine (called REPAIR). Since inosine is functionally equivalent to guanosine, the result is an A->G change in RNA. The catalytically inactive Cas13b ortholog dPspCas13b from Prevotella does not appear to require a specific sequence adjacent to the RNA target, making it a very flexible editing system. Editors based on the second ADAR variant ADAR2 (E488Q / T375G) showed improved specificity, while editors carrying the δ-984-1090 ADAR truncation retained RNA editing ability and were small enough to be packaged in AAV particles.
[0245] In the context of this specification, it should be recognized that the term Cas nuclease includes any nuclease that can site-specifically recognize CRISPR sequences based on gRNA or DNA sequences, and includes Cas9, Cpf1, and other nucleases described below. Many authors have identified CRISPR / Cas genome editing as a preferred way to edit the genomes of complex organisms (Sander and Joung, 2013, Nat Biotech, 2014, 32, 347; Wright et al., 2016, Cell, 164, 29), including plants (Zhang et al., 2016, Journal of Genetics and Genomics, 43, 151; Puchta 2016, Plant J., 87, 5; Khangale and Nadaf, 2016, Plant Biotechnol. Rep., 10, 327). U.S. Patent Application 2016 / 020822 details materials and methods useful for genome editing in plants using the CRISPR / Cas9 system and describes numerous uses of the CRISPR / Cas9 system for genome editing of a range of gene targets in crops.
[0246] It is further recognized that many variations of the CRISPR / Cas system can be used to apply the invention herein, including the use of wild-type Cas9 from Streptococcus pyogenes (type II Cas) (Barakate and Stephens, 2016, Frontiers in Plant Science, 7, 765; Bortesi and Fischer, 2015, Biotechnology Advances 5, 33, 41; Cong et al., 2013, Science, 339, 819; Rani et al., 2016, Biotechnology Letters, 1-16; Tsai et al., 2015, Nature biotechnology, 33, 187). Other examples include Tru-gRNA / Cas9, in which off-target mutations are significantly reduced (Fu et al., 2014, Nature biotechnology, 32, 279; Osakabe et al., 2016, Scientific Reports, 6, 26685; Smith et al., 2016, Scientific Reports, 6, 26685; Zhang et al., 2016, Scientific Reports, 6, 28566), and highly specific Cas9 (mutated Streptococcus pyogenes Cas9) with almost no off-target activity (Kleinstiver et al., 2016, Nature 529, 490; Slaymaker et al., 2016, Science, 351, 84). Further variations include type I and type III systems in which multiple Cas proteins are expressed to achieve editing (Li et al., 2016, Nucleic acids research, 44:e34; Luo et al., 2015, Nucleic acids research, 43, 674), type V Cas systems using the Cpfl enzyme (Kim et al., 2016, Nature biotechnology, 34, 863; Toth et al., 2016, Biology Direct, 11, 46; Zetsche et al., 2015, Cell, 163, 759), DNA-guided editing using the NgAgo alginate enzyme from Natronobacter gregoris, which employs a guide DNA (Xu et al., 2016, Genome Biology, 17, 186), and a dual-vector system in which the Cas9 and gRNA expression cassettes are carried on separate vectors (Cong et al., 2013, Science, 339, 819).The unique nuclease Cpf1 is an alternative to Cas9 and has the advantage of reducing off-target editing (resulting in unwanted mutations in the host genome) compared to the Cas9 system. Examples of crop genome editing using the CRISPR / Cpf1 system include rice (Tang et al., 2017, Nature Plants 3, 1-5; Wu et al., 2017, Molecular Plant, March 16, 2017) and soybean (Kim et al., 2017, Nat Commun. 8, 14406). Other authors have described the use of Argu-related proteins as alternatives to the CRISPR system for gene editing (Hegge et al. Nature Rev. Microbiol. 2017. Epub 2017 / 07 / 25. pmid: 28736447; Swarts et al. Nucleic Acids Res. 2015; 43(10): 5120–9. Epub 2015 / 05 / 01. pmid: 25925567; Swarts et al. Nature. 2014; 507(7491): 258–61. Epub 2014 / 02 / 18. pmid: 24531762. See also PCT Application No. PCT / US2019 / 025163 and / or Publication No. WO2019204266A1.
[0247] Published patent application WO2019195157 lists detailed methods for gene editing in plants to create new crop traits, including selecting cells containing the desired editing, and methods for introducing CRISPR system components into initial target plant cells. As described therein, the "guide polynucleotide" in the CRISPR system also relates to a polynucleotide sequence that can form a complex with the Cas endonuclease and enable the Cas endonuclease to recognize and optionally cleave the DNA target site. The guide polynucleotide can be a single molecule (i.e., a single guide RNA (gRNA), which is a synthetic fusion between crRNA and a portion of the tracrRNA sequence) or two molecules (i.e., crRNA and tracrRNA found in the natural Cas9 system in bacteria). The guide polynucleotide sequence can be provided as an RNA sequence or transcribed from a DNA sequence to produce an RNA sequence. The guide polynucleotide sequence can also be provided as an RNA-DNA sequence combination (see, for example, Yin, H. et al., 2018, Nature Chemical Biology, 14, 311). As used herein, a "guide RNA" sequence comprises a variable targeting domain (referred to as a "guide") complementary to a target site in the genome, and an RNA sequence (referred to as a "guide RNA scaffold") that interacts with Cas9 or Cpf1 endonucleases. A guide polynucleotide comprising only ribonucleic acid is also referred to as a "guide RNA." As used herein, a "guide target sequence" refers to a sequence adjacent to a PAM site in genomic DNA, to which a gRNA will bind to cleave DNA. A "guide target sequence" is typically complementary to the "guide" portion of a gRNA, but depending on its position, several mismatches can be tolerated and still allow Cas-mediated DNA cleavage. The method also provides for introducing a single guide RNA (gRNA) into a plant. A single guide RNA (gRNA) includes a nucleotide sequence complementary to a target chromosomal DNA. The gRNA can be, for example, an engineered single-stranded guide RNA comprising a crRNA sequence (complementary to the target DNA sequence) and a common tracrRNA sequence, or a crRNA-tracrRNA hybridization. The gRNA can be introduced into a cell or organism as DNA with an appropriate promoter, in vitro transcribed RNA, or synthetic RNA. The basic principles for designing guide RNA for any target gene of interest are well known in the art, such as those described by Brazelton et al. (Brazelton, VA et al., 2015, GM Crops & Food, 6, 266-276) and Zhu (Zhu, LJ 2015, Frontiers in Biology, 10, 289-296).
[0248] Published patent applications WO2019195157 and WO2019204266A1 also provide examples of mutation types that can lead to increased activity of transcription factor polypeptides. These mutations include mutations in the coding sequence that result in changes in the amino acids of the encoded protein.
[0249] In certain preferred embodiments of the present invention, the guide polynucleotide / Cas endonuclease system can be used to allow insertion of a promoter or promoter element (such as an enhancer element) of any of the transcription factor sequences of the present invention, wherein the promoter insertion (or promoter element deletion) results in any one or a combination of the following: permanent activation of the locus, increased promoter activity (increased promoter strength), increased promoter tissue specificity, decreased promoter tissue specificity, new promoter activity, extended gene expression window, modification of the time or developmental progression of gene expression, mutation of DNA binding elements, and / or addition of DNA binding elements.
[0250] The guide RNA / Cas endonuclease system can be used to insert promoter elements to increase the expression of the transcription factor sequence of the present invention. Promoter elements (such as enhancer elements) are usually introduced into the promoter of the driving gene expression cassette in the form of multiple copies for trait gene testing or producing transgenic plants expressing specific traits. The enhancer element can be, but is not limited to, a 35S enhancer element (Benfey et al., EMBO J., 1989; 8: 2195-2202). In some plants (events), the enhancer element can lead to a desired phenotype, increased yield, or a change in the expression pattern of the trait of interest. It may be necessary to remove additional copies of the enhancer element while keeping the trait gene cassette intact at its integrated genomic location. The guide RNA / Cas endonuclease can be used to remove unwanted enhancing elements from the plant genome. The guide RNA can be designed to contain a variable targeting region, a target site sequence of 12-30bp adjacent to NGG (PAM) in the targeting enhancer. The Cas endonuclease can be cleaved to insert one or more enhancers.
[0251] To inhibit the function of the target uORF and activate the downstream mORF, bases can be deleted from the uORF, or additional stop codons can be created. Other mutations or edits that can be used to inhibit the function of the target uORF include mutations of the start ATG codon, amino acid deletions, insertions, or frameshift mutations that lead to premature stop codons, or any deleterious mutations within the uORF. In some cases, it may be optimal to replace bases within the uORF to produce an optimized phenotype, thereby obtaining increased yield, increased stress tolerance, or altered biochemical composition without the occurrence of substantial abnormalities such as organ abnormalities or dwarfism.
[0252] Delivery of gene editing components into plant cells and plants:
[0253] Sandhya et al. 2020. J. Genet. Eng. Biotechnol. 18:25. Published online July 7, 2020, doi:10.1186 / s43141-020-00036-8 describes methods for delivering gene-editing tools, such as CRISPR / Cas9 components, into plants to perform gene editing procedures. Efficient delivery of CRISPR / Cas9 components (including a guide sequence, CAS9, and, if applicable, a DNA repair template containing the desired sequence to edit) into plant cells is crucial for efficient editing. Practitioners can choose from a variety of delivery methods to introduce gene-editing components into plant cells. These methods include Agrobacterium-mediated transformation, bombardment or gene gun transformation methods, floral dip, and PEG-mediated protoplast transformation. Additional methods include nanoparticle and pollen magnetofection-mediated delivery systems (Kwak et al., 2019, Nature Nanotechnology, DOI 10.1038 / S41565-019-0375-4) (Demirer et al., 2019, Nature Nanotechnology, DOI 10.1038 / S41565-019-0382-5). CRISPR constructs can be coated on gold particles for introduction into plant cells via gene gun-mediated delivery, CRISPR constructs can be transfected into protoplasts using PEG, or introduced via Agrobacterium strains containing CRISPR vectors. Components can also be introduced by floral dip (Castel et al., 2019. PLoS One 14: e0204778) or pollen tube pathway-based methods. Next, plants or plant cells containing the target genetic modification produced by the introduced CRISPR system are selected. This may involve regenerating cells containing the modifications into explants, plant tissues, or whole plants. In some cases, this process involves selecting explants carrying genome edits on selection plates and regenerating whole plants. Finally, PCR and Sanger sequencing are typically used to confirm that the desired sequence edits have been successfully introduced into the selected plants. The selected plants are then examined to confirm that they exhibit the target traits initially sought by introducing the genome modification.
[0254] Sandhya et al. 2020. J. Genet. Eng. Biotechnol. 18:25. Published online July 7, 2020, doi:10.1186 / s43141-020-00036-8. A table showing methods that have been successfully applied to specific crops is provided. For example, the following plants have been successfully gene-edited using PEG-mediated delivery of CRISPR system components: apple, cabbage (Brassica oleracea), Brassica rapa, watermelon (Citrullus lanatus), soybean (Glycine max), grapevine (Grapevine), rice (Oryza sativa), petunia (Petunia), Physcomitrella patens, tomato (Solanum lycopersicum), common wheat (Triticum aestivum), and corn (Zea mays). By way of further example, the following plants have all been successfully gene-edited using particle bombardment-mediated delivery of CRISPR system components: soybean, barley (Hordeum vulgare), rice, common wheat, and maize. By way of further example, the following plants have all been successfully gene-edited using particle bombardment-mediated delivery of CRISPR system components: Arabidopsis thaliana, banana, sweet orange (Citrus sinensis), cucumber (Cucumis sativum), soybean, kiwifruit, Lotus japonicus, liverwort (Marchantia polymorpha), Medicago truncatula, Nicotiana benthamaina, tobacco (Nicotiana tabacum), rice, Populus, Salvia miltiorrhiza, tomato, sorghum (Sorghum bicolor), common wheat, and maize.
[0255] Activation of polypeptides and their homologs by targeted genetic modifications including uORF mutations:
[0256] Upstream ORFs (uORFs) consist of short segments of mRNA located within the 5' UTR or upstream region of a gene encoding a regulatory protein of interest (the coding sequence is often referred to as the main ORF or mORF). uORFs can be co-framed or out-of-frame with the main coding sequence of the gene of interest. A significant fraction of eukaryotic mRNAs contain uORFs within the 5' leader sequence preceding the primary functional protein-encoding ORF (Kochetov, 2008, BioEssays 30:683–691). uORFs typically encode short peptides that negatively regulate the activity of the regulatory protein encoded by the upstream gene. Furthermore, uORFs sometimes begin with non-canonical codons (e.g., ACG instead of AUG), and the encoded peptides are typically much less than 100 residues in length. These characteristics make uORFs difficult to identify through automated bioinformatics searches, and experimental confirmation is often required to determine whether a putative uORF functions as a negative regulator of downstream coding sequences. For example, Laing et al., 2015. Plant Cell 27:772-786; DOI:10.1105 / tpc.114.133777 removed a uORF encoding a 60- to 65-residue peptide upstream of GGP in lettuce and showed that this was sufficient to produce traits including increased ascorbic acid levels. Indeed, in some cases, the peptides encoded by the uORF are very short; for example, in humans, a functional peptide of only 6 amino acids has been identified as being encoded by a uORF. Studies have shown that various plant uORFs can regulate mORF translation based on the levels of various key metabolites in the cell, such as polyamines, sucrose, phosphocholine, and ascorbic acid. It has been proposed that the function of several peptides encoded by these uORFs is to slow or arrest ribosomes, thereby limiting translation of downstream main ORFs encoding regulatory proteins (see Hellens et al., 2016. Trends Plant Sci., 21:317-328.dx.doi.org / 10.1016 / j.tplants.2015.11.005, and references therein).
[0257] Recently, it has been proposed that uORF gene editing may provide a general method to activate crop genes in a highly targeted manner to produce traits of interest (Zhang and Voytas, 2019, Natl. Sci. Rev. 6:391, doi.org / 10.1093 / nsr / nwy123). In particular, this can avoid many of the shortcomings associated with traditional methods, which typically involve inserting large foreign DNA fragments into the genome, such as sequences of strong promoters, enhancers, or engineered artificial transcriptional activators. In fact, many genetic modification-related problems associated with transgenic integration, including the lack of consumer acceptance of transgenic products, may be eliminated in the next generation of crop traits produced by targeted gene editing to knock out uORFs in regulatory genes.
[0258] An increasing number of uORFs have been identified upstream of genes encoding transcriptional regulators. In many cases, uORFs and / or their encoded short peptides appear to be controlled by metabolic signals (van der Horst 2020. Plant Physiol. 182: 110-122, published online on August 26, 2019. doi: 10.1104 / pp.19.00940 and references therein). These include the S1 group bZIPs, including the (HG1) bZIP transcription factor, which controls amino acid and sugar metabolism, and whose activity is regulated by sucrose. SAC51 (HG15) is a bHLH transcription factor involved in xylem differentiation and regulated by thermospermine. Another example is the HsfB1 / TBF1 (HG18) HSF transcription factor, which is involved in thermotolerance and growth-to-defense transitions regulated by galactosidase.
[0259] Other examples of uORF-regulated transcription factors involve a group of uORF-containing genes found in the AUXIN RESPONSE FACTOR transcription factor family (Hellens 2016. Trend Plant Sci. 21:317-328; Schepetilnikov, M. et al. 2013, EMBO J. 32:1087–1102; Nishimura, T. et al., 2005. Plant Cell 17:2940–2953; Zhou, F. et al., 2010. BMC Plant Biol. 10:193).
[0260] uORFs have also been identified as playing an important role in regulating the activity of transcription factors that control light responses, including transcription factors from the bZIP family (Kurihara et al., 2018. Proc. Natl. Acad. Sci. 115:7831-7836).
[0261] It should be noted that additional confirmation / selection steps are required to select one or more cells from the edited population that contain the desired change. These cells can then be regenerated into cell populations, tissues, organs, or whole plants or plant populations, and optionally, can be further screened to select plants that display the desired trait produced by gene editing of the uORF sequence.
[0262] Traits that can be improved
[0263] Trait improvements of particular interest include improvements to seeds (such as embryos or endosperms), fruits, roots, flowers, leaves, stems, buds, seedlings, and the like, including: enhanced tolerance to environmental conditions (including freezing, cold, high temperature, drought, water saturation, radiation, and ozone); increased tolerance to microbial, fungal, or viral diseases; increased tolerance to pest attacks (including insects, nematodes, mollies, parasitic higher plants (e.g., Ligustrum lucidum)), etc.; reduced herbicide sensitivity; increased tolerance to heavy metals or enhanced ability to absorb heavy metals; improved growth under low light conditions (e.g., low light and / or short daylight), or changes in the expression levels of genes of interest. Other phenotypes that can be altered relate to the production of plant metabolites, such as taxol, tocopherols, tocotrienols, sterols, phytosterols, vitamins, wax monomers, antioxidants, amino acids, lignin, cellulose, tannins, prenyl lipids (such as chlorophyll and carotenoids), glucosinolates and terpenoids, enhanced production or altered composition of proteins or oils (especially in seeds), or modified sugar (insoluble or soluble) and / or starch composition. Physical plant traits that can be modified include cell development (such as the number of trichomes), size and number of fruits and seeds, yield of plant parts (such as stems, leaves, inflorescences and roots), stability of seeds during storage, seed pod characteristics (such as friability), length and number of root hairs, internode distance or seed coat quality. Plant growth characteristics that may be modified include growth rate, seed germination rate, plant and seedling vigor, leaf and flower senescence, male sterility, apomixis, flowering time, flower abscission, nitrogen uptake rate, osmotic sensitivity to soluble sugar concentration, biomass or transpiration characteristics, and plant architectural characteristics such as apical dominance, branching pattern, organ number, organ identity, organ shape or size.
[0264] Example
[0265] The above is a general description of the present invention, which will be more readily understood with reference to the following examples, which are intended only to illustrate certain aspects and embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will recognize that a genetic modification associated with a particular primary trait may also be associated with at least one other, unrelated, inherent secondary trait that is not predictable from the primary trait.
[0266] Example 1. Identification of the presence of an upstream open reading frame.
[0267] 1A. A method for identifying the presence of an upstream open reading frame (uORF) by applying an algorithm to ribosome profiling data, wherein the algorithm identifies the presence of a uORF based on the presence of ribosome enrichment in the interval from one stop codon to the next stop codon within the same open reading frame.
[0268] 1B. The method of claim 1A, wherein the identified uORF, or a uORF located upstream of a major ORF operably linked to the identified uORF, is mutated in the cell and the polypeptide it encodes is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135 7%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identical to the host ORF, and the mutation results in increased translation of the major ORF operably linked to the uORF.
[0269] 1C. The method of claim 1B, wherein the uORF is mutated by introducing at least one gene edit in the uORF.
[0270] Example 2. Reduction of uORF function in plants
[0271] 2A. The method of Representation 1A or Representation 1B, wherein the uORF is mutated in the plant, and the loss or reduction of the function of the uORF results in increased translation of a polypeptide encoding polynucleotide operably linked to the uORF.
[0272] 2B. The method of claim 2A, wherein the polynucleotide encodes a polypeptide whose expression results in cell death, inhibition of cell division, or an improved trait selected from the group consisting of:
[0273] Increased yield, improved flavor, improved texture, altered circadian rhythm, accelerated flowering, accelerated senescence, delayed senescence, increased branching, decreased branching, increased apical dominance, decreased apical dominance, increased shade tolerance, increased root mass, increased number of root hairs, increased nutritional quality, increased fruit mass, improved fruit quality, improved germination rate, increased trichome length, decreased trichome length, decreased trichome length, decreased thorns, decreased thorns, thornless, altered leaf shape, increased leaf number, decreased leaf number, altered leaf angle, altered leaf position, altered branch angle, improved peelability, decreased cell adhesion, increased cell adhesion, decreased pericarp thickness, decreased pericarp thickness, increased pericarp thickness, increased seed coat thickness, decreased seed coat thickness, seedless, decreased seed size, increased seed size, no melting Syngamy, increased embryogenesis, increased sensitivity to transgenic transformation, increased callus formation, increased embryo formation, increased root formation, increased cell division, decreased cell division, sterility, male sterility, pollen inactivity, stamen deficiency, carpel deficiency, increased carpel number, increased petal number, decreased petal number, increased trichome number, decreased trichome number, increased stem width, decreased stem width, increased internode length, decreased internode length, increased floral organ size, altered floral organ shape, decreased floral organ size, decreased fruit abscission, decreased pod shattering, altered organ abscission, mottling, increased hypocotyl length, decreased hypocotyl length, delayed flowering, cessation of flowering, sterility, increased osmotic stress tolerance, improved photosynthesis, improved nitrogen use efficiency, improved phosphorus use efficiency High, improved potassium use efficiency, improved nutrient use efficiency, increased nutrient absorption, increased metal ion absorption, increased heavy metal sequestration, increased tolerance to oxidative stress, increased pigment levels,, improved salt tolerance, improved cold tolerance, improved frost damage tolerance, improved frost damage tolerance, improved dehydration stress tolerance, improved drought tolerance, improved recovery ability after drought, reduced wilting, increased number of plastids, increased chlorophyll content, increased thylakoid density, increased photosynthetic capacity, increased respiration, reduced respiration, increased photorespiration, reduced photorespiration, increased transpiration, reduced transpiration, increased stomatal conductance, reduced stomatal conductance, increased carbon fixation, increased carbon sequestration, increased photosynthetic rate, increased carotenoid levels, reduced carotenoid levels, Increased electron transport, improved non-photochemical quenching, increased ion transport, decreased ion transport, altered carbon-nitrogen balance, increased sensitivity to hormones, decreased sensitivity to hormones, decreased sensitivity to ethylene, increased auxin levels, decreased auxin levels, increased auxin transport, increased auxin sensitivity, decreased auxin sensitivity, increased gibberellin levels, decreased gibberellin levels, increased gibberellin sensitivity, increased abscisic acid levels, decreased abscisic acid levels, increased abscisic acid sensitivity, decreased abscisic acid sensitivity, increased cytokinin levels, decreased cytokinin levels, increased cytokinin levels, decreased cytokinin levels, increased cytokinin sensitivity, decreased cytokinin sensitivity, increased jasmonic acid levels, decreased jasmonic acid levels,Increased sensitivity to jasmonic acid, decreased sensitivity to jasmonic acid, increased salicylic acid levels, decreased salicylic acid levels, decreased salicylic acid sensitivity, increased salicylic acid sensitivity, increased strigolactone levels, decreased strigolactone levels, increased sensitivity to strigolactones, decreased sensitivity to strigolactones, decreased sensitivity to ethylene, increased sensitivity to ethylene, accelerated ripening, delayed ripening, reduced fruit spoilage, extended shelf life, improved heat stress, improved tolerance to low nitrogen conditions, increased seedling growth, improved disease resistance, Improved resistance to fungal pathogens, improved resistance to bacterial pathogens, improved resistance to viral pathogens, improved resistance to Sporodia; improved resistance to powdery mildew; improved resistance to Fusarium; improved resistance to Sclerotinia; increased resistance to rust, improved resistance to Phytophthora, improved resistance to black leaf spot, improved resistance to Xanthomonas, increased resistance to necrotic fungi, increased resistance to biotrophic fungi, increased resistance to nematodes, increased resistance to insects, increased resistance to herbivores, and resistance to soft Enhanced animal performance, increased protein level, increased oil level, decreased lignin level, increased CBD level, increased anthocyanin level, decreased anthocyanin level, increased tissue nutrient level, increased tissue vitamin level, increased carbohydrate level, decreased carbohydrate level, increased starch level, increased sugar level, increased BRIX, increased protein level, decreased protein level, increased metabolite level, increased photosynthetic pigment level, increased lipid level, decreased lipid level, changed fatty acid saturation, increased saturated fat level, decreased saturated fat level, increased tocopherol level, decreased tocopherol level, increased isopentenol level, increased tissue nutrient content, improved processability, increased caloric value, decreased chlorine level, increased alkaloid level, decreased alkaloid level, increased wax level, decreased wax level, increased wax ester level, increased tannin level, increased paclitaxel level, increased lutein level, increased bioplastic level, increased biopolymer level, decreased biopolymer level, changed starch composition, increased latex level, increased rubber level.
[0274] 2C. The method of claim 2B, wherein the uORF regulates translation of a polynucleotide and derepression of the translation results in toxic effects or cell death in the plant.
[0275] 2D. The method of claim 2B, wherein the plants are weeds or other undesirable plants.
[0276] 2E. The method of claim 2B, wherein the uORF regulates translation of the polynucleotide, and the increased translation results in delayed flowering or bolting in the plant compared to a reference or control plant of the same species.
[0277] 2F. The method of claim 2B, wherein the uORF regulates translation of the polynucleotide, and the increased translation results in earlier flowering in the plant compared to a reference or control plant of the same species.
[0278] 2G. The method of claim 2A-2F, wherein the plant is a crop plant, a fruit crop, a cereal crop, a feed crop, a forestry crop, an energy crop, a turf plant, a weed plant, a woody plant, a monocotyledon, a dicotyledon, an algae, a grass, an ornamental plant, a leafy plant, lettuce, a salad green, or a vegetable. green), pine, eucalyptus, tomato, alfalfa, soybean, clover, carrot, celery, parsnip, cabbage, radish, rapeseed, broccoli, melon, cucumber, wheat, corn, cotton, rice, barley, millet, rye, potato, tomato, tobacco, sugar beet, sugarcane, miscanthus, energy cane, bamboo, switchgrass, miscanthus, jatropha, bermudagrass, lentils, chickpeas, peas, beans, pepper, strawberries, blackberries, raspberries, blueberries, banana, pineapple, citrus, nut crops, rubber tree, oil palm, coffee tree, cocoa tree, tea tree, or plants from the Solanaceae, Leguminosae, Apiaceae, Cucurbitaceae, Poaceae, or Cruciferae families.
[0279] 2H. A method of increasing yield, size, grain yield, seed yield or biomass of a plant, wherein the plant is produced by the method of any one of Recipes 2A-2G.
[0280] 2I. A plant produced by the method of any one of Recipient Schemes 2A-2H.
[0281] 2J. A plant or plant cell comprising:
[0282] A targeted genetic modification introduced at a native genomic locus comprising a mutation in a uORF, wherein the native genomic locus comprises a major ORF operably linked downstream of the uORF, wherein the major ORF encodes a polypeptide having regulatory activity comprising an amino acid sequence having a percent identity to a polypeptide selected from the group consisting of SEQ ID NOs: 3999–5227.
[0283] wherein the percent identity is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100%; and
[0284] Targeted genetic modification of the uORF increases the expression level and / or activity of the encoded polypeptide having regulatory activity.
[0285] 2K. A crop, turf, weed, or ornamental plant containing an introduced target genetic modification, the target genetic modification comprising a non-native allele, the non-native allele further comprising a mutation within a uORF, the uORF operably linked to a polynucleotide comprising a major ORF encoding a polypeptide having cell regulatory activity, the polypeptide having an amino acid sequence identical to a sequence selected from the group consisting of SEQ ID NOs: 3999-5227; and
[0286] The genetically modified plant exhibits an improved trait selected from the group consisting of:
[0287] Compared to reference or control plants of the same species lacking the non-native allele,
[0288] Increased yield, improved flavor, improved texture, altered circadian rhythm, accelerated flowering, accelerated senescence, delayed senescence, increased branching, decreased branching, increased apical dominance, decreased apical dominance, increased shade tolerance, increased root mass, increased number of root hairs, increased nutritional quality, increased fruit mass, improved fruit quality, increased germination rate, increased trichome length, decreased trichome length, decreased thorns, decreased spines, thornless, altered leaf shape, increased leaf number, decreased leaf number, altered leaf angle, altered leaf position, altered branch angle, improved peelability, increased cell death, increased leaf senescence, decreased cell adhesion, decreased cell adhesion increased, decreased pericarp thickness, decreased pericarp thickness, increased pericarp thickness, increased seed coat thickness, decreased seed coat thickness, seedless, decreased seed size, increased seed size, apomixis, increased embryogenesis, increased sensitivity to transgenic transformation, increased callus formation, increased embryo formation, increased root formation, increased cell division, decreased cell division, sterility, male sterility, pollen inactivity, stamen deficiency, carpel deficiency, increased number of carpels, increased number of petals, decreased number of petals, increased number of trichomes, decreased number of trichomes, increased stem width, decreased stem width, increased internode length, decreased internode length, increased floral organ size, altered floral organ shape, decreased floral organ size, decreased fruit abscission, decreased pod shattering, Altered organ abscission, mottling, increased hypocotyl length, decreased hypocotyl length, delayed flowering, cessation of flowering, sterility, increased tolerance to osmotic stress, improved photosynthesis, improved nitrogen use efficiency, improved phosphorus use efficiency, improved potassium use efficiency, improved nutrient use efficiency, increased nutrient uptake, increased metal ion uptake, increased heavy metal sequestration, improved tolerance to oxidative stress, increased pigment levels, improved salt tolerance, improved cold tolerance, increased frost tolerance, increased frost tolerance, improved dehydration stress tolerance, improved drought tolerance, improved recovery from drought, reduced wilting, increased number of plastids, increased chlorophyll content, increased thylakoid density, increased photosynthetic capacity, increased respiration, decreased respiration, increased photorespiration, decreased photorespiration weakened, increased transpiration, weakened transpiration, increased stomatal conductance, decreased stomatal conductance, increased carbon fixation, increased carbon sequestration, increased photosynthetic rate, increased carotenoid levels, decreased carotenoid levels, increased electron transport, improved non-photochemical quenching, increased ion transport, decreased ion transport, altered carbon-nitrogen balance, increased sensitivity to hormones, decreased sensitivity to hormones, decreased sensitivity to ethylene, increased auxin levels, decreased auxin levels, increased auxin transport, increased auxin sensitivity, decreased auxin sensitivity, increased gibberellin levels, decreased gibberellin levels, increased gibberellin sensitivity, increased abscisic acid levels, decreased abscisic acid levels, increased abscisic acid sensitivity, decreased abscisic acid sensitivity, increased cytokinin levels,Decreased cytokinin levels, increased cytokinin levels, decreased cytokinin levels, increased cytokinin sensitivity, decreased cytokinin sensitivity, increased jasmonic acid levels, decreased jasmonic acid levels, increased jasmonic acid sensitivity, decreased jasmonic acid sensitivity, increased salicylic acid levels, decreased salicylic acid levels, decreased salicylic acid sensitivity, increased salicylic acid sensitivity, increased strigolactone levels, decreased strigolactone levels, increased sensitivity to strigolactones, decreased sensitivity to strigolactones, decreased sensitivity to ethylene, increased sensitivity to ethylene, accelerated ripening, delayed ripening, reduced fruit spoilage, extended shelf life, improved heat stress, improved tolerance to low nitrogen conditions, seedling vigor vigor), improved disease resistance, improved resistance to fungal pathogens, improved resistance to bacterial pathogens, improved resistance to viral pathogens, improved resistance to Botrytis, improved resistance to Erysiphe, improved resistance to Fusarium, improved resistance to Sclerotinia, enhanced resistance to rust, improved resistance to Phytophthora, and improved resistance to black leaf spot. sigatoka, increased resistance to Xanthomonas, increased resistance to necrotrophic fungi, increased resistance to biotrophic fungi, increased resistance to nematodes, increased resistance to insects, increased resistance to herbivores, increased resistance to mollusks, increased protein levels, increased oil levels, decreased lignin levels, increased CBD levels, increased anthocyanin levels, decreased anthocyanin levels, increased tissue nutrient levels, increased tissue vitamin levels, increased carbohydrate levels, decreased carbohydrate levels, increased starch levels, increased sugar levels, increased BRIX, increased protein levels , decreased protein levels, increased metabolite levels, increased photosynthetic pigment levels, increased lipid levels, decreased lipid levels, altered fatty acid saturation, increased saturated fat levels, decreased saturated fat levels, increased tocopherol levels, decreased tocopherol levels, increased prenyllipid levels, increased tissue nutrient content, improved processability, increased caloric value, decreased chloride levels, increased alkaloid levels, decreased alkaloid levels, increased wax levels, decreased wax levels, increased wax ester levels, increased tannin levels, increased paclitaxel levels, increased lutein levels, increased bioplastic levels, increased biopolymer levels, decreased biopolymer levels, altered starch composition, increased latex levels, increased rubber levels;,
[0289] And, the amino acid sequence identity is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100%.
[0290] 2L. A genetically modified crop, turf, weed or ornamental plant according to claim 2K, wherein the uORF is targeted based on the final nucleotide of the uORF stop codon, which is located between -1 and -1500 nucleotides upstream of the start codon of the polynucleotide encoding the polypeptide having regulatory activity.
[0291] 2M. The genetically modified crop plant of claim 2K, wherein the introduction of the target genetic modification into the plant does not negatively affect the yield, size, organ shape, or vigor of the plant, and the resulting modified crop plant exhibits an improved trait selected from the group consisting of:
[0292] When the genetically modified plants are grown under greenhouse or field conditions, compared to control or reference plants,
[0293] Increased yield, Improved flavor, Improved texture, Altered circadian rhythm, Accelerated flowering, Accelerated senescence, Delayed senescence, Increased branching, Decreased branching, Increased apical dominance, Decreased apical dominance, Increased shade tolerance, Increased root mass, Increased number of root hairs, Increased nutritional quality, Increased fruit mass, Improved fruit quality, Improved germination rate, Increased trichome length, Decreased trichome length, Decreased trichome length, Decreased thorns, Decreased thorns, No thorns, Altered leaf shape, Increased leaf number, Decreased leaf number, Altered leaf angle, Altered leaf position, Altered branch angle, Improved peelability, Decreased cell adhesion, Improved cell adhesion, Decreased pericarp thickness, Decreased pericarp thickness, Increased pericarp thickness, Increased seed coat thickness, Decreased seed coat thickness, Seedless, Decreased seed size, Increased seed size, No fusion Reproduction, increased embryogenesis, increased sensitivity to transgenic transformation, increased callus formation, increased embryo formation, increased root formation, increased cell division, decreased cell division, sterility, male sterility, pollen inactivity, stamen deficiency, carpel deficiency, increased carpel number, increased petal number, decreased petal number, increased trichome number, decreased trichome number, increased stem width, decreased stem width, increased internode length, decreased internode length, increased floral organ size, altered floral organ shape, decreased floral organ size, decreased fruit abscission, decreased pod shattering, altered organ abscission, mottling, increased hypocotyl length, decreased hypocotyl length, delayed flowering, cessation of flowering, sterility, improved osmotic stress tolerance, improved photosynthesis, improved nitrogen use efficiency, improved phosphorus use efficiency , improved potassium use efficiency, improved nutrient use efficiency, increased nutrient absorption, increased metal ion absorption, increased heavy metal sequestration, increased tolerance to oxidative stress, increased pigment levels, , improved salt tolerance, improved cold tolerance, improved frost damage tolerance, improved frost damage tolerance, improved dehydration stress tolerance, improved drought tolerance, improved recovery capacity after drought, reduced wilting, increased number of plastids, increased chlorophyll content, increased thylakoid density, increased photosynthetic capacity, increased respiration, reduced respiration, increased photorespiration, reduced photorespiration, increased transpiration, reduced transpiration, increased stomatal conductance, reduced stomatal conductance, increased carbon fixation, increased carbon sequestration, increased photosynthetic rate, increased carotenoid levels, reduced carotenoid levels, Increased electron transport, improved non-photochemical quenching, increased ion transport, decreased ion transport, altered carbon-nitrogen balance, increased sensitivity to hormones, decreased sensitivity to hormones, decreased sensitivity to ethylene, increased auxin levels, decreased auxin levels, increased auxin transport, increased auxin sensitivity, decreased auxin sensitivity, increased gibberellin levels, decreased gibberellin levels, increased gibberellin sensitivity, increased abscisic acid levels, decreased abscisic acid levels, increased abscisic acid sensitivity, decreased abscisic acid sensitivity, increased cytokinin levels, decreased cytokinin levels, increased cytokinin levels, decreased cytokinin levels, increased cytokinin sensitivity, decreased cytokinin sensitivity, increased jasmonic acid levels, decreased jasmonic acid levels,Increased sensitivity to jasmonic acid, decreased sensitivity to jasmonic acid, increased salicylic acid levels, decreased salicylic acid levels, decreased salicylic acid sensitivity, increased salicylic acid sensitivity, increased strigolactone levels, decreased strigolactone levels, increased sensitivity to strigolactones, decreased sensitivity to strigolactones, decreased sensitivity to ethylene, increased sensitivity to ethylene, accelerated ripening, delayed ripening, reduced fruit spoilage, extended shelf life, improved heat stress, improved tolerance to low nitrogen conditions, increased seedling growth, improved disease resistance, Improved resistance to fungal pathogens, improved resistance to bacterial pathogens, improved resistance to viral pathogens, improved resistance to Sporodia; improved resistance to powdery mildew; improved resistance to Fusarium; improved resistance to Sclerotinia; increased resistance to rust, improved resistance to Phytophthora, improved resistance to black leaf spot, improved resistance to Xanthomonas, increased resistance to necrotic fungi, increased resistance to biotrophic fungi, increased resistance to nematodes, increased resistance to insects, increased resistance to herbivores, and resistance to soft Enhanced animal performance, increased protein level, increased oil level, decreased lignin level, increased CBD level, increased anthocyanin level, decreased anthocyanin level, increased tissue nutrient level, increased tissue vitamin level, increased carbohydrate level, decreased carbohydrate level, increased starch level, increased sugar level, increased BRIX, increased protein level, decreased protein level, increased metabolite level, increased photosynthetic pigment level, increased lipid level, decreased lipid level, changed fatty acid saturation, increased saturated fat level, decreased saturated fat level, increased tocopherol level, decreased tocopherol level, increased isopentenol level, increased tissue nutrient content, improved processability, increased caloric value, decreased chlorine level, increased alkaloid level, decreased alkaloid level, increased wax level, decreased wax level, increased wax ester level, increased tannin level, increased paclitaxel level, increased lutein level, increased bioplastic level, increased biopolymer level, decreased biopolymer level, changed starch composition, increased latex level, increased rubber level.
[0294] 2N. The genetically modified plant according to claim 2K, wherein the non-native allele and / or
[0295] The non-native allele is generated by a method selected from the group consisting of:
[0296] DNA marker-assisted breeding;
[0297] deletion, insertion and / or substitution of one or more polynucleotides;
[0298] site-directed mutagenesis;
[0299] Chemical mutagenesis;
[0300] Targeted Induced Local Damage in Genomes (TILLING); and
[0301] gene editing technology;
[0302] The gene editing technology includes a method based on transcription activator-like effector nuclease (TALEN) or zinc finger nuclease (ZFN), or gene editing using CRISPR-Cas nuclease technology, wherein the CRISPR-Cas nuclease technology uses a nuclease selected from Cas nuclease, Cas9 nuclease, CasX nuclease, CasY nuclease, Cpfl nuclease, C2cl nuclease, C2c2 nuclease (Casl3a nuclease) or C2c3 nuclease, NgAgo nuclease, or a gene editing technology using a base editing deaminase, an engineered site-specific macronuclease, an Argu-related protein or a CreI-related endonuclease.
[0303] 20. The genetically modified plant of claim 2K, wherein the uORF comprises any one of SEQ ID NOs: 5156–5227.
[0304] 2P. A method of producing an improved trait selected from the group consisting of:
[0305] Increased yield in crop plants, improved flavor, improved texture, altered circadian rhythm, accelerated flowering, accelerated senescence, delayed senescence, increased branching, decreased branching, increased apical dominance, decreased apical dominance, increased shade tolerance, increased root mass, increased number of root hairs, increased nutritional quality, increased fruit mass, improved fruit quality, improved germination rate, increased trichome length, decreased trichome length, decreased thorns, decreased thorns, thornlessness, altered leaf shape, increased leaf number, decreased leaf number, altered leaf angle, altered leaf position, altered branch angle, increased peelability, decreased cell adhesion, increased cell adhesion, decreased pericarp thickness, decreased pericarp thickness, increased pericarp thickness, increased seed coat thickness, decreased seed coat thickness, seedlessness, decreased seed size, increased seed size , apomixis, increased embryogenesis, increased sensitivity to transgenic transformation, increased callus formation, increased embryo formation, increased root formation, increased cell division, decreased cell division, sterility, male sterility, pollen inactivity, stamen deficiency, carpel deficiency, increased carpel number, increased number of petals, decreased number of petals, increased number of trichomes, decreased number of trichomes, increased stem width, decreased stem width, increased internode length, decreased internode length, increased floral organ size, altered floral organ shape, decreased floral organ size, decreased fruit abscission, decreased pod shattering, altered organ abscission, mottling, increased hypocotyl length, decreased hypocotyl length, delayed flowering, cessation of flowering, sterility, improved osmotic stress tolerance, improved photosynthesis, improved nitrogen use efficiency, and phosphorus use efficiency Improved nutrient uptake, improved potassium use efficiency, improved nutrient use efficiency, increased nutrient absorption, increased metal ion absorption, increased heavy metal sequestration, improved tolerance to oxidative stress, increased pigment levels, improved salt tolerance, improved cold tolerance, improved frost damage tolerance, improved frost damage tolerance, improved dehydration stress tolerance, improved drought tolerance, improved recovery from drought, reduced wilting, increased number of plastids, increased chlorophyll content, increased thylakoid density, increased photosynthetic capacity, increased respiration, reduced respiration, increased photorespiration, reduced photorespiration, increased transpiration, reduced transpiration, increased stomatal conductance, reduced stomatal conductance, increased carbon fixation, increased carbon sequestration, increased photosynthetic rate, increased carotenoid levels, reduced carotenoid levels , increased electron transport, improved non-photochemical quenching, increased ion transport, decreased ion transport, altered carbon and nitrogen balance, increased sensitivity to hormones, decreased sensitivity to hormones, decreased sensitivity to ethylene, increased auxin levels, decreased auxin levels, increased auxin transport, increased auxin sensitivity, decreased auxin sensitivity, increased gibberellin levels, decreased gibberellin levels, increased gibberellin sensitivity, increased abscisic acid levels, decreased abscisic acid levels, increased abscisic acid sensitivity, decreased abscisic acid sensitivity, increased cytokinin levels, decreased cytokinin levels, increased cytokinin levels, decreased cytokinin levels, increased cytokinin sensitivity, decreased cytokinin sensitivity, increased jasmonic acid levels, decreased jasmonic acid levels,Increased jasmonic acid sensitivity, decreased jasmonic acid sensitivity, increased salicylic acid levels, decreased salicylic acid levels, decreased salicylic acid sensitivity, increased salicylic acid sensitivity, increased strigolactone levels, decreased strigolactone levels, increased sensitivity to strigolactones, decreased sensitivity to strigolactones, decreased sensitivity to ethylene, increased sensitivity to ethylene, accelerated ripening, delayed ripening, reduced fruit spoilage, extended shelf life, improved heat stress, improved tolerance to low nitrogen conditions, improved seedling vigor, improved disease resistance, increased resistance to fungal pathogens, increased resistance to bacterial pathogens, increased resistance to viral pathogens Improved resistance to fungi, improved resistance to Sporodia, improved resistance to powdery mildew, improved resistance to Fusarium, improved resistance to Sclerotinia, enhanced resistance to rust, improved resistance to Phytophthora, improved resistance to black leaf spot, improved resistance to Xanthomonas, improved resistance to nematodes, improved resistance to insects, improved resistance to herbivores, improved resistance to mollusks, increased protein levels, increased oil levels, decreased lignin levels, increased CBD levels, increased anthocyanin levels, decreased anthocyanin levels, tissue Increased nutrient levels, increased tissue vitamin levels, increased carbohydrate levels, decreased carbohydrate levels, increased starch levels, increased sugar levels, increased BRIX, increased protein levels, decreased protein levels, increased metabolite levels, increased photosynthetic pigment levels, increased lipid levels, decreased lipid levels, altered fatty acid saturation, increased saturated fat levels, decreased saturated fat levels, increased tocopherol levels, decreased tocopherol levels, increased levels of prenols, increased tissue nutrient content, improved processability, improved caloric value, decreased chlorine levels, increased alkaloid levels, decreased alkaloid levels, increased wax levels, decreased wax levels, increased wax ester levels, increased tannin levels, increased paclitaxel levels, increased lutein levels, increased bioplastic levels, increased biopolymer levels, decreased biopolymer levels, altered starch composition, increased latex levels, increased rubber levels, comprising introducing a targeted genetic modification into the genome of the crop plant, thereby generating a non-native allele of a gene, the gene further comprising a mutation in a uORF operably linked to a polynucleotide encoding a polypeptide having cell regulatory activity, the polypeptide having an amino acid sequence having a percent identity to a polypeptide selected from the group consisting of: SEQ ID NO: 1 ID NO:3999-5227.
[0306] selecting plants of a crop plant, and wherein the selected plants contain a non-native allele and exhibit an improved trait compared to a reference or control plant of the same species lacking the non-native allele;
[0307] wherein the target genetic modification modulates the expression level and / or activity of an encoded polypeptide having transcriptional regulatory activity; and
[0308] The percent identity is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100%.
[0309] 2Q. The genetically modified plant of claim 2P, wherein the non-native allele is selected and / or produced by a method selected from the group consisting of:
[0310] DNA marker-assisted breeding;
[0311] deletion, insertion and / or substitution of one or more polynucleotides;
[0312] site-directed mutagenesis;
[0313] Chemical mutagenesis;
[0314] Targeted Induced Local Damage in Genomes (TILLING); and
[0315] gene editing technology;
[0316] The gene editing technology includes a method based on transcription activator-like effector nuclease (TALEN) or zinc finger nuclease (ZFN), or gene editing using CRISPR-Cas nuclease technology, wherein the CRISPR-Cas nuclease technology uses a nuclease selected from Cas nuclease, Cas9 nuclease, CasX nuclease, CasY nuclease, Cpfl nuclease, C2cl nuclease, C2c2 nuclease (Casl3a nuclease) or C2c3 nuclease, NgAgo nuclease, or a gene editing technology using a base editing deaminase, an engineered site-specific giant nuclease, an Argu-related protein or a CreI-related endonuclease.
[0317] 2R. The genetically modified crop plant according to claim 2P, wherein the introduction of the target genetic modification into the plant does not negatively affect the size, organ shape, or vigor of the plant, and the resulting modified crop plant exhibits an improved trait selected from the group consisting of:
[0318] Increased yield, Improved flavor, Improved texture, Altered circadian rhythm, Accelerated flowering, Accelerated senescence, Delayed senescence, Increased branching, Decreased branching, Increased apical dominance, Decreased apical dominance, Increased shade tolerance, Increased root mass, Increased number of root hairs, Increased nutritional quality, Increased fruit mass, Improved fruit quality, Improved germination rate, Increased trichome length, Decreased trichome length, Decreased trichome length, Decreased thorns, Decreased thorns, No thorns, Altered leaf shape, Increased leaf number, Decreased leaf number, Altered leaf angle, Altered leaf position, Altered branch angle, Improved peelability, Decreased cell adhesion, Improved cell adhesion, Decreased pericarp thickness, Decreased pericarp thickness, Increased pericarp thickness, Increased seed coat thickness, Decreased seed coat thickness, Seedless, Decreased seed size, Increased seed size, No fusion Reproduction, increased embryogenesis, increased sensitivity to transgenic transformation, increased callus formation, increased embryo formation, increased root formation, increased cell division, decreased cell division, sterility, male sterility, pollen inactivity, stamen deficiency, carpel deficiency, increased carpel number, increased petal number, decreased petal number, increased trichome number, decreased trichome number, increased stem width, decreased stem width, increased internode length, decreased internode length, increased floral organ size, altered floral organ shape, decreased floral organ size, decreased fruit abscission, decreased pod shattering, altered organ abscission, mottling, increased hypocotyl length, decreased hypocotyl length, delayed flowering, cessation of flowering, sterility, improved osmotic stress tolerance, improved photosynthesis, improved nitrogen use efficiency, improved phosphorus use efficiency , improved potassium use efficiency, improved nutrient use efficiency, increased nutrient absorption, increased metal ion absorption, increased heavy metal sequestration, increased tolerance to oxidative stress, increased pigment levels, , improved salt tolerance, improved cold tolerance, improved frost damage tolerance, improved frost damage tolerance, improved dehydration stress tolerance, improved drought tolerance, improved recovery capacity after drought, reduced wilting, increased number of plastids, increased chlorophyll content, increased thylakoid density, increased photosynthetic capacity, increased respiration, reduced respiration, increased photorespiration, reduced photorespiration, increased transpiration, reduced transpiration, increased stomatal conductance, reduced stomatal conductance, increased carbon fixation, increased carbon sequestration, increased photosynthetic rate, increased carotenoid levels, reduced carotenoid levels, Increased electron transport, improved non-photochemical quenching, increased ion transport, decreased ion transport, altered carbon-nitrogen balance, increased sensitivity to hormones, decreased sensitivity to hormones, decreased sensitivity to ethylene, increased auxin levels, decreased auxin levels, increased auxin transport, increased auxin sensitivity, decreased auxin sensitivity, increased gibberellin levels, decreased gibberellin levels, increased gibberellin sensitivity, increased abscisic acid levels, decreased abscisic acid levels, increased abscisic acid sensitivity, decreased abscisic acid sensitivity, increased cytokinin levels, decreased cytokinin levels, increased cytokinin levels, decreased cytokinin levels, increased cytokinin sensitivity, decreased cytokinin sensitivity, increased jasmonic acid levels, decreased jasmonic acid levels,Increased sensitivity to jasmonic acid, decreased sensitivity to jasmonic acid, increased salicylic acid levels, decreased salicylic acid levels, decreased salicylic acid sensitivity, increased salicylic acid sensitivity, increased strigolactone levels, decreased strigolactone levels, increased sensitivity to strigolactones, decreased sensitivity to strigolactones, decreased sensitivity to ethylene, increased sensitivity to ethylene, accelerated ripening, delayed ripening, reduced fruit spoilage, extended shelf life, improved heat stress, improved tolerance to low nitrogen conditions, increased seedling growth, improved disease resistance, Improved resistance to fungal pathogens, improved resistance to bacterial pathogens, improved resistance to viral pathogens, improved resistance to Sporodia; improved resistance to powdery mildew; improved resistance to Fusarium; improved resistance to Sclerotinia; increased resistance to rust, improved resistance to Phytophthora, improved resistance to black leaf spot, improved resistance to Xanthomonas, increased resistance to necrotic fungi, increased resistance to biotrophic fungi, increased resistance to nematodes, increased resistance to insects, increased resistance to herbivores, and resistance to soft Enhanced animal performance, increased protein level, increased oil level, decreased lignin level, increased CBD level, increased anthocyanin level, decreased anthocyanin level, increased tissue nutrient level, increased tissue vitamin level, increased carbohydrate level, decreased carbohydrate level, increased starch level, increased sugar level, increased BRIX, increased protein level, decreased protein level, increased metabolite level, increased photosynthetic pigment level, increased lipid level, decreased lipid level, changed fatty acid saturation, increased saturated fat level, decreased saturated fat level, increased tocopherol level, decreased tocopherol level, increased isopentenol level, increased tissue nutrient content, improved processability, increased caloric value, decreased chlorine level, increased alkaloid level, decreased alkaloid level, increased wax level, decreased wax level, increased wax ester level, increased tannin level, increased paclitaxel level, increased lutein level, increased bioplastic level, increased biopolymer level, decreased biopolymer level, changed starch composition, increased latex level, increased rubber level;
[0319] The modified plants are grown under greenhouse or field conditions compared to control or reference plants not carrying the target genetic modification.
[0320] 2S. A method according to claim 2P, wherein the location of the uORF is first identified by executing a computer algorithm applied to ribosome profile data, whereby the algorithm identifies the presence of the uORF in a polynucleotide encoding a polypeptide or in a polynucleotide encoding a homolog having sequence similarity to the polypeptide based on the presence of ribosome enrichment in the interval from one stop codon to the next stop codon within the same open reading frame.
[0321] 2T. A method for killing plant cells, comprising contacting a portion of the plant with a nanoparticle preparation or a suspension containing cells of an Agrobacterium strain, wherein the cells of the Agrobacterium strain contain a nucleic acid construct comprising a gene editing system, wherein the gene editing system expresses a guide RNA in the cells of the plant, wherein the guide RNA introduces a mutation in a uORF upstream of a main ORF in the genome of the plant, wherein the main ORF encodes a necrosis-inducing polypeptide that triggers plant cell death.
[0322] 2U. The method of claim 2T, wherein the plant is a weed.
[0323] 2V. The method of claim 2U, wherein the weeds are selected from the group consisting of Arabidopsis thaliana, tall amaranth (Amaranthus tuberculatus), Johnson grass (Sorghum halepense), wild oats (Avena fatua), velvet (Abutilon theophrasti), amaranth (Amaranthus palmeri), redroot amaranth (Amaranthus retroflexus), poison sumac (Toxicodendron Vernix), Japanese knotweed (Polygonum cuspidatum), crabgrass (Digitaria), dandelion (Leontodon taraxacum), plantain (Plantago Major), ragweed (Ambrosia artemisiifolia), ragweed (Ambrosia trifida), hedge bindweed (Convolvus arvensis), ivy (Glechoma hederaceae), purslane (Portulaca olearacea), nettle (Urtica dioica), wrinkled sorrel (Rumex Cripus), wild madder (Galium mollugo), clover (Trifolium species)
[0324] 2W. The method of claim 2T, wherein the gene editing system is a CRISPR-CAS system.
[0325] 2X. The method of claim 2T, wherein the nucleic acid construct comprises a DNA sequence encoding a CAS enzyme.
[0326] 2Y. The method of claim 2T, wherein the major ORF encodes a polypeptide comprising SEQ ID NO: 5152, 5153, 5154, or 5155 (AT4G36900, AT2G23340, AT5G67190, or AT3G50260) or a polypeptide having sequence similarity to SEQ ID NO: 5152, 5153, 5154, or 5155.
[0327] 2Z. The method of claim 2T, wherein the major ORF encodes a sequence that is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 240%, 241%, 242%, 243%, 244%, 245%, 246%, 247%, 248%, 249%, 250%, 251%, 252%, 253%, 254%, 255%, 25 98%, or at least 99%, or about 100% identical to a polypeptide.
[0328] 2AA. The method of claim 2T, wherein the uORF comprises any one of SEQ ID NOs: 5211-5227, inclusive.
[0329] 2AB. A herbicidal composition comprising a nanoparticle formulation or a suspension containing cells of an Agrobacterium strain, wherein the composition contains a nucleic acid construct comprising a gene editing system, wherein the gene editing system expresses a guide RNA in cells of a target weed, wherein the guide RNA introduces a mutation in a uORF upstream of a main ORF in the genome of the weed, wherein the main ORF encodes a necrosis-inducing polypeptide that triggers plant cell death.
[0330] 2AC. The composition of claim 2AB, wherein the gene editing system is a CRISPR-CAS system.
[0331] 2AD. The composition of claim 2AB, wherein the nucleic acid construct comprises a DNA sequence encoding a CAS enzyme.
[0332] 2AE. The composition of claim 2AB, wherein the major ORF encodes a polypeptide comprising SEQ ID NO: 5152, 5153, 5154, or 5155 (AT4G36900, AT2G23340, AT5G67190, or AT3G50260) or a polypeptide encoding a homolog of SEQ ID NO: 5152, 5153, 5154, or 5155.
[0333] 2AF. The method of claim 2AB, wherein the major ORF encodes a sequence that is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 240%, 241%, 242%, 243%, 244%, 245%, 246%, 247%, 248%, 249%, 250%, 251%, 252%, 253%, 254%, 255%, 25 98%, or at least 99%, or about 100% identical to a polypeptide.
[0334] 2AG. The composition of claim 2AB, wherein the uORF comprises any one of SEQ ID NOs: 5211-5227, inclusive.
[0335] Example 3. Genetic modification of microorganisms and use of uORF mutations to enhance the production of target molecules and / or enzymes in cells cultured by fermentation.
[0336] 3A. A genetically engineered cell comprising a non-naturally occurring polynucleotide produced by gene editing, wherein the non-naturally occurring polynucleotide encodes a polypeptide that results in increased production levels of a target molecule or enzyme compared to a control microorganism not comprising the non-naturally occurring polynucleotide, and wherein the non-naturally occurring polynucleotide contains a mutation in a uORF located in the same transcript as a major ORF encoding the polypeptide.
[0337] 3B. The genetically modified cell of claim 3A, wherein the uORF is first identified by applying an algorithm to ribosome profiling data, wherein the algorithm identifies the presence of a uORF in a polynucleotide encoding a polypeptide or a polynucleotide encoding a homolog of a polypeptide based on the presence of ribosome enrichment in the interval from one stop codon to the next stop codon within the same open reading frame;
[0338] wherein the homologue has at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% similarity to the polypeptide.
[0339] 3B. The genetically modified cell of claim 3A, wherein the target molecule or enzyme is used for an application selected from the group consisting of: pollutant degradation, plastic degradation, oil degradation, laundry detergent; as a fragrance, as a flavoring, as a pigment, as a material, food processing, wood processing, an antibiotic, cancer treatment, diabetes treatment, heart disease treatment, hypertension treatment, obesity treatment, arthritis treatment, degenerative disease treatment, as a psychoactive substance, anxiety treatment, behavioral disorder treatment, as a food supplement, as a digestive aid, as a herbicide, as an insecticide, as a fungicide, as a rodenticide, as a bactericide, as a nematicide, as an algaecide, or as an antiviral agent.
[0340] 3D. The genetically modified cell of claim 3A, wherein the target molecule or enzyme is produced during a fermentation process.
[0341] 3F. The genetically modified cell of claim 3A, wherein the cell is a fungal cell, a bacterial cell, a mammalian cell, or a plant cell.
[0342] 3G. The genetically modified cell of claim 3A, wherein the molecule is encoded by a biosynthetic gene cluster (BGC), and wherein the uORF is located in the 5' region of a transcript (mRNA) encoding a transcriptional regulatory protein that regulates expression of genes in the biosynthetic gene cluster.
[0343] Example 4. Control of cancer cells
[0344] Typically, practitioners begin by selecting a published ribosome profile dataset or experimentally generating a new ribosome profile dataset by performing ribosome "pulldown" on mRNA samples purified from cancer tissue or cancer cell lines. When performing RNASeq, the pulled-down RNA is reverse transcribed and deeply sequenced (for example, using 50X coverage on an Illumina system). The type of algorithm detailed in this article is run on the sequence data to identify uORFs in stop-stop intervals. Genes containing uORFs are then selected, where the main ORF of the gene encodes a protein that promotes cell death or a cell cycle inhibitor to kill cancer cells.
[0345] In a more specific embodiment of the present invention, ribosome profiles are performed on tumor or cancer cell lines, and data are analyzed by applying the algorithm described in detail herein to identify the locus regulated by uORF. Select a locus with uORF and operatively connected to the main ORF, the main ORF encoding a polypeptide with tumor suppression, cell death or cell division inhibition function. The gene editing construct encodes the same guide RNA as the uORF of the selected locus, which is designed to knock out or mutate uORF. The gene editing construct is then delivered to the tumor or cancer cell in vivo by a delivery system (such as a viral vector). The destruction of uORF in cancer cells or tumor cells causes the translation of the main ORF to increase, thereby leading to control of the tumor or cancer cell.
[0346] Practitioners can apply the methods described herein to identify oncogenes controlled by uORFs by comparing ribosome pull-down data from cancer tissues or cell lines with ribosome pull-down data from control tissues. Typically, practitioners first select a published ribosome profiling dataset from a cancer cell line or experimentally generate a new ribosome profiling dataset by performing ribosome "pull-down" on mRNA samples purified from cancer tissues or cells and control non-cancerous cells. When performing RNASeq, the pulled-down RNA is reverse transcribed and deep sequenced (e.g., using 50X coverage on an Illumina system). The type of algorithm detailed herein is run on sequence data to identify uORFs. Genes containing an upstream uORF identified in control samples, where the uORF contains a mutation in the cancer tissue sample, can be considered candidate oncogenes. Additional evidence that the identified gene is likely an oncogene can be obtained by performing a BLAST comparison of the primary ORF product against public databases; if the primary ORF product shares homology with known cell cycle regulators, it is a strong candidate oncogene that may contribute to the cancerous nature of the cells in which it is active. In contrast, if in cancer samples, a new uORF is clearly located upstream of the main ORF and appears to have been generated by mutation, the generated uORF may suppress an anti-cancer gene compared to control samples.
[0347] 4A. A method for controlling cancer cells or tumor cells, comprising contacting the cancer cells or tumor cells with a delivery vector containing a nucleic acid construct, wherein the nucleic acid construct comprises a gene editing system, wherein the gene editing system expresses a guide RNA in the cell, wherein the guide RNA introduces a mutation in a uORF upstream of a main ORF in the genome of the cell, wherein the main ORF encodes a polypeptide that triggers cancer cell or tumor cell death or inhibits its cell division.
[0348] 4B. The method of claim 4A, wherein the uORF is first identified by applying an algorithm to ribosome profile data, wherein the algorithm identifies the presence of a uORF within a primary ORF encoding a polypeptide or a primary ORF encoding a polypeptide homolog based on the presence of ribosome enrichment in the interval from one stop codon to the next stop codon within the same open reading frame.
[0349] 4B. The method of claim 4A, wherein the major ORF comprises a cancer suppressor gene.
[0350] 4D. The method of claim 4A, wherein the major ORF encodes a polypeptide that inhibits cell division.
[0351] 4E. The method of claim 4A, wherein the delivery vector is a viral vector.
[0352] 4F. As discussed herein, in some cases, uORFs located upstream of the main ORF encode short peptides, or uPEPs, that act to directly or indirectly inhibit the activity of the main ORF. In cases where the uORF is located upstream of a main ORF that promotes cell division or promotes carcinogenesis through other mechanisms, the uPEP can be formulated and delivered as a drug (e.g., orally or intravenously) to inhibit the activity of oncogenes, thereby controlling cancer cells.
[0353] In such cases, uPEP can be synthesized or artificially synthesized by fermentation (eg, in yeast or E. coli), formulated (eg, to promote stability and / or cell entry), and delivered orally or intravenously to the patient to control the cancer.
[0354] Example 5. Control of eukaryotic pests and pathogens
[0355] The following examples have the advantage of using exogenously applied nucleic acids that are specific for the target pest or pathogen, which is superior to the use of chemical agents, which often act non-specifically with broad-spectrum toxicity and have deleterious effects on non-target organisms.
[0356] Typically, practitioners begin by selecting a published ribosome profile dataset or experimentally generating a new ribosome profile dataset by performing ribosome "pulldown" on mRNA samples purified from tissues or cells of the target pathogen. When performing RNASeq, the pulled-down RNA is reverse transcribed and deep sequenced (e.g., 50X coverage using an Illumina system). The type of algorithm detailed in this article is run on the sequence data to identify uORFs. Genes containing uORFs are then selected, where the main ORF of the gene encodes a protein that promotes cell death or a cell cycle inhibitor to control pests or pathogens.
[0357] 5A. A method for controlling a eukaryotic pest or pathogen, comprising contacting a cell of the eukaryotic pest or pathogen with a delivery vector containing a nucleic acid construct, wherein the nucleic acid construct comprises a gene editing system, wherein the gene editing system expresses a guide RNA in the cell, wherein the guide RNA introduces a mutation in a uORF upstream of a main ORF in the genome of the cell, wherein the main ORF encodes a polypeptide that triggers cell death or inhibits cell division of the eukaryotic pathogen.
[0358] 5B. The method of claim 5A, wherein the uORF is first identified by applying an algorithm to ribosome profile data, wherein the algorithm identifies the presence of a uORF in a primary ORF encoding a polypeptide or in a primary ORF encoding a gene having sequence similarity to the polypeptide based on the presence of ribosome enrichment in the interval from one stop codon to the next stop codon within the same open reading frame;
[0359] wherein the genes having sequence similarity are at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identical to the polypeptide.
[0360] 5C. The method of claim 5A, wherein the major ORF encodes a polypeptide that inhibits cell division.
[0361] 5D. The method of claim 5A, wherein the delivery vector is a viral vector or an antibody.
[0362] 5E. The method of claim 5A, wherein the pathogen is a fungus.
[0363] 5F. The method of claim 5A, wherein the pathogen is a protozoa.
[0364] 5G. The method of claim 5A, wherein the pathogen is a malarial cell.
[0365] 5H. The method of claim 5A, wherein the pathogen is a parasite.
[0366] 5I. If an essential gene (also known as a "lethal" gene) in a pest or pathogen is critical for the pest or pathogen to develop or complete its life cycle, and the essential gene has a uORF in its transcript, located upstream of the main ORF, then the uPEP encoded by the uORF can provide an effective pest or pathogen control agent by applying it as an exogenous pesticide. Ideally, an essential gene specific to a particular type of pest will be selected that is not present or non-essential in mammals. For example, if one is trying to control fungi or insects, genes involved in chitin biosynthesis are an example. For herbicides, uORFs are sought in plant-specific essential genes, such as key genes involved in amino acid synthesis, plant hormone production, meristem development, or photosynthesis.
[0367] In these cases, practitioners use the uORF sequence to heterologously produce the encoded uPEP in a fermentation system (e.g., E. coli, yeast, or cell line), formulate the resulting peptide to stabilize it and / or facilitate cell entry, and then apply it exogenously to the pest as a pesticide.
[0368] Example 6. Identification of uORFs in heterologous genes
[0369] Sequence similarity between mORFs from different species can be used to identify uORFs in heterologous genes. AT1G01060.1 (SEQ ID NO:4000; LHY) encodes a putative MYB-related transcription factor involved in circadian rhythms and was identified as a novel candidate uORF-containing gene. The LHY protein sequence from Arabidopsis thaliana was used to identify LHY orthologs in Brassica oleracea. The AT1G01060.1 sequence was then used to perform a BLAST (tblastn) search against the Brassica oleracea genome and a range of other species at genevolution.org / coge / CoGeBlast.pl. The results are shown in Table 2.
[0370] Table 3. Putative orthologs of LHY / CCA1 in Beta vulgaris, Eucalyptus grandis, Medicago truncatula, Brassica oleracea, and A5-DREB in Amaranthus hybridus, identified by BLAST analysis.
[0371]
[0372]
[0373] Example 7. Identification of uORFs in Cell Death-Inducing Genes from Example Weeds
[0374] The protein sequence of AT4G36900 from Arabidopsis thaliana was used to identify orthologous genes from Amaranthus hybridus, as shown in the bottom eight rows of the table above. At4g36900.1 encodes a member of the DREB subfamily A-5 of the ERF / AP2 transcription factor family (RAP2.10) and was identified as a candidate gene containing a uORF in high-throughput analysis.
[0375] The following sequences were used to perform sequence homology searches on the Amaranthus oleraceus genome using BLAST (website: genomevolution.org / coge / CoGeBlast.pl)
[0376] METATEVATVVSTPAVTVAAVATRKRDKPYKGIRMRKWGKWVAEIREPNKRSRIWLGSYSTPEAAARAYDTAVFYLRGPSARLNFPELLAGVTVTGGGGGGVNGG GDMSAAYIRRKAAEVGAQVDALEAAGAGGNRHHHHHQHQRGNHDYVDNHSDYRINDDLMECSKEGFKRCNGSLERVDLNKLPDPETSDDD(AT4G36900,SEQ ID NO:5152).
[0377] The output of this sequence search identified the closest gene with sequence homology to AT4G36900 (SEQ ID NO: 5152) as Ah.03g145670.m01-v1.1.a1.
[0378] Inspection of this locus in the genome using Jbrows revealed the coding sequence corresponding to the leader region and its adjacent upstream sequence.
[0379] Figure 7 Shown are Amaranthus hypochondriacus subsp. (hybrid) (hybrid contigs scaffolded to Amaranthus hypochondriacus): Modified genomic contigs of Amaranthus hypochondriacus scaffolded to the pseudochromosome of Amaranthus hypochondriacus, which is the finished (v1.0, id57429) version. The grey bars represent the putative AUG start codon, while the grey bars represent the stop codons of each of the three reading frames (note that the genes are in reverse order). In the frame, open reading frames that may be uORFs are defined by these stop-stop intervals. In this way, putative uORFs of a gene can be identified and tested as candidates for gene editing.
[0380] Example 8. Identification of a set of genes containing uORFs from Arabidopsis thaliana and subsequent identification of corresponding genes in target crops
[0381] To put the invention detailed herein into practice, codes 1 to 3 were applied to Arabidopsis thaliana ribosome profile data to identify a set of loci identified by the Arabidopsis thaliana genome identifiers that are candidate genes containing uORFs (SEQ ID NOs: 1 to 5155). These loci correspond to the respective loci in SEQ ID NOs: 1 to 5155. <223> Arabidopsis gene identifiers in the annotation lines. Note that in some cases, the locus is represented by a different gene model or multiple different gene models in the sequence listing, while in other cases a given model for a locus has multiple predicted uORFs. Since the presence of uORFs is generally evolutionarily conserved, these data provide a roadmap for generating improved traits by gene editing uORFs in orthologous loci of target crops. In one embodiment of the invention, polypeptide sequences encoded by loci known to produce improved traits of interest when the polypeptide is present at increased levels are compared to a panel of proteins derived from the target crop by applying BLAST or alignment analysis. Identify a crop plant locus that contains a major ORF that encodes a polypeptide that has at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity to the polypeptide used for comparison. Bioinformatics analysis is then performed on the crop genomic DNA region 1-1500 bp upstream of the start codon of the crop's major ORF to identify a stop-stop open reading frame upstream of the crop's major ORF. A polynucleotide construct is designed to encode a guide RNA that introduces a mutation in the identified upstream open reading frame. The guide RNA is delivered to cells of the target crop using a gene editing system, and crop plants expressing the improved trait of interest are regenerated and selected.
[0382] Example 9. Detection of uORFs in Selected Regulatory Genes in Different Plant Species
[0383] In this example, uORFs were detected in target crops and other plants, including sugar beet, Eucalyptus globulus, Brassica oleracea, and Amaranthus oleraceus. Candidate genes were identified by performing homology searches against the gene of interest in Arabidopsis thaliana. Candidate uORF sequences were then extracted upstream of the candidate gene.
[0384] Candidate uORFs are those with any start codon (sometimes a stop codon, sometimes a codon after the stop codon if there are multiple stop codons between this uORF and the previous uORF) and a length of more than 50 nucleotides. These extracted structures are shown in Table 4. Three examples are provided for the late elongating hypocotyl (LHY; encoding a putative MYB-related transcription factor involved in circadian rhythms) and a dehydration response element binding transcription factor (DREB; involved in regulating the expression of many stress-induced genes).
[0385] Table 4. uORFs detected in target plants
[0386]
[0387]
[0388]
[0389] Example 10. Achieving delayed flowering, increased yield, and / or increased biomass-related traits by targeting CCA1-related genes
[0390] In a further embodiment of the present invention, the crop homologs of the Arabidopsis thaliana circadian clock regulatory proteins LATE ELONGATED HYPOCOTYL gene (LHY / AT1G01060) and CIRCADIAN CLOCK ASSOCIATED 1 gene (CCA1 / AT2G46830) are upregulated by knocking out or mutating the operably linked uORF by gene editing or TILLING, wherein the proteins are shown to be regulated by the uORFs herein. Specifically, a genetic modification is introduced into a uORF within an endogenous locus containing a major ORF encoding a polypeptide, wherein the polypeptide or a region thereof is identical to (AT1G01060.1, SEQ ID NO: 4000) or (AT2G46830.1; SEQ ID NO: 4000). %, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% homogeneity. In some embodiments, the crop plants of the target genetic modification that contains introduction show at least 2% yield increase, or at least 3% yield increase, or at least 4% yield increase, or at least 6% yield increase, or at least 8% yield increase, or at least 10% yield increase, or at least 20% yield increase, or at least 50% yield increase.In the specific embodiment of this example, the crop plants of genetic modification are leaf crops or fodder crops, or the nutrient part of plant comprises the crop of required crop.For example, beet just belongs to latter class crop, and it needs large nutrient storage organ, and does not wish to bloom.In another embodiment of the present invention, the crop of genetic modification is tree crop, and its florescence is delayed, or never blooms before results.This is especially favourable for the transgenic trees (for example large eucalyptus and poplar) that are used for biomass for planting.
[0391] Example 11. Inducing flowering by targeting FCA-related genes
[0392] In a further embodiment of the present invention, the crop homolog of the Arabidopsis flowering regulator FCA (AT4G16280.2; SEQ ID NO: 4788) shown herein to be regulated by the uORF is upregulated by knocking out or mutating the operably linked uORF through gene editing. 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identical to SEQ ID NO: 4788. Such genomic modifications produce crop plants that exhibit early flowering in greenhouses, growth chambers, or fields. Under such conditions, crop plants containing the introduced target genetic modification develop floral structures at least 1 day earlier, 5 days earlier, 10 days earlier, 30 days earlier, 60 days earlier, or 180 days earlier than control plants that do not carry the genetic modification.
[0393] Example 12. Methods for controlling weeds by identifying and targeting uORFs of cell death-inducing genes in the control regions encoding AP2 family transcription factors.
[0394] In another embodiment of the present invention, ribosome profiling is performed on weedy plants and the data is analyzed by applying the algorithm detailed herein to identify loci regulated by uORFs. Loci are selected that have uORFs and operably linked major ORFs encoding polypeptides that are operably linked to the major ORFs encoding polypeptides or regions thereof that are operably linked to the major ORFs regulated by AT4G36900, AT2G23340, AT5G67190, or AT3G50260 (SEQ ID NOs: 1 and 2). 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity to the polypeptides encoded by proteins described herein. These proteins form a clade within the AP2 transcription factor family. The gene-editing construct is designed to encode a guide RNA that is identical to a uORF at a selected locus and is designed to knock out or mutate the uORF. The gene-editing construct is then delivered to the weeds via a suspension of Agrobacterium coated on nanoparticles or some other appropriate formulation. Disruption of the uORF in the target weed cells results in increased translation of homologs of AT4G36900, AT2G23340, AT5G67190, or AT3G50260, leading to cell death and control of the target weed.
[0395] Example 13. Increasing BRIX and / or sugar content by targeting uORFs from the bZIP family of transcription factors
[0396] In another embodiment of the present invention, a fruit or vegetable plant is selected and a locus is selected from the genome of the fruit or vegetable plant, wherein the locus has a uORF and an operably linked major ORF, wherein the major ORF encodes a polypeptide, and the polypeptide or a region thereof is identical to the polypeptide encoded by the bZIP protein AT4G34590.1, SEQ ID NO:4871) has at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity to the polypeptide encoded by the gene editing construct. The gene editing construct encodes a guide RNA identical to the uORF of the selected locus, which has a base change designed to knock out or mutate the uORF. The gene editing construct is then delivered to cells of a fruit or vegetable plant. Plants carrying the genetic modification are then regenerated such that uORF repression has been reduced or removed, and the fruit or vegetable plant has an increased sugar content or BRIX content compared to a control plant without the genetic modification when grown in a greenhouse, growth chamber, or field. In a specific embodiment of this example, the plant is a member of the Solanaceae family (such as tomato).
[0397] Example 14. Improving cold tolerance by targeting uORFs of myb family transcription factors
[0398] In another embodiment of the present invention, a crop is selected and a locus is selected from the genome of the crop, wherein the locus has a uORF and an operably linked major ORF, wherein the major ORF encodes a polypeptide, and the polypeptide or a region thereof is homologous to a polypeptide encoded by MYB protein AT1G74650 (SEQ ID NO:4274) has at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity to the polypeptide encoded by the gene editing construct. The gene editing construct encodes a guide RNA identical to the uORF of the selected locus, which has a base change designed to knock out or mutate the uORF. The gene-editing construct is then delivered into the cells of a crop plant. Plants are then regenerated that carry the genetic modification, resulting in reduced or eliminated uORF repression and improved cold tolerance compared to control plants that do not carry the genetic modification when grown in a greenhouse, growth chamber, or field.
[0399] Example 15. Increasing the nutritional content or pigmentation of plant tissues by targeting uORFs in HB family transcription factors required for flavonoid production.
[0400] In another embodiment of the present invention, a crop is selected and a locus is selected from the genome of the crop, wherein the locus has a uORF and an operably linked major ORF, wherein the major ORF encodes a polypeptide, and the polypeptide or a region thereof is homologous to the homeodomain protein ANTHOCYANINLESS2 AT4G00730 (SEQ ID NO:4731) has at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity to the polypeptide encoded by the gene editing construct. The gene editing construct encodes a guide RNA identical to the uORF of the selected locus, which is designed to knock out or mutate the uORF. The gene editing construct is then delivered into the cells of a crop plant. Plants are then regenerated that carry the genetic modification such that uORF repression is reduced or eliminated and the crop has increased pigment levels compared to control plants that do not carry the genetic modification when grown in a greenhouse, growth chamber, or field.
[0401] Example 16. Improving yield, vigor, seedling size, drought tolerance, protein content and / or tolerance to abiotic stress by targeting uORFs in the NF-Y (also known as HAP or CAAT) family of transcription factors.
[0402] The NF-Y family of transcription factors has been shown to regulate a range of key processes, including improved seedling vigor, flowering time, nutritional content and / or stress tolerance (US patents for plants with enhanced size and growth rate (Nelson et al., 2007, PNAS 104(42)16450-16455; Kumimoto et al., 2008, Planta 228,709-723; US Patent 8927811; US Patent 10640781)) through transgenic approaches including overexpression of native forms of the genes encoding these TFs. This example provides a method for obtaining the same or similar traits without the undesirable phenotypes (such as morphological abnormalities, extreme changes in flowering time and / or dwarfing) by alternative approaches to gene editing of endogenous loci encoding genes in crop plants and / or by overexpressing NF-Y genes with modifications or deletions of uORFs in the 5'UTR region of the overexpressed transcript.
[0403] In another embodiment of the present invention, a crop is selected and a locus is selected from the genome of the crop, wherein the locus has a uORF operably linked to a major ORF, wherein the major ORF encodes a polypeptide, wherein the polypeptide or a region thereof is homologous to the polypeptide encoded by the HAP2 protein AT5G12840 (SEQ ID NO:4961) has at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity. Gene editing constructs are established to encode a guide RNA that is identical to a uORF at a selected locus, which is designed to knock out or mutate the uORF. The gene editing construct is then delivered into cells of a crop plant. Plants carrying the genetic modification such that uORF repression is reduced or eliminated are then regenerated and selected for increased drought tolerance, increased seedling size, increased vigor, increased abiotic stress tolerance, and / or increased yield, but lacking any substantial adverse developmental phenotypes compared to control plants not carrying the genetic modification when grown in a greenhouse, growth chamber, or field.
[0404] In a further embodiment of the invention, the NF-YC transcription factor is a member of the NF-YC4 subclade that is orthologous to the Arabidopsis paralogs AT3G48590 and AT5G63470, which regulates beneficial traits including increased vigor, increased tolerance to abiotic stresses, increased nutrient content and enhanced tolerance to biotic stresses including viruses, bacteria, fungi, aphids and nematodes (U.S. Patent 10,640,781; Ling Li et al., PNAS, November 24, 2015, 112(47)14734-14739; Mingsheng Qi et al., Plant Biotechnology Journal (2019) 17, pp. 252-263). Select a crop plant and select a locus from the genome of the crop plant having a uORF operably linked to a major ORF encoding a polypeptide having at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64% similarity to the polypeptide encoded by AT3G48590 or AT5G63470. , 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity. Gene editing constructs are established to encode guide RNAs that are identical to the uORF at the selected locus, which are designed to knock out or mutate the uORF. The gene editing construct is then delivered to the cells of the crop plant. Plants carrying the genetic modification such that uORF repression is reduced or eliminated are then regenerated and selected for increased polypeptide levels, abiotic stress tolerance, increased yield, increased vigour, increased caloric content and / or increased nutrient content compared to control plants not carrying the genetic modification when grown in a greenhouse, growth chamber or field.
[0405] In another embodiment of the present invention, the uORF is mutated in the 5' region, including genetic modification of the endogenous loci encoding the crop homologs of NF-YC4 transcription factors AT3G48590 and AT5G63470. Plants produced by such genomic modifications exhibit increased seedling vigor and / or improved abiotic stress tolerance and / or improved photosynthesis and / or increased protein levels in tissues when grown under greenhouse conditions, field conditions and / or dehydration stress, heat stress and / or salt stress conditions. Under such conditions, crop plants containing such targeted introduced genetic modifications contain a non-native allele of a gene, the non-native allele of the gene is located upstream of the main ORF, and the mutant uORF encodes a polypeptide that has at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, or more similarity to the polypeptide encoded by AT3G48590 or AT5G63470. In some embodiments, the present invention provides a method for producing a crop plant having at least 99%, 100%, 100%, 200%, 200% or more protein content in a plant, animal or plant tissue, or a combination thereof. The method further comprises the step of: obtaining a crop plant having at least 99%, 100%, 100% or more protein content in a plant, animal or plant tissue, or a combination thereof. The method further comprises the step of: obtaining a crop plant having at least 99%, 100% or more protein content in a plant, animal or plant tissue, or a combination thereof. The method further comprises the step of: obtaining a crop plant having at least 99%, 100% or more protein content in a plant, animal or plant tissue, or a combination thereof.
[0406] In a further embodiment of the invention, the plant of choice is a soybean, corn, rice, potato, tomato or wheat plant.
[0407] In a further embodiment of the above invention, the crop plant is a soybean plant and the major ORF encodes a polypeptide that is at least 70% identical to or is identical to the following sequence:
[0408] METNNQQQQQQGAQAQSGPYPVAGAGGSAGAGAGAPPPFQHLLQQQQQQLQMFWSYQRQEIEHVNDFKNHQLPLARIKKIMKADEDVRMISAEAPILFAKACELFILELTIRSWL HAEENKRRTLQKNDIAAAITRTDIFDFLVDIVPRDEIKDDAALVGATASGVPYYYPPIGQPAGMMIGRPAVDPATGVYVQPPSQAWQSVWQSAAEDASYGTGGAGAQRSLDGQS*
[0409] In a further embodiment of the above invention, the crop plant is a maize plant, and the major ORF encodes a polypeptide that is at least 70% identical to the maize NF-YC4 sequence or is identical to the maize NF-YC4 sequence:
[0410] MDNQPLPYSTGQPPAPGGAPVAGMPGAAGLPPVPHHHLLQQQAQLQAFWAYQRQEAERASASDFKNHQLPLARIKKIMKADEDVRMISAEAPVLFAKACELFILELTIRSWLHAEENKRRTL QRNDVAAAIARTDVFDFLVDIVPREEAKEEPGSALGFAAPGTGVVGAGAPGGAPAAGMPYYYPPMGQPAPMMPAWHVPAWDPAWQQGAADVDQSGSFSEEGQGFGAGHGGAASFPPAPPTSE*
[0411] In a further embodiment of the above invention, the crop plant is a wheat plant, and the major ORF encodes a polypeptide having at least 70% identity to the wheat NF-YC4 sequence or being identical to the wheat NF-YC4 sequence:
[0412] MENHQLPYTTQPPATGAAGGAPVPGVPGPPPVPHHHLLQQQQAQLQAFWAYQRQEAERASASDFKNHQLPLARIKKIMKADEDVRMISAEAPVLFAKACELFILELTIRSWLHAEENKRRT LQRNDVAAAIARTDVFDFLVDIVPREEAKEEPGSAALGFAAGGVGAAGGGPAAGLPYYYPPMGQPAAPMMPAWHVPAWEPAWQQGGADVDQGAGSFGEEGQGYTGGHGGSAGFPPGPPSSD*
[0413] In a further embodiment of the above invention, the crop plant is a rice plant, and the major ORF encodes a polypeptide having at least 70% identity to the rice NF-YC4 sequence or being identical to the rice NF-YC4 sequence:
[0414] MDNQQLPYAGQPAAAGAGAPVPGVPGAGGPPAVPHHHLLQQQQAQLQ AFWAYQRQEAERASASDFKNHQLPLARIKKIMKADEDVRMISAEAPVLFAKACELFILELTIRSWLHAEENKRRTLQRNDVAAAIARTDVFDFLVDIVPREEAKEEPGSALGFAAGGPAGAVGAAGPAAGLPYYYPPMGQPAPMMPAWHVPAWDPAWQQGAAPDVDQGAAGSFSEEGQQGFAGHGGAAASFPPAPPSSE*
[0415] In a further embodiment of the above invention, the crop plant is a tomato plant and the major ORF encodes a polypeptide having at least 70% identity to the tomato NF-YC4 sequence or being identical to the tomato NF-YC4 sequence:
[0416] MDNQQLPYAGQPAAAGAGAPVPGVPGAGGPPAVPHHHLLQQQQAQLQAFWAYQRQEAERASASDFKNHQLPLARIKKIMKADEDVRMISAEAPVLFAKACELFILELTIRSWLHAEENKRRTL QRNDVAAAIARTDVFDFLVDIVPREEAKEEPGSALGFAAGGPAGAVGAAGPAAGLPYYYPPMGQPAPMMPAWHVPAWDPAWQQGAAPDVDQGAAGSFSEEGQQGFAGHGGAAASFPPAPPSSE*
[0417] In a further embodiment of the above invention, the crop plant is a potato plant and the major ORF encodes a polypeptide having at least 70% identity to the potato NF-YC4 sequence or is identical to the potato NF-YC4 sequence:
[0418] MDNNPHQSPTEAAAAAAAAAAAAQSATYPPQTPYHHLLQQQQQQLQMFWTYQRQEIEQVNDFKNHQLPLARIKKIMKADEDVRMISAEAPVLFAKACELFILELTIRSWLHAEENK RRTLQKNDIAAAITRTDIFDFLVDIVPRDEIKDEGVVLGPGIVGSTASGVPYYYPPMGQPAPGGVMLGRPAVPGVDPSMYVHPPPSQAWQSVWQTGDDNSYASGGSSGQGNLDGQI*
[0419] In a further embodiment of the above invention, the crop plant is a Gossypium plant, and the major ORF encodes a polypeptide that is at least 70% identical to the Gossypium NF-YC4 sequence or is identical to the Gossypium NF-YC4 sequence:
[0420] MDSNQQTQSTPYPPQPPTSAITPPSSATATAAPPFHHLLQQQQQQLQMFWSYQRQEIEQVNDFKNHQLPLARIKKIMKADEDVRMISAEAPILFAKACELFILELTIRSWLHAEENKR RTLQKNDIAAAITRTDIFDFLVDIVPRDEIKDETGLAPMVGATASGVPYFYPPMGQPAAGGPGGMMIGRPAVDPTGGIYGQPPSQAWQSVWQTAGTDDGSYGSGVTGGQGNLDGQG*
[0421] In a further embodiment of the above invention, the crop plants are plants grown for animal feed or silage, such as alfalfa, sorghum or forage grasses.
[0422] In a further embodiment of the above invention, the plant is a species grown as a source of protein for human consumption, such as peas, beans, beans or chickpeas.
[0423] In another embodiment of the present invention, NF-YC4 group transcription factors are expressed in transgenic plants, but this approach is improved by incorporating a gene encoding the NF-YC4 transcription factor into an expression construct that has a mutation or deletion within the 5'UTR or uORF sequence of the gene upstream of the main ORF encoding the NF-YC4 TF. It is worth noting that previous attempts to overexpress members of this TF family may have been hindered by the implementer inadvertently incorporating the uORF sequence upstream of the main ORF of the target gene being expressed. Therefore, when the transgenic transcript is produced, its translation is inhibited by the presence of the uORF. By intentionally omitting or mutating such uORF, in this example, such unintentional translational inhibition can be avoided and the desired phenotype can be enhanced.
[0424] To further illustrate the above example, researchers have reported phenotypes of transgenic soybean and maize lines expressing NF-YC4 subunits, but the resulting plants had seed protein contents of approximately 40% or less in soybean and approximately 120 mg / g dry weight or less in maize, based on the Lowry assay (O'Conner et al., Chapter 6. "From Arabidopsis to Crops: The Arabidopsis QQS Orphan Gene Modulates Nitrogen Allocation across species." In: Engineering Nitrogen Utilization in Crop Plants. Edited by Shrawat, Zayed, and Lightfoot. Springer 2018). Furthermore, these authors did not note any significant increase in size or vigor of transgenic plants compared to controls. Improved traits can be obtained by overexpressing variants of such transgenes (those lacking uORFs or having mutated uORFs in the 5' UTR). In a specific embodiment of this example, the improved trait in soybeans is a seed protein content greater than about 40% and / or the soybean plants exhibit a larger size than a control. In a further embodiment related to corn, the improved trait is a seed protein content greater than about 120 mg / g fresh weight and / or the corn plants exhibit a larger size than a control.
[0425] Example 17. Use of uORFs to optimize transgene activity in transgenic organisms.
[0426] For decades, there have been relatively few methods for generating transgenic organisms from many species, including plants, animals, and microorganisms. However, practitioners of these methods face a common challenge: optimizing the dosage of the transgenic product. Typically, transformation methods involve constructing a DNA construct containing a transgene or multiple copies of a transgene of interest regulated by a heterologous promoter. The construct is then introduced into cells of the target species, which can optionally be selected and regenerated into tissues or entire organisms. The promoter contained in the transgenic construct will produce higher levels of transgenic RNA in the transformed cells than in control cells, or produce tissue-specific or conditionally induced transgenic RNA expression patterns. However, it is usually impossible to precisely control the translation level of the resulting RNA.
[0427] uORF mutations can enhance the expression of transgenes or target genes at their native loci
[0428] In some cases, uORFs inhibit the translation of the gene's native transcript. In such cases, the implementer can upregulate the gene by mutating the uORF of the native locus or overexpressing the gene using a transgenic method. In such cases, if there is an unidentified uORF upstream of the main ORF in the transgenic transcript, translation may be inhibited and the transgenic cannot achieve the target trait. In such cases, the presence of uORFs in the 5' region of the transgenic can be identified using the methods described herein, and these uORFs can be intentionally ignored, or mutated by TILLING or gene editing (if the native locus is targeted) to weaken or remove uORF function and achieve translation of the transgenic product.
[0429] A specific example of using this approach is to optimize the activity of the REVOLUTA class of HD-ZIP III transcription factors, of which REV / IFL1 is the original member. At least five closely related members of this transcription factor clade are encoded by the Arabidopsis genome (locus identifiers: AT1G30490, AT4G32880, AT2G34710, AT5G60690, and AT1G52150). The activity of these genes and their encoded polypeptides can be upregulated by TILLING or gene editing to obtain alleles that produce high levels of protein, thereby conferring the trait of interest.
[0430] The REVOUTA (REV) branch of transcription factors plays a key role in regulating meristem behavior and development, including adaxial / abaxial patterning. When knocked out in the homozygous state, loss-of-function rev mutants in Arabidopsis exhibit abnormal shoot morphology, including a lack of interfascicular fibers in the stem, reduced growth of secondary shoot meristems, and elongated, distorted leaves. If one attempts to overexpress genes from this group and includes a native uORF upstream of the major ORF downstream of the transgene promoter in the transgenic construct, the resulting transformed plants will typically display a wild-type phenotype. However, if one intentionally omits the native uORF from the DNA clone or includes a weakened uORF variant with sequence changes compared to the native uORF, the resulting transformed plants will typically display one or more improved traits, which may include increased yield or increased biomass production.
[0431] Furthermore, one or more improved traits can be obtained by generating alleles through gene editing or TILLING that contain mutations that disrupt the native uORFs in one or more genes for REV-class transcription factors at their native loci in the plant genome. As a specific example, mutations in the uORFs of tomato genes encoding REV homologs can be mutated by gene editing or TILLING to produce one or more improved traits, which may include altered leaf shape, a more compact shoot system, and / or increased yield.
[0432] Introduction of uORF to repress expression
[0433] Conversely, in transgenics lacking a "strong" uORF, higher-than-optimal levels of translation may result in higher-than-necessary doses of the transgenic polypeptide, leading to undesirable side effects (or "abnormalities") in addition to the trait of interest. Such side effects may include dwarfism, slow growth, and developmental abnormalities such as tissue defects and / or organ malformations. In these cases, a uORF can be introduced into the 5' region of the transgene, upstream of the start codon of the main ORF, to provide a mechanism to inhibit translation of the encoded protein to more optimal levels ( Figure 8 Importantly, uORFs can be introduced during the initial design of a transgenic construct or after the event of selection for a transgenic line or organism that has integrated the transgene at certain specific sites in its genome and displays the desired phenotype but also possesses undesirable abnormalities.
[0434] The above methods can be used to improve or optimize existing transgenic crop events that have been publicly described and / or have been deregulated or submitted for deregulation by the U.S. Department of Agriculture's Animal and Plant Health Inspection Service, which oversees the release of transgenic crops in the United States. Crop events to which this method can be applied include ZmNF-YB drought-tolerant maize [Monsanto (now Bayer ) developed, see: Nelson et al. (2007). PNAS 104, No. 42, 16450-16455], BBX32 soybean [developed by Monsanto Company (now Bayer CropScience)], see: Preuss SB, Meister R, Xu Q, Urwin CP, Tripodi FA et al. (2012) Expression of the Arabidopsis thaliana BBX32 Gene in Soybean Increases Grain Yield. PLoS ONE 7(2):e30717. doi:10.1371 / journal.pone.0030717], ATHB17 corn [developed by Monsanto Company (now Bayer CropScience)], see Rice EA, Khandelwal A, Creelman RA, Griffith C, Ahrens JE et al. (2014) Expression of a Truncated ATHB17 Protein in Maize Increases Ear Weight at Silking. PLoS ONE 9(4):e94238.doi:10.1371 / journal.pone.0094238], ZMM28 maize [Corteva Agriscience Agriscience) development, see: Wu et al. (2019), PNAS Vol. 116, No. 47, 23851], and / or crops transformed with drought-tolerant genes HaHB4, ATHB13 or ATHB7 or their homologs, and / or crops transformed with stress-tolerant genes CBF1-4 or their homologs (CBF1 = AT4G2549, CBF2 = AT4G25470, CBF3 = AT4G25480, CBF4 = AT5G51990; BBX32 = AT3G21150; ATHB17 = AT2G01430; ATHB13 = AT1G69780; ATHB7 = AT2G46680).
[0435] In the above and similar situations, the practitioner selects a sequence that is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identical uORF sequence or create an artificial uORF, i.e., an open reading frame (comprising a stretch of nucleotides starting with a start codon and ending with a stop codon, with a total length of approximately 10-300 bp), located upstream of the main ORF (typically 10-500 bp upstream of the main ORF ATG, but longer or shorter distances may also be effective), and introduce the selected uORF into a transgenic construct (which is subsequently introduced into plant cells) or, in the case of an existing stable crop event being engineered, into the genome through gene editing. The practitioner then selects a plant from the resulting transformants or gene-edited lines that is improved in terms of the improved trait (as defined herein) compared to a control plant (which can be a wild-type plant or, if optimizing an existing transformed line, a plant of the original event). For example, BBX32 soybean lines can be selected from a population that exhibits increased yield without a delay in maturity. Similarly, gene-edited ATHB17 or ZMM28 maize lines that have been introduced with a uORF can be selected to exhibit greater yield improvements, or in the case of ZMM28, reduced delay from heat units to silking, respectively, compared to the original transgenic event, as described by Wu et al., supra. In the case of the ZmNF-YB2 transgenic maize event expressing this transcription factor, yield was significantly increased compared to controls in dry fields without irrigation, but when grown in fields with sufficient water, the event exhibited reduced yield compared to controls (so-called "yield drag"). In the examples presented here, the uORF can be introduced into the ZmNF-YB2 transgene (or a transgene encoding a homologous protein, including the transgene described by Nelson 2007, as described above) to obtain an increase in yield in dryland production while reducing or eliminating the yield drag observed under irrigated conditions.
[0436] Example 18. Generation of a "knockout" allele of a target gene by introducing a uORF through gene editing.
[0437] uORFs can also be used as tools to knock out or knock down target genes by introducing them into the 5' region of the main ORF of the target gene through gene editing.
[0438] A specific example involves the HY5-related transcription factor and its bZIP family homologs that promote photomorphogenesis. Preuss et al. (as mentioned above) and Khanna et al. reported that BBX32 inhibits light signaling by suppressing other BBX family proteins as well as HY5, which results in beneficial traits (such as increased root growth, increased pod number, and / or increased yield) in species such as soybean. Therefore, introducing uORFs into the 5' region of these genes encoding HY5 homologs through gene editing, particularly in soybean, may yield improved traits (such as the aforementioned phenotypes).
[0439] In the above and similar situations, the practitioner selects a sequence that is at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135 %, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identical uORF sequence or insert it into the genome by gene editing to create an artificial uORF, i.e., an open reading frame (comprising a stretch of nucleotides starting with a start codon and ending with a stop codon, with a total length of about 10-300 bp) located upstream of the main ORF. The practitioner then selects plants from the gene-edited lines that are improved in terms of the improved trait (as defined herein) compared to the control plants.
[0440] Example 19. Use of uPEP as a biostimulant
[0441] If a uORF is identified in the 5' transcript of a gene containing a major ORF that promotes a beneficial phenotype when its activity is reduced or knocked out, the uPEP encoded by the uORF can be exogenously applied as a biostimulant to achieve the desired trait. The desired trait may be increased drought tolerance, increased yield, or any of the improved traits detailed herein.
[0442] Example 19A. A method for obtaining an improved trait in a plant, comprising: first selecting a plant that has at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135 The invention relates to a method for producing a uPEP by a method of producing a uPEP in a cell or tissue. The method further comprises: obtaining a uORF sequence with at least 1%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity to a uORF sequence, and introducing the uORF into an expression vector that is capable of producing the encoded uPEP in a cell or tissue by a method such as fermentation. The expression vector is then introduced into the cell or tissue, and the produced uPEP is harvested, processed, formulated, and applied to plants as a biostimulant.
[0443] Example 19B. The method of Example 19, wherein the uORF is selected from a gene whose major ORF encodes a homolog of the bZIP protein HY5 gene (AT5G11260). In a specific embodiment of this example, the uORF is derived from the HY5 locus or a soybean homolog, and the resulting uPEP formulation is sprayed onto soybean plants to increase yield.
[0444] Example 19C. A method for inducing flowering in a crop plant, comprising identifying a uORF in the 5' region of a gene whose major ORF inhibits flowering, introducing the uORF into an expression vector capable of producing the encoded uPEP in a cell or tissue, such as by fermentation, introducing the expression vector into the cell or tissue, harvesting the uPEP produced from the cell or tissue, and applying a formulation containing the uPEP to a vegetatively growing plant.
[0445] The method of Example 19D.19C, wherein the major ORF encodes a homolog of the CCA1 gene (AT2G46830), the LHY gene (AT1G01060), the FLC gene (AT5G10140), or the TERMINAL FLOWER 1 gene (AT5G03840).
[0446] Example 19E. A method for inhibiting or delaying flowering or rendering crops sterile, comprising: identifying a uORF in the 5' region of a gene whose main ORF promotes flowering or floral organ development; introducing the uORF into an expression vector capable of producing the encoded uPEP in cells or tissues, such as by fermentation; introducing the expression vector into cells or tissues; harvesting the uPEP produced from the cells or tissues; and applying a formulation containing the uPEP to a vegetatively growing plant.
[0447] In a further embodiment of this example, practitioners apply one or more uPEP treatments to plants, thereby delaying the floral transition and enabling the plants to accumulate a greater amount of photosynthetic biomass and, therefore, achieve a higher yield after the uPEP treatment is terminated.
[0448] The method of Example 19F.19C, wherein the major ORF encodes a homolog of the CONSTANS gene (AT5G15840), the SOC1 gene (AT2G45660), the FLOWERING LOCUS T gene (AT1G65480), the LEAFY gene (AT5G61850), the FCA gene (AT4G16280), the GIGANTEA gene (AT1G22770), the PISTILLATA gene (AT5G20240), the APETALA3 gene (AT3G54340), the AGAMOUS gene (AT4G18960), the CAULIFLOWER gene (AT1G26310), or the APETALA 1 gene (AT1G69120).
[0449] The present invention is not limited to the specific embodiments described herein. Now that the invention has been fully described, it will be apparent to those skilled in the art that many changes and modifications can be made without departing from the spirit or scope of the claims. Modifications apparent from the foregoing description and accompanying drawings are within the scope of the following claims.
Claims
1. A method for identifying the presence of a putative upstream open reading frame (uORF) in a polynucleotide sequence by applying an algorithm to ribosome profiling data, the method steps comprising: a) identifying existing or putative main open reading frames (main ORFs) in an organism's genome and obtaining ribosome profiling data of mRNA transcribed from the organism's genome, wherein the identification can be de novo or based on existing knowledge; b) evaluating the ribosome occupancy of at least one region of the genome upstream of the main open reading frame; c) identifying genomic positions upstream of the main ORF where there is ribosome enrichment and a sharp drop in ribosome occupancy downstream of the enrichment; d) at or near the sharp drop in ribosome occupancy at the position, identifying a stop codon of a putative uORF of the genome; e) identifying a second stop codon that is upstream of and in the same reading frame as the stop codon of the putative uORF identified in d); and f) wherein the algorithm identifies the presence of a putative uORF in the genome in the interval between the first stop codon of the putative uORF and the second upstream stop codon in the same open reading frame.
2. The method according to claim 1, wherein a target genetic modification is introduced into the putative uORF to produce a modified uORF, and the translation of the main ORF operably linked to the modified uORF is increased.
3. The method according to claim 2, wherein the modified uORF has at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99%, or about 100% identity to the putative uORF.
4. The method according to claim 2, wherein the function of the modified uORF is reduced or knocked out by introducing at least one gene edit in the uORF.
5. The method according to claim 2, wherein the function of the modified uORF is reduced or knocked out in a plant, and the reduction or loss of the function of the uORF nucleotide sequence results in an increase in the translation of the main ORF operably linked to the uORF.
6. The method according to claim 5, wherein the increase in the translation of the main ORF results in cell death, inhibition of cell division, or a modified trait selected from the group consisting of: compared to a reference plant or control plant of the same species, Increased yield, improved flavor, improved texture, altered circadian rhythm, accelerated flowering, accelerated senescence, delayed senescence, increased branching, decreased branching, increased apical dominance, decreased apical dominance, increased shade tolerance, increased root mass, increased root hair number, increased nutritional quality, increased fruit mass, improved fruit quality, increased germination rate, increased trichome length, decreased trichome length, reduced spines, reduced spinules, spineless, altered leaf shape, increased leaf number, decreased leaf number, altered leaf angle, altered leaf position, altered branch angle, improved peelability, decreased cell adhesion, increased cell adhesion, reduced pericarp thickness, reduced exocarp thickness, increased exocarp thickness, increased seed coat thickness, reduced seed coat thickness, seedless, decreased seed size, increased seed size, apomixis, increased embryogenesis, increased sensitivity to genetic transformation, increased callus formation, increased embryo formation, increased root formation, increased cell division, decreased cell division, sterility, male sterility, pollen inactivity, lack of stamens, lack of carpels, increased number of carpels, increased number of petals, decreased number of petals, increased number of trichomes, decreased number of trichomes, increased stem width, decreased stem width, increased internode length, decreased internode length, increased floral organ size, altered floral organ shape, decreased floral organ size, reduced fruit abscission, reduced pod shattering, altered organ abscission, variegation, increased hypocotyl length, decreased hypocotyl length, delayed flowering, halted flowering, sterility, increased osmotic stress tolerance, increased photosynthesis, increased nitrogen use efficiency, increased phosphorus use efficiency, increased potassium use efficiency, increased nutrient use efficiency, increased nutrient uptake, increased metal ion uptake, increased heavy metal sequestration, increased oxidative stress tolerance, increased pigment level, increased salt tolerance, increased cold tolerance, increased frost tolerance, increased frost tolerance, increased dehydration stress tolerance, increased drought tolerance, increased post-drought recovery ability, reduced wilting, increased number of plastids, increased chlorophyll content, increased thylakoid density, increased photosynthetic capacity, increased respiration, decreased respiration, increased photorespiration, decreased photorespiration, increased transpiration, decreased transpiration, increased stomatal conductance, decreased stomatal conductance, increased carbon fixation, increased carbon sequestration, increased photosynthetic rate, increased carotenoid level, decreased carotenoid level, increased electron transport, improved non-photochemical quenching, increased ion transport, decreased ion transport, altered carbon-nitrogen balance, increased sensitivity to hormones, decreased sensitivity to hormones, decreased sensitivity to ethylene, increased auxin level, decreased auxin level, increased auxin transport, increased auxin sensitivity, decreased auxin sensitivity, increased gibberellin level, decreased gibberellin level, increased gibberellin sensitivity, increased abscisic acid level, decreased abscisic acid level, increased abscisic acid sensitivity, decreased abscisic acid sensitivity, increased cytokinin level, decreased cytokinin level, increased cytokinin level, decreased cytokinin level, increased cytokinin sensitivity, decreased cytokinin sensitivity, increased jasmonic acid level, decreased jasmonic acid level,Increased jasmonic acid sensitivity, decreased jasmonic acid sensitivity, increased salicylic acid level, decreased salicylic acid level, decreased salicylic acid sensitivity, increased salicylic acid sensitivity, increased strigolactone level, decreased strigolactone level, increased sensitivity to strigolactone, decreased sensitivity to strigolactone, decreased sensitivity to ethylene, increased sensitivity to ethylene, accelerated ripening, delayed ripening, reduced fruit spoilage, extended shelf life, improved heat stress, improved tolerance to low nitrogen conditions, increased seedling growth vigor, increased disease resistance, increased resistance to fungal pathogens, increased resistance to bacterial pathogens, increased resistance to viral pathogens, increased resistance to Colletotrichum; increased resistance to Erysiphe; increased resistance to Fusarium; increased resistance to Sclerotinia; enhanced rust resistance, increased resistance to Phytophthora, increased resistance to black leaf streak, increased resistance to Xanthomonas, increased resistance to necrotrophic fungi, increased resistance to biotrophic fungi, enhanced nematode resistance, enhanced insect resistance, enhanced herbivore resistance, enhanced mollusk resistance, increased protein level, increased oil level, decreased lignin level, increased CBD level, increased anthocyanin level, decreased anthocyanin level, increased tissue nutrient level, increased tissue vitamin level, increased carbohydrate level, decreased carbohydrate level, increased starch level, increased sugar level, increased BRIX, increased protein level, decreased protein level, increased metabolite level, increased photosynthetic pigment level, increased lipid level, decreased lipid level, altered fatty acid saturation, increased saturated fat level, decreased saturated fat level, increased tocopherol level, decreased tocopherol level, increased isoprenoid level, increased tissue nutrient content, improved processability, increased calorific value, decreased chlorine level, increased alkaloid level, decreased alkaloid level, increased wax level, decreased wax level, increased wax ester level, increased tannin level, increased taxol level, increased lutein level, increased bioplastic level, increased biopolymer level, decreased biopolymer level, altered starch composition, increased latex level, increased rubber level.
7. The method according to claim 5, wherein the uORF regulates the translation of the main ORF, and compared to a reference or control plant of the same species, the increase in the translation of the main ORF results in a toxic effect or cell death or early flowering, delayed flowering or bolting.
8. A plant or plant cell comprising: A target genetic modification introduced at a locus in a native genome, wherein the introduced target genetic modification comprises a mutation in a uORF, wherein the locus in the native genome comprises a main ORF operably linked to the uORF, wherein the main ORF encodes a polypeptide having regulatory activity, and wherein the polypeptide comprises an amino acid sequence having a percent identity with a polypeptide selected from the group consisting of SEQ ID NO: 3999–5155; wherein the percent identity is at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100%; and the target genetic modification of the uORF can increase the expression level and / or activity of the encoded polypeptide having regulatory activity.
9. The plant or plant cell according to claim 8, wherein the introduced target genetic modification of the uORF results in an increase in the translation of the main ORF operably linked to the uORF.
10. The plant or plant cell according to claim 8, wherein the uORF with the target genetic modification has at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, or about 100% identity with the uORF within the locus of the native genome.
11. The plant or plant cell according to claim 8, wherein the uORF comprising the introduced target genetic modification comprises at least one gene editing.
12. The plant or plant cell according to claim 8, wherein the uORF with the introduced target genetic modification has reduced function or is knocked out in the plant, and the modification of the uORF results in an increase in the translation of the main ORF of the encoded polypeptide operably linked to the uORF.
13. The plant or plant cell according to claim 8, wherein the main ORF encodes a polypeptide, and the expression of the polypeptide results in cell death, cell division inhibition, or a modified trait selected from the group consisting of: compared with a reference plant or a control plant of the same species, Increased yield, improved flavor, improved texture, altered circadian rhythm, accelerated flowering, accelerated senescence, delayed senescence, increased branching, decreased branching, increased apical dominance, decreased apical dominance, increased shade tolerance, increased root mass, increased root hair number, increased nutritional quality, increased fruit mass, improved fruit quality, increased germination rate, increased trichome length, decreased trichome length, reduced spines, reduced spinules, spineless, altered leaf shape, increased leaf number, decreased leaf number, altered leaf angle, altered leaf position, altered branch angle, improved peelability, decreased cell adhesion, increased cell adhesion, reduced pericarp thickness, reduced exocarp thickness, increased exocarp thickness, increased seed coat thickness, reduced seed coat thickness, seedless, reduced seed size, increased seed size, apomixis, increased embryogenesis, increased susceptibility to genetic transformation, increased callus formation, increased embryo formation, increased root formation, increased cell division, decreased cell division, sterility, male sterility, pollen inactivity, lack of stamens, lack of carpels, increased number of carpels, increased number of petals, decreased number of petals, increased number of trichomes, decreased number of trichomes, increased stem width, decreased stem width, increased internode length, decreased internode length, increased floral organ size, altered floral organ shape, decreased floral organ size, reduced fruit abscission, reduced pod shattering, altered organ abscission, variegation, increased hypocotyl length, decreased hypocotyl length, delayed flowering, ceased flowering, sterility, increased osmotic stress tolerance, increased photosynthesis, increased nitrogen use efficiency, increased phosphorus use efficiency, increased potassium use efficiency, increased nutrient use efficiency, increased nutrient uptake, increased metal ion uptake, increased heavy metal sequestration, increased oxidative stress tolerance, increased pigment level, increased salt tolerance, increased cold tolerance, increased frost tolerance, increased frost tolerance, increased dehydration stress tolerance, increased drought tolerance, increased post-drought recovery ability, reduced wilting, increased number of plastids, increased chlorophyll content, increased thylakoid density, increased photosynthetic capacity, increased respiration, decreased respiration, increased photorespiration, decreased photorespiration, increased transpiration, decreased transpiration, increased stomatal conductance, decreased stomatal conductance, increased carbon fixation, increased carbon sequestration, increased photosynthetic rate, increased carotenoid level, decreased carotenoid level, increased electron transport, improved non-photochemical quenching, increased ion transport, decreased ion transport, altered carbon-nitrogen balance, increased sensitivity to hormones, decreased sensitivity to hormones, decreased sensitivity to ethylene, increased auxin level, decreased auxin level, increased auxin transport, increased auxin sensitivity, decreased auxin sensitivity, increased gibberellin level, decreased gibberellin level, increased gibberellin sensitivity, increased abscisic acid level, decreased abscisic acid level, increased abscisic acid sensitivity, decreased abscisic acid sensitivity, increased cytokinin level, decreased cytokinin level, increased cytokinin level, decreased cytokinin level, increased cytokinin sensitivity, decreased cytokinin sensitivity, increased jasmonic acid level, decreased jasmonic acid levelIncreased jasmonic acid sensitivity, decreased jasmonic acid sensitivity, increased salicylic acid level, decreased salicylic acid level, decreased salicylic acid sensitivity, increased salicylic acid sensitivity, increased strigolactone level, decreased strigolactone level, increased sensitivity to strigolactone, decreased sensitivity to strigolactone, decreased sensitivity to ethylene, increased sensitivity to ethylene, accelerated ripening, delayed ripening, reduced fruit spoilage, extended shelf life, improved heat stress, improved tolerance to low nitrogen conditions, increased seedling growth vigor, increased disease resistance, increased resistance to fungal pathogens, increased resistance to bacterial pathogens, increased resistance to viral pathogens, increased resistance to Colletotrichum; increased resistance to Erysiphe; increased resistance to Fusarium; increased resistance to Sclerotinia; enhanced rust resistance, increased resistance to Phytophthora, increased resistance to Bipolaris leaf spot, increased resistance to Xanthomonas, increased resistance to necrotrophic fungi, increased resistance to biotrophic fungi, enhanced nematode resistance, enhanced insect resistance, enhanced herbivore resistance, enhanced mollusk resistance, increased protein level, increased oil level, decreased lignin level, increased CBD level, increased anthocyanin level, decreased anthocyanin level, increased tissue nutrient level, increased tissue vitamin level, increased carbohydrate level, decreased carbohydrate level, increased starch level, increased sugar level, increased BRIX, increased protein level, decreased protein level, increased metabolite level, increased photosynthetic pigment level, increased lipid level, decreased lipid level, altered fatty acid saturation, increased saturated fat level, decreased saturated fat level, increased tocopherol level, decreased tocopherol level, increased isoprenoid level, increased tissue nutrient content, improved processability, increased calorific value, decreased chlorine level, increased alkaloid level, decreased alkaloid level, increased wax level, decreased wax level, increased wax ester level, increased tannin level, increased taxol level, increased lutein level, increased bioplastic level, increased biopolymer level, decreased biopolymer level, altered starch composition, increased latex level, increased rubber level.
14. The method according to claim 12, wherein the introduced target genetic modification results in an increase in the translation of the main ORF, and compared with a reference or control plant of the same species, the increase in the translation of the main ORF results in a toxic effect or cell death or an earlier flowering time, a delayed flowering time, or bolting.
15. A plant whose genome contains a non-naturally occurring allele of a gene, the gene containing a mutation in the uORF upstream of the main ORF, the protein encoded by the main ORF being a homolog of the protein encoded by CCA1 (SEQ ID NO: 4483; Arabidopsis locus AT2G46830) or having at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% or about 100% identity with the protein encoded by CCA1 (SEQ ID NO: 4483; Arabidopsis locus AT2G46830), and wherein the plant is selected based on the improved trait.
16. The plant according to claim 15, wherein the improved trait is a delayed flowering time, enhanced photosynthesis, increased vegetative biomass, a more compact shoot architecture, or increased yield.
17. A plant whose genome contains a non-naturally occurring allele of a gene, the gene containing a mutation in the uORF upstream of the main ORF, the protein encoded by the main ORF being a homolog of the protein encoded by REVOLUTA (SEQ ID NO: 5112; Arabidopsis locus AT5G60690) or having at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% or about 100% identity with the protein encoded by REVOLUTA (SEQ ID NO: 5112; Arabidopsis locus AT5G60690), and wherein the plant is selected based on the improved trait.
18. The plant according to claim 17, wherein the improved trait is a change in leaf shape, increased vegetative biomass, a more compact shoot architecture, or increased yield.
19. A plant whose genome contains a non-naturally occurring allele of a gene, said gene comprising a mutation in a uORF upstream of a main ORF, and the protein encoded by the main ORF is a homolog of the protein encoded by Aphaninless 2 (SEQ ID NO: 4731; Arabidopsis locus AT4G00730) or has at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% or about 100% identity to the protein encoded by Aphaninless 2 (SEQ ID NO: 4731; Arabidopsis locus AT4G00730), and wherein the plant is selected based on the improved trait.
20. The plant according to claim 19, wherein the improved trait is an increase in the nutrient content of the plant tissue or an increase in the flavonoid content, or a darker color.
21. A plant whose genome contains a non-naturally occurring allele of a gene, said gene comprising a mutation in a uORF upstream of a main ORF, and the protein encoded by the main ORF is a homolog of the protein encoded by FCA (SEQ ID NO: 4788; Arabidopsis locus AT4G16280) or has at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% or about 100% identity to the protein encoded by FCA (SEQ ID NO: 4788; Arabidopsis locus AT4G16280), and wherein the plant is selected based on the improved trait.
22. The plant according to claim 21, wherein the improved trait is accelerated flowering.
23. A plant whose genome contains a non-naturally occurring allele of a gene, said gene comprising a mutation in a uORF upstream of a main ORF, and the protein encoded by the main ORF is a homolog of the protein encoded by the HAP2 protein encoded by an Arabidopsis locus identified by AT1G72830 (SEQ ID NO: 4266) or AT5G12840 (SEQ ID NO: 4961) or has at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% or about 100% identity to the HAP2 protein encoded by an Arabidopsis locus identified by AT1G72830 (SEQ ID NO: 4266) or AT5G12840 (SEQ ID NO: 4961), and wherein said plant is selected based on an improved trait.
24. The plant according to claim 23, wherein the improved trait is increased tolerance to abiotic stress, increased water use efficiency, or increased tolerance to dehydration.
25. A plant whose genome contains a non-naturally occurring allele of a gene, said gene comprising a mutation in a uORF upstream of a main ORF, and the protein encoded by the main ORF is a homolog of the bZIP protein encoded by AT4G34590.1 (SEQ ID NO: 4871) or has at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% or about 100% identity to the bZIP protein encoded by AT4G34590.1 (SEQ ID NO: 4871), and wherein said plant is selected based on an improved trait.
26. The plant according to claim 25, wherein the improved trait is an increase in the sugar content or BRIX of plant tissues.
27. A plant whose genome contains a non-naturally occurring allele of a gene, said gene comprising a mutation in a uORF upstream of a main ORF, and the polypeptide encoded by the main ORF is a homolog of the protein encoded by the Arabidopsis locus identified by AT5G63470 or AT3G48590 or has at least 30% or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% or about 100% identity, and wherein the plant is selected based on the improved trait.
28. The plant according to claim 27, wherein the polypeptide is the NF-YC4 subunit and the plant is a rice plant.
29. The plant according to claim 28, wherein the polypeptide is the NF-YC4 subunit and the plant is a soybean plant.
30. The plant according to claim 27, 28 or 29, wherein the improved trait is increased protein content, or increased plant size, or increased drought tolerance, or increased abiotic stress tolerance.
Citation Information
Patent Citations
Modification of transcriptional repressor binding site in NF-YC4 promoter for increased protein content and resistance to stress
US10640781B2
Type 1 and type 2 hopping for device-to-device communications
US20160020822A1
Methods for the identification of variant recognition sites for rare-cutting engineered double-strand-break-inducing agents and compositions and uses thereof
US20160032297A1
Polynucleotides and polypeptides in plants
US7345217B2
Polynucleotides and polypeptides in plants
US7511190B2
Cited By
Cream lettuce plant type regulation gene as well as cloning method and application thereof
CN117904147A
Application and method of AtIDD7 gene in regulation and control of plant growth and development
CN119955850A
Application and method of atidd7 gene in regulating plant growth and development
CN119955850B
Panax notoginseng bHLH transcription factor gene PnbHLH2 and application thereof
CN120366334A
Panax notoginseng bHLH transcription factor gene PnbHLH2 and application thereof
CN120366334B