Methods and compositions for reduced stature plants
Targeted genetic modifications in the dwarf8 gene using CRISPR technology induce translation reinitiation, creating a semi-dwarf maize phenotype that enhances yield and resilience, addressing limitations in existing stature modification methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-02
- Publication Date
- 2026-04-09
AI Technical Summary
Existing methods for modifying plant stature, particularly in maize, are limited in effectively achieving a semi-dwarf phenotype that enhances yield and lodging resistance without compromising agronomic performance.
Introduce targeted genetic modifications, specifically premature stop codons in the dwarf8 gene using CRISPR-associated nucleases, to alter the D8 protein expression through translation reinitiation, resulting in a semi-dwarf phenotype with reduced plant height and improved agronomic traits.
The modified maize plants exhibit enhanced resilience to weather conditions, increased planting density, improved yield, and reduced input costs through precise genetic manipulation of the dwarf8 gene, leading to a semi-dwarf phenotype.
Smart Images

Figure US2025049146_09042026_PF_FP_ABST
Abstract
Description
METHODS AND COMPOSITIONS FOR REDUCED STATURE PLANTSFIELD
[0001] This disclosure relates to compositions and methods of modifying stature in plants, including height reduction and other agronomic characteristics.REFERENCE TO A SEQUENCE LISTING SUBMITTTED ELECTRONICALLY
[0002] The official copy of the sequence listing is submitted electronically via Patent Center as an XML formatted sequence listing with a file named 212664 created on September 30, 2025, and having a size of 110,164 bytes and filed concurrently with the specification. The sequence listing comprised in this XML formatted document is part of the specification and is herein incorporated by reference in its entirety.
[0003] The sequence descriptions (Table 1) and sequence listing attached hereto comply with the rules governing nucleotide and amino acid sequence disclosures in patent applications as set forth in 37 C.F.R. §§1.831-1.835.BACKGROUND
[0004] Recent advances in plant genome editing have opened new doors to generating plants with improved characteristics or traits, such as stature, height and other architecture. Plant height is a desirable trait in crop breeding for a variety of crops of commercial interest. Dwarf stature has been used to improve yield and lodging resistance in crop plants, e.g., use of dwarf mutants in wheat and rice that increased harvest index. Height adaptations increase harvest index, favorably partition carbon and nutrients between grain and none-grain biomass, enhance fertilizer use, water use efficiency and play a role in increasing planting density.SUMMARY
[0005] The disclosure provides plants of a semi-dwarf phenotype and methods of generating such plants that have a semi-dwarf phenotype compared to control or wildtype plants. These plants and methods are based on the discovery that the introduction or creation of in-frame or out-of-frame stop codons in a dwarf gene produced a modified polypeptide and a semi-dwarf phenotype in maize plants.
[0006] Provided herein is a genome-edited polynucleotide encoding a polypeptide comprising an amino acid sequence of SEQ ID NO: 59, where the polypeptide sequence does not comprise a functional DELLA motif of SEQ ID NO: 18. In some aspects, the polynucleotide sequence is at least 95% identical to a polynucleotide sequence of SEQ ID NO: 7 or SEQ ID NO: 58.
[0007] Also provided is a genome-edited plant, plant cell, or seed thereof comprising a genome- edited polynucleotide encoding a polypeptide comprising an amino acid sequence of SEQ ID NO: 59, where the polynucleotide may be at least 95% identical to SEQ ID NO: 7 or SEQ ID NO: 58. In some aspects, a polynucleotide that comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 variations, or changes to the polynucleotide sequences set forth in SEQ ID NO: 7 or 58 are part of the inventive concept. In some aspects, the genome-edited plant, plant cell, or seed thereof is a maize plant, plant cell, or seed. In some aspects, the genome-edited plant, plant cell, or seed thereof produces an mRNA transcript that results in generation of a translation reinitiated polypeptide. In further aspects, the genome-edited plant exhibits reduced plant height compared to a control plant not comprising the targeted genome-edited modification that results in the translation reinitiated polypeptide. In further aspects, the plant is a maize inbred plant and exhibits a semidwarf phenotype, when the genome edit is in a heterozygous state. In other aspects, the genome- edited seed is from a genome-edited plant that exhibits a semi-dwarf phenotype. In some aspects, a maize plant cell, seed, or a plant that includes a genome-edited polynucleotide that comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 variations, or changes to the polynucleotide sequences set forth in SEQ ID NO: 7 or 58 are part of the inventive concept.
[0008] Provided is a method of reducing height of a maize plant comprising introducing a targeted modification into an endogenous dwarf8 gene in a maize plant, plant cell or seed thereof, wherein the targeted modification introduces or creates a premature stop codon between a first or original starting ATG codon of the open reading frame of the dwarf8 gene and a subsequent downstream ATG codon. In further aspects, the downstream ATG codon is located upstream or within the 5’ region of a GRAS domain coding sequence, the first or original starting ATG codon is in-frame with the subsequent downstream ATG codon, and the premature stop codon is in-frame or out-of-frame with the downstream ATG codon. The method results in the generation of a modified plant that exhibits reduced height relative to a control plant not comprising the targeted modification, and the reduced height phenotype results from translation reinitiation at the downstream ATG codon.
[0009] In some aspects, prior to modification the dwarfs gene encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6. In further aspects, the targeted modification is introduced using a genomic modification technique comprising a CRISPR-associated nuclease and at least one guide RNA and the targeted modification is an addition, deletion, or substitution of at least one nucleotide that introduces or creates a premature stop codon. In some aspects, the targeted modification results in translation reinitiation from the downstream ATG codon, thereby generating a truncated dwarf8 protein. In additional examples, the premature stop codon is introduced between the first methionine (Ml) of SEQ ID NO: 6 and a second methionine corresponding to methionine at position 53 (M53), methionine at position 65 (M65), methionine at position 67 (M67), methionine at position 69 (M69), methionine at position 106 (Ml 06), methionine at position 184 (M184), methionine at position 201 (M201), methionine at position 271 (M271), or methionine at position 280 (M280) of SEQ ID NO: 6, wherein the second methionine is in-frame with the first methionine and the premature stop codon is in-frame or out-of-frame with the second methionine. For example, the targeted modification comprises an in-frame or out-of-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO: 48, and in certain aspects, the resulting modified maize plant, plant cell or seed thereof comprises SEQ ID NO: 7 or SEQ ID NO: 58. In additional aspects, the method further comprises regenerating a maize plant from the modified plant cell thereby producing a modified plant, wherein the modified plant exhibits a semi-dwarf phenotype due to translation reinitiation at the downstream ATG codon.
[0010] Provided herein is a modified maize plant, plant cell, or seed comprising a targeted genome-edited modification at an endogenous dwarf8 gene in a maize plant, plant cell or seed thereof, wherein the targeted modification introduces or creates a premature stop codon between a first or original starting ATG codon of the open reading frame of the dwarf8 gene and a subsequent downstream ATG codon. In further aspects, the downstream ATG codon is located upstream or within the 5’ region of a GRAS domain coding sequence, the translation initiation ATG codon is in-frame with the subsequent downstream ATG codon, and the premature stop codon is in-frame or out-of-frame with the downstream ATG codon. The method results in the generation of a modified plant that exhibits reduced height relative to a control plant not comprising the targeted genome-edited modification.
[0011] In some examples, the modified maize plant, or plant cell, or seed thereof, comprises a targeted modification that introduces or creates an in-frame or out-of-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48, and in certain aspects, the resulting modified maize plant, plant cell or seed thereof comprises SEQ ID NO:7 or SEQ ID NO: 58. In further aspects, the modified maize plant exhibits a semi -dwarf phenotype due to translation reinitiation at the downstream ATG codon.
[0012] Also provided is a guide polynucleotide molecule that targets an endogenous gene of a plant cell (e.g. maize), wherein the gene comprises a polynucleotide that encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6. A plant cell that includes the guide polynucleotide is provided wherein the guide polynucleotide interacts with a Cas endonuclease at the endogenous gene. In some aspects, the plant cell is a maize plant cell.
[0013] Provided is a plant cell (e.g. maize) including a targeted modification at an endogenous gene that encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6, wherein the targeted genetic modification comprises a premature stop codon between a first or original starting ATG codon of the open reading frame of the endogenous gene and a subsequent downstream ATG codon. In further aspects, the downstream ATG codon is located upstream or within the 5’ region of a GRAS domain coding sequence, the first or original starting ATG codon is in-frame with the subsequent downstream ATG codon; and the premature stop codon is in-frame or out-of-frame with the downstream ATG codon. In particular aspects, the plant cell comprises the targeted modification comprising an addition, deletion, or substitution of at least one nucleotide that introduces or creates the premature stop codon. In further aspects, the targeted modification results in translation reinitiation from the downstream ATG codon, thereby generating a truncated dwarf8 protein.
[0014] In additional examples, the premature stop codon is introduced between a first methionine, Ml, and a second methionine corresponding to methionine at position M53, methionine at position M65, methionine at position M67, methionine at position M69, methionine at position M106, methionine at position M184, methionine at position M201, methionine at position M271, or methionine at position M280 of SEQ ID NO: 6. The second methionine is in-frame with the first methionine and the premature stop codon is in-frame or out-of-frame with the second methionine. For example, the maize plant cell comprises the targeted modification comprising an in-frame or out-of-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO: 48, and in certain aspects, the resulting maize plant cell comprises SEQ ID NO: 7 or SEQ ID NO: 58.
[0015] Also provided is a plant (e.g. maize) comprising a targeted modification at an endogenous gene that encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6, wherein the targeted genetic modification comprises a premature stop codon between a first or original starting ATG codon of the open reading frame of the endogenous gene and a subsequent downstream ATG codon. In further aspects, the downstream ATG codon is located upstream or within the 5’ region of a GRAS domain coding sequence, the first or original starting ATG codon is in-frame with the subsequent downstream ATG codon; and the premature stop codon is in-frame or out-of-frame with the downstream ATG codon. The targeted modification results in the plant having reduced plant height compared to a control plant not comprising the targeted modification, and wherein the reduced height phenotype results from translation reinitiation at the downstream ATG codon.
[0016] In particular aspects, the plant comprises the targeted modification comprising an addition, deletion, or substitution of at least one nucleotide that introduces or creates the premature stop codon. In further aspects, the targeted modification results in translation reinitiation from the downstream ATG codon, thereby generating a truncated dwarf8 protein.
[0017] In additional examples, the premature stop codon is introduced between a first methionine, Ml, and a second methionine corresponding to methionine at position M53, methionine at position M65, methionine at position M67, methionine at position M69, methionine at position M106, methionine at position Ml 84, methionine at position M201, methionine at position M271, or methionine at position M280 of SEQ ID NO: 6. The second methionine is in-frame with the first methionine and the premature stop codon is in-frame or out- of-frame with the second methionine. For example, the maize plant comprises the targeted modification comprising an in-frame or out-of-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO: 48, and in certain aspects, the resulting maize plant comprises SEQ ID NO: 7 or SEQ ID NO: 58.
[0018] Provided is a seed comprising a targeted modification at an endogenous gene that encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%,98%, or 99% identical to SEQ ID NO: 6, wherein the targeted modification comprises a premature stop codon between a first or original starting ATG codon of the open reading frame of the endogenous gene and a subsequent downstream ATG codon. In further aspects, the downstream ATG codon is located upstream or within the 5’ region of a GRAS domain coding sequence, the first or original starting ATG codon is in-frame with the subsequent downstream ATG codon; and the premature stop codon is in-frame or out-of-frame with the downstream ATG codon. In certain aspects, the seed is a maize seed.
[0019] In particular examples, the seed comprises the targeted modification comprising an addition, deletion, or substitution of at least one nucleotide that introduces or creates a stop codon. In further aspects, the targeted modification results in translation reinitiation from the downstream ATG codon, thereby generating a truncated dwarf8 protein.
[0020] In additional examples, the premature stop codon is introduced between a first methionine, Ml, and a second methionine corresponding to methionine at position M53, methionine at position M65, methionine at position M67, methionine at position M69, methionine at position M106, methionine at position M184, methionine at position M201, methionine at position M271, or methionine at position M280 of SEQ ID NO: 6. The second methionine is in-frame with the first methionine and the premature stop codon is in-frame or out- of-frame with the second methionine. For example, the maize seed comprises the targeted modification comprising an in-frame or out-of-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO: 48, and in certain aspects, the resulting maize seed comprises SEQ ID NO: 7 or SEQ ID NO: 58. In further aspects, the maize seed is from a modified maize plant that exhibits a semi-dwarf phenotype. In still further aspects, the maize seed is a maize inbred or hybrid seed produced by a method comprising crossing, selfing, double haploid, and a combination of the foregoing. In other examples, the maize seed is a hybrid maize seed comprising the targeted modification in a heterozygous state.
[0021] Provided herein is a method of identifying a genomic variation in a genomic region of a plant, the method comprising genotyping of one or more isolated polynucleotide samples of one or more maize plants, wherein the polynucleotide samples comprise a portion of the polynucleotide of the genomic region that encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6 and identifying the genomic variation based on the genotyping. In some examples, the genomicvariation comprises an addition, deletion, or substitution of at least one nucleotide that introduces or creates a premature stop codon in the genomic region that encodes the polypeptide. In further examples, the genomic variation comprising the premature stop codon is introduced between a first methionine, Ml, and a second methionine corresponding to methionine at position M53, methionine at position M65, methionine at position M67, methionine at position M69, methionine at position Ml 06, methionine at position Ml 84, methionine at position M201, methionine at position M271, or methionine at position M280 of SEQ ID NO: 6. For example, the method includes identifying a genomic variation that comprises an in-frame or out-of-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO: 48 in a maize plant and the genomic variation results in a semi-dwarf phenotype due to translation reinitiation at the downstream ATG codon.
[0022] Also provided is a recombinant DNA construct comprising a selectable marker for herbicide tolerance, a CRISPR-Cas genome editing system that targets a genomic polynucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6 to introduce a targeted mutation, wherein the targeted mutation is introduced between a first or original starting ATG codon of the open reading frame and a subsequent downstream ATG codon and wherein the downstream ATG codon is located upstream or within the 5’ region of the GRAS domain coding sequence of an endogenous dwarf8 gene.
[0023] Provided herein is a method of targeted creation of an in-frame or out-of-frame translation reinitiation site at a plant genomic locus encoding a gene of interest, the method comprising deleting, inserting, or modifying one or more nucleotides at a genomic sequence comprising an open reading frame such that a premature stop codon after the codon encoding the first or original starting Met of the open reading frame; and creating a translation reinitiation start site that is downstream of the premature stop codon such that the ribosomal translational machinery is capable of initiating translation at the reinitiation start site, wherein the reinitiation start site is positioned at a sufficient distance from the first or the original Met to provide the translational machinery to be associated with and recruit the necessary components to reinitiate translation. In some aspects, the gene of interest is a maize gene involved in plant architecture modulation and confers a dominant phenotype. In further aspects, the targeted modification is performed by a site-directed DNA modifying agent. In further aspects, the translation reinitiationsite may result in the production of a transcript and / or polypeptide that confers an agronomic trait of interest, wherein the agronomic trait of interest is selected from the group consisting of plant height, yield, maturity, moisture content, disease resistance, drought tolerance, and nutrient use efficiency. In some aspects, the plant is a maize plant.BRIEF DESCRIPTION OF DRAWINGS AND SEQUENCE LISTING
[0024] The disclosure can be more fully understood from the following detailed description and the accompanying drawings and Sequence Listing that form a part of this application, which are incorporated herein by reference.
[0025] FIG. 1 shows a partial map of the D8 transcript sequence showing domains, regions, and motifs according to Uniprot, as well as methionine positions.
[0026] FIG. 2 shows partial alignments of ZM-D8 WT and edited protein variants at the N- terminal region.
[0027] FIG. 3 shows plant height of selected T2 plants with different edits and different zygosities at 51 days after planting.
[0028] FIG. 4 shows leaf area of selected T2 plants with different edits and different zygosities at 51 days after planting. Measurements were taken through automated greenhouse imaging.
[0029] FIG. 5 shows percent (%) green data of selected T2 plants with different edits and different zygosities at 51 days after planting.
[0030] FIG. 6 shows the construct design for the protoplast transient assay with the luciferase reporter gene to investigate the occurrence of translation reinitiation in the D8 variants.
[0031] The sequence descriptions summarize the Sequence Listing attached hereto, which is hereby incorporated by reference. The Sequence Listing contains one letter codes for nucleotide sequence characters and the single and three letter codes for amino acids as defined in the IUPAC IUB standards described in Nucleic Acids Research 13:3021 3030 (1985) and in the Biochemical Journal 219(2):345 373 (1984).
[0032] Table 1: Sequence Listing DescriptionDETAILED DESCRIPTION
[0033] The disclosure of all patents, patent applications, and publications cited herein are incorporated by reference in their entirety.
[0034] As used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural reference unless the context clearly dictates otherwise. Thus, for example, reference to “a plant” includes a plurality of such plants, reference to “a cell” includes one or more cells and equivalents thereof known to those skilled in the art, and so forth.I. DELLA proteins and DwarfS (1) }
[0035] DELLA proteins are a subfamily of the GRAS superfamily of proteins and play an important role in the negative regulation of gibberellin (GA) signaling, which affects processes like seed germination, stem elongation, and flowering. In the presence of gibberellin, the DELLA proteins associate with GID1 (Gibberellin insensitive dwarfl) receptors and are then ubiquitinated and degraded through the 26S proteasome pathway, resulting in the subsequent derepression or transactivation of downstream effectors of the gibberellin pathway. DELLAs operate in the nucleus and function as transcriptional regulators.
[0036] FIG. 1 shows a partial map of the D8 transcript, along with the corresponding regions of the translated polypeptide sequence showing domains, regions, and motifs according to Uniprot (Universal Protein Resource), a comprehensive, high-quality, and publicly accessible databasethat provides detailed information about protein sequences and their functions. The DELLA domain is a key feature of DELLA proteins and is located in the N-terminal region of the DELLA proteins. The DELLA domain coding sequence includes nucleotides 1325 to 1546 of the dwarf8 gene sequence set forth in SEQ ID NO: 5. The DELLA domain comprises the DELLA, LExLE, and VHYNP motifs. The primary function of the DELLA domain is to mediate the interaction between DELLA proteins and GID1.
[0037] In addition to the DELLA domain, the GRAS domain is a crucial part of DELLA proteins. The GRAS domain is located in the C-terminal region of these proteins and is responsible for their ability to interact with other proteins. This domain is composed of five conserved subdomains, Leucine receptor 1 (LRI), VHIID, LRII, PFYRE, and SAW1, and the LxCxE, VHIID, and LXXLL motifs. The subdomains help DELLA proteins form complexes with various transcription factors and other regulatory proteins, allowing them to modulate gene expression and influence plant responses to environmental signals.
[0038] The DELLA proteins are highly conserved across land plants, from mosses to angiosperms. Their evolution is marked by gene duplication and functional diversification, especially in vascular plants. For example, Arabidopsis thaliana has five DELLA genes (GAI, RGA, RGL1, RGL2, RGL3), while monocots like rice and wheat typically have one (SLR1 and RHT-1, respectively)
[0039] The maize genome has ten DELLA protein members of the GRAS family. Only two of these DELLAs, dwarf8 (D8) and dwarf9 (D9) have been reported to be involved in gibberellic acid (GA) sensing and plant architecture. The dwarf8 gene encodes a GRAS family transcription factor and has been identified as an orthologue of the gibberellic acid-insensitive (GAI) gene, a negative regulator of GA response in Arabidopsis.
[0040] In some examples, maize wildtype dwarf8 sequences encompass SEQ ID NOs:5 and 6. In particular examples, maize genome edited variants of dwarf8 encompass SEQ ID NOs: 7-16, 50, 51, 53, 56, and 58.II. Reduced stature maize phenotype and benefits for grain and forage production
[0041] The present disclosure relates to maize plants exhibiting a reduced stature phenotype. As used herein, the term "reduced stature" may encompass a range of phenotypic classifications including dwarf, strong dwarf, severe dwarf, semi-dwarf, and intermediate dwarf, depending onthe degree of height reduction relative to a wildtype maize plant. In certain examples, the reduced stature phenotype is achieved through targeted modification of the dwarfS gene, via genome editing. In certain embodiments, “semi-dwarf’ generally refers to milder or less severe dwarf phenotype where the plants exhibit a height reduction of about 25% to about 35%, or about 30% compared to its wild-type control plant. In addition to plant height, in certain embodiments the ear height (e.g., the position of the ear in a maize plant) is also reduced. In certain embodiments, other agronomic characteristics such as leaf size, leaf angle and leaf canopy are modified.
[0042] An important agronomic advantage of reduced stature maize is its enhanced resilience to severe weather conditions. The shortened architecture of the plants results in a lower center of gravity and increased stalk robustness, which collectively reduce susceptibility to stalk lodging and green snap under high wind or storm events. This improved standability contributes to yield preservation and facilitates mechanical harvest.
[0043] In addition to structural benefits, reduced stature maize supports increased planting densities without significantly impacting agronomic performance. Conventional tall maize hybrids often experience interrow competition at high population levels, which can diminish light interception and nutrient acquisition. In contrast, reduced stature hybrids are capable of thriving in narrow row spacing and dense plantings. Their upright leaf orientation and reduced tassel size minimize intra-canopy shading, thereby enhancing light penetration and photosynthetic efficiency throughout the plant canopy.
[0044] Reduced stature maize plants may also be advantageous for forage applications. These plants may exhibit a favorable ear-to-stover biomass ratio, which contributes to the production of high-quality forage characterized by enhanced digestibility and elevated energy content. Such attributes are beneficial for livestock feed systems, where nutritional efficiency and biomass utilization are important.
[0045] Another agronomic benefit of reduced stature maize is improved operational efficiency. Due to shorter height, these hybrids remain accessible to standard ground-based agricultural equipment throughout the growing season. This accessibility enables timely and precise in- season applications of agricultural inputs such as fertilizers, fungicides, and pesticides. As a result, reliance on aerial application methods may be reduced, thereby lowering input costs and improving application accuracy. Furthermore, ear placement in reduced stature maize isselectively bred to align with conventional harvesting equipment specifications, facilitating seamless integration into existing farm operations without the need for specialized machinery.
[0046] Reduced stature maize also contributes to environmental sustainability. The ability to support higher planting densities, combined with robust root architecture, enhances soil coverage and promotes erosion control by forming a dense canopy that mitigates rainwater runoff. Additionally, the deeper and more efficient root systems improve drought tolerance and nutrient uptake, thereby increasing plant resilience under abiotic stress conditions. These features collectively support more consistent yield performance across variable environmental conditions and reduce the likelihood of crop failure.
[0047] Maize plants exhibiting reduced stature possess a number of agronomically beneficial traits that enhance their suitability for grain and forage production. In some examples, modifying plant stature by one or more methods and compositions disclosed herein is characterized by one or more of the following agronomic traits: a shorter stature, reduced internode length, increased stalk / stem diameter, lower ear height, improved lodging resistance, reduced green snap, deeper roots, increased leaf area, earlier canopy closure, altered foliar water content and / or higher stomatai conductance under water / nutrient limiting conditions, increased leaf greenness, improved yield-related traits including a larger female reproductive part of a plant e.g., com ear, panicle, an increase in ear weight, harvest index, yield, seed number, seed size, and / or seed weight, relative to a wild type or control plant. Increased stress tolerance e.g., drought tolerance, nitrogen utilization, and / or tolerance to higher planting density are contemplated. Modified maize plants are provided that have at least one beneficial agronomic trait.III. Translation reinitiation
[0048] Translation reinitiation is a mechanism by which a ribosome, after completing translation of an upstream open reading frame (uORF), resumes scanning the same mRNA and initiates translation at a downstream start codon. This process enables fine-tuned regulation of gene expression and is particularly relevant for genes involved in stress responses and developmental pathways.
[0049] Translation reinitiation impacts gene expression by providing an additional layer of regulation. By allowing ribosomes to re-initiate translation after an upstream open reading frame(uORF), cells can control the amount of protein produced from a single mRNA. This is important for proteins that need to be tightly regulated.
[0050] Translation reinitiation in eukaryotic systems is modulated by a variety of internal molecular factors that influence the ability of ribosomes to resume translation following termination at upstream open reading frames (uORFs). One such factor is the nature of the uORFs themselves. Short uORFs may allow the ribosome’s 40S subunit to remain associated post-termination, facilitating reinitiation at a downstream start codon.
[0051] Cis-acting elements within the mRNA also play a significant role in modulating reinitiation efficiency. Sequences flanking the uORF, including both 5' and 3' enhancer regions, can stabilize the ribosomal subunit and promote reinitiation. For example, AU-rich sequences located near the stop codon have been shown to enhance reinitiation, whereas GC-rich sequences may inhibit this process.
[0052] Initiation factors (elFs) are another class of internal modulators. Eukaryotic initiation factors (elFs), such as eIF3 (particularly subunit eIF3a), eIF2, eIF4A, and eIF4G, play essential roles in stabilizing the ribosome and recruiting the ternary complex (eIF2 GTP Met-tRNAi) necessary for reinitiation.
[0053] The interci stronic distance, defined as the nucleotide spacing between the stop codon of the uORF and the downstream start codon, may also play a role in reinitiation efficiency. A sufficient distance allows time for the reacquisition of initiation factors by the ribosome, thereby enhancing the likelihood of reinitiation.
[0054] Ribosome recycling efficiency, where the 40S subunit remains associated, supports reinitiation proteins such as ABCE1 and eIF3j influence the recycling process, determining whether the ribosome is fully disassembled or not.
[0055] Methods and compositions are provided herein that utilize translation reinitiation to alter D8 protein expression, resulting in an intermediate reduced stature or semi-dwarf phenotype. In some examples, a stop codon introduced or created in-frame or out-of-frame between the first methionine (Ml) and a downstream start codon prior to the GRAS domain supports translation reinitiation, resulting in a truncated D8 protein and a reduced stature phenotype. In particular examples, a stop codon between Ml and M201 may enable reinitiation, provided the downstream ATG is in a favorable context (e.g., strong Kozak sequence or optimal spacing between the stop codon and the reinitiation ATG codon).
[0056] In some examples disclosed herein, targeted genomic modification of the D8 gene is expected to result in translation reinitiation and expression of a partial ZM-D8 protein lacking the N-terminal DRM (Disordered Region Motif) and DELLA motifs, resulting in an intermediate dwarf or semi-dwarf phenotype. In other cases, reinitiation likely did not occur, in part due to incompatible sequence features, resulting in a wildtype phenotype. In particular examples, an inframe or out-of-frame stop codon between Ml and M53 of the D8 protein allows translation reinitiation and the plant exhibits a semi-dwarf phenotype. Results described herein indicate that the presence of translation reinitiation mechanism in plants may be responsible for the underlying the semi-dwarf phenotype. This mechanism may be leveraged to engineer desired traits through targeted gene editing to create (i) premature stop codons, (ii) downstream start codons, or (iii) a combination of the foregoing.
[0057] The methods, composition and guidance provided herein for an exemplary gene d8, can be readily adapted to provide a systematic genome-wide genome editing efforts to promote translation reinitiation of genes of interest. Such a genome editing strategy would take into account some of the translational reinitiation characteristics such as, e.g., the location of the stop codon and the distance to the downstream start codon, the location of the uORF and other translational efficiency factors such as Kozak sequence.IV. Description of terms
[0058] An “isolated polynucleotide” is used herein to refer to a polymer of ribonucleotides (RNA) or deoxyribonucleotides (DNA) that is single or double-stranded, optionally containing synthetic, non-natural or altered nucleotide bases that is no longer in its natural environment, for example in vitro. An isolated polynucleotide in the form of DNA may be comprised of one or more segments of cDNA, genomic DNA or synthetic DNA.
[0059] A "recombinant" nucleic acid molecule (or DNA) is used herein to refer to a nucleic acid sequence that is in a recombinant bacterial or plant host cell. In some embodiments, "isolated" or "recombinant" nucleic acid is free of sequences (preferably protein encoding sequences) that naturally flank the nucleic acid (i.e., sequences located at the 5' and 3' ends of the nucleic acid) in the genomic DNA of the organism from which the nucleic acid is derived. For purposes of the disclosure, "isolated" or "recombinant" when used to refer to nucleic acid molecules excludes isolated chromosomes.
[0060] The terms “polynucleotide”, “polynucleotide sequence”, “nucleic acid sequence”, “nucleic acid fragment”, and “isolated nucleic acid fragment” are used interchangeably herein. These terms encompass nucleotide sequences and the like and may be used interchangeably in the singular or plural. A "polynucleotide" is not intended to limit a polynucleotide of the disclosure to a polynucleotide comprising DNA. Those of ordinary skill in the art will recognize that polynucleotides can comprise ribonucleotides and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogues. The polynucleotides of the disclosure also encompass all forms of sequences including, but not limited to, single-stranded forms, doublestranded forms, hairpins, stem-and-loop structures, and the like. A polynucleotide may be a polymer of RNA or DNA that is single or double-stranded, that optionally contains synthetic, non-natural or altered nucleotide bases. A polynucleotide in the form of a polymer of DNA may be comprised of one or more segments of cDNA, genomic DNA, synthetic DNA, or mixtures thereof. Nucleotides (usually found in their 5’ monophosphate form) are referred to by a single letter designation as follows: “A” for adenylate or deoxyadenylate (for RNA or DNA, respectively), “C” for cytidylate or deoxycytidylate, “G” for guanylate or deoxyguanylate, “U” for uridylate, “T” for deoxythymidylate, “R” for purines (A or G), “Y” for pyrimidines (C or T), “K” for G or T, “H” for A or C or T, “I” for inosine, and “N” for any nucleotide.
[0061] “Regulatory sequences” and “regulatory elements” refer to nucleotide sequences located upstream (5’ non-coding sequences), within, or downstream (3’ non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include, but are not limited to, promoters, enhancers, translation leader sequences, 5’ untranslated sequences or region (5’- UTR), 3’ untranslated sequences or region (3’-UTR), introns, polyadenylation target sequences, RNA processing sites, effector binding sites, and stem-loop structures. Regulatory sequences or regulatory elements may act in "cis" or "trans", and generally it acts in "cis", i.e. it activates expression of genes located on the same nucleic acid molecule, e.g. a chromosome, where the regulatory element is located. The nucleic acid molecule regulated by a regulatory element does not necessarily have to encode a functional peptide or polypeptide, e.g., the regulatory element can modulate the expression of a short interfering RNA or an anti-sense RNA.
[0062] An enhancer element is any nucleic acid molecule that increases transcription of a nucleic acid molecule when functionally linked to a promoter regardless of its relative position. An enhancer may be an innate element of the promoter, or a heterologous element inserted to enhance the level or tissue-specificity of a promoter.
[0063] A repressor (also sometimes called herein silencer) is defined as any nucleic acid molecule which inhibits the transcription when functionally linked to a promoter regardless of relative position.
[0064] ‘ ‘Promoter” generally refers to a nucleic acid fragment involved in recognition and binding of RNA polymerase and other proteins to initiate transcription of another nucleic acid fragment. A promoter generally includes a core promoter (also known as minimal promoter) sequence that includes a minimal regulatory region to initiate transcription, that is a transcription start site. Generally, a core promoter includes a TATA box and a GC rich region associated with a CAAT box or a CCAAT box. These elements act to bind RNA polymerase II to the promoter and assist the polymerase in locating the RNA initiation site. Some promoters may not have a TATA box or CAAT box or a CCAAT box but instead may contain an initiator element for the transcription initiation site. A core promoter is a minimal sequence required to direct transcription initiation and generally may not include enhancers or other UTRs. Promoters may be derived in their entirety from a native gene or be composed of different elements derived from different promoters found in nature or even comprise synthetic DNA segments. It is understood by those skilled in the art that different promoters may direct the expression of a gene in different tissues or cell types, or at different stages of development, or in response to different environmental conditions. Core promoters are often modified to produce artificial, chimeric, or hybrid promoters, and can further be used in combination with other regulatory elements, such as cis-elements, 5’UTRs, enhancers, or introns, that are either heterologous to an active core promoter or combined with its own partial or complete regulatory elements.
[0065] The term "cis-element" generally refers to transcriptional regulatory element that affects or modulates expression of an operably linked transcribable polynucleotide, where the transcribable polynucleotide is present in the same DNA sequence. A cis-element may function to bind transcription factors, which are trans-acting polypeptides that regulate transcription.
[0066] "Promoter functional in a plant" is a promoter capable of initiating transcription in plant cells whether or not its origin is from a plant cell.
[0067] “ Tissue-specific promoter” and “tissue-preferred promoter” are used interchangeably to refer to a promoter that is expressed predominantly but not necessarily exclusively in one tissue or organ, but that may also be expressed in one specific cell.
[0068] "Developmentally regulated promoter" generally refers to a promoter whose activity is determined by developmental events.
[0069] "Constitutive promoter" generally refers to promoters active in all or most tissues or cell types of a plant at all or most developing stages. As with other promoters classified as “constitutive” (e.g., ubiquitin), some variation in absolute levels of expression can exist among different tissues or stages. The term “constitutive promoter” or “tissue-independent” are used interchangeably herein.
[0070] A "domain" generally refers to a contiguous stretch of nucleotides (that can be RNA, DNA, and / or RNA-DNA-combination sequence) or amino acids.
[0071] A “heterologous nucleotide sequence” generally refers to a nucleic acid sequence that originates from a different species, organism, or genetic context than the host organism or genetic locus into which it is introduced. A heterologous sequence may encode a protein, RNA, or regulatory element and is not naturally found at the site of integration or expression in the host genome. The term encompasses sequences that are synthetically designed, modified, or derived from other organisms, and includes coding sequences, promoters, enhancers, untranslated regions (UTRs), and other functional elements that are exogenous to the host system. The terms “heterologous nucleotide sequence”, “heterologous sequence”, “heterologous nucleic acid fragment”, and “heterologous nucleic acid sequence” are used interchangeably herein.
[0072] The heterologous polynucleotide can be stably integrated within the genome such that the polynucleotide is passed on to successive generations. The heterologous polynucleotide may be integrated into the genome alone or as part of a recombinant DNA construct. The alterations of the genome (chromosomal or extra-chromosomal) by conventional plant breeding methods, by genome editing procedures that do not result in an insertion of a foreign polynucleotide, or by naturally occurring events such as random cross-fertilization, non-recombinant viral infection, non-recombinant bacterial transformation, non-recombinant transposition, or spontaneous mutation are also methods of modifying a host genome.
[0073] A “functional fragment” refers to a portion or subsequence of the sequence described in the present disclosure in which, the ability to modulate gene expression is retained. Fragmentscan be obtained via methods such as site-directed mutagenesis and synthetic construction. As with the provided promoter sequences described herein, the functional fragments operate to promote the expression of an operably linked heterologous nucleotide sequence, forming a recombinant DNA construct (also, a chimeric gene). For example, the fragment can be used in the design of recombinant DNA constructs to produce the desired phenotype in a transformed plant. Recombinant DNA constructs can be designed for use in co- suppress! on or antisense by linking a promoter fragment in the appropriate orientation relative to a heterologous nucleotide sequence.
[0074] A nucleic acid fragment that is functionally equivalent to the target sequences of the present disclosure is any nucleic acid fragment that is capable of modulating the expression of a coding sequence or functional RNA in a similar manner to the target sequences of the present disclosure.
[0075] The polynucleotide sequence of the targets of the present disclosure (e.g., SEQ ID NOs: 3 and 4), may be modified or altered to enhance their modulation characteristics. As one of ordinary skill in the art will appreciate, modification or alteration can also be made without substantially affecting the gene expression function. The methods are well known to those of skill in the art. Sequences can be modified, for example by insertion, deletion, or replacement of template sequences through any modification approach.
[0076] A "variant promoter" as used herein, is the sequence of the promoter or the sequence of a functional fragment of a promoter containing changes in which one or more nucleotides of the original sequence is deleted, added, and / or substituted, while substantially maintaining promoter function. One or more base pairs can be inserted, deleted, or substituted internally to a promoter. In the case of a promoter fragment, variant promoters can include changes affecting the transcription of a minimal promoter to which it is operably linked. Variant promoters can be produced, for example, by standard DNA mutagenesis techniques or by chemically synthesizing the variant promoter or a portion thereof.
[0077] Modifying stature of plants by one or more methods and compositions disclosed here are characterized by one or more of the following traits: a shorter stature, severe dwarf plant height, semi-dwarf plant height, or intermediate dwarf plant height, reduced internode length, increased stalk / stem diameter, lower ear height, improved lodging resistance, reduced green snap, deeper roots, increased leaf area, increased leaf greenness, earlier canopy closure, altered foliar watercontent and / or higher stomatai conductance under water / nutrient limiting conditions, improved yield-related traits including a larger female reproductive part of a plant e.g., com ear, panicle, an increase in ear weight, harvest index, yield, seed number / panicle number, and / or seed weight, relative to a wild type or control plant. Increased stress tolerance e.g., drought tolerance, nitrogen utilization, and / or tolerance to higher planting density are contemplated.
[0078] In some examples of the present disclosure, the fragments of polynucleotide sequences disclosed herein can comprise at least about 20 contiguous nucleotides, or at least about 50 contiguous nucleotides, or at least about 75 contiguous nucleotides, or at least about 100 contiguous nucleotides, or at least about 150 contiguous nucleotides, or at least about 200 contiguous nucleotides of nucleic acid sequences or polypeptides encoded designated by the SEQ ID NOs listed in Table 1. In other examples, the fragments can comprise at least about 250 contiguous nucleotides, or at least about 300 contiguous nucleotides, or at least about 350 contiguous nucleotides, or at least about 400 contiguous nucleotides, or at least about 450 contiguous nucleotides, or at least about 500 contiguous nucleotides, or at least about 550 contiguous nucleotides, or at least about 600 contiguous nucleotides, or at least about 650 contiguous nucleotides, or at least about 700 contiguous nucleotides, or at least about 750 contiguous nucleotides, or at least about 800 contiguous nucleotides, or at least about 850 contiguous nucleotides , or at least about 900 contiguous nucleotides, or at least about 950 contiguous nucleotides, or at least about 1000 contiguous nucleotides, or at least about 1050 contiguous nucleotides and further may include a sequence from Table 1 listings.
[0079] The terms “full complement” and “full-length complement” are used interchangeably herein and refer to a complement of a given nucleotide sequence, wherein the complement and the nucleotide sequence consist of the same number of nucleotides and are 100% complementary.
[0080] The terms “substantially similar” and “corresponding substantially” as used herein refer to nucleic acid fragments wherein changes in one or more nucleotide bases do not affect the ability of the nucleic acid fragment to mediate gene expression or produce a certain phenotype. These terms also refer to modifications of the nucleic acid fragments of the instant disclosure such as deletion or insertion of one or more nucleotides that do not substantially alter the functional properties of the resulting nucleic acid fragment relative to the initial, unmodifiedfragment. It is therefore understood, as those skilled in the art will appreciate, that the disclosure encompasses more than the specific exemplary sequences.
[0081] The transitional phrase “consisting essentially of’ generally refers to a composition, method that includes materials, steps, features, components, or elements, in addition to those literally disclosed, provided that these additional materials, steps, features, components, or elements do not materially affect the basic and novel characterise cfs) of the claimed subject matter, e.g., one or more of the claimed sequences.
[0082] The isolated promoter sequence comprised in the recombinant DNA construct of the present disclosure can be modified to provide a range of constitutive expression levels of the heterologous nucleotide sequence. Thus, less than the entire promoter regions may be utilized and the ability to drive expression of the coding sequence retained. However, it is recognized that expression levels of the mRNA may be decreased with deletions of portions of the promoter sequences. Likewise, the tissue-independent, constitutive nature of expression may be changed.
[0083] Modifications of the isolated promoter sequences of the present disclosure can provide for a range of constitutive expression of the heterologous nucleotide sequence. Thus, they may be modified to be weak constitutive promoters or strong constitutive promoters. Generally, by “weak promoter” is intended a promoter that drives expression of a coding sequence at a low level. By “low level” is intended levels about 1 / 10,000 transcripts to about 1 / 100,000 transcripts to about 1 / 500,000 transcripts. Conversely, a strong promoter drives expression of a coding sequence at high level, or at about 1 / 10 transcripts to about 1 / 100 transcripts to about 1 / 1,000 transcripts. Similarly, a “moderate constitutive” promoter is somewhat weaker than a strong constitutive promoter like the maize ubiquitin promoter.
[0084] Planting density in a field, e.g., may range from about at least 36,000 plants per acre, at least 40,000 plants per acre, at least 42,000 plants per acre, at least 44,000 plants per acre, at least 45,000 plants per acre, at least 46,000 plants per acre, at least 48,000 plants per acre, 50,000 plants per acre, at least 52,000 plants per acre, at least 54,000 per acre, or at least 56,000 plants per acre. In an embodiment, com plants may be planted at a higher density, such as in a range from about 36,000 plants per acre to about 60,000 plants per acre, or about 40,000 plants per acre to about 58,000 plants per acre, or about 42,000 plants per acre to about 58,000 plants per acre, or about 40,000 plants per acre to about 45,000 plants per acre, or about 45,000 plants per acre to about 50,000 plants per acre, or about 50,000 plants per acre to about 58,000 plants per acre, orabout 52,000 plants per acre to about 56,000 plants per acre, or about 38,000 plants per acre, about 42,000 plant per acre, about 46,000 plant per acre, or about 48,000 plants per acre, about 50,000 plants per acre, or about 52,000 plants per acre, or about 54,000 plant per acre, as opposed to a standard planting density range, such as about 18,000 plants per acre to about 38,000 plants per acre.
[0085] According to certain embodiments of the present disclosure, genome edited maize plants are provided that have (i) a plant height that is at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% less than the height of a wild-type or control plant, and / or (ii) a stem or stalk diameter that is at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% greater than the stem diameter of the wild-type or control plant.
[0086] According to embodiments of the present disclosure, a genome edited maize plant may have a reduced plant height that is no more than 40%, 50%, 55%, or 60% shorter than the height of a wild-type or control plant.
[0087] For example, a genome edited plant may have (i) a plant height that is at least 30%, at least 25%, or at least 20% less or shorter, but not greater or more than 50% shorter, than a wild type or control plant, and / or (ii) a stem or stalk diameter that is that is at least 5%, at least 10%, or at least 15% greater, but not more than 30%, 35%, or 40% greater, than a wild type or control plant.
[0088] According to embodiments of the present disclosure, modified corn plants are provided that comprise a height between 5% and 50%, between 10% and 40%, between 15% and 35%, between 10% and 30%, between 10% and 25%, between 20% and 50%, between 20% and 40%, between 25% and 35%, between 25% and 40%, between 5% and 30%, between 5% and 25%, between 15% and 30%, between 15% and 35%, or between 30% and 45% less than the height of a wild-type or control plant.
[0089] According to certain embodiments of the present disclosure, genome edited maize plants are provided that have an increase in harvestable grain yield of at least 1 bushel per acre, at least 2 bushels per acre, at least 3 bushels per acre, at least 4 bushels per acre, at least 5 bushels per acre, at least 6 bushels per acre, at least 7 bushels per acre, at least 8 bushels per acre, at least 9 bushels per acre, or at least 10 bushels per acre, relative to a wild-type or control plant. A genome edited maize plant according to the present disclosure may have an increase inharvestable grain yield that is at least 1 %, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 1 1 %, at least 12%, at least 13%, at least 14%, at least 15%, or at least 20% greater than the harvestable yield of a wild-type or control plant.
[0090] In addition to modulating gene expression, the expression modulating elements disclosed herein are also useful as probes or primers in nucleic acid hybridization experiments. The nucleic acid probes and primers hybridize under stringent conditions to a target DNA sequence. A "probe" is generally referred to an isolated / synthesized nucleic acid to which, is attached a conventional detectable label or reporter molecule, such as for example, a radioactive isotope, ligand, chemiluminescent agent, bioluminescent molecule, fluorescent label or dye, or enzyme. Such detectable labels may be covalently linked or otherwise physically associated with the probe. "Primers" generally referred to isolated / synthesized nucleic acids that hybridize to a complementary target DNA strand which is then extended along the target DNA strand by a polymerase, e.g., a DNA polymerase. Primer pairs often used for amplification of a target nucleic acid sequence, e.g., by the polymerase chain reaction (PCR) or other conventional nucleic-acid amplification methods. Primers are also used for a variety of sequencing reactions, sequence captures, and other sequence-based amplification methodologies. Primers are generally about 15, 20, 25 nucleotides or more, and probes can also be longer about 30, 40, 50 and up to a few hundred base pairs. Such probes and primers are used in hybridization reactions to target DNA or RNA sequences under high stringency hybridization conditions or under lower stringency conditions, depending on the need.
[0091] Preferred substantially similar nucleic acid sequences encompassed by this disclosure are those sequences that are 80% identical to the nucleic acid fragments reported herein or which are 80% identical to any portion of the nucleotide sequences reported herein. More preferred are nucleic acid fragments which are 90% identical to the nucleic acid sequences reported herein, or which are 90% identical to any portion of the nucleotide sequences reported herein. Most preferred are nucleic acid fragments which are 95% identical to the nucleic acid sequences reported herein, or which are 95% identical to any portion of the nucleotide sequences reported herein. It is well understood by one skilled in the art that many levels of sequence identity are useful in identifying related polynucleotide sequences. Useful examples of percent identities are those listed above, or also preferred is any integer percentage from 71% to 100%, such as 71%,72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% and 100%.
[0092] In some examples, the isolated sequences of the present disclosure comprises a nucleotide sequence having at least 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% and 100% sequence identity, based on the Clustal Omega method of alignment (Sievers and Higgins. Clustal Omega for making accurate alignments of many protein sequences. Protein Sci. 2018 Jan;27(l):135-145) with pairwise alignment default parameters, when compared to the nucleotide sequence of SEQ ID NO:5. It is known to one of skilled in the art that a 5’ UTR region can be altered (deletion or substitutions of bases) or replaced by an alternative 5’UTR while maintaining promoter activity.
[0093] A “substantially similar sequence” generally refers to variants of the disclosed sequences such as those that result from site-directed mutagenesis, as well as synthetically derived sequences. A substantially similar promoter sequence of the present disclosure also generally refers to those fragments of a particular promoter nucleotide sequence disclosed herein that operate to promote the constitutive expression of an operably linked heterologous nucleic acid fragment. These promoter fragments comprise at least about 20 contiguous nucleotides, at least about 50 contiguous nucleotides, at least about 75 contiguous nucleotides, preferably at least about 100 contiguous nucleotides of the particular promoter nucleotide sequence disclosed herein or a sequence that is at least 95 to about 99% identical to such contiguous sequences. The nucleotides of such fragments will usually include the TATA recognition sequence (or CAAT box or a CCAAT) of the particular promoter sequence. Such fragments may be obtained by use of restriction enzymes to cleave the naturally occurring promoter nucleotide sequences disclosed herein; by synthesizing a nucleotide sequence from the naturally occurring promoter DNA sequence; or may be obtained through the use of PCR technology. Variants of these promoter fragments, such as those resulting from site-directed mutagenesis, are encompassed by the compositions of the present disclosure.
[0094] “ Codon degeneracy” generally refers to divergence in the genetic code permitting variation of the nucleotide sequence without affecting the amino acid sequence of an encoded polypeptide. Accordingly, the instant disclosure relates to any nucleic acid fragment comprising a nucleotide sequence that encodes all or a substantial portion of the amino acid sequences setforth herein. The skilled artisan is well aware of the “codon-bias” exhibited by a specific host cell in usage of nucleotide codons to specify a given amino acid. Therefore, when synthesizing a nucleic acid fragment for improved expression in a host cell, it is desirable to design the nucleic acid fragment such that its frequency of codon usage approaches the frequency of preferred codon usage of the host cell.
[0095] Sequence alignments and percent identity calculations may be determined using a variety of comparison methods designed to detect similar or identical sequences including, but not limited to, the Megalign® program of the LASERGENE® bioinformatics computing suite (DNASTAR® Inc., Madison, WI). Unless stated otherwise, multiple alignment of the sequences provided herein were performed using the Clustal Omega method of alignment (Sievers et al. 2011. Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Molecular systems biology, 7(1), p.539) with the default parameters Dealign Input Sequences=no, Output Alignment Format= ClustalW with character counts, mBed-like Clustering Guide-tree=yes, mBed-like Clustering Iteration=yes, Number of Combined Iterations- 0, Max Guide Tree Iterations=(-1), Max HMM Iterations=(-1), Order=aligned.
[0096] Alternatively, the Clustal V method or Clustal W methods of alignment may be used (Higgins and Sharp (1989) CABIOS. 5: 151-53; Higgins et al. (1992) Comput. Appl. Biosci. 8: 189-191). The Clustal V method of alignment is used with the default parameters (GAP PENAL TY=10, GAP LENGTH PENAL TY=10). Default parameters for pairwise alignments and calculation of percent identity of protein sequences using the Clustal V method are KTUPLE=1, GAP PENALTY-3, WINDOW-5 and DIAGONALS SAVED-5. For nucleic acids these parameters are KTUPLE=2, GAP PENALTY=5, WIND0W=4 and DIAGONALS SAVED=4. After alignment of the sequences, using the Clustal V program, it is possible to obtain “percent identity” and “divergence” values by viewing the “sequence distances” table on the same program; unless stated otherwise, percent identities and divergences provided and claimed herein were calculated in this manner. The Clustal W method of alignment can be found in the MegAlign™ v6.1 program of the LASERGENE® bioinformatics computing suite (DNASTAR® Inc., Madison, Wis.). Default parameters for multiple alignment correspond to GAP PENALTY=10, GAP LENGTH PENALTY=0.2, Delay Divergent Sequences=30%, DNA Transition Weight=0.5, Protein Weight Matrix=Gonnet Series, DNA Weight Matrix=IUB. For pairwise alignments the default parameters are Alignment=Slow-Accurate, Gap Penalty=10.0,Gap Length=0.10, Protein Weight Matrix=Gonnet 250 and DNA Weight Matrix=IUB. After alignment of the sequences using the Clustal W program, it is possible to obtain “percent identity” and “divergence” values by viewing the “sequence distances” table in the same program.
[0097] In one embodiment, the % sequence identity is determined over the entire length of the molecule (nucleotide or amino acid). A “substantial portion” of an amino acid or nucleotide sequence comprises enough of the amino acid sequence of a polypeptide or the nucleotide sequence of a gene to afford putative identification of that polypeptide or gene, either by manual evaluation of the sequence by one skilled in the art, or by computer-automated sequence comparison and identification using algorithms such as BLAST (Altschul, S. F. et al., J. Mol. Biol. 215:403 410 (1993)) and Gapped Blast (Altschul, S. F. et al., Nucleic Acids Res. 25:3389 3402 (1997)). BLASTN generally refers to a BLAST program that compares a nucleotide query sequence against a nucleotide sequence database.
[0098] “Gene” includes a nucleic acid fragment that expresses a functional molecule such as, but not limited to, a specific protein, including regulatory sequences preceding (5’ non-coding sequences) and following (3’ non-coding sequences) the coding sequence. “Native gene” generally refers to a gene as found in nature with its own regulatory sequences.
[0099] A “mutated gene” is a gene that has been altered through human intervention. Such a “mutated gene” has a sequence that differs from the sequence of the corresponding non-mutated gene by at least one nucleotide addition, deletion, or substitution. In certain embodiments of the disclosure, the mutated gene comprises an alteration that results from a guide polynucleotide / Cas endonuclease system as disclosed herein. A mutated plant is a plant comprising a mutated gene.
[0100] “ Chimeric gene” or “recombinant expression construct”, which are used interchangeably, includes any gene that is not a native gene, comprising regulatory and coding sequences that are not found together in nature. Accordingly, a chimeric gene may comprise regulatory sequences and coding sequences that are derived from different sources.
[0101] “Modified protein” is translated from a mutated gene and is different from the protein translated from the non-mutated gene. In some examples, the modified protein may be the result of a frameshift mutation (via short deletion or nucleotide addition / deletion) in a gene. Such a mutation disrupts the normal reading frame which results in the entire gene sequence following the mutation to be incorrectly read. This can result in the addition of the wrong amino acids tothe protein and / or the creation of a codon that stops the protein from growing longer. In particular examples, the modified protein may be the result of an in-frame mutation in a gene resulting in translation reinitiation at a downstream start site (ATG) of the gene. In some examples, a modified protein is made when a CRISPR-Cas edit results in a protein that has an inframe stop codon between the first methionine and a downstream start site (ATG) located before or within the 5’ region of the GRAS domain of a D8 gene.
[0102] “Coding sequence” generally refers to a polynucleotide sequence which codes for a specific amino acid sequence. “Regulatory sequences” refer to nucleotide sequences located upstream (51non-coding sequences), within, or downstream (31non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences may include, but are not limited to, promoters, translation leader sequences, introns, and polyadenylation recognition sequences.
[0103] A “domain” is a contiguous stretch of nucleotides (that can be RNA, DNA, and / or RNA- DNA combination sequence) or amino acids (ex. DELLA, VHYNP, disordered, and GRAS domains). A “region” is a contiguous stretch of nucleotides (that can be RNA, DNA, and / or RNA-DNA combination sequence) or amino acids that is of interest (e.g. Disordered regions I and II). A “motif’ is a short contiguous stretch of up to 20 amino acids of biological interest (e.g. Disordered, DELLA, LExLE, VHYNP, and Poly S / T / V motifs). A “non-functional” DELLA domain or motif or region, in the context of D8 polypeptide generally refers to a polypeptide that is full-length or a partial length polypeptide comprising one or more functional domains of the maize D8 protein, but lacking a functional DELLA domain (e.g., the DELLA domain is absent, or out of frame, or contains one or more mutations that renders the DELLA domain lose its intended function.
[0104] An “intron” is an intervening sequence in a gene that is transcribed into RNA but is then excised in the process of generating the mature mRNA. The term is also used for the excised RNA sequences. An “exon” is a portion of the sequence of a gene that is transcribed and is found in the mature messenger RNA derived from the gene but is not necessarily a part of the sequence that encodes the final gene product.
[0105] The 5' untranslated region (5’UTR) (also known as a translation leader sequence or leader RNA) is the region of an mRNA that is directly upstream from the initiation codon. This regionis involved in the regulation of translation of a transcript by differing mechanisms in viruses, prokaryotes and eukaryotes.
[0106] The “3' non-coding sequences” refer to DNA sequences located downstream of a coding sequence and include polyadenylation recognition sequences and other sequences encoding regulatory signals capable of affecting mRNA processing or gene expression. The polyadenylation signal is usually characterized by affecting the addition of polyadenylic acid tracts to the 3' end of the mRNA precursor.
[0107] ‘ ‘RNA transcript” generally refers to a product resulting from RNA polymerase-catalyzed transcription of a DNA sequence. When an RNA transcript is a perfect complimentary copy of a DNA sequence, it is referred to as a primary transcript or it may be a RNA sequence derived from posttranscriptional processing of a primary transcript and is referred to as a mature RNA. “Messenger RNA” (“mRNA”) generally refers to RNA that is without introns and that can be translated into protein by the cell. “cDNA” generally refers to a DNA that is complementary to and synthesized from an mRNA template using the enzyme reverse transcriptase. The cDNA can be single-stranded or converted into double-stranded by using the KI enow fragment of DNA polymerase I. “Sense” RNA generally refers to RNA transcript that includes mRNA and so can be translated into protein within a cell or in vitro. “Antisense RNA” generally refers to a RNA transcript that is complementary to all or part of a target primary transcript or mRNA and that blocks expression or transcripts accumulation of a target gene. The complementarity of an antisense RNA may be with any part of the specific gene transcript, i.e., at the 5' non-coding sequence, 3’ non-coding sequence, introns, or the coding sequence. “Functional RNA” generally refers to antisense RNA, ribozyme RNA, or other RNA that may not be translated but has an effect on cellular processes.
[0108] The term “operably linked” or “functionally linked” generally refers to the association of nucleic acid sequences on a single nucleic acid fragment so that the function of one is affected by the other. For example, a promoter is operably linked with a coding sequence when it can affect the expression of that coding sequence (i.e., that the coding sequence is under the transcriptional control of the promoter). Coding sequences can be operably linked to regulatory sequences in sense or antisense orientation.
[0109] The terms “initiate transcription”, “initiate expression”, “drive transcription”, and “drive expression” are used interchangeably herein and all refer to the primary function of a promoter.As detailed throughout this disclosure, a promoter is a non-coding genomic DNA sequence, usually upstream (5') to the relevant coding sequence, and its primary function is to act as a binding site for RNA polymerase and initiate transcription by the RNA polymerase. Additionally, there is “expression” of RNA, including functional RNA, or the expression of polypeptide for operably linked encoding nucleotide sequences, as the transcribed RNA ultimately is translated into the corresponding polypeptide.
[0110] The term “expression”, as used herein, generally refers to the production of a functional end-product e.g., an mRNA or a protein (precursor or mature).
[0111] The term “expression cassette” as used herein, generally refers to a discrete nucleic acid fragment into which a nucleic acid sequence or fragment can be cloned or synthesized through molecular biology techniques.
[0112] Expression or overexpression of a gene involves transcription of the gene and translation of the mRNA into a precursor or mature protein. “Antisense inhibition” generally refers to the production of antisense RNA transcripts capable of suppressing the expression of the target protein. “Overexpression” generally refers to the production of a gene product in transgenic organisms that exceeds levels of production in normal or non-transformed organisms. “Cosuppression” generally refers to the production of sense RNA transcripts capable of suppressing the expression or transcript accumulation of identical or substantially similar foreign or endogenous genes (U.S. Patent No. 5,231,020). The mechanism of co-suppression may be at the DNA level (such as DNA methylation), at the transcriptional level, or at post-transcriptional level.
[0113] As stated herein, “suppression” includes a reduction of the level of enzyme activity or protein functionality (e.g., a phenotype associated with a protein) detectable in a transgenic plant when compared to the level of enzyme activity or protein functionality detectable in a non- transgenic or wild type plant with the native enzyme or protein. The level of enzyme activity in a plant with the native enzyme is referred to herein as “wild type” activity. The level of protein functionality in a plant with the native protein is referred to herein as “wild type” functionality. The term “suppression” includes lower, reduce, decline, decrease, inhibit, eliminate and prevent. This reduction may be due to a decrease in translation of the native mRNA into an active enzyme or functional protein. It may also be due to the transcription of the native DNA into decreased amounts of mRNA and / or to rapid degradation of the native mRNA. The term “native enzyme”generally refers to an enzyme that is produced naturally in a non-transgenic or wild type cell. The terms "non-transgenic" and "wild type" are used interchangeably herein.
[0114] “Altering expression” or “modulating expression” generally refers to the production of gene product(s) in plants in amounts or proportions that differ significantly from the amount of the gene product(s) produced by the corresponding wild-type plants (i.e., expression is increased or decreased).
[0115] “ Transformation” as used herein generally refers to both stable transformation and transient transformation.
[0116] “ Stable transformation” generally refers to the introduction of a nucleic acid fragment into a genome of a host organism resulting in genetically stable inheritance. Once stably transformed, the nucleic acid fragment is stably integrated in the genome of the host organism and any subsequent generation. Host organisms containing the transformed nucleic acid fragments are referred to as “transgenic” organisms. “Transient transformation” generally refers to the introduction of a nucleic acid fragment into the nucleus, or DNA-containing organelle, of a host organism resulting in gene expression without genetically stable inheritance.
[0117] The term “introduced” means providing a nucleic acid (e.g., expression construct) or protein into a cell. Introduced includes reference to the incorporation of a nucleic acid into a eukaryotic or prokaryotic cell where the nucleic acid may be incorporated into the genome of the cell and includes reference to the transient provision of a nucleic acid or protein to the cell. Introduced includes reference to stable or transient transformation methods, as well as sexual crossing. Thus, “introduced” in the context of inserting a nucleic acid fragment (e g., a recombinant DNA construct / expression construct) into a cell, means “transfection” or “transformation” or “transduction” and includes reference to the incorporation of a nucleic acid fragment into a eukaryotic or prokaryotic cell where the nucleic acid fragment may be incorporated into the genome of the cell (e.g., chromosome, plasmid, plastid or mitochondrial DNA), converted into an autonomous replicon, or transiently expressed (e.g., transfected mRNA).
[0118] “ Genome” as it applies to plant cells encompasses not only chromosomal DNA found within the nucleus, but organelle DNA found within subcellular components (e.g., mitochondrial, plastid) of the cell.
[0119] “ Genetic modification” generally refers to modification of any nucleic acid sequence or genetic element by insertion, deletion, or substitution of one or more nucleotides in an endogenous nucleotide sequence by mutagenesis, genome editing or by insertion of a recombinant nucleic acid, e.g., as part of a vector or construct in any region of the plant genomic DNA by routine transformation techniques. Examples of modification of genetic components include, but are not limited to, promoter regions, 5' untranslated leaders, introns, genes, 3' untranslated regions, and other regulatory sequences or sequences that affect transcription or translation of one or more nucleic acid sequences.
[0120] Further methods that introduce a modification or mutation randomly in a gene sequence can include, but are not limited to, chemical mutagenesis, such as mutagenic, teratogenic, or carcinogenic organic compounds, for example ethyl methanesulfonate (EMS), that produce random mutations in genetic material, and ionizing radiation mutagenesis, for example fast neutron bombardment. After mutation, screening can be performed to identify mutations that create premature stop codons or otherwise non-functional genes, or biological processes such as translation reinitiation. After mutation, screening can be performed to identify mutations that create functional genes that are capable of being expressed at elevated levels. Screening of mutants can be carried out by sequencing, or by the use of one or more probes or primers specific to the gene or protein. Specific mutations in polynucleotides can also be created that can result in modulated gene expression, modulated stability of mRNA, or modulated stability of protein. Such plants can be referred to as “non-naturally occurring” or “mutant” plants.Mutagenesis may be used to induce mutations in the D8 gene, specifically resulting in translation reinitiation at an in-frame downstream start site (ATG), resulting in a semi-dwarf phenotype.
[0121] ‘ ‘Plant" includes reference to whole plants, plant organs, plant tissues, seeds and plant cells and progeny of same. Plant cells include, without limitation, cells from seeds, suspension cultures, embryos, meristematic regions, callus tissue, leaves, roots, shoots, gametophytes, sporophytes, pollen, and microspores.
[0122] The terms “monocot” and “monocotyledonous plant” are used interchangeably herein. A monocot of the current disclosure includes the Gramineae.
[0123] The terms “dicot” and “dicotyledonous plant” are used interchangeably herein. A dicot of the current disclosure includes the following families: Brassicaceae, Leguminosae, and Solanaceae.
[0124] "Progeny" comprises any subsequent generation of a plant.
[0125] An "inbred" refers to a line that has been bred for genetic homogeneity.
[0126] A "hybrid" refers to the progeny obtained between the crossing of at least two genetically dissimilar parents.
[0127] An "elite plant" or “elite line” is any plant that has resulted from breeding and selection for superior agronomic performance. An elite plant can be an elite inbred or an elite hybrid.
[0128] “ Transient expression” generally refers to the temporary expression of often reporter genes such as -glucuronidase (GUS), fluorescent protein genes ZS-GREEN1, ZS-YELLOW1 Nl, AM-CYAN 1, DS-RED in selected certain cell types of the host organism in which the transgenic gene is introduced temporally by a transformation method. The transformed materials of the host organism are subsequently discarded after the transient gene expression assay.
[0129] Standard recombinant DNA and molecular cloning techniques used herein are well known in the art and are described more fully in Sambrook, J. et al., In Molecular Cloning: A Laboratory Manual; 2nd ed.; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, New York, 1989 (hereinafter “Sambrook et al., 1989”) or Ausubel, F. M., Brent, R., Kingston, R. E., Moore, D. D., Seidman, J. G., Smith, J. A. and Struhl, K., Eds.; In Current Protocols in Molecular Biology; John Wiley and Sons: New York, 1990 (hereinafter “Ausubel et al., 1990”).
[0130] ‘ ‘PCR” or “Polymerase Chain Reaction” is a technique for the synthesis of large quantities of specific DNA segments, consisting of a series of repetitive cycles (Perkin Elmer Cetus Instruments, Norwalk, CT). Typically, the double stranded DNA is heat denatured, the two primers complementary to the 3' boundaries of the target segment are annealed at low temperature and then extended at an intermediate temperature. One set of these three consecutive steps comprises a cycle.
[0131] The terms “plasmid”, “vector” and “cassette” refer to an extra chromosomal element often carrying genes that are not part of the central metabolism of the cell, and usually in the form of circular double-stranded DNA fragments. Such elements may be autonomously replicating sequences, genome integrating sequences, phage or nucleotide sequences, linear or circular, of a single- or double-stranded DNA or RNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construction which is capable of introducing a promoter fragment and DNA sequence for a selected gene product along with appropriate 3' untranslated sequence into a cell.
[0132] The term “recombinant DNA construct” or “recombinant expression construct” is used interchangeably and generally refers to a discrete polynucleotide into which a nucleic acid sequence or fragment can be moved. Preferably, it is a plasmid vector or a fragment thereof comprising the promoters of the present disclosure. The choice of plasmid vector is dependent upon the method that will be used to transform host plants. The skilled artisan is well aware of the genetic elements that must be present on the plasmid vector in order to successfully transform, select and propagate host cells containing the chimeric gene. The skilled artisan will also recognize that different independent transformation events will result in different levels and patterns of expression (Jones et al., EMBO J. 4:2411 2418 (1985); De Almeida et al., Mol. Gen. Genetics 218:78 86 (1989)), and thus that multiple events must be screened in order to obtain lines displaying the desired expression level and pattern. Such screening may be accomplished by PCR and Southern analysis of DNA, RT-PCR and Northern analysis of mRNA expression, Western analysis of protein expression, or phenotypic analysis.
[0133] Various changes in phenotype are of interest including, but not limited to, modifying the fatty acid composition in a plant, altering the amino acid content of a plant, altering a plant’s pathogen defense mechanism, modifying various agronomic characteristics, and the like. These results can be achieved by providing expression of heterologous products or increased expression of endogenous products in plants. Alternatively, the results can be achieved by providing for a reduction of expression of one or more endogenous products, particularly enzymes or cofactors in the plant. In another alternative, the results can be achieved by translation reinitiation leading to expression of a partial protein missing the N-terminal region of the protein. These changes result in a change in phenotype of the transformed plant.
[0134] Genes of interest are reflective of the commercial markets and interests of those involved in the development of the crop. Crops and markets of interest change, and as developing nations open up world markets, new crops and technologies will emerge also. In addition, as understanding of agronomic traits and characteristics such as yield and heterosis increase, the choice of genes for transformation will change accordingly. General categories of genes of interest include, for example, those genes involved in information, such as zinc fingers, those involved in communication, such as kinases, and those involved in housekeeping, such as heat shock proteins. More specific categories of transgenes, for example, include, but are not limited to, genes encoding important traits for agronomics, insect resistance, disease resistance,herbicide resistance, sterility, grain characteristics, and commercial products. Genes of interest include, generally, those involved in oil, starch, carbohydrate, or nutrient metabolism as well as those affecting seed size, sucrose loading, and the like.
[0135] More specific categories, for example, include, but are not limited to, genes encoding important traits for agronomics, insect resistance, disease resistance, herbicide resistance, sterility, grain or seed characteristics, and commercial products. Genes of interest include, generally, those involved in oil, starch, carbohydrate, or nutrient metabolism as well as those affecting seed size, plant development, plant growth regulation, and yield improvement. Plant development and growth regulation also refer to the development and growth regulation of various parts of a plant, such as the flower, seed, root, leaf and shoot.
[0136] Other commercially desirable traits are genes and proteins conferring cold, heat, salt, and drought resistance.
[0137] In certain examples, the present disclosure contemplates the transformation of a recipient cell with more than one advantageous gene. Two or more genes can be supplied in a single transformation event using either distinct gene-encoding vectors, or a single vector incorporating two or more gene coding sequences. Any two or more genes of any description, such as those conferring herbicide, insect, disease (viral, bacterial, fungal, and nematode), or drought resistance, oil quantity and quality, or those increasing yield or nutritional quality may be employed as desired.
[0138] Recombinant DNA constructs comprising an isolated nucleic acid fragment comprising of the targets disclosed herein. In some examples, this disclosure also concerns a recombinant DNA construct comprising a genomic region of interest of the nucleotide sequence set forth in Table 1.
[0139] In some examples, this disclosure concerns a recombinant DNA construct comprising at least one heterologous nucleic acid fragment operably linked to any promoter, or combination of promoter elements, of the present disclosure. Recombinant DNA constructs can be constructed by operably linking the nucleic acid fragment of the disclosure or a fragment that is substantially similar and functionally equivalent to any portion of the nucleotide sequence set forth in Table 1 to a heterologous nucleic acid fragment. Any heterologous nucleic acid fragment can be used to practice the disclosure. The selection will depend upon the desired application or phenotype to be achieved. The various nucleic acid sequences can be manipulated so as to provide for thenucleic acid sequences in the proper orientation. It is believed that various combinations of promoter elements as described herein may be useful in practicing the present disclosure.
[0140] In some examples, this disclosure concerns host cells comprising either the recombinant DNA constructs of the disclosure as described herein or isolated polynucleotides of the disclosure as described herein. Examples of host cells which can be used to practice the disclosure include, but are not limited to, yeast, bacteria, and plants.
[0141] Plasmid vectors comprising the instant recombinant DNA construct can be constructed. The choice of plasmid vector is dependent upon the method that will be used to transform host cells. The skilled artisan is well aware of the genetic elements that must be present on the plasmid vector in order to successfully transform, select and propagate host cells containing the chimeric gene.V. Gene editing
[0142] Site-Directed Nuclease (SDN) genome editing may be facilitated through the induction of a double-stranded break (DSB) or single-strand break, at a predetermined location near the desired alteration by a range of different DNA binding systems. The goal of SDN technology is to take advantage of the targeted DNA break and the host’s natural repair mechanisms to introduce specific small changes at the site of the DNA break.
[0143] DSBs can be induced using any DSB-inducing agent available, including, but not limited to, Meganucleases, Zinc-Finger Nucleases (ZFNs) and Transcription Activator Like Effector Nucleases (TALENs), Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)- associated proteins (CRISPR / Cas), guided cpfl endonuclease systems, and the like. In particular examples, the introduction of a DSB can be combined with the introduction of a polynucleotide modification template.
[0144] The process is divided into three categories, depending on the nature of the desired edit, SDN (Site-Directed Nuclease) 1, SDN 2 and SDN 3. SDN 1 involves making targeted cuts in the DNA without the introduction of a polynucleotide modification template. The cell’s natural repair mechanisms, typically non-homologous end joining (NHEJ), repair the break, often resulting in small insertions or deletions (indels) at the cut site. The result is a targeted, nonspecific genetic mutation. An SDN 2 edit involves using a polynucleotide modification template along with the nuclease. The cell uses this template to repair the break via homology-directed repair (HDR), allowing for precise modifications such as small insertions, deletions, orsubstitutions. The result is a targeted and predetermined mutation. SDN3 also involves a polynucleotide modification template but is designed to introduce larger segments of DNA (i.e. a gene or promoter). The result is the integration of that DNA sequence into the genome.
[0145] A polynucleotide modification template can be introduced into a cell by any method known in the art, such as, but not limited to, transient introduction methods, transfection, electroporation, microinjection, particle mediated delivery, topical application, whiskers mediated delivery, delivery via cell-penetrating peptides, or mesoporous silica nanoparticle (MSN)-mediated direct delivery.
[0146] The polynucleotide modification template can be introduced into a cell as a single stranded polynucleotide molecule, a double stranded polynucleotide molecule, or as part of a circular DNA (vector DNA). The polynucleotide modification template can also be tethered to the guide RNA and / or the Cas endonuclease. Tethered DNAs can allow for co-localizing target and template DNA, useful in genome editing and targeted genome regulation, and can also be useful in targeting post-mitotic cells where function of endogenous HR machinery is expected to be highly diminished (Mali et al. 2013 Nature Methods Vol. 10: 957-963.) The polynucleotide modification template may be present transiently in the cell or it can be introduced via a viral repl icon.
[0147] A “modified nucleotide” or “edited nucleotide” refers to a nucleotide sequence of interest that comprises at least one alteration when compared to its non-modified nucleotide sequence. Such “alterations” include, for example: (i) replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, or (iv) any combination of (i) - (iii).
[0148] The term “polynucleotide modification template” includes a polynucleotide that comprises at least one nucleotide modification when compared to the nucleotide sequence to be edited. A nucleotide modification can be at least one nucleotide substitution, addition or deletion. Optionally, the polynucleotide modification template can further comprise homologous nucleotide sequences flanking the at least one nucleotide modification, wherein the flanking homologous nucleotide sequences provide sufficient homology to the desired nucleotide sequence to be edited.
[0149] The process for editing a genomic sequence combining DSB and modification templates generally comprises: providing to a host cell, a DSB-inducing agent, or a nucleic acid encoding aDSB-inducing agent, that recognizes a target sequence in the chromosomal sequence and is able to induce a DSB in the genomic sequence, and at least one polynucleotide modification template comprising at least one nucleotide alteration when compared to the nucleotide sequence to be edited. The polynucleotide modification template can further comprise nucleotide sequences flanking the at least one nucleotide alteration, in which the flanking sequences are substantially homologous to the chromosomal region flanking the DSB.
[0150] The endonuclease can be provided to a cell by any method known in the art, for example, but not limited to transient introduction methods, transfection, microinjection, and / or topical application or indirectly via recombination constructs. The endonuclease can be provided as a protein or as a guided polynucleotide complex directly to a cell or indirectly via recombination constructs. The endonuclease can be introduced into a cell transiently or can be incorporated into the genome of the host cell using any method known in the art. In the case of a CRISPR-Cas system, uptake of the endonuclease and / or the guided polynucleotide into the cell can be facilitated with a Cell Penetrating Peptide (CPP) as described in WO2016073433 published May 12, 2016.
[0151] In addition to modification by a double strand break technology, modification of one or more bases without such double strand break are achieved using base editing technology, see e.g., Gaudelli et al., (2017) Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature 551 (768 l):464-471 ; Komor et al., (2016) Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage, Nature 533(7603):420-4.
[0152] These fusions contain dCas or Cas nickase and a suitable deaminase, and they can convert e.g., cytosine to uracil without inducing double-strand break of the target DNA. Uracil is then converted to thymine through DNA replication or repair. Improved base editors that have targeting flexibility and specificity are used to edit endogenous locus to create target variations and improve grain yield. Similarly, adenine base editors enable adenine to inosine change, which is then converted to guanine through repair or replication. Thus, targeted base changes i.e., OG to T’A conversion and A»T to G’C conversion at one or more locations made using appropriate site-specific base editors.
[0153] In some examples, base editing is a genome editing method that enables direct conversion of one base pair to another at a target genomic locus without requiring double-stranded DNA breaks (DSBs), homology-directed repair (HDR) processes, or external donor DNA templates. Inan embodiment, base editors include (i) a catalytically impaired CRISPR-Cas mutant that are mutated such that one of their nuclease domains cannot make DSBs; (ii) a single-strand-specific cytidine / adenine deaminase that converts C to U or A to G within an appropriate nucleotide window in the single- stranded DNA bubble created by Cas; (iii) a uracil glycosylase inhibitor (UGI) that impedes uracil excision and downstream processes that decrease base editing efficiency and product purity; and (iv) nickase activity to cleave the non-edited DNA strand, followed by cellular DNA repair processes to replace the G-containing DNA strand.
[0154] As used herein, a “genomic region” is a segment of a chromosome in the genome of a cell that is present on either side of the target site or, alternatively, also comprises a portion of the target site. The genomic region can comprise at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5- 45, 5- 50, 5-55, 5-60, 5-65, 5- 70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5-200, 5-300, 5-400, 5- 500, 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, 5-1200, 5-1300, 5-1400, 5-1500, 5-1600, 5- 1700, 5-1800, 5-1900, 5-2000, 5-2100, 5-2200, 5-2300, 5-2400, 5-2500, 5-2600, 5-2700, 5-2800. 5-2900, 5-3000, 5-3100 or more bases such that the genomic region has sufficient homology to undergo homologous recombination with the corresponding region of homology.
[0155] TAL effector nucleases (TALEN) are a class of sequence-specific nucleases that can be used to make double-strand breaks at specific target sequences in the genome of a plant or other organism. (Miller et al. (2011) Nature Biotechnology 29: 143-148).
[0156] Endonucleases are enzymes that cleave the phosphodiester bond within a polynucleotide chain. Endonucleases include restriction endonucleases, which cleave DNA at specific sites without damaging the bases, and meganucleases, also known as homing endonucleases (HEases), which like restriction endonucleases, bind and cut at a specific recognition site, however the recognition sites for meganucleases are typically longer, about 18 bp or more (patent application PCT / US12 / 30061, filed on March 22, 2012). Meganucleases have been classified into four families based on conserved sequence motifs, the families are the LAGLID ADG, GIY-YIG, H- N-H, and His-Cys box families. These motifs participate in the coordination of metal ions and hydrolysis of phosphodiester bonds. HEases are notable for their long recognition sites, and for tolerating some sequence polymorphisms in their DNA substrates. The naming convention for meganuclease is similar to the convention for other restriction endonuclease. Meganucleases are also characterized by prefix F-, I-, or PI- for enzymes encoded by free-standing ORFs, introns, and inteins, respectively. One step in the recombination process involves polynucleotidecleavage at or near the recognition site. The cleaving activity can be used to produce a doublestrand break. For reviews of site-specific recombinases and their recognition sites, see, Sauer (1994) Curr Op Biotechnol 5:521-7; and Sadowski (1993) FASEB 7:760-7. In some examples the recombinase is from the Integrase or Resolvase families.
[0157] Zinc finger nucleases (ZFNs) are engineered double-strand break inducing agents comprised of a zinc finger DNA binding domain and a double-strand-break-inducing agent domain. Recognition site specificity is conferred by the zinc finger domain, which typically comprising two, three, or four zinc fingers, for example having a C2H2 structure, however other zinc finger structures are known and have been engineered. Zinc finger domains are amenable for designing polypeptides which specifically bind a selected polynucleotide recognition sequence. ZFNs include an engineered DNA-binding zinc finger domain linked to a non-specific endonuclease domain, for example nuclease domain from a Type Ils endonuclease such as Fokl. Additional functionalities can be fused to the zinc-finger binding domain, including transcriptional activator domains, transcription repressor domains, and methylases. In some examples, dimerization of nuclease domain is required for cleavage activity. Each zinc finger recognizes three consecutive base pairs in the target DNA. For example, a 3 -finger domain recognized a sequence of 9 contiguous nucleotides, with a dimerization requirement of the nuclease, two sets of zinc finger triplets are used to bind an 18-nucleotide recognition sequence.
[0158] Genome editing using DSB-inducing agents, such as Cas-gRNA complexes, has been described, for example in U.S. Patent Application US 2015-0082478 Al, published on March 19, 2015, WO2015 / 026886 Al, published on February 26, 2015, WO2016007347, published on January 14, 2016, and WO201625131, published on February 18, 2016, all of which are incorporated by reference herein.
[0159] The term “Cas gene” herein refers to a gene that is generally coupled, associated or close to, or in the vicinity of flanking CRISPR loci in bacterial systems. The terms “Cas gene”, “CRISPR-associated (Cas) gene” are used interchangeably herein. The term “Cas endonuclease” herein refers to a protein encoded by a Cas gene. A Cas endonuclease herein, when in complex with a suitable polynucleotide component, is capable of recognizing, binding to, and optionally nicking or cleaving all or part of a specific DNA target sequence. A Cas endonuclease described herein comprises one or more nuclease domains. Cas endonucleases of the disclosure includes those having a HNH or HNH-like nuclease domain and / or a RuvC or RuvC-like nucleasedomain. Cas endonucleases that have been described include, but are not limited to, for example: Cas9, Casl2f (Cas-alpha, Casl4), Casl21 (Cas-beta), Casl2a (Cpfl), Casl2b (a C2cl protein), Cas 13 (a C2c2 protein), Cas 12c (a C2c3 protein), Cas 12d, Casl2e, Cas 12g, Casl2h, Casl2i, Casl2j, Casl2k, Cas3, Cas3-HD, Cas 5, Cas6, Cas7, Cas8, CaslO, or combinations or complexes of these. In some aspects, the methods and compositions described herein can utilize transposon-associated TnpB, a programmable RNA-guided DNA endonuclease.
[0160] As used herein, the terms “guide polynucleotide / Cas endonuclease complex”, “guide polynucleotide / Cas endonuclease system”, “ guide polynucleotide / Cas complex”, “guide polynucleotide / Cas system”, “guided Cas system” are used interchangeably herein and refer to at least one guide polynucleotide and at least one Cas endonuclease that are capable of forming a complex, wherein said guide polynucleotide / Cas endonuclease complex can direct the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single or double strand break) the DNA target site. A guide polynucleotide / Cas endonuclease complex herein can comprise Cas protein(s) and suitable polynucleotide component(s) of any of the four known CRISPR systems (Horvath and Barrangou, 2010, Science 327: 167-170) such as a type I, II, or III CRISPR system. A Cas endonuclease unwinds the DNA duplex at the target sequence and optionally cleaves at least one DNA strand, as mediated by recognition of the target sequence by a polynucleotide (such as, but not limited to, a crRNA or guide RNA) that is in complex with the Cas protein. Such recognition and cutting of a target sequence by a Cas endonuclease typically occurs if the correct protospacer-adjacent motif (PAM) is located at or adjacent to the 3' end of the DNA target sequence. Alternatively, a Cas protein herein may lack DNA cleavage or nicking activity but can still specifically bind to a DNA target sequence when complexed with a suitable RNA component. (See also U.S. Patent Application US 2015-0082478 Al, published on March 19, 2015, and US 2015-0059010 Al, published on February 26, 2015, both are hereby incorporated in its entirety by reference).
[0161] A guide polynucleotide / Cas endonuclease complex can cleave one or both strands of a DNA target sequence. A guide polynucleotide / Cas endonuclease complex that can cleave both strands of a DNA target sequence typically comprise a Cas protein that has all of its endonuclease domains in a functional state (e.g., wild type endonuclease domains or variants thereof retaining some or all activity in each endonuclease domain). Non-limiting examples ofCas nickases suitable for use herein are disclosed in U.S. Patent Appl. Publ. No. 2014 / 0189896, which is incorporated herein by reference.
[0162] Other Cas endonuclease systems have been described in PCT patent applications PCT / US 16 / 32073, fded May 12, 2016, and PCT / US 16 / 32028 filed May 12, 2016, both applications incorporated herein by reference.
[0163] “Cas9” (formerly referred to as Cas5, Csnl, or Csxl2) herein refers to a Cas endonuclease of a type II CRISPR system that forms a complex with a crNucleotide and a tracrNucleotide, or with a single guide polynucleotide, for specifically recognizing and cleaving all or part of a DNA target sequence. Cas9 protein comprises a RuvC nuclease domain and an HNH (H-N-H) nuclease domain, each of which can cleave a single DNA strand at a target sequence (the concerted action of both domains leads to DNA double-strand cleavage, whereas activity of one domain leads to a nick). In general, the RuvC domain comprises subdomains I, II and III, where domain I is located near the N-terminus of Cas9 and subdomains II and III are located in the middle of the protein, flanking the HNH domain (Hsu et al, Cell 157: 1262-1278). A type II CRISPR system includes a DNA cleavage system utilizing a Cas9 endonuclease in complex with at least one polynucleotide component. For example, a Cas9 can be in complex with a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). In another example, a Cas9 can be in complex with a single guide RNA.
[0164] Any guided endonuclease can be used in the methods disclosed herein. Such endonucleases include but are not limited to Cas9 and Cpfl endonucleases. Many endonucleases have been described to date that can recognize specific PAM sequences (see for example -Jinek et al. (2012) Science 337 p 816-821, PCT patent applications PCT / US 16 / 32073, filed May 12, 2016, and PCT / US 16 / 32028 filed May 12, 2016, and Zetsche B et al. 2015. Cell 163, 1013) and cleave the target DNA at a specific position. It is understood that based on the methods and embodiments described herein utilizing a guided Cas system one can now tailor these methods such that they can utilize any guided endonuclease system.
[0165] The terms “single guide RNA" and “sgRNA” are used interchangeably herein and relate to a synthetic fusion of two RNA molecules, a crRNA (CRISPR RNA) comprising a variable targeting domain (linked to a tracr mate sequence that hybridizes to a tracrRNA), fused to a tracrRNA (trans-activating CRISPR RNA). The single guide RNA can comprise a crRNA or crRNA fragment and a tracrRNA or tracrRNA fragment of the type II CRISPR / Cas system thatcan form a complex with a type II Cas endonuclease, wherein said guide RNA / Cas endonuclease complex can direct the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single or double strand break) the DNA target site.
[0166] The terms “guide RNA / Cas endonuclease complex”, “guide RNA / Cas endonuclease system”, “ guide RNA / Cas complex”, “guide RNA / Cas system”, “gRNA / Cas complex”, “gRNA / Cas system”, “RNA-guided endonuclease” , “RGEN” are used interchangeably herein and refer to at least one RNA component and at least one Cas endonuclease that are capable of forming a complex , wherein said guide RNA / Cas endonuclease complex can direct the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single or double strand break) the DNA target site. A guide RNA / Cas endonuclease complex herein can comprise Cas protein(s) and suitable RNA component(s) of any of the four known CRISPR systems (Horvath and Barrangou, 2010, Science 327: 167-170) such as a type I, II, or III CRISPR system. A guide RNA / Cas endonuclease complex can comprise a Type II Cas9 endonuclease and at least one RNA component (e.g., a crRNA and tracrRNA, or a gRNA). (See also U.S. Patent Application US 2015-0082478 Al, published on March 19, 2015, and US 2015-0059010 Al, published on February 26, 2015, both are hereby incorporated in its entirety by reference).
[0167] The guide polynucleotide can be introduced into a cell transiently, as single stranded polynucleotide or a double stranded polynucleotide, using any method known in the art such as, but not limited to, particle bombardment, Agrobacterium transformation or topical applications. The guide polynucleotide can also be introduced indirectly into a cell by introducing a recombinant DNA molecule (via methods such as, but not limited to, particle bombardment or Agrobacterium transformation) comprising a heterologous nucleic acid fragment encoding a guide polynucleotide, operably linked to a specific promoter that is capable of transcribing the guide RNA in said cell. The specific promoter can be, but is not limited to, a RNA polymerase III promoter, which allow for transcription of RNA with precisely defined, unmodified, 5’ - and 3’-ends (DiCarlo et al., Nucleic Acids Res. 41 : 4336-4343; Ma et al., Mol. Ther. Nucleic Acids 3 :e!61) as described in W02016025131, published on February 18, 2016, incorporated herein in its entirety by reference.
[0168] The terms “target site”, “target sequence”, “target site sequence, ’’target DNA”, “target locus”, “genomic target site”, “genomic target sequence”, “genomic target locus” and “protospacer”, are used interchangeably herein and refer to a polynucleotide sequence such as, but not limited to, a nucleotide sequence on a chromosome, episome, or any other DNA molecule in the genome (including chromosomal, chloroplastic, mitochondrial DNA, plasmid DNA) of a cell, at which a guide polynucleotide / Cas endonuclease complex can recognize, bind to, and optionally nick or cleave . The target site can be an endogenous site in the genome of a cell, or alternatively, the target site can be heterologous to the cell and thereby not be naturally occurring in the genome of the cell, or the target site can be found in a heterologous genomic location compared to where it occurs in nature. As used herein, terms “endogenous target sequence” and “native target sequence” are used interchangeable herein to refer to a target sequence that is endogenous or native to the genome of a cell and is at the endogenous or native position of that target sequence in the genome of the cell. Cells include, but are not limited to, human, non-human, animal, bacterial, fungal, insect, yeast, non-conventional yeast, and plant cells as well as plants and seeds produced by the methods described herein. An “artificial target site” or “artificial target sequence” are used interchangeably herein and refer to a target sequence that has been introduced into the genome of a cell. Such an artificial target sequence can be identical in sequence to an endogenous or native target sequence in the genome of a cell but located in a different position (i.e., a non-endogenous or non-native position) in the genome of a cell.
[0169] An “altered target site”, “altered target sequence”, “modified target site”, “modified target sequence” are used interchangeably herein and refer to a target sequence as disclosed herein that comprises at least one alteration when compared to non-altered target sequence. Such “alterations” include, for example: (i) replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, or (iv) any combination of (i) - (iii).
[0170] Methods for “modifying a target site” and “altering a target site” are used interchangeably herein and refer to methods for producing an altered target site.
[0171] The length of the target DNA sequence (target site) can vary, and includes, for example, target sites that are at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides in length. It is further possible that the target site can be palindromic, that is,the sequence on one strand reads the same in the opposite direction on the complementary strand. The nick / cleavage site can be within the target sequence or the nick / cleavage site could be outside of the target sequence. In another variation, the cleavage could occur at nucleotide positions immediately opposite each other to produce a blunt end cut or, in other cases, the incisions could be staggered to produce single-stranded overhangs, also called “sticky ends”, which can be either 5' overhangs, or 3' overhangs. Active variants of genomic target sites can also be used. Such active variants can comprise at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the given target site, wherein the active variants retain biological activity and hence are capable of being recognized and cleaved by an Cas endonuclease. Assays to measure the single or double-strand break of a target site by an endonuclease are known in the art and generally measure the overall activity and specificity of the agent on DNA substrates containing recognition sites.
[0172] A “protospacer adjacent motif’ (PAM) herein refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that is recognized (targeted) by a guide polynucleotide / Cas endonuclease system described herein. The Cas endonuclease may not successfully recognize a target DNA sequence if the target DNA sequence is not followed by a PAM sequence. The sequence and length of a PAM herein can differ depending on the Cas protein or Cas protein complex used. The PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides long.
[0173] The terms “targeting”, “gene targeting” and “DNA targeting” are used interchangeably herein. DNA targeting herein may be the specific introduction of a knock-out, edit, or knock-in at a particular DNA sequence, such as in a chromosome or plasmid of a cell. In general, DNA targeting can be performed herein by cleaving one or both strands at a specific DNA sequence in a cell with an endonuclease associated with a suitable polynucleotide component. Such DNA cleavage, if a double-strand break (DSB), can prompt NHEJ or HDR processes which can lead to modifications at the target site.
[0174] A targeting method herein can be performed in such a way that two or more DNA target sites are targeted in the method, for example. Such a method can optionally be characterized as a multiplex method. Two, three, four, five, six, seven, eight, nine, ten, or more target sites can be targeted at the same time in certain embodiments. A multiplex method is typically performed by a targeting method herein in which multiple different RNA components are provided, eachdesigned to guide a guide polynucleotide / Cas endonuclease complex to a unique DNA target site.
[0175] The terms “knock-out”, “gene knock-out” and “genetic knock-out” are used interchangeably herein. A knock-out represents a DNA sequence of a cell that has been rendered partially or completely inoperative by targeting with a Cas protein; such a DNA sequence prior to knock-out could have encoded an amino acid sequence, or could have had a regulatory function (e.g., promoter), for example. A knock-out may be produced by an indel (insertion or deletion of nucleotide bases in a target DNA sequence through NHEJ), or by specific removal of sequence that reduces or destroys the function of sequence at or near the targeting site.
[0176] The guide polynucleotide / Cas endonuclease system can be used in combination with a co-delivered polynucleotide modification template to allow for editing (modification) of a genomic nucleotide sequence of interest. (See also U.S. Patent Application US 2015-0082478 Al, published on March 19, 2015, and WO2015 / 026886 Al, published on February 26, 2015, both are hereby incorporated in its entirety by reference.)
[0177] The terms “knock-in”, “gene knock-in, “gene insertion” and “genetic knock-in” are used interchangeably herein. A knock-in represents the replacement or insertion of a DNA sequence at a specific DNA sequence in cell by targeting with a Cas protein (by HR, wherein a suitable donor DNA polynucleotide is also used). Examples of knock-ins are a specific insertion of a heterologous amino acid coding sequence in a coding region of a gene, or a specific insertion of a transcriptional regulatory element in a genetic locus.
[0178] Various methods and compositions can be employed to obtain a cell or organism having a polynucleotide of interest inserted in a target site for a Cas endonuclease. Such methods can employ homologous recombination to provide integration of the polynucleotide of Interest at the target site. In one method provided, a polynucleotide of interest is provided to the organism cell in a donor DNA construct. As used herein, “donor DNA” is a DNA construct that comprises a polynucleotide of Interest to be inserted into the target site of a Cas endonuclease. The donor DNA construct further comprises a first and a second region of homology that flank the polynucleotide of Interest. The first and second regions of homology of the donor DNA share homology to a first and a second genomic region, respectively, present in or flanking the target site of the cell or organism genome. By “homology” is meant DNA sequences that are similar. For example, a “region of homology to a genomic region” that is found on the donor DNA is aregion of DNA that has a similar sequence to a given “genomic region” in the cell or organism genome. A region of homology can be of any length that is sufficient to promote homologous recombination at the cleaved target site. For example, the region of homology can comprise at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5- 50, 5-55, 5-60, 5-65, 5- 70, 5-75, 5-80, 5- 85, 5-90, 5-95, 5-100, 5-200, 5-300, 5-400, 5-500, 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, 5-1200, 5-1300, 5-1400, 5-1500, 5-1600, 5-1700, 5-1800, 5-1900, 5-2000, 5-2100, 5-2200, 5- 2300, 5-2400, 5-2500, 5-2600, 5-2700, 5-2800, 5-2900, 5-3000, 5-3100 or more bases in length such that the region of homology has sufficient homology to undergo homologous recombination with the corresponding genomic region. “Sufficient homology” indicates that two polynucleotide sequences have sufficient structural similarity to act as substrates for a homologous recombination reaction. The structural similarity includes overall length of each polynucleotide fragment, as well as the sequence similarity of the polynucleotides. Sequence similarity can be described by the percent sequence identity over the whole length of the sequences, and / or by conserved regions comprising localized similarities such as contiguous nucleotides having 100% sequence identity, and percent sequence identity over a portion of the length of the sequences.
[0179] The amount of sequence identity shared by a target and a donor polynucleotide can vary and includes total lengths and / or regions having unit integral values in the ranges of about 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, 450-900 bp, 500-1000 bp, 600-1250 bp, 700-1500 bp, 800-1750 bp, 900-2000 bp, 1-2.5 kb, 1.5-3 kb, 2-4 kb, 2.5-5 kb, 3-6 kb, 3.5-7 kb, 4-8 kb, 5-10 kb, or up to and including the total length of the target site. These ranges include every integer within the range, for example, the range of 1-20 bp includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 and 20 bps. The amount of homology can also be described by percent sequence identity over the full aligned length of the two polynucleotides which includes percent sequence identity of about at least 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. Sufficient homology includes any combination of polynucleotide length, global percent sequence identity, and optionally conserved regions of contiguous nucleotides or local percent sequence identity, for example sufficient homology can be described as a region of 75-150 bp having at least 80% sequence identity to a region of the target locus. Sufficient homology can also be described by the predicted ability of twopolynucleotides to specifically hybridize under high stringency conditions, see, for example, Sambrook et al., (1989) Molecular Cloning: A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., Eds (1994) Current Protocols, (Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.); and, Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology— Hybridization with Nucleic Acid Probes, (Elsevier, New York).
[0180] The structural similarity between a given genomic region and the corresponding region of homology found on the donor DNA can be any degree of sequence identity that allows for homologous recombination to occur. For example, the amount of homology or sequence identity shared by the “region of homology” of the donor DNA and the “genomic region” of the organism genome can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity, such that the sequences undergo homologous recombination
[0181] The region of homology on the donor DNA can have homology to any sequence flanking the target site. While in some embodiments the regions of homology share significant sequence homology to the genomic sequence immediately flanking the target site, it is recognized that the regions of homology can be designed to have sufficient homology to regions that may be further 5' or 3' to the target site. In still other embodiments, the regions of homology can also have homology with a fragment of the target site along with downstream genomic regions. In one embodiment, the first region of homology further comprises a first fragment of the target site and the second region of homology comprises a second fragment of the target site, wherein the first and second fragments are dissimilar.
[0182] As used herein, “homologous recombination” includes the exchange of DNA fragments between two DNA molecules at the sites of homology.
[0183] Further uses for guide RNA / Cas endonuclease systems have been described (See U.S. Patent Application US 2015-0082478 Al, published on March 19, 2015, WO2015 / 026886 Al, published on February 26, 2015, US 2015-0059010 Al, published on February 26, 2015, US application 62 / 023246, filed on July 07, 2014, and US application 62 / 036,652, filed on August 13, 2014, all of which are incorporated by reference herein) and include but are not limited to modifying or replacing nucleotide sequences of interest (such as a regulatory elements), insertion of polynucleotides of interest, gene knock-out, gene-knock in, modification of splicingsites and / or introducing alternate splicing sites, modifications of nucleotide sequences encoding a protein of interest, amino acid and / or protein fusions, and gene silencing by expressing an inverted repeat into a gene of interest.
[0184] Methods for transforming dicots, primarily by use of Agrohacterium tumefaciens, and obtaining transgenic plants have been published, among others, for cotton (U.S. Patent No. 5,004,863, U.S. Patent No. 5,159, 135); soybean (U.S. Patent No. 5,569,834, U.S. Patent No. 5,416,011); Brassica (U.S. Patent No. 5,463,174); peanut (Cheng et al., Plant Cell Rep. 15:653 657 (1996), McKently et al., Plant Cell Rep. 14:699 703 (1995)); papaya (Ling et al., Bio / technology 9:752 758 (1991)); and pea (Grant et al., Plant Cell Rep. 15:254 258 (1995)). For a review of other commonly used methods of plant transformation see Newell, C.A., Mol. Biotechnol. 16:53 65 (2000). One of these methods of transformation uses Agrobacterium rhizogenes (Tepfler, M. and Casse-Delbart, F., Microbiol. Sci. 4:24 28 (1987)). Transformation of soybeans using direct delivery of DNA has been published using PEG fusion (PCT Publication No. WO 92 / 17598), electroporation (Chowrira et al., Mol. Biotechnol. 3: 17 23 (1995); Christou et al., Proc. Natl. Acad. Sci. U.S.A. 84:3962 3966 (1987)), microinjection, or particle bombardment (McCabe et al., Biotechnology 6:923-926 (1988); Christou et al., Plant Physiol. 87:671 674 (1988)).
[0185] There are a variety of methods for the regeneration of plants from plant tissues. The particular method of regeneration will depend on the starting plant tissue and the particular plant species to be regenerated. The regeneration, development, and cultivation of plants from single plant protoplast transformants or from various transformed explants is well known in the art (Weissbach and Weissbach, Eds.; In Methods for Plant Molecular Biology; Academic Press, Inc.: San Diego, CA, 1988). This regeneration and growth process typically includes the steps of selection of transformed cells, culturing those individualized cells through the usual stages of embryonic development or through the rooted plantlet stage. Transgenic embryos and seeds are similarly regenerated. The resulting transgenic rooted shoots are thereafter planted in an appropriate plant growth medium such as soil. Preferably, the regenerated plants are selfpollinated to provide homozygous transgenic plants. Otherwise, pollen obtained from the regenerated plants is crossed to seed-grown plants of agronomically important lines. Conversely, pollen from plants of these important lines is used to pollinate regenerated plants. A transgenicplant of the present disclosure containing a desired polypeptide is cultivated using methods well known to a skilled artisan.EXAMPLES
[0186] The present disclosure is further defined in the following Examples. From the above discussion and these Examples, one skilled in the art can ascertain the essential characteristics of this disclosure, and without departing from the spirit and scope thereof, can make various changes and modifications of the disclosure to adapt it to various usages and conditions. Thus, various modifications of the disclosure in addition to those shown and described herein will be apparent to those skilled in the art from the foregoing description. Such modifications are also intended to fall within the scope of the appended claims.
[0187] The disclosure of each reference set forth herein is incorporated herein by reference in its entirety.Example 1 : Target sequences, guide RNAs and construct for D8 edited variants
[0188] Pioneer maize line PH1V69 wildtype genomic (SEQ ID NO: 5) and protein (SEQ ID NO:6) sequences of Dwarf8 (D8) were derived from PHI V69 Chromosomes_v2 genome assembly.
[0189] The constructs used to create edits were assembled using standard protocols. Genomic edits at the D8 genomic loci were created using ZM-D8-CR10 and ZM-CR11 guide RNAs (SEQ ID NOs: 1 and 2, respectively) and target sequences (SEQ ID NOs:3 and 4, respectively) with respective PAM highlighted in underlined (Table 2). The ZM-D8-CR10 / ZM-D8-CR11 edit was expected to produce an in-frame deletion (dropout) of the DELLA motif (SEQ ID NO: 18) of the DELLA domain.
[0190] Table 2
[0191] A partial map of the D8 transcript sequence showing domains, regions and motifs according to Uniprot is shown in FIG. 1. Numbers above the domains and regions indicate corresponding amino acids at the protein level.
[0192] The constructs were delivered to maize explants via Agrobacterium-mediated transformation and resultant plants grown according to methods known in the art.Example 2: TO D8 edited variants
[0193] D8 variants were identified in multiple TO plants. Exemplary D8 variants are shown in Table 3 and additional examples of edited TO variants are listed in Table 4. The EGV.ZM011.342.1.1 and EGV. ZM011.342.1.20 variants in Table 3 corresponded to the expected deletion [-63bp], while other variants corresponded to deletions of larger sizes (e.g. EGV. ZM011.342.1.11 and EGV. ZM011.342.1.3), or deletion and nucleotide additions, (e.g. EGV. ZM011.342.1.18). Variant EGV. ZM011.342.1.10 corresponded to a 35bp deletion at ZM-D8- CR10 and a [-C] deletion at ZM-D8-CR11. The exemplary D8 variants (Table 4) were observed in a selection of TO plants that were crossed with WT PH1V69 plants and grown to T1 plants to study the phenotypic impact of different edits on plant height. No additional work was conducted on the TO variants listed in Table 4; therefore, T1 phenotypic data is not available.
[0194] Table 3: Exemplary edited variant descriptions and genomic modifications
[0195] Table 4: Additional TO edited variants
[0196] The variant descriptions shown in Tables 3 and 4 refer to the in / dels created with the ZM- D8-CR10 and ZM-CR11 guide RNAs and represent the number of base pairs at a position at or near the gRNA target sites that correspond to the wildtype D8 reference sequence (SEQ ID NO: 5). Precise locations of the in / dels are provided as nucleotide positions that correspond to the location of the first ATG codon of a representative wildtype D8 polynucleotide sequence represented by SEQ ID NO: 5. For example, the variant description “[-35bp;-C] edited allele” in Table 3 refers to the 35 bp deletion starting at position 33 to position 67 nucleotides, and a deletion of a cytosine (C) at position 130, of the wildtype D8 sequence represented by SEQ ID NO: 5 The in / del locations can also be readily determined from the wildtype D8 DNA sequence and variant DNA sequences provided within. For example, SEQ ID NO: 7 can be aligned with SEQ ID NO: 5 to indicate the specific in / dels represented by ZM-D8 EGV.ZM011.342.1.10.Example 3: Confirmation of T1 D8 edited variants
[0197] A T1 experiment was conducted to confirm the presence of mutations noted at TO in the progeny and to observe associated phenotypes. A number of plants heterozygous for the D8 mutations were selected based on Gene Edit Molecular Characterization (GEMC) data and plant height and primary ear height data (Svitashev et al., 2016. Nat Commim 7: 13274). The nature ofthe mutations found at TO were confirmed in the T1 plants, except for the EGV.342.1 .20 variant where only WT alleles were recovered in the T1 plants. EGV.342.1.20 variant plants were dropped from further analysis.
[0198] Table 5 shows the mean plant height of heterozygous T1 D8 variants measured 63 days after planting. One-way ANOVA analysis and grouping information using the Tukey method and 95% confidence was performed. Means that do not share a letter are significantly different.
[0199] The data show that plant height of variants EGV.ZM011.342.1.10 (corresponding to [- 35bp;-C] edit) and EGV.ZM011.342.1.1 (corresponding to a [-63bp] edit) showed statistically significant height differences from WT controls. Both variants also showed significantly different plant height between each other, with the heterozygous plants for the [-63bp] edit being about half the size of plants carrying the [-35bp;-C] allele.
[0200] Table 5: Plant height characterization of heterozygous T1 D8 variants
[0201] Selected plants for each of the different edits were sent for Sequencing by Synthesis (SbS) analysis which confirmed the gene edits. The SbS generated sequences of gene edited alleles for variants [-64bp;+3bp], [-65bp], [-64bp], [-35bp;-C], and [-63bp] were used and aligned to the WT PH1V69 map of D8 to build complete maps of the D8 edited allele sequences for each variant. The coding sequences were then translated in silica, starting from ATG to the first stop codon.
[0202] Polypeptide sequences of D8 gene edited alleles (deduced in silica based on the polynucleotide sequence of the edited variants) were then aligned against the wildtype ZM-D8 polypeptide sequence. A partial protein sequence alignment of the N-terminal region was created with Clustal Omega in Geneious (Dotmatics, Boston, MA) and is shown in FIG. 2. The DRM (SEQ ID NO:30; solid line box) was missing in the chimeric proteins resulting from [-35bp;-C] gene edited allele but present in other chimeric proteins resulting from [-63bp], [-64bp], [-64bp;+3bp] and [-65bp] gene edited alleles. The DELLA motif (SEQ ID NO: 18; dotted line box) was absent in the in-frame [-63bp] gene edited allele and all chimeric proteins resulting from the translation of the [-64bp], [-64bp;+3], and [-65bp] gene edited alleles. A summary of the genotypes and phenotypes of the D8 edited variants is shown in Table 6.
[0203] Table 6: Description of the D8-edited variants and their genotypic and phenotypic characterizations
[0204] Results in Table 6 show three variants, [-65bp], [-64bp], and [-64bp;+3bp], that were missing the DELLA motif and showed WT height phenotype, not a dwarf phenotype. None of the three variants with the WT phenotype have a downstream start site that is in-frame with ML The [-35bp;-C] variant that was missing both the DRM and DELLA motif showed an intermediate dwarf phenotype. In the [-35bp;-C] variant, the M53 start site is in-frame with ML The [-63bp] variant which has the DRM but not the DELLA motif and a severe dwarf phenotype is a gain of function mutation that may result in D8 not being degraded and / or may be blocking signaling through D9 which can no longer access the GA receptor.
[0205] The [-35bp;-C] gene edit does not lead to M106-D8 deletion protein with conserved GRAS domain. Instead, the [-35bp;-C] gene edit leads to chimeric protein starting at Ml and allowing translation reinitiation and production of a M53 truncated protein (see FIG. 1).
[0206] Results indicate that translation reinitiation in the [-35bp;-C] gene edit may result in the expression of a partial ZM-D8 protein missing the N-terminal region of the protein including the DRM and the DELLA motifs but exhibited an intermediate dwarf phenotype.Example 4: Confirmation of T2 D8 genome edited variants
[0207] T2 experiments were conducted to confirm the genotypes and phenotypes of the different D8 edited variants that were obtained in PH1V69 to understand the mode of action of D8 in generating intermediate phenotypes.
[0208] A number of heterozygous T1 plants were selfed to produce segregating T2 seeds based on confirmation of the phenotype observed with T1 plants. T2 seeds with different mutations were selected, except for the EGV. ZM011.342.1.1 (SEQ ID NO: 9) severe dwarf [-63bp] variant because selfing of T1 plants was not possible due to extreme dwarfism. T1 seed for the severe dwarf [-63bp] variant was included in the experiment as a reference. The goal of the experiment was to confirm the phenotype of T1 heterozygous plants and to observe the phenotype of homozygous plants for different mutations compared to WT plants.
[0209] Segregating null, heterozygous, and homozygous T2 plants for each mutation were selected from this new experiment using GEMC data, with the exception of the [-63bp] variant plants because only heterozygous and wildtype plants were identified. Plants were imaged in the automated greenhouse to get plant height data at different developmental timepoints. T2 plant heights of the D8 variants at 51 days after planting are graphed in FIG. 4. Individual standard deviations were used to calculate the intervals. Error bars are one standard error from the mean. Table 7 shows one way ANOVA analysis and grouping information of plant height data using the Tukey method and 95% confidence. Means that do not share a letter are significantly different. "Segregating null" or “null”, as used herein, refers to a plant that is part of a segregating population (e.g. T2 population after a transformation step) and does not carry a mutant allele, therefore resulting in a wildtype phenotype.
[0210] Table 7: Plant height characterization of T2 genome-edited variants of maize plants
[0211] The [-63bp] variant (SEQ ID NO: 9) showed a severe dwarf phenotype in the heterozygous plants compared to the WT plants (FIG. 3 and Table 6). The [-63bp] heterozygous variant mean plant height showed a 60.8% reduction in height compared to the [-63bp] wildtype mean plant height. Only heterozygous plants were identified and measured for this variant. Homozygous plants likely did not survive or were discarded as runts. The [-63bp] variant has an in-frame deletion of 21 amino acids including the DELLA motif. Because the deletion is inframe in the [-63bp] variant, there may not be a sufficient opportunity for translation reinitiation to occur because there is no spurious stop codon created. The deletion leads to a “WT-equivalent level” expression of a D8 variant protein missing the DELLA motif, creating a strong dwarf phenotype.
[0212] The [-35bp;-C] variant, with a downstream ATG codon (e.g. coding for M53) that is inframe with first ATG codon coding for Ml, showed an intermediate dwarf phenotype in the heterozygous and homozygous plants compared to the wildtype plants and the [-63bp] plants (FIG. 3 and Table 6). The homozygous variant ([-35bp;-C]-Hom) was statistically shorter than the heterozygous variant ([-35bp;-C]-Het) and both were statistically shorter than the wildtype ([- 35bp;-C]-WT). The [-35bp;-C]-Hom variant mean plant height showed a 47% reduction compared to the wildtype ([-35bp;-C]-WT) mean plant height. The [-35bp;-C]-Het variant meanplant height showed a 30.4% reduction compared to the wildtype ([-35bp;-C]-WT) mean plant height. The [-63bp] variant was statistically shorter than both the heterozygous and homozygous [-35bp;-C] variants.
[0213] The remaining variants, [-64bp;+3bp], [-65bp], and [-64bp], which lack a downstream ATG codon that is in-frame with the first ATG codon coding for Ml, showed a wildtype phenotype in the heterozygous and homozygous plants (FIG. 3 and Table 6). The [-64bp;+3bp], [-65bp], and [-64bp] heterozygous variants had no significant difference in mean plant height compared to the wildtype plants. The [-64bp;+3bp], [-65bp], and [-64bp] homozygous variants also had no significant difference in mean plant height.
[0214] The [-35bp;-C] gene edit results in a change of frame and introduction of a premature stop codon before the DELLA motif. This variant has a semi-dwarf phenotype suggesting that a truncated D8 protein lacking the DELLA motif is produced resulting in the semi-dwarf phenotype. This protein is likely produced through translation reinitiation using one of the downstream ATG codons of the D8 gene, for example M53.
[0215] The [-65bp], [-64bp] and [-64bp;+3bp] gene edits also result in a change of frame and the introduction of premature stop codons. However, these variants have WT phenotypes which indicate that they are unlikely to produce a truncated D8 protein through translation reinitiation using one of the downstream ATG codons of the D8 protein. The first downstream ATG codon in these variants code for either Ml 06 or Ml 84 in D8 WT protein. If these variants were able to produce a truncated D8 protein starting at Ml 06 or Ml 84 through translation reinitiation, it would have resulted in a semi-dwarf phenotype. One hypothesis is that the distance between the introduced stop codons and their relative frame to the next ATG in these variants may be a relevant factor in influencing the efficiency of the translation reinitiation process compared to the [-35bp;-C] variant.Example 5: Secondary phenotypes of D8 edited variants
[0216] Average leaf area (cm2) measurements were taken with automated greenhouse imaging for T2 plants with different D8 variants and zygosity levels. FIG. 4 shows leaf area at 51 days after planting following the same pattern as plant height. The [-63bp] variant showed a significant decrease in leaf area in the heterozygous plants compared to the WT. The [-35bp;-C] variant also showed a significant decrease in leaf area in the heterozygous and homozygousplants compared to the WT, but the decrease was intermediate of the decrease in the [-63bp] variant and WT, similar to what was observed with plant height. The [-64bp;+3bp], [-64bp], and [-65bp] variants showed a leaf area similar to WT for the heterozygous and homozygous plants, which was similar to what was observed with plant height.
[0217] Percent (%) green measurements were taken with automated greenhouse imaging for T2 plants with different D8 variants and zygosity levels at 51 days after planting. The % green data showed that the [-63bp] and [-35bp;-C] variant plants were greener at late stages of development compared to WT plants, while no significant changes were observed with [-64bp;+3bp], [-64bp], and [-65bp] variants compared to WT plants (FIG. 5).Example 6: Translation reinitiation determination using protoplast transient assay
[0218] Translation reinitiation was investigated using maize protoplast transformation and appropriate constructs involving expression of the edited D8 variants fused to the N-terminus of a luciferase reporter gene (FIG. 6). These variants have premature stop codons at different positions that are either in-frame or out-of-frame with a downstream methionine, summarized in Table 8. Constructs of the D8 variants were made with truncated D8 sequences, SEQ ID NOs: 48-56. The protoplast transient assay was conducted according to methods described in Cao, J. et al. "PEG-mediated transient gene expression and silencing system in maize mesophyll protoplasts: a valuable tool for signal transduction study in maize." Acta Physiologiae Plantarum 36 (2014): 1271-1281. Luciferase activity was measured with a luminescence reader, normalized and recorded as the relative reporter protein expression level. The reporter protein expression level is relative to a control protein (Renilla luciferase). The control cassette is the same for all constructs, serving as an internal control to normalize variations introduced, for example, by transfection (copy number), total protein level, and pipetting.
[0219] Constructs resulting in the [-35bp;-C] variant allowed for translation reinitiation and expression of the reporter gene was detected (Table 8). Constructs resulting in the [-64bp;+3bp], [-65bp], and [-64bp] variants did not re-initiate translation and expression of the reporter gene was significantly lower in the [-64bp;+3bp] and [-64bp] variants than in the [-35bp;-C] variant (Table 8). This transient protoplast assay helps understand how translation reinitiation may play a role in the phenotype exhibited by the [-35bp;-C] D8 variant.
[0220] Table 8: Transient protoplast assays for tested variant constructs
[0221] Constructs with additional D8 variants ZM-D8 666 SI1, ZM-D8 666 SO1, ZM- D8 666 SI2, and ZM-D8 666 SI3 were constructed with premature stop codons in similar positions as the [-35bp;-C] ], [-65bp], [-64bp;+3bp] and [-64bp] variants (see descriptions in Table 1) to test in the transient protoplast assay, summarized in Table 9.
[0222] Constructs with the ZM-D8 666 SI1 and ZM-D8 666 SI2 variants were in-frame with their respective downstream start codons, M53 and M106, and likely allowed for translation reinitiation. Expression of the reporter gene was detected with the ZM-D8 666 SI 1 and ZM- D8_666_SI2 variants.
[0223] Constructs with the ZM-D8_666_SI3 variant were in-frame with downstream Ml 84. However, expression of the reporter gene was not detected, indicating that translation reinitiation likely did not occur. Additional factors such as the strength of the Kozak sequence of the downstream start codon may affect whether translation reinitiation may occur.
[0224] Constructs with variant ZM-D8 666 SO1, which has a premature stop codon at the same position as the [-35bp;-C] variant but is out-of-frame with downstream M53 allowed for translation reinitiation and expression of the reporter gene was detected. This was a surprising result and suggests that variants with a premature stop codon at a similar position as in the [- 35bp;-C] variant, either in-frame or out-of-frame with M53, may allow for translation reinitiation.
[0225] Table 9: Transient protoplast assays for tested variant constructs
[0226] SEQ ID NOs provided in Tables 8 and 9 are truncated versions of the full-length variant sequences and are given a different SEQ ID NOs than the full-length sequences. The positions of premature stop codons are provided as the corresponding positions in the truncated wildtype sequence (ZM-D8_WT_trunc; SEQ ID NO: 48).Example 7: TO phenotype of the [-23bp1 D8 edited variant
[0227] The GV.ZM99TK.001.2 [-23bp] D8 edit (SEQ ID NO: 58) was isolated from a SDN1 event created with the ZM-D8-CR25 guide RNA (SEQ ID NO: 57). The edit has a 23 bp deletion of nucleotides 91 to 113 in the wildtype D8 sequence SEQ ID NO: 5 and resulted in a premature stop codon at the same position as in [-35bp;-C] variant. The introduced stop codon was out-of- frame with downstream M53. TO plants with the [-23bp] edit showed a semi-dwarf phenotype with a 26% reduction in plant height compared to the wildtype control plant height (Table 10). This indicates expression of a truncated D8 protein via translation reinitiation may have contributed to the phenotype. This stable data corroborates the observation with the ZM- D8 666 SO1 variant in Table 9 and indicates that variants with premature stop codon [TGA] introduced at nucleotides 114-116 of truncated wildtype D8 sequence SEQ ID NO: 48 may allow for translation reinitiation, whether the stop is in-frame or out-of-frame with downstream M53.
[0228] Table 10: TO plant height and ear height characterization of a D8 genome-edited variantExample 8: Investigating translation reinitiation at downstream start sites
[0229] Appropriate constructs are designed to introduce an in-frame stop codon or create an inframe stop codon due to one or more nucleotide edits between the first methionine (Ml) and a downstream ATG before the GRAS domain (e.g., between Ml and M201; FIG. 1). Resulting mutations may allow translation reinitiation to produce a truncated D8 protein with an appropriate repressor domain and create a reduced stature phenotype. Other constructs are designed to in-frame stop codon between Ml and M271 and between Ml and M280, both in the 5’ end of the GRAS domain, to determine if translation reinitiation occurs and a reduced stature phenotype results. These constructs would help determine if the LRI domain (SEQ ID NO: 22) plays a role in the activity of the truncated GRAS domain.
[0230] Additionally, constructs are designed to investigate any additional molecular conditions for translation reinitiation. These conditions can include but are not limited to determination of an optimal distance between the stop codon and the downstream reinitiation ATG and the presence of a strong Kozak sequence at the reinitiation ATG.Example 9: Genome editing strategies for systematic targeted introduction of pre-mature stop codons at genomic loci in plants to generate translation reinitiation of open reading frames in the plant genome
[0231] Translation reinitiation is a cellular process where ribosome recycling at termination sites is followed by translation reinitiation by the same ribosome machinery on the same mRNA but at a different ORF, often at a downstream location in the same mRNA transcript. The functional outcome of such translation reinitiation is often two protein products translated from the same mRNA. Translation initiation is a highly coordinated and complex process involving numerous eukaryotic initiation factors (elFs). Alternative translation mechanisms can result in similar functional outcomes to those observed during reinitiation - namely, the translation of a downstream coding region. However, these mechanisms do not rely on a preceding translation termination event. A defining feature of the termination-reinitiation process is the generation of two distinct polypeptide chains from a single RNA transcript (excluding start-stop upstream open reading frames, or uORFs), synthesized by the same ribosome.
[0232] Certain internal ribosome entry sites (IRESs), located upstream of coding regions, can bypass many or all of the canonical requirements for translation initiation. Some IRES elementsrequire minimal or no initiation factors to recruit ribosomes and are not restricted to the 5' untranslated region (UTR) for initiating translation. These IRESs can drive translation of a downstream coding region from an intergenic location, functionally resembling reinitiation. However, a key distinction is that IRES-mediated translation of the downstream coding region does not necessitate the same ribosome that translated the upstream region, nor does it require termination at the upstream open reading frame (ORF) stop codon.
[0233] Given the molecular framework of the translation reinitiations at one or more open reading frames in the genome, one can systematically design genome editing strategies to modify genomic loci (e.g., insertion or deletions) to introduce pre-mature stop codons and to insert or delete nucleotides to introduce or create secondary or downstream start codons. This genome editing strategy introduces genomic diversity by modulating the gene expression / translation controls. A method of targeted creation of an in-frame or out-of-frame translation reinitiation site at a plant genomic locus encoding a gene of interest, the method includes for example, deleting, inserting, or modifying one or more nucleotides at a genomic sequence comprising an open reading frame such that a premature stop codon after the codon encoding the first or original starting Met of the open reading frame; and creating a translation reinitiation start site that is downstream of the premature stop codon such that the ribosomal translational machinery is capable of initiating translation at the reinitiation start site, wherein the reinitiation start site is positioned at a sufficient distance from the first or the original Met to provide the translational machinery to be associated with and recruit the necessary components to reinitiate translation. In certain embodiments, the gene of interest confers a dominant phenotype. In certain embodiments, the translation reinitiation site results in the production of a transcript and / or polypeptide that confers an agronomic trait of interest. In certain embodiments, the translation reinitiation results in a semi-dwarf phenotype.
Claims
1. CLAIMS1. A genome-edited polynucleotide encoding a polypeptide comprising an amino acid sequence of SEQ ID NO: 59, the polypeptide sequence not comprising a functional DELLA motif set forth in SEQ ID NO: 18.
2. The polynucleotide of claim 1, wherein the polynucleotide sequence is at least 95% identical to a polynucleotide sequence comprising SEQ ID NO: 7.
3. The polynucleotide of claim 1, wherein the polynucleotide sequence is at least 95% identical to a sequence comprising SEQ ID NO: 58.
4. A genome-edited plant cell comprising the polynucleotide of any one of claims 1-3.
5. The plant cell of claim 4 is a maize plant cell.
6. The plant cell of claims 4 and 5, wherein the plant cell produces an mRNA transcript that results in generation of a translation reinitiated polypeptide.
7. A genome-edited plant comprising the polynucleotide of any one of claims 1-3.
8. The plant of claim 7, wherein the plant produces an mRNA transcript that results in generation of a translation reinitiated polypeptide.
9. The plant of claims 7 and 8, wherein the plant exhibits reduced plant height compared to a control plant not comprising the targeted genome edit modification that results in a translation reinitiation at a downstream ATG codon.
10. The plant of any one of claims 7-9, wherein the plant is a maize inbred plant and exhibits a semi-dwarf phenotype, wherein the genome edit is in a heterozygous state.
11. A genome-edited seed comprising the genome-edited polynucleotide of any one of claims 1- 3.
12. The seed of claim 11, wherein the targeted genome edit modification results in translation reinitiation at a downstream ATG codon.
13. The seed of claims 11 and 12, wherein the seed is from a genome-edited maize plant that exhibits a semi-dwarf phenotype.
14. A method of reducing height of a maize plant, the method comprising introducing a targeted modification into an endogenous dwarf8 gene in a maize plant, plant cell or seed thereof, wherein the targeted modification introduces or creates a premature stop codon between a first or original starting ATG codon of the open reading frame of the dwarfB gene and a subsequent downstream ATG codon, whereina) the downstream ATG codon is located upstream or within the 5’ region of a GRAS domain coding sequence; b) the first or original starting ATG codon is in-frame with the subsequent downstream ATG codon; and c) the premature stop codon is in-frame or out-of-frame with the downstream ATG codon; whereby a modified plant is generated, the modified plant exhibiting reduced height relative to a control plant not comprising the targeted modification, wherein the reduced height phenotype results from translation reinitiation at the downstream ATG codon.
15. The method of claim 14, wherein prior to modification the dwarf8 gene encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6.
16. The method of claims 14 and 15, wherein the targeted modification is introduced using a CRISPR-associated nuclease and at least one guide RNA.
17. The method of any one of claims 14-16, wherein the targeted modification comprises an addition, deletion or a substitution of at least one nucleotide that introduces or creates the premature stop codon.
18. The method of any one of claims 14-17, wherein the targeted modification results in translation reinitiation from the downstream ATG codon, thereby generating a truncated dwarf8 protein.
19. The method of any one of claims 14-18, wherein the premature stop codon is introduced between the first methionine (Ml) of SEQ ID NO: 6 and a second methionine selected from the group consisting of methionine at position 53 (M53), methionine at position 65 (M65), methionine at position 67 (M67), methionine at position 69 (M69), methionine at position 106 (M106), methionine at position 184 (M184), methionine at position 201 (M201), methionine at position 271 (M271), and methionine at position 280 (M280) of SEQ ID NO: 6, wherein the second methionine is in-frame with the first methionine and the premature stop codon is in-frame or out-of-frame with the second methionine.
20. The method of any one of claims 14-19, wherein the targeted modification introduces or creates an in-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.21 . The method of any one of claims 14-19, wherein the targeted modification introduces or creates an out-of-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
22. The method of any one of claims 14-21, wherein the targeted modification is introduced into a maize plant cell genome, and the method further comprises regenerating a maize plant from the modified plant cell thereby producing a modified plant, wherein the modified plant exhibits a semi-dwarf phenotype due to translation reinitiation at the downstream ATG codon.
23. A modified maize plant, plant cell, or seed comprising a targeted genome-edited modification at an endogenous dwarf8 gene in a maize plant, plant cell or seed thereof, wherein the targeted modification introduces or creates a premature stop codon between a first or original starting ATG codon of the open reading frame of the dwarf8 gene and a subsequent downstream ATG codon, wherein a) the downstream ATG codon is located upstream of or within the 5’ region of a GRAS domain coding sequence, b) the translation initiation ATG codon is in-frame with the subsequent downstream ATG codon, and c) the premature stop codon is in-frame or out-of-frame with the downstream ATG codon, wherein the modified plant exhibits reduced height relative to a control plant not comprising the targeted genome-edited modification.
24. The modified maize plant, or plant cell, or seed thereof of claim 23, comprising a targeted modification that introduces or creates an in-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
25. The modified maize plant, or plant cell, or seed thereof of claim 23, comprising a targeted modification that introduces or creates an out-of-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
26. A modified maize plant, plant cell or seed thereof of any one of claims 23-25, comprising SEQ ID NO:7 or SEQ ID NO: 58.
27. The modified maize plant of any one of claims 23-26, wherein the plant exhibits a semidwarf phenotype due to translation reinitiation at the downstream ATG codon.
28. A guide polynucleotide molecule that targets an endogenous gene of a plant cell, wherein the gene comprises a polynucleotide that encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6.
29. A plant cell comprising the guide polynucleotide of claim 28.
30. The guide polynucleotide of claim 28, wherein the guide polynucleotide is provided with a Cas endonuclease at an endogenous dwarjB gene.
31. The plant cell of claim 29, wherein the plant cell is a maize plant cell.
32. A plant cell comprising a targeted modification at an endogenous gene that encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6, wherein the targeted modification comprises a premature stop codon between a first or original starting ATG codon of the open reading frame of the endogenous gene and a subsequent downstream ATG codon, and wherein a) the downstream ATG codon is located upstream of or within the 5’ region of a GRAS domain coding sequence, b) the first or original starting ATG codon is in-frame with the subsequent downstream ATG codon; and c) the premature stop codon is in-frame or out-of-frame with the downstream ATG codon.
33. The plant cell of claim 32, wherein the targeted modification comprises an addition, deletion or substitution of at least one nucleotide that introduces or creates the premature stop codon.
34. The plant cell of claims 32 and 33, wherein the targeted modification results in translation reinitiation at the downstream ATG codon, thereby generating a truncated dwarf8 protein.
35. The plant cell of any one of claims 32-34, wherein the premature stop codon is introduced between the first methionine (Ml) of SEQ ID NO: 6 and a second methionine selected from the group consisting of methionine at position 53 (M53), methionine at position 65 (M65), methionine at position 67 (M67), methionine at position 69 (M69), methionine at position 106 (Ml 06), methionine at position 184 (Ml 84), methionine at position 201 (M201), methionine at position 271 (M271), and methionine at position 280 (M280) of SEQ ID NO: 6, wherein the second methionine is in-frame with the first methionine and the premature stop codon is in-frame or out-of-frame with the second methionine.
36. The plant cell of any one of claims 32-35, wherein the targeted modification comprises an inframe stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
37. The plant cell of any one of claims 32-35, wherein the targeted modification comprises an out-of-frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
38. The plant cell of any one of claims 32-37, wherein the plant cell is a maize plant cell.
39. The plant cell of any one of claims 32-38, comprising SEQ ID NO:7 or SEQ ID NO: 58.
40. A plant comprising a targeted modification at an endogenous gene encoding a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6, wherein the targeted modification comprises a premature stop codon between a first or original starting ATG codon of the open reading frame of the endogenous gene and a subsequent downstream ATG codon, and wherein a) the downstream ATG codon is located upstream of or within the 5’ region of a GRAS domain coding sequence, b) the first or original starting ATG codon is in-frame with the subsequent downstream ATG codon; and c) the premature stop codon is in-frame or out-of-frame with the downstream ATG codon. wherein the targeted modification results in the plant having reduced plant height compared to a control plant not comprising the targeted modification, and wherein the reduced height phenotype results from translation reinitiation at the downstream ATG codon.
41. The plant of claim 40, wherein the targeted modification comprises an addition, deletion or substitution of at least one nucleotide that introduces or creates the premature stop codon.
42. The plant of claims 40 and 41, wherein the targeted modification results in translation reinitiation at the downstream ATG codon and a truncated dwarf8 protein.
43. The plant of any one of claims 40-42, wherein the premature stop codon is introduced between the first methionine (Ml) of SEQ ID NO: 6 and a second methionine selected from the group consisting of methionine at position 53 (M53), methionine at position 65 (M65), methionine at position 67 (M67), methionine at position 69 (M69), methionine at position 106 (M106), methionine at position 184 (M184), methionine at position 201 (M201), methionine at position 271 (M271), and methionine at position 280 (M280) of SEQ ID NO:6, wherein the second methionine is in-frame with the first methionine and the premature stop codon is in-frame or out-of-frame with the second methionine.
44. The plant of any one of claims 40-43, wherein the targeted modification comprises an inframe stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
45. The plant of any one of claims 40-43, wherein the targeted modification comprises an out-of- frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
46. The plant of any one of claims 40-45 wherein the plant is a maize plant and exhibits a semidwarf phenotype due to translation reinitiation.
47. The plant of any one of claims 40-46 comprising SEQ ID NO:7 or SEQ ID NO: 58.
48. The plant of any one of claims 40-47 is a maize inbred.
49. A seed comprising a targeted modification at an endogenous gene that encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6, wherein the targeted modification comprises a premature stop codon between a first or original starting ATG codon of the open reading frame of the endogenous gene and a subsequent downstream ATG codon, and wherein a) the downstream ATG codon is located upstream of or within the 5’ region of a GRAS domain coding sequence, b) the first or original starting ATG codon is in-frame with the subsequent downstream ATG codon; and c) the premature stop codon is in-frame or out-of-frame with the downstream ATG codon.
50. The seed of claim 49, wherein the targeted modification comprises an addition, deletion or substitution of at least one nucleotide that introduces or creates the premature stop codon.
51. The seed of claims 49 and 50, wherein the targeted modification results in translation reinitiation at the downstream ATG codon and a truncated dwarf8 protein.
52. The seed of any one of claims 49-51, wherein the premature stop codon is introduced between the first methionine (Ml) and a second methionine selected from the group consisting of methionine at position 53 (M53), methionine at position 65 (M65), methionine at position 67 (M67), methionine at position 69 (M69), methionine at position 106 (Ml 06), methionine at position 184 (M184), methionine at position 201 (M201), methionine at position 271 (M271), and methionine at position 280 (M280) of SEQ ID NO: 6, wherein thesecond methionine is in-frame with the first methionine and the premature stop codon is inframe or out-of-frame with the second methionine.
53. The seed of any one of claims 49-52, wherein the targeted modification comprises an inframe stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
54. The seed of any one of claims 49-52, wherein the targeted modification comprises an out-of- frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
55. The seed of any one of claims 49-54 comprising SEQ ID NO:7 or SEQ ID NO: 58.
56. The seed of any one of claims 49-55, wherein the seed is from a modified maize plant that exhibits a semi-dwarf phenotype due to translation reinitiation.
57. The seed of any one of claims 49-55 is a maize inbred or a hybrid seed.
58. The seed of any one of claims claim 49-57 produced by a method comprising crossing, selfing, double haploid, and a combination of the foregoing.
59. The seed of claims 49-58 is a maize hybrid seed comprising the targeted modification in a heterozygous state.
60. A method of identifying a genomic variation in a genomic region of a maize plant, the method comprising genotyping of one or more isolated polynucleotide samples of one or more maize plants, wherein the polynucleotide samples comprise a portion of the polynucleotide of the genomic region that encodes a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6 and identifying the genomic variation based on the genotyping.
61. The method of claim 60, wherein the genomic variation comprises an addition, deletion, or substitution of at least one nucleotide that introduces or creates a premature stop codon in the genomic region that encodes the polypeptide.
62. The method of claims 60 and 61, wherein the genomic variation comprising the premature stop codon is introduced between a first methionine, Ml, of SEQ ID NO: 6 and a second methionine corresponding to methionine at position 53 (M53), methionine at position 65 (M65), methionine at position 67 (M67), methionine at position 69 (M69), methionine at position 106 (M106), methionine at position 184 (M184), methionine at position 201 (M201), methionine at position 271 (M271), and methionine at position 280 (M280) of SEQ ID NO: 6, wherein the second methionine is in-frame with the first methionine and the premature stop codon is in-frame or out-of-frame with the second methionine.
63. The method of any one of claims 60-62, wherein the genomic variation comprises an inframe stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
64. The method of any one of claims 60-62, wherein the genomic variation comprises an out-of- frame stop codon at nucleotide positions 114 to 116 of SEQ ID NO:48.
65. The method of any one of claims 60-64, wherein the genomic variation results in a semidwarf phenotype in a maize plant due to translation reinitiation at the downstream ATG codon.
66. A recombinant DNA construct comprising a selectable marker for herbicide tolerance, a CRISPR-Cas genome editing system that targets a genomic polynucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6 to introduce a targeted mutation, wherein the targeted mutation is introduced between a first or original starting ATG codon of the open reading frame and a subsequent downstream ATG codon and wherein the downstream ATG codon is located upstream of or within the 5’ region of the GRAS domain coding sequence of an endogenous dwarf8 gene.
67. A method of targeted creation of an in-frame or out-of-frame translation reinitiation site at a plant genomic locus encoding a gene of interest, the method comprising: a) deleting, inserting, or modifying one or more nucleotides at a genomic sequence comprising an open reading frame such that a premature stop codon after the codon encoding the first or original starting Met of the open reading frame; and b) creating a translation reinitiation start site that is downstream of the premature stop codon such that the ribosomal translational machinery is capable of initiating translation at the reinitiation start site, wherein the reinitiation start site is positioned at a sufficient distance from the first or the original Met to provide the translational machinery to be associated with and recruit the necessary components to reinitiate translation.
68. The method of claim 67, wherein the gene of interest confers a dominant phenotype.
69. The method of claim 67, wherein the translation reinitiation site results in the production of a transcript and / or polypeptide that confers an agronomic trait of interest.
70. The method of claim 67, wherein the gene of interest is a maize gene involved in plant architecture modulation.71 . The method of claim 67, wherein the plant is maize.
72. The method of claim 69, wherein the agronomic trait of interest is selected from the group consisting of plant height, yield, maturity, moisture content, disease resistance, drought tolerance, and nutrient use efficiency.
73. The method of claim 67, wherein the targeted modification is performed by a site-directed DNA modifying agent.
Citation Information
Patent Citations
Genetic control of plant growth and development
US20050060773A1
Compositions and methods for stature modification in plants
US20200199609A1
Wheat With New Alleles of RHT-B1
US20210040497A1