Methods and means for influencing the expression of heteroalleles or alleles in plants by modifying untranslated regions.

JP2026529691APending Publication Date: 2026-09-01MONSANTO TECHNOLOGY LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026511588
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-21
Filing Date
2024-08-19
Publication Date
2026-09-01

Smart Images

  • Figure 2026529691000013
    Figure 2026529691000013
  • Figure 2026529691000014
    Figure 2026529691000014
  • Figure 2026529691000015
    Figure 2026529691000015
Patent Text Reader

Abstract

A composition and method are provided for editing the untranslated region of an endogenous gene using genome editing technology, thereby affecting the expression of an endogenous gene or a modifying allele only in plants such as heterozygous hybrid crops, without affecting the expression of the endogenous gene or the modifying allele in a homozygous state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to compositions and methods for affecting expression of an endogenous gene or allele in a plant by editing the non-coding region of the endogenous gene via genome editing technology only in the heterozygous allelic state without affecting expression of the endogenous gene or allele in the plant in the homozygous state, such as in hybrid crops. The present compositions and methods can also be used to generate moderate, beneficial phenotypes based on gene alleles in plants having strong, potentially deleterious, or non-viable phenotypes, particularly when present in the homozygous state.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS The present application claims priority to U.S. Provisional Applications No. 63 / 520,915 and 63 / 520,898, filed on August 21, 2023, which are incorporated herein by reference in their entireties.

[0003] INCORPORATION OF SEQUENCE LISTING The sequence listing contained in the file named "BCS236339.XML", created on August 17, 2024, with a size of 493 kilobytes (measured in MS-Windows®) and containing 235 sequences, is electronically filed herewith and incorporated herein by reference in its entirety.

Background Art

[0004] Variant alleles of endogenous plant genes have been described that produce potentially commercially interesting phenotypes, such as reduced plant height or reduced seed shattering, but can be harmful or non-viable, or produce undesirable phenotypes during plant development or reproduction, particularly when present in the homozygous state.

[0005] Such variant alleles would hinder the commercial production of hybrid crops containing them in large-scale seed production, because the parent plants used for hybrid seed production would essentially contain these variant alleles in a homozygous state, thereby reducing their ability to produce high-quality hybrid seeds and / or sufficient quantities of hybrid seeds. For example, known dominant variant alleles related to height reduction, such as the reduction of plant height in maize, result in an excessively short phenotype and exhibit reproductive off-type when present in a homozygous state in the plant. As another example, known variant alleles affecting the development of the dehiscence zone in brassica plants, when present in a heterozygous state, can reduce seed or pod crushing, while still allowing pod opening and seed harvesting using conventional harvesting equipment. However, when present in a homozygous state in the parent plants used for hybrid seed production, it results in pods that can no longer be opened using conventional harvesting equipment.

[0006] Therefore, when such alleles are present in a homozygous state, there is still a need for methods and compositions to generate variant alleles with moderate phenotypes in hybrid plants based on variant alleles in plants that have potent and potentially harmful phenotypes, without causing mutant phenotypes that have harmful effects on seed or plant production in plants on a large scale. This problem is solved as described below, including different embodiments, claims and examples. [Overview of the project]

[0007] In summary, various aspects of the present invention are described in the following numbered embodiments.

[0008] Embodiment 1. A method for editing the genome of a plant cell to modify endogenous genes, a) A step of generating a first double-strand break and a second double-strand break in the plant cell without perturbing the coding region, using a targeted editing technique that targets at least one untranslated region of the endogenous gene, b) The method comprising the step of isolating a modified plant cell containing a modifying allele of the endogenous gene, wherein the modifying allele contains an inverted DNA sequence of at least a portion of the at least one untranslated region of the endogenous gene, the modifying allele produces an RNA transcript containing an antisense sequence of the portion of the at least one untranslated region, and the modifying allele does not contain a sense sequence complementary to the antisense sequence of the portion of the at least one untranslated region.

[0009] Embodiment 2) The method according to Embodiment 1, wherein the transcription of the modified allele does not produce an RNA molecule containing a stem-loop structure.

[0010] Embodiment 3) The method according to either Embodiment 1 or 2, wherein the modifying allele of the endogenous gene, when present in the cell in a homozygous state, does not result in reduced expression of the modifying allele of the endogenous gene.

[0011] Embodiment 4) The method according to any one of Embodiments 1 to 3, wherein the inverted DNA sequence of at least a portion of the untranslated region of the endogenous gene encodes an antisense RNA sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleothiodies of the untranslated region of the gene.

[0012] Embodiment 5) The method according to any one of Embodiments 1 to 4, wherein the reduced expression of the gene results in a desired phenotype in a plant containing or essentially consisting of the modified plant cells.

[0013] Embodiment 6) The method according to any one of Embodiments 1 to 5, wherein the untranslated region of the endogenous gene is a 5' untranslated region, a 3' untranslated region, or both.

[0014] Embodiment 7) The method according to any one of Embodiments 1 to 5, wherein the untranslated region is an intron sequence of the endogenous gene.

[0015] Embodiment 8) The method according to any one of Embodiments 1 to 7, wherein the modifying allele of the endogenous gene is present homozygously in the plant cell.

[0016] Embodiment 9) The method according to any one of Embodiments 1 to 7, wherein the modifying allele of the endogenous gene is present heterozygously in the plant cell.

[0017] Embodiment 10) The method according to any one of Embodiments 1 to 7, wherein the plant cell includes the modified allele and the unmodified allele of the endogenous gene.

[0018] Embodiment 11) The method according to any one of Embodiments 1 to 7, wherein the modifying allele of the endogenous gene exists in a heteroallele state within the plant cell, the cell further contains the unmodified allele of the gene, and the expression of the modified and unmodified alleles of the gene is reduced.

[0019] Embodiment 12) The method according to Embodiment 11, wherein the modified allele and the unmodified allele are transcribed into mRNA, and the RNA derived from the modified allele and the RNA derived from the unmodified allele can generate a double-stranded RNA region of at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides.

[0020] Embodiment 13) The method according to any one of Embodiments 9 to 12, wherein the expression of the modified allele and the unmodified allele is reduced.

[0021] Embodiment 14) The method according to Embodiment 13, wherein the reduced expression results in the desired phenotype.

[0022] Embodiment 15) The method according to any one of Embodiments 1 to 14, further comprising the step of regenerating a plant from the modified plant cells.

[0023] Embodiment 16) The method according to Embodiment 14, further comprising the step of crossing a plant containing the modifying allele of the endogenous gene in a homozygous state with another plant containing the unmodifying allele of the endogenous gene in a homozygous state, and harvesting hybrid seeds. The hybrid seeds contain the modifying allele of the endogenous gene and the unmodifying allele of the endogenous gene.

[0024] Embodiment 17) The method according to any one of Embodiments 1 to 16, wherein the plant is rapeseed and the endogenous gene is a non-dehiscing gene derived from rapeseed.

[0025] Embodiment 18) The method according to any one of Embodiments 1 to 16, wherein the plant is maize and the endogenous gene is selected from GA20 oxidase or GA3 oxidase.

[0026] Embodiment 19) The method according to Embodiment 18, wherein the GA20 oxidase is selected from GA20 oxidase_5 or GA20 oxidase_3.

[0027] Embodiment 20) The method according to Embodiment 18, wherein the GA3 oxidase is selected from GA3 oxidase_1, GA3 oxidase_2 or GA3 oxidase_3.

[0028] Embodiment 21) The method according to any one of Embodiments 1 to 16, wherein the plant is maize, and the endogenous gene is selected from Anther Ear1 (GRMZM2G081554), dwarf4 (GRMZM2G065635), brs1-brassinosteroid synthesis 1, nana plant 1 (GRMZM2G057000), brassinosteroid receptor ZmBRI1a / ZmBRI1b (GRMZM2G048294 / GRMZM2G449830), meristem development gene compact plant 2 (GRMZM2G064732) and ZMWRKY60.

[0029] Embodiment 22) The method according to any one of Embodiments 1 to 16, wherein the endogenous gene is Agamous, Bri1, Dwarf1, Pin1 derived from Arabidopsis, or an orthologous gene derived from another plant.

[0030] Embodiment 23) The method according to any one of Embodiments 1 to 16, wherein the endogenous gene encodes a protein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 9, 15, 30, 33, 173, and 214 to 217.

[0031] Embodiment 24) The inverted DNA sequence is: nucleotides 1-29 of SEQ ID NO: 36, nucleotides 1664-1788 of SEQ ID NO: 36, nucleotides 1-38 of SEQ ID NO: 37, nucleotides 1446-1698 of SEQ ID NO: 37, nucleotides 3001-3161 of SEQ ID NO: 168, nucleotides 4796-5406 of SEQ ID NO: 168, nucleotides 3001-3056 of SEQ ID NO: 169, nucleotides 4464-4581 of SEQ ID NO: 169, and nucleotides 3001-313 of SEQ ID NO: 170 0, Nucleotides 4275-4332 of SEQ ID NO: 170, Nucleotides 7621-8029 of SEQ ID NO: 174, Nucleotides 9672-10276 of SEQ ID NO: 174, Nucleotides 7386-7831 of SEQ ID NO: 175, Nucleotides 8862-8967 of SEQ ID NO: 175, Nucleotides 7547-7751 of SEQ ID NO: 176, Nucleotides 8904-9178 of SEQ ID NO: 204, Nucleotides 1-1060 of SEQ ID NO: 204, Nucleotides 5418-5648 of SEQ ID NO: 204, Sequence The method according to any one of Embodiments 1 to 16, comprising a nucleotide sequence having at least 90% sequence identity or complementarity to at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, or at least 1000 consecutive nucleotides from nucleotides 1 to 165 of number 211, nucleotides 3757 to 4167 of SEQ ID NO: 211, nucleotides 664 to 699 of SEQ ID NO: 212, nucleotides 2482 to 2700 of SEQ ID NO: 212, nucleotides 1 to 99 of SEQ ID NO: 213, or nucleotides 3205 to 3506 of SEQ ID NO: 213.

[0032] Embodiment 25) The method according to any one of Embodiments 1 to 24, wherein the targeted editing technique comprises the use of an RNA guide effector protein, or a TALE protein, or a custom meganuclease.

[0033] Embodiment 26) The method according to Embodiment 25, wherein the RNA guide effector protein is a CRISPR-Cas effector protein selected from a CRISPR-Cas system type I, a CRISPR-Cas system type II, a CRISPR-Cas system type III, a CRISPR-Cas system type IV, a CRISPR-Cas system type V, or a CRISPR-Cas system type VI, or a CRISPR-Cas effector protein derived therefrom, or a CRISPR-Cas effector protein containing one or more nuclear localization signals.

[0034] Embodiment 27) The RNA guide endonuclease is Cas9, C2c1, C2c3, Cas12a (also called Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Csnl, Csx12, Cas10, Csyl, Csy2, Csy3, Csel, Cse2, 30Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, C The method according to either embodiment 25 or 26, wherein the CRISPR-Cas effector protein is selected from mr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), Csf5 nuclease, Cas12c(C2c3), Cas12d(CasY), Cas12e(CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and Cas14c effector proteins.

[0035] Embodiment 28) The method according to any one of Embodiments 25 to 28, wherein the RNA guide effector protein is a Cas12a effector protein or an effector protein derived from Cas12a.

[0036] Embodiment 29) The method according to Embodiment 28, wherein the Cas12a effector protein is selected from FnCas12a, LbCas12a, ErCas12a, or AsCas12a or variants thereof.

[0037] Embodiment 30) The method according to Embodiment 29, wherein the Cas12a effector protein comprises an amino acid sequence having at least 95% sequence identity with an amino acid sequence selected from the group consisting of SEQ ID NOs: 194 and 199.

[0038] Embodiment 30) The method according to any one of Embodiments 25 to 29, wherein the target gene editing technique includes the use of one or more guide RNAs containing nucleotide sequences selected from the group of SEQ ID NOs: 177, 178, 179, 180, 205, 206, 207, 208, and 209.

[0039] Embodiment 31) A method for modifying the expression of endogenous genes in hybrid plants without affecting the expression of the genes in the parent plants (multiple parent plants are possible), a) A step of identifying an endogenous gene in a plant, wherein the expression of a variant allele of the gene, when present in a homozygous state, results in an undesirable phenotype; b) A step of providing a first plant comprising a modifying allele of the gene, the inversion of which is a nucleic acid region of a part of the gene, wherein the inversion does not affect the translation of the modifying allele, and the first plant comprises the modifying allele of the gene in a homozygous manner, c) A step of crossing the first plant with a second plant containing an unmodified allele of the gene that does not have the inversion that does not affect the translation of the gene, wherein the unmodified gene is in a homozygous state, d) The method comprising the step of obtaining a hybrid seed containing the modifying allele and the unmodified allele of the gene in a heterozygous or heteroallelic form.

[0040] Embodiment 32) The method according to Embodiment 31, wherein when the modified allele and the unmodified allele are transcribed into an RNA molecule, a double-stranded RNA region may be formed by base pairing between the nucleic acid region, which is an inversion of a portion of the untranslated region of the gene in the RNA transcript of the modified allele, and the nucleic acid region in the RNA transcript of the unmodified allele, and the double-stranded RNA region can inhibit the expression of the modified allele and the unmodified allele by RNA silencing mechanisms such as RNA translation stalling, RNA transcription stalling, resulting destabilization of the RNA molecule, or post-transcriptional degradation of the transcribed RNA molecule.

[0041] Embodiment 33) The method according to either Embodiment 31 or 32, wherein the transcription of the modified allele generates an RNA molecule that does not contain a stem-loop structure.

[0042] Embodiment 34) The method according to any one of Embodiments 31 to 33, wherein the modifying allele of the endogenous gene, when present in a homozygous state in the cells of the first plant, does not result in reduced expression of the modifying allele of the endogenous gene.

[0043] Embodiment 35) The method according to any one of Embodiments 31 to 34, wherein the nucleic acid region which is an inversion of a portion of the gene occurs during transcription in an antisense RNA sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleothiodies of the untranslated region of the gene.

[0044] Embodiment 36) The method according to any one of Embodiments 31 to 35, wherein the expression of the modified allele and the unmodified allele of the endogenous gene is reduced.

[0045] Embodiment 37) The method according to any one of Embodiments 31 to 36, wherein the untranslated region of the endogenous gene is a 5' untranslated region or a 3' untranslated region.

[0046] Embodiment 37) The method according to any one of Embodiments 31 to 36, wherein the untranslated region is an intron sequence of the endogenous gene.

[0047] Embodiment 38) The method according to any one of Embodiments 31 to 37, wherein the plant is rapeseed and the endogenous gene is a non-dehiscing gene derived from rapeseed.

[0048] Embodiment 39) The method according to any one of Embodiments 31 to 37, wherein the plant is maize and the endogenous gene is selected from GA20 oxidase or GA3 oxidase.

[0049] Embodiment 40) The method according to Embodiment 39, wherein the GA20 oxidase is selected from GA20 oxidase_5 or GA20 oxidase_3.

[0050] Embodiment 41) The method according to Embodiment 39, wherein the GA3 oxidase is selected from GA3 oxidase_1, GA3 oxidase_2, or GA3 oxidase_3.

[0051] Embodiment 42) The method according to any one of Embodiments 31 to 37, wherein the plant is maize, and the endogenous gene is selected from Anther Ear1 (GRMZM2G081554)dwarf4 (GRMZM2G065635)brs1 brassinosteroid synthesis 1, nana plant 1 (GRMZM2G057000), brassinosteroid receptor ZmBRI1a / ZmBRI1b (GRMZM2G048294 / GRMZM2G449830), meristematic development gene compact plant 2 (GRMZM2G064732), and ZMWRKY60.

[0052] Embodiment 43) The method according to any one of Embodiments 31 to 37, wherein the endogenous gene is Agamous, Bri1, Dwarf1, Pin1 from Arabidopsis, or an ortholog gene from another plant.

[0053] Embodiment 44) The method according to any one of Embodiments 31 to 37, wherein the endogenous gene encodes a protein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with an amino acid sequence selected from the group consisting of SEQ ID NOs. 9, 15, 30, 33, 173, and 214-217.

[0054] Embodiment 45) The inverted DNA sequence is nucleotides 1-29 of SEQ ID NO: 36, nucleotides 1664-1788 of SEQ ID NO: 36, nucleotides 1-38 of SEQ ID NO: 37, nucleotides 1446-1698 of SEQ ID NO: 37, nucleotides 3001-3161 of SEQ ID NO: 168, nucleotides 4796-5406 of SEQ ID NO: 168, nucleotides 3001-3056 of SEQ ID NO: 169, nucleotides 4464-4581 of SEQ ID NO: 169, Nucleotides 3001-3130 of SEQ ID NO: 170, Nucleotides 4275-4332 of SEQ ID NO: 174, Nucleotides 7621-8029 of SEQ ID NO: 174, Nucleotides 9672-10276 of SEQ ID NO: 175, Nucleotides 7386-7831 of SEQ ID NO: 175, Nucleotides 8862-8967 of SEQ ID NO: 176, Nucleotides 7547-7751 of SEQ ID NO: 176, Nucleotides 8904-9178 of SEQ ID NO: 176, Sequence number The method according to any one of Embodiments 31 to 37, comprising a nucleotide sequence having at least 90% sequence identity or complementarity with at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, or at least 1000 consecutive nucleotides of nucleotides 1 to 1060 of nucleotide 204, nucleotides 5418 to 5648 of SEQ ID NO: 204, nucleotides 1 to 165 of SEQ ID NO: 211, nucleotides 3757 to 4167 of SEQ ID NO: 211, nucleotides 664 to 699 of SEQ ID NO: 212, nucleotides 2482 to 2700 of SEQ ID NO: 212, nucleotides 1 to 99 of SEQ ID NO: 213, or nucleotides 3205 to 3506 of SEQ ID NO: 213.

[0055] Embodiment 46) The method according to any one of Embodiments 31 to 45, wherein the modified allele in the first plant is obtained by a targeted editing technique, such as a targeted editing technique including the use of an RNA guide effector protein or a TALE protein or a custom meganuclease.

[0056] Embodiment 47) The method according to Embodiment 46, wherein the RNA guide effector protein is a CRISPR-Cas effector protein selected from a CRISPR-Cas system type I, a CRISPR-Cas system type II, a CRISPR-Cas system type III, a CRISPR-Cas system type IV, a CRISPR-Cas system type V, or a CRISPR-Cas system type VI, or a CRISPR-Cas effector protein derived therefrom, or optionally a CRISPR-Cas effector protein containing one or more nuclear localization signals.

[0057] Embodiment 48) The RNA guide endonuclease is Cas9, C2c1, C2c3, Cas12a (also called Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Csnl, Csx12, Cas10, Csyl, Csy2, Csy3, Csel, Cse2, 30Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr The method according to Embodiment 46 or 47, wherein the CRISPR-Cas effector protein is selected from 4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), Csf5 nuclease, Cas12c(C2c3), Cas12d(CasY), Cas12e(CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and Cas14c effector proteins.

[0058] Embodiment 49) The method according to any one of Embodiments 46 to 48, wherein the RNA guide effector protein is a Cas12a effector protein or an effector protein derived from Cas12a.

[0059] Embodiment 50) The method according to Embodiment 49, wherein the Cas12a effector protein is selected from FnCas12a, LbCas12a, ErCas12a, or AsCas12a or variants thereof.

[0060] Embodiment 51) The method according to Embodiment 50, wherein the Cas12a effector protein comprises an amino acid sequence having at least 95% sequence identity with an amino acid sequence selected from the group consisting of SEQ ID NOs: 194 and 199.

[0061] Embodiment 52) ​​The method according to any one of Embodiments 46 to 51, wherein the target gene editing technique includes the use of one or more guide RNAs containing nucleotide sequences selected from the group of SEQ ID NOs: 177, 178, 179, 180, 205, 206, 207, 208, and 209.

[0062] Embodiment 53) A plant cell, plant or part thereof or seed comprising a modifying allele of an endogenous gene, wherein the modifying allele comprises an inverted DNA sequence of at least a portion of the untranslated region of the endogenous gene, the modifying allele produces an RNA transcript comprising an antisense sequence of at least a portion of the untranslated region, and the modifying allele does not contain a sense nucleotide sequence of more than 17 nucleotides that is complementary to the antisense sequence of the portion of the untranslated region.

[0063] Embodiment 54) The plant cell, plant or part thereof according to Embodiment 53, wherein the plant cell, plant or part thereof is non-transgenic.

[0064] Embodiment 55) The plant cell, plant or a part thereof or its seed according to either Embodiment 53 or 54, wherein the modification allele of the endogenous gene is obtained by targeted editing technology.

[0065] Embodiment 56) A plant cell, a plant or a part thereof, or a seed according to any one of Embodiments 53 to 55, wherein the transcription of the modified allele produces an RNA molecule that does not contain a stem-loop structure.

[0066] Embodiment 57) A plant cell, plant or part thereof or seed according to any one of Embodiments 53 to 56, wherein the inverted DNA sequence of at least a portion of the untranslated region of the endogenous gene occurs during transcription in an antisense RNA sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleothiodies of the untranslated region of the gene.

[0067] Embodiment 58) The plant cell, plant or part thereof or seed according to any one of Embodiments 53 to 57, wherein the untranslated region of the endogenous gene is a 5' untranslated region or a 3' untranslated region.

[0068] Embodiment 59) The plant cell, plant or part thereof or seed according to any one of Embodiments 53 to 57, wherein the untranslated region is an intron sequence of the endogenous gene.

[0069] Embodiment 60) The plant is rapeseed, and the endogenous gene is a non-dehiscing gene derived from rapeseed, the plant cell, plant or part thereof or seed according to any one of Embodiments 53 to 59.

[0070] Embodiment 61) The plant cell, plant or part thereof or seed according to any one of Embodiments 53 to 59, wherein the plant is maize and the endogenous gene is selected from GA20 oxidase or GA3 oxidase.

[0071] Embodiment 62) The GA20 oxidase is selected from GA20 oxidase_5 or GA20 oxidase_3, as described in Embodiment 61, for plant cells, plants or parts thereof or seeds.

[0072] Embodiment 63) The GA3 oxidase is selected from GA3 oxidase_1, GA3 oxidase_2, or GA3 oxidase_3, as described in Embodiment 61, a plant cell, a plant or part thereof, or a seed.

[0073] Embodiment 64) The plant is maize, and the endogenous gene is selected from Anther Ear1 (GRMZM2G081554)dwarf4 (GRMZM2G065635)brs1-brassinosteroid synthesis 1, Anubias plant 1 (GRMZM2G057000), brassinosteroid receptor ZmBRI1a / ZmBRI1b (GRMZM2G048294 / GRMZM2G449830), meristematic development gene compact plant 2 (GRMZM2G064732), and ZMWRKY60, wherein the plant is a maize, and the endogenous gene is selected from Anther Ear1 (GRMZM2G065635)brs1-brassinosteroid synthesis 1, Anubias plant 1 (GRMZM2G057000), brassinosteroid receptor ZmBRI1a / ZmBRI1b (GRMZM2G048294 / GRMZM2G449830), meristematic development gene compact plant 2 (GRMZM2G064732), and ZMWRKY60, the plant cell, plant or part thereof or seed according to any one of Embodiments 53 to 59.

[0074] Embodiment 65) The plant cell, plant or part thereof or seed according to any one of Embodiments 53 to 59, wherein the endogenous gene is Agamous, Bri1, Dwarf1, Pin1 from Arabidopsis, or an ortholog gene from another plant.

[0075] Embodiment 66) A plant cell, plant or part thereof or seed according to any one of Embodiments 53 to 59, wherein the endogenous gene encodes a protein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with an amino acid sequence selected from the group consisting of SEQ ID NOs. 9, 15, 30, 33, 173, and 214-217.

[0076] Embodiment 67) The inverted DNA sequence is: nucleotides 1-29 of SEQ ID NO: 36, nucleotides 1664-1788 of SEQ ID NO: 36, nucleotides 1-38 of SEQ ID NO: 37, nucleotides 1446-1698 of SEQ ID NO: 37, nucleotides 3001-3161 of SEQ ID NO: 168, nucleotides 4796-5406 of SEQ ID NO: 168, nucleotides 3001-3056 of SEQ ID NO: 169, nucleotides 4464-4581 of SEQ ID NO: Nucleotides 3001-3130 of 170, nucleotides 4275-4332 of SEQ ID NO: 170, nucleotides 7621-8029 of SEQ ID NO: 174, nucleotides 9672-10276 of SEQ ID NO: 174, nucleotides 7386-7831 of SEQ ID NO: 175, nucleotides 8862-8967 of SEQ ID NO: 175, nucleotides 7547-7751 of SEQ ID NO: 176, nucleotides 8904-9178 of SEQ ID NO: 204 A plant cell, plant or part thereof or seed according to any one of embodiments 53 to 59, comprising a nucleotide sequence having at least 90% sequence identity or complementarity with at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides of otinodes 1 to 1060, nucleotides 5418 to 5648 of SEQ ID NO: 204, nucleotides 1 to 165 of SEQ ID NO: 211, nucleotides 3757 to 4167 of SEQ ID NO: 211, nucleotides 664 to 699 of SEQ ID NO: 212, nucleotides 2482 to 2700 of SEQ ID NO: 212, nucleotides 1 to 99 of SEQ ID NO: 213, or nucleotides 3205 to 3506 of SEQ ID NO: 213.

[0077] Embodiment 68) The modified allele in the first plant is obtained by a targeted editing technique, such as a targeted editing technique involving the use of an RNA guide effector protein or a TALE protein or a custom meganuclease, as described in any one of Embodiments 53 to 67, wherein the plant cell, plant or part thereof or seed.

[0078] Embodiment 69) The plant cell, plant or part thereof or seed according to Embodiment 68, wherein the RNA guide effector protein is a CRISPR-Cas effector protein selected from a CRISPR-Cas system type I, a CRISPR-Cas system type II, a CRISPR-Cas system type III, a CRISPR-Cas system type IV, a CRISPR-Cas system type V, or a CRISPR-Cas system type VI, or a CRISPR-Cas effector protein derived therefrom, or optionally a CRISPR-Cas effector protein containing one or more nuclear localization signals.

[0079] Embodiment 70) The RNA guide endonuclease is Cas9, C2c1, C2c3, Cas12a (also called Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Csnl, Csx12, Cas10, Csyl, Csy2, Csy3, Csel, Cse2, 30Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, A CRISPR-Cas effector protein selected from Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), Csf5 nuclease, Cas12c(C2c3), Cas12d(CasY), Cas12e(CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, Cas14c effector protein, plant cell, plant or part thereof or seed, as described in any one of Embodiments 68 or 69.

[0080] Embodiment 71) The plant cell, plant or part thereof or seed according to any one of Embodiments 68 to 70, wherein the RNA guide effector protein is a Cas12a effector protein or an effector protein derived from Cas12a.

[0081] Embodiment 72) The Cas12a effector protein is selected from FnCas12a, LbCas12a, ErCas12a, or AsCas12a or variants thereof, as described in Embodiment 71, a plant cell, a plant or part thereof, or a seed.

[0082] Embodiment 73) The plant cell, plant or part thereof or seed according to Embodiment 72, wherein the Cas12a effector protein comprises an amino acid sequence having at least 95% sequence identity with an amino acid sequence selected from the group consisting of SEQ ID NOs: 194 and 199.

[0083] Embodiment 74) The target gene editing technique comprises the use of one or more guide RNAs containing nucleotide sequences selected from the group of SEQ ID NOs: 177, 178, 179, 180, 205, 206, 207, 208, and 209, as described in any one of Embodiments 68 to 73, wherein the target gene editing technique comprises the use of one or more guide RNAs containing nucleotide sequences selected from the group SEQ ID NOs: 177, 178, 179, 180, 205, 206, 207, 208, and 209, a plant cell, a plant or part thereof, or a seed.

[0084] Embodiment 75) A plant cell, plant or part thereof, or seed according to any one of Embodiments 53 to 74, wherein the modifying allele is present in a homozygous manner.

[0085] Embodiment 76) The plant cell, plant or part thereof, or seed according to any one of Embodiments 53 to 74, wherein the modifying allele is present in a heterozygous or heteroallelic form.

[0086] Embodiment 77) The plant cell, plant or part thereof or seed according to Embodiment 78, wherein the plant cell, plant or part thereof or seed further comprises an unmodified allele of the endogenous gene.

[0087] Embodiment 78) The plant cell, plant or part thereof or seed according to Embodiment 77, wherein the modified allele and the unmodified allele are transcribed into RNA, and the RNA derived from the modified allele and the RNA derived from the unmodified allele can generate a double-stranded RNA region of at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides.

[0088] Embodiment 79) The plant cell, plant or part thereof or seed according to Embodiment 78, wherein the double-stranded RNA region can inhibit the expression of the modified allele and the unmodified allele by an RNA silencing mechanism such as RNA translation stalling, RNA transcription stalling, resulting destabilization of the RNA molecule, or post-transcriptional degradation of the transcribed RNA molecule.

[0089] Embodiment 80) The plant cell, plant or part thereof or seed according to Embodiment 75, wherein the expression of the modified allele is not reduced.

[0090] Embodiment 81) The plant cell, plant or part thereof or seed according to any one of Embodiments 76 to 79, wherein the plant cell, plant or part thereof or seed exhibits the desired phenotype.

[0091] Embodiment 82) The plant according to Embodiment 81, wherein the desired phenotype is characterized by a shorter height compared to a plant that does not contain the modified and unmodified alleles of the endogenous gene.

[0092] Embodiment 83) The plant cell, plant or part thereof or seed according to Embodiment 53, wherein the modifying allele of the endogenous gene containing the inverted DNA sequence is operably linked to its natural or homologous promoter.

[0093] Embodiment 84) A plant regenerated from plant cells according to any one of Embodiments 53 to 83.

[0094] Embodiment 85) A plant comprising or essentially comprising the plant cells described in any one of Embodiments 53 to 83.

[0095] Embodiment 86) A plant or seed obtained by the method described in Embodiment 15 or 16.

[0096] Embodiment 87) The plant according to any one of Embodiments 53 to 86, wherein the plant is selected from monocotyledonous plant species, dicotyledonous plant species, angiosperm species, or gymnosperm species.

[0097] Embodiment 88) The plant according to Embodiment 87, wherein the plant is selected from corn plants, rice plants, sorghum plants, wheat plants, alfalfa plants, barley plants, millet plants, rye plants, sugarcane plants, cotton plants, soybean plants, canola plants, tomato plants, onion plants, cucumber plants, Arabidopsis plants, or potato plants.

[0098] Embodiment 89) The plant according to Embodiment 88, wherein the plant is a low-growing maize plant.

[0099] Embodiment 90) The method according to any one of Embodiments 1 to 30, wherein the modifying allele of the endogenous gene includes an inverted DNA sequence of at least a portion of two untranslated regions.

[0100] Embodiment 91) A method for editing the genome of a plant cell in order to modify endogenous genes, a. A step of generating a double-strand DNA break or a single-strand DNA break (nick) in the plant cell using a targeted editing technique that targets at least one untranslated region of the endogenous gene without perturbing the coding region, b. A step of supplying at least one template nucleic acid to the plant cells, wherein the template nucleic acid includes a portion of the at least one untranslated region of the endogenous gene in reverse orientation, c. A modified plant cell comprising the step of isolating a modified allele of the endogenous gene, wherein the modified allele comprises an inverted DNA sequence of at least a portion of the at least one untranslated region of the endogenous gene, the modified allele produces an RNA transcript comprising an antisense sequence of a portion of the untranslated region, and the modified allele does not contain a sense sequence complementary to the antisense sequence of a portion of the untranslated region.

[0101] Embodiment 92) The method according to Embodiment 91, wherein the template nucleic acid, which includes the portion of at least one untranslated region of the endogenous gene in the reverse direction, is inserted into the target untranslated region of the endogenous gene by non-homologous end joining.

[0102] Embodiment 93) The method according to Embodiment 91, wherein the template nucleic acid includes at least one or two homology arms having homology to the nucleic acid sequence adjacent to the double-strand break, and optionally the homology arms are adjacent in reverse orientation to the portion of the at least one untranslated region.

[0103] Embodiment 94) The method according to Embodiment 93, wherein a portion of the untranslated region of the endogenous gene is introduced in the reverse direction by homology-dependent repair.

[0104] Embodiment 95) A method for editing the genome of a plant cell in order to modify an endogenous gene, a. A step of generating a double-stranded DNA break or a single-stranded DNA break using a CRISPR / CAS fusion protein fused to the reverse transcriptase functional region without perturbation in the coding region of the plant cell and a guide RNA targeting at least one untranslated region of the endogenous gene, wherein the guide RNA further comprises a nucleotide sequence acid containing a portion of the at least one untranslated region of the endogenous gene in an inverted orientation, b. A method comprising the step of isolating a modified plant cell containing a modifying allele of the endogenous gene, wherein the modifying allele contains an inverted DNA sequence of at least a portion of the at least one untranslated region of the endogenous gene, the modifying allele produces an RNA transcript containing an antisense sequence of a portion of the untranslated region, and the modifying allele does not contain a sense sequence complementary to the antisense sequence of a portion of the untranslated region. [Brief explanation of the drawing]

[0105] [Figure 1] Schematic diagram of the method and composition according to the present invention. Panel A of the scheme shows the unedited structure of the gene, where the 3'UTR is oriented in the normal direction. To reverse the orientation of the 3'UTR of a candidate gene, two unique gRNAs are used to excise sections of the 3'UTR in both homologous alleles of the candidate gene and insert them in the reverse direction (Panel B). Such a rearrangement of the 3'UTR should not affect the function of the UTR if present in both alleles, and therefore no phenotype is expected at the homozygous stage. However, in a cross between such a homozygous edited plant and a wild-type plant (both alleles having the 3'UTR oriented in the normal direction), the resulting offspring inherit two alleles with different 3'UTR orientations. This can lead to the formation of a double-stranded RNA region between the complementary sequences that would be formed from both transcripts. As a result, such rearrangements may lead to the generation of siRNA molecules that can affect the stability of the gene's mRNA and the repression of the target gene's function. [Figure 2] Schematic diagram of the recombinant nucleic acid construct used in Example 3. Panels A and B: Recombinant nucleic acid constructs expressing Cas12a under the control of a plant promoter and expressing two guide RNAs that target the untranslated region of Agamous. Panel C: Positive control nucleic acid expressing a miRNA complementary to Agamous. Panel D: Control nucleic acid expressing the agamous1 allele with an inverted 5'UTR. [Figure 3]Schematic diagram of the recombinant nucleic acid constructs used in Example 4. Panel A: Recombinant nucleic acid construct for expressing the reverse 5'UTR of Agamous. Panel B: Recombinant nucleic acid construct for expressing the reverse 3'UTR of Agamous. Panel C: Recombinant nucleic acid constructs for expressing the reverse 5'UTR and 3'UTR of Agamous. [Figure 4] This diagram provides a comparison of the wild-type (WT) and edited alleles of the Zm.GA3ox_1 gene, showing that the edited alleles have deletions and inversions in the 3'UTR region.

[0106] A brief explanation of array entries in an array list. Sequence ID 1: Nucleotide sequence of GA20 oxidase_1 cDNA derived from corn (Zea mays).

[0107] Sequence ID 2: Nucleotide sequence of the GA20 oxidase_1 coding sequence derived from corn (Zea mays).

[0108] Sequence ID 3: Amino acid sequence of GA20 oxidase_1 protein derived from corn (Zea mays).

[0109] Sequence ID 4: Nucleotide sequence of GA20 oxidase_2 cDNA derived from corn (Zea mays).

[0110] Sequence ID 5: Nucleotide sequence of the GA20 oxidase_2 coding sequence derived from corn (Zea mays).

[0111] Sequence ID 6: Amino acid sequence of GA20 oxidase_2 protein derived from corn (Zea mays).

[0112] Sequence ID 7: Nucleotide sequence of GA20 oxidase_3 cDNA derived from corn (Zea mays).

[0113] Sequence ID 8: Nucleotide sequence of the GA20 oxidase_3 coding sequence derived from corn (Zea mays).

[0114] Sequence ID 9: Amino acid sequence of GA20 oxidase_3 protein derived from corn (Zea mays).

[0115] Sequence ID 10: Nucleotide sequence of GA20 oxidase_4 cDNA derived from corn (Zea mays).

[0116] Sequence ID 11: Nucleotide sequence of the GA20 oxidase_4 coding sequence derived from corn (Zea mays).

[0117] Sequence ID 12: Amino acid sequence of GA20 oxidase_4 protein derived from corn (Zea mays).

[0118] Sequence ID 13: Nucleotide sequence of GA20 oxidase_5 cDNA derived from corn (Zea mays).

[0119] Sequence ID 14: Nucleotide sequence of the GA20 oxidase_5 coding sequence derived from corn (Zea mays).

[0120] Sequence ID 15: Amino acid sequence of GA20 oxidase_5 protein derived from corn (Zea mays).

[0121] Sequence ID 16: Nucleotide sequence of GA20 oxidase_6 cDNA derived from corn (Zea mays).

[0122] Sequence ID 17: Nucleotide sequence of the GA20 oxidase_6 coding sequence derived from corn (Zea mays).

[0123] Sequence ID 18: Amino acid sequence of GA20 oxidase_6 protein derived from corn (Zea mays).

[0124] Sequence ID 19: Nucleotide sequence of GA20 oxidase_7 cDNA derived from corn (Zea mays).

[0125] Sequence ID 20: Nucleotide sequence of the GA20 oxidase_7 coding sequence derived from corn (Zea mays).

[0126] Sequence ID 21: Amino acid sequence of GA20 oxidase_7 protein derived from corn (Zea mays).

[0127] Sequence ID 22: Nucleotide sequence of GA20 oxidase_8 cDNA derived from corn (Zea mays).

[0128] Sequence ID 23: Nucleotide sequence of the GA20 oxidase_8 coding sequence derived from corn (Zea mays).

[0129] Sequence ID 24: Amino acid sequence of GA20 oxidase_8 protein derived from corn (Zea mays).

[0130] Sequence ID 25: Nucleotide sequence of GA20 oxidase_9 cDNA derived from corn (Zea mays).

[0131] Sequence ID 26: Nucleotide sequence of the GA20 oxidase_9 coding sequence derived from corn (Zea mays).

[0132] Sequence ID 27: Amino acid sequence of GA20 oxidase_9 protein derived from corn (Zea mays).

[0133] Sequence ID 28: Nucleotide sequence of GA3 oxidase_1 cDNA derived from corn (Zea mays).

[0134] Sequence ID 29: Nucleotide sequence of the GA3 oxidase_1 coding sequence derived from corn (Zea mays).

[0135] Sequence ID 30: Amino acid sequence of GA3 oxidase_1 protein derived from corn (Zea mays).

[0136] Sequence ID 31: Nucleotide sequence of GA3 oxidase_2 cDNA derived from corn (Zea mays).

[0137] Sequence ID 32: Nucleotide sequence of the GA3 oxidase_2 coding sequence derived from corn (Zea mays).

[0138] Sequence ID 33: Amino acid sequence of GA3 oxidase_2 protein derived from corn (Zea mays).

[0139] Sequence ID 34: Nucleotide sequence of the GA20 oxidase_3 genome sequence derived from corn (Zea mays).

[0140] Sequence ID 35: Nucleotide sequence of the GA20 oxidase_5 genome sequence derived from corn (Zea mays).

[0141] Sequence ID 36: Nucleotide sequence of the GA3 oxidase_1 genome sequence derived from corn (Zea mays).

[0142] Sequence ID 37: Nucleotide sequence of the GA3 oxidase_2 genome sequence derived from corn (Zea mays).

[0143] Sequence ID 38: Nucleotide sequence of the GA20 oxidase_4 genome sequence derived from corn (Zea mays).

[0144] Sequence ID 39: GA20 oxidase_3 / 5-1 nucleotide sequence of the cDNA target sequence.

[0145] Sequence ID 40: Nucleotide sequence of the GA20 oxidase_3 / 5-1 miRNA target sequence.

[0146] Sequence ID 41: Nucleotide sequence of GA20 oxidase_3 / 5-2 cDNA target sequence.

[0147] Sequence ID 42: Nucleotide sequence of the GA20 oxidase_3 / 5-2 miRNA target sequence.

[0148] Sequence ID 43: Nucleotide sequence of GA20 oxidase_3 / 5-3 cDNA target sequence.

[0149] Sequence ID 44: Nucleotide sequence of the GA20 oxidase_3 / 5-3 miRNA target sequence.

[0150] Sequence ID 45: GA20 oxidase_3 / 5-4 nucleotide sequence of the cDNA target sequence.

[0151] Sequence ID 46: Nucleotide sequence of the GA20 oxidase_3 / 5-4 miRNA target sequence.

[0152] Sequence ID 47: Nucleotide sequence of GA20 oxidase_1 / 2 cDNA target sequence.

[0153] Sequence ID 48: Nucleotide sequence of the GA20 oxidase_1 / 2 miRNA target sequence.

[0154] Sequence ID 49: GA20 oxidase_3 / 9 nucleotide sequence of the cDNA target sequence.

[0155] Sequence ID 50: Nucleotide sequence of GA20 oxidase_3 / 9 miRNA target sequence.

[0156] Sequence ID 51: Nucleotide sequence of GA20 oxidase_7 / 8 cDNA target sequence.

[0157] Sequence ID 52: Nucleotide sequence of GA20 oxidase_7 / 8 miRNA target sequence.

[0158] Sequence ID 53: Nucleotide sequence of GA20 oxidase_3 individual cDNA target sequence.

[0159] Sequence ID 54: Nucleotide sequence of GA20 oxidase_3 individual miRNA target sequence.

[0160] Sequence ID 55: Nucleotide sequence of GA20 oxidase_5 individual cDNA target sequence.

[0161] Sequence ID 56: Nucleotide sequence of GA20 oxidase_5 individual miRNA target sequence.

[0162] Sequence ID 57: Nucleotide sequence of the GA3 oxidase_1 cDNA target sequence.

[0163] Sequence ID 58: Nucleotide sequence of the GA3 oxidase_1 miRNA target sequence.

[0164] Sequence ID 59: Nucleotide sequence of the GA3 oxidase_2 cDNA target sequence.

[0165] Sequence ID 60: Nucleotide sequence of the GA3 oxidase_2 miRNA target sequence.

[0166] Sequence ID 61: Nucleotide sequence of GA20 oxidase_4 / 6-4 cDNA target sequence.

[0167] Sequence ID 62: Nucleotide sequence of the GA20 oxidase_4 / 6-4 miRNA target sequence.

[0168] Sequence ID 63: GA20 oxidase_4 / 6-6 nucleotide sequence of the cDNA target sequence.

[0169] Sequence ID 64: Nucleotide sequence of the GA20 oxidase_4 / 6-6 miRNA target sequence.

[0170] Sequence ID 65: Nucleotide sequence of the promoter of the rice tunglobacterial virus.

[0171] Sequence ID 66: Nucleotide sequence of the promoter of a cleaved rice tunglobacterial virus.

[0172] Sequence ID 67: Nucleotide sequence of the sucrose synthase (Sus1) promoter derived from corn (Zea mays).

[0173] Sequence ID 68: Nucleotide sequence of the sucrose synthase (Sus1) promoter derived from corn (Zea mays).

[0174] Sequence ID 69: Nucleotide sequence of the sucrose synthase (Sus1) promoter from rice (Oryza sativa).

[0175] Sequence ID 70: Nucleotide sequence of the sucrose synthase (Sut1) promoter from rice (Oryza sativa).

[0176] Sequence ID 71: Nucleotide sequence of the YSL2 promoter derived from rice (Oryza sativa).

[0177] Sequence ID 72: Nucleotide sequence of the PPDK promoter derived from corn (Zea mays).

[0178] Sequence ID 73: Nucleotide sequence of the FDA promoter derived from corn (Zea mays).

[0179] Sequence ID 74: Nucleotide sequence of the Nadh-Gogat promoter from rice (Oryza sativa).

[0180] Sequence ID 75: Nucleotide sequence of Actin1 promoter 1 from rice (Oryza sativa).

[0181] Sequence ID 76: Nucleotide sequence of Actin1 promoter 2 from rice (Oryza sativa).

[0182] Sequence ID 77: Nucleotide sequence of Actin2 promoter 1 from rice (Oryza sativa).

[0183] Sequence ID 78: Nucleotide sequence of Actin2 promoter 2 from rice (Oryza sativa).

[0184] Sequence ID 79: Nucleotide sequence of the cauliflower mosaic virus 35S promoter.

[0185] Sequence ID 80: Nucleotide sequence of the polyubiquitin promoter derived from Coix lacryma-jobi.

[0186] Sequence ID 81: Nucleotide sequence of Gos2 promoter 2 from rice (Oryza sativa).

[0187] Sequence ID 82: Nucleotide sequence of the promoter derived from Mirabilis mosaicariovirus.

[0188] Sequence ID 83: Nucleotide sequence of the promoter derived from peanut leucosma striata kalimovirus.

[0189] Sequence ID 84: Nucleotide sequence of GA20 oxidase 2 cDNA derived from sorghum bicolor.

[0190] Sequence ID 85: Nucleotide sequence of the GA20 oxidase 2 coding sequence derived from sorghum bicolor.

[0191] Sequence ID 86: Amino acid sequence of GA20 oxidase 2 derived from sorghum bicolor.

[0192] Sequence ID 87: Genomic nucleotide sequence of GA20 oxidase 2 from sorghum bicolor.

[0193] Sequence ID 88: Nucleotide sequence of GA20 oxidase 2-like cDNA derived from millet (Setarica italica).

[0194] Sequence ID 89: Nucleotide sequence of the GA20 oxidase 2-like coding sequence derived from millet (Setarica italica).

[0195] Sequence ID 90: GA20 oxidase 2-like amino acid sequence derived from millet (Setarica italica).

[0196] Sequence ID 91: GA20 oxidase 2-like genomic nucleotide sequence from millet (Setarica italica).

[0197] Sequence ID 92: Nucleotide sequence of GA20 oxidase 2 cDNA derived from rice (Oryza sativa).

[0198] Sequence ID 93: Nucleotide sequence of the GA20 oxidase 2 coding sequence derived from rice (Oryza sativa).

[0199] Sequence ID 94: Amino acid sequence of the GA20 oxidase 2 gene from rice (Oryza sativa).

[0200] Sequence ID 95: Genomic nucleotide sequence of GA20 oxidase 2 from rice (Oryza sativa).

[0201] Sequence ID 96: Nucleotide sequence of the GA20 oxidase-D2 coding sequence derived from wheat (Triticum aestivum).

[0202] Sequence ID 97: Nucleotide sequence of GA20 oxidase-D2 derived from wheat (Triticum aestivum).

[0203] Sequence ID 98: Genomic nucleotide sequence of GA20 oxidase-D2 derived from wheat (Triticum aestivum).

[0204] Sequence ID 99: Nucleotide sequence of Fe2OG dioxygenase cDNA derived from barley (Hordeum vulgare).

[0205] Sequence ID 100: Nucleotide sequence of the Fe2OG dioxygenase coding sequence derived from barley (Hordeum vulgare).

[0206] Sequence ID 101: Nucleotide sequence of Fe2OG dioxygenase derived from barley (Hordeum vulgare).

[0207] Sequence ID 102: A nucleotide sequence of a highly reliable 2-ODDcDNA derived from sorghum bicolor.

[0208] Sequence ID 103: A nucleotide sequence of a highly reliable 2-ODD coding sequence derived from sorghum bicolor.

[0209] Sequence ID 104: The amino acid sequence of the 2-ODD gene, which is almost certainly derived from sorghum bicolor.

[0210] Sequence ID 105: The most likely genomic nucleotide sequence of the 2-ODD gene derived from sorghum bicolor.

[0211] Sequence ID 106: Nucleotide sequence of flavonol synthase / flavanone 3-hydroxylase-like cDNA derived from millet (Setarica italica).

[0212] Sequence ID 107: Nucleotide sequence of a flavonol synthase / flavanone 3-hydroxylase-like coding sequence derived from millet (Setarica italica).

[0213] Sequence ID 108: Amino acid sequence of a flavonol synthase / flavanone 3-hydroxylase-like gene derived from millet (Setarica italica).

[0214] Sequence ID 109: Nucleotide sequence of a flavonol synthase / flavanone 3-hydroxylase-like gene derived from millet (Setarica italica).

[0215] Sequence ID 110: Nucleotide sequence of naringenin, 2-oxoglutaric acid 3-dioxygenase cDNA derived from rice (Oryza sativa).

[0216] Sequence ID 111: Nucleotide sequence of the naringenin, 2-oxoglutaric acid 3-dioxygenase coding sequence derived from rice (Oryza sativa).

[0217] Sequence ID 112: Amino acid sequence of the naringenin, 2-oxoglutaric acid 3-dioxygenase gene from rice (Oryza sativa).

[0218] Sequence ID 113: Genomic nucleotide sequence of the naringenin, 2-oxoglutaric acid 3-dioxygenase gene from rice (Oryza sativa).

[0219] Sequence ID 114: Nucleotide sequence of Fe2OG dioxygenase cDNA derived from wheat (Triticum aestivum).

[0220] Sequence ID 115: Nucleotide sequence of the Fe2OG dioxygenase coding sequence derived from wheat (Triticum aestivum).

[0221] Sequence ID 116: Amino acid sequence of the Fe2OG dioxygenase gene derived from wheat (Triticum aestivum).

[0222] Sequence ID 117: Genomic nucleotide sequence of the Fe2OG dioxygenase gene derived from wheat (Triticum aestivum).

[0223] Sequence ID 118: Amino acid sequence of the Fe2OG dioxygenase gene derived from barley (Hordeum vulgare).

[0224] Sequence ID 119: A nucleotide sequence of the most reliable GA3-β-dioxygenase 2-2 cDNA derived from sorghum bicolor.

[0225] Sequence ID 120: Nucleotide sequence of the GA3-β-dioxygenase 2-2 coding sequence derived from sorghum bicolor.

[0226] Sequence ID 121: Amino acid sequence of the GA3-β-dioxygenase 2-2 gene derived from sorghum bicolor.

[0227] Sequence ID 122: Genomic nucleotide sequence of the GA3-β-dioxygenase 2-2 gene from sorghum bicolor.

[0228] Sequence ID 123: Nucleotide sequence of GA3-β-dioxygenase 2-2-like cDNA derived from millet (Setarica italica).

[0229] Sequence ID 124: Nucleotide sequence of a GA3-β-dioxygenase 2-2-like coding sequence derived from millet (Setarica italica).

[0230] Sequence ID 125: Amino acid sequence of the GA3-β-dioxygenase 2-2-like gene derived from millet (Setarica italica).

[0231] Sequence ID 126: Genomic nucleotide sequence of the GA3-β-dioxygenase 2-2-like gene derived from millet (Setarica italica).

[0232] Sequence ID 127: Nucleotide sequence of GA3-β-dioxygenase 2-3 cDNA derived from rice (Oryza sativa).

[0233] Sequence ID 128: Nucleotide sequence of the GA3-β-dioxygenase 2-3 coding sequence derived from rice (Oryza sativa).

[0234] Sequence ID 129: Amino acid sequence of the GA3-β-dioxygenase 2-3 gene from rice (Oryza sativa).

[0235] Sequence ID 130: Genomic nucleotide sequence of the GA3-β-dioxygenase 2-3 gene from rice (Oryza sativa).

[0236] Sequence ID 131: Nucleotide sequence of GA3-β-hydroxylase cDNA derived from barley (Hordeum vulgare).

[0237] Sequence ID 132: Nucleotide sequence of the GA3-β-hydroxylase coding sequence derived from barley (Hordeum vulgare).

[0238] Sequence ID 133: Amino acid sequence of the GA3-β-hydroxylase gene derived from barley (Hordeum vulgare).

[0239] Sequence ID 134: Nucleotide sequence of GA3ox-D2 protein cDNA derived from wheat (Triticum aestivum).

[0240] Sequence ID 135: Nucleotide sequence of the GA3ox-D2 protein coding sequence derived from wheat (Triticum aestivum).

[0241] Sequence ID 136: Amino acid sequence of GA3ox-D2 protein derived from wheat (Triticum aestivum).

[0242] Sequence ID 137: GA3ox-D2 protein genome nucleotide sequence derived from wheat (Triticum aestivum).

[0243] Sequence ID 138: Nucleotide sequence of synthetic construct GA20 oxidase_3-A.

[0244] Sequence ID 139: Nucleotide sequence of synthetic construct GA20 oxidase_3-B.

[0245] Sequence ID 140: Nucleotide sequence of synthetic construct GA20 oxidase_3-C.

[0246] Sequence ID 141: Nucleotide sequence of synthetic construct GA20 oxidase_3-D.

[0247] Sequence ID 142: Nucleotide sequence of synthetic construct GA20 oxidase_3-E.

[0248] Sequence ID 143: Nucleotide sequence of synthetic construct GA20 oxidase_3-F.

[0249] Sequence ID 144: Nucleotide sequence of synthetic construct GA20 oxidase_3-G.

[0250] Sequence ID 145: Nucleotide sequence of synthetic construct GA20 oxidase_3-H.

[0251] Sequence ID 146: Nucleotide sequence of synthetic construct GA20 oxidase_3-I.

[0252] Sequence ID 147: Nucleotide sequence of synthetic construct GA20 oxidase_3-J.

[0253] Sequence ID 148: Nucleotide sequence of synthetic construct GA20 oxidase_5-A.

[0254] Sequence ID 149: Nucleotide sequence of synthetic construct GA20 oxidase_5-B.

[0255] Sequence ID 150: Nucleotide sequence of synthetic construct GA20 oxidase_5-C.

[0256] Sequence ID 151: Nucleotide sequence of synthetic construct GA20 oxidase_5-D.

[0257] Sequence ID 152: Nucleotide sequence of synthetic construct GA20 oxidase_5-E.

[0258] Sequence ID 153: Nucleotide sequence of synthetic construct GA20 oxidase_5-F.

[0259] Sequence ID 154: Nucleotide sequence of synthetic construct GA20 oxidase_5-G.

[0260] Sequence ID 155: Nucleotide sequence of synthetic construct GA20 oxidase_5-H.

[0261] Sequence ID 156: Nucleotide sequence of synthetic construct GA20 oxidase_5-I.

[0262] Sequence ID 157: Nucleotide sequence of synthetic construct GA20 oxidase_5-J.

[0263] Sequence ID 158: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-A.

[0264] Sequence ID 159: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-B.

[0265] Sequence ID 160: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-C.

[0266] Sequence ID 161: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-D.

[0267] Sequence ID 162: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-E.

[0268] Sequence ID 163: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-F.

[0269] Sequence ID 164: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-G.

[0270] Sequence ID 165: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-H.

[0271] Sequence ID 166: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-I.

[0272] Sequence ID 167: Nucleotide sequence of synthetic construct GA20 oxidase_3 / 5-J.

[0273] Sequence ID 168: Nucleotide sequence of the GA3 oxidase_1 genome sequence of maize (Zea mays), including the 3kb promoter and 3kb downstream of the 3'UTR.

[0274] Sequence ID 169: Nucleotide sequence of the GA3 oxidase 2 genome sequence of maize (Zea mays), including the 3kb promoter and 3kb downstream of the 3'UTR.

[0275] Sequence ID 170: Nucleotide sequence of the GA3 oxidase 3 genome sequence of maize (Zea mays), including the 3kb promoter and 3kb downstream of the 3'UTR.

[0276] Sequence ID 171: Nucleotide sequence of GA3 oxidase_3 cDNA derived from corn (Zea mays).

[0277] Sequence ID 172: Nucleotide sequence of the coding sequence for GA3 oxidase 3 derived from corn (Zea mays).

[0278] Sequence ID 173: Amino acid sequence of GA3 oxidase_3 protein derived from corn (Zea mays).

[0279] Sequence ID 174: Nucleotide sequence of the GA3 oxidase_1 genome sequence derived from corn (Zea mays) 01DKD2.

[0280] Sequence ID 175: Nucleotide sequence of the GA3 oxidase_2 genome sequence derived from corn (Zea mays) 01DKD2.

[0281] Sequence ID 176: Nucleotide sequence of the GA3 oxidase_3 genome sequence derived from corn (Zea mays) 01DKD2.

[0282] Sequence ID 177: Nucleotide sequence of SP1 (GA3ox1 3'UTR).

[0283] Sequence ID 178: Nucleotide sequence of SP2 (GA3ox1 3'UTR).

[0284] Sequence ID 179: Nucleotide sequence of SP3 (GA3ox1 5'UTR).

[0285] Sequence ID 180: Nucleotide sequence of SP4 (GA3ox1 5'UTR).

[0286] Sequence ID 181: Nucleotide sequence of SP5 (GA3ox1 promoter).

[0287] Sequence ID 182: Nucleotide sequence of SP6 (GA3ox1 promoter).

[0288] Sequence ID 183: Nucleotide sequence of SP7 (GA3ox1 promoter).

[0289] Nucleotide sequence of sequence number 184:SP8 (GA3ox1 promoter).

[0290] SEQ ID NO: 185: Nucleotide sequence of SP9 (GA3ox1 promoter).

[0291] SEQ ID NO: 186: Nucleotide sequence of SP10 (GA3ox1 promoter).

[0292] SEQ ID NO: 187: Nucleotide sequence of SP11 (GA3ox1 promoter).

[0293] SEQ ID NO: 188: Nucleotide sequence of SP12 (GA3ox1 promoter).

[0294] SEQ ID NO: 189: Nucleotide sequence of SP13 (upstream of GA3ox2).

[0295] SEQ ID NO: 190: Nucleotide sequence of SP14 (downstream of GA3ox2).

[0296] SEQ ID NO: 191: Nucleotide sequence of SP15 (upstream of GA3ox3).

[0297] SEQ ID NO: 192: Nucleotide sequence of SP16 (downstream of GA3ox3).

[0298] SEQ ID NO: 193: Nucleotide sequence of a maize reproductive tissue-preferential promoter (from Zea mays).

[0299] SEQ ID NO: 194: Amino acid sequence of a Cpf1 protein derived from a Lachnospiraceae bacterium.

[0300] SEQ ID NO: 195: Amino acid sequence of an NLS signal from Solanum lycopersicum (HSFA1).

[0301] SEQ ID NO: 196: Nucleotide sequence of a synthetic POL III promoter (GSP2262).

[0302] Sequence ID 197: Nucleotide sequence of scaffold RNA SC1 from Lachnospiraceae bacteria.

[0303] Sequence ID 198: Nucleotide sequence of the constitutive maize ubiquitin promoter (Zea mays).

[0304] Sequence ID 199: Amino acid sequence of the Cpf1 protein derived from Francisella tularensis subsp. novicida.

[0305] Sequence ID 200: Amino acid sequence of the NLS signal (NLS5) derived from potato (Solanum lycopersicum).

[0306] Sequence ID 201: Amino acid sequence of the NLS signal (HSFA1) derived from potato (Solanum lycopersicum).

[0307] Sequence ID 202: Nucleotide sequence of scaffold RNA SC1 derived from Francisella tularensis.

[0308] Sequence ID 203: Nucleotide sequence of the synthetic POL III promoter (GSP2269).

[0309] Sequence ID 204: Nucleotide sequence of the AtAG1 gene from Arabidopsis thaliana.

[0310] Sequence ID 205: Nucleotide sequence of SP17(AtAG1).

[0311] The nucleotide sequence of sequence number 206:SP18(AtAG1).

[0312] Nucleotide sequence of SEQ ID NO: 207: SP19(AtAG1).

[0313] SEQ ID NO: 208: Nucleotide sequence of SP20 (AtAG1).

[0314] SEQ ID NO: 209: Nucleotide sequence of SP21 (AtAG1).

[0315] SEQ ID NO: 210: Nucleotide sequence of CaMV35S promoter.

[0316] SEQ ID NO: 211: Nucleotide sequence of BRI1 gene derived from Arabidopsis thaliana.

[0317] SEQ ID NO: 212: Nucleotide sequence of Dwarf1 gene derived from Arabidopsis thaliana.

[0318] SEQ ID NO: 213: Nucleotide sequence of PIN1 gene derived from Arabidopsis thaliana.

[0319] SEQ ID NO: 214: Amino acid sequence of AtAG1 gene derived from Arabidopsis thaliana.

[0320] SEQ ID NO: 215: Amino acid sequence of BRI1 gene derived from Arabidopsis thaliana.

[0321] SEQ ID NO: 216: Amino acid sequence of Dwarf1 gene derived from Arabidopsis thaliana.

[0322] SEQ ID NO: 217: Amino acid sequence of PIN1 gene derived from Arabidopsis thaliana.

[0323] SEQ ID NO: 218: Codon-optimized nucleotide sequence encoding LbCPf1.

[0324] SEQ ID NO: 219: Codon-optimized nucleotide sequence encoding FnCPf1.

[0325] Sequence ID 220: Promoter, leader, and intron nucleotide sequences of the ubiquitin gene from Medicago truncatula.

[0326] Sequence ID 221: Nucleotide sequence of the terminator sequence derived from Medicago truncatula.

[0327] Sequence ID 222: Nucleotide sequence of the U6 promoter from Arabidopsis thaliana.

[0328] Sequence ID 223: Nucleotide sequence of the inverted 5'UTR of AtAG1.

[0329] Sequence ID 224: Nucleotide sequence of the ORF of AtAG1.

[0330] Sequence ID 225: Nucleotide sequence of the terminator of the FbL2 gene derived from Gossypium barbadense.

[0331] Sequence ID 226: Nucleotide sequence of the inverted 3'UTR of AtAG1.

[0332] Sequence ID 227: Nucleotide sequence of SP22 targeting the At.BRI1 region.

[0333] Sequence ID 228: Nucleotide sequence of SP23 targeting the At.BRI1 region.

[0334] Sequence ID 229: Nucleotide sequence of the inverted 5'UTR of AtBRI1.

[0335] Sequence ID 230: Nucleotide sequence of the ORF of AtBRI1.

[0336] Sequence ID 231: Nucleotide sequence of the inverted 3'UTR of AtBRI1.

[0337] Sequence ID 232: Nucleotide sequence of the GA3ox1 genome sequence derived from zea mays, including a 2000 bp promoter sequence, 5'UTR sequence, coding sequence, and 3'UTR sequence upstream of the transcription start site.

[0338] Sequence ID 233: Nucleotide sequence of the 3'UTR of the GA3ox1 gene derived from corn (Zea mays).

[0339] Sequence ID 234: Nucleotide sequence of the genome sequence of the edited allele (S049) of GA3ox1 from maize (Zea mays), including a 2000 bp promoter sequence upstream of the transcription start site, a 5' UTR sequence, a coding sequence, and an edited 3' UTR sequence.

[0340] Sequence ID 235: Nucleotide sequence of the 3'UTR region of the edited allele (S049) of the GA3ox1 gene derived from corn (Zea mays). [Modes for carrying out the invention]

[0341] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs. Where a term is provided in the singular form, the inventors also intend aspects of this disclosure that are described by the plural form of that term. In the event of any inconsistency between terms and definitions used in references incorporated by reference, the terms used herein shall have the definitions provided herein. Other technical terms used have the common meaning in the art in which they are used, as exemplified in various art-specific dictionaries such as “The American Heritage® Science Dictionary” (Editors of the American Heritage Dictionaries, 2011, Houghton Mifflin Harcourt, Boston and New York), “McGraw-Hill Dictionary of Scientific and Technical Terms” (6th edition, 2002, McGraw-Hill, New York), or “Oxford Dictionary of Biology” (6th edition, 2008, Oxford University Press, Oxford and New York). The inventors do not intend to limit themselves to any mechanism or mode of action. References thereto are provided for illustrative purposes only.

[0342] Implementing this disclosure involves, unless otherwise indicated, conventional techniques within the scope of the skill of a person skilled in the art, in the fields of biochemistry, chemistry, molecular biology, microbiology, cell biology, plant biology, genomics, biotechnology, and genetics. For example, Green and Sambrook, Molecular Cloning: A Laboratory Manual, 4th edition (2012); Current Protocols In Molecular Biology (FMAusubel, et al. eds., (1987)); Plant Breeding Methodology (NFJensen, Wiley-Interscience (1988)); the series Methods In Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach(MJMacPherson,BDHames and GRTaylor eds.(1995));Harlow and Lane,eds.(1988)Antibodies,A Laboratory Manual;Animal Cell Culture(RIFreshney,ed.(1987));Recombinant Protein Purification:Principles And Methods,18-1142-75,GE Healthcare Life Sciences;CNStewart,A.Touraev,V.Citovsky,T.Tzfira See eds. (2011) Plant Transformation Technologies (Wiley-Blackwell) and RHSmith (2013) Plant Tissue Culture: Techniques and Experiments (Academic Press, Inc.).

[0343] For example, any references made herein, including all patents, published patent applications, and non-patent publications, are incorporated in their entirety by reference.

[0344] When a group of alternative items is presented, every conceivable combination of members constituting that group of alternative items is specifically envisioned. For example, if an item is selected from the group consisting of A, B, C, and D, the inventors envision each alternative individually (e.g., A alone, B alone, etc.), as well as A, B, and D, A and C, B and C, and so on.

[0345] As used herein, unless the context clearly indicates otherwise, the singular forms "a," "an," and "the" include, for example, multiple references.

[0346] Any composition, nucleic acid molecule, polypeptide, cell, plant, etc., provided herein is specifically intended to be used in any manner provided herein.

[0347] As used herein, the term "heterozygous" refers to a genetic condition in which different alleles reside at corresponding loci on homologous chromosomes.

[0348] As used herein, the term "homozygosity" refers to a genetic condition in which identical alleles are present at corresponding loci on homologous chromosomes.

[0349] As used herein, the term "allele" refers to one of two or more distinct nucleotides or a sequence of 30 nucleotides that arise at a particular gene locus.

[0350] As used herein, the term “heteroallele” refers to the presence of two different alleles at the same locus.

[0351] The term "gene expression" refers to the process of converting genetic information encoded in genomic DNA into RNA (e.g., mRNA, rRNA, tRNA, or 25-snRNA) through gene transcription via the enzymatic reaction of RNA polymerase, and then into proteins through translation of mRNA.

[0352] As used herein, “inhibition of gene expression,” “gene repression,” or “silencing of a target gene,” as well as similar terms and expressions, refer to the absence or observable reduction of levels of protein and / or mRNA products from a target gene. The results of inhibition, repression, or silencing can be confirmed by the phenotype of the cell or organism, or by biochemical methods.

[0353] As used herein, the terms “dsRNA,” “dsRNA region,” or “double-stranded RNA region” refer to two strands of antiparallel polyribonucleic acid held together by base pairing. dsRNA molecules can be formed by intramolecular hybridization or intermolecular hybridization. In some embodiments, dsRNA may comprise a single strand of RNA that self-hybridizes to form a hairpin or stem-loop structure having at least a partial double-stranded structure, including at least one segment that hybridizes to RNA transcribed from a gene targeted for repression. In some embodiments, dsRNA may comprise two separate RNA strands that hybridize via complementary base pairing. The RNA strands may or may not be polyadenylated. The RNA strands may or may not be translated into polypeptides by the cell’s translation machinery. The two strands may be of the same length or different lengths, provided that there is sufficient sequence homology between them to form a double-stranded structure with at least 80%, 90%, 95%, or 100% complementarity along its entire length.

[0354] As used herein, “inverted DNA sequence” may, depending on the context, refer to either the DNA sequence before inversion by the method described herein, or the DNA sequence that results after inversion. Therefore, any reference to a nucleotide sequence in an inverted DNA region may refer to a nucleotide sequence having at least a certain percentage of sequence identity with at least a portion of non-coding sequences, such as untranslated regions (UTRs) or intron sequences. Alternatively, it may be linked to a nucleotide sequence having at least a certain percentage of sequence complementarity with at least a portion of non-coding sequences, such as untranslated regions (UTRs) or intron sequences. Both methods of reference are used interchangeably. When referring to the sequence of RNA transcribed from an inverted DNA sequence, it is usually stated that such transcribed RNA contains an antisense RNA sequence with a certain percentage of complementarity.

[0355] This disclosure enables the development of RNAi (RNA interference)-based editing systems for suppressing locus dominance, which can result in potent, and sometimes undesirable, phenotypes in hybrid crops, without causing phenotypes linked to the edited locus in a homozygous state.

[0356] This disclosure enables the use of mutant alleles exhibiting a potent and undesirable phenotype in the commercial pipeline (where described, an excessively short homozygous parent plant in a seed production field) because the phenotype is absent in the production pipeline, while simultaneously mitigating the potent phenotype to a moderate level in hybrid / heterozygous plants by mitigating the level of RNAi repression resulting from antisense UTR pairing in mRNA transcripts.

[0357] While the present invention is not intended to limit itself to a specific mode of action, it involves inverting part or all of at least one untranslated region (and the resulting mRNA) of a single gene / locus without perturbing the coding region. For example, inversion of part or all of the 5' or 3' untranslated region (or both) results in a silent mutation in the mRNA when present in a homozygous state. mRNA interactions determine the result that the edited mRNA and the inverted UTR do not interact with each other. Therefore, in a uniform pool of edited mRNA (in homozygous edited plants), no interaction occurs and no phenotype is produced. Only in a pool of edited mRNA and mRNA from non-inverted alleles does interspecies antisense base pairing of mRNA generate a double-stranded RNA region and induce RNA silencing mechanisms (ribosome stalling, RNA transcription stalling, post-transcriptional degradation, destabilization, etc.). As a result, viable transcription levels decrease, leading to knockdown phenotypes or reduced expression in plants (see, for example, Roy B, Jacobson A. The intimate relationships of mRNA decay and translation. Trends Genet. 2013;29:691-699).

[0358] Using the methods and compositions disclosed herein, it is also possible to produce moderate phenotypes from the aforementioned alleles having potent unwanted phenotypes without causing any homozygous mutations / edited phenotypes that would have adverse effects on seed or plant production on a larger scale within the hybrid dominant mechanism.

[0359] In one embodiment, a method for editing the genome of a plant cell to modify an endogenous gene is provided, the method comprising: a) generating a first double-strand break and a second double-strand break in a plant cell using a targeted editing technique that targets the untranslated region of the endogenous gene without perturbing the coding region; and b) isolating a modified plant cell containing a modifying allele of the endogenous gene, wherein the modifying allele contains an inverted DNA sequence of at least a portion of the untranslated region of the endogenous gene, the modifying allele produces an RNA transcript containing an antisense sequence of a portion of the untranslated region, and the modifying allele does not contain a sense sequence complementary to the antisense sequence of the portion of the untranslated region. In one embodiment, the transcription of the modifying allele does not result in an RNA molecule containing a stem-loop structure or an intramolecular double-stranded RNA region. In one embodiment, at least a portion of the inverted DNA sequence of the untranslated region of an endogenous gene encodes an antisense RNA sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to at least 18, at least 19, at least 20, at least 25, at least 30, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides of the untranslated region of the gene. In one embodiment, the modified allele does not contain a sense nucleotide sequence of more than 17 nucleotides complementary to the antisense sequence of the inverted portion of the untranslated region.

[0360] In another embodiment, a method for modifying the expression of an endogenous gene in a hybrid plant without affecting the expression of the gene in the parent plant(s) includes: a) identifying an endogenous gene in a plant, wherein the expression of a variant allele of the gene, if present in a homozygous state, results in an unwanted phenotype; b) providing a first plant containing a gene modification allele that includes a nucleic acid region which is an inversion of a part of the gene, wherein the inversion does not affect the translation of the gene modification allele, and the first plant contains the gene modification allele in a homozygous state; c) crossing the first plant with a second plant which contains an unmodified allele of the gene and does not contain an inversion that does not affect the translation of the gene, wherein the unmodified allele is in a homozygous state; and d) obtaining a hybrid seed containing the gene modification allele and unmodified allele in a heterozygous or heteroallelic form. In one embodiment of this method, when modified and unmodified alleles are transcribed into RNA molecules, a double-stranded RNA region may be formed by base pairing between a nucleic acid region, which is an inversion of a portion of the untranslated region of the gene in the RNA transcript of the modified allele, and a nucleic acid region in the RNA transcript of the unmodified allele. The double-stranded RNA region can inhibit the expression of modified and unmodified alleles through RNA silencing mechanisms, including, for example, RNA translation stalling, RNA transcription stalling, resulting destabilization of the RNA molecule, or post-transcriptional degradation of the transcribed RNA molecule. In one embodiment, the nucleic acid region which is an inversion of a portion of the gene occurs during the transcription of an antisense RNA sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementarity to at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 750, at least 1000 consecutive nucleotides in the untranslated region of the gene.

[0361] In yet another embodiment, a plant cell, plant or part thereof or seed containing a modifying allele of an endogenous gene is provided, wherein the modifying allele contains an inverted DNA sequence of at least a portion of the untranslated region of the endogenous gene, the modifying allele produces an RNA transcript containing an antisense sequence of a portion of the untranslated region, and the modifying allele does not contain a sense nucleotide sequence that is more than 17 nucleotides complementary to the antisense sequence of the portion of the untranslated region. In one embodiment, a plant cell, plant or part thereof or seed contains the modifying allele of the endogenous gene described herein in a homozygous state. In one embodiment, the expression of the modifying allele is not reduced. In another embodiment, the modifying allele is present in a heterozygous state. In yet another embodiment, the modifying allele is present in a heteroallelic state. In one embodiment, the other allele is an unmodified version of the endogenous gene and does not contain an inverted DNA sequence of the untranslated region of the endogenous gene. In one embodiment, the expression of the modified and unmodified alleles is reduced in the heteroallelic plant. In one embodiment, the heteroallelic plant exhibits a phenotype of commercial interest. In one embodiment, the plant is maize, and the commercially interesting phenotype is low-growing.

[0362] Various embodiments of this disclosure share several common features, which will be described in more detail later in this specification. It will be apparent that the following descriptions of features can be combined with each of the main embodiments of this disclosure.

[0363] A common feature of all aspects of this disclosure is that the modifying allele of an endogenous gene contains an inversion of part or all of at least one untranslated region of the endogenous gene. As used herein, “untranslated region” is a region of a gene that is transcribed into RNA but not translated into a polypeptide. A polypeptide encoding an endogenous gene includes a transcribed and translated region (“coding region”), but may include a DNA sequence located upstream or at 5' of the transcribed but untranslated coding region (“5'UTR”), and a DNA sequence located downstream or at 3' of the transcribed but untranslated coding region (“3'UTR”). Another untranslated region of an endogenous gene that may be suitable for various aspects of this disclosure is an intron. As used herein, an “intron” is a nucleotide sequence in the genomic DNA and transcribed RNA (called heterokalRNA or hnRNA) of a gene that does not directly code for a protein, particularly a eukaryotic gene, which is removed by RNA splicing during the precursor messenger RNA (mRNA precursor) stage in mRNA maturation. Introns may be located in the 5' or 3'UTR of an endogenous gene. The methods of this disclosure are also applicable to introns in particular, in cases where introns that are partially or entirely inverted in a modified allele are retained during the RNA maturation process, for example, due to interference with splice signals or inversion, by no longer being recognized as such during the RNA splicing process. Inversions of some or all of the untranslated regions within modifying alleles of endogenous genes may occur during transcription in an RNA molecule containing an antisense RNA sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementarity to at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides in the untranslated region of the gene.In one embodiment, the inversion does not result in an intramolecular double-stranded RNA region or stem-loop structure within the transcribed RNA molecule. In other words, the antisense sequence in the modifying allele of an endogenous gene does not have more than 17 complementary sense nucleotides in the modifying allele of the endogenous gene.

[0364] Another common feature of all aspects of this disclosure is that the modifying allele of an endogenous gene includes an inversion of part or all of the untranslated region of the endogenous gene, which is obtained by targeted genome editing techniques via the generation of first and second double-strand breaks in the untranslated region and isolation of the modifying allele of the endogenous gene, in which a portion of the DNA sequence located between the first and second double-strand breaks is reinserted in reverse via non-homologous end joining. Alternatively, a template or donor nucleic acid in which part or all of the untranslated region is present in antisense orientation can be used. Targeted gene editing techniques are well known in the art and include the use of RNA guide effector proteins and RNA guides, or TALE proteins, or custom meganucleases or Zn finger proteins.

[0365] RNA guide effector proteins include CRISPR-Cas effector proteins selected from the Type I CRISPR-Cas system, Type II CRISPR-Cas system, Type III CRISPR-Cas system, Type IV CRISPR-Cas system, Type V CRISPR-Cas system, or Type VI CRISPR-Cas system, or CRISPR-Cas effector proteins derived therefrom, as optionally detailed below. For example, Cas9, C2c1, C2c3, Cas12a (also called Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Csnl, Csx12, Cas10, Csyl, Csy2, Csy3, Csel, Cse2, 30Cs cl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), Csf5 nuclease, Cas12c(C2c3), Cas12d(CasY), Cas12e(CasX This includes CRISPR-Cas effector proteins containing one or more nuclear localization signals, such as CRISPR-Cas effector proteins selected from Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and Cas14c effector proteins. Guide nucleic acids for use with RNA guide effector proteins are also described in more detail below.

[0366] In one embodiment, inversion of at least a portion of at least one untranslated region of an endogenous gene may result from template editing, thereby a single DNA cleavage, such as one mediated by an RNA guide endonuclease and guide RNA, is sufficient to introduce an inverted portion of the untranslated region of an endogenous gene provided via a template or donor nucleic acid, where part or all of the untranslated region is present in antisense orientation. The parameters of the inverted portion of the untranslated region are as described elsewhere in this document. The template nucleic acid or inverted portion may be inserted via a DNA cleavage, either via non-homologous end joining or via a homology-dependent repair mechanism. In the latter case, it is preferable that the template nucleic acid contains at least one homology arm, preferably two, so that the homology arm is adjacent to the inverted portion of the untranslated region, and so that the homology arm has sufficient sequence identity or complementarity to enable hybridization with the nucleotide sequence of the untranslated region, such as the nucleotide sequence of the untranslated region adjacent to the DNA cleavage.

[0367] Another example of template editing described herein that enables the creation of inversions of at least a portion of the untranslated region of an endogenous gene includes prime editing, in which the inverted portion of the untranslated region is introduced via reverse transcription of an extended guide RNA containing a template nucleotide sequence that identifies a target site and further contains a nucleotide sequence corresponding to the untranslated region, and the portion of the untranslated region is reverse-oriented as described elsewhere herein. Such reverse transcription and introduction into a target gene is achieved by the simultaneous action of an RNA guide endonuclease, such as Cas9 or Cas12 derived from such an RNA guide endonuclease or a nickase that generates single-strand DNA breaks or nicks, and a polymerase such as reverse transcriptase, which is optionally fused to one protein (reviewed in Chen and Liu, 2023 (Prime editing for precise and highly versatile genome manipulation, Nat. Rev. Genetic. 24(3); 161-177) or described in Kim et al., 2022 (A novel mechanistic framework for precise sequence replacement using reverse transcriptase and diverse CRISPR-Cas systems DOI:10.1101 / 2022.12.13.520319)). These methods of prime editing are also described in WO2020191153, WO2020191171, WO2020191233, WO2020191234, WO2020191239, WO2020191241, WO2020191242, WO2020191243, WO2020191245, WO2020191246, WO2020191248, WO2020191249, WO2021 / 092130 and WO2022 / 047135 (all included herein by reference).

[0368] Therefore, in one form, a method is provided for editing the genome of plant cells to modify endogenous genes, and the method a) A step of generating a double-strand break or a single-strand break in a plant cell using a targeted editing technique that targets at least one untranslated region of the endogenous gene without perturbing the coding region, b) A step of supplying at least one template nucleic acid to the plant cell, wherein the template nucleic acid comprises a portion of the at least one untranslated region of the endogenous gene in reverse orientation. c) a modified plant cell containing a modifying allele of the endogenous gene, wherein the modifying allele contains an inverted DNA sequence of at least a portion of the at least one untranslated region of the endogenous gene, the modifying allele produces an RNA transcript containing an antisense sequence of a portion of the untranslated region, and the modifying allele does not contain a sense sequence complementary to the antisense sequence of the portion of the untranslated region.

[0369] The template nucleic acid may contain a portion of at least one untranslated region of an endogenous gene in the reverse direction, and may be inserted into the target untranslated region of the endogenous gene by non-homologous end joining. Otherwise, the template nucleic acid contains at least one or two homology arms homologous to the nucleic acid sequence adjacent to the double-strand break, and optionally the homology arms may be adjacent to a portion of at least one untranslated region in the reverse orientation and introduced into the unmodified region in the reverse direction by homology-dependent repair.

[0370] In another form, a method is provided for prime editing the genome of plant cells to modify endogenous genes, and the method a) A step of generating a double-strand break or a single-strand break in a plant cell without perturbing the coding region, using a CRISPR / CAS fusion protein fused to a reverse transcriptase functional region and a guide RNA targeting at least one untranslated region of the endogenous gene, wherein the guide RNA further comprises a nucleotide sequence acid containing a portion of the at least one untranslated region of the endogenous gene in an inverted orientation; b) a modified plant cell comprising the step of isolating a modifying allele of the endogenous gene, wherein the modifying allele comprises an inverted DNA sequence of at least a portion of the at least one untranslated region of the endogenous gene, the modifying allele produces an RNA transcript comprising an antisense sequence of a portion of the untranslated region, and the modifying allele does not contain a sense sequence complementary to the antisense sequence of the portion of the untranslated region.

[0371] As used herein, “template nucleic acid molecule” refers to a nucleic acid molecule containing a nucleic acid sequence to be inserted into a target DNA molecule. In one embodiment, the template nucleic acid molecule comprises single-stranded DNA. In another embodiment, the template nucleic acid molecule comprises double-stranded DNA. In a further embodiment, the template nucleic acid molecule comprises single-stranded RNA. In yet another further embodiment, the template nucleic acid molecule comprises double-stranded RNA. In yet another embodiment, the template nucleic acid molecule comprises both DNA and RNA.

[0372] In one embodiment, the ribonucleoprotein comprises at least one template nucleic acid molecule. In another embodiment, the ribonucleoprotein comprises at least two template nucleic acid molecules.

[0373] In one embodiment, the template nucleic acid molecule contains at least 10 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 25 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 50 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 75 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 100 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 250 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 500 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 750 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 1000 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 2500 nucleotides.

[0374] In one embodiment, the template nucleic acid molecule contains 10 to 2500 nucleotides. In another embodiment, the template nucleic acid molecule contains 25 to 2500 nucleotides. In another embodiment, the template nucleic acid molecule contains 50 to 2500 nucleotides. In another embodiment, the template nucleic acid molecule contains 75 to 2500 nucleotides. In another embodiment, the template nucleic acid molecule contains 100 to 2500 nucleotides. In another embodiment, the template nucleic acid molecule contains 250 to 2500 nucleotides. In another embodiment, the template nucleic acid molecule contains 500 to 2500 nucleotides. In another embodiment, the template nucleic acid molecule contains 25 to 1000 nucleotides. In another embodiment, the template nucleic acid molecule contains 25 to 500 nucleotides. In another embodiment, the template nucleic acid molecule contains 25 to 250 nucleotides.

[0375] The template nucleic acid may be tethered to a gene editing component, in particular a CRISPR / Cas effector protein, by various methods as described in the art. In one embodiment, the template nucleic acid may be tethered to a gene editing component using a HUH endonuclease described in WO2021 / 025999 (incorporated herein by reference).

[0376] Suitable template nucleic acids for homology-dependent recombination-mediated knock-in of reverse repeats in the untranslated regions of endogenous genes as described herein may be single-stranded template nucleic acids, double-stranded template nucleic acids, or circular template nucleic acids as described in the art.

[0377] As described in detail elsewhere in this disclosure, the methods described herein are particularly suited to reducing the expression of endogenous genes having variants that produce phenotypes of potential commercial interest, such as reduced plant height or reduced seed crushing, but which may be harmful or unviable, or produce undesirable phenotypes during plant development or reproduction, especially when present in a homozygous state.

[0378] An example of a gene that may be used in the methods and compositions described herein is ind (non-dehiscent), which is involved in the normal development of the dehiscent zone of pods in Brassica plants. The nucleotide sequence, coding region, genome sequence, and untranslated region of the ind gene can be found in WO2004 / 113542, WO2009 / 068313, or WO2010 / 006732 (incorporated herein by reference).

[0379] Other examples of suitable genes include GA20 oxidase genes containing GA20 oxidase subtype 3 or 5, which are involved in gibberellin biosynthesis and plant height determination, such as the maize-derived GA20 oxidase genes disclosed in WO2018 / 035354, WO2020 / 243361, or WO2020 / 243363 (incorporated herein by reference).

[0380] Further examples of suitable genes include GA3 oxidase genes, such as the maize-derived GA20 oxidase gene disclosed in PCT / US2023 / 062985 (incorporated herein by reference), which include GA3 oxidase subtypes 1, 2, or 3 involved in gibberellin biosynthesis and plant height determination.

[0381] Some GA oxidases in cereal plants consist of families of related GA oxidase genes. For example, maize has a family of at least nine GA20 oxidase genes, including GA20 oxidase_1, GA20 oxidase_2, GA20 oxidase_3, GA20 oxidase_4, GA20 oxidase_5, GA20 oxidase_6, GA20 oxidase_7, GA20 oxidase_8, and GA20 oxidase_9. However, maize also has three known or potential GA3 oxidase genes: GA3 oxidase_1, GA3 oxidase_2, and GA3 oxidase_3. The DNA and protein sequences for each of these GA20 oxidase genes by sequence number are shown in Table 1, and the DNA and protein sequences for each of these GA3 oxidase genes by sequence number are shown in Table 2. [Table 1] [Table 2]

[0382] The genomic DNA sequence of GA20 oxidase_3 is provided in Sequence ID No. 34, and the genomic DNA sequence of GA20 oxidase_5 is provided in Sequence ID No. 35. For the GA20 oxidase_3 gene, Sequence ID No. 34 provides 3000 nucleotides (nucleotides 1-3000) upstream of the GA20 oxidase_3 5'-UTR, with nucleotides 3001-3096 corresponding to the 5'-UTR, nucleotides 3097-3665 corresponding to the first exon, nucleotides 3666-3775 corresponding to the first intron, nucleotides 3776-4097 corresponding to the second exon, nucleotides 4098-5314 corresponding to the second intron, nucleotides 5315-5584 corresponding to the third exon, and nucleotides 5585-5800 corresponding to the 3'-UTR. Sequence ID 34 also provides 3000 nucleotides downstream of the 3'-UTR terminus (nucleotides 5801-8800). For the GA20 oxidase_5 gene, Sequence ID 35 provides 3000 nucleotides (nucleotides 1-3000) upstream of the GA20 oxidase_5 start codon, with nucleotides 3001-3791 corresponding to the first exon, nucleotides 3792-3906 corresponding to the first intron, nucleotides 3907-4475 corresponding to the second exon, nucleotides 4476-5197 corresponding to the second intron, nucleotides 5198-5473 corresponding to the third exon, and nucleotides 5474-5859 corresponding to the 3'-UTR. Sequence ID 35 also provides 3000 nucleotides downstream of the 3'-UTR terminus (nucleotides 5860-8859).

[0383] The genomic DNA sequences of GA3 oxidase_1 are provided in SEQ ID NOs: 36, 168, and 174; the genomic DNA sequences of GA3 oxidase_2 are provided in SEQ ID NOs: 37, 169, and 175; and the genomic DNA sequences of GA3 oxidase_3 are provided in SEQ ID NOs: 170 and 176. SEQ ID NOs: 36 and 37 provide the 5'-UTR sequence, exon sequence, intron, and 3'-UTR sequence for the GA3 oxidase_1 and GA3 oxidase_2 genes, respectively, while SEQ ID NOs: 168 and 174 and SEQ ID NOs: 169 and 175 further provide the upstream and downstream genomic sequences of the GA3 oxidase_1 and GA3 oxidase_2 genes, as well as additional 5' and 3' UTR sequences, respectively.

[0384] Regarding the GA3 oxidase_1 gene, nucleotides 1-29 of SEQ ID NO: 36 correspond to the 5'-UTR, nucleotides 30-514 of SEQ ID NO: 36 correspond to the first exon, nucleotides 515-879 of SEQ ID NO: 36 correspond to the first intron, nucleotides 880-1038 of SEQ ID NO: 36 correspond to the second exon, nucleotides 1039-1158 of SEQ ID NO: 36 correspond to the second intron, nucleotides 1159-1663 of SEQ ID NO: 36 correspond to the third exon, and nucleotides 1664-1788 of SEQ ID NO: 36 correspond to the 3'-UTR. Alternatively, for the GA3 oxidase_1 gene, Sequence ID No. 168 provides 3000 nucleotides (nucleotides 1-3000) upstream of the GA3 oxidase_1 5'-UTR, nucleotides 3001-3161 of Sequence ID No. 168 correspond to the 5'-UTR, nucleotides 3162-3646 of Sequence ID No. 168 correspond to the first exon, nucleotides 3647-4011 of Sequence ID No. 168 correspond to the first intron, nucleotides 4012-4170 of Sequence ID No. 168 correspond to the second exon, nucleotides 4171-4290 of Sequence ID No. 168 correspond to the second intron, nucleotides 4291-4795 of Sequence ID No. 168 correspond to the third exon, and nucleotides 4796-5406 of Sequence ID No. 168 correspond to the 3'-UTR. Sequence ID 168 also provides 3000 nucleotides downstream of the 3'-UTR terminus (nucleotides 5407-8406).Alternatively, regarding the GA3 oxidase_1 gene, sequence number 174 is GA3 oxidase_1 It provides 7620 nucleotides upstream of the 5'-UTR (nucleotides 1-5620 of SEQ ID NO: 174 correspond to the upstream intergene sequence, nucleotides 5621-7620 of SEQ ID NO: 174 correspond to the GA3 oxidase_1 promoter region), nucleotides 7621-8029 of SEQ ID NO: 174 correspond to the 5'-UTR, nucleotides 8030-8514 of SEQ ID NO: 174 correspond to the first exon, nucleotides 8515-8887 of SEQ ID NO: 174 correspond to the first intron, nucleotides 8888-9046 of SEQ ID NO: 174 correspond to the second exon, nucleotides 9047-9166 of SEQ ID NO: 174 correspond to the second intron, nucleotides 9167-9671 of SEQ ID NO: 174 correspond to the third exon, and nucleotides 9672-10276 of SEQ ID NO: 174 correspond to the 3'-UTR. Sequence ID 174 also provides 3951 nucleotides of an intergenetic sequence downstream of the 3'-UTR (nucleotides 10277-14227 of Sequence ID 174).

[0385] Regarding the GA3 oxidase_2 gene, nucleotides 1-38 of SEQ ID NO: 37 correspond to the 5-UTR, nucleotides 39-532 of SEQ ID NO: 37 correspond to the first exon, nucleotides 533-692 of SEQ ID NO: 37 correspond to the first intron, nucleotides 693-851 of SEQ ID NO: 37 correspond to the second exon, nucleotides 852-982 of SEQ ID NO: 37 correspond to the second intron, nucleotides 983-1445 of SEQ ID NO: 37 correspond to the third exon, and nucleotides 1446-1698 of SEQ ID NO: 37 correspond to the 3'-UTR. Alternatively, for the GA3 oxidase_2 gene, SEQ ID NO: 169 provides 3000 nucleotides (nucleotides 1-3000) upstream of the GA3 oxidase_2 5'-UTR, nucleotides 3001-3056 of SEQ ID NO: 169 correspond to the 5'-UTR, nucleotides 3057-3550 of SEQ ID NO: 169 correspond to the first exon, nucleotides 3551-3710 of SEQ ID NO: 169 correspond to the first intron, nucleotides 3711-3869 of SEQ ID NO: 169 correspond to the second exon, nucleotides 3870-3991 of SEQ ID NO: 169 correspond to the second intron, nucleotides 3992-4463 of SEQ ID NO: 169 correspond to the third exon, and nucleotides 4464-4581 of SEQ ID NO: 169 correspond to the 3'-UTR. Sequence ID 169 also provides 3000 nucleotides downstream of the 3'-UTR terminus (nucleotides 4582-7581).Alternatively, regarding the GA3 oxidase_2 gene, sequence number 175 is GA3 oxidase_2 It provides 7285 nucleotides upstream of the 5'-UTR (nucleotides 1-5385 of SEQ ID NO: 175 correspond to the upstream intergene sequence, and nucleotides 5386-7385 of SEQ ID NO: 175 correspond to the GA3 oxidase_2 promoter region), nucleotides 7386-7831 of SEQ ID NO: 175 correspond to the 5'-UTR, nucleotides 7832-7926 of SEQ ID NO: 175 correspond to the first exon, nucleotides 7927-8086 of SEQ ID NO: 175 correspond to the first intron, nucleotides 8087-8245 of SEQ ID NO: 175 correspond to the second exon, nucleotides 8246-8371 of SEQ ID NO: 175 correspond to the second intron, nucleotides 8372-8861 of SEQ ID NO: 175 correspond to the third exon, and nucleotides 8862-8967 of SEQ ID NO: 175 correspond to the 3'-UTR. Sequence ID 175 also provides 7630 nucleotides of the intergenetic sequence downstream of the 3'-UTR (nucleotides 8968-16597 of Sequence ID 175).

[0386] Alternatively, for the GA3 oxidase_3 gene, SEQ ID NO: 170 provides 3000 nucleotides (nucleotides 1-3000) upstream of the GA3 oxidase_3 5'-UTR, nucleotides 3001-3130 of SEQ ID NO: 170 corresponds to the 5'-UTR, nucleotides 3131-3483 of SEQ ID NO: 170 corresponds to the first exon, nucleotides 3484-3582 of SEQ ID NO: 170 corresponds to the first intron, nucleotides 3583-3907 of SEQ ID NO: 170 corresponds to the second exon, nucleotides 3908-3998 of SEQ ID NO: 170 corresponds to the second intron, nucleotides 3999-4274 of SEQ ID NO: 170 corresponds to the third exon, and nucleotides 4275-4332 of SEQ ID NO: 170 corresponds to the 3'-UTR. Sequence ID 170 also provides 3000 nucleotides downstream of the 3'-UTR terminus (nucleotides 4333-7332). Alternatively, for the GA3 oxidase_3 gene, Sequence ID 176 provides GA3 oxidase_3 It provides 7546 nucleotides upstream of the 5'-UTR (nucleotides 1-5546 of SEQ ID NO: 176 correspond to the upstream intergene sequence, and nucleotides 5547-7546 of SEQ ID NO: 176 correspond to the GA3 oxidase 3 promoter region), nucleotides 7547-7751 of SEQ ID NO: 176 correspond to the 5'-UTR, nucleotides 7752-8104 of SEQ ID NO: 176 correspond to the first exon, nucleotides 8105-8205 of SEQ ID NO: 176 correspond to the first intron, nucleotides 8206-8530 of SEQ ID NO: 176 correspond to the second exon, nucleotides 8531-8621 of SEQ ID NO: 176 correspond to the second intron, nucleotides 8622-8903 of SEQ ID NO: 176 correspond to the third exon, and nucleotides 8904-9178 of SEQ ID NO: 176 correspond to the 3'-UTR. Sequence ID 176 also provides 6176 nucleotides of the intergenetic sequence downstream of the 3'-UTR terminal (nucleotides 9179-15354 of Sequence ID 176).

[0387] For sequence numbers 174, 175, and 176, it should be noted that the nucleotide boundary between the upstream promoter and the intergenic region does not have to be exactly the same as the coordinates described above, and the promoter, expression (e.g., enhancer or repressor), and / or regulator(s) for transcription may be located in the upstream intergenic sequence of each gene.

[0388] Further examples of suitable genes include Anther Ear1 (GRMZM2G081554)dwarf4 (GRMZM2G065635)brs1-brassinosteroid synthesis 1, Anubias plant 1 (GRMZM2G057000), brassinosteroid receptor ZmBRI1a / ZmBRI1b (GRMZM2G048294 / GRMZM2G449830), meristem development gene compact plant 2 (GRMZM2G064732), and ZMWRKY60 (disclosed in CN116217684, which is included herein by reference).

[0389] Other examples of genes that can be used in the methods and compositions disclosed herein include Agamous, Bri1, Dwarf1, Pin1 (disclosed herein) from Arabidopsis, or orthologous genes from other plants.

[0390] The genomic DNA sequence of the Arabidopsis agamous (AG) gene (AT4G18960) is provided in SEQ ID NO: 204, and the protein sequence is provided in SEQ ID NO: 214. The genomic DNA sequence of the Arabidopsis brassinosteroid-insensitive 1 (Bri1) gene is provided in SEQ ID NO: 211, and the protein sequence is provided in SEQ ID NO: 215. The genomic DNA sequence of the Arabidopsis DWARF1 gene (AT3G19820) is provided in SEQ ID NO: 212, and the protein sequence is provided in SEQ ID NO: 216. The genomic DNA sequence of the Arabidopsis PIN-FORMED1 gene (Pin1) (AT1G73590) is provided in SEQ ID NO: 213, and the protein sequence is provided in SEQ ID NO: 217.

[0391] For the AG gene, nucleotides 1-501 and 1057-1060 of sequence number 204 correspond to the 5-UTR, and nucleotides 5418-5648 of sequence number 204 correspond to the 3'UTR. For the Bri1 gene, nucleotides 1-165 of sequence number 211 correspond to the 5-UTR, and nucleotides 3757-4167 of sequence number 211 correspond to the 3'UTR. For the Dwarf1 gene, nucleotides 664-699 of sequence number 212 correspond to the 5-UTR, and nucleotides 2482-2700 of sequence number 212 correspond to the 3'UTR. For the Pin1 gene, nucleotides 1-99 of sequence number 213 correspond to the 5-UTR, and nucleotides 3205-3506 of sequence number 213 correspond to the 3'UTR.

[0392] Nucleic acids and amino acids. The use of the terms “polynucleotide” or “nucleic acid molecule” is not intended to limit this disclosure to polynucleotides containing deoxyribonucleic acid (DNA). For example, ribonucleic acid (RNA) molecules are also envisioned. Those skilled in the art will recognize that polynucleotides and nucleic acid molecules may include deoxyribonucleotides, ribonucleotides, or combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogs. The polynucleotides of this disclosure also encompass all forms of sequences, including but not limited to single-stranded, double-stranded, hairpin, and stem-loop structures. In one embodiment, the nucleic acid molecule provided herein is a DNA molecule. In another embodiment, the nucleic acid molecule provided herein is an RNA molecule. In one embodiment, the nucleic acid molecule provided herein is single-stranded. In another embodiment, the nucleic acid molecule provided herein is double-stranded.

[0393] As used herein, the term “recombinant” with respect to nucleic acid (DNA or RNA) molecules, proteins, constructs, vectors, etc. means a nucleic acid or amino acid molecule or sequence that is artificial and not normally found in nature, and / or exists in a context in which it is not normally found in nature, and includes polynucleotide molecules, proteins, constructs, etc., which do not exist in a contiguous or very close proximity to one another in nature without human intervention, and / or contain at least two polynucleotide or protein sequences that are heterogeneous to each other.

[0394] As used herein, the term “heterogeneous” means a nucleotide / polypeptide derived from an alien species, or, if derived from the same species, substantially modified from its natural form in composition and / or at a genomic locus by intentional human intervention. A “heterogeneous” or “recombinant” nucleotide sequence is a nucleotide sequence that is not naturally associated with the host cell into which it is introduced, and includes multiple copies of a naturally occurring nucleotide sequence that are not naturally occurring.

[0395] As used herein, the term “homology” means a nucleotide / polypeptide that normally occurs in a particular species and that can be functionally linked to other components occurring in that species without intentional human intervention.

[0396] In one embodiment, the methods and compositions provided herein include a vector. As used herein, the term "vector" refers to a DNA molecule used as a medium for transporting exogenous genetic material into a cell.

[0397] In one embodiment, one or more polynucleotide sequences derived from the vector are stably integrated into the plant genome. In another embodiment, one or more polynucleotide sequences derived from the vector are stably integrated into the genome of a plant cell.

[0398] In one embodiment, the first nucleic acid sequence and the second nucleic acid sequence are provided in a single vector. In another embodiment, the first nucleic acid sequence is provided in a first vector, and the second nucleic acid sequence is provided in a second vector.

[0399] As used herein, the term "polypeptide" refers to a chain of at least two covalently linked amino acids. Polypeptides may be encoded by polynucleotides provided herein. An example of a polypeptide is a protein. Proteins provided herein may be encoded by nucleic acid molecules provided herein.

[0400] Nucleic acids can be isolated using techniques conventional in the art. For example, nucleic acids can be isolated using any method, including, but not limited to, recombinant nucleic acid techniques and / or polymerase chain reaction (PCR). Common PCR techniques are described, for example, in PCR Primer: A Laboratory Manual, Dieffenbach & Dveksler, Eds., Cold Spring Harbor Laboratory Press, 1995. Recombinant nucleic acid techniques include, for example, restriction enzyme digestion and ligation, which can be used to isolate nucleic acids. Isolated nucleic acids can also be chemically synthesized as either a single nucleic acid molecule or a series of oligonucleotides. Polypeptides can be purified from natural sources (e.g., biological samples) by known methods such as DEAE ion exchange, gel filtration, and hydroxyapatite chromatography. Polypeptides can also be purified, for example, by expressing nucleic acids in an expression vector. Furthermore, purified polypeptides can be obtained by chemical synthesis. The degree of purity of polypeptides can be measured using any suitable method, for example, column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0401] Nucleic acids can be detected using hybridization, though not limited to this method. Hybridization between nucleic acids is discussed in detail in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).

[0402] Polypeptides can be detected using antibodies. Techniques for detecting polypeptides using antibodies include enzyme-linked immunosorbent assay (ELISA), Western blotting, immunoprecipitation, and immunofluorescence. The antibodies provided herein may be polyclonal or monoclonal antibodies. Antibodies having a specific binding affinity to the polypeptides provided herein can be prepared using methods well known in the art. The antibodies provided herein can be conjugated to a solid support, such as a microtiter plate, using methods known in the art.

[0403] As used herein in relation to two or more nucleotide or protein sequences, the terms “percent identity” or “percent identity” are calculated by (i) comparing two optimally aligned sequences (nucleotides or proteins) across a comparison window; (ii) determining the number of positions in both sequences where identical nucleic acid bases (in the case of nucleotide sequences) or amino acid residues (in the case of proteins) occur to obtain the number of matching positions; (iii) dividing the number of matching positions by the total number of positions in the comparison window; and (iv) multiplying this quotient by 100% to obtain the percentage identity. If “percent identity” is calculated with respect to a reference sequence without specifying a particular comparison window, the percentage identity is determined by dividing the number of matching positions across the alignment region by the sum of the lengths of the reference sequences. Therefore, with respect to the object of this application, if the two sequences (query and subject) are optimally aligned (allowing for gaps in their alignment), the “identity percentage” for the query sequence is equal to the number of identical positions between the two sequences divided by the total number of positions in the query sequence over its length (or comparison window), and then multiplied by 100%. When the sequence identity percentage is used in reference to proteins, it is recognized that non-identical residue positions are often distinguished by conservative amino acid substitutions, where the amino acid residue is substituted with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity), and thus does not change the functional properties of the molecule. If sequences differ in conservative substitutions, the sequence identity percentage may be adjusted upward to compensate for the conservative nature of the substitutions. Sequences that differ by such conservative substitutions are said to have “sequence similarity” or “similarity.”

[0404] Where used herein in relation to two nucleotide sequences, the term “sequence complementarity percentage” or “complementarity percentage” is similar to the concept of identity percentage, but refers to the percentage of nucleotides in the query sequence that best base-pair or hybridize with nucleotides in the subject sequence if the query sequence and subject sequence were arranged linearly and best base-paired without secondary folding structures such as loops, stems, or hairpins. Such complementarity percentages may be between two DNA strands, two RNA strands, or between a DNA strand and an RNA strand. The “complementarity percentage” can be calculated by (i) determining the two nucleotide sequences to best base-pair or hybridize across a comparison window in a linear and fully unfolded arrangement (i.e., without folding or secondary structures), (ii) determining the number of base-pairing positions between the two sequences across the comparison window to find the number of complementary positions, (iii) dividing the number of complementary positions by the total number of positions in the comparison window, and (iv) multiplying this quotient by 100% to obtain the complementarity percentage of the two sequences. The optimal base pairing of two sequences can be determined based on known pairings of nucleotide bases such as GC, AT, and AU by hydrogen bonding. When the "complementarity percentage" is calculated in relation to a reference sequence without specifying a particular comparison window, the identity percentage is determined by dividing the number of complementary positions between the two linear sequences by the total length of the reference sequence. Therefore, for the purposes of this application, if the two sequences (query and subject) are optimally base-paired (mismatched or unbase-paired nucleotides are acceptable), the "complementarity percentage" for the query sequence is equal to the number of base-pair positions between the two sequences divided by the total number of positions in the query sequence over its length, and then multiplied by 100%.

[0405] Various pairwise or multiple sequence alignment algorithms and programs, such as ClustalW or Basic Local Alignment Search Tool (BLAST®), are known in the art and can be used to compare the sequence identity or similarity between two or more nucleotide or protein sequences with respect to the optimal sequence alignment for calculating the sequence identity percentage. Other alignment and comparison methods are known in the art, but the alignment and identity percentage (including the identity percentage ranges described above) between two sequences can be determined by the ClustalW algorithm. For example, Chenna R. et.al., “Multiple sequence alignment with the Clustal series of programs,” Nucleic Acids Research 31:3497-3500 (2003), Thompson JD et.al., “Clustal W: Improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice,” Nucleic Acids Research 22:4673-4680 (1994), Larkin MA et.al., “Clustal W and Clustal See tool. J.Mol.Biol.215:403-410(1990) (its contents and the entire disclosure are incorporated herein by reference).

[0406] As used herein, a first nucleic acid molecule can “hybridize” with a second nucleic acid molecule via non-covalent interactions (e.g., Watson-Crick base pairing) in a sequence-specific and antiparallel manner under appropriate temperature and solution ionic strength in vitro and / or in vivo conditions (i.e., one nucleic acid specifically binds to a complementary nucleic acid). As is known in the art, standard Watson-Crick base pairings include adenine (A) pairing with thymine (T), adenine (A) pairing with uracil (U), and guanine (G) pairing with cytosine (C). Furthermore, with respect to hybridization between two RNA molecules (e.g., dsRNA), it is also known in the art that guanine bases pair with uracil. For example, G / U base pairing is partially involved in the degeneracy (i.e., redundancy) of the genetic code in the context of tRNA anticodon base pairing with codons in mRNA. In the context of this disclosure, guanine in the protein-binding segment (double-stranded RNA) of the target DNA-targeting RNA molecule is considered complementary to uracil, and vice versa. Therefore, if a G / U base pair can be constructed at a given nucleotide position in the protein-binding segment (double-stranded RNA) of the target DNA-targeting RNA molecule, that position is not considered incomplementary, but rather complementary.

[0407] Hybridization and washing conditions are well known and are exemplified in Sambrook, J., Fritsch, E. and Maniatis, T., Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly Chapter 11 and Table 11.1 therein; and Sambrook, J. and Russell, W., Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (2001). Temperature and ionic strength conditions determine the "stringency" of hybridization.

[0408] Hybridization requires that two nucleic acids contain complementary sequences, although mismatches between bases are possible. The appropriate conditions for hybridization between two nucleic acids depend on the length and degree of complementarity of the nucleic acids, which are variable factors well known in the art. The greater the degree of complementarity between two nucleotide sequences, the greater the melting temperature (Tm) for the hybrid of nucleic acids having those sequences. For hybridization between nucleic acids with short-stretch complementarity (e.g., complementarity over 35 nucleotides or less), the location of the mismatch becomes important (see Sambrook et al.). Typically, the length of a hybridizable nucleic acid is at least 10 nucleotides. Examples of minimum hybridizable nucleic acid lengths include at least 15 nucleotides, at least 18 nucleotides, at least 20 nucleotides, at least 22 nucleotides, at least 25 nucleotides, and at least 30 nucleotides). Furthermore, those skilled in the art will recognize that temperature and washing solution salt concentration can be adjusted as needed depending on factors such as the length and degree of complementarity of the complementary region.

[0409] In this art, polynucleotide sequences do not need to be specifically hybridizable or 100% complementary to their target nucleic acid sequence for hybridizability to occur. Furthermore, polynucleotides can hybridize with a target polynucleotide across one or more segments such that intervening or adjacent segments do not participate in the hybridization event (e.g., loop or hairpin structures). For example, an antisense nucleic acid in which 18 out of 20 nucleotides of an antisense compound are complementary to the target region and therefore specifically hybridize exhibits 90 percent complementarity. In this example, the remaining non-complementary nucleotides may be clustered with complementary nucleotides or scattered among complementary nucleotides, and do not need to be in close proximity to each other or to complementary nucleotides. The percentage of complementarity between specific stretches of nucleic acid sequences within a nucleic acid can be determined in a routine procedure using the BLAST® (basic local alignment search tools) and PowerBLAST programs known in the art (see Altschul et.al., J.Mol.Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656) or the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.) using the default settings with the Smith and Waterman algorithm (Adv.Appl.Math., 1981, 2, 482-489).

[0410] promoter In some embodiments, guide RNA and RNA guide effector proteins are provided to plant cells as recombinant nucleic acids, and their coding regions are operably ligated to a plant expression promoter. In one embodiment, the RNA guide effector protein is expressed from nucleic acids operably ligated to a POL II promoter. In another embodiment, the RNA guide effector protein is expressed from nucleic acids operably ligated to a POL II promoter or a POL III promoter.

[0411] As used herein, “promoter” is a nucleotide sequence that controls or regulates the transcription of a nucleotide sequence (e.g., a coding sequence) that is operably associated with a promoter. A coding sequence controlled or regulated by a promoter may encode a polypeptide and / or functional RNA. “Promoter” may also refer to a nucleotide sequence that contains a binding site for RNA polymerase II and directs the initiation of transcription. Generally, promoters are found 5' or upstream of the start of the coding region of the corresponding coding sequence. Promoters may include other elements that act as regulators of gene expression, such as promoter regions. These include TATA box consensus sequences, often CAAT box consensus sequences (Breathnach and Chambon (1981) Annu. Rev. Biochem. 50:349). In plants, CAAT boxes may be replaced with AGGA boxes (Messing et al., (1983) in Genetic Engineering of Plants, T. Kosuge, C. Meredith and A. Hollaender (eds.), Plenum Press, pp. 211-227).

[0412] Useful promoters in the present invention include, for example, constitutive, inducible, temporally restricted, developmentally regulated, chemically regulated, tissue-preferential, and / or tissue-specific promoters used in the preparation of recombinant nucleic acid molecules, such as "synthetic nucleic acid constructs" or "protein-RNA complexes." Various types of these promoters are known in the art.

[0413] The choice of promoter may vary depending on the temporal and spatial requirements for expression, and also on the host cell being transformed. Promoters for many different organisms are well known in the art. Based on the extensive knowledge existing in the art, an appropriate promoter can be selected for a particular host organism of interest. For example, much is known about upstream promoters of highly constitutively expressed genes in model organisms, and such knowledge can be easily accessed and implemented in other systems as needed.

[0414] In some embodiments, promoters that function in plants may be used with the constructs of the present invention. Non-limiting examples of promoters useful for driving expression in plants include the promoter for the RubisCo small subunit gene 1 (PrbcS1), the promoter for the actin gene (Pactin), the promoter for the nitrate reductase gene (Pnr), and the promoter for the replication carbonic anhydrase gene (Pdca1) (see Walker et al. (2005) Plant Cell Rep. 23:727-735; Li et al. (2007) Gene 403:132-142; Li et al. (2010) Mol Biol. Rep. 37:1143-1154). PrbcS1 and Pactin are constitutive promoters, while Pnr and Pdca1 are inductive promoters. Pnr is induced by nitrates and inhibited by ammonium (Li et al. (2007) Gene 403:132-142), and Pdca1 is induced by salts (Li et al. (2010) Mol Biol. Rep. 37:1143-1154). In some embodiments, the promoter useful in the present invention is the RNA polymerase II (Pol II) promoter.

[0415] Examples of constitutive promoters useful for plants include the Cestrum virus promoter (cmp) (US Patent No. 7,166,770), the Rineactin 1 promoter (Wang et al. (1992) Mol. Cell. Biol. 12:3399-3406, and US Patent No. 5,641,876), the CaMV35S promoter (Odell et al. (1985) Nature 313:810-812), the CaMV19S promoter (Lawton et al. (1987) Plant Mol. Biol. 9:315-324), the nos promoter (Ebert et al. (1987) Proc. Natl. Acad. Sci. USA 84:5745-5749), and the Adh promoter (Walker et al. (1987) Proc. Natl. Acad. Sci. USA 84:6624- This includes the sucrose synthase promoter (6629), the sucrose synthase promoter (Yang & Russell (1990) Proc. Natl. Acad. Sci. USA 87:4144-4148), and the ubiquitin promoter. Constitutive promoters derived from ubiquitin accumulate in many cell types. The ubiquitin promoter has been cloned from several transgenic plants, such as sunflower (Binet et al. (1991) Plant Science 79:87-94), maize (Christensen et al. (1989) Plant Molec. Biol. 12:619-632), and Arabidopsis thaliana (Norris et al. (1993) Plant Molec. Biol. 21:895-906). The maize ubiquitin promoter (UbiP) has been developed in transgenic monocotyledonous systems, and its sequence and the vector constructed for monocotyledonous transformation are disclosed in Patent Publication EP0342926. The ubiquitin promoter is suitable for the expression of the nucleotide sequence of the present invention in transgenic plants, particularly monocotyledonous plants.Furthermore, the promoter expression cassette described by McElroy et al. ((1991) Mol. Gen. Genet. 231:150-160) can be readily modified for the expression of the nucleotide sequence of the present invention and is particularly suitable for use in monocotyledonous plant hosts.

[0416] In some embodiments, tissue-specific / tissue-preferential promoters may be used for the expression of heterologous polynucleotides in plant cells. Tissue-specific or preferential expression patterns include, but are not limited to, green tissue-specific or preferential, root-specific or preferential, stem-specific or preferential, flower-specific or preferential, or pollen-specific or preferential expression patterns. Promoters suitable for expression in green tissue include many that regulate genes involved in photosynthesis, many of which have been cloned from both monocots and dicots. In one embodiment, a useful promoter in the present invention is the maize PEPC promoter derived from the phosphenolcarboxylase gene (Hudspeth & Grula (1989) Plant Molec. Biol. 12:579-589). Non-limiting examples of tissue-specific promoters include those related to genes encoding seed storage proteins (such as β-conglycinin, cluciferin, napine, and phaseolin), zein or oil proteins (such as oleosin), or proteins involved in fatty acid biosynthesis (including acyl carrier proteins, stearoyl-ACP desaturase, and fatty acid desaturase (fad2-1)), as well as other nucleic acids expressed during embryonic development (such as Bce4, e.g., Kridl et al. (1991) Seed Sci. Res. 1:209-219, and EP Patent No. 255378). Tissue-specific or tissue-preferential promoters useful for expressing the nucleotide sequences of the present invention in plants, particularly maize, include, but are not limited to, those that direct expression in root, pith, leaves, or pollen. Such promoters are disclosed, for example, in WO93 / 07278, which is incorporated herein by reference in its entirety.Other non-limiting examples of tissue-specific or tissue-preferential promoters useful in conjunction with the present invention include the Wataruvisco promoter disclosed in U.S. Patent No. 6,040,504, the Inescrose synthase promoter disclosed in U.S. Patent No. 5,604,121, the root-specific promoter described by de Framond ((1991) FEBS 290:103-106, Ciba-Geigy EP0452269), the stem-specific promoter described in U.S. Patent No. 5,625,136 (Ciba-Geigy) that drives the expression of the maize trpA gene, the Cestrum yellow leaf curling virus promoter disclosed in WO01 / 73087, and pollen-specific or preferential promoters including, but not limited to, ProOsLPS10 and ProOsLPS11 derived from rice (Nguyen et al. (2015) Plant Biotechnol. Reports). Examples of promoters include those described in 9(5):297-306), ZmSTK2_USP from maize (Wang et al. (2017) Genome 60(6):485-495), LAT52 and LAT59 from tomato (Twell et al. (1990) Development 109(3):705-713), Zm13 (US Patent No. 10,421,972), PLA2-δ promoter from Arabidopsis (US Patent No. 7,141,424), and / or ZmC5 promoter from maize (International PCT Publication No. WO1999 / 042587).

[0417] Further examples of plant tissue-specific / tissue-preferential promoters include, but are not limited to, root hair-specific cis-elements (RHE) (Kim et al. (2006) The Plant Cell 18:2958-2970), root-specific promoters RCc3 (Jeong et al. (2010) Plant Physiol. 153:185-197) and RB7 (US Patent No. 5459252), lectin promoters (Lindstrom et al. (1990) Der. Genet. 11:160-167; and Vodkin (1983) Prog. Clin. Biol. Res. 138:87-98), maize alcohol dehydrogenase 1 promoter (Dennis et al. (1984) Nucleic Acids Res. 12:3983-4000), and S-adenosyl-L-methionine synthetase (SAMS) (Vander Mijnsbrugge et al. (1996) Plant and Cell Physiology 37(8):1108-1115), maize light-harvesting complex promoter (Bansal et al. (1992) Proc. Natl. Acad. Sci. USA 89:3654-3658), maize heat shock protein promoter (O'Dell et al. (1985) EMBO J.5:451-458, and Rochester et al. (1986) EMBO J.5:451-458), pea small subunit RuBP carboxylase promoter (Cashmore, “Nuclear genes encoding the small subunit of ribulose-l,5-bisphosphate carboxylase” pp.29-39 In: Genetic Engineering of Plants (Hollaender ed., Plenum Press 1983; and Poulsen et al.) al. (1986) Mol. Gen. Genet. 205:193-200), Ti plasmid mannopin synthase promoter (Langridge et al. (1989) Proc. Natl. Acad. Sci.USA 86:3219-3223), Ti plasmid nopalin synthase promoter (Langridge et al. (1989) above), petunia chalcone isomerase promoter (van Tunen et al. (1988) EMBO J.7:1257-1263), mame glycine-rich protein 1 promoter (Keller et al. (1989) Genes Dev.3:1639-1646), truncated CaMV 35S promoter (O'Dell et al. (1985) Nature 313:810-812), potato patatin promoter (Wenzler et al. (1989) Plant Mol. Biol.13:347-354), root cell promoter (Yamamoto et al. (1990) Nucleic Acids Res.18:7449), maize zein promoter (Kriz et al. al. (1987) Mol.Gen.Genet. 207:90-98, Langridge et al. (1983) Cell 34:1015-1022, Reina et al. (1990) Nucleic Acids Res. 18:6425, Reina et al. (1990) Nucleic Acids Res. 18:7449, and Wandelt et al. (1989) Nucleic Acids Res. 17:2354), globulin-1 promoter (Belanger et al. (1991) Genetics 129:863-872), α-tubulin cab promoter (Sullivan et al. (1989) Mol.Gen.Genet. 215:431-440), PEPCase promoter (Hudspeth & Grula (1989) Plant Examples include the R gene complex-related promoter (Mol. Biol. 12:579-589), the R gene complex-related promoter (Chandler et al. (1989) Plant Cell 1:1175-1183), and the chalcone synthase promoter (Franken et al. (1991) EMBO J. 10:2605-2612).

[0418] Useful for seed-specific expression include the pea bicillin promoter (Czako et al. (1992) Mol. Gen. Genet. 235:33-40, and seed-specific promoters disclosed in U.S. Patent No. 5,625,136). Promoter useful for expression in mature leaves is a promoter that can be switched at the onset of senescence, such as the Arabidopsis-derived SAG promoter (Gan et al. (1995) Science 270:1986-1988).

[0419] Plant expression promoters useful for the methods and compositions described herein include egg cell-preferential or embryonic tissue-preferential promoters described in WO2022 / 056139 (which is incorporated herein in whole), such as the DSUL1 promoter, EA1 promoter, ES4 promoter, DMC1 promoter, Mps1 promoter, Adf1 promoter, or EAL promoter.

[0420] Other plant expression promoters useful for the present invention include the floral tissue-preferential or floral cell-preferential promoters described in PCT / US2023 / 065042 (which is incorporated herein in its entirety).

[0421] Furthermore, promoters that function in chloroplasts can be used. Non-limiting examples of such promoters include the bacteriophage T3 gene 9 5'UTR and other promoters disclosed in U.S. Patent No. 7,579,516. Other promoters useful in the present invention include, but are not limited to, the S-E9 small subunit RuBP carboxylase promoter and the Kunitz trypsin inhibitor gene promoter (Kti3).

[0422] In some embodiments, different promoters of at least two recombinant constructs or cassettes expressing a guide RNA containing a spacer sequence complementary to a target site, such as a genomic target site described herein, may be selected for the RNA polymerase III (Pol III) promoter. In some embodiments, the POL III promoter may be the U6 promoter, H1 promoter, 5S promoter, adenovirus 2 (Ad2)VAI promoter, tRNA promoter, and 7SK promoter. See, for example, Schramm and Hernandez, 2002, Genes & Development, 16:2593-2620 (the entire work is incorporated herein by reference).

[0423] In some embodiments, the POL III promoter may be derived from a gene encoding nuclear small RNA (sRNA). In some embodiments, the POL III promoter may be selected from the maize, tomato, and soybean U6, U3, U2, U5, and 7SL snRNA promoters disclosed in WO2015 / 131101 (which is incorporated herein by reference in its entirety), and may include the snRNA promoter sequence of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 146-149, 160-166, 201, or 283, as included in the attached sequence listing.

[0424] In some embodiments, the POL III promoter may be a synthetic snRNA promoter, for example, the snRNA promoter described in WO2022 / 232407 (which is incorporated herein by reference in its entirety), and may include the snRNA promoter sequences of SEQ ID NOs. 1 to 10 contained on the attached sequence.

[0425] In some embodiments, the POL III promoter may be a chimeric POL III promoter. In some embodiments, the POL III promoter may be a variant of the chimeric POL III promoter. In some embodiments, a variant of POL III is provided that, when optimally aligned to the reference sequence provided herein, has at least about 85 percent identity, at least about 86 percent identity, at least about 87 percent identity, at least about 88 percent identity, at least about 89 percent identity, at least about 90 percent identity, at least about 91 percent identity, at least about 92 percent identity, at least about 93 percent identity, at least about 94 percent identity, at least about 95 percent identity, at least about 96 percent identity, at least about 97 percent identity, at least about 98 percent identity, or at least about 99 percent identity with respect to the reference sequence, and has the promoter activity disclosed herein.

[0426] regulatory factor Further modulifacts useful in the present invention include, but are not limited to, introns, enhancers, termination sequences, and / or 5' and 3' untranslated regions.

[0427] Introns useful in this invention may be introns identified in plants, isolated from plants, and then inserted into expression cassettes for use in plant transformation. As will be understood by those skilled in the art, introns may contain sequences necessary for autoexcision and are incorporated in-frame into nucleic acid constructs / expression cassettes. Introns may be used as spacers to separate multiple protein-coding sequences within a single nucleic acid construct, or they may be used inside a single protein-coding sequence to stabilize, for example, mRNA. When used within protein-coding sequences, they are inserted "in-frame" with excision sites included. Introns may also associate with promoters to improve or modify expression. As an example, promoter / intron combinations useful in this invention include, but are not limited to, combinations of the maize Ubi1 promoter and introns.

[0428] Non-limiting examples of introns useful in the present invention include introns derived from the ADHI gene (e.g., Adh1-S introns 1, 2, and 6), the ubiquitin gene (Ubi1), the RuBisCO small subunit (rbcS) gene, the RuBisCO large subunit (rbcL) gene, the actin gene (e.g., the actin-1 intron), the pyruvate dehydrogenase kinase gene (pdk), the nitrate reductase gene (nr), the replication carbonic anhydrase gene 1 (Tdca1), the psbA gene, the atpA gene, or any combination thereof.

[0429] Guide nucleic acids As used herein, “guide nucleic acid” refers to a nucleic acid that forms a complex with a guide nuclease (e.g., Cas12a, CasX, but not limited to) and subsequently guides the ribonucleoprotein to a specific sequence of a target nucleic acid molecule, where the guide nucleic acid and the target nucleic acid molecule share complementary sequences. In one embodiment, the ribonucleoprotein comprises at least one guide nucleic acid molecule.

[0430] In one embodiment, the guide nucleic acid comprises DNA. In another embodiment, the guide nucleic acid comprises RNA. In one embodiment, the guide nucleic acid comprises DNA, RNA, or a combination thereof. In one embodiment, the guide nucleic acid is single-stranded. In another embodiment, the guide nucleic acid is at least partially double-stranded.

[0431] If the guide nucleic acid contains RNA, it may be called “guide RNA.” In another embodiment, the guide nucleic acid contains both DNA and RNA. In another embodiment, the guide RNA is single-stranded. In another embodiment, the guide RNA is double-stranded. In yet another embodiment, the guide RNA is partially double-stranded.

[0432] As used herein, “guide nucleic acid,” “guide RNA,” “gRNA,” “CRISPR RNA / DNA,” “crRNA,” or “crDNA” refers to at least one spacer sequence and at least one repeat sequence (e.g., a repeat of the V-Cas12a CRISPR-Cas system, or a fragment or part thereof, a repeat of the II-Cas9 CRISPR-Cas system, or a V-C2c1 CRISPR) that are complementary to the target DNA (e.g., a protospacer). Repeats or fragments of the Cas system, repeats of the CRISPR-Cas system, e.g., C2c3, Cas12a (also called Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Cas1, Cas1B, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Cs This refers to nucleic acids including m2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4(dinG), and / or Csf5), where the repeat sequence may be ligated to the 5' and / or 3' ends of the spacer sequence. The design of the gRNA of the present invention may be based on a type I, type II, type III, type IV, type V, or type VI CRISPR-Cas system.

[0433] In some embodiments, the Cas12a gRNA may include a repeat sequence (full length or a portion thereof ("handle"), e.g., a pseudoknot-like structure) and a spacer sequence from 5' to 3'.

[0434] In some embodiments, the guide nucleic acid may comprise multiple repetitive-spacer sequences (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more repetitive-spacer sequences) (e.g., repetitive-spacer-repetitive, e.g., repetitive-spacer-repetitive-spacer-repetitive-spacer-repetitive-spacer, etc.). The guide nucleic acid of the present invention is synthetic, artificial, and not found in nature. The gRNA can be very long and can be used as an aptamer (as in the MS2 mobilization strategy), or other RNA structures hanging from the spacers can be used. The guide RNA may comprise a donor template for introducing specific modifications to the target sequence.

[0435] As used herein, “repetitive sequence” refers to any repetitive sequence of a wild-type CRISPR-Cas locus (e.g., Cas9, Cas12a, C2c1, etc.) or a repetitive sequence of a synthetic crRNA that functions with a CRISPR-Cas effector protein encoded by the nucleic acid construct of the present invention. Repetitive sequences useful in the present invention may be any known or subsequently identified repetitive sequences of a CRISPR-Cas locus (e.g., type I, type II, type III, type IV, type V, or type VI), or synthetic repetitive sequences designed to function with type I, type II, type III, type IV, type V, or type VI CRISPR-Cas systems. Repetitive sequences may include hairpin and / or stem-loop structures. In some embodiments, repetitive sequences may form pseudoknot-like structures at their 5' end (i.e., “handle”). Therefore, in some embodiments, the repeat sequence may be identical or substantially identical to the repeat sequences derived from the wild-type CRISPR-Cas locus I, II, III, IV, V, and / or VI CRISPR-Cas locus. The repeat sequence derived from the wild-type CRISPR-Cas locus can be determined via established algorithms, such as using CRISPRfinder provided via CRISPRb (see Grissa et al. (2007) Nucleic Acids Res. 35 (Web server issue): W52-7). In some embodiments, the repeat sequence or a portion thereof is ligated at its 3' end to the 5' end of a spacer sequence, thereby forming a repeat-spacer sequence (e.g., guide nucleic acid, guide RNA / DNA, crRNA, crDNA).

[0436] In some embodiments, the repetitive sequence comprises, essentially consists of, or comprises at least 10 nucleotides, depending on whether a guide nucleic acid containing a particular repetitive sequence and the repetitive sequence is being processed (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50-100 or more nucleotides, or any range or value therein). In some embodiments, the repeat sequence substantially consists of, or comprises, approximately 10 to 20, 10 to 30, 10 to 45, 10 to 50, 15 to 30, 15 to 40, 15 to 45, 15 to 50, 20 to 30, 20 to 40, 20 to 50, 30 to 40, 40 to 80, 50 to 100, or more nucleotides.

[0437] The repetitive sequence ligated to the 5' end of the spacer sequence may contain a portion of the repetitive sequence (e.g., consecutive nucleotides of wild-type repetitive sequences 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more). In some embodiments, a portion of the repeat sequence ligated to the 5' end of the spacer sequence may be about 5 to about 10 consecutive nucleotide lengths (e.g., about 5, 6, 7, 8, 9, 10 nucleotides) and may have at least 90% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more) sequence identity with respect to the same region (e.g., the 5' end of the wild-type CRISPR-Cas repeat nucleotide sequence). In some embodiments, a portion of the repeat sequence may include a pseudoknot-like structure at its 5' end (i.e., the "handle").

[0438] As used herein, “spacer sequence” is a nucleotide sequence complementary to a portion of the target nucleic acid (e.g., target DNA) (e.g., a protospacer). The spacer sequence may be fully complementary or substantially complementary to the target nucleic acid (e.g., at least about 70% complementary (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more)). In some embodiments, the spacer sequence may have one, two, three, four, or five mismatches compared to the target nucleic acid, and the mismatches may be continuous or discontinuous. In some embodiments, the spacer sequence may have 70% complementarity to the target nucleic acid. In other embodiments, the spacer nucleotide sequence may have 80% complementarity to the target nucleic acid. In yet another embodiment, the spacer nucleotide sequence may be 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% complementarity to the target nucleic acid (protospacer). In some embodiments, the spacer sequence is 100% complementarity to the target nucleic acid. The spacer sequence may have a length of about 15 to about 30 nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides, or any range or value thereof). Thus, in some embodiments, the spacer sequence may have complete or substantial complementarity over a region of the target nucleic acid (e.g., protospacer) that is at least about 15 to about 30 nucleotides in length. In some embodiments, the spacer is about 20 nucleotides long. In some embodiments, the spacer is about 21, 22, or 23 nucleotides long. In some embodiments, the spacer sequence may include one of the sequences of SEQ ID NOs. 88-90, or any combination thereof.

[0439] In some embodiments, the 5' region of the spacer sequence of the guide nucleic acid may be identical to the target DNA, while the 3' region of the spacer may be substantially complementary to the target DNA (e.g., the spacer in a type V CRISPR-Cas system), or the 3' region of the spacer sequence of the guide nucleic acid may be identical to the target DNA, while the 5' region of the spacer may be substantially complementary to the target DNA (e.g., a type II CRISPR-Cas system), and therefore the overall complementarity of the spacer sequence to the target DNA may be less than 100%. For example, in the guide of a type V CRISPR-Cas system, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 nucleotides in the 5' region (i.e., seed region) of a 20-nucleotide spacer sequence may be 100% complementary to the target DNA, while the remaining nucleotides in the 3' region of the spacer sequence may be substantially complementary to the target DNA (e.g., at least about 70% complementary). In some embodiments, the first 1–8 nucleotides at the 5' end of the spacer sequence (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8 nucleotides and any range therein) may be 100% complementary to the target DNA, and the remaining nucleotides in the 3' region of the spacer sequence may be substantially complementary to the target DNA (e.g., at least about 50% complementary (e.g., 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more)).

[0440] As a further example, in the guide of a type II CRISPR-Cas system, for example, in the 3' region (i.e., seed region) of a 20-nucleotide spacer sequence, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 nucleotides may be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence may be substantially complementary to the target DNA (e.g., at least about 70%). In some embodiments, the first 1 to 10 nucleotides at the 3' end of the spacer sequence (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides, and any range among them) may be 100% complementary to the target DNA, and the remaining nucleotides in the 5' region of the spacer sequence may be substantially complementary to the target DNA (e.g., at least about 50% complementary (e.g., 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more, or any range or value thereof)).

[0441] In some embodiments, the seed region of the spacer may be about 8 to about 10 nucleotides long, about 5 to about 6 nucleotides long, or about 6 nucleotides long.

[0442] In one embodiment, the guide nucleic acid includes guide RNA. In another embodiment, the guide nucleic acid includes at least one guide RNA. In another embodiment, the guide nucleic acid includes at least two guide RNAs. In another embodiment, the guide nucleic acid includes at least three guide RNAs. In another embodiment, the guide nucleic acid includes at least five guide RNAs. In another embodiment, the guide nucleic acid includes at least ten guide RNAs.

[0443] In another embodiment, the guide nucleic acid comprises at least 10 nucleotides. In another embodiment, the guide nucleic acid comprises at least 11 nucleotides. In another embodiment, the guide nucleic acid comprises at least 12 nucleotides. In another embodiment, the guide nucleic acid comprises at least 13 nucleotides. In another embodiment, the guide nucleic acid comprises at least 14 nucleotides. In another embodiment, the guide nucleic acid comprises at least 15 nucleotides. In another embodiment, the guide nucleic acid comprises at least 16 nucleotides. In another embodiment, the guide nucleic acid comprises at least 17 nucleotides. In another embodiment, the guide nucleic acid comprises at least 18 nucleotides. In another embodiment, the guide nucleic acid comprises at least 19 nucleotides. In another embodiment, the guide nucleic acid comprises at least 20 nucleotides. In another embodiment, the guide nucleic acid comprises at least 21 nucleotides. In another embodiment, the guide nucleic acid comprises at least 22 nucleotides. In another embodiment, the guide nucleic acid comprises at least 23 nucleotides. In another embodiment, the guide nucleic acid comprises at least 24 nucleotides. In another embodiment, the guide nucleic acid comprises at least 25 nucleotides. In another embodiment, the guide nucleic acid comprises at least 26 nucleotides. In another embodiment, the guide nucleic acid comprises at least 27 nucleotides. In another embodiment, the guide nucleic acid comprises at least 28 nucleotides. In another embodiment, the guide nucleic acid comprises at least 30 nucleotides. In another embodiment, the guide nucleic acid comprises at least 35 nucleotides. In another embodiment, the guide nucleic acid comprises at least 40 nucleotides. In another embodiment, the guide nucleic acid comprises at least 45 nucleotides. In another embodiment, the guide nucleic acid comprises at least 50 nucleotides.

[0444] In another embodiment, the guide nucleic acid contains 10 to 50 nucleotides. In another embodiment, the guide acid contains 10 to 40 nucleotides. In another embodiment, the guide nucleic acid contains 10 to 30 nucleotides. In another embodiment, the guide nucleic acid contains 10 to 20 nucleotides. In another embodiment, the guide nucleic acid contains 16 to 28 nucleotides. In another embodiment, the guide nucleic acid contains 16 to 25 nucleotides. In another embodiment, the guide nucleic acid contains 16 to 20 nucleotides.

[0445] In one embodiment, the guide nucleic acid contains at least 70% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 75% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 80% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 85% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 90% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 91% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 92% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 93% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 94% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 95% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 96% sequence complementarity to the target site. In one embodiment, the guide nucleic acid contains at least 97% sequence complementarity to the target site. In another embodiment, the guide nucleic acid contains at least 98% sequence complementarity to the target site. In another embodiment, the guide nucleic acid contains at least 99% sequence complementarity to the target site. In another embodiment, the guide nucleic acid contains 100% sequence complementarity to the target site. In yet another embodiment, the guide nucleic acid contains at least 70% to 100% sequence complementarity to the target site. In yet another embodiment, the guide nucleic acid contains at least 80% to 100% sequence complementarity to the target site. In yet another embodiment, the guide nucleic acid contains at least 90% to 100% sequence complementarity to the target site.

[0446] In one embodiment, the guide nucleic acid can hybridize to the target site.

[0447] As described above, some guide nucleases, such as CasX and Cas9, require another non-coding RNA component called trans-activated crRNA (tracrRNA) to have functional activity. The guide nucleic acid molecules provided herein can be formed by combining crRNA and tracrRNA into a single nucleic acid molecule referred to herein as “single guide RNA” (sgRNA). The gRNA guides the active CasX complex to a target site in the target sequence, where CasX can cleave the target site. In other embodiments, crRNA and tracrRNA are provided as separate nucleic acid molecules.

[0448] In one embodiment, the guide nucleic acid includes crRNA. In another embodiment, the guide nucleic acid includes tracrRNA. In yet another embodiment, the guide nucleic acid includes sgRNA.

[0449] The methods and compositions disclosed herein, when expressed as single guide RNAs or from a single recombinant construct, are expected to be useful in combination with RNA guide effector proteins to increase the editing efficiency of guide RNAs that result in poor editing or low editing efficiency at target sites. As used herein, “guide RNAs with low editing efficiency” result in cleavage efficiencies of less than 30%, less than 25%, or less than 20% when an assay is performed, for example, involving the introduction of such guide RNA or a nucleic acid construct encoding such guide RNA in a plant protoplast containing a target site recognized by the guide RNA, and the determination of the percentage of protoplasts containing insertions / deletions at the target site by determining the number of reads, for example, by sequencing. The potential editing efficiency of a particular guide RNA can be estimated in silico as described in US2023091138 (which is incorporated herein in its entirety).

[0450] RNA guide nuclease A guide nuclease is a nuclease that forms a complex (e.g., ribonucleoprotein) with a guide nucleic acid molecule (e.g., guide RNA) and then guides this complex to a target site in a target sequence. One non-limiting example of a guide nuclease is the CRISPR nuclease.

[0451] CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) nucleases (e.g., Cas9, CasX, Cas12a (also known as Cpf1), CasY, MAD7®) are proteins found in bacteria that are guided to a target nucleic acid molecule by a guide RNA ("gRNA"), where the endonuclease can then cleave the single or double strand of the target nucleic acid molecule. Although CRISPR nucleases originate in bacteria, many have been shown to function in eukaryotic cells.

[0452] While not limited to any particular scientific theory, CRISPR nucleases form a complex with a guide RNA (gRNA) that hybridizes with a complementary target site, thereby guiding the CRISPR nuclease to the target site. In class II CRISPR-Cas systems, a CRISPR array containing spacers is transcribed during encounter with recognized invasive DNA and processed into small interfering CRISPR RNA (crRNA). The crRNA contains a repetitive sequence and a spacer sequence complementary to a specific protospacer sequence in the invasive pathogen. The spacer sequence can be designed to be complementary to a target sequence in the eukaryotic genome.

[0453] CRISPR nucleases associate with their respective crRNAs in their active form. CasX, a class II endonuclease similar to Cas9, requires another non-coding RNA component, called trans-activating crRNA (tracrRNA), to have functional activity. The nucleic acid molecules provided herein can be formed by combining crRNA and tracrRNA into a single nucleic acid molecule, referred herein as a “single guide RNA” (sgRNA). Cas12a or MAD7® do not require tracrRNA to be guided to the target site. For Cas12a or MAD7®, crRNA alone is sufficient. The gRNA guides the active CRISPR nuclease complex to the target site, where the CRISPR nuclease can cleave the target site.

[0454] When an RNA guide CRISPR nuclease and guide RNA form a complex, the entire system is called a "ribonucleoprotein." The ribonucleoproteins provided herein may also include additional nucleic acids or proteins.

[0455] A prerequisite for cleavage of a target site by CRISPR ribonucleoproteins is the presence of a conserved protospacer-adjacent motif (PAM) near the target site. Depending on the CRISPR nuclease, cleavage may occur within a specific number of nucleotides from the PAM site (e.g., 18-23 nucleotides in the case of Cas12a). PAM sites are required only for type I and type II CRISPR-related proteins, and different CRISPR endonucleases recognize different PAM sites. Cas12a can recognize at least the following PAM sites, namely TTTN and YTN. CasX can recognize at least the following PAM sites: TTCN, TTCA, and TTC. The MAD7® nuclease recognizes the T-rich PAM sequence YTTN and appears to prefer TTTN over CTTN PAM (wherein T is thymine, C is cytosine, A is adenine, Y is thymine or cytosine, and N is thymine, cytosine, guanine, or adenine).

[0456] Cas12a is a class II, type V CRISPR / Cas system RNA guide nuclease. When Cas12a nucleases cleave double-stranded DNA molecules, they produce staggered cuts. Staggered cuts of double-stranded DNA create an overhang of single-stranded DNA with at least one nucleotide. This is in contrast to blunt-end cuts (such as those produced by Cas9) which do not produce a single-stranded DNA overhang when cleaving double-stranded DNA.

[0457] In one embodiment, the Cas12a nuclease provided herein is the Cas12a (LbCas12a) nuclease of the Lachnospiraceae bacterium. In another embodiment, the Cas12a nuclease provided herein is the Francisella novicida Cas12a (FnCas12a) nuclease. In one embodiment, the Cas12a nuclease is selected from the group consisting of LbCas12a and FnCas12a.

[0458] In one embodiment, Cas12a nuclease, or the nucleic acid encoding Cas12a nuclease, is defined as belonging to Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francis It is derived from a bacterial genera selected from the group consisting of ella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium, Acidaminococcus, Peregrinibacteria, Butyrivibrio, Parcubacteria, Smithella, Candidatus, Moraxella, and Leptospira.

[0459] In one embodiment, the Cas12a nuclease is encoded by a polynucleotide containing a sequence having at least 80% identity to a polynucleotide selected from SEQ ID NO: 218 or SEQ ID NO: 219. In another embodiment, the Cas12a nuclease is encoded by a polynucleotide containing a sequence having at least 85% identity to a polynucleotide selected from SEQ ID NO: 218 or SEQ ID NO: 219. In another embodiment, the Cas12a nuclease is encoded by a polynucleotide containing a sequence having at least 90% identity to a polynucleotide selected from SEQ ID NO: 218 or SEQ ID NO: 219. In another embodiment, the Cas12a nuclease is encoded by a polynucleotide containing a sequence having at least 95% identity to a polynucleotide selected from SEQ ID NO: 218 or SEQ ID NO: 219. In another embodiment, the Cas12a nuclease is encoded by a polynucleotide containing a sequence having at least 96% identity to a polynucleotide selected from SEQ ID NO: 218 or SEQ ID NO: 219. In another embodiment, the Cas12a nuclease is encoded by a polynucleotide containing a sequence having at least 97% identity to a polynucleotide selected from SEQ ID NO: 218 or SEQ ID NO: 219. In another embodiment, the Cas12a nuclease is encoded by a polynucleotide containing a sequence having at least 98% identity to a polynucleotide selected from SEQ ID NO: 218 or SEQ ID NO: 219. In another embodiment, the Cas12a nuclease is encoded by a polynucleotide containing a sequence having at least 99% identity to a polynucleotide selected from SEQ ID NO: 218 or SEQ ID NO: 219. In another embodiment, the Cas12a nuclease is encoded by a polynucleotide containing a sequence having 100% identity to a polynucleotide selected from SEQ ID NO: 218 or SEQ ID NO: 219.

[0460] In one embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence having at least 80% identity to an amino acid sequence selected from SEQ ID NO: 194 or SEQ ID NO: 199. In another embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence having at least 85% identity to an amino acid sequence selected from SEQ ID NO: 194 or SEQ ID NO: 199. In another embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence having at least 90% identity to an amino acid sequence selected from SEQ ID NO: 194 or SEQ ID NO: 199. In another embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence having at least 95% identity to an amino acid sequence selected from SEQ ID NO: 194 or SEQ ID NO: 199. In another embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence having at least 96% identity to an amino acid sequence selected from SEQ ID NO: 194 or SEQ ID NO: 199. In another embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence having at least 97% identity with an amino acid sequence selected from SEQ ID NO: 194 or SEQ ID NO: 199. In another embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence having at least 98% identity with an amino acid sequence selected from SEQ ID NO: 194 or SEQ ID NO: 199. In another embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence having at least 99% identity with an amino acid sequence selected from SEQ ID NO: 194 or SEQ ID NO: 199. In another embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence having 100% identity with an amino acid sequence selected from SEQ ID NO: 194 or SEQ ID NO: 199.

[0461] In one embodiment, the Cas12a provided herein is a variant Lachnospiraceae bacterial Cas12a (LbCas12a) nuclease with enhanced DNA cleavage activity at non-canonical TTTT protospacer-adjacent motifs, such as that described in US2021 / 0348144 (which is incorporated herein by reference in its entirety). In another embodiment, the Cas12a provided herein is a variant Lachnospiraceae bacterial Cas12a (LbCas12a) nuclease with enhanced activity, such as that described in US20230040148 (which is incorporated herein by reference in its entirety), including LbCas12a-ultra, which has the N527R and E795L substitutions in its amino acid sequence (reference amino acid sequence is SEQ ID NO: 194).

[0462] In one embodiment, the Cas12a provided herein is a variant Lachnospiraceae bacterial Cas12a (LbCas12a) nuclease that recognizes the PAM variant TYCV having G532R and K595R substitutions in its amino acid sequence (reference amino acid sequence is Sequence ID No. 4), as disclosed in WO2016205711 (the entirety of which is incorporated herein by reference), or a variant Lachnospiraceae bacterial Cas12a (LbCas12a) nuclease that recognizes the PAM variant TATT having G532R, K538R, and Y524R substitutions in its amino acid sequence (reference amino acid sequence is Sequence ID No. 194).

[0463] CasX is a class II CRISPR-Cas nuclease identified in the phylum Bacteria, specifically in the species Deltaproteobacteria and Planctomyces. Similar to Cas12a, CasX nucleases produce staggered cuts when cleaving double-stranded DNA molecules. However, unlike Cas12a, CasX nucleases require crRNA and tracrRNA, or a single guide RNA, to target and cleave the target nucleic acid.

[0464] In some embodiments, the CasX nucleases provided herein are CasX nucleases derived from the phylum Deltaprotebacteria. In some embodiments, the CasX nucleases provided herein are CasX nucleases derived from the phylum Planctomyces. Additional preferred CasX nucleases, without limitation, are described in WO2019 / 084148, which is incorporated herein by reference in whole.

[0465] MAD7® (also known as ErCas12a) is an engineered nuclease of the class II VA CRISPR-Cas (Cas12a / Cpf1) family with low levels of homology to the canonical Cas12a nuclease. The MAD7® nuclease produces a staggered cleavage when cleaving double-stranded DNA molecules. The MAD7® nuclease was initially identified in Eubacterium rectale. It requires only crRNA, similar to canonical Cas12a. The nucleotide sequence encoding ErCas12a / MAD7® can be found in supplemental data (sequence S1) provided in Lin et al., 2021, Journal of Genetics and Genomics 48, pages 444-451.

[0466] In one embodiment, the guide nuclease capable of generating staggered cuts in a double-stranded DNA molecule is selected from the group consisting of Cas12a, MAD7(registered trademark), and CasX.

[0467] In one embodiment, the guide nuclease is an RNA guide nuclease. In another embodiment, the guide nuclease is a CRISPR nuclease. In another embodiment, the guide nuclease is a Cas12a nuclease. In another embodiment, the guide nuclease is a CasX nuclease. In another embodiment, the guide nuclease is a MAD7® nuclease.

[0468] As used herein, “nuclear localization signal” (NLS) refers to an amino acid sequence that “tags” a protein for uptake into the nucleus of a cell. In one embodiment, the nucleic acid molecules provided herein encode a nuclear localization signal. In another embodiment, the nucleic acid molecules provided herein encode two or more nuclear localization signals.

[0469] In one embodiment, the Cas12a nuclease provided herein includes a nuclear localization signal. In one embodiment, the nuclear localization signal is located at the N-terminus of the Cas12a nuclease. In a further embodiment, the nuclear localization signal is located at the C-terminus of the Cas12a nuclease. In yet another embodiment, the nuclear localization signal is located at both the N-terminus and the C-terminus of the Cas12a nuclease.

[0470] In one embodiment, the CasX nuclease provided herein includes a nuclear localization signal. In one embodiment, the nuclear localization signal is located at the N-terminus of the CasX nuclease. In a further embodiment, the nuclear localization signal is located at the C-terminus of the CasX nuclease. In yet another embodiment, the nuclear localization signal is located at both the N-terminus and the C-terminus of the CasX nuclease.

[0471] In one embodiment, the MAD7® nuclease provided herein includes a nuclear localization signal. In one embodiment, the nuclear localization signal is located at the N-terminus of the MAD7® nuclease. In a further embodiment, the nuclear localization signal is located at the C-terminus of the MAD7® nuclease. In yet another embodiment, the nuclear localization signal is located at both the N-terminus and the C-terminus of the MAD7® nuclease.

[0472] In one embodiment, the ribonucleoprotein includes at least one nuclear localization signal. In another embodiment, the ribonucleoprotein includes at least two nuclear localization signals.

[0473] Different species exhibit specific biases for particular codons of specific amino acids. Codon bias (differences in codon use between organisms) often correlates with messenger RNA (mRNA) translation efficiency, which in turn is thought to depend, among other things, on the characteristics of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs within a cell generally reflects the codons most frequently used in peptide synthesis. Therefore, genes can be tuned for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database" available at www.kazusa.codon or www.codon.jp.codon, and these tables can be adapted to numerous methods. See Nakamura et al., 2000, Nucl. Acids Res. 28:292. Computer algorithms for codon-optimizing specific sequences for expression in specific plant cells are also available, such as Gene Forge (Aptagen; Jacobus, PA).

[0474] As used herein, “codon optimization” refers to the process of modifying a nucleic acid sequence for improved expression in a target plant cell by replacing at least one codon in the sequence (e.g., at least codons 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more) with a codon that is more frequently or most frequently used in the gene of the plant cell, while maintaining the original amino acid sequence (e.g., by introducing silent mutations).

[0475] In one embodiment, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all codons) in the sequence encoding the guide nuclease correspond to the codon most frequently used for a particular amino acid. In another embodiment, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all codons) in the sequence encoding the Cas12a nuclease, CasX nuclease, or MAD7® nuclease correspond to the codon most frequently used for a particular amino acid. The use of codons in plants is described in Campbell and Gowri, 1990, Plant Physiol., 92:1-11, and Murray et al., 1989, Nucleic Acids Res., 17:477-98, each of which is incorporated herein by reference as a whole.

[0476] In one embodiment, the nucleic acid molecule encodes a guide nuclease that is codon-optimized for plants. In another embodiment, the nucleic acid molecule encodes a Cas12a nuclease that is codon-optimized for plants. In another embodiment, the nucleic acid molecule encodes a CasX nuclease that is codon-optimized for plants. In another embodiment, the nucleic acid molecule encodes a MAD7® nuclease that is codon-optimized for plants.

[0477] In another embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for plant cells. In another embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for monocotyledonous plant species. In another embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for dicotyledonous plant species. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for gymnosperm species. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for angiosperm species. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for maize cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for soybean cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for rice cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for wheat cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for cotton cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for sorghum cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for alfalfa cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for sugarcane cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for Arabidopsis cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for tomato cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized guide nuclease for cucumber cells.In a further embodiment, the nucleic acid molecules provided herein encode a guide nuclease that is codon-optimized for potato cells. In a further embodiment, the nucleic acid molecules provided herein encode a guide nuclease that is codon-optimized for onion cells.

[0478] In one embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for plant cells. In another embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for monocotyledonous plant species. In yet another embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for dicotyledonous plant species. In yet another embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for gymnosperm species. In yet another embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for angiosperm species. In yet another embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for maize cells. In yet another embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for soybean cells. In yet another embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for rice cells. In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for wheat cells. In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for cotton cells. In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for sorghum cells. In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for alfalfa cells. In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for sugarcane cells. In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for Arabidopsis cells. In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for tomato cells.In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for cucumber cells. In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for potato cells. In a further embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that is codon-optimized for onion cells.

[0479] In another embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for plant cells. In another embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for monocotyledonous plant species. In another embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for dicotyledonous plant species. In one embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for gymnosperm species. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for angiosperm species. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for maize cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for soybean cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for rice cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for wheat cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for cotton cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for sorghum cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for alfalfa cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for sugarcane cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for Arabidopsis cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for tomato cells.In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for cucumber cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for potato cells. In a further embodiment, the nucleic acid molecule provided herein encodes a CasX nuclease that is codon-optimized for onion cells. In another embodiment, the nucleic acid molecule provided herein encodes a MAD7® nuclease that is codon-optimized for plant cells. In one embodiment, the nucleic acid molecule provided herein encodes a MAD7® nuclease that is codon-optimized for monocotyledonous plant species. In one embodiment, the nucleic acid molecule provided herein encodes a MAD7® nuclease that is codon-optimized for dicotyledonous plant species. In a further embodiment, the nucleic acid molecule provided herein encodes a MAD7® nuclease that is codon-optimized for gymnosperm species. In a further embodiment, the nucleic acid molecule provided herein encodes a MAD7® nuclease that is codon-optimized for angiosperm species. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for maize cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for soybean cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for rice cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for wheat cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for cotton cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for sorghum cells.In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for alfalfa cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for sugarcane cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for Arabidopsis cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for tomato cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for cucumber cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for potato cells. In a further embodiment, the nucleic acid molecule provided herein encodes a codon-optimized MAD7® nuclease for onion cells.

[0480] In some embodiments, the guide nuclease is Cas9, C2c1, C2c3, Cas12a (also called Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Csnl, Csx12, Cas10, Csyl, Csy2, Csy3, Csel, Cse2, 30Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, The following can be selected from Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), Csf5 nuclease, Cas12c(C2c3), Cas12d(CasY), Cas12e(CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and Cas14c effector proteins.

[0481] In some embodiments, CRISPR / Cas effector proteins useful in the present invention may contain mutations in the nuclease active site (e.g., RuvC, HNH, e.g., the RuvC site of the Cas12a nuclease region; e.g., the RuvC site and / or HNH site of the Cas9 nuclease region). CRISPR-Cas effector proteins having mutations in their nuclease active site and therefore no longer possessing nuclease activity are generally referred to as "dead," e.g., dCas. In some embodiments, a CRISPR-Cas effector protein region or polypeptide having mutations in its nuclease active site may exhibit impaired or reduced activity compared to the same CRISPR-Cas effector protein without the mutation, e.g., nikase, e.g., Cas9 nikase, Cas12a nikase.

[0482] In some embodiments, the guide nuclease may include a functional region other than the nuclease, such as an adenine deaminase region, a cytosine deaminase region, or a reverse transcriptase region.

[0483] The adenine deaminase (or adenosine deaminase) useful in the present invention may be any known or later identified adenine deaminase derived from any organism (see, for example, U.S. Patent No. 10,113,163 incorporated herein by reference with respect to the disclosure of adenine deaminase). Adenine deaminase can catalyze the hydrolytic deamination of adenine or adenosine. In one embodiment, adenine deaminase includes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, adenosine deaminase can catalyze the hydrolytic deamination of adenine or adenosine in DNA. In some embodiments, the adenine deaminase encoded by the nucleic acid construct of the present invention may produce an A→G conversion in the sense (e.g., "+", template) strand of the target nucleic acid, or a T→C conversion in the antisense (e.g., "-", complementary) strand of the target nucleic acid.

[0484] In some embodiments, adenosine deaminase may be a variant of naturally occurring adenine deaminase. Therefore, in some embodiments, adenosine deaminase may be about 70% to 100% identical to wild-type adenine deaminase (for example, about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to naturally occurring adenine deaminase, and any range or value within that range). In some embodiments, deaminase or deaminase that does not exist naturally may be referred to as an engineered, mutated, or evolved adenosine deaminase. Therefore, for example, an engineered, mutated, or evolved adenine deaminase polypeptide or adenine deaminase region may be approximately 70% to 99.9% identical to a naturally occurring adenine deaminase polypeptide / region (e.g., approximately 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical, as well as any range or value within that range). In some embodiments, adenosine deaminase may be derived from bacteria (e.g., Escherichia coli, Staphylococcus aureus, Haemophilus influenzae, Caulobacter crescentus, etc.). In some embodiments, the polynucleotide encoding the adenine deaminase polypeptide / region may be codon-optimized for expression in plants.

[0485] In some embodiments, the adenine deaminase region may be a wild-type tRNA-specific adenosine deaminase region, e.g., tRNA-specific adenosine deaminase (TadA) and / or a mutated / evolved adenosine deaminase region, e.g., a mutated / evolved tRNA-specific adenosine deaminase region (TadA*). In some embodiments, the TadA region may originate from E. coli. In some embodiments, TadA may be modified, for example, truncated, and have one or more N-terminal and / or C-terminal amino acids deleted compared to full-length TadA (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal and / or C-terminal amino acid residues may be missing compared to full-length TadA). In some embodiments, the TadA polypeptide or TadA region does not contain an N-terminal methionine. In some embodiments, the polynucleotide encoding TadA / TadA* may be codon-optimized for expression in plants.

[0486] Cytosine deaminases catalyze the deamination of cytosine, resulting in thymidine (via a uracil intermediate), which causes a C-to-T or G-to-A conversion in the complementary strand of the genome. Thus, in some embodiments, the adenine deaminase encoded by the nucleic acid construct of the present invention may cause a C-to-T conversion in the sense (e.g., "+", template) strand of the target nucleic acid, or a G-to-A conversion in the antisense (e.g., "-", complementary) strand of the target nucleic acid.

[0487] In some embodiments, the adenine deaminase encoded by the nucleic acid construct of the present invention may result in an A→G conversion in the sense (e.g., "+", template) strand of the target nucleic acid, or a T→C conversion in the antisense (e.g., "-", complementary) strand of the target nucleic acid.

[0488] The nucleic acid constructs of the present invention encoding a base editor comprising a sequence-specific DNA-binding protein and a cytosine deaminase polypeptide, as well as the nucleic acid constructs / expression cassettes / vectors encoding them, may be used in combination with guide nucleic acids for modifying targets, including but not limited to plasmid sequences, the generation of C→T or G→A mutations in target nucleic acids, the generation of C→T or G→A mutations in coding sequences that alter amino acid identity, the generation of C→T or G→A mutations in coding sequences that generate stop codons, the generation of C→T or G→A mutations in coding sequences that disrupt start codons, the generation of point mutations in genomic DNA that disrupt transcription factor binding, and / or the generation of point mutations in genomic DNA that disrupt splice junctions.

[0489] The nucleic acid constructs of the present invention encoding a base editor comprising a sequence-specific DNA-binding protein and an adenine deaminase polypeptide, as well as expression cassettes and / or vectors encoding them, may be used in combination with guide nucleic acids for modifying targets, including but not limited to plasmid sequences, the generation of A→G or T→C mutations in target nucleic acids, the generation of A→G or T→C mutations in coding sequences that alter amino acid identity, the generation of A→G or T→C mutations in coding sequences that generate stop codons, the generation of A→G or T→C mutations in coding sequences that disrupt start codons, the generation of point mutations in genomic DNA that disrupt function, and / or the generation of point mutations in genomic DNA that disrupt splice junctions.

[0490] target site As used herein, “target sequence” refers to a selected sequence or region of a DNA molecule in which modification (e.g., cleavage, site-specific integration) is desired. The target sequence includes the target site.

[0491] As used herein, “target site” refers to a portion of a target sequence that is cleaved by a guide nuclease, such as a CRISPR nuclease. In contrast to non-target nucleic acids (e.g., non-target ssDNA) or non-target regions, a target site contains significant complementarity with respect to the guide nucleic acid or guide RNA.

[0492] In one embodiment, the target site is 100% complementary to the guide nucleic acid. In another embodiment, the guide nucleic acid is at least 99% complementary to the target site. In another embodiment, the guide nucleic acid is at least 98% complementary to the target site. In another embodiment, the guide nucleic acid is at least 97% complementary to the target site. In another embodiment, the guide nucleic acid is at least 96% complementary to the target site. In another embodiment, the guide nucleic acid is at least 95% complementary to the target site. In another embodiment, the guide nucleic acid is at least 94% complementary to the target site. In another embodiment, the guide nucleic acid is at least 93% complementary to the target site. In another embodiment, the guide nucleic acid is at least 92% complementary to the target site. In another embodiment, the guide nucleic acid is at least 91% complementary to the target site. In another embodiment, the guide nucleic acid is at least 90% complementary to the target site. In another embodiment, the guide nucleic acid is at least 85% complementary to the target site. In another embodiment, the guide nucleic acid is at least 80% complementary to the target site.

[0493] In one embodiment, the target site includes at least one PAM site. In another embodiment, the target site is adjacent to a nucleic acid sequence containing at least one PAM site. In yet another embodiment, the target site is within 5 nucleotides of at least one PAM site. In yet another embodiment, the target site is within 10 nucleotides of at least one PAM site. In yet another embodiment, the target site is within 15 nucleotides of at least one PAM site. In yet another embodiment, the target site is within 20 nucleotides of at least one PAM site. In yet another embodiment, the target site is within 25 nucleotides of at least one PAM site. In yet another embodiment, the target site is within 30 nucleotides of at least one PAM site.

[0494] In one embodiment, the target site is located within gene DNA. In another embodiment, the target site is located within a gene. In another embodiment, the target site is located within the target gene. In another embodiment, the target site is located within an exon of a gene. In another embodiment, the target site is located within an intron of a gene. In another embodiment, the target site is located within the promoter of a gene. In another embodiment, the target site is located within the 5'-UTR of a gene. In another embodiment, the target site is located within the 3'-UTR of a gene. In another embodiment, the target site is located within intergenetic DNA.

[0495] A "protospacer sequence" refers to a target double-stranded DNA sequence, specifically a portion of the target DNA (e.g., a target region in the genome) that is completely or substantially complementary to (and hybridizes with) a CRISPR repeat-spacer sequence (e.g., a guide nucleic acid, CRISPR array, or crRNA).

[0496] In type V CRISPR-Cas (e.g., Cas12a) and type II CRISPR-Cas (Cas9) systems, the protospacer sequence is located adjacent to the protospacer adjacency motif (PAM) (e.g., immediately next to it). In type IV CRISPR-Cas systems, the PAM is located at the 5' end of the non-target strand and the 3' end of the target strand (see below for an example). 5'-NNNNNNNNNNNNNNNNNNNN-3'RNA spacer | | | | | | | | | | | | | | | | | | | | 3'AAANNNNNNNNNNNNNNNNNNNN-5' target strand | | | 5'TTTNNNNNNNNNNNNNNNNNNNNN-3' non-target strand

[0497] In type II CRISPR-Cas systems (e.g., Cas9), the PAM is located 3' immediately after the target region. In type I CRISPR-Cas systems, the PAM is located 5' of the target chain. There is no known PAM for type III CRISPR-Cas systems. Makarova et al. describe the nomenclature for all classes, types, and subtypes of CRISPR systems ((2015) Nature Reviews Microbiology 13:722-736). Guide structures and PAMs are described by R. Barrangou ((2015) Genome Biol. 16:247).

[0498] Canonical Cas12a PAMs are T-rich. In some embodiments, the canonical Cas12a PAM sequence may be 5'-TTN, 5'-TTTN, or 5'-TTTV. In some embodiments, the canonical Cas9 (e.g., S. pyogenes) PAM may be 5'-NGG-3'. In some embodiments, non-canonical PAMs may be used, but this may result in lower efficiency.

[0499] Additional PAM sequences can be determined by those skilled in the art through established experimental and computational approaches. For example, experimental approaches include targeting sequences flanked by all possible nucleotide sequences and identifying sequence members that are not targeted, such as by transformation of the target plasmid DNA (Esvelt et al. (2013) Nat. Methods 10:1116-1121; Jiang et al. (2013) Nat. Biotechnol. 31:233-239). In some embodiments, computational approaches may include performing a BLAST search of natural spacers to identify the original target DNA sequence in the bacteriophage or plasmid, and aligning these sequences to determine conserved sequences adjacent to the target sequence (Briner and Barrangou. (2014) Appl. Environ. Microbiol. 80:994-1001; Mojica et al. (2009) Microbiology 155:733-740).

[0500] In one embodiment, the target DNA molecule is single-stranded. In another embodiment, the target DNA molecule is double-stranded.

[0501] In one embodiment, the target sequence includes genomic DNA. In one embodiment, the target sequence is located in a nucleic acid genome. In one embodiment, the target sequence includes chromatomyne DNA. In one embodiment, the target sequence includes plasmid DNA. In one embodiment, the target sequence is located in a plasmid. In one embodiment, the target sequence includes mitochondrial DNA. In one embodiment, the target sequence is located in a mitochondrial genome. In one embodiment, the target sequence includes plastid DNA. In one embodiment, the target site is located in a plastid genome. In one embodiment, the target sequence includes chloroplast DNA. In one embodiment, the target sequence is located in a genome selected from the group consisting of nuclear genome, mitochondrial genome, and plastid genome.

[0502] In one embodiment, the target sequence includes gene DNA. As used herein, “gene DNA” refers to DNA that codes for one or more genes. In one embodiment, the target sequence includes intergenetic DNA. In contrast to gene DNA, “intergenetic DNA” includes non-coding DNA and lacks DNA that codes for genes. In one embodiment, intergenetic DNA is located between two genes.

[0503] In one embodiment, the target sequence encodes a gene. As used herein, “gene” means a polynucleotide capable of producing a functional unit (e.g., a protein or a non-coding RNA molecule). A gene may include a promoter, an enhancer sequence, a leader sequence, a transcription start site, a transcription termination site, a polyadenylation site, one or more exons, one or more introns, a 5'-UTR, a 3'-UTR, or any combination thereof. A “gene sequence” may include a polynucleotide sequence encoding a promoter, an enhancer sequence, a leader sequence, a transcription start site, a transcription termination site, a polyadenylation site, one or more exons, one or more introns, a 5'-UTR, a 3'-UTR, or any combination thereof. In one embodiment, the gene encodes a non-protein-coding RNA molecule or its precursor. In another embodiment, the gene encodes a protein. In some embodiments, the target sequence is selected from the group consisting of promoters, enhancer sequences, leader sequences, transcription start sites, transcription termination sites, polyadenylation sites, exons, introns, splice sites, 5'-UTR, 3'-UTR, protein-coding sequences, non-protein-coding sequences, miRNAs, pre-miRNAs, and miRNA-binding sites.

[0504] Non-protein-coding RNA molecules include, but are not limited to, microRNAs (miRNAs), miRNA precursors (pre-miRNAs), small interfering RNAs (siRNAs), small RNAs (18-26 nucleotides in length) and their encoding precursors, heterochromatin siRNAs (hc-siRNAs), Piwi-interacting RNAs (piRNAs), hairpin double-stranded RNAs (hairpin dsRNAs), trans-acting siRNAs (ta-siRNAs), naturally occurring antisense siRNAs (nat-siRNAs), CRISPR RNAs (crRNAs), tracer RNAs (tracrRNAs), guide RNAs (gRNAs), and single guide RNAs (sgRNAs). In some embodiments, non-protein-coding RNA molecules include miRNAs. In some embodiments, non-protein-coding RNA molecules include siRNAs. In some embodiments, non-protein-coding RNA molecules include ta-siRNAs. In some embodiments, non-protein-coding RNA molecules are selected from the group consisting of miRNAs, siRNAs, and ta-siRNAs.

[0505] As used herein, “the gene of interest” refers to a polynucleotide sequence encoding a protein or non-protein-coding RNA molecule to be incorporated into a target sequence, or alternatively, an endogenous polynucleotide sequence encoding a protein or non-protein-coding RNA molecule to be edited by each riboprotein. In one embodiment, the gene of interest encodes a protein. In another embodiment, the gene of interest encodes a non-protein-coding RNA molecule. In one embodiment, the gene of interest is exogenous to the targeted DNA molecule. In one embodiment, the gene of interest is a substitute for an endogenous gene in the targeted DNA molecule.

[0506] mutation In one embodiment, the ribonucleoprotein or method provided herein induces at least one mutation in a target sequence.

[0507] In some embodiments, seeds produced from the plants provided herein contain at least one mutation in the gene of interest, including the target site, compared to seeds of a control plant of the same lineage, or seeds of a variant lacking a first nucleic acid sequence encoding a guide nuclease operably linked to a flower cell-preferential promoter, or a second nucleic acid encoding at least one guide nucleic acid operably linked to a second promoter of a different species. In some embodiments, seeds produced from the plants provided herein contain at least one mutation in the gene of interest, including the target site, compared to seeds of a control plant of the same lineage, or seeds of a variant lacking a first nucleic acid sequence encoding a guide nuclease operably linked to a flower tissue-preferential promoter, or a second nucleic acid encoding at least one guide nucleic acid operably linked to a second promoter of a different species.

[0508] In some embodiments, seeds produced from the plants provided herein contain at least one mutation in the gene of interest, including the target site, compared to seeds from a control plant of the same lineage, or seeds from a variant lacking a first nucleic acid sequence encoding a guide nuclease operably linked to a heterologous promoter, or a second nucleic acid encoding at least one guide nucleic acid operably linked to a floral cell-preferential promoter. In some embodiments, seeds produced from the plants provided herein contain at least one mutation in the gene of interest, including the target site, compared to seeds from a control plant of the same lineage, or seeds from a variant lacking a first nucleic acid sequence encoding a guide nuclease operably linked to a heterologous promoter, or a second nucleic acid encoding at least one guide nucleic acid operably linked to a floral tissue-preferential promoter.

[0509] As used herein, “mutation” refers to a change in a nucleic acid or amino acid sequence that does not occur naturally, compared to a naturally occurring reference nucleic acid or amino acid sequence derived from the same organism. When identifying mutations, it will be understood that the reference sequence should be derived from the same nucleic acid (e.g., gene, non-coding RNA) or amino acid (e.g., protein). When determining whether a difference between two sequences constitutes a mutation, it will be understood in the art that such comparison should not be made between homologous sequences of two different species or between homologous sequences of two different varieties of the same species. Rather, the comparison should be made between an edited (e.g., mutated) sequence and an endogenous, unedited (e.g., “wild-type”) sequence of the same organism.

[0510] Several types of mutations are known in the art. In one embodiment, a mutation includes an insertion. “Insertion” means the addition of one or more nucleotides or amino acids to a given polynucleotide or amino acid sequence, compared to an endogenous reference polynucleotide or amino acid sequence. In another embodiment, a mutation includes a deletion. “Deletion” means the removal of one or more nucleotides or amino acids from a given polynucleotide or amino acid sequence, compared to an endogenous reference polynucleotide or amino acid sequence. In another embodiment, a mutation includes a substitution. “Substitution” means the exchange of one or more nucleotides or amino acids from a given polynucleotide or amino acid sequence, compared to an endogenous reference polynucleotide or amino acid sequence. In another embodiment, a mutation includes an inversion. “Inversion” means that the ends of a segment of a polynucleotide or amino acid sequence are reversed. In one embodiment, the mutations provided herein include mutations selected from the group consisting of insertions, deletions, substitutions, and inversions.

[0511] In one embodiment, a plant or seed contains at least one mutation in the gene of interest, the at least one mutation resulting in the deletion of one or more amino acids in the protein encoded by the gene of interest compared to the wild-type protein.

[0512] In one embodiment, a plant or seed contains at least one mutation in the gene of interest, the at least one mutation resulting in the substitution of one or more amino acids in the protein encoded by the gene of interest, compared to the wild-type protein.

[0513] In one embodiment, a plant or seed contains at least one mutation in the gene of interest, the at least one mutation resulting in the insertion of one or more amino acids in the protein encoded by the gene of interest, compared to the wild-type protein.

[0514] Mutations in the coding region of a gene (e.g., exon mutations) can result in truncated proteins or polypeptides when the mutated messenger RNA (mRNA) is translated into a protein or polypeptide. In some embodiments, this disclosure provides mutations that result in truncation of a protein or polypeptide. As used herein, a “truncated” protein or polypeptide contains at least one fewer amino acid than an endogenous control protein or polypeptide. For example, if endogenous protein A contains 100 amino acids, a truncated form of protein A may contain 1 to 99 amino acids.

[0515] Without limiting to any scientific theory, one method for causing truncation of a protein or polypeptide is by introducing an immature stop codon in the mRNA transcript of an endogenous gene. In some embodiments, this disclosure provides mutations that result in an immature stop codon in the mRNA transcript of an endogenous gene. As used herein, “stop codon” refers to a nucleotide triplet in an mRNA transcript that signals the end of protein translation. “Immature stop codon” refers to a stop codon located earlier (e.g., 5') than the normal stop codon position in the endogenous mRNA transcript. Several stop codons are known in the art, including, but are not limited to, “UAG”, “UAA”, “UGA”, “TAG”, “TAA”, and “TGA”.

[0516] In one embodiment, the seed or plant contains at least one mutation, the at least one mutation resulting in the introduction of an immature stop codon in the messenger RNA encoded by the gene of interest, compared to the wild-type messenger RNA.

[0517] In some embodiments, the mutations provided herein include null mutations. As used herein, “null mutation” means a mutation that confers complete loss of function to a protein encoded by the gene containing the mutation, or alternatively, a mutation that confers complete loss of function to a small RNA molecule encoded by a genomic locus. A null mutation may result in the absence of mRNA transcript production, the absence of small RNA transcript production, the absence of protein function, or a combination thereof.

[0518] The mutations provided herein may be located in any part of an endogenous gene. In one embodiment, the mutations provided herein are located within an exon of an endogenous gene. In another embodiment, the mutations provided herein are located within an intron of an endogenous gene. In yet another embodiment, the mutations provided herein are located within the 5'-untranslated region of an endogenous gene. In yet another embodiment, the mutations provided herein are located within the 3'-untranslated region of an endogenous gene. In yet another embodiment, the mutations provided herein are located within the promoter of an endogenous gene.

[0519] In one embodiment, mutations are located at splice sites within a gene. Mutations at splice sites can interfere with exon splicing during mRNA processing. Splicing can be perturbed if one or more nucleotides are inserted, deleted, or substituted at a splice site. Perturbed splicing can result in unspliced ​​introns, exon deletions, or both from a mature mRNA sequence. Typically, but not always, a "GU" sequence is required at the 5' end of an intron and an "AG" sequence is required at the 3' end of an intron for proper splicing. If either of these splice sites is mutated, splicing perturbation can occur.

[0520] In one embodiment, a seed or plant contains at least one mutation, the at least one mutation comprising the deletion of one or more splice sites from the gene of interest. In another embodiment, a seed or plant contains at least one mutation, the at least one mutation located within one or more splice sites from the gene of interest.

[0521] In one embodiment, the mutation includes site-directed integration. In one embodiment, site-directed integration includes the insertion of all or part of the desired sequence into a target sequence.

[0522] As used herein, “site-directed integration” means all or part of a desired sequence (e.g., endogenous gene, edited endogenous gene) to be inserted or integrated into a desired site or locus (e.g., target sequence) in the plant genome. As used herein, “desired sequence” means a DNA molecule containing a nucleic acid sequence to be integrated into the genome of a plant or plant cell. The desired sequence may include a transgene or construct. In some embodiments, the nucleic acid molecule containing the desired sequence includes one or two homology arms adjacent to the desired sequence to facilitate a targeted insertion event via homologous recombination and / or homologous recombination repair.

[0523] In some embodiments, the methods described herein include site-specific incorporation of a desired sequence into a target sequence.

[0524] Any site or locus in the plant genome can be selected for site-specific integration of the transgene or construct of the present disclosure. In one embodiment, the target sequence is located within the chromosome, or supermenulari.

[0525] In site-directed insertion, a double-strand break (DSB) or nick may first occur at the target sequence via a guide nuclease or ribonucleoprotein provided herein. In the presence of a preferred sequence, the DSB or nick may be repaired by homologous recombination (HR) between homology arms (or multiples) of the desired sequence and the target sequence, or by non-homologous end joining (NHEJ), which may generate a targeted insertion event at the site of the DSB or nick by incorporating all or part of the desired sequence into the target sequence.

[0526] In one embodiment, site-specific integration involves the use of an endogenous NHEJ repair mechanism for the cell. In another embodiment, site-specific integration involves the use of an endogenous HR repair mechanism for the cell.

[0527] In one embodiment, double-strand break repair generates at least one mutation in the gene of interest compared to a control plant of the same lineage or variety.

[0528] In one embodiment, the mutation includes the insertion of at least five consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of at least ten consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of at least fifteen consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of at least twenty consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of at least twenty-five consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of at least fifty consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of at least one hundred consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of at least two hundred consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of at least five hundred consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of at least one thousand consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation involves the insertion of at least 2000 consecutive nucleotides of a desired sequence into a target sequence.

[0529] In one embodiment, the mutation includes the insertion of 5 to 3500 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the incorporation of 5 to 2500 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 5 to 1500 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 5 to 750 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 5 to 500 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 5 to 250 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 5 to 150 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 25 to 250 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 25 to 1500 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 25 to 750 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 50 to 2500 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 50 to 1500 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 50 to 750 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 100 to 2500 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation includes the insertion of 100 to 1500 consecutive nucleotides of the desired sequence into the target sequence. In one embodiment, the mutation involves the insertion of 100 to 750 consecutive nucleotides of a desired sequence into a target sequence.

[0530] In some embodiments, the methods provided herein further include detecting editing or mutations in a target sequence. Screening and selection of mutant or edited plants or plant cells can be carried out by any methodology known to those skilled in the art. Examples of screening and selection methodologies include, but are not limited to, Southern analysis, PCR amplification for polynucleotide detection, Northern blotting, RNase protection, primer extension, RT-PCR amplification for RNA transcript detection, Sanger sequencing, next-generation sequencing technologies (e.g., Illumina, PacBio, Ion Torrent, 454), enzyme assays for detecting the enzymatic or ribozyme activity of polypeptides and polynucleotides, as well as protein gel electrophoresis, Western blotting, immunoprecipitation, and enzyme-conjugated immunoassays for detecting polypeptides. Other techniques such as in situ hybridization, enzyme staining, and immunostaining can also be used to detect the presence or expression of polypeptides and / or polynucleotides. Methods for carrying out all of the techniques referenced above are known in the art.

[0531] In one embodiment, the sequences provided herein encode at least one ribozyme. In another embodiment, the sequences provided herein encode at least two ribozymes. In another embodiment, the ribozyme is a self-cleaving ribozyme. Self-cleaving ribozymes are known in the art. See, for example, Jimenez et al., Trends Biochem. Sci., 40:648-661 (2015).

[0532] In some embodiments, the sequence encoding at least one guide nucleic acid is adjacent to the self-cleaving ribozyme. In some embodiments, the sequence encoding at least one guide nucleic acid is immediately adjacent to the sequence encoding the ribozyme (for example, the outermost 5' nucleotide of the guide nucleic acid is abutting the outermost 3' nucleotide of the ribozyme, or the outermost 3' nucleotide of the guide nucleic acid is abutting the outermost 5' nucleotide of the ribozyme). In some embodiments, the sequence encoding at least one guide nucleic acid is at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 250, at least 500, or at least 10,000 nucleotides away from the sequence encoding the ribozyme.

[0533] plant Any plant or plant cell can be used in the methods and compositions provided herein. In some embodiments, the plant is selected from the group consisting of maize plants, rice plants, sorghum plants, wheat plants, alfalfa plants, barley plants, millet plants, rye plants, sugarcane plants, cotton plants, soybean plants, canola plants, tomato plants, onion plants, cucumber plants, Arabidopsis plants, or potato plants. In some embodiments, the plant is an angiosperm. In some embodiments, the plant is a gymnosperm. In some embodiments, the plant is a monocotyledonous plant. In some embodiments, the plant is a dicotyledonous plant. In one embodiment, the plant is a plant belonging to a family selected from the group consisting of the Alliaceae, Anacardiaceae, Apiaceae, Arecaceae, Asteraceae, Brassicaceae, Caesalpinaceae, Cucurbitaceae, Ericaceae, Fabaceae, Juglandaceae, Malvaceae, Mimosoraceae, Moraceae, Musaceae, Orchidaceae, Papilionaceae, Pinaceae, Poaceae, Rosaceae, Rutaceae, Rubiaceae, and Solanaceae.

[0534] In some embodiments, plant cells are selected from the group consisting of maize cells, rice cells, sorghum cells, wheat cells, alfalfa cells, barley cells, millet cells, rye cells, sugarcane cells, cotton cells, soybean cells, canola cells, tomato cells, onion cells, cucumber cells, Arabidopsis cells, and potato cells. In some embodiments, plant cells are angiosperm cells. In some embodiments, plant cells are gymnosperm cells. In some embodiments, plant cells are monocotyledonous plant cells. In some embodiments, plant cells are dicotyledonous plant cells. In one embodiment, the plant cells are plant cells from a family selected from the group consisting of the Alliaceae, Anacardiaceae, Apiaceae, Arecaceae, Asteraceae, Brassicaceae, Caesalpinaceae, Cucurbitaceae, Ericaceae, Fabaceae, Juglandaceae, Malvaceae, Mimosoraceae, Moraceae, Musaceae, Orchidaceae, Papilionaceae, Pinaceae, Poaceae, Rosaceae, Rutaceae, Rubiaceae, and Solanaceae.

[0535] As used herein, “variety” refers to a group of plants within a species (not limited to, for example, zea mays) that share a specific genetic trait that distinguishes them from other possible varieties within that species. Varieties may be inbred or hybrid, but commercially available plants are often hybrids to take advantage of hybrid vigor. Individuals within a hybrid cultivar are uniform, genetically nearly identical, and most loci are heterozygous.

[0536] As used herein, the term “inbred” means a lineage bred for genetic uniformity. In one embodiment, the seeds provided herein are inbred seeds. In another embodiment, the plant cells provided herein are inbred plants.

[0537] As used herein, the term “hybrid” means offspring resulting from a cross between at least two genetically dissimilar parents. Examples of cross schemes, though not limited to them, include single crosses, modified single crosses, double modified single crosses, three-way crosses, modified three-way crosses, and double crosses, where at least one parent in a modified cross is an offspring of a sister line cross. In some embodiments, the seeds provided herein are hybrid seeds. In some embodiments, the plants provided herein are hybrid plants.

[0538] In some jurisdictions, products obtained solely through essentially biological processes, such as plant products, are excluded from patent protection. Therefore, claimed plants, plant parts and cells, and their offspring, can be defined as being directed only to those plants, plant parts and cells, and their offspring obtained through technological intervention (without further propagation via mating and selection). One embodiment of the present invention relates to plant parts or offspring produced or available using the gene editing techniques described herein. Alternatively, subject matter excluded from patentability can be excluded. Embodiments of the present invention relate to plants, plant parts, or offspring containing genomic alterations described elsewhere herein. However, plants, plant parts, or plants or offspring are not obtained solely through essentially biological processes, where essentially biological processes are the processes of production of plants or animals consisting solely of entirely natural phenomena such as mating or selection.

[0539] Transformation The method may include transient transformation or stable incorporation of any nucleic acid molecule into any plant or plant cell provided herein.

[0540] As used herein, “stable integration” or “stably integrated” refers to the introduction of DNA into the genomic DNA of a target cell or plant, thereby enabling the target cell or plant to pass on the introduced DNA to the next generation of the transformed organism. Stable transformation requires that the introduced DNA be integrated into the germ cell(s) of the transformed organism. As used herein, “transiently transformed” or “transient transformation” refers to the introduction of DNA into a cell that is not transmitted to the next generation of the transformed organism. In transient transformation, the transformed DNA is not typically integrated into the genomic DNA of the transformed cell. In one embodiment, the method stably transforms plant cells or plants with one or more nucleic acid molecules provided herein. In another embodiment, the method transiently transforms plant cells or plants with one or more nucleic acid molecules provided herein.

[0541] In one embodiment, a nucleic acid molecule encoding a guide nuclease is stably integrated into the plant genome. In one embodiment, a nucleic acid molecule encoding a Cas12a nuclease is stably integrated into the plant genome. In one embodiment, a nucleic acid molecule encoding a CasX nuclease is stably integrated into the plant genome. In one embodiment, a nucleic acid molecule encoding a guide nucleic acid is stably integrated into the plant genome. In one embodiment, a nucleic acid molecule encoding a guide RNA is stably integrated into the plant genome. In one embodiment, a nucleic acid molecule encoding a single guide RNA is stably integrated into the plant genome.

[0542] Numerous methods for transforming cells with recombinant nucleic acid molecules or constructs are known in the art and can be used in accordance with the methods of this application. Any suitable method or technique for cell transformation known in the art can be used in accordance with the methods of the present invention. Effective methods for plant transformation include bacterial-mediated transformations such as Agrobacterium-mediated or rhizobia-mediated transformations, and microparticle gun-mediated transformations. Various methods are known in the art for transforming explants with transformation vectors via bacterial-mediated transformation or microprojectile bombardment, and then regenerating or developing transgenic plants by culturing those explants, etc.

[0543] In some embodiments, the method includes providing nucleic acid molecules to cells via Agrobacterium-mediated transformation. In some embodiments, the method includes providing nucleic acid molecules to cells via polyethylene glycol-mediated transformation. In some embodiments, the method includes providing nucleic acid molecules to cells via bioristic transformation. In some embodiments, the method includes providing nucleic acid molecules to cells via liposome-mediated transfection. In some embodiments, the method includes providing nucleic acid molecules to cells via viral transduction. In some embodiments, the method includes providing nucleic acid molecules to cells via the use of one or more delivery particles. In some embodiments, the method includes providing nucleic acid molecules to cells via microinjection. In some embodiments, the method includes providing nucleic acid molecules to cells via electroporation.

[0544] In one embodiment, nucleic acid molecules are delivered to cells by a method selected from the group consisting of Agrobacterium-mediated transformation, polyethylene glycol-mediated transformation, bioristic transformation, liposome-mediated transfection, viral transduction, use of one or more delivery particles, microinjection, and electroporation.

[0545] Other conversion methods, such as vacuum impregnation, pressure, sonication, and stirring of silicon carbide fibers, are also known in the art and are intended for use in any manner provided herein.

[0546] Methods for transforming cells are well known to those skilled in the art. For example, specific descriptions of use for transforming plant cells by microprojectile bombardment using recombinant DNA-coated particles (e.g., bioristic transformation) can be found in U.S. Patents 5,550,318, 5,538,880, 6,160,208, 6,399,861, and 6,153,812, and for Agrobacterium-mediated transformation can be found in U.S. Patents 5,159,135, 5,824,877, 5,591,616, 6,384,301, 5,750,871, 5,463,174, and 5,188,958, all of which are incorporated herein by reference. Additional methods for transforming plants can be found, for example, in Compendium of Transgenic Crop Plants (2009), Blackwell Publishing. Any suitable method known to those skilled in the art can be used to transform plant cells using any of the nucleic acid molecules provided herein.

[0547] Lipofection is described, for example, in U.S. Patents 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam® and Lipofectin®). Suitable cationic and neutral lipids for efficient receptor recognition lipofection of polynucleotides include those described in Felgner's WO91 / 17424 and WO91 / 16024. Delivery may be to cells (in vivo or ex vivo administration) or target tissue (in vivo administration).

[0548] Delivery vehicles, vectors, particles, nanoparticles, formulations, and their components for expressing one or more elements of nucleic acid molecules are as used in WO2014 / 093622. In some embodiments, a method for delivering nucleic acid molecules or proteins to cells includes delivery via delivery particles. In some embodiments, a method for delivering nucleic acid molecules or proteins to cells includes delivery via delivery vesicles. In some embodiments, the delivery vesicles are selected from the group consisting of exosomes and liposomes. In some embodiments, a method for delivering nucleic acid molecules to cells includes delivery via viral vectors. In some embodiments, the viral vectors are selected from the group consisting of adenovirus vectors, lentivirus vectors, and adeno-associated virus vectors. In other embodiments, a method for delivering nucleic acid molecules to plant cells or plants includes delivery via nanoparticles. In some embodiments, a method for delivering nucleic acid molecules to plant cells or plants includes microinjection. In some embodiments, a method for delivering nucleic acid molecules to plant cells or plants includes polycations. In some embodiments, a method for delivering nucleic acid molecules to plant cells or plants includes cationic oligopeptides.

[0549] In one embodiment, the delivery particle is selected from the group consisting of exosomes, adenovirus vectors, lentiviral vectors, and adeno-associated virus vectors, nanoparticles, polycations, and cationic oligopeptides. In one embodiment, the method provided herein includes the use of one or more delivery particles. In one embodiment, the method provided herein includes the use of two or more delivery particles. In one embodiment, the method provided herein includes the use of three or more delivery particles.

[0550] Suitable agents for promoting the transfer of nucleic acids into plant cells include those that increase the permeability of the plant to the outside or to plant cells to oligonucleotides or polynucleotides. Such agents for promoting the movement of compositions into plant cells include chemical agents, physical agents, or combinations thereof. Chemicals for conditioning include (a) surfactants, (b) organic solvents, aqueous solutions, or aqueous mixtures of organic solvents, (c) oxidizing agents, (e) acids, (f) bases, (g) oils, (h) enzymes, or combinations thereof.

[0551] Useful organic solvents for conditioning plants to allow polynucleotides to permeate include DMSO, DMF, pyridine, N-pyrrolidine, hexamethylphosphoroamide, acetonitrile, dioxane, polypropylene glycol, and other solvents that are miscible with water or dissolve phosphonucleotides in non-aqueous systems (e.g., used in synthetic reactions). Natural or synthetic oils, such as plant-derived oils and crop oils (e.g., those listed in the 9th Compendium of Harbide Adjuvants, available online at www.herbicide.com), may be used, with or without surfactants or emulsifiers. For example, paraffin oils, polyol fatty acid esters, or oils containing short-chain molecules modified with amides or polyamines such as polyethyleneimine or N-pyrrolidine may be used.

[0552] Examples of useful surfactants include sodium or lithium salts of fatty acids (such as tallow or tallowamine or phospholipids) and organosilicon surfactants. Other useful surfactants include nonionic organosilicon surfactants, such as organosilicon surfactants containing trisiloxane ethoxylate surfactants, or silicone polyether copolymers such as the copolymer of polyalkylene oxide-modified heptamethyltrisiloxane and allyloxypolypropylene glycol methyl ether (commercially available as Silwet® L-77).

[0553] Useful physical agents may include (a) abrasives such as carborundum, corundum, sand, calcite, pumice, and garnet; (b) nanoparticles such as carbon nanotubes; or (c) physical force. Carbon nanotubes are disclosed in Kam et al. (2004) Am. Chem. Soc, 126(22):6850-6851, Liu et al. (2009) Nano Lett, 9(3):1007-1010, and Khodakovkaya et al. (2009) ACS Nano, 3(10):3221-3227. The application of physical force may include heating, cooling, application of positive pressure, or sonication. Embodiments of this method may optionally include an incubation step, a neutralization step (e.g., neutralizing an acid, base, or oxidizing agent, or inactivating an enzyme), a rinsing step, or a combination thereof. The method of the present invention may further include the application of other agents whose effectiveness is enhanced by silencing certain genes. For example, if a polynucleotide is designed to modulate a gene that confers herbicide resistance, subsequent application of the herbicide may have a dramatic effect on the herbicide's potency.

[0554] Chemicals used to condition plant cells in the laboratory to allow polynucleotides to permeate include, for example, the application of chemicals, enzymatic treatment, heating or cooling, treatment under positive or negative pressure, or sonication. Chemicals used to condition plants in the field include surfactants and salts.

[0555] In one embodiment, the transformed or transfected cells are plant cells. Recipient plant cells or explant targets for transformation include, but are not limited to, seed cells, fruit cells, leaf cells, cotyledon cells, hypocotyl cells, meristem cells, embryo cells, endosperm cells, root cells, shoot cells, stem cells, pod cells, flower cells, inflorescence cells, stalk cells, pedicel cells, style cells, stigma cells, receptacle cells, petal cells, sepal cells, pollen cells, anther cells, filament cells, ovary cells, ovule cells, pericarp cells, phloem cells, budding cells, or vascular tissue cells. In another embodiment, the disclosure provides plant chloroplasts. In a further embodiment, the disclosure provides epidermal cells, guard cells, trichome cells, root hair cells, storage root cells, or tuber cells. In another embodiment, the disclosure provides protoplasts. In another embodiment, the disclosure provides plant callus cells. Any cell from which a fertile plant can be regenerated is intended as a recipient cell useful for the implementation of this disclosure. Callus can be initiated from a variety of tissue sources, including but not limited to immature embryos or embryonic portions, seedling apical meristems, and microspores. Cells capable of growing as callus can function as recipient cells for transformation. Practical transformation methods and materials for producing the transgenic plants of this disclosure (e.g., various media and receptor target cells, transformation of immature embryos, and subsequent regeneration of fertile transgenic plants) are disclosed, for example, in U.S. Patent Nos. 6,194,636 and 6,232,526, and U.S. Patent Application Publication No. 2004 / 0216189, all of which are incorporated herein by reference. Transformed explants, cells, or tissues may be subjected to additional culture steps, such as callus induction, selection, and regeneration, as known in the art. Transformed cells, tissues, or explants containing recombinant DNA insertions can be grown, developed, or regenerated into transgenic plants in culture, plugs, or soil according to methods known in the art. In one embodiment, the disclosure provides plant cells that are not reproductive material and do not mediate the natural reproduction of plants. In another embodiment, the disclosure also provides plant cells that are reproductive material and mediate the natural reproduction of plants.In another aspect, the disclosure provides plant cells that are unable to sustain themselves through photosynthesis. In another aspect, the disclosure provides plant somatic cells. Unlike germ cells, somatic cells do not mediate plant reproduction. In one aspect, the disclosure provides non-reproductive plant cells.

[0556] All methods described herein may be performed in any preferred order, unless otherwise shown herein or unless otherwise clearly contradictory by context. The use of any and all examples, or any exemplary language provided herein with respect to a particular embodiment (e.g., “etc.”), is merely intended to clarify the disclosure and does not impose any limitation on the scope of the disclosure unless otherwise stated in the claims. Nothing in this specification should be construed as indicating any unclaimed element that is essential for the practice of the disclosure.

[0557] The grouping of alternative elements or embodiments of the disclosure disclosed herein should not be construed as limiting. Each member of each group may be referenced and claimed individually or in combination with other members of the group or other elements described herein. One or more members of a group may be included in or excluded from a group for convenience or patentability reasons.

[0558] While this disclosure has been described in detail, it will be apparent that modifications, variations, and equivalent embodiments are possible without departing from the scope of this disclosure as defined in the attached claims. Furthermore, it should be understood that all examples in this disclosure are provided as non-limiting examples. [Examples]

[0559] The following embodiments are included to demonstrate embodiments of the present disclosure. Those skilled in the art will understand that many modifications may be made to the specific embodiments disclosed without departing from the concepts, spirit, and scope of the present disclosure, and similar or comparable results may be obtained. More specifically, it will be apparent that certain chemically and physiologically relevant active ingredients may be substituted with the active ingredients described herein, insofar as identical or comparable results are achieved. All such similar substitutions and modifications that are obvious to those skilled in the art are deemed to be within the spirit, scope, and concepts of the present disclosure as defined by the appended claims.

[0560] Example 1. Editing of 5' and 3' UTRs for heteroallelic regulation. Many agriculturally relevant traits are often controlled by genes whose mutations can lead to severe agricultural off-type / phenotypes. This example describes the development of an RNAi-based editing system for strong dominance suppression of heterozygous (hybrid) loci without producing a homozygous phenotype linked to the edited locus. A key aspect of this methodology is to reverse part or all of the untranslated region (UTR) of a candidate gene / locus (and the resulting mRNA) without perturbing the coding region.

[0561] Figure 1 provides a conceptual diagram of this approach to the 3'UTR. Panel A shows a typical unedited structure of the gene, including the 5' untranslated region (5'UTR), coding sequence (CDS), and 3'UTR oriented in the normal direction. To generate an inversion within the 3'UTR of an endogenous gene, two functional guide RNAs (gRNAs) are generated for a CRISPR / RNA guided nuclease system targeting the gene's 3'UTR. See Panel B in Figure 1. Each of the two guide RNAs is selected to be specific to the target site in their respective UTR region and to not contain any known off-target sequences in the target genome (in this case, the off-target sequences are three or fewer mismatches to the 23-mer sequence of the gRNA). Co-delivery of the guide RNAs together with the same nuclease system is expected to result in a CRISPR-mediated double-strand break (DSB) in or near the 3'UTR sequence. Following a double-strand break (DSB) at each of the two target sites, in most cases the intervening region between those target sites is deleted or excised, and the non-homologous end-joining (NHEJ) repair mechanism binds to the adjacent region. Less frequently, in some events, the excised region is reinserted into the genome in the reverse, antisense, or opposite direction by rejoining with the two excised ends of the adjacent genomic DNA. The excised sequence may be completely rejoined with the adjacent genomic DNA in the opposite direction without any sequence changes (except inversion), but one or more nucleotide insertions or deletions (indels) may also occur at one or both of the junctions where the excised fragment and the adjacent genomic DNA rejoin compared to the sequences of the two target sites or their vicinity before excision.

[0562] When present in both homozygous alleles, 3'UTR inversion or inversion within the 3'UTR is not expected to affect protein function, and therefore no phenotype is expected at the homozygous stage, as shown in Figure 1, Panel B. However, when such homozygous edited plants are crossed with wild-type plants, the resulting hybrid offspring inherit two alleles with different 3'UTR orientations (see Figure 1, Panel C). The edited allele results in the production of an RNA transcript containing a sequence complementary to a portion of the natural UTR transcript sequence derived from the wild-type allele. While not bound by any scientific theory, the inverted region of the mRNA from the edited allele and the corresponding region in the mRNA from the wild-type allele are capable of antisense base pairing. This, in turn, can form double-stranded RNA (dsRNA) between the two mRNA populations, potentially causing RNA-mediated repression or silencing of both copies of the gene, producing a knockdown phenotype in the hybrid plant.

[0563] Example 2: Reverse editing of the 5'UTR and 3'UTR of the GA3 oxidase_1 gene. To induce an inversion in either the 5' or 3' UTR region of the endogenous Zm.GA3ox_1 gene, a segment or portion of the 5' or 3' UTR region of the Zm.GA3ox_1 gene can be excised by genome editing using a CRISPR / Cas system (Cpf1 or Cas12a) and reinserted in either the reverse or inverted direction to generate an antisense sequence in the 5' or 3' UTR region of the Zm.GA3ox_1 gene. The inverted antisense sequence provided herein is complementary to the sense sequence of the corresponding segment or portion of the wild-type allele of the Zm.GA3ox_1 gene when an edited Zm.GA3ox_1 allele with UTR inversion is present in a maize plant. This may induce RNA-mediated repression or silencing of both copies of the Zm.GA3ox_1 gene. However, if a maize plant or maize plant cell is homozygous for the edited Zm.GA3ox_1 allele using UTR inversion, the maize plant or maize plant cell does not contain a wild-type copy of the Zm.GA3ox_1 gene, and therefore does not have a complementary UTR sequence, thus preventing RNA-mediated repression or silencing of the Zm.GA3ox_1 gene.

[0564] To generate an inversion in either the 5' or 3' UTR region of the endogenous GA3 oxidase 1 gene in maize, a pair of two 23-mer guide RNAs (gRNAs or spacers) are selected to target their respective UTR regions, generating the inversion by cleaving the intervening DNA sequence and reinserting the cleaved DNA fragments into the same UTR in the reverse, antisense, or opposite direction. In each case, the guide RNAs are selected to be specific to the target site in each UTR region and to not contain any known off-target sequences in the target genome (in this case, off-target sequences are defined as not containing three or fewer mismatches to the 23-mer sequence of the gRNA). The intervening sequence can be excised and reinserted in the reverse, antisense, or reverse direction by performing a double-strand break (DSB) at each of the two target sites within the UTR region, and then reinserted in the reverse, antisense, or reverse direction by rejoining it with the two excised ends of the adjacent genomic DNA. The excised sequence may be able to completely rejoin with the adjacent genomic DNA in the opposite direction without any sequence changes (except inversion), but one or more nucleotide insertions or deletions (indels) may also occur at one or both of the two junctions where the excised fragment and the adjacent genomic DNA rejoin to the two target sites or sequences in their vicinity before excision.

[0565] In this experiment, we identified and selected two pairs of gRNA spacers (SP1 and SP2, see Table 3) to target the excision and inversion of an approximately 450 bp long fragment of the 3' UTR of the Zm.GA3 oxidase_1 gene, and two pairs of gRNA spacers (SP3 and SP4, see Table 3) to target the excision and inversion of an approximately 280 bp long fragment of the 5' UTR of the GA3 oxidase_1 gene. Four plant transformation constructs were created, and genome editing using a CRISPR / Cas12a (Cpf1) system with two guide RNAs adjacent to the UTR sequences intended to be excised and inverted between the two target sites was performed. Two of the four constructs were designed to create 5' UTR inversion edits, and two were designed to create 3' UTR inversion edits. In this example, the vector construct typically has two functional cassettes encoding a gene editing mechanism for creating a targeted mutation in the GA3 oxidase_1 gene: (i) a first cassette for the expression of a Cpf1 (or Cas12a) variant protein, and (ii) a second cassette for the expression of two related guide RNAs targeting the 5'UTR or 3'UTR according to the editing scheme. Each guide RNA contains a common scaffold sequence that fits with the Cpf1 variant, and a unique spacer / targeting sequence complementary to its intended target site in the 5'UTR or 3'UTR of the GA3 oxidase_1 gene.

[0566] In the case of construct pM471, the Cpf1 expression cassette contains a maize germ tissue-preferential promoter (SEQ ID NO: 193) operably ligated to a maize codon-optimized sequence encoding the Lachnospiraceae bacterial Cpf1 RNA-guided endonuclease enzyme (CR-LACba.Cpf1zmCG2:3, SEQ ID NO: 194), which is fused to nuclear localization signals at both the N-terminus and C-terminus of the Cpf1 enzyme. The gRNA expression cassette contains a synthetic promoter (P-Syn.GSP2262_Pol3:1, SEQ ID NO: 196) operably ligated to transcription sequences encoding Cpf1-compatible common scaffold SC1 (GR-LACba.Cpf1:2, SEQ ID NO: 197), spacer SP1, scaffold SC1, spacer SP2, and scaffold SC1 (in order) within the gRNA. The SC1-SP1-SC1-SP2-SC1 portion of the transcript is a pre-crRNA precursor RNA that can be processed into two mature SP1 and SP2 guide RNAs.

[0567] Construct pM313 is similar to the pM471 construct described above, except that the Cpf1 expression cassette contains a constitutive maize ubiquitin promoter (SEQ ID NO: 198) instead of a germline-preferential promoter.

[0568] In the case of construct pM323, the Cpf1 expression cassette includes a maize germ tissue-preferential promoter (SEQ ID NO: 193) operably ligated to a soybean codon-optimized sequence encoding the Francisella tularensis subsp. novicida Cpf1 RNA-induced endonuclease enzyme (CR-Fn.Cpf1-Gm:5, SEQ ID NO: 199), which fuses to a nuclear localization signal (SEQ ID NO: 200) at the N-terminus of the Cpf1 enzyme and to another nuclear localization signal (SEQ ID NO: 201) at the C-terminus. The gRNA expression cassette includes a Cpf1-compatible common scaffold SC2 (GR-Fn.Cpf1 repeat 1, SEQ ID NO: 202), spacer SP3, scaffold SC2, spacer SP4, and a synthetic promoter (P-Syn.GSP2262_Pol3:1, SEQ ID NO: 196) operably ligated to scaffold SC2. The SC2-SP3-SC2-SP4-SC2 portion of the transcript is a pre-crRNA precursor RNA that can be processed into two mature SP3 and SP4 guide RNAs.

[0569] Construct pM314 is similar to the pM323 construct described above, except that the Cpf1 expression cassette contains a constitutive maize ubiquitin promoter (SEQ ID NO: 198) instead of a germline-preferential promoter. Note that the constitutive maize ubiquitin promoter in different constructs can be combined with different leader and intron sequences. [Table 3]

[0570] Example 3: Creation of 5'UTR reverse editing of the Arabidopsis AG gene The design of this gene editing approach is based on the excision and reverse insertion of a fragment of the 5'UTR region derived from the Arabidopsis thaliana Agamous (AG) gene (AT4G18960).

[0571] The Arabidopsis thaliana flower system homeotic gene AGAMOUS (AG) encodes a MADS box transcription factor and plays a central role in the development of reproductive organs (stamens and carpels). AG is thought to regulate developmental pathways by modulating downstream target genes responsible for the uniqueness of stamens and carpels. As a result, ag mutants produce distinctive and scoreable floral phenotypes. Specifically, downregulation of the AG gene results in plants with multiple petal layers (see Yanofsky et al., Nature 346, 35-39 (1990)). Knockout of the AG gene or point mutations in its coding region result in double-flowered phenotype plants with underdeveloped anthers and stigmas due to unregulated petal development in the inner whorled arrangement of the flower. As a result, the strong ag mutation allele produces flowers in which the third whorled stamen is transformed into a petal, while in other flowers that would repeat the same organ pattern, it develops in place of the fourth whorled carpel.

[0572] The genome sequence of the At.AG (AT4G18960, Chr4:10382855..10388539) gene is provided as Sequence ID No. 204. The elements and coordinates in the genome sequence are listed in Table 4. [Table 4]

[0573] The 5'UTR region was scanned, and five FnCas12a target sites were identified. Each contained a TTN PAM site adjacent to a unique spacer sequence that did not contain any known off-target sequences present in the Arabidopsis genome (in this case, an off-target sequence is defined as one that does not contain three or fewer mismatches to the 23-mer sequence of the gRNA). The spacer sequences are listed in Table 5. [Table 5]

[0574] In this experiment, a pair of gRNA spacers (SP18 and SP21, see Table 6) was identified and selected to target the excision and inversion of a fragment approximately 180 bp long from the 5' UTR of the At.AG gene. A second pair of gRNA spacers (SP18 and SP20, see Table 6) was identified and selected to target the excision and inversion of a fragment approximately 780 bp long from the 5' UTR of the At.AG gene. [Table 6]

[0575] Two plant transformation genes were constructed to edit the test construct, and each was designed to create a 5'UTR inversion edit via genome editing using a CRISPR / Cas12a (Cpf1) system, each having two guide RNAs targeting a site adjacent to the intended UTR sequence to be excised and inverted between two target sites. In this example, the vector construct typically includes two functional cassettes encoding the gene editing mechanism for creating targeted mutations in the At.AG gene: (i) a first cassette for the expression of a Cpf1 or Cas12a variant protein, and (ii) a second cassette for the expression of two relevant guide RNAs targeting the 5'UTR according to each editing scheme. Each guide RNA has a common scaffold compatible with the Cpf1 enzyme and a unique spacer / target sequence complementary to its intended target site. See Figure 2.

[0576] In the case of the gene-editing construct pM-01, the Cpf1 expression cassette contained a Medicago truncatula ubiquitin promoter (SEQ ID NO: 220), a dicotyledonous plant codon-optimized sequence encoding the Francisella tularensis subsp. novicida Cpf1 RNA-guided endonuclease enzyme (SEQ ID NO: 199) operably linked thereto, sequences encoding a nuclear localization signal at the N-terminus (SEQ ID NO: 200) and another nuclear localization signal at the C-terminus (SEQ ID NO: 201) fused thereto, followed by a Medicago trunctula terminator sequence (SEQ ID NO: 221). The gRNA expression cassette contained the Fn.Cpf1 compatible common scaffold SC2 (GR-FnCba.Cpf1:2, SEQ ID NO: 202), spacer SP18, scaffold SC2, spacer SP21, and the Arabidopsis plant Pol III promoter (SEQ ID NO: 222) operably bound to scaffold SC2. See Panel A in Figure 2. The gene editing construct pM-02 is similar to the pM-01 construct described above, except that the guide RNA contains SP18 and SP20 spacers. See Panel B in Figure 2. In addition to the above test constructs, two additional control vectors are generated. The miRNA vector control construct pM-03 contains a functional cassette containing a 35S promoter (SEQ ID NO: 210) operably ligated to a synthetic microRNA (miRNA) targeting the 5'UTR region within the At.AG gene for repression. This construct is expected to function as a positive control for gene silencing. Please refer to Figure 2, Panel C.

[0577] The 5'UTR transgenic test control construct pM-04 includes a functional cassette comprising a 35S promoter (SEQ ID NO: 210), operably linked thereto, an inverted fragment of the 5'UTR of At.AG, followed by an AG coding sequence, and a modified At.AG transcribable sequence including the native 3'UTR. See Figure 2, Panel C.

[0578] Arabidopsis plants are transformed with the above construct via the floral dipping method of transformation (see Clough SJ, Bent AF. Plant J. 1998 Dec;16(6):735-43). The treated plants are allowed to produce seeds, which are then sown in a selective medium and the transformants are screened.

[0579] In plants expressing pMON-01 and pMON-02, upon expression of gRNA and Cpf1 nuclease, the gRNA induces the nuclease at each of the two target sites containing the respective UTRs of the AG gene, where the nuclease forms a double-strand break at each target site. In most events, the region between the target sites is deleted, and a non-homologous end-joining repair mechanism binds to the adjacent region. Less frequently, the released UTR DNA fragment may be reoriented and reintegrated into the break site via NHEJ (non-homologous end joining) and innate DNA repair mechanisms, potentially resulting in an inversion of the DNA sequence between the two guide RNA target sites. Transformation events involving a complete UTR inversion are identified using appropriate methods known in the art (e.g., PCR, DNA hybridization, sequencing).

[0580] The expected phenotype of plants homozygous for the edited UTR inversion is expected to be equivalent to that of wild-type Arabidopsis, in that they should exhibit normal flowers with four petals and normal anther and stigma development. The expected phenotype of heterozygous plants with the edited UTR (5'UTR edited allele paired with the wild-type allele) is that they will exhibit an intermediate variant phenotype (more than four petals, reduced or excluded anthers or stigmas) compared to homozygous plants with AG gene knockout or coding region point mutations, which have fully double flowers with uncontrolled development of petals in the inner whorl of the flower, resulting in underdeveloped anthers and stigmas.

[0581] Following phenotypic characterization, small RNA sequencing was performed on several individuals exhibiting both wild-type and altered phenotypes to detect the presence of AG1-specific small RNA. While not bound by any theory, it is predicted that plants exhibiting phenotypic alteration will possess detectable levels of AG1-specific small RNA.

[0582] Example 4: Preparation and characterization of Arabidopsis AG1 transgenic UTR inverted plants Three AG1 UTR transgenic control constructs were generated. The 5'iUTR (inverse UTR) transgenic test control construct pM-04 contained a modified At.AG transcriptolytable sequence. The functional cassette included a 35S promoter (SEQ ID NO: 210), operably ligated to it an inverted fragment of the 5'UTR of At.AG (SEQ ID NO: 223), followed by an AG coding sequence (SEQ ID NO: 224), operably ligated to it a cotton-derived 3' transcriptolyzed terminator sequence (SEQ ID NO: 225). See Figure 3, Panel A.

[0583] The 3'iUTR transgenic test control construct pM-05 contained a functional cassette comprising a 35S promoter (SEQ ID NO: 210), an AG coding sequence operably linked thereto (SEQ ID NO: 224), followed by an inverted fragment of the 3'UTR of At.AG (SEQ ID NO: 226) and a cotton-derived 3' transcriptional terminator sequence (SEQ ID NO: 225). See Figure 3, Panel B.

[0584] The 5' and 3'iUTR transgenic test control construct pM-06 contained a functional cassette comprising a 35S promoter (SEQ ID NO: 210), operably ligated to it an inverted fragment of the 5'UTR of At.AG (SEQ ID NO: 223), followed by an AG coding sequence (SEQ ID NO: 224), operably ligated to it an inverted fragment of the 3'UTR of At.AG (SEQ ID NO: 226), and a cotton-derived 3' transcriptional terminator sequence (SEQ ID NO: 225). See Figure 3, Panel C. All constructs also contained cassettes with an adenylyltransferase (AAD) marker gene conferring resistance to spectinomycin.

[0585] Arabidopsis plants were transformed with the above construct via the floral dip transformation method (see Clough SJ, Bent AF. Plant J. 1998 Dec;16(6):735-43). The treated plants were allowed to produce seeds, which were then seedled in spectinomycin selective medium to screen for transformation. The viable R0 seedlings were advanced, and the Taqman assay was performed to determine the presence and copy number of the LbCas12a expression cassette, which also indicates the copy number of the iUTR cassette.

[0586] iUTR plants contain a natural wild-type copy of the AG1 gene, and therefore, the expected phenotype for these plants (iUTR transcript paired with wild-type allele transcript) is an altered or mutant phenotype whose severity is influenced by the number of transgene copies. Some of the expected altered phenotypes are a reduction or absence of more than four petals, anthers, or stigmas. It has been reported that knockout of the AG gene or homozygous state of coding region point mutations results in fully double flowers with underdeveloped anthers and stigmas due to uncontrolled development of petals in the inner whorl of the flower.

[0587] R0 transformants exhibiting abnormal phenotypes were characterized and summarized as shown in Table 7. The severity of altered floral phenotypes correlated with copy number, with multicopy plants tending to show more severe changes. [Table 7]

[0588] Data from the transgene indicate that an inversion in either the 3'UTR and / or 5'UTR, which does not alter the coding sequence, is sufficient to produce an altered phenotype.

[0589] After phenotypic characterization, small RNA sequencing is performed on several individuals to detect the presence of AG1-specific small RNA. While not bound by any theory, plants exhibiting phenotypic changes are predicted to possess detectable levels of AG1-specific small RNA.

[0590] Example 5: Creation of 5'UTR inverse editing of the Arabidopsis BRI1 gene The design of this gene editing approach is based on the excision and reverse insertion of a fragment of the 3'UTR region derived from the Arabidopsis thaliana brassinosteroid-insensitive 1 (BRI1) gene (AT4G39400).

[0591] The Arabidopsis thaliana BRI1 gene encodes a leucine-rich repeat receptor kinase involved in brassinosteroid signaling. BRI1 is the major receptor for brassinosteroids (see Wang ZY et al., Nature., 410, 6826, March 2001). BRI1 plays a crucial role in plant development, particularly in regulating cell elongation and tolerance to environmental stress. BRI1 enhances cell elongation, promotes pollen development, controls vascular structure development, and promotes cold and frost tolerance. As a result, bri1 null mutants exhibit extreme dwarfism, dark green, downward-curving leaves, male infertility, delayed flowering, altered vascular morphology, and reduced apical dominance, especially in older plants (see review by Clouse, The Arabidopsis Book, 2011 https: / / doi.org / 10.1199 / tab.0151).

[0592] The genome sequence of the At.BRI1 (AT4G39400, Chr4:18324660-18328826) gene is provided as Sequence ID No. 211. The elements and coordinates in the genome sequence are shown in Table 7. [Table 8]

[0593] The 3'UTR region was scanned, and two FnCas12a target sites were identified. Each contained a TTN PAM site adjacent to a unique spacer sequence that did not contain any known off-target sequences present in the Arabidopsis genome (in this case, an off-target sequence is defined as one that does not contain three or fewer mismatches to the 23-mer sequence of the gRNA). The spacer sequences are listed in Table 8. [Table 9]

[0594] In this experiment, we identified a pair of gRNA spacers (SP22 and SP23, see Table 8) and selected them to target the excision and inversion of a fragment approximately 360 bp long in the 3'UTR of the At.BRI1 gene. [Table 10]

[0595] We designed a plant transformation gene editing test construct to create a 3'UTR inversion edit via genome editing using a CRISPR / Cas12a (Cpf1) system having two guide RNA target sites adjacent to the intended UTR sequence, which are excised and inverted between the two target sites. In this example, the vector construct includes two functional cassettes encoding a gene editing mechanism for creating a targeted mutation in the At.BRI1 gene: (i) a first cassette for the expression of a Cpf1 or Cas12a variant protein, and (ii) a second cassette for the expression of two relevant guide RNAs targeting the 3'UTR according to the editing scheme. Each guide RNA contained a common scaffold compatible with the Cpf1 enzyme and a unique spacer / target sequence complementary to its intended target site.

[0596] In the case of the gene editing construct pM-07, the Cpf1 expression cassette contained a Medicago truncatula ubiquitin promoter (SEQ ID NO: 220), a dicotyledonous plant codon-optimized sequence encoding the Francisella tularensis subsp. novicida Cpf1 RNA guide endonuclease (SEQ ID NO: 199) operably linked thereto, a sequence fused thereto encoding a nuclear localization signal at the N-terminus of the Cpf1 enzyme (SEQ ID NO: 200) and another nuclear localization signal at the C-terminus (SEQ ID NO: 201), followed by a Medicago trunctula terminator sequence (SEQ ID NO: 221). The gRNA expression cassette included an Arabidopsis plant Pol III promoter (SEQ ID NO: 222), an Fn.Cpf1-compatible common scaffold SC2 (GR-FnCba.Cpf1:2, SEQ ID NO: 202) operably linked thereto, spacer SP22, scaffold SC2, spacer SP23, and scaffold SC2.

[0597] In addition to the test construct described above, two additional control vectors were generated and described in Example 6. Arabidopsis plants were transformed with the above construct via the floral dip transformation method (see Clough SJ, Bent AF. Plant J. 1998 Dec;16(6):735-43). The treated plants were allowed to produce seeds, which were then seedled in a selective medium and the transformants were screened.

[0598] In plants expressing pMON-07, upon expression of gRNA and Cpf1 nuclease, the gRNA induces the nuclease at each of two target sites within the UTR of the At.BRI1 gene, where the nuclease causes double-strand breaks at each target site. In most events, the region between the target sites is deleted, and a non-homologous end-joint repair mechanism binds to the adjacent region. Less frequently, the released UTR DNA fragment may be reoriented and reintegrated into the cleavage site via NHEJ (non-homologous end joining) and innate DNA repair mechanisms, potentially resulting in an inversion of the DNA sequence between the two guide RNA target sites. Transformation events involving complete UTR inversion are identified using appropriate methods known in the art (e.g., PCR, DNA hybridization, sequencing).

[0599] The expected phenotype of plants homozygous for edited UTR inversion is expected to be equivalent to that of wild-type Arabidopsis. The expected phenotype for heterozygous edited UTR plants (3'UTR edited allele paired with wild-type allele) is that they will exhibit intermediate mutant phenotypes, such as curled and wrinkled leaves, low stature, reduced apical dominance (i.e., increased number of stems / flower stalks), delayed flowering, and reduced fertility, compared to homozygous BRI1 gene knockout or coding region point mutations with severe stunted growth.

[0600] Following phenotypic characterization, small RNA sequencing is performed on several individuals to detect the presence of BRI1-specific small RNA. While not bound by any particular theory, it is predicted that plants exhibiting phenotypic changes will possess detectable levels of BRI1-specific small RNA.

[0601] Example 6: Preparation and characterization of Arabidopsis BRI1 transgenic UTR inverted plants Two BRI1UTR transgenic control constructs were generated. The 5'iUTR (inverse UTR) transgenic test control construct pM-08 contained a modified At.BRI1 transcriptable sequence. The functional cassette included a 35S promoter (SEQ ID NO: 210), operably ligated to it an inverted fragment of the 5'UTR of At.BRI1 (SEQ ID NO: 229), followed by a BRI1 coding sequence (SEQ ID NO: 230), operably ligated to it a cotton-derived 3' transcriptional terminator sequence (SEQ ID NO: 225).

[0602] The 3'iUTR transgenic test control construct pM-09 contained a functional cassette with a 35S promoter (SEQ ID NO: 210), a BRI1 coding sequence operably ligated thereto (SEQ ID NO: 230), followed by an inverted fragment of the 3'UTR of At.BRI1 (SEQ ID NO: 231). Both constructs also contained cassettes with an adenylyltransferase (AAD) marker gene that conferred resistance to spectinomycin.

[0603] Arabidopsis plants were transformed with the above construct via the floral dipping method of transformation (see Clough SJ, Bent AF. Plant J. 1998 Dec;16(6):735-43). The treated plants were allowed to produce seeds, which were then seedled in spectinomycin selective medium to screen for transformation. The viable R0 seedlings were advanced and subjected to a Taqman assay to determine the presence and copy number of LbCas12a expression cassettes.

[0604] iUTR plants contain a natural wild-type copy of the BRI1 gene, and therefore, the expected phenotype for plants containing these transgenes (iUTR transcript paired with wild-type allele transcript) is that they exhibit altered or mutant phenotypes, the severity of which is influenced by the number of transgene copies. Some of the expected altered phenotypes are altered phenotypes that include curled and wrinkled leaves, low plant height, reduced apical dominance (i.e., increased number of stems / flower stalks), delayed flowering, and reduced fertility. Due to higher expression from the CaMV35S promoter than from the natural BRI1 promoter, more potent phenotypic changes may be observed in transgenic plants compared to gene-edited UTR inversions.

[0605] R0 transformants exhibiting abnormal phenotypes were characterized and summarized as shown in Table 8. The severity of the altered phenotype correlated with copy number, with multicopy plants tending to show more severe changes. [Table 11]

[0606] Data from the transgene indicate that an inversion at either the 3'UTR or 5'UTR of BRI1, without altering the coding sequence, is sufficient to produce the altered phenotype.

[0607] Following phenotypic characterization, small RNA sequencing is performed on several individuals to detect the presence of BRI1-specific small RNA. While not bound by any particular theory, it is predicted that plants exhibiting phenotypic changes will possess detectable levels of BRI1-specific small RNA.

[0608] Example 7: Creation of UTR inverse editing of Arabidopsis PIN1 and D genes Using similar methods and processes described in Examples 2, 3, and 5, edited Arabidopsis plants containing UTR inversions in the Dwarf1 gene (At.Dwf1, At3g19820) (SEQ ID NO: 212) and the Pin-formed1 gene (At.PIN1, At1g73590) (SEQ ID NO: 213) are generated and characterized. The At.Dwf1 protein is involved in brassinosteroid hormone synthesis and, in turn, regulates cell elongation. The dwf1 mutant has a dwarf phenotype, as its name suggests (see Clouse (2011) The Arabidopsis Book https: / / doi.org / 10.1199 / tab.0151). Pin-formed1 encodes an auxin efflux carrier involved in shoot and root development. Loss of function severely affects organ initiation, and the pin1 mutant is characterized by inflorescence meristems that do not initiate any flowers, resulting in the formation of a bare stem (see Okada, et al., The Plant Cell, Volume 3, Issue 7, July 1991, pp. 677-684).

[0609] In short, we scan the 5' and 3' UTR regions of the At.Dwf1 and At.PIN1 genes to identify FnCas12a target sites. We identify pairs of gRNA spacers and select them to target excision and inversion of approximately 150 bp long fragments of the 5' UTR of the At.Dwf1 gene, approximately 65 bp long fragments of the 5' UTR of At.PIN1, or approximately 150 bp long fragments of the 3' UTR of At.PIN1. We created plant transformation gene editing test constructs, each designed to create 5' UTR inversion edits via genome editing using a CRISPR / Cas12a (Cpf1) system containing two guide RNA target sites adjacent to the intended UTR sequence, which excise and invert between each of the two target sites.

[0610] Arabidopsis plants are transformed with the above construct via a floral dipping method for transformation (see Clough SJ, Bent AF. Plant J. 1998 Dec;16(6):735-43). The treated plants are allowed to produce seeds, which are then seeded in a selective medium and the transformants are screened. Transformation events, including complete inversion of the target UTR, are identified using appropriate methods known in the art (e.g., PCR, DNA hybridization, sequencing). Homozygous and heterozygous plant phenotypes for inversion are characterized, and small RNA sequencing is performed on plants with wild-type and altered phenotypes to detect the presence of small RNAs targeting the gene of interest.

[0611] Example 8. Editing of the 3'UTR region of the Zm.GA3ox_1 gene and generation and confirmation of conjugation in genome-edited plants. Inbred wild-type maize lines were transformed via Agrobacterium-mediated transformation using the pM313 vector described in Example 2 above. The transformed plant tissues were grown to create mature R0 plants. R1 plants were created by autocrossing R0 plants that had one or more unique genome edits. R2 plants were created by autocrossing R1 plants that were homozygous for an allele containing the edited Zm.GA3ox_1 gene, which has a 3'UTR inversion and lacks a T-DNA sequence.

[0612] To determine whether editing had occurred in the Zm.GA3ox_1 3'UTR region, amplicon sequencing techniques were used to produce mutant sequences of the 750kb 3'UTR region for comparison with the wild-type sequence. Amplicon sequencing involves generating one or more unique PCR products across the entire genomic region of interest for next-generation sequencing. Sequence data from each sample were then mapped to a reference sequence to identify differences in the consensus sequence. Plants with unique inversions were selected. Individual R1 plants produced by self-fertilizing R0 plants with editing were assayed to confirm the editing. For illustrative purposes, one edited plant with the GA3ox_1 gene allele S049, generated using the pM313 transformation vector, was selected and characterized as described in Table 12. In Table 12, "Allele Name" is the identifier of the unique allele, and Wild Type (WT) refers to the unedited plant with the WT GA3ox_1 gene. The nucleotide positions of the edit inversion and deletion start and end refer to a 4800 bp GA3ox_1 wild-type genome sequence (SEQ ID NO: 232), which begins with a 2000 bp promoter sequence upstream of the transcription start site and ends with a 3' UTR sequence (SEQ ID NO: 233) at nucleotide positions 4052-4800. The WT and edited allele sequences and their relationships are shown in Figure 4. Due to the editing process, the WT 3' UTR sequence of SEQ ID NO: 233 was converted to the shorter UTR sequence of SEQ ID NO: 235, and the 4800 bp WT sequence of SEQ ID NO: 232 was converted to the 4642 bp edited genome sequence of SEQ ID NO: 234.

[0613] The conjugation of each editing allele in the R1 plants was further tested to determine whether they were homozygous or heterozygous. Homozygous edited plants were selected and further self-fertilized to produce inbred plants, or crossed with other maize lines to produce hybrid plants. [Table 12]

Claims

1. A method for editing the genome of plant cells in order to modify endogenous genes, a) A step of generating a first double-strand break and a second double-strand break using a targeted editing technique that targets at least one untranslated region of the endogenous gene without perturbation in the coding region within the plant cell, b) The method comprising the step of isolating a modified plant cell containing a modifying allele of the endogenous gene, wherein the modifying allele contains an inverted DNA sequence of at least a portion of the untranslated region of the endogenous gene, the modifying allele produces an RNA transcript containing an antisense sequence of a portion of the untranslated region, and the modifying allele does not contain a sense sequence complementary to the antisense sequence of a portion of the untranslated region.

2. The method according to claim 1, wherein at least a portion of the inverted DNA of the untranslated region of the endogenous gene encodes an antisense RNA sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 99%, or 100% complementary to at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides of the untranslated region of the gene.

3. The method according to claim 1, wherein the untranslated region of the endogenous gene is a 5' untranslated region, a 3' untranslated region, or both.

4. The method according to claim 1, wherein the plant cell comprises a modifying allele and an unmodified allele of the endogenous gene.

5. The method according to claim 4, wherein the modified allele and the unmodified allele are transcribed into mRNA, and the RNA derived from the modified allele and the RNA derived from the unmodified allele can generate a double-stranded RNA region of at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides.

6. The method according to claim 5, wherein the expression of the modified allele and the unmodified allele is reduced.

7. The method according to claim 1, further comprising the step of regenerating a plant from the modified plant cells.

8. The method according to claim 7, further comprising the step of crossing a plant containing the modifying allele of the endogenous gene in a homozygous state with another plant containing the unmodifying allele of the endogenous gene in a homozygous state, and harvesting hybrid seeds.

9. The method according to claim 1, wherein the plant is maize, and the endogenous gene is selected from GA20 oxidase or GA3 oxidase.

10. The aforementioned inverted DNA sequences are nucleotides 1-29 of SEQ ID NO: 36, nucleotides 1664-1788 of SEQ ID NO: 36, nucleotides 1-38 of SEQ ID NO: 37, nucleotides 1446-1698 of SEQ ID NO: 37, nucleotides 3001-3161 of SEQ ID NO: 168, nucleotides 4796-5406 of SEQ ID NO: 168, nucleotides 3001-3056 of SEQ ID NO: 169, nucleotides 4464-4581 of SEQ ID NO: 169, sequence Nucleotides 3001-3130 of sequence number 170, nucleotides 4275-4332 of sequence number 170, nucleotides 7621-8029 of sequence number 174, nucleotides 9672-10276 of sequence number 174, nucleotides 7386-7831 of sequence number 175, nucleotides 8862-8967 of sequence number 175, nucleotides 7547-7751 of sequence number 176, nucleotides 8904-9178 of sequence number 176, The method according to claim 1, comprising a nucleotide sequence having at least 90% sequence identity or complementarity with at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides of sequence number 204, nucleotides 1 to 1060 of sequence number 204, nucleotides 5418 to 5648 of sequence number 204, nucleotides 1 to 165 of sequence number 211, nucleotides 3757 to 4167 of sequence number 211, nucleotides 664 to 699 of sequence number 212, nucleotides 2482 to 2700 of sequence number 212, nucleotides 1 to 99 of sequence number 213, or nucleotides 3205 to 3506 of sequence number 213.

11. The aforementioned targeted editing technique involves the use of RNA effector proteins, wherein the RNA guide endonucleases include Cas9, C2c1, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, and Ca sl, CaslB, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Csnl, Csx12, Cas1 0, Csyl, Csy2, Csy3, Csel, Cse2, 30Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Cs m6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX , Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4 (dinG), Csf5 nuclease, Cas12c (C2c3), Cas1 The method according to claim 1, wherein the CRISPR-Cas effector protein is selected from 2d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and Cas14c effector proteins.

12. The method according to claim 11, wherein the RNA guide effector protein is a Cas12a effector protein or an effector protein derived from Cas12a.

13. The method according to claim 12, wherein the targeted gene editing technique includes the use of one or more guide RNAs comprising a nucleotide sequence selected from the group SEQ ID NOs: 177, 178, 179, 180, 205, 206, 207, 208, and 209.

14. A method for modifying the expression of endogenous genes in hybrid plants, such that the expression of the said genes is not affected in the parent plant(s), a) A step of identifying an endogenous gene in a plant, wherein the expression of a variant allele of the gene, when present in a homozygous state, results in an undesirable phenotype; b) A step of providing a first plant comprising a modifying allele of the gene, the modifying allele of the gene comprising at least one nucleic acid region which is an inversion of a part of the gene, wherein the inversion does not affect the translation of the modifying allele, and the first plant homozygously comprises the modifying allele of the gene, c) A step of crossing the first plant with a second plant containing an unmodified allele of the gene that does not have the inversion that does not affect the translation of the gene, wherein the unmodified gene is in a homozygous state, d) The method comprising the step of obtaining a hybrid seed containing the modifying allele and the unmodified allele of the gene in a heterozygous or heteroallelic form.

15. The method according to claim 14, wherein when the modified allele and the unmodified allele are transcribed into an RNA molecule, a double-stranded RNA region may be formed by base pairing between a nucleic acid region which is an inversion of a part of the untranslated region of the gene in the RNA transcript of the modified allele and a nucleic acid region in the RNA transcript of the unmodified allele, and the double-stranded RNA region can inhibit the expression of the modified allele and the unmodified allele by RNA silencing mechanisms such as RNA translation stalling, RNA transcription stalling, resulting destabilization of the RNA molecule, or post-transcriptional degradation of the transcribed RNA molecule.

16. The method according to claim 14, wherein the nucleic acid region which is an inversion of a portion of the gene is generated during the transcription of an antisense RNA sequence which is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides of the untranslated region of the gene.

17. The method according to claim 14, wherein the GA20 oxidase is selected from GA20 oxidase_5 or GA20 oxidase_3.

18. A plant cell, plant or part thereof or seed containing a modifying allele of an endogenous gene, wherein the modifying allele contains an inverted DNA sequence of at least a portion of the untranslated region of the endogenous gene, the modifying allele produces an RNA transcript containing an antisense sequence of a portion of the untranslated region, and the modifying allele does not contain a sense nucleotide sequence of more than 17 nucleotides that is complementary to the antisense sequence of a portion of the untranslated region.

19. The plant cell, plant or part thereof or seed according to claim 18, wherein the inverted DNA sequence of at least a portion of the untranslated region of the endogenous gene occurs during the transcription of an antisense RNA sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, or at least 1000 consecutive nucleothiodies of the untranslated region of the gene.

20. The plant cell, plant or part thereof or seed according to claim 18, wherein the plant is maize and the endogenous gene is selected from GA20 oxidase or GA3 oxidase.

21. The aforementioned inverted DNA sequences are: nucleotides 1-29 of SEQ ID NO: 36, nucleotides 1664-1788 of SEQ ID NO: 36, nucleotides 1-38 of Hybrid 37, nucleotides 1446-1698 of SEQ ID NO: 37, nucleotides 3001-3161 of SEQ ID NO: 168, nucleotides 4796-5406 of SEQ ID NO: 168, nucleotides 3001-3056 of SEQ ID NO: 169, nucleotides 4464-4581 of SEQ ID NO: 1 Nucleotides 3001-3130 of 70, nucleotides 4275-4332 of SEQ ID NO: 170, nucleotides 7621-8029 of SEQ ID NO: 174, nucleotides 9672-10276 of SEQ ID NO: 174, nucleotides 7386-7831 of SEQ ID NO: 175, nucleotides 8862-8967 of SEQ ID NO: 175, nucleotides 7547-7751 of SEQ ID NO: 176, nucleotides 8904-9178 of SEQ ID NO: 204 The plant cell, plant or part thereof or seed according to claim 18, comprising a nucleotide sequence having at least 90% sequence identity or complementarity with at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750 or at least 1000 consecutive nucleotides of nucleotides 1 to 1060, nucleotides 5418 to 5648 of SEQ ID NO: 204, nucleotides 1 to 165 of SEQ ID NO: 211, nucleotides 3757 to 4167 of SEQ ID NO: 211, nucleotides 664 to 699 of SEQ ID NO: 212, nucleotides 2482 to 2700 of SEQ ID NO: 212, nucleotides 1 to 99 of SEQ ID NO: 213, or nucleotides 3205 to 3506 of SEQ ID NO:

213.

22. The plant cell, plant or part thereof or seed according to claim 18, further comprising an unmodified allele of the endogenous gene.

23. The modified allele and the unmodified allele are transcribed into RNA, and the RNA derived from the modified allele and the RNA derived from the unmodified allele can generate a double-stranded RNA region of at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, or at least 1000 consecutive nucleotides, and the double-stranded RNA region can inhibit the expression of the modified allele and the unmodified allele by RNA silencing mechanisms such as, for example, RNA translation stalling, RNA transcription stalling, resulting destabilization of the RNA molecule, or post-transcriptional degradation of the transcribed RNA molecule, according to claim 22.

24. The plant cell, plant or part thereof or seed according to claim 118, wherein the plant is selected from monocotyledonous plant species, dicotyledonous plant species, angiosperm species, or gymnosperm species.

25. The plant according to claim 18, wherein the plant is selected from corn plants, rice plants, sorghum plants, wheat plants, alfalfa plants, barley plants, millet plants, rye plants, sugarcane plants, cotton plants, soybean plants, canola plants, tomato plants, onion plants, cucumber plants, Arabidopsis plants, or potato plants.