Compositions and methods for increasing genome editing efficiency

JP2025512041A5Pending Publication Date: 2026-04-15JOHN INNES CENT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
JOHN INNES CENT
Filing Date
2023-04-10
Publication Date
2026-04-15

AI Technical Summary

Technical Problem

The prior art has low gene editing efficiency of the CRISPR-Cas9 system in certain plant species, making it difficult to effectively modify the plant genome.

Method used

New Cas12a nucleic acid variants and their corresponding recombinant DNA molecules were developed, which significantly improved the efficiency of plant gene editing by binding to adapted guide RNA.

Benefits of technology

The efficiency and accuracy of plant gene editing is significantly improved, especially in plant species with low editing efficiency, which can generate target mutations more efficiently, and these mutations can be stably inherited in offspring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Compositions and methods are provided for improving gene editing efficiency in plants.Methods and compositions are also provided for producing modifications using novel Cas12a nuclease variants.Modified plant cells and plants are further provided that comprise the DNA and protein compositions of novel Cas12a nuclease variants.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 330,106, filed April 12, 2022, and U.S. Provisional Application No. 63 / 386,452, filed December 7, 2022, the entire contents of each of which are incorporated herein by reference.

[0002] Inclusion of sequence listing The Sequence Listing, including a file named "AGOE008US_ST26.xml", which is 94 kilobytes (measured in MS-Windows™), was created on April 6, 2023, and contains 58 sequences, is incorporated herein by reference in its entirety.

[0003] The present disclosure relates to the field of plant molecular biology and plant genetic engineering, and methods and compositions for genome editing in plants.In particular, the present invention relates to novel Cas12a nuclease variants and methods for improving gene editing efficiency.Using plant genetic engineering methods, Cas12a DNA and encoded protein are modified and these molecules are transferred to agriculturally important plants.More specifically, the present invention includes the DNA and protein compositions of novel LbCas12a nuclease variants, and the plants containing these compositions. [Background technology]

[0004] Precision genome editing technology is a powerful tool for manipulating gene expression and regulating protein function, and has the potential to improve important agricultural properties. In particular, the clustered regularly interspaced short palindromic repeats (CRISPR)-Cas9 system has revolutionized the field of genome editing. However, the editing efficiency of this powerful tool is still very low in some plant species. Therefore, there is a continuing need in the art to develop novel compositions and methods for increasing the efficiency of genome editing in plants. Summary of the Invention

[0005] In one aspect, the disclosure provides a recombinant DNA molecule comprising a polynucleotide sequence selected from the group consisting of: (a) a sequence having at least 85% identity to any of SEQ ID NOs: 1, 3, 5, 7, and 8; (b) a sequence comprising SEQ ID NOs: 1, 3, 5, 7, and 8; (c) a fragment of any of SEQ ID NOs: 1, 3, 5, 7, and 8; and (d) a sequence encoding a protein having at least 85% identity to any of SEQ ID NOs: 2, 4, 6, and 9. In some embodiments, the protein encoded by said polynucleotide sequence comprises a modification at amino acid position 156 compared to a protein comprising the amino acid sequence of SEQ ID NO: 46. For example, the recombinant DNA molecule has at least 90% identity or at least 95% identity to any of SEQ ID NOs: 1, 3, 5, 7, and 8 and encodes a protein having a modification at amino acid position 156 compared to a protein comprising the amino acid sequence of SEQ ID NO: 46. In some embodiments, the recombinant DNA molecule provided herein comprises any of SEQ ID NOs: 1, 3, 5, 7, and 8. In a particular example, the modification at amino acid position 156 compared to SEQ ID NO:46 is further defined as a substitution of an aspartic acid with an arginine.

[0006] In another aspect, the disclosure provides a recombinant DNA molecule comprising a polynucleotide sequence selected from the group consisting of: a) a sequence having at least 85% identity to any of SEQ ID NOs: 1, 3, 5, 7, and 8; b) a sequence including SEQ ID NOs: 1, 3, 5, 7, and 8; c) a fragment of any of SEQ ID NOs: 1, 3, 5, 7, and 8; and d) a sequence encoding a protein having at least 85% identity to any of SEQ ID NOs: 2, 4, 6, and 9, and further comprising at least one intron sequence having a sequence of any of SEQ ID NOs: 10-17. In some embodiments, the polynucleotides provided herein comprise one or more intron sequences of any of SEQ ID NOs: 10-17.

[0007] In yet another aspect, a transgenic plant cell is described that comprises a recombinant DNA molecule provided herein. The transgenic plant cell provided may be a monocotyledonous plant cell, including, but not limited to, barley, B. oleracea, wheat and corn cells. The transgenic plant cell provided may be a dicotyledonous plant cell. A transgenic plant or part thereof that comprises a recombinant DNA molecule described herein is further provided. A progeny plant that comprises a DNA molecule provided herein is further described. The present disclosure further provides a transgenic seed that comprises a recombinant DNA molecule described herein.

[0008] The recombinant DNA molecules described herein may be expressed in plant cells to result in genomic modification and may be operably linked to a vector, said vector being selected from the group consisting of a plasmid, a phagemid, a bacmid, a cosmid, and a bacterial or yeast artificial chromosome.

[0009] The recombinant DNA molecule provided herein can be present in a host cell, which can be any type of cell. The host cell contemplated by the present disclosure includes a cell selected from the group consisting of bacterial cells, animal cells, plant cells, yeast cells, fungal cells and insect cells. For example, the bacterial host cell can be from a bacterial genus selected from the group consisting of Agrobacterium, Rhizobium, Bacillus, Brevibacillus, Escherichia, Pseudomonas, Klebsiella, Pantoea and Erwinia.

[0010] The animal host cell may include a mammalian host cell, such as a fibroblast, an epithelial cell, a lymphocyte, or a macrophage. The animal host cell according to the present disclosure may be an immortalized animal cell line, a primary cell, or a stem cell.

[0011] In another example, the plant cell may be a dicotyledonous or monocotyledonous plant cell, such as a plant cell selected from the group consisting of Fabaceae, sunflower, safflower, sesame, tobacco, potato, cotton, sweet potato, cassava, coffee, tea, apple, pear, fig, citrus, cocoa, avocado, olive, almond, walnut, strawberry, watermelon, pepper, beet, grape, tomato, cucumber, Arabidopsis, Brassica sp., pea, alfalfa, barrel clover, pigeon pea, guar, carob, fenugreek, soybean, kidney bean, cowpea, mung bean, lima bean, broad bean, lentil, peanut, licorice, chickpea, oil palm, coconut, banana, corn, barley, sorghum, rice, and wheat cells.

[0012] In another aspect, the present disclosure provides a method for producing a plant comprising a genome modification, the method comprising: (a) expressing in a plant cell a recombinant DNA molecule according to claim 1 and a guide RNA that is compatible with the protein encoded by said recombinant DNA molecule; (b) introducing a modification into at least one target site in the genome of the plant cell; (c) identifying and selecting one or more plant cells of step (b) that comprise said modification in the plant genome; and (d) regenerating at least one plant from at least one or more cells selected in step (c). In certain examples, the modification can be a substitution, an insertion, an inversion, a deletion, a duplication, and a combination thereof. In some embodiments, the plant for use in the provided method can be a monocotyledonous plant, such as barley, B. oleracea, wheat or corn plant.

[0013] In another aspect, the disclosure provides a method for improving gene targeting using CRISPR-Cas12a gene editing in crop plants, the method comprising: a recombinant DNA molecule comprising a polynucleotide sequence selected from the group consisting of a sequence having at least 85% identity to any of SEQ ID NOs: 1, 3, 5, 7 and 8; a sequence comprising SEQ ID NOs: 1, 3, 5, 7 and 8; a fragment of any of SEQ ID NOs: 1, 3, 5, 7 and 8, and / or a sequence encoding a protein having at least 85% identity to any of SEQ ID NOs: 2, 4, 6 and 9, and expressing a guide RNA compatible with the protein encoded by said recombinant DNA molecule in a plant cell; and / or introducing a modification at at least one target site in the genome of the plant cell, wherein the modification is introduced at a higher rate compared to the rate of introduction of the modification using a method comprising expressing a DNA molecule encoding the amino acid of SEQ ID NO: 46. In some embodiments, the sequence has at least 90% identity to any of SEQ ID NOs: 1, 3, 5, 7, and 8 and encodes a protein having a modification at amino acid position 156 compared to a protein comprising the amino acid sequence of SEQ ID NO: 46. In some embodiments, the sequence has at least 95% identity to any of SEQ ID NOs: 1, 3, 5, 7, and 8 and encodes a protein having a modification at amino acid position 156 compared to a protein comprising the amino acid sequence of SEQ ID NO: 46. In some embodiments, the sequence comprises any of SEQ ID NOs: 1, 3, 5, 7, and 8. In some embodiments, the modification at amino acid position 156 is further defined as a substitution of aspartic acid for arginine. In some embodiments, the polynucleotide sequence further comprises an intron sequence of SEQ ID NOs: 10-17.

[0014] Further provided is a method of producing a progeny seed comprising a recombinant DNA molecule described herein, the method comprising: (a) sowing an initial seed comprising the recombinant DNA molecule of claim 1; (b) growing a plant from the seed of step (a); and (c) harvesting progeny seed from the plant, the harvested seed comprising the recombinant DNA molecule.

[0015] In yet another aspect, the present disclosure provides a method for introducing a genomic modification into a plant, the method comprising: (a) expressing in a plant a protein or a fragment thereof encoded by a DNA molecule provided herein; and (b) expressing in a plant cell a guide RNA that matches the protein or a fragment thereof having nuclease activity.

[0016] The present disclosure further provides a method for detecting the presence of a recombinant DNA molecule provided herein in a sample containing plant genomic DNA, the method comprising: (a) contacting the sample with a DNA probe that hybridizes under stringent hybridization conditions to genomic DNA from a plant containing recombinant nuclear DNA and does not hybridize under such hybridization conditions to genomic DNA from other isogenic plants that do not contain recombinant DNA molecules, the probe being homologous or complementary to a fragment of any of SEQ ID NOs: 1, 3, 5, 7, 8, or a sequence encoding a protein comprising an amino acid sequence having at least 85%, or 90%, or 95%, or 98%, or 99%, or about 100% amino acid sequence identity to any of SEQ ID NOs: 2, 4, 6, and 9; (b) subjecting the sample and the probe to stringent hybridization conditions; and (c) detecting hybridization of the DNA probe to the recombinant DNA molecule.

[0017] In another aspect, the disclosure provides a method of detecting the presence of a nuclease protein or a fragment thereof in a sample containing a protein, wherein the protein comprises the amino acid sequence of any of SEQ ID NOs: 2, 4, 6, and 9, or the protein comprises an amino acid sequence having at least 85%, or 90%, or 95%, or 98% or 99%, or about 100% amino acid sequence identity to any of SEQ ID NOs: 2, 4, 6, and 9, the method comprising: (a) contacting the sample with an immunoreactive antibody; and (b) detecting the presence of the protein or a fragment thereof.

[0018] In further embodiments, methods are provided for modifying a polynucleotide segment encoding a Cas12a protein or a fragment thereof having nuclease activity, the methods comprising: (a) obtaining a polynucleotide sequence of any of SEQ ID NOs: 1, 3, 5, 7 and 8; and (b) introducing a modification at at least one target site in the polynucleotide sequence such that a protein encoded by said polynucleotide sequence comprises a modification at amino acid position 156, compared to a protein comprising the amino acid sequence of SEQ ID NO: 46. In these methods, the protein encoded by the modified polynucleotide sequence comprises an aspartic acid to arginine substitution at amino acid position 156, compared to a polynucleotide segment lacking said modification. The modified polynucleotide sequence may further comprise at least one intron sequence of any of SEQ ID NOs: 10-17, or may comprise one or more intron sequences of any of SEQ ID NOs: 10-17. In further examples, the modified polynucleotide sequence comprises an aspartic acid to arginine modification at amino acid position 156 and further comprises at least one intron sequence of SEQ ID NOs: 10-17.

[0019] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure. The present disclosure may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein. [Brief description of the drawings]

[0020] [Figure 1]A schematic of the editing constructs tested in barley is shown. Briefly, P-ZmUbi refers to the maize ubiquitin promoter, Cas12a refers to the LbCas12a CDS, T-Nos refers to the nopaline synthase terminator, TaU6 refers to the wheat U6 promoter, TaU3 refers to the wheat U3 promoter, DR refers to the direct repeat crRNA, HH / HDV refers to the ribozyme sequence, t refers to the polyT terminator, V1 refers to the V1 array, and V2 refers to the V2 array. The thick black arrow indicates the direction of transcription. [Diagram 2] Figure 1 shows the efficiency of targeting the HORVU.MOREX.r31HG0069960 gene using V1 guide arrays with different LbCas12a constructs. Os refers to OsCas12a, Hs refers to HsCas12a, ttHs refers to ttHsCas12a, ttAt refers to ttAtCas12a, and ttAt+int refers to ttAtCas12a+int. The blue bar indicates the number of T0 lines. The orange bar indicates the number of T0 lines containing targeted mutations. [Diagram 3] Five barley genes, each targeted with ttHsCas12a using the V1 array, are compared to the V2 array. Blue bars indicate TO V1 lines containing the targeted mutations. Orange bars indicate TO V2 lines containing the targeted mutations. The x-axis indicates the order of the array guides. Gene identifiers are indicated. [Figure 4] A representative phenotype comparison of Golden Promise, which has a wild-type 2-line phenotype, compared to Golden Promise TO plants mutagenized with HORVU.MOREX.r3.2HG0184740, which exhibit a 6-line phenotype, is shown. [Diagram 5]Figure 1 shows a sequencing analysis of the HORVU.MOREX.r3.1HG0069960 gene in representative barley lines. Top: Amplicon sequencing revealed the presence of two alleles in the founder T0 generation (-3bp; TTTGGTGCTGCACAATGAAAGCAGACGGC; SEQ ID NO: 50; and -10bp; TTTGGTGCTGCACAACAACAACTGAAAGCAGACGGC; SEQ ID NO: 51). Bottom: In the T1 progeny without T-DNA, the same two alleles were identified, establishing inheritance of the mutation. The bottom left panel shows the unedited sequence along the top (TTTGGTGCTGCACAATGTCAACAACTGAAAGCAGACGGC; SEQ ID NO: 52) compared to the sequence of the T1 homozygous 3bp deletion (SEQ ID NO: 50). The bottom center panel shows the unedited sequence along the top (SEQ ID NO: 52) compared to the T1 homozygous 10bp deletion (SEQ ID NO: 51). The bottom right panel shows the unedited sequence along the top (SEQ ID NO:52) compared to the sequence of the T1 heterozygote (GTTGATGGTTGGTGTTGGGCAATGCCCAATGAAAGCAGACGGC; SEQ ID NO:53). [Figure 6A]A schematic of the construct editing tested in B. Oleracea is shown. Briefly, Nos refers to the nopaline synthase terminator, Npt refers to neomycin phosphotransferase (confers kanamycin resistance for bacterial selection of the plasmid), 35S refers to the cauliflower mosaic virus_35S promoter, and E9 refers to the rbc-E9 terminator (Pisum sativum), ttAtCas12a refers to Arabidopsis codon-optimized LbCas12a with the D156R "temperature tolerance" mutation, ttHsCas12a refers to the Homo sapiens codon-optimized LbCas12a coding sequence with the "temperature tolerance" D156R mutation, and ttAtCas12a+int refers to the Arabidopsis codon-optimized LbCas12a with the D156R "temperature tolerance" mutation. and 8 Arabidopsis introns, Ubi10 refers to the Arabidopsis ubiquitin 10 promoter, U6 refers to the Arabidopsis U626 promoter, HH / HDV refers to the ribozyme sequence, DR refers to the direct repeat crRNA, G_A, _B, _C, and _D refer to the protospacers A, B, C and D, and t refers to the polyT terminator. [Figure 6B] Figure 1 shows a comparison of the mutagenesis efficiency of LbCas12a constructs S5, S6, S7, and S8 targeting Bo2g016480. Comparison of S5, S6, S7, and S8 is possible on target C, where the respective efficiencies were 3%, 50%, 50%, and 68%. [Figure 7]1 shows a sequencing analysis of the Bo2g016480 gene in T-DNA-free T1 B. Oleracea plants. The -3bp, -9bp and -12bp alleles were revealed, establishing inheritance of the mutation. The left panel shows the unedited sequence GAGTTTGGTATGCAGATCAACATTATAAGAATGTACC (SEQ ID NO: 54) along the top compared to the sequence of the T1 homozygous 3bp deletion (GAGTTTTGGTATGCAGATCAACATAAGAATGTACC; SEQ ID NO: 55). The center panel shows the unedited sequence (SEQ ID NO: 54) along the top compared to the sequence of the T1 homozygous 9bp deletion (GAGTTTTGGTATGCAGATCAACATGTACC; SEQ ID NO: 56). The right panel shows the unedited sequence (SEQ ID NO: 54) along the top compared to the sequence of the T1 homozygous 12bp deletion (GAGTTTTGGTATGCAGATCAAGTACC; SEQ ID NO: 57). [Figure 8] A universal genetic code chart showing all possible mRNA triplet codons (T in DNA molecules is replaced by U in RNA molecules) and the amino acid coded for by each codon. [Figure 9] 1 shows the construct structure for evaluating the gene editing efficiency of ttHsCas12a and ttAtCas12a+8 intron nuclease in wheat. [Figure 10] 1 shows the construct structure for evaluating the gene editing efficiency of ttAtCas12a+8 intron nuclease in wheat. [Figure 11] 1 shows the construct structure for evaluating the gene editing efficiency of ttAtCas12a nuclease with and without introns in Arabidopsis thaliana. [Figure 12] 1 shows further construct structures for evaluating the gene editing efficiency of Cas12a variants in barley. [Figure 13] Construct structures of 12 LbCas12a coding sequence variants are shown. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0021] A brief description of the sequence SEQ ID NO:1 is the polynucleotide sequence of the Lachnospiraceae bacterial Cas12a gene, codon-optimized for expression in rice (OsCas12a).

[0022] SEQ ID NO:2 is the amino acid sequence of the Lachnospiraceae bacterial Cas12a protein encoded by SEQ ID NO:1 (OsCas12a).

[0023] SEQ ID NO:3 is the polynucleotide sequence of the Lachnospiraceae bacterial Cas12a gene, codon-optimized for expression in Homo sapiens (HsCas12a).

[0024] SEQ ID NO:4 is the amino acid sequence of the Lachnospiraceae family Cas12a protein encoded by SEQ ID NO:3 (HsCas12a).

[0025] SEQ ID NO:5 is the polynucleotide sequence of the Lachnospiraceae bacterial Cas12a gene, which encodes a protein that is codon-optimized for expression in Homo sapiens and has a D156R mutation compared to the wild-type Cas12a protein (ttHsCas12a).

[0026] SEQ ID NO:6 is the amino acid sequence of the Lachnospiraceae bacterium Cas12a protein encoded by SEQ ID NO:5 (ttHsCas12a).

[0027] SEQ ID NO: 7 is the polynucleotide sequence of the Lachnospiraceae bacterial Cas12a gene, which encodes a protein that is codon-optimized for expression in Arabidopsis and has a D156R mutation compared to the wild-type Cas12a protein (ttAtCas12a).

[0028] SEQ ID NO:8 is the polynucleotide sequence of a Lachnospiraceae bacterial Cas12a gene that is codon-optimized for expression in Arabidopsis, encodes a protein having a D156R mutation compared to the wild-type Cas12a protein, and further contains 8 intron sequences (ttAtCas12a+int).

[0029] SEQ ID NO: 9 is the amino acid sequence of the Lachnospiraceae bacterial Cas12a protein encoded by SEQ ID NOs: 7 and 8 (ttAtCas12a and ttAtCas12a+int, respectively).

[0030] SEQ ID NOs:10 to 17 are polynucleotide sequences of introns within SEQ ID NO:8.

[0031] SEQ ID NO: 18 is the polynucleotide sequence of the V1 guide RNA array construct.

[0032] SEQ ID NO:19 is the polynucleotide sequence of the V2 guide RNA array construct.

[0033] SEQ ID NO:20 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.1HG0069960.

[0034] SEQ ID NO:21 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.1HG0069960.

[0035] SEQ ID NO:22 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.1HG0069960.

[0036] SEQ ID NO:23 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.1HG0069960.

[0037] SEQ ID NO:24 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.2HG0184740.

[0038] SEQ ID NO:25 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.2HG0184740.

[0039] SEQ ID NO:26 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.2HG0184740.

[0040] SEQ ID NO:27 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.2HG0184740.

[0041] SEQ ID NO:28 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.6HG0611290.

[0042] SEQ ID NO:29 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.6HG0611290.

[0043] SEQ ID NO:30 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.6HG0611290.

[0044] SEQ ID NO:31 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.6HG0611290.

[0045] SEQ ID NO:32 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.7HG0640970.

[0046] SEQ ID NO:33 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.7HG0640970.

[0047] SEQ ID NO:34 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.7HG0640970.

[0048] SEQ ID NO:35 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.7HG0640970.

[0049] SEQ ID NO:36 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.2HG0133680.

[0050] SEQ ID NO:37 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.2HG0133680.

[0051] SEQ ID NO:38 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.2HG0133680.

[0052] SEQ ID NO:39 is a polynucleotide sequence encoding a guide RNA targeting the barley gene HORVU.MOREX.r3.2HG0133680.

[0053] SEQ ID NO:40 is a polynucleotide sequence encoding an N-terminal nuclear localization signal.

[0054] SEQ ID NO:41 is the amino acid sequence of the N-terminal nuclear localization signal encoded by SEQ ID NO:40.

[0055] SEQ ID NO: 42 is a polynucleotide sequence encoding a C-terminal nuclear localization signal, codon-optimized for expression in rice.

[0056] SEQ ID NO:43 is the amino acid sequence of the C-terminal nuclear localization signal encoded by SEQ ID NOs:42, 44 and 45.

[0057] SEQ ID NO:44 is a polynucleotide sequence encoding a C-terminal nuclear localization signal, codon-optimized for expression in Homo sapiens.

[0058] SEQ ID NO:45 is a polynucleotide sequence encoding a C-terminal nuclear localization signal, codon-optimized for expression in Arabidopsis.

[0059] SEQ ID NO: 46 is the amino acid sequence of a wild-type Lachnospiraceae Cas12a protein.

[0060] SEQ ID NO: 47 is the DNMT1 guide RNA sequence.

[0061] SEQ ID NO: 48 is the EMX1 guide RNA sequence.

[0062] SEQ ID NO:49 is the FANCF guide RNA sequence.

[0063] SEQ ID NO:50 is the 3 bp deletion allele in the HORVU.MOREX.r3.1HG0069960 gene.

[0064] SEQ ID NO:51 is the 10 bp deletion allele in the HORVU.MOREX.r3.1HG0069960 gene.

[0065] SEQ ID NO:52 is the unedited allele in the HORVU.MOREX.r3.1HG0069960 gene.

[0066] SEQ ID NO: 53 is the sequence of the HORVU.MOREX.r3.1HG0069960 gene in the T1 heterozygote.

[0067] SEQ ID NO:54 is the unedited allele in the Bo2g016480 gene.

[0068] SEQ ID NO:55 is a 3 bp deletion allele in the Bo2g016480 gene.

[0069] SEQ ID NO:56 is a 9 bp deletion allele in the Bo2g016480 gene.

[0070] SEQ ID NO:57 is a 12 bp deletion allele in the Bo2g016480 gene.

[0071] SEQ ID NO: 58 is a polynucleotide sequence encoding a Cas12a variant that is codon-optimized for expression in rice and contains 12 introns (OsCas12a+12 introns).

[0072] The clustered regularly interspaced short palindromic repeats (CRISPR)-Cas9 system represents the most widely used genome editing platform for targeted genome modification in plants. For genome editing applications, the CRISPR / Cas9 system consists of two essential components: the Cas9 effector protein, which induces blunt-ended (i.e., both DNA strands are of equal length) double-strand breaks (DSBs), and a single guide RNA (sgRNA) containing a targeting sequence of approximately 20 nt. DSBs are primarily repaired by either the non-homologous end joining (NHEJ) pathway or the homology-directed repair (HDR) pathway. Loss-of-function mutations are generated by short indels introduced during the NHEJ-mediated repair pathway, whereas specific sequence modifications can be achieved by the HDR pathway in the presence of an appropriate repair template, albeit with much lower efficiency.

[0073] Although the CRISPR-Cas9 system remains the most common plant genome editing tool, the Lachnospiraceae bacterial CRISPR-Cas12a (LbCas12a) nuclease (originally identified as Cpf1) has also been shown to be capable of targeted genome modification in plants. LbCas12a differs in its requirements and outcomes compared to Streptococcus pyogenes Cas9 (SpCas9). First, LbCas12a has a "TTTV" PAM sequence requirement that makes it useful in AT-rich regions, whereas SpCas9 requires "NGG" that makes it useful in GC-rich sequences. Second, SpCas9 typically results in indels of about 1-3 bp, whereas LbCas12a usually results in deletions of about 3-12 bp. Third, SpCas9 cleaves at the PAM-proximal end of the target resulting in blunt ends, whereas LbCas12a cleaves at the PAM-distal region resulting in sticky ends (i.e., one strand is longer than the other). The distinct PAM requirement, mutational profile, and DNA strand structure at the cleavage site of LbCas12a all represent potential advantages in the field of precision genome editing and engineering in plants.

[0074] However, editing using SpCas9 and LbCas12a nucleases is not interchangeable. Modifications shown to increase Cas9 editing efficiency do not necessarily increase efficiency when the corresponding modifications are made to Cas12a. Furthermore, the current efficiency of editing using LbCas12a in various plant species, such as barley, B. oleracea, wheat and maize, is still extremely low (e.g., less than 10%). Therefore, there is still a need to discover and develop new strategies to increase the efficiency of precise genome editing.

[0075] The present disclosure overcomes the limitations of the prior art by providing engineered Cas12a proteins and novel recombinant DNA molecules encoding them, as well as compositions and methods of using them. The novel Cas12a variants are proteins with nuclease activity in plant cells. The novel Cas12a variants, when used in combination with various guide RNA structures, result in a significant increase in editing efficiency in plants compared to control Cas12a proteins. One or more guide RNAs can be utilized. Guide RNAs known in the art (see, for example, Wang, 2021) can be selected by testing mutagenesis of the target gene. Transgenic plants expressing the novel Cas12a sequences demonstrate improved genome editing efficiency for application in plant species that are widely known to show low editing efficiency using CRISPR-Cas9 as well as Cas12a editing technology. Thus, methods and compositions are provided herein for targeted genome editing in plants, which can be used to achieve beneficial results, including, for example, improved reliability of producing edited plants, significantly increasing the number of edited T0 plants, increasing the number of homozygous T0 plants for targeted editing, or combinations thereof.Furthermore, the ability to produce these desirable characteristics in T0 plants with high efficiency provides unique advantages that are not available in other methods in the art.

[0076] To generate such plants, the present disclosure provides, in certain embodiments, methods and compositions for effecting targeted genome modification via the novel Cas12a sequences described herein. For example, recombinant DNA molecules comprising polynucleotide sequences encoding Cas12a proteins in combination with one or more guide RNAs were used to edit plant genomes as disclosed herein. For example, exemplary genes from two plant species known to show low editing efficiency, namely barley and B. oleracea, were targeted for mutagenesis. T0 plants transformed with the novel Cas12a sequences were selected and evaluated for editing efficiency and fidelity. It was shown that edited alleles in target genes can be generated with significantly higher efficiency compared to currently available methods. Both homozygous and heterozygous T0 plants for the edited alleles were generated, and inheritance of the edited alleles was further identified in progeny plants (T1 plants). As described herein, the novel Cas12a sequences using various gRNA structures showed a significant increase in editing efficiency in plant species known to show low editing efficiency using CRISPR-Cas genome editing technology. Thus, the present disclosure represents a significant advance in the art in that it enables the creation of engineered alleles in plants at high frequency.

[0077] I. Engineered Proteins and Recombinant DNA Molecules Provided herein are novel engineered proteins and recombinant DNA molecules encoding them. As used herein, a "Cas12a sequence", "Cas12a variant", or a protein with "nuclease activity" refers to a protein, specifically a Cas12a nuclease. As used herein, the term "engineered" refers to a non-natural DNA, protein, cell, or organism that is not normally found in nature and is created by human intervention. An "engineered protein", "engineered enzyme" or "engineered nuclease" refers to a protein, enzyme, or Cas12a nuclease whose amino acid sequence is devised and created in the laboratory using one or more of the following biotechnology, protein design, or protein engineering techniques: molecular biology, protein biochemistry, bacterial transformation, plant transformation, site-directed mutagenesis, directed evolution using random mutagenesis, genome editing, gene editing, gene cloning, DNA ligation, DNA synthesis, protein synthesis, and DNA shuffling. For example, an engineered protein may have one or more deletions, insertions or substitutions compared to the coding sequence of a wild-type protein, and each deletion, insertion or substitution may consist of one or more amino acids. Genetic engineering can be used to generate DNA molecules that code for engineered proteins, such as engineered Cas12a proteins or Cas12a variants, that contain at least a first amino acid substitution compared to the wild-type Cas12a protein described herein.

[0078] An example of an engineered protein provided herein is an RNA-guided Cas12a nuclease (referred to herein as a "Cas12a protein" or "Cas12a variant") that comprises at least 70% sequence identity to the amino acid sequence of SEQ ID NO: 46, wherein the protein comprises at least one amino acid substitution compared to SEQ ID NO: 46. For example, the protein comprises an arginine (R) at a position corresponding to position 156 of SEQ ID NO: 46. In certain embodiments, the engineered protein provided herein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more substitutions.

[0079] The engineered protein is an enzyme with nuclease activity. As used herein, "nuclease activity" refers to the ability of a protein to introduce a double-strand break (DSB) or a single-strand nick in the nucleic acid backbone of a polynucleotide sequence and / or its complementary DNA strand in a plant genome. Examples of proteins with nuclease activity include RNA-guided nucleases, such as Cas12a. The enzymatic activity of an RNA-guided nuclease can be measured by any means known in the art, for example, by sequencing genomic DNA in the target region of the RNA-guided nuclease after expression of said nuclease and at least the gRNA in the plant cell. In particular, the RNA-guided nuclease activity can be identified based on the occurrence of approximately 1-3 bp or 3-12 bp deletions in the target genomic region.

[0080] The present disclosure provides a polynucleotide sequence encoding a protein having nuclease activity comprising at least 70% sequence identity to the amino acid sequence of SEQ ID NO: 46, wherein the encoded protein comprises at least one amino acid substitution compared to SEQ ID NO: 46. For example, the encoded protein comprises an arginine (R) at a position corresponding to position 156 of SEQ ID NO: 46. In certain embodiments, the engineered proteins provided herein comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more substitutions. Furthermore, the present disclosure provides a polynucleotide sequence encoding a protein having nuclease activity comprising at least 85% sequence identity to the polynucleotide sequence of SEQ ID NO: 46, wherein the protein encoded by said polynucleotide sequence comprises a modification at amino acid position 156 compared to a protein comprising the amino acid sequence of SEQ ID NO: 46. For example, the protein comprises an arginine (R) at a position corresponding to position 156 of SEQ ID NO: 46. The present disclosure also provides a polynucleotide sequence encoding a protein having nuclease activity, comprising at least 70% sequence identity to the amino acid sequence of SEQ ID NO: 46, wherein the polynucleotide sequence further comprises at least one intron sequence of any of SEQ ID NOs: 10-17. In some examples, the polynucleotide of the present disclosure comprises at least one intron taken from an Arabidopsis gene. The splicing efficiency of introns from Arabidopsis genes can be assessed for inclusion in polynucleotides of the present disclosure using bioinformatics methods such as the Netgene splicing tool (Hebsgaard, 1996) or by in vitro or in vivo assays, and one or more introns can be selected for inclusion in polynucleotides of the present disclosure based on such methods. Methods for identifying introns in Arabidopsis have been described (see, e.g., Cheng, 2018).In certain embodiments, the polynucleotide sequence encoding a protein having nuclease activity comprising at least 70% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NO: 46 comprises an arginine (R) at a position corresponding to position 156 of SEQ ID NO: 46, and the polynucleotide sequence further comprises at least one intron sequence of a plant, such as Arabidopsis, or of any of SEQ ID NOs: 10-17, or a combination thereof.

[0081] As used herein, the term "protein-encoding DNA molecule" or "protein-encoding sequence" refers to a DNA molecule that comprises a DNA sequence that encodes a protein. As used herein, the term "protein" refers to a chain of amino acids linked by peptide (amide) bonds and includes both polypeptide chains that are folded or arranged in a biologically functional manner and those that are not. As used herein, "protein-encoding sequence" refers to a DNA sequence that encodes a protein. As used herein, "sequence" refers to a contiguous arrangement of nucleotides or amino acids. A "DNA sequence" may refer to a sequence of nucleotides or a DNA molecule that is composed of a sequence of nucleotides. A "protein sequence" may refer to a sequence of amino acids or a protein that comprises a sequence of amino acids. The boundaries of a protein-encoding sequence are usually determined by a translation start codon at the 5'-terminus and a translation stop codon at the 3'-terminus.

[0082] Engineered proteins are proteins in which the wild-type protein sequence has been altered or modified to provide modified characteristic(s) or useful protein characteristics, such as altered Vmax, Km, Ki, IC, among others. 50Modifications can be made at specific amino acid positions in the protein, or by substituting alternative amino acids in place of the typical amino acids found naturally at the same positions (i.e., in the wild-type protein). Amino acid modifications can be made as single amino acid substitutions in the protein sequence, or in combination with one or more other modifications, such as one or more other amino acid substitutions, deletions or additions. In some embodiments, the engineered protein has altered protein properties, e.g., in the presence of one or more gRNA sequences, resulting in increased editing efficiency compared to the wild-type protein in the presence of the same gRNA sequence. Thus, in other embodiments, the present disclosure provides engineered proteins, such as Cas12a variants, and recombinant DNA molecules encoding same, with one or more amino acid substitutions, e.g., D156R, and the position of the amino acid substitution(s) is compared to the amino acid position(s) shown in SEQ ID NO:46. In certain embodiments, the engineered proteins provided herein contain one, two, three, four, five, six, seven, eight, nine, ten or more of any combination of such substitutions, where the modifications are made at positions relative to the functionally equivalent positions of the amino acid sequence provided as SEQ ID NO: 46. Similar modifications can be made by alignment of the amino acid sequence of the RNA-guided nuclease to be mutated with the amino acid sequence of a RNA-guided nuclease of interest having nuclease activity, e.g., Cas12a, at analogous positions of any RNA-guided nuclease.

[0083] The DNA molecules disclosed herein or fragments thereof can be isolated and manipulated using a number of methods well known to those skilled in the art. For example, polymerase chain reaction (PCR) technology can be used to amplify a particular starting DNA molecule or to create variants of the original molecule. DNA molecules or fragments thereof can also be obtained by other techniques, for example by directly synthesizing the fragments by chemical means, as is commonly practiced by using automated oligonucleotide synthesizers.

[0084] Due to the degeneracy of the genetic code, a variety of different DNA sequences can code for proteins, such as the modified or engineered proteins disclosed herein. For example, FIG. 8 provides a universal genetic code chart showing all possible mRNA triplet codons (T in DNA molecules is replaced by U in RNA molecules), and the amino acids coded for by each codon. DNA sequences coding for Cas12a proteins with amino acid substitutions as described herein can be generated by introducing mutations into DNA sequences coding for wild-type Cas12a proteins using methods known in the art and the information provided in FIG. 8. It is well within the capabilities of one of ordinary skill in the art to generate alternative DNA sequences coding for the same or essentially the same modified or engineered proteins as described herein. These variants or alternative DNA sequences are within the scope of the embodiments described herein. As used herein, reference to "essentially the same" sequences refers to sequences that code for amino acid substitutions, deletions, additions, or insertions that do not substantially alter (i.e., alter function of) the functional activity of the protein encoded by the DNA molecules of the embodiments described herein. Allelic variants of nucleotide sequences coding for wild-type or engineered proteins are also encompassed within the scope of the embodiments described herein. While maintaining the functional activity of the protein encoded by the DNA molecule, such allelic variants can have beneficial effects when expressed in certain plant cells.For example, the results described herein demonstrate that the Cas12a protein and its variants codon-optimized for distantly related plant species or species of separate biological kingdoms surprisingly lead to increased genome editing efficiency in plant species known to be resistant to CRISPR-Cas genome editing, such as barley, B. oleracea, wheat and maize.

[0085] Amino acid substitutions other than those specifically exemplified or naturally occurring in wild-type or engineered Cas12a proteins are contemplated within the scope of the embodiments described herein, as long as the Cas12a protein having the substitution still retains substantially the same functional activity as described herein. These variants or alternative DNA sequences in combination with such amino acid substitutions in the protein encoded by the DNA sequence are also encompassed within the scope of the embodiments described herein, including but not limited to SEQ ID NOs: 1, 3, 5, 7 and 8. Similarly, variants or alternative DNA sequences encoding Cas12a proteins with nuclease activity that further comprise a heterologous intron sequence are also encompassed within the scope of the embodiments described herein. Introns do not contain information encoding a protein or polypeptide. Introns are initially transcribed into an RNA sequence, but are subsequently spliced ​​out of the mature RNA molecule. While maintaining the functional activity of the protein encoded by the DNA molecule further comprising a heterologous intron sequence, such allelic variants containing an intron sequence may have beneficial effects when expressed in certain plant cells.

[0086] For example, the results described herein demonstrate that Cas12a proteins and variants thereof containing at least one intron sequence of any of SEQ ID NOs: 10-17 resulted in increased genome editing efficiency in plant species known to exhibit low editing efficiency using CRISPR-Cas genome editing technology, such as barley, B. oleracea, wheat and maize.

[0087] Polynucleotide sequences encoding Cas12a nucleases provided herein include polynucleotide sequences that include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or more intron sequences. Intron sequences that may be inserted into polynucleotide sequences encoding Cas12a nucleases include, but are not limited to, any of SEQ ID NOs: 10-17, or multiple copies thereof. According to the present disclosure, one or more introns may be inserted at any position within a sequence encoding Cas12a nucleases, such as any of SEQ ID NOs: 1, 3, 5, 7, and 8. Experiments can be performed in which the combined effect of the D156R mutation and the inclusion of one or more introns can be measured (e.g., comparing only the first intron in Cas12a to having any other intron or all eight introns). Other experiments can determine which portion of Cas12a contains an intron that results in improved editing efficiency.

[0088] Recombinant DNA molecules provided herein may be fully or partially synthesized and modified by methods known in the art, where it is desirable to provide sequences useful for DNA manipulation (such as restriction enzyme recognition sites or recombination-based cloning sites), plant-preferred sequences (such as plant-codon usage or Kozak consensus sequences), or sequences useful for DNA construct design (such as spacer or linker sequences). The disclosure includes recombinant DNA molecules and engineered proteins having at least 50% sequence identity, at least 60% sequence identity, at least 70% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 91% sequence identity, at least 92% sequence identity, at least 93% sequence identity, at least 94% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity, and at least 99% sequence identity to any of the recombinant DNA molecules or amino acid sequences provided herein and have nuclease activity. As used herein, the term "percent sequence identity" or "% sequence identity" refers to the percentage of identical nucleotides or amino acids in a linear polynucleotide or amino acid sequence of a reference ("query") sequence (or its complement) compared to a test ("subject") sequence (or its complement) when the two sequences are optimally aligned (insertions, deletions or gaps of appropriate nucleotides or amino acids total less than 20% of the reference sequence over the comparison window).Optimal alignment of sequences to align a comparison window is well known to those of skill in the art and can be performed by tools such as the Smith and Waterman local homology algorithm, the Needleman and Wunsch homology alignment algorithm, the Pearson and Lipman similarity search method, and computer implementations of these algorithms, such as GAP, BESTFIT, FASTA, and TFASTA available as part of the sequence analysis software of the GCG® Wisconsin Package® (Accelrys Inc., San Diego, CA), MEGAlign (DNAStar Inc., 1228 S. Park St., Madison, WI 53715), and MUSCLE (version 3.6) using, for example, default parameters (RC Edgar, "MUSCLE: multiple sequence alignment with high accuracy and high throughput", Nucleic Acids Research 32(5):1792-7 (2004)). The "percent identity" of the aligned segments of the test sequence and the reference sequence is the number of identical elements shared by the two aligned sequences divided by the total number of elements in the portion of the aligned reference sequence segment, i.e., the entire reference sequence or a smaller defined portion of the reference sequence. The percent sequence identity is expressed as the percent identity multiplied by 100. The comparison of one or more sequences can be to full-length sequences or portions thereof, or to longer sequences.

[0089] II. Genome editing The present disclosure provides, in certain embodiments, plants, plant parts, plant cells and seeds that are produced by genome modification using site-specific integration or genome editing. Genome editing can be used to make one or more edits or one or more mutations at a desired target site in the genome of a plant, for example, to change the expression and / or activity of one or more genes, or to integrate an insertion sequence or transgene into a desired location in the genome of a plant. Any site or locus in the genome of a plant can potentially be selected for genome editing (or gene editing) or site-specific integration of a transgene, construct or transcribable DNA sequence. As used herein, a "target site" for genome editing or site-specific integration refers to the location of a polynucleotide sequence in the genome of a plant that is bound and cleaved by a site-specific nuclease to introduce a double-strand break (DSB) or a single-strand nick in the nucleic acid backbone of the polynucleotide sequence and / or its complementary DNA strand in the genome of the plant. The target site may comprise, for example, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 29, or at least 30 consecutive nucleotides. The "target site" of the RNA-guided nuclease may comprise either a complementary strand of a double-stranded nucleic acid (DNA) molecule or a chromosome sequence at the target site. The site-specific nuclease may bind to the target site, for example, via a non-coding guide RNA (such as, but not limited to, a CRISPR RNA (crRNA) or a single guide RNA (sgRNA) as further described herein). The non-coding guide RNA provided herein may be complementary to the target site (e.g., complementary to either strand of a double-stranded nucleic acid molecule or a chromosome at the target site). It will be understood that complete identity or complementarity may not be required for a non-coding guide RNA to bind or hybridize to a target site.For example, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 mismatches (or more) between the target site and the non-coding RNA can be tolerated. "Target site" also refers to the location of a polynucleotide sequence in a plant genome that is bound and cleaved by any other site-specific nuclease that cannot be induced by a non-coding RNA molecule, such as zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), meganuclease, etc., to introduce a DSB or single-stranded nick in the polynucleotide sequence and / or its complementary DNA strand. As used herein, "target region" or "targeted region" refers to a polynucleotide sequence or region adjacent to two or more target sites. Without limitation, in some embodiments, the target region can be subject to mutation, deletion, insertion, substitution, inversion, or duplication. As used herein, "adjacent," when used to describe a target region of a polynucleotide sequence or molecule, refers to two or more target sites of the polynucleotide sequence or molecule surrounding the target region, with one target site on each side of the target region.

[0090] As used herein, "targeted genome editing technology" refers to any method, protocol, or technology that uses site-specific nucleases, such as meganucleases, zinc finger nucleases (ZFNs), RNA-guided endonucleases (e.g., CRISPR / Cas9 or Cas12a systems), TALE (transcription activator-like effector) endonucleases (TALENs), recombinases, or transposases to enable precise and / or targeted editing (i.e., editing is mostly or completely non-random) of specific locations in the genome of a plant. In certain embodiments, "targeted genome editing technology" refers to the RNA-guided Cas12a system. As used herein, "editing" or "genomic editing" refers to making targeted mutations, deletions, insertions, substitutions, inversions or duplications of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1000, at least 2500, at least 5000, at least 10,000, or at least 25,000 nucleotides of an endogenous plant genomic nucleic acid sequence. As used herein, "editing" or "genome editing" can also encompass targeted insertion or site-specific integration of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 10,000, or at least 25,000 nucleotides into the endogenous genome of a plant.The singular "edit" or "genome edit" refers to one such targeted mutation, deletion, insertion, substitution, inversion or duplication, and the singular "edit" or "genome edit" refers to two or more targeted mutations, deletions, insertions, substitutions, inversions and / or duplications, each "edit" being introduced via a targeted genome editing technique.

[0091] According to some embodiments, the site-specific nuclease is co-delivered with a donor template molecule to serve as a template for making the desired edits, mutations, or insertions into the genome at the desired target site via repair of a double-strand break (DSB) or a nick generated by the site-specific nuclease. According to some embodiments, the site-specific nuclease may be co-delivered with a DNA molecule that includes a selectable or screenable marker gene.

[0092] The site-specific nuclease may be an RNA-guided nuclease. According to some embodiments, the RNA-guided endonuclease is selected from the group consisting of Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Csel, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, The RNA-guided endonuclease may be selected from the group consisting of Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, CasX, CasY, and any homologs or modified versions thereof, and Argonaute proteins (non-limiting examples of Argonaute proteins include Thermus thermophilus Argonaute (TtAgo), Pyrococcus furiosus Argonaute (PfAgo), Natronobacterium gregoryi Argonaute (NgAgo), and any homologs or modified versions thereof). According to some embodiments, the RNA-guided endonuclease is a Cas9 or Cpf1 (also referred to herein as Cas12a) enzyme. Further, in some embodiments, the RNA-guided endonuclease is a Cas12a enzyme or variant. In certain embodiments, the RNA-guided endonuclease is a Lachnospiraceae Cas12a (LbCas12a) variant encoded by a sequence having at least 85% identity with any of SEQ ID NOs: 1, 3, 5, 7, and 8. The RNA-guided nuclease can be delivered as a protein with or without a guide RNA, or the guide RNA can be complexed with the RNA-guided nuclease enzyme and delivered as a ribonucleoprotein (RNP).

[0093] In the case of RNA-guided endonucleases, a guide RNA molecule is further provided to direct the endonuclease to a target site in the genome of the plant through base pairing or hybridization to cause a DSB or nick at or near the target site. As described herein, the guide RNA can be transformed or introduced into the plant cell or tissue as a gRNA molecule, or as a recombinant DNA molecule, a construct or vector that includes a transcribable DNA sequence that encodes one or more guide RNAs operably linked to a single promoter or individual promoters. As understood in the art, the guide RNA can include, for example, CRISPR RNA (crRNA), single-stranded guide RNA (sgRNA), or any other RNA molecule that can guide or guide the endonuclease to a specific target site in the genome. The prototypical CRISPR-associated protein, Cas9 from Streptococcus pyogenes, naturally binds to two RNAs, CRISPR RNA (crRNA) guide and trans-acting CRISPR RNA (tracrRNA) CRISPR, to assemble a ribonucleoprotein (crRNP). In comparison, the CRISPR-Cas12a system does not require trans-activating crispr RNA (tracrRNA) for the biogenesis of mature crRNA. Instead, the RuvC endonuclease domain of Cas12a directly processes the mature crRNA. A "single-stranded guide RNA" (or "sgRNA") is an RNA molecule that includes a crRNA covalently linked to a tracrRNA by a linker sequence, which can be expressed as a single RNA transcript or molecule. A guide RNA includes a guide or targeting sequence (also referred to herein as a "spacer sequence") that is identical or complementary to a target site in a plant genome, for example, at or near a gene. A guide RNA is typically a non-coding RNA molecule that does not code for a protein.The guide sequence of the guide RNA can be at least 10 nucleotides in length, e.g., 12-40 nucleotides, 12-30 nucleotides, 12-20 nucleotides, 12-35 nucleotides, 12-30 nucleotides, 15-30 nucleotides, 17-30 nucleotides, or 17-25 nucleotides in length, or about 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more nucleotides in length. The guide sequence can be at least 95%, at least 96%, at least 97%, at least 99%, or 100% identical to or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or more contiguous nucleotides of the DNA sequence at the genomic target site.

[0094] As mentioned above, the target gene for genome editing can be any plant gene of interest. In the knockdown mutation of a gene of interest via genome editing, the RNA-guided endonuclease can be targeted to the upstream or downstream sequence of the gene, such as the promoter and / or enhancer sequence, or the intron, 5'UTR, and / or 3'UTR sequence, to mutate one or more promoters and / or regulatory sequences of the gene to affect or reduce its expression level. Similarly, in the mutation of a gene of interest via genome editing, the RNA-guided endonuclease can be targeted to the transcribable DNA sequence (i.e., the transcribable region) of said gene, such as the region of the gene including the coding sequence, the specific DNA sequence encoding a protein domain, the exon region, the intron region, or a combination thereof. For example, in certain embodiments, the transcribable DNA sequence targeted for genome editing can include or be adjacent to an exon / intron boundary. If the resulting modification spans the exon / intron boundary, the modification can be referred to as a modification in the exon region and the intron region. In gene repair of a gene of interest, a guide RNA can be used, which comprises a guide sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 99% or 100% identical or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25 or more consecutive nucleotides of said gene or the sequence complementary thereto, but alternative splicing and various exon / intron boundaries may occur.As used herein, the term "contiguous" in relation to a polynucleotide or protein sequence means that there are no deletions or gaps in the sequence.

[0095] As used herein, for a given sequence, "complement," "complementary sequence," and "reverse complement" are used interchangeably. All three terms refer to the reverse complement of a nucleotide sequence, i.e., the sequence that is complementary to the given sequence in reverse order of the nucleotides.

[0096] "Ribosome binding site" or "ribosome binding site (RBS)" refers to a sequence of nucleotides upstream of the start codon of an mRNA transcript that is responsible for the recruitment of ribosomes during translation initiation. Generally, RBS refers to a bacterial sequence, but internal ribosome entry sites (IRES) have been described in mRNAs of eukaryotic cells, or viruses that infect eukaryotic organisms. Ribosome recruitment in eukaryotes is generally mediated by the 5' cap present on eukaryotic mRNAs. Ribosome skipping sequences (e.g., 2A sequences such as furin-GSG-T2A) can be used in the construct to prevent covalent ligation of translated amino acid sequences.

[0097] Alternative guide structures for tRNA can also be used, incorporating tRNA sequences instead of ribozymes. One or more tRNAs can be used.

[0098] As used herein, the term "antisense" refers to a DNA or RNA sequence that is complementary to a specific DNA or RNA sequence. An antisense RNA molecule is a single-stranded nucleic acid that can bind to a sense RNA strand or sequence or mRNA to form a duplex by sequence complementarity. The term "antisense strand" refers to a nucleic acid strand that is complementary to the "sense" strand. The "sense strand" of a gene or locus is a strand of DNA or RNA that has the same sequence as the RNA molecule transcribed from the gene or locus (except for uracil in RNA and thymine in DNA).

[0099] A protospacer adjacent motif (PAM) may be present in the genome immediately adjacent to and upstream of the 5' end of the genomic target site sequence that is complementary to the targeting sequence of the guide RNA, i.e., immediately downstream (3') of the sense (+) strand of the genomic target site (relative to the targeting sequence of the guide RNA) as known in the art. See, for example, Wu et al. (Quant Biol. 2(2):59-70, 2014). The genomic PAM sequence on the sense (+) strand adjacent to the target site (relative to the targeting sequence of the guide RNA) may include 5'-NGG-3' for Cas9, or 5'-TTTN-3' for Cas12a. However, the corresponding sequence of the guide RNA (i.e., immediately downstream (3') of the targeting sequence of the guide RNA) may not generally be complementary to the genomic PAM sequence.

[0100] As used herein, a "donor molecule," "donor template," or "donor template molecule" (collectively "donor template") may be a recombinant polynucleotide, DNA or RNA donor template or sequence, and is defined as a nucleic acid molecule having a homologous nucleic acid template or sequence (e.g., a homologous sequence) and / or an insertion sequence for site-specific targeted insertion or recombination into the genome of a plant cell via repair of a nick or DSB in the genome of the plant cell. A donor template may be a separate DNA molecule that contains one or more homologous sequences and / or an insertion sequence for targeted integration, or a donor template may be a sequence portion (i.e., a donor template region) of a DNA molecule that further contains one or more other expression cassettes, genes / transgenes, and / or transcribable DNA sequences. For example, a "donor template" may be used as a template for site-specific integration of a transgene or construct, or to introduce a mutation, e.g., an insertion, deletion, substitution, etc., into a target site in the genome of a plant. The targeted genome editing techniques provided herein may include the use of one or more, two or more, three or more, four or more, or five or more donor molecules or templates. The donor templates provided herein can include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 genes or transgenes and / or transcribable DNA sequences. Alternatively, the donor templates can include no genes, transgenes, or transcribable DNA sequences.

[0101] By way of non-limiting example, the gene / transgene or transcribable DNA sequence of the donor template may include, for example, an insecticide resistance gene, a herbicide resistance gene, a nitrogen use efficiency gene, a water use efficiency gene, a yield enhancing gene, a nutritional quality gene, a DNA binding gene, a selectable marker gene, an RNAi or suppression construct, a site-specific genome modification enzyme gene, a single guide RNA of the CRISPR / Cas9 system, a geminivirus-based expression cassette, or a plant virus expression vector system. According to other embodiments, the insert sequence of the donor template may include a transcribable DNA sequence that encodes a protein-coding sequence or a non-coding RNA molecule that can target an endogenous gene for suppression. The donor template may include a promoter, such as a constitutive promoter, a tissue-specific or tissue-preferential promoter, a developmental stage promoter, or an inducible promoter, operably linked to the coding sequence, gene, or transcribable DNA sequence. The donor template may include a leader, enhancer, promoter, transcription initiation site, 5'-UTR, one or more exons, one or more introns, transcription termination site, region or sequence, 3'-UTR, and / or polyadenylation signal, each of which may be operably linked to a coding sequence, gene (or transgene) or transcribable DNA sequence encoding a non-coding RNA, guide RNA, mRNA and / or protein. The donor template may be a single-stranded or double-stranded DNA or RNA molecule or a plasmid.

[0102] The "insertion sequence" of the donor template is a sequence designed for targeted insertion into the genome of a plant cell, which may be of any suitable length. For example, the length of the insertion sequence of the donor template may be 2-50,000, 2-10,000, 2-5000, 2-1000, 2-500, 2-250, 2-100, 2-50, 2-30, 15-50, 15-100, 15-500, 15-1000, 15-5000, 18-30, 18-26, 20-26, 20-50, 20-100, 20-250, 20-500, 20 The donor template may be 1000, 20-5000, 20-10,000, 50-250, 50-500, 50-1000, 50-5000, 50-10,000, 100-250, 100-500, 100-1000, 100-5000, 100-10,000, 250-500, 250-1000, 250-5000, or 250-10,000 nucleotides or base pairs. The donor template may also have at least one homologous sequence or arm, such as two homologous arms, to introduce a mutation or integration of an insertion sequence via homologous recombination to a target site in the genome of the plant, where the homologous sequence or arm(s) is identical or complementary, or has a percent identity or percent complementarity, to a sequence at or near the target site in the genome of the plant. When the donor template comprises a homology arm(s) and an insertion sequence, the homology arm(s) flank or surround the insertion sequence of the donor template. Each homology arm can be at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 99% or 100% identical or complementary to at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 500, at least 1000, at least 2500, or at least 5000 consecutive nucleotides of the target DNA sequence in the genome of the plant.

[0103] Any method known in the art for site-specific integration can be used with the present disclosure.In the presence of a donor template molecule with an insertion sequence, the DSB or nick can be repaired by homologous recombination between the homologous arm(s) of the donor template and the plant genome, or by non-homologous end joining (NHEJ), resulting in the site-specific integration of the insertion sequence into the plant genome, resulting in a targeted insertion event at the site of the DSB or nick.Therefore, when a transgene, a transducible DNA sequence, a construct or a sequence is located in the insertion sequence of the donor template, the site-specific insertion or integration of the transgene, a transducible DNA sequence, a construct or a sequence can be achieved.

[0104] The introduction of DSBs or nicks can also be used to introduce targeted mutations into the genome of plants. According to this approach, mutations such as deletions, insertions, substitutions, inversions and / or duplications can be introduced into target sites via incomplete repair of DSBs or nicks to result in genetic modifications within genes. Such mutations can be generated by incomplete repair of target loci without the use of donor template molecules. Gene modification can be achieved by inducing DSBs or nicks at or near the endogenous locus of a gene, which results in the expression of a non-functional protein, an interfering protein, or a protein with reduced, disrupted, or altered activity compared to a protein expressed from a gene lacking said modification.

[0105] Similarly, such targeted mutations of genes can be generated using donor template molecules to induce specific or desired mutations at or near the target site via repair of DSBs or nicks. The donor template molecule can contain a homologous sequence with or without an insertion sequence, including one or more mutations, such as one or more deletions, insertions, substitutions, inversions and / or duplications, compared to the target genomic sequence at or near the site of the DSB or nick. For example, targeted mutations of genes can be achieved by deleting, inserting, substituting, inverting or duplicating at least a portion of the gene, such as by introducing a frameshift or premature stop codon into the coding sequence of the gene, or by introducing a modification into a transcribable DNA sequence. Deletion of a portion of a gene can also be introduced by generating DSBs or nicks at two target sites and deleting the intervening target region adjacent to the target sites. Modification of a target gene can result in the expression of a non-functional protein, an interfering protein, or a protein with reduced, disrupted or altered activity compared to a protein expressed from a gene lacking said modification.

[0106] In one aspect, the disclosure provides a plant, or a plant seed, plant part or plant cell thereof, comprising a recombinant DNA molecule, the recombinant DNA molecule comprising a sequence having at least 85% identity to any of SEQ ID NOs: 1, 3, 5, 7, and 8; a sequence comprising any of SEQ ID NOs: 1, 3, 5, 7, and 8; a fragment of any of SEQ ID NOs: 1, 3, 5, 7, and 8; or a sequence encoding a protein having at least 85% identity to any of SEQ ID NOs: 2, 4, 6, and 9. In certain embodiments, the protein encoded by the recombinant DNA molecule comprises (i) a modification at amino acid position 156 compared to a protein comprising the amino acid sequence of SEQ ID NO: 46, and (ii) further comprises one or more intronic sequences of SEQ ID NOs: 10-17, or a combination thereof. When expressed in a plant cell in the presence of one or more guide RNA molecules, the protein encoded by the recombinant DNA molecule described herein may provide a high efficiency of genomic modification within a target region defined by the gRNA(s) compared to a control protein, e.g., compared to a protein comprising the amino acid sequence of SEQ ID NO: 46. The genomic modification can be a deletion of a region comprising at least 1, at least 2, at least 4, at least 6, at least 8, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, or at least 150 consecutive nucleotides within the target region. In one embodiment, the genomic modification can also include a deletion and a nucleotide substitution or insertion of at least 1, at least 2, at least 4, at least 6, at least 8, at least 10, or at least 20 consecutive nucleotides surrounding the deletion.

[0107] In one embodiment, mutant alleles of a gene of interest can comprise two or more modifications in the transcribable region of endogenous gene.The present disclosure provides such mutant alleles, and they can be produced, for example, using a construct that comprises a sequence encoding two or more guide RNAs operably linked to a plant-expressible promoter; or a construct that comprises two gRNA cassettes, each operably linked to a plant-expressible promoter.

[0108] III. Constructs for genome editing Recombinant DNA constructs and vectors are provided that include polynucleotide sequences that code for site-specific nucleases, such as RNA-guided endonucleases, and the coding sequences are operably linked to a plant-expressible promoter.For RNA-guided endonucleases, recombinant DNA constructs and vectors are further provided that include polynucleotide sequences that code for one or more guide RNAs, and the guide RNA(s) include a guide sequence of sufficient length to have percent identity or percent complementarity with a target site in the genome of a plant, such as at or near a target gene of interest.The polynucleotide sequence of the recombinant DNA constructs and vectors that code for site-specific nucleases or guide RNA(s) can be operably linked to a plant-expressible promoter, such as an inducible promoter, a constitutive promoter, or a tissue-specific promoter.

[0109] As used herein, "gene" refers to a nucleic acid sequence that forms a genetic and functional unit and codes for one or more sequence-related RNA and / or polypeptide molecules. A gene generally comprises a coding region operably linked to an appropriate regulatory sequence that regulates the expression of a gene product (e.g., a polypeptide or functional RNA). A gene can have various sequence elements, including, but not limited to, a promoter, an untranslated region (UTR), an exon, an intron, and other upstream or downstream regulatory sequences.

[0110] As used herein, "allele" refers to an alternative nucleic acid sequence of a gene or at a particular locus (e.g., a nucleic acid sequence of a gene or locus that differs from other alleles of the same gene or locus). Such an allele can be considered (i) wild type, or (ii) mutant if one or more mutations or edits are present in the nucleic acid sequence of the mutant allele compared to the wild type allele. A mutant or edited allele for a gene can have a reduced, disrupted, altered, or eliminated activity, or a reduced or eliminated expression level of the gene, compared to the wild type allele. For example, a mutant or edited allele of a gene of interest can have a deletion in a transcribable region of an endogenous gene that reduces, disrupts, or alters the activity of a protein encoded by the mutant allele compared to the activity of a protein encoded by the wild type allele in an otherwise identical plant. In the case of a diploid organism, such as maize, a first allele can occur on one chromosome and a second allele can occur at the same locus on a second homologous chromosome. When one allele at a locus on one chromosome of a plant is a mutant or edited allele, and the other corresponding allele on the homologous chromosome of the plant is wild type, the plant is described as heterozygous for the mutant or edited allele.However, when both alleles at a locus are mutant or edited alleles, the plant is described as homozygous for the mutant or edited allele.The homozygous plant for the mutant or edited allele at a locus can contain the same mutant or edited allele, or different mutant or edited alleles, in the case of hetero alleles or two alleles.

[0111] As used herein, a "wild type gene" or "wild type allele" refers to a gene or allele having the most common sequence or genotype in a particular plant species, or another sequence or genotype that has only natural mutations, polymorphisms, or other silent mutations compared to the most common sequence or genotype that do not significantly affect the expression and activity of the gene or allele. In fact, a "wild type" gene or allele does not contain mutations, polymorphisms, or any other type of mutation that substantially affects the normal function, activity, expression, or phenotypic outcome of that gene or allele compared to the most common sequence or genotype. In general, the term "variant" refers to molecules that have some synthetic or naturally generated differences in their nucleotide or amino acid sequence compared to a reference (natural) polynucleotide or polypeptide, respectively. These differences include substitutions, insertions, deletions, inversions, duplications, or any desired combination of such changes in the natural polynucleotide or amino acid sequence.

[0112] As used herein, the term "expression" refers to the biosynthesis of a gene product, typically the transcription and / or translation of a nucleotide sequence, e.g., an endogenous gene, a heterologous gene, a transgene, or an RNA- and / or protein-coding sequence in a cell, tissue, organ or organism, such as a plant, plant part, or plant cell, tissue or organ.

[0113] The term "recombinant" with respect to a polynucleotide (DNA or RNA) molecule, protein, construct, vector, etc., refers to a polynucleotide or protein molecule or sequence that is man-made, not normally found in nature, and / or exists in a situation not normally found in nature, including a polynucleotide (DNA or RNA) molecule, protein, construct, etc., a combination of two or more polynucleotide or protein sequences that do not naturally occur together in the same manner without human intervention, e.g., a polynucleotide molecule, protein, construct, etc., that includes at least two polynucleotide or protein sequences that are operably linked but heterologous with respect to each other. For example, the term "recombinant" can refer to any combination of two or more DNA or protein sequences in the same molecule (e.g., a plasmid, construct, vector, chromosome, protein, etc.), such a combination is man-made and not normally found in nature. As used in this definition, the phrase "not normally found in nature" means not found in nature without human introduction. A recombinant polynucleotide or protein molecule, construct, etc. can include a polynucleotide or protein sequence(s) that is (i) separated from other(s) polynucleotide or protein sequence(s) that are naturally adjacent to each other, and / or (ii) adjacent to (or adjacent to) other polynucleotide or protein sequence(s) that are not naturally adjacent to each other. Such recombinant polynucleotide molecules, proteins, constructs, etc. can also refer to polynucleotide or protein molecules or sequences that have been genetically engineered and / or constructed outside of a cell. For example, a recombinant DNA molecule can include any engineered or artificial plasmid, vector, etc., and can include linear or circular DNA molecules. Such plasmids, vectors, etc. can contain various maintenance elements, including a prokaryotic origin of replication and a selectable marker, as well as one or more transgenes or expression cassettes, possibly in addition to a plant selectable marker gene, etc.The term "operably linked" refers to a functional link between a promoter or other regulatory element and an associated transcribable DNA sequence or coding sequence of a gene (or transgene) such that the promoter etc. operates or functions to initiate, support, affect, cause and / or promote the transcription and expression of the associated transcribable DNA sequence or coding sequence, at least in a particular cell(s), tissue, developmental stage and / or condition.

[0114] Reference in this application to an "isolated DNA molecule" or "isolated polynucleotide" or equivalent term or phrase is intended to mean that the DNA molecule or polynucleotide is present alone or in combination with other compositions, but is not present in its natural environment. For example, a nucleic acid element such as a coding sequence, an intron sequence, a non-translated leader sequence, a promoter sequence, a transcription termination sequence, etc., naturally found in the DNA of the genome of an organism is not considered to be "isolated" as long as the element is in the genome of the organism and in the location in the genome where it is found in nature. However, each of these elements, and subportions of these elements, would be "isolated" within the scope of this disclosure as long as the element is not in the genome of the organism and in the location in the genome where it is found in nature. Similarly, a nucleotide sequence that encodes a protein or any naturally occurring variant of that protein is an isolated nucleotide sequence, but only if the nucleotide sequence is not in the DNA of the organism in which the sequence encoding the protein is naturally found. A synthetic nucleotide sequence that encodes the amino acid sequence of a naturally occurring protein is considered to be isolated for the purposes of this disclosure. For purposes of this disclosure, any transgenic nucleotide sequence, i.e., a nucleotide sequence of DNA inserted into the genome of a plant or bacterial cell or present in an extrachromosomal vector, is considered to be an isolated nucleotide sequence, whether it is present within a plasmid or similar structure used to transform the cell, present within the genome of the plant or bacteria, or present in detectable amounts in tissues, progeny, biological samples, or commercial products derived from the plant or bacteria.

[0115] As generally understood in the art, the term "promoter" can generally refer to a DNA sequence that contains an RNA polymerase binding site, a transcription initiation site, and / or a TATA box, and that supports or facilitates the transcription and expression of an associated transcribable polynucleotide sequence and / or gene (or transgene). Promoters can be synthetically produced, altered, or derived from known or naturally occurring promoter sequences or other promoter sequences. Promoters can also include chimeric promoters that include a combination of two or more heterologous sequences. Thus, promoters of the present disclosure can include variants or fragments of promoter sequences that are similar in composition, but not identical, to other promoter sequence(s) known or provided herein. Promoters provided herein, or variants or fragments thereof, can include "minimal promoters" that provide a basal level of transcription and are composed of a TATA box or equivalent DNA sequence for recognition and binding of the RNA polymerase II complex for transcription initiation. Promoters can be classified as constitutive, developmental, tissue-specific, inducible, etc., according to various criteria regarding the expression pattern of the associated coding or transcribable sequence or gene (including transgene) operably linked to the promoter. A promoter that drives expression in all or most tissues of a plant is called a "constitutive" promoter. A promoter that drives expression during a specific period or stage of development is called a "developmental" promoter. A promoter that drives enhanced expression in a specific tissue of a plant compared to other plant tissues is called a "tissue-enhanced" or "tissue-preferred" promoter. Thus, a "tissue-preferred" promoter causes relatively high or preferential expression in a specific tissue(s) of a plant, but low expression levels in other tissue(s) of the plant. A promoter that is expressed in a specific tissue(s) of a plant, and has little or no expression in other plant tissues, is called a "tissue-specific" promoter.An "inducible" promoter is a promoter that initiates transcription in response to an environmental stimulus, such as cold, drought or light, or other stimuli, such as wounding or application of a chemical. Promoters can also be classified with respect to their origin, e.g., heterologous, homologous, chimeric, synthetic, etc.

[0116] As used herein, a "plant-expressible promoter" refers to a promoter that is capable of initiating, supporting, influencing, causing, and / or promoting the transcription and expression of its associated transcribable DNA sequence, coding sequence or gene in a plant cell or tissue.

[0117] The term "heterologous" with respect to a promoter or other regulatory sequence associated with an associated polynucleotide sequence (e.g., a transcribable DNA sequence or a coding sequence or a gene) is a promoter or regulatory sequence that is not operably linked to such associated polynucleotide sequence in nature without human introduction, e.g., the promoter or regulatory sequence has a different origin than the associated polynucleotide sequence and / or the promoter or regulatory sequence is not native to the plant species transformed with the promoter or regulatory sequence. Similarly, "heterologous" with respect to a coding sequence can refer to the use of a recombinant DNA molecule that is codon-optimized for a different organism compared to the organism in which the DNA molecule is expressed, e.g., a recombinant DNA sequence encoding Cas12a that is codon-optimized for expression in humans but expressed in a plant cell.

[0118] As used herein, "endogenous gene" or "endogenous locus" refers to a gene or locus in its natural and original chromosomal location. As used herein, in the context of a protein-coding gene, "exon" refers to a segment of a DNA or RNA molecule that contains the information that codes for a protein or polypeptide sequence.

[0119] As used herein, an "intron" of a gene refers to a segment of a DNA or RNA molecule that does not contain any information encoding a protein or polypeptide and that is initially transcribed into an RNA sequence, but is subsequently spliced ​​out of the mature RNA molecule.

[0120] As used herein, the "untranslated region (UTR)" of a gene refers to a segment or sequence of an RNA molecule (e.g., an mRNA molecule) that is expressed from a gene (or transgene), but excluding the exon and intron sequences of the RNA molecule. "Untranslated region (UTR)" also refers to a DNA segment, or a sequence that codes for such a UTR segment of an RNA molecule. An untranslated region can be a 5'-UTR or a 3'-UTR, depending on whether it is located at the 5' or 3' end of a DNA or RNA molecule or sequence relative to the coding region of the DNA or RNA molecule or sequence (i.e., upstream (5') or downstream (3') of the exon and intron sequences, respectively).

[0121] As used herein, "transcribeable region" or "transcribeable DNA sequence" refers to a nucleic acid sequence that is expressed from a gene (or transgene).

[0122] As used herein, a "transcription termination sequence" refers to a nucleic acid sequence that contains a signal that triggers the release of a newly synthesized transcript RNA molecule from the RNA polymerase complex, marking the end of transcription of a gene or locus.

[0123] The term "percent identity", "percent identity" or "percent identical" as used herein with respect to two or more nucleotide or protein sequences is calculated by (i) comparing two optimally aligned sequences (nucleotide or protein) over a comparison window, (ii) determining the number of positions where the same nucleic acid base (for nucleotide sequences) or amino acid residue (for proteins) occurs in both sequences to obtain the number of matching positions, (iii) dividing the number of matching positions by the total number of positions in the comparison window, and (iv) multiplying this quotient by 100% to obtain the percent identity. When "percent identity" is calculated relative to a reference sequence without specifying a specific comparison window, the percent identity is determined by dividing the number of matching positions over the region of alignment by the total length of the reference sequence. Thus, for the purposes of this application, when two sequences (query and subject) are optimally aligned (allowing for gaps in their alignment), the "percent identity" of the query sequence is equal to the number of identical positions between the two sequences divided by the total number of positions in the query sequence over its length (or comparison window), then multiplied by 100%. When using percentages of sequence identity in relation to proteins, it is recognized that non-identical residue positions often differ by conservative amino acid substitutions, where amino acid residues are replaced with other amino acid residues that have similar chemical properties (e.g., charge or hydrophobicity), thus not changing the functional properties of the molecule. When sequences differ in conservative substitutions, the percent sequence identity can be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." A sequence that has a percent identity to a base sequence can indicate the activity of that base sequence.

[0124] Homologs are inferred from sequence similarity by comparison of protein sequences, for example, manually or by using computer-based tools. To optimally align sequences and calculate their percent identity, various pairwise or multiple sequence alignment algorithms and programs, such as ClustalW or Basic Local Alignment Search Tool® (BLAST), are known in the art and can be used to compare sequence identity or similarity between two or more nucleotide or protein sequences. BLAST can also be used to search a query protein sequence of a base organism against a database of protein sequences of various organisms, for example, to find similar sequences. The generated summary expectation value (E-value) can be used to measure the level of sequence similarity. Since the protein with the lowest E-value hit for a particular organism is not necessarily an ortholog or the only ortholog, a reverse query is used to filter hit sequences with significant E-values ​​for ortholog identification. A reverse query requires searching for significant hits against a database of protein sequences of a base organism. If the best hit of the reverse query is the query protein itself or a paralog of the query protein, the hit can be identified as an ortholog. The reverse query process further distinguishes orthologs from paralogs among all homologs, allowing inference of functional equivalence of genes.

[0125] The term "percent complementarity" or "percent complementary" as used herein with respect to two nucleotide sequences is similar to the concept of percent identity, but refers to the percentage of nucleotides in the query sequence that optimally base-pair or hybridize to nucleotides in the subject sequence when the query and subject sequences are aligned linearly and optimally base-paired without secondary folded structures such as loops, stems or hairpins. Such percent complementarity can be between two DNA strands, between two RNA strands, or between a DNA strand and an RNA strand. "Percent complementarity" is calculated by (i) optimally base-pairing or hybridizing two nucleotide sequences in a linear and fully extended alignment (i.e., without folding or secondary structures) over a comparison window, (ii) determining the number of base-pairing positions between the two sequences over the comparison window to obtain the number of complementary positions, (iii) dividing the number of complementary positions by the total number of positions in the comparison window, and (iv) multiplying this quotient by 100% to obtain the percent complementarity of the two sequences. The optimal base pairing of two sequences can be determined based on known pairing of nucleotide bases such as GC, AT, and AU by hydrogen bonding. When "complementarity percentage" is calculated relative to a reference sequence without specifying a specific comparison window, the identity percentage is determined by dividing the number of complementary positions between two linear sequences by the total length of the reference sequence. Thus, for the purpose of this disclosure, when two sequences (query and subject) are optimally base-paired (allowing mismatched or non-base-paired nucleotides, but not including folding or secondary structures), the "complementarity percentage" of a query sequence is equal to the number of base-paired positions between two sequences divided by the total number of positions in the query sequence over its length (or the number of positions in the query sequence over its comparison window), and then multiplied by 100%.

[0126] As used herein, a "fragment" of a polynucleotide refers to a sequence that comprises at least about 50, at least about 75, at least about 95, at least about 100, at least about 125, at least about 150, at least about 175, at least about 200, at least about 225, at least about 250, at least about 275, at least about 300, at least about 500, at least about 600, at least about 700, at least about 750, at least about 800, at least about 900, or at least about 1000 consecutive nucleotides or more of a DNA molecule or protein disclosed herein. Methods for generating such fragments from a starting promoter molecule are well known in the art. A fragment of a DNA molecule or protein can exhibit the activity of the DNA molecule or protein from which it is derived.

[0127] A plant selectable marker transgene in a transformation vector or construct of the present disclosure can be used to assist in the selection of transformed cells or tissues due to the presence of a selective agent, such as an antibiotic or herbicide, where the plant selectable marker transgene provides resistance or tolerance to the selective agent. Thus, the selective agent can bias or favor the survival, development, growth, proliferation, etc., of transformed cells expressing the plant selectable marker gene, for example, to increase the proportion of transformed cells or tissues in an R0 plant. Commonly used plant selectable marker genes include those that confer resistance or tolerance to antibiotics, such as kanamycin and paromomycin (nptll), hygromycin B (aph IV), streptomycin or spectinomycin (aadA) and gentamicin (aac3 and aacC4), or those that confer resistance or tolerance to herbicides, such as glufosinate (bar or pat), dicamba (DMO) and glyphosate (proA or EPSPS). Plant screenable marker genes can also be used, which provide the ability to visually screen for transformants, such as genes expressing luciferase or green fluorescent protein (GFP), or beta-glucuronidase or the uidA gene (GUS), for which various chromogenic substrates are known. Plant transformation can also be performed in the absence of selection during one or more steps or stages of culture, development or regeneration of transformed explants, tissues, plants and / or plant parts.

[0128] IV. Transformation Methods Methods and compositions are provided for transforming plant cells, tissues or explants with recombinant DNA molecules or constructs that code for one or more molecules required for targeted genome editing (e.g., guide RNA(s) and / or site-specific nucleases). Methods suitable for transforming host plant cells include virtually all methods that can introduce DNA or RNA into cells (e.g., recombinant DNA constructs are stably integrated into plant chromosomes, or recombinant DNA constructs or RNA are transiently provided to plant cells) and are well known in the art. Two effective methods for cell transformation are bacteria-mediated transformation, such as Agrobacterium-mediated transformation or Rhizobium-mediated transformation, and microprojectile or particle bombardment transformation. Microprojectile bombardment methods are shown, for example, in U.S. Patent Nos. 5,550,318, 5,538,880, 6,160,208, and 6,399,861. Agrobacterium-mediated transformation methods are described, for example, in U.S. Patent No. 5,591,616, Hinchliffe and Harwood (2019), and Sparrow and Irwin (2015). Other methods for plant transformation, such as microinjection, electroporation, vacuum infiltration, pressurization, sonication, silicon carbide fiber agitation, PEG-mediated transformation, and the like, are also known in the art.

[0129] Transformation of plant material is carried out in tissue culture in a nutrient medium, such as a mixture of nutrients that allows cells to grow in vitro. Recipient cell targets include, but are not limited to, meristematic cells, shoot apices, hypocotyls, calli, immature or mature embryos, and reproductive cells, such as microspores and pollen. Calli can be initiated from tissue sources, including, but not limited to, immature or mature embryos, hypocotyl embryos, seedling apical meristems, microspores, and the like. Cells containing transgenic nuclei are allowed to grow into transgenic plants. Any suitable method or technique for transformation of plant cells known in the art can be used according to the present method. In transformation, DNA is typically introduced into a small percentage of target plant cells in any one transformation experiment. Marker genes are used to provide an efficient system for identifying cells that have received recombinant DNA molecules and are stably transformed by integrating them into their genome.

[0130] As used herein, the terms "regeneration" and "regenerating" refer to the process of growing or developing a plant from one or more plant cells through one or more culture steps. The transformed or edited cells, tissues or explants containing the insertion or editing of DNA sequences can be grown, developed or regenerated into transgenic plants in culture, plugs or soil according to methods known in the art. Thus, certain embodiments of the present disclosure relate to methods and constructs for regenerating plants from cells containing modified genomic DNA resulting from genome editing. The regenerated plants can then be used to propagate additional plants.

[0131] According to one aspect of the present disclosure, regenerated or progeny plants, plant parts or seeds thereof can be screened or selected based on markers, traits or phenotypes created by editing or mutation in the development or regenerated or progeny plants, plant parts or seeds, or by site-specific integration of an insertion sequence, transgene, etc. If a given mutation, editing, trait or phenotype is recessive, one or more generations or crosses (e.g., selfing) from the original R0 plant may be required to create a plant homozygous for the editing or mutation, so that the trait or phenotype can be preserved. Progeny plants, such as plants grown from R1 seeds or in subsequent generations, can be tested for zygosity using any known zygosity assay, for example, by using single nucleotide polymorphism (SNP) assays, DNA sequencing, thermal amplification, or polymerase chain reaction (PCR), and / or Southern blotting, which allows for the distinction between heterozygous, homozygous, and wild-type plants.

[0132] Methods and techniques are provided for screening and / or identifying cells or plants, etc., for the presence of a target editing or transgene, and for selecting cells or plants that contain a target editing or transgene, which can be based on one or more phenotypes or traits, or the presence or absence of a molecular marker or polynucleotide or protein sequence in a cell or plant. As used herein, "molecular techniques" refer to any method known in the art of molecular biology, biochemistry, genetics, plant biology, or biophysics, including the use, manipulation, or analysis of nucleic acids, proteins, or lipids. Molecular techniques useful for detecting the presence of modified sequences in a genome include, but are not limited to, phenotypic screening; molecular marker techniques such as SNP analysis by TaqMan® or Illumina / Infinium technology; Southern blot; PCR; enzyme-linked immunosorbent assay (ELISA); and sequencing (e.g., Sanger, Illumina®, 454, Pac-Bio, Ion Torrent™). In one aspect, the detection methods provided herein include phenotypic screening. In another aspect, the detection methods provided herein include SNP analysis. In a further aspect, the detection method provided herein comprises Southern blot. In a further aspect, the detection method provided herein comprises PCR. In one aspect, the detection method provided herein comprises ELISA. In a further aspect, the detection method provided herein comprises determining the sequence of the nucleic acid or protein. Without being limited thereto, the nucleic acid can be detected using hybridization. Hybridization between nucleic acids is described in detail in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).

[0133] Nucleic acid can be isolated using techniques conventional in the art. For example, nucleic acid can be isolated using any method, including but not limited to recombinant nucleic acid technology and / or PCR. General PCR techniques are described, for example, in PCR Primer: A Laboratory Manual, Dieffenbach&Dveksler, Eds., Cold Spring Harbor Laboratory Press, 1995. Recombinant nucleic acid techniques include, for example, restriction enzyme digestion and ligation, which can be used to isolate nucleic acid. Isolated nucleic acid can also be chemically synthesized as a single nucleic acid molecule or as a series of oligonucleotides.

[0134] Detection (e.g., of amplification products, hybridization complexes, polypeptides) can be accomplished using detectable labels that can be bound or associated with hybridization probes or antibodies. The term "label" is intended to encompass the use of direct labels as well as indirect labels. Detectable labels include enzymes, prosthetic groups, fluorescent materials, luminescent materials, bioluminescent materials, and radioactive materials. Screening and selection of modified (e.g., edited) plants or plant cells can be by any methodology known to those skilled in the art of molecular biology. Examples of screening and selection methodologies include, but are not limited to, Southern analysis, PCR amplification for detection of polynucleotides, Northern blots, RNase protection, primer extension, RT-PCR amplification for detection of RNA transcripts, Sanger sequencing, next-generation sequencing technologies (e.g., Illumina®, PacBio®, Ion Torrent™, etc.), enzyme assays for detecting enzyme or ribozyme activity of polypeptides and polynucleotides, and protein gel electrophoresis, Western blots, immunoprecipitation, and enzyme-linked immunoassays for detecting polypeptides. Other techniques, such as in situ hybridization, enzyme staining, and immunostaining, can also be used to detect the presence or expression of polypeptides and / or polynucleotides. Methods for performing all of the referenced techniques are known in the art.

[0135] As used herein, the term "polypeptide" refers to a chain of at least two covalently linked amino acids. A polypeptide can be encoded by a polynucleotide provided herein. An example of a polypeptide is a protein. A protein provided herein can be encoded by a nucleic acid molecule provided herein. A polypeptide can be purified from a natural source (e.g., a biological sample) by known methods such as DEAE ion exchange, gel filtration, and hydroxyapatite chromatography. A polypeptide can also be purified, for example, by expressing a nucleic acid into an expression vector. In addition, a purified polypeptide can be obtained by chemical synthesis. The degree of purity of a polypeptide can be measured using any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0136] Polypeptides can be detected using antibodies. Techniques for detecting polypeptides using antibodies include enzyme-linked immunosorbent assay (ELISA), Western blot, immunoprecipitation and immunofluorescence. The antibodies provided herein can be polyclonal or monoclonal. Antibodies having specific binding affinity to the polypeptides provided herein can be generated using methods well known in the art. The antibodies provided herein can be attached to a solid support, such as a microtiter plate, using methods well known in the art.

[0137] The recombinant DNA molecules provided herein can be present in a host cell, which can be any type of cell. Host cells contemplated by the present disclosure include cells selected from the group consisting of bacterial cells, animal cells, plant cells, yeast cells, fungal cells and insect cells.

[0138] For example, a bacterial host cell that may be transformed with a recombinant DNA molecule or transformation vector comprising Cas12a, guide RNA(s), or a combination thereof may be from a bacterial genus selected from the group consisting of Agrobacterium, Rhizobium, Bacillus, Brevibacillus, Escherichia, Pseudomonas, Klebsiella, Pantoea, and Erwinia.

[0139] The animal host cells that can be transformed with recombinant DNA molecules or transformation vectors that include Cas12a, guide RNA(s), or a combination thereof can include mammalian host cells, such as fibroblasts, epithelial cells, lymphocytes, or macrophages. The animal host cells according to the present disclosure can be immortalized animal cell lines, primary cells, or stem cells.

[0140] Plant cells that may be transformed with a recombinant DNA molecule or transformation vector comprising Cas12a, guide RNA(s), or a combination thereof, may include a variety of flowering plants or angiosperms, which may be further defined as including a variety of dicotyledonous (dicotyledonous) plant species or monocotyledonous (monocotyledonous) plant species. Dicotyledonous plants include members of the Fabaceae family (such as legumes), sunflowers (Helianthus annuus), safflowers (Carthamus tinctorius), sesame (Sesamum spp.), tobacco (Nicotiana tabacum), potatoes (Solanum tuberosum), cotton (Gossypium barbadense, Gossypium hirsutum), sweet potatoes (Ipomoea batatas), cassava (Manihot esculenta), coffee (Coffea spp.), tea (Camellia spp.), cherry trees (Prunus spp.), such as plums, apricots, peaches, cherries, pears (Pyrus spp.), figs (Ficus carica), citrus trees (Citrus spp.), cocoa (Theobroma spp.), and other plants. cacao), avocado (Persea americana), olive (Olea europaea), almond (Prunus amygdalus), walnut (Juglans spp.), strawberry (Fragaria spp.), watermelon (Citrullus lanatus), pepper (Capsicum spp.), beet (Beta vulgaris), grape (Vitis, Muscadinia), tomato (Lycopersicon esculentum, Solanum lycopersicum), cucumber (Cucumis sativus), as well as members of the Brassicaceae family, such as Arabidopsis thaliana and Brassica sp., e.g. B. napus, B. rapa, B. juncea, particularly Brassica species useful as a source of seed oil.Legumes and leguminous plants include peas (Pisum sativum), alfalfa (Medicago sativa), barrel clover (Medicago truncatula), pigeon pea (Cajanus cajan), guar (Cyamopsis tetragonoloba), carob (Ceratonia siliqua), fenugreek (Trigonella foenum-graecum), soybean (Glycine max), kidney bean (Phaseolus vulgaris), cowpea (Vigna unguiculata), mung bean (Vigna radiata), lima bean (Phaseolus lunatus), broad bean (Vicia faba), lentils (Lens culinaris or Lens esculenta), peanuts (Arachis hypogaea), licorice (Glycyrrhiza glabra and chickpea (Cicer arietinum). Monocotyledons can be oil palms (Elaeis spp.), coconuts (Cocos spp.), bananas (Musa spp.), and cereals such as corn (Zea mays), barley (Hordeum vulgare), sorghum (Sorghum bicolor), rice (Oryza sativa), and wheat (Triticum aestivum). Given that the present disclosure can be applied to a wide range of plant species, the present disclosure also applies to other plant structures similar to the legume pod, such as pods, siliques, fruits, nuts, tubers, and the like.

[0141] V. Genome-modified plants As used herein, "modified" in the context of a plant, plant seed, plant part, plant cell and / or plant genome refers to a plant, plant seed, plant part, plant cell and / or plant genome that contains an engineered change in the expression level and / or sequence of one or more genes of interest compared to a wild-type or control plant, plant seed, plant part, plant cell and / or plant genome. Indeed, the term "modified" may further refer to a plant, plant seed, plant part, plant cell and / or plant genome that has one or more deletions and / or one or more nucleotide substitutions or nucleotide insertions that affect an endogenous gene introduced by genome editing using any of the recombinant DNA molecules described herein. In one embodiment, a modified plant, plant seed, plant part, plant cell and / or plant genome can contain one or more transgenes. Thus, for clarity, a modified plant, plant seed, plant part, plant cell and / or plant genome includes a mutated, edited, and / or transgenic plant, plant seed, plant part, plant cell and / or plant genome having a modified genomic sequence compared to a wild-type or control plant, plant seed, plant part, plant cell and / or plant genome.

[0142] The modified plants, plant parts, seeds, etc. may be subjected to mutagenesis, genome editing or site-specific integration, genetic transformation, or a combination thereof. Such "modified" plants, plant seeds, plant parts, and plant cells include plants, plant seeds, plant parts, and plant cells that are progeny or derived from "modified" plants, plant seeds, plant parts, and plant cells that carry molecular changes (e.g., changes in expression level and / or activity) to a gene of interest. The modified seeds provided herein can generate the modified plants provided herein. The modified plants, plant seeds, plant parts, plant cells, or plant genomes provided herein can include recombinant DNA constructs or vectors or genome editing provided herein. A "modified plant product" can be any product made from the modified plants, plant parts, plant cells, or plant chromosomes provided herein, or any part or component thereof.

[0143] The modified plant can be further crossed with itself or with other plants to produce modified plant seeds and progeny. Modified plants can also be prepared by crossing a first plant containing a DNA sequence or construct or edit (e.g., a genomic deletion) with a second plant lacking the DNA sequence or construct or edit. For example, a DNA sequence or inversion can be introduced into a first plant line suitable for transformation or editing, which can then be crossed with a second plant line to introduce the DNA sequence or edit (e.g., a deletion) into the second plant line. The progeny of these crosses can be backcrossed to a desired line multiple times to produce progeny plants having substantially the same genotype as the original parent line, except for the introduction of the DNA sequence or edit, via, for example, 6-8 generations or backcrosses. The modified plants, plant cells or seeds provided herein can be hybrid plants, plant cells or seeds. As used herein, a "hybrid" is produced by crossing two plants from different varieties, lines, inbreds, or species such that the progeny contains genetic material from each parent. Those skilled in the art will recognize that higher order hybrids may also be generated.

[0144] The modified plants, plant parts, plant cells, or seeds provided herein may be elite varieties or lines. "Elite varieties" or "elite lines" refer to varieties that have arisen from breeding and selection for superior agronomic performance.

[0145] As used herein, the term "control plant" (or similarly, "control" plant seed, plant part, plant cell and / or plant genome) is used for comparison to a modified plant (or modified plant seed, plant part, plant cell and / or plant genome) and refers to a plant (or plant seed, plant part, plant cell and / or plant genome) that has the same or similar genetic background (e.g., same parent line, hybrid cross, inbred line, tester, etc.) as the modified plant (or plant seed, plant part, plant cell and / or plant genome) except for the genome edit(s) (e.g., deletion) affecting the gene of interest. For example, the control plant can be the same inbred line as the inbred line used to generate the modified plant, or the control plant can be the product of a hybrid cross of the same inbred parent line as the modified plant except that any transgenic event or genome edit(s) affecting the gene of interest are absent in the control plant. Similarly, an "unmodified control plant" refers to a plant that shares a substantially similar or essentially identical genetic background with the modified plant, but does not have one or more engineered changes (e.g., mutations or edits) to the genome of the modified plant. For purposes of comparison with a modified plant, plant seed, plant part, plant cell and / or plant genome, a "wild-type plant" (or similarly a "wild-type" plant seed, plant part, plant cell and / or plant genome) refers to a non-transgenic, and non-genome-edited control plant, plant seed, plant part, plant cell and / or plant genome. As used herein, a "control" plant, plant seed, plant part, plant cell and / or plant genome may also be a plant, plant seed, plant part, plant cell and / or plant genome that has a similar (not the same or identical) genetic background to the modified plant, plant seed, plant part, plant cell and / or plant genome, if it is deemed sufficiently similar for the comparison of the analyzed characteristic or trait.

[0146] As used herein, the terms "suppress," "suppress," "inhibit," "inhibiting," "inhibiting," "knock-out," "knock-down," and "down-regulation" refer to the reduction, decrease, or elimination of expression levels of mRNA and / or protein encoded by a target gene in a plant, plant cell, or plant tissue at one or more stages of plant development compared to the expression levels of such target mRNA and / or protein in a wild-type or control plant, cell, or tissue at the same stage(s) of plant development.

[0147] As used herein, the term "activity" refers to the biological function of a gene or protein. A gene or protein may provide one or more different functions. Thus, a decrease, disruption, or change in "activity" refers to a decrease, reduction, or elimination of one or more functions of a gene or protein in a plant, plant cell, or plant tissue at one or more stages of plant development, compared to the activity of the gene or protein in a wild-type or control plant, cell, or tissue at the same stage(s) of plant development. Furthermore, an increase in "activity" refers to an increase in one or more functions of a gene or protein in a plant, plant cell, or plant tissue at one or more stages of plant development, compared to the activity of the gene or protein in a wild-type or control plant, cell, or tissue at the same stage(s) of plant development.

[0148] According to some embodiments, plants are provided having mRNA levels of the recombinant DNA molecules described herein that are reduced or increased in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100% compared to a control plant. According to some embodiments, plants are provided having in at least one plant tissue, mRNA expression levels of a recombinant DNA molecule described herein that are reduced or increased by 5% to 20%, 5% to 25%, 5% to 30%, 5% to 40%, 5% to 50%, 5% to 60%, 5% to 70%, 5% to 75%, 5% to 80%, 5% to 90%, 5% to 100%, 75% to 100%, 50% to 100%, 50% to 90%, 50% to 75%, 25% to 75%, 30% to 80%, or 10% to 75% compared to a control plant. According to some embodiments, plants are provided that have reduced or increased levels of protein expression from the recombinant DNA molecules described herein in at least one plant tissue as compared to a control plant, the levels being at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100% reduced or increased. According to some embodiments, plants are provided having in at least one plant tissue a 5%-20%, 5%-25%, 5%-30%, 5%-40%, 5%-50%, 5%-60%, 5%-70%, 5%-75%, 5%-80%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%-90%, 50%-75%, 25%-75%, 30%-80%, or 10%-75% decreased or increased level of protein expression from a recombinant DNA molecule described herein compared to a control plant.

[0149] According to some embodiments, plants are provided that have gRNA expression levels in at least one plant tissue that are reduced or increased by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100% compared to a control plant.

[0150] According to some embodiments, plants are provided having recombinant DNA molecules that result in an increase in editing efficiency in at least one plant cell of at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100% compared to a control plant.

[0151] Modified plants comprising or derived from plant cells containing the genomic modifications of the present disclosure can be further enhanced with stacked traits, for example modified crop plants having enhanced traits resulting from expression of the DNA disclosed herein in combination with one or more additional genomic modifications that provide beneficial agronomic traits or further improve the enhanced traits.

[0152] The recitation of ranges of values ​​herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each separate value is incorporated herein as if it were individually recited herein. The recitation of discrete values ​​is understood to include the ranges between each value.

[0153] Modified plants comprising or derived from plant cells transformed with recombinant DNA of the present disclosure can be further enhanced with stacked traits, for example modified crop plants having enhanced traits resulting from expression of the DNA disclosed herein in combination with one or more agronomic benefits that provide beneficial agronomic traits to crop plants, such as herbicide and / or pest resistance traits. For example, traits imparted by recombinant DNA constructs of the present disclosure can be stacked with traits of other agronomic benefits, such as traits that provide insect resistance, for example using genes from Bacillus thuringensis to provide insect resistance to lepidopteran, beetle, homopteran, hemipteran and other insects, or quality improvement traits, such as improved nutritional value. Molecules and methods for conferring insect / nematode / virus resistance are disclosed in U.S. Patent Nos. 5,250,515; 5,880,275; 6,506,599; 5,986,175; and U.S. Patent Application Publication No. 2003 / 0150017A1.

[0154] VI.Definition The following definitions are provided to define and clarify the meaning of these terms with respect to the relevant embodiments of the present disclosure as used herein and to guide those of skill in the art in comprehending the present disclosure. Unless otherwise specified, terms are to be understood according to their conventional meaning and usage in the relevant art, particularly in the fields of molecular biology and plant transformation.

[0155] When introducing elements of the disclosure or embodiment(s) thereof, the articles "a," "an," "the," and "said" are intended to mean that there are one or more elements.

[0156] The term "and / or," when used in a list of two or more items, means any one of the items, any combination of the items, or all of the items with which the term is associated.

[0157] The terms "comprising," "including," and "having" are intended to be inclusive, meaning that there may be additional elements other than the listed elements. For example, any method that "comprises," "has," or "includes" one or more steps is not limited to having only those one or more steps, but can also include other unlisted steps. Similarly, any composition or device that "comprises," "has," or "includes" one or more features is not limited to having only those one or more features, but can also include other unlisted features.

[0158] As used herein, a "plant" includes a whole plant, an explant, a plant part, a seedling or a plantlet at any stage of regeneration or development.

[0159] As used herein, "plant part" can refer to any organ or intact tissue of a plant, such as a meristem, shoot organ / structure (e.g., leaves, stems or nodes), root, flower or floral organ / structure (e.g., protective leaves, sepals, petals, stamens, carpels, anthers and ovules), seed, embryo, endosperm, seed coat, fruit, mature ovary, bulbil, or other plant tissue (e.g., vascular tissue, dermal tissue, ground tissue, etc.), or any portion thereof. Plant parts of the present disclosure can be viable, non-viable, regenerable, and / or non-regenerable. A "bulbil" can include any plant part that can grow into a complete plant.

[0160] An "embryo" is a part of a plant seed consisting of precursor tissue (e.g., meristem) that can develop into all or part of an adult plant. An "embryo" can further include parts of a plant embryo.

[0161] "Meristem" or "meristematic tissue" includes undifferentiated or meristematic cells that can differentiate to produce all or part of one or more plant parts, tissues or structures, such as shoots, stems, roots, leaves, seeds.

[0162] As used herein, "genomic DNA" or "gDNA" refers to the chromosomal DNA of an organism.

[0163] As used herein, "genomic modification" (also referred to as "modification") or "genomic editing" (also referred to as "editing") refers to any modification to a genomic nucleotide sequence compared to a wild-type or control plant. Genomic modification or editing includes deletions, insertions, substitutions, inversions, duplications, or any combination thereof.

[0164] As used herein, "T-DNA" or "transfer DNA" refers to the transferred DNA of the tumor-inducing (Ti) plasmid of some bacterial species, such as Agrobacterium tumefaciens.

[0165] As used herein, "editing efficiency" (also referred to as "mutagenesis rate") refers to the number of T0 lines containing a targeted mutation compared to the total number of T0 lines transformed with an applicable construct to produce the targeted mutation.

[0166] As used herein, the "vegetative phase" of plant development is the period of growth between germination and flowering. For corn, a common plant development scale used in the art is known as the V stage. The V stage is defined by the top leaf where the leaf color is visible. VE corresponds to emergence, V1 corresponds to the first leaf, V2 corresponds to the second leaf, V3 corresponds to the third leaf, and V(n) corresponds to the nth leaf. VT occurs when the last branch of the tassel is visible but before silk appears. When staging a corn field, each particular V stage is defined only if more than 50% of the plants in the field are at or above that stage. Other development scales are known to those skilled in the art and can be used with the methods of the present invention. The stages of the reproductive phase of corn are as follows: R1 (silking; silk emerges from the husk); R2 (blister; kernel is white on the outside, internal fluid is clear); R3 (milk; kernel is yellow on the outside, internal fluid is milky white); R4 (glue; accumulation of starch thickens the milky internal fluid); R5 (dent; more than 50% of the kernels are dented); and R6 (physiological maturity; black layer is formed). The dormant and reproductive stages of other crop species are well known to those skilled in the art, and numerous publications describing these stages can be found on the World Wide Web and elsewhere.

[0167] As used herein, the term "isogenic" means genetically identical and non-isogenic means genetically different.

[0168] All methods described herein can be performed in any suitable order unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples or exemplary language (e.g., "etc.") provided with respect to specific embodiments herein is intended merely to clarify the disclosure and does not limit the scope of the disclosure as claimed.

[0169] [Example] [Example 1] Evaluation of novel CAS12A variants with a single promoter-guide construct in barley The editing efficiency of Lachnospiraceae Cas12a nuclease (LbCas12a) variants was evaluated in barley. In particular, the rice-optimized Cas12a coding sequence (CDS) (OsCas12a; SEQ ID NO: 1), the human-optimized Cas12a CDS functional in dicotyledonous plants (HsCas12a; SEQ ID NO: 3), and the Arabidopsis-optimized Cas12a CDS (ttAtCas12a; SEQ ID NO: 5) containing the D156R "temperature-tolerant" mutation were selected for evaluation. Two additional variants, HsCas12a with D156R mutation (ttHsCas12a; SEQ ID NO: 7) and ttAtCas12 with 8 introns (ttAtCas12+int; SEQ ID NO: 8), were also created and evaluated. The constructs containing the Cas12a nuclease variants selected for evaluation each further comprised a C-terminal nuclear localization signal operably linked to the respective codon-optimized Cas12a nuclease variant. Briefly, OsCas12a comprised the polynucleotide of SEQ ID NO: 42 (encoding SEQ ID NO: 43), HsCas12a and ttHsCas12a comprised the polynucleotide of SEQ ID NO: 44 (encoding SEQ ID NO: 43), and ttAtCas12a and ttAtCas12+int comprised the polynucleotide of SEQ ID NO: 45 (encoding SEQ ID NO: 43). The OsCas12a variant further comprised an N-terminal nuclear localization signal (SEQ ID NO: 40; encoding SEQ ID NO: 41). The novel ttAtCas12a+int variant further comprises one synonymous G to A substitution at base 2471 to remove a cryptic splice site after intron insertion.

[0170] The target barley gene used for evaluation was HORVU.MOREX.r3.1HG0069960, using the construct structure shown in Figure 1. A single U6 promoter was used to drive expression of four guide RNA sequences (SEQ ID NOs: 20-23; also referred to herein as the V1 construct or V1 array). LbCas12a can process a single gRNA transcript containing multiple guides into individual guides by recognition and cleavage at its own direct repeat (DR) sequence, thereby forming the invariant portion of the guide. To prevent the formation of spurious additional guides from the final DR, a self-processing hepatitis delta ribozyme (HDV) sequence was placed at the 3' end of the array in front of the terminator. Five constructs (OsCas12a, HsCas12a, ttAtCas12a, ttHsCas12a, and ttAtCas12+int) were generated, each containing a single Cas12a nuclease and four similar gRNA sequences. The five constructs were individually transformed into barley cultivar Golden Promise using Agrobacterium-mediated transformation and T0 plants were regenerated. DNA was extracted from T0 plants and HORVU.MOREX.r3.1HG0069960 locus PCR amplified for sequencing analysis (Sanger sequencing). ABI files were analyzed by looking at chromatograms of alignments against the wild type sequence using Benchling (https: / / www.benchling.com / ) and plants were scored as either positive or negative to mutagenesis, confirming targeted mutations using the ICE tool (Synthego-CRISPR Performance Analysis).

[0171] The number of T0 lines tested / containing mutations is shown in Figure 2. Approximately 20 T0 lines were generated for each of the five constructs, showing a notable difference in the number of lines mutated in the target. Rice-optimized OsCas12a showed no mutated lines (0 / 21), while human-optimized HsCas12a resulted in 6 / 20 (30%) mutated lines. Interestingly, inclusion of the D156 mutation in the human-optimized sequence (ttHsCas12a) increased the mutation rate to 12 / 22 (54%). Even more interestingly, Arabidopsis-optimized Cas12aCDS containing the D156R "temperature tolerance" mutation (ttAtCas12a) resulted in no mutated lines (0 / 17), whereas adding an intron (ttAtCas12a+int) resulted in 20 / 23 (87%) mutated lines. Therefore, an intron was first added to a non-functional Arabidopsis CDS to obtain ttAtCas12+int, which was transformed into the most efficient CDS evaluated in barley. Furthermore, two novel LbCas12a variants, ttHsCas12a and ttAtCas12a+int, both resulted in highly efficient targeted mutagenesis in barley. These results demonstrate the significant and surprising effects that codon usage, the D156 mutation, and the presence of an intron have on the efficiency of Cas12a mutagenesis in barley.

[0172] [Example 2] Evaluation of novel CAS12A variants with multiple promoter-guide structures in barley Four gRNA sequences were used in the LbCas12a comparison described in Example 1, but only two were determined to be active based on sequencing results. To further validate the editing efficiency of the Cas12a variants described herein, constructs were evaluated using additional gRNA constructs, where each guide was driven by a separate TaU6 / TaU3 promoter and flanked by a self-cleaving ribozyme (also referred to herein as a V2 construct or V2 array); a 5' hammerhead (HH) and a 3' HDV (Wolter 2019). Each HDV was followed by a transcription termination signal to prevent read-through. This V2 construct was combined with ttHsCas12a and used to target HORVU.MOREX.r3.1HG0069960. Eight additional constructs (four pairs) containing ttHsCas12a combined with V1 or V2 constructs were generated to target four additional barley genes, each with four guide RNA sequences. This allowed for direct comparison of the V1 / V2 guide structures. 19-25 TO lines were generated for each construct that were PCR / Sanger sequenced, aligned, and ICE tested for the targeted mutations described in Example 1.

[0173] Figure 3 shows the percentage of T0 lines with mutations in individual guide targets and the percentage of lines mutated at any guide target. The V2 array was more efficient overall than the V1 array, resulting in a higher percentage of T0 lines mutated at any guide target (36>23; 90>29; 90>88; 91>65; 85>54). Without wishing to be bound by any particular theory, the difference in editing efficiency between using the V2 array and the V1 array may be due to variations in the abundance of individual gRNAs. For example, a single TaU6 promoter can only transcribe short sequences of approximately equal length as a single guide, thereby causing downstream guides at array positions 2, 3, and 4 to be underrepresented or absent. In the V2 array, each of the four guides can be effectively transcribed by transcription from its own promoter, enriching guide RNAs at array positions 1-4. Notably, the V1 array showed higher mutagenesis in guides at array position 1 than V2 at array position 1 for all five target genes. Nevertheless, these results demonstrate that mutagenesis in approximately 90% of T0 plants for 4 / 5 barley target genes was achieved using ttHsCas12a with V2 guide arrays. These results also indicate that editing efficiency in barley can be further increased using the ttAtCas12a+int variant (87%>54%), which performed best in the Cas12a comparison described in Example 1.

[0174] [Example 3] Phenotypic evaluation of CAS12A variant-edited barley and inheritance of the edit in progeny plants To examine the ability of ttHsCas12a to result in a knockout phenotype in the first generation, mutagenesis of the barley gene HORVU.MOREX.r3.2HG0184740 was evaluated. Specifically, a construct containing ttHsCas12a and a gRNA construct(s) targeting HORVU.MOREX.r3.2HG0184740 were transformed into barley cultivar Golden Promise using Agrobacterium-mediated transformation as described in Examples 1 and 2. Knockout of both copies of HORVU.MOREX.r3.2HG0184740 is known to result in the conversion of two-line Golden Promise spikelets to six-line spikelets (Komatsuda et al., 2007). This phenotype was seen in several active T0 lines when both V1 and V2 guide constructs were used. Exemplary lines containing this phenotype are shown in Figure 4. These results confirm that ttHsCas12a resulted in the expected knockout phenotype in the first generation.

[0175] Further analysis of the T0 lines using the ICE tool calculated one T0 line targeting HORVU.MOREX.r3.1HG0069960 that contained 47% and 42% -10bp and -3bp alleles, respectively. Of the 24 T1 plants generated from it, five did not contain T-DNA, two of which were homozygous for the 3bp deletion, one was homozygous for the 10bp deletion, and two were heterozygous (Figure 5). These results demonstrate that the mutations resulting from ttHsCas12a editing in the T0 plants show inheritance in progeny plants.

[0176] Example 4: Evaluation of novel CAS12A variants with single and multiple promoter-guide constructs in B. oleracea The editing efficiency of Lachnospiraceae Cas12a nuclease (LbCas12a) variants was evaluated in B. oleracea. In particular, as described in Example 1, human-optimized Cas12a CDS (HsCas12a), Arabidopsis-optimized Cas12a CDS (ttAtCas12a) containing D156R "temperature-tolerant" mutation, novel HsCas12a with D156R mutation (ttHsCas12a), and ttAtCas12 with 8 introns (ttAtCas12+int) were selected for evaluation. The target B. oleracea gene used for evaluation was Bo2g016480.

[0177] Constructs were generated as shown in Figure 6A (referred to herein as S5, S6, S7, and S8). Briefly, S5 incorporates a guide structure similar to the V1 array, with four guide RNAs driven by one AtU626 promoter, and processing of the single transcript is carried out by the Cas12a nuclease itself. S6 has an identical LbCas12a expression cassette as S5 (ttAtCas12a), but contains a guide structure similar to the V2 array, with expression of the single guide driven by the AtU626 promoter. Thus, four S6 constructs were generated, each containing a different guide RNA (A, B, C, or D). The V2 guide structure was retained in S7 using guide C with ttHsCas12a. Similarly, S8 contained the V2 structure using guide C, but contained the ttAtCas12+int variant. The constructs were individually transformed into B. oleracea using Agrobacterium-mediated transformation and T0 plants were regenerated.

[0178] Figure 6B shows the percentage of T0 plants mutated for each target locus. From 59 S5 T0 plants screened, only two (3%) carried targeted mutations, both of which were located in the guide C target. T0 plants transformed with S6 containing the same LbCas12a expression cassette with the V2 guide structure had 10% of plants successfully mutagenized at locus A and 50% at locus C. Thus, by changing only the guide structure from V1 to V2, the editing efficiency of targeted mutagenesis increased from 0% to 10% at locus A and from 3% to 50% at locus C.

[0179] With T0 plants transformed with S7, 50% of the plants carried a mutation at locus C, indicating that ttHsCas12a and ttAtCas12a appear to be equally efficient in B. oleracea. Furthermore, the efficiency of targeted mutagenesis increased to 68% at locus C when T0 plants were transformed with S8. These results indicate that simply including eight introns in ttAtCas12a surprisingly increased the efficiency of targeted mutagenesis from 50% to 68%.

[0180] [Example 5] Inheritance of editing in progeny plants of B. oleracea To confirm that the LbCas12a-derived mutations in B. Oleracea could be inherited to the next generation in the absence of T-DNA, two T0 lines carrying mutations at locus C were analyzed in the T1 generation. Twenty-four seeds were germinated for each of the two T0 lines, and progeny without T-DNA were identified using PCR for the NptII marker. From the first line, 9 / 24 progeny did not contain T-DNA and were all homozygous for the 3 bp deletion at locus C. From the second line, 5 / 24 progeny did not contain T-DNA, three of which contained the 9 bp biallelic deletion and two of which had the 12 bp biallelic deletion (Figure 7). These results confirm that ttHsCas12a resulted in the expected knockout phenotype in the first generation. These results also demonstrate that the mutations resulting from LbCas12a editing in B. Oleracea T0 plants show inheritance in progeny plants.

[0181] [Example 6] Evaluation of novel CAS12A variant editing in wheat plants Similar editing efficiency experiments to those described in Examples 1-4 were performed in wheat. Currently, editing efficiency in wheat is believed to be very low (approximately 5%), with only one occurrence of a substantial increase to 24%. Based on the results disclosed herein, it was expected that the ttHsCas12a and ttAtCas12a+int variants could significantly increase the efficiency of Cas12a mutagenesis in wheat to a level similar to that seen in barley.

[0182] Two high-performance versions of LbCas12a, identified in the previous examples, were evaluated in wheat. Guide sequences (Wang, 2021) were used to target various genes in conjunction with human codon-optimized LbCas12a (HsCas12a), described in the previous examples and tested in barley. From these results, guides that led to mutagenesis of target genes that could be used in this experiment were identified. Using the construct structure shown in Figure 9, two guides were used simultaneously to target TaGW7 and one guide to target TaGW2.

[0183] Two constructs were generated that target both GW7 and GW2 and differ only in the LbCas12a version used: construct 1 contained ttHsCas12a (SEQ ID NO: 5) and construct 2 contained ttAtCas12a+8 intron (SEQ ID NO: 8). Forty-eight independent wheat lines were generated for each construct that were assessed by PCR and Sanger sequencing for the presence of the targeted mutation in each of the three subgenomes (A, B, and D) for both the GW7 and GW2 targets.

[0184] Both constructs resulted in mutagenesis in wheat, and as in barley, construct 2 (ttAtCas12a+8 introns) was overall more efficient than construct 1 (ttHsCas12a). At the GW2 locus, 50% of the ttHsCas12a lines were mutated in at least one of the three subgenomes, compared to 83% of the ttAtCas12a+8 intron lines. At the GW7 locus, this figure was 75% and 94%, respectively. At the GW2 locus, 21% of the ttHsCas12a lines were mutated in all three subgenomes, compared to 38% of the ttAtCas12a+8 intron lines. At the GW7 locus, this figure was 38% and 71%, respectively. 19% of the ttHsCas12a lineages were mutated in all three subgenomes at both the GW2 and GW7 loci, and this figure increased to 33% in the ttAtCas12a+8 intron lineage. Of the 288 alleles available at both the GW2 and GW7 loci in the 48 lineages generated for both constructs, 44% were mutated in the ttHsCas12a lineages and 74% were mutated in the ttAtCas12a+8 intron lineages.

[0185] These results indicate that ttAtCas12a+8 intron functions more efficiently than ttHsCas12a in wheat.

[0186] An alternative, more efficient guide structure incorporating a tRNA sequence instead of a ribozyme was also tested in wheat. A third construct was made using the ttAtCas12a+8 intron nuclease with three guide RNAs of this alternative structure, as shown in FIG. 10.

[0187] This structure further improved the results, with 96% of the lines containing mutations in at least one of the GW2 subgenomes and 94% of the lines containing mutations in at least one of the GW7 subgenomes. 90% of all three GW2 and 77% of all GW7 subgenomes were edited in the same line. 73% of the lines contained mutations in all three subgenomes of both GW2 and GW7. Of the 288 alleles available at both the GW2 and GW7 loci, 258 (90%) were edited, with 93% falling into GW2 alleles and 86% into GW7 alleles. Essentially, the greatest improvement from the use of the tRNA guide structure was brought to the GW2 locus, likely by making more GW2T6 guide transcripts available in a form that is readily available for complex formation with Cas12a nuclease.

[0188] The high efficiency of the constructs disclosed herein was quite surprising compared to previous studies performed in protoplasts that reported a maximum efficiency of about 14% (Wang, 2001). Previously reported stable transgenic lines included only 2 / 51 (4%) lines containing mutations in one subgenome at the GW7 locus, but none reported at GW2.

[0189] In summary, the ttAtCas12a+intron construct disclosed herein proved to be highly efficient in wheat. When two tRNA guides were used to target GW7, 86% of the available alleles were mutated. When one tRNA guide was used to target GW2, 93% of the available alleles were mutated.

[0190] [Example 7] Evaluation of novel CAS12A variant editing in maize plants Similar editing efficiency experiments to those described in Examples 1-4 are performed in maize. Currently, the editing efficiency of maize using LbCas12a is thought to be very low. Based on the results disclosed herein, it is expected that the ttHsCas12a and ttAtCas12a+int variants can significantly increase the efficiency of Cas12a mutagenesis in maize to levels similar to those seen in barley and B. Oleracea.

[0191] [Example 8] Comparison of editing efficiency of ttAtCAS12a with and without intron in Arabidopsis thaliana Here, the efficiency of ttAtCas12a with and without introns was compared by targeting the acetolactase synthase (ALS) gene of Arabidopsis (At3g48560) using two guide RNAs in the construct structure shown in Figure 11, where the Cas12a nuclease is driven by an egg cell specific promoter (EC.en). Egg cell expression is expected to be absent in first generation plants (T1) until meiosis, which may occur in egg cells that have separated to contain the transgene.

[0192] Only two transgenic lines were obtained with the intron-containing version of Cas12a, however this gene is likely lethal if completely knocked out due to its role in essential amino acid synthesis, which may lead to inadvertent selection for lines with low editing efficiency.

[0193] For the two intron-containing lines (prefix 3312), 48 plants per line were screened, with 21% and 12.5% ​​edited with guide 1 (approximately 16.7%) and 67% and 52% edited with guide 2 (approximately 59.5%).

[0194] Several lines were obtained for the intron-free Cas12a version. For the no-intron line (prefix 3310), enough seeds were germinated to screen 24 T2 plants per line for 9 randomly selected lines. Efficiency varied between 0%-17% for guide 1 and 4%-58% for guide 2, with an overall average efficiency of 5.1% for guide 1 and 30% for guide 2.

[0195] These results appear to indicate better performance from the intron-containing Cas12a version for the two lines evaluated. Furthermore, the data confirmed that the version of ttCas12a with eight introns disclosed herein functions in Arabidopsis.

[0196] [Example 9] Evaluation of additional CAS12A variants in barley To further test Cas12a variants in barley, additional constructs are assembled. Exemplary variants have the construct structure shown in Figure 12. Using the construct structure of Figure 12, 12 LbCas12a coding sequence (CDS) variants are tested, each construct targets the same three genes, each with only one guide shown to be functional in the previous examples.

[0197] Guide 1 targets HORVU.MOREX.r3.2HG0133680, guide 2 targets HORVU.MOREX.r3.7HG0640970, and guide 3 targets HORVU.MOREX.r3.6HG0611290. The only difference between the constructs is the coding sequence they contain. The 12 CDSs are shown in Figure 13. 20 independent transgenic barley plants are generated for each of the 12 constructs, which are sampled when large enough and screened by PCR and amplicon sequencing for editing at the target locus. The editing efficiency of the 12 CDSs across three different gene targets is determined. The editing efficiency of HsCas12a with and without D156R in barley is measured. The editing efficiency of AtCas12a with and without introns in barley is measured.

[0198] The effect of HsCas12a, ttHsCas12a and ttAtCas12a+8 intron on the editing efficiency in barley is observed for three additional gene targets.In addition, the effect of changing the number of introns in Cas12a variants is determined, including the AtCas12a with D156R (ttAtCas12a; SEQ ID NO: 5) and ttAtCas12a+8 intron comparison with ttAtCas12a+1 intron.The editing efficiency of ttAtCas12a+8 intron, ttAtCas12a+S1 intron (retaining introns 1 / 2 / 3), ttAtCas12a+S2 intron (retaining introns 4 / 5 / 6) and ttAtCas12a+S3 intron (retaining introns 7 / 8) is also evaluated.

[0199] A rice thaliana-optimized Cas12a CDS (OsCas12a+12 intron; SEQ ID NO: 58) is developed using various short Arabidopsis introns, and the gene editing efficiency of this coding sequence is evaluated in comparison with the rice-optimized Cas12a coding sequence (CDS) (OsCas12a; SEQ ID NO: 1).

[0200] [Example 10] Evaluation of additional CAS12A variants in mammalian cells Three Cas12a variants targeting DNMT-1, EXM1, and FANCF genes, L0-Cas12a-HsD156R (human codon optimized), Picsl90022 (Arabidopsis codon optimized), and EC00968 (modified Arabidopsis codons), are provided as glycerol stocks in bacteria. Mammalian cells (FreeStyle™ 293-F cells, QIB Extra, Ltd.) are transfected. Cas12a expression is determined by dot blot, and the efficiency of the reaction is evaluated by flow cytometry and sequencing.

[0201] Recombinant bacterial cells carrying the plasmids with Cas12a are grown and purified. New Cas12a recombinant plasmids are generated by cloning each of the three Cas12a insertions separately into the pcDNA3.1-U6 vector. For the crRNA plasmids, DNMT1 gRNA (SEQ ID NO: 47), EMX1 gRNA (SEQ ID NO: 48) and FANCF gRNA (SEQ ID NO: 49) are synthesized and cloned separately into pcDNA3.1-U6. In total, six recombinant plasmids based on the pcDNA3.1-U6 vector are generated.

[0202] To obtain recombinant plasmids sufficiently purified for mammalian cell transfection, the recombinant plasmids generated above are transformed into competent NEB® 10 beta competent E. coli cells using a heat shock protocol. Catabolite-repressed Super Optimal Broth is added to the cells and incubated at 37°C. The suspension is spread on LB plates containing carbenicillin. Colonies are selected for each transformation reaction and grown in LB broth, recombinant plasmids are purified using PureLink™ HiPure Plasmid Miniprep Kit, and samples are analyzed by agarose gel electrophoresis after restriction digestion to verify the integrity of the recombinant plasmids.

[0203] FreeStyle™ 293-F cells are seeded in 48-well plates with antibiotic-free medium 16 hours prior to transfection (one plate per construct). Cells are co-transfected with each recombinant Cas12a plasmid together with each crRNA recombinant plasmid using Lipofectamine 2000, resulting in nine simultaneous transfections. Cells transfected only with the relevant Cas12a plasmid are used as negative controls. To test transfection efficiency and Cas12a expression, co-transfections of three Cas12a plasmids containing DNMT1 gRNA targets are performed. Control transfections are performed with only Cas12a plasmids. After 8 hours of incubation, the transfection medium is removed and replaced with fresh medium. After 72 hours of incubation, cells are examined for Cas12a expression by antibody detection. Briefly, transfected or control cells are lysed and extracted proteins are analyzed by dot blot using first mouse anti-lbCas12a antibody and anti-mouse IgG-HRP conjugated secondary antibody. Depending on the results, transfection conditions will be optimized before moving to other co-transfection combinations.

[0204] To analyze the target gene cleavage, sequencing is used to monitor EMX1 and FANCF cleavage, while DNMT1 cleavage is determined by both sequencing and flow cytometry (as a suitable commercial antibody for this target is available). For flow cytometry, transfected cells expressing Cas12a (generated from step 3) are first stained with a viability dye (Zombie Fixable Viability), then fixed and permeabilized using fixation / permeabilization buffer, and finally the cells are incubated with anti-DNMT1-PE antibody. For the sequencing approach, FreeStyle™ 293-F cell genomic DNA is purified and used as a template for PCR using specific primers for the gene region of the target site. The PCR product will be further purified using a DNA extraction kit (Qiagen Gel Extraction Kit, Qiagen) and sequenced at our in-house sequencing facility.

Claims

1. Recombinant DNA molecules, a. Sequences having at least 85% identity with any of sequence numbers 7, 8, 1, 3, and 5; b. Sequences including sequence numbers 7, 8, 1, 3, and 5; c. A fragment having at least 85% sequence identity with any of sequence numbers 7, 8, 1, 3, and 5, and possessing nuclease activity; c. Any fragment of sequence numbers 7, 8, 1, 3, and 5; and d. A sequence encoding a protein that has at least 85% identity with any of sequence numbers 2, 4, 6, and 9; It includes a polynucleotide sequence selected from the group consisting of, A recombinant DNA molecule wherein the protein encoded by the polynucleotide sequence includes at least one intron sequence or a functional fragment thereof having a modification at amino acid position 156 compared to the protein containing the amino acid sequence of SEQ ID NO: 46, and a sequence having at least 85% identity with any one of SEQ ID NOs: 13, 14, and 15.

2. The recombinant DNA molecule according to claim 1, wherein the sequence has at least 90% identity with any of sequence numbers 7, 8, 1, 3, and 5, and encodes a protein having a modification at amino acid position 156 compared to a protein containing the amino acid sequence of sequence number 46.

3. The recombinant DNA molecule according to claim 2, wherein the sequence has at least 95% identity with any of sequence numbers 7, 8, 1, 3, and 5, and encodes a protein having a modification at amino acid position 156 compared to a protein containing the amino acid sequence of sequence number 46.

4. The recombinant DNA molecule according to claim 1, wherein the sequence includes any of sequence numbers 7, 8, 1, 3, and 5.

5. The recombinant DNA molecule according to claim 1, wherein the modification at amino acid position 156 is further defined as the substitution of aspartic acid with arginine.

6. The recombinant DNA molecule according to claim 1, wherein the polynucleotide sequence further comprises the intron sequences of SEQ ID NOs: 13, 14, and 15.

7. A transgenic plant cell comprising the recombinant DNA molecule described in claim 1.

8. The transgenic plant cell according to claim 7, wherein the transgenic plant cell is a monocotyledonous plant cell.

9. The transgenic plant cell according to claim 8, wherein the monocotyledonous plant cell is selected from the group consisting of barley, wild cabbage (B. oleracea), wheat, and maize cells.

10. The transgenic plant cell according to claim 7, wherein the transgenic plant cell is a dicotyledonous plant cell.

11. A transgenic plant or a part thereof comprising the recombinant DNA molecule described in claim 1.

12. A progeny of a transgenic plant or a part thereof, comprising the recombinant DNA molecule described above, according to claim 11.

13. A transgenic seed comprising the recombinant DNA molecule described in claim 1.

14. a. The recombinant DNA molecule is expressed in plant cells and causes genome modification, or b. The recombinant DNA molecule is operably linked to a vector, and the vector is selected from the group consisting of plasmids, phagemids, bacmids, cosmids, and bacterial or yeast artificial chromosomes. The recombinant DNA molecule according to claim 1.

15. The recombinant DNA molecule according to claim 14, wherein the host cell is selected from the group consisting of bacterial cells and plant cells.

16. The recombinant DNA molecule according to claim 15, wherein the bacterial host cell is derived from a genus of bacteria selected from the group consisting of Agrobacterium, Rhizobium, Bacillus, Brevibacillus, Escherichia, Pseudomonas, Klebsiella, Pantoea, and Erwinia.

17. The recombinant DNA according to claim 15, wherein the plant cells are dicotyledonous plant cells or monocotyledonous plant cells.

18. The aforementioned plant cells include those of the Fabaceae family, sunflower, safflower, sesame, tobacco, potato, cotton, sweet potato, cassava, coffee, tea, apple, pear, fig, citrus fruits, cocoa, avocado, olive, almond, walnut, strawberry, watermelon, chili pepper, beet, grape, tomato, cucumber, Arabidopsis thaliana, and Brassica genus. Recombinant DNA according to claim 17, selected from the group consisting of cells of ), pea, alfalfa, barrel clover, pigeon pea, guar, carob, fenugreek, soybean, kidney bean, cowpea, mung bean, lima bean, broad bean, lentil, peanut, licorice, chickpea, oil palm, coconut, banana, corn, barley, sorghum, rice, and wheat.

19. A method for creating plants containing genome modifications, a. Expressing the recombinant DNA molecule described in claim 1 and a guide RNA compatible with the protein encoded by the recombinant DNA molecule in plant cells; b. Introducing a modification to at least one target site in the genome of the plant cell; c. Identifying and selecting one or more plant cells from step (b) that contain the modifications in the plant genome; and d. Regenerating at least one plant from at least one cell selected in step (c). A method that includes this.

20. The method according to claim 19, wherein the modification is selected from the group consisting of substitution, insertion, inversion, deletion, duplication, and combinations thereof.

21. The method according to claim 19, wherein the plant is a monocotyledonous plant.

22. The method according to claim 21, wherein the plant is selected from the group consisting of barley, wild cabbage (B. oleracea), wheat, and maize plants.

23. A method for producing progeny seeds containing the recombinant DNA molecule described in claim 1, a. Sowing the first seed containing the recombinant DNA molecule described in claim 1; b. Growing plants from the seeds of process (a); and c. A method comprising harvesting the progeny seeds from the plant, wherein the harvested seeds contain the recombinant DNA molecule.

24. A method for introducing genome modifications into plants, a. Expressing a protein or fragment thereof encoded by the DNA molecule described in claim 1 in a plant; and b. To express in plant cells a guide RNA that is compatible with the protein or fragment having nuclease activity. Methods that include...

25. A method for detecting the presence of recombinant DNA molecules according to claim 1 in a sample containing plant genomic DNA, a. Hybridizing the sample with plant-derived genomic DNA containing the recombinant nuclear DNA described in claim 1 under stringent hybridization conditions, and contacting the probe with a DNA probe that does not hybridize with genomic DNA from other isogeneic plants that do not contain the recombinant DNA molecule described in claim 1 under such hybridization conditions, the probe being homologous or complementary to a protein encoding sequence containing an amino acid sequence having at least 85%, 90%, 95%, 98%, 99%, or about 100% amino acid sequence identity with any of the fragments of SEQ ID NOs: 7, 8, 1, 3, or 5, or any of SEQ ID NOs: 2, 4, 6, and 9; b. Subjecting the sample and the probe to stringent hybridization conditions; and c. Detecting the hybridization between the DNA probe and the recombinant DNA molecule. A method that includes this.

26. A method for detecting the presence of a nuclease protein or a fragment thereof in a sample containing a protein, wherein the protein contains an amino acid sequence or fragment thereof of any of SEQ ID NOs: 2, 4, 6, and 9, or the protein contains an amino acid sequence or fragment thereof having at least 85%, 90%, 95%, 98%, 99%, or about 100% amino acid sequence identity with any of SEQ ID NOs: 2, 4, 6, and 9. a. Contacting the sample with an immunoreactive antibody; and b. Detecting the presence of the protein or its fragments. A method that includes this.

27. A method for modifying a polynucleotide segment encoding a Cas12a protein or a fragment thereof that has nuclease activity, a. Obtain one of the polynucleotide sequences of sequence numbers 7, 8, 1, 3, or 5; and b. The modification involves introducing a modification to at least one target site in the polynucleotide sequence such that the protein encoded by the polynucleotide sequence includes a modification at amino acid position 156 compared to the protein containing the amino acid sequence of SEQ ID NO: 46, A method wherein the modified polynucleotide sequence further comprises at least one intron sequence or a functional fragment thereof having a sequence that is at least 85 percent identical to any one of sequence numbers 13, 14, and 15.

28. The method according to claim 27, wherein the protein encoded by the modified polynucleotide sequence comprises a substitution from aspartic acid to arginine at amino acid position 156, compared to a polynucleotide segment lacking the modification.

29. The method according to claim 28, wherein the modified polynucleotide sequence further comprises the intron sequences of SEQ ID NOs: 13, 14, and 15.

30. The method according to claim 27, wherein the modified polynucleotide sequence comprises a modification from aspartic acid to arginine at amino acid position 156, and further comprises at least one intron sequence of SEQ ID NOs: 13, 14, and 15.

31. A method for improving gene targeting using CRISPR-Cas12a gene editing in crops, a. A step of expressing a recombinant DNA molecule according to claim 1 and a guide RNA compatible with the protein encoded by the recombinant DNA molecule in plant cells; and b. A step comprising introducing a modification to at least one target site in the plant cell genome, A method wherein the modification is introduced at a higher rate than the introduction rate of the modification using a method that includes the expression of a DNA molecule encoding the amino acid of SEQ ID NO:

46.

32. The method according to claim 31, wherein the sequence has at least 90% identity with any of sequence numbers 7, 8, 1, 3, and 5, and encodes a protein having a modification at amino acid position 156 compared to a protein containing the amino acid sequence of sequence number 46.

33. The method according to claim 32, wherein the sequence has at least 95% identity with any of sequence numbers 7, 8, 1, 3, and 5, and encodes a protein having a modification at amino acid position 156 compared to a protein containing the amino acid sequence of sequence number 46.

34. The method according to claim 31, wherein the sequence includes any of sequence numbers 7, 8, 1, 3, and 5.

35. The method according to claim 31, wherein the modification at amino acid position 156 is further defined as the substitution of aspartic acid with arginine.

36. The method according to claim 31, wherein the polynucleotide sequence further comprises the intron sequences of SEQ ID NOs: 13, 14, and 15.