Plant Regulatory Elements and Uses Thereof
Patent Information
- Application Number
- JP2023565870
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-30
- Filing Date
- 2022-04-28
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-04-28
AI Technical Summary
Existing genome editing techniques using CRISPR systems face challenges with multiple U6 snRNA promoters causing recombination events and deletions due to sequence homology, leading to instability in plasmids containing multiple gRNA cassettes.
Development of novel synthetic snRNA promoters with low sequence homology to native U6 promoters, allowing for stable expression of guide RNAs in plant cells, reducing recombination events and promoting efficient genome modification.
The synthetic snRNA promoters enhance the stability and efficiency of CRISPR-mediated genome editing in plants by minimizing construct instability and enabling multiple gRNA cassettes without sequence overlap issues.
Abstract
Description
[Technical field]
[0001] REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 182,288, filed April 30, 2021, and U.S. Provisional Application No. 63 / 295,061, filed December 30, 2021, which are incorporated by reference in their entireties herein.
[0002] Incorporation of sequence listing The sequence listing contained in the file titled "MONS492WO-sequence_listing" is 37 kilobytes (measured in Microsoft Windows®), was created on April 28, 2022, was submitted herewith by electronic application, and is incorporated herein by reference.
[0003] The present disclosure relates to the field of biotechnology.More specifically, the present disclosure provides a novel synthetic plant promoter useful for expressing non-protein-coding small RNA for, for example, CRISPR-mediated genome modification. [Background technology]
[0004] Site-specific recombination has the potential to be applied in a wide range of biotechnology-related fields. Meganucleases, zinc finger nucleases (ZFNs), and transcription activator-like effector nucleases (TALENs), which contain DNA binding and DNA cleavage domains, allow genome modification. Meganucleases, ZFNs, and TALENs are effective and specific, but these techniques require the generation of one or more components for each genome site selected for modification via protein engineering. Advances in the application of clustered regularly interspaced short palindromic repeats (CRISPR) have demonstrated a genome modification method with the advantage of being rapidly engineered.
[0005] The clustered regularly interspaced short palindromic repeats (CRISPR) system constitutes a prokaryotic adaptive immune system that targets endonucleolytic cleavage of invading phages. The system is composed of a protein component (Cas) and a guide RNA (gRNA) that directs the Cas protein to specific loci for endonucleolytic cleavage. The system has been successfully engineered to target specific loci for endonucleolytic cleavage in the genomes of mammals, zebrafish, Drosophila, C. elegans, bacteria, yeast, and plants.
[0006] The DNA sequence encoding the guide RNA is preferably transcribed by RNA polymerase III, which transcribes small nuclear RNAs (snRNAs). Native promoters, such as the U6 snRNA promoter, are often used to drive the expression of gRNAs. Multiplex targeting experiments often rely on the same promoter driving each of the gRNAs. This can lead to technical problems when cloning or maintaining plasmids containing multiple U6 / gRNA cassettes, such as recombination events or deletions resulting from sequence overlap between cassettes. Using multiple snRNA promoters with diverse DNA sequences helps to alleviate this technical problem. Thus, the inventors herein disclose novel synthetic snRNA promoters that have little sequence homology with known native U6 snRNA promoters and with each other. These novel synthetic snRNA promoters are capable of driving the expression of RNA polymerase III transcripts, such as gRNAs, in plant cells. Summary of the Invention
[0007] In one aspect, the present invention provides a synthetic small nuclear RNA (snRNA) promoter comprising a DNA sequence selected from the group consisting of: (a) a sequence having at least 85% sequence identity to any of SEQ ID NOs: 1-10; (b) a sequence comprising any of SEQ ID NOs: 1-10; and (c) a fragment of any of SEQ ID NOs: 1-10. In one embodiment, the synthetic snRNA promoter sequence has at least 90 percent sequence identity to the DNA sequence of any of SEQ ID NOs: 1-10. In another embodiment, the synthetic snRNA promoter sequence has at least 95 percent sequence identity to the DNA sequence of any of SEQ ID NOs: 1-10. In yet another embodiment, the synthetic snRNA promoter fragment comprises gene regulatory activity.
[0008] Another aspect of the invention provides a recombinant DNA construct comprising a synthetic snRNA promoter operably linked to a DNA sequence encoding one or more guide RNAs (gRNAs), wherein the sequence of the synthetic snRNA promoter is selected from the group consisting of: (a) a sequence having at least 85% sequence identity to any of SEQ ID NOs: 1-10; (b) a sequence comprising any of SEQ ID NOs: 1-10; and (c) a fragment of any of SEQ ID NOs: 1-10; wherein the synthetic snRNA promoter is capable of expressing a gRNA. In some embodiments, the recombinant DNA construct further comprises a transcription termination sequence. In some embodiments, the recombinant DNA construct may also further comprise a DNA sequence encoding a promoter operably linked to a DNA sequence encoding a clustered regularly interspaced short palindromic repeats CRISPR-associated protein. In some embodiments, the CRISPR-associated protein is selected from a type I CRISPR-associated protein, a type II CRISPR-associated protein, a type III CRISPR-associated protein, a type IV CRISPR-associated protein, a type V CRISPR-associated protein, or a type VI CRISPR-associated protein. In some embodiments, the CRISPR-associated protein is a synthetic CRISPR-associated protein. In certain embodiments of the recombinant DNA construct, the nucleotide sequence encoding the CRISPR-associated protein can be further operably linked to at least one nuclear localization sequence (NLS).Further, in certain embodiments of the contemplated recombinant DNA constructs, the CRISPR-associated protein is selected from the group consisting of Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12a (also known as Cpf1), Cas12b, Cas12d, Csy1, Csy2, Csy3, Cse1, Cse2, Cse3, Cse4, Cse5, Cse6, Cse7, Cse8, Cse ... 2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, CasX, CasY, and Mad7. In certain embodiments, the construct comprises adjacent left and right homology arms (HA), each of about 2 to 1200 bp in length. In certain embodiments, the length of the homology arms is about 230 to about 1003 bp.
[0009] Another aspect of the present invention provides a recombinant DNA construct comprising a first synthetic snRNA promoter operably linked to a DNA sequence encoding one or more guide RNAs (gRNAs) and a second synthetic snRNA promoter operably linked to a DNA sequence encoding one or more guide RNAs (gRNAs), wherein the sequences of the first and second synthetic snRNA promoters are independently selected from the group consisting of: (a) a sequence having at least 85% sequence identity to any of SEQ ID NOs: 1-10; (b) a sequence comprising any of SEQ ID NOs: 1-10; and (c) a fragment of any of SEQ ID NOs: 1-10, wherein the fragment is capable of expressing a gRNA. In certain embodiments, the first synthetic snRNA promoter is different from the second synthetic snRNA promoter. In certain embodiments, the sequence encoding one or more gRNAs expressed by the first synthetic snRNA promoter is different from the sequence encoding one or more gRNAs expressed by the second synthetic snRNA promoter. In some embodiments, the gRNA-encoding sequence further comprises a sequence encoding one or more tRNAs described in WO / 2016 / 061481, which is incorporated herein by reference in its entirety. In certain embodiments, the construct comprises adjacent left and right homology arms (HA), each of which is about 2 to 1200 bp in length. In certain embodiments, the homology arms are about 230 to about 1003 bp in length. In some embodiments, the recombinant DNA construct further comprises a transcription termination sequence. In some embodiments, the recombinant DNA construct may also further comprise a DNA sequence encoding a promoter operably linked to a DNA sequence encoding a clustered regularly interspaced short palindromic repeats CRISPR-associated protein. In some embodiments, the CRISPR-associated protein is selected from a type I CRISPR-Cas system, a type II CRISPR-Cas system, a type III CRISPR-Cas system, a type IV CRISPR-Cas system, a type V CRISPR-Cas system, or a type VI CRISPR-Cas system. In some embodiments, the CRISPR-associated protein is a synthetic CRISPR-associated protein.In certain embodiments of the recombinant DNA construct, the nucleotide sequence encoding the CRISPR-associated protein may be further operably linked to at least one nuclear localization sequence (NLS). In addition, in certain contemplated embodiments of the recombinant DNA construct, the CRISPR-associated protein may be any of Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12a (also known as Cpf1), Cas12b, Cas12d, Csy1, Csy2, Csy3, Cse1, Cse2, Cse3, Cse4, Cse5, Cse6, Cse7, Cse8, Cse9 (also known as Csn1 and Csx12), Cas10, Cas12a (also known as Cpf1), Cas12b, Cas12d, Csy1, Csy2, Csy3, Cse1, Cse8, Cse9, Cse1 ... 2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, CasX, CasY, and Mad7.
[0010] Another aspect of the invention provides a recombinant DNA construct comprising a synthetic snRNA promoter operably linked to a sequence encoding a non-coding RNA, wherein the sequence of the synthetic snRNA promoter is selected from the group consisting of: (a) a sequence having at least 85% sequence identity to any of SEQ ID NOs: 1-10; (b) a sequence comprising any of SEQ ID NOs: 1-10; and (c) a fragment of any of SEQ ID NOs: 1-10, wherein the fragment comprises gene regulatory activity. In some embodiments, the non-coding RNA is selected from the group consisting of guide RNA (gRNA), microRNA (miRNA), miRNA precursor, mature miRNA, decoy miRNA (described in WO2010 / 002984, which is incorporated herein by reference), small interfering RNA (siRNA), small RNA (22-26 nt in length) and precursors encoding same, heterochromatic siRNA (hc-siRNA), Piwi-interacting RNA (piRNA), hairpin double-stranded RNA (hairpin dsRNA), trans-acting siRNA (ta-siRNA), and naturally occurring antisense siRNA (nat-siRNA). In some embodiments, the recombinant DNA construct comprises a synthetic snRNA promoter operably linked to a sequence encoding two or more non-coding RNAs. In some embodiments, the sequence encoding the two or more non-coding RNAs further comprises a sequence encoding one or more tRNAs.
[0011] Yet another aspect of the present invention includes a recombinant DNA construct comprising: a) a first synthetic snRNA promoter; and b) a second synthetic snRNA promoter; the first synthetic snRNA promoter is selected from the group consisting of: (a) a sequence having at least 85% sequence identity with any of SEQ ID NOs: 1-10; (b) a sequence comprising any of SEQ ID NOs: 1-10; and (c) a fragment of any of SEQ ID NOs: 1-10, wherein the fragment comprises gene regulatory activity and is operably linked to a DNA sequence encoding a non-coding RNA; and the second synthetic snRNA promoter is selected from the group consisting of: (a) a sequence having at least 85% sequence identity with any of SEQ ID NOs: 1-10; (b) a sequence comprising any of SEQ ID NOs: 1-10; and (c) a fragment of any of SEQ ID NOs: 1-10, wherein the fragment comprises gene regulatory activity and is operably linked to a DNA sequence encoding a non-coding RNA, wherein the first synthetic snRNA promoter and the second synthetic snRNA promoter are different. In certain embodiments of the recombinant DNA construct, the sequence encoding the first synthetic snRNA promoter and the sequence encoding the second synthetic snRNA promoter each comprise any of SEQ ID NOs: 1-10, or fragments thereof, wherein the fragment comprises gene regulatory activity. Also contemplated are embodiments in which the recombinant DNA construct further comprises a sequence specifying one or more additional synthetic snRNA promoters selected from the group consisting of SEQ ID NOs: 1-10, or fragments thereof, wherein the fragment comprises gene regulatory activity and is operably linked to a DNA sequence encoding a non-coding RNA, wherein each of the first synthetic snRNA promoter, the second synthetic snRNA promoter, and the one or more additional snRNA promoters are different. In certain embodiments, the recombinant DNA construct sequence specifying the one or more additional synthetic snRNA promoters is selected from the group consisting of SEQ ID NOs: 1-10, or fragments thereof, wherein the fragment comprises gene regulatory activity. In still other embodiments, the recombinant DNA construct comprises three, four, or five synthetic snRNA promoters. In some embodiments, the recombinant DNA construct comprises a non-coding RNA that is a gRNA that targets different selected target sites in a chromosome of the plant cell.In other contemplated embodiments, the recombinant DNA further comprises a DNA sequence encoding a promoter operably linked to a DNA sequence encoding an RNA-guided endonuclease. In further embodiments, the RNA-guided endonuclease is a clustered regularly interspaced short palindromic repeats (CRISPR)-associated protein. In some embodiments, the CRISPR-associated protein is selected from a type I CRISPR-Cas protein, a type II CRISPR-Cas protein, a type III CRISPR-Cas protein, a type IV CRISPR-Cas protein, a type V CRISPR-Cas protein, and a type VI CRISPR-Cas protein. In some embodiments, the CRISPR-associated protein is a synthetic CRISPR-associated protein. In some embodiments, the CRISPR associated protein is Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12a (also known as Cpf1), Cas12b, Cas12d, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1 , Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, CasX, CasY, and Mad7.
[0012] Another aspect of the present invention provides a cell comprising any of the above recombinant DNA constructs. In certain embodiments, the cell is a plant cell. In some embodiments, the plant cell is a monocotyledonous plant cell. In other embodiments, the plant cell is a dicotyledonous plant cell. In yet another embodiment, the plant cell is selected from the group consisting of a corn plant cell, a soybean plant cell, a cotton plant cell, a peanut plant cell, a barley plant cell, an oat plant cell, a orchard grass plant cell, a rice plant cell, a sorghum plant cell, a sugarcane plant cell, a tall fescue plant cell, a turfgrass plant cell, a wheat plant cell, alfalfa plant cell, a canola plant cell, a cabbage plant cell, a mustard plant cell, a rutabaga plant cell, a turnip plant cell, a kale plant cell, a broccoli plant cell, a cauliflower plant cell, a pepper plant cell, a bean plant cell, a cowpea plant cell, a chickpea plant cell, a gourd plant cell, a lettuce plant cell, a cucumber plant cell, a melon plant cell, a carrot plant cell, a tomato plant cell, a radish plant cell, a potato plant cell, and an ornamental plant cell.
[0013] Detailed description of the sequence SEQ ID NO:1 is the DNA sequence of the synthetic snRNA promoter, P-GSP2262.
[0014] SEQ ID NO:2 is the DNA sequence of the synthetic snRNA promoter, P-GSP2268.
[0015] SEQ ID NO:3 is the DNA sequence of the synthetic snRNA promoter, P-GSP2269.
[0016] SEQ ID NO:4 is the DNA sequence of the synthetic snRNA promoter, P-GSP2272.
[0017] SEQ ID NO:5 is the DNA sequence of the synthetic snRNA promoter, P-GSP2273.
[0018] SEQ ID NO: 6 is the DNA sequence of the shortened variant synthetic snRNA promoter P-GSP2262_TR derived from P-GSP2262.
[0019] SEQ ID NO: 7 is the DNA sequence of the shortened variant synthetic snRNA promoter P-GSP2268_TR derived from P-GSP2268.
[0020] SEQ ID NO: 8 is the DNA sequence of the shortened variant synthetic snRNA promoter P-GSP2269_TR derived from P-GSP2269.
[0021] SEQ ID NO: 9 is the DNA sequence of the shortened variant synthetic snRNA promoter P-GSP2272_TR derived from P-GSP2272.
[0022] SEQ ID NO: 10 is the DNA sequence of the shortened variant synthetic snRNA promoter P-GSP2273_TR derived from P-GSP2273.
[0023] SEQ ID NO:11 is the DNA sequence of EXP, EXP-Zm.UbqM1:1:9, consisting of the promoter, leader, and intron derived from the Zea mays subsp. mexicana ubiquitin gene.
[0024] SEQ ID NO: 12 is a DNA sequence encoding the nuclear-targeted Cas12a protein, Cas12a_NLS.
[0025] SEQ ID NO: 13 is the DNA sequence of the 3'UTR, T-Os.LTP:2.
[0026] SEQ ID NO: 14 is the DNA sequence of the guide RNA spacer, NR-Zm.Bmr3_2691.
[0027] SEQ ID NO: 15 is the DNA sequence of the guide RNA, gRNA-Zm.Bmr3_2691.
[0028] SEQ ID NO: 16 is the DNA sequence of the guide RNA spacer, NR-Zm.Bmr3_3170.
[0029] SEQ ID NO: 17 is the DNA sequence of the guide RNA, gRNA-Zm.Bmr3_3170.
[0030] SEQ ID NO: 18 is the DNA sequence of the Zea mays brown midrib 3 (Bmr3) genomic region targeted for genome editing.
[0031] SEQ ID NO:19 is the amino acid sequence of Cas12a_NLS encoded by SEQ ID NO:12.
[0032] SEQ ID NO: 20 is the DNA sequence of the guide RNA spacer, NR-Zm.Bmr3_90.
[0033] SEQ ID NO: 21 is the DNA sequence of the guide RNA spacer, NR-Zm.Bmr3_227.
[0034] SEQ ID NO: 22 is the DNA sequence of the guide RNA spacer, NR-Zm.Bmr3_3279.
[0035] SEQ ID NO: 23 is the DNA sequence of the guide RNA, gRNA-Zm.Bmr3_90_3279.
[0036] SEQ ID NO: 24 is the DNA sequence of the guide RNA, gRNA-Zm.Bmr3_227_3279.
[0037] SEQ ID NO: 25 is the DNA sequence of the guide RNA, gRNA-Zm.Bmr3_2691_2.
[0038] SEQ ID NO: 26 is the DNA sequence of the guide RNA, gRNA-Zm.Bmr3_3170_2.
[0039] SEQ ID NO: 27 is the DNA sequence of the guide RNA, gRNA-Zm.Bmr3_2691_3170.
[0040] SEQ ID NO:28 is the DNA sequence of the Zea mays Zm7 genomic region targeted for genome editing.
[0041] SEQ ID NO:29 is the DNA sequence of the guide RNA spacer, NR-Zm.7.1b.
[0042] SEQ ID NO:30 is the DNA sequence of the guide RNA, gRNA-Zm.7.1b.
[0043] SEQ ID NO:31 is the DNA sequence of the guide RNA spacer, NR-Zm.7.1c.
[0044] SEQ ID NO:32 is the DNA sequence of the guide RNA, gRNA-Zm.7.1c.
[0045] SEQ ID NO:33 is the DNA sequence of the guide RNA, gRNA-7.1c_7.1b. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0046] Provided herein are novel synthetic snRNA (small nuclear RNA) promoters active in plants. Nucleotide sequences of these small nuclear RNA promoters are provided as SEQ ID NOs: 1-10. These small nuclear RNA promoters can affect expression of non-coding RNAs, such as guide RNAs, in plant tissues, and thus can regulate expression of operably linked sequences encoding non-coding RNAs in plants. Also provided are methods of modifying, producing, and using recombinant DNA molecules comprising the provided small nuclear RNA promoters. Also provided are compositions comprising transgenic plant cells, plants, plant parts, and seeds comprising the small nuclear RNA promoters of the invention, as well as methods for preparing and using the same.
[0047] In some embodiments, a variant of a small nuclear RNA promoter selected from SEQ ID NOs: 1-10 is provided. In some embodiments, a variant is provided that has at least about 85 percent identity, at least about 86 percent identity, at least about 87 percent identity, at least about 88 percent identity, at least about 89 percent identity, at least about 90 percent identity, at least about 91 percent identity, at least about 92 percent identity, at least about 93 percent identity, at least about 94 percent identity, at least about 95 percent identity, at least about 96 percent identity, at least about 97 percent identity, at least about 98 percent identity, or at least about 99 percent identity to a reference sequence when optimally aligned to a reference sequence provided herein as any of SEQ ID NOs: 1-10, and has promoter activity as disclosed herein. A variant of any of SEQ ID NOs: 1-10 may have the activity of a base sequence, for example, promoter activity of a base sequence.
[0048] In some embodiments, a fragment of a small nuclear RNA promoter selected from SEQ ID NO: 1-10 is provided that comprises at least about 50, at least about 75, at least about 95, at least about 100, at least about 125, at least about 150, at least about 175, at least about 200, at least about 225, at least about 250, at least about 275, at least about 300, at least about 325, at least about 350, at least about 375, at least about 400 consecutive nucleotides, at least about 425, at least about 450, at least about 475, or more, of a DNA molecule having a promoter activity as disclosed herein. In certain embodiments, a fragment of a small nuclear RNA promoter provided herein is provided that has gene expression activity. Methods for producing such fragments from a starting promoter molecule are well known in the art. A fragment of any of SEQ ID NO: 1-10 may have a base activity, e.g., a promoter activity of a base sequence.
[0049] Compositions derived from any of the promoter elements contained in any of SEQ ID NOs: 1-10 (e.g., internal or 5' deletions) can be produced using methods known in the art to improve or modify expression, including, for example, removal of elements that have either a positive or negative effect on expression, duplication of elements that have a positive or negative effect on expression, and / or duplication or removal of elements that have tissue-specific or cell-specific effects on expression. Compositions derived from any of the promoter elements contained in any of SEQ ID NOs: 1-10 that consist of a 3' deletion, in which the TATA box element or its equivalent and downstream sequences have been removed, can be used to create, for example, enhancer elements. These enhancer elements can be operably linked to other synthetic or native snRNA promoters to enhance expression. Further deletions can be made to remove any elements that have a positive or negative effect on expression. Any of the promoter elements contained in any of SEQ ID NOs: 1-10, and fragments or enhancers derived therefrom, can be used to create chimeric transcriptional regulatory element compositions.
[0050] In some embodiments, the present disclosure provides novel synthetic snRNA (small nuclear RNA) promoters and methods of use thereof, including expression of guide RNAs for targeted gene modification of plant genomes by clustered regularly interspaced short palindromic repeats (CRISPR) editing systems. For example, the present disclosure provides, in one embodiment, a DNA construct encoding at least one expression cassette comprising the synthetic snRNA promoter disclosed herein and a DNA sequence encoding one or more guide RNAs (gRNAs). Methods for allowing the CRISPR system to modify a target genome are also provided, as are genome complements of plants modified by the use of such systems. Thus, the present disclosure provides tools and methods that allow genes, loci, junction blocks, and chromosomes in plant genomes to be inserted, removed, or modified.
[0051] In another embodiment, the present disclosure provides DNA constructs encoding at least one expression cassette comprising a promoter as disclosed herein and a DNA sequence encoding a small non-protein-coding RNA (npcRNA). These constructs are useful for expressing npcRNA molecules.
[0052] CRISPR systems constitute an adaptive immune system in prokaryotes that targets endonucleolytic cleavage of DNA and RNA of invading phages (reviewed in Westra et al., Annu Rev Genet 46:311-39, 2012). Six types of CRISPR systems (type I, type II, type III, type V, and type VI) are known that rely on small RNAs to target foreign nucleic acids for sequence-specific detection and destruction. The components of bacterial CRISPR systems are CRISPR-associated (Cas) proteins and CRISPR array(s) that contain genomic targeting sequences (protospacers) interspersed with short palindromic repeats. In the case of type II CRISPR systems, the protospacer / repeat elements are transcribed into precursor CRISPR RNA (pre-crRNA) molecules, followed by enzymatic cleavage triggered by hybridization between trans-acting CRISPR RNA (tracrRNA) molecules and the pre-crRNA palindromic repeats. The resulting crRNA:tracrRNA molecule consists of one copy of the spacer and a scaffold that can form a complex with the Cas nuclease. The CRISPR / Cas complex is then directed to a DNA sequence (protospacer) that is complementary to the crRNA spacer sequence, and this RNA-Cas protein complex silences the target DNA via enzymatic cleavage of both strands (double-strand breaks; DSBs).
[0053] The native bacterial type II CRISPR system requires four molecular components for targeted cleavage of exogenous DNA: a Cas endonuclease (e.g., Cas9), a housekeeping RNaseIII, a CRISPR RNA (crRNA), and a trans-acting CRISPR RNA (tracrRNA). The latter two components form a dsRNA complex and bind to Cas9 to form an RNA-guided DNA endonuclease complex. For targeted genome modification in eukaryotes, this system has been simplified to two components: Cas9 endonuclease and a guide RNA (gRNA). Experiments first performed in eukaryotic systems found that the RNaseIII component was not required to achieve targeted DNA cleavage. The minimal two-component system, including Cas9 and a gRNA as the only target-specific component, makes this CRISPR system targeted genome modification system more cost-effective and flexible than other targeting platforms such as meganucleases, zinc finger nucleases, and TALE nucleases, which require protein engineering for modification at each targeted DNA site. Furthermore, due to the ease of designing and producing gRNA, the CRISPR system offers several advantages in the application of targeted genome modification. For example, the components of the CRISPR / Cas system (Cas endonuclease, gRNA, and optionally exogenous DNA for integration into the genome) designed for one or more genome target sites can be multiplexed into one transformation, or the introduction of the components of the CRISPR / Cas system can be spatially and / or temporally separated.
[0054] As used herein, "guide nucleic acid" or "guide RNA" or "gRNA" refers to a nucleic acid that includes a spacer sequence that is complementary to (and hybridizes with) a target DNA sequence and a scaffold sequence that binds to a Cas protein. In some embodiments, the scaffold sequence and the spacer sequence are covalently linked and expressed as one RNA transcript or molecule, referred to herein as a "single-stranded guide RNA" (or "sgRNA"). In some embodiments, the scaffold sequence and the spacer sequence are expressed as separate transcripts or molecules, referred to herein as a "dual guide RNA" (or "dgRNA"). The spacer sequence may be either covalently or non-covalently linked to the 5' and / or 3' end of the scaffold sequence. In some embodiments, the guide RNA includes a CRISPR RNA (crRNA) and a trans-activating crRNA (tracrRNA). In other embodiments, the guide RNA includes a crRNA but not a tracrRNA. In some embodiments, the crRNA includes both a spacer sequence and a scaffold sequence. In some embodiments, the gRNA design may be based on a Type I, Type II, Type III, Type IV, Type V, or Type VI CRISPR-Cas system.
[0055] In some embodiments, the array of guide RNAs is expressed from a synthetic snRNA promoter described herein. In some embodiments, the synthetic snRNA promoter described herein can be operably linked to two or more scaffold-spacer (and / or spacer-scaffold) sequences (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more scaffold-spacer (and / or spacer-scaffold) sequences) (e.g., scaffold-spacer-scaffold, e.g., spacer-scaffold-spacer, e.g., scaffold-spacer-scaffold-spacer-scaffold-spacer-scaffold-spacer, e.g., spacer-scaffold-spacer-scaffold-spacer-scaffold-spacer-scaffold-spacer, etc.). In some embodiments, the guide RNA array comprises one or more tRNAs, as described in WO / 2016 / 061481. In some embodiments, the guide RNA array comprises one or more tRNAs separating the scaffold and spacer sequences (e.g., scaffold-spacer-tRNA-scaffold-spacer, e.g., spacer-scaffold-tRNA-spacer-scaffold, e.g., scaffold-spacer-tRNA-scaffold-spacer-tRNA-scaffold-spacer-tRNA-scaffold-spacer, e.g., spacer-scaffold-tRNA-spacer-scaffold-tRNA-spacer-scaffold-tRNA-spacer-scaffold, etc.). In some embodiments, the scaffold sequence is selected from the group consisting of: a repeat sequence of a Cas12a CRISPR-Cas system or a fragment thereof; a repeat sequence of a Cas12b CRISPR-Cas system or a fragment thereof; a repeat sequence of a Cas12c CRISPR-Cas system or a fragment thereof; a repeat sequence of a Cas12d CRISPR-Cas system or a fragment thereof; a repeat sequence of a Cas12e CRISPR-Cas system or a fragment thereof; a repeat sequence of a Cas9 CRISPR-Cas system or a fragment thereof; a repeat sequence of a C2c1 CRISPR Cas system or a fragment thereof; a repeat sequence of a C2c3 CRISPR-Cas system or a fragment thereof; a repeat sequence of a Cas13a CRISPR-Cas system or a fragment thereof;Cas13b CRISPR-Cas system repeat sequence or fragment thereof;Cas13c CRISPR-Cas system repeat sequence or fragment thereof;Cas13d CRISPR-Cas system repeat sequence or fragment thereof;Cas1 CRISPR-Cas system repeat sequence or fragment thereof;Cas1B CRISPR-Cas system repeat sequence or fragment thereof;Cas2 CRISPR-Cas system repeat sequence or fragment thereof;Cas3 CRISPR-Cas system repeat sequence or fragment thereof;Cas3' CRISPR-Cas system repeat sequence or fragment thereof;Cas3'' CRISPR-Cas system repeat sequence or fragment thereof;Cas4 CRISPR-Cas system repeat sequence or fragment thereof;Cas5 CRISPR-Cas system repeat sequence or fragment thereof;Cas6 CRISPR-Cas system repeat sequence or fragment thereof;Cas7 CRISPR-Cas system repeat sequence or fragment thereof;Cas8 CRISPR-Cas system repeat sequence or fragment thereof;Cas10 CRISPR-Cas system repeat sequence or fragment thereof;Csy1 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csy2 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csy3 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Cse1 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Cse2 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csc1 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csc2 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csa5 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csn2 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csm2 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csm3 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csm5 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Csm6 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Cmr1 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Cmr3 a repeat sequence of a CRISPR-Cas system or a fragment thereof;Cmr4 repeat sequence of the CRISPR-Cas system or a fragment thereof;Cmr5 repeat sequence of the CRISPR-Cas system or a fragment thereof;Cmr6 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csb1 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csb2 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csb3 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csx10 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csx14 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csx15 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csx16 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csx17 repeat sequence of the CRISPR-Cas system or a fragment thereof;CsaX repeat sequence of the CRISPR-Cas system or a fragment thereof;Csx1 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csx3 repeat sequence of the CRISPR-Cas system or a fragment thereof;Csf1 a repeat sequence of a CRISPR-Cas system or a fragment thereof; a repeat sequence of a Csf2 CRISPR-Cas system or a fragment thereof; a repeat sequence of a Csf3 CRISPR-Cas system or a fragment thereof; a repeat sequence of a Csf4 CRISPR-Cas system or a fragment thereof; and a repeat sequence of a Csf5 CRISPR-Cas system or a fragment thereof.
[0056] In some embodiments, a guide RNA expressed from a synthetic snRNA promoter described herein may contain two or more crRNA sequences (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more crRNA sequences). In some embodiments, a guide RNA comprises one or more tRNAs separating the crRNA sequences (e.g., crRNA-tRNA-crRNA, e.g., crRNA-tRNA-crRNA-tRNA-crRNA-tRNA-crRNA-tRNA-crRNA).
[0057] In some embodiments, a guideRNA array expressed from a synthetic snRNA promoter described herein may contain two or more tracrRNA sequences (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more tracrRNA sequences). In some embodiments, a guideRNA array contains one or more tRNAs separating the tracrRNA sequences (e.g., tracrRNA-tRNA-tracrRNA, e.g., tracrRNA-tRNA-tracrRNA-tRNA-tracrRNA-tRNA-tracrRNA-tRNA-tracrRNA, etc.).
[0058] In some embodiments, a guide RNA array expressed from a synthetic snRNA promoter described herein may comprise two or more crRNA-tracrRNA (and / or tracrRNA-crRNA) sequences (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more crRNA-tracrRNA (and / or tracrRNA-crRNA) sequences) (e.g., crRNA-tracrRNA-crRNA, e.g., tracrRNA-crRNA-crRNA-tracrRNA-crRNA-tracrRNA-crRNA-tracrRNA-crRNA-tracrRNA, e.g., tracrRNA-crRNA-tracrRNA-crRNA-tracrRNA-crRNA-tracrRNA-crRNA-tracrRNA, etc.). In some embodiments, the guide RNA array comprises one or more tRNAs that separate the crRNA and tracrRNA sequences (e.g., crRNA-tracrRNA-tRNA-crRNA-tracrRNA, e.g., tracrRNA-crRNA-tRNA-tracrRNA-crRNA, e.g., crRNA-tracrRNA-tRNA-crRNA-tracrRNA-tRNA-crRNA-tracrRNA-tRNA-crRNA-tracrRNA-tRNA-tracrRNA, e.g., tracrRNA-crRNA-tRNA-tracrRNA-crRNA-tRNA-tracrRNA-crRNA-tRNA-tracrRNA-crRNA-tRNA-tracrRNA-crRNA, etc.).
[0059] In some embodiments, the guide RNA expressed from the synthetic snRNA promoter described herein may further comprise an aptamer sequence (e.g., an MS2 aptamer). In some embodiments, the aptamer sequence recruits a deaminase. In some embodiments, the aptamer sequence recruits a reverse transcriptase. In some embodiments, the guide RNA may comprise up to one or two or more aptamers (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more aptamers).
[0060] In some embodiments, the guide RNA expressed from the synthetic snRNA promoters described herein may further comprise an RNA template for reverse transcriptase. In some embodiments, the synthetic snRNA promoters described herein are operably linked to a prime edited guide RNA ("PegRNA").
[0061] Cas9 is a class 2 CRISPR effector protein. Class 2 CRISPR-Cas systems rely on a single component effector protein, such as Cas9, where one gRNA-bound Cas protein recognizes and cleaves the target sequence. Cas9 recognizes a G-rich protospacer adjacent motif (PAM) that is 3' to its guide RNA binding site. In some embodiments, the CRISPR Cas9 protein can be, for example, a Cas9 protein from the genus Streptococcus (e.g., S. pyogenes, S. thermophilus), Lactobacillus, Bifidobacterium, Kandleria, Leuconostoc, Oenococcus, Pediococcus, Weissella, and / or Olsenella. An additional family of class 2 Cas effector proteins has been discovered: Cpf1 (also known as Cas12a), C2c1, CasX and CasY (Burstein et al., Nature, 542:237-241, 2017).
[0062] Cas12a belongs to class 2, type V CRISPR systems and utilizes one RNA-guided endonuclease that does not contain tracrRNA. The Cas12a system recognizes T-rich protospacer adjacent motifs (PAMs). The T-rich PAMs allow for application in genome editing, especially in organisms with AT-rich genomes or in AT-rich target regions. The CRISPR array is processed into short mature crRNAs that are 42-44 nucleotides long. Each mature crRNA begins with a 19-nucleotide direct repeat scaffold followed by a 23-25 nucleotide spacer sequence. This arrangement of crRNAs contrasts with that of type II CRISPR-Cas systems, where the mature crRNA begins with a 20-24 nucleotide spacer sequence followed by a ~22 nucleotide direct repeat scaffold (Zetsche et al., Cell 163:759-771, 2015). Cas12a produces staggered cuts when cleaving double-stranded DNA molecules. This is in contrast to blunt end cleavage (such as that generated by Cas9). An example of a Cas12a coding sequence that includes a transit peptide for delivery to the nucleus of a cell is set forth as SEQ ID NO: 12, which encodes a protein set forth as SEQ ID NO: 19.
[0063] CRISPR-Cas nucleases useful in the present invention include, but are not limited to, Cas9, C2c1, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Cas1, Cas1B, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Ca Examples of nucleases that may be mentioned include s10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4(dinG), Csf5 and / or Mad7 nuclease. In some embodiments, the CRISPR-Cas nuclease can be Cas9, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c effector proteins. In some embodiments, the CRISPR-Cas nucleases useful in the invention can include mutations in the nuclease active site (e.g., RuvC, HNH, e.g., the RuvC site of the Cas12a nuclease domain; e.g., the RuvC site and / or the HNH site of the Cas9 nuclease domain). CRISPR-Cas nucleases that have a mutation in the nuclease active site and therefore no longer contain nuclease activity are generally referred to as "dead", e.g., dCas, e.g., dCas9 or dCas12a. In some embodiments, a CRISPR-Cas nuclease domain or polypeptide that has a mutation in the nuclease active site can have impaired or reduced activity compared to the same CRISPR-Cas nuclease, e.g., nickase, e.g., Cas9 nickase, Cas12a nickase, without the mutation.Recently, CRISPR-associated transposases (CASTs) have been discovered and characterized. CASTs are composed of the Tn7-like transposase subunits tnsB, tnsC, and tniQ, and the VK-type CRISPR effector Cas12k catalyzes site-directed DNA transposition. Cas12k forms a complex with partially complementary non-coding RNA species crRNA and tracrRNA, and the tripartite ribonucleoprotein (RNP) complex recognizes chromosomal transposition sites based on the presence of protospacer adjacent motifs (PAMs) and complementarity between the variable portion of the crRNA and the target DNA. The associated transposases, tnsB, tnsC, and tniQ, recognize transposons by their conserved "left-end" (LE) and "right-end" (RE) boundaries and insert the transposon into chromosomal sites near the target sequence recognized by Cas12k (preferentially between TA dinucleotides). Two homologous CAST systems unique to the cyanobacterial species Scytonema hofmanni (UTEX B2349) and Anabaena cylindrica (PCC7122) have been shown to function for transposition in E. coli (Strecker et al., Science 365(6448):48-53, 2019).
[0064] gRNA Expression Strategy The present disclosure provides, in certain embodiments, novel combinations of synthetic snRNA promoters (and functional fragments thereof) with DNA sequences encoding one or more guide nucleic acid molecules. The guide nucleic acid molecules provided herein can be DNA, RNA, or a combination of DNA and RNA.
[0065] In one embodiment, a synthetic snRNA promoter is operably linked to one or more gRNA coding sequences to constitutively express gRNA(s) in transformed cells. This may be desirable, for example, in some embodiments, when the resulting gRNA transcript is retained in the nucleus and thus optimally located in the cell to guide nuclear processes. This may also be desirable, for example, in some embodiments, when the activity of the CRISPR system is low or the frequency of finding and cleaving the target site is low. In some embodiments, it may also be desirable when the promoter of a particular cell type, such as the germline, is not known for a given species of interest.
[0066] In another embodiment, fragments of synthetic snRNA promoters containing the necessary cis elements to drive transcription can be used to express one or more gRNAs. The disclosed full-length synthetic snRNA promoters, shown as SEQ ID NOs: 1-5, are each approximately 500 bp in length. Constructs containing multiple synthetic snRNA promoters can be large as additional expression cassettes are cloned in parallel. This can lead to problems affecting stability and transformation. Thus, in certain instances, synthetic snRNA promoters may be shortened to reduce the size of the construct, as long as the shortened synthetic snRNA promoter retains the ability to drive transcription of the gRNA. Examples of such shortened synthetic snRNA promoters are shown as SEQ ID NOs: 6-10.
[0067] Multiple synthetic snRNA promoters (or functional fragments thereof) with different sequences can be utilized to minimize construct stability issues typically associated with sequence repeats and also to facilitate stacking of multiple gRNA cassettes in the same transformation construct.
[0068] In some embodiments, the synthetic snRNA promoters described herein (or functional fragments thereof) can drive the expression of one gRNA. In some embodiments, the synthetic snRNA promoters described herein (or functional fragments thereof) can drive the expression of an array of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more gRNAs. Each individual guide sequence can target the same or different target sequences. This configuration is suitable for multiplexed gene manipulation (e.g., targeting multiple genes). Several strategies have been described in the art to facilitate the processing of individual guide RNAs from one transcript. In some embodiments, the synthetic snRNA promoters described herein (or functional fragments thereof) can be used to drive gRNA arrays in which the expression cassette comprises at least two or more gRNAs separated by one or more tRNA cleavage sequences (US20190330647). The tRNA cleavage sequences include any sequence and / or structural motif that actively interacts with and is cleaved by the endogenous tRNA system of the cell, such as RNaseP, RNaseZ, and RNaseE (bacteria). This can include structural recognition elements such as acceptor stems, D-loop arms, T Psi C-loops, and specific sequence motifs. In another embodiment, the synthetic snRNA promoters described herein (or functional fragments thereof) can be used to drive gRNA arrays that comprise two or more gRNAs separated by one or more ribozyme cleavage sites (Tang et al., Mol. Plant 9:1088-1091, 2016). In another embodiment, the synthetic snRNA promoters described herein (or functional fragments thereof) can be used to drive gRNA arrays comprising two or more gRNA arrays separated by one or more Csy4 ribonuclease recognition sites (Tsai et al., Nat. Biotechnol, 32(6):569-576, 2014).
[0069] In some embodiments, the synthetic snRNA promoters described herein (or functional fragments thereof) can be used to drive expression of a prime editing gRNA (PEgRNA). Prime editing is a genome editing method that uses a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) working in conjunction with a polymerase (e.g., in the form of a fusion protein or provided in trans with the napDNAbp) to directly write new genetic information at a target DNA site, where the prime editing system is programmed with a specialized prime editing (PE) guide RNA ("PEgRNA") that specifies the target site and serves as a template for the synthesis of the desired edit in the form of a replacement DNA strand with an engineered extension (either DNA or RNA) on the guide RNA (e.g., at the 5' or 3' end or in an internal portion of the guide RNA) (WO2020191248). In some embodiments, a synthetic snRNA promoter (or a functional fragment thereof) described herein is used to drive expression of a PEgRNA comprising a guide RNA and at least one nucleic acid extension arm comprising a DNA synthesis template, where the nucleic acid extension arm is located at the 3' or 5' end of the guide RNA.
[0070] In another embodiment, the synthetic snRNA promoter (or functional fragment thereof) described herein may be used to drive the expression of an enhanced gRNA that further comprises an RNA mobilization sequence that allows RNA to move between cells. The RNA mobilization sequence may be a sequence derived from a plant gene such as Flowering Time (FT) gene, BEL5, GAI, tRNA-like motif, or LeT6 (WO2021041001).
[0071] In other embodiments, the synthetic snRNA promoters described herein (or functional fragments thereof) may be used to drive expression of CRISPR RNA (crRNA), mature crRNA, precursor crRNA, crRNA fragments, trans-activating crRNA (tracrRNA), or tracrRNA fragments.
[0072] In some embodiments, the synthetic snRNA promoters described herein (or functional fragments thereof) may be used to drive gRNAs that are compatible with other forms of CRISPR-mediated gene editing, such as base editing (Komor et al., Nature 533, 420-424, 2016; Gaudelli et.al., Nature 551:464-471, 2017; Komor et.al., Science Advances Vol 3:No.8, 2017; and Rees et.al., Nat Rev Genet.19(12):770-788, 2018).
[0073] In some embodiments, the synthetic snRNA promoters described herein (or functional fragments thereof) can be used to drive gRNAs compatible with CRISPR-associated transposase systems (CAST), such as those derived from Scytonema hofmanni (ShCAST) and Anabaena cylindrica (AcCAST) (Strecker et al., Science 365(6448):48-53, 2019).
[0074] In some embodiments, the synthetic snRNA promoters described herein (or functional fragments thereof) may be used to drive expression of one or more non-protein-coding RNAs (npcRNAs). Non-limiting examples of non-protein-coding RNAs include microRNAs (miRNAs), miRNA precursors, small interfering RNAs (siRNAs), small RNAs (22-26 nt long) and their encoding precursors, heterochromatic siRNAs (hc-siRNAs), Piwi-interacting RNAs (piRNAs), hairpin double-stranded RNAs (hairpin dsRNAs), trans-acting siRNAs (ta-siRNAs), naturally occurring antisense siRNAs (nat-siRNAs), and tRNAs.
[0075] Strategies for expressing CRISPR class 2, type II or type V related genes The present disclosure provides novel synthetic snRNA promoters (and functional fragments thereof) for use in sequence-specific CRISPR-mediated cleavage for molecular breeding, for example by providing transcription of gRNAs containing spacer sequences that are used to target any site for endonuclease cleavage by at least one Cas protein. In certain embodiments, the target site is a genomic target site. In some embodiments, the genomic target site is native or transgenic. Furthermore, the CRISPR system can be customized to catalyze cleavage at one or more genomic target sites.
[0076] One aspect of the present disclosure is to introduce into a plant cell an expression construct comprising one or more cassettes encoding a synthetic snRNA promoter (or functional fragments thereof) described herein operably linked to a nucleotide sequence encoding one or more gRNAs, such as a copy of a spacer sequence complementary to a target site (e.g., a genomic target site), and an expression construct encoding a type I, II, III, IV, V, or VI CRISPR-associated protein, to modify the plant cell in such a way that the plant cell or a plant comprised of such a cell subsequently exhibits a beneficial trait. In one non-limiting example, the trait is a trait such as improved yield, resistance to biotic or abiotic stress, herbicide resistance, or other improvement in agricultural practices. The ability to generate such plant cells derived therefrom is dependent on the introduction of a CRISPR system using the transformation constructs and cassettes described herein.
[0077] The expression constructs encoding CRISPR-associated proteins may include a promoter. In certain embodiments, the promoter is a constitutive promoter, a tissue-specific promoter, a developmentally regulated promoter, or a cell cycle regulated promoter. Particular contemplated promoters include, among others, promoters that are expressed only in germline cells or germ cells. Such developmentally regulated promoters have the advantage that the activity of the CRISPR system is restricted to only cells in which the CRISPR-associated proteins are expressed. In some embodiments, CRISPR-mediated genetic modification (e.g., chromosomal or episomal dsDNA breaks) is restricted to only cells involved in the transmission of genome from one generation to the next. This may be useful in cases where widespread expression of the CRISPR system is genotoxic or has other undesirable effects. Examples of such promoters include promoters of genes encoding DNA ligases, recombinases, replicases, etc.
[0078] In some embodiments, the DNA construct described herein comprises one or more synthetic snRNA promoters or fragments thereof that express one or more gRNA-encoding DNA sequences at high levels.The DNA construct expressing gRNA can be particularly useful to guide CRISPR class 2, type II, or type V-associated proteins with endonuclease activity to specific genomic sequences, so that specific genomic sequences are cut to generate double-strand breaks, which are repaired by double-strand break repair pathways (which can include, for example, non-homologous end joining, microhomology-mediated end joining (MMEJ) homologous recombination, synthesis-dependent strand annealing (SDSA), single-strand annealing (SSA), or combinations thereof, thereby destroying native locus.
[0079] In one embodiment, the CRISPR system includes at least one Type I, II, III, IV, V, or VI CRISPR-associated protein and one gRNA that includes a copy of a spacer sequence complementary to an endogenous target site.
[0080] In some embodiments, the CRISPR system can include a catalytically inactive CRISPR endonuclease. Such an endonuclease will contain a domain that retains its ability to bind to a target nucleic acid but has a reduced or eliminated ability to cleave a nucleic acid molecule, as compared to a control nuclease. In some embodiments, the catalytically inactive nuclease is a catalytically inactive Cas9. In some embodiments, the catalytically inactive Cas9 generates a nick in one of the target DNA strands. In some embodiments, the catalytically inactive Cas9, known as dead Cas9 (dCas9), lacks all nuclease activity. In some embodiments, the catalytically inactive nuclease is a catalytically inactive Cas12a. In some embodiments, the catalytically inactive Cas12a generates a nick in one of the target DNA strands. In some embodiments, the catalytically inactive Cas12a, known as dead Cas12a (dCas12a), lacks all DNase activity.
[0081] The present disclosure also provides the use of CRISPR-mediated double-stranded DNA breaks to genetically modify the expression and / or activity of a gene or gene product of interest in a tissue- or cell-type-specific manner, where the nucleic acid of interest may be endogenous or transgenic in nature, to improve productivity or provide another beneficial trait.Thus, in one embodiment, the CRISPR system is engineered to mediate disruption at a specific site in a gene of interest.A gene of interest includes a gene whose expression level / protein activity is desired to be modified.These DNA break events may be within the coding sequence or within a regulatory element within the gene.
[0082] The present disclosure provides for the introduction of components of the CRISPR system (e.g., CRISPR-associated proteins and their cognate gRNAs) into cells. Examples of CRISPR-associated proteins include natural and engineered (e.g., modified, such as codon redesigned) nucleotide sequences encoding polypeptides having nuclease activity, such as Cas9 from Streptococcus pyogene, Streptococcus thermophilus, or Bradyrhizobium genus; Cpf1 (also known as Cas12a) from Francisella novicida (FnCpf1), Prevotella species, Acidaminococcus species BV3L6, and Lachnospiraceae bacterium ND2006 (LbCpf1); C2c1 from Alicyclobacillus acidoterrestris, Bacilli species, Verrucomicrobia species, α-proteobacteria, or δ-proteobacteria; CasX from Planctomycetes and δ-proteobacteria; or Candidatus Kerfeldbacteria, Candidatus Vogelbacteria, Candidatus Parcubacteria, or Candidatus CasY from Komeilibacteria.
[0083] In certain embodiments, the codon-redesigned FnCpf1 and LbCpf1 nucleotide sequences and expression cassettes comprise the recombinant nucleic acid sequences disclosed in US2020 / 0080096, the contents and disclosure of which are incorporated herein by reference.
[0084] Catalytically active CRISPR-associated genes (e.g., Cas9 endonuclease, C2c1 endonuclease, CasX endonuclease, CasY endonuclease, or Cpf1 endonuclease) can be introduced into or produced by the target cell. As disclosed herein, various methods can be used to accomplish this.
[0085] Transient expression of CRISPR In some embodiments, one or more expression cassettes encoding gRNA and / or CRISPR-associated protein components of type I, type II, type III, type IV, type V, or type VI CRISPR-Cas systems are transiently introduced into the cells. In certain embodiments, the introduced one or more expression cassettes encoding gRNA and / or CRISPR-associated proteins are provided in sufficient amounts to modify the cells, but do not persist after the intended time period or after one or more cell divisions. In such embodiments, no additional steps are required to remove or isolate the one or more expression cassettes encoding gRNA and / or CRISPR-associated proteins from the modified cells. In still other embodiments of the present disclosure, double-stranded DNA fragments are also transiently introduced into the cells along with one or more expression cassettes encoding gRNA and / or CRISPR-associated proteins. In such embodiments, the introduced double-stranded DNA fragments are provided in sufficient amounts to modify the cells, but do not persist after the intended time period or after one or more cell divisions.
[0086] In another embodiment, mRNA encoding CRISPR-associated protein is introduced into cell.In such an embodiment, the mRNA is translated to produce sufficient amount of CRISPR-associated protein to modify cell (in the presence of at least one gRNA whose expression is driven by synthetic snRNA promoter (or functional fragment thereof) described herein), but does not persist after intended time or after one or more cell divisions.In such an embodiment, no additional steps are required to remove or isolate CRISPR-associated protein from modified cell.
[0087] In one embodiment of the present disclosure, catalytically active CRISPR-associated proteins are prepared in vitro prior to introduction into plant cells, comprising at least one gRNA whose expression is driven by a synthetic snRNA promoter (or functional fragment thereof) as described herein. Methods for preparing CRISPR-associated proteins depend on their type and properties and will be known to those skilled in the art. For example, if the CRISPR-associated protein is large and monomeric, active forms of the CRISPR-associated protein can be produced via bacterial expression, in vitro translation, yeast cells, in insect cells, or by other protein production techniques known in the art. After expression, the CRISPR-associated protein is isolated, refolded as necessary, purified, and optionally treated to remove any purification tag, such as a His-tag. Once crudely, partially, or more completely purified CRISPR-associated protein is obtained, the protein can be introduced into plant cells, for example, by electroporation, bombardment with particles coated with the CRISPR-associated protein, chemical transfection, or some other means of transport across the cell membrane. Methods for introducing proteins and nucleic acids into plant cells are well known in the art. Protein can also be delivered using nanoparticles that can deliver the combination of active protein and nucleic acid.When a sufficient amount of CRISPR-associated protein is introduced with suitable gRNA so that there is an effective amount of in vivo activity, the target sequence in genome is cut.It is also recognized that those skilled in the art can create CRISPR-associated protein that is inactive but is activated in vivo by native processing mechanism, and such CRISPR-associated protein is also contemplated by the present disclosure.
[0088] In another embodiment, constructs are created that transiently express gRNA and / or CRISPR-associated proteins and are introduced into plant cells. In yet another embodiment, the construct produces sufficient amounts of gRNA and / or CRISPR-associated proteins to effectively modify the desired episomal or genomic target site(s). For example, the present disclosure contemplates the preparation of constructs that can be delivered into plant cells by bombardment, electroporation, chemical transfection, or some other means. Such structures can have several useful properties. For example, in one embodiment, the constructs can be replicated in a bacterial host such that they can be produced and purified in sufficient amounts for transient expression. In another embodiment, the constructs can encode herbicide resistance genes that allow for selection of the construct in the host, or the constructs can also include expression cassettes to provide for expression of gRNA and / or CRISPR-associated proteins in the plant. In further embodiments, the CRISPR-associated protein expression cassette may include a promoter region, a 5' untranslated region, an optional intron to aid expression, a multiple cloning site to allow for easy introduction of a DNA sequence encoding a CRISPR-associated protein, and a 3' UTR. In certain embodiments, the promoter of the CRISPR-associated protein expression cassette may be a constitutive promoter, a tissue-specific promoter, or other type of promoter that expresses in plant cells. In further embodiments, the gRNA expression cassette may include a snRNA promoter (or a functional fragment thereof) as described herein, a gRNA coding sequence, and a short poly-T region to terminate transcription. In some embodiments, the promoter in the gRNA expression cassette will be a synthetic snRNA promoter selected from SEQ ID NOs: 1-5. In some embodiments, the promoter in the gRNA expression cassette will be a synthetic snRNA promoter selected from SEQ ID NOs: 6-10. In some embodiments, it may be beneficial to include unique restriction sites at one or each end of the expression cassette, which allows for the production and isolation of a linear expression cassette that may then be free of other construction elements. In certain embodiments, the untranslated leader region may be a plant-derived untranslated region.When the expression cassette is transformed or transfected into a monocotyledonous or dicotyledonous plant cell, the use of an intron, which may be of plant origin, is contemplated.
[0089] In other embodiments, one or more elements in the construct contain a spacer complementary to a target site contained within an episomal or genomic sequence, which facilitates CRISPR-mediated modifications within the expression cassette to allow for the removal and / or insertion of elements such as promoters and transgenes.
[0090] In another approach, a bacterial or viral construct host can be used to introduce a transient expression construct into a plant cell. For example, Agrobacterium is one bacterial construct that can be used to introduce a transient expression construct into a host plant cell. When using a bacterial, viral, or other construct host system, the transient expression construct is contained within the host construct system. For example, when using an Agrobacterium host system, the transient expression cassette is flanked by one or more T-DNA borders and cloned into a binary construct. Many such construct systems have been identified in the art (reviewed in Hellens et al., 2000).
[0091] In embodiments where one or more of the gRNA and / or CRISPR-associated protein components of the CRISPR system are transiently introduced in sufficient amounts to modify the plant cells, methods of selecting modified plant cells can be used. In one such method, a second nucleic acid molecule containing a selectable marker is co-introduced with the transient gRNA and / or CRISPR-associated protein. In this embodiment, the co-introduced marker can be part of a molecular strategy to introduce a marker at a target site. For example, the co-introduced marker can be used to disrupt a target gene by inserting between genomic target sites. In another embodiment, the co-introduced nucleic acid can be used to produce a visual marker protein so that transfected cells can be isolated by cell sorting or some other means. In yet another embodiment, the co-introduced marker can be randomly integrated or oriented via the second gRNA:CRISPR-associated protein complex to integrate at a site independent of the primary genomic target site. In yet another embodiment, the co-introduced molecules can target specific loci via double-strand break repair pathways, which may include, for example, non-homologous end joining (NHEJ), microhomology-mediated end joining (MMEJ), homologous recombination, synthesis-dependent strand annealing (SDSA), single-strand annealing (SSA), or combinations thereof, at the genomic target site(s). In the above embodiments, the co-introduced markers can be used to identify or select cells that are likely to have been exposed to gRNA and / or CRISPR-associated proteins and therefore likely to have been modified by CRISPR.
[0092] Stable expression of CRISPR In another embodiment, one or more expression constructs encoding one or more components of the CRISPR system (e.g., a CRISPR-associated protein and its cognate gRNA) are stably transformed into plant cells. In this embodiment, the design of the transformation construct allows flexibility as to when and under what conditions the gRNA and / or CRISPR-associated protein are expressed. Additionally, the transformation construct can be designed to include a selectable or visible marker that provides a means to isolate or efficiently select cell lines that contain one or more expression constructs encoding one or more components of the CRISPR system and / or that have been modified by the CRISPR system.
[0093] Cell transformation systems have been described in the art, including various transformation constructs. For example, for plant transformation, the two main methods include Agrobacterium-mediated transformation and particle bombardment-mediated (e.g., biolistics) transformation. In either case, the nucleotide sequences encoding the components of the CRISPR system are introduced via one or more expression cassettes. In further embodiments, the CRISPR-associated protein expression cassette may include a promoter region, a 5' untranslated region, an optional intron to aid expression, a multiple cloning site that allows for easy introduction of the DNA sequence encoding the CRISPR-associated protein, and a 3' UTR. In certain embodiments, the promoter of the CRISPR-associated protein expression cassette may be a constitutive promoter, a tissue-specific promoter, a developmentally regulated promoter, a cell cycle regulated promoter, or a germline-specific promoter. In further embodiments, the gRNA expression cassette may include a snRNA promoter (or a functional fragment thereof) as described herein, a gRNA coding sequence, and a short poly-T region that terminates transcription. In certain embodiments, the promoter in the gRNA expression cassette will be a synthetic snRNA promoter selected from SEQ ID NOs: 1-5. In some embodiments, the promoter in the gRNA expression cassette will be a synthetic snRNA promoter selected from SEQ ID NOs: 6-10.
[0094] In the case of particle bombardment or protoplast transformation, the expression cassette may be an isolated linear fragment or may be part of a larger construct that may include bacterial replication elements, bacterial selection markers, or other elements. One or more gRNA and / or CRISPR-associated protein expression cassette(s) may be physically linked to the marker cassette or may be mixed with a second nucleic acid molecule encoding the marker cassette. In some embodiments, the marker cassette is composed of elements necessary to express a visual or selectable marker that allows efficient selection of transformed cells. In the case of Agrobacterium-mediated transformation, one or more expression cassettes may be included within a binary construct, adjacent to or between adjacent T-DNA borders. In another embodiment, one or more expression cassettes may be outside the T-DNA. The presence of one or more expression cassettes within a cell may be manipulated by positive or negative selection regime(s). Furthermore, the selectable marker cassette may be within or adjacent to the same T-DNA border, or anywhere else within the second T-DNA on a binary construct (e.g., a 2T-DNA system).
[0095] In some embodiments, the cells that are modified by the CRISPR system, either transiently or stably, are maintained together with unmodified cells. The cells can be subdivided into independent clonally derived lines or used to regenerate independently derived plants. Individual plants or clonal populations regenerated from such cells can be used to generate independently derived lines. At any of these stages, molecular assays can be used to screen the modified cells, plants, or lines. The modified cells, plants, or lines continue to propagate, while the unmodified cells, plants, or lines are discarded. In some embodiments, the presence of an active CRISPR system in the cells is essential to ensure the efficiency of the entire process.
[0096] Transformation method Methods for transforming or transfecting cells are well known in the art. Plant transformation methods using Agrobacterium or DNA-coated particles are well known in the art and are incorporated herein. Methods suitable for transforming host cells for use in the present disclosure may include virtually any method that can introduce DNA into cells, such as Agrobacterium-mediated transformation (U.S. Pat. Nos. 5,563,055; 5,591,616; 5,693,512; 5,824,877; 5,981,840; and 6,384,301), and acceleration of DNA-coated particles (U.S. Pat. Nos. 5,015,580; 5,550,318; 5,538,880; 6,160,208; 6,399,861; and 6,403,865). By applying such techniques, cells of virtually any species can be stably transformed.
[0097] Various methods have been described for selecting transformed cells. For example, drug resistance markers such as neomycin phosphotransferase protein can be utilized to confer resistance to kanamycin, or 5-enolpyruvylshikimate phosphate synthase can be used to confer resistance to glyphosate. In another embodiment, carotenoid synthase is used to create a visually identifiable orange pigment. Each of these three exemplary approaches can be effectively used to isolate cells or plants or tissues thereof that have been transformed and / or modified by CRISPR.
[0098] When the nucleic acid sequence encoding a selectable or screenable marker is inserted into a genome target site, the marker can be used to detect the presence or absence of CRISPR or its activity.This can be useful when cells are modified by CRISPR and it is desired to recover genetically modified cells that no longer contain CRISPR, or plants regenerated from such modified cells.In other embodiments, the marker can be purposefully designed to be integrated into a genome target site, allowing it to be used to track modified cells independently of CRISPR.The marker can be a gene that provides a visually detectable phenotype in seeds, etc., allowing seeds that carry or lack CRISPR expression cassettes to be quickly identified.
[0099] The present disclosure provides a means to regenerate plants from cells that have a repaired double-stranded break within the genomic target site. The regeneration can be used to propagate additional plants.
[0100] The present disclosure further provides novel plant transformation constructs and expression cassettes, including synthetic snRNA promoters and their combination with CRISPR-associated gene(s) and gRNA / expression cassettes. The present disclosure further provides methods for obtaining specifically modified plant cells, whole plants, and seeds or embryos using CRISPR-mediated cleavage. The present disclosure also relates to novel plant cells that include CRISPR-associated Cas endonuclease expression constructs and gRNA expression cassettes.
[0101] Targeting using blunt-ended oligonucleotides In certain embodiments, a CRISPR system (e.g., a CRISPR / Cas9 system or a CRISPR / Cas12a system) can be utilized to target the 5' insertion of a blunt-ended double-stranded DNA fragment to a genomic target site of interest. In some embodiments, CRISPR-mediated endonuclease activity can introduce a double-strand break (DSB) at a selected genomic target site, and DNA repair, such as microhomology-driven non-homologous end joining DNA repair, results in the insertion of a blunt-ended double-stranded DNA fragment into the DSB. In some embodiments, the blunt-ended double-stranded DNA fragment can be designed to have 1-10 bp of microhomology at both the 5' and 3' ends of the DNA fragment, corresponding to the 5' and 3' flanking sequences at the cleavage site of the genomic target site.
[0102] Use of CRISPR systems in molecular breeding In some embodiments, knowledge of genome is utilized for targeted genetic modification of genome. At least one gRNA can be designed to target at least one region of genome and destroy the region from genome. This aspect of the present disclosure can be particularly useful for genetic modification. The resulting plant can have modified phenotype or other characteristics depending on the modified gene(s). Pre-characterized mutant alleles or introduced transgenes can be targeted for CRISPR-mediated modification, thereby allowing for the creation of improved mutants or transgenic lines.
[0103] In another embodiment, the gene targeted for deletion or disruption can be a transgene previously introduced into the target plant or plant cell. This has the advantage of allowing the introduction of an improved version of the transgene or the disruption of a sequence encoding a selectable marker. In yet another embodiment, the gene targeted for disruption by the CRISPR system is at least one transgene that is introduced on the same construct or expression cassette as the other transgene(s) of interest and is present at the same locus as another transgene. Those skilled in the art will appreciate that this type of CRISPR-mediated modification may result in the deletion or insertion of additional sequences. Thus, in certain embodiments, it may be preferable to generate multiple plants or plant cells in which deletions have occurred and screen such plants or plant cells using standard techniques to identify specific plants or plant cells with minimal genome alterations following CRISPR-mediated modification. Such screening may utilize genotypic and / or phenotypic information. In such an embodiment, a specific transgene may be disrupted while leaving the remaining transgene(s) intact. This avoids the need to create new transgenic lines that contain the desired transgene without the undesired transgene.
[0104] In another aspect, the present disclosure includes a method for inserting a DNA fragment of interest into a specific site of a plant genome, the DNA fragment of interest being derived from the genome of the plant or being heterologous to the plant. This disclosure allows for the selection or targeting of a specific region of the genome for nucleic acid (e.g., transgene) stacking (e.g., megalocus). Thus, the target region of the genome may represent the linkage of at least one transgene to a haplotype of interest associated with at least one phenotypic trait, and may result in the generation of a linkage block to facilitate the stacking of transgenes and the integration of transgenic traits, and / or the generation of a linkage block while also allowing the integration of conventional traits.
[0105] Use of CRISPR systems for trait integration Directed insertion of a DNA fragment of interest into at least one genomic target site via CRISPR-mediated cleavage allows for the targeted integration of multiple nucleic acids of interest (e.g., trait stacks) to be added to the genome of a plant, either at the same site or at different sites. Sites for targeted integration can be selected based on knowledge of the underlying breeding value, the performance of the transgene at that location, the underlying recombination rate at that location, existing transgenes at that junction block, or other factors. Once assembled, stacked plants can be used as trait donors for crosses with germplasm being advanced in the breeding pipeline, or advanced directly in the breeding pipeline.
[0106] The present disclosure includes a method for inserting at least one nucleic acid of interest into at least one site, the nucleic acid of interest being derived from the genome of a plant, such as a QTL or allele, or of transgenic origin. Thus, the target region of the genome may represent the linkage of at least one transgene to a haplotype of interest associated with at least one phenotypic trait (as described in US Patent Application Publication No. 2006 / 0282911), stacking of transgenes and generation of linkage blocks to facilitate integration of transgenic traits, stacking of QTLs or haplotypes and generation of linkage blocks to facilitate integration of conventional traits, etc.
[0107] In another embodiment of the present disclosure, multiple unique gRNAs can be used to modify multiple loci within one linkage block contained on one chromosome by utilizing knowledge of genomic sequence information and the ability to design custom gRNAs as described in the art. A gRNA is designed or engineered as needed to be specific to or directed to a genomic target site upstream of the locus containing the non-target allele. A second gRNA is also designed or engineered to be specific to or directed to a genomic target site downstream of the target locus containing the non-target allele. The gRNA can be designed to complement genomic regions that are not homologous to the non-target locus containing the target allele. Either gRNA can be introduced into the cell using one of the methods described above.
[0108] The ability to perform targeted integration is dependent on the action of gRNA:CRISPR associated proteins. This advantage provides a method for engineering a plant of interest, such as a plant or cell, with at least one genomic modification.
[0109] The custom gRNA is utilized in the CRISPR system to generate at least one trait donor to create custom genome modification events, which is then crossed with at least one second plant of interest, such as a plant. Here, CRISPR-associated protein delivery can be combined with the gRNA of interest for genome editing. In other embodiments, one or more plants of interest are directly transformed with the CRISPR system and at least one double-stranded DNA fragment of interest for directional insertion. It is recognized that this method can be performed in various cell, tissue, and developmental types, such as plant gametes. It is further anticipated that one or more of the elements described herein can be combined with the use of promoters specific to certain cells, tissues, plant parts, and / or developmental stages, such as meiosis-specific promoters.
[0110] Furthermore, the present disclosure contemplates targeting transgenic elements already present in genome for deletion or disruption.This allows, for example, the introduction of improved transgenes, removal of selectable markers.In yet another embodiment, the gene targeted for disruption by CRISPR-mediated cleavage is at least one transgene that is introduced on the same construct or expression cassette as other transgene(s) of interest and is present at the same locus as another transgene.
[0111] In one aspect, the disclosure provides a method for modifying a locus of interest in a cell, the method comprising: (a) identifying at least one locus of interest within a DNA sequence; (b) introducing into the cell an expression cassette comprising a synthetic snRNA promoter selected from SEQ ID NOs: 1-10 operably linked to a nucleotide sequence encoding a gRNA, and an expression cassette comprising a plant-expressible promoter operably linked to a nucleic acid sequence encoding a CRISPR-associated protein, where the gRNA and / or the CRISPR-associated protein are expressed transiently or stably; (d) assaying the cell for CRISPR-mediated modifications in DNA constituting or adjacent to the locus of interest; and (e) identifying the cell, or a progeny thereof, as comprising a modification at the locus of interest.
[0112] Another aspect provides a method for modifying multiple loci of interest in a cell, the method comprising: (a) identifying multiple loci of interest in a genome; (b) introducing into at least one cell a multiple expression cassette comprising a synthetic snRNA promoter selected from SEQ ID NOs: 1-10 operably linked to a nucleotide sequence encoding a gRNA, wherein the synthetic snRNA promoters are independently selected, and at least one expression cassette comprises a plant-expressible promoter operably linked to a nucleic acid sequence encoding a CRISPR-associated protein according to the present disclosure, the cell comprising a genomic target site, and wherein the gRNA and the CRISPR-associated protein are transiently or stably expressed, creating a modified locus or multiple loci comprising at least one CRISPR-mediated cleavage event; (d) assaying the cells for CRISPR-mediated modifications in DNA comprising or adjacent to each locus of interest; and (e) identifying the cells or progeny thereof comprising the modified nucleotide sequence at the locus of interest.
[0113] The present disclosure further contemplates sequential modification of a locus of interest with two or more gRNAs and CRISPR-associated proteins according to the present disclosure. Such genes or other sequences added by the action of a first CRISPR-mediated genome modification can be retained, further modified, or removed by the action of a second CRISPR-mediated genome modification.
[0114] The present disclosure includes compositions and methods for modifying a genetic locus of interest in crop plants, such as corn (Zea mays subsp. mays), corn varieties (flower corn (Zea mays var. amylacea), popcorn (Zea mays var. everta), dent corn (Zea mays var. indentata), flint corn (Zea mays var. indurate), sweet corn (Zea mays var. saccharata and Zea mays var. rugose), waxy corn (Zea mays var. ceratina), amylomaize (Zea mays), podcorn (Zea mays var. tunicata Larranaga ex A. St. Hil.), striped maize (Zea mays var. japonica), soybean (Glycine max), cotton (Gossypium hirsutum; Gossypium sp.), peanut (Arachis hypogaea; barley (Hordeum vulgare); oats (Avena sativa); orchard grass (Dactylis glomerata); rice (Oryza sativa, including indica and japonica species); sorghum (Sorghum bicolor); sugarcane (Saccharum sp.); tall fescue (Festuca arundinacea); turfgrass species (e.g., species Agrostis stolonifera, Poa pratensis, Stenotaphrum secundatum); wheat (Triticum aestivum); alfalfa (Medicago sativa); members of the genus Brassica, including but not limited to canola (Brassica napus and Brassica rapa), members of the genus Brassica (e.g., species B. rapa subsp. chinensis), turnip (Brassica rapa var. glabra), Chinese cabbage (Brassica rapa var. glabra), and Chinese mustard greens (Brassica rapa var. glabra). subsp. parachinensis), oilseed rape (Brassica rapa subsp.oleifera), komatsuna (Brassica rapa subsp. perviridis, Chinese cabbage (Brassica rapa subsp. pekinensis), turnip (rapini) (Brassica rapa var. rapifera), tatsoi (Brassica rapa subsp. narinosa), turnip (turnip) (Brassica rapa subsp. rapa), yellow sarson (Brassica rapa subsp. trilocularis), Chinese cabbage, turnip (turnip), turnip (rapini), komatsuna (Brassica rapa (syn. Brassica campestris)), Mallorca cabbage (Brassica balearica), Abyssinian mustard or Abyssinian cabbage, elongated mustard (Brassica elongata), Mediterranean cabbage, used for the production of biodiesel (Brassica carinata), cabbage (Brassica fruticulosa), St Hilarion cabbage (Brassica hilarionis), Indian mustard, brown mustard and leaf mustard, Sarepta mustard (Brassica juncea), rapeseed, canola, rutabaga (swede, swede turnip, Swedish turnip) (Brassica napus), broadbeaked mustard (Brassica narinosa), black mustard (Brassica nigra), kale, cabbage, collard greens, broccoli, cauliflower, Chinese broccoli, Brussels sprouts, kohlrabi (Brassica oleracea), tender greens, komatsuna (Brassica perviridis), brown mustard (Brassica rupestris), seventop turnip (Brassica seticeps), Asian mustard (Brassica. tournefortii), broccoli (B.oleracea); peppers (e.g., species: black pepper, white and green pepper (Piper nigrum), cubeba (Piper cubeba), long pepper (Piper longum), long pepper (Piper retrofractum), Voatsiperifery (Piper borbonense), Ashanti pepper (Piper guineense), banana pepper, bell pepper, cayenne pepper, jalapeno, Florina pepper, (cultivars of Capsicum annuum), chili peppers (cultivars of Capsicum annuum, Capsicum frutescens, Capsicum chinense, Capsicum pubescens, and Capsicum baccatum), and datil pepper (cultivars of Capsicum chinense); legume species (e.g., broad bean or fava bean (Vicia faba), kidney beans; e.g., pinto beans, kidney beans, black beans, Appaloosa beans, and kidney beans, and many others (Phaseolus vulgaris), tepary beans (Phaseolus acutifolius), scarlet beans (Phaseolus coccineus), lima beans (Phaseolus lunatus), also known as P.dumosus, recognised as a separate species in 1995 (Phaseolus polyanthus), moth bean (Vigna aconitifolia), adzuki bean (Vigna angularis), black gram (Vigna mungo), mung bean (Vigna radiata), Bambara bean or ground-bean (Vigna subterranea), bamboo bean (Vigna umbellata), cowpea; also includes black-eyed pea, yardlong bean and others (Vigna unguiculata), chickpea (chickpea or garbanzo bean) (Cicer arietinum), pea (Pisum sativum), grass pea (Lathyrus sativus), Chinese pea (Lathyrus tuberosus), lentil (Lens culinaris), hyacinth bean (Lablab purpureus), winged bean (Psophocarpus tetragonolobus), pigeon pea (Cajanus cajan), mucuna pruriens, guar bean (Cyamopsis tetragonoloba), jack bean (Canavania ensiformis), sword bean (Canavalia gladiata), horse gram (Macrotyloma uniflorum), Lupinus mutabilis, lupin bean (Lupinus albus); members of the Cucurbitaceae family (e.g., squash, pumpkin, zucchini, some gourds (Cucurbita), bottle gourd (Lagenaria), watermelon (Citrullus, such as Citrullus lanatus and Citrullus colocynthis), cucumber (Cucumis sativus), various melons (Cucumis melo, Cucumis metuliferus); spinach (Spinacia oleracea); carrot (Daucus carota subsp. sativus); tomato (Solanum lycopersicum); onion (Allium cepa L.); radish (Raphanus raphanistrum subsp. sativus); potato (Solanum tuberosum); ornamental plants; oil crops such as soybean, canola, oilseed rape, oil palm, sunflower, olive, corn, cottonseed, peanut, linseed, safflower, and coconut.
[0115] Genomic modifications can include modified linkage blocks, linkage of two or more QTLs, disruption of the linkage of two or more QTLs, gene insertion, gene replacement, gene conversion, gene deletion or disruption, transgenic event selection, transgenic trait donor selection, transgene replacement, or targeted insertion of at least one nucleic acid of interest.
[0116] definition The definitions and methods provided define the present disclosure and guide those skilled in the art in practicing the present disclosure. Unless otherwise specified, terms should be understood according to conventional usage by those skilled in the relevant field. Definitions of common terms in molecular biology can also be found in Alberts et al., Molecular Biology of The Cell, 5th Edition, Garland Science Publishing, Inc.: New York, 2007; Rieger et al., Glossary of Genetics: Classical and Molecular, 5th edition, Springer-Verlag: New York, 1991; King et al., A Dictionary of Genetics, 6th ed., Oxford University Press: New York, 15 2247; and Lewin, Genes IX, Oxford University Press: New York, 2007. The nomenclature of DNA bases as defined in 37 CFR § 1.822 is used.
[0117] As used herein, "synthetic nucleotide sequence" or "artificial nucleotide sequence" refers to a nucleotide sequence that is not known to occur in nature or does not occur in nature. The gene regulatory element of the present invention comprises a synthetic nucleotide sequence. Preferably, the synthetic nucleotide sequence shares little or no extended homology with a natural sequence. Extended homology in this context generally refers to 100% sequence identity extending beyond a continuous sequence of about 25 nucleotides.
[0118] Reference in this application to an "isolated DNA molecule" or equivalent term or expression is intended to mean that the DNA molecule is present alone or in combination with other compositions, but not in its natural environment. For example, nucleic acid elements naturally found in the DNA of the genome of an organism, such as coding sequences, intron sequences, untranslated leader sequences, promoter sequences, transcription termination sequences, etc., are not considered to be "isolated" as long as the elements are in the genome of the organism and in the location in the genome where the elements are found in nature. However, each of these elements, and subportions of these elements, are "isolated" within the scope of this disclosure as long as the elements are not in the genome of the organism and in the location in the genome where the elements are found in nature. In one embodiment, the term "isolated" refers to a DNA molecule that is at least partially separated from some nucleic acids that normally flank the DNA molecule in its native or natural state. Thus, a DNA molecule that is fused to regulatory or coding sequences with which it is not normally associated, for example as a result of recombinant techniques, is considered to be isolated herein. Such molecules are considered isolated, i.e., they are not in their native state, even if they are integrated into the chromosome of a host cell or are present in a nucleic acid solution with other DNA molecules. For the purposes of this disclosure, any transgenic nucleotide sequence, i.e., a nucleotide sequence of DNA inserted into the genome of a plant or bacterial cell or present in an extrachromosomal construct, is considered to be an isolated nucleotide sequence, whether it is present in a plasmid or similar structure used to transform the cell, in the genome of the plant or bacteria, or in detectable amounts in tissues, progeny, biological samples, or commercial products derived from the plant or bacteria.
[0119] By "heterologous DNA molecule" is meant that the DNA molecule is heterologous to the polynucleotide sequence to which it is operably linked.
[0120] As used herein, the term "operably linked" refers to a first DNA molecule linked to a second DNA molecule, where the first and second DNA molecules are arranged such that the first DNA molecule affects the function of the second DNA molecule. The two DNA molecules may or may not be part of a single contiguous DNA molecule, and may or may not be adjacent. For example, a promoter is operably linked to a DNA molecule if it controls the transcription of the DNA molecule of interest in a cell. For example, a leader is operably linked to a DNA sequence if it can affect the transcription or translation of the DNA sequence.
[0121] As used herein, a "recombinant DNA molecule" is a DNA molecule that contains a combination of DNA molecules that would not occur together in nature without human intervention. For example, a recombinant DNA molecule may be a DNA molecule that is composed of at least two DNA molecules that are heterologous to each other, a DNA molecule that contains a DNA sequence that deviates from a naturally occurring DNA sequence, a DNA molecule that contains a synthetic DNA sequence, or a DNA molecule that has been incorporated into the DNA of a host cell by genetic transformation or gene editing.
[0122] As used herein, the term "sequence identity" refers to the degree to which two optimally aligned polynucleotide sequences or two optimally aligned polypeptide sequences are identical. An optimal sequence alignment is created by manually aligning two sequences, e.g., a reference sequence and another sequence, to maximize the number of nucleotide matches in the sequence alignment with appropriate internal nucleotide insertions, deletions, or gaps. As used herein, the term "reference sequence" refers to the DNA sequences provided as SEQ ID NOs: 1-10.
[0123] As used herein, the term "percent sequence identity" or "percent identity" or "% identity" refers to the percentage of identity multiplied by 100. The "percent identity" for a sequence optimally aligned to a reference sequence is the number of nucleotide matches in the optimal alignment divided by the total number of nucleotides in the reference sequence, e.g., the total number of nucleotides in the entire full-length reference sequence. Thus, one embodiment of the present invention provides a DNA molecule comprising a sequence that, when optimally aligned to a reference sequence provided herein as any of SEQ ID NOs: 1-10, has at least about 85 percent identity, at least about 86 percent identity, at least about 87 percent identity, at least about 88 percent identity, at least about 89 percent identity, at least about 90 percent identity, at least about 91 percent identity, at least about 92 percent identity, at least about 93 percent identity, at least about 94 percent identity, at least about 95 percent identity, at least about 96 percent identity, at least about 97 percent identity, at least about 98 percent identity, at least about 99 percent identity, or at least about 100 percent identity to the reference sequence. In more specific embodiments, a sequence having a percent identity to any of SEQ ID NOs: 1-10 may be defined as exhibiting the promoter activity possessed by the starting sequence from which it is derived. A sequence having a percent identity to any of SEQ ID NOs: 1-10 may further include a "minimal promoter," which provides a basal level of transcription and is composed of a TATA box or equivalent sequence for recognition and binding of the RNA polymerase III complex to initiate transcription. According to the present invention, a promoter, promoter variant, or promoter fragment may be analyzed for the presence of known promoter elements, i.e., DNA sequence features such as TATA boxes and other known transcription factor binding site motifs. Identification of such known promoter elements may be used by one of skill in the art to design promoter variants that have similar expression patterns as the original promoter.
[0124] The term "genome" encompasses not only chromosomal DNA found in the nucleus, but also organelle DNA found within subcellular components of a cell (eg, mitochondria or plastids).
[0125] As used herein, the term "genome editing" or "editing" refers to any modification of a nucleotide sequence in a site-specific manner. In this disclosure, genome editing techniques include the use of endonucleases, recombinases, transposases, helicases, and any combination thereof. In one aspect, the "modification" comprises hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the sequence-specific editing system comprises adenine deaminase. In one aspect, the "modification" comprises hydrolytic deamination of adenine or adenosine. In one aspect, the "modification" comprises hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In one aspect, the "modification" comprises an insertion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 25, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 3000, at least 4000, at least 5000, or at least 10,000 nucleotides. In another embodiment, the "modification" comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 25, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 3000, at least 4000, at least 5000, or at least 10,000 nucleotides.In further embodiments, the "modification" comprises an inversion of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 25, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 3000, at least 4000, at least 5000, or at least 10,000 nucleotides. In yet another embodiment, the "modification" comprises the substitution of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 25, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 3000, at least 4000, at least 5000, or at least 10,000 nucleotides. In yet another aspect, the "modification" comprises a duplication of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 25, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 3000, at least 4000, at least 5000, or at least 10,000 nucleotides. In some embodiments, the "modification" comprises a substitution of "A" with "C", "G" or "T" in the nucleic acid sequence. In some embodiments, the "modification" comprises a substitution of "C" with "A", "G" or "T" in the nucleic acid sequence. In some embodiments, the "modification" comprises a substitution of "G" with "A", "C" or "T" in the nucleic acid sequence. In some embodiments, the "modification" comprises a substitution of "T" with "A", "C" or "G" in the nucleic acid sequence. In some embodiments, the "modification" comprises the substitution of a "C" for a "U" in a nucleic acid sequence.In some embodiments, the "modification" comprises the substitution of "G" with "A" in the nucleic acid sequence. In some embodiments, the "modification" comprises the substitution of "A" with "G" in the nucleic acid sequence. In some embodiments, the "modification" comprises the substitution of "T" with "C" in the nucleic acid sequence.
[0126] As used herein, "target site" refers to a nucleotide sequence (e.g., protospacer and protospacer adjacent motif (PAM)) located within a DNA sequence selected for targeted modification to which a gRNA / CRISPR-associated protein system binds and / or exerts activity. A target site may be genetic or non-genic. A target site may be on a chromosome, episome, locus, or any other DNA molecule within the genome of a cell (chromosome, chloroplast, mitochondrial DNA, plasmid DNA, etc.). A target site may be an endogenous site in the genome of a cell, or a target site may be heterologous to the cell and thus not naturally occurring in the genome of the cell, or a target site may be found at a heterologous genomic location compared to where it occurs in nature.
[0127] As used herein, "genomic target site" refers to a target site located within the host genome selected for targeted modification (e.g., a protospacer and a protospacer adjacent motif (PAM)).
[0128] As used herein, "protospacer" refers to a short DNA sequence (12-40 bp) that can be targeted by the CRISPR system, guided by complementary base pairing with a spacer sequence in the gRNA.
[0129] As used herein, "microhomology" refers to the presence of the same short sequence of bases (1-10 bp) in different polynucleotide molecules.
[0130] As used herein, "codon optimized" refers to a polynucleotide sequence that has been modified to take advantage of the codon usage bias of a particular plant. The modified polynucleotide sequence still encodes the same or a substantially similar polypeptide as the original sequence, but uses codon nucleotide triplets that are found more frequently in the particular plant.
[0131] As used herein, "non-protein-coding RNA (npcRNA)" refers to non-coding RNA (ncRNA), which is a functional RNA molecule that is not translated into protein, either a precursor small non-protein-coding RNA or a fully processed non-protein-coding RNA.
[0132] As used herein, a "promoter" refers to a nucleic acid sequence located upstream or 5' of the translation start codon of a gene's open reading frame (or protein coding region) and is involved in the recognition and binding of RNA polymerase I, II, or III and other proteins (processing transcription factors) to initiate transcription. A "plant promoter" is a native or non-native promoter that functions in plant cells. A constitutive promoter functions in most or all tissues of a plant throughout plant development. A tissue-, organ-, or cell-specific promoter is expressed only or primarily in a particular tissue, organ, or cell type, respectively. A promoter may not be "specifically" expressed in a particular tissue, plant part, or cell type, but may exhibit "enhanced" expression, i.e., a higher level of expression, in one cell type, tissue, or plant part of a plant compared to other parts of the plant. A temporally regulated promoter functions only or primarily during a particular period of plant development or at a particular time of day, as in the case of genes associated with circadian rhythms, for example. Inducible promoters selectively express an operably linked DNA sequence in response to the presence of an endogenous or exogenous stimulus, for example, by a chemical compound (chemical inducer), or in response to environmental, hormonal, chemical, and / or developmental signals. Inducible or regulated promoters include, for example, promoters that are regulated by light, heat, stress, flood or drought, plant hormones, wounding, or chemicals, such as ethanol, jasmonates, salicylic acid, safeners, etc.
[0133] As used herein, "expression cassette" refers to a polynucleotide sequence comprising at least a first polynucleotide sequence capable of initiating transcription of an operably linked second polynucleotide sequence, and optionally a transcription termination sequence operably linked to the second polynucleotide sequence.
[0134] A palindrome is a nucleic acid sequence that is the same when read 5' to 3' on one strand as it is when read 3' to 5' on the complementary strand with which it forms a double helix. A nucleotide sequence is said to be a palindrome if it is equal to its reverse complement. Palindrome sequences can form hairpins.
[0135] In some embodiments, numbers expressing quantities of ingredients, properties such as molecular weights, reaction conditions, and the like, used to describe and claim certain embodiments of the present disclosure, shall be understood as being modified in some cases by the term "about". In some embodiments, the term "about" is used to indicate that a value includes the average standard deviation of the device or method used to determine that value. In some embodiments, the numerical parameters set forth in the written description and the accompanying claims are approximations that may vary depending on the desired properties sought to be obtained by a particular embodiment. In some embodiments, the numerical parameters should be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Notwithstanding that the numerical ranges and parameters setting forth the broad scope of some embodiments of the present disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as practicable. The numerical values presented in some embodiments of the present disclosure may contain certain errors necessarily resulting from the standard deviation found in their respective testing measurements. Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within that range. Unless otherwise indicated herein, each individual value is incorporated herein as if it were individually set forth herein.
[0136] In some embodiments, the terms "a," "an," and "the," and similar references, as used in the context of describing particular embodiments (particularly in certain contexts of the claims below), may be construed to cover both the singular and the plural, unless specifically indicated otherwise. In some embodiments, the term "or" as used herein, including the claims, is used to mean "and / or," unless expressly indicated to refer to alternatives only or where the alternatives are not mutually exclusive.
[0137] The terms "comprise," "have," and "include" are open-ended linking verbs. Any form or tense of one or more of these verbs, such as "comprises," "comprising," "has," "having," "includes," and "including," are also open-ended. For example, any method that "comprises," "has," or "includes" one or more steps is not limited to only retaining those one or more steps, but may also cover other unlisted steps. Similarly, any composition or device that "comprises," "has," or "includes" one or more features is not limited to only retaining those one or more features, but may also cover other unlisted features.
[0138] All methods described herein may be performed in any suitable order unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples or exemplary language provided with respect to certain embodiments herein (e.g., "such as") is intended merely to make the disclosure more clear and does not pose limitations on the scope of the disclosure unless otherwise stated in the claims. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0139] Groupings of alternative elements or embodiments of the disclosure disclosed herein are not to be construed as limitations. Each group member may be referenced and claimed individually or in combination with other members of the group or other elements described herein. One or more members of a group may be included in, or deleted from, a group for reasons of convenience or patentability.
[0140] Although the present disclosure has been described in detail, it will be apparent that modifications, variations, and equivalents are possible without departing from the scope of the present disclosure as defined in the appended claims. Moreover, it should be understood that all examples in the present disclosure are provided as non-limiting examples. EXAMPLES
[0141] The following examples are included to demonstrate embodiments of the present disclosure. It will be appreciated by those skilled in the art that many changes can be made to the specific embodiments disclosed, without departing from the concept, spirit and scope of the present disclosure, while still achieving the same or similar results. More specifically, it will be apparent that certain agents that are both chemically and physiologically related may be substituted for the agents described herein, while still achieving the same or similar results. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the present disclosure, as defined by the appended claims.
[0142] Example 1 Synthesis of promoter for expressing gRNA The novel synthetic transcriptional regulatory elements are synthetic expression elements designed by algorithmic methods. The synthetic promoter elements of the present invention effect transcription of small nuclear RNA (snRNA) molecules, such as guide RNA (gRNA) molecules. Although the designed synthetic snRNA promoter elements do not have extensive homology to any known nucleic acid sequences occurring in nature, they affect transcription of operably linked DNA sequences similar to naturally occurring snRNA promoters. The full-length synthetic snRNA promoters of the present invention share little sequence identity with each other; ranging from about 38 percent identity to about 47 percent identity. Shortened variants of the synthetic snRNA promoters have also been produced. The truncated synthetic snRNA promoters also share little sequence identity with each other; ranging from about 41 to about 51 percent identity. The low percent identity between the synthetic snRNA promoters reduces the chance of recombination between promoters, making the synthetic snRNA promoters ideal for stacking multiple RNA expression cassettes; where each cassette contains a different synthetic snRNA promoter. As further described in the Examples below, both full-length and truncated synthetic snRNA promoters have demonstrated the ability to drive expression of gRNAs. Table 1 below shows the different synthetic snRNA promoters, their corresponding truncated variants (denoted with "_TR"), and the respective lengths of each synthetic snRNA promoter. [Table 1]
[0143] Example 2 Analysis of synthetic snRNA promoters in transfected maize leaf protoplasts. Maize leaf protoplasts are transfected with a plasmid construct containing an expression cassette for expressing the Cas12a endonuclease driven by a constitutive promoter; and a second expression cassette for expressing a gRNA driven by a synthetic snRNA promoter.
[0144] A plasmid construct was constructed using methods known in the art that contained two transgene cassettes, the first transgene cassette is used for expression of a nuclear-targeted Cas12a protein, which contains EXP, EXP-Zm.UbqM1:1:9 (SEQ ID NO:11), which is operably linked at 5' to a coding sequence, Cas12a_NLS (SEQ ID NO:12), which encodes a nuclear-targeted Cas12a_NLS protein (SEQ ID NO:19), which is operably linked at 5' to a 3'UTR, T-Os.LTP:2 (SEQ ID NO:13); the second transgene cassette contains a synthetic snRNA promoter selected from the group consisting of SEQ ID NOs:1-10, which is operably linked at 5' to a guide RNA, gRNA-Zm.Bmr3_2691 (SEQ ID NO:15), which contains a guide RNA spacer, NR-Zm.Bmr3_2691 (SEQ ID NO:14). The gRNA, gRNA-Zm.Bmr3_2691, is designed to cleave within the brown midrib 3 (Bmr3) genomic sequence (shown as SEQ ID NO: 18) by Cas12a endonuclease. The brown midrib mutation is one of the earliest reported in maize. Plants containing the brown midrib mutation exhibit reddish-brown pigmentation in the midrib of leaves from the time they have 4-6 leaves. These mutations are known to alter the lignin composition and digestibility of plants, thus making them prime candidates in corn silage breeding. The Bmr3 gene encodes the enzyme O-methyltransferase (COMT), which is involved in lignin biosynthesis (Vignols et al., 1995, The Plant Cell, Vol. 7, 407-416).
[0145] Maize leaf protoplasts are transfected using a PEG-based transfection method similar to that known in the art. To evaluate the effectiveness of each of the synthetic snRNA promoters, amplicon fragments are generated from the isolated genomic DNA from a population of transfected protoplast cells using primers that allow the amplification of DNA fragments containing the cleavage site region. The sequences of the amplicon fragments are aligned to identify any fragment sequences that contain mutations, such as DNA deletions, in the cleavage site region. The presence of such mutations indicates the ability of the synthetic snRNA promoter to drive the expression of gRNA.
[0146] Example 3 Analysis of two synthetic snRNA promoters in transfected maize leaf protoplasts. Maize leaf protoplasts are transfected with a plasmid construct containing an expression cassette for expressing the Cas12a endonuclease driven by a constitutive promoter; and two expression cassettes for expressing two different gRNAs, each driven by a synthetic snRNA promoter.
[0147] A plasmid construct was constructed using methods known in the art that contained three transgene cassettes, the first transgene cassette is used for expression of a nuclear-targeted Cas12a protein and contains EXP, EXP-Zm.UbqM1:1:9 (SEQ ID NO:11), which is operably linked 5' to a coding sequence, Cas12a_NLS (SEQ ID NO:12), which encodes a nuclear-targeted Cas12a_NLS protein (SEQ ID NO:19), which is operably linked 5' to a 3'UTR, T-Os.LTP:2 (SEQ ID NO:13); the second transgene cassette contains a synthetic snRNA promoter selected from the group consisting of SEQ ID NOs:1-10. the third transgene cassette comprises a second synthetic snRNA promoter (different from the synthetic snRNA promoter used in the second transgene cassette) selected from the group consisting of SEQ ID NOs: 1-10, which is operably linked at its 5' to a guide RNA, gRNA-Zm.Bmr3_2691 (SEQ ID NO: 15), which comprises a guide RNA spacer, NR-Zm.Bmr3_2691 (SEQ ID NO: 14); the third transgene cassette comprises a second synthetic snRNA promoter (different from the synthetic snRNA promoter used in the second transgene cassette) selected from the group consisting of SEQ ID NOs: 1-10, which is operably linked at its 5' to a guide RNA, gRNA-Zm.Bmr3_3170 (SEQ ID NO: 17), which comprises a guide RNA spacer, NR-Zm.Bmr3_3170 (SEQ ID NO: 16). The gRNA, gRNA-Zm.Bmr3_2691 and gRNA, gRNA-Zm.Bmr3_3170 are designed to cleave within the brown midrib 3 (Bmr3) genomic sequence (shown as SEQ ID NO: 18) by Cas12a endonuclease.
[0148] Maize leaf protoplasts are transfected using a PEG-based transfection method similar to that known in the art. To evaluate the effectiveness of each of the synthetic snRNA promoters in the construct stack, amplicon fragments are generated from isolated genomic DNA from a population of transfected protoplast cells using primers that allow amplification of a DNA fragment containing both cleavage site regions. Mutations or deletions of approximately 480 base pairs detected at each of the cleavage sites indicate the ability of each of the synthetic promoters to drive the expression of their respective gRNAs.
[0149] Example 4 Introducing a targeted double-strand break into the genome of a cell This example demonstrates the use of a synthetic snRNA promoter sequence, when presented with a Cas9 endonuclease, Cas12a endonuclease, or other CRISPR endonuclease, to drive gRNA expression to make a targeted double-stranded break in a cell's genome.
[0150] The synthetic snRNA promoters and truncated variant synthetic snRNA promoters of the invention, shown as SEQ ID NOs: 1-10, can be used to drive gRNA expression in plant cells. When presented to the nucleus of a cell along with Cas9 endonuclease, Cas12a endonuclease, or CRISPR endonuclease, DNA breaks occur at selected target regions that contain sequences complementary to the spacer region of the gRNA.
[0151] There are multiple means by which the necessary components can be introduced into plant cells. The gRNA can be expressed from a DNA fragment that includes a synthetic snRNA promoter or a truncated synthetic snRNA promoter operably linked at 5' to the nucleotide sequence encoding the gRNA, and a 3' poly-T stretch that terminates transcription. Alternatively, the sequence encoding the gRNA can be cloned into a plasmid construct. The plasmid construct can be a construct used to transfect plant-derived protoplasts, or the construct can be a binary plant transformation construct used to stably transform plant cells. The Cas9 endonuclease, Cas12a endonuclease, or other CRISPR endonuclease can be introduced into the plant cell as a protein or via a heterologous DNA that is used to express the Cas9 endonuclease, Cas12a endonuclease, or other CRISPR endonuclease. The Cas9 endonuclease, Cas12a endonuclease, or other CRISPR endonucleases contain at least one nuclear localization signal (NLS) so that endonuclease cleavage occurs more efficiently within the nucleus of a cell.
[0152] Plant cells can be transfected by particle bombardment. In this case, Cas9 endonuclease, Cas12a endonuclease, or other CRISPR endonuclease can be introduced as a protein or as a DNA fragment comprising a plant-expressible promoter operably linked 5' to an intron operably linked 5' to a coding sequence encoding Cas9 endonuclease, Cas12a endonuclease, or other CRISPR endonuclease, optionally including at least one NLS operably linked to the 5' to 3' UTR. DNA encoding gRNA can be introduced into cells with a heterologous DNA fragment comprising a synthetic snRNA promoter or a truncated synthetic snRNA promoter (SEQ ID NO: 1-10) operably linked 5' to a sequence encoding gRNA that also includes a 3' poly-T stretch to terminate transcription. Protoplast cells can also be transfected using the same reagents described above.
[0153] Protoplast cells can also be transfected using one or two plasmid constructs. One such method uses two constructs, as described in Example 2 above, where the first construct contains the transgene cassette for expressing gRNA, and the second construct contains the transgene cassette used for expressing Cas9 endonuclease, Cas12a endonuclease, or other CRISPR endonuclease. Alternatively, both the gRNA transgene cassette and the Cas9 endonuclease, Cas12, or other CRISPR endonuclease transgene cassette can be included in one construct used for transfection.
[0154] To stably transform plant cells, both the gRNA expression cassette and the Cas9 endonuclease, Cas12a endonuclease, or other CRISPR endonuclease expression cassette can be included in one binary plant transformation plasmid construct. Alternatively, two constructs can be used to simultaneously transform plant cells, the first construct containing the gRNA expression cassette; the second construct containing the Cas9 endonuclease, Cas12a endonuclease, or CRISPR endonuclease expression cassette.
[0155] To induce double-strand breaks in DNA without incorporating the transgene cassette into the plant genome, the gRNA and Cas9 endonuclease, Cas12a endonuclease, or other CRISPR endonuclease expression cassette can be excised as linear fragments from the construct(s) that comprised the cassette. The expression cassette and blunt-ended DNA fragments can be delivered to plant cells by particle bombardment. The bombarded cells are induced to form callus. The callus is then used to form whole plants.
[0156] The resulting breaks introduced into the genome of a cell can be used to introduce oligo DNA fragments; or to alter or disrupt sequences by error-prone non-homologous end joining.
[0157] Example 5 Genome modification by integration of blunt-ended double-stranded DNA fragments This example demonstrates how a synthetic snRNA promoter sequence, when presented with a CRISPR endonuclease, can be used to drive gRNA expression and integrate a blunt-ended double-stranded DNA fragment into a selected target site.
[0158] The complementary oligonucleotides are pre-annealed to form blunt-ended double-stranded DNA fragments. The DNA fragments and constructs containing gRNA and CRISPR endonuclease expression cassettes are co-transfected into plant protoplasts. The oligonucleotides can be designed to contain about three base pairs of microhomology regions or no microhomology regions to the corresponding 5' and 3' flanking sequences at the cleavage site in the genomic target site. The microhomology regions can promote the integration of the blunt-ended double-stranded DNA fragments through a mechanism of microhomology-driven non-homologous end joining at the genomic target site.
[0159] One or two constructs may be used to express the gRNA and CRISPR endonuclease. In one construct, both the gRNA and Cas9 expression cassettes are cloned into one plasmid construct. If two constructs are desired, the first construct will contain the gRNA expression cassette and the second construct will contain a cassette for expressing the CRISPR endonuclease. The gRNA expression cassette will contain one of the synthetic snRNA promoters or truncated variant synthetic snRNA promoters of the present invention shown as SEQ ID NOs: 1-10.
[0160] In the case of protoplast transfection, the construct(s) containing the gRNA expression cassette and the CRISPR endonuclease expression cassette are co-transfected with a blunt-ended double-stranded DNA fragment. Detection of the integration of the blunt-ended double-stranded DNA fragment can be performed by amplifying the region surrounding the target site integration and detecting the amplicon using high-resolution capillary electrophoresis and direct sequencing of the amplicon.
[0161] To integrate the blunt-ended DNA fragments into the selected target site and obtain a stably modified plant, the gRNA and CRISPR endonuclease expression cassettes can be excised as linear fragments from the construct(s) that constitute the cassette. The expression cassettes and blunt-ended DNA fragments can be delivered to plant cells by particle bombardment. The bombarded cells are induced to form callus. The callus is then used to form whole plants. The regenerated plants are then assayed using methods known in the art, such as amplification and sequencing, to identify plants that contain the DNA fragments in their genome.
[0162] Example 6 Targeting multiple unique genomic sites by gRNA multiplexing The main advantage of the CRISPR system compared to other genome engineering platforms is that multiple gRNAs directed to individual unique genomic target sites can be delivered as separate components to achieve targeting. Alternatively, multiple gRNAs directed to individual unique genomic target sites can be multiplexed in one expression construct to achieve targeting. An example of an application that may require multiple targeted endonucleolytic cleavage includes the removal of marker genes from transgenic events. The CRISPR system can be used to remove the selection marker from the transgenic insert, leaving the gene(s) of interest.
[0163] Another example of an application where such a CRISPR system may be useful is when multiple targeted endonucleolytic cuts are required, e.g., when the identification of the causative gene behind a quantitative trait is hindered by the absence of meiotic recombination in the QTL region that separates the gene candidates from each other. This can be circumvented by transforming with multiple CRISPR constructs that simultaneously target the genes of interest. These constructs knock out the gene candidates by frameshift mutations or remove the gene candidates by deletion. Such transformation results in random combinations of intact and mutated loci, allowing the identification of the causative gene.
[0164] The gRNA expression cassette comprises two or more synthetic snRNA promoters and / or truncated variant synthetic snRNA promoters of the invention operably linked 5' to a unique gRNA coding sequence designed to direct CRISPR endonuclease activity to a specific site within a plant cell genomic region, as shown as SEQ ID NOs: 1-10. It may be advantageous to use truncated variant synthetic snRNA promoters (SEQ ID NOs: 6-10) in that the smaller size of the truncated variant synthetic snRNA promoters allows for the construction of smaller constructs and reduces the probability of replication errors occurring in the bacterial host prior to transformation of the plant cell.
[0165] Binary plant transformation constructs are constructed similarly to those described above in Example 3, but contain multiple gRNA expression cassettes, each with its own synthetic snRNA promoter or truncated variant synthetic snRNA promoter operably linked 5' to its own gRNA coding sequence. Binary plant transformation constructs also contain expression cassettes used to express CRISPR endonucleases. Plant cells are transformed using Agrobacterium-mediated transformation methods. After transformation, the gRNAs direct the CRISPR endonuclease to genomic regions containing PAM sequences adjacent to sequences complementary to the spacer sequences of each gRNA, causing endonuclease cleavage within the genomic DNA in each target sequence. After cleavage, the genomic DNA between the target sites is excised, and the genomic DNA is repaired by non-homologous end joining. Excision of fragments of genomic DNA can be confirmed via various amplification or sequencing methods available in the art. Depending on the nature of the genomic region targeted for excision, changes in phenotype, metabolism, or other characteristics can be observed.
[0166] Example 7 Targeted integration by homologous recombination Genome modification by targeted integration of desired introduced DNA sequence occurs at the site of double-strand break in chromosome. Integration of DNA sequence is mediated by non-homologous end joining mechanism or homologous recombination utilizing the DNA repair mechanism of host cell. Double-strand break in cell genome can be achieved using CRISPR endonuclease and gRNA that directs CRISPR endonuclease to the target region of genomic DNA. An example of an application that may require homologous recombination is the integration of an expression cassette into plant cell genome within a specific region of the plant genome.
[0167] Integration of a DNA fragment using homologous recombination requires a homologous region (herein referred to as a "homology arm" (HA)) identical to the region where integration is preferred after cleavage by a CRISPR endonuclease. The homology arms are adjacent to the 5' and 3' ends of the DNA fragment. The left HA is designed based on the sequence adjacent to the 5' side of the double-stranded break site for targeted integration. The right HA is designed based on the sequence adjacent to the 3' side of the double-stranded break site for targeted integration. The homology arm can be from about 2 to about 1200 base pairs, although longer homology arms may work more efficiently. A desirable range of homology arm sizes may be 230 base pairs to 1,003 base pairs in length.
[0168] To transfect protoplasts, the construct(s) used for transfection are similar to those described above in Example 5. The gRNA expression cassette comprises one of the synthetic snRNA promoters or truncated variant synthetic snRNA promoters of the present invention, shown as SEQ ID NOs: 1-10. The construct(s) can be co-transfected with a DNA fragment containing homology arms. Alternatively, the expression cassette can be excised from the plasmid construct(s) and the linear expression cassette fragment can be co-transfected with a DNA fragment containing homology arms.
[0169] For stable integration of the DNA fragment containing the homology arms resulting in a stably transformed plant containing the DNA fragment, the expression cassette can be excised from the plasmid construct(s) and the linear expression cassette fragment can be co-transformed with the DNA fragment containing the homology arms by particle bombardment. Alternatively, an expression cassette containing a synthetic snRNA promoter or a truncated synthetic snRNA promoter can be co-transformed with the DNA fragment containing the homology arms by particle bombardment. The transformed tissue is induced to form whole plants and the plants are selected for the presence of the integrated DNA fragment and characterized for insertion into the target site using methods known in the art.
[0170] For Agrobacterium-mediated stable integration of DNA fragments containing homology arms resulting in stable transformed plants containing the DNA fragments, one binary transformation construct can be constructed using methods known in the art. This construct includes a right T-DNA border region; a left homology arm adjacent to, for example, a first transgene cassette (used to select transformed plant cells using either a herbicide or an antibiotic); a second transgene cassette (containing an expression cassette for expression of a gene of interest); a right homology arm; a third transgene cassette (containing a plant-expressible promoter operably linked at 5' to a coding sequence encoding a nuclear-targeting CRISPR endonuclease and operably linked at 5' to a 3'UTR); a fourth transgene cassette (containing a synthetic snRNA promoter or a truncated synthetic snRNA promoter of the invention, shown as SEQ ID NOs: 1-10, operably linked at 5' to a gRNA coding sequence containing a poly-T stretch at the 3' end to terminate transcription); and a right T-DNA border. It may also be preferable that the selectable marker cassette, the CRISPR endonuclease cassette, and the gRNA cassette are flanked by sites that allow for excision of the selectable marker, such as Lox sites that are cleaved by Cre recombinase, after transformants have been selected and characterized.
[0171] By using the two right T-DNA borders, the T-DNA forms double-stranded DNA as a result of the replication process by Agrobacterium. The selection and expression cassettes flanked by homology arms are integrated at the target site. Loss of chromosomal integration of either the full-length T-DNA or the partial T-DNA can be achieved in subsequent generations through breeding and segregation by selecting segregants that only have the selection and expression cassette in the target site. Removal of the selectable marker cassette can be achieved by crossing the plant containing the selection and expression cassette with a transformed plant expressing Cre-recombinase. The Cre-recombinase expression cassette can then also be selected through segregation in the next generation.
[0172] Example 8 P-GSP2262_TR can drive the expression of gRNA Maize plants were transformed with a plasmid construct containing an expression cassette for expression of Cas12a driven by a plant-expressible promoter and an expression cassette for expression of a gRNA driven by the synthetic snRNA promoter GSP2262_TR and assessed for editing within a specific region of the Bmr3 target sequence (SEQ ID NO: 18).
[0173] Maize plants were transformed using two different plasmid constructs, construct-1 and construct-2. Each construct contained an expression cassette for selecting transformed plant cells using glyphosate selection and an expression cassette for expression of Cas12a. Construct-1 also contained an expression cassette for expressing a gRNA, gRNA-Zm.Bmr3_90_3279 (SEQ ID NO:23), driven by the synthetic snRNA promoter GSP2262_TR (SEQ ID NO:6). The gRNA, gRNA-Zm.Bmr3_90_3279, contained two spacer sequences, NR-Zm.Bmr3_90 (SEQ ID NO:20) and NR-Zm.Bmr3_3279 (SEQ ID NO:22), that direct Cas12a to cleave within the Bmr3 target sequence (SEQ ID NO:18). Construct-2 also contained an expression cassette for expressing the gRNA, gRNA-Zm.Bmr3_227_3279 (SEQ ID NO:24), driven by the synthetic snRNA promoter GSP2262_TR (SEQ ID NO:6). The gRNA, gRNA-Zm.Bmr3_227_3279, contained two spacer sequences, NR-Zm.Bmr3_227 (SEQ ID NO:21) and NR-Zm.Bmr3_3279 (SEQ ID NO:22), that direct Cas12a to cleave within the Bmr3 target sequence (SEQ ID NO:18).
[0174] Maize plants were transformed with the above two plasmid constructs using Agrobacterium-mediated transformation method. The transformed cells were induced to form plants by methods known in the art. Leaf tissue samples were taken from the transformed R0 plants, and genomic DNA was extracted from each sample. The regions spanning the target sites were sequenced. The percentage of plants containing at least one edited allele was calculated for each cut site. The percentage of edited target sites is shown in Table 2 below. [Table 2]
[0175] As can be seen in Table 2 above, the synthetic snRNA promoter P-GSP2262_TR (SEQ ID NO: 6) was able to drive gRNA expression, as indicated by the percentage of editing sites specific for each gRNA.
[0176] Example 9 Assay of synthetic snRNA promoters in driving expression of gRNA targeting the Bmr3 genomic locus using transfected protoplasts Maize leaf protoplasts were transfected with the constructs and evaluated for efficacy in inducing editing within the Bmr3 target sequence (SEQ ID NO: 18), where the first construct contains an expression cassette for expression of Cas12a driven by a plant-expressible promoter and the second construct contains an expression cassette for expression of a gRNA designed to target the Bmr3 genomic locus driven by a synthetic snRNA promoter.
[0177] Maize leaf protoplasts were transfected with multiple constructs to assay the ability of synthetic snRNA promoters to drive expression of gRNAs, resulting in editing of specific sequences within the Bmr3 target site (SEQ ID NO: 18). Each protoplast preparation was transfected with four different constructs. The first construct was used to drive expression of Cas12a (Cas12a_NLS, SEQ ID NO: 12) in protoplast cells using a constitutive promoter. The second construct was used to drive expression of gRNAs targeting the Bmr3 locus, driven by a synthetic snRNA promoter selected from the group consisting of SEQ ID NOs: 1-10. The third and fourth constructs were used to drive expression of Renilla luciferase and Firefly luciferase genes, respectively, using constitutive promoters to assess the success of protoplast transfection.
[0178] The second construct used to drive expression of the gRNA, driven by a synthetic snRNA promoter selected from the group consisting of SEQ ID NOs: 1-10, contained one of three different gRNAs: (1) gRNA-Zm.Brm3_2691_2 (SEQ ID NO: 25) (containing the spacer NR-Zm.Brm3_2691 (SEQ ID NO: 14) and directing the Cas12a_NLS protein to cleave within the Bmr3 target sequence); (2) gRNA-Zm.Brm3_3170_2 (SEQ ID NO: 26) (containing the spacer NR-Zm.Brm3_3170 (SEQ ID NO: 16) and directing the Cas12a_NLS protein to cleave within the Bmr3 target sequence); and (3) gRNA-Zm.Brm3_2691_3170 (SEQ ID NO: 27) (directing the Cas12a_NLS to cleave within both positions of the Bmr3 target sequence). A total of 30 constructs were generated to provide all three gRNAs for each of the 10 synthetic snRNA promoters.
[0179] Maize leaf protoplasts were transfected with the above five constructs (1, 2, 3, 4, and 5) using a PEG-based transfection method similar to that known in the art. Genomic DNA was isolated from the protoplast cells after transfection and incubation. DNA sequencing was performed around the target region of the Bmr3 target site. Each transfection was repeated four times, and the average %InDel was calculated based on the four repeats. The percentage of InDel was calculated as follows: %InDel=100x[(In+Del) / (TotalRC)], where "In" is the number of insertion reads; "Del" is the number of deletion reads; and "TotalRC" is the number of all sequence reads from a particular sample, including wild-type and mutant reads. Since each guide RNA differs in terms of the efficiency of inducing double-strand breaks, the %InDel of each gRNA driven by the 10 synthetic snRNA promoters was normalized to 100% for the repeat from any of the 10 snRNA promoters with the highest %InDel. Table 3 shows the average %InDel and average normalized %InDel corresponding to two single-target gRNAs, gRNA-Zm.Brm3_2691_2 (SEQ ID NO: 25) and gRNA-Zm.Brm3_3170_2 (SEQ ID NO: 26). Table 4 shows the average %InDel and average normalized %InDel corresponding to two-target gRNA, gRNA-Zm.Brm3_2691_3170 (SEQ ID NO: 27). [Table 3] [Table 4]
[0180] As seen in Tables 3 and 4, each of the synthetic snRNA promoters was able to drive gRNA expression to direct Cas12a editing at the target site. In these experiments, the gRNA, gRNA-Zm.Brm3_2691_2, appeared to be less efficient than the other two gRNAs, especially with lower average %InDel for promoters GSP2262, GSP2272, and GSP2272_TR. However, these three synthetic snRNA promoters showed similar %InDel to the other synthetic snRNA promoters when driving gRNAs, gRNA-Zm.Brm3_3170_2 and gRNA-Zm.Brm3_2691_3170.
[0181] Example 10 Assay of synthetic snRNA promoters in driving expression of gRNAs targeting the Zm7 genomic locus using transfected protoplasts Maize leaf protoplasts were transfected with the constructs and evaluated for efficacy in inducing editing within the Zm7 target sequence (SEQ ID NO:28), where the first construct contains an expression cassette for expression of Cas12a driven by a plant-expressible promoter and the second construct contains an expression cassette for expression of a gRNA designed to target the Bmr3 genomic locus driven by a synthetic snRNA promoter.
[0182] Maize leaf protoplasts were transfected with multiple constructs to assay the ability of synthetic snRNA promoters to drive expression of gRNAs, resulting in editing of specific sequences within the Zm7 target site (SEQ ID NO: 28). Each protoplast preparation was transfected with four different constructs. The first construct was used to drive expression of Cas12a (Cas12a_NLS, SEQ ID NO: 12) in protoplast cells using a constitutive promoter. The second construct was used to drive expression of gRNAs targeting the Zm7 locus, driven by a synthetic snRNA promoter selected from the group consisting of SEQ ID NOs: 1-10. The third and fourth constructs were used to drive expression of Renilla luciferase and Firefly luciferase genes, respectively, using constitutive promoters to assess the success of protoplast transfection.
[0183] The second construct used to drive expression of the gRNA is driven by a synthetic snRNA promoter selected from the group consisting of SEQ ID NOs: 1-10 and contains one of three different gRNAs: (1) gRNA-Zm.7.1b (SEQ ID NO: 30) (contains the spacer NR-Zm.7.1b (SEQ ID NO: 29) and directs the Cas12a_NLS protein to cleave within the Zm7 target sequence); (2) gRNA-Zm.7.1c (SEQ ID NO: 32) (contains the spacer NR-Zm.7.1c (SEQ ID NO: 31) and directs the Cas12a_NLS protein to cleave within the Zm7 target sequence); and (3) gRNA-7.1c_7.1b (SEQ ID NO: 33) (directs the Cas12a_NLS to cleave within both positions of the Zm7 target sequence). A total of 30 constructs were made to provide all three gRNAs for each of the 10 synthetic snRNA promoters.
[0184] Maize leaf protoplasts were transfected with the above five constructs (first, second, third, fourth, and fifth) using a PEG-based transfection method similar to that known in the art. Genomic DNA was isolated from the protoplast cells after transfection and incubation. DNA sequencing was performed around the target region of the Bmr3 target site. Each transfection was repeated four times, and the average %InDel was calculated based on the four repeats. %InDel and %InDel normalization were calculated as described in Example 9 above.
[0185] Table 5 shows the average %InDel and average normalized %InDel corresponding to two single target gRNAs, gRNA-Zm.7.1b (SEQ ID NO: 30) and gRNA-Zm.7.1c (SEQ ID NO: 32). Table 6 shows the average %InDel and average normalized %InDel corresponding to a two target gRNA, gRNA-7.1c_7.1b (SEQ ID NO: 33). [Table 5] [Table 6]
[0186] As can be seen in Tables 3 and 4, each synthetic snRNA promoter was able to drive gRNA expression and direct Cas12a editing at the target site. Editing of the Zm.7.1b site was less efficient than that of the Zm.7.1c site. However, the synthetic snRNA promoter was able to drive gRNA expression and affect editing by Cas12a.
[0187] Example 11 Assay of synthetic snRNA promoters in driving expression of gRNA targeting the Bmr3 genomic locus in stably transformed maize plants Corn plants were transformed with a plasmid construct containing an expression cassette for expression of Cas12a driven by a plant-expressible promoter and an expression cassette for expression of a gRNA driven by a synthetic snRNA promoter shown as SEQ ID NOs:6-10, and assessed for editing within a specific region of the Bmr3 target sequence (SEQ ID NO:18).
[0188] Maize plants were transformed with five plasmid constructs containing three expression cassettes; a first expression cassette for selecting transformed plant cells using glyphosate selection, a second expression cassette for expressing Cas12a using a plant-expressible promoter, and a third transgene cassette for expression of a gRNA, gRNA-Zm.Brm3_2691_3170 (SEQ ID NO: 27) (driven by a synthetic snRNA promoter shown as SEQ ID NO: 6-10 that directs Cas12a to cleave within two regions of the Bmr3 target sequence (SEQ ID NO: 18)).
[0189] Maize plants were transformed with the above two plasmid constructs using Agrobacterium-mediated transformation method. The transformed cells were induced to form plants by methods known in the art. Leaf tissue samples were taken from the transformed R0 plants, and genomic DNA was extracted from each sample. One and two copy events were selected and sequencing of the region spanning the target site was performed. The average InDel percentage was calculated based on the number of insertions and deletions observed at each target site. Table 7 shows the average InDel percentage calculated for each of the two target sites in the Bmr3 target sequence. [Table 7]
[0190] As can be seen in Table 7 above, each of the synthetic snRNA promoters was able to drive gRNA expression to direct Cas12a editing at the target site of the Bmr3 target sequence.
[0191] Example 12 Assay of synthetic snRNA promoters in driving expression of gRNA targeting the Zm7 genomic locus in stably transformed maize plants
[0192] Corn plants were transformed with a plasmid construct containing an expression cassette for expression of Cas12a driven by a plant-expressible promoter and an expression cassette for expression of a gRNA driven by a synthetic snRNA promoter shown as SEQ ID NOs:6-10, and assessed for editing within a specific region of the Zm7 target sequence (SEQ ID NO:28).
[0193] Maize plants were transformed with five plasmid constructs containing three expression cassettes; a first expression cassette for selecting transformed plant cells using glyphosate selection, a second expression cassette for expressing Cas12a using a plant-expressible promoter, and a third transgene cassette for expression of a gRNA, gRNA-7.1c_7.1b (SEQ ID NO: 33) (driven by a synthetic snRNA promoter shown as SEQ ID NO: 6-10 that directs Cas12a to cleave within two regions of the Zm7 target sequence (SEQ ID NO: 28)).
[0194] Maize plants were transformed with the above two plasmid constructs using Agrobacterium-mediated transformation. The transformed cells were induced to form plants by methods known in the art. Leaf tissue samples were taken from the transformed R0 plants, and genomic DNA was extracted from each sample. One and two copy events were selected and sequencing of the region spanning the target site was performed. The average InDel percentage was calculated based on the number of insertions and deletions observed at each target site. Table 8 shows the average InDel percentage calculated for each of the two target sites in the Zm7 target sequence (SEQ ID NO: 28). [Table 8]
[0195] As can be seen in Table 8 above, each of the synthetic snRNA promoters was able to drive gRNA expression to direct Cas12a editing at the target site of the Zm7 target sequence. *******
[0196] Having illustrated and described the principles of the invention, it should be apparent to those skilled in the art that changes can be made in the arrangement and details of the invention without departing from such principles. The inventors claim all modifications that come within the spirit and scope of the claims. All publications and published patent documents cited in this specification are hereby incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.
Claims
1. A synthetic small nuclear RNA (snRNA) promoter comprising a DNA sequence selected from the group consisting of: a. a sequence having at least 85% sequence identity to SEQ ID NO: 1 or 6; b. a sequence comprising SEQ ID NO: 1 or 6; and c. A fragment of SEQ ID NO: 1 or 6.
2. The synthetic snRNA promoter of claim 1 , wherein the sequence has at least 90% sequence identity to the DNA sequence of SEQ ID NO: 1 or 6.
3. The synthetic snRNA promoter of claim 1 , wherein the sequence has at least 95% sequence identity to the DNA sequence of SEQ ID NO: 1 or 6.
4. The synthetic snRNA promoter of claim 1 , wherein the fragment comprises a gene regulatory activity.
5. 1. A recombinant DNA construct comprising a synthetic snRNA promoter operably linked to a DNA sequence encoding a guide RNA (gRNA), wherein the sequence of said synthetic snRNA promoter is SEQ ID NO: 1 or 6, or a fragment thereof, and wherein said fragment comprises a gene regulatory activity.
6. The recombinant DNA construct of claim 5 , further comprising a transcription termination sequence.
7. 6. The recombinant DNA construct of claim 5, further comprising a DNA sequence encoding a promoter operably linked to a type I CRISPR-associated protein, a type II CRISPR-associated protein, a type III CRISPR-associated protein, a type IV CRISPR-associated protein, a type V CRISPR-associated protein, or a type VI CRISPR-associated protein.
8. 8. The recombinant DNA construct of claim 7, wherein the CRISPR-associated protein is further operably linked to at least one nuclear localization sequence (NLS).
9. The CRISPR-associated protein is Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas12a, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, 8. The recombinant DNA construct of claim 7, selected from the group consisting of Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, CasX, CasY, and Mad7.
10. A recombinant DNA construct comprising a synthetic snRNA promoter operably linked to a sequence specifying a non-coding RNA, wherein the sequence of the synthetic snRNA promoter is SEQ ID NO: 1 or 6 or a fragment thereof, and wherein the fragment comprises a gene regulatory activity.
11. 11. The recombinant DNA construct of claim 10, wherein the non-coding RNA is selected from the group consisting of guide RNA (gRNA), single guide RNA (sgRNA), crRNA, pre-crRNA, tracrRNA, PEGRNA, microRNA (miRNA), miRNA precursor, small interfering RNA (siRNA), small RNA (22-26 nt in length) and precursors encoding same, heterochromatin siRNA (hc-siRNA), Piwi-interacting RNA (piRNA), hairpin double-stranded RNA (hairpin dsRNA), trans-acting siRNA (ta-siRNA), and naturally occurring antisense siRNA (nat-siRNA).
12. 1. A recombinant DNA construct comprising at least a first expression cassette comprising a synthetic snRNA promoter operably linked to a DNA sequence encoding a guide RNA (gRNA), wherein the sequence of said promoter comprises SEQ ID NO: 1 or 6 or a fragment thereof, said fragment comprising a gene regulatory activity.
13. 13. The recombinant DNA construct of claim 12, further comprising at least a second expression cassette, wherein the sequence encoding the first gRNA is different from the sequence encoding the second gRNA.
14. 14. The recombinant DNA construct of claim 13, wherein the synthetic snRNA promoter operably linked to the sequence encoding the first gRNA is different from the synthetic snRNA promoter operably linked to the sequence encoding the second gRNA.
15. 15. The construct of claim 14, comprising adjacent left and right homology arms (HA), each of about 200-1200 bp in length.
16. The construct of claim 15, wherein the length of the homology arms is about 230 to 1003 bp.
17. a. a first synthetic snRNA promoter selected from the group consisting of SEQ ID NO: 1 or 6 or a fragment thereof, wherein said fragment comprises a gene regulatory activity, said first synthetic RNA promoter being operably linked to a DNA sequence encoding a non-coding RNA; b. a second synthetic snRNA promoter selected from the group consisting of SEQ ID NOs: 1-10 or a fragment thereof, wherein the fragment comprises a gene regulatory activity, and the second synthetic snRNA promoter is operably linked to a DNA sequence encoding a non-coding RNA; A recombinant DNA construct comprising: wherein the first synthetic snRNA promoter and the second synthetic snRNA promoter are different.
18. 18. The recombinant DNA construct of claim 17, wherein the sequence encoding the first synthetic snRNA promoter and the sequence encoding the second synthetic snRNA promoter each comprise SEQ ID NO: 1 or 6, or a fragment thereof, wherein the fragment comprises gene regulatory activity.
19. 18. The recombinant DNA construct of claim 17, further comprising a sequence specifying one or more additional synthetic snRNA promoters operably linked to a DNA sequence encoding a non-coding RNA selected from the group consisting of SEQ ID NOs: 1-10 or a fragment thereof, wherein said fragment comprises a gene regulatory activity, and wherein each of the first synthetic snRNA promoter, the second synthetic snRNA promoter, and the one or more additional synthetic snRNA promoters are different.
20. 20. The recombinant DNA construct of claim 19, wherein the sequence specifying the one or more additional synthetic snRNA promoters is selected from the group consisting of SEQ ID NO: 1 or 6 or a fragment thereof, wherein the fragment contains a gene regulatory activity.
21. 20. The recombinant DNA construct of claim 19, wherein the recombinant DNA construct comprises 3, 4, 5, 6, 7, or 8 synthetic snRNA promoters.
22. 18. The recombinant DNA construct of claim 17, wherein the non-coding RNAs are gRNAs that target different selected target sites in a chromosome of a plant cell.
23. 20. The recombinant DNA construct of claim 17, further comprising a DNA sequence encoding a promoter operably linked to the DNA sequence encoding a CRISPR-associated protein.
24. The CRISPR-associated protein is Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas12a, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Csm7, Csm8, Csm9, Csm10, Csm12a, Csyl, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Csm7, Csm8, Csm9, Csm10, Csm12b, Csyl ...
24. The recombinant DNA construct of claim 23, selected from the group consisting of mr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, CasX, CasY, and Mad7.
25. A cell comprising a recombinant DNA construct selected from the group consisting of claim 5, claim 10, claim 12, and claim 17.
26. The cell of claim 25 , wherein the cell is a plant cell.
27. 27. The plant cell of claim 26, wherein the plant cell is a monocotyledonous plant cell.
28. 27. The plant cell of claim 26, wherein the plant cell is a dicotyledonous plant cell.
29. 27. The plant cell of claim 26, wherein the plant cell is selected from the group consisting of a corn plant cell, a soybean plant cell, a cotton plant cell, a peanut plant cell, a barley plant cell, an oat plant cell, a orchard grass plant cell, a rice plant cell, a sorghum plant cell, a sugarcane plant cell, a tall fescue plant cell, a turfgrass plant cell, a wheat plant cell, alfalfa plant cell, a canola plant cell, a cabbage plant cell, a mustard plant cell, a rutabaga plant cell, a turnip plant cell, a kale plant cell, a broccoli plant cell, a cauliflower plant cell, a pepper plant cell, a bean plant cell, a cowpea plant cell, a chickpea plant cell, a gourd plant cell, a lettuce plant cell, a cucumber plant cell, a melon plant cell, a carrot plant cell, a tomato plant cell, a radish plant cell, a potato plant cell, and an ornamental plant cell.