A novel CRISPR-CAS system for genome editing

Cas-alpha endonucleases address the limitations of existing genome editing technologies by offering targeted and efficient DNA cleavage in eukaryotes, enhancing specificity and reducing costs through optimized compositions.

JP7794634B2Active Publication Date: 2026-01-06PIONEER HI BREED INTERNATIONAL INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021533502
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-10
Filing Date
2019-12-13
Publication Date
2026-01-06
Estimated Expiration
2039-12-13

AI Technical Summary

Technical Problem

Existing genome editing technologies, such as ZFNs and TALENs, lack specificity and require redesign for each target site, making them costly and time-consuming, while CRISPR systems need improvement for efficient editing in eukaryotes like animals and plants.

Method used

Development of novel Cas endonucleases, termed Cas-alpha, guided by RNA and comprising specific domains for targeted DNA cleavage, optimized for eukaryotic cells, with compositions derived from various organisms and capable of generating double-strand breaks.

Benefits of technology

The Cas-alpha endonucleases provide targeted and efficient genome editing in eukaryotic cells, including plants and animals, with high specificity and reduced production costs by minimizing the need for redesign.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794634000024
    Figure 0007794634000024
  • Figure 0007794634000025
    Figure 0007794634000025
  • Figure 0007794634000026
    Figure 0007794634000026
Patent Text Reader

Abstract

Compositions and methods are provided for genomic modification of target sequences in the genome of a cell using novel Cas endonucleases. The methods and compositions provide an effective system for modifying or altering target sequences in the genome of a cell or organism using a guide polynucleotide / endonuclease system. Also provided are novel effector and endonuclease systems, such as guide polynucleotide / endonuclease systems comprising an endonuclease, and elements comprising such systems. Also provided are compositions and methods for guide polynucleotide / endonuclease systems comprising at least one endonuclease, optionally covalently or noncovalently linked to or assembled with at least one additional protein subunit, and compositions and methods for direct delivery of the endonuclease as a ribonucleotide protein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 62 / 779,989, filed December 14, 2018, U.S. Provisional Application No. 62 / 794,427, filed January 18, 2019, U.S. Provisional Application No. 62 / 819,409, filed March 15, 2019, U.S. Provisional Application No. 62 / 852,788, filed May 24, 2019, and U.S. Provisional Application No. 62 / 913,492, filed October 10, 2019, all of which are incorporated herein by reference in their entireties.

[0002] Reference to an electronically submitted sequence listing An official copy of the Sequence Listing is submitted electronically via EFS-Web as an ASCII formatted Sequence Listing having a size of 714,386 bytes and filed concurrently herewith with the present specification, with the filename RTS21920B_SequenceListing_ST25.txt, created on December 9, 2019. The Sequence Listing contained in this ASCII format document is a part of the present specification and is incorporated herein by reference in its entirety.

[0003] The present disclosure relates to the field of molecular biology, and in particular to compositions of a novel RNA-guided Cas endonuclease system, as well as compositions and methods for editing or modifying the genome of a cell. [Background technology]

[0004] Recombinant DNA technology has made it possible to insert DNA sequences into targeted genomic locations and / or modify specific endogenous chromosomal sequences. Site-specific integration techniques using site-specific recombination systems, as well as other recombinant technologies, have been used to generate targeted insertions of genes of interest in various organisms. Genome editing technologies such as designer zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), or homing meganucleases can be used to generate targeted genome perturbations, but these systems tend to use designer nucleases that have low specificity and need to be redesigned for each target site, making their production expensive and time-consuming.

[0005] A newer technique using archaeal or bacterial adaptive immune systems, termed CRISPR (clustered regularly interspaced short palindromic repeats), has been identified that contains various domains of effector proteins encompassing diverse activities (DNA recognition, binding, and optional cleavage).

[0006] Despite the identification and characterization of some of these systems, there remains a need to identify novel effectors and systems for editing endogenous and previously introduced heterologous polynucleotides and to demonstrate activity in eukaryotes, particularly animals and plants.

[0007] Described herein are novel Cas endonucleases, "Cas-alpha," exemplary proteins, and methods and compositions for their use. Summary of the Invention [Means for solving the problem]

[0008] Disclosed herein are novel Cas endonuclease compositions and methods of their use. These endonucleases of the novel Cas-alpha class can be guided by a guide polynucleotide to target and cleave double-stranded DNA in a PAM-dependent manner, as demonstrated in prokaryotes (E. coli) and three different eukaryotic kingdoms: plantae, animalae, and fungi.

[0009] In one aspect, a synthetic composition is provided comprising a CRISPR-Cas endonuclease comprising at least one zinc finger-like domain, at least one bridge helix-like domain, three split RuvC domains (comprising non-contiguous RuvC-I, RuvC-II, and RuvC-III domains), and optionally comprising a heterologous polynucleotide.

[0010] In any aspect, any of the compositions and methods provide at least one component that is optimized for expression in a eukaryotic cell, particularly a plant cell, a fungal cell, or an animal cell.

[0011] In one aspect, a synthetic composition is provided comprising a polynucleotide encoding a CRISPR-Cas effector protein and a heterologous polynucleotide derived from an organism selected from the group consisting of Acidibacillus sulfuroxidans, Alicyclobacillus acidoterrestris, Aneurinibacillus danicus, Archaea, Bacillus, Bacillus cereus, Bacillus megaterium, Bacillus pseudomycoides, Bacillus sp.), Bacillus thuringiensis, Bacillus toyonensis, Bacillus wiedmannii, Bacteroides plebeius, Bos taurus, Brevibacillus centrosporus, Candidatus Aureabacteria bacterium, Candidatus Levybacteria bacterium, Candidatus Micrarchaeota archaeon, Cellulosilyticum ruminicola, Clostridioides difficile, Clostridium botulinum, Clostridium fallax, Clostridium hiranonis, Clostridium humii, Clostridium novyi, Clostridium paraputrificum, Clostridium pasteurianum, Clostridium perfringens, Clostridium sp., Clostridium tetani, Clostridium ventriculi, Desulfovibrio fructosivorans, Dorea longicatena, Eubacterium siraeum, Flavobacterium thermophilum, Red jungle fowl (Gallus gallus), Hepatitis delta virus, Homo sapiens, Human betaherpesvirus 5, Hydrogenivirga sp., Mus musculus musculus, Parageobacillus thermoglucosidasius, Peptoclostridium sp., Phascolarctobacterium sp., Prevotella copri, Ruminiclostridium hungatei, Ruminococcus albus, Ruminococcus sp., Saccharomyces cerevisiae, Simian virus 40, Solanum tuberosum, Sulfurihydrogenibium azollense azorense, Syntrophomonas palmitatica, Tobacco etch virus, and Zea mays.

[0012] In one aspect, there is provided a synthetic composition comprising a eukaryotic cell and a heterologous CRISPR-Cas effector, wherein the heterologous CRISPR-Cas effector protein is less than 800, between 790 and 800, less than 790, between 780 and 790, less than 780, between 770 and 780, less than 770, between 760 and 770, less than 760, between 750 and 760, less than 750, between 740 and 750, less than 740, or less than 730. Synthetic compositions are provided that include up to 740, less than 730, 720-730, less than 720, 710-720, less than 710, 700-710, or less than 700 amino acids, e.g., less than 700, less than 790, less than 780, less than 750, less than 700, less than 650, less than 600, less than 550, less than 500, less than 450, less than 400, less than 350, or less than 350 amino acids.

[0013] In one aspect, there is provided a synthetic composition comprising a CRISPR-Cas endonuclease, wherein the CRISPR-Cas endonuclease, when aligned to SEQ ID NO: 17, comprises at least one, at least two, at least three, at least four, at least five, at least six, or seven of the following amino acid positions in SEQ ID NO: 17: glycine (G) at position 337, glycine (G) at position 341, glutamic acid (E) at position 430, leucine (L) at position 432, cysteine ​​(C) at position 487, cysteine ​​(C) at position 490, cysteine ​​(C) at position 507, and / or cysteine ​​(C) or histidine (H) at position 512.

[0014] In one aspect, a synthetic composition is provided comprising a CRISPR-Cas endonuclease, wherein the CRISPR-Cas endonuclease comprises one, two, or three of the following motifs: GxxxG, ExL, and / or one or more Cx n (C,H) where n = one or more amino acids.

[0015] In one aspect, a synthetic composition is provided comprising a CRISPR-Cas endonuclease, wherein the CRISPR-Cas endonuclease comprises one or more zinc finger motifs.

[0016] In one embodiment, the sequences of SEQ ID NOs: 17, 18, 19, 20, 32, 33, 34, 35, 36, 37, 38, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356 , 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, and 371, and at least 250, 250 to 300, at least 300, 300 to 350, at least 350, 350 to 400, at least 400, or more than 400 consecutive amino acids of a sequence selected from the group consisting of at least 50%, 50% to 55%, at least 55%, 55% to 60%, at least 60%, 60% to 65%, at least 65%, 65% to 70%, Synthetic compositions are provided that include CRISPR-Cas effector proteins that share 0%, at least 70%, 70%-75%, at least 75%, 75%-80%, at least 80%, 80%-85%, at least 85%, 85%-90%, at least 90%, 90%-95%, at least 95%, 95%-96%, at least 96%, 96%-97%, at least 97%, 97%-98%, at least 98%, 98%-99%, at least 99%, 99%-100%, or 100% sequence identity.

[0017] In one embodiment, the sequences of SEQ ID NOs: 17, 18, 19, 20, 32, 33, 34, 35, 36, 37, 38, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303 , 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, ​​383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405 At least 250, 250 to 500, at least 500, 500 to 600, at least 600, 600 to 700, at least 700, 700 to 750, at least 750, 750 to 800, at least 800, 800 to 850, at least 850, 850 to 900, at least 900, 900 to 950, at least 950, 950 to 1000, at least 1000, or more than 1000 amino acids of a sequence selected from the group consisting of: Acid and at least 50%, 50%-55%, at least 55%, 55%-60%, at least 60%, 60%-65%, at least 65%, 65%-70%, at least 70%, 70%-75%, at least 75%, 75%-80%, at least 80%, 80%-85%, at least 85%, 85%-90%, at least 90%, 90%-95%, at least 95%, 95%-96%, at least 96%, 96%-97%, at least 97%, 97%-98%, at least 98%, 98%-99%, at least 99%, 99%-100%,Alternatively, synthetic compositions comprising polynucleotides encoding CRISPR-Cas effector proteins that share 100% sequence identity are provided.

[0018] In one embodiment, the sequences of SEQ ID NOs: 57, 58, 59, 64, 65, 66, 67, 68, 73, 74, 75, 76, 77, 102, 103, 104, 105, 177, 178, 179, 180, 181, 182, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 8, 219, 220, 221, 222, 223, 224, 230, 231, 232, 233, 234, 238, 240, 241, 245, 246, 247, 248, 252, and 253. Provided are synthetic compositions comprising a polynucleotide encoding a CRISPR-Cas effector protein capable of hybridizing to a polynucleotide sharing 28, 29, 30, or more than 30 contiguous nucleotides and at least 50%, 50% to 55%, at least 55%, 55% to 60%, at least 60%, 60% to 65%, at least 65%, 65% to 70%, at least 70%, 70% to 75%, at least 75%, 75% to 80%, at least 80%, 80% to 85%, at least 85%, 85% to 90%, at least 90%, 90% to 95%, at least 95%, 95% to 96%, at least 96%, 96% to 97%, at least 97%, 97% to 98%, at least 98%, 98% to 99%, at least 99%, 99% to 100%, or 100% sequence identity.

[0019] Any of the methods or compositions herein may further comprise a heterologous polynucleotide. The heterologous polynucleotide may be selected from the group consisting of: a non-coding expression regulatory element such as a promoter, intron, enhancer, or terminator; a donor polynucleotide; a polynucleotide-modified template optionally containing at least one modification compared to a polynucleotide sequence in a cell; a transgene; a guide RNA; a guide DNA; a guide RNA-DNA hybrid; an endonuclease; a nuclear localization signal; and a cellular transit peptide.

[0020] In one aspect, methods for using any of the compositions disclosed herein are provided. In some embodiments, methods are provided for binding a Cas-alpha endonuclease to a target sequence of a polynucleotide, e.g., in the genome of a cell or in vitro. In some embodiments, the Cas-alpha endonuclease forms a complex with a guide polynucleotide, e.g., a guide RNA. In some embodiments, the complex recognizes, binds to, and optionally creates a nick (single-stranded) or break (double-stranded) in the polynucleotide at or near the target sequence. In some embodiments, the nick or break is repaired by non-homologous end joining (NHEJ). In some embodiments, the nick or break is repaired by homology-directed repair (HDR) or homologous recombination (HR) using a polynucleotide-modified template or donor DNA molecule.

[0021] The novel Cas endonucleases described herein can generate double-stranded breaks in or adjacent to a target polynucleotide that contains an appropriate PAM and is directed by a guide polynucleotide in any prokaryotic or eukaryotic cell. In some examples, the cell is a plant cell, an animal cell, or a fungal cell. In some examples, the plant cell is selected from the group consisting of corn, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oat, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, tobacco, Arabidopsis, safflower, and tomato.

[0022] BRIEF DESCRIPTION OF THE FIGURES AND SEQUENCE LISTING The present disclosure can be more fully understood from the following detailed description and the figures and sequence listing accompanying this specification, which form a part of this application. [Brief explanation of the drawings]

[0023] [Figure 1-1] Figure 1: Figures 1A-1D show complete CRISPR-Cas systems containing all components necessary for acquisition and integration. These include genes encoding all proteins required for spacer acquisition and integration (Cas1 and Cas2) and a novel protein containing a DNA cleavage domain, Cas-alpha (α), in an operon-like structure adjacent to the CRISPR array. Additionally, a gene encoding a protein with homology to Cas4 was also encoded at this locus. Figure 1A shows the locus structure of the Cas-alpha1, Cas-alpha3, and Cas-alpha4 systems. Figure 1B shows the locus structure of the Cas-alpha2 system. Figure 1C shows the locus structure of the Cas-alpha6 system. Figure 1D shows the locus structure of the Cas-alpha5, 7, 8, 9, 10, and 11 systems. [Figure 1-2] (As mentioned above.) [Figure 1-3] (As mentioned above.) [Figure 1-4] (As mentioned above.) [Figure 2] Figure 2 shows a detailed structural examination of the Cas-alpha protein, demonstrating clear differences from previously described class 2 endonucleases. Conserved residues are indicated. Key residues involved in DNA cleavage are marked with an asterisk. Numbers correspond to the Cas-alpha1 protein. [Figure 3] FIG. 3 outlines a method for detecting double-stranded DNA target recognition and cleavage using cell lysates expressing Cas-alpha endonucleases. [Figure 4-1] Figure 4: Figures 4A-4E show cleavage of a target polynucleotide by the Cas-alpha1 endonuclease at nucleotide position 21. Figure 4A shows data for the Cas-alpha1 negative control, Figure 4B shows data for Cas-alpha1 using the entire (complete) CRISPR locus with the CRISPR array engineered to cleave at the target polynucleotide, Figure 4C shows data for the Cas-alpha1 complete locus plus when expression is enhanced with the T7 promoter, Figure 4D shows data for the Cas-alpha1 minimal locus when expression is enhanced with the T7 promoter, and Figure 4E shows data for the reaction without Cas-alpha1 but with the remainder of the CRISPR locus when expression is enhanced with the T7 promoter. [Figure 4-2] (As mentioned above.) [Figure 4-3] (As mentioned above.) [Figure 4-4] (As mentioned above.) [Figure 4-5] (As mentioned above.) [Figure 5-1]Figure 5: Figures 5A-5B show a schematic diagram for determining the orientation of PAM recognition relative to spacer recognition. The guide RNA was designed to base-pair with either the sense or antisense strand of the T2 target. If the guide RNA designed to base-pair with the sense strand results in restoration of PAM preference and a cleavage signal, the protospacer is on the antisense strand, and PAM recognition occurs 3' to it (Figure 5A). Conversely, if the guide RNA designed to base-pair with the antisense strand results in PAM preference and a cleavage signal, the protospacer is on the sense strand, and PAM recognition occurs 5' to it (Figure 5B). [Figure 5-2] (As mentioned above.) [Figure 6-1] Figure 6: Figures 6A-6E show cleavage of a target polynucleotide by the Cas-alpha4 endonuclease at nucleotide position 24. Figure 6A shows data for the Cas-alpha4 negative control. Figure 6B shows data for Cas-alpha4 plus T2-1 sgRNA. Figure 6C shows data for Cas-alpha4 plus T2-2 sgRNA. Figure 6D shows data for Cas-alpha4 plus T2-1 crRNA / tracrRNA. Figure 6E shows data for Cas-alpha4 plus T2-2 crRNA / tracrRNA. [Figure 6-2] (As mentioned above.) [Figure 6-3] (As mentioned above.) [Figure 6-4] (As mentioned above.) [Figure 6-5] (As mentioned above.) [Figure 7-1]Figure 7: Figures 7A-7K show representative Cas-alpha loci, endonucleases, proteins, guide RNA components, and other sequences that have been identified from various bacterial and archaeal organisms, including: Candidatus Micrarchaeota archaeon (Figures 7A, 7B, 7E), Candidatus Aureabacteria bacterium (Figure 7C), various uncultured bacteria (Figures 7D, 7F), Parageobacillus thermoglucosidasius (Figure 7G), Acidibacillus sulfuroxidans (Figure 7H), Ruminococcus sp. (Figure 7I), and Syntrophomonas palmitatica. palmitatica (Figure 7J), and Clostridium novyi (Figure 7K). [Figure 7-2] (As mentioned above.) [Figure 7-3] (As mentioned above.) [Figure 7-4] (As mentioned above.) [Figure 7-5] (As mentioned above.) [Figure 7-6] (As mentioned above.) [Figure 7-7] (As mentioned above.) [Figure 7-8] (As mentioned above.) [Figure 7-9] (As mentioned above.) [Figure 7-10] (As mentioned above.) [Figure 7-11] (As mentioned above.) [Figure 8-1]Figure 8: Figures 8A-8K show distinct structural features in representative Cas-alpha proteins. The protein sequence is shown in bold. Non-bold letters below each amino acid residue indicate possible secondary structure features, where C represents an unstructured element or coil, E represents a beta strand, and H represents an alpha helix. Zinc finger domains are indicated by dashed boxes, and asterisks indicate key amino acid residues involved in zinc ion binding. The RuvC subdomains of split RuvC domains are indicated by solid boxes. The bridge helix is ​​indicated by a dash-dotted box. The coiled coil is depicted as a solid cylinder. Solid plus signs indicate key catalytic residues characteristic of the RuvC domain motif.FIG. 8A shows Cas-alpha 1 (SEQ ID NO: 17) from Candidatus Micrarchaeota archaeon, FIG. 8B shows Cas-alpha 2 from Candidatus Micrarchaeota archaeon (SEQ ID NO: 18), FIG. 8C shows Cas-alpha 3 (SEQ ID NO: 19) from Candidatus Aureabacteria bacterium, FIG. 8D shows Cas-alpha 4 (SEQ ID NO: 20) from an uncultured bacterium, and FIG. 8E shows Cas-alpha 5 from Candidatus Micrarchaeota archaeon. Figure 8F shows Cas-alpha 6 (SEQ ID NO: 33) from an uncultured bacterium, Figure 8G shows Cas-alpha 7 (SEQ ID NO: 34) from Parageobacillus thermoglucosidasius, Figure 8H shows Cas-alpha 8 (SEQ ID NO: 35) from Acidibacillus sulfuroxidans, Figure 8I shows Cas-alpha 9 (SEQ ID NO: 36) from Ruminococcus sp., Figure 8J shows Cas-alpha 10 (SEQ ID NO: 37) from Syntrophomonas palmitatica, which features a unique motif of three zinc finger domains, and Figure 8K shows Cas-alpha 10 (SEQ ID NO: 38) from Clostridium novyi.

[0039] Figure 3 shows Cas-alpha11 (SEQ ID NO: 38) from Cas-alpha11. Whole genome sequencing of the organism containing Cas-alpha11 showed that the Cas-alpha locus is the only CRISPR system in that organism. [Figure 8-2] (As mentioned above.) [Figure 8-3] (As mentioned above.) [Figure 8-4] (As mentioned above.) [Figure 8-5] (As mentioned above.) [Figure 8-6] (As mentioned above.) [Figure 8-7] (As mentioned above.) [Figure 8-8] (As mentioned above.) [Figure 8-9] (As mentioned above.) [Figure 8-10] (As mentioned above.) [Figure 8-11] (As mentioned above.) [Figure 9-1] Figure 9: Figure 9A shows how the Cas-alpha protein subunit interacts with the hybrid duplex of target DNA and guide RNA. Figure 9B is a three-dimensional model of the C-terminal half of Cas-alpha4, showing the helical hairpin / bridge helix region common to Cas proteins, the RuvC domain, and regions identified as zinc finger motifs. [Figure 9-2] (As mentioned above.) [Figure 10-1] Figure 10: Figures 10A-10D show examples of expression constructs for using Cas-alpha endonucleases in eukaryotic cells. Figure 10A is an example of a human cell Cas-alpha DNA expression construct. Figure 10B is an example of a plant cell Cas-alpha DNA expression construct. Figure 10C is an example of a yeast (Saccharomyces cerevisiae) Cas-alpha DNA expression construct. Figure 10D is an example of a yeast (Saccharomyces cerevisiae) Cas-alpha DNA expression construct. [Figure 10-2] (As mentioned above.) [Figure 10-3] (As mentioned above.) [Figure 10-4] (As mentioned above.) [Figure 11-1]Figure 11: Figures 11A-11D show examples of eukaryotic-optimized Cas-alpha guide RNA expression constructs. Figure 11A is an example of a human cell single guide RNA (sgRNA) DNA expression construct. Figure 11B is an example of a plant cell single guide RNA (sgRNA) DNA expression construct. Figure 11C is an example of a yeast (Saccharomyces cerevisiae) single guide RNA (sgRNA) DNA expression construct. Figure 11D is another example of a plant cell single guide RNA (sgRNA) DNA expression construct. [Figure 11-2] (As mentioned above.) [Figure 11-3] (As mentioned above.) [Figure 11-4] (As mentioned above.) [Figure 12] FIG. 12 shows an example of a recombinant gene for the recombinant expression and purification of Cas-alpha endonuclease in E. coli. [Figure 13] Figure 13 shows double-strand break repair mutations in plant cells resulting from Cas-alpha endonuclease activity. Mutations caused by Cas-alpha 4 in Zea mays are shown. WT reference is SEQ ID NO: 120, Mutation 1 is SEQ ID NO: 121, Mutation 2 is SEQ ID NO: 122, Mutation 3 is SEQ ID NO: 123, and Mutation 4 is SEQ ID NO: 124. [Figure 14-1] Figure 14: Figures 14A-14B show double-strand break repair mutations in animal cells derived from Cas-alpha endonuclease activity. Figure 14A shows indel mutations resulting from Cas-alpha 4 RNP electroporation (VEGF A target 2 mutations 1-5 given as SEQ ID NOs: 127-131, compared to SEQ ID NO: 126 of the WT reference; VEGFA target 3 mutation given as SEQ ID NO: 133, compared to SEQ ID NO: 132 of the WT reference). Figure 14B shows indel mutations resulting from Cas-alpha 4 and sgRNA DNA expression cassette lipofection, VEGFA target 3 (mutations 1 and 2 given as SEQ ID NOs: 134-135, compared to SEQ ID NO: 132 of the WT reference). [Figure 14-2](As mentioned above.) [Figure 15-1] Figure 15: Figures 15A-15D demonstrate Cas-alpha4 double-stranded DNA target cleavage. Figure 15A shows that supercoiled (SC) plasmid DNA containing a guide RNA target (approximately 20 bp) adjacent to the 3' end of a PAM (5'-TTTR-3', where R indicates A or Gbp) was converted to a fully linear form (FLL), thus indicating the formation of a dsDNA break. Furthermore, cleavage of the linear DNA generated DNA fragments of the expected size, further confirming Cas-alpha4-mediated dsDNA break formation. Figure 15B shows that Cas-alpha4 requires both a PAM and a guide RNA to cleave the dsDNA target. Figure 15C shows that Cas-alpha4 generates a 5' overhanging DNA cleavage site, with cleavage occurring primarily at positions 20-24 bp from the PAM sequence. Figure 15D demonstrates the trans-acting ssDNase activity of Cas-alpha4 activated by dsDNA only in the presence of a guide RNA. [Figure 15-2] (As mentioned above.) [Figure 15-3] (As mentioned above.) [Figure 15-4] (As mentioned above.) [Figure 16-1]Figure 16: Figures 16A-16T show the double-stranded DNA target cleavage activity of all Cas-alpha endonucleases except Cas-alpha5. Figure 16A is a negative control (-IPTG). Figure 16B is a negative control (+IPTG). Figure 16C shows cleavage of a double-stranded DNA target by Cas-alpha2 (-IPTG) at protospacer position 21. Figure 16D shows cleavage of a double-stranded DNA target by Cas-alpha2 (+IPTG) at protospacer position 21. Figure 16E shows no cleavage of a double-stranded DNA target by Cas-alpha3 (-IPTG). Figure 16F shows cleavage of a double-stranded DNA target by Cas-alpha3 (+IPTG) at protospacer position 21. Figure 16G shows no cleavage of a double-stranded DNA target by Cas-alpha5 (-IPTG). Figure 16H shows no cleavage of double-stranded DNA targets by Cas-alpha5 (-IPTG). Figure 16I shows cleavage of double-stranded DNA targets by Cas-alpha6 (-IPTG). Figure 16J shows no cleavage of double-stranded DNA targets by Cas-alpha6 (+IPTG) at protospacer position 24. Figure 16K shows cleavage of double-stranded DNA targets by Cas-alpha7 (-IPTG) at protospacer position 24. Figure 16L shows cleavage of double-stranded DNA targets by Cas-alpha7 (+IPTG) at protospacer position 24. Figure 16M shows no cleavage of double-stranded DNA targets by Cas-alpha8 (-IPTG). Figure 16N shows cleavage of double-stranded DNA targets by Cas-alpha8 (+IPTG) at protospacer position 24. Figure 16O shows cleavage of double-stranded DNA targets by Cas-alpha9 (-IPTG) at protospacer position 24. Figure 16P shows cleavage of a double-stranded DNA target by Cas-alpha9 (+IPTG) at protospacer position 24. Figure 16Q shows cleavage of a double-stranded DNA target by Cas-alpha10 (-IPTG) at protospacer position 24. Figure 16R shows cleavage of a double-stranded DNA target by Cas-alpha10 (+IPTG) at protospacer position 24. Figure 16S shows cleavage of a double-stranded DNA target by Cas-alpha11 (-IPTG) at protospacer position 24.Figure 16T shows cleavage of a double-stranded DNA target by Cas-alpha11 (+IPTG) at protospacer position 24. [Figure 16-2] (As mentioned above.) [Figure 16-3] (As mentioned above.) [Figure 16-4] (As mentioned above.) [Figure 16-5] (As mentioned above.) [Figure 16-6] (As mentioned above.) [Figure 16-7] (As mentioned above.) [Figure 16-8] (As mentioned above.) [Figure 16-9] (As mentioned above.) [Figure 16-10] (As mentioned above.) [Figure 16-11] (As mentioned above.) [Figure 16-12] (As mentioned above.) [Figure 16-13] (As mentioned above.) [Figure 16-14] (As mentioned above.) [Figure 16-15] (As mentioned above.) [Figure 16-16] (As mentioned above.) [Figure 16-17] (As mentioned above.) [Figure 16-18] (As mentioned above.) [Figure 16-19] (As mentioned above.) [Figure 16-20] (As mentioned above.) [Figure 17-1]Figure 17: Figure 17A shows one method for assessing Cas-alpha double-stranded DNA target cleavage in E. coli cells. Figures 17B-17E show double-stranded DNA target cleavage in E. coli. The "no target" experiment provides a baseline of transformation efficiency in the absence of double-stranded DNA target cleavage. The "target" experiment, PAM+T2, was performed both with and without IPTG (0.5 mM) to test target cleavage under different Cas-alpha endonuclease and guide RNA expression conditions. Figure 17B shows results for Cas-alpha2 and Cas-alpha3. Figure 17C shows results for Cas-alpha6 and Cas-alpha7. Figure 17D shows results for Cas-alpha8 and Cas-alpha9. Figure 17E shows results for Cas-alpha10 and Cas-alpha11. [Figure 17-2] (As mentioned above.) [Figure 17-3] (As mentioned above.) [Figure 17-4] (As mentioned above.) [Figure 17-5] (As mentioned above.) [Figure 18-1] Figure 18: Figures 18A-18B show double-strand break repair mutations in plant cells resulting from Cas-alpha endonuclease activity for particle gun experiments delivering Cas-alpha 10 DNA expression constructs into Zea mays immature embryos. Figure 18A shows restoration of a targeted deletion generated at or near the nuclease cleavage site of the nptII target site. Figure 18B shows restoration of a targeted deletion generated at or near the nuclease cleavage site of the ms26 target site. [Figure 18-2] (As mentioned above.) [Figure 19-1]Figure 19: Figure 19A shows the experimental design for homologous recombination repair in the eukaryotic cell, Saccharomyces cerevisiae. An exogenously supplied DNA repair template (double-stranded) with homology flanking the Cas-alpha10 target site was used to introduce one or two premature stop codons (depending on the DNA repair outcome) into the ade2 gene following a Cas-alpha10-induced double-strand break (DSB). To avoid targeting of the repair template, it also contained a T-to-A change in the PAM region of Cas-alpha10. Figure 19B shows that when both the repair template and the Cas-alpha10 and sgRNA expression constructs were transformed and a double-strand break was created by Cas-alpha endonuclease and repaired (HDR) with the template, a red cell phenotype indicative of ade2 gene disruption was recovered. Figure 19C shows the sequencing results of the Cas-alpha10 ade2 gene target site, confirming the introduction of at least one stop codon in three independent red colonies (labeled "1," "2," and "3"). The stop codon was introduced in the antisense frame. SEQ ID NO: 170. The reference DNA sequence from Saccharomyces cerevisiae is shown as SEQ ID NO: 170, the repair template DNA is SEQ ID NO: 171, red colony 1 repair result 1 is SEQ ID NO: 172, red colony 1 repair result 2 is SEQ ID NO: 173, red colony 2 repair result 1 is SEQ ID NO: 174, red colony 3 repair result 1 is SEQ ID NO: 175, and red colony 3 repair result 2 is SEQ ID NO: 176. [Figure 19-2] (As mentioned above.) [Figure 19-3] (As mentioned above.) [Figure 20]Figure 20 shows the phylogenetic relationships among several Cas-alpha orthologs. Three supergroups were identified (I, II, and III). Group I included Clade 1 (Candidate Archaea and Aureabacteria, whose loci typically encode Cas1, Cas2, and Cas4). Group II consists of clade 2 (Aquificae (genera Sulfurihydrogenibium and Hydrogenivirga) and Deltaproteobacteria (genus Desulfovibrio)), clade 3 (candidate archaea (genera usually encoding Cas1, Cas2, and Cas4)), clade 4 (Bacteroidetes (genera Prevotella and Bacteroides)), clade 5 (candidate Revibacterium)), clade 6 (Candidate Rhizobacterium)), clade 7 (Candidate Rhizobacterium)), clade 8 (Candidate Rhizobacterium)), clade 9 (Candidate Rhizobacterium)), clade 10 (Candidate Rhizobacterium)), clade 11 (Candidate Rhizobacterium)), clade 12 (Candidate Rhizobacterium)), clade 13 (Candidate Rhizobacterium)), clade 14 (Candidate Rhizobacterium)), clade 15 (Candidate Rhizobacterium)), clade 16 (Candidate Rhizobacterium)), clade 17 (Candidate Rhizobacterium)), clade 18 (Candidate Rhizobacterium)), clade 19 (Candidate Rhizobacterium)), clade 20 (Candidate Rhizobacterium)), clade 21 (Candidate Rhizobacterium)), clade 22 (Candidate Rhizobacterium)), clade 23 (Candidate Rhizobacterium)), clade 24 (Candidate Rhizobacterium)), clade 25 (Candidate Rhizobacterium)), clade 26 (Candidate Rhizobacterium)), clade 27 (Candidate Rhizobacterium)), clade 28 Group III included clade 7 (class Bacilli (genus Bacillus, Acidi)), clade 8 (class Clostridia (genus Dorea, genus Ruminococcus, genus Clostridium, genus Clostridioides, genus Peptoclostridium, genus Cellulosilyticym, genus Eubacterium)), clade 9 (class Bacilli (genus Bacillus, Acidi)), clade 10 (class Bacilli (genus Bacillus, Acidi)), clade 11 (class Bacilli (genus Bacillus, Acidi)), clade 12 (class Bacilli (genus Bacillus, Acidi)), clade 13 (class Bacilli (genus Bacillus, Acidi)), clade 14 (class Bacilli (genus Bacillus, Acidi)), clade 15 (class Bacilli (genus Bacillus, Acidi)), clade 16 (class Bacilli (genus Bacillus, Acidi)), clade 17 (class Bacilli (genus Bacillus, Acidi)), clade 18 (class Bacilli (genus Bacillus, Acidi)), clade 19 (class Bacilli (genus Bacillus, Acidi)), clade 20 (class Bacilli (genus Bacillus, Acidi)), clade 21 (class Bacilli (genus Bacillus, Acidi)), clade 22 (class Bacilli (genus Bacillus, Acidi)), clade 23 (class Bacilli (genus Bacillus, Acidi)), clade 24 (class Bacilli (genus Bacillus, Acidi)), clade 25 (class Bacilli (genus Bacillus, Acidi)), clade 26 (class Bacilli (genus Bacillus, Acidi)), clade 27 (class Bacilli (genus Bacillus, The genera included Bacillus (Acidibacillus), Aneurinibacillus, Brevibacillus, Parageobacillus, and Alicyclobacillus, clade 8 (Negativicutes (Phascolarctobacterium)), and clade 9 (Flavobacteriia (Flavobacterium)).Diamond symbols represent the Cas-alpha 1-11 endonucleases described herein. [Figure 21-1] Figure 21: Figure 21A shows a transposase (Tnp)-associated Cas-alpha CRISPR system. In both cases, a Tnp-like protein is encoded upstream of the Cas-alpha endonuclease and CRISPR array. Figure 21B shows the Cas-alpha endonuclease and guide RNA complexed with the Tnp-like protein positioned to integrate a DNA payload (dashed circle) into or near its target site and the Cas-alpha double-stranded DNA target site. [Figure 21-2] (As mentioned above.) DETAILED DESCRIPTION OF THE INVENTION

[0024] The sequence descriptions and the sequence listing attached hereto comply with the rules governing the disclosure of nucleotide and amino acid sequences in patent applications as set forth in 37 C.F.R. §§ 1.821 and 1.825. The sequence descriptions include the three-letter codes for amino acids as set forth in 37 C.F.R. §§ 1.821 and 1.825, which are incorporated herein by reference.

[0025] SEQ ID NO: 1 is Cas1 encoded by the Cas-alpha1 locus PRT sequence from Candidatus Micrarchaeota archaeon.

[0026] SEQ ID NO: 2 is Cas1 encoded by the Cas-alpha2 locus PRT sequence from Candidatus Micrarchaeota archaeon.

[0027] SEQ ID NO: 3 is Cas1 encoded by the Cas-alpha3 locus PRT sequence from Candidatus Aureabacteria bacterium.

[0028] SEQ ID NO: 4 is Cas1 encoded by the Cas-alpha4 locus PRT sequence from an uncultured archaea.

[0029] SEQ ID NO: 5 is the Cas2 encoded by the Cas-alpha1 locus PRT sequence from Candidatus Micrarchaeota archaeon.

[0030] SEQ ID NO: 6 is Cas2 encoded by the Cas-alpha2 locus PRT sequence from Candidatus Micrarchaeota archaeon.

[0031] SEQ ID NO: 7 is the Cas2 encoded by the Cas-alpha3 locus PRT sequence from Candidatus Aureabacteria bacterium.

[0032] SEQ ID NO: 8 is Cas2 encoded by the Cas-alpha4 locus PRT sequence from an uncultured archaea.

[0033] SEQ ID NO: 9 is the Cas4 encoded by the Cas-alpha1 locus PRT sequence from Candidatus Micrarchaeota archaeon.

[0034] SEQ ID NO: 10 is the Cas4 encoded by the Cas-alpha2 locus PRT sequence from Candidatus Micrarchaeota archaeon.

[0035] SEQ ID NO: 11 is the Cas4 encoded by the Cas-alpha3 locus PRT sequence from Candidatus Aureabacteria bacterium.

[0036] SEQ ID NO: 12 is Cas4 encoded by the Cas-alpha4 locus PRT sequence from an uncultured archaea.

[0037] SEQ ID NO: 13 is the DNA sequence of the Cas-alpha 1 endonuclease gene from Candidatus Micrarchaeota archaeon.

[0038] SEQ ID NO: 14 is the DNA sequence of the Cas-alpha 2 endonuclease gene from Candidatus Micrarchaeota archaeon.

[0039] SEQ ID NO: 15 is the DNA sequence of the Cas-alpha3 endonuclease gene from Candidatus Aureabacteria bacterium.

[0040] SEQ ID NO: 16 is the DNA sequence of a Cas-alpha 4 endonuclease gene from an uncultured archaea.

[0041] SEQ ID NO: 17 is the Cas-alpha 1 endonuclease (Cas14b4) PRT sequence from Candidatus Micrarchaeota archaeon.

[0042] SEQ ID NO: 18 is the Cas-alpha 2 endonuclease PRT sequence from Candidatus Micrarchaeota archaeon.

[0043] SEQ ID NO: 19 is the Cas-alpha3 endonuclease PRT sequence from Candidatus Aureabacteria bacterium.

[0044] SEQ ID NO: 20 is the Cas-alpha 4 endonuclease (Cas14a1) PRT sequence from an uncultured archaea.

[0045] SEQ ID NO: 21 is the Cas-alpha 1 locus DNA sequence from Candidatus Micrarchaeota archaeon.

[0046] SEQ ID NO: 22 is the Cas-alpha2 locus DNA sequence from Candidatus Micrarchaeota archaeon.

[0047] SEQ ID NO: 23 is the Cas-alpha3 locus DNA sequence from Candidatus Aureabacteria bacterium.

[0048] SEQ ID NO: 24 is the DNA sequence of the Cas-alpha4 locus from an uncultured archaea.

[0049] SEQ ID NO: 25 is the DNA sequence of the Cas-alpha 5 endonuclease gene from Candidatus Micrarchaeota archaeon.

[0050] SEQ ID NO: 26 is the DNA sequence of a Cas-alpha 6 endonuclease gene from an uncultured archaea.

[0051] SEQ ID NO: 27 is the DNA sequence of the Cas-alpha 7 endonuclease gene from Parageobacillus thermoglucosidasius.

[0052] SEQ ID NO: 28 is the DNA sequence of the Cas-alpha 8 endonuclease gene from Acidibacillus sulfuroxidans.

[0053] SEQ ID NO: 29 is the DNA sequence of the Cas-alpha 9 endonuclease gene from Ruminococcus sp.

[0054] SEQ ID NO: 30 is the DNA sequence of the Cas-alpha 10 endonuclease gene from Syntrophomonas palmitatica.

[0055] SEQ ID NO: 31 is the DNA sequence of the Cas-alpha11 endonuclease gene from Clostridium novyi.

[0056] SEQ ID NO: 32 is the Cas-alpha 5 endonuclease PRT sequence from Candidatus Micrarchaeota archaeon.

[0057] SEQ ID NO: 33 is the sequence of a Cas-alpha 6 endonuclease PRT from an uncultured archaea.

[0058] SEQ ID NO: 34 is the Cas-alpha 7 endonuclease PRT sequence from Parageobacillus thermoglucosidasius.

[0059] SEQ ID NO: 35 is the Cas-alpha 8 endonuclease PRT sequence from Acidibacillus sulfuroxidans.

[0060] SEQ ID NO: 36 is the Cas-alpha 9 endonuclease PRT sequence from Ruminococcus sp.

[0061] SEQ ID NO: 37 is the Cas-alpha 10 endonuclease PRT sequence from Syntrophomonas palmitatica.

[0062] SEQ ID NO: 38 is the Cas-alpha11 endonuclease PRT sequence from Clostridium novyi.

[0063] SEQ ID NO: 39 is the Cas-alpha 5 locus DNA sequence from Candidatus Micrarchaeota archaeon.

[0064] SEQ ID NO: 40 is the Cas-alpha6 locus DNA sequence from an uncultured archaea.

[0065] SEQ ID NO: 41 is the Cas-alpha 7 locus DNA sequence from Parageobacillus thermoglucosidasius.

[0066] SEQ ID NO: 42 is the Cas-alpha 8 locus DNA sequence from Acidibacillus sulfuroxidans.

[0067] SEQ ID NO: 43 is the Cas-alpha9 locus DNA sequence from Ruminococcus sp.

[0068] SEQ ID NO: 44 is the Cas-alpha 10 locus DNA sequence from Syntrophomonas palmitatica.

[0069] SEQ ID NO: 45 is the Cas-alpha11 locus DNA sequence from Clostridium novyi.

[0070] SEQ ID NO: 46 is the Cas-alpha 1 repeat consensus DNA sequence from Candidatus Micrarchaeota archaeon.

[0071] SEQ ID NO: 47 is the Cas-alpha 2 repeat consensus DNA sequence from Candidatus Micrarchaeota archaeon.

[0072] SEQ ID NO: 48 is the Cas-alpha 3 repeat consensus DNA sequence from Candidatus Aureabacteria bacterium.

[0073] SEQ ID NO: 49 is a Cas-alpha 4 repeat consensus DNA sequence from an uncultured archaea.

[0074] SEQ ID NO: 50 is the Cas-alpha 5 repeat consensus DNA sequence from Candidatus Micrarchaeota archaeon.

[0075] SEQ ID NO: 51 is a Cas-alpha 6 repeat consensus DNA sequence from an uncultured archaea.

[0076] SEQ ID NO: 52 is the Cas-alpha 7 repeat consensus DNA sequence from Parageobacillus thermoglucosidasius.

[0077] SEQ ID NO: 53 is the Cas-alpha 8 repeat consensus DNA sequence from Acidibacillus sulfuroxidans.

[0078] SEQ ID NO: 54 is the Cas-alpha 9 repeat consensus DNA sequence from Ruminococcus sp.

[0079] SEQ ID NO: 55 is the Cas-alpha 10 repeat consensus DNA sequence from Syntrophomonas palmitatica.

[0080] SEQ ID NO: 56 is the Cas-alpha11 repeat consensus DNA sequence from Clostridium novyi.

[0081] SEQ ID NO: 57 is the artificially derived Cas-alpha1 crRNA (where N represents any nucleotide) RNA sequence.

[0082] SEQ ID NO: 58 is the artificially derived Cas-alpha2 crRNA (where N represents any nucleotide) RNA sequence.

[0083] SEQ ID NO: 59 is the artificially derived Cas-alpha4 crRNA (where N represents any nucleotide) RNA sequence.

[0084] SEQ ID NO: 60 is the Cas-alpha1 tracrRNA version 1 RNA sequence from Candidatus Micrarchaeota archaeon.

[0085] SEQ ID NO: 61 is the Cas-alpha1 tracrRNA version 2 RNA sequence from Candidatus Micrarchaeota archaeon.

[0086] SEQ ID NO: 62 is the Cas-alpha1 tracrRNA version 3 RNA sequence from Candidatus Micrarchaeota archaeon.

[0087] SEQ ID NO: 63 is the Cas-alpha1 tracrRNA version 4 RNA sequence from Candidatus Micrarchaeota archaeon.

[0088] SEQ ID NO: 64 is the Cas-alpha2 tracrRNA version 1 RNA sequence from Candidatus Micrarchaeota archaeon.

[0089] SEQ ID NO: 65 is the Cas-alpha2 tracrRNA version 2 RNA sequence from Candidatus Micrarchaeota archaeon.

[0090] SEQ ID NO: 66 is the Cas-alpha2 tracrRNA version 3 RNA sequence from Candidatus Micrarchaeota archaeon.

[0091] SEQ ID NO: 67 is the Cas-alpha2 tracrRNA version 4 RNA sequence from Candidatus Micrarchaeota archaeon.

[0092] SEQ ID NO: 68 is the Cas-alpha4 tracrRNA version 1 RNA sequence from an uncultured archaea.

[0093] SEQ ID NO: 69 is the artificially derived Cas-alpha1 sgRNA version 1 RNA sequence.

[0094] SEQ ID NO: 70 is the artificially derived Cas-alpha1 sgRNA version 2 RNA sequence.

[0095] SEQ ID NO: 71 is the artificially derived Cas-alpha1 sgRNA version 3 RNA sequence.

[0096] SEQ ID NO: 72 is the artificially derived Cas-alpha1 sgRNA version 4 RNA sequence.

[0097] SEQ ID NO: 73 is the artificially derived Cas-alpha2 sgRNA version 1 RNA sequence.

[0098] SEQ ID NO: 74 is the artificially derived Cas-alpha2 sgRNA version 2 RNA sequence.

[0099] SEQ ID NO: 75 is the artificially derived Cas-alpha2 sgRNA version 3 RNA sequence.

[0100] SEQ ID NO: 76 is the artificially derived Cas-alpha2 sgRNA version 4 RNA sequence.

[0101] SEQ ID NO: 77 is the artificially derived Cas-alpha4 sgRNA version 1 RNA sequence.

[0102] SEQ ID NO: 78 is the artificially derived T2 spacer DNA sequence.

[0103] SEQ ID NO: 79 is an artificially derived complete Cas-alpha1 locus engineered to target T2 DNA sequences.

[0104] SEQ ID NO: 80 is an artificially derived minimal Cas-alpha1 locus engineered to target T2 DNA sequences.

[0105] SEQ ID NO: 81 is the artificially derived 10x histidine tag PRT sequence.

[0106] SEQ ID NO: 82 is the artificially derived 6x histidine tag PRT sequence.

[0107] SEQ ID NO: 83 is the artificially derived maltose binding protein tag PRT sequence.

[0108] SEQ ID NO: 84 is the Tobacco etch virus cleavage site PRT sequence from Tobacco etch virus.

[0109] SEQ ID NO: 85 is the artificially derived A1 oligonucleotide DNA sequence.

[0110] SEQ ID NO: 86 is the artificially derived A2 oligonucleotide DNA sequence.

[0111] SEQ ID NO: 87 is the artificially derived R0 oligonucleotide DNA sequence.

[0112] SEQ ID NO: 88 is the artificially derived C0 oligonucleotide DNA sequence.

[0113] SEQ ID NO: 89 is the artificially derived F1 oligonucleotide DNA sequence.

[0114] SEQ ID NO: 90 is the artificially derived R1 oligonucleotide DNA sequence.

[0115] SEQ ID NO: 91 is the bridge amplification portion of the artifact-derived F1 oligonucleotide DNA sequence.

[0116] SEQ ID NO: 92 is the bridge amplification portion of the artifact-derived R1 oligonucleotide DNA sequence.

[0117] SEQ ID NO: 93 is the artificially derived F2 oligonucleotide DNA sequence.

[0118] SEQ ID NO: 94 is the artificially derived R2 oligonucleotide DNA sequence.

[0119] SEQ ID NO: 95 is the artificially derived C1 oligonucleotide DNA sequence.

[0120] SEQ ID NO: 96 is the sequence obtained from cleavage of the artifact-derived target DNA sequence at position 21 and adapter ligation.

[0121] SEQ ID NO: 97 is the adapter portion of the DNA sequence of SEQ ID NO: 96, which is derived from an artificial product.

[0122] SEQ ID NO: 98 is the target portion of the DNA sequence of SEQ ID NO: 96 derived from an artifact.

[0123] SEQ ID NO: 99 is the sequence 5' of the artificially derived PAM DNA sequence.

[0124] SEQ ID NO: 100 is an artificially derived fixed double-stranded DNA target DNA sequence.

[0125] SEQ ID NO: 101 is the artificially derived T2 target sequence DNA sequence.

[0126] SEQ ID NO: 102 is the artificially derived Cas-alpha4 T2-1 sgRNA RNA sequence.

[0127] SEQ ID NO: 103 is the artificially derived Cas-alpha4 T2-2 sgRNA RNA sequence.

[0128] SEQ ID NO: 104 is the artificially derived Cas-alpha4 T2-1 crRNA RNA sequence.

[0129] SEQ ID NO: 105 is the artificially derived Cas-alpha4 T2-2 crRNA RNA sequence.

[0130] SEQ ID NO: 106 is the ST-LS1 intron 2 DNA sequence from potato (Solanum tuberosum).

[0131] SEQ ID NO: 107 is the SV40 NLS PRT sequence from Simian virus 40.

[0132] SEQ ID NO: 108 is the Nuc NLS PRT sequence from Mus musculus.

[0133] SEQ ID NO: 109 is the maize UBI promoter DNA sequence from Zea mays.

[0134] SEQ ID NO: 110 is the chicken beta-actin promoter DNA sequence from red jungle fowl (Gallus gallus).

[0135] SEQ ID NO: 111 is the CMV enhancer DNA sequence from human beta-herpesvirus 5.

[0136] SEQ ID NO: 112 is the maize UBI 5 prime untranslated region DNA sequence from Zea mays.

[0137] SEQ ID NO: 113 is the maize UBI intron 1 DNA sequence from Zea mays.

[0138] SEQ ID NO: 114 is an artificially derived hybrid intron DNA sequence.

[0139] SEQ ID NO: 115 is the maize U6 polymerase III promoter DNA sequence from Zea mays.

[0140] SEQ ID NO: 116 is the human U6 polymerase III promoter DNA sequence from Homo sapiens.

[0141] SEQ ID NO: 117 is the artificially derived Strep II tag PRT sequence.

[0142] SEQ ID NO: 118 is the bGH poly(A) terminator DNA sequence from Bos taurus.

[0143] SEQ ID NO: 119 is the potato protease inhibitor II (Pin II) terminator DNA sequence from potato (Solanum tuberosum).

[0144] SEQ ID NO: 120 is the Zea mays Wt Reference (Liguleless Targets 2 and 3) DNA sequence from Zea mays.

[0145] SEQ ID NO: 121 is the Mutant 1 (Liguleless Targets 2 and 3-DNA Exp.) DNA sequence from Zea mays.

[0146] SEQ ID NO: 122 is the Mutant 2 (Liguleless Target 2 and 3-DNA Exp.) DNA sequence from Zea mays.

[0147] SEQ ID NO: 123 is the DNA sequence of Mutant 3 (Liguleless Target 2 and 3-DNA Exp.) from Zea mays.

[0148] SEQ ID NO: 124 is the DNA sequence of Mutant 4 (Liguleless Targets 2 and 3-DNA Exp.) from Zea mays.

[0149] SEQ ID NO: 125 is the DNA sequence of Mutant 5 (Liguleless Targets 2 and 3-DNA Exp.) from Zea mays.

[0150] SEQ ID NO: 126 is the HEK293 Wt Reference (VEGFA Target 2) DNA sequence from Homo sapiens.

[0151] SEQ ID NO: 127 is the mutation 1 (VEGFA target 2-RNP) DNA sequence from Homo sapiens.

[0152] SEQ ID NO: 128 is the mutation 2 (VEGFA target 2-RNP) DNA sequence from Homo sapiens.

[0153] SEQ ID NO: 129 is the mutation 3 (VEGFA target 2-RNP) DNA sequence from Homo sapiens.

[0154] SEQ ID NO: 130 is the mutation 4 (VEGFA target 2-RNP) DNA sequence from Homo sapiens.

[0155] SEQ ID NO: 131 is the DNA sequence of variant 5 (VEGFA target 2-RNP) from Homo sapiens.

[0156] SEQ ID NO: 132 is the HEK293 Wt Reference (VEGFA Target 3) DNA sequence from Homo sapiens.

[0157] SEQ ID NO: 133 is the mutation 1 (VEGFA target 3-RNP) DNA sequence from Homo sapiens.

[0158] SEQ ID NO: 134 is the mutation 1 (VEGFA target 3-DNA Exp) DNA sequence from Homo sapiens.

[0159] SEQ ID NO: 135 is the mutation 2 (VEGFA target 3-DNA Exp) DNA sequence from Homo sapiens.

[0160] SEQ ID NO: 136 is the ROX3 promoter DNA sequence from Saccharomyces cerevisiae.

[0161] SEQ ID NO: 137 is the GAL promoter DNA sequence from Saccharomyces cerevisiae.

[0162] SEQ ID NO: 138 is the DNA sequence of an artificially derived HH ribozyme (where N represents the nucleotide complementary to the 6 nucleotides 3' of the ribozyme).

[0163] SEQ ID NO: 139 is the HDV ribozyme DNA sequence derived from Hepatitis delta virus.

[0164] SEQ ID NO: 140 is the SNR52 promoter DNA sequence from Saccharomyces cerevisiae.

[0165] SEQ ID NO: 141 is the SUP4 terminator DNA sequence from Saccharomyces cerevisiae.

[0166] SEQ ID NO: 142 is the DNA sequence shown in the upper part of Figure 15C, which is derived from an artifact.

[0167] SEQ ID NO: 143 is the DNA sequence shown in the bottom of Figure 15C, which is derived from an artifact.

[0168] SEQ ID NO: 144 is the reference DNA sequence shown in Figure 18A derived from Zea mays.

[0169] SEQ ID NO: 145 is the mutation 1 DNA sequence from Zea mays.

[0170] SEQ ID NO: 146 is the mutation 2 DNA sequence from Zea mays.

[0171] SEQ ID NO: 147 is the mutation 3 DNA sequence from Zea mays.

[0172] SEQ ID NO: 148 is the mutation 4 DNA sequence from Zea mays.

[0173] SEQ ID NO: 149 is the DNA sequence of variant 5 from Zea mays.

[0174] SEQ ID NO: 150 is the DNA sequence of Mutant 6 from Zea mays.

[0175] SEQ ID NO: 151 is the DNA sequence of Mutant 7 from Zea mays.

[0176] SEQ ID NO: 152 is the DNA sequence of Mutant 8 from Zea mays.

[0177] SEQ ID NO: 153 is the DNA sequence of Mutant 9 from Zea mays.

[0178] SEQ ID NO: 154 is the DNA sequence of mutant 10 from Zea mays.

[0179] SEQ ID NO: 155 is the DNA sequence of mutation 11 from Zea mays.

[0180] SEQ ID NO: 156 is the DNA sequence of mutation 12 from Zea mays.

[0181] SEQ ID NO: 157 is the DNA sequence of mutation 13 from Zea mays.

[0182] SEQ ID NO: 158 is the DNA sequence of mutation 14 from Zea mays.

[0183] SEQ ID NO: 159 is the DNA sequence of mutation 15 from Zea mays.

[0184] SEQ ID NO: 160 is the DNA sequence of Mutant 16 from Zea mays.

[0185] SEQ ID NO: 161 is the DNA sequence of Mutant 17 from Zea mays.

[0186] SEQ ID NO: 162 is the DNA sequence of Mutant 18 from Zea mays.

[0187] SEQ ID NO: 163 is the DNA sequence of mutation 19 from Zea mays.

[0188] SEQ ID NO: 164 is the reference DNA sequence shown in Figure 18B from Zea mays.

[0189] SEQ ID NO: 165 is the mutation 1 DNA sequence from Zea mays.

[0190] SEQ ID NO: 166 is the mutation 2 DNA sequence from Zea mays.

[0191] SEQ ID NO: 167 is the mutation 3 DNA sequence from Zea mays.

[0192] SEQ ID NO: 168 is the mutation 4 DNA sequence from Zea mays.

[0193] SEQ ID NO: 169 is the mutation 5 DNA sequence from Zea mays.

[0194] SEQ ID NO: 170 is the reference DNA sequence shown in Figure 19C from Saccharomyces cerevisiae.

[0195] SEQ ID NO: 171 is the artificially derived repair template DNA sequence.

[0196] SEQ ID NO: 172 is the repair result 1 DNA sequence from Saccharomyces cerevisiae.

[0197] SEQ ID NO: 173 is the repair result 2 DNA sequence from Saccharomyces cerevisiae.

[0198] SEQ ID NO: 174 is the repair result 1 DNA sequence from Saccharomyces cerevisiae.

[0199] SEQ ID NO: 175 is the repair result 1 DNA sequence from Saccharomyces cerevisiae.

[0200] SEQ ID NO: 176 is the repair result 2 DNA sequence from Saccharomyces cerevisiae.

[0201] SEQ ID NO: 177 is the artificially derived Cas-alpha3 crRNA (where N represents any nucleotide) RNA sequence.

[0202] SEQ ID NO: 178 is the artificially derived Cas-alpha5 crRNA (where N represents any nucleotide) RNA sequence.

[0203] SEQ ID NO: 179 is the artificially derived Cas-alpha6 crRNA (where N represents any nucleotide) RNA sequence.

[0204] SEQ ID NO: 180 is the artificially derived Cas-alpha7 crRNA (where N represents any nucleotide) RNA sequence.

[0205] SEQ ID NO: 181 is the artificially derived Cas-alpha8 crRNA (where N represents any nucleotide) RNA sequence.

[0206] SEQ ID NO: 182 is the artificially derived Cas-alpha9 crRNA (where N represents any nucleotide) RNA sequence.

[0207] SEQ ID NO: 183 is the artificially derived Cas-alpha10 crRNA (where N represents any nucleotide) RNA sequence.

[0208] SEQ ID NO: 184 is the artificially derived Cas-alpha11 crRNA (where N represents any nucleotide) RNA sequence.

[0209] SEQ ID NO: 185 is the Cas-alpha2 tracrRNA version 5 RNA sequence from Candidatus Micrarchaeota archaeon.

[0210] SEQ ID NO: 186 is the Cas-alpha2 tracrRNA version 6 RNA sequence from Candidatus Micrarchaeota archaeon.

[0211] SEQ ID NO: 187 is the Cas-alpha2 tracrRNA version 7 RNA sequence from Candidatus Micrarchaeota archaeon.

[0212] SEQ ID NO: 188 is the Cas-alpha6 tracrRNA version 1 RNA sequence from an uncultured archaea.

[0213] SEQ ID NO: 189 is the Cas-alpha6 tracrRNA version 2 RNA sequence from an uncultured archaea.

[0214] SEQ ID NO: 190 is the Cas-alpha6 tracrRNA version 3 RNA sequence from an uncultured archaea.

[0215] SEQ ID NO: 191 is the Cas-alpha6 tracrRNA version 4 RNA sequence from an uncultured archaea.

[0216] SEQ ID NO: 192 is the Cas-alpha7 tracrRNA version 1 RNA sequence from Parageobacillus thermoglucosidasius.

[0217] SEQ ID NO: 193 is the Cas-alpha7 tracrRNA version 2 RNA sequence from Parageobacillus thermoglucosidasius.

[0218] SEQ ID NO: 194 is the Cas-alpha8 tracrRNA version 1 RNA sequence from Acidibacillus sulfuroxidans.

[0219] SEQ ID NO: 195 is the Cas-alpha8 tracrRNA version 2 RNA sequence from Acidibacillus sulfuroxidans.

[0220] SEQ ID NO: 196 is the Cas-alpha8 tracrRNA version 3 RNA sequence from Acidibacillus sulfuroxidans.

[0221] SEQ ID NO: 197 is the Cas-alpha9 tracrRNA version 1 RNA sequence from Ruminococcus sp.

[0222] SEQ ID NO: 198 is the Cas-alpha9 tracrRNA version 2 RNA sequence from Ruminococcus sp.

[0223] SEQ ID NO: 199 is the Cas-alpha10 tracrRNA version 1 RNA sequence from Syntrophomonas palmitatica.

[0224] SEQ ID NO: 200 is the Cas-alpha10 tracrRNA version 2 RNA sequence from Syntrophomonas palmitatica.

[0225] SEQ ID NO: 201 is the Cas-alpha10 tracrRNA version 3 RNA sequence from Syntrophomonas palmitatica.

[0226] SEQ ID NO: 202 is the Cas-alpha10 tracrRNA version 4 RNA sequence from Syntrophomonas palmitatica.

[0227] SEQ ID NO: 203 is the Cas-alpha10 tracrRNA version 5 RNA sequence from Syntrophomonas palmitatica.

[0228] SEQ ID NO: 204 is the Cas-alpha11 tracrRNA version 1 RNA sequence from Clostridium novyi.

[0229] SEQ ID NO: 205 is the Cas-alpha11 tracrRNA version 2 RNA sequence from Clostridium novyi.

[0230] SEQ ID NO: 206 is the Cas-alpha11 tracrRNA version 3 RNA sequence from Clostridium novyi.

[0231] SEQ ID NO: 207 is the Cas-alpha11 tracrRNA version 4 RNA sequence from Clostridium novyi.

[0232] SEQ ID NO: 208 is the artificially derived Cas-alpha2 sgRNA version 5 RNA sequence.

[0233] SEQ ID NO: 209 is the artificially derived Cas-alpha2 sgRNA version 6 RNA sequence.

[0234] SEQ ID NO: 210 is the artificially derived Cas-alpha2 sgRNA version 7 RNA sequence.

[0235] SEQ ID NO: 211 is the artificially derived Cas-alpha6 sgRNA version 1 RNA sequence.

[0236] SEQ ID NO: 212 is the artificially derived Cas-alpha6 sgRNA version 2 RNA sequence.

[0237] SEQ ID NO: 213 is the artificially derived Cas-alpha6 sgRNA version 3 RNA sequence.

[0238] SEQ ID NO: 214 is the artificially derived Cas-alpha6 sgRNA version 4 RNA sequence.

[0239] SEQ ID NO: 215 is the artificially derived Cas-alpha7 sgRNA version 1 RNA sequence.

[0240] SEQ ID NO: 216 is the artificially derived Cas-alpha7 sgRNA version 2 RNA sequence.

[0241] SEQ ID NO: 217 is the artificially derived Cas-alpha7 sgRNA version 3 RNA sequence.

[0242] SEQ ID NO: 218 is the artificially derived Cas-alpha8 sgRNA version 1 RNA sequence.

[0243] SEQ ID NO: 219 is the artificially derived Cas-alpha8 sgRNA version 2 RNA sequence.

[0244] SEQ ID NO: 220 is the artificially derived Cas-alpha8 sgRNA version 3 RNA sequence.

[0245] SEQ ID NO: 221 is the artificially derived Cas-alpha8 sgRNA version 4 RNA sequence.

[0246] SEQ ID NO: 222 is the artificially derived Cas-alpha9 sgRNA version 1 RNA sequence.

[0247] SEQ ID NO: 223 is the artificially derived Cas-alpha9 sgRNA version 2 RNA sequence.

[0248] SEQ ID NO: 224 is the artificially derived Cas-alpha9 sgRNA version 3 RNA sequence.

[0249] SEQ ID NO: 225 is the artificially derived Cas-alpha10 sgRNA version 1 RNA sequence.

[0250] SEQ ID NO: 226 is the artificially derived Cas-alpha10 sgRNA version 2 RNA sequence.

[0251] SEQ ID NO: 227 is the artificially derived Cas-alpha10 sgRNA version 3 RNA sequence.

[0252] SEQ ID NO: 228 is the artificially derived Cas-alpha10 sgRNA version 4 RNA sequence.

[0253] SEQ ID NO: 229 is the artificially derived Cas-alpha10 sgRNA version 5 RNA sequence.

[0254] SEQ ID NO: 230 is the artificially derived Cas-alpha11 sgRNA version 1 RNA sequence.

[0255] SEQ ID NO: 231 is the artificially derived Cas-alpha11 sgRNA version 2 RNA sequence.

[0256] SEQ ID NO: 232 is the artificially derived Cas-alpha11 sgRNA version 3 RNA sequence.

[0257] SEQ ID NO: 233 is the artificially derived Cas-alpha11 sgRNA version 4 RNA sequence.

[0258] SEQ ID NO: 234 is the artificially derived Cas-alpha11 sgRNA version 5 RNA sequence.

[0259] SEQ ID NO: 235 is the engineered Cas-alpha 4 Zea mays codon-optimized gene DNA sequence.

[0260] SEQ ID NO: 236 is the engineered Cas-alpha10 Zea mays codon-optimized gene DNA sequence.

[0261] SEQ ID NO: 237 is the artificially derived Cas-alpha10 Saccharomyces cerevisiae (Saccharomyces cerevisiae) codon-optimized gene DNA sequence.

[0262] SEQ ID NO: 238 is the artificially derived Cas-alpha4 sgRNA backbone RNA sequence.

[0263] SEQ ID NO: 239 is the artificially derived Cas-alpha10 sgRNA backbone RNA sequence.

[0264] SEQ ID NO: 240 is the artifact-derived Cas-alpha4 Liguleless2 sgRNA target sequence RNA sequence.

[0265] SEQ ID NO: 241 is the artifact-derived Cas-alpha4 Liguleless3 sgRNA target sequence RNA sequence.

[0266] SEQ ID NO: 242 is the artificially derived Cas-alpha10 nptII sgRNA target sequence RNA sequence.

[0267] SEQ ID NO: 243 is the artificially derived Cas-alpha10 ms26 sgRNA target sequence RNA sequence.

[0268] SEQ ID NO: 244 is the artificially derived Cas-alpha10 ade2 sgRNA target sequence RNA sequence.

[0269] SEQ ID NO: 245 is the artificially derived Cas-alpha4 VEGFA2 sgRNA target sequence RNA sequence.

[0270] SEQ ID NO: 246 is the artificially derived Cas-alpha4 VEGFA3 sgRNA target sequence RNA sequence.

[0271] SEQ ID NO: 247 is the artifact-derived Cas-alpha4 sgRNA targeting Liguleless2 RNA sequence.

[0272] SEQ ID NO: 248 is the artifact-derived Cas-alpha4 sgRNA targeting Liguleless3 RNA sequence.

[0273] SEQ ID NO: 249 is the artifact-derived Cas-alpha10 sgRNA targeting nptII RNA sequence.

[0274] SEQ ID NO: 250 is the artificially derived Cas-alpha10 sgRNA targeting ms26 RNA sequence.

[0275] SEQ ID NO: 251 is the artificially derived Cas-alpha10 sgRNA targeting ade2 RNA sequence.

[0276] SEQ ID NO: 252 is the artificially derived Cas-alpha4 sgRNA targeting VEGFA2 RNA sequence.

[0277] SEQ ID NO: 253 is the artificially derived Cas-alpha4 sgRNA targeting VEGFA3 RNA sequence.

[0278] SEQ ID NO: 254 is the Cas-alpha 12 endonuclease PRT sequence from Clostridioides difficile.

[0279] SEQ ID NO: 255 is the Cas-alpha13 endonuclease PRT sequence from Clostridium paraputrificum.

[0280] SEQ ID NO: 256 is the Cas-alpha14 endonuclease PRT sequence from Clostridium novyi.

[0281] SEQ ID NO: 257 is the Cas-alpha 15 endonuclease PRT sequence from Ruminococcus albus.

[0282] SEQ ID NO: 258 is the Cas-alpha 16 endonuclease PRT sequence from Clostridium hiranonis.

[0283] SEQ ID NO: 259 is the Cas-alpha 17 endonuclease PRT sequence from Clostridium ifumihumii.

[0284] SEQ ID NO: 260 is the Cas-alpha 18 endonuclease PRT sequence from Cellulosilyticum ruminicola.

[0285] SEQ ID NO: 261 is the Cas-alpha19 endonuclease PRT sequence from Eubacterium siraeum.

[0286] SEQ ID NO: 262 is the Cas-alpha 20 endonuclease PRT sequence from Clostridium botulinum.

[0287] SEQ ID NO: 263 is the Cas-alpha 21 endonuclease PRT sequence from Clostridium botulinum.

[0288] SEQ ID NO: 264 is the Cas-alpha 22 endonuclease PRT sequence from Ruminiclostridium hungatei.

[0289] SEQ ID NO: 265 is the Cas-alpha 23 endonuclease PRT sequence from Desulfovibrio fructosivorans.

[0290] SEQ ID NO: 266 is the Cas-alpha 24 endonuclease PRT sequence from Bacillus toyonensis.

[0291] SEQ ID NO: 267 is the Cas-alpha 25 endonuclease PRT sequence from Clostridium paraputrificum.

[0292] SEQ ID NO: 268 is the Cas-alpha 26 endonuclease PRT sequence from Clostridium ventriculi.

[0293] SEQ ID NO: 269 is the Cas-alpha 27 endonuclease PRT sequence from Ruminococcus sp.

[0294] SEQ ID NO: 270 is the Cas-alpha 28 endonuclease PRT sequence from Ruminococcus sp.

[0295] SEQ ID NO: 271 is the Cas-alpha29 endonuclease PRT sequence from Peptoclostridium sp.

[0296] SEQ ID NO: 272 is the Cas-alpha 30 endonuclease PRT sequence from Bacillus sp.

[0297] SEQ ID NO: 273 is the Cas-alpha 31 endonuclease PRT sequence from Clostridioides difficile.

[0298] SEQ ID NO: 274 is the Cas-alpha32 endonuclease PRT sequence from Clostridioides difficile.

[0299] SEQ ID NO: 275 is the sequence of a Cas-alpha33 endonuclease PRT from an uncultured archaea.

[0300] SEQ ID NO: 276 is the sequence of a Cas-alpha34 endonuclease PRT from an uncultured archaea.

[0301] SEQ ID NO: 277 is the sequence of a Cas-alpha35 endonuclease PRT from an uncultured archaea.

[0302] SEQ ID NO: 278 is the sequence of a Cas-alpha36 endonuclease PRT from an uncultured archaea.

[0303] SEQ ID NO: 279 is the sequence of a Cas-alpha37 endonuclease PRT from an uncultured archaea.

[0304] SEQ ID NO: 280 is the sequence of a Cas-alpha38 endonuclease PRT from an uncultured archaea.

[0305] SEQ ID NO: 281 is the sequence of a Cas-alpha39 endonuclease PRT from an uncultured archaea.

[0306] SEQ ID NO: 282 is a Cas-alpha40 endonuclease PRT sequence from an uncultured archaea.

[0307] SEQ ID NO: 283 is the sequence of a Cas-alpha41 endonuclease PRT from an uncultured archaea.

[0308] SEQ ID NO: 284 is the Cas-alpha42 endonuclease PRT sequence from Clostridioides difficile.

[0309] SEQ ID NO:285 is the Cas-alpha43 endonuclease PRT sequence from Desulfovibrio fructosivorans.

[0310] SEQ ID NO: 286 is the Cas-alpha44 endonuclease PRT sequence from Clostridium botulinum.

[0311] SEQ ID NO: 287 is the Cas-alpha45 endonuclease PRT sequence from Clostridioides difficile.

[0312] SEQ ID NO: 288 is the Cas-alpha46 endonuclease PRT sequence from Clostridioides difficile.

[0313] SEQ ID NO: 289 is the Cas-alpha47 endonuclease PRT sequence from Clostridioides difficile.

[0314] SEQ ID NO: 290 is the Cas-alpha48 endonuclease PRT sequence from Clostridioides difficile.

[0315] SEQ ID NO: 291 is the Cas-alpha49 endonuclease PRT sequence from Clostridioides difficile.

[0316] SEQ ID NO: 292 is the Cas-alpha 50 endonuclease PRT sequence from Clostridioides difficile.

[0317] SEQ ID NO: 293 is the Cas-alpha 51 endonuclease PRT sequence from Clostridioides difficile.

[0318] SEQ ID NO: 294 is the Cas-alpha 52 endonuclease PRT sequence from Clostridioides difficile.

[0319] SEQ ID NO: 295 is the Cas-alpha 53 endonuclease PRT sequence from Clostridioides difficile.

[0320] SEQ ID NO: 296 is the Cas-alpha 54 endonuclease PRT sequence from Clostridioides difficile.

[0321] SEQ ID NO: 297 is the Cas-alpha 55 endonuclease PRT sequence from Clostridium hiranonis.

[0322] SEQ ID NO: 298 is the Cas-alpha 56 endonuclease PRT sequence from Clostridioides difficile.

[0323] SEQ ID NO: 299 is the Cas-alpha 57 endonuclease PRT sequence from Aneurinibacillus danicus.

[0324] SEQ ID NO:300 is the Cas-alpha 58 endonuclease PRT sequence from Parageobacillus thermoglucosidasius.

[0325] SEQ ID NO: 301 is the Cas-alpha 59 endonuclease PRT sequence from Brevibacillus centrosporus.

[0326] SEQ ID NO: 302 is the Cas-alpha 60 endonuclease PRT sequence from Clostridium pasteurianum.

[0327] SEQ ID NO: 303 is the Cas-alpha 61 endonuclease PRT sequence from Eubacterium siraeum.

[0328] SEQ ID NO: 304 is the Cas-alpha 62 endonuclease PRT sequence from Bacillus toyonensis.

[0329] SEQ ID NO: 305 is the Cas-alpha 63 endonuclease PRT sequence from Ruminococcus sp.

[0330] SEQ ID NO: 306 is the Cas-alpha 64 endonuclease PRT sequence from Ruminococcus sp.

[0331] SEQ ID NO: 307 is the Cas-alpha 65 endonuclease PRT sequence from Clostridium perfringens.

[0332] SEQ ID NO: 308 is the Cas-alpha 66 endonuclease PRT sequence from Bacillus thuringiensis.

[0333] SEQ ID NO: 309 is the Cas-alpha 67 endonuclease PRT sequence from Clostridium perfringens.

[0334] SEQ ID NO: 310 is the Cas-alpha 68 endonuclease PRT sequence from Bacillus cereus.

[0335] SEQ ID NO: 311 is the Cas-alpha 69 endonuclease PRT sequence from Bacillus toyonensis.

[0336] SEQ ID NO: 312 is the Cas-alpha 70 endonuclease PRT sequence from Bacillus toyonensis.

[0337] SEQ ID NO: 313 is the Cas-alpha 71 endonuclease PRT sequence from Bacillus toyonensis.

[0338] SEQ ID NO:314 is the Cas-alpha 72 endonuclease PRT sequence from Alicyclobacillus acidoterrestris.

[0339] SEQ ID NO: 315 is the Cas-alpha73 endonuclease PRT sequence from Clostridium tetani.

[0340] SEQ ID NO: 316 is the Cas-alpha 74 endonuclease PRT sequence from Candidatus Levybacteria bacterium.

[0341] SEQ ID NO: 317 is the Cas-alpha 75 endonuclease PRT sequence from Bacillus cereus.

[0342] SEQ ID NO: 318 is the Cas-alpha 76 endonuclease PRT sequence from Bacillus cereus.

[0343] SEQ ID NO: 319 is the Cas-alpha 77 endonuclease PRT sequence from Bacillus cereus.

[0344] SEQ ID NO: 320 is the Cas-alpha 78 endonuclease PRT sequence from Clostridium paraputrificum.

[0345] SEQ ID NO: 321 is the Cas-alpha 79 endonuclease PRT sequence from Bacillus cereus.

[0346] SEQ ID NO: 322 is the Cas-alpha 80 endonuclease PRT sequence from Bacillus thuringiensis.

[0347] SEQ ID NO: 323 is the Cas-alpha 81 endonuclease PRT sequence from Bacillus cereus.

[0348] SEQ ID NO: 324 is the Cas-alpha 82 endonuclease PRT sequence from Bacillus toyonensis.

[0349] SEQ ID NO: 325 is the Cas-alpha 83 endonuclease PRT sequence from Bacillus cereus.

[0350] SEQ ID NO: 326 is the Cas-alpha 84 endonuclease PRT sequence from Bacillus toyonensis.

[0351] SEQ ID NO: 327 is the Cas-alpha 85 endonuclease PRT sequence from Bacillus wiedmannii.

[0352] SEQ ID NO: 328 is the Cas-alpha 86 endonuclease PRT sequence from Bacillus cereus.

[0353] SEQ ID NO: 329 is the Cas-alpha 87 endonuclease PRT sequence from Bacillus cereus.

[0354] SEQ ID NO: 330 is the Cas-alpha 88 endonuclease PRT sequence from Bacillus toyonensis.

[0355] SEQ ID NO: 331 is the Cas-alpha 89 endonuclease PRT sequence from Bacillus cereus.

[0356] SEQ ID NO: 332 is the Cas-alpha 90 endonuclease PRT sequence from Bacillus toyonensis.

[0357] SEQ ID NO: 333 is the Cas-alpha91 endonuclease PRT sequence from Bacillus thuringiensis.

[0358] SEQ ID NO: 334 is the Cas-alpha92 endonuclease PRT sequence from Bacillus cereus.

[0359] SEQ ID NO: 335 is the Cas-alpha93 endonuclease PRT sequence from Bacillus cereus.

[0360] SEQ ID NO: 336 is the Cas-alpha94 endonuclease PRT sequence from Bacillus cereus.

[0361] SEQ ID NO: 337 is the Cas-alpha95 endonuclease PRT sequence from Bacillus thuringiensis.

[0362] SEQ ID NO: 338 is the Cas-alpha96 endonuclease PRT sequence from Bacillus sp.

[0363] SEQ ID NO: 339 is the Cas-alpha97 endonuclease PRT sequence from Bacillus cereus.

[0364] SEQ ID NO:340 is the Cas-alpha98 endonuclease PRT sequence from Bacillus cereus.

[0365] SEQ ID NO: 341 is the Cas-alpha 99 endonuclease PRT sequence from Bacillus thuringiensis.

[0366] SEQ ID NO: 342 is the Cas-alpha 100 endonuclease PRT sequence from Bacillus sp.

[0367] SEQ ID NO: 343 is the Cas-alpha 101 endonuclease PRT sequence from Prevotella copri.

[0368] SEQ ID NO: 344 is the Cas-alpha 102 endonuclease PRT sequence from Prevotella copri.

[0369] SEQ ID NO: 345 is the Cas-alpha 103 endonuclease PRT sequence from Clostridioides difficile.

[0370] SEQ ID NO: 346 is the Cas-alpha 104 endonuclease PRT sequence from Clostridioides difficile.

[0371] SEQ ID NO: 347 is the Cas-alpha 105 endonuclease PRT sequence from Clostridioides difficile.

[0372] SEQ ID NO: 348 is the Cas-alpha 106 endonuclease PRT sequence from Clostridioides difficile.

[0373] SEQ ID NO: 349 is the Cas-alpha 107 endonuclease PRT sequence from Clostridioides difficile.

[0374] SEQ ID NO:350 is the Cas-alpha 108 endonuclease PRT sequence from Clostridioides difficile.

[0375] SEQ ID NO: 351 is the Cas-alpha 109 endonuclease PRT sequence from Clostridioides difficile.

[0376] SEQ ID NO:352 is the Cas-alpha110 endonuclease PRT sequence from Flavobacterium thermophilum.

[0377] SEQ ID NO: 353 is the Cas-alpha111 endonuclease PRT sequence from Phascolarctobacterium sp.

[0378] SEQ ID NO: 354 is the Cas-alpha 112 endonuclease PRT sequence from Bacillus pseudomycoides.

[0379] SEQ ID NO: 355 is the Cas-alpha 113 endonuclease PRT sequence from Bacteroides plebeius.

[0380] SEQ ID NO: 356 is the Cas-alpha 114 endonuclease PRT sequence from Clostridium botulinum.

[0381] SEQ ID NO: 357 is the Cas-alpha 115 endonuclease PRT sequence from Bacillus pseudomycoides.

[0382] SEQ ID NO: 358 is the Cas-alpha 116 endonuclease PRT sequence from Bacillus pseudomycoides.

[0383] SEQ ID NO: 359 is the Cas-alpha 117 endonuclease PRT sequence from Clostridium botulinum.

[0384] SEQ ID NO: 360 is the Cas-alpha 118 endonuclease PRT sequence from Clostridium botulinum.

[0385] SEQ ID NO: 361 is the Cas-alpha 119 endonuclease PRT sequence from Clostridium botulinum.

[0386] SEQ ID NO: 362 is the Cas-alpha120 endonuclease PRT sequence from Hydrogenivirga sp.

[0387] SEQ ID NO: 363 is the Cas-alpha 121 endonuclease PRT sequence from Bacillus megaterium.

[0388] SEQ ID NO: 364 is the Cas-alpha 122 endonuclease PRT sequence from Clostridium fallax.

[0389] SEQ ID NO: 365 is the Cas-alpha 123 endonuclease PRT sequence from Bacteroides plebeius.

[0390] SEQ ID NO: 366 is the Cas-alpha 124 endonuclease PRT sequence from Bacillus thuringiensis.

[0391] SEQ ID NO:367 is the Cas-alpha 125 endonuclease PRT sequence from Bacillus cereus.

[0392] SEQ ID NO: 368 is the Cas-alpha 126 endonuclease PRT sequence from Clostridium sp.

[0393] SEQ ID NO: 369 is the Cas-alpha 127 endonuclease PRT sequence from Bacteroides plebeius.

[0394] SEQ ID NO: 370 is the Cas-alpha 128 endonuclease PRT sequence from Dorea longicatena.

[0395] SEQ ID NO:371 is the Cas-alpha 129 endonuclease PRT sequence from Sulfurihydrogenibium azorense.

[0396]

[0003] Compositions and methods are provided for novel CRISPR effector systems and elements comprising such systems, including, but not limited to, novel guide polynucleotide / endonuclease complexes, guide polynucleotides, guide RNA elements, Cas proteins, and endonucleases, as well as proteins comprising endonuclease functionality (domains). Also provided are compositions and methods for the direct delivery of endonucleases, cleavage-ready complexes, guide RNAs, and guide RNA / Cas endonuclease complexes. The present disclosure further includes compositions and methods for modifying target sequences in the genome of a cell, gene editing, and inserting a polynucleotide of interest into the genome of a cell.

[0397] Terms used in the claims and specification are defined as set forth below unless otherwise specified. It must be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0398] definition As used herein, "nucleic acid" refers to a polynucleotide, including single- or double-stranded polymers of deoxyribonucleotide or ribonucleotide bases. Nucleic acids can also include fragments and modified nucleotides. Thus, the terms "polynucleotide," "nucleic acid sequence," "nucleotide sequence," and "nucleic acid fragment" are used interchangeably to refer to polymers of RNA and / or DNA and / or RNA-DNA that are single- or double-stranded, and optionally contain synthetic, non-natural, or modified nucleotide bases. Nucleotides (usually found in their 5'-monophosphate form) are referred to by the single letter designation: "A" for adenosine or deoxyadenosine (for RNA or DNA, respectively), "C" for cytosine or deoxycytosine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "R" for purine (A or G), "Y" for pyrimidine (C or T), "K" for G or T, "H" for A or C or T, "I" for inosine, and "N" for any nucleotide.

[0399] The term "genome" as applied to prokaryotic and eukaryotic cells or organismal cells encompasses not only chromosomal DNA found in the nucleus, but also organelle DNA found within cellular components of the cell (e.g., mitochondria or plastids).

[0400] "Open reading frame" is abbreviated as ORF.

[0401] The term "selectively hybridizes" includes reference to hybridization of a nucleic acid sequence to a detectably greater degree (e.g., at least 2-fold over background) to a particular nucleic acid target sequence under stringent hybridization conditions relative to hybridization to non-target nucleic acid sequences and to the substantial exclusion of non-target nucleic acids. Selectively hybridizing sequences typically have at least about 80% or 90% sequence identity with each other, and up to 100% sequence identity (i.e., fully complementary).

[0402] The terms "stringent conditions" or "stringent hybridization conditions" include reference to conditions under which a probe will selectively hybridize to its target sequence in an in vitro hybridization assay. Stringent conditions are sequence-dependent and will vary under different circumstances. By controlling the stringency of hybridization and / or washing conditions, target sequences that are 100% complementary to a probe can be identified (homologous probing). Alternatively, stringency conditions can be adjusted to tolerate sequence mismatches and allow for sequence mismatches such that lower similarities are detected (heterologous probing). Generally, probes are less than about 1000 nucleotides in length, and optionally less than 500 nucleotides in length. Typically, stringent conditions are those in which the salt concentration is about 1.5 M Na ion concentration, typically about 0.01 to 1.0 M Na ion concentration, at a pH of 7.0 to 8.3, at a temperature of at least about 30°C for short probes (e.g., 10 to 50 nucleotides) and at least about 60°C for long probes (e.g., more than 50 nucleotides). Stringent conditions can also be achieved by adding destabilizing substances such as formamide. Examples of low stringency conditions include hybridization at 37°C in a buffer consisting of 30 to 35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate), followed by washing in 1 to 2× SSC (20× SSC = 3.0 M NaCl / 0.3 M sodium citrate) at 50 to 55°C. Examples of moderate stringency conditions include hybridization in 40-45% formamide, 1 M NaCl, and 1% SDS at 37° C., followed by washing in 0.5-1×SSC at 55-60° C. Examples of high stringency conditions include hybridization in 50% formamide, 1 M NaCl, and 1% SDS at 37° C., followed by washing in 0.1×SSC at 60-65° C.

[0403] "Homology" refers to a similar DNA sequence. For example, a "region homologous to a genomic region" found on donor DNA refers to a region of DNA that has a similar sequence to a given "genomic region" of a cell or organism's genome. The homologous region can be of sufficient length to promote homologous recombination at the target site of cleavage. For example, the homologous region can be at least 5 to 10, 5 to 15, 5 to 20, 5 to 25, 5 to 30, 5 to 35, 5 to 40, 5 to 45, 5 to 50, 5 to 55, 5 to 60, 5 to 65, 5 to 70, 5 to 75, 5 to 80, 5 to 85, 5 to 90, 5 to 95, 5 to 100, 5 to 200, 5 to 300, 5 to 400, 5 to 500, 5 to 600, 5 to 7 ... The term "sufficient homology" refers to a length of 0, 5 to 800, 5 to 900, 5 to 1000, 5 to 1100, 5 to 1200, 5 to 1300, 5 to 1400, 5 to 1500, 5 to 1600, 5 to 1700, 5 to 1800, 5 to 1900, 5 to 2000, 5 to 2100, 5 to 2200, 5 to 2300, 5 to 2400, 5 to 2500, 5 to 2600, 5 to 2700, 5 to 2800, 5 to 2900, 5 to 3000, 5 to 3100, or more bases. "Sufficient homology" refers to two polynucleotide sequences having sufficient structural similarity to act as substrates for a homologous recombination reaction. This structural similarity includes the total length of each polynucleotide fragment as well as the sequence similarity of the polynucleotides. Sequence similarity can be described in terms of percent sequence identity over the entire length of the sequence and / or in terms of conserved regions and percent sequence identity over portions of the length of the sequence that include localized similarities, such as consecutive nucleotides with 100% sequence identity.

[0404] As used herein, a "genomic region" is a segment of a chromosome in the genome of a cell that lies on either side of a target site or that also contains a portion of the target site. This genomic region may be at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5-50, 5-55, 5-60, 5-65, 5-70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5-200, 5-300, 5-400, 5-500, 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, 5-1200, 5-1300, 5- The genomic region may comprise 1400, 5 to 1500, 5 to 1600, 5 to 1700, 5 to 1800, 5 to 1900, 5 to 2000, 5 to 2100, 5 to 2200, 5 to 2300, 5 to 2400, 5 to 2500, 5 to 2600, 5 to 2700, 5 to 2800, 5 to 2900, 5 to 3000, 5 to 3100 or more bases, such that the genomic region has sufficient homology to undergo homologous recombination with a corresponding homologous region.

[0405] As used herein, "homologous recombination" (HR) involves the exchange of DNA fragments between two DNA molecules at homologous sites. The frequency of homologous recombination is influenced by many factors. The amount of homologous recombination and the relative proportions of homologous and non-homologous recombination vary between different organisms. Generally, the length of the homologous region affects the frequency of homologous recombination events; the longer the region of homology, the higher the frequency. The length of the homologous region required to observe homologous recombination also varies between species. In many instances, at least 5 kb of homology is used, although homologous recombination has been observed with 25-50 bp of homology. For example, Singer et al., (1982) Cell 31:25-33; Shen and Huang, (1986) Genetics 112:441-57; Watt et al., (1985) Proc. Natl. Acad. Sci. USA 82:4768-72, Sugawara and Haber, (1992) Mol Cell Biol. 12:563-75; Rubnitz and Subramani, (1984) Mol Cell Biol 4:2253-8; Ayares et al., 1986) Proc. Natl. Acad. Sci. USA 83:5199-203; Liskay et al., (1987) Genetics 115:161-7.

[0406] "Sequence identity" or "identity" in the context of nucleic acid or polypeptide sequences means the nucleobases or amino acid residues in two sequences that are identical when aligned for maximum correspondence over a specified comparison window.

[0407] The term "percentage of sequence identity" refers to a value determined by comparing two optimally aligned sequences over a comparison window, where the portion of the polynucleotide or polypeptide sequence in the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions) in order to optimally align the two sequences. The percentage is calculated by determining the number of positions in both sequences where the same nucleic acid base or amino acid residue occurs to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Useful examples of percent sequence identity include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%, or any percentage between 50% and 100%. These identities can be determined using any of the programs described herein.

[0408] Sequence alignments and percent identity or similarity calculations can be determined using a variety of comparison methods designed to detect homologous sequences, including, but not limited to, the MegAlign™ program in the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, WI). Within the context of this application, when sequence analysis software is used for the analysis, it will be understood that the results of the analysis will be based on the "default values" of the referenced program unless otherwise specified. As used herein, "default values" refers to any set of values ​​or parameters that are initially loaded with the software upon first initialization.

[0409] The "Clustal V method of alignment" is designated Clustal V (as described by Higgins and Sharp, (1989) CABIOS 5:151-153; Higgins et al., (1992) Comput Appl Biosci 8:189-191) and corresponds to the alignment method found in the MegAlign™ program of the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, WI). For multiple alignments, the default values ​​correspond to GAP PENALTY=10 and GAP LENGTH PENALTY=10. The default parameters for pairwise alignment and percent identity calculation of protein sequences using the Clustal method are KTUPLE=1, GAP PENALTY=3, WINDOW=5, and DIAGONALS SAVED=5. For nucleic acids, these parameters are KTUPLE=2, GAP PENALTY=5, WINDOW=4, and DIAGONALS SAVED=4. After alignment of sequences using the Clustal V program, a "percent identity" can be obtained by consulting the "sequence distance" table in the program. The "Clustal W method of alignment" is designated Clustal W (described by Higgins and Sharp, (1989) CABIOS 5:151-153; Higgins et al., (1992) Comput Appl Biosci 8:189-191) and corresponds to the alignment method found in the MegAlign™ v6.1 program of the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, WI). Default parameters for multiple alignment (GAP PENALTY=10, GAP LENGTH PENALTY=0.2, Delay Divergen Sequences (%)=30, DNA Transition Weight=0.5, Protein Weight Matrix=Gonnet Series, DNA Weight Matrix=IUB).After aligning sequences using the Clustal W program, "percent identity" can be obtained by consulting the "sequence distance" table in the program. Unless otherwise specified, the sequence identity / similarity values ​​presented herein refer to values ​​obtained using GAP Version 10 (GCG, Accelrys, San Diego, CA) with the following parameters: for nucleotide sequence identity and similarity, a gap creation penalty weight of 50, a gap extension penalty weight of 3, and the nwsgapdna.cmp scoring matrix are used; for amino acid sequence identity and similarity, a gap creation penalty weight of 8, a gap extension penalty of 2, and the BLOSUM62 scoring matrix are used (Henikoff and Henikoff, (1989) Proc. Natl. Acad. Sci. USA 89:10915). GAP uses the algorithm of Needleman and Wunsch (1970) J Mol Biol 48:443-53 to find a global alignment of two sequences that maximizes the number of matches and minimizes the number of gaps. GAP considers all possible alignments and gap positions and creates an alignment with the maximum number of matched bases and the fewest gaps, using gap creation and extension penalties in units of matched bases. "BLAST" is a search algorithm provided by the National Center for Biotechnology Information (NCBI) used to find regions of similarity between biological sequences. This program compares nucleotide or protein sequences with sequence databases, calculates the statistical significance of matches, and identifies sequences sufficiently similar to a query sequence that the similarity is not expected to have occurred randomly. BLAST reports the identified sequences and their local alignment with the query sequence. Those skilled in the art will appreciate that many levels of sequence identity are useful for identifying polypeptides from other species or that have been modified naturally or synthetically, such that they have the same or similar function or activity.Useful examples of percent identity include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% or any percentage between 50% and 100%. Indeed, any amino acid identity between 50% and 100% may be useful in describing the present disclosure, such as 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%.

[0410] Polynucleotide and polypeptide sequences, their variants, and the structural relationships of these sequences can be described by the terms "homology," "homologous," "substantially identical," "substantially similar," and "substantially corresponding," which are used interchangeably herein. These refer to polypeptide or nucleic acid sequences in which changes in one or more amino acids or nucleotide bases do not affect the function of the molecule, e.g., its ability to mediate gene expression or its ability to produce a particular phenotype. These terms also refer to modifications of nucleic acid sequences in which the functional properties of the resulting nucleic acid are not substantially altered compared to the intended, unmodified nucleic acid. These modifications include deletions, substitutions, and / or insertions of one or more nucleotides in the nucleic acid fragment. Substantially similar nucleic acid sequences can be defined by their ability to hybridize (under moderate stringency conditions, e.g., 0.5×SSC, 0.1% SDS, 60°C) to the sequences exemplified herein or to any portion of the nucleotide sequences disclosed herein, and are functionally equivalent to any of the nucleic acid sequences disclosed herein. Stringency conditions can be adjusted to screen for fragments with moderate similarity, such as homologous sequences from distantly related organisms, against fragments with high similarity, such as genes duplicating functional enzymes from closely related organisms. Post-hybridization washes determine stringency conditions.

[0411] A "centimorgan" (cM) or "map unit" is the distance between two polynucleotide sequences, linked genes, markers, target sites, loci, or any pair thereof, where 1% of meiotic outcomes are recombinations. Thus, a centimorgan is equivalent to the distance that equals the average recombination frequency of 1% between two linked genes, markers, target sites, loci, or any pair thereof.

[0412] An "isolated" or "purified" nucleic acid molecule, polynucleotide, polypeptide, or protein, or biologically active portion thereof, is substantially or essentially free from components that normally co-occur with or interact with the polynucleotide or protein as found in its naturally occurring environment. Thus, an isolated or purified polynucleotide or protein is substantially free of other cellular material and culture medium when produced by recombinant techniques, and is substantially free of chemical precursors and other chemicals when chemically synthesized. Optimally, an "isolated" polynucleotide (optimally, a sequence encoding a protein) is free of sequences that naturally flank the polynucleotide in the genomic DNA of the organism from which the polynucleotide is derived (i.e., sequences located at the 5' and 3' ends of the polynucleotide). For example, in various embodiments, an isolated polynucleotide may contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleotide sequences that naturally flank the polynucleotide in the genomic DNA of the cell from which the polynucleotide is derived. Isolated polynucleotides can be purified from cells in which they naturally occur. Conventional nucleic acid purification methods known to those skilled in the art can be used to obtain isolated polynucleotides. The term also encompasses recombinant and chemically synthesized polynucleotides.

[0413] The term "fragment" refers to a contiguous set of nucleotides or amino acids. In one embodiment, a fragment is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 contiguous nucleotides. In one embodiment, a fragment is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 contiguous amino acids. A fragment may or may not exhibit function of a sequence that shares some percent identity over the entire length of the fragment.

[0414] The terms "functionally equivalent fragment" and "functionally equivalent fragment" are used interchangeably herein. These terms refer to a portion or subsequence of an isolated nucleic acid fragment or polypeptide that exhibits the same activity or function as the longer sequence from which it is derived. In one example, a fragment retains the ability to alter gene expression or produce a particular phenotype, regardless of whether the fragment encodes an active protein. For example, fragments can be used to engineer genes that produce a desired phenotype in an altered plant. A gene can be engineered to be used in an inhibitory manner by linking the nucleic acid fragment, whether it encodes an active enzyme or not, to a plant promoter sequence in a sense or antisense orientation.

[0415] A "gene" includes, but is not limited to, a nucleic acid fragment that expresses a functional molecule, such as a specific protein, including regulatory sequences preceding (5' non-coding sequences) and following (3' non-coding sequences) the coding sequence. A "native gene" refers to a gene found in its natural, endogenous location with its own regulatory sequences.

[0416] The term "endogenous" refers to a sequence or other molecule that is naturally present in a cell or organism. In one aspect, an endogenous polynucleotide is one that is normally found in the genome of a cell; i.e., it is not heterologous.

[0417] An "allele" is one of several alternative forms of a gene that occupies a given locus on a chromosome. If all alleles present at a given locus on a chromosome are the same, the plant is homozygous at that locus. If alleles present at a given locus on a chromosome are different, the plant is heterozygous at that locus.

[0418] "Coding sequence" refers to a polynucleotide sequence that encodes a specific amino acid sequence. "Regulatory sequence" refers to a nucleotide sequence located upstream of a coding sequence (5' non-coding sequence), within a coding sequence, or downstream of a coding sequence (3' non-coding sequence) that influences the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include, but are not limited to, promoters, translation leader sequences, 5' untranslated sequences, 3' untranslated sequences, introns, polyadenylation target sequences, RNA processing sites, effector binding sites, and stem-loop structures.

[0419] A "mutated gene" is a gene that has been altered by human intervention. Such a "mutated gene" has a sequence that differs from that of the corresponding non-mutated gene by the addition, deletion, or substitution of at least one nucleotide. In certain embodiments of the present disclosure, the mutant gene comprises a modification resulting from the guide polynucleotide / Cas endonuclease system disclosed herein. A mutant plant is a plant that comprises a mutant gene.

[0420] As used herein, a "targeted mutation" is a mutation in a gene (referred to as a target gene), including a native gene, that is made by modifying a target sequence within the target gene using any method known to those of skill in the art, including methods involving the inducible Cas endonuclease system disclosed herein.

[0421] The terms "knockout," "gene knockout," and "genetic knockout" are used interchangeably herein. Knockout refers to a cellular DNA sequence that has been partially or completely disabled by targeting with a Cas protein; for example, the DNA sequence before the knockout may have encoded an amino acid sequence or had a regulatory function (e.g., a promoter).

[0422] The terms "knock-in," "gene knock-in," "gene insertion," and "genetic knock-in" are used interchangeably herein. Knock-in refers to the replacement or insertion of a DNA sequence at a specific DNA sequence in a cell by targeting using a Cas protein (e.g., by homologous recombination (HR), also using a suitable donor DNA polynucleotide). Examples of knock-ins are the specific insertion of a heterologous amino acid coding sequence in the coding region of a gene, or the specific insertion of a transcriptional regulatory element at a gene locus.

[0423] "Domain" means a contiguous stretch of nucleotides (which may be RNA, DNA and / or combined RNA-DNA sequences) or a contiguous stretch of amino acids.

[0424] The term "conserved domain" or "motif" refers to a set of amino acids that are conserved at a specific position in the aligned sequences of a set of polynucleotides or evolutionarily related proteins. While amino acids at other positions may vary among homologous proteins, highly conserved amino acids at a specific position represent amino acids that are essential for the structure, stability, or activity of the protein. Because they are identified by the high degree of conservation in the aligned sequences of that protein homolog family, they can be used as identifiers or "signatures" to determine whether a protein having a newly determined sequence belongs to a previously identified protein family.

[0425] A "codon-modified gene" or "codon-preferred gene" or "codon-optimized gene" is a gene whose codon usage is designed to mimic the preferred codon usage of a host cell.

[0426] An "optimized" polynucleotide is a sequence that has been optimized for improved expression in a particular heterologous host cell.

[0427] A "plant-optimized nucleotide sequence" is a nucleotide sequence that has been optimized for improved expression in plants, particularly for increased expression in plants. Plant-optimized nucleotide sequences include codon-optimized genes. Plant-optimized nucleotide sequences can be synthesized by modifying a nucleotide sequence encoding a protein, such as a Cas endonuclease disclosed herein, using one or more plant-preferred codons for improved expression. See, for example, Campbell and Gowri (1990) Plant Physiol. 92:1-11 for a discussion of host-preferred codon usage.

[0428] A promoter is a region of DNA involved in the recognition and binding of RNA polymerase and other proteins that initiate transcription. A promoter sequence consists of proximal and more distal upstream elements, the latter often referred to as enhancers. An "enhancer" is a DNA sequence capable of stimulating promoter activity and may be a native element of the promoter or a heterologous element inserted to enhance the level or tissue specificity of the promoter. A promoter may be derived entirely from a native gene, may be composed of different elements derived from different promoters found in nature, and / or may include synthetic DNA segments. It is understood by those skilled in the art that different promoters can induce the expression of a gene in different tissue or cell types, at different developmental stages, or in response to different environmental conditions. It is further recognized that because the exact boundaries of regulatory sequences in most cases have not been completely defined, DNA fragments of some variation may have identical promoter activity.

[0429] Promoters that most often cause expression of a gene in most cell types are commonly referred to as "constitutive promoters." The term "inducible promoter" refers to a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, for example, by a chemical compound (chemical inducer), or in response to environmental, hormonal, chemical, and / or developmental signals. Inducible or regulatable promoters include, for example, promoters that are induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormones, wounding, or chemicals such as ethanol, abscisic acid (ABA), jasmonate, salicylic acid, or safeners.

[0430] "Translation leader sequence" refers to a polynucleotide sequence located between the promoter sequence and the coding sequence of a gene. Translation leader sequences are located upstream of the translation initiation sequence of an mRNA. Translation leader sequences can affect the processing of the primary transcript into mRNA, mRNA stability, or translation efficiency. Examples of translation leader sequences have been reported (e.g., Turner and Foster, (1995) Mol Biotechnol 3:225-236).

[0431] "3' non-coding sequence," "transcription terminator," or "termination sequence" refers to DNA sequences located downstream of the coding sequence and includes polyadenylation recognition sequences and other sequences encoding regulatory signals that can affect mRNA processing or gene sequences. Polyadenylation signals are usually characterized by affecting the addition of polyadenylic acid to the 3' end of a pre-mRNA. The use of different 3' non-coding sequences is exemplified in Ingelbrecht et al., (1989) Plant Cell 1:671-680.

[0432] "RNA transcript" refers to the product resulting from transcription of a DNA sequence catalyzed by RNA polymerase. When an RNA transcript is a perfectly complementary copy of a DNA sequence, it is called a primary transcript or pre-mRNA. When an RNA transcript is an RNA sequence resulting from post-transcriptional processing of a primary transcript pre-RNA, it is called a mature RNA or mRNA. "Messenger RNA" or "mRNA" refers to RNA that is free of introns and can be translated into protein by a cell. "cDNA" refers to DNA that is complementary to an mRNA template and synthesized from the mRNA template using reverse transcriptase. cDNA can be single-stranded or converted to double-stranded form using the Klenow fragment of DNA polymerase I. "Sense" RNA refers to an RNA transcript that includes mRNA and can be translated into protein in cells or in vitro. "Antisense RNA" refers to an RNA transcript that is complementary to all or part of a target primary transcript or mRNA and blocks expression of a target gene (see, e.g., U.S. Pat. No. 5,107,065). The complementarity of an antisense RNA may be with any part of a specific gene transcript, i.e., at the 5' non-coding sequence, 3' non-coding sequence, introns, or the coding sequence. "Functional RNA" refers to antisense RNA, ribozyme RNA, or other RNA that is not translated yet affects cellular processes. The terms "complement" and "reverse complement" are used interchangeably herein with respect to an mRNA transcript and are intended to define an antisense RNA of the message.

[0433] The term "genome" refers to the entire body of genetic material (genes and non-coding sequences) present in each cell of an organism, or virus or organelle; and / or the set of chromosomes inherited as a (haploid) unit from one parent.

[0434] The term "operably linked" refers to the association of nucleic acid sequences on a single nucleic acid fragment so that the function of one is regulated by the other. For example, a promoter is operably linked to a coding sequence if it is capable of controlling the expression of the coding sequence (i.e., the coding sequence is under the transcriptional control of the promoter). A coding sequence can be operably linked to a regulatory sequence in either sense or antisense orientation. In another example, a complementary RNA region can be operably associated, directly or indirectly, with or within the 5' end of a target mRNA, or the 3' end of a target mRNA, or a first complementary region is at the 5' end of the target mRNA and its complement is at the 3' end of the target mRNA.

[0435] Generally, "host" refers to an organism or cell into which a heterologous component (polynucleotide, polypeptide, other molecule, cell) has been introduced. As used herein, "host cell" refers to a eukaryotic cell, a prokaryotic cell (e.g., a bacterial or archaeal cell), or a cell from a multicellular organism (e.g., a cell line), cultured as a unicellular entity, in vivo or in vitro, into which a heterologous polynucleotide or polypeptide has been introduced. In some embodiments, the cell is selected from the group consisting of: an archaeal cell, a bacterial cell, a eukaryotic cell, a unicellular eukaryotic organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algae cell, an animal cell, an invertebrate cell, a vertebrate cell, a fish cell, a frog cell, an avian cell, an insect cell, a mammalian cell, a porcine cell, a bovine cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell. In some examples, the cell is in vitro. In some examples, the cell is in vivo.

[0436] The term "recombinant" refers to the artificial combination of two otherwise separate segments of sequence, for example, by chemical synthesis or by the manipulation of isolated segments of nucleic acid by genetic engineering techniques.

[0437] The terms "plasmid," "vector," and "cassette" refer to linear or circular supplemental chromosomal elements, usually in the form of double-stranded DNA, that often carry genes that are not part of the cell's central metabolism. Such elements can be autonomously replicating sequences, genome-integrating sequences, phages, or nucleotide sequences of any origin, of single- or double-stranded DNA or RNA, in linear or circular form, in which many nucleotide sequences are linked or recombined into unique constructs capable of introducing a polynucleotide of interest into a cell. A "transformation cassette" refers to a specific vector containing a gene and having elements in addition to the gene that facilitate transformation of a specific host cell. An "expression cassette" refers to a specific vector containing a gene and having elements in addition to the gene that cause the host to express the gene.

[0438] The terms "recombinant DNA molecule," "recombinant DNA construct," "expression construct," "construct," and "recombinant construct" are used interchangeably herein. Recombinant DNA constructs include artificial combinations of nucleic acid fragments, e.g., regulatory and coding sequences, that are not necessarily found together in nature. For example, a recombinant DNA construct may contain regulatory and coding sequences that are derived from different sources, or regulatory and coding sequences that are derived from the same source but arranged in a manner different from how they are found in nature. Such constructs may be used alone or in conjunction with a vector. When a vector is used, the choice of vector will depend on the method used to introduce the vector into a host cell, as is well known to those skilled in the art. For example, a plasmid vector may be used. The skilled artisan is familiar with the genetic elements that must be present on the vector in order to successfully transform, select, and propagate a host cell. The skilled artisan will also recognize that different, independent transformation events can result in different expression levels and patterns (Jones et al., (1985) EMBO J 4:2411-2418; De Almeida et al., (1989) Mol Gen Genetics 218:78-86), and therefore, multiple events are typically screened to obtain strains exhibiting the desired expression levels and patterns. Such screening can be performed by standard molecular biological, biochemical, and other analytical methods, including Southern analysis of DNA, Northern analysis of mRNA expression, PCR, real-time quantitative PCR (qPCR), reverse transcription PCR (RT-PCR), immunoblotting of protein expression, enzyme or activity assays, and / or phenotypic analysis.

[0439] The term "heterologous" refers to a difference between the original environment, location, or composition of a particular polynucleotide or polypeptide sequence and its current environment, location, or composition. Non-limiting examples include differences in taxonomic origin (e.g., a polynucleotide sequence obtained from Zea mays is heterologous if inserted into the genome of a rice (Oryza sativa) plant or a different subspecies or variety of Zea mays; or a polynucleotide obtained from a bacterium is introduced into a plant cell) or sequence differences (e.g., a polynucleotide sequence obtained from Zea mays, isolated, modified, and reintroduced into a maize plant). As used herein, "heterologous" in reference to a sequence can refer to a sequence originating from a different species, subspecies, or alien species, or, if from the same species, a sequence whose composition and / or genomic locus has been substantially altered from its native form by deliberate human intervention. For example, a promoter operably linked to a heterologous polynucleotide may be from a species different from that from which the polynucleotide was derived, or, if from the same / similar species, one or both may be substantially altered from their original form and / or genomic locus, or the promoter may not be the native promoter of the operably linked polynucleotide. Alternatively, one or more of the regulatory regions and / or polynucleotides set forth herein may be entirely synthetic. In another example, a target polynucleotide for cleavage by a Cas endonuclease may be from a different organism than the Cas endonuclease. In another example, the Cas endonuclease and guide RNA may be introduced into a target polynucleotide along with an additional polynucleotide that acts as a template or donor for insertion into the target polynucleotide, where the additional polynucleotide is heterologous to the target polynucleotide and / or Cas endonuclease.

[0440] The term "expression" as used herein refers to the production of a functional end-product (e.g., mRNA, guide RNA, or protein) in precursor or mature form.

[0441] A "mature" protein refers to a post-translationally processed polypeptide (i.e., a polypeptide from which any pre- or propeptides present in the primary translation product have been removed).

[0442] "Precursor" protein refers to the primary product of translation of mRNA (i.e., with pre- and propeptides still present). The pre- and propeptides may be, but are not limited to, subcellular localization signals.

[0443] "CRISPR" (clustered regularly interspaced short palindromic repeats) loci refer to specific loci that encode components of DNA cleavage systems used, for example, by bacterial and archaeal cells to destroy foreign DNA (Horvath and Barrangou, 2010, Science 327:167-170; WO 2007 / 025097, published March 1, 2007). CRISPR loci can be composed of CRISPR arrays containing short direct repeats (CRISPR repeats) separated by short, variable DNA sequences (called spacers), which can be flanked by a variety of Cas (CRISPR-associated) genes.

[0444] As used herein, an "effector" or "effector protein" is a protein whose activities include recognizing, binding, and / or cleaving or nicking a polynucleotide target. An effector or effector protein may also be an endonuclease. The "effector complex" of a CRISPR system includes Cas proteins that are involved in recognizing and binding the crRNA to the target. Some of the component Cas proteins may further include domains involved in target polynucleotide cleavage.

[0445] The term "Cas protein" refers to proteins encoded by Cas (CRISPR-associated) genes. Cas proteins include proteins encoded by genes in the cas locus and include adaptation molecules and interference molecules. Interference molecules of bacterial adaptive immune complexes include endonucleases. Cas endonucleases described herein contain one or more nuclease domains. Cas endonucleases include, but are not limited to, the novel Cas-alpha proteins disclosed herein, Cas9 proteins, Cpf1 (Cas12) proteins, C2c1 proteins, C2c2 proteins, C2c3 proteins, Cas3, Cas3-HD, Cas5, Cas7, Cas8, Cas10, or combinations or complexes thereof. Cas proteins can be "Cas endonucleases" or "Cas effector proteins" that, when complexed with a suitable polynucleotide component, can recognize, bind to, and optionally nick or cleave all or part of a specific polynucleotide target sequence. Cas-alpha endonucleases of the present disclosure include those having one or more RuvC nuclease domains.The Cas protein may be a functional fragment or functional variant of a naturally occurring Cas protein, or a fragment of at least 50, 50-100, at least 100, 100-150, at least 150, 150-200, at least 200, 200-250, at least 250, 250-300, at least 300, 300-350, at least 350, 350-400, at least 400, 400-450, at least 500, or more than 500 consecutive amino acids of a naturally occurring Cas protein, and at least 50%, 50%-55%, at least 55%, 55%-60%, or at least 60% , 60% to 65%, at least 65%, 65% to 70%, at least 70%, 70% to 75%, at least 75%, 75% to 80%, at least 80%, 80% to 85%, at least 85%, 85% to 90%, at least 90%, at least 90% to 95%, at least 95%, 95% to 96%, at least 96%, 96% to 97%, at least 97%, 97% to 98%, at least 98%, 98% to 99%, at least 99%, 99% to 100%, or 100% sequence identity and retains at least some activity of the native sequence.

[0446] The terms "functional fragment," "functionally equivalent fragment," and "functionally equivalent fragment" of a Cas endonuclease are used interchangeably herein to refer to a portion or subsequence of a Cas endonuclease sequence of the present disclosure that retains the ability to recognize, bind to, and optionally unwind, nick, or cleave a target site (introducing a single- or double-stranded break within the target site). This portion or subsequence of a Cas endonuclease can include a complete or partial (functional) peptide of any one of its domains, such as, but not limited to, the entire functional portion of the Cas3 HD domain, the entire functional portion of the Cas3 helicase domain, or the entire functional portion of a protein (such as, but not limited to, Cas5, Cas5d, Cas7, and Cas8b1).

[0447] The terms "functional variant," "functionally equivalent variant," and "functionally equivalent variant" of a Cas endonuclease or Cas effector protein, including Cas-alpha, described herein are used interchangeably herein and refer to a variant of the Cas effector protein disclosed herein that retains the ability to recognize, bind to, and optionally unbind, nick, or cleave all or part of a target sequence.

[0448] Cas endonucleases may also include multifunctional Cas endonucleases. The terms "multifunctional Cas endonuclease" and "multifunctional Cas endonuclease polypeptide" are used interchangeably herein and include reference to a single polypeptide having a Cas endonuclease function (comprising at least one protein domain capable of acting as a Cas endonuclease) and at least one other function, such as, but not limited to, a complexing function (comprising at least a second protein domain capable of forming a complex with another protein). In one embodiment, a multifunctional Cas endonuclease comprises at least one additional protein domain (internal, upstream (5'), downstream (3'), or both internal 5' and 3', or any combination thereof) relative to the domains typical of a Cas endonuclease.

[0449] The terms "cascade" and "cascade complex" are used interchangeably herein and include reference to a multi-subunit protein complex that can assemble with a polynucleotide to form a polynucleotide-protein complex (PNP). A cascade is a PNP that relies on a polynucleotide for complex assembly and stability and for target nucleic acid sequence identification. A cascade functions as a surveillance complex that finds and optionally binds to target nucleic acids complementary to the variable targeting domain of a guide polynucleotide.

[0450] The terms "cleavage-ready cascade," "cr cascade," "cleavage-ready cascade complex," "cr cascade complex," "cleavage-ready cascade system," "CRC," and "cr cascade system" are used interchangeably herein and include reference to a multi-subunit protein complex that can assemble with a polynucleotide to form a polynucleotide-protein complex (PNP), where one of the Cascade proteins is a Cas endonuclease that can recognize, bind to, and optionally unwind, nick, or cleave all or part of a target sequence.

[0451] The terms "5'-cap" and "7-methylguanylate (m7G) cap" are used interchangeably herein. 7-methylguanylate residues are located at the 5' end of messenger RNA (mRNA) in eukaryotic cells. RNA polymerase II (Pol II) transcribes mRNA in eukaryotic cells. Messenger RNA capping generally occurs as follows: the extreme 5' phosphate group of an mRNA transcript is removed by an RNA terminal phosphatase, leaving two terminal phosphate groups. Guanosine monophosphate (GMP) is added to the terminal phosphate group of the transcript by a guanylyltransferase, leaving a 5'-5' triphosphate-linked guanine at the end of the transcript. Finally, the 7-nitrogen of this terminal guanine is methylated by a methyltransferase.

[0452] The term "not having a 5'-cap" is used herein to refer to RNA that has, for example, a 5'-hydroxyl group instead of a 5'-cap. Such RNA can be referred to, for example, as "uncapped RNA." Since 5'-capped RNA undergoes nuclear export, uncapped RNA can be more fully accumulated in the nucleus after transcription. One or more RNA components herein are uncapped.

[0453] As used herein, the term "guide polynucleotide" refers to a polynucleotide sequence that can form a complex with a Cas endonuclease, such as a Cas endonuclease described herein, and enables the Cas endonuclease to recognize, optionally bind to, and optionally cleave a DNA target site. A guide polynucleotide sequence can be an RNA sequence, a DNA sequence, or a combination thereof (a combined RNA-DNA sequence).

[0454] The terms "functional fragment," "functionally equivalent fragment," and "functionally equivalent fragment" of a guide RNA, crRNA, or tracrRNA are used interchangeably herein and refer to a portion or subsequence of a guide RNA, crRNA, or tracrRNA of the present disclosure that retains the ability to function as a guide RNA, crRNA, or tracrRNA, respectively.

[0455] "Functional variant," "functionally equivalent variant," and "functionally equivalent variant" of a guide RNA, crRNA, or tracrRNA (respectively) are used interchangeably herein and refer to a variant of a guide RNA, crRNA, or tracrRNA of the present disclosure that retains the ability to function as a guide RNA, crRNA, or tracrRNA, respectively.

[0456] The terms "single guide RNA" and "sgRNA" are used interchangeably herein and refer to a synthetic fusion of two RNA molecules: a crRNA (CRISPR RNA) containing a variable targeting domain (linked to a tracr mate sequence that hybridizes to the tracrRNA) fused to a tracrRNA (trans-activating CRISPR RNA). The single guide RNA may comprise a crRNA or crRNA fragment of a Type II CRISPR / Cas system and a tracrRNA or tracrRNA fragment that can form a complex with a Type II Cas endonuclease, and the guide RNA / Cas endonuclease complex can guide the Cas endonuclease to a DNA target site, allowing the Cas endonuclease to recognize, optionally bind to, and optionally nick or cleave (introduce a single- or double-strand break) the DNA target site.

[0457] The terms "variable targeting domain" or "VT domain" are used interchangeably herein and comprise a nucleotide sequence that can hybridize to (be complementary to) one strand (nucleotide sequence) of a double-stranded DNA target site. The percent complementarity between the first nucleotide sequence domain (VT domain) and the target sequence can be at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 63%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. The length of a variable targeting domain can be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In some embodiments, the variable targeting domain comprises a contiguous stretch of 12-30 nucleotides. The variable targeting domain can be comprised of a DNA sequence, an RNA sequence, a modified DNA sequence, a modified RNA sequence, or any combination thereof.

[0458] The terms "Cas endonuclease recognition domain" or "CER domain" (of a guide polynucleotide) are used interchangeably herein and include a nucleotide sequence that interacts with a Cas endonuclease polypeptide. A CER domain includes a (trans-acting) tracr nucleotide mate sequence followed by a tracr nucleotide sequence. A CER domain can be composed of a DNA sequence, an RNA sequence, a modified DNA sequence, a modified RNA sequence (see, e.g., U.S. Patent Application Publication No. 2015 / 0059010A1, published February 26, 2015), or any combination thereof.

[0459] As used herein, the terms "guide polynucleotide / Cas endonuclease complex," "guide polynucleotide / Cas endonuclease system," "guide polynucleotide / Cas complex," "guide polynucleotide / Cas system," "inducible Cas system," "polynucleotide-guided endonuclease," and "PGEN" are used interchangeably herein and refer to at least one guide polynucleotide and at least one Cas endonuclease that are capable of forming a complex, wherein the guide polynucleotide / Cas endonuclease complex is capable of directing the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce single- or double-strand breaks) the DNA target site. The guide polynucleotide / Cas endonuclease complex herein may comprise a Cas protein and a suitable polynucleotide component from any known CRISPR system (Horvath and Barrangou, 2010, Science 327:167-170; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15; Zetsche et al., 2015, Cell 163, 1-13; Shmakov et al., 2015, Molecular Cell 60, 1-13).

[0460] The terms "guide RNA / Cas endonuclease complex," "guide RNA / Cas endonuclease system," "guide RNA / Cas complex," "guide RNA / Cas system," "gRNA / Cas complex," "gRNA / Cas system," "RNA-guided endonuclease," and "RGEN" are used interchangeably herein and refer to at least one RNA component and at least one Cas endonuclease capable of forming a complex, wherein the guide RNA / Cas endonuclease complex can guide the Cas endonuclease to a DNA target site, allowing the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single- or double-strand break) the DNA target site.

[0461] The terms "target site," "target sequence," "target site sequence," "target DNA," "target locus," "genomic target site," "genomic target sequence," "genomic target locus," and "protospacer" are used interchangeably herein and refer to a polynucleotide sequence, such as, but not limited to, a nucleotide sequence on a chromosome, episome, locus, or any other DNA molecule in the genome of a cell (e.g., chromosomal DNA, chloroplast DNA, mitochondrial DNA, plasmid DNA), that a guide polynucleotide / Cas endonuclease complex can recognize, bind to, and optionally nick or cleave. A target site can be an endogenous site in the genome of a cell, or the target site can be heterologous to the cell, such that it does not naturally occur in the genome of the cell, or the target site can be found in a genomic location heterologous to where it occurs in nature. As used herein, the terms "endogenous target sequence" and "native target sequence" are used interchangeably herein and refer to a target sequence that is endogenous to or occurs in the genome of a cell and is present in the endogenous or occurring location of that target sequence in the genome of the cell. "Artificial target site" or "artificial target sequence" are used interchangeably herein and refer to a target sequence that has been introduced into the genome of a cell. Such an artificial target sequence may be identical in sequence to an endogenous or native target sequence in the genome of the cell, but may be located at a different location in the genome of the cell (i.e., a non-endogenous or non-occurring location).

[0462] As used herein, the term "protospacer adjacent motif" (PAM) refers to a short nucleotide sequence adjacent to the target sequence (protospacer) recognized (targeted) by the guide polynucleotide / Cas endonuclease system described herein. A target DNA sequence cannot be correctly recognized unless it is followed by a PAM sequence. The sequence and length of a PAM herein can vary depending on the Cas protein or Cas protein complex used. The PAM sequence can be of any length but is generally 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length.

[0463] The terms "altered target site," "altered target sequence," "modified target site," and "modified target sequence" are used interchangeably herein and refer to a target sequence disclosed herein that contains at least one alteration compared to the unaltered target sequence. Such "alteration" may include, for example, (i) a substitution of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, (iv) a chemical modification of at least one nucleotide, or (v) any combination of (i)-(iv).

[0464] "Modified nucleotide" or "edited nucleotide" refers to a nucleotide sequence of interest that contains at least one alteration compared to its unmodified nucleotide sequence. Such alterations include, for example, (i) substitution of at least one nucleotide, (ii) deletion of at least one nucleotide, (iii) insertion of at least one nucleotide, (iv) chemical modification of at least one nucleotide, or (v) any combination of (i)-(iv).

[0465] Methods of "modifying a target site" and "altering a target site" are used interchangeably herein and refer to methods for generating an altered target site.

[0466] As used herein, "donor DNA" is a DNA construct that contains a polynucleotide of interest to be inserted into a target site of a Cas endonuclease.

[0467] The term "polynucleotide modification template" includes a polynucleotide containing at least one nucleotide modification compared to the nucleotide sequence to be edited. The nucleotide modification can be a substitution, addition, or deletion of at least one nucleotide. Optionally, the polynucleotide modification template can further include a homologous nucleotide sequence adjacent to the at least one nucleotide modification, wherein the adjacent homologous nucleotide sequence provides sufficient homology to the desired nucleotide sequence to be edited.

[0468] The term "plant-optimized Cas endonuclease" herein refers to a Cas protein, including a multifunctional Cas protein, that is encoded by a nucleotide sequence that is optimized for expression in a plant cell or plant.

[0469] The terms "plant-optimized nucleotide sequence encoding a Cas endonuclease," "plant-optimized construct encoding a Cas endonuclease," and "plant-optimized polynucleotide encoding a Cas endonuclease" are used interchangeably herein and refer to a nucleotide sequence encoding a Cas protein, or a variant or functional fragment thereof, that has been optimized for expression in a plant cell or plant. Plants comprising a plant-optimized Cas endonuclease include plants that comprise a nucleotide sequence encoding a Cas sequence and / or plants that comprise a Cas endonuclease protein. In one aspect, the plant-optimized Cas endonuclease nucleotide sequence is a maize-optimized, rice-optimized, wheat-optimized, soybean-optimized, cotton-optimized, or canola-optimized Cas endonuclease.

[0470] The term "plant" generally includes whole plants, plant organs, plant tissues, seeds, plant cells, seeds, and their progeny. Plants can be monocotyledons or dicotyledons. Plant cells include, but are not limited to, cells derived from seeds, suspension cultures, embryos, meristematic regions, callus tissue, leaves, roots, shoots, gametophytes, sporophytes, pollen, and microspores. "Plant element" is intended to refer to a whole plant or plant component, which may include, but is not limited to, differentiated and / or undifferentiated tissues, such as plant tissues, parts, and cell types. In one embodiment, the plant element is one of the following: whole plants, seedlings, meristems, ground tissues, vascular tissue, epidermal tissue, seeds, leaves, roots, shoots, stems, flowers, fruits, stolons, bulbs, tubers, corms, stems, shoots, shoots, tumor tissue, and various forms of cells and cultures (e.g., single cells, protoplasts, embryos, callus tissue). It should be noted that because protoplasts lack a cell wall, they are not technically "intact" plant cells (with all their components) found in nature. The term "plant organ" refers to a plant tissue or a group of tissues that constitute a morphologically and functionally independent part of a plant. As used herein, "plant element" is synonymous with plant "part" and refers to any part of a plant, may include separate tissues and / or organs, and may be used interchangeably with the term "tissue" throughout. Similarly, "plant reproductive element" is generally intended to refer to any part of a plant that is capable of starting other plants by sexual or asexual reproduction of that plant, such as, but not limited to, a seed, seedling, root, shoot, cutting, scion, graft, stolon, bulb, tuber, corm, keiki, or sprout. A plant element may be in a plant, or a plant organ, tissue culture, or cell culture.

[0471] "Progeny" includes subsequent generations of the plant.

[0472] As used herein, the term "plant part" refers to plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant callus, plant mass, and intact plant cells in plants or plant parts, such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, grains, ears, cobs, pods, stalks, roots, root tips, anthers, and parts thereof, as well as parts thereof. Grain is intended to mean mature seeds produced by growers for purposes other than the cultivation or propagation of the species. Progeny, variants, and mutants of regenerated plants are also within the scope of the present invention, provided that the parts contain the introduced polynucleotide.

[0473] The term "monocotyledonous" or "monocot" refers to a subclass of angiosperms, also known as the "monocotyledoneae," whose seeds generally contain only one primary leaf or cotyledon. The term includes reference to whole plants, plant elements, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, and their progeny.

[0474] The term "dicotyledonous" or "dicot" refers to a subclass of angiosperms, also known as "dicotyledoneae," whose seeds generally contain only two primary leaves or cotyledons. The term includes reference to whole plants, plant elements, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, and their progeny.

[0475] As used herein, a "male sterile plant" is a plant that does not produce viable or otherwise fertile male gametes. As used herein, a "female sterile plant" is a plant that does not produce viable or otherwise fertile female gametes. It is recognized that male sterile plants and female sterile plants can be female fertile and male fertile, respectively. It is further recognized that male fertile (other than female fertile) plants can produce viable offspring when crossed with female fertile plants, and female fertile (other than male fertile) plants can produce viable offspring when crossed with male fertile plants.

[0476] As used herein, the term "non-conventional yeast" refers to any yeast that is not a Saccharomyces yeast species (e.g., S. cerevisiae) or a Schizosaccharomyces yeast species (see "Non-Conventional Yeasts in Genetics, Biochemistry and Biotechnology: Practical Protocols," K. Wolf, KD Breunig, G. Barth, Eds., Springer-Verlag, Berlin, Germany, 2003).

[0477] In the context of this disclosure, the terms "crossed" or "cross" or "crossing" refer to the fusion of gametes by pollination to produce offspring (i.e., cells, seeds, or plants). The term encompasses both sexual crossing (pollination of one plant by another) and selfing (self-pollination, i.e., where the pollen and ovules (microspores and megaspores) are from the same plant or from genetically identical plants).

[0478] The term "introgression" refers to the transfer of a desired allele at a genetic locus from one genetic background to another. For example, introgression of a desired allele at a particular genetic locus can be transferred to at least one progeny plant by sexual crossing between two parent plants, where at least one parent plant has the desired allele in its genome. Alternatively, for example, transfer of an allele can be achieved by recombination between two donor genomes, e.g., in fused protoplasts, where at least one donor protoplast has the desired allele in its genome. The desired allele can be, for example, a transgene, a modified (mutated or edited) native allele, or a selected allele of a marker or QTL.

[0479] The term "isogenic line" is a comparative term, referring to genetically identical but differently treated reference organisms. In one example, two genetically identical maize plant embryos can be separated into two distinct groups: one that has undergone a treatment (such as the introduction of a CRISPR-Cas effector endonuclease) and the other that has not undergone such a treatment, a control. In this way, phenotypic differences between the two groups can be attributed solely to the treatment and not to any inherent nature of the plant's endogenous genetic makeup.

[0480] "Introducing" is intended to mean providing a polynucleotide or polypeptide or polynucleotide-protein complex to a target, such as a cell or organism, in such a way that the component enters the interior of a cell of the organism or into the cell itself.

[0481] A "polynucleotide of interest" includes a nucleotide sequence that encodes a protein or polypeptide that improves the desirability of a crop, i.e., a trait of agronomic interest. Polynucleotides of interest include, but are not limited to, polynucleotides that encode traits of interest for agronomically important traits, herbicide resistance, insecticide resistance, disease resistance, nematode resistance, herbicide resistance, microbial resistance, fungal resistance, viral resistance, fertility or sterility, crop characteristics, commercial products, phenotypic markers, or other traits of agricultural or commercial importance. Polynucleotides of interest may also be utilized in sense or antisense orientation. Additionally, two or more polynucleotides of interest may be utilized together, or "stacked," to provide additional benefits.

[0482] A "complex trait locus" includes a genomic locus that has multiple transgenes genetically linked to each other.

[0483] The compositions and methods herein may provide plants with improved "agronomic traits" or "agronomically important traits" or "agronomically beneficial traits", which may include, but are not limited to, disease resistance, drought tolerance, heat tolerance, cold tolerance, salt tolerance, salt tolerance, metal tolerance, herbicide tolerance, improved water use efficiency, improved nitrogen utilization, improved nitrogen fixation, pest resistance, herbivore resistance, pathogen resistance, increased yield, enhanced health, improved vigor, improved growth, improved photosynthetic capacity, nutritional enhancement, modified protein content, modified oil content, increased biomass, increased root length, improved root architecture, modulation of metabolites, modulation of the proteome, increased seed weight, modified seed carbohydrate composition, modified seed oil composition, modified seed protein composition, modified seed nutrient composition, compared to an isogenic plant not modified by the methods or compositions described herein.

[0484] "Agronomic trait potential" is intended to mean the ability of a plant element, at a certain point in its life cycle, to exhibit a phenotype, preferably an improved agronomic trait, or to transmit said phenotype to another related plant element of the same plant.

[0485] As used herein, the terms "reduced," "less," "slower," and "increased," "faster," "enhanced," and "greater" refer to a decrease or increase in a property of a modified plant element or resulting plant compared to an unmodified plant element or resulting plant. For example, the decrease in the property may be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, 5% to 10%, at least 10%, 10% to 20%, at least 15%, at least 20%, 20% to 30%, at least 25%, at least 30%, 30% to 40%, at least 35%, at least 40%, 40% to 50%, at least 45%, at least 50%, 50% to 60%, at least about 60%, 60% to 70%, 70% to 80%, at least 75%, at least about 80%, 80% to 90%, at least about 90%, 90% to 100%, at least 100%, 100% to 200%, at least 200%, at least about 300%, at least about 400%) or more less than an untreated control. The increase in the property obtained can be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, 5% to 10%, at least 10%, 10% to 20%, at least 15%, at least 20%, 20% to 30%, at least 25%, at least 30%, 30% to 40%, at least 35%, at least 40%, 40% to 50%, at least 45%, at least 50%, 50% to 60%, at least about 60%, 60% to 70%, 70% to 80%, at least 75%, at least about 80%, 80% to 90%, at least about 90%, 90% to 100%, at least 100%, 100% to 200%, at least 200%, at least about 300%, at least about 400% or more greater than an untreated control.

[0486] As used herein, the term "before," in reference to sequence position, refers to the occurrence of one sequence upstream or 5' of another sequence.

[0487] The abbreviations have the following meanings: "sec" means second, "min" means minute, "h" means hour, "d" means day, "μL" means microliter, "mL" means milliliter, "L" means liter, "μM" means micromolar, "mM" means millimolar, "M" means mole, "mmol" means millimole, "μmole" or "umole" means micromole, "g" means gram, "μg" or "ug" means microgram, "ng" means nanogram, "U" means unit, "bp" means base pair, and "kb" means kilobase.

[0488] Classification of CRISPR-Cas systems CRISPR-Cas systems have been classified according to the sequence and structural analysis of their components. Several CRISPR / Cas systems have been described, including class 1 systems with multisubunit effector complexes (including types I, III, and IV) and class 2 systems with single protein effectors (including types II, V, and VI) (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15; Zetsche et al. 2015, Cell 163, 1-13; Shmakov et al. 2015, Molecular Cell 60, 1-13; Haft et al. 2005, Computational Biology, PLoS Comput Biol 1(6):e60; and Koonin et al. 2017, Curr Opinion Microbiology 37:67-78).

[0489] The CRISPR-Cas system comprises, at a minimum, a CRISPR RNA (crRNA) molecule and at least one CRISPR-associated (Cas) protein, which form a crRNA-ribonucleoprotein (crRNP) effector complex. The CRISPR-Cas locus contains an array of identical repeats interspersed with DNA targeting spacers encoding the crRNA components and an operon-like unit of cas genes encoding the Cas protein components. The resulting ribonucleoprotein complex recognizes polynucleotides in a sequence-specific manner (Jore et al., Nature Structural & Molecular Biology 18, 529-536 (2011)). The crRNA functions as a guide RNA, directing the sequence-specific binding of the effector (protein or complex) to double-stranded DNA sequences by base-pairing with the complementary DNA strand and forming a so-called R-loop, while displacing the non-complementary strand (Jore et al., 2011, Nature Structural & Molecular Biology 18, 529-536).

[0490] The RNA transcripts (pre-crRNA) of the CRISPR locus are specifically cleaved at the repeat sequences by CRISPR-associated (Cas) endoribonucleases in Type I and Type III systems, or by RNase III in Type II systems. The number of CRISPR-associated genes at a given CRISPR locus can vary between species.

[0491] Different cas genes, encoding proteins with distinct domains, exist in different CRISPR systems. The cas operon contains genes encoding one or more effector endonucleases and other Cas proteins. Protein subunits include those described in Makarova et al. 2011, Nat Rev Microbiol. 2011 9(6):467-477; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15; and Koonin et al. 2017, Current Opinion Microbiology 37:67-78. Domain types include those involved in expression (pre-crRNA processing, e.g., Cas6 or RNase III), buffering (including effector modules for crRNA and target binding and domains for target cleavage), adaptation (spacer insertion, e.g., Cas1 or Cas2), and assistance (regulatory, helper, or unknown function). Some domains can serve more than one purpose; for example, Cas9 contains a domain for endonuclease function and a domain for target cleavage.

[0492] Cas endonucleases are guided by a single CRISPR RNA (crRNA) through direct RNA-DNA base pairing and recognize DNA target sites adjacent to a protospacer adjacent motif (PAM) (Jore, MM et al., 2011, Nat. Struct. Mol. Biol. 18:529-536; Westra, ER et al., 2012, Molecular Cell 46:595-605; and Sinkunas, T. et al., 2013, EMBO J. 32:385-394).

[0493] Class I CRISPR-Cas systems Class I CRISPR-Cas systems include types I, III, and IV. A distinctive feature of class I systems is the presence of an effector endonuclease complex instead of a single protein. The Cascade complex contains an RNA recognition motif (RRM) and a nucleic acid-binding domain, which is the core fold of the diverse RAMP (repeat-associated mysterious protein) protein superfamily (Makarova et al. 2013, Biochem Soc Trans 41, 1392-1400; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15). RAMP protein subunits include Cas5 and Cas7 (which comprise the scaffold of the crRNA-effector complex), where the Cas5 subunit binds to the 5' handle of the crRNA and interacts with the large subunit, and Cas6, which is often loosely associated with the effector complex and generally functions as a repeat-specific RNase in pre-crRNA processing (Charpentier et al., FEMS Microbiol Rev 2015, 39:428-441; Niewoehner et al., RNA 2016, 22:318-329).

[0494] Type I CRISPR-Cas systems contain a complex of effector proteins called Cascade (CRISPR-associated complex for antiviral defense), which includes at least Cas5 and Cas7. The effector complex functions with a single CRISPR RNA (crRNA) and Cas3 to defend against invading viral DNA (Brouns, SJJ et al. Science 321:960-964; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15). Type I CRISPR-Cas loci contain the signature gene cas3 (or variants cas3' or cas3"), which encodes a metal-dependent nuclease with a single-stranded DNA (ssDNA)-stimulating superfamily 2 helicase that has demonstrated the ability to unwind double-stranded DNA (dsDNA) and RNA-DNA duplexes (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15). Following target recognition, the Cas3 endonuclease is recruited to the Cascade-crRNA-target DNA complex to cleave and degrade the DNA target (Westra, E.R. et al. (2012) Molecular Cell 46:595-605; Sinkunas, T. et al. (2011) EMBO J. 30:1335-1342; and Sinkunas, T. et al. (2013) EMBO J. 32:385-394). In some type I systems, Cas6 may be the active endonuclease involved in crRNA processing, and Cas5 and Cas7 function as non-catalytic RNA-binding proteins, whereas in type I-C systems, crRNA processing may be catalyzed by Cas5 (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15). Type I systems are divided into seven subtypes (Makarova et al. 2011, Nat Rev Microbiol. 2011 9(6):467-477; Koonin et al. 2017, Curr Opinion Microbiology 37:67-78).A modified type I CRISPR-associated complex (Cascade) for adaptive antiviral defense comprising at least the protein subunits Cas7, Cas5, and Cas6, one of which is synthetically fused to the Cas3 endonuclease or the modified restriction endonuclease FokI, has been described (WO 2013 / 098244, published July 4, 2013).

[0495] Type III CRISPR-Cas systems, containing multiple cas7 genes, target either ssRNA or ssDNA and function as either an RNase or a target RNA-activated DNA nuclease (Tamulaitis et al., Trends in Microbiology 25(10)49-61, 2017). The Csm (type III-A) and Cmr (type III-B) complexes function as RNA-activated single-stranded (ss)DNases that couple target RNA binding / cleavage with ssDNA degradation. Upon infection with foreign DNA, CRISPR RNA (crRNA)-guided binding of the Csm or Cmr complex to nascent transcripts recruits the Cas10 DNase to actively transcribed phage DNA, resulting in degradation of both the transcript and phage DNA, but not host DNA. The Cas10 HD domain is responsible for the ssDNase activity, and the Csm3 / Cmr4 subunits are responsible for the endoribonuclease activity of the Csm / Cmr complex. The 3' flanking sequence of the target RNA is important for the ssDNase activity of Csm / Cmr: base pairing with the 5' handle of the crRNA protects the host DNA from degradation.

[0496] Type IV systems contain typical type I cas5 and cas7 domains in addition to a cas8-like domain, but can lack the CRISPR array that is characteristic of most other CRISPR-Cas systems.

[0497] Class II CRISPR-Cas system Class II CRISPR-Cas systems include types II, V, and VI. A distinctive feature of class II systems is the presence of a single Cas effector protein instead of an effector complex. Type II and V Cas proteins contain a RuvC endonuclease domain that adopts an RNase H fold.

[0498] Type II CRISPR / Cas systems use crRNA and tracrRNA (trans-activating CRISPR RNA) to guide the Cas endonuclease to the DNA target. The crRNA contains a spacer region that is complementary to one strand of the double-stranded DNA target and a region that base-pairs with the tracrRNA (trans-activating CRISPR RNA) to form an RNA duplex that allows the Cas endonuclease to cleave the DNA target, leaving a blunt end. The spacer is obtained through a poorly understood process involving the Cas1 and Cas2 proteins. Type II CRISPR / Cas loci typically contain the cas1 and cas2 genes in addition to the cas9 gene (Chylinski et al., 2013, RNA Biology 10:726-737; Makarova et al., 2015, Nature Reviews Microbiology Vol. 13:1-15). Type II CRISR-Cas loci can encode tracrRNAs that are partially complementary to repeats within each CRISPR array and can contain other proteins such as Csn1 and Csn2. The presence of cas9 adjacent to the cas1 and cas2 genes is a hallmark of type II loci (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15).

[0499] Type V CRISPR / Cas systems contain a single Cas endonuclease, including Cpf1 (Cas12) (Koonin et al., Curr Opinion Microbiology 37:67-78, 2017), which, unlike Cas9, is an active RNA-guided endonuclease that does not necessarily require an additional trans-activating CRISPR (tracr) RNA for target cleavage.

[0500] Type VI CRISPR-Cas systems contain two HEPN (higher eukaryotic and prokaryotic nucleotide-binding) domains, but lack the HNH or RuvC domains and contain the cas13 gene, which is independent of tracrRNA activity. The majority of the HEPN domain contains a conserved motif that constitutes a metal-independent endo-RNase active site (Anantharam et al., Biol Direct 8:15, 2013). This feature suggests that type VI systems act on RNA targets instead of the DNA targets common to other CRISPR-Cas systems.

[0501] Novel CRISPR-Cas system Disclosed herein is a novel CRISPR-Cas system, its components, and methods of using said components. The system includes a novel Cas effector protein, Cas-alpha.

[0502] The novel CRISPR-Cas system components described herein may include one or more subunits from different Cas systems, subunits derived from or modified from two or more different bacterial or archaeal prokaryotes, and / or synthetic or engineered components.

[0503] Described herein is a newly identified CRISPR-Cas system that includes novel sequences for the cas genes. Additionally, novel cas genes and proteins are described.

[0504] One of several features of the novel Cas-alpha system is the locus structure, as shown in Figures 1A-1D. In some embodiments, the Cas-alpha genomic locus includes a cas1 gene, a cas2 gene, a cas4 gene, and a cas-alpha gene encoding the effector protein Cas-alpha. A CRISPR array containing repeats of nucleotide sequences can be found before or after the gene encoding the Cas-alpha endonuclease. In some embodiments, the cas-alpha locus includes a cas-alpha gene encoding the effector protein and a CRISPR array containing repeats, but may not include any one or more of the cas1 gene, cas2 gene, and / or cas4 gene.

[0505] CRISPR-Cas system components Cas proteins Numerous proteins can be encoded in the CRISPR cas operon, including proteins involved in adaptation (spacer insertion), interference (effector module target binding, target nicking or cleavage—e.g., endonuclease activity), expression (pre-crRNA processing), regulation, and more.

[0506] Two proteins, Cas1 and Cas2, are conserved among many CRISPR systems (see, for example, Koonin et al., Curr Opinion Microbiology 37:67-78, 2017). Cas1 is a metal-dependent DNA-specific endonuclease that generates double-stranded DNA fragments. In some systems, Cas1 forms a stable complex with Cas2, which is essential for spacer acquisition and insertion for CRISPR systems (Nunez et al., Nature Str Mol Biol 21:528-534, 2014).

[0507] Many other proteins have been identified in various systems, including Cas4 (which may have similarity to RecB nuclease), and are thought to play a role in capturing new viral DNA sequences for incorporation into CRISPR arrays (Zhang et al., PLOS One 7(10):e47232, 2012).

[0508] Some proteins may encompass multiple functions, for example, Cas9, the signature protein of class 2 type II systems, has been demonstrated to be involved in pre-crRNA processing, target binding, and target cleavage.

[0509] The novel Cas-alpha proteins disclosed herein include effector proteins (endonucleases) and adaptation proteins. Cas endonucleases have been identified from several bacterial and archaeal sources, including those shown in Figures 7A-7K.

[0510] Cas endonucleases and effectors Endonucleases are enzymes that cleave the phosphodiester bonds of polynucleotide chains, including restriction endonucleases that cleave DNA at specific sites without damaging bases. Examples of endonucleases include restriction endonucleases, meganucleases, TAL effector nucleases (TALENs), zinc finger nucleases, and Cas (CRISPR-associated) effector endonucleases.

[0511] Cas endonucleases, either as single effector proteins or in effector complexes with other components, unwind DNA duplexes at target sequences and, optionally, cleave at least one DNA strand, mediated by recognition of the target sequence by a polynucleotide (such as, but not limited to, a crRNA or guide RNA) complexed with the Cas effector protein. Generally, such recognition and cleavage of a target sequence by a Cas endonuclease occurs if the correct protospacer adjacent motif (PAM) is located at or adjacent to the 3' end of the DNA target sequence. Alternatively, the Cas endonucleases herein may lack DNA cleavage or nicking activity when complexed with a suitable RNA component, but still be capable of specifically binding to a DNA target sequence. (See also U.S. Patent Application Publication No. 2015 / 0082478, published March 19, 2015, and U.S. Patent Application Publication No. 2015 / 0059010, published February 26, 2015).

[0512] Cas endonucleases can occur as individual effectors (class 2 CRISPR systems) or as part of larger effector complexes (class I CRISPR systems).

[0513] Described Cas endonucleases include, but are not limited to, Cas3 (characteristic of class 1 type I systems), Cas9 (characteristic of class 2 type II systems), and Cas12 (Cpf1) (characteristic of class 2 type V systems).

[0514] Cas3 (and its variants Cas3' and Cas3") functions as a single-stranded DNA nuclease (HD domain) and an ATP-dependent helicase. Variants of the Cas3 endonuclease can be obtained by abolishing the functional activity of one or both domains of the Cas3 endonuclease polypeptide. ATPase-dependent helicase activity can be abolished (by deletion, knockout of the Cas3-helicase domain, or by mutagenesis of key residues), or by assembling the reaction in the absence of ATP as previously described (Sinkunas, T, et al., 2013, EMBO J. 32:385-394), disabling the HD endonuclease activity can convert a cleavage-ready cascade containing a modified Cas3 endonuclease into a nickase (since the HD domain is still functional). Disabling HD endonuclease activity can be achieved by methods known in the art, including but not limited to, mutagenesis of key residues in the HD domain, and converting a cleavage-ready cascade containing a modified Cas3 endonuclease into a helicase. Disabling both Cas helicase and Cas3 HD endonuclease activity can be achieved by methods known in the art, including but not limited to, mutagenesis of key residues in both the helicase and HD domains, and converting a cleavage-ready cascade containing a modified Cas3 endonuclease into a binding protein that binds to a target sequence.

[0515] "Cas9" (previously called Cas5, Csn1, or Csx12) is a Cas endonuclease that specifically recognizes and cleaves all or part of a DNA target sequence in complex with cr and tracr nucleotides or a single guide polynucleotide. Cas9 recognizes the 3' GC-rich PAM sequence of target dsDNA. The Cas9 protein contains a RuvC nuclease with an HNH (HNH) nuclease adjacent to the RuvC-II domain. The RuvC nuclease and HNH nuclease can each cleave one DNA strand at the target sequence (the concerted activity of both domains results in a DNA double-strand break, while the activity of one domain results in a nick). Generally, the RuvC domain contains subdomains I, II, and III; domain I is located near the N-terminus of Cas9, and subdomains II and III are located in the center of the protein and adjacent to the HNH domain (Hsu et al., Cell 157:1262-1278). Cas9 endonuclease is typically derived from a type II CRISPR system, which includes a DNA cleavage system that utilizes Cas9 endonuclease complexed with at least one polynucleotide component. For example, Cas9 can be complexed with CRISPR RNA (crRNA) and trans-activating CRISPR RNA (tracrRNA). In another example, Cas9 can be complexed with a single guide RNA (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15).

[0516] Cas12 (previously called Cpf1 and variants c2c1, c2c3, CasX, and CasY) contains a RuvC nuclease domain to generate 5' overhangs on dsDNA targets. Some variants do not require tracrRNA, unlike Cas9 functionality. Cas12 and its variants recognize 5' AT-rich PAM sequences on target dsDNA. The insertion domain of the Cas12a protein, called Nuc, has been demonstrated to be involved in target strand cleavage (Yamano et al., Cell 2016, 165:949-962). Further mutation studies in other Cas12 proteins have demonstrated that the Nuc domain, along with the RuvC domain responsible for cleavage, contributes to guidance and target binding (Swarts et al., Mol Cell 2017, 66:221-233 e224).

[0517] Cas endonucleases and effector proteins can be used for targeted genome editing (through single and multiple double-strand breaks and nicks) and targeted genome regulation (by attaching epigenetic effector domains to Cas proteins or sgRNAs). Cas endonucleases can also be engineered to function as RNA-guided recombinases and, through RNA tethers, can serve as scaffolds for the assembly of multiprotein and nucleic acid complexes (Mali et al., 2013, Nature Methods Vol. 10:957-963).

[0518] Cas-alpha endonucleases Cas-alpha endonucleases are defined as functional RNA-guided, PAM-dependent dsDNA cleavage proteins consisting of less than 800 amino acids and divided into three subdomains: a C-terminal RuvC catalytic domain containing a bridge helix and one or more zinc finger motifs; and an N-terminal Rec subunit with a helical bundle, a WED wedge (or "oligonucleotide binding domain," OBD) domain, and optionally a zinc finger motif.

[0519] The Cas-alpha endonuclease, when aligned with SEQ ID NO: 17, contains at least one, at least two, at least three, at least four, at least five, at least six, or seven of the following amino acid positions in SEQ ID NO: 17: glycine (G) at position 337, glycine (G) at position 341, glutamic acid (E) at position 430, leucine (L) at position 432, cysteine ​​(C) at position 487, cysteine ​​(C) at position 490, cysteine ​​(C) at position 507, and / or cysteine ​​(C) or histidine (H) at position 512. The Cas-alpha endonuclease is The following motifs: GxxxG, ExL, Cx n C, Cx n (C or H) (where n = one or more amino acids).

[0520] The RuvC domain has been documented to encompass endonuclease functionality. Cas-alpha endonucleases can be isolated or characterized from loci that contain cas-alpha genes encoding effector proteins and arrays containing multiple repeats. In some embodiments, the cas-alpha locus can further contain part or all of the cas1 gene, the cas2 gene, and / or the cas4 gene.

[0521] Zinc finger motifs are domains that coordinate one or more zinc ions, usually by cysteine ​​and histidine side chains, to stabilize their fold. Zinc fingers are named for the pattern of cysteine ​​and histidine residues that coordinate the zinc ion (e.g., C4 means that the zinc ion is coordinated by four cysteine ​​residues, while CH means that the zinc ion is coordinated by three cysteine ​​residues and one histidine residue).

[0522] Cas-alpha proteins contain one or more zinc finger (ZFN) coordination motifs that can form zinc-binding domains. Zinc finger-like motifs can aid in the separation of target and non-target strands and the loading of guide RNAs onto DNA targets. Cas-alpha proteins containing one or more zinc finger motifs can provide additional stability to ribonucleoprotein complexes on target polynucleotides. Cas-alpha proteins contain C4 or C3H zinc-binding domains.

[0523] Some Cas-alpha proteins and polynucleotides are shown in Figures 7A-7K, and key structural motifs of endonuclease proteins are shown in Figures 8A-8K, respectively.

[0524] Cas-alpha endonucleases are RNA-guided endonucleases that can bind to and cleave double-stranded DNA targets that contain (1) a sequence homologous to the nucleotide sequence of a guide RNA and (2) a PAM sequence. In some embodiments, the PAM is T-rich. In some embodiments, the PAM is C-rich.

[0525] Cas-alpha endonucleases function as double-strand break inducers and may also be nickases or single-strand break inducers. In some embodiments, catalytically inactive Cas-alpha endonucleases can be used to target or guide to a target DNA sequence, but do not induce cleavage. In some embodiments, catalytically inactive Cas-alpha proteins can be used in conjunction with functional endonucleases that cleave the target sequence. In some embodiments, catalytically inactive Cas-alpha proteins can be combined with base-editing molecules such as deaminases. In some embodiments, the deaminase can be cytidine deaminase. In some embodiments, the deaminase can be adenine deaminase. In some embodiments, the deaminase can be ADAR-2.

[0526] Cas-alpha endonucleases are further defined as SEQ ID NOs: 17, 18, 19, 20, 32, 33, 34, 35, 36, 37, 38, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 30 0, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361 , 362, 363, 364, 365, 366, 367, 368, 369, 370, and 371, or a functional fragment or functional variant thereof retaining at least some activity, and , at least 50%, 50% to 55%, at least 55%, 55% to 60%, at least 60%, 60% to 65%, at least 65%, 65% to 70%, at least 70%, 70% to 75%, at least 75%, 75% to 80%, at least 80%, 80% to 85%, at least 85%, 85% to 90%, at least 90%, at least 90% to 95%, at least 95%, 95% to 96%, at least 96%, 96% to 97%, at least 97%, 97% to 98%, at least 98%, 98% to 99%, at least 99%, 99% to 100%,or 100% sequence identity. A "functional fragment" of a Cas-alpha endonuclease retains the ability to recognize, bind to, or nick one strand of a double-stranded polynucleotide, or the ability to cleave both strands of a double-stranded polynucleotide, or any combination thereof.

[0527] The Cas-alpha endonuclease may be selected from the group consisting of at least 50, 50-100, at least 100, 100-150, at least 150, 150-200, at least 200, 200-250, at least 250, 250-300, at least 300, 300-350, at least 350, 350-400, at least 400, 400-450, at least 500, 500-550, at least 600, 600-650, at least 650, 650-700, At least 700, 700-750, at least 750, 750-800, at least 800, 800-850, at least 850, 850-900, at least 900, 900-950, at least 950, 950-1000, at least 1000, or more than 1000 consecutive nucleotides and at least 50%, 50%-55%, at least 55%, 55%-60%, at least 60%, 60%-65%, at least 65%, 65%-70%, at least 70%, 70%-75%, at least 75%, 75%-80%, at least 80%, 80% a polynucleotide having up to 85%, at least 85%, 85% to 90%, at least 90%, 90% to 95%, at least 95%, 95% to 96%, at least 96%, 96% to 97%, at least 97%, 97% to 98%, at least 98%, 98% to 99%, at least 99%, 99% to 100%, or 100% sequence identity to SEQ ID NOs: 17, 18, 19, 20, 32, 33, 34, 35, 36, 37, 38, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 380, 381, 382, ​​383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 8, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299 , 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330,331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, and 371.

[0528] Cas endonucleases, effector proteins, or functional fragments thereof for use in the disclosed methods can be isolated from natural sources or from recombinant sources in which genetically engineered host cells have been modified to express nucleic acid sequences encoding the proteins. Alternatively, Cas proteins can be produced using cell-free protein expression systems or produced synthetically. Effector Cas nucleases can be isolated and introduced into heterologous cells and modified from their native form to exhibit a type or magnitude of activity different from that of their native source. Such modifications include, but are not limited to, fragments, variants, substitutions, deletions, and insertions.

[0529] Fragments and variants of Cas endonucleases and Cas effector proteins can be obtained by methods such as site-directed mutagenesis and synthetic construction. Methods for measuring endonuclease activity are well known in the art, including, but not limited to, those described in WO 2013 / 166113, published November 7, 2013; WO 2016 / 186953, published November 24, 2016; and WO 2016 / 186946, published November 24, 2016.

[0530] Cas endonucleases can include modified forms of Cas polypeptides. Modified forms of Cas polypeptides can include amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the naturally occurring nuclease activity of the Cas protein. For example, in some instances, modified forms of Cas proteins have less than 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nuclease activity of the corresponding wild-type Cas polypeptide (U.S. Patent Application Publication No. 2014 / 0068797, published March 6, 2014). In some instances, modified forms of Cas polypeptides have substantially no nuclease activity and are referred to as catalytically "inactivated Cas" or "deactivated Cas (dCas)." Inactivated Cas / deactivated Cas includes deactivated Cas endonuclease (dCas). Catalytically inactive Cas effector proteins can be fused to heterologous sequences to induce or modify activity.

[0531] The Cas endonuclease may be part of a fusion protein that includes one or more heterologous protein domains (e.g., one, two, three or more domains in addition to the Cas protein). Such a fusion protein may include any additional protein sequences and, optionally, a linker sequence between any two domains, e.g., between the Cas and the first heterologous domain. Examples of protein domains that can be fused to the Cas proteins herein include, but are not limited to, epitope tags (e.g., histidine [His], V5, FLAG, influenza hemagglutinin [HA], myc, VSV-G, thioredoxin [Trx]), reporters (e.g., glutathione-5-transferase [GST], horseradish peroxidase [HRP], chloramphenicol acetyltransferase [CAT], beta-galactosidase, beta-glucuronidase [GUS], luciferase, green fluorescent protein [GFP], HcRed, DsRed, cyan fluorescent protein [CFP], yellow fluorescent protein [YFP], blue fluorescent protein [BFP]), and domains having one or more of the following activities: methylase activity, demethylase activity, transcriptional activator activity (e.g., VP16 or VP64), transcriptional repressor activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Cas proteins can also be fused to proteins that bind to DNA molecules or other molecules, such as maltose binding protein (MBP), S-tag, Lex A DNA binding domain (DBD), GAL4A DNA binding domain, and herpes simplex virus (HSV) VP16.

[0532] Catalytically active and / or inactive Cas endonucleases can be fused to heterologous sequences (U.S. Patent Application Publication No. 2014 / 0068797, published March 6, 2014). Suitable fusion partners include, but are not limited to, polypeptides that provide activities that indirectly increase transcription by acting directly on target DNA or polypeptides associated with target DNA (e.g., histones or other DNA-binding proteins). Additional suitable fusion partners include, but are not limited to, polypeptides that provide methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, or demyristoylating activity. Further suitable fusion partners include, but are not limited to, polypeptides that directly cause increased transcription of a target nucleic acid (e.g., transcriptional activators or fragments thereof, proteins or fragments thereof that induce transcriptional activators, small molecule / drug-responsive transcriptional regulators, etc.) Partially active or catalytically inactive Cas-alpha endonucleases can also be fused to another protein or domain, e.g., Clo51 nuclease or FokI nuclease, to generate double-stranded breaks (Guilinger et al. Nature biotechnology, volume 32, number 6, June 2014).

[0533] Catalytically active or inactive Cas proteins, such as the Cas-alpha proteins described herein, can also be fused to molecules that direct the editing of single or multiple bases in a polynucleotide sequence, e.g., site-specific deaminases that can change the identity of a nucleotide, e.g., from C·G to T·A or from A·T to G·C (Gaudelli et al., "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage," Nature (2017); Nishida et al., "Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems," Science 353 (6305) (2016); Komor et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature (2017)). 533(7603)(2016):420-4. Base editing fusion proteins can include, for example, active (double-strand breaking), partially active (nickases), or inactivated (catalytically inactive) Cas-alpha endonucleases and deaminases (e.g., but not limited to, cytidine deaminase, adenine deaminase, APOBEC1, APOBEC3A, BE2, BE3, BE4, ABE, etc.). Base editing repair inhibitors and glycosylase inhibitors (e.g., uracil glycosylase inhibitors (which prevent the removal of uracil)) are contemplated as other components of the base editing system in some embodiments.

[0534] The Cas endonucleases described herein can be expressed and purified by methods known in the art, for example, as described in WO 2016 / 186953, published November 24, 2016.

[0535] Many Cas endonucleases have been reported to date that can recognize specific PAM sequences (WO 2016 / 186953 published November 24, 2016; WO 2016 / 186946 published November 24, 2016; and Zetsche B et al. 2015. Cell 163, 1013) and cleave target DNA at specific locations. Of course, those skilled in the art will understand that based on the methods and embodiments described herein that use the novel inducible Cas system, these methods can be adapted to use any inducible endonuclease system.

[0536] The Cas effector protein may comprise a heterologous nuclear localization sequence (NLS). The heterologous NLS amino acid sequence herein may be strong enough to drive the accumulation of a detectable amount of the Cas protein in the nucleus of, for example, a yeast cell herein. The NLS may comprise one (monopartite) or multiple (e.g., bipartite) short sequences (e.g., 2-20 residues) of basic, positively charged residues (e.g., lysine and / or arginine) and may be positioned anywhere within the Cas amino acid sequence, provided that it is exposed on the protein surface. The NLS may be operably linked, for example, to the N-terminus or C-terminus of the Cas protein herein. For example, two or more NLS sequences may be linked to the Cas protein, for example, to the N-terminus and C-terminus of the Cas protein. The Cas endonuclease gene can be operably linked to an SV40 nuclear targeting signal upstream of the Cas codon region and a bipartite VirD2 nuclear localization signal (Tinland et al. (1992) Proc. Natl. Acad. Sci. USA 89:7442-6) downstream of the Cas codon region. Non-limiting examples of suitable NLS sequences herein include those disclosed in U.S. Patent Nos. 6,660,830 and 7,309,576.

[0537] Guide polynucleotide The guide polynucleotide enables the Cas endonuclease to recognize, bind to, and optionally cleave the target, and can be a single molecule or a double molecule. The guide polynucleotide sequence can be an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combination sequence). Optionally, the guide polynucleotide can include at least one nucleotide, a phosphodiester bond, or a linkage modification, such as, but not limited to, locked nucleic acid (LNA), 5-methyl dC, 2,6-diaminopurine, 2'-fluoro A, 2'-fluoro U, 2'-O-methyl RNA, phosphorothioate bond, a linkage to a cholesterol molecule, a linkage to a polyethylene glycol molecule, a linkage to a spacer 18 (hexaethylene glycol chain) molecule, or a 5' to 3' covalent linkage that results in cyclization. A guide polynucleotide containing only ribonucleic acid is also called a "guide RNA" or "gRNA" (U.S. Patent Application Publication No. 2015 / 0082478, published March 19, 2015, and U.S. Patent Application Publication No. 2015 / 0059010, published February 26, 2015). Guide polynucleotides can be engineered or synthetically produced.

[0538] Guide polynucleotides include chimeric non-natural guide RNAs (i.e., they are heterologous to each other) that contain regions that are not found together in nature, such as a chimeric non-natural guide RNA that contains a first nucleotide sequence domain (called a variable targeting domain or VT domain) that can hybridize to a nucleotide sequence in a target DNA linked to a second nucleotide sequence that can recognize a Cas endonuclease, where the first and second nucleotide sequences are not found linked together in nature.

[0539] The guide polynucleotide can be a duplex molecule (also called a double-stranded guide polynucleotide) that includes a cr nucleotide sequence (e.g., crRNA) and a tracr nucleotide sequence (e.g., tracrRNA). Optionally, a linker polynucleotide is present that connects the crRNA and tracrRNA to form a single guide, e.g., sgRNA.

[0540] The cr nucleotide comprises a first nucleotide sequence domain (called the variable targeting domain or VT domain) capable of hybridizing to a nucleotide sequence in the target DNA, and a second nucleotide sequence (also called the tracr mate sequence) that is part of the Cas endonuclease recognition (CER) domain. The tracr mate sequence can hybridize to the tracr nucleotide along a complementary region, together forming a Cas endonuclease recognition domain, i.e., a CER domain. The CER domain can interact with a Cas endonuclease polypeptide. The cr nucleotide and tracr nucleotide of the double-stranded guide polynucleotide can be RNA, DNA, and / or RNA-DNA combination sequences. In some embodiments, the cr nucleotide molecule of the double-stranded guide polynucleotide is referred to as "crDNA" (when composed of a continuous stretch of DNA nucleotides), "crNNA" (when composed of a continuous stretch of RNA nucleotides), or "crDNA-RNA" (when composed of a combination of DNA and RNA nucleotides). The cr nucleotide can comprise a fragment of the crRNA naturally occurring in Bacteria and Archaea. The size of the fragments of naturally occurring crRNA in Bacteria and Archaea that can be present in the cr nucleotides disclosed herein can vary from, but are not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. In some embodiments, the crRNA molecule is selected from the group consisting of SEQ ID NOs: 57, 58, and 59.

[0541] In some embodiments, the tracr nucleotide sequence is referred to as "tracrRNA" (when composed of a continuous stretch of RNA nucleotides), or "tracrDNA" (when composed of a continuous stretch of DNA nucleotides), or "tracrDNA-RNA" (when composed of a combination of DNA and RNA nucleotides). In one embodiment, the RNA that guides the RNA / Cas9 endonuclease complex is a double-stranded RNA comprising a double-stranded crRNA-tracrRNA. tracrRNA (trans-activating CRISPR RNA) contains, from 5' to 3', (i) a sequence that anneals to the repeat region of CRISPR type II crRNA, and (ii) a stem-loop-containing portion (Deltcheva et al., Nature 471:602-607). A double-stranded guide polynucleotide can form a complex with a Cas endonuclease, and the guide polynucleotide / Cas endonuclease complex (also called a guide polynucleotide / Cas endonuclease system) can guide the Cas endonuclease to a genomic target site, allowing the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single- or double-strand break) the target site (U.S. Patent Application Publication No. 2015 / 0082478, published March 19, 2015, and U.S. Patent Application Publication No. 2015 / 0059010, published February 26, 2015).

[0542] In some embodiments, the tracrRNA molecule is selected from the group consisting of SEQ ID NOs: 60-68.

[0543] In one embodiment, the guide polynucleotide is a guide polynucleotide capable of forming a PGEN as described herein, wherein said guide polynucleotide comprises a first nucleotide sequence domain that is complementary to a nucleotide sequence of a target DNA and a second nucleotide sequence domain that interacts with said Cas endonuclease polypeptide.

[0544] In one aspect, the guide polynucleotide is a guide polynucleotide described herein, wherein the first nucleotide sequence and the second nucleotide sequence domain are guide polynucleotides selected from the group consisting of DNA sequences, RNA sequences, and combinations thereof.

[0545] In one aspect, the guide polynucleotide is a guide polynucleotide described herein, wherein the first nucleotide sequence and the second nucleotide sequence domain are selected from the group consisting of stability-enhancing RNA backbone modifications, stability-enhancing DNA backbone modifications, and combinations thereof (see Kanasty et al., 2013, Common RNA-backbone modifications, Nature Materials 12:976-977; U.S. Patent Application Publication No. 2015 / 0082478 published March 19, 2015; U.S. Patent Application Publication No. 2015 / 0059010 published February 26, 2015).

[0546] The guide RNA includes a duplex molecule containing a chimeric non-natural crRNA linked to at least one tracrRNA. The chimeric non-natural crRNA includes a crRNA that contains regions that are not found together in nature (i.e., they are heterologous to each other). For example, a crRNA that contains a first nucleotide sequence domain (called a variable targeting domain or VT domain) that can hybridize to a nucleotide sequence in target DNA linked to a second nucleotide sequence (also called a tracrmate sequence), where the first and second sequences are not found linked together in nature.

[0547] The guide polynucleotide may also be a single molecule (also referred to as a single guide polynucleotide) comprising a cr nucleotide sequence linked to a tracr nucleotide sequence. A single guide polynucleotide comprises a first nucleotide sequence domain (called a variable targeting domain or VT domain) capable of hybridizing to a nucleotide sequence in the target DNA, and a Cas endonuclease recognition domain (CER domain) that interacts with a Cas endonuclease polypeptide. In some embodiments, the sgRNA molecule is selected from the group consisting of SEQ ID NOs: 69-77.

[0548] The VT and / or CER domains of a single guide polynucleotide can comprise RNA, DNA, or a combined RNA-DNA sequence. A single guide polynucleotide composed of a sequence derived from cr and tracr nucleotides can be referred to as a "single guide RNA" (when composed of a continuous stretch of RNA nucleotides), a "single guide DNA" (when composed of a continuous stretch of DNA nucleotides), or a "single guide RNA-DNA" (when composed of a combination of RNA and DNA nucleotides). A single guide polynucleotide can form a complex with a Cas endonuclease, and the guide polynucleotide / Cas endonuclease complex (also referred to as a guide polynucleotide / Cas endonuclease system) can guide the Cas endonuclease to a genomic target site, allowing the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single- or double-strand break) the target site. (U.S. Patent Application Publication No. 2015 / 0082478, published March 19, 2015, and U.S. Patent Application Publication No. 2015 / 0059010, published February 26, 2015).

[0549] Chimeric non-naturally occurring single guide RNAs (sgRNAs) include sgRNAs that contain regions that are not found together in nature (i.e., they are heterologous to each other), such as an sgRNA that contains a first nucleotide sequence domain (called a variable targeting domain or VT domain) that can hybridize to a nucleotide sequence in target DNA linked to a second nucleotide sequence (also called a targeting sequence) that is not found linked together in nature.

[0550] The nucleotide sequence linking the cr and tracr nucleotides of the single-stranded guide polynucleotide can comprise an RNA sequence, a DNA sequence, or a combined RNA-DNA sequence. In one embodiment, the nucleotide sequence linking the cr and tracr nucleotides of the single-stranded guide polynucleotide (also referred to as the "loop") is at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108 It can be 3, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides in length. In another embodiment, the nucleotide sequence linking the cr and tracr nucleotides of the single guide polynucleotide can comprise a tetraloop sequence, such as, but not limited to, a GAAA tetraloop sequence.

[0551] Guide polynucleotides can be produced by methods known in the art, including, but not limited to, chemically synthesizing guide polynucleotides (e.g., Hendel et al. 2015, Nature Biotechnology 33, 985-989), generating guide polynucleotides in vitro, and / or self-splicing guide RNAs (e.g., but not limited to, Xie et al., (2015), PNAS 112:3570-3575).

[0552] Protospacer adjacent motif (PAM) As used herein, the term "protospacer adjacent motif" (PAM) refers to a short nucleotide sequence adjacent to the target sequence (protospacer) recognized (targeted) by the guide polynucleotide / Cas endonuclease system. The Cas endonuclease cannot correctly recognize a target DNA sequence unless the target DNA sequence is followed by a PAM sequence. The sequence and length of the PAM herein may vary depending on the Cas protein or Cas protein complex used. The PAM sequence may be of any length but is generally 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length.

[0553] The terms "randomized PAM" and "randomized protospacer adjacent motif" are used interchangeably herein and refer to a random DNA sequence adjacent to a target sequence (protospacer) recognized (targeted) by a guide polynucleotide / Cas endonuclease system. The randomized PAM sequence can be of any length, but is generally 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. The randomized nucleotides include any of the nucleotides A, C, G, or T.

[0554] Guide polynucleotide / Cas endonuclease complex The guide polynucleotide / Cas endonuclease complexes described herein are capable of recognizing, binding to, and optionally nicking, unnicking, or cleaving all or part of a target sequence.

[0555] A guide polynucleotide / Cas endonuclease complex capable of cleaving both strands of a DNA target sequence typically comprises a Cas protein having all of its endonuclease domains functionally (e.g., a wild-type endonuclease domain, or a variant thereof that retains the activity of part or all of each endonuclease domain). Thus, a wild-type Cas protein (e.g., a Cas protein disclosed herein) or a variant thereof that retains the activity of part or all of each endonuclease domain of the Cas protein are suitable examples of Cas proteins that can cleave both strands of a DNA target sequence.

[0556] A guide polynucleotide / Cas endonuclease complex capable of cleaving a single strand of a DNA target sequence may be characterized herein as having nickase activity (e.g., partial cleavage ability). Cas nickases typically contain a single functional endonuclease domain that enables Cas to cleave (i.e., nick) only one strand of a DNA target sequence. For example, a Cas9 nickase may contain (i) a mutated, dysfunctional RuvC domain and (ii) a functional HNH domain (e.g., a wild-type HNH domain). As another example, a Cas9 nickase may contain (i) a functional RuvC domain (e.g., a wild-type RuvC domain) and (ii) a mutated, dysfunctional HNH domain. Non-limiting examples of Cas9 nickases suitable for use herein are disclosed in U.S. Patent Application Publication No. 2014 / 0189896, published July 3, 2014. To increase the specificity of DNA targeting, a pair of Cas nickases can be used. Generally, this can be accomplished by providing two Cas nickases that are linked to RNA components with different guide sequences to target and nick nearby DNA sequences on opposite strands within the region for the desired targeting. Such nearby cleavage of each DNA strand creates a double-strand break (i.e., a DSB with a single-strand overhang), which is then recognized as a substrate for non-homologous end joining (NHEJ) (prone to imperfect repair leading to mutations) or homologous recombination (HR). In these embodiments, the nicks can be separated from one another by, for example, at least about 5, 5-10, at least 10, 10-15, at least 15, 15-20, at least 20, 20-30, at least 30, 30-40, at least 40, 40-50, at least 50, 50-60, at least 60, 60-70, at least 70, 70-80, at least 80, 80-90, at least 90, 90-100, or 100 or more (or any integer between 5 and 100) bases. One or two Cas nickase proteins herein can be used in a Cas nickase pair.For example, a Cas9 nickase with a mutated RuvC domain but a functional HNH domain (i.e., Cas9 HNH / RuvC) can be used (e.g., Streptococcus pyogenes Cas9 HNH / RuvC). Each Cas9 nickase (e.g., Cas9 HNH / RuvC) can be directed to specific DNA sites close together (up to 100 base pairs apart) by using suitable RNA components herein with guide RNA sequences that target each nickase to its specific DNA site.

[0557] In certain embodiments, the guide polynucleotide / Cas endonuclease complex can bind to a DNA target site sequence but does not cleave any strand at the target site sequence. Such a complex can include a Cas protein in which all of its nuclease domains are mutated and dysfunctional. For example, a Cas9 protein that can bind to a DNA target site sequence but does not cleave any strand at the target site sequence can include both a mutated, dysfunctional RuvC domain and a mutated, dysfunctional HNH domain. Cas proteins herein that bind to but do not cleave a target DNA sequence can be used to regulate gene expression; for example, the Cas protein can be fused to a transcription factor (or portion thereof) (e.g., a repressor or activator, such as any of those disclosed herein).

[0558] In one aspect, a guide polynucleotide / Cas endonuclease complex (PGEN) described herein is a PGEN, wherein said Cas endonuclease is optionally covalently or non-covalently linked to or assembled with at least one protein subunit or functional fragment thereof.

[0559] In one embodiment of the present disclosure, the guide polynucleotide / Cas endonuclease complex is a guide polynucleotide / Cas endonuclease complex (PGEN) comprising at least one guide polynucleotide and at least one Cas endonuclease polypeptide, wherein the Cas endonuclease polypeptide comprises at least one protein subunit or functional fragment thereof, the guide polynucleotide is a chimeric non-natural guide polynucleotide, and the guide polynucleotide / Cas endonuclease complex is capable of recognizing, binding to, and optionally nicking, unnicking, or cleaving all or a portion of a target sequence.

[0560] The Cas effector protein can be a Cas-alpha effector protein disclosed herein.

[0561] In one embodiment of the present disclosure, the guide polynucleotide / Cas effector complex is a guide polynucleotide / Cas effector protein complex (PGEN) comprising at least one guide polynucleotide and a Cas-alpha effector protein, wherein the guide polynucleotide / Cas effector protein complex is capable of recognizing, binding to, and optionally nicking, unnicking, or cleaving all or a portion of a target sequence.

[0562] The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein the Cas effector protein further comprises one or more copies of at least one protein subunit or functional fragment thereof. In some embodiments, the protein subunit is selected from the group consisting of a Cas1 protein subunit, a Cas2 protein subunit, a Cas4 protein subunit, and any combination thereof. The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein the Cas effector protein further comprises at least two different protein subunits selected from the group consisting of Cas1, Cas2, and Cas4.

[0563] The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein the Cas effector protein further comprises at least three different protein subunits, or functional fragments thereof, selected from the group consisting of Cas1, Cas2, and optionally one additional Cas protein, including Cas4.

[0564] In one aspect, the guide polynucleotide / Cas endonuclease complex (PGEN) described herein is a PGEN in which the Cas endonuclease is covalently or non-covalently linked to at least one protein subunit or functional fragment thereof. The PGEN can be a guide polynucleotide / Cas effector protein complex in which the Cas effector protein polypeptide is covalently or non-covalently linked to or assembled with one copy or multiple copies of at least one protein subunit or functional fragment thereof selected from the group consisting of a Cas1 protein subunit, a Cas2 protein subunit, one additional Cas protein subunit optionally including a Cas4 protein, and any combination thereof. The PGEN can be a guide polynucleotide / Cas effector protein complex in which the Cas effector protein is covalently or non-covalently linked to or assembled with at least two different protein subunits selected from the group consisting of Cas1, Cas2, and one additional Cas protein optionally including a Cas4 protein. The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein the Cas effector protein is covalently or non-covalently linked to at least three different protein subunits selected from the group consisting of Cas1, Cas2, and optionally one additional Cas protein, including Cas4, and combinations thereof.

[0565] Any component of the guide polynucleotide / Cas effector protein complex, the guide polynucleotide / Cas effector protein complex itself, and the polynucleotide-modified template and / or donor DNA can be introduced into a heterologous cell or organism by any method known in the art.

[0566] Recombinant constructs for cell transformation The disclosed guide polynucleotides, Cas endonucleases, polynucleotide-modified templates, donor DNA, guide polynucleotide / Cas endonuclease systems disclosed herein, and any combination thereof, optionally further comprising one or more polynucleotides of interest, can be introduced into cells, including but not limited to, human, non-human, animal, bacterial, fungal, insect, yeast, non-conventional yeast, and plant cells, as well as plants and seeds produced by the methods described herein.

[0567] Standard recombinant DNA and molecular cloning techniques used herein are well known in the art and are described in more detail in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory: Cold Spring Harbor, NY (1989). Transformation methods are well known to those skilled in the art and are described below.

[0568] Vectors and constructs include circular plasmids and linear polynucleotides, which contain a polynucleotide of interest and, optionally, other components, including linkers, adapters, regulatory elements, or analytical elements. In some embodiments, recognition and / or target sites can be contained within introns, coding sequences, 5'UTRs, 3'UTRs, and / or regulatory regions.

[0569] Components for the expression and utilization of novel CRISPR-Cas systems in prokaryotic and eukaryotic cells The present invention further provides expression constructs for expressing guide RNA / Cas systems capable of recognizing, binding to, and optionally nicking, unnicking, or cleaving all or part of a target sequence in prokaryotic or eukaryotic cells / organisms.

[0570] In one embodiment, an expression construct of the present disclosure comprises a promoter operably linked to a nucleotide sequence encoding a Cas gene (or an optimized plant comprising a Cas endonuclease gene as described herein) and a promoter operably linked to a guide RNA of the present disclosure. The promoter is capable of driving expression of the operably linked nucleotide sequence in a prokaryotic or eukaryotic cell / organism.

[0571] Nucleotide sequence modifications of the guide polynucleotide, VT domain, and / or CER domain may be selected from the group consisting of, but are not limited to, a 5' cap, a 3' polyadenylation tail, a riboswitch sequence, a stability control sequence, a sequence that forms a dsRNA duplex, a modification or sequence that targets the guide polynucleotide to a subcellular location, a modification or sequence that provides tracking, a modification or sequence that provides a binding site for a protein, locked nucleic acid (LNA), 5-methyl dC nucleotides, 2,6-diaminopurine nucleotides, 2'-fluoro A nucleotides, 2'-fluoro U nucleotides; 2'-O-methyl RNA nucleotides, phosphorothioate linkages, linkages to cholesterol molecules, linkages to polyethylene glycol molecules, linkages to spacer 18 molecules, 5' to 3' covalent linkages, or any combination thereof. These modifications may result in at least one additional advantageous characteristic selected from the group of altered or modulated stability, intracellular targeting, tracking, fluorescent labeling, binding sites for proteins or protein complexes, altered binding affinity for complementary target sequences, altered resistance to cellular degradation, and increased cell permeability.

[0572] A method for expressing RNA components, such as gRNAs, in eukaryotic cells for Cas9-mediated DNA targeting is to use an RNA polymerase III (Pol III) promoter, which allows transcription of RNAs with precisely defined, unmodified 5' and 3' ends (DiCarlo et al., Nucleic Acids Res. 41:4336-4343; Ma et al., Mol. Ther. Nucleic Acids 3:e161). This strategy has been successfully applied to cells of several different species, including maize and soybean (U.S. Patent Application Publication No. 2015 / 0082478, published March 19, 2015). A method for expressing RNA components without a 5' cap has been described (WO 2016 / 025131, published February 18, 2016).

[0573] Various methods and compositions can be used to obtain cells or organisms with a polynucleotide of interest inserted into a target site for a Cas endonuclease. Such methods can use homologous recombination (HR) to integrate the polynucleotide of interest into the target site. In one method described herein, the polynucleotide of interest is introduced into the cells of an organism via a donor DNA construct.

[0574] The donor DNA construct further comprises first and second regions of homology flanking the polynucleotide of interest, the first and second regions of homology of the donor DNA having homology to first and second genomic regions, respectively, present in or flanking the target site in the genome of the cell or organism.

[0575] Donor DNA can be ligated to a guide polynucleotide. The ligated donor DNA can enable co-localization of target and donor DNA, which is useful for genome editing, gene insertion, and target genome regulation, and can also be useful for targeting postmitotic cells, which are thought to have greatly reduced function of the endogenous HR machinery (Mali et al., 2013, Nature Methods Vol. 10:957-963).

[0576] The amount of homology or sequence identity shared by the target and donor polynucleotides can vary and can range from about 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, 450-900 bp, 500-600 bp, 500-750 bp, 600-800 bp, 750-800 bp, 800-900 bp, 900-1000 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250- The term "target site" encompasses the entire length and / or entire region of a target site, including all integers within the range of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 bp. The amount of homology can also be described in terms of percent sequence identity over the fully aligned length of two polynucleotides, including percent sequence identities of at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98% to 99%, 99%, 99% to 100%, or 100%. Sufficient homology includes any combination of polynucleotide length, overall percent sequence identity, and optionally conserved regions of consecutive nucleotides or local percent sequence identity; for example, sufficient homology can be described as a 75-150 bp region having at least 80% sequence identity to a region of the target locus.Sufficient homology can also be described by the predictive ability of two polynucleotides to specifically hybridize under high stringency conditions; see, e.g., Sambrook et al., (1989) Molecular Cloning: A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., Eds. (1994) Current Protocols, (Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.); and Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, (Elsevier, New York).

[0577] The structural similarity between a given genomic region and the corresponding homologous region found on the donor DNA can be any degree of sequence identity that allows homologous recombination to occur. For example, the amount of homology or sequence identity shared by a "homologous region" of the donor DNA and a "genomic region" of the organism's genome can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, such that the sequences undergo homologous recombination.

[0578] The homologous region on the donor DNA can have homology to any sequence adjacent to the target site. In some instances, the homologous region shares substantial sequence homology with the genomic sequence immediately adjacent to the target site, but it is recognized that the homologous region can be designed to have sufficient homology to regions that may be further 5' or 3' to the target site. The homologous region can also have homology to a fragment of the target site in addition to downstream genomic regions.

[0579] In one embodiment, the first homologous region further comprises a first fragment of the target site and the second homologous region comprises a second fragment of the target site, wherein the first and second fragments are different.

[0580] target polynucleotide Polynucleotides of interest are described in detail herein and include polynucleotides that reflect commercial markets and commercial market interests related to crop development. Crops of interest and markets will change, and as developing countries expand global markets, new crops and technologies will also emerge. Furthermore, as our understanding of agronomic traits and characteristics such as yield and heterosis increases, gene selection for genetic modification will change accordingly.

[0581] General categories of polynucleotides of interest include, for example, genes involved in signaling such as zinc fingers, genes involved in transduction such as kinases, and genes involved in housekeeping such as heat shock proteins. More specific polynucleotides of interest include, but are not limited to, genes involved in agronomic traits such as those affecting crop yield, grain quality, grain nutrient content, starch and carbohydrate quality and quantity, as well as kernel size, sucrose loading, protein quality and quantity, nitrogen fixation and / or utilization, fatty acid and oil composition, genes encoding proteins that confer resistance to abiotic stresses (such as those that confer resistance to drought, nitrogen, temperature, salinity, toxic metals or trace elements, or toxins such as pesticides and herbicides), genes encoding proteins that confer resistance to biotic stresses (such as attack by fungi, viruses, bacteria, insects and nematodes, and the development of diseases associated with these organisms), and the like.

[0582] Agronomically important traits such as oil, starch, and protein content can be genetically modified in addition to using traditional breeding methods. Modifications include increasing oleic acid, saturated and unsaturated oil content, increasing lysine and sulfur concentrations, providing essential amino acids, and modifying starch. Modifications of the phoridothionin protein are described in U.S. Patent Nos. 5,703,049, 5,885,801, 5,885,802, and 5,990,389.

[0583] The polynucleotide sequence of interest may encode a protein involved in conferring disease resistance or pest resistance. "Disease resistance" or "pest resistance" refers to the avoidance of harmful symptoms in a plant resulting from the interaction of the plant with a pathogen. Pest resistance genes may encode resistance to pests that inhibit high yields, such as cutworms, armyworms, or the European corn borer. Disease and insect resistance genes, such as lysozyme or cecropin for antibacterial protection, or proteins such as defensins, glucanases, or chitinases for antifungal protection, or Bacillus thuringiensis endotoxin for controlling nematodes or insects, protease inhibitors, collagenases, lectins, or glucosidases, are all examples of useful gene products. Genes encoding disease resistance traits include detoxification genes such as those for fumonisins (U.S. Pat. No. 5,792,931), avirulence (avr) genes, and disease resistance (R) genes (Jones et al. (1994) Science 266:789; Martin et al. (1993) Science 262:1432; and Mindrinos et al. (1994) Cell 78:1089).

[0584] Insect resistance genes can encode resistance to pests that inhibit high yields, such as cutworms, armyworms, European corn borers, etc. Examples of such genes include Bacillus thuringiensis virulence protein genes (U.S. Pat. Nos. 5,366,892; 5,747,450; 5,736,514; 5,723,756; 5,593,881; and Geiser et al. (1986) Gene 48:109).

[0585] Proteins obtained by expression of "herbicide tolerance proteins" or "nucleic acid molecules encoding herbicide tolerance" include proteins that confer the ability of cells to tolerate higher concentrations of herbicide than cells that do not express the protein, or to tolerate a particular concentration of herbicide for a longer period of time than cells that do not express the protein. Herbicide tolerance traits can be introduced into plants by genes encoding tolerance to herbicides that act by inhibiting the activity of acetolactate synthase (ALS, also known as acetohydroxyacid synthase, AHAS), particularly sulfonylurea (UK: sulfonylurea)-type herbicides, genes encoding tolerance to herbicides that act by inhibiting the activity of glutamine synthase, such as phosphinothricin or basta (e.g., the bar gene), genes encoding tolerance to glyphosate (e.g., the EPSP synthase gene and the GAT gene), genes encoding tolerance to HPPD inhibitors (e.g., the HPPD gene), or other such genes known in the art. See, e.g., U.S. Patent Nos. 7,626,077, 5,310,667, 5,866,775, 6,225,114, 6,248,876, 7,169,970, 6,867,293, and 9,187,762. The bar gene encodes resistance to the herbicide basta, the nptII gene encodes resistance to the antibiotics kanamycin and geneticin, and the ALS gene mutant encodes resistance to the herbicide chlorsulfuron.

[0586] Furthermore, the polynucleotide of interest may also contain an antisense sequence complementary to at least a portion of the messenger RNA (mRNA) for the target gene sequence of interest. The antisense nucleotide is constructed to hybridize with the corresponding mRNA. The antisense sequence may be modified as long as the sequence hybridizes with the corresponding mRNA and disrupts its expression. In this manner, antisense constructs having 70%, 80%, or 85% sequence identity to the corresponding antisense sequence can be used. Furthermore, portions of the antisense nucleotide can be used to disrupt the expression of the target gene. Generally, sequences of at least 50, 100, 200, or more nucleotides can be used.

[0587] Furthermore, a polynucleotide of interest can also be used in a sense orientation to suppress the expression of an endogenous gene in a plant. Methods for suppressing the expression of a gene in a plant using a polynucleotide in a sense orientation are known in the art. This method generally involves transforming a plant with a DNA construct containing a promoter that drives expression in the plant, operably linked to at least a portion of a nucleotide sequence corresponding to the transcription of the endogenous gene. Typically, such a nucleotide sequence has considerable sequence identity to the transcribed sequence of the endogenous gene, generally greater than about 65% sequence identity, greater than about 85% sequence identity, or greater than about 95% sequence identity. See U.S. Patent Nos. 5,283,184 and 5,034,323.

[0588] The polynucleotide of interest may also be a phenotypic marker. Phenotypic markers are screening or selection markers, including visual markers and selection markers, regardless of whether they are positive or negative selection markers. Any phenotypic marker can be used. In particular, selection or screening markers often contain a DNA segment that allows the identification of a molecule or cells containing this molecule under specific conditions, or the favorable or unfavorable selection of such molecules or cells. These markers can encode an activity, such as, but not limited to, the production of RNA, peptides, or proteins, or can provide binding sites for RNA, peptides, proteins, inorganic and organic compounds, or compositions.

[0589] Examples of selectable markers include, but are not limited to, DNA segments containing restriction enzyme sites; DNA segments encoding products that provide resistance to otherwise toxic compounds, including antibiotics such as spectinomycin, ampicillin, kanamycin, tetracycline, Basta, neomycin phosphotransferase II (NEO), and hygromycin phosphotransferase (HPT); DNA segments encoding products that are otherwise deficient in recipient cells (e.g., tRNA genes, auxotrophic markers); DNA segments encoding products that are easily identifiable (e.g., phenotypic markers such as β-galactosidase, GUS; fluorescent proteins such as green fluorescent protein (GFP), cyan (CFP), yellow (YFP), red (RFP) fluorescent protein, and cell surface proteins); generation of novel primer sites for PCR (e.g., juxtaposition of two DNA sequences not previously juxtaposed); inclusion of DNA sequences that are unaffected or have been acted upon by restriction endonucleases or other DNA-modifying enzymes, chemicals, etc.; and inclusion of DNA sequences required for specific modifications (e.g., methylation) that allow for their identification.

[0590] Additional selectable markers include genes that confer resistance to herbicidal compounds such as sulfonylureas, glufosinate ammonium, bromoxynil, imidazolinones, and 2,4-dichlorophenoxyacetate (2,4-D). See, for example, acetolactate synthase (ALS) for resistance to sulfonylureas, imidazolinones, triazolopyrimidine sulfonamides, pyrimidinyl salicylates, and sulfonylaminocarbonyltriazolinones (Shaner and Singh, 1997, Herbicide Activity: Toxicol Biochem Mol Biol 69-110); 5-enolpyruvylshikimate-3-phosphate (EPSPS) for glyphosate resistance (Saroha et al., 1998, J. Plant Biochemistry & Biotechnology Vol 7:65-72);

[0591] Polynucleotides of interest include genes that can be stacked or used in combination with other traits, such as, but not limited to, herbicide tolerance or any other trait described herein. Polynucleotides of interest and / or traits can be stacked together into composite trait loci, as described in U.S. Patent Application Publication No. 2013 / 0263324, published October 3, 2013, and WO 2013 / 112686, published August 1, 2013.

[0592] A polypeptide of interest includes any protein or polypeptide encoded by a polynucleotide of interest described herein.

[0593] Additionally, methods are provided for identifying at least one plant cell containing in its genome a polynucleotide of interest integrated at a target site. Various methods can be used to identify plant cells with insertion at or near the target site in the genome. Such methods can be considered direct analysis of the target sequence to detect alterations in the target sequence, including, but not limited to, PCR, sequencing, nuclease digestion, Southern blotting, and any combination thereof. See, for example, U.S. Patent Application Publication No. 2009 / 0133152, published May 21, 2009. The method also includes recovering a plant from the plant cell containing the polynucleotide of interest integrated into its genome. The plant can be sterile or fertile. Any polynucleotide of interest can be provided and confirmed to be capable of being integrated into the target site in the genome of the plant and expressed in the plant.

[0594] Optimization of sequences for expression in plants Methods for synthesizing plant-preferred genes are available in the art. See, for example, U.S. Patent Nos. 5,380,831 and 5,436,391, and Murray et al. (1989) Nucleic Acids Res. 17:477-498. Additional sequence modifications are known to enhance gene expression in plant hosts. These modifications include, for example, the removal of one or more sequences encoding spurious polyadenylation signals, the removal of one or more exon-intron splice site signals, the removal of one or more transposon-like repeats, and the removal of other well-characterized sequences that may be deleterious to gene expression. The GC content of the sequence can be adjusted to an average level for a given plant host (calculated with reference to known genes expressed in that host plant cell). When possible, the sequence is modified to avoid predicted hairpin secondary structures in one or more mRNAs. Thus, the "plant-optimized nucleotide sequence" of the present disclosure includes one or more such sequence modifications.

[0595] Expression elements Any polynucleotide encoding a Cas protein or other CRISPR system component disclosed herein can be operably linked to heterologous expression elements to facilitate transcription or regulation in a host cell. Such expression elements include, but are not limited to, promoters, leaders, introns, and terminators. An expression element may be "minimal," meaning a short sequence derived from a natural source that still functions as an expression regulator or modifier. Alternatively, an expression element may be "optimized," meaning that the polynucleotide sequence has been modified from its natural state to function with more desirable characteristics in a particular host cell (e.g., but not limited to, a bacterial promoter may be "corn-optimized" to improve its expression in corn plants). Alternatively, an expression element may be "synthetic," meaning that the expression element is designed in silico and synthesized for use in a host cell. Synthetic expression elements may be wholly synthetic or partially synthetic (including fragments of naturally occurring polynucleotide sequences).

[0596] Certain promoters have been shown to be able to induce RNA synthesis at a higher rate than others. These are called "strong promoters." Certain other promoters have been shown to induce RNA synthesis at high levels only in particular cell or tissue types; if a promoter induces RNA synthesis preferentially in certain tissues and at lower levels in other tissues, it is often called a "tissue-specific promoter" or "tissue-preferred promoter."

[0597] Plant promoters include promoters that can initiate transcription in plant cells. For a description of plant promoters, see Potenza et al., 2004, In vitro Cell Dev Biol 40:1-22; Porto et al., 2014, Molecular Biotechnology (2014), 56(1), 38-49.

[0598] Constitutive promoters include, for example, the core CaMV 35S promoter (Odell et al., (1985) Nature 313:810-2); rice actin (McElroy et al., (1990) Plant Cell 2:163-71); ubiquitin (Christensen et al., (1989) Plant Mol Biol 12:619-32); and the ALS promoter (U.S. Pat. No. 5,659,026).

[0599] Tissue-preferred promoters can be used to target enhanced expression in specific plant tissues. Examples of tissue-preferred promoters include those described in International Publication No. 2013 / 103367 published on July 11, 2013, Kawamata et al., (1997) Plant Cell Physiol 38:792-803; Hansen et al., (1997) Mol Gen Genet 254:337-43; Russell et al., (1997) Transgenic Res 6:157-68; Rinehart et al., (1996) Plant Physiol 112:1331-41; Van Camp et al., (1996) Plant Physiol 112:525-35; Canevascini et al., (1996) Plant Physiol 112:513-524; Lam, (1994) Results Probabilistic Cell Differentiation 20:181-96; and Guevara-Garcia et al., (1993) Plant J 4:495-505. Leaf-preferential promoters include, for example, Yamamoto et al., (1997) Plant J 12:255-65; Kwon et al., (1994) Plant Physiol 105:357-67; Yamamoto et al., (1994) Plant Cell Physiol 35:773-8; Gotor et al., (1993) Plant J 3:509-18;Orozco et al.,(1993)Plant Mol Biol 23:1129-38;Matsuoka et al.,(1993)Proc.Natl.Acad.Sci.USA90:9586-90;Simpson et al.,(1958)EMBO J 4:2723-9;Timko et al. al.,(1988)Nature 318:57-8.Root-preferred promoters include, for example, Hire et al., (1992) Plant Mol Biol 20:207-18 (soybean root-specific glutamine synthetase gene); Miao et al., (1991) Plant Cell 3:11-22 (cytosolic glutamine synthetase (GS)); Keller and Baumgartner, (1991) Plant Cell 3:1051-61 (root-specific regulatory element in the snap bean GRP 1.8 gene); Sanger et al., (1990) Plant Mol Biol 14:433-43 (root-specific promoter of A. tumefaciens mannopine synthase (MAS)); Bogusz et al., (1990) Plant Cell 2:633-41 (Parasponia andersenii) andersonii and Trema tomentosa; Leach and Aoyagi, (1991) Plant Sci 79:69-76 (A. rhizogenes rolC and rolD root-inducible genes); Teeri et al., (1989) EMBO J 8:343-50 (Agrobacterium wound-inducible TR1' and TR2' genes); VfENOD-GRP3 gene promoter (Kuster et al., (1995) Plant Mol Biol 29:759-72); and rolB promoter (Capana et al., (1994) Plant Mol Biol 25:681-91); phaseolin gene (Murai et al., (1983) Science 23:476-82; Sengopta-Gopalen et al. al., (1988) Proc. Natl. Acad. Sci. USA 82:3320-4. See also U.S. Patent Nos. 5,837,876, 5,750,386, 5,633,363, 5,459,252, 5,401,836, 5,110,732, and 5,023,179.

[0600] Seed-preferred promoters include both seed-specific promoters active during seed development and seed germination promoters active during seed germination. See Thompson et al. (1989) BioEssays 10:108. Seed-preferred promoters include, but are not limited to, Cim1 (cytokinin-inducible message); cZ19B1 (maize 19 kDa zein); and milps (myo-inositol-1-phosphate synthase); as well as those disclosed, for example, in International Publication WO 2000 / 011177, published March 2, 2000, and U.S. Patent No. 6,225,529. Seed-preferred promoters for dicotyledonous plants include, but are not limited to, bean-β-phaseolin, napin, β-conglycinin, soybean lectin, and cruciferin. Seed-preferred promoters for monocotyledons include, but are not limited to, maize 15 kDa zein, 22 kDa zein, 27 kDa gamma zein, waxy, shrunken 1, shrunken 2, globulin 1, oleosin, and nuc1. See also WO 2000 / 012733, published March 9, 2000, which discloses seed-preferred promoters from the END1 and END2 genes.

[0601] Chemically inducible (regulated) promoters can be used to regulate gene expression in prokaryotic and eukaryotic cells or organisms by application of exogenous chemical regulators. The promoters can be chemically inducible promoters, in which application of chemicals induces gene expression, or chemically repressible promoters, in which application of chemicals represses gene expression. Chemically inducible promoters include, but are not limited to, the maize In2-2 promoter, which is activated by benzenesulfonamide herbicide safeners (De Veylder et al., (1997) Plant Cell Physiol 38:568-77), the maize GST promoter (GST-II-27, WO 1993 / 001294, published January 21, 1993), which is activated by hydrophobic electrophilic compounds used as pre-emergence herbicides, and the tobacco PR-1a promoter, which is activated by salicylic acid (Ono et al., (2004) Biosci Biotechnol Biochem 68:803-7). Other chemically regulated promoters include steroid-responsive promoters (e.g., glucocorticoid-inducible promoters (see Schena et al., (1991) Proc. Natl. Acad. Sci. USA 88:10421-5; McNellis et al., (1998) Plant J 14:247-257); tetracycline-inducible and tetracycline-repressible promoters (Gatz et al., (1991) Mol Gen Genet 227:229-37; U.S. Pat. Nos. 5,814,618 and 5,789,156).

[0602] Pathogen-inducible promoters that are induced after pathogen infection include, but are not limited to, those that regulate the expression of PR proteins, SA proteins, beta-1,3-glucanase, chitinase, and the like.

[0603] An example of a stress-inducible promoter is the RD29A promoter (Kasuga et al. (1999) Nature Biotechnol. 17:287-91). Those skilled in the art are familiar with protocols for simulating stress conditions such as drought, osmotic stress, salt stress, and temperature stress, and for evaluating the stress tolerance of plants exposed to simulated or naturally occurring stress conditions.

[0604] Another example of an inducible promoter useful in plant cells is the ZmCAS1 promoter, described in U.S. Patent Application Publication No. 2013 / 0312137, published November 21, 2013.

[0605] New promoters of various types useful in plant cells are constantly being discovered, and many examples can be found in Okamuro and Goldberg (1989), The Biochemistry of Plants, Vol. 115, Stumpf and Conn, eds. (New York, NY: Academic Press), pp. 1-82.

[0606] Genome modification by novel CRISPR-Cas system components As described herein, guided Cas endonucleases can recognize and bind to DNA target sequences and introduce single- or double-stranded breaks. When single- or double-stranded breaks are introduced into DNA, the cell's DNA repair machinery is activated to repair the break. Error-prone DNA repair mechanisms can create mutations at the double-stranded break site. The most common repair mechanism for joining broken ends together is the non-homologous end joining (NHEJ) pathway (Bleuyard et al., (2006) DNA Repair 5:1-12). The structural integrity of chromosomes is typically maintained by repair, but deletions, insertions, or other rearrangements (such as chromosomal translocations) are also possible (Siebert and Puchta, 2002, Plant Cell 14:1121-31; Pacher et al., 2007, Genetics 175:21-9).

[0607] DNA double-strand breaks appear to be effective agents for stimulating the homologous recombination pathway (Puchta et al., (1995) Plant Mol Biol 28:281-92; Tzfira and White, (2005) Trends Biotechnol 23:567-9; Puchta, (2005) J Exp Bot 56:1-14). Using DNA-cleaving agents, a 2- to 9-fold increase in homologous recombination was observed between artificially constructed homologous DNA repeats in plants (Puchta et al., (1995) Plant Mol Biol 28:281-92). Experiments using linear DNA molecules in maize protoplasts demonstrated enhanced homologous recombination between plasmids (Lyznik et al., (1991) Mol Gen Genet 230:209-18).

[0608] Homologous recombination repair (HDR) is a mechanism for repairing double-stranded and single-stranded DNA breaks in cells. Homologous recombination repair includes homologous recombination (HR) and single-stranded annealing (SSA) (Lieber. 2010 Annu. Rev. Biochem. 79:181-211). The most common form of HDR is called homologous recombination (HR), which requires the longest sequence homology between the donor and acceptor DNA. Other forms of HDR include single-stranded annealing (SSA) and break-induced replication, which require shorter sequence homology than HR. Homologous recombination repair for nicking (single-stranded breaks) may occur by a different mechanism than HDR for double-stranded breaks (Davis and Maizels. PNAS (0027-8424), 111(10), p. E924-E932).

[0609] For example, modification of the genomes of prokaryotic and eukaryotic living cells or organisms by homologous recombination (HR) is a powerful tool in genetic engineering. Homologous recombination has been demonstrated in plants (Halfter et al., (1992) Mol Gen Genet 231:186-93) and insects (Dray and Gloor, 1997, Genetics 147:689-99). Homologous recombination has also been achieved in other organisms. For example, at least 150-200 bp of homology was required for homologous recombination in the protozoan parasite Leishmania (Papadopoulou and Dumas, (1997) Nucleic Acids Res 25:4278-86). In the filamentous fungus Aspergillus nidulans, gene replacement has been achieved using as little as 50 bp of contiguous homology (Chaveroche et al., (2000) Nucleic Acids Res 28:e97). Targeted gene replacement has also been demonstrated in the ciliate Tetrahymena thermophila (Gaertig et al., (1994) Nucleic Acids Res 22:5391-8). In mammals, homologous recombination has been most successful in mice using pluripotent embryonic stem cell (ES) lines that can be grown in culture, transformed, selected, and introduced into mouse embryos (Watson et al., 1992, Recombinant DNA, 2nd Ed., Scientific American Books distributed by W.H. Freeman & Co.).

[0610] gene targeting The guide polynucleotide / Cas system described herein can be used for gene targeting.

[0611] Generally, DNA targeting can be achieved by cleaving one or both strands at specific polynucleotide sequences in cells using a Cas protein in association with a suitable polynucleotide component. Upon induction of a single- or double-strand break in DNA, the cell's DNA repair machinery can be activated to repair the break by non-homologous end joining (NHEJ) or homology-directed repair (HDR) processes, leading to modification of the target site.

[0612] The length of the DNA sequence of the target site can vary, including, for example, target sites that are at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or 30 nucleotides in length or more. The target site can be palindromic, meaning that a sequence on one strand can be read in the opposite direction on the complementary strand. The nick / cleavage site can be within the target sequence, or the nick / cleavage site can be outside the target sequence. In another variation, the cleavage can occur at nucleotide positions directly opposite each other, resulting in a blunt-end cleavage, or in other cases, the nicks can be offset, resulting in a single-stranded overhang (which can be a 5' overhang or a 3' overhang), also known as a "sticky end." Active variants of genomic target sites can also be used. Such active variants can comprise at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a given target site, and the active variant retains biological activity and thus can be recognized and cleaved by a Cas endonuclease.

[0613] Assays for measuring single- or double-stranded cleavage of target sites by endonucleases are known in the art and generally measure the overall activity and specificity of the agent towards a DNA substrate containing the recognition site.

[0614] The targeting methods herein can be practiced, for example, such that two or more DNA target sites are targeted in the method. Such methods can optionally be characterized as multiplex methods. In certain embodiments, two, three, four, five, six, seven, eight, nine, ten, or more target sites can be targeted simultaneously. Multiplex methods are typically practiced by targeting methods herein that provide multiple different RNA components, each designed to guide the guide polynucleotide / Cas endonuclease complex to a unique DNA target site.

[0615] Gene editing The process of editing a genomic sequence to combine a DSB with a modified template generally involves introducing into a host cell a DSB inducer or a nucleic acid encoding the DSB inducer that can recognize a target sequence in a chromosomal sequence and introduce a DSB into the genomic sequence, and at least one modified polynucleotide template that contains at least one nucleotide modification compared to the nucleotide sequence to be edited. The modified polynucleotide template may further include nucleotide sequences flanking the at least one nucleotide modification. The flanking sequences are substantially homologous to the chromosomal region adjacent to the DSB. Genome editing using DSB inducers, such as Cas-gRNA complexes, is described, for example, in U.S. Patent Application Publication No. 2015 / 0082478, published March 19, 2015; International Publication No. WO 2015 / 026886, published February 26, 2015; International Publication No. WO 2016 / 007347, published January 14, 2016; and International Publication No. WO 2016 / 025131, published February 18, 2016.

[0616] Several uses of the guide RNA / Cas endonuclease system have been described (see, e.g., U.S. Patent Application Publication No. 2015 / 0082478A1, published March 19, 2015; WO 2015 / 026886, published February 26, 2015; and U.S. Patent Application Publication No. 2015 / 0059010, published February 26, 2015), including, but not limited to, altering or substituting a nucleotide sequence of interest (such as a regulatory element), inserting a polynucleotide of interest, gene knockout, gene knockin, altering a splice site and / or introducing an alternative splice site, altering a nucleotide sequence encoding a protein of interest, fusion of amino acids and / or proteins, and gene silencing by expressing an inverted repeat in a gene of interest.

[0617] Proteins can be modified in a variety of ways, including amino acid substitution, deletion, truncation, and insertion. Methods for such manipulations are generally known. For example, amino acid sequence variants of proteins can be prepared by mutations in DNA. Methods for mutagenesis and nucleotide sequence modification include, for example, Kunkel, (1985) Proc. Natl. Acad. Sci. USA 82:488-92; Kunkel et al., (1987) Meth Enzymol 154:367-82; U.S. Pat. No. 4,873,192; Walker and Gaastra, eds. (1983) Techniques in Molecular Biology (MacMillan Publishing Company, New York) and references cited therein. Guidance on amino acid substitutions unlikely to affect the biological activity of a protein can be found, for example, in the model Dayhoff et al. (1978) Atlas of Protein Sequence and Structure (Natl Biomed Res Found, Washington, DC). Conservative substitutions, such as exchanging one amino acid for another with similar properties, may be preferred. Conservative deletions, insertions, and amino acid substitutions are not expected to cause radical changes in the properties of the protein, and the effects of substitutions, deletions, insertions, or combinations thereof can be assessed in routine screening assays. Assays for double-strand break-inducing activity are known and generally measure the overall activity and specificity of an agent on a DNA substrate containing the target site.

[0618] Genome editing methods using Cas endonucleases and complexes comprising Cas endonucleases and guide polynucleotides are described herein. Following characterization of the guide RNA and PAM sequences, the endonucleases and associated CRISPR RNA (crRNA) components can be utilized to modify chromosomal DNA in other organisms, including plants. To facilitate optimal expression and nuclear localization (for eukaryotic cells), genes comprising the complexes can be optimized as described in International Publication No. WO 2016 / 186953, published November 24, 2016, and then delivered to cells as DNA expression cassettes using methods known in the art. The components necessary to comprise an active complex can also be delivered as RNA, with or without modifications that protect the RNA from degradation, i.e., as capped or uncapped mRNA (Zhang, Y. et al., 2016, Nat. Commun. 7:12617), or as a Cas protein-guided polynucleotide complex (WO 2017 / 070032, published April 27, 2017), or any combination thereof. Furthermore, one or more parts of the complex and the crRNA can be expressed from a DNA construct, while the other components can be delivered as RNA, with or without modifications that protect the RNA from degradation, i.e., as capped or uncapped mRNA (Zhang et al. 2016 Nat. Commun. 7:12617), or as a Cas protein-guided polynucleotide complex (WO 2017 / 070032, published April 27, 2017), or any combination thereof. To produce crRNA in vivo, tRNA-derived elements can also be used to recruit endogenous RNAses to cleave the crRNA transcript into a mature form that can guide the complex to its DNA target site, as described, for example, in International Publication No. WO 2017 / 105991, published June 22, 2017. Nickase complexes can be used individually or in concert to generate single or multiple DNA nicks in one or both DNA strands.Furthermore, the cleavage activity of Cas endonucleases can be inactivated by modifying key catalytic residues in their cleavage domains (Sinkunas, T et al., 2013, EMBO J. 32:385-394), resulting in RNA-guided helicases that can be used to enhance homology-directed repair, induce transcriptional activation, or remodel DNA local structures. Furthermore, the activities of both the Cas cleavage and helicase domains can be knocked out and used in combination with other DNA cleavage, DNA nicking, DNA binding, transcriptional activation, transcriptional repression, DNA remodeling, DNA deamination, DNA unwinding, DNA recombination enhancement, DNA integration, DNA inversion, and DNA repair agents.

[0619] The direction of transcription of the tracrRNA (if present) and other components of the CRISPR-Cas system (e.g., variable targeting domain, crRNA repeats, loops, anti-repeats) can be deduced as described in WO 2016 / 186946 published on November 24, 2016 and WO 2016 / 186953 published on November 24, 2016.

[0620] Once appropriate guide RNA requirements are established as described herein, the PAM preference for each new system disclosed herein can be investigated. If the cleavage complex results in degradation of the randomized PAM library, the complex can be converted into a nickase by disabling the ATPase-dependent helicase activity through mutagenesis of key residues or by assembling the reaction in the absence of ATP as previously described (Sinkunas, T. et al., 2013, EMBO J. 32:385-394). Two regions of PAM randomization, separated by two protospacer targets, can be used to generate double-stranded DNA breaks, which can be captured and sequenced to determine the PAM sequences that support cleavage by each complex.

[0621] In one embodiment, the present invention describes a method for modifying a target site in the genome of a cell, the method comprising introducing at least one PGEN described herein into a cell, and identifying at least one cell having a modification at the target, wherein the modification at the target site is selected from the group consisting of: (i) substitution of at least one nucleotide; (ii) deletion of at least one nucleotide; (iii) insertion of at least one nucleotide; chemical modification of at least one nucleotide; and (v) any combination of (i) to (iv).

[0622] The nucleotide to be edited can be located within or outside the target site recognized and cleaved by the Cas endonuclease. In one embodiment, the modification of the at least one nucleotide is not an alteration at the target site recognized and cleaved by the Cas endonuclease. In another embodiment, there are at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 900, or 1000 nucleotides between the at least one nucleotide to be edited and the genomic target site.

[0623] Knockouts can be generated by indels (insertion or deletion of nucleotide bases in the targeted DNA sequence by NHEJ) or by specific removal of sequences that reduce or completely abolish the function of sequences at or near the target site.

[0624] Guide polynucleotide / Cas endonuclease-induced targeted mutations can occur within nucleotide sequences located within or outside the genomic target site recognized and cleaved by the Cas endonuclease.

[0625] The method of editing a nucleotide sequence in the genome of a cell can be a method that does not use an exogenous selectable marker by restoring function to a non-functional gene product.

[0626] In one embodiment, the present invention describes a method for modifying a target site in the genome of a cell. The method comprises introducing at least one PGEN described herein and at least one donor DNA into a cell, wherein the donor DNA comprises a target polynucleotide, and optionally further comprises identifying at least one cell that has the target polynucleotide integrated at or near the target site.

[0627] In one aspect, the methods disclosed herein may use homologous recombination (HR) to provide for integration of a polynucleotide of interest into a target site.

[0628] Various methods and compositions can be used to generate cells or organisms having a polynucleotide of interest inserted into a target site through the activity of the CRISPR-Cas system components described herein. In one method described herein, the polynucleotide of interest is introduced into an organism's cells via a donor DNA construct. As used herein, "donor DNA" refers to a DNA construct containing the polynucleotide of interest to be inserted into the target site of the Cas endonuclease. The donor DNA construct further comprises first and second homologous regions flanking the polynucleotide of interest. The first and second homologous regions of the donor DNA are homologous to first and second genomic regions, respectively, present in or flanking the target site in the genome of the cell or organism.

[0629] Donor DNA can be linked to a guide polynucleotide. The linked donor DNA can enable co-localization of target and donor DNA, which is useful for genome editing, gene insertion, and target genome regulation, and can also be useful for targeting postmitotic cells, which are thought to have greatly reduced function of the endogenous HR machinery (Mali et al., 2013, Nature Methods Vol. 10:957-963).

[0630] The amount of homology or sequence identity shared by the target and donor polynucleotides can vary and can range from about 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, 450-900 bp, 500-600 bp, 500-750 bp, 600-800 bp, 750-800 bp, 800-900 bp, 900-1000 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250- The term "target site" encompasses the entire length and / or entire region of a target site, including all integers within the range of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 bp. The amount of homology can also be described in terms of percent sequence identity over the fully aligned length of two polynucleotides, including percent sequence identity of at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. Sufficient homology includes any combination of polynucleotide length, overall percent sequence identity, and optionally conserved regions of consecutive nucleotides or local percent sequence identity; for example, sufficient homology can be described as a 75-150 bp region having at least 80% sequence identity to a region of the target locus.Sufficient homology can also be described by the predictive ability of two polynucleotides to specifically hybridize under high stringency conditions; see, e.g., Sambrook et al., (1989) Molecular Cloning: A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., Eds. (1994) Current Protocols, (Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.); and Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, (Elsevier, New York).

[0631] Episomal DNA molecules can also ligate into double-strand breaks, for example, integrating T-DNA into chromosomal double-strand breaks (Chilton and Que, (2003) Plant Physiol 133:956-65; Salomon and Puchta, (1998) EMBO J 17:6086-95). When the sequence surrounding the double-strand break is modified, for example, by exonuclease activity involved in double-strand break maturation, the gene conversion pathway can regain its original structure if homologous sequences are available, such as sister chromatids in non-dividing somatic cells or after DNA replication (Molinier et al., (2004) Plant Cell 16:342-52). Ectopic and / or epigenetic DNA sequences can also serve as DNA repair templates for homologous recombination (Puchta, (1999) Genetics 152:1173-81).

[0632] In one embodiment, the present disclosure includes a method for editing a nucleotide sequence in the genome of a cell, the method comprising introducing at least one PGEN described herein into a polynucleotide modified template, wherein the polynucleotide modified template comprises at least one nucleotide modification of the nucleotide sequence, and optionally further comprising selecting at least one cell comprising the edited nucleotide sequence.

[0633] The guide polynucleotide / Cas endonuclease system can be used in combination with at least one polynucleotide modification template to enable editing (modification) of a genomic nucleotide sequence of interest (see also U.S. Patent Application Publication No. 2015 / 0082478, published March 19, 2015, and International Publication No. WO 2015 / 026886, published February 26, 2015).

[0634] Polynucleotides and / or traits of interest can be stacked together at complex trait loci, as described in International Publication No. WO 2012 / 129373, published September 27, 2012, and International Publication No. WO 2013 / 112686, published August 1, 2013. The guide polynucleotide / Cas9 endonuclease system described herein provides a system for efficiently generating double-stranded breaks, allowing for the stacking of traits at complex trait loci.

[0635] The guide polynucleotide / Cas system described herein for mediating gene targeting can be used in methods for directing heterologous gene insertion and / or generating complex trait loci containing multiple heterologous genes, similar to those disclosed in International Publication No. WO 2012 / 129373, published September 27, 2012, where the guide polynucleotide / Cas system disclosed herein replaces the use of double-strand break-inducing agents to introduce genes of interest. Transgenes can be propagated as a single locus by inserting independent transgenes within 0.1, 0.2, 0.3, 0.4, 0.5, 1.0, 2, or 5 centimorgans (cM) of each other (see, e.g., U.S. Patent Application Publication No. 2013 / 0263324, published October 3, 2013, or WO 2012 / 129373, published March 14, 2013). After plants containing the transgenes are selected, plants containing (at least) one of the transgenes can be crossed to form an F1 containing both transgenes. Of the progeny from these F1s (F2 or BC1), 1 / 500 will have the two different transgenes recombined on the same chromosome. This composite locus can then be bred with both transgenes as a single locus. This process can be repeated to accumulate any number of desired traits.

[0636] Further uses of guide RNA / Cas endonuclease systems have been described (e.g., U.S. Patent Application Publication No. 2015 / 0082478 published March 19, 2015; WO 2015 / 026886 published February 26, 2015; U.S. Patent Application Publication No. 2015 / 0059010 published February 26, 2015; WO 2016 / 007347 published January 14, 2016; and PCT application Ser. No. 2016 / 007347 published February 18, 2016). See WO 2016 / 025131), which include, but are not limited to, altering or substituting a nucleotide sequence of interest (such as a regulatory element), inserting a polynucleotide of interest, gene knockout, gene knockin, altering a splice site and / or introducing an alternative splice site, altering a nucleotide sequence encoding a protein of interest, fusion of amino acids and / or proteins, and gene silencing by expressing an inverted repeat in a gene of interest.

[0637] The characteristics obtained from the gene editing compositions and methods described herein can be evaluated. Chromosomal intervals correlated with the desired phenotype or trait can be identified. Various methods well known in the art can be used to identify chromosomal intervals. The boundaries of such chromosomal intervals are drawn to encompass markers that may be linked to genes controlling the desired trait. In other words, chromosomal intervals are defined so that any marker present within the interval (including the terminal markers that define the boundaries of the interval) can be used as a marker for a specific trait. In one embodiment, a chromosomal interval contains at least one QTL, and may actually contain more than one QTL. When multiple QTLs are very close to each other within the same interval, the association of a particular marker with a specific QTL can be unclear, since one marker is linked to more than one QTL. Conversely, for example, if two very close markers cosegregate with a desired phenotypic trait, it is often unclear whether these markers identify the same QTL or two different QTLs. The term "quantitative trait locus" or "QTL" refers to a region of DNA associated with differential expression of a quantitative phenotypic trait in at least one genetic background, e.g., at least one breeding population. A QTL region encompasses or is closely linked to a gene that influences the trait in question. An "allele of a QTL" can include multiple genes or other genetic factors within a contiguous genomic region or linkage group, such as a haplotype. An allele of a QTL can refer to a haplotype within a specific window, which is a contiguous genomic region that can be defined and tracked by a set of one or more polymorphic markers. A haplotype can be defined by the unique fingerprint of the allele at each marker position within a specific window.

[0638] Introduction of CRISPR-Cas system components into cells The methods and compositions disclosed herein do not rely on a particular method for introducing a sequence into an organism or cell, but simply introduce a polynucleotide or polypeptide into at least one cell of the organism. Introduction includes reference to the incorporation of a nucleic acid into a eukaryotic or prokaryotic cell, where the nucleic acid may be incorporated into the genome of the cell, and includes reference to the transient (direct) provision of a nucleic acid, protein, or polynucleotide-protein complex (PGEN, RGEN) to a cell.

[0639] Methods for introducing polynucleotides or polypeptides or polynucleotide-protein complexes into cells or organisms are known in the art and include, but are not limited to, microinjection, electroporation, stable transformation, transient transformation, ballistic particle acceleration (particle bombardment), whisker-mediated transformation, Agrobacterium-mediated transformation, direct gene transfer, virus-mediated transfer, transfection, transduction, cell-penetrating peptides, mesoporous silica nanoparticle (MSN)-mediated direct protein delivery, topical application, sexual crossing, sexual propagation, and any combination thereof.

[0640] For example, guide polynucleotides (guide RNA, cr nucleotides + tracr nucleotides, guide DNA and / or guide RNA-DNA molecules) can be directly (transiently) introduced into cells as single-stranded or double-stranded polynucleotide molecules. Guide RNA (or crRNA + tracrRNA) can also be indirectly introduced into cells by introducing a recombinant DNA molecule comprising a heterologous nucleic acid fragment encoding guide RNA (or crRNA + tracrRNA) operably linked to a specific promoter capable of transcribing guide RNA (crRNA + tracrRNA molecule) in said cells. A particular promoter can be, but is not limited to, an RNA polymerase III promoter, which allows transcription of RNA with strictly defined, unmodified 5' and 3' ends (Ma et al., 2014, Mol. Ther. Nucleic Acids 3:e161; DiCarlo et al., 2013, Nucleic Acids Res. 41:4336-4343; WO 2015 / 026887, published February 26, 2015). Any promoter capable of transcribing the guide RNA within the cell can be used, including heat shock / heat-inducible promoters operably linked to the nucleotide sequence encoding the guide RNA.

[0641] Plant cells differ from animal cells (such as human cells), fungal cells (such as yeast cells), and protoplasts, for example, in that plant cells contain a plant cell wall that can act as a barrier to the delivery of components.

[0642] Delivery of the Cas endonuclease, and / or guide RNA, and / or ribonucleoprotein complex, and / or polynucleotide encoding any one or more of the foregoing into plant cells can be achieved by methods known in the art, including, but not limited to, Rhizobiales-mediated transformation (e.g., Agrobacterium, Ochrobactrum), particle-mediated delivery (particle bombardment), polyethylene glycol (PEG)-mediated transfection (e.g., into protoplasts), electroporation, cell-penetrating peptides, or mesoporous silica nanoparticle (MSN)-mediated direct protein delivery.

[0643] Cas endonucleases, such as those described herein, can be introduced into cells by directly introducing the Cas polypeptide itself (referred to as direct Cas endonuclease delivery), mRNA encoding the Cas protein, and / or the guide polynucleotide / Cas endonuclease complex itself using any method known in the art. Cas endonucleases can also be introduced into cells indirectly by introducing a recombinant DNA molecule encoding the Cas endonuclease. The endonuclease can be transiently introduced into cells or integrated into the genome of the host cell using any method known in the art. Cellular uptake of the endonuclease and / or guide polynucleotide can be facilitated using a cell-penetrating peptide (CPP), as described in International Publication WO 2016 / 073433, published May 12, 2016. Any promoter capable of expressing the Cas endonuclease in the cell can be used, including heat shock / heat inducible promoters operably linked to the nucleotide sequence encoding the Cas endonuclease.

[0644] Direct delivery of polynucleotide-modified templates into plant cells can be achieved by particle-mediated delivery, and any other direct delivery method such as, but not limited to, polyethylene glycol (PEG)-mediated transfection into protoplasts, whisker-mediated transformation, electroporation, particle bombardment, cell-penetrating peptides, or mesoporous silica nanoparticle (MSN)-mediated direct protein delivery can be successfully used to deliver polynucleotide-modified templates in eukaryotic cells such as plant cells.

[0645] Donor DNA can be introduced by any means known in the art. Donor DNA can be provided by any transformation method known in the art, including, for example, Agrobacterium-mediated transformation or biolistic particle bombardment. Donor DNA may be transiently present in the cell or could be introduced by a viral replicon. In the presence of the Cas endonuclease and the target site, the donor DNA is inserted into the transformed plant genome.

[0646] Direct delivery of any one of the inducible Cas system components can be accompanied by direct delivery (co-delivery) of other mRNAs that can facilitate enrichment and / or visualization of cells that receive the guide polynucleotide / Cas endonuclease complex components. For example, direct co-delivery of the guide polynucleotide / Cas endonuclease components (and / or the guide polynucleotide / Cas endonuclease complex itself) with mRNAs encoding phenotypic markers (such as, but not limited to, transcriptional activators such as CRC (Bruce et al. 2000 The Plant Cell 12:65-79)) can restore function to the non-functional gene product, allowing for selection and enrichment of cells without the use of an exogenous selection marker, as described in International Publication WO 2017 / 070032, published April 27, 2017.

[0647] Introduction of the guide RNA / Cas endonuclease complex (representing the cleavage-ready complex described herein) into a cell includes introducing the individual components of the complex into the cell, separately or collectively, directly (direct delivery of the guide RNA and the Cas endonuclease protein and protein subunit, or functional fragments thereof) or by recombinant constructs expressing the components (guide RNA, Cas endonuclease, protein subunit, or functional fragments thereof). Introduction of the guide RNA / Cas endonuclease complex (RGEN) into a cell includes introducing the guide RNA / Cas endonuclease complex into the cell as a ribonucleotide-protein. The ribonucleotide-protein can be assembled before introduction into the cell, as described herein. The components that make up the guide RNA / Cas endonuclease ribonucleotide protein (at least one Cas endonuclease, at least one guide RNA, at least one protein subunit) can be assembled in vitro or by any means known in the art before being introduced into a cell (which will be targeted for genome modification as described herein).

[0648] Direct delivery of RGEN ribonucleoprotein allows genome editing at the target site of the cell genome, with the complex then rapidly degrading, making its presence in the cell transient. This transient presence of the RGEN complex can lead to reduced off-target effects. In contrast, delivery of RGEN components (guide RNA, Cas9 endonuclease) via plasmid DNA sequences results in sustained expression of RGEN from these plasmids, which can increase off-target effects (Cradick, TJ et al. (2013) Nucleic Acids Res 41:9584-9592; Fu, Y et al. (2014) Nat. Biotechnol. 31:822-826).

[0649] Direct delivery can be achieved by combining any one of the components (e.g., at least one guide RNA, at least one Cas protein, and optionally one additional protein) of the guide RNA / Cas endonuclease complex (RGEN) (representing the cleavage-ready complex described herein) with a delivery matrix comprising microparticles (such as, but not limited to, gold particles, tungsten particles, and silicon carbide whisker particles) (see also International Publication No. 2017 / 070032, published April 27, 2017). The delivery matrix can include any one of the components, such as the Cas endonuclease, attached to a solid matrix (e.g., a bombardment particle).

[0650] In one aspect, the guide polynucleotide / Cas endonuclease complex is a complex in which the guide RNA and the Cas endonuclease protein that form the guide RNA / Cas endonuclease complex are introduced into a cell as RNA and protein, respectively.

[0651] In one aspect, a guide polynucleotide / Cas endonuclease complex is one in which the guide RNA and Cas endonuclease protein that form a guide RNA / Cas endonuclease complex, and at least one protein subunit of the complex, are introduced into a cell as RNA and protein, respectively.

[0652] In one aspect, the guide polynucleotide / Cas endonuclease complex is a complex in which the guide RNA and the Cas endonuclease protein that form the guide RNA and Cas endonuclease complex (cleavage-ready complex), and at least one protein subunit of the complex, are pre-assembled in vitro and introduced into a cell as a ribonucleotide-protein complex.

[0653] Protocols for introducing polynucleotides, polypeptides, or polynucleotide-protein complexes (PGENs, RGENs) into eukaryotic cells, such as plants or plant cells, are known and include microinjection (Crossway et al., (1986) Biotechniques 4:320-34 and U.S. Pat. No. 6,300,543), meristem transformation (U.S. Pat. No. 5,736,369), electroporation (Riggs et al., (1986) Proc. Natl. Acad. Sci. USA 83:5602-6, Agrobacterium-mediated transformation (U.S. Pat. Nos. 5,563,055 and 5,981,840), whisker-mediated transformation (Ainley et al. 2013, Plant Biotechnology Journal 11:1126-1134; Shaheen A. and M. Arshad 2011 Properties and Applications of Silicon Carbide (2011), 345-358 Editor(s): Gerhardt, Rosario. Publisher: InTech, Rijeka, Croatia. CODEN: 69PQBP; ISBN: 978-953-307-201-2), direct gene transfer (Paszkowski et al., (1984) EMBO J 3:2717-22), and ballistic particle acceleration (U.S. Pat. Nos. 4,945,050; 5,879,918; 5,886,244; 5,932,782; Tomes et al., (1995) "Direct DNA Transfer into Intact Plant Cells via Microprojectile Bombardment" in Plant Cell, Tissue, and Organ Culture: Fundamental Methods, ed. Gamborg & Phillips (Springer-Verlag, Berlin); McCabe et al., (1988) Biotechnology 6:923-6; Weissinger et al.,(1988) Ann Rev Genet 22:421-77; Sanford et al.,(1987) Particulate Science and Technology 5:27-37 (onion); Christou et al.,(1988) Plant Physiol 87:671-4 (soybean); Finer and McMullen,(1991) In vitro Cell Dev Biol 27P:175-82 (soybean); Singh et al.,(1998) Theor Appl Genet 96:319-24 (soybean); Datta et al.,(1990) Biotechnology 8:736-40 (rice); Klein et al.,(1988) Proc. Natl. Acad. Sci. USA 85:4305-9 (corn); Klein et al.,(1988) Biotechnology 6:559-63 (maize); U.S. Patent Nos. 5,240,855; 5,322,783 and 5,324,646; Klein et al., (1988) Plant Physiol 91:440-4 (maize); Fromm et al., (1990) Biotechnology 8:833-9 (maize); Hooykaas-Van Slogteren et al., (1984) Nature 311:763-4; U.S. Patent No. 5,736,369 (cereals); Bytebier et al., (1987) Proc. Natl. Acad. Sci. USA 84:5345-9 (Liliaceae); De Wet et al., (1985) in The Experimental Manipulation of Ovule Tissues, ed. Chapman et al. al., (Longman, New York), pp. 197-209 (pollen); Kaeppler et al., (1990) Plant Cell Rep 9:415-8) and Kaeppler et al., (1992) Theor Appl Genet 84:560-6 (whisker-mediated transformation); D'Halluin et al.Li et al., (1992) Plant Cell 4:1495-505 (electroporation); Li et al., (1993) Plant Cell Rep 12:250-5; Christou and Ford (1995) Annals Botany 75:407-13 (rice) and Osjoda et al., (1996) Nat Biotechnol 14:745-50 (maize via Agrobacterium tumefaciens).

[0654] Alternatively, polynucleotides can be introduced into plants or plant cells by contacting the cells or organisms with a virus or viral nucleic acid. Generally, such methods involve incorporating the polynucleotide into a viral DNA or RNA molecule. In some embodiments, a polypeptide of interest may be first synthesized as part of a viral polyprotein, which may then be proteolytically processed in vivo or in vitro to produce the desired recombinant protein. Methods for introducing polynucleotides, including viral DNA or RNA molecules, into plants and expressing the encoded proteins therein are known; see, e.g., U.S. Patent Nos. 5,889,191, 5,889,190, 5,866,785, 5,589,367, and 5,316,931.

[0655] Polynucleotides or recombinant DNA constructs can be provided to or introduced into prokaryotic and eukaryotic cells or organisms using a variety of transient transformation methods, including, but not limited to, direct introduction of polynucleotide constructs into plants.

[0656] Nucleic acids and proteins can be provided to cells in any manner, including using molecules that facilitate uptake of any or all components (proteins and / or nucleic acids) of the inducible Cas system, such as cell-penetrating peptides and nanocarriers. See also U.S. Patent Application Publication No. 2011 / 0035836, published February 10, 2011, and European Patent Application Publication No. 2821486A1, published January 7, 2015.

[0657] Other methods for introducing polynucleotides into prokaryotic and eukaryotic cells or organisms or plant parts can be used, such as plastid transformation and methods for introducing polynucleotides into tissues from seedlings or mature seeds.

[0658] Stable transformation is intended to mean that a nucleotide construct introduced into an organism is integrated into the genome of the organism and can be inherited by its progeny. Transient transformation is intended to mean that a polynucleotide is introduced into an organism but is not integrated into the genome of this organism, or that a polypeptide is introduced into an organism. Transient transformation means that the introduced composition is only expressed or present temporarily in the organism.

[0659] A variety of methods are available for identifying those cells with modified genomes at or near the target site without using a screening marker phenotype. Such methods can be considered as direct analysis of the target sequence to detect alterations in the target sequence, and include, but are not limited to, PCR, sequencing, nuclease digestion, Southern blotting, and any combination thereof.

[0660] Cells and plants The polynucleotides and polypeptides of the present disclosure can be introduced into cells, including but not limited to, human, non-human, animal, mammalian, bacterial, fungal, insect, yeast, non-conventional yeast, and plant cells, as well as plants and seeds produced by the methods described herein. Monocotyledonous and dicotyledonous plants, and any plant containing plant elements, can be used with the compositions and methods described herein.

[0661] Examples of monocotyledonous plants that can be used include, but are not limited to, maize (Zea mays), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), barnyard millet (e.g., pearl millet, Pennisetum glaucum), common millet (Panicum miliaceum), foxtail millet (Setaria italica), finger millet (Eleusine coracana), wheat (Triticum species, e.g., Triticum aestivum), and the like. aestivum, Triticum monococcum), sugarcane (Saccharum spp.), oats (Avena), barley (Hordeum), switchgrass (Panicum virgatum), pineapple (Ananas comosus), banana (Musa spp.), palms, ornamentals, turfgrass, and other grasses.

[0662] Examples of dicotyledonous plants that can be used include, but are not limited to, soybean (Glycine max), Brassica species (e.g., but not limited to, rapeseed or canola) (Brassica napus, B. campestris, Brassica rapa, Brassica juncea), alfalfa (Medicago sativa), tobacco (Nicotiana tabacum), Arabidopsis (Arabidopsis thaliana), sunflower (Helianthus annuus), cotton (Gossypium arboreum), and the like. arboreum, Gossypium barbadense), and peanut (Arachis hypogaea), tomato (Solanum lycopersicum), and potato (Solanum tuberosum).

[0663] Additional plants that can be used include safflower (Carthamus tinctorius), sweet potato (Ipomoea batatus), cassava (Manihot esculenta), coffee (Coffea spp.), coconut (Cocos nucifera), citrus trees (Citrus spp.), cocoa (Theobroma cacao), tea plants (Camellia sinensis), bananas (Musa spp.), avocado (Persea americana), fig (Ficus casica), guava (Psidium guajava), and others. guajava), mango (Mangifera indica), olive (Olea europaea), papaya (Carica papaya), cashew (Anacardium occidentale), macadamia (Macadamia integrifolia), almond (Prunus amygdalus), sugar beet (Beta vulgaris), vegetables, ornamental plants, and conifers.

[0664] Vegetables that can be used include tomato (Lycopersicon esculentum), lettuce (e.g., Lactuca sativa), green beans (Phaseolus vulgaris), lima beans (Phaseolus limensis), peas (Lathyrus spp.), and members of the genus Cucumis, such as cucumber (C. sativus), cantaloupe (C. cantalupensis), and muskmelon (C. melo). Ornamental plants include azaleas (Rhododendron spp.), hydrangeas (Macrophylla hydrangea), hibiscus (Hibiscus rosasanensis), roses (Rosa spp.), tulips (Tulipa spp.), daffodils (Narcissus spp.), petunias (Petunia hybrida), carnations (Dianthus caryophyllus), poinsettias (Euphorbia pulcherrima), and chrysanthemums.

[0665] Conifers that can be used include pines, such as loblolly pine (Pinus taeda), slash pine (Pinus elliotii), ponderosa pine (Pinus ponderosa), lodgepole pine (Pinus contorta), and Monterey pine (Pinus radiata); Douglas-fir (Pseudotsuga menziesii); Western hemlock (Tsuga canadensis); Sitka spruce (Picea glauca); Redwood (Sequoia sempervirens); and Norway spruce (Abies amabilis). fir trees, such as Japanese cedar (Thuja plicata) and Alaska yellow cedar (Chamaecyparis nootkatensis); and cedar trees, such as Japanese red cedar (Thuja plicata) and Alaska yellow cedar (Chamaecyparis nootkatensis).

[0666] In certain embodiments of the present disclosure, a fertile plant is one that produces viable male and female gametes and is a self-fertile plant. Such a self-fertile plant can produce progeny plants without the contribution of gametes or genetic material contained therein from other plants. Other embodiments of the present disclosure may use plants that are not self-fertile because the plants do not produce viable or otherwise fertilizable male gametes or female gametes, or both.

[0667] The present disclosure finds use in breeding plants that contain one or more introduced traits.

[0668] A non-limiting example of how two traits can be stacked in the genome, for example, at a genetic distance of 5 cM from each other, is described as follows: A first plant containing a first transgenic target site integrated at a first DSB target site within a genomic window and lacking a first genomic locus of interest is crossed with a second transgenic plant containing a genomic locus of interest at a different genomic insertion site within the genomic window. The second plant does not contain the first transgenic target site. Approximately 5% of the plant progeny from this cross will have both the first transgenic target site integrated at the first DSB target site and the first genomic locus of interest integrated at a different genomic insertion site within the genomic window. The progeny plant containing both sites within the specified genomic window can be further crossed with a third transgenic plant containing a second transgenic target site integrated at a second DSB target site and / or a second genomic locus of interest within the specified genomic window, but lacking the first transgenic target site and the first genomic locus of interest. Progeny are then selected that have the first transgenic target site, the first genomic locus of interest, and the second genomic locus of interest integrated at a different genomic insertion site within the genomic window. Using this method, transgenic plants can be created that contain complex trait loci with at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, or more transgenic target sites integrated at the DSB target site and / or complex trait loci with the genomic loci integrated at different sites within the genomic window. In this manner, a variety of complex trait loci can be generated.

[0669] Cells and animals The polynucleotides and polypeptides of the present disclosure can be introduced into animal cells, including, but not limited to, organisms from phyla including chordates, arthropods, mollusks, annelids, cnidarians, or echinoderms; or from classes including mammals, insects, birds, amphibians, reptiles, or fish. In some embodiments, the animal is a human, mouse, C. elegans, rat, fruit fly (Drosophila spp.), zebrafish, chicken, dog, cat, guinea pig, hamster, chicken, hen, killifish, sea lamprey, pufferfish, tree frog (e.g., Xenopus spp.), monkey, or chimpanzee. Specific cell types contemplated include haploid cells, diploid cells, germ cells, neurons, muscle cells, endocrine or exocrine cells, epithelial cells, muscle cells, tumor cells, embryonic cells, hematopoietic cells, bone cells, germ cells, somatic cells, stem cells, pluripotent stem cells, induced pluripotent stem cells, progenitor cells, meiotic cells, and mitotic cells. In some embodiments, multiple cells from an organism may be used.

[0670] The disclosed novel Cas9 orthologs can be used to edit the genome of animal cells in a variety of ways. In one embodiment, it may be desirable to delete one or more nucleotides. In another embodiment, it may be desirable to insert one or more nucleotides. In one embodiment, it may be desirable to substitute one or more nucleotides. In another embodiment, it may be desirable to modify one or more nucleotides by covalent or non-covalent interaction with another atom or molecule.

[0671] Genome modification with Cas9 orthologs can be used to effect genotypic and / or phenotypic changes in target organisms. Such changes are preferably related to the improvement of a desired phenotype or physiologically significant trait, the correction of an endogenous defect, or the expression of some type of expression marker. In some embodiments, the desired phenotype or physiologically significant trait is related to the overall health, fitness, or reproductive potential of the animal, the animal's ecological fitness, or the animal's relationship or interaction with other organisms in its environment. In some embodiments, the phenotype or physiologically significant trait of interest is selected from the following group: improvement in overall health, reversal of disease, amelioration of disease, stabilization of disease, prevention of disease, treatment of parasitic infection, treatment of viral infection, treatment of retroviral infection, treatment of bacterial infection, treatment of neurological disorders (e.g., but not limited to, multiple sclerosis), correction of an intrinsic genetic defect (e.g., but not limited to, metabolic disease, achondroplasia, alpha 1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Barth syndrome, breast cancer, Charcot-Marie-Tooth disease, colon cancer, cricket cat syndrome, Crohn's disease, cystic fibrosis, Dercum's disease, Down syndrome, Duane's syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter's syndrome, Marfan's syndrome, myotonic dystrophy, neurofibromatosis, Noonan's syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland's syndrome, porphyria, progeria, prostate cancer, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, skin cancer, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner's syndrome, palatocardiofacial syndrome, WAGR syndrome, and Wilson's disease), treatment of congenital immune disorders (such as, but not limited to, immunoglobulin subclass deficiencies), treatment of acquired immune disorders (such as, but not limited to, AIDS and other HIV-related disorders), treatment of cancer, and treatment of diseases including rare or "orphan" conditions for which there are no other effective treatment options.

[0672] Cells genetically engineered using the compositions or methods disclosed herein can be transplanted into a subject for purposes such as gene therapy, e.g., to treat disease, or as antiviral, antipathogen, or anticancer therapeutics, for the production of genetically engineered organisms in agriculture, or for biological research.

[0673] In vitro detection, binding, and modification of polynucleotides The compositions disclosed herein may further be used in in vitro methods in some embodiments involving isolated polynucleotide sequences. The isolated polynucleotide sequences may contain one or more target sequences for modification. In some embodiments, the isolated polynucleotide sequences may be genomic DNA, PCR products, or synthesized oligonucleotides.

[0674] composition The modification of the target sequence can be in the form of nucleotide insertion, nucleotide deletion, nucleotide substitution, addition of an atomic molecule to an existing nucleotide, nucleotide modification, or attachment of a heterologous polynucleotide or polypeptide to the target sequence. The insertion of one or more nucleotides can be achieved by including a donor polynucleotide in the reaction mixture, which is inserted into the double-stranded break generated by the Cas-alpha ortholog polypeptide. The insertion can be achieved by non-homologous end joining or homologous recombination.

[0675] In one embodiment, the sequence of the target polynucleotide is known prior to modification and is compared to the sequence of the polynucleotide resulting from treatment with a Cas-alpha ortholog. In one embodiment, the sequence of the target polynucleotide is not known prior to modification and treatment with a Cas-alpha ortholog is used as part of a method to determine the sequence of said target polynucleotide.

[0676] Polynucleotide modification with a Cas-alpha ortholog can be achieved by using a full-length polypeptide identified from a Cas locus, or a fragment, modification, or variant of a polypeptide identified from a Cas locus. In some embodiments, the Cas-alpha ortholog is obtained or derived from an organism listed in Table 1. In some embodiments, the Cas-alpha ortholog is a polypeptide sharing at least 80% identity with any of SEQ ID NOs: 86-170 or 511-1135. In some embodiments, the Cas-alpha ortholog is a functional variant of any of SEQ ID NOs: 86-170 or 511-1135. In some embodiments, the Cas-alpha ortholog is a functional fragment of any of SEQ ID NOs: 86-170 or 511-1135. In some embodiments, the Cas-alpha ortholog is a Cas-alpha polypeptide encoded by a polynucleotide selected from the group consisting of SEQ ID NOs: 86-170 or 511-1135. In some embodiments, the Cas-alpha ortholog is a Cas-alpha polypeptide that recognizes a PAM sequence listed in any of Tables 4-83. In some embodiments, the Cas-alpha ortholog is a Cas-alpha polypeptide identified from an organism listed in the Sequence Listing.

[0677] In some embodiments, the Cas-alpha ortholog is provided as a Cas-alpha polynucleotide, hi some embodiments, the Cas-alpha polynucleotide is selected from the group consisting of SEQ ID NOs: 1-85, or a sequence that shares at least 80%, 85%, 90%, 95%, 97%, 99%, or 100% with any one of SEQ ID NOs: 1-85.

[0678] In some embodiments, the Cas-alpha ortholog may be selected from the group consisting of an unmodified wild-type Cas-alpha ortholog, a functional Cas-alpha ortholog variant, a functional Cas-alpha ortholog fragment, a fusion protein comprising an active or inactivated Cas-alpha ortholog, a Cas-alpha ortholog further comprising one or more nuclear localization sequences (NLS) at the C-terminus or N-terminus, or at both the N-terminus and C-terminus, a biotinylated Cas-alpha ortholog, a Cas-alpha ortholog nickase, a Cas-alpha ortholog endonuclease, a Cas-alpha ortholog further comprising a histidine tag, and a mixture of any two or more thereof.

[0679] In some embodiments, the Cas-alpha ortholog is a fusion protein that further comprises a nuclease domain, a transcriptional activator domain, a transcriptional repressor domain, an epigenetic modification domain, a cleavage domain, a nuclear localization signal, a cell penetration domain, a translocation domain, a marker, or a transgene that is heterologous to the target polynucleotide sequence or the cell from which the target polynucleotide sequence is obtained or derived.

[0680] In some embodiments, multiple Cas-alpha orthologs may be desired. In some embodiments, the multiple Cas-alpha orthologs may include Cas-alpha orthologs from different biological sources or from different loci within the same organism. In some embodiments, the multiple Cas-alpha orthologs may include Cas-alpha orthologs with different binding specificities for a target polynucleotide. In some embodiments, the multiple Cas-alpha orthologs may include Cas-alpha orthologs with different cleavage efficiencies. In some embodiments, the multiple Cas-alpha orthologs may include Cas-alpha orthologs with different PAM properties. In some embodiments, the multiple Cas-alpha orthologs may include orthologs of different molecular composition, i.e., polynucleotide Cas-alpha orthologs and polypeptide Cas-alpha orthologs.

[0681] The guide polynucleotide may be provided as a single guide RNA (sgRNA), a chimeric molecule comprising tracrRNA, a chimeric molecule comprising crRNA, a chimeric RNA-DNA molecule, a DNA molecule, or a polynucleotide comprising one or more chemically modified nucleotides.

[0682] Storage conditions for Cas-alpha orthologs and / or guide polynucleotides include parameters related to temperature, state of matter, and time. In some embodiments, Cas-alpha orthologs and / or guide polynucleotides are stored at about -80°C, about -20°C, about 4°C, about 20-25°C, or about 37°C. In some embodiments, Cas-alpha orthologs and / or guide polynucleotides are stored as a liquid, frozen liquid, or lyophilized powder. In some embodiments, Cas-alpha orthologs and / or guide polynucleotides are stable for at least one day, at least one week, at least one month, at least one year, or more than one year.

[0683] Any or all of the possible polynucleotide components of a reaction (e.g., guide polynucleotide, donor polynucleotide, and optionally Cas-alpha polynucleotide) can be provided as part of a vector, construct, linearized or circularized plasmid, or as part of a chimeric molecule. Each component can be provided separately or together in the reaction mixture. In some embodiments, one or more polynucleotide components are operably linked to a heterologous non-coding regulatory element that regulates its expression.

[0684] Methods for modifying a target polynucleotide involve combining the minimum elements into a reaction mixture comprising a Cas-alpha ortholog (or a variant, fragment, or other related molecule as described above), a guide polynucleotide comprising a sequence that is substantially complementary to or selectively hybridizes to the target polynucleotide sequence of the target polynucleotide, and a target polynucleotide for modification. In some embodiments, the Cas-alpha ortholog is provided as a polypeptide. In some embodiments, the Cas-alpha ortholog is provided as a Cas-alpha polynucleotide. In some embodiments, the guide polynucleotide is provided as an RNA molecule, a DNA molecule, an RNA:DNA hybrid, or a polynucleotide molecule comprising chemically modified nucleotides.

[0685] The storage buffer or reaction mixture of any one of the components may be optimized for stability, efficacy, or other parameters. Additional components of the storage buffer or reaction mixture may include a buffer composition, Tris, EDTA, dithiothreitol (DTT), phosphate-buffered saline (PBS), sodium chloride, magnesium chloride, HEPES, glycerol, BSA, salt, emulsifier, detergent, chelating agent, redox agent, antibody, nuclease-free water, proteinase, and / or viscosity agent. In some embodiments, the storage buffer or reaction mixture further comprises a buffer comprising at least one of the following components: HEPES, MgCl, NaCl, EDTA, proteinase, proteinase K, glycerol, nuclease-free water.

[0686] Incubation conditions vary according to the desired results. The temperature is preferably at least 10°C, 10-15°C, at least 15°C, 15-17°C, at least 17°C, 17-20°C, at least 20°C, 20-22°C, at least 22°C, 22-25°C, at least 25°C, 25-27°C, at least 27°C, 27-30°C, at least 30°C, 30-32°C, at least 32°C, 32-35°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, or greater than 40°C. The incubation time is at least 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 6 minutes, at least 7 minutes, at least 8 minutes, at least 9 minutes, at least 10 minutes, or greater than 10 minutes.

[0687] The sequence of the polynucleotide in the reaction mixture before, during, or after incubation can be determined by methods known in the art. In one aspect, the modification of the target polynucleotide is , anti The sequence of the polynucleotide purified from the reaction mixture was determined. before combining with Cas-alpha orthologues This can be confirmed by comparing the sequence with that of the target polynucleotide.

[0688] Any one or more of the compositions disclosed herein useful for in vitro or in vivo polynucleotide detection, binding, and / or modification can be included in a kit. The kit includes a Cas-alpha ortholog or a polynucleotide Cas-alpha ortholog encoding such, and optionally further includes buffer components to allow efficient storage and one or more additional compositions that allow the introduction of the Cas-alpha ortholog or polynucleotide Cas-alpha ortholog into a heterologous polynucleotide, where the Cas-alpha ortholog or polynucleotide Cas-alpha ortholog can result in the modification, addition, deletion, or substitution of at least one nucleotide of the heterologous polynucleotide. In a further aspect, the Cas-alpha orthologs disclosed herein can be used for enrichment of one or more polynucleotide target sequences from a mixed pool. In a further aspect, the Cas-alpha orthologs disclosed herein can be immobilized on a matrix for use in in vitro detection, binding, and / or modification of target polynucleotides.

[0689] The Cas-alpha endonuclease may be attached, associated, or immobilized to a solid matrix for purposes of storage, purification, and / or characterization. Examples of solid matrices include, but are not limited to, filters, chromatography resins, assay plates, test tubes, cryogenic vials, etc. The Cas-alpha endonuclease may be substantially purified and stored in an appropriate buffer solution or lyophilized.

[0690] Detection Method Methods for detecting a Cas-alpha:guide polynucleotide complex bound to a target polynucleotide include any method known in the art, including, but not limited to, microscopy, chromatographic separation, electrophoresis, immunoprecipitation, filtration, nanopore separation, microarrays, and those described below.

[0691] The DNA electrophoretic mobility shift assay (EMSA) examines protein binding to a known DNA oligonucleotide probe to assess the specificity of the interaction. This technique is based on the principle that protein-DNA complexes migrate more slowly than free DNA molecules when subjected to electrophoresis in a polyacrylamide or agarose gel. Because DNA migration slows when bound to a protein, this assay is also called a gel retardation assay. Addition of a protein-specific antibody to the binding component generates a larger complex (antibody-protein-DNA) that migrates more slowly during electrophoresis, a phenomenon known as a supershift, which can be used to confirm the identity of the protein.

[0692] In a DNA pull-down assay, a DNA probe labeled with a high-affinity tag, such as biotin, is used to recover or immobilize the probe. The DNA probe is complexed with proteins from cell lysates in a reaction similar to that used in EMSA, which can then be used to purify the complex using agarose or magnetic beads. The protein is then eluted from the DNA and detected by Western blot or identified by mass spectrometry. Alternatively, the protein can be labeled with an affinity tag, or the DNA-protein complex can be isolated using an antibody against the protein of interest (similar to a supershift assay). In this case, the unknown DNA sequence bound to the protein is detected by Southern blot or PCR analysis.

[0693] Reporter assays provide an in vivo real-time readout of the translational activity of a promoter of interest. A reporter gene is a fusion of a target promoter DNA sequence with a reporter gene DNA sequence customized by the researcher. The DNA sequence encodes a protein with a detectable property, such as firefly / renilla luciferase or alkaline phosphatase. These genes produce an enzyme only when the promoter of interest is activated. The enzyme then catalyzes a substrate to produce a light or color change that can be detected by spectroscopy. The signal from the reporter gene is used as an indirect determinant of the translation of an endogenous protein driven from the same promoter.

[0694] The microplate capture and detection assay uses immobilized DNA probes to capture specific protein-DNA interactions and confirm protein identity and relative abundance with target-specific antibodies. Typically, DNA probes are immobilized on the surface of a streptavidin-coated 96- or 384-well microplate. Cell extracts are prepared and added to allow binding proteins to bind to the oligonucleotides. The extract is then removed, and each well is washed several times to remove nonspecifically bound proteins. Finally, proteins are detected using a labeled specific antibody. This method is extremely sensitive, capable of detecting less than 0.2 pg of target protein per well. This method can also be utilized with oligonucleotides labeled with other tags, such as primary amines, which can be immobilized on microplates coated with amino-reactive surface chemistry.

[0695] DNA footprinting is one of the most widely used methods for obtaining detailed information about individual nucleotides in protein-DNA complexes, even in living cells. In such methods, chemicals or enzymes are used to modify or digest DNA molecules. When sequence-specific proteins bind to DNA, they can protect the binding site from modification or digestion. This can then be visualized by denaturing gel electrophoresis, in which unprotected DNA is cleaved more or less randomly. Therefore, it appears as a "ladder" of bands, and sites protected by the protein have no corresponding bands, appearing as footprints in the band pattern. These footprints therefore identify specific nucleosides at the protein-DNA binding site.

[0696] Microscopy techniques include optical microscopy, fluorescence microscopy, electron microscopy, and atomic force microscopy (AFM).

[0697] Chromatin immunoprecipitation analysis (ChIP) allows proteins to be covalently bound to their DNA targets, which can then be unbound and characterized separately.

[0698] Systematic evolution of ligands by in vitro selection (SELEX) exposes a target protein to a random library of oligonucleotides. Those that bind are isolated and amplified by PCR.

[0699] The methods and compositions provided herein include, but are not limited to, the following aspects.

[0700] Embodiment 1: A synthetic composition comprising: (a) a guide polynucleotide; (b) a Cas endonuclease comprising a C-terminal tripartite RuvC domain further comprising a bridge helix and at least one zinc finger domain, an alpha helix bundle, and multiple beta sheets forming a wedge domain, wherein the Cas endonuclease is less than 650 amino acids in length; and (c) a target sequence comprising a nucleotide sequence that shares complementarity with the guide polynucleotide, wherein the guide polynucleotide and the Cas endonuclease form a complex that cleaves a double-stranded DNA polynucleotide comprising the target sequence.

[0701] Aspect 2: (a) guide polynucleotide; (b) a genera selected from the group consisting of Archaea, Microarchaea, Acidibacillus sulfuroxidans, Candidatus Micrarchaeota archaeon, Clostridium novyi, Parageobacillus thermoglucosidasius, Ruminococcus species, and Syntrophomonas palmitatica. palmitatica), wherein the Cas endonuclease forms a complex with a guide polynucleotide; and (c) a double-stranded DNA polynucleotide comprising a target sequence that binds to the guide polynucleotide, wherein the guide polynucleotide and the Cas endonuclease form a complex that cleaves the double-stranded DNA polynucleotide comprising the target sequence.

[0702] Embodiment 3: The synthetic composition of embodiment 1 or embodiment 2, wherein the Cas endonuclease further comprises a zinc finger domain near the N-terminus.

[0703] Embodiment 4: The synthetic composition of embodiment 1 or embodiment 2, wherein the double-stranded DNA polynucleotide further comprises a PAM.

[0704] Embodiment 5: The synthetic composition of embodiment 4, wherein the PAM comprises a plurality of thymine nucleotides.

[0705] Embodiment 6: The synthetic composition of embodiment 1 or embodiment 2, further comprising a heterologous polynucleotide.

[0706] Embodiment 7: The synthetic composition of embodiment 1 or embodiment 2, wherein the guide polynucleotide comprises a 20 nucleotide region having complementarity to the target sequence.

[0707] Embodiment 8: The synthetic composition of embodiment 1 or embodiment 2, wherein the guide polynucleotide is a double-stranded molecule comprising tracrRNA and crRNA.

[0708] Embodiment 9: The synthetic composition of embodiment 1 or embodiment 2, wherein the guide polynucleotide is a single guide polynucleotide comprising a Cas endonuclease recognition domain and a variable targeting domain.

[0709] Embodiment 10: The synthetic composition of embodiment 6, wherein the heterologous polynucleotide is an expression element.

[0710] Embodiment 11: The synthetic composition of embodiment 6, wherein the heterologous polynucleotide is a transgene.

[0711] Embodiment 12: The synthetic composition of embodiment 6, wherein the heterologous polynucleotide is a donor DNA molecule.

[0712] Embodiment 13: The synthetic composition of embodiment 6, wherein the heterologous polynucleotide is a modified polynucleotide template.

[0713] Embodiment 14: The synthetic composition of embodiment 1 or embodiment 2, wherein the CRISPR-Cas endonuclease further comprises a nuclear localization signal.

[0714] Embodiment 15: The synthetic composition of embodiment 1 or embodiment 2, wherein the CRISPR-Cas endonuclease is Cas-alpha or a functional fragment thereof.

[0715] Embodiment 16: The synthetic composition of embodiment 1 or embodiment 2, wherein the CRISPR-Cas endonuclease is catalytically inactive Cas-alf.

[0716] Embodiment 17: The synthetic composition of embodiment 1 or embodiment 2, wherein the CRISPR-Cas endonuclease is a fusion protein comprising a functional fragment of Cas-alpha.

[0717] Embodiment 18: The synthetic composition of embodiment 17, wherein the fusion protein further comprises another nuclease domain.

[0718] Embodiment 19: The synthetic composition of embodiment 1 or embodiment 2, further comprising at least one additional polypeptide.

[0719] Embodiment 20: The synthetic composition of embodiment 19, wherein the further polypeptide is selected from the group consisting of Cas1, Cas2, and Cas4.

[0720] Embodiment 21: The synthetic composition of embodiment 1 or embodiment 2, further comprising a cell.

[0721] Embodiment 22: The synthetic composition of embodiment 21, wherein the cell is a eukaryotic cell.

[0722] Embodiment 23: The synthetic composition of embodiment 21, wherein the cell is a plant cell.

[0723] Embodiment 24: The synthetic composition of embodiment 23, wherein the plant cell is a monocotyledonous plant cell or a dicotyledonous plant cell.

[0724] Aspect 25: The synthetic composition of aspect 23, wherein the plant cell is derived from an organism selected from the group consisting of corn, soybean, cotton, wheat, canola, oilseed rape, sorghum, rice, rye, barley, millet, oat, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis, safflower, and tomato.

[0725] Embodiment 26: The synthetic composition of embodiment 21, further comprising a guide polynucleotide comprising a variable targeting domain that is substantially complementary to a target sequence in the genome of the cell.

[0726] Embodiment 27: A polynucleotide encoding the synthetic composition of embodiment 1 or embodiment 2.

[0727] Embodiment 28: The polynucleotide of embodiment 27, further comprising at least one additional polynucleotide.

[0728] Embodiment 29: The polynucleotide according to embodiment 28, wherein the at least one further polynucleotide is an expression element.

[0729] Embodiment 30: The polynucleotide according to embodiment 28, wherein the at least one further polynucleotide is a gene.

[0730] Embodiment 31: The synthetic composition of embodiment 30, wherein the gene is selected from the group consisting of Cas1, Cas2, and Cas4.

[0731] Embodiment 32: The polynucleotide of embodiment 28, wherein at least one polynucleotide is comprised within a recombinant construct.

[0732] Embodiment 33: The synthetic composition of embodiment 1 or embodiment 2, wherein at least one component is attached to a solid matrix.

[0733] Embodiment 34: A synthetic composition comprising a target double-stranded DNA polynucleotide, a guide polynucleotide that is complementary to a sequence in the double-stranded DNA polynucleotide, and a Cas endonuclease that is at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 17, 18, 19, 20, 32, 33, 34, 35, 36, 37, and 38, or a functional fragment or variant thereof.

[0734] Embodiment 35: A synthetic composition comprising a target double-stranded DNA polynucleotide, a polynucleotide encoding a guide polynucleotide that is complementary to a sequence in the double-stranded DNA polynucleotide, and a cas endonuclease gene that is at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 13, 14, 15, 16, 25, 26, 27, 28, 29, 30, and 31, or a functional fragment or variant thereof.

[0735] Embodiment 36: A method for introducing a site-specific modification into a target sequence in the genome of a cell, the method comprising introducing into the cell the synthetic composition of any of embodiments 1 to 35.

[0736]

[0013] Embodiment 37: A method of generating an organism having a modified genome, comprising: (a) introducing into at least one cell of the organism a heterologous composition comprising: i. a Cas-alpha endonuclease or a cas-alpha polynucleotide encoding the Cas-alpha endonuclease; ii. a guide polynucleotide comprising a variable targeting domain substantially complementary to a target sequence in the genome of the cell, wherein the guide polynucleotide and the Cas-alpha endonuclease are capable of forming a complex that is capable of recognizing, binding to, and optionally nicking or cleaving the target sequence; and iii. a polynucleotide-modified template comprising at least one region that is complementary to a PAM sequence adjacent to a DNA target sequence recognized by the Cas-alpha complex, wherein the at least one region that is complementary to the PAM sequence comprises at least one nucleotide mismatch; (b) incubating the cells; (c) generating a whole organism from the cells; and (d) prior to introducing the heterologous composition of (a), of and identifying the presence of at least one nucleotide modification in the genome of at least one cell of the organism by comparing the genome of the cell to a target sequence.

[0737] Embodiment 38: The method of embodiment 36 or 37, wherein the cell is a eukaryotic cell.

[0738] Embodiment 39: The method according to embodiment 38, wherein the eukaryotic cell is derived from or obtained from an animal or plant.

[0739] Embodiment 40: The method of embodiment 39, wherein the plant is a monocotyledonous or dicotyledonous plant.

[0740] Aspect 41: The method of aspect 39, wherein the plant is selected from the group consisting of corn, soybean, cotton, wheat, canola, oilseed rape, sorghum, rice, rye, barley, millet, oat, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis, safflower, and tomato.

[0741] Embodiment 42: The method of embodiment 36 or 37, further comprising introducing a heterologous polynucleotide.

[0742] Embodiment 43: The method of embodiment 42, wherein the heterologous polynucleotide is a donor DNA molecule.

[0743] Embodiment 44 The method of embodiment 42, wherein the heterologous polynucleotide is a modified polynucleotide template comprising a sequence that is at least 50% identical to a sequence of the cell.

[0744] Embodiment 45: A progeny of an organism obtainable by the method of embodiment 37, which progeny retains at least one nucleotide modification in at least one cell.

[0745] Embodiment 46: A method of modifying a genomic sequence of a target cell, the method comprising providing a Cas endonuclease comprising an amino acid sequence that is at least 95% to 100% identical to one of SEQ ID NOs: 17, 18, 19, 20, 32, 33, 34, 35, 36, 37, and 38, and a guide polynucleotide that targets the genomic sequence of the target cell; and introducing a double-stranded break in the genomic sequence of the target cell, thereby modifying the genomic sequence of the target cell.

[0746] While the present invention has been particularly shown and described with reference to preferred embodiments and various alternative embodiments, those skilled in the relevant art will recognize that various changes in form and detail can be made therein without departing from the spirit and scope of the invention. For example, the following specific examples may illustrate the methods and embodiments described herein using particular target sites or target organisms, but the principles in these examples may be applied to any target site or target organism. Accordingly, it will be understood that the scope of the present invention is encompassed by the embodiments of the invention listed herein rather than the specific examples exemplified below. All cited patents, applications, and publications referenced in this application are incorporated herein by reference in their entirety for all purposes to the same extent as if each were individually and specifically incorporated by reference. [Example]

[0747] Below are examples of specific embodiments of some aspects of the present invention. The examples are provided for illustrative purposes only and are not intended to limit the scope of the present invention in any way. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should, of course, be allowed for.

[0748] Example 1: Identification and characterization of a novel class of Cas-alpha CRISPR-Cas systems This example describes how novel class 2 CRISPR (clustered regularly interspaced short palindromic repeats)-Cas (CRISPR-associated) loci were identified using identification of an operon-like gene structure and protein structure analysis.

[0749] First, CRISPR arrays were detected within microbial sequences using the software programs PILER-CR (Edgar, R. (2007) BMC Bioinformatics, 8:18) and MinCED (Bland, C. et al. (2007) BMC Bioinformatics, 8:209). Known CRISPR-Cas systems were then removed from the dataset by searching for proteins encoded near the CRISPR array (20 kb 5' and 20 kb 3', if possible) for homology to known CRISPR-associated (Cas) proteins using a set of position-specific scoring matrices (PSSMs) encompassing all known Cas protein families, as described in Makarova, K. et al. (2015) Nature Reviews Microbiology, 13:722-736. To aid in the complete removal of known Class 2 CRISPR-Cas systems, multiple sequence alignments of protein sequences from a collection of orthologs of each family of Class 2 CRISPR-Cas endonucleases (e.g., Cas9, Cpf1 (Cas12a), C2c1 (Cas12b), C2c2 (Cas13), and C2c3 (Cas12c)) were performed using MUSCLE (Edgar R. (2004) Nucleic Acids Res. 32:1792-1797). The alignments were inspected, curated, and used to build profile hidden Markov models (HMMs) using HMMER (Eddy, SR (1998) Bioinformatics. 14:755-763; Eddy, SR (2011) PLoS Comp. Biol., 7:e1002195). The resulting HMM model was then used to further identify and remove known Class 2 CRISPR-Cas systems from the dataset. Using a PSSM-specific search as described above, the remaining CRISPR loci were then assessed for the presence of genes encoding Cas1 and Cas2, proteins thought to be important for spacer insertion and adaptation (Makarova, K. et al. 2015, Nature Reviews Microbiology 13:722-736).Next, CRISPR loci containing the cas1 and cas2 genes were selected and further investigated to determine the proximity, order, and orientation of the unidentified genes encoded at the loci relative to the cas1 and cas2 genes and the CRISPR array. Only CRISPR loci in which large (≥1500 bp open reading frame) unidentified genes were located adjacent to the cas1 and cas2 genes and in the same transcriptional direction, forming an operon-like structure, were selected for further analysis. Proteins encoded by the unidentified genes were then analyzed for sequence and structural features indicative of class 2 endonucleases capable of cleaving DNA. Depending on the degree of similarity between the candidate sequences and known proteins, various bioinformatics tools were used to reveal their conserved functional features through pairwise comparisons, family profile searches, structural threading, and manual structural inspection. Generally, homologous sequences of new candidate proteins were first collected against the National Center for Biotechnology Information (NCBI) nonredundant (NR) protein collection by PSI-BLAST (Altschul, S. F. et al. (1997) Nucleic Acids Res. 25:3389-3402) search with an e-value cutoff of 0.01. After reducing redundancy at approximately 90% identity, homologous sequences with various member inclusion thresholds (e.g., greater than 60, 40, or 20% identity) were aligned to reveal conserved motifs using multiple sequence alignment tools, MSAPRobs (Liu, Y. et al. (2010) Bioinformatics. 26:1958-1964) and Clustalw. The most conserved homologous sequences were subjected to sequence-to-family profile searches using HMMER (Eddy, SR (1998) Bioinformatics. 14:755-763) against a number of domain databases, including Pfam, Superfamily, and SCOP (Murzin, A Get al. (1995) J. Mol. Biol. 247:536-540), as well as self-generated structure-based profiles.Additionally, homologous sequence alignments of the resulting candidates were used to generate candidate protein profiles with predicted secondary structures. Furthermore, profile-profile searches were performed using the candidate profiles against the pdb70_hhm and Pfam_hhm profile databases using HHSEARCH (Soding, J. et al. (2006) Nucleic Acids Res. 34:W374-378). In the next step, all of the detected sequence-structure relationships and conserved motifs were threaded into 3D structural templates using MODELLER or manually mapped to known structural references in DiscoveryStudio (BIOVIA) and Pymol (Schrodinger). Finally, to verify and confirm the potential biological relevance of the proteins as class 2 endonucleases, the catalytic or most conserved residues and key structural integrity were manually examined and evaluated against the protein's biochemical function. Following structural identification of key features indicative of a Class 2 endonuclease (e.g., a DNA cleavage domain), other proteins encoded within the locus (5 kb 5' and 5 kb 3', if possible, from the end of the newly defined CRISPR-Cas system) were then examined for homology to known protein families using InterProScan software (EMBL-EBI, UK) and by comparison with the NCBI NR protein collection via the BLAST program (Altschul, S. F. et al. (1990) J. Mol. Biol. 215:403-410). Genes encoding proteins with similarity (at least 30% identity) to known proteins were annotated as such within the CRISPR-Cas locus.

[0750] Initially, we identified four novel class 2 CRISPR-Cas systems from unknown microorganisms (Table 1). As shown in Figures 1A and 1B, each locus encoded a complete CRISPR-Cas system containing all the components necessary for acquisition and interference. These included genes encoding all the proteins required for spacer acquisition and integration (Cas1, Cas2, and optionally Cas4) as well as a novel protein, Cas-alpha (α), containing a DNA cleavage domain, in an operon-like structure flanking the CRISPR array.

[0751] [Table 1]

[0752] Next, we compared the Cas-alpha endonucleases to the NCBI NR protein collection using BLAST, followed by analysis by MinCED. Proteins were found in close proximity (≤5 kb) to the CRISPR array, yielding seven additional CRISPR systems (Table 2). The gene structures of the loci identified for these new proteins are shown in Figures 1C and 1D. The locus encoding Cas-alpha 6 contains the complete cas2 and cas4 genes in addition to a partial cas1 gene (Figure 1C), while Cas-alphas 5, 7, 8, 9, 10, and 11 contain only the endonuclease gene adjacent to the CRISPR array (Figure 1D). The loci of Cas-alphas 18 and 19 are shown in Figure 21A, and their mechanisms of action are shown in Figure 21B.

[0753] [Table 2]

[0754] Structural examination of these proteins revealed that they differed from previously described class 2 CRISPR-Cas endonucleases capable of recognizing and cleaving double-stranded DNA targets. First, the size of the endonucleases (422–613 amino acids) was remarkably compact compared to other known class 2 CRISPR-Cas systems. Second, the first amino (N)-terminal half of the proteins was highly variable in sequence composition, as evidenced by the lack of conservation of even a single amino acid (except the initiating methionine). Despite this, secondary structure prediction (PSIPRED (Jones, JT (1999) J. Mol. Biol. 292:195–202)) showed a mixture of beta-strands and alpha-helices, suggesting the presence of a wedge-shaped (WED) or oligonucleotide-binding domain (OBD) structure and a helical bundle in the N-terminal region of all Cas-alpha proteins. The carboxyl (C)-terminal halves of the proteins retain...

Claims

1. (a) a Cas endonuclease that is at least 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 35, 37, and 38, or a polynucleotide encoding the Cas endonuclease, wherein the Cas endonuclease further comprises: (i) a C-terminal three-split RuvC domain; (ii) the following amino acid motifs: GxxxG, ExL, CxnC, and Cxn(C or H), where G=glycine, E=glutamic acid, C=cysteine, H=histidine, x=any amino acid, and n=an integer from 0 to 11; (iii) an alpha helix, and (iv) multiple beta sheets forming a wedge-shaped domain a Cas endonuclease or a polynucleotide encoding said Cas endonuclease, comprising: (b) a target double-stranded DNA polynucleotide, the target double-stranded DNA polynucleotide being heterologous to the Cas endonuclease source; and (c) a guide polynucleotide comprising a variable targeting domain that includes a region complementary to the target double-stranded DNA polynucleotide. Includes: The Cas endonuclease recognizes a PAM sequence on the target double-stranded DNA polynucleotide, and the guide polynucleotide and the Cas endonuclease form a complex that binds to the target double-stranded DNA polynucleotide. Synthetic composition.

2. 2. The synthetic composition of claim 1, wherein the Cas endonuclease comprises fewer than 800 amino acids.

3. 10. The synthetic composition of claim 1, wherein the Cas endonuclease is provided as a polynucleotide encoding the Cas endonuclease.

4. 10. The synthetic composition of claim 1, wherein the Cas endonuclease cleaves the double-stranded DNA polynucleotide.

5. The synthetic composition of claim 1 further comprising a heterologous polynucleotide.

6. The synthetic composition of claim 5 , wherein the heterologous polynucleotide is an expression element.

7. The synthetic composition of claim 5 , wherein the heterologous polynucleotide is a transgene.

8. The synthetic composition of claim 5 , wherein the heterologous polynucleotide is a donor DNA molecule.

9. The synthetic composition of claim 5 , wherein the heterologous polynucleotide is a modified polynucleotide template.

10. 2. The synthetic composition of claim 1, wherein the CRISPR-Cas endonuclease is catalytically inactive.

11. 2. The synthetic composition of claim 1, wherein the Cas endonuclease recognizes a PAM sequence containing multiple T or C nucleotides.

12. 11. The synthetic composition of claim 10, wherein the PAM sequence is selected from the group consisting of TTAT, TTTR, N(T>V)TTR, N(W>S)TTTR, N(Y>R)N(Y>S>R)TTN(A>G>Y), N(W>S)N(Y>R)TTTR, CTT, N(T>W>C)TTC, and CCD.

13. 10. The synthetic composition of claim 1, wherein the Cas endonuclease is part of a fusion protein.

14. 12. The synthetic composition of claim 11, further comprising a deaminase.

15. The synthetic composition of claim 11 , wherein the fusion protein further comprises a heterologous nuclease domain.

16. 10. The synthetic composition of claim 1, further comprising a non-human eukaryotic cell.

17. 17. The synthetic composition of claim 16, wherein the eukaryotic cell is a plant cell, an animal cell, or a fungal cell.

18. 18. The synthetic composition of claim 17, wherein the plant cell is a monocotyledonous plant cell or a dicotyledonous plant cell.

19. 18. The synthetic composition of claim 17, wherein the plant cells are derived from an organism selected from the group consisting of corn, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oat, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis, safflower, and tomato.

20. 10. The synthetic composition of claim 1, wherein at least one component is attached to a solid matrix.

21. (a) a Cas endonuclease that is at least 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 35, 37, and 38; (b) a target double-stranded DNA polynucleotide, the target double-stranded DNA polynucleotide being heterologous to the Cas endonuclease source; and (c) a guide polynucleotide comprising a variable targeting domain that includes a region complementary to the target double-stranded DNA polynucleotide. Includes: The Cas endonuclease recognizes a PAM sequence on the target double-stranded DNA polynucleotide, and the guide polynucleotide and the Cas endonuclease form a complex that binds to the target double-stranded DNA polynucleotide. Synthetic composition.

22. A method for introducing a targeted edit into a target polynucleotide, comprising: (a) a Cas endonuclease that is at least 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 35, 37, and 38, wherein the Cas endonuclease further comprises: (i) a C-terminal three-split RuvC domain; (ii) the following amino acid motifs: GxxxG, ExL, CxnC, and Cxn(C or H), where G=glycine, E=glutamic acid, C=cysteine, H=histidine, x=any amino acid, and n=an integer from 0 to 11; (iii) an alpha helix, and (iv) multiple beta sheets forming a wedge-shaped domain a Cas endonuclease that recognizes a PAM sequence on the target polynucleotide; (b) a guide polynucleotide comprising a variable targeting domain substantially complementary to a portion of the target polynucleotide, wherein the guide polynucleotide and the Cas-endonuclease form a complex capable of recognizing and binding to the target polynucleotide. providing a heterogeneous composition comprising: The method, wherein the target polynucleotide is in at least one biological cell, and the cell is a non-human cell.

23. introducing at least one nucleotide modification into the target polynucleotide of the at least one biological cell compared to the target polynucleotide of the at least one cell prior to introducing the heterologous composition; incubating said at least one cell and generating an organism from said at least one cell; 23. The method of claim 22, further comprising confirming the presence of said at least one nucleotide modification in said target polynucleotide of at least one cell of said organism.

24. 23. The method of claim 22, wherein the Cas endonuclease recognizes a PAM sequence containing multiple T or C nucleotides.

25. 23. The method of claim 22, wherein the PAM sequence is selected from the group consisting of TTAT, TTTR, N(T>V)TTR, N(W>S)TTTR, N(Y>R)N(Y>S>R)TTN(A>G>Y), N(W>S)N(Y>R)TTTR, CTT, N(T>W>C)TTC, and CCD.

26. 24. The method of claim 23, wherein the cell is a eukaryotic cell.

27. 27. The method of claim 26, wherein the eukaryotic cell is derived from or obtained from an animal, fungus, or plant.

28. 28. The method of claim 27, wherein the plant is a monocotyledonous or dicotyledonous plant.

29. 28. The method of claim 27, wherein the plant is selected from the group consisting of corn, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oat, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis, safflower, and tomato.

30. 24. The method of claim 23, further comprising introducing a heterologous polynucleotide.

31. 31. The method of claim 30, wherein the heterologous polynucleotide is a donor DNA molecule.

32. 31. The method of claim 30, wherein the heterologous polynucleotide is a modified polynucleotide template comprising a sequence that is at least 50% identical to a sequence of the cell.

Citation Information

Patent Citations

  • CRISPR enzyme mutations that reduce off-target effects

    JP2018522546A

  • CasZ Compositions and Methods of Use

    JP2021503278A

  • Methods for identifying class 2 crispr-CAS systems

    WO2018035250A1

  • CASZ compositions and methods of use

    WO2019089820A1