Mb2Cas12a variants with enhanced efficiency

Enhanced Mb2Cas12a variants with specific mutations and domain swaps address the low nuclease activity of wild-type Mb2Cas12a, achieving effective plant genome editing.

JP2026505030APending Publication Date: 2026-02-10SYNGENTA CROP PROTECITON AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025543307
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-27
Filing Date
2024-01-24
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

The wild-type Mb2Cas12a enzyme exhibits lower nuclease activity compared to other Cas12a orthologs, limiting its effectiveness in plant genome editing.

Method used

Development of Mb2Cas12a variants with specific amino acid substitutions and domain swaps, such as D172R, F357W, and WED-MR1 replacements, to enhance DNA cleavage activity in plants.

Benefits of technology

The enhanced Mb2Cas12a variants demonstrate improved nuclease activity, enabling efficient plant genome editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026505030000001_ABST
    Figure 2026505030000001_ABST
Patent Text Reader

Abstract

Described herein are Cas12a mutants derived from Moraxella bovoculi AAX08 and methods for their use. Mb2Cas12a mutants can contain single amino acid substitutions, multiple amino acid substitutions, domain swaps, or all of the above. These mutants have enhanced DNA cleavage activity in plants compared to the wild-type Moraxella bovoculi AAX08 Cas12a enzyme.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cas12a mutants derived from Moraxella bovoculi AAX08 and methods for their use are described. These mutants have enhanced DNA cleavage activity in plants compared to the wild-type enzyme.

[0002] Priority claims This application claims priority under 35 U.S.C. § 119 to PCT Application No. PCT / CN2023 / 073490, filed January 27, 2023, the entire contents of which are incorporated herein by reference.

[0003] Sequence Listing This application is accompanied by a Sequence Listing entitled "Mb2Cas12a Variants with Enhanced Efficiency_ST26.xml," created on November 2, 2022, which is approximately 967 kilobytes in size. This Sequence Listing is incorporated herein by reference in its entirety. This Sequence Listing has been submitted herewith via EFS-Web and complies with 37 C.F.R. § 1.824(a)(2)-(6) and (b). [Background technology]

[0004] Although Mb2Cas12a from Moraxella bovoculi AAX08 has demonstrated plant genome editing capabilities (Zhang et al., 2021), its nuclease activity is lower than that of other Cas12a orthologs widely used in plant or mammalian cells, such as LbCas12a, AsCas12a, or FnCas12a (Zetsche et al., 2020; Zhang et al., 2021). Due to its unique properties of high performance at low temperatures and the potential to recognize shorter PAMs, Mb2Cas12a should have improved nuclease activity in plants. Summary of the Invention [Means for solving the problem]

[0005] Thus, there is a need for wild-type mutant Mb2Cas12a polypeptide variants. These variants can include at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO: 1. The substitution can occur at the following positions: D172, F357, F547, A742, E797, Y819, E913, I914, L917, N918, V921, H939, and / or Y1172. Specific substitutions include D172R, D172K, F357W, F547Y, A742S, E797A, Y819F, I914K, L917V, N918A, N918K, V921K, V921Q, H939Q, and Y1172N. Specific sequences of desirable variants of Mb2Cas12a include SEQ ID NOs: 3, 5, 26, 35, 57, 58, 64, 76, 80, 85, 86, 87, 88, or 127.

[0006] Additional variants can be obtained by swapping domains from Mb2Cas12a with orthologous domains from other Cas12a peptides. Domain swaps can occur in the following domains: WED-MR1, WED-MR2, BH, Up-seq, and Nuc. The WED-MR1 domain can be replaced with SEQ ID NO: 24, 29, or 30. The WED-MR2 domain can be replaced with SEQ ID NO: 25, 32, or 33. The BH domain can be replaced with SEQ ID NO: 40, 41, or 42. The Up-seq domain can be replaced with SEQ ID NO: 59, 60, 61, 62, or 63. The Nuc domain can be replaced with SEQ ID NO: 65, 66, and 67. Domain-swapped variants of Mb2Cas12a can include SEQ ID NO: 37, 59, 60, 61, 62, 63, 74, 75, 77, or 78.

[0007] Variant Mb2Cas12a polypeptides can also include amino acid substitutions and domain swaps. Examples of such variants include SEQ ID NOs: 36, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, and 107.

[0008] These mutants are useful in methods of editing plants when the mutants described herein are used to contact a plant genome. Guide RNAs, such as those found in SEQ ID NOS: 18-21, can also be used in these methods. Performing these methods results in edited plants. Constructs and plasmids encoding these mutant Mb2Cas12a polypeptides can be used to express the desired mutants in appropriate organisms and / or tissues. These organisms can be non-human cells, such as plant cells or tissues.

[0009] A brief description of the sequences in the sequence listing SEQ ID NO: 1 is the amino acid sequence of wild-type Mb2Cas12a.

[0010] SEQ ID NO: 2 is a nucleotide sequence encoding the amino acid sequence of wild-type Mb2Cas12a.

[0011] SEQ ID NO: 3 is the amino acid sequence of Mb2Cas12a containing the D172R mutation.

[0012] SEQ ID NO: 4 is a nucleotide sequence encoding the amino acid sequence of Mb2Cas12a containing the D172R mutation.

[0013] SEQ ID NO: 5 is the amino acid sequence of Mb2Cas12a containing the D172K mutation.

[0014] SEQ ID NO: 6 is a nucleotide sequence encoding the amino acid sequence of Mb2Cas12a containing the D172K mutation.

[0015] SEQ ID NO:7 is the nucleotide sequence of the sugarcane ubiquitin promoter ("prSoUbi4-02").

[0016] SEQ ID NO: 8 is the amino acid sequence of the SV40 nuclear localization signal ("xSV40NLS-06").

[0017] SEQ ID NO: 9 is the amino acid sequence of a 30 amino acid flexible peptide linker with a repeating motif (GGGGS)6 ("xLinker-06").

[0018] SEQ ID NO: 10 is the amino acid sequence of an 8 amino acid short linker ("xSGGSlinker-02") with a repeat motif (SGGS)2.

[0019] SEQ ID NO:11 is the nucleotide sequence of the Agrobacterium tumefaciens nopaline synthase gene terminator ("tNOS-05-01").

[0020] SEQ ID NO: 12 is the nucleotide sequence of the ribozyme ("rHH-05").

[0021] SEQ ID NO: 13 is the nucleotide sequence of the ribozyme ("rHDV-01").

[0022] SEQ ID NO: 14 is the nucleotide sequence of a crRNA array containing ribozymes targeting four maize genes ("rMb2gRNACas12aZmWxy1-01, rMb2gRNACas12aZmBX9-A, rMb2gRNACas12aZmGL2-01, rMb2gRNACas12aZmBINa").

[0023] SEQ ID NO: 15 is the nucleotide sequence of construct 26411.

[0024] SEQ ID NO: 16 is the nucleotide sequence of construct 26363.

[0025] SEQ ID NO: 17 is the nucleotide sequence of construct 26410.

[0026] SEQ ID NO: 18 is a nucleotide sequence encoding the guide RNA portion that hybridizes to the target gene ZmWx1.

[0027] SEQ ID NO: 19 is a nucleotide sequence encoding the guide RNA portion that hybridizes to the target gene ZmBx9.

[0028] SEQ ID NO: 20 is a nucleotide sequence encoding the guide RNA portion that hybridizes to the target gene ZmGL2.

[0029] SEQ ID NO: 21 is a nucleotide sequence encoding the guide RNA portion that hybridizes to the target gene ZmBINa.

[0030] SEQ ID NO: 22 is the amino acid sequence of WED-MR1-Mb2 (same as SEQ ID NO: 31).

[0031] SEQ ID NO: 23 is the amino acid sequence of WED-MR2-Mb2 (same as SEQ ID NO: 32).

[0032] SEQ ID NO: 24 is the amino acid sequence of the WED-MR1 homolog from LbCas12a.

[0033] SEQ ID NO: 25 is the amino acid sequence of the WED-MR2 homolog from LbCas12a.

[0034] SEQ ID NO: 26 is the amino acid sequence of Mb2Cas12a containing the E797A mutation.

[0035] SEQ ID NO: 27 is the nucleotide sequence of construct 27731.

[0036] SEQ ID NO: 28 is the nucleotide sequence of construct 26840.

[0037] SEQ ID NO: 29 is the amino acid sequence of the WED-MR1 homolog from AsCas12a.

[0038] SEQ ID NO: 30 is the amino acid sequence of the WED-MR1 homologue derived from FnCas12a.

[0039] SEQ ID NO: 31 is the amino acid sequence of the WED-MR1 homolog from Mb2Cas12a and Mb2Cas12a-22581.

[0040] SEQ ID NO: 32 is the amino acid sequence of the WED-MR2 homologue derived from AsCas12a.

[0041] SEQ ID NO: 33 is the amino acid sequence of the WED-MR2 homologue derived from FnCas12a.

[0042] SEQ ID NO: 34 is the amino acid sequence of the WED-MR2 homolog from Mb2Cas12a and Mb2Cas12a-22581.

[0043] SEQ ID NO: 35 is the amino acid sequence of Mb2Cas12a containing the D172R and E797A mutations.

[0044] SEQ ID NO: 36 is the amino acid sequence of Mb2Cas12a containing D172R and WED-MR1-Lb.

[0045] SEQ ID NO: 37 is the amino acid sequence of Mb2Cas12a containing WED-MR2-Lb.

[0046] SEQ ID NO: 38 is the nucleotide sequence of construct 26841.

[0047] SEQ ID NO: 39 is the nucleotide sequence of construct 27493.

[0048] SEQ ID NO: 40 is the amino acid sequence of the BH domain of LbCas12a.

[0049] SEQ ID NO: 41 is the amino acid sequence of the BH domain of AsCas12a.

[0050] SEQ ID NO: 42 is the amino acid sequence of the BH domain of FnCas12a.

[0051] Sequence number 43 is the amino acid sequence of the BH domain of Mb2Cas12a and Mb2Cas12a-22581.

[0052] Sequence number 44 is the amino acid sequence of the Up-seq region of LbCas12a.

[0053] SEQ ID NO: 45 is the amino acid sequence of the Up-seq region of AsCas12a.

[0054] SEQ ID NO: 46 is the amino acid sequence of the Up-seq region of FnCas12a.

[0055] SEQ ID NO: 47 is the amino acid sequence of the Up-seq region of Mb2Cas12a-22581.

[0056] Sequence number 48 is the amino acid sequence of the Up-seq region of Mb2Cas12a.

[0057] SEQ ID NO: 49 is the nucleotide sequence of construct 26442.

[0058] SEQ ID NO: 50 is the nucleotide sequence of construct 26443.

[0059] SEQ ID NO:51 is the nucleotide sequence of construct 26623.

[0060] SEQ ID NO: 52 is the nucleotide sequence of construct 26444.

[0061] SEQ ID NO: 53 is the nucleotide sequence of construct 27031.

[0062] SEQ ID NO: 54 is the nucleotide sequence of construct 26445.

[0063] SEQ ID NO: 55 is the nucleotide sequence of construct 27030.

[0064] SEQ ID NO: 56 is the nucleotide sequence of construct 27927.

[0065] SEQ ID NO: 57 is the amino acid sequence of Mb2Cas12a V921Q.

[0066] SEQ ID NO: 58 is the amino acid sequence of Mb2Cas12a V921K.

[0067] SEQ ID NO: 59 is the amino acid sequence of Mb2Cas12a BH1-Lb.

[0068] SEQ ID NO: 60 is the amino acid sequence of Mb2Cas12a BH2-Lb.

[0069] SEQ ID NO: 61 is the amino acid sequence of Mb2Cas12a BH1-As.

[0070] SEQ ID NO: 62 is the amino acid sequence of Mb2Cas12a BH2-As.

[0071] SEQ ID NO: 63 is the amino acid sequence of Mb2Cas12a BH3-As.

[0072] SEQ ID NO: 64 is the amino acid sequence of Mb2Cas12a L917V+N918K.

[0073] SEQ ID NO: 65 is the amino acid sequence of the microregion in the Nuc domain of LbCas12a.

[0074] SEQ ID NO: 66 is the amino acid sequence of a microregion in the Nuc domain of AsCas12a.

[0075] SEQ ID NO: 67 is the amino acid sequence of a microregion in the Nuc domain of FnCas12a.

[0076] SEQ ID NO: 68 is the amino acid sequence of the microregion in the Nuc domain of Mb2Cas12a and Mb2Cas12a-22581.

[0077] SEQ ID NO: 69 is the nucleotide sequence of construct 26440.

[0078] SEQ ID NO: 70 is the nucleotide sequence of construct 26441.

[0079] SEQ ID NO: 71 is the nucleotide sequence of construct 26438.

[0080] SEQ ID NO: 72 is the nucleotide sequence of construct 26553.

[0081] SEQ ID NO: 73 is the nucleotide sequence of construct 27215.

[0082] SEQ ID NO: 74 is the amino acid sequence of Mb2Cas12a Nuc1-Lb.

[0083] SEQ ID NO: 75 is the amino acid sequence of Mb2Cas12a Nuc1-As.

[0084] SEQ ID NO: 76 is the amino acid sequence of Mb2Cas12a Y1172N.

[0085] SEQ ID NO: 77 is the amino acid sequence of Mb2Cas12a Nuc2-Lb.

[0086] SEQ ID NO: 78 is the amino acid sequence of Mb2Cas12a Nuc2-As.

[0087] SEQ ID NO: 79 is the nucleotide sequence of construct 26446.

[0088] SEQ ID NO: 80 is the amino acid sequence of Mb2Cas12a F357W.

[0089] SEQ ID NO: 81 is the nucleotide sequence of construct 27495.

[0090] SEQ ID NO: 82 is the nucleotide sequence of construct 27745.

[0091] SEQ ID NO: 83 is the nucleotide sequence of construct 27747.

[0092] SEQ ID NO: 84 is the nucleotide sequence of construct 27501.

[0093] SEQ ID NO: 85 is the amino acid sequence of Mb2Cas12a F547Y.

[0094] SEQ ID NO: 86 is the amino acid sequence of Mb2Cas12a A742S.

[0095] SEQ ID NO: 87 is the amino acid sequence of Mb2Cas12a Y819F.

[0096] SEQ ID NO: 88 is the amino acid sequence of Mb2Cas12a H939Q.

[0097] SEQ ID NO: 89 is the amino acid sequence of Mb2Cas12a BH2-As+D172R.

[0098] SEQ ID NO: 90 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W.

[0099] SEQ ID NO: 91 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W+D172R.

[0100] SEQ ID NO: 92 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W+V921K.

[0101] SEQ ID NO: 93 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W+V921K+D172R.

[0102] SEQ ID NO: 94 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W+V921K+E797A.

[0103] SEQ ID NO: 95 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W+V921K+E797A+D172R.

[0104] SEQ ID NO: 96 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W+V921K+E797A+WED-MR1.

[0105] SEQ ID NO: 97 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W+V921K+E797A+WED-MR1+D172R.

[0106] SEQ ID NO: 98 is the amino acid sequence of Mb2Cas12a BH2-As+E797A.

[0107] SEQ ID NO: 99 is the amino acid sequence of Mb2Cas12a BH2-As+E797A+D172R.

[0108] SEQ ID NO: 100 is the amino acid sequence of Mb2Cas12a BH2-As+E797A+F357W.

[0109] SEQ ID NO: 101 is the amino acid sequence of Mb2Cas12a BH2-As+E797A+F357W+D172R.

[0110] SEQ ID NO: 102 is the amino acid sequence of Mb2Cas12a V921K+E797A.

[0111] SEQ ID NO: 103 is the amino acid sequence of Mb2Cas12a V921K+E797A+D172R.

[0112] SEQ ID NO: 104 is the amino acid sequence of Mb2Cas12a V921K+E797A+F357W.

[0113] SEQ ID NO: 105 is the amino acid sequence of Mb2Cas12a V921K+E797A+F357W+D172R.

[0114] SEQ ID NO: 106 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W+E797A.

[0115] SEQ ID NO: 107 is the amino acid sequence of Mb2Cas12a BH2-As+Nuc1-As+F357W+E797A+D172R.

[0116] SEQ ID NO: 108 is the nucleotide sequence of construct 27218.

[0117] SEQ ID NO: 109 is the nucleotide sequence of construct 27025.

[0118] SEQ ID NO: 110 is the nucleotide sequence of construct 27216.

[0119] SEQ ID NO:111 is the nucleotide sequence of construct 27219.

[0120] SEQ ID NO: 112 is the nucleotide sequence of construct 27220.

[0121] SEQ ID NO: 113 is the nucleotide sequence of construct 27225.

[0122] SEQ ID NO: 114 is the nucleotide sequence of construct 27228.

[0123] SEQ ID NO: 115 is the nucleotide sequence of construct 27223.

[0124] SEQ ID NO: 116 is the nucleotide sequence of construct 27224.

[0125] SEQ ID NO: 117 is the nucleotide sequence of construct 27310.

[0126] SEQ ID NO: 118 is the nucleotide sequence of construct 27311.

[0127] SEQ ID NO: 119 is the nucleotide sequence of construct 27320.

[0128] SEQ ID NO: 120 is the nucleotide sequence of construct 27321.

[0129] SEQ ID NO: 121 is the nucleotide sequence of construct 27322.

[0130] SEQ ID NO: 122 is the nucleotide sequence of construct 27381.

[0131] SEQ ID NO: 123 is the nucleotide sequence of construct 27382.

[0132] SEQ ID NO: 124 is the nucleotide sequence of construct 27383.

[0133] SEQ ID NO: 125 is the nucleotide sequence of construct 27325.

[0134] SEQ ID NO: 126 is the nucleotide sequence of construct 27323.

[0135] SEQ ID NO: 127 is the amino acid sequence of Mb2Cas12a I914K+L917V+N918A.

[0136] SEQ ID NO: 128 is the nucleotide sequence of construct 27926.

[0137] SEQ ID NO:129 is the nucleotide sequence of the Arabidopsis thaliana EF-1 alpha A1 gene promoter (“prAtEF1aA1”).

[0138] SEQ ID NO: 130 is the nucleotide sequence of the Figwort mosaic virus (FMV) enhancer ("eFMV").

[0139] SEQ ID NO:131 is the nucleotide sequence of the soybean ubiquitin 1 promoter (“prGmUbi1”).

[0140] SEQ ID NO: 132 is the nucleotide sequence encoding the crRNA targeting Δ12-fatty acid desaturase II (GmFAD2). [Brief explanation of the drawings]

[0141] [Figure 1] The general structure of the expression cassette is shown, and this same structure was used for each Mb2Cas12a variant tested. DETAILED DESCRIPTION OF THE INVENTION

[0142] definition All technical and scientific terms used herein are intended to have the same meaning as commonly understood by those skilled in the art, unless otherwise defined below. References to techniques used herein are intended to refer to techniques commonly understood in the art, including variations of those techniques and / or equivalent technique substitutions that would be apparent to those skilled in the art. While the following terms are believed to be well understood by those skilled in the art, definitions are provided below to facilitate description of the subject matter of the present disclosure.

[0143] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to an "enzyme" includes any combination of two or more such molecules, and the like.

[0144] As used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.

[0145] The term "about," as used herein, refers to the normal range of error for the respective value, readily known to one of ordinary skill in the art, e.g., ±20%, ±10%, or ±5% within the intended meaning of the recited value.

[0146] As used herein, the terms "comprising" or "comprises" are open-ended. When used in the context of a subject nucleic acid (or amino acid sequence), it refers to a nucleic acid sequence (or amino acid sequence) that contains the subject sequence as a part or its entire sequence.

[0147] As used herein, the transitional phrase "consisting essentially of" means that the claim should be construed to include the specified materials or steps recited in the claim and materials or steps that do not materially affect the basic novel characteristic(s) of the claimed subject matter. Thus, it is intended that the term "consisting essentially of," when used in the claims of this disclosure, should not be construed as the equivalent of "comprise."

[0148] The term "plurality" refers to more than one entity. Thus, a "plurality of individuals" refers to at least two individuals. In some embodiments, the term "plurality" refers to more than half of a total. For example, in some embodiments, a "plurality of a population" refers to more than half of the members of the population.

[0149] The term "plant," as used herein, refers to any plant at any stage of development, particularly a seed plant. The term "plant cell," as used herein, refers to the structural and physiological unit of a plant, including a protoplast and a cell wall. A plant cell can be in the form of an isolated single cell or a cultured cell, or as part of a more highly organized unit, such as a plant tissue, a plant organ, or a whole plant. A plant cell can be derived from or part of an angiosperm or a gymnosperm. The plant cell can be a monocotyledonous plant cell (e.g., a corn cell, a rice cell, a sorghum cell, a sugarcane cell, a barley cell, a wheat cell, an oat cell, a turfgrass cell, or an ornamental herb cell) or a dicotyledonous plant cell (e.g., a tobacco cell, a pepper cell, a eggplant cell, a sunflower cell, a cruciferous plant cell, a flax cell, a potato cell, a cotton cell, a soybean cell, a sugarbeet cell, or a rapeseed cell. The term "plant cell culture," as used herein, refers to a culture of plant units at various developmental stages, such as protoplasts, cell culture cells, cells of plant tissue, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos. The term "plant tissue," as used herein, refers to a group of plant cells organized into a structural and functional unit. It includes any tissue of a plant whether in planta or in culture. This term includes, but is not limited to, whole plants, plant organs, plant Plant tissues include seeds, tissue cultures, and any group of plant cells organized into a structural and / or functional unit. When used in conjunction with or without any specific type of plant tissue, as listed above or otherwise encompassed by this definition, this term is not intended to exclude any other type of plant tissue. The term "plant part," as used herein, refers to a plant part, including single cells and cellular tissues, such as intact plant cells in a plant, cell clumps that can regenerate plants, and tissue cultures. Examples of plant parts include, but are not limited to, pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, inflorescences, fruits, stems, shoots, cuttings, and seeds, as well as single cells and tissues from pollen, ovules, egg cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, inflorescences, fruits, stems, shoots, cuttings, scions, rootstocks, seeds, protoplasts, calluses, etc.

[0150] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, these terms encompass amino acid chains of any length, including full-length proteins, in which the amino acid residues are linked by covalent peptide bonds.

[0151] The terms "nucleic acid" and "polynucleotide" are used interchangeably and, as used herein, refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) and polymers thereof in either single- or double-stranded form, as well as both sense and antisense strands of RNA, cDNA, genomic DNA, and mitochondrial DNA, as well as synthetic forms and mixed polymers of the above. In higher plants, DNA is the genetic material, while RNA is responsible for transferring the information contained within DNA to proteins. A "genome" is the entire body of genetic material contained in each cell of an organism. When RNA is described, it is understood that its corresponding cDNA is also described, and uridine is represented as thymidine. In certain embodiments, nucleotides refer to ribonucleotides, deoxynucleotides, or modified forms of either type of nucleotide, and combinations thereof. Additionally, polynucleotides disclosed herein can include either or both naturally occurring and modified nucleotides linked together by naturally occurring and / or non-naturally occurring nucleotide linkages. Nucleic acid molecules can be chemically or biochemically modified or can include non-natural or derivatized nucleotide bases, as one of ordinary skill in the art would readily understand. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications, such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, etc.), chelators, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). The above terms are intended to encompass any topological conformation, including single-stranded, double-stranded, partially duplexed, triplexed, hairpinned, circular, and padlock conformations. When referring to a nucleic acid sequence, its complement is included unless otherwise specified. Thus, a reference to a nucleic acid molecule having a particular sequence is understood to encompass its complementary strand with its complementary sequence.Nucleotide sequences are "complementary" if they specifically hybridize in solution (e.g., according to Watson-Crick base pairing rules). The term also includes codon-optimized nucleic acids that encode the same polypeptide sequence. It is also understood that nucleic acids can be crude, purified, or attached to synthetic materials, such as beads or column matrices.

[0152] The term "corresponding," in reference to nucleic acid sequences, means that when the nucleic acid sequences of a given sequence are aligned with each other, those nucleic acids "corresponding" to a given recited position in the present invention are aligned with those positions in the reference sequence, but not necessarily at their exact numerical positions relative to the particular nucleic acid sequence of the present invention. Optimal alignment of sequences for comparison can be performed by computerized implementation of known algorithms or by visual inspection. Ready-made sequence comparison and multiple sequence alignment algorithms are the Basic Local Alignment Search Tool (BLAST) and ClustalW / ClustalW2 / Clustal Omega programs, respectively, available on the Internet (e.g., the EMBL-EBI website). Other suitable programs include, but are not limited to, GAP, BestFit, Plot Similarity, and FASTA, which are part of the Accelrys GCG package available from Accelrys, Inc., San Diego, Calif., United States of America. See also Smith & Waterman, 1981; Needleman & Wunsch, 1970; Pearson & Lipman, 1988; Ausubel et al., 1988 and Sambrook & Russell, 2001.

[0153] Unless otherwise indicated, a particular nucleic acid sequence implicitly encompasses its conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences along with the explicitly indicated sequence. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. See Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994).

[0154] The terms "identity" or "substantial identity," when used in reference to a polynucleotide or polypeptide sequence described herein, refer to a sequence having at least 60% sequence identity with a reference sequence. Alternatively, the percent identity can be any integer between 60% and 100%. Exemplary embodiments include at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity when compared to a reference sequence using a program described herein, preferably BLAST with standard parameters as described below. One of skill in the art will recognize that these values ​​can be appropriately adjusted to determine the corresponding identity of proteins encoded by two nucleotide sequences by taking into account codon degeneracy, amino acid similarity, reading frame alignment, and the like.

[0155] For sequence comparison, typically one sequence serves as a reference sequence, and test sequence is compared with it.When using sequence comparison algorithm, test sequence and reference sequence are input into computer, and partial sequence coordinates are designated as necessary, and sequence algorithm program parameters are designated.Default program parameters can be used, or alternative parameters can be designated.Then, sequence comparison algorithm calculates the sequence identity percentage of test sequence to reference sequence based on program parameters.

[0156] As used herein, the term "comparison window" refers to any segment of a number of consecutive positions selected from the group consisting of 20 to 600, generally about 50 to about 200, and more generally about 100 to about 150, where a sequence can be compared to a reference sequence of the same number of consecutive positions after the two sequences are optimally aligned. Methods for aligning sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be performed using the local homology algorithm of Smith and Waterman Add. APL. Math. 2:482 (1981), the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970), the search for similarity method of Pearson and Lipman Proc. Natl. Acad. Sci. (USA) 85:2444 (1988), computer implementations of these algorithms (e.g., BLAST), or manual alignment and visual inspection.

[0157] A "gene" is a defined region located within a genome that contains, in addition to the aforementioned coding nucleic acid sequence, other primarily regulatory nucleic acid sequences involved in the expression of the coding portion, i.e., the control of transcription and translation. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or a specific protein containing regulatory sequences. A gene may or may not be usable for the production of a functional protein. In some embodiments, a gene refers to only the coding region. The term "native gene" refers to a gene as found in nature. The term "chimeric gene" refers to any gene that 1) contains regulatory and coding sequences that are not found together in nature, or 2) encodes portions of a protein that are not naturally contiguous, or 3) contains portions of a promoter that are not naturally contiguous. Thus, a chimeric gene can contain regulatory and coding sequences from different sources, or regulatory and coding sequences from the same source but arranged in a manner different from that found in nature. A gene may be "isolated," which refers to a nucleic acid molecule that is substantially or essentially free from components normally found associated with the nucleic acid molecule in nature, including other cellular material, culture medium from recombinant production, and / or various chemicals used to chemically synthesize the nucleic acid molecule.

[0158] An "isolated" nucleic acid molecule or nucleotide sequence, or an "isolated" polypeptide, is a nucleic acid molecule, nucleotide sequence, or polypeptide that exists apart from its natural environment by the hand of man and / or has a different, modified, regulated, and / or altered function compared to its function in its natural environment, and is therefore not a product of nature. An isolated nucleic acid molecule or isolated polypeptide can exist in purified form or can exist in a non-native environment (e.g., a recombinant host cell). Thus, for example, with respect to a polynucleotide, the term isolated means that it is separated from the chromosome and / or cell in which it occurs in nature. A polynucleotide is also isolated if it is separated from the chromosome and / or cell in which it occurs in nature and then inserted into a genetic context, chromosome, chromosomal location, and / or cell that does not occur in nature. Recombinant nucleic acid molecules and nucleotide sequences of the present invention can be considered "isolated" as defined above.

[0159] Thus, an "isolated nucleic acid molecule" or "isolated nucleotide sequence" is a nucleic acid molecule or nucleotide sequence that is not immediately adjacent to the nucleotide sequences (one at the 5' end and one at the 3' end) to which it is immediately adjacent in the naturally occurring genome of the organism from which it originates. Thus, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences immediately adjacent to the coding sequence. Thus, the term includes recombinant nucleic acids that are incorporated into, for example, a vector, autonomously replicating plasmid, or virus, or the genomic DNA of a prokaryote or eukaryote, or that exist as a separate molecule independent of other sequences (e.g., a cDNA or genomic DNA fragment produced by PCR or restriction endonuclease treatment). It also includes recombinant nucleic acids that are part of a hybrid nucleic acid molecule that encodes an additional polypeptide or peptide sequence. An "isolated nucleic acid molecule" or "isolated nucleotide sequence" can include nucleotide sequences that are derived from and inserted into the same natural cell type of origin, but that exist in a non-natural state, e.g., in a different copy number and / or under the control of regulatory sequences that differ from those found in the nucleic acid molecule's natural state.

[0160] The term "isolated" can further refer to a nucleic acid molecule, nucleotide sequence, polypeptide, peptide, or fragment (e.g., when produced by recombinant DNA technology) or chemical precursor or other chemical (e.g., when chemically synthesized) that is substantially free of cellular material, viral material, and / or culture medium. Furthermore, an "isolated fragment" is a fragment of a nucleic acid molecule, nucleotide sequence, or polypeptide that is not naturally occurring as a fragment and, as such, would not be found in the natural state. "Isolated" does not necessarily mean that the preparation is technically pure (homogeneous), but is sufficiently pure to provide the polypeptide or nucleic acid in a form that can be used for its intended purpose.

[0161] "Homology-dependent repair," or "homologous recombination repair," or "HDR" refers to a mechanism for repairing ssDNA and double-stranded DNA (dsDNA) damage in cells. This repair mechanism can be used by cells when there is an HDR template with a sequence highly homologous to the damaged site. The term "complete HDR" refers to a situation in which the genomic homology junction in the replaced allele has undergone complete HDR, while "incomplete HDR" refers to a situation in which the genomic homology junction in the replaced allele has undergone partial or incomplete HDR. A donor DNA molecule with homology to the cleaved target DNA sequence is used as a template for repair of the cleaved target DNA sequence, resulting in the transfer of genetic information from the donor polynucleotide to the target DNA. Thus, new nucleic acid material can be inserted / copied into the site. Optionally, the target DNA is contacted with a donor molecule, e.g., a donor DNA molecule. Optionally, the donor DNA molecule is introduced into the cell. Optionally, at least a segment of the donor DNA molecule is integrated into the genome of the cell.

[0162] "Microhomology-mediated end joining," or "MMEJ," or "alternative non-homologous end joining" (Alt-NHEJ) refers to a form of double-strand break repair in DNA. This repair mechanism utilizes microhomology sequences to align the broken strands. "Non-homologous end joining" or "NHEJ" refers to a form of double-strand break repair in DNA. The double-strand break is repaired by directly ligating the broken ends to each other. Generally, there is no insertion of new nucleic acid material at the site, although some nucleic acid material may be lost or added, resulting in small deletions or small insertions.

[0163] Proteins provided herein include site-specific polypeptides. Site-specific modifying polypeptides modify target DNA (e.g., by cleaving or methylating the target DNA) and / or modify polypeptides associated with the target DNA (e.g., methylation or acetylation of histone tails). In some embodiments, the site-specific modifying polypeptide interacts with a guide RNA, either a single RNA molecule or an RNA duplex of at least two RNA molecules, and is guided to a DNA sequence (e.g., a chromosomal sequence or an extrachromosomal sequence, such as an episomal sequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.) by its association with the guide RNA. In some embodiments, the site-specific polypeptide is a site-specific nuclease, which is capable of cleaving one or both strands of DNA at a designated target sequence.

[0164] The term "cleavage" or "cleaving" refers to the breaking of the covalent phosphodiester bond in the ribosyl phosphate diester backbone of a polynucleotide and encompasses both single-strand and double-strand breaks. Double-strand breaks can occur as a result of two separate single-strand cleavage events. Cleavage can result in the creation of either blunt ends or overhanging ends (also known as sticky ends). A "nuclease cleavage site" or "genomic nuclease cleavage site" is a region of nucleotides within which a site-specific nuclease cleaves (e.g., upon binding to a proximal binding site). When the polynucleotide is DNA (e.g., genomic DNA), one or both strands can be cleaved at the nuclease cleavage site. Such cleavage by nuclease enzymes triggers DNA repair mechanisms within the cell, thereby establishing an environment for homologous recombination to occur.

[0165] The site-specific nuclease may be a naturally occurring site-specific nuclease. Exemplary naturally occurring site-specific nucleases are known in the art (see, e.g., Makarova et al., 2017, Cell 168:328-328.e1 and Shmakov et al., 2017, Nat Rev Microbiol 15(3):169-182, both of which are incorporated herein by reference). In some embodiments, the site-specific nuclease binds to a DNA-targeting polynucleotide (e.g., a guide RNA), thereby being guided to a specific sequence within the target DNA and cleaving the target DNA.

[0166] In some embodiments, a site-specific nuclease is modified from its native sequence (e.g., by mutation or one or more amino acid residues) to alter its function. For example, a site-specific nuclease can be modified to be enzymatically inactive. The term "enzymatically inactive" can refer to a site-specific nuclease that can bind to a nucleic acid sequence in a polynucleotide in a sequence-specific manner but does not cleave the target polynucleotide. A polypeptide that targets an enzymatically inactive site can include an enzymatically inactive domain (e.g., a nuclease domain). Enzymatically inactive can refer to no activity. Enzymatically inactive can refer to substantially no activity. Enzymatically inactive can refer to essentially no activity. Enzymatically inactive can refer to 1% or less, 2% or less, 3% or less, 4% or less, 5% or less, 6% or less, 7% or less, 8% or less, 9% or less, or 10% or less of the activity compared to the exemplary wild-type activity.

[0167] In some embodiments, the site-specific nuclease comprises a CRISPR-associated (Cas) protein or Cas nuclease that functions in a CRISPR (clustered regularly interspaced short palindromic repeats) / Cas system. In bacteria, this system can provide adaptive immunity to foreign DNA (Barrangou, R., et al., "CRISPR provides acquired resistance against viruses in prokaryotes," Science (2007) 315:1709-1712; Makarova, K.S., et al., "Evolution and classification of the CRISPR-Cas systems," Nat Rev Microbiol (2011) 9:467-477; Garneau, J.E., et al., "The CRISPR / Cas bacterial immune system cleaves bacteriophage and plasmid DNA," Nature (2010) 468:67-71; Sapranauskas, R., et al., "The Streptococcus thermophilus CRISPR / Cas system provides immunity in Escherichia coli," Nucleic Acids Res (2011) 39:9275-9282). CRISPR / Cas systems (e.g., modified and / or unmodified) can be utilized as genome engineering tools in a wide variety of organisms, including various mammals, animals, plants, microorganisms, and yeast. CRISPR / Cas systems can include a guide nucleic acid, such as a guide RNA (gRNA), complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid editing. RNA-guided Cas proteins (e.g., Cas nucleases, such as Cas9 nuclease) can specifically bind to target polynucleotides (e.g., DNA) in a sequence-dependent manner.Cas proteins can cleave DNA when they possess nuclease activity (Gasiunas, G., et al., “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria,” Proc Natl Acad Sci USA (2012) 109:E2579-E286; Jinek, M., et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science (2012) 337:816-821; Sternberg, SH, et al., “DNA interrogation by the CRISPR RNA-guided endonuclease Cas9,” Nature (2014) 507:62; Deltcheva, E., et al., “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III,” Nature (2011) 471:602-607). DNA breaks (e.g., double-strand breaks) can result in DNA break repair, allowing for the introduction of one or more genetic modifications (e.g., nucleic acid editing). DNA break repair can occur by non-homologous end joining (NHEJ), microhomology-mediated end joining (MMEJ), or homology-directed repair (HDR). In some embodiments, donor nucleic acids are used to facilitate HDR, as described in more detail in the "Systems" section below.The CRISPR-Cas system has been widely used for programmable genome editing in various organisms and model systems (Cong, L., et al., "Multiplex genome engineering using CRISPR-Cas systems," Science (2013) 339:819-823; Jiang, W., et al., "RNA-guided editing of bacterial genomes using CRISPR-Cas systems," Nat. Biotechnol. (2013) 31:233-239; Sander, JD & Joung, JK, "CRISPR-Cas systems for editing, regulating, and targeting genomes," Nature Biotechnol. (2014) 32:347-355).

[0168] In some embodiments, the site-specific nucleases described herein comprise a Cas protein complexed with a guide nucleic acid, such as a guide RNA (further described in the "Systems" section below). In some embodiments, the site-specific nuclease comprises a Cas protein complexed with a single guide nucleic acid, such as a single guide RNA (sgRNA). In some embodiments, the site-specific nuclease comprises an RNA-binding protein (RBP), optionally complexed with a guide nucleic acid, such as a guide RNA (e.g., sgRNA), capable of forming a complex with the Cas protein. In some examples, the RNA-guided Cas protein recognizes a DNA target complementary to a portion of the gRNA known as the CRISPR RNA (crRNA) sequence. The target sequence is often referred to as the protospacer, and the portion of the crRNA sequence complementary to the protospacer is often referred to as the spacer. To function (e.g., to cleave DNA), many Cas nucleases also require a specific protospacer adjacent motif (PAM), a DNA sequence of approximately 2-6 base pairs immediately following the protospacer sequence.

[0169] Cas proteins from various species (e.g., those disclosed in Shmakov et al., 2017, or polypeptides derived therefrom) may require different PAM sequences in target DNA. Therefore, for a particular Cas enzyme of choice, the PAM sequence requirements may differ from the 5'-N GG-3' sequence (where N is A, T, C, or G) known to be required for Cas9 activity. Numerous Cas9 orthologs have been identified from a wide variety of species, and these proteins share only a few identical amino acids. All identified Cas9 orthologs have the same domain architecture, including a central HNH endonuclease domain and split RuvC / RNase H domains. Cas9 proteins share four key motifs of conserved architecture: motifs 1, 2, and 4 are RuvC-like motifs, while motif 3 is an HNH motif. In contrast, Cas12a proteins from various species may have different PAM sequence requirements compared to the standard PAM of TTTV LbCas12a.

[0170] Any suitable CRISPR / Cas system can be used. CRISPR / Cas systems can be referred to using various nomenclature systems. Exemplary nomenclature systems are provided in Makarova, K. Set al., "An updated evolutionary classification of CRISPR-Cas systems," Nat Rev Microbiol (2015) 13:722-736 and Shmakov, S. et al., "Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems," Mol Cell (2015) 60:1-13. A CRISPR / Cas system can be a Type I, Type II, Type III, Type IV, Type V, Type VI system, or any other suitable CRISPR / Cas system. As used herein, a CRISPR / Cas system can be a Class 1, Class 2, or any other suitable classified CRISPR / Cas system. The determination of Class 1 or Class 2 can be based on the genes encoding the effector modules. Class 1 systems generally have a multi-subunit crRNA-effector complex, while Class 2 systems generally have a single protein, such as Cas9, Cpfl, C2c1, C2c2, C2c3, or a crRNA-effector complex. Class 1 CRISPR / Cas systems may use a complex of multiple Cas proteins to effect regulation. Class 1 CRISPR / Cas systems may include, for example, Type I (e.g., Type I, IA, IB, IC, ID, IE, IF, IU), Type III (e.g., Type III, IIIA, IIIB, IIIC, IIID), and Type IV (e.g., Type IV, IVA, IVB) CRISPR / Cas types. Class 2 CRISPR / Cas systems may use a single large Cas protein to effect regulation. Class 2 CRISPR / Cas systems may include, for example, Type II (e.g., Type II, IIA, IIB) and Type V CRISPR / Cas types.CRISPR systems may complement each other and / or provide functional units in trans to facilitate CRISPR gene localization.

[0171] The Cas protein may be derived from any suitable organism, including, but not limited to, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacteria, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii watsonii), Cyanothece sp.), Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium ebestigatum evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp.Examples of suitable organisms include Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, and Francisella novicida. In some embodiments, the organism is Streptococcus pyogenes (S. pyogenes). In some embodiments, the organism is Staphylococcus aureus (S. aureus). In some embodiments, the organism is Streptococcus thermophilus (S. thermophilus).

[0172] Cas proteins are useful in, but not limited to, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, and the like. pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium doricumdolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractor salsagainis salsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygiisyzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinella succinogenes, Campylobacter jejuni subsp. jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum labamentivorans lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. multocida, Sutterella wadsworthensis, Proteobacteria, Legionella pneumophila, Parasterella excrementihominis, Wolinella succinogenes, and Francisella novicida.

[0173] Non-limiting examples of Cas proteins include c2c1, C2c2, c2c3, Casl, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cash, Cas6e, Cas6f, Cas7, Cas8a, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Cpfl, Csyl, Csy2, Csy3, Csel (CasA), and Cse2. (CasB), Cse3 (CasE), Cse4 (CasC), Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx15, Csfl, Csf2, Csf3, Csf4 and Cul966, and homologs or modified versions thereof. In some embodiments, the site-specific nuclease of the fusion proteins provided herein comprises a CRISPR-associated nuclease, wherein the CRISPR-associated nuclease is Cas5, Cas6, Cas7, Cas8, Cas9, Cas12a, Cas12b, Cas12i, Cas12j, Cas12L, Cas12e, Cas12c, Cas12d, Cas12g, Cas12h, TnpB, Cas13a, Cas13b, or Cas14. In some embodiments, the CRISPR-associated nuclease is a Cas9 enzyme. In some embodiments, the CRISPR-associated nuclease is a Cas12a enzyme. In some embodiments, the CRISPR-associated nuclease is a nickase or an inactivated version of a CRISPR-associated nuclease.

[0174] Lachnospiraceae bacterium Cpf1 (LbCpf1) is one of a large group of many Cpf1 proteins. The terms "Cpf1" and "Cas12a" are used interchangeably throughout this disclosure. Cpf1 is a Cas protein. In some embodiments, the site-specific nuclease is catalytically inactive Cas12a from Lachnospiraceae bacterium ("dLbCas12a"). In other embodiments, the site-specific nuclease is catalytically active Cas12a from Lachnospiraceae bacterium ("LbCas12a") or Moraxella bovoculi AAX08_00205 ("Mb2Cas12a"). In some embodiments, the site-specific nuclease domain of the fusion protein is a Cas12a protein from any of Lachnospiraceae bacterium, Acidaminococcus sp., Moraxella bovoculi, Thiomicrospira sp., Moraxella lacunata, Methanomethylophilus alvus, Butyrivibrio sp., or Bacteroidetes oral sp.

[0175] As used herein, the term "domain" refers to a discrete, independently folded unit of amino acid residues. Domain size varies from about 25-30 amino acid residues to about 300 residues, depending on its function and source organism, with the average across domains being about 100 residues. Large proteins usually contain two or more domains. A domain can consist of a combination of motifs. In globular proteins, motifs can be understood as segments of alpha-helices and / or beta-strands connected by loops to form a repeating pattern. See generally PRINCIPLES OF BIOCHEMISTRY (2nd ed.), 92-96.

[0176] A Cas protein may contain one or more domains. Non-limiting examples of domains include a guide nucleic acid recognition and / or binding domain, a nuclease domain (e.g., DNase or RNase domain, RuvC, HNH), a DNA-binding domain, an RNA-binding domain, a helicase domain, a protein-protein interaction domain, and a dimerization domain. The guide nucleic acid recognition and / or binding domain can interact with the guide nucleic acid. The nuclease domain can contain catalytic activity for nucleic acid cleavage. The nuclease domain can lack catalytic activity to prevent nucleic acid cleavage. The Cas protein can be a chimeric Cas protein fused with another protein or polypeptide. For example, the Cas protein can be a chimera of various Cas proteins containing domains from different Cas proteins.

[0177] As used herein, "domain swap" refers to replacing one domain with another. The replaced domain can be a recognized domain, a predicted domain, a microregion, or a motif. As used herein, a "microregion" refers to a region of at least two amino acid residues within a domain. By way of example and not limitation, domains that can be subject to domain swapping include WED-MR1, WED-MR2, BH, Up-seq, or Nuc, or a combination thereof. As a further example, one skilled in the art can clone the LbCas12a WED-MR1 coding domain into a nucleotide sequence encoding Mb2Cas12a, thereby replacing the corresponding Mb2Cas12a WED-MR1 domain.

[0178] As used herein, a Cas protein can be an active variant, an inactive variant, or a fragment of a wild-type or modified Cas protein. The Cas protein can contain amino acid changes, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof, compared to the wild-type version of the Cas protein. The Cas protein can be a polypeptide with at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or similarity to an exemplary wild-type Cas protein. The Cas protein can be a polypeptide with at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or similarity to an exemplary wild-type Cas protein. A variant or fragment may comprise at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or similarity to a wild-type or modified Cas protein or portion thereof. A variant or fragment may lack nucleic acid cleavage activity while being capable of being complexed with a guide nucleic acid and targeted to a nucleic acid locus.

[0179] Cas proteins can be modified to optimize regulation of gene expression. Cas proteins can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to alter any other activity or property of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, or the Cas protein can be truncated to remove domains that are not essential for protein function or to optimize (e.g., enhance or decrease) the activity of the Cas protein for regulating gene expression.

[0180] One or more nuclease domains of a Cas protein (e.g., RuvC, HNH) may be deleted or mutated so that they are no longer functional or contain reduced nuclease activity. For example, in a Cas protein containing at least two nuclease domains (e.g., Cas9), if one of the nuclease domains is deleted or mutated, the resulting Cas protein, known as a nickase, can generate a single-strand break, rather than a double-strand break, at the CRISPR RNA (crRNA) recognition sequence within double-stranded DNA. Such a nickase may be capable of cleaving either the complementary or non-complementary strand, but not both. In some embodiments, the targeting specificity of the double-strand break is improved by targeting the nickase to opposite strands at two nearby loci. If the nickase cleaves a single strand at both loci, a double-strand break is formed and can be repaired by HR as described herein. When all of the nuclease domains of a Cas protein (e.g., both the RuvC and HNH nuclease domains in the Cas9 protein, or the RuvC nuclease domain in the Cpfl protein) are deleted or mutated, the resulting Cas protein may have reduced or no ability to cleave both strands of double-stranded DNA.

[0181] Variants of the polypeptides of the present disclosure are also provided herein. Unless otherwise specified, polypeptide variants retain their respective biological activity. For example, variants of site-specific nuclease polypeptides retain the biological function of full-length native sequence site-specific nucleases. In another example, variants of non-specific endo-processing enzymes retain the biological function of full-length native sequence non-specific endo-processing enzymes.

[0182] Modifications to any of the polypeptides or proteins provided herein can be made by known methods. For example, modifications can be made by site-directed mutagenesis of nucleotides in a nucleic acid encoding the polypeptide, thereby generating DNA encoding the modification, which can then be expressed in a recombinant cell culture to produce the encoded polypeptide. Techniques for making substitution mutations at predetermined sites in DNA having a known sequence are well known. For example, M13 primer mutagenesis and PCR-based mutagenesis methods can be used to make one or more substitution mutations. Any of the nucleic acid sequences provided herein can be codon-optimized to alter, e.g., maximize, expression in a host cell or organism.

[0183] The amino acids in the polypeptides described herein can be any of the 20 naturally occurring amino acids, D stereoisomers of naturally occurring amino acids, unnatural amino acids, and chemically modified amino acids. Unnatural amino acids (i.e., those not naturally found in proteins) are also known in the art, as described, for example, in Zhang et al. "Protein engineering with unnatural amino acids," Curr. Opin. Struct. Biol. 23(4):581-587 (2013); Xie et al. "Adding amino acids to the genetic repertoire," 9(6):548-54 (2005)) and all references cited therein. β- and γ-amino acids are known in the art and are also contemplated herein as unnatural amino acids.

[0184] As used herein, a chemically modified amino acid refers to an amino acid whose side chain has been chemically modified. For example, the side chain can be modified to include a signaling moiety, such as a fluorophore or radiolabel. The side chain can also be modified to include a new functional group, such as a thiol, a carboxylic acid, or an amino group. Post-translationally modified amino acids are also included in the definition of chemically modified amino acids.

[0185] Conservative amino acid substitutions are also contemplated. For example, conservative amino acid substitutions may be made at one or more amino acid residues, such as one or more lysine residues, of any of the polypeptides provided herein. Those skilled in the art will recognize that a conservative substitution is the replacement of an amino acid residue with another amino acid residue that is biologically and / or chemically similar. The following eight groups each contain amino acids that are conservative substitutions for each other: 1) Alanine (A), Glycine (G), 2) Aspartic acid (D), glutamic acid (E), 3) Asparagine (N), Glutamine (Q), 4) Arginine (R), Lysine (K), 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V), 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W), 7) serine (S), threonine (T), and 8) Cysteine ​​(C), methionine (M).

[0186] For example, when arginine is substituted with serine, conservative substitutions of serine (e.g., threonine) are also contemplated. Non-conservative substitutions, such as lysine with asparagine, are also contemplated.

[0187] Also provided are DNA constructs comprising a promoter operably linked to a recombinant nucleic acid encoding a fusion protein or domain thereof as described herein. A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. Many promoters can be used in the constructs described herein. A promoter is a region or sequence located upstream and / or downstream of the start of transcription that is involved in the recognition and binding of RNA polymerase and other proteins to initiate transcription.

[0188] The term "promoter," as used herein, refers to a nucleotide sequence that controls expression of a coding sequence by providing recognition for RNA polymerase and other factors necessary for proper transcription, typically located upstream (5') of that coding sequence. A "promoter regulatory sequence" consists of proximal and more distal upstream elements. Promoter regulatory sequences affect transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, untranslated leader sequences, introns, and polyadenylation signal sequences. These include natural and synthetic sequences, as well as sequences that may be a combination of natural and synthetic sequences. An "enhancer" is a DNA sequence that can stimulate promoter activity and may be an intrinsic element of the promoter or a heterologous element inserted to increase the level or tissue specificity of the promoter. It can operate in both orientations (e.g., forward or reverse) and can function when moved upstream or downstream from the promoter. The term "promoter" includes "promoter regulatory sequence."

[0189] The choice of which promoter to include depends on several factors, including, but not limited to, efficiency, selectability, inducibility, desired expression level, and cell- or tissue-preferred expression. It is routine for one of ordinary skill in the art to regulate the expression of a sequence by appropriate selection and placement of promoters and other regulatory regions relative to that sequence.

[0190] Certain promoters have been shown to be capable of directing RNA synthesis at higher rates than others. These are called "strong promoters." Certain other promoters have been shown to direct RNA synthesis at higher levels only in certain types of cells or tissues; when a promoter preferentially directs RNA synthesis to a particular tissue (whereas lower levels of RNA synthesis may occur in other tissues), it is often referred to as a "tissue-specific promoter" or "tissue-preferred promoter." Because the expression pattern of a chimeric gene (or genes) introduced into plants is controlled using a promoter, there is continuing interest in isolating new promoters capable of controlling the expression of a chimeric gene (or genes) to consistent levels in specific tissue types or at specific plant developmental stages.

[0191] Certain promoters are capable of directing relatively similar levels of RNA synthesis in all tissues of a plant. These are called "constitutive promoters" or "tissue-independent" promoters. Constitutive promoters can be divided into strong, intermediate, and weak categories based on their effectiveness in directing RNA synthesis. Constitutive promoters are particularly useful in this regard, since in many cases, a chimeric gene (or genes) must be expressed simultaneously in different plant tissues to obtain the desired function of the gene (or genes). Although many constitutive promoters have been discovered and characterized from plants and plant viruses, there is still ongoing interest in isolating more novel constitutive promoters, synthetic or natural, capable of controlling the expression of a chimeric gene (or genes) at different levels and expression levels of multiple genes in the same transgenic plant for gene stacking.

[0192] The recombinant nucleic acids provided herein can be included in an expression cassette for expression in a host cell or organism of interest. The cassette will include 5' and 3' regulatory sequences operably linked to the recombinant nucleic acids provided herein, allowing for expression of the fusion protein. The cassette may additionally contain at least one additional gene or genetic element to be cotransformed into the cell or organism. When additional genes or elements are included, the components are operably linked. Alternatively, the additional genes or elements can be provided on multiple expression cassettes. Such expression cassettes comprise multiple restriction and / or recombination sites for insertion of polynucleotides under the transcriptional control of the regulatory regions. The expression cassette may additionally contain a selectable marker gene. The expression cassette will include, in the 5' to 3' transcriptional direction, a transcriptional and translational initiation region (i.e., promoter) functional in the cell or organism of interest, a polynucleotide of the invention, and a transcriptional and translational termination region (i.e., termination region). A promoter of the present invention is capable of directing or driving expression of a coding sequence (i.e., a nucleic acid sequence that is transcribed into RNA, such as mRNA, rRNA, tRNA, snRNA, ncRNA, lncRNA, sense RNA, or antisense RNA, whether or not that RNA is then translated to produce a protein) in a host cell. The regulatory regions (i.e., promoter, transcriptional regulatory region, and translation termination region) can be endogenous or heterologous to the host cell or to each other. As used herein, "heterologous" with respect to a sequence is a sequence that is derived from a foreign species or, if derived from the same species, has been substantially altered from its natural form in composition and / or genomic locus by deliberate human intervention.

[0193] Additional regulatory signals include, but are not limited to, a start site for transcription initiation, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, termination signals, etc. See Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY); Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and references cited therein.

[0194] The expression cassette may also contain a selectable marker gene for selecting transformed cells. Marker genes include genes that confer antibiotic resistance, such as those that confer hygromycin resistance, ampicillin resistance, gentamicin resistance, and neomycin resistance, to name a few. Additional selectable markers are known, and any may be used.

[0195] In preparing expression cassettes, various DNA fragments can be manipulated to provide the DNA sequences in the proper orientation and, if necessary, in the proper reading frame. To this end, adapters or linkers can be used to join the DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. To this end, in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, such as transitions and transversions, can be involved.

[0196] Also provided are vectors containing the recombinant nucleic acids or DNA constructs described herein. It is contemplated that the vectors have the necessary functional elements to direct and regulate the transcription of the inserted nucleic acid. Such functional elements include, but are not limited to, a promoter, a region upstream or downstream of the promoter, such as an enhancer that can regulate the transcriptional activity of the promoter, an origin of replication, a restriction site suitable for facilitating the cloning of an insert adjacent to the promoter, an antibiotic resistance gene or other marker that can be useful for selecting cells containing the vector or a vector containing the insert, an RNA splice junction, a transcription termination region, or any other region that can be useful for promoting the expression of the inserted gene or hybrid gene. Generally, the elements described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, 2012. The vector can be, for example, a plasmid.

[0197] Cellular transformation can be stable or transient. Thus, the transgenic cells, plant cells, plants, and / or plant parts of the present invention can be stably transformed or transiently transformed. Transformation can refer to the transfer of a nucleic acid molecule into the genome of a host cell, resulting in stable genetic inheritance. In some embodiments, introduction into the plant, plant part, and / or plant cell is via bacterial-mediated transformation, particle bombardment transformation, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, liposome-mediated transformation, nanoparticle-mediated transformation, polymer-mediated transformation, virus-mediated nucleic acid delivery, whisker-mediated nucleic acid delivery, microinjection, sonication, infiltration, polyethylene glycol-mediated transformation, protoplast transformation, or any other electrical, chemical, physical, and / or biological mechanism, or any combination thereof, that results in the introduction of a nucleic acid into the plant, plant part, and / or cell thereof.

[0198] Plant transformation procedures are well known and routinely performed in the art and are described throughout this specification. Non-limiting examples of plant transformation methods include transformation by bacterial-mediated nucleic acid delivery (e.g., by bacteria from the genus Agrobacterium), virus-mediated nucleic acid delivery, silicon carbide or nucleic acid whisker-mediated nucleic acid delivery, liposome-mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, infiltration, PEG-mediated nucleic acid uptake, and any other electrical, chemical, physical (mechanical), and / or biological mechanism that results in the introduction of nucleic acid into plant cells, including any combination thereof. General guidelines for various plant transformation methods known in the art include Miki et al. ("Procedures for Introducing Foreign DNA into Plants" in Methods in Plant Molecular Biology and Biotechnology, Glick, BR and Thompson, JE, Eds. (CRC Press, Inc., Boca Raton, 1993), pages 67-88) and Rakowoczy-Trojanowska (Cell. Mol. Biol. Lett. 7:849-858 (2002)).

[0199] Agrobacterium-mediated transformation is a commonly used method for transforming plants due to its high transformation efficiency and its versatility with many different species. Agrobacterium-mediated transformation typically involves transferring a binary vector carrying the foreign DNA of interest into a suitable Agrobacterium strain, which may depend on the complement of vir genes carried by the host Agrobacterium strain either on a coexisting Ti plasmid or on the chromosome (Uknes et al. 1993, Plant Cell 5:159-169). Introduction of the recombinant binary vector into Agrobacterium can be achieved by a triparental mating procedure using Escherichia coli carrying the recombinant binary vector and a helper E. coli strain carrying a plasmid capable of mobilizing the recombinant binary vector into the target Agrobacterium strain. Alternatively, the recombinant binary vector can be transferred to Agrobacterium by nucleic acid transformation (Hoefgen and Willmitzer 1988, Nucleic Acids Res 16:9877).

[0200] Transformation of plants with recombinant Agrobacterium usually involves co-cultivation of the Agrobacterium with an explant from the plant, followed by methods well known in the art. The transformed tissue carries an antibiotic or herbicide resistance marker between the binary plasmid T-DNA borders and is typically regenerated on selective media.

[0201] Another method for transforming plants, plant parts, and plant cells involves projecting inert or biologically active particles into plant tissues and cells. See, e.g., U.S. Patent Nos. 4,945,050, 5,036,006, and 5,100,792. Generally, this method involves projecting inert or biologically active particles into plant cells under conditions effective to penetrate the outer surface of the cells and cause their internalization. When inert particles are used, the vector can be introduced into the cells by coating the particles with a vector containing the nucleic acid of interest. Alternatively, one or more cells can be surrounded by the vector, resulting in the vector being carried into the cells following the particle. Biologically active particles (e.g., dried yeast cells, dried bacteria, or bacteriophage, each containing one or more nucleic acids to be introduced) can also be projected into plant tissue. As used herein, the phrase "biolistic transformation" refers to a method of directly introducing RNA or DNA into a cell (e.g., a plant cell) by mixing the RNA or DNA with heavy metal particles (e.g., tungsten or gold) and releasing them into the cell (e.g., a plant cell) using high-velocity pressure, allowing the RNA or DNA to penetrate the cell (e.g., penetrate the plant cell wall).

[0202] CRISPR / Cas systems can also be used to edit the genome of a host cell or organism. As detailed above, the "CRISPR / Cas" system refers to a broad class of bacterial systems for defense against foreign nucleic acids. Any of the CRISPR / Cas system components described herein can be used to introduce a fusion protein, recombinant nucleic acid, or system into the genome of a host cell or organism. CRISPR / Cas system-mediated genome editing methods are known in the art. It will be understood that the introduction of a fusion protein, recombinant nucleic acid, or system described herein into the genome of a host cell or organism using a CRISPR / Cas system will differ from the detailed methods and systems provided herein.

[0203] In another aspect, provided herein is a system useful for editing one or more nucleic acids. The system comprises one or more of the fusion proteins (or recombinant nucleic acids, constructs, vectors, or host cells) described above. In some embodiments, the system further comprises one or more additional elements useful for editing one or more nucleic acids. For example, the systems provided herein may further comprise a donor polynucleotide. As another example, a system comprising a fusion protein comprising a Cas nuclease may further comprise one or more guide nucleic acids and / or one or more donor polynucleotide sequences.

[0204] Optionally, the systems and methods described herein include at least one guide nucleic acid polynucleotide. Optionally, the systems and methods described herein include multiple guide nucleic acids. In some embodiments, the polynucleotide may be deoxyribonucleic acid (DNA). Optionally, the DNA sequence may be single-stranded or double-stranded. In some embodiments, the at least one guide nucleic acid polynucleotide may be a ribonucleic acid (guide RNA).

[0205] In some embodiments, the nuclease can be complexed with at least one guide RNA polynucleotide. The at least one guide RNA polynucleotide can include a nucleic acid targeting region that includes a sequence complementary to a nucleic acid sequence on a targeted polynucleotide, such as a targeted genomic locus or gene, to confer sequence specificity for nuclease targeting. In some embodiments, the guide nucleic acid is a single guide nucleic acid that includes a crRNA. In some embodiments, the guide nucleic acid is a single guide nucleic acid that includes a crRNA but lacks a tracrRNA. The crRNA can include a nucleic acid targeting segment (e.g., a spacer region) of the guide nucleic acid and a stretch of nucleotides that can form one half of a double-stranded duplex of the Cas protein-binding segment of the guide nucleic acid.

[0206] In some embodiments, the nucleic acid targeting region of the guide nucleic acid (e.g., spacer) is 20 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 19 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 18 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 17 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 16 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 21 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 22 nucleotides in length.

[0207] The nucleotide sequence of the guide nucleic acid complementary to the nucleotide sequence of the target nucleic acid (target sequence) can have a length of, for example, at least about 12 nucleotides (nt) to about 80 nt, about 12 nt to about 50 nt, about 12 nt to about 45 nt, about 12 nt to about 40 nt, about 12 nt to about 35 nt, about 12 nt to about 30 nt, about 12 nt to about 25 nt, about 12 nt to about 20 nt, about 12 nt to about 19 nt, about 19 nt to about 20 ... The length may be about 19 nt to about 25 nt, about 19 nt to about 30 nt, about 19 nt to about 35 nt, about 19 nt to about 40 nt, about 19 nt to about 45 nt, about 19 nt to about 50 nt, about 19 nt to about 60 nt, about 20 nt to about 25 nt, about 20 nt to about 30 nt, about 20 nt to about 35 nt, about 20 nt to about 40 nt, about 20 nt to about 45 nt, about 20 nt to about 50 nt, or about 20 nt to about 60 nt.

[0208] The protospacer sequence of a targeted polynucleotide can be identified by identifying a protospacer adjacent motif (PAM) within the region of interest and selecting a region of desired size upstream or downstream of the PAM as the protospacer. The corresponding spacer sequence can be designed by determining the complementary sequence of the protospacer region.

[0209] Spacer sequences can be identified using a computer program (e.g., machine-readable code) that can use variables such as predicted melting temperature, secondary structure formation and predicted annealing temperature, sequence identity, genomic context, chromatin exposure, %GC, genomic frequency, methylation status, and the presence of SNPs.

[0210] The percent complementarity between a nucleic acid target sequence (e.g., a spacer sequence of at least one guide polynucleotide as disclosed herein) and a target nucleic acid (e.g., a protospacer sequence of one or more target loci as disclosed herein) can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%. The percent complementarity between a nucleic acid target sequence and a target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over about 20 contiguous nucleotides.

[0211] Guide nucleic acids of the disclosed systems may contain modifications or sequences that provide additional desirable characteristics (e.g., modified or modulated stability, intracellular targeting, tracking by fluorescent labels, binding sites for proteins or protein complexes, etc.). Examples of such modifications include, for example, a 5' cap (7-methylguanylate cap (m7G)), a 3' polyadenylation tail (3' poly(A) tail), a riboswitch sequence (e.g., allowing for regulated stability and / or regulated accessibility) by proteins and / or protein complexes, a stability control sequence, a sequence that forms a dsRNA duplex (hairpin), a modification or sequence that targets the RNA to a subcellular location (e.g., nucleus, mitochondria, chloroplast, etc.), a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, etc.), a protein (e.g., a transcriptional activator, a transcriptional repressor, a DNA methyltransferase, a DNA demethylase, a histone acetyltransferase, a histone deacetylase, and combinations thereof), or a modification or sequence that provides binding sites for proteins that act on DNA, including proteins (e.g., transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and combinations thereof).

[0212] A guide nucleic acid may contain one or more modifications (e.g., base modifications, backbone modifications) to provide a nucleic acid with novel or enhanced characteristics (e.g., enhanced stability). A guide nucleic acid may contain a nucleic acid affinity tag. A nucleoside may be a base-sugar combination. The base portion of a nucleotide may be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide may be a nucleoside further comprising a phosphate group covalently linked to the sugar portion of the nucleoside. For nucleosides that include a pentofuranosyl sugar, the phosphate group may be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming a guide nucleic acid, the phosphate group may covalently link adjacent nucleosides to each other to form a linear polymeric compound. The respective ends of this linear polymeric compound can then be further joined to form a circular compound, although linear compounds may be preferred. In addition, linear compounds can have internal nucleotide base complementarity and thus can fold in such a way as to yield fully or partially double-stranded compounds. Furthermore, within a guide nucleic acid, the phosphate groups can generally be referred to as forming the internucleoside backbone of the guide nucleic acid. The linkage or backbone of the guide nucleic acid can be a 3'→5' phosphodiester linkage.

[0213] In some embodiments, the at least one guide RNA polynucleotide of the systems or methods provided herein is capable of binding to at least a portion of a genome (e.g., a plant genome) or a gene (e.g., a plant gene). Optionally, the at least one guide RNA polynucleotide is capable of forming a complex with a site-specific nuclease and directing the site-specific nuclease to target a portion of the target nucleic acid (e.g., a site in the genome or gene).

[0214] In some embodiments, the systems described herein comprise at least two (e.g., at least three, at least four, at least five, or at least six) different guide RNA polynucleotides capable of forming a complex with a site-specific nuclease moiety.

[0215] The following description and examples enumerate various aspects and embodiments of the present compositions and methods. The detailed embodiments are not intended to define the scope of the present compositions and methods. Rather, the embodiments merely provide non-limiting examples of various compositions and methods that fall at least within the scope of the disclosed compositions and methods. The description should be read from the perspective of a person skilled in the art, and therefore may not necessarily include information that is known to a person skilled in the art.

[0216] One aspect of the invention is a mutant Mb2Cas12a polypeptide comprising at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO: 1. In one embodiment, the mutant Mb2Cas12a polypeptide comprises at least one amino acid substitution at a position selected from the group consisting of D172, F357, F547, A742, E797, Y819, E913, I914, L917, N918, V921, H939, and Y1172. In another embodiment, the at least one amino acid substitution is selected from the group consisting of D172R, D172K, F357W, F547Y, A742S, E797A, Y819F, I914K, L917V, N918A, N918K, V921K, V921Q, H939Q, and Y1172N. In a further embodiment, the polypeptide comprises a sequence selected from the group comprising SEQ ID NOs: 3, 5, 26, 35, 57, 58, 64, 76, 80, 85, 86, 87, 88 and 127.

[0217] Another aspect of the invention is a mutant Mb2Cas12a polypeptide comprising at least one domain swap. In one embodiment, the mutant Mb2Cas12a polypeptide comprises a domain swap occurring in a domain selected from the group consisting of WED-MR1, WED-MR2, BH, Up-seq, and Nuc. In another embodiment, the WED-MR1 domain is replaced with a sequence selected from the group consisting of SEQ ID NOs: 24, 29, and 30. In another embodiment, the WED-MR2 domain is replaced with a sequence selected from the group consisting of SEQ ID NOs: 25, 32, and 33. In yet another embodiment, the BH domain is replaced with a sequence selected from the group consisting of SEQ ID NOs: 40, 41, and 42. In another embodiment, the Up-seq domain is replaced with a sequence selected from the group consisting of SEQ ID NOs: 59, 60, 61, 62, and 63. In yet another embodiment, the Nuc domain is replaced with a sequence selected from the group consisting of SEQ ID NOs: 65, 66, and 67. In a further embodiment, the polypeptide comprises a sequence selected from the group comprising SEQ ID NOs: 37, 59, 60, 61, 62, 63, 74, 75, 77 and 78.

[0218] Another aspect of the invention is a mutant Mb2Cas12a polypeptide comprising at least one amino acid substitution and at least one domain swap. In one embodiment, the mutant Mb2Cas12a polypeptide comprises a sequence selected from the group consisting of SEQ ID NOs: 36, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, and 107.

[0219] Another aspect of the invention is a method of editing a plant genome, comprising contacting the plant genome with a mutant Mb2Cas12a polypeptide of the preceding embodiments. In embodiments, the method of editing a plant genome further comprises a guide RNA. In another embodiment, the guide RNA is encoded by a sequence comprising any of SEQ ID NOS: 18-21.

[0220] Another aspect of the invention is an edited plant obtained by the method of the foregoing embodiment.

[0221] Another aspect of the invention is a construct or plasmid comprising a polynucleotide sequence encoding the mutant Mb2Cas12a polypeptide of the foregoing embodiments. One aspect is a non-human cell comprising a construct or plasmid encoding the mutant Mb2Cas12a polypeptide. [Example]

[0222] We achieved improved Mb2Cas12a enzymatic activity through protein engineering. Our approach involved rational design and functional domain swapping based on available protein structures of Cas12a orthologs under different conditions, including the MbCas12a-22581 ortholog from the same Mb2Cas12a species (Table 1). The identity of both orthologs is 94.7% in the primary amino acid sequence. Therefore, MbCas12a-22581 serves as a good reference for studying the functionality of Mb2Cas12a, since the latter does not have a reference structure.

[0223] [Table 1]

[0224] "PDB" refers to the Protein Data Bank, a public database for archiving information about protein structures. See www.rcsb.org.

[0225] In the following examples, all constructs containing engineered or wild-type Mb2Cas12a coding sequences for maize transformation have the same expression cassette and enzyme configuration (Figure 1). Upstream of the coding sequence is a sugarcane ubiquitin promoter ("prSoUbi4", SEQ ID NO: 7). The Mb2Cas12a coding sequence (the sequence depends on the variant used) is fused at the N-terminus to an SV40 nuclear localization signal ("NLS") (SEQ ID NO: 8) via a flexible peptide linker (30 amino acids (GGGGS), SEQ ID NO: 9). The same peptide linker is fused at the C-terminus to two SV40 NLSs separated from each other by short linkers (8 amino acids (SGGS), SEQ ID NO: 10). This coding sequence was codon-optimized and linked to the Agrobacterium tumefaciens nopaline synthase gene terminator ("tNOS", SEQ ID NO: 11). The crRNA array, containing four crRNAs, was flanked by the ribozyme HH (SEQ ID NO: 12) 5-prime and the HDV (SEQ ID NO: 13) 3-prime of the array and controlled by another set of regulatory elements identical to those controlling the expression of the selected Mb2Ca12a sequence. This crRNA array expresses four crRNAs targeting four different maize genes: Waxy1 (ZmWx1), A UDP-glucosyltransferase, benzoxazinoid 9 (ZmBX9), Glossy2 (ZmGL2), and BCL2-associated X (ZmBINa). The natural Mb2Ca12a maturation-directing repeat (DR) was used as the crRNA scaffold for design. The crRNA array, including the ribozyme, is represented by SEQ ID NO: 14.

[0226] The constructs were stably transformed into maize immature embryos using standard transformation protocols (Zhong et al., 2018). Leaf sheath tissue from regenerated plantlets was sampled for DNA extraction, and transgenic plants were identified using TaqMan qPCR assays. Sequencing confirmation of each of the four target sites and analysis by TaqMan qPCR assays were used to determine the editing efficiency of each target site.

[0227] 1. The D172 variant of Mb2Cas12a improves editing efficiency. To determine whether the Mb2Cas12a D172R or D172K mutations improve enzyme activity in maize, two constructs were constructed to express each variant and then transformed into maize for event analysis.

[0228] Table 2 summarizes the SDN1 efficiency at the four target sites. Compared to the control (wild-type Mb2Cas12 (SEQ ID NO: 1)), the D172R variant (SEQ ID NO: 3) significantly improved SDN1 efficiency with only a slight decrease, except for the first crRNA in the array. The D172K variant (SEQ ID NO: 5) significantly improved SDN1 efficiency at all target sites compared to the control, but the increase in the last three crRNAs was less than that of the D172R variant.

[0229] [Table 2]

[0230] Efficiency was measured as the percentage of plants with indel mutations divided by the total number of transgenic plants.

[0231] 2. Mutational variants in the WED-III domain improve Mb2Cas12a editing efficiency. Cas12a possesses an inherent gRNA self-processing ability by processing the 5-prime end of each crRNA from the crRNA array to release individual functional crRNAs. To improve the gRNA processing ability of Mb2Cas12a, we identified a hydrophilic amino acid residue, glutamic acid, conserved throughout Cas12a from the Moraxella bovoculi species at position 797 (E797) of Mb2Cas12a based on alignment analysis. In other Cas12a orthologs, such as Lb, As, and Fn, the corresponding positions are hydrophobic amino acids, namely, L807 in AsCas12a, A766 in LbCas12a, and A850 in FnCas12a. We hypothesized that the hydrophilic side chain of E797 may affect the local spatial structure due to interactions with surrounding atoms. Therefore, we evaluated the D172R-mediated E797A mutation in Mb2Cas12a in stable transformation. Meanwhile, we identified one microregion (MR1) with a lower acidic / basic amino acid content in Moraxella bovoculi (designated WED-MR1, SEQ ID NO: 22) compared with Lb, As, and FnCas12a (Table 3). We found another microregion (MR2) with high diversity across Cas12a orthologs (Table 4) (designated WED-MR2, SEQ ID NO: 23). We performed a microregion swap with the D172R version of Mb2Cas12a by replacing WED-MR1 with its homolog LKKEELVV (SEQ ID NO: 24) from LbCas12a. We further replaced the wild-type version of Mb2Cas12a in WED-MR2 with the LbCas12a homolog TTTLS (SEQ ID NO: 25).

[0232] [Table 3]

[0233] [Table 4]

[0234] As summarized in Table 5, SDN1 efficiency at four target sites is compared with wild-type Mb2Cas12a (SEQ ID NO: 1). An Mb2Cas12a variant with a single amino acid mutation, E797A (SEQ ID NO: 26), significantly improved editing efficiency at four target sites. Meanwhile, this mutation also showed a similar contribution in the D172R background (i.e., D172R and E797A, SEQ ID NO: 35), further confirming the importance of this amino acid for the activity of the Mb2Cas12a enzyme. For construct 26841, encoding an Mb2Cas12a variant containing D172R and WED-MR1 (i.e., SEQ ID NO: 36) and swapping it with the LbCas12a homolog, the efficiency of the first two crRNAs was enhanced, but a slight decrease was observed for the last crRNA. Compared to 26411, 27493, which has a WED-MR2-Lb (i.e., SEQ ID NO: 37) swap with a homolog of LbCas12a, improved efficiency in the first two gRNAs.

[0235] [Table 5]

[0236] Efficiency was measured as the percentage of plants with indel mutations divided by the total number of transgenic plants.

[0237] 3. BH domain swap variants or BH domain extensions significantly improve Mb2Cas12a activity. The bridge helix ("BH") domain, the central helix of the protein that structurally connects the REC and Nuc lobes, influences cleavage accuracy and trimming activity, enhances Cas12a specificity, and allows the apoenzyme to adopt a closed state, thereby facilitating efficient crRNA loading. Disrupting the α-helical nature of the BH in Cas12a alters trimming activity and cleavage rate (Worle et al., 2021). We performed an alignment of the BH domain across Cas12a orthologs (Table 6). The alignment showed that the length of the α-helix in the BH domain in MbCas12a-22581 is shorter than that of other orthologs, which may therefore affect enzymatic activity.

[0238] [Table 6]

[0239] Based on this analysis, we generated a series of Mb2Cas12a variants with domain swaps of microregions from other orthologs with high enzymatic activity, as listed in Table 8. We also generated two variants by rational design, altering amino acids to extend the alpha-helix length of the BH domain (L917V+N918K and I914K+L917V+N918A in Table 8). Meanwhile, we identified a highly diverse region ("Up-seq") located upstream of the BH domain (Table 7). We also identified the hydrophobic amino acid Val921, which has distinct properties compared to the hydrophilic amino acids in AsCas12a (Gln956), LbCas12a (Gln888), or FnCas12a (Lys969) at the same position. The main-chain carbonyl group of Gln956 in the bridge helix can form a hydrogen bond with the side chain of Lys468 in the REC2 domain of AsCas12a (Yamano et al., 2016). Meanwhile, the amide group of the side chain interacts with the sugar-phosphate backbone of the guide RNA via a hydrogen bond. The homologous Val931 in MbCas12a-22581 cannot form such an interaction because the distance between Val931 and the sugar-phosphate backbone exceeds 4 Å. Based on the above analysis and the high consistency between Mb2Cas12a and MbCas12a-22581, we created the single amino acid mutations listed in Table 8.

[0240] [Table 7]

[0241] The SDN1 efficiency at the four target sites is compared as summarized in Table 8. Compared to the control (26411), all of these variants significantly improved SDN1 efficiency at the four target sites. Construct 26445, which carries the BH2-A variant, showed the best performance for editing the four target sites in maize.

[0242] [Table 8]

[0243] 4. Nuc domain variants improve Mb2Cas12a enzymatic activity. The endonuclease active site of FnCas12a is located at the interface of the RuvC and Nuc domains and is responsible for sequential cleavage of target and non-target DNA strands (Swarts et al., 2017). The conserved polar residues Arg1226 and Asp1235, as well as the partially conserved Ser1228, form a cluster near the active site of the RuvC domain in AsCas12a, which is a residue important for endonuclease activity in the Nuc domain (Yamano et al., 2016). Peptide alignment of Cas12a orthologs shows that the surrounding residues in this region are highly diverse (Table 9).

[0244] [Table 9]

[0245] These critical residues are located in a small region, which we replaced with the homologous sequences of LbCas12a and AsCas12a, respectively (Table 10). We also evaluated the point mutation Y1172N in this region (Table 10). Meanwhile, considering that the entire Nuc domain barely interacts with other domains, and following our structural analysis, we also replaced the entire Nuc domain with the homologous sequences of LbCas12a and AsCas12a, respectively (Table 10).

[0246] [Table 10]

[0247] As shown by the editing efficiencies summarized in Table 10, three variants, namely variant Nuc1-As (in construct 26441), which contains the cognate Nuc domain of AsCas12a, variant Y1172N (in construct 26438), which contains a single amino acid mutation, and variant Nuc2-Lb (in construct 26553), which is a replacement of the entire Nuc domain from LbCas12a, significantly improved editing efficiency.

[0248] 5. The Mb2Cas12a F357W variant improves editing efficiency. In the Cas12a system, even though our crRNA design uses a spacer length of 23 nucleotides, the interaction between the target DNA sequence and crRNA forms only 20 base pairs. This is partly because Trp382 interacts with the 20th base pair of DNA and crRNA, which leads to the disruption of base pairs from position 21 onwards. Therefore, this amino acid residue is important for forming the triplex structure of Cas12a-DNA-crRNA.

[0249] In AsCas12a, Trp382 forms a stacking interaction with the C20:dG20 pairing in the heteroduplex format (i.e., "dG20" refers to the base at the 20th position on the DNA target strand, and "C20" refers to the base at the 20th position in the crRNA space), thus preventing base pairing between A21 and dT21. Furthermore, the W382A mutation in AsCas12a reduces enzymatic activity (Yamano et al., 2016), implying that the W382 residue contributes to enzymatic performance. Sequence alignment suggests that there is a functionally conserved aromatic amino acid at this position across most Cas12a orthologs (Table 11).

[0250] [Table 11]

[0251] While further analysis of the side chain revealed that the homolog Phe357 in Mb2Cas12a differs from the conserved Trp or Tyr residues in that it possesses a hydrogen donor or acceptor atom. Because phenylalanine does not possess a hydrogen donor or acceptor atom in its side chain, this suggests that there is little or no interaction between Phe357 in Mb2Cas12a and the RNA-DNA heteroduplex, further implying that Phe357 may reduce enzymatic activity. To test this hypothesis, we constructed a single-residue mutation, F357W, to assess its functionality in maize.

[0252] SDN1 efficiency at the four target sites was compared, as summarized in Table 12. Compared to the control (26411), the F357W variant (26446) significantly improved SDN1 efficiency at the four target sites.

[0253] [Table 12]

[0254] 6. Enhancing the interaction of Mb2Cas12a with the 5'-directed repeats of crRNA improves SDN1 efficiency. To cleave the target DNA sequence, the Cas12a protein binds to the crRNA molecule in the form of a ribonucleoprotein ("RNP") complex. A stronger Cas12a-crRNA interaction, i.e., a more stable Cas12a-crRNA RNP complex, can likely improve DNA binding efficiency and crRNA-dependent DNase activity. Studies of various Cas12a-crRNA complexes suggest that Cas12a binds to a pseudoknot structure formed by the 5'-directed repeat ("DR") in the mature crRNA. We compared the Cas12a-crRNA binary complex structure of LbCas12a with that of MbCas12a-22581. We identified four amino acid residues that behave differently in LbCas12a and MbCas12a-22581 with respect to their interaction with the DR (see Table 13).

[0255] [Table 13]

[0256] While Y516 and S713 of LbCas12a interact with the ribonucleic acid backbone of the DR, the corresponding residues in MbCas12a-22581 (F557 and A752, respectively) do not interact with the DR due to side chain properties that prevent hydrogen bonds from forming. Both Q906 of LbCas12a and its counterpart in MbCas12a-22581 (H949) can potentially interact with nucleobases in the DR. However, they tend to interact with different nucleobases, and the interaction between H949 of MbCas12a-22581 and U7 of the DR (i.e., uracil at position 7 of the directional repeat) likely interferes with base pairing in the pseudoknot structure between U7 and A17 (i.e., adenine at position 17 of the directional repeat). F789 of LbCas12a does not interact with DR, but its counterpart in MbCas12a-22581 (Y829) has a strong potential to form hydrogen bonds with DR due to an additional hydroxyl group in its side chain.

[0257] Based on these analyses, we created a series of single-mutation Mb2Cas12a variants by substituting each of the four residues in Mb2Cas12a with the corresponding residues in LbCas12a (Table 14). As in the previous example, a nuclear localization signal and a polypeptide linker were added to each variant, and binary vectors were constructed to generate transgenic maize plants and target the four maize genes. With the exception of the H939Q variant of 27501, the other variants significantly improved editing efficiency at the four target sites.

[0258] [Table 14]

[0259] 7. Combinations of single-region changes improve SDN1 efficiency at four target sites in maize. To further evaluate the synergistic effect of mutations or small peptide substitutions validated in previous experiments to improve enzyme activity, a series of constructs combining different mutations were generated as shown in Table 15. Compared to the control (26411), all tested combination variants improved efficiency at all target sites, with the exception of one variant, 27381, which showed a slight decrease at the first target site.

[0260] [Table 15]

[0261] 8. Combinations of single region changes improve SDN1 efficiency in soybean. One of the combination variants was further evaluated in soybean for its effect on efficiency. Upstream of the Mb2Cas12a coding sequence is the Arabidopsis thaliana EF-1 alpha A1 gene promoter ("prAtEF1aA1", SEQ ID NO:129) linked to its 5 prime end to the Figwort mosaic virus (FMV) enhancer ("eFMV", SEQ ID NO:130), and downstream of this coding sequence is the Agrobacterium tumefaciens nopaline synthase gene terminator ("tNOS", SEQ ID NO:11). The crRNA (encoded by SEQ ID NO:132) was controlled by the soybean ubiquitin 1 promoter ("prGmUbi1", SEQ ID NO:131) and the Agrobacterium tumefaciens nopaline synthase gene terminator ("tNOS", SEQ ID NO:11). This crRNA targets Δ12-fatty acid desaturase II (GmFAD2). The natural Mb2Cas12a maturation-directing repeat (DR) was employed as the crRNA backbone for design.

[0262] The construct was stably transformed into imbibed mature seeds using standard transformation protocols (Liang, D. et al. (2023) CRISPR / LbCas12a-Mediated Genome Editing in Soybean. In: Yang, B., Harwood, W., Que, Q. (eds) Plant Genome Engineering. Methods in Molecular Biology, vol. 2653. Humana, New York, NY. doi.org / 10.1007 / 978-1-0716-3131-7_3). Leaf tissue from regenerated plantlets was sampled for DNA extraction, and transgenic plants were identified by TaqMan qPCR assay. Sequencing confirmation of each of the four target sites was used to determine the editing efficiency of each target site.

[0263] Compared to the control (wild-type Mb2Cas12, SEQ ID NO: 1), the combined variant (SEQ ID NO: 94) significantly improved SDN1 efficiency. This data indicated that the optimization in maize is applicable to soybean.

[0264] [Table 16]

[0265] Example 8. Construct annotation.

[0266] [Table 17-1]

[0267] [Table 17-2]

[0268] [Table 18-1]

[0269] Table 18-2

[0270] Table 19-1

[0271] Table 19-2

[0272] Table 20-1

[0273] Table 20-2

[0274] Table 21-1

[0275] Table 21-2

[0276] Table 22-1

[0277] Table 22-2

[0278] Table 23-1

[0279] Table 23-2

[0280] Table 24-1

[0281] Table 24-2

[0282] Table 25-1

[0283] Table 25-2

[0284] Table 26-1

[0285] Table 26-2

[0286] Table 27-1

[0287] Table 27-2

[0288] Table 28-1

[0289] Table 28-2

[0290] Table 29-1

[0291] Table 29-2

[0292] Table 30-1

[0293] Table 30-2

[0294] Table 31-1

[0295] Table 31-2

[0296] Table 32-1

[0297] Table 32-2

[0298] Table 33-1

[0299] Table 33-2

[0300] Table 34-1

[0301] Table 34-2

[0302] Table 35-1

[0303] Table 35-2

[0304] Table 36-1

[0305] Table 36-2

[0306] Table 37-1

[0307] Table 37-2

[0308] Table 38-1

[0309] Table 38-2

[0310] Table 39-1

[0311] Table 39-2

[0312] Table 40-1

[0313] Table 40-2

[0314] Table 41-1

[0315] Table 41-2

[0316] Table 42-1

[0317] Table 42-2

[0318] Table 43-1

[0319] Table 43-2

[0320] Table 44-1

[0321] Table 44-2

[0322] Table 45-1

[0323] Table 45-2

[0324] Table 46-1

[0325] Table 46-2

[0326] Table 47-1

[0327] Table 47-2

[0328] Table 48-1

[0329] Table 48-2

[0330] Table 49-1

[0331] Table 49-2

[0332] Table 50-1

[0333] Table 50-2

[0334] Table 51-1

[0335] Table 51-2

[0336] Table 52-1

[0337] Table 52-2

[0338] Table 53-1

[0339] Table 53-2

[0340] Table 54-1

[0341] Table 54-2

[0342] Table 55-1

[0343] Table 55-2

[0344] Table 56-1

[0345] Table 56-2

[0346] Table 57-1

[0347] Table 57-2

[0348] Table 58-1

[0349] Table 58-2

[0350] Table 59-1

[0351] Table 59-2

[0352] Table 60-1

[0353] Table 60-2

[0354] Table 61-1

[0355] Table 61-2

Claims

1. A mutant Mb2Cas12a polypeptide comprising at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO:

1.

2. 2. The mutant Mb2Cas12a polypeptide of claim 1, wherein the at least one amino acid substitution occurs at a position selected from the group consisting of D172, F357, F547, A742, E797, Y819, E913, I914, L917, N918, V921, H939, and Y1172.

3. 3. The mutant Mb2Cas12a polypeptide of claim 2, wherein the at least one amino acid substitution is selected from the group consisting of D172R, D172K, F357W, F547Y, A742S, E797A, Y819F, I914K, L917V, N918A, N918K, V921K, V921Q, H939Q, and Y1172N.

4. The mutant Mb2Cas12a polypeptide of claim 3, comprising a sequence selected from the group consisting of SEQ ID NOs: 3, 5, 26, 35, 57, 58, 64, 76, 80, 85, 86, 87, 88 and 127.

5. A mutant Mb2Casl2a polypeptide comprising at least one domain swap.

6. 6. The mutant Mb2Cas12a polypeptide of claim 5, wherein the domain swap occurs in a domain selected from the group consisting of WED-MR1, WED-MR2, BH, Up-seq, and Nuc.

7. 7. The mutant Mb2Cas12a polypeptide of claim 6, wherein the WED-MR1 domain is replaced with a sequence selected from the group consisting of SEQ ID NOs: 24, 29, and 30.

8. 7. The mutant Mb2Cas12a polypeptide of claim 6, wherein the WED-MR2 domain is replaced with a sequence selected from the group consisting of SEQ ID NOs: 25, 32, and 33.

9. 7. The mutant Mb2Cas12a polypeptide of claim 6, wherein the BH domain is substituted with a sequence selected from the group consisting of SEQ ID NOs: 40, 41, and 42.

10. 7. The mutant Mb2Cas12a polypeptide of claim 6, wherein the Up-seq domain is replaced with a sequence selected from the group consisting of SEQ ID NOs: 59, 60, 61, 62, and 63.

11. 7. The mutant Mb2Cas12a polypeptide of claim 6, wherein the Nuc domain is replaced with a sequence selected from the group consisting of SEQ ID NOs: 65, 66, and 67.

12. The mutant Mb2Cas12a polypeptide of claim 5, comprising a sequence selected from the group comprising SEQ ID NOs: 37, 59, 60, 61, 62, 63, 74, 75, 77 and 78.

13. A mutant Mb2Casl2a polypeptide comprising at least one amino acid substitution and at least one domain swap.

14. 14. The mutant Mb2Cas12a polypeptide of claim 13, wherein the Mb2Cas12a polypeptide sequence is selected from the group consisting of SEQ ID NOs: 36, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106 and 107.

15. 15. A method for editing a plant genome, comprising contacting the plant genome with a mutant Mb2Cas12a polypeptide of claims 1-14.

16. 16. The method of claim 15, further comprising a guide RNA.

17. 17. The method of claim 16, wherein the guide RNA is encoded by a sequence comprising SEQ ID NOs: 18-21.

18. An edited plant obtained by the method of claims 15 to 17.

19. A construct or plasmid comprising a polynucleotide sequence encoding a mutant Mb2Cas12a polypeptide according to claims 1 to 14.

20. 20. A non-human cell comprising the construct or plasmid of claim 19.